A 3D road object detection method based on improved point cloud sparse convolution

By adopting the improved focal sparse convolution module and the large receptive field convolution module in three-dimensional road target detection, the problems of computational redundancy and spatial structure failure are solved, and higher detection accuracy and robustness are achieved.

CN119559588BActive Publication Date: 2025-05-06NANJING TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411714503.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-06
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The prior art has problems of computational redundancy and spatial structure damage in three-dimensional road target detection, especially when dealing with sparse point clouds, it is difficult to maintain a good spatial structure and large receptive fields.

Method used

The improved focus sparse convolution module and the large receptive field convolution module are adopted to obtain the position of the scalable convolution by constructing the offset, avoid computing redundancy, and achieve large receptive field feature extraction through the position weight attention module and submanifold sparse convolution.

Benefits of technology

The detection accuracy and robustness of the object detection model in complex road conditions is improved, computational redundancy and spatial structure damage are avoided, and higher accuracy and processing speed are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559588B_ABST
    Figure CN119559588B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional road target detection method based on improved point cloud sparse convolution, which obtains a point cloud data set, divides and groups the point cloud, extracts features from the grouped voxels, and obtains voxel feature tensors and voxel position tensors; constructs a three-dimensional road target detection model based on point cloud sparse convolution, the three-dimensional road target detection model uses a submanifold sparse convolution layer as an input layer, and the backbone network of the three-dimensional road target detection model involves an improved focal sparse convolution module and a large receptive field convolution module; after the voxel feature encoding is sent to the point cloud sparse convolution backbone network, the output features are sent to a regional generation module RPN to obtain semantic features and a preliminary target bounding box; and then sent to a target detection network for final prediction. The large receptive field convolution layer and the improved focal sparse convolution layer of the present invention complement each other, and improve the detection accuracy and robustness in complex road conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to target detection and deep learning technology, and in particular to a three-dimensional road target detection method based on improved point cloud sparse convolution. Background Art

[0002] With the continuous development of urban transportation and the improvement of its intelligent level, vehicle detection technology, as one of the basic tasks in the fields of autonomous driving, traffic monitoring, and intelligent traffic management, plays an increasingly important role. Faced with complex traffic environments, LiDAR is gradually being used in the field of target detection.

[0003] Image data is easily limited in dealing with vehicle occlusion, lighting changes, and weather effects, and point cloud data can make up for these shortcomings. Therefore, it is necessary to design a feature extraction network based on radar point clouds, and academia and industry have begun to explore feature extraction methods for point clouds. However, due to the unstructured and sparse three-dimensional geometric data characteristics of point clouds, it is impossible to directly migrate the 2D feature extraction paradigm designed for structured data to 3D tasks.

[0004] Existing methods mainly voxelize point clouds and use sparse convolution for feature extraction. Voxelized point clouds can reduce the huge amount of data brought by direct processing of point clouds and give certain structural features, but the empty voxels generated at the same time will bring computational redundancy. Sparse convolution uses hash tables to establish connections between features, feature position encoding, convolution layers, and results, so that the network can calculate at a certain location by controlling the input, thereby achieving the purpose of skipping some voxels. This method has been widely used in autonomous driving point cloud target detection tasks and has inspired a series of improvements.

[0005] The purpose of performing extended convolution on valued features is to seek empty-value features that have a good effect on the spatial structure of features during the convolution process. Sub-manifold sparse convolution always performs convolution calculations with the valued point cloud as the center. Since for point cloud voxels, the convolution always covers 3*3*3 voxel features, in the convolution centered on the empty-value feature, multiple valued features may still be covered. Sub-manifold sparse convolution ignores that some empty voxels can play a role in connecting adjacent non-empty features, destroying the good spatial structure of the voxels. On the contrary, regular sparse convolution performs extended convolution within a 3*3*3 range centered on all valued features. This inevitably adds a lot of computational burden to the model. In addition, the purpose of feature extraction is to deepen important features, and this method will blur the transition between important features and background, so it is usually only used in the first few layers of feature extraction.

[0006] These limitations on feature extraction capabilities stem from treating all features equally. However, for 3D sparse features with different sparsity and importance in space, using a unified processing method to handle non-uniform data is not optimal. To address this problem, it is necessary to focus on the expansion of important features by modifying the input feature sampling so that the kernel shape can adapt to the effective receptive field of the network.

[0007] The latest methods, such as focal coefficient convolution, use self-learning to determine where to perform extended convolution, that is, using the characteristics of convolution training fitting to let it autonomously determine the foreground features and the locations where the foreground features need to be expanded. This method can obtain a good spatial structure while avoiding the large computational burden brought by expansion at all locations. However, convolution is the choice when existing interpretable mathematical methods cannot meet the requirements. Relying solely on a single layer of convolution cannot fit well and is prone to misjudgment. The nonlinear characteristics of convolution also cause a large amount of calculation.

[0008] At the same time, large convolution kernels have been successfully applied in 3D sparse convolution. There is a consensus that a large receptive field contributes positively to many downstream visual tasks, but if the convolution kernel is simply enlarged, the time and space consumption will grow cubically with the increase of the kernel. Existing methods generally pre-fuse features in n×n blocks, which can approximately simulate expanding the size of the convolution kernel by N times without increasing the amount of computation. However, this usually sacrifices the model's understanding of details. Summary of the invention

[0009] Purpose of the invention: The purpose of the present invention is to solve the deficiencies in the prior art and provide a three-dimensional road target detection method with improved point cloud sparse convolution, which involves a large receptive field convolution layer and an improved focal sparse convolution layer. The improved focal sparse convolution layer obtains the positions of all expandable convolutions by constructing an offset through the positional relationship between the null-value point cloud and the valued point cloud that need to be expanded and convolved, thereby avoiding the destruction of the spatial structure caused by convolution only on non-null features and the computational redundancy caused by expanded convolution on all null-value features; the large receptive field convolution concentrates features within a certain range to achieve the purpose of expanding the receptive field without increasing the convolution kernel. The two modules of the present invention complement each other in feature extraction, improving the detection accuracy and robustness in complex road conditions.

[0010] Technical solution: A three-dimensional road target detection method based on improved point cloud sparse convolution of the present invention comprises the following steps:

[0011] Step 1: Construct a data set consisting of point clouds and perform preprocessing operations. The point clouds are cropped according to the three-dimensional spatial layout and the point clouds outside the specified range are removed. The data set here can use the nuScenes three-dimensional target detection data set. The targets of the three-dimensional road target scene are cars, pedestrians and cyclists. The ground truth values ​​are position, size and yaw angle.

[0012] Step 2: point cloud division and grouping, that is, dividing the cropped 3D point cloud space into 3D voxels of uniform size; for 3D voxels with more than N points, randomly sample N points from them, and for 3D voxels with less than N points, fill them with 0;

[0013] Step 3: Use the VFE module to extract features from each grouped voxel to obtain a new voxel feature code, which includes a voxel feature tensor and a voxel position tensor;

[0014] Step 4: construct a 3D road target detection model based on point cloud sparse convolution. The 3D road target detection model uses a submanifold sparse convolution layer as an input layer. The backbone network of the 3D road target detection model involves an improved focal sparse convolution module and a large receptive field convolution module. The backbone network has a total of five layers. The first layer is an improved focal sparse convolution module. Each of the next four layers includes a submanifold sparse convolution module and a large receptive field convolution module, and a downsampling convolution layer is set between two adjacent layers from the second to the fifth layers.

[0015] After the voxel feature tensor obtained in step S2 is fed into the point cloud sparse convolution backbone network, it first passes through the first layer of improved focal sparse convolution module to obtain fine-grained features, and then feeds them into the sub-manifold sparse convolution module and the large receptive field convolution module of the second layer respectively. The sub-manifold sparse convolution module and the large receptive field convolution module extract the features respectively, and then the two features are added and fed into the first downsampling convolution layer until the feature extraction of the fifth layer is completed;

[0016] Step 5: Send the features output by the 3D road target detection model in step 4 to the region generation module RPN, perform downsampling, upsampling and channel connection operations, and extract deeper semantic features and target bounding boxes;

[0017] Step 6: Send the semantic features and target bounding box obtained in step S5 to the target detection network Centerhead for final prediction, and the prediction is classified into the detected target type (ten categories) and the three-dimensional bounding box of the target.

[0018] Furthermore, in the three-dimensional road object detection model, the specific operation after the voxel feature tensor enters the improved focal sparse convolution module of the first layer is as follows:

[0019] Step 1) construct two 0-valued two-dimensional tensors with the number of voxels in the x and y dimensions as length and width, fill the height and 1 values ​​respectively based on the voxel position tensor, and obtain the height-filled tensor and the valued mask;

[0020] Step 2), translate the valued mask in the 3*3 plane according to 16 possible situations (the translation operation here is only to determine whether the valued features overlap in this translation process, that is, to determine whether a valued feature position has a value at the specified offset position), and expand the mask before and after the translation to zero value;

[0021] Step 3), add the expanded valued masks before and after the translation obtained in step 2), take the position codes of all values ​​2 in the added tensor, remove the effects of translation and expansion, and obtain the valued feature position tensor that generates the extended features before translation and the corresponding valued feature position tensor after translation;

[0022] Step 4) Obtain the extended feature offset for each of the 16 possible translation situations (judge the position offset of the null-value feature that needs to be extended convolution by the extended feature offset), add it to the valued feature position tensor, and obtain the extended feature plane position tensor;

[0023] Step 5), use the two valued feature position tensors obtained in step 3) to fill the height tensor with the height tensor to obtain the height tensor, and obtain the height tensor before translation and the height tensor after translation. Subtract the two and divide by two to obtain the remainder, and then determine whether the subtraction value is within the range of (-2, 2) to obtain the extended feature height mask;

[0024] Step 6), concatenate the height tensor before translation with the extended feature plane position tensor, and filter with a mask to obtain the extended feature position tensor;

[0025] Step 7), create a new zero-value feature tensor corresponding to the extended feature position tensor, concatenate the extended feature position tensor with the voxel position tensor, and concatenate the zero-value feature tensor with the voxel feature tensor; and re-encode to obtain the voxel feature code;

[0026] Step 8), the voxel feature encoding obtained in step 7) is put into two layers of sub-manifold sparse convolution to extract the features.

[0027] Furthermore, in the backbone network of the three-dimensional road object detection model, the large receptive field sparse convolution module performs the following operations on the input features:

[0028] First, a position weighted attention module is constructed. The voxel feature tensor and the voxel position tensor are input into the position weighted attention module together to obtain voxel features with position weights. , the specific formula of the weighted attention module is as follows:

[0029] ;

[0030] Where sigmod represents the sigmod activation function, It represents a submanifold sparse convolution with n input channels and 1 output channel; represents the input voxel feature tensor, A tensor representing the input voxel positions;

[0031] Then, the resulting voxel position tensor In three-dimensional space, the quotient is compressed by multiples of 7, and the positions with the same quotient value correspond to Add the eigenvalues ​​of to obtain the compressed voxel feature tensor;

[0032] Next, a submanifold sparse convolution is performed on the compressed voxel feature tensor. Since the voxels are compressed by a multiple of 7, the convolution kernel of this convolution can be regarded as enlarged seven times;

[0033] Secondly, the compressed voxel features after the submanifold sparse convolution are refilled in the compressed encoding order and divided by the position weight to obtain a new voxel feature tensor;

[0034] Finally, the new voxel feature tensor is fed into two layers of submanifold sparse convolution to obtain a large receptive field tensor.

[0035] Furthermore, in the backbone network of the three-dimensional road object detection model, each submanifold sparse convolution module includes two layers of submanifold sparse convolution and normalization layers, and outputs spatial features by performing two layers of submanifold sparse convolution and normalization on the input features.

[0036] Furthermore, in the backbone network of the three-dimensional road object detection model, the number of input channels is 16, and the number of channels of the three downsampling convolutional layers are 32, 64 and 128 respectively.

[0037] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0038] 1. The three-dimensional road target detection model of the present invention provides an improved focal sparse convolution module, which uses the position of important features to determine and expand the operation, avoiding the redundancy caused by the regular sparse convolution expansion calculation at all positions and the sub-manifold sparse convolution without expansion operation to destroy the good structure of the features, thereby improving the accuracy and robustness of the target detection model when processing sparse point clouds. Different from simply using convolution to determine the expansion position, our method is more accurate and more interpretable.

[0039] 2. The present invention introduces a large receptive field convolution module and designs a convolution method that increases the convolution kernel but only slightly increases the amount of calculation, and theoretically the convolution kernel can be infinitely expanded.

[0040] 3. The present invention combines two convolutional modules to construct a sparse point cloud backbone network that can maintain a good spatial structure and have a large receptive field. It can be plug-and-play in point cloud-based object detection tasks. Experimental verification shows that the present invention is superior to the baseline network and some recent inventions in terms of accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is the overall process of the present invention.

[0042] Figure 2 This is the backbone network process of the present invention.

[0043] Figure 3 To provide a visual explanation of the improved focal sparse convolution feature expansion, triangles and circles are valued features, and plus signs are possible locations of null-valued features that have expansion value in the current case.

[0044] Figure 4 To provide a visual explanation of step 2 of the improved focal sparse convolution, both triangles and circles are valued features. When features overlap after translation, there are null-valued features with extended value.

[0045] Table 1 is a comparison of the accuracy of the overall network of the present invention and the baseline network on the validation set.

[0046] Table 2 is the ablation comparison experiment of the improved focal coefficient convolution of the network of the present invention and the sparse convolution network. DETAILED DESCRIPTION

[0047] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0048] like Figure 1 As shown, a three-dimensional road target detection method based on improved point cloud sparse convolution of the present invention comprises the following steps:

[0049] Step 1: Construct a data set consisting of point clouds and perform preprocessing operations to crop the point clouds according to the three-dimensional space layout;

[0050] Step 2: point cloud division and grouping, that is, dividing the cropped 3D point cloud space into 3D voxels of uniform size; for 3D voxels with more than N points, randomly sample N points from them, and for 3D voxels with less than N points, fill them with 0;

[0051] Step 3: Use the VFE module to extract features from each grouped voxel to obtain a new voxel feature code, where the voxel feature code includes a voxel feature tensor and a voxel position tensor. In this embodiment, the voxel position tensor and the voxel feature tensor correspond to each other one by one and together constitute the voxel feature code. For example, for the feature (40, 1440, 1440, m), when it is used as a sparse feature code, it will become a voxel feature tensor (n, m) and a voxel position tensor (n, 3), where n is the number of non-empty voxels, m is the feature, and (n, 3) is the position of the feature in (40, 1440, 1440).

[0052] Step 4: Figure 2 As shown, a 3D road target detection model based on point cloud sparse convolution is constructed. The 3D road target detection model uses a submanifold sparse convolution layer as an input layer. The backbone network of the 3D road target detection model involves an improved focal sparse convolution module and a large receptive field convolution module. The backbone network has a total of five layers, the first layer is an improved focal sparse convolution module, and each of the next four layers includes a submanifold sparse convolution module and a large receptive field convolution module, and a downsampling convolution layer is set between two adjacent layers from the second to the fifth layers.

[0053] After the voxel feature tensor obtained in step S2 is fed into the point cloud sparse convolution backbone network, it first passes through the first layer of improved focal sparse convolution module to obtain fine-grained features, and then feeds them into the sub-manifold sparse convolution module and the large receptive field convolution module of the second layer respectively. The sub-manifold sparse convolution module and the large receptive field convolution module extract the features respectively, and then the two features are added and fed into the first downsampling convolution layer until the feature extraction of the fifth layer is completed;

[0054] Step 5: Send the features output by the 3D road target detection model in step 4 to the region generation module RPN for downsampling, upsampling and channel connection operations to extract deeper semantic features and preliminary target bounding boxes;

[0055] Step 6: Send the semantic features and preliminary target bounding box obtained in step S5 to the target detection network Centerhead for final prediction. The predicted classification includes the detected target type and the three-dimensional bounding box of the target.

[0056] The improved focal sparse convolution module of this embodiment includes a connection feature extraction module and a sub-manifold sparse convolution to obtain features with better spatial structure. Taking the plane as an example, the null-value feature that has a good effect on the spatial structure always has more than two valued features in the outermost layer of the 3*3 range, that is, for a valued feature, the corresponding other valued feature that may produce an extensible feature is always in the outermost layer of the 5*5 range, which can be divided into 5*5-3*3=16 situations where extensible features may be produced. By judging 16 possible situations and collecting the corresponding feature position tensors, the extended features can be obtained. For three-dimensional features, the height feature is compressed first, and then the position tensor of the feature in the plane is collected before judging whether the height is within the range of 3*3*3. In this way, the position tensor of all extended features can be obtained.

[0057] In the three-dimensional road object detection model of this embodiment, the specific operations after the voxel feature tensor enters the improved focal sparse convolution module of the first layer are as follows:

[0058] Step 1) construct two 0-valued two-dimensional tensors with the number of voxels in the x and y dimensions as length and width, fill the height and 1 values ​​respectively based on the voxel position tensor, and obtain the height-filled tensor and the valued mask;

[0059] Step 2), translate the valued mask in the 3*3 plane according to 16 possible situations, and expand the mask before and after the translation to zero value;

[0060] Step 3), add the expanded valued masks before and after the translation obtained in step 2), take the position codes of all values ​​2 in the added tensor, remove the effects of translation and expansion, and obtain the valued feature position tensor that generates the extended features before translation and the corresponding valued feature position tensor after translation;

[0061] Step 4), obtain the extended feature offset in each of the 16 possible translation situations, add it to the valued feature position tensor, and obtain the extended feature plane position tensor;

[0062] Step 5), use the two valued feature position tensors obtained in step 3) to fill the height tensor with the height tensor to obtain the height tensor, and obtain the height tensor before translation and the height tensor after translation. Subtract the two and divide by two to obtain the quotient to obtain the height difference tensor and determine whether the subtraction value is within the range of (-2, 2) to obtain the extended feature height mask;

[0063] Step 6), concatenate the height tensor before translation with the extended feature plane position tensor, and filter with a mask to obtain the extended feature position tensor;

[0064] Step 7), create a new zero-value feature tensor corresponding to the extended feature position tensor, concatenate the extended feature position tensor with the voxel position tensor, and concatenate the zero-value feature tensor with the voxel feature tensor; and re-encode to obtain the voxel feature code;

[0065] Step 8), the voxel feature encoding obtained in step 7) is put into two layers of sub-manifold sparse convolution to extract the features.

[0066] In the backbone network of the three-dimensional road object detection model of this embodiment, the large receptive field sparse convolution module performs the following operations on the input features:

[0067] First, a position weighted attention module is constructed. The voxel feature tensor and the voxel position tensor are input into the position weighted attention module together to obtain voxel features with position weights. , the specific formula of the weighted attention module is as follows:

[0068]

[0069] Where sigmod represents the sigmod activation function, It represents a submanifold sparse convolution with n input channels and 1 output channel; represents the input voxel feature tensor, A tensor representing the input voxel positions;

[0070] Then, the resulting voxel position tensor In three-dimensional space, the quotient is compressed by multiples of 7, and the positions with the same quotient value correspond to Add the eigenvalues ​​of to obtain the compressed voxel feature tensor;

[0071] Next, a submanifold sparse convolution is performed on the compressed voxel feature tensor;

[0072] Secondly, the compressed voxel features after the submanifold sparse convolution are refilled in the compressed encoding order and divided by the position weight to obtain a new voxel feature tensor;

[0073] Finally, the new voxel feature tensor is fed into two layers of submanifold sparse convolution to obtain a large receptive field tensor.

[0074] In the backbone network of the three-dimensional road target detection model of this embodiment, each submanifold sparse convolution module includes two layers of submanifold sparse convolution and normalization layers, and outputs spatial features by performing two layers of submanifold sparse convolution and normalization on the input features.

[0075] In the backbone network of the three-dimensional road object detection model of this embodiment, the number of input channels is 16, and the number of channels of the three downsampling convolutional layers are 32, 64 and 128 respectively.

[0076] To verify the technical effect of the present invention, this embodiment applies the method of the present invention to perform three-dimensional road target detection on the following scene.

[0077] The experimental dataset used in this experiment is nuScenes. The nuScenes dataset is a large-scale autonomous driving benchmark, which contains a total of 10,000 driving scenes, divided into 700, 150 and 150 scenes, respectively, for training, verification and testing. For detection, nuScenes defines a set of evaluation protocols, including nuScenes detection score (NDS), mean average precision (mAP), and five true (TP) indicators, namely mean translation error (mATE), mean scale error (mASE), mean orientation error (mAOE), mean velocity error (mAVE) and mean attribute error (mAAE). In this embodiment, mAP is the average of the average precision of ten classes with distance thresholds of 0.5m, 1m, 2m, and 4m. NDS is a weighted combination of mAP, mATE, mASE, mAOE, mAVE and mAAE. The specific steps of this embodiment are as follows.

[0078] Step 1: Preprocess the dataset and divide it into a training set and a validation set. Limit the sampling range of the point cloud to (-54, 54) on the X axis, (54, 54) on the Y axis, and (5, 3) on the Z axis. Perform data enhancement by random rotation and scaling operations before training.

[0079] Step 2: First, divide the space into stacked voxels of the same size, with the specific size of 0.75 on the X axis, 0.75 on the Y axis, and 0.2 on the Z axis. Then, load the 3D point cloud data into these voxels to achieve grouping, with a maximum number of voxels of 160,000. Randomly sample 10 samples of the point cloud in each voxel. Finally, extract the features of the sampled points in each voxel and encode them into vector representation to obtain voxel feature encoding.

[0080] Step 3: Send the voxel feature encoding into the point cloud sparse convolution backbone network. The input layer is a submanifold sparse convolution, and the first layer is an improved focal sparse convolution. The second, third, fourth, and fifth feature extraction layers are all a submanifold sparse convolution module and a large receptive field sparse convolution. The number of input channels is 16, and each layer is separated by a downsampling layer. The number of channels is (32, 64, 128). Any type of sparse convolution is followed by batch normalization and ReLU activation.

[0081] Step 3: Construct an improved focal sparse convolution, which includes the following steps:

[0082] Construct two 0-valued two-dimensional tensors with the number of voxels in the x and y dimensions as length and width, fill the height and 1 values ​​respectively based on the voxel position tensor, and obtain the height-filled tensor and the valued mask.

[0083] like Figure 3 and Figure 4 As shown, the null-value feature that has a good effect on the spatial structure always has more than two valued features in the outermost layer of the 3*3 range, that is, for one valued feature, the corresponding other valued feature that may produce an extensible feature is always in the outermost layer of the 5*5 range, which can be divided into 5*5-3*3=16 situations in which extensible features may be produced. By judging the 16 possible situations and collecting the corresponding feature position tensors, the extended features can be obtained. For three-dimensional features, the height feature is compressed first, and then the position tensor of the feature in the plane is collected before judging whether the height is within the range of 3*3*3.

[0084] Expand and translate the valued mask. For example, in the 3*3 range, the upper left corner and the lower right corner have values. The specific method is to pad the valued mask (1, 1440, 1440) in the positive direction of the x-axis and the positive direction of the y-axis to obtain the expanded original mask (1, 1442, 1442). Pad in the negative direction of the x-axis and the negative direction of the y-axis to obtain the expanded translation mask (1, 1442, 1442).

[0085] Add the expanded and translated mask to the expanded original mask, take the position codes of all values ​​2 in the added tensor, remove the effects of translation and expansion, and get the valued feature position tensor and the corresponding valued feature position tensor that will generate the extended feature. Obtain the extended feature offset for each of the 16 possible translation situations. This process always assumes that one of the valued features is in the corner. Depending on the position of the other feature, there may be 1 to 3 extended features. Taking the valued features located in the middle of the left and the lower right corner as an example, there are two extended features located in the middle and lower middle. Use the two valued feature position tensors to fill the height tensor to obtain the height tensor, subtract and divide by two to obtain the quotient, obtain the height difference tensor, and judge whether the subtraction value is in the range of (-2, 2) to obtain the extended feature height mask. Concatenate the first height tensor with the extended feature plane position tensor, and filter with the mask to obtain the extended feature position tensor. Create a new zero-value feature tensor corresponding to the extended feature position tensor, concatenate the extended feature position tensor with the voxel position tensor, concatenate the zero-value feature tensor with the voxel feature tensor, and construct a voxel feature code. Put the above voxel feature code into two layers of submanifold sparse convolution to extract important features.

[0086] Step 4: Construct a large receptive field convolution module to obtain global features. It consists of a position weight attention module and a compression tensor and submanifold sparse convolution. The voxel position tensor is divided by multiples of 7 in three-dimensional space. The positions with the same quotient value correspond to The eigenvalues ​​of are added to obtain the compressed voxel feature tensor. The compressed voxel feature tensor is subjected to a submanifold sparse convolution. Since the voxels are compressed by a multiple of 7, the convolution kernel can be regarded as enlarged seven times. The compressed voxel features after convolution are refilled in the compressed encoding order and divided by the position weight to obtain a new voxel feature tensor. The new voxel feature tensor is put into two layers of submanifold sparse convolution to obtain a large receptive field feature tensor.

[0087] Step 5: Send the voxel feature encoding to the point cloud sparse convolution backbone network. The input layer is a submanifold sparse convolution, and the first layer is an improved focal sparse convolution. The second, third, fourth, and fifth feature extraction layers are all a submanifold sparse convolution module and a large receptive field sparse convolution. The number of input channels is 16, and there is a downsampling layer between each of the last four layers. The number of channels is (32, 64, 128). Batch normalization and ReLU activation are performed after any type of sparse convolution.

[0088] In step 6, the features are fed into the RPN module, the number of channels is reduced from 256 to 128 layers by downsampling, the number of channels is increased from 128 to 256 layers by upsampling, and the 128 and 256 layers are connected by channels.

[0089] Step 7: Input the final feature map into the detection head for preliminary detection, and output the first-stage result, i.e., the position information (Cartesian coordinates) and bounding box information (size and yaw angle) of the key point (i.e., the target center point) in Cartesian coordinates. Extract point features from the three-dimensional center of each face of the predicted target bounding box. Specifically, the center of the bounding box, the top surface, and the bottom surface are projected to the same point in the bird's-eye view. For each point, use bilinear interpolation to extract features from the backbone network output. Use these point features to predict the target confidence score and bounding box information refinement that are independent of the category, and output the second-stage result.

[0090] Table 1 Comparison of accuracy with the baseline network on the validation set

[0091] Method model mAP NDS Car Ped CenterPoint L 58.0 65.5 84.6 83.4 CenterPoint + C+L 60.3 67.3 85.2 84.6 FocalsConv L 63.4 68.9 84.6 83.4 Link L 63.2 68.4 84.9 84.2 The present invention L 64.9 69.9 86.3 86.6

[0092] Table 2 Ablation comparison experiment of improved focal coefficient convolution and sparse convolution network

[0093] Method mAP NDS FPS FocalsConv 63.4 68.9 4.9 The present invention 64.2 69.5 6.7

[0094] This embodiment uses the nuscenes standard to verify the model of the present invention on the validation set. The results are shown in Table 1. The model of the present invention exceeds all baseline methods. At the same time, the improved focal sparse convolution module replaces the focal sparse convolution module in focalsconv to perform an ablation experiment. The experimental results are shown in Table 2. The experiment proves that our network is not only more accurate, but also has a faster processing speed.

[0095] In summary, the three-dimensional road target detection model in the present invention is a point cloud feature extraction backbone network with good spatial structure and large receptive field. The improved focal sparse convolution module in the three-dimensional road target detection model obtains the position of all expandable convolutions by constructing an offset through the positional relationship between the expanded convolution and the valued point cloud, and reduces the amount of calculation caused by the three-dimensional features by compressing the height first and then taking the value, thereby avoiding the huge amount of calculation and feature blurring caused by expanding all features. The large receptive field sparse convolution module in the three-dimensional road target detection model aggregates the feature values ​​within a certain range into one value through a layer of sub-manifold sparse convolution, performs a feature extraction, and then expands it. In this way, infinite kernel convolution can be theoretically achieved.

[0096] The present invention achieved a mAP of 64.9 on the validation set of the dataset, surpassing all baseline networks and some of the latest networks.

Claims

1. A three-dimensional road target detection method based on improved point cloud sparse convolution, characterized in that: The following steps are involved: Step 1: Construct a data set consisting of point clouds and perform preprocessing operations to crop the point clouds according to the three-dimensional space layout; Step 2: point cloud division and grouping, that is, dividing the cropped 3D point cloud space into 3D voxels of uniform size; for 3D voxels with more than N points, randomly sample N points from them, and for 3D voxels with less than N points, fill them with 0; Step 3: Use the VFE module to extract features from each grouped voxel to obtain a new voxel feature code, which includes a voxel feature tensor and a voxel position tensor; Step 4: construct a 3D road target detection model based on point cloud sparse convolution. The 3D road target detection model uses a submanifold sparse convolution layer as an input layer. The backbone network of the 3D road target detection model involves an improved focal sparse convolution module and a large receptive field convolution module. The backbone network has a total of five layers. The first layer is an improved focal sparse convolution module. Each of the next four layers includes a submanifold sparse convolution module and a large receptive field convolution module, and a downsampling convolution layer is set between two adjacent layers from the second to the fifth layers. After the voxel feature encoding obtained in step S2 is sent to the point cloud sparse convolution backbone network, it first passes through the first layer of improved focal sparse convolution module to obtain fine-grained features, and then is fed into the sub-manifold sparse convolution module and the large receptive field convolution module of the second layer respectively. The sub-manifold sparse convolution module and the large receptive field convolution module are extracted respectively, and the two features are added and sent to the first downsampling convolution layer until the feature extraction of the fifth layer is completed; Step 5: Send the features output by the three-dimensional road target detection model in step 4 to the region generation module RPN, perform downsampling, upsampling and channel connection operations, and extract semantic features and preliminary target bounding boxes; Step 6: Send the semantic features and preliminary target bounding box obtained in step S5 to the target detection network Centerhead for final prediction. The predicted classification includes the detected target type and the three-dimensional bounding box of the target.

2. The three-dimensional road target detection method based on improved point cloud sparse convolution according to claim 1 is characterized in that: In the three-dimensional road object detection model, the specific operations after the voxel feature tensor enters the improved focal sparse convolution module of the first layer are as follows: Step 1) construct two 0-valued two-dimensional tensors with the number of voxels in the x and y dimensions as length and width, fill the height and 1 values ​​respectively based on the voxel position tensor, and obtain the height-filled tensor and the valued mask; Step 2), translate the valued mask in the 3*3 plane according to 16 possible situations, and expand the mask before and after the translation to zero value; Step 3), add the expanded masks before and after the translation in step 2), take the position codes of all values ​​2 in the added tensor, remove the effects of translation and expansion, and obtain the valued feature position tensor that generates the expanded features before translation and the corresponding valued feature position tensor after translation; Step 4), obtain the extended feature offset in each of the 16 possible translation situations, add it to the valued feature position tensor, and obtain the extended feature plane position tensor; Step 5), use the two valued feature position tensors obtained in step 3) to fill the height tensor with the height tensor to obtain the height tensor, and obtain the height tensor before translation and the height tensor after translation. Subtract the two and divide by two to obtain the quotient to obtain the height difference tensor and determine whether the subtraction value is within the range of (-2, 2) to obtain the extended feature height mask; Step 6), concatenate the height tensor before translation with the extended feature plane position tensor, and filter with a mask to obtain the extended feature position tensor; Step 7), create a new zero-value feature tensor corresponding to the extended feature position tensor, concatenate the extended feature position tensor with the voxel position tensor, concatenate the zero-value feature tensor with the voxel feature tensor, and re-encode to obtain the voxel feature code; Step 8) The voxel feature encoding obtained in step 7) is put into two layers of sub-manifold sparse convolution to extract the features.

3. The three-dimensional road target detection method based on improved point cloud sparse convolution according to claim 1 is characterized in that: In the backbone network of the three-dimensional road object detection model, the large receptive field sparse convolution module performs the following operations on the input features: First, a position weighted attention module is constructed. The voxel feature tensor and the voxel position tensor are input into the position weighted attention module together to obtain voxel features with position weights. , the specific formula of the weighted attention module is as follows: ; Where sigmod represents the sigmod activation function, It represents a submanifold sparse convolution with n input channels and 1 output channel; represents the input voxel feature tensor, Represents the input voxel position tensor; Then, the resulting voxel position tensor In three-dimensional space, the quotient is compressed by multiples of 7, and the positions with the same quotient value correspond to Add the eigenvalues ​​of to obtain the compressed voxel feature tensor; Next, a submanifold sparse convolution is performed on the compressed voxel feature tensor; Secondly, the compressed voxel features after the submanifold sparse convolution are refilled in the compressed encoding order and divided by the position weight to obtain a new voxel feature tensor; Finally, the new voxel feature tensor is fed into two layers of submanifold sparse convolution to obtain a large receptive field tensor.

4. The three-dimensional road target detection method based on improved point cloud sparse convolution according to claim 1, characterized in that: In the backbone network of the three-dimensional road target detection model, each submanifold sparse convolution module includes two layers of submanifold sparse convolution and normalization layers, and outputs spatial features by performing two layers of submanifold sparse convolution and normalization on the input features.

5. The three-dimensional road target detection method based on improved point cloud sparse convolution according to claim 1, characterized in that: In the backbone network of the three-dimensional road object detection model, the number of input channels is 16, and the number of channels of the three downsampling convolutional layers is 32, 64 and 128 respectively.

Citation Information

Patent Citations

  • Point cloud target detection method based on improved SECOND network

    CN115457335A

  • Three-dimensional target detection method based on spatial adaptive sparse convolution

    CN117765242A