Single-stage 3D point cloud object detection method based on neighboring box voting

By using the number of adjacent boxes in the prediction box and the average cross-over information, the category confidence in the three-dimensional point cloud target detection is corrected, and the problem of misalignment of category confidence and positioning accuracy is solved, the detection accuracy is improved, and the detection efficiency is maintained.

CN115953756BActive Publication Date: 2025-05-16XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211663148.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2025-05-16
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

In the existing three-dimensional point cloud object detection method, the category confidence and positioning accuracy are not aligned, resulting in a decrease in detection accuracy.

Method used

By using the number of adjacent boxes of the prediction box and the average intersection of the prediction box and the adjacent boxes, the category confidence is corrected to make it closer to the positioning accuracy, and the adjacent prediction box voting strategy is used for processing without changing the single-stage detector network structure.

Benefits of technology

The alignment of category confidence and positioning accuracy is achieved, the detection accuracy of three-dimensional point cloud object detection is improved, and additional computing overhead is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115953756B_ABST
    Figure CN115953756B_ABST
Patent Text Reader

Abstract

The present invention discloses a single-stage three-dimensional point cloud target detection method based on adjacent frame voting, which mainly solves the problem of misalignment between classification confidence and positioning accuracy in the prior art. The implementation scheme is: 1) obtaining point cloud data, dividing and preprocessing the data set; 2) building a three-dimensional point cloud target detection network and setting a loss function, and iteratively training it; 3) using the trained network to infer the test point cloud samples, filtering the prediction frames with low classification confidence; 4) using the adjacent frame voting strategy to correct the classification confidence; 6) using the non-maximum suppression method to filter the redundant prediction frames to obtain the target detection result. The present invention uses the adjacent frame information of the prediction frame to correct the classification confidence of the prediction frame, making it closer to the positioning accuracy, can effectively filter out low-quality prediction frames, and retain high-quality prediction frames, improve the detection accuracy of the three-dimensional target detector, and can be used to identify and locate targets in three-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a three-dimensional point cloud target detection method, which can be used to identify and locate targets in three-dimensional space. Background Art

[0002] As autonomous driving is becoming more and more popular, 3D object detection, which is closely related to autonomous driving, has become one of the most active research directions in the field of computer vision. 3D object detection aims to locate objects in 3D space and predict the object category. Only by accurately locating the object can the subsequent operation processes of autonomous driving, such as path planning, motion prediction, and collision avoidance, proceed smoothly.

[0003] There are usually two types of sensors used in autonomous driving systems, cameras and lidar. Cameras are divided into monocular cameras and stereo cameras. Stereo cameras can provide depth information compared to monocular cameras. The advantage of cameras is that they are cheap and can provide rich color and texture information, which helps to classify targets. However, cameras are limited by conditions such as lighting and weather, and have limited viewing angles, and cannot provide reliable information under all circumstances. Lidar can get rid of the constraints of the above conditions, can provide reliable sensor data under any lighting and weather conditions, has a 360° viewing angle, and naturally provides depth information without requiring additional calculations like stereo cameras. Therefore, the three-dimensional target detection method based on lidar, that is, the three-dimensional point cloud target detection method, has attracted widespread attention in academia and industry.

[0004] 3D point cloud object detectors usually have the problem of misalignment between positioning accuracy and category confidence, that is, a predicted box with accurate positioning has a lower category confidence, while a predicted box with poor positioning quality has a higher category confidence. This will cause the detector to filter out some predicted boxes with high positioning quality in the post-processing stage, while retaining some predicted boxes with poor positioning quality, which will lead to a decrease in detection accuracy.

[0005] Generally, two-stage detectors are less susceptible to this misalignment problem because they can use region proposal features to further refine the confidence and box positioning in the second stage. However, the introduction of the second-stage network also brings additional computational overhead. The single-stage detector is slightly lower in accuracy than the two-stage detector, but has higher detection efficiency, so it is often more practical.

[0006] In order to solve the problem of misalignment between the category confidence and positioning accuracy of single-stage detectors, Chen-Hang He, Hui Zeng, Jian-Qiang Huang, et al. proposed an interpolation-based confidence correction method in their paper "Structure Aware Single-Stage 3D Object Detection from Point Cloud" (Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition 2020). This method divides the prediction box into a grid, predicts a category confidence map for each grid point, and finally obtains the final category confidence map by averaging the category confidence maps of all grid points. However, due to the complexity of the operation, this method has limited confidence correction capabilities.

[0007] Wu Zheng, Wei-Liang Tang, Si-Jin Chen et al. proposed a confidence correction method based on IoU prediction in their paper "Cia-ssd: Confident iou-aware single-stage object detector from point cloud" (Proceedings of the Annual Conference of the International Association for Advanced Artificial Intelligence 2021). This method adds the IoU prediction task of the predicted box and the true box to the single-stage network, and corrects the classification confidence by the predicted IoU to make it closer to the positioning accuracy. However, since the single-stage network cannot extract information from the region proposal features, the IoU it predicts is not accurate. Summary of the invention

[0008] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a single-stage three-dimensional point cloud target detection method based on adjacent prediction box voting, so as to align the category confidence with the positioning accuracy and improve the detection accuracy.

[0009] The technical idea of ​​the present invention is: according to the characteristics that the number of adjacent boxes of the prediction box and the average intersection-union ratio of the prediction box and the adjacent boxes are often highly positively correlated with the positioning accuracy, the category confidence is corrected by using the two kinds of information, namely the number of adjacent boxes of the prediction box and the average intersection-union ratio of the prediction box and the adjacent boxes, so as to make it closer to the positioning accuracy; at the same time, without changing the existing network structure of the single-stage detector, the detection results are processed by the adjacent prediction box voting strategy.

[0010] To achieve the above object, the technical solution adopted by the present invention mainly includes the following steps:

[0011] (1) Obtain a point cloud dataset through LiDAR and divide it into a training set and a test set at a ratio of 1:1. Then perform data augmentation and voxelization preprocessing on the training set, and voxelization preprocessing on the test set;

[0012] (2) Build a 3D point cloud object detection network consisting of a voxel encoding layer, a 3D backbone network, a 2D multi-scale convolution backbone network, and a multi-task detection head cascade, and set the focal loss as the classification loss function l cls , smooth L1 loss is the box regression loss function l reg , the cross entropy loss is the target direction prediction loss function l dir ;

[0013] (3) Based on the preprocessed training set data, the Adam optimization algorithm is used to train the three-dimensional point cloud target detection network to obtain a trained three-dimensional point cloud target detection network;

[0014] (4) Input the preprocessed test set into the trained 3D point cloud object detection network to locate the objects in the environment and predict their categories in 3D space, and obtain the classification confidence and corresponding prediction box;

[0015] (5) Set the classification confidence threshold t1 to filter out low classification confidence and corresponding prediction boxes;

[0016] (6) Use the adjacent box voting strategy to correct the classification confidence and obtain more accurate positioning confidence:

[0017] 6a) Set the box adjacent threshold t2, calculate the intersection over union (IoU) between the predicted boxes, and compare it with the box adjacent threshold t2. If the IoU is greater than t2, the two predicted boxes are considered adjacent.

[0018] 6b) Count the number of adjacent boxes of each predicted box, cnt, and its average intersection-over-union ratio (iou) with adjacent boxes m ;

[0019] 6c) According to the result of 6b), the classification confidence c of each prediction box is calculated according to the following formula: i To make corrections:

[0020]

[0021] where s i Represents the predicted box b i The final positioning confidence;

[0022] 6g) Set the position confidence threshold t3 and set the final position confidence s i Compare it with its threshold t3 and filter out s i Prediction box b below threshold t3 i , we get the fixed position confidence set S = {s1,…,s i ,…,s n′} and the prediction box set B2 = {b1,…,b i,…,b n′}, where n′ represents the number of existing prediction boxes;

[0023] (7) Use non-maximum suppression to remove redundant prediction boxes in the prediction box set B2 to obtain the final detection result.

[0024] Compared with the prior art, the present invention has the following advantages:

[0025] First, the present invention utilizes the characteristic that the number of adjacent boxes of the prediction box and the average intersection-and-union ratio between the prediction box and the adjacent boxes are positively correlated with the true positioning accuracy, and fuses these two types of information, the number of adjacent boxes of the prediction box and the average intersection-and-union ratio between the prediction box and the adjacent boxes, with the classification confidence, to obtain a more accurate positioning confidence prediction.

[0026] Second, since the adjacent box voting confidence correction strategy used in the present invention is a post-processing technology, it is simpler to implement than the existing confidence correction technology. It does not require changing the network structure and does not bring additional computing overhead. It can correct the classification confidence and improve the final detection accuracy.

[0027] Third, since the proposed adjacent frame voting strategy can be combined with the existing confidence correction technology, the credibility of the confidence prediction can be further improved on the basis of the existing confidence correction technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is the overall flow chart of the implementation of the present invention;

[0029] Figure 2 It is a sub-flow chart for correcting classification confidence in the present invention. DETAILED DESCRIPTION

[0030] The embodiments and effects of the present invention are further described in detail below with reference to the accompanying drawings.

[0031] Reference Figure 1 , the implementation steps of this example are as follows:

[0032] Step 1: Obtain point cloud data and perform data set division and preprocessing.

[0033] 1.1) Use LiDAR to record point cloud data in multiple scenes and multiple time periods, sample the recorded point cloud data at a frequency of 2Hz to obtain 10,000 frames of point cloud data, annotate the three types of targets in the point cloud, namely motor vehicles, pedestrians, and non-motor vehicles, with 3D real frames, and divide the annotated point cloud data set into training and test sets in a 1:1 ratio;

[0034] 1.2) Perform data enhancement preprocessing on the point cloud samples in the training set:

[0035] 1.2.1) Randomly sample n from the training set s A 3D real frame and the target point cloud contained in it are selected, and then collision detection is performed. If the selected 3D real frame and the existing 3D real frame in the current point cloud sample have overlapping areas, it is considered that a collision has occurred. Finally, the target point cloud contained in the 3D real frame without collision is Paste to current point cloud In the point cloud Where N0 and N s Respectively represent the number of points in the original point cloud and the number of points in the sampled point cloud;

[0036] 1.2.2) Randomly flip the point cloud P1 along the x-axis or y-axis to perform random flip data enhancement;

[0037] 1.2.3) Randomly generate a rotation angle θ1 in the interval [-π / 4, π / 4], rotate the point cloud P1 around the z-axis by θ1, and perform random rotation data enhancement;

[0038] 1.2.4) Randomly generate a scale factor s1 in the interval [0.95, 1.05], and multiply the overall coordinate value of point cloud P1 by the scale factor s1;

[0039] 1.2.5) After the data enhancement, the order of the points in the point cloud P1 is disrupted, and the points and 3D real boxes outside the detection range are filtered out;

[0040] 1.3) Perform voxel preprocessing on the point cloud samples in the training set and the point cloud samples in the test set after data enhancement:

[0041] 1.3.1) Define the voxel size as s x ×s y ×s z , the point cloud is divided according to the defined voxel size within the detection range, and each voxel is allocated at most n p points;

[0042] 1.3.2) Define voxels that do not contain points as empty voxels, define voxels that contain points as non-empty voxels, and determine whether the number of points in a voxel exceeds n p indivual:

[0043] If there are more than n points in a voxel p If there are more than 100 points, the extra points will be discarded;

[0044] If less than n p If there are 0s, fill them with 0 to get a non-empty voxel set Among them, N1 is the number of non-empty voxels; represents the i-th non-empty voxel, which contains n ppoints, and the feature dimension of each point is 4.

[0045] Step 2: Build a 3D point cloud target detection network and set the loss function.

[0046] 2.1) Build the voxel encoding layer: This layer is responsible for encoding the point cloud in the voxel. The encoding method is to find the feature mean of the points contained in the voxel, that is, to input a non-empty voxel set Each non-empty voxel is encoded as follows:

[0047] where p j is the voxel v i The feature of the jth point in

[0048] Get the non-empty voxel feature set through the voxel encoding layer

[0049] 2.2) Build a 3D backbone network: It consists of four modules, each of which contains several 3D semi-manifold convolutional layers and a 3D sparse convolutional layer. Each convolutional layer is followed by a batch normalization layer and a nonlinear activation function ReLU. Each module downsamples and extracts 3D voxel features, and finally concatenates the 3D voxel features along the height dimension to obtain a bird's-eye view feature map. Where C is the number of feature channels, L and W are the length and width of the feature map respectively;

[0050] 2.3) Build a two-dimensional multi-scale convolutional backbone network:

[0051] The 2D multi-scale convolution backbone network consists of two modules, each of which consists of multiple convolutional layers. Each convolutional layer is followed by a batch normalization layer and a nonlinear activation function ReLU. The specific operations are as follows:

[0052] 2.3.1) The first module only inputs the bird’s-eye view feature map F BEV Perform feature extraction without changing the scale to obtain the first bird's-eye view feature map Where C1 represents the feature map F BEV1 The number of feature channels, L and W represent the feature map F BEV1 Length and width;

[0053] 2.3.2) The second module performs the first bird’s-eye view feature map F BEV1 Downsampling and further feature extraction are performed to obtain the second bird's-eye view feature map Where C2 represents the feature map F BEV2 The number of feature channels, L′ and W′ represent the feature map F BEV2 Length and width;

[0054] 2.3.3) The second bird's-eye view feature map FBEV2 By transposing the convolution and upsampling back to the original size, we get the third bird’s-eye view feature map Where C3 represents the feature map F BEV3 The number of feature channels, L and W represent the feature map F BEV3 Length and width;

[0055] 3.3.4) The first bird's-eye view feature map F BEV1 and the third bird's-eye view feature map F BEV3 Splicing to obtain a multi-scale bird's-eye view feature map

[0056] 2.4) Build a multi-task detection head:

[0057] The multi-task detection head consists of three convolutional layers with a convolution kernel size of 1×1, which are responsible for the classification task, the box regression task, and the target orientation task respectively. The three convolutional layers output the classification confidence prediction respectively. Box regression prediction And target orientation prediction in:

[0058] n c represents the number of categories, L and W represent the length and width of the network output respectively, the number 7 represents the seven attribute values ​​predicted by the box regression, which are the center point position (x, y, z), length, width and height (dx, dy, dz) of the three-dimensional target box, and the steering angle r, the number 2 represents the two directions of the target orientation prediction, the prediction of 1 represents the positive direction, and the prediction of 0 represents the reverse direction;

[0059] 2.5) Use focal loss as the classification loss function l cls , using smooth L1 loss as the box regression loss function l reg , using cross entropy loss as the target direction prediction loss function l dir , which are calculated as follows:

[0060]

[0061]

[0062]

[0063] Among them, α and β are equal to 0.25 and 2 respectively, N c and N p Represents the number of classification predictions and the number of positive anchor boxes, c i and Respectively represent the true classification value and classification prediction of the i-th position, b ij and Represents the jth attribute value of the predicted box and the real box at the i-th position, di and They represent the true value and predicted target orientation of the i-th position respectively.

[0064] Step 3: Iteratively train the 3D point cloud object detection network.

[0065] 3.1) Assume that the overall loss function of the 3D point cloud target detection network is:

[0066] l=l cls +λ1l reg +λ2l dir ,

[0067] Among them, λ1 and λ1 represent the weights of box regression loss and target orientation loss respectively;

[0068] 3.2) The preprocessed training set point cloud data is batch-inputted into the 3D point cloud object detection network, and the attribute values ​​of the real 3D object box and the attribute values ​​of the predicted 3D object box are used as the input of the loss function l, where the attribute values ​​include the object category, the position of the center point of the object box, the length, width, height and orientation of the object box;

[0069] 3.3) Use the Adam optimization algorithm to update the network parameters until the overall loss function value converges, stop training, and obtain a trained single-stage 3D point cloud target detection network.

[0070] Step 4: Use the trained network to infer the test set point cloud data.

[0071] 4.1) Input the voxelized test set point cloud data into the trained single-stage 3D point cloud object detection network, and obtain the classification confidence prediction, box regression prediction, and target orientation prediction through the network inference;

[0072] 4.2) Use the SECOND 3D target detector to decode the box regression prediction to obtain the predicted box, and correct the steering angle of the predicted box according to the target orientation prediction:

[0073] If the target orientation prediction is 0, the steering angle of the prediction box is negated;

[0074] If the target direction prediction is 1, keep the original value;

[0075] Finally, we get the initial prediction box set and the classification confidence set Where n0 represents the number of initial prediction boxes.

[0076] Step 5: filter prediction boxes with low classification confidence.

[0077] 5.1) Set the classification confidence threshold t1. If the classification confidence of a prediction box is lower than t1, filter out the prediction box.

[0078] 5.2) Collect the prediction boxes that are not filtered out and their corresponding classification confidences to obtain the prediction box set and the classification confidence set Where n1 represents the number of prediction boxes after filtering out prediction boxes with low classification confidence.

[0079] Step 6: Use the neighboring box voting strategy to correct the classification confidence.

[0080] Reference Figure 2 , the specific implementation of this step is as follows:

[0081] 6.1) Set the box adjacent threshold t2, calculate the intersection-and-union ratio between the predicted boxes, and compare it with the box adjacent threshold t2. If the intersection-and-union ratio is greater than t2, the two predicted boxes are considered adjacent.

[0082] 6.2) Count the number of adjacent boxes of each predicted box, cnt, and its average intersection-over-union ratio (iou) with adjacent boxes m ;

[0083] 6.3) According to the results of 6.2), the classification confidence c of each prediction box is calculated according to the following formula: i To make corrections:

[0084]

[0085] where s i Represents the predicted box b i The final positioning confidence;

[0086] 6.4) Set the position confidence threshold t3 and set the final position confidence s i Compare it with its threshold t3 and filter out s i Prediction box b below threshold t3 i , and obtain the location confidence set And the predicted box set Where n2 represents the number of prediction boxes after filtering by the location confidence threshold t3.

[0087] Step 7: Use non-maximum suppression to filter redundant prediction boxes.

[0088] 7.1) Input the location confidence set And the predicted box set According to the positioning reliability, the prediction boxes are sorted. The larger the positioning reliability, the higher the ranking. The sorted positioning reliability set S′ and the corresponding prediction box set B′ are obtained.

[0089] 7.2) Calculate the intersection-over-union ratio between the prediction boxes;

[0090] 7.3) Set the filtering threshold to t4, and take the i-th prediction box b i As a benchmark, judge b i Whether there is a filter tag:

[0091] If b i If it has a filter flag, then skip predicting box b i treatment;

[0092] If b i If there is no filter mark, traverse the prediction box b i+1 to Determine the prediction box With the prediction box b i Is the intersection-union ratio greater than the filtering threshold t4:

[0093] If it is greater than, the prediction box b j Add filter tags;

[0094] Otherwise, no filter mark is added;

[0095] 7.4) Traverse the prediction box b1 to Process each prediction box according to process 7.3)

[0096] 7.5) Collect the prediction boxes without filtering marks and their corresponding location confidence as the output of non-maximum suppression processing to complete the filtering of redundant prediction boxes.

[0097] The effects of the present invention are further described below in conjunction with simulation experiments.

[0098] 1. Simulation experiment conditions:

[0099] The simulation experiment hardware platform of the present invention is: AMD Reyzen 5900X CPU processor, 32GB memory, and Nvidia GeForce RTX 3090 graphics card.

[0100] The simulation experiment software platform of the present invention is: Ubuntu20.04 operating system and python3.8.

[0101] The data used in the simulation experiment of the present invention is: ONCE data set, which contains 8000 frames of public point cloud samples, 5000 frames of point cloud samples are used as training sets, and the other 3000 frames of point cloud samples are used as verification sets.

[0102] The present invention uses average precision AP as the evaluation index and performs evaluation in four detection distance intervals, namely global detection overall, close distance detection of 0-30m, medium distance detection of 30-50m, and long distance detection greater than 50m.

[0103] 2. Simulation content and results analysis:

[0104] A three-dimensional point cloud target detection network constructed using the present invention and four prior arts PointRCNN, Pointpillars, SECOND, and CenterPoints respectively;

[0105] The five networks are trained separately using the same training set data to obtain five trained 3D point cloud object detection networks;

[0106] The same validation set is input into the five trained networks to obtain the corresponding test results. The accuracy of each test result is evaluated and compared using the existing evaluation indicators. The comparison results are shown in Table 1:

[0107] Table 1: Comparison of simulation results between the present invention and the prior art

[0108]

[0109] The four prior art sources in Table 1 are as follows:

[0110] PointRCNN is a 3D point cloud target detection method proposed by Shi et al. in their paper "Pointrcnn: 3D object proposal generation and detection from point cloud" (IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2019), referred to as PointRCNN.

[0111] CenterPoint is a three-dimensional point cloud target detection method proposed by Yin et al. in their paper "Center-based 3D object detection and tracking" (IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2021), referred to as CenterPoint.

[0112] PointPillars is a three-dimensional point cloud target detection method proposed by Lang et al. in their paper "Pointpillars: Fast encoders for object detection from point clouds" (IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2019), referred to as PointPillars.

[0113] SECOND is a three-dimensional point cloud target detection method proposed by Yan et al. in their paper "Second: Sparsely embedded convolutional detection" (Sensors, 2018), referred to as SECOND.

[0114] As can be seen from Table 1, the three-dimensional point cloud target detection method of the present invention is superior to the four prior art methods in multiple evaluation indicators. This shows that the present invention corrects the classification confidence of the prediction frame by using the two types of information, the number of adjacent frames of the prediction frame and the average intersection-over-union ratio between the prediction frame and the adjacent frames, making it closer to the positioning accuracy, thereby effectively filtering low-quality prediction frames and retaining high-quality prediction frames, thereby improving the detection accuracy of the three-dimensional point cloud target detector.

Claims

1. A single-stage 3D point cloud object detection method based on adjacent box voting, characterized in that: These include: (1) Obtain a point cloud dataset through LiDAR and divide it into a training set and a test set in a 1:1 ratio. Then perform data augmentation and voxelization preprocessing on the training set and voxelization preprocessing on the test set. (2) Build a 3D point cloud object detection network consisting of a voxel encoding layer, a 3D backbone network, a 2D multi-scale convolution backbone network, and a multi-task detection head cascade, and set the focal loss as the classification loss function l cls , smooth L1 loss is the box regression loss function l reg , the cross entropy loss is the target direction prediction loss function l dir ; (3) Based on the preprocessed training set data, the Adam optimization algorithm is used to train the 3D point cloud target detection network to obtain a trained single-stage 3D point cloud target detection network; (4) Input the preprocessed test set into the trained 3D point cloud object detection network to locate the objects in the environment and predict their categories in 3D space, and obtain the classification confidence and corresponding prediction box; (5) Set the classification confidence threshold t1 to filter out low classification confidence and corresponding prediction boxes; (6) Use the adjacent box voting strategy to correct the classification confidence and obtain more accurate positioning confidence: 6a) Set the box adjacent threshold t2, calculate the intersection over union (IoU) between the predicted boxes, and compare it with the box adjacent threshold t2. If the IoU is greater than t2, the two predicted boxes are considered adjacent. 6b) Count the number of adjacent boxes of each predicted box, cnt, and its average intersection-over-union ratio (iou) with adjacent boxes m ; 6c) According to the result of 6b), the classification confidence c of each prediction box is calculated according to the following formula: i To make corrections: where s i Represents the predicted box b i The final positioning confidence; 6g) Set the position confidence threshold t3 and set the final position confidence s i Compare it with its threshold t3 and filter out s i Prediction box b below threshold t3 i , we get the fixed position confidence set S = {s1,…,s i ,…,s n′ } and the prediction box set B2 = {b1,…,b i ,…,b n′ }, where n′ represents the number of existing prediction boxes; (7) Use non-maximum suppression to remove redundant prediction boxes in the prediction box set B2 to obtain the final detection result.

2. The method according to claim 1, characterized in that In (1), the training set is preprocessed with data enhancement, which is implemented as follows: 1a) Enhance the 3D real-world box pasting data: Randomly select n from the training set s A 3D real frame and the target point cloud it contains will collide if there is an overlapping area with an existing 3D real frame. The target point cloud contained in the 3D real frame that has not collided will be Paste to current point cloud In the point cloud Where N0 represents the number of points in the original point cloud, N s Represents the number of points in the target point cloud contained in the sampled 3D true box, and 4 represents the feature dimension of the points in the point cloud; 1b) Randomly flip the point cloud P1 along the x-axis or y-axis to perform random flip data enhancement; 1c) Randomly generate a rotation angle θ1 in the interval [-π / 4, π / 4], rotate the point cloud P1 around the z-axis by θ1, and perform random rotation data enhancement; 1d) Randomly generate a scale factor s1 in the interval [0.95, 1.05], multiply the overall coordinate value of the point cloud P1 by the scale factor s1, and perform random scaling data enhancement on the point cloud; 1e) The order of the points in the point cloud P1 after the above data enhancement is disrupted, and the points and three-dimensional real boxes outside its detection range are filtered out.

3. The method according to claim 1, characterized in that In (1), the training set and the test set are preprocessed by voxelization, which is implemented as follows: 1f) Define the voxel size as s x ×s y ×s z , the point cloud is divided according to the defined voxel size within the detection range, and at most n points are assigned to each voxel; 1g) Define voxels that do not contain points as empty voxels, define voxels that contain points as non-empty voxels, and determine whether there are more than n points in the voxel: If there are more than n points in a voxel, the extra points are discarded; If there are less than n voxels, fill them with 0 to get a non-empty voxel set Where: N1 is the number of non-empty voxels; represents the i-th non-empty voxel, which contains n points, and the feature dimension of each point is 4.

4. The method according to claim 1, characterized in that: The voxel encoding layer in (2) is used to obtain a non-empty voxel set after the training set and the test set are voxelized. Each non-empty voxel in is encoded as follows: Get the non-empty voxel feature set through the voxel encoding layer Among them, p j is the nth voxel v i The feature of the jth point in , n is the number of points in each voxel, and N1 is the number of non-empty voxels.

5. The method according to claim 1, characterized in that The three-dimensional backbone network in (2) is used to extract features from three-dimensional voxels. It consists of four modules, each of which contains several three-dimensional semi-manifold convolutional layers and a three-dimensional sparse convolutional layer. Each convolutional layer is followed by a batch normalization layer and a nonlinear activation function ReLU. Each module downsamples and extracts features from three-dimensional voxel features, and concatenates the three-dimensional voxel features along the height dimension to obtain a bird's-eye view feature map. Where C is the number of feature channels, L and W are the length and width of the feature map respectively.

6. The method according to claim 1, characterized in that The two-dimensional multi-scale convolution backbone network in (2) is used to extract multi-scale features from the bird's-eye view feature map. It consists of two feature extraction modules, each of which is composed of multiple convolutional layers, and each convolutional layer is followed by a batch normalization layer and a nonlinear activation function ReLU; The first module only performs BEV feature map F BEV Perform feature extraction without changing the scale to obtain the first bird's-eye view feature map Where C1 represents the number of feature channels, L and W represent the feature maps F BEV1 Length and width; The second module has a BEV feature map F BEV1 Perform downsampling and further feature extraction to obtain the second bird's-eye view feature map And perform transpose convolution and upsample it back to its original size to get the third bird's-eye view feature map C2 and C3 both represent the number of feature channels, L′ and W′ represent the second bird’s-eye view feature map F BEV2 Length and width; The first bird's-eye view feature map F BEV1 With the third bird's-eye view feature F BEV3 Splicing to obtain a multi-scale bird's-eye view feature map 7. The method according to claim 1, characterized in that The multi-task detection head in (2) is composed of multiple independent parallel convolutional layers, which are used to complete the classification task, the box regression task and the target orientation task respectively.

8. The method according to claim 1, characterized in that In (3), based on the preprocessed training set data, the Adam optimization algorithm is used to train the three-dimensional point cloud target detection network, which is implemented as follows: (3a) Assume that the overall loss function of the 3D point cloud object detection network is: l = l cls +λ1l reg +λ2l dir , where l cls represents the classification loss, l reg represents the box regression loss, l dir represents the target orientation loss, λ1 and λ1 represent the weights of the box regression loss and the target orientation loss, respectively; (3b) The preprocessed training set point cloud data is batch-inputted into the 3D point cloud object detection network, and the attribute values ​​of the real 3D object box and the attribute values ​​of the predicted 3D object box are used as the input of the loss function l; (3c) Use the Adam optimization algorithm to update the network parameters until the overall loss function value converges and stop training.

9. The method according to claim 1, characterized in that: In (7), non-maximum suppression is used to remove redundant prediction frames in the prediction frame set B2, which is implemented as follows: 7a) Input the location confidence set S = {s1,…,s i ,…s n′ } and the prediction box set B2 = {b1,…,b i ,…b n′ }, sort the prediction boxes according to the positioning confidence. The larger the positioning confidence, the higher the ranking. Then, the sorted positioning confidence set S′ and the corresponding prediction box set B′ are obtained. 7b) Obtain the intersection-and-union ratio between two prediction frames according to the rotation intersection-and-union ratio calculation formula; 7c) Set the filtering threshold to t4, and take the i-th prediction box b i As a benchmark, judge b i Whether there is a filter tag: If b i If it has a filter flag, then skip predicting box b i If b i If there is no filter mark, traverse the prediction box b i+1 To b n′ , determine the prediction box b j∈[i+1,n′] With the prediction box b i Is the intersection-union ratio greater than the filtering threshold t4? If it is greater than, the prediction box b j Add filter tags; Otherwise, no filter mark is added; 7d) Traverse the prediction boxes b1 to b n′ , process each prediction box b according to 7c) i∈[1,n′] ; 7e) Collect the prediction boxes without filtering marks and their corresponding location confidence as the output of non-maximum suppression processing to complete the filtering of redundant prediction boxes.

Citation Information

Patent Citations

  • Front vehicle detection algorithm based on complex environment

    CN114120246A

  • Collaborative target detection method and system based on collaborative diagram fusion

    CN114913495A