Three-dimensional target detection method and system based on multi-branch feature aggregation
By building a multi-branch feature aggregation network and dynamically aggregating point cloud and image features, the problem of single point cloud feature form and insufficient fusion is solved, and the accuracy and robustness of three-dimensional object detection is improved, especially in sparse object detection.
Patent Information
- Application Number
- CN202510240755.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-08
AI Technical Summary
The existing three-dimensional object detection method has a single form of point cloud features during the fusion stage, and the point cloud features and image features are not fully fused, resulting in low accuracy of sparse object detection.
Build a multi-branch feature aggregation network, including voxel feature extraction module, image feature extraction module, three-branch point voxel image aggregation subnet, foreground point pooling subnet and fine-grained feature fusion enhancement module, and dynamically aggregate point cloud and image features through multi-branching methods to enhance feature interaction.
提高了模型检测精度和鲁棒性,特别是在稀疏目标检测上表现出色,提升了对稀疏点云目标的检测准确率。
Smart Images

Figure CN120279364A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and in particular relates to a three-dimensional target detection method and system, which can be used by artificial intelligence machines and equipment to perceive surrounding 3D scenes. Background Art
[0002] At present, three-dimensional target detection technology is mainly based on two types of data: lidar point cloud and visible light image. The detection methods for three-dimensional targets are mainly divided into three-dimensional target detection methods based on point cloud data and multimodal three-dimensional target detection methods based on the fusion of point cloud and visible light image.
[0003] Typical methods for 3D object detection based on point cloud data include Shi et al. in Pv-rcnn: Point-voxelfeature set abstraction for 3D object detection[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2020:10529-10538., which use the Voxel-SA module to splice the voxel features obtained after voxel division to the corresponding key points, and then use the regional pooling module to aggregate the features of these sampling points to obtain the local features of the region of interest. This detection method based on the fusion of point and voxel features can take into account the advantages of point representation and voxel feature representation, and capture the rich contextual information of the scene. At the same time, this feature representation method also brings a certain amount of computation.
[0004] Since the single-modal 3D target detection method based on point cloud data lacks color and texture information and only uses LiDAR point cloud data for feature extraction, it cannot perform accurate detection when processing sparse targets in LiDAR point clouds.
[0005] In response to the problems of single-modal 3D target detection methods, many researchers have tried to fuse multi-modal data to improve the accuracy of target detection. Multi-modal 3D target detection based on the fusion of point cloud and visible light image has become one of the current hot research topics. Compared with point cloud, under good lighting conditions, images have higher resolution and contain rich semantic information such as color and texture, which can serve as a good information supplement for point cloud data. Fully fusing them can effectively solve the problems of false detection and misdetection caused by the lack of sparse target information in LiDAR point cloud.
[0006] The patent document with the application number CN202111648759.7 discloses "A 3D Object Detection Method and System Based on Multimodal Fusion". It extracts high-dimensional semantic features of the observed point cloud and the image through a multi-branch feature extraction network for feature extraction. The point cloud features and image features extracted are fused through an adaptive fusion module to obtain cross-modal fusion features, and the cross-modal fusion features are sent into the candidate box generation network to obtain the initial detection result. Through the multimodal pooling aggregation module, enhanced local features are constructed for the candidate boxes. Through the optimization network, the generated local enhanced features are sent into the optimization network to optimize and obtain the 3D object detection result.
[0007] Zhu et al. proposed a new architecture in VPFnet: Improving 3d object detection with virtual point - based lidar and stereo data fusion[J].IEEE Transactions on Multimedia,2022. It cleverly aligns and aggregates point cloud and image data in virtual points. In particular, the density of virtual points is between 3D points and 2D pixels, which can well bridge the resolution gap between the two sensors, thus retaining more information for processing.
[0008] Li et al. achieved the deep fusion of point cloud features and image features in LoGoNet: Towards Accurate 3D Object Detection with Local - to - Global Cross - Modal Fusion[C].Proceedings of the IEEE / CVF conference on computer vision and pattern recognition.2023 by local - to - global cross - modal feature fusion and introducing a cross - attention module in a parallel manner.
[0009] Although the above - mentioned multimodal 3D object detection methods based on the fusion of point cloud and visible light image can improve the accuracy of object detection through the fusion of point cloud and image, due to the single way of point cloud feature extraction, the advantages of various feature expression forms of the point cloud are not fully utilized, and a large amount of noise is introduced during fusion, it is easy to cause feature deviation caused by single - modality dominance, resulting in insufficient feature fusion and low accuracy of sparse point cloud object detection. Summary of the Invention
[0010] The object of the present invention is to provide a three-dimensional object detection method and system based on multi-branch feature aggregation, so as to solve the problem in the above-mentioned existing technology that the form of point cloud features is single in the fusion stage, and the fusion of point cloud features and image features is insufficient, resulting in low accuracy of sparse object detection.
[0011] The technical idea for achieving the object of the present invention is as follows: By constructing a voxel feature extraction network to extract voxel features; by constructing an image feature extraction network to extract image features; by constructing a three-branch point-voxel-image aggregation sub-network to fuse voxel and image features and then dynamically aggregate them to the sampled key points; by constructing a foreground point pooling sub-network to fuse the candidate boxes and image features and then dynamically aggregate them to the candidate box grid points and complete max pooling; by constructing a fine-grained feature fusion enhancement module to strengthen the interaction of the three-branch features and enhance the target ROI features.
[0012] According to the above technical idea, the technical solution of the present invention includes:
[0013] 1. A three-dimensional object detection method based on multi-branch feature aggregation, characterized by comprising:
[0014] Obtain a multi-modal data set of laser point clouds and visible light images, perform preprocessing and division into training set and test set on it;
[0015] Construct a three-dimensional object detection network based on multi-branch feature aggregation, which includes a farthest point sampling module, a voxel feature extraction module, an image feature extraction module, a three-branch point-voxel-image aggregation sub-network, a foreground point pooling sub-network and a fine-grained feature fusion enhancement module. After the farthest point sampling module, the voxel feature extraction module and the image feature extraction module are connected in parallel, they are cascaded with the three-branch point-voxel-image aggregation sub-network, the foreground point pooling sub-network and the fine-grained feature fusion enhancement module in sequence;
[0016] Establish a detection head including a linear operation unit and two fully connected operation units;
[0017] Set the loss function of the three-dimensional object detection network: L ALL =L rpn +L rcnn +L seg where L rpn is the loss of the region candidate box generation network, L rcnn is the loss of candidate box refinement, and L seg is the key point segmentation loss;
[0018] Input the training set into the three-dimensional object detection network, use the Adam optimizer and the gradient descent method, and perform iterative training with the goal of minimizing the loss function to obtain a trained three-dimensional object detection network.
[0019] Input the test set into the trained 3D object detection network and output the detection results.
[0020] Further, the three-branch point-voxel-image aggregation sub-network includes a voxel centroid offset projection module, a cross-attention module, a point-voxel-image aggregation module, and a candidate box generation unit. Their structural relationship is as follows: The voxel centroid offset projection module is cascaded with the cross-attention module, and the point-voxel-image aggregation module and the candidate box generation unit are in parallel and then connected after the cross-attention module.
[0021] Further, the foreground point pooling sub-network consists of a predicted key point weight module, a grid center offset projection module, a cross-attention module, and a fused key point ROI pooling unit cascaded in sequence.
[0022] Further, the fine-grained feature fusion and enhancement module includes a fully connected operation unit, a residual channel attention unit, a Transformer encoding unit, and an FFN residual network unit. Their structural relationship is as follows: The fully connected operation unit and the residual channel attention unit are in parallel and then cascaded with the Transformer encoding unit and the FFN residual network unit in sequence.
[0023] 2. A 3D object detection system based on multi-branch feature aggregation, characterized by including:
[0024] A voxel feature extraction module for extracting voxel features of point cloud data;
[0025] An image feature extraction module for extracting image features of image data;
[0026] A farthest point sampling module for extracting sampled key points;
[0027] A three-branch point-voxel-image aggregation module for aggregating voxel features and image features onto the sampled key points to generate key point features and 3D candidate boxes;
[0028] A foreground point pooling module for aggregating key point features onto the 3D candidate box grid to generate pooled features;
[0029] A fine-grained feature fusion and enhancement module for fusing key point features, image features, and pooled features to generate enhanced ROI features;
[0030] A 3D detection head module for outputting the enhanced ROI features as 3D detection boxes and corresponding confidence levels.
[0031] Further, the three-branch point-voxel-image aggregation module includes:
[0032] The body mass center offset projection sub-module is used to calculate the projection of the body mass center in the point cloud coordinate system on the reference point of the image and obtain the image features around the reference point;
[0033] The cross-attention sub-module is used to fuse the voxel features and the image features around their reference points to obtain fused features;
[0034] The point-voxel image feature aggregation module is used to aggregate the fused features onto the sampled key points to obtain key point features;
[0035] The candidate box generation unit sub-module is used to generate 3D candidate boxes according to the fused features.
[0036] Furthermore, the foreground point pooling module includes:
[0037] The predicted key point weight sub-module is used to enhance the foreground point features in the sampled key points;
[0038] The grid center dynamic offset projection sub-module is used to calculate the projection of the 3D candidate box grid center in the point cloud coordinate system on the reference point of the image and obtain the image features around the reference point;
[0039] The cross-attention sub-module is used to fuse the grid features and the image features around their reference points to obtain grid enhanced features;
[0040] The fused key point ROI pooling sub-module is used to aggregate the key point features onto the grid points to generate pooling features;
[0041] Compared with the prior art, the present invention has the following advantages:
[0042] Firstly, by establishing a three-branch point-voxel image aggregation sub-network and a foreground point pooling sub-network, and adopting a three-branch feature extraction method, the present invention makes full use of the point cloud features and the image features, fuses the point cloud features in the form of voxels and points with the image features, and innovatively dynamically aggregates them in the local area, making up for the problem of low accuracy of sparse target detection caused by the single form of point cloud features and the insufficient fusion of point cloud features and image features in the fusion stage, and effectively improving the model detection accuracy and robustness.
[0043] Secondly, by establishing a fine-grained feature fusion enhancement module to finely and adaptively enhance the features with less information interaction, the important channel features are enhanced, enabling the network to focus on the feature channels related to the target, suppressing the information of irrelevant channels, and further improving the robustness and detection accuracy of the model. Brief Description of the Drawings
[0044] Figure 1 is the implementation flowchart of the three-dimensional target detection method based on multi-branch feature aggregation of the present invention;
[0045] Figure 2 Schematic diagram of the 3D object detection network structure based on multi-branch feature aggregation constructed in the method of the present invention;
[0046] Figure 3 is Figure 2 Schematic diagram of the voxel feature extraction module structure in
[0047] Figure 4 is Figure 2 Schematic diagram of the image feature extraction module structure in
[0048] Figure 5 is Figure 2 Schematic diagram of the three-branch voxel image aggregation sub-network in
[0049] Figure 6 is Figure 2 Schematic diagram of the foreground point pooling sub-network in
[0050] Figure 7 is Figure 2 Schematic diagram of the fine-grained feature fusion enhancement module in
[0051] Figure 8 is Figure 2 Schematic diagram of the 3D detection head structure in
[0052] Figure 9 Block diagram of the 3D object detection system based on multi-branch feature aggregation of the present invention;
[0053] Figure 10 is Figure 9 Schematic diagram of the three-branch voxel image aggregation module in
[0054] Figure 11 is Figure 9 Schematic diagram of the foreground point pooling aggregation module;
[0055] Figure 12 Comparison chart of the detection results of the present invention and existing 3D object detection methods. Specific implementation manner
[0056] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] Example 1, 3D object detection method based on multi-branch feature aggregation.
[0058] Reference Figure 1 , the implementation steps of this example are as follows:
[0059] Step 1, obtain the dataset and perform preprocessing and partitioning.
[0060] 1.1) Download the KITTI 3D object detection algorithm evaluation dataset from the public website. This dataset contains 7,481 point cloud files and 7,481 paired visible light images. The image resolution is 1242×375, and the main target categories in the dataset are cars, pedestrians, and bicycles;
[0061] 1.2) Perform preprocessing on the point cloud files and visible light images:
[0062] Set the three-dimensional space range according to the effective detection range of the sensor. In this example, but not limited to: the X-axis is 0 to 70.4m, the Y-axis is -40 to 40m, the Z-axis is -3 to 1m, erase the point cloud beyond this range, and extract the point cloud within the camera's field of view from the erased point cloud according to the KITTI annotation file;
[0063] Crop the visible light images in KITTI. In this example, the cropping size is set to but not limited to 1024×375; and normalize the cropped visible light images, that is, divide their pixel values by 255 to obtain the normalized visible light images;
[0064] 1.3) Save the preprocessed point cloud files and visible light images in different folders respectively. At the same time, save the file paths of the point cloud files, image files, and annotation files, the category and annotation box information of the targets in the KITTI dataset, and the number of point cloud data into a txt file to generate an information file;
[0065] 1.4) Re-partition the point cloud files and visible light images saved in 1.3) into a training set and a test set according to a ratio of 8:2. In this example, randomly select 5,985 point cloud files and the corresponding visible light images as the training set, and the remaining point cloud files and visible light images as the test set.
[0066] Step 2: Select a feature extraction module.
[0067] 2.1) Select the farthest point sampling module:
[0068] The said farthest point sampling module is a functional module existing for farthest point sampling, and its specific sampling implementation is as follows:
[0069] 2.1.1) Initialize the selected set;
[0070] 2.1.2) Randomly select a point from the input point set as the first current selected point and put it into the selected set;
[0071] 2.1.3) Calculate the point in the input point set that is farthest from the currently selected point, and add it to the selected set;
[0072] 2.1.4) Update the current selected point to the farthest point obtained in 2.1.3);
[0073] 2.1.5) Repeat steps 2.1.3) to 2.1.4) until the required number of sampled points is selected.
[0074] 2.2) Select the voxel feature extraction module:
[0075] The voxel feature extraction module is composed of a voxel sampling layer, voxel feature encoding, and multiple 3D convolutional layers connected in sequence. In this example, but not limited to, 6 3D convolutions are set. Each 3D convolutional layer is composed of multiple 3D convolutional blocks cascaded in sequence. Each 3D convolutional block is composed of a 3D submanifold sparse convolution or a 3D conventional sparse convolution, a normalization layer, and a ReLU function layer cascaded. The convolution kernel size of the convolutional block is 3×3×3, and the stride is 1, as Figure 3 shown.
[0076] It should be noted that: if the convolution in the 3D convolutional block is a 3D submanifold sparse convolution, it is called a 3D submanifold convolutional block; if the convolution in the 3D convolutional block is a 3D conventional sparse convolution, it is called a 3D conventional convolutional block.
[0077] This voxel feature extraction module is used to generate voxel features, and its specific implementation is as follows:
[0078] 2.2.1) The voxel sampling layer voxelizes the point cloud under the entire scene with the centroid, and removes empty voxels from the obtained voxels to ensure the robustness of the network. The voxel size is [0.05, 0.05, 0.1];
[0079] 2.2.2) The voxel feature encoding calculates the sum of the features of all points in each voxel according to the number of input voxels, the maximum number of points in each voxel, and the features of each point, and then obtains the preliminary voxel features after normalization processing;
[0080] 2.2.3) Pass the preliminary voxel features through 6 3D convolutional layers in sequence to obtain voxel features;
[0081] 2.3) Select an image feature extraction module for generating image features. It is composed of an image block layer, a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, and a fourth feature extraction layer cascaded in sequence. Among them: the first feature extraction layer consists of a linear embedding module and 2 Swin Transformer blocks; the second feature extraction layer consists of an image stitching module and 2 Swin Transformer blocks; the third feature extraction layer consists of an image stitching module and 6 Swin Transformer blocks; the fourth feature extraction layer consists of an image stitching module and 2 Swin Transformer blocks, as Figure 4 shown.
[0082] Step 3: Construct a three-branch voxel image feature aggregation sub-network.
[0083] Refer to Figure 5 , the implementation of this step includes the following:
[0084] 3.1) Establish a voxel centroid offset projection module:
[0085] The voxel centroid offset projection module is composed of a coordinate calculation layer, a coordinate projection conversion layer, and an image feature stitching layer cascaded in sequence. It is a functional module for obtaining the features around the image projection point of the voxel centroid. Among them:
[0086] The coordinate calculation layer calculates the relative coordinates of the point cloud within the voxel with respect to the voxel centroid point, obtaining the voxel centroid point coordinates in the lidar coordinate system and the relative coordinates of the point cloud within the voxel;
[0087] The coordinate projection conversion layer projects the voxel centroid point coordinates into two-dimensional coordinate points on the image plane through the external parameter matrix and internal parameter matrix provided by the dataset, and converts the relative coordinates of the point cloud within the voxel into relative coordinates on the image plane;
[0088] The image feature stitching layer extracts a set of image features F around it through the two-dimensional coordinate points on the image plane, the relative coordinates on the image plane, and the global image features i ;
[0089] 3.2) Establish a cross-attention module including a fully connected unit and an attention mechanism:
[0090] The fully connected unit is composed of a linear layer, a normalization layer, a ReLU function layer, and a max pooling layer cascaded in sequence, for feature dimension alignment;
[0091] The attention mechanism is used to fuse the image feature F i with the voxel features output by the voxel feature extraction module to generate a fused feature F fuse ;
[0092] The cross-attention module is a functional module for fusing voxel features and image features;
[0093] 3.3) Establish a point-voxel-image feature aggregation module formed by cascading a multi-scale grouping layer, a dynamic aggregation PointNet layer, and a max pooling layer in sequence:
[0094] The multi-scale grouping layer divides the entire point cloud scenic area into N' local domains based on the spherical radius query method using each sampling key point, and each local domain contains the centroid points of several voxels;
[0095] The dynamic aggregation PointNet layer is used to aggregate the centroid point features of voxels in the local area as follows:
[0096] Calculate the voxel feature F l and the image feature F i modal correlation C corr ;
[0097] C corr =σ(MLP([F l ; F i )), where σ is the ReLU function;
[0098] Calculate the dynamic weight W corr of the modal correlation C fuse and the fused feature F agg , and generate a new fused feature
[0099] W agg =Softmax(MLP([F fuse ; C corr ))
[0100]
[0101] Aggregate the new fused feature to the sampling key points in each local domain to obtain the key point features;
[0102] 3.4) Establish a candidate box generation unit formed by cascading ten convolutional blocks and two deconvolutional blocks in sequence, where:
[0103] Each convolutional block is formed by cascading a convolutional layer, a normalization layer, and a ReLU function layer in sequence, and each deconvolutional block is formed by cascading a deconvolutional layer, a normalization layer, and a ReLU function layer in sequence;
[0104] The convolution kernel size of the convolution layer of the sixth convolution block is 3×3, and the stride is 2. The convolution kernel sizes of the convolution layers of the other nine convolution blocks are all 3×3, and the strides are all 1;
[0105] The convolution kernel size of the convolution layer of the first deconvolution block is 1×1, and the stride is 1. The convolution kernel size of the convolution layer of the second deconvolution block is 2×2, and the stride is 2. The fused features are sent to F fuse The candidate box generation unit obtains 3D candidate boxes;
[0106] 3.5) Connect the voxel centroid offset projection module and the cross-attention module in series, connect the point voxel image aggregation module and the candidate box generation unit in parallel and then connect them, and then connect the parallel point voxel image aggregation module and candidate box generation unit after the cross-attention module to form the three-branch point voxel image aggregation sub-network.
[0107] Step 4: Construct the foreground point pooling sub-network.
[0108] Refer to Figure 6 , the implementation of this step includes the following:
[0109] 4.1) Establish a foreground point weight module composed of two fully connected layers, one linear layer, and one Sigmoid function cascaded in sequence. This module is used to determine whether a key point is a foreground point or a background point depending on the 3D detection box label information;
[0110] 4.2) Establish a grid center offset projection module including a voxel sampling layer, two convolution layers, one normalization layer, one ReLU function layer, and a projection calculation layer. Its structural relationship is: the voxel sampling layer, the first convolution layer, the normalization layer, the ReLU function layer, the second convolution layer, and the projection calculation layer are cascaded in sequence. The convolution kernel size of each convolution layer is 3×3, and the stride is 1;
[0111] The voxel sampling layer is used to divide the 3D candidate boxes obtained by the candidate box generation unit into 6×6×6 grids and generate grid features F g ;
[0112] The projection calculation layer converts the three-dimensional point cloud coordinates of the grid points into two-dimensional image coordinates by calculation and splices a set of surrounding image features F i2 ;
[0113] 4.3) Establish a cross-attention module including a fully connected unit and an attention mechanism:
[0114] The fully connected unit is composed of a linear layer, a normalization layer, a ReLU function layer, and a max pooling layer cascaded in sequence, and is used for feature dimension alignment;
[0115] The attention mechanism is used to process the grid features Fg Fuse with the image feature F i2 to obtain the fused grid feature F gfuse ;
[0116] The cross-attention module is a functional module for fusing grid features and image features;
[0117] 4.4) Establish a fused key-point ROI pooling unit including a dynamic aggregation PointNet layer, two convolutional layers, two normalization layers, a ReLU function layer, and a max pooling layer. Their structural relationship is as follows: the dynamic aggregation layer, the first convolutional layer, the first normalization layer, the second convolutional layer, the second normalization layer, the ReLU function layer, and the max pooling layer are cascaded in sequence; the convolutional kernel size of each convolutional layer is 3×3, and the stride is 1;
[0118] The dynamic aggregation PointNet layer is used to aggregate the key points within the grid points as follows:
[0119] Calculate the grid feature F g and the image feature F i2 modal correlation C corr2 ;
[0120] C corr2 =σ(MLP([F g ; F i )), where σ is the ReLU function;
[0121] Calculate the dynamic weights of the modal correlation C corr2 and the grid fusion feature F gfuse and generate a new grid fusion feature
[0122] W agg2 =Softmax(MLP([F gfuse ; C corr ))
[0123]
[0124] Aggregate the key-point features on the grid points through the PointNet aggregation method and add them point by point to the new grid fusion feature to obtain the grid-point aggregation feature F gp ;
[0125] 4.5) Cascade the foreground point weight module, the grid center offset projection module, the cross-attention module, and the fused key-point ROI pooling unit to form the foreground point pooling sub-network.
[0126] Step 5: Construct a fine-grained feature fusion and enhancement module.
[0127] Reference Figure 7 , the implementation of this step includes the following:
[0128] 5.1) Establish a fully-connected operation unit composed of a linear layer, a normalization layer, and a ReLU function layer cascaded in sequence;
[0129] 5.2) Establish a residual channel attention unit including an average pooling layer, two linear layers, a ReLU function layer, and a sigmoid function layer. Its structural relationship is: the average pooling layer, the first linear layer, the ReLU function layer, the second linear layer, and the sigmoid function layer are cascaded in sequence, and the input end of the average pooling layer is connected to the output end of the sigmoid function layer;
[0130] 5.3) Establish an FFN residual network unit including two convolutional layers, a normalization layer, and a ReLU function layer. Its structural relationship is: the first convolutional layer, the normalization layer, the ReLU function layer, and the second convolutional layer are cascaded in sequence, and the input end of the first convolutional layer is connected to the output end of the second convolutional layer;
[0131] 5.4) Establish a Transformer encoding unit for self-attention calculation;
[0132] 5.5) After paralleling the fully-connected operation unit and the residual channel attention unit, cascade them with the Transformer encoding unit and the FFN residual network unit in sequence to form the fine-grained feature fusion and enhancement module.
[0133] Step 6: Construct a 3D detection head.
[0134] Reference Figure 8 , the implementation of this step includes the following:
[0135] 6.1) Establish a linear operation unit including two linear layers, two normalization layers, and two ReLU function layers. Its structural relationship is: the first linear layer, the first normalization layer, the ReLU function layer, the second linear layer, the second normalization layer, and the ReLU function layer are cascaded in sequence;
[0136] 6.2) Establish two fully-connected operation units, each of which is composed of a linear layer, a normalization layer, and a ReLU function layer cascaded in sequence;
[0137] 6.3) After paralleling the first fully-connected operation unit and the second fully-connected operation unit, connect them after the linear operation unit to form the 3D detection head.
[0138] Step 7: Construct a three-dimensional object detection network based on multi-branch feature aggregation.
[0139] After connecting the above-mentioned farthest point sampling module, voxel feature extraction module, and image feature extraction module in parallel, and then cascading them with the three-branch voxel image aggregation sub-network, foreground point pooling sub-network, and fine-grained feature fusion and enhancement module in sequence, a three-dimensional object detection network based on multi-branch feature aggregation is constructed, as Figure 2 described.
[0140] Step 8: Set the loss function L of the three-dimensional object detection network based on multi-branch feature aggregation.
[0141] 8.1) According to the regression parameters predicted by the candidate box The regression target value △r of the candidate box prediction, the hyperparameter β for balancing different losses, and the anchor point classification loss L cls Calculate the loss L of the region candidate box generation network rpn :
[0142]
[0143] 8.2) According to the regression parameters for refining the candidate box The regression target value △r′ of the candidate box refinement and the IoU-guided confidence prediction loss L iou Calculate the candidate box refinement loss L rcnn :
[0144]
[0145] 8.3) According to the prediction result p of the key point segmentation seg and the ground truth label y of the key point segmentation seg Calculate the key point segmentation loss L seg :
[0146] L seg = Focal Loss(p seg , y seg )
[0147] 8.4) Add the above L rpn , L rcnn , L seg to obtain the loss function of the three-dimensional object detection network based on multi-branch feature aggregation:
[0148] L = L rpn + L rcnn + L seg .
[0149] Step 9: Train the three-dimensional object detection network based on multi-branch feature aggregation.
[0150] 9.1) Set the number of training epochs T and the hyperparameter β;
[0151] 9.2) Divide the preprocessed point cloud and visible light image into multiple point cloud image groups on average according to the batch size, and input the first point cloud image group and the corresponding annotation information into the 3D object detection network to obtain the initial weights and bias values of each convolution operation of the network;
[0152] 9.3) Calculate the loss value L corresponding to each point cloud image group in the training set:
[0153] 9.4) Use the Adam optimizer to update the parameters of the 3D object detection network to obtain the 3D object detection network model after the first parameter update;
[0154] 9.5) Input the second point cloud image group into the 3D object detection model after the first parameter update, and repeat steps 9.3) and 9.4) to obtain the 3D object detection network model after the second parameter update, and so on until all point cloud image groups are completely input, completing one training round;
[0155] 9.6) Repeat steps 9.3) to 9.5) until the set number of training rounds T is reached to obtain the trained 3D object detection model. In this example, it is set to but not limited to 80 rounds.
[0156] Step 10: Input the test set into the trained 3D object detection network and output the detection result.
[0157] Embodiment 2, a 3D object detection system based on multi-branch feature aggregation.
[0158] Refer to Figure 9 , this example includes a voxel feature extraction module 1, an image feature extraction module 2, a farthest point sampling module 3, a three-branch point-voxel image aggregation module 4, a foreground point pooling module 5, a fine-grained feature fusion enhancement module 6, and a 3D detection head module 7. Among them:
[0159] The three-branch point-voxel image aggregation module 4 includes a voxel centroid offset projection sub-module 41, a cross-attention sub-module 42, a point-voxel image feature aggregation sub-module 43, and a candidate box generation unit sub-module 44. The structural relationship is: the voxel centroid offset projection sub-module 41 is cascaded with the cross-attention sub-module 42. After the point-voxel image aggregation sub-module 43 and the candidate box generation unit sub-module 44 are connected in parallel, they are then connected after the cross-attention sub-module 42, as Figure 10 shown.
[0160] The foreground point pooling module 5 is composed of a predicted key point weight sub-module 51, a grid center dynamic offset projection sub-module 52, a cross-attention sub-module 53, and a fused key point ROI pooling sub-module 54 cascaded in sequence, as Figure 11 shown.
[0161] After the above voxel feature extraction module 1, image feature extraction module 2, and farthest point sampling module 3 are connected in parallel, they are cascaded with the three-branch point voxel image aggregation module 4, foreground point pooling module 5, fine-grained feature fusion and enhancement module 6, and 3D detection head module 7 in sequence to form a three-dimensional object detection system with multi-branch feature aggregation.
[0162] The working principle of the entire system is as follows:
[0163] The voxel feature extraction module 1, image feature extraction module 2, and farthest point sampling module 3 respectively extract features from the input point cloud data and image data. Among them: the voxel feature extraction module 1 extracts voxel features of the point cloud data, the image feature extraction module 2 extracts image features, and the farthest point sampling module 3 extracts sampling key points of the point cloud data;
[0164] These voxel features and image features extracted above are fused through the three-branch point voxel image aggregation module 4. The specific implementation is as follows: the voxel centroid offset projection sub-module 41 projects the centroid corresponding to the input voxel features onto the reference points of the image and obtains the image features around the reference points; the cross-attention sub-module 42 fuses the voxel features and the image features around their reference points to obtain fused features; the point voxel image feature aggregation sub-module 43 dynamically aggregates the fused features onto the sampling key points to generate key point features; the candidate box generation unit sub-module 44 generates 3D candidate boxes according to the fused features;
[0165] The foreground point pooling module 5 fuses the 3D candidate boxes and image features again, that is, the predicted key point weight sub-module 51 enhances the key point features and generates foreground point features; the grid center dynamic offset projection sub-module 52 projects the 3D candidate box grid onto the reference points of the image and obtains the image features around the reference points; the cross-attention sub-module 53 fuses the grid features and the image features around their reference points to obtain grid-enhanced features; the fused key point ROI pooling sub-module 54 aggregates the foreground point features enhanced by the predicted key point weight sub-module 51 onto the grid points enhanced by the cross-attention sub-module 53 to generate pooled features;
[0166] The fine-grained feature fusion and enhancement module 6 fuses the above pooled features, image features, and foreground point features, and strengthens the three-branch feature interaction to generate ROI-enhanced features; after the ROI-enhanced features are refined by the 3D detection head module 7, 3D detection boxes and corresponding confidence levels are output to complete three-dimensional object detection.
[0167] The effects of the present invention are further illustrated by the following simulation tests:
[0168] I. Simulation test conditions
[0169] Dataset: The KITTI dataset proposed by Andreas et al. was adopted;
[0170] Experimental platform: The CPU is Intel(R) Core(TM) i9-10900X @ 3.70GHz, with 64GB of memory. The operating system is Ubuntu 20.04, the graphics card is NVIDIA GeForce RTX 3090, and the Pytorch version is 1.10.1;
[0171] Parameter settings: In the present invention, the training batch size is set to 8, the initial learning rate is set to 0.01, the number of training epochs T is set to 80, and the hyperparameter β is set to 1.
[0172] II. Simulation test content and results
[0173] Simulation test 1: The present invention and the existing two multi-modal 3D object detection methods based on point cloud and visible light image fusion, namely VPFNe and LoGoNet, were respectively used to perform 3D object detection on three types of objects, namely cars, pedestrians, and bicycles, in the KITTI dataset. The detection results under different categories and different difficulties were obtained, and the average precision index AP of each was calculated. The results are shown in Table 1.
[0174] Table 1 Comparison of average precision of the present invention and 3 existing 3D object detection methods
[0175]
[0176] The calculation method of AP is as follows:
[0177] Define precision Precision and recall Recall as:
[0178]
[0179] TP is the number of positive samples predicted as positive samples, FP is the number of negative samples predicted as positive samples, TN is the number of negative samples predicted as negative samples, and FN is the number of positive samples predicted as negative samples. According to precision Precision and recall Recall, draw the P-R curve, that is, p(r), and obtain AP by calculating the definite integral of p(r). The calculation formula is as follows:
[0180]
[0181] In Table 1, easy, medium, and difficult respectively represent three different situations of the height of the minimum bounding box, the maximum truncation degree, and the occlusion degree of the dataset label box, as shown in Table 2.
[0182] Table 2 Definition of detection difficulty of KITTI dataset
[0183] Detection Difficulty Minimum Bounding Box Height Maximum Truncation Degree Occlusion Degree Easy 40 pixels 15% Fully Visible Medium 25 pixels 30% Partially Occluded Difficult 25 pixels 50% Severely Occluded
[0184] As can be seen from Table 1, the detection accuracy of the present invention for the three types of targets, namely cars, pedestrians, and bicycles, is superior to the other two 3D object detection methods under most difficulties, indicating that the present invention has obvious advantages in detection performance.
[0185] Test 2: Visualize the detection results of the present invention and the existing LoGoNet method in Simulation Test 1. The results are as Figure 12 shown, where:
[0186] Figure 12 (a) Detection results of the present invention and the existing LoGoNet method for the 279th group of data in the KITTI dataset;
[0187] Figure 12 (b) Detection results of the present invention and the existing LoGoNet method for the 2362nd group of data in the KITTI dataset;
[0188] Figure 12 (c) Detection results of the present invention and the existing LoGoNet method for the 308th group of data in the KITTI dataset;
[0189] In each figure, the gray points are lidar point clouds, the 3D boxes are the detected targets, and the white dashed circles indicate the missed detection areas.
[0190] From Figure 12 it can be seen that there are a large number of false detections and missed detections in the LoGoNet method, especially in the case of sparse point clouds. The method of the present invention can not only achieve accurate target positioning, but also improve the detection accuracy of the model for sparse point cloud targets.
[0191] The above simulation results show that: the multi-modal 3D object detection method based on multi-branch feature aggregation proposed by the present invention has very high robustness to target size and scene changes, can effectively improve the detection accuracy of sparse targets, especially for pedestrian and bicycle targets, and has excellent detection performance. Moreover, on the KITTI dataset, the 3D object detection performance of the present invention on lidar point clouds is superior to the other three existing methods.
[0192] It should be noted that the step numbers in the specification and claims of the present invention are only for clearly describing the embodiments of the present invention for easy understanding, and their sequence numbers are not limited.
Claims
1. A three-dimensional object detection method based on multi-branch feature aggregation, characterized in that, Including: Obtain a multi-modal dataset of lidar point clouds and visible light images, preprocess it, and divide it into a training set and a test set; Construct a 3D object detection network based on multi-branch feature aggregation, which includes a farthest point sampling module, a voxel feature extraction module, an image feature extraction module, a three-branch point-voxel-image aggregation sub-network, a foreground point pooling sub-network, and a fine-grained feature fusion enhancement module. After the farthest point sampling module, the voxel feature extraction module, and the image feature extraction module are connected in parallel, they are cascaded with the three-branch point-voxel-image aggregation sub-network, the foreground point pooling sub-network, and the fine-grained feature fusion enhancement module in sequence; Establish a 3D detection head including a linear operation unit and two fully connected operation units; Set the loss function of the 3D object detection network: L ALL = L rpn + L rcnn + L seg , where L rpn is the loss of the region proposal network, L rcnn is the loss of the proposal refinement, and L seg is the loss of the key point segmentation; Input the training set into the 3D object detection network, use the Adam optimizer and the gradient descent method, and perform iterative training with the goal of minimizing the loss function to obtain a trained 3D object detection network. Input the test set into the trained 3D object detection network and output the detection results.
2. The method according to claim 1, characterized in that, The obtaining of the multi-modal dataset of lidar point clouds and visible light images, preprocessing it, and dividing it into a training set and a test set are realized as follows: 2a) Obtain a multi-modal dataset including lidar point clouds and visible light images and their annotation files from a public website; 2b) Preprocess the dataset: For the input point cloud, specify the point cloud ranges of the X, Y, and Z axes, erase the point cloud that exceeds the point cloud range, and obtain the point cloud within the camera's field of view from the erased point cloud through the annotation file; For the input visible light image, crop its size to 1024×375; Normalize the cropped visible light image, that is, divide its pixel values by 255 to obtain the normalized visible light image; 2c) Divide the preprocessed dataset into a training set and a test set according to a ratio of 8:
2.
3. The method according to claim 1, wherein: The farthest point sampling module is used to uniformly sample key points in the point cloud data; The voxel feature extraction module is composed of a voxel sampling layer, a voxel feature encoding layer, an input convolutional layer, four feature convolutional layers, and an output convolutional layer in cascade, and is used for multi-layer voxel features; The image feature extraction module is composed of an image block layer and multiple layers of Swin Transformer in cascade, and is used to extract multi-layer image features.
4. The method according to claim 1, wherein The construction of the three-branch point-voxel-image aggregation sub-network is realized as follows: 4a) Establish a voxel centroid offset projection module: 4a1) Calculate the relative coordinates of the point cloud within the voxel with respect to the voxel centroid to obtain the voxel centroid coordinates in the lidar coordinate system and the relative coordinates of the point cloud within the voxel; 4a2) Convert the voxel centroid coordinates from the lidar coordinate system to the camera coordinate system through the external parameter matrix in the camera calibration parameters to obtain the three-dimensional coordinates of the voxel centroid in the camera coordinate system; 4a3) Convert the voxel centroid coordinates obtained in the second step to the two-dimensional coordinates p on the image plane through the internal parameter matrix in the camera calibration parameters, and convert the relative coordinates of the point cloud within the voxel to the relative coordinates on the image plane; 4a4) Extract a set of image features F around it through the planar two-dimensional coordinates p, the relative coordinates of the point cloud within the voxel, and the global image features i ; 4b) Establish a cross-attention module including a fully connected unit and an attention mechanism: The fully connected unit is composed of a linear layer, a normalization layer, a ReLU function layer, and a max pooling layer cascaded in sequence; The attention mechanism is used to fuse the image feature F i with the voxel features output by the voxel feature extraction module; 4c) Establish a point-voxel image feature aggregation module including a multi-scale grouping layer, a dynamic aggregation PointNet layer, and a max pooling layer operation: The multi-scale grouping layer divides the entire point cloud scenic area into N' local domains based on the spherical radius query using each sampling key point, and each local domain contains the centroid points of several voxels; The dynamic aggregation PointNet layer is used to aggregate the centroid point features of the voxels within the local area; The point-voxel image feature aggregation module is composed of a multi-scale grouping layer, a dynamic aggregation PointNet layer, and a max pooling layer cascaded in sequence; 4d) Establish a candidate box generation unit composed of ten cascaded convolutional blocks and two transposed convolutional blocks in sequence: Among them, each convolutional block is composed of a convolutional layer, a normalization layer, and a ReLU function layer in sequence, and each transposed convolutional block is composed of a transposed convolutional layer, a normalization layer, and a ReLU function layer in sequence. The convolutional kernel size of the convolutional layer of the sixth convolutional block is 3×3, the stride is 2, and the convolutional kernel sizes of the convolutional layers of the remaining nine convolutional blocks are all 3×3, and the strides are all 1; the convolutional kernel size of the convolutional layer of the first transposed convolutional block is 1×1, the stride is 1, and the convolutional kernel size of the convolutional layer of the second transposed convolutional block is 2×2, the stride is 2. Send the fused feature into F fuse The candidate box generation unit obtains 3D candidate boxes; 4e) The voxel centroid offset projection module and the cross-attention module are cascaded, and the point-voxel image aggregation module and the candidate box generation unit are connected in parallel and then connected after the cross-attention module to form the three-branch point-voxel image aggregation sub-network.
5. The method according to claim 4, characterized in that In step (4c), the dynamic aggregation PointNet layer is used to aggregate the centroid point features of the voxels within the local area, and its implementation includes the following: 4c1) Calculate the voxel feature F through the MLP layer and the activation function l and the image feature F i Modal correlation C corr ; C corr = σ(MLP([F l ; F i )), where σ is the ReLU function; 4c2) Calculate the modality correlation C through the MLP layer and the Softmax function corr with the fused feature F fuse for the dynamic weight W agg , and generate a new fused feature W agg = Softmax(MLP([F fuse ; C corr )) 4c3) Aggregate the new fused features using the PointNet aggregation method onto the sampled key points in each local area to obtain key point features.
6. The method according to claim 1, wherein The implementation of the foreground point pooling sub-network includes the following: 6a) Establish a foreground point weight module composed of two fully connected layers, a linear layer, and a Sigmoid function cascaded in sequence: 6b) Establish a grid center offset projection module including a voxel sampling layer, two convolutional layers, a normalization layer, a ReLU function layer, and a projection calculation layer. Its structural relationship is: the voxel sampling layer, the first convolutional layer, the normalization layer, the ReLU function layer, the second convolutional layer, and the projection calculation layer are cascaded in sequence. The convolutional kernel size of each convolutional layer is 3×3, and the stride is 1; The voxel sampling layer is used to divide the 3D candidate boxes obtained by the candidate box generation unit into a 6×6×6 grid to generate grid feature F g ; The projection calculation layer converts the three-dimensional point cloud coordinates of the grid points into two-dimensional image coordinates through calculation, and extracts a set of surrounding image features F i2 ; 6c) Establish a cross-attention module including a fully connected unit and an attention mechanism: The fully connected unit is composed of a linear layer, a normalization layer, a ReLU function layer, and a max pooling layer cascaded in sequence; The attention mechanism is used to cross - fuse the grid feature F g with the image feature F i2 to obtain the fused grid feature F gfuse ; 6d) Establish a fused key point ROI pooling unit including a dynamic aggregation PointNet layer, two convolutional layers, two normalization layers, a ReLU function layer, and a max pooling layer. Its structural relationship is: the dynamic aggregation PointNet layer, the first convolutional layer, the first normalization layer, the second convolutional layer, the second normalization layer, the ReLU function layer, and the max pooling layer are cascaded in sequence; the convolutional kernel size of each convolutional layer is 3×3, and the stride is 1. The dynamic aggregation PointNet layer is used to aggregate the key point features within the grid; 6e) The predicted key point weight module, the grid center offset projection module, the cross-attention module, and the fused key point ROI pooling unit are cascaded in sequence to form the foreground point pooling sub-network.
7. The method according to claim 1, wherein The implementation of the construction of the fine-grained feature fusion enhancement module includes the following: 7a) Establish a fully connected operation unit composed of a linear layer, a normalization layer, and a ReLU function layer cascaded in sequence; 7b) Establish a residual channel attention unit including an average pooling layer, two linear layers, a ReLU function layer, and a sigmoid function layer. The structural relationship is as follows: the average pooling layer, the first linear layer, the ReLU function layer, the second linear layer, and the sigmoid function layer are cascaded in sequence, and the input end of the average pooling layer is connected to the output end of the sigmoid function layer; 7c) Establish an FFN residual network unit including two convolutional layers, a normalization layer, and a ReLU function layer. The structural relationship is: the first convolutional layer, the normalization layer, the ReLU function layer, and the second convolutional layer are cascaded in sequence, and the input end of the first convolutional layer is connected to the output end of the second convolutional layer; 7d) Establish a Transformer encoding unit for self-attention calculation; 7e) After paralleling the fully connected operation unit and the residual channel attention unit, cascade them with the Transformer encoding unit and the FFN residual network unit in sequence to form the fine-grained feature fusion and enhancement module.
8. The method according to claim 1, wherein The 3D detection head is implemented as follows: 8a) Establish a linear operation unit including two linear layers, two normalization layers, and two ReLU function layers. The structural relationship is: the first linear layer, the first normalization layer, the ReLU function layer, the second linear layer, the second normalization layer, and the ReLU function layer are cascaded in sequence; 8b) Establish two fully connected operation units, each of which is composed of a linear layer, a normalization layer, and a ReLU function layer cascaded in sequence; 8c) After paralleling the first fully connected operation unit and the second fully connected operation unit, connect them after the linear operation unit to form the 3D detection head.
9. The method according to claim 1, characterized in that: In the loss function L of the three-dimensional object detection network, the loss L of the region candidate box generation network involved rpn , the loss L of candidate box refinement rcnn , the loss L of key point segmentation seg are respectively expressed as follows: L = L rpn + L rcnn + L seg L seg = Focal Loss(p seg , y seg ) Among them, Δr a is the regression parameter for candidate box prediction, Δr a is the regression target value for candidate box prediction, Δr p is the regression parameter for candidate box refinement, Δr p is the regression target value for candidate box refinement, β is the hyperparameter for balancing different losses, p seg is the prediction result of key point segmentation, y seg is the ground truth label of key point segmentation.
10. The method according to claim 1, characterized in that, Using the Adam optimizer and the gradient descent method, iteratively train the 3D object detection network with the goal of minimizing the loss function. Its implementation includes the following: 10a) Set the number of training epochs T and the hyperparameter β; 10b) Divide the preprocessed point cloud and visible light image into multiple point cloud image groups on average according to the batch size, and input the first point cloud image group and the corresponding annotation information into the 3D object detection network to obtain the initial weights and biases of each convolutional operation in the network; 10c) Use the loss function to calculate the loss value L corresponding to each point cloud image group in the training set; 10d) Use the Adam optimizer to update the parameters of the 3D object detection network to obtain the 3D object detection network model after the first parameter update; 10e) Input the second point cloud image group into the 3D object detection model after the first parameter update, repeat 10c) and 10d), to obtain the 3D object detection network model after the second parameter update, and so on until all image groups are input completely to complete one training epoch; 10f) Repeat 10c) to 10e) until the set number of training epochs T is reached to obtain the trained 3D object detection model.
11. A three-dimensional object detection system based on multi-branch feature aggregation, characterized in that, Including: A voxel feature extraction module (1) for extracting voxel features of point cloud data; An image feature extraction module (2) for extracting image features of image data; A farthest point sampling module (3) for extracting sampled key points; The three-branch voxel image aggregation module (4) is used to aggregate voxel features and image features onto the sampled key points to generate key point features and 3D candidate boxes; The foreground point pooling module (5) is used to aggregate key point features onto the 3D candidate box grid to generate pooled features; The fine-grained feature fusion enhancement module (6) is used to fuse key point features, image features, and pooled features to generate enhanced ROI features; The 3D detection head module (7) is used to output the enhanced ROI features as 3D detection boxes and corresponding confidence levels.
12. The system according to claim 11, wherein: The three-branch voxel image aggregation module (4) includes: The voxel centroid offset projection sub-module (41) is used to calculate the projection of the voxel centroid in the point cloud coordinate system on the reference point of the image and obtain the image features around the reference point; The cross-attention sub-module (42) is used to fuse the voxel features and the image features around its reference point to obtain fused features; The point-voxel image feature aggregation module (43) is used to aggregate the fused features onto the sampled key points to obtain key point features; The candidate box generation unit sub-module (44) is used to generate 3D candidate boxes according to the fused features; The foreground point pooling module (5) includes: The predicted key point weight sub-module (51) is used to enhance the foreground point features in the sampled key points; The grid center dynamic offset projection sub-module (52) is used to calculate the projection of the 3D candidate box grid center in the point cloud coordinate system on the reference point of the image and obtain the image features around the reference point; The cross-attention sub-module (53) is used to fuse the grid features and the image features around its reference point to obtain grid-enhanced features; The fused key point ROI pooling sub-module (54) is used to aggregate the key point features onto the grid points to generate pooled features.
Citation Information
Patent Citations
A three-dimensional target detection method and system based on multimodal fusion
CN114519853B