A target detection method and system based on multi-scale fusion
Through the multi-scale fusion object detection method, virtual point clouds are generated and feature-level fusion is performed, which solves the problems of insufficient information interaction and noise in the prior art, and improves the accuracy and robustness of three-dimensional object detection.
Patent Information
- Application Number
- CN202410753353.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-06-12
AI Technical Summary
The existing multimodal fusion method based on camera and lidar sensors has problems such as insufficient information interaction, noise and data redundancy in target detection, resulting in limited detection accuracy and robustness.
The multi-scale fusion object detection method is adopted to realize data-level fusion by generating virtual point clouds, and feature-level fusion is used to use the first network and the second network, including cross-spatial feature fusion and noise perception, and adaptively fuse cross-layer spatial semantic information to improve detection accuracy.
It improves the accuracy and robustness of three-dimensional object detection, reduces noise problems, and achieves deep interactive feature enhancement and higher object detection accuracy.
Smart Images

Figure CN118628718B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic digital data processing. Specifically, it relates to an object detection method and system based on multi-scale fusion. Background Art
[0002] In the prior art, multi-modal fusion methods based on cameras and lidar sensors include data-level fusion, feature-level fusion, and decision-level fusion. The data-level fusion refers to directly superimposing point cloud data and image data element by element, and its information entropy loss is the smallest. However, this fusion method highly depends on the accuracy of the external parameters of the sensors, and there are also problems of noise and data redundancy in this method. The feature-level fusion method extracts the features of point cloud data and image data respectively, and then maps, correlates, and fuses the features. Therefore, it can make full use of the low-level and high-level features of different modalities, and the calculation speed is also faster. However, this fusion method is not flexible enough, resulting in insufficient information interaction between modalities. The decision-level fusion combines the detection results of different modal models and supplements the bounding boxes through deep learning and other methods, thereby improving the overall object detection accuracy and robustness. However, its effectiveness is restricted by the detection performance of a single model and is not dominant in multi-modal object detection.
[0003] In summary, in order to overcome the defects of the existing fusion methods, the present invention proposes an object detection method based on multi-scale fusion to improve the accuracy of three-dimensional object detection. Summary of the Invention
[0004] The present invention is precisely proposed based on the above requirements of the prior art. The technical problem to be solved by the present invention is to provide an object detection method and system based on multi-scale fusion to improve the accuracy of object detection.
[0005] To solve the above problems, the present invention is implemented by adopting the following technical solutions:
[0006] A target detection method based on multi-scale fusion, the method comprising: obtaining image features corresponding to image data and voxel features corresponding to point cloud data; using a first network to process the image features and voxel features to obtain cross-space fusion features, the first network including a first module, a second module, and a third module; the first module including determining neighborhood features around the voxel features; respectively performing convolution processing on the data projected from the neighborhood features into the image data and the image features, and splicing the processing results of the two to obtain a fusion feature; the second module including performing convolution on the fusion feature, splicing the convolution result and the image features in the feature channels, and performing convolution processing on the splicing result; respectively performing average pooling and max pooling on the convolved splicing result, and splicing the two pooling results to obtain a feature descriptor; performing convolution processing on the feature descriptor to obtain a first spatial attention; multiplying the first spatial attention and the convolved splicing result to obtain a noise-aware feature; the third module including splicing and fusing the voxel features and the noise-aware feature and performing downsampling to obtain cross-space fusion features of multiple different resolutions; performing bird's-eye view transformation on the cross-space fusion features of multiple different resolutions, and performing feature extraction on the transformed bird's-eye view to obtain corresponding bird's-eye features of different resolutions; using a region proposal network to process each bird's-eye feature to respectively obtain a plurality of candidate boxes; calculating local features at different resolutions, including dividing each candidate box at the same resolution into a plurality of sub-regions, and calculating the local features of each sub-region; using a second network to fuse the spatial information and semantic information of the local features with different resolutions to obtain candidate box feature information for detecting a target.
[0007] Optionally, the obtaining image features corresponding to image data and voxel features corresponding to point cloud data includes: obtaining image data, performing feature extraction on the image data to obtain image features; obtaining original point cloud data, using a PENet depth completion network to process the image data and the original point cloud data to generate virtual point clouds; performing hierarchical downsampling on the virtual point clouds according to the distance between the virtual point clouds and the lidar center; combining the sampled virtual point clouds and the original point cloud data to obtain new point cloud data; dividing the point cloud space formed by the new point cloud data into a plurality of voxel units, and extracting a preset number of new point cloud data in each voxel unit; calculating the relative coordinates of the extracted new point cloud data and the centroid of the corresponding voxel unit, and based on the relative coordinates, obtaining an extended feature vector; using a constructed feature extraction network to perform feature extraction on the extended feature vector to obtain voxel features; the feature extraction network including a plurality of modules composed of fully connected layers and max pooling layers.
[0008] Optionally, dividing each candidate box at the same resolution to obtain multiple sub-regions, and calculating local features of each sub-region, including: dividing each candidate box to obtain multiple sub-regions respectively, and determining the central pixel position of the sub-region; aggregating multi-scale neighborhood features of the central pixel position of the sub-region by using a multi-layer perceptron, and performing max pooling along the channel on the aggregation result to obtain local features of each sub-region at each resolution.
[0009] Optionally, fusing the spatial information and semantic information of local features with different resolutions by using a second network to obtain candidate box feature information, including: performing a first operation on the local features of the first resolution and the local features of the second resolution to obtain first fusion data; the first operation includes: compressing the channels of the two input data to obtain third data; normalizing the third data and splitting and calculating to obtain corresponding first weights and second weights; based on the first weights and second weights, performing weighted summation on the two input data to obtain fusion data; performing a first operation on the first fusion data and the local features of the third resolution to obtain second fusion data; splicing the local features of all resolutions and fusing them with the second fusion data to obtain candidate box feature information.
[0010] Optionally, processing the image data and the original point cloud data by using a PENet depth completion network to generate a virtual point cloud, and its expression is: where D f(u,v) represents the pixel depth of the fused virtual point cloud, C cd (u, v) represents the pixel coordinates of the confidence map of the color main branch, e represents a constant, D cd(u,v) represents the pixel coordinates of the depth map of the color main branch, C dd (u, v) represents the pixel coordinates of the confidence map of the depth main branch, D dd(u,v) represents the pixel coordinates of the depth map of the depth main branch.
[0011] Optionally, obtaining an extended feature vector based on relative coordinates, and its expression is: where V′ represents the extended feature vector, represents the i-th point cloud data after extension, x i represents the x-axis coordinate of the i-th point cloud data, y i represents the y-axis coordinate of the i-th point cloud data, z i represents the z-axis coordinate of the i-th point cloud data, r i represents the intensity value of the i-th point cloud data, represents the x-axis coordinate of the centroid, represents the y-axis coordinate of the centroid, represents the z-axis coordinate of the centroid.
[0012] Optionally, the local features of all resolutions are stitched together and fused with the second fusion data to obtain candidate box feature information, and its expression is: η = MLP[MLP(K 3D ([η 2 , η 4 , η 8 )), MLP(K 3D (η 2,4,8 ))], where η represents the candidate box feature information, MLP() represents a multi-layer perceptron network, K 3D represents a standard 3D convolution, η 2 represents the local feature of the first resolution, η 4 represents the local feature of the second resolution, η 8 represents the local feature of the third resolution, η 2,4,8 represents the second fusion data.
[0013] An object detection system based on multi-scale fusion, the system includes: an acquisition unit, which acquires the image features corresponding to the image data and the voxel features corresponding to the point cloud data; a three-dimensional feature extraction unit, which uses a first network to process the image features and voxel features to obtain cross-space fusion features, and the first network includes a first module, a second module and a third module; the first module includes determining the neighborhood features around the voxel features; respectively performing convolution processing on the data projected from the neighborhood features to the image data and the image features, and stitching the processing results of the two to obtain fusion features; the second module includes performing convolution on the fusion features, stitching the convolution results and the image features in the feature channels, and performing convolution processing on the stitched results; respectively performing average pooling and max pooling on the convolved stitched results, and stitching the two pooling results to obtain a feature descriptor; performing convolution processing on the feature descriptor to obtain a first spatial attention; multiplying the first spatial attention and the convolved stitched results to obtain a noise-aware feature; the third module includes stitching and fusing the voxel features and the noise-aware features and performing downsampling to obtain multiple cross-space fusion features with different resolutions; a two-dimensional feature extraction unit, which performs bird's-eye view transformation on the multiple cross-space fusion features with different resolutions, and performs feature extraction on the transformed bird's-eye view to obtain corresponding bird's-eye features with different resolutions; a candidate box extraction unit, which uses a region proposal network to process each bird's-eye feature to obtain multiple candidate boxes respectively; a local feature extraction unit, which calculates the local features at different resolutions, including dividing each candidate box at the same resolution to obtain multiple sub-regions, and calculating the local features of each sub-region; a candidate box feature extraction unit, which uses a second network to fuse the spatial information and semantic information of the local features corresponding to the cross-space fusion features with different resolutions to obtain candidate box feature information for detecting objects.
[0014] Optionally, the candidate box feature extraction unit includes: performing a first operation on the local features of the first resolution and the local features of the second resolution to obtain first fusion data; the first operation includes: compressing the channels of the two input data to obtain third data; normalizing the third data and splitting and calculating to obtain corresponding first weights and second weights; based on the first weights and the second weights, performing weighted summation on the two input data to obtain fusion data; performing a first operation on the first fusion data and the local features of the third resolution to obtain second fusion data; splicing the local features of all resolutions and fusing them with the second fusion data to obtain candidate box feature information.
[0015] Optionally, the local feature extraction unit includes: dividing each candidate box into multiple sub-regions respectively, and determining the central pixel positions of the sub-regions; aggregating the multi-scale neighborhood features of the central pixel positions of the sub-regions by using a multi-layer perceptron, and performing max pooling along the channels on the aggregation result to obtain the local features of each sub-region of each resolution.
[0016] Compared with the prior art, a target detection method and system based on multi-scale fusion provided by the present invention integrates the advantages of data-level fusion and feature-level fusion. Specifically, data-level fusion is achieved by generating virtual point clouds, and feature-level fusion is achieved by a first network and a second network. The first network reduces the noise problem, and depth interactive feature enhancement is achieved through cross-modal fusion, noise sensing, and cross-space fusion. The second network improves the bounding boxes by adaptively fusing cross-layer spatial semantic information, thereby improving the target detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of this specification, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0018] Figure 1 is a flowchart of a target detection method based on multi-scale fusion provided in this embodiment;
[0019] Figure 2 is a schematic diagram of data processing of a target detection method based on multi-scale fusion provided in this embodiment;
[0020] Figure 3 is a schematic diagram of data processing of a method for calculating candidate box feature information provided in this embodiment;
[0021] Figure 4It is a schematic diagram of data processing of a three-dimensional feature extraction module in an object detection system based on multi-scale fusion provided by this embodiment;
[0022] Figure 5 It is a schematic diagram of refined data processing of a three-dimensional feature extraction module in an object detection system based on multi-scale fusion provided by this embodiment. Specific implementation manner
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] For ease of understanding of the embodiments of the present invention, the following will further explain with specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the protection scope of the present invention.
[0025] Embodiment 1
[0026] This embodiment provides an object detection method based on multi-scale fusion, and its process is as Figure 1 and Figure 2 shown, including:
[0027] S1 Obtain the image features corresponding to the image data and the voxel features corresponding to the point cloud data.
[0028] In this step, it includes:
[0029] S100 Obtain the image data, perform feature extraction on the image data, and obtain image features.
[0030] In this embodiment, standard 2D convolution is performed on the image data to obtain the feature map X'.
[0031] S110 Obtain the original point cloud data, and use the PENet depth completion network to process the image data and the original point cloud data to generate virtual point clouds.
[0032] In this embodiment, the PENet depth completion network is used to realize depth completion of the original point cloud guided by high-resolution color image data, and generate a dense virtual point cloud from sparse original lidar point cloud data. The PENet depth completion network includes a color main branch and a depth main branch. The color main branch constructs an encoder-decoder network to extract color-dominated features for depth prediction, so that the depth around the object boundary can be learned through the structural information in the color image. The depth main branch constructs an encoder-decoder network similar to the color main branch, aiming to predict a dense depth map by upsampling a sparse depth map. The PENet network also follows the strategy of FusionNet to fuse the depths from the two branches based on confidence information. The depth maps output by the color main branch and the depth main branch are D cd and D dd , and the confidence maps are C cd and C dd , respectively. Then the depth of the fused virtual point cloud is:
[0033]
[0034] where D f(u,v) represents the depth of the pixel point of the fused virtual point cloud, C cd (u, v) represents the pixel coordinates of the confidence map of the color main branch, e represents a constant, D cd(u,v) represents the pixel coordinates of the depth map of the color main branch, C dd (u, v) represents the pixel coordinates of the confidence map of the depth main branch, and D dd(u,v) represents the pixel coordinates of the depth map of the depth main branch.
[0035] S120 performs hierarchical downsampling on the virtual point cloud according to the distance between the virtual point cloud and the lidar center; combines the sampled virtual point cloud and the original point cloud data to obtain new point cloud data.
[0036] In this step, the purpose of using the virtual point cloud for 3D object detection is to improve the density of the point cloud. The original point cloud data that is dense enough at close range can meet the requirements of the detection model, while the original point cloud data at long range is sparse, which limits the improvement of detection accuracy. In order to enhance the contour representation of distant objects, this embodiment retains fewer virtual point clouds closer to the lidar center and more virtual point clouds farther from the lidar center. This can not only improve the detection accuracy but also reduce the redundant calculation amount.
[0037] Specifically, calculate the horizontal distance of all virtual point clouds from the center of the lidar. Group all virtual point clouds at equal intervals according to the distance. In this embodiment, it is divided into 10 groups, namely [0, 10), [10, 20)... [90, 100), with the unit of meters. Within different groups, retain virtual point clouds according to the set ratio. For example, discard all virtual point clouds within the range of [0, 30), and retain all virtual point clouds within the range of [70, 100). After hierarchical downsampling, the processed virtual point clouds are obtained. Merge the processed virtual point clouds and the original point cloud data to obtain new point cloud data. Through the above processing, the problem of sparse point clouds at long distances is fundamentally solved, the recognition of target contours at long distances can be enhanced, the model can be helped to extract more target feature information, and the detection accuracy can be assisted to improve.
[0038] The new point cloud data and the image data are used as the input data of the first network. The new point cloud data includes 3D spatial coordinate values and intensity values, and the intensity values are derived from the interpolation calculation of the neighborhood original point clouds. The image information includes pixel coordinate values and color channel values.
[0039] S130 divides the point cloud space formed by the new point cloud data into multiple voxel units, and extracts a preset number of new point cloud data in each voxel unit; calculates the relative coordinates of the extracted new point cloud data and the centroid of the corresponding voxel unit, and based on the relative coordinates, obtains an extended feature vector.
[0040] Divide the point cloud space into stacked voxel units of the same size. To reduce the computational complexity, randomly sample a point cloud set with a maximum number of T within each voxel unit. The voxel unit contains non-empty point cloud data with t ≤ T point clouds, denoted as where V represents the voxel unit, p i represents the i-th point cloud data, x i represents the x-axis coordinate of the i-th point cloud data, y i represents the y-axis coordinate of the i-th point cloud data, z i represents the z-axis coordinate of the i-th point cloud data, r i represents the intensity value of the i-th point cloud data, and the centroid coordinates of the point cloud within the voxel are denoted as Calculate the relative coordinates of each point cloud and the centroid, and expand the feature vector of the voxel unit to:
[0041]
[0042] where V′ represents the extended feature vector, represents the i-th point cloud data after expansion, represents the x-axis coordinate of the centroid, represents the y-axis coordinate of the centroid, represents the z-axis coordinate of the centroid.
[0043] S140 uses the constructed feature extraction network to extract features from the augmented feature vectors to obtain voxel features; the feature extraction network includes multiple modules composed of fully connected layers and max pooling layers.
[0044] Each module of the feature extraction network includes: constructing a fully connected network (FCN) to extract the point-wise features f of a single point cloud within a non-empty voxel i , and the FCN is composed of a linear layer, a normalization layer, and a non-linear layer. After obtaining the point-wise features, perform max pooling on all f of the voxel i to obtain the aggregated feature Stack the point-wise features and the aggregated features of the point cloud to obtain the output features as the input data of the fully connected layer of the next module. Multiple modules are connected to gradually increase the dimension of the feature vector, used to extract more advanced depth information, and use the max pooling aggregated feature of the last layer as the input feature of the first network. Let the number of non-empty voxels be N, then the non-empty voxel feature vector set is denoted as X.
[0045] S2 uses the first network to process the image features and voxel features to obtain cross-space fusion features.
[0046] This step fuses voxel features and image features in the 2D space, performs noise perception based on the attention mechanism, and fuses 3D space features and 2D space features. As Figure 3 shown, the first network takes voxel features and image features as inputs and performs multi-level feature encoding and feature fusion with downsampling strides of 1×, 2×, 4×, and 8×.
[0047] Perform geometric feature encoding in the 3D voxel space and 2D image space respectively to obtain and where is the set of real numbers, N represents the size of the image, C in , C mid , C out represent the number of input, intermediate, and output feature channels respectively, and Perform feature encoding in the 2D image space, and splice and in the 2D space to obtain the fused feature Further, obtain the noise perception feature with the help of the attention mechanism; finally, fuse the 3D sparse feature and the 2D sparse noise perception feature across space to obtain the cross-space fusion feature As Figure 4As shown, the first module is a cross-modal feature fusion module, the second module is an attention noise perception module, and the third module is a cross-space feature fusion module, specifically including:
[0048] S200 uses the first module to perform cross-modal feature fusion to obtain fused features.
[0049] Step 1: Determine the neighborhood features around the voxel features.
[0050] Perform sparse convolution on the voxel features to obtain neighborhood features Specifically, for each voxel feature X in X i , based on the corresponding 3D index H i Determine the feature data within the k×k×k spatial neighborhood, and encode the feature data within the neighborhood through the sparse 3D volume K 3D () to obtain neighborhood features The specific calculation is as follows:
[0051]
[0052] Where represents the feature data within the 3D spatial neighborhood of the i-th voxel feature, and Φ represents the non-linear activation function.
[0053] Step 2: Respectively perform convolution processing on the data and image features obtained by projecting the neighborhood features onto the image data, and splice the processing results of the two to obtain fused features.
[0054] Project the neighborhood features into the 2D image space, and use the sparse 2D convolution K 2D Encode through the projected neighborhood features to obtain Specifically, in this embodiment, the neighborhood features are transformed from the 3D lidar coordinate system space to the 2D image space, and the specific formula is as follows:
[0055]
[0056] Where, represents the index vector of the voxel in the 2D image space, P is the projection function, and H represents the index vector of the neighborhood features.
[0057] For each neighborhood feature Based on the corresponding 2D index Further encode the features of the non-empty voxels within the k×k neighborhood to obtain the first data Its expression is:
[0058]
[0059] Where The 2D spatial neighborhood data representing the i-th neighborhood feature, and Φ represents the non-linear activation function.
[0060] In the 2D image space, perform a standard 2D convolution on the image feature X′ to obtain the second data X″.
[0061] In this embodiment, the first data and the second data are concatenated and fused in the 2D space to obtain the 2D space fusion feature Y.
[0062] S210 uses the second module to perform noise perception based on the attention mechanism to obtain the noise perception feature.
[0063] Step 1: Convolve the fusion feature, splice the convolution result and the image feature in the feature channels, and perform convolution processing on the splicing result.
[0064] The method of generating virtual point clouds based on depth estimation will cause noise in the depth values, and this noise will significantly reduce the detection performance. In this embodiment, a 2D convolution module based on the spatial attention mechanism is designed to encode the noise features to perceive the noise, thereby removing the noise.
[0065] Use the sparse 2D convolution K 2D Encode the fusion feature, then splice the encoded fusion feature and the image feature in the feature channels, and then perform a standard 2D convolution on the splicing result to obtain the convolved splicing result Y′. The specific formula is as follows:
[0066] Y′ = Φ(K 2D [Φ(K 2D (Y)), X′])
[0067] Step 2: Perform average pooling and max pooling on the convolved splicing result respectively, and splice the two pooling results to obtain the feature descriptor.
[0068] Perform average pooling and max pooling on the convolved splicing result along the channel dimension respectively to obtain and Splice the two to generate the feature descriptor.
[0069] Step 3: Perform convolution processing on the feature descriptor to obtain the first spatial attention; multiply the first spatial attention and the convolved splicing result to obtain the noise perception feature.
[0070] Use a 2D convolutional layer to generate the first spatial attention map for the spliced feature descriptor The first spatial attention map encodes which regions are highlighted or suppressed, that is, it encodes the degree to which the noise in different voxel regions is suppressed. Its expression is:
[0071]
[0072] Among them, σ represents the sigmoid function, and K 2D represents a standard 2D convolution.
[0073] Multiply the first spatial attention map by the concatenated result after convolution to obtain the noise-aware feature The specific formula is as follows:
[0074] Y" = M s (Y) × Y′
[0075] S220 uses the third module to perform cross-space feature fusion to obtain cross-space fusion features.
[0076] According to the mapping relationship, the voxel features and the noise-aware features are concatenated and fused to obtain multiple cross-space output features with different resolutions Its expression is Then perform downsampling to obtain 4 groups of cross-space fusion features with different resolutions {Z 1X , Z 2X , Z 4X , Z 8X}, and the corresponding spatial indices are {H 1X , H 2X , H 4X , H 8X}.
[0077] S3 performs an aerial view transformation on the cross-space fusion features with multiple different resolutions, and extracts features from the transformed aerial view to obtain the corresponding aerial view features with different resolutions.
[0078] S4 uses the region proposal network to process each aerial view feature to obtain multiple candidate boxes respectively.
[0079] Training the region proposal network includes:
[0080] The region proposal network is a dense detection head. Slide a small window on the aerial view feature map to generate a series of bounding boxes with fixed sizes and aspect ratios; each bounding box is assigned a score to indicate the possibility that the region contains the target object; in addition, perform bounding box regression on each bounding box to correct its position and size.
[0081] Specifically, obtain the intersection over union (IoU) according to the overlap relationship between the ground truth box and the bounding box, and match the bounding box with the largest IoU. The ground truth box is the actual region where the detection target is located. Set the bounding boxes with IoU greater than the foreground threshold as positive samples, and the bounding boxes with IoU less than the background threshold as negative samples. The calculation formula of IoU is as follows:
[0082]
[0083] Among them, area(a) represents the spatial area of the ground truth box, area(b) represents the spatial area of the bounding box, ∩ represents the spatial intersection of different objects, and ∪ represents the spatial union of different objects.
[0084] The sigmoid activation function is used to screen positive and negative samples for the prediction of the classification branch; and the regression deviation between the positive samples and the ground truth box is calculated through the regression function, which is used for the regression branch to predict the translation and scaling parameters.
[0085] In the region proposal network, each bird's-eye view feature is input, and the non-maximum suppression algorithm is used to sort and screen the bounding boxes, removing regions with high overlap and bounding boxes with low confidence to obtain candidate boxes; the candidate boxes are input into the candidate box detection head for more refined bounding box detection. The candidate box detection head includes the following steps S5-S7.
[0086] S5 Calculate the local features at different resolutions, including dividing each candidate box at the same resolution to obtain multiple sub-regions, and calculating the local features of each sub-region.
[0087] In this step, it includes:
[0088] Step 1: Divide each candidate box to obtain multiple sub-regions respectively, and determine the central pixel position of the sub-regions.
[0089] The candidate box is divided into G×G×G sub-regions. For the α-th sub-region, in this embodiment, the central spatial coordinates of each sub-region are extracted as the query center g α .
[0090] Step 2: Use a multi-layer perceptron to aggregate the multi-scale neighborhood features of the central pixel positions of the sub-regions, and perform max pooling along the channels on the aggregation result to obtain the local features of each sub-region at each resolution
[0091] Its expression is:
[0092]
[0093] Among them, in this embodiment, features are extracted from the cross-space fusion features of {Z 2X , Z 4X , Z 8X}, β represents feature query and aggregation for the cross-space fusion feature of the β-th resolution, γ represents the index of the non-empty voxels queried in the current feature layer, represents the relative spatial coordinates, represents the spatial index as The cross - space fusion feature, MLP() represents a multi - layer perceptron network, and MaxPool() represents max - pooling.
[0094] S6 utilizes the second network to fuse the spatial information and semantic information of local features with different resolutions to obtain candidate box feature information for detecting the target.
[0095] Local features at different resolutions map different information. Local features with high resolution and low channel values contain more geometric and spatial information, while local features with low resolution and high channel values include more category and semantic information. To fully fuse spatial information and semantic information, this embodiment designs a progressive adaptive aggregation module, as Figure 5 shown.
[0096] Specifically, it includes:
[0097] S600 performs a first operation on the local features of the first resolution and the local features of the second resolution to obtain first fusion data.
[0098] The first operation includes: compressing the channels of the two input data to obtain third data; normalizing and splitting the calculation of the third data to obtain corresponding first weight and second weight; based on the first weight and the second weight, performing weighted summation on the two input data to obtain fusion data.
[0099] In this embodiment, the channels of local features of two adjacent resolutions are compressed into one:
[0100] δ β =K 3D (η β ),δ β+1 =K 3D (η β+1 )
[0101] Among them, δ β represents the concatenated channel of η β , δ β+1 represents the concatenated channel of η β+1 . The Softmax function is used to normalize the two concatenated channels and split them into two bird's - eye view attention maps, and the first weight and the second weight are obtained from the learned bird's - eye view attention maps. These weighted features are fused by element - wise addition. The specific expression: η β,β+1 =ω β ×η β +ω β+1 ×η β+1 , η β,β+1 represents the fusion data.
[0102] Specifically, the local feature with a resolution of 2X and the local feature with a resolution of 4X are fused to obtain the first fusion data η 2,4 .
[0103] S610 performs a first operation on the first fusion data and the local feature of the third resolution to obtain the second fusion data.
[0104] Continue to fuse the first fusion data η 2,4 with the local feature with a resolution of 8X to obtain the second fusion data η 2,4,8 .
[0105] S620 splices the local features of all resolutions and fuses them with the second fusion data to obtain the candidate box feature information.
[0106] In this embodiment, all the local features of all resolutions are spliced to obtain the fourth data, and the standard 3D convolution K 3D and the multi-layer perceptron are used to process the fourth data to obtain the fifth data; the fifth data and the processing result of the second fusion data by the multi-layer perceptron are spliced in the channel direction to obtain the final candidate box feature information, and its expression is:
[0107] η = MLP[MLP(K 3D ([η 2 , η 4 , η 8 )), MLP(K 3D (η 2,4,8 ))]
[0108] where η represents the candidate box feature information, MLP() represents the multi-layer perceptron network, K 3D represents the standard 3D convolution, η 2 represents the local feature of the first resolution, η 4 represents the local feature of the second resolution, η 8 represents the local feature of the third resolution, η 2,4,8 represents the second fusion data.
[0109] The adaptive fusion module helps to extract more robust features with rich spatial and semantic information, and can detect the bounding box more accurately.
[0110] Based on S1 - S6, a target detection model is constructed; the target detection model is trained, and the trained target detection model is used to identify the input point cloud data and image data to obtain the target detection box.
[0111] The loss function of the object detection model in this embodiment includes the loss of the dense detection head, i.e., the RPN loss, and the loss of the candidate box detection head, i.e., the RoI loss. The RPN loss consists of the positive and negative sample classification loss, the detection box regression loss, and the detection box direction loss. The RoI loss consists of the detection box confidence loss, the detection box regression loss, and the detection box direction loss.
[0112] RPN loss
[0113] Define the positive and negative sample classification loss. Since there are usually only a few true detection boxes in each frame of data, this leads to an extremely unbalanced distribution of positive and negative samples. To solve this problem, this embodiment introduces the Focal Loss function. Focal Loss addresses the problem of uneven distribution of positive and negative samples in object detection. It introduces an adjustable factor to balance the weights of different samples, enabling the model to focus more on key samples, thereby improving the overall detection performance of the model.
[0114] RoI loss
[0115] Define the confidence loss. Recalculate the confidence ground truth based on IoU. The RoI Head loss value is a weighted combination of the confidence loss, the regression loss, and the direction loss.
[0116] Total loss
[0117] The total loss consists of the RPN loss and the RoI loss, and the expression is: L ToTAL =L RPN +L RoI where L ToTAL represents the total loss, L RPN represents the RPN loss, and L RoI represents the ROI loss.
[0118] Compared with the prior art, an object detection method based on multi-scale fusion provided in this embodiment integrates the advantages of data-level fusion and feature-level fusion. Specifically, data-level fusion is achieved by generating virtual point clouds, and feature-level fusion is achieved through the first network and the second network. The first network reduces the noise problem and realizes deep interactive feature enhancement through cross-modal fusion, noise sensing, and cross-space fusion. The second network improves the bounding box by adaptively fusing cross-layer spatial semantic information, enhancing the object detection accuracy.
[0119] Embodiment 2
[0120] This embodiment provides an object detection system based on multi-scale fusion to implement the object detection method based on multi-scale fusion described in Embodiment 1, including:
[0121] An acquisition unit that acquires the image features corresponding to the image data and the voxel features corresponding to the point cloud data;
[0122] A three-dimensional feature extraction unit processes image features and voxel features using a first network to obtain cross-space fusion features. The first network includes a first module, a second module, and a third module. The first module includes determining neighborhood features around voxel features; respectively performing convolutional processing on the data projected from the neighborhood features into the image data and the image features, and concatenating the processing results of the two to obtain a fusion feature. The second module includes performing convolution on the fusion feature, concatenating the convolution result and the image features in the feature channels, and performing convolutional processing on the concatenated result; respectively performing average pooling and max pooling on the concatenated result after convolution, and concatenating the two pooling results to obtain a feature descriptor; performing convolutional processing on the feature descriptor to obtain a first spatial attention; multiplying the first spatial attention and the concatenated result after convolution to obtain a noise-aware feature. The third module includes concatenating and fusing the voxel feature and the noise-aware feature and performing downsampling to obtain multiple cross-space fusion features with different resolutions.
[0123] A two-dimensional feature extraction unit performs an aerial view transformation on multiple cross-space fusion features with different resolutions, and extracts features from the transformed aerial view images to obtain corresponding aerial view features with different resolutions.
[0124] A candidate box extraction unit processes each aerial view feature using a region proposal network to respectively obtain multiple candidate boxes.
[0125] A local feature extraction unit calculates local features at different resolutions, including dividing each candidate box at the same resolution to obtain multiple sub-regions, and calculating the local features of each sub-region.
[0126] A candidate box feature extraction unit fuses the spatial information and semantic information of the local features corresponding to the cross-space fusion features with different resolutions using a second network to obtain candidate box feature information for detecting targets.
[0127] Optionally, the candidate box feature extraction unit includes:
[0128] Performing a first operation on the local features of the first resolution and the local features of the second resolution to obtain first fusion data. The first operation includes: compressing the channels of the two input data to obtain third data; normalizing the third data and splitting and calculating to obtain corresponding first weights and second weights; based on the first weights and second weights, performing weighted summation on the two input data to obtain fusion data.
[0129] Performing a first operation on the first fusion data and the local features of the third resolution to obtain second fusion data.
[0130] Stitch the local features of all resolutions and fuse them with the second fusion data to obtain candidate box feature information.
[0131] Optionally, the local feature extraction unit includes:
[0132] Divide each candidate box into multiple sub-regions and determine the central pixel positions of the sub-regions;
[0133] Use a multi-layer perceptron to aggregate the multi-scale neighborhood features of the central pixel positions of the sub-regions, and perform max pooling on the aggregation results along the channels to obtain the local features of each sub-region at each resolution.
[0134] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A target detection method based on multi-scale fusion, characterized in that, Including: Obtain the image features corresponding to the image data and the voxel features corresponding to the point cloud data; Use a first network to process the image features and voxel features to obtain cross-space fusion features. The first network includes a first module, a second module, and a third module. The first module includes determining the neighborhood features around the voxel features; respectively performing convolutional processing on the data projected from the neighborhood features into the image data and the image features, and splicing the processing results of the two to obtain a fusion feature. The second module includes performing convolution on the fusion feature, splicing the convolution result and the image feature in the feature channels, and performing convolutional processing on the splicing result; respectively performing average pooling and max pooling on the convolved splicing result, and splicing the two pooling results to obtain a feature descriptor; performing convolutional processing on the feature descriptor to obtain a first spatial attention; multiplying the first spatial attention and the convolved splicing result to obtain a noise-aware feature. The third module includes splicing and fusing the voxel feature and the noise-aware feature and performing downsampling to obtain cross-space fusion features with multiple different resolutions; Perform a bird's-eye view visual transformation on the cross-space fusion features with multiple different resolutions, and perform feature extraction on the transformed bird's-eye view to obtain corresponding bird's-eye features with different resolutions; Use a region proposal network to process each bird's-eye feature to respectively obtain a plurality of candidate boxes; Calculate the local features at different resolutions, including dividing each candidate box at the same resolution to obtain a plurality of sub-regions, and calculating the local features of each sub-region; Use a second network to fuse the spatial information and semantic information of the local features with different resolutions to obtain candidate box feature information for detecting the target.
2. The object detection method based on multi-scale fusion according to claim 1, wherein, The obtaining the image features corresponding to the image data and the voxel features corresponding to the point cloud data includes: Obtain the image data, perform feature extraction on the image data to obtain image features; Obtain the original point cloud data, and use the PENet depth completion network to process the image data and the original point cloud data to generate a virtual point cloud; Perform hierarchical downsampling on the virtual point cloud according to the distance between the virtual point cloud and the lidar center; combine the sampled virtual point cloud and the original point cloud data to obtain new point cloud data; Divide the point cloud space formed by the new point cloud data into a plurality of voxel units, and extract a preset number of new point cloud data in each voxel unit; calculate the relative coordinates of the extracted new point cloud data and the centroid of the corresponding voxel unit, and based on the relative coordinates, obtain an extended feature vector; Use the constructed feature extraction network to perform feature extraction on the extended feature vector to obtain voxel features. The feature extraction network includes a plurality of modules composed of fully connected layers and max pooling layers.
3. A multi-scale fusion-based object detection method according to claim 1, characterized in that The dividing each candidate box at the same resolution to obtain a plurality of sub-regions and calculating the local features of each sub-region includes: Divide each candidate box to respectively obtain a plurality of sub-regions, and determine the central pixel position of the sub-region; Use a multi-layer perceptron to aggregate the multi-scale neighborhood features of the central pixel position of the sub-region, and perform max pooling on the aggregation result along the channel to obtain the local features of each sub-region at each resolution.
4. A multi-scale fusion-based object detection method according to claim 1, characterized in that The spatial information and semantic information of local features with different resolutions are fused using the second network to obtain candidate box feature information, including: Performing a first operation on the local features of the first resolution and the local features of the second resolution to obtain first fusion data; the first operation includes: compressing the channels of the two input data to obtain third data; normalizing the third data and splitting and calculating to obtain corresponding first weights and second weights; based on the first weights and the second weights, performing weighted summation on the two input data to obtain fusion data; Performing a first operation on the first fusion data and the local features of the third resolution to obtain second fusion data; Concatenating the local features of all resolutions and fusing them with the second fusion data to obtain candidate box feature information.
5. The object detection method based on multi-scale fusion according to claim 2, wherein The PENet depth completion network is used to process the image data and the original point cloud data to generate a virtual point cloud, and its expression is: Among them, D f (u, v) represents the pixel depth of the fused virtual point cloud, C cd (u, v) represents the pixel coordinates of the confidence map of the color main branch, e represents a constant, D cd(u,v) represents the pixel coordinates of the depth map of the color main branch, C dd (u, v) represents the pixel coordinates of the confidence map of the depth main branch, D dd(u,v) represents the pixel coordinates of the depth map of the depth main branch.
6. A target detection method based on multi-scale fusion according to claim 2, characterized in that, Based on the relative coordinates, an extended feature vector is obtained, and its expression is: Among them, V′ represents the extended feature vector, represents the i-th point cloud data after extension, x i represents the x-axis coordinate of the i-th point cloud data, y i represents the y-axis coordinate of the i-th point cloud data, Z i represents the z-axis coordinate of the i-th point cloud data, r i represents the intensity value of the i-th point cloud data, represents the x-axis coordinate of the centroid, represents the y-axis coordinate of the centroid, represents the z-axis coordinate of the centroid.
7. The object detection method based on multi-scale fusion according to claim 4, wherein The expression for concatenating the local features of all resolutions and fusing them with the second fusion data to obtain candidate box feature information is: η = MLP[MLP(K 3D ([η 2 ,η 4 ,η 8 )),MLP(K 3D (η 2,4,8 ))] Among them, η represents the candidate box feature information, MLP() represents the multi-layer perceptron network, and K 3D represents the standard 3D convolution, and η 2 represents the local feature of the first resolution, and η 4 represents the local feature of the second resolution, and η 8 represents the local feature of the third resolution, and η 2,4,8 represents the second fusion data.
8. An object detection system based on multi-scale fusion, characterized in that, Including: An acquisition unit that acquires the image features corresponding to the image data and the voxel features corresponding to the point cloud data; A three-dimensional feature extraction unit that uses the first network to process the image features and the voxel features to obtain cross-space fusion features. The first network includes a first module, a second module, and a third module; the first module includes determining the neighborhood features around the voxel features; respectively performing convolution processing on the data projected from the neighborhood features to the image data and the image features, and concatenating the processing results of the two to obtain fusion features; the second module includes performing convolution on the fusion features, concatenating the convolution results and the image features in the feature channels, and performing convolution processing on the concatenated results; respectively performing average pooling and max pooling on the concatenated results after convolution, and concatenating the two pooling results to obtain a feature descriptor; performing convolution processing on the feature descriptor to obtain a first spatial attention; multiplying the first spatial attention and the concatenated results after convolution to obtain a noise-aware feature; the third module includes concatenating and fusing the voxel features and the noise-aware features and performing downsampling to obtain multiple cross-space fusion features with different resolutions; A two-dimensional feature extraction unit that performs bird's-eye view transformation on the cross-space fusion features with multiple different resolutions and extracts features from the transformed bird's-eye view to obtain corresponding bird's-eye features with different resolutions; A candidate box extraction unit that uses a region proposal network to process each bird's-eye feature to obtain multiple candidate boxes respectively; A local feature extraction unit that calculates local features at different resolutions, including dividing each candidate box at the same resolution to obtain multiple sub-regions, and calculating the local features of each sub-region; A candidate box feature extraction unit that uses the second network to fuse the spatial information and semantic information of the local features corresponding to the cross-space fusion features with different resolutions to obtain candidate box feature information for detecting targets.
9. The object detection system based on multi-scale fusion according to claim 8, wherein, The candidate box feature extraction unit includes: Perform a first operation on the local features of the first resolution and the local features of the second resolution to obtain first fusion data; the first operation includes: compressing the channels of the two input data to obtain third data; normalizing the third data and splitting and calculating to obtain corresponding first weights and second weights; based on the first weights and the second weights, performing weighted summation on the two input data to obtain fusion data; Perform a first operation on the first fusion data and the local features of the third resolution to obtain second fusion data; Stitch the local features of all resolutions and fuse them with the second fusion data to obtain candidate box feature information.
10. A target detection system based on multi-scale fusion according to claim 8, characterized in that, The local feature extraction unit includes: Divide each candidate box into multiple sub-regions respectively and determine the central pixel positions of the sub-regions; Aggregate the multi-scale neighborhood features of the central pixel positions of the sub-regions using a multi-layer perceptron, and perform max pooling on the aggregation result along the channels to obtain the local features of each sub-region at each resolution.
Citation Information
Patent Citations
Camera and millimeter wave radar front fusion road surface target detection method
CN116259024A
Multimodal 3D target detection method based on semantic propagation and cross-attention mechanism
CN117173655A