3D target detection method based on multi-source multi-representation information fusion
By fusing the voxel representation of 3D point cloud data with the distance image representation, and using the semantic information of image features to enhance the point cloud features, the difficulty of sparse and irregular data processing in 3D object detection is solved, and the detection accuracy is improved, especially when identifying small objects and long-distance objects.
Patent Information
- Application Number
- CN202510051371.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
3D object detection has problems with quantization losses and feature extraction when processing sparse and irregular 3D point cloud data, especially when identifying small objects and long-distance objects, it is difficult to accurately detect.
The 3D object detection method based on multi-source multi-representation information fusion is adopted to fuse the voxel representation of point cloud data with the distance image representation, and the point cloud features are enhanced through the semantic information of image features to reduce information loss and improve detection accuracy.
By combining voxel representation and distance image representation, the information loss in point cloud data processing is reduced, the feature representation ability is improved, and the accuracy of 3D object detection is enhanced, especially when identifying small objects and long-distance objects.
Smart Images

Figure CN119992534A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence deep learning and intelligent driving perception, and specifically relates to a 3D target detection method based on multi-source and multi-representation information fusion. Background Art
[0002] As a basic task of 3D scene perception, 3D object detection plays a vital role in the development of fields such as intelligent driving and mobile robot perception. In a 3D perception system, 3D object detection uses devices such as laser radar and optical cameras to obtain information about the surrounding environment of the vehicle or robot, identify and locate surrounding objects, and provide key information for downstream tasks. With the reduction in the cost of laser scanning and depth acquisition hardware equipment and the increasing development of software technology, laser radar sensors have become the mainstream sensors in 3D perception systems with their advantages of high precision, high frequency and high resolution. Laser radar obtains all-round information about the surrounding scene by emitting laser beams to the environment and receiving laser beams reflected back by objects. The high-precision three-dimensional point cloud map it generates retains the original geometric information in the 3D space and has powerful three-dimensional representation capabilities.
[0003] Due to the working principle of LiDAR sensors and other objective conditions such as occlusion and surface material of objects, 3D point cloud data is sparse and irregular, which means that 3D target detection still faces many challenges. The original 3D point cloud data is large in volume and sparsely distributed. Direct processing requires a lot of downsampling, which can easily lead to the loss of target points and interference from background and noise. Before being sent to the feature extractor, the irregular point cloud data needs to be quantized through a series of regularization operations such as voxelization or projection. In this process, the detailed position and geometric information of the original point cloud will inevitably be lost, resulting in quantization loss.
[0004] In order to improve the shortcomings of 3D target detection methods using a single modality and a single data representation, the present invention proposes a 3D target detection method based on multi-source and multi-representation information fusion, which utilizes the advantages of various representation forms of three-dimensional point clouds, reduces the loss in the process of data processing and feature extraction, and fuses image features to supplement the semantic information of the target, thereby reducing the impact of point cloud sparsity. Summary of the invention
[0005] To solve the above technical problems, the present invention provides a 3D object detection method based on multi-source multi-representation information fusion. The method fuses the voxel representation of point cloud data with the range image representation, while taking advantage of the advantages of different data representation forms, and enhances the point cloud features by fusing the semantic information of the image to improve the 3D object detection accuracy.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A 3D object detection method based on multi-source multi-representation information fusion, the method comprising the following steps:
[0008] Step 1: Voxelize the input 3D point cloud data, encode the voxel features, aggregate the information of the points contained in each voxel, and obtain the corresponding voxel features;
[0009] Step 2: Use the DeepLabv3 model as the 2D backbone network to extract features from the image data and obtain image features;
[0010] Step 3: An image feature indexing method is proposed. For each voxel feature, the voxel coordinates are mapped to the image, and its corresponding image features are indexed. The voxel features and image features of each voxel are fused to obtain multimodal features.
[0011] Step 4: Use a 3D sparse convolutional network to extract deep features from multimodal features and transform the deep features into bird’s-eye view (BEV) features;
[0012] Step 5: Convert the input 3D point cloud data into a range view representation through range projection, and propose a frequency-aware adaptive dilated convolutional network to extract multi-scale features from the range image;
[0013] Step 6: Convert the distance feature to BEV feature, concatenate it with the BEV feature previously obtained from voxel representation and image fusion, and use this multi-source multi-representation feature for 3D object detection.
[0014] Compared with the prior art, the present invention has the following advantages:
[0015] (1) The voxel representation of point cloud data regularizes disordered point clouds and can effectively encode multi-scale features, but it also has shortcomings. First, there is quantization loss in the process of converting point clouds to voxels, including the detailed position and geometric information of the original point cloud. Secondly, the sparsity of point clouds makes most voxels empty, which is not conducive to feature extraction. Considering that range images have the advantages of compact structure and density, they not only retain the spatial information of the original data, but also can be processed by mature two-dimensional convolution by reducing the dimensionality of three-dimensional data. The present invention combines the advantages of voxel representation of point cloud data and range image representation, reduces information loss in the process of point cloud data processing, improves the feature representation ability, and is conducive to the improvement of the final 3D target detection accuracy.
[0016] (2) Point cloud data lacks semantic information such as color and texture, which makes it difficult for the model to accurately detect small objects and distant objects with sparse point clouds, and difficult to distinguish easily confused objects. The features of different modalities are fused, and the semantic features of image data are used to enhance the point cloud features with depth information to make up for the defects of a single sensor in 3D target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall process of 3D object detection method;
[0018] Figure 2 It is a schematic diagram of the network architecture of the 3D target detection method;
[0019] Figure 3 It is a schematic diagram of the image feature indexing process;
[0020] Figure 4 is a schematic diagram of the gated fusion module;
[0021] Figure 5 It is a schematic diagram of the process of extracting distance image features using frequency-aware adaptive dilated convolution. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. The following implementation cases are for the technical personnel to understand the specific process of the present invention, but do not limit the present invention in any form. In addition, changes and improvements can be made on the basis of the present invention without departing from the concept of the present invention.
[0023] The present invention provides a 3D target detection method based on multi-source multi-representation information fusion. The overall process is as follows: Figure 1 As shown, the method comprises the following steps:
[0024] Step 1: Voxelize the input 3D point cloud data, encode the voxel features, aggregate the information of the points contained in each voxel, and obtain the corresponding voxel features;
[0025] Due to the disorder and irregularity of the input raw point cloud data, the point cloud is regularized through voxel operation so as to apply convolution operation processing. The voxelization process is as follows:
[0026] The format of the input 3D point cloud data is set to (N, d), where N is the number of points in the point cloud, and d is the feature dimension of the point cloud, including the three-dimensional coordinates (x, y, z) and intensity value i of the point. The point cloud contains a three-dimensional space with a range of (D, H, W) along the Z, Y, and X axes respectively, and defines the size of each voxel as (v D ,v H ,v W ), so the number of voxels generated is (D / v D ,H / v H ,W / v W ).
[0027] Voxel feature encoding is performed by calculating the average value of all point features in non-empty voxels to obtain the corresponding voxel feature F LiDAR .
[0028] Step 2: Use the DeepLabv3 model as the 2D backbone network to extract features from the image data and obtain image features;
[0029] Step 3: An image feature indexing method is proposed. For each voxel feature, the voxel coordinates are mapped to the image, and its corresponding image features are indexed. The voxel features and image features of each voxel are fused to obtain multimodal features.
[0030] Specifically, the process of indexing image features is as follows: Figure 3 As shown, the specific operations are:
[0031] Step 3.1: According to the spatial coordinate index, point cloud range and voxel size, the corresponding three-dimensional voxel coordinates are converted and calculated from the voxel features;
[0032] Step 3.2: Since data enhancement is performed when processing 3D point cloud data, including flipping, scaling, rotation, etc., the inversion enhancement module is used to obtain the original voxel coordinates. When performing data enhancement, the corresponding enhancement parameters are recorded, and then the recorded parameters are used to invert the data enhancement to obtain the original voxel coordinates;
[0033] Step 3.3: Map the voxel coordinates to the image to obtain the corresponding image pixel coordinates, index from the image feature map according to the pixel position, and integrate them into the same format as the voxel features for subsequent fusion to obtain the corresponding image feature F Camera ;
[0034] After obtaining the corresponding image features, the features of the two modalities are adaptively fused through the gated fusion network, such as Figure 4 As shown, the specific process is:
[0035] First, the voxel feature F LiDAR With image feature F Camera After 3×3 convolution and Sigmoid function, the adaptive weights of different modalities are obtained. The obtained weights are multiplied with the feature maps of each modality to obtain the gated voxel features and gated image features. Finally, the features of different modalities are spliced to obtain the fusion feature F. fusion The specific operations are shown in equations (1) to (3):
[0036]
[0037] In the formula, Represents the channel concatenation operation, Conv L and Conv Cis a 3×3 convolutional layer, σ is the Sigmoid function, and × represents the multiplication operation of elements.
[0038] Step 4: Use a 3D sparse convolutional network to extract deep features from multimodal features and convert the deep features into BEV features;
[0039] Step 5: Convert the input 3D point cloud data into a range view representation through range projection, and propose a frequency-aware adaptive dilated convolutional network to extract multi-scale features from the range image;
[0040] Range image is a raw data format of LiDAR sensor, which is more compact and dense than point cloud representation, and is not affected by the sparsity problem of point cloud to a certain extent. It not only reduces the dimension of data by projecting point cloud into 2D image, but also retains rich original information.
[0041] The distance image projection operation is shown in formula (4):
[0042]
[0043] Where (p, q) is the pixel coordinate in the distance image, (x, y, z) is the three-dimensional coordinate of the point, w and h are the predefined width and height of the distance image, and r is the distance from each point to the origin. f=f up +f down is the vertical viewing angle of the laser radar. f=f up +f down , f is the vertical field of view of the laser radar, f up and f down They are the upward and downward scanning angles of the laser radar in the horizontal direction respectively.
[0044] Through formula 4), the original point cloud data is converted into a distance view (h,w,5), and the five channels are spatial coordinates (x,y,z), distance r, and reflection intensity i.
[0045] In the 3D object detection task, the range image provides a dense and compact representation for the use of two-dimensional convolution, but there is also a scale change problem, that is, the scales of objects at different distances vary greatly, and the traditional convolutional network cannot effectively extract the features of objects of different scales. Therefore, dilated convolution is used to extract the features of the range image. However, a fixed dilation rate will damage the feature representation. When the dilation rate is large, the receptive field is correspondingly expanded, but the bandwidth of the frequency response is also reduced, and the network's ability to process high-frequency components is reduced. Therefore, too large a dilation rate should not be used in the high-frequency part of the feature. In order to better adapt to different scales and obtain a more flexible receptive field while reducing the loss of high-frequency components, in the dilated convolution, the dilation rate is adaptively adjusted through frequency perception. Figure 5Schematic diagram of the process of extracting range image features using frequency-aware adaptive dilated convolution. First, before extracting multi-scale features, the high-frequency and low-frequency components in the feature representation are balanced so that the dilated convolution can be used more effectively to enhance the receptive field. Second, the dilation rate is predicted for each pixel position in the feature map based on the high and low frequency components. Then, the dilated convolution with this adaptive dilation rate is used to extract the range image features.
[0046] Specifically, the distance feature extraction steps described in step 5 are:
[0047] Step 5.1: For the initial distance feature map X obtained by the convolution layer, convert it to the frequency domain through Fourier transform to obtain the feature X in the frequency domain F , calculated as shown in formula 5):
[0048]
[0049] Where H and W are the height and width of the feature map, h and w are the coordinates in the feature map, u and v are the discrete frequencies in the height and width dimensions, and their values are {0,1,...,H-1},{0,1,...,W-1} respectively.
[0050] Step 5.2: In the frequency domain feature X F Different masks M are applied to decompose it into different frequency bands, so as to extract the characteristics of different high and low component frequencies. The calculation formula is as follows:
[0051]
[0052] Where f t , f t+1 is a predefined frequency decomposition threshold whose value is {0, 1 / 16, 1 / 8, 1 / 4, 1 / 2}.
[0053] The mask M(u,v) is combined with the frequency domain feature X F Multiply them together and obtain the feature maps X of four different frequency bands through inverse Fourier transform. b .
[0054] Step 5.3: Adaptively reweight the features of different frequency bands and add them together to obtain high and low frequency balanced features As shown in formula (7):
[0055]
[0056] In the formula, A b is the adaptive weight.
[0057] Step 5.4: Use frequency-aware adaptive dilated convolution to extract range image features. For features with frequency information Through the convolution operation, a dilation rate D(p) is predicted for each pixel position, so as to achieve the purpose of using a larger dilation rate at a lower frequency position to expand the receptive field and a smaller dilation rate at a higher frequency position to reduce the loss of frequency information. The features of the distance image are adaptively extracted by dilation convolution with the predicted dilation rate, as shown in formula (8):
[0058]
[0059] Where Y(p) is the pixel value of the output feature map at position p, K is the size of the convolution kernel, and W is i is the weight, X(p+△p i ×D(p)) is the pixel value at position p of the input feature map, △p i is the offset, and D(p) is the expansion rate.
[0060] Step 6: Considering the large scale variation of the range image, it is difficult to assign detection boxes directly on the range image. The occlusion problem also makes it difficult to remove redundant bounding boxes during non-maximum suppression. In contrast, objects are naturally separated in the BEV view, so the features extracted from the range image are transferred to the BEV view to generate detection boxes. Using the point-by-point representation as the intermediate representation, the full-resolution output features in the range view and the approximate one-to-one correspondence between the range image and the point cloud are utilized to recover the point-by-point features from the range image pixel features, and then convert the point-by-point features into BEV features. After converting the distance features to BEV features, they are concatenated with the BEV features previously obtained from the fusion of the voxel representation and the image, and this multi-source and multi-representation feature is used for 3D object detection. The overall network architecture is shown in the figure below. Figure 2 shown.
[0061] The performance of the method of the present invention is compared with that of different 3D object detection methods on the KITTI dataset, and the results are shown in Table 1. AP is a commonly used evaluation index in 3D object detection tasks. The higher the AP, the higher the detection accuracy. Easy, Moderate, and Hard are the three difficulty levels of the KITTI 3D detection dataset, which are determined according to the size, occlusion state, and truncation level of the target. The results on the KITTI dataset show that the method proposed in the present invention improves the accuracy of 3D object detection.
[0062]
[0063] The above describes the specific implementation cases of the present invention. It should be noted that the present invention is not limited to the above specific implementations, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essential content of the present invention. In the absence of conflict, the implementation cases of the present invention application and the features in the implementation cases can be arbitrarily combined with each other.
Claims
1. A 3D object detection method based on multi-source and multi-representation information fusion, characterized in that: The following steps are involved: Step 1: Voxelize the input 3D point cloud data, encode the voxel features, aggregate the information of the points contained in each voxel, and obtain the corresponding voxel features; Step 2: Use the DeepLabv3 model as the 2D backbone network to extract features from the image data and obtain image features; Step 3: An image feature indexing method is proposed. For each voxel feature, the voxel coordinates are mapped to the image, and its corresponding image features are indexed. The voxel features and image features of each voxel are fused to obtain multimodal features. Step 4: Use a 3D sparse convolutional network to extract deep features from multimodal features and transform the deep features into bird’s-eye view (BEV) features; Step 5: Convert the input 3D point cloud data into a range view representation through range projection, and propose a frequency-aware adaptive dilated convolutional network to extract multi-scale features from the range image; Step 6: Convert the distance feature to BEV feature, concatenate it with the BEV feature previously obtained from voxel representation and image fusion, and use this multi-source multi-representation feature for 3D object detection.
2. The 3D object detection method based on multi-source multi-representation information fusion according to claim 1, characterized in that: The voxelization process in step 1 is as follows: the format of the input 3D point cloud data is set to (N, d), where N is the number of points in the point cloud, and d is the characteristic dimension of the point cloud, including the three-dimensional coordinates (x, y, z) and intensity value i of the point; the point cloud contains a three-dimensional space with a range of (D, H, W) along the Z, Y, and X axes respectively, and the size of each voxel is defined as (v D ,v H ,v W ), so the number of voxels generated is (D / v D ,H / v H ,W / v W ); The voxel feature encoding is performed by calculating the average value of all point features in the non-empty voxel to obtain the corresponding voxel feature F LiDAR .
3. The 3D object detection method based on multi-source multi-representation information fusion according to claim 1, characterized in that: The image feature indexing and multimodal fusion process in step 3 includes: Step 3.1: According to the spatial coordinate index, point cloud range and voxel size, the corresponding three-dimensional voxel coordinates are converted and calculated from the voxel features; Step 3.2: Since data enhancement, including flipping, scaling, rotation, etc., is performed when processing 3D point cloud data, the inversion enhancement module is used to obtain the original voxel coordinates; when performing data enhancement, the corresponding enhancement parameters are recorded, and then the recorded parameters are used to invert the data enhancement to obtain the original voxel coordinates; Step 3.3: Map the voxel coordinates to the image to obtain the corresponding image pixel coordinates, index from the image feature map according to the pixel position, and integrate them into the same format as the voxel features for subsequent fusion to obtain the corresponding image feature F Camera ; Step 3.4: Adaptively fuse the features of the two modalities through the gated fusion network; first, the voxel feature F LiDAR With image feature F Camera After 3×3 convolution and Sigmoid function, the adaptive weights of different modalities are obtained. The obtained weights are multiplied with the feature maps of each modality to obtain the gated voxel features and gated image features. Finally, the features of different modalities are spliced to obtain the fusion feature F. fusion ; The specific operations are shown in equations (1) to (3): In the formula, ⊕ represents the channel concatenation operation, Conv L and Conv C is a 3×3 convolutional layer, σ is the Sigmoid function, and × represents the multiplication operation of elements.
4. The 3D object detection method based on multi-source multi-representation information fusion according to claim 1, characterized in that: The distance projection operation in step 5 is shown in formula (4): Where (p, q) is the pixel coordinate in the distance image, (x, y, z) is the three-dimensional coordinate of the point, w and h are the predefined width and height of the distance image, and r is the distance from each point to the origin. f=f up +f down is the vertical viewing angle of the laser radar; f = f up +f down , f is the vertical field of view of the laser radar, f up and f down are the upward and downward scanning angles of the laser radar in the horizontal direction respectively; Through formula 4), the original point cloud data is converted into a distance view (h,w,5), and the five channels are spatial coordinates (x,y,z), distance r, and reflection intensity i.
5. The 3D object detection method based on multi-source multi-representation information fusion according to claim 1, characterized in that: The specific process of extracting distance features through the frequency-aware adaptive dilated convolutional network in step 5 is: Step 5.1: For the initial distance feature map X obtained by the convolution layer, convert it to the frequency domain through Fourier transform to obtain the feature X in the frequency domain F , calculated as shown in formula 5): Where H and W are the height and width of the feature map, h and w are the coordinates in the feature map, u and v are the discrete frequencies in the height and width dimensions, and their values are {0,1,...,H-1},{0,1,...,W-1} respectively; Step 5.2: In the frequency domain feature X F Different masks M are applied to decompose it into different frequency bands, so as to extract the characteristics of different high and low component frequencies. The calculation formula is as follows: Where f t , f t+1 is the predefined frequency decomposition threshold, and its value is {0, 1 / 16, 1 / 8, 1 / 4, 1 / 2}; The mask M(u,v) is combined with the frequency domain feature X F Multiply them together and obtain the feature maps X of four different frequency bands through inverse Fourier transform. b ; Step 5.3: Adaptively reweight the features of different frequency bands and add them together to obtain high and low frequency balanced features As shown in formula (7): In the formula, A b is the adaptive weight; Step 5.4: Extract range image features using frequency-aware adaptive dilated convolution; For features with frequency information A dilation rate D(p) is predicted for each pixel position through convolution operation; the features of the distance image are adaptively extracted with the predicted dilation rate through dilated convolution, as shown in formula (8): Where Y(p) is the pixel value of the output feature map at position p, K is the size of the convolution kernel, and W is i is the weight, X(p+△p i ×D(p)) is the pixel value at position p of the input feature map, △p i is the offset, and D(p) is the expansion rate.
6. The 3D object detection method based on multi-source multi-representation information fusion according to claim 1, characterized in that: The conversion of the distance features to BEV features in step 6 uses a point-by-point representation as an intermediate representation, utilizes the full-resolution output features in the distance view and the approximate one-to-one correspondence between the distance image and the point cloud, recovers the point-by-point features from the distance image pixel features, and then converts the point-by-point features into BEV features.