A multi-dimensional target detection method based on the fusion of four types of multi-modal data
By fusing multimodal data of radar, infrared, magnetic field and color images, using the attention module to generate suggestions and prediction boxes, the difficulties of occlusion and internal target detection in robots and autonomous driving are solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202111255921.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-10-27
AI Technical Summary
The prior art is difficult to effectively detect occlusion and internal objects in robots and autonomous driving, and there is less detection of three-dimensional targets fused with other data sources.
A method based on four types of multimodal data fusion is adopted, including radar, infrared, magnetic field and color images, and features are extracted through convolutional neural networks, and suggestions and prediction boxes are generated in combination with attention modules to solve the problems of object occlusion and internal target detection difficulties.
The detection accuracy of the inside of the object and the occlusion target is improved, the limitations of a single data source are compensated, and the advantages of multimodal data are complementary.
Smart Images

Figure CN113971801B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning, image recognition, and three-dimensional object detection, and particularly relates to a multi-dimensional object detection method based on the fusion of four types of multi-modal data. Background Art
[0002] In many practical applications such as robots, automatic loading, and autonomous driving, the three-dimensional position information of objects has received increasing attention, and three-dimensional object detection is a key technology for establishing an interaction mechanism between machines and the environment.
[0003] Currently, the three-dimensional object detection methods based on radar point clouds mainly include two types. One is to voxelize the point cloud, such as VoxelNet; the other is to project the point cloud onto a two-dimensional plane, such as PIXOR. The methods for fusing radar point clouds with images as an auxiliary mainly include: MV3D that uses the top view and front view of the point cloud to fuse with the image, AVOD that uses the top view of the point cloud to fuse with the image, etc. Detecting small targets and occluded targets is still the most challenging at present, and there is currently little research on detecting internal targets of objects, and there is also little three-dimensional object detection that fuses information from other data sources. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to solve the technical problems proposed in the above background art. The present invention aims to provide a multi-dimensional object detection method based on the fusion of four types of multi-modal data, fuse the image information of multiple data sources, and integrate the attention network into the multi-modal three-dimensional object detector to solve the problems of object occlusion and difficulty in detecting internal targets of objects.
[0005] To achieve the above technical objectives, the present invention proposes a multi-dimensional object detection method based on the fusion of four types of multi-modal data, including:
[0006] Step 1: Collect the radar, infrared, magnetic field, and color images of the target to be detected, divide them into a training set, a validation set, and a test set, and perform three-dimensional anchor box annotation to generate a data set;
[0007] Step 2: Build four structurally independent convolutional neural networks as the backbone network to extract the feature maps of the four input images, and at the same time use the GAU module to enhance the information of the features;
[0008] Step 3: Preset three-dimensional anchor boxes through clustering on the training set, project them onto the four feature maps, crop and pool them to the same size for fusion, build an RPN network to generate proposal boxes, and at the same time introduce an attention module;
[0009] Step 4: Project the proposal boxes generated by the RPN network onto four feature maps, crop and pool them to the same size for fusion, build a fully connected network to generate the final prediction boxes, and introduce an attention module simultaneously.
[0010] Furthermore, for the multi-dimensional detection method proposed by the present invention, step 1 includes:
[0011] Step 1.1: Integrate the radar, infrared sensor, magnetic sensor, and camera together to ensure the alignment of the four types of images, collect a large number of target images of the four types, remove the unclear images among them, and convert the radar point cloud data into a BEV bird's-eye view.
[0012] Step 1.2: Divide the obtained dataset into a training set, a validation set, and a test set according to a certain ratio, perform three-dimensional anchor box annotation on the training set and the validation set, and the test set is used to evaluate the training effect of the target detection network.
[0013] Furthermore, in step 2 of the multi-dimensional detection method proposed by the present invention, four structurally independent convolutional neural networks are used to extract features from the four input images respectively. The backbone network adopts the VGG16 structure and is truncated at conv-4. The number of filters in each convolutional layer becomes half of the original, and finally four feature maps with 256 channels are extracted. At the same time, the GAU module is used to enhance the information of the feature maps.
[0014] Furthermore, the multi-dimensional detection method proposed by the present invention, step 3 includes:
[0015] Step 3.1: Use a clustering algorithm on the training set to generate a large number of predefined anchor boxes for each category, project them onto the four output feature maps of the backbone network, crop the corresponding parts, and adjust them to feature maps with the same width and height through pooling operations.
[0016] Step 3.2: For each anchor box, fuse the four feature maps through element-wise averaging operation, then input it into a fully connected network, and finally output the regression parameters of the anchor box and the score for the foreground.
[0017] Step 3.3: Introduce an attention module in the RPN network, use the classification recognition localization strategy Grad-CAM to obtain the output feature map of the last convolutional layer, calculate the gradient of the feature map during backpropagation, take the average as the weight of each feature map, and finally perform weighted summation and pass through the LeakyReLU activation function to obtain the class activation map; then use the inverse attention network IAN to generate the inverse attention map in the spatial direction and the inverse attention map in the channel direction, then combine them to generate the inverse attention map, and finally multiply it with the output feature map of the convolutional layer.
[0018] Furthermore, in the multi-dimensional detection method proposed by the present invention, in step 4, the proposed bounding boxes generated in step 3 are projected onto four feature maps, cropped and pooled to the same size, and then fused using element-wise averaging operation, and input into a fully connected network, and finally the regression parameters, orientation estimation, and class classification of each proposed bounding box are output; at the same time, an attention module is also introduced, and the GradCAM and gradient-based IAN are used to calculate the reverse attention map, which is then multiplied element-wise with the fused feature map.
[0019] The present invention adopts the above technical solutions, and has the following technical effects compared with the prior art:
[0020] The present invention combines multiple data sources such as color images, radar, infrared, and magnetic fields, making up for the limitations of single data, and can achieve the effect of complementary advantages. For objects inside an object, the problems of information acquisition are solved through infrared and magnetic fields; in addition, the problem of object occlusion can be solved by integrating an attention network into a multi-modal three-dimensional object detector. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic diagram of the overall architecture of the present invention.
[0022] Figure 2 is a structural diagram of the backbone feature extraction network of the present invention.
[0023] Figure 3 is a structural diagram of the attention module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0024] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings.
[0025] As Figure 1 shown, the present invention proposes a target multi-dimensional detection method based on the fusion of four types of multi-modal data. The method includes the following steps:
[0026] Step 1: Collect the radar, infrared, magnetic field, and color images of the target to be detected, divide them into a training set, a validation set, and a test set, and perform three-dimensional anchor box annotation to generate a data set.
[0027] The specific implementation of this step is as follows:
[0028] Integrate radar, infrared sensors, three-axis magnetic sensors, and cameras together to ensure the alignment of the four types of images, collect sufficient target images of the four types, and remove the unclear images among them. Among them, the radar point cloud data is converted into a BEV (bird's-eye view), and the data measured by the three-axis magnetic sensor can be represented as a quadratic surface and projected onto the plane of the current perspective; the obtained dataset is divided into a training set, a validation set, and a test set according to a ratio of 2:1:1. The training set and the validation set are labeled with 3D anchor boxes, and the test set is used to evaluate the training effect of the target detection network.
[0029] Step 2: Build four structurally independent convolutional neural networks as the backbone networks to extract the feature maps of the four input images for subsequent feature fusion.
[0030] The specific implementation of this step is as follows:
[0031] For the four input image data, four structurally independent backbone feature extraction networks are used. The extraction network consists of two parts: an encoder and a decoder. The encoder is built according to VGG-16 and some modifications are made, mainly reducing the number of channels by half and cutting off the network at the conv-4 layer. The decoder adopts a bottom-up feature pyramid structure, performs global average pooling on the features output by the encoder, then performs 1*1 convolution to change the number of channels to half of the original, that is, the number of channels of the previous-level feature. Then, the sigmoid activation function is used to compress the values to between 0 and 1 as the weights in the channel direction, and then multiplied by the previous-level feature to obtain a new feature map. Finally, the features output by the encoder are upsampled to the same size and channels as the previous-level feature and added to the new feature map for fusion. The finally output feature map has high resolution and representativeness. The structure of the backbone feature extraction network is as Figure 2 shown.
[0032] Step 3: Preset 3D anchor boxes through clustering on the training set, project them onto the four feature maps, crop and pool them to the same size for fusion, build an RPN network to generate proposal boxes, and introduce an attention module at the same time.
[0033] The specific implementation of this step is as follows:
[0034] On the training set, a clustering algorithm is used to generate a large number of predefined 3D anchor boxes with determined sizes for each class, where the anchor boxes are determined by six parameters: the centroid (tx, ty, tz) and the axis-aligned dimensions (dx, dy, dz);
[0035] Using 3D ROI to process such high-dimensional feature maps will greatly increase the computational complexity. Therefore, 1×1 convolution operations are used on the feature maps output by the backbone feature extraction network to reduce the number of channels;
[0036] Project the predefined 3D anchor boxes onto the four feature maps output by the backbone feature extraction network, and crop out the corresponding parts of the anchor boxes. Since the sizes of the anchor boxes are not fixed, in order to facilitate the fusion of the four feature maps, pooling operations are used to unify them to the same size;
[0037] For each anchor box, the cropped parts on the four feature maps have become the same size. Then, the four feature maps are fused through element-wise averaging operation, and then input into the fully connected network. Finally, the regression parameters of the 3D anchor boxes and the score of the current feature map being the foreground are output; In the loss function of the RPN network, the classification loss function uses cross-entropy loss, and the RPN regression loss uses smooth L1 loss;
[0038] On the BEV, distinguish the foreground and background by the IoU between the proposed boxes and the ground truth boxes, and use 2D NMS (Non-Maximum Suppression) on the BEV to remove overlapping proposed boxes;
[0039] An attention module is introduced in the RPN network. The attention module is as Figure 3 shown. Use Grad-CAM (a classification recognition and localization strategy) to obtain the output feature map of the last convolutional layer. When performing backpropagation, obtain the gradient of the feature map, take the sum of the global average and the global maximum as the weight of each feature map, and finally perform weighted summation and pass through the LeakyReLU activation function to obtain the class activation map.
[0040] Calculation of the feature map weight:
[0041]
[0042] Among them, Sc is the score of the c-th class, the size of the feature map is c1*c2, Z = c1*c2, is the pixel value of the i-th feature map at the k-th row and the j-th column;
[0043] Calculation of the class activation map of Grad-CAM:
[0044]
[0045] Use the LeakyReLU activation function to focus on the regions related to the category, that is, the parts of the feature map with values greater than 0, and retain the parts unrelated to the category with smaller values;
[0046] During the standard training process, the gradient descent algorithm will force the attention map to converge to the few most sensitive parts of the object, while ignoring the other less sensitive parts of the object.
[0047] Iteratively invert the original attention tensor through IAN, that is, the reverse attention tensor, thus forcing the network to detect objects based on the less sensitive parts of the objects. Specifically, we generate a reverse attention map in the spatial direction and a reverse attention map in the channel direction, and then combine them to generate the final attention map.
[0048] Calculation of the reverse attention map in the spatial direction:
[0049]
[0050] where T s1 、T s2 are the thresholds of the spatial attention map;
[0051] Calculation of the reverse attention map in the channel direction:
[0052]
[0053] where T c1 、T c2 are the thresholds of the channel attention map;
[0054] Finally, the two are multiplied element-wise to obtain the attention map, which is then multiplied by the fused feature map to complete the addition of the attention module.
[0055] Step 4: Project the proposal boxes generated by RPN onto four feature maps, crop and pool them to the same size for fusion, build a fully connected network to generate the final prediction boxes, and introduce the attention module at the same time.
[0056] The specific implementation of this step is as follows:
[0057] Similar to the operation in Step 3, project the retained proposal boxes in Step 3 onto the four feature maps output by the backbone feature extraction network, crop and pool them to the same size, then perform an element-wise average operation for fusion, input into the fully connected network, and finally output the regression parameters, direction estimation, and class classification of each proposal box;
[0058] Encode the bounding box using four corners and two height values. The two height values represent the vertical angular offsets of the ground plane determined from the sensor height.
[0059] Therefore, the regression target becomes (Δx1…Δx4,Δy1…Δy4,Δh1,Δh2), that is, the offset values of the corners and heights between the proposal box and the ground truth box;
[0060] Use the regression direction vector to solve the ambiguity in the bounding box direction estimation using the four-corner representation adopted. Calculation of the direction vector:
[0061] (xθ, yθ) = (cosθ, sinθ)
[0062] where θ ∈ [-π, π];
[0063] The direction vector is represented as a unique unit vector in the BEV space.
[0064] The attention module is similar to that in step 3; in the loss function of the second-stage detection network, the classification loss function uses softmax loss, and the regression loss function uses L1 loss.
[0065] The embodiments are only for illustrating the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the present invention.
Claims
1. A target multi-dimensional detection method based on the fusion of four types of multi-modal data, characterized in that, It includes the following steps: Step 1: Collect the radar, infrared, magnetic field, and color images of the target to be detected, divide them into a training set, a validation set, and a test set, and perform 3D anchor box annotation to generate a dataset; Step 2: Build four structurally independent convolutional neural networks as the backbone network to extract the feature maps of the four input images; Specifically, the four structurally independent convolutional neural networks are used to extract the features of the four input images respectively. The backbone network adopts the VGG16 structure and is truncated at conv-4. The number of filters in each convolutional layer becomes half of the original. Finally, four feature maps with 256 channels are extracted, and the GAU module is used to enhance the information of the feature maps; Step 3: Preset 3D anchor boxes through clustering on the training set, project them onto the four feature maps, crop and pool them to the same size for fusion, build an RPN network to generate proposal boxes, and introduce an attention module at the same time; Specifically, it includes: Step 3.1: Use the clustering algorithm on the training set to generate a large number of predefined anchor boxes for each category, project them onto the four output feature maps of the backbone network, crop the corresponding parts and adjust them to feature maps with the same width and height through pooling operations; Step 3.2: For each anchor box, fuse the four feature maps through element-wise average operation, then input them into a fully connected network, and finally output the regression parameters of the anchor box and the score for the foreground; Step 3.3: An attention module is introduced in the RPN network. The classification recognition localization strategy Grad-CAM is used to obtain the output feature map of the last convolutional layer. When performing backpropagation, the gradient of the feature map is obtained, and the sum of the average and the maximum value is used as the weight of each feature map. Finally, the weighted sum passes through the LeakyReLU activation function to obtain the class activation map; then use the inverse attention network IAN to generate the inverse attention map in the spatial direction and the inverse attention map in the channel direction, then combine them to generate the inverse attention map, and finally multiply it with the output feature map of the convolutional layer; Step 4: Project the proposal boxes generated by the RPN network onto the four feature maps, crop and pool them to the same size for fusion, build a fully connected network to generate the final prediction boxes, and introduce an attention module at the same time.
2. The multi-dimensional detection method according to claim 1, wherein Step 1 includes: Step 1.1: Integrate the radar, infrared sensor, magnetic sensor, and camera together to ensure the alignment of the four images, collect a sufficient number of target images of the four types, and remove the unclear images among them. The radar point cloud data is converted into a BEV bird's-eye view; Step 1.2: Divide the obtained dataset into a training set, a validation set, and a test set according to a certain ratio, perform 3D anchor box annotation on the training set and the validation set, and the test set is used to evaluate the training effect of the target detection network.
3. The multi-dimensional detection method according to claim 1, wherein, In Step 4, the proposal boxes generated in Step 3 are projected onto the four feature maps, cropped and pooled to the same size, then fused by element-wise average operation, input into a fully connected network, and finally the regression parameters, direction estimation, and class classification of each proposal box are output; at the same time, an attention module is also introduced, and the inverse attention map is calculated using GradCAM and gradient-based IAN, and then multiplied element-wise with the fused feature map.
Citation Information
Patent Citations
Improved SSD small target detection method based on dense feature pyramid
CN111652288A
Image target object real-time detection method and system, terminal and storage medium
CN113222064A