A multi-sensor fusion detection method, model, and model training method based on radar distance view and image

Through a multi-sensor fusion detection method based on radar distance view and image, multi-layer perceptron and neural networks are used to optimize model parameters, the data sparseness and light dependence problems of lidar and camera in driverless car environment perception are solved, and high-precision multi-objective detection is achieved.

CN114359664BActive Publication Date: 2025-07-11JIANGSU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111615306.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-07-11
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

In the prior art, lidar and cameras have problems with sparse point cloud data and light dependence in the perception of driverless car environments, resulting in high probability of missed detection and missed detection when detecting objects at a distance, making it difficult to effectively integrate the advantages of both.

Method used

A multi-sensor fusion detection method based on radar distance view and image is adopted, and camera images and lidar point cloud data are processed through a multi-layer perceptron (MLP), high-resolution feature maps are generated and rasterized. Combined with the decoder's regression of the object center point and 3D Bbox parameters, the loss function is used to optimize the model parameters using neural network training.

Benefits of technology

Without losing detection speed, the detection accuracy and robustness of driverless cars in complex environments is improved, and it can accurately detect targets such as vehicles, pedestrians, bicycles, traffic lights, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359664B_ABST
    Figure CN114359664B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-sensor fusion detection method, model, and model training method based on radar distance view and image. The detection method of the present invention is used for multi-object detection in a driving scenario, and the detection targets mainly include various vehicles, pedestrians, ego vehicles, traffic lights, and traffic signs. The detection model of the present invention is built based on a neural network, which extracts features and associates data from the distance view generated by a lidar and the image generated by a camera, and generates the attribute parameters finally used for detecting targets. The model training method is used for training the fusion detection model, and trains by calculating the gradients of the parameters of each layer of the neural network through a loss function and performing backpropagation to obtain the optimal neural network parameters, thereby determining the parameters of each layer of the fusion detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of driverless vehicles, and particularly relates to a multi-sensor fusion method, model, and model training method based on radar range view and images. Background Art

[0002] With the continuous development of driverless vehicles, their perception ability of the surrounding environment has increasingly become a research hotspot. LiDAR is not restricted by lighting conditions and performs well in detecting the distance and speed of objects. However, the generated point cloud data has the disadvantages of sparsity and discreteness, which increases the difficulty of subsequent data processing; cameras are remarkable in detecting features such as the surface texture and color of objects, have a large pixel density, and can retain a large amount of feature information of objects. However, they are greatly affected by lighting conditions and have poor working ability in relatively harsh weather, and the working conditions are relatively harsh; the fusion technology based on LiDAR and cameras can well combine the advantages of both, improve the ability of driverless vehicles to detect surrounding objects, improve their robustness, and thus improve their safe driving performance.

[0003] In recent years, most of the detection methods based on LiDAR project the point cloud generated by LiDAR onto a bird's-eye view (BEV) for post-processing, which further amplifies the disadvantage of sparse point cloud data. Especially when detecting objects at a relatively far distance, the probability of missed detection and false detection increases. The range view of LiDAR has the advantage of dense point cloud compared with the bird's-eye view, so it can well solve this problem. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a multi-sensor fusion detection method based on radar range view and images, including the following steps:

[0005] Step 1: First, preprocess the image detected by the camera, and then input the preprocessed image into a multi-layer perceptron (MLP) to obtain a high-resolution feature map, and regress the center point of the object and the parameters of the 3D Bbox through different regression heads, where the prediction of the center point of the object is realized through a keypoint heatmap;

[0006] Step 2: Project the LiDAR point cloud onto a range view, generate a high-semantic feature map through a multi-layer perceptron (MLP), and rasterize the high-semantic feature map;

[0007] Step 3: Project the center point predicted from the image in Step 1 onto the feature map generated from the LiDAR range view in Step 2.

[0008] Furthermore, the specific process of Step 1 is as follows:

[0009] Step 1.1: Assume the input image is \(I\in\mathbb{R}\) w*h*3 , where \(w\) and \(h\) are the width and height of the image respectively. Preprocess the input image. Without distorting the image, make it square by filling it with gray bars.

[0010] Perform preprocessing operations on the input image. First, select the long sides of all images, scale the long sides, the scaled size is 512, and the scaling factor is \(r\). Then multiply the other side by \(r\) to get the scaled side length. For the short side, use zero-padding to complete it, and process the final size of the image to be \(512\times512\). Then perform normalization on the image so that the pixel values of the image follow a Gaussian normal distribution.

[0011] Step 1.2: Send the preprocessed image into the ResNet-50 network for convolution and deconvolution operations to obtain a high-resolution feature map. ResNet-50 includes two basic modules, namely Conv Block and Identity Block. The input dimension and output dimension of Conv Block are different, and its function is to change the dimension of the network. Its structure is shown in Figure 3 the left; the input dimension and output dimension of Identity Block are the same, and its function is to deepen the network. Its structure is shown in Figure 3 the right; ConvBlock is mainly composed of two branches: the main branch on the left and the residual branch on the right. The main branch is composed of multiple convolutional layers (Conv 2d), normalization layers (BatchNorm), and activation function ReLu. Its main function is to further extract features. The residual branch is only composed of one convolutional layer (Conv 2d) and one normalization layer (BatchNorm). Its main function is to prevent gradient disappearance during backpropagation. The main process of Conv Block is: the input feature map is processed by the main branch and the residual branch respectively, then the generated data is added element by element, and then the activation function ReLu is used for non-linear processing. Its structure is shown in Figure 3 the left; the input and output dimensions of Identity Block are the same, and its function is to deepen the network. Its specific structure is similar to that of ConvBlock. The main difference is that the residual branch does not perform any processing on the feature map. Its structure is shown in Figure 3Right; after a series of convolutional operations including Conv Block and Identity Block, batch normalization (BN), and max pooling (Maxpool), the image preprocessed in Step 1.1 is first subjected to convolution, normalization, ReLU, and MaxPool operations. The generated feature map is passed into 1 Conv Block and 2 Identity Blocks, then into 1 Conv Block and 3 Identity Blocks, then into 1 Conv Block and 5 Identity Blocks, and then into 1 Conv Block and 2 Identity Blocks to generate feature map C5. Then, it undergoes three deconvolution upsamplings to obtain a higher-resolution output; Step 1.3: Perform 3 convolutional operations on the feature map generated in Step 1.2, which include a 3×3 convolutional layer and a 1×1 convolutional layer, to generate three predictions.

[0012] Furthermore, the output channels of the three deconvolutions in Step 1.2 are 256, 128, and 64 respectively. Each time a deconvolution is performed, the width and height of the feature map will be doubled. After 3 deconvolution upsamplings, the final feature map is generated.

[0013] Furthermore, the three predictions generated in Step 1.3 are as follows:

[0014] Heatmap prediction: At this time, the number of convolutional channels is num_classes, and the final tensor is (128, 128, num_classes), representing whether there is an object at each heat point and the type of the object. The predicted values in the last dimension num_classes represent the probabilities belonging to each class;

[0015] Center point prediction: The number of convolutional channels is 2, and the final tensor is (128, 128, 2), which represents the offset of the center point of each object from the heat point. The predicted values in the last dimension 2 represent the offset of the current feature point towards the lower right corner;

[0016] Width and height prediction: The number of convolutional channels is 2, and the final tensor is (128, 128, 2), which represents the offset of the center point of each object from the heat point. The predicted values in the last dimension 2 represent the width and height of the predicted bounding box corresponding to the current feature point.

[0017] Furthermore, the specific process of Step 2 is as follows:

[0018] Step 2.1: The lidar measures surrounding objects, including: distance (range r), reflectance, azimuth angle (θ) of the sensor, elevation angle of the laser For the wire harness ID m of the radar, each point in the point cloud can be represented by the following formula

[0019]

[0020] P represents each lidar point. By discretizing the ID m of the laser into the ordinate and the azimuth angle into the abscissa, a distance view is constructed. The lidar has 5 input channels: distance (r), height (z), azimuth angle (θ), reflectivity (e), and flag (indicating whether there is a point at this image position). When multiple points fall at the same position, the features of the nearest point are retained;

[0021] Step 2.2: Send the distance view generated in Step 2.1 into DLA34 for feature extraction. The DLA34 network structure includes three layers, and the sizes of the convolutional kernels are 64, 64, and 128 respectively. Each layer contains a feature extraction module and several feature aggregation modules. Downsampling is performed on the horizontal resolution, and the vertical resolution remains unchanged, outputting a high-semantic feature map;

[0022] Step 2.3: Perform a rasterization operation on the high-semantic feature map generated in Step 2.2, dividing it into a 128*128 raster map. The features of each raster are averaged by the features of the point cloud within the raster.

[0023] Furthermore, the specific process of Step 3 is as follows:

[0024] Step 3.1: Project the object center point predicted by the image end onto the high-semantic feature map generated in Step 2. Among them, the features corresponding to the 8 rasters around the center point are used as the features for the final pre-classification and regression of the 3D Bbox parameters;

[0025] Step 3.2: Perform regression and classification based on the features of the 8 rasters where the center point is located. Among them, the mean operation is taken for the 2D parameters of the regressed object and the parameters regressed in Step 1.3, and the weighted average operation is performed for the 3D parameters (such as the depth of the box) of the regression.

[0026] Furthermore, it also includes Step 4: Regress and predict the 3D Bbox and other parameters of the object through the decoder; specifically as follows:

[0027] The fused features are fed into the decoder, which mainly consists of two parts: an object Bbox size regressor and an object class regressor. Both regressors are composed of three convolutional layers. The parameters of the last convolutional layer of the object Bbox size regressor are that the number of convolutional channels is 3, and the final tensor is (128, 128, 3). The last dimension of the tensor represents the length, width, and height information of each object. The parameters of the last convolutional layer of the object class regressor are that the number of convolutional channels is 10, and the final tensor is (128, 128, 10). The last dimension represents the scores of 10 categories, which are: sedan, bus, truck, train, motorcycle, traffic sign, traffic signal, bicycle, cyclist, and pedestrian.

[0028] The present invention also proposes a multi-sensor fusion detection model based on radar range view and image: The neural network feature extraction module ResNet-50 extracts features from image data. The regression head performs different convolutional operations on the neural network extraction module to generate various regression parameters. The neural network extraction module DLA-34 extracts features from radar range view data and performs a rasterization operation, projects the regression parameters generated by the regression head onto the point cloud feature map, and extracts features from the fused data to finally generate the size information and category information of the object.

[0029] The present invention also proposes a training method for the above multi-sensor fusion detection model based on radar range view and image, including the following steps

[0030] Step 1: For each ground truth key point p ∈ R belonging to class C 2 , calculate its corresponding coordinates at low resolution Then use a Gaussian kernel where σ p is the standard deviation of the object size. The main loss function for center point prediction is as follows:

[0031]

[0032] where α and β are hyperparameters, N is the number of key points in image I, represents the heat map of the key points, R is the stride corresponding to the output of the original image, and C is the number of classes corresponding to the object detection;

[0033] Step 2: Let the bounding box corresponding to object k of class C k be The corresponding center point The size of each object k is

[0034] Assume that all classes use the same size prediction The dimensional loss of the object is

[0035]

[0036] A repulsive force loss function for vehicle occlusion is established, and its expression is:

[0037] L R = L Attr + τL RepGT + ωL RepC

[0038] where L Attr is the attraction term, L RepGT is the repulsion term of the predicted center point to other true center points, and τ and ω are hyperparameters.

[0039]

[0040]

[0041]

[0042] where P ∈ P + represents one of all the predicted center points, represents the true center point closest to the predicted point, are the other true center points except the true center that matches the predicted center point p;

[0043] Step 3: The expression of the final loss function is:

[0044] L = L k + λ s L size + λ R L R

[0045] where λ s and λ R are hyperparameters.

[0046] Furthermore, the model trained by this method can be set on an intelligent vehicle to achieve real-time fusion detection.

[0047] The beneficial effects of the present invention are:

[0048] The present invention proposes a multi-sensor fusion detection method, model, and model training method based on radar distance view and image. Without sacrificing the detection speed, the present invention improves the detection accuracy and enhances the robustness of the perception of the surrounding environment by the driverless vehicle in a complex driving environment.

[0049] The detection method is used for multi-target detection in driving environment scenarios, mainly including the detection of various vehicles, pedestrians, bicycles, traffic lights, traffic signs, etc., to obtain the target detection results, and can accurately detect the environment around the driverless vehicle for complex driving scenarios;

[0050] The detection model is built based on a neural network, which extracts features and associates data from the distance view generated by the lidar and the image generated by the camera, and generates the attribute parameters for the final detected target. By combining the distance information collected by the lidar and the image information collected by the camera, it can capture the three-dimensional size, distance, speed and other information of the target more accurately;

[0051] The model training method is used for the training of the fusion detection model. It trains by calculating the gradients of the parameters of each layer of the neural network through the loss function and backpropagation to obtain the optimal neural network parameters, thereby determining the parameters of each layer of the fusion detection model. Through this training method, the detection model can converge at a relatively fast speed, with a short training time and good model optimization effect. Brief Description of the Drawings

[0052] Figure 1 Fusion structure diagram;

[0053] Figure 2 Resnet-50 structure diagram;

[0054] Figure 3 Structure diagrams of Conv Block (left figure) and Identity block (right figure);

[0055] Figure 4 DLA34 (Deep Layer Aggregation Network) structure diagram;

[0056] Figure 5 Structure diagrams of the feature extraction module (left figure) and the feature aggregation module (right figure). Detailed Embodiment

[0057] The following will further illustrate the present invention in conjunction with the drawings, but the protection scope of the present invention is not limited thereto.

[0058] The present invention proposes a multi-sensor fusion detection method based on radar distance view and image, and its specific implementation model and process are as Figure 1 shown, mainly including the following steps:

[0059] Step 1: First, preprocess the image detected by the camera, and then input the preprocessed image into the neural network feature extraction module to obtain a high-resolution feature map. Through different regression heads, the center point of the object and various parameters of the 3D Bbox are regressed. Among them, the prediction of the center point of the object is achieved through the keypoint heatmap;

[0060] Step 1.1: Assume that the input image is I ∈ R w*h*3 , where w and h are the width and height of the image respectively. Preprocess the input image. Without distorting the image, fill it with gray bars to make it a square (512 * 512); specifically:

[0061] Perform preprocessing operations on the input image. First, select the long side of all images, scale the long side, and the scaled size is 512, and the scaling factor is r. Then multiply the other side by r to get the scaled side length. For the short side, use the padding method to complete it, and process the final image size to 512 * 512; then perform normalization on the image to make the pixel values of the image follow a Gaussian normal distribution. Step 1.2: Input the preprocessed image into the ResNet-50 network for convolution and deconvolution operations to obtain a high-resolution feature map. ResNet-50 has two basic modules, namely ConvBlock and Identity Block. The input dimension and output dimension of the Conv Block are different, and its function is to change the dimension of the network. Its structure is shown in Figure 3 the left; the input dimension and output dimension of the Identity Block are the same, and its function is to deepen the network. Its structure is shown in Figure 3 the right

[0062] The Conv Block is mainly composed of two branches: the main branch on the left and the residual branch on the right. The main branch is composed of multiple convolutional layers (Conv 2d), normalization layers (BatchNorm), and activation function ReLu. Its main function is to further extract features. The residual branch is only composed of one convolutional layer (Conv 2d) and one normalization layer (BatchNorm). Its main function is to prevent gradient disappearance during backpropagation. The main process of the Conv Block is: the input feature map passes through the main branch and the residual branch respectively, then the generated data is added element by element, and then the activation function ReLu is used for non-linear processing. Its structure is shown in Figure 3 the left; the input and output dimensions of the Identity Block are the same, and its function is to deepen the network. Its specific structure is similar to that of the Conv Block. The main difference is that the residual branch does not perform any processing on the feature map. Its structure is shown in Figure 3 the right

[0063] After a series of convolutional operations of Conv Block and Identity Block, as well as batch normalization (BN) and max pooling (Maxpool), the feature map C5(16, 16, 2048) is generated. Then, three transposed convolutions are performed for upsampling to obtain a higher-resolution output. To save computational effort, the output channels of these three transposed convolutions are 256, 128, and 64 respectively. Each time a transposed convolution is performed, the width and height of the feature map are doubled. After three times of transposed convolution upsampling, the finally generated feature map is (128, 128, 64). The specific structure is shown in Figure 2 ;

[0064] Step 1.3: Perform three convolutional operations on the feature map generated in the previous step, which includes a 3×3 convolutional layer and a 1×1 convolutional layer, to generate three predictions: Heatmap prediction: At this time, the number of convolutional channels is num_classes, and the final tensor is (128, 128, num_classes), representing whether there is an object at each heat point and the type of the object. The prediction value in the last dimension num_classes represents the probability of belonging to each class; Center point prediction: The number of convolutional channels is 2, and the final tensor is (128, 128, 2), which represents the offset of the center point of each object from the heat point. The prediction value in the last dimension 2 represents the offset of the current feature point towards the lower right corner; Width and height prediction: The number of convolutional channels is 2, and the final tensor is (128, 128, 2), which represents the offset of the center point of each object from the heat point. The prediction value in the last dimension 2 represents the width and height of the predicted bounding box corresponding to the current feature point.

[0065] Step 2: Project the lidar point cloud onto the range view, generate a feature map through a neural network extraction module, and rasterize the feature map;

[0066] Step 2.1: The measurement values of the lidar Velodyne 64E for surrounding objects include: range (r), reflectance (e), the azimuth angle (θ) of the sensor, and the elevation angle of the laser For the beam id m of the radar, each point in the point cloud can be represented by the following formula

[0067]

[0068] P represents each lidar point. By discretizing the id m of the laser into the ordinate and the azimuth angle into the abscissa, a distance view is constructed. Through the relationships of distance, azimuth angle, and pitch angle, the three-dimensional coordinates (x, y, z) of each point in the point cloud in the world coordinate system can be obtained, where x, y, and z represent the lateral distance, longitudinal distance, and vertical distance of each point relative to the origin, respectively. The lidar has 5 input channels: distance (r), height (z), azimuth angle (θ), reflectivity (e), and flag (indicating whether there is a point at the position of the image). When multiple points fall at the same position, the features of the nearest point are retained;

[0069] Step 2.2: To effectively combine multi-scale features, the distance view generated in Step 2.1 is sent into DLA34 (Deep Layer Aggregation structure) for feature extraction. The specific structure is shown in Figure 4 ; This network structure mainly consists of 3 layers, and the sizes of the convolutional kernels are 64, 64, and 128 respectively. Each layer contains a feature extraction module and several feature aggregation modules. The specific structure is shown in Figure 5 , Since the horizontal resolution of the distance view is high and the vertical resolution is low, only downsampling is performed in the horizontal resolution, and the vertical resolution remains unchanged. Finally, a high-semantic feature map is output.

[0070] Step 2.3: Perform a rasterization operation on the high-semantic feature map generated in the previous step, dividing it into a 128*128 raster map, where the features of each raster are averaged by the features of the point cloud within the raster.

[0071] Step 3: Project the center point predicted from the image in Step 1 onto the feature map generated from the lidar distance view in Step 2;

[0072] Step 3.1: Project the center point of the object predicted from the image side onto the high-semantic feature map generated in Step 2, where the features corresponding to the 8 rasters of the feature map around the center point are used as the features for the final pre-classification and regression of 3D Bbox parameters;

[0073] Step 3.2: Perform regression and classification based on the features of the 8 rasters where the center point is located. Among them, the mean operation is taken for the 2D parameters of the object to be regressed and the parameters regressed in Step 1.3 (such as the size of the 2D box), and the weighted average operation is performed for the 3D parameters to be regressed (such as the depth of the box);

[0074] Step 4: Regress and predict the 3D Bbox and other parameters of the object through the decoder;

[0075] Perform the regression of the 3D Bbox of the object through the object size regression head (Bbox head) in the decoder, and perform the regression of the object category through the object category regression head (Class head) in the decoder

[0076] Step 5: Conduct end-to-end training on the Figure 1 shown network model. After training, the model can perform real-time fusion detection;

[0077] The structure of the network model is specifically described as follows:

[0078] As Figure 1 , the neural network feature extraction module ResNet-50 extracts features from the image data, the regression head generates various regression parameters through different convolution operations of the neural network extraction module, the neural network extraction module DLA-34 extracts features from the radar range view data and performs rasterization operations, projects the regression parameters generated by the regression head onto the point cloud feature map, and finally extracts features from the fused data to generate the size information and category information of the object.

[0079] The specific training process of the model is as follows:

[0080] Step 5.1 During training, for each ground truth key point p ∈ R belonging to class C 2 , calculate its corresponding coordinates at low resolution Then use the Gaussian kernel where σ p is the standard deviation of the object size. The main loss function for center point prediction is as follows:

[0081]

[0082] where α and β are hyperparameters, N is the number of key points in image I, represents the heat map of the key points, R is the stride of the output corresponding to the original image, and C is the number of corresponding classes in object detection;

[0083] Step 5.2: Let the bounding box corresponding to object k of class C k be represent the upper left corner coordinates and lower right corner coordinates of the bounding box respectively, and the corresponding center point The size of each object k is

[0084] To reduce the computational complexity, assume that all classes use the same size prediction Then the object size loss is

[0085]

[0086] where S k represent the predicted object size and the true object size respectively.

[0087] To solve the occlusion problem, the present invention innovatively applies the Repulsion Loss in the field of face recognition to the field of automotive occlusion problems. Its expression is as follows:

[0088] L R = L Attr + τL RepGT + ωL size

[0089] where L Attr is the attraction term, L RepGT is the repulsion term of the predicted center point to other true center points, and τ and ω are hyperparameters.

[0090]

[0091]

[0092]

[0093] where Smooth L1 represents the L1 loss function, P ∈ P + represents one of all the predicted center points, represents the true center point closest to the predicted point, is the other true center points except the true center that matches the predicted center point p;

[0094] Step 5.3: The expression of the final loss function is:

[0095] L = L k + λ s L size + λ R L R

[0096] where λ s and λ R are hyperparameters.

[0097] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation modes of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent modes or changes that do not depart from the technology created by the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-sensor fusion detection method based on radar range view and image, characterized in that, It includes the following steps: Step 1: First, preprocess the image detected by the camera, and then input the preprocessed image into a multi-layer perceptron (MLP) to obtain a high-resolution feature map. The center point of the object and the parameters of the 3D Bbox are regressed through different regression heads. The prediction of the center point of the object is achieved through a keypoint heatmap; The specific process of Step 1 is as follows: Step 1.1: Assume the input image is \(I\in\mathbb{R}\) w*h*3 , where \(w\) and \(h\) are the width and height of the image respectively. Preprocess the input image by padding gray bars to make the image square without distortion. Perform preprocessing operations on the input image. First, select the long side of all images, scale the long side, the scaled size is 512, the scaling factor is r, then multiply the other side by r to get the scaled side length, and pad the short side with zeros to make the final size of the image 512*512; then perform normalization operations on the image to make the pixel values of the image follow a Gaussian normal distribution; Step 1.2: Input the preprocessed image into the ResNet-50 network for convolution and deconvolution operations to obtain a high-resolution feature map. ResNet-50 includes two basic modules, namely Conv Block and Identity Block. The input dimension and output dimension of the Conv Block are different, and its function is to change the dimension of the network; the Conv Block mainly includes two branches: the main branch on the left and the residual branch on the right. The main branch is composed of multiple convolutional layers (Conv2d), normalization layers (BatchNorm), and activation function ReLu, and its function is to further extract features. The residual branch includes a convolutional layer (Conv 2d) and a normalization layer (BatchNorm), and its function is to prevent gradient disappearance during backpropagation. The main process of the Conv Block is: the input feature map is processed by the main branch and the residual branch respectively, then the generated data is added element by element, and then the activation function ReLu is used for non-linear processing; the input and output dimensions of the Identity Block are the same, and its function is to deepen the network. Its specific structure is similar to that of the Conv Block, but the difference is that the residual branch does not perform any processing on the feature map; after the input feature map undergoes convolution, batch normalization (BN), and max pooling (Maxpool) operations of the Conv Block and Identity Block, the generated feature map is input into 1 Conv Block and 2 Identity Blocks, then into 1 Conv Block and 3 Identity Blocks, then into 1 Conv Block and 5 Identity Blocks, then into 1 Conv Block and 2 Identity Blocks to generate feature map C5, and then perform three deconvolution upsamplings on it to obtain a higher-resolution output; Step 1.3: Perform 3 convolution operations on the feature map generated in Step 1.2, which includes a 3×3 convolutional layer and a 1×1 convolutional layer to generate three predictions; Step 2: Project the lidar point cloud onto the range view, generate a high-semantic feature map through a multi-layer perceptron (MLP), and rasterize the high-semantic feature map; The specific process of the above Step 2 is as follows: Step 2.1: The lidar measures surrounding objects, including: distance, reflectivity, azimuth angle of the sensor, elevation angle of the laser, and the beam id m of the radar. Then each point in the point cloud can be represented by the following formula P represents each lidar point, r represents the distance, and z represents the height. represents the elevation angle of the laser, θ represents the azimuth angle of the sensor. By discretizing the id m of the laser into the ordinate y and the azimuth angle into the abscissa x, a distance view is constructed. The lidar has 5 input channels: distance, height, azimuth angle, reflectivity, and flag. The flag indicates whether there is a point at the position of the image. When multiple points fall on the same position, the features of the nearest point are retained. Step 2.2: Feed the range view generated in Step 2.1 into the DLA34 network for feature extraction. The DLA34 network structure consists of three layers, and the sizes of the convolutional kernels are 64, 64, and 128 respectively. Each layer contains a feature extraction module and several feature aggregation modules, and downsampling is performed on the horizontal resolution while the vertical resolution remains unchanged to output a high-semantic feature map; Step 2.3: Perform a rasterization operation on the high-semantic feature map generated in Step 2.2, dividing it into a 128*128 raster map, where the feature of each raster is the average of the features of the point cloud within the raster; Step 3: Project the center point predicted from the image in Step 1 onto the feature map generated from the lidar range view in Step 2; Step 4: Regress and predict the 3D Bbox of the object and other parameters through a decoder; specifically as follows: Input the fused features into the decoder. The decoder mainly consists of two parts, an object Bbox size regressor and an object category regressor. Both regressors are composed of three convolutional layers. The parameters of the last convolutional layer of the object Bbox size regressor are that the number of convolutional channels is 3, and the final tensor is (128, 128, 3). The last dimension of the tensor represents the length, width, and height information of each object. The parameters of the last convolutional layer of the object category regressor are that the number of convolutional channels is 10, and the final tensor is (128, 128, 10). The last dimension represents the scores of 10 categories, and the 10 categories are: sedan, bus, truck, train, motorcycle, traffic sign, traffic signal, bicycle, cyclist, pedestrian.

2. The multi-sensor fusion detection method based on radar distance view and image according to claim 1, characterized in that, The output channel numbers of the three transposed convolutions in Step 1.2 are 256, 128, and 64 respectively. Each time a transposed convolution is performed, the width and height of the feature map will be doubled. After 3 times of transposed convolution upsampling, the final feature map is generated.

3. A multi-sensor fusion detection method based on radar distance view and image according to claim 1, characterized in that, The three predictions generated in Step 1.3 are specifically as follows: Heatmap prediction: At this time, the number of convolutional channels is num_classes, and the final tensor is (128, 128, num_classes), representing whether there is an object at each heat point and the type of the object. The predicted values in the last dimension num_classes represent the probabilities belonging to each class; Center point prediction: The number of convolutional channels is 2, and the final tensor is (128, 128, 2), which represents the offset of the center point of each object from the heat point. The predicted values in the last dimension 2 represent the offset of the current feature point towards the lower right corner; Width and height prediction: The number of channels of the convolution is 2, and the final tensor is (128, 128, 2), which represents the offset of the center point of each object from the heat point. The predicted values in the last dimension 2 represent the width and height of the predicted bounding box corresponding to the current feature point.

4. A multi-sensor fusion detection method based on radar distance view and image according to claim 1, characterized in that The specific process of step 3 is as follows: Step 3.1: Project the object center point predicted by the image side onto the high-semantic feature map generated in step 2, and use the features corresponding to the 8 grids of the feature map around the center point as the features for finally pre-classifying and regressing the 3D Bbox parameters. Step 3.2: Perform regression and classification based on the feature maps of the 8 grids where the center point is located. Among them, the mean operation is taken for the 2D parameters of the regressed object and the parameters predicted in step 1.3, and the weighted average operation is performed on the 3D Bbox parameters of the regression.

5. A training method for a model of multi-sensor fusion detection based on radar distance view and image, characterized in that, The model: Use the neural network feature extraction module ResNet-50 to extract features from the image data. The regression head performs different convolution operations on the neural network extraction module to generate various regression parameters. The neural network extraction module DLA-34 extracts features from the radar distance view data and performs rasterization operations, projects the regression parameters generated by the regression head onto the point cloud feature map, extracts features from the fused data, and finally generates the size information and category information of the object. The training method of this model includes the following steps Step 1 For each ground truth key point p ∈ R belonging to category C 2 , calculate its corresponding coordinates at low resolution Then use a Gaussian kernel where σ p is the standard deviation of the object size, and the main loss function for center point prediction is as follows: where α and β are hyperparameters, N is the number of key points of the image I, represents the heatmap of key points, R is the stride corresponding to the output of the original image, and C is the number of corresponding classes in object detection; Step 2: Set class C k The bounding box corresponding to the object k of The corresponding center point The size of each object k is Assume that all classes use predictions of a unified size Then the object size loss is Establish a repulsive loss function for vehicle occlusion, and its expression is: L R = L Attr + tL RepGT + ωL RepC where L Attr is the attraction term, and L RepGT is the repulsion term of the predicted center point with respect to other true center points. t and ω are hyperparameters where P ∈ P + represents one of all the predicted center points, represents the true center point closest to the predicted point, are the other true center points except for the true center that matches the predicted center point p; Step 3: The expression of the final loss function is: L = L k + λ s L size + λ R L R where λ s and λ R are hyperparameters.

6. The model training method according to claim 5, wherein The model trained by this method can be set on an intelligent vehicle to achieve real-time fusion detection.

Citation Information

Patent Citations

  • A 3D vehicle detection method based on multi-sensor fusion

    CN109948661A

  • Three-Dimensional Object Detection

    US20200025931A1