A traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image
By synchronizing 4D millimeter-wave radar with monocular vision images, and combining a pre-encoder and feature extraction module, radar and vision fusion features are generated, solving the problems of wasted computing resources and traffic target detection in complex environments, and achieving efficient and accurate traffic target recognition and drivable area segmentation.
Patent Information
- Application Number
- CN202510005115.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-01-02
AI Technical Summary
In existing technologies, traffic target detection methods using multiple sensors lead to wasted computing resources and processing delays, making it difficult to meet real-time requirements. Furthermore, they struggle to effectively segment drivable areas when identifying dense traffic targets and areas without lane markings in complex environments with occlusion.
By synchronizing the temporal and spatial data of 4D millimeter-wave radar and monocular vision images, and combining a pre-encoder and feature extraction module, radar and vision fusion features are generated. Dynamic convolution and self-attention modules are used for feature extraction, and target boxes and drivable regions are generated through a detection head and a segmentation head. The target distance and velocity are output by combining kernel function density estimation.
It achieves efficient identification of dense traffic targets and accurate segmentation of drivable areas in complex environments, reduces computing resource requirements, and improves detection accuracy and real-time performance.
Smart Images

Figure CN120071266B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to target detection, semantic segmentation, 4D millimeter wave radar and monocular vision sensor technology, and in particular to a traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image. BACKGROUND
[0002] There is a significant problem in the practical application of traffic target detection, that is, there is a lack of an effective network framework to integrate different data features from different sources. Modern vehicles are usually equipped with multiple sensors, including cameras, lidar, millimeter wave radar, etc., each of which captures data of different types and features. For example, cameras provide rich image information, which helps to identify road traffic targets, but the effect is poor in bad weather conditions; millimeter wave radar can provide accurate distance and speed information, but the data is sparse and not intuitive.
[0003] Traffic target detection technology involves multiple complex tasks such as target detection, semantic segmentation, etc., which need to work together to provide effective driving assistance. However, current technology often trains and deploys models for each task separately, such as one model for target detection and another model for semantic segmentation. This way, although it can achieve independent optimization for each task, it causes a great waste of computing resources. The independent running of each model not only occupies a large amount of computing resources and memory, but also can cause processing delays, making it difficult to meet real-time requirements. In addition, multi-model collaboration also brings coordination and communication problems between models, increasing the complexity of the system. Therefore, the present application designs a unified model that can efficiently handle multiple tasks and optimize the utilization of computing resources to improve the practicality and popularity of traffic target detection technology. SUMMARY
[0004] The purpose of the present application is to provide a traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image, to solve the problem of dense traffic target recognition in complex urban environments with occlusion, and the problem of drivable area semantic segmentation for oncoming and outgoing lanes without lane line identification.
[0005] The technical solution to achieve the purpose of the present application is: a traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image, comprising the following steps:
[0006] Step 1, synchronize the start time of 4D millimeter wave radar and vision image, set consistent sampling interval to ensure that the sampling time of 4D millimeter wave radar data and monocular vision data is the same;
[0007] Step 2, synchronize the space of 4D millimeter wave radar and vision image;
[0008] Step 3, using a pre-encoder to coarsely extract 4D millimeter wave radar information features, generate a three-channel radar feature map, and pass the radar feature map through a dynamic convolution module. The picture data is processed using a self-attention module. The outputs of the dynamic convolution module and the self-attention module are connected together for fusion after extraction. The detection head uses an anchor-free method to directly regress the coordinates of the bounding box from the fused radar and image features. The drivable area segmentation head undergoes three upsampling processes to restore the output segmentation feature map to two channels with the same width and height as the original picture. The probability of each pixel in the input image corresponding to the drivable area and the background is generated. Based on this, a traffic target detection model based on the fusion of 4D millimeter wave radar pseudo images and monocular vision images is constructed.
[0009] Step 4, the 4D millimeter wave radar point cloud data collected in the field is preprocessed in the same way as steps 1 and 2, and is sent into the model trained in step 3 together with the monocular vision image data to generate target boxes and drivable areas.
[0010] Step 5, according to the kernel function density estimation method, the distance and speed of the target box identified in step 4 are output.
[0011] Further, step 1, synchronize the time of 4D millimeter wave radar and vision image, the specific method is:
[0012] Start the monocular camera and 4D millimeter wave radar at the same time, read the 4D millimeter wave radar data and monocular camera data in a multi-threaded manner. One sub-thread reads the video image in real time and maintains the number of images in the cache pool to 1. When the next video stream image arrives, update the image in the cache pool to ensure that the image in the cache pool is the current latest frame. Another sub-thread reads the 4D millimeter wave radar data. Sample data based on the low working frequency of the 4D millimeter wave radar sensor. When the 4D millimeter wave radar data arrives, read the latest image in the video sub-thread.
[0013] Further, step 2, synchronize the space of 4D millimeter wave radar and vision image, the specific method is:
[0014] Step 2.1, superimpose the data of the first two frames and the current frame of the 4D millimeter wave radar collected in step 1. When superimposing, only the number of 4D millimeter wave radar information is increased. Each 4D millimeter wave radar information contains five pieces of data: the position (x, y, z) of the target in the 4D millimeter wave radar coordinate system, the speed v, and the radar cross section RCS.
[0015] Step 2.2, measure the position coordinates (x', y', z') of the 4D millimeter wave radar in the camera coordinate system, i.e. the translation vector t = [x' y' z'] T, and 4D millimeter-wave radar rotates around the monocular camera x, y, z axis angle α, β, γ;
[0016] 4D millimeter-wave radar rotates around the monocular camera x axis rotation matrix
[0017] 4D millimeter-wave radar rotates around the monocular camera y axis rotation matrix
[0018] 4D millimeter-wave radar rotates around the monocular camera z axis rotation matrix
[0019] Calculate the total rotation matrix R of 4D millimeter-wave radar to monocular camera R x R y R z ;
[0020] Then the transformation matrix Tr_velo_to_cam of 4D millimeter-wave radar coordinate system to monocular camera coordinate system is as follows:
[0021]
[0022] Step 2.3, according to the position (x, y, z) of the target in the 4D millimeter-wave radar coordinate system, the internal parameter matrix P0 and the correction matrix R0_rect provided by the monocular camera factory, and the transformation matrix Tr_velo_to_cam of the 4D millimeter-wave radar coordinate system to the monocular camera coordinate system, the coordinates of the target in the image coordinate system are calculated, wherein dx and dy represent the physical size of the target in the x and y axes of the image coordinate system respectively, and d represents the physical size of each pixel:
[0023]
[0024] Step 2.4, according to the position (x, y, z) of the target in the 4D millimeter-wave radar coordinate system, the actual distance dis is calculated, and the corresponding velocity vel and radar cross section rcs information are superimposed in the channel dimension to generate a three-channel 4D millimeter-wave radar pseudo image I[h, w, c], (h, w) is a pixel coordinate on the pseudo image, c is the channel dimension, I[h, w, dis] represents the intensity of the coordinate (h, w) in the distance channel, I[h, w, rcs] represents the intensity of the coordinate (h, w) in the radar cross section channel, and I[h, w, vel] represents the intensity of the coordinate (h, w) in the velocity channel:
[0025]
[0026] Further, step 3, a traffic target detection model based on the fusion of 4D millimeter-wave radar pseudo image and monocular vision image is constructed, and the specific method is:
[0027] (a) Network architecture design
[0028] The traffic target detection model based on fusion of 4D millimeter wave radar pseudo image and monocular vision image comprises a pre-encoder module, a four-stage feature extraction module, two parallel detection head and segmentation head modules, and each stage of the feature extraction module comprises two parallel feature extraction modules, i.e., a dynamic convolution module and a self-attention module, and a serial connection module;
[0029] (b) Pre-encoder module design
[0030] 4D millimeter wave radar pseudo image f i After coarse extraction by the pre-encoder module, a three-channel radar feature map f o is generated. i The pre-encoder module takes f a with a shape of (Batch, 3, 320, 320) as input, first uses a 3x3 average pooling layer to take the average value in the distance, speed, and radar cross-sectional area three channels to generate f a with a shape of (Batch, 3, 320, 320), then uses a 1x1 convolution layer, adds a BatchNormalization batch normalization layer and a ReLU activation function layer to reduce the feature dimension of f c to generate f c with a shape of (Batch, 1, 320, 320), then adds a 3x3 max pooling layer to extract the edge features of f d to generate f d with a shape of (Batch, 1, 320, 320), finally uses a 1x1 convolution layer and a BatchNormalization batch normalization layer to restore the dimension of f m to generate f m with a shape of (Batch, 3, 320, 320), and then adds the original 4D millimeter wave radar pseudo image f r to generate a radar feature map f i with a shape of (Batch, 3, 320, 320) after the ReLu layer. o
[0031] f c = ReLu(BN(Conv 1×1 (AvgPool(f i )))
[0032] f d = DeformableConv(f c )
[0033] f m = MaxPool(f d )
[0034] f o = ReLu(f i + BN(Conv 1×1 (f m ))
[0035] (c) Feature extraction module design
[0036] The feature extraction module has four stages, each of which is composed of three sub-modules. In the first stage, on the one hand, the radar feature map f o of (Batch, 3, 320, 320) is generated through a 7x7 convolution module and a dynamic convolution module to generate a radar high-dimensional low-resolution feature map f o of (Batch, 16, 80, 80), and on the other hand, the picture data i of (Batch, 3, 320, 320) is simultaneously generated through a 7x7 convolution module and a self-attention module to generate a visual high-dimensional low-resolution feature map i' of (Batch, 16, 80, 80), and then f o ' and i' are added in the channel dimension to generate a radar-visual fusion feature map r1 of (Batch, 32, 80, 80) through a connection module; the radar-visual fusion feature maps r1, r2, r3 of the subsequent three stages are first divided into two parts in the channel dimension, i.e., f o1 and i1 of (Batch, 16, 80, 80), f o2 and i2 of (Batch, 24, 40, 40), f o3 and i3 of (Batch, 48, 20, 20), and similarly, one part is generated through a 7x7 convolution module and a dynamic convolution module, and the other part is generated through a 7x7 convolution module and a self-attention module, and then the two parts are connected together through a connection module to generate a radar-visual fusion feature map r2 of (Batch, 48, 40, 40), a radar-visual fusion feature map r3 of (Batch, 96, 20, 20), and a radar-visual fusion feature map r4 of (Batch, 176, 10, 10);
[0037] For the (B, C, W, H) of the dynamic convolution module and the self-attention module in each stage of the four stages, it refers to (Batch, 16, 80, 80) generated by f o and i through a 7x7 convolution module, (Batch, 24, 40, 40) generated by f o1 and i1 through a 7x7 convolution module, and fo2 and i2 through a 7x7 convolutional module to generate (Batch, 48, 20, 20), f o3 and i3 through a 7x7 convolutional module to generate (Batch, 88, 10, 10);
[0038] The data f with input shape (B, C, W, H) in the dynamic convolutional module is first passed through a 3x3 average pooling shape invariant, then a fully connected layer and ReLu are used to change the shape to (B, CxWxH), then a fully connected layer and Softmax are used again to generate normalized attention weights Conv1,...,ConvC for C 3x3 convolutional kernels, and the C 3x3 convolutional kernels are added by channel to become a 3x3 grouped convolution with group number C, and f is passed through this grouped convolution and added with BN and ReLu layers to finally generate the output of the dynamic convolutional module with shape (B, C, W, H);
[0039] The self-attention module generates a query Q through a 1x1 convolution, keeping the original shape (B, C, W, H), and generates a key K and a value V through another 1x1 convolution, doubling the channel number to shape (B, 2C, W, H), and splitting into two tensors: K(B, C, W, H), V(B, C, W, H), performing matrix multiplication on Q and K to get attention scores attn, applying Softmax to the last dimension of the attention scores attn to normalize the attention weights, and performing matrix multiplication on the attention weights attn and the value V to get a new feature representation, and finally outputting shape (B, C, W, H);
[0040] The input of each stage in the four-stage connection module is the output of the dynamic convolutional module and the self-attention module added along the channel dimension, i.e., the shape of the input r is (Batch, 32, 80, 80), (Batch, 48, 40, 40), (Batch, 96, 20, 20), (Batch, 176, 10, 10), first performing a 3x3 convolution on each channel of (B, C, W, H) independently, then introducing a batch normalization layer BN and a nonlinear activation function ReLu without changing the shape, then performing a 1x1 pointwise convolution to convert the channel number to 1 / 8 of C, further introducing a batch normalization layer BN and a nonlinear activation function ReLu, applying a 1x1 convolution and a batch normalization layer BN and a nonlinear activation function ReLu to restore the channel number to C, and finally adding the input r to output a (B, C, W, H) multi-sensor fusion feature map;
[0041] (d) Decoder module design
[0042] The decoder module is composed of two modules, the detection head and the drivable area segmentation head, and the radar and vision fusion feature map r4 of (Batch, 176, 10, 10) is changed in dimension to r4' of (Batch, 96, 64, 64) before entering the decoder module.
[0043] In the detection head module, r4' is first changed in dimension to 256 through Conv1x1, and then sent into two branches respectively. The output shape of the class prediction branch is HxWxC, the output shape of the position prediction branch is HxWx4, and the output shape of the IoU prediction branch is HxWx1, where H and W are the sizes of the feature map, and C is the number of object categories.
[0044] The drivable area segmentation head module is composed of three up-sampling modules. The first up-sampling module uses 2x up-sampling to double the spatial resolution of r4' from a low-resolution feature map with a spatial size of 96x64x64 to r'4'; the second up-sampling module up-samples r'4' from 128x128 to 256x256, and reduces the number of channels from 48 to 32 to generate r'4''; and the third up-sampling module up-samples r'4'' to a high-resolution feature map with a spatial size of 512x512, and converts it into a drivable area segmentation output with a final output channel number of 2, where 0 represents that the pixel point is background, and 1 represents that the pixel point is a drivable area.
[0045] Further, in step 5, the distance and speed of the target identified in step 4 are output according to the kernel function density estimation method, and the specific method is as follows:
[0046] Step 5.1, each target box is defined as (x_min, y_min, x_max, y_max), and each 4D millimeter wave radar point is represented by (x, y) coordinates. For each target box, the points located in the box are filtered out;
[0047] Step 5.2, the kernel function K is selected as a Gaussian function, the distance bandwidth h x is 1, the velocity bandwidth h v is also 1, and the density estimation is calculated, where n is the number of radar points in the target box, and e represents the base of natural logarithm;
[0048]
[0049]
[0050] Step 5.3, find the point with the maximum density value as the estimated value of the point in the target box;
[0051] Step 5.4, repeat steps 5.2 and 5.3 until the distance and speed values of all targets are calculated.
[0052] A traffic target detection system based on fusion of 4D millimeter wave radar pseudo image and monocular vision image, based on the traffic target detection method, realizing traffic target detection based on fusion of 4D millimeter wave radar pseudo image and monocular vision image.
[0053] A computer device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, when the processor executes the computer program, based on the traffic target detection method, realizing traffic target detection based on fusion of 4D millimeter wave radar pseudo image and monocular vision image.
[0054] A computer readable storage medium, having a computer program stored thereon, when the computer program is executed by a processor, based on the traffic target detection method, realizing traffic target detection based on fusion of 4D millimeter wave radar pseudo image and monocular vision image.
[0055] Compared with the prior art, the present application has the following advantages: (1) a pre-encoder for feature extraction of sparse radar data is designed, which realizes filtering of noise and redundant information in the original data and retains key features useful for the task; (2) the coarsely extracted radar feature map is again subjected to backbone network extraction together with the RGB image for feature-level fusion, which greatly improves the target detection accuracy; (3) the multi-task model realizes the distinction of drivable area detection task for the difference between the coming and going lanes. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 A traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image according to the present application;
[0057] Figure 2 A detection effect diagram for distinguishing between the coming and going lanes according to the present application;
[0058] Figure 3 A target detection effect diagram in a complex urban occlusion environment according to the present application;
[0059] Figure 4 A pre-encoder module architecture diagram according to the present application;
[0060] Figure 5 A feature extraction module architecture diagram according to the present application;
[0061] Figure 6 A dynamic convolution module architecture diagram according to the present application;
[0062] Figure 7 A self-attention module architecture diagram according to the present application;
[0063] Figure 8 A connection module architecture diagram according to the present application;
[0064] Figure 9 Architecture diagram of the detection head module of the present application;
[0065] Figure 10 Architecture diagram of the segmentation head module of the present application;
[0066] Figure 11a Target detection performance comparison chart of the present application;
[0067] Figure 11b Driveable area detection performance comparison chart of the present application. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0069] As shown in Figure 1 , the present application proposes a traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image. Because two kinds of sensors, 4D millimeter wave radar and monocular camera, are used, the method of the present application will not be affected by shadow and glare on the picture edge. In Figure 2 , when there is no lane line mark, ground information can be supplemented by z-axis and RCS (radar cross section) information on the 4D millimeter wave radar, and the distinction between incoming and outgoing can also be assisted by the speed component on the 4D millimeter wave radar. As shown in Figure 3 , the most intuitive advantage of the method of the present application is that in the case of dense targets, three frames of 4D millimeter wave radar point data can be mapped to 4D millimeter wave radar pseudo image, so that the high robustness of the model pre-encoder and the feature extraction module can be used to detect the completely occluded traffic target in the current picture. The method includes the following steps:
[0070] (I) Time synchronization of 4D millimeter wave radar and vision image
[0071] The monocular camera and the 4D millimeter wave radar are started at the same time, and the 4D millimeter wave radar data and the monocular camera data are read in a multi-threaded manner. One sub-thread reads the video image in real time and maintains the number of images in the cache pool to be 1. When the next video stream image arrives, the image in the cache pool is updated to ensure that the image in the cache pool is the current latest frame. Another sub-thread reads the 4D millimeter wave radar data, and samples the data based on the low working frequency of the 4D millimeter wave radar sensor. When the 4D millimeter wave radar data arrives, the latest image in the video sub-thread is read.
[0072] (II) Spatial synchronization of 4D millimeter wave radar and vision image
[0073] Step 2.1, superimpose the data of the first two frames and the current frame collected in step 1, only increase the number of 4D millimeter wave radar information when superimposed, each 4D millimeter wave radar information contains five data of target position (x, y, z) in 4D millimeter wave radar coordinate system, velocity v, radar cross section RCS;
[0074] Step 2.2, measure the position coordinates (x', y', z') of 4D millimeter wave radar in camera coordinate system, that is, the translation vector t = [x' y' z'] T , and the angles a, b, g of 4D millimeter wave radar rotating around the x, y, z axes of monocular camera.
[0075] Rotation matrix of 4D millimeter wave radar around the x axis of monocular camera
[0076] Rotation matrix of 4D millimeter wave radar around the y axis of monocular camera
[0077] Rotation matrix of 4D millimeter wave radar around the z axis of monocular camera
[0078] Calculate the total rotation matrix R of 4D millimeter wave radar to monocular camera R = R x R y R z
[0079] Then the transformation matrix Tr_velo_to_cam of 4D millimeter wave radar coordinate system to monocular camera coordinate system is as follows:
[0080]
[0081] Step 2.3, according to the position (x, y, z) of the target in the 4D millimeter wave radar coordinate system, the internal parameter matrix P0 and the correction matrix R0_rect provided by the monocular camera factory, and the transformation matrix Tr_velo_to_cam of the 4D millimeter wave radar coordinate system to the monocular camera coordinate system, the coordinates of the target in the image coordinate system are calculated, wherein dx and dy represent the physical size of the target in the x and y axes of the image coordinate system respectively, and d represents the physical size of each pixel:
[0082]
[0083] Step 2.4. Calculate the actual distance dis according to the position (x, y, z) of the target in the 4D millimeter-wave radar coordinate system, and superimpose the corresponding velocity vel and radar cross-section rcs information in the channel dimension to generate a three-channel 4D millimeter-wave radar pseudo-image. (h, w) in I[h, w, c] is a pixel coordinate on the pseudo-image, and c is the channel dimension. I[h, w, dis] represents the intensity of coordinate (h, w) in the distance channel. I[h, w, rcs] represents the intensity of coordinate (h, w) in the radar cross-section channel. I[h, w, vel] represents the intensity of coordinate (h, w) in the velocity channel:
[0084]
[0085] (III) Construction of a traffic target detection model based on fusion of 4D millimeter-wave radar pseudo-images and monocular vision images
[0086] (a) Network architecture design
[0087] The traffic target detection model based on fusion of 4D millimeter-wave radar pseudo-images and monocular vision images includes a pre-encoder module, a four-stage feature extraction module, two parallel detection head and segmentation head modules. Each stage of the feature extraction module consists of two parallel feature extraction modules (dynamic convolution module and self-attention module) and a serial connection module;
[0088] (b) Pre-encoder module design
[0089] 4D millimeter-wave radar pseudo-image f i After coarse extraction by the pre-encoder module, a three-channel radar feature map f o is generated. The pre-encoder module, as shown in Figure 4 , takes f i with an input shape of (Batch, 3, 320, 320) as input. First, a 3x3 average pooling layer is used to take the mean value in the distance, velocity, and radar cross-section channels to generate f a with a shape of (Batch, 3, 320, 320). Batch represents the batch size of the input data, and 320x320 represents the spatial size of the input data. Then, a 1x1 convolution is used with the addition of a Batch Normalization layer and a ReLU activation function layer to reduce the feature dimension of f a to generate f c with a shape of (Batch, 1, 320, 320). Second, DeformableConv is used to deform the convolution kernel of f c to generate f d with a shape of (Batch, 1, 320, 320). Then, a 3x3 max pooling layer is added to extract f df of edge feature generation (Batch, 1, 320, 320) m Finally, a 1x1 convolutional layer and a batch normalization layer are used to recover f. m The dimension of f is generated by (Batch, 3, 320, 320). r Later compared with the original 4D millimeter-wave radar pseudo-image f i The radar feature map f is generated by adding the features and then passing them through a ReLU layer (Batch, 3, 320, 320). o ;
[0090] f c =ReLu(BN(Conv) 1×1 (AvgPool(f i ))))
[0091] f d =DeformableConv(f c )
[0092] f m =MaxPool(f d )
[0093] f o =ReLu(f i +BN(Conv 1×1 (f m )))
[0094] (c) Feature extraction module design
[0095] like Figure 5 As shown, the feature extraction module has four stages, each consisting of three sub-modules. The first stage extracts radar feature maps f from (Batch, 3, 320, 320). o A high-dimensional, low-resolution radar feature map f is generated by a 7×7 convolutional module and a dynamic convolutional module, with a batch size of (Batch, 16, 80, 80). o On the other hand, the image data i of (Batch, 3, 320, 320) is processed through a 7×7 convolutional module and a self-attention module to generate a visual high-dimensional low-resolution feature map i' of (Batch, 16, 80, 80), and then f o After adding ' and i' along the channel dimension, the connection module generates a (Batch, 32, 80, 80) radar-visual fusion feature map r1. The subsequent three stages first divide the radar-visual fusion feature maps r1, r2, and r3 from the previous stage into two equal parts along the channel dimension, namely (Batch, 16, 80, 80) f. o1 And i1, f of (Batch, 24, 40, 40)o2 and i2, (Batch, 48, 20, 20) of f o3 and i3, also one part goes through a 7x7 convolution module and dynamic convolution module, the other part goes through a 7x7 convolution module and self-attention module, then the two parts are connected through the connection module to generate the radar and visual fusion feature map r2 of (Batch, 48, 40, 40), the radar and visual fusion feature map r3 of (Batch, 96, 20, 20), and the radar and visual fusion feature map r4 of (Batch, 176, 10, 10);
[0096] (B, C, W, H) of the dynamic convolution module and the self-attention module for each stage in the four stages respectively refers to f o and i through the 7x7 convolution module to generate (Batch, 16, 80, 80), f o1 and i1 through the 7x7 convolution module to generate (Batch, 24, 40, 40), f o2 and i2 through the 7x7 convolution module to generate (Batch, 48, 20, 20), f o3 and i3 through the 7x7 convolution module to generate (Batch, 88, 10, 10).
[0097] The dynamic convolution module is as shown in Figure 6 The input data f with shape (B, C, W, H) is first passed through a 3x3 average pooling shape invariant, then a fully connected layer and ReLu are used to change the shape to (B, CxWxH), and then a fully connected layer and Softmax are used again to generate normalized attention weights Conv1, …, ConvC for C 3x3 convolution kernels. Adding the C 3x3 convolution kernels by channel becomes a 3x3 grouped convolution with C groups, and the f is passed through this grouped convolution and added with BN and ReLu layers to finally generate the output of the dynamic convolution module with shape (B, C, W, H);
[0098] The self-attention module is as shown in Figure 7 A 1x1 convolution is used to generate the query Q, keeping the original shape (B, C, W, H), and another 1x1 convolution is used to generate the key K and the value V, doubling the number of channels to (B, 2C, W, H). Split into two tensors: K (B, C, W, H), V (B, C, W, H). Perform matrix multiplication on Q and K to get the attention score attn. Apply Softmax to the last dimension of the attention score attn to normalize the attention weight. Perform matrix multiplication between the attention weight attn and the value V to get the new feature representation, and finally output with shape (B, C, W, H);
[0099] The connection module is as shown in Figure 8As shown, the input to each of the four stages is the sum of the outputs of the dynamic convolution module and the self-attention module along the channel dimension, i.e., the shape of the input r is (Batch, 32, 80, 80), (Batch, 48, 40, 40), (Batch, 96, 20, 20), (Batch, 176, 10, 10). Figure 7 As shown, firstly, each channel of (B,C,W,H) is independently convolved with a 3×3 convolution, and then a batch normalization layer (BN) and a non-linear activation function (ReLU) are introduced to maintain the shape. Then, a 1×1 pointwise convolution is performed to convert the number of channels to 1 / 8 of C. A batch normalization layer (BN) and a non-linear activation function (ReLU) are further introduced. The number of channels is restored to C by applying a 1×1 convolution, a batch normalization layer (BN), and a non-linear activation function (ReLU). Finally, the output is the Rado-Vision fusion feature map of (B,C,W,H) by adding it to the input r.
[0100] (d) Decoder module design
[0101] The decoder module consists of two parallel modules: a detection head and a drivable area segmentation head. Before entering the decoder module, the radar-visual fusion feature map r4 of (Batch, 176, 10, 10) is changed to r4' of (Batch, 96, 64, 64).
[0102] Detection head module such as Figure 9 As shown, r4' first changes its dimension to 256 using Conv1×1, and then feeds it into two branches. The output shape of the category prediction branch is H×W×C, the output shape of the position prediction branch is H×W×4, and the output shape of the IoU (Intersection over Union) prediction branch is H×W×1. Here, H and W are the feature map sizes, and C is the number of object categories.
[0103] Driving area segmentation head module such as Figure 10 As shown, the algorithm consists of three upsampling modules. The first upsampling module upsamples r4' from a low-resolution feature map with spatial dimensions of 96×64×64 by a factor of 2 to r'4' with double the spatial resolution. The second upsampling module upsamples r'4' from 128×128 to 256×256 and reduces the number of channels from 48 to 32 to generate r'4". The third upsampling module upsamples r'4" to a high-resolution feature map of 512×512 and converts it into a drivable region segmentation output with a final output of 2 channels, where 0 indicates that the pixel is background and 1 indicates that the pixel is a drivable region.
[0104] (iv) Traffic target detection model prediction based on fusion of 4D millimeter-wave radar pseudo-images and monocular vision images
[0105] The 4D millimeter wave radar data collected on site are preprocessed in the same way as steps 1 and 2, and are sent into the model trained in step 3 together with monocular vision image data to generate a target frame and a drivable area.
[0106] (V) Target distance and speed measurement based on kernel function density estimation
[0107] The distance and speed of the target identified in step 4 are measured. First, each target frame is defined as (x_min, y_min, x_max, y_max), and each 4D millimeter wave radar point is represented by (x, y) coordinates. For each target frame, the points located in the frame are screened out;
[0108] Secondly, the kernel function K is selected as a Gaussian function, the distance bandwidth h x is 1, the velocity bandwidth h v is also 1, and the density estimation is calculated, where n is the number of radar points in the target frame, and e represents the base of natural logarithm;
[0109]
[0110]
[0111] Then, the point with the maximum density value is found as the estimated value of the points in the target frame;
[0112] Finally, the above steps are repeated until the distance and speed values of all targets are calculated.
[0113] The application also proposes a traffic target detection system based on fusion of 4D millimeter wave radar pseudo images and monocular vision images, which realizes traffic target detection based on fusion of 4D millimeter wave radar pseudo images and monocular vision images based on the traffic target detection method.
[0114] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the traffic target detection based on fusion of 4D millimeter wave radar pseudo images and monocular vision images is realized based on the traffic target detection method.
[0115] A computer readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, the traffic target detection based on fusion of 4D millimeter wave radar pseudo images and monocular vision images is realized based on the traffic target detection method.
[0116] Embodiments
[0117] In order to verify the effectiveness of the scheme of the application, the following experiments are performed.
[0118] The application is trained on a VoD dataset containing a total of 6400 images randomly allocated into a training set and a validation set at a ratio of 4:1. In the target detection task, three target detection categories are defined, namely motor vehicles, non-motor vehicles and pedestrians. In the model training stage, an initial learning rate of 0.03 is adopted, and iteration is performed for 100 training cycles. At the same time, the batch size is set to 16, and the learning rate is dynamically adjusted during the training process by using the SGD optimizer and the cosine decay strategy. The training of all models is completed on a hardware platform equipped with a GTX2080Ti graphics card (12G video memory), 128G memory and an Intel(R) Xeon(R) Gold5218 sixteen-core processor.
[0119] The detection performance of the final traffic target detection method based on fusion of 4D millimeter wave radar pseudo images and monocular vision images in an actual scene is shown in Figure 11a and Figure 11b . Wherein Figure 11a represents a target detection performance comparison chart, Figure 11b represents a drivable area detection performance comparison chart. The results show that the detection accuracy of the traffic target detection method based on fusion of 4D millimeter wave radar pseudo images and monocular vision images is 0.3 percentage points higher than that of YOLOv8n for cars, 9.4 percentage points higher for non-motor vehicles, and more excellent for pedestrian detection, and the total mAP is also 5.47 percentage points higher. For drivable area detection, the mIoU can reach as high as 95.7%, and the parameter quantity and required computing performance are still superior to the pure vision segmentation network PSPNet, which is more suitable for embedded devices. Therefore, the experiment proves that the application not only optimizes the parameter quantity and model calculation amount, making it more suitable for edge computing end, no longer requiring a large amount of computing power, and can further reduce deployment costs, but also has higher traffic target detection accuracy.
[0120] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not contradict, they should be considered within the scope of the present application.
[0121] The above-described embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the present application. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image, characterized in that, The method comprises the following steps: Step 1, synchronize the starting time of 4D millimeter wave radar and visual image, set consistent sampling interval to ensure that the sampling time of 4D millimeter wave radar data and monocular visual data is the same; Step 2, synchronize the space of 4D millimeter wave radar and visual image, calculate the actual distance according to the position of the target in the 4D millimeter wave radar coordinate system, and superimpose the corresponding speed and radar cross section area information in the channel dimension to generate a three-channel 4D millimeter wave radar pseudo image; Step 3, use a pre-encoder to coarsely extract the 4D millimeter wave radar information features, generate a three-channel radar feature map, pass the radar feature map through a dynamic convolution module, and process the picture data using a self-attention module, connect the outputs of the dynamic convolution module and the self-attention module together according to the channel to fuse them, obtain the features after the fusion of the radar and the image, and use an anchor-free method to directly regress the coordinates of the bounding box from the features after the fusion of the radar and the image, and a drivable area segmentation head restores the output segmentation feature map to two channel dimensions with the size of the original picture after three upsampling processes, generates the probability of the drivable area and the background corresponding to each pixel in the input image, and constructs a traffic target detection model based on the fusion of the 4D millimeter wave radar pseudo image and the monocular visual image; Step 4, the 4D millimeter wave radar point cloud data collected in the field is preprocessed in the same way as steps 1 and 2, and is sent into the model trained in step 3 together with the monocular visual image data to generate the target box and the drivable area; Step 5, according to the kernel function density estimation method, the distance and speed of the target identified in step 4 are output.
2. The traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image according to claim 1, characterized in that, Step 1, synchronize the time of 4D millimeter wave radar and visual image, the specific method is: Start the monocular camera and 4D millimeter wave radar at the same time, read the 4D millimeter wave radar data and monocular camera data in a multi-threaded manner, one sub-thread reads the video image in real time and maintains the number of images in the cache pool to be 1, updates the image in the cache pool when the next video stream image arrives to ensure that the image in the cache pool is the latest frame; another sub-thread reads the 4D millimeter wave radar data, and samples the data based on the low working frequency of the 4D millimeter wave radar sensor, and reads the latest image in the video sub-thread when the 4D millimeter wave radar data arrives. 3.The traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image according to claim 1, characterized in that, Step 2, synchronize the space of 4D millimeter wave radar and visual image, the specific method is: Step 2.1, superimpose the data of the first two frames and the current frame of the 4D millimeter wave radar collected in step 1, only increase the number of 4D millimeter wave radar information when superimposing, and each 4D millimeter wave radar information includes the position (x, y, z) of the target in the 4D millimeter wave radar coordinate system, the speed v, and the radar cross section area RCS five data; Step 2.2, measure the position coordinates of the 4D mmWave radar in the camera coordinate system , i.e. the translation vector , and the angles of the 4D mmWave radar rotation around the monocular camera x, y, z axes ; 4D millimeter wave radar rotation matrix around monocular camera x-axis ; 4D millimeter wave radar rotation matrix around monocular camera y-axis ; 4D millimeter wave radar rotation matrix around monocular camera z-axis ; Computing the total rotation matrix of a 4D millimeter wave radar to monocular camera ; Then the transformation matrix from the 4D millimeter wave radar coordinate system to the monocular camera coordinate system is in the form of: ; Step 2.
3. Calculate the target's position in the image coordinate system, where and the correction matrix and the transformation matrix from the 4D millimeter-wave radar coordinate system to the monocular camera coordinate system Calculate the target's position in the image coordinate system, where and represent the physical dimensions of the target in the x-axis and y-axis directions of the image coordinate system, respectively, represent the physical size of each pixel point: ; Step 2.4, calculate the actual distance dis according to the position (x, y, z) of the target in the 4D millimeter wave radar coordinate system, and superimpose the corresponding velocity vel and radar cross section rcs information in the channel dimension to generate a three-channel 4D millimeter wave radar pseudo image, (h, w) in the pseudo image is a pixel coordinate, c is the channel dimension, I[h, w, dis] represents the intensity of the coordinate (h, w) in the distance channel, I[h, w, rcs] represents the intensity of the coordinate (h, w) in the radar cross section channel, and I[h, w, vel] represents the intensity of the coordinate (h, w) in the velocity channel: 。 4. The traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image according to claim 1, characterized in that, Step 3, construct a traffic target detection model based on the fusion of 4D millimeter wave radar pseudo image and monocular vision image, the specific method is: (a) Network architecture design The traffic target detection model based on the fusion of 4D millimeter wave radar pseudo image and monocular vision image comprises a pre-encoder module, a four-stage feature extraction module, two parallel detection head and segmentation head modules, and each stage of the feature extraction module comprises two parallel feature extraction modules, i.e. a dynamic convolution module and a self-attention module, and a serial connection module; (b) Pre-encoder module design 4D millimeter wave radar pseudo image After coarse extraction by the pre-encoder module, a three-channel radar feature map is generated ; the pre-encoder module takes an input shape of (Batch, 3, 320, 320) First, a 3x3 average pooling layer is used to take the mean value in the distance, velocity, and radar cross section three channels to generate a (Batch, 3, 320, 320) shape , Batch represents the batch size of the input data, and 320x320 represents the spatial size of the input data; then a 1x1 convolution is used to increase the BatchNormalization batch normalization layer and the ReLU activation function layer to reduce feature dimension to generate a (Batch, 1, 320, 320) shape Second, DeformableConv is used to deform the convolution kernel to generate a (Batch, 1, 320, 320) shape Then, a 3x3 max pooling layer is added to extract the edge features to generate a (Batch, 1, 320, 320) shape Finally, a 1x1 convolution layer and a BatchNormalization batch normalization layer are used to restore the dimension to generate a (Batch, 3, 320, 320) shape After that, it is added to the original 4D millimeter wave radar pseudo image After the ReLu layer, a (Batch, 3, 320, 320) radar feature map is generated ; ; (c) Feature extraction module design The feature extraction module consists of four stages, each composed of three sub-modules. The first stage extracts radar feature maps of size (Batch, 3, 320, 320). A high-dimensional, low-resolution radar feature map (Batch, 16, 80, 80) is generated through a 7×7 convolutional module and a dynamic convolutional module. On the other hand, simultaneously, (Batch, 3, 320, 320) image data Visual high-dimensional low-resolution feature maps (Batch, 16, 80, 80) are generated through a 7×7 convolutional module and a self-attention module. Then and The radar-visual fusion feature map (Batch, 32, 80, 80) is generated by summing the features according to the channel dimension and then passing them through the connection module. The subsequent three stages will first integrate the radar-visual fusion feature map from the previous stage. Divide the data into two equal parts based on the channel dimension, namely (Batch, 16, 80, 80). and (Batch, 24, 40, 40) and (Batch, 48, 20, 20) and One part is processed through a 7×7 convolutional module and a dynamic convolutional module, while the other part is processed through a 7×7 convolutional module and a self-attention module. These two parts are then concatenated together by a connection module to generate a (Batch, 48, 40, 40) radar-visual fusion feature map. The radar-visual fusion feature map of (Batch, 96, 20, 20) The radar-visual fusion feature map of (Batch, 176, 10, 10) ; (B, C, W, H) for the dynamic convolution module and self-attention module of each stage in the four stages respectively refers to and (Batch, 16, 80, 80) generated by the 7x7 convolution module, and (Batch, 24, 40, 40) generated by the 7x7 convolution module, and (Batch, 48, 20, 20) generated by the 7x7 convolution module, and (Batch, 88, 10, 10) generated by the 7x7 convolution module; Data with input shape (B, C, W, H) in dynamic convolution module First, shape is invariant by 3x3 average pooling, then use a fully connected layer and ReLu to change shape to (B, CxWxH), then use again a fully connected layer and Softmax to generate normalized attention weights for C 3x3 convolution kernels Conv1,..., ConvC, add these C 3x3 convolution kernels by channel to become 3x3 grouped convolution with group number C, and Finally generate dynamic convolution module output with shape (B, C, W, H) by this grouped convolution and add BN and ReLu layers; The self-attention module generates a query Q through a 1x1 convolution, maintains the original shape (B, C, W, H), generates a key K and a value V through another 1x1 convolution, multiplies the channel number by 2, and the shape becomes (B, 2C, W, H). Split into two tensors: K(B, C, W, H), V(B, C, W, H), perform matrix multiplication on Q and K to get attention score attn, apply Softmax to the last dimension of attention score attn to normalize attention weight, and perform matrix multiplication on attention weight attn and value V to get new feature representation, and finally output shape (B, C, W, H); The input of each stage of the four-stage connection module is respectively the output of the dynamic convolution module and the self-attention module added along the channel dimension, that is, the input The shape of (Batch, 32, 80, 80), (Batch, 48, 40, 40), (Batch, 96, 20, 20), (Batch, 176, 10, 10) is first independently subjected to 3×3 convolution on each channel of (B, C, W, H), and then a batch normalization layer BN and a nonlinear activation function ReLu are introduced, the shape is not changed, then 1×1 point convolution is performed, the channel number is converted to 1 / 8 of C, a batch normalization layer BN and a nonlinear activation function ReLu are further introduced, 1×1 convolution and a batch normalization layer BN and a nonlinear activation function ReLu are applied to restore the channel number to C, and finally added with the input The fused feature map of the radar and vision is output (B, C, W, H). (d) Decoder module design The decoder module is composed of two modules in parallel, detection head and drivable area segmentation head. Before entering the decoder module, the radar and vision fusion feature map of (Batch, 176, 10, 10) is first processed by the detection head The dimension is changed to (Batch, 96, 64, 64) ; In the detection head module First, the dimension is changed to 256 by Conv1x1, and then sent into two branches respectively. The output shape of the category prediction branch is HxWxC, the output shape of the position prediction branch is HxWx4, and the output shape of the IoU prediction branch is HxWx1, wherein H and W are the size of the feature map, and C is the number of object categories. The drivable area segmentation head module consists of three upsampling modules. The first upsampling module will... From a low-resolution feature map with spatial dimensions of 96×64×64, upsampling was performed by 2x to double the spatial resolution. The second upsampling module will Upsampled from 128×128 to 256×256, and reduced the number of channels from 48 to 32 to generate The third upsampling module will The high-resolution feature map upsampled to 512×512 is converted into a drivable region segmentation output with a final output of 2 channels, where 0 indicates that the pixel is the background and 1 indicates that the pixel is the drivable region.
5. The traffic target detection method based on fusion of 4D millimeter wave radar pseudo image and monocular vision image according to claim 1, characterized in that, Step 5, output the distance and velocity of the target according to the kernel function density estimation method for the target frame identified in step 4, the specific method is: Step 5.1, each target frame is defined as (x_min, y_min, x_max, y_max), and each 4D millimeter wave radar point is represented by (x, y) coordinates. For each target frame, filter out the points located in the frame; Step 5.2, select the kernel function K as Gaussian function, distance bandwidth as 1, velocity bandwidth also as 1, compute the density estimate, where n is the number of radar points in the target box, e denotes the base of the natural logarithm; ; ; Step 5.3, find the point with the maximum density value as the estimated value of the points in the target frame; Step 5.4, repeat steps 5.2 and 5.3 until the distance and velocity values of all targets are calculated.
6. A traffic target detection system based on fusion of 4D millimeter wave radar pseudo image and monocular vision image, characterized in that, The traffic target detection method according to any one of claims 1-5 realizes traffic target detection based on the fusion of 4D millimeter wave radar pseudo image and monocular vision image.
7. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the traffic target detection method according to any one of claims 1-5 is realized to realize traffic target detection based on the fusion of 4D millimeter wave radar pseudo image and monocular vision image. 8.A computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements traffic target detection based on fusion of 4D millimeter wave radar pseudo image and monocular vision image according to the traffic target detection method of any one of claims 1-5.
Citation Information
Patent Citations
Attention-based 4D millimeter wave radar and vision fusion method
CN116129234A
Fusion 3D target detection method based on 4D millimeter wave radar and image
CN117274749A