Target Detection Method Based on the Fusion of Millimeter-Wave Radar and Visible Light Images

By preprocessing radar data and building a converged target detection network, the problems of low detection efficiency and low accuracy in the prior art are solved, and efficient and accurate target detection in complex environments are achieved.

CN115830423BActive Publication Date: 2025-07-01XIDIAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211597596.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-07-01
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

In the prior art, the target detection method of fusing millimeter wave radar with visible light images has problems of low detection efficiency and low accuracy, especially under the influence of environmental factors such as light, rain, snow, and fog, the detection performance is affected.

Method used

Preprocessing the radar data to generate radar images, and build a target detection network based on the fusion of millimeter-wave radar and visible light images, including a radar feature extraction subnet, a visible light image feature extraction subnet and a RetinaNet network. Through the multimodal fusion module, fully utilizes radar information to enhance visible light image features, build a target detection network, and train it.

Benefits of technology

It improves the efficiency and accuracy of target detection, and can perform target detection more accurately in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830423B_ABST
    Figure CN115830423B_ABST
Patent Text Reader

Abstract

The present invention discloses a target detection method for millimeter-wave radar and visible light image fusion. The implementation method is as follows: preprocess the radar data to obtain a radar image; build a target detection network for millimeter-wave radar and visible light image fusion, including a feature extraction sub-network, an image fusion sub-network, and a RetinaNet network; input the preprocessed radar image and visible light image into the target detection network for millimeter-wave radar and visible light image fusion for training to obtain a trained network model; the radar data of the test set is preprocessed in the same way and then input into the trained model together with the visible light image for testing to obtain the target detection result. Compared with the single image detection method, the present invention has higher detection accuracy, can obtain better detection results, and can be used for target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data recognition, and further relates to a target detection method based on the fusion of millimeter-wave radar and visible light image in the field of recognition technology using electronic devices. The present invention can be used to detect targets in visible light images. Background Art

[0002] Target detection technology is a key and hot technology in the field of computer vision at present. In particular, image-based target detection technologies are emerging in an endless stream, and the continuous improvement of models has gradually improved the detection performance. However, affected by some environmental factors, such as light, rain, snow, and fog, the accuracy of target detection is affected to a certain extent. Millimeter-wave radar operates in the millimeter-wave band and has the characteristics of small size, light weight, low resolution, anti-interference, and anti-stealth. Most importantly, millimeter-wave radar has strong ability to penetrate fog, smoke, and dust, and has the characteristics of all-weather and all-time.

[0003] Generally, there are three ways to use the fusion of millimeter-wave radar and visible light image, namely decision-level fusion, data-level fusion, and feature-level fusion. The decision-level fusion method is to fuse the radar data and the prediction results of the visible light image. The data-level fusion method is to convert the radar data into the camera coordinate system, generate an area of interest according to the radar data, extract the corresponding features of the input image for the generated area of interest, and input the obtained features into the detection network to obtain the result. For the feature-level fusion method, it has been used more frequently in recent years. The radar data is converted into a specific form of data, and a feature extraction network is used to extract features from the visible light image and the radar data, and fusion is performed through a fusion network. Common fusion methods include element-level addition, multiplication, and splicing. The fused features are sent into the detection network to obtain target information. At present, for millimeter-wave radar data, common data format types include two-dimensional point cloud, three-dimensional point cloud, Range-Azimuth Map, and Range-Angle-Doppler (RAD) tensor.

[0004] Shanghai Jiao Tong University disclosed a millimeter-wave radar target detection method and system based on fused image features in its patent document "A Millimeter-Wave Radar Target Detection Method and System Based on Fused Image Features" (Patent Application No.: 202111288212.0, Publication No.: 114218999A). This method first obtains a 3D bird's-eye view feature map of the input image through an image feature processing module and inputs it into a radar data feature and image feature fusion module. Then, the radar data feature and image feature fusion module obtains a normalized radar feature map, which is fused with the 3D bird's-eye view feature map to obtain a fused feature map. Finally, based on the fused feature map, the target detection network of the target detection module is trained to obtain a trained model to improve the accuracy of target detection for autonomous vehicles. The disadvantage of this method is that the conversion of image features into a 3D bird's-eye view feature map is a complex process, ignoring the errors and deformations generated during the projection transformation, losing some features of the image, and taking a long time to calculate, reducing the efficiency of target detection.

[0005] Shuo Chang et al. proposed a new target detection method based on spatial attention fusion of millimeter-wave radar point cloud data and visible light images in their published paper "Spatial Attention Fusion for Obstacle Detection Using MmWave Radar and Vision Sensor" (Sensors (Basel, Switzerland), 2020, 20(4)). This method projects the radar point cloud onto the image and expands the points. The expansion method is to draw a circle with a certain point as the center and a specified length as the radius. All the pixel values within the coverage of this circle are made the same as the center point. Then, multiple kernels of different sizes are used for convolution to extract the attention matrix to enhance the image features. The proposed fusion method can be embedded in the feature extraction stage, effectively utilizing the features of millimeter-wave radar and visible light images. The disadvantage of this method is that since the radar point cloud data is generated through Fourier transform and constant false alarm rate (CFAR) process, the features are limited, and some information will be lost after projection, resulting in a decrease in the accuracy of target detection. Summary of the Invention

[0006] The object of the present invention is to propose a target detection method based on the fusion of millimeter-wave radar and visible light images to solve the problems of low target detection efficiency and low target detection accuracy for the above-mentioned existing technologies.

[0007] The idea of achieving the object of the present invention is to preprocess radar data to obtain a radar image, and build an object detection network based on the fusion of millimeter-wave radar and visible light image, including a radar feature extraction sub-network, a visible light image feature extraction sub-network, a radar and visible light image fusion sub-network, and a RetinaNet network. Among them, the radar and visible light image fusion network used in the present invention can make full use of radar information to enhance the features of visible light images. The preprocessed radar image and visible light image are input into the object detection network for the fusion of millimeter-wave radar and visible light image for training to obtain a trained network model. The radar data of the test set is preprocessed in the same way and then input into the trained model together with the visible light image for testing to obtain the object detection result.

[0008] The specific steps implemented by the present invention are as follows:

[0009] Step 1, preprocess the millimeter-wave radar data to generate a radar image:

[0010] Step 1.1, at the same moment when the vehicle-mounted radar obtains the radar echo signal data received by the radar sensor, the vehicle-mounted vision sensor obtains a visible light image corresponding to the radar echo signal data;

[0011] Step 1.2, convert the radar echo signal data into matrix A, where the rows in matrix A represent distance and the columns represent angle; after taking the modulus of matrix A and then normalizing it, a matrix B of size 128*128*1 is obtained, and a matrix C of 128*128*2 is obtained after normalizing matrix A. Matrix B and matrix C are concatenated to obtain matrix D;

[0012] Step 1.3, save matrix D as a radar image;

[0013] Step 2, generate a training set:

[0014] Generate annotation files in json format for both the radar image and the visible light image, and form a training set with the radar image, the visible light image, and their generated annotation files;

[0015] Step 3, construct a radar feature extraction sub-network:

[0016] Build a 10-layer feature extraction sub-network, and its structure is as follows: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer. Set the kernel sizes of the first to fifth convolutional layers to 7×7, 1×1, 1×1, 3×3, 1×1 respectively; set the number of kernels to 64, 256, 64, 64, 256 respectively. The first to fifth batch normalization layers are implemented using the FrozenBatchNorm function;

[0017] Step 4, construct a visible light feature extraction sub-network:

[0018] Build a 22-layer visible light feature extraction sub-network, and its structure is as follows: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer, the sixth convolutional layer, the sixth batch normalization layer, the seventh convolutional layer, the seventh batch normalization layer, the eighth convolutional layer, the eighth batch normalization layer, the ninth convolutional layer, the ninth batch normalization layer, the tenth convolutional layer, the tenth batch normalization layer, the eleventh convolutional layer, the eleventh batch normalization layer. Set the kernel sizes of the first to eleventh convolutional layers to 7×7, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; set the number of kernels of the first to eleventh convolutional layers to 64, 256, 64, 64, 256, 64, 64, 256, 64, 64, 256 respectively. Implement the first to eleventh batch normalization layers using the FrozenBatchNorm function;

[0019] Step 5, construct a radar and visible light image fusion sub-network:

[0020] The structure of the radar and visible light image fusion sub-network is as follows: the first multi-modal fusion module, the second multi-modal fusion module, the first convolutional block, the second convolutional block, the third convolutional block;

[0021] Step 5.1, the structures of the first and second multi-modal fusion modules are the same. The structure of each multi-modal fusion module is as follows: the first linear layer, the second linear layer, the third linear layer, the first activation layer, the second activation layer. Set the number of output neurons of the first to third linear layers of the first multi-modal fusion module to 256, and set the number of output neurons of the first to third linear layers of the second multi-modal fusion module to 512, 256, 256 respectively;

[0022] Step 5.2, the first convolutional block adopts the Stage2 structure of the Resnet50 network, the second convolutional block adopts the Stage3 structure of the Resnet50 network, and the third convolutional block uses the Stage4 structure of the Resnet50 network; the structures of the Stage2, Stage3, and Stage4 are all composed of 4, 6, and 3 Bottleneck structures in series, and each Bottleneck structure is in turn the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer;

[0023] Step 6, construct a target detection network:

[0024] After connecting the radar feature extraction sub-network and the visible light feature extraction sub-network in parallel, they are cascaded with the image fusion sub-network and the RetinaNet sub-network in sequence to form an object detection network for millimeter-wave radar and visible light image fusion;

[0025] Step 7, training the object detection network:

[0026] Input the training set into the object detection network for millimeter-wave radar and visible light image fusion, and use the stochastic gradient descent algorithm to iteratively update the weight values of the network, optimizing the total loss function of the network until it converges to obtain a trained object detection network;

[0027] Step 8, detecting the object:

[0028] Adopt the same processing method as in Step 1 to preprocess the radar echo signal data received by the vehicle-mounted radar sensor to obtain the radar image, and input the radar image and the visible light image generated by the vehicle-mounted vision sensor at the same moment of the radar echo signal data into the trained network to output the object detection result of millimeter-wave radar and visible light image fusion.

[0029] The present invention has the following advantages compared with the existing technologies:

[0030] First, since the present invention preprocesses the radar data, converts the radar data into a radar image, and obtains richer features, it overcomes the problem of low object detection efficiency caused by projection transformation of visible light images or radar data in the existing technologies, enabling the present invention to iterate more quickly during network training and improving the object detection efficiency.

[0031] Second, since the present invention uses a multi-modal fusion module to build an object detection network for millimeter-wave radar and visible light image fusion, it overcomes the problem of incomplete fusion of radar data and visible light images in the existing technologies, resulting in low detection accuracy, enabling the present invention to improve the object detection accuracy and perform object detection more accurately. Description of the Drawings

[0032] Figure 1 is a flowchart of the present invention.

[0033] Figure 2 is a network model diagram of the present invention.

[0034] Figure 3 is a schematic diagram of the multi-modal fusion module in the object detection network for millimeter-wave radar and visible light image fusion of the present invention. Detailed Embodiments

[0035] The present invention will be further described in detail below with reference to the drawings.

[0036] Referring to the attached Figure 1 , the implementation steps of the present invention are further described in detail.

[0037] Step 1, preprocess the millimeter-wave radar data to generate a radar image.

[0038] The millimeter-wave radar data in the embodiments of the present invention is downloaded from the public website CRUW dataset, which contains radar echo signal data received by the radar sensor and visible light images generated by the vision sensor. The visible light images are in the *.jpg format, and the radar echo signal data is in the *.npy format.

[0039] The expression of the radar echo signal data is a matrix A with a size of 128*128*2. After taking the modulus of matrix A and then normalizing it, a matrix B with a size of 128*128*1 is obtained.

[0040] After normalizing matrix A, a matrix C with a size of 128*128*2 is obtained.

[0041] Concatenate matrix B and matrix C to obtain a preprocessed matrix D with a size of 128*128*3.

[0042] Save matrix D as a radar image using the mp.imsave function.

[0043] Step 2, generate a training set and a test set.

[0044] Since the radar data and visible light images used in the embodiments of the present invention are obtained simultaneously by the radar sensor and the vision sensor after being well calibrated and synchronized, the two are in one-to-one correspondence. Therefore, the preprocessed radar images and visible light images are also in one-to-one correspondence.

[0045] Generate annotation files in json format for both the preprocessed radar images and visible light images, and form a training set with the preprocessed radar images, visible light images, and their generated annotation files.

[0046] Divide the samples in the sample set into a training set and a test set according to a ratio of 8:2.

[0047] Step 3, construct a radar feature extraction sub-network, and obtain the feature map of the input radar image through the radar feature extraction network.

[0048] Build a 10-layer feature extraction sub-network, the structure of which is in turn: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer. Set the kernel sizes of the first to fifth convolutional layers to 7×7, 1×1, 1×1, 3×3, 1×1 respectively; and the number of kernels to 64, 256, 64, 64, 256 respectively. Implement the first to fifth batch normalization layers using the FrozenBatchNorm function. For the radar feature extraction sub-network, there is relatively less information, and sufficient features can be obtained without using too many convolutional layers for processing, while the detection efficiency can be improved.

[0049] Step 4, construct a visible light feature extraction sub-network, and obtain the feature map of the input visible light image through the visible light image feature extraction network.

[0050] Build a 22-layer visible light feature extraction sub-network, the structure of which is in turn: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer, the sixth convolutional layer, the sixth batch normalization layer, the seventh convolutional layer, the seventh batch normalization layer, the eighth convolutional layer, the eighth batch normalization layer, the ninth convolutional layer, the ninth batch normalization layer, the tenth convolutional layer, the tenth batch normalization layer, the eleventh convolutional layer, the eleventh batch normalization layer. Set the kernel sizes of the first to eleventh convolutional layers to 7×7, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; and the number of kernels of the first to eleventh convolutional layers to 64, 256, 64, 64, 256, 64, 64, 256, 64, 64, 256 respectively. Implement the first to eleventh batch normalization layers using the FrozenBatchNorm function.

[0051] Step 5, construct a radar and visible light image fusion sub-network, and input the obtained radar features and visible light image features into the radar and visible light image fusion network to obtain the fused feature map.

[0052] The structure of the radar and visible light image fusion sub-network is in turn: the first multi-modal fusion module, the second multi-modal fusion module, the first convolutional block, the second convolutional block, the third convolutional block.

[0053] Refer to Figure 3 Make a further description of the multi-modal fusion module constructed in the embodiment of the present invention.

[0054] The structures of the first and second multi-modal fusion modules are the same. The structure of each multi-modal fusion module is, in sequence: the first linear layer, the second linear layer, the third linear layer, the first activation layer, and the second activation layer. The number of output neurons of the first to third linear layers of the first multi-modal fusion module is set to 256, and the number of output neurons of the first to third linear layers of the second multi-modal fusion module is set to 512, 256, and 256 respectively. Using the multi-modal fusion module can fully fuse the feature information of radar and visible light and improve the accuracy of target detection.

[0055] The first convolutional block uses the Stage2 structure of the Resnet50 network, which is a 26-layer network. Its structure is, in sequence: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer, the sixth convolutional layer, the sixth batch normalization layer, the seventh convolutional layer, the seventh batch normalization layer, the eighth convolutional layer, the eighth batch normalization layer, the ninth convolutional layer, the ninth batch normalization layer, the tenth convolutional layer, the tenth batch normalization layer, the eleventh convolutional layer, the eleventh batch normalization layer, the twelfth convolutional layer, the twelfth batch normalization layer, the thirteenth convolutional layer, and the thirteenth batch normalization layer. The kernel sizes of the first to thirteenth convolutional layers are set to 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; the numbers of kernels of the first to thirteenth convolutional layers are set to 512, 128, 128, 512, 128, 128, 512, 128, 128, 512, 128, 128, 512 respectively. The first to thirteenth batch normalization layers are implemented using the FrozenBatchNorm function.

[0056] The second convolutional block uses the Stage3 structure of the Resnet50 network, which is a network with 38 layers. Its structure is as follows: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer, the sixth convolutional layer, the sixth batch normalization layer, the seventh convolutional layer, the seventh batch normalization layer, the eighth convolutional layer, the eighth batch normalization layer, the ninth convolutional layer, the ninth batch normalization layer, the tenth convolutional layer, the tenth batch normalization layer, the eleventh convolutional layer, the eleventh batch normalization layer, the twelfth convolutional layer, the twelfth batch normalization layer, the thirteenth convolutional layer, the thirteenth batch normalization layer, the fourteenth convolutional layer, the fourteenth batch normalization layer, the fifteenth convolutional layer, the fifteenth batch normalization layer, the sixteenth convolutional layer, the sixteenth batch normalization layer, the seventeenth convolutional layer, the seventeenth batch normalization layer, the eighteenth convolutional layer, the eighteenth batch normalization layer, the nineteenth convolutional layer, the nineteenth batch normalization layer. The kernel sizes of the first to nineteenth convolutional layers are set to 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; the numbers of kernels of the first to nineteenth convolutional layers are set to 1024, 256, 256, 1024, 256, 256, 1024, 256, 256, 1024, 256, 256, 1024, 256, 256, 1024, 256, 256, 1024 respectively. The first to nineteenth batch normalization layers are implemented using the FrozenBatchNorm function.

[0057] The third convolutional block uses the Stage4 structure of the Resnet50 network. It is a network with 20 layers. Its structure is as follows: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer, the sixth convolutional layer, the sixth batch normalization layer, the seventh convolutional layer, the seventh batch normalization layer, the eighth convolutional layer, the eighth batch normalization layer, the ninth convolutional layer, the ninth batch normalization layer, the tenth convolutional layer, the tenth batch normalization layer. The kernel sizes of the first to tenth convolutional layers are set to 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; the numbers of kernels of the first to tenth convolutional layers are set to 2048, 512, 512, 2048, 512, 512, 2048, 512, 512, 2048 respectively. The first to tenth batch normalization layers are implemented using the FrozenBatchNorm function. Resnet50 is an existing and very mature network, which can achieve good performance in the field of object detection. This network is selected to further fuse information.

[0058] Step 6, construct the object detection network.

[0059] After paralleling the radar feature extraction sub-network and the visible light feature extraction sub-network, they are cascaded with the image fusion sub-network and the RetinaNet sub-network in sequence to form an object detection network for millimeter-wave radar and visible light image fusion. The RetinaNet sub-network is built with existing technologies and the parameters scale is set to (32, 64, 128, 256, 512), and ratios is set to (0.5, 1.0, 2.0).

[0060] Generate p3-p7 feature maps according to the outputs of the first convolutional block, the second convolutional block, and the third convolutional block in the radar and visible light image fusion sub-network.

[0061] Input the p3-p7 feature maps into their respective Head parts to obtain class results, centrality, and position results.

[0062] Step 7, train the object detection network.

[0063] Input the training set into the object detection network based on millimeter-wave radar and visible light image fusion, and use the stochastic gradient descent algorithm to iteratively update the weight values of the network, optimize the total loss function of the network until it converges, and obtain a trained object detection network.

[0064] Uniformly set the visible light detection image and the radar image to a size of 1333 in length and 800 in width, randomly flip them during each iteration, and standardize the data.

[0065] Set the training batch size batch to 1, that is, each iteration trains with 1 visible light image and 1 radar image as a group, and the parameters in the model are optimized once each time the model iterates.

[0066] Set the initial learning rate to 0.001 and the weight decay to 0.0001 to reduce the problem of model overfitting to a certain extent.

[0067] Set the maximum number of network iterations to 10000, and obtain a trained fusion network model after multiple rounds of training.

[0068] Step 8, input the test set data for object detection after training is completed.

[0069] Input each test sub-image in the test set into the trained network to obtain the final object detection results based on millimeter-wave radar and visible light image fusion.

[0070] The technical effects of the present invention are further described below through simulation experiments.

[0071] 1. Simulation experiment conditions:

[0072] The hardware platform for the simulation experiment of the present invention is: the processor is 11th Gen Intel(R) Core, the main frequency of the processor is 3.50 GHz, and the graphics card is NVIDIA GeForce RTX 3090.

[0073] The software platform for the simulation experiment of the present invention is: Windows 10 system.

[0074] Under Pytorch of Python3.7, a target detection network based on the fusion of millimeter-wave radar and visible light images is built, and the development language is Python. The Pytorch version is 1.7.1+cu110.

[0075] 2. Simulation experiment content and result analysis:

[0076] Under the above conditions, the data in 20190929_ONRD006 in the CRUW dataset is used. The scenario is the data collected during driving on a road. A target detection network with a single visible light image as the input and the target detection fusion network built by the present invention are used for simulation experiments. The only difference between the detection network of the single visible light image and the network of the present invention is the existence of the radar branch, and other parameters remain the same. The accuracy and recall rate of the target detection network are obtained. Radar target detection refers to the target detection method of RODNet proposed by Yizhou Wang et al. in "RODNet: Radar Object Detection using Cross-Modal Supervision, Workshop on Applications of Computer Vision IEEE, 2021". The detection results are compared as shown in Table 1:

[0077] Table 1: Comparison of target detection results

[0078] Radar target detection Visible light image target detection The present invention Accuracy rate AP 83.76 90.8% 98.2% Recall rate AR 85.62 71.0% 74.5%

[0079] As can be seen from Table 1, the accuracy of the target detection of the present invention is 98.2%, which is 7.4% higher than the accuracy of 90.8% of the visible light image target detection and 14.44% higher than the accuracy of 83.76% of the radar target detection; the recall rate of the target detection of the present invention is 74.5%, which is 3.5% higher than the recall rate of 71.0% of the visible light image target detection and 11.12% lower than the recall rate of 85.62% of the radar target detection. The accuracy of the present invention has a relatively high increase, but the recall rate is slightly lower. The improvement of the accuracy is obtained at the cost of part of the recall rate, and more excellent results are obtained.

[0080] In summary, the object detection method based on the fusion of millimeter-wave radar and visible light images proposed by the present invention can detect visible light images more accurately.

Claims

1. A target detection method based on the fusion of millimeter-wave radar and visible light images, characterized in that, Preprocess the millimeter-wave radar data to generate a radar image, and construct a radar and visible light image fusion sub-network in the target detection network; the steps of this method are as follows: Step 1, preprocess the millimeter-wave radar data to generate a radar image: Step 1.1, at the same moment when the vehicle-mounted radar obtains the radar echo signal data received by the radar sensor, the vehicle-mounted vision sensor obtains a visible light image corresponding to the radar echo signal data; Step 1.2, convert the radar echo signal data into matrix A, where the rows in matrix A represent distance and the columns represent angle; After taking the modulus of matrix A and then normalizing it, a matrix B of size 128*128*1 is obtained. After normalizing matrix A, a matrix C of 128*128*2 is obtained. Concatenate matrix B and matrix C to get matrix D; Step 1.3, save matrix D as a radar image; Step 2, generate a training set: Generate annotation files in json format for both the radar image and the visible light image, and form a training set with the radar image, the visible light image, and their generated annotation files; Step 3, construct a radar feature extraction sub-network: Build a 10-layer feature extraction sub-network, and its structure is as follows: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer. Set the kernel sizes of the first to fifth convolutional layers to 7×7, 1×1, 1×1, 3×3, 1×1 respectively; set the number of kernels of the first to fifth convolutional layers to 64, 256, 64, 64, 256 respectively; implement the first to fifth batch normalization layers using the FrozenBatchNorm function; Step 4, construct a visible light feature extraction sub-network: Build a 22-layer visible light feature extraction sub-network, and its structure is as follows: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, the third batch normalization layer, the fourth convolutional layer, the fourth batch normalization layer, the fifth convolutional layer, the fifth batch normalization layer, the sixth convolutional layer, the sixth batch normalization layer, the seventh convolutional layer, the seventh batch normalization layer, the eighth convolutional layer, the eighth batch normalization layer, the ninth convolutional layer, the ninth batch normalization layer, the tenth convolutional layer, the tenth batch normalization layer, the eleventh convolutional layer, the eleventh batch normalization layer. Set the kernel sizes of the first to eleventh convolutional layers to 7×7, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1, 1×1, 3×3, 1×1 respectively; set the number of kernels of the first to eleventh convolutional layers to 64, 256, 64, 64, 256, 64, 64, 256, 64, 64, 256 respectively; implement the first to eleventh batch normalization layers using the FrozenBatchNorm function; Step 5, construct a radar and visible light image fusion sub-network: The structure of the radar and visible light image fusion sub-network is as follows: the first multi-modal fusion module, the second multi-modal fusion module, the first convolutional block, the second convolutional block, the third convolutional block; Step 5.1, the structures of the first and second multi-modal fusion modules are the same. The structure of each multi-modal fusion module is in turn: the first linear layer, the second linear layer, the third linear layer, the first activation layer, and the second activation layer; Set the number of output neurons of the first to third linear layers of the first multi-modal fusion module to 256, and set the number of output neurons of the first to third linear layers of the second multi-modal fusion module to 512, 256, and 256 respectively; Step 5.2, the first convolutional block adopts the Stage2 structure of the Resnet50 network, the second convolutional block adopts the Stage3 structure of the Resnet50 network, and the third convolutional block uses the Stage4 structure of the Resnet50 network; the Stage2, Stage3, and Stage4 structures are all composed of 4, 6, and 3 Bottleneck structures in series. Each Bottleneck structure is in turn: the first convolutional layer, the first batch normalization layer, the second convolutional layer, the second batch normalization layer, the third convolutional layer, and the third batch normalization layer; Step 6, construct an object detection network: After paralleling the radar feature extraction sub-network and the visible light feature extraction sub-network, and then cascading them with the image fusion sub-network and the RetinaNet sub-network in turn, an object detection network for millimeter-wave radar and visible light image fusion is formed; Step 7, train the object detection network: Input the training set into the object detection network based on millimeter-wave radar and visible light image fusion. Use the stochastic gradient descent algorithm to iteratively update the weight values of the network and optimize the total loss function of the network until it converges, obtaining a trained object detection network; Step 8, detect the target: Adopt the same processing method as in Step 1. Preprocess the radar echo signal data received by the vehicle-mounted radar sensor to obtain a radar image. Input the radar image and the visible light image generated by the vehicle-mounted vision sensor at the same moment of the radar echo signal data into the trained network, and output the object detection result of millimeter-wave radar and visible light image fusion.

2. The object detection method based on millimeter-wave radar and visible light image fusion according to claim 1, characterized in that, The total loss function described in Step 7 is as follows: Among them, $L$ represents the total loss function of the fusion object detection network, $L$ cls represents the class loss of the object bounding boxes output by the fusion object detection network, $L$ reg represents the location loss between the object bounding boxes output by the fusion object detection network and the marked ground truth boxes, $N$ pos is the number of positive samples in the training set, and $\lambda$ is set to 1 to represent the balancing weight of $L$ reg ​

Citation Information

Patent Citations

  • Millimeter wave radar target detection method and system based on fused image features

    CN114218999A

  • Visible light, infrared and radar fusion target detection method based on deep learning

    CN114254696A