A target detection method, device, storage medium and electronic device
By combining camera images and radar maps in the target detection system, using pre-trained detection models for feature extraction and multi-scale encoding, directly fusion between image viewing angle and radar viewing angle, the problem of insufficient long-distance object detection in the existing system is solved, and efficient two-dimensional and three-dimensional object detection is achieved.
Patent Information
- Application Number
- CN202510266414.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-07
AI Technical Summary
When the existing target detection system uses a combination of millimeter-wave radar and cameras, it is difficult to meet the detection needs of long-distance targets. Especially in traffic scenarios, the detection distance is usually less than 200-300m, and the two-dimensional detection of the target cannot be achieved.
A target detection method is adopted to obtain the radar map generated by the camera's image and the point cloud data collected by millimeter-wave radar, and process the camera image and radar map using a pre-trained detection model to achieve two-dimensional and three-dimensional detection of the target. This method includes feature extraction, multi-scale coding and fusion processing, which directly integrates the image viewing angle and radar viewing angle, avoiding information loss caused by BEV viewing angle conversion.
It realizes efficient detection of long-distance targets, can perform 2D and 3D detection simultaneously, improves detection performance and accuracy, expands detection distance, and makes full use of small features in image and radar features.
Smart Images

Figure CN119763101B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to image processing technology, and particularly to an object detection method, device, storage medium, and electronic device. Background Art
[0002] With the progress of millimeter-wave radar technology and image processing technology, in various object detection systems, there has emerged a processing method of combining millimeter-wave radar and camera for object detection.
[0003] For example, the current road monitoring system uses a combination of millimeter-wave radar and camera to form a radar-vision tracking system, which integrates the rich semantics of vision and the advantage of radar in perceiving moving objects, and performs object tracking and detection evidence collection through multi-modal data fusion.
[0004] In the existing object detection that combines millimeter-wave radar and camera, it is usually based on the BEV perspective to fuse the millimeter-wave radar features and the image features collected by the camera, that is, to regress the 3D coordinates in the physical coordinate system to achieve three-dimensional detection of the object. This processing method has high ranging accuracy, but it will limit the detection distance. On the one hand, increasing the detection range will inevitably lead to an increase in the scale of BEV features and also an increase in the amount of calculation; on the other hand, the BEV features of distant objects are sparse, and the detection effect is not good. Especially for traffic scenarios, the demand for object detection distance usually reaches 200 - 300m, while the object detection method based on the BEV perspective cannot meet this business requirement; at the same time, the object detection based on the BEV perspective cannot achieve two-dimensional detection of the object. Summary of the Invention
[0005] The present application provides an object detection method, device, storage medium, and electronic device, which can simultaneously achieve two-dimensional and three-dimensional detection of objects and effectively improve the detection performance for distant objects.
[0006] To achieve the above object, the present application adopts the following technical solutions:
[0007] An object detection method, comprising:
[0008] Obtaining a camera image collected by a camera and a radar map generated from point cloud data collected by a millimeter-wave radar; wherein, the radar map is a point cloud map mapped to the camera image coordinate system;
[0009] Processing the camera image and the radar map image by using a pre-trained detection model to obtain a two-dimensional detection result and a three-dimensional detection result of the object;
[0010] Wherein, the processing of the detection model includes:
[0011] Feature extraction is respectively performed on the camera image and the radar mapping image to obtain image features and radar features, and the image features and radar features are subjected to first fusion processing to obtain primary fusion features;
[0012] The two-dimensional detection head is used to perform two-dimensional detection processing on the primary fusion features to obtain two-dimensional detection results of the target;
[0013] The radar features and the primary fusion features are subjected to second fusion processing to obtain secondary fusion features;
[0014] The three-dimensional detection head is used to perform three-dimensional detection processing on the secondary fusion features to obtain three-dimensional detection results of the target.
[0015] Preferably, before the radar features and the primary fusion features are subjected to second fusion processing, the method further includes: performing multi-scale encoding on the radar features, and using the encoding result as the updated radar features for the second fusion processing.
[0016] Preferably, before the image features and radar features are subjected to second fusion processing, the method further includes: performing multi-scale encoding on the primary fusion features, and using the encoding result as the updated primary fusion features for the two-dimensional detection processing and the second fusion processing.
[0017] Preferably, the second fusion processing of the radar features and the primary fusion features includes:
[0018] Directly fusing the radar features and the primary fusion features to obtain the secondary fusion features;
[0019] Or,
[0020] Converting the respective feature pixel coordinates of the primary fusion features back to the radar coordinate system and performing coordinate encoding;
[0021] After merging the coordinate encoding result and the primary fusion features, the merged features are obtained;
[0022] The merged features and the radar features are directly fused and processed.
[0023] Preferably, the training of the detection model includes a first training stage and a second training stage performed in sequence;
[0024] In the first training stage, only the image training data collected by the camera and the corresponding two-dimensional detection labels are used to train the detection model, and the parameters of the feature extraction and the two-dimensional detection processing performed on the camera image are updated;
[0025] In the second training stage, the detection model is trained using the image training data, the radar training data, and the corresponding 2D detection labels and 3D detection labels to update all parameters of the detection model.
[0026] Preferably, after the first training stage, the training of the detection model further includes:
[0027] When the input training data only includes the image training data collected by the camera and the 2D detection labels, the detection model is trained using the image training data and the 2D detection labels to update the parameters of the feature extraction, the 2D detection processing, and the first fusion processing performed on the camera image;
[0028] When the input training data only includes the image training data collected by the camera, the 2D detection labels, and the 3D detection labels, the detection model is trained using the image training data, the 2D detection labels, and the 3D detection labels to update the parameters of the feature extraction, the first fusion processing, the 2D detection processing, the second fusion processing, and the 3D detection processing performed on the camera image.
[0029] Preferably, the first fusion processing of the image feature and the radar feature includes: splicing or adding the image feature and the radar feature, and performing convolution processing on the splicing or adding result to obtain the first fusion feature.
[0030] Preferably, the direct fusion processing of the radar feature and the first fusion feature includes: splicing or adding the radar feature and the first fusion feature, and performing convolution processing on the splicing or adding result;
[0031] and / or,
[0032] The direct fusion processing of the merged feature and the radar feature includes: splicing or adding the merged feature and the radar feature, and performing convolution processing on the splicing or adding result.
[0033] An object detection device, the device includes: a data acquisition unit, a detection model processing unit, and a result output unit;
[0034] The data acquisition unit is configured to acquire a camera image collected by a camera and a radar mapping image generated from point cloud data collected by a millimeter-wave radar; wherein, the radar mapping map is a point cloud map mapped to the camera image coordinate system;
[0035] The detection model processing unit is configured to process the camera image and the radar map using a pre-trained detection model to obtain two-dimensional and three-dimensional detection results of the target, and output them through the result output unit;
[0036] Among them, the detection model unit includes an image encoding module, a radar encoding module, a first fusion module, a two-dimensional detection head module, a second fusion module, and a three-dimensional detection head module;
[0037] The image encoding module and the radar encoding module are respectively configured to extract features from the camera image and the radar map image to obtain image features and radar features;
[0038] The first fusion module is configured to perform first fusion processing on the image features and the radar features to obtain a first fusion feature;
[0039] The two-dimensional detection head module is configured to perform two-dimensional detection processing on the first fusion feature using a two-dimensional detection head to obtain a two-dimensional detection result of the target;
[0040] The second fusion module is configured to perform second fusion processing on the radar features and the first fusion feature to obtain a second fusion feature;
[0041] The three-dimensional detection head module is configured to perform three-dimensional detection processing on the second fusion feature using a three-dimensional detection head to obtain a three-dimensional detection result of the target.
[0042] Preferably, the detection model processing further includes a multi-scale radar encoding module, which is configured to perform multi-scale encoding on the radar features, and use the encoding result as the updated radar features to input the second fusion module for the second fusion processing with the first fusion feature.
[0043] Preferably, the detection model processing unit further includes a multi-scale image encoding module, which is configured to perform multi-scale encoding on the first fusion feature, and use the encoding result as the updated first fusion feature to input the second fusion module for the second fusion processing with the radar features.
[0044] Preferably, in the second fusion module, the second fusion processing of the radar features and the first fusion feature includes:
[0045] Directly perform fusion processing on the radar features and the first fusion feature to obtain the second fusion feature;
[0046] Or,
[0047] Convert the respective feature pixel coordinates of the first fusion feature back to the radar coordinate system and perform coordinate encoding;
[0048] After merging the coordinate encoding result and the first fusion feature, a merged feature is obtained.
[0049] The merged feature is directly fused with the radar feature.
[0050] Preferably, the device further includes a detection model training unit for training and generating the detection model.
[0051] Among them, the training of the detection model includes a first training stage and a second training stage that are carried out in sequence.
[0052] In the first training stage, the detection model is trained only using the image training data collected by the camera and the corresponding two-dimensional detection labels, and the parameters of the feature extraction and the two-dimensional detection processing performed on the camera image are updated.
[0053] In the second training stage, the detection model is trained using the image training data, the radar training data, and the corresponding two-dimensional detection labels and three-dimensional detection labels, and all parameters of the detection model are updated.
[0054] Preferably, in the detection model training unit, after the first training stage, the training of the detection model further includes:
[0055] When the input training data only includes the image training data collected by the camera and the two-dimensional detection labels, the detection model is trained using the image training data and the two-dimensional detection labels, and the parameters of the image encoding module, the two-dimensional detection processing, and the first fusion processing are updated.
[0056] When the input training data only includes the image training data collected by the camera, the two-dimensional detection labels, and the three-dimensional detection labels, the detection model is trained using the image training data, the two-dimensional detection labels, and the three-dimensional detection labels, and the parameters of the feature extraction, the first fusion processing, the two-dimensional detection processing, the second fusion processing, and the three-dimensional detection processing performed on the camera image are updated.
[0057] Preferably, in the first fusion module, the first fusion processing of the image feature and the radar feature includes: splicing or adding the image feature and the radar feature, and performing convolution processing on the splicing or adding result to obtain the first fusion feature.
[0058] Preferably, in the second fusion module,
[0059] The direct fusion processing of the radar feature and the first fusion feature includes: splicing or adding the radar feature and the first fusion feature, and performing convolution processing on the result of the splicing or adding;
[0060] and / or,
[0061] The direct fusion processing of the merged feature and the radar feature includes: splicing or adding the merged feature and the radar feature, and performing convolution processing on the result of the splicing or adding.
[0062] A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the target detection method described in any one of the above can be implemented.
[0063] An electronic device, which at least includes a computer-readable storage medium and also includes a processor;
[0064] The processor is used to read executable instructions from the computer-readable storage medium and execute the instructions to implement the target detection method described in any one of the above.
[0065] As can be seen from the above technical solutions, in this application, first, a camera image collected by a camera and a radar map generated from point cloud data collected by a millimeter-wave radar are obtained. The radar map is a point cloud map mapped to the camera image coordinate system, that is, a radar image to be processed is obtained through RV mapping. Then, the pre-trained detection model is used to process the camera image and the radar map image to obtain two-dimensional detection results and three-dimensional detection results of the target. In this way, the detection model is directly used to process the camera image and the radar map obtained by RV mapping to achieve two-dimensional and three-dimensional detection, without the need to switch to the BEV perspective for three-dimensional detection. Specifically, the detection model includes two branches: two-dimensional detection and three-dimensional detection. First, feature extraction is respectively performed on the camera image and the radar map to obtain image features and radar features. In the two-dimensional detection branch, the image features and the radar features are subjected to first fusion processing to obtain a first fusion feature, and then the two-dimensional detection head is used to perform two-dimensional detection processing on the first fusion feature to obtain two-dimensional detection results of the target. In the three-dimensional detection branch, the radar features and the first fusion feature in the two-dimensional detection branch are subjected to second fusion processing to obtain a second fusion feature, and then the three-dimensional detection head is used to perform three-dimensional detection processing on the second fusion feature to obtain three-dimensional detection results of the target. Through the above processing, on the one hand, the image features and the radar features are fused and then two-dimensional detection is performed. On the other hand, the fused features and the radar features are subjected to second fusion and then three-dimensional detection is performed. In this way, the information of the radar, including radar depth information, can be more fully utilized to improve the accuracy of three-dimensional detection. At the same time, since the entire processing directly uses the image features and the radar features after RV mapping and does not need to switch to the BEV perspective, the problem that the features of distant targets are not significant in the BEV perspective due to information loss during the BEV conversion is avoided, and the full utilization of fine features in the image and radar features can be effectively ensured, and the detection effect of distant targets can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a schematic diagram of the basic process of the target detection method in this application;
[0067] Figure 2 is a schematic diagram of the specific process of the target detection method in a specific embodiment of this application;
[0068] Figure 3 is a flowchart of the target detection method in a specific embodiment of this application;
[0069] Figure 4 is an implementation example diagram of the first fusion module in a specific embodiment of this application;
[0070] Figure 5 is an example diagram of implementing the second fusion processing by adding coordinate encoding in a specific embodiment of this application;
[0071] Figure 6 This is a schematic diagram of the basic structure of the target detection device in this application;
[0072] Figure 7 This is a schematic diagram of the basic structure of the electronic device provided in this application. Specific embodiments
[0073] In order to make the objectives, technical means, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings.
[0074] The basic idea of this application is as follows: directly utilize the image features and the radar features after RV mapping for two-dimensional and three-dimensional detections, thereby making full use of the fine features in the image and radar features to improve the detection effect of distant targets; at the same time, strengthen the use of radar features in three-dimensional detection to effectively ensure the accuracy of three-dimensional detection.
[0075] Figure 1 This is a schematic diagram of the basic process of the target detection method in this application. As Figure 1 shown, the method includes:
[0076] Step 101, obtain the camera image collected by the camera and the radar mapping map generated from the point cloud data collected by the millimeter-wave radar.
[0077] Among them, the radar mapping map is a point cloud map mapped to the camera image coordinate system, that is, the radar point cloud data collected by the millimeter-wave radar is transformed to the image perspective through the RV matrix to generate the radar mapping map. Each channel of the radar mapping map can be point cloud information such as the coordinates, speed, angle, signal-to-noise ratio, etc. of the radar points in the radar coordinate system.
[0078] Step 102, use the pre-trained detection model to process the camera image and the radar mapping map to obtain the two-dimensional detection result and the three-dimensional detection result of the target.
[0079] Among them, the processing of the camera image and the radar mapping map by the detection model specifically includes steps 102a~102e:
[0080] Step 102a, respectively extract features from the camera image and the radar mapping map to obtain image features and radar features;
[0081] This step is used to obtain image features and radar features. Among them, the image features are obtained by directly extracting features from the camera image, and the radar features are obtained by extracting features from the radar map after RV mapping. That is to say, the features used for target detection in this application are directly the image features obtained from the camera image and the radar features obtained from the radar map after RV mapping of the radar point cloud. The corresponding feature extraction does not need to be converted to the BEV perspective, avoiding the problem that the features of distant targets are not significant in the BEV perspective due to information loss during the BEV conversion process.
[0082] Step 102b: Perform first fusion processing on the image features and radar features to obtain primary fusion features.
[0083] The image features and radar features can be directly fused (to distinguish from subsequent fusion processing, this step of fusion processing is called first fusion processing) to obtain primary fusion features. The specific fusion processing can adopt various existing fusion processing methods. For example, the image features and radar features are added or concatenated, and then convolution processing is performed on the addition or concatenation result.
[0084] Step 102c: Use a two-dimensional detection head to perform two-dimensional detection processing on the primary fusion features to obtain two-dimensional detection results of the target.
[0085] In this step, a two-dimensional detection head can be directly used to perform two-dimensional detection processing on the primary fusion features to obtain two-dimensional detection results of the target. Or, to better reflect the features of targets at various scales and make full use of the features of targets at various scales, preferably, the primary fusion features can also be encoded at multiple scales, and the encoded results are used as the updated primary fusion features for two-dimensional detection processing.
[0086] Among them, the processing of the two-dimensional detection head can be implemented by existing methods. For example, it can be implemented through convolution processing, which can well control the calculation amount on the basis of ensuring detection accuracy. At the same time, it has high efficiency during implementation and is friendly to hardware platforms, supporting deployment on most platforms.
[0087] The processing of the above steps 102b to 102c is actually the processing of the two-dimensional detection branch in the detection model, which is used to perform two-dimensional detection based on image features and radar features and generate two-dimensional detection results. Specifically, it can include identifying the target category and the two-dimensional coordinates, length, width, and scale information of the target. In this branch, the information collected by the camera is fused with the information collected by the millimeter-wave radar, so as to effectively utilize the image information collected by the camera and further supplement the radar information to make up for the defects of the camera image under conditions such as extreme weather, night, and harsh working conditions where the camera image cannot effectively reflect the two-dimensional information of the target, thereby realizing effective two-dimensional detection of the target in various situations.
[0088] Step 102d: Perform second fusion processing on the radar features and the first fusion features to obtain second fusion features.
[0089] This step is for feature preparation for 3D detection. Since the depth information required in 3D detection cannot be directly obtained from camera images, and the radar provides effective complementary information, in this application, when preparing features for 3D detection, to make further full use of the information collected by the millimeter-wave radar, the first fusion features prepared in 2D detection and the radar features are fused to obtain second fusion features for subsequent 3D detection of targets.
[0090] Among them, to distinguish it from the aforementioned first fusion processing, the fusion processing of the radar features and the first fusion features here is called second fusion processing. The radar features obtained in step 102a can be directly fused with the first fusion features, or preferably, to make full use of the radar features of targets at various different scales, the radar features can also be encoded at multiple scales, and the encoding results are used as the updated radar features for second fusion processing.
[0091] The specific fusion processing can adopt various existing fusion processing methods. For example, the first fusion features and the radar features are added or concatenated, and then convolution processing is performed on the added or concatenated results.
[0092] In addition, in the existing 3D detection based on the BEV perspective, since the camera images and radar point clouds need to be converted to the BEV perspective for feature extraction, there is information loss during the conversion, resulting in insignificant features of distant targets in the BEV perspective. The 3D detection based on the weakened features will significantly affect the accuracy of the detection results. In this application, the features for 3D detection are directly obtained by using image features and the radar features after RV mapping, without the need to be converted to the BEV perspective. Therefore, the fine features in the distance can be fully utilized, effectively improving the 3D detection accuracy of distant targets and expanding the 3D detection distance. At the same time, if multi-scale encoding is introduced, it can further make full use of the features at each scale, further improving the utilization rate of fine features and the 3D detection accuracy of distant targets.
[0093] Step 102e: Use a 3D detection head to perform 3D detection processing on the second fusion features to obtain the 3D detection results of the target.
[0094] In this step, an existing 3D detection head is used to perform 3D detection processing on the second fusion features to obtain 3D detection results. The specific 3D detection processing can be implemented by existing methods, such as through convolution processing.
[0095] The processing of steps 102d to 102e above is actually the processing of the 3D detection branch in the detection model, which is used to perform 3D detection based on image features and radar features to generate 3D detection results. Specifically, it can include the 3D coordinates, heading angle, and length, width, and height scale information of the target. In this branch, the information collected by the camera is fused with the information collected by the millimeter-wave radar, especially the radar information is reused, so as to make full use of the radar information to improve the accuracy of 3D detection. At the same time, in 3D detection, there is no need to convert to the BEV perspective, avoiding the problem that the features of distant targets are not obvious in the BEV perspective due to information loss, improving the accuracy of 3D detection of distant targets, and effectively increasing the 3D detection distance.
[0096] So far, Figure 1 the basic process shown ends. The following uses specific embodiments to illustrate the specific implementation of the target detection method in this application.
[0097] Figure 2 It is a schematic diagram of the specific process of the target detection method in a specific embodiment of this application. Figure 3 It is a flowchart of the target detection method in a specific embodiment. As Figure 2 shown, the method includes:
[0098] Step 201, obtain a camera image and a radar map.
[0099] The camera image is obtained by camera acquisition. The radar map is obtained by performing RV mapping on the radar point cloud data collected by the millimeter-wave radar, and can be implemented by existing methods.
[0100] Specifically, assume that the coordinates of a certain radar point A in the radar point cloud in the radar coordinate system are (x, y), and the pre-calibrated RV mapping matrix The pixel coordinates (u, v) of the radar point A in the radar map obtained by RV mapping can be expressed as:
[0101]
[0102] Performing RV mapping on all radar points in the above manner can obtain a radar map with the same size as the camera image.
[0103] After obtaining the camera image and the radar map in step 201, the two images are input into the detection model for processing to obtain the 2D detection result and the 3D detection result of the target. The processing of the detection model starts from step 202 below.
[0104] Step 202, use the image encoding module and the radar encoding module to extract features from the camera image and the radar map respectively to obtain image features and radar features.
[0105] The image coding module is used to extract features from the camera image to obtain image features; the radar coding module is used to extract features from the radar map to obtain radar features.
[0106] Among them, the image coding module and the radar coding module can be implemented in an existing manner. For example, in this embodiment, they are implemented through multiple cascaded convolutional modules.
[0107] Step 203: Use the first fusion module to perform first fusion processing on the image features and the radar features to obtain primary fusion features.
[0108] In this embodiment, the implementation of the first fusion module is as Figure 4 shown. After concatenating or adding the image features and the radar features, the primary fusion features are obtained through a convolutional module. The first fusion processing here realizes the first RV feature fusion.
[0109] Step 204: Use the multi-scale image coding module to perform multi-scale coding on the primary fusion features to obtain updated primary fusion features.
[0110] The specific multi-scale coding can be implemented in an existing manner. Introducing multi-scale coding can fully extract the features of large and small targets at different scales to further improve the detection accuracy.
[0111] Step 205: Use the two-dimensional detection head module to perform two-dimensional detection processing on the updated primary fusion features to obtain the two-dimensional detection results of the target.
[0112] In this embodiment, the two-dimensional detection head is implemented through a convolutional module. Compared with the two-dimensional detection head implemented by using a transformer module, it has a simple logic, a small amount of calculation, a high implementation efficiency, is friendly to the hardware platform support, and supports deployment on most platforms.
[0113] Step 206: Use the multi-scale radar coding module to perform multi-scale coding on the radar features to obtain updated radar features.
[0114] The specific multi-scale coding can be implemented in an existing manner. Introducing multi-scale coding can fully extract the radar features of large and small targets at different scales to further improve the accuracy of three-dimensional detection.
[0115] Step 207: Use the second fusion module to perform second fusion processing on the updated radar features and the updated primary fusion features to obtain secondary fusion features.
[0116] In the second fusion module, the radar features and the updated first fusion features can be directly fused to obtain second fusion features. The fusion process can be implemented in various existing ways, such as weighted addition, weighted concatenation, or convolution after addition or concatenation.
[0117] Alternatively, to further utilize the radar information, the respective feature pixel coordinates of the first fusion features can be converted back to the radar coordinate system and coordinate encoding can be performed; after merging the coordinate encoding results and the first fusion features, the merged features are obtained; then the merged features and the radar features are directly fused to obtain second fusion features. The fusion process can also be implemented in various existing ways, such as weighted addition, weighted concatenation, or convolution after addition or concatenation.
[0118] Specifically, in this embodiment, the second fusion process is implemented by adding coordinate encoding, as Figure 5 shown. Specifically, according to the RV matrix, the radar coordinates (x, y) corresponding to each feature pixel in the first fusion features are calculated, and then coordinate encoding is generated for the radar coordinates. In this embodiment, triangular coordinate encoding is used to obtain the encoding result , where , , is the temperature coefficient, usually set to 10000, d is the feature dimension, and n is the dimension serial number. The coordinate encoding and the first fusion features are merged (the specific merging process can be set according to actual needs, for example, it can be addition or concatenation), the merged result is concatenated or added to the radar features, and then the concatenated or added result is convolved to obtain second fusion features.
[0119] The addition of coordinate encoding here is mainly considered that 3D detection is mainly based on coordinate information. Therefore, adding coordinate encoding in the radar coordinate system can make more full use of coordinate information and improve the accuracy of 3D detection.
[0120] Step 208, use the 3D detection head to perform 3D detection processing on the second fusion features to obtain the 3D detection results of the target.
[0121] The processing of this step can be implemented using a convolution module.
[0122] The processing of steps 202 to 208 is the complete inference process of the detection model. Through this detection model, 2D and 3D detections can be performed on the input camera image and radar map. Among them, the processing of steps 203 to 205 is the processing of the 2D detection branch, and the processing of steps 206 to 208 is the processing of the 3D detection branch.
[0123] So far, Figure 2 andFigure 3 The object detection process in the specific embodiment shown ends.
[0124] The above Figure 2 and Figure 3 The object detection process shown can be applied in various object detection systems. For example, it can be used in a traffic road monitoring system to detect traffic objects. At the same time, the above object detection process realizes the detection of objects based on the data collected on-site by the system. The detection model therein is pre-trained based on various types of training data. The specific training process can adopt the existing mature neural network model training method. For example, in each round of training, camera image training data (i.e., the camera images used for training), radar training data (i.e., the radar maps used for training), corresponding 2D detection labels (i.e., the 2D detection ground truths corresponding to the training data), and 3D detection labels (i.e., the 3D detection ground truths corresponding to the training data) are used. The entire detection model is used to process the camera image training data and radar training data, predict the 2D detection results and 3D detection results, compare them with the 2D detection labels and 3D detection labels respectively to obtain the value of the loss function, and update all the parameters of the entire detection model based on this value, and then perform the next round of training until the training end condition is met, such as reaching the specified number of training rounds, or the value of the loss function meets the preset conditions, etc.
[0125] In addition, considering that in the actual process of training the model, the cost of obtaining image data containing 2D detection boxes (i.e., 2D detection labels) is low, and there is a large amount of data accumulation in the early stage, while the image data containing 3D ground truth boxes (i.e., 3D detection labels) and the corresponding radar data are relatively scarce. In order to make full use of these additional 2D detection data, this application also proposes a hybrid training strategy when the 2D data and 3D data samples are unbalanced to make full use of a large amount of 2D image data. The image data here refers to the camera images generated by the camera.
[0126] Specifically, the training of the detection model can be divided into a first training stage and a second training stage that are carried out in sequence.
[0127] In the first training stage, mainly use 2D image data and 2D detection labels to train and generate a 2D detection model;
[0128] In the second training stage, use 2D image data, radar map data, 2D detection labels and 3D detection labels to train and generate the entire detection model.
[0129] More specifically, in the first training stage, only use 2D image data and 2D detection labels to train the detection model, update the parameters for feature extraction and 2D detection processing of the camera images, and do not update the parameters of other processing, that is, update Figure 3the parameters of the image encoding module and the 2D detection module in; if the detection model includes Figure 3 the multi-scale image encoding module in, the parameters are also updated in the first training stage. When specifically implemented, the image training data collected by the camera can be input into the detection model, feature extraction is performed on the image training data, and the 2D detection head is used to perform 2D detection processing on the feature extraction result to obtain the predicted 2D detection result; the predicted 2D detection result is compared with the 2D detection label corresponding to the image training data, and the parameters in the feature extraction and 2D detection processing are updated based on the comparison result. Thus, a 2D detection model with strong detection ability can be generated in the first training stage.
[0130] In the second training stage, the entire detection model is trained using the image training data, radar training data, and the corresponding 2D detection labels and 3D detection labels to update all the parameters of the detection model.
[0131] Or, further, after the first training stage is completed, the 2D image training data, 2D detection labels, and 3D detection labels can also be used to participate in the training of the detection model to make more full use of the 2D training data. Specifically, when the input training data only includes the image training data collected by the camera and the 2D detection labels (the input of radar training data is 0), the detection model can be trained using the image training data and the 2D detection labels to update the parameters of the feature extraction, 2D detection processing, and the first fusion processing for the camera image, and the parameters of other processing are not updated, that is, update Figure 3 the parameters of the image encoding module, the first fusion module, and the 2D detection module in; if the detection model includes Figure 3 the multi-scale image encoding module in, the parameters are also updated in this training stage. When the input training data only includes the image training data collected by the camera, 2D detection labels, and 3D detection labels (the input of radar training data is 0), the detection model is trained using the image training data, 2D detection labels, and 3D detection labels to update the parameters of the feature extraction, the first fusion processing, 2D detection processing, the second fusion processing, and 3D detection processing for the camera image, and the parameters of other processing are not updated, that is, update Figure 3 the parameters of the image encoding module, the first fusion module, the 2D detection module, the second fusion module, and the 3D detection module in; if the detection model includes Figure 3 the multi-scale image encoding module in, the parameters are also updated in this training stage.
[0132] In specific implementation, when the input training data only includes the image training data collected by the camera and the 2D detection labels (the input of radar training data is 0), the image training data collected by the camera can be input into the detection model to perform feature extraction and first fusion processing on the image training data, and then the 2D detection head is used to perform 2D detection processing on the once-fused features to obtain the predicted 2D detection results; the predicted 2D detection results are compared with the 2D detection labels corresponding to the image training data, and the parameters in the feature extraction, first fusion processing, and 2D detection processing of the camera image are updated based on the comparison results.
[0133] When the input training data only includes the image training data collected by the camera, the 2D detection labels, and the 3D detection labels (the input of radar training data is 0), the image training data collected by the camera can be input into the detection model to perform feature extraction and first fusion processing on the image training data, and then the 2D detection head is used to perform 2D detection processing on the once-fused features to obtain the predicted 2D detection results; the predicted 2D detection results are compared with the 2D detection labels corresponding to the image training data to obtain the 2D comparison results; the once-fused features are subjected to second fusion processing, and then the 3D detection head is used to perform 3D detection processing on the twice-fused features to obtain the predicted 3D detection results; the predicted 3D detection results are compared with the 3D detection labels corresponding to the image training data to obtain the 3D comparison results; the parameters in the feature extraction, first fusion processing, 2D detection processing, second fusion processing, and 3D detection processing of the camera image are updated based on the 2D comparison results and the 3D comparison results.
[0134] The above detection model training method given in this application can make the most of the data and reasonably update the parameters according to the training data. At the same time, since the situation without radar data is included in the training, the trained detection model can also effectively implement target detection in a specific situation without radar data.
[0135] Compared with the existing method based on feature fusion from the BEV perspective, the object detection method in this application can fuse image features and millimeter-wave radar features from the image perspective, and can output 2D and 3D detection results simultaneously to meet various downstream service requirements; a second fusion processing is designed, and coordinate encoding is introduced to further enhance the radar features and obtain more accurate 3D detection results; at the same time, this application provides a method for hybrid training of 2D detection data and 3D detection data, which can effectively utilize the existing 2D detection data when the 3D detection data is insufficient and improve the effect of 3D detection; in addition, the entire process implemented in this application only uses basic convolutional modules, with simple logic, small computational complexity, high implementation efficiency, friendly support for hardware platforms, and support for deployment on most platforms.
[0136] The above is the specific implementation of the object detection method in this application. This application also provides an object detection device, which can be used to implement the above object detection method. Figure 6 It is a schematic diagram of the basic structure of the object detection device in this application. As Figure 6 shown, the device includes: a data acquisition unit, a detection model processing unit, and a result output unit.
[0137] Among them, the data acquisition unit is used to acquire the camera image collected by the camera and the radar map generated from the point cloud data collected by the millimeter-wave radar; wherein, the radar map is a point cloud map mapped to the camera image coordinate system;
[0138] The detection model processing unit is used to process the camera image and the radar map by using a pre-trained detection model to obtain the two-dimensional detection result and the three-dimensional detection result of the object, and output them through the result output unit;
[0139] The result output unit is used to output the two-dimensional detection result and the three-dimensional detection result;
[0140] Among them, the detection model unit includes an image encoding module, a radar encoding module, a first fusion module, a two-dimensional detection module, a second fusion module, and a three-dimensional detection module;
[0141] The image encoding module and the radar encoding module are respectively used to extract features from the camera image and the radar map to obtain image features and radar features;
[0142] The first fusion module is used to perform a first fusion process on the image features and the radar features to obtain a first fusion feature;
[0143] The two-dimensional detection module is used to perform two-dimensional detection processing on the first fusion feature by using a two-dimensional detection head to obtain the two-dimensional detection result of the object;
[0144] The second fusion module is used to perform a second fusion process on the radar features and the first fusion feature to obtain a second fusion feature;
[0145] The three-dimensional detection module is used to perform three-dimensional detection processing on the second fusion feature by using a three-dimensional detection head to obtain the three-dimensional detection result of the object.
[0146] Optionally, the detection model processing further includes a multi-scale radar encoding module, which is used to perform multi-scale encoding on the radar features, and use the encoding result as the updated radar features to input into the second fusion module for the second fusion process with the first fusion feature.
[0147] Optionally, the detection model processing unit further includes a multi-scale image encoding module for performing multi-scale encoding on the first fusion feature, and using the encoding result as the updated first fusion feature to input into the second fusion module for second fusion processing with the radar feature.
[0148] Optionally, in the second fusion module, the second fusion processing of the radar feature and the first fusion feature may specifically include:
[0149] Directly fusing the radar feature and the first fusion feature to obtain a second fusion feature;
[0150] Alternatively, to make full use of the coordinate information and improve the detection accuracy, the respective feature pixel coordinates of the first fusion feature can also be converted back to the radar coordinate system and coordinate encoding can be performed;
[0151] After merging the coordinate encoding result and the first fusion feature, a merged feature is obtained;
[0152] Directly fusing the merged feature and the radar feature for processing.
[0153] Optionally, the apparatus may further include a detection model training unit for training and generating a detection model;
[0154] Wherein, to make more full use of the existing 2D image training data, the training of the detection model may include a first training stage and a second training stage that are carried out in sequence;
[0155] In the first training stage, only the image training data collected by the camera and the corresponding 2D detection labels are used to train the detection model, and the parameters for feature extraction and 2D detection processing of the camera image are updated;
[0156] In the second training stage, the image training data, the radar training data, and the corresponding 2D detection labels and 3D detection labels are used to train the detection model, and all parameters of the detection model are updated.
[0157] Optionally, in the detection model training unit, after the first training stage, the training of the detection model may further include:
[0158] When the input training data only includes the image training data collected by the camera and the 2D detection labels, the detection model is trained using the image training data and the 2D detection labels, and the parameters of the image encoding module, 2D detection processing, and first fusion processing are updated;
[0159] When the input training data only includes the image training data collected by the camera, the 2D detection labels, and the 3D detection labels, the detection model is trained using the image training data, the 2D detection labels, and the 3D detection labels, and the parameters for feature extraction, first fusion processing, 2D detection processing, second fusion processing, and 3D detection processing performed on the camera images are updated.
[0160] Optionally, in the first fusion module, the image features and the radar features are subjected to first fusion processing, which may specifically include: splicing or adding the image features and the radar features, and performing convolution processing on the splicing or addition result to obtain the first fusion features.
[0161] Optionally, in the second fusion module,
[0162] the radar features and the first fusion features are directly fused, which may specifically include: splicing or adding the radar features and the first fusion features, and performing convolution processing on the splicing or addition result;
[0163] and / or,
[0164] the combined features and the radar features are directly fused, which may specifically include: splicing or adding the combined features and the radar features, and performing convolution processing on the splicing or addition result.
[0165] The present application also provides a computer-readable storage medium storing instructions that, when executed by a processor, can execute the steps in the above-mentioned object detection method. In practical applications, the computer-readable medium may be included in each of the devices / apparatuses / systems in the above embodiments, or may exist alone without being assembled into the device / apparatus / system. Among them, instructions are stored in the computer-readable storage medium, and the stored instructions can execute the steps in the above-mentioned object detection method when executed by a processor.
[0166] According to the embodiments disclosed in the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above, but is not used to limit the scope of protection of the present application. In the embodiments disclosed in the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0167] Figure 7 The present application also provides an electronic device. As Figure 7As shown, it shows a schematic structural diagram of an electronic device involved in an embodiment of the present application. Specifically:
[0168] The electronic device may include a processor 701 with one or more processing cores, a memory 702 of one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When executing the program in the memory 702, the object detection method can be implemented.
[0169] Specifically, in practical applications, the electronic device may further include components such as a power supply 703 and an input / output unit 704. Those skilled in the art can understand that Figure 7 the structure of the electronic device shown in does not constitute a limitation to the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0170] The processor 701 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 702, and calling data stored in the memory 702, it executes various functions of the server and processes data, thereby monitoring the entire electronic device.
[0171] The memory 702 can be used to store software programs and modules, that is, the above-mentioned computer-readable storage medium. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the server, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 702 may further include a memory controller to provide the processor 701 with access to the memory 702.
[0172] The electronic device further includes a power supply 703 that powers each component, and can be logically connected to the processor 701 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 703 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0173] The electronic device may further include an input / output unit 704. The input unit 704 may be used to receive input digital or character information, and generate a keyboard, a mouse, a joystick, and an optical signal input related to user settings and function control. The input unit 704 may also be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, which may be composed of graphics, text, icons, videos, and any combination thereof.
[0174] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A target detection method, characterized in that: include: Acquire a camera image captured by a camera and a radar map generated by point cloud data collected by a millimeter-wave radar; wherein the radar map is a point cloud map mapped to a camera image coordinate system; The camera image and the radar mapping image are processed using a pre-trained detection model to obtain a two-dimensional detection result and a three-dimensional detection result of the target; The processing of the detection model includes: Extracting features from the camera image and the radar mapping image respectively to obtain image features and radar features, and performing a first fusion process on the image features and the radar features to obtain a primary fusion feature; Using a two-dimensional detection head to perform two-dimensional detection processing on the primary fusion feature to obtain a two-dimensional detection result of the target; Performing a second fusion process on the radar feature and the primary fusion feature to obtain a secondary fusion feature; Using a three-dimensional detection head to perform three-dimensional detection processing on the secondary fusion feature to obtain a three-dimensional detection result of the target; The second fusion processing of the radar feature and the first fusion feature includes: converting the coordinates of each feature pixel of the first fusion feature back to the radar coordinate system and performing coordinate encoding; merging the coordinate encoding result with the first fusion feature to obtain a merged feature; and directly fusing the merged feature with the radar feature; The training of the detection model adopts a hybrid training strategy of two-dimensional training data and three-dimensional training data, and the training of the detection model includes a first training stage and a second training stage which are performed sequentially; In the first training stage, the detection model is trained using only the image training data collected by the camera and the corresponding two-dimensional detection labels, and the parameters of the feature extraction and the two-dimensional detection processing performed on the camera image are updated; In the second training stage, the detection model is trained using the image training data, the radar training data and the corresponding two-dimensional detection labels and three-dimensional detection labels, and all parameters of the detection model are updated; After the first training stage, the training of the detection model further includes: When the input training data only includes image training data and two-dimensional detection labels acquired by a camera, the detection model is trained using the image training data and the two-dimensional detection labels, and the parameters of the feature extraction, the two-dimensional detection processing, and the first fusion processing performed on the camera image are updated; When the input training data only includes image training data, two-dimensional detection labels and three-dimensional detection labels collected by the camera, the detection model is trained using the image training data, two-dimensional detection labels and three-dimensional detection labels, and the parameters of the feature extraction, the first fusion processing, the two-dimensional detection processing, the second fusion processing and the three-dimensional detection processing performed on the camera image are updated.
2. The method according to claim 1, characterized in that Before performing a second fusion process on the radar feature and the first fusion feature, the method further includes: performing multi-scale encoding on the radar feature, and using the encoding result as an updated radar feature for performing the second fusion process.
3. The method according to claim 1, characterized in that: Before performing the second fusion processing on the image features and the radar features, the method further includes: performing multi-scale encoding on the first fusion features, and using the encoding results as updated first fusion features for performing the two-dimensional detection processing and the second fusion processing.
4. The method according to claim 1, characterized in that: The first fusion processing of the image feature and the radar feature includes: splicing or adding the image feature and the radar feature, and performing convolution processing on the splicing or adding result to obtain the first fusion feature.
5. The method according to claim 1, characterized in that The directly fusing the combined feature with the radar feature includes: splicing or adding the combined feature with the radar feature, and performing convolution processing on the splicing or adding result.
6. A target detection device, characterized in that: The device comprises: a data acquisition unit, a detection model processing unit, a result output unit and a detection model training unit; The data acquisition unit is used to acquire a camera image acquired by a camera and a radar mapping image generated by point cloud data acquired by a millimeter wave radar; wherein the radar mapping image is a point cloud image mapped to a camera image coordinate system; The detection model processing unit is used to process the camera image and the radar map using a pre-trained detection model to obtain a two-dimensional detection result and a three-dimensional detection result of the target, and output them through the result output unit; The detection model training unit is used to train and generate the detection model by adopting a hybrid training strategy of two-dimensional training data and three-dimensional training data; Wherein, the detection model unit includes an image encoding module, a radar encoding module, a first fusion module, a two-dimensional detection head module, a second fusion module and a three-dimensional detection head module; The image encoding module and the radar encoding module are used to extract features from the camera image and the radar mapping image, respectively, to obtain image features and radar features; The first fusion module is used to perform a first fusion process on the image feature and the radar feature to obtain a primary fusion feature; The two-dimensional detection head module is used to perform two-dimensional detection processing on the primary fusion feature using a two-dimensional detection head to obtain a two-dimensional detection result of the target; The second fusion module is used to perform a second fusion process on the radar feature and the first fusion feature to obtain a secondary fusion feature; The three-dimensional detection head module is used to perform three-dimensional detection processing on the secondary fusion feature using a three-dimensional detection head to obtain a three-dimensional detection result of the target; In the second fusion module, the radar feature and the first fusion feature are subjected to a second fusion process, including: converting the coordinates of each feature pixel of the first fusion feature back to the radar coordinate system and performing coordinate encoding; merging the coordinate encoding result with the first fusion feature to obtain a merged feature; and directly fusing the merged feature with the radar feature. In the detection model training unit, the training of the detection model includes a first training stage and a second training stage performed sequentially; In the first training stage, the detection model is trained using only the image training data collected by the camera and the corresponding two-dimensional detection labels, and the parameters of the feature extraction and the two-dimensional detection processing performed on the camera image are updated; In the second training stage, the detection model is trained using the image training data, the radar training data and the corresponding two-dimensional detection labels and three-dimensional detection labels, and all parameters of the detection model are updated; After the first training stage, the training of the detection model further includes: When the input training data only includes image training data and two-dimensional detection labels collected by the camera, the detection model is trained using the image training data and the two-dimensional detection labels, and the parameters of the image encoding module, the two-dimensional detection process and the first fusion process are updated; When the input training data only includes image training data, two-dimensional detection labels and three-dimensional detection labels collected by the camera, the detection model is trained using the image training data, two-dimensional detection labels and three-dimensional detection labels, and the parameters of the feature extraction, the first fusion processing, the two-dimensional detection processing, the second fusion processing and the three-dimensional detection processing performed on the camera image are updated.
7. The device according to claim 6, characterized in that The detection model processing further includes a multi-scale radar encoding module, which is used to perform multi-scale encoding on the radar features, and uses the encoding results as updated radar features, which are input into the second fusion module and the first fusion features for the second fusion processing.
8. The device according to claim 6, characterized in that The detection model processing unit further includes a multi-scale image encoding module, which is used to perform multi-scale encoding on the first fusion feature, and use the encoding result as the updated first fusion feature, which is input into the second fusion module and the radar feature for the second fusion processing.
9. The device according to claim 6, characterized in that In the first fusion module, the first fusion processing of the image features and the radar features includes: splicing or adding the image features and the radar features, and performing convolution processing on the splicing or adding results to obtain the primary fusion features.
10. The device according to claim 6, characterized in that In the second fusion module, The directly fusing the combined feature with the radar feature includes: splicing or adding the combined feature with the radar feature, and performing convolution processing on the splicing or adding result.
11. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by the processor, the target detection method described in any one of claims 1 to 5 can be implemented.
12. An electronic device, characterized in that: The electronic device includes at least a computer-readable storage medium and also includes a processor; The processor is used to read executable instructions from the computer-readable storage medium and execute the instructions to implement the target detection method described in any one of claims 1 to 5 above.
Citation Information
Patent Citations
All-weather target detection method based on vision and millimeter wave fusion
US20220207868A1
Three-dimensional target detection method based on multimodal fusion and depth attention mechanism
US20250037299A1