Method for detecting objects and object detection device

The method addresses radar's limitations in height prediction by fusing radar and camera data with spatial attention, enhancing object detection accuracy and overcoming radar's inaccuracies in height information.

WO2025109045A1PCT designated stage expired Publication Date: 2025-05-30AUMOVIO AUTONOMOUS MOBILITY GERMANY GMBH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/EP2024/083066
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Radar systems in advanced driver assistance systems (ADAS) face challenges in accurately predicting height information, limiting their performance in high-level perception tasks like object detection and semantic segmentation.

Method used

A computer-implemented method for object detection that generates feature matrices from image and radar data, fuses these matrices using a feature extractor network, applies spatial attention, and detects objects based on the fused input data.

Benefits of technology

The method effectively integrates radar and camera data to improve object detection accuracy, particularly in scenarios where radar's depth and velocity information enhance the detection of traffic participants, reducing false positives and refining spatial relations among objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024083066_30052025_PF_FP_ABST
    Figure EP2024083066_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for detecting objects includes generating a first feature matrix based on image data of an environment. The method further includes generating a second feature matrix based on radar point cloud data of the environment. The method further includes fusing the first feature matrix and the second feature matrix using a feature extractor network, to result in a third feature matrix. The method further includes applying spatial attention to the third feature matrix to result in a fused input data. The method further includes detecting objects in the environment based on the fused input data.
Need to check novelty before this filing date? Find Prior Art

Description

202306116 1 METHOD FOR DETECTING AND OBJECT DETECTION DEVICE TECHNICAL FIELD

[0001] Various embodiments relate to methods for detecting objects, and an object detection device. BACKGROUND

[0002] Radar has become a fundamental component for advanced driver assistance systems (ADAS) in vehicles due to its cost-effectiveness and weather robustness. Primarily operating through the emission and reflection of radio millimeter waves, radar determines the distance, direction, and relative velocity of surrounding objects. This makes it valuable in performing tasks such as collision detection and avoidance, velocity estimation, and basic tracking of surroundings. However, radar is rarely used for high-level perception like object detection or semantic segmentation due to the sparsity of its point cloud information. In contrast, cameras are high resolution sensors which can capture dense visual features but the data they capture are generally inaccurate for determining distance. Also, their performance may be dependent on the weather. Radar-camera fusion techniques have been developed, to leverage on the respective advantages of the radar and the camera, for object detection applications. However, these techniques are often limited in performance due to the radar’s inability to predict height information accurately.

[0003] In view of the above, a new method for object detection is proposed to address at least some of the abovementioned challenges. SUMMARY

[0004] According to various embodiments, there is provided a computer-implemented method for detecting objects. The method includes generating a first feature matrix based on image data of an environment. The method further includes generating a second feature matrix based on radar point cloud data of the environment. The method further includes fusing the first feature matrix and the second feature matrix using a feature extractor network,202306116 2 to result in a third feature matrix. The further includes applying spatial attention to the third feature matrix to result in a fused input data. The method further includes detecting objects in the environment based on the fused input data.

[0005] According to various embodiments, there is provided a computer program that includes instructions, which, when the program is executed by a computer, causes the computer to carry out the abovementioned method for detecting objects.

[0006] According to various embodiments, there is provided a data carrier signal that carries the abovementioned computer program.

[0007] According to various embodiments, there is provided an object detection device. The object detection device includes a processor configured to perform the abovementioned method.

[0008] Additional features for advantageous embodiments are provided in the dependent claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the drawings, like reference characters generally refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments are described with reference to the following drawings, in which:

[0010] FIG. 1 shows the architecture of an object detection network according to various embodiments.

[0011] FIG. 2 shows a flow diagram of a computer-implemented method for detecting objects, according to various embodiments.

[0012] FIG. 3A shows a block diagram of an object detection device according to an embodiment.

[0013] FIG. 3B shows a block diagram of the object detection device according to another embodiment.

[0014] FIGS. 4A to 4C show a comparative visualization of 3D object detection results between camera-only baseline (ARC-BEV-C), and the method of FIG.2.202306116 3

[0015] Embodiments described below in context of the devices are analogously valid for the respective methods, and vice versa. Furthermore, it will be understood that the embodiments described below may be combined, for example, a part of one embodiment may be combined with a part of another embodiment.

[0016] It will be understood that any property described herein for a specific device may also hold for any device described herein. It will be understood that any property described herein for a specific method may also hold for any method described herein. Furthermore, it will be understood that for any device or method described herein, not necessarily all the components or steps described must be enclosed in the device or method, but only some (but not all) components or steps may be enclosed.

[0017] The term “coupled” (or “connected”) herein may be understood as electrically coupled or as mechanically coupled, for example attached or fixed, or just in contact without any fixation, and it will be understood that both direct coupling or indirect coupling (in other words: coupling without direct contact) may be provided.

[0018] In this context, the object detection device as described in this description may include a memory which is for example used in the processing carried out in the device. A memory used in the embodiments may be a volatile memory, for example a DRAM (Dynamic Random Access Memory) or a non-volatile memory, for example a PROM (Programmable Read Only Memory), an EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or a flash memory, e.g., a floating gate memory, a charge trapping memory, an MRAM (Magnetoresistive Random Access Memory) or a PCRAM (Phase Change Random Access Memory).

[0019] In order that the invention may be readily understood and put into practical effect, various embodiments will now be described by way of examples and not limitations, and with reference to the figures.

[0020] According to various embodiments, a new method for detecting objects is disclosed. The method may be performed by an object detection network 100, which is also referred herein as ARC-BEV. The new method may involve integrating transformed image bird’s eye view (BEV) features and encoded radar BEV features through a spatial attention module (SAM). This method may optimize feature association and may minimize the impact of radar’s inaccurate height information. The method may further include using radar’s depth202306116 4 and velocity information to improve the accuracy of the object detection as compared to using only camera data.

[0021] FIG.1 shows the architecture of an object detection network 100 according to various embodiments. The object detection network 100 may include a camera branch 110, a radar branch 120, and a fusion block 130. The camera branch 110 may receive image data 102 of an environment. The image data 102 may include stacked multi-view images of the environment. The camera branch 110 may be configured to extract multi-modal features from the image data 102, to output a first feature matrix based on the extracted features. The radar branch 120 may receive radar point cloud data 104 of the same environment. The radar point cloud data 104 may be sparse point cloud data, obtained from a 3D radar, or a low-resolution radar. The radar point cloud data 104 may include aggregated point clouds. The radar branch 120 may be configured to extract multi-modal features from the radar point cloud data 104, to output a second feature matrix based on the extracted features. The fusion block 130 may be configured to fuse the first feature matrix and the second feature matrix, to result in a third feature matrix. The fusion block 130 may be further configured to apply spatial attention to the third feature matrix to result in a fused input data. The detection head 140 may be configured to detect objects in the environment, based on the fused input data generated by the fusion block 130. The detection head 140 may generate a detection output 106 that contains information on the detected objects, for example, their positions, their sizes etc.

[0022] The camera branch 110 may include an image backbone 112 and a view transformation module 114. The image backbone 112 may be configured to extract dense feature maps from the image data 102. The view transformation module 114 may be configured to transform the image data from perspective view to BEV, to obtain a unified BEV representation.

[0023] The image backbone 112 may process the image data 102 to produce an imagefeature Ϝ^ ∈ ℝ^×^×^×^ , where N, C, H, and W are number of camera views, number ofimage feature channels,of the image feature, respectively. The view transformation module 114 may estimate the depth distribution of every pixel on the image plane of the image feature Ϝ^, using a convolutional neural network (CNN). The depthdistribution is denoted herein as Ϝ^ ∈ ℝ^×^×^×^ , where D is the number of depth binsgenerated by uniform depth discretization. Subsequently, the view transformation module 114 may reweight the image feature Ϝ^by taking an outer product between image feature Ϝ^202306116 5and depth distribution Ϝ^ , resulting in voxel feature Ϝ^ ∈ ℝ^×^×^×^×^ according toEquation (1). Ϝ^ ^ Ϝ^ ⊗ Ϝ^ Equation (1)

[0024] Next, the view transformation module 114 may collapse the multi-view voxel feature Ϝ^into a unified BEV feature space. The view transformation module 114 may do so using a BEV pooling operation. The BEV pooling operation may include associating a BEV grid to each voxel feature Ϝ^^,^,^, and then aggregate the features within each BEV grid by average pooling, max-pooling or summation processes, according to Equation (2). Ϝ^ ^ ^^^^^^^^^^^^Ϝ^^^^!" #Equation (2)where Ϝ^ ∈ ℝ^$×%×& is the aggregated BEV feature within X × Y BEV grids. The aggregatedBEV feature is also referred herein as the image features 150 generated by the camera branch 110. The image features 150 is also referred herein as the first feature matrix.

[0025] An example of a suitable image backbone 112 may be Swin-Tiny, which is disclosed in “Swin transformer: Hierarchical vision transformer using shifted windows” by Liu et. al., and published in IEEE International Conference on Computer Vision, 2021. An example of the view transformation module 114 may include the Lift, Splat, Shoot (LSS) model as disclosed in “Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D” by Philion et. al., pages 194-210 of Computer Vision ECCV 2020. An example of a BEV pooling operation is the BEVfusion operation disclosed in “Bevfusion: Multi-task multi sensor fusion with unified bird’s eye view representation” by Liu et. al., pages 2774-2781, IEEE 2023.

[0026] The radar branch 120 may include a voxelizer 122 and a sparse encoder 124. For a radar point cloud with K radar points, the input to the radar stream may include location andradial velocity of each radar point, denoted as ℛ ^ ^(^ ^ )*, +, ,, -., - / 01^!". -.is the x-axis component of the radial velocity while - / is the y-axis component of the radial velocity. The radial velocity is the Doppler velocity of each radar point compensated by ego-motion, which does not necessarily match with the actual velocity vector of the corresponding object. To facilitate feature extraction, the radar point cloud data 104 is converted into a voxelized format using the voxelizer 122. The output of the voxelizer 122 is referred herein as voxelized data output. Subsequently, sparse convolution is applied

[0012] on the voxelized data output using the sparse encoder 124, to extract radar features according to Equation (3). Ϝ2 ^ 3^4^^-^^^*5^^,5^ℛ## Equation (3)202306116 6where Ϝ2 ∈ ℝ^6×%×&×7 and SPConv sparse convolution. Because radar hasinaccurate height information, the radar feature Ϝ2may be flattened along the z-axis to Ϝ28∈ ℝ^6×%×&, which becomes the BEV representation, also referred herein as the radar features 152 generated by the radar branch 120. The radar features 152 is also referred herein as a second feature matrix.

[0027] An example of the sparse encoder 124 is disclosed in “Second: Sparsely embedded convolutional detection” by Yan et.al., Sensors (Basel, Switzerland), 18(10): 3337, 2018.

[0028] The fusion block 130 may include a bird’s eye view (BEV) encoder 132 and a spatial attention module 134. As both the camera branch 110 and the radar branch 120 have generated BEV representation of the image features and the radar features respectively, the image features and the radar features may be fused by concatenation. Simple concatenation may present challenges due to the lack of depth information from the image features and the inherent noise in radar point cloud data, which may result in suboptimal feature alignment. To address this, a spatial attention mechanism is employed to correct potential local misalignment, using the spatial attention module 134. The BEV encoder 132 may be configured to concatenate the image features 150 and the radar features 152. The BEVencoder 132 may further derive multi-scale feature Ϝ^9^ ∈ ℝ^:;<×%×& . The multi-scalefeature may be derived using a feature pyramid network (FPN). To optimize computational efficiency, the BEV encoder 132 may restrict feature extraction to the 1× and 2× scales. Subsequently, spatial attention is applied to Ϝ^9^using the Spatial Attention Module (SAM) 134. The SAM 134 may execute average and maximum pooling on Ϝ^9^along the channel dimension, and then concatenate the resulting feature maps. The SAM 134 may apply a convolutional layer with a sigmoid activation function to Ϝ^9^, to generate the spatial attention weight, as expressed in Equation (4). The SAM 134 may multiply the spatial attention weight with the initial BEV feature map, as expressed in Equation (5). =^ 3^^>^^?^4^^-)@^Ϝ^9^#, ℳ^Ϝ^9^#0# Equation (4)Equation (5)where = ,the reweighted BEV feature, theAverage pooling operation and the Max pooling operation, respectively. The output of the fusion block 130 is Ϝ^9^8, also referred herein as the fused BEV feature map or the fused input data 154.202306116 7

[0029] An example of the FPN is in “Second: Sparsely embedded convolutional detection” by Yan et.al., Sensors (Basel, Switzerland), 18(10): 3337, 2018. An example of the SAM 134 is disclosed in “CBAM: Convolutional Block Attention Module” by Ferrari et. al., Volume 11211 of Computer Vision – ECCV 2018, pages 3–19.

[0030] The detection head 140 may be configured to detect objects based on the fused input data 154 generated by the fusion block 130. The detection head 140 may execute an input- dependent initialization strategy based on a center heatmap, using a single decoder layer, to achieve competitive performance. The detection head 140 may generate the heatmap based on the fused BEV feature map Ϝ^9^8, i.e., the fused input data 154. The heatmap is denotedherein as B ∈ ℝ1×%×& where K denotes the number of categories and X × Y describes the sizeof the BEV feature map, from which top-N local maximum elements are selected to initialize object queries. A single attention decoder layer may be applied to the heatmap to decode the queries into 3D objects. The single attention decoder layer may incorporate both self- attention that reasons pairwise relations between different object candidates, as well as cross attention that aggregates relevant context onto the object candidates. Object queries with rich instance information may be decoded into boxes and class labels by a feed-forward network (FFN). A bipartite matching between the predictions and ground truth objects may be performed through the Hungarian algorithm. A focal loss is computed for classification and the bounding box regression is supervised by an L1 loss for only positive pairs. For the heatmap prediction, a penalty-reduced focal loss is computed. The total loss may be computed as the weighted sum of losses for each component, as expressed in Equation (6). The weighting coefficients of heatmap loss, classification loss, and regression loss may be 1.0, 1.0 and 0.25, respectively. CDEDFG ^ H"C^IFDJFK + HMCNGO + HPCQIR Equation (6)

[0031] An “Bevfusion: Multi-task multi-sensorfusion with unified bird’s eye view representation”, by Liu et. al., published in 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 2774–2781. An example of the single attention decoder layer is disclosed in “End-to-End Object Detection with Transformers” by Carion et. al., pages 213–229, Computer Vision – ECCV 2020. An example of computing the penalty-reduced focal loss is disclosed in “Center-based 3d object detection and tracking”, by Yin et. al., published in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 11784-11793, 2021.202306116 8

[0032] FIG. 2 shows a flow diagram of a implemented method 200 for detecting objects, according to various embodiments. The method 200 may be carried out using the object detection network 100. The method 200 includes processes 202, 204, 206, 208 and 210. Process 202 may include generating a first feature matrix based on image data 102 of an environment. The camera branch 110 may perform the process 202. The first feature matrix may include the image features 150. Process 204 may include generating a second feature matrix based on radar point cloud data 104 of the environment. The second feature matrix may include radar features 152. The radar branch 120 may perform the process 204. Process 206 may include fusing the first feature matrix and the second feature matrix using a feature extractor network, to result in a third feature matrix. The feature extractor network may include the BEV encoder 132. Process 208 may include applying spatial attention to the third feature matrix to result in a fused input data 154. The SAM 134 may perform the process 208. The processes 206 and 208 may be carried out by the fusion block 130. Process 210 may include detecting objects in the environment based on the fused input data 154. The detection head 140 may perform the process 210.

[0033] The method 200 may transform stacked multiple images into BEV features and fuse it efficiently with encoded radar BEV features using an attention block. The attention block may include the SAM 134. Fusing the features in BEV space may minimize alignment loss in feature association and may mitigate the impact of effective height information lack from radar. Also, the method 200 may be extended to incorporate more information or additional modalities. By applying spatial attention fusion using the SAM 134, the method 200 may effectively capture the most relevant features in the spatial domain of camera and radar. The method 200 is also evaluated against state-of-the-art object detection methods and was found to achieve significant improvement compared to camera-only baseline and outperforms other camera-radar fusion detection methods. Further, the method 200 may achieve the improved performance using sparse radar data, unlike conventional methods that may require 4D radar data.

[0034] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, the object detection network 100 may further include at least one other sensor branch configured to process sensor data from a corresponding at least one other sensor, for example, LiDAR. The object detection network 100 may be configured to learn the intermodal association between the sensor data of all the sensors, including the camera, the radar and the at least one other sensor.202306116 9

[0035] According to an embodiment may be combined with the above-described embodiment or with any below described further embodiment, generating the first feature matrix comprises generating a perspective view image feature matrix based on the image data, using an image backbone 112 configured to extract features from the image data. The image data 102 may include perspective view images, and the image backbone 112 may extract features from these perspective view images.

[0036] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, generating the first feature matrix further comprises generating the first feature matrix based on the perspective view image feature matrix using a view transformation module 114 configured to transform perspective view representation of the environment to bird’s eye view representation of the environment.

[0037] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, generating the first feature matrix based on the perspective view image feature matrix comprises estimating depth distribution for every pixel in the perspective view image feature matrix using a CNN, and re- weighting the perspective view image feature matrix based on the estimated depth distribution to result in a voxel feature matrix. Estimation the depth distribution dynamically using the CNN provides a more accurate estimation while re-weighting the perspective view image feature matrix results in the BEV feature matrix.

[0038] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, generating the first feature matrix based on the perspective view image feature matrix further comprises associating a bird’s eye view grid to each voxel feature of the voxel feature matrix, and aggregating features within each bird’s eye view grid.

[0039] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, the image data 102 comprises stacked multi-view images.

[0040] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, the method 200 further includes generating the radar point cloud data based on location and radial velocity of each radar-generated data point.202306116 10

[0041] According to an embodiment may be combined with the above-described embodiment or with any below described further embodiment, the radar point cloud data is generated further based on radar cross section (RCS) information. Including the radar specific attributes such as radial velocity and RCS as information in addition to geometric or positional information, allows the fusion block 130 to better learn the association between camera and radar.

[0042] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, generating the second feature matrix comprises generating a radar feature matrix based on the radar point cloud, and further comprises flattening the radar feature matrix along a height axis. The radar point cloud may be captured from a radar that generates sparse point cloud that lacks accurate height information. As the point cloud lacks accurate height information, the radar feature matrix is flattened in the height dimension to prevent inaccuracies due to errors in the height information.

[0043] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, generating the radar feature matrix comprises converting the radar point cloud to a voxelised radar data matrix, and applying sparse convolution on the voxelised radar data matrix.

[0044] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, the feature extractor network comprises a feature pyramid network.

[0045] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, applying spatial attention to the third feature matrix to result in a fused input data comprises generating a spatial weight map based on the third feature matrix, and multiplying the spatial weight map with the third feature matrix.

[0046] The method according to any preceding claim, wherein detecting objects in the environment comprises generating a heatmap based on the fused input data, identifying a predefined number of top local maximum elements in the heatmap, and initializing object queries based on the identified predefined number of top local maximum elements.

[0047] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, detecting objects in the environment further comprises converting the initialized object queries into 3D object queries202306116 11 using an attention decoder layer, the attention decoder layer is configured to determine relationships between different object candidates and is further configured to aggregate relevant context onto the object candidates. The attention decoder layer may include the SAM 134. Use of the attention decoder layer may refine the accuracy of spatial relations among objects and background, thereby improving the accuracy of object detection.

[0048] According to an embodiment which may be combined with the above-described embodiment or with any below described further embodiment, detecting the objects in the environment further comprises decoding the 3D object queries into boxes and class labels by a feed-forward network.

[0049] FIG. 3A shows a block diagram of an object detection device 300 according to an embodiment. The object detection device 300 may include a processor 302. The processor 302 is configured to perform the method 100.

[0050] FIG. 3B shows a block diagram of the object detection device 300 according to another embodiment. The object detection device 300 may further include at least one of a camera 304 and a radar 306. The camera 304 may be configured to capture the image data 102. The radar 306 may be configured to generate radar data that is used for generating the radar point cloud data 104. For example, positional coordinates (x, y, z) and radar attributes such as RCS and rotational velocity, may be extracted from the radar data to generate the radar point cloud data 104. The processor 302, the camera 304 and the radar 306 may be coupled to one another, for example electrically and / or communicatively, by coupling lines 330.

[0051] According to various embodiments, a computer program may include instructions which can be executed by a computer, so that the computer carries out the method 200.

[0052] According to various embodiments, a data carrier signal may carry the computer program.

[0053] In the following, experiments to validate the performance of the method 200 are described.

[0054] The method 200 was evaluated using the nuScenes dataset. The nuScenes dataset is suitable for the evaluation of the method 200, as it includes both camera images and radar point cloud data with ground truth labels for 3D object detection. The nuScenes dataset is collected from 6 surrounding cameras, 5 radars that cover 360 degrees and 1 Lidar. The dataset has 1000 scenes divided into 700 / 150 / 150 scenes for training, validation and testing respectively. The official mean Average Precision (mAP) and nuScenes detection score202306116 12 (NDS) were used as the evaluation The object detection network 100 was implemented in PyTorch. The input image size was set as 256x704 resolution, which has comparable performance and does not require high computation power. The data points from all 5 radars were gathered, and previous 6 radar sweeps were transformed into the current frame considering the ego-motion to generate a denser point cloud.

[0055] During the training process, the initial weights of the camera branch 110 were loaded. The Swin-Tiny pre-trained weights with nuimages were used as the initial weights. The AdamW optimizer with cyclic learning policy was used as the training strategy, for training 10 epochs. To minimize the potential performance loss of camera branch and to ensure asmooth gradient descent, the initial learning rate is set as 4 × 10VW, the target ratio is set as 6and the step ratio up is set as 0.3 in the cyclic learning policy. For data augmentation, the copy-and paste augmentation strategy is used, in addition to except the general random flipping, scaling and rotation. The number of generated ground truth (GT) objects is specifically optimized based on the radar ability for different categories. The settings used for GT generation for data augmentation are shown in Table 1. Category Car Truck Bus Trailer Construct- Pedes- Motor- Bicycle Traffic Barrier ion vehicle trian cycle cone Number of 2 3 4 6 7 10 6 6 2 2 generated GT Table 1. Setting of GT generation for data augmentation

[0056] For a fair comparison with state-of-the-art methodologies, all of the methods were tested under the same conditions, using single-frame input trained on the nuScenes dataset only. Table 2 shows a comparison of the metrics achieved by each method. Method Modal mAP ↑ mATE ↓ mASE ↓ mAOE↓ mAVE ↓ mAAE↓ NDS ↑ ARC-BEV-C Camera 0.362 0.628 0.261 0.591 1.090 0.226 0.411 CenterFusion Radar + 0.326 0.631 0.261 0.516 0.614 0.115 0.449 Camera CRAFT Radar + 0.411 0.467 0.268 0.456 0.519 0.114 0.523 Camera Method 200 Radar + 0.437 0.496 0.279 0.680 0.719 0.195 0.481 Camera Table 2. State-of-the-art comparison on the nuScenes test set.202306116 13

[0057] The method 200 outperforms the competing camera-radar methods tested, by at least 2.6% in the main metric mAP for object detection. This indicates that compared to previous methods, which decorated radar features and performed camera-radar fusion in the image view, the method 200 can achieve more effective fusion in the BEV space. Moreover, the method 200 surpasses ARC-BEV-C (i.e., using image data 102 only), by 7.5%, which reveals the value of automotive radar for the multi-sensor perception system despite its sparse point cloud.

[0058] The per-class map comparison with other methods on the validation (val) set is shown in Table 3. The method 200 beats other camera-radar fusion method in most categories and outperforms the camera-only baseline in all categories. With radar’s velocity and depth information, the method 200 achieves a remarkable performance boost for most traffic participators like +16.9% for car and +13.2% for bus comparing with camera-only baseline ARC-BEV-C. Meanwhile, the method 200 compared to the state-of-the-art camera-radar method CRAFT, has substantial performance boost in most categories like +5.2% for bicycle and +4.6% for truck. Method Modal Car Truck Bus Trai- Constru- Pede- Motor Bicycle Traffic Barrier mAP ler ction strian -cycle cone vehicle ARC-BEV-C Camera 53.8 29.0 40.4 17.7 9.2 38.6 34.1 26.4 54.4 51.8 35.5 CenterFusion Radar + 52.4 26.50 36.2 15.4 5.5 38.9 30.5 22.9 56.3 47.0 33.2 Camera CRAFT Radar + 69.6 37.60 47.3 20.1 10.7 46.2 39.5 31.0 57.1 51.1 41.1 Camera Method 200 Radar + 70.7 42.20 50.4 19.2 14.6 44.4 41.9 36.2 54.7 58.9 43.3 Camera Improvement vs. +16.9 +13.2 +10.0 +1.5 +5.4 +5.8 +7.8 +9.8 +0.3 +7.1 +7.8 Camera Improvement vs. +1.1 +4.6 +3.1 -0.9 +3.9 -1.8 +2.4 +5.2 -2.4 +7.8 +2.2 Radar + Camera Table 3. Per-class AP comparison on the nuScenes val set.

[0059] The nuScenes dataset and 3D object detection benchmark are disclosed in “nuscenes: A multimodal dataset for autonomous driving”, by Caesar et. al., in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 11621-11631, 2020, and https: / / www.nuscenes.org / nuscenes.202306116 14

[0060] The CenterFusion method is in “Centerfusion: Center-based radar and camera fusion for 3d object detection” by Nabati et. al., Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, pages 1527–1536, 2021.

[0061] The CRAFT method is disclosed in “Craft: Camera radar 3d object detection with spatio-contextual fusion transformer” by Kim et. al., Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 1160-1168, 2023.

[0062] FIGS. 4A to 4C is a comparative visualization of 3D object detection results between camera-only baseline (ARC-BEV-C), and the method 200. The first row 410 shows the results from the camera-only baseline (ARC-BEV-C), the second row 420 shows the results from the method 200, and the third row 430 shows the ground truth. In each result figure, camera view (up) and BEV (down) are shown. Failure cases for camera-only approach are highlighted with white dashed circles. Lower confidence score 3D bounding boxes was filtered out for clarity. The method 200 displays notable detection performance across diverse scenarios. Compared to the camera-only baseline, the method 200 is capable of detecting objects from longer distance as well as enhancing the depth estimation for bounding boxes, which is attributed to radar’s long detection distance and depth information. Furthermore, the incorporation of radar data and the spatial attention module reduces the occurrence of false positives and refines the accuracy of spatial relations among objects.

[0063] While embodiments of the invention have been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims. The scope of the invention is thus indicated by the appended claims and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced. It will be appreciated that common numerals, used in the relevant drawings, refer to components that serve a similar or the same purpose.

[0064] It will be appreciated to a person skilled in the art that the terminology used herein is for the purpose of describing various embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements,202306116 15 and / or components, but do not preclude presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0065] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed is an illustration of exemplary approaches. Based upon design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged. Further, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not meant to be limited to the specific order or hierarchy presented.

[0066] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, the term “some” refers to one or more. Combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof” include any combination of A, B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof” may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combinations may contain one or more member or members of A, B, or C. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims.

Claims

202306116 16 CLAIMS 1. A computer-implemented method (200) for detecting objects, the method (200) comprising: generating a first feature matrix based on image data (102) of an environment; generating a second feature matrix based on radar point cloud data (104) of the environment; fusing the first feature matrix and the second feature matrix using a feature extractor network, to result in a third feature matrix; applying spatial attention to the third feature matrix to result in a fused input data (154); and detecting objects in the environment based on the fused input data (154).

2. The method (200) according to any preceding claim, wherein generating the first feature matrix comprises generating a perspective view image feature matrix based on the image data (102), using an image backbone configured to extract features from the image data (102).

3. The method (200) according to claim [0035] , wherein generating the first feature matrix further comprises generating the first feature matrix based on the perspective view image feature matrix using a view transformation module configured to transform perspective view representation of the environment to bird’s eye view representation of the environment.

4. The method (200) according to claim [0036] , wherein generating the first feature matrix based on the perspective view image feature matrix comprises estimating depth distribution for every pixel in the perspective view image feature matrix using a convolutional neural network, and re-weighting the perspective view image feature matrix based on the estimated depth distribution to result in a voxel feature matrix.

5. The method (200) according to claim [0037] , wherein generating the first feature matrix based on the perspective view image feature matrix further comprises202306116 17 associating a bird’s eye view grid to voxel feature of the voxel feature matrix, and aggregating features within each bird’s eye view grid.

6. The method (200) according to any preceding claim, wherein the image data (102) comprises stacked multi-view images.

7. The method (200) according to any preceding claim, further comprising: generating the radar point cloud data (104) based on location and radial velocity of each radar-generated data point.

8. The method (200) according to claim 7, wherein the radar point cloud data (104) is generated further based on radar cross section information.

9. The method (200) according to any preceding claim, wherein generating the second feature matrix comprises generating a radar feature matrix based on the radar point cloud data (104), and further comprises flattening the radar feature matrix along a height axis.

10. The method (200) according to claim [0041] , wherein generating the radar feature matrix comprises converting the radar point cloud data (104) to a voxelised radar data matrix, and applying sparse convolution on the voxelised radar data matrix.

11. The method (200) according to any preceding claim, wherein the feature extractor network comprises a feature pyramid network.

12. The method (200) according to any preceding claim, wherein applying spatial attention to the third feature matrix to result in a fused input data (154) comprises generating a spatial weight map based on the third feature matrix, and multiplying the spatial weight map with the third feature matrix.

13. The method (200) according to any preceding claim, wherein detecting objects in the environment comprises generating a heatmap based on the fused input data (154), identifying a predefined number of top local maximum elements in the heatmap, and202306116 18 initializing object queries based on identified predefined number of top local maximum elements.

14. The method (200) according to claim [0046] , wherein detecting objects in the environment further comprises converting the initialized object queries into 3D object queries using an attention decoder layer, wherein the attention decoder layer is configured to determine relationships between different object candidates and is further configured to aggregate relevant context onto the object candidates.

15. The method (200) of any claim [0047] , wherein detecting the objects in the environment further comprises decoding the 3D object queries into boxes and class labels by a feed-forward network.

16. A computer program comprising instructions, which, when the program is executed by a computer, causes the computer to carry out the method (200) according to any preceding claims.

17. A data carrier signal carrying the computer program of claim 16.

18. An object detection device (300) comprising: a processor (302) configured to perform the method (200) according to any one of claims 1 to [0048] .

19. The object detection device (300) of claim 18, further comprising: a camera (304) configured to capture the image data (102); and a radar (306) configured to generate radar data for generating the radar point cloud data (104).

Citation Information

Cited By

  • Camouflage target detection method

    CN121213900A