3D target detection method based on instance-scene fusion enhancement
By adopting cross-view angle correlation mechanism and an example-scene fusion method of geometric and semantic dual attention mechanism in 3D object detection, the problem of insufficient accuracy and robustness of 3D object detection in the prior art is solved, and efficient detection performance improvement in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510275683.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing multimodal fusion methods are insufficient in accuracy and robustness in 3D object detection, especially in complex scenarios, and lack methods to effectively combine scene-level and instance-level fusion.
A 3D object detection method based on instance-scene fusion enhancement is proposed. The preliminary fusion of instance features and scene features is achieved through cross-view angle correlation mechanisms, and the geometric and semantic dual attention mechanisms are used to achieve deep enhancement, obtain the enhanced instance-scene BEV features, and finally realize the detection and positioning of 3D objects through decoding.
It significantly improves the detection performance of multimodal 3D object detection algorithm in complex scenarios, effectively combining the advantages of scene-level and instance-level fusion.
Smart Images

Figure CN120107536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a 3D object detection method based on instance-scene fusion enhancement, aiming to improve the 3D object detection performance of multi-modal sensor fusion. Background Art
[0002] With the rapid advancement of autonomous driving technology, 3D object detection plays a key role in autonomous driving systems. It enables vehicles to accurately perceive the surrounding environment and obtain important information such as position and speed to ensure driving safety. At present, the fusion of LiDAR and cameras is the main method to achieve three-dimensional object detection. LiDAR provides high-precision geometric depth data, while cameras capture rich semantic information. The complementarity of the two significantly improves the perception capability of 3D object detection. However, LiDAR sensors are expensive and their performance may degrade in severe weather conditions.
[0003] In contrast, millimeter-wave radar (4D Radar) has gradually become an alternative in the field of autonomous driving perception because it can still work stably in bad weather and has low cost. 4D millimeter-wave radar can not only provide distance, speed and angle information of objects, but also obtain height data of objects, giving it higher resolution in complex scenes. However, due to the sparseness and high noise of radar point cloud data, these factors limit the performance of millimeter-wave radar in target detection tasks.
[0004] Existing multimodal fusion methods use rich camera image semantic features to improve the sparse and noisy information of millimeter-wave radar to a certain extent, and thus have made good progress in 3D object detection. However, existing radar-camera fusion methods mainly rely on scene-level fusion or instance-level fusion, and each method has inherent limitations. Scene-level fusion can provide global scene perception, but due to inaccurate perspective conversion and insufficient attention to instance-related features, it leads to feature ambiguity. For example, BEVFusion realizes the fusion of cross-modal information in BEV space, and the LXL method further uses the depth estimated by monocular to distinguish image features in space, but neither of them considers using instance features that are easier to detect in perspective to enhance. Although instance-level fusion performs well in instance feature optimization, it lacks global perception and deep multimodal interaction, which leads to insufficient features. For example, CeterFusion uses 2D object detection to screen the visual cone space, and CRAFT uses an attention mechanism to fuse instances obtained by multimodal detection, but both lack perception of the entire 3D space. Although the two methods have their own advantages, there is currently no suitable method to effectively combine the two. Summary of the invention
[0005] Considering the lack of accuracy and robustness of the existing 3D target detection methods that fuse images and radars, the present invention proposes a 3D target detection method based on instance-scene fusion enhancement. This method realizes the preliminary fusion of instance features and scene features through a cross-view correlation mechanism to obtain instance-scene BEV features; realizes the deep enhancement of instance-scene fusion through a geometric and semantic dual attention mechanism to obtain enhanced instance-scene BEV features; decodes the enhanced instance-scene BEV features to realize the detection and positioning of 3D targets.
[0006] The present invention is specifically implemented by the following technical solutions:
[0007] A 3D object detection method based on instance-scene fusion enhancement, the specific steps are as follows:
[0008] S1: Through the cross-view correlation mechanism, the 2D instance features are transferred to the bird's-eye view to achieve the initial fusion of instance and scene features and obtain instance-scene BEV features. The details are as follows.
[0009] S11: Extract radar BEV features from the millimeter-wave radar point cloud, extract semantics and convert the semantics into a bird's-eye view to obtain camera BEV features, and fuse the two to obtain scene features. The image features are obtained by encoding F 2D Further semantics is obtained through feature extractor Where C is the number of channels in the feature map, H and W represent the height and width of the image respectively, and n represents the downsampling rate. 2D Get depth through feature extractor Where D is the number of predefined discrete depth intervals. Next, the outer product of semantics C and depth D is taken to obtain the three-dimensional spatial feature Finally, F 3D By performing voxel pooling voxelpool(·), we can obtain the camera BEV features and extract the radar BEV features of the millimeter-wave radar point cloud. The camera BEV features and radar BEV features are fused to obtain the scene features. Where X and Y represent the length and width of the BEV space; the three-dimensional spatial feature F 3D and scene feature F S The calculation of is expressed as follows:
[0010] F 3D =D×C,
[0011] F S =voxelpool(F 3D ).
[0012] S12: Perform 2D object detection and aggregation for the semantics described in S11 to obtain aggregated instance features. Specifically, for semantics C, a 2D object detector is used to detect N instances, and feature pooling is used to pool the features in the target boxes of the N instances to obtain instance features. Next, a learnable vector is used to aggregate all instance features to obtain the aggregate instance features Specifically, F I After being concatenated with the vector, a multi-head self-attention mechanism is used to aggregate instance information, and then the vector is taken out to obtain the aggregated instance feature v.
[0013] S13: For the scene feature F in S11 S As well as the aggregated instance feature v in S12, the scene feature and the aggregated instance feature are fused using the cross-view correlation mechanism to obtain the instance-scene BEV feature, which can realize the transfer of the perspective instance feature to the bird's-eye view scene feature. The specific method is:
[0014] F S Obtain foreground scene features through foreground encoder and background encoder Background scene features The aggregate instance features v are respectively related to F o and F b Perform matrix correlation operations to obtain the foreground segmentation mask Segmentation mask with background m o and m b The calculation process is expressed as:
[0015] m o =σ(F o v T ),
[0016] m b =σ(F b v T ),
[0017] where σ(·) represents the sigmoid function.
[0018] The cosine similarity of the foreground segmentation mask and the background segmentation mask is calculated using the foreground scene features and the background scene features respectively. o With F o Calculating the cosine similarity can get the foreground similarity vector v o , m b With F b Calculating the cosine similarity can get the background similarity vector v b , each element in the vector represents the similarity between the scene feature and the mask.
[0019] Finally, v o With F o By performing dot multiplication, we can achieve a preliminary fusion of instance and scene features and obtain instance-scene BEV features.
[0020] In this process, two kinds of losses are used for supervision. On the one hand, the foreground segmentation mask and the background segmentation mask are supervised by using the true foreground occupancy and the true background occupancy. On the other hand, the foreground similarity vector and the background similarity vector are supervised to make their element values close to 1, so as to encourage the foreground features and background features to be close to the shape of the true occupancy during the encoding process.
[0021] S2: Through the dual attention mechanism of geometry and semantics, the instance-related features in the instance-scene BEV features are further enhanced to achieve deep enhancement of instance-scene fusion and obtain enhanced instance-scene BEV features. Specifically, the following steps are included:
[0022] S21: By using the attention mechanism, the instance-scene BEV feature described in S1 is semantically enhanced to obtain the semantically enhanced instance-scene BEV feature. Specifically, first, the instance-scene BEV feature F is used IS Initialize the query in the attention mechanism; then, generate the three-dimensional spatial feature F based on the semantics 3D As the value in the attention mechanism, the semantic attention mechanism is initialized; finally, the instance-scene BEV features are enhanced by the semantic attention mechanism to obtain the semantically enhanced instance-scene BEV features. S21 specifically includes the following steps: (1) Each query is lifted from the BEV plane to the column according to its position in the BEV space, and 3D reference points are sampled from the column, and the 3D reference points are assigned with the corresponding query; (2) The 3D reference points are used as queries and the three-dimensional spatial features are used as values, and these 3D reference points are projected onto F 3D The three-dimensional deformable cross-attention operation is performed in ; the above process is specifically described as:
[0023] For a 3D query Q at position q q , obtained through the three-dimensional deformable attention mechanism 3DDCA operation after semantic enhancement
[0024]
[0025] Where N represents the number of 3D queries at position q sampled in the deformable attention mechanism; represents the camera projection function; A n ∈[0,1] is the learnable attention weight; W represents the feature projection weight; represents the predicted offset to position q; represents trilinear interpolation;
[0026] Finally, all semantically enhanced By reintegrating into the shape of the BEV feature space, we can obtain the semantically enhanced instance-scene BEV features.
[0027] S22: Through the attention mechanism, the instance-scene BEV feature described in S1 is enhanced using the point cloud to obtain a geometrically enhanced instance-scene BEV feature. Specifically, first, the instance-scene BEV feature F is used IS Initialize the query in the attention mechanism; then, use the radar BEV features generated from the point cloud as the value in the attention mechanism to complete the initialization of the geometric attention mechanism; finally, enhance F through the geometric attention mechanism IS , and obtain the geometrically enhanced instance-scene BEV features. S22 specifically includes the following steps: downsampling the instance-scene BEV features and the radar BEV features using convolution respectively, and then using the neighborhood cross attention mechanism (NCA) to interactively fuse the query and value to update the query, and repeating until the downsampling size reaches 8 times; then, upsampling the instance-scene BEV features and the radar BEV features using convolution respectively, and then using the neighborhood cross attention mechanism to interactively fuse the query and value to update the query, and repeating until the size is the same as the input size; the final output query is the geometrically enhanced instance-scene BEV feature.
[0028] S23: Add the semantically enhanced instance-scene BEV feature described in S21 and the geometrically enhanced instance-scene BEV feature described in S22 to obtain an enhanced instance-scene BEV feature.
[0029] S3: Decode the enhanced instance-scene BEV features obtained in step S2 to achieve 3D target detection and positioning.
[0030] The beneficial effects of the present invention are:
[0031] The method of the present invention effectively transfers instance information from the perspective view to the bird's-eye view (BEV) through a cross-view correlation mechanism, thereby achieving a preliminary fusion of instance features and scene features. Secondly, the present invention further enhances the instance-related features in the instance-scene BEV features through a geometric attention mechanism and a semantic attention mechanism, obtains enhanced instance-scene BEV features, and achieves deep enhancement of instance-scene fusion. The enhanced instance-scene BEV features are decoded to achieve detection and positioning of 3D targets. The method of the present invention effectively combines the paradigms of scene-level fusion and instance-level fusion, giving full play to the advantages of both, and experimental results show that the method of the present invention significantly improves the detection performance of the current multimodal 3D target detection algorithm based on deep learning in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of a 3D object detection method based on instance-scene fusion enhancement provided by an embodiment of the present invention.
[0033] Figure 2 4 is a diagram showing the generalization enhancement effect of an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following is further described in conjunction with specific embodiments and drawings.
[0035] Example 1
[0036] The present embodiment provides a 3D target detection device based on instance-scene fusion enhancement, which is used to implement a 3D target detection method based on instance-scene fusion enhancement of the present invention. In the 3D target detection process, 2D instance features are transferred to the bird's-eye view perspective through a cross-perspective correlation mechanism to achieve a preliminary fusion of instance features and scene features, and obtain instance-scene BEV features; through a geometric and semantic dual attention mechanism, the instance-related features in the instance-scene BEV features are further enhanced to achieve deep enhancement of instance-scene fusion and obtain enhanced instance-scene BEV features; the enhanced instance-scene BEV features are decoded to achieve 3D target detection and positioning. Ultimately, deep interaction and fusion of multimodal features can be achieved, thereby improving the performance of 3D target detection.
[0037] The detection device includes a feature extraction module, a feature fusion module, an instance-scene fusion module and a decoding head;
[0038] The feature extraction module is used to extract original millimeter wave radar and monocular image features;
[0039] The feature fusion module is used to obtain BEV features at a bird's-eye view based on the original millimeter-wave radar and monocular image features, and fuse radar BEV features and camera BEV features at a bird's-eye view to obtain scene features;
[0040] The instance-scene fusion module achieves the preliminary fusion of instance features and scene features based on the cross-view correlation mechanism to obtain instance-scene BEV features. And based on the geometric and semantic dual attention mechanism, it achieves the deep enhancement of instance-scene fusion to obtain the enhanced instance-scene BEV features.
[0041] The decoding head is used to decode the enhanced instance-scene BEV features to achieve detection and positioning of 3D targets.
[0042] like Figure 1 The detection process of the detection device can be divided into the following four stages:
[0043] In the first stage, the feature extraction module is used to extract the original millimeter wave radar and monocular image features, such as Figure 1 For the image The image features are obtained by encoding Here C is the number of channels of the feature map, H and W represent the height and width of the image respectively, and n represents the downsampling rate. For millimeter-wave radar, the point cloud encoder is used to encode the millimeter-wave radar point cloud. The process includes steps such as cylinder voxelization, shallow feature extraction, and deep feature extraction to obtain radar BEV features.
[0044] In the second stage, the feature fusion module is used to fuse the radar BEV features and the camera BEV features under the bird’s-eye view to obtain scene features, such as Figure 1 As shown. 2D Further semantics is obtained through feature extractor Where C is the number of channels in the feature map, H and W represent the height and width of the image respectively, and n represents the downsampling rate. 2D Get depth through feature extractor Where D is the number of predefined discrete depth intervals. Next, the outer product of semantics C and depth D is taken to obtain the three-dimensional spatial feature Finally, F 3D By performing voxel pooling voxelpool(·), the camera BEV features can be obtained. Then, the radar BEV features of the first stage are fused with the camera BEV features through a cross-modal fusion network to obtain the scene features. Where X and Y represent the length and width of the BEV space. 3D and scene feature F S The calculation of is expressed as follows:
[0045] F 3D =D×C,
[0046] F S =voxelpool(F 3D ).
[0047] In the third stage, the instance-scene fusion module realizes the preliminary fusion of instance features and scene features based on the cross-view correlation mechanism to obtain the instance-scene BEV features. At this time, the instance-scene fusion is further enhanced based on the geometric and semantic dual attention mechanism to obtain the enhanced instance-scene BEV features. The process is as follows: Figure 1 shown.
[0048] First, the instance and scene features are initially fused based on the cross-view correlation mechanism. Specifically, for semantic C, a 2D object detector is used to detect N instances, and feature pooling is used to pool the features in the target boxes of the N instances to obtain instance features. Next, a learnable vector is used to aggregate all instance features to obtain the aggregate instance features Specifically, F I After being concatenated with the vector, a multi-head self-attention mechanism is used to aggregate instance information, and then the vector is taken out to obtain the aggregated instance feature v. S And the aggregated instance features v, use the cross-view correlation mechanism to initially fuse the scene features with the aggregated instance features to obtain the instance-scene BEV features. Specifically, F S Obtain foreground scene features through foreground encoder and background encoder Background scene features The aggregate instance features v are respectively related to F o and F b Perform matrix correlation operations to obtain the foreground segmentation mask Segmentation mask with background m o and m b The calculation process is expressed as:
[0049] m o =σ(F o v T ),
[0050] m b =σ(F b v T ),
[0051] Where σ(·) represents the sigmoid function. Further, the cosine similarity of the foreground segmentation mask and the background segmentation mask is calculated using the foreground scene features and the background scene features, respectively. o With Fo Calculating the cosine similarity can get the foreground similarity vector v o , m b With F b Calculating the cosine similarity can get the background similarity vector v b , each element in the vector represents the similarity between the scene feature and the mask. Finally, v o With F o By performing dot multiplication, we can achieve a preliminary fusion of instance and scene features and obtain instance-scene BEV features. In this process, two kinds of losses are used for supervision. On the one hand, the foreground segmentation mask and the background segmentation mask are supervised by using the true foreground occupancy and the true background occupancy. On the other hand, the foreground similarity vector and the background similarity vector are supervised to make their element values close to 1, so as to encourage the foreground features and background features to be close to the shape of the true occupancy during the encoding process.
[0052] Then, the instance-related features in the instance-scene BEV features are further enhanced through the geometric and semantic dual attention mechanism to achieve deep enhancement of instance-scene fusion and obtain the enhanced instance-scene BEV features. IS The fully enhanced instance-scene BEV feature is obtained by enhancing it through the geometric attention mechanism and the semantic attention mechanism respectively, and then adding the obtained features. For the semantic attention mechanism, first, use the instance-scene BEV feature F IS Initialize the query in the attention mechanism; then, generate the three-dimensional spatial feature F based on the semantics 3D As the value in the attention mechanism, the semantic attention mechanism is initialized; finally, the instance-scene BEV feature is enhanced by the semantic attention mechanism to obtain the semantically enhanced instance-scene BEV feature. For the geometric attention mechanism, first, the instance-scene BEV feature F is used IS Initialize the query in the attention mechanism; then, use the radar BEV features generated from the point cloud as the value in the attention mechanism to complete the initialization of the geometric attention mechanism; finally, enhance F through the geometric attention mechanism IS , and obtain the geometrically enhanced instance-scene BEV features. Add the semantically enhanced instance-scene BEV features and the geometrically enhanced instance-scene BEV features to obtain the enhanced instance-scene BEV features.
[0053] In the fourth stage, the anchor-based feature decoding head is used to decode the enhanced instance-scene BEV features to achieve 3D object detection and localization.
[0054] Figure 2The results of the method of the present invention on the View-of-Delft dataset and the TJ4DRadSet dataset. Each image corresponds to a data frame containing images and radar points (gray), and the red triangle marks the position of the vehicle. The orange and yellow boxes represent the real boxes in the perspective view and the bird's-eye view, respectively. The green and blue boxes represent the bounding boxes predicted by the method of the present invention, and the upper right corner of the figure shows the visualization of the BEV feature map. The figure shows the detection performance of the method of the present invention for cars, pedestrians, and bicycles in the dataset.
Claims
1. A 3D object detection method based on instance-scene fusion enhancement, characterized in that: Firstly, the 2D instance features are transferred to the bird's-eye view to achieve the preliminary fusion of instance features and scene features, and obtain instance-scene BEV features; then, the instance-related features in the instance-scene BEV features are further enhanced to achieve deep enhancement of instance-scene fusion and obtain enhanced instance-scene BEV features; finally, the enhanced instance-scene BEV features are decoded to achieve 3D target detection and positioning.
2. The 3D object detection method based on instance-scene fusion enhancement according to claim 1, characterized in that: The specific steps are as follows: S1: Through the cross-view correlation mechanism, the 2D instance features are transferred to the bird's-eye view to achieve the initial fusion of instance and scene features and obtain instance-scene BEV features; S11: Extract radar BEV features from the millimeter-wave radar point cloud, extract semantics and convert the semantics into a bird's-eye view to obtain camera BEV features, and fuse the two to obtain scene features; S12: Perform 2D object detection and aggregation for the semantics described in S11 to obtain aggregated instance features; S13: Preliminarily fuse the scene features described in S11 and the aggregated instance features described in S12 through a cross-view correlation mechanism to obtain instance-scene BEV features; S2: Through the dual attention mechanism of geometry and semantics, the instance-related features in the instance-scene BEV features are further enhanced to achieve deep enhancement of instance-scene fusion and obtain enhanced instance-scene BEV features; S21: The instance-scene BEV features described in S1 are semantically enhanced through the attention mechanism to obtain semantically enhanced instance-scene BEV features; S22: The instance-scene BEV feature described in S1 is enhanced using the point cloud through the attention mechanism to obtain the geometrically enhanced instance-scene BEV feature; S23: adding the semantically enhanced instance-scene BEV feature described in S21 and the geometrically enhanced instance-scene BEV feature described in S22 to obtain an enhanced instance-scene BEV feature; S3: Decode the enhanced instance-scene BEV features obtained in step S2 to achieve 3D target detection and positioning.
3. The 3D object detection method based on instance-scene fusion enhancement according to claim 2, characterized in that: Step S11 specifically includes: The image features are obtained by encoding F 2D Get semantics through feature extractor Where C is the number of channels in the feature map, H and W represent the height and width of the image respectively, and n represents the downsampling rate; at the same time, F 2D Get depth through feature extractor Where D is the number of predefined discrete depth intervals; then, the outer product of semantic C and depth D is performed to obtain the three-dimensional spatial feature Finally, F 3D Perform voxel pooling voxelpool(·) to obtain the camera BEV features, and extract the radar BEV features of the millimeter-wave radar point cloud. The camera BEV features and radar BEV features are fused to obtain the scene features. Where X and Y represent the length and width of the BEV space; the three-dimensional spatial feature F 3D and scene feature F S The calculation of is expressed as follows: F 3D =D×C, F S =voxelpool(F 3D )。 4. The 3D object detection method based on instance-scene fusion enhancement according to claim 2, characterized in that: Step S12 is as follows: for semantic C, a 2D object detector is used to detect N instances, and feature pooling is used to pool the features in the target boxes of the N instances to obtain instance features Next, a learnable vector is used to aggregate all instance features to obtain the aggregate instance features Specifically, F I After being concatenated with the vector, a multi-head self-attention mechanism is used to aggregate instance information, and then the vector is taken out to obtain the aggregated instance feature v.
5. The 3D object detection method based on instance-scene fusion enhancement according to claim 2, characterized in that: In step S13: for the scene feature F S And the aggregated instance feature v, use the cross-view correlation mechanism to initially fuse the scene feature and the aggregated instance feature to obtain the instance-scene BEV feature; the specific method is: F S Obtain foreground scene features through foreground encoder and background encoder Background scene features The aggregate instance features v are respectively related to F o and F b Perform matrix correlation operations to obtain the foreground segmentation mask Segmentation mask with background m o and m b The calculation process is expressed as: m o =σ(F o v T ), m b =σ(F b v T ), Where σ(·) represents the sigmoid function; The cosine similarity of the foreground segmentation mask and the background segmentation mask is calculated using the foreground scene features and the background scene features respectively; m o With F o Calculate the cosine similarity to get the foreground similarity vector v o , m b With F b Calculating the cosine similarity can get the background similarity vector v b , each element in the vector represents the similarity between the scene feature and the mask; Finally, v o With F o Perform point multiplication to achieve the initial fusion of instance and scene features and obtain instance-scene BEV features In this process, two kinds of losses are used for supervision; on the one hand, the foreground segmentation mask and the background segmentation mask are supervised by the true foreground mask and the true background mask; on the other hand, the foreground similarity vector and the background similarity vector are supervised to make their element values close to 1.0, so as to encourage the foreground scene features and the background scene features to be close to the shape of the true mask during the encoding process.
6. The 3D object detection method based on instance-scene fusion enhancement according to claim 2, characterized in that: S21 is specifically as follows: First, use the instance-scenario BEV feature F IS Initialize the query in the attention mechanism; then, generate the three-dimensional spatial feature F based on the semantics 3D As the value in the attention mechanism, the semantic attention mechanism is initialized. Finally, the instance-scene BEV features are enhanced through the semantic attention mechanism to obtain the semantically enhanced instance-scene BEV features.
7. The 3D object detection method based on instance-scene fusion enhancement according to claim 2, characterized in that: Step S21 specifically includes the following steps: (1) lifting each query from the BEV plane to the column according to its position in the BEV space, sampling 3D reference points from the column, and assigning values to the 3D reference points with the corresponding query; (2) taking the 3D reference points as queries and the 3D spatial features as values, and projecting these 3D reference points onto F 3D The three-dimensional deformable cross-attention operation is performed in ; the above process is specifically described as: For a 3D query Q at position q q , obtained through the operation of three-dimensional deformable cross-attention mechanism after semantic enhancement Where N represents the number of samples of the 3D query at position q in the 3D deformable cross-attention mechanism; represents the camera projection function; A n ∈[0,1] is the learnable attention weight; W represents the feature projection weight; represents the predicted offset to position q; represents trilinear interpolation; Finally, all semantically enhanced By reintegrating into the shape of the BEV feature space, we can obtain the semantically enhanced instance-scene BEV features.
8. The 3D object detection method based on instance-scene fusion enhancement according to claim 2, characterized in that: S22 is specifically as follows: First, use the instance-scenario BEV feature F IS Initialize the query in the attention mechanism; then, use the radar BEV features generated from the point cloud as the value in the attention mechanism to complete the initialization of the geometric attention mechanism; finally, enhance F through the geometric attention mechanism IS , and obtain the geometrically enhanced instance-scene BEV features.
9. The 3D object detection method based on instance-scene fusion enhancement according to claim 2, characterized in that: Step S22 specifically includes the following steps: downsampling the instance-scene BEV features and the radar BEV features respectively using convolution, and then interactively fusing the query and the value using the neighborhood attention mechanism to update the query, and repeating until the downsampling size reaches 8 times; then, upsampling the instance-scene BEV features and the radar BEV features respectively using convolution, and then interactively fusing the query and the value using the neighborhood attention mechanism to update the query, and repeating until the size is the same as the input size; the final output query is the geometrically enhanced instance-scene BEV feature.
Citation Information
Patent Citations
Double-attention network model training method and vehicle control method
CN118485995A
Automatic driving streetscape unsupervised 4D automatic labeling method based on three-dimensional reconstruction
CN118967934A
3D target detection method based on three-dimensional deformable attention mechanism enhancement
CN119478370A
Target Detection Method and Apparatus
US20230045519A1
Scenario aware method and related device thereof
WO2024217411A1
Cited By
Target detection method combining vision and lidar applied to embodied intelligent agents
CN122574831A