Multi-point radar camera fusion detection method based on aerial view
By introducing a multi-point radar camera fusion detection method with radar-assisted view conversion and multimodal deformable attention mechanism, the problem of insufficient 3D perception accuracy in low light or inclement weather by autonomous driving systems is solved, and the accuracy of target detection and system robustness are improved.
Patent Information
- Application Number
- CN202510492566.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
In existing autonomous driving systems, the camera-based 3D perception scheme is susceptible to interference in low light or inclement weather, with large positioning errors, resulting in inaccurate detection of obstacles, affecting path planning and driving safety.
The multi-point radar camera fusion detection method based on bird's eye view is adopted, and the radar feature maps in BEV are aggregated to enhance the accuracy of 3D target recognition by introducing the radar-assisted view conversion method RVT and the multi-modal deformable attention mechanism.
It improves the target detection accuracy and generalization capabilities in harsh environments, enhances the system's ability to identify 3D targets, reduces the environmental impact, and improves the safety and reliability of autonomous driving.
Smart Images

Figure CN120356186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving target perception, and in particular to a multi-point radar camera fusion detection method based on a bird's-eye view. Background Art
[0002] Autonomous driving systems require an accurate and efficient 3D perception system that covers 3D object detection, tracking, and segmentation. Although camera-based solutions have shown some potential in recent years, they are easily disturbed by low-light environments or bad weather and have large positioning errors. The consequences of these errors may threaten the safety and reliability of autonomous driving systems. For example, in low light or bad weather, the system may not be able to accurately detect obstacles or other traffic participants, resulting in an increased risk of collision. In addition, large positioning errors may prevent the vehicle from accurately judging its position on the road, affecting the accuracy of path planning and decision-making, which in turn affects the driving experience or causes traffic accidents.
[0003] As a top-down perspective, BEV can better represent the real world. When integrating different views, modalities, time series and spatial information, it provides a principle of physical interpretability, and has great development prospects in the field of autonomous driving. In the existing technology, radar and camera BEV perception fusion is mostly used to effectively reduce costs while reducing environmental impact and maintaining qualified perception capabilities. However, this method generally lacks spatial information, and there is a spatial misalignment between different sensor data such as cameras and radars, which in turn affects the perception accuracy, may cause misjudgment in scenarios such as autonomous driving, and bring certain hidden dangers to the driver's personal safety. Summary of the invention
[0004] Purpose of the invention: To solve the problems mentioned in the background technology, the present invention discloses a multi-point radar camera fusion detection method based on a bird's-eye view. On the basis of adopting the BEVDepth model, a radar-assisted view conversion method RVT is introduced to make up for the shortcomings of the image in depth perception, and a multimodal deformable attention mechanism is used to further aggregate the image and radar feature map in the BEV, so as to further enhance the recognition accuracy of the model and the system for 3D targets.
[0005] Technical solution:
[0006] The present invention discloses a multi-point radar camera fusion detection method based on a bird's-eye view, the method comprising the following steps:
[0007] S1 obtains the target image and radar data to be detected and divides them into a training set and a test set;
[0008] S2 constructs an improved BEVDepth network model, introduces the radar-assisted view transformation method RVT, and generates a semantically rich and spatially accurate BEV feature map by fusing the complementary characteristics of camera and radar sensors;
[0009] S3 introduces a cross-modal dynamic feature guidance mechanism DCMG in the multi-modal feature aggregation layer of the improved BEVDepth network model: through cross-modal dynamic gating, it realizes weight adaptive allocation, uses deformable spatial convolution to align the image features and radar features of the BEV feature map, extracts channel-level global context descriptions, combines with the Sigmoid function to generate dynamic channel weights, and obtains the fused features of radar and camera through per-channel weighted fusion;
[0010] S4 inputs the dataset to train the improved object detection model;
[0011] S5 uses the trained network model to detect the image to be detected.
[0012] Furthermore, the specific steps of S2 are as follows:
[0013] S2.1 Image feature encoding and depth distribution: Given a set of data containing N panoramic images, use an image backbone network combined with a feature pyramid network to extract a 16-fold downsampled feature map F1 for each perspective image, and further extract the image context feature C1 in the perspective space through an additional stacked convolutional layer according to the LSS method PV ∈R N×C×H×W and the depth distribution D of each pixel ∈R N×D×H×W;
[0014] S2.2 Introduce radar-assisted view transformation, radar feature encoding, and radar occupancy grid generation: Project the radar points onto N camera perspectives to find the corresponding image pixels while retaining their depth information, and then obtain the camera frustum view voxel V FV (d, u, v), where u and v are pixel units in the width and height directions of the image respectively, d is the metric unit in the depth direction, set v = 1 to adopt a columnar structure, and perform feature encoding on the non-empty radar columns through PointNet and sparse convolution to obtain F2 ∈R N×C×D×W , extract the radar context feature C2 in the frustum view space FV ∈R N×C×D×W and the radar occupancy grid O ∈R N×1×D×W , and the convolution operation here is applied to the top view coordinates (d, u);
[0015] S2.3 Transform the image context feature map into the camera frustum view through frustum view transformation; transform the camera and radar context feature maps finally located in N camera frustum views into a single BEV feature map through bird's-eye view transformation.
[0016] Furthermore, the specific processes of the frustum view transformation and the bird's-eye view transformation are as follows:
[0017] Frustum view transformation:
[0018] Given the depth distribution D1, the radar occupancy O, and the image context feature map C1 PV is converted into the camera frustum view C1 FV ∈R N×C×D×H×W , and its conversion method is as follows:
[0019]
[0020] where [·;·] represents the concatenation operation along the channel dimension, represents the outer product operation;
[0021] Bird's-eye view transformation:
[0022] The camera and radar context feature map F finally located in N camera frustum views FV ={C1 FV , C2 FV}∈R N ×C×D×H×W is converted into a single BEV feature map R C×1×X×Y :
[0023]
[0024] where is the feature map of the i-th frustum view, where M(·) represents the view transformation module.
[0025] Furthermore, the specific content of the view transformation module M is as follows: Adopt the voxel pooling method supported by CUDA acceleration, and use average pooling to replace the summation operation to aggregate the features within each BEV grid. By normalizing the feature mean within each BEV grid, the feature magnitude imbalance caused by the difference in the number of frustum grids is eliminated.
[0026] Furthermore, the specific steps of S3 are as follows:
[0027] Cross-modal dynamic gating adaptively assigns modal weights according to the scene context based on the spatially aligned features in the BEV feature map: cross-modal channel-level statistics are extracted through global average pooling, and lightweight MLP is used to generate dynamic gating weights to fuse image and radar features in a channel-wise weighted form. The cross-modal consistency loss constrains the semantic alignment of the above process through contrastive learning during the training phase, forcing the image and radar feature pairs of the same object to closely aggregate in the embedding space, while the feature pairs of different objects are significantly separated. Through semantic-level constraints on the intermediate features after spatial alignment, if the fused features suffer from semantic drift due to environmental interference, the contrastive loss will adjust the spatial alignment offset prediction and channel weight allocation strategy through gradient backpropagation.
[0028] Furthermore, the cross-modal dynamic gating operates as follows:
[0029] Cross-modal dynamic gating adopts an adaptive weight allocation mechanism for the credibility differences between camera and radar features in different scenarios. The input is the image features aligned in the spatial dimension and radar features After that, they are concatenated along the channel dimension into Channel-level global context descriptors are extracted through global average pooling The middle layer dimension of the two-layer MLP is compressed to C / 4, the activation function is ReLU, and combined with the Sigmoid function to generate dynamic channel weights W gate ∈[0,1] C , and finally obtained through channel-wise weighted fusion. The acquisition formula is as follows:
[0030] F fused =W gate ⊙F cam +(1 - W gate )⊙F radar
[0031] where F cam is the image feature aligned in the spatial dimension, F radar is the radar feature, W gate is the dynamic channel weight, F fused is the fused feature, and ⊙ is matrix dot multiplication.
[0032] Furthermore, the specific process of the spatial alignment is as follows:
[0033] Deformable spatial alignment is based on the high-precision spatial information of radar features. The deformable convolutional layer predicts 9 offsets and modulation factor Δm n ∈[0,1], corresponding to the sampling points of the 3×3 convolutional kernel. Through bilinear interpolation resampling of the image features, spatially aligned features are generated:
[0034]
[0035] Among them, The coordinate offset of the nth sampling point is predicted by radar features; Δm n ∈[0,1]: The modulation factor weight of the nth sampling point; p: The target position coordinates in the BEV space, which are made to exactly match the radar features in the BEV space, so as to eliminate the geometric deviation between cross-modal feature maps and overcome the sub-pixel level spatial misalignment.
[0036] Furthermore, the cross-modal consistency loss constrains the semantic consistency between modalities through a contrastive learning framework. Using the same target image-radar features associated with the annotation box as the positive sample pair and different target features randomly sampled within the batch as the negative sample pair, the cosine similarity is calculated and a loss function is constructed. The formula is as follows:
[0037]
[0038] Among them, s(·): Cosine similarity, τ: Temperature coefficient, which controls the sharpness of the similarity distribution, B: Batch size, i, j: The i / jth sample in the batch.
[0039] Beneficial effects:
[0040] 1. The present invention improves the BEVDepth network model, introduces the Radar-aided View Transformation (RVT) in the image branch as a spatio-temporal feature extractor, improves the feature extraction ability, obtains a BEV feature map with rich semantics and accurate space, strengthens the multi-point radar-camera fusion detection and fusion ability, and improves the accuracy and generalization ability of the model for detecting targets such as vehicles and pedestrians.
[0041] 2. The present invention introduces the Dynamic Cross-modal Guidance (DCMG), uses the channel gating mechanism to dynamically fuse the image semantic information and the radar spatial features, and the gating weights are dynamically adjusted through the temporal confidence of the historical fusion features, significantly enhancing the adaptability to harsh environments such as rain and fog. Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the specific process of the method of the present invention;
[0043] Figure 2 It is a schematic diagram of the improved BEVDepth network model structure of the present invention;
[0044] Figure 3 It is a schematic diagram of the Radar-aided View Transformation (RVT) of the present invention;
[0045] Figure 4 It is a schematic diagram of the cross-modal dynamic feature guidance of the multi-modal feature aggregation layer of the present invention;
[0046] Figure 5 This is the schematic diagram of the final detection result of the embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] As Figure 1 shown, the present invention discloses a multi-point radar-camera fusion detection method based on an aerial view. The method steps are as follows:
[0049] S1: Obtain the target image to be detected and radar data, preprocess the image data and point cloud data to remove noise, improve the accuracy and stability of subsequent processing, and divide them into a training set and a test set;
[0050] S2: Construct an improved BEVDepth network model. As Figure 2 shown, introduce the radar-aided view transformation method RVT, and generate a BEV feature map with rich semantics and accurate space by fusing the complementary characteristics of the camera and radar sensors;
[0051] S2.1: Image feature encoding and depth distribution. Given a set of data containing N panoramic images, use an image backbone network combined with a feature pyramid network to extract a feature map F1 with 16 times downsampling for each perspective image. Subsequently, through an additional stacked convolutional layer, further extract the image context feature C1 in the perspective space according to the LSS method PV ∈R N ×C×H×W and the depth distribution D ∈ R of each pixel N×D×H×W :
[0052]
[0053] D(u, v) = Softmax(Conv(F1)(u, v))
[0054] where (u, v) represents the pixel coordinates in the image plane, and D is the number of depth intervals;
[0055] S2.2: As Figure 3As shown, radar feature encoding and radar occupancy grid generation: Different from previous methods that directly "elevate" image features to the bird's-eye view (BEV) by estimating depth distribution, the present invention uses radar measurements with noise but high precision for perspective transformation. First, radar points are projected onto N camera perspectives to find corresponding image pixels while retaining their depth information. Subsequently, through voxelization, the camera frustum view voxel V FV (d, u, v) is obtained. Here, u and v are pixel units in the width and height directions of the image respectively, and d is the metric unit in the depth direction. Since the radar cannot provide reliable height measurements, v = 1 is set to adopt a columnar structure. Feature encoding of non-empty radar columns is performed through PointNet and sparse convolution to obtain F2 ∈ R N×C×D×W . Radar context feature C2 FV ∈ R N×C×D×W and radar occupancy grid O ∈ R N ×1×D×W are extracted in the frustum view space. The convolution operation here is applied to the top view coordinates (d, u) rather than (u, v). Therefore, the radar context feature C2 FV and the radar occupancy grid O(d, u) in the frustum view space are as follows:
[0056]
[0057] O(d, u) = σ(Conv(F2)(d, u))
[0058] where σ represents the sigmoid activation function. Here, the sigmoid function is used instead of the softmax function because radar occupancy does not require one-hot encoding like depth distribution.
[0059] S2.3 Transform the image context feature map into the camera frustum view through frustum view transformation; Transform the camera and radar context feature maps finally located in N camera frustum views into a single BEV feature map through bird's-eye view transformation.
[0060] Frustum view transformation:
[0061] Given the depth distribution D1 and the radar occupancy rate O, the image context feature map C1 PV is transformed into the camera frustum view C1 FV ∈ R N×C×D×H×W , and its transformation method is as follows:
[0062]
[0063] where [·; ·] represents the concatenation operation along the channel dimension, represents the outer product operation;
[0064] Bird's-eye view transformation:
[0065] The camera and radar context feature maps F are finally located in the N camera frustum views. FV ={C1 FV ,C2 FV}∈R N ×C×D×H×W It is converted into a single BEV feature map R through the view transformation module M C×1×X×Y :
[0066]
[0067] in, is the feature map of the i-th cone view. Where M(·) represents the view conversion module (e.g. voxel pooling)
[0068] In the S2.4 bird's eye view (BEV) conversion, the details of the view transformation module M are improved CUDA voxel pooling, which is a core operation that maps feature maps from multiple perspectives to a unified BEV space. This method achieves feature fusion across perspectives by dividing the three-dimensional space into regular voxel grids and aggregating the features within each grid. In order to improve computational efficiency, a voxel pooling technique that supports CUDA acceleration is used to quickly project pixels or radar points in the cone view to the BEV grid using the parallel computing power of the GPU, and process features in groups. However, traditional sum pooling (SumPooling) will introduce distance bias due to the "near large and far small" characteristic of perspective projection - objects closer to the vehicle occupy more pixels in the image, causing their corresponding BEV grids to be associated with more cone grids. Direct summation will make the feature values near significantly higher than those far away, destroying the model's robustness to distance.
[0069] To this end, the improved solution uses average pooling instead of the summation operation, and normalizes the feature mean in each BEV grid to eliminate the imbalance in feature magnitude caused by the difference in the number of cone grids. This improvement not only Maintaining consistency at near and far distances reduces the model's sensitivity to distance, improves training stability, and is more in line with the physical distribution characteristics of sensor data (such as high resolution at near distances and sparseness at far distances), namely:
[0070]
[0071] in, is the number of frustum primes corresponding to the BEV grid, represents the set of voxels mapped to (x, y) in the i-th viewing cone, is the feature map of the i-th cone view, and each voxel (d, h, w) is projected to the position (x, y) of the BEV coordinate system through the sensor calibration parameters, that is: is the projection function of the i-th sensor.
[0072] By normalizing the sum result of the features, the dependence on K is eliminated, making x,y the dependence on K is eliminated, making only the feature mean is retained, which is independent of the distance. In actual implementation, a counting tensor K ∈ R needs to be maintained X×Y to record the number of voxels in each BEV grid and perform two steps: after sum pooling, divide each element by K (special handling for zero values). Combined with CUDA acceleration, average pooling further enhances the representation ability of BEV features in tasks such as object detection and road segmentation in the autonomous driving scenario on the basis of ensuring real-time performance, and is more robust especially in dealing with the problems of false detection and missed detection of near and far objects.
[0073] S3 introduces a cross-modal dynamic feature guidance mechanism DCMG in the multi-modal feature aggregation layer of the improved BEVDepth network model: through cross-modal dynamic gating, adaptive weight allocation is achieved, deformable spatial convolution aligns the image features and radar features of the BEV feature map, extracts channel-level global context descriptions, generates dynamic channel weights in combination with the Sigmoid function, and obtains the fused features of radar and camera through per-channel weighted fusion.
[0074] As Figure 4 shown, in the multi-modal feature aggregation layer of camera-radar fusion, the key challenge lies in how to effectively combine complementary multi-modal information while avoiding the inherent defects of each modality. Although image features contain rich semantic clues, there are inherent deviations in their spatial positions; while radar features have accurate spatial positioning, but have the shortcomings of insufficient context information and more noise. Traditional direct fusion methods such as channel-level concatenation or simple summation can neither solve the cross-modal spatial misalignment problem nor handle the uncertain associations between features. Therefore, a cross-modal dynamic feature guidance mechanism DCMG is introduced, which adopts an adaptive weight allocation mechanism for the credibility differences between camera and radar features in different scenarios. Input the image features and radar features aligned in the spatial dimension, extract channel-level global context descriptions, generate dynamic channel weights in combination with the Sigmoid function, and finally obtain the fused features of radar and camera through per-channel weighted fusion, enabling the model to dynamically adjust the modality contributions according to the environment. Dynamic cross-modal guidance aims to solve the semantic ambiguity and spatial misalignment problems caused by sensor characteristic differences in camera-radar multi-modal fusion, and achieves efficient fusion through dynamic feature selection and geometric alignment. Its core design includes three key technical points:
[0075] S3.1 Cross-modal dynamic gating For the credibility differences between camera and radar features in different scenarios, an adaptive weight allocation mechanism is adopted, and the input is the image features aligned in the spatial dimension and radar features After that, they are concatenated along the channel dimension into Extract the channel-level global context descriptor through global average pooling The middle layer dimension of the two-layer MLP is compressed to C / 4, the activation function is ReLU, and the dynamic channel weight W is generated by combining with the Sigmoid function gate ∈[0,1] C , and finally obtained through per-channel weighted fusion. The acquisition formula is as follows:
[0076] F fused =W gate ⊙F cam +(1 - W gate )⊙F radar
[0077] Among them, F cam is the image feature aligned in the spatial dimension, F radar is the radar feature, W gate is the dynamic channel weight, F fused is the fused feature, and ⊙ is matrix dot multiplication.
[0078] S3.2 Deformable spatial alignment is based on the high-precision spatial information of radar features. The deformable convolutional layer predicts 9 offsets and the modulation factor Δm n ∈[0,1], corresponding to the sampling points of the 3×3 convolutional kernel. Through bilinear interpolation resampling of the image features, the spatially aligned features are generated:
[0079]
[0080] Among them, The coordinate offset of the nth sampling point (predicted by radar features), Δm n ∈[0,1]: The modulation factor (weight) of the nth sampling point, p: The target position coordinates in the BEV space. It makes it exactly match the radar features in the BEV space, thereby eliminating the geometric deviation between cross-modal feature maps and overcoming sub-pixel level spatial misalignment.
[0081] S3.3 Cross-modal consistency loss constrains the semantic consistency between modalities through a contrastive learning framework. The image-radar features of the same target associated with the annotation box are used as positive sample pairs, and the different target features randomly sampled within the batch are used as negative sample pairs. Calculate the cosine similarity and construct the loss function. The formula is as follows:
[0082]
[0083] Among them, s(·): cosine similarity, τ: temperature coefficient (default 0.1), which controls the sharpness of the similarity distribution, B: batch size, i, j: the i / j-th sample in the batch. This loss function forces multi-modal features of the same target to closely cluster in the embedding space, while features of different targets are significantly separated, effectively alleviating the semantic drift phenomenon.
[0084] S4 Input the dataset to train the improved object detection model;
[0085] S5 As Figure 5 shown, use the trained network model to detect the image to be detected.
[0086] S6 This embodiment conducts experiments on the nuScenes dataset. This fusion scheme improves the mAP of dynamic object detection to 58.34%. By jointly optimizing the temporal consistency loss and the cross-modal contrast loss, the recall rate of the model in the scenario of temporary vehicle occlusion and reproduction is improved, fully verifying the effectiveness of the collaborative modeling of multi-modal spatio-temporal features. The results of this embodiment and the comparison results of the conventional model are shown in Table 1.
[0087] Table 1
[0088]
[0089]
[0090] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to the embodiments will be obvious to those skilled in the art. The general principles of the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention should not be limited to the embodiments shown herein, but should cover the widest range that conforms to the principles and novel features disclosed in the present invention.
Claims
1. A multi-point radar-camera fusion detection method based on an aerial view, characterized in that, The method includes the following steps: S1 Obtain the target image to be detected and radar data, and divide them into a training set and a test set; S2 Construct an improved BEVDepth network model, introduce the radar-aided view transformation method RVT, and generate a semantically rich and spatially accurate BEV feature map by fusing the complementary characteristics of camera and radar sensors; S3 Introduce a cross-modal dynamic feature guidance mechanism DCMG in the multi-modal feature aggregation layer of the improved BEVDepth network model: achieve weight adaptive allocation through cross-modal dynamic gating, use deformable spatial convolution to align the image features and radar features of the BEV feature map, extract channel-level global context descriptions, generate dynamic channel weights by combining with the Sigmoid function, and obtain the fused features of radar and camera through channel-wise weighted fusion; S4 Input the dataset to train the improved object detection model; S5 Use the trained network model to detect the image to be detected.
2. The multi-point radar camera fusion detection method based on an aerial view according to claim 1, wherein The specific steps of S2 are as follows: S2.1 Image Feature Encoding and Depth Distribution: Given a set of data containing N panoramic images, an image backbone network combined with a Feature Pyramid Network is used to extract a feature map F1 with 16x downsampling for each perspective image. Through an additional stacked convolutional layer, the image context feature C1 in the perspective space is further extracted according to the LSS method PV ∈R N×C×H×W and the depth distribution D of each pixel ∈ R N×D×H×W; S2.2 Introduce radar-assisted view transformation, radar feature encoding, and radar occupancy grid generation: Project radar points onto N camera views to find corresponding image pixels while retaining their depth information, and then obtain the camera frustum view voxel V FV (d, u, v), where u and v are pixel units in the width and height directions of the image, respectively, and d is the metric unit in the depth direction. Set v = 1 to adopt a columnar structure, and perform feature encoding on non-empty radar columns through PointNet and sparse convolution to obtain F2 ∈ R N×C×D×W , and extract the radar context feature C2 FV ∈ R N×C×D×W and the radar occupancy grid O ∈ R N×1×D×W , where the convolution operation is applied to the top view coordinates (d, u); S2.3 Transform the image context feature map into a camera frustum view through frustum view transformation; transform the camera and radar context feature maps finally located in N camera frustum views into a single BEV feature map through bird's-eye view transformation.
3. The multi-point radar camera fusion detection method based on the bird's-eye view according to claim 2, wherein, The specific processes of the frustum view transformation and the bird's-eye view transformation are as follows: Frustum view transformation: Given a depth distribution D1 and a radar occupancy O, an image context feature map C1 PV is converted to a camera frustum view C1 FV ∈R N ×C×D×H×W , and the conversion is as follows: where [·;·] represents the concatenation operation along the channel dimension, represents the outer product operation; Bird's-eye view transformation: The camera and radar context feature map F finally located in the N camera frustum views FV ={C1 FV , C2 FV} ∈ R N×C×D×H×W is transformed into a single BEV feature map R by the view transformation module M C×1×X×Y : Among them, is the feature map of the i-th cone view, where M(·) is the view conversion module.
4. The method for multi-point radar camera fusion detection based on an aerial view according to claim 3, wherein The specific content of the view transformation module M is as follows: adopt a voxel pooling method that supports CUDA acceleration, and use average pooling to replace the summation operation to aggregate the features within each BEV grid, and eliminate the feature magnitude imbalance caused by the difference in the number of frustum grids by normalizing the feature mean within each BEV grid.
5. The multi-point radar camera fusion detection method based on an aerial view according to claim 1, wherein The specific steps of S3 are as follows: Based on the spatially aligned features of the BEV feature map, cross-modal dynamic gating adaptively allocates modal weights according to the scene context: extract cross-modal channel-level statistics through global average pooling, and use a lightweight MLP to generate dynamic gating weights to fuse image and radar features in a channel-wise weighted form. The cross-modal consistency loss constrains the semantic alignment of the above process through contrastive learning during the training stage, forcing the image and radar feature pairs of the same target to be closely aggregated in the embedding space, while the feature pairs of different targets are significantly separated. The semantic-level constraint acts on the intermediate features after spatial alignment. If the fused features cause semantic drift due to environmental interference, the contrastive loss will adjust the spatial alignment offset prediction and channel weight allocation strategy through gradient backpropagation.
6. The multi-point radar camera fusion detection method based on an aerial view according to claim 5, characterized in that, The specific operation of the cross-modal dynamic gating is as follows: Cross-modal dynamic gating addresses the credibility differences between camera and radar features in different scenarios, adopts an adaptive weight allocation mechanism, and inputs image features aligned in the spatial dimension and radar features After that, they are concatenated along the channel dimension into Extract the channel-level global context descriptor through global average pooling The dimension of the middle layer of the two-layer MLP is compressed to C / 4, the activation function is ReLU, and the dynamic channel weight W is generated in combination with the Sigmoid function gate ∈[0,1] C Finally, it is obtained through per-channel weighted fusion, and the acquisition formula is as follows: F fused = W gate ⊙ F cam +(1 - W gate ) ⊙ F radar Among them, F cam is the image feature aligned in the spatial dimension, F radar is the radar feature, W gate is the dynamic channel weight, F fused is the fusion feature, and ⊙ is the matrix dot product.
7. The multi-point radar camera fusion detection method based on an aerial view according to claim 5, wherein, The specific process of the spatial alignment is as follows: Deformable spatial alignment is based on high-precision spatial information of radar features. The deformable convolutional layer predicts 9 offsets for each image feature position and a modulation factor Δm n ∈[0,1], corresponding to the sampling points of the 3×3 convolutional kernel. By bilinearly interpolating and resampling the image features, spatially aligned features are generated: Among them, The coordinate offset of the nth sampling point, predicted by radar features; Δm n ∈[0,1]: The modulation factor weight of the nth sampling point; p: The target position coordinates in the BEV space, which makes it exactly match the radar features in the BEV space, so as to eliminate the geometric deviation between cross-modal feature maps and overcome the sub-pixel level spatial misalignment.
8. The method for multi-point radar camera fusion detection based on an aerial view according to claim 5, characterized in that The cross-modal consistency loss constrains the semantic consistency between modalities through a contrastive learning framework. Use the image-radar features of the same target associated with the annotation box as positive sample pairs, and randomly sampled different target features within the batch as negative sample pairs, calculate the cosine similarity and construct a loss function, and the formula is as follows: where s(·): cosine similarity, τ: temperature coefficient, which controls the sharpness of the similarity distribution, B: batch size, i, j: the i / j-th sample in the batch.
Citation Information
Cited By
Multi-modal fusion sensing method, device and system for automatic driving
CN120635850A
Multi-mode BEV perception feature dynamic fusion method based on pixel-level confidence gating
CN121708436A
A pixel-level confidence-gated dynamic fusion method for multimodal BEV perception features
CN121708436B