Radar / camera fusion 3D target detection method based on bidirectional cross-modal attention

By adopting a radar/camera fusion method with bidirectional cross-modal attention in the autonomous driving perception system, the camera image and lidar point cloud data are effectively fused, and the problem of insufficient target detection accuracy and robustness is solved, and 3D object detection with high accuracy and high robustness is achieved.

CN119992065APending Publication Date: 2025-05-13SOUTHEAST UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510165773.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate camera images and lidar point cloud data in an autonomous driving perception system, resulting in reduced target detection accuracy and insufficient robustness.

Method used

The radar/camera fusion 3D object detection method based on bidirectional cross-modal attention is adopted, and the image and point cloud data are compressed into the BEV space through feature extraction and enhancement modules, and adaptive weighted fusion is used for use with the bidirectional cross-modal attention mechanism, and finally input the 3D object detection task head for detection.

Benefits of technology

It realizes the provision of accurate and reliable detection results in 3D object detection tasks, improves the robustness and detection accuracy of the system, and adapts to the target detection needs in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992065A_ABST
    Figure CN119992065A_ABST
Patent Text Reader

Abstract

The invention discloses a radar / camera fusion 3D target detection method based on bidirectional cross-modal attention, and the method comprises the steps: inputting original image data into a feature extractor, carrying out the enhancement of image features, and compressing the image features into a BEV space; voxelizing original point cloud data, inputting the original point cloud data into a feature extractor, compressing the original point cloud data into a BEV space, and performing feature enhancement; performing adaptive weighted fusion on the point cloud features and the image features in the BEV space by using a bidirectional cross-modal attention mechanism; and finally, inputting the fused features into a 3D target detection task head to obtain a 3D target detection result. According to the method, the attention mechanism is applied to a self-adaptive fusion framework, self-adaptive weighted fusion is carried out on point cloud data and image data of different modalities, an accurate and reliable detection result can be provided in a 3D target detection task, and good adaptability is achieved; the system also has double advantages of detection precision and robustness, and meets actual requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-sensor information fusion, and in particular to the fusion of camera images and lidar point cloud data, which is commonly used in autonomous driving technology. It mainly involves a radar / camera fusion 3D target detection method based on bidirectional cross-modal attention. Background Art

[0002] Autonomous driving technology includes three parts: perception, planning, and control. The core task of perception is to perceive the environment around the vehicle and provide necessary information for decision-making and control. The perception system includes a variety of visual tasks, among which 3D target detection is used to obtain the size and position of targets in the real world. Cameras and lidars are two important sensors in the field of perception. The images collected by camera sensors contain rich texture and color information, which can effectively identify the geometric shape of objects, but they lack depth information and cannot reflect the distance between the object and the camera. They are also easily affected by external conditions such as weather changes and light brightness. The point cloud data collected by lidar can accurately reflect the distance between the object and the sensor, but it is less sensitive to external factors such as light changes. In general, it is difficult for a single sensor to fully obtain environmental information.

[0003] With the advancement of technology, more and more sensor fusion solutions are being applied to autonomous driving perception systems. By fusing multimodal sensor data, the lack of information from a single sensor can be compensated, thereby achieving more comprehensive and accurate target detection and positioning. The best fusion method should make full use of point cloud information and image information. However, due to the differences in the targets detected in different scenarios, simple splicing may greatly reduce the accuracy of detection, leading to problems such as missed and false detection of targets. Therefore, for different scenarios, it is necessary to perform weighted fusion of point clouds and images to give full play to their respective advantages and avoid wasting effective information.

[0004] The introduction of the attention mechanism can realize the adaptive weighting of image and point cloud features and improve the fusion effect. When facing long-distance information interaction, convolutional neural networks require more network layers or use fully connected networks, which will consume more computing resources. The attention mechanism can dynamically generate weights of different connections, effectively saving computing resources and focusing more on processing more important information. Although the attention mechanism can enhance the ability of feature extraction and improve the robustness of the network structure, it is still a challenge to reasonably allocate the weights of different modalities when fusing cross-modal features. In addition, features that are not obvious during the fusion process may also affect the performance of the entire network structure. Therefore, it is of great practical significance to study an adaptive fusion method so that the system has stronger robustness while maintaining high detection accuracy. Summary of the invention

[0005] The present invention is aimed at the problems existing in the prior art, and proposes a radar / camera fusion 3D target detection method based on bidirectional cross-modal attention. First, the original image data is input into the feature extractor, the image features are enhanced, and the enhanced features are compressed into the BEV space; then the original point cloud data is voxelized, input into the feature extractor, and compressed into the BEV space for feature enhancement; then the bidirectional cross-modal attention mechanism is used to adaptively weighted fuse the point cloud features and image features in the BEV space; finally, the fused features are input into the 3D target detection task head to obtain the 3D target detection results. The method of the present invention applies the attention mechanism to the adaptive fusion framework, combines point cloud features and image features, realizes the fusion of the bidirectional cross-modal attention mechanism, and adaptively weighted fuses point cloud data and image data of different modalities, which can provide accurate and reliable detection results in 3D target detection tasks and has good adaptability; the system of the present invention also has the dual advantages of detection accuracy and robustness, which meets actual needs.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is: a 3D target detection system based on radar / camera fusion with bidirectional cross-modal attention, which at least includes a feature extraction and enhancement module, a fusion module and a detection module.

[0007] The feature extraction and enhancement module is used to receive the original image data and the laser radar point cloud data, extract features from the two types of data respectively, perform feature enhancement, and compress them into the BEV space;

[0008] The fusion module: at least includes a cross-modal cross-attention fusion layer, which is used to realize the attention interaction and fusion of point cloud data and image data in the BEV space; the feature fusion process of the two types of data includes at least two times;

[0009] The detection module at least includes a 3D target detection task head, which is used to receive the fusion features output by the fusion module and output the 3D target detection results.

[0010] In order to achieve the above object, the present invention also adopts a technical solution: a radar / camera fusion 3D target detection method based on bidirectional cross-modal attention, which at least includes the following steps:

[0011] S1, image feature extraction and enhancement: the original image data is input into the feature extractor, the image features are enhanced, and the enhanced features are compressed into the BEV space;

[0012] S2, radar feature extraction and enhancement: the original point cloud data is voxelized, input into the feature extractor, compressed into the BEV space, and feature enhancement is performed;

[0013] S3, data fusion: using the bidirectional cross-modal attention mechanism to adaptively weight the point cloud features in the BEV space enhanced in step S2 and the image features enhanced in step S1;

[0014] S4, target detection: The features fused in step S3 are sent to the 3D target detection head to complete the target detection task.

[0015] As an improvement of the present invention, the step S1 of image feature extraction and enhancement includes the following steps:

[0016] S11. Image feature extraction: Through the ResNet50 residual network and feature pyramid network, the multi-scale semantic information features of the original input image are extracted to obtain the deep feature information F FPN , the depth feature information F FPN Specifically:

[0017] F FPN =FPN[Res(I)]

[0018] Where Res represents the ResNet50 residual network, and FPN represents the feature pyramid network as a multi-scale semantic information extractor;

[0019] S12, multi-scale semantic feature enhancement: The multi-scale semantic information obtained in step S11 is enhanced by an improved effective multi-scale attention method to obtain the enhanced image feature F Image :

[0020] F Image =EMA(F FPN );

[0021] S13, feature compression to BEV space: associate the position coordinates of the enhanced image feature discrete point cloud obtained in step S12 in 3D space to the two-dimensional image plane, use Softmax to predict the discrete depth distribution α of each pixel, scatter each feature pixel along the camera light to discrete points, rescale the features by depth probability, and generate the depth feature F d ; Use BEV pooling to aggregate all features in the BEV space and flatten the features along the z-axis to finally obtain the image feature F in the BEV space ImBEV :

[0022] F ImBEV =BEVpooling[F d ].

[0023] As another improvement of the present invention, in step S13, the position coordinates of the discrete point cloud in the 3D space are: Associate the coordinates of d discrete point clouds from the 3D space to the pixels (u, v) on the 2D image plane, specifically:

[0024]

[0025] Where P2 is the camera projection matrix, R0 is the distortion correction matrix, and R L2C The rotation matrix from the coordinate system to the camera coordinate system.

[0026] As another improvement of the present invention, the step S2 of radar feature extraction and enhancement specifically includes the following steps:

[0027] S21. Voxelization of point cloud data: The point cloud data is converted into a regular 3D voxel grid to obtain voxel features, and then the voxel features are encoded using the average voxel feature encoding method to obtain the encoded voxel feature F Voxel :

[0028]

[0029] In the formula, x i ,y i ,z i Represents the position coordinates of the i-th point; r i Indicates the intensity return value of the i-th point; n indicates the number of valid point clouds contained in the voxel;

[0030] S22, extract voxel features and compress them along the z-axis to BEV space: extract the voxel features F obtained in step S21 Voxel Perform convolution operation, extract point cloud features by 8 times downsampling, and then compress along the z-axis to obtain the lidar feature F under BEV representation Li ;

[0031] S23, feature enhancement: The lidar feature F under the BEV representation obtained in step S22 is enhanced through the SE channel attention mechanism Li Perform feature enhancement, specifically:

[0032] F LiBEV =F Li [MLP(GAP(F Li ))]

[0033] Where GAP stands for global average pooling, and MLP includes two fully connected layers, ReLU and sigmoid functions.

[0034] As another improvement of the present invention, in step S21, the point cloud data Converted into a regular 3D voxel grid, the voxel features obtained are:

[0035]

[0036] Where V j represents the feature of the jth voxel; t represents the total number of voxels.

[0037] As a further improvement of the present invention, in the data fusion of step S3, the image feature F in the BEV space is ImBEV and the lidar feature F LiBEV Perform alignment operation to obtain image feature F ImF and the lidar feature F LiF, The two features are fused twice through the cross-modal cross-attention fusion layer; after the first fusion, feature F is obtained. FF :

[0038] F FF =Conv(Concat(Remap(BCAFL(Q ImF ,K LiF ,V LiF ))+F ImF ));

[0039] The feature F FF Through a full convolution layer, we get the corresponding K FF and V FF , Q value is determined by F LiF Generate, get Q LiF ; Then perform the second fusion, and obtain the feature F after the second fusion SF :

[0040] F SF =Conv(Concat(Remap(BCAFL(Q LiF ,K FF ,V FF ))+F LiF ))

[0041] Compared with the prior art, the present invention has the following beneficial effects: the present invention discloses a radar / camera fusion 3D target detection method based on bidirectional cross-modal attention, introduces image feature and point cloud feature enhancement in the BEV-based autonomous driving perception algorithm, and can effectively prevent information loss after compression into the BEV space; the bidirectional cross-modal attention mechanism can effectively fuse information of two different modalities, image and point cloud, to achieve adaptive weighted fusion, and ensure the accuracy and reliability of the 3D target detection task results. The rational application of the two improves the robustness and detection accuracy of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flowchart of the steps of the radar / camera fusion 3D target detection method based on bidirectional cross-modal attention of the present invention;

[0043] Figure 2 Schematic diagram of the structure of the cross-modal cross-attention fusion layer of the present invention;

[0044] Figure 3 This is the KITTI test set result diagram in the test example of the present invention;

[0045] Figure 4 This is a comparison chart of the results of different algorithms in the test example of the present invention on the KITTI validation set. DETAILED DESCRIPTION

[0046] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0047] Example 1

[0048] A 3D target detection system based on radar / camera fusion with bidirectional cross-modal attention includes at least a feature extraction and enhancement module, a fusion module and a detection module.

[0049] The feature extraction and enhancement module is used to receive the original image data and lidar point cloud data, extract features from the two types of data respectively, perform feature enhancement, and both are compressed into the BEV space; the fusion module includes at least a cross-modal cross-attention fusion layer, which is used to realize the attention interaction and fusion of point cloud data and image data in the BEV space; the detection module includes at least a 3D target detection task head, which is used to receive the fusion features output by the fusion module and output the 3D target detection results.

[0050] Using the system in this embodiment, a radar / camera fusion 3D target detection method based on bidirectional cross-modal attention uses the original input image and point cloud data to perform feature extraction and enhancement respectively, and uses the bidirectional cross-modal attention mechanism to achieve adaptive weighted fusion, and finally sends it to the detection head to complete the target detection task. Figure 1 As shown, the specific steps include:

[0051] Step S1: Input the original image data into the feature extractor, then enhance the image features and compress the enhanced features into the BEV space.

[0052] S11: Image feature extraction.

[0053] Extract multi-scale semantic information from the original input image, use ResNet50 as the backbone network, and use FPN as the multi-scale semantic information extractor to generate deep feature information, thereby effectively solving the problem of small target detection and low accuracy. After passing through the backbone network and the multi-scale semantic information extractor, the final deep feature information FFPN for:

[0054] F FPN =FPN[Res(I)]

[0055] In the formula, Res represents the ResNet50 network, and FPN represents the multi-scale semantic information extractor.

[0056] S12: Multi-scale semantic feature enhancement.

[0057] Since point cloud features do not have semantic information and deep image features are easily affected by background noise, in order to improve the expressiveness of deep image features, an improved EMA method is used to suppress background noise and enhance the expressiveness of image features:

[0058] F Image =EMA(F FPN )

[0059] S13: Feature compression to BEV space.

[0060] The position coordinates of the discrete point cloud in 3D space are Associate d discrete point clouds from 3D space coordinates to 2D image plane pixels (u, v):

[0061]

[0062] Where P2 is the camera projection matrix, R0 is the distortion correction matrix, and R L2C The rotation matrix from the coordinate system to the camera coordinate system.

[0063] Use Softmax to predict the discrete depth distribution α of each pixel, scatter each feature pixel into d discrete points along the camera light, and rescale the relevant features by the corresponding depth probability, thus generating a depth feature F of size H×W×d d Next, BEV pooling is used to aggregate all features in the BEV space and flatten the features along the z-axis to finally obtain the image feature F in the BEV space. ImBEV :

[0064] F ImBEV =BEVpooling[F d ]

[0065] Step S2: The original point cloud data is voxelized, input into the feature extractor, compressed into the BEV space, and feature enhanced.

[0066] S21, voxelization of point cloud data;

[0067] Since point cloud data It is disordered. In order to extract point cloud features, the point cloud data must first be converted into a regular 3D voxel grid:

[0068] D′=D / P D ,H′=H / P H ,W′=W / P W

[0069] Where D, H, and W are the depth, height, and width of the point cloud space, respectively. D ,P H ,P W are the depth, height and width of each voxel, so a total of D′×H′×W′ voxels can be divided. Thus, the characteristics of each voxel can be obtained:

[0070]

[0071] Where V j represents the feature of the jth voxel; x i ,y i ,z i Represents the position coordinates of the i-th point; r i Represents the intensity return value of the i-th point; n represents the number of valid point clouds contained in the voxel; t represents the total number of voxels.

[0072] Then the voxel features are encoded, and the average voxel feature encoding method is adopted, so the encoded feature F is obtained Voxel :

[0073]

[0074] S22, voxel features are extracted and compressed along the z-axis to BEV space;

[0075] In order to further extract the features of sparse point clouds, Voxel Perform convolution operation, extract point cloud features by 8 times downsampling, and then compress along the z-axis to obtain the lidar feature F under BEV representation LiBEV .

[0076] S23, feature enhancement;

[0077] The point cloud features in the BEV space have rich spatial distance and geometric structure information, and are obtained by compressing and transforming the point cloud features along the z-axis, so it is necessary to filter the spatial information and remove redundant spatial information. The SE channel attention mechanism can filter and aggregate more representative BEV spatial features.

[0078] First, the input lidar feature F LiBEVPerform global average pooling to transform the two-dimensional feature channel into a 1×1×c feature vector with a global receptive field, which represents the global response distribution on the feature channel. Use an MLP to capture the correlation of spatial information between feature channels, thereby obtaining the weight of each channel. Finally, let the input feature F LiBEV Perform channel-by-channel multiplication weighting to complete F in BEV space LiBEV The spatial information of the feature is enhanced. This process can be expressed as follows:

[0079] F LiBEV =F LiBEV [MLP(GAP(F LiBEV ))]

[0080] Where GAP stands for global average pooling, and MLP includes two fully connected layers, ReLU and sigmoid functions.

[0081] Step S3: Use the bidirectional cross-modal attention mechanism to adaptively weight the point cloud features and image features in the BEV space.

[0082] In order to achieve adaptive weighted fusion of LiDAR and camera, a bidirectional cross attention fusion layer is constructed to realize the attention interaction and fusion of point cloud data and image data in BEV space. The structure of BCAFL is as follows: Figure 2 As shown, a BCAFL schematic diagram of radar-image is provided to demonstrate the cross-attention fusion process between point cloud data and image data features in BEV space.

[0083] In the BCAFL layer, the point cloud information and image information in the BEV space are fused. From the input image features, we can extract From the input radar features, we can extract and Where M and N represent the number of points of image features and the number of points of point cloud features respectively, and D1 and D2 represent the number of channels of image features and the number of channels of point cloud features in BEV space respectively. Usually, the number of channels of point cloud features and the number of channels of image features in BEV space are the same, which is represented by D here:

[0084]

[0085] In the formula, Represents the pairwise similarity between point cloud features and image feature elements in BEV space.

[0086] In different environments, the perception capabilities of LiDAR and camera are different, so they cannot be simply spliced ​​together to complete the fusion operation. The features of LiDAR and camera should be weighted fused in a weighted manner. The fusion process adopts the proposed BCAFL. In order to improve the robustness of the model for small target detection, a convolution operation is required before fusion, which also saves computing resources for subsequent fusion.

[0087] Because F ImBEV and F LiBEV The number of feature channels of the two features is different, so we first need to align the two features before fusion, that is, use MLP to align the number of feature channels after expansion to obtain F ImF and F LiF Then use BCAFL twice to complete the fusion operation of information features. The feature F obtained after the first fusion FF for:

[0088] F FF =Conv(Concat(Remap(BCAFL(Q ImF ,K LiF ,V LiF ))+F ImF ))

[0089] First, use F ImF and F LiF Generate Q ImF , K LiF and V LiF , passing them through the BCAFL layer, we can get the first fused feature F FF The trainable attention score matrix can give greater weight to point cloud information with more information, so as to filter out some weakly correlated features in the initial fusion process and ensure that the output features are more obvious. In order to avoid the loss of original features in the subsequent fusion process, F LiF The features are concatenated with the Remapped features to ensure that the information is not distorted. Subsequently, a convolutional layer, a normalization layer, and a ReLu function are used to normalize the features.

[0090] In the first feature fusion process, the obtained F FF It will be flattened again and then pass through a full convolution layer to get the corresponding K FF and V FF , and the corresponding Q value is given by F LiF Generate, and then get Q LiF Therefore, after the second fusion, feature F can be obtained SF :

[0091] F SF=Conv(Concat(Remap(BCAFL(Q LiF ,K FF ,V FF ))+F LiF ))

[0092] Step S4: Target detection: The features fused in step S3 are sent to the 3D target detection head to complete the target detection task.

[0093] Test Case

[0094] This test case verifies the effectiveness of the method proposed in the present invention through simulation experiments. The experimental parameters are set as follows: CPU is Intel(R) Xeon(R) Platinum 8474C, GPU is RTX 4090D, Ubuntu version is 22.04, Python version is 3.10, PyTorch version is 2.1.0, and Cuda version is 12.1. The dataset used is KITTI, the batch size is set to 4, and a total of 80 rounds are trained. The learning rate optimizer uses the Adam (Adaptive Moment Estimation) optimizer, and the initial learning rate is set to 0.002. The experimental results are tested on the official website of the KITTI dataset, and the corresponding 3D target detection results and BEV detection results are obtained.

[0095] like Figure 3 As shown in FIG. 1 , the performance of the present invention on the KITTI validation set is given. For the automobile category, the detection accuracy of the present invention in simple, medium and difficult situations is 89.21%, 79.27% ​​and 78.62% respectively, indicating that the present invention has a high detection accuracy.

[0096] like Figure 4 As shown, the results of several algorithms on the KITTI test set are compared. For the car category, the detection accuracy of the present invention in simple, medium and difficult situations is 86.45%, 78.22% and 73.63%, respectively. Compared with the performance on the validation set, it is reduced by about 1% to 5%, indicating that the present invention has good robustness. Compared with the VoxelNet and PointPillars methods that only use radar, the detection accuracy of the present invention is about 15% higher than that of the method that only uses radar. Compared with the simple fusion methods MVXNet and AVOD, the adaptive fusion method mentioned in the present invention is also improved by about 5%, which illustrates the effectiveness of the present invention.

[0097] In summary, the present invention applies the attention mechanism to an adaptive fusion framework, combines point cloud features and image features, and realizes the fusion of a bidirectional cross-modal attention mechanism. The method of the present invention can perform adaptive weighted fusion on point cloud data and image data of different modalities, so that the system has both detection accuracy and robustness.

[0098] It should be noted that the above content only illustrates the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.

Claims

1. A 3D object detection system based on radar / camera fusion with bidirectional cross-modal attention, characterized by: At least including feature extraction and enhancement module, fusion module and detection module, The feature extraction and enhancement module is used to receive the original image data and the laser radar point cloud data, extract features from the two types of data respectively, perform feature enhancement, and compress them into the BEV space; The fusion module: at least includes a cross-modal cross-attention fusion layer, which is used to realize the attention interaction and fusion of point cloud data and image data in the BEV space; the feature fusion process of the two types of data includes at least two times; The detection module at least includes a 3D target detection task head, which is used to receive the fusion features output by the fusion module and output the 3D target detection results.

2. A radar / camera fusion 3D target detection method based on bidirectional cross-modal attention using the system as claimed in claim 1, characterized in that: At least the following steps are included: S1, image feature extraction and enhancement: the original image data is input into the feature extractor, the image features are enhanced, and the enhanced features are compressed into the BEV space; S2, radar feature extraction and enhancement: the original point cloud data is voxelized, input into the feature extractor, compressed into the BEV space, and feature enhancement is performed; S3, data fusion: using the bidirectional cross-modal attention mechanism to adaptively weight the point cloud features in the BEV space enhanced in step S2 and the image features enhanced in step S1; S4, target detection: The features fused in step S3 are sent to the 3D target detection head to complete the target detection task.

3. The radar / camera fusion 3D target detection method based on bidirectional cross-modal attention as claimed in claim 2, characterized in that: The step S1 of image feature extraction and enhancement comprises the following steps: S11. Image feature extraction: Through the ResNet50 residual network and feature pyramid network, the multi-scale semantic information features of the original input image are extracted to obtain the deep feature information F FPN , the depth feature information F FPN Specifically: F FPN =FPN[Res(I)] Where Res represents the ResNet50 residual network, and FPN represents the feature pyramid network as a multi-scale semantic information extractor; S12, multi-scale semantic feature enhancement: The multi-scale semantic information obtained in step S11 is enhanced by an improved effective multi-scale attention method to obtain the enhanced image feature F Image : F Image =EMA(F FPN ); S13, feature compression to BEV space: associate the position coordinates of the enhanced image feature discrete point cloud obtained in step S12 in 3D space to the two-dimensional image plane, use Softmax to predict the discrete depth distribution α of each pixel, scatter each feature pixel along the camera light to discrete points, rescale the features by depth probability, and generate the depth feature F d ; Use BEV pooling to aggregate all features in the BEV space and flatten the features along the z-axis to finally obtain the image feature F in the BEV space ImBEV : F ImBEV =BEVpooling[F d ]。 4. The radar / camera fusion 3D target detection method based on bidirectional cross-modal attention as claimed in claim 3, characterized in that: In step S13, the position coordinates of the discrete point cloud in the 3D space are [X Lidar ,Y Lidar ,Z Lidar ] T , associate the coordinates of d discrete point clouds from the 3D space to the pixels (u, v) in the 2D image plane, specifically: Where P2 is the camera projection matrix, R0 is the distortion correction matrix, and R L2C The rotation matrix from the coordinate system to the camera coordinate system.

5. The radar / camera fusion 3D target detection method based on bidirectional cross-modal attention as claimed in claim 2, characterized in that: The step S2 of radar feature extraction and enhancement specifically includes the following steps: S21. Voxelization of point cloud data: The point cloud data is converted into a regular 3D voxel grid to obtain voxel features, and then the voxel features are encoded using the average voxel feature encoding method to obtain the encoded voxel feature F Voxel : In the formula, x i ,y i ,z i Represents the position coordinates of the i-th point; r i Indicates the intensity return value of the i-th point; n indicates the number of valid point clouds contained in the voxel; S22, extract voxel features and compress them along the z-axis to BEV space: extract the voxel features F obtained in step S21 Voxel Perform convolution operation, extract point cloud features by 8 times downsampling, and then compress along the z-axis to obtain the lidar feature F under BEV representation Li ; S23, feature enhancement: The lidar feature F under the BEV representation obtained in step S22 is enhanced through the SE channel attention mechanism Li Perform feature enhancement, specifically: F LiBEV =F Li [MLP(GAP(F Li ))] Where GAP stands for global average pooling, and MLP includes two fully connected layers, ReLU and sigmoid functions.

6. The radar / camera fusion 3D target detection method based on bidirectional cross-modal attention as claimed in claim 5, characterized in that: In step S21, the point cloud data Converted into a regular 3D voxel grid, the voxel features obtained are: Where V j represents the feature of the jth voxel; t represents the total number of voxels.

7. The radar / camera fusion 3D target detection method based on bidirectional cross-modal attention according to claim 4 or 6, characterized in that: In the step S3 data fusion, the image feature F in the BEV space is ImBEV and the lidar feature F LiBEV Perform alignment operation to obtain image feature F ImF and the lidar feature F LiF, The two features are fused twice through the cross-modal cross-attention fusion layer; after the first fusion, feature F is obtained. FF : F FF =Conv(Concat(Remap(BCAFL(Q ImF ,K LiF ,V LiF ))+F ImF )); The feature F FF Through a full convolution layer, we get the corresponding K FF and V FF , Q value is determined by F LiF Generate, get Q LiF ; Then perform the second fusion, and obtain the feature F after the second fusion SF : F SF =Conv(Concat(Remap(BCAFL(Q LiF ,K FF ,V FF ))+F LiF ))。

Citation Information

Cited By

  • Semantic occupancy prediction method and system based on bidirectional multi-modal residual fusion

    CN121304982A

  • Polarization-event-radar multi-mode cooperative sensing method based on Mama

    CN121883822A

  • Mamba-based polarization-event-radar multi-modal collaborative perception method

    CN121883822B

  • Intelligent traffic target detection system and method based on 4D millimeter wave and laser radar point cloud fusion

    CN122218707A