A multi-modal fusion target detection method and system based on images, laser radars and millimeter wave radars
By constructing enhanced radar point cloud features and evidence-guided residual fusion, the problems of insufficient utilization of millimeter-wave radar information and insufficient velocity correction are solved, improving the stability and applicability of multimodal target detection, especially the detection performance under low light and severe weather conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-28
AI Technical Summary
Existing technologies suffer from insufficient utilization of millimeter-wave radar information, lack of physical correction for velocity, and insufficient overall stability in the fusion of weak radar semantic interference, resulting in poor detection performance, especially in low-light environments, adverse weather conditions, and complex traffic target motion scenarios.
By constructing enhanced radar point cloud features, radar semantic BEV features and radar velocity BEV feature maps are generated. An evidence-guided residual fusion method is used to independently encode the millimeter-wave radar velocity information and fuse it with image and lidar features. Multi-frame time difference and azimuth angle are used to enhance the temporal and directional information of historical radar points for target detection and velocity correction.
It improves the ability to detect the speed of moving targets, enhances the detection stability in low-light environments and adverse weather conditions, effectively utilizes the characteristics of millimeter-wave radar, avoids the impact of weak radar characteristics on the backbone network, and is suitable for complex traffic scenarios.
Smart Images

Figure CN122469339A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of target detection technology, specifically relating to a multimodal fusion target detection method and system based on image, lidar and millimeter-wave radar. Background Technology
[0002] Existing autonomous driving target detection technologies mainly include single-modal or multi-modal fusion methods such as pure image processing, pure LiDAR, and image and LiDAR. Image-based methods are relatively low-cost and provide rich image information, but their performance and accuracy are significantly affected in scenarios such as long target distances, target occlusion, low light conditions, and adverse weather conditions due to inaccurate depth prediction. LiDAR-based methods can model the surrounding environment (modeling accuracy depends on LiDAR accuracy), but point cloud density decreases and noise increases in rainy / foggy weather, occluded scenarios, and low-reflection target scenarios. Furthermore, LiDAR itself cannot directly provide measurement information on the radial velocity of the target. Millimeter-wave radar has advantages such as strong resistance to adverse weather conditions, speed perception, and rich Doppler information. Therefore, more and more studies are trying to incorporate millimeter-wave radar into target detection methods. However, millimeter-wave radar point clouds are fewer in number, sparser in spatial distribution, and have more noise and ghosting compared to lidar point clouds. This makes it easy for radar features to be submerged by strong features after being concatenated with image or lidar features when directly added to multimodal network fusion. Other problems include: multi-frame radar data being simply stacked or concatenated, historical frame displacement information not being well utilized; radar semantic features and radar physical velocity information being mixed in the same encoding path, resulting in Doppler velocity being weakened within the modality, making it difficult for the network to take full advantage of the complete radar velocity information; and most fusion layers using direct concatenation or uniform convolutional fusion, where the introduction of weak radar branches can destroy the already trained strong image and strong lidar features.
[0003] While existing technologies, such as image and lidar fusion detection under unified BEV representation, radar and image fusion detection in BEV space, image BEV detection based on depth estimation, and general multimodal feature fusion techniques, each have their own characteristics, they still suffer from the following technical problems:
[0004] (1) Millimeter-wave radar information is not fully utilized.
[0005] (2) The speed lacks physical correction.
[0006] (3) Overall stability of weak radar semantic interference fusion.
[0007] (4) Not suitable for scenarios such as low light environment, bad weather and complex traffic target movement.
[0008] Therefore, there is an urgent need for a new multimodal 3D target detection method that enables the radar branch to separate the velocity branch, providing semantic features for the overall network while preserving and integrating the radar's Doppler velocity features into the network. Summary of the Invention
[0009] To address the aforementioned technical problems in the prior art, the purpose of this invention is to resolve the issues of insufficient utilization of millimeter-wave radar information, lack of physical velocity correction, and overall stability in the fusion of weak radar semantic interference in the prior art. The technical solution is as follows:
[0010] A multimodal fusion target detection method based on image, lidar, and millimeter-wave radar includes the following steps:
[0011] Step 1: Acquire image data, LiDAR point cloud data, and millimeter-wave radar point cloud data;
[0012] Step 2: Construct enhanced radar point cloud features and input the semantic radar coding path to perform radar point cloud voxelization and generate radar semantic BEV features.
[0013] Step 3: Input the millimeter-wave radar point cloud data into the millimeter-wave radar velocity encoding path, encode the features of radar point-level velocity, and generate a radar velocity BEV feature map.
[0014] Step 4: Perform feature extraction and depth view transformation on the image data to generate image BEV spatial features; perform feature extraction on the lidar point cloud data to generate lidar BEV spatial features;
[0015] Step 5: Perform residual fusion of image BEV spatial features, lidar BEV spatial features, and radar semantic BEV map features through evidence-guided analysis to generate fused BEV spatial features.
[0016] Step 6: Perform target detection and velocity detection on the fused BEV spatial features;
[0017] Step 7: Perform confidence correction on the radar velocity based on the target query location to obtain the final target velocity result.
[0018] Furthermore, in step 1, the image data is multi-view image data with 6 surround views; the lidar point cloud data is the current frame and 9 historical lidar point cloud data; and the millimeter-wave radar point cloud data is the current frame and 5 historical millimeter-wave radar point cloud data.
[0019] Furthermore, the lidar point cloud serves as an explicit depth indicator to supervise depth prediction for image branches.
[0020] Furthermore, the specific process of step 2 is as follows:
[0021] Step 2.1: Calculate the planar modulus of each radar point cloud data point. The calculation expression is as follows:
[0022]
[0023] in, Indicates the radar point is at The distance from the origin of the reference coordinate system within the plane; and These represent the points at... shaft and Coordinates along the axis;
[0024] Step 2.2: Construct the azimuth feature, whose expression is:
[0025]
[0026]
[0027] in, This represents the azimuth angle of the radar point relative to the origin of the reference coordinate system. Represents the sine value of the azimuth angle; Represents the cosine value of the azimuth angle;
[0028] Step 2.3: When When it is less than the preset threshold, and Setting it to 0 constructs the enhanced radar point cloud features, expressed as:
[0029]
[0030] in, as well as Represents spatial location coordinates; Radar scattering intensity; and These represent the radar points at... direction and The velocity component after directional compensation; and These represent the radar points at... direction and Uncertainty in directional velocity measurement; This refers to the time difference between multiple frames.
[0031] Step 2.4: Perform radar point cloud voxelization on the enhanced radar point cloud features to obtain radar semantic BEV features.
[0032] Furthermore, the aforementioned The multi-frame time difference is the time difference between radar points in historical frames and the current frame, and its calculation expression is:
[0033]
[0034] in, Indicates the current frame timestamp; This represents the timestamp of a historical radar frame.
[0035] Furthermore, the specific process of step 3 is as follows:
[0036] Step 3.1: Construct a velocity-coded input from point-level radar features, wherein the velocity-coded input expression is:
[0037]
[0038] in, and They represent the first radar points at direction and The velocity component in the direction; 10 is the velocity normalization coefficient; and This indicates the azimuth triangle feature of the point; and They represent the first radar points at direction and The root mean square error of velocity measurement in the direction; 1.0 is Normalization coefficient;
[0039] Step 3.2: Project the radar velocity onto the BEV spatial grid through multiple sensing layers. For multiple radar points within the same grid, perform weighted aggregation based on RMS confidence to obtain the radar velocity BEV feature map. The weight expression for aggregation and the expression for the radar velocity BEV feature map are as follows:
[0040]
[0041]
[0042] in, Indicates the first radar points at Root mean square error of velocity measurement in the direction; Indicates the first radar points at The root mean square error of velocity measurement in the direction; 0.1 is a stability constant set to prevent the denominator from being too small; The characteristics indicate that they originate from the radar branch; This indicates that the feature is a velocity-based BEV spatial feature; This represents a BEV spatial grid cell; Indicates the number of elements belonging to this grid. One radar point; Indicates the first The feature vector of each point after passing through the velocity encoding network; Indicates the grid Summing all points within the range.
[0043] Furthermore, in step 3, the radar velocity BEV feature map includes the target's velocity components in the plane, velocity uncertainty, and velocity confidence.
[0044] Furthermore, the specific process of step 5 is as follows:
[0045] Step 5.1: Use the image BEV spatial features and the lidar BEV spatial features as the main fusion path, and use the radar semantic BEV features as the radar residual path;
[0046] Step 5.2: Calculate the evidence value for each modality feature. Confidence level and uncertainty Generate modal fusion weights The calculation expressions for the evidence value, confidence level, uncertainty, and modality fusion weight of each modality feature are as follows:
[0047]
[0048] in, Indicates the first Dirichlet parameters for each mode; Indicates the modality number; 1 represents prior unit evidence;
[0049]
[0050] in, Indicates a stable term;
[0051]
[0052] in, This represents the summation of belief levels across all modalities. To sum the subscripts;
[0053] Step 5.3: Generate the main path fusion result by combining the image BEV features and the LiDAR BEV features through main fusion convolution;
[0054] Step 5.4: Perform independent projection branch transformation on the radar semantic BEV features to generate radar residual features;
[0055] Step 5.5: Adjust the injection intensity of the radar residual features by using learnable radar gating parameters and evidence weights;
[0056] Step 5.6: Add the adjusted radar residual features to the main path fusion result to obtain the fused BEV spatial features, expressed as:
[0057]
[0058] in, This indicates the main path convolution fusion operation; This indicates that the BEV spatial features of the image and the BEV spatial features of the LiDAR are stitched together in the channel dimension; Represents the BEV spatial features of the image branch; This represents the spatial characteristics of the BEV branch of lidar; This represents a mapping function that projects radar features onto a channel. Represents the semantic BEV space features of radar; This represents the learnable gating coefficient.
[0059] Furthermore, the specific process of step 7 is as follows:
[0060] Step 7.1: Based on the position of each candidate target or query target in the BEV space, read the radar velocity information corresponding to the target from the radar velocity BEV map features;
[0061] Step 7.2: Map the target query location to the sampled coordinates in the radar velocity BEV map features;
[0062] Step 7.3: Sample the radar velocity BEV map features to obtain the radar velocity components and velocity uncertainty of the target position;
[0063] Step 7.4: Generate radar confidence based on the velocity uncertainty. If there is no radar point at the sampling location, set the radar confidence to zero.
[0064] Step 7.5: Based on the radar confidence level, perform weighted fusion of the radar velocity and the network predicted velocity, and output the final target velocity result, expressed as:
[0065]
[0066] in, Indicates the final output speed of the target; Indicates the confidence level of physical velocity measurement; This represents the physical velocity obtained from the radar velocity BEV spatial feature sampling; This indicates the network prediction speed obtained from the detection head regression; The complementary weights represent the network prediction speed.
[0067] A multimodal fusion target detection system based on image, lidar, and millimeter-wave radar, for performing the method of any one of claims 1 to 9, comprising:
[0068] Data acquisition module: used to acquire image data, LiDAR point cloud data, and millimeter-wave radar point cloud data;
[0069] Radar enhancement module: used to generate radar semantic BEV features;
[0070] Radar velocity BEV map feature generation module: used to generate radar velocity BEV map features;
[0071] Image and LiDAR BEV spatial feature generation module: used to generate image BEV spatial features and LiDAR BEV spatial features;
[0072] Evidence-guided residual fusion module: used to generate fused BEV spatial features;
[0073] Target detection and velocity detection module: used to perform target detection and velocity detection on fused BEV spatial features;
[0074] Velocity confidence correction module: Used to perform confidence correction on radar velocity to obtain the final velocity result of the target.
[0075] Beneficial effects: (1) By enhancing the time difference and azimuth of multiple radar frames, the temporal and directional information of historical radar points can be enhanced, which is conducive to improving the speed detection capability of moving targets. (2) By using radar dual-path coding, speed information is avoided from being masked by other features during the fusion process. (3) By correcting the network's predicted speed based on the radar speed BEV feature map and confidence level, the target speed estimation can be made more stable. (4) By using evidence-guided residual radar fusion, millimeter-wave radar features can be effectively introduced while preserving the stability of the image plus lidar main path, avoiding the impact of weak millimeter-wave radar features on the performance of the strong backbone network. (5) Applicable to scenarios such as low-light environment, severe weather and complex traffic target movement. Attached Figure Description
[0076] Figure 1 This is a flowchart of the multimodal fusion target detection method based on image, lidar, and millimeter-wave radar of the present invention;
[0077] Figure 2 This is a flowchart illustrating the construction of radar enhancement features according to the present invention;
[0078] Figure 3 This is a diagram of the radar dual-path coding structure of the present invention;
[0079] Figure 4 This invention provides evidence-guided residual radar injection fusion maps.
[0080] Figure 5 This is the velocity confidence correction diagram based on the target query location of the present invention;
[0081] Figure 6 This is a framework diagram of the multimodal fusion target detection system based on image, lidar, and millimeter-wave radar of the present invention. Detailed Implementation
[0082] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0083] like Figure 1 As shown, the multimodal fusion target detection method based on image, lidar, and millimeter-wave radar of the present invention includes the following steps:
[0084] S1: Acquire multi-view image data from 6 surround views, current frame and 9 frames of historical LiDAR point cloud data, and current frame and 5 frames of historical millimeter-wave radar point cloud data. The point cloud range is set as follows:
[0085]
[0086] in, and They respectively represent the detection space in Minimum and maximum values along the axis. and Indicates the detection space is in Minimum and maximum values along the axis. and They respectively represent the detection space in Minimum and maximum values along the axis;
[0087] The following dimensional parameters are selected for the original radar features:
[0088]
[0089] in, , , These represent the three-dimensional position coordinates of the radar point in the reference coordinate system; This indicates the scattering intensity of the radar point; and These represent the radar points at... direction and The compensated velocity component in the direction; and These represent the radar points at... direction and Root mean square error of velocity measurement in the direction.
[0090] After concatenating the time difference dimension, a 9-dimensional feature is obtained:
[0091]
[0092] in, The time difference between the radar points in the historical frame and the current frame, and , Indicates the current frame timestamp; This represents the timestamp of a historical radar frame.
[0093] S2: As Figures 2-3 As shown, enhanced radar point cloud features are constructed, and the semantic radar coding path is input to perform radar point cloud voxelization to generate radar semantic BEV features. The specific process of constructing enhanced radar point cloud features is as follows:
[0094] S2.1: Calculate the planar modulus of each radar point cloud data point. The calculation expression is as follows:
[0095]
[0096] in, Indicates the radar point is at The distance from the origin of the reference coordinate system within the plane; and These represent the points at... shaft and Coordinates along the axis;
[0097] S2.2: Construct the azimuth feature, the expression of which is:
[0098]
[0099]
[0100] in, This represents the azimuth angle of the radar point relative to the origin of the reference coordinate system. Represents the sine value of the azimuth angle; Represents the cosine value of the azimuth angle;
[0101] S2.3: When When it is less than the preset threshold, and Setting it to 0 constructs the enhanced radar point cloud features, expressed as:
[0102]
[0103] in, as well as Represents spatial location coordinates; Radar scattering intensity; and These represent the radar points at... direction and The velocity component after directional compensation; and These represent the radar points at... direction and Uncertainty in directional velocity measurement; This refers to the time difference between multiple frames.
[0104] S2.4: Perform radar point cloud voxelization on the enhanced radar point cloud features to obtain radar semantic BEV features;
[0105] Set the voxel size of the radar point cloud voxelization to:
[0106]
[0107] The first 0.6 indicates Orientational voxel dimensions, in meters; the second 0.6 indicates... Orientational voxel dimensions, in meters; the third 8.0 indicates... Orientational voxel dimensions, in meters;
[0108] The corresponding BEV space grid size is 180×180, with a maximum of 10 radars retained within each voxel. The maximum number of voxels during training and testing is 30,000 and 40,000, respectively. The voxel features are processed by a columnar coding network to obtain 64-channel radar semantic BEV space features, expressed as follows:
[0109]
[0110] in, Represents feature map; subscript This indicates that the feature originates from the millimeter-wave radar branch; superscript This indicates that the feature is a semantic BEV space feature; 64 is the number of channels; These represent the space of the BEV. direction and The number of grids in the direction.
[0111] S3: As Figure 3 As shown, the millimeter-wave radar point cloud data is input into the millimeter-wave radar velocity encoding path to encode the features of radar point-level velocity, generating a radar velocity BEV feature map. The specific process is as follows:
[0112] S3.1: Construct a velocity-coded input from point-level radar features, wherein the velocity-coded input expression is:
[0113]
[0114] in, and They represent the first radar points at direction and The velocity component in the direction; 10 is the velocity normalization coefficient; and This indicates the azimuth triangle feature of the point; and They represent the first radar points at direction and The root mean square error of velocity measurement in the direction; 1.0 is Normalization coefficient.
[0115] S3.2: Projected onto the BEV spatial grid through multiple sensing layers, multiple radar points within the same grid are weighted and aggregated using RMS-based confidence scores to obtain the radar velocity BEV feature map. The weighting expressions for aggregation and the expressions for the radar velocity BEV feature map are as follows:
[0116]
[0117]
[0118] in, Indicates the first radar points at Root mean square error of velocity measurement in the direction; Indicates the first radar points at The root mean square error of velocity measurement in the direction; 0.1 is a stability constant set to prevent the denominator from being too small; The characteristics indicate that they originate from the radar branch; This indicates that the feature is a velocity-based BEV spatial feature; This represents a BEV spatial grid cell; Indicates the number of elements belonging to this grid. One radar point; Indicates the first The feature vector of each point after passing through the velocity encoding network; Indicates the grid Summing all points within the range;
[0119] Thus, the 4-channel radar velocity BEV characteristic map is obtained, expressed as:
[0120]
[0121] Where 4 represents the number of output channels;
[0122] The radar velocity BEV feature map includes the target's velocity components in the plane, velocity uncertainty, and velocity confidence.
[0123] S4: Perform feature extraction and depth view transformation on the image data to generate BEV space features of the image. The specific process is as follows:
[0124] The enhanced image resolution is set to 256×704, and the BEV space transformation boundary parameters are set as follows:
[0125]
[0126]
[0127]
[0128]
[0129] in, express The BEV space discrete boundary parameters for the direction correspond to the minimum coordinates, maximum coordinates, and grid step size, respectively; express The BEV space discrete boundary parameters for the direction correspond to the minimum coordinate, maximum coordinate, and grid step size, respectively; the unit of 0.3 is meters. express The spatial discrete boundary parameters of the direction correspond to the minimum height, maximum height, and step size, respectively; The discrete boundary parameters representing the depth dimension correspond to the minimum depth, maximum depth, and depth interval, respectively.
[0130] The number of discrete depth layers is expressed as:
[0131]
[0132] in, This indicates the number of discrete depth layers; 60 is the maximum depth value; 1 is the minimum depth value; and 0.5 is the interval between adjacent depth layers.
[0133] For each camera, a 25-dimensional parameter vector is extracted, including 4-dimensional intrinsic parameters, 12-dimensional extrinsic parameters, and 9-dimensional image enhancement matrix parameters. This vector is then encoded as camera parameters and added to the network. The acquired image features are subjected to channel modulation, and the discrete depth distribution is predicted. The expression for the discrete depth distribution is as follows:
[0134]
[0135] in, Represents the depth dimension; This represents the depth logits of the deep branch output; This represents the normalized exponential function, used to convert depth logits into a probability distribution across each depth layer;
[0136] The view transformation features are generated by cross-productting the context features, and finally the 80-channel image BEV spatial features are obtained.
[0137] Feature extraction is performed on the LiDAR point cloud data. The LiDAR point cloud serves as an explicit depth indicator for supervising depth prediction in the image branch, generating LiDAR BEV spatial features. The specific process is as follows:
[0138] The lidar branch is represented using voxel dimensions as follows:
[0139]
[0140] The first 0.075 represents Orientation voxel size; the second 0.075 indicates Orientation voxel size; the third 0.2 indicates Dimensions of directional voxels, all in meters;
[0141] Through voxel encoding, sparse encoding, and a BEV backbone network, the spatial features of the 256-channel lidar BEV are obtained; this branch mainly provides stable spatial geometric information.
[0142] S5: As Figure 4 As shown, evidence-guided residual fusion is performed on image BEV spatial features, lidar BEV spatial features, and radar semantic BEV map features to generate fused BEV spatial features. The specific process is as follows:
[0143] S5.1: Use image BEV spatial features and lidar BEV spatial features as the main fusion path, and use radar semantic BEV features as the radar residual path.
[0144] S5.2: Calculate the evidence value for each modality feature separately. Confidence level and uncertainty Generate modal fusion weights The calculation expressions for the evidence value, confidence level, uncertainty, and modality fusion weight of each modality feature are as follows:
[0145]
[0146] in, Indicates the first Dirichlet parameters for each mode; Indicates the modality number; 1 indicates prior unit evidence;
[0147]
[0148] in, Indicates a stable term;
[0149]
[0150] in, This represents the summation of belief levels across all modalities. To sum the subscripts;
[0151] S5.3: Generate the main path fusion result by combining the image BEV features and the LiDAR BEV features through main fusion convolution;
[0152] S5.4: Perform independent projection branch transformation on the radar semantic BEV features to generate radar residual features;
[0153] S5.5: The injection intensity of the radar residual features is adjusted by combining learnable radar gating parameters and evidence weights;
[0154] S5.6: Add the adjusted radar residual features to the main path fusion result to obtain the fused BEV spatial features, expressed as:
[0155]
[0156] in, This indicates the main path convolution fusion operation; This indicates that the BEV spatial features of the image and the BEV spatial features of the LiDAR are stitched together in the channel dimension; Represents the BEV spatial features of the image branch; This represents the spatial characteristics of the BEV branch of lidar; This represents a mapping function that projects radar features onto a channel. Represents the semantic BEV space features of radar; This represents the learnable gating coefficient.
[0157] S6: Perform target detection and velocity detection based on fused BEV spatial features.
[0158] The fused BEV spatial features are input into the detection head, which outputs the target center, size, orientation, category, and network prediction speed.
[0159] S7: As Figure 5 As shown, confidence correction is performed on the radar velocity based on the target query location to obtain the final target velocity result. The specific process is as follows:
[0160] S7.1: Based on the position of each candidate target or query target in the BEV space, read the radar velocity information corresponding to the target from the radar velocity BEV map features;
[0161] S7.2: Map the target query location to the sampled coordinates in the radar velocity BEV map feature;
[0162] S7.3: The radar velocity BEV map features are sampled by bilinear interpolation to obtain the radar velocity components and velocity uncertainty of the target position;
[0163] S7.4: Let the velocity uncertainty obtained from the sampling be... The radar confidence level is generated based on the velocity uncertainty, and its expression is as follows:
[0164]
[0165] in, Represents the Sigmoid function; This represents the velocity uncertainty or RMS value sampled from the radar velocity BEV space; 0.5 is the uncertainty threshold center; 5 is the slope control parameter of the Sigmoid function, used to enhance the distinction between high and low confidence levels.
[0166] If no radar point exists at the sampling location, the radar confidence level is set to zero.
[0167] S7.5: Based on the radar confidence level, the radar velocity and the network predicted velocity are weighted and fused to output the final target velocity result, expressed as:
[0168]
[0169] in, Indicates the final output speed of the target; Indicates the confidence level of physical velocity measurement; This represents the physical velocity obtained from the radar velocity BEV spatial feature sampling; This indicates the network prediction speed obtained from the detection head regression; The complementary weights represent the network prediction speed.
[0170] When radar velocity measurement is reliable, it relies more on physical velocity; when radar points are sparse, targets are stationary, or the confidence level is low, it automatically degenerates into network-predicted velocity.
[0171] This invention employs phased training. The first phase primarily trains the radar branch and fusion module; the second phase enables evidence supervision and joint fine-tuning. Preferably, the training batch size in the first phase is 6, and the training batch size in the second phase is 7; the optimizer uses AdamW, and the learning rate is set to... or .
[0172] like Figure 6 As shown, a multimodal fusion target detection system based on image, lidar, and millimeter-wave radar includes the following modules:
[0173] Data acquisition module: used to acquire image data, LiDAR point cloud data, and millimeter-wave radar point cloud data;
[0174] Radar enhancement module: used to generate radar semantic BEV features;
[0175] Radar velocity BEV map feature generation module: used to generate radar velocity BEV map features;
[0176] Image and LiDAR BEV spatial feature generation module: used to generate image BEV spatial features and LiDAR BEV spatial features;
[0177] Evidence-guided residual fusion module: used to generate fused BEV spatial features;
[0178] Target detection and velocity detection module: used to perform target detection and velocity detection on fused BEV spatial features;
[0179] Velocity confidence correction module: Used to perform confidence correction on radar velocity to obtain the final velocity result of the target.
Claims
1. A multi-modal fusion target detection method based on image, lidar, and millimeter-wave radar, characterized in that, Includes the following steps: Step 1: Acquire image data, LiDAR point cloud data, and millimeter-wave radar point cloud data; Step 2: Construct enhanced radar point cloud features and input the semantic radar coding path to perform radar point cloud voxelization and generate radar semantic BEV features. Step 3: Input the millimeter-wave radar point cloud data into the millimeter-wave radar velocity encoding path, encode the features of radar point-level velocity, and generate a radar velocity BEV feature map. Step 4: Perform feature extraction and depth view transformation on the image data to generate image BEV spatial features; perform feature extraction on the lidar point cloud data to generate lidar BEV spatial features; Step 5: Perform residual fusion of image BEV spatial features, lidar BEV spatial features, and radar semantic BEV map features through evidence-guided analysis to generate fused BEV spatial features. Step 6: Perform target detection and velocity detection on the fused BEV spatial features; Step 7: Perform confidence correction on the radar velocity based on the target query location to obtain the final target velocity result.
2. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 1, characterized in that, In step 1, the image data is multi-view image data with 6 surround views; the lidar point cloud data is the current frame and 9 historical lidar point cloud data; and the millimeter-wave radar point cloud data is the current frame and 5 historical millimeter-wave radar point cloud data.
3. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 1, characterized in that, The lidar point cloud serves as an explicit depth indicator to supervise depth prediction of image branches.
4. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 1, characterized in that, The specific process of step 2 is as follows: Step 2.1: Calculate the planar modulus of each radar point cloud data point. The calculation expression is as follows: in, Indicates the radar point is at The distance from the origin of the reference coordinate system within the plane; and These represent the points at... shaft and Coordinates along the axis; Step 2.2: Construct the azimuth feature, whose expression is: in, This represents the azimuth angle of the radar point relative to the origin of the reference coordinate system. Represents the sine value of the azimuth angle; Represents the cosine value of the azimuth angle; Step 2.3: When When it is less than the preset threshold, and Setting it to 0 constructs the enhanced radar point cloud features, expressed as: in, as well as Represents spatial location coordinates; Radar scattering intensity; and These represent the radar points at... direction and The velocity component after directional compensation; and These represent the radar points at... direction and Uncertainty in directional velocity measurement; This refers to the time difference between multiple frames. Step 2.4: Perform radar point cloud voxelization on the enhanced radar point cloud features to obtain radar semantic BEV features.
5. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 4, characterized in that, The The multi-frame time difference is the time difference between radar points in historical frames and the current frame, and its calculation expression is: in, Indicates the current frame timestamp; This represents the timestamp of a historical radar frame.
6. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 1, characterized in that, The specific process of step 3 is as follows: Step 3.1: Construct a velocity-coded input from point-level radar features, wherein the velocity-coded input expression is: in, and They represent the first radar points at direction and The velocity component in the direction; 10 is the velocity normalization coefficient; and This indicates the azimuth triangle feature of the point; and They represent the first radar points at direction and The root mean square error of velocity measurement in the direction; 1.0 is Normalization coefficient; Step 3.2: Project the radar velocity onto the BEV spatial grid through multiple sensing layers. For multiple radar points within the same grid, perform weighted aggregation based on RMS confidence to obtain the radar velocity BEV feature map. The weight expression for aggregation and the expression for the radar velocity BEV feature map are as follows: in, Indicates the first radar points at Root mean square error of velocity measurement in the direction; Indicates the first radar points at The root mean square error of velocity measurement in the direction; 0.1 is a stability constant set to prevent the denominator from being too small; The characteristics indicate that they originate from the radar branch; This indicates that the feature is a velocity-based BEV spatial feature; This represents a BEV spatial grid cell; Indicates the number of elements belonging to this grid. One radar point; Indicates the first The feature vector of each point after passing through the velocity encoding network; Indicates the grid Summing all points within the range.
7. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 1, characterized in that, In step 3, the radar velocity BEV feature map includes the target's velocity components in the plane, velocity uncertainty, and velocity confidence.
8. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 1, characterized in that, The specific process of step 5 is as follows: Step 5.1: Use the image BEV spatial features and the lidar BEV spatial features as the main fusion path, and use the radar semantic BEV features as the radar residual path; Step 5.2: Calculate the evidence value for each modality feature. Confidence level and uncertainty Generate modal fusion weights The calculation expressions for the evidence value, confidence level, uncertainty, and modality fusion weight of each modality feature are as follows: in, Indicates the first Dirichlet parameters for each mode; Indicates the modality number; 1 represents prior unit evidence; in, Indicates a stable term; in, This represents the summation of belief levels across all modalities. To sum the subscripts; Step 5.3: Generate the main path fusion result by combining the image BEV features and the LiDAR BEV features through main fusion convolution; Step 5.4: Perform independent projection branch transformation on the radar semantic BEV features to generate radar residual features; Step 5.5: Adjust the injection intensity of the radar residual features by using learnable radar gating parameters and evidence weights; Step 5.6: Add the adjusted radar residual features to the main path fusion result to obtain the fused BEV spatial features, expressed as: in, This indicates the main path convolution fusion operation; This indicates that the BEV spatial features of the image and the BEV spatial features of the LiDAR are stitched together in the channel dimension; Represents the BEV spatial features of the image branch; This represents the spatial characteristics of the BEV branch of lidar; This represents a mapping function that projects radar features onto a channel. Represents the semantic BEV space features of radar; This represents the learnable gating coefficient.
9. The multimodal fusion target detection method based on image, lidar, and millimeter-wave radar according to claim 1, characterized in that, The specific process of step 7 is as follows: Step 7.1: Based on the position of each candidate target or query target in the BEV space, read the radar velocity information corresponding to the target from the radar velocity BEV map features; Step 7.2: Map the target query location to the sampled coordinates in the radar velocity BEV map features; Step 7.3: Sample the radar velocity BEV map features to obtain the radar velocity components and velocity uncertainty of the target position; Step 7.4: Generate radar confidence based on the velocity uncertainty. If there is no radar point at the sampling location, set the radar confidence to zero. Step 7.5: Based on the radar confidence level, perform weighted fusion of the radar velocity and the network predicted velocity, and output the final target velocity result, expressed as: in, Indicates the final output speed of the target; Indicates the confidence level of physical velocity measurement; This represents the physical velocity obtained from the radar velocity BEV spatial feature sampling; This indicates the network prediction speed obtained from the detection head regression; The complementary weights represent the network prediction speed.
10. A multi-modal fusion target detection system based on image processing, lidar, and millimeter-wave radar, characterized in that, For performing the method according to any one of claims 1 to 9, comprising: Data acquisition module: used to acquire image data, LiDAR point cloud data, and millimeter-wave radar point cloud data; Radar enhancement module: used to generate radar semantic BEV features; Radar velocity BEV map feature generation module: used to generate radar velocity BEV map features; Image and LiDAR BEV spatial feature generation module: used to generate image BEV spatial features and LiDAR BEV spatial features; Evidence-guided residual fusion module: used to generate fused BEV spatial features; Target detection and velocity detection module: used to perform target detection and velocity detection on fused BEV spatial features; Velocity confidence correction module: Used to perform confidence correction on radar velocity to obtain the final velocity result of the target.