Multi-modal fusion perception method, device and system for automatic driving
By performing spatiotemporal alignment and feature fusion on image and millimeter-wave radar point cloud data in autonomous driving systems, and utilizing cross-attention and lightweight hierarchical gating mechanisms, the problem of single sensors being affected by lighting and inclement weather is solved, achieving efficient multimodal fusion perception, which is suitable for autonomous driving in complex dynamic scenarios.
Patent Information
- Application Number
- CN202511120471.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-12
AI Technical Summary
In existing autonomous driving systems, single sensors are susceptible to the effects of lighting and severe weather. The fusion method of millimeter-wave radar and vision does not fully consider the temporal characteristics, resulting in the breakage of dynamic target trajectories or positional drift. The lack of selective suppression during multimodal feature fusion leads to a waste of computing resources.
By acquiring image and millimeter-wave radar point cloud data, performing preprocessing and feature extraction, and then spatiotemporally aligning the data, a multimodal fusion perception method is constructed using cross-attention mechanism and lightweight hierarchical gating mechanism for feature fusion, including temporal fusion, voxelization, 3D sparse convolution and Doppler velocity compensation.
It effectively alleviates sparsity, improves the perception accuracy and speed of autonomous driving in complex dynamic scenarios, reduces computational complexity, and is suitable for real-time applications.
Smart Images

Figure CN120635850B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic driving, and in particular to a multi-modal fusion perception method, device and system for automatic driving. BACKGROUND
[0002] In the automatic driving system, a single sensor has significant limitations: the camera is easily disturbed by light changes (such as strong light, backlight) and bad weather (such as rain and fog), resulting in target missed detection or false detection; although the millimeter wave radar has all-weather detection capability, the recognition accuracy of static objects and small size targets (such as pedestrians and cones) is insufficient.
[0003] Although the existing millimeter wave radar and vision fusion scheme can partially make up for the above defects, the traditional radar and camera fusion method (such as direct feature splicing or post-fusion) does not fully consider the time sequence characteristics of the millimeter wave radar point cloud, resulting in trajectory breakage or position drift of dynamic targets between consecutive frames; the existing 3D convolution network usually separates the time and space dimensions for processing (such as temporal filtering first and then spatial convolution), resulting in the loss of motion target spatio-temporal context information; when multi-modal feature fusion, the traditional attention mechanism (such as full connection layer weighting) lacks selective inhibition of irrelevant areas, resulting in waste of computing resources.
[0004] Therefore, how to solve the problem of automatic driving perception that cannot be applied to complex dynamic scenes due to the current multi-modal fusion defects has become a technical problem to be solved by the technical personnel in the field. SUMMARY
[0005] The present application provides a multi-modal fusion perception method for automatic driving, a multi-modal fusion perception device for automatic driving, and a multi-modal fusion perception system for automatic driving, which solves the problem of automatic driving perception that cannot be applied to complex dynamic scenes due to the current multi-modal fusion defects in the related art.
[0006] As a first aspect of the present application, a multi-modal fusion perception method for automatic driving is provided, comprising:
[0007] Respectively acquiring image data information and millimeter wave radar point cloud data information of an automatic driving vehicle;
[0008] The millimeter wave radar point cloud data information is preprocessed, and feature extraction is performed to obtain a millimeter wave bird's eye view feature;
[0009] The image data information is image preprocessed, and feature extraction is performed to obtain a camera bird's eye view feature;
[0010] According to the calibration of the camera and the millimeter wave radar, the millimeter wave bird's eye view features and the camera bird's eye view features are spatio-temporally aligned, to obtain the millimeter wave bird's eye view features and the camera bird's eye view features after spatio-temporal alignment;
[0011] According to the cross attention mechanism and the lightweight hierarchical gating mechanism, the millimeter wave bird's eye view features and the camera bird's eye view features after spatio-temporal alignment are fused, to obtain the multi-modal fusion perception result.
[0012] Further, the millimeter wave radar point cloud data information is preprocessed, and feature extraction is performed to obtain millimeter wave bird's eye view features, including:
[0013] The point cloud information in the millimeter wave radar point cloud data information is extracted, and the point cloud information at least includes point cloud frames, point cloud numbers, point cloud positions, point cloud intensities, radial velocities and radar cross-sectional areas;
[0014] The point cloud information is processed by time series fusion to realize point cloud alignment;
[0015] The aligned point cloud information is voxelized to obtain voxelized point cloud information;
[0016] The dynamic features of the moving target are extracted from the voxelized point cloud information to obtain millimeter wave bird's eye view features.
[0017] Further, the dynamic features of the moving target are extracted from the voxelized point cloud information to obtain millimeter wave bird's eye view features, including:
[0018] The voxelized point cloud information is input into a 3D sparse convolutional network for spatio-temporal joint sparse convolution to extract the dynamic features of the moving target, wherein a time channel convolution kernel is embedded in the 3D sparse convolutional network;
[0019] The dynamic features of the moving target are processed by pooling reduction to obtain millimeter wave bird's eye view features.
[0020] Further, the point cloud information is processed by time series fusion to realize point cloud alignment, including:
[0021] According to the Doppler velocity, the point cloud information of the historical frame is back projected to the coordinate system of the point cloud information of the current frame, wherein the projection calculation expression is:
[0022] ,
[0023] wherein, represents the coordinates of the point cloud information of the current frame, represents the coordinates of the point cloud information at the previous time, the coordinates of the point cloud information at the previous time, represents a Doppler velocity, represents a time interval;
[0024] The point cloud target existing in the historical frame and missing in the current frame is supplemented by a history Kalman gain prediction method to realize point cloud alignment.
[0025] Further, the millimeter wave bird's eye view feature and the camera bird's eye view feature are spatio-temporally aligned according to the calibration of the camera and the millimeter wave radar, including:
[0026] The point cloud information of the millimeter wave radar is converted from a radar coordinate system to a camera coordinate system, and the point cloud coordinates are projected to an image plane according to a camera intrinsic matrix;
[0027] For the unaligned regions in the millimeter wave bird's eye view feature and the camera bird's eye view feature, the feature values are filled according to a bilinear interpolation method to realize spatio-temporal alignment.
[0028] Further, the millimeter wave bird's eye view feature and the camera bird's eye view feature after spatio-temporal alignment are fused according to a cross-attention mechanism and a lightweight hierarchical gating mechanism, including:
[0029] The millimeter wave bird's eye view feature after spatio-temporal alignment is taken as Query in the cross-attention mechanism, and the camera bird's eye view feature after spatio-temporal alignment is taken as Key and Value in the cross-attention mechanism;
[0030] The cross-attention weight is calculated according to Query and Key, and the preliminary fusion perception result is obtained by weighted fusion according to the cross-attention weight and Value;
[0031] The preliminary fusion perception result is screened according to the lightweight hierarchical gating mechanism to obtain a multi-modal fusion perception result.
[0032] Further, screening the preliminary fusion perception result according to the lightweight hierarchical gating mechanism includes:
[0033] The region of interest is screened according to the coarse-grained gating;
[0034] The channel weight in the region of interest is adjusted according to the fine-grained gating.
[0035] Further, screening the region of interest according to the coarse-grained gating includes:
[0036] A spatial mask is generated according to a depth separable convolution;
[0037] The region of interest is obtained by screening the preliminary fusion perception result according to the spatial mask.
[0038] As another aspect of the present application, a multi-modal fusion perception device for automatic driving is provided for implementing the multi-modal fusion perception method for automatic driving described above, wherein it comprises:
[0039] an acquisition module for acquiring image data information and millimeter wave radar point cloud data information respectively;
[0040] a point cloud preprocessing module for preprocessing the millimeter wave radar point cloud data information and extracting features to obtain a millimeter wave bird's eye view feature;
[0041] an image preprocessing module for preprocessing the image data information and extracting features to obtain a camera bird's eye view feature;
[0042] a space-time alignment module for aligning the millimeter wave bird's eye view feature and the camera bird's eye view feature in space-time according to the calibration of the camera and the millimeter wave radar, to obtain the millimeter wave bird's eye view feature and the camera bird's eye view feature after space-time alignment;
[0043] a feature fusion module for fusing the millimeter wave bird's eye view feature and the camera bird's eye view feature after space-time alignment according to a cross-attention mechanism and a lightweight hierarchical gating mechanism, to obtain a multi-modal fusion perception result.
[0044] As another aspect of the present application, a multi-modal fusion perception system for automatic driving is provided, wherein it comprises a millimeter wave radar, a camera and the multi-modal fusion perception device for automatic driving described above, the millimeter wave radar and the camera are both in communication connection with the multi-modal fusion perception device for automatic driving;
[0045] the millimeter wave radar is used for collecting millimeter wave radar point cloud data information of an automatic driving vehicle in real time;
[0046] the camera is used for collecting image data information of an automatic driving vehicle in real time;
[0047] the multi-modal fusion perception device for automatic driving is used for fusing features based on a cross-attention mechanism and a lightweight hierarchical gating mechanism according to the millimeter wave radar point cloud data information and the image data information, to obtain a multi-modal fusion perception result.
[0048] The application provides a multi-modal fusion perception method for automatic driving. BRIEF DESCRIPTION OF DRAWINGS
[0049] The accompanying drawings are included to provide a further understanding of the application, and are incorporated in and constitute a part of this specification, illustrate embodiments of the application, and together with the description serve to explain the principles of the application.
[0050] Figure 1 The flow chart of the multi-modal fusion perception method for automatic driving provided by the application.
[0051] Figure 2 The six ring-view camera image data schematic diagram provided by the application.
[0052] Figure 3 The flow chart of the multi-modal fusion perception method for automatic driving provided by the application.
[0053] Figure 4 The flow chart of the multi-modal fusion perception method for automatic driving provided by the application.
[0054] Figure 5 The flow chart of the multi-modal fusion perception method for automatic driving provided by the application.
[0055] Figure 6 The flow chart of the multi-modal fusion perception method for automatic driving provided by the application.
[0056] Figure 7 The flow chart of the multi-modal fusion perception method for automatic driving provided by the application.
[0057] Figure 8 The final result schematic diagram of the millimeter wave bird's eye view feature and the camera bird's eye view feature provided by the present application is shown in the figure.
[0058] Figure 9 The structural block diagram of the multi-modal fusion perception device for automatic driving provided by the present application is shown in the figure.
[0059] Figure 10 The structural block diagram of the multi-modal fusion perception system for automatic driving provided by the present application is shown in the figure. DETAILED DESCRIPTION
[0060] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0061] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.
[0062] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0063] In the present embodiment, a multi-modal fusion perception method for automatic driving is provided, Figure 1 The flow chart of the multi-modal fusion perception method for automatic driving provided by the present embodiment of the present application is shown in the figure, which includes: Figure 1
[0064] S100, image data information and millimeter wave radar point cloud data information of an automatic driving vehicle are acquired respectively;
[0065] In the embodiment of the present application, a millimeter wave radar and a camera are installed on an autonomous vehicle, the millimeter wave radar collects millimeter wave radar point cloud data information in real time, and the camera collects image data information in real time. Specifically, based on an open-source large-scale autonomous driving dataset, Bosch Radar Dataset, K-Radar, etc. can be obtained. 6 ring-view camera image data and millimeter wave radar point cloud data are obtained, as shown in Figure 2
[0066] S200, pre-processing the millimeter wave radar point cloud data information, and performing feature extraction to obtain millimeter wave bird's eye view features;
[0067] Specifically, the millimeter wave radar point cloud data information is pre-processed, the purpose is to perform time series fusion, so as to enhance the feature expression of the moving target in the point cloud data, thereby extracting the moving target feature in the millimeter wave radar point cloud data information, and obtaining the millimeter wave bird's eye view feature.
[0068] S300, image pre-processing the image data information, and performing feature extraction to obtain camera bird's eye view features;
[0069] In the embodiment of the present application, the image data information is image pre-processed, the features of the camera image are obtained in advance, and the camera bird's eye view features are obtained.
[0070] S400, according to the calibration of the camera and the millimeter wave radar, the millimeter wave bird's eye view features and the camera bird's eye view features are spatio-temporally aligned, and the millimeter wave bird's eye view features and the camera bird's eye view features after spatio-temporal alignment are obtained;
[0071] Specifically, the millimeter wave bird's eye view features and the camera bird's eye view features are spatio-temporally aligned, and the feature difference of the unaligned region is calculated, so as to obtain the millimeter wave bird's eye view features and the camera bird's eye view features after spatio-temporal alignment.
[0072] S500, according to the cross-attention mechanism and the lightweight hierarchical gating mechanism, the millimeter wave bird's eye view features and the camera bird's eye view features after spatio-temporal alignment are fused, and a multi-modal fusion perception result is obtained.
[0073] In the embodiment of the present application, based on the millimeter wave bird's eye view features and the camera bird's eye view features after spatio-temporal alignment, the cross-attention mechanism-based feature fusion is performed to obtain a preliminary feature fusion result, and then the lightweight hierarchical gating mechanism is used to screen the preliminary feature fusion result to obtain a multi-modal fusion perception result. The screening of the lightweight hierarchical gating mechanism can reduce invalid calculation and reduce the computational complexity, so as to meet the real-time requirement.
[0074] In summary, the multi-modal fusion perception method for automatic driving provided by the application obtains image data information and millimeter wave radar point cloud data information of an automatic driving vehicle respectively, and obtains camera bird's eye view features and millimeter wave bird's eye view features after pre-processing and feature extraction of the two, then performs spatio-temporal fusion on the two bird's eye view features, finally realizes feature fusion based on a cross-attention mechanism and a lightweight hierarchical gating mechanism to obtain multi-modal fusion perception results. The multi-modal fusion perception method for automatic driving can effectively alleviate sparsity by compensating multiple frames of point clouds and spatio-temporal convolution through time sequence fusion when pre-processing millimeter wave radar point cloud data, and can realize efficient cross-modal fusion through channel grouping and dynamic re-labeling through the cross-attention mechanism and the lightweight hierarchical gating mechanism, and balance between accuracy and speed, thereby providing a practical solution for millimeter wave and camera fusion, especially suitable for automatic driving perception tasks in complex dynamic scenes. Therefore, the multi-modal fusion perception method for automatic driving provided by the application can solve the automatic driving perception problem in complex dynamic scenes caused by the defects of current multi-modal fusion.
[0075] As a specific implementation, the millimeter wave radar point cloud data information is pre-processed and feature extraction is performed to obtain millimeter wave bird's eye view features, as shown in Figure 3 , including:
[0076] S210, extracting point cloud information in the millimeter wave radar point cloud data information, the point cloud information at least including point cloud frames, point cloud numbers, point cloud positions, point cloud intensities, radial velocities and radar cross-sectional areas;
[0077] For the input millimeter wave radar point cloud data, the point cloud information therein is read, including the number of frames (N), the number of point clouds (Npoint), the point cloud position (x, y, z), the point cloud intensity (intensity), the radial velocity (Vr), and the radar cross-sectional area (RCS).
[0078] S220, performing time sequence fusion processing on the point cloud information to realize point cloud alignment;
[0079] Specifically, as shown in Figure 4 , multiple frames of point clouds are aligned before voxelization through time sequence point cloud association, enhancing the feature expression of moving targets.
[0080] In the embodiment of the application, the time sequence fusion processing is performed on the point cloud information to realize point cloud alignment, including:
[0081] 1) The point cloud information of the historical frame is projected to the coordinate system of the point cloud information of the current frame according to the Doppler velocity, and the projection calculation expression is:
[0082] ,
[0083] wherein, denotes the coordinates of the point cloud information of the current frame, denotes the coordinates of the point cloud information of the previous frame at the time point, the coordinates of the point cloud information of the previous frame at the time point, denotes the Doppler velocity, denotes the time interval;
[0084] 2) the point cloud target existing in the historical frame but missing in the current frame is supplemented by a historical Kalman gain prediction manner to realize point cloud alignment.
[0085] Specifically, for the target detected in the continuous multiple frames but missing in the current frame, the position and confidence thereof are predicted by Kalman gain, and the prediction equation is:
[0086] ,
[0087] wherein, F k denotes a state transition matrix, K k denotes a Kalman gain, Z k denotes an observation value, H k denotes an observation matrix.
[0088] S230, voxelizing the aligned point cloud information to obtain voxelized point cloud information;
[0089] It should be understood that the voxelization of the aligned point cloud information is specifically a process of converting discrete point cloud data in a three-dimensional space into a regular three-dimensional grid (voxel) representation. Each voxel is a small cubic unit, similar to a pixel in a two-dimensional image, except in a three-dimensional space. Voxelization is a preprocessing step in point cloud processing, used to simplify data, extract structural features, or adapt to subsequent algorithm requirements. The specific voxelization process can include space division, point cloud assignment, and attribute calculation, wherein the space division is to divide the bounding box (Bounding Box) in which the point cloud is located into a uniform voxel grid according to the set resolution (voxel size); the point cloud assignment is to attribute each point to the corresponding voxel, possibly by nearest neighbor or weighted assignment; the attribute calculation is to assign voxel attributes according to the distribution statistical characteristics (such as the number of points, mean value, normal vector, etc.) of the points in the voxel. The specific voxelization implementation process is well known to those skilled in the art, and will not be described here.
[0090] S240, extracting the dynamic features of the moving target from the voxelized point cloud information to obtain a millimeter wave bird's eye view feature.
[0091] In the embodiment of the present application, the enhanced millimeter wave radar point cloud is voxelized, and in the process of pooling and dimension reduction, a time channel convolution kernel is embedded in the 3D sparse convolution network to extract the dynamic feature response of the moving target through spatio-temporal joint convolution.
[0092] Specifically, the dynamic feature of the moving target is extracted from the voxelized point cloud information to obtain the millimeter wave bird's eye view feature, including:
[0093] 1) input the voxelized point cloud information into the 3D sparse convolution network for spatio-temporal joint sparse convolution to extract the dynamic feature of the moving target, wherein a time channel convolution kernel is embedded in the 3D sparse convolution network;
[0094] The obtained voxelized data is normally sparse-convoluted and then pooled and dimension-reduced to obtain the BEV (bird's eye view) feature of the millimeter wave radar data. In the embodiment of the present application, a time channel convolution kernel is embedded in the process of sparse convolution to enhance the expression feature of the moving object.
[0095] 2) pool and dimension-reduce the dynamic feature of the moving target to obtain the millimeter wave bird's eye view feature.
[0096] Specifically, let the current time be t, the time one moment ago be t-τ, and the input data feature be , then the output feature is ; the size of the spatio-temporal joint convolution kernel is defined as , wherein kt represents the time dimension kernel size, kx and ky represent the spatial dimension kernel size; and the relationship is:
[0097] ,
[0098] , wherein W(τ, i, j) represents a learnable weight matrix.
[0099] In the embodiment of the present application, for the data of the camera, the architecture of BEVFusion can be specifically used. After the camera features are extracted through the camera Encoder, the BEV features of the camera are obtained after the improved LSS (Lift-splat-shoot). In order to avoid the unreliability of the depth estimation obtained by LSS, the embodiment of the present application performs spatio-temporal alignment on the BEV features of the millimeter wave radar and the BEV features of the camera through spatial alignment, projects the radar point cloud to the image plane by using the calibration external parameters, and constructs the geometric correspondence relationship between the two modalities.
[0100] Specifically, the millimeter wave bird's eye view feature and the camera bird's eye view feature are spatio-temporally aligned according to the calibration of the camera and the millimeter wave radar, as shown in Figure 5 , including:
[0101] S410, convert the point cloud information of the millimeter wave radar from a radar coordinate system to a camera coordinate system, and project the point cloud coordinates to an image plane according to a camera intrinsic matrix;
[0102] Specifically, as shown in the following formula, the radar point cloud is converted from the radar coordinate system to the camera coordinate system, and then the point cloud coordinates are projected to the image plane through the camera intrinsic matrix K: Figure 6
[0103] ,
[0104] ,
[0105] wherein, represents the point cloud information in the millimeter wave radar coordinate system, represents the point cloud information in the camera coordinate system, R 3×3 represents a rotation matrix, t 3×1 represents a translation vector, (u, v) represents pixel coordinates, and z cam represents the depth in the camera coordinate system.
[0106] S420, for the region where the millimeter wave bird's eye view feature and the camera bird's eye view feature are not aligned, the feature values are filled according to the bilinear interpolation method to realize the spatio-temporal alignment.
[0107] In the embodiment of the application, for the region where the radar BEV feature and the image BEV feature are not aligned, the bilinear interpolation method is used to fill the feature values, and the expression is as follows:
[0108] ,
[0109] wherein, represents the interpolation weight of the four adjacent grid points, represents the feature point that is not aligned, represents the point cloud information around
[0110] In the embodiment of the application, the millimeter wave bird's eye view feature and the camera bird's eye view feature after spatio-temporal alignment are fused according to the cross attention mechanism and the lightweight hierarchical gating mechanism, as shown in the following formula, including: Figure 7
[0111] S510, the millimeter wave bird's eye view feature after spatio-temporal alignment is taken as Query in the cross attention mechanism, and the camera bird's eye view feature after spatio-temporal alignment is taken as Key and Value in the cross attention mechanism;
[0112] Specifically, when fusing the millimeter wave bird's eye view feature and the camera bird's eye view feature, the millimeter wave radar feature is taken as the Query, the camera feature is taken as the Key and the Value, and the feature fusion is realized through the cross attention mechanism; and then a lightweight hierarchical gating mechanism is introduced, and first, the region of interest is screened based on the BEV spatial position through the coarse-grained gating, and then the channel weight is adjusted in the screened region through the fine-grained gating.
[0113] In the embodiment of the application, the cross attention mechanism is adopted, and specifically, the millimeter wave feature Q∈R N×d is taken as the Query, the camera feature K∈R M×d and V∈R M×d are taken as the key and the value respectively, wherein N and M represent the number of spatial positions, and d represents the feature dimension.
[0114] S520, the cross attention weight is calculated according to the Query and the Key, and the Value is weighted and fused according to the cross attention weight, to obtain a preliminary fusion perception result;
[0115] Specifically, the attention weight calculation adopts the scaled dot product form, and is normalized through the Softmax function to obtain the attention matrix A; and the multiplication with V is weighted summation to obtain the fused feature F fusion , that is, the expression feature after single-head attention fusion. Subsequently, the multi-head attention mechanism is introduced, the Query, the Key and the Value are split into groups for parallel calculation, and the final feature Fmuti-head is the splicing of the outputs of each head.
[0116] Specifically, the implementation formula is as follows:
[0117] ,
[0118] ,
[0119] ,
[0120] Wherein, Q represents the millimeter wave bird's eye view feature after space-time alignment, which is obtained after being multiplied by a weight matrix; K and V represent the camera bird's eye view feature after space-time alignment, which are obtained after being multiplied by two weight matrices respectively; Concat represents the splicing function; d represents the matrix dimension of Q or K; W0∈R hd×d represents the output projection matrix.
[0121] S530, the preliminary fusion perception result is screened according to the lightweight hierarchical gating mechanism, to obtain a multi-modal fusion perception result.
[0122] In the embodiment of the present application, by adopting a lightweight hierarchical gating mechanism, specifically by two-level screening of spatial mask and channel attention, effective features are screened level by level, and invalid calculation is reduced.
[0123] Specifically, screening the preliminary fusion perception result according to the lightweight hierarchical gating mechanism comprises:
[0124] 1) screening a region of interest from the preliminary fusion perception result according to coarse-grained gating;
[0125] It should be understood that the region of interest in the preliminary fusion perception result can be obtained by screening the preliminary fusion perception result through coarse-grained gating.
[0126] Further specifically, screening the region of interest from the preliminary fusion perception result according to coarse-grained gating comprises:
[0127] 11) generating a spatial mask according to a depth separable convolution;
[0128] 12) screening the preliminary fusion perception result according to the spatial mask to obtain the region of interest.
[0129] In the embodiment of the present application, coarse-grained spatial gating: let the input feature map be , a depth separable convolution (using channel-wise convolution, without spatial fusion) is used to generate a spatial mask , let the screened feature be :
[0130] ,
[0131] ,
[0132] ,
[0133] wherein, , represents a depth separable convolution function, and compared with a traditional convolution, is used to significantly reduce the amount of calculation and the number of parameters; represents an activated spatial mask; represents element-wise multiplication.
[0134] 2) adjusting the channel weight in the region of interest according to fine-grained gating.
[0135] In the embodiment of the present application, fine-grained channel gating: let the input feature map after preliminary screening be , a weight vector is calculated through a channel attention mechanism, let the final output feature be :
[0136] ,
[0137] ,
[0138] ,
[0139] in, This represents the average pooling function. and All represent the weights of the fully connected layer, and r represents the compression ratio. This indicates that the channel dimension is broadcast multiplied.
[0140] Finally, the radar channel network in the sensor backbone network is improved to make it suitable for lower-cost millimeter-wave radar. Furthermore, a lightweight hierarchical gating attention mechanism is used, specifically through two-level filtering of spatial masks and channels, to reduce invalid computation by 70%, thereby reducing computational complexity and meeting real-time requirements. A schematic diagram of the final result of feature fusion of millimeter-wave bird's-eye view features and camera bird's-eye view features in this embodiment of the invention is shown below. Figure 8 As shown in Table 1, the performance of BEVFusio is compared with its parameter indicators.
[0141] Table 1. Performance Comparison of the Fusion Algorithm of the Present Invention and BEVFusio
[0142]
[0143] While maintaining high detection accuracy (mAP improved by 3.2%), this invention, through motion feature enhancement and lightweight temporal fusion design, is suitable for real-time autonomous driving applications in dynamic and complex scenarios, with improvements in speed / direction estimation (MAVE / MAOE), tracking stability, and embedded deployment efficiency (latency reduced by 33.3%).
[0144] In summary, the multi-modal fusion perception method for automatic driving provided by the application is characterized in that, in view of the characteristics of the sparse point cloud of the millimeter wave radar, multiple frames of point clouds are associated before voxelization to enhance the feature extraction capability of the moving target; in the process of convolution, pooling and dimension reduction, a time channel convolution kernel is embedded to extract the feature response of the moving target through convolution operation. In the spatial alignment module of the Bev features of the millimeter wave radar and the Bev features of the camera, the radar point cloud is projected to the image plane by using the calibration parameters of the camera and the millimeter wave radar, and the geometric correspondence between the two modalities is constructed. In the modal fusion, a cross-attention mechanism is used, the millimeter wave features are taken as Query, the camera features are taken as Key and Value, feature fusion is performed, and a lightweight hierarchical gating mechanism is introduced, a coarse-grained gating is based on the BEV spatial position to screen the region of interest, and a fine-grained gating adjusts the feature weight in the screening region. The time sequence fusion process effectively alleviates the sparseness through multi-frame point cloud compensation and spatio-temporal convolution, the grouping gating module realizes efficient cross-modal fusion through channel grouping and dynamic re-calibration, and the balance between precision and speed is achieved through the cooperation of the two, thereby providing a practical solution for millimeter wave and camera fusion, which is especially suitable for automatic driving perception tasks in complex dynamic scenes.
[0145] As another embodiment of the application, a multi-modal fusion perception device 100 for automatic driving is provided for implementing the multi-modal fusion perception method for automatic driving described above, wherein, as shown in Figure 9 the device comprises:
[0146] The acquisition module 110 is configured to acquire image data information and millimeter wave radar point cloud data information respectively.
[0147] The point cloud preprocessing module 120 is configured to preprocess the millimeter wave radar point cloud data information and extract features to obtain millimeter wave bird's eye view features.
[0148] The image preprocessing module 130 is configured to preprocess the image data information and extract features to obtain camera bird's eye view features.
[0149] The space-time alignment module 140 is configured to perform space-time alignment on the millimeter wave bird's eye view features and the camera bird's eye view features according to the calibration of the camera and the millimeter wave radar, and obtain the space-time aligned millimeter wave bird's eye view features and the space-time aligned camera bird's eye view features.
[0150] The feature fusion module 150 is configured to perform feature fusion on the space-time aligned millimeter wave bird's eye view features and the space-time aligned camera bird's eye view features according to the cross-attention mechanism and the lightweight hierarchical gating mechanism, and obtain the multi-modal fusion perception result.
[0151] The application provides a multi-modal fusion perception device for automatic driving, which obtains image data information and millimeter wave radar point cloud data information of an automatic driving vehicle respectively, and obtains camera bird's eye view features and millimeter wave bird's eye view features after pre-processing and feature extraction of the two, then performs spatio-temporal fusion on the two bird's eye view features, and finally realizes feature fusion based on a cross attention mechanism and a lightweight hierarchical gating mechanism to obtain a multi-modal fusion perception result. The multi-modal fusion perception method for automatic driving can effectively alleviate sparsity by compensating multiple frames of point clouds and spatio-temporal convolution through time sequence fusion when pre-processing millimeter wave radar point cloud data, and can realize efficient cross-modal fusion through channel grouping and dynamic re-scaling through the cross attention mechanism and the lightweight hierarchical gating mechanism. The two mechanisms balance between accuracy and speed and provide a feasible solution for millimeter wave and camera fusion, and are especially suitable for automatic driving perception tasks in complex dynamic scenes. Therefore, the multi-modal fusion perception device for automatic driving can solve the automatic driving perception problem in complex dynamic scenes caused by the defects of current multi-modal fusion.
[0152] The specific working principle of the multi-modal fusion perception device for automatic driving provided by the application can be referred to the description of the multi-modal fusion perception method for automatic driving in the foregoing, which will not be repeated here.
[0153] As another embodiment of the application, a multi-modal fusion perception system 10 for automatic driving is provided, as shown in the accompanying drawings, comprising a millimeter wave radar 200, a camera 300 and the multi-modal fusion perception device 100 for automatic driving described in the foregoing, wherein the millimeter wave radar 200 and the camera 300 are in communication connection with the multi-modal fusion perception device 100 for automatic driving. Figure 10
[0154] The millimeter wave radar 200 is used to collect millimeter wave radar point cloud data information of an automatic driving vehicle in real time.
[0155] The camera 300 is used to collect image data information of an automatic driving vehicle in real time.
[0156] The multi-modal fusion perception device 100 for automatic driving is used to perform feature fusion based on a cross attention mechanism and a lightweight hierarchical gating mechanism according to the millimeter wave radar point cloud data information and the image data information to obtain a multi-modal fusion perception result.
[0157] The application provides a multi-modal fusion perception system for automatic driving, which adopts the multi-modal fusion perception device for automatic driving described above, obtains image data information and millimeter wave radar point cloud data information of an automatic driving vehicle respectively, and obtains camera bird's eye view features and millimeter wave bird's eye view features after pre-processing and feature extraction of the two, then performs space-time fusion on the two bird's eye view features, finally realizes feature fusion based on a cross attention mechanism and a lightweight layered gating mechanism, and obtains a multi-modal fusion perception result. The multi-modal fusion perception method for automatic driving can effectively alleviate sparseness by compensating multiple frames of point clouds and space-time convolution through time sequence fusion when pre-processing millimeter wave radar point cloud data, and can realize efficient cross-modal fusion through channel grouping and dynamic rescaling by the cross attention mechanism and the lightweight layered gating mechanism, so that balance between precision and speed is achieved, and a practical solution for millimeter wave and camera fusion is provided, which is especially suitable for automatic driving perception tasks in complex dynamic scenes. Therefore, the multi-modal fusion perception system for automatic driving provided by the application can solve the automatic driving perception problem in complex dynamic scenes caused by the defects of current multi-modal fusion.
[0158] The specific working principle of the multi-modal fusion perception system for automatic driving provided by the application can be referred to the description of the multi-modal fusion perception method for automatic driving described above, and will not be repeated here.
[0159] It can be understood that the above embodiments are only exemplary embodiments adopted for illustrating the principles of the application, and the application is not limited thereto. Various modifications and improvements can be made by those skilled in the art without departing from the spirit and essence of the application, and these modifications and improvements are also considered to be within the protection scope of the application.
Claims
1. A multi-modal fusion perception method for autonomous driving, characterized in that, The method comprises the following steps: respectively acquiring image data information and millimeter wave radar point cloud data information of an autonomous vehicle; preprocessing the millimeter wave radar point cloud data information and performing feature extraction to obtain a millimeter wave bird's eye view feature; performing image preprocessing on the image data information and performing feature extraction to obtain a camera bird's eye view feature; spatiotemporally aligning the millimeter wave bird's eye view feature and the camera bird's eye view feature according to the calibration of the camera and the millimeter wave radar, to obtain the spatiotemporally aligned millimeter wave bird's eye view feature and the camera bird's eye view feature; performing feature fusion on the spatiotemporally aligned millimeter wave bird's eye view feature and the camera bird's eye view feature according to a cross-attention mechanism and a lightweight hierarchical gating mechanism, to obtain a multi-modal fusion perception result; wherein the preprocessing of the millimeter wave radar point cloud data information and the feature extraction to obtain the millimeter wave bird's eye view feature comprise: extracting point cloud information in the millimeter wave radar point cloud data information, the point cloud information at least comprising a point cloud frame, a point cloud number, a point cloud position, a point cloud intensity, a radial velocity and a radar cross-sectional area; performing time series fusion processing on the point cloud information to realize point cloud alignment; performing voxelization processing on the aligned point cloud information to obtain voxelized point cloud information; extracting dynamic features of a moving target from the voxelized point cloud information to obtain the millimeter wave bird's eye view feature; wherein the extracting of the dynamic features of the moving target from the voxelized point cloud information to obtain the millimeter wave bird's eye view feature comprises: inputting the voxelized point cloud information into a 3D sparse convolution network to perform spatiotemporal joint sparse convolution, to extract the dynamic features of the moving target, wherein a time channel convolution kernel is embedded in the 3D sparse convolution network; performing pooling dimension reduction processing on the dynamic features of the moving target to obtain the millimeter wave bird's eye view feature.
2. The multi-modal fusion perception method for automatic driving according to claim 1, wherein, The time series fusion processing on the point cloud information to realize point cloud alignment comprises: projecting the point cloud information of a historical frame to the coordinate system of the point cloud information of a current frame in a reverse direction according to a Doppler velocity, wherein a projection calculation expression is: , wherein, coordinates of the point cloud information of the current frame, coordinates of the point cloud information of the current frame, coordinates of the point cloud information of the current frame, coordinates of the point cloud information of the current frame, indicates a Doppler velocity, indicates a time interval; complementing the point cloud target existing in the historical frame but missing in the current frame by a historical Kalman gain prediction, to realize point cloud alignment.
3. The multi-modal fusion perception method for automatic driving according to claim 1, wherein, The spatiotemporal alignment of the millimeter wave bird's eye view feature and the camera bird's eye view feature according to the calibration of the camera and the millimeter wave radar comprises: converting the point cloud information of the millimeter wave radar from a radar coordinate system to a camera coordinate system, and projecting the point cloud coordinates to an image plane according to a camera intrinsic matrix; for the misaligned regions in the millimeter wave bird's eye view feature and the camera bird's eye view feature, filling feature values according to a bilinear interpolation method to realize spatiotemporal alignment.
4. The multi-modal fusion perception method for automatic driving according to claim 1, wherein, The feature fusion on the spatiotemporally aligned millimeter wave bird's eye view feature and the camera bird's eye view feature according to the cross-attention mechanism and the lightweight hierarchical gating mechanism comprises: taking the spatiotemporally aligned millimeter wave bird's eye view feature as Query in the cross-attention mechanism, and taking the spatiotemporally aligned camera bird's eye view feature as Key and Value in the cross-attention mechanism. According to the Query and the Key, cross-attention weights are calculated, and a preliminary fusion perception result is obtained by weighted fusion according to the cross-attention weights and the Value; According to the lightweight hierarchical gating mechanism, the preliminary fusion perception result is screened to obtain a multi-modal fusion perception result.
5. The multi-modal fusion perception method for automatic driving according to claim 4, wherein, According to the lightweight hierarchical gating mechanism, the preliminary fusion perception result is screened, including: According to the coarse-grained gating, a region of interest is screened from the preliminary fusion perception result; According to the fine-grained gating, the channel weights in the region of interest are adjusted.
6. The multi-modal fusion perception method for automatic driving according to claim 5, characterized in that, According to the coarse-grained gating, a region of interest is screened from the preliminary fusion perception result, including: According to the depth separable convolution, a spatial mask is generated; According to the spatial mask, the preliminary fusion perception result is screened to obtain a region of interest.
7. A multi-modal fusion perception device for autonomous driving, configured to implement the multi-modal fusion perception method for autonomous driving according to any one of claims 1 to 6. Including: An acquisition module is configured to acquire image data information and millimeter wave radar point cloud data information respectively; A point cloud preprocessing module is configured to preprocess the millimeter wave radar point cloud data information and extract features to obtain a millimeter wave bird's eye view feature; An image preprocessing module is configured to preprocess the image data information and extract features to obtain a camera bird's eye view feature; A space-time alignment module is configured to perform space-time alignment on the millimeter wave bird's eye view feature and the camera bird's eye view feature according to the calibration of the camera and the millimeter wave radar, to obtain space-time aligned millimeter wave bird's eye view features and camera bird's eye view features; A feature fusion module is configured to perform feature fusion on the space-time aligned millimeter wave bird's eye view features and camera bird's eye view features according to a cross-attention mechanism and a lightweight hierarchical gating mechanism, to obtain a multi-modal fusion perception result.
8. A multi-modal fusion perception system for use in autonomous driving, comprising: Including: A millimeter wave radar, a camera, and the multi-modal fusion perception device for automatic driving according to claim 7, wherein the millimeter wave radar and the camera are in communication connection with the multi-modal fusion perception device for automatic driving; The millimeter wave radar is configured to collect millimeter wave radar point cloud data information of an automatic driving vehicle in real time; The camera is configured to collect image data information of the automatic driving vehicle in real time; The multi-modal fusion perception device for automatic driving is configured to perform feature fusion based on a cross-attention mechanism and a lightweight hierarchical gating mechanism according to the millimeter wave radar point cloud data information and the image data information, to obtain a multi-modal fusion perception result.
Citation Information
Patent Citations
Three-dimensional sensing method based on millimeter wave radar and camera aerial view fusion
CN118038396A
3D target detection method and system based on multi-modal fusion under BEV visual angle
CN118365891A