Automatic driving multi-mode perception data information fusion method

Through the multimodal fusion method based on sparse features, the problem of huge computing power of multimodal perceptual data fusion is solved, and more efficient and accurate data fusion is achieved, making full use of the advantages of sensors such as lidar, camera and millimeter wave radar.

CN120451724APending Publication Date: 2025-08-08ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510614076.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing multimodal perceptual data fusion method has a large amount of calculation during dense feature processing, resulting in high computing power requirements and limited accuracy, and it is impossible to effectively utilize the advantages of multiple sensors.

Method used

A multimodal fusion method based on sparse features is adopted, and the feature maps of different modal perceptual data are obtained, reference points and their derivative points are generated, and feature sampling and weight configuration are performed under their respective coordinate systems, ultimately realizing the fusion of multimodal data.

Benefits of technology

It effectively reduces the computing power demand for multimodal fusion, improves the accuracy of fusion, and makes full use of the advantages of different sensors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451724A_ABST
    Figure CN120451724A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving multi-mode perception data information fusion method, and the main conception is that a dense algorithm is abandoned, and multi-mode fusion based on sparse features is adopted. Specifically, the method comprises the following steps: pre-acquiring feature maps of different modal sensing data on a vehicle; based on instance features of a target point in the sensing data, obtaining a reference point which is shared by each mode and represents a target object; generating a plurality of derivative points corresponding to the reference points, and projecting the reference points and the position information of the derivative points to respective coordinate systems of different modal sensing data; performing feature sampling on the corresponding feature maps based on coordinate systems of the reference points and the derivative points thereof in each mode; and according to sources of different modal perception data, respectively configuring weights for each modal feature and executing fusion. According to the method, after the reference point shared by different modes is projected back to each mode, feature sampling and fusion are carried out respectively, so that the problem of huge multi-mode fusion computing power can be effectively solved, and the method has the effect of improving the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method for fusing multimodal perception data information for autonomous driving. Background Art

[0002] In the field of autonomous driving (intelligent driving and assisted driving), vehicles often utilize multiple sensor systems. For example, 4D imaging millimeter wave sensors, lidar, and cameras each have distinct advantages and disadvantages. Cameras can capture color information and, using deep learning models, extract texture and color features. However, they lack depth information, resulting in higher 3D perception errors among many sensors. LiDAR point clouds can extract position and shape features through deep learning models, providing relatively accurate 3D spatial perception. However, they lack texture features, making them prone to misjudging objects with similar shapes. 4D imaging millimeter wave point clouds provide 3D spatial perception with velocity information, allowing deep learning models to extract shape and object movement features. They are less susceptible to environmental influences, resulting in accuracy that falls between the previous two sensors. Therefore, the question of how to integrate features extracted by multiple sensors to leverage their respective strengths and complement each other has become a key focus in this field.

[0003] Based on this, the industry has proposed two types of deep learning perception multimodal feature-level pre-fusion methods in the dense feature track: Transformer-based and Lift Splat Shoot (LSS)-based:

[0004] First, the TransFusion series belongs to the Transformer-based category. It requires projecting the camera into a Bird's Eye View (BEV) and then fusing it with the LiDAR. After the LiDAR is expressed as voxels or pillars, the voxel or pillar features are used as transformer queries, and the camera features are used as values and keys. In other words, the camera features are projected into voxel or pillar feature expressions of the LiDAR through the Transformer, all of which are dense features.

[0005] Secondly, the BevFusion series belongs to the LSS-based category. Although it does not use the transformer architecture, it also needs to project the camera features into voxel and then flatten it into BEV using the bevpool algorithm. It is also a pre-fusion method at the dense feature deep learning perception multimodal feature level.

[0006] Regardless of the method, it needs to be projected into the BEV algorithm and then fused under dense features and the same viewing angle. The BEV algorithm belongs to dense features, and dense features have a large amount of calculation. In a large range, the BEV resolution affects the accuracy. The higher the resolution, the higher the accuracy, but the more computing power is required. Figure 1 As shown in the figure, when the image is projected into the BEV dense algorithm, the black part in the figure indicates that the feature is 0, and the non-black part indicates that there is a feature value. The brighter the feature value, the higher the feature value. It can be seen that in the feature map, most of the feature values are 0, so most of them are useless features when performing convolution operations on the feature map. Summary of the Invention

[0007] In view of the above, the present invention aims to provide a method for fusing multimodal perception data information for autonomous driving to solve the technical problems mentioned above.

[0008] The technical solution adopted in the present invention is as follows:

[0009] The present invention provides a method for fusing multimodal perception data information for autonomous driving, comprising:

[0010] Pre-acquire feature maps of different modal perception data on the vehicle;

[0011] Based on the instance features of the target point in the perception data, a reference point representing a target object shared by all modalities is obtained;

[0012] generating a plurality of derivative points corresponding to the reference point;

[0013] Projecting the position information of the reference point and its derivative points into respective coordinate systems of different modal perception data;

[0014] Based on the coordinate systems of the reference point and its derivative points in each mode, feature sampling is performed on the corresponding feature maps respectively;

[0015] According to the sources of different modal perception data, weights are configured for each modal feature and fusion is performed.

[0016] In at least one possible implementation, multi-position compensation is performed on each of the reference points to obtain multiple dynamic positions to represent derived points, and the derived point corresponding to each of the reference points is used as a key point of the target object.

[0017] In at least one possible implementation, after a reference point and its derivative points are projected into the geometric space corresponding to each mode, position information and feature vectors of the reference point and its derivative points in the geometric space of each mode are extracted respectively.

[0018] In at least one possible implementation, the fusion method further includes: before generating the derived points, each reference point is subjected to feature matching clustering via a self-attention mechanism.

[0019] In at least one possible implementation, the sources of the modal perception data include at least two of the following: lidar, camera, and millimeter-wave radar.

[0020] In at least one possible implementation manner, the coordinate system includes: a 3D Cartesian coordinate system, an image coordinate system, and a 2D top-view Cartesian coordinate system.

[0021] Compared with the existing technology, the main design concept of the present invention is to abandon the dense algorithm and adopt multimodal fusion based on sparse features. Specifically, the feature maps of the different modal perception data on the vehicle are obtained in advance; based on the instance features of the target point in the perception data, a reference point representing a target object shared by each modality is obtained; multiple derivative points corresponding to the reference point are generated, and the position information of the reference point and its derivative points are projected into the respective coordinate systems of the different modal perception data; based on the coordinate system of the reference point and its derivative points in each modality, the corresponding feature maps are sampled respectively; according to the source of the different modal perception data, the weights of each modal feature are configured and fusion is performed. After the present invention projects the reference point shared by different modalities back to each modality, feature sampling is performed separately and multimodal data fusion is completed. This method can effectively solve the problem of huge multimodal fusion computing power and has the effect of improving accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described below with reference to the accompanying drawings, in which:

[0023] Figure 1 Project the existing image into a representation diagram of the BEV dense algorithm in the feature map;

[0024] Figure 2 Flowchart of a method for fusing multimodal perception data information for autonomous driving provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0026] Before elaborating on the following solutions of the present invention, further analysis and deduction are required: The essence of current conventional sparse algorithms lies in using a deformable transformer to select only valid features for convolution-like operations. This reduces computing power requirements compared to dense algorithms, and feature strength is less likely to be diluted by zero feature values within the receptive field. However, sparse algorithms only extract valid locations for convolution-like operations, which results in feature map corruption and creates problems in multimodal fusion. Even with the proposed SparseFusion, this approach uses an instance query for each modality to encode the desired output features as a query. The query is then matched against the key (point cloud or image features) using attention. Features are then extracted from the value (point cloud or image features) to generate instances. Finally, the instance fusion operation uses transformers to perform multimodal feature matching. Specifically, the number of transformers used increases with the number of modalities, and a self-attention transformer is added for instance fusion. This process does not significantly improve computing power requirements.

[0027] Taking laser point cloud and camera data as an example, the sparse features of the point cloud are located in the Cartesian coordinate system, while the sparse features of the camera image are located in the image coordinate system. Therefore, if these two perception modalities are to be fused, they need to be fused in the same domain, that is, the laser point cloud should be projected into the image coordinate system, or the image should be projected into the Cartesian coordinate system. Only when each modality is in the same coordinate system can it be more convenient to fuse them.

[0028] The present invention directly performs query fusion (query fusion) of instance query features in different coordinate systems. The same query vector query simultaneously extracts values (data features) from different modalities, encodes them, and fuses them. Therefore, only a single transformer is required to achieve perceptual information fusion across multiple modalities, significantly reducing computing power requirements.

[0029] Therefore, the present invention proposes an embodiment of a method for fusing multimodal perception data information for autonomous driving. Specifically, Figure 2 shown, including:

[0030] Step S1: pre-acquire feature maps of different modal perception data on the vehicle;

[0031] Different modalities may have a 3D Cartesian coordinate system for lidar, an image coordinate system for a camera, and a 2D overhead Cartesian coordinate system for a millimeter-wave radar. Each modality has a separate backbone (backbone network, referring to the basic network used to extract features). Each backbone performs feature extraction based on the characteristics of the different modalities. Subsequently, the backbone of each perception modality will output its own feature map, such as voxels for lidar, pixels for cameras, and pillars for millimeter-wave radar. The deformable transformer mentioned in the subsequent steps of the embodiment will perform feature sampling on the feature maps of different forms corresponding to each modality.

[0032] Step S2: Based on the instance features of the target point in the perception data, a reference point representing a target object shared by different modalities is obtained;

[0033] In some implementation examples, the encoding of the generated 3D point can be used as an instance feature, and after generating several fixed point encodings from the 3D anchor / key points, the encodings are added to the aforementioned instance feature to obtain a query, and a query is regarded as a reference point. In subsequent steps, each reference point will derive multiple points for use as key point searches for the target object, that is, each reference point generates several 3D key points. We will not go into details here, but it can be further pointed out that, preferably, each reference point as a 3D point encoding can first pass through a self-attention module to perform feature matching between the feature encodings of the 3D points, that is, similar point features will be matched together to form a clustering effect.

[0034] Step S3: generating a plurality of derivative points corresponding to the reference point;

[0035] Step S4: projecting the position information of the reference point and its derived points into the respective coordinate systems of the different modal perception data;

[0036] Step S5: based on the coordinate systems of the reference point and its derivative points in each mode, perform feature sampling on the corresponding feature maps respectively;

[0037] Step S6: According to the sources of different modal perception data, weights are assigned to the features of each modality and fusion is performed.

[0038] Regarding steps S3 through S6, in some preferred embodiments of the present invention, a deformable transformer is used to transform 2D images into 3D images and incorporate multimodal fusion. The deformable transformer is illustrated by the following formula: x is a feature map. Features are sampled at corresponding locations in x from a reference point p and its delta_p derivatives. Multiple derivatives are then fused using weights A to produce a feature vector for the target instance instance.

[0039]

[0040] The above process mainly involves two aspects:

[0041] First, key points are derived.

[0042] Specifically, the key points generation module can perform multiple position compensations on each reference point to derive several dynamic positions, and decode the features of the reference point and its several derived points into a 3D coordinate system x, y, z representation. In the above example, the reference point and its derived points can all be in the 3D Cartesian coordinate system and used to search for the key feature positions of the target object in 3D space. In other words, each target object corresponds to a reference point, and the derived points of each reference point are the key points of the target object itself.

[0043] The second aspect is query fusion.

[0044] The main idea of query fusion is to share the same query. Instead of generating multiple queries due to multimodality, the shared query is used to search for key points of the target object, which are then projected back into various perceptual modalities. Based on each modality, each is sampled separately and then multiplied by learnable weights for multimodal feature fusion.

[0045] To elaborate, after each reference point in the present invention generates several 3D points, these 3D point positions are projected back to each modality. Each coordinate system has its own different projection function, such as projecting 3D points into voxel coordinate system features for lidar features, pixel coordinate system features for camera features, and pillar coordinate system features for millimeter wave radar.

[0046] After projecting several 3D points to the corresponding positions in different coordinate systems, the aforementioned feature sampling is performed. This mainly refers to extracting the feature vector from the corresponding position after any 3D point is projected: the same 3D point is projected back into the geometric space corresponding to multiple modes, and the feature vector of the position is extracted respectively, indicating the corresponding position of the key point of the target object in the geometric space of each mode, and extracting the feature vector from it.

[0047] At this point, each reference point generates several 3D keypoints, each of which is projected onto a different modality. Features are then extracted from the corresponding positions of the projections in each modality's different coordinate systems. Those skilled in the art will understand that the various sensors installed on the vehicle exist in the same space, and therefore each sensor (lidar, camera, millimeter-wave radar, and other similar sensors) has a corresponding projection relationship with each other. Therefore, it is important to emphasize that this invention does not project other modalities onto a single modality, but rather projects queries shared by different modalities back onto each modality.

[0048] It's also important to note that after feature sampling is performed for each modality, each modality's weights are generated by a corresponding MLP (Multi-Layer Perception) and normalized using a softmax operator. In short, the number of weights corresponds to the number of modalities, effectively assigning weights to each modality. This mechanism addresses the fact that the data source (sensor) for each modality is located in a different position and has different perceptual angles. This can cause a modality's view to be partially obstructed. Therefore, when a shared query is projected back onto that modality, the extracted features are null. Therefore, weights must be estimated for the features sampled from each modality. Finally, the features from multiple modalities are fused based on the weighted ratios, completing the query fusion process.

[0049] In this invention, each perceptual modality shares a single query. A multi-dimensional search in 3D space or time and space is performed under this shared query. The searched location is projected back to each modality, its features are sampled, and the multiple modalities are fused based on learnable weights. Finally, combined with conventional sparse algorithm operations, different levels of fusion are performed. Furthermore, each reference point generates multiple key points of the target object. Key point feature fusion, multi-level fusion (such as Hierarchy Fusion), and multi-head fusion at a single reference point are all conventional operations of sparse algorithms. After fusion at various levels, each reference point (also known as a query) ultimately generates fused instance / candidate features, and the number of queries corresponds to the number of instances.

[0050] In combination with the above-mentioned specific embodiments, the present invention utilizes only a single deformable transformer for cross attention, allowing multiple modalities to share a query. This makes model inference more efficient than other instance fusion solutions. Furthermore, compared to the previously mentioned BEV algorithm, which requires full-domain projection, the present invention employs a sparse algorithm, capturing only key features for projection, further reducing computational power. This method has proven superior accuracy to the BEV algorithm, effectively addressing the massive computational overhead of multimodal fusion while also improving accuracy.

[0051] In summary, the main design concept of the present invention is to abandon the dense algorithm and adopt multimodal fusion based on sparse features. Specifically, the feature maps of different modal perception data on the vehicle are obtained in advance; based on the instance features of the target point in the perception data, a reference point representing a target object shared by each modality is obtained; multiple derivative points corresponding to the reference point are generated, and the position information of the reference point and its derivative points are projected into the respective coordinate systems of the different modal perception data; based on the coordinate system of the reference point and its derivative points in each modality, the corresponding feature maps are sampled respectively; according to the source of the different modal perception data, the weights of each modal feature are configured and fusion is performed. After the present invention projects the reference point shared by different modalities back to each modality, feature sampling is performed separately and multimodal data fusion is completed. This method can effectively solve the problem of huge multimodal fusion computing power and has the effect of improving accuracy.

[0052] If the expressions expressing directions are mentioned in the embodiments of the present invention, they are relative concepts based on the embodiments. In addition, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a, b and c, where a, b, c can be single or multiple.

[0053] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings, but the above is only a preferred embodiment of the present invention. It should be noted that the technical features involved in the above embodiments and their preferred modes can be reasonably combined and matched into a variety of equivalent schemes by those skilled in the art without departing from or changing the design ideas and technical effects of the present invention; therefore, the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.

Claims

1. A method for fusion of multimodal perception data for autonomous driving, characterized in that: include: Pre-acquire feature maps of different modal perception data on the vehicle; Based on the instance features of the target point in the perception data, a reference point representing a target object shared by all modalities is obtained; generating a plurality of derivative points corresponding to the reference point; Projecting the position information of the reference point and its derivative points into respective coordinate systems of different modal perception data; Based on the coordinate systems of the reference point and its derivative points in each mode, feature sampling is performed on the corresponding feature maps respectively; According to the sources of different modal perception data, weights are configured for each modal feature and fusion is performed.

2. The method for fusion of multimodal perception data for autonomous driving according to claim 1, characterized in that: Multi-position compensation is performed on each of the reference points to obtain multiple dynamic positions to represent derived points, and the derived points corresponding to each of the reference points are used as key points of the target object.

3. The method for fusion of multimodal perception data for autonomous driving according to claim 1, characterized in that: After projecting a reference point and its derivative points to the geometric space corresponding to each mode, the position information in the geometric space of each mode is extracted respectively and the feature vector is extracted.

4. The method for fusion of multimodal perception data for autonomous driving according to claim 1, characterized in that: The fusion method further includes: before generating the derived points, each reference point is subjected to feature matching clustering via a self-attention mechanism.

5. The method for fusion of multimodal perception data for autonomous driving according to any one of claims 1 to 4, characterized in that: The sources of the modal perception data include at least the following two: lidar, camera, and millimeter-wave radar.

6. The method for fusion of multimodal perception data for autonomous driving according to claim 5, characterized in that: The coordinate system includes: a 3D Cartesian coordinate system, an image coordinate system, and a 2D top-view Cartesian coordinate system.