Multi-modal fusion method based on cross self-attention and feature induction
Through the multimodal fusion method of cross-self attention and feature induction, the problem of limited object detection performance in complex scenarios is solved, and the improvement of small object detection and model convergence acceleration are achieved.
Patent Information
- Application Number
- CN202510425692.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing multimodal visual fusion method is difficult to capture the context information of long-distance small targets in complex scenarios, and the Transformer converges slowly, which requires high data set size, resulting in limited object detection performance.
A multimodal fusion method based on cross-modal self-attention and feature induction is adopted, and feature interaction and global learning are performed through cross-modal cross-attention mechanism and self-attention mechanism, combined with splicing convolution operations, feature association and global modeling are enhanced, and the semantic understanding ability of the model is improved.
It significantly improves the integrity and accuracy of object detection in complex traffic scenarios, especially small object detection performance, and accelerates the convergence speed of the model.
Smart Images

Figure CN120339771A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multimodal fusion method based on cross self-attention and feature induction, belonging to the technical field of image processing. Background Art
[0002] The success of convolutional neural networks (CNNs) has promoted the application of deep learning in the field of object detection, significantly improving the accuracy and robustness of detection. However, in the presence of lighting changes and complex backgrounds, the performance of unimodal methods is limited. To meet the requirements of complex scene perception, multimodal visual object detection has become a research hotspot, aiming to enhance the perception ability by fusing multi-sensor information. The core of multimodal object detection lies in the effective fusion of different modal data. Multimodal feature fusion faces two major challenges: (1) Spatial association (feature space registration): Different modal data need to be mapped to a unified space to achieve spatial alignment. (2) Feature association (multimodal feature interaction and remapping): Through modal feature interaction and integration, multi-sensor information should be fully utilized to ensure efficient feature combination and information complementarity. Mainstream multimodal fusion networks are usually based on CNN or Transformer architectures. BEVFusion maps image and LiDAR features to the bird's-eye view (BEV) space to achieve high-precision spatial association through coordinate system transformation, and its fusion network uses CNN for local feature fusion. However, the receptive field of CNN is limited, lacking global modeling ability, making it difficult to capture the context information of small targets at long distances and affecting complex scene perception. TransFusion uses Transformer to fuse image and LiDAR features globally, extracting key fusion features through adaptive weight assignment, showing excellent performance under conditions such as occlusion and lighting changes, demonstrating the global feature association, interaction, and perception ability of Transformer. However, Transformer has a slow convergence rate and high requirements for the scale of the dataset. In the task of general road object detection, the diversity and complexity of visual information pose challenges to unimodal perception. Although multimodal visual fusion-based solutions can improve detection performance, existing feature fusion methods still have bottlenecks. Summary of the Invention
[0003] The purpose of the present invention is to provide a multimodal fusion method based on cross self-attention and feature induction, which processes road objects through multimodal feature interaction, captures deep interactions between modalities, enhances feature association, and performs global learning on the fused features, thereby significantly improving the integrity and accuracy of object detection in complex traffic scenes.
[0004] To achieve the above object / to solve the above technical problems, the present invention is implemented by the following technical solutions.
[0005] On the one hand, the present invention provides a multi-modal fusion method based on cross self-attention and feature induction, comprising the following steps: Obtain the LiDAR feature data and image feature data of road targets, and preprocess the LiDAR feature data; Parallelly input the preprocessed LiDAR feature data and image feature data into a pre-trained fusion network model to obtain road target detection results; Among them, the processing process of the pre-trained fusion network model is as follows: Perform multi-modal interaction processing on the preprocessed LiDAR feature data and image feature data through a cross-modal cross-attention mechanism to respectively obtain image enhanced features and LiDAR enhanced features, so that the image feature data fuses the spatial position information in the LiDAR feature data, and the LiDAR feature data fuses the rich semantic information of the image features, Introduce the inductive bias information of the LiDAR feature data and image feature data into the LiDAR enhanced features and image enhanced features to respectively obtain LiDAR balanced features and image balanced features; Perform a concatenation operation on the LiDAR balanced features and image balanced features in the channel dimension to obtain fused features; Perform global feature learning processing on the fused features to obtain globally modeled features; Perform a matrix addition operation on the concatenated features of the LiDAR feature data and image feature data and the globally modeled features to obtain final fused features, and perform target detection on the final fused features.
[0006] Further, the preprocessing of the LiDAR feature data is specifically: perform 1x1 convolution processing on the LiDAR feature data to achieve compression of the LiDAR feature data in the channel dimension.
[0007] Further, the multi-modal interaction processing of the preprocessed LiDAR feature data and image feature data through a cross-modal cross-attention mechanism specifically includes: the LiDAR feature data to image feature data fusion stage and the image feature data to LiDAR feature data fusion stage; Among them, the LiDAR feature data to image feature data fusion stage is specifically: use the preprocessed LiDAR feature data as the query vector and the image feature data as the key vector and value vector to input into the cross-attention layer to obtain image enhanced features; The image feature data to LiDAR feature data fusion stage is specifically: use the image enhanced features as the query vector and the preprocessed LiDAR feature data as the key vector and value vector to input into the cross-attention layer to obtain LiDAR enhanced features.
[0008] Further, the preprocessed LiDAR feature data and image feature data are subjected to multi-modal interaction processing through a cross-modal cross-attention mechanism at least twice.
[0009] Further, the method for obtaining the inductive bias information of the LiDAR feature data and the image feature data specifically includes: The LiDAR feature data and the image feature data are concatenated in the channel dimension to obtain a first feature A, then the first feature A is subjected to a 3x3 convolution process to obtain a second feature B, then the second feature B is subjected to a 1x1 convolution process to obtain a third feature C, and finally the third feature C is subjected to a 1x1 convolution process to obtain a fourth feature D, and the fourth feature D is the inductive bias information.
[0010] Further, the process of performing global feature learning on the fused feature to obtain a globally modeled feature specifically includes: The fused feature is respectively mapped into a query vector, a key vector, and a value vector, and then global modeling is performed through a self-attention mechanism to capture the dependencies between features within the global range. The self-attention mechanism calculates the similarity between the query vector and the key vector to obtain attention weights, enabling each feature point to adaptively receive information from all other feature points. The attention weights and the value vector are weighted and summed to obtain a globally modeled feature, realizing the global modeling of the fused feature.
[0011] Further, the method for obtaining the concatenated feature of the LiDAR feature data and the image feature data specifically includes: The LiDAR feature data and the image feature data are concatenated in the channel dimension to obtain a first feature A, then the first feature A is subjected to a 3x3 convolution process to obtain a second feature B, then the second feature B is subjected to a 1x1 convolution process to obtain a third feature C, and the fourth feature C is the concatenated feature.
[0012] In a second aspect, the present invention provides a multi-modal fusion road target detection device, including: An acquisition module, configured to acquire LiDAR feature data and image feature data of a road target, and preprocess the LiDAR feature data; A detection module, configured to parallelly input the preprocessed LiDAR feature data and image feature data into a pre-trained fusion network model to obtain a road target detection result; Wherein, the processing process of the pre-trained fusion network model is as follows: The preprocessed LiDAR feature data and image feature data are subjected to multimodal interaction processing through a cross-modal cross-attention mechanism to obtain image-enhanced features and LiDAR-enhanced features respectively, so that the image feature data can fuse the spatial position information in the LiDAR feature data, and the LiDAR feature data can fuse the rich semantic information of the image features. The inductive bias information of the LiDAR feature data and the image feature data is introduced into the LiDAR-enhanced features and the image-enhanced features to obtain LiDAR-balanced features and image-balanced features respectively; The LiDAR-balanced features and the image-balanced features are concatenated in the channel dimension to obtain fused features; The fused features are subjected to global feature learning processing to obtain global modeling features; The concatenated features of the LiDAR feature data and the image feature data and the global modeling features are subjected to matrix addition operation processing to obtain final fused features, and the final fused features are used for object detection.
[0013] In a third aspect, the present invention provides a computer system, including: A memory for storing computer programs / instructions; A processor for executing the computer programs / instructions to implement the steps of the above-mentioned multimodal fusion method based on cross self-attention and feature induction.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned multimodal fusion method based on cross self-attention and feature induction are implemented.
[0015] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention proposes a multimodal fusion method based on cross self-attention and feature induction for road general object detection tasks. By using the cross-attention mechanism, the present invention realizes deep interaction between LiDAR and image features, enhances the spatial position and semantic expression of modal features, and bridges the semantic differences between modalities; through the self-attention mechanism, global modeling is performed on the fused features to capture the long-range dependence relationship between features, improve the global semantic understanding and reasoning ability of the model, and introduce inductive bias information through concatenated convolution operations to make up for the deficiencies of Transformer in processing absolute position information and local details, balance local and global modeling, accelerate model convergence, and optimize the learning effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flow schematic diagram of the present invention; Figure 2 It is a schematic diagram of the processing process of the fusion network model of the present invention; Figure 3 Schematic diagram of the influence of introducing inductive bias information on the convergence speed of the network model; Figure 4 Aerial view of the ground truth in a complex traffic scenario; Figure 5 Aerial view of the prediction results of the baseline network in a complex traffic scenario; Figure 6 Aerial view of the prediction results of the fusion network model of the present invention in a complex traffic scenario; Figure 7 Aerial view of the ground truth in a small target detection scenario; Figure 8 Aerial view of the prediction results of the baseline network in a small target detection scenario; Figure 9 Aerial view of the prediction results of the fusion network model of the present invention in a small target detection scenario; Figure 10 Aerial view of the ground truth in a complex occlusion scenario; Figure 11 Aerial view of the prediction results of the baseline network in a complex occlusion scenario; Figure 12 Aerial view of the prediction results of the fusion network model of the present invention in a complex occlusion scenario. Detailed implementation manners
[0017] It should be noted that: The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0018] The term "and / or" merely describes the associated relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after. Embodiment
[0019] As Figures 1 - 2 shown, this embodiment provides a multi-modal fusion method based on cross self-attention and feature induction, including the following steps: Step 1: Obtain the LiDAR feature data F L of the road target and the image feature data F C , and preprocess the LiDAR feature data F L . The preprocessing of the LiDAR feature data F LPerform preprocessing, specifically: for the LiDAR feature data F L Perform compression in the channel dimension; Step 2: Input the preprocessed LiDAR feature data F ~ L and the image feature data F C into the pre-trained fusion network model in parallel to obtain the road target detection result; Among them, the processing process of the pre-trained fusion network model is as follows: Step 2.1: Input the preprocessed LiDAR feature data F ~ L and the image feature data F C through the cross-modal cross-attention mechanism for two multi-modal interaction processes to obtain the image enhanced feature and the LiDAR enhanced feature respectively, so that the image feature data fuses the spatial position information in the LiDAR feature data, and the LiDAR feature data fuses the rich semantic information of the image feature. Specifically: Step 2.11: The first multi-modal interaction process: The LiDAR feature data to image feature data fusion stage is specifically: using the preprocessed LiDAR feature data F ~ L as the query vector and the image feature data F C as the key vector and value vector to input into the cross-attention layer to obtain the first image enhanced feature F; Based on the image features, according to the degree of association between each LiDAR feature data and the image features, the image features are selectively aggregated, which is equivalent to re-weighting the image features with the LiDAR features, meaning the incorporation of LiDAR information, and generating the first image enhanced feature F.
[0020] The image feature data to LiDAR feature data fusion stage is specifically: using the first image enhanced feature F as the query vector and the preprocessed LiDAR feature data F ~ L as the key vector and value vector to input into the cross-attention layer to obtain the first LiDAR enhanced feature G; Based on the LiDAR features, according to the degree of association between each LiDAR feature data and the image enhanced features, the LiDAR features themselves are selectively aggregated, and the image information is incorporated into them, thus obtaining the first LiDAR enhanced feature G; Step 2.12: Introduce the LiDAR feature data F L and the image feature data F CThe inductive bias information is used to obtain the first LiDAR balance feature I and the first image balance feature H respectively, specifically as follows: The LiDAR feature data F L is concatenated with the image feature data Fc in the channel dimension to obtain the first feature A. Then, the first feature A is processed by a 3x3 convolution to obtain the second feature B. Then, the second feature B is processed by a 1x1 convolution to obtain the third feature C. Finally, the third feature C is processed by a 1x1 convolution to obtain the fourth feature D, and the fourth feature D is the first inductive bias information; The first image enhancement feature F and the fourth feature D are subjected to a matrix addition operation, and the inductive bias information regarding the LiDAR and image fusion features is added to the first image enhancement feature F to obtain the first image balance feature H; The first LiDAR enhancement feature G and the fourth feature D are subjected to a matrix addition operation, and the inductive bias information regarding the LiDAR and image fusion features in the fourth feature D is added to the first LiDAR enhancement feature G to obtain the first LiDAR balance feature I; Step 2.13: Second multi-modal interaction processing: The LiDAR feature data to image feature data fusion stage is specifically as follows: The first LiDAR balance feature I is used as the query vector, and the first image balance feature H is used as the key vector and value vector and input into the cross-attention layer to obtain the second image enhancement feature J; The image feature data to LiDAR feature data fusion stage is specifically as follows: The second image enhancement feature J is used as the query vector, and the first LiDAR balance feature I is used as the key vector and value vector and input into the cross-attention layer to obtain the second LiDAR enhancement feature K; Step 2.14: The fourth feature D is processed by a 1x1 convolution to obtain the fifth feature E, which is the second inductive bias information, The second image enhancement feature J and the fifth feature E are subjected to a matrix addition operation, and the inductive bias information regarding the LiDAR and image fusion features is added to the second image enhancement feature J to obtain the second image balance feature L; The second LiDAR enhancement feature K and the fifth feature E are subjected to a matrix addition operation, and the second inductive bias information regarding the LiDAR and image fusion features is added to the second LiDAR enhancement feature K to obtain the second LiDAR balance feature M; Step 2.2: The second LiDAR balance feature M and the second image balance feature L are concatenated in the channel dimension to obtain the fusion feature N; Step 2.3: The fusion feature N is subjected to global feature learning processing to obtain the global modeling feature O, specifically as follows: The fused feature N is separately mapped into a query vector, a key vector, and a value vector, and then global modeling is performed through a self-attention mechanism to capture the dependencies between features within the global scope. The self-attention mechanism calculates the similarity between the query vector and the key vector to obtain attention weights, enabling each feature point to adaptively receive information from all other feature points. The attention weights and the value vector are weighted and summed to obtain the globally modeled feature, achieving global modeling of the fused feature and obtaining the globally modeled feature O.
[0021] Step 2.4: Perform a matrix addition operation on the concatenated feature C of the LiDAR feature data F L and the image feature data F C and the globally modeled feature O to obtain the road target detection result, specifically: Perform a concatenation operation on the LiDAR feature data F L and the image feature data F C in the channel dimension to obtain the first feature A, then perform a 3x3 convolution operation on the first feature A to obtain the second feature B, and then perform a 1x1 convolution operation on the second feature B to obtain the third feature C. The fourth feature C is the concatenated feature. Then, the concatenated feature C and the globally modeled feature O are subjected to a matrix addition operation to obtain the final feature P.
[0022] The above method for training the fusion network model: The nuScenes dataset is used for experimental evaluation. nuScenes is a multi-modal dataset widely used in the field of autonomous driving, containing synchronous data from 6 cameras, 1 lidar, 5 radars, IMU, and GPS; the dataset covers 1000 urban scenes, each scene lasting 20 - 40 seconds, with a total capacity of approximately 130GB, containing approximately 1 million images and 2 million frames of LiDAR point cloud data, and providing detailed object annotations and trajectory information for training and evaluation of perception tasks such as object detection, tracking, and semantic segmentation.
[0023] During the training process, the resolution of the BEV feature map is unified to 90x90, the total number of training epochs is set to 10, and the Cyclic learning rate scheduling strategy is adopted. The initial learning rate is set to 0.0001, and the maximum learning rate is set to 0.0002. The Cyclic learning rate strategy dynamically adjusts the learning rate.
[0024] To verify the effectiveness of the proposed fusion network model for multi-modal features, first, under the condition of a 90x90-sized BEV feature map, it is compared with the BEVFusion baseline network, and the results are shown in Table 1: ;
[0025] The results show that the improved fusion network (Ours) outperforms the original BEVFusion in both mAP and NDS. The mAP is increased by about 3% and the NDS is increased by about 2.6%. The detection performance of all categories has been improved, especially for small objects such as pedestrians (Ped.), bicycles (Bike), and traffic cones (T.C.), with the improvement rate approaching or exceeding 5%. This indicates that the improved network has significant advantages in dealing with small object detection tasks.
[0026] To further verify the superiority of the improved network and make a fair comparison with other advanced methods, experiments were conducted under the condition of BEV feature maps with a size of 180x180. Since the improved fusion network adopts the Transformer architecture, the computational cost of directly processing the original resolution feature maps is too high. Therefore, a Feature Pyramid Network (FPN) structure is introduced in the network output stage to integrate the original size feature maps into the FPN to achieve performance comparison with other methods at the same scale. The results are shown in Table 2: ;
[0027] As can be seen from Table 2, the result of BEVFusion (baseline) (mAP is 65.14%) is significantly higher than that of BEVFusion in Table 1 (mAP is 61.34%). This difference is mainly caused by the different sizes of the input feature maps. The improved network in this embodiment still achieves the best performance, further verifying its effectiveness and robustness.
[0028] To deeply analyze the contributions of cross-modal feature interaction processing (CAM), global feature learning processing (SAM), and feature induction processing (FIM) to the overall performance, a series of ablation experiments were conducted. The performance under different module combinations is shown in Table 3 below: ;
[0029] CAM analysis: When using CAM alone (the first row), the performance is slightly lower than that of the baseline model (BEVFusion, Table 1). This indicates that although CAM enhances the cross-modal feature interaction, it lacks the ability of global modeling.
[0030] SAM analysis: When using SAM alone (the third row), the performance is better than that of using only CAM and is improved compared with the baseline model. This shows that global feature modeling is crucial for multi-modal fusion.
[0031] The synergistic effect of CAM and SAM: When using CAM and SAM simultaneously (the second row), the performance is significantly better than using CAM or SAM alone, indicating their complementarity. CAM is responsible for local cross-modal feature interaction, and SAM is responsible for global feature modeling. The combination of the two can achieve more effective multi-modal feature fusion.
[0032] FIM analysis: Only the combination of FIM and SAM (the third row) exceeds the combination of CAM and SAM in terms of mAP and NDS metrics, indicating that the inductive bias information provided by FIM can be effectively utilized by SAM to improve performance.
[0033] When CAM, SAM, and FIM are used simultaneously (the last row), the model performance is the best.
[0034] To further explore the impact of the number of attention layers in CAM and SAM on model performance, additional ablation experiments were conducted, and the results are shown in Table 4: ;
[0035] The experimental results show that when CAM uses 2 layers and SAM uses 2 layers, i.e., (2,2), the performance of the fusion network model reaches the best; when the number of layers of SAM is increased to 3 layers, the performance decreases instead, which may be due to model overfitting or optimization difficulties; when the number of layers of CAM is increased to 4 layers, the performance does not show significant improvement, indicating that 2 layers of CA are sufficient to capture the interaction relationships between modalities; therefore, we finally choose the configuration with 2 layers of CAM and 2 layers of SAM.
[0036] As Figure 3 shown, the impact of the introduction of inductive bias information (FIM) on the convergence speed of the fusion network model; it can be seen that after the introduction of inductive bias information, the decline rate of the model loss value is significantly improved, and this phenomenon of accelerated convergence is mainly attributed to the internal inductive bias modeling mechanism; specifically, this mechanism improves the training efficiency in the following ways: 1) Adaptive learning guidance: The inductive bias modeling mechanism can adaptively guide the parameter update direction during training, providing a more explicit learning signal for the model; 2) Reducing ineffective fluctuations: Different from the ineffective fluctuations that easily occur in traditional Transformer models at the beginning of training, the prior learning framework provided by the inductive bias modeling mechanism can effectively calibrate the update path of network parameters and reduce unnecessary oscillations. 3) Efficient gradient propagation: Since the parameter update direction is more explicit and the path is more stable, the gradient propagation process of the model becomes more efficient, thus accelerating the overall convergence speed.
[0037] In summary, by introducing inductive bias, an effective prior knowledge is provided for the Transformer architecture, enabling it to be optimized rapidly and stably towards a better solution at the beginning of training, and significantly improving the convergence speed of the model.
[0038] To visually evaluate the detection performance of the proposed fusion network (Ours) in this paper, three typical scenarios are selected, and the prediction results of Ours are compared and analyzed with those of the baseline model fusion network (Baseline, i.e., BEVFusion); by comparing the prediction results of the two fusion networks with the ground truth boxes from the BEV perspective, the problems existing in the prediction results of different models are observed, so as to highlight the performance differences between the models.
[0039] As Figures 4 - 6 shown, the comparison chart of prediction results in complex traffic scenarios: It can be seen that Figure 5 compared with Figure 6 Through the differential annotation of the bird's-eye view, the object detection performance of the baseline network and the network proposed in this paper in complex traffic scenarios is visually compared; the results show that the Ours network is superior to the baseline network in terms of detection integrity and accuracy; specifically, the Ours network effectively reduces the missed detection rate and successfully detects the objects that the baseline network fails to recognize, which is mainly attributed to the following advantages of the Transformer architecture: 1) Modal feature correlation fusion: Transformer can effectively fuse the features of two modalities, i.e., images and LiDAR point clouds, make up for the deficiencies of single-modal information, and enhance the perception ability of objects. Especially for small objects at long distances, the fusion of multi-modal information is more critical; 2) Global context modeling: The self-attention mechanism of Transformer can capture global context information, enhance the model's understanding of the mutual relationships between objects in complex traffic scenarios, and thus more accurately identify objects.
[0040] However, the Ours network has misdetected small objects in some occluded areas; this phenomenon may be caused by the following factors: 1) Feature artifacts caused by occlusion: Occluders may cause sparse and irregular reflections in LiDAR point clouds, or form textures and edges similar to objects in images. These "artifact" features may be misrecognized as small objects; 2) Feature extraction and fusion errors: During the feature extraction process of the image and LiDAR branches, incorrect feature representations may be generated for occluded areas, and these errors are amplified in the multi-modal fusion stage, resulting in the model having hallucinations and misjudging the existence of small objects.
[0041] Generally speaking, the Ours network significantly improves the object detection integrity and accuracy in complex traffic scenarios through the modal feature fusion and global modeling capabilities of the Transformer architecture, but there is still room for improvement in dealing with feature artifacts and misjudgments caused by occlusion.
[0042] As Figures 7 - 9 shown, the comparison chart of prediction results in small object detection scenarios: It can be seen that: Figure 8 and Figure 9The detection results of the baseline network and the network proposed in this paper are respectively shown in the scenario focusing on the detection performance of small targets; the comparative analysis shows that Ours is significantly better than the baseline network in small target detection, and can effectively detect multiple small targets missed by the baseline network, whether they are at a long distance or occluded; this advantage is mainly due to the self-attention mechanism in the Transformer architecture, which is specifically reflected in the following two aspects: 1) Global semantic association: The features of small targets are usually sparsely distributed in the image and are easily interfered by background noise; the self-attention mechanism can establish global semantic connections between small targets and their surrounding environments (such as roads, adjacent targets, etc.), thereby enhancing the feature expression of small targets and effectively suppressing background noise; 2) Cross-modal feature alignment: The cross-modal attention mechanism can adaptively align the features of small targets from different sensors (such as cameras and LiDAR), thereby improving the robustness and consistency of the features, and further improving the performance of small target detection.
[0043] Such as Figures 10 - 12 shown, the comparison chart of prediction results in complex occlusion scenarios: It can be seen that Figure 11 and Figure 12 respectively show the detection results of the baseline network and the network proposed in this paper in complex occlusion scenarios; the comparative analysis shows that the network designed in this paper performs better than the baseline network in complex occlusion scenarios by using global context information and multi-modal information complementarity, and can improve the missed detection phenomenon to a certain extent; however, due to reasons such as too high occlusion degree and insufficient feature information, there are still cases of missed detection of occluded targets.
[0044] Based on the visual comparative analysis of the above three typical scenarios, it fully verifies the improvement of the fusion network model (Ours) proposed in this paper for the target detection performance compared with the baseline network in the case of small targets and occlusion conditions; the experiment shows that by introducing the Transformer architecture, Ours effectively enhances the interaction and fusion between modal features, improves the global semantic understanding ability of the model, and thus achieves more accurate and robust target detection in complex scenarios. Embodiment
[0045] This embodiment provides a multi-modal fusion road target detection device, including: An acquisition module, configured to acquire LiDAR feature data and image feature data of road targets, and preprocess the LiDAR feature data; A detection module, configured to parallelly input the preprocessed LiDAR feature data and image feature data into a pre-trained fusion network model to obtain road target detection results; Wherein, the processing process of the pre-trained fusion network model is: The preprocessed LiDAR feature data and image feature data are subjected to multimodal interaction processing through a cross-modal cross-attention mechanism to obtain image enhanced features and LiDAR enhanced features respectively, so that the image feature data can fuse the spatial position information in the LiDAR feature data, and the LiDAR feature data can fuse the rich semantic information of the image features. The inductive bias information of the LiDAR feature data and the image feature data is introduced into the image enhanced features and the LiDAR enhanced features to obtain LiDAR balanced features and image balanced features respectively; The LiDAR balanced features and the image balanced features are concatenated in the channel dimension to obtain fused features; The fused features are subjected to global feature learning processing to obtain globally modeled features; The concatenated features of the LiDAR feature data and the image feature data and the globally modeled features are subjected to matrix multiplication operation processing to obtain road target detection results. Embodiment
[0046] This embodiment provides a computer system, including: A memory for storing computer programs / instructions; A processor for executing the computer programs / instructions to implement the steps of the above multimodal fusion method based on cross self-attention and feature induction. Embodiment
[0047] This embodiment provides a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the above multimodal fusion method based on cross self-attention and feature induction are implemented.
[0048] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0049] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0050] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0052] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims. All of these fall within the protection scope of the present invention.
Claims
1. A multimodal fusion method based on cross self-attention and feature induction, characterized in that It includes the following steps: Obtain the LiDAR feature data and image feature data of the road target, and preprocess the LiDAR feature data; Parallelly input the preprocessed LiDAR feature data and image feature data into a pre-trained fusion network model to obtain the road target detection result; Among them, the processing process of the pre-trained fusion network model is: Perform multi-modal interaction processing on the preprocessed LiDAR feature data and image feature data through a cross-modal cross-attention mechanism to respectively obtain image enhanced features and LiDAR enhanced features, so that the image feature data fuses the spatial position information in the LiDAR feature data, and the LiDAR feature data fuses the rich semantic information of the image features; Introduce the inductive bias information of the LiDAR feature data and image feature data into the LiDAR enhanced features and image enhanced features to respectively obtain LiDAR balanced features and image balanced features; Perform a concatenation operation on the LiDAR balanced features and image balanced features in the channel dimension to obtain fused features; Perform global feature learning processing on the fused features to obtain global modeling features; Perform a matrix addition operation on the concatenated features of the LiDAR feature data and image feature data and the global modeling features to obtain the final fused features, and perform target detection on the final fused features.
2. The multimodal fusion method based on cross self-attention and feature induction according to claim 1, wherein The preprocessing of the LiDAR feature data is specifically: Perform 1x1 convolution processing on the LiDAR feature data to achieve compression of the LiDAR feature data in the channel dimension.
3. The multimodal fusion method based on cross self-attention and feature induction according to claim 1, wherein The multi-modal interaction processing of the preprocessed LiDAR feature data and image feature data through a cross-modal cross-attention mechanism specifically includes: the LiDAR feature data to image feature data fusion stage and the image feature data to LiDAR feature data fusion stage; Among them, the LiDAR feature data to image feature data fusion stage is specifically: use the preprocessed LiDAR feature data as the query vector, and the image feature data as the key vector and value vector to input into the cross-attention layer to obtain image enhanced features; The image feature data to LiDAR feature data fusion stage is specifically: use the image enhanced features as the query vector, and the preprocessed LiDAR feature data as the key vector and value vector to input into the cross-attention layer to obtain LiDAR enhanced features.
4. The multimodal fusion method based on cross self-attention and feature induction according to claim 1, characterized in that The multi-modal interaction processing of the preprocessed LiDAR feature data and image feature data through a cross-modal cross-attention mechanism is set at least twice.
5. The multimodal fusion method based on cross self-attention and feature induction according to claim 3, wherein The method for obtaining the inductive bias information of the LiDAR feature data and image feature data specifically includes: Perform a concatenation operation on the LiDAR feature data and image feature data in the channel dimension to obtain the first feature A, then perform 3x3 convolution processing on the first feature A to obtain the second feature B, then perform 1x1 convolution processing on the second feature B to obtain the third feature C, and finally perform 1x1 convolution processing on the third feature C to obtain the fourth feature D, and the fourth feature D is the inductive bias information.
6. The multimodal fusion method based on cross self-attention and feature induction according to claim 1, wherein The global feature learning processing of the fused features to obtain global modeling features specifically includes: The fused features are respectively mapped into query vectors, key vectors, and value vectors, and then global modeling is performed through the self-attention mechanism. The self-attention mechanism calculates the similarity between the query vectors and key vectors to obtain attention weights, and performs weighted summation on the attention weights and value vectors to obtain global modeling features, realizing global modeling of the fused features.
7. The multimodal fusion method based on cross self-attention and feature induction according to claim 1, wherein The method for obtaining the concatenated features of the LiDAR feature data and the image feature data specifically includes: Performing a concatenation operation on the LiDAR feature data and the image feature data in the channel dimension to obtain a first feature A, then performing a 3x3 convolution operation on the first feature A to obtain a second feature B, and then performing a 1x1 convolution operation on the second feature B to obtain a third feature C. The fourth feature C is the concatenated feature.
8. A multi-modal fusion road target detection device, characterized in that, It includes: An acquisition module for acquiring the LiDAR feature data and the image feature data of the road target and preprocessing the LiDAR feature data; A detection module for parallelly inputting the preprocessed LiDAR feature data and the image feature data into a pre-trained fusion network model to obtain a road target detection result; Among them, the processing process of the pre-trained fusion network model is: Performing multi-modal interaction processing on the preprocessed LiDAR feature data and the image feature data through a cross-modal cross-attention mechanism to respectively obtain an image enhancement feature and a LiDAR enhancement feature, so that the image feature data fuses the spatial position information in the LiDAR feature data, and the LiDAR feature data fuses the rich semantic information of the image features. Introducing the inductive bias information of the LiDAR feature data and the image feature data into the LiDAR enhancement feature and the image enhancement feature to respectively obtain a LiDAR balanced feature and an image balanced feature; Performing a concatenation operation on the LiDAR balanced feature and the image balanced feature in the channel dimension to obtain a fused feature; Performing global feature learning processing on the fused feature to obtain a global modeling feature; Performing a matrix addition operation on the concatenated feature of the LiDAR feature data and the image feature data and the global modeling feature to obtain a final fused feature, and performing target detection on the final fused feature.
9. A computer system, characterized in that, It includes: A memory for storing computer programs / instructions; A processor for executing the computer programs / instructions to implement the steps of the multi-modal fusion method based on cross self-attention and feature induction according to any one of claims 1-7.
10. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer programs / instructions are executed by the processor, the steps of the multi-modal fusion method based on cross self-attention and feature induction according to any one of claims 1-7 are implemented.