A Multimodal Data Scene Recognition Method Based on Multi-level Interactive Fusion

By employing a multi-level interactive fusion method, non-visual data and visual data are fused in a multi-modal manner. This multi-level interactive fusion method for scene recognition solves the problem of insufficient visual data in autonomous driving and improves the accuracy and speed of scene recognition.

CN115878983BActive Publication Date: 2025-10-28HANGZHOU DIANZI UNIV
0 Cites 0 Cited by

Patent Information

Application Number
CN202211597492.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-10-28
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

In existing autonomous driving technologies, scene recognition mainly relies on visual data, lacking the supplementation of non-visual data and multi-level fusion, resulting in insufficient recognition accuracy and speed.

Method used

This paper proposes a multi-level interactive fusion method to integrate non-visual and visual data in a multimodal manner. Non-visual data is used to assist decision-making. The method adopts a multi-modal data scene recognition method based on multi-level interactive fusion, including feedforward neural network, two-stage attention mechanism, multi-layer spatiotemporal attention network and self-attention mechanism, to extract 2D and 3D features, and optimize feature interaction through contrastive learning loss.

Benefits of technology

It improves the accuracy and speed of autonomous driving scene recognition, reduces information redundancy, and enhances the ability to comprehensively describe scenes.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention discloses a multimodal data scene recognition method based on multi-level interactive fusion. Using video data and vehicle data collected by onboard sensors in autonomous driving scenarios, three single-modal features are extracted: 2D-level features from the video are obtained through multi-instance learning based on a two-stage attention mechanism; 3D spatiotemporal features from the scene video are extracted through a multi-layer spatiotemporal attention network, and onboard information feature vectors are added for training and interaction; and the onboard information feature vectors are trained. After feature extraction of the three modalities, similarity loss is calculated. During training, the similarity of the three modalities is maximized, and the features of the three modalities are interacted based on a multi-layer self-attention network. Finally, classification is performed. This method can utilize existing video and onboard information interaction to supplement information and improve the speed and accuracy of scene recognition.
Need to check novelty before this filing date? Find Prior Art