基于跨模态动态卷积的视频多模态情感识别方法、装置及计算机设备
By combining cross-modal dynamic convolution and multi-head attention mechanisms, the problem of modal information weight adjustment and fusion in multimodal sentiment analysis is solved, and more accurate sentiment recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2022-01-20
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multimodal sentiment analysis technologies struggle to effectively combine audio and image modal information to dynamically adjust the weight of text information. Furthermore, existing fusion strategies are ill-equipped to handle intermodal conflicts and redundant information, impacting the accuracy of sentiment recognition.
We employ a cross-modal dynamic convolution-based approach to acquire primary features, word-level alignment features, and high-level features from the video. We then utilize a bidirectional GRU network and cross-modal dynamic convolution for multimodal interaction, and combine this with a multi-head attention mechanism for feature fusion to output emotion recognition results.
It improves the accuracy of sentiment classification, effectively models the local information of the temporal dimension of modality, avoids important information being buried, and enhances the effect of sentiment recognition.
Smart Images

Figure CN114511906B_ABST