A multi-modal sentiment recognition method and device based on agent distillation
By employing a surrogate distillation-based multimodal emotion recognition method, this approach utilizes self-attention and cross-modal attention mechanisms to extract multimodal semantic features. Combined with a soft-label filtering distillation model, it addresses the challenges of data labeling and information conflict in multimodal emotion recognition, achieving more accurate and stable emotion recognition.
CN122196724APending Publication Date: 2026-06-12GUANGDONG POLYTECHNIC NORMAL UNIV +2
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG POLYTECHNIC NORMAL UNIV
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-12
Smart Images

Figure CN122196724A_ABST
Abstract
The application discloses a kind of multi-modal sentiment recognition method and device based on agent distillation, comprising: S1. based on multi-modal emotional video, extract multi-modal emotional features and convert into the same dimension multi-modal semantic feature vector, including image mode, text mode and audio mode;S2. each modal semantic feature vector is learned through self-attention mechanism Context, while using the rest two modal semantic vectors through cross-modal attention operation Feature enhancement;S3. based on the enhanced multi-modal semantic feature vector, use cross-modal double agent model, learn the distillation label in mode and the soft label filtering distillation learning between modes;S4. after the weighted fusion of the distilled multi-modal semantic vector Linear mapping is the dimension of emotional type, after Softmax, the dimension with the largest value is taken as the corresponding emotional type;The application can stably and effectively capture complex human emotion, improve the accuracy of emotion recognition.
Need to check novelty before this filing date? Find Prior Art