A multi-modal sentiment recognition method and device based on agent distillation

By employing a surrogate distillation-based multimodal emotion recognition method, this approach utilizes self-attention and cross-modal attention mechanisms to extract multimodal semantic features. Combined with a soft-label filtering distillation model, it addresses the challenges of data labeling and information conflict in multimodal emotion recognition, achieving more accurate and stable emotion recognition.

CN122196724APending Publication Date: 2026-06-12GUANGDONG POLYTECHNIC NORMAL UNIV +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG POLYTECHNIC NORMAL UNIV
Filing Date
2026-01-30
Publication Date
2026-06-12

Smart Images

  • Figure CN122196724A_ABST
    Figure CN122196724A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal sentiment recognition method and device based on agent distillation, comprising: S1. based on multi-modal emotional video, extract multi-modal emotional features and convert into the same dimension multi-modal semantic feature vector, including image mode, text mode and audio mode;S2. each modal semantic feature vector is learned through self-attention mechanism Context, while using the rest two modal semantic vectors through cross-modal attention operation Feature enhancement;S3. based on the enhanced multi-modal semantic feature vector, use cross-modal double agent model, learn the distillation label in mode and the soft label filtering distillation learning between modes;S4. after the weighted fusion of the distilled multi-modal semantic vector Linear mapping is the dimension of emotional type, after Softmax, the dimension with the largest value is taken as the corresponding emotional type;The application can stably and effectively capture complex human emotion, improve the accuracy of emotion recognition.
Need to check novelty before this filing date? Find Prior Art