一种双模态语音情感识别方法及系统
By combining voice and text information in a dual-modal recognition method, the problem of information limitation in single-modal recognition is solved, and more accurate emotion recognition is achieved, which is applicable to fields such as emotional intelligence assistants and intelligent customer service.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FOURTH MILITARY MEDICAL UNIVERSITY
- Filing Date
- 2024-02-07
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, speech emotion recognition methods rely on data from a single modality, which leads to information limitations, fails to fully extract the speaker's emotional information, and has insufficient representational capabilities for the extracted acoustic and spectral features.
A dual-modal speech emotion recognition method is adopted. By combining speech and text information, high-level features are extracted using a pre-trained model, and temporal information is added through a self-attention mechanism and a long short-term memory network. Finally, a modal fusion algorithm is used to fuse emotion features.
It improves the accuracy and robustness of emotion recognition, reduces reliance on labeled data, enhances the scalability of the method, and provides a more accurate user experience and decision support.
Smart Images

Figure CN118038901B_ABST