语音情感识别方法、装置、设备以及存储介质
By constructing a neural network model and combining Mel spectrogram transformation and attention feature extraction, the problem of poor feature extraction in traditional speech emotion recognition methods is solved, and more efficient speech signal emotion recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTH CHINA NORMAL UNIV
- Filing Date
- 2023-03-01
- Publication Date
- 2026-07-17
AI Technical Summary
Traditional speech emotion recognition methods rely on handcrafted acoustic features, resulting in poor recognition performance and difficulty in effectively capturing emotion-related features.
A neural network model is constructed, including a time-frequency channel attention feature extraction module, a deep feature extraction module, an empirical feature extraction module, and a feature fusion module. Through Mel spectrogram transformation, channel attention extraction, and time-frequency attention extraction, combined with deep learning and human prior knowledge, detailed speech signal features are extracted and fused.
It improves the accuracy and efficiency of emotion recognition in speech signals, realizing the complementary advantages of human prior knowledge and deep learning.
Smart Images

Figure CN116312642B_ABST