基于改进视觉Transformer模型的语音特征识别方法及系统
By improving the P2T module and SparseTransformer network of the visual Transformer model, the problems of local information loss and redundant information in speech emotion recognition are solved, and more efficient speech emotion recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF INFORMATION SCI & TECH
- Filing Date
- 2023-07-19
- Publication Date
- 2026-07-17
AI Technical Summary
Existing visual Transformer models suffer from problems such as loss of local information frames and introduction of redundant information in speech emotion recognition, resulting in high computational complexity and high complexity of the correlation matrix.
An improved visual Transformer model is adopted, which extracts local features through the P2T module and performs sparse operations using the SparseTransformer network. The SMHA module is combined to enhance contextual information understanding and reduce computational complexity.
It effectively extracts local feature information of speech emotion, reduces computational complexity, and improves speech emotion recognition rate.
Smart Images

Figure CN116778912B_ABST