基于改进视觉Transformer模型的语音特征识别方法及系统

By improving the P2T module and SparseTransformer network of the visual Transformer model, the problems of local information loss and redundant information in speech emotion recognition are solved, and more efficient speech emotion recognition is achieved.

CN116778912BActive Publication Date: 2026-07-17NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2023-07-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing visual Transformer models suffer from problems such as loss of local information frames and introduction of redundant information in speech emotion recognition, resulting in high computational complexity and high complexity of the correlation matrix.

Method used

An improved visual Transformer model is adopted, which extracts local features through the P2T module and performs sparse operations using the SparseTransformer network. The SMHA module is combined to enhance contextual information understanding and reduce computational complexity.

Benefits of technology

It effectively extracts local feature information of speech emotion, reduces computational complexity, and improves speech emotion recognition rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778912B_ABST
    Figure CN116778912B_ABST
Patent Text Reader

Abstract

本发明公开了基于改进视觉Transformer模型的语音特征识别方法及系统,涉及语音特征识别技术领域,方法包括以下步骤:接收原始语音信号,对原始语音信号进行预处理得到语音处理信号;对语音处理信号提取声学特征,得到log‑Mel语谱图;将log‑Mel语谱图输入至预先建立的P2T模块内,得到特征向量;将特征向量输入至预先建立的SparseTransformer网络内,得到输出结果;将输出结果导入预先建立的Softmax分类器后,得到识别结果。
Need to check novelty before this filing date? Find Prior Art