语音情感识别方法、装置、设备以及存储介质

By constructing a neural network model and combining Mel spectrogram transformation and attention feature extraction, the problem of poor feature extraction in traditional speech emotion recognition methods is solved, and more efficient speech signal emotion recognition is achieved.

CN116312642BActive Publication Date: 2026-07-17SOUTH CHINA NORMAL UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA NORMAL UNIV
Filing Date
2023-03-01
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Traditional speech emotion recognition methods rely on handcrafted acoustic features, resulting in poor recognition performance and difficulty in effectively capturing emotion-related features.

Method used

A neural network model is constructed, including a time-frequency channel attention feature extraction module, a deep feature extraction module, an empirical feature extraction module, and a feature fusion module. Through Mel spectrogram transformation, channel attention extraction, and time-frequency attention extraction, combined with deep learning and human prior knowledge, detailed speech signal features are extracted and fused.

Benefits of technology

It improves the accuracy and efficiency of emotion recognition in speech signals, realizing the complementary advantages of human prior knowledge and deep learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312642B_ABST
    Figure CN116312642B_ABST
Patent Text Reader

Abstract

本发明涉及语音信号情感识别领域,特别涉及一种语音情感识别方法,包括:构建神经网络模型,获得待识别的语音信号,将语音信号输入至时频通道注意力特征提取模块中进行特征提取,获取时频通道注意力特征图;将时频通道注意力特征图输入至深度特征提取模块中进行特征提取,获得深度特征图;获得语音信号的音频特征图,将音频特征图输入至经验特征提取模块中进行特征提取,获得经验特征图;将深度特征图以及经验特征图输入至特征融合模块中进行特征融合,获得融合特征图;将融合特征图输入至情感识别模块中进行情感识别,获得情感识别结果。充分利用人类先验知识和深度学习的互补优势,提高了对语音信号的情感识别的精准度以及效率。
Need to check novelty before this filing date? Find Prior Art