一种双模态语音情感识别方法及系统

By combining voice and text information in a dual-modal recognition method, the problem of information limitation in single-modal recognition is solved, and more accurate emotion recognition is achieved, which is applicable to fields such as emotional intelligence assistants and intelligent customer service.

CN118038901BActive Publication Date: 2026-07-17FOURTH MILITARY MEDICAL UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FOURTH MILITARY MEDICAL UNIVERSITY
Filing Date
2024-02-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, speech emotion recognition methods rely on data from a single modality, which leads to information limitations, fails to fully extract the speaker's emotional information, and has insufficient representational capabilities for the extracted acoustic and spectral features.

Method used

A dual-modal speech emotion recognition method is adopted. By combining speech and text information, high-level features are extracted using a pre-trained model, and temporal information is added through a self-attention mechanism and a long short-term memory network. Finally, a modal fusion algorithm is used to fuse emotion features.

Benefits of technology

It improves the accuracy and robustness of emotion recognition, reduces reliance on labeled data, enhances the scalability of the method, and provides a more accurate user experience and decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118038901B_ABST
    Figure CN118038901B_ABST
Patent Text Reader

Abstract

本发明公开一种双模态语音情感识别方法及系统,涉及情绪识别技术领域,包括:获取待识别语音数据的语音信号,并提取其中的文本信息;对语音信号进行分帧处理后输入语音预训练模型中进行编码,获得语音信号的高级特征;提取语音信息中的声学特征,将高级特征与声学特征按帧拼接,获得语音特征序列;使用文本预训练模型提取出文本特征序列;提取语音特征序列和文本特征序列中的关键情感特征并添加时序信息,获得语音深度情感特征和文本深度情感特征;采用模态融合算法将语音深度情感特征和文本深度情感特征进行融合,获得语音情感特征来对待识别语音数据进行情感识别;将语音与文本两种模态的信息有机地融合,提高了情感识别的准确性和鲁棒性。
Need to check novelty before this filing date? Find Prior Art