Emotion-Aware Speaker Identification for Emotional Speech Variation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speaker identification technologies fail to accurately identify speakers when the evaluated utterance contains emotional variations such as laughter or angry shouts, as they do not account for the emotion in the utterance, leading to decreased accuracy.
Innovation Solution
A speaker identification apparatus that utilizes a trained deep neural network (DNN) to estimate the emotion in an utterance and adjusts the speaker identification process accordingly, using a speaker identification processor to output a score based on the emotion estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speaker identification technology calculates similarity between speaker feature vectors without considering emotion, then the identification process is simple, but the speaker identification accuracy decreases when emotional variations are present
Solution Approach 1:
The identification process is segmented into two independent modules: an emotion estimator that detects emotional states from acoustic features, and a speaker identification processor that uses both acoustic features and emotion information to calculate speaker similarity. This segmentation allows the system to account for emotional variations without creating a monolithic complex system.
Solution Approach 2:
An emotion estimator acts as an intermediary component between the acoustic feature extraction and speaker identification processing. This intermediary detects emotional states and provides correction information to the speaker identification processor, enabling accurate speaker identification despite emotional variations in the utterance.
2Reliability
If the system accounts for emotion in utterance to improve identification accuracy, then speaker identification accuracy improves, but the device complexity increases due to additional emotion estimation components
Solution Approach 1:
The speaker identification processor is designed to handle multiple functions: it processes both normal acoustic feature-based speaker identification and emotion-corrected speaker identification. By making the processor universal, the system achieves high reliability across different utterance types without requiring entirely separate processing paths.
Solution Approach 2:
The system changes the parameters used in speaker identification based on detected emotion. When emotion is detected, the system adjusts the feature vectors by applying correction amounts derived from emotion information, thereby improving identification reliability without fundamentally changing the core identification algorithm.
Data Source
AI summary
A speaker identification apparatus that identifies a speaker of utterance data indicating a voice of an utterance subjected to identification includes: an emotion estimator that estimates, from an acoustic feature value calculated from the utterance data, an emotion contained in the voice of the utterance indicated by the utterance data, using a trained deep neural network (DNN); and a speaker identification processor that outputs, based on the acoustic feature value calculated from the utterance data, a score for identifying the speaker of the utterance data, using an estimation result of the emotion estimator.


