Emotion-Aware Speaker Identification for Emotional Speech Variation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker identification technologies fail to accurately identify speakers when the evaluated utterance contains emotional variations such as laughter or angry shouts, as they do not account for the emotion in the utterance, leading to decreased accuracy.

Innovation Solution

A speaker identification apparatus that utilizes a trained deep neural network (DNN) to estimate the emotion in an utterance and adjusts the speaker identification process accordingly, using a speaker identification processor to output a score based on the emotion estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speaker identification technology calculates similarity between speaker feature vectors without considering emotion, then the identification process is simple, but the speaker identification accuracy decreases when emotional variations are present

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoididentification process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The identification process is segmented into two independent modules: an emotion estimator that detects emotional states from acoustic features, and a speaker identification processor that uses both acoustic features and emotion information to calculate speaker similarity. This segmentation allows the system to account for emotional variations without creating a monolithic complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An emotion estimator acts as an intermediary component between the acoustic feature extraction and speaker identification processing. This intermediary detects emotional states and provides correction information to the speaker identification processor, enabling accurate speaker identification despite emotional variations in the utterance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system accounts for emotion in utterance to improve identification accuracy, then speaker identification accuracy improves, but the device complexity increases due to additional emotion estimation components

Engineering Contradiction:
Improveidentification reliabilityVSAvoidsystem structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The speaker identification processor is designed to handle multiple functions: it processes both normal acoustic feature-based speaker identification and emotion-corrected speaker identification. By making the processor universal, the system achieves high reliability across different utterance types without requiring entirely separate processing paths.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the parameters used in speaker identification based on detected emotion. When emotion is detected, the system adjusts the feature vectors by applying correction amounts derived from emotion information, thereby improving identification reliability without fundamentally changing the core identification algorithm.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12394421B2Speaker identification apparatus, speaker identification method, and recording medium
Publication Date: 2025.08.19 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US12394421B2 patent drawing
  • US12394421B2 patent drawing
  • US12394421B2 patent drawing

AI summary

A speaker identification apparatus that identifies a speaker of utterance data indicating a voice of an utterance subjected to identification includes: an emotion estimator that estimates, from an acoustic feature value calculated from the utterance data, an emotion contained in the voice of the utterance indicated by the utterance data, using a trained deep neural network (DNN); and a speaker identification processor that outputs, based on the acoustic feature value calculated from the utterance data, a score for identifying the speaker of the utterance data, using an estimation result of the emotion estimator.