Emotional Speech Processing Using PLDA Normalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current emotional speech processing technologies face challenges in distinguishing emotional speech from read/conversational speech and struggle with speaker variability, leading to poor performance in emotion recognition and classification.

Innovation Solution

The use of Probabilistic Linear Discriminant Analysis (PLDA) to normalize speech representation, making it more emotion/speaking style dependent and less speaker dependent, by transforming Gaussian Mixture Model supervectors and applying it for emotion clustering and classification across different languages and speaking styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If statistical voice recognition models trained with read speech are used, then recognition performance on read speech is good, but recognition performance on emotional speech deteriorates

Engineering Contradiction:
Improverecognition performanceVSAvoidperformance across different speech types
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments speech data by emotional state and speaker, creating separate Gaussian Mixture Models for each emotion-speaker combination. This segmentation allows the system to handle emotional speech characteristics differently from read speech, improving recognition performance across diverse speech types while maintaining reliability on standard speech.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms speech features using Probabilistic Linear Discriminant Analysis (PLDA) to change the parameter space from speaker-dependent to emotion-dependent representations. This parameter transformation enables the same model to effectively process both read speech and emotional speech by adjusting the feature representation rather than requiring separate models.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If traditional emotion recognition methods are used, then some emotion detection is possible, but accuracy deteriorates due to speaker variability

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidhandling speaker variability
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes speaker-specific information from speech features using PLDA, isolating the emotional characteristics from speaker identity. This extraction process eliminates the confounding effect of speaker variability, allowing accurate emotion recognition across different speakers without requiring complex speaker adaptation mechanisms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the feature representation parameters from speaker-dependent to emotion-dependent through PLDA transformation. This parameter change simplifies the problem by focusing only on emotion-related variations in the speech signal, improving measurement precision while reducing the complexity of handling speaker variability.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If speaker-dependent speech models are used, then performance on specific speakers is good, but generalization to new speakers and emotions deteriorates

Engineering Contradiction:
Improveperformance on trained speakersVSAvoidgeneralization capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal emotion recognition framework using PLDA that can handle multiple speakers and emotional states with a single model. The PLDA transformation produces speaker-independent emotion features that generalize across different speakers and languages, making the system multi-functional for various speech types without requiring speaker-specific training.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10127927B2Emotional speech processing
Publication Date: 2018.11.13 SONY INTERACTIVE ENTERTAINMENT LLC
  • US10127927B2 patent drawing
  • US10127927B2 patent drawing
  • US10127927B2 patent drawing

AI summary

A method for emotion or speaking style recognition and/or clustering comprises receiving one or more speech samples, generating a set of training data by extracting one or more acoustic features from every frame of the one or more speech samples, and generating a model from the set of training data, wherein the model identifies emotion or speaking style dependent information in the set of training data. The method may further comprise receiving one or more test speech samples, generating a set of test data by extracting one or more acoustic features from every frame of the one or more test speeches, and transforming the set of test data using the model to better represent emotion/speaking style dependent information, and use the transformed data for clustering and/or classification to discover speech with similar emotion or speaking style. It is emphasized that this abstract is provided to comply with the rules requiring an abstract that will allow a searcher or other reader to quickly ascertain the subject matter of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims.