Textless Speech Emotion Conversion via Discrete Representations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing technologies face challenges in accurately recognizing and generating speech signals with emotions, leading to higher error rates in automatic speech recognition and struggling to produce convincing emotional expressions.

Innovation Solution

A speech-processing system that decomposes speech signals into phonetic-content units, prosodic features, speaker identity, and emotion, allowing for the translation of emotions while preserving lexical content and generating expressive speech signals by applying a neural vocoder to predicted representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If emotion is removed from speech signals to improve automatic speech recognition accuracy, then ASR error rate decreases, but the ability to generate expressive speech signals is lost

Engineering Contradiction:
Improveautomatic speech recognition accuracyVSAvoidemotional expression capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The speech signal is segmented into distinct components: phonetic content units and emotional features. This segmentation allows the system to separately process lexical content and emotional expression, enabling accurate ASR by isolating phonetic information while preserving the ability to generate expressive speech by manipulating emotional features independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts emotional features from speech signals as separate entities from the phonetic content. By taking out emotional features (such as pitch contours, energy patterns, and temporal dynamics) as independent elements, the system can remove them for accurate ASR or reinsert modified versions for expressive speech generation, resolving the contradiction between accuracy and expression.

Inventive Principle:
Principle #2Taking out (Extraction)

2Ease of manufacture

If conventional speech processing methods are used to generate emotional speech, then processing simplicity is maintained, but the convincing quality of emotional expressions deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoidemotional expression quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The system introduces phonetic content units as intermediary representations between raw speech and emotional features. These units serve as a bridge that preserves lexical meaning while allowing independent manipulation of emotional characteristics. This intermediary approach maintains processing simplicity by working with discrete units while achieving high emotional expression quality through controlled modification of prosodic and spectral features.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system achieves convincing emotional expressions by systematically changing key speech parameters including pitch contours, energy distribution, temporal dynamics, and spectral characteristics. By applying parameter changes to the phonetic content units in a controlled manner, the system generates high-quality emotional speech without requiring complex processing architectures, thus maintaining ease of manufacture while improving expression quality.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240105207A1Textless Speech Emotion Conversion Using Discrete and Decomposed Representations
Publication Date: 2024.03.28 META PLATFORMS INC
  • US20240105207A1 patent drawing
  • US20240105207A1 patent drawing
  • US20240105207A1 patent drawing

AI summary

In one embodiment, a method includes accessing a speech signal corresponding to a source emotion, generating content units based on the speech signal, generating altered content units for the content units based on a target emotion, determining a respective duration for each of the altered content units based on the target emotion, generating a respective pitch curve for each of the altered content units based on the target emotion and the respective altered duration, and generating an altered speech signal corresponding to the target emotion based on the target emotion, speech characteristics associated with a speaker, the altered content units based on their respective altered durations, and the pitch curves for the altered content units.