Textless Speech Emotion Conversion via Discrete Representations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech processing technologies face challenges in accurately recognizing and generating speech signals with emotions, leading to higher error rates in automatic speech recognition and struggling to produce convincing emotional expressions.
Innovation Solution
A speech-processing system that decomposes speech signals into phonetic-content units, prosodic features, speaker identity, and emotion, allowing for the translation of emotions while preserving lexical content and generating expressive speech signals by applying a neural vocoder to predicted representations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If emotion is removed from speech signals to improve automatic speech recognition accuracy, then ASR error rate decreases, but the ability to generate expressive speech signals is lost
Solution Approach 1:
The speech signal is segmented into distinct components: phonetic content units and emotional features. This segmentation allows the system to separately process lexical content and emotional expression, enabling accurate ASR by isolating phonetic information while preserving the ability to generate expressive speech by manipulating emotional features independently.
Solution Approach 2:
The system extracts emotional features from speech signals as separate entities from the phonetic content. By taking out emotional features (such as pitch contours, energy patterns, and temporal dynamics) as independent elements, the system can remove them for accurate ASR or reinsert modified versions for expressive speech generation, resolving the contradiction between accuracy and expression.
2Ease of manufacture
If conventional speech processing methods are used to generate emotional speech, then processing simplicity is maintained, but the convincing quality of emotional expressions deteriorates
Solution Approach 1:
The system introduces phonetic content units as intermediary representations between raw speech and emotional features. These units serve as a bridge that preserves lexical meaning while allowing independent manipulation of emotional characteristics. This intermediary approach maintains processing simplicity by working with discrete units while achieving high emotional expression quality through controlled modification of prosodic and spectral features.
Solution Approach 2:
The system achieves convincing emotional expressions by systematically changing key speech parameters including pitch contours, energy distribution, temporal dynamics, and spectral characteristics. By applying parameter changes to the phonetic content units in a controlled manner, the system generates high-quality emotional speech without requiring complex processing architectures, thus maintaining ease of manufacture while improving expression quality.
Data Source
AI summary
In one embodiment, a method includes accessing a speech signal corresponding to a source emotion, generating content units based on the speech signal, generating altered content units for the content units based on a target emotion, determining a respective duration for each of the altered content units based on the target emotion, generating a respective pitch curve for each of the altered content units based on the target emotion and the respective altered duration, and generating an altered speech signal corresponding to the target emotion based on the target emotion, speech characteristics associated with a speaker, the altered content units based on their respective altered durations, and the pitch curves for the altered content units.


