Two-Level Speech Prosody Transfer via Intermediate Representation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-speech (TTS) models struggle to effectively model a full variety of prosodic styles, resulting in synthesized speech that lacks expressiveness, especially when transferring prosody from one speaker to another or across different domains.
Innovation Solution
A two-level speech prosody transfer system is employed, where a first TTS model generates an intermediate synthesized speech representation that captures the intended prosody, and a second TTS model encodes this representation into an utterance embedding to generate expressive speech with the intended prosody and speaker characteristics of a target voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single TTS model is used to generate speech with both prosody and speaker characteristics, then the system complexity is reduced, but the ability to accurately transfer prosody across different domains and speakers deteriorates
Solution Approach 1:
The patent divides the TTS system into two separate models: a first TTS model that generates intermediate speech representations with accurate prosody, and a second TTS model that transfers the prosody to the target speaker's voice. This segmentation allows each model to specialize in one aspect (prosody generation and speaker adaptation), thereby improving overall prosody transfer capability while managing system complexity through modular design
Solution Approach 2:
The patent introduces an intermediate synthesized speech representation as a mediator between the prosody generation stage and the speaker adaptation stage. This intermediate representation contains the prosodic features extracted from the first TTS model and serves as input to the second TTS model, enabling effective prosody transfer to different speakers and domains without requiring direct optimization between disparate components
2Adaptability or versatility
If prosody is transferred to a target voice with insufficient training data, then the versatility of the target voice is improved, but the naturalness and quality of the synthesized speech deteriorates
Solution Approach 1:
The patent performs preliminary prosody encoding by extracting prosodic features from the intermediate synthesized speech representation before applying them to the target voice. The encoder portion processes the intermediate representation to create an utterance embedding that captures the prosodic characteristics, which is then used by the decoder to generate natural-sounding speech in the target voice, even when training data is limited
Solution Approach 2:
The patent copies the prosodic patterns from the intermediate speech representation into the target voice through the encoder-decoder architecture. The utterance embedding serves as a copied representation of the prosody that can be applied to different speakers, allowing the target voice to inherit natural prosodic characteristics from the source without requiring extensive target-specific training data
Data Source
AI summary
A method includes receiving an input text utterance to be synthesized into expressive speech having an intended prosody and a target voice and generating, using a first text-to-speech (TTS) model, an intermediate synthesized speech representation for the input text utterance. The intermediate synthesized speech representation possesses the intended prosody. The method also includes providing the intermediate synthesized speech representation to a second TTS model that includes an encoder portion and a decoder portion. The encoder portion is configured to encode the intermediate synthesized speech representation into an utterance embedding that specifies the intended prosody. The decoder portion is configured to process the input text utterance and the utterance embedding to generate an output audio signal of expressive speech that has the intended prosody specified by the utterance embedding and speaker characteristics of the target voice.


