Cross-Speaker Style Transfer in Text-to-Speech via PPG and Prosody Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech (TTS) models require extensive and time-consuming data collection and training to produce speech in multiple speaking styles, as they are speaker-dependent and require vast amounts of data, making it inefficient and costly to train models for various styles.
Innovation Solution
The method involves generating spectrogram data that combines the voice timbre of a target speaker with the prosody style of a source speaker, using phonetic posterior gram (PPG) data and additional prosody features to create a versatile training dataset for neural TTS models, allowing for efficient cross-speaker style transfer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a TTS model is trained using extensive target speaker data to produce speech in multiple styles, then the model can accurately replicate target speaker voice timbre, but the data collection and training process becomes extremely time-consuming and costly
Solution Approach 1:
The patent introduces an intermediary approach by using source speaker data combined with style transfer techniques to generate target speaker speech. Instead of directly collecting extensive target speaker data, the system uses source speaker acoustic models and applies style transfer to transform the speech into target speaker voice timbre, thereby reducing the time and cost of data collection while maintaining voice replication accuracy
Solution Approach 2:
The patent employs copying by creating synthetic training data that replicates target speaker characteristics. The system generates synthetic speech data by copying the voice timbre and prosody style of target speakers through style transfer models, which then train the TTS model without requiring actual recordings from target speakers, thus significantly reducing data collection time
2Adaptability or versatility
If a TTS model is trained with vast amounts of style-specific training data, then the model can produce diverse speaking styles, but the training cost and complexity increase proportionally
Solution Approach 1:
The patent applies universality by creating a single training framework that can generate multiple speaking styles through style transfer. Instead of training separate models for each style, the system uses a universal approach where source speaker data is transformed into various target speaker styles using style transfer techniques, allowing one training process to produce multiple speaking capabilities
Solution Approach 2:
The patent utilizes parameter changes by modifying the acoustic model parameters through style transfer. The system adjusts prosody parameters, pitch contours, and spectral characteristics of source speaker speech to transform it into different target speaker styles, enabling versatile style production through parameter transformation rather than requiring separate training data for each style
3Reliability
If the acoustic model is made speaker-dependent to accurately replicate target speaker voice, then the model can produce authentic target speaker speech, but the training process becomes more difficult and time-consuming
Solution Approach 1:
The patent applies preliminary action by pre-training a source acoustic model on diverse speaker data before fine-tuning for specific target speakers. This preliminary training creates a robust foundation that can efficiently adapt to different target speakers through style transfer, reducing the time and difficulty of subsequent speaker-specific training while maintaining authentic voice replication
Data Source
AI summary
Systems are configured for generating spectrogram data characterized by a voice timbre of a target speaker and a prosody style of source speaker by converting a waveform of source speaker data to phonetic posterior gram (PPG) data, extracting additional prosody features from the source speaker data, and generating a spectrogram based on the PPG data and the extracted prosody features. The systems are configured to utilize/train a machine learning model for generating spectrogram data and for training a neural text-to-speech model with the generated spectrogram data.


