Cross-Speaker Style Transfer in Text-to-Speech via PPG and Prosody Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech (TTS) models require extensive and time-consuming data collection and training to produce speech in multiple speaking styles, as they are speaker-dependent and require vast amounts of data, making it inefficient and costly to train models for various styles.

Innovation Solution

The method involves generating spectrogram data that combines the voice timbre of a target speaker with the prosody style of a source speaker, using phonetic posterior gram (PPG) data and additional prosody features to create a versatile training dataset for neural TTS models, allowing for efficient cross-speaker style transfer.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a TTS model is trained using extensive target speaker data to produce speech in multiple styles, then the model can accurately replicate target speaker voice timbre, but the data collection and training process becomes extremely time-consuming and costly

Engineering Contradiction:
Improveaccuracy of voice timbre replicationVSAvoidtime required for data collection and training
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent introduces an intermediary approach by using source speaker data combined with style transfer techniques to generate target speaker speech. Instead of directly collecting extensive target speaker data, the system uses source speaker acoustic models and applies style transfer to transform the speech into target speaker voice timbre, thereby reducing the time and cost of data collection while maintaining voice replication accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent employs copying by creating synthetic training data that replicates target speaker characteristics. The system generates synthetic speech data by copying the voice timbre and prosody style of target speakers through style transfer models, which then train the TTS model without requiring actual recordings from target speakers, thus significantly reducing data collection time

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If a TTS model is trained with vast amounts of style-specific training data, then the model can produce diverse speaking styles, but the training cost and complexity increase proportionally

Engineering Contradiction:
Improvecapability to produce multiple speaking stylesVSAvoidtraining process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by creating a single training framework that can generate multiple speaking styles through style transfer. Instead of training separate models for each style, the system uses a universal approach where source speaker data is transformed into various target speaker styles using style transfer techniques, allowing one training process to produce multiple speaking capabilities

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent utilizes parameter changes by modifying the acoustic model parameters through style transfer. The system adjusts prosody parameters, pitch contours, and spectral characteristics of source speaker speech to transform it into different target speaker styles, enabling versatile style production through parameter transformation rather than requiring separate training data for each style

Inventive Principle:
Principle #35Parameter changes

3Reliability

If the acoustic model is made speaker-dependent to accurately replicate target speaker voice, then the model can produce authentic target speaker speech, but the training process becomes more difficult and time-consuming

Engineering Contradiction:
Improveauthenticity of target speaker speechVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-training a source acoustic model on diverse speaker data before fine-tuning for specific target speakers. This preliminary training creates a robust foundation that can efficiently adapt to different target speakers through style transfer, reducing the time and difficulty of subsequent speaker-specific training while maintaining authentic voice replication

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11361753B2System and method for cross-speaker style transfer in text-to-speech and training data generation
Publication Date: 2022.06.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11361753B2 patent drawing
  • US11361753B2 patent drawing
  • US11361753B2 patent drawing

AI summary

Systems are configured for generating spectrogram data characterized by a voice timbre of a target speaker and a prosody style of source speaker by converting a waveform of source speaker data to phonetic posterior gram (PPG) data, extracting additional prosody features from the source speaker data, and generating a spectrogram based on the PPG data and the extracted prosody features. The systems are configured to utilize/train a machine learning model for generating spectrogram data and for training a neural text-to-speech model with the generated spectrogram data.