Cross-lingual Speaker Adaptation via Universal Speech Model Transform
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis systems struggle to accurately replicate the voice characteristics of a speaker in one language when generating speech in a different language, limiting their ability to provide personalized and culturally appropriate output in multilingual settings.
Innovation Solution
The method involves using a universal speech model to estimate a speaker transform from input speech data, which is then applied to a speaker-independent speech model to create a speaker-specific model, allowing for the generation of speech in a second language that mimics the voice characteristics of the original speaker, even if they do not speak that language.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a speaker-independent speech model is used for multi-lingual synthesis, then the system can generate speech in multiple languages, but it cannot replicate the voice characteristics of a specific speaker across different languages
Solution Approach 1:
The patent transforms speaker-specific parameters from the source language into target language parameters using a transformation matrix. This allows the voice characteristics (spectral envelope, pitch contour, timing) to be adapted across languages while maintaining speaker identity. The core innovation is changing the parameter representation from language-specific to language-invariant through linear transformation of acoustic features.
Solution Approach 2:
The patent creates a virtual copy of the speaker's voice characteristics by extracting and storing speaker-specific parameters from source language speech. These parameters are then copied and transformed for synthesis in target languages, enabling the same speaker identity to be reproduced across multiple languages without requiring actual speech samples in each target language.
2Adaptability or versatility
If speech data from multiple speakers and languages is used to train a universal model, then the model becomes more versatile, but it loses the ability to capture individual speaker characteristics
Solution Approach 1:
The patent segments the speech synthesis task into two independent components: a language-independent speaker characteristic model and a language-specific phonetic model. The speaker characteristics are extracted and stored separately from language-specific content, allowing the speaker model to be trained on limited data while maintaining high fidelity to individual voice characteristics across multiple languages.
Solution Approach 2:
The patent creates a universal speaker adaptation framework where a single speaker model can be applied across multiple languages through parameter transformation. The speaker characteristic extractor and transformation matrix serve universal functions, enabling the same speaker model to generate speech in different languages while maintaining consistent voice characteristics.
Data Source
AI summary
The subject matter of the disclosure is embodied in a method that includes receiving input speech data from a speaker in a first language, and estimating, based on a universal speech model, a speaker transform representing speaker characteristics associated with the input speech data. The method also includes accessing a speaker-independent speech model for generating speech data in a second language that is different from the first language. The method further includes modifying the speaker-independent speech model using the speaker transform to obtain a speaker-specific speech model, and generating speech data in the second language using the speaker-specific speech model.


