Cross-lingual Speaker Adaptation via Universal Speech Model Transform

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis systems struggle to accurately replicate the voice characteristics of a speaker in one language when generating speech in a different language, limiting their ability to provide personalized and culturally appropriate output in multilingual settings.

Innovation Solution

The method involves using a universal speech model to estimate a speaker transform from input speech data, which is then applied to a speaker-independent speech model to create a speaker-specific model, allowing for the generation of speech in a second language that mimics the voice characteristics of the original speaker, even if they do not speak that language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a speaker-independent speech model is used for multi-lingual synthesis, then the system can generate speech in multiple languages, but it cannot replicate the voice characteristics of a specific speaker across different languages

Engineering Contradiction:
Improvemulti-lingual capabilityVSAvoidvoice characteristics replication
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transforms speaker-specific parameters from the source language into target language parameters using a transformation matrix. This allows the voice characteristics (spectral envelope, pitch contour, timing) to be adapted across languages while maintaining speaker identity. The core innovation is changing the parameter representation from language-specific to language-invariant through linear transformation of acoustic features.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a virtual copy of the speaker's voice characteristics by extracting and storing speaker-specific parameters from source language speech. These parameters are then copied and transformed for synthesis in target languages, enabling the same speaker identity to be reproduced across multiple languages without requiring actual speech samples in each target language.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If speech data from multiple speakers and languages is used to train a universal model, then the model becomes more versatile, but it loses the ability to capture individual speaker characteristics

Engineering Contradiction:
Improvelanguage coverageVSAvoidspeaker characteristic fidelity
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the speech synthesis task into two independent components: a language-independent speaker characteristic model and a language-specific phonetic model. The speaker characteristics are extracted and stored separately from language-specific content, allowing the speaker model to be trained on limited data while maintaining high fidelity to individual voice characteristics across multiple languages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal speaker adaptation framework where a single speaker model can be applied across multiple languages through parameter transformation. The speaker characteristic extractor and transformation matrix serve universal functions, enabling the same speaker model to generate speech in different languages while maintaining consistent voice characteristics.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9922641B1Cross-lingual speaker adaptation for multi-lingual speech synthesis
Publication Date: 2018.03.20 GOOGLE LLC
  • US9922641B1 patent drawing
  • US9922641B1 patent drawing
  • US9922641B1 patent drawing

AI summary

The subject matter of the disclosure is embodied in a method that includes receiving input speech data from a speaker in a first language, and estimating, based on a universal speech model, a speaker transform representing speaker characteristics associated with the input speech data. The method also includes accessing a speaker-independent speech model for generating speech data in a second language that is different from the first language. The method further includes modifying the speaker-independent speech model using the speaker transform to obtain a speaker-specific speech model, and generating speech data in the second language using the speaker-specific speech model.