Text-to-Speech Voice Synthesis Using Reference Speaker Vector Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for building high-quality statistical parametric speech synthesis (SPSS) systems face challenges in scalability and cost-effectiveness, particularly when using recordings from multiple speakers in diverse environments, and there is a lack of large, uniform speech databases for long-tail languages, leading to reduced quality and availability of text-to-speech (TTS) services.

Innovation Solution

A method and system that condition diverse speech recordings from multiple speakers using a reference speaker's database, replacing colloquial speaker vectors with optimally-matched reference speaker vectors through a 'matching under transform' technique to create a high-quality aggregated speech database for training SPSS systems, enabling the use of diverse recordings for building effective TTS systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If diverse recordings from multiple speakers are used to build SPSS systems, then scalability and cost-effectiveness improve, but speech quality and consistency deteriorate

Engineering Contradiction:
ImprovescalabilityVSAvoidspeech quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces a reference speaker as an intermediary to mediate between diverse colloquial speakers and the target SPSS system. Colloquial speaker vectors are transformed and matched against reference speaker vectors, using the reference speaker as a standard mediator to ensure quality consistency while allowing diverse input sources. This resolves the contradiction by enabling scalability through multiple speakers while maintaining quality through the reference speaker mediation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter transformation by converting colloquial speaker feature vectors into reference speaker feature vectors through statistical parametric transformation. The transformation changes the parameter space from diverse colloquial characteristics to standardized reference characteristics, enabling quality consistency while accepting diverse recordings. This resolves the contradiction by transforming parameters to maintain quality standards across scalable diverse inputs.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If diverse recordings from multiple speakers are used to build SPSS systems, then cost-effectiveness improves, but speech consistency deteriorates

Engineering Contradiction:
Improvecost-effectivenessVSAvoidspeech consistency
Core Design Contradiction:
Ease of manufactureVSStability of the object's composition

Solution Approach 1:

The reference speaker serves as a quality mediator that standardizes diverse colloquial inputs. By transforming colloquial speaker vectors to match reference speaker vectors, the system maintains consistent speech characteristics while accepting cost-effective diverse recordings from multiple speakers. This resolves the contradiction by using the reference speaker as a mediator that ensures consistency without requiring expensive controlled recording conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates copies of reference speaker vectors to replace colloquial speaker vectors in the training database. Instead of using original diverse recordings directly, the system creates standardized copies based on the reference speaker, ensuring speech consistency while maintaining the cost benefits of using diverse colloquial sources. This resolves the contradiction by copying reference characteristics onto diverse inputs.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If colloquial speaker vectors are replaced with reference speaker vectors, then speech quality improves, but processing complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary transformation of colloquial speaker vectors into reference speaker vectors before aggregating them into the training database. By pre-processing and standardizing vectors before aggregation, the system improves speech quality while managing processing complexity through organized preliminary steps rather than complex real-time processing. This resolves the contradiction by applying transformation in advance in a structured manner.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the processing into distinct steps: extracting colloquial speaker vectors, transforming them to reference speaker vectors, matching them optimally, and aggregating into the training database. This segmentation of the complex transformation process into manageable stages improves speech quality while making the processing complexity more tractable and organized. This resolves the contradiction by breaking down the complex transformation into sequential manageable steps.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9542927B2Method and system for building text-to-speech voice from diverse recordings
Publication Date: 2017.01.10 GOOGLE LLC
  • US9542927B2 patent drawing
  • US9542927B2 patent drawing
  • US9542927B2 patent drawing

AI summary

A method and system is disclosed for building a speech database for a text-to-speech (TTS) synthesis system from multiple speakers recorded under diverse conditions. For a plurality of utterances of a reference speaker, a set of reference-speaker vectors may be extracted, and for each of a plurality of utterances of a colloquial speaker, a respective set of colloquial-speaker vectors may be extracted. A matching procedure, carried out under a transform that compensates for speaker differences, may be used to match each colloquial-speaker vector to a reference-speaker vector. The colloquial-speaker vector may be replaced with the matched reference-speaker vector. The matching-and-replacing can be carried out separately for each set of colloquial-speaker vectors. A conditioned set of speaker vectors can then be constructed by aggregating all the replaced speaker vectors. The condition set of speaker vectors can be used to train the TTS system.