Text-to-Speech Voice Synthesis Using Reference Speaker Vector Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques for building high-quality statistical parametric speech synthesis (SPSS) systems face challenges in scalability and cost-effectiveness, particularly when using recordings from multiple speakers in diverse environments, and there is a lack of large, uniform speech databases for long-tail languages, leading to reduced quality and availability of text-to-speech (TTS) services.
Innovation Solution
A method and system that condition diverse speech recordings from multiple speakers using a reference speaker's database, replacing colloquial speaker vectors with optimally-matched reference speaker vectors through a 'matching under transform' technique to create a high-quality aggregated speech database for training SPSS systems, enabling the use of diverse recordings for building effective TTS systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If diverse recordings from multiple speakers are used to build SPSS systems, then scalability and cost-effectiveness improve, but speech quality and consistency deteriorate
Solution Approach 1:
The patent introduces a reference speaker as an intermediary to mediate between diverse colloquial speakers and the target SPSS system. Colloquial speaker vectors are transformed and matched against reference speaker vectors, using the reference speaker as a standard mediator to ensure quality consistency while allowing diverse input sources. This resolves the contradiction by enabling scalability through multiple speakers while maintaining quality through the reference speaker mediation.
Solution Approach 2:
The patent applies parameter transformation by converting colloquial speaker feature vectors into reference speaker feature vectors through statistical parametric transformation. The transformation changes the parameter space from diverse colloquial characteristics to standardized reference characteristics, enabling quality consistency while accepting diverse recordings. This resolves the contradiction by transforming parameters to maintain quality standards across scalable diverse inputs.
2Ease of manufacture
If diverse recordings from multiple speakers are used to build SPSS systems, then cost-effectiveness improves, but speech consistency deteriorates
Solution Approach 1:
The reference speaker serves as a quality mediator that standardizes diverse colloquial inputs. By transforming colloquial speaker vectors to match reference speaker vectors, the system maintains consistent speech characteristics while accepting cost-effective diverse recordings from multiple speakers. This resolves the contradiction by using the reference speaker as a mediator that ensures consistency without requiring expensive controlled recording conditions.
Solution Approach 2:
The patent creates copies of reference speaker vectors to replace colloquial speaker vectors in the training database. Instead of using original diverse recordings directly, the system creates standardized copies based on the reference speaker, ensuring speech consistency while maintaining the cost benefits of using diverse colloquial sources. This resolves the contradiction by copying reference characteristics onto diverse inputs.
3Manufacturing precision
If colloquial speaker vectors are replaced with reference speaker vectors, then speech quality improves, but processing complexity increases
Solution Approach 1:
The patent performs preliminary transformation of colloquial speaker vectors into reference speaker vectors before aggregating them into the training database. By pre-processing and standardizing vectors before aggregation, the system improves speech quality while managing processing complexity through organized preliminary steps rather than complex real-time processing. This resolves the contradiction by applying transformation in advance in a structured manner.
Solution Approach 2:
The patent segments the processing into distinct steps: extracting colloquial speaker vectors, transforming them to reference speaker vectors, matching them optimally, and aggregating into the training database. This segmentation of the complex transformation process into manageable stages improves speech quality while making the processing complexity more tractable and organized. This resolves the contradiction by breaking down the complex transformation into sequential manageable steps.
Data Source
AI summary
A method and system is disclosed for building a speech database for a text-to-speech (TTS) synthesis system from multiple speakers recorded under diverse conditions. For a plurality of utterances of a reference speaker, a set of reference-speaker vectors may be extracted, and for each of a plurality of utterances of a colloquial speaker, a respective set of colloquial-speaker vectors may be extracted. A matching procedure, carried out under a transform that compensates for speaker differences, may be used to match each colloquial-speaker vector to a reference-speaker vector. The colloquial-speaker vector may be replaced with the matched reference-speaker vector. The matching-and-replacing can be carried out separately for each set of colloquial-speaker vectors. A conditioned set of speaker vectors can then be constructed by aggregating all the replaced speaker vectors. The condition set of speaker vectors can be used to train the TTS system.


