Distributed Voice Data Collection System for TTS Phoneme Coverage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Collecting high-quality voice data from a diverse range of individuals to create personalized text-to-speech (TTS) voices is challenging due to the need for extensive and varied speech samples, especially from individuals with limited speaking ability or diverse accents, and ensuring data quality in various environments.

Innovation Solution

A system for collecting voice data from multiple donors using a network-based platform that allows for distributed data collection, user-friendly interfaces, and advanced processing techniques to ensure high-quality data storage and phoneme representation, including multiple sessions and feedback mechanisms to improve donor engagement and data accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice data is collected from a large number of diverse voice donors, then the quality and representativeness of the voice bank is improved, but the complexity of data collection and processing increases

Engineering Contradiction:
Improvevoice bank qualityVSAvoiddata collection system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the voice bank creation process into distinct components: voice data collection from multiple donors, phoneme extraction and labeling, voice quality assessment, and TTS model training. This segmentation allows each component to be optimized independently while maintaining overall system reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The voice bank system is designed to handle multiple types of voice data from diverse donors simultaneously, processing different languages, accents, and speech patterns through a unified framework. This multi-functional approach improves voice bank quality without requiring separate systems for each donor type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Quantity of substance

If extensive voice data is collected to represent all speech sounds, then the completeness of phoneme coverage is improved, but the time and resources required for data collection increase

Engineering Contradiction:
Improvephoneme coverageVSAvoiddata collection time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs preliminary phoneme extraction and labeling from raw voice data before full processing. By pre-identifying and organizing phoneme units from diverse donors, the system reduces the time required for subsequent TTS model training and ensures comprehensive phoneme coverage is achieved efficiently.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates phoneme representations and voice models that can be copied and reused across different TTS applications. Once phoneme data is collected and processed from diverse donors, these standardized representations can be replicated for multiple voice synthesis tasks, reducing redundant data collection time.

Inventive Principle:
Principle #26Copying

3Measurement precision

If high quality voice data is collected from individuals with limited speaking ability, then the personalization accuracy of TTS voices is improved, but the difficulty of data collection increases

Engineering Contradiction:
ImproveTTS voice accuracyVSAvoiddata collection difficulty
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The system incorporates automated voice quality assessment and phoneme identification that operates without requiring expert annotators. The automated processing extracts phoneme data from voice recordings independently, reducing the difficulty of collecting high-quality data from donors with limited speaking ability while maintaining TTS voice accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual voice analysis and phoneme labeling by experts is replaced with automated speech processing algorithms. This substitution reduces the difficulty of data collection from vulnerable populations by eliminating the need for complex manual intervention while preserving measurement precision for TTS synthesis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9336782B1Distributed collection and processing of voice bank data
Publication Date: 2016.05.10 VERITONE INC
  • US9336782B1 patent drawing
  • US9336782B1 patent drawing
  • US9336782B1 patent drawing

AI summary

Voice data may be collected by a plurality of voice donors and stored in a voice bank. A voice donor may authenticate to a voice collection system to start a session to provide voice data. During the voice collection session, the voice donor may be presented with a sequence of prompts to speak and voice data may be transferred to a server. The received voice data may be processed to determine the speech units spoken by the voice donor and a count of speech units received from the voice donor may be updated. Feedback may be provided to the voice donor indicating, for example, a progress of the voice collection, a quality level of the voice data, or information about speech unit counts. The voice bank may be used to create TTS voices for voice recipients, create a model of voice aging, or for other applications.