Distributed Voice Data Collection System for TTS Phoneme Coverage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Collecting high-quality voice data from a diverse range of individuals to create personalized text-to-speech (TTS) voices is challenging due to the need for extensive and varied speech samples, especially from individuals with limited speaking ability or diverse accents, and ensuring data quality in various environments.
Innovation Solution
A system for collecting voice data from multiple donors using a network-based platform that allows for distributed data collection, user-friendly interfaces, and advanced processing techniques to ensure high-quality data storage and phoneme representation, including multiple sessions and feedback mechanisms to improve donor engagement and data accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice data is collected from a large number of diverse voice donors, then the quality and representativeness of the voice bank is improved, but the complexity of data collection and processing increases
Solution Approach 1:
The system segments the voice bank creation process into distinct components: voice data collection from multiple donors, phoneme extraction and labeling, voice quality assessment, and TTS model training. This segmentation allows each component to be optimized independently while maintaining overall system reliability.
Solution Approach 2:
The voice bank system is designed to handle multiple types of voice data from diverse donors simultaneously, processing different languages, accents, and speech patterns through a unified framework. This multi-functional approach improves voice bank quality without requiring separate systems for each donor type.
2Quantity of substance
If extensive voice data is collected to represent all speech sounds, then the completeness of phoneme coverage is improved, but the time and resources required for data collection increase
Solution Approach 1:
The system performs preliminary phoneme extraction and labeling from raw voice data before full processing. By pre-identifying and organizing phoneme units from diverse donors, the system reduces the time required for subsequent TTS model training and ensures comprehensive phoneme coverage is achieved efficiently.
Solution Approach 2:
The system creates phoneme representations and voice models that can be copied and reused across different TTS applications. Once phoneme data is collected and processed from diverse donors, these standardized representations can be replicated for multiple voice synthesis tasks, reducing redundant data collection time.
3Measurement precision
If high quality voice data is collected from individuals with limited speaking ability, then the personalization accuracy of TTS voices is improved, but the difficulty of data collection increases
Solution Approach 1:
The system incorporates automated voice quality assessment and phoneme identification that operates without requiring expert annotators. The automated processing extracts phoneme data from voice recordings independently, reducing the difficulty of collecting high-quality data from donors with limited speaking ability while maintaining TTS voice accuracy.
Solution Approach 2:
Manual voice analysis and phoneme labeling by experts is replaced with automated speech processing algorithms. This substitution reduces the difficulty of data collection from vulnerable populations by eliminating the need for complex manual intervention while preserving measurement precision for TTS synthesis.
Data Source
AI summary
Voice data may be collected by a plurality of voice donors and stored in a voice bank. A voice donor may authenticate to a voice collection system to start a session to provide voice data. During the voice collection session, the voice donor may be presented with a sequence of prompts to speak and voice data may be transferred to a server. The received voice data may be processed to determine the speech units spoken by the voice donor and a count of speech units received from the voice donor may be updated. Feedback may be provided to the voice donor indicating, for example, a progress of the voice collection, a quality level of the voice data, or information about speech unit counts. The voice bank may be used to create TTS voices for voice recipients, create a model of voice aging, or for other applications.


