Voice Data Collection Platform for High-Quality TTS Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for collecting high-quality voice data from diverse individuals for text-to-speech (TTS) systems face challenges in ensuring data quality and representation, particularly from unexperienced donors in various environments, leading to inconsistent and noisy recordings.
Innovation Solution
A system and method for collecting voice data from multiple donors using a network-based platform that provides user-friendly interfaces, authentication, and feedback mechanisms to ensure high-quality data collection, including multiple sessions and adaptive prompts to cover a wide range of phonemes and environments, with processing techniques for noise reduction and speaker recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice data is collected from a large number of diverse donors in various environments, then the versatility and representation of TTS voices is improved, but the data quality and consistency deteriorates due to noise and unexperienced donors
Solution Approach 1:
The system performs preliminary actions by providing donors with guidance instructions before recording, implementing authentication procedures in advance, and preparing adaptive prompt selections beforehand. This ensures that even unexperienced donors can provide high-quality data while maintaining diverse environmental representation.
Solution Approach 2:
The system implements feedback mechanisms by analyzing recorded voice data quality in real-time, providing immediate guidance to donors on improving their recordings, and using speaker recognition to verify authentic samples. This feedback loop maintains data quality while allowing diverse environmental collection.
2Ease of manufacture
If simple collection methods are used to ease donor participation, then the quantity of voice data increases, but the manufacturing precision and quality control of the TTS voice deteriorates
Solution Approach 1:
The system enables self-service by allowing donors to independently complete authentication, receive automated guidance instructions, record voice samples using their own devices in various environments, and obtain immediate feedback. This maintains ease of participation while ensuring quality through automated procedures.
Solution Approach 2:
The system dynamically adjusts recording parameters, prompt selections, and guidance levels based on individual donor characteristics and environmental conditions. This adaptive parameter adjustment maintains simple participation processes while ensuring high-quality data collection for TTS voice synthesis.
3Quantity of substance
If multiple recording sessions and adaptive prompts are implemented to cover wide range of phonemes, then the completeness of voice data improves, but the time and complexity of the collection process increases
Solution Approach 1:
The system dynamically adapts the recording process by adjusting prompt selections based on phoneme coverage needs, modifying session structures according to donor performance, and flexibly allocating recording time across multiple sessions. This dynamic approach ensures complete phoneme coverage while optimizing collection time.
Solution Approach 2:
The system implements periodic action by distributing voice data collection across multiple recording sessions rather than requiring all data in one session. Adaptive prompts are periodically updated based on progress, and phoneme coverage is systematically reviewed at intervals to ensure completeness without excessive time investment.
Data Source
AI summary
A voice recipient may request a text-to-speech (TTS) voice that corresponds to an age or age range. An existing TTS voice or existing voice data may be used to create a TTS voice corresponding to the requested age by encoding the voice data to voice parameter values, transforming the voice parameter values using a voice-aging model, synthesizing voice data using the transformed parameter values, and then creating a TTS voice using the transformed voice data. The voice-aging model may model how one or more voice parameters of a voice change with age and may be created from voice data stored in a voice bank.


