Custom Text-to-Speech Voice Generation Using Neural Network Copying
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing process of creating custom text-to-speech voices for spoken dialog systems is costly, time-consuming, and requires significant human expertise and effort, limiting the deployment of spoken dialog services due to high costs and inflexibility.
Innovation Solution
A method to automatically generate high-quality, application-dependent custom text-to-speech voices with minimal human interaction, using a pre-existing inventory of speech units and active learning to enhance quality, which reduces recording time and cost by leveraging pre-existing text data sources and statistical adaptation algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional speech corpus recording method is used, then voice quality is improved, but time consumption and cost increase significantly
Solution Approach 1:
The patent creates a virtual copy of the target speaker's voice by training a neural network on audio data from a different speaker who reads the same text. This allows the system to generate synthetic speech that mimics the target speaker's voice characteristics without requiring actual recording of the target speaker, thereby eliminating the time-consuming recording process while maintaining voice quality.
Solution Approach 2:
The patent replaces the mechanical recording and manual processing system with an automated neural network-based synthesis system. The neural network learns the relationship between text, acoustic features, and pronunciation patterns, then automatically generates synthetic speech without human intervention, substituting manual recording processes with computational modeling.
2Manufacturing precision
If traditional speech corpus recording method is used, then voice quality is improved, but cost increases significantly
Solution Approach 1:
The system creates a virtual voice copy using neural networks trained on publicly available audio data, eliminating the need for expensive professional recording sessions. The neural network models can be trained using free or low-cost datasets, significantly reducing the cost of creating customized TTS voices while maintaining high quality.
Solution Approach 2:
The system enables self-service voice generation where users can create customized TTS voices without requiring professional recording studios or expert annotators. The automated neural network process allows individuals to generate their own voice models using minimal input data, eliminating the need for expensive human expertise in the process.
3Ease of manufacture
If recorded prompts approach is used, then cost is reduced, but flexibility and adaptability decrease
Solution Approach 1:
The patent creates a dynamic TTS system where the neural network can adapt to different speaking styles, accents, and contexts by training on diverse audio data. The model learns the underlying patterns of speech and can generate appropriate responses for various scenarios, providing flexibility and adaptability without requiring separate recorded prompts for each situation.
Solution Approach 2:
The system changes the parameters of voice generation by adjusting the neural network's training data and model configurations to match different speaker characteristics and domains. This allows the same base model to adapt to various applications and contexts by modifying training parameters rather than requiring separate recorded prompts for each scenario.
Data Source
AI summary
A system and method are disclosed for generating customized text-to-speech voices for a particular application. The method comprises generating a custom text-to-speech voice by selecting a voice for generating a custom text-to-speech voice associated with a domain, collecting text data associated with the domain from a pre-existing text data source and using the collected text data, generating an in-domain inventory of synthesis speech units by selecting speech units appropriate to the domain via a search of a pre-existing inventory of synthesis speech units, or by recording the minimal inventory for a selected level of synthesis quality. The text-to-speech custom voice for the domain is generated utilizing the in-domain inventory of synthesis speech units. Active learning techniques may also be employed to identify problem phrases wherein only a few minutes of recorded data is necessary to deliver a high quality TTS custom voice.


