Reduced Script for Concatenative TTS Voice Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing concatenative text-to-speech (TTS) voice generation methods require extensive recording time and costs due to the need for a large script and uniform recording conditions to achieve full phonetic coverage, making it prohibitive for creating custom TTS voices, especially when a large vocabulary is needed.
Innovation Solution
A reduced script is used, which includes a minimal set of phrases that, when read by a voice talent, generates a reduced set of speech assets that, when combined with pre-recorded assets, provides full phonetic coverage for a TTS voice, thereby minimizing recording time and costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a large script is used to achieve full phonetic coverage for unit selection synthesis, then the vocabulary breadth and naturalness of the TTS voice is improved, but the recording time and cost increase significantly
Solution Approach 1:
The patent segments the speech corpus into different types of assets: pre-recorded assets from existing databases and reduced assets from new recordings. This segmentation allows the system to use pre-recorded assets for common phonetic contexts and only record new assets for missing or rare contexts, significantly reducing the total recording time while maintaining full phonetic coverage
Solution Approach 2:
The patent performs preliminary analysis to identify which phonetic assets are already available in pre-recorded databases before initiating new recording sessions. This preliminary action allows the system to minimize the recording script to only the necessary phrases that fill gaps in existing coverage, rather than recording all possible phrases from scratch
2Reliability
If a large script is used to ensure complete phonetic coverage, then the quality and naturalness of the TTS voice is improved, but the recording cost increases
Solution Approach 1:
The patent changes the parameter of script size by using a reduced script that is specifically designed to cover only the missing phonetic contexts. This parameter change reduces recording costs while maintaining complete phonetic coverage through the combination of pre-recorded and reduced assets, rather than requiring a comprehensive script covering all possible contexts
3Manufacturing precision
If uniform recording conditions are enforced to generate a clean speech corpus, then the quality of the TTS voice is improved, but the complexity and cost of the recording process increases
Solution Approach 1:
The patent uses pre-recorded speech assets from existing databases as templates or copies for common phonetic contexts. These pre-recorded assets were captured under controlled uniform conditions, and the system reuses them rather than requiring new recordings under identical conditions, thereby reducing the complexity of maintaining uniform recording conditions while preserving speech corpus quality
Data Source
AI summary
The present invention discloses a system and a method for creating a reduced script, which is read by a voice talent to create a concatenative text-to-speech (TTS) voice. The method can automatically process pre-recorded audio to derive speech assets for a concatenative TTS voice. The pre-recording audio can include sets of recorded phrases used by a speech user interface (Sill). A set of unfulfilled speech assets needed for foil phonetic coverage of the concatenative TTS voice can be determined. A reduced script can be constructed that includes a set of phrases, which when read by a voice talent result in a reduced corpus. When the reduced corpus is automatically processed, a reduced set of speech assets result. The reduced set includes each of the unfulfilled speech assets. When this reduced corpus is combined with existing speech assets the result will be a voice with a complete set of speech assets.


