Cohesive Script Generation for Concatenative TTS Corpora
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for creating speech corpora for concatenative Text-To-Speech synthesis are inefficient, result in scripts with grammatical errors and poor writing, and lack natural prosody, making it difficult for speakers to read and generate high-quality recordings.
Innovation Solution
An intelligent software system generates cohesive scripts by selecting phoneme sequences from a text database and using templates to create coherent text documents that include required phoneme sequences, improving fluency and prosody in speech recordings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data mining is used to generate scripts by searching through large text databases, then phonemic coverage is improved, but script coherence and reading fluency deteriorate
Solution Approach 1:
The script generation process is segmented into distinct phases: phoneme sequence identification, template selection, and coherent text assembly. This allows independent optimization of phonemic coverage and reading fluency in different stages of the process.
Solution Approach 2:
Templates serve as an intermediary between the raw data mining results and the final script. These templates provide a coherent structural framework that organizes phonemically-rich content into naturally flowing text, resolving the conflict between comprehensive phoneme coverage and readable coherence.
2Quantity of substance
If sentences are chosen independently to maximize phoneme diversity, then phonemic representation is improved, but script coherence deteriorates
Solution Approach 1:
Templates are prepared in advance with established coherent structures and thematic connections. By pre-defining these organizational frameworks, the system ensures that independently selected sentences will be integrated into a cohesive whole, maintaining script coherence while preserving phoneme diversity.
Solution Approach 2:
The system changes the organizational parameter from random independent sentence selection to template-based structured assembly. This parameter change allows sentences to be arranged according to coherent thematic and structural principles while maintaining diverse phoneme representation.
3Quantity of substance
If the script includes many sentences to ensure phoneme coverage, then phonemic completeness is improved, but recording time increases
Solution Approach 1:
Multiple phoneme sequences are merged into unified template structures that can be efficiently read in context. By combining several phoneme-rich sentences into coherent thematic units, the system reduces the total number of discrete reading segments while maintaining comprehensive phoneme coverage.
Solution Approach 2:
The system changes from exhaustive enumeration of all required phoneme sequences to selective integration of phoneme-rich content within coherent templates. This parameter change optimizes the balance between phoneme completeness and recording efficiency by including only the necessary phoneme sequences within naturally flowing text.
Data Source
AI summary
A method (and system) which autonomously generates a cohesive script from a text database for creating a speech corpus for concatenative text-to-speech, and more particularly, which generates cohesive scripts having fluency and natural prosody that can be used to generate compact text-to-speech recordings that cover a plurality of phonetic events.


