Cohesive Script Generation for Concatenative TTS Corpora

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for creating speech corpora for concatenative Text-To-Speech synthesis are inefficient, result in scripts with grammatical errors and poor writing, and lack natural prosody, making it difficult for speakers to read and generate high-quality recordings.

Innovation Solution

An intelligent software system generates cohesive scripts by selecting phoneme sequences from a text database and using templates to create coherent text documents that include required phoneme sequences, improving fluency and prosody in speech recordings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data mining is used to generate scripts by searching through large text databases, then phonemic coverage is improved, but script coherence and reading fluency deteriorate

Engineering Contradiction:
Improvephonemic coverageVSAvoidreading fluency
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The script generation process is segmented into distinct phases: phoneme sequence identification, template selection, and coherent text assembly. This allows independent optimization of phonemic coverage and reading fluency in different stages of the process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Templates serve as an intermediary between the raw data mining results and the final script. These templates provide a coherent structural framework that organizes phonemically-rich content into naturally flowing text, resolving the conflict between comprehensive phoneme coverage and readable coherence.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If sentences are chosen independently to maximize phoneme diversity, then phonemic representation is improved, but script coherence deteriorates

Engineering Contradiction:
Improvephoneme diversityVSAvoidscript coherence
Core Design Contradiction:
Quantity of substanceVSStability of the object's composition

Solution Approach 1:

Templates are prepared in advance with established coherent structures and thematic connections. By pre-defining these organizational frameworks, the system ensures that independently selected sentences will be integrated into a cohesive whole, maintaining script coherence while preserving phoneme diversity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the organizational parameter from random independent sentence selection to template-based structured assembly. This parameter change allows sentences to be arranged according to coherent thematic and structural principles while maintaining diverse phoneme representation.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If the script includes many sentences to ensure phoneme coverage, then phonemic completeness is improved, but recording time increases

Engineering Contradiction:
Improvephoneme completenessVSAvoidrecording time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

Multiple phoneme sequences are merged into unified template structures that can be efficiently read in context. By combining several phoneme-rich sentences into coherent thematic units, the system reduces the total number of discrete reading segments while maintaining comprehensive phoneme coverage.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system changes from exhaustive enumeration of all required phoneme sequences to selective integration of phoneme-rich content within coherent templates. This parameter change optimizes the balance between phoneme completeness and recording efficiency by including only the necessary phoneme sequences within naturally flowing text.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8155963B2Autonomous system and method for creating readable scripts for concatenative text-to-speech synthesis (TTS) corpora
Publication Date: 2012.04.10 CERENCE OPERATING CO
  • US8155963B2 patent drawing
  • US8155963B2 patent drawing
  • US8155963B2 patent drawing

AI summary

A method (and system) which autonomously generates a cohesive script from a text database for creating a speech corpus for concatenative text-to-speech, and more particularly, which generates cohesive scripts having fluency and natural prosody that can be used to generate compact text-to-speech recordings that cover a plurality of phonetic events.