Concatenative Text-To-Speech Database Reorganization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Concatenative Text-To-Speech (CTTS) systems require time-consuming and expensive updates to adapt to new domains or applications, necessitating additional human speech recordings and expert phonetic-linguistic skills, limiting their ability to produce high-quality synthetic speech across varied contexts.
Innovation Solution
The method involves reorganizing speech segments in the base database using application-specific decision tree modifications, creating new context classes, and applying a weighting function to improve runtime segment selection, allowing for adaptive updates without additional recordings, thus enabling efficient adaptation to arbitrary domains and applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If additional human speech recordings are performed to adapt CTTS systems to new domains, then speech quality for new domains is improved, but time consumption and cost increase
Solution Approach 1:
The patent creates virtual speech recordings by copying and transforming existing speech data through acoustic context manipulation. Instead of recording new speech, the system generates synthetic speech samples by modifying acoustic contexts of existing recordings, thereby achieving domain adaptation without additional human recording time.
Solution Approach 2:
The patent changes acoustic context parameters of existing speech segments to adapt them to new domains. By modifying parameters such as acoustic context labels and reorganizing speech segments according to new context classifications, the system achieves domain-specific adaptation without requiring new recordings.
2Reliability
If expert phonetic-linguistic skills are required for CTTS system updates, then speech synthesis quality is maintained, but ease of operation deteriorates
Solution Approach 1:
The patent enables the CTTS system to perform self-adaptation through automated acoustic context reorganization. The system automatically analyzes new domain data, identifies required context classes, and reorganizes speech segments without requiring expert phonetic-linguistic intervention, making updates accessible to non-experts while maintaining quality.
3Adaptability or versatility
If the speech database is reorganized with application-specific context classes, then adaptability to new domains is improved, but device complexity increases
Solution Approach 1:
The patent segments the speech database into distinct acoustic context classes that can be independently organized and managed. By dividing the database into context-specific segments rather than treating it as a monolithic structure, the system achieves better domain adaptability while maintaining manageable complexity through modular organization.
4Reliability
If a large number of synthesis units are stored to cover various applications, then speech quality across domains is improved, but productivity during runtime search deteriorates
Solution Approach 1:
The patent performs preliminary organization of speech segments into application-specific acoustic context classes before runtime. By pre-segmenting and indexing speech units according to their acoustic contexts and target applications, the system enables fast retrieval during runtime without compromising speech quality, resolving the contradiction between comprehensive coverage and search efficiency.
Data Source
AI summary
The present invention relates to computer-generated text-to-speech conversion. It relates in particular to a method and system for updating a Concatenative Text-To-Speech (CTTS) system with a speech database from a base version to a new version. The present invention performs an application-specific re-organization of a synthesizer's speech database by means of certain decision tree modifications. By that reorganization, certain synthesis units are made available for the new application, which are not available in prior art without a new speech session. This allows the creation of application-specific synthesizers with improved output speech quality for arbitrary domains and applications at very low cost.


