Speech Database Enhancement via Phonetic Segment Substitution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Unit selection concatenative synthesis for speech synthesis is limited by the variations within the recorded voice database, making it difficult to accurately pronounce foreign words or dialects without additional expensive recordings.
Innovation Solution
A system and method that enhance a speech database by labeling and substituting audio segments with varying pronunciations from a secondary database, allowing for phonetic expansion and accurate synthesis of foreign words or dialects, using techniques such as segment substitution and speech representation models like harmonic plus noise models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If unit selection synthesis uses a fixed recorded voice database, then the synthesis quality is natural and spontaneous, but the phonetic coverage is limited to the variations within the database
Solution Approach 1:
The patent segments speech into phonetic units (phones, diphones, triphones) and organizes them in hierarchical database structures. This segmentation allows the system to selectively combine units from different recording sessions and speakers, expanding phonetic coverage while maintaining natural synthesis quality through proper unit selection and concatenation.
Solution Approach 2:
The patent creates a universal speech database structure that can serve multiple functions: storing recordings from single speakers and multiple speakers, supporting different languages and dialects, and accommodating various phonetic variations. This multi-functional database design allows the same system to maintain high synthesis quality while adapting to diverse phonetic requirements.
2Adaptability or versatility
If additional recordings are made to expand phonetic coverage, then foreign words and dialects can be pronounced accurately, but the cost and time required increase significantly
Solution Approach 1:
The patent performs preliminary organization and indexing of speech segments during database construction, creating structured metadata that enables efficient retrieval and combination of phonetic units. This preliminary action allows the system to expand phonetic coverage by integrating new recordings without requiring extensive additional processing time, as the infrastructure is already in place to handle diverse phonetic content.
Solution Approach 2:
The patent uses phonetic transcription and symbolic representation to create abstract models of speech sounds that can be replicated and combined from different sources. Instead of requiring complete new recordings for every phonetic variation, the system copies and recombines existing phonetic units based on linguistic rules and phonetic equivalence, significantly reducing the need for additional recording time.
3Adaptability or versatility
If additional recordings are made to expand phonetic coverage, then foreign words and dialects can be pronounced accurately, but the expense increases significantly
Solution Approach 1:
The patent merges speech segments from multiple recording sessions and multiple speakers into a unified database structure. By combining existing recordings rather than requiring entirely new recording projects, the system expands phonetic coverage while minimizing additional costs. The merging process integrates phonetic units from diverse sources into a cohesive database that maintains natural synthesis quality.
Solution Approach 2:
The patent changes the organizational parameters of the speech database from speaker-specific to phonetic-unit-based structures. This parameter change allows the system to reuse existing recordings in multiple contexts and combinations, extracting maximum value from existing audio assets. By organizing data around phonetic features rather than recording sessions, the system reduces the need for additional expensive recordings while expanding phonetic coverage.
Data Source
AI summary
A system, method and computer readable medium that enhances a speech database for speech synthesis is disclosed. The method may include labeling audio files in a primary speech database, identifying segments in the labeled audio files that have varying pronunciations based on language differences, identifying replacement segments in a secondary speech database, enhancing the primary speech database by substituting the identified secondary speech database segments for the corresponding identified segments in the primary speech database, and storing the enhanced primary speech database for use in speech synthesis.


