Speech Database Enhancement via Phonetic Segment Substitution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Unit selection concatenative synthesis for speech synthesis is limited by the variations within the recorded voice database, making it difficult to accurately pronounce foreign words or dialects without additional expensive recordings.

Innovation Solution

A system and method that enhance a speech database by labeling and substituting audio segments with varying pronunciations from a secondary database, allowing for phonetic expansion and accurate synthesis of foreign words or dialects, using techniques such as segment substitution and speech representation models like harmonic plus noise models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If unit selection synthesis uses a fixed recorded voice database, then the synthesis quality is natural and spontaneous, but the phonetic coverage is limited to the variations within the database

Engineering Contradiction:
Improvesynthesis qualityVSAvoidphonetic coverage
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments speech into phonetic units (phones, diphones, triphones) and organizes them in hierarchical database structures. This segmentation allows the system to selectively combine units from different recording sessions and speakers, expanding phonetic coverage while maintaining natural synthesis quality through proper unit selection and concatenation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal speech database structure that can serve multiple functions: storing recordings from single speakers and multiple speakers, supporting different languages and dialects, and accommodating various phonetic variations. This multi-functional database design allows the same system to maintain high synthesis quality while adapting to diverse phonetic requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If additional recordings are made to expand phonetic coverage, then foreign words and dialects can be pronounced accurately, but the cost and time required increase significantly

Engineering Contradiction:
Improvephonetic coverageVSAvoidrecording time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary organization and indexing of speech segments during database construction, creating structured metadata that enables efficient retrieval and combination of phonetic units. This preliminary action allows the system to expand phonetic coverage by integrating new recordings without requiring extensive additional processing time, as the infrastructure is already in place to handle diverse phonetic content.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses phonetic transcription and symbolic representation to create abstract models of speech sounds that can be replicated and combined from different sources. Instead of requiring complete new recordings for every phonetic variation, the system copies and recombines existing phonetic units based on linguistic rules and phonetic equivalence, significantly reducing the need for additional recording time.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If additional recordings are made to expand phonetic coverage, then foreign words and dialects can be pronounced accurately, but the expense increases significantly

Engineering Contradiction:
Improvephonetic coverageVSAvoidcost
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent merges speech segments from multiple recording sessions and multiple speakers into a unified database structure. By combining existing recordings rather than requiring entirely new recording projects, the system expands phonetic coverage while minimizing additional costs. The merging process integrates phonetic units from diverse sources into a cohesive database that maintains natural synthesis quality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent changes the organizational parameters of the speech database from speaker-specific to phonetic-unit-based structures. This parameter change allows the system to reuse existing recordings in multiple contexts and combinations, extracting maximum value from existing audio assets. By organizing data around phonetic features rather than recording sessions, the system reduces the need for additional expensive recordings while expanding phonetic coverage.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8977552B2Method and system for enhancing a speech database
Publication Date: 2015.03.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8977552B2 patent drawing
  • US8977552B2 patent drawing
  • US8977552B2 patent drawing

AI summary

A system, method and computer readable medium that enhances a speech database for speech synthesis is disclosed. The method may include labeling audio files in a primary speech database, identifying segments in the labeled audio files that have varying pronunciations based on language differences, identifying replacement segments in a secondary speech database, enhancing the primary speech database by substituting the identified secondary speech database segments for the corresponding identified segments in the primary speech database, and storing the enhanced primary speech database for use in speech synthesis.