Corpus-Based Speech Synthesis Using Diphone Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current corpus-based speech synthesis systems face challenges in achieving high-quality, natural-sounding speech due to limitations in segment selection and database scalability, leading to issues like data scarcity, increased production cycles, and computational overhead.

Innovation Solution

The approach involves using compound speech units derived from carefully selected and annotated speech data, which are added to the database to enhance segment length and quality, reducing dependency on recordings and speaker variability, and incorporating perceptual validation to improve unit selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If corpus-based unit selection synthesis is used to achieve high speech quality, then speech naturalness is improved, but database size and selection complexity increase

Engineering Contradiction:
Improvespeech qualityVSAvoidselection complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments speech into diphone units (combinations of two phonemes) as the basic building blocks. This segmentation allows the system to construct arbitrary speech from a limited set of reusable units, reducing database requirements while maintaining synthesis quality. The diphone segmentation strikes a balance between unit size and coverage, avoiding the complexity of phone-level segmentation while providing finer control than word-level units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a virtual copy mechanism where a single recorded diphone can be reused multiple times in different contexts through parameter modification. Instead of recording every possible speech variation, the system copies base diphone units and applies prosodic modifications to generate context-appropriate versions, dramatically reducing the required database size while maintaining naturalness.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If smaller speech units (phones, diphones) are used in corpus-based synthesis, then speech flexibility and quality are improved, but database size grows exponentially

Engineering Contradiction:
Improvespeech flexibilityVSAvoiddatabase size
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent makes diphone units universal by designing them to function in multiple contexts. Each diphone is annotated with contextual features (phonetic context, prosodic information) that allow it to be selected and modified for various speech situations. This multi-functionality enables a compact database to generate diverse speech output, achieving flexibility without exponential database growth.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs parameter modification to adapt stored diphone units to different speech contexts. By changing prosodic parameters (pitch, duration, energy) of recorded units rather than storing separate recordings for each context, the system achieves high speech flexibility with a limited database. This parameter-based adaptation replaces the need for exhaustive unit collection.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If canned speech synthesis with large units (phrases, words) is used, then database footprint is reduced, but speech flexibility and vocabulary are limited

Engineering Contradiction:
Improvedatabase footprintVSAvoidvocabulary flexibility
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent segments speech into diphone units, which are smaller than word-level units used in canned speech. This finer segmentation enables the construction of arbitrary words and phrases from a limited set of diphones, dramatically expanding vocabulary flexibility while keeping the database compact. The diphone level provides the right granularity for both compression and flexibility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic prosodic modification capabilities that allow stored diphone units to adapt to different speech contexts in real-time. This dynamic adjustment of pitch, duration, and energy parameters enables the system to generate natural-sounding speech for unseen inputs, achieving vocabulary flexibility comparable to much larger canned speech systems while maintaining a compact database footprint.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS7567896B2Corpus-based speech synthesis based on segment recombination
Publication Date: 2009.07.28 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7567896B2 patent drawing
  • US7567896B2 patent drawing
  • US7567896B2 patent drawing

AI summary

A system and method generate synthesized speech through concatenation of speech segments that are derived from a large prosodically-rich corpus of speech segments including using an additional dictionary of speech segment identifier sequences.