Speech Synthesis Using Diphone Segmentation for Intonation Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis methods struggle to produce high-quality, natural-sounding speech with precise intonation reproduction due to limitations in database capacity and computational processing, leading to insufficient prosodic variability and intonation overtones.

Innovation Solution

A method of text-based speech synthesis that determines physical parameters of target speech sounds based on intonation, searching for the most suitable sounds in a database that match these parameters, and using allophones as the minimal units for synthesis, incorporating linguistic and intonation models to ensure accurate intonation reproduction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If phonemes are used as synthesis units, then database capacity is reduced, but speech quality deteriorates due to coarticulation boundary effects

Engineering Contradiction:
Improvedatabase capacityVSAvoidspeech quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments speech into diphones (pairs of adjacent phonemes) rather than individual phonemes. This segmentation approach reduces the number of connection points and boundary effects compared to phoneme-level synthesis, while maintaining manageable database size. Each diphone captures the transition between two phonemes, preserving coarticulation information without requiring separate storage for every possible phoneme combination.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent pre-processes and stores diphone units with their coarticulation characteristics already embedded. By preparing these intermediate units in advance with smoothed transitions, the system avoids the need for complex real-time coarticulation modeling during synthesis, thereby maintaining speech quality while controlling database requirements.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If diphones are used as synthesis units, then coarticulation information is preserved, but the number of connection points increases requiring complex smoothing algorithms

Engineering Contradiction:
Improvecoarticulation informationVSAvoidsmoothing algorithms
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies smoothing and coarticulation processing during the diphone recording and storage phase rather than during real-time synthesis. By pre-smoothing the transitions and embedding coarticulation characteristics in the stored diphone units, the system eliminates the need for complex smoothing algorithms during the actual speech generation process.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If single variation of each diphone is stored, then database capacity is reduced, but prosodic variability is lost requiring additional control techniques

Engineering Contradiction:
Improvedatabase capacityVSAvoidprosodic variability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent stores diphones with multiple variations of acoustic parameters including duration, pitch contour, and intensity. Instead of storing multiple complete diphone recordings, the system stores base diphone units with parameter sets that allow dynamic adjustment of prosodic features during synthesis, achieving variability without proportional increase in database size.

Inventive Principle:
Principle #35Parameter changes

4Device complexity

If speech units of natural speech are used, then fewer connection points are required, but the number of units increases requiring larger database capacity

Engineering Contradiction:
Improvenumber of connection pointsVSAvoiddatabase capacity
Core Design Contradiction:
Device complexityVSQuantity of substance

Solution Approach 1:

The patent uses diphone-level segmentation as an intermediate approach between phoneme and word-level synthesis. This segmentation creates sufficiently long units to reduce connection points compared to phoneme synthesis, while keeping the database manageable by reusing diphone units across different contexts through parameter variation rather than storing complete word recordings.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP2462586B1A method of speech synthesis
Publication Date: 2017.08.02 SPEECH TECH CENT
  • EP2462586B1 patent drawingFigure 1
  • EP2462586B1 patent drawing
  • EP2462586B1 patent drawing

AI summary

The present invention relates to a method of text-based speech synthesis, wherein at least one portion of a text is specified; the intonation of each portion is determined; target speech sounds are associated with each portion; physical parameters of the target speech sounds are determined; speech sounds most similar in terms of the physical parameters to the target speech sounds are found in a speech database; and speech is synthesized as a sequence of the found speech sounds. The physical parameters of said target speech sounds are determined in accordance with the determined intonation. The present method, when used in a speech synthesizer, allows improved quality of synthesized speech due to precise reproduction of intonation.