Speech Synthesis Using Diphone Segmentation for Intonation Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis methods struggle to produce high-quality, natural-sounding speech with precise intonation reproduction due to limitations in database capacity and computational processing, leading to insufficient prosodic variability and intonation overtones.
Innovation Solution
A method of text-based speech synthesis that determines physical parameters of target speech sounds based on intonation, searching for the most suitable sounds in a database that match these parameters, and using allophones as the minimal units for synthesis, incorporating linguistic and intonation models to ensure accurate intonation reproduction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If phonemes are used as synthesis units, then database capacity is reduced, but speech quality deteriorates due to coarticulation boundary effects
Solution Approach 1:
The patent segments speech into diphones (pairs of adjacent phonemes) rather than individual phonemes. This segmentation approach reduces the number of connection points and boundary effects compared to phoneme-level synthesis, while maintaining manageable database size. Each diphone captures the transition between two phonemes, preserving coarticulation information without requiring separate storage for every possible phoneme combination.
Solution Approach 2:
The patent pre-processes and stores diphone units with their coarticulation characteristics already embedded. By preparing these intermediate units in advance with smoothed transitions, the system avoids the need for complex real-time coarticulation modeling during synthesis, thereby maintaining speech quality while controlling database requirements.
2Manufacturing precision
If diphones are used as synthesis units, then coarticulation information is preserved, but the number of connection points increases requiring complex smoothing algorithms
Solution Approach 1:
The patent applies smoothing and coarticulation processing during the diphone recording and storage phase rather than during real-time synthesis. By pre-smoothing the transitions and embedding coarticulation characteristics in the stored diphone units, the system eliminates the need for complex smoothing algorithms during the actual speech generation process.
3Quantity of substance
If single variation of each diphone is stored, then database capacity is reduced, but prosodic variability is lost requiring additional control techniques
Solution Approach 1:
The patent stores diphones with multiple variations of acoustic parameters including duration, pitch contour, and intensity. Instead of storing multiple complete diphone recordings, the system stores base diphone units with parameter sets that allow dynamic adjustment of prosodic features during synthesis, achieving variability without proportional increase in database size.
4Device complexity
If speech units of natural speech are used, then fewer connection points are required, but the number of units increases requiring larger database capacity
Solution Approach 1:
The patent uses diphone-level segmentation as an intermediate approach between phoneme and word-level synthesis. This segmentation creates sufficiently long units to reduce connection points compared to phoneme synthesis, while keeping the database manageable by reusing diphone units across different contexts through parameter variation rather than storing complete word recordings.
Data Source
Figure 1

AI summary
The present invention relates to a method of text-based speech synthesis, wherein at least one portion of a text is specified; the intonation of each portion is determined; target speech sounds are associated with each portion; physical parameters of the target speech sounds are determined; speech sounds most similar in terms of the physical parameters to the target speech sounds are found in a speech database; and speech is synthesized as a sequence of the found speech sounds. The physical parameters of said target speech sounds are determined in accordance with the determined intonation. The present method, when used in a speech synthesizer, allows improved quality of synthesized speech due to precise reproduction of intonation.