Variable Expansion Rate Phoneme Synthesis for Natural Voice

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis technologies struggle to produce aurally natural voices when expanding phonetic pieces, as they maintain a fixed expansion and contraction rate within phonetic piece ranges, leading to unnatural voice synthesis.

Innovation Solution

A voice synthesis apparatus that adjusts the expansion rate of phonetic pieces based on phoneme types, using different expansion processes for consonant phonemes, and interpolates envelope data for voiced consonants to create synthesized phonetic piece data with varying expansion rates, ensuring natural voice synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a fixed expansion and contraction rate is maintained within a phonetic piece range, then the processing simplicity is improved, but the aural naturalness of the synthesized voice deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoidaural naturalness
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent divides the phonetic piece into multiple sections (normal part and transition part) and applies different expansion/contraction rates to each section. The transition part uses a smaller expansion rate than the normal part, creating local quality variations that match natural speech patterns while maintaining overall processing simplicity.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces dynamic expansion rate adjustment based on the position within the phonetic piece. By making the expansion rate variable (different for normal vs. transition parts) rather than fixed, the system achieves more natural speech synthesis while keeping the adjustment rule-based and computationally manageable.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If the expansion rate is uniformly applied across all sections of a phonetic piece, then the device complexity is reduced, but the synthesis accuracy deteriorates

Engineering Contradiction:
Improvedevice complexityVSAvoidsynthesis accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent implements local quality by assigning different expansion rates to different sections (normal part vs. transition part) of the phonetic piece. This section-based differentiation improves synthesis accuracy by matching natural speech characteristics without requiring complex global adjustment mechanisms.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the phonetic piece into distinct parts (normal part and transition part) with different expansion characteristics. This segmentation allows for more accurate synthesis by treating different sections differently, while the segmentation itself is simple and based on predefined boundaries.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9230537B2Voice synthesis apparatus using a plurality of phonetic piece data
Publication Date: 2016.01.05 YAMAHA CORP
  • US9230537B2 patent drawing
  • US9230537B2 patent drawing
  • US9230537B2 patent drawing

AI summary

A voice signal is synthesized using a plurality of phonetic piece data each indicating a phonetic piece containing at least two phoneme sections corresponding to different phonemes. In the apparatus, a phonetic piece adjustor forms a target section from first and second phonetic pieces so as to connect the first and second phonetic pieces to each other such that the target section includes a rear phoneme section of the first piece and a front phoneme section of the second piece, and expands the target section by a target time length to form an adjustment section such that a central part is expanded at an expansion rate higher than that of front and rear parts of the target section, to thereby create synthesized phonetic piece data having the target time length. A voice synthesizer creates a voice signal from the synthesized phonetic piece data.