Speech Synthesis Device Phoneme Boundary Update

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis methods using statistical models, such as HMM and MSD-HMM, face challenges in representing phoneme durations shorter than the analytical frame duration, as increasing analysis resolution leads to increased computational costs and difficulties in reproducing short phoneme durations, especially when trained with data from rapid speech.

Innovation Solution

A speech synthesis device and method that utilize a voiced utterance likelihood index to update phoneme boundary positions between neighboring phonemes, allowing for shorter phoneme durations by determining the degree of voiced utterance likelihood at each state and adjusting the phoneme boundaries accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Duration of action of moving object

If the resolution of analysis is increased by narrowing frame width, then the shortest phoneme duration can be shortened, but the calculation amount increases

Engineering Contradiction:
Improveshortest phoneme durationVSAvoidcalculation amount
Core Design Contradiction:
Duration of action of moving objectVSPower

Solution Approach 1:

The phoneme is divided into multiple states (e.g., 5 states), and the phoneme boundary is determined by selecting the appropriate state boundary based on voiced/unvoiced likelihood. This segmentation allows the system to represent short phonemes without requiring extremely narrow frame widths, thus avoiding excessive computational complexity while achieving sub-frame duration representation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention changes the parameter used for boundary determination from fixed frame boundaries to state boundaries weighted by voiced/unvoiced likelihood. By introducing this probabilistic parameter, the system can represent phoneme durations shorter than the analytical frame duration without increasing calculation complexity, as it reuses existing HMM state structures rather than requiring finer temporal sampling.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If training data from rapid speech is used, then short phoneme durations can be captured, but it becomes difficult to reproduce accurate phoneme durations

Engineering Contradiction:
Improvephoneme duration reproduction accuracyVSAvoidapplicability to different speech rates
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system uses voiced/unvoiced likelihood as feedback to determine phoneme boundaries. By calculating the likelihood at each state and selecting the boundary where this likelihood changes, the system can accurately reproduce phoneme durations regardless of the speech rate in the training data. This feedback mechanism adapts to different speech rates without requiring rate-specific training data.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The invention introduces voiced/unvoiced likelihood as a new parameter for boundary determination, changing from fixed duration-based boundaries to probability-based boundaries. This parameter change enables the system to handle varying speech rates effectively, as the likelihood-based approach naturally adapts to different temporal scales without sacrificing precision.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9520125B2Speech synthesis device, speech synthesis method, and speech synthesis program
Publication Date: 2016.12.13 NEC CORP
  • US9520125B2 patent drawing
  • US9520125B2 patent drawing
  • US9520125B2 patent drawing

AI summary

There are provided a speech synthesis device, a speech synthesis method and a speech synthesis program which can represent a phoneme as a duration shorter than a duration upon modeling according to a statistical method. A speech synthesis device 80 according to the present invention includes a phoneme boundary updating means 81 which, by using a voiced utterance likelihood index which is an index indicating a degree of voiced utterance likelihood of each state which represents a phoneme modeled by a statistical method, updates a phoneme boundary position which is a boundary with other phonemes neighboring to the phoneme.