Speech Synthesis Device Phoneme Boundary Update
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis methods using statistical models, such as HMM and MSD-HMM, face challenges in representing phoneme durations shorter than the analytical frame duration, as increasing analysis resolution leads to increased computational costs and difficulties in reproducing short phoneme durations, especially when trained with data from rapid speech.
Innovation Solution
A speech synthesis device and method that utilize a voiced utterance likelihood index to update phoneme boundary positions between neighboring phonemes, allowing for shorter phoneme durations by determining the degree of voiced utterance likelihood at each state and adjusting the phoneme boundaries accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Duration of action of moving object
If the resolution of analysis is increased by narrowing frame width, then the shortest phoneme duration can be shortened, but the calculation amount increases
Solution Approach 1:
The phoneme is divided into multiple states (e.g., 5 states), and the phoneme boundary is determined by selecting the appropriate state boundary based on voiced/unvoiced likelihood. This segmentation allows the system to represent short phonemes without requiring extremely narrow frame widths, thus avoiding excessive computational complexity while achieving sub-frame duration representation.
Solution Approach 2:
The invention changes the parameter used for boundary determination from fixed frame boundaries to state boundaries weighted by voiced/unvoiced likelihood. By introducing this probabilistic parameter, the system can represent phoneme durations shorter than the analytical frame duration without increasing calculation complexity, as it reuses existing HMM state structures rather than requiring finer temporal sampling.
2Manufacturing precision
If training data from rapid speech is used, then short phoneme durations can be captured, but it becomes difficult to reproduce accurate phoneme durations
Solution Approach 1:
The system uses voiced/unvoiced likelihood as feedback to determine phoneme boundaries. By calculating the likelihood at each state and selecting the boundary where this likelihood changes, the system can accurately reproduce phoneme durations regardless of the speech rate in the training data. This feedback mechanism adapts to different speech rates without requiring rate-specific training data.
Solution Approach 2:
The invention introduces voiced/unvoiced likelihood as a new parameter for boundary determination, changing from fixed duration-based boundaries to probability-based boundaries. This parameter change enables the system to handle varying speech rates effectively, as the likelihood-based approach naturally adapts to different temporal scales without sacrificing precision.
Data Source
AI summary
There are provided a speech synthesis device, a speech synthesis method and a speech synthesis program which can represent a phoneme as a duration shorter than a duration upon modeling according to a statistical method. A speech synthesis device 80 according to the present invention includes a phoneme boundary updating means 81 which, by using a voiced utterance likelihood index which is an index indicating a degree of voiced utterance likelihood of each state which represents a phoneme modeled by a statistical method, updates a phoneme boundary position which is a boundary with other phonemes neighboring to the phoneme.


