Speech Suitability Prediction for TTS Model Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In Text-to-Speech Synthesis, the hybrid Multi-Form Segment approach faces challenges in maintaining natural speech quality due to differences in voice character between concatenated template and model segments, and the speaker selection process for statistical TTS systems is labor-intensive and time-consuming.
Innovation Solution
A system that automatically determines the suitability of a speech signal for statistical modeling by estimating temporal stationarity, allowing for dynamic selection of segment representation type and efficient speaker selection based on statistical modelability scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If template segments and model segments are concatenated in MFS synthesis, then system flexibility and natural speech quality are improved, but voice character inconsistency between segments deteriorates human perception
Solution Approach 1:
The patent introduces an acoustic model as an intermediary to evaluate and select template segments. The acoustic model computes similarity scores between template segments and target speech, acting as a mediator to ensure voice character consistency while maintaining system flexibility in segment selection
Solution Approach 2:
The system implements feedback through acoustic evaluation where the acoustic model continuously assesses template segments and provides similarity scores. This feedback mechanism allows the system to adjust segment selection dynamically, ensuring consistent voice character across concatenated segments while maintaining adaptability
2Weight of stationary object
If template segments are replaced with model segments in MFS synthesis, then system footprint is reduced, but speech naturalness deteriorates
Solution Approach 1:
The patent applies local quality by selectively replacing template segments with model-generated segments based on acoustic similarity evaluation. Only segments with sufficient acoustic match are replaced, maintaining naturalness in critical regions while reducing system footprint through selective model segment usage
Solution Approach 2:
The system changes the parameter of segment representation dynamically - using template segments when acoustic similarity is high and model segments when appropriate. This parameter switching allows optimization of both system footprint and speech naturalness based on acoustic evaluation results
3Reliability
If manual speaker selection is performed for statistical TTS system build, then speaker quality can be optimized, but time and labor consumption increase significantly
Solution Approach 1:
The patent implements self-service by enabling the system to automatically evaluate and select speakers based on acoustic modelability criteria. The acoustic model autonomously assesses speech recordings and identifies suitable speakers without human intervention, maintaining quality optimization while eliminating time-consuming manual selection
Solution Approach 2:
The system replaces the mechanical process of manual speaker evaluation with an automated acoustic modeling approach. The acoustic model computationally assesses speaker suitability based on statistical properties, substituting human expert evaluation with automated analysis to reduce time and labor while maintaining or improving selection quality
4Reliability
If extensive speech recording is performed for statistical TTS model training, then model quality can be improved, but recording time and resources increase
Solution Approach 1:
The patent applies partial action by using a limited subset of speech recordings for statistical model training rather than extensive recordings. The acoustic model identifies and utilizes only the most informative segments for training, achieving good model quality with reduced recording time and resources
Data Source
AI summary
An embodiment according to the invention provides a capability of automatically predicting how favorable a given speech signal is for statistical modeling, which is advantageous in a variety of different contexts. In Multi-Form Segment (MFS) synthesis, for example, an embodiment according to the invention uses prediction capability to provide an automatic acoustic driven template versus model decision maker with an output quality that is high, stable and depends gradually on the system footprint. In speaker selection for a statistical Text-to-Speech synthesis (TTS) system build, as another example context, an embodiment according to the invention enables a fast selection of the most appropriate speaker among several available ones for the full voice dataset recording and preparation, based on a small amount of recorded speech material.


