Outlier Detection for Speech Synthesis Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text-to-speech systems face decreased synthesis quality due to mis-alignments and mis-pronunciations caused by poor alignments between audio and transcription, leading to robotic-sounding speech and spectral variations.
Innovation Solution
The system employs fundamental frequency and group delay based outlier detection methods to identify and remove mis-alignments and mis-pronunciations during the training phase of Hidden Markov Model-based speech synthesis, ensuring accurate phoneme and syllable alignments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated alignment methods are used in text-to-speech systems, then processing efficiency is improved, but alignment accuracy deteriorates due to mis-alignments and mis-pronunciations
Solution Approach 1:
The system performs preliminary outlier detection and removal during the training phase before synthesis. By identifying and removing mis-aligned audio-transcription pairs using fundamental frequency and group delay analysis beforehand, the system ensures high alignment accuracy in the synthesized output without compromising processing efficiency during actual use.
Solution Approach 2:
The system introduces fundamental frequency analysis and group delay calculation as intermediary steps between automated alignment and synthesis. These intermediary analyses serve as quality filters that detect mis-alignments without requiring manual intervention, thus maintaining processing efficiency while improving alignment accuracy.
2Productivity
If training data with poor alignments is used, then training speed is improved, but synthesis quality deteriorates due to robotic-sounding speech and spectral variations
Solution Approach 1:
The system performs preliminary quality assessment of training data using fundamental frequency and group delay analysis before the training process. By identifying and removing outliers in advance, the system ensures that only high-quality aligned data is used for training, thereby maintaining synthesis quality without significantly impacting training speed.
Solution Approach 2:
The system changes the parameters used for data selection by analyzing fundamental frequency contours and group delay characteristics. This parameter-based filtering approach automatically identifies mis-aligned segments without requiring manual review, thus maintaining training speed while improving synthesis quality by excluding poor-quality data.
3Reliability
If outlier detection and removal processes are added, then synthesis quality is improved, but system complexity increases
Solution Approach 1:
The system performs self-service quality control by automatically detecting outliers using fundamental frequency and group delay analysis without requiring external manual intervention. The system identifies its own mis-aligned data and removes it autonomously, thereby improving synthesis quality while adding minimal operational complexity.
Solution Approach 2:
The system replaces manual alignment verification with automated fundamental frequency and group delay analysis. This substitution eliminates the need for human reviewers while maintaining high detection accuracy, thus improving synthesis quality without proportionally increasing system complexity.
Data Source
Figure 1a~1b
Figure 1c
Figure 2a~2b
AI summary
A system and method are presented for outlier identification to remove poor alignments in speech synthesis. The quality of the output of a text-to-speech system directly depends on the accuracy of alignments of a speech utterance. The identification of mis-alignments and mis-pronunciations from automated alignments may be made based on fundamental frequency methods and group delay based outlier methods. The identification of these outliers allows for their removal, which improves the synthesis quality of the text-to-speech system.