Outlier Identification for Speech Synthesis Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text-to-speech systems face decreased synthesis quality due to mis-alignments and mis-pronunciations in automated alignments, which affect the accuracy of fundamental frequency and duration values, leading to poor prosody and spectral variations in synthesized speech.
Innovation Solution
The system employs fundamental frequency and group delay based outlier identification methods to detect and remove mis-alignments by extracting phoneme and syllable level alignments, identifying outliers based on predetermined criteria, and discarding sentences with excessive outliers from model training, thereby improving synthesis quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated alignment methods are used in text-to-speech systems, then processing efficiency is improved, but alignment accuracy deteriorates due to mis-alignments and mis-pronunciations
Solution Approach 1:
The system performs preliminary outlier detection and removal on the training data before model training. By identifying and removing mis-aligned phoneme instances based on fundamental frequency and duration criteria beforehand, the training model learns from cleaner data, thus improving alignment accuracy without sacrificing processing efficiency during the actual speech synthesis operation
Solution Approach 2:
The patent introduces fundamental frequency and duration analysis as intermediary measures to detect mis-alignments. These intermediate metrics serve as mediators between the automated alignment process and the final speech output, allowing the system to identify and remove poor alignments based on physiological speech characteristics before training the synthesis model
2Quantity of substance
If all audio data is used for model training, then training data quantity is maximized, but synthesis quality deteriorates due to inclusion of poor alignments
Solution Approach 1:
The system discards audio instances identified as outliers based on fundamental frequency and duration analysis before model training. By removing mis-aligned phoneme instances that would otherwise degrade synthesis quality, the training process uses only high-quality data, thereby improving reliability without unnecessarily reducing the effective training data quantity
Solution Approach 2:
The patent changes the parameter quality by filtering training data based on fundamental frequency and duration parameters. Instances that fall outside acceptable ranges for these parameters are identified as outliers and removed, thus improving the overall quality of training data while maintaining sufficient quantity for effective model training
Data Source
AI summary
A system and method are presented for outlier identification to remove poor alignments in speech synthesis. The quality of the output of a text-to-speech system directly depends on the accuracy of alignments of a speech utterance. The identification of mis-alignments and mis-pronunciations from automated alignments may be made based on fundamental frequency methods and group delay based outlier methods. The identification of these outliers allows for their removal, which improves the synthesis quality of the text-to-speech system.


