Outlier Detection for Speech Synthesis Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech systems face decreased synthesis quality due to mis-alignments and mis-pronunciations caused by poor alignments between audio and transcription, leading to robotic-sounding speech and spectral variations.

Innovation Solution

The system employs fundamental frequency and group delay based outlier detection methods to identify and remove mis-alignments and mis-pronunciations during the training phase of Hidden Markov Model-based speech synthesis, ensuring accurate phoneme and syllable alignments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated alignment methods are used in text-to-speech systems, then processing efficiency is improved, but alignment accuracy deteriorates due to mis-alignments and mis-pronunciations

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidalignment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary outlier detection and removal during the training phase before synthesis. By identifying and removing mis-aligned audio-transcription pairs using fundamental frequency and group delay analysis beforehand, the system ensures high alignment accuracy in the synthesized output without compromising processing efficiency during actual use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces fundamental frequency analysis and group delay calculation as intermediary steps between automated alignment and synthesis. These intermediary analyses serve as quality filters that detect mis-alignments without requiring manual intervention, thus maintaining processing efficiency while improving alignment accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If training data with poor alignments is used, then training speed is improved, but synthesis quality deteriorates due to robotic-sounding speech and spectral variations

Engineering Contradiction:
Improvetraining speedVSAvoidsynthesis quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary quality assessment of training data using fundamental frequency and group delay analysis before the training process. By identifying and removing outliers in advance, the system ensures that only high-quality aligned data is used for training, thereby maintaining synthesis quality without significantly impacting training speed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the parameters used for data selection by analyzing fundamental frequency contours and group delay characteristics. This parameter-based filtering approach automatically identifies mis-aligned segments without requiring manual review, thus maintaining training speed while improving synthesis quality by excluding poor-quality data.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If outlier detection and removal processes are added, then synthesis quality is improved, but system complexity increases

Engineering Contradiction:
Improvesynthesis qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs self-service quality control by automatically detecting outliers using fundamental frequency and group delay analysis without requiring external manual intervention. The system identifies its own mis-aligned data and removes it autonomously, thereby improving synthesis quality while adding minimal operational complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces manual alignment verification with automated fundamental frequency and group delay analysis. This substitution eliminates the need for human reviewers while maintaining high detection accuracy, thus improving synthesis quality without proportionally increasing system complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP3308378B1System and method for outlier identification to remove poor alignments in speech synthesis
Publication Date: 2019.09.11 INTERACTIVE INTELLIGENCE GROUP INC
  • EP3308378B1 patent drawingFigure 1a~1b
  • EP3308378B1 patent drawingFigure 1c
  • EP3308378B1 patent drawingFigure 2a~2b

AI summary

A system and method are presented for outlier identification to remove poor alignments in speech synthesis. The quality of the output of a text-to-speech system directly depends on the accuracy of alignments of a speech utterance. The identification of mis-alignments and mis-pronunciations from automated alignments may be made based on fundamental frequency methods and group delay based outlier methods. The identification of these outliers allows for their removal, which improves the synthesis quality of the text-to-speech system.