Outlier Identification for Speech Synthesis Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech systems face decreased synthesis quality due to mis-alignments and mis-pronunciations in automated alignments, which affect the accuracy of fundamental frequency and duration values, leading to poor prosody and spectral variations in synthesized speech.

Innovation Solution

The system employs fundamental frequency and group delay based outlier identification methods to detect and remove mis-alignments by extracting phoneme and syllable level alignments, identifying outliers based on predetermined criteria, and discarding sentences with excessive outliers from model training, thereby improving synthesis quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated alignment methods are used in text-to-speech systems, then processing efficiency is improved, but alignment accuracy deteriorates due to mis-alignments and mis-pronunciations

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidalignment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary outlier detection and removal on the training data before model training. By identifying and removing mis-aligned phoneme instances based on fundamental frequency and duration criteria beforehand, the training model learns from cleaner data, thus improving alignment accuracy without sacrificing processing efficiency during the actual speech synthesis operation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces fundamental frequency and duration analysis as intermediary measures to detect mis-alignments. These intermediate metrics serve as mediators between the automated alignment process and the final speech output, allowing the system to identify and remove poor alignments based on physiological speech characteristics before training the synthesis model

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If all audio data is used for model training, then training data quantity is maximized, but synthesis quality deteriorates due to inclusion of poor alignments

Engineering Contradiction:
Improvetraining data quantityVSAvoidsynthesis quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system discards audio instances identified as outliers based on fundamental frequency and duration analysis before model training. By removing mis-aligned phoneme instances that would otherwise degrade synthesis quality, the training process uses only high-quality data, thereby improving reliability without unnecessarily reducing the effective training data quantity

Inventive Principle:
Principle #34Discarding and recovering

Solution Approach 2:

The patent changes the parameter quality by filtering training data based on fundamental frequency and duration parameters. Instances that fall outside acceptable ranges for these parameters are identified as outliers and removed, thus improving the overall quality of training data while maintaining sufficient quantity for effective model training

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10497362B2System and method for outlier identification to remove poor alignments in speech synthesis
Publication Date: 2019.12.03 GENESYS CLOUD SERVICES INC
  • US10497362B2 patent drawing
  • US10497362B2 patent drawing
  • US10497362B2 patent drawing

AI summary

A system and method are presented for outlier identification to remove poor alignments in speech synthesis. The quality of the output of a text-to-speech system directly depends on the accuracy of alignments of a speech utterance. The identification of mis-alignments and mis-pronunciations from automated alignments may be made based on fundamental frequency methods and group delay based outlier methods. The identification of these outliers allows for their removal, which improves the synthesis quality of the text-to-speech system.