Speech Suitability Prediction for TTS Model Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In Text-to-Speech Synthesis, the hybrid Multi-Form Segment approach faces challenges in maintaining natural speech quality due to differences in voice character between concatenated template and model segments, and the speaker selection process for statistical TTS systems is labor-intensive and time-consuming.

Innovation Solution

A system that automatically determines the suitability of a speech signal for statistical modeling by estimating temporal stationarity, allowing for dynamic selection of segment representation type and efficient speaker selection based on statistical modelability scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If template segments and model segments are concatenated in MFS synthesis, then system flexibility and natural speech quality are improved, but voice character inconsistency between segments deteriorates human perception

Engineering Contradiction:
Improvesystem flexibilityVSAvoidhuman perception quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an acoustic model as an intermediary to evaluate and select template segments. The acoustic model computes similarity scores between template segments and target speech, acting as a mediator to ensure voice character consistency while maintaining system flexibility in segment selection

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback through acoustic evaluation where the acoustic model continuously assesses template segments and provides similarity scores. This feedback mechanism allows the system to adjust segment selection dynamically, ensuring consistent voice character across concatenated segments while maintaining adaptability

Inventive Principle:
Principle #23Feedback

2Weight of stationary object

If template segments are replaced with model segments in MFS synthesis, then system footprint is reduced, but speech naturalness deteriorates

Engineering Contradiction:
Improvesystem footprintVSAvoidspeech naturalness
Core Design Contradiction:
Weight of stationary objectVSReliability

Solution Approach 1:

The patent applies local quality by selectively replacing template segments with model-generated segments based on acoustic similarity evaluation. Only segments with sufficient acoustic match are replaced, maintaining naturalness in critical regions while reducing system footprint through selective model segment usage

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the parameter of segment representation dynamically - using template segments when acoustic similarity is high and model segments when appropriate. This parameter switching allows optimization of both system footprint and speech naturalness based on acoustic evaluation results

Inventive Principle:
Principle #35Parameter changes

3Reliability

If manual speaker selection is performed for statistical TTS system build, then speaker quality can be optimized, but time and labor consumption increase significantly

Engineering Contradiction:
Improvespeaker qualityVSAvoidspeaker selection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the system to automatically evaluate and select speakers based on acoustic modelability criteria. The acoustic model autonomously assesses speech recordings and identifies suitable speakers without human intervention, maintaining quality optimization while eliminating time-consuming manual selection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces the mechanical process of manual speaker evaluation with an automated acoustic modeling approach. The acoustic model computationally assesses speaker suitability based on statistical properties, substituting human expert evaluation with automated analysis to reduce time and labor while maintaining or improving selection quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Reliability

If extensive speech recording is performed for statistical TTS model training, then model quality can be improved, but recording time and resources increase

Engineering Contradiction:
Improvemodel qualityVSAvoidrecording time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by using a limited subset of speech recordings for statistical model training rather than extensive recordings. The acoustic model identifies and utilizes only the most informative segments for training, achieving good model quality with reduced recording time and resources

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9484045B2System and method for automatic prediction of speech suitability for statistical modeling
Publication Date: 2016.11.01 CERENCE OPERATING CO
  • US9484045B2 patent drawing
  • US9484045B2 patent drawing
  • US9484045B2 patent drawing

AI summary

An embodiment according to the invention provides a capability of automatically predicting how favorable a given speech signal is for statistical modeling, which is advantageous in a variety of different contexts. In Multi-Form Segment (MFS) synthesis, for example, an embodiment according to the invention uses prediction capability to provide an automatic acoustic driven template versus model decision maker with an output quality that is high, stable and depends gradually on the system footprint. In speaker selection for a statistical Text-to-Speech synthesis (TTS) system build, as another example context, an embodiment according to the invention enables a fast selection of the most appropriate speaker among several available ones for the full voice dataset recording and preparation, based on a small amount of recorded speech material.