Dynamic TTS Selection via Multi-Engine Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech synthesis systems often produce discontinuities and varying quality due to differences in algorithms and settings, with no effective method to dynamically select the best output for a given text across multiple systems.

Innovation Solution

A method and system that dynamically select among text-to-speech systems by synthesizing text using multiple TTS engines, generating candidate waveforms, calculating a cost function score for each, and selecting the waveform with the lowest score as the output, allowing for automatic and adaptive speech output generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a single TTS system is used with fixed algorithm and parameters, then the system is simple and easy to operate, but the synthesis quality varies and produces discontinuities for different texts

Engineering Contradiction:
Improvesynthesis quality consistencyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple TTS systems (first TTS system, second TTS system, third TTS system) into a unified architecture where each system processes the same input text independently and produces a waveform output. The systems are merged through a common selection mechanism that evaluates and chooses from their outputs, resolving the contradiction by integrating multiple quality sources while maintaining a unified interface.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces dynamic selection among TTS systems based on text-specific performance evaluation. Instead of using a fixed system, the architecture dynamically chooses which system's output to use based on predicted quality metrics for each text, allowing the system to adapt to varying text characteristics and optimize synthesis quality for different inputs.

Inventive Principle:
Principle #15Dynamics

2Reliability

If multiple TTS systems are used to synthesize text, then the synthesis quality and naturalness improve, but the system complexity and computational cost increase

Engineering Contradiction:
Improvesynthesis qualityVSAvoidnumber of systems
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the TTS processing into independent parallel systems, where each system processes text independently through separate algorithmic paths. This segmentation allows quality evaluation of each system's output without requiring complex interactions between systems, managing complexity through modular independence while enabling quality comparison and selection.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback through quality prediction and score generation for each TTS system's output. The system evaluates each waveform's quality using predicted metrics and uses this feedback to determine the best output to select, creating a closed-loop system that continuously optimizes based on quality assessment while managing complexity through automated decision-making.

Inventive Principle:
Principle #23Feedback

3Reliability

If text is synthesized by multiple TTS systems with different algorithms, then the intelligibility and naturalness improve, but the processing time and computational resources increase

Engineering Contradiction:
ImproveintelligibilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by predicting quality scores for each TTS system's output before final selection. The system pre-evaluates each system's expected performance on the given text using quality metrics, allowing the selection to be made based on predicted outcomes rather than requiring full processing and comparison of all systems' complete outputs, thus reducing computational time while maintaining intelligibility.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7702510B2System and method for dynamically selecting among TTS systems
Publication Date: 2010.04.20 CERENCE OPERATING CO
  • US7702510B2 patent drawing
  • US7702510B2 patent drawing
  • US7702510B2 patent drawing

AI summary

Systems and methods for dynamically selecting among text-to-speech (TTS) systems. Exemplary embodiments of the systems and methods include identifying text for converting into a speech waveform, synthesizing said text by three TTS systems, generating a candidate waveform from each of the three systems, generating a score from each of the three systems, comparing each of the three scores, selecting a score based on a criteria and selecting one of the three waveforms based on the selected of the three scores.