Text to Speech Synthesis with Alternative Unit Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Text-to-Speech (TTS) systems face challenges in producing optimal voice quality and speaking style due to underspecification of input text information, difficulty in detecting voice quality and speaking style changes, and sparseness in unit databases, leading to suboptimal unit selection and synthesis.

Innovation Solution

A unit selection system that generates multiple acoustic realizations of a linguistic description, allowing operators to choose the most suitable option, by deriving target unit sequences, selecting alternative unit sequences, and concatenating them into speech waveforms, with an intelligent construction algorithm to simplify the process and reduce manual iteration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional unit selection TTS systems are used, then speech synthesis can be performed, but the voice quality and speaking style are suboptimal due to underspecification of input text information

Engineering Contradiction:
Improvevoice qualityVSAvoidinput text information
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent segments the speech database into multiple overlapping units (phones, syllables, words, phrases) with different granularities. This segmentation allows the system to select units at appropriate levels of detail to capture voice quality and speaking style information that would be lost in traditional phoneme-based approaches alone.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a new dimension to unit selection by incorporating prosodic features (stress, intonation, rhythm) and speaking style attributes as explicit selection criteria. This transforms the selection process from purely phoneme-matching to multi-dimensional optimization that preserves voice quality information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If manual iteration is used to optimize speech prompts, then optimal voice quality can be achieved, but the process is time-consuming and requires expert knowledge

Engineering Contradiction:
Improvevoice qualityVSAvoidevaluation effort
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by automatically generating multiple candidate unit sequences with different prosodic and stylistic variations before presentation to the user. This pre-computation of alternatives eliminates the need for iterative manual optimization, as the best options are already prepared in advance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service by automatically evaluating and ranking multiple unit sequence candidates based on predefined criteria for voice quality and speaking style. This automated self-evaluation reduces dependence on expert operators and minimizes manual iteration time.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If multiple alternative unit sequences are generated, then operator choice is improved, but the number of alternatives may become excessive

Engineering Contradiction:
Improveoperator choiceVSAvoidnumber of alternatives
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent controls the number of alternatives by adjusting parameters such as the diversity of prosodic variations, the granularity of unit segmentation, and the selection criteria weights. These parameter changes allow the system to generate an optimal number of meaningful alternatives without overwhelming the operator.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by generating alternatives that differ specifically in prosodic and stylistic features while maintaining phonetic accuracy. This ensures that each alternative has distinct local characteristics in terms of voice quality, making the selection process more efficient and less overwhelming.

Inventive Principle:
Principle #3Local quality

4Manufacturing precision

If unit databases are expanded to include more voice quality information, then synthesis quality improves, but database sparseness increases

Engineering Contradiction:
Improvesynthesis qualityVSAvoidunit database density
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent uses a nested structure where units at different levels of granularity (phones within syllables, syllables within words, words within phrases) are organized hierarchically. This nesting allows the system to leverage information from multiple levels simultaneously, improving synthesis quality without requiring a proportionally larger database at each level.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent makes the unit database multi-functional by creating units that serve multiple purposes: they represent phonetic content, prosodic features, and speaking style simultaneously. This universality allows the same database to support various synthesis quality requirements without needing separate specialized databases for each feature.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7979280B2Text to speech synthesis
Publication Date: 2011.07.12 CERENCE OPERATING CO
  • US7979280B2 patent drawing
  • US7979280B2 patent drawing
  • US7979280B2 patent drawing

AI summary

An input linguistic description is converted into a speech waveform by deriving at least one target unit sequence corresponding to the linguistic description, selecting from a waveform unit database for the target unit sequences a plurality of alternative unit sequences approximating the target unit sequences, concatenating the alternative unit sequences to alternative speech waveforms and presenting the alternative speech waveforms to an operating person and enabling the choice of one of the presented alternative speech waveforms. There are no iterative cycles of manual modification and automatic selection, which enables a fast way of working. The operator does not need knowledge of units, targets, and costs, but chooses from a set of given alternatives. The fine-tuning of TTS prompts therefore becomes accessible to non-experts.