Speech Unit Selection via Dual-Cost Lattice Construction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech systems face challenges in producing natural-sounding speech synthesis due to the inability to effectively balance target costs and join costs during lattice construction, leading to local minima or maxima issues and suboptimal quality, especially with large corpora of speech units.

Innovation Solution

A text-to-speech system that uses both target costs and join costs to select speech units for lattice construction, considering acoustic and linguistic parameters to ensure a more natural and coherent synthesis by evaluating the fit of speech units within the entire sequence and using the Viterbi algorithm to maintain diversity of paths.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If only target cost is used to select speech units for lattice construction, then the selection process is simpler and faster, but the synthesized speech quality deteriorates due to inability to ensure natural transitions between units

Engineering Contradiction:
Improvelattice construction speedVSAvoidspeech synthesis quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces a dual-cost parameter system (target cost and join cost) instead of using only target cost. This parameter change enables the system to evaluate speech units based on both their matching accuracy to phonetic targets and their compatibility with adjacent units, thereby improving speech quality while maintaining computational efficiency through the Viterbi algorithm optimization.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If join cost is incorporated into speech unit selection, then speech synthesis quality improves through better acoustic continuity, but computational complexity increases

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidlattice construction complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary calculation and storage of join costs for all speech unit pairs before the actual lattice construction process. This preliminary action pre-processes the compatibility information, allowing the main Viterbi algorithm to efficiently retrieve and use these pre-computed values without performing complex real-time calculations, thus reducing computational complexity during the critical synthesis phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a copied and simplified representation of the cost structure by separating target cost and join cost into distinct, pre-computed arrays. This copying approach allows the Viterbi algorithm to work with simplified data structures rather than complex nested calculations, improving computational efficiency while maintaining the ability to evaluate both target matching and acoustic continuity.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If a large corpus of speech units is used, then the versatility and naturalness of synthesized speech improves, but the difficulty of selecting optimal units increases due to more candidate options

Engineering Contradiction:
Improvespeech corpus coverageVSAvoidoptimal unit selection difficulty
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent transforms the selection problem from an unmanageable combinatorial optimization problem into a tractable dynamic programming problem by introducing the join cost parameter. This parameter change enables efficient pruning of the search space by eliminating candidate units that have poor acoustic continuity with their neighbors, even if they perfectly match their target phoneme, thus making large corpus processing feasible.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback through the Viterbi algorithm that continuously evaluates both target cost and join cost as it progresses through the lattice construction. This feedback mechanism allows the system to adjust its selections dynamically, using the accumulated cost information to guide subsequent unit selections and ensure overall path optimality rather than greedy local selections.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3376498B1Speech synthesis unit selection
Publication Date: 2023.11.15 GOOGLE LLC
  • EP3376498B1 patent drawingFigure 1
  • EP3376498B1 patent drawingFigure 2
  • EP3376498B1 patent drawingFigure 3

AI summary

A method of selecting units for speech synthesis includes receiving, by one or more computers of a text-to-speech system, data indicating text for speech synthesis; determining, by the one or more computers, a sequence of text units that each represent a respective portion of the text, the sequence of text units including at least a first text unit followed by a second text unit; determining, by the one or more computers, multiple paths of speech units that each represent the sequence of text units, wherein determining the multiple paths of speech units includes: selecting, from a speech unit corpus, a first speech unit that includes speech synthesis data representing the first text unit; selecting, from the speech unit corpus, multiple second speech units including speech synthesis data representing the second text unit, each of the multiple second speech units being determined based on (i) a join cost to concatenate the second speech unit with a first speech unit and (ii) a target cost indicating a degree that the second speech unit corresponds to the second text unit; and defining paths from the selected first speech unit to each of the multiple second speech units to include in the multiple paths of speech units; and providing, by the one or more computers of the text-to-speech system, synthesized speech data according to a path selected from among the multiple paths.