Speech Unit Selection via Dual-Cost Lattice Construction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech systems face challenges in producing natural-sounding speech synthesis due to the inability to effectively balance target costs and join costs during lattice construction, leading to local minima or maxima issues and suboptimal quality, especially with large corpora of speech units.
Innovation Solution
A text-to-speech system that uses both target costs and join costs to select speech units for lattice construction, considering acoustic and linguistic parameters to ensure a more natural and coherent synthesis by evaluating the fit of speech units within the entire sequence and using the Viterbi algorithm to maintain diversity of paths.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If only target cost is used to select speech units for lattice construction, then the selection process is simpler and faster, but the synthesized speech quality deteriorates due to inability to ensure natural transitions between units
Solution Approach 1:
The patent introduces a dual-cost parameter system (target cost and join cost) instead of using only target cost. This parameter change enables the system to evaluate speech units based on both their matching accuracy to phonetic targets and their compatibility with adjacent units, thereby improving speech quality while maintaining computational efficiency through the Viterbi algorithm optimization.
2Manufacturing precision
If join cost is incorporated into speech unit selection, then speech synthesis quality improves through better acoustic continuity, but computational complexity increases
Solution Approach 1:
The patent performs preliminary calculation and storage of join costs for all speech unit pairs before the actual lattice construction process. This preliminary action pre-processes the compatibility information, allowing the main Viterbi algorithm to efficiently retrieve and use these pre-computed values without performing complex real-time calculations, thus reducing computational complexity during the critical synthesis phase.
Solution Approach 2:
The patent creates a copied and simplified representation of the cost structure by separating target cost and join cost into distinct, pre-computed arrays. This copying approach allows the Viterbi algorithm to work with simplified data structures rather than complex nested calculations, improving computational efficiency while maintaining the ability to evaluate both target matching and acoustic continuity.
3Adaptability or versatility
If a large corpus of speech units is used, then the versatility and naturalness of synthesized speech improves, but the difficulty of selecting optimal units increases due to more candidate options
Solution Approach 1:
The patent transforms the selection problem from an unmanageable combinatorial optimization problem into a tractable dynamic programming problem by introducing the join cost parameter. This parameter change enables efficient pruning of the search space by eliminating candidate units that have poor acoustic continuity with their neighbors, even if they perfectly match their target phoneme, thus making large corpus processing feasible.
Solution Approach 2:
The patent implements feedback through the Viterbi algorithm that continuously evaluates both target cost and join cost as it progresses through the lattice construction. This feedback mechanism allows the system to adjust its selections dynamically, using the accumulated cost information to guide subsequent unit selections and ensure overall path optimality rather than greedy local selections.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method of selecting units for speech synthesis includes receiving, by one or more computers of a text-to-speech system, data indicating text for speech synthesis; determining, by the one or more computers, a sequence of text units that each represent a respective portion of the text, the sequence of text units including at least a first text unit followed by a second text unit; determining, by the one or more computers, multiple paths of speech units that each represent the sequence of text units, wherein determining the multiple paths of speech units includes: selecting, from a speech unit corpus, a first speech unit that includes speech synthesis data representing the first text unit; selecting, from the speech unit corpus, multiple second speech units including speech synthesis data representing the second text unit, each of the multiple second speech units being determined based on (i) a join cost to concatenate the second speech unit with a first speech unit and (ii) a target cost indicating a degree that the second speech unit corresponds to the second text unit; and defining paths from the selected first speech unit to each of the multiple second speech units to include in the multiple paths of speech units; and providing, by the one or more computers of the text-to-speech system, synthesized speech data according to a path selected from among the multiple paths.