Text to Speech Synthesis with Alternative Unit Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Text-to-Speech (TTS) systems face challenges in producing optimal voice quality and speaking style due to underspecification of input text information, difficulty in detecting voice quality and speaking style changes, and sparseness in unit databases, leading to suboptimal unit selection and synthesis.
Innovation Solution
A unit selection system that generates multiple acoustic realizations of a linguistic description, allowing operators to choose the most suitable option, by deriving target unit sequences, selecting alternative unit sequences, and concatenating them into speech waveforms, with an intelligent construction algorithm to simplify the process and reduce manual iteration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional unit selection TTS systems are used, then speech synthesis can be performed, but the voice quality and speaking style are suboptimal due to underspecification of input text information
Solution Approach 1:
The patent segments the speech database into multiple overlapping units (phones, syllables, words, phrases) with different granularities. This segmentation allows the system to select units at appropriate levels of detail to capture voice quality and speaking style information that would be lost in traditional phoneme-based approaches alone.
Solution Approach 2:
The patent adds a new dimension to unit selection by incorporating prosodic features (stress, intonation, rhythm) and speaking style attributes as explicit selection criteria. This transforms the selection process from purely phoneme-matching to multi-dimensional optimization that preserves voice quality information.
2Manufacturing precision
If manual iteration is used to optimize speech prompts, then optimal voice quality can be achieved, but the process is time-consuming and requires expert knowledge
Solution Approach 1:
The patent performs preliminary action by automatically generating multiple candidate unit sequences with different prosodic and stylistic variations before presentation to the user. This pre-computation of alternatives eliminates the need for iterative manual optimization, as the best options are already prepared in advance.
Solution Approach 2:
The system performs self-service by automatically evaluating and ranking multiple unit sequence candidates based on predefined criteria for voice quality and speaking style. This automated self-evaluation reduces dependence on expert operators and minimizes manual iteration time.
3Ease of operation
If multiple alternative unit sequences are generated, then operator choice is improved, but the number of alternatives may become excessive
Solution Approach 1:
The patent controls the number of alternatives by adjusting parameters such as the diversity of prosodic variations, the granularity of unit segmentation, and the selection criteria weights. These parameter changes allow the system to generate an optimal number of meaningful alternatives without overwhelming the operator.
Solution Approach 2:
The patent applies local quality by generating alternatives that differ specifically in prosodic and stylistic features while maintaining phonetic accuracy. This ensures that each alternative has distinct local characteristics in terms of voice quality, making the selection process more efficient and less overwhelming.
4Manufacturing precision
If unit databases are expanded to include more voice quality information, then synthesis quality improves, but database sparseness increases
Solution Approach 1:
The patent uses a nested structure where units at different levels of granularity (phones within syllables, syllables within words, words within phrases) are organized hierarchically. This nesting allows the system to leverage information from multiple levels simultaneously, improving synthesis quality without requiring a proportionally larger database at each level.
Solution Approach 2:
The patent makes the unit database multi-functional by creating units that serve multiple purposes: they represent phonetic content, prosodic features, and speaking style simultaneously. This universality allows the same database to support various synthesis quality requirements without needing separate specialized databases for each feature.
Data Source
AI summary
An input linguistic description is converted into a speech waveform by deriving at least one target unit sequence corresponding to the linguistic description, selecting from a waveform unit database for the target unit sequences a plurality of alternative unit sequences approximating the target unit sequences, concatenating the alternative unit sequences to alternative speech waveforms and presenting the alternative speech waveforms to an operating person and enabling the choice of one of the presented alternative speech waveforms. There are no iterative cycles of manual modification and automatic selection, which enables a fast way of working. The operator does not need knowledge of units, targets, and costs, but chooses from a set of given alternatives. The fine-tuning of TTS prompts therefore becomes accessible to non-experts.


