Text-to-Speech Phoneme Insertion for Spontaneous Rhythm
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-to-speech (TTS) systems generate speech that is stiffer and less smooth compared to human spontaneous speech, lacking simulations of pauses, repetitions, and diverse rhythms.
Innovation Solution
Generate an initial phoneme sequence from text, insert additional phonemes related to spontaneous speech characteristics, determine phoneme durations using expert models, and create a second phoneme sequence to produce spontaneous-style speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional TTS systems generate speech from text, then speech output is produced efficiently, but the speech is stiffer and less smooth compared to spontaneous speech
Solution Approach 1:
The system segments the speech generation process into distinct phoneme sequences: an initial phoneme sequence from text conversion, and a modified phoneme sequence with inserted additional phonemes. This segmentation allows independent optimization of each sequence's characteristics to achieve naturalness while managing complexity.
Solution Approach 2:
The patent introduces an intermediary processing stage that inserts additional phonemes between the initial phoneme sequence and the final speech output. This intermediary layer modifies the phoneme sequence to incorporate spontaneous speech characteristics without completely redesigning the entire TTS system.
2Reliability
If additional phonemes are inserted to simulate spontaneous speech characteristics, then speech naturalness is improved, but processing time and computational complexity increase
Solution Approach 1:
The system performs preliminary actions by pre-identifying and inserting additional phonemes into the phoneme sequence before the main speech synthesis process. This preliminary modification ensures that spontaneous speech characteristics are built into the structure early, avoiding time-consuming adjustments during later processing stages.
3Reliability
If multiple expert models are used to determine phoneme durations, then rhythm variation and speech naturalness are improved, but computational resources and system complexity increase
Solution Approach 1:
The expert models serve multiple functions: they determine both the duration of phonemes and the rhythm patterns of spontaneous speech. By making the expert models multi-functional, the system achieves improved rhythm variation without proportionally increasing the number of separate components.
Solution Approach 2:
The system changes parameters by adjusting phoneme durations through the expert models to create rhythm variation. Instead of adding complex structural elements, the patent achieves naturalness by dynamically modifying temporal parameters of existing phonemes.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
According to implementations of the subject matter described herein, a solution is proposed for text to speech. In this solution, an initial phoneme sequence corresponding to text is generated, the initial phoneme sequence comprising feature representations of a plurality of phonemes. A first phoneme sequence is generated by inserting a feature representation of an additional phoneme into the initial phoneme sequence, the additional phoneme being related to a characteristic of spontaneous speech. The duration of a phoneme among the plurality of phonemes and the additional phoneme is determined by using an expert model corresponding to the phoneme, and a second phoneme sequence is generated based on the first phoneme sequence. Spontaneous-style speech corresponding to the text is determined based on the second phoneme sequence. In this way, spontaneous-style speech with more varying rhythms can be generated based on spontaneous-style additional phonemes and multiple expert models.