Text-to-Speech Phoneme Insertion for Spontaneous Rhythm

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech (TTS) systems generate speech that is stiffer and less smooth compared to human spontaneous speech, lacking simulations of pauses, repetitions, and diverse rhythms.

Innovation Solution

Generate an initial phoneme sequence from text, insert additional phonemes related to spontaneous speech characteristics, determine phoneme durations using expert models, and create a second phoneme sequence to produce spontaneous-style speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional TTS systems generate speech from text, then speech output is produced efficiently, but the speech is stiffer and less smooth compared to spontaneous speech

Engineering Contradiction:
Improvenaturalness of speechVSAvoidcomplexity of phoneme sequence generation
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the speech generation process into distinct phoneme sequences: an initial phoneme sequence from text conversion, and a modified phoneme sequence with inserted additional phonemes. This segmentation allows independent optimization of each sequence's characteristics to achieve naturalness while managing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing stage that inserts additional phonemes between the initial phoneme sequence and the final speech output. This intermediary layer modifies the phoneme sequence to incorporate spontaneous speech characteristics without completely redesigning the entire TTS system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If additional phonemes are inserted to simulate spontaneous speech characteristics, then speech naturalness is improved, but processing time and computational complexity increase

Engineering Contradiction:
Improveauthenticity of spontaneous speechVSAvoidprocessing time for phoneme sequence generation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-identifying and inserting additional phonemes into the phoneme sequence before the main speech synthesis process. This preliminary modification ensures that spontaneous speech characteristics are built into the structure early, avoiding time-consuming adjustments during later processing stages.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If multiple expert models are used to determine phoneme durations, then rhythm variation and speech naturalness are improved, but computational resources and system complexity increase

Engineering Contradiction:
Improverhythm variation of speechVSAvoidnumber of expert models
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The expert models serve multiple functions: they determine both the duration of phonemes and the rhythm patterns of spontaneous speech. By making the expert models multi-functional, the system achieves improved rhythm variation without proportionally increasing the number of separate components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes parameters by adjusting phoneme durations through the expert models to create rhythm variation. Instead of adding complex structural elements, the patent achieves naturalness by dynamically modifying temporal parameters of existing phonemes.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4715807A2Text-based speech generation
Publication Date: 2026.03.25 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4715807A2 patent drawingFigure 1
  • EP4715807A2 patent drawingFigure 2
  • EP4715807A2 patent drawingFigure 3

AI summary

According to implementations of the subject matter described herein, a solution is proposed for text to speech. In this solution, an initial phoneme sequence corresponding to text is generated, the initial phoneme sequence comprising feature representations of a plurality of phonemes. A first phoneme sequence is generated by inserting a feature representation of an additional phoneme into the initial phoneme sequence, the additional phoneme being related to a characteristic of spontaneous speech. The duration of a phoneme among the plurality of phonemes and the additional phoneme is determined by using an expert model corresponding to the phoneme, and a second phoneme sequence is generated based on the first phoneme sequence. Spontaneous-style speech corresponding to the text is determined based on the second phoneme sequence. In this way, spontaneous-style speech with more varying rhythms can be generated based on spontaneous-style additional phonemes and multiple expert models.