Phonemic-Level TOBI Prosody for Natural Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis technologies struggle to naturally and smoothly combine prosodic features of text, leading to uncontrollable intensity in synthesized audio.

Innovation Solution

A speech synthesis method that generates phonemic-level TOBI representation sequences and prosodic-acoustic features, incorporating both language and acoustic levels of prosody to control audio intensity and naturalness, using a pre-trained model with encoding, attention, and decoding networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing speech synthesis technologies are used, then speech synthesis can be performed, but the prosodic features cannot be naturally and smoothly combined, and the intensity is uncontrollable

Engineering Contradiction:
Improvenaturalness of synthesized audioVSAvoidcontrollability of audio intensity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent segments the prosodic feature processing into multiple distinct modules: a prosodic feature analysis module that extracts features like pitch, intensity, and duration; a prosodic feature combination module that integrates these features; and a prosodic feature adjustment module that controls intensity. This segmentation allows each module to handle specific aspects of prosody independently, enabling both natural combination and precise control of audio intensity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic adjustment of prosodic features through the prosodic feature adjustment module, which can modify intensity, pitch, and duration parameters in real-time based on input text and desired output characteristics. This dynamic control enables the system to adapt prosodic features to achieve naturalness while maintaining controllability over audio intensity.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If prosodic features are combined to improve naturalness, then the synthesized audio becomes more natural, but the system complexity increases

Engineering Contradiction:
Improvenaturalness of synthesized audioVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent employs a speech synthesis model that performs multiple functions: it processes phoneme sequences, extracts prosodic features, combines prosodic features, and generates acoustic features. By making this model multi-functional, the patent reduces the need for separate dedicated components for each processing stage, thereby managing system complexity while achieving natural prosodic combination.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces a prosodic feature combination module as an intermediary between the prosodic feature analysis module and the acoustic feature generation. This intermediary module seamlessly integrates multiple prosodic features (pitch, intensity, duration) into a unified representation, simplifying the overall system architecture by providing a clear interface between feature extraction and acoustic synthesis.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If phonemic-level TOBI representation and prosodic-acoustic features are generated, then audio intensity and naturalness are controlled, but the processing time increases

Engineering Contradiction:
Improvecontrol of audio intensityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing by generating phonemic-level TOBI representation sequences and prosodic-acoustic features from the input text before the main acoustic synthesis. This preliminary action organizes and pre-processes prosodic information in advance, enabling more efficient subsequent processing and reducing overall computation time during actual speech synthesis operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12444401B2Method, apparatus, computer readable medium, and electronic device of speech synthesis
Publication Date: 2025.10.14 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US12444401B2 patent drawing
  • US12444401B2 patent drawing
  • US12444401B2 patent drawing

AI summary

A method, apparatus, a computer readable medium, and an electronic device of speech synthesis. The method includes: obtaining a phoneme sequence corresponding to text to be synthesized; generating a phonemic-level TOBI representation sequence and a prosodic-acoustic feature corresponding to the text to be synthesized based on the phoneme sequence and the text to be synthesized, and generating acoustic feature information corresponding to the text to be synthesized based on the TOBI representation sequence and the prosodic-acoustic feature; and generating first audio information corresponding to the text to be synthesized based on the acoustic feature information. The method enables the synthesized audio to be more natural, cadenced, and aligned with the intended semantics of a speaker.