TTS Style Conditioning With Disentangled Audio Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI speech generation systems struggle to produce speech with human-understandable and controllable style and environmental characteristics, such as emotional content and acoustic environments, which are crucial for realistic and efficient audio production.

Innovation Solution

A style model is trained to encode style information in a disentangled format using exclusionary data and disentanglement strategies, allowing for human-understandable and controllable audio characteristics to be embedded in the generated speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If AI models generate speech with human-like characteristics, then the realism of generated speech is improved, but the controllability and understandability of style and environmental characteristics deteriorate

Engineering Contradiction:
Improverealism of generated speechVSAvoidcontrollability of style characteristics
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent segments the style information into distinct dimensions including speaker identity, emotional content, acoustic environment, and prosody. Each dimension is encoded separately in the style embedding, allowing independent control and manipulation of each characteristic while maintaining overall human-like realism in the generated speech.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If style information is encoded in a disentangled format, then the controllability of audio characteristics is improved, but the complexity of the training process deteriorates

Engineering Contradiction:
Improvecontrollability of audio characteristicsVSAvoidcomplexity of training process
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary style encoder that processes audio clips and extracts style information into a standardized embedding format. This intermediary component simplifies the training process by providing a clear interface between the audio data and the speech generation model, while enabling disentangled control of audio characteristics through the encoded style dimensions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If multiple style dimensions are encoded simultaneously, then the versatility of generated speech is improved, but the difficulty of detecting and measuring specific characteristics deteriorates

Engineering Contradiction:
Improveversatility of generated speechVSAvoiddetection of style characteristics
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies local quality by assigning specific regions or dimensions within the style embedding to represent specific audio characteristics such as speaker identity, emotion, acoustic environment, and prosody. This localized encoding strategy enables simultaneous encoding of multiple style dimensions while maintaining detectability and measurability of individual characteristics through their dedicated spatial locations in the embedding space.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12567421B2Text to audio conversion with disentangled style conditioning
Publication Date: 2026.03.03 NARO CORP
  • US12567421B2 patent drawing
  • US12567421B2 patent drawing
  • US12567421B2 patent drawing

AI summary

A style encoder can be trained to encode audio style and audio characteristics into selected regions of a style vector. The style vector can be used to condition a text to speech (TTS) model to generate speech with human-understandable and controllable styles. Various training strategies of the style encoder are described, including a first, second and third training strategy that can be used to disentangle audio styles into selected regions of a style vector. The distinct regions of the style vector can be used to provide numerous customization options to a user of the described system, along with tools to generate speech with a speaker identity and using selected audio styles and characteristics.