Interactive Synthesized Speech Style Control Without Retraining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis systems require training data that reflects a single speaking style, limiting the ability to control and adjust speaking styles without additional recordings or retraining.

Innovation Solution

A method for controlling speaking style in text-to-speech systems by configuring a summarizing unit and synthesizing unit with configurable parameters, allowing interactive adjustment of characteristics like speed, kindness, and pitch without retraining, using a style basis to map quality targets to desired speaking styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If training data includes only a single speaking style, then the TTS system can be trained efficiently with consistent style characteristics, but the system cannot control or adjust speaking styles without additional recordings or retraining

Engineering Contradiction:
ImproveTraining efficiencyVSAvoidSpeaking style control
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent extracts speaking style characteristics from training data into separate style embeddings that are independent of the main TTS model training. This allows style information to be taken out and manipulated separately, enabling style control without retraining the entire system.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces style embeddings as an intermediary between the text input and speech synthesis process. These embeddings act as a mediator that carries style information independently, allowing style control without affecting the core TTS model training.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If training data includes multiple speaking styles, then the system can select from different styles, but it requires additional input representations and complex training processes

Engineering Contradiction:
ImproveStyle selection capabilityVSAvoidTraining process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent changes the parameter representation of speaking styles by using continuous style embeddings instead of discrete style labels. This allows smooth transitions between styles and simplifies the training process by treating style as a continuous parameter rather than requiring separate training for each style.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The style embeddings serve multiple functions: they enable style selection, style mixing, and style interpolation all through the same mechanism. This universal approach handles multiple style-related tasks without requiring separate processing paths.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If the TTS system allows interactive adjustment of speaking style characteristics, then users can achieve desired style without retraining, but the system requires a mechanism to map quality targets to style representations

Engineering Contradiction:
ImproveInteractive style adjustmentVSAvoidStyle mapping mechanism
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements feedback by allowing users to interactively adjust style characteristics and immediately hear the effect on synthesized speech. This real-time feedback loop enables users to fine-tune style parameters without requiring complex retraining processes.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The style basis is pre-computed from training data before interactive adjustment begins. This preliminary action creates a mapping between quality targets and style embeddings in advance, so that interactive adjustments only require simple parameter modifications rather than complex computations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250225976A1Interactive Modification of Speaking Style of Synthesized Speech
Publication Date: 2025.07.10 CERENCE OPERATING CO
  • US20250225976A1 patent drawing
  • US20250225976A1 patent drawing
  • US20250225976A1 patent drawing

AI summary

Control over speaking style of a text-to-speech (TTS) system is provided without necessarily requiring that the training of the TTS conversion process (e.g., the ANN used for the conversion) take into account the speaking styles of the training data. For example, the TTS system may allow adjustment of characteristics of speaking styles, such as, speed, perceivable degree of “kindness”, average pitch, pitch variation, and duration of pauses. In some examples, a voice designer may have a number of independent controls that vary corresponding characteristics without necessarily varying others. Once the designer has configured a desired overall speaking style based on those controllable characteristics, the TTS system can be configured to use that speaking style for deployments of the TTS system. For example, the TTS system may be used for audio output in a voice assistant, for instance, for an in-vehicle voice assistant.