Natural Language Style Tags for Emotion-Aware Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies fail to effectively replicate the personality and emotional nuances of human speakers, leading to subpar audio content in broadcast programs, and lack user-friendly interfaces for generating and editing audio content based on text.

Innovation Solution

A method and system that utilize a text-to-speech synthesis model trained with reference voice data and style tags in natural language to generate synthesis voices and images, allowing users to input style tags through a user interface, recommend candidate tags, and adjust voice style features, enabling the creation of audio content that reflects desired emotions and talking styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech synthesis technology is used to generate audio content, then production convenience is improved, but the audio content does not reflect speaker personality and emotions

Engineering Contradiction:
Improveproduction convenienceVSAvoidemotional accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies parameter changes by introducing style tags that modify the emotional and stylistic parameters of synthesized speech. The system adjusts prosodic parameters (pitch, duration, energy) and spectral parameters based on style tag embeddings, enabling the same text to be synthesized with different emotional characteristics without changing the underlying synthesis model.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses style tags as an intermediary between the text input and the speech synthesis process. These style tags act as mediators that carry emotional and stylistic information, transforming the synthesis process from direct text-to-speech into a three-stage process: text encoding, style tag processing, and combined synthesis, thereby enabling emotional expression without complex model architectures.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If traditional speech synthesis methods are used, then technical complexity is reduced, but user interface intuitiveness deteriorates

Engineering Contradiction:
Improvetechnical complexityVSAvoiduser interface intuitiveness
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The patent uses natural language style tags as a simplified copy or representation of complex emotional states. Instead of requiring users to manipulate multiple technical parameters (pitch, duration, energy independently), the system copies the natural way humans describe emotions in text and uses that as input, maintaining simplicity while improving intuitiveness.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The style tag mechanism serves multiple functions simultaneously: it controls emotional expression, adjusts prosodic features, and provides a unified interface for style transfer. This multi-functionality allows a single simple input mechanism to handle various aspects of speech style control that would otherwise require separate controls.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If style tags are added to control voice features, then emotional accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveemotional accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical system of multiple independent parameter adjustments with a neural network-based embedding system. Instead of manually configuring multiple parameters to achieve desired emotional expression, the system uses learned embeddings from style tags that automatically encode the appropriate parameter combinations, reducing the need for complex mechanical control structures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary action by pre-training the style tag embedding model on large datasets of speech and text. This pre-training captures complex emotional and stylistic patterns in advance, so that during actual synthesis, the system only needs to retrieve and apply pre-computed embeddings rather than computing emotional parameters in real-time, thereby managing complexity through advance preparation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240105160A1Method and system for generating synthesis voice using style tag represented by natural language
Publication Date: 2024.03.28 NEOSAPIENCE INC
  • US20240105160A1 patent drawing
  • US20240105160A1 patent drawing
  • US20240105160A1 patent drawing

AI summary

A method for generating a synthesis voice is provided, which is performed by one or more processors, and includes acquiring a text-to-speech synthesis model trained to generate a synthesis voice for a training text, based on reference voice data and a training style tag represented by natural language, receiving a target text, acquiring a style tag represented by natural language, and inputting the style tag and the target text into the text-to-speech synthesis model and acquiring a synthesis voice for the target text reflecting voice style features related to the style tag.