Natural Language Style Tags for Emotion-Aware Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies fail to effectively replicate the personality and emotional nuances of human speakers, leading to subpar audio content in broadcast programs, and lack user-friendly interfaces for generating and editing audio content based on text.
Innovation Solution
A method and system that utilize a text-to-speech synthesis model trained with reference voice data and style tags in natural language to generate synthesis voices and images, allowing users to input style tags through a user interface, recommend candidate tags, and adjust voice style features, enabling the creation of audio content that reflects desired emotions and talking styles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech synthesis technology is used to generate audio content, then production convenience is improved, but the audio content does not reflect speaker personality and emotions
Solution Approach 1:
The patent applies parameter changes by introducing style tags that modify the emotional and stylistic parameters of synthesized speech. The system adjusts prosodic parameters (pitch, duration, energy) and spectral parameters based on style tag embeddings, enabling the same text to be synthesized with different emotional characteristics without changing the underlying synthesis model.
Solution Approach 2:
The patent uses style tags as an intermediary between the text input and the speech synthesis process. These style tags act as mediators that carry emotional and stylistic information, transforming the synthesis process from direct text-to-speech into a three-stage process: text encoding, style tag processing, and combined synthesis, thereby enabling emotional expression without complex model architectures.
2Device complexity
If traditional speech synthesis methods are used, then technical complexity is reduced, but user interface intuitiveness deteriorates
Solution Approach 1:
The patent uses natural language style tags as a simplified copy or representation of complex emotional states. Instead of requiring users to manipulate multiple technical parameters (pitch, duration, energy independently), the system copies the natural way humans describe emotions in text and uses that as input, maintaining simplicity while improving intuitiveness.
Solution Approach 2:
The style tag mechanism serves multiple functions simultaneously: it controls emotional expression, adjusts prosodic features, and provides a unified interface for style transfer. This multi-functionality allows a single simple input mechanism to handle various aspects of speech style control that would otherwise require separate controls.
3Reliability
If style tags are added to control voice features, then emotional accuracy is improved, but system complexity increases
Solution Approach 1:
The patent replaces the mechanical system of multiple independent parameter adjustments with a neural network-based embedding system. Instead of manually configuring multiple parameters to achieve desired emotional expression, the system uses learned embeddings from style tags that automatically encode the appropriate parameter combinations, reducing the need for complex mechanical control structures.
Solution Approach 2:
The system performs preliminary action by pre-training the style tag embedding model on large datasets of speech and text. This pre-training captures complex emotional and stylistic patterns in advance, so that during actual synthesis, the system only needs to retrieve and apply pre-computed embeddings rather than computing emotional parameters in real-time, thereby managing complexity through advance preparation.
Data Source
AI summary
A method for generating a synthesis voice is provided, which is performed by one or more processors, and includes acquiring a text-to-speech synthesis model trained to generate a synthesis voice for a training text, based on reference voice data and a training style tag represented by natural language, receiving a target text, acquiring a style tag represented by natural language, and inputting the style tag and the target text into the text-to-speech synthesis model and acquiring a synthesis voice for the target text reflecting voice style features related to the style tag.


