Emotional Text to Speech Using Rhythm Piece Emotion Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing emotional Text to Speech (TTS) technologies fail to produce natural and smooth emotional transitions, as they often assign unified emotion categories to sentences, leading to abrupt changes, and require manual intervention, which is costly and not suitable for batch processing.
Innovation Solution
A method and system that generate emotion tags based on rhythm pieces expressed as emotion vectors, allowing for multiple emotion categories and automatic emotion scoring, enabling more nuanced and realistic emotional expression without manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If unified emotion categories are assigned to sentences, then the TTS system is simple to operate, but the emotional transitions become abrupt and unnatural
Solution Approach 1:
The patent segments sentences into multiple rhythm pieces (e.g., phrases or clauses) and assigns emotion tags to each rhythm piece individually rather than uniformly to the entire sentence. This segmentation allows for gradual emotional transitions between different parts of the text, resolving the contradiction by maintaining operational simplicity while achieving nuanced emotional expression through localized emotion assignments.
Solution Approach 2:
The patent applies different emotion tags to different rhythm pieces within the same sentence based on their local semantic content. For example, in the sentence 'Mr. Ding suffers severe paralysis since he is young, but he learns through self-study and finally wins the heart of Ms. Zhao with the help of network', the first rhythm piece receives a sad emotion tag while the second receives a joy emotion tag. This local differentiation resolves the contradiction by enabling precise emotional control where needed while maintaining overall system simplicity.
2Manufacturing precision
If manual emotion specification is used, then the emotional expression can be controlled, but the processing cost increases and batch processing becomes difficult
Solution Approach 1:
The patent implements an automatic emotion tag generation mechanism that analyzes the semantic content of rhythm pieces and assigns appropriate emotion tags without manual intervention. The system uses predefined emotion categories (happy, sad, angry, neutral) and automatically determines which category applies to each rhythm piece based on its content. This self-service approach resolves the contradiction by maintaining emotional expression control through systematic rules while enabling efficient batch processing of large text volumes.
3Manufacturing precision
If emotion tags are generated by rhythm piece, then the emotional transitions become smooth and natural, but the system complexity increases
Solution Approach 1:
The patent divides text into rhythm pieces as the basic unit for emotion tag assignment. This segmentation strategy resolves the contradiction by creating a manageable intermediate structure that enables smooth emotional transitions without requiring complex continuous emotion modeling. Each rhythm piece is processed independently with a discrete emotion tag, simplifying the overall system architecture while achieving natural-sounding emotional progression.
Solution Approach 2:
The patent uses discrete emotion category parameters (happy, sad, angry, neutral) assigned to different rhythm pieces to achieve smooth emotional transitions. By changing the emotion parameter at rhythm piece boundaries rather than continuously, the system achieves natural emotional progression while maintaining relatively simple system structure. The emotion tag for each rhythm piece serves as a control parameter that guides the TTS synthesis.
Data Source
AI summary
A method and system for achieving emotional text to speech. The method includes: receiving text data; generating emotion tag for the text data by a rhythm piece; and achieving TTS to the text data corresponding to the emotion tag, where the emotion tags are expressed as a set of emotion vectors; where each emotion vector includes a plurality of emotion scores given based on a plurality of emotion categories. A system for the same includes: a text data receiving module; an emotion tag generating module; and a TTS module for achieving TTS, wherein the emotion tag is expressed as a set of emotion vectors; and wherein emotion vector includes a plurality of emotion scores given based on a plurality of emotion categories.


