Dynamic Text Reader Emotion and Speaker Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-speech applications lack emotional and tonal variations, making readings less engaging and vivid for listeners, as they typically use a plain voice without emphasis or emotional expression.
Innovation Solution
A dynamic text reader that performs pre-processing on text documents to identify emotional statements, speakers, and punctuation, generating a voice with tonal modulation to convey emotions, using a role-to-voice map and emotion database to create a more engaging listening experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a plain voice is used for text-to-speech conversion, then the system is simple and easy to operate, but the reading lacks emotional and tonal variations making it less engaging
Solution Approach 1:
The patent segments the text into different emotional statements and identifies distinct speakers, assigning different digital voice representations to each. This segmentation allows the system to apply different emotional tones and vocal characteristics to different portions of the text, transforming a monotonous plain reading into an engaging multi-voiced performance with emotional variation.
Solution Approach 2:
The patent changes the parameters of the voice output by selecting different digital voice representations with varying tonal qualities, pitches, and emotional characteristics. By modifying these voice parameters based on the identified emotional content and speakers, the system transitions from a static plain voice to a dynamic emotionally expressive reading.
2Adaptability or versatility
If emotional and tonal variations are added to text readings, then the listening experience becomes more engaging and vivid, but the system complexity increases
Solution Approach 1:
The patent performs pre-processing of the text document before actual speech conversion, including identifying emotional statements, determining speakers, and creating a role-to-voice map. This preliminary action prepares all necessary emotional and vocal information in advance, allowing the speech synthesis phase to simply retrieve and apply the predetermined voice representations, thereby managing complexity through staged processing.
Solution Approach 2:
The patent introduces a role-to-voice map as an intermediary data structure that connects identified speakers and emotional statements to appropriate digital voice representations. This intermediary layer decouples the complex analysis of emotional content from the speech synthesis process, making the system more manageable by creating a clear mapping layer between text analysis and voice generation.
Data Source
AI summary
Embodiments are disclosed for a method for dynamic text reading. The method includes performing pre-processing for a text document. Pre-processing includes determining the text document comprises an emotional statement based on an indicator of an emotion associated with the emotional statement. Pre-processing also includes identifying a speaker of the emotional statement. Further, pre-processing includes generating a role-to-voice map that associates the speaker with a digital representation of a voice for the speaker. The method additionally includes generating, based on the pre-processing, the voice for the speaker reading aloud a text of the text document using the digital representation of the voice with a tonal modulation that conveys the emotion.


