Dynamic Text Reader Emotion and Speaker Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech applications lack emotional and tonal variations, making readings less engaging and vivid for listeners, as they typically use a plain voice without emphasis or emotional expression.

Innovation Solution

A dynamic text reader that performs pre-processing on text documents to identify emotional statements, speakers, and punctuation, generating a voice with tonal modulation to convey emotions, using a role-to-voice map and emotion database to create a more engaging listening experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a plain voice is used for text-to-speech conversion, then the system is simple and easy to operate, but the reading lacks emotional and tonal variations making it less engaging

Engineering Contradiction:
Improveease of operationVSAvoidemotional expression capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent segments the text into different emotional statements and identifies distinct speakers, assigning different digital voice representations to each. This segmentation allows the system to apply different emotional tones and vocal characteristics to different portions of the text, transforming a monotonous plain reading into an engaging multi-voiced performance with emotional variation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of the voice output by selecting different digital voice representations with varying tonal qualities, pitches, and emotional characteristics. By modifying these voice parameters based on the identified emotional content and speakers, the system transitions from a static plain voice to a dynamic emotionally expressive reading.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If emotional and tonal variations are added to text readings, then the listening experience becomes more engaging and vivid, but the system complexity increases

Engineering Contradiction:
Improveemotional expression capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs pre-processing of the text document before actual speech conversion, including identifying emotional statements, determining speakers, and creating a role-to-voice map. This preliminary action prepares all necessary emotional and vocal information in advance, allowing the speech synthesis phase to simply retrieve and apply the predetermined voice representations, thereby managing complexity through staged processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a role-to-voice map as an intermediary data structure that connects identified speakers and emotional statements to appropriate digital voice representations. This intermediary layer decouples the complex analysis of emotional content from the speech synthesis process, making the system more manageable by creating a clear mapping layer between text analysis and voice generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11282497B2Dynamic text reader for a text document, emotion, and speaker
Publication Date: 2022.03.22 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11282497B2 patent drawing
  • US11282497B2 patent drawing
  • US11282497B2 patent drawing

AI summary

Embodiments are disclosed for a method for dynamic text reading. The method includes performing pre-processing for a text document. Pre-processing includes determining the text document comprises an emotional statement based on an indicator of an emotion associated with the emotional statement. Pre-processing also includes identifying a speaker of the emotional statement. Further, pre-processing includes generating a role-to-voice map that associates the speaker with a digital representation of a voice for the speaker. The method additionally includes generating, based on the pre-processing, the voice for the speaker reading aloud a text of the text document using the digital representation of the voice with a tonal modulation that conveys the emotion.