Emotion Metadata for Voice-Text Channel Preservation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies fail to effectively preserve and translate the emotional content of voice and text communications across different languages and cultures, leading to loss of emotional nuance in communication transformations.

Innovation Solution

A system that analyzes voice and text communications to extract emotional metadata, translates this metadata into target languages, and adjusts voice synthesis to match the emotional tone of the original communication, using emotion dictionaries and context profiles to ensure emotional consistency across channels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If voice communication is transformed into text using automatic speech recognition, then text content can be extracted and processed, but emotional content and delivery characteristics are lost

Engineering Contradiction:
Improveemotional contentVSAvoidtext processing capability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the speech signal into two independent components: phonemic content (for text recognition) and emotional delivery characteristics (for emotion preservation). This allows text extraction while maintaining emotional information through separate processing channels, resolving the contradiction between text processing capability and emotional content preservation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces emotion metadata as an intermediary element that bridges the gap between voice communication and text representation. This metadata carries emotional delivery characteristics and is attached to the transcribed text, enabling emotional content to be preserved and transmitted across modalities without sacrificing text processing capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If emotion recognition systems analyze acoustic representations of sub-emotion units, then emotional state can be identified, but the complexity of the analysis increases

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidanalysis complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts emotional delivery characteristics from the full speech signal by focusing on specific acoustic features (pitch, tone, cadence, amplitude) that carry emotional information. This extraction approach isolates the emotional component from the phonemic content, enabling accurate emotion recognition while reducing overall analysis complexity through targeted feature selection

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the complex acoustic signal into simplified emotional parameters (sub-emotion units) that represent emotional state. By changing the representation from raw acoustic waveforms to standardized emotional parameters, the system achieves high emotion recognition accuracy while reducing the complexity of subsequent processing and computation

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If speech is filtered into gender-neutral monotonic audio stream for text recognition, then word recognition accuracy improves, but emotional delivery characteristics are altered

Engineering Contradiction:
Improveword recognition accuracyVSAvoiddelivery characteristics
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the speech processing into two parallel paths: one that filters and processes phonemic content for accurate word recognition, and another that preserves and analyzes emotional delivery characteristics. This segmentation allows each path to optimize for its specific function without compromising the other, resolving the contradiction between recognition accuracy and delivery characteristic preservation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces emotion metadata as an intermediary that captures delivery characteristics without interfering with the text recognition process. This metadata serves as a bridge that conveys emotional information while allowing the phonemic processing to proceed independently with optimal filtering and normalization for accurate word recognition

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If text translation is performed without emotion metadata, then translation speed increases, but emotional nuance and cultural context are lost

Engineering Contradiction:
Improvetranslation speedVSAvoidemotional nuance
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent performs preliminary emotion extraction and metadata generation on the source language speech before translation occurs. This preliminary action captures emotional delivery characteristics in advance, allowing them to be attached to the translated text without adding computational delay to the translation process itself, thus maintaining high productivity while preserving emotional nuance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses emotion metadata as an intermediary that travels with the text through the translation process. This metadata is attached to the source text, processed through translation, and attached to the target text, enabling emotional nuance to be preserved across languages without requiring complex emotion translation algorithms that would slow down the process

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7983910B2Communicating across voice and text channels with emotion preservation
Publication Date: 2011.07.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US7983910B2 patent drawing
  • US7983910B2 patent drawing
  • US7983910B2 patent drawing

AI summary

Communicating across channels with emotion preservation includes: receiving, by a processor in a communication device, a voice communication; analyzing, by the processor in the communication device, the voice communication for first emotion content; analyzing, by the processor in the communication device, textual content of the voice communication for second emotion content; and marking up, by the processor in the communication device, the textual content with emotion metadata for one of the first emotion content and the second emotion content.