Emotion Metadata Extraction in Speech Translation Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional language translation systems, such as ASR, fail to effectively convey the emotional state of a speaker beyond phonemes, lacking the ability to utilize tone, timbre, and speech gender to determine emotions, resulting in a limited understanding of delivery in human communication.

Innovation Solution

A system and method that employ non-speaker dependent models to analyze acoustic information like pitch, tone, cadence, and amplitude to recognize human emotions, generating emotion metadata that can be used to adjust the delivery of translated language, incorporating user-specific and general filters to refine emotional interpretation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional ASR systems filter speech into phonemes and translate to gender-neutral monotonic audio, then speech recognition accuracy is improved, but emotional information and delivery characteristics are lost

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidemotional information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system segments speech analysis into multiple parallel pathways: one pathway extracts phonemes for accurate speech recognition, while another pathway analyzes acoustic features (pitch, tone, timbre, cadence) for emotion detection. This segmentation allows both speech accuracy and emotional information to be preserved simultaneously without interference.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds an emotional dimension to the traditional speech recognition output by generating emotion metadata that complements the phoneme-based translation. This transforms the monodimensional phoneme stream into a multidimensional output that includes both linguistic content and emotional delivery characteristics.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If speaker-dependent models are used for emotion recognition, then emotion detection accuracy is improved, but system complexity and data requirements increase

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system employs universal acoustic feature extraction that works across different speakers and languages without requiring speaker-specific training data. The same acoustic analysis pipeline (analyzing pitch, tone, timbre, cadence) applies universally to all speakers, eliminating the need for complex speaker-dependent models while maintaining effectiveness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the approach from speaker-specific parameter modeling to speaker-independent acoustic parameter analysis. By focusing on fundamental acoustic characteristics (pitch contours, tone patterns, timbre variations, cadence rhythms) that manifest similarly across speakers, the system achieves emotion recognition without the complexity of individual speaker adaptation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11587561B2Communication system and method of extracting emotion data during translations
Publication Date: 2023.02.21 WEIR MARY LEE
  • US11587561B2 patent drawing
  • US11587561B2 patent drawing
  • US11587561B2 patent drawing

AI summary

A communication system is provided that generates human emotion metadata during language translation of verbal content. The communication system includes a media control unit that is coupled to a communication device and a translation server that receive verbal content from the communication device in a first language. An adapter layer having a plurality of filters determines emotion associated with the verbal content, wherein the adapter layer associates emotion metadata with the verbal content based on the determined emotion. The plurality of filters may include user-specific filters and non-user-specific filters. An emotion lexicon is provided that links an emotion value to the corresponding verbal content. The communication system may include a display that graphically displays emotions alongside the corresponding verbal content.