Audio Augmentation with Visual Features for Remote Interpretation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional over-the-phone interpretation (OPI) systems rely solely on auditory cues, limiting the efficacy of human-spoken language interpretation due to remoteness and limited data availability, which can lead to inaccuracies and context misunderstandings.

Innovation Solution

A configuration that augments voice-based interpretation with real-time visual features, such as facial expressions and hand gestures, to enhance language interpretation by simulating a video remote interpretation (VRI) session, either using a human or machine interpreter, thereby improving interpretation accuracy and context understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional OPI configurations are used, then device complexity is reduced, but interpretation accuracy deteriorates due to limited auditory cues

Engineering Contradiction:
Improveinterpretation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines audio data with visual features (facial expressions, gestures, body language) into a unified interpretation system. The processor merges multiple data types from different sources to create a comprehensive set of augmented audio data that enhances interpretation accuracy while managing system complexity through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to handle multiple types of input data (audio, visual features, contextual information) through a single processor that performs multiple functions: receiving audio data, obtaining visual features, augmenting audio with visual information, and generating interpretation output. This multi-functional approach improves accuracy without proportionally increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If visual features are added to augment audio, then interpretation accuracy is improved, but data processing requirements increase

Engineering Contradiction:
Improveinterpretation accuracyVSAvoiddata processing load
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system extracts only the most relevant visual features (facial expressions, gestures, body language) from the visual data stream, rather than processing all visual information. This selective extraction reduces the data processing load while maintaining the interpretative value needed for accurate language interpretation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Visual features are obtained and prepared in advance to accompany the audio data before interpretation occurs. The system pre-processes visual information to identify and organize relevant cues, reducing the computational burden during real-time interpretation and minimizing data processing requirements during critical moments.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If real-time visual augmentation is implemented, then context understanding is enhanced, but network latency increases

Engineering Contradiction:
Improvecontext understandingVSAvoidnetwork latency
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

Visual features are captured and pre-processed before being transmitted to the interpretation system. By preparing visual data in advance and buffering it appropriately, the system reduces the time required for real-time processing and transmission, thereby minimizing network latency while maintaining context understanding.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system transmits only the essential visual features needed for context understanding rather than complete visual data streams. This partial transmission approach provides sufficient contextual information for accurate interpretation while reducing the data volume and associated network latency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10613827B2Configuration for simulating a video remote interpretation session
Publication Date: 2020.04.07 LANGUAGE LINE SERVICES INC
  • US10613827B2 patent drawing
  • US10613827B2 patent drawing
  • US10613827B2 patent drawing

AI summary

A configuration receives, with a processor, a request for a voice-based, human-spoken language interpretation from a first human-spoken language to a second human-spoken language. Further, the configuration routes, with the processor, the request to a device associated with a remotely-located human interpreter. In addition, the configuration receives, with the processor, audio in the first human-spoken language from a telecommunication device. The configuration also augments, in real-time with the processor, the audio with one or more visual features corresponding to the audio. Further, the configuration sends, with the processor, the augmented audio to the device associated with the human interpreter for the voice-based, human-spoken language interpretation to be based on the augmented audio in a simulated video remote interpretation session.