Real-Time Speech-to-Speech Interpretation via Inversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional language interpretation methods, particularly in customer care environments, are inefficient and cumbersome due to consecutive dialogue, which wastes computing resources and disrupts communication between users speaking different languages.

Innovation Solution

A remote multi-channel language interpretation system that performs simultaneous interpretation using a processor to translate spoken language queries and responses, generating audio and image data corresponding to the target language, allowing users to communicate naturally without waiting for interpretations, by utilizing a language interpretation platform that includes a processor, memory, and databases for imagery and gesture manipulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If consecutive language interpretation is used, then language translation is achieved, but communication efficiency deteriorates and computing resources are wasted

Engineering Contradiction:
Improvelanguage translation accuracyVSAvoidcommunication efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

Instead of translating speech to text then text to speech (consecutive interpretation), the system inverts the approach by directly translating speech to speech in real-time using audio data processing and synthesis, eliminating the intermediate text conversion step and enabling simultaneous interpretation

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system maintains continuous communication flow by performing real-time speech-to-speech translation without interrupting the dialogue, allowing both parties to speak and be understood simultaneously rather than taking turns in consecutive interpretation

Inventive Principle:
Principle #20Continuity of useful action

2Reliability

If consecutive dialogue interpretation is implemented, then language barriers are overcome, but time consumption increases

Engineering Contradiction:
Improveinterpretation accuracyVSAvoidwaiting time for interpretation
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-processing audio data and preparing translation models in advance, enabling real-time speech translation without requiring waiting time during the actual conversation, thus eliminating the pause-and-translate pattern of consecutive interpretation

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If extensive language pair libraries are stored, then comprehensive language support is provided, but computing resource requirements increase

Engineering Contradiction:
Improvelanguage pair coverageVSAvoidcomputing resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system achieves universality by using a single multi-lingual speech translation model that can handle multiple language pairs simultaneously, eliminating the need to store and process separate extensive libraries for each language pair, thus reducing computing resource requirements while maintaining comprehensive language support

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10839801B2Configuration for remote multi-channel language interpretation performed via imagery and corresponding audio at a display-based device
Publication Date: 2020.11.17 LANGUAGE LINE SERVICES INC
  • US10839801B2 patent drawing
  • US10839801B2 patent drawing
  • US10839801B2 patent drawing

AI summary

A configuration is implemented to receive, with a processor from a customer care platform, a request for spoken language interpretation of a user query from a first spoken language to a second spoken language. The first spoken language is spoken by a user situated at a display-based device that is remotely situated from the customer care platform. The user query is sent from the display-based device by the user to the customer care platform. The configuration performs, at a language interpretation platform, a first spoken language interpretation of the user query from the first spoken language to the second spoken language. Further, the configuration transmits, from the language interpretation platform to the customer care platform, the first spoken language interpretation so that a customer care representative speaking the second spoken language understands the first spoken language being spoken by the user.