Interactive Speech Translation Interface for Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine translation systems lack user interaction capabilities for manipulating and correcting translated sentences, limiting the accuracy and user satisfaction in language translation processes.

Innovation Solution

A system that allows users to interact with machine translations through a graphical user interface (GUI), enabling the breakdown of sentences into phrases and words, allowing selection, correction, and modification of translated text using touch or voice commands, and providing dictionary definitions and alternative translations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine translation systems operate automatically without user interaction, then productivity is improved, but translation accuracy and user satisfaction deteriorate

Engineering Contradiction:
Improvetranslation speedVSAvoidtranslation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system implements feedback loops where user corrections are captured and fed back into the translation system. The GUI collects user edits and selections, which are then used to refine and improve future translations, creating a continuous improvement cycle that maintains both speed and accuracy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The graphical user interface acts as an intermediary between the automatic translation system and the user. It provides tools for sentence manipulation, word selection, and correction while maintaining the efficiency of automated translation, mediating between speed and accuracy requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If machine translation systems provide only automatic output without interaction capabilities, then device complexity is reduced, but ease of operation deteriorates

Engineering Contradiction:
Improvesystem structureVSAvoiduser control capability
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The system segments the translation process into distinct interactive components: sentence-level manipulation, phrase-level editing, and word-level correction. This segmentation allows users to control translations at different granularities, improving ease of operation while maintaining manageable system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system provides dynamic interaction capabilities where users can adjust their level of involvement based on needs. The GUI enables flexible manipulation of translated text at various levels (sentence, phrase, word), allowing the system to adapt to different user requirements without requiring complete system redesign.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If machine translation systems allow extensive user interaction and manipulation, then translation accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvetranslation accuracyVSAvoidinterface complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The graphical user interface is designed with multi-functional elements that perform multiple operations through unified interactions. For example, the same interface mechanisms handle sentence breakdown, phrase selection, word editing, and correction submission, reducing interface complexity while providing comprehensive translation control capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11900072B1Quick lookup for speech translation
Publication Date: 2024.02.13 AMAZON TECH INC
  • US11900072B1 patent drawing
  • US11900072B1 patent drawing
  • US11900072B1 patent drawing

AI summary

Offered is a system that presents on a display screen a translation of a sentence together with an untranslated version of the sentence, and that can cause both of the displayed sentences to break apart into component parts in response to a simple user action, e.g., double-tapping on one of them. When the user selects (e.g., taps on) any portion of either version of the sentence, the system can identify a corresponding portion of the other version (in the other language). In some implementations, a user device can include both a microphone and a display screen, and an automatic speech recognition (ASR) engine can be used to transcribe the user's speech in one language (e.g., English) into text. The system can translate the resulting text into another language (e.g., Spanish) and display the translated text on the display screen along with the untranslated text. When a user selects a portion of a sentence, the system can also present information about the selected portion (e.g., a dictionary definition) and/or enable selection of a modification to the selected portion (e.g., a different one of the N-best results that were generated by a translation service or an automatic speech recognition (ASR) service).