Speech Translation System with Interactive User Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine translation systems are unreliable for real-time cross-lingual communication due to issues like word-sense disambiguation, limited semantically rich dictionary information, and incompatible interfaces, leading to errors and a lack of convenient, cost-effective solutions for bridging language gaps, especially in informal or impromptu situations.

Innovation Solution

A method and system for real-time speech and text translation that involves mapping word senses across lexical resources by selecting a target term, identifying possible matches, calculating semantic distance, and ranking them for relevance, along with a multi-modal user interface that allows users to verify and correct translations, using a database of Meaning Cues to facilitate accurate communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If currently available machine translation tools are used, then translation capability is provided, but reliability and accuracy deteriorate due to errors in speech recognition and translation processing

Engineering Contradiction:
Improvetranslation accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where the system presents translation results to users for verification and correction. Users can review translated text, identify errors, and provide corrections that are fed back into the system. This feedback loop continuously improves translation accuracy by learning from user corrections and refining the translation model over time, directly addressing the reliability issue while maintaining manageable complexity through iterative improvement.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary verification step between speech recognition and final translation output. Users act as intermediaries who review and validate the translation results before final acceptance. This intermediary layer allows the system to benefit from human judgment in resolving ambiguous cases without requiring complete redesign of the underlying complex processing systems, thus improving reliability while leveraging existing infrastructure.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If human interpreters are used for live meetings and phone calls, then communication accuracy is improved, but cost and scheduling difficulty increase

Engineering Contradiction:
Improvecommunication accuracyVSAvoidaccessibility
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements a self-service machine translation system that users can access independently without requiring human interpreters. The system provides automated speech-to-speech translation capabilities that users can invoke on-demand for live meetings and phone calls. This self-service approach eliminates the need for scheduling and coordinating human interpreters while maintaining accessibility through automated, on-demand translation services.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a digital copy of human interpretation capability through automated speech recognition and machine translation systems. Instead of relying on physical human interpreters, the system uses software-based translation engines that replicate interpretation functions. This digital copying enables widespread accessibility without the logistical constraints of human interpreter availability, directly resolving the contradiction between accuracy and ease of access.

Inventive Principle:
Principle #26Copying

3Reliability

If written translation services are used, then translation quality can be improved, but time delays become unacceptably long

Engineering Contradiction:
Improvetranslation qualityVSAvoidtranslation delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements periodic, real-time translation processing that occurs continuously during speech communication. Rather than batch processing written text, the system performs periodic speech-to-speech translation as speech flows, providing near-real-time translation output. This periodic action maintains translation quality through systematic processing while eliminating unacceptable time delays by translating speech as it is spoken rather than waiting for written input.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent replaces the mechanical process of written translation with automated speech-to-speech translation technology. Instead of relying on human translators working with written text, the system uses speech recognition, automated translation, and speech synthesis to convert spoken language directly to spoken language. This substitution eliminates the time-consuming written translation process while maintaining or improving quality through automated real-time processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Productivity

If speech recognition systems are combined with machine translation tools, then real-time translation capability is achieved, but error rates compound at each processing step

Engineering Contradiction:
Improvereal-time translation capabilityVSAvoiderror rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms at each processing stage to detect and correct errors before they compound. The system provides intermediate results to users for verification, allowing error detection and correction at the speech recognition stage before translation processing. This feedback loop prevents error propagation by enabling users to correct recognition errors early in the process, thereby maintaining productivity while reducing cumulative error rates through iterative validation.

Inventive Principle:
Principle #23Feedback

5Ease of operation

If users are permitted to rewrite input sentences, then some control over translation is achieved, but the system still permits little or no user intervention in the recognition or translation processes

Engineering Contradiction:
Improveuser controlVSAvoidsystem control complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a dynamic user control system that adapts to user needs during the translation process. Users can intervene at multiple stages including speech recognition verification, translation option selection, and result validation. The system dynamically adjusts its level of automation based on user input, allowing greater user control when needed while maintaining automated operation for routine cases. This dynamic approach enables flexible user control without requiring complex manual configuration or system reconfiguration.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS7539619B1Speech-enabled language translation system and method enabling interactive user supervision of translation and speech recognition accuracy
Publication Date: 2009.05.26 ZAMA INNOVATIONS LLC
  • US7539619B1 patent drawing
  • US7539619B1 patent drawing
  • US7539619B1 patent drawing

AI summary

A system and method for a highly interactive style of speech-to-speech translation is provided. The interactive procedures enable a user to recognize, and if necessary correct, errors in both speech recognition and translation, thus providing robust translation output than would otherwise be possible. The interactive techniques for monitoring and correcting word ambiguity errors during automatic translation, search, or other natural language processing tasks depend upon the correlation of Meaning Cues and their alignment with, or mapping into, the word senses of third party lexical resources, such as those of a machine translation or search lexicon. This correlation and mapping can be carried out through the creation and use of a database of Meaning Cues, i.e., SELECT. Embodiments described above permit the intelligent building and application of this database, which can be viewed as an interlingua, or language-neutral set of meaning symbols, applicable for many purposes. Innovative techniques for interactive correction of server-based speech recognition are also described.