Simultaneous Speech Translation via Acoustic Resegmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Language barriers hinder the dissemination of knowledge in globalized settings, as traditional human translation services are costly and often prohibitively expensive, preventing lectures and presentations from being translated in real-time for diverse audiences.

Innovation Solution

A real-time open domain speech translation system that uses automatic speech recognition, resegmentation, and machine translation to simultaneously translate spoken presentations from one language to another, adapting to individual speakers and domains through acoustic and language model tuning, and employing targeted audio and visual output devices for efficient communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human translators are used to overcome language barriers, then translation quality is improved, but cost increases prohibitively

Engineering Contradiction:
Improvetranslation qualityVSAvoidcost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses machine translation systems that create textual copies of the original speech in the target language, replacing human translators. The system transcribes spoken language, translates the text, and outputs translated speech, providing a cost-effective copy-based solution rather than human translation services.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of human translation with an automated speech-to-speech translation system comprising acoustic modeling, language modeling, and text-to-speech synthesis components, eliminating the need for human translators while maintaining translation functionality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If machine translation is used to reduce cost, then cost decreases, but translation quality and accuracy deteriorate

Engineering Contradiction:
ImprovecostVSAvoidtranslation quality
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent divides the translation task into distinct segments: acoustic modeling for speech recognition, language modeling for text generation, and text-to-speech synthesis for output. This segmentation allows each component to be optimized independently, improving overall translation quality while maintaining cost-effectiveness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces text as an intermediary between source speech and target speech. The system transcribes speech to text, translates the text, and synthesizes text to speech, using the textual representation as a mediator to improve translation accuracy through language modeling.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If traditional translation services are used, then translation accuracy is improved, but real-time translation capability is lost

Engineering Contradiction:
Improvetranslation accuracyVSAvoidtranslation delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-training acoustic models, language models, and text-to-speech systems before actual translation. This preparation enables the system to perform real-time translation without delay during the actual speech-to-speech conversion process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous speech recognition and translation processing, where the system continuously transcribes and translates speech in real-time rather than processing discrete segments with delays, maintaining continuous useful action throughout the translation process.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS9830318B2Simultaneous translation of open domain lectures and speeches
Publication Date: 2017.11.28 META PLATFORMS INC
  • US9830318B2 patent drawing
  • US9830318B2 patent drawing
  • US9830318B2 patent drawing

AI summary

Speech translation systems and methods for simultaneously translating speech between first and second speakers, wherein the first speaker speaks in a first language and the second speaker speaks in a second language that is different from the first language. The speech translation system may comprise a resegmentation unit that merge at least two partial hypotheses and resegments the merged partial hypotheses into a first-language translatable segment, wherein a segment boundary for the first-language translatable segment is determined based on sound from the second speaker.