Speech Translation Apparatus Handling Overlapping Utterances

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users may misinterpret translation results due to timing issues in presenting translated speech, as current speech translation systems typically display translations after the segment of speech has ended, leading to potential confusion from overlapping utterances.

Innovation Solution

A speech translation apparatus that includes a speech recognizer, a detector, a machine translator, and a controller, which acquires and processes speech signals, detects segments of meaning, and controls the display order of translated texts based on speaker information and time information to handle overlapping utterances, ensuring timely and accurate presentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If translations are displayed after speech segments end, then translation accuracy is improved, but timing responsiveness deteriorates causing user misinterpretation

Engineering Contradiction:
Improvetranslation accuracyVSAvoidtiming responsiveness
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary segmentation of speech into meaning units before translation completion, preparing the structure for display in advance. This allows the translation to be displayed as soon as it becomes available without waiting for the entire speech segment to end, resolving the timing conflict while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech is divided into multiple meaning segments rather than treating it as a single unit. Each segment can be translated and displayed independently at its optimal time, allowing some translations to appear earlier while others wait for segment completion, thus balancing accuracy and responsiveness.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If translations are displayed in chronological order, then processing simplicity is improved, but understanding of overlapping utterances deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoidspeaker identification clarity
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

Different display strategies are applied to different translation segments based on their temporal and speaker characteristics. Translations from overlapping utterances are distinguished through visual differentiation (such as color coding or positioning), while non-overlapping translations use simple chronological display, optimizing both simplicity and clarity locally.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system introduces an intermediary display layer that sits between the chronological translation output and the user. This intermediary organizes and presents translations with speaker identification information, using visual cues to indicate which speaker made each utterance even when translations arrive in chronological order.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If all translations are displayed simultaneously, then information completeness is improved, but user comprehension deteriorates due to information overload

Engineering Contradiction:
Improveinformation completenessVSAvoiduser comprehension
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The complete set of translations is segmented and displayed in a structured hierarchy rather than all at once. Translations are organized by speaker, time, or meaning unit, allowing users to comprehend information progressively while still receiving the complete information set, thus balancing completeness with ease of understanding.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9600475B2Speech translation apparatus and method
Publication Date: 2017.03.21 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9600475B2 patent drawing
  • US9600475B2 patent drawing
  • US9600475B2 patent drawing

AI summary

According to one embodiment, a speech translation apparatus includes a speech recognizer, a detector, a machine translator and a controller. The speech recognizer performs a speech recognition processing in chronological order on utterances of at least one first language made by a plurality of speakers to obtain a recognition text as a speech recognition result. The detector detects segments of meaning of the recognition text to obtain segments of text. The machine translator translates the segments of text into a second language different from the first language to obtain translated texts. The controller controls, if an utterance overlaps with another utterance in the chronological order, an order of displaying the translated texts corresponding to the overlapped utterances.