Speech Translation Apparatus Handling Overlapping Utterances
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users may misinterpret translation results due to timing issues in presenting translated speech, as current speech translation systems typically display translations after the segment of speech has ended, leading to potential confusion from overlapping utterances.
Innovation Solution
A speech translation apparatus that includes a speech recognizer, a detector, a machine translator, and a controller, which acquires and processes speech signals, detects segments of meaning, and controls the display order of translated texts based on speaker information and time information to handle overlapping utterances, ensuring timely and accurate presentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If translations are displayed after speech segments end, then translation accuracy is improved, but timing responsiveness deteriorates causing user misinterpretation
Solution Approach 1:
The system performs preliminary segmentation of speech into meaning units before translation completion, preparing the structure for display in advance. This allows the translation to be displayed as soon as it becomes available without waiting for the entire speech segment to end, resolving the timing conflict while maintaining accuracy.
Solution Approach 2:
The speech is divided into multiple meaning segments rather than treating it as a single unit. Each segment can be translated and displayed independently at its optimal time, allowing some translations to appear earlier while others wait for segment completion, thus balancing accuracy and responsiveness.
2Device complexity
If translations are displayed in chronological order, then processing simplicity is improved, but understanding of overlapping utterances deteriorates
Solution Approach 1:
Different display strategies are applied to different translation segments based on their temporal and speaker characteristics. Translations from overlapping utterances are distinguished through visual differentiation (such as color coding or positioning), while non-overlapping translations use simple chronological display, optimizing both simplicity and clarity locally.
Solution Approach 2:
The system introduces an intermediary display layer that sits between the chronological translation output and the user. This intermediary organizes and presents translations with speaker identification information, using visual cues to indicate which speaker made each utterance even when translations arrive in chronological order.
3Quantity of substance
If all translations are displayed simultaneously, then information completeness is improved, but user comprehension deteriorates due to information overload
Solution Approach 1:
The complete set of translations is segmented and displayed in a structured hierarchy rather than all at once. Translations are organized by speaker, time, or meaning unit, allowing users to comprehend information progressively while still receiving the complete information set, thus balancing completeness with ease of understanding.
Data Source
AI summary
According to one embodiment, a speech translation apparatus includes a speech recognizer, a detector, a machine translator and a controller. The speech recognizer performs a speech recognition processing in chronological order on utterances of at least one first language made by a plurality of speakers to obtain a recognition text as a speech recognition result. The detector detects segments of meaning of the recognition text to obtain segments of text. The machine translator translates the segments of text into a second language different from the first language to obtain translated texts. The controller controls, if an utterance overlaps with another utterance in the chronological order, an order of displaying the translated texts corresponding to the overlapped utterances.


