Speech Interpretation Apparatus Reducing Translation Delay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech interpretation apparatuses experience delays in translating speech audio to text, leading to lagged display of subtitles, making it difficult for viewers to understand the meaning, especially in continuous lectures, as the output duration of translation results is often insufficient or excessively reduced.
Innovation Solution
An interpretation apparatus comprising a speech recognition unit, a translator, a calculator, and a generator that calculates the number of words to be omitted from the machine translation result based on time and output duration, generating abridged sentences to reduce delay and improve understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the output duration of the translation result is extended to ensure viewer understanding, then the comprehension quality improves, but the delay from speech start to translation output accumulatively increases
Solution Approach 1:
The translation output is segmented into multiple subtitle displays rather than showing the entire translation at once. Each subtitle displays a portion of the translation result for a limited duration, then transitions to the next segment. This segmentation allows the system to maintain shorter display durations (reducing delay accumulation) while still conveying the complete meaning through sequential segments.
Solution Approach 2:
The system employs periodic action by continuously updating and replacing translation subtitles at regular intervals. Instead of displaying one long translation for an extended period, the system periodically generates new translation results and displays them in sequence, maintaining a rhythm that balances viewer comprehension with timely updates that prevent excessive delay accumulation.
2Loss of time
If the output duration of the translation result is reduced to decrease delay, then the timeliness improves, but the viewer may not finish reading and understanding the meaning
Solution Approach 1:
By dividing the full translation into multiple subtitle segments displayed sequentially, the system can present each segment within a shorter time window. This allows the overall translation to be delivered over time without requiring any single subtitle to remain on screen for an excessively long duration, thus maintaining timeliness while ensuring complete comprehension through the sequence of segments.
Solution Approach 2:
The system dynamically adjusts the display duration and timing of each subtitle segment based on the ongoing speech and translation generation process. This dynamic approach allows the subtitles to be updated in real-time, adapting to the speech pace and ensuring that each segment is displayed long enough to be read while maintaining overall synchronization with the speech flow.
3Reliability
If the translation result is displayed continuously for a long time, then the viewer can understand the meaning, but the delay accumulatively increases making it difficult to understand
Solution Approach 1:
The translation is segmented into multiple manageable subtitle displays that are presented sequentially rather than as one continuous long display. This segmentation makes the information more digestible for viewers, allowing them to process each segment individually while still understanding the complete meaning through the accumulation of segments, thereby improving ease of understanding.
Solution Approach 2:
The system uses periodic action by continuously refreshing and updating translation subtitles at regular intervals. This periodic update mechanism ensures that the translation remains synchronized with the speech while presenting information in manageable chunks, making it easier for viewers to follow and understand without being overwhelmed by a single continuous long display.
Data Source
AI summary
According to one embodiment, an interpretation apparatus includes a translator, a calculator and a generator. The translator performs machine translation on a speech recognition result corresponding to an input speech audio from a first language into a second language to generate a machine translation result. The calculator calculates a word number based on a first time when the machine translation result is generated and a second time when output relating to a prior machine translation result generated prior to the machine translation result ends, the word number being 0 or larger. The generator omits at least the word number of words from the machine translation result to generate an abridged sentence output while being associated with the speech audio.


