Real-Time Speech Translation Terminal Using Segmented Voice Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional translation apparatuses experience significant time lags in translating speech, hindering smooth conversation between users speaking different languages.
Innovation Solution
A terminal equipment system that includes a voice input unit, speech recognition unit, and translation unit, which processes voice data in real-time to convert speech into character information and translate it into another language, allowing for simultaneous display of both original and translated text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If conventional translation apparatuses process speech translation, then translation functionality is provided, but large time lag occurs before translation starts
Solution Approach 1:
The patent segments the speech recognition process into continuous intervals (e.g., 200ms) rather than waiting for complete speech utterances. The speech recognition unit processes voice data in overlapping time windows, allowing translation to start before the user finishes speaking. This segmentation of the processing timeline directly reduces the time lag between speech input and translation output.
Solution Approach 2:
The system performs preliminary speech recognition on voice data segments before the complete speech is uttered. By continuously processing voice data in advance (predictive processing), the translation apparatus can prepare and display translations before the user completes their sentence, effectively reducing the perceived time lag and enabling more natural conversational flow.
2Speed
If speech is processed in continuous real-time, then translation responsiveness is improved, but processing complexity increases
Solution Approach 1:
The patent divides the continuous speech stream into discrete time segments (e.g., 200ms intervals) that are processed independently. This segmentation allows the system to manage complexity by handling small, manageable units of speech data rather than processing entire continuous streams, while still achieving real-time responsiveness through continuous interval processing.
Solution Approach 2:
The system dynamically adjusts the processing intervals and speech recognition parameters based on the speaking rate and context. The predetermined time interval for speech processing can be adjusted to balance responsiveness with processing complexity, allowing the system to adapt to different speaking conditions while maintaining manageable computational load.
Data Source
AI summary
A speech recognition terminal includes: a voice input unit to accept an input of a voice; a speech recognition command unit to command a speech recognition unit to convert voices of joined voice data acquired by the voice input unit joining the voice data of the voice accepted by the voice input unit to the voice data of the voice accepted previously into character information of a first language at an interval of predetermined time; a translation command unit to command a translation unit to translate first character information of a first language into a second language whenever receiving the first character information of the first language converted by the voice recognition unit; and a display unit to display the first character information of the second language translated by the translation unit together with the first character information of the first language.


