Real-Time Speech Translation Terminal Using Segmented Voice Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional translation apparatuses experience significant time lags in translating speech, hindering smooth conversation between users speaking different languages.

Innovation Solution

A terminal equipment system that includes a voice input unit, speech recognition unit, and translation unit, which processes voice data in real-time to convert speech into character information and translate it into another language, allowing for simultaneous display of both original and translated text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If conventional translation apparatuses process speech translation, then translation functionality is provided, but large time lag occurs before translation starts

Engineering Contradiction:
Improvetime lagVSAvoidtranslation speed
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The patent segments the speech recognition process into continuous intervals (e.g., 200ms) rather than waiting for complete speech utterances. The speech recognition unit processes voice data in overlapping time windows, allowing translation to start before the user finishes speaking. This segmentation of the processing timeline directly reduces the time lag between speech input and translation output.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary speech recognition on voice data segments before the complete speech is uttered. By continuously processing voice data in advance (predictive processing), the translation apparatus can prepare and display translations before the user completes their sentence, effectively reducing the perceived time lag and enabling more natural conversational flow.

Inventive Principle:
Principle #10Preliminary action

2Speed

If speech is processed in continuous real-time, then translation responsiveness is improved, but processing complexity increases

Engineering Contradiction:
Improvetranslation responsivenessVSAvoidprocessing complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent divides the continuous speech stream into discrete time segments (e.g., 200ms intervals) that are processed independently. This segmentation allows the system to manage complexity by handling small, manageable units of speech data rather than processing entire continuous streams, while still achieving real-time responsiveness through continuous interval processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts the processing intervals and speech recognition parameters based on the speaking rate and context. The predetermined time interval for speech processing can be adjusted to balance responsiveness with processing complexity, allowing the system to adapt to different speaking conditions while maintaining manageable computational load.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10489516B2Speech recognition and translation terminal, method and non-transitory computer readable medium
Publication Date: 2019.11.26 FUJITSU LTD
  • US10489516B2 patent drawing
  • US10489516B2 patent drawing
  • US10489516B2 patent drawing

AI summary

A speech recognition terminal includes: a voice input unit to accept an input of a voice; a speech recognition command unit to command a speech recognition unit to convert voices of joined voice data acquired by the voice input unit joining the voice data of the voice accepted by the voice input unit to the voice data of the voice accepted previously into character information of a first language at an interval of predetermined time; a translation command unit to command a translation unit to translate first character information of a first language into a second language whenever receiving the first character information of the first language converted by the voice recognition unit; and a display unit to display the first character information of the second language translated by the translation unit together with the first character information of the first language.