Service Server Apparatus for Real-Time Speech Translation Error Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current translation systems during telephone calls do not allow speakers to easily recognize or correct translation errors, as the translated content is only transmitted to the called party, preventing the speaker from knowing how their speech is being translated.
Innovation Solution
A service server apparatus that records verbal speeches during calls, performs speech recognition, translation, and synthesis, and provides both text and speech data to the speaker's terminal device, enabling them to see and hear the translation process and correct errors in real-time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If translated speeches are transmitted only to the called party, then the translation service can be provided, but the speaker cannot recognize or correct translation errors
Solution Approach 1:
The patent applies feedback by transmitting task data (translated speeches and related information) back to the speaker's terminal device. This allows the speaker to receive feedback about how their speech was translated, enabling them to recognize errors and request corrections, thus improving translation reliability through iterative refinement
Solution Approach 2:
The patent segments the translation service into multiple components: speech recording, speech recognition, text translation, speech synthesis, and task data generation. By providing task data that includes both the translated speech and source speech information separately, the system enables the speaker to review and correct specific portions of the translation
2Productivity
If speech recognition, translation, and synthesis are performed during telephone call, then translation service is provided, but the speaker cannot easily check or correct translation content
Solution Approach 1:
The system provides feedback to the speaker by transmitting task data that includes the translated speech and source speech information. This feedback mechanism enables the speaker to review the translation content and request corrections, making error correction easy while maintaining high translation efficiency
Solution Approach 2:
The patent creates a copy of the translation process output by generating task data that mirrors the translated speech and source speech information. This copy is transmitted to the speaker's terminal device, allowing them to review and verify the translation without interfering with the real-time translation service efficiency
3Adaptability or versatility
If only translated speeches are transmitted to called party, then communication function is maintained, but speaker cannot know how their speech is translated
Solution Approach 1:
The patent segments the translation output into multiple information components: the translated speech itself and the task data containing source speech information. Both components are transmitted to the speaker's terminal device, preserving complete translation process information while maintaining the communication function
Solution Approach 2:
The system implements multi-functionality by making the terminal device serve multiple purposes: it acts as both a speech communication device for real-time translation and a information display device for showing task data. This allows the speaker to both communicate and review translation information using the same device
Data Source
AI summary
A service server apparatus is provided which can easily cope with a correction of an error of a task performed based on the content of verbal speeches of a speaker. The service server apparatus includes a service activating unit that receives an instruction for performing a different task from a task performed by an application relating to a speech communication, a telephone/call control enabler that records verbal speeches of the speaker during a speech communication between a plurality of speech communication terminal device, a speech recognizing enabler which performs a task based on the recorded speeches and which generates task data including text data representing the result of the performance and speech data representing the result of the performance, a text translating enabler, and a speech synthesizing enabler.


