Real-time Voice-to-Text Correction in Conferencing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional conferencing systems face difficulties in smoothly transmitting and receiving accurate text information when errors occur in voice-to-text conversions, leading to potential miscommunication between users.
Innovation Solution
An information processing system that includes a voice receiver, voice recognizer, display controller, and correction reception portion, allowing for real-time display and correction of text information across connected devices via a network, enabling users to correct errors and ensure accurate communication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If voice-to-text conversion is performed automatically, then text information can be transmitted efficiently, but errors in the converted text may occur leading to miscommunication
Solution Approach 1:
The patent implements a feedback mechanism where the converted text is displayed to the user before transmission, allowing the user to review and correct any recognition errors. The system provides visual feedback of the converted text and awaits user confirmation or correction, ensuring accuracy while maintaining efficient transmission workflow
Solution Approach 2:
The patent introduces the user as an intermediary between the automatic voice-to-text conversion and the final text transmission. The user acts as a mediator to verify and correct the converted text, bridging the gap between automated conversion efficiency and human judgment accuracy
2Device complexity
If text information is displayed only in a single area, then the interface is simple, but the user cannot easily compare original and corrected text
Solution Approach 1:
The patent divides the display area into multiple sections: a first display area showing the original converted text and a second display area showing the corrected text. This segmentation allows users to easily compare original and modified content while maintaining a structured, organized interface layout
3Reliability
If correction functionality is added to the conferencing system, then text accuracy improves, but the system complexity increases
Solution Approach 1:
The patent integrates the correction functionality into the existing conferencing system's display interface, making the display component serve multiple functions: showing converted text, displaying corrected text, and enabling user editing. This multi-functionality reduces the need for separate correction modules, thereby limiting complexity increase
Data Source
AI summary
An information processing system according one embodiment includes: a voice receiver which receives a first voice uttered by a first user of a first information processing device; a voice recognizer which recognizes the first voice received by the voice receiver; a display controller which causes a first text, which corresponds to the first voice recognized by the voice recognizer, to be displayed in each of first display areas of the first information processing device and a second information processing device, and a second display area of the first information processing device; and a correction reception portion which receives a correction operation of the first user for the first text displayed in the second display area.


