Chat Voice Message Text Conversion for Silent Confirmation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In chat systems, voice messages are only displayed as icons and cannot be immediately confirmed without reproduction, and converting all voice data to characters degrades visibility.
Innovation Solution
A control method and apparatus that converts voice messages into characters for display in a chat room, allowing users to confirm the content even when voice reproduction is inhibited, using a system with a chat application server and mobile terminals that process and display the converted text alongside the voice message.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If voice messages are displayed only as icons for reproduction, then the system maintains simplicity and audio functionality, but the content cannot be confirmed without reproduction and visibility is poor
Solution Approach 1:
The patent creates a text copy of the voice message content through speech recognition, allowing users to read the message content without needing to reproduce the audio. The text representation serves as an alternative copy of the information, enabling content confirmation in environments where audio output is inhibited or inconvenient.
Solution Approach 2:
The patent introduces text as an intermediary representation between the voice message and the user. Instead of requiring direct audio reproduction, the system converts the voice message into text form, which then serves as the medium for content confirmation. This intermediary text layer resolves the contradiction by providing accessibility without requiring audio playback.
2Loss of information
If all voice data is converted to characters for display, then content visibility is improved, but the original audio functionality and reproduction capability are degraded
Solution Approach 1:
The patent segments the voice message into two independent representations: the original audio data for reproduction and the text conversion for content confirmation. This segmentation allows both functions to coexist without interfering with each other. Users can access the text content when needed while the audio reproduction capability remains intact for when audio output is required.
Solution Approach 2:
The patent implements a dynamic display system where the voice message can be presented in different forms based on user needs or contextual conditions. The system can switch between audio reproduction mode and text display mode, or show both simultaneously. This dynamic adaptability resolves the contradiction by allowing the system to optimize for different requirements without permanently sacrificing either function.
3Reliability
If voice messages require reproduction to confirm content, then audio fidelity is maintained, but time is lost and productivity decreases in environments where audio output is inhibited
Solution Approach 1:
The patent performs preliminary conversion of the voice message into text form before the user needs to confirm the content. The speech recognition processing is executed in advance, creating the text representation that can be immediately displayed and read. This preliminary action eliminates the need for time-consuming audio reproduction and allows instant content confirmation through text reading.
Solution Approach 2:
The patent substitutes the mechanical/audio-based message confirmation process with a text-based processing system. Instead of requiring audio playback and listening, the system uses speech recognition to convert audio to text, which then can be processed and displayed. This substitution replaces the audio mechanism with a text processing mechanism, dramatically improving confirmation efficiency while maintaining audio quality for reproduction when needed.
Data Source
AI summary
A storage medium storing a program for causing a computer to execute a control program for controlling an information processing apparatus, which makes it possible to confirm the content of a voice message in a chat room even in a case where the voice message cannot be reproduced. The information processing apparatus is caused to receive an instruction for converting content of a voice message in a chat room where chats are posted by a plurality of users into characters, execute processing for converting the content of the voice message into characters, and display the characters in a state associated with the voice message in the chat room, based on the instruction.


