Chat Voice Message Text Conversion for Silent Confirmation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In chat systems, voice messages are only displayed as icons and cannot be immediately confirmed without reproduction, and converting all voice data to characters degrades visibility.

Innovation Solution

A control method and apparatus that converts voice messages into characters for display in a chat room, allowing users to confirm the content even when voice reproduction is inhibited, using a system with a chat application server and mobile terminals that process and display the converted text alongside the voice message.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If voice messages are displayed only as icons for reproduction, then the system maintains simplicity and audio functionality, but the content cannot be confirmed without reproduction and visibility is poor

Engineering Contradiction:
Improvemessage content accessibilityVSAvoidoperation complexity
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent creates a text copy of the voice message content through speech recognition, allowing users to read the message content without needing to reproduce the audio. The text representation serves as an alternative copy of the information, enabling content confirmation in environments where audio output is inhibited or inconvenient.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces text as an intermediary representation between the voice message and the user. Instead of requiring direct audio reproduction, the system converts the voice message into text form, which then serves as the medium for content confirmation. This intermediary text layer resolves the contradiction by providing accessibility without requiring audio playback.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If all voice data is converted to characters for display, then content visibility is improved, but the original audio functionality and reproduction capability are degraded

Engineering Contradiction:
Improvemessage content visibilityVSAvoidaudio reproduction capability
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent segments the voice message into two independent representations: the original audio data for reproduction and the text conversion for content confirmation. This segmentation allows both functions to coexist without interfering with each other. Users can access the text content when needed while the audio reproduction capability remains intact for when audio output is required.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a dynamic display system where the voice message can be presented in different forms based on user needs or contextual conditions. The system can switch between audio reproduction mode and text display mode, or show both simultaneously. This dynamic adaptability resolves the contradiction by allowing the system to optimize for different requirements without permanently sacrificing either function.

Inventive Principle:
Principle #15Dynamics

3Reliability

If voice messages require reproduction to confirm content, then audio fidelity is maintained, but time is lost and productivity decreases in environments where audio output is inhibited

Engineering Contradiction:
Improveaudio qualityVSAvoidmessage confirmation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary conversion of the voice message into text form before the user needs to confirm the content. The speech recognition processing is executed in advance, creating the text representation that can be immediately displayed and read. This preliminary action eliminates the need for time-consuming audio reproduction and allows instant content confirmation through text reading.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes the mechanical/audio-based message confirmation process with a text-based processing system. Instead of requiring audio playback and listening, the system uses speech recognition to convert audio to text, which then can be processed and displayed. This substitution replaces the audio mechanism with a text processing mechanism, dramatically improving confirmation efficiency while maintaining audio quality for reproduction when needed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240283761A1Storage medium storing program for displaying voice message, control method, and information processing apparatus
Publication Date: 2024.08.22 CANON KK
  • US20240283761A1 patent drawing
  • US20240283761A1 patent drawing
  • US20240283761A1 patent drawing

AI summary

A storage medium storing a program for causing a computer to execute a control program for controlling an information processing apparatus, which makes it possible to confirm the content of a voice message in a chat room even in a case where the voice message cannot be reproduced. The information processing apparatus is caused to receive an instruction for converting content of a voice message in a chat room where chats are posted by a plurality of users into characters, execute processing for converting the content of the voice message into characters, and display the characters in a state associated with the voice message in the chat room, based on the instruction.