Speech Imaging Apparatus Reducing Cognitive Load via Keyword Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies for real-time conversion of utterance details into text during conversations impose a high cognitive load on listeners, making it difficult to promote mutual understanding, especially in self-introduction scenarios.
Innovation Solution
An utterance imaging device that extracts key character strings from recognized sounds, acquires images based on these strings, and outputs them to a position corresponding to the speaker, thereby reducing cognitive load by visualizing conversation details.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If utterance details are displayed as text during conversation, then information completeness is improved, but cognitive load increases
Solution Approach 1:
The patent extracts only key character strings (keywords) from the complete utterance text rather than displaying all text. This selective extraction reduces the information volume presented to users while maintaining the essential meaning, thereby reducing cognitive load while preserving important information.
Solution Approach 2:
The patent segments the continuous utterance text into discrete key character strings or keywords. By dividing the information into manageable segments rather than presenting a continuous block of text, the system makes it easier for users to process and understand the conversation content without being overwhelmed.
2Measurement precision
If all utterance details are converted to text in real time, then conversation accuracy is improved, but processing time increases
Solution Approach 1:
The system extracts only essential key character strings from the utterance rather than converting and processing the entire text. This extraction approach maintains accuracy for the most important information while significantly reducing the total processing time required.
Solution Approach 2:
The patent applies partial action by converting only the necessary portions of speech (key character strings) into text rather than the complete utterance. This partial conversion achieves sufficient accuracy for conversation understanding while minimizing processing time.
Data Source
AI summary
An utterance imaging device reduces a cognitive load on utterance details by extracting some character strings from character strings recognized as sounds from an utterance, acquiring an image based on the some character strings, and outputting the image to a position corresponding to a speaker related to the utterance.


