Speech Imaging Apparatus Reducing Cognitive Load via Keyword Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies for real-time conversion of utterance details into text during conversations impose a high cognitive load on listeners, making it difficult to promote mutual understanding, especially in self-introduction scenarios.

Innovation Solution

An utterance imaging device that extracts key character strings from recognized sounds, acquires images based on these strings, and outputs them to a position corresponding to the speaker, thereby reducing cognitive load by visualizing conversation details.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If utterance details are displayed as text during conversation, then information completeness is improved, but cognitive load increases

Engineering Contradiction:
Improveinformation completenessVSAvoidcognitive load
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent extracts only key character strings (keywords) from the complete utterance text rather than displaying all text. This selective extraction reduces the information volume presented to users while maintaining the essential meaning, thereby reducing cognitive load while preserving important information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the continuous utterance text into discrete key character strings or keywords. By dividing the information into manageable segments rather than presenting a continuous block of text, the system makes it easier for users to process and understand the conversation content without being overwhelmed.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If all utterance details are converted to text in real time, then conversation accuracy is improved, but processing time increases

Engineering Contradiction:
Improveconversation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts only essential key character strings from the utterance rather than converting and processing the entire text. This extraction approach maintains accuracy for the most important information while significantly reducing the total processing time required.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by converting only the necessary portions of speech (key character strings) into text rather than the complete utterance. This partial conversion achieves sufficient accuracy for conversation understanding while minimizing processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12205610B2Speech imaging apparatus, speech imaging method and program
Publication Date: 2025.01.21 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12205610B2 patent drawing
  • US12205610B2 patent drawing
  • US12205610B2 patent drawing

AI summary

An utterance imaging device reduces a cognitive load on utterance details by extracting some character strings from character strings recognized as sounds from an utterance, acquiring an image based on the some character strings, and outputting the image to a position corresponding to a speaker related to the utterance.