Visual Speech Map for Hearing Impaired Users

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals with hearing loss face difficulties in accurately identifying speakers during conversations with multiple participants due to reduced auditory direction perception, leading to challenges in communication.

Innovation Solution

An information processing system comprising a multi-microphone device and a controller that generates a map image displaying speech content at positions corresponding to the direction of sound sources, allowing users to visually distinguish speakers and their contributions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If speeches of other users are displayed in an aggregated state in an image display region set for a certain user, then the display region can accommodate multiple users' speeches, but it becomes difficult to immediately ascertain by whom each speech was made and what each person said

Engineering Contradiction:
Improvenumber of speeches displayedVSAvoidspeaker identification information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent divides the aggregated speech display into separate individual speech displays. Each speech is presented as a distinct element with its own display region, rather than combining multiple speeches into a single aggregated view. This segmentation allows users to clearly identify each speaker and their corresponding speech content while still displaying multiple speeches simultaneously.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If a conversation support apparatus displays text speech recognition results in an image display region, then speech content can be visually presented, but individuals with hearing loss still struggle to identify speakers due to reduced auditory direction perception

Engineering Contradiction:
Improvespeech content informationVSAvoidspeaker direction perception
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent introduces a camera as an intermediary device to capture images of speakers. The captured images are then superimposed with speech recognition results and displayed in correspondence with each speaker's location. This intermediary visual channel compensates for the reduced auditory direction perception, allowing users to identify speakers through visual cues rather than relying solely on auditory direction.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If the apparatus uses only audio-based speech recognition, then the system remains simple, but users with hearing loss cannot accurately perceive the direction of arrival of sound

Engineering Contradiction:
Improvesystem configurationVSAvoidsound direction perception
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges multiple information sources including camera images, speech recognition results, and display elements into a unified presentation system. By combining visual information from cameras with audio-based speech recognition, the system maintains relative simplicity while overcoming the limitation of audio-only direction perception. The merged information is presented in an integrated display that shows both visual and textual information corresponding to each speaker.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240410969A1Information processing apparatus and information processing method
Publication Date: 2024.12.12 PIXIE DUST TECH INC
  • US20240410969A1 patent drawing
  • US20240410969A1 patent drawing
  • US20240410969A1 patent drawing

AI summary

An information processing apparatus includes: acquiring information indicating a direction of a sound source with respect to at least one multi-microphone device; acquiring information regarding content of a speech emitted from the sound source and collected by the multi-microphone device; generating a map image in which the information regarding the content of the speech is arranged at a position corresponding to the direction of the sound source of the sound with respect to the multi-microphone device; and displaying the map image on a display unit of a display device.