Sound Source Image Overlay for Utterance Direction Display

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound data recording and reproducing devices struggle to intuitively convey utterance situations, as they separately display sound source directions and content, making it difficult for viewers to understand which sound represents what utterance, especially in multi-speaker scenarios.

Innovation Solution

An information processing device and system that creates display data with characters representing utterance content and symbols indicating sound direction, combining this data with images of sound sources to match the sound radiation orientation, allowing viewers to easily understand utterance situations by positioning and formatting the display based on sound source positions and emotions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If sound data is separated and displayed separately from sound source direction, then sound data processing is simplified, but it becomes difficult for viewers to intuitively understand utterance situations

Engineering Contradiction:
Improvesound data processing complexityVSAvoidviewer understanding of utterance situations
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The patent merges sound data display with sound source direction indication by overlaying text representing utterance content onto images of sound sources. This integration allows viewers to simultaneously see both the speaker and the transcribed content in a single visual field, eliminating the confusion that arises from separate displays while maintaining processing simplicity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary display layer that combines audio information (transcribed text) with visual information (sound source images) through spatial positioning. This intermediary composite display serves as a bridge between the separate audio processing components and the viewer's understanding, making the connection between sound sources and utterances immediately apparent.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If display data is shown separately from sound source images, then text processing is simplified, but the connection between sound and visual information is lost

Engineering Contradiction:
Improvetext processing complexityVSAvoidconnection between sound and visual information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent merges text display with sound source imaging by positioning transcribed text directly on or near the corresponding speaker image. This visual combination preserves the connection between audio content and visual source while keeping text processing independent and simple, as the text generation remains a separate module that simply overlays information rather than processing complex relationships.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If multiple sound sources are displayed separately, then sound separation is improved, but it becomes difficult to understand which sound represents which utterance

Engineering Contradiction:
Improvesound source separation accuracyVSAvoidviewer ability to match sound with utterance
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent combines separated sound source identification with utterance display by assigning unique visual identifiers (images and spatial positions) to each sound source and overlaying corresponding transcribed text on those same identifiers. This allows the system to maintain precise sound separation for processing while providing clear visual matching for viewer understanding through consistent spatial and visual association.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent applies local quality differentiation by displaying distinct visual characteristics (different images, positions, or visual styles) for different sound sources. Each sound source receives customized visual treatment that matches its identity, making it easier for viewers to distinguish and match specific sounds with their corresponding utterances while maintaining the benefits of sound separation.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8886530B2Displaying text and direction of an utterance combined with an image of a sound source
Publication Date: 2014.11.11 HONDA MOTOR CO LTD
  • US8886530B2 patent drawing
  • US8886530B2 patent drawing
  • US8886530B2 patent drawing

AI summary

An information processing device includes a display data creating unit configured to create display data including characters representing the content of an utterance based on a sound and a symbol surrounding the characters and indicating a first direction, and an image combining unit configured to determine the position of the display data based on a display position of an image representing a sound source of the utterance, and to combine the display data and the image of the sound source so that an orientation in which the sound is radiated is matched with the first direction.