Sound Source Separation for Information Presentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In existing information presentation systems, it is difficult for users to distinguish and understand the content from multiple sound sources near a robot device, as stereo microphones record superimposed audio signals, making it challenging to separate and identify individual audio information.

Innovation Solution

An information presentation device that includes an audio signal input unit, image signal input unit, sound source localization unit, sound source separation unit, and sound source selection unit, which estimates direction information for each sound source, separates audio signals, and selects sound sources based on user input, allowing for easier identification of utterance content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a stereo microphone records audio information from multiple sound sources, then the system can capture audio from various directions, but the audio signals of each sound source become superimposed making it difficult to distinguish and understand individual utterance content

Engineering Contradiction:
Improveaudio recording capabilityVSAvoidaudio signal separation
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent applies segmentation by dividing the mixed audio signal into separate sound source components. The sound source separation unit segments the superimposed audio information into individual sound source signals, allowing each speaker to be identified and processed separately despite being recorded simultaneously in a multi-directional environment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing system consisting of the sound source localization unit and sound source separation unit. These units act as mediators between the stereo microphone and the user, transforming the superimposed audio signals into separated, identifiable sound source information through signal processing algorithms.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the system displays image information corresponding to sound source directions, then users can visually identify speakers, but the complexity of processing and coordinating audio-visual information increases

Engineering Contradiction:
Improveutterance content understandingVSAvoidaudio-visual processing system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent merges audio and visual processing functions into a coordinated system. The sound source localization unit processes audio directional information while the image display unit simultaneously displays corresponding visual information, combining both modalities to enhance speaker identification and utterance understanding through multi-sensory integration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements a multi-functional system where the same processing infrastructure serves multiple purposes: sound source separation for audio processing, direction estimation for spatial localization, and image display for visual feedback. This universal approach handles various tasks (speaker identification, utterance recognition, visual attention guidance) through a unified system architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8990078B2Information presentation device associated with sound source separation
Publication Date: 2015.03.24 HONDA MOTOR CO LTD
  • US8990078B2 patent drawing
  • US8990078B2 patent drawing
  • US8990078B2 patent drawing

AI summary

An information presentation device includes an audio signal input unit configured to input an audio signal, an image signal input unit configured to input an image signal, an image display unit configured to display an image indicated by the image signal, a sound source localization unit configured to estimate direction information for each sound source based on the audio signal, a sound source separation unit configured to separate the audio signal to sound-source-classified audio signals for each sound source, an operation input unit configured to receive an operation input and generates coordinate designation information indicating a part of a region of the image, and a sound source selection unit configured to select a sound-source-classified audio signal of a sound source associated with a coordinate which is included in a region indicated by the coordinate designation information, and which corresponds to the direction information.