Sound Source Separation for Information Presentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In existing information presentation systems, it is difficult for users to distinguish and understand the content from multiple sound sources near a robot device, as stereo microphones record superimposed audio signals, making it challenging to separate and identify individual audio information.
Innovation Solution
An information presentation device that includes an audio signal input unit, image signal input unit, sound source localization unit, sound source separation unit, and sound source selection unit, which estimates direction information for each sound source, separates audio signals, and selects sound sources based on user input, allowing for easier identification of utterance content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a stereo microphone records audio information from multiple sound sources, then the system can capture audio from various directions, but the audio signals of each sound source become superimposed making it difficult to distinguish and understand individual utterance content
Solution Approach 1:
The patent applies segmentation by dividing the mixed audio signal into separate sound source components. The sound source separation unit segments the superimposed audio information into individual sound source signals, allowing each speaker to be identified and processed separately despite being recorded simultaneously in a multi-directional environment.
Solution Approach 2:
The patent introduces an intermediary processing system consisting of the sound source localization unit and sound source separation unit. These units act as mediators between the stereo microphone and the user, transforming the superimposed audio signals into separated, identifiable sound source information through signal processing algorithms.
2Ease of operation
If the system displays image information corresponding to sound source directions, then users can visually identify speakers, but the complexity of processing and coordinating audio-visual information increases
Solution Approach 1:
The patent merges audio and visual processing functions into a coordinated system. The sound source localization unit processes audio directional information while the image display unit simultaneously displays corresponding visual information, combining both modalities to enhance speaker identification and utterance understanding through multi-sensory integration.
Solution Approach 2:
The patent implements a multi-functional system where the same processing infrastructure serves multiple purposes: sound source separation for audio processing, direction estimation for spatial localization, and image display for visual feedback. This universal approach handles various tasks (speaker identification, utterance recognition, visual attention guidance) through a unified system architecture.
Data Source
AI summary
An information presentation device includes an audio signal input unit configured to input an audio signal, an image signal input unit configured to input an image signal, an image display unit configured to display an image indicated by the image signal, a sound source localization unit configured to estimate direction information for each sound source based on the audio signal, a sound source separation unit configured to separate the audio signal to sound-source-classified audio signals for each sound source, an operation input unit configured to receive an operation input and generates coordinate designation information indicating a part of a region of the image, and a sound source selection unit configured to select a sound-source-classified audio signal of a sound source associated with a coordinate which is included in a region indicated by the coordinate designation information, and which corresponds to the direction information.


