Wearable Audio Source Separation and Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality solutions do not effectively enhance audio scenes by separating and classifying audio sources within a real-world environment, limiting the ability to provide relevant additional information to users.

Innovation Solution

A method and apparatus that capture audio signals, process them through filtering and beamforming, separate audio sources, classify them, and present additional information related to the classified sources, using a wearable device with microphones, cameras, and sensors to enhance the audio experience by localizing and tracking audio sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If audio sources are separated and classified to provide additional information, then the usefulness of audio scene augmentation is improved, but the device complexity increases due to multiple microphones and processing modules

Engineering Contradiction:
Improveinformation completenessVSAvoiddevice complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the audio scene into multiple independent audio sources through source separation processing. The processor divides the mixed audio signal from multiple microphones into distinct audio sources, classifies them, and processes them independently. This allows selective enhancement of specific sources while maintaining overall system functionality, resolving the contradiction between information completeness and device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The wearable device integrates multiple functions including audio capture, source separation, classification, and information presentation in a single system. The processor performs multiple operations (filtering, beamforming, source separation, classification) and the device can present information through multiple channels (visual display, audio output), making the complex device versatile and justifying its complexity through multi-functionality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple processing steps (filtering, beamforming, source separation) are applied, then the measurement precision of audio sources is improved, but the processing time increases

Engineering Contradiction:
Improvesource localization precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary processing steps (filtering and beamforming) to audio signals before source separation and classification. By pre-processing the signals to enhance specific characteristics and reduce noise, the subsequent source separation and classification operations can be performed more efficiently with reduced computational complexity, thus maintaining high precision while reducing overall processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The processing pipeline is designed to be dynamic and adaptive. The system can adjust the level of processing applied based on the audio scene characteristics, user preferences, and computational resources available. This allows the system to maintain high measurement precision when needed while reducing processing time in less demanding situations, resolving the contradiction between precision and processing time.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9949056B2Method and apparatus for presenting to a user of a wearable apparatus additional information related to an audio scene
Publication Date: 2018.04.17 ECOLE POLYTECHNIQUE FEDERALE DE LAUSANNE (EPFL)
  • US9949056B2 patent drawing
  • US9949056B2 patent drawing
  • US9949056B2 patent drawing

AI summary

A method for modifying an audio scene and/or presenting additional information relevant to the audio scene includes capturing audio signals from the audio scene with a plurality of microphones; outputting an audio signal with a plurality of acoustical transducers; processing the captured audio signals, where the processing comprises one or more of filtering, equalization, echoes processing, and beamforming; separating and distinguishing audio signal sources using the processed audio signals; selecting at least one separated audio signal source; classifying the at least one selected separated audio signal source; retrieving additional information related to the classified audio signal source; and presenting the additional information in a perceptible form.