Audio Source Separation Using Angular Location

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio source separation technologies face challenges in identifying and isolating a target signal from a mixture of voices in real-world scenarios, particularly in environments with multiple concurrent speakers, where the target speaker's voice may not be the loudest or most prominent.

Innovation Solution

A deep learning-based system utilizing a neural network that determines the angle of arrival of the target signal from a microphone array, allowing for the separation and enhancement of the desired audio signal by steering a virtual microphone direction towards the selected speaker, while canceling out other sounds through spectral and spatial characteristics analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional audio source separation methods are used, then the system can process mixed audio signals, but it fails to accurately identify and isolate the target speaker's voice when multiple speakers are present simultaneously

Engineering Contradiction:
Improvetarget signal identification accuracyVSAvoidperformance in real-world multi-speaker environments
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces spatial dimension (angle of arrival) as an additional feature dimension for source separation. By incorporating angular information from microphone arrays, the system transforms the problem from purely spectral analysis to spatio-spectral analysis, enabling accurate identification of target speakers in multi-speaker scenarios through directional discrimination

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent employs angle of arrival estimation as an intermediary step between raw microphone signals and final source separation. This intermediate spatial parameter serves as a mediator that guides the separation process by indicating the directional location of the target speaker, thereby improving identification accuracy in complex acoustic environments

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system enhances the target speaker's voice, then communication clarity improves, but the system complexity increases due to the need for angle of arrival determination and spatial filtering

Engineering Contradiction:
Improvecommunication clarityVSAvoidsystem architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the source separation task into distinct functional segments: angle of arrival estimation module, spectral analysis module, and spatial filtering module. This segmentation allows each component to specialize in a specific aspect of the problem, improving overall reliability while making the complex system more manageable and modular

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional mechanical beamforming approaches with deep learning-based spatial filtering. The neural network learns optimal spatial filters directly from data, substituting complex mechanical signal processing with adaptive computational methods that achieve similar or better results with improved flexibility

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240274148A1Sound source separation using angular location
Publication Date: 2024.08.15 INTEL CORP
  • US20240274148A1 patent drawing
  • US20240274148A1 patent drawing
  • US20240274148A1 patent drawing

AI summary

Systems and methods for audio source separation. A deep learning-based system uses an azimuth angle location to separate an audio signal originating from a selected location from other sound. Techniques are disclosed for steering a virtual direction of a microphone towards a selected speaker. A deep-learning based audio regression method, which can be implemented as a neural network, learns to separate out various speakers by leveraging spectral and spatial characteristics of all sources. The neural network can focus on multiple sources in multiple respective target directions, and cancel out other sounds. A user can choose which source to listen to. The network can use the time-domain signal and a frequency-domain signal to separate out the target signal and generate a separated audio output. The direction of the selected speaker relative to the microphone array can be input to the system as a vector.