Deep Neural Network Audio Spectral Mask for Voice Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing technologies fail to effectively isolate and identify voice data from ambient sound in crowded environments, making it difficult for individuals to locate specific speakers in noisy settings.

Innovation Solution

A method and system utilizing a deep neural network with a predictive audio spectral mask, combined with a multi-microphone device, to separate target speech from ambient noise by determining the location of the speaker through amplitude and phase data processing, displayed on a user device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional audio signal processing methods are used, then device complexity is low, but the ability to isolate and identify voice data from ambient sound in crowded environments is insufficient

Engineering Contradiction:
Improvevoice isolation accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a deep neural network as an intermediary computational layer between the raw audio signals from multiple microphones and the final voice isolation output. This neural network acts as a mediator that learns complex mappings from training data, enabling effective voice separation in crowded environments without requiring complex manual signal processing algorithms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary training of the deep neural network offline using large datasets of audio recordings. This preliminary action pre-learns the complex patterns of voice isolation, so that during actual deployment, the network can quickly process new audio inputs without requiring complex real-time computations. The spectral mask generation is also prepared in advance through the training process.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If deep neural networks with predictive audio spectral masks are deployed, then target speech can be effectively separated from crowd noise, but computational resources and processing time increase

Engineering Contradiction:
Improvespeech separation reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The deep neural network is trained offline in advance using large datasets, performing the computationally intensive learning process before deployment. This preliminary training action stores the learned knowledge in the network weights, allowing fast inference during actual speech separation tasks without requiring real-time complex computations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The audio processing is segmented into distinct stages: the deep neural network processes spectral information to generate separation masks, while a separate beamforming module uses amplitude and phase data for spatial filtering. This segmentation allows each component to be optimized independently and processed in parallel, reducing overall processing time.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If multi-microphone devices with deep neural networks are used, then location of origin can be determined accurately, but device complexity and cost increase

Engineering Contradiction:
Improvelocation determination accuracyVSAvoidmulti-microphone system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from analyzing only the temporal amplitude spectrum to incorporating the spatial dimension by using phase information from multiple microphones. By adding this spatial dimension through phase-based beamforming, the system can determine the location of origin of target speech, providing directional awareness in addition to voice isolation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The multi-microphone device with integrated deep neural network performs multiple functions simultaneously: it isolates target speech from crowd noise, determines the location of origin through phase analysis, and can potentially identify speakers. This multi-functionality is achieved through a unified system architecture that processes both spectral and spatial information.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11514892B2Audio-spectral-masking-deep-neural-network crowd search
Publication Date: 2022.11.29 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11514892B2 patent drawing
  • US11514892B2 patent drawing
  • US11514892B2 patent drawing

AI summary

A system includes a memory having instructions therein and at least one processor in communication with the memory. The at least one processor is configured to execute the instructions to communicate, into a user device, a deep neural network comprising a predictive audio spectral mask. The at least one processor is also configured to execute the instructions to: generate data corresponding to ambient sound via a multi-microphone device; separate amplitude data and/or phase data from the data via the deep neural network comprising the predictive audio spectral mask; and determine, via the user device and based on the amplitude data and/or phase data, a location of origin of target speech relative to the user device. The at least one processor is configured to execute the instructions to display, via the user device, the location of origin of the target speech relative to the user device.