Deep Neural Net Filter Prediction for Audio Event Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio extraction techniques fail in dynamic environments with distributed microphone setups, where static sensor placement is not possible, leading to ineffective removal of undesired signals from target signals in applications like wearable devices and automotive systems.

Innovation Solution

The use of deep neural networks for filter gain prediction, which involves collecting audio event data, extracting features, labeling them with filter information, and training the network to predict filter gains for improving audio event segmentation, even in noisy conditions, using distributed microphones and normalized spectral band energies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional single microphone stationary noise suppression is used, then the system is simple and low cost, but it cannot effectively remove undesired signals in dynamic environments with distributed microphone setups

Engineering Contradiction:
Improveaudio event extraction accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces conventional mechanical signal processing methods (stationary noise suppression, beamforming) with a data-driven deep neural network system. The DNN learns optimal filtering strategies from training data and applies them dynamically, substituting complex mechanical sensor arrangements with an intelligent software-based solution that adapts to varying acoustic environments.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system dynamically changes filter parameters (filter gains) based on input acoustic conditions. The DNN processes spectral features and generates time-varying filter gains that adapt to different signal-to-noise ratios and acoustic scenarios, allowing the same hardware setup to perform optimally across diverse environments without requiring physical reconfiguration.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If distributed microphone arrangements are used, then the system can work in dynamic environments, but static sensor placement assumptions are violated and signal extraction becomes difficult

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidsignal extraction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary training offline using diverse acoustic data collected from distributed microphone arrangements. During this preliminary phase, the DNN learns the relationships between spectral features and optimal filter gains for various acoustic scenarios. This pre-learning enables the system to accurately extract signals in dynamic environments during operation without requiring real-time adaptation or precise sensor placement.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If iterative filter gain prediction is performed, then segmentation accuracy improves, but computational time increases

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system employs iterative filter gain prediction where the DNN processes the input spectrogram, applies filter gains, and reprocesses the enhanced output to check for further improvements. This feedback loop continues until convergence or a maximum iteration limit is reached, ensuring high segmentation accuracy while preventing excessive computational overhead through the convergence criterion.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9666183B2Deep neural net based filter prediction for audio event classification and extraction
Publication Date: 2017.05.30 QUALCOMM INC
  • US9666183B2 patent drawing
  • US9666183B2 patent drawing
  • US9666183B2 patent drawing

AI summary

Disclosed is a feature extraction and classification methodology wherein audio data is gathered in a target environment under varying conditions. From this collected data, corresponding features are extracted, labeled with appropriate filters (e.g., audio event descriptions), and used for training deep neural networks (DNNs) to extract underlying target audio events from unlabeled training data. Once trained, these DNNs are used to predict underlying events in noisy audio to extract therefrom features that enable the separation of the underlying audio events from the noisy components thereof.