Deep Neural Net Filter Prediction for Audio Event Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio extraction techniques fail in dynamic environments with distributed microphone setups, where static sensor placement is not possible, leading to ineffective removal of undesired signals from target signals in applications like wearable devices and automotive systems.
Innovation Solution
The use of deep neural networks for filter gain prediction, which involves collecting audio event data, extracting features, labeling them with filter information, and training the network to predict filter gains for improving audio event segmentation, even in noisy conditions, using distributed microphones and normalized spectral band energies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional single microphone stationary noise suppression is used, then the system is simple and low cost, but it cannot effectively remove undesired signals in dynamic environments with distributed microphone setups
Solution Approach 1:
The patent replaces conventional mechanical signal processing methods (stationary noise suppression, beamforming) with a data-driven deep neural network system. The DNN learns optimal filtering strategies from training data and applies them dynamically, substituting complex mechanical sensor arrangements with an intelligent software-based solution that adapts to varying acoustic environments.
Solution Approach 2:
The system dynamically changes filter parameters (filter gains) based on input acoustic conditions. The DNN processes spectral features and generates time-varying filter gains that adapt to different signal-to-noise ratios and acoustic scenarios, allowing the same hardware setup to perform optimally across diverse environments without requiring physical reconfiguration.
2Adaptability or versatility
If distributed microphone arrangements are used, then the system can work in dynamic environments, but static sensor placement assumptions are violated and signal extraction becomes difficult
Solution Approach 1:
The system performs preliminary training offline using diverse acoustic data collected from distributed microphone arrangements. During this preliminary phase, the DNN learns the relationships between spectral features and optimal filter gains for various acoustic scenarios. This pre-learning enables the system to accurately extract signals in dynamic environments during operation without requiring real-time adaptation or precise sensor placement.
3Measurement precision
If iterative filter gain prediction is performed, then segmentation accuracy improves, but computational time increases
Solution Approach 1:
The system employs iterative filter gain prediction where the DNN processes the input spectrogram, applies filter gains, and reprocesses the enhanced output to check for further improvements. This feedback loop continues until convergence or a maximum iteration limit is reached, ensuring high segmentation accuracy while preventing excessive computational overhead through the convergence criterion.
Data Source
AI summary
Disclosed is a feature extraction and classification methodology wherein audio data is gathered in a target environment under varying conditions. From this collected data, corresponding features are extracted, labeled with appropriate filters (e.g., audio event descriptions), and used for training deep neural networks (DNNs) to extract underlying target audio events from unlabeled training data. Once trained, these DNNs are used to predict underlying events in noisy audio to extract therefrom features that enable the separation of the underlying audio events from the noisy components thereof.


