DNN Reverberation Time Estimation Using Multichannel Spatial Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing reverberation time estimation methods using single feature vectors are not robust and struggle with noise, particularly in real-life environments, leading to degraded accuracy in voice recognition and acoustic signal processing.
Innovation Solution
A method and apparatus using a deep neural network (DNN) for multichannel microphone-based reverberation time estimation, which derives a feature vector including spatial information from voice signals through a multichannel microphone, applying it to model nonlinear distributions and estimate reverberation time, incorporating negative-side variance and cross-correlation functions to enhance accuracy and robustness against noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single feature vector is used for reverberation time estimation, then the method is simple, but the accuracy and robustness degrade in noisy real-life environments
Solution Approach 1:
The patent segments the feature extraction process into multiple independent feature vectors including spectral centroid, spectral rolloff, zero-crossing rate, and spatial information features. Each feature vector captures different acoustic characteristics, and their combination provides comprehensive representation of the acoustic environment, thereby improving estimation accuracy without excessive complexity increase
Solution Approach 2:
The patent combines multiple feature vectors into a composite feature set that integrates spectral features, temporal features, and spatial information. This composite approach is analogous to using composite materials - each feature type contributes unique properties that complement each other, creating a robust estimation system that maintains accuracy across diverse noisy environments
2Device complexity
If polynomial regression is used for reverberation time estimation, then the method is computationally simple, but it cannot model nonlinear distributions and performs poorly in varied environments
Solution Approach 1:
The patent replaces traditional polynomial regression (mechanical/mathematical system) with a deep neural network model. The DNN automatically learns nonlinear mappings between feature vectors and reverberation time through training data, capturing complex environmental variations without requiring explicit mathematical modeling, thereby achieving high adaptability while managing computational complexity through efficient network architecture
Solution Approach 2:
The patent changes the modeling approach from fixed polynomial parameters to learnable DNN parameters that adapt to different environments. The model parameters are trained on diverse acoustic data, enabling the system to automatically adjust to various environmental conditions, noise types, and reverberation characteristics, thus achieving universal applicability across different scenarios
3Device complexity
If spatial information is not incorporated, then the processing is simpler, but the robustness against noise and environmental variability decreases
Solution Approach 1:
The patent adds spatial information as an additional dimension to the feature vectors by incorporating inter-microphone time delay (ITD) and inter-microphone level difference (ILD) features from multichannel microphone arrays. This dimensional expansion provides geometric context about sound source location and room geometry, enabling the system to distinguish between direct and reverberant paths more effectively, thereby improving noise robustness while maintaining manageable processing complexity through efficient spatial feature computation
Data Source
AI summary
A multichannel microphone-based reverberation time estimation method and device which use a deep neural network (DNN) are disclosed. A multichannel microphone-based reverberation time estimation method using a DNN, according to one embodiment, comprises the steps of: receiving a voice signal through a multichannel microphone; deriving a feature vector including spatial information by using the inputted voice signal; and estimating the degree of reverberation by applying the feature vector to the DNN.


