Neural Network Impulsive Noise Detection in Voice Applications
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional voice activity detection systems often misclassify impulsive noise as speech due to similarities in signal envelopes and spectra, leading to performance deterioration in applications like beamsteering, where incorrect detection can result in directional errors.
Innovation Solution
An integrated circuit with a processor implementing an impulsive noise detector using a neural network to differentiate between speech and noise events by analyzing feature vectors derived from characteristics such as sudden onset, harmonicity, and spectral flatness measures, combined with data augmentation for robust training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If traditional VAD relies on changes in signal level on full-band or sub-band basis, then the detection method is simple and computationally efficient, but impulsive noise is misdetected as speech leading to false detection
Solution Approach 1:
The patent segments the audio signal into multiple frequency sub-bands and extracts different features (spectral flatness, harmonicity, zero-crossing rate) for each sub-band. This segmentation allows the system to analyze specific characteristics of impulsive noise versus speech in different frequency regions, improving detection accuracy without requiring a completely complex system architecture.
Solution Approach 2:
The patent transitions from traditional single-dimensional signal level analysis to multi-dimensional feature space analysis by incorporating spectral flatness, harmonicity, zero-crossing rate, and sub-band energy ratios. This dimensional expansion enables the system to distinguish between impulsive noise and speech by evaluating multiple characteristics simultaneously, resolving the false detection problem.
2Use of energy by moving object
If traditional VAD uses signal envelope analysis, then the processing is computationally efficient, but the system cannot distinguish between impulsive noise and speech events
Solution Approach 1:
The patent performs preliminary analysis by computing spectral flatness, harmonicity, and zero-crossing rate features before making the final speech/noise determination. These pre-computed features capture essential characteristics of the signal that differentiate speech from impulsive noise, allowing the system to make reliable decisions with moderate computational effort.
Solution Approach 2:
The patent introduces intermediate feature representations (spectral flatness measure, harmonicity measure, zero-crossing rate) that serve as mediators between the raw audio signal and the final detection decision. These intermediate features transform the complex discrimination problem into a series of simpler comparisons, maintaining computational efficiency while improving reliability.
3Device complexity
If impulsive noise spectrum is averaged over various occurrences and compared with averaged speech spectrum, then the comparison is simplified, but the spectra are not significantly different leading to misclassification
Solution Approach 1:
The patent applies local quality analysis by evaluating spectral characteristics in specific frequency sub-bands rather than analyzing the entire spectrum uniformly. By computing spectral flatness and energy ratios in different sub-bands, the system identifies localized differences between speech and impulsive noise spectra that are not apparent in overall averaged spectra, improving differentiation accuracy.
Solution Approach 2:
The patent changes the parameters used for spectral comparison from simple averaged power spectra to multiple derived parameters including spectral flatness, harmonicity, and zero-crossing rate. These parameter transformations highlight distinctive features of speech (periodicity, harmonics) versus impulsive noise (randomness, flat spectrum), enabling accurate classification despite spectral similarities.
Data Source
AI summary
In accordance with embodiments of the present disclosure, an integrated circuit for implementing at least a portion of an audio device may include an audio output configured to reproduce audio information by generating an audio output signal for communication to at least one transducer of the audio device, a microphone input configured to receive an input signal indicative of ambient sound external to the audio device, and a processor configured to implement an impulsive noise detector. The impulsive noise detector may comprise a plurality of processing blocks for determining a feature vector based on characteristics of the input signal and a neural network for determining based on the feature vector whether the impulsive event comprises a speech event or a noise event.


