Voice Activity Detection Using PNCC and Chroma Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Voice Activity Detection (VAD) techniques face limitations in real-time processing, memory efficiency, and accuracy in noisy environments, particularly in low power devices, due to high computational complexity and sensitivity to non-stationary noise, reverberation, and varying microphone-speaker distances.
Innovation Solution
The proposed system employs an audio processing device with a local feature extraction component, a global feature extraction component, and a neural network that integrates Power-Normalized Coefficients (PNCC) and chroma/harmonicity features to determine the probability of a target audio signal, using a single-layer feedforward network with a cost function based on Signal to Noise Ratio (SNR), enabling real-time processing with minimal latency and robustness to noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional VAD techniques are used, then voice activity detection can be performed, but processing resources and memory requirements are too high for real-time detection in low power devices
Solution Approach 1:
The audio signal is divided into multiple frames, and features are extracted and processed in a segmented manner. The neural network processes each frame independently, allowing for real-time detection with reduced computational burden compared to analyzing the entire audio stream at once.
Solution Approach 2:
The patent extracts specific relevant features (Power-Normalized Coefficients and chroma/harmonicity features) from the audio signal, focusing only on the most important characteristics for voice activity detection. This selective extraction reduces the dimensionality and complexity of the processing required.
2Measurement precision
If conventional VAD techniques are used, then voice activity detection can be performed, but accuracy decreases in noisy environments with non-stationary noise
Solution Approach 1:
The patent uses Power-Normalized Coefficients which normalize the power of spectral components, making the features less sensitive to variations in noise levels. The chroma and harmonicity features also capture intrinsic properties of speech that remain relatively stable under varying noise conditions, improving robustness.
Solution Approach 2:
The patent replaces conventional signal processing methods with a neural network-based approach. The neural network learns optimal decision boundaries from training data, enabling it to distinguish speech from various types of noise more effectively than traditional algorithmic approaches.
Data Source
AI summary
An audio processing device or method includes an audio transducer operable to receive audio input and generate an audio signal based on the audio input. The audio processing device or method also includes an audio signal processor operable to extract local features from the audio signal, such as Power-Normalized Coefficients (PNCC) of the audio signal. The audio signal processor also is operable to extract global features from the audio signal, such as chroma features and harmonicity features. A neural network is provided to determine a probability that a target audio is present in the audio signal based on the local and global features. In particular, the neural network is trained to output a value indicating whether the target audio is present and locally dominant in the audio signal.


