Voice Activity Detection Using PNCC and Chroma Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Voice Activity Detection (VAD) techniques face limitations in real-time processing, memory efficiency, and accuracy in noisy environments, particularly in low power devices, due to high computational complexity and sensitivity to non-stationary noise, reverberation, and varying microphone-speaker distances.

Innovation Solution

The proposed system employs an audio processing device with a local feature extraction component, a global feature extraction component, and a neural network that integrates Power-Normalized Coefficients (PNCC) and chroma/harmonicity features to determine the probability of a target audio signal, using a single-layer feedforward network with a cost function based on Signal to Noise Ratio (SNR), enabling real-time processing with minimal latency and robustness to noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional VAD techniques are used, then voice activity detection can be performed, but processing resources and memory requirements are too high for real-time detection in low power devices

Engineering Contradiction:
Improvereal-time processing capabilityVSAvoidprocessing resources and memory requirements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple frames, and features are extracted and processed in a segmented manner. The neural network processes each frame independently, allowing for real-time detection with reduced computational burden compared to analyzing the entire audio stream at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts specific relevant features (Power-Normalized Coefficients and chroma/harmonicity features) from the audio signal, focusing only on the most important characteristics for voice activity detection. This selective extraction reduces the dimensionality and complexity of the processing required.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If conventional VAD techniques are used, then voice activity detection can be performed, but accuracy decreases in noisy environments with non-stationary noise

Engineering Contradiction:
Improvedetection accuracy in noisy conditionsVSAvoidsensitivity to non-stationary noise
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent uses Power-Normalized Coefficients which normalize the power of spectral components, making the features less sensitive to variations in noise levels. The chroma and harmonicity features also capture intrinsic properties of speech that remain relatively stable under varying noise conditions, improving robustness.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces conventional signal processing methods with a neural network-based approach. The neural network learns optimal decision boundaries from training data, enabling it to distinguish speech from various types of noise more effectively than traditional algorithmic approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10504539B2Voice activity detection systems and methods
Publication Date: 2019.12.10 SYNAPTICS INC
  • US10504539B2 patent drawing
  • US10504539B2 patent drawing
  • US10504539B2 patent drawing

AI summary

An audio processing device or method includes an audio transducer operable to receive audio input and generate an audio signal based on the audio input. The audio processing device or method also includes an audio signal processor operable to extract local features from the audio signal, such as Power-Normalized Coefficients (PNCC) of the audio signal. The audio signal processor also is operable to extract global features from the audio signal, such as chroma features and harmonicity features. A neural network is provided to determine a probability that a target audio is present in the audio signal based on the local and global features. In particular, the neural network is trained to output a value indicating whether the target audio is present and locally dominant in the audio signal.