Voice Activity Detection Using Spatial and Conventional Detectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice activity detection systems are limited in their ability to effectively utilize multiple microphones, leading to suboptimal performance in noisy environments with non-stationary noise, such as public places, where conventional detectors struggle to distinguish speech from background noise.

Innovation Solution

The implementation of a spatial voice activity detection system that uses two microphones to form main and anti-beam signals, calculating power ratios and differences to determine the presence of speech, thereby improving accuracy and noise suppression by considering the directionality and noise content of the audio signal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single microphone is used for voice activity detection, then the device complexity is low, but the measurement precision of speech in noisy environments deteriorates

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidmicrophone array complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the voice activity detection function into two separate detectors: a spatial VAD that processes multi-microphone signals to detect speech presence, and a conventional VAD that processes single-microphone signals. This segmentation allows each detector to specialize in specific aspects of voice activity detection, improving overall accuracy while managing system complexity through functional division.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a spatial dimension to voice activity detection by using multiple microphones arranged in a specific geometry and processing signals in both time and frequency domains. The spatial VAD utilizes directional information from the microphone array to distinguish speech from noise, adding a spatial dimension that enhances detection accuracy in noisy environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Object-affected harmful factors

If multiple microphones are used to form beam-formed signals, then the noise suppression capability is improved, but the device complexity increases

Engineering Contradiction:
Improvenoise and interferenceVSAvoidsignal processing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent merges the outputs of the spatial VAD and conventional VAD through a classification stage that combines their detection results. The classifier integrates the spatial information from the beam-formed signals and the temporal information from the conventional VAD to make a final voice activity decision, thereby improving noise suppression while managing processing complexity through intelligent signal fusion.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a classification stage as an intermediary between the spatial VAD and conventional VAD. This classifier acts as a mediator that processes the outputs from both detectors, reconciles their results, and produces the final voice activity decision. This intermediary layer simplifies the overall system architecture by providing a clear separation between spatial processing and temporal processing functions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If spatial processing is applied to detect speech direction, then the ability to distinguish speech from noise is improved, but the computational requirements increase

Engineering Contradiction:
Improvespeech vs noise distinctionVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies spatial processing selectively rather than continuously. The spatial VAD processes signals only when needed for voice activity detection decisions, and the beam-forming operations are performed at specific frequency bins where speech is likely to be present. This partial application of spatial processing reduces computational energy consumption while maintaining reliable speech vs. noise distinction.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent applies different processing strategies to different parts of the signal spectrum. The spatial VAD focuses computational resources on frequency regions where speech is most likely to occur, while using simpler processing for other frequency regions. This local quality approach optimizes the balance between speech detection reliability and computational energy consumption by allocating processing power where it is most needed.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3392668B1Method and apparatus for voice activity determination
Publication Date: 2023.04.12 NOKIA TECHNOLOGIES OY
  • EP3392668B1 patent drawingFigure 1
  • EP3392668B1 patent drawingFigure 2~5
  • EP3392668B1 patent drawingFigure 4a~4b

AI summary

In accordance with an example embodiment of the invention, there is provided an apparatus for detecting voice activity in an audio signal. The apparatus comprises a first voice activity detector (6b) for making a first voice activity detection decision (D2) based at least in part on the voice activity of a first audio signal (A1) received from a first microphone (1a). The apparatus also comprises a second voice activity detector (6a) for making a second voice activity detection decision (D1) based at least in part on an estimate of a direction of the first audio signal (A1) and an estimate of a direction of a second audio signal (A2) received from a second microphone (1b). The apparatus further comprises a classifier (6c) for making a third voice activity detection decision (D3) based at least in part on the first and second voice activity detection decisions.