Acoustic and Lip Signal Fusion for Voice Activity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice activity detection systems face challenges in maintaining accuracy under various noise environments, particularly in complex and dynamic noise conditions encountered with the increasing use of voice recognition technologies.
Innovation Solution
A voice activity detection apparatus that utilizes both acoustic and non-acoustic signals, such as lip image signals, to enhance detection accuracy by calculating integrated features and voice existence scores through a series of neural networks trained to reduce losses in noise environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If only acoustic signals are used for voice activity detection, then the system complexity is low, but detection accuracy deteriorates in noise environments
Solution Approach 1:
The patent combines acoustic signals and non-acoustic signals (such as lip images) into a unified multimodal input system. The voice activity detection apparatus integrates features from both acoustic and non-acoustic domains, processing them through shared and modality-specific neural network layers to achieve improved detection accuracy in noise environments while managing system complexity through efficient feature fusion mechanisms.
Solution Approach 2:
The patent introduces non-acoustic signal dimensions (visual/lip image data) to complement the traditional acoustic signal dimension. By adding this another dimension of information, the system overcomes the limitations of acoustic-only detection in noisy conditions, as lip movements provide complementary evidence for voice activity that is independent of acoustic interference.
2Reliability
If multimodal signals are integrated for voice detection, then detection accuracy in noise environments improves, but processing complexity increases
Solution Approach 1:
The patent segments the processing architecture into distinct components: acoustic signal processing pathways, non-acoustic signal processing pathways, and feature fusion layers. This segmentation allows each modality to be processed independently through specialized networks before integration, reducing overall processing complexity while maintaining the reliability benefits of multimodal integration.
Solution Approach 2:
The patent introduces an intermediate feature representation layer that acts as a mediator between acoustic and non-acoustic inputs. This intermediate layer processes and aligns features from different modalities before final integration, simplifying the complexity of directly combining raw multimodal signals while preserving detection reliability through structured feature fusion.
3Measurement precision
If deep neural networks are used to process integrated features, then voice existence score accuracy improves, but computational cost increases
Solution Approach 1:
The patent performs preliminary processing of acoustic and non-acoustic signals through dedicated feature extraction networks before integration. By pre-processing each modality separately and extracting relevant features in advance, the system reduces the computational burden on subsequent deep neural network layers that determine voice existence scores, thereby lowering overall computational energy requirements while maintaining accuracy.
Data Source
AI summary
According to one embodiment, a voice activity detection apparatus includes a processing circuit. The processing circuit acquires an acoustic signal and a non-acoustic signal, calculates an acoustic feature based on the acoustic signal, calculates a non-acoustic feature based on the non-acoustic signal, calculates a voice emphasized feature based on the acoustic signal and the non-acoustic signal, calculates a voice existence/non-existence feature on the basis of the acoustic feature and the non-acoustic feature, calculates a voice existence score based on the voice emphasized feature and the voice existence/non-existence feature, detects a voice section and/or a non-voice section based on comparison of the voice existence score with a threshold.


