Acoustic and Lip Signal Fusion for Voice Activity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice activity detection systems face challenges in maintaining accuracy under various noise environments, particularly in complex and dynamic noise conditions encountered with the increasing use of voice recognition technologies.

Innovation Solution

A voice activity detection apparatus that utilizes both acoustic and non-acoustic signals, such as lip image signals, to enhance detection accuracy by calculating integrated features and voice existence scores through a series of neural networks trained to reduce losses in noise environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If only acoustic signals are used for voice activity detection, then the system complexity is low, but detection accuracy deteriorates in noise environments

Engineering Contradiction:
Improvevoice detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines acoustic signals and non-acoustic signals (such as lip images) into a unified multimodal input system. The voice activity detection apparatus integrates features from both acoustic and non-acoustic domains, processing them through shared and modality-specific neural network layers to achieve improved detection accuracy in noise environments while managing system complexity through efficient feature fusion mechanisms.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces non-acoustic signal dimensions (visual/lip image data) to complement the traditional acoustic signal dimension. By adding this another dimension of information, the system overcomes the limitations of acoustic-only detection in noisy conditions, as lip movements provide complementary evidence for voice activity that is independent of acoustic interference.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If multimodal signals are integrated for voice detection, then detection accuracy in noise environments improves, but processing complexity increases

Engineering Contradiction:
Improvedetection reliabilityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the processing architecture into distinct components: acoustic signal processing pathways, non-acoustic signal processing pathways, and feature fusion layers. This segmentation allows each modality to be processed independently through specialized networks before integration, reducing overall processing complexity while maintaining the reliability benefits of multimodal integration.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediate feature representation layer that acts as a mediator between acoustic and non-acoustic inputs. This intermediate layer processes and aligns features from different modalities before final integration, simplifying the complexity of directly combining raw multimodal signals while preserving detection reliability through structured feature fusion.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If deep neural networks are used to process integrated features, then voice existence score accuracy improves, but computational cost increases

Engineering Contradiction:
Improvevoice existence score accuracyVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary processing of acoustic and non-acoustic signals through dedicated feature extraction networks before integration. By pre-processing each modality separately and extracting relevant features in advance, the system reduces the computational burden on subsequent deep neural network layers that determine voice existence scores, thereby lowering overall computational energy requirements while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12387747B2Voice activity detection apparatus, learning apparatus, and storage medium
Publication Date: 2025.08.12 KK TOSHIBA
  • US12387747B2 patent drawing
  • US12387747B2 patent drawing
  • US12387747B2 patent drawing

AI summary

According to one embodiment, a voice activity detection apparatus includes a processing circuit. The processing circuit acquires an acoustic signal and a non-acoustic signal, calculates an acoustic feature based on the acoustic signal, calculates a non-acoustic feature based on the non-acoustic signal, calculates a voice emphasized feature based on the acoustic signal and the non-acoustic signal, calculates a voice existence/non-existence feature on the basis of the acoustic feature and the non-acoustic feature, calculates a voice existence score based on the voice emphasized feature and the voice existence/non-existence feature, detects a voice section and/or a non-voice section based on comparison of the voice existence score with a threshold.