Voice Activity Detection Using Lip Movement Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice activity detection methods in complex noise environments, especially with ambient voices, suffer from poor accuracy due to incorrect detection of voice start and end points, leading to increased interaction delays and recognition errors in human-machine interactions.

Innovation Solution

Combining voice detection models with lip movement detection technology, where voice start and end points are corrected using lip movement detection results, ensuring more accurate timing and reducing noise interference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice detection model is used for voice activity detection, then detection speed is maintained, but detection accuracy deteriorates in complex noise environments

Engineering Contradiction:
Improvevoice detection accuracyVSAvoidnoise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces lip movement detection as an intermediary mechanism to bridge voice detection and verification. The lip movement detection model acts as a mediator that provides independent verification of voice activity, helping to filter out noise interference and improve overall detection accuracy in complex acoustic environments

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent combines voice detection model with lip movement detection model into a unified voice activity detection system. By merging audio-based detection with visual-based detection, the system leverages multiple sensing modalities to overcome the limitations of single-modality detection in noisy environments

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If voice detection model is used, then interaction speed is maintained, but interaction success rate deteriorates due to incorrect detection

Engineering Contradiction:
Improveinteraction success rateVSAvoidinteraction delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a feedback mechanism where lip movement detection results are used to verify and correct voice detection outcomes. This feedback loop ensures that only accurately detected voice activities trigger interactions, reducing false positives and improving interaction success rates while maintaining timely responses

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary lip movement detection and verification before finalizing voice activity detection results. By conducting this preliminary check, the system prevents incorrect interactions from being triggered, thereby improving reliability without significantly increasing interaction delays

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4086905B1Voice activity detection method and apparatus, electronic device and storage medium
Publication Date: 2023.12.13 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP4086905B1 patent drawingFigure 1
  • EP4086905B1 patent drawingFigure 2~3
  • EP4086905B1 patent drawingFigure 4~5

AI summary

The present disclosure discloses a voice activity detection method and apparatus, an electronic device and a storage medium, and relates to the field of artificial intelligence, such as deep learning, intelligent voices, or the like. The method may include: acquiring time-aligned voice data and video data; performing a first detection of a voice start point and a voice end point of the voice data using a voice detection model obtained by a training operation; performing a second detection of a lip movement start point and a lip movement end point of the video data; and correcting a result of the first detection using a result of the second detection, and taking a corrected result as a voice activity detection result. The solution of the present disclosure may improve accuracy of the voice activity detection result, or the like.