Voice Activity Detection Using Lip Movement Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice activity detection methods in complex noise environments, especially with ambient voices, suffer from poor accuracy due to incorrect detection of voice start and end points, leading to increased interaction delays and recognition errors in human-machine interactions.
Innovation Solution
Combining voice detection models with lip movement detection technology, where voice start and end points are corrected using lip movement detection results, ensuring more accurate timing and reducing noise interference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice detection model is used for voice activity detection, then detection speed is maintained, but detection accuracy deteriorates in complex noise environments
Solution Approach 1:
The patent introduces lip movement detection as an intermediary mechanism to bridge voice detection and verification. The lip movement detection model acts as a mediator that provides independent verification of voice activity, helping to filter out noise interference and improve overall detection accuracy in complex acoustic environments
Solution Approach 2:
The patent combines voice detection model with lip movement detection model into a unified voice activity detection system. By merging audio-based detection with visual-based detection, the system leverages multiple sensing modalities to overcome the limitations of single-modality detection in noisy environments
2Reliability
If voice detection model is used, then interaction speed is maintained, but interaction success rate deteriorates due to incorrect detection
Solution Approach 1:
The patent implements a feedback mechanism where lip movement detection results are used to verify and correct voice detection outcomes. This feedback loop ensures that only accurately detected voice activities trigger interactions, reducing false positives and improving interaction success rates while maintaining timely responses
Solution Approach 2:
The system performs preliminary lip movement detection and verification before finalizing voice activity detection results. By conducting this preliminary check, the system prevents incorrect interactions from being triggered, thereby improving reliability without significantly increasing interaction delays
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
The present disclosure discloses a voice activity detection method and apparatus, an electronic device and a storage medium, and relates to the field of artificial intelligence, such as deep learning, intelligent voices, or the like. The method may include: acquiring time-aligned voice data and video data; performing a first detection of a voice start point and a voice end point of the voice data using a voice detection model obtained by a training operation; performing a second detection of a lip movement start point and a lip movement end point of the video data; and correcting a result of the first detection using a result of the second detection, and taking a corrected result as a voice activity detection result. The solution of the present disclosure may improve accuracy of the voice activity detection result, or the like.