Lip Movement Analysis for Voice Section Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice section detection methods face challenges in high-noise environments like vehicle-driving situations, where illumination changes and noise interfere with accurate detection using audio or image signals.
Innovation Solution
A method and device that detect voice sections based on the characteristics and movement features of the lip area, using image analysis to identify lip movement and shape differences, robust against environmental changes like illumination variations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional voice section detection methods using audio signals are used, then the detection process is simple, but the detection accuracy deteriorates in high-noise environments such as vehicle-driving situations
Solution Approach 1:
The patent introduces image signals as an intermediary to detect lip movement characteristics, which serve as a visual mediator to identify voice sections. This intermediary approach allows the system to bypass audio noise by using visual information from the lip area to determine when a person is speaking, thereby resolving the contradiction between simple detection and accuracy in noisy environments.
2Reliability
If conventional image-based voice section detection methods are used, then the detection can be performed visually, but the detection accuracy deteriorates due to continuous illumination changes in vehicle environments
Solution Approach 1:
The patent changes the detection parameter from absolute pixel values (which are sensitive to illumination changes) to motion information derived from differences between consecutive image frames. By using differential parameters that cancel out common illumination variations, the system maintains detection accuracy despite continuous illumination changes in vehicle environments.
3Productivity
If voice section detection is performed to improve voice recognition efficiency, then the recognition time is reduced, but the detection reliability deteriorates when using only audio signals in noisy environments
Solution Approach 1:
The patent creates a multi-functional detection system that processes both audio signals and image signals to determine voice sections. The image processing component serves as a universal indicator of speech activity that works reliably across different acoustic environments, thereby improving detection reliability without compromising voice recognition efficiency.
Data Source
AI summary
Provided is a method of detecting a voice section, including detecting from at least one image an area where lips exist, obtaining a feature value of movement of the lips in the detected area based on a difference between pixel values of pixels included in the detected area, and detecting the voice section from the at least one image based on the feature value.


