Multimodal Lip Shape Detection for Motion-Robust Utterance Sections
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting utterance sections in sound data suffer from reduced precision when the utterer is in motion, leading to erroneous detection due to low extraction accuracy of lip regions from camera images.
Innovation Solution
A device and method that utilize both sound and image data to estimate lip shapes, using a first lip shape estimation module based on sound data and a second lip shape estimation module based on image data, to detect utterance sections by analyzing changes in both lip shapes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If only image data from a camera is used to detect utterance sections, then the device complexity is low, but the measurement precision of utterance section detection drops when the utterer is in motion
Solution Approach 1:
The patent combines image data from a camera and sound data from a microphone into a integrated detection system. The utterance section detection section simultaneously processes both image data (for lip shape extraction) and sound data (for vocalization detection) to determine utterance sections, thereby improving detection precision while maintaining reasonable system complexity through unified processing.
Solution Approach 2:
The patent introduces an utterance section detection section as an intermediary that correlates lip shape changes from image data with vocalization information from sound data. This intermediary component resolves the contradiction by using the relationship between lip movement and sound production to achieve high-precision utterance detection without requiring overly complex individual processing modules.
2Adaptability or versatility
If the utterer is in motion (walking or moving head), then the adaptability of the detection system to real-world conditions improves, but the manufacturing precision of lip region extraction from image data deteriorates
Solution Approach 1:
The patent uses feedback by correlating lip shape changes detected from image data with actual vocalization information from sound data. The utterance section detection section continuously adjusts lip region extraction by comparing expected lip movements (from sound data) with actual extracted lip shapes (from image data), thereby maintaining extraction precision even when the utterer is in motion.
Solution Approach 2:
The patent adds the sound data dimension to the traditional image-only approach. By incorporating audio information about vocalization timing and characteristics, the system creates a multi-dimensional detection framework that compensates for motion-induced degradation in image-based lip extraction precision.
3Reliability
If only sound data is used to detect utterance sections, then the device complexity is low, but the reliability of detection deteriorates due to background noise and non-vocal sounds
Solution Approach 1:
The patent merges sound data processing with image data processing in a unified detection framework. The utterance section detection section combines vocalization detection from sound data with lip shape change detection from image data, where the two data sources mutually validate each other to improve reliability while avoiding the need for separate independent detection systems.
Solution Approach 2:
The patent implements feedback mechanisms where sound data provides expected utterance timing and the image data validates actual lip movements, and vice versa. This cross-validation feedback loop significantly improves detection reliability by filtering out false positives from background noise or non-vocal sounds, without requiring overly complex verification systems.
Data Source
AI summary
Provided is an utterance section detection device including: a first lip shape estimation module configured to estimate a first lip shape of an utterer, based on sound data including a voice of the utterer, a second lip shape estimation module configured to estimate a second lip shape of the utterer; based on image data in which an image of at least a face of the utterer is photographed; and an utterance section detection module configured to detect an utterance section in which the utterer is vocalizing in the sound data, based on changes in the first lip shape and changes in the second lip shape.


