Multimodal Lip Shape Detection for Motion-Robust Utterance Sections

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting utterance sections in sound data suffer from reduced precision when the utterer is in motion, leading to erroneous detection due to low extraction accuracy of lip regions from camera images.

Innovation Solution

A device and method that utilize both sound and image data to estimate lip shapes, using a first lip shape estimation module based on sound data and a second lip shape estimation module based on image data, to detect utterance sections by analyzing changes in both lip shapes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If only image data from a camera is used to detect utterance sections, then the device complexity is low, but the measurement precision of utterance section detection drops when the utterer is in motion

Engineering Contradiction:
Improveutterance section detection precisionVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines image data from a camera and sound data from a microphone into a integrated detection system. The utterance section detection section simultaneously processes both image data (for lip shape extraction) and sound data (for vocalization detection) to determine utterance sections, thereby improving detection precision while maintaining reasonable system complexity through unified processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an utterance section detection section as an intermediary that correlates lip shape changes from image data with vocalization information from sound data. This intermediary component resolves the contradiction by using the relationship between lip movement and sound production to achieve high-precision utterance detection without requiring overly complex individual processing modules.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the utterer is in motion (walking or moving head), then the adaptability of the detection system to real-world conditions improves, but the manufacturing precision of lip region extraction from image data deteriorates

Engineering Contradiction:
Improvedetection system adaptability to motionVSAvoidlip region extraction precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent uses feedback by correlating lip shape changes detected from image data with actual vocalization information from sound data. The utterance section detection section continuously adjusts lip region extraction by comparing expected lip movements (from sound data) with actual extracted lip shapes (from image data), thereby maintaining extraction precision even when the utterer is in motion.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent adds the sound data dimension to the traditional image-only approach. By incorporating audio information about vocalization timing and characteristics, the system creates a multi-dimensional detection framework that compensates for motion-induced degradation in image-based lip extraction precision.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If only sound data is used to detect utterance sections, then the device complexity is low, but the reliability of detection deteriorates due to background noise and non-vocal sounds

Engineering Contradiction:
Improveutterance section detection reliabilityVSAvoiddetection system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges sound data processing with image data processing in a unified detection framework. The utterance section detection section combines vocalization detection from sound data with lip shape change detection from image data, where the two data sources mutually validate each other to improve reliability while avoiding the need for separate independent detection systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback mechanisms where sound data provides expected utterance timing and the image data validates actual lip movements, and vice versa. This cross-validation feedback loop significantly improves detection reliability by filtering out false positives from background noise or non-vocal sounds, without requiring overly complex verification systems.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12387728B2Utterance section detection device, utterance section detection method, and storage medium
Publication Date: 2025.08.12 PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
  • US12387728B2 patent drawing
  • US12387728B2 patent drawing
  • US12387728B2 patent drawing

AI summary

Provided is an utterance section detection device including: a first lip shape estimation module configured to estimate a first lip shape of an utterer, based on sound data including a voice of the utterer, a second lip shape estimation module configured to estimate a second lip shape of the utterer; based on image data in which an image of at least a face of the utterer is photographed; and an utterance section detection module configured to detect an utterance section in which the utterer is vocalizing in the sound data, based on changes in the first lip shape and changes in the second lip shape.