Lip Movement Analysis for Voice Section Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice section detection methods face challenges in high-noise environments like vehicle-driving situations, where illumination changes and noise interfere with accurate detection using audio or image signals.

Innovation Solution

A method and device that detect voice sections based on the characteristics and movement features of the lip area, using image analysis to identify lip movement and shape differences, robust against environmental changes like illumination variations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional voice section detection methods using audio signals are used, then the detection process is simple, but the detection accuracy deteriorates in high-noise environments such as vehicle-driving situations

Engineering Contradiction:
Improvevoice section detection accuracyVSAvoidnoise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces image signals as an intermediary to detect lip movement characteristics, which serve as a visual mediator to identify voice sections. This intermediary approach allows the system to bypass audio noise by using visual information from the lip area to determine when a person is speaking, thereby resolving the contradiction between simple detection and accuracy in noisy environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If conventional image-based voice section detection methods are used, then the detection can be performed visually, but the detection accuracy deteriorates due to continuous illumination changes in vehicle environments

Engineering Contradiction:
Improvevoice section detection accuracyVSAvoidillumination changes
Core Design Contradiction:
ReliabilityVSIllumination intensity

Solution Approach 1:

The patent changes the detection parameter from absolute pixel values (which are sensitive to illumination changes) to motion information derived from differences between consecutive image frames. By using differential parameters that cancel out common illumination variations, the system maintains detection accuracy despite continuous illumination changes in vehicle environments.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If voice section detection is performed to improve voice recognition efficiency, then the recognition time is reduced, but the detection reliability deteriorates when using only audio signals in noisy environments

Engineering Contradiction:
Improvevoice recognition efficiencyVSAvoiddetection reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent creates a multi-functional detection system that processes both audio signals and image signals to determine voice sections. The image processing component serves as a universal indicator of speech activity that works reliably across different acoustic environments, thereby improving detection reliability without compromising voice recognition efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10923126B2Method and device for detecting voice activity based on image information
Publication Date: 2021.02.16 SAMSUNG ELECTRONICS CO LTD
  • US10923126B2 patent drawing
  • US10923126B2 patent drawing
  • US10923126B2 patent drawing

AI summary

Provided is a method of detecting a voice section, including detecting from at least one image an area where lips exist, obtaining a feature value of movement of the lips in the detected area based on a difference between pixel values of pixels included in the detected area, and detecting the voice section from the at least one image based on the feature value.