Speech Recognition Using Depth Information and Lip Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated speech recognition systems fail in noisy environments and are limited by the need for orthogonal frontal images, making them unsuitable for unconstrained and noisy conditions such as industrial or automotive settings.

Innovation Solution

The use of depth information to detect lip movements via lip descriptor points, generating scale, translation, and rotation invariant visual patterns, allowing for speech recognition without audio input and with partial occlusions, using a combination of image receivers, landmark detectors, descriptor calculators, pattern generators, and speech recognizers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional automated speech recognition systems are used, then speech can be recognized in controlled environments, but they fail in noisy environments and require orthogonal frontal images

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidadaptability to noisy and unconstrained environments
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent replaces traditional audio-based speech recognition with a visual-based system using depth cameras and lip descriptor analysis. This substitution allows the system to recognize speech through visual lip movement patterns rather than acoustic signals, enabling reliable operation in noisy environments where audio recognition fails

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms speech recognition from audio domain to visual domain by extracting lip descriptor features from depth images. This parameter change from acoustic frequencies to spatial-temporal lip movement patterns enables the system to overcome noise interference and work in unconstrained environments

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If audio-based speech recognition is used, then speech can be recognized accurately in controlled settings, but it cannot work with partial occlusions or head movements

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidoperation under occlusion and head movement
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent adds the depth dimension by using depth cameras instead of standard 2D images. This third dimension provides robustness to scale, translation, and rotation variations, allowing accurate lip descriptor extraction even when the user moves their head or when parts of the face are occluded

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system uses dynamic lip descriptor points that track and adapt to lip movements in real-time. These descriptors are updated frame-by-frame to follow lip contours during speech, maintaining measurement precision even with head movements and partial occlusions

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If depth information and lip descriptors are used, then speech recognition works in noisy and unconstrained environments, but the system complexity increases

Engineering Contradiction:
Improveadaptability to noisy and unconstrained environmentsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the face into specific regions of interest (lips) and extracts descriptors only from these segmented areas. By focusing computation on lip regions rather than processing the entire face or full audio signal, the system achieves high adaptability while managing computational complexity through selective feature extraction

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10515636B2Speech recognition using depth information
Publication Date: 2019.12.24 HYUNDAI MOTOR CO LTD
  • US10515636B2 patent drawing
  • US10515636B2 patent drawing
  • US10515636B2 patent drawing

AI summary

An example apparatus for detecting speech includes an image receiver to receive depth information corresponding to a face. The apparatus also includes a landmark detector to detect the face comprising lips and track a plurality of descriptor points comprising lip descriptor points located around the lips. The apparatus further includes a descriptor computer to calculate a plurality of descriptor features based on the tracked descriptor points. The apparatus includes a pattern generator to generate a visual pattern of the descriptor features over time. The apparatus also further includes a speech recognition engine to detect speech based on the generated visual pattern.