Lip-Language Identification in AR Devices via Facial Image Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality systems lack the ability to effectively identify and translate lip language in real-time, limiting user interaction and communication, especially in environments where visual or auditory cues are unclear or unavailable.

Innovation Solution

A lip-language identification method and apparatus that acquire a sequence of face images, perform lip-language identification by determining semantic information corresponding to lip actions, and output this information as text or audio, utilizing existing AR device components without additional hardware, thereby expanding AR device functions and improving user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If lip-language identification is implemented in AR systems, then communication capability is improved, but device complexity increases

Engineering Contradiction:
Improvecommunication capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the communication function by separating lip-language identification from core AR operations. A dedicated module captures facial images, extracts lip movement features, and translates them to text/audio independently, while the main AR system continues handling visual overlay and spatial rendering. This modular segmentation allows communication capability enhancement without significantly increasing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The AR device leverages its existing multi-functional capabilities to support lip-language identification. The camera system, already used for AR scene capture, is also utilized for facial image acquisition. The processing unit, designed for real-time AR rendering, is extended to handle lip movement analysis. This universal use of existing components enables new communication functionality without proportionally increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If real-time lip-language identification is performed, then communication speed is improved, but processing time increases

Engineering Contradiction:
Improvecommunication speedVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously capturing and pre-processing facial images even before translation is requested. Lip movement features are extracted in advance from the video stream, and potential speech segments are pre-identified. When translation is needed, the system only processes the pre-prepared feature data, significantly reducing actual translation latency while maintaining real-time communication capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements selective processing by skipping unnecessary analysis steps. Instead of analyzing every frame in detail, it identifies key frames containing lip movements and processes only those. The processing pipeline rushes through essential feature extraction and translation steps while bypassing redundant operations, achieving fast translation without processing every piece of visual data thoroughly.

Inventive Principle:
Principle #21Skipping (Rushing through)

3Measurement precision

If additional hardware is added for lip-language identification, then identification accuracy is improved, but device portability deteriorates

Engineering Contradiction:
Improveidentification accuracyVSAvoiddevice portability
Core Design Contradiction:
Measurement precisionVSWeight of moving object

Solution Approach 1:

The system achieves accurate lip-language identification without additional hardware by making the existing AR device components perform multiple functions. The camera, already present for AR scene capture, is used for facial image acquisition. The processing unit, designed for real-time AR rendering, is extended to handle lip movement analysis. This universal use of existing components eliminates the need for dedicated hardware while maintaining identification accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The AR device serves itself by using its own existing resources for lip-language identification. The device's camera captures the necessary facial images, its processor performs the lip movement analysis, and its display system presents the translated output. No external or additional hardware is required - the system leverages its inherent capabilities to provide the new function, maintaining portability while achieving accurate identification.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11527242B2Lip-language identification method and apparatus, and augmented reality (AR) device and storage medium which identifies an object based on an azimuth angle associated with the AR field of view
Publication Date: 2022.12.13 BEIJING BOE TECH DEV CO LTD
  • US11527242B2 patent drawing
  • US11527242B2 patent drawing
  • US11527242B2 patent drawing

AI summary

A lip-language identification method and an apparatus thereof, an augmented reality device and a storage medium. The lip-language identification method includes: acquiring a sequence of face images for an object to be identified; performing lip-language identification based on a sequence of face images so as to determine semantic information of speech content of the object to be identified corresponding to lip actions in a face image; and outputting the semantic information.