Avatar Viseme Rendering Using Camera-Based Mouth Movement Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In speech animation, avatars incorrectly render visemes corresponding to phonemes detected by a microphone, even if the wearer did not utter the speech, leading to inaccurate representation of the user's mouth movements.

Innovation Solution

The system determines whether the wearer uttered the speech by detecting mouth movement using cameras and sensors, and if not, prevents the rendering of visemes, ensuring accurate synchronization of avatar movements with the wearer's speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the system renders visemes for all detected phonemes, then the avatar appears to be speaking continuously, but the avatar incorrectly represents speech that the wearer did not utter

Engineering Contradiction:
Improveaccuracy of avatar speech representationVSAvoidcomplexity of speech detection system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary verification mechanism using camera-based mouth movement detection to mediate between the microphone's speech detection and the avatar's viseme rendering. The camera acts as a mediator that validates whether the wearer actually uttered the detected speech before triggering avatar mouth movements, thereby resolving the contradiction between reliable speech representation and system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by continuously monitoring the wearer's actual mouth movements via camera and comparing them against the detected speech phonemes. This feedback loop allows the system to verify whether rendered visemes accurately correspond to the wearer's actual speech, improving reliability by correcting mismatches between detected speech and actual utterance.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the system uses only microphone detection, then the implementation is simple, but the avatar renders visemes for ambient speech not uttered by the wearer

Engineering Contradiction:
Improveprecision of speech attributionVSAvoidcomplexity of multi-sensor system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech detection function into two independent components: microphone-based phoneme detection and camera-based mouth movement verification. This segmentation allows each sensor to perform its specialized function independently, with the microphone detecting phonemes and the camera verifying actual utterance, thereby improving measurement precision through multi-modal verification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system achieves multi-functionality by using the camera for dual purposes: capturing the wearer's environment and verifying mouth movements for speech attribution. This universal use of the camera sensor improves speech detection precision without proportionally increasing system complexity, as the same hardware serves multiple functions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If the system captures detailed facial images for verification, then the accuracy of mouth movement detection improves, but the processing load and privacy concerns increase

Engineering Contradiction:
Improveprecision of mouth movement detectionVSAvoidcomplexity of image processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential mouth movement information from captured facial images, rather than processing complete high-resolution facial data. By extracting specific mouth region features and movement patterns, the system achieves sufficient measurement precision for speech verification while reducing processing complexity and privacy concerns through selective data extraction.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240312093A1Rendering Avatar to Have Viseme Corresponding to Phoneme Within Detected Speech
Publication Date: 2024.09.19 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US20240312093A1 patent drawing
  • US20240312093A1 patent drawing
  • US20240312093A1 patent drawing

AI summary

Speech is detected using a microphone of a head-mountable display (HMD). The speech includes a phoneme. Whether a wearer of the HMD uttered the speech is determined. In response to determining that the wearer uttered the speech, an avatar representing the wearer is rendered to have a viseme corresponding to the phoneme.