Avatar Viseme Rendering Using Camera-Based Mouth Movement Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In speech animation, avatars incorrectly render visemes corresponding to phonemes detected by a microphone, even if the wearer did not utter the speech, leading to inaccurate representation of the user's mouth movements.
Innovation Solution
The system determines whether the wearer uttered the speech by detecting mouth movement using cameras and sensors, and if not, prevents the rendering of visemes, ensuring accurate synchronization of avatar movements with the wearer's speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the system renders visemes for all detected phonemes, then the avatar appears to be speaking continuously, but the avatar incorrectly represents speech that the wearer did not utter
Solution Approach 1:
The patent introduces an intermediary verification mechanism using camera-based mouth movement detection to mediate between the microphone's speech detection and the avatar's viseme rendering. The camera acts as a mediator that validates whether the wearer actually uttered the detected speech before triggering avatar mouth movements, thereby resolving the contradiction between reliable speech representation and system complexity.
Solution Approach 2:
The system implements feedback by continuously monitoring the wearer's actual mouth movements via camera and comparing them against the detected speech phonemes. This feedback loop allows the system to verify whether rendered visemes accurately correspond to the wearer's actual speech, improving reliability by correcting mismatches between detected speech and actual utterance.
2Measurement precision
If the system uses only microphone detection, then the implementation is simple, but the avatar renders visemes for ambient speech not uttered by the wearer
Solution Approach 1:
The patent segments the speech detection function into two independent components: microphone-based phoneme detection and camera-based mouth movement verification. This segmentation allows each sensor to perform its specialized function independently, with the microphone detecting phonemes and the camera verifying actual utterance, thereby improving measurement precision through multi-modal verification.
Solution Approach 2:
The system achieves multi-functionality by using the camera for dual purposes: capturing the wearer's environment and verifying mouth movements for speech attribution. This universal use of the camera sensor improves speech detection precision without proportionally increasing system complexity, as the same hardware serves multiple functions.
3Measurement precision
If the system captures detailed facial images for verification, then the accuracy of mouth movement detection improves, but the processing load and privacy concerns increase
Solution Approach 1:
The patent extracts only the essential mouth movement information from captured facial images, rather than processing complete high-resolution facial data. By extracting specific mouth region features and movement patterns, the system achieves sufficient measurement precision for speech verification while reducing processing complexity and privacy concerns through selective data extraction.
Data Source
AI summary
Speech is detected using a microphone of a head-mountable display (HMD). The speech includes a phoneme. Whether a wearer of the HMD uttered the speech is determined. In response to determining that the wearer uttered the speech, an avatar representing the wearer is rendered to have a viseme corresponding to the phoneme.


