Facial Mesh User Masks for Low-Bandwidth Online Expression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing real-time interactive communication technologies face bandwidth constraints and lack effective methods to provide useful contextual information without relying on high-quality video or photo-realistic representations, especially in professional settings.
Innovation Solution
Employing low-resolution user masks or graphical representations of participants, generated through facial mesh analysis, to convey real-time interaction and contextual cues with minimal bandwidth requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If real-time video of participants is used, then contextual information about participants is provided, but bandwidth constraints are exceeded and interaction quality deteriorates
Solution Approach 1:
The patent extracts only the essential facial feature points from complete video feeds, transmitting merely the coordinates of key facial landmarks rather than full video data. This extraction approach maintains the ability to convey contextual information through facial expressions while dramatically reducing bandwidth consumption.
Solution Approach 2:
Instead of transmitting actual video feeds, the patent creates simplified graphical representations (masks) that copy only the essential facial expression information. These masks are generated by mapping facial landmark coordinates onto template faces, providing a bandwidth-efficient copy of the relevant contextual information.
2Measurement precision
If photo-realistic representations are used, then participant appearance is accurately represented, but bandwidth requirements increase and privacy concerns arise
Solution Approach 1:
The patent applies local quality by focusing computational and transmission resources only on the locally important facial feature points rather than the entire face. By identifying and tracking specific landmarks (eyes, eyebrows, mouth corners), the system achieves accurate facial expression representation while minimizing data transmission requirements.
Solution Approach 2:
The patent segments the face into discrete feature points or landmarks that can be independently tracked and transmitted. This segmentation allows the system to capture essential facial expression information through a small set of coordinate points rather than transmitting continuous video data.
3Reliability
If static avatars are used, then privacy concerns are addressed, but real-time contextual information about participant interaction is lost
Solution Approach 1:
The patent transforms static avatars into dynamic masks that update in real-time based on facial landmark coordinates. These masks can dynamically change their appearance to reflect current facial expressions, maintaining privacy protection while conveying real-time interaction context through animated facial features.
4Loss of information
If full video feeds are transmitted, then complete visual information is provided, but bandwidth constraints are violated and interaction quality deteriorates under bandwidth limitations
Solution Approach 1:
The patent extracts only the essential facial feature points from complete video feeds, transmitting merely the coordinates of key facial landmarks rather than full video data. This extraction approach maintains the ability to convey contextual information through facial expressions while dramatically reducing bandwidth consumption.
Data Source
Figure 1
Figure 2A~2B
Figure 3
AI summary
The technology provides enhanced co-presence of interactive media participants without high quality video or other photo-realistic representations of the participants. A low-resolution graphical representation (318) of a participant provides real-time dynamic co-presence. Face detection captures a maximum amount of facial expression with minimum detail in order to construct the low-resolution graphical representation. A set of facial mesh data (304) is generated by the face detection to include a minimal amount of information about the participant's face per frame. The mesh data, such as facial key points, is provided to one or more user devices so that the graphical representation of the participant can be rendered in a shared app at the other device(s) (308, 804). The rendering can include generating a hull (314, 806) that delineates a perimeter of the user mask and generating a set of facial features (316, 808), in which the graphical representation is assembled by combining the hull and the set of facial features (318, 810).