Facial Landmark Map Compression for Low Bandwidth Video Chat
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Many users worldwide are unable to engage in video-chat communication due to prohibitive data costs or reliance on outdated technologies and infrastructures, which often result in insufficient bandwidth, such as 2G networks that can only support up to 30 kbits/s, while current video-call quality requires at least 200 kbits/s.
Innovation Solution
The method involves compressing video data by generating a feature map and landmark maps that represent key facial features, allowing for reduced bandwidth transmission by sending only these maps instead of full image and pixel information, and using machine-learning models to decode and reconstruct the video on the receiving device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional video compression methods are used, then video quality can be maintained, but bandwidth consumption increases making video-chat unavailable on low-bandwidth networks
Solution Approach 1:
The patent extracts only the essential facial feature information (landmark maps) from complete video frames for transmission. Instead of sending full image data, the system identifies and transmits only the coordinates and characteristics of key facial landmarks, dramatically reducing bandwidth consumption while preserving sufficient information for video-chat functionality
Solution Approach 2:
The patent applies different processing quality levels to different parts of the video data. Critical facial regions are captured with high precision through landmark detection, while non-critical areas are either omitted or represented with minimal data, optimizing the balance between video quality and bandwidth usage
2Loss of information
If full image data is transmitted, then video quality is maintained, but data costs become prohibitive for users with limited data plans
Solution Approach 1:
The system extracts only the necessary facial landmark coordinates from complete video frames. By identifying key facial features (eyes, nose, mouth, etc.) and transmitting only their positional data rather than full pixel information, the patent reduces data volume by orders of magnitude while retaining essential visual information for communication
Solution Approach 2:
The patent creates simplified representations (landmark maps) that copy only the essential structural information of facial features. These landmark maps serve as compact substitutes for full images, preserving the geometric relationships and relative positions of facial landmarks without requiring complete pixel data
3Quantity of substance
If video compression is applied to reduce bandwidth, then bandwidth consumption decreases, but video quality and recognition accuracy may deteriorate
Solution Approach 1:
The patent concentrates computational and transmission resources on capturing facial landmark positions with high precision. By focusing on specific critical regions (key facial landmarks) rather than attempting to preserve all image details, the system achieves accurate facial feature recognition while using minimal bandwidth
Solution Approach 2:
The patent replaces traditional mechanical/image-based compression methods with a computational approach using machine learning models. The encoder uses trained neural networks to detect and encode landmark positions, while the decoder uses corresponding models to reconstruct facial representations, achieving better precision than conventional compression at equivalent bandwidth levels
Data Source
AI summary
In one embodiment, a first device may receive, from a second device, a reference landmark map identifying locations of facial features of a user of the second device depicted in a reference image and a feature map, generated based on the reference image, representing an identity of the user. The first device may receive, from the second device, a current compressed landmark map based on a current image of the user and decompress the current compressed landmark map to generate a current landmark map. The first device may update the feature map based on a motion field generated using the reference landmark map and the current landmark map. The first device may generate scaling factors based on a normalization facial mask of pre-determined facial features of the user. The first device may generate an output image of the user by decoding the updated feature map using the scaling factors.


