Facial Landmark Map Compression for Low Bandwidth Video Chat

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Many users worldwide are unable to engage in video-chat communication due to prohibitive data costs or reliance on outdated technologies and infrastructures, which often result in insufficient bandwidth, such as 2G networks that can only support up to 30 kbits/s, while current video-call quality requires at least 200 kbits/s.

Innovation Solution

The method involves compressing video data by generating a feature map and landmark maps that represent key facial features, allowing for reduced bandwidth transmission by sending only these maps instead of full image and pixel information, and using machine-learning models to decode and reconstruct the video on the receiving device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional video compression methods are used, then video quality can be maintained, but bandwidth consumption increases making video-chat unavailable on low-bandwidth networks

Engineering Contradiction:
Improvevideo-chat availabilityVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential facial feature information (landmark maps) from complete video frames for transmission. Instead of sending full image data, the system identifies and transmits only the coordinates and characteristics of key facial landmarks, dramatically reducing bandwidth consumption while preserving sufficient information for video-chat functionality

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different processing quality levels to different parts of the video data. Critical facial regions are captured with high precision through landmark detection, while non-critical areas are either omitted or represented with minimal data, optimizing the balance between video quality and bandwidth usage

Inventive Principle:
Principle #3Local quality

2Loss of information

If full image data is transmitted, then video quality is maintained, but data costs become prohibitive for users with limited data plans

Engineering Contradiction:
Improvevideo information completenessVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only the necessary facial landmark coordinates from complete video frames. By identifying key facial features (eyes, nose, mouth, etc.) and transmitting only their positional data rather than full pixel information, the patent reduces data volume by orders of magnitude while retaining essential visual information for communication

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates simplified representations (landmark maps) that copy only the essential structural information of facial features. These landmark maps serve as compact substitutes for full images, preserving the geometric relationships and relative positions of facial landmarks without requiring complete pixel data

Inventive Principle:
Principle #26Copying

3Quantity of substance

If video compression is applied to reduce bandwidth, then bandwidth consumption decreases, but video quality and recognition accuracy may deteriorate

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidfacial feature recognition accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent concentrates computational and transmission resources on capturing facial landmark positions with high precision. By focusing on specific critical regions (key facial landmarks) rather than attempting to preserve all image details, the system achieves accurate facial feature recognition while using minimal bandwidth

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent replaces traditional mechanical/image-based compression methods with a computational approach using machine learning models. The encoder uses trained neural networks to detect and encode landmark positions, while the decoder uses corresponding models to reconstruct facial representations, achieving better precision than conventional compression at equivalent bandwidth levels

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12026921B2Systems and method for low bandwidth video-chat compression
Publication Date: 2024.07.02 META PLATFORMS INC
  • US12026921B2 patent drawing
  • US12026921B2 patent drawing
  • US12026921B2 patent drawing

AI summary

In one embodiment, a first device may receive, from a second device, a reference landmark map identifying locations of facial features of a user of the second device depicted in a reference image and a feature map, generated based on the reference image, representing an identity of the user. The first device may receive, from the second device, a current compressed landmark map based on a current image of the user and decompress the current compressed landmark map to generate a current landmark map. The first device may update the feature map based on a motion field generated using the reference landmark map and the current landmark map. The first device may generate scaling factors based on a normalization facial mask of pre-determined facial features of the user. The first device may generate an output image of the user by decoding the updated feature map using the scaling factors.