Video Conferencing Facial Landmark Extraction for Bandwidth Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing technologies face challenges in maintaining high-quality video calls over low-bandwidth networks due to the need for high-bandwidth, low-latency data connections, and they lack features for users to control their visual presentation during calls.

Innovation Solution

A method for video conferencing that involves recording video frames, detecting facial landmarks, and transmitting these landmarks instead of raw video feeds, allowing for reconstruction of photorealistic synthetic video frames on the receiving device using local face reconstruction models, enabling low-bandwidth, low-latency, and high-quality video calls with user-controlled visual presentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If raw video feeds are transmitted during video calls, then video quality is maintained, but bandwidth consumption increases and latency increases

Engineering Contradiction:
Improvevideo qualityVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent extracts only the essential facial landmark coordinates from complete video frames for transmission. Instead of sending entire video feeds, the system identifies and transmits only the key facial feature points (landmarks) that capture the essential visual information needed for reconstruction, thereby dramatically reducing bandwidth consumption while preserving video quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The receiving device creates a synthetic copy of the sender's video feed by mapping transmitted facial landmarks onto a pre-stored 3D face model. This synthetic video copy reproduces the sender's appearance and expressions without requiring transmission of the original video data, achieving quality preservation with minimal bandwidth usage.

Inventive Principle:
Principle #26Copying

2Reliability

If raw video feeds are transmitted during video calls, then video quality is maintained, but latency increases

Engineering Contradiction:
Improvevideo qualityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the essential facial landmark coordinates from complete video frames for transmission. Instead of sending entire video feeds, the system identifies and transmits only the key facial feature points (landmarks) that capture the essential visual information needed for reconstruction, thereby dramatically reducing bandwidth consumption while preserving video quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The receiving device performs preliminary actions by pre-storing 3D face models and reconstruction algorithms before video calls begin. When landmarks are received, the system can immediately map them to the pre-prepared models and generate synthetic video frames without delay, significantly reducing processing latency compared to analyzing and rendering video frames in real-time.

Inventive Principle:
Principle #10Preliminary action

3Loss of energy

If facial landmark containers are transmitted instead of video feeds, then bandwidth is reduced, but video reconstruction complexity increases

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidvideo reconstruction complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The receiving device creates a synthetic copy of the sender's video feed by mapping transmitted facial landmarks onto a pre-stored 3D face model. This synthetic video copy reproduces the sender's appearance and expressions without requiring transmission of the original video data, achieving quality preservation with minimal bandwidth usage.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the video representation from pixel-based continuous data to parameter-based discrete landmarks (x, y, z coordinates of key facial points). This parameterization simplifies transmission while the receiving system uses these parameters to control a 3D model, converting them back to visual form through model deformation and rendering.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If users transmit their actual video feeds, then visual authenticity is maintained, but user control over visual presentation is lost

Engineering Contradiction:
Improvevisual authenticityVSAvoiduser control over visual presentation
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent enables dynamic control of visual presentation by allowing users to modify parameters of their 3D face model in real-time. Users can adjust lighting conditions, background scenes, virtual accessories, and other visual attributes of their synthetic video representation, providing adaptability and versatility while maintaining the authenticity of their facial expressions and movements through landmark-based tracking.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11889230B2Video conferencing method
Publication Date: 2024.01.30 PRESENT COMM INC
  • US11889230B2 patent drawing
  • US11889230B2 patent drawing
  • US11889230B2 patent drawing

AI summary

One variation of a method for video conferencing includes, at a first device associated with a first user: capturing a first video feed; representing constellations of facial landmarks, detected in the first video feed, in a first feed of facial landmark containers; and transmitting the first feed of facial landmark containers to a second device. The method further includes, at the second device associated with a second user: accessing a first face model representing facial characteristics of the first user; accessing a synthetic face generator; transforming the first feed of facial landmark containers and the first face model into a first feed of synthetic face images according to the synthetic face generator; and rendering the first feed of synthetic face images.