Video Conferencing Facial Landmark Extraction for Bandwidth Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing technologies face challenges in maintaining high-quality video calls over low-bandwidth networks due to the need for high-bandwidth, low-latency data connections, and they lack features for users to control their visual presentation during calls.
Innovation Solution
A method for video conferencing that involves recording video frames, detecting facial landmarks, and transmitting these landmarks instead of raw video feeds, allowing for reconstruction of photorealistic synthetic video frames on the receiving device using local face reconstruction models, enabling low-bandwidth, low-latency, and high-quality video calls with user-controlled visual presentation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If raw video feeds are transmitted during video calls, then video quality is maintained, but bandwidth consumption increases and latency increases
Solution Approach 1:
The patent extracts only the essential facial landmark coordinates from complete video frames for transmission. Instead of sending entire video feeds, the system identifies and transmits only the key facial feature points (landmarks) that capture the essential visual information needed for reconstruction, thereby dramatically reducing bandwidth consumption while preserving video quality.
Solution Approach 2:
The receiving device creates a synthetic copy of the sender's video feed by mapping transmitted facial landmarks onto a pre-stored 3D face model. This synthetic video copy reproduces the sender's appearance and expressions without requiring transmission of the original video data, achieving quality preservation with minimal bandwidth usage.
2Reliability
If raw video feeds are transmitted during video calls, then video quality is maintained, but latency increases
Solution Approach 1:
The patent extracts only the essential facial landmark coordinates from complete video frames for transmission. Instead of sending entire video feeds, the system identifies and transmits only the key facial feature points (landmarks) that capture the essential visual information needed for reconstruction, thereby dramatically reducing bandwidth consumption while preserving video quality.
Solution Approach 2:
The receiving device performs preliminary actions by pre-storing 3D face models and reconstruction algorithms before video calls begin. When landmarks are received, the system can immediately map them to the pre-prepared models and generate synthetic video frames without delay, significantly reducing processing latency compared to analyzing and rendering video frames in real-time.
3Loss of energy
If facial landmark containers are transmitted instead of video feeds, then bandwidth is reduced, but video reconstruction complexity increases
Solution Approach 1:
The receiving device creates a synthetic copy of the sender's video feed by mapping transmitted facial landmarks onto a pre-stored 3D face model. This synthetic video copy reproduces the sender's appearance and expressions without requiring transmission of the original video data, achieving quality preservation with minimal bandwidth usage.
Solution Approach 2:
The patent transforms the video representation from pixel-based continuous data to parameter-based discrete landmarks (x, y, z coordinates of key facial points). This parameterization simplifies transmission while the receiving system uses these parameters to control a 3D model, converting them back to visual form through model deformation and rendering.
4Reliability
If users transmit their actual video feeds, then visual authenticity is maintained, but user control over visual presentation is lost
Solution Approach 1:
The patent enables dynamic control of visual presentation by allowing users to modify parameters of their 3D face model in real-time. Users can adjust lighting conditions, background scenes, virtual accessories, and other visual attributes of their synthetic video representation, providing adaptability and versatility while maintaining the authenticity of their facial expressions and movements through landmark-based tracking.
Data Source
AI summary
One variation of a method for video conferencing includes, at a first device associated with a first user: capturing a first video feed; representing constellations of facial landmarks, detected in the first video feed, in a first feed of facial landmark containers; and transmitting the first feed of facial landmark containers to a second device. The method further includes, at the second device associated with a second user: accessing a first face model representing facial characteristics of the first user; accessing a synthetic face generator; transforming the first feed of facial landmark containers and the first face model into a first feed of synthetic face images according to the synthetic face generator; and rendering the first feed of synthetic face images.


