Synthetic Video Feed Generation via Facial Landmark Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing technologies face challenges in maintaining a seamless and authentic video feed when a user steps away or is temporarily unavailable, leading to disruptions and increased bandwidth requirements due to the need for continuous high-definition video transmission.

Innovation Solution

A method that captures facial landmark constellations from a live video feed, transforms them into synthetic face images using a face model, and switches to prerecorded sequences during the user's absence, allowing for a seamless transition and reduced bandwidth usage by transmitting lightweight facial landmark containers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If continuous high-definition video transmission is used to maintain seamless video feed, then video quality is improved, but bandwidth requirements increase

Engineering Contradiction:
Improvevideo qualityVSAvoidbandwidth requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential facial landmark constellation data from the full video feed for transmission. Instead of sending complete high-definition video frames, the system identifies and transmits only the key facial feature points (landmarks) that define the user's facial geometry and expressions. This extraction approach maintains video quality perception while dramatically reducing the data quantity transmitted over the network.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a simplified digital copy of the user's face using only landmark coordinates rather than transmitting the actual video image data. The receiving端 reconstructs the visual representation from these landmark copies, which are minimal data representations containing only the essential geometric information needed to render the user's facial appearance and expressions.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If lightweight facial landmark containers are transmitted instead of full video feeds, then bandwidth usage is reduced, but video feed authenticity may be compromised

Engineering Contradiction:
Improvebandwidth usageVSAvoidvideo feed authenticity
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent performs preliminary action by pre-identifying and pre-transmitting the facial landmark constellation that defines the user's facial geometry before the actual video content is needed. The receiving端 uses these pre-transmitted landmarks to accurately reconstruct the user's facial representation, ensuring authenticity is maintained from the outset rather than attempting to compress or simplify video data after capture.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If synthetic face images are generated using facial landmarks, then bandwidth requirements are reduced, but system complexity increases

Engineering Contradiction:
Improvebandwidth requirementsVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical system of transmitting and processing large volumes of video pixel data with a computational system that transmits minimal landmark coordinates and uses algorithms to synthesize the visual representation. Instead of moving heavy video data through the network, the system substitutes this with lightweight coordinate data and computational reconstruction at the receiving end.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11671562B2Method for enabling synthetic autopilot video functions and for publishing a synthetic video feed as a virtual camera during a video call
Publication Date: 2023.06.06 PRESENT COMM INC
  • US11671562B2 patent drawing
  • US11671562B2 patent drawing
  • US11671562B2 patent drawing

AI summary

A method for publishing a synthetic video feed during a video call including, during an operating period: tracking a computational load of a first device; and receiving a sequence of frames in a video feed from a camera facing a first user. The method also includes, responsive to the computational load of the first device falling below a first computational load threshold: detecting the first user's face in the sequence of frames; generating facial landmark containers representing facial actions of the first user; inserting the facial landmark containers and a look model, into a synthetic face generator to generate a first synthetic video feed; and publishing the first synthetic video feed for access by a second device. The method further includes, responsive to the computational load of the first device exceeding the first computational load threshold, offloading generation of a second synthetic video feed to the second device.