Face Video Decoder Using Background-Preserving Generative Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding technologies face challenges in efficiently processing large amounts of digital video data while maintaining image quality, particularly in facial re-enactment applications where background distortion occurs when giving motion to a fundamental image containing a face.
Innovation Solution
A decoder system that utilizes a generative model to synthesize a face video by combining a fundamental image with background information, using geometric attributes to reduce background distortion and maintain image quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If conventional video coding technologies are used to process facial re-enactment, then processing speed is improved, but background distortion occurs and image quality deteriorates
Solution Approach 1:
The patent segments the video processing into two distinct parts: fundamental images containing facial information and separate background images. This segmentation allows independent processing of face motion (using geometric attributes) and background preservation, thereby maintaining image quality while achieving efficient processing speed.
2Productivity
If conventional video coding technologies are used to compress video data, then processing amount is reduced, but coding efficiency is insufficient for high-quality facial re-enactment
Solution Approach 1:
The patent extracts and separates the background component from the fundamental image, encoding them independently. The background image is extracted and preserved separately, while only the essential facial geometric attributes are encoded from the fundamental image. This extraction approach improves coding efficiency by focusing compression resources on the most important visual information.
3Measurement precision
If geometric attributes are applied to give motion to fundamental images, then motion accuracy is improved, but background distortion occurs
Solution Approach 1:
The patent introduces a background image as an intermediary component that mediates between the motion-corrected fundamental image and the final synthesized output. The geometric attributes accurately drive face motion in the fundamental image, while the separate background image is composited afterward, preventing background distortion while maintaining motion accuracy.
Data Source
AI summary
A decoder includes memory and circuitry coupled to the memory. In operation, the circuitry: decodes, from one or more streams, (i) a fundamental image that is an image including a face, (ii) geometric information indicating geometric attributes of a subject and corresponding to each of frames of a captured video by a camera, and (iii) background information regarding a background image; and generates a synthesized face video using a generative model from the fundamental image, the geometric information, and the background information. The synthesized face video is a video including the face and synthesized with the background image.


