Face Video Decoder Using Background-Preserving Generative Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding technologies face challenges in efficiently processing large amounts of digital video data while maintaining image quality, particularly in facial re-enactment applications where background distortion occurs when giving motion to a fundamental image containing a face.

Innovation Solution

A decoder system that utilizes a generative model to synthesize a face video by combining a fundamental image with background information, using geometric attributes to reduce background distortion and maintain image quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If conventional video coding technologies are used to process facial re-enactment, then processing speed is improved, but background distortion occurs and image quality deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidimage quality
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The patent segments the video processing into two distinct parts: fundamental images containing facial information and separate background images. This segmentation allows independent processing of face motion (using geometric attributes) and background preservation, thereby maintaining image quality while achieving efficient processing speed.

Inventive Principle:
Principle #1Segmentation

2Productivity

If conventional video coding technologies are used to compress video data, then processing amount is reduced, but coding efficiency is insufficient for high-quality facial re-enactment

Engineering Contradiction:
Improvecoding efficiencyVSAvoidprocessing amount
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts and separates the background component from the fundamental image, encoding them independently. The background image is extracted and preserved separately, while only the essential facial geometric attributes are encoded from the fundamental image. This extraction approach improves coding efficiency by focusing compression resources on the most important visual information.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If geometric attributes are applied to give motion to fundamental images, then motion accuracy is improved, but background distortion occurs

Engineering Contradiction:
Improvemotion accuracyVSAvoidbackground quality
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent introduces a background image as an intermediary component that mediates between the motion-corrected fundamental image and the final synthesized output. The geometric attributes accurately drive face motion in the fundamental image, while the separate background image is composited afterward, preventing background distortion while maintaining motion accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260045014A1Decoder, encoder, bitstream generator, decoding method, and encoding method
Publication Date: 2026.02.12 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US20260045014A1 patent drawing
  • US20260045014A1 patent drawing
  • US20260045014A1 patent drawing

AI summary

A decoder includes memory and circuitry coupled to the memory. In operation, the circuitry: decodes, from one or more streams, (i) a fundamental image that is an image including a face, (ii) geometric information indicating geometric attributes of a subject and corresponding to each of frames of a captured video by a camera, and (iii) background information regarding a background image; and generates a synthesized face video using a generative model from the fundamental image, the geometric information, and the background information. The synthesized face video is a video including the face and synthesized with the background image.