Generative Face Video SEI Matching for Compressed Face Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding standards face challenges in efficiently encoding and decoding face information in video data, particularly in high-compression scenarios, leading to suboptimal quality and efficiency in applications like surveillance, conferencing, and live broadcasting.

Innovation Solution

The use of generative face video supplemental enhancement information (SEI) messages to enhance the encoding and decoding process, allowing for the reconstruction of face pictures using generative networks, by incorporating face information parameters and identifying network compatibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional video coding standards are used to encode face information, then compression is achieved, but encoding and decoding efficiency deteriorates and quality is suboptimal

Engineering Contradiction:
Improveface information qualityVSAvoidencoding and decoding efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent segments face information from general video content by introducing a dedicated face detection module and separate encoding path. Face regions are identified and extracted using bounding boxes or segmentation masks, then processed independently through SEI messages while the rest of the video follows conventional coding. This segmentation allows specialized handling of face regions to improve both quality and efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation for face information by introducing SEI messages that carry dedicated face attributes (landmarks, expressions, pose) alongside traditional video parameters. This parameter transformation enables more efficient encoding of face-specific characteristics while maintaining compatibility with existing video coding standards, thereby improving both encoding efficiency and reconstruction quality.

Inventive Principle:
Principle #35Parameter changes

2Loss of energy

If compression is increased to reduce bandwidth and storage, then transmission efficiency improves, but face information quality deteriorates

Engineering Contradiction:
Improvebandwidth and storage requirementsVSAvoidface information quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent introduces SEI messages as an intermediary layer between the video encoder and decoder, specifically for face information. These messages carry compressed yet essential face attributes (landmarks, expressions, pose parameters) that act as a mediator to preserve face quality even when the main video stream is heavily compressed. The intermediary SEI layer ensures that critical face characteristics are maintained independently of the compression level applied to the overall video.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms face information into a compact parameter representation within SEI messages, encoding essential face characteristics (landmarks, expressions, pose) in a compressed format. This parameter transformation allows high-quality face reconstruction from reduced data, enabling significant bandwidth savings while maintaining face information quality through efficient parameter-based representation rather than full-resolution face data.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If conventional encoding methods are used, then implementation is simple, but adaptability to different face scenarios deteriorates

Engineering Contradiction:
Improveencoding implementation complexityVSAvoidadaptability to different face scenarios
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal face encoding framework through SEI messages that can handle multiple face scenarios (different expressions, poses, lighting conditions, face sizes) within a single standardized structure. The SEI message format is designed to be multi-functional, accommodating various face attributes and generative network requirements, thereby providing broad adaptability across different surveillance, conferencing, and broadcasting scenarios without requiring separate encoding solutions for each case.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces dynamic adaptability by allowing the SEI message content and face detection parameters to adjust automatically based on scene requirements. The system can dynamically select which face attributes to encode, adjust compression levels, and adapt to varying face conditions (distance, angle, expression) while maintaining a relatively simple base implementation. This dynamic behavior enables the system to handle diverse face scenarios without proportionally increasing implementation complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20260012646A1Supplemental enhancement information (SEI) message for generative face video
Publication Date: 2026.01.08 ALIBABA (CHINA) CO LTD
  • US20260012646A1 patent drawing
  • US20260012646A1 patent drawing
  • US20260012646A1 patent drawing

AI summary

A method for decoding a bitstream includes: receiving a bitstream and decoding, using coded information of the bitstream, one or more pictures. The decoding of the one or more pictures includes: determining whether a generative face video supplemental enhancement information (SEI) message matches with a generative network; and in response to the generative face video SEI message matches with the generative network, decoding the SEI message. The decoding of the SEI message includes: determining a face information parameter and a base picture associated with the SEI message; and reconstructing a face picture based on the face information parameter and the base picture.