Face Video SEI Compression for Accurate Low-Bitrate Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding standards, such as HEVC and VVC, face challenges in efficiently compressing and decompressing video data, particularly in applications requiring high compression efficiency and accurate reconstruction of facial features, which are not adequately addressed by current supplemental enhancement information (SEI) messages.

Innovation Solution

The method involves decoding a bitstream to identify the use of a face video generative compression scheme, extracting facial information from a supplemental enhancement information (SEI) message, and reconstructing a face picture based on this information and a base picture, while encoding a video sequence to signal the use of this scheme.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional video compression standards (HEVC/VVC) are used, then general video compression is achieved, but face video reconstruction quality deteriorates at ultra-low bitrates

Engineering Contradiction:
Improveface video reconstruction qualityVSAvoidbitrate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments the video compression process by introducing a dedicated face video generative compression SEI message that specifically targets facial regions. This SEI message carries facial information (landmarks, expressions, head movements) separately from the base video stream, allowing specialized reconstruction of face portions while maintaining standard compression for the rest of the video content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation by encoding facial information as discrete parameters (landmark coordinates, expression coefficients, head pose angles) rather than full pixel data. These parameters are transmitted through the SEI message and used to drive generative models that synthesize high-quality face regions at ultra-low bitrates.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If standard video coding technologies are used, then coding efficiency is improved, but facial expressions and head movements are not accurately rendered

Engineering Contradiction:
Improvecoding efficiencyVSAvoidfacial expression rendering accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary SEI message that carries facial information between the encoder and decoder. This SEI message acts as a mediator that supplements the base video stream with dedicated facial parameters, enabling the decoder to accurately reconstruct facial expressions and head movements without compromising the overall coding efficiency of the video stream.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies dynamics by using generative models that can adaptively reconstruct facial regions based on the transmitted parameters. The system dynamically adjusts the reconstruction process according to the facial information contained in the SEI message, allowing accurate rendering of varying facial expressions and head movements while maintaining efficient compression.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12621492B2Method and apparatuses for using face video generative compression SEI message
Publication Date: 2026.05.05 ALIBABA (CHINA) CO LTD
  • US12621492B2 patent drawing
  • US12621492B2 patent drawing
  • US12621492B2 patent drawing

AI summary

A method of decoding a bitstream to output one or more pictures for a video stream, includes: receiving a bitstream; and decoding, using coded information of the bitstream, one or more pictures. The decoding includes: determining, based on an identifying number, whether a face video generative compression scheme is used; in response to a determination that the face video generative compression scheme is used, decoding a supplemental enhancement information (SEI) message, the SEI message comprising facial information; and reconstructing a face picture based on the facial information and a base picture associated with the SEI message.