Face Video Encoding With Editable Emotion Attributes at Ultra-Low Bitrate

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing face video coding frameworks struggle to achieve ultra-low bitrate communication while supporting user-specified emotional editing, as they either lack flexibility in emotion editing or incur additional bitrate costs due to high-dimensional representations.

Innovation Solution

The proposed EmoCodec framework uses a three-level motion information approach, including ultra-compact face motion representation, finer-grained pose and expression motions, and semantic-level expression attributes, enabling flexible emotion editing and ultra-low bitrate communication through a conditional continuous normalizing flow mechanism and windowed smoothing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If face video coding frameworks use end-to-end animation models to achieve ultra-low bitrate communication, then bitrate is reduced, but emotion editing capability is lost

Engineering Contradiction:
ImprovebitrateVSAvoidemotion editing capability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent segments the facial video representation into three distinct components: ultra-compact motion codes for bitrate reduction, pose information for structural accuracy, and expression attributes for emotion editing. This segmentation allows each component to serve its specific function independently, resolving the contradiction between compression and editability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts expression attributes as separable entities from the compressed video stream. By taking out the expression information as distinct parameters that can be independently manipulated, the system enables emotion editing without requiring full decompression or sacrificing the ultra-low bitrate achievement.

Inventive Principle:
Principle #2Taking out (Extraction)

2Manufacturing precision

If face emotion editing schemes separate expression from other information, then emotion editing quality is improved, but bitrate increases due to high-dimensional representations

Engineering Contradiction:
Improveemotion editing qualityVSAvoidbitrate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies local quality by using ultra-compact representations specifically for the motion codes while maintaining high-dimensional detailed representations only for the expression attributes that require editing. This localized approach to representation quality ensures bitrate efficiency in compression-critical areas while preserving editability where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent creates a composite representation system that combines three different types of data structures: compact motion codes, pose information, and expression attributes. This composite approach leverages the strengths of each representation type, achieving both compression efficiency and editing capability without the bitrate penalty of uniformly high-dimensional representations.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20260067488A1Encoding and decoding for face videos
Publication Date: 2026.03.05 CITY UNIVERSITY OF HONG KONG
  • US20260067488A1 patent drawing
  • US20260067488A1 patent drawing
  • US20260067488A1 patent drawing

AI summary

There is provided an encoder for face videos, which includes a video codec for compressing an initial frame of a face video sequence, a compact feature extraction module for extracting a motion code and expression attributes across subsequent inter frames, and a feature encoding module for feature compression with feature-level inter prediction based on the motion code and the expression attributes.