Face Video Encoding With Editable Emotion Attributes at Ultra-Low Bitrate
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing face video coding frameworks struggle to achieve ultra-low bitrate communication while supporting user-specified emotional editing, as they either lack flexibility in emotion editing or incur additional bitrate costs due to high-dimensional representations.
Innovation Solution
The proposed EmoCodec framework uses a three-level motion information approach, including ultra-compact face motion representation, finer-grained pose and expression motions, and semantic-level expression attributes, enabling flexible emotion editing and ultra-low bitrate communication through a conditional continuous normalizing flow mechanism and windowed smoothing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If face video coding frameworks use end-to-end animation models to achieve ultra-low bitrate communication, then bitrate is reduced, but emotion editing capability is lost
Solution Approach 1:
The patent segments the facial video representation into three distinct components: ultra-compact motion codes for bitrate reduction, pose information for structural accuracy, and expression attributes for emotion editing. This segmentation allows each component to serve its specific function independently, resolving the contradiction between compression and editability.
Solution Approach 2:
The patent extracts expression attributes as separable entities from the compressed video stream. By taking out the expression information as distinct parameters that can be independently manipulated, the system enables emotion editing without requiring full decompression or sacrificing the ultra-low bitrate achievement.
2Manufacturing precision
If face emotion editing schemes separate expression from other information, then emotion editing quality is improved, but bitrate increases due to high-dimensional representations
Solution Approach 1:
The patent applies local quality by using ultra-compact representations specifically for the motion codes while maintaining high-dimensional detailed representations only for the expression attributes that require editing. This localized approach to representation quality ensures bitrate efficiency in compression-critical areas while preserving editability where needed.
Solution Approach 2:
The patent creates a composite representation system that combines three different types of data structures: compact motion codes, pose information, and expression attributes. This composite approach leverages the strengths of each representation type, achieving both compression efficiency and editing capability without the bitrate penalty of uniformly high-dimensional representations.
Data Source
AI summary
There is provided an encoder for face videos, which includes a video codec for compressing an initial frame of a face video sequence, a compact feature extraction module for extracting a motion code and expression attributes across subsequent inter frames, and a feature encoding module for feature compression with feature-level inter prediction based on the motion code and the expression attributes.


