Joint Expression Coding for Facial Emotion Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current emotion recognition systems based on facial images and deep learning struggle to effectively combine static and dynamic facial expression information, leading to poor representation capability due to independent extraction of features and high computational complexity.
Innovation Solution
A joint expression coding system that combines static and dynamic expression images through preprocessing, dynamic expression image generation, dynamic weight image calculation, and joint expression coding, allowing simultaneous representation of static and dynamic information in a single image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If 3D convolution is used to extract space and time information simultaneously, then emotion recognition performance is improved, but computational complexity increases and processing efficiency decreases
Solution Approach 1:
The patent segments the video into multiple sub-videos and processes each sub-video independently through parallel channels. This segmentation reduces the computational burden on each processing unit while maintaining the ability to extract both spatial and temporal features through separate static and dynamic feature extraction pathways, thereby resolving the contradiction between comprehensive feature extraction and computational complexity
Solution Approach 2:
The patent transforms the 4D tensor processing problem of 3D convolution into a combination of 2D spatial processing and 1D temporal processing. By separating spatial feature extraction (through 2D convolutions on static frames) and temporal feature extraction (through 1D convolutions on dynamic feature sequences), the system achieves comparable emotion recognition performance with reduced computational complexity
2Productivity
If static and dynamic features are extracted independently through separate channels, then processing efficiency is improved, but the representation capability of facial features deteriorates due to lack of internal correlation
Solution Approach 1:
The patent introduces dynamic feature maps as an intermediary that bridges static and dynamic feature extraction channels. The dynamic feature maps, generated from optical flow information, serve as a mediator that provides temporal context to the static feature extraction process, enabling the system to maintain independent processing channels for efficiency while establishing internal correlations through the intermediary dynamic features
Solution Approach 2:
The patent merges static and dynamic feature representations at multiple stages of the processing pipeline. By combining static facial appearance features with dynamic motion features through concatenation and fusion operations, the system maintains processing efficiency through separate channels while achieving enhanced representation capability through the integration of complementary feature types
Data Source
AI summary
A joint expression coding system based on static and dynamic expression images is disclosed, including an image preprocessing module, a dynamic expression image generation module, a dynamic weight image generation module, and a joint expression coding image generation module. A joint expression coding method based on the joint expression coding system is also disclosed. A static expression image and a dynamic expression image are combined into one image according to the coding method by adopting the joint expression coding system and method based on the static and dynamic expression images, whereby static expression information and dynamic expression information can be represented at the same time, thus improving the emotion recognition capability based on facial expressions.


