Facial Expression Recognition Using Hybrid Attention and Segmented CNN-RNN

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current facial expression recognition technologies face challenges in natural environments due to factors like head posture shifting, illumination changes, occlusion, and motion blur, leading to reduced recognition rates.

Innovation Solution

A facial expression recognition method and system that incorporates an attention mechanism, involving face detection, alignment, spatial feature extraction using a residual neural network, hybrid attention module for feature weighting, and temporal feature extraction with a recurrent neural network, to improve recognition accuracy by correlating information between video frames and eliminating irrelevant features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep cascade network structure (CNN+RNN) is used to extract spatial and temporal features, then recognition accuracy is improved, but gradient explosion or gradient disappearance occurs

Engineering Contradiction:
Improverecognition accuracyVSAvoidtraining stability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the deep cascade network into two independent modules: a CNN module for spatial feature extraction and an RNN module for temporal feature extraction. This segmentation allows each module to be trained separately with appropriate loss functions, avoiding the gradient propagation issues that occur in deeply cascaded architectures while maintaining the ability to capture both spatial and temporal dependencies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary feature representation that serves as the output of the CNN module and input to the RNN module. This intermediary layer acts as a buffer that decouples the gradient flow between CNN and RNN, preventing gradient explosion or disappearance while still enabling effective feature transfer between spatial and temporal processing stages.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If 3DCNN is used to add time dimension for temporal feature extraction, then time series information is obtained, but training parameters and calculation amount increase

Engineering Contradiction:
Improvetime series informationVSAvoidtraining parameters and calculation amount
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments temporal feature extraction from spatial feature extraction by using separate CNN and RNN modules. Instead of using 3DCNN which processes space and time together with high computational complexity, the patent first extracts spatial features with 2D CNN, then extracts temporal features from the spatial feature sequences using RNN, significantly reducing the number of parameters and computational requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the approach from adding a time dimension to the convolution operation (3DCNN) to processing temporal sequences through a separate recurrent computation dimension. This dimensional transformation allows temporal modeling without the computational burden of 3D convolutions, as RNN processes sequences step-by-step with fixed parameter sets.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If optical flow method is used to extract time series information from video sequence, then temporal correlation is obtained, but preprocessing time increases and real-time performance deteriorates

Engineering Contradiction:
Improvetemporal correlationVSAvoidpreprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts only the necessary temporal correlation information directly from the spatial feature sequences produced by CNN, without performing the extensive preprocessing required by optical flow methods. The RNN module processes the spatial feature sequences to capture temporal dependencies, eliminating the need for computationally intensive optical flow calculation while still obtaining effective temporal correlation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical optical flow computation process with a learned feature-based temporal modeling approach using RNN. Instead of calculating pixel-level motion fields through complex mathematical operations, the system uses the RNN to learn temporal patterns directly from the spatial feature representations, achieving comparable temporal correlation with much lower computational cost and faster processing speed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11967175B2Facial expression recognition method and system combined with attention mechanism
Publication Date: 2024.04.23 HUAZHONG NORMAL UNIV
  • US11967175B2 patent drawing
  • US11967175B2 patent drawing
  • US11967175B2 patent drawing

AI summary

Provided are a facial expression recognition method and system combined with an attention mechanism. The method comprises: detecting faces comprised in each video frame in a video sequence, and extracting corresponding facial ROIs, so as to obtain facial pictures in each video frame; aligning the facial pictures in each video frame on the basis of location information of facial feature points of the facial pictures; inputting the aligned facial pictures into a residual neural network, and extracting spatial features of facial expressions corresponding to the facial pictures; inputting the spatial features of the facial expressions into a hybrid attention module to acquire fused features of the facial expressions; inputting the fused features of the facial expressions into a gated recurrent unit, and extracting temporal features of the facial expressions; and inputting the temporal features of the facial expressions into a fully connected layer, and classifying and recognizing the facial expressions.