Multi-scale Spatiotemporal Neural Network for Micro-expression Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current micro-expression recognition methods face challenges in extracting sufficient and abstract features from video data, leading to superficial information extraction and inadequate representation, especially due to the short duration and small motion amplitude of micro-expressions.

Innovation Solution

A micro-expression recognition method based on a multi-scale spatiotemporal feature neural network is developed, which involves converting video data into image frame sequences, extracting face images, normalizing time scales, and constructing a spatiotemporal neural network that combines spatial and temporal feature extraction layers to improve feature representation and recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional manual feature extraction methods (LBP, LBP-TOP, directional average of optical flow) are used, then the extraction process is simple, but the extracted features are superficial and insufficient for accurate micro-expression recognition

Engineering Contradiction:
Improverecognition accuracyVSAvoidfeature extraction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional manual feature extraction methods with a deep learning-based automatic feature extraction system. The convolutional neural network automatically learns and extracts features from video sequences, substituting the manual mechanical process of hand-crafted feature extraction with an intelligent automated system that achieves superior recognition accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transitions from traditional 2D image processing to 3D spatiotemporal feature extraction by incorporating temporal dimensions. The model processes video sequences across multiple frames and scales, extracting features in three dimensions (width, height, time) to capture dynamic micro-expression patterns that static 2D methods cannot detect.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If deep learning models are applied to extract features, then feature representation capability is improved, but the computational complexity and processing time increase significantly

Engineering Contradiction:
Improvefeature information completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent segments the video input into multiple scale levels (e.g., different temporal resolutions and spatial regions). Each segment is processed independently by the neural network to extract features at different granularities, allowing comprehensive feature capture while distributing computational load across multiple parallel processing streams rather than analyzing the entire video sequence at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts features from multiple scales and regions, potentially processing more data than strictly necessary for basic recognition. This excessive action ensures comprehensive feature capture and robust recognition accuracy by analyzing the video from multiple perspectives and temporal resolutions, compensating for the brief duration and subtle nature of micro-expressions.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If multiple scales and temporal sequences are analyzed, then recognition accuracy is improved, but the computational load and model complexity increase

Engineering Contradiction:
Improvemicro-expression recognition accuracyVSAvoidneural network structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent designs a universal neural network architecture that handles multiple functions: it processes different video scales, extracts spatiotemporal features, and performs classification. This multi-functional model reduces the need for separate specialized networks for each task, achieving high recognition accuracy while managing complexity through a unified framework that can adapt to various input configurations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11908240B2Micro-expression recognition method based on multi-scale spatiotemporal feature neural network
Publication Date: 2024.02.20 INST OF AUTOMATION CHINESE ACAD OF SCI
  • US11908240B2 patent drawing
  • US11908240B2 patent drawing
  • US11908240B2 patent drawing

AI summary

Disclosed is a micro-expression recognition method based on a multi-scale spatiotemporal feature neural network, in which spatial features and temporal features of micro-expression are obtained from micro-expression video frames, and combined together to form more robust micro-expression features, at the same time, since the micro-expression occurs in local areas of a face, active local areas of the face during occurrence of the micro-expression and an overall area of the face are combined together for micro-expression recognition.