Multi-scale Spatiotemporal Neural Network for Micro-expression Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current micro-expression recognition methods face challenges in extracting sufficient and abstract features from video data, leading to superficial information extraction and inadequate representation, especially due to the short duration and small motion amplitude of micro-expressions.
Innovation Solution
A micro-expression recognition method based on a multi-scale spatiotemporal feature neural network is developed, which involves converting video data into image frame sequences, extracting face images, normalizing time scales, and constructing a spatiotemporal neural network that combines spatial and temporal feature extraction layers to improve feature representation and recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional manual feature extraction methods (LBP, LBP-TOP, directional average of optical flow) are used, then the extraction process is simple, but the extracted features are superficial and insufficient for accurate micro-expression recognition
Solution Approach 1:
The patent replaces traditional manual feature extraction methods with a deep learning-based automatic feature extraction system. The convolutional neural network automatically learns and extracts features from video sequences, substituting the manual mechanical process of hand-crafted feature extraction with an intelligent automated system that achieves superior recognition accuracy.
Solution Approach 2:
The patent transitions from traditional 2D image processing to 3D spatiotemporal feature extraction by incorporating temporal dimensions. The model processes video sequences across multiple frames and scales, extracting features in three dimensions (width, height, time) to capture dynamic micro-expression patterns that static 2D methods cannot detect.
2Loss of information
If deep learning models are applied to extract features, then feature representation capability is improved, but the computational complexity and processing time increase significantly
Solution Approach 1:
The patent segments the video input into multiple scale levels (e.g., different temporal resolutions and spatial regions). Each segment is processed independently by the neural network to extract features at different granularities, allowing comprehensive feature capture while distributing computational load across multiple parallel processing streams rather than analyzing the entire video sequence at once.
Solution Approach 2:
The patent extracts features from multiple scales and regions, potentially processing more data than strictly necessary for basic recognition. This excessive action ensures comprehensive feature capture and robust recognition accuracy by analyzing the video from multiple perspectives and temporal resolutions, compensating for the brief duration and subtle nature of micro-expressions.
3Measurement precision
If multiple scales and temporal sequences are analyzed, then recognition accuracy is improved, but the computational load and model complexity increase
Solution Approach 1:
The patent designs a universal neural network architecture that handles multiple functions: it processes different video scales, extracts spatiotemporal features, and performs classification. This multi-functional model reduces the need for separate specialized networks for each task, achieving high recognition accuracy while managing complexity through a unified framework that can adapt to various input configurations.
Data Source
AI summary
Disclosed is a micro-expression recognition method based on a multi-scale spatiotemporal feature neural network, in which spatial features and temporal features of micro-expression are obtained from micro-expression video frames, and combined together to form more robust micro-expression features, at the same time, since the micro-expression occurs in local areas of a face, active local areas of the face during occurrence of the micro-expression and an overall area of the face are combined together for micro-expression recognition.


