Efficient micro-expression emotion recognition method

By combining multi-level and multi-scale terraced feature extraction with the spatiotemporal genetic attention module, the problems of ignoring the secondary area features of micro-expressions and insufficient capture of motion changes in existing technologies are solved, efficient micro-expression emotion recognition is achieved, the recognition accuracy is improved and it is suitable for multiple application scenarios.

CN120823635APending Publication Date: 2025-10-21CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510998052.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing micro-expression recognition methods often only focus on the most significant regional features in micro-expressions, ignore secondary regional features, and cannot effectively capture the weak and short-term motion changes in micro-expression video sequences.

Method used

A multi-level and multi-scale terraced feature extraction module and a spatiotemporal genetic attention module are used, combined with ResNet50 and Bi-LSTM networks, to perform intra-frame feature extraction and inter-frame feature capture, enhance the perception of key and secondary areas of micro-expressions, and capture subtle dynamic changes.

Benefits of technology

It significantly improves the accuracy of micro-expression emotion recognition, especially on the SAMM, CASME II and CAS(ME)3 datasets, where its performance far exceeds existing methods. It is suitable for fields such as psychological analysis, interview screening, anti-fraud, medical assistance and human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823635A_ABST
    Figure CN120823635A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of machine vision, and particularly relates to an efficient micro-expression emotion recognition method, which comprises the following steps of: 1, performing preprocessing operation on SAMM, CASME II and CAS (ME) 3 data sets to generate data for training; 2, dividing the preprocessed data into a training set, a verification set and a test set; the multi-stage multi-scale terrace feature extraction module and the space-time genetic attention module act on an intra-frame feature extraction stage and inter-frame tiny motion perception, so that a network can effectively perceive an area where main features occur and an area where secondary features occur, and loss of micro-expression features is reduced. According to the provided micro-expression emotion recognition model, the micro-expression emotion recognition capability is remarkably improved, and effective data support can be provided for scenes such as psychological analysis, medical assistance and human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision technology, and in particular to an efficient micro-expression emotion recognition method. Background Art

[0002] Facial expressions are one of the most direct forms of human expression, conveying 55% of all information. Facial expressions can be categorized into macro- and micro-expressions. Macro-expressions occur spontaneously, are generally frequent, and are often seen in interpersonal communication. They are an objective reflection of a person's inner emotions and can often be used to conceal their true feelings. In contrast, micro-expressions occur involuntarily, are characterized by low intensity, small amplitude, and a duration of only 0.04-0.2 seconds. Micro-expressions often reflect a person's inner emotions, whether conscious or unconscious. These expressions typically appear in multiple facial regions, such as the eyes and corners of the mouth, with minimal amplitude. Micro-expression emotion recognition helps reveal an individual's underlying emotional state and is widely used in scenarios such as psychological analysis, interview screening, fraud prevention, medical assistance, and human-computer interaction. In the process of micro-expression emotion recognition, extracting discriminative micro-expression features significantly improves the recognition accuracy of the model. However, existing micro-expression recognition methods often only focus on the most prominent features in a micro-expression, while ignoring features in other less significant regions. This results in an inability to extract effective features from multiple relevant micro-expression facial regions. Second, existing micro-expression recognition techniques often focus only on the start, apex, and end frames, failing to detect subtle, brief motion changes in micro-expression video sequences.

[0003] Therefore, we propose a multi-level and multi-scale terraced feature extraction and spatiotemporal genetic attention method for micro-expression emotion recognition to solve the above problems. Summary of the Invention

[0004] (1) Technical problems solved

[0005] In view of the shortcomings of the existing technology, the present invention provides an efficient micro-expression emotion recognition method, which solves the problems raised in the above background technology.

[0006] (2) Technical solution

[0007] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0008] An efficient micro-expression emotion recognition method comprises the following steps:

[0009] S1: SAMM, CASME II and CAS(ME) 3 The dataset is preprocessed to generate data for training;

[0010] S2: Split the preprocessed data into training set, validation set and test set;

[0011] S3: Use ResNet50 and the multi-level multi-scale terrace feature extraction module to extract intra-frame features, realizing feature extraction of the main and secondary occurrence areas of micro-expressions;

[0012] S4: Use Bi-LSTM network and spatiotemporal genetic attention module to extract inter-frame features and capture the subtle dynamic changes of micro-expressions between frames;

[0013] S5: The entire network structure was trained for 100 rounds to construct a micro-expression emotion recognition model, which integrates the efficient micro-expression emotion recognition method TM-SGANet;

[0014] S6: Comparative analysis and optimization of the detection performance of TM-SGANet with other micro-expression models.

[0015] Furthermore, the S1 is for SAMM, CASME II and CAS (ME) 3 During data preprocessing, MediaPipe was used for face detection and key point annotation, and faces were aligned and cropped. Temporal interpolation was then used to lengthen the sequence to a specified length of 20 frames. Finally, mirror flipping, random angle rotation, and scaling were performed to enhance the data to prevent overfitting and alleviate the problems of small sample size and severe class unevenness.

[0016] Furthermore, the processed data in S2 is randomly divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0017] Furthermore, in S3, we innovatively introduced a multi-level and multi-scale terraced feature extraction module in the last four stages of ResNet50. This module enhances the model's perception of key and secondary areas of micro-expressions while retaining richer detail information. For multi-level feature extraction, three parallel 2D convolutions are used to capture features of different receptive fields. The parallel approach reduces the information loss caused by serial convolution. Each convolution is followed by a batch normalization operation to stabilize the data distribution within a certain range, alleviate internal covariate shift, accelerate model convergence, and improve generalization ability. Since convolution parameter errors may lead to feature estimation bias, the maximum pooling layer retains significant facial texture information while reducing the dimensionality, further reducing the shift error and enhancing the robustness of the feature. For multi-scale feature extraction, the network uses a cross-layer connection-based method to optimize the underlying feature maps, obtaining feature representations at different scales. Specifically, four convolutions with a stride of 2 are performed to increase the scale, resulting in feature sequences at five scales. High-level features with rich semantic information are then upsampled and fused with the features of the previous layer as the feature representation of the current scale. This method enables the model to obtain representative semantic information from the upsampled features and combine it with the underlying feature maps to obtain finer detail information, reducing the detail loss caused by convolution. The multi-level, multi-scale terraced feature extraction module achieves precise positioning of key micro-expression areas and multi-level feature expression through multi-level parallel convolution and multi-scale cross-layer fusion. Experiments show that this module effectively improves the model's focus on the primary and secondary features of micro-expressions, laying a better feature foundation for subsequent time series modeling and classification tasks.

[0018] Furthermore, the inter-frame data of the video in S4 is subjected to inter-frame feature extraction using a Bi-LSTM network and a spatiotemporal genetic attention module, so that the model can better handle small movements between frames of the video sequence. This method uses 6 micro-motion attention modules. In order to avoid the performance degradation problem caused by the deep structure, a residual connection structure is introduced before each micro-motion attention module. This structure effectively alleviates the problem of gradient vanishing or information loss in the deep network by adding the module input to the residual output of the convolution branch. Specifically, for the i-th residual structure, its input is recorded as , after 1×1 convolution and 3×3 convolution After the operation, the output of the residual structure is obtained In order to construct a more comprehensive spatial attention map, each incremental attention module combines the maximum pooling and average pooling strategies when generating the attention map, which are used to extract significant area features and global context features respectively. Through the splicing operation, combined with the 1×1 convolution layer and the Mish activation function, an attention map is generated that can guide the network to focus on subtle facial motion areas. Next, the currently generated attention map is multiplied point by point with the attention map of the previous layer to achieve spatial attention weighting of the current layer feature map and introduce temporal consistency in attention propagation. This strategy of introducing the previous layer attention map can not only dynamically update the attention distribution, but also spatially weight the previous attention map, so that the network's attention to the motion change area is more accurate. It is worth noting that the first incremental attention module does not introduce the previous attention map when generating the attention map. The formula for the generated attention map is:

[0019]

[0020] In the above formula, is the input of the first micro-movement attention module; is the Mish function; and They are maximum pooling and average pooling respectively. Other modules introduce the front layer attention map when generating the attention map, and the attention map is

[0021]

[0022] in, Denotes an element-by-element dot product operation. By element-by-element multiplication of the generated attention map with the previous frame's attention map, we obtain a feature map capable of perceiving subtle local motion. This multi-frame concatenation mechanism enables the network, guided by the dynamic updates of the micro-motion attention map, to gradually focus on regions with subtle motion changes. This robustly extracts and learns motion features from micro-expression images, achieving more accurate and robust feature representation.

[0023] Furthermore, the S5 model was trained using an NVIDIA GeForce RTX 4090, Python 3.9, and the PyTorch framework. SGD was used as the optimizer for training parameters; the momentum factor was set to 0.9; the learning rate was set to 1x10⁻4 and dynamically adjusted, using a cosine annealing strategy; the dropout rate was set to 0.4 in the dropout layer and 0.5 in the fully connected layer, using the cross-entropy loss function.

[0024] Furthermore, S6 was compared and analyzed with existing micro-expression recognition networks, using the same dataset and number of classification categories, consistent evaluation criteria, and a "leave-one-subject" cross-validation method to ensure consistency among the experimental benchmark, evaluation criteria, and validation methods. This leave-one-subject method uses the data of one subject as the test set and the data of all remaining subjects as the training set. The average of N test results is then used as the final performance metric. This effectively evaluates the generalization ability of a micro-expression recognition model on unseen individuals, avoiding data leakage and overfitting.

[0025] (3) Beneficial effects

[0026] Compared with the existing technology, the present invention provides an efficient micro-expression emotion recognition method with the following beneficial effects:

[0027] We retain the effective intra-frame features of micro-expressions by accurately identifying the primary and secondary regions where micro-expressions occur, and use spatiotemporal genetic attention methods to capture subtle motion changes between video frames, which greatly improves the emotion recognition ability of micro-expressions. Our method achieves 91.58% and 90.36% accuracy in five categories on the SAMM and CASME II micro-expression datasets, respectively. 3 This method achieved superior performance on the dataset, surpassing existing experimental methods, with an accuracy of 64.89% for five emotions, significantly exceeding baseline methods. This method can be applied to fields such as psychological analysis, interview screening, fraud prevention, medical assistance, and human-computer interaction, providing effective support for emotional data. The model has a very small number of parameters and can be used for real-time detection on mobile devices, thus meeting instant judgment needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 Schematic diagram of the method flow of the present invention

[0029] Figure 2 This is a schematic diagram of the overall network structure of the present invention;

[0030] Figure 3 This is a schematic diagram of the multi-level and multi-scale terrace feature extraction structure of the present invention;

[0031] Figure 4 Schematic diagram of the spatiotemporal genetic attention structure of the present invention;

[0032] Figure 5 Schematic diagram of the data processing structure of the present invention

[0033] Figure 6 This is the confusion matrix diagram of the recognition effect of the present invention DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0035] Example

[0036] like Figure 1-6 As shown, an efficient micro-expression emotion recognition method proposed in one embodiment of the present invention includes the following steps:

[0037] S1: SAMM, CASME II and CAS(ME) 3 The dataset is preprocessed to generate data for training;

[0038] Micro-expression video data contains variations in facial pose and background. To minimize their impact on model recognition performance, this paper uses MediaPipe for face detection and keypoint annotation, as well as face alignment and cropping. To address the varying lengths of micro-expression video clips, temporal interpolation is employed to interpolate sequences to a specified length of 20 frames. Finally, during training, data augmentation is performed using mirror flipping, random angle rotation, and scaling to prevent overfitting and mitigate the issues of small sample sizes and severe class imbalance.

[0039] S2: Split the preprocessed data into training set, validation set and test set randomly in the ratio of 8:1:1.

[0040] S3: Use ResNet50 and the multi-level and multi-scale terraced feature extraction module to perform intra-frame feature extraction to achieve feature extraction of the main and secondary occurrence areas of micro-expressions.

[0041] The multi-level, multi-scale terraced feature extraction module uses a multi-level, multi-scale feature extraction approach to enhance the model's perception of key and secondary areas of micro-expressions while retaining richer detailed information. Specifically, multi-level feature extraction uses three parallel 2D convolutions to capture features with different receptive fields. This parallel approach reduces information loss caused by serial convolutions. Each convolution is followed by a batch normalization operation to stabilize the data distribution within a certain range, alleviating internal covariate shift, accelerating model convergence, and improving generalization. Because convolution parameter errors can lead to feature estimation bias, the maximum pooling layer retains significant facial texture information while reducing dimensionality, further reducing offset errors and enhancing feature robustness. Specifically for multi-scale feature extraction, the network uses a cross-layer connection-based method to optimize the underlying feature maps and obtain feature representations at different scales. Specifically, four convolutions with a stride of 2 are performed as a way to increase the scale, obtaining feature sequences at five scales. The high-level features with rich semantic information are then upsampled and fused with the features of the previous layer as the feature representation of the current scale. This method enables the model to obtain representative semantic information from the upsampled features and combine it with the underlying feature maps to integrate finer detail information, reducing the loss of detail caused by convolution. The TMMFE module achieves precise positioning of key areas of micro-expressions and multi-level feature expression through multi-level parallel convolution and multi-scale cross-layer fusion. Experiments show that this module effectively improves the model's attention to the main and secondary features of micro-expressions, laying a better feature foundation for subsequent time series modeling and classification tasks.

[0042] S4: Use Bi-LSTM network and spatiotemporal genetic attention module to extract spatial features between frames and capture the subtle dynamic changes of micro-expressions between frames.

[0043] The network uses spatiotemporal genetic attention to process small movements between frames of video sequences. Six micro-motion attention modules are used. To avoid the performance degradation caused by deep structures, a residual connection structure is introduced before each micro-motion attention module. This structure effectively alleviates the problem of gradient vanishing or information loss in deep networks by adding the module input to the residual output of the convolution branch. Specifically, for the i-th residual structure, its input is recorded as , after 1×1 convolution and 3×3 convolution After the operation, the output of the residual structure is obtained In order to construct a more comprehensive spatial attention map, each incremental attention module combines the maximum pooling and average pooling strategies when generating the attention map, which are used to extract significant area features and global context features respectively. Through the splicing operation, combined with the 1×1 convolution layer and the Mish activation function, an attention map is generated that can guide the network to focus on subtle facial motion areas. Next, the currently generated attention map is multiplied point by point with the attention map of the previous layer to achieve spatial attention weighting of the current layer feature map and introduce temporal consistency in attention propagation. This strategy of introducing the previous layer attention map can not only dynamically update the attention distribution, but also spatially weight the previous attention map, so that the network's attention to the motion change area is more accurate. It is worth noting that the first incremental attention module does not introduce the previous attention map when generating the attention map. The formula for the generated attention map is:

[0044]

[0045] In the above formula, is the input of the first micro-movement attention module; is the Mish function; and They are maximum pooling and average pooling respectively. Other modules introduce the front layer attention map when generating the attention map, and the attention map is

[0046]

[0047] in, Denotes an element-by-element dot product operation. By element-by-element multiplication of the generated attention map with the previous frame's attention map, we obtain a feature map capable of perceiving subtle local motion. This multi-frame concatenation mechanism enables the network, guided by the dynamic updates of the micro-motion attention map, to gradually focus on regions with subtle motion changes. This robustly extracts and learns motion features from micro-expression images, achieving more accurate and robust feature representation.

[0048] S5: The entire network structure was trained for 100 rounds to construct a micro-expression emotion recognition model, which integrates the efficient micro-expression emotion recognition method TM-SGANet;

[0049] The model was trained on an NVIDIA GeForce RTX 4090, using Python 3.9 and the PyTorch framework to build the neural network model. For training parameters, SGD was used as the optimizer; the momentum factor was set to 0.9; the learning rate was set to 1x10⁻4, with dynamic adjustment, using a cosine annealing strategy; the dropout rate was set to 0.4 in the dropout layer and 0.5 in the fully connected layer, and the cross-entropy loss function was used.

[0050] S6: Comparative analysis and optimization of the detection performance of TM-SGANet with other micro-expression models.

[0051] The same dataset and number of classification categories are used, consistent evaluation criteria are adopted, and the "leave one subject out" method is used for cross-validation to ensure the consistency of the experimental benchmark, evaluation criteria, and verification methods. The leave-one-subject method uses the data of one subject as the test set each time, and the data of all other subjects as the training set. Finally, the average of N test results is taken as the final performance indicator. It can effectively evaluate the generalization ability of the micro-expression recognition model on unseen individuals and avoid data leakage and overfitting problems. According to experimental data, our network model method performed best on the CASME II and SAMM datasets, achieving a UF1 score of 91.23% and a UAR score of 89.68% on the CASME II, and a UF1 score of 90.46% and a UAR score of 88.26% on the SAMM dataset. In CAS(ME) 3 It achieved superior performance on the dataset, far surpassing existing experimental methods. The accuracy of the five emotions reached 64.89%, far exceeding the baseline method.

[0052] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An efficient micro-expression emotion recognition method, comprising a multi-level and multi-scale terrace feature extraction module and a spatiotemporal genetic attention module, characterized in that: The steps include: S1: SAMM, CASME II and CAS(ME) 3 The dataset is preprocessed to generate data for training; S2: Split the preprocessed data into training set, validation set and test set; S3: Use ResNet50 and the multi-level multi-scale terrace feature extraction module to extract intra-frame features, realizing feature extraction of the main and secondary occurrence areas of micro-expressions; S4: Use Bi-LSTM network and spatiotemporal genetic attention module to extract inter-frame features and capture the subtle dynamic changes of micro-expressions between frames; S5: The entire network structure was trained for 100 rounds to construct a micro-expression emotion recognition model, which integrates the efficient micro-expression emotion recognition method TM-SGANet; S6: Comparative analysis and optimization of the detection performance of TM-SGANet with other micro-expression models.

2. The efficient micro-expression emotion recognition method according to claim 1, characterized in that: The S1 is in SAMM, CASME II and CAS(ME) 3 During data preprocessing, MediaPipe was used for face detection and key point annotation, and faces were aligned and cropped. Temporal interpolation was then used to lengthen the sequence to a specified length of 20 frames. Finally, mirror flipping, random angle rotation, and scaling were performed to enhance the data to prevent overfitting and alleviate the problems of small sample size and severe class unevenness.

3. The efficient micro-expression emotion recognition method according to claim 1, characterized in that: The pre-processed data in S2 are randomly divided into a training set, a validation set, and a test set in a ratio of 8:1:

1.

4. The efficient micro-expression emotion recognition method according to claim 1, characterized in that: In the S3, ResNet50 and the multi-level multi-scale terrace feature extraction module are used to extract intra-frame spatial features. We introduce the multi-level multi-scale terrace feature extraction module in the last four stages of ResNet50. The multi-level and multi-scale terrace feature extraction module enhances the model's perception of key areas and secondary areas of micro-expressions, while retaining richer detail information. For multi-level feature extraction, three parallel 2D convolutions are used to capture different receptive field features. The parallel approach reduces the information loss caused by serial convolution. Each convolution is followed by a batch normalization operation to stabilize the data distribution within a certain range, alleviate internal covariate shifts, accelerate model convergence and improve generalization capabilities. Since convolution parameter errors may lead to feature estimation deviations, the maximum pooling layer retains significant facial texture information while reducing dimensionality, reduces offset errors, and enhances feature robustness. For multi-scale feature extraction, the network adopts A method based on cross-layer connection was used to optimize the underlying feature maps and obtain feature representations of different scales. Specifically, four convolutions were performed with a step size of 2 as a way to increase the scale, and feature sequences of five scales were obtained. Then, the high-level features with rich semantic information were upsampled and fused with the features of the previous layer as the feature representation of the current scale, so that the model can obtain representative semantic information from the upsampled features and combine it with the underlying feature maps to obtain finer detail information, reducing the detail loss caused by convolution. The multi-level and multi-scale terraced feature extraction module achieves precise positioning of key areas of micro-expressions and multi-level feature expression through multi-level parallel convolution and multi-scale cross-layer fusion.

5. The efficient micro-expression emotion recognition method according to claim 1, characterized in that: The S4 uses Bi-LSTM network and spatiotemporal genetic attention module to extract inter-frame features, which can better handle small movements between frames of video sequences. The method adopts 6 micro-motion attention modules. In order to avoid the performance degradation problem caused by deep structure, a residual connection structure is introduced before each micro-motion attention module. This structure effectively alleviates the problem of gradient disappearance or information loss in deep network by adding the module input and the residual output of the convolution branch. Specifically, for the i-th residual structure, its input is recorded as , after 1×1 convolution and 3×3 convolution After the operation, the output of the residual structure is obtained In order to construct a more comprehensive spatial attention map, each incremental attention module combines the maximum pooling and average pooling strategies when generating the attention map, which are used to extract significant area features and global context features respectively. Through the splicing operation, combined with the 1×1 convolution layer and the Mish activation function, an attention map is generated that can guide the network to focus on subtle facial motion areas. Next, the currently generated attention map is multiplied point by point with the attention map of the previous layer to achieve spatial attention weighting of the feature map of the current layer and introduce temporal consistency in the attention propagation. This strategy of introducing the previous layer attention map can not only dynamically update the attention distribution, but also spatially weight the previous attention map, making the network's attention to the motion change area more accurate. It is worth noting that the first incremental attention module does not introduce the previous attention map when generating the attention map. The generated attention map formula is: In the above formula, is the input of the first micro-movement attention module; is the Mish function; and They are maximum pooling and average pooling respectively. Other modules introduce the front-layer attention map when generating the attention map, and the obtained attention map is: in, Represents the element-by-element dot product operation. After multiplying the generated attention map with the previous frame attention map element-by-element, we can obtain a feature map that can perceive local subtle movements. This multi-frame concatenation mechanism enables the network to gradually focus on areas with subtle motion changes under the guidance of the dynamic update of the micro-motion attention map, thereby robustly extracting and learning motion features in micro-expression images and achieving more accurate and robust feature expression.

6. The efficient micro-expression emotion recognition method according to claim 1, characterized in that: The S5 includes the following steps: using NVIDIA GeForce RTX 4090, using Python 3.9 language, and PyTorch framework to build a neural network model. In terms of training parameters, SGD is used as the optimizer; the momentum factor is set to (0.9); the learning rate is set to 1x10-4 and dynamically adjusted, using a cosine annealing strategy; the dropout rate in the dropout layer is set to 0.4, and in the fully connected layer it is set to 0.5, using the cross entropy loss function. In CASME II, SAMM and CAS (ME) 3 100 rounds of training were performed on the dataset.

7. The efficient micro-expression emotion recognition method according to claim 1, characterized in that: The S6 experimental results analysis used the same dataset and number of classification categories, consistent evaluation criteria, and a "leave-one-subject" cross-validation method to ensure consistency in the experimental benchmark, evaluation criteria, and validation methods. This leave-one-subject method uses the data of one subject as the test set and the data of all other subjects as the training set. The average of N test results is then used as the final performance metric. This effectively evaluates the generalization ability of the micro-expression recognition model on unseen individuals, avoiding data leakage and overfitting.