Dynamic and static two-way interactive and collaborative micro-expression recognition method
Through the feature fusion and interactive attention module of dynamic branches and static branches, the noise interference and data scarcity problems of micro-expression feature extraction in complex classroom scenarios are solved, and accurate micro-expression recognition in real scenarios is achieved.
Patent Information
- Application Number
- CN202510653914.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively extract micro-expression features in complex classroom scenarios, especially to suppress light changes and head motion noise interference. In addition, deep learning methods perform poorly in real scenarios, making it difficult to achieve efficient feature fusion.
The dual attention guide feature fusion module in dynamic branches and the gated cross-layer feature delivery mechanism are adopted, combined with the hybrid attention Transformer block and token of the static branch, the Transformer block is redistributed. The dynamic static features are fused through the two-way interactive attention module to achieve accurate recognition of micro-expressions.
It enhances the sensitivity to weak micro-expression motion signals, suppresses noise interference, avoids loss of deep network details, alleviates overfitting caused by data scarcity, and realizes accurate micro-expression recognition in complex scenarios.
Smart Images

Figure CN120388410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image recognition technology, in particular to micro-expression recognition technology in image recognition, and specifically to a dynamic and static two-way interactive collaborative micro-expression recognition method. Background Art
[0002] Microexpressions are spontaneous, brief, and subtle facial movements lasting between 1 / 25 and 1 / 5 of a second, typically reflecting an individual's true emotions. In classroom settings, teachers can use microexpressions to understand students' true emotional states and, in turn, gauge their understanding of the content. This allows them to adjust their teaching strategies and pacing, improving teaching effectiveness. Therefore, microexpression recognition plays a crucial role in education. However, in complex classroom settings, students' microexpressions are both more natural and more complex. While some researchers have used hybrid deep convolutional networks to extract optical flow features from students' microexpressions or employed multi-network architectures to optimize microexpression feature extraction, these efforts have struggled to effectively address the noise interference and degradation of dynamic microexpression features caused by head movement and illumination changes in classroom environments. This has limited the performance of microexpression recognition in classroom settings. Therefore, effectively extracting microexpression features remains a key issue in microexpression recognition research.
[0003] To effectively extract micro-expression features, early research primarily relied on manual feature extraction methods. While these methods have gradually improved the performance of micro-expression recognition, they still face challenges such as a strong reliance on domain expertise, cumbersome parameter optimization, high computational complexity, and limitations in generalization and robustness. Consequently, research in the field of micro-expression recognition has gradually shifted toward a deep learning paradigm. However, most deep learning methods are limited to a single input, such as optical flow or a sequence of micro-expression frames, and the backbone feature extraction networks typically use ResNet, Vision Transformer, or hybrid architectures. While these methods perform well in controlled experimental environments, they face a series of challenges in complex real-world scenarios. First, optical flow features are susceptible to interference from lighting changes and can easily introduce head motion noise. Second, while micro-expression sequence frames can capture temporal information, they also contain a large amount of redundant data. Furthermore, the residual connections of ResNet can lead to noise accumulation, obscuring important subtle motion features. Meanwhile, the single-scale self-attention mechanism of the Vision Transformer is limited in its efficiency in capturing local details and has high computational complexity. Furthermore, existing multi-stream networks often use fusion strategies such as simple feature concatenation or weighted summation, making it difficult to effectively fuse multiple features. Therefore, efficient input representation, optimized network architecture, and the design of effective feature fusion strategies remain core issues that need to be addressed in the field of micro-expression recognition. Summary of the invention
[0004] The object of the present invention is to provide a micro-expression recognition method with dynamic and static two-way interactive collaboration in view of the deficiencies of the prior art. This method can enhance the sensitivity to weak micro-expression motion signals and suppress noise interference, while avoiding the loss of micro-expression details in deep networks, and can accurately extract multi-granularity features while alleviating overfitting caused by data scarcity, so as to achieve accurate recognition of micro-expressions.
[0005] The technical solution for realizing the object of the present invention is as follows:
[0006] A micro-expression recognition method with dynamic and static two-way interactive collaboration, comprising the following steps:
[0007] 1) Quantify the motion features of micro-expressions: Use the optical flow method and the frame difference algorithm to calculate the optical flow F flow and the pixel difference F diff between the starting frame and the vertex frame of the micro-expression respectively;
[0008] 2) Accurately model the subtle muscle motion features of micro-expressions: Input the optical flow F flow and the pixel difference F diff into the dynamic branch, and adopt the dual attention guided feature fusion module DAGFFM (Dual Attention Guided Feature Fusion Module, hereinafter referred to as DAGFFM) and the gated cross-layer feature transfer mechanism GCFTM (Gated Cross-layer Feature Transfer Mechanism, hereinafter referred to as GCFTM), including:
[0009] 2-1) Construct a dual attention guided feature fusion module: Adopt the dual attention mechanism to highlight the significant regions and key channels of the micro-expression motion features and fuse them, including:
[0010] 2-1-1) Respectively perform 1×1 convolution operations on the optical flow map F flow and the pixel difference map F diff to obtain feature maps and respectively, and then perform 3×3 convolution operations to obtain feature maps and respectively. Concatenate with with respectively in the channel dimension to obtain X 12 and X 21 ;
[0011] 2-1-2) Highlight the significant regions of micro-expression motion: Take X 21Through the spatial attention mechanism SA (Spatial Attention, abbreviated as SA), specifically: First, X 21 Pass through the average pooling layer and the max pooling layer, and concatenate the output of the average pooling layer with the output of the max pooling layer. Then, successively pass through a 7×7 convolutional layer with a stride of 1 and a Sigmoid activation function to generate the spatial attention weight map W 21 , and finally multiply the attention weight map W 21 element-wise with the original input feature Figure X 21 to obtain the weighted spatial attention feature map A 21 , as shown in formula (1):
[0012] W 21 = Sigmoid(f 7×7 (Concat(MaxPool(AvgPool(X 21 )), AvgPool(X 21 )))) (1),
[0013]
[0014] In formula (1), MaxPool(·) represents the max pooling operation, AvgPool(·) represents the average pooling operation, Concat(·) represents the concatenation operation, f 7×7 (·) represents the 7×7 convolutional operation, Sigmoid(·) represents the Sigmoid activation function, represents element-wise multiplication;
[0015] 2-1-3) Highlight the key channels of micro-expression movements: Pass X 12 through the dual-path channel attention mechanism DPCA (Dual-Patch Channel Attention, abbreviated as DPCA) consisting of two branches. Specifically: In the max pooling branch, after performing the max pooling operation on X 12 , pass through a fully connected layer to learn the non-linear relationship between channels to obtain the attention weight X' 12 of this branch. In the average pooling branch, obtain the attention weight X'' 12 of this branch according to the operation in formula (2):
[0016] X'' 12 = δ(f 1×1 (AvgPool(δ(β(X 12 ))))) (2),
[0017] In formula (2), β(·) represents channel normalization, δ(·) represents the ReLU activation function, f 1×1Denote the 1×1 convolution operation;
[0018] 2-1-4) Add X′ 12 and X″ 12 and then activate the result using the Sigmoid function to obtain the attention weight W output by the dual-channel attention mechanism 12 , and multiply it element-wise with X 12 to obtain the channel attention feature map A 12 ;
[0019] 2-1-5) Average split the feature map A 12 into two parts along the channel dimension to obtain the split features and Similarly, split the feature map A 21 into features and feature and obtain the micro-expression motion feature F M :
[0020]
[0021] In formula (3), ⊕ represents the addition operation;
[0022] 2-2) Construct the gated cross-layer feature transfer mechanism: The gated cross-layer feature transfer mechanism consists of 8 convolutional layers of size 3×3 and 5 feature importance gating units FGIU (Feature Importance Gating Unit, abbreviated as FGIU), including:
[0023] 2-2-1) The gated cross-layer feature transfer mechanism performs two consecutive 3×3 convolution operations on the motion feature F M . The stride of the first convolution is 2, which downsamples the feature map and doubles the number of channels at the same time; the stride of the second convolution is 1 to obtain the feature map L1;
[0024] 2-2-2) The first feature importance gating unit groups the channels of L1, and then performs a normalization operation on each group according to formula (4) to obtain the normalized feature L gn :
[0025]
[0026] In formula (4), GN(·) represents the group normalization operation, μ and σ are the mean and standard deviation of L1, ε is a small positive constant added to ensure the stability of division, γ and β are trainable affine transformations, where γ is used to measure the spatial pixel variance of each channel. When the spatial information of the micro-expression feature is richer, the value of γ is larger. Normalize γ to obtain the weight matrix W of different feature maps γ ;
[0027] 2-2-3) Screen key motion features: Multiply the weight matrix W γ element-wise with L gn to obtain a weighted feature map, and use the Sigmoid function to map it to the range (0, 1), then perform gating through the set threshold T to obtain the importance weights and multiply them element-wise with the input feature L1. The specific process is shown in Equation (5):
[0028]
[0029] In Equation (5), G1 represents the screened key motion features, and Gate(·) represents the gating operation;
[0030] 2-2-4) Retain the original complete information and enhance the response of the key region: Jump L1 to add the output of the importance feature gating unit and G1 to obtain the significant feature F1;
[0031] 2-2-5) Establish a cross-layer feature transfer path: After selecting the key motion features from the shallow layer, directly inject them into the deep network after adaptive pooling alignment. The specific process is shown in Equation (6):
[0032]
[0033] In Equation (6), i = 2, 3, 4, f 3×3 (·) represents a 3×3 convolution operation with a stride of 1, represents a 3×3 convolution operation with a stride of 2, represents the feature map after passing through two 3×3 convolution layers, GU i (·) represents the i-th FIGU screening operation, represents the key motion features after passing through the i-th FIGU screening, represents the feature map obtained by pooling G i after pooling, represents the updated significant feature;
[0034] 2-2-6) Further screen the significant feature F4 through the fifth importance feature gating unit to obtain the key feature G5, and add it to the key feature G4 screened by the fourth importance feature gating unit to obtain the dynamic feature F D output by the dynamic branch finally, as shown in Equation (7):
[0035]
[0036] 3) Precise extraction of multi-granularity static features of original micro-expression vertex frames: Construct a static branch composed of a patch embedding layer, four alternately stacked hybrid attention Transformer blocks, and a token reassignment Transformer block to extract multi-granularity static features of original micro-expression vertex frames, including:
[0037] 3-1) The patch embedding layer divides the vertex frame F with a size of 192×192 apex into 36 patches, with an embedding dimension of 512, to obtain the embedded feature X;
[0038] 3-2) Input the feature X into the hybrid attention Transformer block to extract local fine-grained and global coarse-grained information, and generate a micro-expression saliency map: The self-attention mechanism of the hybrid attention Transformer block evenly divides the number of attention heads into two groups, with each group containing 2 heads. Among them, the first group uses global attention, and the second group uses local attention, including:
[0039] 3-2-1) Global attention calculation: First, perform a 6×6 convolution operation on the input feature X to compress it into a single feature representation, and then calculate self-attention. The calculation process is shown in formula (8):
[0040]
[0041] In formula (8), Q1, K1, and V1 respectively represent the query, key, and value matrices generated using global attention; SR(X, 6×6) represents the compression operation of using a 6×6 convolution operation on X; W1 Q , W1 K , W1 V is the projection matrix; atten1 represents the weight matrix generated by global attention; d is the dimension of the attention head; O1 represents the feature map generated by global attention;
[0042] 3-2-2) Local attention calculation: X is projected into the query, key, and value matrices Q2, K2, and V2 respectively, and then K2 and V2 are divided using a fixed-size window M to make the window division of Q2 match that of K2 and V2. The window size of Q2 is set to sM, where s represents the scaling factor of the receptive field, and then self-attention is calculated within the window, as shown in formula (9):
[0043]
[0044] In formula (9), Window(·) represents the local window partitioning operation. Q′2, K′2, and V′2 respectively represent the query, key, and value matrices generated by local attention. atten2 represents the weight matrix generated by local attention, and O2 represents the feature map generated by local attention.
[0045] 3-2-3) Generate the saliency map: Concatenate the feature maps O1 and O2 generated by global attention and local attention to obtain the feature map O generated by hybrid attention. At the same time, perform an average operation on the weight matrices atten1 and atten2 generated by global attention and local attention and then aggregate them to obtain the saliency map S. The process is shown in formula (10):
[0046] O = Concat(O1, O2)
[0047]
[0048] In formula (10), represents the attention score of the i-th head of global attention at position p, represents the attention score of the j-th head of local attention at position q;
[0049] 3-3) Input the feature map O and the saliency map S into the token reassignment Transformer block, including:
[0050] 3-3-1) Sort the importance weights of the saliency map S in ascending order and divide them into 3 sub-regions S 1 , S 2 , S 3 , where S 1 is the least important region, and S 3 is the most important region. At the same time, all tokens in the feature map O are divided into O 1 , O 2 , O 3 corresponding to S 1 , O 2 , O 3 , and then use different aggregation rates r 1 , r 2 , r 3 to aggregate the tokens of O 1 , O 2 , O 3 in different saliency regions to obtain the aggregated where r represents the aggregation rate at which every r tokens are merged into one token. The more important the region, the smaller the aggregation rate. Finally, concatenate to obtain the reassigned token sequence. The process is shown in formula (11):
[0051]
[0052] In formula (11), F(·) represents the aggregation function, which is implemented by a fully connected layer with an input dimension of r and an output dimension of 1. F(O i ,r i ) represents the token sequence of the i-th group O i aggregated at an aggregation rate of r i ; represents the token sequence after reallocation;
[0053] 3-3-2) Project the token sequence after reallocation onto the key matrix K and the value matrix V, while the original token sequence is projected onto the query matrix Q. The specific process of calculating the attention is shown in formula (12):
[0054]
[0055] In formula (12), atten represents the attention weights generated by the token reallocation Transformer block, and R represents the feature map generated by the token reallocation Transformer block;
[0056] 3-3-3) Update the saliency map and the feature map according to steps 3-2-1) to 3-2-3) for the feature map R generated by the token reallocation Transformer block. The updated saliency map and feature map then go through steps 3-3-1) to 3-3-2) to obtain the multi-granularity static feature F S ;
[0057] 4) Bidirectional interactive fusion of dynamic and static features: Construct a bidirectional interactive attention fusion module BIAFM (Bidirectional Interactive Attention Fusion Module, abbreviated as BIAFM) to fuse the dynamic feature F D from the dynamic branch and the multi-granularity static feature F S from the static branch, including:
[0058] 4-1) Project the dynamic feature F D onto the query matrix Q D , the key-value matrix K D and the value matrix V D , as shown in formula (13):
[0059]
[0060] In formula (13), is the projection matrix;
[0061] 4 - 2) Project the static feature F S onto the query matrix Q S , key matrix K S and value matrix V S , as shown in formula (14):
[0062]
[0063] In formula (14), and are projection matrices;
[0064] 4 - 3) Calculate the attention weights of the dynamic feature, where the query matrix comes from the static feature, and then multiply it with the value matrix V D , as shown in formula (15):
[0065]
[0066] F′ D = Atten D V D
[0067] In formula (15), is the scaling factor, Atten D is the attention weight matrix of F D , and F′ D is the feature map of F D after being calculated by the attention mechanism;
[0068] 4 - 4) Calculate the attention weights of the static feature, where the query matrix comes from the dynamic feature, and then multiply it with the value matrix V S , as shown in formula (16):
[0069]
[0070] F′ S = Atten S V S
[0071] In formula (16), Atten S is the attention weight matrix of F S , and F′ S is the feature map of F S after being calculated by the attention mechanism;
[0072] 4 - 5) Add F′ D and F′ S , retain the main feature information, and then obtain the final fused feature F fusion :
[0073] F fusion = AvgPool(F' D ⊕ F' S )(17),
[0074] In formula (17), AvgPool(·) represents the average pooling operation;
[0075] 5) Input F fusion into the fully connected layer to achieve accurate classification of micro - expressions.
[0076] This technical solution first uses a dynamic branch to suppress noise interference, which can avoid the loss of details in the deep network and enhance the model's ability to model subtle movements; then uses a static branch to accurately extract multi - granular static features while alleviating overfitting caused by data scarcity and reducing computational complexity; finally, uses a bidirectional interactive attention fusion module to deeply interactively fuse the dynamic branch and the static branch, realizing complementary enhancement and collaborative optimization of dynamic and static features, thereby achieving accurate recognition of micro - expressions.
[0077] Compared with the existing technology, this technical solution has the following characteristics:
[0078] 1) Constructed a dynamic branch, in which a feature fusion module guided by dual attention and a gated cross - layer feature transfer mechanism are innovatively designed, effectively fusing the optical flow and pixel differences between the starting frame and the vertex frame, and suppressing noise interference and alleviating the problem of loss of motion details caused by continuous downsampling of the network;
[0079] 2) Constructed a static branch, in which a hybrid attention Transformer block is innovatively designed to generate a saliency map of the micro - expression vertex frame and a token re - allocation Transformer block guided by the saliency map, accurately extracting multi - granular static features while alleviating overfitting caused by data scarcity;
[0080] 3) Designed a bidirectional interactive attention fusion module, which effectively fuses the key motion features of the dynamic branch and the spatial multi - granular information of the static branch through bidirectional information flow and deep interaction between dynamic and static features, promoting complementary enhancement and collaborative optimization of dynamic and static features;
[0081] 4) This technical solution has high robustness and effectiveness in complex classroom scenarios, providing a new technical means for the fields of intelligent education and emotion analysis.
[0082] This method can enhance the sensitivity to weak micro - expression motion signals and suppress noise interference, while avoiding the loss of micro - expression details in the deep network, and can accurately extract multi - granular features while alleviating overfitting caused by data scarcity, thereby achieving accurate recognition of micro - expressions. Description of the Drawings
[0083] Figure 1 Schematic diagram of the process of the embodiment method;
[0084] Figure 2 Schematic diagram of the DSBICNet model framework in the embodiment;
[0085] Figure 3 Schematic diagram of the structure of the dual-attention-guided feature fusion module in the embodiment;
[0086] Figure 4 Schematic diagram of the structure of the importance feature gating unit in the embodiment;
[0087] Figure 5 Schematic diagram of the structure of the bidirectional interactive attention fusion module in the embodiment;
[0088] Figure 6 Schematic diagram of the comparison of the unweighted F1 scores (UF1) of CapsuleNet, RCN-A, MMNet, HTNet, HFA-Net, and DSBICNet on the Full dataset in the embodiment;
[0089] Figure 7 Schematic diagram of the comparison of the unweighted F1 scores (UF1) of CapsuleNet, RCN-A, MMNet, HTNet, HFA-Net, and DSBICNet on the SMIC dataset in the embodiment;
[0090] Figure 8 Schematic diagram of the comparison of the unweighted F1 scores (UF1) of CapsuleNet, RCN-A, MMNet, HTNet, HFA-Net, and DSBICNet on the CASME II dataset in the embodiment;
[0091] Figure 9 Schematic diagram of the comparison of the unweighted F1 scores (UF1) of CapsuleNet, RCN-A, MMNet, HTNet, HFA-Net, and DSBICNet on the SAMM dataset in the embodiment;
[0092] Figure 10 Schematic diagram of the comparison of the unweighted F1 scores (UF1) of CapsuleNet, RCN-A, MMNet, HTNet, HFA-Net, and DSBICNet on the GUCME dataset in the embodiment. Detailed implementation manners
[0093] The content of the present invention will be further described below in conjunction with the accompanying drawings and embodiments, but it is not a limitation to the present invention.
[0094] Embodiment:
[0095] Refer to Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 A dynamic and static two-way interactive collaborative micro-expression recognition method, comprising the following steps:
[0096] 1) Quantify the motion features of micro-expressions: Use the optical flow method and the frame difference algorithm to calculate the optical flow F flow and pixel difference F diff ;
[0097] 2) Accurately model the subtle muscle motion features of micro-expressions: Input the optical flow F flow and pixel difference F diff into the dynamic branch, and adopt the dual attention-guided feature fusion module DAGFFM and the gated cross-layer feature transfer mechanism GCFTM, including:
[0098] 2-1) Construct a dual attention-guided feature fusion module: Adopt the dual attention mechanism to highlight the significant regions and key channels of the micro-expression motion features and fuse them, including:
[0099] 2-1-1) Perform 1×1 convolution operations on the optical flow map F flow and the pixel difference map F diff respectively to obtain feature maps and Then perform 3×3 convolution operations on them respectively to obtain feature maps and Concatenate with with on the channel dimension respectively to obtain X 12 and X 21 ;
[0100] 2-1-2) Highlight the significant regions of micro-expression motion: Pass X 21 through the spatial attention mechanism SA, specifically: First, pass X 21 through the average pooling layer and the maximum pooling layer, and concatenate the output of the average pooling layer and the output of the maximum pooling layer, then successively pass through a 7×7 convolutional layer with a stride of 1 and a Sigmoid activation function to generate the spatial attention weight map W 21 , and finally multiply the attention weight map W 21 element-wise with the original input feature Figure X 21 to obtain the weighted spatial attention feature map A 21 , as shown in formula (1):
[0101] W 21 = Sigmoid(f 7×7 (Concat(MaxPool(AvgPool(X 21 )), AvgPool(X 21 )))) (1),
[0102]
[0103] In formula (1), MaxPool(·) represents the max pooling operation, AvgPool(·) represents the average pooling operation, Concat(·) represents the concatenation operation, f 7×7 (·) represents the 7×7 convolution operation, Sigmoid(·) represents the Sigmoid activation function, represents element-wise multiplication;
[0104] 2-1-3) Highlight the key channels of micro-expression movement: Pass X 12 through the dual-path channel attention mechanism DPCA composed of two branches. Specifically: In the max pooling branch, after performing the max pooling operation on X 12 and passing through the fully connected layer to learn the non-linear relationship between channels, obtain the attention weight X' 12 of this branch. In the average pooling branch, obtain the attention weight X'' 12 of this branch according to the operation of formula (2):
[0105] X'' 12 = δ(f 1×1 (AvgPool(δ(β(X 12 ))))) (2),
[0106] In formula (2), β(·) represents channel normalization, δ(·) represents the ReLU activation function, f 1×1 represents the 1×1 convolution operation;
[0107] 2-1-4) Add X' 12 and X'' 12 and then activate using the Sigmoid function to obtain the attention weight W 12 output by the dual-path channel attention mechanism, and perform element-wise multiplication with X 12 to obtain the channel attention feature map A 12 ;
[0108] 2-1-5) Average split the feature map A 12 into two parts in the channel dimension to obtain the split features and Similarly, split the feature map A 21 into features and features and obtain the micro-expression motion feature F according to the operation of formula (3) M :[[]]
[0109]
[0110] In formula (3), ⊕ represents the addition operation;
[0111] 2-2) Construct a gated cross-layer feature transfer mechanism: The gated cross-layer feature transfer mechanism consists of 8 convolutional layers of size 3×3 and 5 importance feature gating units FGIU, including:
[0112] 2-2-1) The gated cross-layer feature transfer mechanism performs two consecutive 3×3 convolutional operations on the motion feature F M where the stride of the first convolution is 2, downsampling the feature map is completed, and the number of channels is doubled at the same time; the stride of the second convolution is 1 to obtain the feature map L1;
[0113] 2-2-2) The first importance feature gating unit performs channel grouping on L1, and then performs a normalization operation on each group according to formula (4) to obtain the normalized feature L gn :[[]]
[0114]
[0115] In formula (4), GN(·) represents the group normalization operation, μ and σ are the mean and standard deviation of L1, ε is a small normal constant added to ensure the stability of division, γ and β are trainable affine transformations, where γ is used to measure the spatial pixel variance of each channel. When the spatial information of the micro-expression feature is richer, the value of γ is larger. Normalize γ to obtain the weight matrix W of different feature maps γ ;
[0116] 2-2-3) Screen key motion features: Multiply the weight matrix W γ element-wise with L gn to obtain a weighted feature map, and use the Sigmoid function to map it to the range (0,1), then perform gating through the set threshold T to obtain the importance weight and multiply it element-wise with the input feature L1. The specific process is shown in formula (5):
[0117]
[0118] In formula (5), G1 represents the screened key motion feature, and Gate(·) represents the gating operation;
[0119] 2-2-4) Preserve the original complete information and enhance the response of key regions: Connect L1 to the output of the importance feature gating unit and add it to G1 to obtain the significant feature F1;
[0120] 2-2-5) Establish a cross-layer feature transfer path: After selecting key motion features from the shallow layer, they are directly injected into the deep network after being aligned by adaptive pooling. The specific process is shown in Equation (6):
[0121]
[0122] In Equation (6), i = 2, 3, 4, f 3×3 (·) represents a 3×3 convolution operation with a stride of 1, represents a 3×3 convolution operation with a stride of 2, represents the feature map after passing through two 3×3 convolution layers, GU i (·) represents the i-th FIGU screening operation, represents the key motion feature after being screened by the i-th FIGU, represents G i The pooled feature map, represents the updated significant feature;
[0123] 2-2-6) Further screen the significant feature F4 through the fifth importance feature gating unit to obtain the key feature G5, and add it to the key feature G4 screened by the fourth importance feature gating unit to obtain the dynamic feature F output by the final dynamic branch D , as shown in Equation (7):
[0124]
[0125] 3) Precisely extract the multi-granularity static features of the original micro-expression vertex frame: Construct a static branch composed of a patch embedding layer, four alternately stacked hybrid attention Transformer blocks, and a token reallocation Transformer block to extract the multi-granularity static features of the original micro-expression vertex frame, including:
[0126] 3-1) The patch embedding layer divides the vertex frame F with a size of 192×192 apex into 36 patches, with an embedding dimension of 512, to obtain the embedded feature X;
[0127] 3-2) Input the feature X into the hybrid attention Transformer block to extract local fine-grained and global coarse-grained information, and generate a micro-expression saliency map: The self-attention mechanism of the hybrid attention Transformer block evenly divides the number of attention heads into two groups, with each group containing 2 heads. Among them, the first group uses global attention, and the second group uses local attention, including:
[0128] 3-2-1) Global attention calculation: First, perform a 6×6 convolutional operation on the input feature X to compress it into a single feature representation, and then calculate self-attention. The calculation process is shown in formula (8):
[0129]
[0130] In formula (8), Q1, K1, and V1 respectively represent the query, key, and value matrices generated using global attention; SR(X, 6×6) represents the compression operation of using a 6×6 convolutional operation on X; W1 Q , W1 K , W1 V is the projection matrix; atten1 represents the weight matrix generated by global attention; d is the dimension of the attention head; O1 represents the feature map generated by global attention;
[0131] 3-2-2) Local attention calculation: X is projected into the query, key, and value matrices Q2, K2, and V2 respectively. Then, the fixed-size window M is used to divide K2 and V2 to make the window division of Q2 match that of K2 and V2. The window size of Q2 is set to sM, where s represents the scaling factor of the receptive field. Then, calculate self-attention within the window, as shown in formula (9):
[0132]
[0133] In formula (9), Window(·) represents the local window division operation, Q′2, K′2, and V′2 respectively represent the query, key, and value matrices generated using local attention, atten2 represents the weight matrix generated by local attention, and O2 represents the feature map generated by local attention;
[0134] 3-2-3) Generate the saliency map: Concatenate the feature maps O1 and O2 generated by global attention and local attention to obtain the feature map O generated by hybrid attention. At the same time, perform an average operation on the weight matrices atten1 and atten2 generated by global attention and local attention and then aggregate them to obtain the saliency map S. The process is shown in formula (10):
[0135] O = Concat(O1, O2)
[0136]
[0137] In formula (10), represents the attention score of the i-th global attention head at position p, represents the attention score of the j-th local attention head at position q;
[0138] 3-3) The feature map O and the saliency map S are input into the token and redistributed into the Transformer block, including:
[0139] 3-3-1) Sort the importance weights of the saliency map S in ascending order and divide it into three sub-regions S 1 ,S 2 ,S 3 , where S 1 is the least important area, S 3 is the most important area, and all tokens in the feature map O are sorted according to S 1 ,S 2 ,S 3 The corresponding points are O 1 ,O 2 ,O 3 , and then use different polymerization rates r 1 ,r 2 ,r 3 For different significant regions, O 1 ,O 2 ,O 3 The tokens are aggregated to obtain the aggregated Where r represents the aggregation rate of r tokens merged into one token. The more important the region, the smaller the aggregation rate. Finally, The token sequence after redistribution is obtained by splicing. The process is shown in formula (11):
[0140]
[0141] In formula (11), F(·) represents the aggregation function, which is implemented by a fully connected layer with an input dimension of r and an output dimension of 1. i ,r i ) represents the i-th group O i The token sequence has an aggregation rate of r i Perform aggregation, Represents the token sequence after redistribution;
[0142] 3-3-2) The token sequence after redistribution Projected to the key matrix K and the value matrix V, and the original token sequence is projected to the query matrix Q. The specific process of calculating attention is shown in formula (12):
[0143]
[0144] In formula (12), atten represents the attention weight generated by the token reallocation Transformer block, and R represents the feature map generated by the token reallocation Transformer block;
[0145] 3 - 3 - 3) Update the saliency map and the feature map with the feature map R generated by the token reallocation Transformer block according to steps 3 - 2 - 1) to 3 - 2 - 3). The updated saliency map and feature map then go through steps 3 - 3 - 1) to 3 - 3 - 2) to obtain the multi - granularity static feature F S ;
[0146] 4) Dynamic - static feature two - way interactive fusion: Construct a two - way interactive attention fusion module BIAFM to fuse the dynamic feature F D from the dynamic branch and the multi - granularity static feature F S from the static branch, including:
[0147] 4 - 1) Project the dynamic feature F D onto the query matrix Q D , key - value matrix K D and value matrix V D , as shown in formula (13):
[0148]
[0149] In formula (13), is the projection matrix;
[0150] 4 - 2) Project the static feature F S onto the query matrix Q S , key - value matrix K S and value matrix V S , as shown in formula (14):
[0151]
[0152] In formula (14), and are the projection matrices;
[0153] 4 - 3) Calculate the attention weight of the dynamic feature, where the query matrix comes from the static feature, and then multiply it with the value matrix V D , as shown in formula (15):
[0154]
[0155] F′ D = Atten D VD
[0156] In formula (15), is the scaling factor, Atten D is the attention weight matrix of F D , and F′ D is the feature map of F D after being calculated by the attention mechanism;
[0157] 4 - 4) Calculate the attention weight of the static feature, where the query matrix comes from the dynamic feature, and then multiply it with the value matrix V S as shown in formula (16):
[0158]
[0159] F′ S = Atten S V S
[0160] In formula (16), Atten S is the attention weight matrix of F S , and F′ S is the feature map of F S after being calculated by the attention mechanism;
[0161] 4 - 5) Add F′ D and F′ S , retain the main feature information, and then obtain the final fused feature F fusion :
[0162] F fusion = AvgPool(F′ D ⊕ F′ S ) (17),
[0163] In formula (17), AvgPool(·) represents the average pooling operation;
[0164] 5) Input F fusion into the fully - connected layer to achieve accurate classification of micro - expressions.
[0165] To verify the effectiveness of the method in this example, it will be further illustrated by comparing the experimental results:
[0166] This example uses UF1 and UAR as evaluation indicators, and compares with five advanced methods on the public micro - expression datasets Full, SMIC, CASME II, SAMM and the micro - expression dataset GUCME in the teaching scenario, including CapsuleNet, RCN - A, MMNet, HTNet, HFA - Net, as follows:
[0167] 1) CapsuleNet: Extract local features through ResNet18, and then use the dynamic routing mechanism and vector representation of Capsule to effectively capture the local features of micro-expression vertex frames and the local and global relationships.
[0168] 2) RCN-A: Adopt low-resolution input and shallow network architecture, embed parameter-free attention units into the recursive convolutional network, and strengthen the feature response to sensitive areas of micro-expressions.
[0169] 3) MMNet: Model local subtle muscle movement patterns through continuous attention blocks, and combine a ViT-based position calibration module to achieve more accurate localization and recognition of micro-expressions.
[0170] 4) HTNet: Consists of Transformer layers and aggregation layers, focusing on modeling the complex dependencies between the left and right lip regions and the left and right eye regions, thus significantly improving the recognition accuracy of micro-expressions.
[0171] 5) HFA-Net: Extract micro-expression features through a two-branch architecture, and combine multi-level feature aggregation and adaptive attention feature fusion modules to enhance the model's discriminative ability for key features.
[0172] In this experiment, Dlib is used to detect the key points of human faces in images, and the detected key points are used for alignment and cropping. The images on all datasets are uniformly adjusted to 192×192 pixels to match the input sizes of the dynamic branch and the static branch. In addition, data augmentation methods such as random cropping and random horizontal flipping are used to avoid overfitting. To optimize the training of the model, the AdamW optimizer is used in this example, the weight decay parameter is set to 0.1, the batch size is 16, and the initial learning rate is 7×10 -4 , and 150 epochs are used for training. The experiment is carried out using PyTorch 1.10.0 and Python 3.7 on the Ubuntu 20.04 operating system.
[0173] Under the LOSO protocol, the UF1 score curves of each method on the Full, SMIC, CASME II, SAMM, and GUCME datasets are as Figure 6 , Figure 7 , Figure 8 , Figure 9 and Figure 10 shown, and are composed of Figure 6 , Figure 7 , Figure 8 , Figure 9 and Figure 10As can be seen, DSBICNet demonstrates significant performance advantages on the vast majority of datasets. On the Full composite dataset, DSBICNet shows stable performance advantages among 68 subjects, indicating its robustness in handling diverse data. On the SMIC dataset containing 16 subjects, the UF1 curve of DSBICNet is always above that of other methods, demonstrating its effectiveness in processing small-scale datasets. On the CASME II dataset containing 25 subjects, the UF1 value of DSBICNet is close to 1, showing significant advantages and indicating that DSBICNet has excellent adaptability to this dataset. On the SAMM dataset containing 28 subjects, the UF1 value of DSBICNet is approximately 0.86. Although it is slightly lower than HFANet, it is still significantly better than other advanced methods. For the self-constructed classroom micro-expression dataset GUCME, DSBICNet still maintains the lead among 18 subjects. Compared with other datasets, the performance of all methods on GUCME is generally lower, reflecting the challenges of micro-expression recognition in a real classroom environment. However, DSBICNet still shows the best environmental adaptability. These results fully prove that DSBICNet exhibits stronger robustness and accuracy than existing methods, indicating that the proposed DSBICNet in this example has significant generalization ability in the micro-expression recognition task.
[0174] Tables 1 and 2 show the UF1 and UAR metrics of DSBICNet and advanced methods on the Full, SMIC, CASME II, SAMM, and GUCME datasets. As can be seen from Tables 1 and 2, the proposed DSBICNet in this example performs excellently on the experimental datasets, demonstrating the strong robustness and effectiveness of DSBICNet in processing low frame rate and high consistency annotation as well as complex scene data. Compared with various advanced methods, this example solves four problems:
[0175] 1) DSBICNet fuses optical flow and pixel difference features through a dynamic-branch dual-attention-guided feature fusion module, making full use of dynamic information and enhancing the robustness of feature representation;
[0176] 2) DSBICNet transfers key motion features through a dynamic-branch gated cross-layer feature mechanism, effectively suppressing noise interference, avoiding signal attenuation caused by continuous downsampling, and accurately modeling the weak motion of micro-expressions;
[0177] 3) Through the static branch, DSBICNet reduces the computational amount by guiding the compression of tokens in non-critical regions with a saliency map and effectively models spatial fine-grained details;
[0178] 4) The complementary enhancement and collaborative optimization of dynamic and static features are achieved through the two-way interactive fusion of dynamic and static features. By solving the above problems, the method in this example achieves better performance.
[0179] Table 1 UF1 and UAR metrics of DSBICNet and advanced methods on the public dataset
[0180]
[0181] Table 2 UF1 and UAR metrics of DSBICNet and advanced methods on the GUCME dataset
[0182]
Claims
1. A method for dynamic and static two-way interactive collaboration in micro-expression recognition, characterized in that, It includes the following steps: 1) Quantify the motion features of microexpressions: Use the optical flow method and the frame difference algorithm to calculate the optical flow F flow and pixel difference F diff ; 2) Precise modeling of the subtle muscle movement characteristics of microexpressions: Input the optical flow F flow and pixel difference F diff into the dynamic branch, and adopt the dual-attention-guided feature fusion module DAGFFM and the gated cross-layer feature transfer mechanism GCFTM, including: 2-1) Construct a dual-attention-guided feature fusion module: Adopt a dual-attention mechanism to highlight the significant regions and key channels of micro-expression motion features and fuse them, including: 2-1-1) Perform 1×1 convolution operations on the optical flow map F flow and the pixel difference map F diff respectively to obtain feature maps and Then perform 3×3 convolution operations on them respectively to obtain feature maps and Concatenate with and along the channel dimension respectively to obtain X 12 and X 21 ; 2-1-2) Highlight the significant regions of micro-expression movements: Let X 21 Pass through the spatial attention mechanism SA, specifically: First, let X 21 Pass through the average pooling layer and the max pooling layer, and concatenate the output of the average pooling layer and the output of the max pooling layer. Then, successively pass through a 7×7 convolutional layer with a stride of 1 and the Sigmoid activation function to generate the spatial attention weight map W 21 , and finally, let the attention weight map W 21 Multiply element-wise with the original input feature map X 21 To obtain the weighted spatial attention feature map A 21 , as shown in formula (1): In formula (1), MaxPool(·) represents the max pooling operation, AvgPool(·) represents the average pooling operation, Concat(·) represents the concatenation operation, and f 7×7 (·) represents the 7×7 convolution operation, Sigmoid(·) represents the Sigmoid activation function, represents element-wise multiplication; 2-1-3) Highlight the key channels of micro-expression movements: X 12 Through the dual-path channel attention mechanism DPCA composed of two branches, specifically: in the max-pooling branch, X 12 After performing the max-pooling operation, it passes through a fully connected layer to learn the non-linear relationship between channels, and obtains the attention weight X' of this branch 12 , and in the average-pooling branch, the attention weight X'' of this branch is obtained according to the operation of formula (2) 12 : X″ 12 = δ(f 1×1 (AvgPool(δ(β(X 12 ))))) (2) In formula (2), β(·) represents channel normalization, δ(·) represents the ReLU activation function, and f 1×1 represents a 1×1 convolution operation; (2-1-4) Add X′ 12 and X″ 12 After addition, use the Sigmoid function for activation to obtain the attention weight W output by the dual-channel attention mechanism 12 , and perform element-wise multiplication with X 12 to obtain the channel attention feature map A 12 ; (2-1-5) Split the feature map A 12 into two parts on average in the channel dimension to obtain the split features and Similarly, split the feature map A 21 into feature and feature and obtain the micro-expression motion feature F according to the operation in formula (3) M : In formula (3), represents an addition operation; 2-2) Construct a gated cross-layer feature transfer mechanism: The gated cross-layer feature transfer mechanism consists of 8 convolutional layers with a size of 3×3 and 5 importance feature gating units FGIU, including: 2-2-1) Gated Cross-Layer Feature Transfer Mechanism for Motion Feature F M Performs two consecutive 3×3 convolution operations on it. The stride of the first convolution is 2, which downsamples the feature map and doubles the number of channels at the same time. The stride of the second convolution is 1, and the feature map L1 is obtained; 2-2-2) The first important feature: The gating unit groups the channels of L1, and then performs a normalization operation on each group according to formula (4) to obtain the normalized feature L gn : In formula (4), GN(·) represents the group normalization operation, μ and σ are the mean and standard deviation of L1, ε is a small normal constant added to ensure the stability of division, and γ and β are trainable affine transformations. Among them, γ is used to measure the spatial pixel variance of each channel. When the spatial information of micro-expression features is richer, the value of γ is larger. The weight matrix W of different feature maps is obtained by normalizing γ γ ; 2-2-3) Screening key motion features: Multiply the weight matrix W γ element-wise with L gn to obtain a weighted feature map, map it to the range (0, 1) using the Sigmoid function, then gate it with the set threshold T to obtain the importance weights and multiply them element-wise with the input feature L1. The specific process is shown in Equation (5): In formula (5), G1 represents the filtered key motion features, and Gate(·) represents the gating operation; 2-2-4) Retain the original complete information and enhance the response of the key region: Jump L1 to the output of the importance feature gating unit and add it to G1 to obtain the significant feature F1; 2-2-5) Establish a cross-layer feature transfer path: After selecting the key motion features from the shallow layer, they are directly injected into the deep network after being aligned by adaptive pooling. The specific process is shown in formula (6): In formula (6), i = 2, 3, 4, f 3×3 (·) represents a 3×3 convolution operation with a step size of 1, represents a 3×3 convolution operation with a step size of 2, represents the feature map after passing through two 3×3 convolutional layers, GU i (·) represents the i-th FIGU screening operation, represents the key motion features after the i-th FIGU screening, represents G i the pooled feature map, represents the updated significant features; (2-2-6) Further screen the significant feature F4 through the fifth importance feature gating unit to obtain the key feature G5, and add it to the key feature G4 screened by the fourth importance feature gating unit to obtain the dynamic feature F finally output by the dynamic branch D , as shown in formula (7): 3) Accurately extract the multi-granularity static features of the original micro-expression vertex frame: Construct a static branch composed of a patch embedding layer, four alternately stacked hybrid attention Transformer blocks, and a token reallocation Transformer block to extract the multi-granularity static features of the original micro-expression vertex frame, including: 3-1) The patch embedding layer divides the vertex frame F with a size of 192×192 into 36 patches, with an embedding dimension of 512, to obtain the embedded feature X; apex 3-2) Input the feature X into the hybrid attention Transformer block to extract local fine-grained and global coarse-grained information and generate a micro-expression saliency map: The self-attention mechanism of the hybrid attention Transformer block evenly divides the number of attention heads into two groups, each group containing 2 heads. Among them, the first group uses global attention, and the second group uses local attention, including: 3-2-1) Global attention calculation: First, perform a 6×6 convolution operation on the input feature X to compress it into a single feature representation, and then calculate self-attention. The calculation process is shown in formula (8): In formula (8), Q1, K1, and V1 respectively represent the query, key, and value matrices generated using global attention; SR(X, 6×6) represents the compression operation of applying a 6×6 convolutional operation to X; W1 Q , W1 K , W1 V is the projection matrix; atten1 represents the weight matrix generated by global attention; d is the dimension of the attention head; O1 represents the feature map generated by global attention; 3-2-2) Local attention calculation: X is projected into the query, key, and value matrices Q2, K2, V2 respectively. Then, the key matrix K2 and value matrix V2 are divided using a fixed-size window M to make the window division of Q2 match that of K2 and V2. The window size of Q2 is set to sM, where s represents the scaling factor of the receptive field. Then, calculate self-attention within the window, as shown in formula (9): In formula (9), Window(·) represents the local window division operation, Q2′, K2′, V2′ respectively represent the query, key, and value matrices generated by local attention, atten2 represents the weight matrix generated by local attention, and O2 represents the feature map generated by local attention; 3-2-3) Generate a saliency map: Concatenate the feature maps O1 and O2 generated by global attention and local attention to obtain the feature map O generated by hybrid attention. At the same time, perform an average operation on the weight matrices atten1 and atten2 generated by global attention and local attention and then aggregate them to obtain the saliency map S. The process is shown in formula (10): In formula (10), represents the attention score of the i-th head of global attention at position p, represents the attention score of the j-th head of local attention at position q; 3-3) Input the feature map O and the saliency map S into the token reallocation Transformer block, including: 3-3-1) Sort the importance weights of the saliency map S in ascending order and divide it into 3 sub-regions S 1 , S 2 , S 3 , where S 1 is the least important region, S 3 is the most important region. At the same time, all tokens in the feature map O are divided into O 1 , S 2 , S 3 correspondingly, and then different aggregation rates r 1 , r 2 , r 3 are used to aggregate the tokens of O 1 , O 2 , O 3 in different saliency regions to obtain the aggregated 1 , O 2 , O 3 . Among them, r represents the aggregation rate at which every r tokens are combined into one token. The more important the region is, the smaller the aggregation rate. Finally, is concatenated to obtain the reallocated token sequence. The process is shown in Equation (11): is concatenated to obtain the reallocated token sequence. The process is shown in Equation (11): In formula (11), F(·) represents an aggregation function, which is implemented by a fully connected layer with an input dimension of r and an output dimension of 1. F(O i ,r i ) represents the token sequence of the i-th group O i aggregated at an aggregation rate of r i ; represents the token sequence after reallocation. 3-3-2) Project the reallocated token sequence onto the key matrix K and the value matrix V, while the original token sequence is projected onto the query matrix Q. The specific process of calculating the attention is shown in Equation (12): In Equation (12), atten represents the attention weights generated by the token reallocation Transformer block, and R represents the feature map generated by the token reallocation Transformer block; 3-3-3) Reassign the feature map R generated by the token reallocation Transformer block to update the saliency map and the feature map according to steps 3-2-1) to 3-2-3). The updated saliency map and feature map then go through steps 3-3-1) to 3-3-2) to obtain the multi-granularity static feature F S ; 4) Bidirectional interactive fusion of dynamic and static features: Construct a bidirectional interactive attention fusion module BIAFM to fuse the dynamic feature F from the dynamic branch and the multi-granularity static feature F from the static branch, including: D and the multi-granularity static feature F from the static branch S , including: 4-1) Project the dynamic feature F D onto the query matrix Q D , the key matrix K D and the value matrix V D , as shown in formula (13): In formula (13), is the projection matrix; (4-2) Project the static feature F S onto the query matrix Q S , the key matrix K S and the value matrix V S , as shown in Equation (14): In formula (14), and are projection matrices; 4-3) Calculate the attention weights of the dynamic features, where the query matrix comes from the static features and then multiplies with the value matrix V D as shown in formula (15): In formula (15), is the scaling factor, and Atten D is the attention weight matrix of F D , and F D ' is the feature map of F D after being calculated by the attention mechanism; 4-4) Calculate the attention weights of the static features, where the query matrix comes from the dynamic features and then multiplies with the value matrix V S as shown in formula (16): In formula (16), Atten S is the attention weight matrix of F S , and F S ' is the feature map obtained by calculating F S through the attention mechanism; (4-5) Add F D ′ and F S ′, retain the main feature information, and then obtain the final fused feature F fusion : In Equation (17), AvgPool(·) represents the average pooling operation; 5) Input F fusion into the fully connected layer to achieve accurate classification of microexpressions.
Citation Information
Cited By
Facial paralysis grading method, system and equipment fusing multi-modal data and medium
CN120876479A
A method, system, device, and medium for grading facial paralysis by integrating multimodal data.
CN120876479B
Micro-expression recognition method and device based on multi-scale optical flow attention guiding mechanism
CN121661698A
Micro-expression recognition method and device based on multi-scale optical flow attention mechanism
CN121661698B