Micro-expression recognition method based on cross-source double-branch dynamic space-time diagram convolutional network model
Through the cross-original dual-branch dynamic spatiotemporal graph convolution network model, combining global information with dynamic spatiotemporal feature extraction network and attention-enhanced twin spatiotemporal graph fusion network, the problem of micro-expression recognition in the existing technology is difficult to capture subtle motion differences and has a great impact on non-expression factors, achieving more efficient micro-expression recognition and stronger generalization capabilities.
Patent Information
- Application Number
- CN202510302940.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
Existing micro-expression recognition technology is difficult to accurately capture subtle differences in facial muscle movements, and has a great impact on non-expression factors such as lighting changes and head posture offsets during video preprocessing, resulting in limited model generalization ability.
The cross-original dual-branch dynamic spatiotemporal graph convolution network model is adopted, and the domain invariant features are learned through twin structures, and the network is extracted and attention-enhanced twin spatiotemporal graph fusion network is extracted to extract subtle motion characteristics of facial structures when expression changes, and the model is optimized through cross-domain joint loss.
It improves the accuracy and generalization ability of micro-expression recognition, can more effectively extract and distinguish the subtle motor characteristics of micro-expression, and reduces the impact on non-expression factors.
Smart Images

Figure CN120236311A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a micro-expression recognition method based on a cross-source dual-branch dynamic spatio-temporal graph convolutional network model, belonging to the technical fields of deep learning and pattern recognition. Background Art
[0002] Micro-expressions are the instantaneous leakage of human inner emotions, which can reflect the emotions that an individual attempts to suppress or hide. They are a subtle, short-lived, and spontaneous emotional representation, usually occurring when a person deliberately or unconsciously hides his true emotions. This provides a basis for revealing people's true psychology or emotions. Therefore, the accurate recognition of micro-expressions has extremely important application value in security interrogation, clinical diagnosis, business negotiation, and daily social interaction.
[0003] In the field of manually extracting features, several common feature extraction methods are widely used. In particular, Guo et al. introduced a feature called LBP-TOP, which is extended from the LBP algorithm, calculates the local binary pattern on three orthogonal planes, introduces the three-dimensional space into feature calculation, and adds the information of the time dimension. In addition, Liu proposed a main direction average optical flow feature (MDMO), which is based on a specific region of interest in the image. MDMO first divides the face into 36 regions of interest, then calculates the average optical flow of all frames in each region and uses it as a feature representation. This method reduces the feature dimension to 72 dimensions, significantly reducing the computational complexity. On this basis, Liu et al. further proposed a sparse MDMO feature, proposed a new distance metric method and combined it with graph-regularized sparse coding, which can more effectively reveal the potential manifold structure of the features. The above several schemes do not consider the impact of head movement on expression recognition. Therefore, Xu et al. proposed a Facial Dynamic Map (FDM) method, which extracts a dense optical flow field, regards it as a spatio-temporal cuboid and calculates the main direction, effectively statistically analyzing the facial expression movement pattern to reduce the impact caused by head movement.
[0004] Traditional feature engineering methods have technical bottlenecks. These methods rely on manually designed feature operators, and their performance is highly restricted by prior knowledge and specific scenario assumptions, making it difficult to fully represent the complex and variable dynamic features of facial expressions. Especially in the task of micro-expression recognition, due to the extremely small amplitude of facial muscle movement (usually less than 500 milliseconds), traditional machine learning frameworks based on handcrafted features are difficult to capture the subtle differences in the combination changes of AUs (Action Units). In addition, existing methods have a strong dependence on the video preprocessing process. Non-expression factors such as illumination changes and head pose offsets will be amplified through the preprocessing link, inevitably introducing redundant noise and restricting the generalization ability of the classification model. In recent years, the end-to-end feature learning framework based on deep learning has gradually become the research focus. Its characteristic of automatically extracting deep semantic features through multi-layer non-linear transformations provides a new technical path for solving the technical difficulties of micro-expression recognition. Deep learning has developed rapidly in recent years and has achieved remarkable results in many fields. Micro-expression recognition, as a frontier topic in the field of computer vision, has also made continuous breakthroughs with the help of deep learning. In 2016, Dae Hoe Kim proposed a micro-expression recognition method that combines a convolutional neural network (CNN) and a long short-term memory network (LSTM). This method can simultaneously capture spatial and temporal features in the video sequence, providing a new idea for micro-expression recognition. In 2018, Wang further proposed a micro-expression recognition method using the transfer learning strategy. By using the model pre-trained on a large-scale dataset, the performance of micro-expression recognition was effectively improved. In 2020, Yante Li proposed a micro-expression peak frame recognition algorithm based on three-dimensional Fourier transform and combined it with the LGCcon model. By fusing local and global features to train the network, the accuracy of micro-expression recognition was further improved. Xie proposed a method using facial action units (AUs) to assist micro-expression recognition, enhancing the ability of micro-expression recognition by identifying subtle action units on the face. Ling et al. proposed the MER-GCN model, which is the first end-to-end micro-expression recognition architecture based on action units (AUs) and GCN. This method extracts AU features through a 3D convolutional neural network, constructs the GCN adjacency matrix in a data-driven manner, represents the mutual relationship using the co-occurrence of AUs in the training set, and uses a one-hot encoder to convert AUs into machine-understandable labels. This model is superior to the micro-expression recognition network based on CNN, demonstrating the positive effect of GCN learning the relationship of facial physical muscle movements on micro-expression understanding. In 2021, Ben conducted an in-depth investigation and analysis of the field of micro-expression detection and recognition based on video datasets and prospected the future prospects. Aiming at the problem of the lack of micro-expression datasets, he proposed a new dataset MMEW, providing valuable resources for micro-expression recognition research.Ankith et al. used Eulerian motion magnification (EMM) to enhance video signals, automatically selected frames with high expression intensity through the optical flow magnitude threshold, detected facial landmark points and constructed a graph structure, calculated the optical flow patch feature information at the landmark points, and finally classified them through a two-stream GACNN. In 2022, Xiong et al. proposed a micro-expression recognition algorithm based on GCN and long short-term memory network (LSTM), which used GCN to extract the spatial features of micro-expression images and then used LSTM to extract the temporal features. During the calculation process of GCN, the output of each layer of the graph convolutional network was updated through a specific formula; LSTM processed the time series information through a complex gating mechanism.
[0005] Micro-expression features are subtle, instantaneous, and difficult to extract. Existing micro-expression datasets have a small number of samples and class imbalance. When performing recognition tasks, it is difficult to identify effective expression features, and irrelevant interference information such as background light will also cause obvious performance losses to the model. Considering the biological principle of the movement process of facial muscles, and in the preprocessing process of existing methods, the same facial area is wrongly mapped to different pixel positions in the image, resulting in the network being difficult to accurately locate the position where the subtle movement of micro-expressions occurs, thereby affecting the model performance. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention provides a micro-expression recognition method based on a cross-source dual-branch dynamic spatio-temporal graph convolutional network model. The present invention proposes a cross-source dual-branch siamese spatio-temporal graph structure network to mine the subtle movement features of the facial structure during expression changes and learn domain-invariant features through the siamese structure. This design models the global facial information through a global information and dynamic spatio-temporal feature extraction network, extracts the subtle movement information of the key facial structures through an attention-enhanced siamese spatio-temporal graph fusion network, fully models the facial structure association during expression occurrence, and performs domain adaptation by constructing a cross-domain joint loss.
[0007] Term Explanation:
[0008] 1. Dlib Visual Library: Dlib is an open-source C++ toolkit containing machine learning algorithms, which can be used to solve many practical problems in the field of machine learning. Currently, Dlib has been widely used in the industrial and academic fields.
[0009] 2. Detection of Facial Key Feature Points: The 68 key facial feature points are mainly distributed on the eyebrows, eyes, nose, mouth, and facial contour. As Figure 1 shown, it is detected through the Dlib visual library, which is the prior art.
[0010] 3. Spatio-Temporal Graph Convolutional Network: A spatio-temporal graph convolutional network is a neural network model for processing data with spatio-temporal structures. It models time and space information in the form of a graph, where nodes can represent different positions or objects in space, and edges are used to describe the relationships between nodes in space and time.
[0011] 4. Attention Mechanism: The attention mechanism is inspired by the way the human visual system processes information. That is, when humans observe a scene, they automatically focus their attention on the regions of interest and ignore other irrelevant information. In deep learning, the attention mechanism aims to enable the model to automatically learn the important information in the data and assign different attention weights to different input elements.
[0012] 5. Pre-trained Model: A pre-trained model is a model that has been previously trained on a large-scale dataset. These models are usually based on deep learning architectures such as Transformer, ResNet, etc., and are trained in an unsupervised or supervised manner on a vast amount of data to learn the general features and patterns in the data.
[0013] 6. Triplet Loss: Triplet loss is a loss function used in metric learning and is commonly used in deep neural networks. Its basic idea is to learn the feature representation of data by constructing triplet data, so that similar data is closer in the feature space, and dissimilar data is farther apart. Specifically, a triplet consists of an anchor sample, a positive sample, and a negative sample.
[0014] The technical solution of the present invention is as follows:
[0015] A micro-expression recognition method based on a cross-source dual-branch dynamic spatio-temporal graph convolutional network model, comprising the following steps:
[0016] A. Preprocessing of micro-expression video sequences, including: obtaining a video frame sequence, face detection and localization, face alignment, and facial graph cropping;
[0017] B. Constructing a facial key point graph;
[0018] C. Constructing a global motion information and spatio-temporal feature extraction network (Global Information and Dynamic Spatio-Temporal Feature Extraction Network, GIDST); using a lightweight network to extract global features from the original input image sequence, and mapping the feature map fused with global information into the facial key point structure through a key point mapping network as node features, and then entering step D. Using two modules of the network in step D to construct a spatio-temporal graph feature extraction network to obtain features for calculating classification loss and triplet loss;
[0019] D. Construct an Attention-Enhanced Siamese Spatio-Temporal Graphs Fusion Network (AESSTG); this network includes two layers of feature extraction networks, and each layer of feature extraction network includes a Static Spatio-Temporal Graph Convolution Module (SSTGC) and a Dynamic Spatio-Temporal Graph Feature Fusion Module (DSTGFF); residual connections are added between each layer of feature extraction networks to make the model optimization easier. The node input feature of the network is the pixel feature of the fixed area size corresponding to the key point; after the features extracted by the network pass through spatio-temporal pooling, they are output through two different Fully Connected Layers (FCL) to obtain the features for constructing the classification loss and the triplet loss respectively, completing the extraction of the global motion features.
[0020] E. Construct a cross-domain joint loss, calculate the classification loss in the macro-expression domain and the micro-expression domain and the cross-domain triplet loss respectively, which are used to optimize the global motion information and the spatio-temporal feature extraction network and the attention-enhanced siamese spatio-temporal graph fusion network, so that the model can learn domain-invariant features, and the obtained final model is used for micro-expression recognition.
[0021] Preferably according to the present invention, in step A, the preprocessing of the micro-expression video sequence includes the following steps:
[0022] 1) Obtain the video frame sequence: perform frame splitting on the micro-expression video sequence to obtain the video frame sequence and store it.
[0023] 2) Obtain the starting frame, peak frame and ending frame: according to the information annotated by experts in the micro-expression dataset, select the starting frame, peak frame and ending frame of the micro-expression in each video sequence file, and uniformly sample within the interval of the three special frames to obtain a frame sequence with a length of 8.
[0024] The starting frame refers to the first frame in the micro-expression video sequence where the micro-expression appears.
[0025] The peak frame refers to the frame in the micro-expression video sequence where the facial muscle changes are most obvious and contains the most micro-expression information.
[0026] The ending frame refers to the last frame in the micro-expression video sequence where the micro-expression ends.
[0027] 3) Face detection and localization: Use the Dlib library to perform face detection and localization on the frame sequence obtained in step 2), and detect the number of faces in the video frame and the distance of the face from the image boundary;
[0028] 4) Face alignment: On the basis of face localization, use the Dlib library to determine 68 key feature points on the face, and complete face segmentation, face correction, and facial map cropping;
[0029] Face segmentation refers to: Using the Dlib vision library to segment the face with a rectangular box;
[0030] Face correction refers to: Among the 68 key feature points detected on the face, the line connecting the key feature point marked at the left corner of the left eye and the key feature point marked at the right corner of the right eye has an angle a with the horizontal line. Through this angle a, the corresponding rotation matrix is obtained, and the segmented face is rotated to make the line connecting the key feature point marked at the left corner of the left eye and the key feature point marked at the right corner of the right eye parallel to the horizontal line, realizing the correction of the face pose, and scaling the face;
[0031] Facial map cropping refers to: Scaling the segmented and corrected face to obtain a picture of a fixed size.
[0032] Further preferably, in step 4), the fixed size of the picture obtained during facial map cropping is set to 128 * 128 pixels.
[0033] According to the preference of the present invention, in step B, the selection and preprocessing of the facial key point map are as follows:
[0034] By calculating the optical flow features and image entropy, it can be found that the main facial movements during the occurrence of micro-expression emotional changes are reflected in areas such as the eyebrows and mouth, and a small part is reflected in areas such as the eyes and cheeks. The key facial structures in these areas are the basic units we focus on. An overly large adjacency matrix will cause a significant increase in the model's computational complexity. Therefore, in order to balance the computational complexity and the feature contribution degree of the key points, finally, 12 points on the mouth contour, 5 points on the left eyebrow, 5 points on the right eyebrow, 4 points on the nose, and 3 points on the face contour, a total of 29 points, are selected as the nodes of the graph, and an initial adjacency matrix is constructed according to the facial structure relationship.
[0035] Calculate the optical flow features of the image sequence by the Farneback method. Its core idea is to process large displacements through image pyramid hierarchical processing and approximate the optical flow field using polynomial expansion. In a local window, represent the image brightness as a quadratic polynomial: I(x, y, t) = a + bx + cy + dx 2 + exy + fy 2, where abcdef are all constants. For adjacent frames I1 and I2, solve for the displacement (u, v) such that: I2(x + u, y + v, t + 1) = I1(x, y, t). Estimate the optical flow by minimizing the following error function, and this process can be expressed as Equation (1):
[0036]
[0037] where λ is the smoothing weight, I(x, y, t) represents the luminance of the image at the position (x, y) and at time t, and I x 、I y represent the gradients of the image in the x and y directions, and I t represents the rate of change of the image at time t. u is the velocity of the optical flow in the x direction, and v is the velocity of the optical flow in the y direction;
[0038] Image entropy is a statistic used to describe the information content of an image and reflects the distribution of pixel values in the image. For a grayscale image, the formula for calculating its entropy is:
[0039]
[0040] where L is the number of gray levels of the image (for example, for an 8-bit grayscale image, (L = 256)), and p i is the probability that a pixel with a gray value of i appears in the image; according to the calculation of image entropy and optical flow features, the present invention selects 29 facial key points with obvious motion change characteristics during micro-expression occurrence to form a facial key point map.
[0041] The structure of the facial key point map can be expressed as In the facial key point map, the nodes represent the facial landmark points located such as the outer end of the eyebrow and the corner of the mouth, and are expressed as The total number of nodes is N; ε is the edge representing the explicit connection of the key points, and can be represented by the adjacency matrix For the element a i,j in the i-th row and j-th column of it, if there is an edge connection between v i and v j , a i,j = 1, otherwise a i,j = 0; The expression recognition based on the facial key point map uses such a set of graph sequences as input, and its key point feature set can be represented by a tensor as where x t,n is the C-dimensional feature vector of the key point v n at time t in the T-frame facial key point map, which is usually the feature after global feature extraction processing or the original pixel value at the corresponding position in the original data; thus, the graph sequence structure used in the present invention can be described by the node feature set X and the adjacency matrix A.
[0042] Preferably according to the present invention, in step C, a global motion information and spatio-temporal feature extraction network is constructed, and the steps are as follows:
[0043] The present invention constructs a Global Information and Dynamic Spatio-Temporal Feature Extraction Network (GIDST). A lightweight network is used to extract global features from the expression sequence, and the feature map integrating global information is mapped into the facial key point structure through the key point mapping network as node features. Specifically, first, the pre-trained model MobileNet is used to extract features from the original input image sequence, and the global semantic information of the image is obtained through the pre-trained model; then, through the key point mapping network, specifically implemented as node feature mapping based on the self-attention mechanism, the grid features of the global feature map are mapped into the node features of the facial key points; finally, the facial key point map integrating global semantic information enters step D, and the attention-enhanced twin spatio-temporal graph fusion network proposed in step D of the present invention is further used to extract detailed information between nodes, thereby improving the recognition performance of the model for global information and detailed spatio-temporal motion features. Specifically, two modules of the network in step D are used to form a layer of feature extraction network for feature extraction, and two features for cross-entropy and triplet loss are obtained for global information.
[0044] Furthermore, first, the pre-trained MobileNet model is used to model the global information of the input sample sequence from the spatial dimension. After the input RGB image sequence is input, the feature map F is obtained. Map This feature map is changed into graph node features through the key point mapping network. First, the attention mechanism is applied to weight the feature map, enabling the network to adaptively focus on the important parts in the feature map and enhancing the expression ability of the features. Then, the feature map passes through the convolutional network. In this way, each channel corresponds to a node, thereby associating the channel information of the feature map with the nodes; Input feature map: C is the number of channels, and H and W are the height and width of the feature map respectively; Query vector Key vector Value vector
[0045]
[0046]
[0047] ⊙ represents element-wise multiplication, and the obtained Q, K, and V are transposed and flattened for calculating the attention score. The formula is:
[0048]
[0049] att = Softmax(Scores) (7)
[0050]
[0051] The above process represents the following formula to obtain the attention-weighted x att Perform channel-node feature mapping through convolution and obtain node features of size 7*7 through pooling;
[0052] x att = Attention(F Map ) (9)
[0053]
[0054] where w conv is the convolution weight, mapping the channel features to the node dimension, Reshape() is the dimension adjustment operation, and Attention refers to calculating the input F Map through the process of formulas (3)-(8), x node and x avg respectively represent the features after node mapping and pooling;
[0055] Finally, the obtained graph sequence X out enters step D. Through the attention-enhanced twin spatio-temporal graph fusion network (AESSTG) constructed in step D, further spatio-temporal feature extraction is performed. After spatio-temporal pooling, it is output through two different fully connected layers (FCL), and the features for constructing the classification loss and the triplet loss are obtained respectively, completing the extraction of the global motion features.
[0056] According to the preference of the present invention, in step D, constructing the attention-enhanced twin spatio-temporal graph fusion network includes the following steps:
[0057] For the input facial key graph sequence C, T, and N are the number of channels, the number of frames, and the number of nodes respectively. The pixel features of the facial key points are obtained through the facial key point data preprocessing module as the network input features of the static spatio-temporal graph convolution C in is the number of input feature channels;
[0058] The Attention-Enhanced Siamese Spatio-Temporal Graph Fusion Network includes two layers of feature extraction networks. Each layer of the feature extraction network includes a Static Spatio-Temporal Graph Convolution Module (SSTGC) and a Dynamic Spatio-Temporal Graph Feature Fusion Module (DSTGFF); Residual connections are added between each layer of the feature extraction network to make the model optimization easier.
[0059] The static spatial graph convolution layer is used to fuse the information between the fixed adjacent nodes in the same frame. The original adjacency matrix represents the relationship between the directly connected nodes and is denoted as Rich spatial features are extracted through the K-order adjacency matrix; for the k-hop adjacency matrix A k , the element a k(i,j) in its i-th row and j-th column is defined as:
[0060]
[0061] where dis(v i , v j ) represents the distance (shortest path length) between the key points v i and v j in the skeleton graph, k ∈ {0,..., K - 1}; therefore, A0 = I, A1 = A; for After normalizing it, it is used as the k-hop structural adjacency matrix
[0062]
[0063] where D k represents the degree matrix of the k-order link;
[0064]
[0065] where C out is the number of output feature channels, is the learnable parameter matrix, K represents the order of the k-hop adjacency matrix, F Sout is the result of the spatial graph convolution obtained, F in is the input feature, and σ(·) is the sigmoid activation function.
[0066] The static temporal convolution layer processes the state information of the same node at different time steps. After the static spatial graph convolution layer, the output feature dimension is T × N × C out, the size of the temporal convolution kernel is m×1, and the stride is s. After performing convolution on the nodes at the same position in m frames, skip s frames along the temporal dimension and continue convolution until the convolution of nodes with the same number is completed, and then perform convolution on the next numbered node;
[0067] The above process can be expressed by the formula:
[0068] F Tout = TCN(F Sout )(16)
[0069] where TCN(·) represents temporal convolution, represent the input and output of the static temporal convolution layer respectively.
[0070] The static spatio-temporal graph samples features from the facial key point graph based on a specific pattern and is difficult to learn the implicit link relationships between different nodes. For example, when an angry expression occurs, there is no explicit link in the fixed adjacency matrix between the corners of the mouth and the eyebrows, but there is a co-movement relationship during the actual movement of the expression, which is difficult to capture through the adjacency matrix of the static facial position relationship. Therefore, a dynamic spatio-temporal graph feature fusion module is designed for the generation of an adaptive adjacency matrix, which can capture the movement relationships between nodes.
[0071] On the basis of the operations of the static spatio-temporal graph convolution, the dynamic spatio-temporal graph feature fusion module introduces the generation of a feature-based dynamic adjacency matrix and performs convolution operations on the basis of the new adjacency matrix;
[0072] For the input feature F in , calculate the temporal dimension attention, node dimension attention, and feature channel dimension attention respectively, and obtain a multi-dimensional attention tensor through the broadcast mechanism;
[0073] The calculation process for the temporal dimension attention is as follows:
[0074] a t = σ(W time ·AvgPool (V,1) (x))(17)
[0075] where, AvgPool (V,1) (·) is the average pooling function in the spatial dimension, V is the node dimension, W time is a one-dimensional convolution with a kernel size of 3, used to expand the temporal receptive field, and σ(·) is the Sigmoid activation function.
[0076] The calculation process for the node dimension attention is as follows:
[0077] a s = σ(W spatial ·AvgPool(T) (Reshape(x))) (18)
[0078] Among them, AvgPool (T) (·) is the average pooling function in the time dimension, is the learnable parameter matrix of the fully connected layer.
[0079] The calculation process for the feature channel dimension attention is as follows:
[0080] a c = σ(W channel ·AvgPool (T,V) (x)) (19)
[0081] Among them, AvgPool (T,V) (·) is the average pooling function in the time and space dimensions, are the learnable parameter matrices of two linear layers, r is the dimension transformation factor, set to 16;
[0082] a = σ(a t ·a s ·a c ) (20)
[0083] After calculating the multi-dimensional attention tensor through the broadcasting mechanism and performing element-wise multiplication with the input features, the enhanced features are obtained for subsequent convolution processing;
[0084] F a = F in ⊙a (21)
[0085] The obtained attention tensor is mapped to a low-dimensional space through two linear layers and undergoes dimension transformation to obtain a dynamically generated adjacency matrix. On the basis of retaining the key facial structures, the k-th order adjacency matrix is added to the output k-th order matrix and weighted, and the weight is a learnable gating parameter for balancing the stability of the adjacency matrix;
[0086] A d = γ[0]·A[0] + γ[k]·(W2(ReLU(W1(a))) + A[k]) (22)
[0087] Among them, W2, W2 are the weights of the fully connected layer, A is the input adjacency matrix; γ is the gating parameter for controlling the weighted balance between the k-th order matrix and the adaptive adjacency matrix; through the generated dynamic adjacency matrix A d and the enhanced feature F a perform spatio-temporal convolution operations to obtain the output F Gout .
[0088] The output layer includes a temporal pooling layer, a spatial pooling layer, and a fully connected layer, and the output features of the last temporal feature extraction network perform temporal pooling and spatial pooling operations respectively. The pooling process can be expressed by the formula:
[0089] F T =MaxPool t (F Gout ) (23)
[0090] F S =MaxPool s (F T )+AvgPool s (F T ) (24)
[0091] where represent the output features after temporal and spatial pooling respectively; MaxPool t (·) represents max pooling in the temporal dimension, and MaxPool s (·) and AvgPool s (·) represent max pooling in the spatial dimension and average pooling in the spatial dimension respectively; finally, F S passes through two different fully connected layers to obtain the features for calculating the cross-entropy loss and F CE and F Tri represent the feature outputs for calculating cross-entropy and triplet loss respectively, and C class is the number of classes, that is, the number of emotion classes.
[0092] According to the preference of the present invention, in step E, constructing a cross-domain joint loss includes the following steps:
[0093] The cross-domain joint loss combines the classification task and the cross-domain task, optimizes the network through the classification cross-entropy loss and the triplet loss, and the global features local features obtained through steps C and D respectively
[0094]
[0095] where y i is the true label and p i is the predicted value.
[0096] Taking each epoch as the update unit, the global features local features obtained through steps C and D respectively construct triplet loss respectively
[0097] Specifically, taking each micro-expression sample as an anchor point, through clustering, randomly select samples with the same label that are closest to the sample as positive samples, and samples with different labels as negative samples to construct triplets; by calculating the Euclidean distances between the anchor point and the positive and negative samples, samples with the same label are made closer in the feature space, while samples with different labels are made farther apart, thereby enhancing the model's ability to distinguish features.
[0098] For the micro-expression anchor sample F micro 、positive sample F POS and negative sample F NEG , its triplet loss can be calculated by the following formula:
[0099]
[0100] where d(·) is the distance metric (such as Euclidean distance); m is a margin constant used to control the distance between positive and negative samples, represents the triplet loss calculated from global features, represents the triplet loss calculated from local features.
[0101] Calculate the triplet loss and the cross-entropy loss and use the average value of all branch losses to supervise the training process of the model; the above process can be expressed by the following formula:
[0102]
[0103] where β is a hyperparameter used to balance the weights of the two losses.
[0104] The beneficial effects of the present invention are as follows:
[0105] In view of the problem of micro-expression recognition, the present invention proposes a cross-source dual-branch dynamic spatio-temporal graph convolutional network model, which maps features in different domains to the same space through a siamese structure to learn domain-invariant features. The global motion features are obtained through the global information and dynamic spatio-temporal feature extraction network (GIDST), and the subtle motion features are learned based on the facial structure through the attention-enhanced siamese spatio-temporal graph fusion network (AESSTG), thereby improving the performance of the recognition algorithm. Brief Description of the Drawings
[0106] Figure 1 It is a schematic diagram of 68 key feature points on the face of the present invention;
[0107] Figure 2 It is a schematic diagram of the structure of the global motion information and spatio-temporal feature extraction network of the present invention;
[0108] Figure 3 Schematic diagram of the attention-enhanced twin spatio-temporal graph fusion network structure of the present invention;
[0109] Figure 4 Schematic diagram of the cross-source dual-branch dynamic spatio-temporal graph convolutional network framework proposed by the present invention;
[0110] Figure 5 Schematic diagram of the facial key point graph construction process of the present invention;
[0111] Figure 6 Schematic diagram of the dynamic adjacency matrix (order 0) defined by the present invention;
[0112] Figure 7 Schematic diagram of the dynamic adjacency matrix (order 1) defined by the present invention;
[0113] Figure 8 Schematic diagram of the confusion matrix of the five-classification experimental results of the present invention on the CASMEII dataset, and the five-classification labels are ("surprise", "happiness", "disgust", "sadness", "fear");
[0114] Figure 9 Schematic diagram of the confusion matrix of the three-classification experimental results of the present invention on the CASMEII dataset, and the three-classification labels are ("surprise", "happiness", "negative");
[0115] Figure 10 Schematic diagram of the confusion matrix of the five-classification experimental results of the present invention on the SAMM dataset, and the five-classification labels are ("surprise", "happiness", "disgust", "sadness", "fear");
[0116] Figure 11 Schematic diagram of the confusion matrix of the three-classification experimental results of the present invention on the SAMM dataset, and the three-classification labels are ("surprise", "happiness", "negative");
[0117] Figure 12 Comparison of ablation experiment indicators of each module proposed by the present invention. Detailed implementation manners
[0118] The present invention will be further described below by way of examples in conjunction with the accompanying drawings, but is not limited thereto.
[0119] Example 1:
[0120] A micro-expression recognition method based on a cross-source dual-branch dynamic spatio-temporal graph convolutional network model, as Figure 4 shown, includes the following steps:
[0121] A. Preprocessing of the micro-expression video sequence, including: obtaining the video frame sequence, face detection and localization, face alignment, and facial graph cropping;
[0122] In step A, preprocess the micro-expression video sequence, and the steps are as follows:
[0123] 1) Obtain the video frame sequence: Perform frame splitting on the micro-expression video sequence to obtain the video frame sequence and store it;
[0124] 2) Obtain the starting frame, peak frame, and ending frame: According to the information annotated by experts in the micro-expression dataset, select the starting frame, peak frame, and ending frame of the micro-expression in each video sequence file, and uniformly sample within the three special frame intervals to obtain a frame sequence with a length of 8;
[0125] The starting frame refers to: The first frame in the micro-expression video sequence where the micro-expression appears;
[0126] The peak frame refers to: The frame in the micro-expression video sequence where the facial muscle changes are most obvious and contains the most micro-expression information;
[0127] The ending frame refers to: The last frame in the micro-expression video sequence where the micro-expression ends;
[0128] 3) Face detection and localization: Use the Dlib library to perform face detection and localization on the frame sequence obtained in step 2), and detect the number of faces in the video frame and the distance between the face and the image boundary;
[0129] 4) Face alignment: On the basis of face localization, use the Dlib library to determine 68 key facial feature points, as Figure 5 shown, and complete face segmentation, face correction, and facial map cropping;
[0130] Face segmentation refers to: Using the Dlib vision library to segment the face with a rectangular box;
[0131] Face correction refers to: Among the 68 key facial feature points detected, the line connecting the key feature point of the left eye corner and the key feature point of the right eye corner has an angle a with the horizontal line. Through this angle a, the corresponding rotation matrix is obtained, and the segmented face is rotated to make the line connecting the key feature point of the left eye corner and the key feature point of the right eye corner parallel to the horizontal line, realizing the correction of the face pose, and scaling the face;
[0132] Facial map cropping refers to: Scaling the segmented and corrected face to obtain a picture with a fixed size. In this embodiment, the fixed size of the picture is set to 128*128 pixels.
[0133] B. Construct the facial key point map;
[0134] In step B, the selection and preprocessing of the facial key point map are as follows:
[0135] By calculating the optical flow features and image entropy, it can be found that the main facial movements during the occurrence of micro-expression emotional changes are reflected in areas such as the eyebrows and mouth, and a small part is reflected in areas such as the eyes and cheeks. The key facial structures in these areas are the basic units we focus on. An overly large adjacency matrix will cause a significant increase in the model's computational complexity. Therefore, in order to balance the computational complexity and the feature contribution degree of the key points, 29 key points are finally selected, namely the mouth contour (48 - 59), the left eyebrow (22 - 26), the right eyebrow (17 - 21), the nose (27, 31, 33, 35), and the contour (3, 13, 8). Twelve points of the mouth contour, five points of the left eyebrow, five points of the right eyebrow, four points of the nose, and three points of the face contour, a total of 29 points, are selected as the nodes of the graph, and an initial adjacency matrix is constructed according to the facial structure relationship.
[0136] The optical flow features of the image sequence are calculated by the Farneback method. Its core idea is to process large displacements through image pyramid layering and approximate the optical flow field using polynomial expansion. Within a local window, the image brightness is expressed as a quadratic polynomial: I(x, y, t) = a + bx + cy + dx 2 + exy + fy 2 , where a, b, c, d, e, and f are all constants. For adjacent frames I1 and I2, the displacement (u, v) is solved such that: I2(x + u, y + v, t + 1) = I1(x, y, t). The optical flow is estimated by minimizing the following error function, and this process can be expressed as Equation (1):
[0137]
[0138] where λ is the smoothing weight, I(x, y, t) represents the brightness of the image at the position (x, y) and time t, I x 、I y represent the gradients of the image in the x and y directions, I t represents the change rate of the image at time t, u is the velocity of the optical flow in the x direction, and v is the velocity of the optical flow in the y direction;
[0139] Image entropy is a statistic used to describe the information content of an image and is used to reflect the distribution of pixel values in the image. For a grayscale image, the formula for its entropy is:
[0140]
[0141] where L is the number of gray levels of the image (for example, for an 8-bit grayscale image, (L = 256)), and p i is the probability that a pixel with a gray value of i appears in the image; According to the calculation of image entropy and optical flow features, the present invention selects 29 facial key points with obvious movement change features during micro-expression occurrence to form a facial key point graph, as Figure 5 shown.
[0142] The facial key-point map structure can be expressed as In the facial key-point map, nodes represent the landmark points of the human face located such as the outer end of the eyebrow and the corner of the mouth, and are expressed as The total number of nodes is N; ε is the edge representing the explicit connection of key points, and can be represented by an adjacency matrix For the element a at the i-th row and j-th column of it i,j , if there is an edge connection between v i and v j , a i,j = 1, otherwise a i,j = 0; The expression recognition based on the facial key-point map uses such a set of graph sequences as input, and its key-point feature set can be represented by a tensor as where x t,n is the C-dimensional feature vector of the key point v n at the t-th moment in the T-frame facial key-point map. In the original data, it is usually the feature after global feature extraction processing or the original pixel value at the corresponding position; Thus, the graph sequence structure used in the present invention can be described by the node feature set X and the adjacency matrix A.
[0143] C. Construct a Global Information and Dynamic Spatio-Temporal Feature Extraction Network (GIDST); use a lightweight network to perform global feature extraction on the original input image sequence, and map the feature map integrating global information into the facial key-point structure through a key-point mapping network as node features, and enter step D. Use two modules of the network in step D to construct a spatio-temporal graph feature extraction network to obtain the features for calculating the classification loss and the triplet loss;
[0144] In step C, to construct a global information and dynamic spatio-temporal feature extraction network, the steps are as follows:
[0145] The present invention constructs a Global Information and Dynamic Spatio-Temporal Feature Extraction Network (GIDST), as Figure 2As shown in the figure, a lightweight network is used to extract global features of the expression sequence, and the feature map integrating global information is mapped into the facial key-point structure through the key-point mapping network as node features. Specifically, first, the pre-trained model MobileNet is used to extract features from the original input image sequence to obtain the global semantic information of the image; then, through the key-point mapping network, specifically implemented as node feature mapping based on the self-attention mechanism, the grid features of the global feature map are mapped into the node features of the facial key points; finally, the facial key-point map integrating global semantic information enters step D, and the attention-enhanced twin spatio-temporal graph fusion network proposed in step D of the present invention is further used to extract detailed information between nodes, thereby improving the recognition performance of the model for global information and detailed spatio-temporal motion features. Specifically, two modules of the network in step D are used to form a layer of feature extraction network for feature extraction, and two features for cross-entropy and triplet loss are obtained for global information.
[0146] Further, first, the pre-trained MobileNet model is used to model the global information of the input sample sequence in the spatial dimension, and the input RGB image sequence is input to obtain the feature map F. Map , and the feature map is changed into graph node features through the key-point mapping network. First, the attention mechanism is applied to weight the feature map to enable the network to adaptively focus on the important parts in the feature map and enhance the expression ability of the features. Then, the feature map passes through the convolutional network. In this way, each channel corresponds to a node, so as to associate the channel information of the feature map with the nodes; input feature map: C is the number of channels, and H and W are the height and width of the feature map respectively; query vector key vector value vector
[0147]
[0148] ⊙ represents element-wise multiplication, and the obtained Q, K, and V are transposed and flattened for calculating the attention scores. The formula is:
[0149]
[0150] att = Softmax(Scores) (7)
[0151]
[0152] The above process represents the following formula to obtain the attention-weighted x. att Channel-node feature mapping is performed through convolution, and node features of size 7*7 are obtained through pooling.
[0153] xatt = Attention(F Map ) (9)
[0154]
[0155] where w conv is the convolutional weight, mapping the channel features to the node dimension, Reshape() is the dimension adjustment operation, and Attention refers to calculating the input F Map through the process of formulas (3)-(8), and x node , x avg represent the features after node mapping and pooling respectively;
[0156] Finally, the obtained graph sequence X out enters step D. Through the attention-enhanced siamese spatio-temporal graph fusion network (AESSTG) constructed in step D, further spatio-temporal feature extraction is performed. After spatio-temporal pooling, it is output through two different fully connected layers (FCL), and the features for constructing the classification loss and the triplet loss are obtained respectively, completing the extraction of the global motion features.
[0157] D. Construct an attention-enhanced siamese spatio-temporal graph fusion network (Attention-Enhanced Siamese Spatio-Temporal Graphs Fusion Network, AESSTG), as Figure 3 , 4 shown; this network includes two layers of feature extraction networks, and each layer of feature extraction network includes a static spatio-temporal graph convolution module (Static Spatio-Temporal GraphConvolution Module, SSTGC) and a dynamic spatio-temporal graph feature fusion module (Dynamic Spatio-TemporalGraph Feature Fusion Module, DSTGFF); residual connections are added between each layer of feature extraction networks to make the model optimization easier. The node input feature of the network is the pixel feature of the fixed region size corresponding to the key point; after the features extracted by the network pass through spatio-temporal pooling, they are output through two different fully connected layers (FCL), and the features for constructing the classification loss and the triplet loss are obtained respectively, completing the extraction of the global motion features;
[0158] In step D, constructing an attention-enhanced siamese spatio-temporal graph fusion network includes the following steps:
[0159] For the input facial key graph sequence C, T, and N represent the number of channels, the number of frames, and the number of nodes respectively. The pixel features of facial key points are obtained through the facial key point data preprocessing module as the network input features of static spatio-temporal graph convolution. C in is the number of input feature channels;
[0160] The attention-enhanced twin spatio-temporal graph fusion network includes two layers of feature extraction networks. Each layer of the feature extraction network includes a static spatio-temporal graph convolution module (Static Spatio-Temporal Graph Convolution Module, SSTGC) and a dynamic spatio-temporal graph feature fusion module (Dynamic Spatio-Temporal Graph Feature Fusion Module, DSTGFF); Residual connections are added between each layer of the feature extraction network to make the model optimization easier.
[0161] The static spatial graph convolution layer is used to fuse the information between fixed adjacent nodes in the same frame. The original adjacency matrix represents the relationship between directly connected nodes and is denoted as Rich spatial features are extracted through the K-order adjacency matrix; for the k-hop adjacency matrix A k , the element a k(i,j) in its i-th row and j-th column is defined as:
[0162]
[0163] where dis(v i , v j ) represents the distance (shortest path length) between key points v i and v j in the skeleton graph, k ∈ {0,..., K - 1}; thus A0 = I, A1 = A; for After normalizing it, it is used as the k-hop structural adjacency matrix
[0164]
[0165] where D k represents the degree matrix of the k-hop link;
[0166]
[0167] where C out is the number of output feature channels, is the learnable parameter matrix, K represents the order of the k-hop adjacency matrix, F Sout is the result of the spatial graph convolution obtained, F in is the input feature, and σ(·) is the sigmoid activation function.
[0168] The static temporal convolutional layer processes the state information of the same node at different time steps, and outputs features with dimensions of T×N×C after passing through the static spatial graph convolutional layer out , where the size of the temporal convolutional kernel is m×1 and the stride is s. After performing convolution on the nodes at the same position in m frames, it jumps s frames along the temporal dimension and continues convolution until the convolution of the nodes with the same number is completed, and then proceeds to the convolution of the next numbered node;
[0169] The above process can be expressed by the formula:
[0170] F Tout = TCN(F Sout ) (16)
[0171] where TCN(·) represents temporal convolution, represent the input and output of the static temporal convolutional layer respectively.
[0172] The static spatio-temporal graph samples features from the facial key point graph based on a specific pattern, and it is difficult to learn the implicit link relationships between different nodes. For example, when an angry expression occurs, there is no explicit link between the corners of the mouth and the eyebrows in the fixed adjacency matrix, but there is a co-movement relationship during the actual movement of the expression, which is difficult to capture through the adjacency matrix of the static facial position relationship. Therefore, a dynamic spatio-temporal graph feature fusion module is designed for the generation of an adaptive adjacency matrix, which can capture the movement relationships between nodes.
[0173] On the basis of the operations of the static spatio-temporal graph convolution, the dynamic spatio-temporal graph feature fusion module introduces the generation of a feature-based dynamic adjacency matrix and performs convolution operations on the basis of the new adjacency matrix;
[0174] For the input feature F in , calculate the temporal dimension attention, node dimension attention, and feature channel dimension attention respectively, and obtain a multi-dimensional attention tensor through the broadcast mechanism;
[0175] The calculation process for the temporal dimension attention is as follows:
[0176] a t = σ(W time ·AvgPool (V,1) (x)) (17)
[0177] where, AvgPool (V,1) (·) is the average pooling function in the spatial dimension, V is the node dimension, W time is a one-dimensional convolution with a kernel size of 3, used to expand the temporal receptive field, and σ(·) is the Sigmoid activation function.
[0178] The calculation process of node dimension attention is as follows:
[0179] a s =σ(W spatial ·AvgPool (T) (Reshape(x))) (18)
[0180] Wherein, AvgPool (T) (·) is the average pooling function in the time dimension, is the learnable parameter matrix of the fully connected layer.
[0181] The calculation process of feature channel dimension attention is as follows:
[0182] a c =σ(W channel ·AvgPool (T,V) (x)) (19)
[0183] Wherein, AvgPool (T,V) (·) is the average pooling function in the time and space dimensions, is the learnable parameter matrices of two linear layers, r is the dimension transformation factor, set to 16;
[0184] a=σ(a t ·a s ·a c ) (20)
[0185] After calculating the multi-dimensional attention tensor through the broadcast mechanism, it is element-wise multiplied with the input features to obtain the enhanced features for subsequent convolution processing;
[0186] F a =F in ⊙a (21)
[0187] The obtained attention tensor is mapped to a low-dimensional space through two linear layers and undergoes dimension transformation to obtain a dynamically generated adjacency matrix. On the basis of retaining the key facial structures, the k-order adjacency matrix is added to and weighted with the output k-order matrix, and the weight is the learnable gating parameter, which is used to balance the stability of the adjacency matrix;
[0188] A d =γ[0]·A[0]+γ[k]·(W2(ReLU(W1(a)))+A[k]) (22)
[0189] Wherein, W2, W2 are the weights of the fully connected layer, A is the input adjacency matrix; γ is the gating parameter, which is used to control the weighted balance between the k-order matrix and the adaptive adjacency matrix; through the generated dynamic adjacency matrix Ad Perform spatio-temporal convolution operation with the enhanced feature F a to obtain the output F Gout .
[0190] The output layer includes a temporal pooling layer, a spatial pooling layer, and a fully connected layer. The output features of the last temporal feature extraction network perform temporal pooling and spatial pooling operations respectively. The pooling process can be expressed by the formula:
[0191] F T = MaxPool t (F Gout ) (23)
[0192] F S = MaxPool s (F T ) + AvgPool s (F T ) (24)
[0193] where represent the output features after temporal and spatial pooling respectively; MaxPool t (·) represents the maximum pooling in the temporal dimension, and MaxPool s (·) and AvgPool s (·) represent the maximum pooling and average pooling in the spatial dimension respectively; finally, pass F S through two different fully connected layers to obtain the features for calculating the cross-entropy loss and F CE and F Tri represent the feature outputs for calculating the cross-entropy and triplet losses respectively, and C class is the number of classes, that is, the number of emotion classes.
[0194] E. Construct a cross-domain joint loss, calculate the classification losses in the macro-expression domain and the micro-expression domain and the cross-domain triplet loss respectively, to optimize the global motion information, the spatio-temporal feature extraction network, and the attention-enhanced siamese spatio-temporal graph fusion network, so that the model can learn domain-invariant features, and the obtained final model is used for micro-expression recognition.
[0195] In step E, constructing the cross-domain joint loss includes the following steps:
[0196] The cross-domain joint loss combines the classification task and the cross-domain task, optimizes the network through the classification cross-entropy loss and the triplet loss, and calculates the cross-entropy loss for the global features local features obtained respectively through step C and step D
[0197]
[0198] Among them, y i is the true label, and p i is the predicted value.
[0199] Taking each epoch as the update unit, the global features local features obtained through step C and step D respectively are used to construct triplet losses
[0200] Specifically, taking each micro-expression sample as an anchor point, through clustering, randomly select samples with the same label that are closest to the sample as positive samples, and samples with different labels as negative samples to construct triplets; by calculating the Euclidean distances between the anchor point and the positive and negative samples, samples with the same label are made closer in the feature space, while samples with different labels are made farther apart, thereby improving the model's ability to distinguish features.
[0201] For the micro-expression anchor sample F micro , positive sample F POS and negative sample F NEG , its triplet loss can be calculated by the following formula:
[0202]
[0203] Among them, d(·) is the distance metric (such as Euclidean distance); m is a margin constant used to control the distance between positive and negative samples, represents the triplet loss calculated from the global features, represents the triplet loss calculated from the local features.
[0204] Calculate the triplet loss and the cross-entropy loss and use the average value of all branch losses to supervise the training process of the model; the above process can be expressed by the following formula:
[0205]
[0206] Among them, β is a hyperparameter used to balance the weights of the two losses.
[0207] Experimental Example 1
[0208] A micro-expression recognition method based on a cross-source dual-branch dynamic spatio-temporal graph convolutional network model. The steps are as described in Example 1. Micro-expression five-classification and three-classification experiments were respectively carried out on CASME II and SAMM, and ablation experiments on each module were carried out on the CASME II dataset. Figures 8 - 11Shows the confusion matrices of the corresponding datasets when each sample is cross-validated by the leave-one-out method on two datasets.
[0209] Ablation experiments of each module were carried out on the CASME II dataset. Figure 12 Shows the changes in the accuracy and F1 values of different modules missing the corresponding datasets when each sample is cross-validated by the leave-one-out method on the CASME II dataset. It can be seen that the present invention has high performance in facial expression recognition accuracy, and the proposed global information feature extraction network, dynamic spatio-temporal feature information fusion network, dynamic adjacency matrix, and cross-domain joint loss in the invention all play a key role in improving the model performance.
Claims
1. A micro-expression recognition method based on a cross-source dual-branch dynamic spatiotemporal graph convolutional network model, characterized in that: The steps include: A. Preprocessing of micro-expression video sequences, including: obtaining video frame sequences, face detection and positioning, face alignment, and facial image cropping; B. Construct facial key point map; C. Construct a global motion information and spatiotemporal feature extraction network; use a lightweight network to extract global features from the original input image sequence, map the feature map that integrates the global information to the facial key point structure through the key point mapping network as the node feature, and enter step D to use the two modules of the network in step D to construct a spatiotemporal graph feature extraction network to obtain features for calculating classification loss and triplet loss; D. Construct an attention-enhanced twin spatiotemporal graph fusion network; the network includes two layers of feature extraction networks, each of which includes a static spatiotemporal graph convolution module and a dynamic spatiotemporal graph feature fusion module; residual connections are added between each layer of feature extraction networks, and the node input features of the network are pixel features of a fixed area size corresponding to the key points; the features extracted by the network are output through two different fully connected layers after spatiotemporal pooling, and the features used to construct classification loss and triplet loss are obtained respectively, completing the extraction of global motion features; E. Construct a cross-domain joint loss, calculate the classification loss of the macro-expression domain and the micro-expression domain, and the cross-domain triplet loss respectively, which are used to optimize the global motion information and spatiotemporal feature extraction network and the attention-enhanced twin spatiotemporal graph fusion network, so that the model can learn domain-invariant features, and the final model is used for micro-expression recognition.
2. The micro-expression recognition method based on the cross-source dual-branch dynamic spatiotemporal graph convolutional network model according to claim 1 is characterized in that: In step A, the micro-expression video sequence is preprocessed, including the following steps: 1) Obtaining a video frame sequence: performing frame processing on the micro-expression video sequence to obtain and store a video frame sequence; 2) Obtaining the start frame, peak frame, and end frame: According to the information annotated by experts in the micro-expression dataset, the micro-expression start frame, peak frame, and end frame in each video sequence file are selected, and sampled within the three-frame special frame interval to obtain a frame sequence length of 8; The starting frame refers to: the first frame in the micro-expression video sequence where the micro-expression appears; The peak frame refers to the frame in the micro-expression video sequence where the facial muscle changes are most obvious and contains the most micro-expression information; The end frame refers to the last frame in the micro-expression video sequence where the micro-expression ends; 3) Face detection and positioning: Use the Dlib library to perform face detection and positioning on the frame sequence obtained in step 2), and detect the number of faces in the video frame and the distance between the face and the image boundary; 4) Face alignment: Based on face positioning, the Dlib library is used to determine 68 key facial feature points to complete face segmentation, face correction and face image cropping; Face segmentation refers to: using the Dlib visual library to segment faces using rectangular boxes; Face correction means: among the 68 key feature points detected on the face, the line connecting the key feature points marking the left corner of the left eye and the key feature points marking the right corner of the right eye is at an angle a with the horizontal line. The corresponding rotation matrix is obtained through the angle a, and the segmented face is rotated to make the line connecting the key feature points marking the left corner of the left eye and the key feature points marking the right corner of the right eye parallel to the horizontal line, so as to achieve correction of the face posture; Facial image cropping means scaling the segmented and corrected face to obtain an image of a fixed size.
3. The micro-expression recognition method based on the cross-source dual-branch dynamic spatiotemporal graph convolutional network model according to claim 2 is characterized in that: In step 4), the fixed size of the image obtained when the face image is cropped is set to 128*128 pixels.
4. The micro-expression recognition method based on the cross-source dual-branch dynamic spatiotemporal graph convolutional network model according to claim 1 is characterized in that: In step B, the selection and preprocessing of the facial key point map are as follows: A total of 29 points, including 12 points of the mouth contour, 5 points of the left eyebrow, 5 points of the right eyebrow, 4 points of the nose, and 3 points of the face contour, were selected as nodes of the graph, and the initial adjacency matrix was constructed according to the facial structure relationship; The Farneback method is used to calculate the optical flow features of the image sequence, the large displacement is processed by layering the image pyramid, and the optical flow field is approximated by polynomial expansion. In the local window, the image brightness is expressed as a quadratic polynomial: I(x, y, t) = a + bx + cy + dx 2 +exy+fy 2 , abcdef are all constants. For adjacent frames I1 and I2, solve the displacement (u, v) so that: I2(x+u,y+v,t+1)=I1(x,y,t). The optical flow is estimated by minimizing the following error function. This process is expressed as formula (1): ∑x,y(I2(x+u,y+v)-I1(x,y)) 2 +λ∑x,y(▽u 2 +▽v 2 ) (1) Among them, λ is the smoothing weight, I(x,y,t) represents the brightness of the image at the (x,y) position and time t, and I x ,I y Represents the gradient of the image in the x and y directions, I t represents the rate of change of the image at time t, u is the speed of the optical flow in the x direction, and v is the speed of the optical flow in the y direction; Image entropy is a statistic used to describe the information content of an image. It is used to reflect the distribution of pixel values in an image. For a grayscale image, the entropy calculation formula is: Among them, L is the gray level of the image, p i is the probability of a pixel with gray value i appearing in the image; based on the calculation of image entropy and optical flow features, 29 facial key points are selected to form a facial key point map; The facial key point graph structure is represented as In the facial key point graph, the nodes represent the located facial landmark points, expressed as The total number of nodes is N; ε is the edge that represents the explicit connection of key points, using the adjacency matrix It means that for the element a in the i-th row and j-th column i,j , if v i and v j There are edges connecting the i,j =1, otherwise a i,j =0; facial key point graph-based expression recognition uses such a set of graph sequences as input, and its key point feature set The tensor is represented as where x t,n is the key point v at time t in the facial key point map of frame T n C-dimensional feature vector; therefore, the graph sequence structure used in the present invention is described by the node feature set X and the adjacency matrix A.
5. The micro-expression recognition method based on the cross-source dual-branch dynamic spatiotemporal graph convolutional network model according to claim 1 is characterized in that: In step C, a global motion information and spatiotemporal feature extraction network is constructed, including the following steps: First, the pre-trained model MobileNet is used to extract features from the original input image sequence, and the global semantic information of the image is obtained through the pre-trained model; then, the key point mapping network is used, which is specifically implemented as a node feature mapping based on the self-attention mechanism, to map the grid features of the global feature map to the node features of the facial key points; finally, the facial key point map that is fused with the global semantic information enters step D, and the attention-enhanced twin spatiotemporal graph fusion network proposed in step D is used to further extract detailed information between nodes. The two modules of the network in step D are used to form a feature extraction network for feature extraction, and the global information is used to obtain two features of cross entropy and triplet loss.
6. The micro-expression recognition method based on the cross-source dual-branch dynamic spatiotemporal graph convolutional network model according to claim 5 is characterized in that: First, the pre-trained MobileNet model is used to model the global information of the input sample sequence from the spatial dimension, and the input RGB image sequence is input to obtain the feature map F Map , the feature map is transformed into graph node features through the key point mapping network. First, the attention mechanism is applied to weight the feature map, and then the feature map is passed through the convolutional network. In this way, each channel corresponds to a node, thereby associating the channel information of the feature map with the node; input feature map: C is the number of channels, H and W are the height and width of the feature map respectively; the query vector Key Vector Value vector ⊙ represents element-by-element multiplication. The obtained Q, K, and V are flattened by transposition to calculate the attention score. The formula is: att=Softmax(Scores) (7) The above process is expressed as the following formula to obtain the attention-weighted x att Channel-node feature mapping is performed through convolution, and node features of size 7*7 are obtained through pooling; x att =Attention(F Map ) (9) Among them, w conv is the convolution weight, mapping the channel features to node dimensions, Reshape() is the dimension adjustment operation, and Attention refers to converting the input F Map Through the process of formula (3)-(8), x node , x avg Respectively represent the features after node mapping and pooling; Finally, the obtained graph sequence X out Go to step D, and perform further spatiotemporal feature extraction through the attention-enhanced twin spatiotemporal graph fusion network constructed in step D. After spatiotemporal pooling, the output is passed through two different fully connected layers to obtain features for constructing classification loss and triplet loss respectively, thereby completing the extraction of global motion features.
7. The micro-expression recognition method based on the cross-source dual-branch dynamic spatiotemporal graph convolutional network model according to claim 1 is characterized in that: In step D, an attention-enhanced twin spatiotemporal graph fusion network is constructed, including the following steps: For the input facial key map sequence C, T, and N are the number of channels, frames, and nodes respectively. The pixel features of facial key points are obtained through the facial key point data preprocessing module as the network input features of the static spatiotemporal graph convolution. C in is the number of input feature channels; The attention-enhanced twin spatiotemporal graph fusion network includes two layers of feature extraction networks, each of which includes a static spatiotemporal graph convolution module and a dynamic spatiotemporal graph feature fusion module; Residual connections are added between each layer of feature extraction network; The static spatial graph convolution layer is used to fuse the information between fixed adjacent nodes in the same frame. The original adjacency matrix represents the relationship between directly connected nodes and is recorded as Extract rich spatial features through K-order adjacency matrix; For the k-th hop adjacency matrix A k , the element a in the i-th row and j-th column k(i,j) Defined as: Among them, dis(v i ,v j ) represents the key point v in the skeleton graph i and v j The shortest path length between them, k∈{0,...,K-1}; therefore A0=I, A1=A; After normalization, it is used as the k-th hop structure adjacency matrix Among them, D k represents the degree matrix of k-order links; Among them, C out is the number of output feature channels, is a learnable parameter matrix, K represents the order of the k-order adjacency matrix, F Sout is the result of the spatial graph convolution, F in is the input feature, σ(·) is the sigmoid activation function; The static time convolution layer processes the state information of the same node at different time steps, and the static spatial graph convolution layer outputs feature dimensions of T×N×C out , the time convolution kernel size is m×1, the step length is s, after completing the convolution of the node at the same position on the m frame, jump s frames along the time dimension to continue the convolution until the convolution of the node with the same number is completed, and then perform the convolution of the next numbered node; The above process is expressed as follows: F Tout =TCN(F Sout ) (16) Among them, TCN(·) represents temporal convolution, Represent the input and output of the static temporal convolutional layer respectively; The dynamic spatiotemporal graph feature fusion module introduces the generation of a feature-based dynamic adjacency matrix based on the convolution operation of the static spatiotemporal graph, and performs convolution operations on the basis of the new adjacency matrix. For the input feature F in , respectively calculate the time dimension attention, node dimension attention and feature channel dimension attention, and obtain the multi-dimensional attention tensor through the broadcast mechanism; The calculation process for time dimension attention is as follows: a t =σ(W time ·AvgPool (V,1) (x)) (17) in, AvgPool (V,1) (·) is the average pooling function in the spatial dimension, V is the node dimension, and W time is a one-dimensional convolution with a kernel size of 3, which is used to expand the temporal receptive field, and σ(·) is the Sigmoid activation function; The calculation process for node dimension attention is as follows: in, AvgPool (T) (·) is the average pooling function in the time dimension, is the learnable parameter matrix of the fully connected layer; The calculation process for feature channel dimension attention is as follows: a c =σ(W channel ·AvgPool (T,V) (x)) (19) in, AvgPool (T,V) (·) is the average pooling function in time and space dimensions, are the learnable parameter matrices of the two linear layers, r is the dimension transformation factor, which is set to 16; a=σ(a t ·a s ·a c ) (20) The multi-dimensional attention tensor is calculated through the broadcast mechanism and then multiplied element by element with the input feature to obtain the enhanced feature for subsequent convolution processing; F a =F in ⊙a (21) The obtained attention tensor is mapped to a low-dimensional space through two linear layers and transformed into a dynamically generated adjacency matrix. On the basis of retaining the key structure of the face, the k-order adjacency matrix is added to the output k-order matrix and then weighted. The weight is a learnable gating parameter used to balance the stability of the adjacency matrix. AND d =γ[0] A[0]+γ[k] (W2(ReLU(W1(a)))+A[k]) (22) Among them, W2, W2 are the weights of the fully connected layer, A is the input adjacency matrix; γ is the gating parameter used to control the weighted balance between the k-order matrix and the adaptive adjacency matrix; the dynamic adjacency matrix A generated by d With the enhanced feature F a Perform spatiotemporal convolution operation to obtain the output F Gout ; The output layer includes a time pooling layer, a spatial pooling layer, and a fully connected layer. The last layer extracts the output features of the network. Temporal pooling and spatial pooling operations are performed separately, and the pooling process is expressed by the formula: F T =MaxPool t (F Gout ) (23) F S =MaxPool s (F T )+AvgPool s (F T ) (24) in, Respectively represent the output features after time and space pooling; MaxPool t (·) represents the maximum pooling in the time dimension, MaxPool s (·) and AvgPool s (·) represent the maximum pooling of the spatial dimension and the average pooling of the spatial dimension respectively; finally, F S Through two different fully connected layers, the features used to calculate the cross entropy loss are obtained and F CE and F Tri They represent the feature outputs used to calculate the cross entropy and triplet loss, respectively. class is the number of categories, that is, the number of emotion categories.
8. The micro-expression recognition method based on the cross-source dual-branch dynamic spatiotemporal graph convolutional network model according to claim 1 is characterized in that: In step E, a cross-domain joint loss is constructed, including the following steps: Cross-domain joint loss combines the classification task with the cross-domain task, and optimizes the network through classification cross entropy loss and triple loss. The global features obtained in step C and step D are Local features Calculating cross entropy loss Among them, y i is the true label, p i is the predicted value; Taking each epoch as the update unit, the global features obtained by step C and step D are Local features Construct triplet losses separately Specifically, each micro-expression sample is used as an anchor point. Through clustering, the sample with the same label that is closest to the sample is randomly selected as the positive sample, and the sample with different labels is selected as the negative sample to construct a triplet. By calculating the Euclidean distance between the anchor point and the positive sample and the negative sample, the samples with the same label are closer in the feature space, and the samples with different labels are farther away from each other. For micro-expression anchor sample F micro , positive sample F POS and negative samples F NEG , whose triplet loss is calculated as follows: Among them, d(·) is the distance metric; m is a margin constant used to control the distance between positive and negative samples. represents the triplet loss calculated by the global features, represents the triplet loss calculated by local features; Calculating triplet loss and cross entropy loss And use the average of all branch losses to supervise the model training process; the above process can be expressed as follows: Among them, β is a hyperparameter used to balance the weights of the two losses.
Citation Information
Cited By
Facial micro-expression recognition method
CN120783379A
Emotion detection method based on spatial-temporal feature fusion of facial key points
CN120877354A
An emotion detection method based on facial key point space-time feature fusion
CN120877354B
Photovoltaic array fault diagnosis and positioning method and system based on digital twinning and deep learning
CN121352769A
Micro-expression analysis method based on facial key point recognition
CN121640549A