Micro-expression recognition method based on staged adaptive course learning

Through the phased adaptive curriculum learning method, the noise interference problem in the micro-expression dataset was solved, the neural network training was optimized, and the accuracy of micro-expression recognition and the generalization ability of the model were improved.

CN120599676APending Publication Date: 2025-09-05SHANDONG UNIV +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510664683.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-05

Smart Images

  • Figure CN120599676A_ABST
    Figure CN120599676A_ABST
Patent Text Reader

Abstract

The invention relates to a micro-expression recognition method based on staged adaptive course learning. The method comprises the following steps: A, preprocessing a micro-expression video sequence and a macro-expression video sequence; b, constructing a spatial-temporal feature fusion model, performing deep feature extraction on the micro-expression data set obtained by preprocessing, and pre-training a macro-expression recognition teacher model; c, constructing a deep learning algorithm based on staged adaptive curriculum learning, and introducing a micro-expression recognition-oriented deep learning algorithm based on staged adaptive curriculum learning in the training process of the constructed spatio-temporal feature fusion model to optimize the training process; and D, carrying out classification identification on the macro expression identification teacher model obtained by training on a test set. According to the method, more effective and discriminative micro-expression features are obtained, the generalization ability of the model is improved, and the problems that in the existing micro-expression recognition field, available data sets are lacked, redundant information contained in the data sets is large, and the recognition accuracy is not high are further solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a micro-expression recognition method based on stage-based adaptive course learning, and belongs to the technical field of deep learning and pattern recognition. Background Art

[0002] Microexpressions are spontaneous and unconscious facial expressions of emotion that are short-lived, triggering small, low-intensity facial muscle movements, in stark contrast to more common macroexpressions. Because microexpressions can reveal a person's mental state and true thoughts, they have considerable application value in monitoring the true psychological activities of target individuals. They have been widely used in fields such as national security, public safety, and psychotherapy, and hold immense practical value for specific groups such as psychologists and public security criminal investigators. However, microexpression datasets inevitably contain samples with significant noise. This noise includes motion information caused by the subject's facial shaking, interference caused by occlusion and shaking of objects such as glasses, clothing, and hair, and noise points introduced by the recording equipment itself, posing a challenge to microexpression recognition.

[0003] In recent years, computer-based micro-expression recognition technologies and methods have achieved considerable success. Currently, automatic micro-expression recognition algorithms can be primarily categorized into two categories: traditional algorithms and deep learning-based algorithms. Within the traditional algorithm category, local binary patterns based on three orthogonal planes, used to describe dynamic textures, were first applied to micro-expression recognition. Meanwhile, the principal direction mean optical flow (MDMO) proposed by Liu et al. is also a representative spatiotemporal description method, classified as an optical flow method. Subsequently, Liu et al. further integrated MDMO into classic graph regularized sparse coding to generate sparse MDMO features. Furthermore, Xu et al. proposed the facial dynamics map (FDM) feature. This feature utilizes an iterative optimization strategy to calculate the principal optical flow directions of a spatiotemporal cuboid obtained from segmented micro-expression sequences, thereby better representing local facial dynamics. Meanwhile, Happy et al. proposed the fuzzy histogram of optical flow directions (FHOFO). This method uses histogram fuzzification to construct an angle histogram based on the direction of the optical flow vector, thereby encoding the temporal pattern characteristics of micro-expressions. In addition, Liong proposed a new optical strain weighted feature extraction scheme, which extracts weighted spatiotemporal information from the block optical flow map.

[0004] Deep learning, a key technology in the field of artificial intelligence, has flourished in recent years, achieving significant results in numerous fields. Micro-expression recognition, a hot research area in computer vision, has continued to advance with the help of deep learning. In 2016, Dae Hoe Kim developed a micro-expression recognition method that combines CNN and LSTM, capable of simultaneously extracting spatial and temporal features from video sequences. In 2018, Wang proposed a micro-expression recognition approach using transfer learning strategies. In 2020, Yante Li constructed a micro-expression peak frame recognition algorithm based on a three-dimensional Fourier transform and built the LGCcon model. This model extracts and fuses local and global features to train the network. Xia used Euclidean video magnification technology to pre-process the dataset, aiming to amplify the dynamic information of micro-expressions. He also designed a masking algorithm to extract information from key facial areas, significantly reducing the data dimensionality. In terms of neural network design, Xia used a recursive convolutional neural network (RCNN) with the ability to extract spatiotemporal features. Xie proposed using facial action units (AUs) to assist in micro-expression recognition. In 2021, Ben conducted a comprehensive and in-depth investigation and analysis of the field of micro-expression detection and recognition based on video datasets, and provided an outlook on its future development prospects. Considering the scarcity of datasets in the micro-expression field, Ben proposed a new dataset, MMEW. In 2022, Chen proposed a micro-expression recognition algorithm that only uses the start and peak frames. He combined optical flow features extracted from four different methods and designed a block-wise convolutional neural network, effectively optimizing the optical flow feature extraction process. Zhao et al. first proposed a method for micro-expression recognition using a visual Transformer model and demonstrated its successful application to micro-expression recognition tasks. Mao et al. explored the task of micro-expression recognition when the face is occluded by objects in real-world settings and proposed a Region-Reflected Relational Reasoning Network (RRRN) to capture the complementary relationships between facial regions. Wei designed a novel attention-based magnification adaptive network (AMAN) that can adapt to different magnification levels during dynamic changes in micro-expression sequences. In 2023, Nguyen et al. proposed a micro-expression recognition framework based on the BERT model (Micron-BERT). Leveraging the BERT framework's powerful advantages in sequence modeling, Micron-BERT can accurately locate blocks of interest within micro-expression frames. This approach effectively reduces the negative impact of interference factors such as background noise on the model's recognition performance, significantly improving the accuracy and reliability of micro-expression recognition. In 2024, Bao et al. proposed an innovative method that simultaneously integrates supervised prototype-based memory contrast learning to mine discriminative features and adds self-expression reconstruction as an auxiliary task and regularization method, effectively improving the recognition accuracy of micro-expressions.

[0005] However, micro-expression datasets inevitably contain some samples with high noise content. This noise information covers many aspects, such as motion information generated by the subject's facial shaking, interference caused by occlusion and shaking of objects such as glasses, clothing, and hair, and noise points generated by the recording equipment itself. Deep learning-based models have powerful feature extraction capabilities but are easily affected and interfered with by noise. The presence of this noise poses many challenges to micro-expression recognition. Summary of the Invention

[0006] In view of the shortcomings of the existing technology, the present invention provides a micro-expression recognition method based on staged adaptive curriculum learning. Summary of the invention:

[0008] The present invention aims to address the existing problem of noise samples in datasets interfering with feature learning in the field of micro-expression recognition. It provides a micro-expression recognition method based on staged adaptive curriculum learning, comprising: dataset preprocessing, neural network construction, and a micro-expression recognition algorithm based on staged adaptive curriculum learning. By using the staged adaptive curriculum learning method during deep neural network training, the present invention designs a new micro-expression recognition algorithm that effectively reduces the impact of noise in the dataset, optimizes the neural network training process, and enables the network to extract more effective micro-expression features from limited samples.

[0009] Explanation of terms:

[0010] 1. Dlib Visual Library: Dlib is an open-source C++ toolkit that includes machine learning algorithms and can be used to solve many practical problems in machine learning. Dlib has been widely used in industry and academia.

[0011] 2. Facial key feature point detection: 68 key feature points on the face are mainly distributed in eyebrows, eyes, nose, mouth and facial contours, such as Figure 1 As shown, detection is performed using the Dlib visual library, which is an existing technology.

[0012] 3. Curriculum Learning: When training a model, we gradually transition from simple examples to more difficult examples, aiming to simulate the cognitive process of the human brain. Theoretical analysis has shown that this can effectively enhance the model's generalization ability.

[0013] 4. Farneback algorithm: an optical flow calculation method that is widely used in the field of computer vision.

[0014] The technical solutions of the present invention are as follows:

[0015] A micro-expression recognition method based on staged adaptive curriculum learning, comprising:

[0016] A. Preprocessing of micro-expression and macro-expression video sequences, including: obtaining video frame sequences, face detection and positioning, and face alignment;

[0017] B. Build a spatiotemporal feature fusion model, perform deep feature extraction on the micro-expression dataset preprocessed in step A, and pre-train a macro-expression recognition teacher model;

[0018] C. Construct a deep learning algorithm based on staged adaptive curriculum learning. Introduce a deep learning algorithm based on staged adaptive curriculum learning for micro-expression recognition into the training process of the spatiotemporal feature fusion model constructed in step B to optimize the training process.

[0019] D. Classify and recognize the macro expression recognition teacher model trained in step C on the test set.

[0020] According to the preferred embodiment of the present invention, in step A, pre-processing the micro-expression and macro-expression video sequences includes the following steps:

[0021] 1) Obtaining a video frame sequence: performing frame processing on the video sequence to obtain and store the video frame sequence;

[0022] 2) Face detection and localization: Use the Dlib visual library to perform face detection and localization on each frame obtained in step 1), detecting the number of faces in the video frame and the distance between the face and the image boundary;

[0023] 3) Face alignment: Use the Dlib visual library to determine 68 key facial feature points to complete face segmentation and face correction;

[0024] Face segmentation refers to: using the Dlib visual library to segment faces using rectangular boxes;

[0025] Face correction means: among the 68 key feature points detected on the face, the line connecting the key feature point 37 marked at the left corner of the left eye and the key feature point 46 marked at the right corner of the right eye is at an angle a with the horizontal line. The corresponding rotation matrix is ​​obtained through angle a, and the segmented face is rotated to make the line connecting the key feature point 37 marked at the left corner of the left eye and the key feature point 46 marked at the right corner of the right eye parallel to the horizontal line, thereby correcting the facial posture and scaling the face.

[0026] Preferably, according to the present invention, in step B, the spatiotemporal feature fusion model structure includes a spatial appearance feature extraction branch and a temporal subtle dynamic feature extraction branch; in the spatiotemporal feature fusion model, first, the downsampled peak frame image is sent to the spatial appearance feature extraction branch to extract the spatial appearance feature; secondly, the absolute difference between the starting frame and the peak frame is sent to the temporal subtle dynamic feature extraction branch to extract the temporal subtle dynamic feature; finally, the obtained spatial appearance feature and temporal subtle dynamic feature are fused to obtain the final micro-expression feature, which will be sent to the classification layer for micro-expression category classification.

[0027] Preferably, according to the present invention, the downsampled peak frame image is sent to the spatial appearance feature extraction branch to extract the spatial appearance features; including:

[0028] First, the peak frame of the micro-expression is downsampled to obtain a peak frame apex of size 28×28×3. i :

[0029]

[0030] Among them, i represents the i-th sample, is the real number domain, downsample is the downsampling operation; Refers to the peak frame of the i-th sample, apex refers to the peak frame, that is, a single frame image containing the richest micro-expression features in a micro-expression sample video sequence;

[0031] The downsampled image is expanded into a 784×3 tensor in width and height dimensions, and the primary features are extracted through the facial pixel block embedding layer FaceEmbed:

[0032]

[0033] The facial pixel block embedding layer is a two-dimensional convolution operation with 3 input channels, 768 output channels, a convolution kernel size of 1×1, and a stride of 1. i refers to the peak frame image after downsampling obtained by the operation in formula (1); flatten() refers to the expansion operation, which is to expand the downsampled image into a tensor of size 784×3 in the width and height dimensions; z i is the expanded 784×3 tensor; FaceEmbed() refers to the facial pixel block embedding layer operation, that is, a two-dimensional convolution operation with 3 input channels, 768 output channels, a convolution kernel size of 1×1, and a stride of 1; is the extracted primary feature;

[0034] After obtaining the primary features, add position encoding:

[0035]

[0036] Among them, E pos is the position code, Refers to position encoding and Primary features obtained by element-by-element addition;

[0037] Next, the primary features pass through three stacked Transformer encoders (TEncoders) with an embedding vector dimension of 512, and then through a multi-layer perceptron (MLP) with an output dimension of 512 to extract spatial appearance features:

[0038]

[0039] Among them, TEncoder1, TEncoder2, and TEncoder3 are Transformer encoders with the same structure but no shared parameters; MLP is a multi-layer perceptron, that is, a fully connected layer with an output dimension of 512; Refers to the spatial appearance features finally extracted by the spatial appearance feature extraction branch.

[0040] Preferably, according to the present invention, the absolute difference between the start frame and the peak frame is fed into a temporal subtle dynamic feature extraction branch to extract the temporal subtle dynamic feature; including:

[0041] The temporal subtle dynamic feature extraction branch uses the absolute difference D between the starting frame and the peak frame as input:

[0042]

[0043] Among them, D i Refers to the absolute difference between the starting frame and the peak frame of the i-th sample; refer to the peak frame and starting frame of the i-th sample respectively;

[0044] A block-based convolutional neural network is used to extract features from the absolute difference D;

[0045] First, the absolute difference map of size H×H×3 is evenly divided into pixel blocks of size d×d×3 and tiled in order from top to bottom and from left to right into a list:

[0046]

[0047] Among them, D list i It refers to the list obtained by evenly dividing the absolute difference map of size H×H×3 into pixel blocks of size d×d×3 and tiling them in order from top to bottom and from left to right. kRefers to the kth pixel block, H refers to the height and width of the absolute difference map, and d refers to the height and width of each pixel block obtained after blocking;

[0048] Next, each pixel block will pass through the subtle dynamic feature perception layer SDFL to obtain K feature tensors and splice them in the channel dimension:

[0049] F list i = [ f t1 ; f t2 ;…; f tK ] = [SDFL(block1);SDFL(block2);…;SDFL(block K )] (8);

[0050]

[0051] SDFL includes two-dimensional convolutional layers and maximum pooling; f tk It refers to the feature tensor obtained after the k-th pixel block passes through the subtle dynamic feature perception layer SDFL; F list i Refers to the K features obtained by Arrange the list in order; dim=1 means the concatenation operation is performed on the first dimension; concat means the concatenation operation; Refers to the features obtained after the splicing operation;

[0052] Finally, the total feature tensor The time-subtle dynamic features will be obtained through a mapping layer ML consisting of a two-dimensional convolution layer and a multi-layer perceptron

[0053]

[0054] After obtaining the spatial appearance features and temporal subtle dynamic features, the spatial appearance features and temporal subtle dynamic features are fused. For the i-th sample, first, the query vector q and key vector k are extracted from the spatial appearance features extracted from the spatial appearance feature extraction branch, while the value vector v is extracted from the temporal subtle dynamic features:

[0055]

[0056]

[0057] Among them, fc q 、fc k 、fc vThey are all fully connected layers with input and output dimensions of 512, but the parameters are not shared. The subscripts q, k, and v represent the three used to extract the query vector q, key vector k, and value vector v respectively; q i 、k i 、v i Respectively represent the query vector, key vector, and value vector extracted from the i-th sample;

[0058] Next, the correlation strength between the query vector and the corresponding key vector is calculated and normalized using the Softmax function to obtain the spatial correlation matrix:

[0059]

[0060] Among them, T represents the transposition operation, and the Softmax function is normalized along the direction of each column; the spatial correlation matrix M i Each element in represents the degree of attention that should be paid to the dynamic features of different positions, and v represents the dynamic feature information corresponding to each position;

[0061] Finally, the spatial correlation matrix is ​​multiplied by the temporal dynamic eigenvalue vector to obtain the spatiotemporal features weighted by the attention mechanism, and then the overall scaling is performed by the learnable weight parameter ρ and the original temporal subtle dynamic features are added element by element to obtain the final spatiotemporal fusion feature f i :

[0062]

[0063] Finally, the obtained spatiotemporal fusion features are sent to the classification layer for final expression category classification. The classification layer includes a layer normalization layer and a fully connected layer whose output dimension is the number of classification task categories.

[0064] Preferably, according to the present invention, the macro expression recognition teacher model, i.e., the macro expression model, is trained separately using the CK+ macro expression data set. The architecture of the macro expression model is a spatiotemporal feature fusion model, which uses the first frame of each macro expression video sequence as the starting frame, and all other frames are used as the peak frames of the macro expression, to help the macro expression model fully learn the spatiotemporal related features of the macro expression of the human face. The cross entropy loss function is used during the pre-training of the macro expression model.

[0065] The macro expression model will assist in the initial training of the micro expression model, using KL divergence as the loss function for the macro expression model to guide the micro expression model:

[0066]

[0067] Among them, C is the number of expression categories, a=[a1,a2,…,a C] is the predicted output of the macro expression model, b=[b1,b2,…,b C ] is the predicted output of the micro-expression model; L g refers to the KL divergence;

[0068] Use the cross entropy loss function L CE As the classification loss of the micro-expression model:

[0069]

[0070] in, is the predicted probability distribution of the micro-expression model, y k is the true label of the sample, C is the number of categories; in the early stage of micro-expression model training, the total loss function is:

[0071] L CE +αL g (18);

[0072] Where α is the intensity coefficient of the guided loss;

[0073] Total loss function L total for:

[0074]

[0075] Preferably, in step C, a deep learning algorithm based on staged adaptive curriculum learning for micro-expression recognition is introduced into the training process of the spatiotemporal feature fusion model constructed in step B to optimize the training process; the steps include:

[0076] Use the Dlib visual library 68-point face detection algorithm to perform mask extraction on the left eye area, right eye area and mouth area; calculate the average optical flow field intensity amplitude mag of the left eye area, right eye area and mouth area i :

[0077]

[0078] Where M(x,y) represents the optical flow amplitude obtained at the pixel point (x,y) according to the Farneback algorithm, W is the image width, H is the image height, i represents the i-th micro-expression sample; the apparent motion pattern difficulty score of the i-th sample is It is defined as the inverse of the average optical flow field intensity amplitude of the sample:

[0079]

[0080] For the i-th micro-expression sample, the difficulty score of the macro-expression model evaluation is defined as:

[0081]

[0082] in, is the predicted probability distribution of the macro expression model for the i-th micro expression sample, is the true label of the i-th micro-expression sample, and C is the number of categories;

[0083] For the i-th micro-expression sample, the initial difficulty score before the micro-expression model training begins It is defined as the sum of the difficulty score of the apparent movement pattern and the difficulty score of the macro expression model evaluation:

[0084]

[0085] Among them, the subscript 1 represents the first round, represents the difficulty score of the i-th sample before the start of the t-th round of training, 1≤t≤EP, t∈Z; 1≤i≤N, i∈Z, Z represents the integer domain, and N is the total number of samples in the dataset;

[0086] Through the adaptive course adjustment strategy, the difficulty scores of all samples are re-evaluated after each training round, and the new difficulty information is updated to the difficulty score of the sample through the momentum-based difficulty value update function; the adaptive course adjustment strategy is shown in the following equations (24) to (26):

[0087] The cross entropy loss is shown in formula (24):

[0088]

[0089] in, It represents the loss of the parameters of the current micro-expression model corresponding to the i-th sample after the end of the t-th training round. is the predicted probability distribution of the micro-expression model for the i-th sample, is the true label of the i-th sample, C is the number of categories; after the t-th training round, the update difficulty score of the i-th sample is the loss value of the sample for the micro-expression model state:

[0090]

[0091] in, It refers to the loss value of the i-th sample for the current micro-expression model state. refers to the update difficulty score of the i-th sample;

[0092] The momentum method is used to update the difficulty score of the sample. For the i-th sample, the difficulty score at the beginning of the t+1-th training round depends on the difficulty score of the t-th round:

[0093]

[0094] in, They refer to the difficulty scores of the i-th sample in the t-th training round and the t+1-th training round respectively; ω is the momentum coefficient, which is used to balance the weight between historical difficulty information and current difficulty information; before calculating the sampling probability in each round, the difficulty score is normalized by Min-Max:

[0095]

[0096] Among them, min(score t ) represents the minimum difficulty score of all samples before the start of round t, max(score t ) represents the maximum difficulty score of all samples before the start of round t. The normalized difficulty score will be used to calculate the sampling probability. Refers to the normalized difficulty score obtained after Min-Max normalization;

[0097] The step function controls the size s of the sampling pool. The sampling pool size s refers to the sample set consisting of the first s samples in the entire dataset queue. The dataset queue refers to the sample index queue obtained by sorting the samples from small to large according to their difficulty scores, that is, the first s samples in the sample queue after the entire dataset is sorted from easy to difficult.

[0098] Use exponential function as step function s(t):

[0099]

[0100] Where N is the size of the original dataset, t is the number of training rounds, r0 controls the initial size, K and T jointly control the step rate, and min() is the minimum function. Suppose all samples of the original micro-expression dataset are arranged in order:

[0101]

[0102] Before the start of round t of training, the difficulty score queue is:

[0103]

[0104] Sort from small to large to get a sequential queue:

[0105]

[0106] Get the index queue:

[0107] index = {k1,k2,k3,…,k N} (32);

[0108] Take the samples corresponding to the first s indexes of the index queue to form a sampling pool:

[0109]

[0110] Dynamic sampling probability mapping function Defined as:

[0111]

[0112] in:

[0113]

[0114] When t=1, focus on learning simple samples:

[0115]

[0116] Where i represents the i-th sample, t represents that the sampling probability is calculated before the start of the t-th training round, t∈Z, 1≤t≤EP, Z is an integer domain, EP represents the total number of training rounds, and e is a natural constant;

[0117] Finally, the sampling pool is obtained by combining the sampling pool size s calculated by the step function. For the set of s samples in the sampling pool

[0118]

[0119] Set The difficulty scores corresponding to the s samples in :

[0120]

[0121] Substituting into formula (38) respectively, the sampling probability values ​​obtained are:

[0122]

[0123] Normalize the sampling probability values:

[0124]

[0125] Get a list of normalized sampling probability values:

[0126]

[0127] Calculate the probability mass function of the discrete sampling probability distribution of s samples in the sampling pool:

[0128]

[0129] Let X1, X2, …, X L is the result of L independent and identically distributed sampling from discrete probability distributions, then the sample set is obtained

[0130]

[0131] Among them, the same sample is sampled repeatedly, and the sample set obtained after sampling is The training set of the tth round is sent to the spatiotemporal feature fusion model for training the micro-expression recognition task, and L is set to the number of samples in the original dataset.

[0132] Preferably, according to the present invention, in step D, the target micro-expression spatiotemporal feature fusion model trained in step C is subjected to a classification and recognition test on a test set.

[0133] The beneficial effects of the present invention are:

[0134] The present invention introduces a staged adaptive curriculum learning training strategy into the micro-expression recognition algorithm, designs and implements a new curriculum learning algorithm for micro-expression recognition, effectively reduces the influence of noise in the data set, optimizes the training process of the neural network, obtains more effective and discriminative micro-expression features, improves the generalization ability of the model, and further solves the problems existing in the field of micro-expression recognition, such as the lack of available data sets, the large amount of redundant information in the data sets, and the low recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0135] Figure 1 Schematic diagram of 68 key feature points on the face of the present invention;

[0136] Figure 2 Schematic diagram of the spatiotemporal feature fusion model in the present invention;

[0137] Figure 3 This is a branch architecture diagram for extracting spatial appearance features in the present invention;

[0138] Figure 4 This is a branch architecture diagram for extracting subtle temporal dynamic features in the present invention;

[0139] Figure 5 This is the architecture diagram of the spatiotemporal feature fusion module in the present invention;

[0140] Figure 6 This is the overall architecture diagram of the staged adaptive course learning algorithm in the present invention;

[0141] Figure 7 The following are example graphs of the dynamic sampling probability mapping function under different training rounds t. DETAILED DESCRIPTION

[0142] In order to facilitate understanding of the present invention, the present invention will be further described below with reference to embodiments and accompanying drawings, but the present invention is not limited thereto.

[0143] Example 1

[0144] A micro-expression recognition method based on staged adaptive curriculum learning, comprising:

[0145] A. Preprocessing of micro-expression and macro-expression video sequences, including: obtaining video frame sequences, face detection and positioning, and face alignment;

[0146] B. Build a spatiotemporal feature fusion model, perform deep feature extraction on the micro-expression dataset preprocessed in step A, and pre-train a macro-expression recognition teacher model;

[0147] C. Construct a deep learning algorithm based on staged adaptive curriculum learning. Introduce a deep learning algorithm based on staged adaptive curriculum learning for micro-expression recognition into the training process of the spatiotemporal feature fusion model constructed in step B to optimize the training process.

[0148] D. Classify and recognize the macro expression recognition teacher model trained in step C on the test set.

[0149] Example 2

[0150] The micro-expression recognition method based on staged adaptive curriculum learning described in Example 1 is different in that:

[0151] In step A, the micro-expression and macro-expression video sequences are pre-processed, including the following steps:

[0152] 1) Obtaining a video frame sequence: performing frame processing on the video sequence to obtain and store the video frame sequence;

[0153] 2) Face detection and localization: Use the Dlib visual library to perform face detection and localization on each frame obtained in step 1), detecting the number of faces in the video frame and the distance between the face and the image boundary;

[0154] 3) Face alignment: Based on face positioning, the Dlib visual library is used to determine 68 key facial feature points to complete face segmentation and face correction.

[0155] Face segmentation refers to: using the Dlib visual library to segment faces using rectangular boxes;

[0156] Face correction means: among the 68 key feature points detected on the face, the line connecting the key feature point 37 marked at the left corner of the left eye and the key feature point 46 marked at the right corner of the right eye is at an angle a with the horizontal line. The corresponding rotation matrix is ​​obtained through angle a, and the segmented face is rotated to make the line connecting the key feature point 37 marked at the left corner of the left eye and the key feature point 46 marked at the right corner of the right eye parallel to the horizontal line, thereby correcting the facial posture and scaling the face.

[0157] In step B, the spatiotemporal feature fusion model structure includes a spatial appearance feature extraction branch and a temporal subtle dynamic feature extraction branch. In the spatiotemporal feature fusion model, first, the downsampled peak frame image is sent to the spatial appearance feature extraction branch to extract the spatial appearance features. Second, the absolute difference between the starting frame and the peak frame is sent to the temporal subtle dynamic feature extraction branch to extract the temporal subtle dynamic features. Finally, the obtained spatial appearance features and temporal subtle dynamic features are fused to obtain the final micro-expression features, which will be sent to the classification layer for micro-expression category classification.

[0158] This method designs a micro-expression recognition model that integrates spatiotemporal features as the backbone network. Figure 2 As shown in the figure, this architecture directly simulates the process of humans' comprehensive perception of temporal and spatial information when observing facial expressions. Combined with the staged adaptive curriculum learning strategy, it can effectively imitate the process of humans learning to recognize micro-expression features in reality.

[0159] The downsampled peak frame image is sent to the spatial appearance feature extraction branch to extract the spatial appearance features, including:

[0160] The spatial appearance feature extraction branch imitates the way humans focus on the relative positions of key facial organs while ignoring details such as identity when observing facial expressions. Leveraging the concept of the visual Transformer architecture, a lightweight network branch was designed. This architecture not only understands the relationship between various parts of the face from a global perspective, but also assists the temporal subtle dynamic feature extraction branch in identifying the local locations where micro-expression dynamics appear on the face. Figure 3 As shown, first, the peak frame of the micro-expression is downsampled to obtain a peak frame apex of size 28×28×3 i :

[0161] Among them, i represents the i-th sample, is in the real domain, and downsample is the downsampling operation. Downsampling not only prevents the network from over-learning interfering information such as identity information and facial texture unrelated to micro-expressions, but also significantly reduces computational effort. Since peak frames contain the strongest expressive information in a micro-expression sequence, downsampling the peak frames allows this branch to extract spatial information about the expression's appearance. Combined with the unique position encoding module of the ViT structure, it can effectively identify associations between various facial positions, assisting the temporal subtle dynamic feature extraction branch in locating the occurrence points of subtle micro-expression dynamics, thereby extracting more effective features. It refers to the peak frame of the i-th sample, and apex refers to the peak frame, that is, the single frame image in a micro-expression sample video sequence that contains the richest micro-expression features. It is annotated by experts and is the label that comes with the public micro-expression dataset.

[0162] The downsampled image is expanded into a 784×3 tensor in width and height dimensions, and the primary features are extracted through the facial pixel block embedding layer FaceEmbed:

[0163]

[0164] The facial pixel block embedding layer is a two-dimensional convolution operation with 3 input channels, 768 output channels, a convolution kernel size of 1×1, and a stride of 1. i refers to the peak frame image after downsampling obtained by the operation in formula (1); flatten() refers to the expansion operation, which is to expand the downsampled image into a tensor of size 784×3 in the width and height dimensions; z i is the expanded 784×3 tensor; FaceEmbed() refers to the facial pixel block embedding layer operation, that is, a two-dimensional convolution operation with 3 input channels, 768 output channels, a convolution kernel size of 1×1, and a stride of 1; is the extracted primary feature;

[0165] After obtaining the primary features, add position encoding:

[0166]

[0167] Among them, the position encoding is a set of trainable parameters with the same dimension as the primary features. The position encoding provides a position identifier for each element in the input pixel sequence, enabling the model to perceive the sequential relationship and relative position relationship between various positions on the face. pos is the position code, Refers to position encoding and Primary features obtained by element-by-element addition;

[0168] Next, the primary features pass through three stacked Transformer encoders (TEncoders) with an embedding vector dimension of 512, and then through a multi-layer perceptron (MLP) with an output dimension of 512 to extract spatial appearance features:

[0169]

[0170] The Transformer encoder is highly sensitive to high-frequency information, such as the edge texture of the face. This helps extract subtle surface features such as skin wrinkles caused by muscle contraction in micro-expression peak frames. At the same time, through the self-attention mechanism, it can capture the relative dependencies between various positions on the face, thereby extracting global features and assisting the temporal subtle dynamic feature extraction branch to locate the position of dynamic information on the face. Among them, TEncoder1, TEncoder2, and TEncoder3 are Transformer encoders with the same structure but no shared parameters; MLP is a multi-layer perceptron, that is, a fully connected layer with an output dimension of 512. Refers to the spatial appearance features finally extracted by the spatial appearance feature extraction branch.

[0171] The absolute difference between the start frame and the peak frame is fed into the temporal subtle dynamic feature extraction branch to extract temporal subtle dynamic features, including:

[0172] The temporal subtle dynamic feature extraction branch uses the absolute difference D (absolute difference) between the starting frame and the peak frame as input:

[0173]

[0174] The absolute difference is the difference between the two key frames, which is the subtraction of pixel intensities and the calculation of the absolute value. It represents the difference between the two key frames and can characterize the dynamic characteristics of micro-expressions. The dynamic characteristics of micro-expressions are very subtle and mainly concentrated in the key positions near the mouth and eyes. In order to better extract the temporal dynamic features and combine them with the position encoding information obtained in the spatial appearance feature extraction branch, D i Refers to the absolute difference between the starting frame and the peak frame of the i-th sample; They refer to the peak frame and starting frame of the i-th sample respectively; the peak frame is a single frame image containing the richest micro-expression features in a micro-expression sample video sequence, and the starting frame is a single frame image at the time point when the micro-expression begins to occur in a micro-expression sample video sequence. Both are annotated by experts and are the labels that come with the public micro-expression dataset.

[0175] The block-based convolutional neural network is used to extract features from the absolute difference D; Figure 4 As shown;

[0176] First, the absolute difference map of size H×H×3 is evenly divided into pixel blocks of size d×d×3 and tiled in order from top to bottom and from left to right into a list:

[0177]

[0178] Among them, Dlist i It refers to the list obtained by evenly dividing the absolute difference map of size H×H×3 into pixel blocks of size d×d×3 and tiling them in order from top to bottom and from left to right. k Refers to the kth pixel block, H refers to the height and width of the absolute difference map, and d refers to the height and width of each pixel block obtained after blocking;

[0179] In this method, H is 224, d is 16, and therefore K is 196. Next, each pixel block passes through the Subtle Dynamic Feature Perception Layer (SDFL) to obtain K feature tensors and concatenate them in the channel dimension:

[0180] F list i = [ f t1 ; f t2 ;…; f tK ] = [SDFL(block1);SDFL(block2);…;SDFL(block K )] (8);

[0181]

[0182] SDFL includes two-dimensional convolutional layers and maximum pooling; f tk It refers to the feature tensor obtained after the k-th pixel block passes through the subtle dynamic feature perception layer SDFL; F list i Refers to the K features f obtained by t1 ;f t2 ;…;f tK Arrange the list in order; dim=1 means the concatenation operation is performed on the first dimension (i.e., the channel dimension); concat refers to the concatenation operation; Refers to the features obtained after the splicing operation;

[0183] Finally, the total feature tensor The time-subtle dynamic features will be obtained through a mapping layer ML consisting of a two-dimensional convolution layer and a multi-layer perceptron

[0184]

[0185] The specific structures and parameters of SDFL and ML are shown in Table 1 below.

[0186] Table 1

[0187]

[0188] After obtaining the spatial appearance features and temporal subtle dynamic features, the spatial appearance features and temporal subtle dynamic features are fused; based on the attention mechanism, the macroscopic spatial appearance features are used to guide the fusion of subtle temporal dynamic features, and residual connections are performed to help the network locate key position encoding information from the spatial appearance features, so as to enhance the information located in certain key facial regions in the temporal subtle dynamic features and suppress the dynamic features of facial regions with weaker correlation. Figure 5 As shown in the figure, for the i-th sample, first, the query vector q and key vector k are extracted from the spatial appearance feature extraction branch, and the value vector v is extracted for the temporal subtle dynamic features:

[0189]

[0190] Among them, fc represents a fully connected layer with both input and output dimensions of 512, and different subscripts indicate that the corresponding branch parameters are not shared. q 、fc k 、fc v All of them are fully connected layers with input and output dimensions of 512, but the three parameters are not shared. The subscripts q, k, and v represent the three layers used to extract the query vector q, key vector k, and value vector v respectively; q i 、k i 、v i Respectively represent the query vector, key vector, and value vector extracted from the i-th sample;

[0191] Next, the correlation strength between the query vector and the corresponding key vector is calculated and normalized using the Softmax function to obtain the spatial correlation matrix:

[0192]

[0193] Among them, T represents the transposition operation, and the Softmax function is normalized along the direction of each column; the spatial correlation matrix M i Each element in represents the degree of attention that should be paid to the dynamic features of different positions, and v represents the dynamic feature information corresponding to each position;

[0194] Finally, the spatial correlation matrix is ​​multiplied by the temporal dynamic eigenvalue vector to obtain the spatiotemporal features weighted by the attention mechanism, and then the overall scaling is performed by the learnable weight parameter ρ and the original temporal subtle dynamic features are added element by element to obtain the final spatiotemporal fusion feature f i :

[0195]

[0196] Since the information related to facial expression features is mainly contained in the subtle temporal dynamic features, the above-mentioned residual connection can fully utilize the spatial features, i.e., the features at different positions of the face, to adjust the position weights of the temporal dynamic features while effectively retaining some of the information contained in the original dynamic features, thereby improving the robustness of the model.

[0197] Finally, the obtained spatiotemporal fusion features are sent to the classification layer for final expression category classification. The classification layer includes a layer normalization layer and a fully connected layer whose output dimension is the number of classification task categories.

[0198] The above-mentioned spatiotemporal feature fusion model architecture simulates the process of human eyes observing facial expressions. Therefore, it can not only be used as the backbone network architecture for micro-expression recognition, but also for macro-expression recognition. The macro-expression recognition teacher model, that is, the macro-expression model, is trained separately using the CK+macro-expression dataset. The architecture of the macro-expression model is a spatiotemporal feature fusion model. The network has relevant knowledge about macro-expressions. In the early stage of micro-expression model training, relevant knowledge about the dynamic features of different expressions, dynamic areas, and macro-patterns of facial muscle movements can be transferred to the micro-expression model to help the model train. Specifically, the first frame of each macro-expression video sequence is used as the starting frame, and all other frames are used as peak frames of macro-expressions to help the macro-expression model fully learn the spatiotemporal related features of facial macro-expressions. The cross-entropy loss function is used during pre-training of the macro-expression model;

[0199] The macro-expression model will assist in the initial training of the micro-expression model. This method uses the KL divergence (Kullback-Leibler Divergence) as the loss function for the macro-expression model to guide the micro-expression model:

[0200] Both the macro-expression model and the micro-expression model fuse spatiotemporal features, but they do not share parameters. The macro-expression model is trained using a macro-expression dataset and possesses relevant macro-expression knowledge. The micro-expression model, on the other hand, is designed to recognize micro-expressions. The goal of training the macro-expression model is to leverage the easily recognizable macro-expression knowledge to guide the more challenging micro-expression model.

[0201]

[0202] Among them, C is the number of expression categories, a=[a1,a2,…,a C ] is the predicted output of the macro expression model, b=[b1,b2,…,b C ] is the predicted output of the micro-expression model; L g Refers to KL divergence; it is a loss function commonly used in the field of deep learning, which minimizes the loss function L g,The micro-expression model is encouraged to learn relevant knowledge about the dynamic areas and dynamic features of ,expressions from the macro-expression model in the early stages of training.

[0203] Use the cross entropy loss function L CE As the classification loss of the micro-expression model:

[0204]

[0205] in, is the predicted probability distribution of the micro-expression model, y k is the true label of the sample, C is the number of categories; in the early stage of micro-expression model training, the total loss function is:

[0206] L CE +αL g (18);

[0207] Where α is the intensity coefficient of the guided loss;

[0208] This method defines the "initial stage" as the first 20% of the total training rounds EP. After the initial stage, the teacher model no longer plays a guiding role, that is, the total loss function L total for:

[0209]

[0210] In step C, a deep learning algorithm based on staged adaptive curriculum learning for micro-expression recognition is introduced into the training process of the spatiotemporal feature fusion model constructed in step B to optimize the training process. The steps include the following:

[0211] The entire course is designed as a phased structure, divided into several stages, such as Figure 6The figure shows the overall architecture of the staged adaptive curriculum learning algorithm. In the initial stage, the difficulty score of each example in the entire training set is evaluated based on a predefined curriculum based on human prior knowledge and a curriculum evaluated by a macro-expression recognition teacher model, thus initializing the curriculum. During this stage, the teacher model synchronously guides the output of the student model (the micro-expression recognition model), but ceases guidance after the initial stage. The step function in the early stage starts slowly, initially including only the simplest examples in the sampling pool. The size of the sampling pool is then gradually increased until the entire dataset is covered. The step function controls the sample size in the sampling pool, thereby controlling the diversity and complexity of the examples in each training round, achieving staged sample scheduling. Throughout the training process, the adaptive curriculum adjustment strategy reassesses the difficulty score of the entire training set after each round of training and updates it using a momentum-based difficulty update function. Finally, a dynamic sampling probability mapping function is used to form a sampling probability distribution and pass it to the sampler, which then probabilistically samples examples from the sampling pool based on this probability distribution. During the entire training process, the dynamic sampling probability mapping function will continue to change. It converts the difficulty score into sampling probability, so as to gradually focus on simple samples in the early stage and difficult samples in the later stage, thereby helping the neural network complete the learning process from easy to difficult and imitate human learning behavior.

[0212] The Dlib visual library's 68-point face detection algorithm is used to extract masks for the left eye area, right eye area, and mouth area. The mask area is a polygon formed by connecting the key points in Table 2 below.

[0213] Table 2

[0214]

[0215] Calculate the average optical flow field intensity amplitude mag of the left eye area, right eye area and mouth area i :

[0216]

[0217] Where M(x,y) represents the optical flow amplitude obtained by the Farneback algorithm at the pixel point (x,y), W is the image width, H is the image height, and i represents the i-th micro-expression sample; the optical flow field amplitude can represent the intensity of the movement. Since the movement intensity and the difficulty score should be inversely proportional, that is, micro-expression samples with greater movement intensity should be easier to recognize, the apparent movement pattern difficulty score of the i-th sample is It is defined as the inverse of the average optical flow field intensity amplitude of the sample:

[0218]

[0219] For the i-th micro-expression sample, the difficulty score of the macro-expression model evaluation is defined as:

[0220]

[0221] in, is the predicted probability distribution of the macro expression model for the i-th micro expression sample, is the true label of the i-th micro-expression sample, and C is the number of categories;

[0222] For the i-th micro-expression sample, the initial difficulty score before the micro-expression model training begins It is defined as the sum of the difficulty score of the apparent movement pattern and the difficulty score of the macro expression model evaluation:

[0223]

[0224] Among them, the subscript 1 represents the first round, represents the difficulty score of the i-th sample before the start of the t-th round of training, 1≤t≤EP, t∈Z; 1≤i≤N, i∈Z, Z represents the integer domain, and N is the total number of samples in the dataset; the difficulty score is used to construct the curriculum, that is, to construct the set of training samples used in each round of training.

[0225] The initialization score cannot accurately describe the recognition difficulty of the sample during the entire training process. The difficulty score of each sample should be dynamically adjusted according to the training status of the model. Therefore, an adaptive course adjustment strategy is proposed. Through the adaptive course adjustment strategy, the difficulty scores of all samples are re-evaluated after each training round, and the new difficulty information is updated to the difficulty score of the sample through the momentum-based difficulty value update function; the loss of a sample for the current model state can be regarded as its difficulty value. A large loss value of the sample means that the current model has difficulty in effectively learning the features of the sample or in distinguishing it from other categories. The sample may contain more noise information or its features are more similar to those of samples of other categories and are easily confused. Therefore, the sample corresponds to a greater difficulty. The adaptive course adjustment strategy is shown in the following equations (24) to (26):

[0226] The cross entropy loss is shown in formula (24):

[0227]

[0228] in, It represents the loss of the parameters of the current micro-expression model corresponding to the i-th sample after the end of the t-th training round. is the predicted probability distribution of the micro-expression model for the i-th sample, is the true label of the i-th sample, C is the number of categories; after the t-th training round, the update difficulty score of the i-th sample is the loss value of the sample for the micro-expression model state:

[0229]

[0230] in, It refers to the loss value of the i-th sample for the current micro-expression model state. refers to the update difficulty score of the i-th sample;

[0231] In order to smoothly update the difficulty score, we introduce new difficulty information while retaining some historical difficulty information. We use the momentum method to update the difficulty score of the sample. For the i-th sample, the difficulty score at the beginning of the t+1-th training round depends on the difficulty score of the t-th round:

[0232]

[0233] in, refers to the difficulty score of the i-th sample in the t-th training round and the t+1-th training round respectively; ω is the momentum coefficient, which is used to balance the weight between historical difficulty information and current difficulty information; in order to ensure the stability of numerical calculations and retain the information about the relative difficulty of each sample, the difficulty score is normalized by Min-Max before calculating the sampling probability in each round:

[0234]

[0235] Among them, min(score t ) represents the minimum difficulty score of all samples before the start of round t, max(score t ) represents the maximum difficulty score of all samples before the start of round t. The normalized difficulty score will be used to calculate the sampling probability. Refers to the normalized difficulty score obtained after Min-Max normalization;

[0236] The step function controls the size s of the sampling pool. The sampling pool size s refers to the sample set consisting of the first s samples in the entire dataset queue. The dataset queue refers to the sample index queue obtained by sorting the samples from small to large according to their difficulty scores, that is, the first s samples in the sample queue after the entire dataset is sorted from easy to difficult.

[0237] According to the idea of ​​curriculum learning, the model should learn a smaller number of samples in the early stages of training. As learning progresses, the number of samples should be gradually increased to allow the model to learn more diverse and complex feature distributions, thereby optimizing the feature learning process. This method uses the exponential function form commonly used in the field of curriculum learning research as the step function s(t):

[0238]

[0239] Among them, N is the size of the original dataset, t is the number of training rounds, r0 controls the initial size, K and T jointly control the step rate, and min() is the minimum value function; in this method, considering that the number of samples in the micro-expression dataset is small and the total number of training rounds EP is set to 100, r0 is set to 0.2, K is set to 1.15, and T is set to 5. This not only ensures a slow start in the early stages of training, that is, a slow increase in sample complexity, but also ensures that all samples can be included in the sampling pool in the later stages of training. Suppose all samples of the original micro-expression dataset are arranged in order:

[0240]

[0241] Before the start of round t of training, the difficulty score queue is:

[0242]

[0243] Sort from small to large to get a sequential queue:

[0244]

[0245] Get the index queue:

[0246] index = {k1,k2,k3,…,k N} (32);

[0247] Take the samples corresponding to the first s indexes of the index queue to form a sampling pool:

[0248]

[0249] After obtaining the difficulty score of the sample, the sampling probability of each sample needs to be calculated based on the difficulty score of each sample. The sampler will sample the data set sample pool according to this probability distribution. Based on the idea of ​​curriculum learning, this method proposes a dynamic sampling probability mapping function. Unlike the fixed calculation method in traditional curriculum learning methods, this function can change with the training round t, and focus on samples of different difficulty levels at different stages of training, that is, as the training progresses, it gradually switches from focusing on simple samples to focusing on difficult samples. In the early stages of training, simple samples with smaller difficulty scores should be paid more attention, and therefore correspond to a larger sampling probability. As training continues, samples with larger difficulty scores should gradually have their sampling probability increased, while simple samples that have been learned by the model in the early stages are gradually ignored. The dynamic sampling probability mapping function designed in this method Defined as:

[0250]

[0251] in:

[0252]

[0253] When t=1, focus on learning simple samples:

[0254]

[0255] Here, i represents the i-th sample, t represents the sampling probability calculated before the start of the t-th training round, t∈Z, 1≤t≤EP, Z is an integer domain, EP represents the total number of training rounds, and e is a natural constant. This function dynamically adjusts the mapping function from difficulty score to sampling probability as training progresses. In particular, the marginal benefit of simple samples in optimizing model performance may decrease in the later stages of training, but they still play an important role. Continuously exposing the model to simple samples can effectively retain basic and important micro-expression feature information, preventing the model from overfitting to the specific characteristics of difficult samples in the later stages. At any stage t>1, this function ensures that simple samples with normalized difficulty scores close to 0 still obtain a certain unnormalized sampling probability value, preventing simple samples from being abandoned.

[0256] Finally, the sampling pool is obtained by combining the sampling pool size s calculated by the step function. For the set of s samples in the sampling pool

[0257]

[0258] Set The difficulty scores corresponding to the s samples in :

[0259]

[0260] Substituting into formula (38) respectively, the sampling probability values ​​obtained are:

[0261]

[0262] Normalize the sampling probability values:

[0263]

[0264] Get a list of normalized sampling probability values:

[0265]

[0266] Calculate the probability mass function of the discrete sampling probability distribution of s samples in the sampling pool:

[0267]

[0268] Let X1, X2, …, X L is the result of L independent and identically distributed sampling from discrete probability distributions, then the sample set is obtained

[0269]

[0270] Among them, the same sample is sampled repeatedly, and the sample set obtained after sampling is The training set of the tth round is sent to the spatiotemporal feature fusion model for training the micro-expression recognition task, and L is set to the number of samples in the original dataset.

[0271] In step D, the target micro-expression spatiotemporal feature fusion model trained in step C is subjected to classification and recognition testing on the test set.

[0272] In this embodiment, the experiment was conducted using the pytorch deep learning framework under the Ubuntu operating system, with CUDA version 11.7, torch version 2.0.1, and the GPU model NVIDIA GeForce RTX 3090. ω was set to 0.6, α was set to 0.2, r0 was set to 0.2, K was set to 1.15, T was set to 5, the number of sampling times L was set to 5 times the number of samples in the original dataset, that is, the number of samples in the dataset after data enhancement, the total number of training rounds EP was set to 100 rounds, the initial stage was the first 20 rounds, the optimizer used the Adam optimizer, the learning rate μ was set to 0.00001, the sample batch size was set to 16, the mean and standard deviation in the data normalization preprocessing were both set to 0.5, all cropped face images had a pixel size of 224×224, the sample images were three-channel RGB images with a bit depth of 24, and the two frames before and after the peak frame of all micro-expression samples, a total of four frames, were used in the data enhancement method to expand the sample size. The LOSO (Leave-One-Subject-Out Cross-Validation) cross-validation method was used.

[0273] In this embodiment, in order to verify the advanced nature of the present invention, the recognition results of the present invention under the above settings are compared with the recognition results of the staged adaptive curriculum learning method in step C. The macro expression dataset uses the CK+ dataset and the micro expression dataset uses the CASMEII dataset. The results are shown in Table 3 below.

[0274] Table 3

[0275] Experimental results Accuracy F1 score Unweighted average recall Use this method 89.9 87.8 87.2 Do not use this method 79.8 63.8 62.5

[0276] The accuracy of the test set using the method of the present invention is effectively improved, which proves the advancement of the micro-expression recognition method based on staged adaptive curriculum learning of the present invention.

[0277] Example 3

[0278] The difference between the micro-expression recognition method based on staged adaptive curriculum learning described in Example 2 is that:

[0279] In this embodiment, experiments were conducted using the pytorch deep learning framework under the Ubuntu operating system, with CUDA version 11.7, torch version 2.0.1, and an NVIDIA GeForce RTX 3090 GPU. ω was set to 0.6, α was set to 0.2, r0 was set to 0.2, K was set to 1.15, T was set to 5, the number of sampling times L was set to 5 times the number of samples in the original dataset, that is, the number of samples in the dataset after data augmentation. The total number of training rounds EP was set to 100, with the initial stage being the first 20 rounds. The optimizer used the Adam optimizer, the learning rate μ was set to 0.00001, the sample batch size was set to 16, and the mean and standard deviation in the data normalization preprocessing were both set to 0.5. All cropped face images had a pixel size of 224×224, and the sample images were three-channel RGB images with a bit depth of 24. The two frames before and after the peak frame of all micro-expression samples, a total of four frames, were used in the data augmentation method to expand the sample size. The LOSO (Leave-One-Subject-Out Cross-Validation) cross-validation method was used.

[0280] In this embodiment, in order to verify the advanced nature of the present invention, the recognition results of the present invention under the above settings are compared with the recognition results of the staged adaptive curriculum learning method in step C. The micro-expression dataset uses SMIC, and the results are shown in Table 4 below.

[0281] Table 4

[0282] Experimental results Accuracy F1 score Unweighted average recall Use this method 75.6 75.2 74.8 Do not use this method 66.5 65.6 65.7

[0283] The accuracy of the test set using the method of the present invention is effectively improved, which proves the advancement of the micro-expression recognition method based on staged adaptive curriculum learning of the present invention.

Claims

1. A micro-expression recognition method based on staged adaptive curriculum learning, characterized in that: include: A. Preprocessing of micro-expression and macro-expression video sequences, including: obtaining video frame sequences, face detection and positioning, and face alignment; B. Build a spatiotemporal feature fusion model, perform deep feature extraction on the micro-expression dataset preprocessed in step A, and pre-train a macro-expression recognition teacher model; C. Construct a deep learning algorithm based on staged adaptive curriculum learning. Introduce a deep learning algorithm based on staged adaptive curriculum learning for micro-expression recognition into the training process of the spatiotemporal feature fusion model constructed in step B to optimize the training process. D. Classify and recognize the macro expression recognition teacher model trained in step C on the test set.

2. A micro-expression recognition method based on staged adaptive curriculum learning according to claim 1, characterized in that: In step A, the micro-expression and macro-expression video sequences are pre-processed, including the following steps: 1) Obtaining a video frame sequence: performing frame processing on the video sequence to obtain and store the video frame sequence; 2) Face detection and localization: Use the Dlib visual library to perform face detection and localization on each frame obtained in step 1), detecting the number of faces in the video frame and the distance between the face and the image boundary; 3) Face alignment: Use the Dlib visual library to determine 68 key facial feature points to complete face segmentation and face correction; Face segmentation refers to: using the Dlib visual library to segment faces using rectangular boxes; Face correction means: among the 68 key feature points detected on the face, the line connecting the key feature point 37 marked at the left corner of the left eye and the key feature point 46 marked at the right corner of the right eye is at an angle a with the horizontal line. The corresponding rotation matrix is ​​obtained through angle a, and the segmented face is rotated to make the line connecting the key feature point 37 marked at the left corner of the left eye and the key feature point 46 marked at the right corner of the right eye parallel to the horizontal line, thereby correcting the facial posture and scaling the face.

3. The micro-expression recognition method based on staged adaptive curriculum learning according to claim 1, characterized in that: In step B, the spatiotemporal feature fusion model structure includes a spatial appearance feature extraction branch and a temporal subtle dynamic feature extraction branch. In the spatiotemporal feature fusion model, first, the downsampled peak frame image is fed into the spatial appearance feature extraction branch to extract the spatial appearance features. Secondly, the absolute difference between the starting frame and the peak frame is sent to the temporal subtle dynamic feature extraction branch to extract the temporal subtle dynamic features; finally, the obtained spatial appearance features and temporal subtle dynamic features are fused to obtain the final micro-expression features, which will be sent to the classification layer for micro-expression category classification.

4. A micro-expression recognition method based on staged adaptive curriculum learning according to claim 3, characterized in that: The downsampled peak frame image is sent to the spatial appearance feature extraction branch to extract the spatial appearance features, including: First, the peak frame of the micro-expression is downsampled to obtain a peak frame apex of size 28×28×3. i : Among them, i represents the i-th sample, is the real number domain, downsample is the downsampling operation; Refers to the peak frame of the i-th sample, apex refers to the peak frame, that is, a single frame image containing the richest micro-expression features in a micro-expression sample video sequence; The downsampled image is expanded into a 784×3 tensor in width and height dimensions, and the primary features are extracted through the facial pixel block embedding layer FaceEmbed: The facial pixel block embedding layer is a two-dimensional convolution operation with 3 input channels, 768 output channels, a convolution kernel size of 1×1, and a stride of 1. i refers to the peak frame image after downsampling obtained by the operation in formula (1); flatten() refers to the expansion operation, which is to expand the downsampled image into a tensor of size 784×3 in the width and height dimensions; z i is the expanded 784×3 tensor; FaceEmbed() refers to the facial pixel block embedding layer operation, that is, a two-dimensional convolution operation with 3 input channels, 768 output channels, a convolution kernel size of 1×1, and a stride of 1; is the extracted primary feature; After obtaining the primary features, add position encoding: Among them, E pos is the position code, Refers to position encoding and Primary features obtained by element-by-element addition; Next, the primary features pass through three stacked Transformer encoders (TEncoders) with an embedding vector dimension of 512, and then through a multi-layer perceptron (MLP) with an output dimension of 512 to extract spatial appearance features: Among them, TEncoder1, TEncoder2, and TEncoder3 are Transformer encoders with the same structure but no shared parameters; MLP is a multi-layer perceptron, that is, a fully connected layer with an output dimension of 512; Refers to the spatial appearance features finally extracted by the spatial appearance feature extraction branch.

5. The micro-expression recognition method based on staged adaptive curriculum learning according to claim 3, characterized in that: The absolute difference between the start frame and the peak frame is fed into the temporal subtle dynamic feature extraction branch to extract temporal subtle dynamic features, including: The temporal subtle dynamic feature extraction branch uses the absolute difference D between the starting frame and the peak frame as input: Among them, D i Refers to the absolute difference between the starting frame and the peak frame of the i-th sample; refer to the peak frame and starting frame of the i-th sample respectively; A block-based convolutional neural network is used to extract features from the absolute difference D; First, the absolute difference map of size H×H×3 is evenly divided into pixel blocks of size d×d×3 and tiled in order from top to bottom and from left to right into a list: Among them, D list i It refers to the list obtained by evenly dividing the absolute difference map of size H×H×3 into pixel blocks of size d×d×3 and tiling them in order from top to bottom and from left to right. k Refers to the kth pixel block, H refers to the height and width of the absolute difference map, and d refers to the height and width of each pixel block obtained after blocking; Next, each pixel block will pass through the subtle dynamic feature perception layer SDFL to obtain K feature tensors and splice them in the channel dimension: F list i = [ f t1 ; f t2 ;…; f tK ] = [SDFL(block1);SDFL(block2);…;SDFL(block K )](8); SDFL includes two-dimensional convolutional layers and maximum pooling; f tk It refers to the feature tensor obtained after the k-th pixel block passes through the subtle dynamic feature perception layer SDFL; F list i Refers to the K features f obtained by t1 ;f t2 ;…;f tK Arrange the list in order; dim=1 means the concatenation operation is performed on the first dimension; concat means the concatenation operation; Refers to the features obtained after the splicing operation; Finally, the total feature tensor The time-subtle dynamic features will be obtained through a mapping layer ML consisting of a two-dimensional convolution layer and a multi-layer perceptron 6. A micro-expression recognition method based on staged adaptive curriculum learning according to claim 5, characterized in that: After obtaining the spatial appearance features and temporal subtle dynamic features, the spatial appearance features and temporal subtle dynamic features are fused. For the i-th sample, first, the query vector q and key vector k are extracted from the spatial appearance features extracted from the spatial appearance feature extraction branch, while the value vector v is extracted from the temporal subtle dynamic features: Among them, fc q 、fc k 、fc v They are all fully connected layers with input and output dimensions of 512, but the parameters are not shared. The subscripts q, k, and v represent the three used to extract the query vector q, key vector k, and value vector v respectively; q i 、k i 、v i Respectively represent the query vector, key vector, and value vector extracted from the i-th sample; Next, the correlation strength between the query vector and the corresponding key vector is calculated and normalized by the Softmax function to obtain the spatial correlation matrix: Among them, T represents the transposition operation, and the Softmax function is normalized along the direction of each column; the spatial correlation matrix M i Each element in represents the degree of attention that should be paid to the dynamic features of different positions, and v represents the dynamic feature information corresponding to each position; Finally, the spatial correlation matrix is ​​multiplied by the temporal dynamic eigenvalue vector to obtain the spatiotemporal features weighted by the attention mechanism, and then the overall scaling is performed by the learnable weight parameter ρ and the original temporal subtle dynamic features are added element by element to obtain the final spatiotemporal fusion feature f i : Finally, the obtained spatiotemporal fusion features are sent to the classification layer for final expression category classification. The classification layer includes a layer normalization layer and a fully connected layer whose output dimension is the number of classification task categories.

7. The micro-expression recognition method based on staged adaptive curriculum learning according to claim 1, characterized in that: The macro-expression recognition teacher model, or macro-expression model, is trained separately using the CK+ macro-expression dataset. The architecture of the macro-expression model is a spatiotemporal feature fusion model. The first frame of each macro-expression video sequence is used as the starting frame, and all other frames are used as the peak frames of the macro-expression. This helps the macro-expression model fully learn the spatiotemporal correlation features of facial macro-expressions. The cross-entropy loss function is used during pre-training of the macro-expression model. The macro expression model will assist in the initial training of the micro expression model, using KL divergence as the loss function for the macro expression model to guide the micro expression model: Among them, C is the number of expression categories, a=[a1,a2,…,a C ] is the predicted output of the macro expression model, b=[b1,b2,…,b C ] is the predicted output of the micro-expression model; L g refers to the KL divergence; Use the cross entropy loss function L CE As the classification loss of the micro-expression model: in, is the predicted probability distribution of the micro-expression model, y k is the true label of the sample, C is the number of categories; in the early stage of micro-expression model training, the total loss function is: L CE +αL g (18); Where α is the intensity coefficient of the guided loss; Total loss function L total for:

8. The micro-expression recognition method based on staged adaptive curriculum learning according to claim 1, characterized in that: In step C, a deep learning algorithm based on staged adaptive curriculum learning for micro-expression recognition is introduced into the training process of the spatiotemporal feature fusion model constructed in step B to optimize the training process; The steps are as follows: Use the Dlib visual library 68-point face detection algorithm to perform mask extraction on the left eye area, right eye area and mouth area; calculate the average optical flow field intensity amplitude mag of the left eye area, right eye area and mouth area i : Where M(x,y) represents the optical flow amplitude obtained at the pixel point (x,y) according to the Farneback algorithm, W is the image width, H is the image height, i represents the i-th micro-expression sample; the apparent motion pattern difficulty score of the i-th sample is It is defined as the inverse of the average optical flow field intensity amplitude of the sample: For the i-th micro-expression sample, the difficulty score of the macro-expression model evaluation is defined as: in, is the predicted probability distribution of the macro expression model for the i-th micro expression sample, is the true label of the i-th micro-expression sample, and C is the number of categories; For the i-th micro-expression sample, the initial difficulty score before the micro-expression model training begins It is defined as the sum of the difficulty score of the apparent movement pattern and the difficulty score of the macro expression model evaluation: Among them, the subscript 1 represents the first round, represents the difficulty score of the i-th sample before the start of the t-th round of training, 1≤t≤EP, t∈Z; 1≤i≤N, i∈Z, Z represents the integer domain, and N is the total number of samples in the dataset; Through the adaptive course adjustment strategy, the difficulty scores of all samples are re-evaluated after each training round, and the new difficulty information is updated to the difficulty score of the sample through the momentum-based difficulty value update function; the adaptive course adjustment strategy is shown in the following equations (24) to (26): The cross entropy loss is shown in formula (24): in, It represents the loss of the parameters of the current micro-expression model corresponding to the i-th sample after the end of the t-th training round. is the predicted probability distribution of the micro-expression model for the i-th sample, is the true label of the i-th sample, C is the number of categories; after the t-th training round, the update difficulty score of the i-th sample is the loss value of the sample for the micro-expression model state: in, It refers to the loss value of the i-th sample for the current micro-expression model state. refers to the update difficulty score of the i-th sample; The momentum method is used to update the difficulty score of the sample. For the i-th sample, the difficulty score at the beginning of the t+1-th training round depends on the difficulty score of the t-th round: in, They refer to the difficulty scores of the i-th sample in the t-th training round and the t+1-th training round respectively; ω is the momentum coefficient, which is used to balance the weight between historical difficulty information and current difficulty information; before calculating the sampling probability in each round, the difficulty score is normalized by Min-Max: Among them, min(score t ) represents the minimum difficulty score of all samples before the start of round t, max(score t ) represents the maximum difficulty score of all samples before the start of round t. The normalized difficulty score will be used to calculate the sampling probability. Refers to the normalized difficulty score obtained after Min-Max normalization; The step function controls the size s of the sampling pool. The sampling pool size s refers to the sample set consisting of the first s samples in the entire dataset queue. The dataset queue refers to the sample index queue obtained by sorting the samples from small to large according to their difficulty scores, that is, the first s samples in the sample queue after the entire dataset is sorted from easy to difficult. Use exponential function as step function s(t): Where N is the size of the original dataset, t is the number of training rounds, r0 controls the initial size, K and T jointly control the step rate, and min() is the minimum function. Suppose all samples of the original micro-expression dataset are arranged in order: Before the start of round t of training, the difficulty score queue is: Sort from small to large to get a sequential queue: Get the index queue: index = {k1,k2,k3,…,k N } (32); Take the samples corresponding to the first s indexes of the index queue to form a sampling pool: Dynamic sampling probability mapping function Defined as: in: When t=1, focus on learning simple samples: Where i represents the i-th sample, t represents that the sampling probability is calculated before the start of the t-th training round, t∈Z, 1≤t≤EP, Z is an integer domain, EP represents the total number of training rounds, and e is a natural constant; Finally, the sampling pool is obtained by combining the sampling pool size s calculated by the step function. For the set of s samples in the sampling pool Set The difficulty scores corresponding to the s samples in : Substituting into formula (38) respectively, the sampling probability values ​​obtained are: Normalize the sampling probability values: Get a list of normalized sampling probability values: Calculate the probability mass function of the discrete sampling probability distribution of s samples in the sampling pool: Let X1, X2, …, X L is the result of L independent and identically distributed sampling from discrete probability distributions, then the sample set is obtained Among them, the same sample is sampled repeatedly, and the sample set obtained after sampling is The training set of the tth round is sent to the spatiotemporal feature fusion model for training the micro-expression recognition task, and L is set to the number of samples in the original dataset.

9. A micro-expression recognition method based on staged adaptive curriculum learning according to any one of claims 1 to 8, characterized in that: In step D, the target micro-expression spatiotemporal feature fusion model trained in step C is subjected to classification and recognition testing on the test set.

Citation Information

Cited By

  • Expression recognition model training method and system based on head posture and facial key points

    CN120913009A

  • Step-by-step data enhancement method and device based on course learning rule and meta-learner

    CN121392463A

  • Progressive data augmentation method and apparatus based on curriculum learning rules and meta-learner

    CN121392463B

  • Micro-expression recognition method and system based on balanced adaptive grouped sampling

    CN121768057A