Micro-expression recognition method for cross-domain feature center assisted emotional intensity invariance feature extraction

Through the emotional intensity invariance feature extraction network assisted by the cross-domain feature center, the macro-expression data set assists in micro-expression recognition, and feature optimization is performed on the hypersphere, which solves the problem of data set size limited and sample imbalance in micro-expression recognition, and achieves more efficient micro-expression emotional feature extraction and recognition accuracy.

CN120220210APending Publication Date: 2025-06-27SHANDONG UNIV SHENZHEN RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510298455.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

There are problems in the field of micro-expression recognition that data set size is limited and sample imbalance is not balanced, making it difficult for algorithms and models to learn high discriminative micro-expression emotional characteristics.

Method used

A micro-expression recognition method is proposed to assist in emotional intensity invariance feature extraction with cross-domain feature center. By constructing a network of emotional intensity invariance feature extraction assisted by cross-domain feature center, a macro-expression data set assists in micro-expression recognition is used to guide the learning of micro-expression features, and feature optimization is carried out on the hypersphere to extract more discriminant angle-related emotional features.

Benefits of technology

The number of available cross-domain data samples during training of micro-expression model is effectively expanded, the model's ability to learn discriminant micro-expression emotional characteristics, and the accuracy and generalization ability of micro-expression recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220210A_ABST
    Figure CN120220210A_ABST
Patent Text Reader

Abstract

The invention relates to a micro-expression recognition method for cross-domain feature center assisted emotional intensity invariance feature extraction, belongs to the technical field of deep learning and pattern recognition, designs a cross-domain feature center assisted emotional intensity invariance feature extraction network, and fully utilizes an existing macro-expression data set to assist micro-expression recognition. Macro-expression related features are utilized to guide learning of micro-expression features, a neural network is helped to learn more emotion related features, hyper-spherical constraint is carried out on the micro-expression features and the macro-expression features, multiple angle optimization strategies are designed, intensity information of the features is weakened, angle information is focused, and the effect of improving the accuracy of the micro-expression features is achieved. In addition, a macro-expression feature center and a micro-expression feature center are designed to cooperatively guide training of the network, and a model is helped to learn more compact emotional feature distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a micro-expression recognition method for cross-domain feature center-assisted extraction of emotion intensity-invariant features, belonging to the technical fields of deep learning and pattern recognition. Background Art

[0002] Currently, there are problems in micro-expression recognition, such as the limited scale of the micro-expression dataset and the imbalance in the number of samples among different categories, and it is difficult for various algorithms and models to learn highly discriminative micro-expression emotion features. In recent years, the integration of artificial intelligence technology and micro-expression recognition has led to the rapid development of micro-expression recognition technology. Since expressions can reveal a person's mental state and true thoughts and can be applied to various different fields, extensive research has been conducted in recent decades, including research in psychology, sociology, neuroscience, and computer science. Psychologists classify expressions into two categories, namely macro-expressions and micro-expressions. Among them, common macro-expressions can be deliberately concealed and have a relatively long duration, with a large facial muscle movement area. In comparison, micro-expressions have a very short duration, are involuntary and uncontrollable small and rapid movements of facial muscles, are not easily detectable by the naked eye, and cannot be consciously suppressed or concealed by the subject. Therefore, one of the biggest differences between micro-expressions and macro-expressions lies in the intensity of emotions. The objective manifestation of the difference in emotion intensity between the two in samples is the different duration of the expression sequence and the different amplitudes of facial muscle movements. Although the difference in emotion intensity between the two is large, the emotion features expressed by macro-expression samples and micro-expression samples of the same category are consistent. Currently, micro-expression recognition has considerable application value and has been widely applied in fields such as national security, public security, and psychotherapy.

[0003] In recent years, micro-expression recognition technologies and methods relying on computer technology have achieved certain results. Currently, automatic micro-expression recognition algorithms are mainly divided into two categories, namely traditional algorithms and deep learning-based algorithms. In terms of traditional algorithms, local binary patterns based on three orthogonal planes for describing dynamic textures were first applied to the field of micro-expression recognition. At the same time, the main direction mean optical flow (MDMO) proposed by Liu is a typical spatio-temporal description method, which belongs to a kind of optical flow method. Subsequently, Liu et al. further incorporated the main direction mean optical flow into classical graph-regularized sparse coding, thereby generating sparse main direction mean optical flow features. In addition, Xu et al. proposed the facial dynamics map (FDM) feature, using an iterative optimization strategy to calculate the main optical flow direction of the spatio-temporal cuboid obtained from the segmented micro-expression sequence in order to more effectively present local facial dynamics features. Besides, Happy et al. proposed the fuzzy histogram of optical flow directions (FHOFO), which constructs an angular histogram from the optical flow vector direction by means of histogram fuzzification to encode the temporal pattern features of micro-expressions. Liong proposed a brand-new optical strain weighted feature extraction scheme, which extracts weighted spatio-temporal information from the segmented optical flow map.

[0004] Meanwhile, with the development of the current field of artificial intelligence and deep learning, more and more researchers have started using deep learning-based micro-expression recognition methods. In 2016, Dae Hoe Kim proposed a micro-expression recognition method that combines CNN and LSTM, which can extract spatial and temporal features from video sequences simultaneously. In 2018, Wang introduced a micro-expression recognition method using transfer learning. In 2020, Yante Li proposed a micro-expression peak frame recognition algorithm based on three-dimensional Fourier transform, and simultaneously constructed the LGCcon model. By extracting and fusing local and global features, the network was then trained. Xia preprocessed the dataset using the Euclidean video magnification technique to magnify the dynamic information of micro-expressions, and also designed a mask algorithm to extract information on key facial parts, significantly reducing the data dimension. In the design of the neural network, the recursive convolutional neural network RCNN with spatio-temporal feature extraction capabilities was selected. Xie proposed using facial action unit AU to assist micro-expression recognition. In 2021, Ben conducted a detailed investigation and analysis as well as a prospect for the field of micro-expression detection and recognition based on video datasets. Given the current shortage of datasets in the micro-expression field, Ben proposed a new dataset, MMEW. In 2022, Chen proposed a micro-expression recognition algorithm that only uses the starting frame and the peak frame, fused the optical flow features extracted by four different methods, and designed a block convolutional neural network, effectively optimizing the extraction of optical flow features. Zhao et al. were the first to propose a micro-expression recognition method using the Vision Transformer model, demonstrating that the model can be successfully applied to micro-expression recognition tasks. Mao et al. explored the micro-expression recognition task when the face is occluded by an object in a real environment and proposed a region heuristic relationship reasoning network (RRRN) to capture the complementary relationships between different regions of the face. Wei designed a novel amplification adaptive network (AMAN) based on the attention mechanism, which can adapt to different amplification levels of the dynamic changes in micro-expression sequences. In 2023, Nguyen et al. proposed a micro-expression recognition framework (Micron-BERT) based on the BERT model. Micron-BERT utilizes the powerful sequence modeling ability of the BERT framework to accurately locate the region of interest (POI) in micro-expression frames. This effectively reduces the impact of interference factors such as background noise on the model's recognition performance and significantly improves the accuracy and reliability of micro-expression recognition. In 2024, Bao et al. proposed a new method that simultaneously integrates supervised prototype-based memory contrast learning for discriminative feature mining and adds self-expression reconstruction as an auxiliary task and regularization method, effectively improving the recognition accuracy of micro-expressions.

[0005] Micro-expression recognition algorithms based on learning features usually rely on deep learning methods, and the training of deep neural networks cannot be achieved without the support of a large amount of data. However, the collection of micro-expression datasets faces many challenges. It has extremely stringent requirements on conditions such as light, environment, excitation source, and the subject's own state, and the collection cost is relatively high. This results in the current available datasets not only being scarce in number, but also containing a large amount of redundant information, which has become a major obstacle in the field of micro-expression recognition research. Therefore, it is necessary to consider using cross-domain datasets to effectively expand the datasets required for learning features of micro-expression recognition models. Summary of the invention

[0006] In view of the deficiencies in the prior art, the present invention provides a micro-expression recognition method for extracting emotion intensity invariant features assisted by cross-domain feature centers. Summary of the invention:

[0008] The purpose of the present invention is to solve the problem that the data set size in the field of micro-expression recognition is limited and the number of samples between categories is unbalanced, and various algorithms and models are difficult to learn high-discriminative micro-expression emotional features. Combining the objective characteristics of micro-expression and macro-expression data sets, and utilizing contemporary artificial intelligence technology and multi-task learning ideas, the present invention proposes a micro-expression recognition algorithm for cross-domain feature center-assisted emotion intensity invariant feature extraction, designs a cross-domain feature center-assisted emotion intensity invariant feature extraction network, makes full use of the existing macro-expression data set to assist micro-expression recognition, uses macro-expression related features to guide the learning of micro-expression features, helps the neural network learn more emotion-related features, and at the same time, performs hyperspherical constraints on micro-expression features and macro-expression features, and designs a variety of angle optimization strategies to weaken the intensity information of the features and focus on the angle information, so as to help the neural network extract more discriminative angle-related emotional features. In addition, the macro-expression feature center and the micro-expression feature center are designed to collaboratively guide the training of the network, so as to help the model learn a more compact distribution of emotional features.

[0009] Terminology explanation:

[0010] 1. Dlib visual library: Dlib is an open source toolkit that contains machine learning algorithms and can be used to solve many practical problems in the field of machine learning. Currently, Dlib has been widely used in industry and academia.

[0011] 2. Facial key feature point detection: The 68 key feature points of the face are mainly distributed in eyebrows, eyes, nose, mouth and facial contours, such as Figure 1 As shown, detection is performed through the Dlib visual library, which is a prior art.

[0012] 3. TVL1 Optical Flow Extraction Method: An algorithm for estimating the motion of objects in an image sequence. Based on the principle of variational method, it has strong robustness and is suitable for dealing with complex motion scenes. It is an existing technology and has been widely applied in the field of optical flow calculation currently.

[0013] 4. ConvNext Neural Network: An advanced new convolutional neural network architecture. It redesigns the classic convolutional neural network module by drawing on the Transformer architecture and has achieved excellent performance in computer vision tasks. It has been widely applied in the field of deep learning currently.

[0014] The technical solution of the present invention is as follows:

[0015] A micro-expression recognition method for cross-domain feature center-assisted emotion intensity invariant feature extraction, including the following steps:

[0016] A. Preprocess the micro-expression and macro-expression video sequences, including: obtaining the video frame sequence, obtaining the starting frame and peak frame, face detection and positioning, face alignment, and obtaining the TVL1 optical flow feature map;

[0017] B. Construct a cross-domain feature center-assisted emotion intensity invariant feature extraction network to perform deep feature extraction on the optical flow feature map extracted in step A;

[0018] C. Construct a cross-domain loss function and optimization strategy based on the emotion intensity invariant principle to optimize the training process;

[0019] D. According to the dataset constructed in step A, the deep learning neural network constructed in step B, and the loss function used in the backpropagation algorithm specified in step C, train the cross-domain feature center-assisted emotion intensity invariant feature extraction network, and use the trained deep network model to perform classification and recognition on the dataset.

[0020] Preferably according to the present invention, in step A, preprocess the micro-expression and macro-expression video sequences, including the following steps:

[0021] 1) Obtain the video frame sequence: Perform frame splitting on the micro-expression and macro-expression video sequences to obtain the video frame sequence and store it;

[0022] 2) Obtain the starting frame and peak frame: According to the information marked by experts in the micro-expression dataset, select the starting frame and peak frame of the micro-expression in each video sequence file and store them; for the macro-expression dataset, select the starting frame and the second to fifth frames in each video sequence file as emotion frames and store them;

[0023] The starting frame of the micro-expression refers to: the first frame when the micro-expression appears in the micro-expression video sequence;

[0024] The micro-expression peak frame refers to: the frame with the most obvious facial muscle changes in the micro-expression video sequence, which is marked by experts in the dataset and contains the most micro-expression information;

[0025] The macro-expression start frame refers to: the first frame in the macro-expression video sequence;

[0026] The macro-expression emotion frame refers to: the second to fifth frames in the macro-expression video sequence, whose contained expression and emotion information are similar to those of micro-expressions, but with a higher intensity than micro-expressions;

[0027] 3) Face detection and localization: Use the Dlib library to perform face detection and localization on all the video frame images obtained in step 2), and detect the number of faces in the video frames and the distance of the faces from the image boundaries;

[0028] 4) Face alignment: On the basis of face localization, use the Dlib library to determine 68 key facial feature points to complete face segmentation and face correction;

[0029] Face segmentation refers to: using the Dlib vision library to segment the face with a rectangular box;

[0030] Face correction refers to: among the 68 key facial feature points detected, the line connecting the key feature point marked at the left corner of the left eye and the key feature point marked at the right corner of the right eye has an angle a with the horizontal line. Through this angle a, the corresponding rotation matrix is obtained, and the segmented face is rotated to make the line connecting the key feature point marked at the left corner of the left eye and the key feature point marked at the right corner of the right eye parallel to the horizontal line, realizing the correction of the face pose, and scaling the face, and finally cropping it into a face image of size 224×224×3;

[0031] 5) Obtain the TVL1 optical flow feature map: Use the TVL1 optical flow extraction method to perform optical flow feature extraction on the start frame and the peak frame in each aligned micro-expression video sequence, and at the same time perform optical flow feature extraction on the start frame and the emotion frame in each aligned macro-expression video sequence; The TVL1 optical flow extraction algorithm outputs two parts including the horizontal optical flow information component and the vertical optical flow information component. For micro-expression samples, select the optical flow information between the start frame and the peak frame and as the model input, so as to maximize the extraction of the emotion information of facial muscle movement representation; For macro-expression samples, select the optical flow information between the start frame and the emotion frame and as the model input, where, micro_i represents the serial number of the i-th micro-expression sample, macro_i represents the serial number of the i-th macro-expression sample, and flowh represents horizontal optical flow information, flow v represents vertical optical flow information, and the size of all optical flow maps is 224×224×3.

[0032] According to a preferred embodiment of the present invention, in step B, the cross-domain feature center-assisted emotion intensity invariance feature extraction network includes a two-branch emotion intensity invariance feature extraction network and an angle optimization network:

[0033] First, a two-branch emotion intensity invariance feature extraction network is designed, and the emotion feature extraction parameters of this network are shared among cross-domain samples. Cross-domain samples refer to the set of micro-expression samples and macro-expression samples; this network includes two backbone networks with the same structure but non-shared parameters, namely the horizontal emotion feature extraction network EFEN h (horizontalEmotionFeature Extraction Network) and the vertical emotion feature extraction network EFEN v (verticalEmotion Feature Extraction Network), and an Adaptive Emotion feature Focusing Module (AEFM) is added to the end of each branch network respectively, which is used to extract horizontal emotion features and vertical emotion features

[0034]

[0035] Among them, the backbone network is the horizontal emotion feature extraction network EFEN h and the vertical emotion feature extraction network EFEN v Refer to the current advanced ConvNeXt convolutional neural network structure for architecture design, and the specific structure is shown in the following table.

[0036] Table 1 ConvNeXt Convolutional Neural Network Architecture Design

[0037]

[0038] Among them, the specific structure of the ConvNeXt block in the above table is shown in the following table.

[0039] Table 2 ConvNeXt Block Structure

[0040]

[0041]

[0042] The Adaptive Emotion Feature Focusing Module (AEFM) is designed based on the channel attention mechanism. It highlights certain most important eigenvalue of the input 64-dimensional feature f while suppressing other eigenvalues, thereby balancing the relationship between the horizontal and vertical feature components, so as to better extract effective emotion features. The output of this module is:

[0043] f emotion = Sigmoid(Linear2(RELU(Linear1(Avgpool(f)))))⊙ f (3)

[0044] Where Sigmoid is the Sigmoid activation function, RELU is the RELU activation function, Avgpool represents global average pooling, Linear1 represents a fully connected layer with input and output dimensions of 64 and 4, Linear2 represents a fully connected layer with input and output dimensions of 4 and 64, and the dot multiplication symbol ⊙ represents element-wise multiplication.

[0045] To obtain the final emotion feature tensor, the feature and are added and fused element by element at the corresponding positions, and then passed through the Emotion Feature Mapping Layer (EML), that is, a fully connected layer. In this method, the size of the emotion feature dimension is taken as 32 to obtain the emotion feature x emotion :

[0046]

[0047] After the sample is extracted by the above double-branch network, the emotion feature x emotion is obtained and sent into the angle optimization network; due to the cross-domain characteristics, hereinafter x i is used to represent the emotion feature obtained after the i-th cross-domain expression sample passes through the double-branch network, N is the size of the emotion feature dimension, is the real number field; use to represent the emotion feature of the i-th micro-expression sample, to represent the emotion feature of the i-th macro-expression sample. The superscript micro represents micro-expression, and macro represents macro-expression, so as to distinguish the target domain and the source domain.

[0048] The emotion feature center W=[W1,W2,…,W C refers to a set of trainable parameters, which consists of C N-dimensional feature tensors. C represents the number of categories in the classification task, where W c represents the feature center of the c-th type of emotion, and its feature dimension is consistent with that of x i , that is 1 ≤ c ≤ C, c ∈ Z, where Z is the set of integers; the feature center is located at the last layer of the angle optimization network and can represent the representative center points of the feature of each category of samples; on the hypersphere, the smaller the angle between the emotion feature of a sample and the feature center of a certain category, the higher the emotion similarity between the two, and vice versa; in this method, the micro-expression emotion feature center and the macro-expression emotion feature center During the backpropagation process, the macro-expression samples do not directly participate in the parameter update of the micro-expression emotion feature center to ensure the representativeness of the micro-expression emotion feature center for the emotion features of micro-expression samples.

[0049] Different from directly outputting the predicted probability distribution of each category in the classification layer of the traditional neural network, this method calculates the cosine similarity between all expression sample features and the emotion feature centers of the corresponding domains, and normalizes the emotion features of all samples and the feature centers of each category respectively, so as to weaken the intensity property of the features, thereby enabling the model to focus on the angular property of the features. The final output of the angle optimization network is the cosine similarity between the sample emotion feature and each emotion feature center:

[0050]

[0051] where represents the cosine similarity between the i-th micro-expression sample and the c-th micro-expression feature center, represents the cosine similarity between the i-th macro-expression sample and the c-th macro-expression feature center, and the radian values θ of all angles are limited to [0, π]; represents the micro-expression feature center corresponding to the c-th emotion, represents the macro-expression feature center corresponding to the c-th emotion, 1 ≤ c ≤ C, c ∈ Z, where C represents the number of categories in the emotion classification task and Z represents the set of integers; y i represents the true emotion label category of the i-th sample, and j represents other emotion label categories that are inconsistent with the true emotion label of the i-th sample, that is, 1 ≤ j ≤ C, j ≠ y i , j ∈ Z; in the above formula, for any N-dimensional feature T = (t1, t2, t3…, t N-1 , t N ), the symbol ‖T‖ represents the Euclidean norm of T:

[0052]

[0053] The dot product symbol A·B represents the inner product operation of tensors. For any two tensors A = (a1, a2,…, a N ) and B = (b1, b2,…, bN ) There is:

[0054]

[0055] The predicted category finally output by the angle optimization network is:

[0056]

[0057] Since cosθ is monotonically decreasing on the domain [0, π], the predicted category is also:

[0058]

[0059] According to the above formula, the cross-domain feature center-assisted emotion intensity invariance feature extraction network can give a predicted category for each cross-domain sample.

[0060] According to the preference of the present invention, in step C, constructing a cross-domain loss function and an optimization strategy based on the emotion intensity invariance principle includes the following steps:

[0061] 1) Construct the emotion classification cosine similarity optimization loss on the micro-expression and macro-expression discrimination branches. In order to make a preliminary classification of the cross-domain expression categories output by the network, combining the cosine similarity metric with the Softmax loss function form, design the cosine similarity-based micro / macroexpression emotion classification optimization loss. On the micro-expression discrimination branch, the loss function is:

[0062]

[0063] On the macro-expression discrimination branch, the loss function is:

[0064]

[0065] Among them, B micro +B macro =B; B represents the batch size during network training, B micro represents the number of micro-expression samples in this batch, B macro represents the number of macro-expression samples in this batch; e is the natural constant, and all logarithms in the function are based on e; S represents the radius size of the hypersphere, that is, all sample features and feature centers are mapped to tensors with an Euclidean norm of S, S>0.

[0066] 2) Construct a cross-domain sample emotion feature alignment loss function. The fundamental purpose of setting the dual feature centers is to use the macro-expression sample features to guide the learning of the same-class micro-expression features, and at the same time guide the micro-expression features of other different classes to be as far away from each other as possible, so as to learn more discriminative features. Design a cross-domain sample emotion feature alignment loss function L cross (cross-domain emotion feature alignment loss function) to constrain the angular relationship between the micro-expression features and the macro-expression feature centers, so that the micro-expression sample emotion features are aligned with the macro-expression emotion feature centers of the same class, in order to achieve the cross-domain guiding effect of the macro-expression feature centers on the micro-expression sample features:

[0067]

[0068] Among them, the functions and are the scalar of the within-class loss coefficient and the scalar of the between-class loss coefficient respectively:

[0069]

[0070] Among them, k, a, and b are hyperparameters, and represent the within-class cross-domain cosine similarity and the between-class cross-domain cosine similarity respectively:

[0071]

[0072] The cross-domain cosine similarity represents the cosine similarity between the micro-expression sample and the macro-expression feature center. By minimizing this loss function, the model can be helped to learn cross-domain features and improve the generalization ability of the model for micro-expression samples. and G(θ j ) imitate the characteristics of the Sigmoid function, and their role is to control the optimization speed, that is, to control the parameter update rate in the backpropagation process of the neural network. Among them, a and b are the angular bounds (in radians), which are two hyperparameters used to control how sharply the loss function changes when the angular relationship between the feature and the feature center reaches a certain bound.

[0073] 3) Construct a cross-domain emotion feature centers regularization loss function. The emotion feature centers of the same-class micro-expressions and macro-expressions should be consistent. Apply a center regularization loss L center (cross-domain emotion feature centers regularization loss function) to guide the emotion feature centers of the same-class micro-expressions and macro-expressions to approach each other in the feature space:

[0074]

[0075] This loss encourages the model to pay attention to the consistency of the within-class feature centers between macro-expressions and micro-expressions during the learning process, that is, to minimize the angular gap between the emotion feature centers of macro-expressions and micro-expressions of the same class.

[0076]

[0077] This regularization loss term helps the feature centers of macro-expressions and micro-expressions of the same class to be more compact in the feature space, and this loss term can cooperate with the alignment loss function L. cross Fully exploit the invariant features of emotion intensity shared between macro-expressions and micro-expressions, which helps to prevent the model from overfitting, promotes the sharing and transfer of features, and further strengthens the feature compactness within the same class across domains.

[0078] 4) Construct a macro-expression emotion feature compactness loss function. Macro-expression data has a more compact emotion feature distribution than micro-expression data. To better play the guiding role of the macro-expression emotion feature center on the emotion features of micro-expression samples, apply the macro-expression emotion feature compactness loss function L compact (macro-expressionemotion feature compactness loss function) in the macro-expression domain to enhance its within-domain separability and strengthen the representativeness of the macro-expression feature center for various emotion features:

[0079]

[0080] Among them, the functions H and G are the same as in formulas (18) and (19).

[0081] 5) Construct a joint loss function. The joint loss function L total in this method is:

[0082] L total = L micro + L macro + αL cross + βL center + γL compact (25)

[0083] Among them, α, β, and γ are loss coefficient hyperparameters used to balance the importance between different loss terms, and the three cooperate to control the regularization strength.

[0084] According to the preference of the present invention, in step D, the backpropagation formula during the training of the cross-domain feature center-assisted emotion intensity invariant feature extraction network is:

[0085]

[0086] Among them, Θ represents all network parameters of the cross-domain feature center-assisted emotion intensity invariance feature extraction network, and W macro is the macro-expression emotion feature center parameter, and W micro is the micro-expression emotion feature center parameter, and μ is the learning rate; finally, the trained micro-expression recognition model can predict the expression category of the sample.

[0087] The beneficial effects of the present invention are as follows:

[0088] The present invention designs a novel cross-domain feature center-assisted emotion intensity invariance feature extraction network, proposes a new method for aligning macro-expression and micro-expression sample features, first proposes to optimize and align the angular features of cross-domain expression samples on the hypersphere, and proposes to use the feature center idea to assist in optimizing the feature learning process. At the same time, four optimization strategies are proposed, including five loss functions, to extract the emotion intensity invariance angular features and enhance the intra-class compactness and discriminability of the extracted features. The present invention effectively expands the number of available cross-domain data samples during micro-expression model training, enhances the model's ability to learn discriminative micro-expression emotion features, enhances the generalization ability of the model, and improves the accuracy of micro-expression recognition. Brief Description of the Drawings

[0089] Figure 1 is a schematic diagram of 68 key feature points of the face according to the present invention;

[0090] Figure 2 is a schematic diagram of the cross-domain feature center-assisted emotion intensity invariance feature extraction network according to the present invention;

[0091] Figure 3 is a schematic diagram of the model predicting the sample category according to the present invention;

[0092] Figure 4 is a schematic diagram of the optimization principle of the cosine similarity optimization loss for macro-expression emotion classification according to the present invention;

[0093] Figure 5 is a schematic diagram of the optimization principle of the cosine similarity optimization loss for micro-expression emotion classification according to the present invention;

[0094] Figure 6 is a schematic diagram of the optimization principle of the cross-domain sample emotion feature alignment loss function according to the present invention;

[0095] Figure 7 is a schematic diagram of the optimization principle of the cross-domain emotion feature center regularization loss function according to the present invention;

[0096] Figure 8This is a schematic diagram of the optimization principle of the compact loss function for macro-expression emotion features in the present invention. Detailed implementation manners

[0097] The present invention will be further described below through embodiments in conjunction with the accompanying drawings, but is not limited thereto.

[0098] Embodiment 1:

[0099] A micro-expression recognition method for cross-domain feature center-assisted emotion intensity invariance feature extraction, including the following steps:

[0100] A. Preprocess the micro-expression and macro-expression video sequences, including: obtaining the video frame sequence, obtaining the starting frame and the peak frame, face detection and positioning, face alignment, and obtaining the TVL1 optical flow feature map.

[0101] In step A, preprocess the micro-expression and macro-expression video sequences, including the following steps:

[0102] 1) Obtain the video frame sequence: Perform frame splitting on the micro-expression and macro-expression video sequences to obtain the video frame sequence and store it.

[0103] 2) Obtain the starting frame and the peak frame: According to the information annotated by experts in the micro-expression dataset, select the starting frame and the peak frame of the micro-expression in each video sequence file and store them. For the macro-expression dataset, select the starting frame and the second to fifth frames in each video sequence file as the emotion frames and store them.

[0104] The micro-expression starting frame refers to: the first frame in the micro-expression video sequence where the micro-expression appears.

[0105] The micro-expression peak frame refers to: the frame in the micro-expression video sequence where the facial muscle changes are most obvious, which is annotated by experts in the dataset, and this frame contains the most micro-expression information.

[0106] The macro-expression starting frame refers to: the first frame in the macro-expression video sequence.

[0107] The macro-expression emotion frame refers to: the second to fifth frames in the macro-expression video sequence, whose contained expression and emotion information are similar to those of the micro-expression, but the intensity is higher than that of the micro-expression. Selecting the second to fifth frames can introduce macro-expression features with different emotion intensities while making the emotion features of the macro-expression as close as possible to those of the micro-expression, thereby enhancing the generalization ability of the model.

[0108] 3) Face detection and positioning: Use the Dlib library to perform face detection and positioning on all the video frame pictures obtained in step 2), and detect the number of faces in the video frame and the distance of the faces from the image boundary.

[0109] 4) Face alignment: Based on face localization, use the Dlib library to determine 68 key facial feature points, and complete face segmentation and face correction.

[0110] Face segmentation means: Use the Dlib vision library to segment the face with a rectangular box.

[0111] Face correction means: Among the 68 key facial feature points detected, mark the key feature point at the left corner of the left eye (i.e., Figure 1 36 in Figure 1 ) and the key feature point at the right corner of the right eye (i.e.,

[0112] ) and the key feature point at the right corner of the right eye (i.e., 45 in ). The line connecting these two points has an angle a with the horizontal line. Obtain the corresponding rotation matrix through this angle a, and perform a rotation transformation on the segmented face to make the line connecting the key feature point at the left corner of the left eye and the key feature point at the right corner of the right eye parallel to the horizontal line, realizing the correction of the face pose, and scale the face, and finally crop it into a face image of size 224×224×3. and as the model input. This can maximize the extraction of emotional information representing facial muscle movement. For macro-expression samples, select the optical flow information between the starting frame and the emotion frame and as the model input. Among them, micro_i represents the serial number of the i-th micro-expression sample, macro_i represents the serial number of the i-th macro-expression sample, flow h represents the horizontal optical flow information, flow v represents the vertical optical flow information, and the size of all optical flow maps is 224×224×3.

[0113] B. Construct an emotion intensity invariance feature extraction network assisted by a cross-domain feature center, and perform deep feature extraction on the optical flow feature maps extracted in step A.

[0114] In step B, the emotion intensity invariance feature extraction network assisted by a cross-domain feature center includes a two-branch emotion intensity invariance feature extraction network and an angle optimization network.

[0115] The overall architecture diagram of the emotion intensity invariance feature extraction network assisted by cross-domain feature centers is as follows Figure 2 shown. First, a two-branch emotion intensity invariance feature extraction network is designed, and the emotion feature extraction parameters of this network are shared among cross-domain samples, where cross-domain samples refer to the set of micro-expression samples and macro-expression samples. This network contains two backbone networks with the same structure but non-shared parameters, namely the horizontal emotion feature extraction network EFEN h (horizontal EmotionFeatureExtraction Network) and the vertical emotion feature extraction network EFEN v (vertical Emotion FeatureExtraction Network), and an Adaptive Emotion feature Focusing Module (AEFM) is added to the end of each branch network respectively, which is used to extract horizontal emotion features and vertical emotion features

[0116]

[0117] Among them, the backbone network is the horizontal emotion feature extraction network EFEN h and the vertical emotion feature extraction network EFEN v The architecture design refers to the current advanced ConvNeXt convolutional neural network structure, and the specific structure is shown in the following table

[0118] Table 1 ConvNeXt Convolutional Neural Network Architecture Design

[0119]

[0120] Among them, the specific structure of the ConvNeXt block structure in the above table is as follows

[0121] Table 2 ConvNeXt Block Structure

[0122]

[0123] The Adaptive Emotion feature Focusing Module (AEFM) is designed based on the channel attention mechanism. It highlights the most important certain feature values of the input 64-dimensional feature f, while suppressing other feature values, so as to balance the relationship between the horizontal and vertical feature components, and thus better extract effective emotion features. The output of this module is

[0124] f emotion= Sigmoid(Linear2(RELU(Linear1(Avgpool(f))))) ⊙ f (3)

[0125] Where Sigmoid is the Sigmoid activation function, RELU is the RELU activation function, Avgpool represents global average pooling, Linear1 represents a fully connected layer with input and output dimensions of 64 and 4, Linear2 represents a fully connected layer with input and output dimensions of 4 and 64, and the dot product symbol ⊙ represents element-wise multiplication.

[0126] To obtain the final emotion feature tensor, the features and are element-wise added and fused at the corresponding positions, and then passed through the emotion feature mapping layer EML (Emotion feature Mapping Layer), which is a fully connected layer. In this method, the size of the emotion feature dimension is taken as 32 to obtain the emotion feature x emotion :

[0127]

[0128] After the samples are extracted by the above double-branch network, the emotion feature x emotion is obtained and sent into the angle optimization network. Due to the cross-domain characteristics, hereinafter, x i is used to represent the emotion feature obtained by the double-branch network for the i-th cross-domain expression sample, N is the size of the emotion feature dimension, is the real number field. Use to represent the emotion feature of the i-th micro-expression sample, to represent the emotion feature of the i-th macro-expression sample. The superscript micro represents micro-expression, and macro represents macro-expression, so as to distinguish between the target domain and the source domain.

[0129] The emotion feature center W = [W1, W2, …, W C refers to a set of trainable parameters, which consists of C N-dimensional feature tensors. C represents the number of categories in the classification task, where W c represents the feature center of the c-th category of emotion, and its feature dimension is consistent with that of x i , that is 1 ≤ c ≤ C, c ∈ Z, Z is the integer field. The feature center is located at the last layer of the angle optimization network and can represent the representative center points of the features of each category of samples. On the hypersphere, the smaller the angle between the emotion feature of a sample and the feature center of a certain category, the higher the emotion similarity between the two, and vice versa. In this method, the micro-expression emotion feature center and the macro-expression emotion feature center During the backpropagation process, the macro-expression samples do not directly participate in the parameter update of the micro-expression emotion feature center, so as to ensure the representativeness of the micro-expression emotion feature center for the emotion features of micro-expression samples.

[0130] Different from the traditional neural network where the classification layer directly outputs the predicted probability distribution of each category, this method calculates the cosine similarity between all expression sample features and the emotion feature centers of the corresponding domains, and normalizes the emotion features of all samples and the feature centers of each category respectively, weakening the intensity property of the features, so that the model focuses on the angular property of the features. The final output of the angular optimization network is the cosine similarity between the sample emotion features and each emotion feature center:

[0131]

[0132]

[0133] where, represents the cosine similarity between the i-th micro-expression sample and the micro-expression feature center of the c-th category, represents the cosine similarity between the i-th macro-expression sample and the macro-expression feature center of the c-th category. The radian values θ of all angles are limited to [0, π]. represents the micro-expression feature center corresponding to the c-th emotion, represents the macro-expression feature center corresponding to the c-th emotion, 1 ≤ c ≤ C, c ∈ Z, where C represents the number of categories in the emotion classification task, and Z represents the set of integers. y i represents the true emotion label category of the i-th sample, and j represents other emotion label categories that are inconsistent with the true emotion label of the i-th sample, that is, 1 ≤ j ≤ C, j ≠ y i , j ∈ Z. In the above formula, for any N-dimensional feature T = (t1, t2, t3…, t N-1 , t N ), the symbol ‖T‖ represents the Euclidean norm (L2 norm) of T:

[0134]

[0135] The dot product symbol A·B represents the inner product operation of tensors. For any two tensors A = (a1, a2,…, a N ) and B = (b1, b2,…, b N ) in the N-dimensional feature space, there is:

[0136]

[0137] The predicted category finally output by the angular optimization network is:

[0138]

[0139] Since cosθ is monotonically decreasing in the domain [0, π], the predicted category is also:

[0140]

[0141] According to the above formula, the cross-domain feature center-assisted emotion intensity invariance feature extraction network can give the predicted category for each cross-domain sample, such as Figure 3 This is a schematic diagram of the predicted sample category of the method model.

[0142] C. Construct a cross-domain loss function and optimization strategy based on the principle of emotion intensity invariance to optimize the training process.

[0143] In step C, constructing a cross-domain loss function and optimization strategy based on the principle of emotion intensity invariance includes the following steps:

[0144] 1) Construct the optimization loss of the emotion classification cosine similarity on the micro-expression and macro-expression discrimination branches. Figure 4 This is a schematic diagram of the optimization principle of the cosine similarity optimization loss for macro-expression emotion classification of the present invention. Figure 5 This is a schematic diagram of the optimization principle of the cosine similarity optimization loss for micro-expression emotion classification of the present invention. In order to make a preliminary classification of the cross-domain expression categories output by the network, combining the cosine similarity metric and the Softmax loss function form, design the cosine similarity-based micro / macro expression emotion classification optimization loss. On the micro-expression discrimination branch, the loss function is:

[0145]

[0146] On the macro-expression discrimination branch, the loss function is:

[0147]

[0148] Among them, B micro +B macro = B. B represents the batch size in the network training process, B micro represents the number of micro-expression samples in this batch, and B macro represents the number of macro-expression samples in this batch. e is the natural constant, and all logarithms in the function are logarithms with base e.

[0149] S represents the radius size of the hypersphere, that is, all sample features and feature centers are mapped to tensors with an Euclidean norm of S, where S > 0.

[0150] 2) Construct a cross-domain sample emotion feature alignment loss function. Figure 6 This is a schematic diagram of the optimization principle of the cross-domain sample emotion feature alignment loss function of the present invention. The fundamental purpose of setting the dual feature centers is to use the macro-expression sample features to guide the learning of the same-class micro-expression features, and at the same time guide the micro-expression features of other different classes to be as far away from each other as possible, so as to learn more discriminative features. Design the cross-domain sample emotion feature alignment loss function L cross (cross-domain emotion feature alignment loss function) to constrain the angular relationship between the micro-expression feature and the macro-expression feature center, so that the micro-expression sample emotion feature is aligned with the macro-expression emotion feature center of the same class, in order to achieve the cross-domain guiding effect of the macro-expression feature center on the micro-expression sample features:

[0151]

[0152] Among them, the functions and are the within-class loss coefficient scalar and the between-class loss coefficient scalar respectively:

[0153]

[0154] Among them, k, a, and b are hyperparameters. and represent the within-class cross-domain cosine similarity and the between-class cross-domain cosine similarity respectively:

[0155]

[0156]

[0157] The cross-domain cosine similarity represents the cosine similarity between the micro-expression sample and the macro-expression feature center. By minimizing this loss function, it can help the model learn cross-domain features and improve the generalization ability of the model for micro-expression samples. and G(θ j ) imitate the characteristics of the Sigmoid function, and their role is to control the optimization speed, that is, to control the parameter update rate in the backpropagation process of the neural network. Among them, a and b are angular bounds (in radians), which are two hyperparameters used to control how sharply the loss function changes when the angular relationship between the feature and the feature center reaches a certain bound.

[0158] 3) Construct a cross-domain emotion feature center regularization loss function. Figure 7This is a schematic diagram of the optimization principle of the cross-domain emotion feature center regularization loss function. There should be consistency between the emotion feature centers of micro-expressions and macro-expressions of the same category. Apply the center regularization loss L center (cross-domain emotion feature centersregularization loss function) to guide the emotion feature centers of micro-expressions and macro-expressions of the same category to approach each other in the feature space:

[0159]

[0160] This loss encourages the model to pay attention to the consistency of the same-category feature centers between macro-expressions and micro-expressions during the learning process, that is, to minimize the angular gap between the emotion feature centers of macro-expressions and micro-expressions of the same category

[0161]

[0162] This regularization loss term helps the feature centers of macro-expressions and micro-expressions of the same category to be more compact in the feature space. This loss term can cooperate with the alignment loss function L cross fully mine the invariant features of emotion intensity shared between macro-expressions and micro-expressions, which helps to prevent the model from overfitting, and promotes the sharing and transfer of features, further strengthening the feature compactness within the cross-domain same category.

[0163] 4) Construct the macro-expression emotion feature compactness loss function, Figure 8 This is a schematic diagram of the optimization principle of the macro-expression emotion feature compactness loss function of the present invention. Macro-expression data has a more compact emotion feature distribution than micro-expression data. In order to better play the guiding role of the macro-expression emotion feature center on the emotion features of micro-expression samples, apply the macro-expression emotion feature compactness loss function L compact (macro-expression emotion feature compactnessloss function) to enhance its within-domain separability and strengthen the representativeness of the macro-expression feature center for various emotion features:

[0164]

[0165] Among them, the functions H and G are the same as in formulas (18) and (19).

[0166] 5) Construct the joint loss function. The joint loss function L in this method total is:

[0167] L total = L micro+ L macro + αL cross + βL center + γL compact (25)

[0168] Among them, α, β, and γ are loss coefficient hyperparameters, which are used to balance the importance of different loss terms. The three work together to control the regularization strength.

[0169] D. Based on the data set constructed in step A, the deep learning neural network constructed in step B, and the loss function used in the back propagation algorithm specified in step C, the emotion intensity invariance feature extraction network assisted by the cross-domain feature center is trained, and the trained deep network model is used for classification and recognition on the data set.

[0170] In step D, the back propagation formula for deep learning training of the emotion intensity invariance feature extraction network assisted by the cross-domain feature center is:

[0171]

[0172] Where Θ is all the network parameters of the emotion intensity invariance feature extraction network assisted by the cross-domain feature center, W macro is the central parameter of macro expression emotion feature, W micro is the central parameter of the micro-expression emotion feature, and μ is the learning rate. Finally, the trained micro-expression recognition model can predict the expression category of the sample.

[0173] In this embodiment, in order to verify the advancement of the micro-expression recognition method of the present invention for cross-domain feature center-assisted emotion intensity invariant feature extraction, the recognition results of the present invention under the above settings are compared with the recognition results without using the optimization strategy in step C, using the CASMEII dataset as the micro-expression dataset, the CK+ dataset as the macro-expression dataset, α is set to 0.1, β is set to 0.1, γ is set to 0.1, S is set to 32, a is set to 0.9, b is set to 1.3, k is set to 80, the emotion feature dimension is set to 32, the number of training rounds is set to 100 rounds, and the optimizer uses A dam optimizer, learning rate μ is set to 0.0001, sample batch size is set to 16, mean and standard deviation are set to 0.5 in data standardization preprocessing, pixel size of all cropped face images and input sample optical flow images is 224×224, original expression sample image is a three-channel RGB image with a bit depth of 24, TVL1 optical flow component feature map is a single-channel grayscale image with a bit depth of 8, LOSO leave-one-sample cross-validation method is used, macro expressions are only used for auxiliary training and are not included in the experimental results. The experimental results are the test results on the micro-expression dataset, and the results are shown in the following table.

[0174] Table 3 Comparison results

[0175] Experimental results Accuracy F1 score Unweighted average recall Using this method 95.3 95.3 93.7 Without using this method 82.9 55.7 54.6

[0176] It can be seen from the experimental results that all indicators using the optimization strategy in the method of the present invention are higher than the results when the optimization strategy in the method of the present invention is not used, which proves the advanced nature of the micro-expression recognition method of the present invention that uses cross-domain feature centers to assist in the extraction of emotion intensity invariant features.

[0177] Example 2

[0178] The difference between the micro-expression recognition method using cross-domain feature center-assisted emotion intensity invariant feature extraction described in Example 1 is that:

[0179] In step A, the micro-expression and macro-expression video sequences are preprocessed, including the following steps:

[0180] 1) Obtaining video frame sequences: performing frame processing on micro-expression and macro-expression video sequences to obtain and store video frame sequences.

[0181] 2) Get the start frame and peak frame: According to the information annotated by experts in the micro-expression dataset, select the micro-expression start frame and peak frame in each video sequence file and store them. For the macro-expression dataset, select the start frame and the second to fifth frames in each video sequence file as emotion frames and store them.

[0182] The micro-expression start frame refers to the first frame in which a micro-expression appears in a micro-expression video sequence.

[0183] The micro-expression peak frame refers to the frame in the micro-expression video sequence where the facial muscle changes are most obvious. It is annotated by experts in the dataset and contains the most micro-expression information.

[0184] The macro expression start frame refers to: the first frame in the macro expression video sequence.

[0185] Macro-expression peak frames refer to the second to fifth frames of the macro-expression video sequence, which contain expression and emotional information similar to micro-expressions, but with higher intensity than micro-expressions.

[0186] 3) Face detection and positioning: Use the Dlib library to perform face detection and positioning on all video frame images obtained in step 2), and detect the number of faces in the video frame and the distance between the face and the image boundary.

[0187] 4) Face alignment: Based on face positioning, the Dlib library is used to determine 68 key feature points on the face to complete face segmentation and face correction.

[0188] Face segmentation refers to using the Dlib visual library to segment faces using rectangular boxes.

[0189] Facial correction means that: among the 68 key feature points of the detected face, the line connecting the key feature point marked at the left corner of the left eye and the key feature point marked at the right corner of the right eye has an angle a with the horizontal line. The corresponding rotation matrix is obtained through this angle a, and the segmented face is rotationally transformed so that the line connecting the key feature point marked at the left corner of the left eye and the key feature point marked at the right corner of the right eye is parallel to the horizontal line, realizing the correction of the face pose, and the face is scaled and finally cropped into a face image of size 224×224×3.

[0190] 5) Obtain the TVL1 optical flow map: Use the TVL1 optical flow extraction method to perform optical flow feature extraction on the starting frame and the peak frame in each aligned micro-expression video sequence, and at the same time perform optical flow feature extraction on the starting frame and the emotion frame in each aligned macro-expression video sequence. The TVL1 optical flow extraction algorithm outputs two parts, including the horizontal optical flow information component and the vertical optical flow information component. For micro-expression samples, select the optical flow information between the starting frame and the peak frame and as the model input, so that the emotional information of facial muscle movement representation can be extracted to the greatest extent. For macro-expression samples, select the optical flow information between the starting frame and the emotion frame and as the model input, where micro_i represents the serial number of the i-th micro-expression sample, macro_i represents the serial number of the i-th macro-expression sample, flow h represents the horizontal optical flow information, flow v represents the vertical optical flow information, and the size of all optical flow maps is 224×224×3.

[0191] In this embodiment, a total of 164 samples of positive, negative, and surprised categories in the SMIC dataset are used as the micro-expression dataset, and a total of 309 samples of positive (i.e., happy category), negative (including angry, disgusted, fearful, and sad categories), and surprised categories in the CK+ dataset are used as the macro-expression dataset.

[0192] In step B, construct an emotion intensity invariance feature extraction network assisted by cross-domain feature centers, and set the emotion feature dimension to 32.

[0193] In step C, construct a joint loss function, set α to 0.1, β to 0.1, γ to 0.1, S to 32, a to 0.9, b to 1.3, and k to 80.

[0194] In this embodiment, the method of the present invention is implemented on the PyTorch framework under Ubuntu, the number of training rounds is set to 100 rounds, the optimizer uses the Adam optimizer, the learning rate μ is set to 0.0001, and the sample batch size is set to 16.

[0195] In this embodiment, in order to verify the advancement of the micro-expression recognition method of the present invention for cross-domain feature center-assisted emotion intensity invariant feature extraction, the recognition results of the present invention under the above-mentioned settings are compared with the results without using this method, and the LOSO leave-one-sample cross-validation method is used. The results are shown in the following table.

[0196] Table 4 Comparison results

[0197] Experimental results Accuracy F1 score Unweighted average recall Using this method 80.5 80.3 81.0 Without using this method 71.3 54.1 52.6

[0198] It can be seen from the experimental results that all indicators using the optimization strategy in the method of the present invention are higher than the results when the optimization strategy in the method of the present invention is not used, which proves the advanced nature of the micro-expression recognition method of the present invention that uses cross-domain feature centers to assist in the extraction of emotion intensity invariant features.

Claims

1. A micro-expression recognition method for extracting emotion intensity invariant features assisted by cross-domain feature centers, characterized in that: The steps include: A. Preprocessing of micro-expression and macro-expression video sequences, including: obtaining video frame sequences, obtaining starting frames and peak frames, face detection and positioning, face alignment, and obtaining TVL1 optical flow feature maps; B. Construct a cross-domain feature center-assisted emotion intensity invariance feature extraction network to perform deep feature extraction on the optical flow feature map extracted in step A. C. Construct a cross-domain loss function and optimization strategy based on the principle of emotion intensity invariance to optimize the training process; D. Based on the data set constructed in step A, the deep learning neural network constructed in step B and the loss function specified in step C, the emotion intensity invariance feature extraction network assisted by the cross-domain feature center is trained, and the trained deep network model is used for classification and recognition on the data set.

2. The micro-expression recognition method for extracting emotion intensity invariant features assisted by cross-domain feature centers according to claim 1 is characterized in that: In step A, the micro-expression and macro-expression video sequences are preprocessed, including the following steps: 1) Obtaining video frame sequences: performing frame processing on micro-expression and macro-expression video sequences to obtain and store video frame sequences; 2) Obtaining the starting frame and peak frame: According to the information annotated by experts in the micro-expression dataset, the micro-expression starting frame and peak frame in each video sequence file are selected and stored; for the macro-expression dataset, the starting frame and the second to fifth frames in each video sequence file are selected as emotion frames and stored; The micro-expression start frame refers to: the first frame in the micro-expression video sequence where the micro-expression appears; The micro-expression peak frame refers to the frame in the micro-expression video sequence where the facial muscle changes are most obvious. It is annotated by experts in the dataset and contains the most micro-expression information. The macro expression start frame refers to: the first frame in the macro expression video sequence; Macro-expression emotion frames refer to the second to fifth frames of the macro-expression video sequence, which contain stronger expression and emotion information than micro-expressions; 3) Face detection and positioning: Use the Dlib library to perform face detection and positioning on all video frame images obtained in step 2), and detect the number of faces in the video frame and the distance between the face and the image boundary; 4) Face alignment: Based on face positioning, the Dlib library is used to determine 68 key facial feature points to complete face segmentation and face correction; Face segmentation refers to: using the Dlib visual library to segment faces using rectangular boxes; Face correction means: among the 68 key feature points detected on the face, mark the key feature points of the left eye corner and the Note that there is an angle a between the line connecting the key feature points at the right corner of the right eye and the horizontal line. The corresponding rotation matrix is ​​obtained through the angle a, and the segmented face is rotated to make the line connecting the key feature points at the left corner of the left eye and the key feature points at the right corner of the right eye parallel to the horizontal line, so as to correct the face posture and scale the face; 5) Obtain TVL1 optical flow feature map: Use TVL1 optical flow extraction method to extract optical flow features from the start frame and peak frame of each aligned micro-expression video sequence, and extract optical flow features from the start frame and emotion frame of each aligned macro-expression video sequence; The output of TVL1 optical flow extraction algorithm includes two parts: horizontal optical flow information component and vertical optical flow information component. For micro-expression samples, the starting frame is selected. and peak frame The optical flow information between and As model input; for macro expression samples, select the starting frame Frames with emotions The optical flow information between and As the model input, micro_i represents the number of the i-th micro-expression sample, macro_i represents the number of the i-th macro-expression sample, and flow h Represents horizontal optical flow information, flow v Represents vertical optical flow information.

3. The micro-expression recognition method for extracting emotion intensity invariant features assisted by cross-domain feature centers according to claim 2 is characterized in that: In step 4), the face is scaled and finally cropped into a face image of size 224×224×3; in step 5), the size of all optical flow feature maps is 224×224×3.

4. The micro-expression recognition method for extracting emotion intensity invariant features assisted by cross-domain feature centers according to claim 1 is characterized in that: In step B, the emotion intensity invariance feature extraction network assisted by the cross-domain feature center includes a dual-branch emotion intensity invariance feature extraction network and an angle optimization network: A dual-branch emotion intensity invariance feature extraction network, whose emotion feature extraction parameters are shared between cross-domain samples. Cross-domain samples refer to the collection of micro-expression samples and macro-expression samples. The network contains two branch backbone networks with the same structure but no shared parameters, namely the horizontal emotion feature extraction network EFEN. h and vertical emotion feature extraction network EFEN v , and add the emotion feature adaptive focusing module AEFM at the end of each branch network to extract the horizontal emotion features and vertical emotional features The emotional feature adaptive focusing module AEFM is designed based on the channel attention mechanism. The output of this module is: f emotion =Sigmoid(Linear2(RELU(Linear1(Avgpool(f)))))⊙f(3) Among them, Sigmoid is the Sigmoid activation function, RELU is the RELU activation function, and Avgpool represents the global average pooling. Linear1 represents a fully connected layer with input and output dimensions of 64 and 4, Linear2 represents a fully connected layer with input and output dimensions of 4 and 64, and the dot product symbol ⊙ represents element-by-element multiplication; The characteristics and After adding and fusing the corresponding elements one by one, they are passed through the emotion feature mapping layer EML, that is, a fully connected layer. In this method, the emotion feature dimension size is 32 to obtain the emotion feature x emotion : After the sample is extracted through the above two-branch network, the emotional feature x is obtained emotion And sent to the angle optimization network; due to the cross-domain characteristics, use x i represents the emotional features obtained after the i-th cross-domain expression sample passes through the dual-branch network, N is the dimension of the emotional feature, is the field of real numbers; use represents the emotional characteristics of the i-th micro-expression sample, represents the emotional features of the i-th macro-expression sample, the superscript micro represents micro-expression, and macro represents macro-expression; Emotion feature center W=[W1,W2,…,W C ] refers to a set of trainable parameters, consisting of C N-dimensional feature tensors, where C represents the number of categories in the classification task, and W c represents the feature center of the c-th type of emotion, and its feature dimension is the same as x i Keep consistent, that is 1≤c≤C, c∈Z, Z is an integer domain; the feature center is located in the last layer of the angle optimization network, representing the representative center point of the sample features of each category; on the hypersphere, the smaller the angle between the emotional feature of a sample and the feature center of a certain category, the higher the emotional similarity between the two, and vice versa; set the micro-expression emotional feature center Hehong Expression and Emotion Feature Center The cosine similarity between all expression sample features and the emotion feature centers of the corresponding domain is calculated. The emotion features of all samples and the feature centers of each category are normalized respectively. The final output of the angle optimization network is the cosine similarity between the sample emotion features and the centers of each emotion feature: represents the cosine similarity between the i-th micro-expression sample and the c-th micro-expression feature center, represents the cosine similarity between the i-th macro expression sample and the c-th macro expression feature center, and the radian values ​​θ of all angles are limited to [0,π]; represents the center of the micro-expression feature corresponding to the c-th emotion, represents the macro expression feature center corresponding to the c-th emotion, 1≤c≤C,c∈Z, where C represents the number of categories in the emotion classification task, and Z represents an integer set; i represents the true emotion label category of the i-th sample, and j represents other emotion label categories that are inconsistent with the true emotion label of the i-th sample, that is, 1≤j≤C,j≠y i ,j∈Z; In the above formula, for any N-dimensional feature T=(t1,t2,t3…,t N-1 ,t N ), the symbol ‖T‖ represents the Euclidean norm of T: The dot product symbol A·B represents the tensor inner product operation. For any two tensors A=(a1, a2,…, a N ) and B=(b1,b2,…,b N ),have: The predicted category of the final output of the angle optimization network is: Since cosθ is monotonically decreasing on the domain [0,π], the predicted category is: According to the above formula, the emotion intensity invariance feature extraction network assisted by the cross-domain feature center gives a predicted category for each cross-domain sample.

5. The micro-expression recognition method for extracting emotion intensity invariant features assisted by cross-domain feature centers according to claim 1 is characterized in that: In step C, a cross-domain loss function and optimization strategy based on the principle of emotion intensity invariance is constructed, including the following steps: 1) Construct the cosine similarity optimization loss for emotion classification on the micro-expression and macro-expression discrimination branches, and design the cosine similarity optimization loss for emotion classification of micro-expression and macro-expression. On the micro-expression discrimination branch, the loss function is: On the macro expression discrimination branch, the loss function is: Among them, B micro +B macro =B; B represents the batch size during network training, B micro represents the number of micro-expression samples in the batch, B macro Indicates the number of macro expression samples in the batch; e is a natural constant, and the log in the function is the logarithm with base e; S represents the radius of the hypersphere, that is, all sample features and feature centers are mapped to a tensor with Euclidean norm S, S>0; 2) Constructing the cross-domain sample emotion feature alignment loss function: Among them, the function and They are the intra-class loss coefficient scalar and inter-class loss coefficient scalar respectively: Among them, k, a, and b are hyperparameters. and Respectively represent the same cross-domain cosine similarity and heterogeneous cross-domain cosine similarity: 3) Construct a cross-domain emotion feature center regularization loss function and apply a center regularization loss L to the two feature centers of micro-expressions and macro-expressions. center : This loss encourages the model to focus on the consistency of the centers of the same category features between macro-expressions and micro-expressions during the learning process, that is, to minimize the angular gap between the centers of the same category macro-expressions and micro-expression emotional features. 4) Construct a macro-expression emotion feature compaction loss function and apply the macro-expression emotion feature compaction loss function L to the macro-expression domain. compact : Among them, functions H and G are the same as formulas (18) and (19); 5) Construct the joint loss function L total for: L total =L micro + L macro + αL cross + βL center + γL compact (25) Among them, α, β, and γ are loss coefficient hyperparameters.

6. The micro-expression recognition method for extracting emotion intensity invariant features assisted by cross-domain feature centers according to claim 1 is characterized in that: In step D, the back propagation formula for training the emotion intensity invariance feature extraction network assisted by the cross-domain feature center is: Where Θ is all the network parameters of the emotion intensity invariance feature extraction network assisted by the cross-domain feature center, W macro is the central parameter of macro expression emotion feature, W micro is the central parameter of the micro-expression emotional feature, and μ is the learning rate; finally, the trained micro-expression recognition model can predict the expression category of the sample.

Citation Information

Cited By

  • Facial micro-expression recognition method

    CN120783379A