Cross-domain micro-expression recognition method based on gradient inversion domain adaptation and double-flow optimization
By adopting gradient inversion domain adaptation and dual-stream optimization methods in cross-domain micro-expression recognition, combining the dual-domain semantic adversarial learning network and a single-domain characterization network, the inter-domain offset problem is solved, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510296186.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
In the cross-domain micro-expression recognition process, there are differences in the feature distribution of training and test data, resulting in inter-domain offset problems, affecting the generalization ability and recognition accuracy of the model.
Using a method based on gradient inversion domain adaptation and dual-stream optimization, a dual-domain semantic adversarial learning network and a single-domain representation network are used, combining global dynamic extractors, macro-micro knowledge sharing modules, gradient inversion modules, local variation extractors and microdomain expression modules to reduce inter-domain offsets and improve recognition performance.
It effectively reduces inter-domain offsets, improves the model's adaptability in the target domain and the accuracy and robustness of micro-expression recognition.
Smart Images

Figure CN120220209A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross-domain micro-expression recognition method based on gradient inversion domain adaptation and two-stream optimization, belonging to the technical fields of deep learning and pattern recognition. Background Art
[0002] Micro-expression recognition technology has shown important application potential in multiple fields. For example, in the aspect of mental health, it can assist in diagnosing mental illnesses such as depression, anxiety, and post-traumatic stress disorder (PTSD), thereby providing more accurate support for personalized psychological intervention. In the process of security monitoring and law enforcement, this technology can be used to analyze the potential threat behaviors of criminal suspects, improving the scientific nature of interrogation and the accuracy of judgment. In addition, in the field of intelligent interaction, micro-expression recognition can enhance the perception of human emotions by artificial intelligence, making human-computer communication more natural and intelligent.
[0003] In response to the challenges faced by micro-expression recognition, researchers have proposed various optimization strategies, mainly including the following three categories of methods. First is the method based on manual feature extraction, which relies on artificially designed feature descriptions, such as local binary pattern (LBP), second-order features, and optical flow features, etc., using local image information, statistical analysis, or dynamic change patterns to capture micro-expressions. However, this method has relatively high requirements for prior knowledge and is difficult to adapt to the facial feature changes of different individuals, so it has certain limitations.
[0004] Second is the strategy based on single-domain deep learning. This method constructs a neural network to automatically learn the spatio-temporal features of micro-expressions, thereby reducing the dependence on artificial features. For example, the local and global joint feature extraction method based on convolutional neural network (CNN) has improved the recognition accuracy to a certain extent. However, due to the small scale of the micro-expression dataset, the performance of the deep learning model is still limited by the amount of data and it is difficult to fully exert its potential capabilities.
[0005] Finally is the cross-domain transfer learning method. The core idea of this strategy is to use data in related fields to make up for the deficiency of micro-expression data. Since macro-expressions and micro-expressions may share some action units (AUs) under the same emotion category, macro-expression data is often used to enhance the effect of micro-expression recognition. For example, singular value decomposition (SVD) is used to achieve feature transfer from macro-expressions to micro-expressions.
[0006] However, the scale of the micro-expression dataset is still an important factor restricting the effect of transfer learning, and at the same time, how to effectively extract and transfer high-quality information is still an urgent problem to be solved. Summary of the Invention
[0007] In the cross-domain micro-expression recognition task, due to the differences in the feature distributions between the training data and the test data, traditional methods often struggle to maintain performance in the target domain. To address the deficiencies of the existing technologies, the present invention proposes a cross-domain micro-expression recognition method based on gradient inversion domain adaptation and two-stream optimization. This method mainly consists of two core networks: a dual-domain semantic adversarial learning network and a single-domain representation network, so as to improve the adaptability and accuracy of cross-domain micro-expression recognition.
[0008] The focus of the present invention is to mitigate the domain shift problem that often occurs in cross-domain recognition, aiming to maximize the integration of the shared knowledge between domains and the knowledge of the micro-expression domain, enabling the network to pay attention to the dynamics and detailed changes of micro-expressions under the guidance of macro-expressions, and further improving the performance of micro-expression recognition. Summary of the Invention:
[0010] A cross-domain micro-expression recognition method based on gradient inversion domain adaptation and two-stream optimization, including a dual-domain semantic adversarial learning network, a single-domain representation network, and a bilinear pooling fusion unit. Among them, the dual-domain semantic adversarial learning network includes a global dynamics extractor, a macro-micro knowledge sharing module, and a gradient inversion module. The single-domain representation network includes a local variation extractor and a micro-domain expression module.
[0011] The technical problem solved by the present invention is that although there are certain similarities in facial features between macro-expressions and micro-expressions, significant inter-domain shift problems will still be encountered in the process of cross-domain micro-expression recognition. Specifically, there are often differences in the acquisition sources of the training and test data, such as different camera devices, changes in lighting conditions, and differences in individual facial features. These factors lead to a large deviation in the distributions of the features learned in the macro-expression domain and the features in the micro-expression domain, making the effect of directly migrating the features of macro-expressions to micro-expression recognition limited. As shown in the figure, even though there are certain commonalities in their facial structures, due to the differences in the forms of expression and scales, directly applying the features of macro-expressions for micro-expression classification cannot significantly improve the recognition effect. Therefore, how to effectively solve the inter-domain shift problem becomes a key challenge.
[0012] In order to alleviate this problem, the present invention proposes a cross-domain micro-expression recognition method combining gradient inversion domain adaptation and dual-stream optimization. The method consists of three key components: a dual-domain semantic adversarial learning network, a single-domain feature modeling network, and a bilinear pooling fusion unit. The core goal of the dual-domain semantic adversarial learning network is to promote the migration of knowledge from the source domain (macro-expression domain) to the target domain (micro-expression domain), thereby reducing the negative impact of inter-domain offset on the generalization ability of the model. The network includes a global dynamic extraction module for capturing global motion information in macro-expressions and micro-expressions, and identifying the similarities between the two in facial expression dynamics. The macro-micro knowledge sharing module shares global dynamic features between the source domain and the target domain to help the model maintain consistent feature representations between different domains. At the same time, the gradient inversion module performs adversarial training through the gradient inversion mechanism GRL, effectively reducing the feature distribution differences between the source domain and the target domain, thereby improving the adaptability of the model in the target domain.
[0013] The single-domain representation network focuses on fine-tuning the micro-expression features in the target domain (i.e., the micro-expression domain). The network includes a local change extraction module, which is used to extract subtle local change features from the inter-frame differences of micro-expressions. These features can reveal the subtle differences in the time scale and motion amplitude of micro-expressions. Next, the micro-domain expression module optimizes the expression of the features of the target domain to ensure that the detailed information of the micro-expressions is fully preserved. Finally, the bilinear fusion unit fuses the shared knowledge of the source domain with the micro-expression features of the target domain through a bilinear pooling operation, so that the model can better capture the dynamics and detailed changes of micro-expressions with the guidance of macro-expressions, and further improve the effect of micro-expression recognition.
[0014] Terminology explanation:
[0015] 1. Peak frame: It is the most significant moment in the process of micro-expression change, usually showing the maximum intensity of expression. This frame corresponds to the maximum amplitude of expression change and can effectively reflect the sudden change of emotion.
[0016] 2. Start frame: It is the starting point of micro-expression changes. Usually the initial stage of expression changes is relatively smooth.
[0017] 3. Loss function: The loss function is used to evaluate the degree of inconsistency between the model's predicted value and the true value. The smaller the loss function, the better the robustness of the model. The loss function can guide model learning.
[0018] 4. Gradient inversion mechanism GRL: Gradient inversion mechanism is a deep learning technology used to solve domain adaptation problems, especially in cross-domain learning. GRL introduces reverse gradient updates in the network, allowing the model to learn features that have good generalization capabilities for both the source domain and the target domain. During the training process, GRL forces the model to learn features shared between the two domains through adversarial training, while suppressing domain-specific differences. This mechanism is widely used in tasks such as cross-domain feature learning and domain adaptation, which helps to improve the performance and robustness of the model in the target domain.
[0019] The technical solution of the present invention is as follows:
[0020] A cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization, comprising:
[0021] Construct and train a cross-domain micro-expression recognition model; the cross-domain micro-expression recognition model includes a dual-domain semantic adversarial learning network and a single-domain representation network; the dual-domain semantic adversarial learning network includes a global dynamic extractor, a macro-micro knowledge sharing module, and a gradient inversion module; the single-domain representation network includes a local change extractor, a micro-domain expression module, a bilinear pooling fusion unit, and a classifier;
[0022] The data to be recognized is pre-processed and then fed into the trained cross-domain micro-expression recognition model to perform cross-domain micro-expression recognition; including:
[0023] A. Input the starting frame and peak frame of the macro-expression and micro-expression video frames into the global dynamic extractor to extract the dynamic optical flow features of each macro-expression and micro-expression video frame;
[0024] B. Input the dynamic optical flow features obtained from step A into the macro-micro knowledge sharing module to obtain macro-micro knowledge features;
[0025] C. The macro-micro knowledge features obtained from step B are input into the gradient inversion module to further purify the shared features of the two domains by reversing the gradient; at the same time, the shared features are input into the domain discriminator for domain discrimination;
[0026] D. Input the starting frame and peak frame of the micro-expression into the local change extractor to obtain the local change feature descriptor of the micro-expression;
[0027] E. Input the local change feature descriptor of the micro-expression obtained in step D into the micro-domain expression module to further capture the local dynamic changes of the micro-expression and obtain the single-domain feature;
[0028] F. Fuse the shared features of the two domains obtained in step C and the single-domain features obtained in step E through a bilinear pooling fusion unit to obtain a high-dimensional feature representation with rich spatio-temporal information and domain generalization ability; then, input the obtained high-dimensional feature representation into the classifier of the single-domain characterization network for final micro-expression recognition and classification.
[0029] Preferably according to the present invention, in step A, for each video sequence x in micro-expressions and macro-expressions mi / ma is segmented into a series of video frames, expressed as formula (1):
[0030] x mi / ma = [f mi / ma,1 , f mi / ma,2 , f mi / ma,3 ,..., f mi / ma,ap ,..., f mi / ma,n , E(f mi / ma,ap ) = E max (1);
[0031] where n is the total number of micro-expression and macro-expression video frames, f mi / ma,ap is the peak frame of the micro-expression and macro-expression video sequences, E is the expression intensity of the expression frame sequence; for the input micro-expression start frames f mi,1 and f mi,ap , the TV-L1 optical flow obtains the micro-expression optical flow E i (u, v) and the macro-expression optical flow E a (u, v) as shown in formula (2) and formula (3):
[0032]
[0033]
[0034] where x, y represent pixel points, u, v respectively represent the displacement components of the optical flow field in the x and y directions, λ is the regularization parameter of the smoothness term, used to balance the data term and the regularization term; at the same time, the L1 constraint |f(x + u, y + v) - f(x, y)| of the data term makes the optical flow calculation more robust to illumination calculation, and the regularization term controls the smoothness of the optical flow;
[0035] Solve this optimization problem by the method of Lagrange multipliers and use Gaussian-Seidel iteration to solve it, as shown in formula (4):
[0036]
[0037] Finally, obtain the dynamic optical flow features, u t and v tDenote the optical flow components at the t-th iteration, where u t represents the optical flow component in the horizontal direction, and v t represents the optical flow component in the vertical direction; the optical flow field finally obtains the motion information of the pixel at time t, that is, the displacement of the pixel from the previous frame to the current frame where u t represents the optical flow component in the horizontal direction, and v t represents the optical flow component in the vertical direction.
[0038] According to the preferred embodiment of the present invention, in step B, in each macro-micro knowledge sharing module, downsampling is performed through an initial convolutional layer, and then residual connection is performed with the second convolutional layer. The function of each convolutional layer is to implement feature extraction; the structure B of each macro-micro knowledge sharing module is expressed by formulas (5) to (7):
[0039] Z m.k = Β k (Z m,k-1 ; θ m,k ) = Conv2(BN(ReLU(Conv1(Z m,k-1 ; W 1,k ))) ; W 2,k ) + Z m,k-1 (5);
[0040] Z iu = Β K (Β K-1 (... Β2(Β1(Ei(u,v))))) (6);
[0041] Z au = Β K (Β K-1 (... Β2(Β1(Ea(u,v))))) (7);
[0042] where B k represents the k-th macro-micro knowledge sharing module, Z m,k is the output of the k-th macro-micro knowledge sharing module, Z iu and Z au respectively represent the outputs of the optical flow features of micro-expression and macro-expression passing through the macro-micro shared knowledge module. Conv1 and Conv2 respectively represent two convolutional layers, W 1,k and W 2,k are convolutional kernels, θ m,k is the parameter of the convolutional block, K is the number of macro-micro knowledge sharing modules, and the outputs Z iu and Z au are collectively referred to as Z s , that is, the information features of the output macro-expression domain and micro-expression domain.
[0043] Preferably according to the present invention, in step C, the macro-micro knowledge sharing module is reversely trained in an adversarial manner through the gradient reversal layer in the gradient reversal module to obtain the shared features of the two domains. The role of the gradient reversal layer is to reverse the gradient during the training process, so that while the feature extraction network optimizes the classification task, it reduces the inter-domain distribution difference through an adversarial manner, thereby learning domain-invariant features. The specific role is shown in formulas (8) to (10):
[0044]
[0045] Among them, GRL represents the gradient reversal layer, λ is the reversal coefficient, which is dynamically updated during training, p is the current training progress; Z s is the output feature of the macro-micro sharing module, which is input into the gradient reversal layer to obtain the output I represents the identity matrix. After that, the output feature representation is continuously input into the domain classifier D(·) for classification. The role of the domain classifier D(·) is to classify the expressions of micro-expressions and macro-expressions into their respective domains, as shown in formula (11):
[0046]
[0047] Among them, θ d is the parameter of the domain classifier; is the output of the gradient reversal layer the classification vector output after being input into the domain classifier D(·);
[0048] For the domain classifier D(·), as shown in formula (12):
[0049]
[0050] Among them, is the domain label probability predicted by the domain classifier for the current sample i to be tested, that is, the output of the domain classifier; d i is the domain label of sample i. Specifically, as shown in formula (13):
[0051]
[0052] The gradient in the macro-micro shared knowledge module is expressed as formula (14):
[0053]
[0054] Preferably according to the present invention, in step D, for each video sequence x in the micro-expression mi is segmented into a series of video frames; the local change feature of the micro-expression is expressed as formula (15)
[0055]
[0056] Among them, and respectively refer to the peak frame and the starting frame of the current micro-expression sequence. I represents the number of frames of the micro-expression sequence. refers to the output after the micro-expression sequence passes through the local variation extractor, that is, the local variation feature descriptor of the micro-expression.
[0057] According to the preference of the present invention, in step E, through multi-level feature extraction of the input local variation data, from low-level local variation to high-level abstract features, the specific structure of each micro-domain expression module is shown in formulas (16) to (18):
[0058] D k =N k (D k-1 ; θ k ), k = 1, 2,..., K (16);
[0059] X l =σ(W l *X l-1 +b l ), l = 1, 2,..., L k (17);
[0060] X l.pool =Pool(X l ), l = 1, 2,..., L k (18);
[0061] Formula (16) indicates that the output D k of the k-th micro-domain expression module is obtained according to the input D k-1 of the (k - 1)-th micro-domain expression module. N k (·; θ k ) represents the k-th micro-domain expression module, and θ k is the parameter of the k-th micro-domain expression module. The k-th micro-domain expression module has L k layers, and each layer includes a convolutional layer and an activation function, as shown in formula (17). σ represents the activation function, W l and b l respectively represent the convolution kernel and the bias of the first convolutional layer. * represents the convolution operation. X l-1 and X l respectively represent the input and output of the first convolutional layer. After the convolution operation, a pooling operation is also performed on each layer. The pooling operation is expressed as formula (18), Pool represents the pooling operation, and X l.poolDenotes the output after the pooling operation is performed on the output of each micro-domain expression module after the convolution operation. Each layer of the convolution operation not only performs spatial feature extraction but also downsamples the feature map through the pooling operation, and finally obtains the output D of the k-th micro-domain expression module k 。
[0062] According to the preference of the present invention, in step F, a bilinear pooling method is introduced. By performing a product operation on different features, more rich and accurate cross-domain and single-domain interaction features are extracted; specifically expressed as formula (19):
[0063]
[0064] Among them, Z s Denotes the dual-domain feature passing through the macro-micro knowledge sharing module, D k Denotes the single-domain feature of the output of the k-th micro-domain expression module, Denotes the outer product operation, generating a matrix of d s ×d k Introduce a dimensionality reduction operation to map the generated high-dimensional features to a manageable dimensional space; specifically, through a linear transformation matrix To perform feature dimensionality reduction; the dimensionality-reduced feature F fuse Is expressed as formula (20):
[0065] F fuse = W·vec(F bilinear ) (20);
[0066] Among them, vec(·) means expanding the matrix F bilinear Column by column into a vector, W is the linear transformation matrix, d fuse Is the dimensionality of the dimensionality-reduced feature;
[0067] The feature F fuse After linear dimensionality reduction undergoes a non-linear transformation; the LeakyReLU activation function is used to introduce non-linear factors, and the specific transformation formula is shown in formula (21):
[0068] F nonlinear = σ(F fuse + b) (21);
[0069] Among them, σ(·) is the LeakyReLU activation function, b is the bias term; F nonlinear Denotes the output of the dimensionality-reduced feature F fuse After passing through the activation function;
[0070] Next, input F nonlinear Into the classifier for classification, specifically expressed as formula (22):
[0071] Y cls = W cls ·F nonlinear + b cls (22);
[0072] Wherein, is the weight matrix of the classifier, is the bias term, C is the number of categories. After linear transformation, the obtained output represents the scores of each category of micro-expression categories "positive", "negative", and "surprise";
[0073] Perform Softmax activation on the result of the linear transformation. The Softmax function converts the category scores into a probability distribution, as shown in formula (23):
[0074]
[0075] Wherein, is a vector of length C, representing the prediction probabilities of each category of micro-expression categories "positive", "negative", and "surprise".
[0076] According to the preference of the present invention, the cross-domain micro-expression recognition model is optimized and iteratively updated by using the single-domain sentiment classification recognition loss and the double-domain sentiment domain adaptation loss respectively; including:
[0077] The output of the sentiment classifier in the single-domain representation network is p i ∈P, representing the prediction probabilities that the current i-th micro-expression to be measured belongs to each sentiment category; the goal of sentiment classification is to minimize the cross-entropy loss between the true label y i and the predicted label p i ; the mathematical expression of the cross-entropy loss is as shown in formula (24):
[0078]
[0079] Wherein, N is the number of samples, K is the number of sentiment categories, is the true label of sample i in sentiment category k, is the probability that the model predicts sample i belongs to sentiment category k;
[0080] In the gradient inversion module, after the input feature Z undergoes an inversion operation, it is input into the domain classifier D v ; the goal of the domain classifier is to determine whether the current sample belongs to the macro-expression domain or the micro-expression domain; the output of the domain classifier is representing the probability that the current input sample belongs to the micro-expression domain, and the loss function is as shown in formula (25):
[0081]
[0082] Wherein: is the domain label probability predicted by the domain classifier for the current sample i to be tested, d i is the domain label of sample i.
[0083] The beneficial effects of the present invention are as follows:
[0084] (1) Combination of global and local motion features: By introducing a global dynamic extractor and a local variation extractor, motion features are extracted in the macro-expression domain and the micro-expression domain respectively. This multi-level feature extraction method can capture both the global dynamics of macro-expressions and the local detail changes of micro-expressions simultaneously, thus providing a more comprehensive and detailed feature representation for cross-domain micro-expression recognition.
[0085] (2) Application of the gradient inversion module in shared knowledge modeling: Using the gradient inversion module (GRL) for adversarial training, the aim is to narrow the feature distribution differences between the macro-expression and micro-expression domains, thereby achieving effective inter-domain knowledge transfer. This not only alleviates the inter-domain shift problem but also improves the recognition performance of the model in the target domain (micro-expression domain).
[0086] (3) Innovative design of the bilinear fusion unit: A bilinear fusion unit is proposed to deeply fuse the shared knowledge of the macro-expression domain and the micro-expression domain with the specific knowledge of the micro-expression domain. This design enhances the model's sensitivity to micro-expression details, making cross-domain micro-expression recognition more accurate and robust. Brief Description of the Drawings
[0087] Figure 1 is a schematic flow chart of a cross-domain micro-expression recognition method based on gradient inversion domain adaptation and two-stream optimization;
[0088] Figure 2 is a visualization diagram of the optical flow features obtained from the global dynamic extractor;
[0089] Figure 3 is a schematic diagram of the confusion matrix of the CASMEII dataset;
[0090] Figure 4 is a schematic diagram of the confusion matrix of the SAMM dataset;
[0091] Figure 5 is a schematic diagram of the confusion matrix of the SMIC dataset; Detailed Embodiments
[0092] The present invention will be further described below through embodiments in conjunction with the drawings, but is not limited thereto.
[0093] Embodiment 1
[0094] A cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization, such as Figure 1 As shown, including:
[0095] Construct and train a cross-domain micro-expression recognition model; the cross-domain micro-expression recognition model includes a dual-domain semantic adversarial learning network and a single-domain representation network; the dual-domain semantic adversarial learning network includes a global dynamic extractor, a macro-micro knowledge sharing module, and a gradient inversion module; the single-domain representation network includes a local change extractor, a micro-domain expression module, a bilinear pooling fusion unit, and a classifier;
[0096] The interaction between the two networks is completed by fusing domain features of the macro-micro knowledge sharing module and the micro-domain expression module through a bilinear pooling fusion unit, as follows: The dual-domain semantic adversarial learning network is designed to transfer knowledge from the source domain (macro-expression domain) to the target domain (micro-expression domain) to reduce the impact of inter-domain offset on the generalization performance of the model. The dual-domain semantic adversarial learning network includes a global dynamic extractor, which is used to extract global motion information in macro-expressions and micro-expressions, and capture similar facial expression dynamics between the two. The macro-micro knowledge sharing module is responsible for sharing global dynamic features between the macro-expression domain and the micro-expression domain, so that the model can maintain consistent representations between different domains. In addition, the gradient inversion module implements adversarial training through the gradient inversion mechanism (GRL), further narrowing the feature distribution difference between the macro-expression domain and the micro-expression domain, and enhancing the adaptability of the model in the micro-expression domain. The single-domain representation network focuses on accurately modeling micro-expression features in the target domain (micro-expression domain). The single-domain representation network includes a local change extractor, which extracts local change features from the inter-frame differences of micro-expressions. These features can capture the subtle changes in micro-expressions in time scale and motion amplitude. Then, the micro-domain expression module is used to express and optimize the features of the micro-expression domain to ensure that the detailed information of the micro-expression can be fully preserved. Finally, a bilinear pooling fusion unit is designed to fuse the shared knowledge between domains with the knowledge of the micro-expression domain, so that the network can pay attention to the dynamic and detailed changes of micro-expressions through the distinguishing features between domains, further improving the performance of micro-expression recognition.
[0097] The data to be recognized is pre-processed and then fed into the trained cross-domain micro-expression recognition model to perform cross-domain micro-expression recognition; including:
[0098] A. Input the starting frame and peak frame of the macro-expression and micro-expression video frames into the global dynamic extractor to extract the dynamic optical flow features of each macro-expression and micro-expression video frame;
[0099] B. Input the dynamic optical flow features obtained from step A into the macro-micro knowledge sharing module to obtain macro-micro knowledge features;
[0100] C. The macro-micro knowledge features obtained from step B are input into the gradient inversion module to further purify the shared features of the two domains by reversing the gradient; meanwhile, the shared features are input into the domain discriminator for domain discrimination;
[0101] D. The starting frame and peak frame of the micro-expression are input into the local variation extractor to obtain the local variation feature descriptor of the micro-expression;
[0102] E. The local variation feature descriptor of the micro-expression obtained in step D is input into the micro-domain expression module to further capture the local dynamic changes of the micro-expression and obtain the single-domain features;
[0103] F. The shared features of the two domains obtained in step C and the single-domain features obtained in step E are fused through the bilinear pooling fusion unit to obtain a high-dimensional feature representation with rich spatio-temporal information and domain generalization ability; then, the obtained high-dimensional feature representation is input into the classifier of the single-domain characterization network for final micro-expression recognition and classification.
[0104] Example 2
[0105] A cross-domain micro-expression recognition method based on gradient inversion domain adaptation and two-stream optimization according to Example 1, characterized in that:
[0106] In step A, for each video sequence x in micro-expressions and macro-expressions mi / ma is segmented into a series of video frames, expressed as formula (1):
[0107] x mi / ma = [f mi / ma,1 , f mi / ma,2 , f mi / ma,3 ,..., f mi / ma,ap ,..., f mi / ma,n , E(f mi / ma,ap ) = E max (1);
[0108] where n is the total number of micro-expression and macro-expression video frames, f mi / ma,ap is the peak frame of the micro-expression and macro-expression video sequence, and E is the expression intensity of the expression frame sequence; since the facial changes of micro-expressions are very subtle, in order to let the model fully perceive the feature changes of micro-expressions, the optical flow features between the starting frame and the peak frame in micro-expressions are calculated, as Figure 2 shown, and the same applies to macro-expressions. For the input micro-expression starting frames f mi,1 and f mi,ap , the TV-L1 optical flow obtains the micro-expression optical flow E i (u, v) and the macro-expression optical flow E a (u, v) by solving the following variational optimization problem, as shown in formula (2) and formula (3):
[0109]
[0110]
[0111] Among them, x and y represent pixel points, u and v respectively represent the displacement components of the optical flow field in the x and y directions, and λ is the regularization parameter of the smoothness term, which is used to balance the data term and the regularization term; at the same time, through the L1 constraint |f(x+u,y+v)-f(x,y)| of the data term, the optical flow calculation is more robust to the illumination calculation, and the regularization term controls the smoothness of the optical flow;
[0112] Solve this optimization problem by the Lagrange multiplier method and use the Gauss-Seidel iteration to solve it, as shown in formula (4):
[0113]
[0114] Finally, the dynamic optical flow feature, u t and v t represent the optical flow components at the t-th iteration. Among them, u t represents the optical flow component in the horizontal direction (x-axis direction), and v t represents the optical flow component in the vertical direction (y-axis direction); the optical flow field finally obtains the motion information of the pixel point at time t, that is, the displacement amount of the pixel from the previous frame to the current frame Among them, u t represents the optical flow component in the horizontal direction (x-axis direction), and v t represents the optical flow component in the vertical direction (y-axis direction).
[0115] In step B, in order to further map the global features of macro and micro expressions to the same feature space to extract information, multiple identical macro and micro knowledge sharing modules are designed to gradually abstract features, using the optical flow features E i (u, v) and E a (u, v) as inputs. Its specific structure is as follows. In each macro and micro knowledge sharing module, downsampling is performed through an initial convolutional layer, and then a residual connection is made with the second convolutional layer. The role of each convolutional layer is to implement feature extraction; to ensure the efficient propagation of information. The structure B of each macro and micro knowledge sharing module is expressed as formulas (5) to (7):
[0116] Z m.k =Β k (Z m,k-1 ; θ m,k ) = Conv2(BN(ReLU(Conv1(Z m,k-1 ; W 1,k ))); W2,k ) + Z m,k-1 (5);
[0117] Z iu = Β K (Β K-1 (...Β2(Β1(E i (u, v))))) (6);
[0118] Z au = Β K (Β K-1 (...Β2(Β1(E a (u, v))))) (7);
[0119] Among them, B k represents the k-th macro-micro knowledge sharing module, Z m,k is the output of the k-th macro-micro knowledge sharing module, Z iu and Z au respectively represent the outputs of the optical flow features of micro-expression and macro-expression passing through the macro-micro shared knowledge module. Conv1 and Conv2 respectively represent two convolutional layers, W 1,k and W 2,k are convolutional kernels, θ m,k is the parameter of the convolutional block, K is the number of macro-micro knowledge sharing modules. For the convenience of subsequent representation, the outputs Z iu and Z au are collectively referred to as Z s , that is, the information features of the output macro-expression domain and micro-expression domain. Next, adversarial training will be carried out through the gradient inversion module to eliminate the distribution difference between the source domain and the target domain, and the macro-micro shared knowledge module will learn domain-invariant feature representations.
[0120] In step C, after the feature extraction of the macro domain and the micro domain through the above-mentioned macro-micro shared knowledge module, there are still large differences in the features of the two domains, and the knowledge of the macro domain cannot be directly applied to the micro-expression domain. Therefore, it is necessary to reverse-train the macro-micro knowledge sharing module in an adversarial manner through the gradient reversal layer in the gradient inversion module to obtain the shared features of the two domains. The role of the gradient reversal layer is to reverse the gradient during the training process, so that the feature extraction network can reduce the domain distribution difference through an adversarial manner while optimizing the classification task, thereby learning domain-invariant features to improve the cross-domain transfer ability. The specific role is shown in formulas (8) to (10):
[0121]
[0122]
[0123]
[0124] Among them, GRL represents the Gradient Reversal Layer, λ is the reversal coefficient, which is dynamically updated during training, p is the current training progress; Z s is the output feature of the macro-micro shared module, which is input into the Gradient Reversal Layer to obtain the output I represents the identity matrix. After that, the output feature representation continues to be input into the domain classifier D(·) for classification. The role of the domain classifier D(·) is to classify the expressions of micro-expressions and macro-expressions into their respective domains, as shown in formula (11):
[0125]
[0126] Among them, θ d is the parameter of the domain classifier; is the output of the Gradient Reversal Layer is the classification vector output after being input into the domain classifier D(·);
[0127] In order to more clearly show the change of the propagated gradient, the mathematical representation in the process is given here. For the domain classifier D(·), as shown in formula (12):
[0128]
[0129] Among them, is the domain label probability predicted by the domain classifier for the current sample i to be measured, that is, the output of the domain classifier; d i is the domain label of sample i. Specifically, as shown in formula (13):
[0130]
[0131] For the macro-micro shared knowledge module, for the sake of simplicity here, f(·) is used to represent the entire macro-micro shared knowledge module. Due to the role of the gradient inversion module GRL, the gradient in the macro-micro shared knowledge module is expressed as formula (14):
[0132]
[0133] The gradient inversion module enables the gradient to propagate in the opposite direction at the macro-micro shared knowledge module, and finally narrows the distributions of the macro-expression domain and the micro-expression domain.
[0134] In step D, a local change extractor is constructed. Specifically, in the process of micro-expressions, tiny expression changes usually present three stages: rapid start, peak, and decay. Therefore, in order to accurately capture the local changes of the peak frame and the start frame, the present invention proposes to extract the local change features by calculating the inter-frame difference between the peak frame and the start frame to capture the subtle dynamics of the expression change. For each video sequence x in the micro-expression miAll are segmented into a series of video frames; the local change features of micro - expressions are represented by formula (15).
[0135]
[0136] Wherein, And respectively refer to the peak frame and the starting frame of the current micro - expression sequence, I represents the number of frames of the micro - expression sequence, refers to the output after the micro - expression sequence passes through the local change extractor, that is, the local change feature descriptor of the micro - expression.
[0137] The output micro - expression change will enter the micro - domain expression module for further modeling.
[0138] In step E, in order to extract more fine - grained and high - dimensional feature representations from the local changes of micro - expressions and further strengthen the spatio - temporal characteristic expression of micro - expressions, the present invention designs a multi - layer convolutional neural network structure, which is stacked by k micro - domain expression modules.
[0139] By performing multi - level feature extraction on the input local change data, from low - level local changes to high - level abstract features, various dynamic changes of micro - expressions are fully captured. The specific structure of each micro - domain expression module is shown in formulas (16) to (18):
[0140] D k =N k (D k-1 ; θ k ), k = 1, 2,..., K (16);
[0141] X l =σ(W l *X l-1 +b l ), l = 1, 2,..., L k (17);
[0142] X l.pool =Pool(X l ), l = 1, 2,..., L k (18);
[0143] Formula (16) indicates that the output D k of the k - th micro - domain expression module is obtained according to the input D k-1 of the (k - 1) - th micro - domain expression module, N k (·; θ k ) represents the k - th micro - domain expression module, and θ k is the parameter of the k - th micro - domain expression module; the k - th micro - domain expression module has L klayer, each layer includes a convolutional layer and an activation function, as shown in formula (17), σ represents the activation function, W l and b l represent the convolution kernel and bias of the first convolutional layer respectively, * represents the convolution operation, X l-1 and X l represent the input and output of the first convolutional layer respectively. After the convolution operation, each layer also performs a pooling operation, and the pooling operation is expressed as formula (18), Pool represents the pooling operation, X l.pool represents the output after the pooling operation of each micro-domain expression module after the convolution operation. Each layer of convolution operation not only extracts features spatially, but also downsamples the feature map through the pooling operation to reduce the computational amount and extract more robust features. Finally, the output D k of the k-th micro-domain expression module is obtained.
[0144] In step F, in order to fuse the dual-domain (macro-expression domain and micro-expression domain) features and single-domain (micro-expression domain) features to obtain a high-dimensional feature representation with rich spatio-temporal information and domain generalization ability. Bilinear Pooling Fusion (BPF) can not only retain the unique information of each domain, but also effectively model the interaction relationship between the macro-expression domain and the micro-expression domain, further improving the accuracy and robustness of cross-domain recognition.
[0145] Traditional feature fusion methods often cannot fully capture the complex interaction relationships between different features, especially in cross-domain tasks where the relationships between features are more complex. Therefore, the present invention introduces a bilinear pooling method, which extracts richer and more accurate cross-domain and single-domain interaction features by performing a product operation on different features; specifically expressed as formula (19):
[0146]
[0147] where Z s represents the dual-domain features passing through the macro-micro knowledge sharing module, D k represents the single-domain features of the output of the k-th micro-domain expression module, represents the outer product operation, generating a matrix of d s ×d k . The core of this operation is to calculate the interaction relationship of each dimension between the dual-domain shared features and the single-domain features. Through this interaction, the similarities and differences between different domains can be captured, thus generating richer feature expressions. Bilinear pooling can generate a high-order feature representation, but this direct outer product operation will cause a sharp increase in the feature dimension. To overcome this problem, a dimensionality reduction operation is introduced to map the generated high-dimensional features to a manageable dimensional space; specifically, through a linear transformation matrix to perform feature dimensionality reduction; the feature F after dimensionality reduction fuse is expressed as formula (20):
[0148] F fuse = W·vec(F bilinear ) (20);
[0149] where, vec(·) means expanding the matrix F bilinear by columns into a vector, W is a linear transformation matrix, and d fuse is the feature dimension after dimensionality reduction; through this transformation, a fused feature with a smaller dimension will be generated which can effectively reduce the complexity while retaining important interaction information.
[0150] To further enhance the expressive ability of the bilinear fused feature, the feature F after linear dimensionality reduction fuse is subjected to a non - linear transformation; the LeakyReLU activation function is used to introduce non - linear factors to ensure that the model can fit complex non - linear relationships. The specific transformation formula is shown in formula (21):
[0151] F nonlinear = σ(F fuse + b) (21);
[0152] where, σ(·) is the LeakyReLU activation function and b is the bias term; F nonlinear represents the output of the feature F after dimensionality reduction fuse after passing through the activation function; this non - linear transformation makes the bilinear fused feature more adaptable to complex patterns and relationships, effectively improving the fitting ability of the model.
[0153] Next, F nonlinear is input into the classifier for classification, which is specifically expressed as formula (22):
[0154] Y cls = W cls ·F nonlinear + b cls (22);
[0155] where, is the weight matrix of the classifier, is the bias term, C is the number of classes, and after linear transformation, the obtained output represents the scores of each class of micro - expression categories "positive", "negative", "surprise";
[0156] To obtain the classification probability distribution, the result after linear transformation is then activated by Softmax. The Softmax function converts the class scores into a probability distribution, as shown in Equation (23):
[0157]
[0158] where is a vector of length C, representing the predicted probabilities for each of the micro-expression categories "positive", "negative", and "surprised".
[0159] The cross-domain micro-expression recognition model is optimized and iteratively updated using the single-domain sentiment classification recognition loss and the dual-domain sentiment domain adaptation loss; including:[[]]
[0160] The output of the sentiment classifier in the single-domain representation network is p i ∈P, representing the predicted probabilities that the i-th micro-expression to be tested belongs to each sentiment category; the goal of sentiment classification is to minimize the cross-entropy loss between the true label y i and the predicted label p i The mathematical expression of the cross-entropy loss is as shown in Equation (24):
[0161]
[0162] where N is the number of samples, K is the number of sentiment categories, is the true label of sample i in sentiment category k, is the probability that the model predicts sample i belongs to sentiment category k; this loss function aims to minimize the difference between the true label and the predicted label, thereby improving the accuracy of sentiment classification. To achieve cross-domain transfer learning and eliminate the distribution difference between the macro-expression domain and the micro-expression domain, the domain adaptation loss is introduced. In this algorithm, a gradient reversal module is used, which achieves adversarial training by reversing the gradient during backpropagation, enabling the macro-micro knowledge sharing module to learn domain-invariant features, thereby achieving the alignment of the feature distributions of the macro-expression domain and the micro-expression domain.[[]]
[0163] In the gradient reversal module, after the input feature Z undergoes the reversal operation, it is input into the domain classifier D v The goal of the domain classifier is to determine whether the current sample belongs to the macro-expression domain or the micro-expression domain; the output of the domain classifier is representing the probability that the current input sample belongs to the micro-expression domain, and the loss function is as shown in Equation (25):
[0164]
[0165] where:[[]] is the domain label probability predicted by the domain classifier for the current sample i to be tested, d iIt is the domain label of sample i.
[0166] The goal of the dual-domain semantic adversarial learning network is to minimize this loss, prompting the feature distributions of the macro-expression domain and the micro-expression domain to be as close as possible. The task of the macro-micro knowledge sharing module is to make the domain classifier unable to distinguish between the macro-expression domain and the micro-expression domain, and thereby learn domain-invariant features for final classification through this adversarial process.
[0167] In this embodiment, the CK+ dataset released by the Kanade, Cohn, and Tian teams at Carnegie Mellon University is used as the macro-expression dataset, and micro-expression recognition tests are respectively conducted on the original videos of the CASMEⅡ micro-expression database released by the Fu Xiaolan team at the Institute of Psychology, Chinese Academy of Sciences, the SAMM database released by the Davison team at Manchester Metropolitan University, UK, and the SMIC database released by the research team at the University of Oulu, Finland. Table 1 shows the test results.
[0168] Table 1
[0169]
[0170] In addition, the present invention also draws the confusion matrices of the three datasets of CASMEII, SAMM, and SMIC to prove the superiority of the algorithm of the present invention. As Figure 3 , Figure 4 and Figure 5 shown, it can be seen from these three confusion matrices that the algorithm has achieved high classification accuracy on different datasets (CASME II, SAMM, SMIC), especially showing obvious advantages in key categories (such as Positive in CASME II and Negative in SAMM and SMIC). Taking CASME II as an example, the recognition accuracies of Surprise, Positive, and Negative are all around 90% or above, indicating that the model has good generalization ability in capturing the subtle differences of micro-expressions. Even on more challenging datasets such as SAMM and SMIC, the model can still maintain a high correct classification ratio in most categories, and the misclassifications are mainly concentrated between adjacent emotion categories, indicating that this method can effectively distinguish categories with large differences. Generally speaking, these results prove that the present invention can achieve relatively robust performance under the conditions of cross-different datasets and multi-emotion categories, reflecting its good modeling ability for the spatio-temporal features of micro-expressions and cross-domain adaptability.
[0171] In addition, a fine classification (five-classification) experiment for CASMEII is also conducted in this paper, as shown in Table 2:
[0172] Table 2
[0173]
[0174] It can be found from the data that even with the increased classification difficulty, the accuracy of the fine classification of CASMEII has reached 81.9%.
[0175] In addition, the present invention also conducted ablation experiments on key modules (gradient inversion module, macro-domain fusion), as shown in Table 3 and Table 4:
[0176] Table 3
[0177]
[0178] It was found through experiments that when these modules were removed, the experimental results decreased significantly (3% - 7%), proving the effectiveness of each structure of the present invention.
Claims
1. A cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization, characterized in that: include: Build and train a cross-domain micro-expression recognition model; The cross-domain micro-expression recognition model includes a dual-domain semantic adversarial learning network and a single-domain representation network; the dual-domain semantic adversarial learning network includes a global dynamic extractor, a macro-micro knowledge sharing module, and a gradient inversion module; the single-domain representation network includes a local change extractor, a micro-domain expression module, a bilinear pooling fusion unit, and a classifier; The data to be recognized is pre-processed and then fed into the trained cross-domain micro-expression recognition model to perform cross-domain micro-expression recognition; including: A. Input the starting frame and peak frame of the macro-expression and micro-expression video frames into the global dynamic extractor to extract the dynamic optical flow features of each macro-expression and micro-expression video frame; B. Input the dynamic optical flow features obtained from step A into the macro-micro knowledge sharing module to obtain macro-micro knowledge features; C. The macro-micro knowledge features obtained from step B are input into the gradient inversion module to further purify the shared features of the two domains by reversing the gradient; at the same time, the shared features are input into the domain discriminator for domain discrimination; D. Input the starting frame and peak frame of the micro-expression into the local change extractor to obtain the local change feature descriptor of the micro-expression; E. Input the local change feature descriptor of the micro-expression obtained in step D into the micro-domain expression module to further capture the local dynamic changes of the micro-expression and obtain the single-domain feature; F. The shared features of the two domains obtained in step C and the single domain features obtained in step E are fused through a bilinear pooling fusion unit to obtain a high-dimensional feature representation with rich spatiotemporal information and domain generalization capability; then, the obtained high-dimensional feature representation is input into the classifier of the single domain representation network for the final micro-expression recognition and classification.
2. According to claim 1, a cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization is characterized in that: In step A, for each video sequence x in micro-expressions and macro-expressions mi / ma They are divided into a series of video frames, expressed as formula (1): x mi / ma =[f mi / ma,1 ,f mi / ma,2 ,f mi / ma,3 ,...,f mi / ma,ap ,...,f mi / ma,n ],E(f mi / ma,ap )=E max (1); Where n is the total number of micro-expression and macro-expression video frames, f mi / ma,ap is the peak frame of the micro-expression and macro-expression video sequence, E is the expression intensity of the expression frame sequence; for the input micro-expression start frame f mi,1 and f mi,ap , TV-L1 optical flow obtains the micro-expression optical flow E by solving the following variational optimization problem i (u,v) and macro expression optical flow E a (u,v), as shown in formula (2) and formula (3): Among them, x, y represent pixel points, u, v represent the displacement components of the optical flow field in the x and y directions respectively, and λ is the regularization parameter of the smoothness term, which is used to balance the data term and the regularization term. At the same time, the L1 constraint |f(x+u,y+v)-f(x,y)| of the data term makes the optical flow calculation more robust to the illumination calculation, and the regularization term Control the smoothness of optical flow; The optimization problem is solved by the Lagrange multiplier method and the Gauss-Seidel iteration is used as shown in formula (4): Finally, the dynamic optical flow feature, u t and v t represents the optical flow component at the tth iteration, where u t Represents the horizontal optical flow component, v t Represents the optical flow component in the vertical direction; the optical flow field finally obtains the motion information of the pixel at time t, that is, the displacement of the pixel from the previous frame to the current frame Among them, u t Represents the horizontal optical flow component, v t Represents the optical flow component in the vertical direction.
3. According to claim 1, a cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization is characterized in that: In step B, in each macro-micro knowledge sharing module, downsampling is performed through an initial convolutional layer, and then residual connection is performed with the second convolutional layer. The function of each convolutional layer is to realize feature extraction; the structure B of each macro-micro knowledge sharing module is expressed as formula (5) to formula (7): WITH m.k =Β k (WITH m,k-1 ;θ m,k )=Conv2(BN(ReLU(Conv1(Z m,k-1 ;IN 1,k )));IN 2,k )+Z m,k-1 (5); Z iu =B K (B K-1 (...B2(B1(E i (u,v))))) (6); Z au =B K (B K-1 (...B2(B1(E a (u,v))))) (7); Among them, B k represents the kth macro-micro knowledge sharing module, Z m,k is the output of the kth macro-micro knowledge sharing module, Z iu and Z au The optical flow features representing micro-expressions and macro-expressions are output through the macro-micro shared knowledge module. Conv1 and Conv2 represent two convolutional layers, respectively. 1,k and W 2,k is the convolution kernel, θ m,k is the parameter of the convolutional block, K is the number of macro-micro knowledge sharing modules, and the output Z iu and Z au Collectively referred to as Z s , that is, the information features of the output macro-expression domain and micro-expression domain.
4. The cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization according to claim 1 is characterized in that: In step C, the macro-micro knowledge sharing module is reversely trained in an adversarial manner through the gradient reversal layer in the gradient reversal module to obtain the shared features of the two domains. The function of the gradient reversal layer is to reverse the gradient during the training process, so that the feature extraction network can reduce the distribution difference between domains in an adversarial manner while optimizing the classification task, thereby learning the domain invariant features. The specific functions are shown in formulas (8) to (10): Among them, GRL represents the gradient reversal layer, λ is the reversal coefficient, which is dynamically updated during training, and p is the current training progress; Z s The output feature of the macro-micro shared module is input into the gradient reversal layer to obtain the output I represents the unit matrix. After that, the output feature expression is further input into the domain classifier D(·) for classification. The function of the domain classifier D(·) is to classify the expressions of micro-expressions and macro-expressions into their respective domains, as shown in formula (11): Among them, θ d are the parameters of the domain classifier; is the output of the gradient reversal layer The classification vector output after being input into the domain classifier D(·); For the domain classifier D(·), as shown in formula (12): in, is the domain label probability predicted by the domain classifier for the current sample i, that is, the output of the domain classifier; d i is the domain label of sample i. Specifically, as shown in formula (13): The gradient in the macro-micro shared knowledge module is expressed as formula (14):
5. The cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization according to claim 1 is characterized in that: In step D, for each video sequence x in micro-expression mi are divided into a series of video frames; the local change feature of micro-expression is expressed as formula (15) in, and They refer to the peak frame and starting frame of the current micro-expression sequence respectively, I represents the number of frames in the micro-expression sequence, Refers to the output of the micro-expression sequence after passing through the local change extractor, that is, the local change feature descriptor of the micro-expression.
6. The cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization according to claim 1 is characterized in that: In step E, multi-level feature extraction is performed on the input local change data, from low-level local changes to high-level abstract features. The specific structure of each micro-domain expression module is shown in formulas (16) to (18): Formula (16) represents the output D of the kth micro-domain expression module: k is based on the input D of the k-1th micro-domain expression module k-1 Obtained, N k (·;θ k ) represents the kth micro-domain expression module, θ k is the parameter of the kth micro-domain expression module; the kth micro-domain expression module has L k layers, each layer includes a convolutional layer and an activation function, as shown in formula (17), σ represents the activation function, W l and b l They represent the convolution kernel and bias of the first convolutional layer, * represents the convolution operation, and X l-1 and X l They represent the input and output of the first convolutional layer respectively. After the convolution operation, each layer also performs a pooling operation. The pooling operation is expressed as formula (18). Pool represents the pooling operation, and X l.pool It represents the output of each micro-domain expression module after the convolution operation and then the pooling operation. Each layer of convolution operation not only extracts spatial features, but also downsamples the feature map through the pooling operation. Finally, the output D of the kth micro-domain expression module is obtained. k .
7. The cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization according to claim 1 is characterized in that: In step F, a bilinear pooling method is introduced to extract richer and more accurate cross-domain and single-domain interaction features by multiplying different features. The specific expression is formula (19): Among them, Z s represents the dual-domain features after the macro-micro knowledge sharing module, D k represents the single domain feature of the output of the k-th micro-domain expression module, Represents the outer product operation, generating a d s ×d k , introduces a dimensionality reduction operation to map the generated high-dimensional features into a manageable dimensional space; specifically, through a linear transformation matrix To perform feature dimensionality reduction; the feature F after dimensionality reduction fuse It is expressed as formula (20): F fuse =W·vec(F bilinear ) (20); Among them, vec(·) represents the matrix F bilinear Expand it into a vector by column, W is the linear transformation matrix, d fuse is the feature dimension after dimensionality reduction; After linear dimensionality reduction, the feature F fuse After nonlinear transformation, the LeakyReLU activation function is used to introduce nonlinear factors. The specific transformation formula is shown in formula (21): F nonlinear =σ(F fuse +b) (21); Where σ(·) is the LeakyReLU activation function, b is the bias term; F nonlinear Represents the feature F after dimensionality reduction fuse Output after activation function; Next, F nonlinear Input into the classifier for classification, which is specifically expressed as formula (22): Y cls =W cls ·F nonlinear +b cls (22); in, is the weight matrix of the classifier, is the bias term, C is the number of categories, and after linear transformation, the output Indicates the scores of each category of micro-expression categories "positive", "negative", and "surprise"; The result after linear transformation is activated by Softmax. The Softmax function converts the category score into a probability distribution, as shown in formula (23): in, is a vector of length C, representing the predicted probability of each category of micro-expression categories "positive", "negative", and "surprised".
8. A cross-domain micro-expression recognition method based on gradient inversion domain adaptation and dual-stream optimization according to any one of claims 1-7, characterized in that: The cross-domain micro-expression recognition model is optimized and iteratively updated using the single-domain emotion classification recognition loss and the dual-domain emotion domain adaptation loss; including: The output of the sentiment classifier in the single domain representation network is p i ∈P, which indicates the predicted probability that the i-th micro-expression currently under test belongs to each emotion category; the goal of emotion classification is to minimize the true label y i With the predicted label p i The cross entropy loss between ; the mathematical expression of the cross entropy loss is shown in formula (24): Among them, N is the number of samples, K is the number of emotion categories, is the true label of sample i in sentiment category k, The probability that the sample i predicted by the model belongs to the sentiment category k; In the gradient inversion module, the input feature Z is inverted and then input to the domain classifier D v In the above example, the goal of the domain classifier is to determine whether the current sample belongs to the macro expression domain or the micro expression domain; the output of the domain classifier is represents the probability that the current input sample belongs to the micro-expression domain. The loss function is shown in formula (25): in: is the domain label probability predicted by the domain classifier for the current test sample i, d i is the domain label of sample i.
Citation Information
Cited By
Facial micro-expression recognition method
CN120783379A