An Automatic Micro-Expression Classification Model Training Method Based on Neural Processes

Through the automatic micro-expression classification model training method based on neural processes, the problems of one-sidedness and time uncertainty of time information processing in the prior art are solved, and more accurate and efficient micro-expression recognition is achieved.

CN115019366BActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210587464.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-05-27
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

Existing micro-expression recognition technology is more one-sided when processing time information, and fails to effectively consider time uncertainty, resulting in the inability to ensure the time uncertainty of the context in modeling.

Method used

The automatic micro-expression classification model training method based on neural processes is adopted. By extracting micro-expression image nodes from micro-expression video frames and dividing them into context nodes and target nodes, the latent feature information and expression prediction tags are calculated using an encoder and decoder, and the regularized loss function that maximizes the lower limit of evidence and output distribution is optimized.

Benefits of technology

This method can effectively learn and process the distributed information on time, estimate time uncertainty, and show high efficiency in computing efficiency, improving the accuracy of micro-expression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019366B_ABST
    Figure CN115019366B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for training an automatic micro-expression classification model based on neural processes, comprising the following steps: extracting a number of micro-expression images as micro-expression image nodes, and dividing them into context nodes and target nodes; performing feature extraction on the micro-expression image nodes to obtain micro-expression feature information; inputting the micro-expression feature information and true labels of the context nodes into an encoder to calculate latent feature information, and aggregating the latent feature information to obtain a global latent variable; using the global latent variable to calculate the expression prediction labels of the target nodes in a decoder; setting a joint optimization objective and minimizing the joint optimization objective. The method of the present invention uses neural processes. Neural processes combine neural networks and stochastic processes, can learn the distribution over functions, and can also estimate the uncertainty of its predictions based on contextual observations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of facial expression recognition, and particularly relates to a method for training an automatic micro-expression classification model based on neural processes. Background Art

[0002] In the research of micro-expression recognition, based on traditional image recognition methods such as LBP-TOP and directly using mature neural networks, methods regarding temporal features and expression enhancement are proposed for the characteristics of small amplitude and short duration of micro-expressions.

[0003] Temporal context information is one of the keys to recognizing emotional expressions. Existing micro-expression recognition technologies rely more on optical flow characteristics in terms of time information. Some researchers have disclosed technologies based on the optical flow characteristics of the OFF-ApexNet (Optical Flow based on Framelets for ApexNet). This is a new feature descriptor that combines context information guided by optical flow with intermediate information of a convolutional neural network. However, this technology only compares the optical flow information between the expressionless frame and the apex frame, which is relatively one-sided.

[0004] There are also related technologies that perform deep learning operations on the time axis using 3D convolutional kernels for the temporal domain features and optical flow features of micro-expression videos, or extract optical flow information and use the optical flow information as the object of deep learning. Although the problem of relatively one-sided information mentioned above is solved, the uncertainty of time is not considered.

[0005] Existing models relying on the recurrent self-attention mechanism can strengthen temporal consistency. However, if the method in terms of the dimension of task feature information ignores the temporal dependence of task features, it is impossible to guarantee the temporal uncertainty of the context in modeling. Summary of the Invention

[0006] Based on the above-mentioned drawbacks and deficiencies existing in the prior art, one of the purposes of the present invention is to at least solve one or more of the above problems existing in the prior art. In other words, one of the purposes of the present invention is to provide a method for training an automatic micro-expression classification model based on neural processes that meets one or more of the foregoing requirements.

[0007] In order to achieve the above-mentioned invention purpose, the present invention adopts the following technical solutions:

[0008] A method for training an automatic micro-expression classification model based on neural processes specifically includes the following steps:

[0009] S1. Extract a plurality of micro-expression images from the micro-expression video frames of the training dataset, respectively as micro-expression image nodes, and divide the plurality of micro-expression image nodes into a plurality of context nodes and a plurality of target nodes;

[0010] S2. Extract features from the micro-expression image nodes to obtain a plurality of micro-expression feature information;

[0011] S3. Input the micro-expression feature information and true labels of several context nodes into the encoder, calculate the latent feature information of the micro-expression feature information of each context node respectively, and aggregate all the latent feature information to obtain a global latent variable;

[0012] S4. Use the global latent variable to calculate the expression prediction label of the micro-expression feature information of each target node in the decoder;

[0013] S5. Set the maximization of the evidence lower bound and the regularization loss function of the output distribution as the joint optimization objective, and minimize the joint optimization objective.

[0014] As a preferred solution, step S3 specifically includes:

[0015] S31. Use a multi-layer perceptron as the encoder, and calculate the corresponding latent feature information according to the micro-expression feature information and true labels of each context node;

[0016] S32. Sum and average several latent feature information to obtain global latent feature information;

[0017] S33. Use a latent encoder to parameterize the global latent feature information into a Gaussian distribution, take the Gaussian distribution as the global latent variable, and any value of the global latent variable distribution is a latent representation.

[0018] As a further preferred solution, the calculation method of the feature representation vector in step S31 is specifically:

[0019] Convert the true label to 512 dimensions through a linear layer and a ReLU activation function, combine it with the micro-expression feature information of the context node, and obtain the feature representation vector after passing through two linear layers and two activation functions.

[0020] As a further preferred solution, the parameterization process of the global latent feature information in step S33 is specifically:

[0021] Pass the global latent feature information through a linear layer and an activation function, and then enter two linear layers to obtain the mean and standard deviation of the Gaussian distribution respectively.

[0022] As a further preferred solution, step S4 specifically includes:

[0023] S41. Randomly take values from the global latent variable distribution to obtain a latent representation, and combine the micro-expression feature information of each target node with a latent representation to obtain combined information;

[0024] S42. Use a multi-layer perceptron as the decoder, and calculate the mean and standard deviation corresponding to the expression prediction label of each target node according to the combined information;

[0025] S43. Use the mean value as the expression prediction label of the target node.

[0026] As a further preferred solution, the calculation process of the mean value and standard deviation corresponding to the expression prediction label of each target node in step S42 is specifically as follows:

[0027] Pass the combined information through three linear layers and activation functions, enter two linear layers, and obtain the mean value and standard deviation corresponding to the expression prediction label of each target node respectively.

[0028] As a further preferred solution, the functional form of maximizing the evidence lower bound is:

[0029] L 1 = -E qΦ(z|C) [logp(y T |x T ,z)];

[0030] Among them, p(y T |x T ,z) is the probability of the value of the expression prediction label y T when the micro-expression feature information on the total set T of target nodes is x T and the latent variable sampling is z, and q Φ (z|C) is the probability of obtaining the latent representation sampling as z when the context node set is C;

[0031] The regularization loss function of the output distribution is:

[0032]

[0033] Among them, D KL (·||·) refers to the KL divergence of two distributions, refers to the average normal distribution of the predicted label y t of the target node when the number of samples of the target node is n t , and N(y,1) refers to the normal distribution form of the true label of the target node.

[0034] As a preferred solution, before step S2, there is also step S20:

[0035] S20. Perform preprocessing operations on the micro-expression image, and the preprocessing operations include face extraction, face alignment, and image scaling processing.

[0036] Compared with the prior art, the beneficial effects of the present invention are:

[0037] The method of the present invention uses a neural process, which combines a neural network and a stochastic process, can learn the distribution on a function, and can also estimate the temporal uncertainty of its prediction according to the observations of the context. In addition, the neural process can also generate predictions in a computationally efficient manner. Description of the Drawings

[0038] Figure 1 is a flowchart of an automatic micro-expression classification model based on a neural process according to an embodiment of the present invention;

[0039] Figure 2 is a structural diagram of an encoder according to an embodiment of the present invention;

[0040] Figure 3 is a structural diagram of a latent encoder according to an embodiment of the present invention;

[0041] Figure 4 is a structural diagram of a decoder according to an embodiment of the present invention. Detailed Embodiments

[0042] To more clearly illustrate the embodiments of the present invention, the specific embodiments of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings, and other embodiments can be obtained.

[0043] Embodiment: This embodiment specifically provides an implementation manner of a training method for an automatic micro-expression classification model based on a neural process:

[0044] The flowchart of the automatic micro-expression classification model based on a neural process in this embodiment is as Figure 1 shown. First, in order to implement the training of the model, a training data set of micro-expression videos needs to be prepared. Assume that a series of micro-expression image frames are currently provided Use I g to represent an image in the total set G, and want to predict the type of its facial micro-expression. Based on each frame, y g ∈R 3 , that is, the type of the micro-expression label y g may be three categories, namely positive, negative, and surprised. Using this data set, perform step S1. Extract several micro-expression images from the micro-expression video frames of the training data set. Each frame image is used as a micro-expression image node, and all the micro-expression image nodes are divided into two groups: context nodes and target nodes;

[0045] Then, step S2 is performed to extract features of the micro-expression image nodes through a pre-trained backbone convolutional neural network, i.e., a feature extractor Φ, to obtain micro-expression feature information of each micro-expression image node. The feature extractor Φ is provided with independent images and true labels of each micro-expression image node for training in a standard supervised manner. The purpose of training the feature extractor Φ is to simplify the input information of the model.

[0046] First, the feature extractor Φ is used to calculate x g =Φ(I g ), x g The micro-expression feature information x will be used as the micro-expression feature information of each micro-expression image node. Then, the micro-expression feature information x g Instead of taking the image sequence directly as input, the sequence of images is used as the input of the neural process in the subsequent method.

[0047] During subsequent model training, the feature extractor will be frozen and, in some cases, may be fine-tuned again.

[0048] The micro-expression feature information extracted according to step S2 is divided into two parts. The context node is extracted as {x c ,y c} c∈C , where x c is the micro-expression feature information, y c is the true label; the target node is extracted as {x t} t∈T , where x t It is the micro-expression feature information.

[0049] After the micro-expression feature information is extracted, step S3 is performed to input the micro-expression feature information and the true label of the context node into the encoder, and the latent feature information of the micro-expression feature information of each context node is calculated respectively, and all the latent feature information is aggregated to obtain a global latent variable.

[0050] Specifically, S3 includes the following steps:

[0051] S31, use a three-layer multilayer perceptron as the encoder q Φ , extract {x from the context node c ,y c} c∈C Input encoder q Φ First, the true label y of the input c Upsample to match the micro-expression feature information x c The encoder structure is shown in Figure 2 Shown: Feature information x c The dimension of M is 512, so the true label y of dimension Mc It is also transformed into 512 dimensions through a linear layer and a ReLU activation function. The combination of the two results in 1024-dimensional information, and the 1024-dimensional information passes through two linear layers and activation functions to obtain the latent feature information r. c 。

[0052] S32. Perform an aggregation operation between the encoder and the latent encoder to average the latent feature information r obtained from numerous context nodes, and calculate a global latent feature information r. c Take the average value to calculate a global latent feature information r.

[0053] S33. Input the global latent feature information r into the latent encoder. The structure diagram of the latent encoder is as shown. After passing through a linear layer and an activation function in the latent encoder, it enters two linear layers respectively, so as to calculate the mean μ(r) and standard deviation σ(r) of the Gaussian distribution corresponding to the global latent feature information r in the two linear layers respectively. Among them, the standard deviation calculation will perform a sigmoid function calculation one more time to control its value range. Figure 3 The Gaussian distribution in step S3 above is used as the global latent variable. A latent representation z can be obtained by sampling from the distribution of the global latent variable, z ∼ N(μ, σ). Use the above global latent variable to perform step S4, and use the global latent variable to calculate the expression prediction label of the micro-expression feature information of each target node in the decoder;

[0054] Specifically, step S4 is implemented in the following process:

[0055] Specifically, step S4 is implemented as follows:

[0056] S41. Randomly sample from the distribution of the 128-dimensional global latent variable to obtain a 128-dimensional latent representation z, and then combine the micro-expression feature information x of each target node with a latent representation z to obtain combined information. The micro-expression feature information x is 512 dimensions, the latent representation z is 128 dimensions, and the combined information obtained by combining the two is 640 dimensions. t Combine the micro-expression feature information x of each target node with a latent representation z to obtain combined information. t The micro-expression feature information x is 512 dimensions, the latent representation z is 128 dimensions, and the combined information obtained by combining the two is 640 dimensions.

[0057] S42. Use a multi-layer perceptron as the decoder. The structure diagram of the decoder is as shown. After passing the combined information through three linear layers and activation functions, it enters two linear layers respectively, so as to obtain the mean Figure 4 corresponding to the prediction label and the standard deviation in the two linear layers respectively. Among them, the mean is directly obtained by the linear layer, while the variance will pass through a sigmoid function to control its range. Among them, the mean is directly obtained by the linear layer, while the variance will pass through a sigmoid function to control its range.

[0058] In the test stage, the mean is directly used as the expression prediction label of the target node.

[0059] To improve the recognition accuracy, it is necessary to optimize the above-mentioned neural process model. Perform step S5, set the optimization objective as the joint optimization objective of maximizing the evidence lower bound and the regularization term loss function of the output distribution, and achieve the final minimization of the optimization objective.

[0060] Specifically, the evidence in maximizing the evidence lower bound refers to the probability density of data or observable variables. The functional form of maximizing the evidence lower bound is:

[0061] L 1 =-E qΦ(z|C) [logp(y T |x T ,z)];

[0062] Among them, p(y T |x T ,z) is the probability of the expression prediction label y T taking a value under the condition that the micro-expression feature information on the total set T of target nodes is x T and the latent variable sampling is z, and q Φ (z|C) is the probability of obtaining the latent representation sampling as z under the condition that the context node set is C.

[0063] The optimization objective of the distributed output design is the regularization term loss function of the output distribution, and this function is specifically:

[0064]

[0065] Among them, D KL (·||·) refers to the KL divergence of two distributions, refers to the average normal distribution of the prediction label y t of the target nodes when the number of samples of the target nodes is n t , and N(y,1) refers to the normal distribution form of the true label of the target nodes.

[0066] Combine the above two optimization objectives and use λ as a free parameter to control the balance of the loss function. The final minimization optimization objective is:

[0067] L = L 1 + λL 2 .

[0068] The micro-expression recognition model trained by the above method is used as follows:

[0069] Input the micro-expression image to be recognized into the feature extractor Φ to extract features. The extracted features are used as the input of the decoder, and at the same time, randomly sample from the latent feature distribution obtained during training as the latent representation input to the decoder, and obtain the micro-expression classification result through the output of the decoder.

[0070] After implementing according to the above embodiments, compare this embodiment with the implementations of other methods. The comparison results are shown in the following table. The references for other methods in the table are as follows:

[0071] [1] Xiaohua H, Wang SJ, Liu X, et al. Discriminative spatiotemporal local binary pattern with revisited integral projection for spontaneous facial micro-expression recognition[C]. IEEE Trans. Affect. Comput., 2019, 10(1): 32-47.

[0072] [2] Liong S T, See J, Wong K S, et al. Less is more: Micro-expression recognition from video using apex frame[J]. Signal Processing: Image Communication, 2018, 62: 82-92.

[0073] [3] Gan Y S, Liong S T, Yau W C, et al. OFF-ApexNet on micro-expression recognition system[J]. Signal Processing: Image Communication, 2019, 74: 129-139.

[0074] [4] Quang NV, Chun J, Tokuyama T. Capsulenet for micro-expression recognition[C]. 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition, 2019: 1-7.

[0075] [5] Zhou L, Mao Q, Xue L. Dual-inception network for cross-database micro-expression recognition[C]. 2019 14th IEEE International Conference on Automatic Face&Gesture Recognition, 2019: 1-5.

[0076] [6] Liong ST, Gan YS, See J, et al. Shallow triple stream three-dimensional cnn (ststnet) for micro-expression recognition[C]. 2019 14th IEEE International Conference on Automatic Face&Gesture Recognition, 2019: 1-5.

[0077] [7] Liu Y, Du H, Zheng L, et al. A neural micro-expression recognizer[C]. 2019 14th IEEE international conference on automatic face&gesture recognition. IEEE, 2019: 1-4.

[0078] The experimental results are listed in the following table:

[0079]

[0080] As can be seen from the table, overall, the method of the present invention, NP-ResNet50 and NP-VGG16, can reach the same level as the current excellent methods compared with other methods. Among them, on CASMEⅡ, the method model with ResNet-50 as the backbone network can reach a UF1 of 0.8711 and a UAR of 0.8835. Compared with other methods, the UAR has exceeded the results of other methods, while the UF1 has not exceeded other methods, but is very close to the best method. On SAMM, the situation of the method model with ResNet-50 as the backbone network is similar, and the UAR reaches 0.7342, which has exceeded other methods. In short, the method of the present invention has good performance in UAR and UF1 on a single database and can reach or exceed other methods. However, in the case of cross-database, that is, the results of the combined database, it does not exceed the best result of other methods.

[0081] Comparing the method models with different backbone networks, the experimental results of using ResNet-50 as the backbone network and the experimental results of using VGG-16 as the backbone network both perform well on the CASMEⅡ dataset and the SAMM dataset. Overall, the experimental results of using ResNet-50 as the backbone network are better than those of using VGG-16 as the backbone network. Especially in the case of cross-database, the UAR of using ResNet-50 as the backbone network is about 3% higher than that of using VGG-16 as the backbone network.

[0082] It should be noted that the above embodiments only elaborate on the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.

Claims

1. A training method for an automatic micro-expression classification model based on neural processes, characterized in that, specifically includes the following steps: S1. Extract a number of micro-expression images from the micro-expression video frames of the training data set, and use them as micro-expression image nodes respectively. Then divide the number of micro-expression image nodes into a number of context nodes and a number of target nodes; S2. Extract features from the micro-expression image nodes to obtain a number of micro-expression feature information; S3. Input the micro-expression feature information and true labels of the number of context nodes into the encoder, calculate the latent feature information of the micro-expression feature information of each context node respectively, and aggregate all the latent feature information to obtain a global latent variable; S4. Use the global latent variable to calculate the expression prediction label of the micro-expression feature information of each target node in the decoder; S5. Set the maximized evidence lower bound and the regularization loss function of the output distribution as the joint optimization objective, and minimize the joint optimization objective; The functional form of the maximized evidence lower bound is: where p(y T |x T , z) is the probability of the expression prediction label y T taking a value when the feature micro-expression feature information on the total set T of the target nodes is x T and the latent variable sampling is z, and q Φ (z|C) is the probability of obtaining the latent representation sampling as z when the context node set is C; The regularization loss function of the output distribution is: Among them, D KL (·||·) represents the KL divergence of two distributions, indicating that the number of samples taken by the target node is n t and the predicted label y of the target node t is the average normal distribution. N(y, 1) represents the normal distribution form of the true label of the target node.

2. The training method for an automatic micro-expression classification model based on neural processes according to claim 1, characterized in that, the step S3 specifically includes the following steps: S31. Use a multi-layer perceptron as the encoder, and calculate the corresponding latent feature information according to the micro-expression feature information and true label of each context node; S32. Sum and average the number of latent feature information to obtain global latent feature information; S33. Use a latent encoder to parameterize the global latent feature information into a Gaussian distribution, and use the Gaussian distribution as the global latent variable. Any value of the global latent variable distribution is a latent representation.

3. The training method for an automatic micro-expression classification model based on neural processes according to claim 2, characterized in that, the calculation method of the feature representation vector in the step S31 is specifically: Convert the true label into 512 dimensions through a linear layer and a ReLU activation function, combine it with the micro-expression feature information of the context node, and obtain the feature representation vector after passing through two linear layers and two activation functions.

4. The training method for an automatic micro-expression classification model based on neural processes according to claim 2, characterized in that, the parameterization process of the global latent feature information in the step S33 is specifically: After passing the global latent feature information through a linear layer and an activation function, enter two linear layers to obtain the mean and standard deviation of the Gaussian distribution respectively.

5. The training method for an automatic micro-expression classification model based on neural processes according to claim 1, characterized in that , the step S4 specifically includes the following steps: S41. Randomly take values from the global latent variable distribution to obtain a latent representation, and combine the micro-expression feature information of each target node with a latent representation to obtain combined information; S42. Use a multi-layer perceptron as the decoder, and calculate the mean and standard deviation corresponding to the expression prediction label of each target node according to the combined information; S43. Use the mean value as the expression prediction label of the target node.

6. A method for training an automatic micro-expression classification model based on neural processes according to claim 5, wherein, the calculation process of the mean value and standard deviation corresponding to the expression prediction label of each target node in step S42 is specifically as follows: Pass the combined information through three linear layers and an activation function, and enter two linear layers to respectively obtain the mean value and standard deviation corresponding to the expression prediction label of each target node.

7. A method for training an automatic micro-expression classification model based on neural processes according to claim 1, wherein, before step S2, there is also step S20: S20. Perform preprocessing operations on the micro-expression image, and the preprocessing operations include face extraction, face alignment, and image scaling processing.

Citation Information

Patent Citations

  • Micro-expression type discrimination method based on transfer learning and auto-encoder data enhancement

    CN111767842A

  • Discriminative feature learning method and system for micro-expression recognition

    CN112800891A