Micro-expression Recognition Method Based on Video Motion Magnification and Optical Flow Features

Through the combination of video motion amplification and optical flow characteristics, the airspace characteristics of micro-expressions are extracted using LVMM and RAFT networks, and the VGG16 network is used for classification, which solves the problems of low facial motion intensity and short duration in micro-expression recognition, achieving a high recognition accuracy rate.

CN115331289BActive Publication Date: 2025-07-11XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210948759.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-07-11
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

In the prior art, micro-expression recognition methods are difficult to effectively extract subtle movement changes of human faces in video frames, resulting in low recognition accuracy, limiting the promotion and application of micro-expression recognition.

Method used

Using a method based on video motion amplification and optical flow characteristics, the facial muscle movement amplification method (LVMM) is used to enhance the amplification of facial muscles by learning video motion amplification method (LVMM), optical flow characteristics are extracted in combination with deep learning RAFT network, and feature extraction and classification are used using the VGG16 network model.

Benefits of technology

The accuracy of micro-expression recognition has been improved to reach 67.98%, which is better than other mainstream methods, and solves the problems of low facial movement intensity and short duration in micro-expression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331289B_ABST
    Figure CN115331289B_ABST
Patent Text Reader

Abstract

The present invention discloses a micro-expression recognition method based on video motion magnification and optical flow features, specifically: select a data set and classify it according to emotions; preprocess all the original image frame sequences of the selected data set, and use all the obtained single-channel grayscale image sequences as a part of the input of the network model; adopt the "RAFT" network structure based on deep learning to calculate the optical flow features of all the obtained image frame sequences and use the visualized optical flow map as another part of the input of the network model; stack all the single-channel grayscale image sequences and all the visualized RGB optical flow map sequences into a four-channel image, input it into the designed VGG16 network to extract the spatial domain features of micro-expressions and classify to obtain the final recognition accuracy. This method solves the key problems existing in the prior art, such as low facial motion intensity, short duration, and difficulty in extracting the subtle motion changes of human faces in video frames, in the micro-expression recognition method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing and recognition, and particularly relates to a micro-expression recognition method based on video motion magnification and optical flow features. Background Technique

[0002] In recent years, micro-expression recognition has important research value in fields such as criminal investigation lie detection and depression analysis. However, due to the characteristics of micro-expressions themselves, such as small action amplitude and short duration, it is very difficult to manually recognize micro-expressions. Even professional trained psychology researchers have an accuracy rate of only about 47% in recognizing micro-expressions. Because relying on the human eye to recognize micro-expressions is limited by professional training and a large amount of time cost, and at the same time the recognition accuracy rate is low, the large-scale popularization of micro-expression recognition has been seriously hindered. With the rapid development of computer vision and deep learning, more and more researchers apply machine learning algorithms to micro-expression recognition, thus solving many difficulties existing in manual recognition, and the recognition accuracy rate has also been significantly improved. However, due to the characteristics of micro-expressions, so far, how to extract the subtle motion changes of the human face in video frames is still a key issue in this field. Therefore, micro-expression recognition is still in the stage of rapid development, and realizing micro-expression recognition has gradually become an important research topic in the field of affective computing. Summary of the Invention

[0003] The purpose of the present invention is to provide a micro-expression recognition method based on video motion magnification and optical flow features, which solves the key problems existing in the prior art of low facial motion intensity, short duration, and difficulty in extracting the subtle motion changes of the human face in the micro-expression recognition method.

[0004] The technical solution adopted by the present invention is a micro-expression recognition method based on video motion magnification and optical flow features, which is specifically implemented according to the following steps:

[0005] Step 1, select a data set and classify it according to emotions;

[0006] Step 2, preprocess all the original image frame sequences of the selected data set, and use all the obtained single-channel grayscale image sequences as a part of the input of the network model;

[0007] Step 3, adopt the "RAFT" network structure based on deep learning to calculate the optical flow features of all the image frame sequences obtained in Step 2 and use the visualized optical flow map as another part of the input of the network model;

[0008] Step 4, stack all the single-channel grayscale image sequences obtained in Step 2 and all the visualized RGB optical flow map sequences obtained in Step 3 into a four-channel image, input it into the VGG16 network designed by us to extract the spatial domain features of micro-expressions and classify to obtain the final recognition accuracy.

[0009] The features of the present invention also lie in that

[0010] Step 2 is specifically implemented according to the following steps:

[0011] Step 2.1: Adopt a learning-based video motion magnification method to magnify the subtle facial muscle movement amplitudes in all the original image frame sequences of the selected data set;

[0012] Step 2.2: Use the model for detecting 68 key-point information provided by the dlib library to implement the facial alignment operation, crop the facial region, and uniformly adjust its resolution to 224 pixels × 224 pixels;

[0013] Step 2.3: Select the peak frame and the 4 frames before and after it from each micro-expression image sequence, a total of 9 frames of images as key frames to reduce the influence of redundant information in all the image frame sequences obtained in Step 2.2 on recognition;

[0014] Step 2.4: Use the cv2.imread() function to perform grayscale processing on all the image frame sequences obtained in Step 2.3 to obtain single-channel grayscale images, which are used as part of the input to the network model.

[0015] Step 2.1 is specifically implemented according to the following steps:

[0016] First, all adjacent frames (X t-1 , X t ) in all the input original image frame sequences are passed through the encoder H e (·) to obtain their respective shape features (M t-1 , M t ) and texture features (V t-1 , V t );

[0017] Then, the shape features (M t-1 , M t ) of the front and back frames are sent to the amplifier for action amplitude amplification; among them, the amplifier H m (·) is expressed as:

[0018] H m (M t-1 , M t , α) = M t-1 + h(α × g(M t - M t-1 )) (1)

[0019] In formula (1), g(·) is represented by a 3×3 convolution followed by a ReLU activation function, and h(·) is a 3×3 convolution followed by a 3×3 residual block;

[0020] Finally, the decoder reconstructs the changed shape information and the unchanged texture information to generate an enlarged sequence of image frames.

[0021] In step 3, the "RAFT" network structure based on deep learning is adopted to calculate the optical flow features of all the sequences of image frames obtained in step 2.3. The specific steps are as follows:

[0022] First, the feature encoder P θ extracts optical flow features pixel by pixel from adjacent frames (T1, T2) of all the sequences of image frames obtained in step 2.3 and outputs them at 1 / 8 resolution, where the number of channels D of the output feature map is 256; meanwhile, it also includes a context encoder C θ , which only extracts optical flow features from T1; the feature encoder P θ and the context encoder C θ together constitute the feature extraction stage of RAFT, which only needs to be executed once;

[0023] Then, given the image features P θ (T1) and P θ (T2) obtained by feature extraction, by taking the dot product of all pairs of feature vectors (P θ (T1), P θ (T2)), a complete correlation quantity Q is obtained to calculate visual similarity, where the expression of the correlation quantity Q is as follows:

[0024]

[0025] Finally, a recurrent update structure based on gated recurrent units is used to iteratively update the optical flow to generate the final visual optical flow map.

[0026] Step 4 is specifically implemented according to the following steps:

[0027] In step 4.1, a VGG16 network model is designed. The designed VGG16 network model uses 13 convolutional layers and 5 max-pooling layers to be responsible for feature extraction, and zero-padding is used to fill the feature edges before each convolutional layer; the last 3 fully connected layers are responsible for completing the classification task, and the dropout method is applied to the fully connected layers, and its ratio is set to 0.5, that is, dropout = 0.5;

[0028] In step 4.2, the initial learning rate used by the VGG16 network model designed in step 4.1 during training is 10 -5 , and the decay is 10 -6, the epoch is set to 100 and the batch_size is set to 3. After all the parameters are set, all the single-channel grayscale image sequences obtained in step 2.4 and all the visualized RGB optical flow image sequences obtained in step 3 are superimposed into a four-channel image sequence, which is input into the VGG16 network model designed by us to extract its spatial features and use softmax to achieve emotion classification;

[0029] Step 4.3, first, randomly divide the obtained four-channel image sequence into two parts, where 80% is the training set and 20% is the test set;

[0030] Then, use the training set to train the model and the test set to test the accuracy of the model. The calculation method is shown in formula (3) to verify the effectiveness of the model;

[0031]

[0032] Next, in order to reduce errors, we adopt the method of 10 groups of simple cross-validation, shuffle the samples, reselect the training set and the test set, and continue to train and verify the model; repeat this 10 times, obtain the accuracy of 10 groups of models and take the average value as the final accuracy of the model.

[0033] The beneficial effects of the present invention are:

[0034] The method of the present invention combines LVMM and RAFT to process microexpressions, uses the VGG16 network to extract the spatial domain features of microexpressions and classifies them to obtain the microexpression recognition results. At the same time, in order to reduce the influence of redundant information in the microexpression image frame sequence on recognition, we select the key 9 frames of the microexpression sequence on the CASME II dataset for experiments and compare them with 7 other mainstream methods. The experimental results show that our method has obtained better performance, and the recognition accuracy reaches 67.98%. Description of the Drawings

[0035] Figure 1 is the algorithm framework flow chart used in the microexpression recognition method based on video motion amplification and optical flow features of the present invention;

[0036] Figure 2 is the effect comparison diagram after LVMM in the method of the present invention adopts different amplification factors α. Detailed Embodiments

[0037] The present invention will be described in detail below in conjunction with the drawings and specific embodiments.

[0038] The present invention provides a microexpression recognition method based on video motion amplification and optical flow features, as Figure 1 shown, and is specifically implemented according to the following steps:

[0039] Step 1: Select a dataset and classify it into 5 categories according to emotions (happiness, disgust, surprise, repression, others). The publicly available spontaneous micro-expression dataset CASME II released by the team of Fu Xiaolan from the Institute of Psychology, Chinese Academy of Sciences is used. The dataset division is shown in Table 1 as follows:

[0040] Table 1 Division of CASME II Dataset

[0041]

[0042] Step 2: Preprocess all the original image frame sequences of the selected dataset. The specific implementation steps are as follows:

[0043] Step 2.1: Adopt the learning-based video motion magnification method (LVMM) to magnify the amplitude of subtle facial muscle movements in all the original image frame sequences of the selected dataset and enhance the visual features. LVMM mainly consists of three parts: encoder H e (·), amplifier H m (·), and decoder H d (·). In the experiment of magnification using LVMM, first, all adjacent frames (X t-1 , X t ) in all the input original image frame sequences are passed through the encoder H e (·) to obtain their respective shape features (M t-1 , M t ) and texture features (V t-1 , V t ); the obtained texture features are not magnified in terms of motion but are mainly used to constrain the noise caused by subsequent intensity magnification; then, the shape features (M t-1 , M t ) of the front and back frames are sent to the amplifier for action amplitude magnification. The amplifier H m (·) can be expressed as:

[0044] H m (M t-1 , M t , α) = M t-1 + h(α × g(M t - M t-1 )) (1)

[0045] In formula (1), g(·) is represented by a 3×3 convolution followed by a ReLU activation function, and h(·) is a 3×3 convolution followed by a 3×3 residual block; finally, the decoder reconstructs the changed shape information and the unchanged texture information to generate the magnified image frame sequence.

[0046] After repeated experimental comparisons, we finally selected a relatively reasonable amplification factor α = 15. As Figure 2 shown are the results obtained with different amplification factors (α = 5, α = 10, α = 15, α = 20, α = 25) for a certain frame in all the original image frame sequences. We found that when α = 15, while amplifying the image frame effect, the image quality was not affected.

[0047] Step 2.2, in order to minimize the impact of non-face regions in all the amplified image frame sequences obtained in Step 2.1 on micro-expression recognition, we used the model for detecting 68 key-point information provided by the dlib library to implement the face alignment operation, cropped to obtain the face region, and uniformly adjusted its resolution to 224 pixels × 224 pixels so that the input spatial dimension matches the VGG16 network model;

[0048] Step 2.3, considering that the facial movement changes in the micro-expression image sequence are extremely subtle, and the changes between two consecutive frames are almost imperceptible. If all the image frame sequences obtained in Step 2.2 are directly input into the network model for training, it will contain a large amount of redundant features. At the same time, since the shortest duration of a micro-expression is about 1 / 25 second and the frame rate of the samples in the CASME II dataset is 200 frames per second, the shortest continuous frame sequence of a micro-expression can be calculated to be 8 frames. Therefore, the peak frame of each micro-expression image sequence and 4 frames before and after it, a total of 9 frames of images, are selected as key frames to reduce the impact of redundant information in all the image frame sequences obtained in Step 2.2 on recognition;

[0049] Step 2.4, use the cv2.imread() function to perform grayscale processing on all the image frame sequences obtained in Step 2.3 to obtain single-channel grayscale images, which are used as part of the input to the network model.

[0050] Step 3, since optical flow can capture representative motion features between adjacent image frame sequences of micro-expressions, can obtain a higher signal-to-noise ratio, and provide rich and key input features for the network. Therefore, we first adopted the "RAFT" network structure based on deep learning to calculate the optical flow features of all the image frame sequences obtained in Step 2.3 and use the visualized optical flow map as another part of the input to the network model. RAFT extracts optical flow in three steps: First, the feature encoder P θ extracts optical flow features pixel by pixel from adjacent frames (T1, T2) of all the image frame sequences obtained in Step 2.3 and outputs them at 1 / 8 resolution, where the number of channels D of the output feature map is 256. At the same time, it also includes a context encoder C θ , which only extracts optical flow features from T1. The feature encoder P θ and the context encoder C θTogether, they constitute the feature extraction stage of RAFT, which only needs to be executed once; then, given the image features P obtained by feature extraction θ (T1) and P θ (T2), by taking the dot product of all pairs of feature vectors (P θ (T1), P θ (T2)), a complete correlation quantity Q is obtained to calculate visual similarity:

[0051]

[0052] Finally, a recurrent update structure based on the gated recurrent unit (GRU) is used to iteratively update the optical flow to generate the final visual optical flow map.

[0053] Step 4: Stack all the single-channel grayscale image sequences obtained in Step 2.4 and all the visual RGB optical flow map sequences obtained in Step 3 into four-channel images, and input them into the VGG16 network we designed to extract the spatial features of micro-expressions and classify them to obtain the final recognition accuracy. The specific implementation steps are as follows:

[0054] Step 4.1: After the above three steps, we already have all the single-channel grayscale image sequences obtained in Step 2.4 and all the visual RGB optical flow map sequences obtained in Step 3. To complete the classification and recognition of micro-expressions, we designed a VGG16 network model. The VGG16 network is simple and regular, with 13 convolutional layers and 5 max-pooling layers responsible for feature extraction. To ensure that the size of the input image does not change, zero-padding is used to fill the feature edges before each convolutional layer. The last 3 fully connected layers are responsible for completing the classification task. To reduce the overfitting phenomenon, we applied the dropout method to the fully connected layers. It can randomly mask the number of neurons according to the set parameters, improve the generalization ability of the network model, and also speed up the training speed of the network. Referring to the empirical value, we set its ratio to 0.5, that is, dropout = 0.5.

[0055] Step 4.2: The initial learning rate (lr) used by the VGG16 network model we designed during training is 10 -5 , the decay is 10 -6 , the epoch is set to 100, and the batch_size is set to 3. After all the parameter settings are completed, stack all the single-channel grayscale image sequences obtained in Step 2.4 and all the visual RGB optical flow map sequences obtained in Step 3 into a four-channel image sequence, input it into the VGG16 network model to extract its spatial features and use softmax to achieve emotion classification.

[0056] Step 4.3. Since the cross-validation of the model can, to a certain extent, avoid the error phenomenon caused by dataset division, 10 groups of simple cross-validation are adopted in our experiment to reduce this error. Specific operation: First, randomly divide the obtained four-channel image sequence into two parts (80% training set and 20% test set); then, train the model with the training set and test the accuracy of the model with the test set (the calculation method is shown in Formula 3) to verify the effectiveness of the model; next, adopt the method of 10 groups of simple cross-validation, shuffle the samples, re-select the training set and test set, and continue to train and verify the model. As shown in Table 2, the final recognition result is the average value of 10 groups of experiments, which is 67.98%.

[0057]

[0058] Table 2 Training Results of 10 Groups

[0059]

[0060] Comparison and verification of experimental results: The average value of the 10 groups of experimental results obtained in Step 4.3 (as shown in Table 2) was compared with other existing methods, as shown in Table 3, including traditional methods LBP-TOP, STLBP-IP, Bi-WOOF, and deep learning methods ELRCN-SE, CNN+LSTM, CNNCapsNet, MSCNN. The results show that the recognition accuracy of our method is 3.35% higher than that of the sub-optimal method CNNCapsNet, and a better micro-expression recognition effect is obtained. This method solves the key problems existing in the existing micro-expression recognition methods, such as low facial movement intensity, short duration, and difficulty in extracting subtle movement changes of the human face in video frames.

[0061] Table 3 Performance Comparison between the Method in This Paper and Existing Methods on CASME II

[0062]

Claims

1. A micro-expression recognition method based on video motion magnification and optical flow features, characterized in that, The implementation is carried out according to the following steps: Step 1: Select a dataset and classify it according to emotions; Step 2: Preprocess all the original image frame sequences of the selected dataset, and all the obtained single-channel grayscale image sequences are used as part of the input to the network model; Step 2 is specifically implemented according to the following steps: Step 2.1: Adopt a learning-based video motion magnification method to magnify the subtle facial muscle movement amplitudes in all the original image frame sequences of the selected dataset; Step 2.1 is specifically implemented according to the following steps: First, all adjacent frames (X t-1 , X t ) in all the input sequences of original image frames are passed through the encoder H e (·) to obtain their respective shape features (M t-1 , M t ) and texture features (V t-1 , V t ); Then, the shape features of the front and rear frames (M t-1 , M t ) are sent to an amplifier for amplifying the action amplitude; among them, the amplifier H m (·) is expressed as: (1) In formula (1), g(·) is represented by a 3×3 convolution followed by a ReLU activation function, and h(·) is a 3×3 convolution followed by a 3×3 residual block; Finally, the decoder reconstructs the changed shape information and the unchanged texture information to generate an amplified image frame sequence; Step 2.2: Use the model for detecting 68 key point information provided by the dlib library to implement facial alignment operations, crop to obtain the facial region, and uniformly adjust its resolution to 224 pixels × 224 pixels; Step 2.3: Select the peak frame of each micro-expression image sequence and 4 frames before and after it, a total of 9 frames of images as key frames to reduce the influence of redundant information in all the image frame sequences obtained in Step 2.2 on recognition; Step 2.4: Use the cv2.imread() function to perform grayscale processing on all the image frame sequences obtained in Step 2.3 to obtain single-channel grayscale images, which are used as part of the input to the network model; Step 3: Adopt a deep learning-based "RAFT" network structure to calculate the optical flow features of all the image frame sequences obtained in Step 2, and the visualized optical flow map is used as another part of the input to the network model; Step 4: Stack all the single-channel grayscale image sequences obtained in Step 2 and all the visualized RGB optical flow map sequences obtained in Step 3 into a four-channel image, and input it into the designed VGG16 network to extract the spatial domain features of micro-expressions and classify to obtain the final recognition accuracy.

2. The micro-expression recognition method based on video motion magnification and optical flow features according to claim 1, characterized in that, In Step 3, the specific steps for adopting a deep learning-based "RAFT" network structure to calculate the optical flow features of all the image frame sequences obtained in Step 2.3 are as follows: First, the feature encoder P θ extracts optical flow features pixel by pixel from adjacent frames (T1, T2) of all image frame sequences obtained in step 2.3 and outputs them at 1 / 8 resolution, where the number of channels D of the output feature map is 256; at the same time, it also includes a context encoder C θ , which only extracts optical flow features from T1; Feature encoder P θ and the context encoder C θ together constitute the feature extraction stage of RAFT, which is only executed once; Then, given the image features obtained by feature extraction and , by taking the dot product of all feature vector pairs ( ), a complete correlation quantity Q is obtained to calculate visual similarity, and the expression of the correlation quantity Q is as follows: (2) Finally, use a recurrent update structure based on a gated recurrent unit to iteratively update the optical flow to generate the final visualized optical flow map.

3. The micro-expression recognition method based on video motion magnification and optical flow features according to claim 2, characterized in that Step 4 is specifically implemented according to the following steps: Step 4.1: Design a VGG16 network model. The designed VGG16 network model uses 13 convolutional layers and 5 max-pooling layers to be responsible for feature extraction, and zero-padding is used to fill the feature edges before each convolutional layer; the last 3 fully connected layers are responsible for completing the classification task, and the dropout method is applied to the fully connected layers, and its ratio is set to 0.5, that is, dropout = 0.5; Step 4.2, the initial learning rate used by the VGG16 network model designed in Step 4.1 during training is 10 -5 , and the decay is 10 -6 , the epoch is set to 100, and the batch_size is set to 3; after all the parameters are set, all the single-channel grayscale image sequences obtained in Step 2.4 and all the visualized RGB optical flow map sequences obtained in Step 3 are superimposed into a four-channel image sequence, which is input into the designed VGG16 network model to extract its spatial features and use softmax to achieve emotion classification; Step 4.3: First, randomly divide the obtained four-channel image sequence into two parts, among which, 80% is the training set and 20% is the test set; Then, use the training set to train the model and the test set to test the accuracy of the model. The calculation method is shown in formula (3) to verify the effectiveness of the model; (3) Next, using the method of 10-fold simple cross-validation, shuffle the samples, reselect the training set and the test set, continue to train the data and validate the model; repeat this 10 times, obtain the accuracies of 10 groups of models and take the average value as the final accuracy of the model.

Citation Information

Patent Citations

  • Micro-expression recognition method based on adaptive motion amplification and convolutional neural network

    CN113537008A

  • Face emotion recognition method based on dual-stream convolutional neural network

    US20190311188A1