Face micro-expression recognition method, system and terminal
Through the pre-trained face micro-expression recognition model, the problem of weak generalization ability of traditional methods is solved by using dense optical flow and feature fusion technology, and the micro-expression recognition with high accuracy and robustness is achieved, which improves the visualization effect.
Patent Information
- Application Number
- CN202510516412.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-22
AI Technical Summary
When facing a diverse face, traditional micro-expression recognition methods have weak generalization capabilities and are difficult to ensure the accuracy of recognition, resulting in incorrect judgments and affecting their application effects in various fields.
A pre-trained face micro-expression recognition model is used to obtain the starting frame, intermediate frame and end frame images of the target video, and dense optical flow is extracted, and feature fusion and classification decisions are made in combination with the CNN layer, the deep feature extraction layer, the feature fusion layer and the Softmax function layer to identify seven types of micro-expressions.
It improves the generalization ability of micro-expression recognition, ensures the accuracy and robustness of recognition, improves visualization effect, reduces misjudgment, and improves recognition accuracy.
Smart Images

Figure CN120526464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of facial micro-expressions, and in particular to a facial micro-expression recognition method, system and terminal. Background Art
[0002] Facial microexpressions, fleeting facial expressions, are widely used in fields such as psychology, clinical medicine, security, and interrogation. For example, their application in student psychology can facilitate the identification of emotions such as anxiety and aversion to learning, enabling timely intervention. Furthermore, microexpressions, combined with behavioral analysis, can help assess student motivation and stress levels, providing a basis for personalized education.
[0003] However, most traditional micro-expression recognition methods rely on a simple classification strategy for each frame in a video. This involves using the corresponding labels to train deep learning models, enabling them to learn the characteristic patterns of different expressions. This approach has significant limitations and poor generalization. Furthermore, facial structure and muscle orientation vary from person to person, and the visual characteristics of the same micro-expression can also vary from person to person. This makes traditional methods prone to misjudgment due to feature extraction bias when dealing with diverse faces, making it difficult to ensure accurate micro-expression recognition, which in turn affects the effectiveness of their application in various fields. Summary of the Invention
[0004] The object of the present invention is to provide a method, system and terminal for recognizing facial micro-expressions, so as to solve at least one of the above problems.
[0005] In a first aspect, the technical solution of the present invention provides a method for recognizing facial micro-expressions, the method comprising: Step L1, obtaining the target video; Step L2: extract the first frame, middle frame, and last frame images of the target video as its start frame, peak frame, and end frame; Step L3, extracting dense optical flow between the start frame, the vertex frame and the end frame; Step L4: input the dense optical flow and the vertex frame into a pre-trained facial micro-expression recognition model for calculation to obtain the facial micro-expressions in the target video; The pre-trained facial micro-expression recognition model is a neural network model that can recognize seven pre-set micro-expression categories, including: Input layer, used for data input of the model; The CNN layer is used to extract features from the vertex frame and dense optical flow input by the input layer to obtain the first preliminary feature and the second preliminary feature; A deep feature extraction layer, configured to extract spatial features of the first preliminary features and temporal features of the second preliminary features; A feature fusion layer is used to fuse the spatial features and the temporal features to obtain fused features; A fully connected layer, configured to map the fused features into a low-dimensional classification space and output score vectors for the seven pre-set micro-expression categories; Softmax function layer, used to convert the output of the fully connected layer into a probability distribution and output it; The classification decision layer is used to select the micro-expression category with the highest probability as the model's prediction result based on the probability distribution output by the Softmax function layer, and output it.
[0006] Furthermore, the method for obtaining the target video includes: Obtain the video to be used for facial micro-expression recognition; According to the preset reading granularity, the segments of the video to be subjected to facial micro-expression recognition are read one by one to obtain the target video; the value range of the preset reading granularity is: 1 second 2 seconds.
[0007] Furthermore, step L4 also includes: after obtaining the facial micro-expressions in the target video, marking the obtained facial micro-expressions on the corresponding video segment of the video to be subjected to facial micro-expression recognition.
[0008] Furthermore, the method further comprises: Detecting a face in the video to be subjected to facial micro-expression recognition, and drawing a face frame for the detected face; Step L4 also includes: after obtaining the facial micro-expressions in the target video, marking the obtained facial micro-expressions on the top of the face frame on the corresponding video segment in the video to be subjected to facial micro-expression recognition.
[0009] Furthermore, step L3 specifically includes: Detect whether the starting frame, the vertex frame, and the ending frame all contain facial images. If so, execute step J; if not, stop recognizing facial micro-expressions in the current target video; Step J: extracting dense optical flow between the start frame, the vertex frame and the end frame.
[0010] Furthermore, in step J, the dense optical flow between the start frame, the vertex frame, and the end frame is extracted, including: Step J1: For each frame image in the starting frame, the vertex frame, and the ending frame, the following steps are performed respectively to obtain the corresponding frame face image of each frame image: Convert the frame image into a grayscale image; Detect and extract all faces in a grayscale image: If the number of extracted face images is 1, the extracted face image is saved as the corresponding frame face image; If the number of extracted face images 2. Calculate the proportion of each extracted face image in the corresponding frame image, and then select the face image with the largest proportion and save it as the corresponding frame face image; The corresponding frame face images of the obtained starting frame, vertex frame and ending frame are the starting frame face image, vertex frame face image and ending frame face image respectively; Step J2: extract dense optical flows between the vertex frame face image and the start frame face image and between the vertex frame face image and the end frame face image, and obtain the dense optical flows to be extracted between the start frame, the vertex frame and the end frame.
[0011] Furthermore, in step J2, the method for extracting dense optical flow between the vertex frame face image and the start frame face image and between the vertex frame face image and the end frame face image includes: Step 1: Perform the following processing on the face images in the starting frame face image, the vertex frame face image, and the ending frame face image, respectively, to obtain eighteen target face regions in the starting frame face image, the vertex frame face image, and the ending frame face image: Use the 68-point annotation method to annotate the face in the face image, and obtain 68-point annotations of the face in the face image; Based on the 68-point annotations, extract the areas corresponding to the left eye, right eye, mouth, nose, left eyelid, and right eyelid in the face image; Divide the extracted area corresponding to the left eye into three areas according to a preset first division method, divide the area corresponding to the right eye into three areas according to a preset second division method, divide the extracted area corresponding to the mouth into four areas according to a preset third division method, and divide the extracted area corresponding to the nose into a left area of the nose bridge and a right area of the nose bridge; Gather the extracted regions and the divided regions to obtain eighteen face regions corresponding to the face image; The obtained 18 face regions of the starting frame face image are the 18 target face regions of the starting frame face image; Step 2: Calculate the dense optical flow between the face image of the starting frame and the face image of the vertex frame, which is recorded as the first dense optical flow; Step 3: Calculate the horizontal component of the first dense optical flow And the component y1 in the vertical direction, then use Calculate the modulus m of the first dense optical flow; if , then the vertex frame face image is used as the target vertex frame face image, and the obtained 18 face regions of the vertex frame face image are used as the 18 target face regions of the vertex frame face image, and then step five is executed; if , then the vertex frame face image is translated horizontally , pan in the vertical direction , obtain the adjusted vertex frame face image, and translate the eighteen face regions of the vertex frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the adjusted vertex frame face image, and then execute step 4; M is a preset modulus length threshold; Step 4: Calculate the dense optical flow between the face image of the starting frame and the face image of the vertex frame after the latest adjustment, which is recorded as the second dense optical flow. Then calculate the horizontal component of the second dense optical flow. and the vertical component , and calculate the modulus of the second dense optical flow , : if , then the most recently adjusted vertex frame face image is used as the target vertex frame face image, and the eighteen face regions of the most recently adjusted vertex frame face image are used as the eighteen target face regions of the vertex frame face image, and then the process goes to step five to continue; if , then the face image of the vertex frame that has been adjusted the most recently is translated in the horizontal direction , pan in the vertical direction , get the next adjusted vertex frame face image, and translate the eighteen face areas of the latest adjusted vertex frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the next adjusted vertex frame face image, and then execute step 4 again; Step 5: Calculate the dense optical flow between the target vertex frame face image and the end frame face image, which is recorded as the third dense optical flow; Step 6: Calculate the horizontal component of the third dense optical flow and the component in the vertical direction , and then use Calculate the modulus of the third dense optical flow : if , the obtained 18 face regions of the end frame face image are used as the 18 target face regions of the end frame face image, and then step 8 is executed; if , then the end frame face image is translated horizontally , pan in the vertical direction , obtain the adjusted end frame face image, and translate the eighteen face regions of the obtained end frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the adjusted end frame face image, and then execute step seven; Step 7: Calculate the dense optical flow between the target vertex frame face image and the current most recently adjusted end frame face image, which is recorded as the fourth dense optical flow. Then calculate the horizontal component of the fourth dense optical flow. and the component in the vertical direction , and calculate the modulus of the fourth dense optical flow : if , then the eighteen face regions of the most recently adjusted ending frame face image are used as the eighteen target face regions of the ending frame face image, and then step eight is executed; if , then the face image of the end frame after the latest adjustment is horizontally shifted , pan in the vertical direction , get the next adjusted end frame face image, and translate the eighteen face areas of the latest adjusted end frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the next adjusted end frame face image, and then execute step seven again; Step eight, calculating dense optical flows between the eighteen target face regions of the starting frame face image and the eighteen target face regions of the vertex frame face image, to obtain dense optical flows between the starting frame face image and the vertex frame face image; The dense optical flow between the eighteen target face regions of the vertex frame face image and the eighteen target face regions of the end frame face image is calculated to obtain the dense optical flow between the vertex frame face image and the end frame face image.
[0012] Furthermore, the method for obtaining the trained facial micro-expression recognition model includes: Step S1, constructing a training set; Step S2, constructing a network framework for a facial micro-expression recognition model; Step S3, using the training set to train the network framework to obtain the trained facial micro-expression recognition model; In step S1, a training set is constructed, including: Selecting facial data of the seven pre-set micro-expressions from the CASME facial data set to construct a first data set; each facial data in the first data set corresponds to a micro-expression; each facial data in the first data set is composed of three frames of images: a start frame, a peak frame, and an end frame corresponding to the micro-expression; Normalizing the facial data in the first data set to obtain normalized facial data; For each normalized face data, the dense optical flow between the start frame, vertex frame and end frame is extracted; Setting micro-expression classification labels for the seven types of micro-expressions; For each normalized face data, its vertex frame and the extracted dense optical flow between its three frames are taken as a sample; Use pre-set micro-expression classification labels to annotate each sample with micro-expressions; The samples marked with micro-expression classification labels are collected to obtain the training set.
[0013] In a second aspect, the present invention provides a facial micro-expression recognition system, the system comprising: Acquisition module, used to acquire target video; An image extraction module is used to extract the first frame, middle frame and last frame images of the target video as the start frame, top frame and end frame of the target video; an optical flow extraction module, configured to extract dense optical flow between the start frame, the vertex frame, and the end frame; A micro-expression recognition module is used to input the dense optical flow and the vertex frame into a pre-trained facial micro-expression recognition model for calculation to obtain facial micro-expressions in the target video; The pre-trained facial micro-expression recognition model is a neural network model that can recognize seven pre-set micro-expression categories, including: Input layer, used for data input of the model; The CNN layer is used to extract features from the vertex frame and dense optical flow input by the input layer to obtain the first preliminary feature and the second preliminary feature; A deep feature extraction layer, configured to extract spatial features of the first preliminary features and temporal features of the second preliminary features; A feature fusion layer is used to fuse the spatial features and the temporal features to obtain fused features; A fully connected layer, configured to map the fused features into a low-dimensional classification space and output score vectors for the seven pre-set micro-expression categories; Softmax function layer, used to convert the output of the fully connected layer into a probability distribution and output it; The classification decision layer is used to select the micro-expression category with the highest probability as the model's prediction result based on the probability distribution output by the Softmax function layer, and output it.
[0014] In a third aspect, the present invention provides a terminal, comprising: processor; A memory for storing execution instructions of the processor; The processor is configured to execute the methods described in the above aspects.
[0015] It can be seen from the above technical solutions that the present invention has the following advantages: The present invention recognizes facial micro-expressions based on a pre-trained facial micro-expression recognition model. The pre-trained facial micro-expression recognition model includes an input layer, a CNN layer, a deep feature extraction layer, a feature fusion layer, a fully connected layer, a Softmax function layer and a classification decision layer. When used, the vertex frame of the target video and the extracted dense optical flow between the start frame, vertex frame and end frame of the target video are input, and the facial micro-expressions in the target video are output after calculation. It can be seen that the present invention does not recognize by simply classifying each frame of the video, but recognizes micro-expressions based on the input of dense optical flow and vertex frame of the face, that is, the present invention recognizes micro-expressions based on changes in the face, which helps to improve the generalization ability of micro-expression recognition to a certain extent.
[0016] After obtaining the facial micro-expressions in the target video, the present invention can mark the obtained facial micro-expressions on the corresponding video segment of the video to be used for facial micro-expression recognition corresponding to the target video, which helps to improve the visualization effect of facial micro-expressions to a certain extent.
[0017] The present invention can detect faces in the video to be subjected to facial micro-expression recognition, draw face frames for the detected faces, and mark the obtained facial micro-expressions in the target video on the top of the face frame on the corresponding video segment in the video to be subjected to facial micro-expression recognition, which further helps to improve the visualization effect of facial micro-expressions.
[0018] The present invention adopts a face detection mechanism. Only when the starting frame, vertex frame and ending frame all contain face images, the dense optical flow between the starting frame, vertex frame and ending frame is extracted for facial micro-expression recognition. This helps the present invention avoid misjudgment caused by the failure to detect faces in the video to be used for facial micro-expression recognition, and to a certain extent helps to ensure the reliability and robustness of recognition.
[0019] The present invention extracts optical flows of different facial regions (i.e., eighteen regions) on the starting frame facial image, the vertex frame facial image, and the ending frame facial image, and uses the extracted dense optical flows between the starting frame, the vertex frame, and the ending frame for facial micro-expression recognition, which helps to better reflect the subtle changes in various facial regions and thus helps to improve recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.
[0022] Figure 2 FIG. 4 is a schematic block diagram of a system according to an embodiment of the present invention.
[0023] Figure 3 A schematic diagram of the structure of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0024] To make the technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them.
[0025] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention, its application, or use. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0026] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0027] It should be noted that the present invention uses words such as "first" and "second" to limit the corresponding parts only to facilitate the distinction of the corresponding parts. Unless otherwise stated, the above words have no special meaning and therefore cannot be understood as limiting the scope of protection of the present invention.
[0028] The key terms appearing in the present invention are explained below.
[0029] CNN, which is a convolutional neural network model.
[0030] ViT, or Vision Transformer, is a Transformer-based neural network model architecture that captures spatial information in input images by introducing self-attention mechanism and position encoding.
[0031] LSTM, which stands for Long Short-Term Memory, is a long short-term memory network.
[0032] The CASME (Chinese Academy of Sciences Micro-Expressions) face dataset is a public facial micro-expression dataset. Each data in the dataset is a micro-expression sample, and each micro-expression sample is annotated with the micro-expression, including the start frame, vertex frame, and end frame.
[0033] Dense optical flow is a representation of optical flow. It generates an optical flow field by calculating the motion offset of each pixel in the image, where each pixel corresponds to a two-dimensional velocity vector, which is used to describe the instantaneous pixel motion speed of a moving object in space on the imaging plane.
[0034] Figure 1 A schematic flow chart of a facial micro-expression recognition method provided by an embodiment of the present invention. Figure 1 The execution subject may be a facial micro-expression recognition system.
[0035] The facial micro-expression recognition method provided by the embodiment of the present invention is executed by a computer device. Accordingly, the facial micro-expression recognition system runs in the computer device.
[0036] Please refer to Figure 1 This embodiment provides a method for recognizing facial micro-expressions, including: Step 110: Obtain target video; Step 120: extract the first frame, middle frame, and last frame images of the target video as its start frame, top frame, and end frame; Step 130: extracting dense optical flow between the start frame, the vertex frame, and the end frame; Step 140: Input the dense optical flow and the vertex frame into a pre-trained facial micro-expression recognition model for calculation to obtain facial micro-expressions in the target video.
[0037] The above-mentioned pre-trained facial micro-expression recognition model is a neural network model that can recognize seven pre-set categories of micro-expressions.
[0038] In this embodiment, the pre-trained facial micro-expression recognition model specifically includes: Input layer, used for data input of the model; The CNN layer is used to extract features from the vertex frame and dense optical flow input by the input layer to obtain the first preliminary feature and the second preliminary feature; A deep feature extraction layer, configured to extract spatial features of the first preliminary features and temporal features of the second preliminary features; A feature fusion layer is used to fuse the spatial features and the temporal features to obtain fused features; A fully connected layer, configured to map the fused features into a low-dimensional classification space and output score vectors for the seven pre-set micro-expression categories; Softmax function layer, used to convert the output of the fully connected layer into a probability distribution and output it; The classification decision layer is used to select the micro-expression category with the highest probability as the model's prediction result based on the probability distribution output by the Softmax function layer, and output it.
[0039] In this embodiment, the CNN layer includes a first CNN layer and a second CNN layer. The first CNN layer is used to extract features from the vertex frame input by the input layer to obtain preliminary features of the frame image input by the input layer, namely, first preliminary features. The second CNN layer is used to extract features from the dense optical flow input by the input layer to obtain preliminary features of the dense optical flow input by the input layer, namely, second preliminary features.
[0040] In this embodiment, the deep feature extraction layer includes a ViT layer and an LSTM layer. The ViT layer is used to extract the spatial features of the first preliminary features. The LSTM layer is used to extract the temporal features of the second preliminary features.
[0041] The present invention uses a fully connected layer as a classifier, which first converts the features fused in the model into a 96-dimensional feature vector through a flattening operation, and then uses formula ① to compress the 96-dimensional feature vector into a 7-dimensional feature vector. Each element in the feature vector corresponds to the score of the seven pre-set micro-expression categories mentioned above.
[0042] Formula ① is: , Where x is the 96-dimensional feature vector, W is the weight matrix, and b is the bias term.
[0043] Substituting the above 96-dimensional feature vector into the above formula ①, the calculated y is the feature vector with a dimension of 7 mentioned above.
[0044] The Softmax function layer uses the Softmax function to convert the 7-dimensional output of the fully connected layer into the conditional probability distribution of the category. The probability of each category is between (0, 1), and the sum of the probabilities of all categories is 1.
[0045] When the trained facial micro-expression recognition model is used, the vertex frames of the read video segments and the corresponding extracted dense optical flows are taken as a set of data and input through the input layer. Then, the preliminary features of the input vertex frames are extracted through the first CNN layer, and the preliminary features of the input dense optical flows are extracted through the second CNN layer. Subsequently, the extracted preliminary features of the vertex frames are sent to ViT to extract the spatial features of the image (tensor array), and the preliminary features of the optical flows (numpy arrays) are sent to LSTM to extract the temporal features (a set of tensor arrays). Then, the feature fusion layer fuses the spatial features extracted by ViT with the temporal features extracted by LSTM to obtain a high-dimensional feature, and inputs the feature into the fully connected layer. The fully connected layer maps the input features to a low-dimensional classification space and outputs the scores of each category (i.e., the seven categories of micro-expression classification corresponding to the seven pre-set micro-expressions mentioned above) to the Softmax function layer (this layer is implemented using the Softmax function). The Softmax function layer converts the output of the fully connected layer into a probability distribution and outputs it to the classification decision layer. The classification decision layer outputs the prediction result based on the probability distribution output by the Softmax function layer.
[0046] Optionally, the feature fusion layer concatenates and fuses the spatial features extracted by ViT and the temporal features extracted by LSTM in the column direction.
[0047] In an optional embodiment, the method for acquiring the target video includes: Obtain the video to be used for facial micro-expression recognition; According to the preset reading granularity, the segments of the video to be subjected to facial micro-expression recognition are read one by one to obtain the target video.
[0048] The value range of the above preset reading granularity is: 1 second 2 seconds.
[0049] It is understandable that the video to be used for facial micro-expression recognition can be video data stored locally, or can be video data collected in real time by a camera.
[0050] It can be understood that in the present invention, reading the segments of the video to be subjected to facial micro-expression recognition one by one according to the preset reading granularity is equivalent to segmenting the video to be subjected to facial micro-expression recognition.
[0051] It can be understood that the present invention reads the segments of the video to be subjected to facial micro-expression recognition one by one according to a preset reading granularity, and each read video segment is a target video.
[0052] In this embodiment, the preset reading granularity is 1 second.
[0053] The present invention reads a video of 1 second length from the video to be used for facial micro-expression recognition each time, and starts reading from the beginning of the video.
[0054] It can be understood that in this embodiment, the video to be used for facial micro-expression recognition is first obtained, and then the segments of the video to be used for facial micro-expression recognition are read one by one according to the preset reading granularity (1 second) to obtain the target video. For each target video read: the first frame, middle frame and last frame image of the target video are extracted as its starting frame, vertex frame and ending frame; then the dense optical flow between the starting frame, vertex frame and ending frame of the current target video is extracted; then the extracted dense optical flow and the vertex frame of the current target video are input into a pre-trained facial micro-expression recognition model for calculation to obtain the facial micro-expression in the current target video.
[0055] In an optional implementation, step 140 further includes: after obtaining the facial micro-expressions in the target video, marking the obtained facial micro-expressions on the corresponding video segment of the video to be subjected to facial micro-expression recognition.
[0056] That is, after obtaining the facial micro-expressions in the target video, step 140 marks the obtained facial micro-expressions on the video segment corresponding to the target video in the video to be subjected to facial micro-expression recognition.
[0057] In an optional embodiment, the method further includes: detecting a face in the video to be subjected to facial micro-expression recognition, and drawing a face frame for the detected face.
[0058] In an optional embodiment, step 140 further includes: after obtaining the facial micro-expressions in the target video, marking the obtained facial micro-expressions on the top of the face frame on the corresponding video segment in the first video. The first video is the video to be subjected to facial micro-expression recognition.
[0059] After identifying facial micro-expressions in the target video, the present invention marks the identified facial micro-expressions on the top of the face frame on the corresponding video segment in the first video. This helps to improve the visualization effect, making it easier to intuitively understand the micro-expressions, thereby helping to quickly identify abnormal emotions and take timely intervention measures.
[0060] In a specific implementation, the face detector of dlib can be used to detect the face in the video to be subjected to facial micro-expression recognition, and draw the face frame.
[0061] Optionally, when the target video is a real-time video, faces in the video to be subjected to facial micro-expression recognition are detected in real time, and face frames are drawn for the detected faces. In a specific implementation, a dlib face detector can be used to dynamically detect faces in the video to be subjected to facial micro-expression recognition and draw face frames.
[0062] In this embodiment, step 130 specifically includes: Detect whether the starting frame, the vertex frame, and the ending frame all contain facial images. If so, execute step J; if not, stop recognizing facial micro-expressions in the current target video; Step J: extracting dense optical flow between the start frame, the vertex frame and the end frame.
[0063] Illustratively, in step 140 , shape_predictor_68_face_landmarks (a facial landmark detection model) is used to detect facial images in the start frame, the vertex frame, and the end frame.
[0064] The present invention extracts dense optical flow between the three frames only when a face can be detected in the start frame, the vertex frame, and the end frame. This helps the present invention avoid misjudgment caused by failure to detect a face in the target video.
[0065] For example, each time the recognition of facial micro-expressions in a video segment is stopped, “No face detected” may be displayed on the corresponding video segment of the target video.
[0066] In an optional embodiment, extracting dense optical flow between the start frame, the vertex frame, and the end frame in step J includes the following steps J1 and J2.
[0067] Step J1: For each frame image in the starting frame, the vertex frame, and the ending frame, the following steps are performed respectively to obtain the corresponding frame face image of each frame image: Convert the frame image into a grayscale image; Detect and extract all faces in a grayscale image: If the number of extracted face images is 1, the extracted face image is saved as the corresponding frame face image; If the number of extracted face images 2. Calculate the proportion of each extracted face image in the corresponding frame image, and then select the face image with the largest proportion and save it as the corresponding frame face image.
[0068] Among them, the corresponding frame face image of the obtained starting frame is the starting frame face image, the corresponding frame face image of the obtained vertex frame is the vertex frame face image, and the corresponding frame face image of the obtained ending frame is the ending frame face image.
[0069] Step J2: extract dense optical flows between the vertex frame face image and the start frame face image and between the vertex frame face image and the end frame face image, and obtain the dense optical flows to be extracted between the start frame, the vertex frame and the end frame.
[0070] Extract dense optical flow between the vertex frame face image and the starting frame face image and between the vertex frame face image and the ending frame face image, that is, extract dense optical flow between the starting frame face image and the vertex frame face image, and extract dense optical flow between the vertex frame face image and the ending frame face image.
[0071] Optionally, the above step J1 specifically includes: (1) Detect and extract all face images in the starting frame: If the number of extracted face images is 1, the extracted face image is saved as the starting frame face image; If the number of extracted face images 2. Calculate the proportion of each extracted face image in the entire starting frame, and then select the face image with the largest proportion and save it as the starting frame face image; (2) Detect and extract all face images in the vertex frame: If the number of extracted face images is 1, the extracted face image is saved as a vertex frame face image; If the number of extracted face images 2. Calculate the proportion of each extracted face image in the entire vertex frame, and then select the face image with the largest proportion and save it as the vertex frame face image; (3) Detect and extract all face images in the end frame: If the number of extracted face images is 1, the extracted face image is saved as the end frame face image; If the number of extracted face images 2. Calculate the proportion of each extracted face image in the entire end frame, and then select the face image with the largest proportion and save it as the end frame face image.
[0072] In an optional embodiment, in step J2, the method for extracting dense optical flow between the vertex frame face image and the start frame face image and between the vertex frame face image and the end frame face image includes: Step 1: Perform the following processing on the face images in the starting frame face image, the vertex frame face image, and the ending frame face image, respectively, to obtain eighteen target face regions in the starting frame face image, the vertex frame face image, and the ending frame face image: Use the 68-point annotation method to annotate the face in the face image, and obtain 68-point annotations of the face in the face image; Extracting the areas corresponding to the left eye, right eye, mouth, nose, left eyelid, and right eyelid in the face image based on the 68-point annotations; Divide the extracted area corresponding to the left eye into three areas according to a preset first division method, divide the area corresponding to the right eye into three areas according to a preset second division method, divide the extracted area corresponding to the mouth into four areas according to a preset third division method, and divide the extracted area corresponding to the nose into a left area of the nose bridge and a right area of the nose bridge; Gather the extracted regions and the divided regions to obtain eighteen face regions corresponding to the face image; The obtained 18 face regions of the starting frame face image are the 18 target face regions of the starting frame face image; Step 2: Calculate the dense optical flow between the face image of the starting frame and the face image of the vertex frame, which is recorded as the first dense optical flow; Step 3: Calculate the horizontal component of the first dense optical flow And the component y1 in the vertical direction, then use Calculate the modulus m of the first dense optical flow; if , then the vertex frame face image is used as the target vertex frame face image, and the obtained 18 face regions of the vertex frame face image are used as the 18 target face regions of the vertex frame face image, and then step five is executed; if , then the vertex frame face image is translated horizontally , pan in the vertical direction , obtain the adjusted vertex frame face image, and translate the eighteen face regions of the vertex frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the adjusted vertex frame face image, and then execute step 4; M is a preset modulus length threshold; Step 4: Calculate the dense optical flow between the face image of the starting frame and the face image of the vertex frame after the latest adjustment, which is recorded as the second dense optical flow. Then calculate the horizontal component of the second dense optical flow. and the vertical component , and calculate the modulus of the second dense optical flow , : if , then the most recently adjusted vertex frame face image is used as the target vertex frame face image, and the eighteen face regions of the most recently adjusted vertex frame face image are used as the eighteen target face regions of the vertex frame face image, and then the process goes to step five to continue; if , then the face image of the vertex frame that has been adjusted the most recently is translated in the horizontal direction , pan in the vertical direction , get the next adjusted vertex frame face image, and translate the eighteen face areas of the latest adjusted vertex frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the next adjusted vertex frame face image, and then execute step 4 again; Step 5: Calculate the dense optical flow between the target vertex frame face image and the end frame face image, which is recorded as the third dense optical flow; Step 6: Calculate the horizontal component of the third dense optical flow and the component in the vertical direction , and then use Calculate the modulus of the third dense optical flow : if , the obtained 18 face regions of the end frame face image are used as the 18 target face regions of the end frame face image, and then step 8 is executed; if , then the end frame face image is translated horizontally , pan in the vertical direction , obtain the adjusted end frame face image, and translate the eighteen face regions of the obtained end frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the adjusted end frame face image, and then execute step seven; Step 7: Calculate the dense optical flow between the target vertex frame face image and the current most recently adjusted end frame face image, which is recorded as the fourth dense optical flow. Then calculate the horizontal component of the fourth dense optical flow. and the component in the vertical direction , and calculate the modulus of the fourth dense optical flow : if , then the eighteen face regions of the most recently adjusted ending frame face image are used as the eighteen target face regions of the ending frame face image, and then step eight is executed; if , then the face image of the end frame after the latest adjustment is horizontally shifted , pan in the vertical direction , get the next adjusted end frame face image, and translate the eighteen face areas of the latest adjusted end frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the next adjusted end frame face image, and then execute step seven again; Step eight, calculating dense optical flows between the eighteen target face regions of the starting frame face image and the eighteen target face regions of the vertex frame face image, to obtain dense optical flows between the starting frame face image and the vertex frame face image; The dense optical flow between the eighteen target face regions of the vertex frame face image and the eighteen target face regions of the end frame face image is calculated to obtain the dense optical flow between the vertex frame face image and the end frame face image.
[0073] In this embodiment, the method for calculating dense optical flow between eighteen target face regions of a starting frame face image and eighteen target face regions of a vertex frame face image includes: Traversing each of the eighteen target face regions of the face image of the starting frame; For each traversed face region, the dense optical flow between it and the corresponding region in the eighteen target face regions of the vertex frame face image is calculated to obtain an optical flow field; All the obtained optical flow fields are collected to obtain the dense optical flow between the starting frame face image and the vertex frame face image.
[0074] In this embodiment, the method for calculating dense optical flow between eighteen target face regions of a vertex frame face image and eighteen target face regions of an end frame face image includes: Traversing each of the eighteen target face regions of the vertex frame face image; For each face region traversed, the dense optical flow between it and the corresponding region in the eighteen target face regions of the end frame face image is calculated to obtain an optical flow field; All the obtained optical flow fields are collected to obtain the dense optical flow between the vertex frame face image and the end frame face image (a total of 18 optical flow fields).
[0075] A corresponding relationship is established between the eighteen target face regions of the starting frame face image and the eighteen target face regions of the vertex frame face image according to the relative positions of the target face regions in the face.
[0076] Similarly, the eighteen target face regions of the vertex frame face image and the eighteen target face regions of the end frame face image are also correspondingly related according to their relative positions in the face.
[0077] Taking the starting frame face image and the vertex frame face image as an example: the left eye area among the eighteen target face areas of the starting frame face image corresponds to the left eye area among the eighteen target face areas of the vertex frame face image; the left eyelid area among the eighteen target face areas of the starting frame face image corresponds to the left eyelid area among the eighteen target face areas of the vertex frame face image; the right eyelid area among the eighteen target face areas of the starting frame face image corresponds to the right eyelid area among the eighteen target face areas of the vertex frame face image.
[0078] It can be understood that by calculating the dense optical flow between the 18 target facial regions of the starting frame face image and the vertex frame face image, a total of 18 optical flow fields can be obtained. By calculating the dense optical flow between the 18 target facial regions of the vertex frame face image and the 18 target facial regions of the ending frame face image, a total of 18 optical flow fields can also be obtained. That is, in the above step J2: the dense optical flow between the starting frame, the vertex frame, and the ending frame is extracted, and a total of 36 optical flow fields are obtained. Each optical flow field corresponds to a set of optical flow vectors, represented as a two-dimensional matrix.
[0079] That is, when a trained facial micro-expression recognition model is used, the dense optical flow input through the input layer is 36 optical flow fields. The preliminary features of the 36 optical flow fields are extracted through the second CNN layer to obtain the preliminary features of each of the 36 optical flow fields (corresponding to 36 numpy arrays). The preliminary features of the 36 optical flow fields are sent to the LSTM, and the corresponding time features of the preliminary features of the 36 optical flow fields can be extracted (corresponding to 36 tensor arrays).
[0080] In an optional embodiment, the method for obtaining the trained facial micro-expression recognition model comprises the following steps: Step S1, constructing a training set; Step S2, constructing a network framework for a facial micro-expression recognition model; Step S3: using the training set to train the network framework to obtain the trained facial micro-expression recognition model.
[0081] In an optional embodiment, in the above step S1, the construction of the training set includes the following steps Q1 to Q7.
[0082] Step Q1: Select the facial data of the seven preset micro-expression categories from the CASME facial dataset to construct a first dataset.
[0083] In this embodiment, each facial data in the first data set corresponds to a micro-expression sample in the CASME facial data set.
[0084] In this embodiment, each facial data in the first data set corresponds to one of the seven pre-set micro-expressions.
[0085] Each facial data in the first dataset consists of three frames of images: the starting frame, the apex frame, and the ending frame of the corresponding micro-expression.
[0086] Step Q2: normalize the facial data in the first data set to obtain normalized facial data.
[0087] For each face data after normalization, the sizes of the start frame, vertex frame, and end frame are all adjusted to 256 pixels × 256 pixels.
[0088] Step Q3: For each normalized face data, extract the dense optical flow between its three frames (i.e., its starting frame, vertex frame, and end frame) (the method for extracting the dense optical flow is the same as the method for extracting the dense optical flow between the starting frame, vertex frame, and end frame in step J).
[0089] Step Q4: setting micro-expression classification labels for the seven types of micro-expressions respectively.
[0090] Optionally, different integers are used to set the micro-expression classification labels of the seven types of micro-expressions.
[0091] Step Q5: For each normalized face data, take its vertex frame and the extracted dense optical flow between its three frames as a sample.
[0092] Step Q6: Use the pre-set micro-expression classification labels to annotate the micro-expression of each sample.
[0093] Step Q7: Collect all samples marked with micro-expression classification labels to obtain the training set.
[0094] In an optional embodiment, the seven pre-set micro-expressions are: anger, disgust, fear, happy, neutral, sad, and surprise.
[0095] In this embodiment, the micro-expression classification labels of anger, disgust, fear, happy, neutral, sad, and surprise are: "0", "1", "2", "3", "4", "5", and "6" respectively.
[0096] In an optional embodiment, step S3 uses the training set to train the network framework to obtain the trained facial micro-expression recognition model, specifically including: Step S31: Initialize the network framework to obtain an initialized network framework; Step S32: input the data in the training set into the initialized network framework for iterative training. In each iterative training, the cross entropy loss function is used to calculate the loss between the predicted result and the true label. Then, the model gradient is calculated by the back propagation algorithm based on the calculated loss, and the model parameters are updated using the Adam optimizer until the model converges or reaches a preset number of iterations. The training is completed to obtain the trained facial micro-expression recognition model.
[0097] Exemplarily, in step S31, the network framework is initialized, including initializing the CNN convolution kernels of the CNN layers (i.e., the first CNN layer and the second CNN layer), initializing the ViT embedding matrix of the ViT layer, initializing the LSTM weights and bias of the LSTM layer, and initializing the weight matrix and bias term of the fully connected layer.
[0098] The dense optical flow extraction algorithms mentioned in this specification may all be the Gunnar Farneback algorithm.
[0099] The facial micro-expression recognition model used in the present invention is a deep learning model.
[0100] This invention extracts dense optical flow from different facial regions (the eighteen regions listed above). Optical flow features can better reflect subtle changes in each facial region, which can be learned by deep learning models. Based on this learned model, this invention can recognize the seven types of micro-expressions mentioned above in any video containing a human face.
[0101] It should be noted that if a face image can be detected in a video, it means that a face exists in the video.
[0102] An embodiment of a facial micro-expression recognition method has been described in detail above. Based on the facial micro-expression recognition method described in the above embodiment, an embodiment of the present invention also provides a facial micro-expression recognition system corresponding to the method.
[0103] Please refer to Figure 2 , the system comprises: Acquisition module 201, used to acquire target video; An image extraction module 202 is configured to extract the first frame, the middle frame, and the last frame of the target video as the start frame, the top frame, and the end frame of the target video; An optical flow extraction module 203 is configured to extract dense optical flows between the start frame, the vertex frame, and the end frame; A micro-expression recognition module 204 is configured to input the dense optical flow and the vertex frame into a pre-trained facial micro-expression recognition model for calculation to obtain facial micro-expressions in the target video; The pre-trained facial micro-expression recognition model is a neural network model that can recognize seven pre-set micro-expression categories, including: Input layer, used for data input of the model; The CNN layer is used to extract features from the vertex frame and dense optical flow input by the input layer to obtain the first preliminary feature and the second preliminary feature; A deep feature extraction layer, configured to extract spatial features of the first preliminary features and temporal features of the second preliminary features; A feature fusion layer is used to fuse the spatial features and the temporal features to obtain fused features; A fully connected layer, configured to map the fused features into a low-dimensional classification space and output score vectors for the seven pre-set micro-expression categories; Softmax function layer, used to convert the output of the fully connected layer into a probability distribution and output it; The classification decision layer is used to select the micro-expression category with the highest probability as the model's prediction result based on the probability distribution output by the Softmax function layer, and output it.
[0104] The facial micro-expression recognition system of this embodiment is used to implement the aforementioned facial micro-expression recognition method. Therefore, the specific implementation method of the system can be seen in the embodiment part of the facial micro-expression recognition method in the previous text. Therefore, its specific implementation method can refer to the description of the corresponding embodiments of each part and will not be elaborated here.
[0105] Figure 3 This is a structural diagram of a terminal 300 provided in an embodiment of the present invention. The terminal 300 can be used to execute the method provided in an embodiment of the present invention.
[0106] The terminal 300 may include a processor 310, a memory 320, and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention. The server structure may be a bus structure or a star structure, and may include more or fewer components than shown, or may combine certain components or arrange the components differently.
[0107] Memory 320 can be used to store execution instructions of processor 310. Memory 320 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in memory 320 are executed by processor 310, terminal 300 can perform some or all of the steps in the above-described method embodiments.
[0108] The processor 310 is the control center of the storage terminal. It uses various interfaces and lines to connect various parts of the entire electronic terminal. It executes various functions of the electronic terminal and / or processes data by running or executing software programs and / or modules stored in the memory 320, and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 310 can only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.
[0109] The communication unit 330 is configured to establish a communication channel so that the storage terminal can communicate with other terminals, receive user data sent by other terminals, or send user data to other terminals.
[0110] The technical effects that can be achieved by this embodiment can be found in the description above and will not be repeated here.
[0111] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.
[0112] The present invention combines optical flow extraction with CNN, ViT, and LSTM models, which helps to capture subtle facial expressions, improve recognition accuracy, and ensure model generalization.
[0113] This invention combines video segmentation with optical flow extraction technology, which helps ensure detection stability and efficiency to a certain extent. In particular, for micro-expression recognition in real-time video, it helps dynamically annotate facial micro-expressions in videos, improves visualization, and facilitates rapid recognition of facial micro-expressions in videos.
[0114] When the video to be used for facial micro-expression recognition is a real-time video of students in school, it helps to dynamically detect changes in students' micro-expressions, thereby helping to promptly discover students' negative emotions, provide real-time warnings for students' mental health management, and then help to discover students' psychological problems early.
[0115] Although the present invention has been described in detail with reference to the accompanying drawings and in combination with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and such modifications or substitutions shall be within the scope of the present invention. Any person skilled in the art who is familiar with the present invention may easily conceive of changes or substitutions within the technical scope disclosed in the present invention, and such changes or substitutions shall be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A facial micro-expression recognition method, characterized in that: Methods include: Step L1, obtaining the target video; Step L2: extract the first frame, middle frame, and last frame images of the target video as its start frame, peak frame, and end frame; Step L3, extracting dense optical flow between the start frame, the vertex frame and the end frame; Step L4: input the dense optical flow and the vertex frame into a pre-trained facial micro-expression recognition model for calculation to obtain the facial micro-expressions in the target video; The pre-trained facial micro-expression recognition model is a neural network model that can recognize seven pre-set micro-expression categories, including: Input layer, used for data input of the model; The CNN layer is used to extract features from the vertex frame and dense optical flow input by the input layer to obtain the first preliminary feature and the second preliminary feature; A deep feature extraction layer, configured to extract spatial features of the first preliminary features and temporal features of the second preliminary features; A feature fusion layer is used to fuse the spatial features and the temporal features to obtain fused features; A fully connected layer, configured to map the fused features into a low-dimensional classification space and output score vectors for the seven pre-set micro-expression categories; Softmax function layer, used to convert the output of the fully connected layer into a probability distribution and output it; The classification decision layer is used to select the micro-expression category with the highest probability as the model's prediction result based on the probability distribution output by the Softmax function layer, and output it.
2. The facial micro-expression recognition method according to claim 1, wherein The method for obtaining the target video includes: Obtain the video to be used for facial micro-expression recognition; According to the preset reading granularity, the segments of the video to be subjected to facial micro-expression recognition are read one by one to obtain the target video; the value range of the preset reading granularity is: 1 second 2 seconds.
3. The facial micro-expression recognition method according to claim 2, wherein: Step L4 also includes: after obtaining the facial micro-expressions in the target video, marking the obtained facial micro-expressions on the corresponding video segment of the video to be subjected to facial micro-expression recognition.
4. The facial micro-expression recognition method according to claim 3, wherein: The method also includes: Detecting a face in the video to be subjected to facial micro-expression recognition, and drawing a face frame for the detected face; Step L4 also includes: after obtaining the facial micro-expressions in the target video, marking the obtained facial micro-expressions on the top of the face frame on the corresponding video segment in the video to be subjected to facial micro-expression recognition.
5. The facial micro-expression recognition method according to claim 1, wherein: Step L3 specifically includes: Detect whether the starting frame, the vertex frame, and the ending frame all contain facial images. If so, execute step J; if not, stop recognizing facial micro-expressions in the current target video; Step J: extracting dense optical flow between the start frame, the vertex frame and the end frame.
6. The facial micro-expression recognition method according to claim 5, wherein: In step J, dense optical flow between the start frame, vertex frame, and end frame is extracted, including: Step J1: For each frame image in the starting frame, the vertex frame, and the ending frame, the following steps are performed respectively to obtain the corresponding frame face image of each frame image: Convert the frame image into a grayscale image; Detect and extract all faces in a grayscale image: If the number of extracted face images is 1, the extracted face image is saved as the corresponding frame face image; If the number of extracted face images 2. Calculate the proportion of each extracted face image in the corresponding frame image, and then select the face image with the largest proportion and save it as the corresponding frame face image; The corresponding frame face images of the obtained starting frame, vertex frame and ending frame are the starting frame face image, vertex frame face image and ending frame face image respectively; Step J2: extract dense optical flows between the vertex frame face image and the start frame face image and between the vertex frame face image and the end frame face image, and obtain the dense optical flows to be extracted between the start frame, the vertex frame and the end frame.
7. The facial micro-expression recognition method according to claim 6, wherein: In step J2, the method for extracting dense optical flow between the vertex frame face image and the start frame face image and between the vertex frame face image and the end frame face image includes: Step 1: Perform the following processing on the face images in the starting frame face image, the vertex frame face image, and the ending frame face image, respectively, to obtain eighteen target face regions in the starting frame face image, the vertex frame face image, and the ending frame face image: Use the 68-point annotation method to annotate the face in the face image, and obtain 68-point annotations of the face in the face image; Based on the 68-point annotations, extract the areas corresponding to the left eye, right eye, mouth, nose, left eyelid, and right eyelid in the face image; Divide the extracted area corresponding to the left eye into three areas according to a preset first division method, divide the area corresponding to the right eye into three areas according to a preset second division method, divide the extracted area corresponding to the mouth into four areas according to a preset third division method, and divide the extracted area corresponding to the nose into a left area of the nose bridge and a right area of the nose bridge; Gather the extracted regions and the divided regions to obtain eighteen face regions corresponding to the face image; The obtained 18 face regions of the starting frame face image are the 18 target face regions of the starting frame face image; Step 2: Calculate the dense optical flow between the face image of the starting frame and the face image of the vertex frame, which is recorded as the first dense optical flow; Step 3: Calculate the horizontal component of the first dense optical flow And the component y1 in the vertical direction, then use Calculate the modulus m of the first dense optical flow; if , then the vertex frame face image is used as the target vertex frame face image, and the obtained 18 face regions of the vertex frame face image are used as the 18 target face regions of the vertex frame face image, and then step five is executed; if , then the vertex frame face image is translated horizontally , pan in the vertical direction , obtain the adjusted vertex frame face image, and translate the eighteen face regions of the vertex frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the adjusted vertex frame face image, and then execute step 4; M is a preset modulus length threshold; Step 4: Calculate the dense optical flow between the face image of the starting frame and the face image of the vertex frame after the latest adjustment, which is recorded as the second dense optical flow. Then calculate the horizontal component of the second dense optical flow. and the vertical component , and calculate the modulus of the second dense optical flow , : if , then the most recently adjusted vertex frame face image is used as the target vertex frame face image, and the eighteen face regions of the most recently adjusted vertex frame face image are used as the eighteen target face regions of the vertex frame face image, and then the process goes to step five to continue; if , then the face image of the vertex frame that has been adjusted the most recently is translated in the horizontal direction , pan in the vertical direction , get the next adjusted vertex frame face image, and translate the eighteen face areas of the latest adjusted vertex frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the next adjusted vertex frame face image, and then execute step 4 again; Step 5: Calculate the dense optical flow between the target vertex frame face image and the end frame face image, which is recorded as the third dense optical flow; Step 6: Calculate the horizontal component of the third dense optical flow and the component in the vertical direction , and then use Calculate the modulus of the third dense optical flow : if , the obtained 18 face regions of the end frame face image are used as the 18 target face regions of the end frame face image, and then step 8 is executed; if , then the end frame face image is translated horizontally , pan in the vertical direction , obtain the adjusted end frame face image, and translate the eighteen face regions of the obtained end frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the adjusted end frame face image, and then execute step seven; Step 7: Calculate the dense optical flow between the target vertex frame face image and the current most recently adjusted end frame face image, which is recorded as the fourth dense optical flow. Then calculate the horizontal component of the fourth dense optical flow. and the component in the vertical direction , and calculate the modulus of the fourth dense optical flow : if , then the eighteen face regions of the most recently adjusted ending frame face image are used as the eighteen target face regions of the ending frame face image, and then step eight is executed; if , then the face image of the end frame after the latest adjustment is horizontally shifted , pan in the vertical direction , get the next adjusted end frame face image, and translate the eighteen face areas of the latest adjusted end frame face image in the horizontal direction , pan in the vertical direction , obtain the eighteen face regions of the next adjusted end frame face image, and then execute step seven again; Step eight, calculating dense optical flows between the eighteen target face regions of the starting frame face image and the eighteen target face regions of the vertex frame face image, to obtain dense optical flows between the starting frame face image and the vertex frame face image; The dense optical flow between the eighteen target face regions of the vertex frame face image and the eighteen target face regions of the end frame face image is calculated to obtain the dense optical flow between the vertex frame face image and the end frame face image.
8. The method for recognizing facial micro-expressions according to any one of claims 1 to 7, characterized in that: The method for obtaining a trained facial micro-expression recognition model includes: Step S1, constructing a training set; Step S2, constructing a network framework for a facial micro-expression recognition model; Step S3, using the training set to train the network framework to obtain the trained facial micro-expression recognition model; In step S1, a training set is constructed, including: Selecting facial data of the seven pre-set micro-expressions from the CASME facial data set to construct a first data set; each facial data in the first data set corresponds to a micro-expression; each facial data in the first data set is composed of three frames of images: a start frame, a peak frame, and an end frame corresponding to the micro-expression; Normalizing the facial data in the first data set to obtain normalized facial data; For each normalized face data, the dense optical flow between the start frame, vertex frame and end frame is extracted; Setting micro-expression classification labels for the seven types of micro-expressions; For each normalized face data, its vertex frame and the extracted dense optical flow between its three frames are taken as a sample; Use pre-set micro-expression classification labels to annotate each sample with micro-expressions; The samples marked with micro-expression classification labels are collected to obtain the training set.
9. A facial micro-expression recognition system, characterized in that: The system includes: Acquisition module, used to acquire target video; An image extraction module is used to extract the first frame, middle frame and last frame images of the target video as the start frame, top frame and end frame of the target video; an optical flow extraction module, configured to extract dense optical flow between the start frame, the vertex frame, and the end frame; A micro-expression recognition module is used to input the dense optical flow and the vertex frame into a pre-trained facial micro-expression recognition model for calculation to obtain facial micro-expressions in the target video; The pre-trained facial micro-expression recognition model is a neural network model that can recognize seven pre-set micro-expression categories, including: Input layer, used for data input of the model; The CNN layer is used to extract features from the vertex frame and dense optical flow input by the input layer to obtain the first preliminary feature and the second preliminary feature; A deep feature extraction layer, configured to extract spatial features of the first preliminary features and temporal features of the second preliminary features; A feature fusion layer is used to fuse the spatial features and the temporal features to obtain fused features; A fully connected layer, configured to map the fused features into a low-dimensional classification space and output score vectors for the seven pre-set micro-expression categories; Softmax function layer, used to convert the output of the fully connected layer into a probability distribution and output it; The classification decision layer is used to select the micro-expression category with the highest probability as the model's prediction result based on the probability distribution output by the Softmax function layer, and output it.
10. A terminal, characterized in that: include: processor; A memory for storing execution instructions of the processor; The processor is configured to execute the facial micro-expression recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Cross-library micro-expression recognition method and device based on optical flow attention neural network
CN110516571A
Micro-expression detection method based on optical flow
CN111461021A