Micro-expression recognition ability training system based on action unit driven expression simulation

By establishing a correspondence between facial action units and muscle activity areas and generating simulated video frames, the problems in existing systems affected by facial lighting and posture are solved, and efficient training of micro-expression recognition capabilities is achieved.

CN119942618BActive Publication Date: 2025-10-10JIANGSU POLICE INST +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510076710.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-10-10
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing micro-expression recognition ability training system is affected by factors such as different facial lighting, interference from irrelevant facial muscles, and different head postures, resulting in poor training efficiency and results.

Method used

By establishing the correspondence between facial action units and muscle activity areas, simulated video frames are generated to highlight the facial muscle activity areas, eliminate unfavorable factors, and use simulated video frames for training.

Benefits of technology

The accuracy and efficiency of micro-expression recognition training are improved, and trainees can quickly capture the movements of facial muscle activity areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942618B_ABST
    Figure CN119942618B_ABST
Patent Text Reader

Abstract

The present application provides a kind of micro-expression recognition ability training system based on action unit driven expression simulation.The system detects the activity of different muscle activity regions in the face image in the original video frame, and uses the correlation between the activity of different muscle activity regions when the face makes micro-expression and facial action unit to reconstruct simulation video frame, thereby eliminating the factors in the original video frame that are not conducive to micro-expression recognition, highlighting the action of relevant facial muscle activity regions when the face makes micro-expression. Through the micro-expression ability recognition training system, the trainer can quickly and accurately capture the action of facial muscle activity regions when the face makes micro-expression, thereby improving the training effect of the micro-expression recognition ability training system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a micro-expression recognition ability training system based on action unit driven expression simulation. BACKGROUND

[0002] Micro-expression is an expression that occurs when people try to consciously or subconsciously hide their true feelings, and its duration is 1 / 25 to 1 / 5 seconds. Therefore, it is more sensitive to capture micro-expression than ordinary expression. The existing micro-expression recognition ability training system usually requires the trainee to repeatedly watch a picture sequence or a video clip of a specified type of micro-expression. Some systems also provide the type of micro-expression in the picture sequence or video clip in the form of a text narration. During the training process, the different face illuminations, irrelevant facial muscle interference, different head poses in the picture sequence or video clip, and the non-intuitive text description often affect the speed and accuracy of the trainee's capture of micro-expression features, reducing the efficiency and effectiveness of the entire training. Therefore, how to improve the efficiency and effectiveness of the entire micro-expression recognition ability training system has become a problem to be solved. SUMMARY

[0003] The present application aims to provide a micro-expression recognition ability training system based on action unit driven expression simulation. The system uses the correlation between micro-expression and facial action units to eliminate the adverse factors in the original face image, and reconstructs a simulation image that can intuitively represent the facial muscle action under micro-expression, thereby improving the training effect of the micro-expression recognition ability training system.

[0004] SUMMARY To achieve the above-mentioned purpose, the present application proposes the following technical solutions:

[0005] A micro-expression recognition ability training system based on action unit driven expression simulation, the system comprising:

[0006] A data acquisition module configured to acquire an original video frame, the original video frame being a real face image;

[0007] A face modeling module configured to identify a face region image from the original video frame, divide a muscle activity region in the face region image, and establish a corresponding relationship between the muscle activity region and a facial action unit, each facial action unit representing a facial action;

[0008] A feature extraction module configured to extract features of the face region image to obtain facial features;

[0009] a facial action unit detection module configured to detect the facial features using a pre-trained classification model to determine an activated facial action unit, wherein the activated facial action unit is used to indicate that a muscle activity area corresponding to the facial action unit has undergone a shape change specified by the corresponding facial action;

[0010] a simulated video generation module configured to highlight the muscle activity area corresponding to the activated facial action unit in the facial area image to obtain a simulated video frame;

[0011] The video playing module is configured to play the simulated video frame.

[0012] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the facial modeling module is specifically used to use a pre-trained face target detection model to perform facial area detection on the original video frame to obtain the facial area image.

[0013] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the facial modeling module is specifically used to detect feature points of the facial area image and divide the muscle activity areas in the facial area image according to the detected facial feature points.

[0014] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the feature extraction module is specifically used to perform feature point detection and local feature extraction on the facial area image, and construct the facial features based on the extracted facial feature points and local feature vectors.

[0015] Specifically, the feature extraction module is specifically used to perform oriented gradient histogram feature extraction on the facial area image to obtain the local features.

[0016] Specifically, the feature extraction module is specifically used to use a constrained local neural field algorithm to perform feature point detection on the facial area image to obtain the facial feature points.

[0017] As an optional embodiment of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the facial modeling module is further configured to define a relationship between the color value of the muscle activity area and the intensity value of the facial action unit, and a relationship between the transparency of the muscle activity area and the intensity value of the facial action unit;

[0018] The facial action unit detection module is further configured to quantitatively describe the facial features using a pre-trained regression model to obtain an intensity value of the facial action unit;

[0019] The analog video generation module is further configured to adjust, in the face region image, a color value and a transparency of a muscle activity region corresponding to the activated facial action unit based on the intensity value of the facial action unit, so as to highlight the corresponding muscle activity region.

[0020] Specifically, the regression model is obtained by pre-training a support vector machine with a linear kernel.

[0021] As an optional implementation of the micro-expression recognition ability training system based on the action unit driven expression simulation, the classification model is obtained by pre-training a support vector machine with a linear kernel.

[0022] As an optional implementation of the micro-expression recognition ability training system based on the action unit driven expression simulation, the video playing module is specifically configured to simultaneously play the analog video frame and the original video frame of the analog video frame in the same picture.

[0023] Beneficial effects: the embodiment of the specification provides a micro-expression recognition ability training system based on action unit driven expression simulation, which detects the activity of different muscle activity regions in the face image in the original video frame, and uses the correlation between the activity of different muscle activity regions when a person makes a micro-expression and the facial action unit to reconstruct an analog video frame, thereby eliminating factors that are not conducive to micro-expression recognition in the original video frame, and highlighting the relevant facial muscle activity region action when a person makes a micro-expression. Through the system for micro-expression ability recognition training, the trainer can quickly and accurately capture the action of the facial muscle activity region when a person makes a micro-expression, thereby improving the training effect of the micro-expression recognition ability training system. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the specification or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the specification, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0025] Figure 1 It is a micro-expression display page screenshot of the existing micro-expression recognition ability training system.

[0026] Figure 2 It is a structure diagram of a micro-expression recognition ability training system based on action unit driven expression simulation according to an embodiment.

[0027] Figure 3 Schematic diagram of the correspondence between micro-expressions and action units involved in an embodiment.

[0028] Figure 4 Schematic diagram of the position distribution of 68 feature points obtained after detecting a facial area image using a constrained local neural field algorithm according to an embodiment.

[0029] Figure 5 A schematic diagram of a sample division of muscle activity areas involved in an embodiment.

[0030] Figure 6 Schematic diagram of a micro-expression recognition ability training screen involved in an embodiment. DETAILED DESCRIPTION

[0031] First, it should be noted that the terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used in the embodiments of the present invention and the appended claims, the singular forms "a," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise.

[0032] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0033] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0034] like Figure 1The figure shows a screenshot of the micro-expression display page of an existing micro-expression recognition ability training system. The trainee observes the facial image in the micro-expression display page and selects the micro-expression option on the right. Existing micro-expression recognition ability training systems usually require trainees to repeatedly watch a sequence of pictures or video clips of a specified type of micro-expression. Some systems also provide text narration to describe the type of micro-expression in the picture sequence or video clip. Since these micro-expression picture sequences or video clips are often real images, these picture sequences or video clips inevitably have unfavorable factors such as different facial lighting, interference from irrelevant facial muscles, different head postures, and non-intuitive text descriptions. These unfavorable factors will affect the speed and accuracy of the trainee's capture of micro-expression features, resulting in poor training results of the existing micro-expression recognition ability training system.

[0035] In view of this, this embodiment provides a micro-expression recognition ability training system based on action unit driven expression simulation.

[0036] The micro-expression recognition ability training system based on action unit driven expression simulation described in one or more embodiments of this specification will be further described in detail below in combination with the drawings and specific embodiments of the specification, but this detailed description does not constitute a limitation to the embodiments of this specification.

[0037] Please refer to Figure 2 , Figure 2 This is a structural diagram of a micro-expression recognition ability training system based on action unit driven expression simulation proposed in one or more embodiments of this specification. Figure 2 As shown, the system includes:

[0038] The data acquisition module 201 is configured to acquire an original video frame, which is a real face image.

[0039] The facial modeling module 202 is configured to identify a facial region image from an original video frame, divide the muscle activity regions in the facial region image, and establish a correspondence between the muscle activity regions and facial action units, where each facial action unit represents a facial action.

[0040] The feature extraction module 203 is configured to extract features from the facial region image to obtain facial features.

[0041] The facial action unit detection module 204 is configured to detect facial features using a pre-trained classification model to determine an activated facial action unit. The activated facial action unit is used to indicate that the muscle activity area corresponding to the facial action unit has undergone a shape change specified by the corresponding facial action.

[0042] The simulated video generation module 205 is configured to highlight the muscle activity area corresponding to the activated facial action unit in the facial area image to obtain a simulated video frame.

[0043] The video playing module 206 is configured to play the simulated video frames.

[0044] Below, the working principles of each module in the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation will be explained in combination with specific implementation scenarios.

[0045] First, with respect to the data acquisition module 201, the module can collect the original video stream frame by frame to obtain the original video frame. It should be noted that the original video stream is composed of a group of real face images.

[0046] Next, with respect to the facial modeling module 202, the module is used to identify the facial region image from the original video frame, then divide the facial region image into different muscle activity regions, and establish a correspondence between each muscle activity region and a facial action unit.

[0047] In some embodiments, the facial modeling module 202 may use a pre-trained face object detection model to perform facial region detection on the original video frame to obtain a facial region image.

[0048] Specifically, the aforementioned face detection model can be implemented using a multi-task cascaded convolutional neural network (MTCNN). MTCNN is a deep learning algorithm for face detection and landmark location. It consists of three cascaded convolutional neural networks: the Proposal Network (P-Net), the Refine Network (R-Net), and the Output Network (O-Net), each responsible for a different task. MTCNN is designed to achieve efficient and accurate face detection and simultaneously output the position of the face bounding box, pose estimation, and key landmarks (such as the coordinates of the eyes, nose, and mouth).

[0049] Before deploying the face target detection model, the MTCNN model needs to be trained based on the face target detection task. The training process is as follows:

[0050] Collect face image samples, mark the facial area position frame, key feature point positions, and posture information in the face image samples as training labels.

[0051] The face image sample is input into the MTCNN model to obtain the predicted facial area detection frame, key feature point location and posture information.

[0052] A loss function is constructed based on the predicted values ​​and the corresponding training labels. This loss function is used to update the parameters of the MTCNN model until a facial object detection model that meets the requirements is obtained. Using this facial object detection model, facial regions can be detected from the original video frames, resulting in facial region images.

[0053] It should be noted that the above-mentioned face target detection model can also be implemented using other detection algorithms / models, and this embodiment does not limit this.

[0054] After the facial modeling module 202 uses the face target detection model to detect the facial area of ​​the original video frame, it can also preprocess the facial area image, such as position correction, brightness, clarity adjustment, etc. The specific preprocessing method can be selected according to needs, and this embodiment does not limit this.

[0055] In some implementations, the facial modeling module 202 may perform feature point detection on the facial region image, and then divide the muscle activity areas in the facial region image according to the detected facial feature points.

[0056] Specifically, the facial modeling module 202 can use the Constrained Local Neural Fields (CLNF) algorithm to accurately detect 68 feature points in facial region images. The CLNF algorithm is a variation of the Constrained Local Model (CLM) for facial feature point detection and tracking. CLNF utilizes more advanced patch experts and optimization functions to achieve more accurate facial feature point detection. Due to its high accuracy and robustness to interference, the CLNF algorithm is preferred for feature point detection in facial region images in this embodiment.

[0057] Please refer to Figure 4 , Figure 4 The figure shows 68 feature points obtained after detecting the facial region image using the constrained local neural field algorithm. The facial modeling module 202 can define the facial muscle activity area based on the 68 facial feature points. Figure 5 , Figure 5 An example of muscle activity area division is shown, in which the line area connected by three points (x

[27] , y

[27] ), (x

[39] , y

[39] ), and (x

[21] , y

[21] ) represents the medial part of the left frontalis muscle (e.g. Figure 5 As shown in the red area), (x

[39] ,y

[39] ), (x

[40] ,y

[40] ), (x

[41] ,y

[41] ), (x

[36] ,y

[36] ) and their combination form the lateral side of the left orbicularis oculi muscle (as shown in the red area). Figure 5As shown in the green area, the three points (x

[48] ,y

[48] ), (x[3],y[3]), and (x[2],y[2]) represent the left zygomatic muscle (as shown in the green area). Figure 5 shown in yellow area).

[0058] It should be noted that the detection algorithm used by the facial modeling module 202 to detect feature points on the facial region image is not limited to the above-mentioned CLNF algorithm. The selected detection algorithm can be set according to requirements, and this embodiment does not impose any restrictions on this.

[0059] After dividing the muscle activity areas in the facial area image, the facial modeling module 202 can establish the correspondence between the facial action units and the muscle activity areas. The above-mentioned facial action unit refers to a specific facial action, such as cheek lifting, lip corner pulling, etc. When a face makes a micro-expression, it will make at least one facial action, and each facial action is completed by one or more muscle activity areas. Based on this, the facial modeling module 202 can construct the above-mentioned facial action units according to the facial actions corresponding to different micro-expressions. Please refer to Figure 3 , Figure 3 A schematic diagram showing the correspondence between micro-expressions and action units is shown. Figure 3 Figure 2 shows seven facial micro-expressions (happiness, anger, sadness, surprise, fear, disgust, and contempt) and their associated facial action units. Based on the correspondence between micro-expressions and facial action units, the aforementioned facial action units can be constructed. After constructing the facial action units, the facial modeling module 202 can establish a correspondence between the facial action units and muscle activity regions based on the muscle activity region segmentation results of the facial region image. For example, the muscle activity regions associated with facial action unit AU6 (orbicularis oculi contraction) are primarily the left and right lateral orbicularis oculi muscles, while the muscle activity regions associated with AU12 (upper lip corner elevation) are primarily the left zygomaticus and right lateral zygomaticus muscles.

[0060] Next, with respect to the above-mentioned feature extraction module 203, the feature extraction module 203 can be specifically used to perform feature point detection and local feature extraction on the facial area image, and construct facial features based on the extracted facial feature points and local feature vectors. For the detection of facial feature points, the above-mentioned constrained local neural field algorithm can be specifically used to detect and obtain 68 feature points of the facial area image. For local feature extraction, the histogram of oriented gradients (HOG) feature can be used to describe it. For example, 2x2 cell blocks of 8x8 pixels can be used to extract local features. The HOG feature dimension is large, and the principal component analysis (PCA) method can be used to reduce the dimension. For the construction of facial features, the facial feature points and local feature vectors can be directly used as facial features, or the facial feature points and local feature vectors can be fused and the fused features can be used as facial features.

[0061] The facial action unit detection module 204 can utilize a pre-trained classification model to detect facial features to determine activated facial action units. This classification model can be implemented using a support vector machine (SVM) with a linear kernel. Specifically, a single SVM can be used to perform multi-classification tasks for different facial action units. Specifically, the facial features are input into the SVM, which directly outputs the category corresponding to the facial features. This category is used to characterize the activated facial action units. The classification model can also be implemented using multiple SVMs. Specifically, a binary classification model can be trained for each facial action unit. This binary classification model determines whether the corresponding facial action unit is activated based on the input facial features. An activated facial action unit indicates that the muscle activity area corresponding to the facial action unit has undergone a shape change specified by the corresponding facial action. Conversely, an inactivated facial action unit indicates that the muscle activity area corresponding to the facial action unit has not undergone a shape change specified by the corresponding facial action.

[0062] Regarding the above-mentioned simulated video generation module 205, this module is specifically used to highlight the muscle activity areas corresponding to the activated facial action units in the facial area image, such as adjusting the color, brightness, transparency, etc. of these muscle activity areas, so that these muscle activity areas are visually distinguished from the muscle activity areas corresponding to the inactivated facial action units, thereby obtaining a simulated video frame.

[0063] Regarding the above-mentioned video playing module 206, this module is mainly used to play the simulated video frames. For example, the video playing module 206 can play the simulated video frames separately, that is, the original video frames are not used during training, but the simulated video frames are used to replace the original video frames for training. The video playing module 206 can also play the original video frames and the simulated video frames of the original video frames in one picture at the same time, such as Figure 6As shown, during training, the original video frames are played on the left side of the screen, and the simulated video frames are played synchronously on the right side of the screen. By comparing the real pictures with the simulated video frames, the trainee can quickly capture the changes in micro-expressions and their key features, thereby improving the accuracy and efficiency of micro-expression recognition training.

[0064] In some more specific embodiments, the facial modeling module 202 may further define a relationship between the color value of the muscle activity area and the intensity value of the facial action unit, and a relationship between the transparency of the muscle activity area and the intensity value of the facial action unit.

[0065] Correspondingly, the facial action unit detection module 204 can also be used to quantitatively describe facial features using a pre-trained regression model to obtain intensity values ​​of facial action units.

[0066] The above-mentioned regression model is obtained by pre-training a support vector machine (SVM) with a linear kernel. During regression model training, the collected facial feature samples are scored for the intensity values ​​of different facial action units (FAUs). A higher intensity value indicates a higher correlation between the facial feature sample and the corresponding FAU. These facial feature samples are then used to perform regression training on the SVM, enabling the SVM to fit the relationship between the facial features and the intensity values ​​of different FAUs, thereby obtaining the above-mentioned regression model. This regression model can then predict the intensity values ​​of FAUs based on the input facial features.

[0067] Correspondingly, the above-mentioned simulated video generation module 205 is also specifically used to adjust the color value and transparency of the muscle activity area corresponding to the activated facial action unit in the facial area image based on the intensity value of the facial action unit, and display the muscle activity area with color intensity to achieve highlighting of the corresponding muscle activity area.

[0068] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0069] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.

Claims

1. A micro-expression recognition ability training system based on action unit driven expression simulation, characterized by: include: A data acquisition module is configured to acquire an original video frame, wherein the original video frame is a real face image; a facial modeling module configured to identify a facial region image from the original video frame, divide muscle activity regions in the facial region image, and establish a correspondence between the muscle activity regions and facial action units, each facial action unit representing a facial action; the facial modeling module is further configured to define a relationship between a color value of the muscle activity region and an intensity value of the facial action unit, and a relationship between a transparency of the muscle activity region and the intensity value of the facial action unit; A feature extraction module is configured to extract features from the facial region image to obtain facial features; A facial action unit detection module is configured to detect the facial features using a pre-trained classification model to determine an activated facial action unit, wherein the activated facial action unit is used to indicate that a muscle activity area corresponding to the facial action unit has undergone a shape change specified by the corresponding facial action; the facial action unit detection module is further configured to quantitatively describe the facial features using a pre-trained regression model to obtain an intensity value of the facial action unit; a simulated video generation module configured to adjust, in the facial region image, the color value and transparency of the muscle activity region corresponding to the activated facial action unit based on the intensity value of the facial action unit, so as to highlight the corresponding muscle activity region and obtain a simulated video frame; The video playing module is configured to play the simulated video frame.

2. The system according to claim 1, wherein: The facial modeling module is specifically used to perform facial region detection on the original video frame using a pre-trained face target detection model to obtain the facial region image.

3. The system according to claim 1, wherein: The facial modeling module is specifically used to detect feature points on the facial region image and divide the muscle activity areas in the facial region image according to the detected facial feature points.

4. The system according to claim 1, wherein: The feature extraction module is specifically used to perform feature point detection and local feature extraction on the facial region image, and construct the facial features based on the extracted facial feature points and local feature vectors.

5. The system according to claim 4, characterized in that The feature extraction module is specifically used to perform directional gradient histogram feature extraction on the facial area image to obtain the local features.

6. The system according to claim 4, characterized in that The feature extraction module is specifically used to use a constrained local neural field algorithm to perform feature point detection on the facial area image to obtain the facial feature points.

7. The system according to claim 1, wherein: The regression model is obtained by pre-training a support vector machine with a linear kernel.

8. The system according to claim 1, wherein: The classification model is obtained by pre-training a support vector machine with a linear kernel.

9. The system according to claim 1, wherein: The video playback module is specifically configured to simultaneously play the simulated video frame and the original video frame of the simulated video frame in the same picture.

Citation Information

Patent Citations

  • Facial expression interaction method, interaction device and computer storage medium

    CN111638784A

  • Depression recognition system based on micro-expression analysis

    CN112232191A

  • Awareness detection method and equipment based on olfactory stimulation and facial expression, and medium

    CN117633606A