Micro-expression recognition ability training system based on action unit driving expression simulation

Through the micro-expression recognition ability training system based on action unit-driven expression simulation, the training efficiency and effect reduction in the existing system due to facial light, muscle interference, different postures and unintuitive text descriptions is solved, and a faster and more accurate micro-expression feature capture and training effect is achieved.

CN119942618AActive Publication Date: 2025-05-06JIANGSU POLICE INST +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510076710.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing micro-expression recognition ability training system has reduced the speed and accuracy of the trainer to capture micro-expression features due to different facial light, facial-independent muscle interference, different head postures, and unintuitive text descriptions, reducing the efficiency and effect of the entire training.

Method used

A micro-expression recognition ability training system based on action unit-driven expression simulation is adopted. The system obtains the original video frame through the data acquisition module. The facial modeling module recognizes the facial area and divides the muscle activity area, establishes the correspondence between the facial action unit and the muscle activity area. The feature extraction module extracts facial features, the facial action unit detection module detects the activated facial action unit, and the simulation video generation module highlights the muscle activity area corresponding to the activated facial action unit, and generates simulated video frames.

Benefits of technology

By eliminating the unfavorable factors in the original video frame and reconstructing intuitive simulated video frames, trainers can quickly and accurately capture the movements of facial muscles when faces make micro-expressions, thereby improving the training effect of the micro-expression recognition ability training system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942618A_ABST
    Figure CN119942618A_ABST
Patent Text Reader

Abstract

The invention provides a micro-expression recognition ability training system based on action unit driven expression simulation. The system detects activity conditions of different muscle activity areas in a face image in an original video frame, and reconstructs an analog video frame by using an association relationship between the activity conditions of different muscle activity areas and a facial action unit when a face makes a micro-expression, thereby eliminating factors which are not beneficial to micro-expression recognition in the original video frame, and improving the recognition accuracy of the micro-expression. And the related action condition of the facial muscle activity area when the human face makes the micro-expression is highlighted. Micro-expression ability recognition training is carried out through the system, a trainee can quickly and accurately capture the action condition of a facial muscle activity area when a human face makes a micro-expression, and therefore the training effect of the micro-expression recognition ability training system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a micro-expression recognition ability training system based on action unit driven expression simulation. Background Art

[0002] Micro-expressions occur when people try to consciously or subconsciously hide their true feelings, and last for 1 / 25 to 1 / 5 seconds. Therefore, capturing micro-expressions requires higher sensitivity than ordinary expressions. Existing micro-expression recognition training systems usually require trainees to repeatedly watch a specified type of micro-expression picture sequence or video clip. Some systems also describe the type of micro-expression in the picture sequence or video clip in the form of text narration. During the training process, due to different facial lighting in the picture sequence or video clip, irrelevant facial muscle interference, different head postures, and unintuitive text descriptions, the speed and accuracy of trainees in capturing micro-expression features are often affected, reducing the efficiency and effectiveness of the entire training. Therefore, how to improve the efficiency and effectiveness of the training of the entire micro-expression recognition training system has become an urgent problem to be solved. Summary of the invention

[0003] Purpose of the invention: The present invention aims to propose a micro-expression recognition ability training system based on action unit driven expression simulation. The system utilizes the correlation between micro-expressions and facial action units, eliminates the unfavorable factors in the original face image, and reconstructs a simulated image that can intuitively show the facial muscle movements under micro-expressions, thereby improving the training effect of the micro-expression recognition ability training system.

[0004] Invention content: To achieve the above objectives, the present invention proposes the following technical solutions: A micro-expression recognition ability training system based on action unit driven expression simulation, the system comprising: A data acquisition module is configured to acquire an original video frame, wherein the original video frame is a real face image; A facial modeling module, configured to identify a facial region image from the original video frame, divide muscle activity regions in the facial region image, and establish a correspondence between the muscle activity regions and facial action units, each facial action unit representing a facial action; A feature extraction module, configured to extract features from the facial region image to obtain facial features; A facial action unit detection module is configured to detect the facial features using a pre-trained classification model to determine an activated facial action unit, wherein the activated facial action unit is used to indicate that a muscle activity area corresponding to the facial action unit undergoes a shape change specified by a corresponding facial action; a simulated video generation module configured to highlight the muscle activity area corresponding to the activated facial action unit in the facial area image to obtain a simulated video frame; The video playing module is configured to play the simulated video frame.

[0005] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the facial modeling module is specifically used to use a pre-trained face target detection model to perform facial area detection on the original video frame to obtain the facial area image.

[0006] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the facial modeling module is specifically used to detect feature points of the facial area image and divide the muscle activity areas in the facial area image according to the detected facial feature points.

[0007] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the feature extraction module is specifically used to perform feature point detection and local feature extraction on the facial area image, and construct the facial features based on the extracted facial feature points and local feature vectors.

[0008] Specifically, the feature extraction module is specifically used to extract directional gradient histogram features from the facial area image to obtain the local features.

[0009] Specifically, the feature extraction module is specifically used to use a constrained local neural field algorithm to perform feature point detection on the facial area image to obtain the facial feature points.

[0010] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the facial modeling module is further used to define the relationship between the color value of the muscle activity area and the intensity value of the facial action unit, and the relationship between the transparency of the muscle activity area and the intensity value of the facial action unit; The facial action unit detection module is further specifically used to quantitatively describe the facial features using a pre-trained regression model to obtain the intensity value of the facial action unit; The simulated video generation module is specifically used to adjust the color value and transparency of the muscle activity area corresponding to the activated facial action unit in the facial area image based on the intensity value of the facial action unit, so as to highlight the corresponding muscle activity area.

[0011] Specifically, the regression model is obtained by pre-training a support vector machine with a linear kernel.

[0012] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the classification model is obtained by pre-training a support vector machine with a linear kernel.

[0013] As an optional implementation of the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation, the video playback module is specifically used to simultaneously play the simulated video frame and the original video frame of the simulated video frame in the same screen.

[0014] Beneficial effects: The embodiment of this specification provides a micro-expression recognition ability training system based on action unit driven expression simulation, which detects the activity of different muscle activity areas in the face image in the original video frame, and uses the correlation between the activity of different muscle activity areas and facial action units when the face makes micro-expressions to reconstruct the simulated video frame, thereby eliminating the factors in the original video frame that are not conducive to micro-expression recognition, and highlighting the movement of the facial muscle activity areas related to the face when making micro-expressions. Through the system for micro-expression ability recognition training, the trainee can quickly and accurately capture the movement of the facial muscle activity areas when the face makes micro-expressions, thereby improving the training effect of the micro-expression recognition ability training system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 A screenshot of the micro-expression display page of the existing micro-expression recognition training system.

[0017] Figure 2 The present invention is a structural schematic diagram of a micro-expression recognition ability training system based on action unit driven expression simulation according to an embodiment.

[0018] Figure 3 The figure is a schematic diagram of the correspondence between a micro-expression and an action unit involved in an embodiment.

[0019] Figure 4 The figure is a schematic diagram of the position distribution of 68 feature points obtained after detecting a facial area image using a constrained local neural field algorithm according to an embodiment.

[0020] Figure 5 A schematic diagram of a sample of muscle activity area division involved in an embodiment.

[0021] Figure 6 A schematic diagram of a micro-expression recognition ability training screen involved in an embodiment. DETAILED DESCRIPTION

[0022] First of all, it should be noted that the terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise.

[0023] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0024] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0025] like Figure 1 As shown, it is a screenshot of the micro-expression display page of the existing micro-expression recognition ability training system. The trainee observes the face image in the micro-expression display page and selects the micro-expression option on the right. The existing micro-expression recognition ability training system usually asks the trainee to repeatedly watch the picture sequence or video clip of the specified type of micro-expression. Some systems also describe the type of micro-expression in the picture sequence or video clip in the form of text narration. Since these micro-expression picture sequences or video clips are often real images, these picture sequences or video clips will inevitably have unfavorable factors such as different facial lighting, irrelevant facial muscle interference, different head postures, and unintuitive text descriptions. These unfavorable factors will affect the speed and accuracy of the trainee's capture of micro-expression features, resulting in poor training results of the existing micro-expression recognition ability training system.

[0026] In view of this, the present embodiment provides a micro-expression recognition ability training system based on action unit driven expression simulation.

[0027] The micro-expression recognition ability training system based on action unit driven expression simulation described in one or more embodiments of this specification will be further described in detail below in conjunction with the drawings and specific embodiments of the specification, but this detailed description does not constitute a limitation on the embodiments of this specification.

[0028] Please refer to Figure 2 , Figure 2 Schematic diagram of a micro-expression recognition ability training system based on action unit driven expression simulation proposed in one or more embodiments of this specification. Figure 2 As shown, the system includes: The data acquisition module 201 is configured to acquire an original video frame, which is a real face image.

[0029] The facial modeling module 202 is configured to identify a facial region image from an original video frame, divide the muscle activity regions in the facial region image, and establish a correspondence between the muscle activity regions and facial action units, each facial action unit representing a facial action.

[0030] The feature extraction module 203 is configured to extract features from the facial region image to obtain facial features.

[0031] The facial action unit detection module 204 is configured to detect facial features using a pre-trained classification model to determine an activated facial action unit. The activated facial action unit is used to indicate that a muscle activity area corresponding to the facial action unit has undergone a shape change specified by the corresponding facial action.

[0032] The simulated video generation module 205 is configured to highlight the muscle activity area corresponding to the activated facial action unit in the facial area image to obtain a simulated video frame.

[0033] The video playing module 206 is configured to play the simulated video frames.

[0034] Next, the working principles of each module in the above-mentioned micro-expression recognition ability training system based on action unit driven expression simulation will be explained in combination with specific implementation scenarios.

[0035] First, with respect to the data acquisition module 201, the module can collect the original video stream frame by frame to obtain the original video frame. It should be noted that the original video stream is composed of a group of real face images.

[0036] Next, with regard to the facial modeling module 202, the module is used to identify the facial region image from the original video frame, then divide the facial region image into different muscle activity regions, and establish a corresponding relationship between each muscle activity region and a facial action unit.

[0037] In some implementations, the facial modeling module 202 may use a pre-trained facial object detection model to perform facial region detection on the original video frame to obtain a facial region image.

[0038] Specifically, the above-mentioned face target detection model can be implemented using a multi-task cascaded convolutional neural network (MTCNN), which is a deep learning algorithm for face detection and feature point location. It consists of three cascaded convolutional neural networks, namely Proposal Network (P-Net), Refine Network (R-Net) and OutputNetwork (O-Net), each of which is responsible for different tasks. The design goal of MTCNN is to achieve efficient and accurate face detection, and simultaneously output the position of the face frame, posture estimation, and key feature points (such as the coordinates of the eyes, nose, and mouth).

[0039] Before deploying the face target detection model, the MTCNN model needs to be trained based on the face target detection task. The training process is as follows: Collect face image samples, mark the facial area location frame, key feature point locations, and add posture information in the face image samples as training labels.

[0040] The face image samples are input into the MTCNN model to obtain the predicted facial area detection frame, key feature point locations and posture information.

[0041] A loss function is constructed based on the predicted value and the corresponding training label, and the loss function is used to update the parameters of the MTCNN model until a face target detection model that meets the requirements is obtained. Using this face target detection model, the face area can be detected from the original video frame to obtain a face area image.

[0042] It should be noted that the above-mentioned face target detection model can also be implemented using other detection algorithms / models, and this embodiment does not limit this.

[0043] After using the face target detection model to perform facial area detection on the original video frame, the above-mentioned facial modeling module 202 can also pre-process the facial area image, such as position correction, brightness, clarity adjustment, etc. The specific pre-processing method can be selected according to needs, and this embodiment does not limit this.

[0044] In some implementations, the facial modeling module 202 may perform feature point detection on the facial region image, and then divide the muscle activity regions in the facial region image according to the detected facial feature points.

[0045] Specifically, the facial modeling module 202 can use the Constrained Local Neural Fields (CLNF) algorithm to accurately detect 68 feature points of the facial region image. The CLNF algorithm is a technology for facial feature point detection and tracking, which is a variant of the Constrained Local Model (CLM). CLNF uses more advanced patch experts and optimization functions to achieve more accurate facial feature point detection. Due to the high accuracy and strong interference resistance of the algorithm, the present embodiment preferably uses the CLNF algorithm to perform feature point detection on the facial region image.

[0046] Please refer to Figure 4 , Figure 4 The figure shows 68 feature points obtained after detecting the facial region image using the constrained local neural field algorithm. The facial modeling module 202 can define the facial muscle activity area based on the 68 facial feature points. Figure 5 , Figure 5 An example of muscle activity area division is shown, in which the line area connected by three points (x

[27] , y

[27] ), (x

[39] , y

[39] ), and (x

[21] , y

[21] ) represents the medial side of the left frontalis muscle (e.g. Figure 5 The red area shows that (x

[39] , y

[39] ), (x

[40] , y

[40] ), (x

[41] , y

[41] ), (x

[36] , y

[36] ) and their combination form the lateral side of the left orbicularis oculi muscle (as shown in the red area). Figure 5 As shown in the green area, the three points (x

[48] , y

[48] ), (x[3], y[3]), and (x[2], y[2]) represent the left zygomatic muscle (as shown in Figure 5 yellow area).

[0047] It should be noted that the detection algorithm used by the facial modeling module 202 to detect feature points on the facial region image is not limited to the above-mentioned CLNF algorithm. The selected detection algorithm can be set according to requirements, and this embodiment does not impose any limitation on this.

[0048] After dividing the muscle activity areas in the facial area image, the facial modeling module 202 can establish the correspondence between the facial action units and the muscle activity areas. The above-mentioned facial action unit refers to a specific facial action, such as cheek lifting, lip corner pulling, etc. When a face makes a micro-expression, it will make at least one facial action, and each facial action is completed by one or more muscle activity areas. Based on this, the facial modeling module 202 can construct the above-mentioned facial action units according to the facial actions corresponding to different micro-expressions. Please refer to Figure 3 , Figure 3 A schematic diagram showing the correspondence between micro-expressions and action units is shown. Figure 3 , 7 kinds of facial micro-expressions (happiness, anger, sadness, surprise, fear, disgust, contempt) and facial action units related to these micro-expressions are shown. Based on the correspondence between micro-expressions and facial action units, the above-mentioned facial action units can be constructed. After constructing the facial action units, the facial modeling module 202 can establish the correspondence between the facial action units and the muscle activity areas according to the muscle activity area division result of the above-mentioned facial area image. For example, the muscle activity areas mainly corresponding to the facial action unit AU6 (orbicularis oculi contraction) are the outer side of the left orbicularis oculi and the outer side of the right orbicularis oculi, and the muscle activity areas mainly corresponding to AU12 (upper lip corner lifting) are the left zygomatic muscle and the left side of the right side.

[0049] Next, with respect to the above-mentioned feature extraction module 203, the feature extraction module 203 can be specifically used to perform feature point detection and local feature extraction on the facial area image, and construct facial features based on the extracted facial feature points and local feature vectors. For the detection of facial feature points, the above-mentioned constrained local neural field algorithm can be specifically used to detect 68 feature points of the facial area image. For local feature extraction, the oriented gradient histogram (HOG) feature can be used to describe it. For example, 2x2 cell blocks of 8x8 pixels can be used to extract local features. The HOG feature dimension is large, and the principal component analysis (PCA) method can be used to reduce the dimension. For the construction of facial features, the facial feature points and local feature vectors can be directly used as facial features, or the facial feature points and local feature vectors can be feature fused, and the fused features can be used as facial features.

[0050] For the facial action unit detection module 204, the module can use a pre-trained classification model to detect facial features to determine the activated facial action unit. The above classification model can be implemented by a support vector machine with a linear kernel. Specifically, a multi-classification task of different facial action units can be implemented by a support vector machine, that is, the above facial features are input into the support vector machine, and the support vector machine directly outputs the category corresponding to the facial features, and the category is used to characterize the activated facial action unit. The above classification model can also be implemented by multiple support vector machines, that is, a binary classification model can be trained for each facial action unit, and this binary classification model determines whether the corresponding facial action unit is activated based on the input facial features. The above activated facial action unit is used to characterize that the muscle activity area corresponding to the facial action unit has undergone a shape change specified by the corresponding facial action, and conversely, the inactivated facial action unit characterizes that the muscle activity area corresponding to the facial action unit has not undergone a shape change specified by the corresponding facial action.

[0051] Regarding the above-mentioned simulated video generation module 205, the module is specifically used to highlight the muscle activity areas corresponding to the activated facial action units in the facial area image, such as adjusting the color, brightness, transparency, etc. of these muscle activity areas, so that these muscle activity areas are visually distinguished from the muscle activity areas corresponding to the inactivated facial action units, thereby obtaining a simulated video frame.

[0052] Regarding the above-mentioned video playback module 206, this module is mainly used to play the simulated video frame. For example, the video playback module 206 can play the simulated video frame alone, that is, the original video frame is not used during training, but the simulated video frame is used to replace the original video frame for training. The video playback module 206 can also play the original video frame and the simulated video frame of the original video frame in one picture at the same time, such as Figure 6 As shown, during training, the original video frames are played on the left side of the screen, and the simulated video frames are played synchronously on the right side of the screen. By comparing the real pictures with the simulated video frames, the trainee can quickly capture the changes in micro-expressions and their key features, thereby improving the accuracy and efficiency of micro-expression recognition training.

[0053] In some more specific implementations, the facial modeling module 202 may further define a relationship between a color value of a muscle activity area and an intensity value of a facial action unit, and a relationship between a transparency of a muscle activity area and an intensity value of a facial action unit.

[0054] Correspondingly, the facial action unit detection module 204 can also be used to quantitatively describe facial features using a pre-trained regression model to obtain the intensity value of the facial action unit.

[0055] The above regression model is obtained by pre-training a support vector machine with a linear kernel. When training the regression model, the collected facial feature samples are scored for the intensity values ​​of different facial action units. The higher the intensity value, the higher the correlation between the facial feature sample and the corresponding facial action unit. The above support vector machine is trained for regression using these facial feature samples, so that the support vector machine can fit the relationship between the facial features and the intensity values ​​of different facial action units, thereby obtaining the above regression model. Using this regression model, the intensity value of the facial action unit can be predicted based on the input facial features.

[0056] Correspondingly, the above-mentioned simulation video generation module 205 is also specifically used to adjust the color value and transparency of the muscle activity area corresponding to the activated facial action unit in the facial area image based on the intensity value of the facial action unit, and display the muscle activity area with color intensity to achieve highlighting of the corresponding muscle activity area.

[0057] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0058] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A micro-expression recognition ability training system based on action unit driven expression simulation, characterized in that: include: A data acquisition module is configured to acquire an original video frame, wherein the original video frame is a real face image; A facial modeling module, configured to identify a facial region image from the original video frame, divide muscle activity regions in the facial region image, and establish a correspondence between the muscle activity regions and facial action units, each facial action unit representing a facial action; A feature extraction module, configured to extract features from the facial region image to obtain facial features; A facial action unit detection module is configured to detect the facial features using a pre-trained classification model to determine an activated facial action unit, wherein the activated facial action unit is used to indicate that a muscle activity area corresponding to the facial action unit undergoes a shape change specified by a corresponding facial action; a simulated video generation module configured to highlight the muscle activity area corresponding to the activated facial action unit in the facial area image to obtain a simulated video frame; The video playing module is configured to play the simulated video frame.

2. The system according to claim 1, characterized in that The facial modeling module is specifically used to perform facial region detection on the original video frame using a pre-trained face target detection model to obtain the facial region image.

3. The system according to claim 1, characterized in that The facial modeling module is specifically used to detect feature points of the facial region image, and divide the muscle activity areas in the facial region image according to the detected facial feature points.

4. The system according to claim 1, characterized in that The feature extraction module is specifically used to perform feature point detection and local feature extraction on the facial region image, and construct the facial features based on the extracted facial feature points and local feature vectors.

5. The system according to claim 4, characterized in that The feature extraction module is specifically used to extract directional gradient histogram features from the facial area image to obtain the local features.

6. The system according to claim 4, characterized in that The feature extraction module is specifically used to use a constrained local neural field algorithm to perform feature point detection on the facial area image to obtain the facial feature points.

7. The system according to claim 1, characterized in that The facial modeling module is further specifically used to define the relationship between the color value of the muscle activity area and the intensity value of the facial action unit, and the relationship between the transparency of the muscle activity area and the intensity value of the facial action unit; The facial action unit detection module is further specifically used to quantitatively describe the facial features using a pre-trained regression model to obtain the intensity value of the facial action unit; The simulated video generation module is specifically used to adjust the color value and transparency of the muscle activity area corresponding to the activated facial action unit in the facial area image based on the intensity value of the facial action unit, so as to highlight the corresponding muscle activity area.

8. The system according to claim 7, characterized in that The regression model is obtained by pre-training a support vector machine with a linear kernel.

9. The system according to claim 1, characterized in that The classification model is obtained by pre-training a support vector machine with a linear kernel.

10. The system according to claim 1, characterized in that The video playing module is specifically used to play the simulated video frame and the original video frame of the simulated video frame simultaneously in the same picture.

Citation Information

Patent Citations

  • Robot and method for simulating human facial movements by robot

    CN106326980A

  • Image generation method and device, storage medium and electronic equipment

    CN110349232A

  • Facial expression interaction method, interaction device and computer storage medium

    CN111638784A

  • Depression recognition system based on micro-expression analysis

    CN112232191A

  • Awareness detection method and equipment based on olfactory stimulation and facial expression, and medium

    CN117633606A