Facial Micro-Expression Recognition Method, Device, Equipment and Storage Medium

By learning face semantics on the pre-trained model and training on it using a small number of face images marked with face micro-expressions, the problem of difficulty in obtaining training samples of the micro-expression recognition model is solved, and the recognition accuracy is improved.

CN113065512BActive Publication Date: 2025-07-01ONE CONNECT SMART TECH CO LTD SHENZHEN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110432485.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-21
Publication Date
2025-07-01
Estimated Expiration
2041-04-21

AI Technical Summary

Technical Problem

In the prior art, it is difficult to obtain training samples of micro-expression recognition models, resulting in low recognition accuracy.

Method used

By inputting the face images without face information to the pre-trained model for face semantic learning, a face semantic recognition model is obtained, and on the basis of this, a face image marked with face micro-expressions is used to train, and a face micro-expressions recognition model is obtained.

Benefits of technology

Without increasing the calculation overhead, the accuracy of facial micro-expression recognition is improved, and the difficulty and error of training samples acquisition of micro-expression recognition models are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113065512B_ABST
    Figure CN113065512B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and discloses a method, device, equipment and storage medium for facial micro-expression recognition. The method includes: inputting face images of a first preset number of unlabeled face information into a pre-trained model, and re-training the pre-trained model based on at least two preset tasks to obtain a face semantic recognition model; inputting face images of a second preset number of labeled facial micro-expressions into the face semantic recognition model for re-training to obtain a facial micro-expression recognition model; performing micro-expression recognition on a target image based on the facial micro-expression recognition model to obtain the facial micro-expression in the target image. This avoids excessive manual participation in the annotation of training samples during the training process of the facial micro-expression recognition model, and improves the training efficiency and recognition accuracy of the micro-expression recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for facial micro-expression recognition. Background Art

[0002] Micro-expressions are short involuntary facial expressions that people produce when they try to suppress their true inner emotions and deep feelings, and are triggered by emotions and habits without thought control. If the information contained in micro-expressions can be effectively utilized, social phenomena such as fraud can be effectively reduced. For example, in the financial field, fraud prevention can be effectively carried out through micro-expression recognition. At present, common micro-expression recognition technologies are mostly based on micro-expression recognition models. However, since the micro-expression recognition model needs to be trained with a large number of images of human face micro-expressions that have been labeled, and the process of labeling images of human face micro-expressions requires professional psychologists for labeling, and the labeling error is relatively large.

[0003] Therefore, in the prior art, there is a problem that it is difficult to obtain training samples for the micro-expression recognition model, resulting in low recognition accuracy of the micro-expression recognition model. Summary of the Invention

[0004] The present application provides a method, device, equipment and storage medium for facial micro-expression recognition, which can ensure the search accuracy without increasing the computational cost.

[0005] In a first aspect, the present application provides a method for facial micro-expression recognition, the method comprising:

[0006] Inputting face images of a first preset number of unlabeled face information into a pre-trained model, so that the pre-trained model performs face semantic learning on the face images of unlabeled face information based on at least two preset face feature prediction tasks, to obtain a face semantic recognition model;

[0007] Inputting face images of a second preset number of labeled facial micro-expressions into the face semantic recognition model, so that the face semantic recognition model learns from the face images of labeled facial micro-expressions, to obtain a facial micro-expression recognition model; wherein, the second preset number is less than the first preset number;

[0008] Performing face recognition on a target image based on the facial micro-expression recognition model to obtain the facial micro-expression in the target image.

[0009] In a second aspect, the present application further provides a facial micro-expression recognition device, comprising:

[0010] A first obtaining module, configured to input face images of a first preset number of unlabeled face information into a pre-trained model, so that the pre-trained model performs face semantic learning on the face images of unlabeled face information based on at least two preset face feature prediction tasks, and obtain a face semantic recognition model;

[0011] A second obtaining module, configured to input face images of a second preset number of face micro-expressions into the face semantic recognition model, so that the face semantic recognition model learns the face images of face micro-expressions, and obtain a face micro-expression recognition model; wherein, the second preset number is less than the first preset number;

[0012] A recognition module, configured to perform face recognition on a target image based on the face micro-expression recognition model, and obtain the face micro-expression in the target image.

[0013] In a third aspect, the present application further provides a face micro-expression recognition device, including:

[0014] A memory and a processor;

[0015] The memory is used to store a computer program;

[0016] The processor is configured to execute the computer program and implement the steps of the face micro-expression recognition method described in the first aspect above when executing the computer program.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the steps of the face micro-expression recognition method described in the first aspect above.

[0018] The present application discloses a face micro-expression recognition method, device, device and storage medium. By retraining a pre-trained model based on at least two preset face feature prediction tasks, the pre-trained face recognition model performs face semantic learning on face images of a first preset number of unlabeled face information during training, and obtains a face semantic information recognition model. Furthermore, based on the face semantic information recognition model, by training with face images of face micro-expressions with a number less than the first preset number, a face micro-expression recognition model can be obtained. It enables the face micro-expression recognition model to not require a large number of training samples with micro-expressions labeled during training, solves the problem that it is difficult to obtain micro-expression training samples and there are large errors in the prior art, improves the training efficiency of the micro-expression recognition model, and based on the trained face micro-expression recognition model, recognizes the face micro-expression in the target image, improving the accuracy of face micro-expression recognition. Description of the Drawings

[0019] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart of a human face micro-expression recognition method provided by an embodiment of the present application;

[0021] Figure 2 It is a schematic structural diagram of a human face micro-expression recognition device provided by an embodiment of the present application;

[0022] Figure 3 It is a schematic block diagram of the structure of a human face micro-expression recognition device provided by an embodiment of the present application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0024] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all the content and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0025] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0026] It should also be understood that the term " / and" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0027] Embodiments of the present application provide a method, apparatus, device, and storage medium for facial micro-expression recognition. The facial micro-expression recognition method provided by the embodiments of the present application can be used to retrain a pre-trained model based on a first preset number of unlabeled facial information face images to obtain a facial semantic recognition model, and then train the facial semantic recognition model based on a number of face images less than the first preset number that are labeled with facial micro-expressions to obtain a facial micro-expression recognition model. Since the training of the facial micro-expression recognition model is carried out on the basis of the facial semantic recognition model, it is possible to reduce the number of manually labeled facial micro-expression training samples, improve the training efficiency of the micro-expression recognition model, and at the same time, based on the trained facial micro-expression recognition model, recognize the facial micro-expressions in the target image, thereby improving the accuracy of facial micro-expression recognition.

[0028] For example, the facial micro-expression recognition method provided by the embodiments of the present application can be applied to a terminal or a server. By retraining a pre-trained model based on images of a first preset number of unlabeled facial information to obtain a facial semantic recognition model, and then, based on a number of face images less than the first preset number that are labeled with facial micro-expressions, a facial micro-expression recognition model is obtained on the basis of the facial semantic recognition model.

[0029] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0030] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a facial micro-expression recognition method provided by an embodiment of the present application. This facial micro-expression recognition method can be implemented by a terminal or a server. The terminal can be a handheld terminal, a laptop computer, a wearable device, or a robot, etc.; the server can be a single server or a server cluster.

[0031] As Figure 1 shown, the facial micro-expression recognition method provided in this embodiment specifically includes: steps S101 to S103. Details are as follows:

[0032] S101. Input face images of a first preset number of unlabeled facial information into a pre-trained model, so that the pre-trained model performs facial semantic learning on the face images of unlabeled facial information based on at least two preset facial feature prediction tasks to obtain a facial semantic recognition model.

[0033] Among them, the face images are face images of unlabeled facial information, which are used to retrain the pre-trained model based on at least two preset facial feature prediction tasks, so that the pre-trained model learns the facial semantics based on at least two preset facial feature prediction tasks to obtain a facial semantic information recognition model.

[0034] In some embodiments of the present application, the pre-trained model includes two parts; one part is the main body of the model, which is used to extract the features of training samples (in this embodiment, face images without labeled face information); the other part is the fully connected layer and the loss function connected to the main body of the model. For example, in one embodiment, the pre-trained model is a deep neural network model, which includes the main body of the model and the fully connected layer and the loss function connected to the main body of the model. Generally, the main body of the model includes a preset number of convolutional layers, which are used to extract sample features and perform convolutional calculations on the extracted sample features; the structure of the main body of the model does not change during the retraining process of the model, and the parameters of each convolutional layer will be continuously iteratively updated. The fully connected layer and the loss function connected to the main body of the model will change during the retraining process of the model. Specifically, with the different preset face feature prediction tasks, the fully connected layer and the loss function connected to the main body of the model will correspondingly change differently.

[0035] Specifically, the input of the fully connected layer is the weight matrix obtained after the main body of the model (convolutional layer) performs convolutional calculations on the extracted sample features; the fully connected layer adds the weight matrix and the bias term related to the preset face feature prediction task and performs mapping, so as to map n real numbers in (-∞, +∞) to K real numbers in (-∞, +∞) (the real numbers after the mapping of the fully connected layer represent the prediction scores corresponding to the face feature prediction task); the loss function further maps K real numbers in (-∞, +∞) to K real numbers in (0, 1) (the real numbers after the mapping of the loss function are prediction probability values), and at the same time, the sum of the real numbers after the mapping of the loss function is 1. Specifically as follows:

[0036] y = softmax(z) = softmax(W^Tx + b)

[0037] Where x is the vector obtained after the main body part (convolutional layer) of the pre-trained model performs convolutional calculations on the extracted training samples; W^Tx is the weight matrix input to the fully connected layer, b is the bias term, and y is the prediction score output by Softmax. Specifically, the calculation method of Softmax can be expressed as follows:

[0038] softmax(zj) = e^zj / ∑_K e^zj

[0039] Softmax is the mapping function of the fully connected layer of the pre-trained model. After this mapping function, there is a loss function connected. The loss function is related to the preset face feature prediction task, and with the different preset face feature prediction tasks, the loss function is also different. Specifically, the loss function is used to map each prediction score obtained by the mapping function of the fully connected layer to the probability value of the prediction category.

[0040] In some embodiments of the present application, the first preset number of face images with unlabeled face information are input into a pre-trained model, so that the pre-trained model performs face semantic learning on the face images with unlabeled face information based on at least two preset face feature prediction tasks, and a face semantic recognition model is obtained.

[0041] In one embodiment, the preset face semantic prediction task includes at least one of a face feature segmentation task, a face feature ranking task, and a face feature conversion task.

[0042] Exemplarily, the preset face semantic prediction task includes two tasks, namely a first preset face semantic prediction task and a second preset face semantic prediction task. The step of making the pre-trained model perform face semantic learning on the face images with unlabeled face information based on at least two preset face feature prediction tasks to obtain a face semantic recognition model includes: based on the first preset face semantic prediction task, performing face feature ranking training on the pre-trained model based on the input face images until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold; based on the second preset face semantic prediction task, performing three-channel color prediction training on the pre-trained model based on the input face images until the second prediction value output by the pre-trained model is greater than a preset second prediction threshold.

[0043] Among them, the first preset face semantic prediction task can be a face feature cutting and sorting task. The pre-trained model is based on the face feature cutting and sorting task for unlabeled image data containing frontal face information. First, face features are extracted through a convolutional layer, and then based on the face feature cutting and sorting task, the face features output by the convolutional layer are subjected to regional cutting prediction through a fully connected layer to obtain the prediction results (scores) of the parts containing different face features belonging to each region. According to the prediction results, the parts of different face features are cut into different regions to obtain multiple sub-images, the sub-images are re-sorted, and the score of the sorting result is predicted. After obtaining the scores of the sorting prediction results of the different sub-images after cutting, the score of the sorting result is mapped to a probability value through a loss function. Among them, the face feature cutting task can be cutting based on the midpoint of the image data, corresponding to obtaining four sub-images; then, based on the four sub-images after cutting, the pre-trained model (mainly the values of the fully connected layer and the loss function) is further trained for face feature sorting. Specifically, based on the first preset face feature prediction task, the pre-trained model is trained for face feature sorting based on the input face image until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold, including: based on the face feature segmentation task, the pre-trained model is trained for face segmentation based on the input face image to obtain multiple segmented face sub-images output by the pre-trained model; based on the face feature sorting task, the pre-trained model is trained for face feature sorting based on the face sub-images until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold.

[0044] Specifically, based on the face feature sorting task, the pre-trained model is trained for face feature sorting based on the face sub-images until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold, including: based on the face feature sorting task, the fully connected layer of the pre-trained model is trained for sorting according to the face features in the four sub-images until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold. In addition, the sorting method adopted during the face feature sorting training process can be any combination of the four sub-images. The fully connected layer predicts the order of each combination of the four sub-images to obtain the probability value of each combination order, and determines the arrangement order of the four sub-images according to the size of the probability value.

[0045] In some embodiments of the present application, the second preset face feature prediction task includes a face feature conversion task. Correspondingly, based on the second preset face feature prediction task, the pre-trained model is trained for three-channel color prediction based on the input face image until the second prediction value output by the pre-trained model is greater than a preset second prediction threshold, including: based on the face feature conversion task, performing channel conversion and color filling on the input face image in the pre-trained model until the second prediction value output by the pre-trained model is greater than the preset second prediction threshold.

[0046] Specifically, each convolutional layer of the pre-trained model extracts face features from the unlabeled face image, and the fully connected layer performs channel conversion on the extracted image containing different face features. For example, the extracted image containing different face features is a multi-channel color picture, and the fully connected layer converts the multi-channel color picture into a single-channel black and white picture, and further performs color filling prediction on the black and white picture through the fully connected layer. Specifically, the convolutional layer of the pre-trained model continuously extracts face features from the face image, and the fully connected layer performs channel conversion on the extracted image containing different face features and performs color filling on the black and white picture. Among them, when the pre-trained model predicts and further converts the single-channel black and white picture into a three-channel image through the fully connected layer, the predicted score corresponding to the three-channel value filled in each pixel point is determined, and according to the predicted score size of the three-channel value, the predicted score corresponding to the predicted three-channel color of each pixel point is determined, and further, the predicted score of the three-channel color is mapped to a probability value through the loss function.

[0047] In addition, the process of predicting the score of the three-channel color by the pre-trained model is a process of learning and understanding the face semantic information, and in this process, the ability of the pre-trained model to understand the semantic information of the face image can be enhanced.

[0048] Exemplarily, the face image semantic information includes the model's understanding information of the facial features in the face image and the judgment information of the positions of the facial features, such as the information of the facial color blocks represented in the face image, the information of the facial lines, and the position information of the facial features such as eyes, ears, nose, and mouth.

[0049] It should be noted that during the retraining process of the pre-trained model, as the preset face feature prediction task is different, the fully connected layer and the loss function of the pre-trained model change accordingly with the preset face feature prediction task. For example, during the training of the pre-trained model based on the first preset face feature prediction task, the output of the corresponding fully connected layer is It represents the predicted value of the arrangement order of the face features included in each sub-image, and the corresponding loss function is used to map the predicted value of the arrangement order of the face features included in each sub-image to the predicted probability value of the arrangement order of the face features included in each sub-image. Specifically, the value obtained by subtracting the predicted value of the arrangement order of each sub-image from the sequence value corresponding to the position order of each sub-image in the original face image and then normalizing it is used as the predicted probability value of the arrangement order of the face features included in each sub-image.

[0050] Exemplarily, the corresponding loss function can be expressed as:

[0051]

[0052] Wherein, represents the sequence value corresponding to the position order of each sub-image in the original face image, represents the predicted value of the arrangement order of the face features included in each sub-image output by the pre-trained model.

[0053] In one embodiment, during the process of training the pre-trained model based on the second preset face feature prediction task, the output of the corresponding fully connected layer is Y, which represents the score of the predicted three-channel color corresponding to each pixel point; the corresponding loss function is used to map the score of the predicted three-channel color corresponding to each pixel point to the probability value of the three-channel color of each pixel point. Specifically, based on the sum of the squared residuals of the pixels in the original image and the pixels in the re-filled image as the probability value of the three-channel color after mapping, that is, the loss value of the corresponding loss function is the loss function with the sum of the squared residuals of the pixels in the original image and the pixels in the re-filled image as the loss value, which can be expressed as:

[0054]

[0055] Wherein, n is the number of filled images, Y - f(X) represents the residual between the pixels in the original image and the pixels in the re-filled image, L represents the probability value after mapping, and f(x) is the three-channel pixel value of each pixel point in the original image.

[0056] In addition, after retraining the pre-trained model based on different preset face feature prediction tasks, the obtained face semantic recognition model can be a multi-task target recognition model, and the loss function of the multi-task target recognition model includes the sum of the loss functions corresponding to each preset face feature prediction task.

[0057] For example, in the above embodiment, the preset face feature prediction tasks include a first preset face feature prediction task and a second preset face feature prediction task. Correspondingly, the pre-trained model is retrained based on the first preset face feature prediction task and the second preset face feature prediction task respectively. After the retraining is completed, the first preset face feature prediction task and the second preset face feature prediction task respectively correspond to different loss functions; and the loss function of the finally obtained face semantic recognition model is the sum of the loss function corresponding to the first preset face feature prediction task and the loss function corresponding to the second preset face feature prediction task.

[0058] S102. Input face images with a second preset number of labeled face micro-expressions into the face semantic recognition model, so that the face semantic recognition model learns from the face images with labeled face micro-expressions to obtain a face micro-expression recognition model; wherein, the second preset number is less than the first preset number.

[0059] Among them, the face images with a second preset number of labeled face micro-expressions are a small number of face images with unlabeled face information less than the first preset number, and are used to retrain the face semantic recognition model, so that the face semantic recognition model strengthens the recognition of face micro-expressions to obtain a face micro-expression recognition model. Since the face semantic recognition model can already effectively recognize the semantic information in the face image, the face semantic recognition model can be retrained based on the fine-tuning method or the transfer learning method to obtain the final face micro-expression recognition model.

[0060] In one embodiment, the process of retraining the face semantic recognition model based on the fine-tuning method includes:

[0061] Adjust the face semantic recognition model based on the samples with labeled face micro-expressions, so that the parameters in the middle of the face semantic recognition model are as close as possible, or when adjusting the face semantic recognition model based on the samples with labeled face micro-expressions, only adjust the parameters of the preset layer and do not adjust the parameters of other layers, so that the parameters of other layers do not change.

[0062] In one embodiment, the process of retraining the face semantic recognition model based on the transfer learning method includes at least one of retraining the face semantic recognition model based on the instance-based transfer learning method, retraining the face semantic recognition model based on the feature representation-based transfer learning method, retraining the face semantic recognition model based on the parameter-based transfer learning method, or retraining the semantic recognition model based on the relationship knowledge transfer learning method.

[0063] In one embodiment, the facial micro-expression recognition model includes a linear classifier for predicting facial micro-expressions. The facial micro-expression recognition model also includes a preset number of convolutional layers, fully connected layers, and a loss function; the fully connected layer and the loss function are different from the above-mentioned pre-trained model and the facial semantic information recognition model. The fully connected layer of the facial micro-expression recognition model is a linear classifier. Among them, the linear classifier can be expressed as f(x) = W×X + b. Here, W is a 10*3072-dimensional matrix, and the columns of the matrix represent the label categories of preset micro-expressions. x is a 3072*1-dimensional flattened input image vector, representing the micro-expression category represented by the image; b is a 10*1-dimensional scalar deviation value, representing the acceptable deviation for micro-expression prediction. For the training of this linear classifier, it is to learn W and b. After W and b are trained and fixed, the classification test process is a simple matrix multiplication (and addition).

[0064] S103, perform face recognition on the target image based on the facial micro-expression recognition model to obtain the facial micro-expression in the target image.

[0065] Through the above analysis, it can be seen that the facial micro-expression recognition method provided by the embodiment of the present application retrains the pre-trained model based on at least two preset tasks, so that the pre-trained face recognition model learns facial semantic information based on the face images of the first preset number of unlabeled face information during the training process, and obtains a facial semantic information recognition model. Furthermore, based on the facial semantic information recognition model, it is possible to train the facial semantic information recognition model based on the face images of the standard facial micro-expressions that are less than the first preset number, and obtain a facial micro-expression recognition model. It is realized that by training with a small number of samples labeled with micro-expressions, a facial micro-expression recognition model can be obtained. The facial micro-expression recognition model does not require a large number of training samples labeled with micro-expressions during the training process, solves the problem that it is difficult to obtain micro-expression training samples and there are large errors in the prior art, improves the training efficiency of the micro-expression recognition model, and based on the trained facial micro-expression recognition model, recognizes the facial micro-expression in the target image, improving the accuracy of facial micro-expression recognition.

[0066] Please refer to Figure 2 , Figure 2 is a schematic structural diagram of a facial micro-expression recognition device provided by an embodiment of the present application. The facial micro-expression recognition device is used to execute Figure 1 the facial micro-expression recognition method shown. The facial micro-expression recognition device can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device.

[0067] As Figure 2 shown, the facial micro-expression recognition device 200 includes:

[0068] The first obtaining module 201 is configured to input face images of a first preset number of unlabeled face information into a pre-trained model, so that the pre-trained model performs face semantic learning on the face images of unlabeled face information based on at least two preset face feature prediction tasks, and obtain a face semantic recognition model;

[0069] The second obtaining module 202 is configured to input face images of a second preset number of face microexpressions into the face semantic recognition model, so that the face semantic recognition model learns the face images of face microexpressions, and obtain a face microexpression recognition model; wherein, the second preset number is less than the first preset number;

[0070] The recognition module 203 is configured to perform face recognition on a target image based on the face microexpression recognition model, and obtain the face microexpression in the target image.

[0071] In one embodiment, the pre-trained model includes a model main body part, a fully connected layer connected to the model main body part, and a loss function; the first obtaining module 201 includes:

[0072] The first obtaining unit is configured to input face images of a first preset number of unlabeled face information into the pre-trained model, so that the model main body part extracts the face features included in the face images, performs convolution calculation on the face features, and obtain a weight matrix;

[0073] The second obtaining unit is configured to add the weight matrix and a bias term based on at least two of the preset face feature prediction tasks through the fully connected layer and perform mapping, so as to obtain a prediction score corresponding to the face feature prediction task, where the bias term is a value related to the preset face feature prediction task;

[0074] The third obtaining unit is configured to perform secondary mapping on the prediction score through the loss function, and obtain a prediction probability value corresponding to the preset face feature prediction task.

[0075] In one embodiment, the preset task includes at least one of a face feature segmentation task, a face feature sorting task, and a face feature conversion task.

[0076] In one embodiment, the at least two preset face feature prediction tasks include a first preset face feature prediction task and a second preset face feature prediction task; the first obtaining module 201 includes:

[0077] The first training unit is configured to perform face feature sorting training on the pre-trained model based on the input face images based on the first preset face feature prediction task until a first prediction value output by the pre-trained model is greater than a preset first prediction threshold;

[0078] A second training unit, configured to perform three-channel color prediction training on the pre-trained model based on the input face image according to the second preset face feature prediction task until the second prediction value output by the pre-trained model is greater than a preset second prediction threshold.

[0079] In one embodiment, the first preset face feature prediction task includes a face feature segmentation task and a face feature sorting task;

[0080] The first training unit includes:

[0081] A first obtaining subunit, configured to perform face segmentation training on the pre-trained model based on the input face image according to the face feature segmentation task, to obtain a plurality of segmented face sub-images output by the pre-trained model;

[0082] A first training subunit, configured to perform face feature sorting training on the pre-trained model based on the face sub-images according to the face feature sorting task until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold.

[0083] In one embodiment, the second preset task includes a face feature conversion task;

[0084] The second training unit is specifically configured to:

[0085] Perform channel conversion and color filling on the input face image in the pre-trained model according to the face feature conversion task until the second prediction value output by the pre-trained model is greater than a preset second prediction threshold.

[0086] In one embodiment, the second obtaining module 202 is specifically configured to:

[0087] Input a second preset number of face images labeled with face micro-expressions into the face semantic recognition model, and perform re-learning training on the face semantic recognition model based on fine-tuning or transfer learning to obtain the face micro-expression recognition model.

[0088] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described terminal and each module can refer to Figure 1 The corresponding processes in the embodiments of the face micro-expression recognition method described above, which will not be elaborated here.

[0089] The above face micro-expression recognition method can be implemented in the form of a computer program, and the computer program can run on a device as shown in Figure 2 shown.

[0090] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of the structure of the face micro-expression recognition device provided by an embodiment of the present application. The face micro-expression recognition device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.

[0091] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can be made to execute any face micro-expression recognition method.

[0092] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0093] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can be made to execute any face micro-expression recognition method.

[0094] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 3 the structure shown in

[0095] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the terminal to which the solution of the present application is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0096] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0096] Wherein, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:

[0097] Input face images of a first preset number of unlabeled face information into a pre-trained model, so that the pre-trained model performs face semantic learning on the face images of unlabeled face information based on at least two preset face feature prediction tasks to obtain a face semantic recognition model;

[0098] Input a second preset number of face images with micro facial expressions marked into the face semantic recognition model, so that the face semantic recognition model learns from the face images with micro facial expressions marked to obtain a micro facial expression recognition model; wherein, the second preset number is less than the first preset number;

[0099] Based on the micro facial expression recognition model, perform face recognition on the target image to obtain the micro facial expression in the target image.

[0100] In one embodiment, the pre-trained model includes a model main body part, a fully connected layer connected to the model main body part, and a loss function; the step of inputting a first preset number of face images without marked face information into the pre-trained model, so that the pre-trained model performs face semantic learning on the face images without marked face information based on at least two preset face feature prediction tasks to obtain a face semantic recognition model includes:

[0101] Input a first preset number of face images without marked face information into the pre-trained model, so that the model main body part extracts the face features included in the face images, performs convolution calculation on the face features to obtain a weight matrix;

[0102] Based on at least two of the preset face feature prediction tasks through the fully connected layer, add the weight matrix and a bias term and perform mapping to obtain a prediction score corresponding to the face feature prediction task, where the bias term is a value related to the preset face feature prediction task;

[0103] Perform a second mapping on the prediction score through the loss function to obtain a prediction probability value corresponding to the preset face feature prediction task.

[0104] In one embodiment, the preset face feature prediction tasks include at least one of a face feature segmentation task, a face feature ranking task, and a face feature conversion task.

[0105] In one embodiment, the at least two preset face feature prediction tasks include a first preset face feature prediction task and a second preset face feature prediction task; the step of making the pre-trained model perform face semantic learning on the face images without marked face information based on at least two preset face feature prediction tasks to obtain a face semantic recognition model includes:

[0106] Based on the first preset face feature prediction task, perform face feature ranking training on the pre-trained model based on the input face images until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold;

[0107] Based on the second preset facial feature prediction task, perform three-channel color prediction training on the pre-trained model based on the input face image until the second prediction value output by the pre-trained model is greater than the preset second prediction threshold.

[0108] In one embodiment, the first preset facial feature prediction task includes a facial feature segmentation task and a facial feature sorting task;

[0109] Based on the first preset facial feature prediction task, performing facial feature sorting training on the pre-trained model based on the input face image until the first prediction value output by the pre-trained model is greater than the preset first prediction threshold includes:

[0110] Based on the facial feature segmentation task, perform facial segmentation training on the pre-trained model based on the input face image to obtain multiple segmented facial sub-images output by the pre-trained model;

[0111] Based on the facial feature sorting task, perform facial feature sorting training on the pre-trained model based on the facial sub-images until the first prediction value output by the pre-trained model is greater than the preset first prediction threshold.

[0112] In one embodiment, the second preset facial feature prediction task includes a facial feature conversion task;

[0113] Based on the second preset facial feature prediction task, performing three-channel color prediction training on the pre-trained model based on the input face image until the second prediction value output by the pre-trained model is greater than the preset second prediction threshold includes:

[0114] Based on the facial feature conversion task, perform channel conversion and color filling on the input face image in the pre-trained model until the second prediction value output by the pre-trained model is greater than the preset second prediction threshold.

[0115] In one embodiment, inputting the second preset number of face images labeled with facial micro-expressions into the face semantic recognition model, so that the face semantic recognition model learns the face images labeled with facial micro-expressions to obtain a facial micro-expression recognition model, includes:

[0116] Input the second preset number of face images labeled with facial micro-expressions into the face semantic recognition model, and perform re-learning training on the face semantic recognition model based on fine-tuning or transfer learning to obtain the facial micro-expression recognition model.

[0117] An embodiment of the present application further provides a computer-readable storage medium storing a computer program including program instructions, and when the processor executes the program instructions, the method for recognizing human face micro-expressions provided by the embodiment shown in the present application is implemented. Figure 1 The method for recognizing human face micro-expressions provided by the embodiment shown.

[0118] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.

[0119] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for facial micro-expression recognition, characterized in that, The method includes: Inputting face images with the first preset number of unlabeled face information into a pre-trained model, so that the pre-trained model performs face semantic learning on the face images with unlabeled face information based on at least two preset face feature prediction tasks to obtain a face semantic recognition model, where the preset face feature prediction tasks include at least one of a face feature segmentation task, a face feature sorting task, and a face feature conversion task; Among them, the pre-trained model, based on the face feature cutting and sorting task, for the unlabeled image data containing frontal face information, first extracts face features through a convolutional layer, and then, based on the face feature cutting and sorting task, through a fully connected layer, predicts the regional cutting of the face features output by the convolutional layer to obtain the prediction results of the parts containing different face features belonging to each region, and cuts the parts of different face features of the face into different regions according to the prediction results to obtain multiple sub-images, reorders each of the sub-images, and predicts the score of the sorting result. After obtaining the scores of the sorting prediction results of the different sub-images after cutting, the scores of the sorting results are mapped to probability values through a loss function; each convolutional layer of the pre-trained model extracts face features from the unlabeled face image based on the face feature conversion task, and performs channel conversion on the extracted images containing different face features through a fully connected layer. The convolutional layer of the pre-trained model continuously extracts face features from the face image, and the fully connected layer performs channel conversion on the extracted images containing different face features and fills in colors for black and white pictures; among them, when the pre-trained model predicts and further converts a single-channel black and white picture into a three-channel image through a fully connected layer, predicts the prediction scores corresponding to the three-channel values filled in each pixel point, determines the scores of the predicted three-channel colors corresponding to each pixel point according to the magnitudes of the prediction scores of the predicted three-channel values, and further maps the scores of the predicted three-channel colors to probability values through a loss function; Inputting face images with the second preset number of labeled face micro-expressions into the face semantic recognition model, so that the face semantic recognition model learns the face images with labeled face micro-expressions to obtain a face micro-expression recognition model; where the second preset number is less than the first preset number; Performing face recognition on a target image based on the face micro-expression recognition model to obtain the face micro-expression in the target image.

2. The face micro-expression recognition method according to claim 1, wherein The pre-trained model includes a model main body part, a fully connected layer connected to the model main body part, and a loss function; the step of inputting face images with the first preset number of unlabeled face information into the pre-trained model, so that the pre-trained model performs face semantic learning on the face images with unlabeled face information based on at least two preset face feature prediction tasks to obtain a face semantic recognition model includes: Inputting face images with the first preset number of unlabeled face information into the pre-trained model, so that the model main body part extracts the face features included in the face images, performs convolutional calculation on the face features to obtain a weight matrix; Based on at least two of the preset face feature prediction tasks, the fully connected layer adds the weight matrix and the bias term and performs mapping to obtain a prediction score corresponding to the face feature prediction task, where the bias term is a value related to the preset face feature prediction task; The loss function performs a second mapping on the prediction score to obtain a prediction probability value corresponding to the preset face feature prediction task.

3. The face micro-expression recognition method according to claim 1, characterized in that The at least two preset face feature prediction tasks include a first preset face feature prediction task and a second preset face feature prediction task; the step of enabling the pre-trained model to perform face semantic learning on a face image of unlabeled face information based on at least two preset face feature prediction tasks to obtain a face semantic recognition model includes: Based on the first preset face feature prediction task, perform face feature ranking training on the pre-trained model based on the input face image until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold; Based on the second preset face feature prediction task, perform three-channel color prediction training on the pre-trained model based on the input face image until the second prediction value output by the pre-trained model is greater than a preset second prediction threshold.

4. The face micro-expression recognition method according to claim 3, characterized in that, The first preset face feature prediction task includes a face feature segmentation task and a face feature ranking task; Based on the first preset face feature prediction task, performing face feature ranking training on the pre-trained model based on the input face image until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold includes: Based on the face feature segmentation task, perform face segmentation training on the pre-trained model based on the input face image to obtain multiple segmented face sub-images output by the pre-trained model; Based on the face feature ranking task, perform face feature ranking training on the pre-trained model based on the face sub-images until the first prediction value output by the pre-trained model is greater than a preset first prediction threshold.

5. The face micro-expression recognition method according to claim 3, characterized in that, The second preset face feature prediction task includes a face feature conversion task; Based on the second preset face feature prediction task, performing three-channel color prediction training on the pre-trained model based on the input face image until the second prediction value output by the pre-trained model is greater than a preset second prediction threshold includes: Based on the face feature conversion task, perform channel conversion and color filling on the input face image in the pre-trained model until the second prediction value output by the pre-trained model is greater than a preset second prediction threshold.

6. The face micro-expression recognition method according to claim 1, characterized in that The step of inputting a second preset number of face images labeled with face micro-expressions into the face semantic recognition model to enable the face semantic recognition model to learn from the face images labeled with face micro-expressions to obtain a face micro-expression recognition model includes: Input a second preset number of face images labeled with face micro-expressions into the face semantic recognition model, and perform re-learning training on the face semantic recognition model based on fine-tuning or transfer learning to obtain the face micro-expression recognition model.

7. A micro-expression recognition model training device, characterized in that Includes: A first obtaining module, configured to input face images of a first preset number of unlabeled face information into a pre-trained model, so that the pre-trained model performs face semantic learning on the face images of unlabeled face information based on at least two preset face feature prediction tasks, and obtains a face semantic recognition model, wherein the preset face feature prediction tasks include at least one of a face feature segmentation task, a face feature sorting task, and a face feature conversion task; Among them, the pre-trained model is based on the face feature cutting and sorting task for the unlabeled image data containing frontal face information. First, face features are extracted through a convolutional layer, and then based on the face feature cutting and sorting task, the face features output by the convolutional layer are predicted for region cutting through a fully connected layer, and the prediction results of the parts containing different face features belonging to each region are obtained. According to the prediction results, the parts of different face features of the face are cut into different regions to obtain a plurality of sub-images, and each of the sub-images is re-sorted, and the score of the sorting result is predicted. After obtaining the scores of the sorting prediction results of the different sub-images after cutting, the score of the sorting result is mapped to a probability value through a loss function; Each convolutional layer of the pre-trained model performs face feature extraction on the unlabeled face image based on the face feature conversion task, and performs channel conversion on the extracted image containing different face features through a fully connected layer. The convolutional layer of the pre-trained model continuously extracts face features from the face image, and the fully connected layer performs channel conversion on the extracted image containing different face features and fills the black and white picture with colors; Among them, when the pre-trained model predicts through the fully connected layer to further convert a single-channel black and white picture into a three-channel image, the predicted scores of the three-channel values filled for each pixel point are determined. According to the size of the predicted scores of the three-channel values, the scores of the predicted three-channel colors corresponding to each pixel point are determined, and further the scores of the predicted three-channel colors are mapped to probability values through a loss function; A second obtaining module, configured to input face images of a second preset number of labeled face micro-expressions into the face semantic recognition model, so that the face semantic recognition model learns the face images of labeled face micro-expressions to obtain a face micro-expression recognition model; wherein, the second preset number is less than the first preset number; A recognition module, configured to perform face recognition on a target image based on the face micro-expression recognition model to obtain the face micro-expression in the target image.

8. A micro-expression recognition model training device, characterized in that Comprising: A memory and a processor; The memory is used for storing a computer program; The processor is configured to execute the computer program and implement the steps of the face micro-expression recognition method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the processor is caused to implement the steps of the face micro-expression recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Facial micro-expression recognition method

    CN107679526A

  • Training method of face recognition model and online education system

    CN112329735A