Multimodal fusion facial expression intelligent perception method and system

CN115862099BActive Publication Date: 2026-09-01TSINGHUA UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211501478.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2026-09-01
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

然而,相关表情感知相关的计算机视觉研究中,分类模型往往是有明确的分类边界的,从而导致表情感知的准确性较低

Benefits of technology

[0062]在本公开一个或多个实施例中,通过获取待测人脸图像;将待测人脸图像输入至目标多模融合表情智能感知模型中,得到至少一种表情对应的输出预测分数集合;根据输出预测分数集合,确定待测人脸图像对应的表情感知结果。因此,通过将表情图像的多分类问题分解为多个二分类问题,获取至少一种表情对应的输出预测分数集合,可以无需考虑“表情判定有固定边界”这一缺乏合理性的假设,可以模糊分类边界,提高表情感知的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862099B_ABST
    Figure CN115862099B_ABST
Patent Text Reader

Abstract

This disclosure relates to the fields of computer vision and neuroscience, and particularly to a multimodal fusion intelligent expression perception method and system. The multimodal fusion intelligent expression perception method includes: acquiring a face image to be tested; inputting the face image to be tested into a target multimodal fusion intelligent expression perception model to obtain a set of output prediction scores corresponding to at least one expression; and determining the expression perception result corresponding to the face image to be tested based on the set of output prediction scores. Using this disclosure can blur classification boundaries and improve the accuracy of expression perception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of computer vision and neuroscience, and in particular to a multimodal fusion facial expression intelligent perception method and system. Background Technology

[0002] In reality, facial expressions often convey complex or even mixed emotions, and their discrimination boundaries are inherently vague and even overlapping. In other words, in brain-like perception, there shouldn't be completely defined boundaries between expression categories. However, in computer vision research related to facial expression perception, classification models often have clear classification boundaries, leading to lower accuracy in expression perception. Therefore, how to blur classification boundaries and improve the accuracy of facial expression perception has become a key focus. Summary of the Invention

[0003] This disclosure provides a multi-modal fusion facial expression intelligent perception method and system, the main purpose of which is to blur the classification boundaries and improve the accuracy of facial expression perception.

[0004] According to one aspect of this disclosure, a multi-modal fusion facial expression intelligent perception method is provided, comprising:

[0005] Acquire the face image of the person to be tested;

[0006] The face image to be tested is input into the target multimodal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression;

[0007] Based on the set of output prediction scores, the expression perception result corresponding to the face image to be tested is determined.

[0008] Optionally, before inputting the face image to be tested into the target multimodal fusion expression intelligent perception model, the method further includes:

[0009] Obtain an initial training dataset, wherein the initial training dataset includes at least one face image;

[0010] The initial training dataset is preprocessed to obtain the target training dataset;

[0011] The initial multimodal fusion facial expression intelligent perception model is trained using the target training dataset to obtain the target multimodal fusion facial expression intelligent perception model.

[0012] Optionally, the preprocessing of the initial training dataset to obtain the target training dataset includes:

[0013] Face localization is performed on at least one face image in the initial training dataset to obtain at least one face detection and localization result corresponding to the at least one face image;

[0014] Based on the at least one face detection and localization result, data augmentation is performed on the at least one face image to obtain at least one data-augmented face image;

[0015] Determine the eye region dataset, mouth region dataset, and overall face dataset corresponding to the at least one data-enhanced face image;

[0016] The target training dataset is determined based on the eye region dataset, the mouth region dataset, and the overall facial dataset.

[0017] Optionally, the data augmentation method includes at least one of the following:

[0018] Small-angle random rotation;

[0019] Flip horizontally;

[0020] Translation;

[0021] Cutting.

[0022] Optionally, determining the eye region dataset, mouth region dataset, and overall facial dataset corresponding to the at least one data-enhanced face image includes:

[0023] Rotate the at least one data-enhanced face image to the target angle to obtain the overall face dataset;

[0024] At least one overall facial image from the overall facial dataset is masked to obtain the eye region dataset and the mouth region dataset.

[0025] Optionally, the step of inputting the face image to be tested into the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression includes:

[0026] The face image to be tested is input into at least one expression discrimination branch of the target multimodal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression.

[0027] Optionally, the expression discrimination branch includes three classifier paths. The step of inputting the test face image into at least one expression discrimination branch of the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression includes:

[0028] The face image to be tested is input into the three classifier paths corresponding to any one of the expression discrimination branches in at least one expression discrimination branch to obtain three probability scores corresponding to the face image to be tested;

[0029] Based on the activity parameters corresponding to the three classifier paths, the three probability scores are weighted and concatenated to obtain the output prediction score corresponding to any expression discrimination branch, so as to obtain the set of output prediction scores corresponding to at least one expression.

[0030] According to another aspect of this disclosure, a multi-modal fusion facial expression intelligent perception system is provided, comprising:

[0031] The image acquisition unit is used to acquire the face image of the person to be tested;

[0032] The set acquisition unit is used to input the face image to be tested into the target multimodal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression;

[0033] The result determination unit is used to determine the expression perception result corresponding to the face image to be tested based on the output prediction score set.

[0034] Optionally, the system further includes a dataset acquisition unit, a training set processing unit, and a model training unit, used before inputting the test face image into the target multimodal fusion expression intelligent perception model:

[0035] The dataset acquisition unit is used to acquire an initial training dataset, wherein the initial training dataset includes at least one face image;

[0036] The training set processing unit is used to preprocess the initial training dataset to obtain the target training dataset;

[0037] The model training unit is used to train the initial multimodal fusion facial expression intelligent perception model using the target training dataset to obtain the target multimodal fusion facial expression intelligent perception model.

[0038] Optionally, the training set processing unit is used to preprocess the initial training dataset to obtain the target training dataset, specifically for:

[0039] Face localization is performed on at least one face image in the initial training dataset to obtain at least one face detection and localization result corresponding to the at least one face image;

[0040] Based on the at least one face detection and localization result, data augmentation is performed on the at least one face image to obtain at least one data-augmented face image;

[0041] Determine the eye region dataset, mouth region dataset, and overall face dataset corresponding to the at least one data-enhanced face image;

[0042] The target training dataset is determined based on the eye region dataset, the mouth region dataset, and the overall facial dataset.

[0043] Optionally, the data augmentation method includes at least one of the following:

[0044] Small-angle random rotation;

[0045] Flip horizontally;

[0046] Translation;

[0047] Cutting.

[0048] Optionally, when the training set processing unit determines the eye region dataset, mouth region dataset, and overall face dataset corresponding to the at least one data-augmented face image, it is specifically used for:

[0049] Rotate the at least one data-enhanced face image to the target angle to obtain the overall face dataset;

[0050] At least one overall facial image from the overall facial dataset is masked to obtain the eye region dataset and the mouth region dataset.

[0051] Optionally, when the set acquisition unit inputs the face image to be tested into the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression, it is specifically used for:

[0052] The face image to be tested is input into at least one expression discrimination branch of the target multimodal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression.

[0053] Optionally, the expression discrimination branch includes three classifier paths. The set acquisition unit, when inputting the test face image into at least one expression discrimination branch of the target multi-modal fusion expression intelligent perception model to obtain the output prediction score set corresponding to at least one expression, specifically uses the following:

[0054] The face image to be tested is input into the three classifier paths corresponding to any one of the expression discrimination branches in at least one expression discrimination branch to obtain three probability scores corresponding to the face image to be tested;

[0055] Based on the activity parameters corresponding to the three classifier paths, the three probability scores are weighted and concatenated to obtain the output prediction score corresponding to any expression discrimination branch, so as to obtain the set of output prediction scores corresponding to at least one expression.

[0056] According to another aspect of this disclosure, a terminal is provided, comprising:

[0057] At least one processor; and

[0058] A memory communicatively connected to the at least one processor; wherein,

[0059] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in any one of the preceding aspects.

[0060] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described in any one of the preceding aspects.

[0061] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in any one of the preceding aspects.

[0062] In one or more embodiments of this disclosure, a face image to be tested is acquired; the face image is input into a target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression; and the expression perception result corresponding to the face image to be tested is determined based on the set of output prediction scores. Therefore, by decomposing the multi-classification problem of expression images into multiple binary classification problems and obtaining a set of output prediction scores corresponding to at least one expression, the unreasonable assumption that "expression judgment has fixed boundaries" can be disregarded, the classification boundaries can be blurred, and the accuracy of expression perception can be improved.

[0063] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0064] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0065] Figure 1 A flowchart illustrating the first multimodal fusion facial expression intelligent perception method provided in this disclosure embodiment is shown.

[0066] Figure 2 A flowchart illustrating the second multimodal fusion facial expression intelligent perception method provided in this disclosure embodiment is shown.

[0067] Figure 3 This diagram illustrates a preprocessing flow for an initial training dataset provided in an embodiment of this disclosure.

[0068] Figure 4 This diagram illustrates a training flowchart for an expression discrimination branch provided in an embodiment of the present disclosure;

[0069] Figure 5 A smooth trend graph showing the training accuracy and validation accuracy in the Jeffe dataset provided in this embodiment of the disclosure is shown.

[0070] Figure 6 A smooth trend graph showing the training accuracy and validation accuracy of the CK+ dataset provided in this embodiment of the disclosure is shown.

[0071] Figure 7 This diagram illustrates a workflow of a target multimodal fusion facial expression intelligent perception model provided in an embodiment of this disclosure.

[0072] Figure 8 This diagram illustrates an example of an active region distribution visualization provided by an embodiment of this disclosure.

[0073] Figure 9 This diagram illustrates the structure of a first multimodal fusion facial expression intelligent perception system provided in an embodiment of this disclosure.

[0074] Figure 10 This diagram illustrates the structure of a second multimodal fusion facial expression intelligent perception system provided in an embodiment of this disclosure.

[0075] Figure 11 This is a block diagram of a terminal used to implement the multimodal fusion facial expression intelligent perception method of the embodiments of this disclosure. Detailed Implementation

[0076] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0077] Current research on facial expression recognition is primarily based on single-channel processing. However, the human visual system often involves multiple neural circuits related to recognition when processing a simple visual task. Different facial stimuli generate different active areas in the brain, and the human brain's facial expression perception model exhibits a degree of parallelism. This view is supported by numerous studies in neuroimaging, psychology, and neurology. For example, patients with autism spectrum disorder show low accuracy in recognizing fear, sadness, and disgust, while performing similarly to normal individuals in recognizing other facial expressions. Chronic alcoholics exhibit significant deficits in recognizing happiness and anger. It is evident that some neurological diseases or injuries can affect the brain's facial expression perception process, potentially leading to difficulties in recognizing certain categories of expressions. Therefore, it is understood that the human brain possesses independent perceptual pathways for different categories of facial expressions.

[0078] One manifestation of the insufficient robustness of traditional facial expression perception is its high dependence on the overall intensity of facial expressions. Furthermore, mainstream facial expression perception methods face significant challenges when processing large amounts of facial expression data, as real-life facial expressions often convey complex or even mixed emotions, and their discrimination boundaries are inherently blurred or overlapping. Evidence suggests that infants' ability to perceive facial expressions develops gradually with age, progressing from binary classification (positive and negative emotions). This indicates that, assuming facial expressions are defined in a two-dimensional space, the binary classification boundaries in the initial human recognition system may be more pronounced. Over time, children's facial expression perception abilities gradually emerge, with the earliest and most accurate recognition of happiness, followed by sadness and anger, and then surprise and fear. Moreover, children aged three to five still struggle to recognize neutral expressions. This suggests that the human brain's learning of each emotion category possesses a degree of independence, progressing from the most primitive binary perception to the ability to perceive all categories of facial expressions. Therefore, in brain-like perception, there should be no absolute boundaries between facial expression categories.

[0079] It's easy to understand that in computer vision research related to facial expression perception, classification models often have clear classification boundaries, leading to low accuracy in facial expression perception. Therefore, how to blur the classification boundaries and improve the accuracy of facial expression perception has become a key focus.

[0080] Furthermore, research on facial expressions in related technologies treats the face as a whole, as there is indeed some interrelationship among different facial regions in expression and recognition. However, several studies involving neuroimaging, psychology, and neuroscience have found that different facial regions are responsible for expressing different facial expressions. This indicates that different facial regions contribute differently to specific expressions, and in single-category expression perception, the human brain exhibits a preference for different facial regions in the expression perception process.

[0081] It is easy to understand that the relevant models all treat the face as a whole for expression recognition, which leads to low accuracy in expression recognition. Therefore, how to build a brain-inspired expression perception neural network model to recognize expressions based on the degree of contribution of facial regions to specific expressions, in order to improve the accuracy of expression recognition, has become a focus of attention.

[0082] The present disclosure will now be described in detail with reference to specific embodiments.

[0083] In the first embodiment, such as Figure 1 As shown, Figure 1 The diagram illustrates a flowchart of a first multimodal fusion facial expression intelligent perception method provided in this embodiment. This method can be implemented using a computer program and can run on a system performing multimodal fusion facial expression intelligent perception. The computer program can be integrated into an application or run as a standalone utility application.

[0084] The multi-mode fusion facial expression intelligent perception system can be a terminal with multi-mode fusion facial expression intelligent perception function. This terminal includes, but is not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. In different networks, the terminal may be called by different names, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), 5G network, 4G network, 3G network, or terminals in future evolved networks.

[0085] Specifically, this multimodal fusion facial expression intelligent perception method includes:

[0086] S101, acquire the face image of the person to be tested;

[0087] According to some embodiments, the face image to be tested refers to a face image for which expression perception is required.

[0088] It is easy to understand that when the terminal performs multi-modal fusion facial expression intelligent perception, the terminal can acquire the face image of the person to be tested.

[0089] S102, Input the face image to be tested into the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression;

[0090] According to some embodiments, the target multimodal fusion facial expression intelligent perception model refers to a trained multimodal fusion facial expression intelligent perception model. This target multimodal fusion facial expression intelligent perception model does not specifically refer to a fixed model. For example, when the pre-training multimodal fusion facial expression intelligent perception model, i.e., the initial multimodal fusion facial expression intelligent perception model, changes, the target multimodal fusion facial expression intelligent perception model can also change.

[0091] In some embodiments, at least one expression includes, but is not limited to, surprise, happiness, sadness, etc.

[0092] According to some embodiments, the output prediction score set refers to a set comprised of at least one output prediction score. Each output prediction score corresponds one-to-one with an facial expression.

[0093] It is easy to understand that when the terminal acquires the face image to be tested, the terminal can input the face image to be tested into the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression.

[0094] S103, Based on the output prediction score set, determine the expression perception result corresponding to the face image to be tested.

[0095] It is easy to understand that when the terminal obtains a set of output prediction scores corresponding to at least one expression, the terminal can determine the expression perception result corresponding to the face image to be tested based on the set of output prediction scores.

[0096] In summary, the method provided in this disclosure acquires a face image to be tested; inputs the face image to be tested into a target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression; and determines the expression perception result corresponding to the face image to be tested based on the set of output prediction scores. Therefore, by decomposing the multi-classification problem of expression images into multiple binary classification problems and obtaining a set of output prediction scores corresponding to at least one expression, the unreasonable assumption that "expression judgment has fixed boundaries" can be disregarded, thus blurring the classification boundaries and improving the accuracy of expression perception.

[0097] Please see Figure 2 , Figure 2 This diagram illustrates a flowchart of a second multimodal fusion facial expression intelligent perception method provided in an embodiment of this disclosure. Specifically, the multimodal fusion facial expression intelligent perception method includes:

[0098] S201, Obtain the initial training dataset;

[0099] According to some embodiments, the initial training dataset includes at least one face image.

[0100] In some embodiments, a face image refers to an image containing a complete human face. This face image does not specifically refer to any particular image.

[0101] It is easy to understand that when the terminal performs multi-modal fusion facial expression intelligent perception, the terminal can obtain the initial training dataset.

[0102] S202, preprocess the initial training dataset to obtain the target training dataset;

[0103] According to some embodiments, the target training dataset includes at least one face region training dataset. This at least one face region training dataset includes, but is not limited to, an eye region dataset, a mouth region dataset, and a whole face dataset.

[0104] According to some embodiments, Figure 3 This diagram illustrates a preprocessing flow for an initial training dataset provided in an embodiment of this disclosure. Figure 3 As shown, firstly, face localization can be performed on at least one face image in the initial training dataset to obtain at least one face detection and localization result corresponding to at least one face image. Next, based on the at least one face detection and localization result, data augmentation can be performed on at least one face image to obtain at least one data-augmented face image. Secondly, the eye region dataset, mouth region dataset, and overall facial dataset corresponding to at least one data-augmented face image can be determined. Finally, based on the eye region dataset, mouth region dataset, and overall facial dataset, the target training dataset can be determined.

[0105] In some embodiments, such as Figure 3 As shown, when performing face localization on at least one face image in the initial training dataset, face localization can be performed through feature point detection.

[0106] In some embodiments, during the training phase, a large difference in the number of samples between categories can lead to overfitting in the trained model, causing significant variations in classification performance across different categories. Therefore, after the face detection and localization stage, data with an excessive number of categories can be augmented with limited augmentation, while data with the fewest categories can be fully augmented. This minimizes the difference in the number of samples between categories in the training set. After data augmentation, the face images in each dataset can be scaled to the same size.

[0107] The data augmentation methods include at least one of the following:

[0108] Small-angle random rotation, for example, a small-angle random rotation of -30° to 30°.

[0109] Flip horizontally;

[0110] Translation;

[0111] Cutting.

[0112] In some embodiments, OpenCV's Cascade classifier can be used to perform face detection and localization. The face detection and localization results are then cropped and normalized to ensure that each image contains only facial information and is of uniform size.

[0113] According to some embodiments, such as Figure 3 As shown, to determine the eye region dataset, mouth region dataset, and overall face dataset corresponding to at least one data-augmented face image, firstly, the at least one data-augmented face image can be rotated to the target angle to obtain the overall face dataset. Next, region segmentation can be performed on at least one overall face image in the overall face dataset to obtain the eye region dataset and the mouth region dataset.

[0114] In some embodiments, region subdivision can be achieved, for example, through masking.

[0115] In some embodiments, considering the rotation angle issue, the data-augmented face images can be uniformly rotated to be parallel to the horizontal line. Then, masks of half the image size are placed in the upper and lower parts of the face image, respectively, to obtain eye region datasets and mouth region datasets.

[0116] In some embodiments, the overall facial dataset serves not only to enable the classifier to obtain global features, but also to be used as a comparison group for masked data, thereby evaluating the impact of local occlusion on recognition accuracy.

[0117] It is easy to understand that when the terminal obtains the initial training dataset, it can preprocess the initial training dataset to obtain the target training dataset.

[0118] S203, the initial multimodal fusion facial expression intelligent perception model is trained using the target training dataset to obtain the target multimodal fusion facial expression intelligent perception model;

[0119] According to some implementations, an initial multimodal fusion facial expression intelligent perception model can be trained using a target training dataset via backpropagation of stochastic gradient descent. Stochastic gradient descent (SGD) involves randomly selecting a set of samples, training it, updating it once according to the gradient, then selecting another set and updating it again. With a very large sample size, it may be possible to obtain a model with an acceptable loss value without training all the samples. Here, "random" means that the samples are randomly shuffled during each iteration.

[0120] In some embodiments, when training the initial multimodal fusion expression intelligent perception model using the target training dataset, at least one expression discrimination branch in the initial multimodal fusion expression intelligent perception model can be controlled to iteratively learn from at least one face region training dataset to obtain at least one set of active parameters corresponding to at least one expression discrimination branch. Then, the target multimodal fusion expression intelligent perception model can be determined based on the at least one set of active parameters.

[0121] In some embodiments, the initial multi-modal fusion facial expression intelligent perception model consists of n single-class facial expression recognition models operating in parallel, where n is the total number of facial expression categories and is a positive integer. Each single-class facial expression recognition model includes an facial expression discrimination branch. Each facial expression discrimination branch may include at least one classifier path, and each classifier path corresponds to an active parameter. When controlling at least one facial expression discrimination branch in the initial multi-modal fusion facial expression intelligent perception model to iteratively learn on at least one face region training dataset, specifically, at least one classifier path in each of the at least one facial expression discrimination branches in the initial multi-modal fusion facial expression intelligent perception model can be controlled to iteratively learn. The parameter updates between each classifier path are independent of each other.

[0122] It is easy to understand that the target multimodal fusion facial expression intelligent perception model provided in the embodiments of this disclosure can be constructed by the fusion of n deep neural network channels, and local discrimination components can be added on the basis of single-channel discrimination.

[0123] According to some embodiments, when iteratively learning any expression discrimination branch, in any iteration process, at least one face region training dataset can be input into at least one classifier path to obtain the loss value corresponding to the initial multimodal fusion expression intelligent perception model; if the loss value meets the parameter update condition, the model parameters corresponding to the initial multimodal fusion expression intelligent perception model are updated, and the next round of iteration is carried out; if the loss value does not meet the parameter update condition, the next round of iteration is carried out directly until the number of iterations reaches the iteration threshold.

[0124] In some embodiments, the parameter update condition can be that the loss value generated in the current iteration period (Epoch) is smaller than that in the previous period. That is, if the loss value generated in the current iteration period (Epoch) is smaller than that in the previous period, then all model parameters corresponding to the initial multimodal fusion facial expression intelligent perception model are updated; if the loss value generated in the current period is larger than that in the previous period, then the parameter update step is skipped.

[0125] In some embodiments, the classifier path does not specifically refer to a particular fixed classifier path. For example, at least one classifier path may include a first classifier path, a second classifier path, and a third classifier path.

[0126] According to some embodiments, Figure 4 This diagram illustrates a training flowchart for an expression discrimination branch provided in an embodiment of this disclosure. Figure 4 As shown, when the terminal acquires the raw data, it can perform data augmentation on the raw data to obtain preprocessed data 1 (overall face dataset), preprocessed data 2 (eye region dataset), and preprocessed data 3 (mouth region dataset). Then, the terminal can input preprocessed data 1 into the first classifier channel of classifier 1 to obtain the first activity parameter. The preprocessed data 2 is input into the second classifier channel of classifier 2 to obtain the second activity parameter. The preprocessed data 3 is input into the third classifier channel of classifier 3 to obtain the third activity parameter. Next, based on the first, second, and third activity parameters, the loss value and accuracy of the initial facial expression perception neural network model are determined using a loss function. If the loss value and accuracy do not improve compared to the previous iteration, the parameters are not updated in the next iteration; conversely, if the loss value and accuracy improve compared to the previous iteration, the parameters are updated in the next iteration.

[0127] In some embodiments, the loss value corresponding to the initial multimodal fusion facial expression intelligent perception model can be determined based on a loss function. This loss function can, for example, be in the form of a cross-entropy loss function.

[0128] In some embodiments, active parameters can be defined during the recognition of a specific facial expression category; that is, each facial expression category in the training set can correspond to a set of active parameters. The active parameters can take the form of... Where k is the activity parameter, i is the expression category index, and r is the facial local region index. In a single classifier path, the activity parameter is related to the classification accuracy of the trained local part-based expression classifier, and its initial value can be obtained by the following formula:

[0129]

[0130] Where P is the prediction accuracy of the prediction network (initial multimodal fusion facial expression intelligent perception model) without active parameters. j Let be the prediction accuracy of the classifier for the j-th type of expression in the validation set.

[0131] In some embodiments, the validation set refers to the dataset used to validate an expression-aware neural network model, such as a target multimodal fusion expression intelligent perception model. The expression labels in both the training set (face region training dataset) and the validation set remain consistent with the original data (initial training dataset).

[0132] In some implementations, activity parameters represent the relative importance of each facial component during the classification process. Acquiring these activity parameters must occur before final recognition; that is, a complete set of activity parameters needs to be obtained through the learning of an expression classifier—that is, through the learning of the classifier pathway—and then fed back to the classifier to obtain the target multimodal fusion expression intelligent perception model. Furthermore, by comparing the acquired activity parameters, the inherent attributes of expression data from different background groups can be studied and estimated. This includes the reliability of the expression data, the differences in expression attributes among different ethnic groups or cultures, and so on.

[0133] For example, the backbone network used in the initial multimodal fusion facial expression intelligent perception model could be a VGG-16 network pre-trained on ImageNet. The learning rate is set to 10. -3 Updated to 10 after the 300th cycle. -4 The weight decay is set to 3×10. -4 The central processing unit (CPU) in the terminal is an Intel(R) Core(TM) i7-10700K, and the graphics processing unit (GPU) is an Nvidia GeForce RTX 3080.

[0134] To investigate the impact of different population data on activity parameters, this embodiment of the disclosure uses datasets with indistinct population characteristics and datasets with distinct population characteristics for experiments. Specifically, the datasets used are the FER2013 dataset, the CK+ dataset, and the JAFFE dataset.

[0135] In some embodiments, the training data is pre-processed in three batches, each time reducing all images to contain only a single facial region. Sub-models for each classifier pathway are then trained using the pre-processed FER2013, CK+, and JAFFE datasets, respectively, to obtain initial values ​​for the activity parameters. The loss function used in the initial training does not include the activity parameters.

[0136] In some embodiments, Figure 5 The diagram shows a smoothed trend of the training accuracy and validation accuracy in the Jeffe dataset provided in this embodiment of the disclosure. Figure 6 This diagram shows a smoothed trend of the training accuracy and validation accuracy in the CK+ dataset provided in this embodiment of the disclosure. Figure 5 and Figure 6 As shown, to avoid excessive overlap in the discounting, mean smoothing is performed. Here, Epochs refers to the number of iterations, Accuracy refers to the accuracy, "ori" is the overall facial data, "mask_1" is the mouth region data, "mask_2" is the eye region data, val_acc refers to the model's accuracy on the validation set, and train_acc refers to the model's accuracy on the training set.

[0137] To verify the accuracy of the target multimodal fusion facial expression intelligent perception model trained on the FER2013 dataset, the test set recommended by the Kaggle facial expression recognition challenge can be used as a public test set, which contains 3589 samples. The corresponding test accuracy is shown in Table (1).

[0138] Table (1)

[0139]

[0140] Five-fold cross-validation was used to test the accuracy of the target multimodal fusion facial expression intelligent perception model trained by Jaffe and CK+, and the corresponding test accuracy is shown in Table (2).

[0141] Table (2)

[0142]

[0143] In some embodiments, the method provided in this disclosure does not set clear boundaries between facial expressions. Since the classification outputs of multiple channels are independent of each other, the model may generate multiple facial expression outputs during the expression recognition process. During model testing, if the output contains ground truth (GT), the classification is considered accurate. Despite its complexity, this form of multi-class output model is more similar to the human brain's efficacious reflex (FER) process compared to a single-channel model.

[0144] Table (3) shows a comparison of the accuracy of the model constructed using the embodiments of this disclosure and related schemes (Bag of words, CNN+SVM, ARM, Inception, VGG, Fine-tuned VGG) when performing facial expression classification tests using the test set provided by Kaggle. The recognition accuracy obtained by the embodiments of this disclosure is the weighted average of the recognition rates of the seven different facial expressions.

[0145] Table (3)

[0146]

[0147] The comparison shows that the overall accuracy of the model constructed in this embodiment is higher than that of related schemes. This is mainly due to the model's multi-class output, which avoids the situation where the classification results for facial expression image data at the classification boundary differ from the ground truth (GT).

[0148] It is easy to understand that when the terminal obtains the target training dataset, the terminal can use the target training dataset to train the initial multi-modal fusion expression intelligent perception model to obtain the target multi-modal fusion expression intelligent perception model.

[0149] S204, Obtain the face image to be tested;

[0150] It is easy to understand that when the terminal obtains the target multi-modal fusion expression intelligent perception model, the terminal can obtain the face image of the person to be tested.

[0151] S205, input the face image to be tested into at least one expression discrimination branch in the target multimodal fusion expression intelligent perception model to obtain the set of output prediction scores corresponding to at least one expression;

[0152] According to some embodiments, when the expression discrimination branch includes three classifier paths, the test face image is input into at least one expression discrimination branch in the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression. Figure 7 This diagram illustrates a workflow of a target multimodal fusion facial expression intelligent perception model provided in an embodiment of this disclosure. Figure 7As shown, firstly, the face image to be tested can be input into the three classifier paths corresponding to any one of the expression discrimination branches (backbone network) to obtain three probability scores for the face image. Then, based on the activity parameters of the three classifier paths, the three probability scores are weighted and concatenated to obtain the output prediction score corresponding to any one expression discrimination branch, thus obtaining a set of output prediction scores for at least one expression. It is easy to understand that when the terminal acquires the face image to be tested, it can input the face image into at least one expression discrimination branch in the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores for at least one expression.

[0153] S206, Based on the output prediction score set, determine the expression perception result corresponding to the face image to be tested.

[0154] According to some embodiments, when the target multimodal fusion facial expression intelligent perception model recognizes the test image, the final integration method of the recognition result includes, but is not limited to, feature-level fusion, decision-level fusion, etc. Feature-level fusion can be, for example, stitching together features of different scales and modalities. Decision-level fusion can be, for example, a combined decision strategy of the classification outputs of multiple classifiers.

[0155] In some embodiments, such as Figure 7 As shown, when using decision-level fusion, the final output can be obtained by weighted averaging of single-channel classification probabilities.

[0156] According to some embodiments, to ensure the capture of key facial details, the pooling degree of the expression feature extractor should not be too large when extracting expression features. Therefore, the expression feature extractor can use VGG-16 as the backbone network of the expression feature extractor, which facilitates the adjustment of parameters of each layer at any time.

[0157] In some embodiments, each classifier path corresponds to a classifier, and each classifier consists of a fully connected (FC) layer and a softmax layer, performing a binary classification task (e.g., happy expression and unhappy expression), responsible for identifying one type of expression.

[0158] In some embodiments, Softmax is an activation function that normalizes a numerical vector into a probability distribution vector, where the sum of the probabilities is 1. Softmax can be used as the final layer of a neural network for outputting multi-class classification problems.

[0159] In some embodiments, each classifier path outputs a different classification probability, and each classifier path outputs only two probabilities: the probability of whether or not it is a specific category of expression (e.g., happiness). The softmax output of each classifier path is a probability score, and each expression discrimination branch corresponds to a set of probability scores.

[0160] In some embodiments, the active parameter {k} corresponding to each classifier path in the expression discrimination branch can be activated. i r The probability scores are weighted and concatenated from at least one probability score in the probability score set, and finally fused into a single probability output. Finally, the expression perception result corresponding to the test image can be determined based on at least one fused probability output corresponding to at least one expression discrimination branch.

[0161] In some embodiments, when each expression discrimination branch includes three classifier paths, the output score of Softmax can be represented as P. a1 =(a1,a2,...a n ), P a2 = (b1, b2, ... b n ), P a3 =(c1,c2,...c n ), where P a This outputs the feature vector for each facial feature region of interest. Each classification decision transforms multi-class classification into a single binary classification. For example, the Softmax output for "happiness" is (a1, b1, c1), and the Softmax output score is represented in binary form: P a1 =(a1,1-a1), P a2 =(b1,1-b1), P a3 = (c1, 1-c1), all categories in the data except "happiness" are considered to be in the same category (not "happiness").

[0162] In some embodiments, when determining the expression perception result corresponding to the image under test based on at least one fusion probability output corresponding to at least one expression discrimination branch, if any fusion probability output does not meet the output condition, then that fusion probability output is discarded. For example... Figure 7 As shown, the fusion probability outputs corresponding to backbone network #1 and backbone network #n need to be discarded to obtain the fusion probability outputs corresponding to backbone network #2 to backbone network #k (output category #2 to output category #k). Finally, based on the fusion probability outputs corresponding to backbone network #2 to backbone network #k, the expression perception result of the original image, happy / neutral, is obtained.

[0163] In some embodiments, the classifier may also choose whether to update the activity parameters during the expression recognition process. Specific update strategies can be found in step S203, and will not be elaborated here.

[0164] According to some embodiments, by extracting active parameters from the target multimodal fusion facial expression intelligent perception model and combining them with pooling layers and key pixel locations (localization information) in the feature mapping of a specific expression (e.g., surprise), a regional activity heatmap in the facial expression image can be obtained, identifying the active regions in the facial expression image during the perception process of that category of expression. Therefore, visualization of the feature intensity of facial expression regions for a specific expression can be achieved.

[0165] In some embodiments, key pixels can be obtained by propagating the loss back to the pixel values. Key pixels can be selected only from pixels that have a large impact on the loss value; these pixels represent the visual features that the convolutional neural network can capture from the input. For ease of observation, a 576×576 pixel expression image can be used, thereby capturing more active pixels.

[0166] In some embodiments, Figure 8 This diagram illustrates an example of an active region distribution visualization provided by an embodiment of this disclosure. Figure 8 As shown, darker colors represent higher activity levels. Areas with high activity provide more evidence for the corresponding category of facial expression recognition, while areas with low activity play a less significant role in the recognition process. In this way, it's relatively easy to see where the brain-like model primarily focuses its attention on specific facial expressions. For example, the recognition of "happy" and "surprised" expressions relies more on the mouth area, while the recognition of "angry" and "sad" expressions relies more on the eye area. The recognition of "disgust," "neutral," and "fear" expressions relies almost evenly on these two areas.

[0167] It is easy to understand that when the terminal acquires the image to be tested, the terminal can determine the expression perception result corresponding to the face image to be tested based on the output prediction score set.

[0168] In summary, the method provided in this disclosure involves: acquiring an initial training dataset; preprocessing the initial training dataset to obtain a target training dataset; training an initial multi-modal fusion expression intelligent perception model using the target training dataset to obtain a target multi-modal fusion expression intelligent perception model; acquiring a test face image; inputting the test face image into at least one expression discrimination branch of the target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression; and determining the expression perception result corresponding to the test face image based on the set of output prediction scores. Therefore, it can simulate the human brain's perception process of expressions and effectively integrate local-global information of the face. Furthermore, by decomposing the multi-classification problem of expression images into multiple binary classification problems and obtaining a set of output prediction scores corresponding to at least one expression, the classification process can avoid considering the unreasonable assumption that "expression judgment has fixed boundaries," thus blurring the classification boundaries and improving the accuracy of expression perception. In addition, learning local information can significantly improve the robustness of intra-class classification tasks. Secondly, the activity parameter can effectively represent the classification significance of different facial regions of the same expression category. Furthermore, the activity parameter obtained through learning can visualize the feature intensity of facial expression regions of different expressions during the recognition process.

[0169] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0170] The following are system embodiments of this disclosure, which can be used to execute the method embodiments of this disclosure. For details not disclosed in the system embodiments of this disclosure, please refer to the method embodiments of this disclosure.

[0171] Please see Figure 9 This diagram illustrates the structure of a first type of multimodal fusion intelligent facial expression perception system provided in this embodiment. This multimodal fusion intelligent facial expression perception system can be implemented as all or part of a system through software, hardware, or a combination of both. The multimodal fusion intelligent facial expression perception system 900 includes an image acquisition unit 901, an image set acquisition unit 902, and a result determination unit 903, wherein:

[0172] Image acquisition unit 901 is used to acquire the face image of the person to be tested;

[0173] The set acquisition unit 902 is used to input the face image to be tested into the target multi-modal fusion expression intelligent perception model to obtain the set of output prediction scores corresponding to at least one expression;

[0174] The result determination unit 903 is used to determine the expression perception result corresponding to the face image to be tested based on the output prediction score set.

[0175] Optional, Figure 10 This diagram illustrates the structure of a second multimodal fusion facial expression intelligent perception system provided in an embodiment of this disclosure. Figure 10 As shown, the multimodal fusion facial expression intelligent perception system 900 also includes a dataset acquisition unit 904, a training set processing unit 905, and a model training unit 906, used to: before inputting the test face image into the target multimodal fusion facial expression intelligent perception model:

[0176] Data set acquisition unit 904 is used to acquire the initial training dataset, wherein the initial training dataset includes at least one face image;

[0177] The training set processing unit 905 is used to preprocess the initial training dataset to obtain the target training dataset;

[0178] The model training unit 906 is used to train the initial multimodal fusion facial expression intelligent perception model using the target training dataset to obtain the target multimodal fusion facial expression intelligent perception model.

[0179] Optionally, the training set processing unit 905 is used to preprocess the initial training dataset to obtain the target training dataset, specifically for:

[0180] Perform face localization on at least one face image in the initial training dataset to obtain at least one face detection and localization result corresponding to at least one face image;

[0181] Based on at least one face detection and localization result, data augmentation is performed on at least one face image to obtain at least one data-augmented face image;

[0182] Identify at least one augmented face image corresponding to the eye region dataset, mouth region dataset, and overall face dataset;

[0183] The target training dataset is determined based on the eye region dataset, mouth region dataset, and overall facial dataset.

[0184] Optionally, data augmentation methods include at least one of the following:

[0185] Small-angle random rotation;

[0186] Flip horizontally;

[0187] Translation;

[0188] Cutting.

[0189] Optionally, when the training set processing unit 905 is used to determine the eye region dataset, mouth region dataset, and overall face dataset corresponding to at least one data-augmented face image, it is specifically used for:

[0190] Rotate at least one data-augmented face image to the target angle to obtain the overall face dataset;

[0191] Mask at least one full-face image from the full-face dataset to obtain eye region datasets and mouth region datasets.

[0192] Optionally, the set acquisition unit 902 is used to input the face image to be tested into the target multi-modal fusion expression intelligent perception model to obtain the set of output prediction scores corresponding to at least one expression, specifically for:

[0193] The face image to be tested is input into at least one expression discrimination branch in the target multimodal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression.

[0194] Optionally, the expression discrimination branch includes three classifier paths. The set acquisition unit 902 is used to input the face image to be tested into at least one expression discrimination branch in the target multi-modal fusion expression intelligent perception model to obtain the output prediction score set corresponding to at least one expression. Specifically, it is used for:

[0195] The face image to be tested is input into the three classifier paths corresponding to any one of the expression discrimination branches in at least one expression discrimination branch to obtain three probability scores corresponding to the face image to be tested;

[0196] Based on the activity parameters corresponding to the three classifier paths, the three probability scores are weighted and concatenated to obtain the output prediction score corresponding to any expression discrimination branch, so as to obtain a set of output prediction scores corresponding to at least one expression.

[0197] It should be noted that the multi-modal fusion intelligent expression perception system provided in the above embodiments is only illustrated by the division of the above functional modules when executing the multi-modal fusion intelligent expression perception method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the multi-modal fusion intelligent expression perception system and the multi-modal fusion intelligent expression perception method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0198] In summary, the system provided in this embodiment acquires a face image to be tested through an image acquisition unit; the set acquisition unit inputs the face image to be tested into a target multi-modal fusion expression intelligent perception model to obtain a set of output prediction scores corresponding to at least one expression; and the result determination unit determines the expression perception result corresponding to the face image to be tested based on the set of output prediction scores. Therefore, by decomposing the multi-classification problem of expression images into multiple binary classification problems and obtaining a set of output prediction scores corresponding to at least one expression, the system can avoid considering the unreasonable assumption that "expression judgment has fixed boundaries," thus blurring the classification boundaries and improving the accuracy of expression perception.

[0199] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0200] According to embodiments of this disclosure, this disclosure also provides a terminal, a readable storage medium, and a computer program product.

[0201] Figure 11 A schematic block diagram of an example terminal 1100 that can be used to implement embodiments of the present disclosure is shown. The terminal is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The terminal may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0202] like Figure 11 As shown, terminal 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded from storage unit 1108 into random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of terminal 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0203] Multiple components in terminal 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of displays, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows terminal 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0204] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the multimodal fusion facial expression intelligent perception method. For example, in some embodiments, the multimodal fusion facial expression intelligent perception method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on terminal 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the multimodal fusion facial expression intelligent perception method described above can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform a multimodal fusion facial expression intelligent perception method by any other suitable means (e.g., by means of firmware).

[0205] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0206] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0207] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0208] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0209] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0210] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0211] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0212] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A multi-modal fusion expression intelligent perception method, characterized in that, include: Acquire the face image of the person to be tested; The face image to be tested is input into three classifier paths corresponding to any one of the expression discrimination branches in at least one expression discrimination branch to obtain three probability scores corresponding to the face image to be tested. The expression discrimination branch includes three classifier paths, each classifier path corresponds to a classifier, and each classifier consists of a fully connected layer and a Softmax layer to perform a binary classification task to identify a type of expression. The three classifier paths respectively process the overall facial information, eye region information and mouth region information of the face image to be tested. Using the active parameters obtained through iterative learning during the training process of the three classifier pathways, which reflect the contribution of different facial regions to the recognition of specific expressions, the three probability scores are weighted and concatenated to obtain the output prediction score corresponding to any expression discrimination branch. The training process uses backpropagation of stochastic gradient descent for training, and the parameter updates between each classifier pathway are independent of each other. Whether to update the model parameters is determined based on whether the loss value meets the parameter update condition. The expression perception result corresponding to the face image to be tested is determined based on the set of output prediction scores corresponding to at least one expression discrimination branch.

2. The method according to claim 1, characterized in that, Before inputting the face image to be tested into the target multimodal fusion expression intelligent perception model, the method further includes: Obtain an initial training dataset, wherein the initial training dataset includes at least one face image; The initial training dataset is preprocessed to obtain the target training dataset; The initial multimodal fusion facial expression intelligent perception model is trained using the target training dataset to obtain the target multimodal fusion facial expression intelligent perception model.

3. The method according to claim 2, characterized in that, The preprocessing of the initial training dataset to obtain the target training dataset includes: Face localization is performed on at least one face image in the initial training dataset to obtain at least one face detection and localization result corresponding to the at least one face image; Based on the at least one face detection and localization result, data augmentation is performed on the at least one face image to obtain at least one data-augmented face image; Determine the eye region dataset, mouth region dataset, and overall face dataset corresponding to the at least one data-enhanced face image; The target training dataset is determined based on the eye region dataset, the mouth region dataset, and the overall facial dataset.

4. The method according to claim 3, characterized in that, The data augmentation method includes at least one of the following: Small-angle random rotation; Flip horizontally; Translation; Cutting.

5. The method according to claim 3, characterized in that, The process of determining the eye region dataset, mouth region dataset, and overall facial dataset corresponding to the at least one data-enhanced face image includes: Rotate the at least one data-enhanced face image to the target angle to obtain the overall face dataset; At least one overall facial image from the overall facial dataset is masked to obtain the eye region dataset and the mouth region dataset.

6. A multi-modal fusion facial expression intelligent perception system, characterized in that, include: The image acquisition unit is used to acquire the face image of the person to be tested; The set acquisition unit is used to input the face image to be tested into three classifier paths corresponding to any one of the expression discrimination branches in at least one expression discrimination branch, and obtain three probability scores corresponding to the face image to be tested. The expression discrimination branch includes three classifier paths, each corresponding to a classifier. Each classifier consists of a fully connected layer and a Softmax layer, performing a binary classification task to identify one type of expression. The three classifier paths respectively process the overall facial information, eye region information, and mouth region information of the face image to be tested. Using the active parameters obtained through iterative learning during the training process of the three classifier paths, which reflect the contribution of different facial regions to the recognition of a specific expression, the three probability scores are weighted and concatenated to obtain the output prediction score corresponding to any expression discrimination branch. The training process uses backpropagation with stochastic gradient descent, and the parameter updates between each classifier path are independent. Whether to update the model parameters is determined based on whether the loss value meets the parameter update condition. The result determination unit is used to determine the expression perception result corresponding to the face image to be tested based on the set of output prediction scores corresponding to at least one expression discrimination branch.

7. A terminal, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Facial expression recognition method and device

    CN111144266A