Training Method of Facial Emotion Recognition Model, Emotion Recognition Method and Related Devices

By training a facial emotion recognition model, using emotional vocabulary collection and annotated face images, the problem that traditional methods cannot accurately identify complex facial emotions is solved, and a more accurate emotion recognition effect is achieved.

CN114202791BActive Publication Date: 2025-06-03NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111461044.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-06-03
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Traditional facial emotion recognition methods cannot accurately identify complex facial emotions because they rely only on the basic seven emotions representation models and cannot fully capture the diversity of human emotions.

Method used

A training method for facial emotion recognition model is proposed. By obtaining multiple emotional vocabulary related to emotions, integrating them to form a collection of emotional vocabulary, and collecting the face images corresponding to each emotional vocabulary for annotation, forming a training sample set. These images are then inputted to the pre-constructed initial network model for training, and the model parameters are adjusted according to the output results until the preset convergence conditions are reached.

Benefits of technology

Through this method, a facial emotion recognition model consistent with the natural language expression space is obtained, which can more accurately identify facial emotions, and the results are more in line with the true emotions perceived by humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202791B_ABST
    Figure CN114202791B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method for a facial emotion recognition model, an emotion recognition method, and related devices. It can collect some emotion-related words in natural language, then collect the face images corresponding to each emotion word, and use the emotion words to label the face images to form a training sample set. The pre-constructed initial network model is trained to obtain a facial emotion recognition model, and then the facial emotion recognition model is used to perform facial emotion recognition processing. In this way, since the facial emotion recognition model is consistent with the natural language expression space, the result of facial emotion recognition using the facial emotion recognition model is more in line with the real emotion of human natural perception, and the emotion recognition is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a method for training a facial emotion recognition model, an emotion recognition method, and related devices. Background Art

[0002] A person's facial expression can reflect a great deal of their inner emotions. Therefore, observing the visual changes of a person's face has always been the best way to identify a person's emotional state. In the fields of computer vision, human-computer interaction, and computational psychology, constructing a manually annotated face emotion dataset and using a deep network model for image classification learning is a work with research value and practical significance.

[0003] Traditional face emotion recognition methods / datasets are all based on basic emotion representation models, that is, including seven basic emotions: neutral, happy, sad, surprised, afraid, angry, disgusted, etc. However, a person's facial emotions are complex, and relying solely on these seven basic emotions cannot fully express a person's facial emotions, which will lead to inaccurate face emotion recognition. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to propose a method for training a facial emotion recognition model, an emotion recognition method, and related devices to solve or partially solve the above technical problems.

[0005] Based on the above purpose, the first aspect of the present disclosure provides a method for training a facial emotion recognition model, including:

[0006] Obtain a plurality of emotion-related words, and integrate the plurality of emotion-related words to form an emotion word set;

[0007] Collect face images corresponding to each emotion word in the emotion word set, annotate the corresponding face images with the emotion words, and use the annotated face images as a training sample set, where one emotion word corresponds to a plurality of collected face images;

[0008] Input the face images in the training sample set into a pre-constructed initial network model in sequence for training, and adjust the parameters of the initial network model according to the output result of each round of training of the initial network model and the emotion words corresponding to the annotations in the training sample set; when the initial network model reaches a preset convergence condition, use the initial network model as a facial emotion recognition model for facial emotion recognition.

[0009] In some exemplary embodiments, the obtaining a plurality of emotion-related words and integrating the plurality of emotion-related words to form an emotion word set includes:

[0010] Screen multiple candidate words related to emotions from the language vocabulary library;

[0011] Select multiple of the candidate words through crowdsourcing, remove the candidate words that cannot express the corresponding emotions, and use the remaining candidate words as emotion words;

[0012] Integrate the emotion words to form an emotion word set.

[0013] In some exemplary embodiments, the selecting multiple of the candidate words through crowdsourcing, removing the candidate words that cannot express the corresponding emotions, and using the remaining candidate words as emotion words includes:

[0014] Generate corresponding survey data according to each of the candidate words, publish the survey data through the network for respondents to receive the survey data transmitted by the network through a terminal device, and vote on whether the survey data can express the corresponding emotion to generate corresponding voting information;

[0015] Receive the voting information feedback by each respondent through the terminal device, remove the candidate words in the voting information that cannot express the corresponding emotions, and use the remaining candidate words as emotion words.

[0016] In some exemplary embodiments, the collecting face images corresponding to each of the emotion words in the emotion word set, using the emotion words to label the corresponding face images, and using the labeled face images as a training sample set includes:

[0017] Use each of the emotion words in the emotion word set as a search tag to search, obtain multiple face images corresponding to the search tag, use the emotion words corresponding to the search tag to label the multiple face images, and store the labeled face images in a database;

[0018] Use a pre-constructed face expression encoding model to filter the labeled face images in the database to obtain filtered face images;

[0019] Randomly sample the filtered face images, and output the sampling results to the display terminal of the judge for the judge to judge whether the emotion words corresponding to the filtered face images are matched through the display terminal to generate a judgment result;

[0020] Receive the judgment results feedback by each judge through the display terminal, calculate the proportion of the judgment results that are matched, delete the filtered face images corresponding to the proportion less than a predetermined ratio, and use the remaining filtered face images and the corresponding labeled emotion words as a training sample set.

[0021] In some exemplary embodiments, a pre - constructed facial expression encoding model is used to filter the labeled facial images in the database to obtain filtered facial images, including:

[0022] Sequentially determine corresponding target emotion words from multiple emotion words;

[0023] Obtain multiple labeled facial images corresponding to the target emotion words from the database as images to be filtered;

[0024] Use a pre - constructed facial expression encoding module to perform clustering processing on the images to be filtered to obtain at least one clustering result;

[0025] Retain the clustering result with the largest number, remove other clustering results, and use the clustering result with the largest number as the filtered facial image.

[0026] In some exemplary embodiments, inputting the facial images in the training sample set into a pre - constructed initial network model for training processing, and adjusting the parameters of the initial network model according to the emotion words corresponding to the labels in the training sample set specifically includes:

[0027] Pre - construct an initial network model with an input layer, multiple hidden layers, and an output layer based on a convolutional operator deep neural network;

[0028] Input the facial images in the training sample set into the input layer of the initial network model. The input layer pre - processes the input facial images and sends them to the hidden layer. After analysis by multiple hidden layers, an analysis result is generated and sent to the output layer. The output layer processes the analysis result to generate prediction probability values for various emotions, selects the target emotion corresponding to the maximum prediction probability value from the prediction probability values of various emotions, and the output layer outputs the target emotion;

[0029] Calculate a loss function according to the difference between the target emotion and the emotion words corresponding to the labels of the input facial images, adjust the parameters of each layer of the initial network model according to the loss function, and obtain the next facial image from the training sample set and input it into the initial network model for training processing.

[0030] In some exemplary embodiments, embed a pre - obtained similarity matrix between different emotions in the first hidden layer of the multiple hidden layers;

[0031] After the input layer pre - processes the input facial images and sends them to the hidden layer, and after analysis by multiple hidden layers, an analysis result is generated and sent to the output layer, including:

[0032] The input layer preprocesses the input face image and then sends it to the first hidden layer;

[0033] The first hidden layer extracts emotional features from the input face image according to the similarity matrix, sends the extracted emotional features to the remaining hidden layers for emotional analysis in sequence, and the last hidden layer sends the analysis result to the output layer.

[0034] In some exemplary embodiments, the input layer and multiple hidden layers of the initial network model are composed of two twin VGGNets in parallel.

[0035] Based on the same inventive concept, a second aspect of the present disclosure provides an emotional recognition method for a facial emotion recognition model, including:

[0036] Receiving a facial image to be recognized, and inputting the facial image into the facial emotion recognition model obtained by using the training method of the facial emotion recognition model in the first aspect;

[0037] Using the facial emotion recognition model to perform emotional analysis processing on the facial image to be recognized, determining the pending probability values of each emotional word corresponding to the facial image to be recognized, and screening out the emotional words whose pending probability values exceed the set threshold as the output emotions for output.

[0038] In some exemplary embodiments, the facial emotion recognition model includes: an input layer, multiple hidden layers and an output layer, and a similarity matrix between different emotions obtained in advance is embedded in the first hidden layer among the multiple hidden layers;

[0039] The facial emotion recognition model performs emotional analysis processing on the facial image to be recognized, determines the pending probability values of each emotional word corresponding to the facial image to be recognized, and screens out the emotional words whose pending probability values exceed the set threshold as the output emotions for output, including:

[0040] The facial image to be recognized is input to the input layer, the input layer preprocesses the facial image to be recognized, and inputs the preprocessed facial image to the first hidden layer;

[0041] The first hidden layer extracts emotional features from the preprocessed facial image according to the similarity matrix, sends the extracted emotional features to the remaining hidden layers for emotional analysis in sequence to obtain the pending probability values of each emotional word corresponding to the facial image to be recognized, and the last hidden layer sends each of the pending probability values to the output layer;

[0042] The output layer screens out emotional words with pending probability values exceeding a set threshold from each of the pending probability values, and outputs them as output emotions.

[0043] Based on the same inventive concept, a third aspect of the present disclosure provides a training device for a facial emotion recognition model, including:

[0044] A vocabulary acquisition module, configured to acquire a plurality of words related to emotions, correspond the plurality of words to corresponding emotions respectively to form a plurality of emotional words, and integrate the plurality of emotional words to form an emotional word set;

[0045] A face image collection module, configured to collect face images corresponding to each of the emotional words in the emotional word set, label the corresponding face images with the emotional words, and use the labeled face images as a training sample set, wherein one emotional word corresponds to a plurality of the collected face images;

[0046] A training processing module, configured to sequentially input the face images in the training sample set into a pre-constructed initial network model for training processing, and adjust the parameters of the initial network model according to the output result of each round of training of the initial network model and the emotional words corresponding to the labels in the training sample set; when the initial network model reaches a preset convergence condition, use the initial network model as a facial emotion recognition model for facial emotion recognition.

[0047] Based on the same inventive concept, a fourth aspect of the present disclosure provides an emotion recognition device for a facial emotion recognition model, including:

[0048] A receiving module, configured to receive a face image to be recognized, and input the face image into the facial emotion recognition model obtained by using the training method of the facial emotion recognition model in the first aspect;

[0049] An emotion recognition module, configured to perform emotion analysis processing on the face image to be recognized by using the facial emotion recognition model, determine the pending probability values of the respective emotional words corresponding to the face image to be recognized, and screen out emotional words with pending probability values exceeding a set threshold as output emotions for output.

[0050] Based on the same inventive concept, a fifth aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the training method of the facial emotion recognition model as described in the first aspect, or the emotion recognition method of the facial emotion recognition model as described in the second aspect.

[0051] Based on the same inventive concept, a sixth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the training method of the facial emotion recognition model described in the first aspect or the emotion recognition method of the facial emotion recognition model described in the second aspect.

[0052] As can be seen from the above, the training method, emotion recognition method and related devices of the facial emotion recognition model provided by the present disclosure can collect some emotion-related words in natural language, then collect the face images corresponding to each emotion word, and use the emotion words to label the face images to form a training sample set, and train the pre-constructed initial network model to obtain a facial emotion recognition model. In this way, since the facial emotion recognition model is consistent with the natural language expression space, the result of facial emotion recognition using the facial emotion recognition model is more in line with the real emotion of human natural perception, and the emotion recognition is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 Schematic diagram of the application scenario of the exemplary embodiment of the present disclosure;

[0055] Figure 2 Flowchart of the training method of the facial emotion recognition model of the embodiment of the present disclosure;

[0056] Figure 3 Flowchart of the emotion recognition method of the facial emotion recognition model of the embodiment of the present disclosure;

[0057] Figure 4 Block diagram of the structure of the training device of the facial emotion recognition model of the embodiment of the present disclosure;

[0058] Figure 5 Block diagram of the structure of the emotion recognition device of the facial emotion recognition model of the embodiment of the present disclosure;

[0059] Figure 6 Schematic diagram of the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and thereby implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0061] The face emotion representation models in the related art can be classified into the following three categories according to the methods and quantities of representing emotions:

[0062] 1. Basic emotion representation model: Proposed by researchers in the field of psychology in the late last century, it generally includes several common basic emotion categories, such as: neutral, happy, sad, surprised, afraid, angry, disgusted, etc.

[0063] 2. Composite emotion representation model: On the basis of the basic emotion representation model, some researchers proposed using two different basic emotions to depict more detailed composite emotions, such as "both happy and surprised", "both afraid and disgusted", etc.

[0064] 3. Multi-dimensional representation model: Based on several emotion expression dimensions summarized by psychologists, such as pleasure degree, arousal degree, etc., any emotion type can be expressed as a set composed of continuous values of each dimension.

[0065] Based on the above three types of descriptions, the corresponding disadvantages include:

[0066] 1. The basic emotion model can only be used to depict several sparse emotion categories, which is far from the ever-changing emotional states of human beings' true inner hearts.

[0067] 2. Although the composite emotion model makes up for the disadvantage of the small number of emotions depicted by the basic emotion model to a certain extent, it is still limited by several basic emotion definitions itself, and not all basic emotions can be reasonably combined to form new emotion instances.

[0068] 3. Although the multi-dimensional representation model can theoretically represent any emotion type, it must itself master the multi-dimensional scores of the target emotion; in cognitive psychology and emotion theory research, there is currently no unanimously agreed emotion evaluation method in the academic community. Therefore, setting multi-dimensional scores for any emotion type still lacks a reasonable reference standard.

[0069] In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0070] For the convenience of understanding, the following explains the nouns involved in the embodiments of the present disclosure:

[0071] Artificial Neural Networks (ANNs): Build practical artificial neural network models according to the principles of biological neural networks and the needs of practical applications, design corresponding learning algorithms, simulate certain intelligent activities of the human brain, and then implement them technically to solve practical problems.

[0072] VGGNet (Visual Geometry Group Net): In 2014, researchers from the Visual Geometry Group at the University of Oxford and Google DeepMind developed a new deep convolutional neural network: VGGNet explored the relationship between the depth of the convolutional neural network and its performance, successfully constructed a convolutional neural network with 16 - 19 layers deep, proved that increasing the depth of the network can affect the final performance of the network to a certain extent, significantly reduce the error rate, and at the same time has strong scalability and very good generalization when migrated to other image data.

[0073] Emotional words, which represent words that can express human psychological feelings.

[0074] Crowdsourcing method: A practice where a company or organization outsources work tasks that were previously performed by employees to a non - specific (and usually large) public network in a free and voluntary form.

[0075] The solution of the present disclosure aims to provide a training method for a facial emotion recognition model, an emotion recognition method and related devices, which can obtain a facial emotion recognition model consistent with the natural language expression space, and the result of emotion recognition is more in line with the real emotion of human natural perception, and the emotion recognition is more accurate.

[0076] Reference Figure 1, which is a schematic diagram of the application scenario of the training method and emotion recognition method of the facial emotion recognition model provided by the embodiments of the present disclosure. The application scenario includes a terminal device 101, a server 102, and a data storage system 103. Among them, the terminal device 101, the server 102, and the data storage system 103 can all be connected through a wired or wireless communication network. The terminal device 101 includes, but is not limited to, a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a personal digital assistant (PDA), or other electronic devices capable of implementing the above functions. The server 102 and the data storage system 103 can both be independent physical servers, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0077] The server 102 is used to provide an emotion recognition service to the user of the terminal device 101. A client for communicating with the server 102 is installed in the terminal device 101, and the user can input a facial image to be recognized through the client. The user sends the facial image to be recognized to the server 102 through the client. The server 102 inputs the facial image to be recognized into a pre-trained facial emotion recognition model, obtains the emotion recognition result output by the facial emotion recognition model, and then sends the emotion recognition result to the client. The client displays the processed emotion recognition result to the user to complete the task of emotion recognition for the facial image to be recognized.

[0078] A large amount of training data is stored in the data storage system 103, and the training data includes face images labeled with corresponding emotion words. The server 102 can train the initial network model based on the large amount of training data, so that the trained facial emotion recognition model can perform emotion recognition on facial images, making the emotion recognition result more in line with the real emotions of human natural perception and the emotion recognition more accurate.

[0079] Next, in combination with Figure 1 the application scenario, the training method and emotion recognition method of the facial emotion recognition model according to the exemplary embodiments of the present disclosure will be described. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0080] Referring to Figure 2 , the training method of the facial emotion recognition model according to the embodiments of the present disclosure includes the following steps:

[0081] Step 201: Obtain multiple emotion-related vocabulary words, and integrate the multiple emotion-related vocabulary words to form an emotion vocabulary set.

[0082] Specifically, this step 201 includes:

[0083] Step 2011: Screen multiple candidate words related to emotions from a language vocabulary library.

[0084] Among them, the language vocabulary library can be at least one of the following: Chinese library, English library, Japanese library, Korean library, Russian library, French library, German library, and Arabic library, etc. Specifically, it can be selected according to the corresponding environmental requirements and the corresponding local language. Select each emotion-related word from them and use these words as candidate words.

[0085] Step 2012: Select multiple candidate words through a crowdsourcing method, remove the candidate words that cannot express the corresponding emotion, and use the remaining candidate words as emotion vocabulary words.

[0086] Specifically, since there are many candidate words related to emotions screened from the language vocabulary library, and there are some words that are not related to emotions among them, these irrelevant words need to be deleted. However, it is difficult to determine which of these words are irrelevant words. Therefore, the crowdsourcing method is used to let the public make a choice, so as to know the emotional cognitions of people of different ages, different industries, and different personalities towards these candidate words.

[0087] Among them, when the crowdsourcing method is specifically implemented, the specific process includes:

[0088] Step 20121: Generate corresponding survey data according to each candidate word, and publish the survey data through the network for the surveyors to receive the survey data transmitted by the network through a terminal device and vote on whether the survey data can express the corresponding emotion to generate corresponding voting information.

[0089] For example, the survey data of one of the candidate words is: "Are you extremely angry" a word that can correctly express the emotion of anger? If so, please confirm the corresponding level representing anger "from low to high, respectively, level one, level two, and level three"; if not, please directly click "No".

[0090] Step 20122: Receive the voting information feedback by each surveyor through the terminal device, remove the candidate words in the voting information that cannot express the corresponding emotion, and use the remaining candidate words as emotion vocabulary words.

[0091] In specific implementation, the generated survey data can be distributed to terminals corresponding to various groups of people of different ages, industries, and genders. The people can choose whether to participate in the survey by themselves. If they participate in the survey, they can directly vote according to the prompts. In this way, the voting information feedback by each investigator through the terminal device will be received, the voting results of the voting information will be summarized, the number of votes for "yes" corresponding to each candidate word will be counted, and the candidate words with the number of votes lower than the minimum threshold (for example, 10 votes) will be deleted. The remaining candidate words are the emotion words that can represent the corresponding emotions and have been unanimously confirmed by the public vote.

[0092] In this way, the emotion words obtained through the crowdsourcing method are more in line with the emotional cognition of the public.

[0093] Step 2013, integrate the emotion words to form an emotion word set.

[0094] In specific implementation, the obtained multiple emotion words are arranged and integrated in alphabetical order of the first letter, or in the order of the number of strokes, or in other set ways to form an emotion word set.

[0095] Step 202, collect face images corresponding to each emotion word in the emotion word set, label the corresponding face images with the emotion words, and use the labeled face images as a training sample set, where one emotion word corresponds to multiple collected face images.

[0096] In specific implementation, this step includes:

[0097] Step 2021, use each emotion word in the emotion word set as a search tag to search, obtain multiple face images corresponding to the search tag, label the multiple face images with the emotion word corresponding to the search tag, and store the labeled face images in the database.

[0098] Step 2022, perform filtering processing on the labeled face images in the database by using a pre-constructed face expression encoding model to obtain filtered face images.

[0099] Among them, the face expression encoding model can be pre-constructed by using a neural network.

[0100] The specific process includes:

[0101] Step 20221, sequentially determine the corresponding target emotion words from multiple emotion words.

[0102] Step 20222: Obtain multiple of the labeled face images corresponding to the target emotion vocabulary from the database as the images to be filtered.

[0103] Step 20223: Use the pre-constructed face expression encoding module to perform clustering processing on the images to be filtered, and obtain at least one clustering result.

[0104] Step 20224: Retain the clustering result with the largest number, remove other clustering results, and use the clustering result with the largest number as the filtered face images.

[0105] Through the above solution, it is possible to use the face expression encoding model to automatically filter the set of face images under each emotion vocabulary, remove the noisy face images that are inconsistent with the expressions of most face images, and then obtain the filtered face images composed of most face images. In this way, denoising and filtering processing is performed on the multiple labeled face images of each emotion vocabulary, so that the obtained filtered face images can better represent the corresponding emotion vocabulary and have stronger representativeness.

[0106] Among them, each emotion vocabulary is filtered once. If there are K emotion vocabularies, it is necessary to use the face expression encoding model for denoising and filtering K times.

[0107] Step 2023: Randomly sample the filtered face images, and output the sampling results to the display terminal of the judge, so that the judge can judge whether the emotion vocabulary corresponding to the filtered face images matches through the display terminal, and generate a judgment result.

[0108] Step 2024: Receive the judgment results feedback by each judge through the display terminal, calculate the proportion of the judgment results that are matched, delete the filtered face images corresponding to the proportion less than the predetermined ratio, and use the remaining filtered face images and the corresponding labeled emotion vocabulary as the training sample set.

[0109] Through the above solution, the face images under each emotion vocabulary are automatically filtered to remove the noisy images that are inconsistent with the expressions of most images. Then, manual random spot checks are carried out. The filtered face images under each emotion label are randomly sampled, and a group of double-blind testers are used to judge whether the sampled face images match the corresponding emotion vocabulary, and the images with poor manual judgment consistency are eliminated. In this way, the finally obtained face images and the corresponding labeled emotion vocabulary can be used as training samples.

[0110] Step 203: Input the face images in the training sample set into the pre-constructed initial network model in turn for training processing, and adjust the parameters of the initial network model according to the output results of each round of training of the initial network model and the emotion vocabulary corresponding to the training sample set.

[0111] In specific implementation, it includes:

[0112] Step 2031: Based on the convolutional operator deep neural network, an initial network model with an input layer, multiple hidden layers, and an output layer is pre-constructed.

[0113] In specific implementation, a pre-obtained similarity matrix between different emotions is pre-embedded in the first hidden layer among the multiple hidden layers.

[0114] Among them, the input layer and multiple hidden layers of the initial network model are composed of two twin VGGNets in parallel.

[0115] Step 2032: Input the face images in the training sample set into the input layer of the initial network model. After the input layer preprocesses the input face images, it sends them to the hidden layer. After being analyzed by multiple hidden layers, an analysis result is generated and sent to the output layer. The output layer processes the analysis result to generate prediction probability values for various emotions, selects the target emotion corresponding to the maximum prediction probability value from the prediction probability values of various emotions, and the output layer outputs the target emotion.

[0116] In specific implementation, the input layer preprocesses the input face images and sends them to the first hidden layer; the first hidden layer extracts emotion features from the input face images according to the similarity matrix, and sends the extracted emotion features to the remaining hidden layers for emotion analysis in sequence. Calculate the 512-dimensional difference features obtained by the two twin VGGNets as the analysis result, and the last hidden layer sends the analysis result to the output layer composed of several fully connected network layers.

[0117] Further process using several fully connected network layers to obtain K (where K represents the number of emotion types) -dimensional prediction probability values. Select the target emotion corresponding to the maximum prediction probability value from the K -dimensional prediction probability values as the result output.

[0118] Step 2033: Calculate the loss function according to the difference between the target emotion and the emotion vocabulary labeled corresponding to the input face image, adjust the parameters of each layer of the initial network model according to the loss function, and obtain the next face image from the training sample set and input it into the initial network model for training processing.

[0119] In specific implementation, the loss function used is the cross-entropy loss function. The initial network model is supervised and constrained using the cross-entropy loss function, and the parameters of each layer of the initial network model are adjusted by means of backpropagation. The above process is continuously repeated using each labeled face image in the training samples, and thus the initial network model is continuously trained to make the cross-entropy loss function continuously converge.

[0120] Step 204, when the initial network model reaches a preset convergence condition, use the initial network model as a facial emotion recognition model for facial emotion recognition.

[0121] In specific implementation, the preset convergence condition can be that all training is completed, or the loss value obtained by the corresponding cross-entropy loss function is less than or equal to a preset convergence value. Among them, the smaller the loss value, the higher the accuracy of emotion recognition.

[0122] Through the solution described in the above embodiments, some emotion-related emotion words in natural language can be collected, then the face images corresponding to each emotion word are collected, and the face images are labeled with emotion words to form a training sample set, and the pre-constructed initial network model is trained to obtain a facial emotion recognition model. In this way, since the facial emotion recognition model is consistent with the natural language expression space, the result of facial emotion recognition using the facial emotion recognition model is more in line with the real emotions perceived by humans naturally, and the emotion recognition is more accurate.

[0123] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0124] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0125] Based on the same inventive concept, the embodiments of the present disclosure also provide an emotion recognition method for a facial emotion recognition model, refer to Figure 3 , the emotion recognition method for the facial emotion recognition model includes the following steps:

[0126] Step 301: Receive a facial image to be recognized, and input the facial image into a facial emotion recognition model, where the facial emotion recognition model is obtained by using the training method of the facial emotion recognition model in any of the above embodiments.

[0127] In specific implementation, the facial emotion recognition model includes an input layer, multiple hidden layers, and an output layer. The first hidden layer among the multiple hidden layers embeds a similarity matrix obtained in advance between different emotions. In this way, the similarity matrix can be used to extract emotion features from the facial image.

[0128] Step 302: Use the facial emotion recognition model to perform emotion analysis processing on the facial image to be recognized, determine the pending probability values of each emotion vocabulary corresponding to the facial image to be recognized, and screen out the emotion vocabulary whose pending probability value exceeds a set threshold as the output emotion for output.

[0129] In specific implementation, it includes:

[0130] Step 3021: Input the facial image to be recognized into the input layer, and the input layer preprocesses the facial image to be recognized, and inputs the preprocessed facial image into the first hidden layer.

[0131] Step 3022: The first hidden layer extracts emotion features from the preprocessed facial image according to the similarity matrix, and sends the extracted emotion features to the remaining hidden layers for emotion analysis in sequence to obtain the pending probability values of each emotion vocabulary corresponding to the facial image to be recognized. The last hidden layer sends each of the pending probability values to the output layer.

[0132] Step 3023: The output layer screens out the emotion vocabulary whose pending probability value exceeds a set threshold from each of the pending probability values as the output emotion for output.

[0133] Through the above solution, by using the pre-trained facial emotion recognition model, it is possible to perform an emotion recognition process of real emotions that is more in line with human natural perception, accurately recognize the emotions corresponding to the facial image, and improve the emotion recognition effect.

[0134] Reference Figure 4 , based on the same inventive concept as the embodiments of the training method of any of the above facial emotion recognition models, the embodiments of the present disclosure also provide a training device for a facial emotion recognition model, including:

[0135] A vocabulary acquisition module 501, configured to acquire multiple vocabulary related to emotions, respectively correspond the multiple vocabulary to the corresponding emotions to form multiple emotion vocabulary, and integrate the multiple emotion vocabulary to form an emotion vocabulary set;

[0136] The face image collection module 502 is configured to collect face images corresponding to each of the emotion vocabulary in the emotion vocabulary set, label the corresponding face images with the emotion vocabulary, and use the labeled face images as a training sample set, wherein, one emotion vocabulary corresponds to multiple collected face images;

[0137] The training processing module 503 is configured to sequentially input the face images in the training sample set into a pre-constructed initial network model for training processing, and adjust the parameters of the initial network model according to the output result of each round of training of the initial network model and the emotion vocabulary corresponding to the label in the training sample set; when the initial network model reaches a preset convergence condition, use the initial network model as a facial emotion recognition model for facial emotion recognition.

[0138] In some alternative embodiments, the vocabulary acquisition module 501 includes:

[0139] A screening unit configured to screen multiple candidate words related to emotions from a language vocabulary library;

[0140] A crowdsourcing selection unit configured to select multiple candidate words by crowdsourcing, remove the candidate words that cannot express the corresponding emotions, and use the remaining candidate words as emotion vocabulary;

[0141] An integration unit configured to integrate the emotion vocabulary to form an emotion vocabulary set.

[0142] In some alternative embodiments, the crowdsourcing selection unit is specifically configured to:

[0143] Generate corresponding survey data according to each candidate word, publish the survey data through the network for surveyors to receive the survey data transmitted by the network through a terminal device, and vote on whether the survey data can express the corresponding emotion to generate corresponding voting information; receive the voting information feedback by each surveyor through the terminal device, remove the candidate words in the voting information that cannot express the corresponding emotion, and use the remaining candidate words as emotion vocabulary.

[0144] In some alternative embodiments, the face image collection module 502 includes:

[0145] A search unit configured to use each emotion vocabulary in the emotion vocabulary set as a search tag for searching, obtain multiple face images corresponding to the search tag, label the multiple face images with the emotion vocabulary corresponding to the search tag, and store the labeled face images in a database;

[0146] A filtering unit, configured to perform filtering processing on the labeled face images in the database by using a pre-constructed face expression encoding model to obtain filtered face images;

[0147] A judging unit, configured to randomly sample the filtered face images, output the sampling results to the display end of the judge, so that the judge can judge whether the emotional words corresponding to the filtered face images match through the display end, and generate a judging result;

[0148] A calculation unit, configured to receive the judging results fed back by each judge through the display end, calculate the proportion of the judging results that are matched, delete the filtered face images corresponding to the proportion less than a predetermined ratio, and use the remaining filtered face images and the corresponding labeled emotional words as a training sample set.

[0149] In some alternative embodiments, the filtering unit is specifically configured to:

[0150] Sequentially determine corresponding target emotional words from multiple emotional words; obtain multiple labeled face images corresponding to the target emotional words from the database as images to be filtered; perform clustering processing on the images to be filtered by using a pre-constructed face expression encoding module to obtain at least one clustering result; retain the clustering result with the largest number, remove other clustering results, and use the clustering result with the largest number as the filtered face image.

[0151] In some alternative embodiments, the training processing module 503 specifically includes:

[0152] A construction unit, configured to pre-construct an initial network model with an input layer, multiple hidden layers, and an output layer based on a convolutional operator deep neural network;

[0153] A training processing unit, configured to input the face images in the training sample set into the input layer of the initial network model, the input layer preprocesses the input face images and then sends them to the hidden layer, after analysis by multiple hidden layers, generates an analysis result and sends it to the output layer, the output layer processes the analysis result to generate prediction probability values of various emotions, selects the target emotion corresponding to the maximum prediction probability value from the prediction probability values of various emotions, and the output layer outputs the target emotion;

[0154] A training adjustment unit, configured to calculate a loss function according to the difference between the target emotion and the emotional words corresponding to the input face images, adjust the parameters of each layer of the initial network model according to the loss function, and obtain the next face image from the training sample set and input it into the initial network model for training processing.

[0155] In some alternative embodiments, the building unit is further configured to embed a similarity matrix between pre-obtained different emotions in the first hidden layer among the multiple hidden layers;

[0156] The training processing unit is further configured to:

[0157] The input layer preprocesses the input face image and sends it to the first hidden layer; the first hidden layer extracts emotion features from the input face image according to the similarity matrix, and sends the extracted emotion features to the remaining hidden layers for emotion analysis in sequence, and the last hidden layer sends the analysis result to the output layer.

[0158] In some alternative embodiments, the input layer and the multiple hidden layers of the initial network model are composed of two twin VGGNets in parallel.

[0159] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0160] The device in the above embodiment is used to implement the training method of the corresponding facial emotion recognition model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0161] Reference Figure 5 , based on the same inventive concept as the emotion recognition method embodiment of any of the above facial emotion recognition models, the embodiment of the present disclosure further provides an emotion recognition device for a facial emotion recognition model, including:

[0162] A receiving module 601, configured to receive a face image to be recognized and input the face image into the facial emotion recognition model obtained by using the training method of the facial emotion recognition model in the above embodiment;

[0163] An emotion recognition module 602, configured to perform emotion analysis processing on the face image to be recognized by using the facial emotion recognition model, determine the pending probability values of each emotion vocabulary corresponding to the face image to be recognized, and screen out the emotion vocabulary whose pending probability value exceeds a set threshold as the output emotion for output.

[0164] In some alternative embodiments, the facial emotion recognition model includes: an input layer, multiple hidden layers, and an output layer, and a similarity matrix between pre-obtained different emotions is embedded in the first hidden layer among the multiple hidden layers;

[0165] The emotion recognition module 602 is further configured to:

[0166] The face image to be recognized is input into the input layer, and the input layer preprocesses the face image to be recognized and inputs the preprocessed face image into the first hidden layer; the first hidden layer extracts emotion features from the preprocessed face image according to the similarity matrix, and sends the extracted emotion features to the remaining hidden layers for emotion analysis in sequence to obtain the pending probability values of each emotion word corresponding to the face image to be recognized. The last hidden layer sends each of the pending probability values to the output layer; the output layer screens out the emotion words whose pending probability values exceed the set threshold from each of the pending probability values as the output emotions for output.

[0167] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0168] The device in the above embodiment is used to implement the emotion recognition method of the corresponding face emotion recognition model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0169] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the training method of the face emotion recognition model in any of the above embodiments, or the emotion recognition method of the face emotion recognition model in any of the above embodiments.

[0170] Figure 6 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 710, a memory 720, an input / output interface 730, a communication interface 740, and a bus 750. Among them, the processor 710, the memory 720, the input / output interface 730, and the communication interface 740 are communicatively connected to each other inside the device through the bus 750.

[0171] The processor 710 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0172] The memory 720 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 720 can store the operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 720 and called and executed by the processor 710.

[0173] The input / output interface 730 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input devices can include keyboards, mice, touchscreens, microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.

[0174] The communication interface 740 is used to connect to the communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0175] The bus 750 includes a path for transmitting information between various components of the device (such as the processor 710, the memory 720, the input / output interface 730, and the communication interface 740).

[0176] It should be noted that although the above device only shows the processor 710, the memory 720, the input / output interface 730, the communication interface 740, and the bus 750, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification and do not necessarily include all the components shown in the figure.

[0177] The electronic device in the above embodiments is used to implement the training method of the corresponding facial emotion recognition model in any of the foregoing embodiments, or the emotion recognition method of the corresponding facial emotion recognition model in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0178] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the training method of the facial emotion recognition model as described in any one of the above embodiments, or the emotion recognition method of the facial emotion recognition model as described in any one of the above embodiments.

[0179] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0180] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the training method of the facial emotion recognition model as described in any one of the above embodiments, or the emotion recognition method of the facial emotion recognition model as described in any one of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0181] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.

[0182] Additionally, for simplicity of explanation and discussion, and so as not to render the embodiments of the present disclosure difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid rendering the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.

[0183] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations thereof will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0184] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A training method for a facial emotion recognition model, characterized in that, it includes: Obtain multiple emotion-related words, and integrate the multiple emotion-related words to form an emotion word set; Collect face images corresponding to each emotion word in the emotion word set, label the corresponding face images with the emotion words, and use the labeled face images as a training sample set. Among them, one emotion word corresponds to multiple collected face images; Input the face images in the training sample set into a pre-constructed initial network model in sequence for training, and adjust the parameters of the initial network model according to the output result of each round of training of the initial network model and the emotion words corresponding to the labels in the training sample set; When the initial network model reaches a preset convergence condition, use the initial network model as a facial emotion recognition model for facial emotion recognition; The collecting face images corresponding to each emotion word in the emotion word set, labeling the corresponding face images with the emotion words, and using the labeled face images as a training sample set includes: Use each emotion word in the emotion word set as a search tag for searching, obtain multiple face images corresponding to the search tag, label the multiple face images with the emotion word corresponding to the search tag, and store the labeled face images in a database; Use a pre-constructed face expression encoding model to filter the labeled face images in the database to obtain filtered face images; Randomly sample the filtered face images, and output the sampling result to the display terminal of the judge for the judge to judge whether the emotion words corresponding to the filtered face images are matched through the display terminal, and generate a judgment result; Receive the judgment results feedback by each judge through the display terminal, calculate the proportion of the judgment results that are matched, delete the filtered face images corresponding to the proportion less than a predetermined ratio, and use the remaining filtered face images and the corresponding labeled emotion words as a training sample set.

2. The training method for a facial emotion recognition model according to claim 1, characterized in that, The obtaining multiple emotion-related words and integrating the multiple emotion-related words to form an emotion word set includes: Screen multiple candidate words related to emotions from a language vocabulary library; Select the multiple candidate words by means of crowdsourcing, remove the candidate words that cannot express the corresponding emotions, and use the remaining candidate words as emotion words; Integrate the emotion words to form an emotion word set.

3. The training method for a facial emotion recognition model according to claim 2, characterized in that, The selecting the multiple candidate words by means of crowdsourcing, removing the candidate words that cannot express the corresponding emotions, and using the remaining candidate words as emotion words includes: Generate corresponding survey data for each of the candidate words, and publish the survey data through the network for respondents to receive the survey data transmitted over the network via a terminal device and vote on whether the survey data can express the corresponding emotion, generating corresponding voting information; Receive the voting information feedback by each respondent via a terminal device, remove the candidate words in the voting information that cannot express the corresponding emotion, and use the remaining candidate words as emotion words.

4. The method for training a facial emotion recognition model according to claim 1, characterized in that, Filter the labeled face images in the database by using a pre-constructed face expression encoding model to obtain filtered face images, including: Sequentially determine corresponding target emotion words from multiple emotion words; Obtain multiple labeled face images corresponding to the target emotion word from the database as images to be filtered; Perform clustering processing on the images to be filtered by using a pre-constructed face expression encoding module to obtain at least one clustering result; Retain the clustering result with the largest number, remove other clustering results, and use the clustering result with the largest number as the filtered face image.

5. The method for training a facial emotion recognition model according to claim 1, characterized in that, Input the face images in the training sample set into a pre-constructed initial network model for training processing, and adjust the parameters of the initial network model according to the emotion words corresponding to the labels in the training sample set, specifically including: Pre-construct an initial network model with an input layer, multiple hidden layers, and an output layer based on a convolutional operator deep neural network; Input the face images in the training sample set into the input layer of the initial network model. The input layer preprocesses the input face images and then sends them to the hidden layer. After analysis by multiple hidden layers, an analysis result is generated and sent to the output layer. The output layer processes the analysis result to generate predicted probability values for various emotions, selects the target emotion corresponding to the largest predicted probability value from the predicted probability values for various emotions, and the output layer outputs the target emotion; Calculate a loss function according to the difference between the target emotion and the emotion word corresponding to the label of the input face image, adjust the parameters of each layer of the initial network model according to the loss function, and obtain the next face image from the training sample set and input it into the initial network model for training processing.

6. The method for training a facial emotion recognition model according to claim 5, characterized in that, Embed a pre-obtained similarity matrix between different emotions in the first hidden layer of the multiple hidden layers; The input layer preprocesses the input face images and then sends them to the hidden layer. After analysis by multiple hidden layers, an analysis result is generated and sent to the output layer, including: The input layer preprocesses the input face images and then sends them to the first hidden layer; The first hidden layer extracts emotional features from the input face image according to the similarity matrix, and sends the extracted emotional features to the remaining hidden layers for emotional analysis in sequence. The last hidden layer sends the analysis result to the output layer.

7. The training method of the facial emotion recognition model according to claim 5 or 6, wherein, The input layer and multiple hidden layers of the initial network model are composed of two twin VGGNets in parallel.

8. An emotion recognition method of a facial emotion recognition model, wherein, including: Receiving a face image to be recognized, and inputting the face image into the facial emotion recognition model obtained by using the training method of the facial emotion recognition model according to any one of claims 1-7; Using the facial emotion recognition model to perform emotion analysis processing on the face image to be recognized, determining the pending probability values of each emotion word corresponding to the face image to be recognized, and screening out the emotion words whose pending probability values exceed the set threshold as the output emotion for output.

9. The emotion recognition method of the facial emotion recognition model according to claim 8, wherein, The facial emotion recognition model includes: an input layer, multiple hidden layers and an output layer. The first hidden layer among the multiple hidden layers embeds a similarity matrix between different emotions obtained in advance; The facial emotion recognition model performs emotion analysis processing on the face image to be recognized, determines the pending probability values of each emotion word corresponding to the face image to be recognized, and screens out the emotion words whose pending probability values exceed the set threshold as the output emotion for output, including: Inputting the face image to be recognized into the input layer, and the input layer preprocesses the face image to be recognized and inputs the preprocessed face image into the first hidden layer; The first hidden layer extracts emotional features from the preprocessed face image according to the similarity matrix, and sends the extracted emotional features to the remaining hidden layers for emotional analysis in sequence to obtain the pending probability values of each emotion word corresponding to the face image to be recognized. The last hidden layer sends each of the pending probability values to the output layer; The output layer screens out the emotion words whose pending probability values exceed the set threshold from each of the pending probability values as the output emotion for output.

10. A training device for a facial emotion recognition model, wherein, including: A vocabulary acquisition module, configured to acquire multiple words related to emotions, correspond the multiple words to the corresponding emotions respectively to form multiple emotion words, and integrate the multiple emotion words to form an emotion word set; A face image collection module, configured to collect face images corresponding to each emotion word in the emotion word set, label the corresponding face images with the emotion words, and use the labeled face images as a training sample set, wherein one emotion word corresponds to multiple collected face images; A training processing module, configured to sequentially input the face images in the training sample set into a pre-constructed initial network model for training processing, and adjust the parameters of the initial network model according to the output result of each round of training of the initial network model and the corresponding labeled emotion words in the training sample set; when the initial network model reaches a preset convergence condition, use the initial network model as a face emotion recognition model for face emotion recognition; The face image collection module includes: A search unit, configured to use each of the emotion words in the emotion word set as a search tag to perform a search, obtain multiple face images corresponding to the search tag, label the multiple face images with the emotion words corresponding to the search tag, and store the labeled face images in a database; A filtering unit, configured to perform filtering processing on the labeled face images in the database by using a pre-constructed face expression encoding model to obtain filtered face images; A judging unit, configured to randomly sample the filtered face images, output the sampling result to the display end of a judge, so that the judge can judge whether the emotion words corresponding to the filtered face images match through the display end, and generate a judging result; A calculation unit, configured to receive the judging results fed back by each judge through the display end, calculate the proportion of the judging results that are matched, delete the filtered face images corresponding to the proportion less than a predetermined ratio, and use the remaining filtered face images and the corresponding labeled emotion words as a training sample set.

11. An emotion recognition device for a face emotion recognition model, characterized in that, it includes: A receiving module, configured to receive a face image to be recognized, and input the face image into the face emotion recognition model obtained by using the training method of the face emotion recognition model according to any one of claims 1 to 7; An emotion recognition module, configured to perform emotion analysis processing on the face image to be recognized by using the face emotion recognition model, determine the pending probability values of the respective emotion words corresponding to the face image to be recognized, and screen out the emotion words whose pending probability values exceed a set threshold as output emotions for output.

12. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, it implements the training method of the face emotion recognition model according to any one of claims 1 to 7, or the emotion recognition method of the face emotion recognition model according to claim 8 or 9.

13. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium stores computer instructions, characterized in that, the computer instructions are used to cause a computer to execute the training method of the face emotion recognition model according to any one of claims 1 to 7, or the emotion recognition method of the face emotion recognition model according to claim 8 or 9.

Citation Information

Patent Citations

  • Emotion recognition method and device, computer equipment and storage medium

    CN109784153A