Facial Expression Recognition Method, Apparatus, Electronic Device, and Storage Medium
Through self-supervised learning methods, the expression recognition network is trained, and expression extraction is optimized using expression feature similarity, and combined with the expression classification network, the problem of low facial expression recognition accuracy is solved, achieving efficient and accurate expression recognition.
Patent Information
- Application Number
- CN202210100976.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-01-27
AI Technical Summary
In the prior art, facial expression recognition accuracy is low and cannot be implemented on a large scale. This is mainly due to the subjectivity of emotions and the difficulty of collecting expression data, which makes manual labeling time-consuming and labor-intensive.
The self-supervised learning method is adopted to extract image features through the encoding network, and the expression feature similarity between the reference image and the first image is trained using the expression feature similarity between the reference image and the second image, and the expression feature similarity between the reference image and the second image, and the expression feature similarity is trained. Expression recognition is combined with the expression classification network, reducing manual annotation, and improving the recognition accuracy.
It effectively improves the accuracy and reliability of facial expression recognition, reduces the cost of manual labeling, and achieves efficient expression recognition.
Smart Images

Figure CN114495230B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and in particular, to a method, device, electronic device and storage medium for facial expression recognition. Background Art
[0002] As an important part of human-computer interaction, emotion recognition can improve the user experience during the interaction process. In recent years, with the development of deep learning technology, facial expression recognition has also developed accordingly.
[0003] However, due to the subjectivity of emotions themselves and the similarity between different emotions, it is difficult to collect facial expression data, and manual annotation is time-consuming and laborious, resulting in low accuracy of current facial expression recognition and inability to be used on a large scale. Summary of the Invention
[0004] The present invention provides a method, device, electronic device and storage medium for facial expression recognition, so as to solve the defect that the accuracy of facial expression recognition in the prior art is low and it cannot be used on a large scale.
[0005] The present invention provides a method for facial expression recognition, including:
[0006] Determining image features of a face image to be recognized based on an encoding network;
[0007] Performing facial expression feature extraction on the image features based on a facial expression extraction network to obtain facial expression features;
[0008] Performing facial expression recognition on the facial expression features based on a facial expression classification network;
[0009] The facial expression extraction network is trained based on a first facial expression feature similarity between a reference image and a first image, and a second facial expression feature similarity between the reference image and a second image. The reference image, the first image and the second image correspond to the same face in a sample face video, and the time interval between the reference image and the first image is less than the time interval between the reference image and the second image.
[0010] According to a method for facial expression recognition provided by the present invention, the encoding network and the facial expression extraction network are trained based on the following steps:
[0011] Determining an initial extraction network, where the initial extraction network includes an initial encoding network and an initial facial expression extraction branch;
[0012] Based on the initial extraction network, respectively determining predicted facial expression features of the reference image, the first image and the second image;
[0013] Determine the first expression feature similarity based on the predicted expression features of the reference image and the first image, and determine the second expression feature similarity based on the predicted expression features of the reference image and the second image;
[0014] Determine an expression extraction loss based on the first expression feature similarity and the second expression feature similarity;
[0015] Train the initial extraction network based on the expression extraction loss to obtain the encoding network and the expression extraction network.
[0016] According to an expression recognition method provided by the present invention, the determining an expression extraction loss based on the first expression feature similarity and the second expression feature similarity includes:
[0017] Determine an expression extraction loss based on the difference between the first expression feature similarity and the second expression feature similarity.
[0018] According to an expression recognition method provided by the present invention, the initial extraction network further includes an initial identity extraction branch;
[0019] The training the initial extraction network based on the expression extraction loss includes:
[0020] Determine the predicted image features of a sample image based on the initial encoding network;
[0021] Perform identity feature extraction on the predicted image features of the sample image based on the initial identity extraction branch to obtain the predicted identity features of the sample image;
[0022] Determine a first identity feature similarity based on the predicted identity features of sample images corresponding to the same face, and determine a second identity feature similarity based on the predicted identity features of sample images corresponding to different faces;
[0023] Determine an identity extraction loss based on the first identity feature similarity and the second identity feature similarity;
[0024] Train the initial extraction network based on the expression extraction loss and the identity extraction loss.
[0025] According to an expression recognition method provided by the present invention, the initial extraction network further includes an initial decoding network;
[0026] The training the initial extraction network based on the expression extraction loss and the identity extraction loss includes:
[0027] Based on the initial decoding network, perform feature decoding on the predicted expression features and predicted identity features of the sample image to obtain the reconstructed image features of the sample image;
[0028] Based on the predicted image features and reconstructed image features of the sample image, determine the decoding loss;
[0029] Based on the expression extraction loss, the identity extraction loss, and the decoding loss, train the initial extraction network.
[0030] According to an expression recognition method provided by the present invention, the training of the initial extraction network based on the expression extraction loss, the identity extraction loss, and the decoding loss includes:
[0031] Based on the expression extraction loss, the identity extraction loss, and the decoding loss, perform the first-stage training on the initial extraction network to obtain the initial extraction network after the first-stage training as the first extraction network;
[0032] Based on the predicted identity features and predicted expression features of the sample images corresponding to different faces, generate perturbation features;
[0033] Based on the first decoding network in the first extraction network, perform feature decoding on the perturbation features to obtain perturbation reconstruction features;
[0034] Based on the first expression extraction branch in the first extraction network, perform expression feature extraction on the perturbation reconstruction features to obtain the reconstructed expression features of the perturbation features, and / or, based on the first identity extraction branch in the first extraction network, perform identity feature extraction on the perturbation reconstruction features to obtain the reconstructed identity features of the perturbation features;
[0035] Based on the reconstructed expression features and predicted expression features of the perturbation features, determine the expression perturbation loss, and / or, based on the reconstructed identity features and predicted identity features of the perturbation features, determine the identity perturbation loss;
[0036] Based on the expression extraction loss and the expression perturbation loss, and / or, based on the identity extraction loss and the identity perturbation loss, train the first expression extraction branch and / or the first identity extraction branch.
[0037] According to an expression recognition method provided by the present invention, the determining of the expression perturbation loss based on the reconstructed expression features and predicted expression features of the perturbation features, and / or, the determining of the identity perturbation loss based on the reconstructed identity features and predicted identity features of the perturbation features includes:
[0038] Determine an expression perturbation loss based on the difference between the reconstructed expression features and the predicted expression features of the perturbation features; and / or,
[0039] Determine an identity perturbation loss based on the difference between the reconstructed identity features and the predicted identity features of the perturbation features.
[0040] According to an expression recognition method provided by the present invention, after training the initial extraction network based on the expression extraction loss, it further includes:
[0041] Based on the sample images carrying expression classification labels, jointly fine-tune the trained initial encoding network, the initial expression extraction branch, and the initial classification network, and use the fine-tuned initial encoding network, the initial expression extraction branch, and the initial classification network as the encoding network, the expression extraction network, and the expression classification network respectively.
[0042] According to an expression recognition method provided by the present invention, the reference image, the first image, and the second image are determined based on the following steps:
[0043] Determine a sample face video, where each frame image in the sample face video corresponds to the same face;
[0044] Extract two adjacent frame images from the image frame sequence of the sample face video as the reference image and the first image, and extract a frame image with the farthest time interval from the reference image as the second image.
[0045] The present invention also provides an expression recognition device, including:
[0046] An image feature determination unit, configured to determine the image features of a face image to be recognized based on an encoding network;
[0047] An expression feature extraction unit, configured to perform expression feature extraction on the image features based on an expression extraction network to obtain expression features;
[0048] An expression recognition unit, configured to perform expression recognition on the expression features based on an expression classification network;
[0049] The expression extraction network is trained based on the first expression feature similarity between the reference image and the first image, and the second expression feature similarity between the reference image and the second image. The reference image, the first image, and the second image correspond to the same face in the sample face video, and the time interval between the reference image and the first image is less than the time interval between the reference image and the second image.
[0050] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned expression recognition methods are implemented.
[0051] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned expression recognition methods are implemented.
[0052] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned expression recognition methods are implemented.
[0053] The expression recognition method, device, electronic device, and storage medium provided by the present invention train an expression extraction network based on the first expression feature similarity between a reference image and a first image, and the second expression feature similarity between the reference image and a second image, so that the expression extraction network can better distinguish the differences between expression features under the same facial expression change of a person, and the extracted expression features focus more on expression information; and perform expression recognition based on the expression features, avoiding the collection of a large amount of facial expression data, reducing the manual annotation cost, and effectively improving the accuracy and reliability of the facial expression recognition result. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 is one of the flow diagrams of the expression recognition method provided by the present invention;
[0056] Figure 2 is one of the flow diagrams of the training method of the initial extraction network provided by the present invention;
[0057] Figure 3 is the second flow diagram of the training method of the initial extraction network provided by the present invention;
[0058] Figure 4 is the third flow diagram of the training method of the initial extraction network provided by the present invention;
[0059] Figure 5 is the fourth flow diagram of the training method of the initial extraction network provided by the present invention;
[0060] Figure 6 It is a schematic flowchart of the method for determining the reference image, the first image, and the second image provided by the present invention;
[0061] Figure 7 It is a schematic diagram of the data flow of the self-supervised learning network provided by the present invention;
[0062] Figure 8 It is a schematic structural diagram of the facial expression recognition device provided by the present invention;
[0063] Figure 9 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0064] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the scope of protection of the present invention.
[0065] In the existing facial expression recognition technical solution, a self-supervised learning solution is used to learn a representation information and then transfer it to the supervised data of human face expressions. Finally, the learned representation may focus more on the identity information and has a weak representation ability for the expression information. Therefore, it is not conducive to the downstream facial expression recognition task and results in a low accuracy of facial expression recognition.
[0066] Based on this, an embodiment of the present invention provides a facial expression recognition method, Figure 1 It is a schematic flowchart of the facial expression recognition method provided by the present invention. As Figure 1 shown, the method includes:
[0067] Step 110, based on an encoding network, determine the image features of the face image to be recognized.
[0068] Specifically, the face image includes the facial expressions of a human face. The face image to be recognized is the image that needs to perform human face expression recognition. The face image to be recognized can be obtained by an image acquisition device such as a mobile phone or a camera, or can be obtained by web crawling or downloading. The embodiments of the present invention do not make specific limitations on this.
[0069] Here, the encoding network can be a convolutional neural network or a residual network. The encoding network can be used to encode the face image to be recognized, that is, compress the face image into the image feature space to obtain the image features. This process is also called feature extraction. The method of feature extraction can be implemented by any one of the face image feature extraction algorithms in the prior art, and will not be elaborated here in this embodiment.
[0070] The image features obtained by feature extraction can represent the global or local information of the face image to be recognized. For example, they can specifically include face identity features and face expression features.
[0071] Step 120: Based on the expression extraction network, perform expression feature extraction on the image features to obtain expression features.
[0072] The expression extraction network is trained based on the first expression feature similarity between the reference image and the first image, and the second expression feature similarity between the reference image and the second image. The reference image, the first image, and the second image correspond to the same face in the sample face video. The time interval between the reference image and the first image is less than the time interval between the reference image and the second image.
[0073] Specifically, the expression extraction network is used to extract the expression features in the image features, so as to obtain an expression representation vector that can represent the face image, that is, the expression features. The expression features can be represented as a vector and output as a result.
[0074] Before performing step 120, the expression extraction network can also be pre-trained. Specifically, the network can be trained by the following method:
[0075] First, collect a large number of sample face videos containing faces, and select three images from the sample face videos of the same face, namely the reference image, the first image, and the second image. Among them, the time interval between the reference image and the first image is less than the time interval between the reference image and the second image. For example, the reference image is at the 1st second in the sample face video, the first image is at the 2nd second in the sample face video, and the second image is at the 5th second in the sample face video. The time interval between the reference image and the first image is 1 second, which is less than the time interval of 4 seconds between the reference image and the second image. That is to say, compared with the reference image and the second image, the reference image and the first image are closer in the sample face video.
[0076] Then, input the reference image features, the first image features, and the second image features encoded by the encoding network into the initial expression extraction network for training. During the training process, the initial expression extraction network can amplify and learn the common features of the expression features of the reference image and the first image, and at the same time amplify and learn the differential features of the expression features of the reference image and the second image. The expression extraction network trained in this way can better distinguish the differences between the expression features under the expression changes of the same person.
[0077] Among them, the first facial expression feature similarity between the reference image and the first image is the similarity between the facial expression features corresponding to the reference image and the facial expression features corresponding to the first image. Since the time interval between the reference image and the first image is relatively short, the higher the facial expression feature similarity between the two, the more the corresponding facial expression features of the two can reflect the common features of the facial expression features of the reference image and the first image.
[0078] The second facial expression feature similarity between the reference image and the second image is the similarity between the facial expression features corresponding to the reference image and the facial expression features corresponding to the second image. Since the time interval between the reference image and the second image is relatively long, the lower the facial expression feature similarity between the two, the more the corresponding facial expression features of the two can reflect the differential features of the facial expression features of the reference image and the second image.
[0079] Step 130, perform facial expression recognition on the facial expression features based on the facial expression classification network.
[0080] Specifically, the facial expression classification network is used to perform classification prediction on the extracted facial expression features of a human face. The facial expression features of the human face are input into the facial expression classification network, and facial expression classification recognition is performed through the facial expression classification network.
[0081] Among them, the facial expression classification network can be independently trained for the initial classification network based on the facial expression features of each sample image. Alternatively, it can also be obtained by fine-tuning the encoding network and the trained facial expression extraction network in combination with the initial classification network. During the fine-tuning process, only a small number of sample images with facial expression classification labels are required to achieve stable training loss of the facial expression classification network. Compared with directly training using a large number of sample images with facial expression classification labels of human faces, this method avoids collecting a large amount of facial expression data of human faces, reduces the manual annotation cost, and effectively improves the accuracy and reliability of the facial expression recognition result of human faces.
[0082] The method provided by the embodiment of the present invention trains a facial expression extraction network based on the first facial expression feature similarity between the reference image and the first image, and the second facial expression feature similarity between the reference image and the second image, so that the facial expression extraction network can better distinguish the differences between the facial expression features under the same facial expression change of the same human face, and the extracted facial expression features focus more on the facial expression information; and perform facial expression recognition based on the facial expression features, avoiding collecting a large amount of facial expression data of human faces, reducing the manual annotation cost, and effectively improving the accuracy and reliability of the facial expression recognition result of human faces.
[0083] Based on the above embodiment, Figure 2 is one of the schematic flowcharts of the training method of the initial extraction network provided by the present invention, as Figure 2 shown, the training method includes:
[0084] Step 210: Determine an initial extraction network, which includes an initial encoding network and an initial expression extraction branch;
[0085] Step 220: Based on the initial extraction network, determine the predicted expression features of the reference image, the first image, and the second image respectively;
[0086] Step 230: Based on the predicted expression features of the reference image and the first image, determine the first expression feature similarity, and based on the predicted expression features of the reference image and the second image, determine the second expression feature similarity;
[0087] Step 240: Based on the first expression feature similarity and the second expression feature similarity, determine the expression extraction loss;
[0088] Step 250: Based on the expression extraction loss, train the initial extraction network to obtain an encoding network and an expression extraction network.
[0089] Specifically, the reference image, the first image, and the second image are all expression sample images selected from a sample face video of the same person. The initial encoding network is used to extract the image features of the expression sample images, and the initial expression extraction branch is used to perform expression feature extraction based on the extracted image features, so as to obtain the predicted expression features of the expression sample images. The initial encoding network and the initial expression extraction branch constitute the initial extraction network.
[0090] After inputting the reference image, the first image, and the second image into the initial encoding network in the initial extraction network respectively, predicted image features are obtained; then, after inputting the predicted image features into the initial expression extraction branch in the initial extraction network respectively, the predicted expression features of the reference image, the first image, and the second image are obtained.
[0091] The first expression feature similarity can be determined according to the predicted expression features of the reference image and the first image, and specifically can be obtained by using a feature similarity calculation function. For example, the Euclidean distance function, the cosine similarity function, or the correlation coefficient, etc. The embodiments of the present invention do not make specific limitations in this regard.
[0092] Correspondingly, the second expression feature similarity can be determined according to the predicted expression features of the reference image and the second image, and similarly can be obtained by using a feature similarity calculation function.
[0093] The expression extraction loss determined by the first expression feature similarity and the second expression feature similarity is used to maximize the first expression feature similarity and minimize the second expression feature similarity. For example, the expression extraction loss can constrain the difference or ratio between the first expression feature similarity and the second expression feature similarity to be maximized.
[0094] The determined expression extraction loss enables the encoding network and the expression extraction network to learn as many common features of two frames of images with small expression changes and differential features of two frames of images with large expression changes as possible during the training process, so that the expression features extracted by the trained expression extraction network can fully represent the expression information in the face image to be recognized.
[0095] The method provided by the embodiment of the present invention determines the expression extraction loss through the first expression feature similarity and the second expression feature similarity, and trains the initial extraction network based on the expression extraction loss, so that the trained encoding network and expression extraction network can better distinguish the expression differences of the same face, and improve the accuracy and reliability of expression feature extraction.
[0096] Based on any of the above embodiments, step 240 specifically includes:
[0097] Determine the expression extraction loss based on the difference between the first expression feature similarity and the second expression feature similarity.
[0098] Specifically, in order to maximize the first expression feature similarity and minimize the second expression feature similarity, the expression extraction loss can be determined according to the difference between the first expression feature similarity and the second expression feature similarity. The greater the difference between the two, the smaller the expression extraction loss; the smaller the difference between the two, the greater the expression extraction loss.
[0099] Further, for each sample face video, multiple expression sample images can be selected, and then an expression sample image group can be formed according to the multiple expression sample images. Each group contains three images, namely a reference image, a first image, and a second image. Then the loss of the initial extraction network can be determined by the expression extraction losses respectively determined by the expression sample image group.
[0100] In one embodiment, during the training of the initial extraction network, each time K video segments are randomly sampled, then the loss of the initial extraction network can be expressed in the following form:
[0101]
[0102] In the formula, That is, the first expression feature similarity between the reference image and the first image in the first group of expression sample images, That is, the second expression feature similarity between the reference image and the second image in the first group of expression sample images; That is, the first expression feature similarity between the reference image and the first image in the second group of expression sample images, That is, the second expression feature similarity between the reference image and the second image in the second group of expression sample images;
[0103] γ is a hyperparameter that needs to be manually adjusted during the training process. The ultimate optimization goal of network training is to minimize the loss function L Exp .
[0104] Among them, and can be specifically expressed in the following form:
[0105]
[0106]
[0107]
[0108]
[0109] In the formula, x k1 to x k5 respectively represent the expression sample images selected from the Kth video segment. In the first group of expression sample images, the x in k1 can represent the reference image, and x k2 can represent the first image, the x in k5 can represent the second image; in the second group of expression sample images, the x in k5 can represent the reference image, and x k4 can represent the first image, the x in k1 can represent the second image. f(x) represents the initial encoding network, and g EXP (x) represents the initial expression extraction branch. Here, the initial expression extraction branch can include a projection module for realizing the conversion of the input features from the image coding dimension to the expression feature dimension.
[0110] Based on any of the above embodiments, the initial extraction network further includes an initial identity extraction branch, Figure 3 is the second flow diagram of the training method of the initial extraction network provided by the present invention. As Figure 3 shown, in step 250, the initial extraction network is trained based on the expression extraction loss, including:
[0111] Step 251, based on the initial encoding network, determine the predicted image features of the sample image;
[0112] Step 252, based on the initial identity extraction branch, perform identity feature extraction on the predicted image features of the sample image to obtain the predicted identity features of the sample image;
[0113] Step 253: Determine the first identity feature similarity based on the predicted identity features of the sample images corresponding to the same face, and determine the second identity feature similarity based on the predicted identity features of the sample images corresponding to different faces;
[0114] Step 254: Determine the identity extraction loss based on the first identity feature similarity and the second identity feature similarity;
[0115] Step 255: Train the initial extraction network based on the expression extraction loss and the identity extraction loss.
[0116] Specifically, since the image features extracted by the encoding network can include expression features and identity features, in order to make the expression features extracted by the expression extraction network focus more on the expression information and minimize the influence of the identity features, while extracting the expression features, the identity features are extracted to decouple the expression features and the identity features. Therefore, the initial extraction network further includes an initial identity extraction branch for extracting features of the identity features.
[0117] The sample images here can be video frames selected from the sample face videos. After inputting the sample images into the initial encoding network in the initial extraction network, the predicted image features are obtained; then, after inputting the predicted image features into the initial identity extraction branch in the initial extraction network, the predicted identity features of the sample images are obtained.
[0118] It can be understood that the initial expression extraction branch and the initial identity extraction branch share the initial encoding part in the initial extraction network, and realize the information sharing between the expression features and the identity features by sharing the image feature extraction part.
[0119] The first identity feature similarity is the similarity between the predicted identity features of the sample images corresponding to the same face. Since the sample images correspond to the same face, the higher the first identity feature similarity, the more the predicted identity features of the sample images can reflect the common features of the identity features of the sample images.
[0120] The second identity feature similarity is the similarity between the predicted identity features of the sample images corresponding to different faces. Since the sample images correspond to different faces, the lower the second identity feature similarity, the more the predicted identity features of the sample images can reflect the differential features of the identity features of the sample images.
[0121] The first identity feature similarity and the second identity feature similarity can be determined according to the corresponding predicted identity features, and specifically can be obtained by using a feature similarity calculation function. For example, the Euclidean distance function, the cosine similarity function or the correlation coefficient, etc. The embodiments of the present invention do not make specific limitations on this.
[0122] The identity extraction loss determined by the first identity feature similarity and the second identity feature similarity is used to maximize the first identity feature similarity and minimize the second identity feature similarity. The determined identity extraction loss enables the encoding network and the identity extraction network to learn as many common features of sample images with the same identity and differential features of sample images with different identities as possible during the training process.
[0123] Subsequently, the initial extraction network can be trained by combining the expression extraction loss and the identity extraction loss, so that the expression features extracted by the trained expression extraction network can fully represent the expression information in the face image to be recognized, and the identity features extracted by the trained identity extraction network can fully represent the identity information in the face image to be recognized, enabling the expression features and the identity features to be fully decoupled.
[0124] The method provided by the embodiment of the present invention determines the identity extraction loss based on the first identity feature similarity and the second identity feature similarity; and trains the initial extraction network based on the expression extraction loss and the identity extraction loss, enabling the expression features and the identity features to be fully decoupled.
[0125] In one embodiment, during the network training process, K video segments are randomly sampled each time. For each video segment, 5 sample images are selected, so a total of 5K video frames are obtained. Then the loss of the initial identity extraction branch can be expressed in the following form:
[0126]
[0127] where x ki and x kj respectively represent the sample images corresponding to the same face, x pn and x qn respectively represent the sample images corresponding to different faces, f(x) represents the initial encoding network, and g ID (x) represents the initial identity extraction branch. Here, the initial identity extraction branch may include a projection module for realizing the conversion of the input features from the image encoding dimension to the expression feature dimension. The ultimate optimization goal of the identity extraction branch is to minimize the loss function L ID .
[0128] Therefore, the loss function of the initial extraction network can be jointly determined by the expression loss and the identity loss, and can be expressed in the following form:
[0129] L = L ID + λ1L EXP
[0130] where λ1 is a hyperparameter that can be adjusted according to specific experiments.
[0131] Based on any of the above embodiments, the initial extraction network further includes an initial decoding network. Figure 4 It is the third schematic flow chart of the training method of the initial extraction network provided by the present invention. As Figure 4 shown, step 255 specifically includes:
[0132] Step 2551: Based on the initial decoding network, perform feature decoding on the predicted expression feature and predicted identity feature of the sample image to obtain the reconstructed image feature of the sample image.
[0133] Step 2552: Determine the decoding loss based on the predicted image feature and reconstructed image feature of the sample image.
[0134] Step 2553: Train the initial extraction network based on the expression extraction loss, identity extraction loss, and decoding loss.
[0135] Specifically, in order to further verify the independence and integrity of the expression feature extraction and identity feature extraction, the initial extraction network further includes an initial decoding network, and the initial decoding network is used to perform feature decoding on the predicted expression feature and predicted identity feature of the sample image.
[0136] Specifically, any frame of the sample image can be sent into the initial encoding network to obtain the predicted image feature of the sample image.
[0137] After the predicted image feature passes through the initial expression extraction branch and the initial identity extraction branch respectively, the decoupled predicted expression feature and predicted identity feature are obtained.
[0138] After fusing the predicted expression feature and the predicted identity feature, they are sent into the initial decoding network to perform feature decoding on the fused predicted expression feature and predicted identity feature to obtain the reconstructed image feature of the sample image.
[0139] On this basis, the decoding loss can be determined based on the predicted image feature and the reconstructed image feature. Here, the decoding loss can characterize the difference between the predicted image feature and the reconstructed image feature of the sample image. The greater the difference, the greater the decoding loss; the smaller the difference, the smaller the decoding loss. In order to ensure the independence and integrity of the decoupling of the image feature into the expression feature and the identity feature, the decoding loss can be constrained to be the smallest.
[0140] Furthermore, the difference between the predicted image feature and the reconstructed image feature of the sample image can be used to characterize the difference between the two.
[0141] Thus, according to the emotion extraction loss, identity extraction loss, and decoding loss, the initial extraction network can be trained so that the trained initial extraction network can simultaneously learn emotion features and identity features, and can restore to the image features before decoupling through feature decoding based on the decoupled emotion features and identity features.
[0142] In one embodiment, the loss of the initial decoding network can be expressed in the following form:
[0143]
[0144] In the formula, x kn represents any frame of sample image, and h(x) represents the initial decoding network. The ultimate optimization goal of the initial decoding network is to minimize the loss function L Decoder .
[0145] Therefore, the loss function of the initial extraction network can be jointly determined by the emotion loss, identity loss, and decoding loss, and can be expressed in the following form:
[0146] L = L ID + λ1L EXP + λ2L Decoder
[0147] In the formula, both λ1 and λ2 are hyperparameters and can be adjusted according to specific experiments.
[0148] The method provided by the embodiments of the present invention, by combining the emotion loss, identity loss, and decoding loss, enables the trained initial extraction network to simultaneously learn emotion features and identity features, and can restore to the image features before decoupling through feature decoding based on the decoupled emotion features and identity features, further improving the accuracy and reliability of emotion feature extraction.
[0149] Based on any of the above embodiments, Figure 5 is the fourth flowchart of the training method of the initial extraction network provided by the present invention. As Figure 5 shown, step 2553 specifically includes:
[0150] Based on the emotion extraction loss, identity extraction loss, and decoding loss, perform the first-stage training on the initial extraction network to obtain the initial extraction network after the first-stage training as the first extraction network;
[0151] Generate perturbation features based on the predicted identity features and predicted emotion features of the sample images corresponding to different faces;
[0152] Based on the first decoding network in the first extraction network, perform feature decoding on the perturbation features to obtain perturbation reconstruction features;
[0153] Based on the first expression extraction branch in the first extraction network, perform expression feature extraction on the perturbed reconstruction features to obtain the reconstructed expression features of the perturbed features, and / or based on the first identity extraction branch in the first extraction network, perform identity feature extraction on the perturbed reconstruction features to obtain the reconstructed identity features of the perturbed features;
[0154] Based on the reconstructed expression features and predicted expression features of the perturbed features, determine the expression perturbation loss, and / or based on the reconstructed identity features and predicted identity features of the perturbed features, determine the identity perturbation loss;
[0155] Based on the expression extraction loss and the expression perturbation loss, and / or based on the identity extraction loss and the identity perturbation loss, train the first expression extraction branch and / or the first identity extraction branch.
[0156] Specifically, after the initial extraction network has been trained for a certain period of time, at this time, the initial face expression branch and the initial face identity branch have already been able to learn better representation information respectively, and the initial decoding network also has strong decoding ability. At this time, the initial extraction network has completed the first stage of training and obtained the first extraction network.
[0157] In order to enable the expression extraction branch and the identity extraction branch to learn more focused feature information of their respective branches and make the learned feature information more robust, continue to perform perturbation training on the expression extraction branch and the identity extraction branch based on the first extraction network. Specifically, perturbations can be added to the two branches to each other.
[0158] For the expression extraction branch, the feature vector of each frame of image represents the expression information extracted from that frame. Therefore, even if identity information perturbation is added to the expression extraction branch, this branch should still be able to extract expression information; similarly, for the identity extraction branch, if expression information perturbation is added, the identity branch can still extract identity information.
[0159] Send the sample images corresponding to different faces into the first extraction network respectively to obtain the predicted identity features and predicted expression features. Here, the predicted identity features are the identity features before adding perturbations, and the predicted expression features are the expression features before adding perturbations.
[0160] Based on the predicted identity features and predicted expression features, generate perturbed features. The perturbed features can be the fusion features of the predicted identity features and predicted expression features, such as the fusion feature obtained by adding the two.
[0161] It should be noted that for the expression extraction branch, the perturbed features can represent adding identity information perturbation to the expression extraction branch; and for the identity extraction branch, the perturbed features can represent adding expression information perturbation to the identity extraction branch.
[0162] Send the perturbation features into the first decoding network in the first extraction network to perform feature decoding on the perturbation features and obtain the reconstructed perturbation features.
[0163] For the expression extraction branch, the reconstructed perturbation features can be sent into the first expression extraction branch in the first extraction network to perform expression feature extraction on the reconstructed perturbation features, and the reconstructed expression features of the perturbation features are obtained. The expression perturbation loss is used to characterize the difference between the reconstructed expression features of the perturbation features and the predicted expression features, that is, the difference between the reconstructed expression features and the expression features before adding the perturbation. The greater the difference between the two, the greater the expression perturbation loss; the smaller the difference between the two, the smaller the expression perturbation loss. In order to constrain the reconstructed expression features to approximate the expression features before adding the perturbation, so that the expression extraction branch can learn expression information that is not affected by identity information and improve the robustness of this branch, the expression perturbation loss can be constrained to be minimized.
[0164] Similarly, for the identity extraction branch, the reconstructed perturbation features can be sent into the first identity extraction branch in the first extraction network to perform identity feature extraction on the reconstructed perturbation features, and the reconstructed identity features of the perturbation features are obtained. The identity perturbation loss is used to characterize the difference between the reconstructed identity features of the perturbation features and the predicted identity features, that is, the difference between the reconstructed identity features and the identity features before adding the perturbation. The greater the difference between the two, the greater the identity perturbation loss; the smaller the difference between the two, the smaller the identity perturbation loss. In order to constrain the reconstructed identity features to approximate the identity features before adding the perturbation, so that the identity extraction branch can learn identity information that is not affected by expression information and improve the robustness of this branch, the identity perturbation loss can be constrained to be minimized.
[0165] Therefore, the first expression extraction branch can be trained by combining the expression extraction loss and the expression perturbation loss; the first identity extraction branch can be trained by combining the identity extraction loss and the identity perturbation loss.
[0166] It should be noted that when training the first expression extraction branch and / or training the first identity extraction branch, the parameters of the first decoding network can be fixed and not participate in the training, and only the parameters of the first expression extraction branch and / or the first identity extraction branch are adjusted.
[0167] The method provided by the embodiment of the present invention, by continuing to perform perturbation training on the expression extraction branch and / or the identity extraction branch on the basis of the first extraction network, adding perturbations to each other for the two branches respectively, enables the feature information learned by the expression extraction branch and the identity extraction branch to be more robust.
[0168] Based on any of the above embodiments, determining the expression perturbation loss based on the reconstructed expression features and predicted expression features of the perturbation features, and / or determining the identity perturbation loss based on the reconstructed identity features and predicted identity features of the perturbation features, includes:
[0169] Determining the expression perturbation loss based on the difference between the reconstructed expression features and predicted expression features of the perturbation features; and / or determining the identity perturbation loss based on the difference between the reconstructed identity features and predicted identity features of the perturbation features.
[0170] Specifically, the expression perturbation loss is used to characterize the difference between the reconstructed expression features and predicted expression features of the perturbation features. Specifically, the expression perturbation loss can be determined according to the difference between the reconstructed expression features and predicted expression features; the identity perturbation loss is used to characterize the difference between the reconstructed identity features and predicted identity features of the perturbation features. Specifically, the identity perturbation loss can be determined according to the difference between the reconstructed identity features and predicted identity features.
[0171] Furthermore, the loss function of the expression perturbation loss can be expressed in the following form:
[0172]
[0173] sub to: k′ = ran int(1, K) and k′ ≠ k; n′ = ran int(1, 5)
[0174] In the formula, x kn and x k′n′ respectively represent sample images corresponding to different faces, g EXP (f(x kn )) represents the predicted expression feature of the nth sample image in the kth video; g EXP (f(x kn )) + g ID (f(x k′n′ )) represents the sum of the predicted expression feature of the nth sample image in the kth video and the predicted identity feature of the n′th sample image in the k′th video, obtaining the perturbation feature; h(g EXP (f(x kn )) + g ID (f(x k′n′ ))) represents the perturbation reconstruction feature obtained after feature decoding of the perturbation feature; g EXP (h(g EXP (f(x kn )) + g ID (f(x k′n′ )))) represents the reconstructed expression feature of the perturbation feature. The ultimate optimization goal of the expression perturbation loss is to minimize the loss function L Exp1 .
[0175] The loss function of the identity perturbation loss can be expressed in the following form:
[0176]
[0177] sub to: k′ = ran int(1, K) and k′ ≠ k; n′ = ran int(1, 5)
[0178] Where, g ID (f(x kn )) represents the predicted identity feature of the nth sample image in the kth video; g ID (f(x kn )) + g EXP (f(x k′n′ )) represents the perturbation feature obtained by adding the predicted identity feature of the nth sample image in the kth video and the predicted expression feature of the n′th sample image in the k′th video; h(g ID (f(x kn )) + g EXP (f(x k′n′ ))) represents the perturbation reconstruction feature obtained after decoding the perturbation feature; g ID (h(g ID (f(x kn )) + g EXP (f(x k′n′ )))) represents the reconstructed identity feature of the perturbation feature. The ultimate optimization goal of the identity perturbation loss is to minimize the loss function L Exp1 .
[0179] Therefore, the loss function of the first extraction network can be expressed in the following form:
[0180] L = L ID + λ3L ID1 + λ4L EXP + λ5L EXP1
[0181] Where, L ID represents the identity extraction loss, L ID1 represents the identity perturbation loss, L EXP represents the expression extraction loss, L EXP1 represents the expression perturbation loss, and λ3, λ4, and λ5 are all hyperparameters that can be adjusted according to specific experiments.
[0182] When the loss value of network training remains stable, stop training.
[0183] Based on any of the above embodiments, in step 250, the initial extraction network is trained based on the expression extraction loss, and afterwards, it further includes:
[0184] Based on the sample images carrying expression classification labels, the trained initial encoding network and initial expression extraction branch are jointly fine-tuned with the initial classification network, and the fine-tuned initial encoding network, initial expression extraction branch, and initial classification network are respectively used as the encoding network, expression extraction network, and expression classification network.
[0185] Specifically, the initial extraction network is trained based on the expression extraction loss to obtain the trained initial encoding network and initial expression extraction branch. Since the trained initial encoding network and initial expression extraction branch are not trained using real labeled face expression data, in order to better utilize the expression information and further improve the accuracy of face expression recognition, it is necessary to perform transfer learning on this network using labeled data until the training loss is stable.
[0186] The sample images carrying expression classification labels here are the sample images that have been marked with the expression classification indicated by the sample images, for example, happy, surprised, or sad, etc.
[0187] The initial classification network can be a classifier or a neural network, for example, a VGG16 network or a residual network, etc.
[0188] The fine-tuned initial encoding network can be used as the encoding network to extract image features from the image to be recognized, obtaining image features; the fine-tuned initial expression extraction branch can be used as the expression extraction network to further extract expression features from the image features, obtaining expression features; the fine-tuned initial classification network can be used as the expression classification network to classify the expression features, obtaining the expression recognition result.
[0189] It should be noted that, compared with directly using labeled images for network training, using sample images carrying expression classification labels to fine-tune the network can greatly reduce the number of labeled samples, reduce the annotation cost, and save time and effort.
[0190] The method provided by the embodiments of the present invention can further improve the accuracy of face expression recognition while greatly reducing the number of labeled samples, reducing the annotation cost, saving time and effort, and improving the efficiency of expression recognition by using sample images carrying expression classification labels to fine-tune the network.
[0191] Based on any of the above embodiments, Figure 6 is a schematic flowchart of the method for determining the reference image, the first image, and the second image provided by the present invention, as Figure 6 shown, the reference image, the first image, and the second image are determined based on the following steps:
[0192] Step 610: Determine a sample face video, where each frame image in the sample face video corresponds to the same face.
[0193] Step 620: From the image frame sequence of the sample face video, extract two adjacent frame images as a reference image and a first image, and extract a frame image with the farthest time interval from the reference image as a second image.
[0194] Specifically, the sample face video is a sample video containing a face. Each frame image in each sample face video corresponds to the same face. That is to say, for the frame images from the same sample face video, the faces in them are all the same person.
[0195] Image frames at different times can be selected from the sample face video, arranged in chronological order into an image frame sequence, and two adjacent frame images are extracted as a reference image and a first image, and a frame image with the farthest time interval from the reference image is extracted as a second image. The time interval between the thus-extracted reference image and the first image is less than the time interval between the reference image and the second image.
[0196] In one embodiment, for any sample face video, 5 frame images can be selected at equal intervals, arranged in chronological order as x1, x2, x3, x4, x5. The first two frames x1, x2 and the last frame x5 form a sample image group, where x1 and x2 are used as the reference image and the first image respectively, and x5 is used as the second image; the first frame x1 and the last two frames x4, x5 can also form another sample image group, where x4 and x5 are used as the reference image and the first image respectively, and x1 is used as the second image.
[0197] Based on any of the above embodiments, the sample face video can be determined through the following steps:
[0198] Prepare an initial data set: Download a large number of short videos or long videos such as movies and TV dramas that may contain faces from the network, or they can also be obtained through video capture devices.
[0199] For the obtained large number of long videos, it is necessary to cut out the discontinuous video segments of each person from them. Therefore, a face detection and face tracking model can be used. The specific steps are as follows:
[0200] ① Perform face detection on the video frames of each long video, use the first frame with a detected face as the starting frame, and determine the position of the face frame therein.
[0201] ② Track the determined face until the face is no longer detected or the identity of the person changes, and cut out the video segment during this process.
[0202] ③ Process all the cut video segments according to their durations. For those with a duration greater than 5 seconds, split them every 5 seconds, and discard all videos with a duration less than 3 seconds.
[0203] According to the above steps, all the final video segments for training, that is, the sample face videos, can be obtained.
[0204] Based on any of the above embodiments, the expression recognition can also be achieved through an expression recognition model. Input the image to be recognized into the expression recognition model to obtain the expression recognition result. Among them, the expression recognition model includes an encoding network, an expression extraction network, and an expression classification network.
[0205] A self-supervised learning network can be constructed to train the initial expression recognition model to obtain a trained expression recognition model. Figure 7 It is a schematic diagram of the data flow of the self-supervised learning network provided by the present invention. As Figure 7 shown, the self-supervised learning network includes an initial encoding network, an initial expression extraction branch, an initial identity extraction branch, and an initial decoding network. Among them, the initial expression extraction branch and the initial identity extraction branch share the initial encoding network.
[0206] During the training process of the expression recognition model, the facial expression features and facial identity features are learned through self-supervised learning, and different constraint conditions are imposed on the initial expression extraction branch and the initial identity extraction branch respectively, so that the model can learn the corresponding representation information and decouple the expression features and identity features.
[0207] At the same time, an initial decoding network is added to enable the restoration of the decoupled features after the fusion of the expression features and identity features.
[0208] First, in order to enable the initial identity extraction branch to learn the person identity information well, a corresponding contrastive learning loss is designed to constrain this branch. The design idea of the loss function for identity extraction is as follows: For video frames from the same video segment, it is default that the faces in them are the same person, while the faces in video frames from different video segments come from different people. Therefore, for each video segment, 5 frames of images are selected at equal intervals. Based on the video frames from the same video segment, the first identity feature similarity is determined; based on the video frames from different video segments, the second identity feature similarity is determined. And based on the first identity feature similarity and the second identity feature similarity, the identity extraction loss function is determined.
[0209] Meanwhile, for the initial expression extraction branch, since human expression changes are a continuous process, in a video segment, the closer the video frames are in time sequence, the higher the expression similarity between them, and the farther apart they are, the lower the expression similarity. Therefore, the feature similarity between two frames that are close in time sequence is greater than that between two frames that are far apart in time sequence. Based on this, an expression extraction loss function is designed.
[0210] For the initial decoding network, the goal is to send any frame of image into the initial encoding network. After decoupling through the initial expression extraction branch and the initial identity extraction branch, the face expression representation vector and the face identity representation vector are added together, and then sent into the initial decoding network, and the feature representation before decoupling can still be restored. Based on this, a decoding loss function is designed.
[0211] In the first stage of network training, the identity extraction loss function and the expression extraction loss function are used to supervise the training of the initial face identity extraction branch, the initial face expression extraction branch, and the initial encoding network respectively, and the decoding loss function is used to supervise the training of the initial decoding network. The first encoding network, the first face identity extraction branch, the first face expression extraction branch, and the first decoding network are obtained respectively.
[0212] After that, the second stage of network training is designed. In the second stage, the parameters of the first decoding network are fixed and not involved in the training. In order to enable the two extraction branches to learn more robust features, perturbations are added to each other for perturbation training. The design idea of the perturbation loss function is as follows: for the first expression extraction branch, for each video segment, a random identity representation vector extracted from an image of any other video segment is added to the expression representation vectors extracted from 5 equally spaced frames of images, and the two are added together and sent into the first decoding network. Then, the feature representation decoded by the first decoding network is sent into the first face expression extraction branch to obtain a face expression vector. By constraining this vector to approximate the expression vector before adding the perturbation, an expression perturbation loss is designed. Similarly, for the first identity extraction branch, an identity perturbation loss is designed.
[0213] After the above two stages of training are completed, in order to better utilize the expression information, the first encoding network and the first face expression extraction branch are extracted, combined with the expression classification network, and fine-tuned using labeled data until the training loss is stable, obtaining an expression recognition model for performing expression recognition on the image to be recognized.
[0214] The self-supervised learning scheme for decoupling face identity information and face expression information provided by the embodiments of the present invention learns face identity information and face expression information simultaneously, decouples the two, and perturbs each other during the learning process, so as to learn more robust identity information and expression information. Therefore, it can be better transferred to the downstream face expression recognition task.
[0215] The following describes the expression recognition device provided by the present invention. The expression recognition device described below can be correspondingly referred to the expression recognition method described above. Figure 8 is a schematic structural diagram of the expression recognition device provided by the present invention, as Figure 8 shown, the device includes:
[0216] An image feature determination unit 810, configured to determine image features of a face image to be recognized based on an encoding network;
[0217] An expression feature extraction unit 820, configured to extract expression features from the image features based on an expression extraction network to obtain expression features;
[0218] An expression recognition unit 830, configured to perform expression recognition on the expression features based on an expression classification network;
[0219] The expression extraction network is trained based on a first expression feature similarity between a reference image and a first image, and a second expression feature similarity between the reference image and a second image. The reference image, the first image, and the second image correspond to the same face in a sample face video. The time interval between the reference image and the first image is less than the time interval between the reference image and the second image.
[0220] The expression recognition device provided by the present invention trains an expression extraction network based on a first expression feature similarity between a reference image and a first image, and a second expression feature similarity between a reference image and a second image, so that the expression extraction network can better distinguish the differences between expression features under the same face expression change, and the extracted expression features focus more on expression information; and performs expression recognition based on the expression features, avoiding collecting a large amount of face expression data, reducing the manual annotation cost, and effectively improving the accuracy and reliability of the face expression recognition result.
[0221] Based on any of the above embodiments, the device further includes a network training unit, and the network training unit is configured to:
[0222] Determine an initial extraction network, where the initial extraction network includes an initial encoding network and an initial expression extraction branch;
[0223] Based on the initial extraction network, respectively determine the predicted expression features of the reference image, the first image, and the second image;
[0224] Based on the predicted expression features of the reference image and the first image, determine the first expression feature similarity, and based on the predicted expression features of the reference image and the second image, determine the second expression feature similarity;
[0225] Determine an expression extraction loss based on the first expression feature similarity and the second expression feature similarity;
[0226] Train the initial extraction network based on the expression extraction loss to obtain the encoding network and the expression extraction network.
[0227] Based on any of the above embodiments, the network training unit is further configured to:
[0228] Determine an expression extraction loss based on the difference between the first expression feature similarity and the second expression feature similarity.
[0229] Based on any of the above embodiments, the network training unit is further configured to:
[0230] Determine the predicted image features of the sample image based on the initial encoding network;
[0231] Perform identity feature extraction on the predicted image features of the sample image based on the initial identity extraction branch to obtain the predicted identity features of the sample image;
[0232] Determine a first identity feature similarity based on the predicted identity features of each sample image corresponding to the same face, and determine a second identity feature similarity based on the predicted identity features of each sample image corresponding to different faces;
[0233] Determine an identity extraction loss based on the first identity feature similarity and the second identity feature similarity;
[0234] Train the initial extraction network based on the expression extraction loss and the identity extraction loss.
[0235] Based on any of the above embodiments, the network training unit is further configured to:
[0236] Perform feature decoding on the predicted expression features and predicted identity features of the sample image based on the initial decoding network to obtain the reconstructed image features of the sample image;
[0237] Determine a decoding loss based on the predicted image features and the reconstructed image features of the sample image;
[0238] Train the initial extraction network based on the expression extraction loss, the identity extraction loss, and the decoding loss.
[0239] Based on any of the above embodiments, the network training unit is further configured to:
[0240] Based on the expression extraction loss, the identity extraction loss, and the decoding loss, perform the first-stage training on the initial extraction network to obtain the initially extracted network after the first-stage training as the first extraction network;
[0241] Generate perturbation features based on the predicted identity features and predicted expression features of sample images corresponding to different human faces;
[0242] Based on the first decoding network in the first extraction network, perform feature decoding on the perturbation features to obtain perturbation reconstruction features;
[0243] Based on the first expression extraction branch in the first extraction network, perform expression feature extraction on the perturbation reconstruction features to obtain the reconstructed expression features of the perturbation features, and / or, based on the first identity extraction branch in the first extraction network, perform identity feature extraction on the perturbation reconstruction features to obtain the reconstructed identity features of the perturbation features;
[0244] Based on the reconstructed expression features and predicted expression features of the perturbation features, determine the expression perturbation loss, and / or, based on the reconstructed identity features and predicted identity features of the perturbation features, determine the identity perturbation loss;
[0245] Based on the expression extraction loss and the expression perturbation loss, and / or, based on the identity extraction loss and the identity perturbation loss, train the first expression extraction branch and / or the first identity extraction branch.
[0246] Based on any of the above embodiments, the network training unit is further configured to:
[0247] Determine the expression perturbation loss based on the difference between the reconstructed expression features and predicted expression features of the perturbation features; and / or, determine the identity perturbation loss based on the difference between the reconstructed identity features and predicted identity features of the perturbation features.
[0248] Based on any of the above embodiments, the expression recognition device provided by the embodiments of the present invention further includes a network fine-tuning unit, and the network fine-tuning unit is configured to:
[0249] Based on the sample images carrying expression classification labels, jointly fine-tune the trained initial encoding network and the initial expression extraction branch with the initial classification network, and use the fine-tuned initial encoding network, the initial expression extraction branch, and the initial classification network as the encoding network, the expression extraction network, and the expression classification network respectively.
[0250] Based on any of the above embodiments, the expression recognition device provided by the embodiments of the present invention further includes an image extraction unit, and the image extraction unit is configured to:
[0251] Determine a sample face video, where each frame image in the sample face video corresponds to the same face;
[0252] Extract two adjacent frame images from the image frame sequence of the sample face video as a reference image and a first image, and extract a frame image with the farthest time interval from the reference image as a second image.
[0253] Figure 9 An example of a schematic physical structure diagram of an electronic device is shown as Figure 9 shown. The electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communication interface 920, and the memory 930 complete mutual communication through the communication bus 940. The processor 910 can call the logical instructions in the memory 930 to execute an image recognition method, which includes: determining the image features of a face image to be recognized based on an encoding network; extracting expression features from the image features based on an expression extraction network to obtain expression features; performing expression recognition on the expression features based on an expression classification network; the expression extraction network is trained based on a first expression feature similarity between the reference image and the first image, and a second expression feature similarity between the reference image and the second image. The reference image, the first image, and the second image correspond to the same face in the sample face video, and the time interval between the reference image and the first image is less than the time interval between the reference image and the second image.
[0254] In addition, when the logical instructions in the above-mentioned memory 930 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0255] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image recognition method provided by each of the above methods. The method includes: determining the image features of a face image to be recognized based on an encoding network; extracting expression features from the image features based on an expression extraction network to obtain expression features; performing expression recognition on the expression features based on an expression classification network; the expression extraction network is trained based on the first expression feature similarity between a reference image and a first image, and the second expression feature similarity between the reference image and a second image. The reference image, the first image, and the second image correspond to the same face in a sample face video, and the time interval between the reference image and the first image is less than the time interval between the reference image and the second image.
[0256] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the image recognition method provided by each of the above methods. The method includes: determining the image features of a face image to be recognized based on an encoding network; extracting expression features from the image features based on an expression extraction network to obtain expression features; performing expression recognition on the expression features based on an expression classification network; the expression extraction network is trained based on the first expression feature similarity between a reference image and a first image, and the second expression feature similarity between the reference image and a second image. The reference image, the first image, and the second image correspond to the same face, and the time interval between the reference image and the first image is less than the time interval between the reference image and the second image.
[0257] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0258] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0259] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for facial expression recognition, characterized in that, Including: Based on an encoding network, determine the image features of a face image to be recognized; Based on an expression extraction network, perform expression feature extraction on the image features to obtain expression features; Based on an expression classification network, perform expression recognition on the expression features; The expression extraction network is trained based on the first expression feature similarity between a reference image and a first image, and the second expression feature similarity between the reference image and a second image. The reference image, the first image, and the second image correspond to the same face in a sample face video, and the time interval between the reference image and the first image is less than the time interval between the reference image and the second image; Wherein, the encoding network and the expression extraction network are trained based on the following steps: Determine an initial extraction network, where the initial extraction network includes an initial encoding network and an initial expression extraction branch; Based on the initial extraction network, respectively determine the predicted expression features of the reference image, the first image, and the second image; Based on the predicted expression features of the reference image and the first image, determine the first expression feature similarity, and based on the predicted expression features of the reference image and the second image, determine the second expression feature similarity; Based on the first expression feature similarity and the second expression feature similarity, determine an expression extraction loss; Based on the expression extraction loss, train the initial extraction network to obtain the encoding network and the expression extraction network; Wherein, the initial extraction network further includes an initial identity extraction branch; The training of the initial extraction network based on the expression extraction loss includes: Based on the initial encoding network, determine the predicted image features of a sample image; Based on the initial identity extraction branch, perform identity feature extraction on the predicted image features of the sample image to obtain the predicted identity features of the sample image; Based on the predicted identity features of each sample image corresponding to the same face, determine a first identity feature similarity, and based on the predicted identity features of each sample image corresponding to different faces, determine a second identity feature similarity; Based on the first identity feature similarity and the second identity feature similarity, determine an identity extraction loss; Based on the expression extraction loss and the identity extraction loss, train the initial extraction network.
2. The facial expression recognition method according to claim 1, characterized in that The determination of the expression extraction loss based on the first expression feature similarity and the second expression feature similarity includes: Based on the difference between the first expression feature similarity and the second expression feature similarity, determine the expression extraction loss.
3. The facial expression recognition method according to claim 1, wherein The initial extraction network further includes an initial decoding network; The training of the initial extraction network based on the expression extraction loss and the identity extraction loss includes: Based on the initial decoding network, perform feature decoding on the predicted expression features and predicted identity features of the sample image to obtain the reconstructed image features of the sample image; Based on the predicted image features and the reconstructed image features of the sample image, determine a decoding loss; Based on the expression extraction loss, the identity extraction loss, and the decoding loss, train the initial extraction network.
4. The facial expression recognition method according to claim 3, wherein The training of the initial extraction network based on the expression extraction loss, the identity extraction loss, and the decoding loss includes: Based on the expression extraction loss, the identity extraction loss, and the decoding loss, perform a first-stage training on the initial extraction network to obtain the initially extracted network after the first-stage training as the first extraction network; Generate perturbation features based on the predicted identity features and predicted expression features of sample images corresponding to different human faces; Based on the first decoding network in the first extraction network, perform feature decoding on the perturbation features to obtain perturbation reconstruction features; Based on the first expression extraction branch in the first extraction network, perform expression feature extraction on the perturbation reconstruction features to obtain the reconstructed expression features of the perturbation features, and / or, based on the first identity extraction branch in the first extraction network, perform identity feature extraction on the perturbation reconstruction features to obtain the reconstructed identity features of the perturbation features; Based on the reconstructed expression features and predicted expression features of the perturbation features, determine the expression perturbation loss, and / or, based on the reconstructed identity features and predicted identity features of the perturbation features, determine the identity perturbation loss; Based on the expression extraction loss and the expression perturbation loss, and / or, based on the identity extraction loss and the identity perturbation loss, train the first expression extraction branch and / or the first identity extraction branch.
5. The facial expression recognition method according to claim 4, wherein The determination of the expression perturbation loss based on the reconstructed expression features and predicted expression features of the perturbation features, and / or, the determination of the identity perturbation loss based on the reconstructed identity features and predicted identity features of the perturbation features includes: Determine the expression perturbation loss based on the difference between the reconstructed expression features and predicted expression features of the perturbation features; and / or, Determine the identity perturbation loss based on the difference between the reconstructed identity features and predicted identity features of the perturbation features.
6. The facial expression recognition method according to claim 1, wherein After the training of the initial extraction network based on the expression extraction loss, it further includes: Based on the sample images carrying expression classification labels, jointly fine-tune the trained initial encoding network and initial expression extraction branch with the initial classification network, and use the fine-tuned initial encoding network, initial expression extraction branch, and initial classification network as the encoding network, the expression extraction network, and the expression classification network respectively.
7. The facial expression recognition method according to any one of claims 1 to 6, characterized in that The reference image, the first image, and the second image are determined based on the following steps: Determine a sample human face video, where each frame image in the sample human face video corresponds to the same human face; From the sequence of image frames of the sample human face video, extract two adjacent frames of images as the reference image and the first image, and extract the frame image with the farthest time interval from the reference image as the second image.
8. An expression recognition device, characterized in that, It includes: An image feature determination unit for determining the image features of a human face image to be recognized based on an encoding network; An expression feature extraction unit for performing expression feature extraction on the image features based on an expression extraction network to obtain expression features; An expression recognition unit, configured to perform expression recognition on the expression features based on an expression classification network; The expression extraction network is trained based on a first expression feature similarity between a reference image and a first image, and a second expression feature similarity between the reference image and a second image, where the reference image, the first image, and the second image correspond to the same human face in a sample human face video, and a time interval between the reference image and the first image is less than a time interval between the reference image and the second image; Wherein, the encoding network and the expression extraction network are trained based on the following steps: Determine an initial extraction network, where the initial extraction network includes an initial encoding network and an initial expression extraction branch; Based on the initial extraction network, respectively determine predicted expression features of the reference image, the first image, and the second image; Based on the predicted expression features of the reference image and the first image, determine the first expression feature similarity, and based on the predicted expression features of the reference image and the second image, determine the second expression feature similarity; Based on the first expression feature similarity and the second expression feature similarity, determine an expression extraction loss; Based on the expression extraction loss, train the initial extraction network to obtain the encoding network and the expression extraction network; Wherein, the initial extraction network further includes an initial identity extraction branch; The training of the initial extraction network based on the expression extraction loss includes: Based on the initial encoding network, determine predicted image features of a sample image; Based on the initial identity extraction branch, perform identity feature extraction on the predicted image features of the sample image to obtain predicted identity features of the sample image; Based on the predicted identity features of each sample image corresponding to the same human face, determine a first identity feature similarity, and based on the predicted identity features of each sample image corresponding to different human faces, determine a second identity feature similarity; Based on the first identity feature similarity and the second identity feature similarity, determine an identity extraction loss; Based on the expression extraction loss and the identity extraction loss, train the initial extraction network.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the expression recognition method according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the expression recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Face attribute recognition method and device
CN111666846A
Expression recognition and classroom state evaluation method and device, and medium
CN113239916A
Model training method and device, equipment and storage medium
CN113515980A