A method and system for automatic annotation of facial expressions
The self-supervised learning method using Efficient-CapsNet for automated face expression annotation addresses inefficiencies and inconsistencies in manual annotation, enhancing annotation quality and efficiency by leveraging unannotated data for improved feature representation.
Patent Information
- Application Number
- CN202210564154.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-23
AI Technical Summary
In the prior art, the annotation of face expression recognition data sets relies on manual annotation, resulting in low efficiency and uneven results, and there are subjective differences.
The self-supervised learning method is adopted, and the automatic annotation model is constructed using the Efficient-CapsNet encoder. Through data augmentation, comparison learning and supervised training, combined with self-supervisation and downstream tasks, automatic annotation is achieved.
It improves the efficiency and accuracy of the annotation of facial expression data sets, reduces the subjective differences in manual annotation, makes full use of the intrinsic attribute information of unlabeled data, and is suitable for a variety of data source scenarios.
Smart Images

Figure CN116012903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of face expression recognition in affective computing, and more particularly to a method and system for automatic annotation of face expressions. Background Art
[0002] Automatic annotation of face expressions is a research based on face expression recognition, which is a very important part in the field of emotion recognition. Humans express their emotions in various ways, and facial expressions are the most extensive expression channel among all ways. Currently, the recognition and analysis based on face expressions have been studied and applied in many fields such as medical treatment, education, and customer service. In the research of computer vision and machine learning, various facial expression recognition (FER) systems have encoded expression information from facial representations. Although the face expression recognition technology is constantly developing, there are not many datasets that promote the research of face expressions. Currently, common face expression datasets include JAFFE dataset, CK+ dataset, MMI dataset, Oulu-CASIA dataset, etc. To conduct more in-depth and extensive research on face expression recognition, the quantity and extensiveness of the datasets become increasingly important. Currently, there are numerous channels and methods for obtaining face expressions, but it is relatively difficult to obtain a labeled face expression dataset.
[0003] The development of face expression recognition depends on good datasets, but the annotation work of the datasets used for face expression recognition has not been well developed. Currently, the annotation work of face expression datasets mostly relies on human labor, and the automatic annotation technology has been continuously improved with the development of deep learning technology. Currently, most of the image data annotation methods have problems with uneven annotation effects. In terms of manual annotation, on the one hand, the subjective differences between people will lead to the inconsistency and low accuracy of the annotation results. On the other hand, manual data annotation requires the coordination of multiple aspects of work such as manually obtaining the dataset, manually annotating, manually checking, and manually verifying. This series of cumbersome work will greatly reduce the annotation efficiency, making the datasets used for emotion recognition research always in a state of low sample size. Currently, there are also face expression datasets with a large sample size among the publicly available ones, but such datasets are often network images automatically crawled using keyword-based web crawler technology, with a large amount of non-standard labeled data and very poor annotation quality, which has a great interference on the network training process. Therefore, some researchers use machine learning techniques and methods to semi-automatically annotate the datasets to improve the annotation efficiency. In recent years, with the continuous attention and development of deep learning technology, the annotation work of the datasets has gradually shifted from manual to fully automatic.
[0004] The existing invention patent application document "A Method for Annotating Facial Attribute Data, Computer Equipment and Storage Medium" with the publication number CN114332136A establishes a facial color image dataset; detects the facial region mask of the images in the facial color image dataset; uses a three-dimensional deformation model to randomly initialize the parameters of the images in the facial color image dataset; renders the initialized parameters to obtain a rendered image; annotates all the image data in the facial color image dataset to obtain an annotated illumination dataset and a head pose dataset; inputs the facial images into a facial attribute prediction model for training; iteratively optimizes the model; performs face detection on the to-be-tested facial images, crops the images in the facial region, and inputs them into the trained facial attribute prediction model to predict the illumination parameters and head pose of the face at this time. From the content of the specification of this existing document, it can be seen that the technical solution and logical implementation disclosed in this existing document are significantly different from those of the present application and cannot achieve the technical effects of the present application. The existing invention patent application document "A Multi-Dimensional Emotion Recognition Method and System" with the publication number CN113780341A trains an emotion recognition model and a label mapping model based on a first sample set with labels; inputs the second sample set without labels into the emotion recognition model to obtain the predicted labels of the physiological features in each emotional dimension; inputs the predicted labels into the label mapping model to obtain the mapped labels of the corresponding physiological features in the current dimension; determines whether the consistency between the predicted labels and the mapped labels meets the preset conditions, selects the emotional dimensions whose consistency meets the preset conditions for automatic annotation, and the automatic annotation value of each emotional dimension is the weighted average of the predicted labels and the mapped labels of the corresponding dimension; continues to train the emotion recognition model based on the newly annotated data to obtain the final emotion recognition model. From the content of the embodiments in this existing document, it can be known that the specific application scenario of the technical solution disclosed in this existing document is different from that of the present application, and this existing document does not disclose the technical solution of using the Efficient-CapsNet model for automatic annotation in the present application and cannot achieve the technical effects of the present application.
[0005] Existing automatic annotation methods mostly construct a good deep learning-based model to recognize and annotate the content of images. With the continuous innovation of artificial intelligence, deep learning technology has been constantly evolving. In terms of traditional machine learning, deep learning networks have powerful feature self-learning capabilities, and their model recognition effects and robustness have natural advantages. The research on automatic annotation of facial expressions hopes to greatly save the time and cost of manual annotation through deep learning technology. However, deep learning is highly dependent on large-scale labeled data, which poses obstacles to the exploration of small samples in deep learning, and sufficient theories are needed to improve the expression ability of deep learning.
[0006] In summary, there are technical problems in the existing expression recognition and annotation, such as relying on manual annotation and the uneven results caused by subjective differences in manual annotation. Summary of the invention
[0007] The technical problem to be solved by the present invention is how to solve the technical problem in the prior art that the results are uneven due to reliance on manual labeling and subjective differences in manual labeling.
[0008] The present invention adopts the following technical solution to solve the above technical problem: A method for automatically marking facial expressions comprises:
[0009] S1, using a preset image acquisition device to acquire image frames of facial expression images, forming a data set with the image frames, removing abnormal facial acquisition images in the data set, selecting facial images corresponding to peak expressions in the data set as facial expression image data sets, and preprocessing the facial expression image data sets;
[0010] S2, dividing the facial expression image dataset according to a preset division ratio, wherein the subsets obtained by the above division operation include: a dataset to be labeled and an unlabeled dataset, manually labeling the dataset to be labeled with emotions to obtain a supervised training dataset, and using the unlabeled dataset as a training dataset for self-supervised learning;
[0011] S3, constructing a self-supervised annotation model based on Efficient-CapsNet, using the Efficient-CapsNet encoder as the representation extraction encoder in the self-supervised learning auxiliary task, and performing comparative learning to obtain the optimal pre-training model, the step S3 includes:
[0012] S31, data enhancement processing of the facial expression image to obtain an image to be encoded, and processing the image to be encoded with an Efficient-CapsNet encoder to obtain image feature representation data;
[0013] S32, performing comparative learning according to the image feature representation data, thereby constructing an auxiliary task of the self-supervised annotation model, setting auxiliary training parameters, and inputting the training data set of the self-supervised learning into the auxiliary task to perform iterative comparative training, thereby obtaining and saving the optimal pre-trained model;
[0014] S4, in the self-supervised downstream task, the optimal pre-trained model is combined with a preset classifier, and supervised training and preset adjustment operations are performed on the supervised training data set to obtain an automatic labeling model, and the step S4 includes:
[0015] S41, constructing a downstream task of the self-supervised labeling model, wherein the downstream task includes: a downstream task encoder and a downstream task classifier;
[0016] S42, setting downstream training parameters, inputting the supervised training data set into the downstream task, combining the downstream task classifier and the optimal pre-training model to perform supervised iterative training, and obtaining and saving the automatic labeling model accordingly;
[0017] S5. Automatically annotate the facial expression image with emotions using the automatic annotation model to obtain an automatic facial expression annotation result.
[0018] The present invention adopts a self-supervised method to train an automatic annotation model, which overcomes the defects of low efficiency and uneven results caused by subjective differences between different annotation personnel in the current facial expression data set in pure manual annotation. The present invention uses a self-supervised learning method. When faced with a large amount of unlabeled data, the auxiliary task of self-supervised learning can learn a large amount of attribute information inherent in the data from the unsupervised data, making full use of data resources to make full use of the superior performance of the pre-trained model on a small amount of labeled data in downstream tasks. The method provided by the present invention has universal application, is not targeted at a specific hardware environment, and can meet the basic software dependency package. In addition, the method of the present invention has good scalability and is not limited to specific data source scenarios.
[0019] In a more specific technical solution, step S1 includes:
[0020] S11, eliminating non-frontal face images and non-face images in the data set, and selecting peak expression face images from the remaining data set as the final facial expression data set;
[0021] S12, cropping the facial images in the data set, and uniformly adjusting the facial images to a preset size;
[0022] S13, using a face detector to detect faces in the face image, performing an alignment operation on the faces, and using the aligned faces to generate the facial expression image dataset.
[0023] In a more specific technical solution, step S2 includes:
[0024] S21, dividing the facial expression data set into a small-ratio data set and a large-ratio data set according to the preset division ratio, wherein the preset division ratio includes: 4:1;
[0025] S22, using the small-scale dataset as the dataset to be labeled;
[0026] S23, using the large-scale dataset as the unlabeled dataset;
[0027] S24. Manually annotate the dataset to be annotated with artificial emotions, and use the unannotated dataset as the training dataset for the self-supervised learning.
[0028] In a more specific technical solution, step 31 includes:
[0029] S311. The first part is the data augmentation part. The input image of the model will be randomly augmented twice, and the two augmented input images will be input into the preset network simultaneously for parallel learning.
[0030] S312. Use the Efficient-CapsNet encoder to extract features from the input image to obtain the feature representations of two types of images. Among them, the Efficient-CapsNet encoder includes: a convolutional layer, a depth convolutional layer, a primary capsule layer, and an FCCaps layer. Efficient-CapsNet also includes a self-attention mechanism routing. After passing through the FCCaps layer, an image characterization matrix is output, and the size of the image characterization matrix is: the number of categories × 16.
[0031] In the data input stage of the contrastive learning of the present invention, the input data examples need to be randomly augmented to obtain two related views of the same example. In the self-supervised learning used in the present invention, the encoder network of the auxiliary task adopts the Efficient-CapsNet encoder.
[0032] In a more specific technical solution, step S312 includes:
[0033] S3121. The input image enters the convolutional layer of the Efficient-CapsNet encoder. The input image is grayscaled and sent to four preset convolutional layers for processing to obtain the encoder convolutional output feature map.
[0034] S3122. Use the batch normalization method to normalize each neuron in the preset network layer, and use the following transformation reconstruction algorithm to process to obtain the normalization result:
[0035]
[0036] Among them, is the input after k-layer normalization, and γ and β are a pair of introduced parameters that are learned together with the model parameters.
[0037] S3123. Use the following logic to restore the feature distribution learned by each layer:
[0038]
[0039]
[0040] Embed a batch normalization layer between the convolutional layers of the Efficient-CapsNet, and adopt a weight sharing method on the batch normalization layer during the convolutional operation, and process the convolutional output feature map of the encoder through the processing method of neurons to uniform the data distribution within the layer;
[0041] S3124. Perform a depthwise separable convolution operation on the convolutional output feature map of the encoder to construct the primary capsules;
[0042] S3125. Adopt the self-attention mechanism routing in the FFCaps layer to process and obtain the image feature representation data accordingly.
[0043] The Efficient-CapsNet of the present invention adds an attention mechanism routing and a depthwise separable convolution operation on the basis of the capsule network. While ensuring the recognition accuracy, it greatly reduces the network parameters and improves the training efficiency of the network. The capsules in the capsule network are a form of feature representation. They can store the attribute information of different targets from different perspectives and have equivariance. The capsules store the attribute information of the target in the form of vectors, such as the size and direction angle of the target entity. At the same time, the capsule vector can also represent the existence or non-existence of the target.
[0044] In a more specific technical solution, in the step S3125, the following logic is adopted to obtain the feature representation:
[0045]
[0046]
[0047] Among them, B l is the prior matrix, and all capsules of the l+1 layer are calculated by the following logic
[0048]
[0049] Squeeze the lengths of all capsule vectors of the l+1 layer to between 0 and 1 through the following squeezing function to obtain
[0050]
[0051] Among them, C l is the coupling coefficient matrix generated by the self-attention mechanism algorithm, n l represents that there are n l capsules in the lth layer, n l+1 represents that there are n l+6 capsules in the lth layer, and d 1 is the dimension of the capsules in the lth layer.
[0052] In a more specific technical solution, the step S32 includes:
[0053] S321. First input the feature representation into a preset non-linear projection transformation layer for non-linear projection transformation to eliminate redundant and irrelevant information in the feature representation, so as to obtain sample representation attribute data;
[0054] S322. Perform contrast learning on the sample representation attribute data, and continuously update the learning parameters of the Efficient-CapsNet encoder through contrast feedback, so as to obtain the optimal pre-trained model.
[0055] In the present invention, the feature representation is first input into a non-linear projection transformation layer to eliminate redundant and irrelevant information in the feature, so as to reveal the essential attributes of the sample data; then, the representation after non-linear projection transformation is subjected to contrast learning, and the learning parameters of the encoder are continuously updated through contrast feedback. The self-supervised learning method of the present invention adopts a discriminative self-supervised learning method. Discriminative self-supervised learning expects that the data representation contains enough information, and finds the differences between data through discriminative tasks, and then finds the classification boundary.
[0056] In a more specific technical solution, the step S322 includes:
[0057] S3221. Input the image representation matrix into a non-linear MLP (Dense->Relu->Dense) layer with two layers to map the image representation matrix into the space of contrast loss;
[0058] S3222. Use a sliding window to block the matrix, and calculate the variance of each window block separately:
[0059]
[0060]
[0061] where ω i is the Gaussian kernel weight, and N is the number of elements in the window block;
[0062] S3223. Calculate the covariance of the corresponding window blocks b and b′ of the two image representation matrices according to the following logic:
[0063]
[0064] S3224. Process the variance and covariance of the window blocks according to the following logic to obtain the SSIM value:
[0065]
[0066] where c1 = (k1L) 2 and c2 = (k2L) 2 are two variables used to stabilize division. Here, L in c1 and c2 is the dynamic range of matrix element values, and k1 and k2 are hyperparameters;
[0067] S3225. Average the SSIM of all the window blocks to obtain an average value as the overall similarity of the image representation matrix:
[0068]
[0069] where B is the number of matrix sliding window blocks, z and z′ are the input representation matrices, and z i and z′ i are the representation matrices of the i-th window block corresponding to the two representation matrices;
[0070] S3226. Use the adjustable temperature-normalized cross-entropy loss based on the SSIM algorithm and the following cosine similarity metric transformation logic to calculate the overall similarity by comparing losses to obtain the SSIM matrix similarity metric comparison loss function:
[0071]
[0072] S3227. Feed back the update network according to the SSIM matrix similarity metric comparison loss function to obtain the optimal pre-trained model.
[0073] The SSIM of the present invention uses the weighted mean and variance of matrix elements to describe the structural information of the matrix, and uses covariance to describe the mutual relationship between the element distributions of two matrices. According to the idea of NCE, the positive and negative samples can be classified as positive and negative using the data distribution relationship between positive and negative samples. SSIM jointly calculates the similarity between two matrices using the weighted mean, variance, and covariance.
[0074] Since the difference between the matrix and the vector in the present invention is mainly reflected in that the span of matrix elements is very large and the mean and variance of elements cannot be calculated as a whole in the form of a vector, in the SSIM algorithm, a sliding window is used to divide the matrix into blocks, calculate the SSIM for each block separately, and finally average the SSIM values of each block, avoiding the phenomenon of large fluctuations in the mean and variance.
[0075] In a more specific technical solution, the step S5 includes:
[0076] S51. Input the preprocessed face expression image dataset into the input of the automatic annotation model of the network;
[0077] S52, performing auxiliary task learning on the unlabeled data set in the facial expression data set, and saving the encoder model with the best comparison effect through iterative training;
[0078] S53: In the downstream task, the data set to be labeled is input into the pre-trained model, and iterative supervised training is performed through a supervised learning strategy to obtain and save the optimal labeling model;
[0079] S54. Combining the encoder model with the annotation model to obtain the optimal automatic annotation model, and obtaining the facial expression automatic annotation result accordingly.
[0080] In a more specific technical solution, a system for automatically labeling facial expressions includes:
[0081] An expression data set module is used to obtain image frames of facial expression images with a preset image acquisition device, form a data set with the image frames, remove abnormal facial acquisition images in the data set, select facial images corresponding to peak expressions in the data set as facial expression image data sets, and pre-process the facial expression image data sets;
[0082] A data set partitioning module is used to partition the facial expression image data set according to a preset partitioning ratio, wherein the subsets obtained by the aforementioned partitioning operation include: a data set to be labeled and a data set without labels, the data set to be labeled is manually labeled with emotions to obtain a data set for supervised training, and the data set without labels is used as a training data set for self-supervised learning, and the data set partitioning module is connected to the expression data set module;
[0083] The optimal pre-training model acquisition module is used to construct a self-supervised annotation model based on Efficient-CapsNet, use the Efficient-CapsNet encoder as the representation extraction encoder in the self-supervised learning auxiliary task, and perform comparative learning to obtain the optimal pre-training model. The optimal pre-training model acquisition module is connected to the data set partitioning module, and the optimal pre-training model acquisition module includes:
[0084] A feature representation module, used for processing the facial expression image by data enhancement to obtain an image to be encoded, and processing the image to be encoded by an Efficient-CapsNet encoder to obtain image feature representation data;
[0085] A self-supervised learning module, used for performing comparative learning according to the image feature representation data, thereby constructing an auxiliary task of a self-supervised annotation model, setting auxiliary training parameters, inputting the training data set of the self-supervised learning into the auxiliary task to perform iterative comparative training, thereby obtaining and saving the optimal pre-training model, and the self-supervised learning module is connected to the feature representation module;
[0086] An automatic annotation model acquisition module, which is used to combine the optimal pre-trained model with a preset classifier in a self-supervised downstream task, and perform supervised training and preset adjustment operations on the supervised training dataset to obtain an automatic annotation model. The automatic annotation model acquisition module is connected to the dataset division module. The automatic annotation model acquisition module includes:
[0087] A downstream task construction module, which is used to construct the downstream task of the self-supervised annotation model. Among them, the downstream task includes: a downstream task encoder and a downstream task classifier;
[0088] A supervised iterative training module, which is used to set downstream training parameters, input the supervised training dataset into the downstream task, and combine the downstream task classifier and the optimal pre-trained model to perform supervised iterative training, so as to obtain and save the automatic annotation model. The supervised iterative training module is connected to the downstream task construction module;
[0089] An automatic annotation module, which is used to automatically annotate the emotion of the facial expression image with the automatic annotation model to obtain an automatic facial expression annotation result. The automatic annotation module is connected to the expression dataset module, the optimal pre-trained model acquisition module and the automatic annotation model acquisition module.
[0090] The present invention has the following advantages compared with the prior art: The present invention uses a self-supervised method to train an automatic annotation model, which overcomes the defects of low efficiency and uneven results caused by subjective differences between different annotators in the pure manual annotation of the current facial expression dataset. The method used in the present invention is a self-supervised learning method. When facing a large amount of unannotated data, the auxiliary tasks of self-supervised learning can learn a large amount of inherent attribute information in the unsupervised data, making full use of the data resources, so as to make full use of the superior performance of the pre-trained model on a small amount of annotated data in the downstream task. The method provided by the present invention has universality in application, does not target a specific hardware environment, and only needs to meet the basic software dependency packages. And the method of the present invention has good scalability and is not limited to a specific data source scenario.
[0091] In the data input stage of the contrastive learning of the present invention, the input data example needs to be randomly augmented to obtain two related views of the same example. In the self-supervised learning used in the present invention, the encoder network of its auxiliary task adopts an Efficient-CapsNet encoder.
[0092] The Efficient-CapsNet of the present invention adds an attention mechanism routing and depthwise separable convolution operation on the basis of the capsule network. While ensuring the recognition accuracy, it greatly reduces the network parameters and improves the training efficiency of the network. Capsules in the capsule network are a form of feature representation. They can store the attribute information of different objects from different perspectives and have equivariance. Capsules store the attribute information of objects in vector form, such as the size and direction angle of the target entity. At the same time, the capsule vector can also represent the existence or non-existence of the target.
[0093] The present invention first inputs the feature representation into a non-linear projection transformation layer to eliminate redundant and irrelevant information in the features, thereby revealing the essential attributes of the sample data. Then, the representation after non-linear projection transformation is subjected to contrastive learning, and the learning parameters of the encoder are continuously updated through contrastive feedback. The self-supervised learning method of the present invention adopts a discriminative self-supervised learning method. Discriminative self-supervised learning expects that the data representation contains enough information to find the differences between data through discriminative tasks, and then find the classification boundary.
[0094] The SSIM of the present invention uses the weighted mean and variance of matrix elements to describe the structural information of the matrix, and uses covariance to describe the mutual relationship of the distribution of two matrix elements. According to the idea of NCE, the positive and negative samples can be classified as positive and negative using the data distribution relationship between positive and negative samples. SSIM uses the weighted mean, variance, and covariance to jointly calculate the similarity of two matrices.
[0095] Since the difference between the matrix and the vector of the present invention is mainly reflected in the large span of matrix elements and cannot calculate the mean and variance of elements in the form of a vector as a whole, in the SSIM algorithm, a sliding window is used to divide the matrix into blocks, calculate the SSIM for each block separately, and finally average the SSIM values of each block, avoiding the phenomenon of large fluctuations in the mean and variance. The present invention solves the technical problems existing in the prior art, such as relying on manual annotation and the uneven results caused by the subjective differences of manual annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1 It is a schematic diagram of the basic steps of a method for automatic facial expression annotation in Embodiment 1 of the present invention;
[0097] Figure 2 It is a schematic diagram of the face image acquisition and preprocessing process in Embodiment 1 of the present invention;
[0098] Figure 3 It is a schematic diagram of the process for obtaining a pre-trained model in Embodiment 1 of the present invention;
[0099] Figure 4It is the overall step flowchart of a method for automatic facial expression annotation provided in Embodiment 2 of the present invention;
[0100] Figure 5 It is the schematic diagram of the architecture of Efficient-CapsNet in Embodiment 2;
[0101] Figure 6 It is the schematic diagram of the self-attention mechanism routing structure in Embodiment 2;
[0102] Figure 7 It is the schematic diagram of the structure of contrastive learning in the auxiliary task in Embodiment 2;
[0103] Figure 8 It is the schematic diagram of the architecture of the downstream classification task in Embodiment 2;
[0104] Figure 9 It is the schematic diagram of the change curve of the loss value in the contrastive training process in Embodiment 3;
[0105] Figure 10 It is the schematic diagram of the change curves of the cosine similarity and matrix similarity of the output representation in Embodiment 3;
[0106] Figure 11 It is the schematic diagram of the change curves of the training loss and validation loss of classification in Embodiment 3;
[0107] Figure 12 It is the schematic diagram of the change curves of the training accuracy and validation accuracy of classification in Embodiment 3;
[0108] Figure 13 It is the schematic diagram of the confusion matrix of automatic annotation on a new unlabeled facial expression dataset in Embodiment 3. Detailed implementation manners
[0109] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0110] Embodiment 1
[0111] As Figure 1 shown, a method for automatic facial expression annotation provided by the present invention includes:
[0112] S1: Collect facial images using devices such as cameras in a specific application scenario, and preprocess the collected facial expression image dataset.
[0113] Furthermore, as Figure 2 shown, the preprocessing of the facial expression image dataset in step S1 is specifically as follows:
[0114] S11: First, non-frontal face images and no-face images in the dataset are removed, and peak-expression face images are selected from the remaining dataset as the final facial expression dataset.
[0115] S12: Then, the face images in the dataset are cropped, and the size of the images is uniformly adjusted to 64×64.
[0116] S13: Use a face detector to detect the faces in the images and perform face alignment operations to generate a new aligned facial expression image dataset.
[0117] S2: Select a certain amount of face image data from the facial expression image dataset collected in step S1 at a certain ratio for manual annotation of emotion labels.
[0118] Furthermore, a certain amount of face image data is selected from the facial expression image dataset in step S2 at a certain ratio for manual annotation of its emotion labels. Specifically: First, the facial expression dataset is randomly divided at a ratio of 4:1, and then the smaller part of the divided ratio is manually annotated with emotions as the training dataset for supervised learning, and the larger dataset is used as the training dataset for self-supervised learning.
[0119] S3: Construct a self-supervised annotation model based on Efficient-CapsNet, use the Efficient-CapsNet encoder as the feature extraction encoder in the self-supervised learning auxiliary task, and perform contrastive learning to obtain a pre-trained model.
[0120] Furthermore, the self-supervised learning method in step S3 uses a discriminative self-supervised learning method. Discriminative self-supervised learning expects the data representation to contain enough information, and finds the differences between data through discriminative tasks, and then finds the classification boundary. The self-supervised learning method mainly consists of an auxiliary task and a downstream task. The construction of the auxiliary task of its self-supervised annotation model is divided into three parts.
[0121] As Figure 3 shown, step S3 also includes the following steps:
[0122] S31: The first part is the data augmentation part. The input images of the model will be randomly augmented twice, and the two augmented input images will be input into the network simultaneously for parallel learning.
[0123] In this embodiment, in the data input stage of contrastive learning in step S31, the input data examples need to be randomly augmented to obtain two related views of the same example. Random cropping, random color distortion, and random Gaussian blur are adopted as three data augmentation methods.
[0124] S32: The second part is the encoder part. The two input images obtained in step S31 are used by the encoder for feature extraction to obtain the feature representations of the two images.
[0125] In this embodiment, the encoder in step S32 adopts the encoder of the Efficient-CapsNet model. It includes a convolutional layer, a depth convolutional layer, a primary capsule layer, and an FCCaps layer, and its output representation is a matrix of size number of classes × 16.
[0126] S33: The third part is the contrastive learning part. The feature representations obtained in step S32 are first input into a non-linear projection transformation layer to remove redundant and irrelevant information in the features, so as to expose the essential attributes of the sample data; then the representations after non-linear projection transformation are subjected to contrastive learning, and the learning parameters of the encoder are continuously updated through contrastive feedback to obtain a pre-trained model with automatic annotation.
[0127] In this embodiment, in the contrast stage of contrastive learning in step S33, a matrix similarity measurement method based on the SSIM algorithm is adopted for the two input representation matrices.
[0128] SSIM uses the weighted mean and variance of matrix elements to describe the structural information of the matrix, and uses covariance to describe the mutual relationship between the distributions of two matrix elements. According to the idea of NCE, the positive and negative samples can be classified as positive or negative by using the data distribution relationship between positive and negative samples. SSIM jointly calculates the similarity of two matrices by using weighted mean, variance, and covariance.
[0129] The difference between a matrix and a vector is mainly reflected in that the span of matrix elements is very large, and the mean and variance of elements cannot be calculated as a whole in the form of a vector, otherwise there will be a phenomenon of large fluctuations in the mean and variance. Therefore, in the SSIM algorithm, a sliding window is used to divide the matrix into blocks, the SSIM of each block is calculated separately, and finally the SSIM values of each block are averaged. SSIM uses a Gaussian convolution kernel with a variance of 1.5 to calculate the weighted average of each window block, and its calculation is as shown in formula (1), where ω i is the Gaussian kernel weight and N is the number of elements in the window block. The variance calculation of the window block is as shown in formula (2).
[0130]
[0131]
[0132] The covariance of the window blocks b and b′ corresponding to the two matrices is calculated as shown in formula (3).
[0133]
[0134] Finally, the SSIM is calculated as shown in formula (4), where b and b′ are the corresponding sliding window blocks of the two matrices.
[0135]
[0136] where c1 = (k1L) 2 and c2 = (k2L) 2 are two variables used to stabilize the division and prevent the denominator from being zero. L in c1 and c2 is the dynamic range of the matrix element values, and k1 and k2 are hyperparameters, taking 0.01 and 0.03 respectively.
[0137] Formula (4) calculates the SSIM of each window block. Therefore, it is also necessary to average the SSIM of all window blocks and use the average value as the overall similarity of the matrix. As shown in formula (5), where B is the number of matrix sliding window blocks, z and z′ are the input representation matrices, and z i and z i ′ are the representation matrices of the i-th window block corresponding to the two representation matrices.
[0138]
[0139] In this embodiment, the loss function in the contrastive learning training process in step S22 adopts the temperature-adjustable normalized cross-entropy loss (NT-Xent) based on the SSIM algorithm.
[0140] Taking the representation matrices of the two enhanced images extracted by the encoder as the input of the SSMI algorithm, the similarity of the two representation matrices is calculated. According to the idea of the NT-Xent loss function, the cosine similarity metric is transformed to obtain the contrastive loss function of the matrix similarity metric based on the SSIM algorithm, as shown in formula (6).
[0141]
[0142] S4: In the self-supervised downstream task, apply the pre-trained model obtained in step S33 to the downstream task through parameter-based transfer learning, connect the pre-trained encoder model and the classifier to form a complete facial expression classification model, and use the supervised method to fine-tune the model on a small amount of labeled data to obtain the final automatic annotation model.
[0143] S5: Automatically annotate the facial expression images obtained by the automatic annotation model for the same scene to obtain the annotation results.
[0144] Embodiment 2
[0145] As Figure 4 shown, this embodiment provides a method for automatic annotation of facial expressions, and the method includes the following processes:
[0146] S1': Obtain image frames from a specific application scenario through a camera to form a data set, eliminate non-frontal face images and no-face images therein, and select the face images with peak expressions as the facial expression image data set;
[0147] In this embodiment, facial images of a specific scene are collected by a camera. By means of extracting video frames, one video image frame is extracted every five frames, and a face detector is used to detect the faces in the images to obtain face images. Then, non-frontal face images and no-face images in the initially obtained facial image data set are eliminated, and the face images with peak expressions are selected from the remaining data set as the final facial expression data set.
[0148] S2': Preprocess the facial expression image data set, including size normalization and face alignment;
[0149] In this embodiment, the facial images in the data set are further cropped, and the size of the images is uniformly adjusted to 64×64 by using the bilinear interpolation method, and then face alignment operation is performed to generate a new aligned facial expression image data set.
[0150] S3': Randomly divide the facial expression image data set into two parts according to a ratio of 4:1, and perform manual emotion annotation on the data set with a small ratio;
[0151] In this embodiment, a certain amount of facial image data is selected from the facial expression image data set in the previous step to perform manual annotation on its emotion labels. First, the facial expression data set is randomly divided according to a ratio of 4:1, and then the part with a small division ratio is manually annotated with emotions as the training data set for supervised training, and the data set with a large ratio is used as the training data set for self-supervised learning.
[0152] S4': Construction of the auxiliary task of the self-supervised annotation model, including a data augmentation module, an Efficient-CapsNet encoder, and contrast learning;
[0153] The annotation model adopts a discriminative self-supervised learning method, which mainly consists of an auxiliary task and a downstream task.
[0154] First, the auxiliary tasks of the self-supervised annotation model are constructed. The process is divided into three parts: data augmentation, encoder, and contrastive learning. The processing steps of the input image are as follows:
[0155] 1. For data augmentation of the input image, three methods are adopted: random cropping, random color distortion, and random Gaussian blur. Each image is randomly augmented twice, and the two randomly augmented images are simultaneously input into the encoder.
[0156] 2. As Figure 5 shown, the encoder uses the Efficient-CapsNet encoder. Efficient-CapsNet improves the training efficiency of its model by adding attention mechanism routing and depthwise separable convolution operations on the basis of the capsule network.
[0157] Efficient-CapsNet is mainly divided into three parts: convolutional layer, primary capsule layer, and self-attention mechanism routing. In the first part of Efficient-CapsNet, in the convolutional layer, the input is mapped to a higher-dimensional space through multiple convolutional operations and batch normalization operations to prepare for the creation of capsules. In the second part, the high-dimensional feature map is further used to create a vector representation of the represented features through depthwise separable convolution, resulting in the primary capsule layer. Depthwise separable convolution consists of depth convolution and pointwise convolution. In depth convolution, instead of performing multi-channel convolution like the original convolution, the multi-channel feature map is first disassembled into single channels, and then each single channel is convolved separately. Each input channel corresponds to a filter, and then the obtained feature map is subjected to pointwise convolution. Pointwise convolution is an operation using a 1×1 convolutional kernel for convolution, which can perform linear output for deep networks and greatly reduces the number of parameters required by the network compared to traditional convolution operations. In the last part, the self-attention mechanism routing is used to route low-level capsules to the whole they represent.
[0158] Its encoder includes a convolutional layer, a depth convolutional layer, a primary capsule layer, and an FCCaps layer. The encoding process of the input face image dataset includes the following steps:
[0159] 1) First, the input image enters the convolutional layer of the encoder. After grayscaling, it is sent to four convolutional layers for processing. The first convolutional layer uses 32 channels, a 7×7 convolutional kernel, and a stride of 2. For a face image with an input size of 64×64, it outputs 32 feature maps of size 29×29. The second and third convolutional layers both use 64 channels, a 3×3 convolutional kernel, and a stride of 1, and output 64 feature maps of size 25×25. The fourth convolutional layer uses 128 channels, a 3×3 convolutional kernel, and a stride of 2, and outputs 128 feature maps of size 12×12.
[0160] During the network training process, small changes in the data of each layer will be amplified in the deeper layers. If the data distribution for each training is uneven, then each iterative learning has to adapt to the new distribution rule, which will affect both the training speed and the convergence speed of the network. At the same time, it will also lead to a significant reduction in the generalization ability of the network. Batch Normalization improves the network training effect from the perspective of uneven training data distribution. It normalizes each neuron in the network layer, and to solve the problem that the features learned in each layer are destroyed due to the normalization operation, a transformation reconstruction algorithm is proposed to obtain a new normalization result. The transformation calculation is shown in formula (7), and then the feature distribution learned in each layer is restored through formula (8) and formula (9).
[0161]
[0162] Among them is the input after normalization in the k-th layer, and γ, β are a pair of introduced parameters, which are learned together with the model parameters.
[0163]
[0164]
[0165] In the convolution operation, in order to reduce the number of γ and β parameters generated during the transformation reconstruction process, a method of weight sharing is adopted on the batch normalization layer, and the feature map is processed in the way of neurons. Therefore, in these four convolutional layers of Efficient-CapsNet, a batch normalization layer is embedded between each layer to evenly distribute the data within the layer.
[0166] 2) For the feature map obtained from the previous convolutional layer, perform a depthwise separable convolution operation on it. Use 128 channels, a 12×12 convolutional kernel, a stride of 1, output 128 neurons, and group these 128 neurons into capsules in the shape of (16, 8), output 16 capsules, with 8 neurons in each capsule. So far, the main capsules are constructed.
[0167] 3) As Figure 6 shown, a self-attention mechanism routing is adopted in the FFCaps part, which is similar to a fully connected network. The input of the upper-layer capsules is the weighted sum of all "prediction vectors" from the lower-layer capsules, and the number of output capsules is equal to the number of classification categories.
[0168] Among them, represents that there are 16 capsules in the l-th layer, and each capsule has 8 dimensions, represents that there are 7 capsules in the (l + 1)-th layer, and each capsule has 16 dimensions, Represents the weight matrix, whose dimension is (16, 7, 8, 16). It is also the matrix for affine transformation between the front and back two layers of capsules, and makes predictions on the attributes of the lower-layer capsules according to the attributes of the current capsule. Are all the predictions of the previous layer of capsules, C l Is the coupling coefficient matrix generated by the self-attention mechanism algorithm, see Formulas (10) and (11). In Formulas (3 - 7), n l Represents that there are n l Capsules in the l-th layer, n l+1 Represents that there are n l+1 Capsules in the l-th layer, f l Is the dimension of the capsules in the l-th layer, The role of is to balance the coupling coefficient and the logarithmic priority, so as to stabilize the training process. A l Is the self-attention matrix, each capsule corresponds to a self-attention matrix, and it contains the consistency scores of each combination of predictions.
[0169]
[0170]
[0171] B l Is the prior matrix, which contains the discriminant information of all weights, and calculates all the capsules in the l + 1 layer through Formula (12)
[0172]
[0173] Then, through the squashing function, the lengths of all capsule vectors in the l + 1 layer are squashed between 0 and 1 to obtain The squashing function of this network is shown in Formula (13).
[0174]
[0175] After passing through the FFCaps layer, the feature matrix of the output image is obtained, and the size of the feature matrix is: the number of categories × 16.
[0176] 3. The framework of contrastive learning is as Figure 7 Shown, and the processing process includes the following steps:
[0177] 1) Input the two feature matrices in the previous step into a two-layer non-linear MLP (Dense->Relu->Dense) layer to map the feature matrices into the space of contrastive loss.
[0178] 2) In the contrastive task, the goal is to maximize the different feature vectors Z of the same image i And Z jThe similarity between them is calculated by first computing the similarity between two input feature matrices using a matrix similarity measurement method based on the SSIM algorithm.
[0179] SSIM uses the weighted mean and variance of matrix elements to describe the structural information of the matrix, and covariance to describe the mutual relationship of the element distributions of two matrices. According to the idea of NCE, the positive and negative samples can be classified as positive or negative using the data distribution relationship between positive and negative samples. SSIM jointly calculates the similarity of two matrices using the weighted mean, variance, and covariance.
[0180] The difference between a matrix and a vector is mainly reflected in the fact that the span of matrix elements is very large, and the mean and variance of elements cannot be calculated as a whole in the form of a vector, otherwise there will be a phenomenon of large fluctuations in the mean and variance. Therefore, in the SSIM algorithm, a sliding window is used to block the matrix, and the SSIM is calculated separately for each block, and finally the SSIM values of each block are averaged. SSIM uses a Gaussian convolution kernel with a variance of 1.5 to calculate the weighted average of each window block, and its calculation is as shown in formula (14), where ω i is the Gaussian kernel weight, and N is the number of elements in the window block. The variance calculation for the window block is as shown in formula (15).
[0181]
[0182]
[0183] The covariance of the corresponding window blocks b and b′ of two matrices is calculated as shown in formula (16).
[0184]
[0185] Finally, the calculation of SSIM is as shown in formula (17), where b and b′ are the corresponding sliding window blocks of two matrices.
[0186]
[0187] where c1 = (k1L) 2 and c2 = (k2L) 2 are two variables used to stabilize the division and prevent the denominator from being zero. L in c1 and c2 is the dynamic range of matrix element values, and k1 and k2 are hyperparameters, taking 0.01 and 0.03 respectively.
[0188] What formula (4) calculates is the SSIM of each window block. Therefore, it is also necessary to average the SSIM of all window blocks and use the average value as the overall similarity of the matrix. As shown in formula (18), where B is the number of matrix sliding window blocks, z and z′ are the input feature matrices, z i and z′ iis the feature matrix of the i-th window block corresponding to the two feature matrices.
[0189]
[0190] 3) Calculate the contrast loss of the similarity between the two input features to feedback and update the network. The loss function uses the temperature-adjustable normalized cross-entropy loss (NT-Xent) based on the SSIM algorithm. Using the similarity of the two feature matrices obtained in the previous step, according to the idea of the NT-Xent loss function, transform the cosine similarity metric in it to obtain the contrast loss function of the matrix similarity metric based on the SSIM algorithm, as shown in formula (19).
[0191]
[0192] After the above steps of the auxiliary task, the pre-trained network model is constructed.
[0193] S5’. Construct the downstream task of the self-supervised annotation model, including an encoder and a classifier;
[0194] In this embodiment, in the self-supervised downstream task, apply the pre-trained model obtained from the auxiliary task to the downstream task through parameter-based transfer learning, connect the pre-trained encoder model and the classifier to form a complete facial expression classification model, and use the supervised method to fine-tune the model on a small amount of labeled data to obtain the final automatic annotation model. The process of the model is as Figure 8 shown.
[0195] S6’. In the auxiliary task training stage, set relevant training parameters, including learning rate, number of iterations, decay strategy, τ value, window size, etc. Input the unlabeled dataset into the auxiliary task for iterative contrast training, and save the optimal pre-trained model;
[0196] S7’. In the downstream task training stage, set relevant training parameters, including learning rate, number of iterations, decay strategy, batch size, etc. Input the labeled dataset into the downstream classification task for supervised iterative training, and save the optimal annotation model;
[0197] S8’. Use the generated automatic annotation model to automatically annotate the facial expression images obtained in the same scene to obtain the annotation results.
[0198] In this embodiment, the auxiliary task model and the downstream task model have been constructed in the previous step. At this stage, they need to be trained to obtain a high-annotation-rate automatic face expression annotation model. Relevant parameters for the training stage need to be set according to specific training metric requirements. These relevant parameters include learning rate, number of iterations, and decay strategy, etc. In this embodiment, the preprocessed face expression dataset is used as the input of the network. First, for the unannotated dataset in the face expression dataset, auxiliary task learning is carried out. After the iterative training is completed, the encoder model with the best comparison effect is saved. In the downstream task, a small amount of annotated face expression dataset is used as the input of the pre-trained model, and iterative supervised training is carried out through a supervised learning strategy, and the annotation model with the optimal training effect is saved.
[0199] Embodiment 3
[0200] This embodiment is a simulation experiment of Embodiment 1. In other embodiments, neither a simulation experiment may be conducted, nor other experimental schemes may be adopted to determine relevant parameters and the effect of automatic face expression annotation.
[0201] In this embodiment, the relevant operating environment is configured. The hardware support is an Intel(R) Core(TM) i7-6850K CPU@3.60GHz processor, 32GB of memory, and an NVIDIA GeForce GTX 1070 (8GB) graphics card. The cuda version is 10.1, and the cudnn version is 7.6.5; the TensorFlow 2.0 deep learning framework is used.
[0202] In the preparatory stage of model training, the training learning rate of the contrast learning network in the auxiliary task is set to 0.01, the classification training learning rate of the downstream task is set to 0.001, and the batch size is set to 128. The number of training iterations is set to 1000 times. Set τ = 0.1 in the NT-Xent loss function. The window size in the SSIM algorithm is set to 11×11. The size of the face expression image dataset collected in this embodiment is 3450. After random division in a ratio of 4:1, the size of the unannotated face expression dataset is 2760, and the size of the annotated face expression dataset is 690.
[0203] During the auxiliary task learning process, the training logs of the contrast learning network are visually analyzed, and the loss change of the contrast learning is as Figure 9 shown.
[0204] From the change of the contrast loss, it can be seen that the loss value gradually decreases with the iteration. This shows that based on the SSIM algorithm, the contrast learning based on the Efficient-CapsNet encoder is effective. Its contrast loss starts to converge around 400 iterations and fluctuates between 0.5 and 0.9 until then.
[0205] Furthermore, in this embodiment, a comparative study was also conducted on the matrix similarity and the cosine similarity of the vector after flattening the matrix during the contrast training process. The change curves of the cosine similarity and the matrix similarity of the output representation are as Figure 10 shown. The similarity gradually increases with the increase of the number of iterations, indicating that the similarity measurement of the representation matrix based on the SSIM algorithm is effective. The specific experimental results are shown in Table 1. The similarity of the representation matrix almost remains at about 0.74 at convergence, while the cosine similarity measurement after flattening the matrix is far less than the similarity measurement of the representation matrix.
[0206] Table 1 Cosine similarity and matrix similarity of the output representation
[0207]
[0208]
[0209] Note: The symbol ± indicates the floating range of the average value under multiple experiments.
[0210] In the downstream classification task, the changes in the training loss and validation loss of the classification are as Figure 11 shown. The loss of the model converges after more than twenty times, indicating that during the contrast training process, the model has learned certain target attribute features, and in the current supervised training, more learning is to make up for and correct the attribute features of the current specific task. In the classification training, a supervised training method is adopted. When the sample size of the classification training is not large, it is difficult to learn more target features. Therefore, in the classification training, the loss value will converge earlier. That is to say, in the supervised model fine-tuning training, when the sample size is not large, more training iteration times will not bring better results.
[0211] The change in the classification accuracy is as Figure 12 shown. The change in the accuracy is consistent with the change in the loss, and it converges after more than twenty iterations. The highest validation accuracy reaches 70.4%.
[0212] At the same time, in order to better evaluate the generalization ability of the automatic facial expression annotation model of the present invention, in this embodiment, the trained automatic facial expression annotation model is also applied to an unlabeled facial expression dataset in the same acquisition environment, with the expectation of achieving an automatic annotation target of 70%. The new unlabeled facial expression dataset has a size of 600 images, and the male-female ratio is 1:1. After experiments, the automatic annotation results are shown in Table 2 below. In order to better evaluate the effect of the automatic annotation model, the members of the laboratory were also organized to manually adjust and correct the pre-labeled elderly expression dataset again, and the class distribution of the manually adjusted labeled dataset is shown in Table 3. Comparing the manually adjusted labeled dataset with the initial automatic annotation results, the confusion matrix of the comparison results is as Figure 13 shown, and the accuracy of the model's automatic annotation reaches 70.8%, meeting the expected results of the automatic annotation task.
[0213] Table 2 Distribution of the number of automatically annotated classes on the new unlabeled facial expression dataset
[0214]
[0215] Table 3 Distribution of the number of manually annotated classes on the new unlabeled facial expression dataset
[0216]
[0217] In summary, the present invention trains an automatic annotation model using a self-supervised method, overcoming the defects of low efficiency and uneven results caused by subjective differences between different annotators in the pure manual annotation of the current facial expression dataset. The method used in the present invention is a self-supervised learning method. When facing a large amount of unlabeled data, the auxiliary tasks of self-supervised learning can learn a large amount of intrinsic attribute information of the data from the unsupervised data, making full use of the data resources, so as to make full use of the superior performance of the pre-trained model on a small amount of labeled data in downstream tasks. The method provided by the present invention has universality in application, does not target a specific hardware environment, and only needs to meet the basic software dependency packages. And the method of the present invention has good scalability and is not limited to a specific data source scenario.
[0218] In the data input stage of the contrastive learning of the present invention, the input data examples need to be randomly augmented to obtain two related views of the same example. In the self-supervised learning used in the present invention, the encoder network of the auxiliary task adopts an Efficient-CapsNet encoder.
[0219] The Efficient-CapsNet of the present invention adds an attention mechanism routing and depthwise separable convolution operation on the basis of the capsule network. While ensuring the recognition accuracy, it greatly reduces the network parameters and improves the training efficiency of the network. Capsules in the capsule network are a form of feature representation. They can store the attribute information of different objects from different perspectives and have equivariance. Capsules store the attribute information of the object in the form of vectors, such as the size and direction angle of the target entity. At the same time, the capsule vector can also represent the existence or non-existence of the object.
[0220] The present invention first inputs the feature representation into a non-linear projection transformation layer to remove redundant and irrelevant information in the features, thereby revealing the essential attributes of the sample data. Then, the representation after non-linear projection transformation is subjected to contrastive learning, and the learning parameters of the encoder are continuously updated through contrastive feedback. The self-supervised learning method of the present invention adopts a discriminative self-supervised learning method. Discriminative self-supervised learning expects that the data representation contains enough information to find the differences between data through discriminative tasks, and then find the classification boundary.
[0221] The SSIM of the present invention uses the weighted mean and variance of matrix elements to describe the structural information of the matrix, and uses covariance to describe the mutual relationship between the distributions of two matrix elements. According to the idea of NCE, the positive and negative samples can be classified as positive and negative using the data distribution relationship between positive and negative samples. SSIM jointly calculates the similarity of two matrices using the weighted mean, variance, and covariance.
[0222] Since the difference between the matrix and the vector of the present invention is mainly reflected in the large span of matrix elements and cannot calculate the mean and variance of elements in the form of a vector as a whole, in the SSIM algorithm, a sliding window is used to divide the matrix into blocks, calculate the SSIM for each block separately, and finally average the SSIM values of each block, avoiding the phenomenon of large fluctuations in the mean and variance. The present invention solves the technical problems existing in the prior art, such as relying on manual annotation and the uneven results caused by the subjective differences of manual annotation.
[0223] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatic facial expression annotation, characterized in that, The method includes: S1. Obtain image frames of facial expression images using a preset image acquisition device, form a data set with the image frames, eliminate abnormally acquired facial images in the data set, select the facial images corresponding to the peak expressions in the data set as the facial expression image data set, and preprocess the facial expression image data set; S2. Divide the facial expression image data set according to a preset division ratio. Among them, the subsets obtained by the foregoing division operation include: a data set to be labeled and an unlabeled data set. Manually perform emotional labeling on the data set to be labeled to obtain a supervised training data set, and use the unlabeled data set as the training data set for self-supervised learning; S3. Construct a self-supervised labeling model based on Efficient-CapsNet, use the Efficient-CapsNet encoder as the feature extraction encoder in the self-supervised learning auxiliary task, and perform contrastive learning to obtain an optimal pre-trained model. The step S3 includes: S31. Perform data augmentation processing on the facial expression images to obtain images to be encoded, and process the images to be encoded using the Efficient-CapsNet encoder to obtain image feature representation data; S32. Perform contrastive learning based on the image feature representation data to construct an auxiliary task of the self-supervised labeling model, set auxiliary training parameters, input the training data set of the self-supervised learning into the auxiliary task, and perform iterative contrastive training to obtain and save the optimal pre-trained model; S4. In the downstream task of self-supervision, combine the optimal pre-trained model with a preset classifier, and perform supervised training and preset adjustment operations on the supervised training data set to obtain an automatic labeling model. The step S4 includes: S41. Construct the downstream task of the self-supervised labeling model, where the downstream task includes: a downstream task encoder and a downstream task classifier; S42. Set downstream training parameters, input the supervised training data set into the downstream task, combine the downstream task classifier and the optimal pre-trained model, and perform supervised iterative training to obtain and save the automatic labeling model; S5. Use the automatic labeling model to perform automatic emotional labeling on the facial expression images to obtain the automatic facial expression labeling result.
2. The method for automatically annotating human face expressions according to claim 1, characterized in that, The step S1 includes: S11. Eliminate non-frontal face images and face-less images in the data set, and select the peak expression facial images in the remaining data set as the final facial expression data set; S12. Crop the facial images in the data set and uniformly adjust the facial images to a preset size; S13. Use a face detector to detect the faces in the facial images, perform alignment operations on the faces, and generate the facial expression image data set using the aligned faces.
3. A method for automatic facial expression annotation according to claim 1, characterized in that, The step S2 includes: S21. Divide the facial expression data set into a small-ratio data set and a large-ratio data set according to the preset division ratio, where the preset division ratio includes: 4:1; S22. Use the small - scale dataset as the dataset to be labeled; S23. Use the large - scale dataset as the unlabeled dataset; S24. Manually perform sentiment annotation on the dataset to be labeled, and use the unlabeled dataset as the training dataset for self - supervised learning.
4. A method for automatic facial expression annotation according to claim 1, characterized in that, The step 31 includes: S311. The first part is the data augmentation part. The input image of the model will be randomly augmented twice, and the two augmented input images will be input into the pre - set network simultaneously for parallel learning; S312. Use the Efficient - CapsNet encoder to extract features from the input image to obtain feature representations of two types of images. Among them, the Efficient - CapsNet encoder includes: a convolutional layer, a depth - wise convolutional layer, a primary capsule layer, an FCCaps layer. The Efficient - CapsNet also includes a self - attention mechanism routing. After passing through the FCCaps layer, an image representation matrix is output, and the size of the image representation matrix is: the number of classes×16.
5. The method for automatically annotating human face expressions according to claim 4, characterized in that The step S312 includes: S3121. The input image enters the convolutional layer of the Efficient - CapsNet encoder. The input image is grayscaled and sent to four pre - set convolutional layers for processing to obtain the encoder convolutional output feature map; S3122. Use the batch normalization method to normalize each neuron in the pre - set network layer, and use the following transformation reconstruction algorithm to process to obtain the normalization result: Among them, is the input after k-layer normalization, and γ and β are a pair of introduced parameters, which are learned together with the model parameters; S3123. Use the following logic to restore the feature distribution learned by each layer: Embed a batch normalization layer between the convolutional layers of the Efficient - CapsNet, and adopt a weight - sharing method on the batch normalization layer during the convolutional operation. Process the encoder convolutional output feature map through the way of neurons to uniform the data distribution within the layer; S3124. Perform a depth - wise separable convolution operation on the encoder convolutional output feature map to construct primary capsules; S3125. Adopt the self - attention mechanism routing in the FCCaps layer to process and obtain the image feature representation data.
6. A method for automatic annotation of facial expressions according to claim 5, characterized in that, In the step S3125, the following logic is used to process and obtain the feature representation: Among them, B l is a prior matrix, and all capsules in the (l + 1)-th layer are calculated using the following logic Squeeze the lengths of all capsule vectors of the l+1 layer to between 0 and 1 through the following squeezing function to obtain Among them, C l is the coupling coefficient matrix generated by the self-attention mechanism algorithm, and n l represents that there are n l capsules in the l-th layer, and n l+1 represents that there are n l+1 capsules in the l-th layer, and d l is the dimension of the capsules in the l-th layer.
7. A method for automatic annotation of human facial expressions according to claim 1, characterized in that, The step S32 includes: S321. First, input the feature representation into a pre - set non - linear projection transformation layer for non - linear projection transformation to eliminate redundant and irrelevant information in the feature representation, so as to obtain sample representation attribute data; S322. Perform contrast learning on the sample representation attribute data, update the network through contrast feedback, and continuously update the learning parameters of the Efficient - CapsNet encoder to obtain the optimal pre - trained model.
8. A method for automatic facial expression annotation according to claim 7, characterized in that, The step S322 includes: S3221. Input the image representation matrix into a two - layer non - linear MLP (Dense->Relu->Dense) layer to map the image representation matrix into the space of the contrast loss; S3222. Use a sliding window to block the matrix, and calculate the variance of each window block separately: where ω i is the Gaussian kernel weight, and N is the number of elements in the window block; S3223. Calculate the covariance of the window blocks b and b′ corresponding to the two image representation matrices with the following logic: S3224. Process the variance and the covariance of the window blocks with the following logic to obtain the SSIM value: where c1 = (k1L) 2 and c2 = (k2L) 2 are two variables used to stabilize division, where L in c1 and c2 is the dynamic range of matrix element values, and k1 and k2 are hyperparameters; S3225. Average the SSIM of all the window blocks, and obtain an average value based on this, which is used as the overall similarity of the image representation matrix: Where B is the number of matrix sliding window blocks, z and z′ are the input feature matrices, and z i and z′ i are the feature matrices of the i-th window block corresponding to the two feature matrices; S3226. Use the adjustable temperature-normalized cross-entropy loss based on the SSIM algorithm, and with the following cosine similarity metric transformation logic, calculate the overall similarity by comparing the loss to obtain the SSIM matrix similarity metric comparison loss function: S3227. Feed back the updated network according to the SSIM matrix similarity metric comparison loss function, and obtain the optimal pre-trained model based on this:
9. A method for automatic annotation of facial expressions according to claim 1, characterized in that The step S5 includes: S51. Input the pre-processed face expression image dataset into the input of the network, the automatic annotation model; S52. Perform auxiliary task learning on the unlabeled dataset in the face expression dataset, and save the encoder model with the optimal comparison effect through iterative training; S53. In the downstream task, input the dataset to be labeled into the pre-trained model, and perform iterative supervised training through a supervised learning strategy to obtain and save the optimal annotation model; S54. Combine the encoder model and the annotation model to obtain the optimal automatic annotation model, and obtain the face expression automatic annotation result based on this:
10. A system for automatic annotation of facial expressions, characterized in that, The system includes: An expression dataset module, which is used to obtain the image frames of face expression images with a preset image acquisition device, form a dataset with the image frames, eliminate the abnormal face acquisition images in the dataset, select the face images corresponding to the peak expressions in the dataset as the face expression image dataset, and pre-process the face expression image dataset; A dataset division module, which is used to divide the face expression image dataset according to a preset division ratio. Among them, the subsets obtained from the above division operation include: a dataset to be labeled and an unlabeled dataset. Manually perform sentiment annotation on the dataset to be labeled to obtain a supervised training dataset, and use the unlabeled dataset as the training dataset for self-supervised learning. The dataset division module is connected to the expression dataset module; An optimal pre-trained model acquisition module, which is used to construct a self-supervised annotation model based on Efficient-CapsNet, use the Efficient-CapsNet encoder as the feature extraction encoder in the self-supervised learning auxiliary task, and perform contrast learning to obtain the optimal pre-trained model. The optimal pre-trained model acquisition module is connected to the dataset division module. The optimal pre-trained model acquisition module includes: A feature representation module, which is used to perform data augmentation on the face expression image to obtain an image to be encoded, and process the image to be encoded with the Efficient-CapsNet encoder to obtain image feature representation data; A self-supervised learning module, which is used to perform contrastive learning based on the data of the image feature representation, construct an auxiliary task of the self-supervised annotation model accordingly, set auxiliary training parameters, input the training data set of the self-supervised learning into the auxiliary task to perform iterative contrastive training, and obtain and save the optimal pre-trained model accordingly. The self-supervised learning module is connected to the feature representation module; An automatic annotation model acquisition module, which is used to combine the optimal pre-trained model with a preset classifier in a self-supervised downstream task, and perform supervised training and preset adjustment operations on the supervised training data set to obtain an automatic annotation model. The automatic annotation model acquisition module is connected to the data set division module. The automatic annotation model acquisition module includes: A downstream task construction module, which is used to construct the downstream task of the self-supervised annotation model. Among them, the downstream task includes: a downstream task encoder and a downstream task classifier; A supervised iterative training module, which is used to set downstream training parameters, input the supervised training data set into the downstream task, and combine the downstream task classifier and the optimal pre-trained model to perform supervised iterative training, and obtain and save the automatic annotation model accordingly. The supervised iterative training module is connected to the downstream task construction module; An automatic annotation module, which is used to automatically annotate the emotion of the facial expression image with the automatic annotation model to obtain an automatic facial expression annotation result. The automatic annotation module is connected to the expression data set module, the optimal pre-trained model acquisition module and the automatic annotation model acquisition module.
Citation Information
Patent Citations
Multi-dimensional emotion recognition method and system
CN113780341A
Face attribute data labeling method, computer equipment and storage medium
CN114332136A
Facial expression migration method based on self-supervised learning and generative adversarial mechanism
CN111243066A
Pulmonary nodule auxiliary diagnosis method based on three-dimensional multi-resolution attention capsule network
CN113208641A