Asymmetric facial expression recognition

CN117769725BActive Publication Date: 2026-09-25FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280046259.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-13
Filing Date
2022-07-26
Publication Date
2026-09-25
Estimated Expiration
2042-07-26

Smart Images

  • Figure CN117769725B_ABST
    Figure CN117769725B_ABST
Patent Text Reader

Abstract

This disclosure describes techniques for facial expression recognition. A first loss function can be determined based on a first set of feature vectors associated with a first set of images depicting facial expressions and a first set of labels indicative of the facial expressions. A second loss function can be determined based on a second set of feature vectors associated with a second set of images depicting asymmetric facial expressions and a second set of labels indicative of the asymmetric facial expressions. The first loss function and the second loss function can be used to determine a maximum loss function. The maximum loss function can be applied during training of a model. The trained model can be configured to predict at least one asymmetric facial expression in a subsequently received image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Application No. 17 / 402,344, filed August 13, 2021, entitled ASYMMETRIC FACIAL EXPRESSION RECOGNITION, the entire contents of which are incorporated herein by reference. Background Technology

[0003] Image recognition refers to a set of automated methods for detecting and analyzing images to support specific tasks. It is a technique capable of identifying places, people, objects, and many other types of elements within an image and drawing conclusions based on their analysis. Improvements in image recognition technology are desirable. Attached Figure Description

[0004] The following detailed description can be better understood when read in conjunction with the accompanying drawings. Example embodiments of various aspects of this disclosure are shown in the drawings for illustrative purposes; however, the invention is not limited to the specific methods and tools disclosed.

[0005] Figure 1 An example system for distributing content is shown.

[0006] Figure 2 An example of facial expression analysis is shown.

[0007] Figure 3 Example sets of facial expressions are shown, some symmetrical and some asymmetrical.

[0008] Figure 4 An example method for facial expression recognition is shown.

[0009] Figure 5 Example images of the training dataset used to enhance facial expression recognition models are shown.

[0010] Figure 6 A set of example images depicting various facial expressions is shown.

[0011] Figure 7 An example application of the facial expression recognition model is shown.

[0012] Figure 8 An example computing device is shown that can be used to perform any of the techniques disclosed herein. Detailed Implementation

[0013] Facial expressions play a crucial role in conveying nonverbal information about human feelings and / or emotions. Therefore, facial expression recognition is becoming increasingly prevalent. Facial expression recognition can be used in many practical applications, including but not limited to human-computer interaction and facial animation. As the adoption rate of facial expression recognition continues to increase, improvements in facial expression recognition technology are likely to follow.

[0014] Many facial expressions are asymmetrical. For example, for many facial expressions, the left side of the face may be more expressive than the right side (and vice versa). As a result, when making an expression, the left side of the face may look different from the right side. For example, the right eye may move in a different way than the left eye, and the right side of the mouth may move in a different way than the left side of the mouth, and so on. Current facial expression recognition technologies may struggle to analyze such asymmetrical facial expressions. Therefore, improvements in asymmetrical facial expression recognition technology are desirable.

[0015] Some facial expression recognition technologies attempt to improve the performance of asymmetric facial expression recognition by augmenting the dataset used to train models (for predicting facial expressions). Image data augmentation is perhaps the most well-known type of data augmentation and involves creating transformed versions of images in the training dataset that belong to the same category as the original images. Transformations can include a range of operations from the realm of image manipulation, such as shifting, flipping, scaling, etc. The intention is to expand the training dataset with new, plausible examples. This means changes that the model might see in the training set images. For example, a horizontal flip of a photo of a cat might make sense because the photo could have been taken from the left or right. Modern deep learning algorithms (such as convolutional neural networks or CNNs) can learn features in an image that are unaffected by its position. Augmentation can further assist this transformation-invariant approach to learning and can also help the model learn transformation-invariant features, such as left-to-right to top-to-bottom ordering, lighting levels in a photograph, etc.

[0016] To increase the number of asymmetric samples in the training dataset, the training set can be augmented. For example, data augmentation can be achieved through spatial transformations, such as horizontally flipping a face image and swapping attributes associated with the left and right sides of the face (e.g., swapping the activation signals of left_eye_closed and right_eye_closed). By augmenting the training set, the model gains more learning experience and thus performs better in predicting asymmetric expressions.

[0017] However, augmenting the training dataset in this way doubles the storage required for the training data. Therefore, an improved asymmetric facial expression recognition technique with good performance but requiring less storage is desired. Asymmetric loss can be introduced during model training to help refine asymmetric expressions. The asymmetric loss can act as an indirect augmentation of the original input. A first loss function and ground truth labels associated with the original expression parameters can be determined. A second loss function and ground truth labels associated with the asymmetric expression parameters can be determined. Using the first and second loss functions, a maximum loss function can be determined. This maximum loss function (i.e., the asymmetric loss function) can be used to train the model for predicting facial expressions. By training the model using the maximum loss function, it minimizes the risk of the model reaching a local minimum during application (e.g., when predicting facial expressions).

[0018] Facial expression recognition models (such as those trained using asymmetric loss) can be used by a variety of different systems or entities. For example, content distributors can leverage such models for facial expression recognition. Figure 1 An example system 100 for distributing content is illustrated. System 100 may include a cloud network 102 and multiple client devices 104a to 104d. The cloud network 102 and the multiple client devices 104a to 104d may communicate with each other via one or more networks 120.

[0019] Cloud network 102 may be located at a data center, such as in a single location, or distributed across different geographical locations (e.g., in multiple locations). Cloud network 102 may provide services via one or more networks 120. Network 120 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, etc. Network 120 may include physical links, such as coaxial cable links, twisted-pair cable links, fiber optic links, and combinations thereof. Network 120 may include wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0020] Cloud network 102 may include multiple computing nodes 118 hosting various services. In one embodiment, node 118 hosts content service 112. Content service 112 may include content streaming services, such as Internet Protocol video streaming services. Content service 112 may be configured to distribute content 116 via various transport technologies. Content service 112 is configured to provide content 116, such as video, audio, text data, combinations thereof. Content 116 may include content streams (e.g., video streams, audio streams, information streams), content files (e.g., video files, audio files, text files), and / or other data. Content 116 may be stored in database 114. For example, content service 112 may include video sharing services, video hosting platforms, content distribution platforms, collaborative gaming platforms, etc.

[0021] In this embodiment, the content 116 distributed or provided by the content service 112 includes short videos. The duration of a short video may be less than or equal to a predetermined time limit, such as one minute, five minutes, or other predetermined minutes. By way of example, and not limitation, a short video may include at least one, but no more than four, 15-second clips strung together. The short duration of the video can provide viewers with a rapid burst of entertainment, allowing users to watch a large amount of video within a short timeframe. This rapid burst of entertainment may become popular on social media platforms.

[0022] In this embodiment, content 116 can be output to different client devices 104 via network 120. Content 116 can be streamed to client devices 104. The content stream can be a short video stream received from content service 112. Multiple client devices 104 can be configured to access content 116 from content service 112. In this embodiment, client devices 104 may include content application 106. Content application 106 outputs (e.g., displays, renders, presents) content 116 to a user associated with client device 104. The content may include video, audio, comments, text data, etc.

[0023] Multiple client devices 104 can include any type of computing device, such as mobile devices, tablets, laptops, desktop computers, smart TVs or other smart devices (e.g., smartwatches, smart speakers, smart glasses, smart helmets), gaming devices, set-top boxes, digital streaming devices, robots, etc. Multiple client devices 104 can be associated with one or more users. A single user can use one or more of the multiple client devices 104 to access the cloud network 102. Multiple client devices 104 can travel to various locations and use different networks to access the cloud network 102.

[0024] In this embodiment, a user can use a content application 106 on client device 104 to create content and upload short videos to cloud network 102. Client device 104 can access interface 108 of content application 106. Interface 108 may include input elements. For example, the input elements may be configured to allow a user to create content. To create content, the user can grant content application 106 permission to access image capture devices (such as cameras) or microphones on client device 104. After the user has created content, the user can use content application 106 to upload the content to cloud network 102 and / or save the content locally to user device 104. Content service 112 may store the uploaded content and any metadata associated with that content in one or more databases 114.

[0025] Multiple compute nodes 118 can handle tasks associated with content service 112. Multiple compute nodes 118 can be implemented as one or more compute devices, one or more processors, one or more virtual compute instances, or combinations thereof. Multiple compute nodes 118 can be implemented by one or more compute devices. One or more compute devices can include virtualized compute instances. Virtualized compute instances can include virtual machines, such as simulations of computer systems, operating systems, servers, etc. Virtual machines can be loaded by the compute device based on virtual images and / or other data defining specific software used for simulation (e.g., operating systems, dedicated applications, servers). As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more compute devices. A hypervisor can be implemented to manage the use of different virtual machines on the same compute device.

[0026] In this embodiment, content service 112 includes model 110. Model 110 may be, for example, a machine learning model. Model 110 may be used at least in part to predict and / or analyze facial expressions, including asymmetrical facial expressions. As discussed above, conventional facial expression recognition techniques perform poorly in predicting asymmetrical facial expressions and / or require significant storage space.

[0027] Model 110 may be able to accurately predict asymmetric facial expressions without training on augmentation datasets that require significant storage. Instead, model 110 can be trained using asymmetric loss. Asymmetric loss can act as an indirect augmentation of the original input. A first loss function and ground truth labels associated with the original expression parameters can be determined. A second loss function and ground truth labels associated with the asymmetric expression parameters can be determined. Using the first and second loss functions, a maximum loss function can be determined. This maximum loss function can be used to train the model for predicting facial expressions. By training the model using the maximum loss function, it minimizes the risk of the model reaching local minima during application (e.g., when predicting facial expressions). As a result, the quality of facial expression recognition can be improved.

[0028] Figure 2 An example facial expression analysis 200 is illustrated. Facial expression analysis 200 can be generated, for example, from a facial expression model (e.g., model 110). Facial expression analysis 200 can indicate an analysis 206 of the expression associated with face 202. The analysis 206 of the expression associated with face 202 can be generated, for example, based on the analysis of various facial features 204a to 204f. For example, facial expression analysis 206 can indicate that the expression is one of happiness, sadness, anger, surprise, pain, or any other emotion or feeling. Whether facial expression analysis 206 indicates that the expression is one of happiness, sadness, anger, surprise, pain, etc., can depend on how the various facial features 204a to 204f move or are positioned on face 202.

[0029] However, Figure 2 The expressions depicted on the face 202 are largely symmetrical (e.g., the left side of the face looks the same as the right side). As discussed above, models may have more difficulty predicting asymmetrical facial expressions (e.g., expressions where the left side of the face does not look the same as the right side). For example, when a person makes an expression indicating happiness or surprise, the left side of his / her face may look slightly different from the right side—one eye / eyebrow may be slightly raised than the other, and one corner of the mouth may be more upturned than the other. Due to the asymmetry that can be found in facial expressions, traditional facial expression recognition models may struggle to accurately predict the facial expressions being made.

[0030] Throughout the video, the individual may make a variety of different facial expressions. Figure 3 The illustration shows an example set of facial expressions 300. The set of facial expressions 300 can include any number of symmetrical and / or asymmetrical facial expressions, such as... Figure 3The four facial expressions 302, 304, 306, and 308 depicted in the video are possible. Facial expressions 302, 304, 306, and 308 can be facial expressions made by an individual throughout the video. For example, during the video, an individual could first make facial expression 302, then facial expression 304, then facial expression 306, and then facial expression 308. However, it should be understood that an individual can make facial expressions in any other order and / or can make different or additional facial expressions throughout the video (some of which may be asymmetrical or not asymmetrical).

[0031] Some facial expressions in the facial expression set 300 may be symmetrical, while others may be asymmetrical. For example, expressions 302 and 304 are largely symmetrical, while expressions 306 and 308 are asymmetrical. In expression 306, the left side of an individual's mouth moves in a different manner than the right side of their mouth. Similarly, in expression 308, the left side of an individual's mouth moves in a different manner than the right side of their mouth. As discussed above, traditional facial expression recognition models may struggle to accurately predict facial expressions 306 and 308.

[0032] As discussed above, some facial expression recognition techniques attempt to improve the performance of asymmetric facial expression recognition by augmenting the dataset used to train the model (for predicting facial expressions). To increase the number of asymmetric samples in the training dataset, the training set can be augmented. For example, data augmentation can be achieved through spatial transformations, such as horizontally flipping a face image and swapping attributes associated with the left and right sides of the face (e.g., swapping the activation signals of left_eye_closed and right_eye_closed). By augmenting the training set, the model gains more learning experience and therefore performs better in predicting asymmetric expressions.

[0033] However, augmenting the training dataset in this way doubles the storage required for the training data. Therefore, an improved asymmetric facial expression recognition technique that performs well without requiring a large amount of storage space is desired. Figure 4 The illustration depicts an example process 400 performed by a computing device. The computing device can execute process 400 to train a facial expression recognition model (e.g., model 110) in a manner that does not require a large amount of storage space. Once trained, the facial expression recognition model can be used to predict facial expressions, including asymmetric facial expressions (such as...). Figure 3 Facial expressions shown (306, 308). Although in Figure 4The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.

[0034] Process 400 introduces an asymmetric loss function during model training to help refine asymmetric expressions. The asymmetric loss function can act as an indirect enhancement of the original input. The asymmetric loss function is associated with and determined based on two different loss functions. In 402, the first loss function can be determined based on a first set of feature vectors associated with a first set of images depicting facial expressions and a first set of labels indicating facial expressions. For example, the first set of images can be taken from a video. The facial expressions depicted by the first set of feature vectors can include asymmetric facial expressions (such as...). Figure 3 Facial expressions shown are 306 and 308.

[0035] The first feature vector set can be generated by a model based on a first image set. The first feature vector set may include a list of numbers used to characterize facial expressions in the first image set recognized by the model. A first label set (i.e., ground truth data) indicates realistic or true facial expressions in the first image set. The first feature vector set and the first label set may include a set of blending shape coefficients. Each blending shape coefficient may be associated with a specific facial feature or facial region. Blending shape coefficients may include, but are not limited to, blending shape coefficients associated with eyes, mouth / chin, eyebrows / cheeks / nose, or tongue.

[0036] The mixed shape coefficients associated with the eyes may include: coefficients describing the closure of the upper left eyelid, coefficients describing the movement of the left eyelid in accordance with downward gaze, coefficients describing the movement of the left eyelid in accordance with right gaze, coefficients describing the movement of the left eyelid in accordance with left gaze, coefficients describing the movement of the left eyelid in accordance with upward gaze, coefficients describing facial contraction around the left eye, coefficients describing eyelid widening around the left eye, coefficients describing the closure of the upper right eyelid, coefficients describing the movement of the right eyelid in accordance with downward gaze, coefficients describing the movement of the right eyelid in accordance with left gaze, coefficients describing the movement of the right eyelid in accordance with right gaze, coefficients describing the movement of the right eyelid in accordance with upward gaze, coefficients describing facial contraction around the right eye, and / or coefficients describing eyelid widening around the right eye.

[0037] The mixed shape coefficients associated with the mouth / chin may include: coefficients describing forward jaw movement, coefficients describing leftward jaw movement, coefficients describing rightward jaw movement, coefficients describing jaw opening, coefficients describing lip closure regardless of chin position, coefficients describing the contraction of both lips into an open shape, coefficients describing the contraction and compression of both closed lips, coefficients describing both lips moving together to the left, coefficients describing both lips moving together to the right, coefficients describing upward movement of the left corner of the mouth, coefficients describing upward movement of the right corner of the mouth, coefficients describing downward movement of the left corner of the mouth, and coefficients describing downward movement of the right corner of the mouth. The coefficients describing the backward movement of the left corner of the mouth, the backward movement of the right corner of the mouth, the left left corner of the mouth moving to the left, the left right corner of the mouth moving to the right, the coefficients describing the inward movement of the lower lip, the inward movement of the upper lip, the outward movement of the lower lip, the outward movement of the upper lip, the upward pressing of the left upper lip, the upward pressing of the right lower lip, the downward movement of the left lower lip, the downward movement of the right lower lip, the upward movement of the left upper lip, and / or the downward movement of the right upper lip.

[0038] The blending shape coefficients associated with eyebrows / cheeks / nose or tongue may include: coefficients describing the downward movement of the outer left eyebrow, coefficients describing the downward movement of the outer right eyebrow, coefficients describing the upward movement of the inner sides of both eyebrows, coefficients describing the upward movement of the inner left eyebrow, coefficients describing the downward movement of the outer right eyebrow, coefficients describing the outward movement of both cheeks, coefficients describing the upward movement of the cheeks around and below the left eye, coefficients describing the upward movement of the cheeks around and below the right eye, coefficients describing the elevation of the left side of the nose around the nostrils, coefficients describing the elevation of the right side of the nose around the nostrils, and / or coefficients describing the extension of the tongue.

[0039] By using examples rather than restrictions, the first loss function L δ(a) We can use L1 smoothed Huber loss, as follows:

[0040]

[0041] Here, 'a' represents the difference between the first label space (i.e., the ground truth data) and the first feature vector set (e.g., the predicted mixed shape result). The L-1 smooth Hubel loss function is a loss function used for robust regression that is less sensitive to outliers in the data than the squared error loss. This function is quadratic for small values ​​of 'a' and linear for large values; at two points where the absolute value of 'a' equals δ, the values ​​and slopes of different cross sections are equal. The first feature vector set (i.e., the predicted result) can be generated based on the first image set. It can be passed through the model via F(A,L) during each step of the network (e.g., a neural network).δ(a) Generate the first feature vector set, where Let L represent the input image matrix associated with the first image set, and L δ(a) This represents the first loss function.

[0042] In 404, the second loss function L can be determined based on the second feature vector set associated with the second image set depicting asymmetrical facial expressions and the second label set indicating asymmetrical facial expressions. δ(b) The second loss function L is presented through examples rather than restrictions. δ(b) This can be based on L1 smoothed Hubel loss. The second feature vector set can be generated by the model based on a second image set. For example, the second feature vector set can be generated via F(A i(n+1-j) ,L δ(b) ) generated, where A i(n+1-j) Let represent the input image matrix associated with a second set of images depicting asymmetrical facial expressions (e.g., horizontally flipped images of a subset of the first image set), b represent the difference between the second label set (i.e., ground truth data) and the second feature vector set (e.g., the predicted blended shape result), and L δ(b) This represents the second loss function.

[0043] The asymmetric loss function (i.e., the global loss function) can be generated based on the first loss function and the second loss function. In 406, the maximum loss function (i.e., the asymmetric loss function) can be determined based on the first and second loss functions. For example, the maximum loss function (i.e., the asymmetric loss function) L asymmetryδ(a,b) It can be represented as follows:

[0044] L asymmetryδ(a,b) =max(L δ(a) L δ(b) )

[0045] Where L δ(a) Let L represent the first loss function, and L... δ(b) This represents the second loss function.

[0046] By way of example rather than limitation, the second tag set can be generated based on at least one flipped portion of an image from a subset of the first image set. For example, the second tag set C' gt (e) can be generated based on the following

[0047]

[0048] Where e represents the asymmetric mixed-shape expression associated with the second image set, and B represents all possible mixed-shape expressions associated with the first image set.

[0049] In one embodiment, a second image set can be generated via data augmentation of a subset of a first image set including asymmetrical facial expressions. For example, the subset of the first image set depicts asymmetrical facial expressions (such as...). Figure 3 The facial expressions shown are 306 and 308. The second image set may be similar to an image in a subset of the first image set, but may have at least one portion flipped (e.g., horizontally flipped) relative to an image in the subset of the first image set. The second tag set may be generated based on at least one flipped portion of an image in a subset of the first image set. In another embodiment, the second image set may be further enhanced by flipping (e.g., horizontally flipping) at least one portion of at least one image in the second image set. The second tag set may also be enhanced based on at least one flipped portion of at least one image.

[0050] Figure 5 Example images and example flipped images are depicted. For example, first image 502 depicts an asymmetrical facial expression. The face in image 502 raises the right corner of the mouth higher than the left corner. This is indicated by the blending shape coefficients: the coefficient describing the upward movement of the right corner of the mouth (0.735) is greater than the coefficient describing the upward movement of the left corner of the mouth (0.12). First image 502 is also associated with a coefficient of 0.56 describing the upward movement of the cheek around and below the right eye. Second image 504 is a horizontally flipped version of first image 502. Blending shape coefficients associated with image 504 are generated based on at least some of the blending shape coefficients associated with flipping image 502. For example, in second image 504, the coefficient describing the upward movement of the left corner of the mouth is 0.735, the coefficient describing the upward movement of the right corner of the mouth is 0.12, and the coefficient 0.56 describes the upward movement of the cheek around and below the left eye.

[0051] The second feature vector set may include a list of numbers used to characterize asymmetric facial expressions in the second image set recognized by the model. The second label set (i.e., ground truth data) indicates realistic or true asymmetric facial expressions in the second image set. Both the second feature vector set and the second label set may include a set of blending shape coefficients. Each blending shape coefficient may be associated with a specific facial feature or facial region. Blending shape coefficients may include, but are not limited to, blending shape coefficients associated with eyes, mouth / chin, eyebrows / cheeks / nose, or tongue.

[0052] As described above, the maximum loss function (i.e., the asymmetric loss function) can be determined based on the first and second loss functions. The maximum loss function can be applied to models such as facial expression recognition models to train them to accurately predict facial expressions including asymmetric facial expressions. In section 408, the maximum loss function can be applied to the model during training. The trained model can be configured to predict at least one asymmetric facial expression in subsequently received images (such as video frames). By training the model using the maximum loss function, it minimizes the risk of the model reaching a local minimum during application (e.g., when predicting facial expressions). Once the maximum loss function is found, the model can be trained on any set of images including asymmetric facial expressions (these images may or may not include the set of images used to derive the maximum loss function).

[0053] Figure 6 An exemplary image set 600 is shown, which can be fed into a system for training a facial expression recognition model. The image set 600 may include images acquired from a video (e.g., various frames acquired from a video). The image set 600 may include any number of images, such as images 602a to 602c. One or more images in the image set 600 may depict faces with asymmetrical expressions. For example, in… Figure 6 In the image 602c, a face with an asymmetrical expression is depicted, as the boy's mouth is upturned at only one corner. The rest of the images (602a to 602b) are largely symmetrical.

[0054] One or more facial features may be detected and / or extracted from each of images 602a to 602c. Facial features may be detected and / or extracted in any suitable manner. For example, any existing algorithms for facial feature extraction may be used, including but not limited to: Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), skin color, wavelet, and / or Artificial Neural Network (ANN).

[0055] In an embodiment, multiple facial feature points can be generated for each of the images 602a to 602c. Each facial feature point may correspond to a specific facial feature and may indicate the location of that facial feature on the image (e.g., a pair of coordinates). Images 604a to 604c depict the facial feature points generated for each of the images 602a to 602c. For example, image 604a depicts the facial feature points generated for image 602a, image 604b depicts the facial feature points generated for image 602b, and image 604c depicts the facial feature points generated for image 602c.

[0056] In an embodiment, each of a plurality of facial feature points may be associated with a number overlaid on a facial image. Each number corresponds to a specific facial region, such as the left eye, right eye, left pupil, right pupil, left eyebrow, right eyebrow, nose, upper lip, lower lip, or the remainder of the face. More than one feature point may correspond to a single facial region. For example, nine feature points in a set of facial feature points may correspond to the right eyebrow, and another nine feature points in the set of facial feature points may correspond to the left eyebrow. Similarly, multiple feature points may correspond to each of the left eye, right eye, left pupil, right pupil, nose, upper lip, lower lip, or the remainder of the face.

[0057] Figure 7 The illustration shows an example system 700 for controlling animation based on predictions of facial expressions output from a trained facial expression recognition model. The system can receive images as input 702. System 700 can apply a neural network (e.g., a trained facial expression model) to generate a feature representation (e.g., a feature vector) associated with each image. By way of example, and not limitation, the neural network can include VGGnet (which has three fully connected layers: the first two layers each have 4076 channels, and the third layer has 1000 channels, one for each class), AlexNet (AlexNet is a convolutional neural network containing eight layers with weights; the first five layers are convolutional, and the remaining three are fully connected. The output of the last fully connected layer is fed into a 1000-way softmax, which produces a distribution over 1000 class labels), GoogLeNet (GoogLeNet is a 22-layer deep convolutional neural network), and / or any other suitable type of neural network.

[0058] For each input image, the initial output of system 700 can be a high-dimensional (e.g., 1000-dimensional) feature vector. Regression can be performed on (multiple) high-dimensional feature vectors. For example, linear regression can be performed on (multiple) high-dimensional feature vectors. Linear regression is a supervised learning algorithm used to predict real-valued outputs. A linear regression model is a linear combination of features from the input examples. For example, the real-valued output can be a 24-dimensional feature vector. A 24-dimensional feature vector can, for example, be a facial action unit (AU) vector. Each 24-dimensional feature vector can indicate a predicted expression associated with an image in the input image.

[0059] The output of system 700 can be a high-dimensional (e.g., 1000-dimensional) feature vector. Regression can be performed on (multiple) high-dimensional feature vectors. For example, linear regression can be performed on (multiple) high-dimensional feature vectors. Linear regression is a supervised learning algorithm used to predict real-valued outputs. A linear regression model is a linear combination of features of the input examples. For example, the real-valued output can be a 24-dimensional feature vector. A 24-dimensional feature vector can, for example, be a facial action unit (AU) vector. Each 24-dimensional feature vector can indicate a predicted expression associated with an image selected from a flipped image set. System 700 can output a prediction 704 of the facial expression associated with the input image.

[0060] Application 706 of the trained model may include using the output of the trained model to control animation. For example, application 706 of the trained model may include using the output of the trained model to control facial animation. The trained model may output predictions associated with one or more input images. Each input image may depict a face. In at least some input images, the face may have asymmetrical facial expressions. In at least some other input images, the face may have symmetrical facial expressions. The trained model may generate predictions associated with the expressions in each input image. For example, the trained model may predict whether the expression associated with each input image is one of admiration, worship, happiness, anger, anxiety, awe, embarrassment, boredom, calmness, confusion, longing, disgust, empathic pain, reverie, excitement, fear, bliss, dread, interest, joy, nostalgia, ease, sadness, satisfaction, surprise, and / or any other emotion or feeling. The predictions associated with each image may be used to control facial animation. For example, if a trained model predicts that the facial expression depicted in an image is fear, then facial animation can be controlled to depict the expression of fear.

[0061] Figure 8 The illustration shows computing devices that can be used in various aspects, such as Figure 1 The services, networks, modules, and / or devices described herein. About Figure 1 In the example architecture, cloud network 102, model 110, and multiple client devices 104a to 104d can be respectively... Figure 8 This is achieved through one or more instances of the computing device 800. Figure 8 The computer architecture shown illustrates conventional server computers, workstations, desktop computers, laptop computers, tablet computers, networked appliances, PDAs, e-readers, digital cellular phones, or other computing nodes, and can be used to perform any aspect of the computer described herein, such as implementing the methods described herein.

[0062] The computing device 800 may include a substrate or “motherboard,” which is a printed circuit board in which multiple components or devices can be connected via a system bus or other electrical communication paths. One or more central processing units (CPUs) 804 may operate in conjunction with a chipset 806. The CPUs 804(s) may be standard programmable processors that perform the arithmetic and logic operations necessary for the operation of the computing device 800.

[0063] Multiple CPUs (804) can perform necessary operations by manipulating switching elements that distinguish and change these states, transitioning from one discrete physical state to the next. Switching elements typically include electronic circuitry (such as flip-flops) that maintains one of two binary states, and electronic circuitry (such as logic gates) that provides an output state based on a logical combination of the states of one or more other switching elements. These basic switching elements can be combined to create more complex logic circuits, including registers, adder-subtractor units, arithmetic logic units, floating-point units, etc.

[0064] The (multiple) CPUs 804 can be enhanced or replaced by other processing units (such as, (multiple) GPUs 805). The (multiple) GPUs 805 may include processing units specifically designed for, but not necessarily limited to, highly parallel computing, such as graphics processing and other visualization-related processing.

[0065] Chipset 806 provides an interface between CPU(s) 804 and the remaining components and devices on the substrate. Chipset 806 provides an interface to Random Access Memory (RAM) 808, which serves as the main memory in computing device 800. Chipset 806 may also provide an interface to computer-readable storage media such as Read-Only Memory (ROM) 820 or Non-Volatile RAM (NVRAM) (not shown), for storing basic routines that help boot computing device 800 and transfer information between various components and devices. ROM 820 or NVRAM may also store other software components necessary for the operation of computing device 800 according to the various aspects described herein.

[0066] Computing device 800 can operate in a networked environment using a logical connection to remote computing nodes and computer systems via a local area network (LAN). Chipset 806 may include functionality for providing network connectivity via a network interface controller (NIC) 822 (such as a Gigabit Ethernet adapter). NIC 822 may be able to connect computing device 800 to other computing nodes via network 816. It is understood that multiple NICs 822 may be present in computing device 800, thereby connecting the computing device to other types of networks and remote computer systems.

[0067] Computing device 800 can be connected to mass storage device 828, which provides non-volatile storage for the computer. Mass storage device 828 can store system programs, application programs, other program modules, and data, which have been described in more detail herein. Mass storage device 828 can be connected to computing device 806 via storage controller 824, which is connected to chipset 806. Mass storage device 828 can consist of one or more physical storage units. Mass storage device 828 may include management components. Storage controller 824 can interface with physical storage units via a Serial Attached SCSI (SAS) interface, a Serial Advanced Technology Attachment (SATA) interface, a Fibre Channel (FC) interface, or other types of interfaces used for physical connections and data transfer between the computer and physical storage units.

[0068] The computing device 800 can store data on the mass storage device 828 by changing the physical state of the physical storage units to reflect the information being stored. The specific changes in physical state may depend on various factors and the different implementations described herein. Examples of such factors may include, but are not limited to, the technology used to implement the physical storage units and whether the mass storage device 828 is characterized as a primary storage device or a secondary storage device.

[0069] For example, computing device 800 can issue instructions via storage controller 824 to store information in mass storage device 828 to change the magnetic characteristics of a specific location within a disk drive unit, the reflection or refraction characteristics of a specific location in an optical storage unit, or the electrical characteristics of a specific capacitor, transistor, or other discrete component in a solid-state storage unit. Other transformations of the physical medium are possible without departing from the scope and spirit of this description, and the foregoing examples are provided only to facilitate the description. Computing device 800 can also read information from mass storage device 828 by detecting the physical state or characteristics of one or more specific locations within a physical storage unit.

[0070] In addition to the aforementioned mass storage device 828, the computing device 800 can access other computer-readable storage media to store and retrieve information, such as program modules, data structures, or other data. Those skilled in the art will understand that a computer-readable storage medium can be any available medium that provides storage for non-transient data and can be accessed by the computing device 800.

[0071] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, transient and non-transient computer-readable storage media, as well as removable and non-removable media, implemented in any method or technology. Computer-readable storage media include, but are not limited to, RAM, ROM, erasable programmable ROM (“EPROM”), electrically erasable programmable ROM (“EEPROM”), flash memory or other solid-state memory technologies, optical disc ROM (“CD-ROM”), digital universal disc (“DVD”), high-definition DVD (“HD-DVD”), Blu-ray disc or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage devices, other magnetic storage devices, or any other medium that can be used to store desired information in a non-transient manner.

[0072] Massive storage devices (such as, Figure 8 The mass storage device 828 described herein can store an operating system used to control the operation of the computing device 800. The operating system may include a version of the LINUX operating system. The operating system may include a version of the WINDOWS server operating system from Microsoft. Depending on other aspects, the operating system may include a version of the UNIX operating system. Various mobile phone operating systems, such as iOS and Android, may also be used. It should be understood that other operating systems may also be used. The mass storage device 828 may store other systems, applications, and data used by the computing device 800.

[0073] Mass storage device 828 or other computer-readable storage medium may also be encoded with computer-executable instructions that, when loaded into computing device 800, transform the computing device from a general-purpose computing system into a special-purpose computer capable of implementing the various aspects described herein. As described above, these computer-executable instructions transform computing device 800 by specifying how CPU(s) 804(s) transition between states. Computing device 800 can access the computer-readable storage medium storing the computer-executable instructions, which, when executed by computing device 800, can perform the methods described herein.

[0074] Computing devices (such as, Figure 8 The computing device 800 depicted may also include an input / output controller 832 for receiving and processing input from various input devices, such as a keyboard, mouse, touchpad, touchscreen, electronic pen, or other types of input devices. Similarly, the input / output controller 832 may provide output to a display, such as a computer monitor, flat panel display, digital projector, printer, plotter, or other types of output device. It should be understood that the computing device 800 may not include... Figure 8 All components shown may include Figure 8 Other components not explicitly shown, or those that can be used with Figure 8 The architecture shown is completely different.

[0075] As described in this article, a computing device can be a physical computing device, such as... Figure 8 The computing device 800. A computing node may also include a virtual machine host process and one or more virtual machine instances. Computer-executable instructions can be indirectly executed by the physical hardware of the computing device through the interpretation and / or execution of instructions stored and executed in the context of a virtual machine.

[0076] It is important to understand that the methods and systems are not limited to any particular method, component, or implementation. It is also important to understand that the terminology used herein is for describing particular embodiments only and is not intended to be restrictive.

[0077] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural indicators unless explicitly specified otherwise. A range may be expressed herein as from “about” one particular value and / or to “about” another particular value. When such a range is expressed, another embodiment includes from one particular value and / or to another particular value. Similarly, when a value is expressed as an approximation using the antecedent “about,” it is to be understood that the particular value forms another embodiment. It should also be understood that each endpoint of a range is significant with respect to the other endpoint and is independent of the other endpoint.

[0078] "Optional" or "optionally" means that the event or situation described below may or may not occur, and the description includes instances where the event or situation occurs as well as instances where it does not occur.

[0079] In the description and claims of this specification, the word "comprise" and variations thereof (such as "comprising" or "comprises") mean "including but not limited to" and are not intended to exclude, for example, other components, integers, or steps. "Exemplary" means "an example of..." and is not intended to convey indications of preferred or ideal embodiments. "Like" is not used in a limiting sense but for interpretive purposes.

[0080] Components that can be used to perform the described methods and systems are described. When describing combinations, subsets, interactions, groups, etc., of these components, it is to be understood that while specific references to each of the various individual and collective combinations and permutations of these components may not be explicitly described, each is specifically contemplated and described herein for all methods and systems. This applies to all aspects of this application, including but not limited to operations in the described methods. Therefore, if various additional operations exist that can be performed, it is to be understood that each of these additional operations can be performed using any specific embodiment or combination of embodiments of the described methods.

[0081] The method and system can be more readily understood by referring to the following detailed description of preferred embodiments and examples included therein, as well as the accompanying drawings and their descriptions.

[0082] As those skilled in the art will understand, the methods and systems may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the methods and systems may take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) implemented therein. More specifically, the methods and systems may take the form of computer software implemented on the web. Any suitable computer-readable storage medium may be used, including hard disks, CD-ROMs, optical storage devices, or magnetic storage devices.

[0083] Embodiments of methods and systems are described below with reference to block diagrams and flowcharts of methods, systems, apparatuses, and computer program products. It should be understood that each block in the block diagrams and flowcharts, as well as combinations of blocks in the block diagrams and flowcharts, can be implemented, respectively, by computer program instructions. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute on the computer or other programmable data processing apparatus, create components for implementing the functions specified in one or more flowchart blocks.

[0084] These computer program instructions may also be stored in a computer-readable storage medium that can instruct a computer or other programmable data processing apparatus to operate in a particular manner, causing the instructions stored in the computer-readable storage medium to produce an article of writing comprising computer-readable instructions for implementing the functions specified in one or more flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowchart blocks.

[0085] The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Additionally, in some implementations, certain method or process blocks may be omitted. The methods and processes described herein are not limited to any particular sequence, and the associated blocks or states may be executed in other suitable sequences. For example, the described blocks or states may be executed in a different order than specifically described, or multiple blocks or states may be combined in a single block or state. Example blocks or states may be executed serially, in parallel, or in some other manner. Blocks or states may be added to or removed from the described example embodiments. The example systems and components described herein may be configured differently from those described. For example, elements may be added, removed, or rearranged compared to the described example embodiments.

[0086] It should also be understood that various items are illustrated as being stored in memory or on storage devices during use, and these items, or portions thereof, may be transferred between memory and other storage devices for memory management and data integrity purposes. Alternatively, in other embodiments, some or all of the software modules and / or systems may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Furthermore, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways, such as being implemented or provided at least in part as firmware and / or hardware, including, but not limited to, one or more application-specific integrated circuits (“ASICs”), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field-programmable gate arrays (“FPGAs”), complex programmable logic devices (“CPLDs”), etc. Some or all modules, systems, and data structures may also be stored (e.g., as software instructions or structured data) on computer-readable media, such as hard disks, memory, networks, or portable media articles, for retrieval by appropriate devices or via appropriate connections. The system, modules, and data structures can also be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) over various computer-readable transmission media, including wireless and wired / cable-based media, and can take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). In other embodiments, such computer program products can also take other forms. Therefore, the invention can be practiced with other computer system configurations.

[0087] While methods and systems have been described in conjunction with preferred embodiments and specific examples, the scope is not intended to be limited to the specific embodiments stated, as the embodiments herein are intended in all respects to be illustrative rather than restrictive.

[0088] Unless otherwise expressly stated, no method described herein is intended to be construed as requiring its steps to be performed in a specific order. Therefore, unless a method claim actually enumerates the order in which its operations are performed, or unless otherwise specifically stated in the claims or description that these operations will be limited to a specific order, the order is not intended to be inferred in any way. This applies to any possible non-express basis of interpretation, including: logical questions relating to the arrangement of steps or operational flows; the general meaning derived from grammatical organization or punctuation; and the number or type of embodiments described in the specification.

[0089] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the spirit or scope of this disclosure. Other embodiments will be apparent to those skilled in the art upon consideration of the specification and practice described herein. The specification and example drawings are intended to be considered exemplary only, wherein the true scope and spirit are indicated by the following claims.

Claims

1. A method for facial expression recognition, comprising: A first loss function is determined based on a first set of feature vectors associated with a first set of images depicting facial expressions and a first set of labels indicating the facial expressions, wherein the first set of labels includes a first number of mixed shape coefficients; A second loss function is determined based on a second feature vector set associated with a second set of images depicting asymmetrical facial expressions and a second label set indicating the asymmetrical facial expressions, the second label set including a second number of mixed shape coefficients, the second number being less than the first number, and the second set of images corresponding to a subset of the first set of images; Determine the maximum loss function between the first loss function and the second loss function; as well as The maximum loss function is applied during model training, wherein the trained model is configured to predict at least one asymmetric facial expression in subsequently received images.

2. The method of claim 1, wherein the second set of tags includes a mixed shape coefficient indicating the asymmetrical facial expression associated with at least one of the eyes, mouth, chin, eyebrows, cheeks, nose, or tongue.

3. The method according to claim 1, further comprising: The second image set is enhanced by flipping at least a portion of at least one image in the second image set.

4. The method according to claim 3, further comprising: The second tag set is enhanced based on the flipped portion of the at least one image.

5. The method according to claim 1, further comprising: The first feature vector set is determined by inputting the first image set into a neural network and performing regression on different feature vector sets that have a higher dimension than the first feature vector set, the different feature vector sets being associated with the first image set.

6. The method according to claim 1, further comprising: The animation is controlled based on predictions of at least one asymmetrical facial expression in the subsequently received image output from the trained model.

7. A system for facial expression recognition, comprising: At least one computing device communicating with a computer memory, the computer memory including computer-readable instructions that, when executed by the at least one computing device, configure the system to perform operations including: A first loss function is determined based on a first set of feature vectors associated with a first set of images depicting facial expressions and a first set of labels indicating the facial expressions, wherein the first set of labels includes a first number of mixed shape coefficients; A second loss function is determined based on a second feature vector set associated with a second set of images depicting asymmetrical facial expressions and a second label set indicating the asymmetrical facial expressions, the second label set including a second number of mixed shape coefficients, the second number being less than the first number, and the second set of images corresponding to a subset of the first set of images; Determine the maximum loss function among the first loss function and the second loss function; and The maximum loss function is applied during model training, wherein the trained model is configured to predict at least one asymmetric facial expression in subsequently received images.

8. The system of claim 7, wherein the second set of tags includes a mixed shape coefficient indicating the asymmetrical facial expression associated with at least one of the eyes, mouth, chin, eyebrows, cheeks, nose, or tongue.

9. The system according to claim 7, wherein the operation further comprises: The second image set is enhanced by flipping at least a portion of at least one image in the second image set.

10. The system according to claim 9, wherein the operation further comprises: The second tag set is enhanced based on the flipped portion of the at least one image.

11. The system according to claim 7, wherein the operation further comprises: The first feature vector set is determined by inputting the first image set into a neural network and performing regression on different feature vector sets that have a higher dimension than the first feature vector set, the different feature vector sets being associated with the first image set.

12. The system according to claim 7, wherein the operation further comprises: The animation is controlled based on predictions of at least one asymmetrical facial expression in the subsequently received image output from the trained model.

13. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a processor, cause the processor to perform operations, the operations including: A first loss function is determined based on a first set of feature vectors associated with a first set of images depicting facial expressions and a first set of labels indicating the facial expressions, wherein the first set of labels includes a first number of mixed shape coefficients; A second loss function is determined based on a second feature vector set associated with a second set of images depicting asymmetrical facial expressions and a second label set indicating the asymmetrical facial expressions, the second label set including a second number of mixed shape coefficients, the second number being less than the first number, and the second set of images corresponding to a subset of the first set of images; Determine the maximum loss function between the first loss function and the second loss function; as well as The maximum loss function is applied during model training, wherein the trained model is configured to predict at least one asymmetric facial expression in subsequently received images.

14. The non-transient computer-readable storage medium of claim 13, further comprising: The first feature vector set is determined by inputting the first image set into a neural network and performing regression on different feature vector sets that have a higher dimension than the first feature vector set, the different feature vector sets being associated with the first image set.

15. The non-transient computer-readable storage medium of claim 13, further comprising: The animation is controlled based on predictions of at least one asymmetrical facial expression in the subsequently received image output from the trained model.

Citation Information

Patent Citations

  • Expression recognition method and system based on local and overall feature adaptive fusion

    CN112784763A