Image synthesis for personalized facial expression classification
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2022-07-25
- Publication Date
- 2026-07-31
Smart Images

Figure 0007898103000001 
Figure 0007898103000002 
Figure 0007898103000003
Abstract
Description
Technical Field
[0004] , , , , ,
[0003] , ,
[0001] Embodiments of the present disclosure relate to image synthesis for personalized expression classification.
Background Art
[0002] Image analysis may be performed on a face image to identify which expression is being made. Expressions can convey emotion, intent, pain, and may be used in interpersonal behavior. Often, these expressions are characterized based on FACS (Facial Action Coding System) using AUs (Action Units), where each AU may correspond to the relaxation or contraction of a specific muscle or muscle group. Each AU may often be further characterized by categories of intensity labeled 0 and A - E, where 0 indicates no category of intensity or the absence of the AU, and A - E each span a range from trace to maximum intensity. A given emotion may be characterized as a combination of AUs, which may include variations in intensity, such as AU 6B + 12B (each being the raising of the cheek and the pulling up of the lip corner at a light level of intensity).
Summary of the Invention
[0003] One or more embodiments of the present disclosure may include a method that includes obtaining a face image of a subject and identifying the number of new images to be synthesized with a combination of target AUs and categories of intensity. The method may also include synthesizing that number of new images with that number of combinations of target AUs and categories of intensity using the face image of the subject as a base image, where a number of the new images have AU combinations different from the face image of the subject. The method may additionally include adding that number of new images to a dataset and using the dataset to train a machine learning system to identify the expression of the subject. <The objectives and advantages of this embodiment are realized and achieved by at least the elements, features, and combinations specifically pointed out in the claims.
[0005] Please understand that both the general description above and the embodiments for carrying out the invention described below are merely examples and explanatory, and not limiting. [Brief explanation of the drawing]
[0006] Exemplary embodiments are described and illustrated with additional specificity and detail through the use of the accompanying drawings.
[0007] [Figure 1] This figure shows an exemplary environment that can be used for image analysis of facial images.
[0008] [Figure 2A] Examples of face images containing synthesized face images using different synthesis techniques are shown. [Figure 2B] Examples of face images containing synthesized face images using different synthesis techniques are shown.
[0009] [Figure 3] An illustrative flowchart of an exemplary method for personalized facial expression classification is shown.
[0010] [Figure 4] This diagram illustrates how facial expressions are classified using a personalized dataset.
[0011] [Figure 5] This illustrates an exemplary computing system. [Modes for carrying out the invention]
[0012] This disclosure relates to the generation of personalized datasets that can be used to train a machine learning system based on combinations of AUs and / or categories of their intensities in training images. A machine learning system trained on personalized datasets may be used to classify facial expressions in input images. In some potential training datasets, the images used to train the dataset are general across all faces, which may not be accurate. This disclosure provides personalized datasets for training a machine learning system to classify individual facial expressions more accurately. Although the term "images" is used, it will be understood that it is equally applicable to other representations of faces.
[0013] In some embodiments, an individual's input image may be analyzed to determine the combinations and intensity categories of AUs present in the input image, and based on the determination, additional images may be identified to be synthesized to provide a sufficient number of images to train the machine learning system (e.g., images may be identified to provide a diverse range of AU combinations and intensity categories for images in the training dataset). A personalized training dataset may be used to train a machine learning system for image classification, using images that are all based on an individual and therefore sometimes referred to as "personalized." To do so, the machine learning system may first be trained generally to be applicable to any person, and then be tailored or further trained based on images of a specific individual to be personalized.
[0014] After training, the machine learning system may be used to label input images of the same individual with the combinations and / or categories of AUs in the input images. For example, the machine learning system may identify which AUs are present (e.g., binary decision) and / or the categories of intensity of the present AUs (e.g., multiple potential intensity levels). The identified combinations and / or intensity categories of AUs may then be used to classify the facial expression of the subject in the input image. For example, if the input image is identified by the trained machine learning system as having 6+12 combinations of AUs, the expression in the input image may be classified as a smile or as containing a smile.
[0015] Certain embodiments of the present disclosure may offer improvements over previous iterations of machine learning systems for facial image analysis. For example, embodiments of the present disclosure may provide a more personalized dataset for training, allowing the machine learning system to better identify and classify facial expressions in input images to the machine learning system, since the system is trained on individual images rather than being trained on a variety of images of various individuals in general. Additionally, because the present disclosure synthesizes specific images, certain embodiments allow the machine learning system to operate with a training set having fewer initial input images, reducing the cost (both computational and economic) of preparing larger training datasets. Furthermore, because the present disclosure can provide a superior training set to the machine learning system, the machine learning system itself can operate more efficiently and reach decisions more quickly, thus saving computational resources and time spent on longer analyses compared to the present disclosure.
[0016] One or more exemplary embodiments are described with reference to the accompanying drawings.
[0017] Figure 1 shows an exemplary environment 100 that may be used for image analysis on facial images according to at least one embodiment of the present disclosure. As shown in Figure 1, the environment 100 may include a dataset 110 of images that may be used to train a machine learning system 130. After training, the machine learning system 130 may analyze the images 120 and generate a labeled image 140 with labels 145. For example, labels 145 may be applied to the image 120 to generate the labeled image 140.
[0018] The dataset 110 may include one or more labeled images. For example, the dataset 110 may include facial images of individuals that can be labeled to identify which AUs are expressed in the image and / or the category of AU intensity in the image. In some embodiments, one or more of the images in the dataset 110 may be artificially synthesized rather than native images, such as images captured by a camera or other image sensor. In some embodiments, the images in the dataset 110 may be manually labeled or automatically labeled. In these embodiments and other embodiments, the images in the dataset 110 may all be of the same individual so that when the machine learning system 130 is trained using the dataset 110, it is personalized for that individual.
[0019] Image 120 may be any image including a face. Image 120 may be provided as input to the machine learning system 130.
[0020] The machine learning system 130 may include any system, device, network, etc. configured to be trained based on the dataset 110, enabling the machine learning system 130 to identify the categories of AUs and / or their respective intensities in the image 120. In some embodiments, the machine learning system 130 may include a deep learning architecture such as a deep neural network, artificial neural network, convolutional neural network (CNN), etc. The machine learning system 130 may output a label 145 that identifies one or more of the AUs in the image 120 and / or the categories of their respective intensities. For example, the machine learning system 130 may identify which AUs are present (e.g., binary decision) and / or the intensity of the present AUs (e.g., multiple potential intensity levels). Additionally or alternatively, the machine learning system 130 may identify which categories of AUs and / or intensities are absent (e.g., absence of combination 6+12).
[0021] In some embodiments, the machine learning system 130 may generally be trained to perform image labeling across any face image. An example of such training may be described in U.S. Patent Application No. 16 / 994,530 (「IMAGE SYNTHESIS FOR BALANCED DATASETS」), the entire disclosure of which is incorporated herein by reference in its entirety. After being generally trained for any face, the machine learning system 130 may be further trained, adjusted, etc. using an image of a single individual, such that the performance of the machine learning system 130 regarding that person is improved compared to the performance of the generally trained machine learning system.
[0022] The labeled image 140 may represent the image 120 when labeled with a label 145 indicating the categories of AUs and / or their respective intensities determined by the machine learning system 130.
[0023] Without departing from the scope of the present disclosure, modifications, additions, or omissions may be made to the environment 100. For example, the designation of different elements in the described method is intended to assist in the explanation of the concepts described herein and is not limiting. Further, the environment 100 may include any number of other elements and may be implemented with other systems or environments other than those described.
[0024] Figures 2A and 2B show examples of face images 200a and 200b including synthetic face images 230a and 230b using different synthetic techniques according to one or more embodiments of the present disclosure. The synthetic image 230a in Figure 2A is synthesized based on the two-dimensional (2D) registration of the input image 210a, and the synthetic image 230b in Figure 2B is synthesized based on the three-dimensional (3D) registration of the input image 210b.
[0025] The face image 200a in Figure 2A includes an input image 210a, a target image (220a), and a synthetic image of 230a. The input image 210a may be selected as the image that forms the basis of the synthetic image. In some embodiments, the input image 210a may include a face image with few or no wrinkles and / or an expressionless face. The input image 210a may include a face that is facing almost straight ahead.
[0026] In some embodiments, a 2D registration of the input image (210a) may be performed on the input image 210a. For example, the 2D registration may map points in the 2D image to various facial features, landmarks, muscle groups, etc. In some embodiments, the 2D registration may map various facial features, landmarks, muscle groups, etc. of the input image 210a to the target image 220a. The synthetic image 230a may be based on the 2D registration of the input image 210a.
[0027] Please note that there seems to be a small error in the original text where "ターゲット画像220a" was written as "ターゲット画像220aおよび合成画像230aを含む" in ID=8, which might be a duplication. I've corrected it in the translation for better understanding. Also, the "(220a)" in ID=8 is added to make the text more complete as it seems some part might be missing in the original regarding this reference.The target image 220a may represent a desired facial expression (e.g., a face image showing a desired combination of AUs and intensity categories to be synthesized to balance the dataset). The input image 210a may or may not have the same identity as the target image 220a (e.g., representing the same person).
[0028] Referring to Figure 2A, the composite image 230a may have various artifacts based on 2D registration. For example, there may be holes or gaps in the face, certain facial features may be distorted, or otherwise have an unhuman appearance.
[0029] In Figure 2B, the input image 210b and target image 220b may be the same as the input image 210a and target image 220a in Figure 2A. 3D registration of the input image 210b and / or target image 220b may be performed. For example, instead of 2D images, 3D projections of the faces shown in the input image 210b and target image 220b may be generated. By doing so, a more complete, robust, and / or accurate mapping between the input image 210b and target image 220b may be obtained.
[0030] Based on 3D registration, a composite image 230b may be created using the input image 210b as a basis. As observed, the composite image 230b in Figure 2B is of higher quality than the composite image 230a in Figure 2A. For example, it has fewer artifacts and facial features are closer to the target image 220b.
[0031] Facial images 200a / 200b may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the described methods is intended to aid in explaining the concepts described herein, and is not limiting. Furthermore, facial images 200a / 200b may include any number of other elements and may be implemented with other systems or environments other than those described herein. For example, any number of input images, target images, and / or composite images may be used.
[0032] Figure 3 shows an exemplary flowchart of an exemplary method 300 for image synthesis for personalized facial expression classification according to one or more embodiments of the present disclosure. For example, method 300 may be performed to generate a personalized dataset for training a machine learning system to identify facial expressions on input images of a given subject (e.g., by identifying combinations of AUs and categories of their respective intensities). One or more operations of method 300 may be performed by a system or device, or a combination thereof, such as any computing device hosting any component of environment 100 or 400 in Figure 1 and / or Figure 4, for example, a computing device hosting the training dataset 110, a machine learning system 130, etc. Although shown as discrete blocks, the various blocks of method 300 may be split into additional blocks, combined into fewer blocks, or removed, depending on the desired implementation.
[0033] In block 310, the image of the subject may be obtained such that it includes at least the subject's face. The image of the subject may be obtained by any method in which the final result is at least a 2D image of the subject's face. For example, the image of the subject may be obtained by 3D rendering of the subject's face, which may be mapped and then rasterized into a 2D image.
[0034] In block 320, identification may be made regarding the number of new images to be synthesized in order to generate categories of different AU combinations and intensities in order to generate a personalized dataset for the subject. In these and other embodiments, the number of images may be sufficient to train the machine learning system. For example, the number of images synthesized may be a number of images that represent each AU combination and intensity category for the input image. In some embodiments, the number of new images may be a discrete number having a predetermined number of AU combinations and intensity categories. Additionally or alternatively, the number of new images may depend on the input image and the AU combinations and intensity categories already present in the input image. In some embodiments, the number of new images may be determined based on the purpose or application to which the machine learning system is applied. For example, if an end user, application, algorithm, etc., is used to identify an individual only if they are smiling, the number of images synthesized may primarily focus on AU combinations and intensity categories related to smiles.
[0035] In block 330, the number of new images identified in block 320 may be synthesized with the associated AU combinations and intensity categories. In some embodiments, a neutral expression may be used as the base image when synthesizing the new images. Additionally or alternatively, 3D registration of the input images and / or new images (e.g., images showing the AU combinations and intensity categories into which the additional images are synthesized) may be performed to facilitate the synthesis of high-quality images. In some embodiments, one or more loss parameters may be used when synthesizing images to facilitate the generation of high-quality images.
[0036] In block 340, new images synthesized in block 330 may be added to the dataset. For example, synthesized images with input images and their labeled AU combinations and intensity categories may be added to the dataset, and the dataset will include input images and images synthesized in block 330. In an alternative embodiment, the dataset may include images of facial expressions unrelated to the input images, labeled with AU combinations and intensity categories. For example, a number of public face images of a subject may be collected (e.g., from the subject's social media page or from the subject's electronic device), labeled (e.g., automatically or manually) with the AU combinations and intensity categories present in each image of the subject, and grouped into the dataset together with input images from blocks 310, 320, and 330 and that number of synthesized images. Additionally or alternatively, the dataset may include a number of input images related to different users collected and synthesized according to blocks 310, 320, and 330.
[0037] In block 350, a machine learning system may be trained using the dataset generated in block 340. For example, the machine learning system may be trained to identify facial expressions of a subject in an input image of the subject. For example, a CNN may be trained using the dataset to facilitate image labeling using a CNN. After training, the CNN may be provided with unlabeled input images of the subject's face. Using the trained CNN, the input images may be labeled with the identified facial expressions (e.g., by identifying combinations of AUs and / or associated intensity categories). In some embodiments, the machine learning system may already be generally trained on any face, and the training in block 350 may be personalization of the machine learning system. An example of such generally trained machine learning may be described in U.S. Patent Application No. 16 / 994,530 ("IMAGE SYNTHESIS FOR BALANCED DATASETS"), the entire disclosure of which is incorporated herein by reference.
[0038] In block 360, the subject may be identified. In some embodiments, subject identification and / or verification may be performed by fingerprint authentication, password verification, passcode verification, iris scanning, or multi-factor authentication, etc. For example, the subject may use an electronic device with a camera to log in to an application running on a computer system that performs one or more of the actions of method 300 using facial recognition. Additionally or alternatively, the machine learning system may assume the subject's identity based on the device being used. For example, instead of providing a passcode, password, facial recognition, etc., the subject may prefer that the machine learning system identify the subject based on the Internet Protocol (IP) address of the device being used (e.g., the subject's mobile device). Additionally or alternatively, the machine learning system may identify the subject's face through a captured image in which the entire face is substantially visible. In some embodiments, the same image acquired in block 310 may be used to identify the combination of AUs and the associated intensity categories.
[0039] In block 370, the machine learning system may identify categories of AU combinations and intensities in input images of a subject (e.g., a person whose identity was identified / verified in block 360). For example, if the subject was identified in block 360 using face recognition, the machine learning system may use images collected to perform face recognition as input images to identify categories of AU combinations and intensities. Additionally or alternatively, images of the subject may be obtained from the subject's device's camera or video monitoring the subject. For example, after identifying the subject, the device may instruct another camera or imaging device monitoring the subject to collect images of the subject to classify the subject's facial expressions. In some embodiments, the subject identified in block 360 may have one or more captured images of the subject's face. For example, a machine learning system trained in block 350 using a dataset from block 340 made up of at least input images in block 310 and composite images in block 330 may obtain a number of unlabeled images from the subject and identify categories of AU combinations and intensities associated with one or more unlabeled images.
[0040] In block 380, the subject's facial expression may be classified. For example, if the machine learning system identifies AU combination 6+12 (upturned cheeks, upturned corners of the mouth) which represents happiness, the machine learning system may classify the subject's facial expression as a smile. Additionally or alternatively, the machine learning system may classify the user's emotional level based on the AU combination and the associated intensity category. For example, if AU combination 6+12 is used again and the combination is associated with the maximum intensity E, the machine learning system may classify the subject's smile differently than if the AU combination were associated with the minimum intensity A.
[0041] In some embodiments, facial expressions may be classified as emotions (such as happiness) rather than as descriptive facial features (such as a smile). In some embodiments, facial expressions may be used as surrogate input to another process. For example, facial expressions may be classified as "like," "dislike," "somewhat like," or "somewhat dislike" on a social media post. For example, if a user watches a TikTok® stream and smiles, an electronic device may automatically process the smile as a "like" on the stream. Another example is monitoring the facial expressions of a medical patient to facilitate the determination of the patient's pain level or discomfort over time (for example, if the facial expression progresses to a more severe, grim expression, it is increasingly likely that the patient is experiencing discomfort). As an additional example, facial expression classifications may be used to indicate how receptive a subject is to a given advertisement. For example, if an identified classification of a facial image indicates enjoyment, surprise, attention, etc., the classification may indicate that the subject is likely to be receptive to similar advertisements and / or products offered or displayed in the advertisements or similar products.
[0042] Without departing from the scope of this disclosure, the environment 300 may be modified, added to, or omitted. For example, some of the operations of method 300 may be implemented in a different order. Additionally or alternatively, two or more operations may be performed simultaneously. Furthermore, the outlined operations and actions are provided as examples, and some of the operations and actions may be optional, combined with fewer operations and actions, or extended to additional operations and actions without impairing the essence of the disclosed embodiments.
[0043] Figure 4 shows an exemplary environment 400 for personalized facial expression classification according to one or more embodiments of the present disclosure. The environment 400 may include an imaging system 410B from which a subject 410A and at least one input image 420 of subject 410A can be captured. A machine learning system 430, trained on a dataset including at least one input image (as described in block 310) and one or more composite images of subject 410A, identifies categories 440 of AU combinations and intensities present in at least one input image 420. Based on the categories 440 of AU combinations and intensities, facial expressions may be classified, for example, as a frown 450A, a neutral expression 450B, a smile 450C, or other related facial expressions 450D.
[0044] As an example of operation in relation to environment 400, subject 410A may view an advertisement on an imaging system 410B (e.g., a smartphone), or subject 410A may view an advertisement on a video streaming service (e.g., Hulu®). The imaging system 410B may collect at least one image 420 from subject 410A while viewing the advertisement, or the imaging system 410B may collect multiple consecutive images of subject 410A throughout the advertisement. The machine learning system 430 may identify an AU combination and associated intensity category 440 from at least one image 420. If the AU combination and associated intensity category 440 identifies subject 410A as smiling, for example, as shown in a smiling expression 450C, information that subject 410A is open to an advertisement, either from that company or showing a similar subject, may be passed to the advertiser. As another example, if a combination of AUs and associated intensity categories 440 identifies subject 410A as having a frown, for example, as shown in a frowning expression 450A, the advertiser may infer that subject 410A is not open to the subject of the advertisement. The overall inferences drawn from the information gathered from the identified combination of AUs and associated intensity categories 440 may depend, for example, on the subject of the advertisement, subject 410A's average facial expression, the overall clarity of at least one image 420, the subject of a video subject 410A was watching before the advertisement, and changes in subject 410A's facial expression throughout the advertisement.
[0045] The machine learning system 430 may include any system, device, network, etc., configured to be trained on a dataset (such as the dataset identified in block 340), so that the machine learning system can identify categories of AUs and / or their respective intensities in at least one image 420. In some embodiments, the machine learning system 430 may be trained using a personalized dataset that includes at least input images of subject 410A and composite images of subject 410A. For example, one or more images of the subject may be synthesized, as described in Method 300 (such as blocks 310-340), to provide a more robust spectrum of AU combinations and / or intensities in the training set used to train the machine learning system 430. In some embodiments, the training dataset for the machine learning system 430 includes only images of subject 410A.
[0046] The imaging system 410B may be any system configured to capture images, store images, and / or transmit images over a network to a remote server or to store them on another device or location capable of electronic storage. In some embodiments, the imaging system 410B may be the same imaging system from which images are acquired to synthesize images for training a machine learning system 430 (e.g., the imaging system associated with block 310, and / or the same imaging system from which a subject is identified in block 360). For example, a subject 410A may log in to a mobile application that can perform one or more of the operations of method 300 in Figure 3, and in the process of starting the mobile application, the device may identify the subject using a facial recognition process, password authentication, passcode authentication, iris scanning, etc. Additionally or alternatively, the imaging system 410B may capture at least one image 420 of the identified subject 410A. In some embodiments, the imaging system 410B may instruct another device to monitor the subject and / or acquire at least one input image 420 of the subject 410A. In some embodiments, the imaging system 410B may play advertisements from a media platform, and simultaneously, while presenting the advertisements, the imaging system 410B may capture at least one image 420 of an identified subject 410A. In these embodiments and other embodiments, such captured images 420 may be used to classify the facial expressions of the subject 410A, thereby interpreting the subject 410A's response to the advertisements.
[0047] At least one image 420 of subject 410A may include at least the face of subject 410A, with the subject facing the imaging device 410B such that the entire face is substantially visible to the imaging system 410B. In some embodiments, the captured image 420 may be an image acquired in block 310 and / or an image captured in block 360 to identify the subject in method 300 of Figure 3.
[0048] The AU combination and intensity categories 440 may include only AU combinations, and / or categories of AU combinations and their associated intensities. In some embodiments, the AU combination and intensity categories 440 may be identified in their entirety in at least one image 420. For example, the machine learning system 430 may identify the AU combination and intensity categories 440 such that AU combinations 1, 10, 20, and 25 (archetypal AUs that may be classified as “terribly repulsive”) can be determined together with the associated intensities for each of the AUs. Additionally or alternatively, the machine learning system 430 may generate the AU combination and intensity categories 440 such that only the presence or absence of a physical expression or facial feature is identified based on the presence or absence of a given AU in the combination. For example, the AU combination and intensity categories 440 may enumerate only the presence or absence of AU6 and AU12.
[0049] Facial expression classifications 450A-D may be classified according to descriptive physical characteristics. For example, an expression may be classified as a frown 450A, a neutral expression 450B, a smile 450C, or any other expression 450D that can be reasonably inferred from the combination and intensity categories 440 of AUs. In some embodiments, the classification of a physical characteristic (e.g., a smile 450C) may be classified as a "like" on a social media post, and depending on the combination and intensity categories 440 of AUs, a smile 450C may be classified as being receptive to or open to an advertisement. Additionally or alternatively, a frown 450A may be classified as being unreceptive or closed to an advertisement, or as pain, discomfort, depression, or sadness towards a medical patient.
[0050] Environment 400 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the described method is intended to aid in explaining the concepts described herein, and is not limited thereto. Furthermore, Environment 400 may include any number of other elements and may be implemented in conjunction with other systems or environments not described herein.
[0051] Figure 5 shows an exemplary computing system 500 according to at least one embodiment of the present disclosure. The system 500 may include a processor 510, memory 520, a communication unit 530, data storage 530, and / or a user interface unit 540, all of which may be communicatively coupled. Any or all of the environments 100 and 400 in Figures 1 and 4, their components, or computing systems hosting those components may be implemented as computing systems that harmonize with computing system 500.
[0052] Generally, the processor 510 may include any computing entity or processing device, including various computer hardware or software modules, and may be configured to execute instructions stored in any applicable computer-readable storage medium. For example, the processor 510 may include a microprocessor, microcontroller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or any other digital or analog circuitry configured to interpret and / or execute program instructions and / or process data.
[0053] Although shown as a single processor in Figure 5, it will be understood that the processor 510 may include any number of processors distributed across any number of networks or physical locations, configured to individually or collectively perform any number of operations described in this disclosure. In some embodiments, the processor 510 may interpret and / or execute program instructions stored in memory 520, data storage 530, or memory 520 and data storage 530, and / or process data. In some embodiments, the processor 510 may fetch program instructions from data storage 530 and load them into memory 520.
[0054] After the program instructions are loaded into memory 520, the processor 510 may execute program instructions such as the instructions for performing method 300 in Figure 3. For example, the processor 510 may obtain instructions for determining the number of images to be synthesized to personalize the dataset and for synthesizing the images.
[0055] The memory 530 and data storage 530 may include computer-readable storage media that carry or store computer-executable instructions or data structures, or one or more computer-readable storage media. Such computer-readable storage media may be any available media that can be accessed by a general-purpose or dedicated computer, such as a processor 510. In some embodiments, the computing system 500 may or may not include either the memory 520 or the data storage 530.
[0056] For example, but not limited to, such computer-readable storage media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disk read-only memory (CD-ROM), or other optical disk storage, magnetic disk storage, or other magnetic storage, flash memory devices (e.g., solid-state memory devices), or non-temporary computer-readable storage media that can be used to carry or store desired program code in the form of computer-executable instructions or data structures and can be accessed by a general-purpose or special-purpose computer. The above combinations may also be included within the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause the processor 510 to perform a particular operation or group of operations.
[0057] The communication unit 540 may include any components, devices, systems, or combinations thereof configured to transmit or receive information over a network. In some embodiments, the communication unit 540 may communicate with other locations, other devices in the same location, or even other components within the same system. For example, the communication unit 540 may include modems, network cards (wireless or wired), optical communication devices, infrared communication devices, wireless communication devices (such as antennas), and / or chipsets (such as Bluetooth® devices, 802.6 devices (e.g., Metropolitan Area Network (MAN)), WiFi® devices, WiMAX devices, cellular communication equipment, etc.), and / or similar. The communication unit 540 may allow data to be exchanged with the networks described in this disclosure and / or any other devices or systems. For example, the communication unit 540 may enable system 500 to communicate with other systems, such as computing devices and / or other networks.
[0058] A person skilled in the art may, after reviewing this disclosure, recognize that System 500 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, System 500 may include more or fewer components than those expressly illustrated and described.
[0059] The foregoing disclosure is not intended to limit the disclosure to the exact form or specific field of use disclosed. Thus, it is intended that various alternative embodiments and / or modifications to the disclosure, whether expressly described or implied herein, are possible in light of the disclosure. While embodiments of the disclosure are described herein, it will be recognized that modifications may be made in form and detail without departing from the scope of the disclosure. Therefore, the disclosure is limited only by the claims.
[0060] In some embodiments, the different components, modules, engines, and services described herein may be implemented as objects or processes that run on a computing system (for example, as separate threads). While some of the systems and processes described herein are generally described as being implemented in software (stored on and / or run by general-purpose hardware), specific hardware implementations or combinations of software and specific hardware implementations are also possible and intended.
[0061] In this specification, terms used in particular in the appended claims (e.g., the text of the appended claims) are generally intended to be “open” terms (for example, the term “contains” should be interpreted as “contains, but is not limited to,” the term “has” should be interpreted as “has at least,” and the term “includes” should be interpreted as “contains, but is not limited to,” etc.).
[0062] Furthermore, if a particular number of introduced claims are intended to be provided, such intent is explicitly stated in the claim; if there is no such provision, such intent does not exist. For example, to aid understanding, the claims attached below may contain the use of the introductory phrases “at least one” and “one or more” to introduce the provisions of the claims. However, the use of such phrases should not be interpreted as suggesting that the introduction of a provision by the indefinite article “a” or “an” limits any particular claim containing such introduced provision to only one embodiment containing such provision, and this is also true when the same claim contains the introductory phrase “one or more” or “at least one” and an indefinite article such as “a” or “an” (for example, “a” and / or “an” should be interpreted as meaning “at least one” or “one or more”), and the same is true in the case of the use of an indefinite article used to introduce the provisions of a claim.
[0063] Additionally, even if a specific number of provisions in an introduced claim is explicitly specified, a person skilled in the art will recognize that such provisions should be interpreted as meaning at least the specified number (for example, the mere provision of “two” without other modifiers means at least two provisions, or two or more provisions). Furthermore, where a similar convention is used, such a structure is generally intended to include A alone, B alone, C alone, A and B, A and C, B and C, or A, B and C, etc. For example, the use of the term “and / or” is intended to be interpreted in this way.
[0064] Furthermore, any word or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to intend one of those terms, either of those terms, or both of those terms. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B" or "A and B".
[0065] However, the use of such phrases should not be interpreted as suggesting that the introduction of a claim provision by the indefinite article “a” or “an” limits any particular claim containing such introduced provision to only one embodiment containing such provision, and this is also true when the same claim contains the introductory phrase “one or more” or “at least one” and an indefinite article such as “a” or “an” (for example, “a” and / or “an” should be interpreted as meaning “at least one” or “one or more”), and the same is true in the case of the use of an indefinite article used to introduce a claim provision.
[0066] Additionally, the use of terms such as “first,” “second,” and “third” is not necessarily used in this specification to imply a particular order. Generally, terms such as “first,” “second,” and “third” are used to distinguish different elements. Unless it is specifically indicated that terms such as “first,” “second,” and “third” imply a particular order, it should not be understood that these terms imply a particular order.
[0067] All examples and conditional language set forth herein are intended for educational purposes to assist the reader in understanding the present invention and the concepts to which the inventors have contributed to advancing the art, and should be construed as not being limited to such specifically defined examples and conditions. While embodiments of this disclosure are described in detail, it should be understood that various modifications, substitutions, and exchanges may be made thereto without departing from the spirit and scope of this disclosure.
[0068] The preceding description of the disclosed embodiments is provided to enable those skilled in the art to manufacture or use the disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the disclosure. Accordingly, the disclosure is not intended to be limited to the embodiments shown herein, but should be granted to the broadest extent that is in harmony with the principles and novel features disclosed herein.
[0069] This disclosure includes the following inventions. (Note 1) It is a method, Obtaining a facial image of the subject, Identifying the number of new images synthesized with target AU combinations and intensity categories, The synthesis involves using the subject's face image as the base image in the synthesis to synthesize the aforementioned number of new images, wherein two or more of the aforementioned number of new images have a different combination of AUs than the subject's face image. Adding the aforementioned number of new images to the dataset so that the dataset contains only images of the subject, A method comprising training a machine learning system using the aforementioned dataset, wherein the machine learning system is trained to identify the facial expressions of the subject. (Note 2) The method according to Appendix 1, wherein the dataset includes only the facial images of the subject and the number of new images synthesized from the facial images of the subject. (Note 3) The method described in Appendix 1, wherein the facial image of the subject includes an expressionless face. (Note 4) The method according to Appendix 1, wherein capturing the facial image of the subject includes capturing two or more images of the subject. (Note 5) The method according to Appendix 1, wherein identifying the number of novel images synthesized with the number of target AU combinations and intensity categories includes verifying that at least one image represents each intensity category for each AU. (Note 6) Identifying at least one combination of AUs and at least one intensity category in the facial image of the subject, The method according to Appendix 1, further comprising determining a set of target AU combinations and intensity categories based on the facial image of the subject. (Note 7) Identifying the subject, The method according to Appendix 1, further comprising classifying the facial expressions of the subject by identifying the categories of each AU combination and intensity using the machine learning system. (Note 8) The method according to Appendix 7, wherein the identification of the subject is performed via an identification technique including at least one of facial recognition, password verification, passcode verification, fingerprint verification, iris scanning, or multi-factor authentication. (Note 9) One or more non-temporary computer-readable media configured to store one or more instructions, wherein the one or more instructions cause the system to perform an action in response to being executed by one or more programs, and the action is Obtaining a facial image of the subject, Identifying the number of new images synthesized with target AU combinations and intensity categories, The process involves synthesizing a certain number of new images using the subject's face image as the base image, wherein two or more of the new images have a different combination of AUs than the subject's face image. Add the aforementioned number of new images to the dataset, Training a machine learning system using the aforementioned dataset, wherein the machine learning system is trained to identify the facial expressions of the subject, and training one or more computer-readable media. (Note 10) The dataset is one or more computer-readable media as described in Appendix 9, which includes only one or more facial images of the subject and the aforementioned number of new images synthesized from the facial images of the subject. (Note 11) Identifying the aforementioned number of novel images synthesized with the aforementioned number of target AU combinations and intensity categories includes verifying that at least one image represents each intensity category for each AU, in one or more computer-readable media as described in Appendix 9. (Note 12) The aforementioned operation is, Identifying at least one combination of AUs and at least one intensity category in the facial image of the subject, One or more computer-readable media as described in Appendix 9, further comprising determining a set of target AU combinations and intensity categories based on the facial image of the subject. (Note 13) The aforementioned operation is, Identifying the subject, One or more computer-readable media as described in Appendix 9, further comprising classifying the facial expressions of the subject by identifying the categories of each AU combination and intensity using the machine learning system. (Note 14) Identifying the subject is performed via identification technology including at least one of facial recognition, password verification, passcode verification, fingerprint verification, iris scanning, or multi-factor authentication, as described in one or more computer-readable media as described in Appendix 13. (Note 15) It is a system, One or more processors, A system comprising one or more non-temporary computer-readable media configured to store one or more instructions, wherein the one or more instructions cause the system to perform an action in response to being executed by one or more programs, and the action is Obtaining a facial image of the subject, Identifying the number of new images synthesized with target AU combinations and intensity categories, The process involves synthesizing a number of new images using the subject's face image as the base image, with the aforementioned number of target AU combinations and intensity categories, wherein two or more of the aforementioned new images have different AU combinations than the subject's face image. Add the aforementioned number of new images to the dataset, A system comprising training a machine learning system using the aforementioned dataset, wherein the machine learning system is trained to identify the facial expressions of the subject. (Note 16) The system as described in Appendix 15, wherein the dataset includes only one or more facial images of the subject and the aforementioned number of new images synthesized from the facial images of the subject. (Note 17) Identifying the aforementioned number of novel images synthesized with the aforementioned number of target AU combinations and intensity categories includes verifying that at least one image represents each intensity category for each AU, as described in Appendix 15. (Note 18) The aforementioned operation is, Identifying at least one combination of AUs and at least one intensity category in the facial image of the subject, The system according to Appendix 15, further comprising determining a set of target AU combinations and intensity categories based on the facial image of the subject. (Note 19) The aforementioned operation is, Identifying the subject, The system according to Appendix 15, further comprising classifying the facial expressions of the subject by identifying the categories of each AU combination and intensity using the machine learning system. (Note 20) The system as described in Appendix 15, wherein the identification of the subject is performed via identification technology including at least one of facial recognition, password verification, passcode verification, fingerprint verification, iris scanning, or multi-factor authentication. [Explanation of Symbols]
[0070] 145 labels 500 Systems 510 Processor 520 memory 530 Data Storage 540 Communication Unit
Claims
1. It is a method, Obtaining a facial image of the subject, Identifying the number of new images synthesized with target AU combinations and intensity categories, The synthesis involves using the subject's face image as the base image in the synthesis, and synthesizing the aforementioned number of new images, wherein two or more of the aforementioned number of new images have a different AU combination than the subject's face image. Adding the aforementioned number of new images to the dataset so that the dataset contains only images of the subject, A method comprising training a machine learning system using the aforementioned dataset, wherein the machine learning system is trained to identify the facial expressions of the subject.
2. The method according to claim 1, wherein the dataset includes only the face image of the subject and the number of new images synthesized from the face image of the subject.
3. The method according to claim 1, wherein the facial image of the subject includes a face with no expression.
4. The method according to claim 1, wherein capturing the facial image of the subject includes capturing two or more images of the subject.
5. The method according to claim 1, wherein identifying the number of new images synthesized with the aforementioned number of target AU combinations and intensity categories includes verifying that at least one image represents each intensity category for each AU.
6. Identifying at least one combination of AUs and at least one intensity category in the facial image of the subject, The method according to claim 1, further comprising determining a set of target AU combinations and intensity categories based on the facial image of the subject.
7. Identifying the subject, The method according to claim 1, further comprising classifying the facial expressions of the subject by using the machine learning system to identify the categories of each AU combination and intensity.
8. One or more non-temporary computer-readable media configured to store one or more instructions, wherein the one or more instructions cause the system to perform an action in response to being executed by one or more programs, and the action is Obtaining a facial image of the subject, Identifying the number of new images synthesized with target AU combinations and intensity categories, The process involves synthesizing a certain number of new images using the subject's face image as the base image, wherein two or more of the new images have a different AU combination than the subject's face image. Add the aforementioned number of new images to the dataset, Training a machine learning system using the aforementioned dataset, wherein the machine learning system is trained to identify the facial expressions of the subject, in one or more non-temporary computer-readable media.
9. It is a system, One or more processors, A system comprising one or more non-temporary computer-readable media configured to store one or more instructions, wherein the one or more instructions cause the system to perform an action in response to being executed by one or more programs, and the action is Obtaining a facial image of the subject, Identifying the number of new images synthesized with target AU combinations and intensity categories, The process involves synthesizing a number of new images using the subject's face image as a base image, with the aforementioned number of target AU combinations and intensity categories, wherein two or more of the aforementioned new images have different AU combinations than the subject's face image. Add the aforementioned number of new images to the dataset, A system comprising training a machine learning system using the aforementioned dataset, wherein the machine learning system is trained to identify the facial expressions of the subject.