Facial Emotion Recognition System

JP2025511116A5Pending Publication Date: 2026-04-07HUMINTELL LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Conventional facial emotion recognition systems fail to accurately differentiate between real and fake facial expressions, as they rely on static prototypes and frame-by-frame analysis, neglecting the dynamics and timing characteristics of facial muscle movements.

Method used

The system analyzes video data to identify appearance changes and timing characteristics of facial muscle movements, determining whether the facial behavior is real or fake, and classifies emotions based on these dynamics, using machine learning models to process and interpret the data.

Benefits of technology

This approach enables more accurate classification of emotional facial expressions by considering the temporal dynamics and symmetry of facial movements, effectively distinguishing between genuine and simulated emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The computing system identifies video data capturing an expresser exhibiting facial behavior. The computing system analyzes the video data to determine a type of emotion exhibited by the expresser in the video data by identifying appearance changes produced by facial muscle movements in the video data and determining timing characteristics of the facial muscle movements in the video data, the timing characteristics indicating whether the facial behavior exhibited by the expresser is a genuine or fake facial expression. The computing system generates a classification of the type of emotion exhibited by the expresser based on the facial muscle movements and the timing characteristics of the movements. The computing system outputs the classification.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD OF THE DISCLOSURE The embodiments disclosed herein relate generally to facial emotion recognition systems. [Background technology]

[0002] Facial emotion recognition is a method of measuring facial emotions by identifying facial expressions of emotion produced by muscle contractions, which can be detected by comparing them with known facial muscle movements to categorize or classify the emotion exhibited by an individual. Summary of the Invention

[0003] In some embodiments, a method is disclosed herein: A computing system identifies video data capturing an expresser exhibiting a facial behavior; The computing system analyzes the video data to determine a type of emotion exhibited by the expresser in the video data by the steps of identifying appearance changes produced by facial muscle movements in the video data and determining timing characteristics of the facial muscle movements in the video data, the timing characteristics indicating whether the facial behavior exhibited by the expresser is a genuine or fake facial expression; The computing system generates a classification of the type of emotion exhibited by the expresser based on the facial muscle movements and the timing characteristics of the movements, and the computing system outputs the classification.

[0004] In some embodiments, a non-transitory computer readable medium is disclosed herein. The non-transitory computer readable medium comprises one or more sets of instructions that, when executed by a processor, cause a computing system to perform operations. The operations include: the computing system identifies video data capturing an expresser exhibiting a facial behavior. The operations further include: the computing system analyzes the video data to determine a type of emotion exhibited by the expresser in the video data by identifying appearance changes produced by facial muscle movements in the video data and determining timing characteristics of the facial muscle movements in the video data, the timing characteristics indicating whether the facial behavior exhibited by the expresser is a genuine or fake facial expression. The operations further include: the computing system generates a classification of a type of emotion exhibited by the expresser based on the facial muscle movements and the timing characteristics of the movements. The operations further include: the computing system outputs the classification.

[0005] In some embodiments, a system is disclosed herein. The system comprises a processor and a memory. The memory has programming instructions stored thereon, the instructions, when executed by the processor, cause the system to perform an operation. The operations comprise identifying video data capturing an expresser exhibiting a facial behavior. The operations further comprise analyzing the video data to determine a type of emotion exhibited by the expresser in the video data by identifying appearance changes produced by facial muscle movements in the video data and determining timing characteristics of the facial muscle movements in the video data, the timing characteristics indicating whether the facial behavior exhibited by the expresser is a genuine or fake facial expression. The operations further comprise generating a classification of a type of emotion exhibited by the expresser based on the facial muscle movements and the timing characteristics of the movements. The operations further comprise outputting the classification.

[0006] So that the above-mentioned features of the present disclosure can be understood in detail, a more particular description of the present disclosure briefly summarized above can be had by reference to embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical embodiments of the present disclosure and therefore should not be considered as limiting the scope of the present disclosure, since the present disclosure may admit of other equally effective embodiments. [Brief description of the drawings]

[0007] [Figure 1] FIG. 1 is a block diagram illustrating a computing environment according to one exemplary embodiment. [Figure 2A] 1 is a block diagram illustrating a facial analysis system according to one exemplary embodiment. [Figure 2B] 1 is a block diagram illustrating a facial analysis system according to one exemplary embodiment. [Diagram 3] 1 is a flow diagram illustrating a method for classifying a facial configuration of an expresser according to an exemplary embodiment. [Figure 4] 6 is a chart illustrating exemplary intensity levels for action units, according to an exemplary embodiment. [Figure 5A] FIG. 1 illustrates a system bus computing system architecture according to an exemplary embodiment. [Figure 5B] FIG. 1 illustrates a computer system having a chipset architecture according to an exemplary embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0008] For ease of understanding, the same reference numbers have been used, where possible, to designate identical elements common to the figures. It is contemplated that elements disclosed in one embodiment may be beneficially utilized in other embodiments without specific recitation.

[0009] Conventional approaches to assessing facial expressions of emotion typically involve employing a list of facial expressions and appearance changes due to estimated facial muscle movements that have been empirically shown to be associated with emotion. Such conventional approaches typically rely on a prototypical configuration of the whole face of facial behavior displayed in a still image (e.g., a photograph). In particular, conventional approaches describe what a prototypical facial configuration of emotion looks like, but do not describe the wide variety of emotional expressions that a face can produce, and do not consider the dynamics of such expressions related to timing and symmetry.

[0010] Most conventional approaches to facial expression recognition technology typically involve analyzing videos on a frame-by-frame basis, identifying appearance changes resulting from movements of putative facial muscles innervated by the facial muscles in each frame, and classifying the putative facial behaviors as emotional expressions based on whether those appearance changes match known facial muscle appearance changes in published prototypes. Such approaches typically fail to take into account the behavioral dynamics of muscle movements and the many different kinds of deformations or permutations of those holistic facial configurations that signal emotions.

[0011] In addition, because many conventional approaches classify estimated emotions on a frame-by-frame basis, many conventional approaches then generate data on estimated emotional expressions on a frame-by-frame basis, typically expressed as ratios or probabilities of emotional expressions.

[0012] One or more techniques disclosed herein provide an improvement over conventional systems by supporting many different types of facial muscle configurations that signal emotions and go far beyond the prototype facial expressions of the whole face utilized by conventional systems.For example, one or more techniques described herein provide a fundamentally different way of evaluating facial behaviors to classify behaviors as emotional expressions compared to conventional approaches, for example, by utilizing the known facial dynamics of true emotional expressions and identifying the nature of the behavioral dynamics of the expressions.In this way, the present technique can classify an increased number of variations of each emotion compared to conventional approaches.

[0013] The term "user" as used herein includes, for example, a person or entity that owns a computing or wireless device, a person or entity that operates or utilizes a computing or wireless device, or a person or entity that is associated with a computing or wireless device. It is contemplated that the term "user" is not intended to be limiting and may include a variety of examples other than those described.

[0014] The term "expressor" as used herein may refer to the person generating the facial behavior being analyzed. In some embodiments, the "user" may be the "expressor." In some embodiments, the "user" may be different from the "expressor."

[0015] As used herein, the term "action unit" may refer to basic muscle movements of facial muscles based on the facial action coding system. As used herein, the term "behavior" may refer to unclassified facial movements of a subject.

[0016] As used herein, the term "expression" may refer to the facial movements of a classified expressor. The term "excursion" or "excursions" as used herein may refer to the movement of action units in any facial behavior. In some embodiments, an excursion may be defined as a movement (e.g., innervation) of a facial muscle or combination thereof from a starting point to another point at a greater intensity and then back to a baseline or resting state of lower intensity. These excursions may move the skin on the expresser's face and cause a change in appearance. In some embodiments, the intensity change (e.g., maximum contraction of a facial muscle) from the starting position to the apex may be such that it produces an observable result (e.g., one or more intensity level differences) between the starting position and the apex. In some embodiments, the facial expression may start from a neutral intensity (e.g., the muscle was not contracted at all to begin with) or from some degree of pre-existing contraction. In those embodiments where the facial expression starts with some degree of pre-existing contraction, the new facial expression may be superimposed on top of the one that was previously present. In some embodiments, the lower intensity state to which the excursion may return may be neutral (no evidence of innervation) or non-neutral. The lower intensity condition may be sufficient to produce an observable decrease (eg, one or more intensity level differences).

[0017] 1 is a block diagram illustrating a computing environment 100 according to one embodiment. The computing environment 100 may include a client device 102 and a back-end computing system 104 that communicate over a network 105.

[0018] Network 105 may be of any suitable type, including individual connections over the Internet, such as a cellular network or a Wi-Fi network. In some embodiments, network 105 may connect terminals, services, and mobile devices using radio frequency identification (RFID), near field communication (NFC), Bluetooth, low energy Bluetooth (BLE), Wi-Fi, ZigBee, ambient backscatter communication (ABC) protocol, USB, direct connections, such as WAN, or LAN. Because the information transmitted may be personal or confidential, security concerns may dictate that one or more of these types of connections be encrypted or otherwise secured. However, in some embodiments, the information transmitted may be less personal, and thus a network connection may be selected for convenience over security.

[0019] Network 105 may include any type of computer networking configuration used to exchange data. For example, network 105 may be the Internet, a private data network, a virtual private network using a public network, and / or other suitable connection that allows components in computing environment 100 to send and receive information between components of computing environment 100.

[0020] The client device 102 may be operated by a user. The client device 102 may represent a mobile device, a tablet, a desktop computer, or any computing system having the capabilities described herein. The client device 102 may include at least an application 112.

[0021] The application 112 may represent an application associated with the backend computing system 104. In some embodiments, the application 112 may be a stand-alone application associated with the backend computing system 104. In some embodiments, the application 112 may represent a web browser configured to communicate with the backend computing system 104. In some embodiments, the client device 102 may communicate over the network 105 to access functionality of the backend computing system 104 via a web client application server 114 of the backend computing system 104. Content to be displayed on the client device 102 may be transmitted from the web client application server 114 to the client device 102 and subsequently processed by the application 112 for display through a graphical user interface (GUI) of the client device 102.

[0022] As shown, the client device 102 may be associated with a camera 108. In some embodiments, the camera 108 may be integrated within the client device 102 (e.g., a front camera, a rear camera, etc.). In such embodiments, the user may grant the application 112 access to the camera 108 so that the client device 102 can upload image or video data of the exhibitor to the backend computing system 104.

[0023] In some embodiments, the camera 108 may be a separate component from the client device 102. For example, the camera 108 may communicate with the client device 102 via one or more wired (e.g., USB) or wireless (e.g., WiFi, Bluetooth, etc.) connections. In some embodiments, the camera 108 may transmit or provide captured image or video data to the client device 102 for uploading to the backend computing system 104 via the application 112. In some embodiments, the camera 108 may transmit or provide captured image or video data directly to the backend computing system 104 over the network 105.

[0024] More generally, a user may utilize the camera 108 to capture image data and / or video data of a subject's facial behavior and expressions. The back-end computing system 104 may analyze the image data and / or video data to evaluate facial behavior. In some embodiments, the back-end computing system 104 may be configured to analyze the image data and / or video data in real-time (or near real-time). In some embodiments, the back-end computing system 104 may be configured to analyze pre-stored image data and / or video data.

[0025] The back-end computing system 104 may communicate with the client device 102 and / or the camera 108. The back-end computing system 104 may include a web client application server 114 and a facial analysis system 120. The facial analysis system 120 may be comprised of one or more software modules. The one or more software modules are a collection of codes or instructions stored on a medium (e.g., memory of the back-end computing system 104) that represent a series of machine instructions (e.g., program code) that implement one or more algorithmic steps. Such machine instructions may be actual computer code that a processor of the back-end computing system 104 interprets to implement the instructions, or alternatively, may be higher-level coding of instructions that are interpreted to obtain the actual computer code. The one or more software modules may also include one or more hardware components. One or more aspects of the exemplary algorithm may be executed by the hardware components (e.g., circuits) themselves, rather than as a result of instructions.

[0026] Facial analysis system 120 may be configured to analyze the uploaded image and / or video data to determine the emotion or range of emotions exhibited by the expresser. For example, facial analysis system 120 may utilize one or more machine learning techniques to determine the emotion or range of emotions exhibited by the expresser.

[0027] Emotions may refer to responses to stimuli. Stimuli may be real, physical, and external to the individual, or imagined within the individual (e.g., a thought or memory). Emotional responses typically have physiological, psychological, and social components. Emotional facial expressions are observable behaviors that may signal the presence of an emotional state.

[0028] To determine the emotion of the expresser, in some embodiments, the facial analysis system 120 may analyze the expresser's facial behavior based on the image data and / or video data. Facial behavior may refer to any movement of facial muscles. Many facial behaviors cause appearance changes by moving the skin and generating wrinkled movement patterns. However, some facial behaviors may not produce observable appearance changes.

[0029] 2A is a block diagram illustrating an example embodiment of a backend computing system 104. As shown, the backend computing system 104 includes a repository 202 and one or more computer processors 204.

[0030] Repository 202 may represent any type of storage unit and / or device for storing data (eg, a file system, a database, a collection of tables, or any other storage mechanism). Additionally, the repository 202 may include multiple different storage units and / or devices that may or may not be of the same type, or that may or may not be at the same physical site. As shown, the repository 202 includes at least a facial analysis system 120.

[0031] As shown, the facial analysis system 120 may include a pre-processing module 206, a database 208, a training module 212, and a machine learning model 214. Each of the pre-processing module 206 and the training module 212 may be composed of one or more software modules. The one or more software modules are a collection of codes or instructions stored on a medium (e.g., memory of the back-end computing system 104) that represent a series of machine instructions (e.g., program code) that implement one or more algorithmic steps. Such machine instructions may be actual computer code that a processor of the back-end computing system 104 interprets to implement the instructions, or alternatively, may be higher-level coding of instructions that are interpreted to obtain the actual computer code. The one or more software modules may also include one or more hardware components. One or more aspects of the exemplary algorithm may be executed by the hardware components (e.g., circuits) themselves, rather than as a result of instructions.

[0032] The pre-processing module 206 may be configured to generate a training data set from data stored in the database 208. In general, the database 208 may include videos and / or images of high quality depictions of facial expressions of known emotions portrayed by individuals from multiple racial / ethnic groups and genders. In some embodiments, the database 208 may also include images and / or videos of partial facial images.

[0033] In some embodiments, the training data may include still photographs or videos of known facial expressions of the seven emotions (e.g., anger, contempt, disgust, fear, happiness, sadness, and surprise). Such photographs or videos may have been previously coded using a Facial Action Coding System (FACS). The FACS codes are then used to select the best examples of images for the characterization of the facial expressions of the known published emotions, based on decades of research. In some embodiments, these still photographs may be combined with neutral images of the same expressers and then rendered to generate a moving face in a video, where the target expression of the emotion moves from the onset, through the apex, and then to the offset. Thus, these videos contain the most hygienic moving versions of the facial expressions of the seven emotions.

[0034] The training module 212 may be configured to train the machine learning model 214 to classify the emotional facial expression of the expresser based on the training dataset. In some embodiments, the training may include the training module 212 training the machine learning model 214 to identify action units and excursions of the action units and to learn timing characteristics of the action units and excursions that may be indicative of the emotional facial expression.

[0035] The training module 212 may be configured to train the machine learning model 214 to identify action units and excursions of action units. In general, facial expressions of emotions are facial behaviors with one or more features. For example, the expressions may involve the movement of specific configurations of targeted facial muscles. In some embodiments, the configurations may occur in response to a stimulus. The muscle movements may meet certain criteria regarding their behavioral dynamics. The facial analysis system 120 may be trained to determine an emotion or range of emotions of an expresser by analyzing the expresser's basic muscle movements.

[0036] The training module 210 may be configured to train the machine learning model 214 to identify and monitor action units of an expressor. In some embodiments, the training module 210 may train the machine learning model 214 to monitor all action units of an expressor. For example, the facial analysis system 120 may analyze all action units occurring simultaneously at each vertex. A combination of action units at the same vertex may be classified as a configuration by the facial analysis system 120.

[0037] The training module 212 may further be configured to train the machine learning model 214 to classify the expresser's emotional facial expressions based on the learned features that may be indicative of an emotional facial class. For example, using the training dataset, the training module 212 may train the machine learning model 214 to match the detected configurations with a set of known or deterministic action units. Based on the set of deterministic action units, the machine learning model 214 may identify possible emotional facial expressions exhibited by the expresser. Exemplary known emotional facial expressions may include, but are not limited to, anger, contempt, disgust, fear, joy, smile, etc.

[0038] In some embodiments, the expresser's possible facial expressions may be classified as "possible facial expressions of emotion." Based on the possible emotional facial expressions, the training module 212 may train the machine learning model 214 to compare the inconclusive action units in the facial configuration to an allowed list and a disallowed list of inconclusive action units to further refine the classification of the possible emotional facial expressions. The inconclusive action units in the "allowed" list may represent action units that do not interfere with the classification of the configuration as a possible emotional facial expression. The inconclusive action units in the disallowed list may qualify the classification of the possible emotional facial expressions. In some embodiments, the machine learning model 214 may be trained to retain the inconclusive action units in its analysis or output. For example, the machine learning model 214 may mark the inconclusive action units with a question mark.

[0039] In some embodiments, the training process may include a machine learning model 214 that learns the timing characteristics of movements that may indicate certain expressions. As mentioned above, a limitation of conventional facial analysis programs is that they rely on frame-by-frame analysis of topographical changes in the face, for example, as detected by pixel changes, coloration, spectroscopy, or other such visual cues in images. Such conventional approaches ignore the important temporal dynamics associated with true and spontaneous emotional facial expressions. One of the largest and more important factors in deciphering between true and fake emotional facial expressions is their timing characteristics. Humans can simulate most of the emotional facial expressions that research has demonstrated to occur across cultures. For example, when an expresser simulates an emotional facial expression, the simulation may have the same facial movements (or facial actions) as the true and spontaneous facial expressions, but the simulated facial expressions differ from the true facial expressions in their timing characteristics. Because the current state of the art utilizes a frame-by-frame analysis of the face and calculates the probability of a match with known facial expressions on each image, such approaches cannot take into account the temporal dynamics of facial expressions that occur across many images.

[0040] Machine learning and artificial intelligence offer the unique ability to address this limitation in the current state of the art by enabling the processing of frame-by-frame data and, through various learned timing features, detecting the temporal dynamics of facial expressions exhibited by the expresser.

[0041] The training module 212 may be configured to train the machine learning model 214 to learn timing characteristics that may represent true spontaneous facial expressions. Exemplary characteristics may include, but are not limited to, simultaneous onset of multiple muscle movements in the same configuration, smooth onset acceleration of multiple muscle movements in the same configuration, simultaneous apex of multiple muscle movements in the same configuration, shared apex duration of multiple muscle movements in the same configuration, symmetry of multiple muscle movements in the same configuration with limited exceptions, and smooth end declaration of multiple muscle movements in the same configuration.

[0042] When facial expressions involve multiple muscles, synchrony can occur. Generally, the excursions of multiple muscles start relatively simultaneously. Smooth onset acceleration may refer to the relative smoothness of the onset of a muscle, i.e., the smooth acceleration of the muscle movement to the apex. Generally, the acceleration to the apex is smooth, i.e., equal or comparable rates of change divided by time, so as to be similar for multiple muscle movements.

[0043] Simultaneous peaks may refer to muscles involved in facial expressions that peak at approximately the same time. Shared peak duration may refer to the duration of a shared peak. For example, when an emotion is responding to a single elicitor, multiple muscles of the same configuration have equal or comparable peak durations and do not last more than a certain number of seconds (e.g., 4 or 5 seconds). Exceptions to this rule may be made on an individual basis.

[0044] Symmetry may refer to the shape of the facial expression. In general, a person's facial expressions are symmetric for many expressions, i.e., the muscles on the right and left sides of the face share the same timing characteristics as defined herein, with some limited exceptions.

[0045] A smooth end declaration may refer to the end of a muscle, i.e., a drop to the baseline. Generally, the end of a muscle is smooth, i.e., an equal or comparable rate of change divided by time, as is the case for multiple muscle movements. However, as noted above, the expresser's facial expression does not necessarily return to the baseline. For an event that includes a combination of action units, the end of the event may be determined based on the end of the first target action unit from the event.

[0046] Thus, the training module 212 may train the machine learning model 214 to detect and / or analyze timing characteristics of the expresser's action unit movements when determining the expresser's emotional facial expression.

[0047] Once training is complete, the trained machine learning model ("trained model 216") may be deployed to a computing environment. In some embodiments, the trained model 216 may be deployed locally. In some embodiments, the trained model 216 may be deployed to a cloud-based environment.

[0048] In some embodiments, the machine learning model 214 may represent one or more of machine learning models or algorithms, which may include, but are not limited to, a random forest model, a support vector machine, a neural network, a deep learning model, a Bayesian algorithm, a convolutional neural network, etc.

[0049] 2B is a block diagram illustrating the backend computing system 104, according to an example embodiment. FIG 2B may represent a deployment computing environment for the trained model 216 after training.

[0050] As shown, the capture module 250 may receive video data of the expresser. In some embodiments, the capture module 250 may be configured to pre-process the video data. For example, the capture module 250 may be configured to upsample or downsample the video data before inputting it to the trained model 216.

[0051] The trained model 216 may receive as input video data of the expressor. Based on the video data of the expressor, the trained model 216 may generate an output 252. In general, the output 252 may represent a facial expression of the expressor's emotion.

[0052] The output 252 may be represented in a variety of ways. In some embodiments, the output 252 may be represented as a JSON file that includes the probability of occurrence of various spontaneous emotional facial expressions. In some embodiments, such a JSON file may be stored in the database 208 and used in a feedback loop to train or retrain the machine learning model 214.

[0053] In some embodiments, the output 252 may be provided regarding the particular emotion category detected. For example, because the trained model 216 is trained to consider timing characteristics, and facial expression and intensity of emotional responses are highly correlated, the temporal data may be used to represent the intensity of the detected emotion. Such an output is an improvement over the current state of the art, which is limited to simply detecting the probability of the presence or absence of an emotion based on a still image and the intensity or strength of its emotional response.

[0054] 3 is a flow diagram illustrating a method 300 for classifying an expresser's facial configuration, according to an example embodiment. The method 300 may begin at step 302. In step 302, the facial analysis system 120 may receive image data and / or video data of a person. In some embodiments, the image data and / or video data may be uploaded by a user from a client device 102. In some embodiments, the image data and / or video data may be uploaded by a camera 108.

[0055] In some embodiments, method 300 may include step 304. In step 304, facial analysis system 120 may generate a video based on a still image of an expressor. For example, facial analysis system 120 may receive a neutral image of the same expressor and then render the neutral image along with the expression image to generate a moving face in the video.

[0056] In step 306, the facial analysis system 120 may analyze the video data of the speaker to determine the facial expressions of emotions exhibited by the speaker. For example, the trained model 216 may receive the video data as input.

[0057] In some embodiments, determining the expresser's facial expression of emotion may include the trained model 216 identifying and tracking any facial muscle excursions in the video data.

[0058] Facial analysis system 120 may identify and / or classify facial configurations based on one or more parameters. For example, facial analysis system 120 may identify all action units that occur simultaneously during a vertex period and classify these action units as configurations.

[0059] The facial analysis system 120 may identify and / or classify possible emotional facial expressions based on the facial configuration. For example, the facial analysis system 120 may divide the action units in the configuration into "significant" and "non-significant" action units. The facial analysis system 120 may compare the "significant" action units in the facial configuration to a list of significant action units associated with emotional facial expressions to identify matches with known emotional facial expressions. The facial analysis system 120 may compare the non-significant action units in the facial configuration to "allowed" and "not allowed" lists of non-significant action units to further refine the classification of the possible emotional facial expressions. In some embodiments, the non-significant action units in the "allowed" list may not interfere with the classification of the configuration as a possible emotional facial expression. In some embodiments, the non-significant action units in the "not allowed" list may qualify for classification of the possible emotional facial expression. In some embodiments, such classification may be marked with a question mark.

[0060] In some embodiments, when identifying and / or classifying possible emotional facial expressions based on facial configurations, facial analysis system 120 may ignore some action units if they occur simultaneously (or temporally) with speech. However, if the timing or intensity of these action units is distinctly different from movements that occur during speech, these action units may be counted in the configuration.

[0061] The facial analysis system 120 may classify significant action units in a facial configuration that match known significant action units that match emotional facial expressions and that include "acceptable" non-significant action units as "possible emotional facial expressions." Significant action units in a facial configuration that match known significant action units that match emotional facial expressions but that include "not allowed" non-significant action units may be classified as possible emotional facial expressions but may be marked with a question mark. In some embodiments, the facial analysis system 120 may generate an initial classification of emotional facial expressions based on the possible emotional facial expressions.

[0062] Facial analysis system 120 may further analyze possible emotional facial expressions according to whether they meet behavioral dynamics criteria. For example, facial analysis system 120 may determine simultaneous onset based on generated start times. Facial analysis system 120 may compare start times of all significant units in a composition. Based on this comparison, facial analysis system 120 may make a decision about simultaneous onset. For example, if the starts (e.g., onsets) of significant action units in a composition are within a threshold tolerance (e.g., 15-30 ms) of each other, facial analysis system 120 may classify such movements as simultaneous.

[0063] To determine the emotional facial expression exhibited by the expresser, the facial analysis system 120 may identify timing characteristics of the excursion using the trained model 216. For example, the trained model 216 may identify a start time (e.g., the time when the muscle first begins to move), a peak start time (e.g., the start time of maximum muscle contraction), a peak end time (e.g., the end time of maximum muscle contraction), and an end time (e.g., the time when the muscle returns to a baseline or resting state).

[0064] The facial analysis system 120 may fit a function through the data during the beginning and end periods to derive the speed or acceleration of the muscle movement and the smoothness of the muscle movement. In some embodiments, the function may be a Gauss-Newton algorithm. The facial analysis system 120 may do so for each facial muscle that may be moving at the same time or for muscle excursions. For example, if the muscle movement is asymmetric, the facial analysis system 120 may do so for the right and left sides of the muscle.

[0065] The facial analysis system 120 may further determine a smooth start. In some embodiments, to determine a smooth onset, facial analysis system 120 may utilize a function (e.g., a Gauss-Newton algorithm) to describe the onset period and associated statistics (e.g., derivatives) to calculate the goodness of fit between the smoothed curve and the actual data points. Facial analysis system 120 may determine a smooth onset using the first and / or second derivatives of the function, which reflect the rate of change and smoothness of the function. Facial analysis system 120 may perform such analysis on a per motion unit basis. In some embodiments, the function may be an acceleration function, where the accelerations of different muscles are analyzed.

[0066] In some embodiments, to determine a smooth start, facial analysis system 120 may identify or determine the start and apex start times, as well as the acceleration, of every action unit. In some embodiments, if the start and apex start times are the same for every art unit (and for both the right and left sides of the same action unit), facial analysis system 120 may determine that the muscle had the same trajectory from start to apex. In some embodiments, if the acceleration is the same for every art unit (and for both the right and left sides of the same action unit), facial analysis system 120 may determine that the muscle had the same trajectory from start to apex.

[0067] The facial analysis system 120 may further determine a shared vertex duration. For example, the facial analysis system 120 may identify a vertex start time and an end time for each significant action unit. The facial analysis system 120 may make the determination of the shared vertex duration. For example, the facial analysis system 120 may consider an action unit to have a shared vertex duration if the vertex start time and end time of each significant action unit are within a threshold tolerance (e.g., 15-30 ms) of each other.

[0068] Facial analysis system 120 may further determine symmetry. In some embodiments, facial analysis system 120 may determine intensity symmetry and / or timing symmetry. To determine intensity symmetry, facial analysis system 120 may identify the intensity levels of significant action units in the configuration at the vertices separately for both the left and right sides of the face. If the intensity levels of both the right and left sides of the same action units are the same, facial analysis system 120 may classify them as symmetric, otherwise facial analysis system 120 may classify them as asymmetric.

[0069] In some embodiments, to determine timing symmetry, if the difference in apex start time between the right and left sides of any muscle is less than a threshold (e.g., 15-30 ms), facial analysis system 120 may classify it as symmetric, otherwise facial analysis system 120 may classify it as asymmetric. In some embodiments, to determine timing symmetry, if the difference in acceleration rate between the right and left sides of any muscle is less than a threshold, facial analysis system 120 may classify it as symmetric, otherwise facial analysis system 120 may classify it as asymmetric.

[0070] The facial analysis system 120 may further determine a smooth exit. In some embodiments, to determine a smooth exit, the facial analysis module may use a function generated to describe the exit period and associated statistics (derivatives) to calculate the goodness of fit between the smoothed curve and the actual data points. For example, the facial analysis system 120 may determine a smooth exit using the first and second derivatives of the function, which reflect the rate of change and smoothness of the function. In some embodiments, the function may be an acceleration function, where the accelerations of different muscles are analyzed. The facial analysis system 120 may perform such processing for all motion units.

[0071] In some embodiments, to determine a smooth end, facial analysis system 120 may determine that the vertex end time and end time are the same for all action units and for both the right and left sides of the same action unit. If facial analysis system 120 determines that the vertex end time and end time are the same, facial analysis system 120 may conclude that the muscle had the same trajectory from vertex to end.

[0072] Facial analysis system 120 may generate a final classification of the emotional facial expression based on the initial classification. For example, facial analysis system 120 may classify possible emotional facial expressions having simultaneous onset, smooth onset, shared vertices, symmetrical and smooth ending as emotional facial expressions. Facial analysis system 120 may retain possible emotional facial expression classifications for possible emotional facial expressions having simultaneous onset, smooth onset, shared vertices, symmetrical or smooth ending. In some embodiments, facial analysis system 120 may determine that the emotional facial expression was spontaneous based on the determined duration of the emotional facial expression.

[0073] In step 308, the facial analysis system 120 may output an expression classification. 4 is a chart 400 illustrating example intensity levels and timing characteristics of action units, according to an example embodiment. As shown, chart 400 may show sample smile dynamics for action unit 6 (AU6) and action unit 12 (AU12) over a five second period.

[0074] According to the chart 400, AU6 and AU12 are identified as co-climaxing (BC) and are therefore classified as "configuration" by the facial analysis system 120. The facial analysis module may divide the action units into "significant" and "non-significant" action units. As shown in this example, both AU6 and AU12 are significant action units. The facial analysis system 120 may compare the "significant" action units in the facial configuration with known significant action units associated with emotional facial expressions to identify a match with the known emotional facial expressions. In the current example, the facial analysis system 120 may determine that the significant action units match the significant action units for the classification of "happy smile".

[0075] If facial analysis system 120 determines that non-significant action units are present (in this example, they are not), facial analysis system 120 compares the non-significant action units in the facial configuration to "acceptable" and "not acceptable" lists of non-significant action units to further refine the classification of possible emotional facial expressions.

[0076] The facial analysis system 120 may classify significant action units in a facial configuration that match known significant action units that match emotional facial expressions and that include "acceptable" non-significant AUs as "possible emotional facial expressions."

[0077] Facial analysis system 120 may further analyze possible emotional facial expressions according to whether they meet behavioral dynamics criteria. In doing so, facial analysis system 120 may perform one or more of the following operations. For example, facial analysis system 120 may determine simultaneous start, smooth start (e.g., both action units have the same smooth start acceleration), shared vertex duration (e.g., both action units have the same vertex time), strength symmetry (e.g., right and left sides of both action units have the same strength), timing symmetry (e.g., right and left sides of both action units have the same timing characteristics), smooth end (e.g., both action units have the same smooth end deceleration).

[0078] Based on the one or more actions, facial analysis system 120 may determine a final classification of the facial expression of emotion. For example, the configuration identified in Figure 4 may have a simultaneous onset, a smooth onset, shared vertices, symmetry, and a smooth end. Thus, facial analysis system 120 classifies this configuration as a "happy smile."

[0079] FIG. 5A illustrates an architecture of a system bus computing system 500, according to an exemplary embodiment. One or more components of the system 500 may communicate electronically with each other using a bus 505. The system 500 may include a processor (e.g., one or more CPUs, GPUs, or other types of processors) 510 and a system bus 505 that couples various system components to the processor 510, including a system memory 515, such as a read-only memory (ROM) 520 and a random access memory (RAM) 525. The system 500 may include a cache of high-speed memory directly connected to, adjacent to, or integrated as part of the processor 510. The system 500 may copy data from the memory 515 and / or storage device 530 to the cache 512 for quick access by the processor 510. In this manner, the cache 512 may provide a performance boost that avoids delays of the processor 510 while waiting for data. These and other modules may control or be configured to control the processor 510 to perform various actions. Other system memories 515 may be available as well. The memory 515 may include multiple different types of memory with different performance characteristics. The processor 510 may represent a single processor or multiple processors. The processor 510 may include one or more of general purpose processors or hardware or software modules, such as service 1 532, service 2 534, and service 5 536 stored in the storage device 530, configured to control the processor 510, and special purpose processors where software instructions are integrated into the actual processor design. The processor 510 may essentially be a fully self-contained computing system including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0080] To allow user interaction with the system 500, the input device 545 can be any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. The output device 535 (e.g., a display) can also be one or more of several output mechanisms known to those skilled in the art. In some examples, a multimodal system can allow a user to provide multiple types of input to communicate with the system 500. The communication interface 540 can generally govern and manage user input and system output. There is no restriction to operating on any particular hardware configuration, and thus the basic features herein may be easily substituted with improved hardware or firmware configurations as they are developed.

[0081] The storage device 530 may be a non-volatile memory and can be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as a magnetic cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cartridge, a random access memory (RAM) 525, a read-only memory (ROM) 520, and hybrids thereof.

[0082] The storage device 530 may include services 532, 534, and 536 for controlling the processor 510. Other hardware or software modules are also contemplated. The storage device 530 may be connected to the system bus 505. In one aspect, a hardware module performing a particular function may comprise a software component stored in a computer-readable medium in association with the necessary hardware components, such as the processor 510, the bus 505, an output device 535 (e.g., a display), etc., to perform that function.

[0083] FIG. 5B illustrates a computer system 550 having a chipset architecture, according to an exemplary embodiment. The computer system 550 may be an example of computer hardware, software, and firmware that may be used to implement the disclosed techniques. The system 550 may include one or more processors 555, which represent any number of physically and / or logically distinct resources capable of executing software, firmware, and hardware configured to perform the identified computations. The one or more processors 555 may communicate with a chipset 560, which may control input to and output from the one or more processors 555. In this example, the chipset 560 may output information to an output 565, such as a display, and may read and write information to a storage device 570, which may include, for example, magnetic media and solid-state media. The chipset 560 may also read data from and write data to the storage device 575 (e.g., RAM). A bridge 580 for interfacing with various user interface components 585 may be provided for interfacing with the chipset 560. Such user interface components 585 may include a keyboard, a microphone, touch detection and processing circuitry, a pointing device such as a mouse, etc. In general, input to the system 550 may come from any of a variety of machine-generated and / or human-generated sources.

[0084] The chipset 560 can also interface with one or more communication interfaces 590, which can have different physical interfaces. Such communication interfaces can include interfaces for wired and wireless local area networks, broadband wireless networks, and personal area networks. Some applications of the methods for generating, displaying, and using GUIs disclosed herein can involve receiving an ordered data set via a physical interface, or can be generated by the machine itself by one or more processors 555 analyzing data stored in the storage device 570 or 575. Additionally, the machine can receive inputs from a user via a user interface component 585 and perform appropriate functions, such as browsing functions, by interpreting these inputs using one or more processors 555.

[0085] It will be appreciated that the exemplary systems 500 and 550 can have more than one processor 510 or be part of a group or cluster of computing devices networked together to provide greater processing power.

[0086] The above is directed to the embodiments described herein, although other and further embodiments may be devised without departing from the basic scope thereof. For example, aspects of the disclosure may be implemented by hardware or software, or a combination of hardware and software. One embodiment described herein may be implemented as a program product for use with a computer system. The program of the program product defines the functions of the embodiments (including the methods described herein) and can be included in various computer-readable storage media. Exemplary computer-readable storage media include, but are not limited to, (i) non-writable storage media in which information is permanently stored (e.g., a read-only memory (ROM) device in a computer, such as a CD-ROM disk readable by a CD-ROM drive, a flash memory, a ROM chip, or any type of solid-state non-volatile memory), and (ii) writable storage media in which changeable information is stored (e.g., a floppy disk in a diskette drive or hard disk drive, or any type of solid-state random access memory). Such computer-readable storage media, when carrying computer-readable instructions that direct the functions of the disclosed embodiments, are embodiments of the disclosure.

[0087] Those skilled in the art will understand that the above examples are illustrative and not limiting. All permutations, extensions, equivalents, and improvements that are apparent to those skilled in the art upon reading this specification and examining the drawings are intended to be included within the true spirit and scope of the present disclosure. Accordingly, the following appended claims are intended to include all such modifications, permutations, and equivalents that fall within the true spirit and scope of these teachings.

Claims

1. It is a method, A video data identification step, comprising: a computing system identifying video data capturing one or more expressors showing the facial behavior of the faces of one or more expressors, wherein the video data may be any one of the following: previously recorded video data uploaded to the computing system for analysis, video data acquired in real time from a live video camera, and video data acquired by any other means in which facial behavior can be input to the system; The computing system comprises the steps of analyzing the video data to determine one or more types of emotions expressed by the person expressing them in the video data, A step of identifying changes in appearance generated by facial muscle movements in the aforementioned video data, A step of determining the timing characteristics of the movement of the facial muscles in the video data, wherein the timing characteristics indicate whether the facial behavior shown by the person expressing it is a genuine or fake expression, and the timing characteristics determination step is A step of determining whether the movement of a particular facial muscle or a particular combination of facial muscles includes a simultaneous and smooth initiation, A step of determining whether the movement of the aforementioned specific facial muscle or a particular combination of facial muscles reaches its peak almost simultaneously, that is, whether it reaches the apex of muscle contraction. A step of determining whether the movement of the aforementioned particular facial muscle or particular combination of facial muscles simultaneously relaxes, i.e., decreases from its peak, A step performed by a timing characteristic determination step, which includes the step of determining whether the movement and apex of the particular facial muscle or a particular combination of facial muscles are symmetrical. An emotion classification generation step in which the computing system evaluates whether the movement of the facial muscles or combination of muscles includes simultaneous and smooth movement from start to peak, whether the movement of the facial muscles reaches its peak almost simultaneously, whether the movement of the facial muscles or combination of muscles relaxes simultaneously, and whether the movement of the facial muscles or combination of muscles is symmetrical on both sides of the face, thereby generating a classification of one or more types of emotions expressed by the expresser; A method comprising the steps of: the computing system outputting the classification.

2. The aforementioned video data identification step is: The method according to claim 1, comprising the step of receiving a plurality of still images of the exhibitor, wherein the plurality of still images include a first set of images including an image capturing the activation of a particular facial muscle or combination of facial muscles, and a second set of images including a neutral image showing the facial muscle or combination of facial muscles in an inactive or prior resting state.

3. The emotion classification generation step is: A step of determining simultaneous termination by comparing the timing of the specific facial muscle or combination of facial muscles in the neutral image or the second set of images including the prior static state, A step of determining a smooth start by comparing the change from the neutral image or the prior stationary state to the vertex across the second set of images, A step of comparing the time it takes for each of the aforementioned specific facial muscles or combinations of facial muscles to return from the peak to a neutral or resting state, The method according to claim 2, comprising the step of determining symmetry by evaluating the amplitude of muscle activation on the left and right sides of the face for each facial muscle or combination of facial muscles.

4. A non-temporary computer-readable medium comprising one or more sets of instructions, wherein the instructions, when executed by a processor, are used by a computing system. A video data identification step, comprising: a computing system identifying video data capturing one or more expressors showing the facial behavior of the faces of one or more expressors, wherein the video data may be any one of the following: previously recorded video data uploaded to the computing system for analysis, video data acquired in real time from a live video camera, and video data acquired by any other means in which facial behavior can be input to the system; The computing system comprises the steps of analyzing the video data to determine one or more types of emotions expressed by the person expressing them in the video data, A step of identifying changes in appearance generated by facial muscle movements in the aforementioned video data, A step of determining the timing characteristics of the movement of the facial muscles in the video data, wherein the timing characteristics indicate whether the facial behavior shown by the person expressing it is a genuine or fake expression, and the timing characteristics determination step is A step of determining whether the movement of a particular facial muscle or a particular combination of facial muscles includes a simultaneous and smooth initiation, A step of determining whether the movement of the aforementioned specific facial muscle or a particular combination of facial muscles reaches its peak almost simultaneously, that is, whether it reaches the apex of muscle contraction. A step of determining whether the movement of the aforementioned particular facial muscle or particular combination of facial muscles simultaneously relaxes, i.e., decreases from its peak, A step performed by a timing characteristic determination step, which includes the step of determining whether the movement and apex of the particular facial muscle or a particular combination of facial muscles are symmetrical. An emotion classification generation step in which the computing system evaluates whether the movement of the facial muscles or combination of muscles includes simultaneous and smooth movement from start to peak, whether the movement of the facial muscles reaches its peak almost simultaneously, whether the movement of the facial muscles or combination of muscles relaxes simultaneously, and whether the movement of the facial muscles or combination of muscles is symmetrical on both sides of the face, thereby generating a classification of one or more types of emotions expressed by the expresser; A non-temporary computer-readable medium that causes the computing system to perform an operation comprising the step of outputting the classification.

5. The aforementioned video data identification step is: A non-temporary computer-readable medium according to claim 4, comprising the step of receiving a plurality of still images of the exhibitor, wherein the plurality of still images include a first set of images including an image capturing the activation of a particular facial muscle or combination of facial muscles, and a second set of images including a neutral image showing the facial muscle or combination of facial muscles in an inactive or prior resting state.

6. The emotion classification generation step is: A step of determining simultaneous termination by comparing the timing of the specific facial muscle or combination of facial muscles in the neutral image or the second set of images including the prior static state, A step of determining a smooth start by comparing the change from the neutral image or the prior stationary state to the vertex across the second set of images, A step of comparing the time it takes for each of the aforementioned specific facial muscles or combinations of facial muscles to return from the peak to a neutral or resting state, A non-temporary computer-readable medium according to claim 5, comprising the step of determining symmetry by evaluating the amplitude of muscle activation on the left and right sides of the face for each facial muscle or combination of facial muscles.

7. It is a system, Processor and The system comprises a memory having programming instructions stored within it, and the programming instructions, when executed by the processor, are transmitted to the system. A video data identification step, comprising: a computing system identifying video data capturing one or more expressors showing the facial behavior of the faces of one or more expressors, wherein the video data may be any one of the following: previously recorded video data uploaded to the computing system for analysis, video data acquired in real time from a live video camera, and video data acquired by any other means in which facial behavior can be input to the system; The computing system comprises the steps of analyzing the video data to determine one or more types of emotions expressed by the person expressing them in the video data, A step of identifying changes in appearance generated by facial muscle movements in the aforementioned video data, A step of determining the timing characteristics of the movement of the facial muscles in the video data, wherein the timing characteristics indicate whether the facial behavior shown by the person expressing it is a genuine or fake expression, and the timing characteristics determination step is A step of determining whether the movement of a particular facial muscle or a particular combination of facial muscles includes a simultaneous and smooth initiation, A step of determining whether the movement of the aforementioned specific facial muscle or a particular combination of facial muscles reaches its peak almost simultaneously, that is, whether it reaches the apex of muscle contraction. A step of determining whether the movement of the aforementioned particular facial muscle or particular combination of facial muscles simultaneously relaxes, i.e., decreases from its peak, A step performed by a timing characteristic determination step, which includes the step of determining whether the movement and apex of the particular facial muscle or a particular combination of facial muscles are symmetrical. An emotion classification generation step in which the computing system evaluates whether the movement of the facial muscles or combination of muscles includes simultaneous and smooth movement from start to peak, whether the movement of the facial muscles reaches its peak almost simultaneously, whether the movement of the facial muscles or combination of muscles relaxes simultaneously, and whether the movement of the facial muscles or combination of muscles is symmetrical on both sides of the face, thereby generating a classification of one or more types of emotions expressed by the expresser; A system that causes the computing system to perform an operation comprising the step of outputting the classification.

8. The aforementioned video data identification step is: The system according to claim 7, comprising the step of receiving a plurality of still images of the exhibitor, wherein the plurality of still images include a first set of images including an image capturing the activation of a particular facial muscle or combination of facial muscles, and a second set of images including a neutral image showing the facial muscle or combination of facial muscles in an inactive or prior resting state.

9. The emotion classification generation step is: A step of determining simultaneous termination by comparing the timing of the specific facial muscle or combination of facial muscles in the neutral image or the second set of images including the prior static state, A step of determining a smooth start by comparing the change from the neutral image or the prior stationary state to the vertex across the second set of images, A step of comparing the time it takes for each of the aforementioned specific facial muscles or combinations of facial muscles to return from the peak to a neutral or resting state, The system according to claim 8, comprising the step of determining symmetry by evaluating the amplitude of muscle activation on the left and right sides of the face for each facial muscle or combination of facial muscles.

10. A method, A computing system includes the steps of identifying video data capturing the expresser showing the facial behavior of the expresser's face, The computing system comprises the steps of analyzing the video data to determine one or more types of emotions expressed by the person expressing them in the video data, A step of identifying changes in appearance generated by facial muscle movements in the aforementioned video data, A step of determining the timing characteristics of the movement of the facial muscles in the video data, wherein the timing characteristics indicate whether the facial behavior shown by the person expressing it is a genuine or fake expression, and the timing characteristics determination step is A step of determining whether the aforementioned movement of a particular facial muscle includes a simultaneous and smooth initiation, A step of determining whether the movements of the aforementioned specific facial muscles reach their peak almost simultaneously, A step of determining whether the movement of the specific facial muscle relaxes simultaneously, A step performed by a timing characteristic determination step, which includes a step of determining whether the movement of the particular facial muscle is symmetrical, An emotion classification generation step in which the computing system generates a classification of one or more types of emotions expressed by the expresser by evaluating whether the movement of the facial muscle includes simultaneous and smooth movement from start to apex, whether the movement of the facial muscle reaches its peak almost simultaneously, whether the movement of the facial muscle relaxes simultaneously, and whether the movement of the facial muscle is symmetrical on both sides of the face, A method comprising the steps of: the computing system outputting the classification.

11. A non-temporary computer-readable medium comprising one or more sets of instructions, wherein the instructions, when executed by a processor, are used by a computing system. The computing system includes the step of identifying video data capturing the expresser showing the facial behavior of the expresser's face, The computing system comprises the steps of analyzing the video data to determine one or more types of emotions expressed by the person expressing them in the video data, A step of identifying changes in appearance generated by facial muscle movements in the aforementioned video data, A step of determining the timing characteristics of the movement of the facial muscles in the video data, wherein the timing characteristics indicate whether the facial behavior shown by the person expressing it is a genuine or fake expression, and the timing characteristics determination step is A step of determining whether the aforementioned movement of a particular facial muscle includes a simultaneous and smooth initiation, A step of determining whether the movements of the aforementioned specific facial muscles reach their peak almost simultaneously, A step of determining whether the movement of the specific facial muscle relaxes simultaneously, A step performed by a timing characteristic determination step, which includes a step of determining whether the movement of the particular facial muscle is symmetrical, An emotion classification generation step in which the computing system generates a classification of one or more types of emotions expressed by the expresser by evaluating whether the movement of the facial muscle includes simultaneous and smooth movement from start to apex, whether the movement of the facial muscle reaches its peak almost simultaneously, whether the movement of the facial muscle relaxes simultaneously, and whether the movement of the facial muscle is symmetrical on both sides of the face, A non-temporary computer-readable medium that causes the computing system to perform an operation comprising the step of outputting the classification.