Emotion estimation device
The emotion estimation device uses a database of valence and activity levels to extract facial movements, allowing for more accurate classification of complex human emotions beyond simple categories.
Patent Information
- Application Number
- JP2024024339
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-21
- Publication Date
- 2025-09-02
AI Technical Summary
Existing emotion estimation devices can only categorize human emotions into a limited number of types, such as 'neutral', 'happy', 'irritated', and 'depressed', failing to accurately capture the complexity of human emotions.
An emotion estimation device that utilizes a database storing emotion data defined by emotional valence and activity levels, extracting facial movements from image data to estimate emotions as a combination of valence and activity, enabling more accurate emotion classification.
Enables the objective and quantitative estimation of complex human emotions by defining emotions through valence and activity levels, improving the accuracy of emotion recognition.
Smart Images

Figure 2025127571000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a device for estimating a person's emotions from image data of the person's face. [Background technology]
[0002] 2. Description of the Related Art There is known a device that estimates a person's emotions based on imaging data obtained by imaging the person's face.
[0003] As an example of such a device, Patent Document 1 discloses an emotion estimation device that detects microexpressions appearing on a person's face from a facial image obtained by capturing the person's face and estimates the person's emotion based on the proportion of each type of detected microexpression. Specifically, the emotion estimation device described in Patent Document 1 detects six types of microexpressions, "happiness," "anger," "sadness," "disgust," "fear," and "surprise," from a facial image obtained by capturing the face, and estimates the subject's emotion as either "neutral," "happy," "irritated," or "depressed" based on the proportion of each detected microexpression. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-189703 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the emotion estimation device described in Patent Document 1 can only estimate complex human emotions into one of four types: "neutral," "happy," "irritated," and "depressed."
[0006] The present invention is intended to solve the above-mentioned problems, and has an object to provide a technology for more accurately estimating human emotions. [Means for solving the problem]
[0007] The emotion estimation device of the present invention comprises: a database storing emotion data defining multiple types of human emotions based on emotional valence, which is an index representing the quality of an emotion, and activity, which is an index representing the strength of an emotion; an imaging data acquisition unit that acquires imaging data of the face of a subject; an estimation unit that extracts facial movements from the image capture data of the subject's face acquired by the image capture data acquisition unit, and estimates the subject's emotion as an emotion defined by the valence and the activity level based on the extracted facial movements and the emotion data stored in the database; The present invention is characterized by comprising: [Effects of the Invention]
[0008] According to the emotion estimation device of the present invention, the emotion of a subject is estimated as an emotion defined by valence, which is an index representing the quality of the emotion, and activity, which is an index representing the intensity of the emotion, based on emotion data stored in a database and the facial movements of the subject extracted from image capture data, thereby enabling more accurate estimation of complex human emotions. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram schematically illustrating a configuration of a feeling estimation device according to a first embodiment of the present invention. [Figure 2] FIG. 1 is a diagram for explaining emotion data defined by five levels of emotional valence and five levels of activity. [Figure 3] FIG. 1 is a diagram showing an example of a circular model that defines various emotions according to emotional valence and activation level. [Figure 4] FIG. 10 is a diagram showing an example of the positions of a plurality of image capturing units when capturing an image of a facial expression shown by a subject. [Figure 5]1A and 1B are diagrams for explaining an example of the positional relationship of ten imaging devices, where FIG. 1A is a diagram for explaining the positional relationship of multiple imaging devices placed at different heights, and FIG. 1B is a diagram for explaining the positional relationship of seven imaging devices placed at approximately the same height as the subject's face. [Figure 6] FIG. 10 is a diagram showing an example of a still image obtained from a moving image of a subject's face captured by ten imaging units. [Figure 7] FIG. 10 is a diagram schematically illustrating a configuration of a feeling estimation device according to a second embodiment of the present invention. [Figure 8] This is a diagram to explain events that refer to the affect grid (Russell, Weiss & Mendelsohn, 1989). [Figure 9] (a) is a perspective view showing an example of the arrangement of multiple imaging units when capturing the facial expressions of a subject, (b) is a front view showing an example of the arrangement of multiple imaging units, and (c) is a diagram showing an example of a still image obtained from a video of the subject's face captured by 10 imaging units. [Figure 10] FIG. 10 is a diagram showing the peak intensity of each facial movement. [Figure 11] FIG. 10 is a diagram showing a heat map of the loadings of each component for facial expressions of all events. [Figure 12] This figure shows the time changes of the four components of facial expressions for 25 events, and the ribbons represent ±1 standard error. DETAILED DESCRIPTION OF THE INVENTION
[0010] The features of the present invention will be described in more detail below by showing embodiments of the present invention.
[0011] First Embodiment 1 is a diagram schematically illustrating the configuration of a feeling estimation device 10 according to a first embodiment of the present invention. The feeling estimation device 10 according to the first embodiment includes a database 1, an imaging data acquisition unit 2, and an estimation unit 3.
[0012] Database 1 stores emotion data that defines multiple types of human emotions using valence, an index representing the quality of an emotion, and arousal, an index representing the intensity of an emotion. More specifically, database 1 stores emotion data that defines multiple types of human emotions in a matrix format using valence and arousal. Valence represents the quality of an emotion, i.e., "pleasant" or "unpleasant." Here, it is assumed that database 1 stores emotion data (see FIG. 2 ) that defines human emotions using five levels of valence and five levels of arousal. The five levels of valence can be defined, for example, as "very unpleasant," "unpleasant," "neutral," "pleasant," and "very pleasant." The five levels of arousal can be defined, for example, as "very low arousal," "low arousal," "moderate arousal," "high arousal," and "very high arousal." However, the levels defining the valence of the emotion data stored in database 1 are not limited to five levels, and the levels defining arousal are not limited to five levels.
[0013] Figure 3 shows an example of a circular model that defines various emotions according to valence and activation level. The circular model shown in Figure 3 shows 24 emotions: "activation," "surprise," "excitement," "happy," "happiness," "delight," "satisfaction," "fulfillment," "calm," "calm," "relaxation," "drowsiness," "fatigue," "boredom," "depression," "sad," "unpleasant," "frustration," "suffering," "anger," "fear," "tension," and "alertness." However, emotions according to valence and activation level are not limited to the 24 emotions shown in Figure 3.
[0014] In the circular model shown in Figure 3, "activation," "surprise," "excitement," "happy," "happiness," "delight," "satisfied," "fulfilled," "calm," "relaxed," and "sleepiness" are emotions with a "pleasant" valence, while "fatigue," "boredom," "depression," "sad," "unpleasant," "frustration," "suffering," "anger," "fear," "tension," and "alertness" are emotions with an "unpleasant" valence. Furthermore, when activity levels are classified into "high" and "low," "frustration," "suffering," "anger," "fear," "tension," "alertness," "activation," "surprise," "excitement," "happy," "happiness," and "joy" are emotions with a high activation level, while "satisfaction," "fulfillment," "calm," "relaxed," "sleepiness," "fatigue," "boredom," "depression," "depressed," "sad," and "unpleasant" are emotions with a low activation level.
[0015] Valence and activation are the core of emotions, and it is believed that all human emotions can be expressed through a combination of valence and activation.
[0016] The emotion data stored in the database 1 can be prepared by a method described later.
[0017] The imaging data acquisition unit 2 acquires imaging data of the face of a subject whose emotion is to be estimated. The imaging data acquisition unit 2 acquires the imaging data of the face of the subject via, for example, an internet line. Note that in this specification, the manner in which the imaging data acquisition unit 2 acquires imaging data also includes receiving transmitted imaging data.
[0018] The image data of the subject's face acquired by the image data acquisition unit 2 is sent to the estimation unit 3.
[0019] The estimation unit 3 extracts facial movements from the imaging data of the subject's face acquired by the imaging data acquisition unit 2, and estimates the subject's emotion as an emotion defined by valence and activity level based on the extracted facial movements and the emotion data stored in the database 1. A detailed method for estimating the subject's emotion by the estimation unit 3 will be described later.
[0020] The estimation unit 3 is, for example, a computer, and includes a CPU, a memory, an input / output interface, etc. The estimation unit 3 is configured to be able to access at least the database 1.
[0021] The following describes a method for creating emotion data to be stored in database 1. Here, the explanation is based on the assumption that emotion data defined by five levels of valence and five levels of activation as shown in FIG.
[0022] To create emotion data, multiple subjects (for example, 48 subjects of various ages) are prepared, and each of the subjects is asked to imagine in advance an event that evokes one of 25 emotions (5 × 5) corresponding to five levels of valence and five levels of activation (see Figure 2). That is, subjects are asked to imagine in advance an event that evokes one of the 25 emotions corresponding to five levels of valence and five levels of activation, such as an event that evokes an emotion with a valence of "very unpleasant" and an activation level of "very low arousal," and an event that evokes an emotion with a valence of "very unpleasant" and an activation level of "low arousal."
[0023] For example, one subject imagined "working at a job they dislike" as an event that would evoke feelings of "very unpleasant" valence and "medium arousal," and "passing a university entrance exam" as an event that would evoke feelings of "very pleasant" valence and "very high arousal."
[0024] Next, the subject is verbally told about the assumed event, and asked to express a facial expression corresponding to the event. The facial expressions are then captured by multiple image capture devices. If the subject is wearing a mask, it is preferable to capture the facial image with the mask removed. If the subject is wearing glasses, it is preferable to capture the facial image with the glasses removed.
[0025] It is also possible to have subjects express facial expressions without imagining in advance the events that will evoke each of the 25 emotions corresponding to the five levels of valence and five levels of activation. For example, subjects are asked to "make an expression of very pleasantness and very high arousal," and the expressions are captured using multiple imaging devices. In this case, too, subjects are asked to express 25 emotions corresponding to the five levels of valence and five levels of activation, and each facial expression is captured using multiple imaging devices.
[0026] Alternatively, the subject may be shown sad or funny images according to the emotional valence and activity level, and the facial expressions of the subject at that time may be captured by a plurality of image capturing devices.
[0027] There are no particular restrictions on the type of imaging device used to capture the subject's face; for example, a Microsoft device called "Azure Kinect" or a camera from Intel's "Real Sense" series can be used. The number of imaging devices is, for example, 10. However, the number of imaging devices is not limited to 10, and the types of imaging devices are not limited to "Azure Kinect" or a camera from Intel's "Real Sense" series.
[0028] The imaging device may be provided with a depth sensor, which allows distance information to be acquired when capturing an image of the subject's face, thereby enabling more accurate extraction of the subject's facial movements, as will be described later.
[0029] 4 is a diagram showing an example of the positions of multiple image capture devices 4 used when capturing an image of a facial expression displayed by a subject. The multiple image capture devices 4 are preferably placed within the range of motion of the subject's neck and at positions where they can capture an image of the facial expression. The range of motion of the subject's neck is, for example, 0° to 50° inclusive when flexing (bending forward), 0° to 80° inclusive when extending (bending backward), 0° to 80° inclusive when rotating right, and 0° to 80° inclusive when rotating left.
[0030] FIG. 4 shows an example in which ten imaging devices 4 are arranged. Of the ten imaging devices 4, one imaging device 4 is arranged at a position higher than the subject's face, seven imaging devices 4 are arranged at a position approximately at the same height as the subject's face, and two imaging devices 4 are arranged at a position lower than the subject's face. The imaging device 4 arranged at a position higher than the subject's face images the subject's face from diagonally above. The imaging device 4 arranged at a position lower than the subject's face images the subject's face from diagonally below. The seven imaging devices 4 arranged at a position approximately at the same height as the subject's face each image the subject's face from a different direction. Each of the ten imaging devices 4 is configured so that the orientation of the lens can be adjusted so that the subject's face can be imaged.
[0031] 5(a) and 5(b) are diagrams illustrating an example of the positional relationship of ten imaging devices 4. Fig. 5(a) is a diagram illustrating the positional relationship of multiple imaging devices 4 placed at different heights, and Fig. 5(b) is a diagram illustrating the positional relationship of seven imaging devices 4 placed at approximately the same height as the face of the subject.
[0032] 5(a), the angle θ1 between the horizontal line L1 and an image capture device 4 positioned higher than the subject's face is, for example, 80°±10°. The angle θ2 between the horizontal line L1 and an image capture device 4 positioned lower than the subject's face is, for example, 50°±10°. In FIG. 5(b), the angles θ3 between the image capture device 4 located on the left side of the subject facing forward and the image capture device 4 positioned adjacent to it, and between the image capture device 4 located on the right side of the subject facing forward and the image capture device 4 positioned adjacent to it, are, for example, 45°±10°, and the angle θ4 between any two other adjacent image capture devices 4 is, for example, 22.5°±10°.
[0033] When capturing an image of the subject's face using an imaging device 4, it is preferable to provide lights 5 to illuminate the subject's face, as shown in Fig. 4. Fig. 4 shows an example in which a total of three lights 5 are provided: two lights 5 to illuminate the subject's face from diagonally above, and one light 5 to illuminate the subject's face from diagonally below. By capturing an image of the subject's face with the lights 5 illuminating it, clear image data of the face without shadows can be obtained. However, the number of lights 5 provided is not limited to three.
[0034] Each of the multiple image capture devices 4 captures an image of the subject's face for a predetermined time, for example, four seconds. As an example, the subject expresses the emotion evoked by the conveyed event in the first second, maintains the expressed facial expression for two seconds, and then returns to the original expression in the next second. Each of the multiple image capture devices 4 captures, as a video, the changes in the subject's face when they are informed of an event that evokes an emotion corresponding to the emotional valence and activation level.
[0035] Fig. 6 is a diagram showing an example of a still image obtained from a moving image of a subject's face captured by ten imaging devices 4. As shown in Fig. 6, by capturing an image of the subject's face using multiple imaging devices 4, imaging data of the subject's face captured from various directions can be obtained.
[0036] Next, facial movements are extracted from the captured video data of the subject's face to evaluate facial expression patterns. The facial movements to be extracted preferably include at least one of the following, and more preferably all of them: "inner eyebrow lift," "outer eyebrow lift," "eyebrow lowering," "upper eyelid lift," "cheek lift," "eyelid tightening," "nose wrinkles," "upper lip lift," "lip corners lift," "dimples," "lip corners lowering," "chin lift," "lip stretching to the side," "lip closing," "lip pressing," "chin lowering," and "blinking." These facial movements are sometimes called action units. However, the facial movements to be extracted are not limited to those described above.
[0037] For example, when the emotional valence is on the "unpleasant" side, facial movements such as "lowering eyebrows" and "tightening eyelids" are strongly expressed. Also, when the emotional valence is on the "pleasant" side, facial movements such as "lifting cheeks," "tightening eyelids," "upper lip lift," "lip corners lift," and "dimples" are expressed. Also, when the emotional valence is on the "pleasant" side, the higher the arousal level, the more frequently the mouth is opened.
[0038] The facial movement extraction described above is performed for each of 25 emotions (5 × 5) corresponding to five levels of emotional valence and five levels of activity. Once the facial movement is extracted for each of the 25 emotions corresponding to five levels of emotional valence and five levels of activity, the extracted facial movement data is stored in database 1 as emotion data. As described above, this emotion data is facial movement data extracted from multiple videos captured from different angles showing changes in the faces of multiple subjects. To facilitate the emotion estimation process of the subject by estimation unit 3, the extracted facial movement data may be quantified and the quantified data may be stored in database 1 as emotion data.
[0039] Next, a method for estimating the subject's emotion by the estimation unit 3 will be described. The subject's emotion estimated by the estimation unit 3 is a subjective emotion. Reactions to external stimuli include subjective reactions that appear as changes in mood, physiological reactions that are reactions of the autonomic nervous system and immune system, and behavioral reactions that appear as changes in facial expression, posture, etc., and the estimation unit 3 estimates the subjective emotion of the subject based on the subject's behavioral reactions.
[0040] First, imaging data of the subject's face is acquired by the imaging data acquisition unit 2. The acquired imaging data of the subject's face is video data. The imaging data of the subject's face acquired by the imaging data acquisition unit 2 is sent to the estimation unit 3.
[0041] The estimation unit 3 extracts facial movements from the imaging data of the subject's face acquired by the imaging data acquisition unit 2. That is, the estimation unit 3 extracts facial movements of the subject from video data of the subject's face acquired by the imaging data acquisition unit 2. The extracted facial movements of the subject are the same as the facial movements extracted from the imaging data of the subject's face to store emotion data in the database 1, and are, for example, "inner eyebrows rise," "outer eyebrows rise," "eyebrows drop," "upper eyelids rise," "cheeks lift," "eyelids tighten," "nose wrinkles," "upper lip rise," "lip corners rise," "dimples form," "lip corners drop," "chin rise," "lips stretch sideways," "lips close," "lips press together," "chin drops," and "blink."
[0042] Next, the estimation unit 3 estimates the emotion of the subject as an emotion defined by valence and activation level based on the extracted facial movement of the subject and the emotion data stored in the database 1. Specifically, the estimation unit 3 compares the extracted facial movement of the subject with the emotion data, i.e., facial movement data, stored in the database 1, and identifies the valence and activation level corresponding to the extracted facial movement. Then, the estimation unit 3 estimates the emotion defined by the identified valence and activation level as the emotion of the subject.
[0043] For example, if the emotion data stored in database 1 contains a facial movement that is identical to the extracted facial movement of the subject, or is similar enough to be considered identical, the emotion defined by the valence and activity level corresponding to that facial movement is estimated to be the emotion of the subject. If the emotion data stored in database 1 is defined by five levels of valence, from level 1 to level 5, and five levels of activity, from level 1 to level 5, the emotion of the subject can be estimated to be, for example, an emotion with a valence level of 3 and an activity level of 4.
[0044] In the emotion estimation device 10 of this embodiment, the estimation unit 3 can estimate the emotion of a subject as an emotion defined by a greater number of valence levels than the valence levels defining the emotion data and a greater number of activity levels than the activity levels defining the emotion data. For example, if the emotion data stored in the database 1 is defined by five valence levels, from level 1 to level 5, and five activity levels, from level 1 to level 5, the estimation unit 3 can estimate the emotion of the subject as an emotion with a valence level of 3.5 and an activity level of 4.3, based on the extracted facial movement of the subject and the emotion data stored in the database 1. In this way, by estimating the emotion of a subject as an emotion defined by a greater number of valence levels than the valence levels defining the emotion data and a greater number of activity levels than the activity levels defining the emotion data, it is possible to more accurately estimate complex human emotions.
[0045] The estimation unit 3 can estimate the emotion of the subject by, for example, machine learning. For example, by clustering the emotion data of multiple subjects, the data is divided into groups based on similar data groups, and the emotion of the subject is estimated based on the grouped data groups and the facial movement data of the subject. For example, Time Series K-means or K-Shape can be used for the clustering. Furthermore, for example, GRU (Gated Recurrent Unit) or SVR (Support Vector Regression) can be used to estimate the emotional valence and activity level of the subject.
[0046] However, the method for estimating the emotion of the subject by the estimation unit 3 is not limited to the above-mentioned method. The estimation unit 3 can estimate the emotion of the subject by any method as long as it extracts facial movements from the image data of the subject's face captured by the imaging unit 2 and estimates the emotion of the subject based on the extracted facial movements and the emotion data stored in the database 1.
[0047] The target person for emotion estimation is not limited to one person, but may be multiple people. In other words, if the imaging data acquired by the imaging data acquisition unit 2 includes the faces of multiple people, the estimation unit 3 can estimate the emotions of multiple people.
[0048] Furthermore, when the imaging data acquired by the imaging data acquisition unit 2 includes the faces of multiple people, the estimation unit 3 may estimate the emotion of only the subject who is considered to have been in the closest position at the time of imaging. In order to accurately estimate the emotion of the subject, it is preferable that the estimation unit 3 removes noise such as natural light from the imaging data.
[0049] The imaging data acquisition unit 2 may acquire imaging data of the subject's face that has been captured and saved in advance, or may immediately acquire imaging data of the subject's face captured by the imaging unit. When the imaging data acquisition unit 2 immediately acquires imaging data of the subject's face that has been captured, the estimation unit 3 can estimate the subject's emotion in real time based on the imaging data of the subject's face acquired by the imaging data acquisition unit 2.
[0050] According to the emotion estimation device 10 of the present embodiment, facial movements are extracted from the imaging data of the subject's face acquired by the imaging data acquisition unit 2, and the emotion of the subject is estimated as an emotion defined by valence and activity level based on the extracted facial movements and the emotion data stored in the database 1, thereby enabling more accurate estimation of human emotions. That is, by defining human emotions by valence, which is an index representing the quality of the emotion, and activity level, which is an index representing the intensity of the emotion, complex human emotions can be estimated objectively and quantitatively.
[0051] The emotion data stored in database 1 is facial movement data extracted from video footage of the facial changes of multiple subjects who are informed of events that evoke emotions corresponding to valence and activation level. Since each person's face is different, the facial expressions expressed according to emotions also vary from person to person; however, the facial movements corresponding to emotions according to valence and activation level tend to be the same or similar across different people. Therefore, since the emotion data stored in database 1 is facial movement data, the emotions of the subject can be more accurately estimated. Furthermore, because the facial changes of multiple subjects are video-recorded to prepare the emotion data to be stored in database 1, more appropriate emotion data can be stored in database 1 compared to when the face of a single subject is recorded, allowing for more accurate estimation of the subject's emotion.
[0052] The emotion data stored in database 1 is data on the facial movements of multiple subjects, and the facial movements include at least one of "inner eyebrows rise," "outer eyebrows rise," "eyebrows drop," "upper eyelids rise," "cheeks lift," "eyelids tighten," "nose wrinkles," "upper lip lift," "lip corners rise," "dimples," "lip corners drop," "chin lift," "lips stretch sideways," "lips close," "lips press," "chin drop," and "blink," which allows the subject's emotion to be more accurately estimated. In other words, since the above-mentioned facial movements are expressed in accordance with various emotions, including data on these facial movements in the emotion data stored in database 1 allows the subject's emotion to be more accurately estimated from among the various emotions.
[0053] The emotion estimation device 10 in this embodiment can be used in various places. For example, by estimating the emotions of a person who is worried about their mental health and needs objective information about their emotions and communicating that information to that person, it is possible to prevent illnesses caused by unconscious stress accumulation. Furthermore, by estimating the emotions of elderly people receiving care, particularly those with dementia, in nursing homes and homes, it is possible to provide care based on objective emotion data.
[0054] In the future, it is also conceivable that the emotion estimation device 10 of the present invention will be installed in social robots that live with people as partners. That is, the social robot will be able to estimate a person's emotion and provide appropriate social behavior based on the estimated emotion. For example, if the estimated emotion of a person is an emotion with an emotional valence of "unpleasant" and a high level of activity, the social robot will be able to induce the person's emotion to a "pleasant" state by engaging in a relaxing conversation with the person.
[0055] In the future, it is also conceivable to build a smart house equipped with the emotion estimation device 10 of the present invention. That is, it is possible to estimate the emotions of people living in the smart house and to comprehensively control the living environment based on the estimated emotions. As an example, if the estimated emotion of a resident is an emotion with an emotional valence of "unpleasant" and a high level of activity, it is possible to guide the resident's emotion to a "pleasant" state by displaying an image of a relaxing natural environment in a virtual window.
[0056] Furthermore, the emotion estimation device 10 of the present invention can be used to estimate the emotions of soldiers, astronauts, athletes, etc., or to estimate the emotions of people who cannot accurately explain their condition, such as people with dementia, people with autism, and elderly people who are not good at speaking.
[0057] As an example of an embodiment of the emotion estimation device 10 of the present invention, an application that is installed on a smartphone and used can be used. That is, an emotion estimation app is prepared for using the emotion estimation device 10 of the present invention. A user can estimate emotions by using the emotion estimation app installed on their smartphone.
[0058] <Second embodiment> 7 is a diagram schematically illustrating the configuration of a feeling estimation device 10A in the second embodiment. The feeling estimation device 10A in the second embodiment further includes an imaging unit 6 in addition to the configuration of the feeling estimation device 10 in the first embodiment.
[0059] The imaging unit 6 is for capturing an image of the face of the subject whose emotion is to be estimated. The imaging unit 6 may be any device capable of capturing an image of the subject's face, and there are no particular restrictions on the type. The number of imaging units 6 may be one or two or more. When the number of imaging units 6 is two or more, the two or more imaging units 6 can capture images of the subject's face from different angles, allowing the facial movements of the subject to be extracted more accurately, and enabling the estimation unit 3 to more accurately estimate the subject's emotion.
[0060] When obtaining video data of the subject's face, the imaging unit 6 captures the subject's face for a predetermined period of time. It is preferable that the imaging unit 6 captures the subject's face for at least the time it takes to capture the subject's face to obtain emotion data. For example, if the subject's face is captured for four seconds, it is preferable that the imaging unit 6 captures the subject's face for at least four seconds.
[0061] The imaging data acquisition unit 2 acquires imaging data of the subject's face captured by the imaging unit 6. When the imaging data acquisition unit 2 acquires the imaging data immediately after the subject's face is captured by the imaging unit 6, the estimation unit 3 can estimate the subject's emotion in real time.
[0062] Like the feeling estimation device 10 in the first embodiment, the feeling estimation device 10A in the second embodiment can also estimate complex human emotions objectively and quantitatively.
[0063] In the feeling estimation device 10A of the present embodiment, the imaging unit 6 captures an image of the subject's face for at least the time it takes to image the subject's face to acquire emotion data, which allows the estimation unit 3 to more accurately estimate the subject's emotion. That is, if the time it takes to image the subject's face is shorter than the time it takes to image the subject's face, sufficient data on the subject's facial movement may not be obtained, making it impossible to accurately estimate the subject's emotion. However, if the time it takes to image the subject's face is equal to or longer than the time it takes to image the subject's face, the subject's emotion can be more accurately estimated.
[0064] =Paper on emotion estimation device= (Summary) Development of facial expression analysis using sensing information is progressing in interdisciplinary fields such as psychology, affective computing, and cognitive science. Previous facial expression datasets have not simultaneously addressed multiple theoretical views of emotions, individualized contexts, and multiple angle / depth information. Therefore, we developed a new facial database (the RIKEN Facial Expression Database) that includes multiple theoretical views of emotions, individual events of expressors, and multi-angle / depth information. The RIKEN Facial Expression Database consists of recordings of 25 different events of 48 Japanese subjects captured with 10 Kinect cameras. In this study, we identified facial expression patterns associated with several valences and found that these patterns are consistent with previous studies investigating the coherence between facial movements and internal states. This database represents a further step toward developing new sensing systems, conducting psychological experiments, and understanding the complexity of emotional events.
[0065] (Introduction) The development of new facial expression databases will contribute to advances in many fields, including psychology, affective computing, and cognitive science. Recent research has provided facial expression databases suitable for many contexts and research purposes, including deception detection (Lloyd et al., 2019), free speeches (Senturk et al., 2022), group discussions (Girard et al., 2017), spontaneous tears (Kuster, Baker & Krumhuber, 2021), pain-related faces (Lucey et al., 2011; Fernandes-Magalhaes et al., 2022), and social stigma (Workman & Chatterjee, 2021). Currently, the most prominent example is the communication of emotional states (see the systematic survey, Dawel et al., 2021; review, Krumhuber et al., 2017; and meta-database, Workman & Chatterjee, 2021). Based on these findings, numerous applied research studies have been carried out (see reviews by Li and Deng, 2020; Ekundayo & Viriri, 2021). In this paper, we introduce a new facial expression database called the RIKEN Facial Expression Database. This database offers the possibility of achieving multiple objectives depending on the individual research goals of emotional expression.
[0066] There are three issues that must be overcome in order to develop a new emotional expression database: the lack of multiple theoretical concepts of emotions, the lack of individualized context, and the lack of multi-angle and depth information.
[0067] (assignment) (a) Many databases rely on the theory of basic emotions (i.e., six basic emotion categories: Calvo & Lundqvist, 2008; Chung et al., 2019; Ebner et al., 2010; Holland et al., 2019; Langner et al., 2010; LoBue & Thrasher, 2015; Van Der Schalk et al., 2011). Other databases exist that use elicitation-based procedures (e.g., spontaneous expressions: Mavadati et al., 2013; Sneddon et al., 2011; Zhang et al., 2014), which involve expressing emotions in situations that are expected to evoke the categories assumed in the theory of basic emotions. However, facial expression databases based on other theories are lacking. Core Affect Theory uses two vectors, valence and activation, to represent emotional states (Barrett & Russell, 1999). Databases based on this theory are few compared to those based on basic emotion theory or simplified emotion labels (AffectNet: Mollahosseini, Hasani, & Mahoor, 2017; AFEW-VA: Kossaifi, Tzimiropoulos, Todorovic, & Pantic, 2017). Furthermore, these databases use observer-perspective ratings as annotations (except for the Stanford Emotional Narratives Dataset: Ong et al., 2019). Furthermore, the Component Process Model (CPM; Scherer, 2001) assumes that the evaluation results along each evaluation dimension sequentially drive the dynamics of emotion. The only face database that relies on this approach is the actor database developed by Seuss et al. (2019). Theoretical discussions on emotions recommend multiple scientific theories on how to define emotions ( Scarantino, 2012 ).The database, which takes into account various theoretical backgrounds, can be used flexibly depending on research purposes and practice.
[0068] (b) Contextual information is important for perceiving emotional expressions (Barrett et al., 2019; Barrett, Mesquita & Gendron, 2011; Chen & Whitney, 2019). Le Mau et al. (2021) emphasize the role of contextual information in inferring internal states from facial movements. Different people react differently to the same event. When insulted, one person may feel anger, while another may feel contempt or fear. Given that individual developmental history, including cultural learning, influences emotion-related facial movements (Barrett, 2017; Griffiths, 1997), expanding face databases requires setting contexts and conditions tailored to each individual.
[0069] (c) Many face stimuli have been created using RGB images and videos (Calvo & Lundqvist, 2008; Chung et al., 2019; Dawel et al., 2021; Langner et al., 2010; LoBue & Thrasher, 2015; Tottenham et al., 2009; Van Der Schalk et al.). While face databases that include multi-angle data and depth information exist, they are limited in number (Cudeiro et al.). However, multi-angle data and depth information offer several advantages in individual research practices. Psychologists have pointed out the importance of the angle of the face (e.g., whether viewed from the front or in profile) in studying face perception. Guo and Shaw (2015) showed that profile views significantly reduce perceived intensity compared to frontal views. Furthermore, multi-angle information has also attracted attention in computer science. From multiple 2D images taken from multiple angles, state-of-the-art algorithms can generate 3D representations (Mihajlovic, Bansal, Zollhoefer, Tang & Saito, 2022). Similar to 3D representations, 4D information, which adds dynamic information to 3D faces, has attracted attention in psychological research, such as face perception (Burt & Crewther, 2020). For example, Chelnokova and Laeng (2011) demonstrated that 3D faces are more recognizable than 2D faces. To enrich face database science, scholars are developing new databases that directly measure depth information using tools such as Kinect (e.g., Aly et al., 2015; Boccignone, Conte, Cuculo & Lanzarotti, 2017). Therefore, acquiring multi-angle data and depth information is becoming increasingly standard in the affective computing field (Cheng, Kotsia, Pantic & Zafeiriou, 2018; Li et al., 2022; Matuszewski et al., 2012; Seuss et al., 2019).Collecting such information is important for conducting psychological experiments and developing automatic emotion estimation systems.
[0070] In this study, we developed a new face database with multi-angle and depth information, incorporating individualized context and multiple theoretical emotional concepts. More specifically, we aimed to create a face database corresponding to 25 individual events corresponding to emotional valence and arousal (Russell, Weiss, & Mendelsohn, 1989). Furthermore, we obtained free-text labeling data (Haidt & Keltner, 1999) and ratings related to the evaluative dimensions (Scherr, 1997; Scherer et al., 2018) from the 25 events prepared by participants. While the primary purpose of this paper is to report on the developed database, we also quantified facial expressions using automated Facial Action Coding System (FACS) analysis. We analyzed the relationship between the emotional valence, arousal, and arousal dimensions and facial action units (AUs). Since substantial psychophysiological studies using facial electromyography have shown that activity in the frenulum labii superioris (associated with AU4) and zygomaticus major (associated with AU12) is negatively and positively associated with subjective value experiences, respectively (Bradley & Lang, 2000; Greenwald, Cook & Lang, 1989; Lang, Greenwald, Bradley & Hamm, 1993; Larsen, Norris & Cacioppo, 2003; Sato, Fujimura, Kochiyama & Suzuki, 2013; Sato, Kochiyama & Yoshikawa, 2020; Sato, Murata, Uraoka, Shibata, Yoshikawa & Furuta, 2021; Tan et al., 2012), we anticipate that similar characteristics will be observed in this database and conduct an analysis to examine this.
[0071] (method) Participants: Forty-eight Japanese participants (22 women, 26 men) aged 20 to 30 years (mean age 23.33, SD = 3.65) participated in this study. Prior to recording facial movements, all participants provided written informed consent (including information on purpose, methods, risks, right of withdrawal, handling of personal information, and voluntariness of participation). The informed consent also included a statement regarding whether participants agreed to their footage being made public for academic purposes, including psychological experiments and emotion calculations. Participants were paid 13,000 yen for participation and database creation. The RIKEN Ethics Committee (Protocol Number: Wako3 2020-21) approved the experimental procedures and the protocol submitted to participants. The recorded facial expressions were prepared for use in creating a face database, not for estimating population index effect sizes. As a result, power analysis was not possible. However, the number of participants in this study is larger than the typical number used in studies of actors' facial expressions (Scherr, Dieckmann, Unfried, Ellgring, & Mortillaro, 2021), so the statistical results below based on these data are expected to have a certain degree of validity.
[0072] Procedure: One week before the filming session, all participants were instructed to recall and write down 25 events that occurred in their lives from five emotional valence levels (strongly unpleasant, unpleasant, neutral, pleasant, and strongly pleasant) and five activation levels (very low arousal or sleepiness, low arousal, medium arousal, high arousal, and very high arousal). Each individual was assigned a corresponding emotional event (Figure 8). Qualtrics was used as the platform for collecting events. Participants also rated each described event on a scale from 1 (strongly disagree) to 5 (strongly agree) for novelty (predictability: "The event was predictable"; familiarity: "The event was common"), goal importance ("The event was important to you"), and manageability ("By taking appropriate actions, you could control or avoid the event"). Participants were also asked to freely describe any possible labels they could think of for each event. The order of the events requiring responses across valence and activation levels was randomized.
[0073] On the day of the recording, participants were briefed on the experiment and then moved to the recording location. Figures 9(a)–(c) show the recording environment. Participants sat in a chair and fixed their face position. Three filming lights (AL-LED-SQA-W: Toshiba) were set up to illuminate the participants' faces from the top right, top left, and bottom, making them clear and eliminating shadows. The participants were then asked to remove their masks and glasses. Ten Kinect cameras (RGB: 1920 × 1080, Depth: 640 × 576) were used to create a recording environment for facial movements, and video clips of the participants' facial movements were recorded at 30 frames per second. The horizontal camera spacing was 22.5°, and the cameras were used at 22.5°, 45°, and 90° (67.5° was skipped in this database). A program to record facial movements was created using the SDK program. Depth information was limited to eight cameras, one on top, two on bottom, one facing forward, two on the left, and two on the right, to reduce processing load and avoid equipment errors. A green carpet covered the background as much as possible (Figure 9(a)-(c)).
[0074] The timing of each segmentation was indicated by a beep (onset: 880Hz, peak: 1174Hz, offset: 880Hz) emitted from a speaker system, and participants were instructed to make facial expressions according to the time course. The model was instructed to express an emotion based on the previously explained event for the first second, maintain the intended emotion for two seconds, and then return to a neutral expression for one second.
[0075] Data Analysis: We used OpenFace (Baltrusaitis et al., 2018) to extract 17 facial movements to evaluate facial expression patterns: AU1 (eyebrow raise), AU2 (eyebrow lower), AU4 (eyebrow lower), AU5 (upper eyelid raise), AU6 (cheek raise), AU7 (eyelid tighten), AU9 (nose wrinkle), AU10 (upper lip raise), AU12 (mouth corner lower), AU14 (mouth corner lower), AU15 (mouth corner lower), AU17 (chin lift), AU20 (lip stretch), AU23 (lip tighten), AU25 (lip part), AU26 (chin lower), and AU45 (blink). Among automatic facial movement detection systems, OpenFace performed relatively well (Namba et al., 2018). Namba, Sato, and Yoshikawa (2021) also found that OpenFace yields the highest accuracy for frontal facial images, so here we focused only on facial expressions captured by a frontal camera. Given the nature of the procedure, where the strength of facial combinations is expected to be greatest during the peak beep, we focused primarily on the intermediate frames (i.e., frame 61). We used R (4.1.2, R Core Team, 2021) for statistical analysis. We performed text mining for each event using the tm and openxlsx packages (Feinerer & Hornik, 2022; Feinerer, Hornik, & Meyer, 2008; Schauberger & Walker, 2022). We used the psych package (Revelle, 2022) to check for correlations between several evaluation dimensions. We used the nnTensor package (Tsuyuzaki, Ishii, & Nikaido, 2022) to reduce the dimensionality of the data extracted by OpenFace. We used the tidyverse package (Wickham et al., 2019) to visualize the data. As mentioned above, based on ample psychophysiological evidence (e.g., Greenwald et al., 1989), we predicted and analyzed the relationship between emotional valence and AU4 / 12 by hierarchical linear regression modeling using the lmerTest package (Kuznetsova, Brockhoff, & Christensen, 2017). Results were considered significant at p < 0.05.All codes are available at Gakunin RDM (https: / / dmsgrdm.riken.jp:5000 / uphvb / ). The design and analysis of this study were not preregistered.
[0076] (result) Event details As described in the Methods section, we obtained 1,200 events (48 participants × 5 valences × 5 activation levels). All Japanese events were translated into English and back-translated using TEXT (https: / / www.text-edit.com / english-page / ). Table 1 shows the top three most frequently occurring English words for each event obtained through text mining. In the resulting database, words that appeared to be common events (frequency 10 / 48 or higher) appeared in events with a valence of 4*activation level of 4 (friends) and a valence of 5*activation level of 5 (university entrance exam pass). The latter, in particular, indicates that the university entrance exam significantly influences emotional events, as this study was limited to younger participants.
[0077] [Table 1]
[0078] The "..." in Table 1 indicates that there are multiple words with the same frequency as the word in the table. V represents valence, and A represents activation. "Get / got" was removed because it is used as a "be verb" in Japanese.
[0079] The correspondence between emotional valence, activation, and evaluation is also shown in Table 2. Emotional valence and activation appeared to be positively correlated with the evaluation of the importance of each event (rs > .21). It was also revealed that the more predictable the event, the higher the activation of that event (r = 0.22). Furthermore, the more unfamiliar the event, the higher the activation (r = -.19). Regarding the correlations between evaluation dimensions, positive correlations were found between predictability and familiarity, and between predictability and controllability (rs > .29).
[0080] [Table 2]
[0081] Tables 3 and 4 show the most frequently occurring emotion labels using the free-text data. Only the top 18 modes (N = 730 / 1200) are listed. Positive emotion labels such as joy, happiness, and pleasure showed high emotional valence, while negative emotion labels such as anger, sadness, and displeasure showed low emotional valence. Activation was high for surprise and impatience. Interesting correspondences were also found between the controllability element and frustration, predictability, and pleasure, but this study did not go into detail as it was beyond the scope of this study's objective to provide an overview of the database.
[0082] [Table 3] [Table 4]
[0083] (Facial movement details) Although facial data for some events was missing due to camera failure or participant issues, the events themselves were recorded as described above. The number of usable frames was 142,865. Extracting only peak frames yielded 1,190 frames, with six missing expressions for men and four for women. Ultimately, 1,190 data points were available for analysis. Figure 10 shows the facial patterns for individual events related to emotional valence and activation. Visual inspection revealed that AU4 (lowering brows) and AU7 (tightening eyelids) were strongly expressed during negative events (V1-V2). Positive events (V4-V5) evoked AU6 (raising cheeks), AU7, AU10 (raising upper lip), AU12 (pushing corners of the mouth), and AU14 (dimples), which are considered to represent strong smiling expressions. Neutral events (V3) may have relatively low facial movement intensity. Furthermore, for positive events, Figure 10 shows that the higher the activation level, the more frequently the mouth-opening behavior (AU25: lip part, AU26: chin down) was observed. We also confirmed the correlation between the estimated action units and the evaluation dimensions in the peak intensity frames (Table 5). Compared to the correlation between emotional valence and some facial movements, such as AU12 (lip corner down: r = 0.49), the correlations for all combinations of facial movements and evaluation dimensions were relatively low (|r|s < 0.25). We used a hierarchical linear regression model to examine the relationship between emotional valence / activation and AUs. Consistent with our prediction, emotional valence significantly predicted the intensity of AU4 (brow down) negatively (β = -0.11, t = 6.38, p < 0.001) and the intensity of AU12 (lip pulling) positively (β = 0.28, t = 12.64, p < 0.001). Furthermore, activity significantly predicted AU12 strength scores (β = 0.10, t = 9.19, p < 0.001).Furthermore, a post-hoc sensitivity power analysis using the simr package (Green & MacLeod, 2016) indicated that our sample size (i.e., N = 1190) was sufficient to detect all coefficients in a hierarchical linear regression model with a significance level of α = 0.05 and 99% power.
[0084] [Table 5]
[0085] Nonnegative matrix decomposition was applied to reduce dimensionality and extract spatiotemporal features (Lee and Seung, 1999). This approach can identify dynamic facial patterns (e.g., Delis, Panzeri, Pozzo, & Berret, 2014; Perusquia-Hernandez et al.). Factorization rank was determined using the co-ordering coefficient (Brunet, Tamayo, Golub, & Mesirov, 2004) and the dispersion index (Kim and Park, 2007). Information about factorization rank is available at Gakunin RDM (https: / / dmsgrdm.riken.jp:5000 / uphvb / ).
[0086] Figure 11 shows the AU profiles of the top four components. In Figure 11, the color of the numbers represents the contribution of each facial movement to each element score. By visually examining the relative contribution of each AU to the independent components, we interpreted component 1 as the Duchenne marker (AUs 6 and 7), component 2 as blinks and other facial movements (AUs 1, 14, 17, and 45), component 3 as the lowered brow (AU 4), and component 4 as a smile (AUs 6, 10, 12, and 14). These results were consistent with the peak intensity of each facial movement (Figure 10).
[0087] Figure 12 shows how spatial components changed over time for each combination of emotional valence and activation. For component 1 (Duchenne marker), visual inspection revealed larger movements for negative events (V1-V2) and positive events (V4-V5) (e.g., V1A1 and V5A5). This result is consistent with the finding that eye contractions are systematically associated with both negative and positive emotional expressions (Mattson et al.). Component 2 (eye blinks and other facial movements) can be interpreted as relaxation movements associated with the expression of intentional facial manipulation or as noise unrelated to the primary emotional expression. However, this movement increases during the offset duration (frames = 91-120) following the peak duration (frames = 31-90). For component 3 (brow down), negative expressions (V1-V2) induced stronger facial changes than other expressions (V3-V5). Component 4 (smile) occurred more frequently during positive events (V4-V5) than during other events (V1-V3). In summary, facial movements such as blinking, inward brow raising, and chin lifting (i.e., component 2) were characteristic of the offset of intentional facial expressions in naive Japanese participants. More interestingly, smiles corresponded to positive values (component 4), lowered brows corresponded to negative values (component 3), and eye constriction (component 1) corresponded to both values.
[0088] (Consideration) In this study, we developed a new face database that includes expressive annotations, such as individualized emotional events, evaluation checks, and free-text labels with multi-angle and depth information. The results (Tables 3 and 4) show that there is little agreement between the words for each event, suggesting a large degree of variability in the emotional events in the database. This database, which includes a variety of events and individual evaluations, can be validated for academic purposes. For example, researchers can use the data as a starting point to investigate questions such as the typical elements of events labeled as anger and the evaluation components that comprise them.
[0089] Analysis of frontal facial expressions revealed facial movements associated with pleasant and unpleasant emotions. For example, lowering the eyebrows was associated with negative emotional valence, while pulling in the lip corners was associated with positive emotional valence. These results are consistent with previous findings investigating the coherence between emotional valence and facial myoelectric activity (e.g., Greenwald et al., 1989). In addition to the data presented here, we are currently conducting manual facial movement coding by certified FACS coders and plan to release the annotation data in the future. In recent years, efforts to extract facial movements have become more active and developed amid the debate over emotions (e.g., Cordaro, Fridlund, Keltner, Russell, & Scarantino, 2015) (see Cohn et al., 2019). The release of a database containing manual FACS annotations, including depth information, could pave the way for further development in affective computing research.
[0090] While this study provides a new facial database of emotions, it does have certain limitations. First, the number of participants was small, given the diversity of facial movements and emotional events. In future research using environments like those shown in Figure 9, we plan to further develop databases for younger and older adults, and to expand the database to cultures and ethnicities other than Japanese. Second, we did not investigate how depth and infrared information can be utilized. Compared to RGB images, this information is less affected by lighting. This database will serve as an important foundation for developing a facial movement sensing system that is robust to room conditions. We will use these databases to provide an internal state estimation algorithm via an API in combination with devices such as smartphones. Furthermore, we hope to utilize this technology to develop solutions for quickly and appropriately communicating reports to people who are unable to communicate.
[0091] This database, which includes the expresser's events, labels, and the strength of evaluation checks, has been made available to the academic community as the RIKEN Facial Expression Database. The notable features of this database are: (a) the availability of multiple theoretical perspectives on emotions (emotional valence / activation, evaluation dimensions, and free emotion labels), (b) the rich variety of events present for each individual, and (c) the abundance of feature information due to the images captured from 10 multi-angle and depth cameras.
[0092] (Funding) This research was supported by the Telecommunications Advancement Foundation (to SN) and the Japan Science and Technology Agency-Mirai Program (Grant No. JPMJMI20D7 to WS).
[0093] (Conflict of interest) The authors declare that there are no conflicts of interest related to this manuscript.
[0094] (Ethics approval) The Ethics Committee of RIKEN (Protocol number: Wako3 2020-21) approved all experimental procedures and protocols submitted to participants.
[0095] (Participant consent) All participants submitted a consent form at the start of the experiment.
[0096] (Consent to Publication) Participants provided written informed consent for the publication of their aggregated data at the start of the study.
[0097] (Data availability) The RIKEN facial expression database is freely available to the research community (https: / / dmsgrdm.riken.jp:5000 / uphvb / ). To access the database, you must complete an End User License Agreement (EULA).
[0098] (References) Aly, S., Trubanova, A., Abbott, AL, White, SW, & Youssef, AE (2015). VT-KFER: A Kinect-based RGBD+ time dataset for spontaneous and non-spontaneous facial expression recognition. 2015 International Conference on Biometrics (ICB), 90-97. https: / / doi.org / 10.1109 / ICB.2015.7139081 Baltrusaitis, T., Zadeh, A., Lim, YC, & Morency, LP (2018). Openface 2.0: Facial behavior analysis toolkit. 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018), 59-66. https: / / doi.org / 10.1109 / FG.2018.00019 Barrett, LF (2017). How emotions are made: The secret life of the brain. Pan Macmillan, 448. Barrett, L. F., Adolphs, R., Marsella, S., Martinez, A. M., & Pollak, S. D. (2019). Emotional expressions reconsidered: Challenges to inferring emotion from human facial movements. Psychological Science in the Public Interest, 20(1), 1-68. https: / / doi.org / 10.1177 / 1529100619832930 Barrett, L. F., Mesquita, B., & Gendron, M. (2011). Context in emotion perception. Current Directions in Psychological Science, 20(5), 286-290. https: / / doi.org / 10.1177 / 0963721411422522 Barrett, L. F., & Russell, J. A. (1999). The structure of current affect: Controversies and emerging consensus. Current Directions in Psychological Science, 8(1), 10-14. https: / / doi.org / 10.1111 / 1467-8721.00003 Boccignone, G., Conte, D., Cuculo, V., & Lanzarotti, R. (2017). AMHUSE: A multimodal dataset for HUmour SEnsing. Proceedings of the 19th ACM International Conference on Multimodal Interaction, 438-445. https: / / doi.org / 10.1145 / 3136755.3136806 Bradley, M. M., & Lang, P. J. (2000). Affective reactions to acoustic stimuli. Psychophysiology, 37, 204215. https: / / doi.org / 10.1111 / 1469-8986.3720204 Brunet, J. P., Tamayo, P., Golub, T. R., & Mesirov, J. P. (2004). Metagenes and molecular pattern discovery using matrix factorization. Proceedings of the National Academy of Sciences, 101(12), 4164-4169. https: / / doi.org / 10.1073 / pnas.0308531101 Burt, A. L., & Crewther, D. P. (2020). The 4D space-time dimensions of facial perception. Frontiers in Psychology, 11, 1842. https: / / doi.org / 10.3389 / fpsyg.2020.01842 Calvo, M. G., & Lundqvist, D. (2008). Facial expressions of emotion (KDEF): Identification under different display-duration conditions. Behavior Research Methods, 40(1), 109-115. https: / / doi.org / 10.3758 / BRM.40.1.109 Chelnokova, O., & Laeng, B. (2011). Three-dimensional information in face recognition: An eye-tracking study. Journal of Vision, 11(13), 27. https: / / doi.org / 10.1167 / 11.13.27 Chen, Z., & Whitney, D. (2019). Tracking the affective state of unseen persons. Proceedings of the National Academy of Sciences, 116(15), 7559-7564. https: / / doi.org / 10.1073 / pnas.1812250116 Cheng, S., Kotsia, I., Pantic, M., & Zafeiriou, S. (2018). 4dfab: A large scale 4d database for facial expression analysis and biometric applications. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5117-5126. https: / / doi.org / 10.1109 / CVPR.2018.00537 Chung, K. M, Kim, S.J., Jung, W. H., & Kim, V. Y. (2019). Development and validation of the Yonsei Face Database (Yface DB). Frontiers in Psychology, 10, 2626. https: / / doi.org / 10.3389 / fpsyg.2019.02626 Cohn, J. F., Ertugrul, I. O., Chu, W. S., Girard, J. M., Jeni, L. A., & Hammal, Z. (2019). Affective facial computing: Generalizability across domains. Multimodal Behavior Analysis in the Wild, 407-441. https: / / doi.org / 10.1016 / B978-0-12-814601-9.00026-2 Cordaro,D., Fridlund, A. J., Keltner, D., Russell, J. A., & Scarantino, A. (2015). Debate: Keltner and Cordaro vs. Fridlund vs. Russell. Retrieved from http: / / emotionresearcher.com / the-great-expressions-debate / Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., & Black, M. J. (2019). Capture, learning, and synthesis of 3D speaking styles. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 10101-10111. https: / / doi.org / 10.48550 / arXiv.1905.03079 Dawel, A., Miller, E. J., Horsburgh, A., & Ford, P. (2021). A systematic survey of face stimuli used in psychological research 2000-2020. Behavior Research Methods, 1-13. https: / / doi.org / 10.3758 / s13428-021-01705-3 Delis, I., Panzeri, S., Pozzo, T., & Berret, B. (2014). A unifying model of concurrent spatial and temporal modularity in muscle activity. Journal of Neurophysiology, 111(3), 675-693. https: / / doi.org / 10.1152 / jn.00245.2013 Ebner, N., Riediger, M., & Lindenberger, U. (2010). FACES-A database of facial expressions in young, middle-aged, and older women and men: Development and validation. Behavior Research Methods, 42, 351-362. https: / / doi.org / 10.3758 / BRM.42.1.351 Ekundayo, O. S., & Viriri, S. (2021). Facial expression recognition: A review of trends and techniques. IEEE Access, 9, 136944-136973. https: / / doi.org / 10.1109 / ACCESS.2021.3113464 Feinerer I., & Hornik, K. (2022). tm: Text Mining Package. R package version 0.7-10. Retrieved from https: / / CRAN.R-project.org / package=tm Feinerer, I., Hornik, K., & Meyer, D. (2008). Text mining infrastructure in R. Journal of Statistical Software, 25(5), 1-54. https: / / doi.org / 10.18637 / jss.v025.i05 Fehr, B., & Russell, J. A. (1984). Concept of emotion viewed from a prototype perspective. Journal of Experimental Psychology: General, 113(3), 464-486. https: / / doi.org / 10.1037 / 0096-3445.113.3.464 Fernandes-Magalhaes, R., Carpio, A., Ferrera, D., Van Ryckeghem, D., Pelaez, I., Barjola, P., De Lahoz , M. E., Martin-Buro, M. C., Hinojosa, J. A., Van Damme, S., Carretie, L., & Mercado, F. (2022). Pain E-motion Faces Database (PEMF): Pain-related micro-clips for emotion research. Behavior Research Methods, 1-14. https: / / doi.org / 10.3758 / s13428-022-01992-4 Fujimura, T., & Umemura, H. (2018). Development and validation of a facial expression database based on the dimensional and categorical model of emotions. Cognition and Emotion, 32(8), 1663-1670. https: / / doi.org / 10.1080 / 02699931.2017.1419936 Girard, J. M., Chu, W. S., Jeni, L. A., & Cohn, J. F. (2017). Sayette group formation task (gft) spontaneous facial expression database. 2017 12th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2017), 581-588. https: / / doi.org / 10.1109 / FG.2017.144 Green, P., & MacLeod, C. J. (2016). SIMR: An R package for power analysis of generalized linear mixed models by simulation. Methods in Ecology and Evolution, 7(4), 493-498. https: / / doi.org / 10.1111 / 2041-210X.12504 Greenwald, M. K., Cook, E. W., & Lang, P. J. (1989). Affective judgment and psychophysiological response: Dimensional covariation in the evaluation of pictorial stimuli. Journal of Psychophysiology, 3, 51-64. Griffiths, P. E. (1997). What emotions really are: The problem of Psychological categories. University of Chicago Press, 293. https: / / doi.org / 10.7208 / chicago / 9780226308760.001.0001 Haidt, J., & Keltner, D. (1999). Culture and facial expression: Open-ended methods find more expressions and a gradient of recognition. Cognition & Emotion, 13(3), 225-266. https: / / doi.org / 10.1080 / 026999399379267 Holland, C. A. C., Ebner, N. C., Lin, T., & Samanez-Larkin, G. R. (2019). Emotion identification across adulthood using the Dynamic FACES database of emotional expressions in younger, middle aged, and older adults. Cognition and Emotion, 33, 245-257. https: / / doi.org / 10.1080 / 02699931.2018.1445981 Kim, H., & Park, H. (2007). Sparse non-negative matrix factorizations via alternating non-negativity-constrained least squares for microarray data analysis. Bioinformatics, 23(12), 1495-1502. https: / / doi.org / 10.1093 / bioinformatics / btm134 Kossaifi, J., Tzimiropoulos, G., Todorovic, S., & Pantic, M. (2017). AFEW-VA database for valence and arousal estimation in-the-wild. Image and Vision Computing, 65, 23-36. https: / / doi.org / 10.1016 / j.imavis.2017.02.001 Krumhuber, E. G., Skora, L., Kuster, D., & Fou, L. (2017). A review of dynamic datasets for facial expression research. Emotion Review, 9(3), 280-292. https: / / doi.org / 10.1177 / 1754073916670022 Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. B. (2017). lmerTest Package: Tests in Linear Mixed Effects Models. Journal of Statistical Software, 82(13), 1-26 https: / / doi.org / 10.18637 / jss.v082.i13 Kuster, D., Baker, M., & Krumhuber, E. G. (2021). PDSTD-The Portsmouth Dynamic Spontaneous Tears Database. Behavior Research Methods, 1-15. https: / / doi.org / 10.3758 / s13428-021-01752-w Lang, P. J. (1994). The varieties of emotional experience: a meditation on James-Lange theory. Psychological Review, 101, 211-221. https: / / doi.org / 10.1037 / 0033-295X.101.2.211 Lang, P. J., Greenwald, M. K., Bradley, M. M., & Hamm, A. O. (1993). Looking at pictures: Affective, facial, visceral, and behavioral reactions. Psychophysiology, 30(3), 261-273. https: / / doi.org / 10.1111 / j.1469-8986.1993.tb03352.x Langner, O., Dotsch, R., Bijlstra, G., Wigboldus, D. H. J., Hawk, S. T., & van Knippenberg, A. (2010). Presentation and validation of the Radboud Faces Database. Cognition & Emotion, 24, 1377-1388. https: / / doi.org / 10.1080 / 02699930903485076 Larsen, J. T., Norris, C. J., & Cacioppo, J. T. (2003). Effects of positive and negative affect on electromyographic activity over zygomaticus major and corrugator supercilii. Psychophysiology, 40, 776-785. https: / / doi.org / 10.1111 / 1469-8986.00078 Le Mau, T., Hoemann, K., Lyons, S. H., Fugate, J., Brown, E. N., Gendron, M., & Barrett, L. F. (2021). Professional actors demonstrate variability, not stereotypical expressions, when portraying emotional states in photographs. Nature Communications, 12(1), 1-13. https: / / doi.org / 10.1038 / s41467-021-25352-6 Lee, D. D., & Seung, H. S. (1999). Learning the parts of objects by non-negative matrix factorization. Nature, 401(6755), 788-791. https: / / doi.org / 10.1038 / 44565 Li, X., Cheng, S., Li, Y., Behzad, M., Shen, J., Zafeiriou, S., Pantic, M., & Zhao, G. (2022). 4DME: A spontaneous 4d micro-expression dataset with multimodalities. IEEE Transactions on Affective Computing, 1-18. https: / / doi.org / 10.1109 / TAFFC.2022.3182342 Li, S. & W. Deng, W. (2022). Deep facial expression recognition: A survey. IEEE Transactions on Affective Computing, 13(3), 1195-1215. https: / / doi.org / 10.1109 / TAFFC.2020.2981446 Lloyd, E. P., Deska, J. C., Hugenberg, K., McConnell, A. R., Humphrey, B. T., & Kunstman, J. W. (2019). Miami University deception detection database. Behavior Research Methods, 51(1), 429-439. https: / / doi.org / 10.3758 / s13428-018-1061-4 LoBue, V., & Thrasher, C. (2015). The child affective facial expression (CAFE) set: Validity and reliability from untrained adults. Frontiers in Psychology, 5, 1532. https: / / doi.org / 10.3389 / fpsyg.2014.01532 Lucey, P., Cohn, J. F., Prkachin, K. M., Solomon, P. E., & Matthews, I. (2011). Painful data: The UNBC-McMaster shoulder pain expression archive database. 2011 IEEE International Conference on Automatic Face & Gesture Recognition (FG), 57-64. https: / / doi.org / 10.1109 / FG.2011.5771462 Mattson, W. I., Cohn, J. F., Mahoor, M. H., Gangi, D. N., & Messinger, D. S. (2013). Darwin’s Duchenne: Eye constriction during infant joy and distress. PLoS One, 8(11), e80161. https: / / doi.org / 10.1371 / journal.pone.0080161 Matuszewski, B. J., Quan, W., Shark, L. K., McLoughlin, A. S., Lightbody, C. E., Emsley, H. C., & Watkins, C. L. (2012). Hi4D-ADSIP 3-D dynamic facial articulation database. Image and Vision Computing, 30(10), 713-727. https: / / doi.org / 10.1016 / j.imavis.2012.02.002 Mavadati, S. M., Mahoor, M. H., Bartlett, K., Trinh, P., & Cohn, J. F. (2013). Disfa: A spontaneous facial action intensity database. IEEE Transactions on Affective Computing, 4(2), 151-160. https: / / doi.org / 10.1109 / T-AFFC.2013.4 Mihajlovic, M., Bansal, A., Zollhoefer, M., Tang, S., & Saito, S. (2022). KeypointNeRF: Generalizing image-based volumetric avatars using relative spatial encoding of keypoints. European Conference on Computer Vision, 179-197. https: / / doi.org / 10.1007 / 978-3-031-19784-0_11 Mollahosseini, A., Hasani, B., & Mahoor, M. H. (2017). Affectnet: A database for facial expression, valence, and arousal computing in the wild. IEEE Transactions on Affective Computing, 10(1), 18-31. https: / / doi.org / 10.1109 / TAFFC.2017.2740923 Namba, S., Nakamura, K., & Watanabe, K. (2022). The spatio-temporal features of perceived-as-genuine and deliberate expressions. PLoS One, 17(7), e0271047. https: / / doi.org / 10.1371 / journal.pone.0271047 Namba, S., Sato, W., Osumi, M., & Shimokawa, K. (2021). Assessing automated facial action unit detection systems for analyzing cross-domain facial expression databases. Sensors, 21(12), 4222. https: / / doi.org / 10.3390 / s21124222 Namba, S., Sato, W., & Yoshikawa, S. (2021). Viewpoint Robustness of Automated Facial Action Unit Detection Systems. Applied Sciences, 11(23), 11171. https: / / doi.org / 10.3390 / app112311171 Ong, D. C., Wu, Z., Tan, Z. X., Reddan, M., Kahhale, I., Mattek, A., & Zaki, J. (2019). Modeling emotion in complex stories: the Stanford Emotional Narratives Dataset. IEEE Transactions on Affective Computing, 12(3), 579-594. https: / / doi.org / 10.1109 / TAFFC.2019.2955949 Perusquia-Hernandez, M., Dollack, F., Tan, C. K., Namba, S., Ayabe-Kanamura, S., & Suzuki, K. (2021, December). Smile action unit detection from distal wearable electromyography and computer vision. 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021), 1-8. https: / / doi.org / 10.1109 / FG52635.2021.9667047 R Core Team. (2021). R: A language and environment for statistical computing. Vienna, Austria: R Foundation for Statistical Computing. Retrieved from http: / / www.R-project.org / . Revelle, W. (2022) psych: Procedures for Personality and Psychological Research. Retrieved from https: / / CRAN.R-project.org / package=psych. Russell, J. A., Weiss, A., & Mendelsohn, G. A. (1989). Affect grid: a single-item scale of pleasure and arousal. Journal of Personality and Social Psychology, 57(3), 493-502. https: / / doi.org / 10.1037 / 0022-3514.57.3.493 Sato, W., Fujimura, T., Kochiyama, T., & Suzuki, N. (2013). Relationships among facial mimicry, emotional experience, and emotion recognition. PLoS One, 8, e57889. https: / / doi.org / 10.1371 / journal.pone.0057889 Sato, W., Kochiyama, T., & Yoshikawa, S. (2020). Physiological correlates of subjective emotional valence and arousal dynamics while viewing films. Biological Psychology, 157, 107974. https: / / doi.org / 10.1016 / j.biopsycho.2020.107974 Sato, W., Murata, K., Uraoka, Y., Shibata, K., Yoshikawa, S., & Furuta, M. (2021). Emotional valence sensing using a wearable facial EMG device. Scientific Reports, 11, 5757. https: / / doi.org / 10.1038 / s41598-021-85163-z Scarantino, A. (2012). How to define emotions scientifically. Emotion review, 4(4), 358-368. https: / / doi.org / 10.1177 / 1754073912445810 Schauberger P., & Walker, A. (2022). openxlsx: Read, Write and Edit xlsx Files. R package version 4.2.5.1. Retrieved from https: / / CRAN.R-project.org / package=openxlsx Scherer, K. R. (1997) Profiles of emotion-antecedent appraisal: Testing theoretical predictions across cultures. Cognition and Emotion, 11(2), 113-150. https: / / doi.org / 10.1080 / 026999397379962 Scherer, K. R. (2001). Appraisal considered as a process of multi-level sequential checking. In K. R. Scherer, A. Schorr, & T. Johnstone (Eds.). Appraisal processes in emotion: Theory, Methods, Research (pp. 92-120). New York and Oxford: Oxford University Press. Scherer, K. R., Dieckmann, A., Unfried, M., Ellgring, H., & Mortillaro, M. (2021). Investigating appraisal-driven facial expression and inference in emotion communication. Emotion, 21(1), 73-95. https: / / doi.org / 10.1037 / emo0000693 Scherer, K. R., Mortillaro, M., Rotondi, I., Sergi, I., & Trznadel, S. (2018). Appraisal-driven facial actions as building blocks for emotion inference. Journal of Personality and Social Psychology, 114(3), 358-379. https: / / doi.org / 10.1037 / pspa0000107 Scherer, A. Schorr, & T. Johnstone (Eds.), Appraisal processes in emotion: Theory, methods, research. New York: Oxford University Press, 92-120. Senturk, Y. D., Tavacioglu, E. E., Duymaz, I., Sayim, B., & Alp, N. (2022). The Sabancl University Dynamic Face Database (SUDFace): Development and validation of an audiovisual stimulus set of recited and free speeches with neutral facial expressions. Behavior Research Methods, 1-22. https: / / doi.org / 10.3758 / s13428-022-01951-z Seuss, D., Dieckmann, A., Hassan, T., Garbas, J. U., Ellgring, J. H., Mortillaro, M., & Scherer, K. (2019, September). Emotion expression from different angles: A video database for facial expressions of actors shot by a camera array. 2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII), 35-41. https: / / doi.org / 10.1109 / ACII.2019.8925458 Sneddon, I., McRorie, M., McKeown, G., & Hanratty, J. (2011). The Belfast induced natural emotion database. IEEE Transactions on Affective Computing, 3(1), 32-41. https: / / doi.org / 10.1109 / T-AFFC.2011.26 Tan, J. W., Walter, S., Scheck, A., David, H., Hoffmann, H., Kessler, H., & Traue, H. C. (2012). Repeatability of facial electromyography (EMG) activity over corrugator supercilii and zygomaticus major on differentiating various emotions. Journal of Ambient Intelligence and Humanized Computing, 3, 3-10. https: / / doi.org / 10.1007 / s12652-011-0084-9 Tottenham, N., Tanaka, J. W., Leon, A. C., McCarry, T., Nurse, M., Hare, T. A., Marcus, D. J., Westerlund, A., Casey, B. J., & Nelson, C. (2009). The NimStim set of facial expressions: Judgments from untrained research participants. Psychiatry Research, 168(3), 242-249. https: / / doi.org / 10.1016 / j.psychres.2008.05.006 Tsuyuzaki K, Ishii M and Nikaido I. (2022). nnTensor: Non-Negative Tensor Decomposition. R package version 1.1.9. Retrieved from https: / / github.com / rikenbit / nnTensor Ueda, Y., Nunoi, M., & Yoshikawa, S. (2019). Development and validation of the Kokoro Research Center (KRC) facial expression database. Psychologia, 61(4), 221-240. https: / / doi.org / 10.2117 / psysoc.2019-A009 Van Der Schalk, J., Hawk, S. T., Fischer, A. H., & Doosje, B. (2011). Moving faces, looking places: Validation of the Amsterdam Dynamic Facial Expression Set (ADFES). Emotion, 11(4), 907-920. https: / / doi.org / 10.1037 / a0023853 Wickham H, Averick M, Bryan J, Chang W, McGowan LD, Francois R, Grolemund G, Hayes A, Henry L, Hester J, Kuhn M, Pedersen TL, Miller E, Bache SM, Muller K, Ooms J, Robinson D, Seidel DP, Spinu V, Takahashi K, Vaughan D, Wilke C, Woo K, Yutani H (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. https: / / doi.org / 10.21105 / joss.01686 Workman, C. I., & Chatterjee, A. (2021). The Face Image Meta-Database (fIMDb) & ChatLab Facial Anomaly Database (CFAD): Tools for research on face perception and social stigma. Methods in Psychology, 5, 100063. https: / / doi.org / 10.1016 / j.metip.2021.100063 Zhang, X., Yin, L., Cohn, J. F., Canavan, S., Reale, M., Horowitz, A., Liu, P., & Girard, J. M. (2014). Bp4d-spontaneous: A high-resolution spontaneous 3d dynamic facial expression database. Image and Vision Computing, 32(10), 692-706. https: / / doi.org / 10.1016 / j.imavis.2014.06.002
[0099] The present invention is not limited to the above-described embodiments, and various applications and modifications can be made within the scope of the present invention. For example, in the feeling estimation device 10A of the second embodiment, the imaging unit 6 has been described as capturing an image of the subject's face for a predetermined period of time to obtain video data capturing the subject's face. However, the imaging unit 6 may continuously capture an image of the subject's face to obtain multiple still image data. In this case, the estimation unit 3 may extract facial movements from the multiple still image data of the subject's face captured by the imaging unit 6. Similarly, in the feeling estimation device 10 of the first embodiment, the imaging data acquisition unit 2 may acquire multiple still image data capturing the subject's face.
[0100] When extracting the facial movements of the subject, the estimation unit 3 can also create a 3D model of the subject's face and extract the facial movements from the created 3D model. The 3D model of the face can be created from video data of the subject's face or still image data of multiple faces. By extracting facial movements from the 3D model of the face, it becomes possible to extract facial movements even from image data of the subject's profile, for example, and to extract facial movements more accurately.
[0101] The subject whose emotion is to be estimated may be wearing glasses or a mask. In order to accurately estimate the emotion of a subject wearing glasses, it is preferable that the emotion data stored in database 1 includes emotion data of a person wearing glasses. Similarly, in order to accurately estimate the emotion of a subject wearing a mask, it is preferable that the emotion data stored in database 1 includes emotion data of a person wearing a mask.
[0102] The emotion estimation device in this application is as follows.
[0103] <1> A database storing emotion data that defines multiple types of human emotions based on emotional valence, which is an index representing the quality of an emotion, and activation, which is an index representing the strength of an emotion; an imaging data acquisition unit that acquires imaging data of the face of a subject; an estimation unit that extracts facial movements from the image capture data of the subject's face acquired by the image capture data acquisition unit, and estimates the subject's emotion as an emotion defined by the valence and the activity level based on the extracted facial movements and the emotion data stored in the database; An emotion estimation device comprising:
[0104] <2> Further, an imaging unit is provided to capture an image of the face of the subject, The imaging data acquisition unit acquires imaging data of the face of the subject captured by the imaging unit. <1> The emotion estimation device according to claim 1.
[0105] <3> The emotion data is characterized in that the facial movements of a plurality of subjects are captured as video images when they are informed of an event that evokes an emotion corresponding to the emotional valence and the activation level, and the facial movements are extracted from the captured video images. <1> or <2> The emotion estimation device according to claim 1.
[0106] <4> The emotion data is facial movement data extracted from multiple videos captured from different angles showing changes in the faces of multiple subjects. <3> The emotion estimation device according to claim 1.
[0107] <5> The facial movements include at least one of "inner eyebrows rising," "outer eyebrows rising," "eyebrows lowering," "upper eyelids lifting," "cheeks lifting," "eyelid tightening," "nose wrinkling," "upper lip lifting," "lip corners lifting," "dimples," "lip corners lowering," "chin lifting," "lips stretching sideways," "lips closing," "lips pressing together," "chin lowering," and "blinking." <3> or <4> The emotion estimation device according to claim 1.
[0108] <6> The emotion data is characterized in that it defines human emotions by five levels of emotional valence and five levels of activity. <1> ~ <5> 10. The emotion estimation device according to claim 9, wherein
[0109] <7> The emotional valence on the 0.5 scale is "very unpleasant," "unpleasant," "neutral," "pleasant," and "very pleasant." The five levels of activity are characterized as "very low arousal," "low arousal," "medium arousal," "high arousal," and "very high arousal." <6> The emotion estimation device according to claim 1.
[0110] <8> The estimation unit is characterized in that it can estimate the emotion of the subject as an emotion defined by the valence having more stages than the valence stages that define the emotion data and the activity level having more stages than the activity level that defines the emotion data. <1> ~ <7> 10. The emotion estimation device according to claim 9, wherein
[0111] <9> The emotion data is facial movement data extracted from a video of facial changes of a plurality of subjects who are informed of an event that evokes an emotion corresponding to the emotional valence and the activation level, and The imaging unit captures the image of the subject's face for at least the time it takes to capture the image of the subject's face to acquire the emotion data. <2> The emotion estimation device according to claim 1.
[0112] <10> The number of the imaging units is two or more. <2> or <9> The emotion estimation device according to claim 1. [Explanation of symbols]
[0113] 1 Database 2. Image data acquisition unit 3 Estimation part 4. Imaging device 5 Light 6. Imaging unit 10, 10A Emotion estimation device
Claims
1. a database storing emotion data defining multiple types of human emotions based on emotional valence, which is an index representing the quality of an emotion, and activity, which is an index representing the strength of an emotion; an imaging data acquisition unit that acquires imaging data of the face of a subject; an estimation unit that extracts facial movements from the image capture data of the subject's face acquired by the image capture data acquisition unit, and estimates the subject's emotion as an emotion defined by the valence and the activity level based on the extracted facial movements and the emotion data stored in the database; An emotion estimation device comprising:
2. further comprising an imaging unit that images a face of the subject, The feeling estimation device according to claim 1 , wherein the imaging data acquisition unit acquires imaging data of the face of the subject captured by the imaging unit.
3. 2. The emotion estimation device according to claim 1, wherein the emotion data is facial movement data extracted from a video of changes in the faces of a plurality of subjects who are informed of an event that evokes an emotion corresponding to the valence and the activity level.
4. 4. The emotion estimation device according to claim 3, wherein the emotion data is facial movement data extracted from a plurality of videos captured from different angles showing changes in the faces of the plurality of subjects.
5. 4. The feeling estimation device according to claim 3, wherein the facial movements include at least one of “inner eyebrows rising,” “outer eyebrows rising,” “eyebrows lowering,” “upper eyelids lifting,” “cheeks lifting,” “eyelids tightening,” “nose wrinkling,” “upper lip lifting,” “corners of the lip lifting,” “dimples forming,” “corners of the lip lowering,” “chin lifting,” “lips stretching sideways,” “lips closing,” “lips pressing together,” “chin lowering,” and “blinking.”
6. 2. The emotion estimation device according to claim 1, wherein the emotion data is data that defines human emotions using five levels of valence and five levels of activity.
7. The five levels of emotional valence are "very unpleasant," "unpleasant," "neutral," "pleasant," and "very pleasant." 7. The emotion estimation device according to claim 6, wherein the five levels of activity are "very low arousal," "low arousal," "medium arousal," "high arousal," and "very high arousal."
8. 2. The emotion estimation device according to claim 1, wherein the estimation unit is capable of estimating the emotion of the subject as an emotion defined by a greater number of valence stages than the valence stages defining the emotion data and a greater number of activity stages than the activity stages defining the emotion data.
9. the emotion data is facial movement data extracted from a video captured by capturing facial changes of a plurality of subjects who are informed of an event that evokes an emotion corresponding to the emotional valence and the activity level, The emotion estimation device according to claim 2 , wherein the imaging unit captures the subject's face for at least the same time as capturing the subject's face to acquire the emotion data.
10. The emotion estimation device according to claim 2 or 9, wherein the number of the imaging units is two or more.
Citation Information
Patent Citations
Facial expression determination system, program and facial expression determination method
JP2020057111A
Emotion estimation device, emotion estimation method, and program
JP2022189703A