Emotion estimation device, emotion estimation method, emotion estimation program and emotion estimation system
The emotion estimation device addresses the time-consuming calibration requirement by using singularity information to set thresholds, enabling early and efficient emotion estimation without recalibration.
Patent Information
- Application Number
- JP2024080907
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Conventional emotion estimation techniques require calibration and setting of threshold values before each processing session, which is time-consuming and delays the start of emotion estimation.
An emotion estimation device that extracts singularity information from time-series data during calibration and uses it to set adjusted thresholds, allowing emotion estimation to start without recalibration.
Enables early initiation of emotion estimation processing by setting thresholds based on pre-extracted singularity information, eliminating the need for recalibration.
Smart Images

Figure 2025174495000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an emotion estimation device, an emotion estimation method, an emotion estimation program, and an emotion estimation system. [Background technology]
[0002] Conventionally, a technique has been disclosed in which an emotional index value is calculated based on biometric data measured from a subject, and the subject's emotion is estimated from the calculated index value (see, for example, Patent Document 1). Furthermore, with this type of technique, the way the biometric data is expressed varies depending on individual differences and differences in the surrounding environment, which causes the index value to vary accordingly, and therefore calibration is required to adjust the threshold used for emotion estimation in accordance with the variation. For example, Patent Document 1 discloses a technique in which the subject is asked to perform a task that can identify changes in emotion, and calibration is performed based on the biometric data during the task execution. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2023-071507 Summary of the Invention [Problem to be solved by the invention]
[0004] However, with conventional techniques, it is necessary to perform calibration and set a threshold value each time before starting emotion estimation processing, which means that it takes time each time before starting emotion estimation processing.
[0005] The present application has been made in view of the above, and aims to provide an emotion estimation device, an emotion estimation method, an emotion estimation program, and an emotion estimation system that are capable of starting emotion estimation processing early. [Means for solving the problem]
[0006] The emotion estimation device according to the present application is an emotion estimation device that performs emotion estimation from emotion index values related to a biological state of an emotion estimation subject based on estimation conditions adjusted by calibration, and includes a controller. During calibration, the controller extracts and stores singularity information that is a feature of time-series data of the detected emotion index values, and during emotion estimation, extracts correspondence information that corresponds to the singularity information in the time-series data of the detected emotion index values, and performs emotion estimation based on the difference between the singularity information and the correspondence information. [Effects of the Invention]
[0007] In the present disclosure, singularity information of emotion index values is extracted in advance, and from the next time onwards, emotion estimation processing is performed by setting values to be adjusted, such as thresholds, based on the singularity information without performing calibration. Therefore, according to the present disclosure, thresholds and the like can be set without performing calibration from the next time onwards, allowing the emotion estimation processing to start earlier. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing the relationship between a vehicle and a person assumed in the embodiment. [Figure 2] FIG. 2 is a diagram showing a schematic block diagram of the in-vehicle system. [Figure 3] FIG. 3 is a diagram showing an example of a two-dimensional model related to emotion estimation. [Figure 4] FIG. 4 is a diagram for explaining the neutral region. [Figure 5] FIG. 5 is a diagram illustrating an example of the configuration of the feeling estimation device according to the embodiment. [Figure 6] FIG. 6 is a diagram showing an example of the psychological plane table. [Figure 7] FIG. 7 is a diagram showing an example of the psychological plane table. [Figure 8] FIG. 8 is a diagram illustrating an example of the special task table. [Figure 9] FIG. 9 is a diagram illustrating an example of the neutral region table. [Figure 10] FIG. 10 is a diagram showing a schematic diagram of changes over time in the index values of physiological responses when the first task and the second task are being performed. [Figure 11] FIG. 11 is a diagram showing a schematic diagram of changes in physiological responses over time. [Figure 12] FIG. 12 is a schematic diagram showing the change over time of the estimated upper limit value of the neutral region. [Figure 13] FIG. 13 is a diagram illustrating an example of the singularity information. [Figure 14] FIG. 14 is a flowchart showing the procedure of the initial calibration process. [Figure 15] FIG. 15 is a flowchart showing the procedure of the emotion estimation process from the next time onwards. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of an emotion estimation device, an emotion estimation method, an emotion estimation program, and an emotion estimation system will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the embodiments described below.
[0010] <<Embodiment>> An embodiment of the present invention will be described. A technology for estimating a person's emotion will be described in the embodiment. FIG. 1 is a diagram showing the relationship between a vehicle V1 and a person U1 assumed in the embodiment. The person U1 is a target of emotion estimation, and will be referred to as the target person U1 hereinafter. The vehicle V1 may be any type of vehicle. Herein, the vehicle V1 is assumed to be an automobile or the like traveling on a road. The target person U1 is an occupant of the vehicle V1. In the present embodiment, it is assumed that the target person U1 is the driver of the vehicle V1. Hereinafter, when simply referring to the driver, this refers to the driver of the vehicle V1 (hence the target person U1). However, the target person U1 may also be an occupant other than the driver (i.e., a passenger in the vehicle V1). In line with the person U1 being referred to as the target person U1, the vehicle V1 may also be referred to as the target vehicle V1. An in-vehicle system SYS is installed in the vehicle V1. The in-vehicle system SYS is an example of an emotion estimation system.
[0011] In this disclosure, an example is given in which the target U1 of emotion estimation is a driver, but the target U1 is not limited to a driver. For example, the target U1 may be an e-sports player, a patient at a medical institution, a student at an educational institution, or a viewer of content such as video or music.
[0012] 2 is a diagram showing a schematic block diagram of the in-vehicle system SYS. The in-vehicle system SYS includes an emotion estimation device 10, a vehicle control device 20, an actuator unit 30, a vehicle sensor unit 40, a biosensor 60, an exterior camera 71, an interior camera 72, a display unit 73, a speaker 74, and a microphone 75. The components of the in-vehicle system SYS can transmit and receive any signals and information to and from each other through an in-vehicle network formed in the vehicle V1. The in-vehicle network includes, for example, a CAN (Controller Area Network) and an AVCLAN (Audio Visual Communication Local Area Network).
[0013] The emotion estimation device 10 has an emotion estimation model 131 and estimates the emotion of the subject U1 using the emotion estimation model 131. The vehicle control device 20 controls the driving of the vehicle V1 using an actuator unit 30. The actuator unit 30 has various driving components such as a motor that realizes the driving of the vehicle V1. The actuator unit 30 includes an engine and a motor that generate driving force for the vehicle V1, a steering actuator that drives the steering of the vehicle V1, and a brake actuator that drives the brakes of the vehicle V1.
[0014] The vehicle sensor unit 40 has sensors that detect the details of the driving operation of the vehicle V1 by the driver of the vehicle V1 and sensors that detect various states of the vehicle V1. The vehicle sensor unit 40 outputs vehicle sensor information containing these detection results. The vehicle control device 20 realizes driving control of the vehicle V1 by driving and controlling the actuator unit 30 in accordance with the vehicle sensor information. At this time, the vehicle control device 20 can also perform driving control in accordance with the emotion estimation result. The biometric sensor 60 detects biometric information of the subject U1. The emotion of the subject U1 is estimated based on the detection result of the biometric information by the biometric sensor 60. Details of the emotion estimation device 10, the vehicle control device 20, the vehicle sensor unit 40, and the biometric sensor 60 will be described later.
[0015] The exterior camera 71 consists of one or more cameras that capture images of the scene outside the vehicle V1. The exterior camera 71 has a capture area set outside the vehicle V1, and generates an exterior camera image by capturing images of the scene within the capture area. The exterior camera image is an image captured within the capture area of the exterior camera 71. Image information showing the exterior camera image is referred to as exterior image information. The exterior camera 71 captures images at a predetermined frame rate.
[0016] Vehicle V1 is assumed to be traveling in a forward direction. A vehicle located in front of vehicle V1 is referred to as a leading vehicle, and a vehicle located behind vehicle V1 is referred to as a trailing vehicle. Although it depends on the distance between vehicle V1 and the leading vehicle, it is assumed here that the leading vehicle is located within the shooting range of exterior camera 71, and therefore the exterior camera image includes an image of the leading vehicle. Although it also depends on the distance between vehicle V1 and the trailing vehicle, it is assumed here that the trailing vehicle is located within the shooting range of exterior camera 71, and therefore the exterior camera image includes an image of the trailing vehicle. The exterior camera image may be composed of a captured image of the area in front of vehicle V1, a captured image of the area behind vehicle V1, a captured image of the area to the right of vehicle V1, and a captured image of the area to the left of vehicle V1. Alternatively, the exterior camera image may be a wide-angle image containing information from these four captured images. The exterior camera image may have any configuration.
[0017] The in-vehicle camera 72 consists of one or more cameras that capture images of the interior of the vehicle V1. The in-vehicle camera 72 has a capture area set inside the vehicle V1 (i.e., the interior of the vehicle V1), and generates an in-vehicle camera image by capturing images of the capture area. The in-vehicle camera image is an image captured in the capture area by the in-vehicle camera 72. Image information showing the in-vehicle camera image is referred to as in-vehicle image information. The in-vehicle camera 72 captures images at a predetermined frame rate.
[0018] The subject U1 is located within the imaging area of the in-vehicle camera 72, and therefore the in-vehicle camera image includes an image of the subject U1 (particularly a facial image). That is, if the subject U1 is the driver, for example, the in-vehicle camera 72 is installed so as to capture an image of the area around the driver's seat from the front of the vehicle interior. In particular, it is assumed here that the imaging area of the in-vehicle camera 72 includes the face of the subject U1, and therefore the in-vehicle camera image includes a facial image of the subject U1 (a captured image of the face of the subject U1).
[0019] The display unit 73 is a display device having a liquid crystal display panel or the like, and displays any image under the control of the emotion estimation device 10, the vehicle control device 20, or a display control device (not shown). The display unit 73 is installed at an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can see the display content of the display unit 73. A plurality of display units 73 may be installed in the cabin of the vehicle V1. The display unit 73 may be a component of a car navigation system installed in the vehicle V1. The car navigation system may be included in the in-vehicle system SYS.
[0020] The speaker 74 outputs any sound (message, music, etc.) under the control of the emotion estimation device 10, the vehicle control device 20, or an audio device (not shown). The speaker 74 is installed at an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can hear the output sound from the speaker 74. A plurality of speakers 74 may be installed in the cabin of the vehicle V1.
[0021] The microphone 75 converts ambient sounds around the microphone 75 into an electrical audio signal. The audio signal obtained by the conversion of the microphone 75 is referred to as a microphone signal. In the in-vehicle system SYS, the microphone signal is transmitted to the emotion estimation device 10 and the vehicle control device 20. The microphone 75 is installed at an appropriate location in the cabin of the vehicle V1 so that the signal components of the speech sounds of each occupant of the vehicle V1 are included in the microphone signal (i.e., so that the speech sounds are included in the content picked up by the microphone 75). Multiple microphones 75 may be installed in the cabin of the vehicle V1. Hereinafter, the speech sounds picked up by the microphone 75 are assumed to be the speech sounds of the target person U1.
[0022] The emotion estimation device 10 estimates the emotion of the subject U1 based on the subject U1's biometric information. A biosensor 60 is attached to the subject U1 to acquire the subject U1's biometric information. The biosensor 60 includes at least an electroencephalogram (EEG) sensor and a heartbeat sensor, and the EEG sensor and heartbeat sensor are attached to the subject U1. The EEG sensor detects the subject U1's brain waves and outputs EEG data indicating the results of the brainwave detection. That is, the EEG data includes information on the detected EEG. The heartbeat sensor detects the subject U1's heartbeat and outputs heartbeat data indicating the results of the heartbeat detection. That is, the heartbeat data includes information on the detected heartbeat. For example, a headgear-type EEG sensor is used as the EEG sensor. For example, a chest belt-type electrocardiogram heartbeat sensor is used as the heartbeat sensor. An optical heartbeat sensor can also be used as the heartbeat sensor.
[0023] The biosensor 60 further includes a sweat sensor and a body temperature sensor, which are attached to the subject U1. The sweat sensor detects the amount of sweat (amount of sweat per predetermined time) of the subject U1 and outputs sweat data indicating the detected amount of sweat. The body temperature sensor detects the body temperature of the subject U1 and outputs body temperature data indicating the detected body temperature.
[0024] Data obtained from the biosensor 60 attached to the subject U1 is referred to as biodata. The biodata includes brain wave data, heart rate data, sweat data, and body temperature data, and represents bioinformation of the subject U1. Depending on the type of bioinformation to be acquired or the wearability, other sensors may be attached to the subject U1 as the biosensor 60. Examples of other sensors include a blood pressure monitor or a NIRS (Near Infrared Spectroscopy) device. Note that it is not essential for the biosensor 60 to include a sweat sensor, and therefore the biodata may not include sweat data. Similarly, it is not essential for the biosensor 60 to include a body temperature sensor, and therefore the biodata may not include body temperature data.
[0025] The biometric data obtained by the biometric sensor 60 is provided to the emotion estimation device 10. The emotion estimation device 10 references the biometric data and estimates an emotion based on two emotion indices, which are indices indicating the mental and physical state of the subject U1. The mental and physical state of any given person includes a psychological state, and therefore each emotion indices can also be said to be an index indicating a psychological state. The value of an emotion index is referred to as an emotion index value. One of the emotion indices used in this embodiment is the arousal level of the central nervous system (hereinafter referred to as arousal level), and the other emotion index used in this embodiment is the activity level of the autonomic nervous system (hereinafter referred to as activity level). That is, one emotion index value is expressed by arousal level, and the other emotion index value is expressed by activity level. Arousal level can be derived, for example, based on electroencephalogram data. Activity level can be derived, for example, based on heart rate data. Specifically, arousal level is calculated based on beta waves and alpha waves of the electroencephalogram. The activity level is calculated from the standard deviation of the heartbeat LF (Low Frequency) component (the low frequency component of the heartbeat waveform signal).
[0026] The emotion estimation model 131 provided in the emotion estimation device 10 is a model for emotion estimation based on arousal and activity, and may be configured with a calculation formula or a conversion data table that identifies an emotion from arousal and activity. However, as will be described later, the emotion estimation model 131 estimates one or more emotion candidates as candidates for the emotion of the subject U1, and a single emotion estimation model 131 may not narrow down the emotion of the subject U1 to one.
[0027] The emotion estimation model 131 is a multidimensional model with arousal and activity as parameters, and in this case, it is a two-dimensional model with arousal and activity as two axes. The two-dimensional model is created based on medical evidence (such as papers) showing the relationship between arousal and activity and emotions. Alternatively, the two-dimensional model is created based on the results of a questionnaire administered to many subjects. The subject questionnaire results include the subject's arousal and activity identified from the measured values of the subject's electroencephalogram and heart rate, and the subject's declared emotions at the time of measuring the electroencephalogram and heart rate. Note that the emotion estimation model 131 may be a multidimensional model with three or more dimensions, that is, a model that further includes parameters (emotion indices) other than arousal and activity for estimating emotions.
[0028] FIG. 3 is a diagram showing an example of a two-dimensional model related to emotion estimation. FIG. 3 shows a two-dimensional psychological plane PP defined in the two-dimensional model. According to various medical evidence related to psychology, psychology can be estimated based on two types of indices that indicate mental and physical states. In the psychological plane PP shown in FIG. 3, a vertical axis corresponding to arousal level and a horizontal axis corresponding to activity level are defined.
[0029] Points indicating the arousal and activity of subject U1 can be plotted (placed) on the psychological plane PP. The coordinates of the plotted (placed) points move in the positive direction on the vertical axis as subject U1's arousal level increases, and move in the negative direction on the vertical axis as subject U1's arousal level decreases. The coordinates of the plotted (placed) points move in the positive direction on the horizontal axis as subject U1's activity level increases, and move in the negative direction on the horizontal axis as subject U1's activity level decreases. From the origin of the psychological plane PP, the positive side of the vertical axis corresponds to an aroused state, and the negative side of the vertical axis corresponds to an unconscious state. From the origin of the psychological plane PP, the positive side of the horizontal axis corresponds to a state in which the sympathetic nervous system is activated (a state in which relatively strong emotions are present), and the negative side of the horizontal axis corresponds to a state in which the parasympathetic nervous system is activated (a state in which relatively weak emotions are present).
[0030] The point where the vertical and horizontal axes intersect on the psychological plane PP is the origin. The psychological plane PP has first to fourth quadrants separated by the vertical and horizontal axes. The first quadrant is the area where, as viewed from the origin of the psychological plane PP, the vertical axis component has a positive value and the horizontal axis component also has a positive value. The second quadrant is the area where, as viewed from the origin of the psychological plane PP, the vertical axis component has a positive value and the horizontal axis component has a negative value. The third quadrant is the area where, as viewed from the origin of the psychological plane PP, the vertical axis component has a negative value and the horizontal axis component also has a negative value. The fourth quadrant is the area where, as viewed from the origin of the psychological plane PP, the vertical axis component has a negative value and the horizontal axis component has a positive value.
[0031] In the psychological plane PP, a corresponding psychological state (in other words, emotion) is assigned to each quadrant. The psychological states of "fun, joy, anger, sadness" are assigned to the first quadrant. The psychological state of "moderate tension" may also be assigned to the first quadrant. The psychological state of "melancholy, crying" is assigned to the second quadrant. "Crying" in the second quadrant can also be considered a form of "sadness." The psychological states of "boredom, relaxation, calmness" are assigned to the third quadrant. The psychological states of "anxiety, fear, unpleasantness" are assigned to the fourth quadrant. The distance from the origin indicates the intensity of the corresponding psychological state. For example, when points indicating the arousal and activity levels of subject U1 are plotted in the second quadrant, the greater the distance between the plotted point and the origin, the stronger the psychological state of "melancholy, crying" for subject U1.
[0032] In this way, the emotions of any person depend on the arousal level and activity level of that person. The position of each axis on the psychological plane PP may be set appropriately based on experiments, etc. Each axis may also be set based on the position of the neutral area described below.
[0033] The emotion estimation model 131 can estimate the psychological state of the subject U1 from coordinates obtained by plotting (arranging) points indicating the arousal and activity of the subject U1 on the psychological plane PP. The emotion estimation model 131 can estimate the psychological state and its intensity of the subject U1 based on which quadrant of the psychological plane PP the coordinates of the plotted point are in, the position within the quadrant, and the distance between the plotted point and the origin. The first to fourth quadrants on the psychological plane PP can also be considered as first to fourth classes, respectively. In this case, it can also be said that the emotion estimation model 131 classifies the psychological state of the subject U1 into one of the first to fourth classes. Note that when the emotion estimation model 131 is a multidimensional model with three or more dimensions, a multidimensional psychological space with three or more dimensions is used depending on the number of types of emotion index values used.
[0034] Incidentally, when emotion intensity is strong, that is, when the emotion index value swings significantly toward the maximum or minimum value, emotion estimation accuracy tends to increase. On the other hand, when emotion intensity is weak, that is, when the emotion index value is near the median, emotion estimation accuracy tends to decrease. For this reason, the region near the median of the emotion index value is set as the neutral region. Figure 4 is a diagram for explaining the neutral region. As shown in Figure 4, a method can be adopted in which the region near the median of the emotion index value is set as the neutral region (corresponding to the diagonal line region in Figure 4), and an emotion estimation is determined to be impossible or no emotion is estimated for emotion index values that fall into the neutral region.
[0035] There are an arousal neutral region NR1 and an activity neutral region NR2 as neutral regions. In the psychological plane PP, a region where the distance from the horizontal axis is less than or equal to a predetermined first boundary distance is the arousal neutral region NR1, and a region where the distance from the vertical axis is less than or equal to a predetermined second boundary distance is the activity neutral region NR2. The arousal neutral region NR1 and the activity neutral region NR2 overlap with each other in a region including the origin of the psychological plane PP. The combined region of the arousal neutral region NR1 and the activity neutral region NR2 is hereinafter referred to as the neutral region NR. Each position within the neutral region NR may be understood as not belonging to any of the first to fourth quadrants of the psychological plane PP.
[0036] <Claim 1> In the present disclosure, singularity point information of the emotion index value is extracted from the data of the emotion index value at the time of the first calibration, and hereinafter, a method of setting a neutral region based on the singularity point information without performing calibration is provided. Specifically, the emotion estimation device 10 accumulates the singularity point information of the emotion index value, and sets a neutral region (threshold value of the emotion index value) based on the accumulated singularity point information. A singularity point is a point where the value changes specifically in the time-series emotion index value, and details will be described later.
[0037] That is, in the present disclosure, by accumulating singularity point information where there is a high possibility that the emotion of the subject U1 has changed, the neutral region is set from the singularity point information without performing calibration. As a result, since the neutral region can be set without performing calibration after the singularity point information is acquired, the emotion estimation process can be started early. Details of the method for setting the neutral region in the present disclosure will be described later.
[0038] Next, a configuration example of the feeling estimation device 10 will be described. FIG. 5 is a diagram showing a configuration example of the feeling estimation device 10 according to the embodiment. Note that FIG. 5 shows components necessary for explaining the features of this embodiment, and general components are omitted. As shown in FIG. 5, the feeling estimation device 10 includes a communication unit 11 and a storage unit 12. The feeling estimation device 10 also includes a controller 13. The feeling estimation device 10 may be a so-called computer device. Note that the feeling estimation device 10 may be configured to include an input device such as a keyboard and an output device such as a display.
[0039] The communication unit 11 is an interface for communicating data with other devices via the network N. The communication unit 11 is, for example, a network interface card (NIC).
[0040] The storage unit 12 is configured to include a volatile memory and a non-volatile memory. The volatile memory may include, for example, a random access memory (RAM). The non-volatile memory may include, for example, a read-only memory (ROM), a flash memory, or a hard disk drive. The non-volatile memory stores computer-readable programs and data. Note that at least some of the programs and data stored in the non-volatile memory may be obtained from another computer device connected via a wired or wireless connection, or from a portable recording medium.
[0041] As shown in FIG. 5, in this embodiment, the storage unit 12 includes a table (data table) 121. In detail, the table 121 includes a plurality of tables for various processes. For example, the table 121 includes a psychological plane table 121a, a special task table 121b, and a neutral area table 121c. FIGS. 6 and 7 are diagrams showing an example of the psychological plane table 121a. FIG. 8 is a diagram showing an example of the special task table 121b. FIG. 9 is a diagram showing an example of the neutral area table 121c.
[0042] First, a search-type psychological plane table 121a1 and a formula-type psychological plane table 121a2, which are examples of psychological plane table 121a, will be described with reference to Fig. 6 and Fig. 7. Fig. 6 shows search-type psychological plane table 121a1, and Fig. 7 shows formula-type psychological plane table 121a2.
[0043] As shown in FIG. 6, the psychological plane table 121a1 is a two-dimensional matrix table with index types as parameters on the vertical and horizontal axes. In the search-type psychological plane table 121a, emotion data identified from two types of index type data are stored in memory cells (memory frames) determined by data (values) of the two types of index types. Note that the arousal level and activity level in the psychological plane table 121a1 are values indicating ranges. Then, using each calculated value of the two types of index types, a corresponding memory cell in the psychological plane table 121a1 is searched, and the data (emotion type) stored in the searched memory cell becomes the emotion estimation result (including "neutral"). In FIG. 6, when the arousal level is A1 and the activity level is B1, the emotion is "melancholy / sadness." Furthermore, in the psychological plane table 121a1, "neutral" refers to the neutral region. For example, when the arousal level is A3 and the activity level is B1, an emotion is estimated to be indeterminate. In addition, each emotion type and "neutral" stored in each storage cell in the psychological plane table 121a1 are determined at the time of the initial calibration.
[0044] Next, as shown in FIG. 7, the formula-type psychological plane table 121a2 is a table that stores upper and lower limits for each index type (arousal level, activity level) for each emotion type. Since a neutral region exists between each emotion type, the upper and lower limits for each emotion type are the upper and lower limits of the neutral region. In the psychological plane table 121a2, the boundary value (upper or lower limit) between the neutral region of arousal level and the neutral region of activity level is determined during the initial calibration, and the boundary value (upper or lower limit) with respect to the neutral region is set as the boundary value (upper or lower limit) for each emotion type. For example, the lower limit of arousal level in the region of the emotion type "melancholy / sadness" is the upper limit of the neutral region of arousal level. Furthermore, the upper limit of activity level in the region of the emotion type "melancholy / sadness" is the lower limit of the neutral region of activity level. Then, it is determined based on psychological plane table 121a2 which emotion type the calculated values of the two index types fall within, and the corresponding emotion type (including "neutral") is the emotion estimation result.
[0045] 8 is a diagram showing an example of the special task table 121b. As shown in Fig. 8, the items of the special task table 121b include "task ID," "sensor type," "corresponding indicator type," "task type," and "task content."
[0046] The item "task ID" of the special task table 121b stores ID data, which is identification information for identifying task information in the special task table 121b.
[0047] The item "task type" in the special task table 121b stores whether the task is a first task that increases the physiological response (indicator value is large) or a second task that decreases the physiological response (indicator value is small). Whether the task increases or decreases the physiological response is determined based on medical evidence. For example, if the type of physiological response is arousal level, the task that increases arousal level is the first task, and the task that decreases arousal level is the second task. Furthermore, if the type of physiological response is autonomic nervous system activity level, the task that activates the sympathetic nervous system is the first task, and the task that activates the parasympathetic nervous system (i.e., deactivates the sympathetic nervous system) is the second task.
[0048] The "task content" item in the special task table 121b stores the specific content of the task to be performed by the user U1. The task content is determined based on the above-mentioned task type and medical evidence. For example, when the physiological response type is arousal level, the task content of the first task is "mental addition of displayed numbers." The number of mental additions may be, for example, 10 times. For example, when the physiological response type is arousal level, the task content of the second task is "rest and clench your fist." The time spent resting with your fist clenched may be, for example, three minutes. For example, when the physiological response type is autonomic nervous system activity level, the task content of the first task is "standing test." The standing test requires, for example, the user to rest in a supine position with their eyes open for a predetermined time, and then sit upright with their upper body open and rest for a predetermined time. The predetermined time may be three minutes. For example, when the physiological response type is autonomic nervous system activity level, the task content of the second task is "Aschner's test." The Aschner test, for example, requires a subject to gently press the eyelid of one eye with the pads of their index and middle fingers for a predetermined period of time while keeping their eyes closed. Another example of the Aschner test is to wear an eye mask for a predetermined period of time, such as 3 minutes.
[0049] As shown in FIG. 9, the neutral region table 121c includes the following items: "neutral data ID," "user ID," "sensor type," "corresponding index type," "upper limit value," "lower limit value," and "acquisition date and time."
[0050] The item "Neutral Data ID" of the neutral area table 121c stores neutral data ID data, which is identification information for identifying each neutral data in the neutral area table 121c. Note that each data record forming the neutral area table 121c is created for each neutral data ID data.
[0051] The neutral area table 121c has an item "user ID" that stores user ID data, which is identification information for identifying a user. The neutral area table 121c also has items "corresponding index type" that indicates the type of neutral index value and "sensor type" that indicates the sensor type corresponding to the index value, and the values (data) of each item described below are stored for each type of index value (sensor type).
[0052] The "Upper limit" and "Lower limit" items of the neutral area table 121c store information on the upper and lower limit values of the neutral area. The "Acquisition date and time" item of the neutral area table 121c stores information on the date and time when the information on the upper and lower limits was updated and stored. The upper and lower limit values determined for each neutral data ID are updated each time the latest information is acquired. Furthermore, when the information on the upper and lower limits is updated, the acquisition date and time is also updated. In other words, the neutral area table 121c stores, for each user, the upper and lower limit values of the neutral area for each index value (arousal level and activity level), as well as the data update date and time of these upper and lower limit values.
[0053] Returning to FIG. 5, the controller 13 includes a processor that performs arithmetic processing and the like. The processor may be configured to include, for example, a CPU (Central Processing Unit). The controller 13 may be configured with one processor or multiple processors. When configured with multiple processors, the processors may be connected to each other so that they can communicate with each other. Note that when the emotion estimation device 10 is configured as a cloud server, the CPU that configures the processor may be a virtual CPU.
[0054] In this embodiment, the functions of the controller 13 are realized by the processor executing arithmetic processing in accordance with a program stored in the storage unit 12.
[0055] The scope of this embodiment may include a computer program that causes a processor (computer) to realize at least a portion of the functions of the emotion estimation device 10. The scope of this embodiment may also include a computer-readable nonvolatile recording medium that records such a computer program. The nonvolatile recording medium may be, for example, the nonvolatile memory described above, an optical recording medium (e.g., an optical disk), a magneto-optical recording medium (e.g., a magneto-optical disk), a USB memory, an SD card, or the like.
[0056] Furthermore, each function of the controller 13 may be realized by a single program, or, for example, a configuration in which each function is realized by a separate program. Each function may also be realized as a separate server. As described above, each function may be realized by having a processor execute a program, i.e., by software, but may also be realized by other methods. At least a portion of each function may be realized using, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). That is, each function may be realized by hardware using a dedicated IC or the like. Each function may also be realized by a combination of software and hardware. Each function is a conceptual component. A function performed by one component may be distributed among multiple components. Furthermore, functions possessed by multiple components may be integrated into one component.
[0057] Next, we will explain the emotion estimation process performed by the controller 13. Prior to the emotion estimation process, the controller 13 performs a calibration process to set a neutral region. As the calibration process, the controller 13 has the subject U1 perform a special task prepared in the special task table 121b, and determines the neutral region according to the measurement results of the biological data when the task is performed.
[0058] As can be seen from the contents of the special task table 121b, determining the neutral region for each physiological response (emotion index value) requires the subject U1 to perform a first task and a second task for each type of physiological response. The controller 13 performs the following process to determine the neutral region for each physiological response. First, the controller 13 has the subject U1 perform a first task in which the emotion index value obtained during the execution of the tasks is a value in the positive region (e.g., an aroused state) and a second task in which the emotion index value is a value in the negative region (e.g., a non-aroused state). The controller 13 then determines the upper limit of the neutral region based on the emotion measurement value obtained by executing the first task, and determines the lower limit of the neutral region based on the emotion measurement value obtained by executing the second task, thereby determining the neutral region. Once the neutral region is determined, the positive and negative regions for the emotion index value are automatically determined. That is, the physiological response can be divided into three regions. Furthermore, there is a time difference between the timing at which the first task and the second task are performed for each physiological response type, and the physiological response is measured for each task.
[0059] In this configuration, before emotion estimation, the subject U1 is made to perform a special task in advance, and the positive, neutral, and negative regions for the emotion index value are determined based on the results. In other words, the physiological response (index) can be divided into three regions, appropriately reflecting the environment in which the user U1 is currently placed. In other words, it is expected that emotion estimation will be performed appropriately.
[0060] A specific example of the above-mentioned method for determining the neutral region will be described with reference to FIG. 10. FIG. 10 is a diagram schematically illustrating the change over time in the index value of a physiological response when a first task and a second task are being executed. In FIG. 10, the horizontal axis represents time, and the vertical axis represents the physiological response (index) value. In FIG. 10, the index value curve α corresponds to the change over time in the index value obtained when the first task is being executed. The index value curve β corresponds to the change over time in the index value obtained when the second task is being executed. In FIG. 10, the special tasks are executed in the order of the first task and the second task, but this is an example, and the special tasks may be executed in the order of the second task and the first task. Furthermore, in FIG. 10, the area indicated by the arrow NR (a part of the area on the vertical axis) corresponds to the neutral region.
[0061] The controller 13 statistically processes the time-varying data of the index values obtained during execution of the first task to determine the upper limit of the neutral region NR. For example, the controller 13 smooths the time-varying data of the index values obtained during execution of the first task, and sets the minimum value of the data after the smoothing process as the upper limit of the neutral region NR. The controller 13 also statistically processes the time-varying data of the index values obtained during execution of the second task to determine the lower limit of the neutral region NR. For example, the controller 13 smooths the time-varying data of the index values obtained during execution of the second task, and sets the maximum value of the data after the smoothing process as the lower limit of the neutral region NR. The smoothing process is performed for the purpose of removing noise components such as minute peak signals.
[0062] In other words, the controller 13 determines the boundary between the neutral and positive regions of the index based on the index value obtained when the first task is executed. The controller 13 also determines the boundary between the neutral and negative regions of the index based on the index value obtained when the second task is executed. This type of calibration method is suitable when two special tasks (the first task and the second task) can be prepared that reliably result in a specific state of the emotion index value (either a positive value or a negative value).
[0063] The controller 13 determines the upper and lower limits of the neutral region for each type of physiological response. To this end, in this embodiment, the subject U1 is requested to perform a special task consisting of a first task and a second task for each of the two physiological responses (arousal level and autonomic nervous system activity level). Then, for each of the two physiological responses, a process of determining the neutral region is performed according to the measurement results of the biological signals during the execution of each special task.
[0064] In the above description, the first fluctuation range, which is the fluctuation range of the index value obtained when the first task is executed, and the second fluctuation range, which is the fluctuation range of the index value obtained when the second task is executed, are assumed to be separated from each other, and the neutral region NR is defined as the region between the two fluctuation ranges (see FIG. 10). The fluctuation range of the emotion index value refers to the range of value fluctuation accompanying changes in the index value over time. However, for example, if there is a small difference in the magnitude of the physiological response between the first task and the second task, the first fluctuation range and the second fluctuation range may overlap. In other words, rather than a special task that reliably puts the emotion index value into a specific state, there may be a task in which the emotion index value is biased toward the specific state (a case in which the emotion index value does not fit completely into the positive or negative range). In such a case, for example, the controller 13 may set the minimum value of the data obtained by smoothing the time-varying data of the index value obtained when the first task is executed as the lower limit of the neutral region NR. The controller 13 may then set the maximum value of the data obtained by smoothing the time-varying data of the index value obtained when the second task is executed as the upper limit of the neutral region NR.
[0065] In other words, the controller 13 determines the neutral region NR as the region where the first variation range of the index value during execution of the first task and the second variation range of the index value during execution of the second task overlap. By adopting such a neutral region method configuration, it is possible to ease the constraints on determining the first task and the second task to be used, making it easier to set the special task.
[0066] Next, a second example of the region determination process for determining the neutral region will be described. In the second example, the subject U1 is not required to perform a special task. In the second example, the neutral region is determined by focusing on the fact that as the accumulation of time-series data of physiological responses progresses, the upper and lower limit values of the neutral region estimated by statistical processing converge to certain values. Once the neutral region is determined, the positive and negative regions for the physiological responses (emotion index values) are automatically determined. In other words, once the neutral region is determined, the stratified region for the physiological responses can be divided into three regions.
[0067] In other words, in the second example, the controller 13 accumulates time-series data of the index values and determines the neutral region based on the results of statistical processing of the time-series data. This method does not require the subject U1 to perform a special task before estimating the emotion, thereby reducing the burden on the subject U1.
[0068] A specific method for determining a neutral region without a specific task will be described with reference to FIGS. 11 and 12. FIG. 11 is a diagram schematically showing changes over time in a physiological response. In FIG. 11, the horizontal axis represents time, and the vertical axis represents a physiological response index (FIG. 11 uses arousal level as an example). FIG. 12 is a diagram schematically showing changes over time in an estimated upper limit value of the neutral region (see FIG. 7). In FIG. 12, the horizontal axis represents time, and the vertical axis represents the estimated upper limit value of the neutral region. Note that the changes over time in an estimated lower limit value of the neutral region show the same changes over time as the upper limit value shown in FIG. 12. In other words, the vertical axis of the graph shown in FIG. 12 may be interpreted as the lower limit value.
[0069] As shown in Figure 11, a person's physiological responses (which can also be considered mental and physical states) change over time (based on changes in the environment, etc., which change over time) even without a specific task. Specifically, physiological responses fluctuate between high and low states, or an ambiguous state (neutral state) over time. Because physiological responses fluctuate up and down in response to external stimuli, the distribution of physiological response values is statistically distinctive. For example, since there are many states without external stimuli, physiological response values are usually within the neutral region. When stimuli are present, physiological response values tend to fluctuate suddenly. Time-series data of physiological responses is accumulated, and cluster analysis is performed to divide the accumulated data into the three states mentioned above, taking into account the characteristic trends in physiological response fluctuations in response to external stimuli. The upper and lower limits of the neutral region can be estimated by cluster analysis. A known method may be used for the cluster analysis. For example, k-means, Otsu's multilevel thresholding method, Gaussian mixture model, etc. may be used as the cluster analysis method.
[0070] In FIG. 12, it is assumed that the amount of accumulated time-series data of physiological responses increases over time. That is, in FIG. 12, it is assumed that the amount of data used to estimate the upper limit value of the neutral region using a statistical method (cluster analysis in a detailed example) increases over time. As the amount of data increases over time, the upper limit value estimated by statistical processing converges to a certain value over time, as shown in FIG. 12. This tendency also applies to the lower limit value, as described above. The neutral region identified by the converged upper and lower limit values is considered to be highly reliable. For this reason, in this example, the neutral region is determined when it is determined that the estimated upper and lower limit values have converged.
[0071] Whether the upper limit value has converged may be determined based on a comparison between the current estimated upper limit value and the previous estimated upper limit value. For example, if the ratio of the two is within a predetermined range relative to 1, it may be determined that the fluctuations are small and therefore convergence has occurred. Also in this example, the controller 13 determines the upper limit value and the lower limit value of the neutral region for each type of physiological response.
[0072] The controller 13 stores the upper and lower limit values of the neutral region determined by performing a special task or without performing a special task in the storage unit 12. In detail, the controller 13 stores the determined upper and lower limit values of the neutral region in a data table in the storage unit 12. That is, the table 121 includes a neutral region table 121c including neutral region information for each subject U1.
[0073] <Claim 6> Next, the controller 13 extracts, for each subject, singularity information, which is information relating to singularities in the emotion index values, based on the time-series data of the emotion index values at the time of calibration. That is, during the initial calibration in which the subject U1 is instructed to perform a special task, the controller 13 extracts singularity information from the emotion index values when the special task is performed. This results in extracting singularity information from the emotion index values when the special task is performed, which allows changes in the emotional state to be identified, and therefore allows for highly significant singularity information to be obtained. Here, using FIG. 13, a description will be given of singularities, which are points at which the time-series emotion index values become unique or change uniquely.
[0074] FIG. 13 is a diagram showing an example of singularity information. The singularity information shown in FIG. 13 is stored and accumulated in, for example, the storage unit 12. As shown in FIG. 13, the controller 13 extracts information on how the biometric information (emotion index value) appears, the width and position of the neutral region, emotions that tend to appear in the biometric information, emotions that do not tend to appear in the biometric information, and subjective symptoms of emotions, as singularity information, and records the extracted information. Specifically, the singularity information has items such as "singularity ID," "singularity type," "data content," "target processing," and "processing content." Note that the data on "singularity ID," "singularity type," "target processing," and "processing content," excluding "data content," are set and stored in advance as appropriate by a system designer or the like.
[0075] "Singularity ID" stores singularity ID data, which is identification information for identifying each piece of singularity information. Note that each data record that stores each piece of singularity information is created for each piece of singularity ID data.
[0076] The "singularity type" is information indicating the type of singularity information, i.e., the type of characteristics of the singularity in the singularity information. The singularity type is selected from among predetermined singularity types. The "data content" is data indicating the specific content of the singularity, i.e., data indicating the specific content and characteristics of the singularity extracted from the time-series data. Note that in FIG. 13, the data content is expressed abstractly, such as "A1," but in reality, data with specific content, as described below, is stored. The "target processing" is information indicating the type of processing to be performed when a singularity such as that shown in the singularity data is detected. The "processing content" is information indicating the specific processing content of the target processing. Note that in FIG. 13, the data content is expressed abstractly, such as "amplitude correction a," but in reality, data with specific content, as described below, is stored.
[0077] For example, singularity information identified by singularity ID "S1" has a singularity type of "manifestation of biometric information," data content of "time-series data of emotion index values," and target processing to be performed when such a singularity is detected is "correction of time-series data," the specific processing content of which is "fluctuation range correction a." The "manifestation of biometric information" in data identified by singularity ID "S1" is information (an example of correspondence information) relating to the range of fluctuation in the time-series data of emotion index values, and is information indicating the extent of the fluctuation of the emotion index value when the emotion of subject U1 changes. For example, during the initial calibration, the controller 13 records the amount of change in the emotion index value when subject U1 performs a special task (the difference between the emotion index value when a special task is performed that increases the emotion index value and the emotion index value when a special task is performed that decreases the emotion index value, or the difference between these emotion index values and the emotion index value when a special task is performed that is neutral) as the manifestation of biometric information. For example, if the amount of change in the emotion index value is equal to or greater than a threshold, the controller 13 records "strongly expressed" as data content A1, and if the amount of change is less than the threshold, records "weakly expressed" as data content A1. Note that instead of the two levels of "strongly expressed" and "weakly expressed," stratification may be into multiple levels of three or more. Furthermore, the controller 13 may record the maximum amount of change in the emotion index value when a special task is performed as the manifestation of the biometric information. Note that the manifestation of the biometric information is not limited to the maximum amount of change in the emotion index value, and may also be an average value. Then, after completing the initial calibration, the controller 13 corrects the emotion index value based on that calibration and performs emotion estimation.
[0078] <Claim 2> Then, at the next opportunity for emotion estimation (for example, in a situation where the environment has changed after some time has passed, such as the next driving opportunity in the case of a vehicle), the controller 13 does not perform calibration, but corrects the time-series data of emotion index values based on the results of a feature comparison between the recorded data content A1 and the detected emotion index values (data acquired over a period of time that is appropriately set in advance based on experiments, etc., from the start of data collection). Specifically, the controller 13 compares the time-series data of emotion index values newly acquired for emotion estimation with the appearance of the biometric information recorded as singularity information, and determines a method of correcting the time-series data of emotion index values based on the comparison results—in this case, whether or not to correct the time-series data of emotion index values. Specifically, if the amplitude of fluctuation in the time-series data of emotion index values is equal to or greater than a threshold (a value appropriately set so as to correspond to “strongly expressed”) and the appearance of the biometric information in the singularity information is “strongly expressed,” the controller 13 performs emotion estimation processing without correcting the time-series data. In other words, if the fluctuation range of the time-series data matches the manifestation of the biometric information in the singularity information, the controller 3 infers that the state is the same as at the time of calibration and does not correct the time-series data (the calibrated settings are used as is). On the other hand, if the fluctuation range in the time-series data of the emotion index value is less than the threshold and the manifestation of the recorded biometric information is "strong," the controller 13 corrects the time-series data to increase the fluctuation range. In other words, the controller 13 infers that the situation is one in which the emotion index value will be detected as smaller than the state at the time of calibration, and corrects the emotion index value to increase by multiplying the entire time-series data by a predetermined coefficient (a value exceeding 1 that is appropriately set based on experiments, etc.). Furthermore, if the controller 13 has recorded a maximum value as data on the amount of change in the emotion index value as the manifestation of the biometric information, it may multiply the emotion index value of the time-series data by a coefficient (maximum value recorded as the manifestation of the biometric information / maximum value of the newly acquired time-series data) so that the maximum value of the newly acquired time-series data matches the maximum value recorded as the manifestation of the biometric information.
[0079] Furthermore, if the fluctuation range in the time-series data of the emotion index value is equal to or greater than a threshold value (a value appropriately set so as to correspond to "weakly expressed") and the recorded manifestation of the biometric information is "weakly expressed", the controller 13 corrects the time-series data to reduce the fluctuation range overall. In other words, the controller 13 infers that the situation is such that the emotion index value will be detected as being greater than the state at the time of calibration, and corrects the emotion index value to be smaller by multiplying the entire time-series data by a predetermined coefficient (a value less than 1 that is appropriately set based on experiments, etc.). Furthermore, if the controller 13 has recorded the maximum value of the amount of change in the emotion value described above as the manifestation of the biometric information, it may multiply the emotion index value of the time-series data by a coefficient (maximum value recorded as the manifestation of the biometric information / maximum value of the newly acquired time-series data) so that the maximum value of newly acquired time-series data is the same as the maximum value recorded as the manifestation of the biometric information.
[0080] As a result, even if a situation arises in which the emotion index value is detected with a tendency different from that at the time of the initial calibration due to environmental changes such as changes in the driving environment or the in-vehicle environment, the controller 13 will make a correction appropriate to this situation based on the singularity information. As a result, the controller 13 can improve the accuracy of the estimation results in the emotion estimation process at a later stage.
[0081] The "neutral area width and position" in the singularity information identified by the singularity ID "S2" is information (an example of correspondence information) that determines the threshold used to stratify the emotion index values. The width and position, which are parameters of the neutral area set by the initial calibration, are recorded as data content B2. Specifically, the width of the neutral area is the length from the upper limit to the lower limit of the neutral area. The position of the neutral area is the median value of the upper and lower limits, i.e., the position of the central axis of the neutral area. The position of the neutral area is represented, for example, by the coordinates (emotion index values) of the psychological plane PP model. Note that the position of the neutral area may also be represented by the coordinates (emotion index values) of the upper and lower limits (once the upper and lower limits of the neutral area are determined, the position is determined from these).
[0082] <Claim 3> The controller 13 then compares the width and position of the neutral region, which is the singularity information recorded during the initial calibration, with the width and position of the neutral region estimated based on the time-series data of emotion index values newly acquired for emotion estimation, and determines a correction method for the time-series data of emotion index values based on the comparison results—here, the correction value to be used in the emotion index value correction process. Specifically, the controller 13 first estimates the width and position of the neutral region based on the time-series data of emotion index values newly acquired for emotion estimation. Because the emotional state is maintained (slowly transitioning to a normal state) until a new external stimulus is encountered, situations arise in which the emotion index is maintained in a high, neutral, or low state. In other words, the distribution of emotion index values tends to be dominated by data in high, neutral, and low states. The controller 13 utilizes this characteristic to estimate the width and position of the neutral region based on the distribution characteristics of the data for the high, neutral, and low states. For example, the average value of the emotion index value at the distribution center of gravity in a high state and the emotion index value at the distribution center of gravity in a low state is taken as the position (median) of the neutral region. Then, a predetermined weighted average value (weighting coefficients are set to appropriate values based on experiments, etc.: for example, 0.3 on the high state side and 0.7 on the neutral region side) of the emotion index value at the distribution center of gravity in a high state and the emotion index value at the position of the neutral region is taken as the upper boundary value of the neutral region. Also, a predetermined weighted average value (weighting coefficients are set to appropriate values based on experiments, etc.: for example, 0.3 on the low state side and 0.7 on the neutral region side) of the emotion index value at the distribution center of gravity in a low state and the emotion index value at the position of the neutral region is taken as the lower boundary value of the neutral region. Then, the controller 13 corrects the newly acquired time-series data for emotion estimation so that the width and position of the neutral region estimated based on the emotion index value newly acquired for emotion estimation are aligned with the width and position of the neutral region recorded as singularity information.
[0083] Specifically, the time series data is corrected using the following calculation formula, for example. F1 (corrected emotion index value) = F0 × α + β α=(Lu0-Ld0) / (Lu1-Ld1) β=Lo0-Lo1 [Variable definition in formula] A. Emotion index value in the acquired (calculated) time series data: F0 B. Neutral area setting value set by calibration (calculated separately using the method described above) Upper boundary value, lower boundary value, position (center axis): Lu0, Ld0, Lo0 (= (Lu0 + Ld0) / 2) C. The neutral region setting value estimated based on the time series data acquired during emotion estimation (calculated separately using the method described above) Upper boundary value, lower boundary value, position (center axis): Lu1, Ld1, Lo1 (= (Lu1 + Ld1) / 2)
[0084] In other words, even if a situation arises in which the emotion index value is detected with a tendency different from that at the time of the initial calibration due to environmental changes such as changes in the driving environment or the in-vehicle environment, the controller 13 will make a correction appropriate to this situation based on the singularity information (features of the neutral region).As a result, the controller 13 can improve the accuracy of the estimation results in the emotion estimation process at a later stage.
[0085] The corrections based on the singularity information identified by the singularity IDs "S1" and "S2" above are corrections to the detected emotion index value, but the corrections based on the singularity information identified by the following singularity IDs "S3" and "S4" are corrections (changes) to the estimated emotion and emotion-based control processing. Furthermore, this singularity information is generated and stored based on data (biological signals, emotion index values, estimated emotions) acquired as needed during emotion estimation processing, rather than during calibration.
[0086] The "emotion likely to appear in biometric information" in the singularity information identified by singularity ID "S3" is information (an example of corresponding information) indicating an emotion in which the emotion index value of subject U1 changes significantly when in that emotional state. Conversely, the "emotion unlikely to appear in biometric information" in the singularity information identified by singularity ID "S4" is information (an example of corresponding information) indicating an emotion in which the emotion index value of subject U1 changes little when in that emotional state. For example, the following combinations of types (combinations of A, B, C and D, E, F) are possible types of human sensitivity to external stimuli. Type A: When in a positive emotional state, biological information reacts strongly to stimuli. Type B: When in a negative emotional state, biological signals do not react very well to stimuli. Type C: When in a negative emotional state, biological signals react normally to stimuli. Type D: When in a positive emotional state, biological information reacts strongly to stimuli. Type E: When in a negative emotional state, biological signals do not react very well to stimuli. F type: When in a negative emotional state, biological information reacts normally to stimuli.
[0087] For emotions to which the biometric information responds strongly, the reliability of the estimated emotion based on the discrimination result using the biometric signal threshold is high. Conversely, for emotions to which the biometric information does not respond much, the reliability of the estimated emotion based on the discrimination result using the biometric signal threshold is low. Taking this characteristic into consideration, processing and control corresponding to the emotional state are performed during emotion estimation and various control operations using the estimated emotion. Specifically, for an "emotion likely to appear in biometric information," the system loosens restrictions on vehicle control using the estimated emotion (actively performs control based on the estimated emotion). Alternatively, the system performs processing such as including the "emotion likely to appear in biometric information" in the candidate estimated emotion types for emotion estimation based on biometric information. Conversely, for an "emotion unlikely to appear in biometric information," the system notifies the user (subject) of the estimated emotion at appropriate intervals, and the user inputs a judgment evaluation of the estimated emotion. Based on the result, the judgment threshold for the "emotion unlikely to appear in biometric information" is adjusted (if the judgment result is incorrect, the threshold is changed so that the emotion type of the "emotion unlikely to appear in biometric information" is difficult to infer). Alternatively, the system performs processing such as excluding the "emotion unlikely to appear in biometric information" from the candidate estimated emotion types for emotion estimation based on biometric information.
[0088] The "emotional awareness" in the singularity information identified by the singularity ID "S5" is information (an example of correspondence information) indicating the degree to which the subject U1 is aware of the emotion they are experiencing (at what intensity the emotion becomes aware). If the subject U1 is easily aware of emotions, the subject U1 will become aware of the emotion early on as the emotion's intensity increases, allowing them to respond to the emotion quickly and, as a result, be able to avoid dangerous situations caused by the emotion before they occur. Conversely, if the subject U1 is not easily aware of emotions, the subject U1 is likely to be slow to initiate avoidance actions in the event of a dangerous situation caused by the emotion. Taking this characteristic into consideration, processing and control are performed according to the subject U1's emotional awareness characteristics when estimating emotions and when various controls are performed using the estimated emotions. Specifically, if the subject U1 is "emotionally aware," the start of vehicle control restrictions using the estimated emotion is delayed, for example, driving advice based on the estimated emotion is delayed (advice is provided when a stronger emotion is detected than usual). Conversely, if the driver is "difficult to recognize emotions," the system will start restricting vehicle control using estimated emotions sooner, for example, by providing driving advice based on estimated emotions sooner (by detecting emotions weaker than usual).
[0089] <Claim 4> Furthermore, the controller 13 accumulates information such as time-series data, correction process details, emotion estimation results, etc. during normal operation (emotion estimation process, emotion-based control, etc.), analyzes trends in singular points based on the accumulated information, and performs processes such as adding and updating the singular point information shown in Fig. 13. For example, when the controller 13 identifies through analysis of the accumulated time-series data that biometric information or emotions undergo a unique change under specific conditions, the controller 13 links the specific condition with the details of the unique change and records (adds) it as singular point information. Note that specific conditions are, for example, conditions that cause unique changes in biometric information, as well as conditions related to the driving environment (occurrence of an accident, sudden braking, sudden steering, etc.), conditions related to the in-vehicle environment (output status of content such as music, presence or absence of passengers, etc.), and other conditions that cause unique changes in biometric information or emotions. Furthermore, when time series data corresponding to already stored singularity information is detected, the time series data in the singularity information is updated based on the difference between the time series data in the singularity information and the newly detected time series data (for example, an appropriate weighted average of these pieces of information is updated as new time series data in the singularity information), or when the accuracy rate of the emotion estimated based on the singularity information falls below an appropriately set reference value, the singularity information is deleted, etc. Then, for example, when a specific condition is detected by the vehicle sensor unit 40, the controller 13 performs emotion estimation processing using singularity information corresponding to that condition.
[0090] In this way, the controller 13 can improve the accuracy of the singularity information by analyzing the newly obtained time-series data of emotion index values and updating the singularity information.
[0091] Next, the flow of processing executed by the feeling estimation device 10 will be described with reference to Figs. 14 and 15. Fig. 14 is a flowchart showing the processing procedure of calibration processing at a first emotion processing opportunity, etc. Fig. 15 is a flowchart showing the processing procedure of emotion estimation processing. Note that the processing shown in Fig. 14 is executed only once, such as when the feeling estimation device 10 is started for the first time or when a calibration execution operation is performed by a user, and the processing shown in Fig. 15 is executed repeatedly in a state in which the in-vehicle system SYS is activated, for example, from when the IG of the vehicle V1 is turned on until it is turned off.
[0092] 14, the controller 13 first issues a special task instruction to the subject U1 wearing the biosensor 60 (step S101). Next, the controller 13 acquires biodata from the biosensor 60 when the subject U1 executes the special task (step S102). The biodata includes, for example, brain wave data, heart rate data, sweat data, and body temperature data.
[0093] Next, the controller 13 calculates the emotion index value of the subject U1 based on the acquired biological data (step S103). For example, the controller 13 calculates the level of arousal based on the electroencephalogram data, and calculates the activity level based on the heartbeat data. More specifically, the controller 13 calculates the level of arousal based on the beta / alpha waves of the electroencephalogram. The controller 13 also calculates the activity level based on the standard deviation of the LF component in the heartbeat.
[0094] Next, the controller 13 sets a neutral region based on the calculated emotion index value (step S104). Specifically, the controller 13 sets the upper and lower limit values (width) and the central axis (position) of the neutral region.
[0095] Next, controller 13 extracts singularity information from the time-series data of emotion index values, the estimated emotion type, etc., and stores it in storage unit 12 (step S105), then ends the process. Specifically, controller 13 records, as singularity information, how the biometric information appears, the width and position of the neutral region, emotions that are likely to appear in the biometric information, emotions that are unlikely to appear in the biometric information, and subjective symptoms of the emotions.
[0096] Next, the emotion estimation process will be described. As shown in Fig. 15, the controller 13 first acquires biometric data measured by the biometric sensor 60 worn by the subject U1 (step S201). The biometric data includes, for example, brain wave data, heart rate data, sweat data, and body temperature data.
[0097] Next, the controller 13 calculates the emotion index value of the subject U1 based on the acquired biological data (step S202). For example, the controller 13 calculates the level of arousal based on the electroencephalogram data, and calculates the activity level based on the heartbeat data. More specifically, the controller 13 calculates the level of arousal based on the beta / alpha waves of the electroencephalogram. The controller 13 also calculates the activity level based on the standard deviation of the LF component in the heartbeat.
[0098] Next, the controller 13 determines whether or not the time-series data of emotion index values has accumulated to a predetermined amount (an appropriate value is set based on experiments, etc.) or more (step S203). Specifically, the controller 13 determines whether or not the length of collection time of the time-series data of emotion index values (the number of data constituting the time-series data) is equal to or greater than a threshold value.
[0099] If the controller 13 has accumulated more than a predetermined amount of time-series data of emotion index values (step S203: Yes), it reads out singularity information in which features extracted based on information related to the acquired and accumulated time-series data of emotion index values correspond to features (data content) in the singularity information (step S204). Specifically, it reads out singularity information having features that correspond to the features of the accumulated time-series data from the singularity information recorded in the storage unit 12 (step S204). On the other hand, if the controller 13 has accumulated less than the predetermined amount of time-series data of emotion index values (step S203: No), it returns to step S201.
[0100] Next, controller 13 corrects the time-series data based on the read-out singularity information (step S205). Specifically, controller 13 corrects the time-series data based on the appearance of the biometric information, which is the singularity information, and the width and position of the neutral region. Note that, in the case of singularity information that does not include time-series data correction processing, controller 13 does not perform any particular processing in step S205, and after processing in step S206 (described below), controller 13 will correct the emotion estimation result, adjust control based on the estimated emotion, and so on.
[0101] Next, the controller 13 performs emotion estimation processing based on the corrected time-series data (or the original uncorrected time-series data if no correction was made in step S205) (step S206), and ends the processing. Note that as a process using singularity information other than the correction of time-series data, specifically, the controller 13 may perform processing such as asking the subject U1 to confirm the emotion estimation result based on emotions that are likely (or unlikely) to appear in the biometric information, which is the singularity information, and subjective symptoms of the emotions, and correcting the emotion estimation result in accordance with the response result.
[0102] As described above, the feeling estimation device 10 performs the feeling estimation process using singularity information based on previously acquired information. Therefore, the feeling estimation device 10 does not need to perform calibration every time the feeling estimation device 10 is started, and the feeling estimation process can be started quickly after the feeling estimation device 10 is started.
[0103] <Claim 7> <Modification> The method of acquiring biometric data is not limited to acquisition from the body contact type biometric sensor 60, and for example, a technology (hereinafter referred to as a first inverse estimation technology) may be adopted in which the data is acquired from a biometric information estimation model that estimates the brain waves and heart rate of the subject U1 from image information of the subject U1 (the biometric information estimation model is installed in the emotion estimation device 10).
[0104] When the first inverse estimation technique is adopted, a biometric information estimation model is stored in the storage unit 12 of the emotion estimation device 10. The biometric information estimation model is an AI (artificial intelligence) model that has a person's facial image information as input information and has trained data in which the person's brain waves and heart rate information at the time the facial image was acquired is used as correct answer data. Therefore, by inputting facial image information (camera-captured image) of subject U1 into the biometric information estimation model, the brain waves and heart rate of subject U1 are estimated (inferred) by the biometric information estimation model.
[0105] The controller 13 extracts facial image information of the subject U1 from the in-vehicle image information captured by the in-vehicle camera 72 and inputs it into the biological information estimation model. As a result, the obtained estimation results of the subject U1's brain waves and heart rate are supplied to the controller 13 as brain wave data and heart rate data of the subject U1. The controller 13 derives emotion index values (typically, arousal level and activity level) of the subject U1 based on the brain wave data and heart rate data obtained by estimation using the biological information estimation model, and further estimates the emotion based on these emotion index values.
[0106] The biological information estimation model is created, for example, in a learning device (not shown). The learning device may be configured with any one or more server devices. A method for creating the biological information estimation model in the learning device will be described.
[0107] First, in the data collection step, a plurality of sets of training data including image information of a subject and electroencephalogram data and heart rate data of the subject are collected. A large amount of training data (for example, tens of thousands of sets) is collected.
[0108] Image information of the subject includes image information of the subject's face and is obtained by photographing the subject with any camera. An EEG sensor equivalent to the EEG sensor is attached to the subject. The EEG sensor attached to the subject measures the subject's brain waves to obtain the subject's brain wave data (hereinafter referred to as measured EEG data). A heart rate sensor equivalent to the heart rate sensor is attached to the subject. The heart rate sensor attached to the subject measures the subject's heart rate to obtain the subject's heart rate data (hereinafter referred to as measured heart rate data).
[0109] A unit period of an appropriate length (set based on experiments, etc.) for estimating brain waves and heart rate from facial images is set, and facial image information of the subject obtained by photographing the subject during that unit period is matched with the measured brain wave data and measured heart rate data during that unit period. One set of learning data is composed of the matched subject's image information and the subject's measured brain wave data and measured heart rate data. In other words, one set of learning data consists of the subject's facial image information, measured brain wave data, and measured heart rate data obtained by photographing and measuring during the same period. The subject may be subject U1 himself (advantageous for brain wave and heart rate estimation AI dedicated to subject U1), or it may be a subject other than subject U1, or multiple subjects (advantageous for data collection and general-purpose AI generation).
[0110] After the data collection process, a learning process is performed in the learning device. In the learning process, learning data is input to the learning device. A pre-learning AI model consisting of a neural network is provided in the learning device. The learning device provides the pre-learning AI model with image information of the subject as input data and with the subject's measured electroencephalogram data and measured heart rate data as correct answer data. The AI model estimates (infers) the subject's electroencephalogram data and heart rate data from the subject's image information. In the learning process, the learning device performs learning of the AI model using an error backpropagation method or the like to reduce errors between the electroencephalogram data and heart rate data estimated by the AI model and the measured electroencephalogram data and measured heart rate data provided as correct answer data. The learning here is supervised machine learning, and the parameters (weights, etc.) of the AI model are adjusted during the learning. The learning ends when a predetermined learning termination condition is met, such as when the error converges to a sufficiently small value. The AI model after learning is incorporated into the emotion estimation device 10 as a biological information estimation model.
[0111] <Claim 8> Furthermore, in the above modification, an example of estimating biometric data from image information has been described, but an emotion index value may be estimated from image information. In this case, the emotion estimation device 10 employs a technology (hereinafter referred to as a second inverse estimation technology) in which an emotion index estimation model that estimates the arousal and activity level of the subject U1 from image information of the subject U1 is installed in the emotion estimation device 10. Note that the emotion estimation device 10 does not need to perform processing to calculate an emotion index value from biometric data.
[0112] When the second inverse estimation technique is adopted, an emotion index estimation model is stored in the storage unit 12 of the emotion estimation device 10. The emotion index estimation model is an AI model that has a person's facial image information as input information and has trained data that uses the person's emotional index values (arousal and activity) at the time the facial image was acquired as correct answer data. Therefore, by inputting the facial image information of subject U1 into the emotion index estimation model, the arousal and activity of subject U1 are estimated (inferred) by the emotion index estimation model.
[0113] The controller 13 extracts facial image information of the subject U1 from the in-vehicle image information captured by the in-vehicle camera 72 and inputs it into the emotion index estimation model. As a result, the emotion index estimation model estimates and derives the subject U1's arousal and activity as the subject U1's emotion index value. The controller 13 estimates the emotion of the subject U1 based on the emotion index value (typically arousal and activity) of the subject U1 obtained by estimation using the biological information estimation model.
[0114] The emotion index estimation model is created in a learning device (not shown). The learning device may be configured with any one or more server devices. A method for creating an emotion index estimation model in the learning device will be described below.
[0115] First, in the data collection step, a large amount of learning data (for example, tens of thousands of sets) is collected, each set including a facial image of a subject and the subject's level of arousal and activity.
[0116] Image information of the subject includes image information of the subject's face, and is obtained by photographing the subject with any camera. An EEG sensor equivalent to the EEG sensor is attached to the subject. The EEG sensor attached to the subject measures the subject's brain waves, thereby obtaining the subject's brain wave data (referred to as measured EEG data as described above). A heart rate sensor equivalent to the heart rate sensor is attached to the subject. The heart rate sensor attached to the subject measures the subject's heart rate, thereby obtaining the subject's heart rate data (referred to as measured heart rate data as described above).
[0117] In the data collection step, the level of alertness is estimated using a first model (e.g., a model that calculates the level of alertness from beta waves / alpha waves of the brain waves) based on the measured electroencephalogram data. In the data collection step, the level of activity is estimated using a second model (e.g., a model that calculates the level of activity from the standard deviation of the LF component of the heart rate) based on the measured heart rate data. The level of alertness and the level of activity of the subject estimated by the first and second models are included in the training data.
[0118] More specifically, a unit period of time appropriate for estimating emotion index values from facial images (set based on experiments, etc.) is set, and facial image information of the subject obtained by photographing the subject during that unit period is associated with the subject's alertness and activity level during that unit period. The subject's alertness and activity level during a certain unit period are estimated using first and second models based on the subject's measured electroencephalogram data and measured heart rate data during that unit period. The associated image information of the subject and the subject's alertness and activity level form a set of learning data. The subject may be subject U1 itself (advantageous for an emotion index value estimation AI dedicated to subject U1), or may be a subject other than subject U1 or multiple subjects (advantageous for data collection and general-purpose AI generation). The first and second models may be created in advance using known supervised machine learning methods.
[0119] After the data collection step, a learning step is performed in the learning device. In the learning step, learning data is input to the learning device. An AI model consisting of a neural network is provided in the learning device. The learning device provides the AI model before learning with image information of the subject in the learning data as input data and the subject's arousal and activity levels in the learning data as ground truth data. The AI model estimates (infers) the subject's arousal and activity levels from the subject's image information. In the learning step, the learning device performs learning of the AI model using backpropagation or the like to reduce errors between the arousal and activity levels estimated by the AI model and the arousal and activity levels provided as ground truth data. The learning here is supervised machine learning, and the parameters (weights, etc.) of the AI model are adjusted during the learning. Learning ends when a predetermined learning termination condition is met, such as when the error converges to a sufficiently small value. The AI model after learning is incorporated into the emotion estimation device 10 as an emotion index estimation model.
[0120] Further advantages and modifications will readily occur to those skilled in the art. Therefore, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and their equivalents. [Explanation of symbols]
[0121] 10 Emotion estimation device 11 Communications Department 12 Storage section 13 Controller 20 Vehicle control device 30 Actuator section 40 Vehicle sensor unit 60 Biometric Sensor 71 Exterior camera 72 In-car camera 73 Display section 74 Speaker 75 microphone 121 Table 131 Emotion estimation model PP psychological plane SYS In-vehicle system
Claims
1. An emotion estimation device that estimates emotions from emotion index values related to a biological state of an emotion estimation subject based on estimation conditions adjusted by calibration, the emotion estimation device comprising: a controller; The controller During calibration, singularity information that is a feature of the time-series data of the detected emotion index values is extracted and stored, and during emotion estimation, corresponding information that corresponds to the singularity information in the time-series data of the detected emotion index values is extracted, and emotion estimation is performed based on the difference between the singularity information and the corresponding information. Emotion estimation device.
2. The controller Correcting the detected emotion index value according to the difference between the singularity information and the corresponding information, and estimating emotion based on the corrected emotion index value The emotion estimation device according to claim 1 .
3. The singularity information and the correspondence information are the amplitude of fluctuation in the time-series data of emotion index values. The emotion estimation device according to claim 1 or 2.
4. If the emotion index value is in the positive region, the emotion index value is determined to be in the first emotional state; If the emotion index value is in the negative region, the emotion index value is determined to be in the second emotional state; a feeling estimation device that determines that the emotion index value is in an indeterminate emotion state when the emotion index value is located in a neutral region between the positive region and the negative region, The controller During calibration, a neutral region threshold that defines a neutral region is determined based on time-series data of emotion index values detected using a neutral region determination method, and during emotion estimation, a comparison region threshold that defines a comparison region is determined based on time-series data of emotion index values detected using the neutral region determination method, the emotion index value detected is corrected in accordance with the difference between the neutral region threshold and the comparison region threshold, and emotion estimation is performed based on the corrected emotion index value. The emotion estimation device according to claim 1 or 2.
5. The controller The singularity information is updated based on the corrected time series data. The emotion estimation device according to claim 1 or 2.
6. The controller During the calibration, the singularity information is extracted from the emotion index value when the emotion estimation subject performs a special task. The emotion estimation device according to claim 1 or 2.
7. The controller The emotion index value is generated by an AI model that estimates an emotion index value from facial image information of a person whose emotion is to be estimated. The emotion estimation device according to claim 1 or 2.
8. The controller The biological state of the emotion estimation subject is estimated using an AI model that estimates the biological state from facial image information, and the emotion index value is calculated based on the estimated biological state. The emotion estimation device according to claim 1 or 2.
9. An emotion estimation method for estimating emotions from emotion index values related to a biological state of an emotion estimation target person based on estimation conditions adjusted by calibration, comprising: During calibration, singularity information that is a feature of the time-series data of the detected emotion index values is extracted and stored, and during emotion estimation, corresponding information that corresponds to the singularity information in the time-series data of the detected emotion index values is extracted, and emotion estimation is performed based on the difference between the singularity information and the corresponding information. Emotion estimation method.
10. An emotion estimation program that estimates emotions from emotion index values related to a biological state of an emotion estimation target based on estimation conditions adjusted by calibration, During calibration, singularity information that is a feature of the time-series data of the detected emotion index values is extracted and stored, and during emotion estimation, corresponding information that corresponds to the singularity information in the time-series data of the detected emotion index values is extracted, and emotion estimation is performed based on the difference between the singularity information and the corresponding information. Emotion estimation program.
11. a biosensor for detecting a biometric state of a person who is a target of emotion estimation; an emotion estimation device that estimates emotions from emotion index values related to the biological state of an emotion estimation target person based on estimation conditions adjusted by calibration; Including, the emotion estimation device, During calibration, singularity information that is a feature of the time-series data of the detected emotion index values is extracted and stored, and during emotion estimation, corresponding information that corresponds to the singularity information in the time-series data of the detected emotion index values is extracted, and emotion estimation is performed based on the difference between the singularity information and the corresponding information. Emotion estimation system.
Citation Information
Patent Citations
Determination device, work system, and determination method
JP2023071507A