Emotion estimation device and emotion estimation method
By generating facial images using a camera mounted on a vehicle, setting a threshold for facial features and standardizing features greater than the threshold, the problem of needing to obtain specific expression images in advance in existing technologies is solved, enabling appropriate emotion estimation even in the absence of facial expressions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, when inferring emotions by detecting facial feature points of a vehicle driver, it is necessary to obtain facial images containing specific expressions in advance, and it is difficult to properly infer emotions when there is no expression.
Facial images are generated by mounting a camera on the vehicle. A threshold for facial feature quantities is set, and feature quantities greater than the threshold are identified and standardized. The standardized feature quantities are then used to infer the driver's emotions.
It enables the appropriate inference of a driver's emotions even when the driver is expressionless, thus improving the accuracy of emotion inference.
Smart Images

Figure CN116363634B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to an emotion estimation device and an emotion estimation method that estimate an emotion of a driver of a vehicle. BACKGROUND
[0002] A technique is known in which an emotion is estimated based on a facial feature amount that indicates a departure from a state of a reference from each part of a face, which is detected from a facial image that indicates a face of a person. A driving control of a vehicle can be changed in accordance with an emotion of a driver of the vehicle that is estimated from a facial image that indicates a face of the driver.
[0003] Patent Document 1 describes an image processing device that calculates a variation amount of a feature point in each of a predetermined part group of a face in an image from a variation amount of the feature point in a face of a predetermined expression, normalizes the variation amount in accordance with a variation in a size of the face and a rotation of the face, and judges an expression of the face based on the normalized variation amount.
[0004] PRIOR ART DOCUMENTS
[0005] Patent Document 1: Japanese Patent Application Publication No. 2005-056388 SUMMARY
[0006] PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] According to the image processing device of Patent Document 1, a feature point of a predetermined part group that becomes a reference is detected from an image that contains a face of a predetermined expression in advance, and the feature point is compared with a feature point detected from an image, thereby judging an expression of the face. It is troublesome to obtain an image that contains a face of a specific expression in advance from a person who becomes an object of judgment. In addition, sometimes, due to an influence of a feature point detected from an image when the person who becomes the object of judgment has no expression, it is difficult to appropriately estimate an emotion.
[0008] An object of the present disclosure is to provide an emotion estimation device that can appropriately estimate an emotion of a driver of a vehicle.
[0009] TECHNICAL SOLUTION FOR SOLVING THE PROBLEM
[0010] The mood estimation device according to the present disclosure includes: a threshold value setting section that sets a face feature value threshold value based on a plurality of face feature values that indicate a magnitude of deviation from a reference state of a predetermined portion of a face, which are respectively detected from a plurality of face images that represent the face of a driver of a vehicle, which are generated by an imaging section mounted on the vehicle at times included in a predetermined time range; a face feature value determination section that determines a face feature value that is greater than the face feature value threshold value, from among a plurality of face feature values that are respectively detected from a plurality of face images that are generated by the imaging section at times not included in the predetermined time range; and a mood estimation section that estimates a mood of the driver using a normalized feature value obtained by normalizing the determined face feature value downward.
[0011] In the mood estimation device according to the present disclosure, it is preferable that the predetermined time range include a time after the mood estimation device has just been started.
[0012] In the mood estimation device according to the present disclosure, it is preferable that the threshold value setting section set, as the face feature value threshold value, a value of a face feature value located at a position of a predetermined ratio from a minimum value when a plurality of face feature values that are respectively detected from a plurality of face images generated at times included in the predetermined time range are arranged in ascending order.
[0013] In the mood estimation device according to the present disclosure, it is preferable that the threshold value setting section set the face feature value threshold value for each of a plurality of face motion units included in the plurality of face images, and the face feature value determination section determine, for each of the plurality of face motion units, a face feature value that is greater than the face feature value threshold value set for the face motion unit.
[0014] The mood estimation method according to the present disclosure includes: setting a face feature value threshold value based on a plurality of face feature values that indicate a magnitude of deviation from a reference state of a predetermined portion of a face, which are respectively detected from a plurality of face images that represent the face of a driver of a vehicle, which are generated by an imaging section mounted on the vehicle at times included in a predetermined time range; determining a face feature value that is greater than the face feature value threshold value, from among a plurality of face feature values that are respectively detected from a plurality of face images that are generated by the imaging section at times not included in the predetermined time range; and estimating a mood of the driver using a normalized feature value obtained by normalizing the determined face feature value downward.
[0015] According to the mood estimation device according to the present disclosure, it is possible to appropriately estimate a mood of a driver of a vehicle. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a schematic configuration diagram of a vehicle on which a mood estimation device is mounted.
[0017] Figure 2 This is a hardware schematic diagram of an emotion estimation device.
[0018] Figure 3 This is a functional block diagram of the processor in the emotion estimation device.
[0019] Figure 4 (a) is a diagram illustrating multiple facial images. Figure 4 (b) is a diagram illustrating the multiple facial feature quantities detected from multiple facial images.
[0020] Figure 5 (a) is a graph showing the distribution of facial feature quantities over a predetermined time period. Figure 5 (b) indicates that it is related to Figure 5 The distribution of facial feature quantities in different time ranges in (a) is shown in the figure. Figure 5 (c) is a graph representing the distribution of facial features obtained by standardizing facial features larger than the facial feature threshold downwards.
[0021] Figure 6 This is a flowchart for emotion presumption processing.
[0022] Label Explanation
[0023] 1 vehicle
[0024] 3. Emotion estimation device
[0025] 331 Threshold Setting Section
[0026] 332 Facial Feature Measurement Determination Section
[0027] 333 Emotional Prediction Department Detailed Implementation
[0028] The following is a detailed description, with reference to the accompanying drawings, of an emotion estimation device capable of appropriately estimating the emotions of a vehicle driver. The emotion estimation device sets a facial feature quantity threshold based on facial feature quantities detected from facial images representing the driver's face generated by a camera unit mounted on the vehicle within a predetermined time range, which express the magnitude of the deviation from a state that is a reference to a predetermined part of the face. Furthermore, the emotion estimation device determines facial feature quantities from facial images detected by the camera unit at times not included in the predetermined time range that are greater than the facial feature quantity threshold. Finally, the emotion estimation device uses a standardized feature quantity obtained by downwardly standardizing the determined facial feature quantity to estimate the driver's emotions.
[0029] Figure 1 This is a schematic diagram of a vehicle equipped with an emotion estimation device.
[0030] The vehicle 1 has a driver monitoring camera 2 and an emotion estimation device 3. The driver monitoring camera 2 and the emotion estimation device 3 are communicably connected via an in-vehicle network that complies with a standard such as Controller Area Network.
[0031] The driver monitoring camera 2 is one example of a photographing section for generating a face image that represents a face of a driver of the vehicle. The driver monitoring camera 2 has a two-dimensional detector that is constituted by an array of photoelectric conversion elements such as CCD or C-MOS that have sensitivity to infrared light, and an imaging optical system that images an image of a region that becomes a photographing target on the two-dimensional detector. In addition, the driver monitoring camera 2 has a light source that emits infrared light. The driver monitoring camera 2 is installed on a front upper portion of a vehicle cabin, for example, toward a face of a driver who sits on a driver seat. The driver monitoring camera 2 irradiates the driver with infrared light at a predetermined photographing cycle (for example, 1 / 30 to 1 / 10 seconds), and outputs an image that represents the face of the driver.
[0032] The emotion estimation device 3 is an ECU (Electronic Control Unit) that has a communication interface, a memory, and a processor. The emotion estimation device 3 estimates an emotion of a driver of the vehicle 1 using an image generated by the driver monitoring camera 2.
[0033] The vehicle 1 causes each section of the vehicle 1 to act appropriately in accordance with an emotion of the driver estimated by the emotion estimation device 3. For example, in a case where uneasiness is detected as the emotion of the driver, the vehicle 1 reduces a travel speed in automatic driving based on a travel control device that is not illustrated. In addition, in a case where boredom is detected as the emotion of the driver, the vehicle 1 outputs music that the driver likes from a speaker that is not illustrated.
[0034] Figure 2 is a hardware diagram of the emotion estimation device 3. The emotion estimation device 3 has a communication interface 31, a memory 32, and a processor 33.
[0035] The communication interface 31 is one example of a communication section, and has a communication interface circuit for connecting the emotion estimation device 3 to the in-vehicle network. The communication interface 31 provides received data to the processor 33. In addition, the communication interface 31 outputs data provided from the processor 33 to the outside.
[0036] The memory 32 has a volatile semiconductor memory and a non-volatile semiconductor memory. The memory 32 holds various data used in processing by the processor 33, such as parameters of a neural network used as a recognizer that detects facial feature amounts from facial images, and the like. In addition, the memory 32 holds various application programs, such as an emotion estimation program that executes emotion estimation processing, and the like.
[0037] The processor 33 is one example of a control unit, and has one or more processors and peripheral circuits thereof. The processor 33 can also have other arithmetic circuits such as a logic operation unit, a numerical operation unit, or a graphics processing unit.
[0038] Figure 3 is a functional block diagram of the processor 33 that the emotion estimation device 3 has.
[0039] The processor 33 of the emotion estimation device 3 has a threshold value setting section 331, a facial feature amount determination section 332, and an emotion estimation section 333 as functional blocks. These sections that the processor 33 has are functional modules installed by a program executed on the processor 33. Alternatively, these sections that the processor 33 has can also be installed in the emotion estimation device 3 as independent integrated circuits, microprocessors, or firmware.
[0040] The threshold value setting section 331 sets a facial feature amount threshold value based on facial feature amounts that represent the magnitude of deviation from a reference state of a predetermined portion of a face, which are detected from facial images that represent the face of the driver of the vehicle, which are generated at times included in a predetermined time range by the photographing section mounted on the vehicle. In the present embodiment, the threshold value setting section 331 acquires a plurality of facial images that represent the face of the driver of the vehicle 1, which are generated at times included in a predetermined time range by the driver monitoring camera 2. The threshold value setting section 331 detects the positions of facial action units by inputting the acquired plurality of facial images to a recognizer that is learned to detect facial action units (FAU: Facial Action Unit) that act in conjunction with changes in emotion in the face, such as the outer corner of the eye, the inner corner of the eye, and the corner of the mouth. Furthermore, the threshold value setting section 331 detects the amount of change in the positions of the facial action units from the positions of the facial action units in the model of the face that becomes the reference from the positions of the facial action units detected from the plurality of facial images as a plurality of facial feature amounts. The facial action unit is one example of a predetermined portion of the face.
[0041] The recognizer can be configured, for example, as a convolutional neural network (CNN) with multiple convolutional layers connected in series from the input side to the output side. By pre-feeding facial images containing facial action units as teacher data into the CNN and learning from them, the CNN acts as a recognizer to determine facial action units and perform actions.
[0042] The threshold setting unit 331 sets the value of the facial feature quantity that is at a predetermined proportion (e.g., 95%) from the minimum value when multiple facial feature quantities detected from multiple facial images are arranged in ascending order.
[0043] The threshold setting unit 331 can set a facial feature quantity threshold based on facial feature quantities detected from multiple facial images generated at times encompassing the time range after the start of the emotion estimation device 3. The time range encompassing the time range after the start of the emotion estimation device 3 is, for example, the range from the start of the emotion estimation device 3 to 5 minutes later.
[0044] Figure 4 (a) is a diagram illustrating multiple facial images. Figure 4 (b) is a diagram illustrating the multiple facial feature quantities detected from multiple facial images.
[0045] like Figure 4 As shown in (a), facial images P1 to Pn are generated within the time range of time t1 to time tn. Multiple facial feature quantities are detected from the multiple facial images generated in this way.
[0046] Multiple facial feature quantities corresponding to multiple facial action units are detected from each facial image. In this embodiment, the facial feature quantity for a certain facial action unit is detected as a real number greater than or equal to 0 and less than 10. Figure 4 In (b), for example, facial features corresponding to FAU1, FAU2, ..., FAU9 are detected from a facial image generated at time t1. Thus, the facial features detected from multiple facial images generated within a certain time range can be represented as tensors.
[0047] Figure 5 (a) is a graph representing the distribution of facial feature quantities over a predetermined time period.
[0048] Figure 5The graph 500 of (a) indicates, for the plurality of facial feature amounts regarding one facial action unit detected from the plurality of facial images generated in the predetermined time range, the number of facial feature amounts included in each range of facial feature amounts of the bar in the vertical axis direction. That is, for example, the height of the bar indicating the position of "0" in the horizontal axis indicates the number of facial feature amounts having a value of 0 or more and less than 1 among the plurality of facial feature amounts detected from the plurality of facial images generated in the predetermined time range.
[0049] The threshold setting section 331 sets, as the facial feature amount threshold, the value of the facial feature amount at the position in ascending order of the detected plurality of facial feature amounts at a predetermined ratio from the smallest value. In the example of (a) of FIG. 5, 7 is set as the facial feature amount threshold, which is the value of the facial feature amount at the position becoming 0 to 95%. Figure 5
[0050] In the example of (a) of FIG. 5, the facial feature amount determination section 332 determines the facial feature amount having a value of 7 or more among the plurality of facial feature amounts regarding one facial action unit generated in the time range becoming the object. Figure 5 The example of the facial feature amount regarding one facial action unit is shown in (a) of FIG. 5, but the facial feature amount determination section 331 also similarly determines the facial feature amount regarding other facial action units.
[0051] The facial feature amount determination section 332 acquires the facial feature amount greater than the facial feature amount threshold from among the plurality of facial feature amounts respectively detected by the driver monitoring camera 2 from the plurality of facial images generated at the time not included in the predetermined time range.
[0052] Figure 5 (b) of FIG. 5 is a graph indicating the distribution of the facial feature amount in the time range different from the time range of (a) of FIG. 5. Figure 5 (b) of FIG. 5 is a graph indicating the distribution of the facial feature amount in the time range different from the time range of (a) of FIG. 5.
[0053] In the example of (b) of FIG. 5, the facial feature amount determination section 332 determines the facial feature amount having a value greater than 7 from among the plurality of facial feature amounts regarding one facial action unit generated in the time range becoming the object. Figure 5 The example of the facial feature amount regarding one facial action unit is shown in (b) of FIG. 5, but the facial feature amount determination section 332 also similarly determines the facial feature amount regarding other facial action units.
[0054] Figure 5 The emotion estimation section 333 estimates the emotion of the driver of the vehicle 1 by inputting a standardized feature amount obtained by standardizing the determined facial feature amount downward to an estimator learned so as to estimate the emotion of a person represented based on the facial feature amount detected from an image representing the face of the person having an expression.
[0055] The emotion estimation section 333 estimates the emotion of the driver of the vehicle 1 by inputting a standardized feature amount obtained by standardizing the determined facial feature amount downward to an estimator learned so as to estimate the emotion of a person represented based on the facial feature amount detected from an image representing the face of the person having an expression.
[0056] Figure 5 (c) is a graph indicating a distribution of the face feature amount obtained by normalizing the face feature amount larger than the face feature amount threshold value downward.
[0057] The emotion estimation section 333 changes each value of the determined face feature amount so that a range from the face feature amount threshold value to a maximum value of the acceptable face feature amount corresponds to a range from a minimum value of the acceptable face feature amount to the maximum value of the acceptable face feature amount. Figure 5 In the example of (c) of FIG. 7, the emotion estimation section 333 changes each value of the face feature amount determined to be included in the range of 7-9 in (b) of FIG. 7 so that the range of 7-9 corresponds to the range of 0-9. Figure 6
[0058] The estimator can be provided as, for example, a neural network having a plurality of layers connected in series from an input side to an output side. The CNN acts as an estimator that estimates an emotion by inputting a face feature amount detected from an image representing a face of a person with an expression as teacher data to the neural network and performing learning.
[0059] is a flowchart of the emotion estimation process. The emotion estimation device 3 executes the emotion estimation process each time it is activated.
[0060] First, the threshold value setting section 331 of the emotion estimation device 3 sets a face feature amount threshold value based on a plurality of face feature amounts respectively detected from a plurality of face images generated by the driver monitoring camera 2 at times included in a predetermined time range after the emotion estimation device 3 is activated (step S1).
[0061] Next, the face feature amount determination section 332 of the emotion estimation device 3 determines a face feature amount larger than the face feature amount threshold value from among a plurality of face feature amounts respectively detected from face images generated by the driver monitoring camera 2 at times not included in the predetermined time range (step S2).
[0062] Then, the emotion estimation section 333 of the emotion estimation device 3 generates a normalized feature amount by normalizing the determined face feature amount downward (step S3), and estimates an emotion of the driver of the vehicle 1 using the normalized feature amount (step S4).
[0063] After the process of step S4 is executed, the process of the emotion estimation device 3 returns to step S2, and the processes of steps S2 to S4 are repeated.
[0064] By the emotion estimation device 3 performing the emotion estimation process as described above, the emotion estimation device 3 is not easily affected by the facial feature amount whose value is small detected from the facial image at the time when the driver is expressionless, and thus, it is possible to appropriately estimate the emotion of the driver of the vehicle.
[0065] It is to be understood that various alterations, modifications, and improvements can be made to the disclosed subject matter without departing from the spirit and scope thereof.
Claims
1. An emotion estimation device comprising: a threshold value setting section that sets a facial feature value threshold value based on a plurality of facial feature values that represent a magnitude of deviation from a reference state of a predetermined portion of a face of a driver of a vehicle, the plurality of facial feature values being detected from a plurality of facial images that represent the face of the driver, the plurality of facial images being generated by an imaging section mounted on the vehicle at times included in a predetermined time range; a facial feature value determination section that determines a facial feature value that is greater than the facial feature value threshold value, from among a plurality of facial feature values that are detected from a plurality of the facial images generated by the imaging section at times not included in the predetermined time range; and an emotion estimation section that estimates an emotion of the driver using a normalized feature value obtained by changing each value of the determined facial feature value so that a range from the facial feature value threshold value to a maximum value of the facial feature value corresponds to a range from a minimum value of the facial feature value to the maximum value of the facial feature value.
2. The emotion estimation device according to claim 1, wherein the predetermined time range includes a time after the emotion estimation device is just started.
3. The emotion estimation device according to claim 1 or 2, wherein the threshold value setting section sets, as the facial feature value threshold value, a value of the facial feature value that is located at a position of a predetermined proportion from a minimum value when a plurality of the facial feature values that are detected from a plurality of the facial images generated at times included in the predetermined time range are arranged in ascending order.
4. The emotion estimation device according to any one of claims 1 to 3, wherein the threshold value setting section sets the facial feature value threshold value for each of a plurality of facial motion units included in each of the plurality of facial images, and the facial feature value determination section determines, for each of the plurality of facial motion units, a facial feature value that is greater than the facial feature value threshold value set for the facial motion unit.
5. An emotion estimation method comprising: setting a facial feature value threshold value based on a plurality of facial feature values that represent a magnitude of deviation from a reference state of a predetermined portion of a face of a driver of a vehicle, the plurality of facial feature values being detected from a plurality of facial images that represent the face of the driver, the plurality of facial images being generated by an imaging section mounted on the vehicle at times included in a predetermined time range; determining a facial feature value that is greater than the facial feature value threshold value, from among a plurality of facial feature values that are detected from a plurality of the facial images generated by the imaging section at times not included in the predetermined time range; and estimating an emotion of the driver using a normalized feature value obtained by changing each value of the determined facial feature value so that a range from the facial feature value threshold value to a maximum value of the facial feature value corresponds to a range from a minimum value of the facial feature value to the maximum value of the facial feature value.
Citation Information
Patent Citations
Image processing apparatus, image processing method and imaging device
JP2005056388A
Image Processing System for Extracting a Behavioral Profile from Images of an Individual Specific to an Event
US20210326586A1