Learning data generation device, learning device, emotion estimation device, learning data generation method, and learning data generation program
The training data generation device efficiently generates high-quality learning data for AI by extracting and binarizing data within specific thresholds, addressing the challenge of generating accurate AI training data.
Patent Information
- Application Number
- JP2023199914
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-06
AI Technical Summary
Generating highly accurate AI requires a large amount of high-quality training data, and existing technologies lack efficient methods for generating suitable training data.
A training data generation device that extracts data equal to or greater than a first threshold and data equal to or less than a second threshold, and generates training data from these extracted values and corresponding input values.
This approach efficiently generates learning data suitable for AI training, improving the accuracy of AI models by excluding data with low reliability and balancing data distribution.
Smart Images

Figure 2025086098000001_ABST
Abstract
Description
[Technical field]
[0001] The disclosed embodiments relate to a training data generation device, a learning device, a feeling estimation device, a training data generation method, and a training data generation program. [Background technology]
[0002] Various devices using AI (Artificial Intelligence) have been put to practical use. In order to generate AI, it is necessary to generate a learning data set according to the AI's target function and train a pre-learning AI model using the learning data. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2020-119322 A Summary of the Invention [Problem to be solved by the invention]
[0004] To generate highly accurate AI, a large amount of high-quality training data is required, and technology that can efficiently generate training data suitable for learning is desirable.
[0005] The present invention has been made in consideration of the above, and has an object to provide a training data generation device, a learning device, an emotion estimation device, a training data generation method, and a training data generation program that are capable of efficiently generating training data suitable for learning. [Means for solving the problem]
[0006] In order to solve the above-mentioned problems and achieve the object, the training data generation device of the present invention is a training data generation device for training data in which the binarized value of first data is a correct value and second data corresponding to the first data is an input value, the device extracts the first data that is equal to or greater than a first threshold and the first data that is equal to or less than a second threshold that is smaller than the first threshold, and generates training data from the extracted first data and the second data that corresponds to the first data. Effect of the Invention
[0007] According to the present invention, learning data suitable for learning can be efficiently generated. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an outline of the emotion estimation system. [Diagram 2] FIG. 2 is a diagram illustrating an outline of the emotion estimation method. [Diagram 3] FIG. 3 is a diagram showing a specific example of the index estimation model. [Figure 4] FIG. 4 is an explanatory diagram regarding the suitability of adoption as learning data. [Diagram 5] FIG. 5 is a block diagram of the training data generating device. [Figure 6] FIG. 6 is a diagram illustrating an example of the learning data storage unit. [Figure 7] FIG. 7 is a diagram showing an example of the distribution. [Figure 8] FIG. 8 is a diagram showing an example of the distribution. [Figure 9] FIG. 9 is a block diagram of the learning device. [Figure 10] FIG. 10 is a block diagram of the emotion estimation device. [Figure 11] FIG. 11 is a flowchart showing a processing procedure executed by the training data generating device. [Figure 12] FIG. 12 is a flowchart showing a processing procedure executed by the learning device. [Figure 13]FIG. 13 is a flowchart showing a processing procedure executed by the feeling estimation device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, a training data generation device, a learning device, an emotion estimation device, a training data generation method, and a training data generation program according to the present application (hereinafter, referred to as "embodiments") will be described in detail with reference to the drawings. Note that the training data generation device, the learning device, the emotion estimation device, the training data generation method, and the training data generation program according to the present application are not limited to the embodiments.
[0010] First, an overview of the emotion estimation system according to the embodiment will be described with reference to Fig. 1. Fig. 1 is a schematic configuration diagram of the emotion estimation system according to the embodiment. As shown in Fig. 1, the emotion estimation system S includes a plurality of learning data generation devices 1, a learning device 10, and a plurality of emotion estimation devices 100.
[0011] The emotion estimation system S uses an index estimation model to estimate an emotion based on two emotion index values, which are indexes indicating a mental and physical state related to the emotion of a user. The learning data generation device 1 is a device that generates learning data for the index estimation model. For example, the learning data generation device 1 is a PC (Personal Computer) or the like installed in a laboratory or the like for collecting learning data or in a real environment where emotion estimation is performed.
[0012] The learning device 10 is a device that learns (generates) an index estimation model based on generated learning data, and is configured, for example, with a PC, a server, etc. The index estimation model is, for example, AI (Artificial Intelligence).
[0013] The feeling estimation device 100 is a device that estimates the user's feeling by using the index estimation model generated by the learning device 10. For example, the feeling estimation device 100 is a drive recorder mounted on each vehicle, and estimates the driver's feeling from an image capturing the driver's face. The feeling estimated by the feeling estimation device 100, for example, the feeling of the driver of the vehicle, is used in a process of determining advice content in driving assistance (for example, when the driver is excited, advice to calm the driver is output).
[0014] Next, an outline of the emotion estimation method will be described with reference to Fig. 2. Fig. 2 is a diagram for explaining the outline of the emotion estimation method. First, prior to the emotion estimation method according to the present application, prerequisites for the emotion estimation method will be described.
[0015] The prerequisite emotion estimation method estimates emotion based on two emotion index values, which are indices indicating the user's emotional and physical state. One of the emotion index values used in this embodiment is the central nervous system arousal level (hereinafter also referred to as arousal level), whose index value can be calculated using "β waves / α waves of the brainwaves." The other emotion index value is the autonomic nervous system activity level (hereinafter also referred to as activity level), whose index value can be calculated using the "standard deviation of the heartbeat LF (Low Frequency) component (low frequency component of the heartbeat waveform signal)."
[0016] As shown in the upper part of Fig. 2, the prerequisite emotion estimation method estimates the user's emotion from the brainwaves and heart rate in the following procedure: First, using the above-mentioned calculation formula, the central nervous system arousal level (corresponding to the above-mentioned arousal level) is calculated from the brainwaves, and the autonomic nervous system activity level (corresponding to the above-mentioned activity level) is calculated from the heart rate.
[0017] Next, the calculated activity level and activity level are each coded (stratified using a threshold), the awakening level is converted into an awakening level coded value, and the activity level is converted into an activity level coded value.
[0018] Next, emotions are estimated by plotting the obtained arousal level encoded value and activity level encoded value on a psychological plane. As shown in the upper part of Figure 2, the psychological plane (psychological plane model) defines "arousal level (awakened-unaroused)" on the vertical axis and "activity level of the autonomic nervous system (sympathetic nervous activity (strong emotion)-parasympathetic nervous activity (weak emotion)" on the horizontal axis.
[0019] In this psychological plane, the corresponding psychological states are assigned to each of the four quadrants separated by the vertical and horizontal axes (corresponding to the thresholds in encoding). The first quadrant is assigned to the psychological states of "fun, joy, anger, and sadness." The second quadrant is assigned to the psychological state of "melancholy." The third quadrant is assigned to the psychological state of "relaxation and calm." The fourth quadrant is assigned to the psychological states of "anxiety, fear, and discomfort." The index values of the two types of mental and physical states (encoded arousal value and encoded activity value) are plotted on the psychological plane to obtain coordinates (quadrant), from which the psychological state can be estimated. Specifically, the psychological state can be estimated based on which quadrant of the psychological plane the plotted coordinates are in.
[0020] As shown in the upper part of FIG. 2, the psychological plane may include an uncertain area Ar near each axis. The uncertain area Ar is an area where the psychological state is uncertain. In other words, if the coordinates of the arousal level and activity level plotted on the psychological plane are in the uncertain area Ar, the psychological state is uncertain. Specifically, the coding of the arousal level and activity level is converted into three values, positive, uncertain, and negative, respectively, and the coded values are plotted on a psychological plane including the uncertain area Ar (a psychological plane of nine areas (assigned to a type of feeling) where the arousal level and activity level are divided into positive, uncertain, and negative, respectively. Note that an uncertain emotion is assigned to an area where the coded value of the arousal level or activity level is uncertain)).
[0021] In this way, by setting the uncertainty area Ar for the psychological plane, it is possible to suppress erroneous determination of the psychological state, and to improve the accuracy of estimation of the psychological state.
[0022] In such emotion estimation, the user needs to wear an electroencephalogram sensor or a heart rate sensor. However, in an actual product, it is not preferable from the viewpoint of practicality for the user to always wear an electroencephalogram sensor or a heart rate sensor.
[0023] Therefore, in this embodiment, in the learning stage of the index estimation model, the degree of arousal obtained from the electroencephalogram sensor, the degree of activity obtained from the heart rate sensor, and image data obtained from a camera, which is a non-contact sensor, can be estimated from the user's facial expression characteristics that are considered to be correlated with emotions (emotion index values), such as gaze and face direction, and the relationship with the model is learned by the index estimation model. Note that image data other than gaze and face direction that are correlated with emotions, such as facial expression, blinking information, and various facial image features, can also be used.
[0024] In this embodiment, the index estimation model is a model that learns the relationship between emotions (emotion index values) and the user's gaze and facial direction. As shown in Fig. 2A, the index estimation model is a model that outputs an alertness coding value and an activity coding value from the user's gaze and facial direction obtained by image recognition of a camera image, or as shown in Fig. 2B, a model that outputs a central nervous system alertness level and an autonomic nervous system activity level from the user's gaze and facial direction.
[0025] In this way, by using the index estimation model, it is possible to estimate the awakening level encoded value and the activity level encoded value from the user's line of sight and face direction. Therefore, it is possible to estimate the user's emotions without the user having to wear an electroencephalogram sensor or a heart rate sensor.
[0026] Here, the index estimation model will be further described with reference to Fig. 3. Fig. 3 is a diagram showing a specific example of the index estimation model. First, the learning process of the index estimation model will be described. As shown in Fig. 3, in the learning stage of the index estimation model, learning data is generated from the output results of an electroencephalogram sensor and a heart rate sensor attached to a subject, and a camera that captures an image of the subject.
[0027] The electroencephalogram sensor data output from the electroencephalogram sensor is converted into a wakefulness level and then encoded, and the heartbeat sensor data output from the heartbeat sensor is converted into an activity level and then encoded.
[0028] The encoded arousal level coded value and the activity level coded value are used as correct values and as learning data for the index estimation model M. In this embodiment, as will be described later, the arousal level coded value and the activity level coded value are stratified by a threshold value, so that learning data suitable for learning can be efficiently generated.
[0029] Furthermore, the camera image output from the camera is subjected to image processing to identify the gaze and face direction, and the identified gaze and face direction are used as input values for learning data of the index estimation model M. In this way, the learning data of the index estimation model M is data in which the gaze and face direction are associated with the awakening degree encoded value and activity degree encoded value of the subject. Note that in the same learning data, the original brain waves, heart rate, and camera images are data acquired at the same timing.
[0030] After learning, the index estimation model M is then able to output an alertness coded value and an activity coded value in response to input values of the user's gaze and facial direction identified from the camera image input from the camera.
[0031] However, it is necessary to generate more reliable training data for the index estimation model M. For example, as described above, the training data may include cases where the distribution of the data amount for each label (correct value) varies, or where ambiguous data exists near the threshold value when encoding is included.
[0032] To address this issue, it is necessary to extract data with high reliability for each label in a balanced manner. Therefore, the training data generation device 1 according to the embodiment generates training data by excluding data with ambiguous label reliability.
[0033] Fig. 4 is an explanatory diagram regarding the suitability of adoption as learning data. Fig. 4 shows an example of time-series change in the degree of arousal calculated based on electroencephalograms. It also shows a state in which the degree of arousal is stratified into a large area A1, a medium area A2, and a small area A3.
[0034] The large area A1 is an area showing high values of arousal level, and is an area where the subject's arousal level is high. The small area A3 is an area showing low values of arousal level, and is an area where the subject's arousal level is low. On the other hand, the medium area A2 is an area showing medium values of arousal level. Therefore, it can be said that the peaks included in the large area A1 and the small area A3 are data suitable for learning.
[0035] The training data generation device 1 decreases the first threshold T1 for the large region A1 and the medium region A2 while increasing the second threshold T2 for the medium region A2 and the small region A3 until a stopping condition described later is satisfied. That is, the training data generation device 1 expands the large region A1 and the small region A3 toward the medium region A2 until the stopping condition is satisfied. The training data generation device 1 accumulates each piece of data collected during a training data collection period set by an operator or the like, and performs these processes on many pieces of data collected after data collection.
[0036] When the stopping condition is satisfied, the training data generation device 1 extracts data included in the large area A1 and the small area A3 to generate training data. For example, the training data generation device 1 sets the stopping condition as the number of data included in the large area A1 and the small area A3 reaching a threshold value.
[0037] In this case, data is extracted from the large region A1 in order of increasing arousal level, and data is extracted from the small region A3 in order of decreasing arousal level. The data extracted here indicates data from a section where the arousal level was higher than the first threshold T1 for several seconds or more, or data from a section where the arousal level was lower than the second threshold T2 for several seconds or more. In other words, the learning data generation device 1 extracts, as a single piece of data, data from a section where the arousal level was included in the large region A1 for a predetermined time or a section where the arousal level was included in the small region A3 for a predetermined time.
[0038] The learning data generating device 1 then generates learning data consisting of correct answer values obtained by binarizing the awakening level into the large area A1 and the small area A3, and input values of the line of sight and face direction corresponding to each time.
[0039] Specifically, the learning data generation device 1 binarizes the data extracted from the large area A1 as an awake state and the data extracted from the small area A3 as a non-awake state, and generates learning data in which these binarized values are regarded as correct values and second data (gaze and facial direction) corresponding to each EEG measurement time are used as input values.
[0040] In other words, the training data generation device 1 can generate training data excluding data with low reliability of labels (correct values) by generating training data excluding data that is neither a clear awake state nor a clear non-awake state. In addition, since the training data generation device 1 generates training data in order of increasing and decreasing arousal levels, it is also possible to suppress variation in the distribution of data amounts for each label.
[0041] Therefore, the learning data generation device 1 according to the embodiment can efficiently generate learning data suitable for learning.
[0042] Next, a configuration example of the training data generating device 1 according to the embodiment will be described with reference to Fig. 5. Fig. 5 is a block diagram of the training data generating device 1.
[0043] 5, the training data generation device 1 includes a control unit 2 and a storage unit 3. The storage unit 3 is realized by a storage device such as a read only memory (ROM), a random access memory (RAM), a flash memory, etc. For example, the storage unit 3 includes a training data storage unit.
[0044] Here, a specific example of the learning data storage unit will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the learning data storage unit. As shown in Fig. 7, the learning data storage unit of the storage unit 3 stores information on items such as "brain wave", "heart rate", "face direction", and "gaze" in association with "time" (data measurement time).
[0045] "Time" indicates the time when each data (such as brain waves) for generating the corresponding learning data was measured. "Breast waves" indicates the brain waves measured at the corresponding time, and "heart rate" indicates the heart rate measured at the corresponding time. As shown in FIG. 6, "brain waves" and "heart rate" each have "input value" and "correct value" items.
[0046] The "input value" is the value (measurement value) of the brainwave or heart rate, and the "correct value" is the binarized value (i.e., the coded value) of the alertness level obtained from the brainwave or the activity level obtained from the heart rate. Note that a "- (hyphen)" in the input value and correct value indicates that the label (correct value) is data with low reliability, for example, data in the above-mentioned middle area A2.
[0047] "Facial direction" indicates the facial direction of the subject measured at the corresponding time, and "gaze" indicates the gaze of the subject measured at the corresponding time. The learning data storage unit of the storage unit 3 stores learning data selected by the control unit 2, which will be described later.
[0048] Returning to the explanation of FIG. 5, the control unit 2 will be explained. The control unit 2 corresponds to a so-called processor. The control unit 2 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphical Processing Unit), or the like. The control unit 2 executes a program according to an embodiment (not shown) stored in the storage unit 3, using a RAM as a working area. The control unit 2 can also be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0049] The control unit 2 generates learning data using the binarized values (alertness encoded value, activeness encoded value) of the first data (alertness, activeness) as correct values and the second data (face direction and gaze) corresponding to the first data as input values. The control unit 2 extracts the first data (alertness, activeness) equal to or greater than a first threshold and the first data (alertness, activeness) equal to or less than a second threshold that is smaller than the first threshold, and generates learning data (input values: face direction and gaze, correct values: alertness, alertness encoded value generated based on activeness, activeness encoded value) using the extracted first data (alertness, activeness) and the second data (face direction and gaze) corresponding to the first data (alertness, activeness).
[0050] The control unit 2 acquires, as first data, the brain waves and heart rate of the subject from the brain wave sensor 30 and the heart rate sensor 31. The control unit 2 also acquires a camera image of the subject's face taken by the camera 32.
[0051] The control unit 2 generates data on the facial direction and gaze of the subject from the camera image as second data by image analysis of the camera image. The facial direction and gaze correspond to an example of facial expression feature data.
[0052] Furthermore, the control unit 2 calculates the level of alertness and the level of activity from the brain waves and heartbeats acquired from the brain wave sensor 30 and the heartbeat sensor 31, and encodes the level of alertness and the level of activity.
[0053] The control unit 2 defines the amount of learning data that can be extracted from the first data, and moves the boundary between the large, medium, and small regions until the amount of learning data is reached. For example, the control unit 2 defines the amount of learning data as the number of time-series data that can be extracted.
[0054] 4, the control unit 2 gradually decreases the first threshold T1 for the large region A1 and the medium region A2, while gradually increasing the second threshold T2 for the small region A3 of the medium region A2. Then, the control unit 2 moves (decreases, increases) the first threshold T1 and the second threshold T2 until the number of data included in the large region A1 and the small region A3 reaches the amount of training data.
[0055] In addition, a method can also be applied in which the amount of learning data is a predetermined set ratio to the total number of data in the first data, and the first threshold T1 and the second threshold T2 are moved (decreased, increased) until the number of data contained in the large area A1 and the small area A3 reach their respective set ratios to the total number of data in the first data.
[0056] The control unit 2 may define the amount of learning data based on the normal distribution of the first data. Fig. 7 and Fig. 8 are diagrams showing examples of distribution. Fig. 7 shows the distribution of the activity encoded value obtained from the heart rate, and Fig. 8 shows the distribution of the awakening encoded value obtained from the brain wave. In Fig. 7, the vertical axis shows the frequency of occurrence of data, and the horizontal axis shows the numerical value of data, while in Fig. 8, the vertical axis shows the numerical value of data, and the horizontal axis shows the frequency of occurrence.
[0057] When the control unit 2 defines the amount of training data based on a normal distribution, the control unit 2 obtains the standard deviation by statistical processing of all acquired data and fits the data to the normal distribution. Then, the control unit 2 generates training data by excluding data within a range of a predetermined set value * σ from the average (range of -δ to +δ).
[0058] In the example of FIG. 7, a first threshold T1 and a second threshold T2 are set to a predetermined set value *σ from the average, and data included in a middle region A2 between the first threshold T1 and the second threshold T2 is excluded from the learning data.
[0059] Similarly, in the example of Figure 8, the first threshold T1 and the second threshold T2 are set to a predetermined set value *σ from the average, and data included in the intermediate region A2 from the first threshold T1 to the second threshold T2 is excluded from the learning data.
[0060] In other words, the control unit 2 sets the first threshold T1 and the second threshold T2 based on the standard deviation of the first data, and generates learning data by preferentially extracting data far from the average value determined by the standard deviation. This allows the control unit 2 to generate learning data suitable for learning while suppressing the variation in the distribution of the data amount.
[0061] The control unit 2 may define the amount of learning data based on a judgment axis used when estimating emotions. Specifically, for example, the control unit 2 may determine the first threshold T1 and the second threshold T2 based on each index value based on an electroencephalogram or a heart rate observed when a subject performs a specific task.
[0062] Here, a specific task is a task that puts the subject in a specific mental and physical state (awake state, non-awake state, active state, non-active state). A task that puts the subject in an awakened state is calculation (for example, adding displayed numbers in mental arithmetic), and a task that puts the subject in a non-awakened state is resting and clenching a fist, etc. By having the subject perform these tasks, data on the brain waves and heart rate when the subject is in a specific mental and physical state can be obtained.
[0063] Then, the control unit 2 sets a first threshold T1 and a second threshold T2 based on the data obtained by the task, and binarizes the data included in the large area A1 and the data included in the small area A3. At this time, the control unit 2 may provide an offset to the first threshold T1 and the second threshold T2 in consideration of noise, etc.
[0064] The control unit 2 may also determine the first threshold T1 and the second threshold T2 based on the results of a questionnaire given to the subject. In this case, the questionnaire asks the subject about his or her own mental and physical state in each time period, and the control unit 2 determines the first threshold T1 and the second threshold T2 based on the brain waves and heart rates at the time when the subject answers as a specific mental and physical state (awake state, non-awake state, active state, inactive state).
[0065] In this way, by defining the amount of learning data based on the judgment axis used when estimating emotions, the control unit 2 can generate learning data based on data when the subject is in a specific mental and physical state, thereby generating learning data suitable for learning.
[0066] Then, the control unit 2 binarizes the first data (index value based on electroencephalogram or heartbeat) according to the presence state in the large area A1 and the small area A3. That is, if the first data is electroencephalogram (wakefulness level based on electroencephalogram), the control unit 2 encodes the data in the large area A1 into an encoded value indicating an wakefulness state and the data in the small area A3 into an encoded value indicating a non-wakefulness state, and if the first data is heartbeat (activity level based on heartbeat), the control unit 2 encodes the data in the large area A1 into an encoded value indicating an activity state and the data in the small area A3 into an encoded value indicating a non-activity state.
[0067] Then, the control unit 2 generates learning data in which the encoded value is set as a correct value and the corresponding second data is set as an input value, stores the learning data in the learning data storage unit, and passes the learning data to the learning device 10.
[0068] Next, a configuration example of the learning device 10 according to the embodiment will be described with reference to Fig. 9. Fig. 9 is a block diagram of the learning device 10. As shown in Fig. 9, the learning device 10 has a control unit 12 and a storage unit 13.
[0069] The storage unit 13 is realized by a storage device such as a ROM (Read Only Memory), a RAM (Random Access Memory), or a flash memory. The control unit 12 corresponds to a so-called processor. The control unit 12 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphical Processing Unit), or the like. The control unit 12 executes a program according to an embodiment (not shown) stored in the storage unit 13, using the RAM as a working area. The control unit 12 can also be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0070] The control unit 12 generates an index estimation model M using the learning data generated by the learning data generation device 1. The control unit 12 performs a learning process for the index estimation model M using a data set of supervised learning data consisting of ground truth values of the first data (encoded alertness and activity levels) and second data (gaze and face direction), and generates a learned index estimation model M. Note that the learning method for the index estimation model M may be any of various known methods such as an error backpropagation method.
[0071] Then, the control unit 12 stores the learned index estimation model M in the storage unit 13 and provides it to each feeling estimation device 100. Note that the control unit 12 may generate a trained model for each attribute of the subject. The attribute of the subject is a demographic attribute, etc.
[0072] For example, the control unit 12 classifies the learning data according to the attributes of the subjects, learns the index estimation model M using the learning data for each attribute, and generates a learned model for each attribute. Then, the feeling estimation device 100 uses the index estimation model M according to the attributes of the user who is the target of feeling estimation. This is expected to improve the accuracy of feeling estimation.
[0073] Furthermore, the control unit 12 may learn the index estimation model M using different second data depending on the purpose of use of the feeling estimation device 100. For example, when each feeling estimation device 100 is mounted on a vehicle and used to estimate the driver's feeling, the second data may include data related to driving such as accelerator, brake, and steering.
[0074] In this case, the learning data generation device 1 generates learning data in which data related to driving, such as the accelerator, brake, and steering, is used as input values and associated with correct values related to arousal and activity. The control unit 12 of the learning device 10 uses the learning data to learn the index estimation model M. In this case, by using the index estimation model M, it is also possible to estimate the user's emotions using the data related to driving as input values.
[0075] Next, a configuration example of the feeling estimation device 100 according to the embodiment will be described with reference to Fig. 10. Fig. 10 is a block diagram of the feeling estimation device 100. As shown in Fig. 10, the feeling estimation device 100 includes a control unit 102 and a storage unit 103.
[0076] The storage unit 103 is realized by a storage device such as a ROM, a RAM, or a flash memory. The control unit 102 corresponds to a so-called processor. The control unit 102 is realized by a CPU, an MPU, a GPU, or the like. The control unit 102 executes a program according to an embodiment (not shown) stored in the storage unit 103, using the RAM as a working area. The control unit 102 can also be realized by an integrated circuit such as an ASIC or an FPGA.
[0077] The control unit 102 estimates the user's emotion using the index estimation model M. For example, the control unit 102 estimates the user's emotion from a camera image captured by the camera 32 of the user.
[0078] Specifically, the control unit 102 acquires a camera image of the user's face input from the camera 32, and detects the user's facial direction and line of sight from the camera image. Note that the method of detecting the facial direction and line of sight may be any of various known methods.
[0079] Next, the control unit 102 uses the data related to the face direction and gaze as input values to obtain output values output from the index estimation model M. The output values are an alertness encoded value and an activity encoded value. Next, the control unit 102 plots the alertness encoded value and the activity encoded value output from the index estimation model M on a psychological plane (see FIG. 2) and estimates the user's emotion from the coordinates.
[0080] For example, when the feeling estimation device 100 is mounted on a vehicle, the feeling of the user estimated by the control unit 102 is used for vehicle control. Specifically, when the driver is excited, the feeling can be applied to control such as outputting advice to calm the driver, lowering the accelerator sensitivity (to make acceleration harder), increasing the brake sensitivity (to make stopping easier), and lowering the upper limit of maximum speed control.
[0081] Next, a process procedure executed by the learning data generating device 1 will be described with reference to Fig. 11. Fig. 11 is a flowchart showing the process procedure executed by the learning data generating device 1. Note that the process procedure shown below is repeatedly executed by the control unit 2 of the learning data generating device 1 when collecting learning data.
[0082] 11, the control unit 2 first acquires first data and second data at the same time (step S101). The control unit 2 acquires brain waves and heart rate as the first data, and acquires the subject's facial direction and line of sight as the second data.
[0083] Next, the control unit 2 defines the amount of learning data that can be extracted from the first data (step S102). The control unit 2 defines the amount of learning data by the number of data extracted in the large area A1 and the small area A3, percentiles and α intervals in a normal distribution, judgment axes used for emotion estimation, etc. The definition of the amount of learning data may be specified by an administrator, etc.
[0084] Next, the control unit 2 extracts the first data whose value is the maximum / minimum value (step S103), and determines whether or not a stop condition is satisfied (step S104). The control unit 2 determines, as the stop condition, whether or not the extracted first data has reached the learning data amount.
[0085] If the control unit 2 determines that the stop condition is satisfied (step S104; Yes), the control unit 2 proceeds to the process of step S105, and if the control unit 2 determines that the stop condition is not satisfied (step S104; No), the control unit 2 returns to the process of step S103. In this case, the control unit 2 extracts the first data of the maximum and minimum values from the data that has not yet been extracted in step S103.
[0086] Next, for each extracted data, the control unit 2 converts the first data into a correct value (step S105), sets this converted value as the correct value, and generates learning data with the corresponding second data as an input value (step S106), and then terminates the processing.
[0087] Next, the processing procedure executed by the learning device 10 will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the processing procedure executed by the learning device 10. Note that the processing procedure shown below is executed by the control unit 12 of the learning device 10 when generating the index estimation model M (for example, when an operator or the like performs an operation to start learning). Note that the index estimation model M before learning is stored in the memory unit 13 in advance by an operator or the like.
[0088] 12, the control unit 12 acquires learning data from the learning data generation device 1 (step S201). Next, the control unit 12 trains the index estimation model M before learning (completion) using the learning data (step S202). Then, the control unit 12 distributes the index estimation model M after learning (completion) to each feeling estimation device 100 (step S203), and ends the process.
[0089] Next, a processing procedure executed by the feeling estimation device 100 will be described with reference to Fig. 13. Fig. 13 is a flowchart showing the processing procedure executed by the feeling estimation device 100. Note that the processing procedure shown below is repeatedly executed by the control unit 102 while the feeling estimation device 100 is running. In addition, an index estimation model M that has been trained in advance by an operator or the like is stored (installed) in the storage unit 13 of the feeling estimation device 100.
[0090] 13, the feeling estimation device 100 acquires a camera image (Step S301). Next, the feeling estimation device 100 analyzes the user's line of sight and facial direction from the camera image (Step S302).
[0091] Next, the feeling estimation device 100 estimates the arousal level and the activity level by applying the index estimation model M (Step S303). Then, the feeling estimation device 100 applies the arousal level and the activity level estimated by the index estimation model M to the psychological plane model (see FIG. 2) to estimate the user's emotion (Step S304), and ends the process.
[0092] Incidentally, the present invention is not limited to the above-described embodiment and may be applied to other fields as long as the field involves generating learning data using correct values of first data and second data as input values.
[0093] Further advantages and modifications may readily occur to those skilled in the art. Thus, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and equivalents thereof. [Explanation of symbols]
[0094] 1. Training data generation device 2. Control section 3 Storage section 10 Learning Device 12 Control section 13 Storage section 30 Brainwave Sensor 31 Heart Rate Sensor 32 Camera 100 Emotion estimation device 102 Control section 103 Storage section
Claims
1. A learning data generating device for learning data in which a binarized value of first data is a correct answer value, and second data corresponding to the first data is an input value, extracting the first data that is equal to or greater than a first threshold and the first data that is equal to or less than a second threshold that is smaller than the first threshold; generating learning data from the extracted first data and second data corresponding to the first data; Training data generation device.
2. The first threshold is decreased and the second threshold is increased until the number of extracted first data reaches the set number of generated learning data. The training data generating device according to claim 1 .
3. The number of extracted first data items equal to or greater than the first threshold is set to be equal to the number of extracted first data items equal to or less than the second threshold. The training data generating device according to claim 2 .
4. A ratio of the number of extracted first data items equal to or greater than the first threshold value to the total number of first data items is set as a first ratio; A ratio of the number of extracted first data items that are equal to or smaller than the second threshold value to the total number of first data items is defined as a second ratio. The training data generating device according to claim 2 .
5. The first threshold and the second threshold are determined according to a normal distribution state of the first data. The training data generating device according to claim 1 .
6. The first data is a central nervous system arousal level and an autonomic nervous system activity level, The second data is facial expression feature data. The training data generating device according to any one of claims 1 to 5.
7. In a learning device that learns an AI model using learning data, The learning data is A binarized value of a first data is a correct value, and second data corresponding to the first data is an input value, The first data that is equal to or greater than a first threshold value and the first data that is equal to or less than a second threshold value that is smaller than the first threshold value are regarded as correct values; The second data corresponding to the first data set as a correct answer is used as an input value, and the second data is learning data. Learning device.
8. 1. An emotion estimation device that calculates a binarized value of a central nervous system arousal level and a binarized value of an autonomic nervous system activity level using an AI model based on facial expression data of a person, and estimates an emotion based on a combination of the binarized value of the central nervous system arousal level and the binarized value of the autonomic nervous system activity level, The AI model is Taking facial expression data of a person as input, A central nervous system arousal level equal to or higher than a first arousal level threshold is defined as a first arousal level binary value, and a central nervous system arousal level equal to or lower than a second arousal level threshold that is smaller than the first arousal level threshold is defined as a second arousal level binary value. The autonomic nervous system activity level equal to or higher than the first activity level threshold is set as a first activity level binary value, and the autonomic nervous system activity level equal to or lower than a second activity level threshold that is smaller than the first activity level threshold is set as a second activity level binary value. Emotion estimation device.
9. A method for generating learning data, the method including: setting a binarized value of first data as a correct value; and setting second data corresponding to the first data as an input value, the method comprising the steps of: extracting the first data that is equal to or greater than a first threshold and the first data that is equal to or less than a second threshold that is smaller than the first threshold; generating learning data from the extracted first data and second data corresponding to the first data; Training data generation method.
10. A program for generating learning data in which a binarized value of first data is a correct value and second data corresponding to the first data is an input value, extracting the first data that is equal to or greater than a first threshold and the first data that is equal to or less than a second threshold that is smaller than the first threshold; generating learning data from the extracted first data and second data corresponding to the first data; A learning data generation program that causes a computer to execute the above.
Citation Information
Patent Citations
Learning requesting device and method for requesting learning
JP2020119322A