Learning method, detection device, learning data generation device, and learning data generation program

The AI model improves the detection accuracy of emotional changes during near misses by using biological signals and appearance information, overcoming the challenges of contact sensors, such as discomfort and practical issues.

JP2025087114APending Publication Date: 2025-06-10DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023201541
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The challenge is to improve the detection accuracy of changes in emotions during near misses without the discomfort and practical issues associated with wearing contact sensors, such as electrodes, on individuals like drivers.

Method used

An AI model is trained using biological signals and appearance information, where the emotion index value is calculated based on biological signals, and the change amount of this value is used to generate supervised learning data. This AI model learns to detect changes in emotions by classifying the change amount as a correct value, improving detection accuracy for near misses.

Benefits of technology

The trained AI model enhances the detection accuracy of emotional changes during near misses by utilizing appearance information and biological signals, addressing the discomfort and practical issues of contact sensors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087114000001_ABST
    Figure 2025087114000001_ABST
Patent Text Reader

Abstract

To provide an AI model contributing to improvement of detection accuracy of emotion changes when near-miss and the like occur.SOLUTION: An exemplary learning method acquires a biological signal and appearance information of a subject for learning data generation at the same timing, calculates emotion index value on the basis of the acquired biological signal, calculates a quantity of change of the calculated emotion index value, generates learning data with a teacher that has the acquired appearance information as input value and a label categorizing the calculated quantity of change as correct answer value, and learns an AI model by the generated learning data with teacher.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for training an AI model, a detection device using a trained AI model, a training data generation device for generating training data used for training the AI model, and a training data generation program for causing a computer to generate training data used for training the AI model.

Background Art

[0002] Conventionally, an emotion estimation device that estimates emotions based on biological signals (for example, heartbeat and brain waves) is known (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in order to measure biological signals such as heartbeat and brain waves, wearing work such as attaching electrodes of a contact sensor (biological sensor) for detecting biological signals to the skin of the emotion estimation target person is required. Due to the trouble of wearing work and the discomfort felt by the emotion estimation target person due to wearing the contact sensor, it is not very realistic to attach the contact sensor to passengers such as drivers, especially when applying it to in-vehicle devices.

[0005]

[0006] Therefore, it is expected to adopt a method of estimating (calculating) emotions using an artificial intelligence (AI) model from the appearance information of the emotion estimation target person. Further, as an example of the usage of the emotions estimated by the emotion estimation device, detection of occurrences such as near misses is assumed. For this reason, generation of an AI model suitable for detection of occurrences such as near misses is desired.In view of the above circumstances, an object of the present invention is to provide an AI model that contributes to improving the detection accuracy of changes in emotions when a near miss or the like occurs.

Means for Solving the Problems

[0007] An exemplary learning method of the present invention acquires a biological signal and appearance information of a subject for generating learning data at the same timing, calculates an emotion index value based on the acquired biological signal, calculates a change amount of the calculated emotion index value, generates supervised learning data using the acquired appearance information as an input value and a label obtained by classifying the calculated change amount as a correct value, and learns an AI model with the generated supervised learning data.

Effects of the Invention

[0008] According to an exemplary aspect of the present invention, learning of an AI model is performed using, as a correct value of AI learning data, a label obtained by classifying a change amount of an emotion index value based on a biological signal. Since emotions in near misses or the like are presumed to have a high correlation with the change amount of the emotion, the learned AI model becomes an AI model that contributes to improving the detection accuracy of changes in emotions when a near miss or the like occurs.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 10

Mode for Carrying Out the Invention

[0010] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the content of the embodiments shown below.

[0011] <1. Conceptual Configuration of Hiccup Detection Device> First, the hiccup detection device will be described. FIG. 1 is a conceptual explanatory diagram showing the detection process by the hiccup detection device 10.

[0012] In the usage example of the hiccup detection device 10 shown in FIG. 1, the user U1 (the observed person) is the driver of the vehicle V1. The AI model 121M used in the detection process by the hiccup detection device 10 estimates the change amount of the emotion index value, which is an index indicating the mental and physical state related to the emotion of the driver (user U1). The AI model 121M is separately learned and provided as a learned model, and is installed in the hiccup detection device 10. Details of the hiccup detection device 10 will be described later.

[0013] Note that the hiccup detection device 10 can be realized by a computer device installed in the vehicle V1, or can also be realized by a server connected to the vehicle V1 via a network. Further, the server may be a physical server or a virtual server.

[0014] Vehicle V1 is equipped with a camera C and a microphone M as in-vehicle sensors. Then, the gaze data and face orientation data of user U1 (driver) based on the captured image of camera C and the voice data of user U1 (driver) acquired by microphone M are input into AI model 121M.

[0015] In the illustrated example, the behavior determination unit 132 processes the captured image of camera C through image recognition processing or the like, and processes it into gaze data (data processed into an image of the eye area or text / numerical data indicating the gaze (direction)) and face orientation data (data processed into an image of the face or text / numerical data indicating the orientation of the face), and inputs it into AI model 121M. However, the captured image of camera C may also be input. Note that an AI model 121M (an AI model designed and learned according to the input data type) corresponding to the input data type will be installed in the near-miss detection device 10.

[0016] AI model 121M is an AI model that inputs the gaze data, face orientation data, and voice data of user U1 and outputs a change amount label of the emotion index value, and is designed with a structure suitable for the input and output. And AI model 121M is learned with a dataset of a lot of learning data using the gaze data, face orientation data, and voice data as input data and the change amount label of the emotion index value corresponding to the input data as correct data. That is, AI model 121M is a learned AI model. The details of the learning method will be described later.

[0017] Then, AI model 121M outputs the change amount label of the emotion index value estimated (output) based on the gaze data, face orientation data, and voice data of user U1 to the near-miss detection unit 133. The near-miss detection unit 133 detects a near-miss that user U1 (driver) is facing based on the input change amount label of the emotion index value.

[0018] Here, the relationship between emotion and emotion index value will be explained.

[0019] The emotion estimation model is a model that estimates emotions based on emotion index values, which are indicators showing the psychosomatic state related to emotions. One of the emotion index values used in this embodiment is the central nervous system arousal level (hereinafter referred to as arousal level), and its index value can be calculated as "β wave / α wave of electroencephalogram". Another emotion index value is the autonomic nervous system activity level (hereinafter referred to as activity level), and its index value can be calculated as "standard deviation of the heart rate LF (Low Frequency) component (low frequency component of the heart rate waveform signal)".

[0020] The emotion estimation model used in this embodiment is composed of a model (calculation formula and conversion data table) for emotion estimation based on the arousal level and the activity level. The emotion estimation model used in this embodiment is a multi-dimensional model (here, a two-dimensional model with the arousal level and the activity level as two axes) that estimates emotions using the arousal level and the activity level as parameters. The two-dimensional model is created based on medical evidence (papers, etc.) showing the relationship between each of the multiple indicators and emotions (the relationship between the arousal level, activity level, and emotions). Alternatively, the two-dimensional model is created based on the questionnaire results of many subjects (data consisting of the emotion declarations by the subjects and the arousal level and activity level (based on electroencephalogram and heart rate measurement values) at that time). Note that the emotion estimation model used in this embodiment can also be a multi-dimensional model of three dimensions or more, not just a two-dimensional model.

[0021] Figure 2 is a diagram showing an example of a multi-dimensional (two-dimensional) model (psychological plane) for emotion estimation. According to various medical evidences related to psychology, it is said that psychology can be estimated based on two types of indicators (emotion index values) showing the physical state. In the psychological plane shown in Figure 2, the vertical axis is "arousal level (aroused - non-aroused)", and the horizontal axis is "activity level of the autonomic nervous system (sympathetic nerve activity (strong emotion) - parasympathetic nerve activity (weak emotion)".

[0022] In this psychological plane, the corresponding emotions (types) are assigned to each of the four quadrants separated by the vertical and horizontal axes. The distance from each axis indicates the intensity of the corresponding emotion. The emotions of "happy, joy, anger, sadness" are assigned to the first quadrant. Also, the emotion of "depression" is assigned to the second quadrant. Also, the emotions of "relaxed, calm" are assigned to the third quadrant. Also, the emotions of "uneasy, fear, unpleasant" are assigned to the fourth quadrant.

[0023] Note that the positions of the axes are appropriately set based on experiments such as measuring the arousal level and activity level of the subject and performing statistical processing.

[0024] And, based on the two types of emotion index values (arousal level and activity level) obtained from the biological signals, emotions can be estimated from the coordinates obtained by plotting them on the psychological plane. Specifically, based on which quadrant of the psychological plane the plotted coordinates are in, what position within the quadrant they are in, and what the distance from the origin is, the emotion and its intensity can be estimated. Note that the emotion estimation model shown in Figure 2 is a two-dimensional plane, but it becomes a multi-dimensional space of three dimensions or more depending on the number of indices used.

[0025] Also, when the emotion intensity is strong, that is, when the emotion index values fluctuate greatly towards the maximum value side or the minimum value side, the emotion estimation accuracy becomes high. However, when the emotion intensity is weak, that is, when the emotion index values are near the median value, the emotion estimation accuracy becomes low. For this reason, a method of making a determination such as no estimated emotion or emotion estimation impossible by setting the region near the median value of the emotion index values as a neutral region can be considered.

[0026] Figure 3 is a diagram showing an example of a psychological plane including a neutral region. In Figure 3, the regions Rn1 and Rn2 indicated by the diagonal lines are neutral regions.

[0027] The upper and lower limit values of the neutral region in the emotional index of "arousal level" are YP and YN, and the region sandwiched between the upper limit value YP and the lower limit value YN is the arousal neutral region Rn1 for "arousal level". Also, the upper and lower limit values of the neutral region in the emotional index of "activity level" are XP and XN, and the region sandwiched between the upper limit value XP and the lower limit value XN is the activity neutral region Rn2 for "activity level".

[0028] The settings of these neutral regions (upper limit value YP, lower limit value YN, upper limit value XP, and lower limit value XN) can also be appropriately set by experiments or the like.

[0029] By fitting the emotional indices "arousal level" and "activity level" to the emotional estimation model shown in the psychological plane as described above (by plotting as coordinates), the emotion can be estimated. Therefore, if the amount of change in the emotional indices "arousal level" and "activity level" is large (if the position change of the coordinates plotted on the psychological plane is large), it can be estimated that the change in emotion is large.

[0030] When the user U1 (driver) faces a near miss, since it is considered that the emotion of the user U1 (driver) changes greatly, the near miss detection device 10, which is configured to detect a near miss based on the amount of change in the emotional index value, can be expected to improve the detection accuracy of a near miss compared to a configuration that detects a near miss based on the emotional index value itself.

[0031] <2. Near Miss Detection Device> Next, the near miss detection device 10 will be described with reference to FIG. 4. In FIG. 4, the components necessary for explaining the features of the present embodiment are shown, and the description of general components is omitted. The near miss detection device 10 includes a communication unit 11, a storage unit 12, and a controller 13.

[0032] The storage unit 12 is provided with an AI model storage unit 121 that stores the learned AI model 121M shown in FIG. 1.

[0033] Note that the AI model 121M is stored in the AI model storage unit 121 as follows, for example.

[0034] When the manufacturer of the near miss detection device 10 or the like assembles the device, etc., they install the memory in which the learned AI model 121M is written as the AI model storage unit 121. Alternatively, the manufacturer of the near miss detection device 10 or the like connects an external storage device in which the learned AI model 121M is stored to the near miss detection device 10, reads the AI model 121M, and writes it to the AI model storage unit 121. Alternatively, the manufacturer of the near miss detection device 10 or the like receives the learned AI model 121M from a learning device described later via a network and writes it to the AI model storage unit 121.

[0035] The controller 13 controls various operations of the near miss detection device 10 and includes a processor that performs arithmetic processing and the like, and is configured by, for example, a CPU (Central Processing Unit). As its functions, the controller 13 includes an acquisition unit 131, a behavior determination unit 132, a near miss detection unit 133, and a provision unit 134. The functions of the controller 23 are realized by the processor executing arithmetic processing according to a program stored in the storage unit 22.

[0036] The acquisition unit 131 acquires various appearance information (image information, voice information) of the user U1, who is the observer, captured and sound-collected by the camera C and the microphone M via the communication unit 11. Here, the appearance information includes not only image data but also information detectable by non-contact sensors (remote detection sensors: camera, microphone, etc.) such as voice. The acquisition unit 131 stores the acquired various appearance information in a data table formed in the storage unit 12 as needed for subsequent processing. Note that the acquisition unit 131 acquires these appearance information at substantially the same time and stores them in the data table as one data set. And these data are used to detect the behavior state and near misses at that time.

[0037] The behavior determination unit 132 performs analysis processing on the image information and voice information of the user U1 acquired by the acquisition unit 131 to determine predetermined types of behaviors of the user U1, such as the user U1's line of sight, face orientation, expression, voice, etc. The predetermined types of behaviors are behaviors corresponding to the data types input by the AI model 121M. In the case of this embodiment, specifically, they are the line of sight, face orientation, expression, and voice (content).

[0038] Therefore, the behavior determination unit 132 will perform various processes according to the input specifications (determined by design conditions, learning conditions, etc.) of the AI model 121M. The various processes are, for example, the clipping process of images (still image groups or videos) and voice data with a processing target time length, the clipping process and orientation detection process of the user U2's face and eyeball parts from the captured image of the camera C, and the frequency analysis process of the output signal of the microphone M, etc.

[0039] And the AI model 121M outputs an appropriate estimation (correct) label for the input of data with the same type (content and style) as the input data during learning. For this reason, the output data type (content type and style) of the behavior determination unit 132 needs to be the same as the input data during the learning of the AI model 121M. The behavior determination unit 132 will perform such data processing. The programs and data for realizing these operations of the behavior determination unit 132 are stored in the storage unit 12.

[0040] The near miss detection unit 133 detects the near misses faced by the user U1 (driver) based on the labels output from the AI model 121M.

[0041] The providing unit 134 provides the information on the near misses faced by the user U1 (driver) detected by the near miss detection unit 133 to external devices of the near miss detection device 10, such as a safe driving support device installed in a vehicle (for example, vocalizing or displaying driving advice corresponding to the near miss information), or a server for providing a service for providing a near miss road map using the near miss information, etc.

[0042] <3. Processing of the Near Miss Detection Device> FIG. 5 is a flowchart showing the detection process executed by the controller 13 of the near miss detection device in FIG. 4. This flowchart shows the technical content of a computer program that causes a computer device to implement the detection process. Further, the computer program is stored in various readable non-volatile recording media and provided (sold, distributed, etc.). The computer program may be composed of only one program, or may be composed of a plurality of cooperating programs.

[0043] The process shown in FIG. 5 starts when the vehicle V1 starts and starts at the timing when the vehicle V1 becomes drivable, and is then repeatedly executed over a time period when near miss detection is required.

[0044] In step S101, the controller 13 (acquisition unit 131) acquires data (appearance information) indicating the behavior of the user U1, specifically, a camera captured image of the user U1 and microphone collected sound, and proceeds to step S102. The acquired various appearance information is stored in the data table of the storage unit 12 as necessary.

[0045] In step S102, the controller 13 (behavior determination unit 132) determines the behavior of the user U1, specifically, the line of sight, face orientation, expression, and voice (content) based on the various data (camera captured image, microphone collected sound) acquired by the acquisition unit 131 in step S101, and proceeds to step S103.

[0046] In step S103, the controller 13 (near miss detection unit 133) detects a near miss faced by the user U1 based on the label output from the AI model 121M when inputting each data based on the behavior determination result by the behavior determination unit 132 determined in step S102 as an input value, and proceeds to step S104. The storage unit 12 stores a data table showing the relationship between the type of label output from the AI model 121M and the presence or absence of near miss detection, and the controller 13 (near miss detection unit 133) uses the data table to detect a near miss.

[0047] In step S104, the controller 13 (provision unit 134) provides (outputs) the detection information of the near miss faced by the user U1 (driver) detected by the near miss detection unit 133 (presence or absence of near miss detection, detection date and time, detection position, line of sight, face orientation, camera captured image data, microphone recorded voice data, etc., request data of an external device) outside the near miss detection device 10, and ends the process. In the process shown in FIG. 5, even when a near miss is not detected in step S103, the detection information indicating that a near miss has not been detected is provided (output) outside the near miss detection device 10, but when a near miss is not detected in step S103, the provision of the detection information may be omitted.

[0048] <4. Learning Method of AI Model> Next, the learning method of the AI model 121m will be described. In order to easily distinguish between the pre-learning and post-learning models, the pre-learning model is denoted as AI model 121m, the post-learning model is denoted as AI model 121M, and the pre-learning and post-learning are distinguished by the last character of the symbol. FIG. 6 is a conceptual explanatory diagram showing the learning process of the AI model by the learning device 20. Note that the learning data generation device 30 in the figure is a part of the learning device 20 and includes an index value calculation unit 232, a behavior determination unit 233, and a learning data set generation unit 234.

[0049] In this embodiment, the user U2 is a subject for generating learning data (a subject for generating learning data), and members of the development team of the near miss detection device 10 or the like serve as the subject. In the case of the near miss detection device 10 dedicated to an individual whose user is limited, it is preferable that the subject is the user, but even if it is another person, there is a similarity, so it may be the user or another person. Further, in the case of the general-purpose near miss detection device 10 where the user is not limited, the subjects are the user and other people, but if learning data is generated by a plurality of subjects, learning data by subjects with various characteristics can be obtained. Thereby, it can be expected that the near miss detection device 10 will be a highly versatile device. Further, in this embodiment, since an in-vehicle near miss detection device 10 is assumed, learning is performed using an in-vehicle device with a similar environment, and the subject user U2 is the driver of the vehicle V2.

[0050] Note that the learning data used for the learning of the AI model is determined by appropriate information according to the use of the AI model. For example, in the case of an AI application device for a specific driver, the information of the corresponding driver is appropriate information, and in the case of an AI application device for various drivers, the information of various drivers is appropriate information. Further, in the case of an AI application device for a driver, the information of the driver is appropriate information, and in the case of an AI application device for each passenger other than the driver in the vehicle, the information of various passengers other than the driver is appropriate information. And a lot of information is required for the learning data in order to improve the accuracy of emotion estimation by the AI model.

[0051] In this embodiment, it is assumed that the AI application device can be applied to all drivers (not a specific individual but general drivers). In order to be a device that controls the vehicle based on the emotions of those drivers, it is preferable that the experimental driving test runs for AI learning are experimental driving test runs in various patterns by a plurality of subjects of various types. Since the learning by each subject is the same, the learning by a certain subject (driver) will be described.

[0052] Also, for example, when the AI application device is a diagnostic device in a medical institution, the user U2 (the target of information collection) suitable for the subject becomes a patient, a doctor, etc. in the medical institution. When the AI application device is an educational guidance device in an educational institution, the user U2 suitable for the subject becomes a student, a teacher, etc. When the AI application device is an e-sports related device or an entertainment content related device, the user U2 suitable for the subject becomes an e-sports player, a viewer of the content, etc.

[0053] It is also possible to mount the learning device 20 on the vehicle V2 to perform learning, or to install the learning device 20 in a place other than the vehicle V2 such as a research and development room to perform learning. In the latter case, each biometric data etc. (measurement data such as brain waves, heartbeats, images, voices) of the user U2 of the subject who has boarded the vehicle V2 is input to the learning device 20 via a recording medium or via a communication line. In this embodiment, an example of mounting the learning device 20 on the vehicle V2 to perform learning will be described.

[0054] A storage unit for the AI model is provided in the storage device (memory etc.) of the learning device 20 mounted on the vehicle V2, and the AI model 121m before learning is stored. Then, by the learning device 20 executing the learning process described below, the learning of the AI model before learning is performed, and a learned AI model, that is, the AI model 121M to be implemented in the AI application device is generated.

[0055] The learned AI model 121M will be used in the hat detection device 10 shown in FIGS. 1 and 4 above. Therefore, the learning data of the AI model 121m has input data as behavior information such as the user U2's line of sight, face orientation, expression, voice, etc. (appearance information, data of the same type and format as the input data of the AI model 121M of the hat detection device 10), and the correct answer data is the change amount label of the emotion index value.

[0056] Therefore, the vehicle V2 equipped with the learning device 20 is provided with various in-vehicle sensors that output data used by the near-miss detection device 10 (input to the near-miss detection device 10). In the present embodiment, specifically, the in-vehicle sensors include a camera C and a microphone M. The camera C outputs captured image information in which a face is reflected so that data on the line of sight and face orientation of the driver U2 can be obtained. The microphone M detects voice information such as generated sounds made by the driver U2. The camera C and the microphone M are installed, for example, near the front glass or the dashboard of the vehicle V2 so that the direction of the driver U2 is the shooting direction and the sound collection direction.

[0057] Then, the outputs of the camera C and the microphone M are processed to convert them into behavior data by a behavior determination unit 233 having the same configuration as the behavior determination unit 132 in FIG. 4, and are processed into behavior data on the line of sight, face orientation, expression, and voice. Then, these behavior data are input to the AI model 121m of the learning device 20 as input values.

[0058] A biological sensor is worn on the driver U2. In the present embodiment, the biological sensor is an electroencephalogram sensor S11 that detects electroencephalograms and a heart rate sensor S12 that detects heartbeats. For the electroencephalogram sensor S11, for example, a headgear-type electroencephalogram sensor is used. For the heart rate sensor S12, for example, a chest belt-type electrocardiogram heart rate sensor is used. Note that other sensors may be added or changed to the biological sensor according to the biological information to be acquired, wearability, etc. Other biological sensors may be, for example, an optical heart rate (pulse) sensor, a sphygmomanometer, or a NIRS (Near Infrared Spectroscopy) device.

[0059] Then, the electroencephalogram data output by the electroencephalogram sensor S11 and the heart rate data output by the heart rate sensor S12 are converted into arousal level and activity level, which are emotional index values, in the index value calculation unit 232. Specifically, in the same manner as the above-described method for calculating the arousal level and the activity level, the arousal level is calculated as "β wave / α wave of electroencephalogram", and the activity level is calculated as "standard deviation of the LF component of the heart rate (low-frequency component of the heart rate waveform signal)".

[0060] Furthermore, the change amount label of the emotion index value calculated by the index value calculation unit 232 is input into the AI model 121m as the correct answer value.

[0061] That is, the AI model 121m is trained with a supervised learning dataset in which the input value is the behavior data (appearance information) of the line of sight, face orientation, expression, and voice, and the correct answer value is the change amount label of the emotion index value. Therefore, the learned AI model 121M is generated as an AI model that uses the input values such as the line of sight, face orientation, expression, and voice behavior data as shown in FIG. 1, and the output value is the change amount label of the emotion index value.

[0062] Subsequently, referring to FIG. 6, the flow of the learning method (AI learning process) of the AI model 121m will be described. Note that specific operations such as wearing the sensor will be performed by the operator in charge of learning the AI model, the subject (driver U2), etc.

[0063] In the storage device (memory, etc.) of the learning device 20, a pre-learning AI model 121m that estimates the change amount label of the emotion index value using the information of the line of sight, face direction, expression, and voice as input is stored. Note that the pre-learning AI model 121m will be separately created in advance by AI design developers, etc. using a computer for AI design and according to the specifications (usage form, required performance, etc.) of the AI model.

[0064] The driver U2 gets into the vehicle equipped with the learning device 20 and adjusts the position, orientation, sensitivity, etc. of the camera C and the microphone M mounted on the vehicle. In addition, the driver U2 wears the electroencephalogram sensor S11 and the heart rate sensor S12. When the preparations for learning are complete as described above, the driver U2 etc. start the learning device 20 and learning begins. Note that, if necessary (such as when vehicle driving is a learning condition), the driver U2 starts driving (operating) the vehicle.

[0065] The camera C and the microphone M detect an image and the voice emitted by the user U2, which are the appearance information, and output the image information and the voice information to the learning device 20.

[0066] The electroencephalogram sensor S11 and the heart rate sensor S12 detect the electroencephalogram and heart rate, which are the biological signals of the user U2, and output the electroencephalogram and heart rate to the learning device 20.

[0067] Then, the learning device 20 performs subsequent processing using the image information and voice information of the user U2 at the same timing and the electroencephalogram and heart rate as paired data (one piece of learning data).

[0068] In the index value calculation unit 232 of the learning device 20, the electroencephalogram and heart rate of the input biological signals are converted into an emotion index value representing arousal (for example, β wave / α wave of the electroencephalogram) and an emotion index value representing activity (for example, standard deviation of the heart rate LF component) by emotion index value conversion processing.

[0069] In the behavior determination unit 233 of the learning device 20, the input image information and voice information of the user U2 are converted into behavior data to generate information on the line of sight, face orientation, expression, and voice. At this time, since the image information and voice information are time-series data that are temporally continuous, the image information and voice information are converted into behavior data in units of an appropriate period (appropriate timing and time length based on the emotion index value calculation timing) for learning data generation with respect to the emotion index value calculation timing. This period may be determined based on experiments or the like so that the change amount of the emotion index value of the AI model becomes an appropriate value. Then, the behavior determination unit 233 outputs the behavior data of this period length as one piece of learning data.

[0070] In the learning data set generation unit 234 of the learning device 20, learning data for the AI model 121m is generated, where the input value is behavior data (appearance information) such as the line of sight, face orientation, expression, and voice, and the correct answer value is the change amount label of the emotion index value. Then, in the learning data set generation unit 234 of the learning device 20, a large number of learning data generated at each timing are aggregated to generate a supervised learning data set for the AI model 121m. After that, the learning device 20 performs learning of the AI model 121m using the generated respective learning data sets. Note that the learning data may be sequentially input to the AI model 121m and learned at the timing when the learning data was created.

[0071] Specifically, the learning device 20 sequentially inputs the learning data of each generated learning dataset into the AI model 121m. Then, the learning device 20 performs learning such as adjusting parameters such as weights in the AI model 121m using a learning algorithm such as the error backpropagation learning method.

[0072] As a result, the AI model 121m is trained with a large number of supervised learning data generated based on the data sequentially detected by each sensor. Then, the learned AI model 121M used in the near-miss detection device 10 is generated. As described above, since the label classifying the change amount of the emotion index value based on the biological signal is used as the correct value of the AI learning data and the AI model is trained, the learned AI model 121M becomes an AI model that contributes to improving the detection accuracy of the change in emotion when a near-miss or the like occurs.

[0073] <5. Configuration of the learning device> FIG. 7 is a block diagram showing the configuration of the learning device 20. In FIG. 7, the components necessary for explaining the features of the present embodiment are shown, and the description of general components is omitted.

[0074] As shown in FIG. 7, the learning device 20 includes a communication unit 21, a storage unit 22, and a controller 23. The learning device 20 can be configured by a so-called computer device. Although not shown in the figure, the learning device 20 includes an input device such as a keyboard and an output device such as a display.

[0075] The communication unit 21 is an interface for performing data communication with other devices and various sensors via a communication network. The communication unit 21 is configured by, for example, a NIC (Network Interface Card).

[0076] The storage unit 22 includes a volatile memory and a non-volatile memory. The volatile memory is composed of, for example, RAM (Random Access Memory). The non-volatile memory is composed of, for example, ROM (Read Only Memory), flash memory, and hard disk drive. Programs and data that can be read by the controller 23 are stored in the non-volatile memory. At least a part of the programs and data stored in the non-volatile memory may be acquired from other computer devices (server devices) connected by wire or wirelessly, or from portable recording media.

[0077] An AI model storage unit 221 is provided in the storage unit 22. The pre-learning AI model 121m, which is the learning target, is stored in the AI model storage unit 221.

[0078] The controller 23 realizes various functions of the learning device 20 and includes a processor that performs arithmetic processing and the like. The processor is composed of, for example, a CPU. The controller 23 may be composed of one processor or a plurality of processors. When composed of a plurality of processors, those processors are communicably connected to each other and cooperate to execute processing. Note that the learning device 20 can also be configured as a cloud server. In that case, the CPU that constitutes the processor may be a virtual CPU.

[0079] As its functions, the controller 23 includes an acquisition unit 231, an index value calculation unit 232, a behavior determination unit 233, a learning data set generation unit 234, a generation unit 235, and a provision unit 236. In the present embodiment, the functions of the controller 23 are realized by the processor executing arithmetic processing according to a program stored in the storage unit 22.

[0080] The acquisition unit 231 acquires various types of information (image (video) information, audio information, electroencephalogram information, and heartbeat information) detected by the camera C, the microphone M, the electroencephalogram sensor S11, and the heartbeat sensor S12 via the communication unit 21. The acquisition unit 231 stores the acquired various types of information in a data table formed in the storage unit 22 as necessary for subsequent processing. Note that the acquisition unit 231 acquires these pieces of information at substantially the same time and stores them in one data record in the data table as one data set. And these data are used to generate one piece of supervised learning data.

[0081] The index value calculation unit 232 calculates an emotional index value representing the arousal level and the activity level based on the electroencephalogram and heartbeat data of the biological information. As described above, the emotional index value representing the arousal level can be calculated by "β wave / α wave of the electroencephalogram". Also, the emotional index value representing the activity level can be calculated by "standard deviation of the LF component of the heartbeat". Note that data such as calculation formulas necessary for calculating these emotional index values are stored in the storage unit 22.

[0082] The behavior determination unit 233 performs analysis processing on the image information and the generated audio including the face of the user U2 (see FIG. 6) to determine predetermined types of behaviors such as the line of sight, the direction of the face, the expression, and the voice of the user U2. Note that the processing performed by the behavior determination unit 233 is equivalent to the processing performed by the behavior determination unit 132 in FIG. 4, and each outputs the same format of data on the line of sight, the direction of the face, the expression, and the voice.

[0083] For example, regarding the line of sight of the user U3, the behavior determination unit 233 performs recognition processing such as feature amount calculation and shape discrimination with the left and right eyeballs of the user U2 as detection objects from the image including the face of the user U2. Based on the result of the recognition processing, the behavior determination unit 233 determines the behavior of the line of sight and the fixation point of the user U2 by a predetermined line of sight detection process using, for example, the position of the inner canthus, the center positions of the iris and the pupil of the eye, the center position of the corneal reflection image (Purkinje image) by near-infrared illumination, and the center position of the eyeball. The line of sight of the user U2 can be represented, for example, by two-dimensional coordinates at the position where the line-of-sight vector of the user U2 penetrates a virtual plane provided in front of the user U2 and facing the user U2.

[0084] Also, for example, regarding the face orientation of user U2, the behavior determination unit 233 performs recognition processing such as feature amount calculation and shape discrimination with the face of user U2 as the detection target from an image including the face of user U2. Based on the result of the recognition processing, the behavior determination unit 233 determines the behavior of the face orientation of user U2 by a predetermined face orientation detection process using, for example, the positions of respective parts such as eyes, nose, and mouth, the position of the top of the nose, the face contour, and the central position in the width direction of the face contour.

[0085] Also, for example, regarding the expression of user U2, the behavior determination unit 233 performs recognition processing such as feature amount calculation and shape discrimination with the face of user U2 as the detection target from an image including the face of user U2. Based on the result of the recognition processing, the behavior determination unit 233 determines the behavior of the expression of user U2 by a predetermined expression detection process using, for example, the angle of the corners of the mouth, the angle of the eyebrows, and the degree of eye opening.

[0086] Also, for example, regarding the voice information including the pronunciation such as the voice of user U2, the behavior determination unit 233 determines the behavior of the voice of user U2 based on, for example, the voice volume, the voice speed, the speech frequency, and the language analysis result of the uttered content based on the voice recognition result (e.g., the appearance of words or sentences related to joy, anger, anxiety, etc.).

[0087] The learning dataset generation unit 234 generates a supervised learning dataset with the appearance information of user U2 based on the information detected by the remote sensors (non-contact sensors: cameras, microphones) as input value data and the change amount of the emotion index values (arousal level and activity level) calculated based on the biological signals detected by the contact biological signal sensors (contact sensors: electroencephalogram sensors, heart rate sensors) as the correct value data.

[0088] Then, the generation unit 235 uses, as input values, the line-of-sight, face orientation, expression, and voice appearance information (extracted appearance information) of the user U2 determined by the behavior determination unit 233 in the same learning data, and uses the change amount label of the emotion index value as the correct answer value to input to the AI model 121m, and performs learning of the AI model 121m. Through this learning, the learned AI model 121M becomes an AI model that estimates, as input, the appearance information of the user (driver) U2 and labels indicating the change amount of the emotion index value (arousal level and activity level) of the user.

[0089] Note that the generation unit 235 performs learning of the AI model 121m by adjusting parameters such as weights in the AI model 121m using a learning algorithm such as the error backpropagation learning method.

[0090] The providing unit 236 provides the learned AI model 121m (data thereof) generated by the generation unit 235 to the near miss detection device 10 (see FIGS. 1 and 4) used in the vehicle V1 described above via a network. As a result, the near miss detection device 10 captures the AI model 121M (data thereof) provided by the providing unit 236 and stores it in the AI model storage unit 121. Therefore, the near miss detection device 10 can detect near misses based on the image information and voice information acquired by the camera C and the microphone M.

[0091] Note that although the providing unit 236 provides the learned AI model 121M (data thereof) generated by the generation unit 235 to the near miss detection device 10 via a network, it is also possible to provide it by another method. For example, by writing the data of the AI model 121M into an LSI forming an AI platform to generate an AI execution LSI and incorporating the AI execution LSI into the near miss detection device 10, the AI model 121M (data thereof) can also be provided. It is also possible to provide it to the near miss detection device 10 using a data transmission medium other than communication, such as a memory card or an optical disk recording medium.

[0092] <6. Operation Example of Learning Device > FIG. 8 is a flowchart showing the AI learning process executed by the controller 23 of the learning device 20 in FIG. 7. This flowchart shows the technical content of a computer program that causes a computer device to realize the learning process of an AI model. Further, the computer program is stored in various readable non-volatile recording media and provided (sold, distributed, etc.). The computer program may be composed of only one program, or may be composed of a plurality of cooperating programs.

[0093] The process shown in FIG. 8 is executed when, for example, a start operation of the learning process is performed by an operation unit such as a keyboard when the designer or the like of the learning device 20 executes the learning process of the AI model. If learning data sets are to be collected and generated during vehicle driving, the process is executed based on a start operation of collecting the learning data sets by an operator or a subject for the in-vehicle learning device.

[0094] In step S201, the controller 23 (acquisition unit 231) acquires data of the camera captured image and the microphone collected voice from the camera C and the microphone M, specifically, the appearance information and the biological signal of the user U2, and also acquires biological signals (brain wave data and heart rate data) from the brain wave sensor S11 and the heart rate sensor S12 and stores them in the storage unit 22, and then proceeds to step S202.

[0095] When learning is performed using each data collected by the in-vehicle device in the learning device 20 installed in a research and development room or the like, steps S201 and subsequent processes are performed using each data collected and stored by the in-vehicle device provided to the learning device 20 by wireless communication or a recording medium.

[0096] In step S202, the controller 23 (behavior determination unit 233) determines the behavior (gaze, face orientation, voice) based on each data of the camera image and the microphone voice acquired by the acquisition unit 231 in step S201, and then proceeds to step S203.

[0097] In step S203, the controller 23 (index value calculation unit 232) calculates the arousal level and activity level, which are emotion index values, based on the electroencephalogram data and heart rate data of the biological signal acquired by the acquisition unit 231 in step S201, and proceeds to step S204.

[0098] In step S204, the controller 23 (learning data set generation unit 234) generates supervised learning data with the gaze, face orientation, expression, and voice data of the behavior determined in step S202 as input values and the label indicating the change amount of the emotion index values (arousal level and activity level) calculated in step S203 as correct answer value data, and proceeds to step S205.

[0099] In step S205, the controller 23 (generation unit 235) provides the supervised learning data generated in step S204 to the AI model 121m before learning completion to perform learning of the AI model 121m, and proceeds to step S206.

[0100] In step S206, it is determined whether the learning of the AI model 121m is completed, or here, whether the learning driving of the vehicle V2 is completed, based on, for example, the ignition switch state, the operations of the vehicle driver (learning worker, subject), etc. If it is completed, the process ends; if not, it returns to step S201 to continue learning.

[0101] After that, the learned AI model 121m for which learning is completed is provided to an AI model utilization device such as the near-miss detection device 10 used in the vehicle V1 shown in FIG. 1 based on an instruction from an operator operating the learning device 20 or the like.

[0102] In the learning process shown in FIG. 8, supervised learning data is generated and the AI model is learned each time data is collected from each sensor. However, supervised learning data may be generated and accumulated each time data is collected from each sensor, and the AI model 121m before learning completion may be learned using a learning data set composed of the accumulated learning data.

[0103] In that case, step S205 accumulates the supervised learning data generated in step S204. Then, step S206 determines whether the creation of the learning data set is completed (for example, determines whether the learning driving of vehicle V3 has ended based on, for example, the ignition switch state, the operations of the vehicle driver (learning worker, subject), etc.). If it has ended, the process ends; if it has not ended, the process returns to step S201 to continue the generation process of the learning data.

[0104] After that, in a development and design office or the like, the accumulated learning data set is used to train the AI model 121m.

[0105] Also, in the training process of the above AI model, learning data is generated in units of one measurement data from each sensor. However, it is also possible to generate learning data using data obtained by statistically processing a plurality of each data within an appropriate period (statistically processing sensor output data, behavior data, or emotional index values) and use it as learning data.

[0106] During vehicle driving, bio-signal data such as electroencephalogram and heartbeat, and appearance information data such as camera-captured images and microphone-collected voices are recorded, and the recorded data is used as input data. In a development and design office or the like, the learning device 20 can generate learning data and train the AI model 121m.

[0107] <7. Example of calculating the change amount of the emotional index value > Figures 9A to 9C are diagrams showing the temporal changes of the emotional index value. Figure 9A is a diagram for explaining the first calculation example of the emotional index value, Figure 9B is a diagram for explaining the second calculation example of the emotional index value, and Figure 9C is a diagram for explaining the third calculation example of the emotional index value.

[0108] In the first calculation example of the change amount of the emotion index value shown in FIG. 9A, the change amount of the emotion index value at the first timing is obtained by dividing the difference (Y) between the latest consecutive extreme values (EVn, EVn-1) by the time difference (X) between these extreme values, and this value is defined as the change amount (Ta). When the difference (Y) between the extreme values is less than the threshold value (a threshold value for determining whether to handle inappropriate data such as noise), it is preferably not adopted as the change amount of the emotion index value.

[0109] In this first calculation example, since the change amount of the emotion index value between the extreme values is calculated, the change of the emotion index value when the change direction of the emotion index value changes can be surely captured.

[0110] In the second calculation example of the change amount of the emotion index value shown in FIG. 9B, the change amount of the emotion index value at the first timing is the difference between the maximum value and the minimum value of the emotion index value in the first time (period) (the difference between the maximum value and the minimum value at the first time indicated by the black-filled circle in FIG. 9B), or the difference between the emotion index value at the start point and the emotion index value at the end point in the first time. The first time is, for example, about several seconds, and it is preferably set to a time when it is assumed that the change of the emotion index value continues due to a close call. The change amount of the emotion index value is sequentially calculated at this first time interval.

[0111] In the second calculation example of the change amount of the emotion index value, even for short-time emotion index value data, the change amount of the emotion index value at the first timing can be calculated. Also, since it is a process of calculating the change amount of the emotion index value based on the maximum value and the minimum value in a predetermined time (the first time) period, or the start point value and the end point value in the predetermined time period, the process is relatively simple.

[0112] The third calculation example of the change amount of the emotion index value shown in FIG. 9C is the same as the first calculation example shown in FIG. 9A in terms of the calculation method of the change amount of the emotion index value itself. However, in the third calculation example, the value obtained by dividing the change amount Ta of the emotion index value at the first time Tx by the average Taav of the change amounts of the respective emotion index values at the second time T, that is, the ratio, is used. The first time is, for example, about several seconds, and it is preferably set to the time during which the change in the emotion index value is assumed to continue due to a hiccup. The second time is longer than the first time (for example, the average time between adjacent extreme values), and is, for example, several minutes to several tens of minutes or several hours. That is, the second time is preferably set to an appropriate time in order to set the average change amount (the change amount serving as a criterion for various determinations) of the change amount of the emotion index value that occurs naturally and the change amount of the emotion index value in response to some stimulus.

[0113] When calculating the change amount of the emotion index value at the first timing in real time, the start and end timings of the second time are set so that the second time ends before the first timing. That is, the average Taav of the change amounts of the emotion index value becomes a moving average value. When calculating the change amount of the emotion index value at the first timing retrospectively, the second time may span before and after the first time.

[0114] Note that, also in the second calculation example of the change amount of the emotion index value shown in FIG. 9B, the change amount of the emotion index value to be calculated may be normalized using the average Taav of the change amounts of the emotion index value during the second time period in the same manner.

[0115] In the third calculation example of the change amount of the emotion index value, the change amount of the emotion index value at the first timing is calculated based on the average of the emotion index values at the second hour. As a result, since the average of the emotion index values in the normal state (a state where no near misses, etc. are faced) is used as a reference, the change amount of the emotion index value at the first timing can be normalized with respect to individual differences, differences in the environment (e.g., vehicle type), etc. Also, since the change amount of the emotion index value at the first timing is calculated from the average of the emotion index values at the first hour and the average of the emotion index values at the second hour, even if noise is included in the emotion index value at at least one of the first hour and the second hour, the influence of the noise on the change amount of the emotion index value at the first timing can be suppressed.

[0116] <8. Examples of Labels for the Change Amount of the Emotion Index Value > The change amount of the emotion index value calculated by the above-described calculation method is a label of a continuous value (per unit), but as the label of the output value when using the AI model 121M and also as the label of the correct value when the AI model 121m is being trained, either a continuous value or a discrete value can be used.

[0117] As an example where the label of the change amount of the emotion index value is a discrete value, for example, three types, "rise", "small change", and "fall", are provided.

[0118] In the first and third calculation examples of the change amount of the emotion index value, if the change amount of the emotion index value is, for example, +20% or more, a label of "rise" is assigned, if the change amount of the emotion index value is greater than -20% and less than +20%, a label of "small change" is assigned, and if the change amount of the emotion index value is, for example, -20% or less, a label of "fall" is assigned. In the second calculation example of the change amount of the emotion index value, if the change amount of the emotion index value is, for example, +10 (the set rise threshold value in the appropriate index value unit) or more, a label of "rise" is assigned, if the change amount of the emotion index value is greater than -10 (the set fall threshold value in the appropriate index value unit) and less than +10, a label of "small change" is assigned, and if the change amount of the emotion index value is, for example, -10 or less, a label of "fall" is assigned.

[0119] As an example where the change amount label of the emotion index value is a continuous value, the change value itself of the emotion index value is used (second calculation example). Also, in the first and third calculation examples of the change amount of the emotion index value, if the change amount of the emotion index value is, for example, +23%, a label of "+23% change" is assigned. In the second calculation example of the change amount of the emotion index value, if the change amount of the emotion index value is, for example, +7, a label of "+7 change" is assigned.

[0120] In the present embodiment, as described above, two types of emotion index values, arousal level and activity level, are used. Therefore, the learned AI model 121M outputs the change amount label of the emotion index value related to the arousal level and the change amount label of the emotion index value related to the activity level. Therefore, the near miss detection unit 133 detects a near miss based on, for example, a data table of near miss determination results using the discrete values of arousal level and activity level shown in FIG. 10 as parameters.

[0121] When using one type of emotion index value, it is also possible to assign labels of near miss detection and non-detection of near miss according to the value range of the change amount of the emotion index value.

[0122] <9. Usage Example of Near Miss Detection Result> When the vehicle V1 shown in FIG. 1 is equipped with a drive recorder, for example, using the timing when a near miss is detected by the near miss detection device 10 as a trigger, various data related to the near miss detected by the drive recorder, such as the location, date and time of near miss occurrence, surrounding image data, etc. are stored (data storage at the time of event occurrence). Then, various information displays and the like are performed based on the stored near miss-related data. For example, when the vehicle V1 shown in FIG. 1 is equipped with a navigation device, the position of the vehicle V1 at the timing when a near miss is detected by the near miss detection device 10 is specified, and the near miss occurrence location is superimposed and displayed on the map displayed by the navigation device. The map with the near miss occurrence location superimposed (near miss map) can be used for post-analysis (review of driving, utilization for safety education).

[0123] Also, using the timing when a near miss is detected by the near miss detection device 10 as a trigger, in-vehicle equipment (for example, a voice guidance device included in a navigation device) mounted on the vehicle V1 shown in FIG. 1 may start psychological support (for example, a greeting such as "Are you okay?").

[0124] A server connected to the vehicle V1 via a network may create the aforementioned near miss map. Further, when the server connected to the vehicle V1 via the network receives data output from a drive recorder mounted on the vehicle V1 shown in FIG. 1, among the received data, the data to be stored may be selected using the detection result of the near miss detection device 10.

[0125] <10. Precautions, etc.> Various technical features disclosed as embodiments in this specification can be variously modified without departing from the gist of the technical creation. That is, the above embodiments are illustrative in all respects and not restrictive. The technical scope of the present invention is shown not by the description of the above embodiments but by the claims, and includes all modifications belonging to the meaning and scope equivalent to the claims. Also, a plurality of embodiments shown in this specification may be appropriately combined and implemented within the possible range.

[0126] Also, in the above embodiment, it has been described that various functions are realized software-wise by the arithmetic processing of the CPU according to a program, but at least some of these functions may be realized by electrical hardware resources. Examples of the hardware resources may include ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array), etc. Conversely, at least some of the functions assumed to be realized by hardware resources may be realized software-wise.

[0127] Further, a computer program that causes a processor (computer) to realize at least some of the functions of the near miss detection device 10 and the learning device 20 may be included. Such a computer program can be stored and provided (sold, etc.) in a computer-readable non-volatile recording medium (for example, in addition to the aforementioned non-volatile memory, an optical recording medium (for example, an optical disk), a magneto-optical recording medium (for example, a magneto-optical disk), a USB memory, or an SD card, etc.), and can also be provided by a server device via a communication line such as the Internet, that is, provided by so-called downloading.

[0128] Also, in the above embodiment, the change amount label of the emotion index value output from the learned AI model 121M is used for the detection of near misses, but it may also be used for the detection of events assuming changes in emotions other than near misses.

[0129] Also, in the above embodiment, the near miss detection device 10 uses the learned AI model 121M that outputs the change amount label of the emotion index value related to arousal level and the change amount label of the emotion index value related to activity level, but a learned AI model that outputs the change amount label of the emotion index value related to arousal level and a learned AI model that outputs the change amount label of the emotion index value related to activity level may also be used. That is, in the near miss detection device, a learned AI model may be provided for each type of emotion index value.

[0130] In this case, the near-miss detection device classifies and calculates a first emotional index value (e.g., central nervous system arousal level) using a first AI model based on the appearance information of the emotion estimation target (user), calculates a first label, classifies and calculates a second emotional index value (e.g., autonomic nervous system activity level) using a second AI model based on the appearance information of the emotion estimation target, calculates a second label, and is configured to detect the occurrence of a near-miss event based on the change state of the combined state of the first label and the second label. For example, in the psychological plane of FIG. 2 or FIG. 3, when there is a change in the direction from a quadrant other than the fourth quadrant towards the fourth quadrant (e.g., when at least one of the increase amount of the autonomic nervous system activity level and the increase amount of the central nervous system arousal level is equal to or greater than a threshold value), the near-miss detection device may detect that a near-miss event has occurred. The first AI model is trained using first supervised learning data, and the second AI model is trained using second supervised learning data. The first supervised learning data acquires the first biological signal and appearance information of the subject for learning data generation at the same timing, calculates a first emotional index value based on the acquired first biological signal, calculates a first change amount of the calculated first emotional index value, uses the acquired appearance information as an input value, and is supervised learning data with the first label classified from the calculated first change amount as the correct value. The second supervised learning data acquires the second biological signal and appearance information of the subject for learning data generation at the same timing, calculates a second emotional index value based on the acquired second biological signal, calculates a second change amount of the calculated second emotional index value, uses the acquired appearance information as an input value, and is supervised learning data with the second label classified from the calculated second change amount as the correct value.

[0131] By providing a learned AI model for each type of emotional index value, the relationship between the input and output of each learned AI model becomes simple, so an improvement in the estimation accuracy of each emotional index value (e.g., central nervous system arousal level and autonomic nervous system activity level) can be expected. In addition, various data processes can be performed on the output of each learned AI model, and the versatility of using the output of each learned AI model is increased.

[0132] For example, based on the outputs of each trained AI model (e.g., the amount of change in central nervous system arousal and the amount of change in autonomic nervous system activity), in the psychological plane of FIG. 2 or FIG. 3, determine the change in the direction from which quadrant to which quadrant, and configure an emotion estimation device that estimates the emotion of the emotion estimation target person (user) from the determination result. In this case, rather than the coordinates at which it is located in the psychological plane of FIG. 2 or FIG. 3 (static information), the emotion of the emotion estimation target person (user) is estimated based on what coordinate changes are occurring in the psychological plane of FIG. 2 or FIG. 3 (dynamic information). For example, when there is a change in the direction from a quadrant other than the fourth quadrant to the fourth quadrant (e.g., when at least one of the increase amount of autonomic nervous system activity and the increase amount of central nervous system arousal is equal to or greater than the threshold value), even if the coordinates after the change have not reached the fourth quadrant, it may be estimated that the emotion of the emotion estimation target person (user) is the emotion described in the fourth quadrant of the psychological plane of FIG. 2 or FIG. 3.

Explanation of Signs

[0133] 10 Hat Detection Device 20 Learning Device 30 Learning Data Generation Device 11, 21 Communication Unit 12, 22 Memory Unit 13, 23 Controller 121, 221 AI Model Memory Unit 121m, 121M AI Model 131, 231 Acquisition Unit 132, 233 Behavior Judgment Unit 133 Hat Detection Unit 134 Provision Unit 232 Index Value Calculation Unit 234 Learning Data Set Generation Unit 235 Generation Unit 236 Provision Unit C Camera M Microphone U1 User V1 Vehicle

Claims

1. Obtain the biological signal and appearance information of the subject for generating learning data at the same timing, Calculate an emotion index value based on the obtained biological signal, Calculate the change amount of the calculated emotion index value, Generate supervised learning data with the obtained appearance information as the input value and the label obtained by classifying the calculated change amount as the correct value, Train an AI model with the generated supervised learning data, Learning method.

2. The change amount of the emotion index value is the ratio between the change amount of the emotion index value at a first time related to the calculation of the change amount of the emotion index value and the statistical average value of the difference between consecutive extreme values of the emotion index value at a second time longer than the first time. The learning method according to Claim 1.

3. The change amount of the emotion index value is the difference between the maximum value and the minimum value of the emotion index value at a first time related to the calculation of the change amount of the emotion index value, or the difference between the emotion index value at the starting point and the emotion index value at the ending point at the first time. The learning method according to Claim 1.

4. The change amount of the emotion index value is the ratio between the average of the emotion index value at a first time related to the calculation of the change amount of the emotion index value and the average of the emotion index value at a second time longer than the first time. The learning method according to Claim 1.

5. The label is a discrete value label. The learning method according to any one of Claims 1 to 4.

6. Based on the appearance information of the subject for emotion estimation, calculate and classify a first emotion index value with a first AI model, and calculate a first label. Based on the appearance information of the subject for emotion estimation, calculate and classify a second emotion index value with a second AI model, and calculate a second label. An emotion estimation device that estimates the emotion of the subject for emotion estimation based on the combination state of the first label and the second label. The first AI model is trained with first supervised learning data. The second AI model is trained with second supervised learning data. The first supervised learning data is Obtain the first biological signal and appearance information of the subject for generating learning data at the same timing, Calculate a first emotion index value based on the obtained first biological signal, Calculate the first change amount of the calculated first emotion index value, Supervised learning data with the obtained appearance information as the input value and the first label obtained by classifying the calculated first change amount as the correct value. The second supervised learning data is At the same timing, acquire the second biological signal of the subject for generating learning data and the appearance information, Calculate a second emotion index value based on the acquired second biological signal, Calculate a second change amount of the calculated second emotion index value, The supervised learning data uses the acquired appearance information as an input value and the calculated second label obtained by classifying the second change amount as a correct value. Emotion estimation device.

7. The first emotion index value is the central nervous system arousal level, The second emotion index value is the autonomic nervous system activity level, The emotion estimation device according to claim 6.

8. Based on the appearance information of the subject for whom emotion is to be estimated, calculate and classify a first emotion index value using a first AI model, and calculate a first label, Based on the appearance information of the subject for whom emotion is to be estimated, calculate and classify a second emotion index value using a second AI model, and calculate a second label, A near miss detection device that detects the occurrence of a near miss event based on the change state of the combination state of the first label and the second label, The first AI model is trained using first supervised learning data, The second AI model is trained using second supervised learning data, The first supervised learning data is At the same timing, acquire the first biological signal of the subject for generating learning data and the appearance information, Calculate a first emotion index value based on the acquired first biological signal, Calculate a first change amount of the calculated first emotion index value, The supervised learning data uses the acquired appearance information as an input value and the calculated first label obtained by classifying the first change amount as a correct value. The second supervised learning data is At the same timing, acquire the second biological signal of the subject for generating learning data and the appearance information, Calculate a second emotion index value based on the acquired second biological signal, Calculate a second change amount of the calculated second emotion index value, The supervised learning data uses the acquired appearance information as an input value and the calculated second label obtained by classifying the second change amount as a correct value. Near miss detection device.

9. The first emotion index value is the central nervous system arousal level, The second emotion index value is the autonomic nervous system activity level, The near miss detection device according to claim 8.

10. At the same timing, acquire the biological signal and the appearance information of the subject for generating learning data, Calculate an emotion index value based on the acquired biological signal, Calculate the change amount of the calculated emotion index value, A learning data generation device that generates supervised learning data using the acquired appearance information as an input value and the label obtained by classifying the calculated change amount as a correct value.

11. Acquiring a biological signal and appearance information of a subject for learning data generation at the same timing, Calculating an emotion index value based on the acquired biological signal, Calculating a change amount of the calculated emotion index value, Generating supervised learning data using the acquired appearance information as an input value and the label obtained by classifying the calculated change amount as a correct value, A learning data generation program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Emotion estimation system and emotion estimation device

    JP2020185138A