Estimation device and learning device

By extracting a partial range including the baseline from electrocardiogram waveforms and using it for emotion estimation, the device improves the accuracy of emotion detection, specifically in distinguishing discomfort from emotionlessness.

JP2025094996APending Publication Date: 2025-06-26SUBARU CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023210742
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing techniques face challenges in accurately distinguishing between user discomfort and emotionlessness when estimating emotions based solely on electrocardiogram potential, leading to decreased estimation accuracy.

Method used

An estimation device that extracts a partial range including the baseline from the electrocardiogram waveform and uses this extracted data for emotion estimation, improving the accuracy by focusing on the vicinity of the baseline during machine learning.

Benefits of technology

The proposed solution enhances the estimation accuracy of user emotions, particularly in differentiating between discomfort and emotionlessness, resulting in improved correct answer rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094996000001_ABST
    Figure 2025094996000001_ABST
Patent Text Reader

Abstract

To increase estimation accuracy of user's feelings.SOLUTION: An estimation device comprises an extraction part that extracts a partial range, including a baseline, of an electrocardiographic waveform indicated by time series data on a cardiac potential, and an estimation part that estimates a feeling on the basis of the extracted electrocardiographic waveform.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of an estimation device and a learning device.

Background Art

[0002] Conventionally, a technique has been proposed in which a vital feature amount is acquired based on a vital signal indicating the action potential (electrocardiogram potential) of a user's heart, and the psychological index (emotion) of the user is estimated by machine learning based on the acquired vital feature amount (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, if one tries to simply estimate the user's emotion based on the electrocardiogram potential, it is difficult to distinguish whether the user is uncomfortable or is emotionless without feeling anything, and there is a problem that the estimation accuracy of the user's emotion decreases.

[0005] The present invention has been made in view of the above circumstances, and an object thereof is to improve the estimation accuracy of the user's emotion.

Means for Solving the Problems

[0006] An estimation device according to an embodiment of the present invention includes an extraction unit that extracts a partial range including a baseline in an electrocardiogram waveform represented by time-series data of an electrocardiogram potential, and an estimation unit that estimates an emotion based on the extracted electrocardiogram waveform.

Effects of the Invention

[0007] According to the present invention, the estimation accuracy of the user's emotion can be improved.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0009] <1. Configuration of the Learning Device> FIG. 1 is a diagram showing a schematic configuration of the learning device 1. As shown in FIG. 1, the learning device 1 is a computer including a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, and a GPU (Graphics Processing Unit) 14.

[0010] The CPU 11 controls the entire learning device 1 by executing programs stored in the ROM 12 and programs read from the storage unit 17 and expanded in the RAM 13. The CPU 11 functions as an extraction unit 21 and a learning unit 22. These functional units will be described later.

[0011] The GPU 14 performs calculations in machine learning, which will be described later. Therefore, it can be said that the learning unit 22 functions by both the CPU 11 and the GPU 14. Note that the learning unit 22 may function by only one of the CPU 11 and the GPU 14.

[0012] In addition to the CPU 11, ROM 12, RAM 13, and GPU 14, the learning device 1 includes an operation unit 15, a display unit 16, a storage unit 17, a communication unit 18, and an electrocardiogram sensor 19.

[0013] The operation unit 15 is composed of a keyboard, mouse, buttons, dials, touch panel, etc., and accepts user operations. When the operation unit 15 accepts a user operation, it outputs a signal corresponding to the operation to the CPU 11.

[0014] The display unit 16 is a liquid crystal display or an organic EL display, and displays various images (screens) based on the control of the CPU 11.

[0015] The storage unit 17 is a non-volatile memory such as an HDD (Hard Disk Drive) or a flash memory, and stores various information. Also, a program executed by the CPU 11 may be stored in the storage unit 17.

[0016] The communication unit 18 communicates with an external device (for example, the estimation device 41) via a network.

[0017] The electrocardiogram sensor 19 has, for example, three electrodes, and measures the action potential of a person's myocardium (hereinafter referred to as the electrocardiogram potential) at predetermined intervals by the electrodes attached to a predetermined position of the user. The time-series data of the electrocardiogram potential indicating the measurement result of the electrocardiogram sensor 19 is stored in the storage unit 17. Note that the electrocardiogram sensor 19 may be provided separately from the learning device 1, and the time-series data of the electrocardiogram obtained by the electrocardiogram sensor 19 may be transmitted to the learning device 1 via the communication unit 18.

[0018] FIG. 2 is a diagram for explaining the extraction of the electrocardiogram image 34. As shown in FIG. 2, it is possible to obtain an electrocardiogram waveform 31 by graphing the time-series data of the electrocardiogram. In the electrocardiogram waveform 31, the electrocardiogram changes periodically.

[0019] When the extraction unit 21 acquires the time-series data of the electrocardiogram, it performs band-pass filter processing to pass the frequency components from 0.4 Hz to 50 Hz in order to remove the noise derived from electromyogram and respiration from the time-series data of the electrocardiogram. Thereby, noise can be removed from the electrocardiogram waveform 31.

[0020] The extraction unit 21 detects an R wave 32 in which the electrocardiogram changes rapidly from the electrocardiogram waveform 31 based on the time-series data from which noise has been removed. Since a known detection method can be used for the detection method of the R wave 32, the description thereof is omitted here.

[0021] Next, the extraction unit 21 sets the midpoint between adjacent R waves 32 as the midpoint 33, sets the period between adjacent midpoints 33 as one period, and cuts out the time-series data corresponding to the electrocardiogram waveform 31 for three periods as shown by the broken line in FIG. 2. Hereinafter, the cut-out time-series data is referred to as partial time-series data. Note that here, the time-series data for three periods is cut out, but the time-series data for an arbitrary period such as one period, two periods, four periods or more may be cut out.

[0022] When the extraction unit 21 cuts out the partial time-series data, it cuts out the partial time-series data for three periods with the next midpoint 33 as a reference in time series. In this way, the extraction unit 21 cuts out the partial time-series data while shifting the midpoint 33 one by one in time series.

[0023] Here, it is assumed that an electrocardiogram image 34 is generated by imaging an electrocardiogram waveform 31 represented by partial time-series data. Then, an estimation result of estimating a user's emotion from another electrocardiogram image 34 using a learning model generated by performing machine learning based on the generated electrocardiogram image 34 is shown in FIG. 3.

[0024] In the estimation of the user's emotion based on the electrocardiogram image 34, as shown in FIG. 3, when the user is actually emotionless, the number of times of estimating emotionless is 54 times, and the number of times of estimating discomfort is 10 times. Therefore, the correct answer rate when the user is actually emotionless is 54 / (54 + 10)×100 = 84.4%, and it was possible to accurately estimate that the user is emotionless. On the other hand, when the user is actually uncomfortable, the number of times of estimating emotionless is 21 times, and the number of times of estimating discomfort is 43 times. Therefore, the correct answer rate when the user is actually uncomfortable is 43 / (21 + 43)×100 = 67.2%, and the estimation accuracy when the user is uncomfortable is low. Also, the overall correct answer rate was (54 + 43) / (54 + 10 + 21 + 43)×100 = 75.8%.

[0025] FIG. 4 is a diagram showing a region of interest when the confidence level of discomfort is 100%. FIG. 5 is a diagram showing a region of interest when the confidence level of discomfort is 0.01%. In FIGS. 4 and 5, it is a visualization of the region that the estimation device (computer) pays attention to when estimating emotion using a learning model generated by machine learning based on the electrocardiogram image 34, and it shows that the higher the color density, the higher the degree of attention.

[0026] As shown in FIG. 4, it can be seen that when the confidence level of discomfort is 100%, the vicinity of the baseline (baseline: for example, a line connecting the start point of the P wave to the start point of the next P wave) in the electrocardiogram waveform 31 is mainly being paid attention to. On the other hand, as shown in FIG. 5, when the confidence level of discomfort is 0.01%, it can be seen that not only the vicinity of the baseline in the electrocardiogram waveform 31 but also the lower side away from the baseline is being noted. In other words, when the confidence level of discomfort is 0.01%, it can be seen that the attention area cannot be narrowed down and is being noted over a wide range.

[0027] Therefore, when estimating discomfort, it is considered that the correct answer rate (confidence level) can be improved by mainly focusing on the vicinity of the baseline.

[0028] Therefore, the extraction unit 21 extracts a partial range including the baseline in the electrocardiogram waveform 31 based on the cut-out partial time-series data, so that only the extracted part is used in the subsequent machine learning.

[0029] FIG. 6 is a diagram for explaining the extraction of the extracted electrocardiogram image 37. As shown on the left side of FIG. 6, in the electrocardiogram waveform 31 based on the partial time-series data, a trend of gently rising or falling can be seen in the vicinity of the baseline. These trends may expand the attention area when learning and estimating the user's emotion.

[0030] Therefore, the extraction unit 21 generates difference time-series data by taking the difference between the front and back in the time series order in the partial time-series data. For example, when the partial time-series data is represented as [x t , x t+1 , x t+2 , ···], the difference time-series data can be represented as [x t+1 - x t , x t+2 - x t+1 , x t+3 - x t+2 , ···].

[0031] As shown in the center of FIG. 6, in the electrocardiogram waveform 35 based on the differential time-series data, the gentle rising or falling trend that existed near the baseline is generally removed, and the waveform near the baseline becomes flat. As a result, attention areas tend to gather near the baseline in learning and estimation, making it possible to improve the correct rate of estimating the user's emotion.

[0032] Next, the extraction unit 21 normalizes the differential time-series data so that the largest value becomes 1 and the smallest value becomes 0. Further, the extraction unit 21 calculates the average value of the normalized differential time-series data, and extracts (crops) the range of ±0.2 based on the average value from the electrocardiogram waveform 36 based on the normalized differential time-series data and images it to generate the extracted electrocardiogram image 37. Here, the extraction unit 21 generates a square extracted electrocardiogram image 37 by expanding the extracted electrocardiogram waveform 36 in the vertical direction. The extraction unit 21 stores the generated extracted electrocardiogram image 37 in the storage unit 17 as a learning image. At this time, the extracted electrocardiogram image 37 is stored in association with the user's emotion (unhappiness, apathy, etc.) when the electrocardiogram potential was measured.

[0033] Here, the differential time-series data is normalized, but normalization is not essential. Also, the range of ±0.2 is extracted based on the average value, but the extraction range is not limited to this and may be other ranges. At this time, in the electrocardiogram waveform 36, it is preferable that the range is determined so that the vicinity of the baseline is included and the portion corresponding to the peak of the R wave where the electrocardiogram potential changes rapidly is outside the range.

[0034] The learning unit 22 acquires from the storage unit 17 the extracted electrocardiogram image 37 stored in the storage unit 17 and the user's emotion (unhappiness, apathy, etc.) at that time. Then, the learning unit 22 sets the user's current emotion as the correct answer, and performs learning processing by known machine learning using the extracted electrocardiogram image 37 at that time as the learning image to generate a learning model. The machine learning performed here can use, for example, Deep Learning (DL). As algorithms used in deep learning, known methods such as Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) can be used.

[0035] Figure 7 is a flowchart showing the processing flow of the learning device 1. As shown in Figure 7, in step S1, the extraction unit 21 acquires time-series data of the electrocardiogram potential from the electrocardiogram sensor 19. In step S2, the extraction unit 21 performs band-pass filter processing to pass frequency components from 0.5 Hz to 40 Hz on the time-series data. In step S3, the extraction unit 21 detects the R wave 32 from the electrocardiogram waveform 31 based on the time-series data with noise removed. In step S4, the extraction unit 21 cuts out partial time-series data for three cycles from the midpoint 33 between adjacent R waves 32.

[0036] In step S5, the extraction unit 21 generates differential time-series data by taking the difference between the front and back in time series order in the partial time-series data. In step S6, the extraction unit 21 normalizes the differential time-series data. In step S7, the extraction unit 21 draws an electrocardiogram waveform 36 based on the normalized differential time-series data. In step S8, the extraction unit 21 generates an extracted electrocardiogram image 37 by extracting and imaging the range from ±0.2 from the average value in the electrocardiogram waveform 36. In step S9, the learning unit 22 uses the current emotion of the user as the correct answer, and generates a learning model by performing learning processing by known machine learning using the extracted electrocardiogram image 37 at that time as the learning image.

[0037] <2. Configuration of a vehicle equipped with an estimation device> FIG. 8 is a diagram showing an outline of the configuration of the vehicle 100 equipped with the estimation device 41. As shown in FIG. 8, the vehicle 100 includes an estimation device 41, an operation unit 42, a display unit 43, an audio output unit 44, a storage unit 45, a communication unit 46, and an electrocardiogram sensor 47. Note that the operation unit 42, the display unit 43, the audio output unit 44, the storage unit 45, the communication unit 46, and the electrocardiogram sensor 47 may be provided as part of the estimation device 41.

[0038] The estimation device 41 is a computer such as an ECU (Electronic Control Unit) and controls each part mounted on the vehicle 100. The estimation device 41 functions as an extraction unit 51, an estimation unit 52, and an operation control unit 53. These functional units will be described later.

[0039] The operation unit 42 is composed of buttons, dials, touch panels, etc., and receives the user's operations. When the operation unit 42 receives the user's operation, it outputs a signal corresponding to the operation to the estimation device 41.

[0040] The display unit 43 is a liquid crystal display or an organic EL display, and displays various images (screens) based on the control of the estimation device 41.

[0041] The audio output unit 44 is, for example, a speaker, and outputs various sounds (such as music and voices) based on the control of the estimation device 41.

[0042] The storage unit 45 is a non-volatile memory such as an HDD (Hard Disk Drive) or a flash memory, and stores various information (for example, a learning model). Also, a program executed by the estimation device 41 may be stored in the storage unit 45. Note that the learning model may be stored in the storage unit 45 in advance, or may be acquired via the communication unit 46 and stored in the storage unit 45.

[0043] The communication unit 46 communicates with an external device (for example, the learning device 1) via a network.

[0044] The electrocardiogram sensor 47 has, for example, three electrodes, measures the electrocardiogram potential at predetermined intervals by the electrodes attached to a predetermined position of a user (for example, a driver), and outputs time-series data of the electrocardiogram potential indicating the measurement results to the estimation device 41.

[0045] Similar to the extraction unit 21 of the learning device 1, when the extraction unit 51 acquires the time-series data of the electrocardiogram potential measured by the electrocardiogram sensor 47, it performs band-pass filter processing to pass the frequency components from 0.5 Hz to 40 Hz on the time-series data, and detects the R wave 32 from the electrocardiogram waveform 31 based on the time-series data with noise removed. Further, the extraction unit 51 cuts out partial time-series data for three cycles from the midpoint 33 between adjacent R waves 32, and generates difference time-series data by taking the difference between the front and back in the time series order in the cut-out partial time-series data. Then, the extraction unit 51 normalizes the difference time-series data, extracts a range of ±0.2 from the average value of the electrocardiogram waveform 36 based on the normalized difference time-series data, and generates an extracted electrocardiogram image 37 by imaging, and outputs it to the estimation unit 52.

[0046] The estimation unit 52 performs an estimation process of estimating the user's emotion using the learning model stored in the storage unit 45 on the extracted electrocardiogram image 37.

[0047] FIG. 9 is a diagram showing the estimation result by machine learning using the extracted electrocardiogram image 37. As shown in FIG. 9, in the machine learning based on the extracted electrocardiogram image 37, when the user is actually emotionless, the number of times estimated as emotionless is 47 times, and the number of times estimated as unhappy is 17 times. Therefore, the correct answer rate when the user is actually emotionless is 47 / (47 + 17)×100 = 73.4%. Also, when the user is actually unhappy, the number of times estimated as emotionless is 11 times, and the number of times estimated as unhappy is 53 times. Therefore, the correct answer rate when the user is actually unhappy is 53 / (11 + 53)×100 = 82.8%. Therefore, the overall correct answer rate is (47 + 53) / (47 + 17 + 11 + 53)×100 = 78.1%.

[0048] Here, the estimation result by machine learning using the electrocardiogram image 34 (see FIG. 3) is compared with the estimation result by machine learning using the extracted electrocardiogram image 37 (FIG. 9). As a result, it can be seen that in the estimation result by machine learning using the extracted electrocardiogram image 37, the correct answer rate when the user is actually uncomfortable has been significantly improved from 67.2% to 82.8% compared to the estimation result by machine learning using the electrocardiogram image 34. Also, it can be seen that in the estimation result by machine learning using the extracted electrocardiogram image 37, the overall correct answer rate has been improved from 75.8% to 78.1% compared to the estimation result by machine learning using the electrocardiogram image 34.

[0049] In this way, by performing the learning process and the estimation process based on the extracted electrocardiogram image 37 obtained by cutting out the vicinity of the baseline that is focused on when the confidence level is high in machine learning, it becomes possible to accurately estimate whether the user is uncomfortable or emotionless.

[0050] When the estimation unit 52 estimates that the user is uncomfortable, the operation control unit 53 issues an alert by video from the display unit 43 and also issues an alert by voice from the voice output unit 44. Thereby, it becomes possible to notify the user that they are uncomfortable and cause the user to take actions such as taking a break.

[0051] FIG. 10 is a flowchart showing the processing flow of the estimation device 41. Note that the same processes as those performed by the learning device 1 shown in FIG. 7 are denoted by the same reference numerals and detailed descriptions thereof are omitted.

[0052] As shown in FIG. 10, in step S1, the extraction unit 51 acquires time-series data of the cardiac potential from the electrocardiogram sensor 47. In step S2, the extraction unit 51 performs band-pass filter processing on the time-series data. In step S3, the extraction unit 51 detects the R wave 32 from the electrocardiogram waveform 31 based on the time-series data from which noise has been removed. In step S4, the extraction unit 21 cuts out partial time-series data for three cycles from the midpoints 33 between adjacent R waves 32.

[0053] In step S5, the extraction unit 51 generates difference time-series data by taking the difference between the front and back in time series order in the partial time-series data. In step S6, the extraction unit 21 normalizes the difference time-series data. In step S7, the extraction unit 51 draws an electrocardiogram waveform 36 based on the normalized difference time-series data. In step S8, the extraction unit 51 generates an extracted electrocardiogram image 37 by extracting a range of ±0.2 from the average value of the electrocardiogram waveform 36 and imaging it.

[0054] In step S11, the estimation unit 52 performs an estimation process of estimating the user's emotion using the learning model stored in the storage unit 45 for the extracted electrocardiogram image 37. In step S12, the operation control unit 53 determines whether the user's emotion estimated by the estimation unit 52 is unpleasant. And when the user's emotion is not unpleasant (No in step S12), the operation control unit 53 ends the process. On the other hand, when the user's emotion is unpleasant (Yes in step S12), in step S13, the operation control unit 53 issues an alert by video from the display unit 43 and issues an alert by voice from the voice output unit 44.

[0055] <3. Modification example> As described above, the embodiments according to the present invention have been described, but the present invention is not limited to the above-described specific examples and can adopt various configurations. For example, in the above embodiment, the estimation device 41 is mounted on the vehicle 100, but the estimation device 41 may be provided outside the vehicle 100.

[0056] Also, in the above embodiment, the learning device 1 and the estimation device 41 are configured by different hardware, but they may be configured by the same hardware.

[0057] Also, in the above embodiment, either non-emotional or unpleasant is learned and estimated as the user's emotion, but other emotions (such as pleasant) may be learned and estimated in addition to non-emotional and unpleasant as the user's emotion.

[0058] Also, in the above-described embodiment, when it is estimated that the user is uncomfortable, an alert is issued by video and audio. However, when it is estimated that the user is uncomfortable, the operation control unit 53 may cause other operations other than issuing an alert to be performed.

[0059] <4. Summary of Embodiment> As described above, the estimation device 41 of the embodiment includes an extraction unit 51 that extracts a partial range including a baseline in the electrocardiogram waveform 31 indicated by the time-series data of the electrocardiogram potential, and an estimation unit 52 that estimates an emotion based on the extracted electrocardiogram waveform 35 (36). Thereby, by performing learning processing and estimation processing based on the time-series data near the baseline that is noted when the confidence level is high in machine learning, it is possible to accurately estimate whether the user is uncomfortable or non-emotional.

[0060] The extraction unit 51 cuts out one or more cycles from the time-series data, and extracts a partial range including the baseline in the electrocardiogram waveform 31 based on the cut-out time-series data as an electrocardiogram image (extracted electrocardiogram image 37), and the estimation unit 52 estimates an emotion based on the electrocardiogram image (extracted electrocardiogram image 37). Thereby, since the electrocardiogram waveform 31 can be treated as the extracted electrocardiogram image 37, the estimation accuracy of the user's emotion can be improved.

[0061] The extraction unit 51 calculates differential time-series data by taking a difference between the front and back in the time-series data, and extracts a partial range including the baseline in the electrocardiogram waveform 35 (36) based on the differential time-series data. Thereby, the gentle rising or falling trend existing near the baseline is removed, and attention areas tend to gather near the baseline in learning and estimation, and the correct answer rate of estimating the user's emotion can be improved.

[0062] The extraction unit 51 extracts a partial range including the baseline so as not to include the peak component of the R wave in the electrocardiogram waveform 36. As a result, since the R wave component that greatly deviates from the baseline in the electrocardiogram waveform 36 is removed, it becomes possible to pay more attention to the vicinity of the baseline in the estimation process, and it is possible to more accurately estimate whether the user is in a state of discomfort or apathy.

[0063] Also, as described above, the learning device 1 of the embodiment includes an extraction unit 21 that extracts a partial range including the baseline in the electrocardiogram waveform 31 represented by the time-series data of the electrocardiogram potential, the extracted electrocardiogram waveform 31, and a learning unit 22 that learns the user's emotion based on the user's emotion when the time-series data is measured. As a result, by performing the learning process and the estimation process based on the time-series data near the baseline that is focused on when the confidence level is high in machine learning, it is possible to accurately learn whether the user is in a state of discomfort or apathy.

Explanation of Reference Numerals

[0064] 1 Learning device 21 Extraction unit 22 Learning unit 41 Estimation device 51 Extraction unit 52 Estimation unit 53 Operation control unit

Claims

1. An extraction unit that extracts a partial range including a baseline in an electrocardiogram waveform represented by time-series data of electrocardiographic potentials, An estimation unit that estimates an emotion based on the extracted electrocardiogram waveform, An estimation device comprising the above.

2. The extraction unit cuts out one or more cycles from the time-series data and extracts, as an electrocardiogram image, a partial range including a baseline in the electrocardiogram waveform based on the cut-out time-series data, The estimation unit estimates an emotion based on the electrocardiogram image The estimation device according to Claim 1.

3. The extraction unit calculates difference time-series data by taking a difference between before and after in the time-series data, and extracts a partial range including a baseline in the electrocardiogram waveform based on the difference time-series data The estimation device according to Claim 1 or Claim 2.

4. The extraction unit extracts a partial range including a baseline so as not to include the peak component of the R wave in the electrocardiogram waveform The estimation device according to Claim 1 or Claim 2.

5. An extraction unit that extracts a partial range including a baseline in an electrocardiogram waveform represented by time-series data of electrocardiographic potentials, A learning unit that learns the user's emotion based on the extracted electrocardiogram waveform and the user's emotion when the time-series data was measured, A learning device comprising the above.

Citation Information

Patent Citations

  • Estimation device, learning device, estimation method and computer program

    JP2020130344A