System for estimating emotion and stress index from brainwave signals and facial image
A system using EEG sensors and facial image analysis with deep learning models in complex domains improves the accuracy of emotional and stress index estimation, facilitating continuous remote monitoring.
Patent Information
- Application Number
- PCT/KR2024/020935
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
Existing technologies have not effectively provided objective, reliable, and highly accurate methods for simultaneously estimating emotional and stress indices using human biosignals.
A system utilizing an EEG sensor, camera, and deep learning models to process EEG and facial images, converting them into complex domains for improved accuracy in estimating emotional and stress indices.
The system significantly enhances the accuracy of emotional and stress index estimation, enabling continuous remote monitoring and visualization of mental and physical health status.
Smart Images

Figure KR2024020935_03072025_PF_FP_ABST
Abstract
Description
A system that estimates emotional and stress indices from brainwave signals and facial images.
[0001] The present invention relates to a system, device and method for estimating emotional and stress indices, and more particularly, to a system, device and method for estimating emotional and stress indices from electroencephalogram (EEG) signals and facial images captured of a user's face.
[0002] Emotions and stress are constantly intertwined. Emotions are our subjective experiences of a situation, while stress is our physical or mental response to a factor or situation. These two interact to profoundly impact our physical health. Persistent or excessively intense negative emotions can trigger stress, which in turn can intensify and prolong negative emotions, potentially leading to other physical health problems such as heart disease, diabetes, cancer, and abdominal pain. For these reasons, both emotions and stress require ongoing and objective monitoring.
[0003] "Stress Detection With a Single PPG Sensor by Orchestrating Multiple Denoising and Peak-Detecting Methods": This study proposes a method for stress detection using precise signal processing based on PPG data. This method utilizes PPG signals from a wearable device to detect stress levels and manage them.
[0004] "Development of a Wearable System for Real-Time Blood Pressure-Stress Monitoring and Workplace Stress Management": This study aims to develop a wearable device for multimodal, multi-biosignal measurement for stress management, including the development of a low-power integrated chip for measuring various biosignals, such as EEG, ECG, and PPG. This also includes the development of a relative blood pressure-encephalography-based stress measurement algorithm.
[0005] "Stress status classification based on EEG signals": This study explores the relationship between brainwave signals and stress. Using a 32-channel wired EEG device, the study analyzes frequency signatures of brainwaves and derives a quantitative stress index from these data. The study attempts to improve the accuracy of stress classification by analyzing power values in specific EEG bands.
[0006] However, despite these efforts to date, research to derive objective, reliable, and highly accurate stress indices using various human biosignals has not yet made significant progress, and in particular, there are no efforts to provide emotional states and stress indices together.
[0007] The technical problem to be achieved in the present invention is to provide a device capable of providing emotional and stress indices to a user.
[0008] Another technical challenge to be achieved by the present invention is to provide a remote device capable of providing emotional and stress indices to a user.
[0009] Another technical challenge to be achieved in the present invention is to provide a method for providing emotional and stress indices to a user.
[0010] Another technical problem to be achieved by the present invention is to provide a computer-readable recording medium or a computer-readable recording medium having recorded thereon a computer program for performing a method of providing an emotional and stress index to a user.
[0011] The technical problems to be achieved in the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.
[0012] In order to achieve the above technical task, a device capable of estimating an emotion and stress index according to the present invention may include an EEG sensor for sensing an electroencephalogram (EEG) signal of a user; a camera for photographing a face of the user to obtain a facial image; and a processor for detecting a first image of a forehead region and a second image of a cheek region from the obtained facial image, complexizing the detected first image and the detected second image and applying them to a predetermined trained first deep learning model to output a remote PPG (remote Photoplethysmography, rPPG) signal waveform of the user, and applying the sensed EEG signal waveform and the output rPPG signal waveform to a predetermined trained second deep learning model to calculate an emotion and stress index of the user.
[0013] The processor can convert the detected first image into a complex domain by forming a complex number with the real part and the second image as the imaginary part, and then apply the converted complex number to the predetermined trained first deep learning model.
[0014] The processor may convert the detected first image into a complex domain by forming a complex number with the imaginary part of the first image and the real part of the second image, and then apply the converted complex number to the predetermined trained first deep learning model.
[0015] The above-described first deep learning model may be a convolutional neural network (CNN) model.
[0016] The above-described second deep learning model may be a multi-task learning algorithm model based on a convolutional neural network (CNN).
[0017] The above-mentioned acquired facial image may be acquired using ambient light around the user as a light source.
[0018] It may further include a communication unit that transmits information about the user's emotional and stress indices calculated above to a user terminal or server.
[0019] In order to achieve the above technical task, a remote device capable of estimating an emotion and stress index may include a communication unit that receives an electroencephalogram (EEG) signal sensed for a user and a facial image captured for the user's face; and a processor that detects a first image for a forehead region and a second image for a cheek region from the acquired images, complexizes the detected first image and the detected second image and applies them to a predetermined trained first deep learning model to output a remote PPG (remote Photoplethysmography, rPPG) signal waveform of the user, and applies the sensed EEG signal waveform and the output rPPG signal waveform to a predetermined trained second deep learning model to calculate an emotion and stress index of the user.
[0020] The above processor can apply the sensed EEG signal waveform as a real part and the output rPPG signal waveform as an imaginary part, or can complexize the sensed EEG signal waveform as an imaginary part and the output rPPG signal waveform as a real part to the predetermined learned second deep learning model.
[0021] In order to achieve the above technical task, a method for estimating an emotion and stress index may include the steps of sensing an electroencephalogram (EEG) signal of a user; obtaining a face image by photographing a face of the user; detecting a first image of a forehead region from the obtained image; detecting a second image of a cheek region from the obtained image; complexizing the detected first and second images; applying the complexed data to a predetermined trained first deep learning model to output a remote PPG (remote Photoplethysmography, rPPG) signal waveform of the user; and applying the sensed EEG signal waveform and the output rPPG signal waveform to a predetermined trained second deep learning model to calculate an emotion and stress index of the user.
[0022] The method may further include a step of transmitting information about the user's emotional and stress indices calculated above to a user terminal or server.
[0023] An emotion and stress index estimation device according to one embodiment of the present invention calculates a user's emotion and stress index based on the user's brain wave signal and facial image, and can significantly improve accuracy.
[0024] An emotion and stress index estimation device according to one embodiment of the present invention remotely receives a user's brainwave signal and facial image and calculates an emotion and stress index, thereby remotely estimating the user's emotion and stress index and enabling continuous monitoring of the user's mental and physical health status.
[0025] An emotion and stress index estimation device according to one embodiment of the present invention calculates the emotion and stress index of a user based on the user's brainwave signal and facial image, and transmits the calculation result to the user's terminal or server, and visualizes and conveys the result to the user in various ways, including past calculation results.
[0026] The effects that can be obtained from the invention are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.
[0027] Figure 1 is a diagram illustrating the layer structure of an artificial neural network.
[0028] Figure 2 is a diagram illustrating an example of a deep neural network.
[0029] FIG. 3 is a block diagram illustrating the function of a device capable of estimating emotional and stress indices according to the present invention.
[0030] FIG. 4 is a diagram for explaining input data of a predetermined learned first deep learning model according to the present invention.
[0031] FIG. 5 is a flowchart illustrating a method for outputting an rPPG signal waveform from a facial image according to one embodiment of the present invention.
[0032] Figure 6 is an exemplary drawing for explaining the complexization of an image for a detected forehead area and an image for a detected cheek area.
[0033] FIG. 7 is an example diagram showing a convolutional neural network (CNN) model applied to a predetermined learned first deep learning model according to the present invention.
[0034] FIG. 8 is a diagram schematically illustrating a method for estimating emotional and stress indices according to one embodiment of the present invention.
[0035] Figure 9 is an exemplary diagram for explaining the complexization of EEG signal waveforms and rPPG signal waveforms.
[0036] FIG. 10 is an example diagram of a multi-task learning algorithm model based on a convolutional neural network (CNN) applied to a predetermined learned second deep learning model according to the present invention.
[0037] FIG. 11 is a block diagram illustrating a configuration of a remote device capable of estimating emotional and stress indices according to another embodiment of the present invention.
[0038] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. The detailed description set forth below, together with the accompanying drawings, is intended to illustrate exemplary embodiments of the present invention and is not intended to represent the only embodiments in which the present invention may be practiced. The following detailed description includes specific details to provide a thorough understanding of the present invention. However, one of ordinary skill in the art will appreciate that the present invention may be practiced without these specific details.
[0039] In some cases, to avoid ambiguity in the concepts of the present invention, well-known structures and devices may be omitted or illustrated in block diagram form focusing on the core functions of each structure and device. Furthermore, the same components are described using the same reference numerals throughout this specification.
[0040] Predicting PPG based on facial images is safe for contact-transmitted diseases like COVID-19. It's also particularly beneficial for diseases like diabetes, which require continuous monitoring because only images are required. With the aging population and the rise in chronic diseases, this technology can measure biomarkers and information anytime, anywhere for diseases that place a significant burden on public and household healthcare costs.
[0041] The present invention proposes an algorithmic model that estimates emotional and stress indices using a user's brainwave and facial information. The proposed algorithmic model analyzes the collected data and accurately estimates emotional and stress indices.
[0042] Before explaining the present invention, let's explain artificial intelligence (AI), machine learning, and deep learning. The easiest way to understand the relationship between these three concepts is to imagine three concentric circles. AI is the largest circle, followed by machine learning, and deep learning, which is driving the current AI boom, is the smallest circle.
[0043] The concept of artificial intelligence first emerged in 1956 at the Dartmouth Conference hosted by Professor John McCarthy at Dartmouth College, and has experienced explosive growth in recent years. This growth has been particularly accelerated since 2015 by the introduction of GPUs, which offer fast and powerful parallel processing capabilities. The advent of the big data era, with its explosive growth in storage capacity and the resulting overflow of data across all domains, including images, text, and mapping data, has also significantly influenced this growth.
[0044] Artificial Intelligence - Implementing human intelligence in machines
[0045] Back in 1956, the pioneers of artificial intelligence dreamed of ultimately creating a complex computer with characteristics similar to human intelligence. While this type of AI possesses human senses, reasoning, and human-like thinking, it's called "general AI." However, the AI currently achievable with technological advancement falls under the umbrella of "narrow AI." Narrow AI is characterized by its ability to perform specific tasks, such as image classification services for social media or facial recognition, with superior human capabilities.
[0046] Machine Learning - A Specific Approach to Implementing Artificial Intelligence
[0047] Machine learning automatically filters spam from your inbox. Meanwhile, machine learning fundamentally uses algorithms to analyze data, learn from the analysis, and then make judgments or predictions based on what it learns. Therefore, the ultimate goal is not to directly code specific instructions for decision-making criteria into software, but to "learn" the computer itself through massive amounts of data and algorithms, allowing it to learn how to perform tasks. Machine learning originated from concepts pioneered by early AI researchers, and algorithmic approaches include decision tree learning, inductive logic programming, clustering, reinforcement learning, and Bayesian networks. However, none of these achieved the ultimate goal of general AI, and it's true that early machine learning approaches often struggled to achieve even narrow AI.
[0048] While machine learning has achieved significant results in fields like computer vision, it faces the limitation of requiring a certain amount of coding throughout the entire process of implementing artificial intelligence, even without specific guidelines. For example, when recognizing an image of a stop sign using a machine learning system, developers must manually code edge detection filters that programmatically identify the object's start and end, shape detection to identify the object's faces, and classifiers that recognize characters like "STO-P." In this way, machine learning recognizes images from "coded" classifiers and "learns" about stop signs through algorithms.
[0049] While machine learning's image recognition rate is sufficient for commercial use, it can sometimes be poor in certain situations, such as fog or trees obscuring signs. Until recently, computer vision and image recognition have struggled to reach human-level performance due to these recognition rate issues and frequent errors.
[0050] Deep Learning - The Technology That Realizes Full Machine Learning
[0051] Artificial neural networks (ANNs), another algorithm developed by early machine learning researchers, were inspired by the biological properties of the human brain, specifically the interconnected structure of its neurons. However, unlike the brain, where any physically adjacent neurons can be interconnected, ANNs have consistent layer connections and data propagation directions.
[0052] For example, if an image is sliced into numerous tiles and input into the first layer of a neural network, the neurons in that tile repeatedly pass the data to the next layer until the final output is generated at the last layer. Each neuron is assigned a weight representing the accuracy of the input based on the task it performs, and the final output is then determined by adding all the weights together. In the case of a stop sign, the characteristics of the image, such as the octagonal shape, the red color, the text displayed, the size, and the presence of movement, are sliced and "examined" by the neurons, and the neural network's task is to identify whether it is a stop sign. Here, a "probability vector" is utilized, which predicts the outcome based on the weights based on sufficient data.
[0053] Deep learning is a form of artificial intelligence developed from artificial neural networks. It learns data by utilizing information input / output layers similar to neurons in the brain. However, even basic neural networks require enormous computational power, hindering the commercialization of deep learning from the beginning. Nevertheless, researchers continued their research and, using supercomputers, successfully parallelized algorithms that proved the concept of deep learning. The advent of GPUs, optimized for parallel computing, dramatically accelerated neural network computation, ushering in the emergence of true deep learning-based artificial intelligence.
[0054] Neural networks are likely to make numerous errors during the "learning" process. Returning to the stop sign example, adjusting the weights of neuron inputs precisely enough to consistently produce the correct answer regardless of weather conditions or day / night changes might require learning hundreds, thousands, or even millions of images. Only when this level of accuracy is achieved can the neural network be considered to have properly learned the stop sign classification. In 2012, Google and Stanford University Professor Andrew Ng implemented a "deep neural network" consisting of over a billion neural networks on 16,000 computers. Using this network, they analyzed 10 million images from YouTube and successfully taught the computer to classify photos of people and cats. They taught the computer to recognize and judge the shape and appearance of cats in the videos on its own.
[0055] Systems trained with deep learning already have image recognition capabilities that surpass those of humans. Deep learning also encompasses the ability to identify cancer cells in blood and tumors in MRI scans. Google's AlphaGo learned the fundamentals of Go and further strengthened its neural network through repeated matches against AI systems like itself. The advent of deep learning has enhanced the practicality of machine learning and expanded the scope of artificial intelligence. Deep learning subdivides tasks into every possible way a computer system can support. Technologies based on deep learning, such as driverless cars, better preventative medicine, and more accurate movie recommendations, are already being used in our daily lives or are on the verge of becoming practical. Deep learning is considered both the present and the future of artificial intelligence, with the potential to realize general AI, once a fantasy of science fiction.
[0056] Below, we will look at deep learning in more detail.
[0057] Deep learning is a type of artificial neural network (ANN) that utilizes the theory of the human neural network (Neural Network). It is a set of machine learning models or algorithms that refer to a deep neural network (DNN) that is structured in a layer structure and has one or more hidden layers (hereinafter referred to as intermediate layers) between the input layer and the output layer. Simply put, deep learning can be said to be an artificial neural network with a deep layer.
[0058] The human brain is estimated to be composed of 25 billion nerve cells. The brain is made up of nerve cells, and each nerve cell (neuron) is a single nerve cell that forms a neural network. A nerve cell consists of a cell body, an axon (or nucleus), and usually multiple dendrites (or protoplasmic processes). Information is transmitted between nerve cells through synapses, the connections between nerve cells. While a single nerve cell appears simple in isolation, when these nerve cells come together, they are capable of human intelligence. Dendrites are the part that receives signals from other nerve cells (input), while the axon is the very long extension from the cell body that transmits signals to other nerve cells (output). Synapses, the connections between axons and dendrites that transmit signals between nerve cells, do not transmit signals unconditionally. Instead, they only transmit signals if the signal strength exceeds a certain value (threshold). In other words, not only does each synapse have a different connection strength, but it also determines whether or not a signal will be transmitted.
[0059] Artificial neural networks (ANNs), a branch of artificial intelligence, are mathematical models modeled after the structure of the biological (typically human) brain. In other words, ANNs mimic the information processing and transmission processes of biological neurons. Similar to how the human brain solves problems, ANNs exhibit excellent parallelism because each neuron operates independently. Furthermore, because information is distributed across numerous connections, problems in a few neurons do not significantly impact the overall system. Consequently, ANNs are robust to a certain level of error and possess the ability to learn from a given environment.
[0060] Deep neural networks (DNNs) can be considered descendants of artificial neural networks (ANNs). They transcend existing limitations and achieve success in areas where numerous AI technologies have failed in the past. They are the latest version of ANNs. Looking at the modeling of ANNs based on biological neural networks, the processing units are modeled as nodes, and the connections are modeled as synapses, which are weights, as shown in Table 1.
[0061] (Synapse) was modeled with weights as shown in Table 1 below.
[0062] Biological neural network Artificial neural network Cell body Node Dendrite Input Axon Output Synapse Weight
[0063] Figure 1 illustrates the layer structure of an artificial neural network. Just as human biological neurons are interconnected in multiple layers to perform meaningful tasks, individual neurons in an artificial neural network are interconnected through synapses. Multiple layers are interconnected, and the connection strengths between each layer can be updated using weights. This multilayer structure and connection strengths are utilized in fields such as learning and cognition.
[0064] Each node is connected by weighted links, and the entire model learns by repeatedly adjusting the weights. Weights are the basic means of long-term memory and express the importance of each node. Simply put, an artificial neural network trains the entire model by initializing these weights and updating and adjusting them with the training data set. After training is complete, when a new input value is received, the appropriate output value is inferred. The learning principle of an artificial neural network can be viewed as a process in which intelligence is formed through the generalization of experience and is performed in a bottom-up manner. In Figure 1, when there are two or more intermediate layers (i.e., 5 to 10), it is considered deep and is called a deep neural network. Learning and inference models achieved through such a deep neural network can be referred to as deep learning.
[0065] While artificial neural networks can perform to some extent with a single intermediate layer (commonly referred to as a hidden layer) beyond the input and output layers, as the problem complexity increases, the number of nodes or layers must be increased. While increasing the number of layers to achieve a multilayered model is effective, its application is limited due to the impossibility of efficient learning and the large computational load required to train the network.
[0066] However, by overcoming these limitations, artificial neural networks have been able to achieve deep structures. This has enabled the construction of complex, highly expressive models, leading to groundbreaking results in diverse fields such as speech recognition, face recognition, object recognition, and character recognition.
[0067] Figure 2 is a diagram illustrating an example of a deep neural network.
[0068] A deep neural network (DNN) is an artificial neural network (ANN) with multiple hidden layers between the input layer and the output layer. It is a collection of machine learning models or algorithms that refer to deep neural networks (DNNs) with one or more hidden layers between the input layer and the output layer. The connections in the neural network are made from the input layer to the hidden layer, and from the hidden layer to the output layer.
[0069] Deep neural networks, like typical artificial neural networks, can model complex non-linear relationships. For example, in a deep neural network architecture for object recognition, each object can be represented as a hierarchical structure of basic image elements. Additional layers can then gradually aggregate features from lower layers. This characteristic of deep neural networks allows them to model complex data with a smaller number of units (nodes) compared to similarly implemented artificial neural networks.
[0070] While previous deep neural networks were typically designed as feed-forward neural networks, recent research has successfully applied deep learning structures to recurrent neural networks (RNNs). For example, deep neural network architectures have been applied to language modeling. Convolutional neural networks (CNNs) have been successfully applied to computer vision, with each successful application well documented. More recently, CNNs have been applied to acoustic modeling for Automatic Speech Recognition (ASR), and are considered more successful than existing models. Deep neural networks can be trained using the standard error backpropagation algorithm, where weights are updated using stochastic gradient descent using equations.
[0071] Many attempts have been made to predict physiological information using deep neural networks and machine learning. In the implementation of mobile healthcare systems, multi-task learning (MTL) is a crucial approach for performing multiple tasks with limited resources. In this invention, we propose a complex value-based multi-task learning (MTL) algorithm model that simultaneously processes EEG and rPPG signal waveforms. EEG and rPPG signal waveforms consist of complex numerical data that are simultaneously processed by a complex value-based neural network architecture. This complex process allows for more efficient and accurate extraction of user emotion and stress indices compared to real-valued single-task learning algorithms. This approach can be applied to the development of real-time emotion and stress monitoring systems and personalized mental stress assessment forms.
[0072] Briefly describe the telemetry method.
[0073] Development of a health monitoring system utilizing biometric information recognition sensors: By continuously collecting and monitoring pulse / movement information while the user is going about his or her daily life using a 3-axis acceleration sensor, the user's health status can be continuously monitored even when the user is not in a medical facility and is enjoying his or her daily life.
[0074] Wearable sensor unit for biometric information monitoring: A wearable sensor unit for biometric information monitoring is a wearable sensor unit composed of a main sensor layer and a supplementary sensor layer for monitoring biometric information such as temperature and movement status of a dynamic object including a human body.
[0075] Analysis of CNN-based remote-PPG to understand its limitations and sensitivities: This is a camera-based vital signs monitoring method based on a deep learning neural network, the convolutional neural network (CNN).
[0076] The Impact of Makeup on Remote PPG Monitoring: Camera-based remote PPG can non-contactly measure blood flow and pulse rate from human skin. Skin visibility is essential for remote PPG because the camera must penetrate deep into the skin tissue to capture light reflected from the skin, which transmits blood flow information. Facial makeup can affect this measurement by reducing the amount of light penetrating and reflecting from the skin.
[0077] Digital Remote Blood Pressure Management (DRM): This mobile blood pressure monitoring device eliminates the need for a doctor's visit, saving patients time. It can also improve personal health management and increase patient engagement.
[0078] Remote monitoring to help control high blood pressure: This is a technology that monitors blood pressure remotely through a monitor, as telemedicine technology can help lower the numbers in people with high blood pressure, which can reduce the risk of heart disease and stroke in the long term.
[0079] New insights into the origin of remote PPG signals in the visible and infrared: Remote photoplethysmography (remote PPG) is an optical measurement technique with potential applications in vital signs monitoring. Recent understanding of blood volume (BV) as a source of PPG signals has also led to consideration of the validity of remote SpO2 methodology. We demonstrate that remote PPG systems truly probe arterial blood. Green wavelengths probe cutaneous arteries, while red IR wavelengths also reach subcutaneous BVV. Because of their consistent penetration depth, the red IR diagnostic window has also been studied for SpO2 measurements in various skin conditions.
[0080] Remote photoplethysmography method Photoplethysmography method Electrocardiography measurement method Main principle Principle that uses the fact that hemoglobin reflects red light and absorbs green light Principle that uses the fact that hemoglobin reflects red light and absorbs green light Measures the electrical activity of the heart like an electrocardiograph, detects the electrical signal generated by the heartbeat Measurement device Camera LED, optical sensor ECG sensor Environment All body parts that can be captured by the camera Body extremities, wrist Near the heart User environment No attached or worn device Light source and sensor attached to the body Sensor mounted near the heart
[0081] Referring to Table 2, the remote photoplethysmography method applied in the present invention utilizes the principle that hemoglobin reflects red light and absorbs green light, and the camera captures the face. Multi-Task Learning (MTL) is a model learning method that simultaneously learns and predicts two or more tasks through a shared layer. By learning related tasks simultaneously, learned representations are shared, and thus, with good representations, each task can contribute to model learning. Useful information gained during learning can positively influence other tasks, contributing to a better model. Furthermore, by predicting multiple tasks simultaneously, a generalized model is learned that is more resistant to overfitting. Since it is created as a single model instead of having to develop two separate tasks in the past, it also greatly contributes to model weight reduction, making it more advantageous for application to mobile devices such as smartphones.
[0082] FIG. 3 is a block diagram illustrating the function of a device capable of estimating emotional and stress indices according to the present invention.
[0083] Referring to FIG. 3, a device (300) capable of estimating an emotional and stress index may include a processor (310), an EEG sensor (320), a memory (330), and a camera (340).
[0084] The EEG sensor (320) can sense the user's EEG signal and store the sensed EEG signal information in the memory (330).
[0085] The camera (340) can capture a user's face to obtain a facial image and store the obtained facial image in the memory (330). Here, the obtained facial image may be, for example, an RGB image.
[0086] The memory (330) can store EEG signal information received from the EEG sensor (320), facial image information received from the camera (340), information required for the processor (310) to output the rPPG result, and information required for the processor (310) to calculate the emotion and stress index.
[0087] In the present invention, the memory (330) may store a first deep learning model for outputting rPPG results and a second deep learning model for calculating emotional and stress indices.
[0088] The processor (310) can perform various necessary operations, such as applying the acquired facial image information and EEG signal information stored in the memory (330) to the first deep learning model and the second deep learning model stored in the memory (330) to output the rPPG result and calculate the emotion and stress index or store the result in the memory (330).
[0089] FIG. 4 is a diagram for explaining input data of a predetermined learned first deep learning model according to the present invention.
[0090] Referring to FIG. 4, the camera (340) can obtain original data (410) that captures a user's face and a face image (440) within the original data. The processor (310) can separate the original data (410) into a forehead region (420) and a cheek region (430), and detect a region of interest (ROI) in each of the forehead region (420) and the cheek region (430). The processor (310) can detect an image (450) of the forehead region and an image (460) of the cheek region from the face image (440) (here, the image (450) of the region of interest in the forehead region may be referred to as a first image, and the image (460) of the region of interest in the cheek region may be referred to as a second image).
[0091] This remote PPG (rPPG) method, which measures changes in photoperiod based on images, enables acquisition of biosignals using only ambient light as a light source, without a separate light source. LED green light is transmitted to the skin, where it is partially absorbed and reflected by the skin. This change can be measured and used to predict PPG and other biosignals. Using this characteristic, a convolutional neural network (CNN), a deep learning network primarily used in image processing, can be used to learn facial images and predict PPG using a pre-trained first deep learning model. Filters (kernels) are repeatedly applied to all areas of the forehead and cheek images extracted from the facial image to identify and learn patterns. The forehead and cheek regions are selected to train the model by utilizing the differences between the two body parts as meaningful information due to the time difference in blood flow from the heart. The acquired facial image can be acquired using the ambient light surrounding the user as a light source.
[0092] FIG. 5 is a flowchart illustrating a method for outputting an rPPG signal waveform from a facial image according to one embodiment of the present invention.
[0093] Referring to FIG. 5, the processor (310) detects a region of interest (ROI) (450, 460) from a face image / face video acquired from a camera (340). Here, the region of interest (ROI) is an image (450) of a region of interest in the forehead area in the face image / face video, and an image (460) of a region of interest in the cheek area. The processor (310) can perform image processing on an image (450) of a region of interest in the forehead area and an image (460) of a region of interest in the cheek area. As an example, the processor (310) can complexize (or convert into a complex domain) the image (450) of a region of interest in the detected forehead area and the image (460) of a region of interest in the detected cheek area and use them as input values for a predetermined trained first deep learning model. Herein, the predetermined trained first deep learning model can use a CNN model. The details of converting the image (450) of a region of interest in the forehead area and the image (460) of a region of interest in the detected cheek area into a complex domain are described below with reference to FIG. 6. As a specific example of the first deep learning model learned in FIG. 5, specific details of applying a convolutional neural network (CNN) model are shown in the lower right part of FIG. 5, but for specific details, please refer to FIG. 7.
[0094] FIG. 6 is an exemplary drawing for explaining the complexization of an image (450) for a region of interest of a detected forehead area and an image (460) for a region of interest of a detected cheek area, and FIG. 7 is an exemplary drawing for applying a convolutional neural network (CNN) model to a predetermined learned first deep learning model according to the present invention.
[0095] As illustrated in FIGS. 6 and 7, when a convolutional neural network (CNN) model is used as a predetermined learned first deep learning model according to the present invention, information on the cheeks and cheeks detected in a single face image is converted into a complex valued domain rather than a real-valued domain and trained together, thereby improving inference performance.
[0096] An image (450) for a region of interest in the forehead area and an image (460) for a region of interest in the cheek area are complexized and used as input values of a predetermined trained first deep learning model to which a convolutional neural network (CNN) model is applied. As illustrated in FIG. 6, the processor (310) converts the image (450) for a region of interest in the forehead area detected by the predetermined trained first deep learning model to which a convolutional neural network (CNN) model is applied into a complex domain by forming a complex number (z) having a real part (x) and an image (460) for a region of interest in the cheek area as an imaginary part (y). Alternatively, the processor (310) may convert the image (450) of the region of interest in the forehead area detected by a predetermined trained first deep learning model to which a convolutional neural network (CNN) model is applied into a complex domain by forming a complex number (z) with the imaginary part (y) and the image (460) of the region of interest in the cheek area as the real part (x).
[0097] As illustrated in FIG. 7, the processor (310) complexifies (or converts into a complex domain) the image (450) of the detected region of interest of the forehead area and the image (460) of the detected region of interest of the cheek area and inputs them into a predetermined trained first deep learning model to which a convolutional neural network (CNN) model is applied. That is, the processor (310) uses the image (450) of the forehead area and the image (460) of the cheek area together and converts them into a complex domain.
[0098] Unlike the existing deep learning in the real domain, we designed a deep learning algorithm model in the complex domain. The complex domain expresses phase signals, and has advantages in expressing and predicting biometric information compared to the real domain. The convolutional neural network (CNN) model simultaneously learns images of two areas of the face (450, 460) through a complex-valued CNN structure, so more accurate rPPG signal prediction is expected.
[0099] The processor (310) outputs the user's rPPG signal waveform based on the input complex data and the CNN through a predetermined learned first deep learning model to which a convolutional neural network (CNN) model is applied.
[0100] FIG. 8 is a diagram schematically illustrating a method for estimating emotional and stress indices according to one embodiment of the present invention.
[0101] As illustrated in FIG. 8, a first image for a region of interest in the forehead area and a second image for a region of interest in the cheek area are detected from an EEG signal waveform acquired through an EEG sensor (320) and a face image acquired through a camera, and the detected first image and the detected second image are complexized and applied to a predetermined trained first deep learning model, and the output user's rPPG signal is used as an input value for a predetermined trained second deep learning model to calculate an emotion and stress index.
[0102] FIG. 9 is an exemplary diagram for explaining the complexization of an EEG signal waveform (910) and an rPPG signal waveform (920), and FIG. 10 is an exemplary diagram for applying a convolutional neural network (CNN)-based multi-task learning algorithm model to a predetermined learned second deep learning model according to the present invention.
[0103] As illustrated in FIGS. 9 and 10, when a multi-task learning algorithm model based on a convolutional neural network (CNN) is used as a second deep learning model according to the present invention, the waveform information detected from the EEG signal waveform and the rPPG signal waveform acquired at the same time are converted into a complex valued domain rather than a real-valued domain and trained together, thereby improving inference performance.
[0104] The EEG signal waveform (910) and the rPPG signal waveform (920) are complexized and used as input values of a second deep learning model to which a convolutional neural network (CNN)-based multi-task learning algorithm model is applied. As illustrated in FIG. 9, the processor (310) converts the sensed EEG signal waveform (910) into a complex domain by forming a complex number (z) having a real part (x) and the output rPPG signal waveform (920) as an imaginary part (y) in the second deep learning model to which a convolutional neural network (CNN)-based multi-task learning algorithm model is applied. Alternatively, the processor (310) may convert the sensed EEG signal waveform (910) into a complex domain by configuring a complex number (z) with an imaginary part (y) and the output rPPG signal waveform (920) as a real part (x) in a predetermined learned second deep learning model to which a convolutional neural network (CNN)-based multi-task learning algorithm model is applied.
[0105] As illustrated in FIG. 10, the processor (310) complexifies (or converts to a complex domain) the EEG signal waveform (910) and the rPPG signal waveform (920) acquired through the EEG sensor (320) and inputs them into a second deep learning model to which a convolutional neural network (CNN)-based multi-task learning algorithm model (or may also be referred to as a CVMT model) has been trained.
[0106] That is, it is a proposed method to create a number of one complex domain using the EEG signal waveform (910) and the rPPG signal waveform (920) and use it as one input value during learning of a convolutional neural network (CNN)-based multi-task learning model, so that the multi-task learning model can be expected to learn additional information in the process.
[0107] Unlike the existing deep learning in the real domain, we designed a deep learning algorithm model in the complex domain. The complex domain expresses phase signals, and has advantages in expressing and predicting emotional and stress indices compared to the real domain. The multi-task learning algorithm model based on a convolutional neural network (CNN) simultaneously learns two biosignals (910, 920) through a complex-valued CNN structure, and more accurate emotional and stress index predictions can be expected.
[0108] The processor (310) calculates the user's emotion and stress index based on the input complex data and the convolutional neural network (CNN)-based multi-task learning algorithm model through a predetermined learned second deep learning model to which the convolutional neural network (CNN)-based multi-task learning algorithm model is applied.
[0109] Neural networks in the complex domain can improve the speed of backpropagation, the basic learning method of deep learning models, by 2 to 3 times, and require only half the factors such as weights and thresholds that are generally required.
[0110] Table 3 below is a table illustrating the accuracy of emotional and stress indices using the learned multi-task learning algorithm model according to the present invention by the processor (310).
[0111] Emotion Classification Stress Index Parameter MAE_emoRMSEMAE_strRMSE Our Model 0.43 25 0.50 31 0.53 60 0.58 0 4 1,609 547 Siamese Model 0.55 46 0.70 19 0.58 37 0.72 28 1,722 215
[0112] Referring to Table 3, the mean square error (MAE) of the emotion classification of the learned convolutional neural network (CNN)-based multi-task learning algorithm model (Our Model) according to the present invention is 0.12 less than the mean square error of the Siamese network model, and the root mean square deviation (RMSE) of the emotion classification of the learned multi-task learning algorithm model (Our Model) according to the present invention is also 0.20 less than the root mean square deviation of the Siamese network model, so that the accuracy of the emotion classification prediction is evaluated to be better than the Siamese model. In addition, the mean square error (MAE) of the stress index of the learned multi-task learning algorithm model (Our Model) according to the present invention is 0.05 less than the mean square error of the Siamese network model, and the root mean square deviation (RMSE) of the stress index of the learned multi-task learning algorithm model (Our Model) according to the present invention is also 0.05 less than the root mean square deviation of the Siamese model. As the value was as low as 0.14, the accuracy of stress index prediction was evaluated to be better than that of the Siamese network model.
[0113] In this way, we can see an improvement in performance compared to the Siamese network model that learns by sharing two inputs and weights, and we can confirm that we have succeeded in reducing parameters rather than actually creating two models.
[0114] As illustrated in FIGS. 7 and 10, the predetermined learned first deep learning model and the predetermined learned second deep learning model of the device capable of estimating the emotional and stress index according to the present invention can be input with the values input to the model changed to a complex domain, and a convolutional neural network (CNN) model can be applied, and in particular, a convolutional neural network (CNN)-based multi-task learning algorithm model can be applied to the predetermined learned second deep learning model.
[0115] In a situation where bio-information prediction through a CNN model of a single task in the existing general real-world domain is mainly done, this method helped to improve the learning speed by using the complex domain and to better calculate and use the brainwave signal and rPPG signal, which are the inputs of the learning data, and improved the learning quality by sharing the beneficial information of the two tasks by learning and predicting the emotion and stress index simultaneously through multi-task learning. In addition, it also improved in terms of model compression by creating one model rather than two models separately.
[0116] FIG. 11 is a block diagram illustrating a configuration of a remote device (1100) capable of estimating emotional and stress indices as another embodiment according to the present invention.
[0117] Referring to FIG. 11, the EEG sensor (320) and camera (340) constituting the device (300) capable of estimating the emotional and stress index of FIG. 3 can be replaced with a communication unit (1130).
[0118] The communication unit (1130) can receive the user's EEG signal waveform sensed remotely and the facial image acquired remotely and store them in the memory (1120).
[0119] The memory (1120) can store EEG signal information received from the communication unit (1130), facial image information, information required for the processor (1110) to output rPPG results, and information required for the processor (1110) to calculate an emotion and stress index.
[0120] In FIG. 11, the memory (1120) may include a first deep learning model for outputting rPPG results and a second deep learning model for calculating emotional and stress indices.
[0121] The processor (1110) can perform various necessary operations, such as applying the acquired facial image information and EEG signal information stored in the memory (1120) to the first deep learning model and the second deep learning model stored in the memory (1120) to output the rPPG result and calculate the emotion and stress index or store the result in the memory (1120).
[0122] The embodiments described above are combinations of components and features of the present invention in a predetermined form. Each component or feature should be considered optional unless explicitly stated otherwise. Each component or feature may be implemented without being combined with other components or features. Furthermore, it is also possible to form an embodiment of the present invention by combining some components and / or features. The order of operations described in the embodiments of the present invention may be changed. Some components or features of one embodiment may be included in another embodiment or may be replaced with corresponding components or features of another embodiment. It is self-evident that claims that do not have an explicit citation relationship in the patent claims may be combined to form an embodiment or may be incorporated as a new claim through a post-application amendment.
[0123] In the present invention, the processor (310, 1110) may be implemented by hardware, firmware, software, or a combination thereof. When implementing an embodiment of the present invention using hardware, ASICs (Application Specific Integrated Circuits) or DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), etc. configured to perform the present invention may be provided in the processor (310, 1110). The method for preventing user information leakage during user authentication according to the present invention may also be implemented as a computer-readable recording medium recording a program for executing the method on a computer.
[0124] It will be apparent to those skilled in the art that the present invention can be embodied in other specific forms without departing from the essential characteristics thereof. Therefore, the above detailed description should not be construed as limiting in any respect, but rather as illustrative. The scope of the present invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the scope of equivalents of the present invention are intended to be included within the scope of the present invention.
[0125] A system that estimates emotional and stress indices from brainwave signals and facial images can be used industrially in various industries, including the healthcare industry.
Claims
1. In a device capable of estimating emotional and stress indices, An EEG sensor for sensing the user's electroencephalogram (EEG) signals; A camera for capturing the face of the user and obtaining a facial image; and From the acquired facial image, a first image for the forehead region and a second image for the cheek region are detected, The detected first image and the detected second image are complexized and applied to a predetermined learned first deep learning model to output the user's remote PPG (remote Photoplethysmography, rPPG) signal waveform, An emotion and stress estimation device, comprising a processor that applies the sensed EEG signal waveform and the output rPPG signal waveform to a predetermined learned second deep learning model to calculate the emotion and stress index of the user.
2. In paragraph 1, The above processor, An emotion and stress estimation device that converts the first image detected above into a complex domain by forming a complex number having a real part and the second image as an imaginary part, and then applies it to the first deep learning model that has been trained.
3. In paragraph 1, The above processor, An emotion and stress estimation device that converts the detected first image into a complex domain by forming a complex number with the imaginary part of the first image and the real part of the second image, and then applies the converted complex number to the first deep learning model that has been trained.
4. In paragraph 1, An emotion and stress estimation device, wherein the first deep learning model learned above is a convolutional neural network (CNN) model.
5. In paragraph 1, The above-mentioned second deep learning model is an emotion and stress estimation device that is a multi-task learning algorithm model based on a convolutional neural network (CNN).
6. In paragraph 1, An emotion and stress estimation device, wherein the acquired facial image is acquired using ambient light around the user as a light source.
7. In paragraph 1, An emotion and stress estimation device further comprising a communication unit that transmits information on the user's emotion and stress index calculated above to a user terminal or server.
8. In paragraph 1, The above processor is an emotion and stress estimation device that applies the sensed EEG signal waveform as a real part and the output rPPG signal waveform as an imaginary part, or the sensed EEG signal waveform as an imaginary part and the output rPPG signal waveform as a real part, to the predetermined learned second deep learning model.
9. In a device capable of estimating emotional and stress indices, A communication unit that receives an electroencephalogram (EEG) signal sensed about a user and a facial image captured about the user's face; and From the acquired images, a first image for the forehead region and a second image for the cheek region are detected, The detected first image and the detected second image are complexized and applied to a predetermined learned first deep learning model to output the user's remote PPG (remote Photoplethysmography, rPPG) signal waveform, An emotion and stress estimation device, comprising a processor that applies the sensed EEG signal waveform and the output rPPG signal waveform to a predetermined learned second deep learning model to calculate the emotion and stress index of the user.
10. In a method for estimating emotional and stress indices, A step of sensing a user's electroencephalogram (EEG) signal; A step of photographing the face of the user to obtain a facial image; A step of detecting a first image for the forehead region from the acquired image; A step of detecting a second image for a cheek area from the acquired image; A step of complexizing the detected first image and second image; A step of applying the complex data to a predetermined learned first deep learning model to output a remote PPG (remote Photoplethysmography, rPPG) signal waveform of the user; and A method for estimating emotion and stress, comprising a step of applying the sensed EEG signal waveform and the output rPPG signal waveform to a predetermined learned second deep learning model to calculate the emotion and stress index of the user.
11. In paragraph 10, An emotion and stress estimation device further comprising a step of transmitting information on the user's emotion and stress index calculated above to a user terminal or a server.
12. A computer-readable recording medium having recorded thereon a program for executing the emotion and stress estimation method described in either of Article 10 or Article 11 on a computer.
Citation Information
Patent Citations
Method for providing NFT issuance survice based on location data
KR1020240177037A
Camping car usage guide system and method
KR1020250018648A
System for estimating emotion and stress from electroencephalogram and face image
KR102699988B1
KR20230148073A