Apparatus for estimating biosignals in non-contact manner

KR103001026B1Inactive Publication Date: 2026-08-05KWANGWOON UNIVERSITY INDUSTRY ACADEMIC COLLABORATION FOUNDATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020220183613
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2026-08-05
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure 112022139373521-PAT00003_ABST
    Figure 112022139373521-PAT00003_ABST
Patent Text Reader

Abstract

The device for estimating biosignals in a non-contact manner according to the present invention may include: a face image acquisition unit for acquiring a face image of a user; and a processor for detecting an image of a region of interest in the face image, converting the detected image of the region of interest into a quaternion-valued domain, applying it to a predetermined learned multi-task learning model using a Convolutional Neural Network (CNN) model, and outputting a photoplethysmography (PPG) signal waveform, an oxygen saturation signal waveform, and a blood pressure signal waveform of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a device for estimating biosignals, and more specifically, to a device and method for estimating biosignals from a face image in a non-contact or remote manner. Background Technology

[0002] Mental health is important to humans. Mental stress can lead to heart disease, diabetes, cancer, and other physical disorders such as abdominal pain. Methods such as stress assessments by medical professionals are not suitable for continuously monitoring stress because they depend on the expert evaluation. Consequently, a continuous and objective method for monitoring stress is also required.

[0003] Remote health monitoring technology is based on communication systems such as mobile phones or online health portals. Such technology may be in high demand for continuous patient monitoring even after pandemics like COVID-19 have ended. It is used to measure users' physiological signals based on facial video streams using cameras. This technology can be used not only for infectious diseases but also for monitoring the vital signs of infants and young children, as well as for the elderly or mental health monitoring.

[0004] Furthermore, the paradigm of medical services is shifting from receiving treatment at a hospital after the onset of disease to a model where individuals manage their own health and prevent disease. While wearable devices and smart devices are currently widely used, measuring biometric information using smart devices has clear limitations due to the conditions of purchasing, wearing, and continuous use of healthcare wearable devices. Therefore, we believe that non-contact biosignal measurement through video can serve as a solution to this problem.

[0005] Research exists on assessing stress levels by sensing biosignals using PPG sensors in a non-invasive manner. However, due to the current COVID-19 pandemic, technologies capable of remotely monitoring health in a non-invasive manner have become significantly important. The Centers for Disease Control and Prevention (CDCP) in Korea recommends using remote health strategies whenever possible to reduce the risk of COVID-19 in medical settings. Therefore, new methods are required that go beyond existing biometric monitoring methods using sensors that require physical contact.

[0006] However, to date, there are no studies or products available for estimating human vital signs and calculating stress levels as remote health monitoring technology. The problem to be solved

[0007] The technical problem to be solved by the present invention is to provide a device that estimates biosignals in a non-contact manner.

[0008] Another technical objective of the present invention is to provide a method for estimating biosignals in a non-contact manner.

[0009] Another technical objective of the present invention is to provide a computer-readable recording medium that records a computer program for executing a method of estimating biosignals in a non-contact manner.

[0010] The technical problems to be solved by the present invention are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which the present invention belongs from the description below. means of solving the problem

[0011] A device for estimating biosignals in a non-contact manner according to the present invention, for achieving the above technical objectives, may include: a face image acquisition unit for acquiring a face image of a user; and a processor for detecting an image of a region of interest in the face image, converting the detected image of the region of interest into a quaternion-valued domain, applying it to a predetermined learned multi-task learning model using a Convolutional Neural Network (CNN) model, and outputting a photoplethysmography (PPG) signal waveform, an oxygen saturation signal waveform, and a blood pressure signal waveform of the user.

[0012] The processor can convert the image of the detected region of interest in the predetermined learned multi-task learning model into a quaternion value domain and then use the CNN to output the user's PPG signal waveform, oxygen saturation signal waveform, and blood pressure signal waveform. The processor can calculate the user's stress index based on the output PPG signal waveform, the output oxygen saturation signal waveform, and the output blood pressure signal waveform.

[0013] The image of the region of interest may be a full face image. The acquired face image may be obtained using ambient light surrounding the user as a light source.

[0014] The processor may calculate the user's heart rate from the user's PPG signal waveform, calculate oxygen saturation from the oxygen saturation signal waveform, and calculate the user's blood pressure value from the blood pressure signal waveform. The blood pressure signal waveform includes a systolic blood pressure signal waveform and a diastolic blood pressure signal waveform, and the calculated blood pressure value may include a systolic blood pressure value and a diastolic blood pressure value.

[0015] The above device may further include a communication unit that transmits information regarding the calculated user's heart rate or information regarding the calculated blood pressure value to a user terminal or a server.

[0016] A method for estimating biosignals in a non-contact manner according to the present invention, for achieving other technical objectives as described above, may include: a step of acquiring a face image of a user; a step of detecting an image of a region of interest in the face image; a step of converting the detected image of the region of interest into a quaternion-valued domain and applying it to a predetermined learned multi-task learning model using a Convolutional Neural Network (CNN) model; and a step of outputting the user's Photoplethysmography (PPG) signal waveform, oxygen saturation signal waveform, and blood pressure signal waveform from the multi-task learning model.

[0017] The above method may include a step of converting the image of the detected region of interest in the above-described learned multi-task learning model into a quaternion value domain, and then using the CNN to output the user's PPG signal waveform, the oxygen saturation signal waveform, and the blood pressure signal waveform.

[0018] The above method may further include the step of calculating the user's stress index based on the output PPG signal waveform, the output oxygen saturation signal waveform, and the output blood pressure signal waveform.

[0019] In the face image acquisition step of the above method, the acquired face image can be acquired by using ambient light around the user as a light source.

[0020] The above method may further include the steps of calculating the user's heart rate from the user's PPG signal waveform, calculating the oxygen saturation from the oxygen saturation signal waveform, and calculating the user's blood pressure value from the blood pressure signal waveform. Effects of the invention

[0021] In situations where infectious diseases such as COVID-19 persist, it is effective to monitor patients' health by estimating biosignals non-contact and based on video.

[0022] It is a biosignal measurement technology designed for non-contact measurement convenience, capable of simultaneously measuring multiple biometric data in real time, and can be applied without restrictions in various fields such as stress management, drowsiness detection while driving, and biometric monitoring of infants and young children using biometric data.

[0023] Through the cutter-based multitask learning according to the present invention, PPG, oxygen saturation, and blood pressure are learned and predicted simultaneously, thereby helping to share beneficial information from the two tasks for learning. In addition, creating a single model rather than creating two separate models offers advantages in terms of model compression. Thus, there is an advantage in that the model can be improved and applied more freely to various fields.

[0024] The Cutterion-based multitask learning according to the present invention has the effect of significantly higher accuracy when predicting or estimating PPG, oxygen saturation, and blood pressure values ​​compared to existing models.

[0025] The effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below. Brief explanation of the drawing

[0026] The accompanying drawings, which are included as part of the detailed description to aid in understanding the present invention, provide embodiments of the present invention and explain the technical concept of the present invention together with the detailed description. Figure 1 is a diagram illustrating the layer structure of an artificial neural network. Figure 2 is a diagram illustrating an example of a deep neural network. FIG. 3 is a block diagram illustrating the configuration of a biosignal estimation device in a non-contact manner according to an embodiment of the present invention. FIG. 4 is a diagram for schematically explaining a non-contact biosignal estimation method according to one embodiment of the present invention. FIG. 5 is a diagram for specifically explaining a non-contact biosignal estimation method according to one embodiment of the present invention. FIG. 6 is a diagram illustrating a multi-task learning model for estimating biosignals in a non-contact manner according to the present invention. Figure 7 is a diagram illustrating the simulation results of the mean absolute error for blood pressure, oxygen saturation, and PPG between the quaternion-based multitask learning model according to the present invention and existing models. Specific details for implementing the invention

[0027] Hereinafter, preferred embodiments according to the present invention will be described in detail with reference to the accompanying drawings. The detailed description disclosed below, together with the accompanying drawings, is intended to describe exemplary embodiments of the present invention and is not intended to represent the only embodiment in which the present invention may be practiced. The following detailed description includes specific details to provide a complete understanding of the present invention. However, those skilled in the art will know that the present invention may be practiced without such specific details.

[0028] In some cases, to avoid obscuring the concept of the present invention, known structures and devices may be omitted or illustrated in the form of block diagrams focusing on the core functions of each structure and device. Additionally, throughout this specification, the same components are described using the same reference numerals.

[0029] Predicting PPG and respiratory rate (RR) based on facial images is safe for contact-transmitted diseases such as COVID-19. It is particularly advantageous from the perspective of diseases requiring continuous monitoring, such as diabetes, as it relies solely on imaging. This technology allows for the measurement of biometric indices and information anytime and anywhere for diseases that place a significant burden on both the public and households in terms of medical costs due to an aging population and the increase in chronic illnesses.

[0030] This invention proposes an algorithmic model that predicts biometric information using only facial data in a non-contact state, and proposes a method for predicting stress levels and various diseases using such biometric information. Generally, PPG (or heart rate) signals from a PPG (Photoplethysmography) sensor have been used to estimate stress levels. In this invention, respiratory rate signals are used in addition to PPG signals. The algorithmic model proposed in this invention analyzes collected data and accurately estimates stress levels.

[0031] Before describing the present invention, we will explain artificial intelligence (AI), machine learning, and deep learning. The easiest way to understand the relationship between these three concepts is to visualize three concentric circles. Artificial intelligence is the largest circle, followed by machine learning, and deep learning, which is leading the current AI boom, can be considered the smallest circle.

[0032] The concept of artificial intelligence first emerged at the Dartmouth Conference hosted by Professor John McCarthy at Dartmouth College in the United States in 1956, and it has been growing explosively in recent years. This growth has been further accelerated, particularly since 2015, by the introduction of GPUs that provide rapid and powerful parallel processing capabilities. The advent of the Big Data era, characterized by explosively increasing storage capacity and a flood of data across all domains—including images, text, and mapping data—has also had a significant impact on this growth trend.

[0033] Artificial Intelligence - Realizing human intelligence in machines

[0034] In 1956, the pioneers of artificial intelligence dreamed of ultimately creating complex computers with characteristics similar to human intelligence. While artificial intelligence that thinks like a human, possessing human senses and thinking abilities, is called 'General AI,' the artificial intelligence achievable at the current level of technological development falls under the concept of 'Narrow AI.' Narrow AI is characterized by its ability to perform specific tasks with capabilities exceeding those of humans, such as image classification services on social media or facial recognition functions.

[0035] Machine Learning - A Specific Approach to Implementing Artificial Intelligence

[0036] Machine learning serves the role of automatically filtering spam from your inbox. Meanwhile, machine learning fundamentally uses algorithms to analyze data, learns through analysis, and performs judgments or predictions based on what it has learned. Therefore, its ultimate goal is not to directly code specific guidelines for decision criteria into the software, but rather to 'train' the computer itself through massive amounts of data and algorithms to learn how to perform tasks. Machine learning originated from concepts directly proposed by early artificial intelligence researchers, and its algorithmic methods include decision tree learning, inductive logic programming, clustering, reinforcement learning, and Bayesian networks. However, none of these have achieved general AI, which can be considered the ultimate goal, and it is true that early machine learning approaches often struggled to complete even narrow AI.

[0037] Currently, machine learning is achieving significant results in fields such as computer vision, but it has encountered a limitation in that a certain amount of coding work is involved throughout the entire process of implementing artificial intelligence, even without specific guidelines. For instance, when recognizing an image of a stop sign based on a machine learning system, the developer must directly code boundary detection filters that programmatically identify the start and end points of an object, shape detection systems that verify the surface of an object, and classifiers that recognize characters such as 'STO-P'. In this way, machine learning operates by recognizing images from 'coded' classifiers and 'learning' stop signs through algorithms.

[0038] While machine learning achieves sufficient performance for commercialization in image recognition, the accuracy can drop in specific situations where signs are obscured by fog or trees. The reason computer vision and image recognition have not yet reached human levels until recently is due to these recognition rate issues and frequent errors.

[0039] Deep Learning - A technology that enables complete machine learning

[0040] The biological characteristics of the human brain, particularly the connection structure of neurons, inspired artificial neural networks, another algorithm created by early machine learning researchers. However, unlike the brain, where any physically adjacent neurons can be interconnected, artificial neural networks have fixed layer connections and data propagation directions.

[0041] For example, when an image is cut into numerous tiles and input into the first layer of a neural network, the neurons repeat the process of passing data to the next layer until a final output is generated at the last layer. Each neuron is assigned a weight representing the accuracy of the input based on the task performed, and the final output is determined by summing all the weights. In the case of a stop sign, the image's characteristics—such as its octagonal shape, red color, text, size, and movement—are finely cut and 'inspected' by the neurons, and the neural network's task is to identify whether it is a stop sign. Here, a 'probability vector' is utilized to predict the result based on weights derived from sufficient data.

[0042] Deep learning is a form of artificial intelligence that has evolved from artificial neural networks, utilizing information input and output layers similar to the neurons in the brain to learn data. However, because even basic neural networks require a tremendous amount of computation, the commercialization of deep learning faced obstacles from the beginning. Nevertheless, researchers continued their work and succeeded in parallelizing algorithms that prove the concept of deep learning based on supercomputers. Furthermore, the emergence of GPUs, which are optimized for parallel processing, dramatically accelerated the computational speed of neural networks, leading to the arrival of true deep learning-based artificial intelligence.

[0043] Neural networks are highly likely to produce numerous incorrect answers during the 'learning' process. Returning to the example of the stop sign, to precisely adjust the weights of neuron inputs to always produce the correct answer regardless of weather conditions or day-night cycles, one might need to learn from hundreds, thousands, or perhaps even millions of images. Only when this level of accuracy is reached can the neural network be considered to have properly learned the stop sign. In 2012, Google and Stanford University Professor Andrew Ng implemented a 'Deep Neural Network' consisting of over 1 billion neural networks using 16,000 computers. Through this, they extracted and analyzed 10 million images from YouTube and succeeded in having the computer classify photos of people and cats. They enabled the computer to independently learn the process of recognizing and judging the shape and appearance of cats appearing in videos.

[0044] The image recognition capabilities of systems trained with deep learning have already surpassed those of humans. Furthermore, the scope of deep learning extends to areas such as identifying cancer cells in the blood and tumors in MRI scans. Google's AlphaGo learned the fundamentals of Go and further strengthened its neural network through the process of repeatedly playing matches against AIs similar to itself. The emergence of deep learning has enhanced the practicality of machine learning and expanded the scope of artificial intelligence. Deep learning subdivides tasks in every way possible that can be supported by computer systems. Deep learning-based technologies, such as driverless cars, improved preventive medicine, and more accurate movie recommendations, are already being used in our daily lives or are on the verge of practical application. Deep learning is regarded as both the present and the future of artificial intelligence, possessing the potential to realize the general AI that once appeared in science fiction.

[0045] Below, we will take a closer look at deep learning.

[0046] Deep learning is a type of artificial neural network (ANN) based on human neural network theory, and is a set of machine learning models or algorithms that refer to a deep neural network (DNN) composed of a layer structure and having one or more hidden layers (hereinafter referred to as intermediate layers) between the input layer and the output layer. Simply put, deep learning can be described as an artificial neural network with deep layers.

[0047] The human brain is estimated to be composed of 25 billion nerve cells. The brain consists of nerve cells, and each nerve cell (neuron) refers to a single nerve cell that forms a neural network. A nerve cell contains a cell body, a single axon (or nurite) which is a projection of the cell body, and usually several dendrites (or protoplasmic processes). Information exchange between these nerve cells is transmitted through junctions between nerve cells called synapses. While a single nerve cell appears very simple when viewed in isolation, when these nerve cells come together, they are capable of possessing human intelligence. The dendrites are the part that receives signals sent by other nerve cells (Input), while the axon is the long extension from the cell body that transmits signals to other nerve cells (Output). There is a connection called a synapse that links the axon and dendrite, which transmit signals between nerve cells; however, the signal is not transmitted unconditionally, but is only transmitted when the signal strength exceeds a certain value (threshold). In other words, not only is the connection strength different for each synapse, but it also determines whether or not to transmit a signal.

[0048] Artificial neural networks (ANNs), a field of artificial intelligence, are mathematical models modeled by mimicking the structure of the biological (typically human) brain (neural networks). In other words, artificial neural networks are implemented by imitating the information processing and transmission processes of these biological neurons. As they are implemented similarly to how the human brain solves problems, neural networks possess excellent parallelism because each neuron operates independently. Furthermore, since information is distributed across numerous connections, problems in a few neurons do not significantly affect the entire network; consequently, they are resilient to a certain level of error and possess the ability to learn from a given environment.

[0049] Deep neural networks can be viewed as descendants of artificial neural networks. They are the latest version of artificial neural networks, having overcome existing limitations and achieved success in areas where numerous artificial intelligence technologies had previously failed. When examining the modeling of artificial neural networks that mimic biological neural networks, biological neurons are modeled as nodes in terms of processing units, and synapses are modeled as weights in terms of connections, as shown in Table 1 below.

[0050] biological neural networks artificial neural networks cell body node dendrites input Axon output synapse weight

[0051] Figure 1 is a diagram illustrating the layer structure of an artificial neural network.

[0052] Just as human biological neurons perform meaningful tasks through the connection of multiple cells rather than a single one, artificial neural networks connect individual neurons to one another via synapses, creating multiple interconnected layers where the connection strength between layers can be updated using weights. In this way, these multi-layered structures and connection strengths are utilized in fields such as learning and cognition.

[0053] Each node is connected by weighted links, and the entire model learns by repeatedly adjusting these weights. Weights represent the importance of each node as a fundamental means for long-term memory. Simply put, an artificial neural network trains the entire model by initializing these weights and updating and adjusting them using the data set to be trained. Once training is complete, when a new input is received, it infers an appropriate output value. The learning principle of an artificial neural network can be viewed as the process by which intelligence is formed from the generalization of experience, and it operates in a bottom-up manner. In Figure 1, when there are two or more intermediate layers (i.e., 5 to 10), the layers are considered to be deep, and it is called a Deep Neural Network; the learning and inference model achieved through such a Deep Neural Network can be referred to as Deep Learning.

[0054] Artificial neural networks can perform a certain role even with only one intermediate layer (commonly referred to as a 'hidden layer') in addition to inputs and outputs, but as the complexity of the problem increases, the number of nodes or layers must be increased. Among these, adopting a multi-layered model by increasing the number of layers is effective, but its scope of application is limited due to the limitations that efficient learning is impossible and the amount of computation required to train the network is large.

[0055] However, as the existing limitations mentioned above have been overcome, artificial neural networks have become capable of adopting deep structures. This has enabled the construction of complex and highly expressive models, leading to the 발표 of groundbreaking results in various fields such as speech recognition, face recognition, object recognition, and character recognition.

[0056] Figure 2 is a diagram illustrating an example of a deep neural network.

[0057] A Deep Neural Network (DNN) is an Artificial Neural Network (ANN) composed of multiple hidden layers between an input layer and an output layer. It is a set of machine learning models or algorithms referring to a Deep Neural Network (DNN) that has one or more hidden layers between an input layer and an output layer. Connections in a neural network are formed from the input layer to the hidden layer, and from the hidden layer to the output layer.

[0058] Deep neural networks, like general artificial neural networks, can model complex non-linear relationships. For example, in a deep neural network structure for an object identification model, each object can be represented as a hierarchical composition of the basic elements of an image. In this case, additional layers can combine features from progressively gathered lower layers. This characteristic of deep neural networks enables the modeling of complex data with fewer units (nodes) compared to similarly performed artificial neural networks.

[0059] Previous deep neural networks were typically designed as feedforward networks, but recent research has successfully applied deep learning structures to Recurrent Neural Networks (RNNs). Examples include the application of deep neural network structures in the field of language modeling. In the case of Convolutional Neural Networks (CNNs), not only have they been successfully applied in the field of computer vision, but their respective successful applications are also well-documented. More recently, CNNs have been applied to the field of acoustic modeling for Automatic Speech Recognition (ASR), and are evaluated as having been applied more successfully than existing models. Deep neural networks can be trained using the standard backpropagation algorithm. In this case, weights can be updated through stochastic gradient descent using the equation below.

[0060] Many attempts are being made to predict physiological information using deep neural networks and machine learning. In the implementation of mobile medical systems, Multi-Task Learning (MTL) is an important approach for performing multiple tasks with limited resources. PPG signals and respiratory rate signals are simultaneously extracted from facial video streams using MTL. This invention proposes a complex value-based Multi-Task Learning (MTL) algorithm model that processes video streams simultaneously. Two facial regions consist of complex numerical data that is processed simultaneously within a neural network architecture of complex values. Through this complex process, PPG signals and respiratory rate signals can be extracted more efficiently and accurately by comparing them with actual value single-task learning algorithms. This approach can be applied to the development of real-time mental stress monitoring systems and personalized mental stress assessment forms.

[0061] Briefly explain the telemetry method.

[0062] Development of a health monitoring system utilizing biometric recognition sensors: By using a 3-axis accelerometer, pulse / exercise information is continuously collected and monitored during the user's daily life, allowing for continuous monitoring of health status even when the user is not in a medical facility and is enjoying their daily life.

[0063] Wearable sensor unit for monitoring biological information: The wearable sensor unit for monitoring biological information is a wearable sensor unit composed of a main sensor layer and a complementary sensor layer, intended to monitor biological information such as the temperature and movement state of a dynamic object including the human body.

[0064] Analysis of CNN-based remote-PPG to understand limitations and sensitivities: A camera-based vital signs monitoring method based on a deep learning neural network, the Convolutional Neural Network (CNN).

[0065] Impact of makeup on remote-PPG monitoring: Camera-based remote-PPG can non-contactly measure blood volume and pulse from a person's skin. Skin visibility is essential for remote PPG because the camera must penetrate deep into the skin tissue to capture light reflected from the skin, which transmits blood pulsation information. Using facial makeup can affect this measurement method by reducing the amount of light that penetrates and reflects from the skin.

[0066] Digital Remote Blood Pressure Management: As a portable blood pressure monitoring device, it eliminates the need for doctor visits, saving patients time. It can improve personal health maintenance and increase patient engagement.

[0067] Remote monitoring to help control high blood pressure: This is a technology that monitors blood pressure remotely via a monitor, as it helps lower levels in people suffering from high blood pressure through remote medical technology and can reduce the risk of heart disease and stroke in the long term.

[0068] New Insights into the Origin of Remote PPG Signals in Visible and Infrared Light: Remote photoplethysmography (remote PPG) is an optical measurement technique applicable to vital sign monitoring. With the recent understanding of blood volume changes (BV) as the source of PPG signals, the feasibility of remote SpO2 methodologies is also being considered. This demonstrates that remote PPG systems actually irradiate arterial blood. Green wavelengths irradiate cutaneous arteries, while red IR wavelengths also reach subcutaneous BVV. Due to its stable penetration depth, the red IR diagnostic window has also been studied for SpO2 measurements in various skin conditions.

[0069] Remote photoplethysm measurement method Photocirculatory flow measurement method ECG measurement method Main principles The principle utilizing the fact that hemoglobin reflects red light and absorbs green light The principle utilizing the fact that hemoglobin reflects red light and absorbs green light Measures the electrical activity of the heart like an electrocardiogram and detects electrical signals generated by heartbeats. Measuring device camera LED, optical sensor ECG sensor Usage environment All body parts that can be photographed with a camera extremities, wrists near the heart User environment No attachment or wearable device Light source and sensor attached to the body Sensor installed near the heart

[0070] Referring to Table 2, the remote photoplethysm measurement method to be applied in the present invention utilizes the principle that hemoglobin reflects red light and absorbs green light, and a camera photographs the face.

[0071] Accordingly, the present invention relates to remote measurement technologies for monitoring biometric information remotely or non-contactually. In particular, since infectious diseases such as COVID-19 are currently spreading and people are avoiding visiting hospitals or using contact sensors, the invention proposes an algorithm model that predicts biometric information using only facial information in a non-contact state.

[0072] Recently, there have been many attempts to predict physiological information using deep neural networks and machine learning. In the implementation of mobile medical systems, multi-task learning is an important approach for performing multiple tasks with limited resources. Using MTL from facial video streams, it was possible to simultaneously extract PPG, oxygen saturation (SpO2), and blood pressure (SBP, DBP) signals. In this invention, we propose a quaternion-based multi-task learning (algorithm) model that processes video streams simultaneously. The facial region consists of data processed simultaneously within a neural network architecture of quaternion values. Through this process, PPG, oxygen saturation, and blood pressure signals could be extracted more efficiently and accurately by comparing them with real-value single-task learning algorithms. This approach can be applied to the development of real-time mental stress monitoring systems and personalized mental stress assessment forms.

[0073] Multi-Task Learning (MTL) is a model training method that predicts by simultaneously learning two or more tasks through shared layers. By training related tasks concurrently, learned representations are shared, allowing each task to aid in the model's training by utilizing their superior representations. The useful information gained from training can positively influence other tasks, contributing to the development of a better model. Furthermore, by predicting multiple tasks simultaneously, the model is trained to be more generalized and resistant to overfitting. Additionally, since the two tasks were previously created separately but are now combined into a single model, it significantly aids in model lightweighting, making it more advantageous for application on mobile devices such as smartphones.

[0074] Methods that have been widely studied to date require contact sensors or lighting, and there has been much research on models that predict a single type of biometric information. However, through the present invention, it becomes possible to predict multiple types of biometric information simultaneously by learning and predicting two or more types of biometric information using only a face image.

[0075] Furthermore, by utilizing multi-task learning according to the present invention, two or more related tasks can be learned more efficiently together. Additionally, by reducing the model size and applying it to mobile devices, users can easily access and frequently use it. In fact, the paradigm of medical services is shifting from receiving treatment at a hospital after the onset of a disease to a model where individuals manage their own health and prevent disease. Although wearable devices and smart devices are currently widely used, there are clear limitations to measuring biometric information using smart devices, which involve conditions such as purchasing, wearing, and continuous use of healthcare wearable devices. Therefore, non-contact biosignal measurement via video can serve as a solution to this. As a biosignal measurement technology that considers the convenience of non-contact measurement, it can simultaneously measure multiple types of biometric information in real time. It is expected to be easily applicable without restrictions in various fields, such as stress management for office workers, drowsiness detection while driving, and biometric monitoring of infants and toddlers.

[0076] FIG. 3 is a block diagram for explaining the configuration of a biosignal estimation device in a non-contact manner according to an embodiment of the present invention, FIG. 4 is a diagram for schematically explaining a biosignal estimation method in a non-contact manner according to an embodiment of the present invention, and FIG. 5 is a diagram for specifically explaining a biosignal estimation method in a non-contact manner according to an embodiment of the present invention.

[0077] Referring to FIG. 3, a non-contact biosignal estimation device (300) may include a processor (310), a face image acquisition unit (320), a memory (330), and a communication unit (340). The memory (330) is electrically coupled to the processor (310) so as to store a deep learning model according to the present invention or various information such as results output from a deep learning model. The non-contact biosignal estimation device (300) may be in the form of a wearable device, a mobile device, etc.

[0078] Referring to FIGS. 3 and 4, the face image acquisition unit (320) can acquire a user's face image. Here, the user's face image may be an image received from a camera that has been captured by an external camera, or a face image acquired by a non-contact biosignal estimation device (300) that has captured the user's face. If the user's face image is an image captured through an external camera, the face image acquisition unit (320) may be replaced by a communication unit (330). However, if the user's face image is a face image acquired by a non-contact biosignal estimation device (300) that has captured the user's face, the face image acquisition unit (320) may be a camera, a shooting unit, etc., equipped in the biosignal estimation device (300). Here, the user's face image may be an RGB image.

[0079] Referring to FIGS. 3 and FIGS. 5, the face image acquisition unit (320) can acquire original data in which a user's face is captured and a face image within the original data. Referring to FIG. 5, the processor (310) can detect an image (500) of a region of interest (ROI) from the original data. Here, the image of the region of interest may be an image of the entire face, rather than an image of a specific part of the face.

[0080] The remote PPG (rPPG) method, which measures changes in photocirculation based on such images, enables the acquisition of biosignals using only ambient light as a light source without a separate light source. LED green light is transmitted to the skin, where it is partially absorbed and reflected; the resulting changes can be measured and used to predict PPG and other biosignals. As shown in Fig. 4, PPG can be predicted by training a face image using a CNN model in image processing.

[0081] The processor (310) learns by repeatedly applying a filter (kernel) to all areas of the Red, Green, and Blue channels extracted from the detected region of interest image (500) to find patterns. The reason for using RGB channels is that this learning method has the advantage of being able to make predictions by better utilizing information from the three color channels through learning through color differences.

[0082] The deep learning model that receives and performs the quaternion value shown on the right side of FIG. 5 as input may be a predetermined trained multi-task learning model using a CNN model. The multi-task learning model according to the present invention shown on the right side of FIG. 5 is specifically illustrated in FIG. 6.

[0083] FIG. 6 is a diagram illustrating a multi-task learning model for estimating biosignals in a non-contact manner according to the present invention.

[0084] The details regarding the conversion of the image (500) of the region of interest into a quaternion-valued domain are described in detail with reference to FIG. 6. The processor (310) converts the detected image (500) of the region of interest into a quaternion-valued domain. Then, as shown in FIG. 6, the converted quaternion values ​​are applied to a trained multitask learning model using a CNN model to output the PPG (Photoplethysmography) signal waveform, oxygen saturation (SpO2) signal waveform, and blood pressure signal waveform of the user (or subject, test subject).

[0085] The processor (310) can calculate the user's stress index based on at least one of the output PPG signal waveform, the output oxygen saturation signal waveform, and the output blood pressure signal waveform. The processor (310) can calculate the user's heart rate from the output PPG signal waveform and calculate the user's blood pressure value from the output blood pressure signal waveform. The blood pressure signal waveform includes a systolic blood pressure signal waveform and a diastolic blood pressure signal waveform, and the processor (310) can calculate the blood pressure value as a systolic blood pressure value (SBP) and a diastolic blood pressure value (DBP).

[0086] As shown in FIG. 6, the multitask learning (algorithm) model according to the present invention is trained by converting it into a quaternion-valued domain rather than a real-valued domain because it is desired to improve the performance of existing general CNN models and lighten the model.

[0087] In general, the number of parameters such as required weights and thresholds can be reduced to less than half. Additionally, a multitask learning method was applied to utilize RGB color channel information more efficiently.

[0088] The quaternion domain to be applied in the present invention can represent rotation based on the X-axis, Y-axis, and Z-axis. This means that three channels can be represented simultaneously. This is advantageous for representing the amount of information contained in data without loss during machine learning training.

[0089] In a situation where the prediction of biological information using single-task CNN models in the general real-valued domain is predominant, this method utilizes the quaternion domain to improve training speed and enable better computation and utilization of input data. Furthermore, by simultaneously learning and predicting PPG, oxygen saturation, and blood pressure through multi-task learning, it facilitates the sharing of beneficial information from both tasks. Creating a single model rather than two separate ones offers advantages in terms of model compression. Thus, it offers the advantage of improving the model and enabling more unrestricted application across various fields.

[0090] The processor (310) can control the communication unit (340) to transmit at least one of the information regarding the calculated user's heart rate, the calculated oxygen saturation information, and the calculated blood pressure value to the user terminal or server. In particular, the processor (310) can control the communication unit (340) to transmit the biosignal information that is out of the normal range to the user terminal or server when the calculated heart rate, oxygen saturation, or blood pressure value deviates from a predetermined normal range.

[0091] Figure 7 is a diagram illustrating the simulation results of the mean absolute error for blood pressure, oxygen saturation, and PPG between the quaternion-based multitask learning model according to the present invention and existing models.

[0092] The Quaternion multi-task model illustrated in FIG. 7 shows the mean absolute error (MAE) values ​​of the PPG signal waveform, blood pressure signal waveform, and oxygen saturation (SpO2) signal waveform using a quaternion-based multi-task learning model according to the present invention by the processor (310).

[0093] In the rPPGNET model, which was proposed for inputting face images and predicting PPG, the average absolute error of PPG was predicted to be 5.51, and in the CNN-based Deep-phy model, the average absolute error of PPG was predicted to be 4.86; however, in the Quaternion multi-task model according to the present invention, the average absolute error of PPG was predicted to be 4.61, outputting the PPG signal waveform with the highest accuracy. In addition, in the existing rPPGNET model, which had the best blood pressure measurement, the average absolute error of blood pressure (BP) was 2.95, but in the Quaternion multi-task model according to the present invention, the average absolute error was 1.18, confirming that the accuracy of blood pressure (BP) estimation is higher.

[0094] In addition, the average absolute error value for oxygen saturation was calculated to be 2.75 in the existing CNN-based Deep-phy model, which had the best oxygen saturation (SpO2) learning, but in the Quaternion multi-task model according to the present invention, the average absolute error value for oxygen saturation was calculated to be 1.24, confirming that the accuracy for oxygen saturation measurement is very high.

[0095] As such, compared to existing models, it can be seen that the accuracy is significantly higher when predicting or estimating PPG, oxygen saturation, and blood pressure values ​​using the Quaternion multi-task model according to the present invention.

[0096] The embodiments described above are combinations of the components and features of the present invention in a specific form. Each component or feature should be considered optional unless otherwise explicitly stated. Each component or feature may be implemented in a form not combined with other components or features. Additionally, it is possible to construct embodiments of the present invention by combining some components and / or features. The order of operations described in the embodiments of the present invention may be changed. Some components or features of one embodiment may be included in another embodiment, or may be replaced with corresponding components or features of another embodiment. It is obvious that embodiments may be constructed by combining claims that do not have an explicit citation relationship in the claims, or that new claims may be included by amendment after filing.

[0097] In the present invention, the processor (110) may be implemented by hardware, firmware, software, or a combination thereof. When implementing an embodiment of the present invention using hardware, the processor (110) may be equipped with ASICs (application specific integrated circuits) or DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), etc., configured to perform the present invention. It may also be implemented as a computer-readable recording medium that records a program for executing a method to prevent user information leakage during user authentication according to the present invention on a computer.

[0098] It is obvious to those skilled in the art that the present invention may be embodied in other specific forms without departing from the essential features of the invention. Accordingly, the foregoing detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.

Claims

Claim 1 A biosignal estimation device comprising: a face image acquisition unit for acquiring a face image of a user; and a processor for detecting an entire face area in the face image as an image of a region of interest and converting the entire face area corresponding to the detected image of the region of interest into a quaternion-valued domain, wherein the processor applies the converted quaternion value to a predetermined learned quaternion-based multi-task learning model using a Convolutional Neural Network (CNN) model to simultaneously output a photoplethysmography (PPG) signal waveform, an oxygen saturation signal waveform, and a blood pressure signal waveform of the user, and wherein the processor simultaneously learns the quaternion-based multi-task learning model by sharing information between three tasks, including PPG signal waveform prediction, oxygen saturation signal waveform prediction, and blood pressure signal waveform prediction, through a shared layer. Claim 2 delete Claim 3 A biosignal estimation device according to claim 1, wherein the processor calculates the user's stress index based on the output PPG signal waveform, the output oxygen saturation signal waveform, and the output blood pressure signal waveform. Claim 4 delete Claim 5 delete Claim 6 A biosignal estimation device according to claim 1, wherein the processor calculates the user's heart rate from the user's PPG signal waveform, calculates oxygen saturation from the oxygen saturation signal waveform, and calculates the user's blood pressure value from the blood pressure signal waveform. Claim 7 A biosignal estimation device according to claim 1, wherein the blood pressure signal waveform includes a systolic blood pressure signal waveform and a diastolic blood pressure signal waveform, and the calculated blood pressure value includes a systolic blood pressure value and a diastolic blood pressure value. Claim 8 A biosignal estimation device according to claim 6, further comprising a communication unit that transmits information regarding the user's heart rate calculated above or information regarding the blood pressure value calculated above to a user terminal or server. Claim 9 delete Claim 10 delete Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 delete

Citation Information

Patent Citations

  • A method for measuring a physiological parameter of a subject in a contactless manner

    KR1020210062534A