Generalist healthcare artificial intelligence system based on multimodal unconstrained sensor
The multimodal AI-based healthcare device addresses integration and scalability issues by preprocessing biosignals from non-restraining sensors, generating embedding vectors, and applying them to pre-trained models, enhancing efficiency and adaptability in healthcare item provision.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KWANGWOON UNIVERSITY INDUSTRY ACADEMIC COLLABORATION FOUNDATION
- Filing Date
- 2025-09-12
- Publication Date
- 2026-05-21
AI Technical Summary
Existing specialized AI systems in healthcare are limited by their inability to integrate information efficiently, leading to high costs and inflexibility in responding to new situations or complex problems, and they struggle with scalability when new input or output data is added.
A multimodal AI-based healthcare device that preprocesses biosignals using non-restraining sensors, generates embedding vectors through tokenization and embedding processes, and applies these vectors to pre-trained AI models to output healthcare items without requiring redesign when new data is added.
Enables efficient, flexible, and scalable healthcare item provision with high accuracy, reducing costs and improving adaptability to new situations.
Smart Images

Figure KR2025095580_21052026_PF_FP_ABST
Abstract
Description
Multimodal non-restrictive sensor-based generalist healthcare AI system
[0001] The present invention relates to a generalist healthcare artificial intelligence device and method capable of performing health status monitoring, risk detection, and personalized health management using multimodal data measured from a non-restraining sensor.
[0002] Artificial intelligence technology has made remarkable advancements in the medical field and healthcare sector. In particular, specialized AI is demonstrating results that surpass human capabilities in various areas, such as medical image analysis, patient condition assessment, pathology testing, and new drug development. This development of specialized AI has increased diagnostic accuracy and significantly improved the work efficiency of medical professionals in the medical and healthcare fields.
[0003] However, specialized artificial intelligence has several significant limitations. Since each AI system operates independently, integrating information is difficult, and costs increase because multiple systems must be operated simultaneously. Additionally, because it is optimized for specific tasks, it has the disadvantage of being difficult to respond flexibly to new situations or complex problems.
[0004] Furthermore, the paradigm of modern healthcare is shifting from experiential medicine and evidence-based medicine, which are based on the experience of medical professionals, to data-driven medicine, which analyzes and evaluates various data obtained from inside and outside the hospital.
[0005] In order to overcome the limitations of existing specialized artificial intelligence in the medical and healthcare fields and to respond to changes in the modern healthcare paradigm, the present invention proposes a generalist healthcare artificial intelligence device and a method thereof.
[0006] Generalist AI refers to a multi-purpose AI system capable of performing various tasks and domains. The generalist healthcare AI device and method proposed in this invention enable comprehensive diagnosis and treatment using various biometric information of a user, provide personalized medical care, and flexibly process various healthcare-related tasks, thereby allowing for expected effects of cost reduction and efficiency.
[0007] The technical problem to be solved by the present invention is to provide a device that efficiently provides one or more healthcare items using one or more unrestrained biosignals of a user.
[0008] Another technical objective of the present invention is to provide a method for efficiently providing one or more healthcare items using one or more unrestrained biosignals of a user.
[0009] Another technical objective of the present invention is to provide a computer-readable recording medium that records a program for executing on a computer a method of efficiently providing one or more healthcare items using one or more unrestrained biosignals of a user.
[0010] The technical problems to be solved by the present invention are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which the present invention belongs from the description below.
[0011] A multimodal artificial intelligence-based healthcare device according to the present invention, for achieving the above technical objectives, may include a processor that generates one or more tokenized preprocessed data separated by identification information of the unbounded sensor unit and the identification information of the unbounded sensor unit that has acquired the biometric information of a user, by adding the identification information of the unbounded sensor unit that has acquired the biometric information to the respective acquired biometric information, generates one or more embedding vectors separated by identification information of the unbounded sensor unit through a tokenization and embedding process of the generated tokenized preprocessed data, generates one or more embedding vectors separated by identification information of the unbounded sensor unit, applies the generated one or more embedding vectors to a first artificial intelligence model that has been trained in advance to generate one or more integrated embedding vectors separated by a second artificial intelligence model that has been trained in advance, and applies the generated one or more integrated embedding vectors to the respective second artificial intelligence models that have been trained in advance to output one or more healthcare items.
[0012] The processor can segment one or more of the acquired biometric information into the same number of data points, and for each segment, add identification information of the unbound sensor unit that acquired the biometric information to generate one or more tokenized preprocessed data separated by identification information of the unbound sensor unit.
[0013] A multimodal artificial intelligence-based healthcare device according to the present invention, for achieving the above technical objectives, comprises one or more unrestrained sensor units for acquiring biometric information of a user; applying the acquired one or more biometric information to a pre-trained first artificial intelligence model to generate one or more augmented biometric information for the acquired one or more biometric information; adding identification information of the unrestrained sensor units that acquired the biometric information to the corresponding generated augmented biometric information to generate one or more tokenized preprocessed data separated by identification information of the unrestrained sensor units; generating one or more embedding vectors separated by identification information of the unrestrained sensor units through a tokenization and embedding process of the generated tokenized preprocessed data; applying the generated one or more embedding vectors to a pre-trained second artificial intelligence model to generate one or more integrated embedding vectors separated by a pre-trained third artificial intelligence model; and applying the generated one or more integrated embedding vectors to the corresponding pre-trained third artificial intelligence model to output one or more healthcare items. It may include a processor.
[0014] A multimodal artificial intelligence-based healthcare device according to the present invention for achieving the above technical objectives may include: a communication unit that receives one or more biometric information of a user and identification information of an unbound sensor unit that has acquired the one or more biometric information; and a processor that generates one or more tokenized preprocessed data separated by identification information of the unbound sensor unit by adding the received identification information of the unbound sensor unit to the corresponding received biometric information, generates one or more embedding vectors separated by identification information of the unbound sensor unit through a tokenization and embedding process of the generated tokenized preprocessed data, generates one or more integrated embedding vectors separated by identification information of the unbound sensor unit by applying the generated one or more embedding vectors to a pre-trained first artificial intelligence model to generate one or more integrated embedding vectors separated by a pre-trained second artificial intelligence model, and outputs one or more healthcare items by applying the generated one or more integrated embedding vectors to the corresponding pre-trained second artificial intelligence model.
[0015] A multimodal artificial intelligence-based healthcare device according to the present invention for achieving the above technical objectives comprises: a communication unit that receives one or more biometric information of a user and identification information of an unrestrained sensor unit that has acquired the one or more biometric information; The system may include a processor that applies one or more received biometric information to a first artificial intelligence model trained in advance to generate one or more augmented biometric information for the one or more acquired biometric information, adds identification information of the received unbound sensor unit to the corresponding generated augmented biometric information to generate one or more tokenized preprocessed data separated by identification information of the unbound sensor unit, generates one or more embedding vectors separated by identification information of the unbound sensor unit through the tokenization and embedding process of the generated tokenized preprocessed data, applies the generated one or more embedding vectors to a second artificial intelligence model trained in advance to generate one or more integrated embedding vectors separated by a third artificial intelligence model trained in advance, and applies the generated one or more integrated embedding vectors to the corresponding third artificial intelligence model trained in advance to output one or more healthcare items.
[0016] A method for outputting a multimodal artificial intelligence-based healthcare item according to the present invention for achieving the above technical problem may include: a step of acquiring one or more biometric information of a user through an unbound sensor unit; a step of generating one or more tokenized preprocessed data separated by identification information of the unbound sensor unit by adding identification information of the unbound sensor unit that acquired the biometric information to the corresponding biometric information; a step of generating one or more embedding vectors separated by identification information of the unbound sensor unit through a tokenization and embedding process of the generated tokenized preprocessed data; a step of generating one or more integrated embedding vectors separated by a prebound second artificial intelligence model by applying the generated one or more embedding vectors to a pre-trained first artificial intelligence model; and a step of outputting one or more healthcare items by applying the generated one or more integrated embedding vectors to the corresponding pre-trained second artificial intelligence model.
[0017] The step of generating the above tokenized preprocessed data may be a step of segmenting the one or more acquired biometric information into the same number of data, and adding identification information of the unbound sensor unit that acquired the biometric information for each segment to generate one or more tokenized preprocessed data separated by identification information of the unbound sensor unit.
[0018] A method for outputting a multimodal artificial intelligence-based healthcare item according to the present invention, for achieving the above technical objectives, comprises the steps of: acquiring one or more biometric information of a user through an unbound sensor unit; applying the acquired one or more biometric information to a pre-trained first artificial intelligence model to generate one or more augmented biometric information for the acquired one or more biometric information; adding identification information of the unbound sensor unit that acquired the biometric information to the corresponding generated augmented biometric information to generate one or more tokenized preprocessed data classified by identification information of the unbound sensor unit; generating one or more embedding vectors classified by identification information of the unbound sensor unit through a tokenization and embedding process of the generated tokenized preprocessed data; applying the generated one or more embedding vectors to a pre-trained second artificial intelligence model to generate one or more integrated embedding vectors classified by a pre-trained third artificial intelligence model; and applying the generated one or more integrated embedding vectors to the corresponding pre-trained third artificial intelligence model to generate one or more It may include a step to output healthcare items.
[0019] A method for outputting a multimodal artificial intelligence-based healthcare item according to the present invention for achieving the above technical problem may include the steps of: receiving one or more biometric information of a user and identification information of an unbound sensor unit that has acquired the one or more biometric information through a communication unit; adding the received identification information of the unbound sensor unit to the corresponding received biometric information to generate one or more tokenized preprocessed data separated by the identification information of the unbound sensor unit; generating one or more embedding vectors separated by the identification information of the unbound sensor unit through a tokenization and embedding process of the generated tokenized preprocessed data; applying the generated one or more embedding vectors to a pre-trained first artificial intelligence model to generate one or more integrated embedding vectors separated by a pre-trained second artificial intelligence model; and applying the generated one or more integrated embedding vectors to the corresponding pre-trained second artificial intelligence model to output one or more healthcare items.
[0020] The multimodal artificial intelligence-based healthcare device and method according to the present invention preprocess one or more biosignals of a user, generate an embedding vector based on the preprocessed biosignals using an artificial intelligence model, and apply the generated embedding data to one or more healthcare artificial intelligence models respectively, thereby enabling the provision of high-accuracy healthcare items.
[0021] A multimodal artificial intelligence-based healthcare device and method according to the present invention preprocesses one or more biosignals of a user, generates an embedding vector based on the preprocessed biosignals using an artificial intelligence model, and applies the generated embedding data to one or more healthcare artificial intelligence models, thereby providing high scalability for adding / changing / deleting biosignals and adding / changing / deleting healthcare items to be provided.
[0022] The multimodal artificial intelligence-based healthcare device and method according to the present invention utilize an artificial intelligence model to generate an augmented biosignal based on a user's biosignal, preprocess the augmented biosignal, generate an embedding vector based on the preprocessed biosignal through another artificial intelligence model, and apply the generated embedding data to one or more healthcare artificial intelligence models respectively, thereby enabling the provision of high-accuracy healthcare items.
[0023] The effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below.
[0024] The accompanying drawings, which are included as part of the detailed description to aid in understanding the present invention, provide embodiments of the present invention and explain the technical concept of the present invention together with the detailed description.
[0025] Figure 1 is a diagram illustrating the layer structure of an artificial neural network.
[0026] Figure 2 is a diagram illustrating an example of a deep neural network.
[0027] Figure 3 is a diagram illustrating the data processing process of an existing specialized artificial intelligence model.
[0028] Figure 4 is a diagram illustrating the data processing process of an existing specialized artificial intelligence model with added input data.
[0029] Figure 5 is a diagram illustrating the data processing process of an existing specialized artificial intelligence model with added output results.
[0030] FIG. 6 is a diagram illustrating a block diagram for explaining the functions of a multimodal artificial intelligence-based healthcare device (600) according to the present invention.
[0031] FIG. 7 is an example diagram illustrating the data processing process of a multimodal artificial intelligence-based healthcare device (600) according to the present invention.
[0032] FIG. 8 is another example drawing for explaining the data processing process of a multimodal artificial intelligence-based healthcare device (600) according to the present invention.
[0033] FIG. 9 is another example drawing illustrating a block diagram to explain the function of a multimodal artificial intelligence-based healthcare device (900) according to the present invention.
[0034] FIG. 10 is an example diagram illustrating the data processing process of a multimodal artificial intelligence-based healthcare device (900) according to the present invention.
[0035] FIG. 11 is another example drawing for explaining the data processing process of a multimodal artificial intelligence-based healthcare device (900) according to the present invention.
[0036] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.
[0037] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0038] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0039] The terms used herein are merely for describing specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0040] Additionally, terms such as “…part,” “…unit,” “…module,” and “…device” described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.
[0041] In some cases, to avoid obscuring the concept of the present invention, known structures and devices may be omitted or illustrated in the form of block diagrams focusing on the core functions of each structure and device. Additionally, throughout this specification, the same components are described using the same reference numerals.
[0042] Furthermore, the components of the embodiments described with reference to each drawing are not limited to the respective embodiments and may be implemented to be included in other embodiments within the scope of maintaining the technical spirit of the present invention. It is also obvious that multiple embodiments may be re-implemented as a single embodiment that integrates multiple embodiments, even if a separate description is omitted.
[0043] Before describing the present invention, we will explain artificial intelligence (AI), machine learning, and deep learning. The easiest way to understand the relationship between these three concepts is to visualize three concentric circles. Artificial intelligence is the largest circle, followed by machine learning, and deep learning, which is leading the current AI boom, can be considered the smallest circle.
[0044] The concept of artificial intelligence first emerged at the Dartmouth Conference hosted by Professor John McCarthy at Dartmouth College in the United States in 1956, and it has been growing explosively in recent years. This growth has been further accelerated, particularly since 2015, by the introduction of GPUs that provide rapid and powerful parallel processing capabilities. The advent of the Big Data era, characterized by explosively increasing storage capacity and a flood of data across all domains—including images, text, and mapping data—has also had a significant impact on this growth trend.
[0045] Artificial Intelligence - Realizing human intelligence in machines
[0046] In 1956, the pioneers of artificial intelligence dreamed of ultimately creating complex computers with characteristics similar to human intelligence. While artificial intelligence that thinks like a human, possessing human senses and thinking abilities, is called 'General AI,' the artificial intelligence achievable at the current level of technological development falls under the concept of 'Narrow AI.' Narrow AI is characterized by its ability to perform specific tasks with capabilities exceeding those of humans, such as image classification services on social media or facial recognition functions.
[0047] Machine Learning - A Specific Approach to Implementing Artificial Intelligence
[0048] Machine learning serves the role of automatically filtering spam from your inbox. Meanwhile, machine learning fundamentally uses algorithms to analyze data, learns through analysis, and performs judgments or predictions based on what it has learned. Therefore, its ultimate goal is not to directly code specific guidelines for decision criteria into the software, but rather to 'train' the computer itself through massive amounts of data and algorithms to learn how to perform tasks. Machine learning originated from concepts directly proposed by early artificial intelligence researchers, and its algorithmic methods include decision tree learning, inductive logic programming, clustering, reinforcement learning, and Bayesian networks. However, none of these have achieved general AI, which can be considered the ultimate goal, and it is true that early machine learning approaches often struggled to complete even narrow AI.
[0049] Currently, machine learning is achieving significant results in fields such as computer vision, but it has encountered a limitation in that a certain amount of coding work is involved throughout the entire process of implementing artificial intelligence, even without specific guidelines. For instance, when recognizing an image of a stop sign based on a machine learning system, the developer must directly code boundary detection filters that programmatically identify the start and end points of an object, shape detection systems that verify the surface of an object, and classifiers that recognize characters such as 'STO-P'. In this way, machine learning operates by recognizing images from 'coded' classifiers and 'learning' stop signs through algorithms.
[0050] Machine learning training methods find the most suitable model by adjusting the model parameters in a direction that minimizes the error between the target value and the predicted value. Here, the predicted value refers to the value produced when an input value is fed into the model, i.e., the output value. For example, a model composed of an arbitrary number of convolution layers, bidirectional LSTMs, and feedforward layers modifies each convolution layer, bidirectional LSTM, and feedforward layer as training progresses so that it becomes a model with the smallest error from the target value.
[0051] While machine learning achieves sufficient performance for commercialization in image recognition, the accuracy can drop in specific situations where signs are obscured by fog or trees. The reason computer vision and image recognition have not yet reached human levels until recently is due to these recognition rate issues and frequent errors.
[0052] Deep Learning - A technology that enables complete machine learning
[0053] The biological characteristics of the human brain, particularly the connection structure of neurons, inspired artificial neural networks, another algorithm created by early machine learning researchers. However, unlike the brain, where any physically adjacent neurons can be interconnected, artificial neural networks have fixed layer connections and data propagation directions.
[0054] For example, when an image is cut into numerous tiles and input into the first layer of a neural network, the neurons repeat the process of passing data to the next layer until a final output is generated at the last layer. Each neuron is assigned a weight representing the accuracy of the input based on the task performed, and the final output is determined by summing all the weights. In the case of a stop sign, the image's characteristics—such as its octagonal shape, red color, text, size, and movement—are finely cut and 'inspected' by the neurons, and the neural network's task is to identify whether it is a stop sign. Here, a 'probability vector' is utilized to predict the result based on weights derived from sufficient data.
[0055] Deep learning is a form of artificial intelligence that has evolved from artificial neural networks, utilizing information input and output layers similar to the neurons in the brain to learn data. However, because even basic neural networks require a massive amount of computation, the commercialization of deep learning faced obstacles from the beginning. Nevertheless, researchers continued their work and succeeded in parallelizing algorithms that prove the concept of deep learning based on supercomputers. Furthermore, the emergence of GPUs, which are optimized for parallel processing, dramatically accelerated the computational speed of neural networks, leading to the advent of true deep learning-based artificial intelligence.
[0056] Neural networks are highly likely to produce numerous incorrect answers during the 'learning' process. Returning to the example of the stop sign, to precisely adjust the weights of neuron inputs to always produce the correct answer regardless of weather conditions or day-night cycles, one might need to learn from hundreds, thousands, or perhaps even millions of images. Only when this level of accuracy is reached can the neural network be considered to have properly learned the stop sign. In 2012, Google and Stanford University Professor Andrew Ng implemented a 'Deep Neural Network' consisting of over 1 billion neural networks using 16,000 computers. Through this, they extracted and analyzed 10 million images from YouTube and succeeded in having the computer classify photos of people and cats. They enabled the computer to independently learn the process of recognizing and judging the shape and appearance of cats appearing in the videos.
[0057] The image recognition capabilities of systems trained with deep learning have already surpassed those of humans. Furthermore, the scope of deep learning extends to areas such as identifying cancer cells in the blood and tumors in MRI scans. Google's AlphaGo learned the fundamentals of Go and further strengthened its neural network through the process of repeatedly playing matches against AIs similar to itself. The emergence of deep learning has enhanced the practicality of machine learning and expanded the scope of artificial intelligence. Deep learning subdivides tasks in every way possible that can be supported by computer systems. Deep learning-based technologies, such as driverless cars, improved preventive medicine, and more accurate movie recommendations, are already being used in our daily lives or are on the verge of practical application. Deep learning is regarded as both the present and the future of artificial intelligence, possessing the potential to realize the general AI that once appeared in science fiction.
[0058] Below, we will take a closer look at deep learning.
[0059] Deep learning is a type of artificial neural network (ANN) based on human neural network theory. It is a set of machine learning models or algorithms that refer to a deep neural network (DNN) composed of a layer structure and having one or more hidden layers (hereinafter referred to as intermediate layers) between the input layer and the output layer. Simply put, deep learning can be described as an artificial neural network with deep layers.
[0060] The human brain is estimated to be composed of 25 billion nerve cells. The brain consists of nerve cells, and each nerve cell (neuron) refers to a single nerve cell that forms a neural network. A nerve cell contains a cell body, a single axon (or nurite) which is a projection of the cell body, and usually several dendrites (or protoplasmic processes). Information exchange between these nerve cells is transmitted through junctions between nerve cells called synapses. While a single nerve cell appears very simple when viewed in isolation, when these nerve cells come together, they are capable of possessing human intelligence. The dendrites are the part that receives signals sent by other nerve cells (Input), while the axon is the long extension from the cell body that transmits signals to other nerve cells (Output). There is a connection called a synapse that links the axon and dendrite, which transmit signals between nerve cells; however, the signal is not transmitted unconditionally, but is only transmitted when the signal strength exceeds a certain value (threshold). In other words, not only is the connection strength different for each synapse, but it also determines whether or not to transmit a signal.
[0061] Artificial neural networks (ANNs), a field of artificial intelligence, are mathematical models modeled by mimicking the structure of the biological (typically human) brain (neural networks). In other words, artificial neural networks are implemented by imitating the information processing and transmission processes of these biological neurons. As they are implemented similarly to how the human brain solves problems, neural networks possess excellent parallelism because each neuron operates independently. Furthermore, since information is distributed across numerous connections, problems in a few neurons do not significantly affect the entire network; consequently, they are resilient to a certain level of error and possess the ability to learn from a given environment.
[0062] Deep neural networks can be viewed as descendants of artificial neural networks. They are the latest version of artificial neural networks, having overcome existing limitations and achieved success in areas where numerous artificial intelligence technologies had previously failed. When examining the modeling of artificial neural networks that mimic biological neural networks, biological neurons are modeled as nodes in terms of processing units, and synapses are modeled as weights in terms of connections, as shown in Table 1 below.
[0063] Biological Neural Network Artificial Neural Network Cell Body Node Dendrite Input Axon Output Synapse Weight
[0064] Figure 1 is a diagram illustrating the layer structure of an artificial neural network.
[0065] Just as human biological neurons perform meaningful tasks by connecting multiple cells rather than just one, artificial neural networks connect individual neurons to one another through synapses, creating multiple interconnected layers where the connection strength between layers can be updated using weights. In this way, they are utilized in fields for learning and cognition through their multi-layered structure and connection strengths.
[0066] Each node is connected by weighted links, and the entire model learns by repeatedly adjusting these weights. Weights represent the importance of each node as a fundamental means for long-term memory. Simply put, an artificial neural network trains the entire model by initializing these weights and updating and adjusting them with the data set to be trained. Once training is complete, when a new input is received, it infers an appropriate output value. The learning principle of an artificial neural network can be viewed as the process by which intelligence is formed from the generalization of experience, and it operates in a bottom-up manner. In Figure 1, when there are two or more intermediate layers (i.e., 5 to 10), the layers are considered to be deep, and it is called a Deep Neural Network; the learning and inference model achieved through such a Deep Neural Network can be referred to as Deep Learning.
[0067] Artificial neural networks can perform a certain role even with only one intermediate layer (commonly referred to as a 'hidden layer') in addition to inputs and outputs, but as the complexity of the problem increases, the number of nodes or layers must be increased. Among these, adopting a multi-layered model by increasing the number of layers is effective, but its scope of application is limited due to the limitations that efficient learning is impossible and the amount of computation required to train the network is large.
[0068] However, as the existing limitations mentioned above have been overcome, artificial neural networks have become capable of adopting deep structures. This has enabled the construction of complex and highly expressive models, leading to the 발표 of groundbreaking results in various fields such as speech recognition, face recognition, object recognition, and character recognition.
[0069] Figure 2 is a diagram illustrating an example of a deep neural network.
[0070] A Deep Neural Network (DNN) is an Artificial Neural Network (ANN) composed of multiple hidden layers between an input layer and an output layer. It is a set of machine learning models or algorithms referring to a Deep Neural Network (DNN) that has one or more hidden layers between an input layer and an output layer. Connections in a neural network are formed from the input layer to the hidden layer, and from the hidden layer to the output layer.
[0071] Deep neural networks, like general artificial neural networks, can model complex non-linear relationships. For example, in a deep neural network structure for an object identification model, each object can be represented as a hierarchical composition of the basic elements of an image. In this case, additional layers can combine features from progressively gathered lower layers. This characteristic of deep neural networks enables the modeling of complex data with fewer units (nodes) compared to similarly performed artificial neural networks.
[0072] Previous deep neural networks were typically designed as feedforward networks, but recent research has successfully applied deep learning structures to Recurrent Neural Networks (RNNs). Examples include the application of deep neural network structures in the field of language modeling. In the case of Convolutional Neural Networks (CNNs), not only have they been successfully applied in the field of computer vision, but their successful applications are also well-documented. More recently, CNNs have been applied to acoustic modeling for Automatic Speech Recognition (ASR) and are considered to have been more successful than existing models. Deep neural networks can be trained using the standard backpropagation algorithm. In this process, weights can be updated through stochastic gradient descent.
[0073] Various signals from the surrounding environment received by humans through their sensory organs can be expressed via a computer in the form of text, audio, images, and videos, and stored as data in the computer's internal storage device.
[0074] The high-dimensional data corresponding to the aforementioned text, audio, images, and videos stored in a computer consists of combinations of continuous '0's and '1's from a low-dimensional perspective, while from the perspective of a slightly higher-dimensional computer program, it consists of various structures, objects, or class instances defined in the programming language used by each program.
[0075] For existing AI technologies to learn, they must extract features—data capable of effectively representing high-dimensional data—from high-dimensional data such as text, audio, images, and videos that humans can process via computers; however, the methods for implementing and the terminology used to refer to this feature data differ across various AI models and the programming languages capable of implementing them.
[0076] The following describes key terms related to artificial intelligence technology used in the present invention.
[0077] Generalist AI
[0078] Generalist AI refers to an AI system capable of performing various domains and tasks using a single model. Unlike existing specialized AI systems that were tailored to specific tasks (e.g., image recognition, translation, gameplay), Generalist AI comprehensively possesses multiple capabilities, including language understanding, visual recognition, reasoning, and decision-making.
[0079] The core characteristic of generalist AI is its ability to flexibly adapt to new situations and tasks. Similar to human general intelligence, this refers to the ability to solve previously unencountered problems based on existing knowledge and experience. Large-scale language models such as GPT-4 and PaLM demonstrate some of these characteristics of generalist AI and can perform various tasks including coding, creation, analysis, and conversation.
[0080] However, current generalist AI remains at a limited level and is far from Artificial General Intelligence (AGI) in the true sense. To become a perfect generalist AI, capabilities such as common-sense reasoning, long-term memory, self-awareness, and emotion understanding must be further developed, and this remains one of the primary goals of AI research.
[0081] Multi-Modal AI
[0082] Multimodal AI refers to an AI system capable of simultaneously processing and understanding various types of data, such as text, images, voice, and video. Unlike existing AI systems that processed only a single form of data, multimodal AI can integrally interpret and process diverse forms of input. This is similar to how humans receive and process information through multiple senses, such as sight, hearing, and language.
[0083] The major application fields of multimodal AI are very extensive. For example, in the medical field, it can make diagnoses by comprehensively analyzing patients' medical images, clinical records, and voice data, while autonomous vehicles understand the driving environment by integrating camera images, LiDAR sensor data, and GPS information. Furthermore, recent large-scale language models such as GPT-4V (Vision) and Claude 3 are equipped with multimodal capabilities that can process images and text together.
[0084] The development of multimodal AI presents various technical challenges. Effectively integrating different data forms and identifying interrelationships is a core task, requiring appropriate preprocessing and representation learning that consider the characteristics of each modality. Furthermore, constructing large-scale multimodal datasets, developing efficient training methods, and improving inference speed are currently major topics of research.
[0085] Unrestrained sensing system
[0086] A non-restraint sensing system refers to a system composed of non-restraint sensors that measure biosignals in a non-contact manner without attaching any devices to the user's body. It utilizes non-contact sensors, such as radar, infrared sensors, and cameras, to detect and analyze the user's activities and biosignals.
[0087] Unlike wearable sensing devices, the non-restraint sensing system measures biosignals indirectly through sensors installed in the environment without the need for the user to wear it directly on their body. Additionally, it does not require separate management such as battery charging.
[0088] Due to these characteristics, non-restraining sensing systems are highly useful in applications requiring long-term monitoring, such as patient or elderly care and sleep analysis. In particular, by utilizing various spaces within a home environment and integrating with multimodal artificial intelligence, it is possible to continuously obtain complex user data, enabling personalized healthcare.
[0089] The present invention proposes a multimodal AI-based healthcare device and method that preprocesses one or more biosignals of a user, generates an embedding vector based on the preprocessed biosignals using an AI model, and applies the generated embedding data to one or more healthcare AI models.
[0090] Generally, the data processing process of an AI model that receives a user's biosignals and outputs healthcare items consists largely of a "biometric signal data input stage," a "tokenization stage," an "embedding stage," an "encoding stage," a "decoding stage," and a "healthcare item output stage."
[0091] In the biosignal data input stage, various biosignals are collected from the user. Basic biomarkers such as heart rate, blood pressure, and body temperature, as well as detailed signals like respiratory rate, skin conductivity, and electromyography, are acquired in real time. During this process, preprocessing tasks such as sensor noise removal and signal correction are performed.
[0092] The tokenization stage is the process of converting continuously incoming biosignals into an analyzable form. Long time-series data is divided into segments of a fixed length, and meaningful features are extracted from each segment. For example, results of heart rate variability frequency analysis or statistical characteristics of skin conductivity are calculated. The extracted features typically undergo a normalization process to enable effective processing by the model.
[0093] In the embedding stage, tokenized data is transformed into a vector space that AI models can understand. It is mapped into low-dimensional vectors capable of expressing the relationships between signals while preserving the unique characteristics of each biosignal. Additionally, positional embedding is performed to preserve the sequential information of time-series data, and temporal embedding is conducted to reflect the temporal context of the measurement point.
[0094] In the encoding stage, complex patterns in the embedded data are learned. Interactions and dependencies between different biosignals are captured through mechanisms such as multi-head attention, and more abstract features are extracted using a feedforward neural network.
[0095] In the decoding stage, target healthcare items are output based on the encoded information. The relationship between the encoded features and the healthcare items to be output is learned through cross-attention, and the final features are extracted through multiple layers of a neural network.
[0096] In the result output stage, the decoded information is converted into healthcare item indicators in a format that is easy for the user to understand. Healthcare items can be various indicators related to the user's healthcare, such as stress index, fatigue level, heart rate, respiratory rate, sleep stages, or fall history.
[0097] The artificial intelligence model mentioned in this invention refers to an artificial intelligence module that processes the encoding and embedding steps in the data processing process of the aforementioned artificial intelligence model.
[0098] Figure 3 is a diagram illustrating the data processing process of an existing specialized artificial intelligence model.
[0099] Referring to Figure 3, the data processing process of three artificial intelligence models that output three different healthcare items is shown.
[0100] AI MODEL #1 is a healthcare AI model that receives camera sensor data, a form of biometric information, as input and measures the user's stress index, a healthcare item; it is an existing specialized AI model designed to measure the Stress Index.
[0101] AI MODEL #2 is a healthcare AI model that receives UWB Radar (Ultra Wideband Radar) sensor data, a form of biometric information, as input and detects user falls, a healthcare function; it is an existing specialized AI model designed for fall detection.
[0102] AI MODEL #3 is a healthcare AI model that receives PVDF (Polyvinylidene Fluoride) piezoelectric sensor data, a form of biometric information, as input to determine the user's heart rate, a healthcare parameter; it is an existing specialized AI model designed for heart rate determination.
[0103] The data processing process of the specialized artificial intelligence model illustrated in Fig. 3 may cause various difficulties in scalability if used in an unrestricted sensing system.
[0104] First, the data processing process of specialized AI models faces difficulties in scalability when new input data is added.
[0105] Figure 4 is a diagram illustrating the data processing process of an existing specialized artificial intelligence model with added input data.
[0106] Figure 4 illustrates a case where UWB Radar data is newly added to the data processing process of the AI model (AI MODEL #1) for measuring the Stress Index of Figure 3, thereby expanding it into the data processing process of a multimodal AI model.
[0107] When new input data is added to such a specialized artificial intelligence model, it is necessary to redesign and modify the overall stages, including the data input stage, tokenization stage, embedding stage, and encoding / decoding stage. Therefore, in fact, AI MODEL #1 of FIG. 3 should be referred to as the new artificial intelligence model AI MODEL #4 in the extended structure of FIG. 4.
[0108] Secondly, the data processing process of specialized AI models faces difficulties in scalability when new output results are added.
[0109] Figure 5 is a diagram illustrating the data processing process of an existing specialized artificial intelligence model with added output results.
[0110] Figure 5 illustrates an expanded case in which a new fall detection output result is added during the data processing of the artificial intelligence model (AI MODEL #4) for measuring the stress index of Figure 4.
[0111] When new input data is added to such a specialized artificial intelligence model, redesign and modification are required up to the data embedding stage and the encoding / decoding stage. Therefore, in fact, AI MODEL #4 of FIG. 4 should be referred to as a new artificial intelligence model AI MODEL #5 in the extended structure of FIG. 5.
[0112] FIG. 6 is a diagram illustrating a block diagram for explaining the functions of a multimodal artificial intelligence-based healthcare device (600) according to the present invention.
[0113] Referring to FIG. 6, a multimodal artificial intelligence-based healthcare device (600) may include a processor (610), one or more non-restraining sensor units (620) and a memory (630). The non-restraining sensor unit (620) performs the function of acquiring a user's biometric information through a non-restraining sensor and, more preferably, may be composed of a plurality of non-restraining sensor units.
[0114] FIG. 7 is an example diagram illustrating the data processing process of a multimodal artificial intelligence-based healthcare device (600) according to the present invention.
[0115] Referring to FIG. 7, the multimodal artificial intelligence-based healthcare device according to the present invention includes one or more unrestrained sensor units. The unrestrained sensor units can acquire biometric information of a user and may be unrestrained sensors such as the illustrated camera sensor, UWB radar sensor, or PVDF piezoelectric element sensor.
[0116] Each unrestrained sensor unit can acquire camera sensor data containing the user's biometric information acquired from a camera sensor, UWB radar sensor data containing the user's biometric information acquired from a UWB radar sensor, and PVDF piezoelectric sensor data containing the user's biometric information acquired from a PVDF piezoelectric sensor.
[0117] The processor (610) can obtain identification information of one or more unconstrained sensor parts.
[0118] The processor (610) performs the function of generating tokenized preprocessed data by distinguishing according to the identification information of one or more unbound sensor parts acquired, adding the identification information of the unbound sensor part that acquired biometric information to the acquired biometric information, generating an embedding vector through the process of tokenization and embedding, and applying this to a pre-trained artificial intelligence model to generate an integrated embedding vector.
[0119] In addition, the function of outputting one or more healthcare items is performed by separately applying the aforementioned integrated embedding vector to each pre-trained artificial intelligence model capable of outputting healthcare items.
[0120] The memory (630) can store the acquired user's biometric information, tokenization preprocessed data, tokenized data output through the tokenization process, embedding data output through the embedding process, integrated embedding vector data, output results of healthcare items, a pre-trained artificial intelligence model, and various computational information generated during the processing of artificial intelligence.
[0121] Below, the functions of the processor (610) are described in more detail.
[0122] Referring to FIG. 7, the processor (610) can perform tokenization preprocessing, tokenization, and embedding processes included in the encoding unit using one or more acquired biometric information.
[0123] The processor (610) can generate one or more tokenized preprocessed data separated by identification information of the unbound sensor unit by adding identification information of the unbound sensor unit, which has acquired biometric information during the tokenization preprocessing process, to the corresponding acquired biometric information. That is, the processor (610) can generate tokenized preprocessed data by adding identification information, such as an ID that can identify the camera sensor, to camera sensor data containing the acquired user's biometric information. Similarly, tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the UWB radar sensor, to UWB radar sensor data containing the acquired user's biometric information, and tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the PVDF piezoelectric sensor, to PVDF piezoelectric sensor data containing the acquired user's biometric information.
[0124] More specifically, when the processor (610) adds identification information of the camera sensor to camera sensor data containing acquired user biometric information, it can generate tokenized preprocessed data by segmenting the acquired user biometric information into equal numbers of data and adding identification information of the camera sensor for each segment.
[0125] Subsequently, the processor (610) can generate one or more embedding vectors separated by identification information of the unconstrained sensor part through the tokenization and embedding process of the generated tokenization preprocessed data. That is, the processor (610) can perform the tokenization and embedding process of the generated tokenization preprocessed data separated by the camera sensor, UWB radar sensor, and PVDF piezoelectric element sensor of FIG. 7.
[0126] Additionally, the processor (610) can perform a data processing process of transmitting one or more embedding vectors, separated by identification information of the unconstrained sensor unit included in the encoding unit, as input to the Generalist AI Unit.
[0127] The generalist AI department may include a pre-trained AI model.
[0128] The artificial intelligence model included in the generalist artificial intelligence unit can provide the effect of not requiring the redesign or modification of each healthcare artificial intelligence model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the existing decoding unit when new input data is added or new output results are added during the data processing process of the multimodal artificial intelligence-based healthcare device according to the present invention. Therefore, in the present invention, the artificial intelligence model included in the generalist artificial intelligence unit is referred to as the generalist artificial intelligence model.
[0129] The processor (610) can apply all embedding vectors received from the encoding unit to the aforementioned pre-trained generalist artificial intelligence model to generate and transmit one or more integrated embedding vectors to be applied to each healthcare artificial intelligence model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0130] The above integrated embedding vector may include the contents of one or more embedding vectors separated by identification information of each unconstrained sensor part. The data of the integrated embedding vector may differ for each healthcare artificial intelligence model in the decoding part, and even if an embedding vector is applied to a specific healthcare artificial intelligence model in the decoding part, the weights of the embedding vectors for each unconstrained sensor part within the integrated embedding vector may change as learning progresses.
[0131] The processor (610) can output the pre-determined healthcare items by applying the integrated embedding vector received through the generalist AI unit to one or more healthcare AI models (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0132] FIG. 8 is another example drawing for explaining the data processing process of a multimodal artificial intelligence-based healthcare device (600) according to the present invention.
[0133] Referring to Fig. 8, an Augmentation AI Unit may be added to the front of the encoding unit of Fig. 7.
[0134] Referring to FIG. 8, the processor (610) can perform a data processing process to transmit one or more acquired biometric information as input to an Augmentation AI Unit.
[0135] The above-mentioned augmented artificial intelligence unit may include a pre-trained artificial intelligence model.
[0136] The artificial intelligence model included in the augmented artificial intelligence unit can generate augmented bio-information by augmenting bio-information transmitted as input during the data processing process of the multimodal artificial intelligence-based healthcare device according to the present invention, and in the present invention, the artificial intelligence model included in the augmented artificial intelligence unit is referred to as an augmented artificial intelligence model.
[0137] The processor (610) can transmit only one or more specific biometric information as input to the Augmentation AI Unit according to pre-set information.
[0138] Referring to FIG. 8, the processor (610) applies only camera sensor data and PVDF piezoelectric sensor data containing acquired user biometric information to the augmented artificial intelligence unit according to pre-set information to generate augmented camera sensor data and augmented PVDF piezoelectric sensor data.
[0139] The processor (610) can perform tokenization preprocessing, tokenization, and embedding processes included in the encoding unit using bio-information augmented according to pre-set information and acquired bio-information not augmented according to pre-set information.
[0140] The processor (610) can generate one or more tokenized preprocessed data separated by identification information of the unbound sensor unit by adding identification information of the unbound sensor unit, which has acquired biometric information during the tokenization preprocessing process, to the corresponding acquired biometric information. That is, the processor (610) can generate tokenized preprocessed data by adding identification information, such as an ID that can identify the camera sensor, to the augmented camera sensor data. Similarly, tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the UWB radar sensor, to UWB radar sensor data containing the acquired user's biometric information, and tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the PVDF piezoelectric sensor, to the augmented PVDF piezoelectric sensor data.
[0141] More specifically, when the processor (610) adds identification information of the camera sensor to the augmented camera data, it can generate tokenized preprocessed data by segmenting the acquired user's biometric information into the same number of data and adding identification information of the camera sensor for each segment.
[0142] Subsequently, the processor (610) can generate one or more embedding vectors separated by identification information of the unconstrained sensor part through the tokenization and embedding process of the generated tokenization preprocessed data. That is, the processor (610) can perform the tokenization and embedding process of the generated tokenization preprocessed data separated by the camera sensor, UWB radar sensor, and PVDF piezoelectric element sensor of FIG. 7.
[0143] Additionally, the processor (610) can perform a data processing process of transmitting one or more embedding vectors, separated by identification information of the unconstrained sensor unit included in the encoding unit, as input to the Generalist AI Unit.
[0144] The generalist AI department may include a pre-trained AI model.
[0145] The AI model included in the generalist AI unit can provide the effect of not requiring the redesign or modification of each healthcare AI model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the existing Decoding Unit when new input data is added or new output results are added during the data processing process of the multimodal AI-based healthcare device according to the present invention.
[0146] The processor (610) can apply all embedding vectors received from the encoding unit to the aforementioned pre-trained generalist artificial intelligence model to generate and transmit one or more integrated embedding vectors to be applied to each healthcare artificial intelligence model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0147] The above integrated embedding vector may include the contents of one or more embedding vectors separated by identification information of each unconstrained sensor part. The data of the integrated embedding vector may differ for each healthcare artificial intelligence model in the decoding part, and even if an embedding vector is applied to a specific healthcare artificial intelligence model in the decoding part, the weights of the embedding vectors for each unconstrained sensor part within the integrated embedding vector may change as learning progresses.
[0148] The processor (610) can output the pre-determined healthcare items by applying the integrated embedding vector received through the generalist AI unit to one or more healthcare AI models (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0149] FIG. 9 is another example drawing illustrating a block diagram to explain the function of a multimodal artificial intelligence-based healthcare device (900) according to the present invention.
[0150] Referring to FIG. 9, one or more unconstrained sensor units (630) constituting the multimodal artificial intelligence-based healthcare device (600) of FIG. 6 can be replaced with a communication unit (920).
[0151] The communication unit (920) can receive one or more biometric information and identification information of the unrestrained sensor unit that acquired the one or more biometric information from a remote source and store them in the memory (930).
[0152] FIG. 10 is an example diagram illustrating the data processing process of a multimodal artificial intelligence-based healthcare device (900) according to the present invention.
[0153] Referring to FIG. 10, a multimodal artificial intelligence-based healthcare device (900) according to the present invention receives camera sensor identification information, acquired camera sensor biometric information, UWB radar sensor identification information, acquired UWB radar sensor biometric information, PVDF piezoelectric element sensor identification information, and acquired PVDF piezoelectric element sensor biometric information through a communication unit.
[0154] The processor (910) performs the function of generating tokenized preprocessed data by distinguishing according to the identification information of one or more unbound sensor parts received, adding the identification information of the unbound sensor part that acquired biometric information to the acquired biometric information, generating an embedding vector through the process of tokenization and embedding, and applying this to a pre-trained artificial intelligence model to generate an integrated embedding vector.
[0155] In addition, the function of outputting one or more healthcare items is performed by separately applying the aforementioned integrated embedding vector to each pre-trained artificial intelligence model capable of outputting healthcare items.
[0156] The memory (930) can store the received user's biometric information, tokenization preprocessed data, tokenized data output through the tokenization process, embedding data output through the embedding process, integrated embedding vector data, output results of healthcare items, a pre-trained artificial intelligence model, and various computational information generated during the processing of artificial intelligence.
[0157] Below, the functions of the processor (910) are described in more detail.
[0158] Referring to FIG. 10, the processor (910) can perform tokenization preprocessing, tokenization, and embedding processes included in the encoding unit using one or more acquired biometric information.
[0159] The processor (910) can generate one or more tokenized preprocessed data separated by identification information of the unbound sensor unit by adding identification information of the unbound sensor unit, which has acquired biometric information during the tokenization preprocessing process, to the corresponding acquired biometric information. That is, the processor (910) can generate tokenized preprocessed data by adding identification information, such as an ID that can identify the camera sensor, to camera sensor data containing the acquired user's biometric information. Similarly, tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the UWB radar sensor, to UWB radar sensor data containing the acquired user's biometric information, and tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the PVDF piezoelectric sensor, to PVDF piezoelectric sensor data containing the acquired user's biometric information.
[0160] More specifically, when the processor (910) adds identification information of the camera sensor to camera sensor data containing the received user's biometric information, it can generate tokenized preprocessed data by segmenting the acquired user's biometric information into the same number of data and adding identification information of the camera sensor for each segment.
[0161] Subsequently, the processor (910) can generate one or more embedding vectors separated by identification information of the unconstrained sensor part through the tokenization and embedding process of the generated tokenization preprocessed data. That is, the processor (910) can perform the tokenization and embedding process on the generated tokenization preprocessed data separated by the received camera sensor identification information, UWB radar sensor identification information, and PVDF piezoelectric element sensor identification information of FIG. 10.
[0162] Additionally, the processor (910) can perform a data processing process of transmitting one or more embedding vectors, separated by identification information of the unconstrained sensor unit included in the encoding unit, as input to the Generalist AI Unit.
[0163] The AI model included in the generalist AI unit can provide the effect of not requiring the redesign or modification of each healthcare AI model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the existing Decoding Unit when new input data is added or new output results are added during the data processing process of the multimodal AI-based healthcare device according to the present invention.
[0164] The processor (910) can apply all embedding vectors received from the encoding unit to the aforementioned pre-trained generalist artificial intelligence model to generate and transmit one or more integrated embedding vectors to be applied to each healthcare artificial intelligence model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0165] The above integrated embedding vector may include the contents of one or more embedding vectors separated by identification information of each unconstrained sensor part. The data of the integrated embedding vector may differ for each healthcare artificial intelligence model in the decoding part, and even if an embedding vector is applied to a specific healthcare artificial intelligence model in the decoding part, the weights of the embedding vectors for each unconstrained sensor part within the integrated embedding vector may change as learning progresses.
[0166] The processor (910) can output the pre-determined healthcare items by applying the integrated embedding vector received through the generalist AI unit to one or more healthcare AI models (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0167] FIG. 11 is another example drawing for explaining the data processing process of a multimodal artificial intelligence-based healthcare device (900) according to the present invention.
[0168] Referring to Fig. 11, an Augmentation AI Unit may be added to the front of the encoding unit of Fig. 10.
[0169] Referring to FIG. 11, the processor (910) can perform a data processing process to transmit one or more acquired biometric information as input to an Augmentation AI Unit.
[0170] The above-mentioned augmented artificial intelligence unit may include a pre-trained artificial intelligence model.
[0171] The artificial intelligence model included in the above-mentioned augmented artificial intelligence unit can generate augmented bio-information by augmenting bio-information transmitted as input during the data processing process of the multimodal artificial intelligence-based healthcare device according to the present invention.
[0172] The processor (910) can transmit only one or more specific biometric information as input to the Augmentation AI Unit according to pre-set information.
[0173] Referring to FIG. 11, the processor (910) applies only the camera sensor data and PVDF piezoelectric sensor data containing the received user's biometric information to the augmented artificial intelligence unit according to pre-set information to generate augmented camera sensor data and augmented PVDF piezoelectric sensor data.
[0174] The processor (910) can perform tokenization preprocessing, tokenization, and embedding processes included in the encoding unit using biometric information augmented according to pre-set information and received biometric information not augmented according to pre-set information.
[0175] The processor (910) can generate one or more tokenized preprocessed data separated by identification information of the unbound sensor unit by adding identification information of the unbound sensor unit, which has acquired biometric information during the tokenization preprocessing process, to the corresponding acquired biometric information. That is, the processor (910) can generate tokenized preprocessed data by adding received identification information, such as an ID that can identify the camera sensor, to the augmented camera sensor data. Similarly, tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the received UWB radar sensor, to UWB radar sensor data containing the received user's biometric information, and tokenized preprocessed data can be generated by adding identification information, such as an ID that can identify the received PVDF piezoelectric sensor, to the augmented PVDF piezoelectric sensor data.
[0176] More specifically, when the processor (910) adds identification information of the camera sensor to the augmented camera sensor data, it can generate tokenized preprocessed data by segmenting the received user's biometric information into the same number of data and adding identification information of the received camera sensor for each segment.
[0177] Subsequently, the processor (910) can generate one or more embedding vectors separated by identification information of the unconstrained sensor part through the tokenization and embedding process of the generated tokenization preprocessed data. That is, the processor (910) can perform the tokenization and embedding process of the generated tokenization preprocessed data separated by the camera sensor, UWB radar sensor, and PVDF piezoelectric element sensor of FIG. 7.
[0178] Additionally, the processor (910) can perform a data processing process of transmitting one or more embedding vectors, separated by identification information of the unconstrained sensor unit included in the encoding unit, as input to the Generalist AI Unit.
[0179] The generalist AI department may include a pre-trained AI model.
[0180] The AI model included in the generalist AI unit can provide the effect of not requiring the redesign or modification of each healthcare AI model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the existing Decoding Unit when new input data is added or new output results are added during the data processing process of the multimodal AI-based healthcare device according to the present invention.
[0181] The processor (910) can apply all embedding vectors received from the encoding unit to the aforementioned pre-trained generalist artificial intelligence model to generate and transmit one or more integrated embedding vectors to be applied to each healthcare artificial intelligence model (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0182] The above integrated embedding vector may include the contents of one or more embedding vectors separated by identification information of each unconstrained sensor part. The data of the integrated embedding vector may differ for each healthcare artificial intelligence model in the decoding part, and even if an embedding vector is applied to a specific healthcare artificial intelligence model in the decoding part, the weights of the embedding vectors for each unconstrained sensor part within the integrated embedding vector may change as learning progresses.
[0183] The processor (910) can output the pre-determined healthcare items by applying the integrated embedding vector received through the generalist AI unit to one or more healthcare AI models (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) included in the decoding unit.
[0184] A multimodal artificial intelligence-based healthcare device and method according to the present invention preprocesses one or more biosignals of a user and generates an integrated embedding vector that can be applied to one or more healthcare artificial intelligence models through a generalist artificial intelligence model.
[0185] Due to this difference from existing specialized AI models, the multimodal AI-based healthcare device and method according to the present invention provide flexible scalability that eliminates the need to redesign or modify existing healthcare AI models (HELTHCARE AI MODEL #1, HELTHCARE AI MODEL #2, HELTHCARE AI MODEL #3) when adding new input data or new output results in a non-constrained sensing environment composed of various non-constrained sensors.
[0186] Based on this scalability, users' biometric information can be easily augmented through an augmented artificial intelligence model, and based on this, more accurate healthcare results can be output, and furthermore, it can provide utility in building personalized healthcare devices and methods.
[0187] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.
[0188] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computing devices and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0189] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0190] The embodiments described above are combinations of the components and features of the present invention in a specific form. Each component or feature should be considered optional unless otherwise explicitly stated. Each component or feature may be implemented in a form not combined with other components or features. Additionally, it is possible to construct embodiments of the present invention by combining some components and / or features. The order of operations described in the embodiments of the present invention may be changed. Some components or features of one embodiment may be included in another embodiment, or may be replaced with corresponding components or features of another embodiment. It is obvious that embodiments may be constructed by combining claims that do not have an explicit citation relationship in the claims, or that new claims may be included by amendment after filing.
[0191] In the present invention, the processor (610, 910) may be implemented by hardware, firmware, software, or a combination thereof. When implementing an embodiment of the present invention using hardware, the processor (610, 910) may be equipped with ASICs (application specific integrated circuits) or DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), etc., configured to perform the present invention. It may also be implemented as a computer-readable recording medium that records a program for executing a method to prevent user information leakage during user authentication according to the present invention on a computer.
[0192] It is obvious to those skilled in the art that the present invention may be embodied in other specific forms without departing from the essential features of the invention. Accordingly, the foregoing detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention shall be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.
[0193] A multimodal, non-restrictive sensor-based generalist healthcare AI system is industrially available to provide high-accuracy healthcare items using biosignals.
Claims
1. In a multimodal artificial intelligence-based healthcare device, One or more non-restrained sensor units for acquiring the user's biometric information; and One or more tokenized preprocessed data separated by identification information of the unbound sensor unit that acquired the above biometric information are generated by adding identification information of the unbound sensor unit to the corresponding biometric information. One or more embedding vectors are generated for each identification information of the unconstrained sensor unit through the tokenization and embedding processes of the above-mentioned generated tokenization preprocessed data, and One or more of the above-mentioned generated embedding vectors are applied to a pre-trained first artificial intelligence model to generate one or more integrated embedding vectors distinguished by a pre-trained second artificial intelligence model, and A multimodal artificial intelligence-based healthcare device characterized by including a processor that outputs one or more healthcare items by applying one or more of the generated integrated embedding vectors to corresponding prior-trained second artificial intelligence models.
2. In Paragraph 1, The above processor is, Segmenting the above one or more acquired biometric information into the same number of data, and A multimodal artificial intelligence-based healthcare device that adds identification information of an unbound sensor unit that has acquired the biometric information for each of the above segmentations, and generates one or more tokenized preprocessed data separated by identification information of the unbound sensor unit.
3. In a multimodal artificial intelligence-based healthcare device, One or more non-restrained sensor units for acquiring the user's biometric information; and One or more of the above-mentioned acquired biometric information is applied to a pre-trained first artificial intelligence model to generate one or more augmented biometric information for the one or more of the above-mentioned acquired biometric information, and One or more tokenized preprocessed data separated by identification information of the unbound sensor unit that acquired the above biometric information are generated by adding the identification information of the unbound sensor unit to the corresponding generated augmented biometric information, and One or more embedding vectors are generated for each identification information of the unconstrained sensor unit through the tokenization and embedding processes of the above-mentioned generated tokenization preprocessed data, and One or more of the above-mentioned generated embedding vectors are applied to a pre-trained second artificial intelligence model to generate one or more integrated embedding vectors distinguished by a pre-trained third artificial intelligence model, and A multimodal artificial intelligence-based healthcare device characterized by including a processor that outputs one or more healthcare items by applying one or more of the generated integrated embedding vectors to corresponding third artificial intelligence models.
4. In a multimodal artificial intelligence-based healthcare device, A communication unit that receives one or more biometric information of a user and identification information of an unrestrained sensor unit that has acquired the one or more biometric information; and One or more tokenized preprocessed data separated by identification information of the received unrestrained sensor unit are generated by adding the identification information of the received unrestrained sensor unit to the corresponding received biometric information, and One or more embedding vectors are generated for each identification information of the unconstrained sensor unit through the tokenization and embedding processes of the above-mentioned generated tokenization preprocessed data, and One or more of the above-mentioned generated embedding vectors are applied to a pre-trained first artificial intelligence model to generate one or more integrated embedding vectors distinguished by a pre-trained second artificial intelligence model, and A multimodal artificial intelligence-based healthcare device characterized by including a processor that outputs one or more healthcare items by applying one or more of the generated integrated embedding vectors to corresponding prior-trained second artificial intelligence models.
5. In a multimodal artificial intelligence-based healthcare device, A communication unit that receives one or more biometric information of a user and identification information of an unrestrained sensor unit that has acquired the one or more biometric information; and One or more of the received biometric information is applied to a pre-trained first artificial intelligence model to generate one or more augmented biometric information for the one or more of the acquired biometric information, and The identification information of the received unbound sensor unit is added to the corresponding generated augmented biometric information to generate one or more tokenized preprocessed data separated by the identification information of the unbound sensor unit, and One or more embedding vectors are generated for each identification information of the unconstrained sensor unit through the tokenization and embedding processes of the above-mentioned generated tokenization preprocessed data, and One or more of the above-mentioned generated embedding vectors are applied to a pre-trained second artificial intelligence model to generate one or more integrated embedding vectors distinguished by a pre-trained third artificial intelligence model, and A multimodal artificial intelligence-based healthcare device characterized by including a processor that outputs one or more healthcare items by applying one or more of the generated integrated embedding vectors to corresponding third artificial intelligence models.
6. Regarding the method of outputting multimodal AI-based healthcare items, A step of acquiring one or more biometric information of a user through a non-restrained sensor unit; A step of generating one or more tokenized preprocessed data separated by identification information of the unrestrained sensor unit, by adding identification information of the unrestrained sensor unit that acquired the above biometric information to the corresponding biometric information; A step of generating one or more embedding vectors classified by identification information of the unconstrained sensor unit through the tokenization and embedding processes of the tokenized preprocessed data generated above; A step of applying one or more of the generated embedding vectors to a pre-trained first artificial intelligence model to generate one or more integrated embedding vectors distinguished by a pre-trained second artificial intelligence model, and A method for outputting a multimodal AI-based healthcare item, characterized by including the step of applying one or more of the generated integrated embedding vectors to each of the corresponding pre-trained second AI models to output one or more healthcare items.
7. In Paragraph 6, The step of generating the above tokenized preprocessed data is, Segmenting the above one or more acquired biometric information into the same number of data, and A method for outputting multimodal artificial intelligence-based healthcare items, comprising the step of adding identification information of an unbound sensor unit that has acquired the biometric information for each of the above segmentations, and generating one or more tokenized preprocessed data separated by identification information of the unbound sensor unit.
8. Regarding the method of outputting multimodal AI-based healthcare items, A step of acquiring one or more biometric information of a user through a non-restrained sensor unit; A step of applying one or more of the acquired biometric information to a pre-trained first artificial intelligence model to generate one or more augmented biometric information for the one or more of the acquired biometric information; A step of generating one or more tokenized preprocessed data separated by identification information of the unrestrained sensor unit, wherein the identification information of the unrestrained sensor unit that acquired the above biometric information is added to the corresponding generated augmented biometric information; A step of generating one or more embedding vectors classified by identification information of the unconstrained sensor unit through the tokenization and embedding processes of the tokenized preprocessed data generated above; A step of applying one or more of the generated embedding vectors to a pre-trained second artificial intelligence model to generate one or more integrated embedding vectors distinguished by a pre-trained third artificial intelligence model; A method for outputting a multimodal AI-based healthcare item, characterized by including the step of applying one or more of the generated integrated embedding vectors to each of the corresponding pre-trained third AI models to output one or more healthcare items.
9. Regarding the method of outputting multimodal AI-based healthcare items, A step of receiving one or more biometric information of a user and identification information of an unrestrained sensor unit that has acquired the one or more biometric information through a communication unit; A step of generating one or more tokenized preprocessed data separated by identification information of the unrestrained sensor unit by adding the identification information of the received unrestrained sensor unit to the corresponding received biometric information; A step of generating one or more embedding vectors classified by identification information of the unconstrained sensor unit through the tokenization and embedding processes of the tokenized preprocessed data generated above; A step of applying one or more of the generated embedding vectors to a pre-trained first artificial intelligence model to generate one or more integrated embedding vectors distinguished by a pre-trained second artificial intelligence model; A method for outputting a multimodal AI-based healthcare item, characterized by including the step of applying one or more of the generated integrated embedding vectors to each of the corresponding pre-trained second AI models to output one or more healthcare items.
10. A computer-readable recording medium storing a program for executing on a computer a multimodal artificial intelligence-based healthcare item output method described in any one of paragraphs 6 through 9.