Periodic variable disease risk prediction device and method
The integration of genetic, lifestyle, and medical data through a comprehensive prediction system addresses the limitations of existing methods, improving disease risk prediction accuracy and enabling real-time risk assessment.
Patent Information
- Application Number
- PCT/KR2025/095448
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-12
- Filing Date
- 2025-07-04
- Publication Date
- 2026-02-19
AI Technical Summary
Existing disease risk prediction methods based on genetic factors have limitations in accuracy and usability, while methods using medical records fail to account for genetic factors, leading to incomplete disease risk assessments and hindering real-time risk evaluation due to outdated medical information.
A device and method that integrates genetic information, lifestyle habits, and medical examination records to predict periodic variable disease risk by using a data collection module, prediction module, and learning module to calculate comprehensive disease risk through pre-learned models and models trained on patient and control group data.
Enhances the accuracy and commercialization potential of disease risk prediction by accounting for dynamic lifestyle and medical examination data, enabling real-time assessment and personalized risk evaluation.
Smart Images

Figure KR2025095448_19022026_PF_FP_ABST
Abstract
Description
Device and method for predicting periodic variable disease risk
[0001] The present invention relates to a device and method for predicting periodic variable disease risk.
[0002] Specifically, the present invention relates to a disease risk prediction device and method capable of predicting a periodic variable disease risk for a specific disease by predicting a user's medical examination record or disease risk based on information about the user's personal information and lifestyle habits at a predetermined period, and then synthesizing the predicted result and the disease risk related to the user's genes and lifestyle habits.
[0003]
[0004] The content described in this section merely provides background information for the present embodiment and does not constitute prior art.
[0005] The development of disease is comprehensively influenced by genetic and environmental factors (lifestyle factors, medical information, etc.). Except for rare genetic disorders, which are absolutely influenced by specific genetic factors, most diseases are influenced by lifestyle and environmental factors in addition to genetics.
[0006] Recently, methods for predicting disease risk (susceptibility) based on genetic factors have been widely used by genetic testing service companies. These methods predict disease risk by examining genetic mutations associated with disease. However, predicting disease risk based solely on these genetic factors has limitations in terms of prediction accuracy and usability.
[0007] Additionally, some methods for predicting disease risk using medical records and medical information have recently been commercialized. These methods generally demonstrate higher accuracy than disease susceptibility tests based on genetic factors. However, disease risk prediction based on medical records, such as health checkups, has limitations: because they do not account for genetic factors, they fail to account for individual differences in disease incidence due to genetic predisposition.
[0008] Therefore, there is a significant need to increase the accuracy and commercialization potential of disease risk prediction by synthesizing a vast amount of healthcare big data on specific diseases, such as genetic information, lifestyle information on lifestyle habits, and medical examination information on medical records, to predict disease risk.
[0009] Meanwhile, genetic factors are inherently immutable, whereas lifestyle habits are dynamic and subject to change. Medical examination records, which reflect the current state of genetic factors and lifestyle habits, also exhibit this volatility. Lifestyle factors and medical information differ in their volatility. While individuals can immediately recognize or intentionally change lifestyle changes, medical information is typically updated once or twice a year through health checkups. This means that medical information remains fixed for a period of one to two years, depending on the frequency of new health checkups. This limitation arises when using medical information alone or in combination with other factors to predict disease risk, as it prevents the use of up-to-date medical information. This hinders the real-time assessment of disease risk and, consequently, hinders the effectiveness of inducing lifestyle changes (e.g., diet, exercise) that take disease risk into account.
[0010] Meanwhile, the present invention used human resources from the National Central Bank of the National Institute of Health, Korea Centers for Disease Control and Prevention (NBK-2023-061).
[0011]
[0012] The purpose of the present invention is to provide a periodic variable disease risk prediction device and method capable of predicting the disease risk for a specific disease by comprehensively utilizing healthcare big data including information about a user's genes, information about lifestyle habits, and information about medical examination records.
[0013] In addition, the purpose of the present invention is to provide a disease risk prediction device and method capable of predicting a periodic variable disease risk for a specific disease by predicting a user's medical examination record or disease risk based on information about the user's personal information and lifestyle habits at a predetermined period, and then synthesizing the predicted result and the disease risk related to the user's genes and lifestyle habits.
[0014] The objectives of the present invention are not limited to those mentioned above. Other objectives and advantages of the present invention not mentioned above can be understood through the following description and will be more clearly understood through the embodiments of the present invention. Furthermore, it will be readily apparent that the objectives and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims.
[0015]
[0016] According to some embodiments of the present invention, a device for predicting a periodic variable disease risk may include a data collection module that receives genetic information about a user's genes, lifestyle information about lifestyle habits, and basic information about personal information, and a prediction module that predicts a comprehensive disease risk, which is a risk of the user related to a predefined disease, based on the genetic information, the lifestyle information, and the basic information, wherein the prediction module may include a first calculation unit that calculates a genetic risk related to the disease based on the genetic information, a second calculation unit that calculates a lifestyle risk related to the disease based on the lifestyle information, a prediction unit that predicts a checkup risk related to a medical checkup record of the user based on the basic information and the lifestyle information, and an integrated calculation unit that calculates the comprehensive disease risk based on the genetic risk, the lifestyle risk, and the checkup risk.
[0017] In addition, the data collection module can receive the genetic information, the lifestyle information, and the basic information from an external database including a genetic database linked to the comprehensive disease risk prediction device, a user terminal carried by the user, and a medical record database storing the basic information.
[0018] In addition, the data collection module can receive, from the user terminal, the results of a survey related to the user's disease entered into the user terminal as the lifestyle information.
[0019] In addition, the data collection module receives the survey results as the lifestyle information from the user terminal according to a predefined cycle, the prediction unit predicts the examination risk based on each of the lifestyle information received according to the cycle, and the integrated calculation unit can calculate the comprehensive disease risk according to the cycle based on the genetic risk, the lifestyle risk, and the examination risk generated according to the cycle.
[0020] In addition, the first calculation unit calculates the genetic risk using a pre-learned genetic risk calculation model and the genetic information, the second calculation unit calculates the lifestyle risk using a pre-learned lifestyle risk calculation model and the lifestyle information, the prediction unit calculates the screening risk using a pre-learned screening risk prediction model, the basic information, and the lifestyle information, and the integrated calculation unit can calculate the comprehensive disease risk using a pre-learned comprehensive disease risk calculation model, the genetic risk, the lifestyle risk, and the screening risk.
[0021] In addition, the periodic variable disease risk prediction device may further include a learning module that learns the genetic risk calculation model, the lifestyle risk calculation model, the screening risk prediction model, and the comprehensive disease risk calculation model.
[0022] In addition, the learning module, when training the screening risk prediction model, can determine associated information related to associated lifestyle habits that are related to the disease, and train the screening risk prediction model to predict the screening risk based on the basic information and the associated information.
[0023] In addition, the screening risk prediction model may include at least one of an individual model including a first model that predicts screening information regarding the user's medical screening record when the basic information and the lifestyle information are input, a second model that calculates the screening risk based on the predicted screening information, and an integrated model that calculates the user's screening risk when the basic information and the lifestyle information are input.
[0024] In addition, the present invention may further include an output module that outputs the comprehensive disease risk in the form of a probability value or converts the probability value into a rank compared to a comparison group and outputs it.
[0025] A method for predicting a periodic variable disease risk, performed by a periodic variable disease risk predicting device according to some embodiments of the present invention, includes a receiving step of receiving genetic information about a user's genes, lifestyle information about lifestyle habits, and basic information about personal information, and a prediction step of predicting a comprehensive disease risk, which is a risk of the user related to a predefined disease, based on the genetic information, the lifestyle information, and the basic information, wherein the prediction step may include a step of calculating a genetic risk related to the disease based on the genetic information, a step of calculating a lifestyle risk related to the disease based on the lifestyle information, a step of predicting a checkup risk related to a medical checkup record of the user based on the basic information and the lifestyle information, and a step of calculating the comprehensive disease risk based on the genetic risk, the lifestyle risk, and the checkup risk.
[0026]
[0027] The periodic variable disease risk prediction device and method according to some embodiments of the present invention can further improve the accuracy and commercialization potential of disease risk prediction by comprehensively using information about a user's genes, information about lifestyle habits, and information about medical examination records to predict the disease risk for a specific disease. In other words, the periodic variable disease risk prediction device and method according to some embodiments of the present invention can predict a disease risk related to genes using information about a user's genes, a disease risk related to lifestyle habits using information about lifestyle habits, and a disease risk related to medical examination records using information about medical examination records, and then calculate a disease risk by integrating these individual disease risks, thereby having a remarkable effect of more accurately predicting disease risks.
[0028] In addition, since disease risks related to lifestyle habits and disease risks related to medical examination records have characteristics that can fluctuate over time, the periodic variable disease risk prediction device and method according to some embodiments of the present invention can separately predict disease risks related to genes, disease risks related to lifestyle habits, and disease risks related to medical examination records, and then calculate a comprehensive disease risk by integrating them, thereby distinguishing and processing the variability characteristics of each risk.
[0029] At this time, the periodic variable disease risk prediction device and method according to some embodiments of the present invention have a new effect of being able to predict the periodic variable disease risk for a specific disease by predicting the user's medical examination record or the disease risk resulting therefrom based on information about the user's personal information and lifestyle habits at a predetermined cycle and then using the predicted medical examination record. That is, information about lifestyle habits and medical records has the characteristic of changing over time, and in this case, there exists a problem in that it is difficult to reflect the information about medical records in real time compared to information about lifestyle habits. To solve this, the periodic variable disease risk prediction device and method according to some embodiments of the present invention can improve the real-time nature of comprehensive disease risk prediction by predicting the user's medical examination record or the disease risk resulting therefrom based on information about the user's personal information and lifestyle habits at a predetermined cycle (e.g., day, week, month, quarter, year) and then using the predicted medical examination record.
[0030] In addition to the above-described contents, the specific effects of the present invention are described together with the specific matters for carrying out the invention below.
[0031]
[0032] FIG. 1 illustrates a periodic variable disease risk prediction system according to some embodiments of the present invention.
[0033] FIG. 2 is a block diagram of a periodic variable disease risk prediction device according to some embodiments of the present invention.
[0034] FIG. 3 is a diagram illustrating the structure of a neural network model according to some embodiments of the present invention.
[0035] FIG. 4 is a detailed block diagram of a prediction module according to some embodiments of the present invention.
[0036] FIGS. 5A to 5D are diagrams for explaining the learning control process of each model and its learning module according to some embodiments of the present invention.
[0037] FIG. 6 is a flowchart of a method for predicting periodic variable disease risk according to some embodiments of the present invention.
[0038] FIG. 7 is a diagram illustrating a hardware implementation of a periodic variable disease risk prediction device that performs a periodic variable disease risk prediction method according to some embodiments of the present invention.
[0039]
[0040] The terms and words used in this specification and claims should not be interpreted based on their general or dictionary meanings. In accordance with the principle that inventors can define the concepts of terms and words to best describe their inventions, they should be interpreted in a way that is consistent with the technical concept of the present invention. Furthermore, the embodiments described in this specification and the configurations depicted in the drawings are merely examples of how the present invention can be realized and do not fully represent the technical concept of the present invention. Therefore, it should be understood that various equivalents, modifications, and applicable examples may exist as of the time of filing.
[0041] The terms first, second, A, B, etc. used in this specification and claims may be used to describe various components, but the components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term "and / or" includes any combination of a plurality of related listed items or any item among a plurality of related listed items.
[0042] The terminology used in this specification and claims is for the purpose of describing specific embodiments only and is not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. It should be understood that terms such as "comprise" or "have" in this application do not preclude the presence or addition of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification.
[0043] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.
[0044] Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined in this application.
[0045] In addition, each configuration, process, procedure or method included in each embodiment of the present invention may be shared within a scope that is not technically inconsistent with each other.
[0046] Hereinafter, with reference to FIGS. 1 to 7, a device and method for predicting periodic variable disease risk according to some embodiments of the present invention and a periodic variable disease risk prediction system including the same will be described.
[0047]
[0048] FIG. 1 illustrates a periodic variable disease risk prediction system according to some embodiments of the present invention.
[0049] Referring to FIG. 1, a periodic variable disease risk prediction system (1) may include an external database (100), a periodic variable disease risk prediction device (200), and a communication network (300).
[0050] The external database (100) is a device that transmits input data regarding the user's disease risk to a periodic variable disease risk prediction device (200).
[0051] As some examples, the external database (100) may include a genetic database (101), a user terminal (102), a medical record database (103), etc. However, the embodiments of the present invention are not limited thereto, and it is obvious that the external database (100) may include more types of objects.
[0052] The genetic database (101) can store genetic information about the user and transmit such genetic information to the periodic variable disease risk prediction device (200). At this time, the genetic information may include information related to genes in the user's body. For example, the genetic information may include DNA sequencing (Deoxyribo Nucleic Acid Sequencing), DNA chip (Deoxyribo Nucleic Acid Chip), PCR (Polymerase Chain Reaction) results targeting the user's skin, blood, etc., but the embodiments of the present invention are not limited thereto. At this time, when the genetic database (101) receives a control signal regarding genetic information transmission of the periodic variable disease risk prediction device (200), it collects genetic information according to the signal and transmits it to the comprehensive disease risk prediction device (200), or when it already has genetic information about the user, it can transmit the pre-stored genetic information to the periodic variable disease risk prediction device (200) without a process of collecting genetic information. Meanwhile, this genetic database (101) may be in the form of a workstation, a data center, an internet data center (IDC), a direct attached storage (DAS) system, a storage area network (SAN) system, a network attached storage (NAS) system, and a redundant array of inexpensive disks (RAID) system, but the embodiments of the present invention are not limited thereto.
[0053] The user terminal (102) can store lifestyle information about the user and transmit such lifestyle information to the periodic variable disease risk prediction device (200). At this time, the lifestyle information may include information about the user's daily lifestyle habits. For example, the lifestyle information may include the user's smoking information, drinking information, exercise information, food intake information, body mass index, sleep time, stress index, etc., but the embodiments of the present invention are not limited thereto. For example, the user terminal (102) can output a survey pop-up, etc., which can obtain the user's lifestyle information, on the screen, and receive the user's response thereto as a survey result, and determine the received survey result as the lifestyle information. As another example, the user terminal (102) can include a sensor (e.g., a sensor measuring a biosignal mounted on a smartphone or a wearable device, etc.) that can obtain the user's lifestyle information, sense the user's lifestyle habits, etc. using the sensor, and then determine the sensing result as the lifestyle information. Meanwhile, the user terminal (102) may be in the form of various types of electronic devices such as a smartphone, a computer, a laptop PC, a wearable device, an IoT device, etc., but the embodiments of the present invention are not limited thereto.
[0054] The medical record database (103) can store basic information about the user and transmit such basic information to the periodic variable disease risk prediction device (200). At this time, the basic information can include information about the user's personal information or basic facts related to a medical examination. For example, the basic information can include the user's age, height, weight, family history, etc., but the embodiment of the present invention is not limited thereto. Meanwhile, the medical record database (103) can be in the form of a workstation, a data center, an internet data center (IDC), a direct attached storage (DAS) system, a storage area network (SAN) system, a network attached storage (NAS) system, and a redundant array of inexpensive disks (RAID) system, but the embodiment of the present invention is not limited thereto.
[0055] However, the embodiment of the present invention is not limited thereto, and in the periodic variable disease risk prediction system (1), the medical record database (103) may be omitted, and at this time, the user's basic information may be transmitted from the genetic database (101) or the user terminal (102) to the periodic variable disease risk prediction device (200).
[0056] The periodic variable disease risk prediction device (200) can calculate a comprehensive disease risk based on input data received from an external database (100) and output the calculated comprehensive disease risk as output data. In this case, the comprehensive disease risk may also be referred to as a periodic variable disease risk. For convenience of explanation, the periodic variable disease risk will be referred to as a comprehensive disease risk and explained below.
[0057] The comprehensive disease risk may be a disease risk prediction result for a specific disease that takes into account all genetic information, lifestyle information, and basic information received from an external database (100). In other words, the periodic variable disease risk prediction device (200) can calculate a comprehensive disease risk by combining, integrating, and / or synthesizing genetic information received from a genetic database (101), lifestyle information received from a user terminal (102), and basic information received from a medical record database (103). This will be described in detail later.
[0058] The communication network (300) refers to a communication means that performs data exchange between an external database (100) and a periodic variable disease risk prediction device (200).
[0059] At this time, the communication network (300) may include a network based on wired Internet technology, wireless Internet technology, and short-range communication technology. The wired Internet technology may include, for example, at least one of a local area network (LAN) and a wide area network (WAN). The wireless Internet technology may include, for example, at least one of wireless LAN (WLAN), Digital Living Network Alliance (DMNA), Wireless Broadband (Wibro), World Interoperability for Microwave Access (Wimax), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), IEEE 802.16, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), Wireless Mobile Broadband Service (WMBS), and 5G NR (New Radio) technology. However, the present embodiment is not limited thereto. Short-range communication technologies may include, for example, at least one of Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra-Wideband (UWB), ZigBee, Near Field Communication (NFC), Ultra Sound Communication (USC), Visible Light Communication (VLC), Wi-Fi, Wi-Fi Direct, and 5G NR (New Radio).However, this embodiment is not limited thereto.
[0060] Hereinafter, the structure and operation of a periodic variable disease risk prediction device (200) according to some embodiments of the present invention will be described in more detail with reference to FIGS. 2 to 7.
[0061]
[0062] FIG. 2 is a block diagram of a periodic variable disease risk prediction device according to some embodiments of the present invention.
[0063] Referring to FIGS. 1 and 2, the periodic variable disease risk prediction device (200) may include a data collection module (210), a prediction module (220), a learning module (230), and an output module (240).
[0064] The data collection module (210) can receive input data from an external database (100). At this time, the input data can include genetic information (Genetic Data, hereinafter referred to as “GD”), lifestyle information (Lifestyle Data, hereinafter referred to as “LD”), and basic information (Basic Data, hereinafter referred to as “BD”). In other words, the data collection module (210) can receive genetic information (GD) from a genetic database (101), lifestyle information (LD) from a user terminal (102), and basic information (BD) from a medical record database (103). However, the embodiment of the present invention is not limited thereto, and the medical record database (103) can be omitted in the periodic variable disease risk prediction system (1), and at this time, the data collection module (210) can also receive basic information (BD) from the genetic database (101) or the user terminal (102).
[0065] Genetic information (GD) may include information related to genes within the user's body. For example, genetic information (GD) may include DNA sequencing (Deoxyribo Nucleic Acid Sequencing), DNA chip (Deoxyribo Nucleic Acid Chip), PCR (Polymerase Chain Reaction) results, etc. targeting the user's skin, blood, etc., but embodiments of the present invention are not limited thereto.
[0066] Lifestyle information (LD) may include information about the user's daily living habits. For example, the lifestyle information (LD) may include the user's smoking information, drinking information, exercise information, food intake information, body mass index, sleep time, stress index, etc., but the embodiments of the present invention are not limited thereto. Food intake information may include information about the category of food consumed (e.g., meat, fish, processed food, fried food, caffeine), intake amount, salinity, etc., but it is to be understood that the embodiments of the present invention are not limited thereto. For example, the lifestyle information (LD) may include data in the form of survey results or text data obtained by post-processing the survey results. As another example, the lifestyle information (LD) may include the results of automatically sensing the user's lifestyle habits, etc. using a sensor included in the user terminal (102) as described above in FIG. 1 (e.g., a sensor measuring a biosignal mounted on a smartphone or a wearable device).
[0067] Meanwhile, the data collection module (210) can periodically receive such life information (LD) from the user terminal (102) according to a predefined cycle (e.g., daily, weekly, monthly, quarterly, semi-annually, annually, etc.).
[0068] For example, the data collection module (210) can output a survey pop-up, etc. to the user through the user terminal (102) at a predetermined period, and can receive lifestyle information (LD) in the form of survey results by receiving the user's response thereto.
[0069] As another example, the data collection module (210) can automatically receive sensing data sensed by a sensor mounted on a user terminal (102) at a predetermined cycle as lifestyle information (LD). At this time, the sensor mounted on the user terminal (102) can be preset to generate sensing data at a predefined cycle or can generate the corresponding sensing data under the control of the data collection module (210).
[0070] At this time, the prediction module (210) can periodically output an integrated disease risk score (hereinafter referred to as “IRS”), as described below, according to the cycle in which the data collection module (210) receives lifestyle information (LD).
[0071] Basic information (BD) may include information about the user's personal details or basic facts related to a medical examination. For example, basic information (BD) may include the user's age, height, weight, family history, etc. However, embodiments of the present invention are not limited thereto, and basic information (BD) may further include information regarding the user's medical examination records. For example, basic information (BD) may further include hospital visit records, prescription records, examination records, health examination records, etc., but embodiments of the present invention are not limited thereto.
[0072] The data collection module (210) can transmit genetic information (GD), lifestyle information (LD), and basic information (BD) to other components within the periodic variable disease risk prediction device (200). For example, the data collection module (210) can transmit genetic information (GD), lifestyle information (LD), and basic information (BD) to the prediction module (220) and the learning module (230), but the embodiments of the present invention are not limited thereto.
[0073] The prediction module (220) can calculate an integrated disease risk score (IRS) based on genetic information (GD), lifestyle information (LD), and basic information (BD). The IRS may be a predicted risk score for a given user related to a predefined disease (e.g., kidney disease, lung disease, etc.).
[0074] As some examples, the prediction module (220) can calculate the comprehensive disease risk (IRS) by inputting genetic information (GD), lifestyle information (LD), and basic information (BD) into at least one risk model used to calculate the comprehensive disease risk (IRS). At this time, the risk model may include a genetic risk calculation model, a lifestyle risk calculation model, a screening risk prediction model, and a comprehensive disease risk calculation model, as described below, but embodiments of the present invention are not limited thereto.
[0075] These risk models can be trained through various methods. For example, the risk model may be a pre-known risk prediction algorithm, or a model trained through statistical analysis (e.g., association analysis, Cox regression, Mendelian randomization), artificial intelligence, machine learning, or deep learning.
[0076] In this case, if the risk model is trained using deep learning, the risk model may include a pre-trained neural network structure. To explain in more detail, deep learning, a type of machine learning technology, learns at multiple levels based on data, going down to a deeper level. In other words, deep learning refers to a set of machine learning algorithms that extract core data from multiple data sets by increasing the level.
[0077] As examples, neural networks can utilize various well-known deep learning architectures. For example, neural networks can utilize structures such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep belief networks (DBNs), graph neural networks (GNNs), generative adversarial networks (GANs), transformers, and autoencoders.
[0078] Specifically, a Convolutional Neural Network (CNN) is a model that mimics the function of the human brain, based on the assumption that when recognizing an object, humans extract its basic features, then perform complex computations in the brain to recognize the object based on the results. CNNs can include, but are not limited to, well-known structures such as LeNet, AlexNet, VGGNet, GoogleNet, and ResNet.
[0079] RNN (Recurrent Neural Network) is widely used in natural language processing, etc., and is an effective structure for processing time-series data that changes over time. It can construct an artificial neural network structure by stacking layers at each moment.
[0080] A DBN (Deep Belief Network) is a deep learning structure constructed by stacking multiple layers of Restricted Boltzman Machines (RBMs), a deep learning technique. By repeatedly training RBMs (Restricted Boltzman Machines), a certain number of layers can be created, creating a DBN (Deep Belief Network) with that number of layers.
[0081] GNN (Graphic Neural Network, hereinafter referred to as GNN) represents an artificial neural network structure implemented in a way that derives similarity and feature points between modeling data by using modeling data modeled based on data mapped between specific parameters.
[0082] A Generative Adversarial Network (GAN) is an artificial neural network structure that uses a generative neural network and a discriminative neural network to generate new data in a similar form to the input data. GANs may include the well-known DCGAN (Deep Convolutional GAN), CGAN (Conditional GAN), WGAN (Wasserstein GAN), StyleGAN (Style-Based GAN), CycleGAN, etc., but embodiments of the present invention are not limited thereto.
[0083] Transformer is an artificial neural network with an attention-based encoder-decoder structure that can understand the overall meaning between input and output sequences. Transformer uses the attention mechanism to ensure that all elements of the input sequence influence the output sequence, allowing both the encoder and decoder to consider the entire sequence. Transformer can use natural language, time-series data, and even patched images as input.
[0084] An autoencoder is a deep learning architecture that extracts and reconstructs data features. Typically, an autoencoder comprises an encoder, which compresses input values, and a decoder, which restores the compressed data. The encoder transforms the input values into a low-dimensional latent representation, and the decoder reconstructs the latent representation to the same dimensionality as the input values. Each encoder and decoder can be configured as a multilayer perceptron (MLP). When training an autoencoder, input data is input, and weights and biases are trained to minimize the difference between the output and the input values. This trained autoencoder can effectively extract input data features and reconstruct noisy input data. Autoencoders are primarily used in fields such as data compression, dimensionality reduction, noise removal, and data generation, and can also be utilized in areas such as image recognition, natural language processing, and speech recognition.
[0085] Meanwhile, artificial neural network learning can be achieved by adjusting the weights of connections between nodes (and, if necessary, bias values) to ensure the desired output for a given input. Furthermore, artificial neural networks can continuously update their weight values through learning. Furthermore, methods such as backpropagation can be used for artificial neural network learning.
[0086] At this time, machine learning methods for artificial neural networks can include unsupervised learning, semi-supervised learning, and supervised learning. Furthermore, the neural network can be controlled to automatically update its structure to output post-learning analysis data, depending on settings.
[0087] Hereinafter, a neural network structure according to some embodiments of the present invention will be described with reference to FIG. 3.
[0088]
[0089] FIG. 3 is a diagram illustrating the structure of a neural network model according to some embodiments of the present invention.
[0090] Referring to FIG. 3, a neural network (hereinafter referred to as “NN”) according to some embodiments of the present invention may include an input layer, an output layer, and M hidden layers positioned between the input layer and the output layer.
[0091] Here, weights can be assigned to the edges connecting the nodes of each layer. These weights or edges can be added, removed, or updated during the learning process. Therefore, the weights of the nodes and edges between the k input nodes and i output nodes can be updated during the learning process.
[0092] Before a neural network (NN) begins learning, all nodes and edges can be set to initial values. However, as information accumulates, the weights of nodes and edges change, and this process can create a match between the parameters input as learning factors and the values assigned to output nodes.
[0093] Additionally, when using a cloud server, neural networks (NNs) can receive and process a large number of parameters. Therefore, neural networks (NNs) can learn based on massive amounts of data.
[0094] The weights of the nodes and edges between the input and output nodes that make up a neural network (NN) can be updated through the NN's learning process. Furthermore, the parameters input or output from a neural network (NN) can be further expanded with various data.
[0095]
[0096] Referring again to FIGS. 1 and 2, the detailed process of calculating the comprehensive disease risk (IRS) by inputting genetic information (GD), lifestyle information (LD) and basic information (BD) into at least one risk model by the prediction module (220) will be described later.
[0097] Meanwhile, the comprehensive disease risk score (IRS) output by the prediction module (220) may be in the form of a probability value. In other words, the prediction module (220) may output the probability value that the user will have the disease as the comprehensive disease risk score (IRS). In this case, the prediction module (220) may provide the comprehensive disease risk score (IRS) to an output module (240), etc.
[0098] The learning module (230) can control the learning of at least one risk model used by the prediction module (220). In other words, the learning module (230) can use predefined learning data to perform and control the learning process for the genetic risk calculation model, lifestyle risk calculation model, screening risk prediction model, and comprehensive disease risk calculation model used by the prediction module (220).
[0099] At this time, the training data may include patient group data and control group data. Patient group data may refer to data on individuals with a specific disease, while control group data may refer to data on individuals without the disease.
[0100] The specific process by which the learning module (230) controls learning for the risk model will be described later.
[0101] The output module (240) can output output data (hereinafter referred to as “OD”) based on the comprehensive disease risk score (IRS).
[0102] At this time, as described above, the comprehensive disease risk (IRS) provided by the output module (240) from the prediction module (220) may be in the form of a probability value.
[0103] For example, the output module (240) can output the comprehensive disease risk score (IRS) in the form of a probability value as output data (OD). In other words, the output module (240) can provide the absolute score, which is the comprehensive disease risk score (IRS) itself in the form of a probability value provided from the prediction module (220), as output data (OD) to the user terminal (102), etc.
[0104] As another example, the output module (240) may post-process the comprehensive disease risk (IRS) in the form of a probability value to convert it into a ranking compared to a comparison group, and output the converted data as output data (OD). In other words, the output module (240) may compare the comprehensive disease risk (IRS) in the form of a probability value provided from the prediction module (220) with similar users in the comparison group, thereby calculating the risk ranking in the comparison group, and may provide the calculated ranking or the relative score based on the comparison with the comparison group as output data (OD) to the user terminal (102), etc.
[0105] Hereinafter, the operation of the prediction module (220) and the learning module (230) according to some embodiments of the present invention will be described in more detail with reference to FIGS. 4 to 5d.
[0106]
[0107] FIG. 4 is a detailed block diagram of a prediction module according to some embodiments of the present invention. FIGS. 5a to 5d are diagrams illustrating the learning control process of each model and its learning module according to some embodiments of the present invention.
[0108] Referring to FIGS. 2 and 4, a prediction module (220) according to some embodiments of the present invention may include a first calculation unit (221), a second calculation unit (222), a prediction unit (223), and an integrated calculation unit (224).
[0109] The first generating unit (221) can calculate a genetic risk score (hereinafter referred to as "GS") by inputting genetic information (GD) into a genetic risk score calculating model (Generic Risk Score Calculating Model, hereinafter referred to as "M_GS"). At this time, the genetic risk score (GS) may include the probability that a user will genetically have a specific disease defined in advance. In other words, the genetic risk score (GS) may include a result of predicting the probability that a user will have a specific disease (incidence prediction result) based on the genetic information (GD).
[0110] This genetic risk calculation model (M_GS) can be trained by a learning module (230) to output a genetic risk (GS) when genetic information (GD) is input.
[0111] More specifically, referring to FIGS. 2, 4, and 5A, the learning module (230) can determine a genetic biomarker (hereinafter referred to as “GB”) related to a specific predefined disease, and train the genetic risk calculation model (M_GS) to output a genetic risk (GS) based on the genetic information (GD) and the genetic biomarker (GB). At this time, the genetic biomarker (GB) may include information about a biomarker in the human body that has a correlation with the specific disease. For example, the genetic biomarker (GB) may include information about a genetic variant or polymorphism (single nucleotide polymorphism, SNP) related to the specific disease, but the embodiments of the present invention are not limited thereto. For example, the learning module (230) can determine the degree of influence of the genetic mutation of gene A on the disease by comparing the rate at which the genetic mutation of gene A occurs in patient group data with the rate at which the genetic mutation of gene A occurs in control group data, and can determine which associated genes have a correlation higher than a predefined threshold with the disease through this method, and then determine the determined associated genes as genetic biomarkers (GB).
[0112] Specifically, first, the learning module (230) can determine a genetic biomarker (GB) associated with at least one disease. For example, the learning module (230) can receive information about what biomarkers are known to be associated with each of a plurality of diseases in advance (e.g., the high probability of occurrence of a specific disease depending on the type and presence of gene A) from an external source (e.g., a medical-related academic DB), and determine the received information as a genetic biomarker (GB) for each disease. As another example, the learning module (230) can derive a genetic biomarker (GB) associated with each of a plurality of diseases by performing a genome-wide association study and / or a case control study on learning data including patient group data and control group data prepared in advance.
[0113] Next, the learning module (230) can train the genetic risk calculation model (M_GS) to output the genetic risk (GS) for the specific user when the genetic biomarker (GB) and the genetic information (GD) of the specific user are input. At this time, as described above, the learning module (230) can train the genetic risk calculation model (M_GS) through statistical analysis (e.g., association analysis, Cox regression analysis, Mendelian randomization), artificial intelligence, machine learning, deep learning, etc. For example, the learning module (230) can train the genetic risk calculation model (M_GS) to compare the genetic information (GD) and the genetic biomarker (GB) of the specific user and then output the genetic risk (GS) as the probability that the user will have the disease based on the comparison result. For example, the genetic risk calculation model (M_GS) can be trained by inputting genetic information (GD) of patient group data and control group data, determining inclusion information (e.g., whether included, type of inclusion, etc.) for genetic biomarkers (GB) of each genetic information (GD), and then comparing the inclusion information in each of the patient group data and the control group data to output a genetic risk (GS) corresponding to the genetic information (GD). At this time, the genetic risk (GS) may be data in the form of a probability value. However, the embodiments of the present invention are not limited thereto.
[0114] Referring back to FIGS. 2 and 4, the second calculation unit (222) can calculate a Lifestyle Risk Score (hereinafter referred to as “LS”) by inputting lifestyle information (LD) into a Lifestyle Risk Score Calculating Model (hereinafter referred to as “M_LS”). At this time, the Lifestyle Risk Score (LS) may include the probability of the user having a specific disease defined in advance based on the user’s lifestyle habits. In other words, the Lifestyle Risk Score (LS) may include a result of predicting the probability of the user having a specific disease (incidence rate prediction result) based on the lifestyle information (LD).
[0115] The life risk calculation model (M_LS) can be trained by a learning module (230) to output a life risk (LS) when life information (LD) is input.
[0116] More specifically, referring to FIGS. 2, 4, and 5b, the learning module (230) can determine a related index of lifestyle (hereinafter referred to as “RL”) related to a predefined specific disease, and train the lifestyle risk calculation model (M_LS) to output a lifestyle risk (LS) based on the lifestyle information (LD) and the lifestyle risk correlation (RL). At this time, the lifestyle risk correlation (RL) may include information related to diet, sleep, activity, etc., which are presumed to have a correlation with the specific disease. For example, the lifestyle risk correlation (RL) may include a correlation analysis result between a group, range, or value in each item included in the lifestyle information (LD) and the lifestyle risk (LS), but the embodiments of the present invention are not limited thereto. That is, as described above, lifestyle information (LD) may include the user's smoking information, drinking information, exercise information, food intake information, body mass index, sleep time, and stress index, and lifestyle correlation (RL) may include the results of analyzing the correlation between each of these multiple items and lifestyle risk (LS).
[0117] Specifically, first, the learning module (230) can determine a lifestyle association (RL) related to at least one disease. For example, the learning module (230) can receive information from an external source (e.g., a medical academic DB) regarding lifestyle habits known in advance to be related to each of a plurality of diseases (e.g., a high risk of lung disease when smoking), and determine the received information as a lifestyle association (RL) for each disease. As another example, the learning module (230) can derive a lifestyle association (RL) related to each of a plurality of diseases by performing a genome-wide association study and / or a case-control study on learning data including pre-prepared patient group data and control group data.
[0118] Next, the learning module (230) can train the lifestyle risk calculation model (M_LS) to output the lifestyle risk (LS) for a specific user when the lifestyle association (RL) and the lifestyle information (LD) of the specific user are input. At this time, as described above, the learning module (230) can train the lifestyle risk calculation model (M_LS) through statistical analysis (e.g., association analysis, Cox regression analysis, Mendelian randomization), artificial intelligence, machine learning, deep learning, etc. For example, the learning module (230) can train the lifestyle risk calculation model (M_LS) to compare the lifestyle information (LD) and the lifestyle association (RL) of a specific user, and then output the lifestyle risk (LS) as the probability that the user will have the disease based on the comparison result. For example, the lifestyle risk calculation model (M_LS) can be trained by inputting lifestyle information (LD) of patient group data and control group data, determining the similarity information between each lifestyle information (LD) and lifestyle habit association (RL), and then comparing the similarity information in each of the patient group data and control group data to output a lifestyle risk (LS) corresponding to the lifestyle information (LD). At this time, the lifestyle risk (LS) may be data in the form of a probability value. However, the embodiments of the present invention are not limited thereto.
[0119] Referring back to FIGS. 2 and 4, the prediction unit (223) can calculate a medical risk score (hereinafter referred to as “MS”) by inputting lifestyle information (LD) and basic information (BD) into a medical risk score forecasting model (hereinafter referred to as “M_MS”). At this time, the medical risk score (MS) may include a probability of having a specific disease defined in advance based on the user’s medical examination record. In other words, the medical risk score (MS) may include a result of predicting the probability of the user having a specific disease based on the medical examination record (incidence prediction result) based on the basic information (BD) and lifestyle information (LD).
[0120] As some examples, the screening risk prediction model (M_MS) can use lifestyle information (LD) and basic information (BD) as input data and output screening risk (MS) as output data.
[0121] At this time, the screening risk prediction model (M_MS) may include an individual model that predicts screening information based on lifestyle information (LD) and basic information (BD) and then outputs screening risk (MS) based on the predicted screening information, and an integrated model that directly predicts screening risk (MS) based on lifestyle information (LD) and basic information (BD).
[0122] In more detail, FIGS. 2, 4 and 5c <a1>Referring to FIG. 5c, the screening risk prediction model (M_MS) may include an individual model (hereinafter referred to as "M_MS_IN") and an integrated model (hereinafter referred to as "M_MS_IT"). <a1>Although the screening risk prediction model (M_MS) is shown to include both the individual model (M_MS_IN) and the integrated model (M_MS_IT), this is only for convenience of explanation, and either the individual model (M_MS_IN) or the integrated model (M_MS_IT) may be omitted and implemented in the screening risk prediction model (M_MS).
[0123] The individual model (M_MS_IN) may include a first model (1st Individual Model, hereinafter referred to as "M_MS_IN1") and a second model (2nd Individual Model, hereinafter referred to as "M_MS_IN2"). When basic information (BD) and lifestyle information (LD) are input, the first model (M_MS_IN1) may predict and output medical examination information (Medical Data, hereinafter referred to as "MD") regarding the medical examination record of the corresponding user. The second model (M_MS_IN2) may receive medical examination information (MD) and output a medical examination risk (MS) corresponding to the medical examination information (MD).
[0124] The integrated model (M_MS_IT) can directly output the screening risk (MS) for a user when basic information (BD) and lifestyle information (LD) are entered.
[0125] At this time, the screening risk prediction model (M_MS) can be trained by the learning module (230) to output the screening risk (MS) when lifestyle information (LD) and basic information (BD) are input. In other words, the learning module (230) can train the first model (M_MS_IN1) included in the individual model (M_MS_IN) to predict and output screening information (MD) regarding the medical screening record of the corresponding user when basic information (BD) and lifestyle information (LD) are input, can train the second model (M_MS_IN2) included in the individual model (M_MS_IN) to output the screening risk (MS) corresponding to the screening information (MD) through the screening information (MD), and can train the integrated model (M_MS_IT) to directly output the screening risk (MS) for the corresponding user based on the basic information (BD) and lifestyle information (LD).
[0126] In more detail, the learning module (230) can train the screening risk prediction model (M_MS) to output a screening risk (MS) when basic information (BD) and lifestyle information (LD) are input. At this time, the learning module (230) can train the screening risk prediction model (M_MS) through statistical analysis (e.g., association analysis, Cox regression analysis, Mendelian randomization), artificial intelligence, machine learning, deep learning, etc. For example, the learning module (230) can train the screening risk prediction model (M_MS) to output the probability that the user will have the disease as a screening risk (MS) based on the basic information (BD) and lifestyle information (LD). For example, the screening risk prediction model (M_MS) can be trained by comparing the basic information (BD) and lifestyle information (LD) of the patient group data and the control group data, respectively, when the basic information (BD) and lifestyle information (LD) of the patient group data and the control group data are input, thereby outputting the screening risk (MS) corresponding to the basic information (BD) and lifestyle information (LD).
[0127] For example, the first model (M_MS_IN1) can be trained to predict and output screening information (MD) regarding the medical examination records of each user by comparing the basic information (BD) and lifestyle information (LD) of the patient group data and the control group data, respectively, and the actual medical examination records of the patient group data and the control group data, when the basic information (BD) and lifestyle information (LD) of the patient group data and the control group data are input. In addition, the second model (M_MS_IN2) can be trained to predict and output the screening risk (MS) of each user by comparing the screening information (MD) of the patient group data and the control group data, respectively, when the screening information (MD) of the patient group data and the control group data are input. At this time, the first model (M_MS_IN1) and the second model (M_MS_IN2) can be trained to compare and analyze the patient group data and the control group data through the aforementioned statistical analysis, artificial intelligence, machine learning, deep learning, etc.
[0128] As another example, the integrated model (M_MS_IT) can be trained to directly predict and output the screening risk (MS) of each user by comparing the basic information (BD) and lifestyle information (LD) of the patient group data and the control group data, respectively, when the basic information (BD) and lifestyle information (LD) of the patient group data and the control group data are input. At this time, the integrated model (M_MS_IT) can be trained to compare and analyze the patient group data and the control group data through the aforementioned statistical analysis, artificial intelligence, machine learning, deep learning, and other methods.
[0129] Meanwhile, in Figs. 2, 4 and 5c <a2>Referring to , the learning module (230) can determine related information (Related Data, hereinafter referred to as "RD") related to related lifestyle habits that are associated with a specific disease, and train the screening risk prediction model (M_MS) to output a screening risk (MS) based on the basic information (BD), lifestyle information (LD), and related information (RD). At this time, the related information (RD) can include information on the type and range of lifestyle habits (related lifestyle habits) that are associated with a specific disease.
[0130] Specifically, first, the learning module (230) can determine associated information (RD) related to at least one specific disease. For example, the learning module (230) can receive information about associated information (RD) known in advance to be related to each disease (e.g., irregular sleep increases the risk of developing diabetes) from an external source (e.g., a medical academic database) and determine the received information as associated information (RD) for each disease.
[0131] Next, the learning module (230) can train the screening risk calculation model (M_MS) to output the screening risk (MS) for a specific user when the related information (RD) and the basic information (BD) and lifestyle information (LD) of a specific user are input. At this time, the learning module (230) can train the screening risk calculation model (M_MS) in the same manner as described above, but can train the screening risk calculation model (M_MS) so that the screening risk prediction model outputs the screening risk (MS) using only the related information (RD) among the lifestyle information (LD) together with the basic information (BD).
[0132] Referring back to FIGS. 2 and 4, on the one hand, such lifestyle information (LD) can be periodically and regularly received by the data collection module (210), and at this time, the prediction unit (223) can update or renew the screening risk (MS) based on each lifestyle information (LD) received periodically. In other words, when lifestyle information (LD) in the form of a survey result is received by the data collection module (210) according to a predefined cycle (e.g., daily, weekly, monthly, quarterly, semi-annually, yearly, etc.), the prediction unit (223) can periodically output the screening risk (MS) based on each of the plurality of lifestyle information (LD) according to the cycle.
[0133] The integrated calculation unit (224) can calculate the integrated disease risk (IRS) by inputting the genetic risk (GS), lifestyle risk (LS), and screening risk (MS) into the integrated disease risk calculation model (Integrated Risk Score Calculating Model, hereinafter referred to as "M_IRS"). Accordingly, the integrated disease risk (IRS) can include the disease risk prediction results for a specific disease in which genetic information (GD), lifestyle information (LD), and basic information (BD) are all taken into account. At this time, the integrated disease risk (IRS) output by the integrated calculation unit (224) can be in the form of a probability value.
[0134] The comprehensive disease risk calculation model (M_IRS) can be trained by a learning module (230) to output a comprehensive disease risk (IRS) when genetic risk (GS), lifestyle risk (LS), and screening risk (MS) are input.
[0135] For example, referring to FIGS. 2, 4, and 5d, the learning module (230) can train the comprehensive disease risk calculation model (M_IRS) to output the comprehensive disease risk (IRS) when the genetic risk (GS), lifestyle risk (LS), and screening risk (MS) are input. At this time, as described above, the learning module (230) can train the comprehensive disease risk calculation model (M_IRS) through statistical analysis (e.g., association analysis, Cox regression analysis, Mendelian randomization), artificial intelligence, machine learning, deep learning, etc. For example, the learning module (230) can train the comprehensive disease risk calculation model (M_IRS) to output the probability that the user will have the disease as the comprehensive disease risk (IRS) by combining the genetic risk (GS), the lifestyle risk (LS), and the screening risk (MS). For example, the comprehensive disease risk calculation model (M_IRS) can be trained in a manner that, when the genetic risk (GS), the lifestyle risk (LS), and the screening risk (MS) of the patient group data and the control group data are input, the comprehensive disease risk calculation model (M_IRS) outputs the comprehensive disease risk (IRS) by comparing the genetic risk (GS), the lifestyle risk (LS), and the screening risk (MS) of the patient group data and the control group data, respectively, with the actual onset. At this time, the comprehensive disease risk (IRS) may be data in the form of a probability value. However, the embodiments of the present invention are not limited thereto.
[0136] As another example, the learning module (230) may determine additional parameters (hereinafter referred to as “AP”) related to a predefined specific disease, and train the comprehensive disease risk calculation model (M_IRS) to output a comprehensive disease risk (IRS) based on the genetic risk (GS), the lifestyle risk (LS), the screening risk (MS), and the additional parameters (AP). The additional parameters (AP) may include information related to human body conditions, diet, sleep, activity, etc., which are predefined to be highly correlated with a specific disease (e.g., age is a highly correlated parameter in the case of Parkinson’s disease).
[0137] More specifically, first, the learning module (230) can receive information from an external source (e.g., a medical-related academic DB) about what additional parameters (AP) are known in advance to have a high correlation with each disease for each of a plurality of diseases, and determine the received information as the additional parameters (AP) for each disease.
[0138] Next, the learning module (230) can train the comprehensive disease risk calculation model (M_IRS) to output the comprehensive disease risk (IRS) when the genetic risk (GS), the lifestyle risk (LS), the screening risk (MS), and the additional parameter (AP) are input. At this time, as described above, the learning module (230) can train the comprehensive disease risk calculation model (M_IRS) through statistical analysis (e.g., association analysis, Cox regression analysis, Mendelian randomization), artificial intelligence, machine learning, deep learning, etc. For example, the learning module (230) can train the comprehensive disease risk calculation model (M_IRS) to output the probability that the user will have the disease as the comprehensive disease risk (IRS) by combining the genetic risk (GS), the lifestyle risk (LS), the screening risk (MS), and the additional parameter (AP). For example, the comprehensive disease risk calculation model (M_IRS) can be trained in a way that, when the genetic risk (GS), lifestyle risk (LS), screening risk (MS) and additional parameter (AP) of patient group data and control group data are input, the genetic risk (GS), lifestyle risk (LS), screening risk (MS) and additional parameter (AP) of the patient group data and the control group data are compared with the actual onset of the disease, thereby outputting the comprehensive disease risk (IRS). At this time, the comprehensive disease risk (IRS) may be data in the form of a probability value. However, the embodiments of the present invention are not limited thereto.
[0139] Meanwhile, the integrated calculation unit (224) can periodically output the comprehensive disease risk (IRS) according to the renewal and update of the lifestyle information (LD). That is, as described above, the lifestyle information (LD) can be periodically and regularly received, renewed and updated by the data collection module (210), and the prediction unit (223) can output the screening risk (MS) based on each periodically received lifestyle information (LD), and at this time, the integrated calculation unit (224) can newly output the comprehensive disease risk (IRS) based on the updated and renewed screening risk (MS).
[0140]
[0141] Referring again to FIGS. 2, 4 and 5c, the following describes experimental data predicting the screening risk (MS) based on basic information (BD) and lifestyle information (LD) through the prediction unit (223) of the periodic variable disease risk prediction device (200) of the present invention.
[0142] The experimental data includes the results of experiments on the types of diseases specified as age-related cataract (hereinafter referred to as “cataract”), hypertension, osteoporosis, stroke, and type 2 diabetes (hereinafter referred to as “diabetes”), and the experimental groups were divided into male and female groups.
[0143] At this time, the experimental data includes the experimental results for each of the individual models (M_MS_IN) and the integrated model (M_MS_IT) included in the screening risk prediction model (M_MS).
[0144] First, the composition of the experimental group in the experimental data for the individual model (M_MS_IN) was as follows: for male cataract, 743 patients were tested and 7,344 controls were tested; for female cataract, 1,079 patients were tested and 10,521 controls were tested; for male hypertension, 3,673 patients were tested and 15,429 controls were tested; for female hypertension, 4,399 patients were tested and 30,803 controls were tested; for male osteoporosis, 139 patients were tested and 1,399 controls were tested; for female osteoporosis, 2,599 patients were tested and 25,737 controls were tested; for male stroke, 389 patients were tested and 3,838 controls were tested; for female stroke, 353 patients were tested and 3,436 controls were tested; for male diabetes, 1,524 patients were tested, The experiment was conducted with 15,126 control subjects, and in the case of female diabetes, the experiment was conducted with 1,499 patients and 14,777 control subjects.
[0145] For the first model (M_MS_IN1) of the individual model (M_MS_IN), age, height, weight, and family history (stroke, heart disease, diabetes, hypertension, and cancer in family and relatives) were used as basic information (BD), and smoking level, drinking level, body mass index, meat intake, processed food intake, fried food intake, salt level in eating habits, exercise level, caffeine intake, sleep time, stress level, depression level, etc. were used as lifestyle information (LD). At this time, the first model (M_MS_IN1) of the individual model (M_MS_IN) used in the experiment was trained through the non-negative multiple linear regression method.
[0146] The correlation results (R) between the prediction results of the screening information (MD) by the first model (M_MS_IN1) and the actual screening information values 2 ) are shown in and below. shows the experimental results for men, and shows the experimental results for women.
[0147] MaleASTALTHBgGTPWAISTCreatinineTCHLTGHDLPRT16_UDBPSBPFastingGlucoseStress0.0000.4390.0002.6300.0000.0000.5591.9730.0000.0070.1900.0000.576Smoking0.0000.2530.2019.1290.3440.0000.65024. 2070.0000.0200.0000.0000.084Alcohol1.4360.0000.05314.2680.1860.0001.8698.3843.1500.0051. 2161.6591.880Meat0.0000.0000.0890.2810.0480.0002.9440.0000.5260.0000.0000.0000.000Instant food0.4620.6660.0481.9470.0850.0002.4718.3040.3620.0040.0070.2960.081Salty food0.0000.3060.0000.0000.0000.0000.2275.2660.0000.0000.6470.5320.079Fried food0.4100.9600.0040.0000.0740.0041.4490.0000.0000.0300.2800.2100.000Caf feine0.0000.0000.0210.0000.1090.0032.3440.0000.0000.0040.0000.0000.000Ex ercise0.0000.3000.0620.4670.1360.0000.9123.7310.0000.0020.1060.0000.000B MI1.9044.2210.0308.9582.3260.0001.04616.3020.0000.0241.1962.4141.698Sleep duration0.0040.0010.0000.0100.0000.0000.0000.0020.0000.0000.0000.0050.000Depression0.0000.0000.0000.0 000.0000.0000.0000.0000.0000.0000.0000.0000.000AGE0.0790.0000.0000.0000.1920.0010.0000.0000.0950.0020.0000.2500.245HEIGHT0.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.000Family history of cancer0.0000.0000.0170.0000.0000.0010.5750.0000.2990.0060.0000.0000 .000Family history of stroke0.0000.0000.0060.0000.1690.0000.0000.0000.4270.0010.2480.2440.000Family history of diabetes0.1311.2440.0352.8230.0860.0040.0009.0350.0000.0140.0000.0008.453High blood pressure Family History0.0000.5420.0200.0000.0000.0000.0002.7020.0000.0051.8772.8660.000Family History of Heart Disease0.0000.0000.0000.0000.0000.0082.2720.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000WEIGHT0.000 0.2090.0220.0080.5680.0030.1561.3910.0000.0010.1410.1770.197Intercept14.496 -0.15412.684-21.42130.5280.669159.140-71.08339.0070.74760.17787.94257.336R. 2 0.0090.0440.0540.0640.6480.0110.0210.0730.0410.0050.0610.0680.044
[0148]
[0149] FemaleASTALTHBgGTPWAISTCreatinineTCHLTGHDLPRT16_UDBPSBPFastingGlucoseStress0.0850.0600.0000.2780.1920.0000.0002.1460.0000.0000.0350.0000.996Smoking0.2281.0450.28411.0200.3180.0000.00021 .4290.0000.0300.0000.0002.223Alcohol0.0080.0000.0532.6040.1590.0001.3050.0003.3810.0040. 7650.5910.432Meat0.0000.0000.0330.1490.0000.0001.2190.0000.9970.0000.0000.0000.000Instant food0.0000.0000.0000.1860.0130.0000.5951.4000.6120.0000.2650.4680.000Salty food0.0000.0000.0000.0000.0230.0050.0000.0000.0000.0000.0000.0000.000Fried food0.0000.0000.0030.0000.0000.0000.0000.0000.4710.0000.0000.0000.000Caf feine0.0000.0000.0010.0000.0000.0072.4150.0000.9500.0040.0000.0000.000Ex ercise0.0000.0000.0120.0000.0210.0000.1731.7110.0000.0000.0040.0000.000B MI1.2162.4150.0702.7573.2050.0011.9028.3290.0000.0201.5862.7781.819Sleep duration0.0060.0060.0000.0010.0000.0000.0030.0000.0010.0000.0000.0000.003Depression0.5140.3100.0060.1830 .0000.0030.0001.0310.0000.0000.2000.1340.391AGE0.2060.1810.0180.1970.2650.0020.6591.0660.0000.0000.1890.4720.300HEIGHT0.0000.0000.0000.0000.0000.0020.0000.0000.0770.0000.0000.0000.000Family history of cancer0.7280.5480.0050.2390.0000.0000.7240.0000.2880.0000.0000.0000.00 0Family history of stroke0.0000.0000.0000.0220.1570.0000.0001.8740.0000.0010.4030.5040.000Family history of diabetes0.3950.6380.0200.5820.0000.0000.0004.1120.0000.0240.0000.0005.555Family history of hypertension 0.0130.0360.0170.0080.0000.0050.0001.1590.1380.0102.2353.3950.000Family history of heart disease0.0000.0180.0010.8820.0000.0002.9021.0650.0000.0000.0000.0000.0000.000WEIGHT0.0500 .1980.0080.1870.6260.0000.1131.0700.0000.0010.1740.2370.262Intercept4.802- 7.41111.192-20.88022.8450.295145.375-68.81141.9510.92846.80471.63250.759R. 2 0.0060.0170.0260.0320.6270.0120.0290.0530.0240.0030.0720.1210.057
[0150]
[0151] Next, the second model (M_MS_IN2) calculated the screening risk (MS) based on the predicted screening information (MD). In the case of below, the results of the second model (M_MS_IN2) predicting the screening risk (MS) based on the predicted screening information (MD) from the first model (M_MS_IN1) and the correlation results (R) between the actual screening risk values 2 ) is shown.
[0152] Diseases Male Female Senile cataract 0.24000.5358 Hypertension 0.65090.7292 Osteoporosis 0.42650.7437 Stroke 0.23560.2607 Type 2 diabetes 0.44080.4852 Average R2 0.39880.5509
[0153] Next, the composition of the experimental group in the experimental data for the integrated model (M_MS_IT) was as follows: for male cataract, 795 patients and 7,933 controls were tested; for female cataract, 1,158 patients and 11,610 controls were tested; for male hypertension, 4,002 patients and 16,824 controls were tested; for female hypertension, 4,783 patients and 33,988 controls were tested; for male osteoporosis, 149 patients and 1,499 controls were tested; for female osteoporosis, 2,836 patients and 28,442 controls were tested; for male stroke, 420 patients and 4,196 controls were tested; for female stroke, 375 patients and 3,771 controls were tested; and for male diabetes, 375 patients and 3,771 controls were tested. The experiment was conducted with 1,611 patients and 16,118 controls, and in the case of female diabetes, the experiment was conducted with 1,596 patients and 15,963 controls. In the case of the integrated model (M_MS_IT), the experiment was conducted by using different basic information (BD) and lifestyle information (LD) depending on the type of disease.
[0154] First, in the case of cataracts, age, height, weight, and family history (whether family members or relatives have had stroke, heart disease, diabetes, high blood pressure, or cancer) were used as basic information (BD), and smoking level, drinking level, body mass index, meat intake, processed food intake, salt intake in eating habits, exercise level, and sleep duration were used as lifestyle information (LD).
[0155] Next, for hypertension, age, height, weight, and family history (whether family members or relatives have had stroke, heart disease, diabetes, high blood pressure, or cancer) were used as basic information (BD), and smoking level, drinking level, body mass index, meat intake, processed food intake, fried food intake, salt intake in eating habits, exercise level, stress level, and depression level were used as lifestyle information (LD).
[0156] Next, for osteoporosis, age, height, weight, and family history (stroke, heart disease, diabetes, hypertension, and cancer in family and relatives) were used as basic information (BD), and smoking level, drinking level, processed food intake, exercise level, caffeine intake, sleep time, and depression level were used as lifestyle information (LD).
[0157] Next, in the case of stroke, age, height, weight, and family history (stroke, heart disease, diabetes, hypertension, and cancer in family and relatives) were used as basic information (BD), and smoking level, body mass index, meat consumption, salt content in diet, exercise level, caffeine intake, sleep time, stress level, and depression level were used as lifestyle information (LD).
[0158] Next, for diabetes, age, height, weight, and family history (family and relatives' history of stroke, heart disease, diabetes, high blood pressure, and cancer) were used as basic information (BD), and smoking level, drinking level, body mass index, meat consumption, fried food consumption, exercise level, caffeine consumption, sleep time, and depression level were used as lifestyle information (LD).
[0159] At this time, the integrated model (M_MS_IT) used in the experiment was trained through the non-negative multiple linear regression method.
[0160] Under these experimental conditions, the integrated model (M_MS_IT) directly calculated the screening risk (MS) based on basic information (BD) and lifestyle information (LD).
[0161] In the case of below, the results of the integrated model (M_MS_IT) predicting the screening risk (MS) and the correlation results (R) of the actual screening risk value 2 ) is shown.
[0162] DiseaseMaleFemaleAge-related cataract0.45680.6781Hypertension0.46560.6112Osteoporosis0.33490.6268Stroke0.27560.3071Type 2 diabetes0.41620.4517Average R20.38980.5350
[0163] As shown in the above and , both the individual model (M_MS_IN) and the integrated model (M_MS_IT) according to some embodiments of the present invention have the association results (R 2 ) has an average value of 0.39 or more and 0.55 or less, showing a high level of predictive power. Through this, it can be confirmed that the screening risk prediction model (M_MS) including the individual model (M_MS_IN) and the integrated model (M_MS_IT) according to some embodiments of the present invention can secure a high level of screening risk (MS) predictive power.
[0164]
[0165] FIG. 6 is a flowchart of a method for predicting periodic variable disease risk according to some embodiments of the present invention. Each step (S100 to S600) of FIG. 6 can be performed by the periodic variable disease risk prediction device (200 of FIGS. 1 and 2) of FIGS. 1 and 2 . Hereinafter, overlapping details will be briefly described.
[0166] Referring to FIGS. 1, 2, 4, 5c and 6, first, the periodic variable disease risk prediction device (200) can receive genetic information (GD) about the user's genes, lifestyle information (LD) about lifestyle habits and basic information (BD) about personal information (S100).
[0167] At this time, the genetic information (GD) may include information related to the genes in the user's body. For example, the genetic information (GD) may include DNA sequencing (Deoxyribo Nucleic Acid Sequencing), DNA chip (Deoxyribo Nucleic Acid Chip), PCR (Polymerase Chain Reaction) results for the user's skin, blood, etc., but the embodiments of the present invention are not limited thereto. The lifestyle information (LD) may include information about the user's daily living habits. For example, the lifestyle information (LD) may include the user's smoking information, drinking information, exercise information, food intake information, body mass index, sleep time, stress index, etc., but the embodiments of the present invention are not limited thereto. The food intake information may include information about the category of food consumed (e.g., meat, fish, processed food, fried food, caffeine), intake amount, salinity, etc., but it is obvious that the embodiments of the present invention are not limited thereto. At this time, the lifestyle information (LD) may be in the form of a survey result or in the form of text data obtained by post-processing the survey result. Basic information (BD) may include information about the user's personal details or basic facts related to a medical examination. For example, basic information (BD) may include the user's age, height, weight, family history, etc. However, embodiments of the present invention are not limited thereto, and basic information (BD) may further include information regarding the user's medical examination records. For example, basic information (BD) may further include hospital visit records, prescription records, examination records, health examination records, etc., but embodiments of the present invention are not limited thereto.
[0168] Next, the periodic variable disease risk prediction device (200) can calculate genetic risk (GS) based on genetic information (GD) (S200), can calculate lifestyle risk (LS) based on lifestyle information (LD) (S300), and can predict screening risk (MS) based on basic information (BD) and lifestyle information (LD) (S400).
[0169] For example, a periodic variable disease risk prediction device (200) can calculate a genetic risk (GS) by inputting genetic information (GD) into a genetic risk calculation model (M_GS), can calculate a lifestyle risk (LS) by inputting lifestyle information (LD) into a lifestyle risk calculation model (M_LS), and can calculate a screening risk (MS) by inputting basic information (BD) and lifestyle information (LD) into a screening risk prediction model (M_MS). At this time, the genetic risk calculation model (M_GS), lifestyle risk calculation model (M_LS), and screening risk prediction model (M_MS) can be pre-trained through statistical analysis (e.g., association analysis, Cox regression, Mendelian randomization), artificial intelligence, machine learning, deep learning, etc. to output genetic risk (GS), lifestyle risk (LS), and screening risk (MS) respectively when input data is entered.
[0170] At this time, the screening risk prediction model (M_MS) may include an individual model (M_MS_IN) that predicts screening information (MD) based on lifestyle information (LD) and basic information (BD) and then outputs screening risk (MS) based on the predicted screening information (MD), and an integrated model (M_MS_IT) that directly predicts screening risk (MS) based on lifestyle information (LD) and basic information (BD). A detailed description is omitted.
[0171] Next, the periodic variable disease risk prediction device (200) can calculate the comprehensive disease risk (IRS) based on the genetic risk (GS), lifestyle risk (LS), and screening risk (MS) (S500).
[0172] For example, a periodic variable disease risk prediction device (200) can calculate a comprehensive disease risk (IRS) by inputting genetic risk (GS), lifestyle risk (LS), and screening risk (MS) into a comprehensive disease risk calculation model (M_IRS). At this time, the comprehensive disease risk calculation model (M_IRS) can be learned in advance through statistical analysis (e.g., association analysis, Cox regression, Mendelian randomization), artificial intelligence, machine learning, deep learning, etc., so that when genetic risk (GS), lifestyle risk (LS), and screening risk (MS) are input, it outputs a comprehensive disease risk (IRS).
[0173] Next, the periodic variable disease risk prediction device (200) can output an integrated disease risk (IRS) (S600).
[0174] For example, the output module (240) can output the comprehensive disease risk score (IRS) in the form of a probability value as output data (OD). In other words, the output module (240) can provide the absolute score, which is the comprehensive disease risk score (IRS) itself in the form of a probability value provided from the prediction module (220), as output data (OD) to the user terminal (102), etc.
[0175] As another example, the output module (240) may post-process the comprehensive disease risk (IRS) in the form of a probability value to convert it into a ranking compared to a comparison group, and output the converted data as output data (OD). In other words, the output module (240) may compare the comprehensive disease risk (IRS) in the form of a probability value provided from the prediction module (220) with similar users in the comparison group, thereby calculating the risk ranking in the comparison group, and may provide the calculated ranking or the relative score based on the comparison with the comparison group as output data (OD) to the user terminal (102), etc.
[0176]
[0177] FIG. 7 is a diagram illustrating a hardware implementation of a periodic variable disease risk prediction device that performs a periodic variable disease risk prediction method according to some embodiments of the present invention.
[0178] Referring to FIGS. 1 and 7, a periodic variable disease risk prediction device (200) according to some embodiments of the present invention may be implemented as an electronic device (1000). The electronic device (1000) may include a controller (1010), an input / output device (1020), a memory device (1030), an interface (1040), and a bus (1050). The controller (1010), the input / output device (1020), the memory device (1030), and / or the interface (1040) may be coupled to each other via a bus (1050). In this case, the bus (1050) corresponds to a path through which data is transferred.
[0179] Specifically, the controller (1010) may include at least one of a CPU (Central Processing Unit), an MPU (Micro Processor Unit), an MCU (Micro Controller Unit), a GPU (Graphics Processing Unit), a microprocessor, a digital signal processor, a microcontroller, an application processor (AP), and logic elements capable of performing functions similar thereto.
[0180] The input / output device (1020) may include at least one of a keypad, a keyboard, a touchscreen, and a display device.
[0181] The memory device (1030) can store data and / or programs, etc.
[0182] The interface (1040) may perform a function of transmitting data to or receiving data from a communication network. The interface (1040) may be wired or wireless. For example, the interface (1040) may include an antenna or a wired / wireless transceiver. Although not illustrated, the memory device (1030) may further include high-speed DRAM and / or SRAM as an operating memory for improving the operation of the controller (1010). The memory device (1030) may store programs or applications therein.
[0183] The periodic variable disease risk prediction device (200) according to embodiments of the present invention may be a system formed by connecting multiple electronic devices (1000) via a network. In this case, each module or combination of modules may be implemented as an electronic device (1000). However, the present embodiment is not limited thereto.
[0184] Additionally, the periodic variable disease risk prediction device (200) may be implemented as at least one of a workstation, a data center, an internet data center (IDC), a direct attached storage (DAS) system, a storage area network (SAN) system, a network attached storage (NAS) system, a redundant array of inexpensive disks (RAID) system, and an electronic document management (EDMS) system, but the present embodiment is not limited thereto.
[0185] Additionally, the periodic variable disease risk prediction device (200) can transmit data to an external database (100) via a network. The network may include a network based on wired Internet technology, wireless Internet technology, and short-range communication technology. For example, the wired Internet technology may include at least one of a local area network (LAN) and a wide area network (WAN).
[0186] The wireless Internet technology may include, for example, at least one of Wireless LAN (WLAN), Digital Living Network Alliance (DMNA), Wireless Broadband (Wibro), World Interoperability for Microwave Access (Wimax), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), IEEE 802.16, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), Wireless Mobile Broadband Service (WMBS), and 5G NR (New Radio) technologies. However, the present embodiment is not limited thereto.
[0187] Short-range communication technologies may include, for example, at least one of Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra-Wideband (UWB), ZigBee, Near Field Communication (NFC), Ultra Sound Communication (USC), Visible Light Communication (VLC), Wi-Fi, Wi-Fi Direct, and 5G NR (New Radio). However, the present embodiment is not limited thereto.
[0188] A periodic variable disease risk prediction device (200) communicating through a network may comply with technical standards and standard communication methods for mobile communication. For example, the standard communication method may include at least one of GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTEA (Long Term Evolution-Advanced), and 5G NR (New Radio). However, the present embodiment is not limited thereto.
[0189] The above description is merely an example of the technical idea of the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of the present embodiment, but rather to explain it, and the scope of the technical idea of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.
Claims
1. A data collection module that receives genetic information about the user's genes, lifestyle information about lifestyle habits, and basic information about personal information; and Including a prediction module that predicts the comprehensive disease risk, which is the user's risk related to a predefined disease, based on the genetic information, the lifestyle information, and the basic information, The above prediction module, A first calculation unit that calculates a genetic risk related to the disease based on the genetic information; A second calculation unit that calculates the lifestyle risk related to the disease based on the above lifestyle information, A prediction unit that predicts the examination risk related to the user's medical examination record based on the above basic information and the above lifestyle information; Including an integrated calculation unit that calculates the comprehensive disease risk based on the genetic risk, the lifestyle risk, and the examination risk. A device for predicting the risk of periodic fluctuations in disease.
2. In paragraph 1, The above data collection module, Receiving the genetic information, the lifestyle information and the basic information from an external database including a genetic database linked to the comprehensive disease risk prediction device, a user terminal carried by the user and a medical record database storing the basic information. A device for predicting the risk of periodic fluctuations in disease.
3. In paragraph 2, The above data collection module, From the user terminal, the results of a survey related to the user's disease entered into the user terminal are received as the lifestyle information. A device for predicting the risk of periodic fluctuations in disease.
4. In paragraph 3, The above data collection module receives the survey results as the lifestyle information from the user terminal according to a predefined cycle, The above prediction unit predicts the examination risk based on each of the lifestyle information received according to the cycle, The above integrated calculation unit calculates the comprehensive disease risk according to the cycle based on the genetic risk, the lifestyle risk, and the examination risk generated according to the cycle. A device for predicting the risk of periodic fluctuations in disease.
5. In paragraph 1, The first output unit calculates the genetic risk using a pre-learned genetic risk calculation model and the genetic information, The second output unit calculates the life risk using a pre-learned life risk calculation model and the life information, The above prediction unit calculates the screening risk using a pre-learned screening risk prediction model, the basic information, and the lifestyle information, The above integrated calculation unit calculates the comprehensive disease risk using a pre-learned comprehensive disease risk calculation model, the genetic risk, the lifestyle risk, and the examination risk. A device for predicting the risk of periodic fluctuations in disease.
6. In paragraph 5, Further comprising a learning module for learning the genetic risk calculation model, the lifestyle risk calculation model, the screening risk prediction model, and the comprehensive disease risk calculation model. A device for predicting the risk of periodic fluctuations in disease.
7. In paragraph 6, The above learning module, when learning the screening risk prediction model, Determine relevant information related to lifestyle habits that are associated with the above diseases, The above screening risk prediction model is trained to predict the screening risk based on the basic information and the related information. A device for predicting the risk of periodic fluctuations in disease.
8. In paragraph 5, The above screening risk prediction model is, When the above basic information and the above lifestyle information are entered, an individual model including a first model that predicts examination information regarding the user's medical examination record, and a second model that calculates the examination risk based on the predicted examination information, When the above basic information and the above lifestyle information are entered, at least one of the integrated models that calculate the user's screening risk is included. A device for predicting the risk of periodic fluctuations in disease.
9. In paragraph 1, It further includes an output module that outputs the above comprehensive disease risk in the form of a probability value or converts the probability value into a rank compared to a comparison group and outputs it. A device for predicting the risk of periodic fluctuations in disease.
10. In a method for predicting a periodic variable disease risk performed by a periodic variable disease risk prediction device, A receiving step for receiving genetic information about the user's genes, lifestyle information about lifestyle habits, and basic information about personal information; and Including a prediction step of predicting the comprehensive disease risk, which is the user's risk related to a predefined disease, based on the genetic information, the lifestyle information and the basic information, The above prediction step is, A step of calculating the genetic risk related to the disease based on the genetic information; A step of calculating the lifestyle risk related to the disease based on the above lifestyle information, A step of predicting the examination risk related to the user's medical examination record based on the above basic information and the above lifestyle information, A step of calculating the comprehensive disease risk based on the genetic risk, the lifestyle risk, and the examination risk. A method for predicting the risk of comprehensive diseases with periodic fluctuations.
Citation Information
Patent Citations
Method, server and system for generating disease prediction models
KR1020180099185A
Apparatus and method for predicting disease risk score combining genetic risk score of related phenotypes
KR102087613B1
Method and System for personalized healthcare
KR102131973B1
Device and method for predicting periodically updated disease risk
KR102814030B1
Manufacturing apparatus for kimbugak
KR102894222B1