Deep learning-based device for predicting presence or absence of disease, pre-trained using language-format data converted from human microbiome data, and method thereof

The deep learning-based method addresses the generalization issue of existing models by pre-training on language-formatted microbiome data and fine-tuning for disease prediction, enhancing predictive accuracy and generalization.

WO2026089411A1PCT designated stage Publication Date: 2026-04-30IMMUNOBIOME INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/016619
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-10-17
Filing Date
2025-10-20
Publication Date
2026-04-30

Smart Images

  • Figure KR2025016619_30042026_PF_FP_ABST
    Figure KR2025016619_30042026_PF_FP_ABST
Patent Text Reader

Abstract

This method for predicting the presence or absence of a specific disease by using microbiome data, which is performed by a disease presence / absence prediction device, may comprise: a data acquisition step of acquiring multiple pieces of human microbiome data; a preprocessing step of converting the multiple pieces of human microbiome data into language-format data; a pre-training step of training a language model including at least one transformer encoder block by using the language-format data; a fine-tuning step of configuring a disease prediction model in which a prediction layer for predicting the presence or absence of a specific disease is combined with the language model, and fine-tuning the disease prediction model by using microbiome data related to the specific disease; and a prediction step of predicting the presence or absence of the specific disease by using the disease prediction model for which the fine-tuning has been completed.
Need to check novelty before this filing date? Find Prior Art

Description

Deep learning-based device and method for predicting the presence or absence of disease by pre-training on language-formatted data converted from human microbiome data

[0001] The present invention relates to a deep learning-based device for predicting the presence or absence of disease and a method thereof, which pre-learns data in a language format converted from human microbiome data.

[0002] Recently, active research has been conducted on the impact of the gut microbiome on human health. The gut microbiome is an aggregate of microorganisms residing in the human intestines, known to play important roles in digestive function, immune system regulation, and metabolic function. Imbalances in this microbiome are reported to be associated with various diseases, particularly chronic conditions such as metabolic disorders, inflammatory bowel disease, and obesity.

[0003] Live biotherapeutic products (LBPs) are garnering attention as a method for regulating the gut microbiome. LBPs are drugs that exert therapeutic effects by utilizing live microorganisms to regulate the intestinal environment and strengthen beneficial microbial communities. LBPs are being developed to improve the intestinal environment and enhance patient health in specific disease states by increasing the amount of beneficial microorganisms or inhibiting the proliferation of pathogenic microorganisms.

[0004] In addition, research applying machine learning technology to analyze gut microbiome data has recently been actively conducted. Machine learning is an excellent technology for analyzing and predicting large amounts of data, providing useful insights for identifying correlations between changes in microbiome composition and diseases, or for developing personalized treatments. For example, machine learning models can be used to analyze the gut microbial composition and gene expression patterns of specific patient groups to select probiotic therapies suitable for those patients or suggest optimal treatment methods.

[0005] Therefore, research on the gut microbiome combining probiotics and machine learning is expected to play a significant role in the development of next-generation drugs and personalized treatments.

[0006] Meanwhile, existing machine learning models for disease prediction based on gut microbiome data suffer from a generalization problem; while they demonstrate good predictive performance for specific cohorts, their performance declines for new samples. This appears to be because the models failed to learn the general environmental conditions of an individual's gut microbiome and instead trained only on data related to specific disease situations.

[0007] Accordingly, the present invention proposes a deep learning-based device and method for predicting the presence or absence of disease, which pre-learns language-formatted data converted from human microbiome data.

[0008] The present invention aims to provide a deep learning-based device and method for predicting the presence or absence of disease, which pre-trains on language-formatted data converted from human microbiome data, in order to solve the generalization problem of existing models as described above.

[0009] According to the first aspect of the present invention, a method for predicting the presence or absence of a specific disease using microbiome data performed in a disease presence or absence prediction device may include: a data acquisition step of acquiring a plurality of human microbiome data; a preprocessing step of converting the plurality of human microbiome data into language format data; a pre-training step of training a language model including at least one transformer encoder block using the language format data; a fine-tuning step of constructing a disease prediction model by combining a prediction layer for predicting the presence or absence of a specific disease with the language model and fine-tuning the disease prediction model using microbiome data related to the specific disease; and a prediction step of predicting the presence or absence of the specific disease using the fine-tuned disease prediction model. The preprocessing step may include: a step of extracting information on a plurality of metabolic pathways and the abundance of each metabolic pathway from the plurality of human microbiome data; and a step of converting each metabolic pathway into words and arranging each metabolic pathway in order based on the abundance of each metabolic pathway to generate the language format data.

[0010] According to the second aspect of the present invention, a disease presence prediction device for predicting the presence or absence of a specific disease using microbiome data comprises at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code enable the device to perform the method of the first aspect through the at least one processor.

[0011] The above means of solution are merely examples and should be interpreted as including all means within the scope that a person skilled in the art can understand from the description in this application.

[0012] The present invention can provide a deep learning-based device and method for predicting the presence or absence of disease, which pre-learns data in a language format converted from human microbiome data.

[0013] The above effects are merely examples and should be interpreted as including all effects within the scope that a person skilled in the art can understand from the description in this application.

[0014] Figure 1 shows a block diagram of a device for predicting the presence or absence of disease.

[0015] Figure 2 shows a flowchart of a method for predicting the presence or absence of disease.

[0016] Figure 3 illustrates the process of preprocessing microbiome data into a language format according to Embodiment 1.

[0017] Figure 4 shows a learning method of a pre-trained model according to Example 2.

[0018] Figure 5 shows the structure of a pre-training model according to Example 2.

[0019] Figure 6 shows a learning method for a fine-tuning model for configuring a disease prediction model according to Embodiment 2.

[0020] Figure 7 shows the structure of a fine-tuning model for configuring a disease prediction model according to Embodiment 2.

[0021] Figure 8a shows the loss of the microbiome data pre-training model according to Example 1. Here, the loss refers to the difference between the result predicted by the model for a metabolic pathway left blank and the actual metabolic pathway in the blank. Additionally, Epoch represents the number of training iterations, train represents the data used for training, and validation represents the data separated from training to verify the training results.

[0022] Figure 8b shows the accuracy of the microbiome data pre-training model according to Example 1. Here, accuracy refers to how accurately the metabolic pathway to be entered in the blank is matched. Additionally, Epoch represents the number of training iterations, train represents the data used for training, and validation represents the data separated from training to verify the training results.

[0023] Figures 9a to 9l show the results of distinguishing microbiome data characteristics according to an individual's age after birth before and after prior training of the learning model according to Example 2.

[0024] FIGS. 10a to 10d show the results of learning the characteristics according to the amount of intestinal Bacteroides depending on whether the learning model according to Example 2 was prior trained.

[0025] FIGS. 11a to 11d show the results of learning characteristics according to the intestinal microbiome type (enterotype) based on whether the learning model according to Example 2 was prior trained.

[0026] FIGS. 12a to 12d show the results of learning characteristics according to the degree of intestinal environmental imbalance (dysbiosis score) based on whether the learning model according to Example 2 was previously trained.

[0027] FIGS. 13a to 13d show the results of learning characteristics according to menopause status based on whether the learning model according to Example 2 was previously trained.

[0028] FIGS. 14a to 14d show the results of learning characteristics according to lifestyle based on whether the learning model according to Example 2 has been prior trained.

[0029] FIGS. 15a to 15d show the results of learning characteristics according to the village type based on whether the learning model according to Example 2 has been prior trained.

[0030] Figure 16 shows the results of verifying the performance of a disease presence prediction model through fine-tuning and comparison with existing models using data from various disease patients and non-patients.

[0031] Figures 17 to 22 show the results of performance verification of a colorectal cancer (CRC) presence or absence prediction model.

[0032] Figures 23 to 25 show the results of performance verification of a model for predicting the presence or absence of obesity (OB).

[0033] Figures 26 to 28 show the results of performance verification of a model for predicting the presence or absence of type 2 diabetes (T2D).

[0034] Figure 29 shows the results of performance verification of a model for predicting the presence or absence of atherosclerotic cardiovascular disease (ACVD).

[0035] Figure 30 shows the results of performance verification of the hypertension (HT) presence or absence prediction model.

[0036] Figure 31 shows the results of performance verification of a model for predicting the presence or absence of liver cirrhosis (LC).

[0037] Figure 32 shows the results of performance verification of a model for predicting the presence or absence of rheumatoid arthritis (RA).

[0038] Figures 33 and 34 show the results of performance verification of a model for predicting the presence or absence of Crohn's disease (CD).

[0039] Figures 35 and 36 show the results of performance verification of the ulcerative colitis (UC) presence or absence prediction model.

[0040] Figure 37 shows the results of performance verification of a neuroblastoma (NB) presence or absence prediction model.

[0041] Figures 38 to 41 show the results of performance verification of a model for predicting the presence or absence of autism spectrum disorder (ASD).

[0042] Embodiments of the present invention are described below with reference to the attached drawings to enable those skilled in the art to easily implement the invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0043] Throughout this specification, when a part is described as 'comprising' a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0044] Throughout the entire specification, ‘steps’ or ‘steps of’ do not mean ‘steps for’.

[0045] Throughout this specification, the term 'combination(s) of these' included in the Markush-type expression means one or more mixtures or combinations selected from the group consisting of the components described in the Markush-type expression, and means including one or more selected from the group consisting of said components.

[0046] Throughout this specification, the description of 'A and / or B' means 'A or B, or A and B'.

[0047] Throughout this specification, the term "microbiome" is a compound word formed from "microbiota" (or "microbe") and "genome," and refers to a microbial community that includes microorganisms and their entire genetic information. More recently, the term "microbiome" has also come to refer to microbial communities existing inside or outside the human body.

[0048] The functions realized by the components described herein may be implemented in a general-purpose processor, a specific-purpose processor, an integrated circuit, an Application Specific Integrated Circuit (ASIC), a Central Processing Unit (CPU), a circuit, and / or a combination thereof, which are programmed to realize the described functions. A processor may include transistors or other circuits and is considered to be a circuit or a processing circuit. A processor may be a programmed processor that executes a program stored in memory.

[0049] In this specification, circuits, parts, units, and means are hardware programmed to perform or execute the described functions. Such hardware may be any hardware disclosed in this specification or any hardware known to be programmed or execute the described functions.

[0050] If the hardware is a processor considered to be a circuit type, the circuit, the part, means, or unit is a combination of the hardware and the software used to constitute the hardware and / or processor.

[0051] Throughout the entire specification, "Live Biotherapeutic Products" (LBP) refers to pharmaceuticals containing live microorganisms, such as bacteria, that are suitable for human use, and used for the purpose of preventing or treating human diseases.

[0052] Throughout this specification, 'language models' and 'disease prediction models' are machine learning utilizing deep learning, and 'machine learning' may refer to an artificial intelligence application in which a computer program uses algorithms to find patterns in given data. It may primarily refer to a field in which a computer is trained to learn from data and improve through experience. The algorithms or learning methods of the 'language models' and 'disease prediction models' used herein are merely examples and should be interpreted to include all machine learning methods or types that can be used for the present invention. For example, machine learning methods may include (1) supervised learning, (2) unsupervised learning, (3) reinforcement learning, (4) semi-supervised learning, and more specifically, may include Naive Bayes Classification, Logistic Regression, Decision Tree, Random Forest, Boosting (XGBoost / ensemble boosting / AdaBoost / Gradient Boost / LightGBM / CatBoost, etc.), Perceptron, Support Vector Machine, Quadratic classifiers, Clustering (K-means clustering, Bayesian network clustering, etc.), but are not limited thereto.

[0053] Hereinafter, embodiments and examples of the present invention will be described in detail with reference to the attached drawings. However, the present invention may not be limited to these embodiments and examples and drawings.

[0054] The present invention aims to predict health status through the analysis of an individual's gut microbiome data and to develop a customized probiotic therapeutic agent by utilizing the prediction results.

[0055] This invention was developed based on the observation that existing microbiome data-based disease prediction models suffer from a problem of model generalization, which shows good predictive performance for specific cohorts but suffers from poor predictive performance for new samples.

[0056] Hereinafter, specific details for implementing the present invention will be described with reference to the attached block diagram or flowchart.

[0057] FIG. 1 is a block diagram of a device for predicting the presence or absence of disease according to one embodiment of the present invention.

[0058] Referring to FIG. 1, a disease presence or absence prediction device (1) may include a processor (100) and a memory (110) containing computer program code. However, since the disease presence or absence prediction device (1) of FIG. 1 is merely one embodiment of the present invention, the present invention is not to be interpreted as being limited by FIG. 1.

[0059] According to one embodiment of the present invention, a disease presence or absence prediction device (1) can perform the disease presence or absence prediction method described below through at least one memory (110) and at least one processor (100).

[0060] Below, a method for predicting the presence or absence of a specific disease using microbiome data performed by the disease presence or absence prediction device (1) of FIG. 1 is described in more detail.

[0061] FIG. 2 is a flowchart of a method for predicting the presence or absence of disease using the disease presence or absence prediction device (1) shown in FIG. 1, according to one embodiment of the present invention.

[0062] Referring to FIGS. 1 and 2, a processor (100) can perform a method for predicting the presence or absence of disease, comprising: a data acquisition step (S210) for acquiring a plurality of human microbiome data; a preprocessing step (S220) for converting the plurality of human microbiome data into language format data; a pre-training step (S230) for training a language model including at least one transformer encoder block using the language format data; a fine-tuning step (S240) for configuring a disease prediction model by combining a prediction layer for predicting the presence or absence of a specific disease with the language model and fine-tuning the disease prediction model using microbiome data related to the specific disease; and a prediction step (S250) for predicting the presence or absence of the specific disease using the disease prediction model after the fine-tuning is completed.

[0063] Here, the preprocessing step may include the step of extracting information on a plurality of metabolic pathways and the abundance of each metabolic pathway from the plurality of human microbiome data, and the step of converting each metabolic pathway into words and arranging each metabolic pathway in order based on the abundance of each metabolic pathway to generate the language format data.

[0064] In one embodiment of the present invention, the language model may be a transformer-based language model comprising at least one of BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (Robustly Optimized BERT Pretraining Approach), DistilBERT, ALBERT (A Lite BERT for Self-supervised Learning of Language Representations), and ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately).

[0065] In one embodiment of the present invention, the prior training step may involve masking a portion of the language format data through masked language modeling and training the language model to predict the masked portion.

[0066] In one embodiment of the present invention, the at least one transformer encoder block may include a self-attention mechanism and a feedforward neural network and may be learned in a self-supervised learning manner through the masked language modeling.

[0067] In one embodiment of the present invention, the prior learning step involves using the language format data to train the language model to learn the general characteristics of the human microbiome, and the general characteristics of the human microbiome may include information on the human gut environment.

[0068] In one embodiment of the present invention, the general characteristics of the human microbiome may include at least one of the following: characteristics according to age after birth, characteristics according to intestinal Bacteroides abundance, characteristics according to intestinal enterotype, characteristics according to intestinal dysbiosis score, characteristics according to menopause status, characteristics according to lifestyle, and characteristics according to village type.

[0069] In one embodiment of the present invention, the prediction layer may be composed of a neural network structure including at least one of a multilayer perceptron, a convolutional neural network, or linear regression.

[0070] In one embodiment of the present invention, the specific disease may include one or more of autism spectrum, colorectal cancer, obesity, colorectal adenoma, type 2 diabetes, ulcerative colitis, Crohn's disease, neuroblastoma, hypertension, impaired glucose tolerance, atherosclerotic cardiovascular disease, liver cirrhosis, rheumatoid arthritis, arteriosclerosis, type 1 diabetes, gastric cancer, liver cancer, breast cancer, cervical cancer, thyroid cancer, lung cancer, and prostate cancer.

[0071] Below, the method for predicting the presence or absence of disease according to the present invention will be described in detail.

[0072]

[0073] Implementation Example 1. Data preprocessing (steps S210, S220)

[0074] Human gut microbiome data can be expressed in terms of 'metabolic pathways' and 'metabolic pathway richness,' which are functions included in the microbiome. Since the microbiome refers to the entire community of microorganisms inhabiting a specific environment or within a host, 'metabolic pathways' here refer to the chemical reactions occurring within the microbial community, while 'metabolic pathway richness' refers to the quantity of such chemical reactions.

[0075] The aforementioned 'metabolic pathways' and 'metabolic pathway richness' can be extracted by methods such as extracting DNA from human stool samples, obtaining a FASTQ file containing nucleotide sequence information of DNA fragments via Next-Generation Sequencing (NGS), and applying various analytical methods to the file. For example, if the Biobakery pipeline is used as an analytical method, the 'metabolic pathways' and 'metabolic pathway richness' of the human gut microbiome can be measured based on the number of stool sample DNA fragments that match the microbial metabolic pathway gene database.

[0076] Additionally, referring to Figure 3, microbiome data can be converted into data with the attributes of a sentence by arranging the 'richness of metabolic pathways' from highest to lowest. For example, since a sentence is a sequence of words, the names of 'metabolic pathways' can be used as words in the sentence, and the order of words can be assigned by sorting the 'richness of metabolic pathways' from highest to lowest values. In particular, to distinguish the beginning and end of the sentence, special words "CLS" and "SEP," which can be recognized by a language model, are placed at the beginning and end of the sentence, thereby converting information about an individual's microbiome into a single sentence.

[0077] The present invention can preprocess human gut microbiome data into a language format in the same manner as the above example and use it for training a language model for pre-training (hereinafter, language model).

[0078] Implementation Example 2. Language Model Training (Steps S230, S240)

[0079] The learning of a language model according to the present invention includes a 'pre-learning stage' and a 'fine-tuning stage'.

[0080] 1) Pre-learning stage

[0081] Referring to Figure 4, pre-training is a process in which a language model learns the general characteristics of the human microbiome using a large amount of human microbiome information. Here, general characteristics refer to information regarding the human gut environment, and include, for example, characteristics according to age after birth, characteristics according to the abundance of Bacteroides in the gut, characteristics according to the enterotype, characteristics according to the dysbiosis score, characteristics according to menopause status, characteristics according to lifestyle, and characteristics according to the village type in which one resides. In addition, it should be interpreted as including all gut information that can be used to predict the presence or absence of any disease.

[0082] In addition, referring to Fig. 5, for the pre-learning method, for example, a masked language modeling method can be used, which is a method of learning to understand the characteristics of a sentence by learning the relationships between words by making some of the words (dialogue paths) of a sentence blank in the language model and learning to infer the word to go in the blank.

[0083] When using the masked language modeling method above, 15% of the 'metabolic pathways' in an individual's microbiome data are randomly treated as 'blanks' and the learning model is instructed to predict the 'metabolic pathways' that fill in the blanks. By learning to identify the relationships between 'metabolic pathways' within the microbiome data and predict the 'metabolic pathways' in the blanks, the model learns the general characteristics of the human microbiome.

[0084] As a preprocessing step for pre-training, a step that represents microbiome data preprocessed into a language format as numerical values ​​(e.g., an embedding module, etc.) and a step that learns the relationships between 'metabolic pathways' within the microbiome (e.g., a transformer encoder block) may be additionally included.

[0085] 2) Fine-tuning

[0086] Referring to Figures 6 and 7, fine-tuning is a process in which a pre-trained language model is adjusted to predict which disease a person has based on the general characteristics of the human microbiome learned. In this process, the pre-trained model is trained by cross-referencing microbiome information of patients with a specific disease with microbiome information of non-patients without the disease, thereby enabling the model to learn the microbiome characteristics specific to a specific disease among the general characteristics.

[0087] For such fine-tuning, a disease prediction model can be constructed by combining a prediction layer that forecasts the presence or absence of a specific disease with a language model, and then transfer learning can be performed to retrain parameters using microbiome data related to the specific disease. For example, the disease prediction model can utilize multilayer perceptrons or convolutional neural networks (CNNs), which can convert microbiome characteristic information into disease probability values, by combining them as adapters with the pre-trained model.

[0088] Example 1. Verification of pre-training using preprocessed data

[0089] The inventors verified through two methods whether the language model could actually pre-learn the preprocessed data according to the present invention.

[0090] The first verification method is to check whether the loss decreases. Loss refers to the difference between the result predicted by the language model for a 'metabolic pathway' created as a 'blank' and the 'metabolic pathway' that actually existed in the 'blank'. If the training is proceeding accurately, the loss will decrease as the number of training epochs increases.

[0091] Referring to Fig. 8a, the verification results showed that the loss value of the language model decreased as the number of training iterations increased, confirming that pre-training is possible using the preprocessed data according to the present invention in the language model.

[0092] The second validation method is to check the degree of accuracy (accuracy) of accurately predicting the 'metabolic pathways' that were left as 'blanks'. If prior training was not performed properly, the training model will randomly predict the 'metabolic pathways' of the 'blanks', resulting in an accuracy of 0.03% (1 / 3228 (total number of metabolic pathways)); if training was performed properly, the accuracy for the 'blanks' will increase as the number of training iterations increases. In Figure 8b, training accuracy refers to the accuracy on the data used for training, while validation accuracy refers to the accuracy on the data separated from training to validate the training results.

[0093] Referring to Fig. 8b, the verification results showed that the prediction accuracy for blanks increased as the number of training iterations increased, confirming that pre-training is possible using the preprocessed data according to the present invention in the language model.

[0094] Through the two methods described above, the inventor confirmed that a language model can actually be pre-trained using the preprocessed data according to the present invention.

[0095] Example 2. Verification of whether general characteristics of the human microbiome were learned

[0096] The inventors verified whether the language model according to the present invention actually learned the general characteristics of the human microbiome. The verification was confirmed by determining whether the general characteristics of the human microbiome, known through prior research, were naturally learned during the learning phase without any input values.

[0097] The general characteristics of the human microbiome used for verification are as follows:

[0098] [1] Characteristics based on age after birth: Since the degree of engraftment of the gut microbiome begins to form after birth, the gut microbiome environment may vary depending on the age of birth.

[0099] [2] Characteristics based on intestinal Bacteroides abundance: Since intestinal Bacteroides abundance is highly correlated with an individual's health status, differences in characteristics based on intestinal Bacteroides abundance may occur depending on health status.

[0100] [3] Characteristics of the gut microbiome type (enterotype): According to prior research, individuals have three types of microbiome types depending on their dietary characteristics and lifestyle habits. Therefore, differences in the gut microbiome may occur depending on individual characteristics, etc.

[0101] [4] Characteristics based on the degree of intestinal dysbiosis score: It is known that an individual's degree of intestinal dysbiosis is closely related to their health status. Therefore, differences in the degree of intestinal dysbiosis may occur depending on health status.

[0102] [5] Characteristics based on menopause status: A woman's menopausal status is closely linked to hormonal changes, and these hormonal changes can affect the composition and function of the gut microbiome. Therefore, differences in the characteristics of the gut microbiome may occur depending on menopausal status.

[0103] [6] Characteristics based on lifestyle: An individual's lifestyle habits, such as diet, exercise, sleep, and drinking habits, can have a direct impact on the diversity and stability of the gut microbiome. Therefore, the gut microbiome environment can vary depending on differences in lifestyle.

[0104] [7] Characteristics based on village type: The local environment in which an individual lives can affect the gut microbiome through external factors such as drinking water, access to food ingredients, and hygiene conditions. Therefore, the characteristics of the gut microbiome may vary depending on the residential area (e.g., rural areas, slums on the outskirts of cities, etc.).

[0105] Since the general characteristics of the human microbiome are influenced by an individual's age, lifestyle, and surrounding environment, these individual-specific characteristics are reflected in each person's gut environment. Therefore, in this embodiment, to evaluate whether the pre-trained model has learned the general characteristics of the human microbiome, the criterion was based on whether the model could distinguish each of the seven characteristics related to an individual's gut environment. Additionally, the performance of the pre-trained model according to the present invention was further verified by using a model that had not undergone the pre-training step as a comparison group.

[0106] 2-1) Characteristics according to age after birth (Figs. 9a to 9l)

[0107] As a result of verification, as illustrated in Figures 9a to 9l, a comparison of the T-SNE dimensionality reduction results before and after pre-training (Figures 9a, 9b, 9e, 9f, 9i, and 9j) revealed a tendency for samples with similar postnatal ages to be more densely distributed after pre-training. This suggests that the model independently learned general characteristics associated with postnatal age through the relationships between metabolic pathways within the microbiome, even though no label information regarding postnatal age was input during the pre-training process. This trend was confirmed not only by differences in visual distribution within the same cohort but also by a consistent increase in quantitative indicators, such as the silhouette score (Figures 9c, 9g, and 9k) and the Calinski-Harabasz score (Figures 9d, 9h, and 9l), after pre-training compared to before. In addition, the same improvement pattern was observed in all three independent analyses using different cohorts (Figs. 9a to 9d, Figs. 9e to 9h, Figs. 9i to 9l), proving that the pre-training model according to the present invention stably improves the performance of distinguishing age characteristics after birth.

[0108] 2-2) Characteristics according to intestinal Bacteroides abundance (Figs. 10a to 10d)

[0109] As a result of verification, as illustrated in FIGS. 10a to 10d, a comparison of the T-SNE dimensionality reduction results before and after pre-training (Figs. 10a and 10b) confirmed that samples with similar intestinal Bacteroides abundance tended to be distributed more densely with one another after pre-training. This implies that the model independently learned general characteristics associated with Bacteroides abundance through the relationships between metabolic pathways within the microbiome, even though label information regarding Bacteroides abundance was not separately input during the pre-training process. This trend was confirmed to increase after pre-training compared to before, not only in visual distribution differences but also in quantitative indicators such as the silhouette score (Fig. 10c) and the Calinski-Harabasz score (Fig. 10d). Therefore, it can be seen that the pre-training model according to the present invention secures discrimination performance regarding Bacteroides abundance.

[0110] 2-3) Characteristics according to enterotype (Figs. 11a to 11d)

[0111] As a result of verification, as illustrated in FIGS. 11a to 11d, a comparison of the T-SNE dimensionality reduction results (Figs. 11a and 11b) before and after pre-training revealed a tendency for samples with the same enterotype to be more densely distributed after pre-training. This implies that the model independently learned general characteristics associated with enterotypes through the relationships between metabolic pathways within the microbiome, even though label information regarding enterotypes was not separately input during the pre-training process. This trend was confirmed to consistently increase after pre-training compared to before, not only in visual distribution differences but also in quantitative indicators such as the silhouette score (Fig. 11c) and the Calinski-Harabasz score (Fig. 11d). Therefore, it can be seen that the pre-training model according to the present invention secures the performance to distinguish enterotypes.

[0112] 2-4) Characteristics according to the degree of intestinal environmental imbalance (dysbiosis score) (Figs. 12a to 12d)

[0113] As a result of verification, as illustrated in FIGS. 12a to 12d, a comparison of the T-SNE dimensionality reduction results (Figs. 12a and 12b) before and after pre-training revealed a tendency for samples with similar dysbiosis scores to be more densely distributed after pre-training. This implies that the model independently learned general characteristics associated with dysbiosis scores through the relationships between metabolic pathways within the microbiome, even though label information regarding dysbiosis scores was not separately input during the pre-training process. This trend was confirmed to consistently increase after pre-training compared to before, not only in visual distribution differences but also in quantitative indicators such as the silhouette score (Fig. 12c) and the Calinski-Harabasz score (Fig. 12d). Therefore, it can be seen that the pre-training model according to the present invention secures discrimination performance regarding dysbiosis scores.

[0114] 2-5) Characteristics according to menopause status (Figs. 13a to 13d)

[0115] As a result of verification, as illustrated in FIGS. 13a to 13d, a comparison of the T-SNE dimensionality reduction results (Figs. 13a and 13b) before and after pre-training revealed a tendency for samples with the same individual menopausal status to be distributed more densely with one another after pre-training. This implies that the model independently learned general characteristics associated with menopausal status through the relationships between metabolic pathways within the microbiome, even though label information regarding menopausal status was not separately input during the pre-training process. This trend was confirmed to consistently increase after pre-training compared to before, not only in visual distribution differences but also in quantitative indicators such as the silhouette score (Fig. 13c) and the Calinski-Harabasz score (Fig. 13d). Therefore, it can be seen that the pre-training model according to the present invention secures the performance to distinguish individual menopausal status.

[0116] 2-6) Characteristics according to lifestyle (Figs. 14a to 14d)

[0117] As a result of verification, as illustrated in FIGS. 14a to 14d, a comparison of the T-SNE dimensionality reduction results (Figs. 14a and 14b) before and after pre-training revealed a tendency for samples with identical personal lifestyle habits to be distributed more densely with one another after pre-training. This implies that the model independently learned general characteristics associated with lifestyle habits through the relationships between metabolic pathways within the microbiome, even though label information regarding lifestyle habits was not separately input during the pre-training process. This trend was confirmed to consistently increase after pre-training compared to before, not only in visual distribution differences but also in quantitative indicators such as the silhouette score (Fig. 14c) and the Calinski-Harabasz score (Fig. 14d). Therefore, it can be seen that the pre-training model according to the present invention secures performance in distinguishing lifestyle habits.

[0118] 2-7) Characteristics according to the village type inhabited (Figs. 15a to 15d)

[0119] As a result of verification, as illustrated in FIGS. 15a to 15d, a comparison of the T-SNE dimensionality reduction results (Figs. 15a and 15b) before and after pre-training revealed a tendency for samples with the same type of village where individuals reside to be more densely distributed after pre-training. This implies that the model independently learned general characteristics associated with the type of village through the relationships between metabolic pathways within the microbiome, even though label information regarding the type of village resided was not separately input during the pre-training process. This trend was confirmed to consistently increase after pre-training compared to before, not only in visual distribution differences but also in quantitative indicators such as the silhouette score (Fig. 15c) and the Calinski-Harabasz score (Fig. 15d). Therefore, it can be seen that the pre-training model according to the present invention secures the performance to distinguish the type of village where individuals reside.

[0120] Therefore, it was confirmed that the pre-learning model according to the present invention can independently distinguish general characteristics of the human microbiome as a result of learning, even though it did not directly receive information such as age after birth, abundance of intestinal Bacteroides, type of intestinal microbiome, degree of intestinal environmental imbalance, menopausal status, lifestyle habits, and type of residential village during the pre-learning process. This implies that the pre-learning model of the present invention effectively learns the general characteristics of the human microbiome.

[0121] Example 3. Verification of whether a disease presence / absence prediction model can be established through fine-tuning

[0122] The inventors verified whether the pre-trained model according to the present invention could establish a model capable of predicting the presence or absence of a specific disease (a disease presence or absence prediction model) through a fine-tuning step, and verified the prediction accuracy of the said disease presence or absence prediction model. To this end, using an actual disease dataset, they compared the impact of fine-tuning on disease presence or absence prediction with a model that had not undergone fine-tuning and with existing widely used machine learning models. For verification, a Monte-Carlo cross-validation technique was applied in which the entire dataset was randomly split in a 7:3 ratio, with 70% used as the training dataset and 30% as the test dataset for evaluating predictive power. The stability of the model was confirmed by conducting a total of 15 iterated tests.

[0123] By comparing the case where a disease presence prediction model was established based on a fine-tuned model with the case where a model was established based on a model that had not undergone the fine-tuning step, it was confirmed that fine-tuning contributes to improving the accuracy of disease presence prediction. In addition, the performance of the model of the present invention was verified by comparing it with nine widely used machine learning models and two existing indicator-based models. The comparison subjects included logistic regression, random forest (RF), and support vector machine (SVM) with pathway abundance (pathwayAB), rank, and z-score normalized values, respectively, as inputs. Additionally, the Gut Microbiome Health Index (GMHI) developed by Gupta et al. and the Shannon index, one of the alpha diversity (a-diversity) indicators, were included as comparison models.

[0124] The disease presence prediction model was established by classifying it into two main cases. First, a prediction model was established that distinguishes between patients and non-patients without distinguishing the type of disease. To this end, a total of 36 datasets were integrated, and data from 3,903 patients and 3,471 non-patients were used for training and validation. The patient group included patients with autism spectrum disorder (ASD), colorectal cancer (CRC), obesity (OB), colorectal adenoma, type 2 diabetes (T2D), ulcerative colitis (UC), Crohn's disease (CD), neuroblastoma (NB), hypertension (HT), impaired glucose tolerance, atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), rheumatoid arthritis (RA), symptomatic atherosclerosis (SA), and type 1 diabetes (T1D). The number of datasets and patients by disease type analyzed is shown in Table 1 below.

[0125] Number of datasets and patients by disease type analyzed Kinds of diseases No. of datasets N_patients Autism (ASD) 6 199 Colorectal cancer (CRC) 6 1288 Obesity (OB) 5 628 Colorectal adenoma 3 188 Type 2 diabetes (T2D) 3 244 Ulcerative colitis (UC) 2 320 Crohn's disease (CD) 2 226 Neuroblastoma (NB) 2 132 Hypertension (HT) 199 Impaired glucose tolerance 149 Atherosclerotic cardiovascular disease (ACVD) 1214 Liver Cirrhosis (LC) 1165 Rheumatoid arthritis (RA) 195 Symptomatic atherosclerosis (SA) 114 Type 1 diabetes (T1D) 142

[0126] Second, a disease-specific prediction model was established to distinguish between patients and non-patients for each individual disease. The diseases analyzed included colorectal cancer (CRC), obesity (OB), type 2 diabetes (T2D), ulcerative colitis (UC), Crohn's disease (CD), neuroblastoma (NB), hypertension (HT), atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), rheumatoid arthritis (RA), and autism spectrum disorder (ASD). Independent prediction models were established for each disease, and their performance was verified.

[0127] 3-1) Result of establishing a prediction model that distinguishes between patients and non-patients without distinguishing by type of disease (Fig. 16)

[0128] The inventors established a model for predicting the presence or absence of disease based on a pre-training model according to the present invention and verified the prediction accuracy of the model. To this end, a microbiome dataset consisting of 3,903 patients and 3,471 non-patients was utilized.

[0129] Figure 16A illustrates the results of comparing the performance of a disease presence prediction model built based on a fine-tuned model with that of a model built based on a non-fine-tuned model. In the case of fine-tuning, the predictive power continuously improved, recording a high Area Under the Receiver Operating Characteristic curve (AUROC) value, whereas in the case of non-fine-tuning, the predictive power remained limited. From this, it was found that the fine-tuning step directly contributes to the improvement of disease presence prediction power.

[0130] Figure 16B illustrates the results of comparing the performance of the disease presence prediction model (IMBERT) according to the present invention with existing machine learning-based models (logistic regression, random forest, support vector machine, etc.) and existing indicator-based methods (GMHI, Shannon diversity index, etc.). As a result, the pre-trained model (IMBERT) of the present invention showed a superior AUROC value compared to existing machine learning-based models and indicator-based models, confirming that the disease presence prediction performance is significantly improved.

[0131] Therefore, the pre-training model according to the present invention can effectively distinguish between patients and non-patients based on microbiome data, and it has been proven that the fine-tuning step plays an important role in improving the accuracy of predicting the presence or absence of disease.

[0132] 3-2) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Colorectal cancer (Figs. 17 to 22)

[0133] The inventors verified whether the pre-training model according to the present invention can establish a model for predicting the presence or absence of colorectal cancer (CRC) and verified the prediction accuracy of said model. To this end, a total of six independent colorectal cancer datasets were utilized.

[0134] Figures 17A, 18C, 19E, 20G, 21I, and 22K respectively illustrate the results of comparing the performance of a colorectal cancer presence prediction model built based on fine-tuning with that of a colorectal cancer presence prediction model built based on a model without fine-tuning. In all datasets, the predictive power was superior when based on fine-tuning, which means that fine-tuning contributes to the improvement of colorectal cancer presence prediction performance.

[0135] Figures 17B, 18D, 19F, 20H, 21J, and 22L illustrate the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0136] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance across multiple independent colorectal cancer datasets, which demonstrates that the model can effectively predict the presence or absence of a specific disease (colorectal cancer).

[0137] 3-3) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Obesity (Figs. 23 to 25)

[0138] The inventors verified whether the pre-training model according to the present invention can establish a model that predicts the presence or absence of obesity (OB) and verified the prediction accuracy of said model. To this end, a total of three independent obesity datasets were utilized.

[0139] Figures 23A, 24C, and 25E respectively illustrate the results of comparing the performance of an obesity prediction model built based on fine-tuning with that of an obesity prediction model built based on a model without fine-tuning. In all datasets, the predictive power was superior when based on fine-tuning, which means that fine-tuning contributes to the improvement of obesity prediction performance.

[0140] FIGS. 23B, FIGS. 24D, and FIGS. 25F illustrate the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0141] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance across multiple independent obesity datasets, which demonstrates that the model can effectively predict the presence or absence of a specific disease (obesity).

[0142] 3-4) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Type 2 diabetes (Figs. 26 to 28)

[0143] The inventors verified whether the pre-training model according to the present invention can establish a model that predicts the presence or absence of type 2 diabetes (T2D) and verified the prediction accuracy of said model. To this end, a total of three independent type 2 diabetes datasets were used.

[0144] Figures 26A, 27C, and 28E respectively illustrate the results of comparing the performance of a model for predicting the presence of type 2 diabetes based on fine-tuning with a model for predicting the presence of type 2 diabetes based on a model without fine-tuning. In all datasets, the predictive power was superior when based on fine-tuning, which means that fine-tuning contributes to improving the performance of predicting the presence of type 2 diabetes.

[0145] FIGS. 26B, FIGS. 27D, and FIGS. 28F illustrate the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0146] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance on multiple independent Type 2 diabetes datasets, which demonstrates that the model can effectively predict the presence or absence of a specific disease (Type 2 diabetes).

[0147] 3-5) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Atherosclerotic cardiovascular disease (Fig. 29)

[0148] The inventors verified whether the pre-training model according to the present invention can establish a model for predicting the presence or absence of atherosclerotic cardiovascular disease (ACVD) and verified the prediction accuracy of said model. To this end, an independent ACVD dataset was utilized.

[0149] Figure 29A illustrates the results of comparing the performance of an ACVD presence prediction model built based on fine-tuning with that of an ACVD presence prediction model built based on a model without fine-tuning. When based on fine-tuning, the predictive power was significantly improved, which means that fine-tuning contributes to the improvement of ACVD presence prediction performance.

[0150] Figure 29B illustrates the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0151] Therefore, the pre-trained model according to the present invention secured excellent predictive performance on an independent ACVD dataset, which demonstrates that the model can effectively predict the presence or absence of a specific disease (ACVD).

[0152] 3-6) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Hypertension (Fig. 30)

[0153] The inventors verified whether the pre-training model according to the present invention can establish a model that predicts the presence or absence of hypertension (HT) and verified the prediction accuracy of said model. To this end, an independent hypertension dataset was utilized.

[0154] Figure 30A illustrates the results of comparing the performance of a hypertension presence prediction model built based on fine-tuning with that of a hypertension presence prediction model built based on a model without fine-tuning. When based on prior training, the predictive power was relatively improved, which means that fine-tuning contributes to the improvement of hypertension presence prediction performance.

[0155] Figure 30B illustrates the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0156] Therefore, the pre-trained model according to the present invention secured a predictive performance of a certain level or higher in the hypertension dataset, which demonstrates that the model can effectively predict the presence or absence of a specific disease (hypertension).

[0157] 3-7) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Liver cirrhosis (Fig. 31)

[0158] The inventors verified whether the pre-training model according to the present invention can establish a model that predicts the presence or absence of liver cirrhosis (LC) and verified the prediction accuracy of said model. To this end, an independent liver cirrhosis dataset was used.

[0159] Figure 31A illustrates the results of comparing the performance of a liver cirrhosis presence prediction model built based on fine-tuning with a liver cirrhosis presence prediction model built based on a model without fine-tuning. When based on fine-tuning, the predictive power was significantly improved, which means that fine-tuning contributes to the improvement of liver cirrhosis presence prediction performance.

[0160] Figure 31B illustrates the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0161] Therefore, the pre-trained model according to the present invention achieved a high level of predictive performance in the liver cirrhosis dataset, which demonstrates that the model can effectively predict the presence or absence of a specific disease (liver cirrhosis).

[0162] 3-8) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Rheumatoid arthritis (Fig. 32)

[0163] The inventors verified whether the pre-training model according to the present invention can establish a model that predicts the presence or absence of rheumatoid arthritis (RA) and verified the prediction accuracy of said model. To this end, an independent rheumatoid arthritis dataset was used.

[0164] Figure 32A illustrates the results of comparing the performance of a rheumatoid arthritis presence prediction model built based on fine-tuning with a rheumatoid arthritis presence prediction model built based on a model without fine-tuning. When based on fine-tuning, the predictive power was significantly improved, which means that fine-tuning contributes to the improvement of rheumatoid arthritis presence prediction performance.

[0165] Figure 32B illustrates the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0166] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance in the rheumatoid arthritis dataset, which demonstrates that the model can effectively predict the presence or absence of a specific disease (rheumatoid arthritis).

[0167] 3-9) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Crohn's disease (Figs. 33 and 34)

[0168] The inventors verified whether the pre-training model according to the present invention can establish a model for predicting the presence or absence of Crohn's disease (CD) and verified the prediction accuracy of said model. To this end, a total of two independent Crohn's disease datasets were utilized.

[0169] Figures 33A and 34C respectively illustrate the results of comparing the performance of a Crohn's disease presence prediction model built based on fine-tuning and a Crohn's disease presence prediction model built based on a model without fine-tuning. In both datasets, the predictive power was significantly improved when based on fine-tuning, which means that fine-tuning contributes to the improvement of Crohn's disease presence prediction performance.

[0170] FIGS. 33B and FIGS. 34D illustrate the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0171] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance on the Crohn's disease dataset, which demonstrates that the model can effectively predict the presence or absence of a specific disease (Crohn's disease).

[0172] 3-10) Results of establishing a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Ulcerative colitis (Figs. 35 and 36)

[0173] The inventors verified whether the pre-training model according to the present invention can establish a model that predicts the presence or absence of ulcerative colitis (UC) and verified the prediction accuracy of said model. To this end, a total of two independent ulcerative colitis datasets were used.

[0174] Figures 35A and 36C respectively illustrate the results of comparing the performance of an ulcerative colitis presence prediction model built based on fine-tuning and an ulcerative colitis presence prediction model built based on a model without fine-tuning. In both datasets, the predictive power was improved when based on fine-tuning, which means that fine-tuning contributes to the improvement of ulcerative colitis presence prediction performance.

[0175] FIGS. 35B and FIGS. 36D illustrate the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0176] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance in the ulcerative colitis dataset, which demonstrates that the model can effectively predict the presence or absence of a specific disease (ulcerative colitis).

[0177] 3-11) Establishment of a disease-specific prediction model distinguishing between patients and non-patients by individual disease - Neuroblastoma (Fig. 37)

[0178] The inventors verified whether the pre-training model according to the present invention can establish a model for predicting the presence or absence of neuroblastoma (NB) and verified the prediction accuracy of said model. To this end, a total of two independent neuroblastoma datasets were used.

[0179] Figure 37A illustrates the results of comparing the performance of a neuroblastoma presence prediction model built based on fine-tuning with that of a neuroblastoma presence prediction model built based on a model without fine-tuning. Predictive power was improved when based on fine-tuning, which means that fine-tuning contributes to the improvement of neuroblastoma presence prediction performance.

[0180] Figure 37B illustrates the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0181] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance in the neuroblastoma dataset, which demonstrates that the model can effectively predict the presence or absence of a specific disease (neuroblastoma).

[0182] 3-12) Results of establishing a disease-specific predictive model distinguishing patients and non-patients by individual disease - Autism Spectrum (Figs. 38 to 41)

[0183] The inventors verified whether the pre-training model according to the present invention can establish a model for predicting the presence or absence of autism spectrum disorder (ASD) and verified the prediction accuracy of said model. To this end, a total of four independent ASD datasets were utilized.

[0184] Figures 38A, 39C, 40E, and 41G respectively illustrate the results of comparing the performance of an ASD presence prediction model built based on fine-tuning with that of an ASD presence prediction model built based on a model without fine-tuning. In all datasets, the predictive power was superior when based on fine-tuning, which means that pre-training contributes to the improvement of ASD presence prediction performance.

[0185] FIGS. 38B, FIGS. 39D, FIGS. 40F, and FIGS. 41H illustrate the results of comparing the performance of a pre-trained model (IMBERT) according to the present invention with that of existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based indicators (GMHI, Shannon diversity index, etc.). As a result of evaluation based on AUROC, the pre-trained model of the present invention showed overall equivalent or superior performance compared to other models.

[0186] Therefore, the pre-trained model according to the present invention has consistently achieved excellent predictive performance across multiple independent ASD datasets, which demonstrates that the model can effectively predict the presence or absence of a specific disease (autism spectrum disorder).

[0187] In conclusion, the fine-tuning step was confirmed to play an important role in predicting the presence or absence of disease. In tests for various diseases (colorectal cancer, obesity, type 2 diabetes, atherosclerotic cardiovascular disease, hypertension, liver cirrhosis, rheumatoid arthritis, Crohn's disease, ulcerative colitis, neuroblastoma, autism spectrum disorder, etc.) illustrated in FIGS. 16 to 41, the predictive power was consistently improved when fine-tuning was performed compared to when fine-tuning was not performed. Furthermore, the pre-trained model (IMBERT) according to the present invention demonstrated overall equivalent or superior predictive performance when compared to existing machine learning models (logistic regression, random forest, support vector machine, etc.) and existing gut microbiome-based health indices (GMHI, Shannon diversity index, etc.). These results demonstrate the technical feasibility and excellence of establishing a disease presence or absence prediction model based on fine-tuning, which can effectively predict the presence or absence of various diseases by utilizing microbiome data.

[0188] Furthermore, the diseases verified in this embodiment are not limited to all diseases that can be predicted by utilizing information regarding the general characteristics of the human microbiome, namely the intestinal environment, such as stomach cancer, liver cancer, breast cancer, cervical cancer, thyroid cancer, lung cancer, prostate cancer, etc.

[0189] [Explanation of the symbol]

[0190] 1: Disease Presence Prediction Device

[0191] 100: Processor

[0192] 110: Memory

Claims

1. A method for predicting the presence or absence of a specific disease using microbiome data performed in a disease presence or absence prediction device, A data acquisition step for acquiring multiple human microbiome data; A preprocessing step of converting the above plurality of human microbiome data into language format data; A pre-training step of training a language model including at least one transformer encoder block using the above language format data; A fine-tuning step for constructing a disease prediction model by combining a prediction layer that predicts the presence or absence of a specific disease with the above language model, and fine-tuning the disease prediction model using microbiome data related to the specific disease; and A method comprising a prediction step of predicting the presence or absence of a specific disease using the disease prediction model with the fine-tuning completed.

2. In Paragraph 1, The above preprocessing step is, A step of extracting information on a plurality of metabolic pathways and the abundance of each metabolic pathway from the plurality of human microbiome data; and A method comprising the step of converting each of the above-mentioned metabolic pathways into words and arranging each of the above-mentioned metabolic pathways in order based on the richness of each of the above-mentioned metabolic pathways to generate the above-mentioned language format data.

3. In Paragraph 1, The method wherein the language model is a transformer-based language model comprising at least one of BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (Robustly Optimized BERT Pretraining Approach), DistilBERT, ALBERT (A Lite BERT for Self-supervised Learning of Language Representations), and ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately).

4. In Paragraph 1, The above pre-learning step is, A method of masking a portion of the language format data through masked language modeling and training the language model to predict the masked portion.

5. In Paragraph 4, A method wherein at least one transformer encoder block comprises a self-attention mechanism and a feedforward neural network, and is learned in a self-supervised learning manner through the masked language modeling.

6. In Paragraph 1, The above pre-learning step is, The method involves using the above language format data to train the above language model on the general characteristics of the human microbiome, and A method in which the general characteristics of the human microbiome described above include information on the human gut environment.

7. In Paragraph 6, A method wherein the general characteristics of the human microbiome described above include at least one of the following: characteristics according to age after birth, characteristics according to intestinal Bacteroides abundance, characteristics according to intestinal enterotype, characteristics according to intestinal dysbiosis score, characteristics according to menopause status, characteristics according to lifestyle, and characteristics according to village type.

8. In Paragraph 1, A method in which the prediction layer is composed of a neural network structure including at least one of a multilayer perceptron, a convolutional neural network, or a linear regression.

9. In Paragraph 1, A method wherein the specific disease mentioned above includes one or more of autism spectrum, colorectal cancer, obesity, colorectal adenoma, type 2 diabetes, ulcerative colitis, Crohn's disease, neuroblastoma, hypertension, impaired glucose tolerance, atherosclerotic cardiovascular disease, liver cirrhosis, rheumatoid arthritis, arteriosclerosis, type 1 diabetes, gastric cancer, liver cancer, breast cancer, cervical cancer, thyroid cancer, lung cancer, and prostate cancer.

10. A disease presence or absence prediction device that predicts the presence or absence of a specific disease using microbiome data, At least one processor; and It includes at least one memory containing computer program code, and The above at least one memory and the above computer program code are through the above at least one processor the device, Acquire multiple human microbiome data, and The above multiple human microbiome data are preprocessed to convert them into language format data, and A language model including at least one transformer encoder block is pre-trained using the above language format data, and A disease prediction model is constructed by combining a prediction layer that predicts the presence or absence of a specific disease with the above language model, and the above disease prediction model is fine-tuned using microbiome data related to the specific disease. A disease presence or absence prediction device that predicts the presence or absence of the specific disease using the fine-tuned disease prediction model.

11. In Paragraph 10, The above preprocessing is, Information on multiple metabolic pathways and the abundance of each metabolic pathway is extracted from the above multiple human microbiome data, and A device that converts each of the above-mentioned metabolic pathways into words and sorts each of the above-mentioned metabolic pathways in order based on the richness of each of the above-mentioned metabolic pathways to generate the above-mentioned language format data.

12. In Paragraph 10, The device, wherein the language model is a transformer-based language model comprising at least one of BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (Robustly Optimized BERT Pretraining Approach), DistilBERT, ALBERT (A Lite BERT for Self-supervised Learning of Language Representations), and ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately).

13. In Paragraph 10, The above prior learning is, An apparatus that masks a portion of the language format data through masked language modeling and trains the language model to predict the masked portion.

14. In Paragraph 13, A device wherein at least one transformer encoder block comprises a self-attention mechanism and a feedforward neural network, and is learned in a self-supervised learning manner through the masked language modeling.

15. In Paragraph 10, The above prior learning is, The method involves using the above language format data to train the above language model on the general characteristics of the human microbiome, and A device in which the general characteristics of the human microbiome described above include information on the human intestinal environment.

16. In Paragraph 15, A device wherein the general characteristics of the human microbiome described above include at least one of the following: characteristics according to age after birth, characteristics according to intestinal Bacteroides abundance, characteristics according to intestinal enterotype, characteristics according to intestinal dysbiosis score, characteristics according to menopause status, characteristics according to lifestyle, and characteristics according to village type.

17. In Paragraph 10, A device in which the above prediction layer is composed of a neural network structure including at least one of a multilayer perceptron, a convolutional neural network, or a linear regression.

18. In Paragraph 10, A device wherein the specific diseases mentioned above include one or more of autism spectrum, colorectal cancer, obesity, colorectal adenoma, type 2 diabetes, ulcerative colitis, Crohn's disease, neuroblastoma, hypertension, impaired glucose tolerance, atherosclerotic cardiovascular disease, liver cirrhosis, rheumatoid arthritis, arteriosclerosis, type 1 diabetes, gastric cancer, liver cancer, breast cancer, cervical cancer, thyroid cancer, lung cancer, and prostate cancer.