Methods for using machine learning models to derive microbial community data associated with specific diseases

A machine learning model is used to derive disease-associated microbial communities, addressing the limitations of single-strain probiotics by identifying comprehensive bacterial communities associated with diseases, including novel species.

JP7748113B2Active Publication Date: 2025-10-02IMMUNOBIOME INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023210030
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2023-12-13
Publication Date
2025-10-02
Estimated Expiration
2043-12-13

AI Technical Summary

Technical Problem

Current probiotics and live biotherapeutic products (LBPs) often require specific nutritional requirements and have difficulty in maintaining viability, and using single candidate strains may not effectively address the complex microbial community changes associated with diseases.

Method used

A method using a machine learning model to derive disease-associated microbial community data by acquiring gut microbiota data, deriving candidate microbial community data, and then deriving disease-related microbial community data through a computing device.

Benefits of technology

Enables the identification of entire bacterial communities associated with specific diseases, including novel species, rather than individual microbial species, enhancing the effectiveness of probiotics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007748113000009
    Figure 0007748113000009
  • Figure 0007748113000010
    Figure 0007748113000010
  • Figure 0007748113000011
    Figure 0007748113000011
Patent Text Reader

Abstract

To provide a method for deriving a disease specific microorganism community for improving probiotics effective for a disease.SOLUTION: The method includes the steps of: (1) deriving microorganism community data (hereinafter referred to as a candidate group consortium) of a specific disease-related candidate, obtained by pre-treating bacterial flora data (a taxonomy abundance table; and (2) causing the machine learning model to learn the candidate group consortium, selecting a model with the most excellent performance, and deriving a disease-microorganism consortium.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for deriving microbial community data associated with a particular disease using machine learning models. [Background technology]

[0002] Changes in the structure of the intestinal microbiota have various effects on the human body, including aging and health. In particular, imbalances in the normal intestinal microbiota (dysbiosis) in disease are associated with a wide range of systemic symptoms, including gastrointestinal conditions such as inflammatory bowel disease (IBD), obesity, and atopy. Therefore, probiotics, which are live microorganisms that can beneficially affect host health, have emerged as a way to regulate the structure of the intestinal microbiota. Currently, probiotics and live biotherapeutic products (LBPs), available as foods or food supplements (e.g., functional health foods), are attracting attention as novel therapeutic modalities for diseases.

[0003] Next-generation probiotics (NGPs), which treat diseases through the intestinal microbiota, target only specific microorganisms eliminated by disease, rather than transplanting the entire bacterial environment of a healthy individual. This allows for disease-specific treatment. While numerous candidate bacteria have been identified as NGPs, most of them require specific nutritional requirements. Furthermore, obtaining a large biomass of these bacteria is difficult, and even if successfully obtained, maintaining their viability over the long term is challenging. However, even if these issues are addressed, a single candidate strain may not be effective enough. This is due to the nature of the intestinal microbiota, which is an ecosystem of various microorganisms. In other words, the specific microorganisms eliminated by a disease are not individual microorganisms but a microbial community (microbiota), so a single candidate strain may not be effective enough. Therefore, it is expected that better efficacy can be achieved by using a disease-relevant microbe consortium (a group of two or more symbiotic microorganisms) as an LBP rather than individual strains.

[0004] As a similar example of using microbial information to determine disease, European Patent No. 3097211 presents a method for obtaining microbial information from a patient using a sampling kit and analyzing the microbial information within the patient based on this, and Chinese Patent No. 114854847 provides a method for creating a machine learning model for determining the presence or absence of disease based on the host's genetic or microbial information. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] European Patent No. 3097211 [Patent Document 2] China Patent No. 114854847 Summary of the Invention [Problem to be solved by the invention]

[0006] As part of the development of next-generation probiotics, the present applicants aim to provide a method for deriving microbial community data associated with specific diseases (hereinafter referred to as disease-microbial consortium). [Means for solving the problem]

[0007] The present application provides a method for deriving disease-associated microbial community data utilizing a machine learning model through a computing device, the method comprising: (1) acquiring gut microbiota data and deriving candidate microbial community data from the gut microbiota data; (2) deriving disease-related microbial community data from the candidate microbial community data.

[0008] The present application further provides an apparatus for deriving disease-related microbial community data using a machine learning model through a computing device, the apparatus including: an acquisition unit for acquiring gut microbiota data; a candidate consortium derivation unit for deriving candidate microbial community data from the gut microbiota data; and a disease-microbial consortium derivation unit for deriving disease-related microbial community data from the candidate microbial community data. [Effects of the Invention]

[0009] Advantages of the present invention include the ability to derive the entire collection of bacteria associated with a particular disease, such as a community of associated microorganisms, a bacterial flora or a microbiome, rather than simply finding information about individual microbial species associated with a particular disease. [Brief explanation of the drawings]

[0010] [Figure 1] This shows the process according to the present invention. [Figure 2] This shows the process of deriving a candidate consortium. [Figure 3] This shows the process of deriving a disease-microbe consortium. [Figure 4] This shows the process of dividing the data into a training set and a test set. [Figure 5] (A) Comparison of the performance of each algorithm in the process of deriving obesity-related consortia from amplicon sequencing data, (B) Comparison of the performance difference between statistically based methods and the algorithm used in this application in the process of deriving obesity-related consortia from amplicon sequencing data, (C) Information on the best-performing machine learning model in the process of deriving obesity-related consortia from amplicon sequencing data. [Figure 6] (A) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in Cohort 1 during the process of deriving obesity-related consortia from amplicon sequencing data. (B) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in C0 during the process of deriving obesity-related consortia from amplicon sequencing data. [Figure 7] (A) The distances between the constituent bacteria of C0 and between bacteria in consortia different from C0 during the process of deriving obesity-related consortia using amplicon sequencing data are shown. (B) The distances in A are visualized using a PCA plot. [Figure 8](A) Comparison of the performance of each algorithm in the process of deriving CDI-related consortia from amplicon sequence data, (B) Comparison of the performance difference between statistically based methods and the algorithm used in this application in the process of deriving CDI-related consortia from amplicon sequence data, and (C) Information on the best-performing machine learning model in the process of deriving CDI-related consortia from amplicon sequence data. [Figure 9] (A) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in Cohort 2 during the process of deriving CDI-associated consortia from amplicon sequencing data. (B) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in C17 during the process of deriving CDI-associated consortia from amplicon sequencing data. [Figure 10] (A) Distances between the constituent bacteria of C17 and bacteria in consortia different from C17 during the process of deriving CDI-associated consortia using amplicon sequencing data. (B) PCA plot visualizing the distances in Figure 9a. [Figure 11] (A) A comparison of the performance of each algorithm in the process of deriving RA-related consortia from amplicon sequence data; (B) A comparison of the performance difference between statistically based methods and the algorithm used in this application in the process of deriving RA-related consortia from amplicon sequence data; (C) Information on the best-performing machine learning model in the process of deriving RA-related consortia from amplicon sequence data. [Figure 12](A) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in Cohort 3 during the process of deriving RA-associated consortia from amplicon sequencing data. (B) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in C1 during the process of deriving RA-associated consortia from amplicon sequencing data. [Figure 13] (A) Distances between the constituent bacteria of C1 and between bacteria in consortia different from C1 during the process of deriving RA-related consortia using amplicon sequencing data are shown. (B) Distances in A are visualized using a PCA plot. [Figure 14] (A) Comparison of the performance of each algorithm in the process of deriving disease-microorganism consortia from whole metagenomic sequence data; (B) Comparison of the performance difference between statistically based methods and the algorithm used in this application in the process of deriving disease-microorganism consortia from whole metagenomic sequence data; (C) Information on the best-performing machine learning model in the process of deriving disease-microorganism consortia from whole metagenomic sequence data. [Figure 15] This shows cross-cohort prediction results for the best-performing machine learning model in the process of deriving disease-microbial consortia from whole metagenomic sequencing data. [Figure 16] (A) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in Cohort 4 during the process of deriving RA-associated consortia from whole metagenomic sequencing data. (B) The abundance of disease-microbe consortia versus feature importance for the best-performing machine learning model in C3 during the process of deriving RA-associated consortia from whole metagenomic sequencing data. [Figure 17](A) Distances between the constituent bacteria of C3 and between bacteria in consortia different from C3 during the process of deriving RA-related consortia using whole metagenomic data. (B) Distances in A visualized using a PCA plot. [Figure 18] 1 is a flowchart of the present method. [Figure 19] FIG. 1 is a block diagram of the present device. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present application will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily carry out the present application. However, the present application may be embodied in various different forms and is not limited to the embodiments described herein. In the drawings, parts that are not relevant to the description are omitted for clarity, and similar parts are designated by similar reference numerals throughout the specification.

[0012] Throughout this specification, when a member is said to be "on" another member, this includes not only when the member is in contact with the other member, but also when there is another member between the two members.

[0013] Throughout the specification of this application, when a part "comprises" a certain element, this means that it may further include other elements, not excluding other elements, unless otherwise specified. Terms of degree such as "about," "substantially," and the like used throughout the specification of this application are used to mean a numerical value or a close vicinity of a numerical value when inherent manufacturing and material tolerances are presented, and are used to prevent unscrupulous infringers from unfairly taking advantage of disclosures in which precise or absolute numerical values ​​are recited to aid in the understanding of this application. Terms of degree such as "steps for" or "steps of" used throughout the specification of this application do not mean "steps for."

[0014] Throughout the specification of this application, the term "combination(s) of these" contained in a Markush expression means a mixture or combination of one or more selected from the group of elements set forth in the Markush expression, and is meant to include one or more selected from the group of elements.

[0015] Throughout the specification of this application, the phrase "A and / or B" means "A or B, or A and B."

[0016] Throughout this specification, the term "machine learning" refers to an artificial intelligence application in which a computer program uses an algorithm to find patterns in given data. It primarily refers to the field of training a computer to learn from data and improve through experience. The machine learning algorithm used in this application is merely an example and should be construed to include all machine learning methods or types that can be used for the present invention. For example, machine learning methods may include (1) supervised learning, (2) unsupervised learning, (3) reinforcement learning, and (4) semi-supervised learning. More specifically, machine learning methods may include, but are not limited to, Naive Bayes Classification, Logistic Regression, Decision Tree, Random Forest, Boosting (XGBoost / ensemble boosting / AdaBoost / Gradient Boost / LightGBM / CatBoost, etc.), Perceptron, Support Vector Machine, Quadratic Classifiers, Clustering (K-means clustering, Bayesian network clustering, etc.), etc.

[0017] Throughout this specification, "gut microbiota" refers to the complex community of microorganisms that live in the digestive tract of humans and other animals, including insects, such as the gut, gut microbiota, or gastrointestinal microbiota.

[0018] Throughout this specification, "Supervised Learning" refers to the learning process of a machine learning model in which the model labels a specific set of data and learns according to a purpose, while "Unsupervised Learning" refers to the opposite, in which the model clusters similar features within a specific set of data without being given a purpose, and predicts results for new data.

[0019] Throughout this specification, "clustering" means dividing the entire data set into groups of similar individuals (data) within the given data.

[0020] Throughout this specification, "Quality Control" refers to the process of establishing quality standards and ensuring that all manufacturing / fabrication processes comply with the established quality standards in order to maintain consistent quality of the results.

[0021] Throughout the specification of this application, the term "consortium" refers to community data. For example, a "candidate consortium" refers to microbial community data that can be candidates for the microbial community data ultimately to be derived in this application, and a "disease-microbial consortium" refers to microbial community data that can distinguish disease patients from non-patients. In other words, it refers to a collection of symbiotic bacteria discovered for a specific disease.

[0022] The present invention is aimed at selecting bacteria that can be used to develop probiotics effective against a certain disease. Specifically, the present invention involves the steps of (1) deriving candidate microbial community data (hereinafter referred to as a candidate consortium) related to a specific disease by preprocessing the microbiota data of a specific cohort (a taxonomy abundance table), (2) training this data in a machine learning model, selecting the best-performing model from the results, and deriving a disease-microbial consortium (Figure 1).

[0023] The above process (1) is a process of deriving a candidate group consortium using unsupervised learning, and the process (2) is a process of deriving microbial data (disease-microbial consortium) most relevant to a specific disease from the candidate group consortium using supervised learning. Specifically, the specific disease includes, but is not limited to, obesity, CDI (Clostridioides Difficile Infection), and RA (Rheumatoid Arthritis).

[0024] Hereinafter, embodiments and examples of the present application will be described in detail with reference to the accompanying drawings, but the present application is not limited to these embodiments and examples or the drawings.

[0025] Example 1. Derivation of candidate microbial consortia Various clustering algorithms based on the similarity between bacteria were applied to derive candidate cluster consortia (Fig. 2).

[0026] 1-1. Calculation of pairwise taxonomic similarity between bacteria For feature scaling, relevant taxonomic abundance values ​​were converted to percent values ​​per sample. To define the similarity between fungi, six similarity metrics were applied to all pairs of fungi. For similarity calculations, the SciPy (v1.8.0) Python library was used to calculate the distance between fungi, and the similarity was calculated by subtracting the distance from 1 (calculated as 1-distance).

[0027] 1-2.Clustering To derive the candidate cluster consortium, clustering was performed using three clustering algorithms, hierarchical clustering, K-means clustering, and Gaussian mixture model, along with methodological variations. These algorithms were applied to a matrix of pairwise taxanomic similarities calculated in 1-1. Clustering was performed while varying Ncluster, the number of clusters, from 21 to 60. Each cluster represents a microbial community data (consortium) that could potentially become a candidate cluster. Specific information about the similarity metric, algorithm variables, and Ncluster variable used in the above process can be found in Table 1 below. Through this process, 1,680 attempts were made to form candidate cluster consortia (initial cluster data) for each disease. This entire process was done using the Scikit-learn package (v0.24.1).

[0028] [Table 1]

[0029] 1-3. Removal of low-quality consortium, QC (Quality control) Low-quality consortia were selected and removed from the 1,680 candidate consortia. The quality of each consortium was assessed based on the number of bacterial species present. Specifically, consortia containing one, two, or more than half of the total microbial taxa of the entire consortium were classified as low-quality.

[0030] 1-4. Derivation of candidate consortium Consortial richness was defined as the arithmetic sum of taxa within the community.

[0031] Example 2. Derivation of a disease-microbe consortium This process derives the microbial community data (consortium) most closely associated with a specific disease. To do this, we selected the microbial consortium with the highest feature importance in the machine learning model that showed the highest predictive performance. During the model training process, the machine learning model adjusts feature importance, and a machine learning model with higher predictive performance indicates that the adjusted feature importance can contribute to predicting future data (Figure 3).

[0032] 2-1. Model training The inventors trained a machine learning model using all types of candidate group consortia derived in Example 1 above. During training, the candidate group consortium data was divided into a training set and an experimental set using Monte Carlo random sampling. The training set was used to train the machine learning model, and each data set was labeled as healthy (control) or diseased (case). Four algorithms were used: logistic regression, naive Bayes, random forest, and support vector machines (SVM). The hyperparameters for each machine learning algorithm were determined using a grid search strategy. This strategy involves performing k-fold cross-validation (CV) with various combinations of hyperparameters to find the best-performing hyperparameter based on CV performance. The proportion of healthy / disease samples was preserved in all MC sampling and k-fold CV. The above learning process was repeated five times, each with a different learning set (Figure 4). The model hyperparameters are summarized in Table 2.

[0033] [Table 2]

[0034] 2-2.Evaluation of the model The predictive performance of machine learning was evaluated using the experimental group. Model performance was evaluated through AUROC (area under the ROC curve).

[0035] 2-3. Selection of the best-performing machine learning model The best-performing machine learning model was selected through the following two steps:

[0036] (1) Identify the best-performing machine learning algorithms (2) Select the best-performing machine learning model from the algorithm.

[0037] The process of selecting the best-performing machine learning algorithm involved comparing the median predictive performance of each algorithm with that of other algorithms. The algorithm with the highest predictive performance was selected from the best-performing machine learning algorithms. Ultimately, the machine learning model with the best predictive performance within the selected algorithm was determined to be the best-performing machine learning model.

[0038] Prediction performance was evaluated using cross-validation on amplicon sequencing data. For the whole metagenomic data, the best-performing model was first evaluated through cross-validation, and its performance in cross-cohort prediction was evaluated by training it using the entire training cohort. This evaluation was performed to ensure the generality of the model across other cohorts. To achieve this, only taxa found in the training cohort were used.

[0039] 2-4. Derivation of disease-microbe consortium Disease-microbial consortia were derived using the absolute feature importance of the best-performing machine learning model. To assess the association between microbial consortia and disease, we compared the abundance of the consortium between healthy and disease groups. To determine the degree of interconnection between consortium members, we also compared the distance between microorganisms within and outside the consortium.

[0040] 2-5. Classification using methods based on statistical tests (STAT) As a baseline model, we generated a classifier based on STAT. Specifically, in the training set, we identified the consortium abundance that best distinguished between the healthy and diseased groups using the Mann-Whitney U test (MWU). Consequently, the critical point for distinguishing between the healthy and diseased groups was determined based on the best predictive performance observed in the training set. The predictive performance of the STAT-based method was evaluated using the test set.

[0041] Example 3. Verification process 3-1. Amplicon sequencing data For three other disease cohorts, 16Sr RNA amplicon sequencing data were obtained from the MicrobiomeHD database (https: / / doi.org / 10.6084 / m9.figshare.14531724.v1). Information for each cohort is shown in Table 3.

[0042] [Table 3]

[0043] 3-2.Metagenomic shotgun sequencing data The metagenomics shotgun sequencing data were analyzed by Gupta et al. 1) Disease information was assigned to samples based on the data in Liu R et al., Le Chatelier E et al., Jie Z et al., and raw fastq files were analyzed. 2)3)4)The species abundance tables were obtained from the BioBakery workflow. 5) The information for each cohort is shown in Table 4.

[0044] [Table 4]

[0045] 3-3. Learning parameters To train the machine learning model, we performed MC sampling with a training group:experimental group ratio of 8:2. This process was repeated five times. Furthermore, stratified 2-fold division was used for GridSearch cross-validation. For the STAT method, the MC sampling with the 8:2 ratio was applied 50 times. Performance evaluation was based on AUROC. Cohort 4 was selected to select the best-performing machine learning algorithm and model. For inter-cohort prediction, Cohort 4 was used as the training cohort because it showed an appropriate balance between the diseased and healthy groups, and Cohorts 5 and 6 were used as experimental cohorts.

[0046] Example 4. Derivation of disease-microbial consortia using 16S rRNA amplicon sequencing data 4-1. Obesity-related consortium The logistic regression algorithm was identified as the best-performing machine learning algorithm. This algorithm demonstrated the highest predictive performance (median AUROC: 0.698), outperforming other machine learning algorithms (Figure 5A). It was able to more accurately predict the patient's disease status compared to statistically based methods (Figure 5B). Cohort 1 was used to train and evaluate the machine learning model.

[0047] The best-performing machine learning model showed an AUROC of 0.796 for the logistic regression algorithm (Figure 5C). In this machine learning model, taxonomic similarity was measured through correlation, and a consortium of candidates was identified using the kmeans clustering algorithm with parameters set as "algorithms=full" and "Ncluster=42" (Figure 5C).

[0048] The best-performing machine learning model identified C0 as the most disease-related consortium among the consortia. This showed the highest absolute value for feature importance (Figure 6A). Furthermore, C0 abundance showed a significant difference between the obese group (median: 0.196) and the healthy group (median: 0.114) (P = 4.27x10 -8 , Figure 6B).

[0049] We confirmed that the bacteria in C0 are indeed related to each other. We confirmed that the constituent bacteria of C0 are closer to each other than the distance between bacteria in consortia different from C0 (Fig. 7A). The distances were visualized by a PCA plot (Fig. 7B).

[0050] The accuracy of the confirmed disease-associated consortium for obesity is supported by the results of other previous studies. C0 includes the family Ruminococcaceae (Table 5), which is consistent with Peters et al. (2018) 6) Consistent with the findings of previous studies, the study confirmed that bacteria within the Ruminococcaceae family, such as Oscillibacter, were largely absent in obese patients. Furthermore, C0 includes other genera, such as Incertae Sedis XIII, Desulfovibrionaceae, and unclassified species. These results demonstrate the comprehensive capabilities of the present invention to identify not only previously identified species in obesity-associated bacteria, but also novel species within the disease-related consortium.

[0051] [Table 5]

[0052] 4-2.CDI (Clostridioides difficile infection) related consortium The random forest algorithm was confirmed as the best-performing machine learning algorithm. This algorithm demonstrated the highest predictive performance (median AUROC: 0.994), outperforming the other four machine learning algorithms (Figure 8A). This enabled more accurate prediction of patient disease status than statistically-based methods (Figure 8B).

[0053] The best-performing machine learning model, the random forest algorithm, showed an AUROC of 1.0 (Figure 8C). In this machine learning model, taxonomic similarity was measured through correlation, and a consortium of candidates was identified using the GMM clustering algorithm with parameters set to covariance=full and Ncluster=22 (Figure 8C).

[0054] The best-performing machine learning model identified C17 as a disease-associated consortium, which exhibited the highest absolute value for feature importance (Figure 9A). Furthermore, C17 abundance showed a significant difference between the CDI group (median: 0.001) and the healthy group (median: 0.062) (P = 1.06x10 -30 , Figure 9B).

[0055] We confirmed that the bacteria within C17 are indeed closely related to each other. The constituent bacteria of C17 are even closer to each other than the distance between C17 and bacteria from different consortia (Figure 10A). The distances were visualized by a PCA plot (Figure 10B).

[0056] The accuracy of the disease-associated consortium for confirmed CDI is supported by other previous studies. C17 includes Lachnospiraceae and Ruminococcus (Table 6). This is consistent with Martinez et al. (2022) 7) Consistent with the findings of a study by

[1999] , the study confirmed that Lachnospiraceae and Ruminococcus bacteria were rarely present in CDI patients. Furthermore, C17 includes Acholeplasma and Anaerovorax. These results demonstrate the comprehensive capability of the present invention to identify not only species previously identified in previous studies but also new species within the disease-related consortium of bacteria associated with CDI.

[0057] [Table 6]

[0058] 4-3.RA (Rheumatoid arthritis) related consortium The logistic regression algorithm was identified as the best-performing machine learning algorithm. This algorithm demonstrated the highest predictive performance (median AUROC: 0.907), outperforming the other four machine learning algorithms (Figure 11A). This enabled more accurate prediction of patient disease status than statistically-based methods (Figure 11B).

[0059] The best-performing machine learning model showed an AUROC of 1.0 for the logistic regression algorithm (Figure 11C). In this machine learning model, taxonomic similarity was measured using "dice," and a candidate consortium was identified using the "kmeans" clustering algorithm with parameters set as "algorithms=full" and "Ncluster=22" (Figure 11C).

[0060] C1 was identified as the most RA-associated consortium, showing the highest absolute value for feature importance (Fig. 12A). Furthermore, C1 abundance showed a significant difference between the RA group (median: 0.444) and the healthy group (median: 0.012) (P = 8.38 x 10 -05 , Figure 12B).

[0061] We confirmed that the bacteria within C1 are indeed closely related to each other (Figure 13). We confirmed that the constituent bacteria of C1 are even closer to each other than the distance between C1 and bacteria from different consortia (Figure 13A). The distances were visualized by a PCA plot (Figure 13B).

[0062] The accuracy of the disease-associated consortium for confirmed RA is supported by the results of other previous studies. C1 includes Prevotella (Table 7). Many studies 8)9)10) confirmed an increase in Prevotella in RA (rheumatoid arthritis) patients compared to healthy controls. Furthermore, C1 contains bacteria such as Anaerotruncus, Pseudoflavonifractor, and Dialister. These results demonstrate that the present invention has comprehensive capabilities in identifying bacteria associated with RA, not only bacteria identified in previous studies but also new bacteria within the disease-related consortium.

[0063] [Table 7]

[0064] Example 5. Derivation of disease-microbial consortia from whole metagenome sequencing data The present applicant demonstrated that the present invention can derive disease-microbial consortia using whole metagenomic sequencing data. The data has the advantage of providing species-level information, unlike 16srRNA amplicon sequencing data, which provides genus-level information. This data was used to confirm whether the present invention can effectively derive disease-microbial consortia regardless of the type of sequencing data.

[0065] 5-1. Selecting the best performing machine learning model The random forest algorithm was identified as the best-performing machine learning algorithm (Figure 14). This algorithm demonstrated the highest predictive performance (median AUROC: 0.854), outperforming the other four machine learning algorithms (Figure 14A). This enabled more accurate prediction of patient disease status than statistically-based methods (Figure 14B).

[0066] The best-performing machine learning model, a random forest algorithm, showed an AUROC of 0.959 (Figure 14C). In this machine learning model, taxonomic similarity was measured through correlation, and a consortium of candidates was identified using the GMM clustering algorithm with parameters set to covariance=spherical and Ncluster=48 (Figure 14C).

[0067] The best-performing algorithm also proved capable of predicting disease status in obese patients. In cross-cohort prediction, the best-performing machine learning algorithm predicted disease status more accurately than statistically based predictions across two independent datasets (Figure 15).

[0068] 5-2. Identifying obesity-related consortia C3 was identified as the most obesity-related consortium. It showed the highest absolute value in feature importance (Figure 16A). Furthermore, C3 abundance showed a significant difference between the obese group (median: 0.012) and the healthy group (median: 0.004) (P = 2.55x10 -12 , Figure 16B).

[0069] We confirmed that the bacteria within C3 are indeed closely related to each other. The constituent bacteria of C3 are even closer than the distance between C3 and bacteria from different consortia (Figure 17A). The distances were visualized by a PCA plot (Figure 17B).

[0070] The accuracy of the confirmed disease-related consortium for obesity is supported by the results of other previous studies. C3 includes Collinsella aerofaciens, Eubacterium hallii, and Dorea longicatena (Table 8). Liu et al. 15) identified increased numbers of Collinsella aerofaciens, Eubacterium hallii, and Dorea longicatena in the obese group. Furthermore, C3 included Streptococcus salivarius, Blautia obeum, Solobacterium moorei, and other previously unreported bacteria. These results demonstrate the comprehensive ability of the present invention to identify not only previously identified species in obesity-associated bacteria, but also new species within the disease-related consortium.

[0071] [Table 8]

[0072] The present invention suggests that disease-microorganism consortia can be accurately derived between many diseases and sequencing platforms, which means that it can be applied to a variety of diseases.

[0073] Furthermore, the present invention has demonstrated that it is possible to discover new disease-associated bacteria within the disease-microbe consortium. This not only expands our understanding of disease-associated microbial communities, but also aids in the development of strategies to discover new microorganisms that can be used to alleviate disease. Therefore, the present invention not only contributes to the advancement of knowledge about disease-associated microbial communities, but also aids in new discoveries for disease management.

[0074] 1)Gupta VK et al.,A predictive index for health status using species-level gut microbiome profiling,Nature Communications,2020,11,4635 2)Liu R et al., Gut microbiome and serum metabolome alterations in obesity and after weight-loss intervention,Nature Medicine,2017,23,859-868 3)Le Chatelier E et al.,Richness of human gut microbiome correlates with metabolic markers,Nature,2013,500,541-546 4)Jie Z et al.,The gut microbiome in atherosclerotic cardiovascular disease, Nature Communications,2017,8,845 5)Beghini F et al.,Integrating taxonomic,functional,and strain-level profiling of diverse microbial communities with bioBakery 3,eLife,2021 6)14.Peters BA et al.,A taxonomic signature of obesity in a large study of American adults,Nature scientific reports,2018,8,9749 7)Martinez E et al.,Gut Microbiota Composition Associated with Clostridioides difficile Colonization and Infection,Pathogens,2022,11(7),781 8)Scher JU et al.,Expansion of intestinal Prevotella copri correlates with enhanced susceptibility to arthritis,Elife,2013,2,e01202 9)Xu et al.,Interactions between Gut Microbiota and Immunomodulatory Cells in Rheumatoid Arthritis,Mediators of Inflammation,2020 10)Zhao T et al.,Gut microbiota and rheumatoid arthritis:From pathogenesis to novel therapeutic opportunities,Frontiers Immunology,2022,13,1007165 11)Liu R et al.,Gut microbiome and serum metabolome alterations in obesity and after weight-loss intervention,Nature Medicine,2017,23,859-868 12)Schubert AM et al., Microbiome Data Distinguish Patients with Clostridium difficile Infection and Non-C.difficile-Associated Diarrhea from Healthy Controls,mBio,2014,5(3),e01021-14 13)Scher JU et al.,Expansion of intestinal Prevotella copri correlates with enhanced susceptibility to arthritis,Elife,2013,2,e01202 14) Liu R et al., Gut microbiome and serum metabolome alterations in obesity and after weight-loss intervention, Nature Medicine, 2017, 23, 859-868 15) Le Chatelier E et al., Richness of human gut microbiome correlates with metabolic markers, Nature, 2013, 500, 541-546 16)Jie Z et al.,The gut microbiome in atherosclerotic cardiovascular disease,Nature Communications,2017,8,845 [Explanation of symbols]

[0075] 19: Device for deriving disease-related microbial community data 1910: Acquisition Department 1920: Candidate Group Consortium Derivation Department 1930: Disease-Microbe Consortium Derivation Department

Claims

1. 1. A method executed by a computing device for deriving disease-associated microbial community data utilizing a machine learning model, comprising: (1) acquiring gut microbiota data and deriving candidate microbial community data from the gut microbiota data; (2) deriving disease-related microbial community data from the candidate microbial community data; The step (1) includes: Calculating the taxonomic similarity between bacteria in the gut microbiota data; forming initial community data by clustering bacteria in the intestinal microbiota data based on the similarity; and deriving the candidate microbial community data from the initial community data through quality control; The step (2) is Dividing the candidate microbial community data into a training set and an experimental set; training a machine learning model using the training corpus; Selecting an algorithm that shows the highest median predictive performance using the experimental group, and selecting a machine learning model that shows the highest predictive performance among a plurality of machine learning models that use the algorithm; and deriving the disease-related microbial community data using the selected machine learning model.

2. The method of claim 1 , wherein the clustering is done using an unsupervised learning algorithm.

3. 3. The method of claim 2, wherein the unsupervised learning algorithm is one or more of hierarchical clustering, K-means clustering, and Gaussian mixture model.

4. The method of claim 1, wherein the quality control removes data from the initial community data that contains one, two, or more than half of the total number of microbial taxa in the acquired gut microbiota data.

5. 2. The method of claim 1, wherein the training group and the experimental group are partitioned using Monte Carlo Random Sampling.

6. The method of claim 1 , wherein the algorithm is a supervised learning algorithm.

7. 7. The method of claim 6, wherein the supervised learning algorithm is one or more of logistic regression, naive Bayes, random forest, and support vector machines (SVM).

8. The method of claim 1, wherein the disease-related microbial community data is selected as the community data that shows the highest absolute value in feature importance among the candidate microbial community data used in the selected machine learning model.

9. 2. The method according to claim 1, wherein the disease is obesity, CDI (Clostridioides difficile infection), or RA (Rheumatoid arthritis).

10. A computing device that utilizes a machine learning model to derive disease-related microbial community data, comprising: an acquisition unit that acquires gut microbiota data; a candidate consortium derivation unit that derives candidate microbial community data from the intestinal microbiota data; a disease-microbial consortium derivation unit that derives disease-related microbial community data from the candidate microbial community data; The candidate group consortium derivation unit Calculating the taxonomic similarity between bacteria in the intestinal microbiota data; forming initial community data by clustering the bacteria in the intestinal microbiota data based on the similarity; Deriving the candidate microbial community data from the initial community data through quality control; The disease-microorganism consortium derivation unit The candidate microbial community data is divided into a training set and an experimental set; training a machine learning model using the training swarm and an algorithm; Using the experimental group, an algorithm showing the highest median predictive performance is selected, and a machine learning model showing the highest predictive performance is selected from a plurality of machine learning models using the algorithm; A computing device that uses the selected machine learning model to derive the disease-related microbial community data.

11. The computing device of claim 10 , wherein the clustering is done using an unsupervised learning algorithm.

12. 12. The computing device of claim 11, wherein the unsupervised learning algorithm is one or more selected from the group consisting of hierarchical clustering, K-means clustering, and Gaussian mixture model.

13. The computing device of claim 10, wherein the quality control removes data from the initial community data that contains one, two, or more than half of the total number of microbial taxa in the acquired gut microbiota data.

14. The computing device of claim 10 , wherein the training group and the experimental group are partitioned using Monte Carlo Random Sampling.

15. The computing device of claim 10 , wherein the algorithm is a supervised learning algorithm.

16. 16. The computing device of claim 15, wherein the supervised learning algorithm is one or more selected from the group consisting of logistic regression, naive Bayes, random forest, and support vector machines (SVM).

17. The computing device of claim 10, wherein the disease-related microbial community data is selected from the candidate microbial community data used in the selected machine learning model, the community data showing the highest absolute value in feature importance.

18. The computing device of claim 10 , wherein the disease is obesity or Clostridioides difficile infection (CDI) or Rheumatoid arthritis (RA).

Citation Information

Patent Citations

  • Marker for early intestinal cancer screening and adenoma diagnosis and applications thereof

    CN109852714A

  • Method for constructing machine learning model for identifying infectious diseases and non-infectious diseases

    CN114854847A

  • Method for identifying immune-related microorganisms in colorectal cancer based on multiple omics characteristics

    CN115472214A

  • Method and system for microbiome analysis

    EP3097211A2

  • Methods and systems for microbiome characterization, monitoring, and treatment

    JP2016525355A