Methods, compositions, and kits useful for inflammatory bowel disease (IBD) and IBD subtypes

The method of determining bacterial species in stool samples using machine learning models offers a non-invasive approach for diagnosing and treating IBD, addressing the invasiveness of current methods and improving disease management by distinguishing between CD and UC.

WO2025185637A1PCT designated stage Publication Date: 2025-09-11MICROBIOTA I CENT MAGIC LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/080596
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-03
Filing Date
2025-03-05
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Current methods for diagnosing inflammatory bowel diseases (IBD) such as Crohn's disease (CD) and ulcerative colitis (UC) are invasive and lack a single reference standard, necessitating the development of non-invasive and accurate methods for risk assessment, diagnosis, and treatment.

Method used

A method involving the determination of specific bacterial species in stool samples, using machine learning models to generate risk scores based on relative abundances, and optionally treating individuals with increased risk or diagnosis of IBD.

Benefits of technology

Provides a cost-effective, non-invasive method for assessing IBD risk and subtypes, enabling targeted treatment and improving disease management by differentiating between CD and UC for appropriate medical interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025080596_12092025_PF_FP_ABST
    Figure CN2025080596_12092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are methods for determining the risk of, diagnosing, preventing, or treating Crohn's Disease (CD), ulcerative colitis (UC) and / or inflammatory bowel disease (IBD) in an individual. Some embodiments provide kits and computer program products for determining the risk of or diagnosing Crohn's Disease (CD), ulcerative colitis (UC) and / or inflammatory bowel disease (IBD) in an individual. Other example embodiments are described herein. In certain embodiments, the disclosed methods and kits are accurate, cost-effective and non-invasive.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS, COMPOSITIONS, AND KITS USEFUL FOR INFLAMMATORY BOWEL DISEASE (IBD) AND IBD SUBTYPESCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to, and the benefit of, U.S. Provisional Application No. 63 / 562,232 filed on Mar 6, 2024, U.S. Provisional Application No. 63 / 562,233 filed on Mar 6, 2024, U.S. Provisional Application No. 63 / 675,266 filed on Jul 25, 2024 and U.S. Provisional Application No. 63 / 689,864 filed on Sep 3, 2024. The entire contents of the foregoing applications are hereby incorporated by reference in its entirety for all purposes.REFERENCE TO SEQUENCE LISTING

[0002] This application contains a sequence listing which has been submitted electronically in ST. 26 (xml) format and is hereby incorporated by reference in its entirety. Said ST. 26 copy, created on Mar 4, 2025, is named “M057001PCTCN. xml” and is 182 kilobytes in size.FIELD OF INVENTION

[0003] This application relates to methods and kits for determining the risk of, diagnosing, preventing or treating Crohn’s Disease (CD) , ulcerative colitis (UC) and / or inflammatory bowel disease (IBD) in an individual, and compositions for maintaining health, preventing or treating Crohn’s Disease (CD) , ulcerative colitis (UC) and / or inflammatory bowel disease (IBD) in an individual.BACKGROUND OF INVENTION

[0004] Inflammatory bowel disease (IBD) , which includes Crohn’s disease (CD) and ulcerative colitis (UC) , is a chronic and relapsing inflammatory disorder of the gastrointestinal (GI) tract. Globally, over seven million people are estimated to be living with IBD. Although IBD used to be more prevalent in Western countries, newly industrialized countries have experienced a rise in disease incidence in the past few decades largely attributed to the influence of lifestyle and environmental factors. Delayed diagnosis is often associated with disease progression and intestinal surgery, whereas early diagnosis followed by timely intervention can lead to improved outcomes.

[0005] Ulcerative colitis (UC) is a chronic intestinal inflammatory disease that mainly occurs in the rectum and colon. Crohn’s disease (CD) is another subtype of inflammatory bowel disease that can result in progressive bowel damage and disability. The precise etiology of CD remains unknown. Currently, there is no single reference standard for the diagnosis of CD. Conventionally, diagnosis of CD and UC requires a combination of invasive procedures such as colonoscopy, oesophago-gastro-duodenoscopy (OGD) , computed tomography (CT) scan. Accordingly, there is a need for new / alternative methods for determining the risk of, diagnosing, preventing or treating CD, UC and / or IBD.SUMMARY OF INVENTION

[0006] In the light of the foregoing background, in certain embodiments, it is an object to provide novel kits, methods and uses that are useful for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) , ulcerative colitis (UC) and / or inflammatory bowel disease (IBD) .

[0007] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating Inflammatory Bowel Disease (IBD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from IBD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from IBD, optionally treating the individual;wherein the IBD is an active or inactive IBD.

[0008] In some embodiments, a method for determining the risk of, diagnosing, preventing, or treating inflammatory bowel disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:(a) determining the relative abundance of a first set of bacterial species, optionally the level of fecal calprotectin, and optionally the relative abundance of a second set of bacterial species in a stool sample from the individual, wherein the first set of bacterial species and the second set of bacterial species each independently comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;(b) comparing the relative abundance of each bacterial species in the first set of bacterial species from the individual with the relative abundance of each bacterial species in the first set of bacterial species in a first reference data set and optionally comparing the level of fecal calprotectin from the individual with the level of fecal calprotectin in the first reference data set using a first machine learning model to generate a first risk score;(c) determining the individual as having increased risk for or is suffering from IBD if the first risk score is higher than a first cutoff value;(d) if the individual is determined to have an increased risk for or is suffering from IBD, optionally comparing the relative abundance of each bacterial species in the second set of bacterial species from the individual with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score;(e) optionally determining the individual as having increased risk for or is suffering from CD if the second risk score is higher than a second cutoff value, or determining the individual as having increased risk for or is suffering from UC if the risk score is lower than or equal to the second cutoff value; and(f) if the individual is determined to have an increased risk for or is suffering from IBD in step (c) , or if steps (d) and (e) are performed and if the individual is determined to have an increased risk for or is suffering from UC or CD, optionally treating the individual.

[0009] In some embodiments, provided is a method for diagnosing, preventing, or treating an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:a) determining if the individual has increased risk for or is suffering from IBD;b) if the individual is determined to have an increased risk for or is suffering from IBD, determining the relative abundance of a second set of bacterial species in a stool sample from the individual, wherein the second set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;c) comparing the relative abundance of each bacterial species in the second set of bacterial species from the individual with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score;d) determining the individual as having increased risk for or is suffering from CD if the second risk score is higher than a second cutoff value, or determining the individual as having increased risk for or is suffering from UC if the risk score is lower than or equal to the second cutoff value; ande) if the individual is determined to have an increased risk for or is suffering from UC or CD, optionally treating the individual.

[0010] In some embodiments, provided is a method for identifying a stool sample with an altered Inflammatory Bowel Disease’s (IBD) microbiome and / or an altered IBD subtype microbiome, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:a) determining the relative abundance of a first set of bacterial species, optionally the level of fecal calprotectin, and optionally the relative abundance of a second set of bacterial species in the stool sample, wherein the first set of bacterial species and the second set of bacterial species each independently comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the first set of bacterial species from the stool sample with the relative abundance of each bacterial species in the first set of bacterial species in a first reference data set and optionally comparing the level of fecal calprotectin from the stool sample with the level of fecal calprotectin in a the first reference data set using a first machine learning model to generate a risk score; andc) determining the stool sample as having an altered IBD microbiome if the first risk score is higher than a first cutoff value;d) if the stool sample is determined to have the altered IBD microbiome, optionally comparing the relative abundance of each bacterial species in the second set of bacterial species from the stool sample with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score; ande) optionally determining the stool sample as having an altered CD microbiome if the second risk score is higher than a second cutoff value, or determining the stool sample as having an altered UC microbiome if the risk score is lower than or equal to the second cutoff value.

[0011] In some embodiments, provided is a method for identifying a stool sample with an altered IBD subtype microbiome, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:a) determining if the stool sample has an altered IBD microbiome;b) if the stool sample is determined to have an altered IBD microbiome, determining the relative abundance of a second set of bacterial species in a stool sample from the individual, wherein the second set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;c) comparing the relative abundance of each bacterial species in the second set of bacterial species from the stool sample with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score; andd) determining the stool sample as having an altered CD microbiome if the second risk score is higher than a second cutoff value, or determining the stool sample as having an altered UC microbiome if the risk score is lower than or equal to the second cutoff value.

[0012] In some embodiments, provided is a kit for determining the risk of or diagnosing Inflammatory Bowel Disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said kit comprising a reagent for detecting a first set of bacterial species and optionally a reagent for detecting a second set of bacterial species, wherein the first set of bacterial species and the second set of bacterial species each independent comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0013] In some embodiments, provided is a computer program product for determining the risk of or diagnosing Inflammatory Bowel Disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of embodiments 2-47.

[0014] In some embodiments, provided is a computer program product for determining the risk of, diagnosing, or suggesting treatment for Inflammatory Bowel Disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of embodiments 2-47, and optionally further configured to enable the execution of the following step:if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, suggesting an appropriate IBD treatment, UC treatment or CD treatment, respectively to the individual.

[0015] In some embodimets, provided is a computer program product for determining the risk of or diagnosing an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of embodiments 48-78.

[0016] In some embodiments, provided is a computer program product for determining the risk of, diagnosing, or suggesting treatment for an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of embodiments 48-78, and optionally further configured to enable the execution of the following step:if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, suggesting an appropriate IBD treatment, UC treatment or CD treatment, respectively to the individual.

[0017] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from CD, optionally treating the individual.

[0018] In some embodiments, provided is a method for identifying a stool sample with an altered Crohn’s Disease microbiome comprising the steps of:a) determining the relative abundance of a set of bacterial species in the stool sample, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the stool sample as having an altered CD microbiome if the risk score is higher than a cutoff value.

[0019] In some embodiments, provided is a method of increasing amino acid biosynthesis in a cell, wherein the amino acid is L-arginine, L-ornithine, or L-valine, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.

[0020] In some embodiments, provided is a method of increasing carbohydrate degradation in a cell, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.

[0021] In some embodiments, provided is a method of increasing cofactor, carrier, and vitamin biosynthesis in a cell, wherein the vitamin is thiamine phosphate, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.

[0022] In some embodiments, provided is a method of increasing amino acid biosynthesis in a cell, wherein the amino acid is L-tryptophan, comprising providing an effective amount of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, or Roseburia intestinalis to the cell.

[0023] In some embodiments, provided is a composition comprising Lawsonibacter asaccharolyticus and Eubacterium sp. CAG: 274.

[0024] In some embodiments, provided is a composition comprising Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, and Roseburia inulinivorans.

[0025] In some embodiments, provided is a kit for determining the risk of or diagnosing Crohn’s Disease (CD) in an individual, comprising a reagent for detecting a set of bacterial species, wherein the set of bacterial species is selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3.

[0026] In some embodiments, provided is a computer program product for determining the risk of or diagnosing Crohn’s Disease (CD) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value.

[0027] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Gemmiger formicilis, Eubacterium hallii, Blautia obeum, Roseburia inulinivorans, Fusicatenibacter saccharivorans, Eubacterium rectale, Roseburia faecis, Asaccharobacter celatus, Collinsella aerofaciens, Faecalibacterium prausnitzii, Anaerostipes hadrus, Lachnospira pectinoschiza, Ruminococcus torques, Clostridium leptum, Parabacteroides merdae, Ruminococcus bromii, Roseburia intestinalis, Adlercreutzia equolifaciens, Alistipes putredinis, Eubacterium sp. CAG: 38, Roseburia hominis, Agathobaculum butyriciproducens, Dorea formicigenerans, Bacteroides stercoris, Odoribacter splanchnicus, Oscillibacter sp. 57_20, Coprococcus catus, Bacteroides vulgatus, Dorea longicatena, Alistipes shahii, Lawsonibacter asaccharolyticus, Firmicutes bacterium CAG: 83, Oscillibacter sp. CAG: 241, Eubacterium ramulus, Alistipes finegoldii, Bacteroides caccae, Akkermansia muciniphila, Coprococcus comes, Butyricimonas virosa, Alistipes indistinctus, Bacteroides uniformis, Bacteroides coprocola, Eubacterium sp. CAG: 274, Desulfovibrio piger, Bacteroides massiliensis, Clostridium sp. CAG: 299, Eubacterium eligens, Romboutsia ilealis, Eubacterium ventriosum, Anaerostipes hadrus, Lactobacillus rogosae, Ruminococcus obeum, Gemella haemolysans, Streptococcus mitis, Actinomyces odontolyticus, Actinomyces sp. S6-Spd3, Actinomyces graevenitzii, Lactobacillus mucosae, Streptococcus parasanguinis, Rothia mucilaginosa, Actinomyces sp. oral taxon 181, Veillonella parvula, Escherichia coli, Bacteroides fragilis, and Eubacterium sulci;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from CD, optionally treating the individual.

[0028] In some embodiments, provided is a method for monitoring disease activity of Crohn’s Disease (CD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, and Escherichia coli;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased disease activity if the risk score is higher than a cutoff value.

[0029] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from UC, optionally treating the individual.

[0030] In some embodiments, provided is a method for identifying a stool sample with an altered ulcerative colitis microbiome comprising the steps of:a) determining the relative abundance of a set of bacterial species in the stool sample, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the stool sample as having an altered UC microbiome if the risk score is higher than a cutoff value.

[0031] In some embodiments, provided is a method of increasing nucleoside and nucleotide biosynthesis in a cell, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, and Ruminococcus torques to the cell.

[0032] In some embodiments, provided is a composition comprising Fusicatenibacter saccharivorans, Clostridium leptum, and Gemmiger formicilis.

[0033] In some embodiments, provided is a kit for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, comprising a reagent for detecting a set of bacterial species, wherein the set of bacterial species is selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0034] In some embodiments, provided is a computer program product for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value.

[0035] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Collinsella aerofaciens, Clostridium leptum, Ruminococcus torques, Asaccharobacter celatus, Gemmiger formicilis, Fusicatenibacter saccharivorans, Alistipes putredinis, Dorea longicatena, Ruminococcus bromii, Odoribacter splanchnicus, Lachnospira pectinoschiza, Coprococcus comes, Adlercreutzia equolifaciens, Roseburia inulinivorans, Blautia obeum, Eubacterium rectale, Dorea formicigenerans, Alistipes shahii, Bacteroides stercoris, Parabacteroides merdae, Oscillibacter sp. CAG: 241, Akkermansia muciniphila, Alistipes finegoldii, Eubacterium hallii, Oscillibacter sp. 57_20, Roseburia hominis, Phascolarctobacterium faecium, Collinsella stercoris, Bifidobacterium adolescentis, Roseburia intestinalis, Butyricimonas virosa, Alistipes indistinctus, Eubacterium sp. CAG: 38, Anaerostipes hadrus, Bacteroides caccae, Clostridium sp. CAG: 242, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Coprococcus catus, Bacteroides massiliensis, Roseburia faecis, Clostridium sp. CAG: 58, Ruthenibacterium lactatiformans, Bacteroides plebeius, Bilophila wadsworthia, Firmicutes bacterium CAG: 145, Eubacterium ramulus, Enterorhabdus caecimuris, Gemella morbillorum, Veillonella infantium, Haemophilus sp. HMSC71H05, Actinomyces sp. oral taxon 181, Blautia producta, Lactobacillus mucosae, Enterococcus avium, Veillonella atypica, Eubacterium sulci, Blautia hansenii, Rothia mucilaginosa, Clostridium spiroforme, Tyzzerella nexilis, Veillonella parvula, and Bacteroides fragilis;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from UC, optionally treating the individual.

[0036] In some embodiments, provided is a method for monitoring disease activity of ulcerative colitis (UC) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Odoribacter splanchnicus, Gemmiger formicili, Actinomyces sp. oral taxon 181 and Clostridium spiroforme;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the individual as having increased disease activity if the risk score is higher than a cutoff value.

[0037] In some embodiments, provided is a computer program product for determining the risk of, diagnosing, or suggesting treatment for Crohn’s Disease (CD) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from CD, optionally suggesting an appropriate CD treatment to the individual.

[0038] In some embodiments, provided is a computer program product for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; andif the individual is determined to have an increased risk for or is suffering from UC, optionally suggesting an appropriate UC treatment to the individual.

[0039] In some embodiments, provided is a computer readable medium including instructions which, when executed by at least one computer, causes the at least one computer to carry out the steps of the method according to any of the examples as described herein.

[0040] In some embodiments, provided is a computer system including at least one memory, at least one processor and a computer program stored in the at least one memory, wherein the at least one processor executes the computer program to carry out the steps of the method according to any of the examples as described herein.

[0041] Other example embodiments will be described below.Advantages

[0042] There are many advantages of the present disclosure.

[0043] In some embodiments, the present disclosure provides novel methods and kits for assessing risk and diagnosing IBD in a subject, and / or to identify / diagnose the IBD subtype (CD or UC) of the subject by using one or more bacterial markers. By determining the relative abundance of these bacterial markers in fecal samples and generating a risk or diagnostics score, the disclosed techniques can provide a cost-effective, non-invasive method to support clinical decision making, and hence to help improve disease management. For example, the choice of medical treatment differs according to disease subtype. UC and CD have different pathogenesis and differential responses to drug treatment. Medical therapy is often different, depending on disease severity, and / or whether a step-up approach is used. For example, in mild to moderate UC, 5-aminosalicylates (5-ASA) is the recommended treatment. Biologics are commonly used for the treatment of moderate to severe CD and UC but certain drugs are more targeted for one disease. For example, tofacitinib is approved by the FDA for UC but not CD. Mirikizumab is a drug approved for the treatment of UC, and Guselkumab is approved for UC and is still under review for CD. Natalizumab is approved for CD but not UC. Hence, it is important to determine diagnosis of UC or CD for more appropriate drug use which is influenced by reimbursement of drugs for specific disease subtype.

[0044] In some embodiments, the provided methods and kits are useful in the diagnosis of IBD subtypes, which is crucial for long-term management. For example, UC and CD patients may require different long-term management strategies. UC management focuses on maintaining remission with medical therapy and regular monitoring to prevent flare-ups and complications like colorectal cancer. For CD, long-term management may involve a more aggressive approach to prevent complications due to the risk of strictures and penetrating disease and patients require regular imaging and endoscopic evaluations. In addition, in subjects with severe or refractory disease whereby surgical therapies are needed, definitive and accurate diagnosis is critical in determining the type of surgical procedures which differs for CD and UC as the former is not curative with surgery. In some embodiments, stratification for further workup differs with CD and UC. CD, being a more complex condition, can involve the upper gastrointestinal tract, as well as lead to strictures and fistulas. In some embodiments, this necessitates a greater number of diagnostic procedures compared to UC, including computed tomography, ultrasonography, magnetic resonance imaging, and upper gastrointestinal endoscopy to locate and assess affected areas in the small intestine.

[0045] In some embodiments, the novel methods and kits are useful to hospitals, institutions or individuals who are willing to do fecal / intestinal microbiota transplantation (FMT or IMT) treatment for IBD, CD or UC.

[0046] In some embodiments, the present disclosure provides compositions comprising a set of underrepresented bacteria in CD (also referred to as “CD depleted species” in some embodiments) that can be used to treat CD patients. After an individual is being determined as having increased risk for or is suffering from CD, the individual can undergo treatment through IMT or supplementation of the disclosed compositions in the form of probiotics / microbial consortium pills.

[0047] In some embodiments, the present disclosure provides compositions comprising a set of underrepresented bacteria in UC (also referred to as “UC depleted species” in some embodiments) that can be used to treat UC patients. After an individual is being determined as having increased risk for or is suffering from UC, the individual can undergo treatment through IMT or supplementation of the disclosed compositions in the form of probiotics / microbial consortium pills.BRIEF DESCRIPTION OF FIGURES

[0048] FIG. 1A is a diagram showing the top bacterial species associated with Crohn’s disease (CD) identified in the study, according to an example embodiment.

[0049] FIG. 1B is a plot showing the associations between CD, gender, age and the relative abundance of the 18 selected CD-depleted and CD-enriched bacterial species, according to an example embodiment.

[0050] [Corrected under Rule 26, 14.03.2025]FIGs. 2A-2C are plots showing the relative abundance of 18 bacterial species markers determined by metagenomic sequencing in CD patients and normal individuals in Hong Kong, China (HK, China) discovery cohort (FIG. 2A) , Hong Kong, China (HK, China) validation cohort (FIG. 2B) and Australia (AUS) validation cohort (FIG. 2C) , respectively, according to an example embodiment.

[0051] [Corrected under Rule 26, 14.03.2025]FIG. 3A is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the machine learning model (random forest model) using the nine selected bacterial species markers determined by metagenomic sequencing in the training set, test set, the Hong Kong, China (HK, China) validation cohort and Australia (AUS) validation cohort, according to an example embodiment.

[0052] [Corrected under Rule 26, 14.03.2025]FIG. 3B is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the machine learning model (random forest model) using the nine selected bacterial species markers determined by metagenomic sequencing in three public datasets including subjects from the United States, Netherlands and China, according to an example embodiment.

[0053] FIGs. 4A-4C are plots showing the relative abundance (normalized abundance) of the nine selected bacterial species markers determined by droplet digital PCR in CD patients and normal individuals (controls) in a subgroup of discovery cohort, according to an example embodiment.

[0054] FIG. 5 is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the random forest model using the nine selected bacterial species markers determined by droplet digital PCR from the training set and test set, according to an example embodiment.

[0055] FIG. 6A is a stacked bar plot comparing the relative abundance (%) of bacterial species in starch degradation between healthy controls and CD patients, according to an example embodiment.

[0056] FIG. 6B is a stacked bar plot comparing the relative abundance (%) of bacterial species in L-tryptophan biosynthesis between healthy controls and CD patients, according to an example embodiment.

[0057] FIG. 6C is a stacked bar plot comparing the relative abundance (%) of bacterial species in thiamine phosphate formation from pyrithiamine and oxythiamine (yeast) between healthy controls and CD patients, according to an example embodiment.

[0058] FIG. 6D is a stacked bar plot comparing the relative abundance (%) of bacterial species in the superpathway of L-lysine, L-theonine and L-methionine biosynthesis II, between healthy controls and CD patients, according to an example embodiment.

[0059] FIG. 7 is a plot showing the correlation between functional dysbiosis scores and probability of disease generated by the machine learning model based on the nine selected bacterial species biomarkers, according to an example embodiment.

[0060] FIGs. 8A-8C are plots showing the relative abundance of the nine selected bacterial species in healthy controls and patients at inactive and active status, according to an example embodiment.

[0061] FIG. 8D is a plot showing the comparison of the probability of disease calculated by the random forest model between CD patients at inactive (N = 69) and active (N = 9) status, according to an example embodiment.

[0062] FIG. 8E is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the random forest model using the nine selected bacterial species markers determined by metagenomics in CD patients in remission stage and healthy controls, according to an example embodiment.

[0063] FIGs. 9A-9C are plots showing the performance of the CD diagnostic model in differentiating CD patients from other subjects with non-IBD GI diseases, according to an example embodiment.

[0064] FIGs. 10A-10B are plots showing the performance of the CD diagnostic model in differentiating CD from other non-IBD, non-GI diseases, according to an example embodiment.

[0065] FIG. 10C is a plot showing the ROC curve of classifying CD patients from patients with all other non-IBD (GI and non-GI) diseases using the nine selected CD biomarkers, according to an example embodiment.

[0066] FIG. 11 is a plot showing the ROC curve of classifying CD patients from UC patients using the nine selected CD biomarkers, according to an example embodiment.

[0067] FIG. 12 shows an example method for determining the risk of, diagnosing, preventing, or treating Crohn's Disease (CD) in an individual according to an example embodiment.

[0068] FIG. 13A is a diagram showing the top bacterial species associated with ulcerative colitis (UC) identified in the study, according to an example embodiment.

[0069] FIG. 13B is a plot showing the associations between UC, gender, age and the relative abundance of the 18 selected UC-depleted and UC-enriched bacterial species calculated by general linear model MaAsLin2, according to an example embodiment.

[0070] [Corrected under Rule 26, 14.03.2025]FIGs. 14A-14B are plots showing the relative abundance of 18 bacterial species markers determined by metagenomic sequencing in UC patients and normal individuals in Hong Kong, China (HK, China) discovery cohort (FIG. 14A) and Hong Kong, China (HK, China) validation cohort (FIG. 14B) , respectively, according to an example embodiment.

[0071] [Corrected under Rule 26, 14.03.2025]FIG. 15A is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the machine learning model (random forest model) using 10 selected bacterial species markers determined by metagenomic sequencing in the training set, test set, and the Hong Kong, China (HK, China) validation cohort according to an example embodiment.

[0072] [Corrected under Rule 26, 14.03.2025]FIG. 15B is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the machine learning model (random forest model) using 10 selected bacterial species markers determined by metagenomic sequencing in three public datasets including subjects from the United States, Netherlands and China according to an example embodiment.

[0073] FIGs. 16A-16C are plots showing the relative abundance (normalized abundance) of 10 selected bacterial species markers determined by droplet digital PCR in UC patients and normal individuals (controls) in a subgroup of discovery cohort according to an example embodiment.

[0074] FIG. 17 is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the random forest model using 10 selected markers determined by droplet digital PCR from the training set and test set according to an example embodiment.

[0075] FIG. 18A is a stacked bar plot comparing the relative abundance (%) of selected bacterial species in L-arginine biosynthesis I (via L-ornithine) between healthy controls and UC patients according to an example embodiment.

[0076] FIG. 18B is a stacked bar plot comparing the relative abundance (%) of selected bacterial species in L-valine biosynthesis between healthy controls and UC patients according to an example embodiment.

[0077] FIG. 18C is a stacked bar plot comparing the relative abundance (%) of selected bacterial species in L-ornithine biosynthesis I between healthy controls and UC patients according to an example embodiment.

[0078] FIG. 18D is a stacked bar plot comparing the relative abundance (%) of selected bacterial species in adenosine ribonucleotides de novo biosynthesis between healthy controls and UC patients according to an example embodiment.

[0079] FIG. 19 is a plot showing the correlation between functional dysbiosis scores and probability of disease generated by the machine learning model based on the 10 selected bacterial species biomarkers, according to an example embodiment.

[0080] FIGs. 20A-20C are plots showing the relative abundance of the 10 selected bacterial species in healthy controls and patients at inactive and active status, according to an example embodiment.

[0081] FIG. 20D is a plot showing the comparison of the probability of disease calculated by the random forest model between UC patients with inactive (N = 110) and active (N = 12) status, according to an example embodiment.

[0082] FIG. 20E is a plot showing the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the random forest model using the 10 selected bacterial species markers determined by metagenomics in UC patients in remission stage and healthy controls, according to an example embodiment.

[0083] FIGs. 21A-21C are plots showing the performance of the UC diagnostic model in differentiating UC patients from other subjects with non-IBD GI diseases, according to an example embodiment.

[0084] FIGs. 22A-22B are plots showing the performance of the UC diagnostic model in differentiating UC from other non-IBD, non-GI diseases, according to an example embodiment.

[0085] FIG. 22C is a plot showing the ROC curve of classifying UC patients from patients with all other non-IBD (GI and non-GI) diseases using the 10 selected UC biomarkers, according to an example embodiment.

[0086] FIG. 23 shows an example method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual according to an example embodiment.

[0087] FIG. 24A is a diagram showing an overview of the study workflow.

[0088] FIG. 24B are Violin plots showing the Shannon index and observed species of fecal microbiome in patients with UC (N=205) , CD (N=174) , and controls (N=118) .

[0089] FIG. 24C is a principal Coordinates Analysis (PCoA) plot showing the varied microbial composition among groups (174 CD patients, 205 UC patients and 118 controls) .

[0090] FIG. 24C’ is a chart of multivariate analysis showing the amount of explained variance and the respective P value determined by PERMANOVA based on Bray-Curtis dissimilarity at species level.

[0091] FIG. 24D is a stacked bar chart showing the relative abundance of the six most abundant phyla in patients with UC (N=205) , CD (N=174) , and controls (N=118) .

[0092] FIG. 24D’ is a comparison chart of the relative abundance of six phyla among patients with UC, CD, and healthy controls.

[0093] FIG. 25C is a chart showing the relative abundance of ten bacterial species biomarkers in UC (N=205) and control group (N=118) .

[0094] FIG. 25D is a chart showing the relative nine bacterial species biomarkers in CD (N=174) and control group (N=118) .

[0095] FIG. 25E is a chart showing the performance of model with ten UC bacterial species biomarkers for classifying UC.

[0096] FIG. 25F is a chart showing the performance of model with nine CD bacterial species biomarkers for classifying CD.

[0097] FIG. 25G is a chart showing the Shapley Additive Explanations (SHAP) values of the ten UC bacterial species biomarkers for each sample.

[0098] FIG. 25H is a chart showing the Shapley Additive Explanations (SHAP) values of the nine CD bacterial species biomarkers for each sample.

[0099] FIG. 25K is a chart showing the performance of model with different number of used features to discriminate patients with UC from healthy controls.

[0100] FIG. 25L is a chart showing the performance of model with different number of used features to discriminate patients with CD from healthy controls.

[0101] FIGs. 26A-26D are charts showing the differential functional pathways between UC / CD patients and healthy controls, and their correlation with bacterial species biomarkers.

[0102] FIGs. 26E is a chart showing the distribution of functional dysbiosis scores determined by median Bray-Curtis dissimilarity between a sample and controls.

[0103] FIGs. 26F is a chart showing the comparison of functional dysbiosis scores among UC (N = 205) , CD (N = 174) and controls (N = 118) .

[0104] FIG. 27A is a chart showing the relative abundance of ten UC bacterial species biomarkers in healthy controls and UC patients at inactive and active status.

[0105] FIG. 27B is a chart showing the comparison of the probability of disease calculated by the random forest model between UC patients at inactive (N = 110) and active (N = 11) status.

[0106] FIG. 27C is a chart showing the model performance in distinguishing inactive UC patients (N = 40) and healthy controls (N = 42) .

[0107] FIG. 27D is a chart showing the relative abundance of nine CD bacterial species biomarkers in healthy controls and CD patients at inactive and active status.

[0108] FIG. 27E is a chart showing the comparison of the probability of disease calculated by the random forest model between CD patients at inactive (N = 69) and active (N = 9) status.

[0109] FIG. 27F is a chart showing model performance in distinguishing inactive CD patients (N = 69) and healthy controls (N = 108) .

[0110] FIG. 28A is a chart showing the signature of bacterial species biomarkers for UC diagnosis in patients and healthy individuals of discovery cohort, validation cohort, and three downloaded public datasets.

[0111] FIG. 28B is a chart showing the signature of bacterial species biomarkers for CD diagnosis in patients and healthy individuals of discovery cohort, validation cohort, and three downloaded public datasets.

[0112] [Corrected under Rule 26, 14.03.2025]FIG. 28C are charts showing the probability of disease calculated by the random forest model between UC / CD patients and healthy controls in Hong Kong, China discovery cohort, validation cohort from Hong Kong, China and Australia, and public datasets (USA, Netherland, and China) .

[0113] FIG. 28D is a chart showing the correlation among the ten UC bacterial species markers.

[0114] FIG. 28E is a chart showing the correlation among the nine CD bacterial species markers.

[0115] [Corrected under Rule 26, 14.03.2025]FIG. 29A is a chart showing the performance of model with ten UC selected bacterial species biomarkers for classifying UC patients with controls in Hong Kong, China validation cohort.

[0116] [Corrected under Rule 26, 14.03.2025]FIGs. 29B-29C are charts showing the performance of model with nine CD selected bacterial species biomarkers for classifying CD patients with controls in Hong Kong, China and Australia validation cohorts.

[0117] FIG. 29D is a chart showing the performance of model with the selected bacterial species biomarkers for classification of UC patients with controls in the three downloaded public datasets.

[0118] FIG. 29E is a chart showing the performance of model with the selected bacterial species biomarkers for classification of CD patients with controls in the three downloaded public datasets.

[0119] FIGs. 29F-29G are charts showing the associations between disease group (UC or CD) , geography, ethnicity and the relative abundance of bacterial species biomarkers were calculated by MaAsLin2 in all IBD cohorts.

[0120] FIG. 29H is a chart showing the Performance of model with the selected bacterial species biomarkers for classifying UC patients (n=817) with controls (n=1746) in all UC validation cohorts.

[0121] FIG. 29I is a chart showing the performance of a model with the selected bacterial species biomarkers for classifying CD patients (n=1065) with controls (n=1873) in all CD validation cohorts.

[0122] FIG. 29J is a chart showing the model performance in distinguishing treated and  UC patients from controls in two downloaded public datasets.

[0123] FIG. 29K is a chart showing the model performance in distinguishing treated and  CD patients from controls in two downloaded public datasets.

[0124] FIG. 29L is a chart showing the model performance in distinguishing UC patients from controls compared with fecal calprotectin test in two downloaded public datasets.

[0125] FIG. 29M is a chart showing the model performance in distinguishing CD patients from controls compared with fecal calprotectin test in two downloaded public datasets.

[0126] FIG. 30A are charts showing the relative abundance of 10 UC bacterial species biomarkers in UC and other non-IBD disease group.

[0127] FIG. 30B are charts showing the relative abundance of 9 CD bacterial species biomarkers in CD and other non-IBD disease group.

[0128] FIG. 31A is a chart showing the composition of international multi-disease datasets from different countries and regions.

[0129] FIG. 31B is a chart showing the comparison of their probability of disease generated by UC model based on ten UC bacterial species biomarkers in controls (n=2391) , CVD (n = 143) , obesity (n=318) , CA (n=230) , CRC (n=372) , IBS-D (n=146) and UC patients (n = 817) .

[0130] FIG. 31C is a chart showing the performance of UC model in classifying UC patients (n = 817) from other non-IBD subjects (n = 3600) .

[0131] FIG. 31D is a chart showing the comparison of their probability of disease generated by CD model based on nine CD bacterial species biomarkers in controls (n=2391) , CVD (n = 143) , obesity (n=318) , CA (n=230) , CRC (n=372) , IBS-D (n=146) and CD patients (n = 1065) .

[0132] FIG. 31E is a chart showing the performance of CD model in classifying CD patients (n = 1065) from other non-IBD subjects (n = 3600) .

[0133] FIG. 32A is a chart showing the ROC of general IBD model in classifying IBD from controls and non-IBD in test set, IBD validation cohort, and non-IBD cohort.

[0134] FIG. 32B is a chart showing the prototypical standards for reporting diagnostic accuracy studies (STARD) diagram reporting the flow of participants in independent international IBD cohort (IBD=1882, Controls=2027) .

[0135] FIG. 32C is a chart showing the comparison of diagnostic performance of general IBD model and fecal calprotectin in classifying IBD from and IBS subjects.

[0136] FIG. 33A are charts showing the panel design of multiplex droplet digital PCR for UC and CD bacterial species markers.

[0137] FIG. 33B are charts showing the correlation between the abundance of the ten UC bacterial species biomarkers determined by metagenomics and multiplex droplet digital PCR method.

[0138] FIG. 33C are charts showing the correlation between the abundance of the nine CD bacterial species biomarkers determined by metagenomics and multiplex droplet digital PCR method.

[0139] FIG. 34A are charts showing the relative abundance of ten bacterial species biomarkers in UC and healthy control group in discovery cohort (205 UC; 84 controls) .

[0140] [Corrected under Rule 26, 14.03.2025]FIG. 34B is a chart showing the diagnostic performance to discriminate patients with UC from healthy control with ten bacterial species biomarkers determined by m-ddPCR in discovery cohort (testset, N=62) and Hong Kong, China cohort (N=108) .

[0141] FIG. 34C are charts showing the relative abundance of nine bacterial species biomarkers in CD and healthy control group in discovery cohort (172 CD; 86 controls) .

[0142] [Corrected under Rule 26, 14.03.2025]FIG. 34D is a chart showing the diagnostic performance to discriminate patients with CD from healthy control with nine bacterial species biomarkers determined by m-ddPCR in discovery cohort (testset, N=66) , Hong Kong, China cohort (N=153) and Australia cohort (N=177) .

[0143] [Corrected under Rule 26, 14.03.2025]FIG. 34E are charts showing the diagnostic performance of fecal calprotectin test and UC model with ten bacterial species biomarkers determined by m-ddPCR in Canada cohort (100 UC, 53 Controls) and Taiwan, China cohort (40 UC, 40 Controls) .

[0144] [Corrected under Rule 26, 14.03.2025]FIG. 34F are charts showing the diagnostic performance of fecal calprotectin test and CD model with ten bacterial species biomarkers determined by m-ddPCR in Canada cohort (100 CD, 53 Controls) and Taiwan, China cohort (40 CD, 40 Controls) .

[0145] [Corrected under Rule 26, 14.03.2025]FIG. 34G are charts showing the comparison of the probability of disease calculated by UC / CD model using m-ddPCR data and fecal calprotectin results between UC / CD patients at inactive and active status, and healthy controls in Canada and Taiwan, China cohort.

[0146] FIG. 34H is a chart showing the diagnostic performance of fecal calprotectin test and UC model with ten bacterial species biomarkers determined by m-ddPCR in distinguishing inactive UC patients (N=81) and healthy controls (N=93) .

[0147] FIG. 35A is a chart showing the difference of probability of disease (POD) calculated by metagenomics-based model and m-ddPCR-based model in UC.

[0148] FIG. 35B is a chart showing the difference of probability of disease (POD) calculated by metagenomics-based model and m-ddPCR-based model in CD.

[0149] FIG. 36 is a diagram showing the workflow of a test algorithm using IBD Model and IBD Subtype Model according to an example embodiment.DETAILED DESCRIPTION

[0150] As used herein and in the claims, the terms “comprising” (or any related form such as “comprise” and “comprises” ) , “including” (or any related forms such as “include” or “includes” ) , “containing” (or any related forms such as “contain” or “contains” ) , means including the following elements but not excluding others. It shall be understood that for every embodiment in which the term “comprising” (or any related form such as “comprise” and “comprises” ) , “including” (or any related forms such as “include” or “includes” ) , or “containing” (or any related forms such as “contain” or “contains” ) is used, this disclosure / application also includes alternate embodiments where the term “comprising” , “including, ” or “containing, ” is replaced with “consisting essentially of” or “consisting of” . These alternate embodiments that use “consisting of” or “consisting essentially of” are understood to be narrower embodiments of the “comprising” , “including, ” or “containing, ” embodiments.

[0151] For example, alternate embodiments of “a composition comprising A, B, and C” would be “a composition consisting of A, B, and C” and “a composition consisting essentially of A, B, and C. ” Even if the latter two embodiments are not explicitly written out, this disclosure / application includes those embodiments. Furthermore, it shall be understood that the scopes of the three embodiments listed above are different.

[0152] For the sake of clarity, “comprising” , including, and “containing” , and any related forms are open-ended terms which allows for additional elements or features beyond the named essential elements, whereas “consisting of” is a closed end term that is limited to the elements recited in the claim and excludes any element, step, or ingredient not specified in the claim.

[0153] For the sake of clarity, “characterized by” or “characterized in” (together with their related forms as described above) , does not limit or change the nature of whether the list of terms following it are open or closed. For example, in a claim directed towards “a composition comprising A, B, C, and characterized in D, E, and F” , the elements D, E, and F are still open-ended terms and the claim is meant to include other elements due to the use of the word “comprising” earlier in the claim.

[0154] “Consisting essentially of” limits the scope of a claim to the specified materials, components, or steps ( “essential elements” ) that do not materially affect the essential characteristic (s) of the claimed invention. In some embodiments, the essential characteristics are the basic and novel characteristic (s) of the claimed invention. For example, in some embodiments, the essential elements of a composition of the disclosure can be “Xmg to Ymg” of compound A. Even if the composition includes additional excipients, as long as the additional excipients do not materially affect the essential characteristics of the compound, e.g., in compound A’s ability to bind to XX target or to treat YY disease, then such embodiment that “consists essentially of compound A” still includes compositions with the aforementioned additional excipients.

[0155] As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Where a range is referred in the specification, the range is understood to include each discrete point within the range. For example, 1-7 means 1, 2, 3, 4, 5, 6, and 7.

[0156] As used herein and in the claims, an "effective amount" , is an amount that is effective to achieve at least a measurable amount of a desired effect. In some embodiments, the amount may be effective to increasing amino acid biosynthesis, carbohydrate degradation, cofactor, carrier, and / or vitamin biosynthesis in a cell. In some embodiments, providing an effective amount includes, but is not limited to, administering a therapeutic to a patient (e.g., orally, intravanously, subcutaneously, topically, via inhalation or intranasal, intramuscular) , or contacting a cell with the composition, either in vitro, ex vivo, or in vivo.

[0157] As used herein, the term “composition” refers to a formulation containing one or more active ingredient (s) . In some examples, a composition includes one or more bacterial species and optionally an acceptable carrier. In some examples, the composition is used as a medicament capable of treating CD, UC and / or IBD in a subject in need. In some embodiments, the acceptable carrier is a pharmaceutical acceptable carrier. In some embodiments, the carrier is a dietary formulation, food, or beverage. In some examples, the composition is used as a nutritional supplement, dietary formulation, functional food, medical food, food, or beverage for use in maintaining general health in a subject. In some examples, the composition is in the form of probiotics / microbial consortium pills.

[0158] As used herein and in the claims, the term “prevent” , “preventing” , “preventive” , “preventative” or “prevention” refers the methods of reducing the risk of the onset, relapse or spread of a disease or disorder or one or more of their symptoms.

[0159] As used herein, the term "treat, " "treating" or "treatment" refers to methods of alleviating, abating or ameliorating a disease or condition symptoms, preventing additional symptoms, ameliorating or preventing the underlying metabolic causes of symptoms, inhibiting the disease or condition, arresting the development of the disease or condition, relieving the disease or condition, causing regression of the disease or condition, relieving a condition caused by the disease or condition, or stopping the symptoms of the disease or condition either prophylactically and / or therapeutically.

[0160] As used herein, the term “sequencing” refers to a technology that includes obtaining sequence information from one or more nucleic acid molecules by determining the identity and / or order of at least some nucleotides within each nucleic acid molecule.

[0161] As used herein, the term “relative abundance” when used in the context of describing the presence of a particular bacterial species in relation to all bacterial species present in the same sample (e.g. in a stool sample of an individual) , refers to the relative amount of the bacterial species out of the amount of all bacterial species in that particular sample. For instance, the relative abundance of one particular bacterial species can be determined by comparing the quantity of DNA specific for this species (e.g., determined by quantitative polymerase chain reaction) in one given sample with the quantity of all bacterial DNA (e.g., determined by quantitative polymerase chain reaction (PCR) and / or sequencing based on the 16s rRNA sequence) in the same sample.

[0162] As used herein, the terms “amplification” or “amplify” as used herein includes methods for copying a target nucleic acid sequence, thereby increasing the number of copies of the target nucleic acid sequence. Amplification may be exponential or linear, for example, by polymerase chain reaction (PCR) , real-time PCR, digital droplet PCR, or other nucleic-acid amplification methods known in the art. The sequences amplified in this manner form an “amplification product. ”

[0163] As used herein, the term “set of primers” refers to a pair of primers that are both directed to a target nucleic acid sequence. In some embodiments, a set of primers that is directed to a particular target nucleic acid sequence contains a forward primer and a reverse primer, each of which hybridizes under a suitable condition to a different strand (sense or antisense) of the target nucleic acid sequence.

[0164] As used herein, the term “probe” refers to a nucleic acid molecule attached to at least one detectable label, wherein the nucleic acid molecule comprises a probe sequence capable of hybridizing to at least a portion of a target nucleic acid sequence. Example of detectable labels include, but are not limited to, radioactive isotopes, enzyme substrates, co-factors, ligands, chemiluminescent or fluorescent agents, haptens, and enzymes.

[0165] As used herein, the term “fluorophore-quencher probe” refers to a nucleic acid molecule labelled with a fluorophore and a quencher, wherein the nucleic acid molecule comprises a probe sequence capable of hybridizing to at least a portion of a target nucleic acid sequence. In some embodiments, the fluorophore is attached to the 5’-end of the nucleic acid molecule and the quencher is attached to the 3’-end of the nucleic acid molecule. During the PCR amplification process, if the fluorophore-quencher probe binds to the target nucleic acid sequence, the fluorophore is separated from the quencher, leading to the release of fluorescence that can be detected and quantified.

[0166] As used herein, the term “markers” , “bacterial markers” , “bacterial species markers” , “bacterial biomarkers” or “bacterial species biomarkers” refers to pre-selected bacterial species and / or genetic sequences to be used in the methods / kits as described herein. For example, the relative abundance of deoxyribonucleic acid (DNA) fragments can be measured for each bacterial biomarkers in a fecal sample of each of the subjects.

[0167] As used herein, the term “reference data set” refers to dataset generated from a cohort of subjects having a clear diagnosis of CD, UC and / or IBD and another cohort of subjects who have not been diagnosed with CD, UC and / or IBD and do not present any signs or symptoms of CD, UC and / or IBD or gastrointestinal illness. It can be generated by measuring the relative abundance for each of the “bacterial markers” for each subject in the cohort. The diagnosis of subjects in the cohort can be determined by a standard method. For example, a diagnosis can be made by combination of clinical, biochemical, stool, endoscopic, cross-sectional imaging, and histological investigations according to international guidelines. In order to properly establish a “reference cohort” , a sufficient number of individuals (e.g., at least 10, 12, 15, 20, 24 or more individuals) without CD, UC and / or IBD must be included in the control group to provide samples for determination of the average level (s) of one or more pre-selected bacterial species.

[0168] As used herein, the term “normal individuals” , “healthy controls” or “controls” refers to individuals who have no severe diseases or existing gut disorders (such as inflammatory bowel diseases, cancer, advanced adenoma, irritable bowel syndrome, or other GI symptoms) .

[0169] As used herein, the term “individuals without IBD” refers to individuals who have no IBD, but may have other diseases or conditions such as gastrointestinal diseases or conditions, irritable bowel syndrome, obesity, acid reflux, colon polyps, functional dyspepsia, gastrointestinal stromal tumor (GIST) , or other GI symptoms.

[0170] As used herein, the term “active IBD” refers to IBD patient with severe IBD-related symptoms, and the term “inactive IBD” refers to IBD patient with mild-to-moderate IBD-related symptoms. Likewise, the term “active UC” refers to UC patient with severe UC-related symptoms, and the term “inactive UC” refers to UC patient with mild-to-moderate UC-related symptoms; the term “active CD” refers to CD patient with severe CD-related symptoms, and the term “inactive CD” refers to CD patient with mild-to-moderate CD-related symptoms.

[0171] As used herein, the term “training data set” refers to a subset of the “reference data set” . It includes the known health classifications (CD, UC and / or IBD or control) for the subjects and the feature vectors for the subjects, which can be used to train a machine learning model.

[0172] As used herein, the term “tuning” refers to the process of adjusting the parameters or hyperparameters of a machine learning algorithm or model to optimize its performance on a given task or dataset. In some examples, during the tuning process, different values or combinations of values for the parameters are explored and evaluated to find the settings that yield the best results. The goal is to fine-tune the model's performance by finding the optimal configuration that minimizes errors, maximizes accuracy, or achieves the desired outcome.

[0173] As used herein, the terms “active” or “inactive” refers to the alternation of periods of disease activity of UC, CD and / or IBD. In some embodiments, UC is characterized by the alternation of periods of disease activity including flares (i.e., active UC) and remissions (i.e., inactive UC) . In some embodiments, CD is characterized by the alternation of periods of disease activity including flares (i.e., active CD) and remissions (i.e., inactive CD) . In some embodiments, the disease activities of UC patients were evaluated using the Mayo score, where patients were categorized to those with active disease and inactive disease; the disease activities of CD patients were evaluated using the Crohn’s Disease Activity Index or Harvey-Bradshaw Index, where patients were categorized to those with active disease and inactive disease. In some embodiments, the term “active IBD” refers to active UC or active CD. In some embodiments, the term “inactive IBD” refers to inactive UC or inactive CD.

[0174] When it is described herein that “if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, treating the individual with an appropriate IBD treatment, UC treatment or CD treatment, respectively” , the use of the term “respectively” means if the individual is determined to have an increased risk for or is suffering from one disease (e.g., UC) , treating the individual with an appropriate treatment for that disease (UC treatment) . In some embodiments, there may be an overlap in treatments that work for UC and another disease (such as CD, IBD, or others) . In some embodiments, the appropriated disease treatment is a treatment that has obtained regulatory approval for that disease. In some embodiments, the appropriate disease treatment is specific to that one disease (e.g., UC) and is not approved for other diseases (e.g., CD) . In some embodiments, the treatment is approved for both UC and another disease, such as CD.

[0175] When it is described herein that an individual is “without a XX treatment” (wherein XX is CD, UC or IBD) , it means the individual is currently not undergoing XX treatment, or the individual has not recently received XX treatment (for example, for 1-3x months before sample collection) .

[0176] When it is described herein that an individual is “XX  ” (wherein XX is CD, UC or IBD) , it means the individual has never received XX treatment before sample collection.

[0177] When it is described herein that an individual is “exposed to a XX treatment” (wherein XX is CD, UC or IBD) , it means the individual is currently undergoing XX treatment, the individual has recently received XX treatment (for example, for 1-3 months before sample collection) , or the individual has ever received XX treatment any time before sample collection.

[0178] Although the description referred to particular embodiments, the disclosure should not be construed as limited to the embodiments set forth herein.

[0179] Provided herein are examples that describe in more detail certain embodiments of the present disclosure. The examples provided herein are merely for illustrative purposes and are not meant to limit the scope of the invention in any way. All references given below and elsewhere in the present application are hereby included by reference.EMBODIMENT 1

[0180] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating Crohn's Disease (CD) in an individual, including the steps of: (a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3; (b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; (c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; and (d) if the individual is determined to have an increased risk for or is suffering from CD, optionally treating the individual.

[0181] In some embodiments, the set of bacterial species includes one or more bacterial species selected from the group consisting of Gemmiger formicilis, Eubacterium hallii, Blautia obeum, Roseburia inulinivorans, Fusicatenibacter saccharivorans, Eubacterium rectale, Roseburia faecis, Asaccharobacter celatus, Collinsella aerofaciens, Faecalibacterium prausnitzii, Anaerostipes hadrus, Lachnospira pectinoschiza, Ruminococcus torques, Clostridium leptum, Parabacteroides merdae, Ruminococcus bromii, Roseburia intestinalis, Adlercreutzia equolifaciens, Alistipes putredinis, Eubacterium sp. CAG: 38, Roseburia hominis, Agathobaculum butyriciproducens, Dorea formicigenerans, Bacteroides stercoris, Odoribacter splanchnicus, Oscillibacter sp. 57_20, Coprococcus catus, Bacteroides vulgatus, Dorea longicatena, Alistipes shahii, Lawsonibacter asaccharolyticus, Firmicutes bacterium CAG: 83, Oscillibacter sp. CAG: 241, Eubacterium ramulus, Alistipes finegoldii, Bacteroides caccae, Akkermansia muciniphila, Coprococcus comes, Butyricimonas virosa, Alistipes indistinctus, Bacteroides uniformis, Bacteroides coprocola, Eubacterium sp. CAG: 274, Desulfovibrio piger, Bacteroides massiliensis, Clostridium sp. CAG: 299, Eubacterium eligens, Romboutsia ilealis, Eubacterium ventriosum, Anaerostipes hadrus, Lactobacillus rogosae, Ruminococcus obeum, Gemella haemolysans, Streptococcus mitis, Actinomyces odontolyticus, Actinomyces sp. S6-Spd3, Actinomyces graevenitzii, Lactobacillus mucosae, Streptococcus parasanguinis, Rothia mucilaginosa, Actinomyces sp. oral taxon 181, Veillonella parvula, Escherichia coli, Bacteroides fragilis, and Eubacterium sulci.

[0182] In some embodiments, the reference data set is obtained by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects including patients with CD and normal individuals in a reference cohort.

[0183] In some embodiments, the comparing step in step (b) is performed using a machine learning model to generate the risk score.

[0184] In some embodiments, the machine learning model is constructed by (i) obtaining the reference data set by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects including patients with CD and normal individuals in a reference cohort; (ii) dividing the reference data set into a training data set and a test data set; (iii) dividing the training data set into a training subset and a validation subset; (iv) training a candidate model with the training subset and evaluating the performance of the candidate model using the validation subset; (v) optionally tuning the candidate model by adjusting hyperparameters in the candidate model; (vi) repeating step (iv) and optionally step (v) to produce a set of candidate models, and selecting an optimal candidate model among the set of candidate models; (vii) evaluating the performance of the optimal candidate model using the test data set; (viii) optionally training a final model with the reference data set using a set of hyperparameters corresponding to the optimal candidate model to produce the final model; and (ix) using the optimal candidate model of step (vii) or if step (viii) is used to train the final model, using the final model of step (viii) as the machine learning model for the comparing step in step (b) to generate the risk score. For sake of clarity, the reference data set used in step (viii) is the complete reference data set which includes both the training data set and the test data set.

[0185] In some embodiments, the machine learning model is a random forest model, and wherein the risk score in step (b) is generated by the following steps: (b1) generating an ensemble of decision trees by the random forest model using the relative abundance of the set of bacterial species in the reference data set; and (b2) running the relative abundance of the set of bacterial species from the individual along the ensemble of decision trees to generate the risk score.

[0186] In some embodiments, the optimal candidate model is the candidate model with the highest performance among the set of candidate models.

[0187] In some embodiments, the performance of the candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the candidate model. In some embodiments, the Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) is also referred to as the “AUROC” .

[0188] In some embodiments, the cutoff value is in the range of 0.4 to 0.7.

[0189] In some embodiments, the cutoff value is 0.5.

[0190] In some embodiments, the cutoff value is determined by the following steps: (c1) calculating a Youden's index using the risk score in the reference data set; and (c2) determining the cutoff value based on the Youden's index.

[0191] In some embodiments, the set of bacterial species includes one or more combinations of the bacterial species, wherein said combinations are selected from the group consisting of Table 1.4 and Table 1.6.

[0192] In some embodiments, the set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Lawsonibacter asaccharolyticus and Eubacterium sp. CAG: 274.

[0193] In some embodiments, the set of bacterial species includes Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans, and optionally Actinomyces sp. oral taxon 181 and / or Escherichia coli.

[0194] In some embodiments, the set of bacterial species includes Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, and Dorea formicigenerans.

[0195] In some embodiments, the set of bacterial species consists of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, and Dorea formicigenerans.

[0196] In some embodiments, the set of bacterial species includes Bacteroides fragilis, Escherichia coli, Roseburia inulinivorans and Ruminococcus obeum (B. obeum) , whereby the method identifies the individual's risk for CD over other IBD subtypes or non-IBD diseases.

[0197] In some embodiments, the other IBD subtypes comprise Ulcerative Colitis (UC) , and wherein the non-IBD diseases comprise colorectal cancer, colorectal adenomas, irritable bowel syndrome (diarrhea subtype) , obesity and cardiovascular disease.

[0198] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing or droplet digital PCR (ddPCR) .

[0199] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 1.4.

[0200] In some embodiments, the relative abundance of the set of bacterial species is determined by ddPCR, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 1.6.

[0201] In some embodiments, the ddPCR is performed by using one or more sets of primers selected from Table 1.1.

[0202] In some embodiments, if the individual is determined to have an increased risk for or is suffering from CD, treating the individual with an appropriate CD treatment, wherein the CD treatment is selected from the group consisting of a therapeutic treatment, fecal microbiota transplantation (FMT) , intestinal microbiota transplantation (IMT) , lifestyle and diet modifications, and surgical intervention. In some embodiments, the therapeutic treatment is selected from the group consisting of 5-Aminosalicylates (5-ASA, such as mesalazine, sulphasalazine) , steroids (such as budesonide, prednisolone) , thiopurines (such as azathioprine, mercaptopurine, methotrexate) , TNF inhibitors (such as infliximab, adalimumab, and certolizumab pegol) , anti-integrins (such as vedolizumab) , and anti-interleukin (IL) 12 / 23 (such as ustekinumab) .

[0203] In some embodiments, the 5-Aminosalicylates are mesalazine or sulphasalazine; the steroids are budesonide or prednisolone; the thiopurines are azathioprine, mercaptopurine or methotrexate; the TNF inhibitors are infliximab, adalimumab or certolizumab pegol; the anti-integrins are vedolizumab; and the anti-interleukin (IL) 12 / 23 is ustekinumab.

[0204] In some embodiments, the CD is in remission stage.

[0205] In some embodiments, provided is a method for identifying a stool sample with an altered Crohn's Disease microbiome including the steps of: (a) determining the relative abundance of a set of bacterial species in the stool sample, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3; (b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; and (c) determining the stool sample as having an altered CD microbiome if the risk score is higher than a cutoff value.

[0206] In some embodiments, the comparing step in step (b) is performed using a machine learning model to generate the risk score.

[0207] In some embodiments, the machine learning model is constructed by (i) obtaining the reference data set by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects including patients with CD and normal individuals in a reference cohort; (ii) dividing the reference data set into a training data set and a test data set; (iii) dividing the training data set into a training subset and a validation subset; (iv) training a candidate model with the training subset and evaluating the performance of the candidate model using the validation subset; (v) optionally tuning the candidate model by adjusting hyperparameters in the candidate model; (vi) repeating step (iv) and optionally step (v) to produce a set of candidate models, and selecting an optimal candidate model among the set of candidate models; (vii) evaluating the performance of the optimal candidate model using the test data set; (viii) optionally training a final model with the reference data set using a set of hyperparameters corresponding to the optimal candidate model to produce the final model; and (ix) using the optimal candidate model in step (vii) or if step (viii) is used to train the final model, using the final model in step (viii) as the machine learning model for the comparing step in step (b) to generate the risk score.

[0208] In some embodiments, the machine learning model is a random forest model, and wherein the risk score in step (b) is generated by the following steps: (b1) generating an ensemble of decision trees by the random forest model using the relative abundance of the set of bacterial species in the reference data set; and (b2) running the relative abundance of the set of bacterial species from the individual along the ensemble of decision trees to generate the risk score. In some embodiments, the risk score is generated by averaging prediction across the decision trees or by majority voting of class label predicted across the decision trees, wherein the proportion of votes in each class across the ensemble is the predicted probability (i.e. the risk score) .

[0209] In some embodiments, the optimal candidate model is the candidate model with the highest performance among the set of candidate models.

[0210] In some embodiments, the performance of the candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the candidate model.

[0211] In some embodiments, the cutoff value is in the range of 0.4 to 0.7.

[0212] In some embodiments, the cutoff value is 0.5.

[0213] In some embodiments, the cutoff value is determined by the following steps: (c1) calculating a Youden's index using the risk score in the reference data set; and (c2) determining the cutoff value based on the Youden's index.

[0214] In some embodiments, provided is a method, wherein the set of bacterial species includes one or more combinations of the bacterial species, wherein said combinations are selected from the group consisting of Table 1.4 and Table 1.6.

[0215] In some embodiments, the set of bacterial species includes Actinomyces sp. oral taxon 181, Lawsonibacter asaccharolyticus and Eubacterium sp. CAG: 274. In some embodiments, the set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Lawsonibacter asaccharolyticus and Eubacterium sp. CAG: 274.

[0216] In some embodiments, the set of bacterial species includes Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans, and optionally Actinomyces sp. oral taxon 181 and / or Escherichia coli. In some embodiments, the set of bacterial species consists of Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans. In some embodiments, the set of bacterial species consists of Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans, Actinomyces sp. oral taxon 181 and Escherichia coli. In some embodiments, the set of bacterial species consists of Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans, and a bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181 and Escherichia coli.

[0217] In some embodiments, the set of bacterial species includes Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, and Dorea formicigenerans.

[0218] In some embodiments, the set of bacterial species consists of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, and Dorea formicigenerans.

[0219] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing or droplet digital PCR (ddPCR) .

[0220] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 1.4.

[0221] In some embodiments, the relative abundance of the set of bacterial species is determined by ddPCR, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 1.6.

[0222] In some embodiments, the ddPCR is performed by using one or more sets of primers selected from Table 1.1.

[0223] In some embodiments, provided is a method of increasing amino acid biosynthesis in an individual, wherein the amino acid is L-arginine, L-ornithine, or L-valine, including providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the individual.

[0224] In some embodiments, provided is a method of increasing amino acid biosynthesis in a cell, wherein the amino acid is L-arginine, L-ornithine, or L-valine, including providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.

[0225] In some embodiments, provided is a method of increasing carbohydrate degradation in an individual, including providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the individual.

[0226] In some embodiments, provided is a method of increasing carbohydrate degradation in a cell, including providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.

[0227] In some embodiments, provided is a method of increasing cofactor, carrier, and vitamin biosynthesis in an individual, wherein the vitamin is thiamine phosphate, including providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the individual.

[0228] In some embodiments, provided is a method of increasing cofactor, carrier, and vitamin biosynthesis in a cell, wherein the vitamin is thiamine phosphate, including providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.

[0229] In some embodiments, provided is a method of increasing amino acid biosynthesis in an individual, wherein the amino acid is L-tryptophan, including providing an effective amount of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, or Roseburia intestinalis to the individual.

[0230] In some embodiments, provided is a method of increasing amino acid biosynthesis in a cell, wherein the amino acid is L-tryptophan, including providing an effective amount of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, or Roseburia intestinalis to the cell.

[0231] In some embodiments, the cell is part of a living organism.

[0232] In some embodiments, provided is a composition including Lawsonibacter asaccharolyticus and Eubacterium sp. CAG: 274.

[0233] In some embodiments, provided is a composition including Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, and Roseburia inulinivorans.

[0234] In some embodiments, provided is a composition as described herein further including one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia intestinalis, Eubacterium ventriosum, Anaerostipes hadrus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, and Romboutsia ilealis.

[0235] In some embodiments, provided is a dietary composition including the composition of any one of the foregoing aspects or embodiments, e.g., wherein the dietary composition is chosen from a medical food, a functional food, or a supplement. In some embodiments, provided is a nutritional supplement, dietary formulation, functional food, medical food, food, or beverage comprising a composition described herein for use in maintaining general health in a subject. Another embodiment provides a nutritional supplement, dietary formulation, functional food, medical food, food, or beverage comprising a composition described herein for use in the management of CD in a subject.

[0236] In some embodiments, provided is a kit for determining the risk of or diagnosing Crohn's Disease (CD) in an individual, including reagents for detecting a set of bacterial species, wherein the set of bacterial species is selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3.

[0237] In some embodiments, the reagents include one or more sets of primers, wherein each set of primers is configured to amplify a nucleic acid sequence belonging to a bacterial species in the set of bacterial species.

[0238] In some embodiments, the one or more sets of primers is selected from Table 1.1.

[0239] In some embodiments, the amplification is PCR.

[0240] In some embodiments, the PCR is ddPCR, and the reagents further comprise one or more probes selected from Table 1.1.

[0241] In some embodiments, the set of bacterial species includes one or more combinations of the bacterial species, wherein said combinations are selected from the group consisting of Table 1.6.

[0242] In some embodiments, the set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Lawsonibacter asaccharolyticus and Eubacterium sp. CAG: 274.

[0243] In some embodiments, the set of bacterial species includes Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans, and optionally Actinomyces sp. oral taxon 181 and / or Escherichia coli.

[0244] In some embodiments, the set of bacterial species includes Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, and Dorea formicigenerans.

[0245] In some embodiments, the set of bacterial species consists of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, and Dorea formicigenerans.

[0246] In some embodiments, provided is a computer program product for determining the risk of or diagnosing Crohn’s Disease (CD) in an individual, wherein the computer program product includes a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of: a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3; b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; and c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value.

[0247] In some embodiments, the reference data set is obtained by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects comprising patients with CD and normal individuals in a reference cohort.

[0248] In some embodiments, the comparing step in step (b) is performed using a machine learning model to generate the risk score.

[0249] In some embodiments, the machine learning model is constructed by i) obtaining the reference data set by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects comprising patients with CD and normal individuals in a reference cohort; ii) dividing the reference data set into a training data set and a test data set; iii) dividing the training data set into a training subset and a validation subset; iv) training a candidate model with the training subset and evaluating the performance of the candidate model using the validation subset; v) optionally tuning the candidate model by adjusting hyperparameters in the candidate model; vi) repeating step (iv) and optionally step (v) to produce a set of candidate models, and selecting an optimal candidate model among the set of candidate models; vii) evaluating the performance of the optimal candidate model using the test data set; viii) optionally training a final model with the reference data set using a set of hyperparameters corresponding to the optimal candidate model to produce the final model; and ix) using the optimal candidate model of step (vii) or if step (viii) is used to train the final model, using the final model of step (viii) as the machine learning model for the comparing step in step (b) to generate the risk score.

[0250] In some embodiments, the machine learning model is a random forest model, and wherein the risk score in step (b) is generated by the following steps: (b1) generating an ensemble of decision trees by the random forest model using the relative abundance of the set of bacterial species in the reference data set; and (b2) running the relative abundance of the set of bacterial species from the individual along the ensemble of decision trees to generate the risk score.

[0251] In some embodiments, the optimal candidate model is the candidate model with the highest performance among the set of candidate models.

[0252] In some embodiments, the performance of the candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the candidate model.

[0253] In some embodiments, the cutoff value is in the range of 0.4 to 0.7.

[0254] In some embodiments, the cutoff value is 0.5.

[0255] In some embodiments, the cutoff value is determined by the following steps: (c1) calculating a Youden’s index using the risk score in the reference data set; and (c2) determining the cutoff value based on the Youden’s index.

[0256] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) in an individual, comprising the steps of: a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Gemmiger formicilis, Eubacterium hallii, Blautia obeum, Roseburia inulinivorans, Fusicatenibacter saccharivorans, Eubacterium rectale, Roseburia faecis, Asaccharobacter celatus, Collinsella aerofaciens, Faecalibacterium prausnitzii, Anaerostipes hadrus, Lachnospira pectinoschiza, Ruminococcus torques, Clostridium leptum, Parabacteroides merdae, Ruminococcus bromii, Roseburia intestinalis, Adlercreutzia equolifaciens, Alistipes putredinis, Eubacterium sp. CAG: 38, Roseburia hominis, Agathobaculum butyriciproducens, Dorea formicigenerans, Bacteroides stercoris, Odoribacter splanchnicus, Oscillibacter sp. 57_20, Coprococcus catus, Bacteroides vulgatus, Dorea longicatena, Alistipes shahii, Lawsonibacter asaccharolyticus, Firmicutes bacterium CAG: 83, Oscillibacter sp. CAG: 241, Eubacterium ramulus, Alistipes finegoldii, Bacteroides caccae, Akkermansia muciniphila, Coprococcus comes, Butyricimonas virosa, Alistipes indistinctus, Bacteroides uniformis, Bacteroides coprocola, Eubacterium sp. CAG: 274, Desulfovibrio piger, Bacteroides massiliensis, Clostridium sp. CAG: 299, Eubacterium eligens, Romboutsia ilealis, Eubacterium ventriosum, Anaerostipes hadrus, Lactobacillus rogosae, Ruminococcus obeum, Gemella haemolysans, Streptococcus mitis, Actinomyces odontolyticus, Actinomyces sp. S6-Spd3, Actinomyces graevenitzii, Lactobacillus mucosae, Streptococcus parasanguinis, Rothia mucilaginosa, Actinomyces sp. oral taxon 181, Veillonella parvula, Escherichia coli, Bacteroides fragilis, and Eubacterium sulci; b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; and d) if the individual is determined to have an increased risk for or is suffering from CD, optionally treating the individual.

[0257] In some embodiments, the set of bacterial species includes Gemmiger formicilis, Eubacterium hallii, Blautia obeum, Roseburia inulinivorans, Fusicatenibacter saccharivorans, Eubacterium rectale, Roseburia faecis, Asaccharobacter celatus, Collinsella aerofaciens, Faecalibacterium prausnitzii, Anaerostipes hadrus, Lachnospira pectinoschiza, Ruminococcus torques, Clostridium leptum, Parabacteroides merdae, Ruminococcus bromii, Roseburia intestinalis, Adlercreutzia equolifaciens, Alistipes putredinis, Eubacterium sp. CAG: 38, Roseburia hominis, Agathobaculum butyriciproducens, Dorea formicigenerans, Bacteroides stercoris, Odoribacter splanchnicus, Oscillibacter sp. 57_20, Coprococcus catus, Bacteroides vulgatus, Dorea longicatena, Alistipes shahii, Lawsonibacter asaccharolyticus, Firmicutes bacterium CAG: 83, Oscillibacter sp. CAG: 241, Eubacterium ramulus, Alistipes finegoldii, Bacteroides caccae, Akkermansia muciniphila, Coprococcus comes, Butyricimonas virosa, Alistipes indistinctus, Bacteroides uniformis, Bacteroides coprocola, Eubacterium sp. CAG: 274, Desulfovibrio piger, Bacteroides massiliensis, Clostridium sp. CAG: 299, Gemella haemolysans, Streptococcus mitis, Actinomyces odontolyticus, Actinomyces sp. S6-Spd3, Actinomyces graevenitzii, Lactobacillus mucosae, Streptococcus parasanguinis, Rothia mucilaginosa, Actinomyces sp. oral taxon 181, Veillonella parvula, Escherichia coli, and Bacteroides fragilis.

[0258] In some embodiments, provided is a method for monitoring disease activity of Crohn's Disease (CD) in an individual, including the steps of: (a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Dorea formicigenerans, Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, and Escherichia coli; (b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; (c) determining the individual as having increased disease activity if the risk score is higher than a cutoff value.

[0259] In some embodiments, the reference data set is obtained by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects including active CD patients and inactive CD patients in a reference cohort.

[0260] The present invention provides a set of bacterial markers (used as single marker or in combination as indicated in Table 1.4, 5 and 9) useful for determining the presence or assessing risk of CD in a subject. The set of microbial markers can include (1) all the microbial markers listed in Table 1.1-1.5, (2) one or more microbial markers selected from Table 1.1-1.5.

[0261] Risk assessment or diagnosis of Crohn’s Disease

[0262] 1. Workflow

[0263] In one embodiment, to determine the risk of Crohn’s disease in an individual or whether the individual is suffering from Crohn’s disease, the following steps are carried out:(1) Obtain a stool sample from the individual and determine the level / relative abundance of one or more bacterial species selected from Table 1.2 and Table 1.3 (e.g. in a combination specified in Table 1.4 or Table 1.5) .(2) Obtain stool samples taken from a reference cohort comprising subjects with or without Crohn’s disease and determine the level / relative abundance of the same bacterial species in (1) .(3) Compare the level / relative abundance of the bacterial species obtained in (1) to the level / relative abundance of the bacterial species obtained in (2) .

[0264] In some embodiments, provided is a set of bacterial markers (used as single marker or in combination as indicated in Table 1.4, 5 and 6) useful for determining the presence or assessing risk of CD in a subject. The set of microbial markers can include (1) all the microbial markers listed in Table 1.1-1.5, (2) one or more microbial markers selected from Table 1.1-1.5.

[0265] 2. Determining the level of bacterial markers

[0266] In one embodiment, provided is a method to measure the level or amount of a signature DNA or RNA for one or more bacterial species found in a person’s stool sample as a mean to diagnose or assess the risk of Crohn’s disease. Thus, the first steps of the method are to obtain a stool sample from a test subject and extract microbial DNA or RNA from the sample.

[0267] Acquisition and Preparation of Stool Samples

[0268] A stool sample is obtained from a person to be tested or monitored for Crohn’s disease using an example method of the present disclosure. Collection of a stool sample from an individual can be easily achieved either in a clinic or at patient’s home. An appropriate amount of stool is collected and may be stored according to standard procedures prior to further preparation. The analysis of bacterial DNA or RNA found in a patient's stool sample according to the example method may be performed using established techniques. The methods for preparing stool samples for nucleic acid extraction are well known among those of skill in the art.

[0269] Extraction and Quantitation of DNA and RNA

[0270] Methods for extracting DNA from a biological sample are well-known and routinely practiced in the art of molecular biology. RNA contamination should be eliminated to avoid interference with DNA analysis.

[0271] Likewise, there are numerous methods for extracting DNA from a biological sample. The most common DNA extraction methods are mechanical, chemical and enzymatic lysis, precipitation, purification, and concentration. Other specific methods used to extract the DNA includes phenol-chloroform extraction, alcohol precipitation, or silica-based purification. Commercial kits, e.g. MO BIO PowerSoil DNA Isolation Kit (MO BIO Laboratories, Carlsbad, CA, USA) , QIAamp DNA Stool Mini Kit (Qiagen, Hilden, Germany) ,  RSC PureFood GMO, and Authentication Kit (Promega) , may also be used to obtain DNA from a biological sample from a test subject. The quality and concentration of the extracted DNA will be further assessed.

[0272] PCR-Based Quantitative Determination of Bacterial Level

[0273] Once DNA is extracted from a sample, the amount of a predetermined bacterial DNA (such as 16s rDNA or DNA encoded by a bacterial gene unique to the bacterial species) is quantified. In some embodiments, the method for determining the DNA level is an amplification-based method, e.g., by polymerase chain reaction (PCR) such as real-time quantitative PCR or Droplet Digital PCR (ddPCR) for DNA quantitative analysis.

[0274] The general methods of PCR are well-known in the art and are thus not described in detail herein. PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems.

[0275] PCR is most usually carried out as an automated process with a thermostable enzyme. In this process, the temperature of the reaction mixture is cycled through a denaturing region, a primer annealing region, and an extension reaction region automatically. Machines specifically adapted for this purpose are commercially available.

[0276] Real-time quantitative PCR, a semiquantitative PCR method, measures the fluorescence after each cycle and the intensity of the fluorescent signal reflects the momentary amount of DNA amplicons in the sample at that specific time.

[0277] Droplet-based digital PCR (ddPCR) is a refinement of the conventional polymerase chain reaction (PCR) methods. In ddPCR, DNA / RNA is encapsulated stochastically inside the microdroplets as reaction chambers. It provides more sensitive nuclei acid detection and absolute quantification.

[0278] Although PCR amplification of the target bacterial DNA or RNA is used in practicing certain embodiments of the present disclosure, one of skill in the art will recognize, however, that amplification of these DNA or RNA species in a sample may be accomplished by any known method, such as ligase chain reaction (LCR) , transcription-mediated amplification, and self-sustained sequence replication or nucleic acid sequence-based amplification (NASBA) , each of which provides sufficient amplification. More recently developed branched-DNA technology may also be used to quantitatively determine the amount of DNA or mRNA in the sample.

[0279] Metagenomics-based Quantitative Determination of Bacterial Level

[0280] The metagenomics dataset can include sequencing data generated from the same standardized protocol including steps from fecal DNA extraction to sequencing, raw data quality filter, host reads decontamination, to microbiome interpretation. Various techniques, such as shotgun metagenomic sequencing, can be used to characterize a sample’s species. In shotgun metagenomic sequencing, DNA is obtained from a heterogenous sample of cells and segmented into DNA fragments that can be aligned with the genomes of multiple microbial species to identify species in the sample.

[0281] 3. Generation of Disease Risk Score (probability of disease)

[0282] Machine Learning Model

[0283] In one embodiment, the disease risk score (probability of disease) can be generated using machine learning model, e.g. the random forest model. Random forest is a supervised learning approach used in machine learning for classification and regression. It is a supervised machine learning algorithm that averages the results or makes the final decision based on the majority voting of many decision trees applied to distinct subsets of a dataset to improve the dataset’s projected accuracy.

[0284] In this regard, in one embodiment, the method comprises constructing a machine learning model by:(1) obtaining a set of reference data set by determining in fecal samples the relative abundance of bacterial species selected from Table 1.2 and Table 1.3 (e.g. in a combination specified in Table 1.4 or Table 1.5) in a reference cohort;(2) dividing the data into training data set and test set;(3) under the training data set, dividing the data into training and validation subset, training a candidate model with the training subset and evaluate the performance of the candidate model using the validation subset;(4) tuning the candidate model by using different combinations of hyperparameters;(5) repeating step (3) - (4) and choosing the best candidate models with best performance;(6) evaluating the candidate models with test set; and(7) as an optional step, training a final model with the combinations of hypermeters of the best candidate model using the whole reference data set to obtain the final model.(8) using the optimal candidate model in step (6) or if step (7) is used to train the final model, using the final model in step (7) as the machine learning model.

[0285] After constructing the machine learning model, it can be deployed to generate a risk score (probability of disease) for the subject whose risk of Crohn’s disease is to be determined. In one embodiment, the method comprises:(1) determining the relative abundance of the corresponding bacterial species (same as those used to construct the machine learning model) in an individual whose diagnosis of Crohn’s disease is to be determined;(2) inputting the relative abundance of these species obtained from step (1) from the individual to the machine learning model constructed to generate a risk score (probability of disease) ; if random forest model is used, the relative abundance of the species listed in Table 1.2 and Table 1.3 obtained in step (1) from the individual are run down the decision trees in the random forest model;(3) determining the individual as being at risk for Crohn’s disease or is suffering from Crohn’s disease when the risk score is larger than a cutoff value, and determining the individual as not being at risk for Crohn’s disease or is not suffering from Crohn’s disease when the risk score is less than a cutoff value.

[0286] In some embodiments, the cutoff value can be determined as follow.

[0287] Determination of cutoff value

[0288] After obtaining the disease risk score (probability of disease) for the subject, either the default cutoff value 0.5 or optimized cutoff value can be used to determine whether or not the subject is suffering or at risk for Crohn’s disease.

[0289] In some embodiments, the optimized cutoff value can be calculated based on Youden’s index using the risk score data from reference cohort. In some embodiments, the risk score data is obtained by inputting the relative abundance of the set of bacterial species in the test set of the reference cohort into the machine learning model to generate the risk score for each bacterial species. Youden’s index can integrate sensitivity and specificity information, and by using Youden’s index analysis, the optimal cutoff value which provides the best tradeoff between sensitivity and specificity can be obtained. The Youden’s index (Y) can be calculated with the following formula:Y=sensitivity+specificity-1

[0290] Example method of determining the risk of CD in a subject

[0291] In one embodiment, to determine the risk score of Crohn’s disease in a subject, the following steps were carried out:(1) measuring the relative abundance of a set of selected bacterial species in the subject;(2) importing the relative abundance data into the pre-trained random forest model to generate the disease risk score (probability of disease) for the subject;(3) If a subject's risk score (probability of disease) is higher than the optimized cutoff value, the subject is regarded as suffering or at risk of Crohn’s disease. Conversely, if the risk score (probability of disease) is not higher than the optimized cutoff value, the individual is considered not to be suffering from or at risk of Crohn's disease.Example 1: Methods (Sample collection, DNA extraction and sequencing)Cohort Description and Study Subjects

[0292] In this study, the shotgun metagenomic profiling of fecal microbiomes from three diverse cohorts were analyzed to discover and validate the diagnostic models constructed with CD-specific bacterial species biomarkers. In discovery cohort, a total of 292 Chinese subjects (aged between 18 and 87 years) were recruited, including 174 subjects with Crohn’s disease (CD) and 118 normal individuals. Male accounted for 53.1%of the total subjects. In validation cohort, a total of 200 Chinese subjects (aged between 20 and 73) were recruited, including 92 CD patients and 108 normal individuals. Male accounted for 59.6%of total subjects. In another validation cohort, a total of 119 Australian subjects (aged between 18 and 78) were recruited, including 98 CD patients and 81 normal individuals. Male accounted for 46.6%of total subjects.

[0293] Patients were included if they were 18 years or older and had a diagnosis of CD defined by endoscopy, radiology, and histology; were on stable medication. Patients were excluded if they had infection with an enteric pathogen; had short bowel syndrome; or had significant hepatic, renal, endocrine, respiratory, neurologic, or cardiovascular disease.

[0294] Normal individuals who have no severe diseases (such as inflammatory bowel diseases, cancer, advanced adenoma) were included as controls.

[0295] All subjects consented to donate fecal samples and to the questionnaire investigation, where written informed consents were obtained. Fecal samples from the study subjects were stored at -80℃ for downstream microbiome analyses.

[0296] [Corrected under Rule 26, 14.03.2025]Besides, three public datasets comprising data from a total of 248 subjects were downloaded for model validation. These subjects included 102 subjects from United States (68 CD patients and 34 normal individuals, aged between 21 and 82) , 44 subjects from Netherlands (20 CD patients and 22 normal controls, aged between 21 and 71) and 102 subjects from China (48 CD patients and 54 normal controls, aged between 13 and 51) .Fecal DNA Extraction and DNA Sequencing

[0297] Fecal bacterial DNA was extracted by RSC PureFood GMO and Authentication Kit (Promega) with modifications to standard protocol to increase the yield of DNA. Approximately 100 mg from each stool sample was pretreated: stool sample suspended in 1 ml ddH2O and pelleted by centrifugation at 13,000×g for 1 min. Washed sample added with 800ul TE buffer (PH 7.5) , 16ul beta-Mercaptoethanol and 250U lyticase was sufficiently mixed and digested at 37℃ for 90 minutes, which was then pelleted by centrifugation at 13,000×g for 3 minutes.

[0298] [Corrected under Rule 26, 14.03.2025]After pretreatment, precipitate was re-suspended in 800ul CTAB buffer ( RSC PureFood GMO and Authentication Kit following manufacturer’s instructions) and mixed well. After samples were heated at 95℃ for 5 minutes and cooled down, nucleic acid was released from the samples by vortexing with 0.5mm and 0.1mm beads at 2850 rpm for 15 minutes. Following this, 40ul Proteinase K and 20ul RNase A were added and nucleic acid digested at 70℃ for 10 minutes. Finally, supernatant was obtained after centrifugation at 13,000×g, 5 minutes and placed in a RSC instrument for DNA extraction. The extracted fecal DNA was used for in-house (MagIC, Hong Kong, China) or outsourced (Novogene, Beijing, China) ultra-deep metagenomics sequencing via Ilumina Novaseq 6000.Quality Control of Raw Sequences

[0299] Raw sequence reads were trimmed by Trimmomatic (v0.39) . Non-human reads was then separated from contaminant host reads. Steps to acquire clean reads include: 1) Remove adapters; 2) Scan the read with a 4-base wide sliding window and remove reads when the average quality per base drop below 20; 3) Drop reads below the 50 bases long. Trimmed sequence reads were mapped to human genome (Reference database: hg37decv0.1) by KneadData (v0.10.0) to remove reads originated from the host. Pair-end two reads were concatenated together.Example 2: Methods for determining human gut bacteria composition and differential bacterial species between CD patients and normal individualsAnalysis of the Bacterial Microbiome

[0300] Profiling of the composition of bacterial communities was performed on metagenomic trimmed reads via MetaPhlAn3 (v3.0.13) . Mapping reads to clade-specific markers gene and annotation of species pangenomes was done through Bowtie2 (v2.2.4.2) . The output table contained bacterial species and its relative abundance in different levels, from kingdom to strain level. The resulting data were analyzed in R v3.6.1 using ggpubr (v0.2) and phyloseq (v1.24.2) . Human gut bacteria composition and the selected differential bacterial species were compared between CD patients and normal individuals via Microbiome Multivariable Associations with Linear Models (MaAsLin2) .Machine Learning Model

[0301] MaAsLin2 was used for the identification of discriminative features. Random forest (RF) was chosen to build CD patients versus normal controls prediction model using the selected fecal microbes because of its superior performance for classification with binary features. Random Forest is one of the approaches in metagenomic data analysis to build prediction models. As a widely used ensemble learning algorithm, Random Forest consists of a series of classification and regression trees (CARTs) to form a strong classifier. A subset of data randomly sampled from the original dataset with replacement is known as bootstrap sampling, applying to build the trees. When the training dataset for the current tree is drawn by the bootstrap method,  observations are left out from the overall dataset. With infinite N, there are 36.8%data not occurred in the training samples called out-of-bag (OOB) observations, which would not be used for constructing the trees. In addition, extra randomness introduced to the random forest as each decision tree splits nodes based on a random subset of features selected from the overall features. The features with the least Gini (Gini are used to evaluate the purity of the node) would be utilized to split the nodes in each iteration to generate the trees. With different subsets of data and features, the algorithm is able to train different trees and obtain the final classification by averaging or majority voting the result from the tree models.

[0302] A total of 174 CD patients and 118 normal controls were included as the discovery cohort for modeling. The selected species determined by MaAsLin2 were imported for the random forest model construction. The performance of the model was evaluated using 5-fold cross-validation. The performance of the model was evaluated in terms of binary classifiers with Area Under the Curve (AUC) in Receiver Operating Characteristic (ROC) curves. The parameters for model construction were tuned to achieve the best accuracy and kappa. These analyses were done using R packages MaAsLin2 v2_1.4.0, randomForest v4.6-14 and pROC v1.18.2.Droplet Digital PCR (ddPCR)Primers and probe design for the nine bacterial species

[0303] The specific gene sequence of each bacterium was downloaded from the MetaPhlAn3 database, and the specificity of each bacterium’s gene sequence was confirmed by performing the BLAST program using a public database such as GenBank. Based on the gene sequence, primers and probes were designed in Primer3Plus (https:  / / www. primer3plus. com / index. html) by evaluating their Tm value, GC content, and possible secondary structure. Primer and probe sequences for the 16s rDNA internal control were the same as in reported studies. Primers and probes are synthesized in BGI Bio-Solutions HongKong Co., Limited.Detection of selected bacterial species using droplet digital PCR (ddPCR)

[0304] Droplet digital PCR reactions were designed to detect the selected bacterial species (listed in Table 1.2 and Table 1.3) . The first reaction was designed to detect Bacteroides fragilis (probe conc. 120 nM) , Eubacterium sp. CAG: 274 (probe conc. 400 nM), Escherichia coli (probe conc. 120 nM) , and Actinomyces sp. oral taxon 181 (probe conc. 400nM) . The second reaction was designed to detect Roseburia inulinivorans (probe conc. 170 nM) , Ruminococcus obeum (probe conc. 480 nM) , Roseburia intestinalis (probe conc. 180 nM) , and Lawsonibacter asaccharolyticus (probe conc. 450 nM) . The third reaction was designed to detect Dorea formicigenerans (probe conc. 250 nM) and 16s reference gene (probe conc. 350 nM) . The sequences of primers and probes for detecting certain identified bacterial species (including the 9 selected bacterial species described herein) and the reference gene are listed in Table 1.1. The ddPCR mixture consisted of 10 μL ddPCR Supermix for Probes (No dUTP, Bio-rad Cat No. 1863024) , primers (900 nM) , probes (concentration was listed herein) , 2 μL DNA (diluted to 5 ng / μL in the reaction 1 and 2, diluted to 0.05 ng / μL in the reaction 3) , and nuclease-free water (to 20 μL) . The 20 μL ddPCR mixture and 70 μL Droplet Generation Oil for Probes (Bio-rad Cat No. 1863005) were then loaded into the DG8TM Cartridges (Bio-rad Cat No. 1864008) . DG8TM Gaskets (Bio-rad Cat No. 1863009) was hooked over the cartridge holder. The droplets for each sample would be generated by the QX200 Droplet Generator (Bio-rad) , and then transferred into the ddPCRTM 96-Well Plates (Bio-rad Cat No. 12001925) . After covering the plate with foil seal (Bio-rad Cat No. 1814040) and sealing in PX1 PCR Plate Sealer (Bio-Rad) , the 96-well plates were run on Bio-Rad T100 PCR System. The PCR program was: (1) initial denaturation at 95 ℃ for 10 min; (2) 40 cycles of denaturation at 94 ℃ for 30 s, and annealing and extension at 59 ℃ for 1 min; (3) enzyme deactivation at 98 ℃ for 10 min. After PCR, the fluorescence of droplet was detected in the QX200 Droplet Reader (Bio-rad) and the data would be analyzed by the QuantaSoftTM Analysis Pro (v 1.0.596) . The concentration of each bacterial species was determined and then normalized by the concentration of 16s reference gene (Normalized abundance of target = Concentration of the target  / Concentration of the 16s) .Table 1.1: Nucleotide Sequences of Primers And Probes For Bacterial Species Nucleic acid sequences of PCR products amplified by primers in Table 1.1 for the selected bacterial species

[0305] The nucleic acid sequences of PCR products amplified by primers in Table 1.1 ( “amplified fragments” ) for the selected bacterial species and the nucleic acid sequences of genes where the amplified fragments are located ( “marker genes” ) are listed below.

[0306] Bacteroides fragilis

[0307] Amplified fragment: Region from 1564367 to 1564483 of Bacteroides fragilis NCTC 9343, complete genome (GenBank: CR626927.1)

[0308] Nucleic acid sequence:

[0309]

[0310] Marker Gene: menE (1095 nt)

[0311] Gene product: putative O-succinylbenzoate--CoA ligase

[0312] Nucleic acid sequence:

[0313]

[0314] Escherichia coli

[0315] Amplified fragment: Region from 2558008 to 2558115 of Escherichia coli strain ATCC 25922 chromosome, complete genome (GenBank: CP117235.1)

[0316] Nucleic acid sequence:

[0317]

[0318] Marker Gene: hipA (1323 nt)

[0319] Gene product: type II toxin-antitoxin system serine / threonine protein kinase toxin HipA

[0320] Nucleic acid sequence:

[0321]

[0322] Dorea formicigenerans

[0323] Amplified fragment: Region from 1799237 to 1799333 of Dorea formicigenerans strain ATCC 27755 chromosome, complete genome (GenBank: CP102279.1)

[0324] Nucleic acid sequence:

[0325]

[0326] Marker gene: NQ560_08855 (471 nt)

[0327] Gene product: MogA / MoaB family molybdenum cofactor biosynthesis protein

[0328] Nucleic acid sequence:

[0329]

[0330] Eubacterium eligens (Lachnospira eligens)

[0331] Amplified fragment: Region from 1247080 to 1247192 of Lachnospira eligens strain FDAARGOS_1570 chromosome, complete genome (GenBank: CP085938.1)

[0332] Nucleic acid sequence:

[0333]

[0334] Marker gene: LK414_08735 (429 nt)

[0335] Nucleic acid sequence:

[0336]

[0337] Roseburia inulinivorans

[0338] Amplified fragment: Region from 1160270 to 1160376 of Roseburia inulinivorans DSM 16841 strain FDAARGOS_1587 ctg. s1.000001F_arrow_pilon, whole genome shotgun sequence (GenBank: JAJFOJ010000002.1)

[0339] Nucleic acid sequence:

[0340]

[0341] Marker gene: LK408_RS18325 (567 nt)

[0342] Gene product: 5-formyltetrahydrofolate cyclo-ligase

[0343] Nucleic acid sequence:

[0344]

[0345] Roseburia intestinalis

[0346] Amplified fragment: Region from 1246072 to 1246169 of Roseburia intestinalis L1-82 genome assembly, chromosome: 1 (GenBank: LR027880.1)

[0347] Nucleic acid sequence:

[0348]

[0349] Marker gene: RIL182_01152 (456 nt)

[0350] Nucleic acid sequence:

[0351]

[0352] Ruminococcus obeum (Blautia obeum)

[0353] Amplified fragment: Region from 3464821 to 3464916 of Blautia obeum ATCC 29174 chromosome, complete genome (GenBank: CP102265.1)

[0354] Nucleic acid sequence:

[0355]

[0356] Marker gene: NQ503_16555 (1236 nt)

[0357] Gene product: M20 family metallo-hydrolase

[0358]

[0359] Eubacterium ventriosum

[0360] Amplified fragment: Region from 1481183 to 1481301 of Eubacterium ventriosum strain ATCC 27560 chromosome, complete genome (GenBank: CP102282.1)

[0361] Nucleic acid sequence:

[0362]

[0363] Marker gene: NQ558_06695 (432 nt)

[0364] Gene product: stage V sporulation protein AB

[0365] Nucleic acid sequence:

[0366]

[0367] Anaerostipes hadrus

[0368] Amplified fragment: Region from 2653461 to 2653547 of Anaerostipes hadrus JCM 17467 DNA, complete genome (GenBank: AP028031.1)

[0369] Nucleic acid sequence:

[0370]

[0371] Marker gene: Ahadr17467_25260 (1041 nt)

[0372] Gene product: alpha / beta hydrolase

[0373] Nucleic acid sequence:

[0374]

[0375] Lawsonibacter asaccharolyticus

[0376] Amplified fragment: Region from 627394 to 627504 of Lawsonibacter asaccharolyticus strain OA10 chromosome, complete genome (GenBank: CP091870.1)

[0377] Nucleic acid sequence:

[0378]

[0379] Marker gene: L9O85_02835 (633nt)

[0380] Gene product: TetR / AcrR family transcriptional regulator

[0381] Nucleic acid sequence:

[0382]

[0383] Oscillibacter sp. CAG: 241

[0384] Amplified fragment: Region from 1167 to 1283 of Oscillibacter sp. CAG: 241, WGS project CBDI01000000 data, contig, whole genome shotgun sequence (GenBank: CBDI010000003.1)

[0385] Nucleic acid sequence:

[0386]

[0387] Marker gene: BN557_01301 (1506 nt)

[0388] Nucleic acid sequence:

[0389]

[0390] Oscillibacter sp. 57_20

[0391] Amplified fragment: Region from 33675 to 33760 of MAG: Oscillibacter sp. 57_20 Ley3_66761_scaffold_1931, whole genome shotgun sequence (GenBank: MNTE01000025.1)

[0392] Nucleic acid sequence:

[0393]

[0394] Marker gene: BHW41_06875 (522 nt)

[0395] Nucleic acid sequence:

[0396]

[0397] Lactobacillus rogosae

[0398] Amplified fragment: Region from 22 to 132 of Lactobacillus rogosae strain ATCC 27753 16S ribosomal RNA, partial sequence NCBI Reference Sequence: NR_104836.1)

[0399] Nucleic acid sequence:

[0400]

[0401] Marker gene: rRNA-16S ribosomal RNA (1452 nt)

[0402] Gene product: rRNA-16S ribosomal RNA

[0403] Nucleic acid sequence:

[0404]

[0405] Eubacterium sp. CAG: 274

[0406] Amplified fragment: Region from 7396 to 7500 of Eubacterium sp. CAG: 274 WGS project CBEX01000000 data, contig, whole genome shotgun sequence (GenBank: CBEX010000022.1)

[0407] Nucleic acid sequence:

[0408]

[0409] Marker gene: BN582_00070 (861 nt)

[0410] Gene product: prolipoprotein diacylglyceryl transferase

[0411] Nucleic acid sequence:

[0412]

[0413] Romboutsia ilealis

[0414] Amplifed fragment: Region from 3958 to 4076 of Romboutsia ilealis strain CRIB genome assembly, chromosome: chr1 (GenBank: LN555523.1)

[0415] Nucleic acid sequence:

[0416]

[0417] Marker gene: 23s_rRNA (2896 nt)

[0418] Gene product: 23s_rRNA

[0419] Nucleic acid sequence:

[0420]

[0421] Eubacterium sulci

[0422] Amplified fragment: Region from 1611276 to 1611387 of Eubacterium sulci ATCC 35585, complete genome (GenBank: CP012068.1)

[0423] Nucleic acid sequence:

[0424]

[0425] Marker gene: ADJ67_07525 (624 nt)

[0426] Nucleic acid sequence:

[0427]

[0428] Actinomyces sp. oral taxon 181

[0429] Amplified fragment: Region from 458 to 544 of Actinomyces sp. oral taxon 181 str. F0379 A_sporaltaxon181-1.0_Cont73.5, whole genome shotgun sequence (GenBank: AMEW01000023.1)

[0430] Nucleic acid sequence:

[0431]

[0432] Marker gene: HMPREF9061_01114 (545 nt)

[0433] Nucleic acid sequence:

[0434]

[0435] Actinomyces sp. S6-Spd3

[0436] Amplified fragment: Region from 162210 to 162329 of Actinomyces sp. S6-Spd3 contig286, whole genome shotgun sequence (GenBank: JRMV01000281.1)

[0437] Nucleic acid sequence:

[0438]

[0439] Marker gene: HMPREF1627_08235 (1098 nt)

[0440] Nucleic acid sequence:

[0441] Example 3: Determination of PerformanceGut bacterial profile is different between CD patients and normal individuals.

[0442] Now referring to FIG. 1A, a diagram of the top bacterial species associated with Crohn’s disease (CD) identified in the study is shown. The lollipop plot in the left panel shows the coefficient of each species with disease calculated by MaAsLin2 with age and gender adjusted. The box in the middle panel indicates the phylum of each species. The bar plot in the right panel demonstrated the proportion of each species present in CD and healthy groups.

[0443] As shown in FIG. 1A, with MaAsLin2 analysis, 46 bacterial species are found to be negatively correlated with CD, i.e., CD-depleted species, namely Gemmiger formicilis, Eubacterium hallii, Blautia obeum, Roseburia inulinivorans, Fusicatenibacter saccharivorans, Eubacterium rectale, Roseburia faecis, Asaccharobacter celatus, Collinsella aerofaciens, Faecalibacterium prausnitzii, Anaerostipes hadrus, Lachnospira pectinoschiza, Ruminococcus torques, Clostridium leptum, Parabacteroides merdae, Ruminococcus bromii, Roseburia intestinalis, Adlercreutzia equolifaciens, Alistipes putredinis, Eubacterium sp. CAG: 38, Roseburia hominis, Agathobaculum butyriciproducens, Dorea formicigenerans, Bacteroides stercoris, Odoribacter splanchnicus, Oscillibacter sp. 57_20, Coprococcus catus, Bacteroides vulgatus, Dorea longicatena, Alistipes shahii, Lawsonibacter asaccharolyticus, Firmicutes bacterium CAG: 83, Oscillibacter sp. CAG: 241, Eubacterium ramulus, Alistipes finegoldii, Bacteroides caccae, Akkermansia muciniphila, Coprococcus comes, Butyricimonas virosa, Alistipes indistinctus, Bacteroides uniformis, Bacteroides coprocola, Eubacterium sp. CAG: 274, Desulfovibrio piger, Bacteroides massiliensis, Clostridium sp. CAG: 299. Twelve bacterial species are also found to be positively correlated with CD, i.e., CD-enriched species, namely Gemella haemolysans, Streptococcus mitis, Actinomyces odontolyticus, Actinomyces sp. S6-Spd3, Actinomyces graevenitzii, Lactobacillus mucosae, Streptococcus parasanguinis, Rothia mucilaginosa, Actinomyces sp. oral taxon 181, Veillonella parvula, Escherichia coli, Bacteroides fragilis.

[0444] On further selection, 13 bacterial species (including those selected from the 46 CD-depleted species as described herein and additional bacterial species) are identified as the potential bacterial markers for CD diagnosis, found to be negatively correlated with CD, i.e., CD-depleted species, namely Eubacterium eligens, Romboutsia ilealis, Eubacterium ventriosum, Eubacterium sp. CAG: 274, Dorea formicigenerans, Anaerostipes hadrus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Roseburia intestinalis, Lactobacillus rogosae, Lawsonibacter asaccharolyticus, Ruminococcus obeum, and Roseburia inulinivorans. The mean relative abundance in normal individuals (%) of the CD-depleted species are shown in Table 1.2.

[0445] Five species (including those selected from the 12 CD-enriched species as described herein and additional bacterial species) are also identified as the potential bacterial markers for CD diagnosis, found to be positively correlated with CD, i.e., CD-enriched species, namely Actinomyces sp. oral taxon 181, Escherichia coli, Actinomyces sp. S6-Spd3, Bacteroides fragilis, and Eubacterium sulci. The mean relative abundance in normal individuals (%) of the CD-enriched species are shown in Table 1.3. As shown in FIG. 1B, the associations between CD, gender, age and the relative abundance of the 13 CD-depleted bacterial species and the 5 CD-enriched species calculated by machine learning model MaAsLin2 are shown. Significant associations (FDR < 0.05) were marked with a plus for positive correlations and a minus for negative correlations, respectively. False discovery rate (FDR) was computed by Benjamini–Hochberg correction. These bacteria can be used to determine the risk of or diagnose Crohn’s disease in a subject.

[0446] As such, bacteria listed in Table 1.2 and Table 1.3 can be used in different combinations to build an assessment model to determine the risk of or diagnose Crohn’s disease in a subject, and whether microbiome restoration therapy or supplementation is required. The relative abundance can be determined by qPCR or ddPCR using a panel of primers, or by metagenomics sequencing to determine the risk of or diagnose Crohn’s disease in the subject.

[0447] Bacteria listed in Table 1.2 can be administered to CD patients for relieving symptoms of CD.Table 1.2: Bacterial Species Enriched in Normal Individuals Compared to CD patients (CD-depleted species) . Table 1.3: Bacterial Species Enriched in CD patients Compared to Normal Individuals (CD-enriched species) Performance of machine learning model based on single bacterial marker or different bacterial markers combination using metagenomic data.

[0448] Single bacterial marker or different bacterial markers combination were used in the machine learning model to evaluate the performance of CD diagnosis. The model performance ranged from 0.391 to 0.773 with single bacterial biomarker (No. 1-18 in Table 1.4) . To enhance the model capability for disease diagnosis, different bacterial markers combination was tested, with the AUC ranging from 0.705 to 0.936 in the discovery cohort (No. 19-136 in Table 1.4) . The combination of nine bacterial biomarkers achieved the best diagnostic performance with AUC of 0.936 (No. 19 in Table 1.4) .Table 1.4: Performance of machine learning model based on different bacterial markers combination for Risk Prediction of CD using metagenomic data. Performance of machine learning model based on the nine selected bacterial species biomarkers using metagenomic data.

[0449] [Corrected under Rule 26, 14.03.2025]Now referring to FIGs. 2A-2C, which show the relative abundance (%) of 18 bacterial species markers (as listed in Table 1.2 and Table 1.3) determined by metagenomic sequencing in CD patients and normal individuals in Hong Kong, China (HK, China) discovery cohort, Hong Kong, China (HK, China) validation cohort and Australia (AUS) validation cohort, respectively.

[0450] [Corrected under Rule 26, 14.03.2025]To obtain the model with best performance in different population, different combinations of bacterial biomarkers were evaluated. Nine selected bacterial species biomarkers (also referred to as “biomarkers” ) were finally used in the machine learning model, including Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, Dorea formicigenerans (Table 1.5) . As shown in FIGs. 2A-2C, the relative abundance of these 9 biomarkers is consistent in HK, China discovery cohort, HK, China validation cohort and Australian validation cohort.

[0451] [Corrected under Rule 26, 14.03.2025]Now referring to FIG. 3A, which shows the receiver operating characteristic (ROC) curve and the area under the curve (AUC) of the machine learning model (random forest model) using the 9 biomarkers determined by metagenomic sequencing in the training set, test set, the Hong Kong, China (HK, China) validation cohort and Australia (AUS) validation cohort. The final model using these 9 markers has an Area Under the Curve (AUC) in Receiver Operating Characteristic (ROC) curve of 0.9357 (95%CI: 0.8871-0.9844; sensitivity of 88.33%, specificity of 89.47%with optimized threshold of 0.456) in HK, China discovery cohort (labeled as ‘Testset’ as shown in FIG. 3A) . In HK, China validation cohort and Australian validation cohort, the AUC is 0.8314 and 0.7264 with these 9 markers respectively.

[0452] [Corrected under Rule 26, 14.03.2025]Now referring to FIG. 3B, which shows the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the machine learning model (random forest model) using 9 markers determined by metagenomic sequencing in three public datasets including subjects from the United States, Netherlands and China. In the three public datasets, the AUCs of diagnostic model in United States, Netherlands and China are 0.8932, 0.8636 and 0.9676 respectively.

[0453] The results indicate that the machine learning model using the 9 biomarkers is accurate in predicting the risk of Crohn’s disease in a subject.Table 1.5: Bacterial Species Included in Machine Learning Model for Risk Prediction of CD Determination of Crohn’s Disease risk using different combinations of bacterial biomarkers

[0454] To determine the risk of or diagnose Crohn’s Disease in a subject, the following steps are carried out:(1) obtaining a set of training data by determining the relative abundance of species selected from Table 1.2 or Table 1.3 in a cohort of normal individuals and Crohn’s disease patients;(2) determining the relative abundance of these species in the subject whose risk of Crohn’s disease is to be determined;(3) comparing the relative abundance of these species in the subject with the training data using random forest model; and(4) generating decision trees by random forest from the training data. The relative abundances will be run down the decision trees and generate a risk score. If more than 50%trees in the model consider the subjects have Crohn’s disease, the subject being tested is deemed to be at an increased risk for Crohn’s disease. If less than 50%trees in the model consider the subject as normal individual, the subject being tested is deemed to not have an increased risk for Crohn’s disease.Bacterial species determined by droplet digital PCR (ddPCR)

[0455] Now referring to FIGs. 4A-4C. Using the specific primers and probes listed in the Table 1.1, the level of 6 species that are negatively correlated with CD and 3 species that are positively correlated with CD (Table 1.5) was determined by ddPCR in a subgroup of discovery cohort (CD = 172, Controls = 86) .Performance of machine learning model based on single bacterial marker or different bacterial markers combination using ddPCR data.

[0456] Single bacterial marker or different bacterial markers combination were used in the machine learning model to evaluate the performance of CD diagnosis. The model performance ranged from 0.608 to 0.772 with single bacterial biomarker (No. 1-9 in Table 1.6) . To enhance the model capability for disease diagnosis, different bacterial markers combination was tested, with the AUC ranging from 0.59 to 0.877 (No. 10-95 in Table 1.6, where only the combination whose AUC is greater than 0.6 was kept) . The combination of nine selected bacterial biomarkers yielded a superior AUC, with a relatively balanced sensitivity and specificity, resulting in an overall performance surpassing that of other combinations (No. 10 in Table 1.6) .Table 1.6: Performance of machine learning model based on different bacterial markers combination for Risk Prediction of CD using ddPCR data. Performance of machine learning model using ddPCR data

[0457] Nine selected bacterial species biomarkers were finally used in the machine learning model, including Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, Dorea formicigenerans (Table 1.5) . The subgroup of discovery cohort (CD = 172, Controls = 86) was randomly divided into training set (CD = 131, Controls = 61) and test set (CD = 41, Controls = 25) .

[0458] FIG. 5 shows the Receiver operating characteristic (ROC) curve and the area under the curve (AUC) of the machine learning model. AUC of random forest model using the 9 markers determined by droplet digital PCR. The final model using these 9 markers achieved an AUC of 0.999 in training set and 0.8688 in the test set (95%CI: 0.7779-0.9596; sensitivity of 90.24%, specificity of 76.00%with optimized threshold of 0.637) .CD-depleted bacterial species biomarkers contributed to the beneficial pathways which depleted in CD patients

[0459] FIGs. 6A-6D are stacked bar plots comparing the relative abundance (%) of different bacterial species (i.e., Actinomyces sp. oral taxon 181, Bacteroides fragilis, Blautia obeum or Ruminococcus obeum, Dorea formicigenerans, Eubacterium sp. CAG: 274, Lawsonibacter asaccharolyticus, Roseburia intestinalis, Roseburia inulinivorans and others) in the functional pathways of starch degradation, L-tryptophan biosynthesis, thiamine phosphate formation from pyrithiamine and oxythiamine (yeast) , and superpathway of L-lysine, L-theonine and L-methionine biosynthesis II, respectively, between healthy controls and CD patients. Each stacked bar plot indicates the contribution of bacteria species biomarkers and other bacteria in each functional pathway.

[0460] As shown in FIGs. 6A-6D, the abundance of functional pathways including amino acid biosynthesis (L-tryptophan biosynthesis, superpathway of L-lysine, L-threonine and L-methionine biosynthesis II) , carbohydrate degradation (Starch degradation) , and cofactor, carrier, and vitamin biosynthesis (Thiamine phosphate formation from purithiamine and exythiamine) were decresed in the CD patients. These pathways have higher abundance in healthy controls. The higher abundance of these pathways in healthy controls were mainly contributed by the CD-depleted species, such as B. obeum, R. inulinivorans, D. formicigenerans, Eubacterium sp. CAG: 274, and R. intestinalis. As such, by providing an effective amount of B. obeum, R. inulinivorans, D. formicigenerans, Eubacterium sp. CAG: 274, and / or R. intestinalis to a CD patient, the abundance of these functional pathways can be increased, which is beneficial in the treatment of CD.Example 4: Risk Prediction for CD on Test Subjects

[0461] In this example, to determine the risk score of Crohn’s disease in a subject, the following steps were carried out:(1) measuring the relative abundance of nine selected bacterial species biomarkers as shown in Table 1.5 in the subject by metagenomics sequencing or ddPCR;(2) importing the relative abundance data into the pre-trained machine learning model (random forest model) to generate the disease risk score (probability of disease) for the subject;(3) If a subject's risk score (probability of disease) is higher than the optimized cutoff value of 0.456 for metagenomics data or the optimized cutoff value of 0.637 for ddPCR data, the subject is regarded as suffering or at risk of Crohn’s disease. Conversely, if the risk score (probability of disease) is not higher than the optimized cutoff value, the individual is considered not to be suffering from or at risk of Crohn’s disease.

[0462] Table 1.7 shows the risk prediction (probability of disease) for CD on 98 subjects (including known CD patients and controls, as shown in “original group” ) using the machine learning model based on metagenomic data. Table 1.8 shows the risk prediction (probability of disease) for CD on 66 subjects (including known CD patients and controls, as shown in “original group” ) using the machine learning model based on ddPCR data.

[0463] The “original group” column indicates the condition evaluated by health professions by other means (e.g., colonoscopy) in each subject, where “CD” refers to a subject diagnosed with “Crohn’s disease” and “Controls” refers to a subject without Crohn’s disease. The risk score (probability of disease or risk prediction) was calculated after analyzing the sample from the subjects according to the disclosed method as described in the preceding examples. If the risk score was higher than the optimized cutoff value, the subject was labeled as “CD” in the “test results” column. Conversely, if the risk score was not higher than the optimized cutoff value, the subject was labeled as “Controls” in the “test results” column.

[0464] As seen in Table 1.7, some 34 out of 38 subjects in the original group of “Controls” were correctly identified as “Controls” in the test results using the machine learning model based on metagenomic data, and 53 out of 60 subjects in the original group of “CD” were correctly identified as “CD” in the test results using the machine learning model based on metagenomic data. As seen in Table 1.8, 19 out of 25 subjects in the original group of “Controls” were correctly identified as “Controls” in the test results using the machine learning model based on ddPCR data, and 37 out of 41 subjects in the original group of “CD” were correctly identified as “CD” in the test results using the machine learning model based on ddPCR data.

[0465] The sensitivity is 88.33%and specificity is 89.47%for the machine learning model based on metagenomic data. The sensitivity is 90.24%and specificity is 76%for the machine learning model based on ddPCR data.

[0466] These results demonstrate that the machine learning models based on metagenomic data and ddPCR data are both highly accurate in predicting the risk of CD in a subject.Table 1.7: Risk Score (probability of disease) for CD on subjects using the Machine Learning Model based on metagenomic data. Table 1.8: Risk Score (probability of disease) for CD on subjects using the Machine Learning Model based on ddPCR data. Example 5: Risk Prediction for CD patient with different disease activityDisease risk scores generated by machine learning model with the nine selected CD bacterial markers can reflect the metabolic dysregulations in patients with CD.

[0467] Metabolic dysfunctional score for each subject was determined by calculating the median Bray-Curtis dissimilarity to the reference control group based on the metabolic functional pathways. Now referring to FIG. 7, which shows a plot showing the correlation between functional dysbiosis scores and probability of disease generated by the machine learning model based on the nine selected CD bacterial species biomarkers according to Table 1.5. The correlation coefficient R and p value were given by Spearman correlation (*, p<0.05) . As shown in FIG. 7, the disease risk scores generated by the machine learning model using the 9 selected CD bacterial biomarkers (listed in Table 1.5) positively correlated with the dysfunctional scores. This indicates that the machine learning model can reflect the metabolic dysregulations in patients with CD.Diagnostic models based on the nine selected CD bacterial markers can diagnose CD patients in remission stage

[0468] CD is characterized by the alternation of periods of disease activity including flares and remissions. The accuracy and stability of the 9 selected bacterial markers (listed in Table 1.5) in CD patients with different disease activity were evaluated. Using the Crohn’s disease activity index (CDAI) , CD patients were categorized to those with active disease (CD: CDAI>150) and inactive disease (CD: CDAI≤150) .

[0469] FIGs. 8A-8C showed that the six CD depleted bacterial markers Roseburia inulinivorans, Blautia obeum, Lawsonibacter asaccharolyticus, Roseburia intestinalis, Dorea formicigenerans, Eubacterium sp. CAG: 274 were decreased in inactive CD patients compared to healthy controls, while the three CD enriched bacterial markers Bacteroides fragilis, Escherichia coli, Actinomyces sp. oral taxon 181 were increased in inactive CD patients compared to healthy controls. P values were given by the Wilcoxon rank sum test. (*p<0.05, **p<0.01, ***p<0.001, NS no significance) . FIG. 8D showed that the disease risk scores generated by the machine learning model using the nine selected CD bacterial markers showed no significant difference between CD patients with active and inactive disease. FIG. 8E showed that the model achieved an AUC of 0.837 in classifying CD patients in remission from controls. This indicates that the machine learning model can be used to determine the risk of Crohn’s disease in an individual regardless of disease activity.

[0470] Among the 9 selected bacterial markers, B. obeum, L. asaccharolyticus, and D. formicigenerans showed lower levels in active CD patients compared to inactive CD patients, whereas E. coli exhibited higher levels in active CD patients than in inactive CD patients. As such, these species are useful as bacterial markers for monitoring disease activity in CD patients. For example, a machine learning model can be trained in a cohort of CD patients with active and inactive disease utilizing a panel of markers comprising these species in order to be used as a biomarker for disease monitoring and for predicting disease flare.Example 6: Differentiating CD from other diseasesDiagnostic models based on the nine selected CD bacterial markers can differentiate CD from other gastrointestinal (GI) and non-GI diseases

[0471] In light of shared microbiota alterations across various diseases, it is important to verify disease specificity for the identified bacteria markers, thereby ensuring a low false positive rate for inflammatory bowel disease (IBD) diagnosis. For this purpose, several non-IBD disease datasets were assessed, consisting of subjects with gastrointestinal (GI) diseases (n=439) including colorectal cancer (CRC, n=160) , colorectal adenomas (CA, n=162) , irritable bowel syndrome (diarrhea subtype, IBS-D, n=117) and non-gastrointestinal (non-GI) diseases (n=291) including obesity (Body mass index>28; n=148) and cardiovascular disease (CVD; n=143) .

[0472] Now referring to FIGs. 9A-9B. FIG. 9A is a plot showing ROC curves of classifying CD patients from patients with other GI diseases, including CA (n=162) , IBS-D (n=117) , and CRC (n=160) . As shown in FIG. 9A, the CD diagnostic model / machine learning model (random forest model) using the 9 selected bacterial markers (listed in Table 1.5) determined by metagenomic sequencing can differentiate CD subjects from patients with other GI diseases with an AUROCs of 0.8963, 0.8283, 0.8078 for colorectal adenomas (CA) , irritable bowel syndrome (diarrhea subtype, IBS-D) and colorectal cancer (CRC) , respectively. FIG. 9B is the comparison of probability of disease (risk score) generated by the model based on the nine selected CD bacterial species biomarkers for CD patients and patients with other GI diseases. P values were calculated using the Wilcoxon rank-sum test (p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001; ns, no significance) . As shown in FIG. 9B, the probability of disease (risk score) for patients with other GI diseases is significantly lower than the probability of disease for CD patients, showing that the machine learning model is accurate in discriminating CD patients from patients with other GI diseases. FIG. 9C is a plot showing the ROC curve of classifying CD patients from patients with non-IBD GI diseases, where the CD diagnostic model differentiated CD subjects from patients with other GI diseases with an average AUROC of 0.8459.

[0473] Now referring to FIGs. 10A-10B. FIG. 10A is a plot showing ROC curves of classifying CD patients from patients with other non-GI diseases, including obesity (n=148) and CVD (n=143) , and the comparison of their probability of disease generated by model based on the nine selected CD bacterial species biomarkers. As shown in FIG. 10A, the same model can also differentiate CD subjects from patients with other non-GI diseases with AUROCs of 0.85 and 0.918 for obesity and cardiovascular disease (CVD) , respectively. FIG. 10B is a plot showing the comparison of probability of disease (risk score) generated by model based on nine selected CD bacterial species biomarkers among CD and other non-GI diseases. P values were calculated using the Wilcoxon rank-sum test (p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001; ns, no significance) . As shown in FIG. 10B, the probability of disease (risk score) for patients with other non-GI diseases is significantly lower than the probability of disease for CD patients, showing that the machine learning model is accurate in discriminating CD patients from patients with other non-GI diseases.

[0474] FIG. 10C is a plot showing the ROC curve of classifying CD patients from patients with all other non-IBD diseases (GI and non-GI diseases) (n = 730) using the machine learning model based on the nine selected bacterial species biomarkers. As shown in FIG. 10C, the machine learning model can differentiate CD subjects from patients with all other non-IBD diseases with an AUROC of 0.8609.

[0475] These results suggest that the diagnostic model encompassing the nine selected CD-associated biomarkers exhibited satisfactory performance in differentiating CD from all other non-IBD diseases, i.e., GI diseases and / or other non-GI diseases.

[0476] Among the nine selected bacterial markers, R. inulinivorans and B. obeum are depleted in CD patients compared with other diseases while B. fragilis and E. coli are enriched in CD patients compared with other diseases.Diagnostic models based on the nine selected CD bacterial markers can differentiate CD from another IBD subtype, ulcerative colitis

[0477] Now referring to FIG. 11, which shows the ROC curve of classifying CD patients from UC patients using the CD diagnostic model / machine learning model based on the nine selected bacterial species biomarker according to Table 1.5. Among the bacteria biomarkers, depletion of R. inulinivorans and B. obeum, and enrichment of B. fragilis and E. coli were characteristic of CD. As shown in FIG. 11, the CD diagnostic model can differentiate CD subjects from ulcerative colitis (UC) subjects with AUROC of 0.7014.

[0478] The results of the model based on the nine selected CD bacterial species biomarkers in differentiating CD from GI diseases, other non-GI (non-IBD) diseases and from UC are summarized in Table 1.9.Table 1.9: Performance of CD diagnostic model in differentiating CD from other gastrointestinal (GI) and non-GI diseases Example 7: Method for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) in an individual

[0479] Now referring to FIG. 12, which shows an example method 100 for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) in an individual with the following steps involved:

[0480] Step 110: determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;

[0481] Step 120: comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;

[0482] Step 130: determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value;

[0483] Step 140: if the individual is determined to have an increased risk for or is suffering from CD, optionally treating the individual.EMBODIMENT 2

[0484] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual, including the steps of: (a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii; (b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; (c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; and (d) if the individual is determined to have an increased risk for or is suffering from UC, optionally treating the individual.

[0485] In some embodiments, the set of bacterial species includes one or more bacterial species selected from the group consisting of Collinsella aerofaciens, Clostridium leptum, Ruminococcus torques, Asaccharobacter celatus, Gemmiger formicilis, Fusicatenibacter saccharivorans, Alistipes putredinis, Dorea longicatena, Ruminococcus bromii, Odoribacter splanchnicus, Lachnospira pectinoschiza, Coprococcus comes, Adlercreutzia equolifaciens, Roseburia inulinivorans, Blautia obeum, Eubacterium rectale, Dorea formicigenerans, Alistipes shahii, Bacteroides stercoris, Parabacteroides merdae, Oscillibacter sp. CAG: 241, Akkermansia muciniphila, Alistipes finegoldii, Eubacterium hallii, Oscillibacter sp. 57_20, Roseburia hominis, Phascolarctobacterium faecium, Collinsella stercoris, Bifidobacterium adolescentis, Roseburia intestinalis, Butyricimonas virosa, Alistipes indistinctus, Eubacterium sp. CAG: 38, Anaerostipes hadrus, Bacteroides caccae, Clostridium sp. CAG: 242, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Coprococcus catus, Bacteroides massiliensis, Roseburia faecis, Clostridium sp. CAG: 58, Ruthenibacterium lactatiformans, Bacteroides plebeius, Bilophila wadsworthia, Firmicutes bacterium CAG: 145, Eubacterium ramulus, Enterorhabdus caecimuris, Gemella morbillorum, Veillonella infantium, Haemophilus sp. HMSC71H05, Actinomyces sp. oral taxon 181, Blautia producta, Lactobacillus mucosae, Enterococcus avium, Veillonella atypica, Eubacterium sulci, Blautia hansenii, Rothia mucilaginosa, Clostridium spiroforme, Tyzzerella nexilis, Veillonella parvula, and Bacteroides fragilis.

[0486] In some embodiments, the reference data set is obtained by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects including patients with UC and normal individuals in a reference cohort.

[0487] In some embodiments, the comparing step in step (b) is performed using a machine learning model to generate the risk score.

[0488] In some embodiments, the machine learning model is constructed by (i) obtaining the reference data set by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects including patients with UC and normal individuals in a reference cohort; (ii) dividing the reference data set into a training data set and a test data set; (iii) dividing the training data set into a training subset and a validation subset; (iv) training a candidate model with the training subset and evaluating the performance of the candidate model using the validation subset; (v) optionally tuning the candidate model by adjusting hyperparameters in the candidate model; (vi) repeating step (iv) and optionally step (v) to produce a set of candidate models, and selecting an optimal candidate model among the set of candidate models; (vii) evaluating the performance of the optimal candidate model using the test data set; (viii) optionally training a final model with the reference data set using a set of hyperparameters corresponding to the optimal candidate model to produce the final model; and (ix) using the optimal candidate model of step (vii) or if step (viii) is used to train the final model, using the final model of step (viii) as the machine learning model for the comparing step in step (b) to generate the risk score. For sake of clarity, the reference data set used in step (viii) is the complete reference data set which includes both the training data set and the test data set.

[0489] In some embodiments, the machine learning model is a random forest model, and wherein the risk score in step (b) is generated by the following steps: (b1) generating an ensemble of decision trees by the random forest model using the relative abundance of the set of bacterial species in the reference data set; and (b2) running the relative abundance of the set of bacterial species from the individual along the ensemble of decision trees to generate the risk score.

[0490] In some embodiments, the optimal candidate model is the candidate model with the highest performance among the set of candidate models.

[0491] In some embodiments, the performance of the candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the candidate model. In some embodiments, the Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) is also referred to as the “AUROC” .

[0492] In some embodiments, the cutoff value is in the range of 0.4 to 0.7.

[0493] In some embodiments, the cutoff value is 0.5.

[0494] In some embodiments, the cutoff value is determined by the following steps: (c1) calculating a Youden's index using the risk score in the reference data set; and (c2) determining the cutoff value based on the Youden's index.

[0495] In some embodiments, the set of bacterial species includes one or more combinations of the bacterial species, wherein said combinations are selected from the group consisting of Table 2.4 and Table 2.6.

[0496] In some embodiments, the set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Clostridium spiroforme, and Gemella morbillorum.

[0497] In some embodiments, the set of bacterial species includes Fusicatenibacter saccharivorans, Clostridium leptum, Gemmiger formicilis, and optionally Ruminococcus torques and / or Odoribacter splanchnicus.

[0498] In some embodiments, the set of bacterial species includes Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0499] In some embodiments, provided is a method, wherein the set of bacterial species consists of Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0500] In some embodiments, the set of bacterial species includes Clostridium leptum, Fusicatenibacter saccharivorans, Odoribacter splanchnicus, Gemmiger formicilis, Clostridium spiroforme, and Ruminococcus torques, whereby the method identifies the individual's risk for UC over other non-IBD diseases.

[0501] In some embodiments, the other non-IBD diseases comprise colorectal cancer, colorectal adenomas, irritable bowel syndrome (diarrhea subtype) , obesity and cardiovascular disease.

[0502] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing or droplet digital PCR (ddPCR) .

[0503] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 2.4.

[0504] In some embodiments, the relative abundance of the set of bacterial species is determined by ddPCR, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 2.6.

[0505] In some embodiments, the ddPCR is performed by using one or more sets of primers selected from Table 2.1.

[0506] In some embodiments, if the individual is determined to have an increased risk for or is suffering from UC, treating the individual with an appropriate UC treatment, wherein the UC treatment is selected from the group consisting of a therapeutic treatment, fecal microbiota transplantation (FMT) , intestinal microbiota transplantation (IMT) , lifestyle and diet modifications, and surgical intervention. In some embodiments, the therapeutic treatment is selected from the group consisting of 5-Aminosalicylates (5-ASA, such as mesalazine, sulphasalazine) ; steroids (such as budesonide, prednisolone) ; thiopurines (such as azathioprine, mercaptopurine, methotrexate) ; TNF inhibitors (such as infliximab, adalimumab, and certolizumab pegol) ; anti-integrins (such as vedolizumab) , and anti-interleukin (IL) 12 / 23 (such as Ustekinumab) .

[0507] In some embodiments, the 5-Aminosalicylates are mesalazine or sulphasalazine; the steroids are budesonide or prednisolone; the thiopurines are azathioprine, mercaptopurine or methotrexate; the TNF inhibitors are infliximab, adalimumab or certolizumab pegol; the anti-integrins are vedolizumab; and the anti-interleukin (IL) 12 / 23 is ustekinumab.

[0508] In some embodiments, the UC is in remission stage.

[0509] In some embodiments, provided is a method for identifying a stool sample with an altered ulcerative colitis microbiome including the steps of: (a) determining the relative abundance of a set of bacterial species in the stool sample, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii; (b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; and (c) determining the stool sample as having an altered UC microbiome if the risk score is higher than a cutoff value.

[0510] In some embodiments, the comparing step in step (b) is performed using a machine learning model to generate the risk score.

[0511] In some embodiments, the machine learning model is constructed by (i) obtaining the reference data set by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects including patients with UC and normal individuals in a reference cohort; (ii) dividing the reference data set into a training data set and a test data set; (iii) dividing the training data set into a training subset and a validation subset; (iv) training a candidate model with the training subset and evaluating the performance of the candidate model using the validation subset; (v) optionally tuning the candidate model by adjusting hyperparameters in the candidate model; (vi) repeating step (iv) and optionally step (v) to produce a set of candidate models, and selecting an optimal candidate model among the set of candidate models; (vii) evaluating the performance of the optimal candidate model using the test data set; (viii) optionally training a final model with the reference data set using a set of hyperparameters corresponding to the optimal candidate model to produce the final model; and (ix) using the optimal candidate model in step (vii) or if step (viii) is used to train the final model, using the final model in step (viii) as the machine learning model for the comparing step in step (b) to generate the risk score.

[0512] In some embodiments, the machine learning model is a random forest model, and wherein the risk score in step (b) is generated by the following steps: (b1) generating an ensemble of decision trees by the random forest model using the relative abundance of the set of bacterial species in the reference data set; and (b2) running the relative abundance of the set of bacterial species from the individual along the ensemble of decision trees to generate the risk score. In some embodiments, the risk score is generated by averaging prediction across the decision trees or by majority voting of class label predicted across the decision trees, wherein the proportion of votes in each class across the ensemble is the predicted probability (i.e. the risk score) .

[0513] In some embodiments, the optimal candidate model is the candidate model with the highest performance among the set of candidate models.

[0514] In some embodiments, the performance of the candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the candidate model.

[0515] In some embodiments, the cutoff value is in the range of 0.4 to 0.7.

[0516] In some embodiments, the cutoff value is 0.5.

[0517] In some embodiments, the cutoff value is determined by the following steps: (c1) calculating a Youden's index using the risk score in the reference data set; and (c2) determining the cutoff value based on the Youden's index.

[0518] In some embodiments, the set of bacterial species includes one or more combinations of the bacterial species, wherein said combinations are selected from the group consisting of Table 2.4 and Table 2.6.

[0519] In some embodiments, the set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Clostridium spiroforme, and Gemella morbillorum.

[0520] In some embodiments, the set of bacterial species includes Fusicatenibacter saccharivorans, Clostridium leptum, Gemmiger formicilis, and optionally Ruminococcus torques and / or Odoribacter splanchnicus.

[0521] In some embodiments, the set of bacterial species includes Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0522] In some embodiments, the set of bacterial species consists of Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0523] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing or droplet digital PCR (ddPCR) .

[0524] In some embodiments, the relative abundance of the set of bacterial species is determined by metagenomic sequencing, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 2.4.

[0525] In some embodiments, the relative abundance of the set of bacterial species is determined by ddPCR, and the set of bacterial species includes one or more combinations of the bacterial species selected from Table 2.6.

[0526] In some embodiments, the ddPCR is performed by using one or more sets of primers selected from Table 2.1.

[0527] In some embodiments, provided is a method of increasing amino acid biosynthesis in an individual, wherein the amino acid is L-arginine, L-ornithine, or L-valine, including providing an effective amount of one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, and Ruminococcus torques to the individual.

[0528] In some embodiments, provided is a method of increasing amino acid biosynthesis in a cell, wherein the amino acid is L-arginine, L-ornithine, or L-valine, including providing an effective amount of one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, and Ruminococcus torques to the cell.

[0529] In some embodiments, provided is a method of increasing nucleoside and nucleotide biosynthesis in an individual, including providing an effective amount of one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, and Ruminococcus torques to the individual.

[0530] In some embodiments, provided is a method of increasing nucleoside and nucleotide biosynthesis in a cell, including providing an effective amount of one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, and Ruminococcus torques to the cell.

[0531] In some embodiments, the cell is part of a living organism.

[0532] In some embodiments, provided is a composition including Fusicatenibacter saccharivorans, Clostridium leptum, and Gemmiger formicilis.

[0533] In some embodiments, the composition further includes Ruminococcus torques.

[0534] In some embodiments, the composition further includes one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, and Bilophila wadsworthia.

[0535] In some embodiments, provided is a kit for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, including a reagents for detecting a set of bacterial species, wherein the set of bacterial species is selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0536] In some embodiments, the reagents include one or more sets of primers, wherein each set of primers is configured to amplify a nucleic acid sequence belonging to a bacterial species in the set of bacterial species.

[0537] In some embodiments, the one or more sets of primers are selected from Table 2.1.

[0538] In some embodiments, the amplification is PCR.

[0539] In some embodiments, the PCR is ddPCR, and the reagents further include one or more probes selected from Table 2.1.

[0540] In some embodiments, the set of bacterial species includes one or more combinations of the bacterial species, wherein said combinations are selected from the group consisting of Table 2.6.

[0541] In some embodiments, the set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Clostridium spiroforme, and Gemella morbillorum.

[0542] In some embodiments, the set of bacterial species includes Fusicatenibacter saccharivorans, Clostridium leptum, Gemmiger formicilis, and optionally Ruminococcus torques and / or Odoribacter splanchnicus.

[0543] In some embodiments, the set of bacterial species includes Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0544] In some embodiments, the set of bacterial species consists of Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0545] In some embodiments, provided is a computer program product for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, wherein the computer program product includes a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of: a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii; b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; and c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value.

[0546] In some embodiments, the reference data set is obtained by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects comprising patients with UC and normal individuals in a reference cohort.

[0547] In some embodiments, the comparing step in step (b) is performed using a machine learning model to generate the risk score.

[0548] In some embodiments, the machine learning model is constructed by i) obtaining the reference data set by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects comprising patients with UC and normal individuals in a reference cohort; ii) dividing the reference data set into a training data set and a test data set; iii) dividing the training data set into a training subset and a validation subset; iv) training a candidate model with the training subset and evaluating the performance of the candidate model using the validation subset; v) optionally tuning the candidate model by adjusting hyperparameters in the candidate model; vi) repeating step (iv) and optionally step (v) to produce a set of candidate models, and selecting an optimal candidate model among the set of candidate models; vii) evaluating the performance of the optimal candidate model using the test data set; viii) optionally training a final model with the reference data set using a set of hyperparameters corresponding to the optimal candidate model to produce the final model; and ix) using the optimal candidate model of step (vii) or if step (viii) is used to train the final model, using the final model of step (viii) as the machine learning model for the comparing step in step (b) to generate the risk score.

[0549] In some embodiments, the machine learning model is a random forest model, and wherein the risk score in step (b) is generated by the following steps: (b1) generating an ensemble of decision trees by the random forest model using the relative abundance of the set of bacterial species in the reference data set; and (b2) running the relative abundance of the set of bacterial species from the individual along the ensemble of decision trees to generate the risk score.

[0550] In some embodiments, the optimal candidate model is the candidate model with the highest performance among the set of candidate models.

[0551] In some embodiments, the performance of the candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the candidate model.

[0552] In some embodiments, the cutoff value is in the range of 0.4-0.7.

[0553] In some embodiments, the cutoff value is 0.5.

[0554] In some embodiments, the cutoff value is determined by the following steps: (c1) calculating a Youden’s index using the risk score in the reference data set; and (c2) determining the cutoff value based on the Youden’s index.

[0555] In some embodiments, provided is a method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual, including the steps of: a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Collinsella aerofaciens, Clostridium leptum, Ruminococcus torques, Asaccharobacter celatus, Gemmiger formicilis, Fusicatenibacter saccharivorans, Alistipes putredinis, Dorea longicatena, Ruminococcus bromii, Odoribacter splanchnicus, Lachnospira pectinoschiza, Coprococcus comes, Adlercreutzia equolifaciens, Roseburia inulinivorans, Blautia obeum, Eubacterium rectale, Dorea formicigenerans, Alistipes shahii, Bacteroides stercoris, Parabacteroides merdae, Oscillibacter sp. CAG: 241, Akkermansia muciniphila, Alistipes finegoldii, Eubacterium hallii, Oscillibacter sp. 57_20, Roseburia hominis, Phascolarctobacterium faecium, Collinsella stercoris, Bifidobacterium adolescentis, Roseburia intestinalis, Butyricimonas virosa, Alistipes indistinctus, Eubacterium sp. CAG: 38, Anaerostipes hadrus, Bacteroides caccae, Clostridium sp. CAG: 242, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Coprococcus catus, Bacteroides massiliensis, Roseburia faecis, Clostridium sp. CAG: 58, Ruthenibacterium lactatiformans, Bacteroides plebeius, Bilophila wadsworthia, Firmicutes bacterium CAG: 145, Eubacterium ramulus, Enterorhabdus caecimuris, Gemella morbillorum, Veillonella infantium, Haemophilus sp. HMSC71H05, Actinomyces sp. oral taxon 181, Blautia producta, Lactobacillus mucosae, Enterococcus avium, Veillonella atypica, Eubacterium sulci, Blautia hansenii, Rothia mucilaginosa, Clostridium spiroforme, Tyzzerella nexilis, Veillonella parvula, and Bacteroides fragilis; b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; and d) if the individual is determined to have an increased risk for or is suffering from UC, optionally treating the individual.

[0556] In some embodiments, provided is a method for monitoring disease activity of ulcerative colitis (UC) in an individual, including the steps of: a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Odoribacter splanchnicus Gemmiger formicili, Actinomyces sp. oral taxon 181 and Clostridium spiroforme; b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; and c) determining the individual as having increased disease activity if the risk score is higher than a cutoff value.

[0557] In some embodiments, the reference data set is obtained by determining relative abundances of each bacterial species in the set of bacterial species in stool samples from a group of subjects comprising active UC patients and inactive UC patients in a reference cohort.

[0558] Risk assessment or diagnosis of ulcerative colitis

[0559] 1. Workflow

[0560] In one embodiment, to determine the risk of ulcerative colitis in an individual or whether the individual is suffering from ulcerative colitis, the following steps will be carried out:1. Obtain a stool sample from the individual and determine the level / relative abundance of one or more bacterial species selected from Table 2.2 and Table 2.3 (e.g. in a combination specified in Table 2.4 or 5) .2. Obtain stool samples taken from a reference cohort comprising subjects with or without ulcerative colitis (UC) and determine the level / relative abundance of the same bacterial species in step (1) .3. Compare the level / relative abundance of the bacterial species obtained in step (1) to the level / relative abundance of the bacterial species obtained in step (2) .

[0561] Reference cohort as used herein refers a group of individuals with ulcerative colitis and without ulcerative colitis (healthy control or normal individuals) who provided samples (e.g. stool samples) for the determination of level of pre-selected bacterial species in such samples, wherein diagnosis of ulcerative colitis was diagnosed by endoscopy, radiology, and histology examinations; and control are individuals who did not suffer from ulcerative colitis or present any signs or symptoms of gastrointestinal illness. In order to properly establish a “reference cohort” , a sufficient number of individuals (e.g., at least 10, 12, 15, 20, 24 or more individuals) with and without ulcerative colitis must be included in each of the ulcerative colitis group and the control group to provide samples for determination of the average level (s) of one or more pre-selected bacterial species.

[0562] 2. Determining the level of bacterial markers

[0563] In one embodiment, provided is a method to measure the level or amount of a signature DNA or RNA for one or more bacterial species found in a person’s stool sample as a mean to diagnose or assess the risk of ulcerative colitis. Thus, the first steps of the method are to obtain a stool sample from a test subject and extract microbial DNA or RNA from the sample.

[0564] Acquisition and Preparation of Stool Samples

[0565] A stool sample is obtained from a person to be tested or monitored for ulcerative colitis using an example method of the present disclosure. Collection of a stool sample from an individual can be easily achieved either in a clinic or at patient’s home. An appropriate amount of stool is collected and may be stored according to standard procedures prior to further preparation. The analysis of bacterial DNA or RNA found in a patient's stool sample according to the example method may be performed using established techniques. The methods for preparing stool samples for nucleic acid extraction are well known among those of skill in the art.

[0566] Extraction and Quantitation of DNA and RNA

[0567] Methods for extracting DNA from a biological sample are well-known and routinely practiced in the art of molecular biology. RNA contamination should be eliminated to avoid interference with DNA analysis.

[0568] Likewise, there are numerous methods for extracting DNA from a biological sample. The most common DNA extraction methods are mechanical, chemical and enzymatic lysis, precipitation, purification, and concentration. Other specific methods used to extract the DNA includes phenol-chloroform extraction, alcohol precipitation, or silica-based purification. Commercial kits, e.g. MO BIO PowerSoil DNA Isolation Kit (MO BIO Laboratories, Carlsbad, CA, USA) , QIAamp DNA Stool Mini Kit (Qiagen, Hilden, Germany) ,  RSC PureFood GMO, and Authentication Kit (Promega) , may also be used to obtain DNA from a biological sample from a test subject. The quality and concentration of the extracted DNA will be further assessed.

[0569] PCR-Based Quantitative Determination of Bacterial Level

[0570] Once DNA is extracted from a sample, the amount of a predetermined bacterial DNA (such as 16s rDNA or DNA encoded by a bacterial gene unique to the bacterial species) is quantified. In some embodiments, the method for determining the DNA level is an amplification-based method, e.g., by polymerase chain reaction (PCR) , such as real-time quantitative PCR or Droplet Digital PCR (ddPCR) for DNA quantitative analysis.

[0571] The general methods of PCR are well-known in the art and are thus not described in detail herein. PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems.

[0572] PCR is most usually carried out as an automated process with a thermostable enzyme. In this process, the temperature of the reaction mixture is cycled through a denaturing region, a primer annealing region, and an extension reaction region automatically. Machines specifically adapted for this purpose are commercially available.

[0573] Real-time quantitative PCR, a semiquantitative PCR method, measures the fluorescence after each cycle and the intensity of the fluorescent signal reflects the momentary amount of DNA amplicons in the sample at that specific time.

[0574] Droplet-based digital PCR (ddPCR) is a refinement of the conventional polymerase chain reaction (PCR) methods. In ddPCR, DNA / RNA is encapsulated stochastically inside the microdroplets as reaction chambers. It provides more sensitive nuclei acid detection and absolute quantification.

[0575] Although PCR amplification of the target bacterial DNA or RNA is typically used in practicing certain embodiments of the present disclosure, one of skill in the art will recognize, however, that amplification of these DNA or RNA species in a sample may be accomplished by any known method, such as ligase chain reaction (LCR) , transcription-mediated amplification, and self-sustained sequence replication or nucleic acid sequence-based amplification (NASBA) , each of which provides sufficient amplification. More recently developed branched-DNA technology may also be used to determine the amount of DNA or mRNA in the sample quantitatively.

[0576] Metagenomics-based Quantitative Determination of Bacterial Level

[0577] The metagenomics dataset can include sequencing data generated from the same standardized protocol including steps from fecal DNA extraction to sequencing, raw data quality filter, host reads decontamination, to microbiome interpretation. Various techniques, such as shotgun metagenomic sequencing, can be used to characterize a sample’s species. In shotgun metagenomic sequencing, DNA is obtained from a heterogenous sample of cells and segmented into DNA fragments that can be aligned with the genomes of multiple microbial species to identify species in the sample.

[0578] 3. Generation of disease risk score (probability of disease)

[0579] Construction of Machine Learning Model

[0580] In one embodiment, the disease risk score (probability of disease) can be generated using machine learning model, e.g. the random forest model. Random forest is a supervised learning approach used in machine learning for classification and regression. It is a supervised machine learning algorithm that averages the results or makes the final decision based on the majority voting of many decision trees applied to distinct subsets of a dataset to improve the dataset’s projected accuracy.

[0581] In this regard, in one embodiment, the method comprises constructing a machine learning model by(1) obtaining a set of reference data set by determining in fecal samples the relative abundance of bacterial species selected from Table 2.2 and Table 2.3 (e.g. in a combination specified in Table 2.4 or Table 2.5) in a reference cohort;(2) dividing the data into training and test set(3) under the training data set, dividing the data into training and validation subset, training a candidate model with the training subset and evaluate the performance of the candidate model using the validation subset;(4) tuning the candidate model by using different combinations of hyperparameters;(5) repeating steps (3) - (4) and choosing the best candidate models with best performance; and(6) evaluating the candidate models with test set; and(7) as an optional step, training a final model with the combinations of hypermeters of the best candidate model using the whole reference data set to obtain the final model.(8) using the optimal candidate model in step (6) or if step (viii) is used to train the final model, using the final model in step (7) as the machine learning model.

[0582] After constructing the machine learning model, it can be deployed to generate a risk score (probability of disease) for the subject whose risk of ulcerative colitis is to be determined. In one embodiment, the method comprises:(1) determining the relative abundance of the corresponding bacterial species (same as those used to construct the machine learning model) in an individual whose diagnosis of ulcerative colitis is to be determined. In one embodiment, the method comprises:(2) inputting the relative abundance of these species obtained from step (1) from the individual to the machine learning model constructed to generate a risk score (probability of disease) ; if random forest model is used, the relative abundance of the species listed in Table 2.2 and Table 2.3 obtained in step (1) from the individual are run down the decision trees in the random forest model;(3) determining the individual as being at risk for Ulcerative colitis or is suffering from Ulcerative colitis when the risk score is larger than a cutoff value, and determining the individual as not being at risk for Ulcerative colitis or is not suffering from Ulcerative colitis when the risk score is less than a cutoff value.

[0583] In some embodiments, the cutoff value can be determined as follow.

[0584] Determination of cutoff value

[0585] After obtaining the disease risk score (probability of disease) for each subject, either the default cut-off value 0.5 or optimized cut-off value can be used to determine whether or not the subject is suffering or at risk for Ulcerative colitis.

[0586] In some embodiments, the optimized cutoff value can be calculated based on Youden’s index using the risk score data from reference cohort. In some embodiments, the risk score data is obtained by inputting the relative abundance of the set of bacterial species in the test set of the reference cohort into the machine learning model to generate the risk score for each bacterial species. Youden’s index can integrate sensitivity and specificity information, and by using Youden’s index analysis, the optimal cutoff value which provides the best tradeoff between sensitivity and specificity can be obtained. The Youden’s index (Y) can be calculated with the following formula:

[0587] Y=sensitivity+specificity-1

[0588] If a subject's risk score (probability of disease) is higher than the optimized cutoff value, the subject is regarded as suffering or at risk of Ulcerative colitis. On the other hand, when the risk score (probability of disease) is no higher than the optimized cutoff value, the individual is deemed as not suffering from or at risk of Ulcerative colitis.

[0589] Example method of determining the risk of UC in a subject

[0590] In one embodiment, to determine the risk score of ulcerative colitis (UC) in a subject, the following steps were carried out:(1) measuring the relative abundance of a set of selected bacterial species in the subject;(2) importing the relative abundance data into the pre-trained random forest model to generate the disease risk score (probability of disease) for the subject;(3) If a subject's risk score (probability of disease) is higher than the optimized cutoff value, the subject is regarded as suffering or at risk of ulcerative colitis. Conversely, if the risk score (probability of disease) is not higher than the optimized cutoff value, the individual is considered not to be suffering from or at risk of ulcerative colitis.Example 1: Methods (Sample collection, DNA extraction and sequencing)

[0591] Cohort Description and Study Subjects

[0592] In this study, the shotgun metagenomic profiling of fecal microbiomes from two diverse cohorts were analyzed to discover and validate the diagnostic models constructed with ulcerative colitis-specific (UC-specific) bacterial species biomarkers. In discovery cohort, a total of 323 Chinese subjects (aged between 21 and 87 years) were recruited, including 205 subjects with Ulcerative colitis (UC) and 118 normal individuals. The male accounts for 44.3%. In validation cohort, a total of 278 Chinese subjects (aged between 20 and 88) were recruited, including 139 UC patients and 139 normal individuals. The male accounts for 54.3%. Patients were included if they were 18 years or older and had a diagnosis of UC defined by endoscopy, radiology, and histology; were on stable medication. Patients were excluded if they had infection with an enteric pathogen; had short bowel syndrome; or had significant hepatic, renal, endocrine, respiratory, neurologic, or cardiovascular disease. Normal individuals who have no severe diseases (such as inflammatory bowel diseases, cancer, advanced adenoma) were included as controls. All subjects consented to donate fecal samples and to the questionnaire investigation, where written informed consents were obtained. Fecal samples from the study subjects were stored at -80℃ for downstream microbiome analyses.

[0593] [Corrected under Rule 26, 14.03.2025]Besides, a total of 248 subjects come from three public datasets were downloaded for model validation, including 87 subjects from United States (53 UC patients and 34 normal individuals, aged between 20 and 82) , 45 subjects from Netherlands (23 UC patients and 22 normal controls, aged between 19 and 80) and 40 subjects from China (25 UC patients and 15 normal controls) .

[0594] Fecal DNA Extraction and DNA Sequencing

[0595] Fecal bacterial DNA was extracted by RSC PureFood GMO and Authentication Kit (Promega) with modifications to standard protocol to increase the yield of DNA. Approximately 100 mg from each stool sample was pretreated: stool sample suspended in 1 ml ddH2O and pelleted by centrifugation at 13,000×g for 1 min. Washed sample added with 800ul TE buffer (PH 7.5) , 16ul beta-Mercaptoethanol and 250U lyticase was sufficiently mixed and digested at 37℃ for 90 minutes, which was then pelleted by centrifugation at 13,000×g for 3 minutes.

[0596] [Corrected under Rule 26, 14.03.2025]After pretreatment, precipitate was re-suspended in 800ul CTAB buffer ( RSC PureFood GMO and Authentication Kit following manufacturer’s instructions) and mixed well. After samples were heated at 95℃ for 5 minutes and cooled down, nucleic acid was released from the samples by vortexing with 0.5mm and 0.1mm beads at 2850 rpm for 15 minutes. Following this, 40ul Proteinase K and 20ul RNase A were added and nucleic acid digested at 70℃ for 10 minutes. Finally, supernatant was obtained after centrifugation at 13,000×g, 5 minutes and placed in a RSC instrument for DNA extraction. The extracted fecal DNA was used for in-house (MagIC, Hong Kong, China) or outsourced (Novogene, Beijing, China) ultra-deep metagenomics sequencing via Ilumina Novaseq 6000.

[0597] Quality Control of Raw Sequences

[0598] Raw sequence reads were trimmed by Trimmomatic (v0.39) . Non-human reads were then separated from contaminant host reads. Steps to acquire clean reads include: 1) Remove adapters; 2) Scan the read with a 4-base wide sliding window and remove reads when the average quality per base drop below 20; 3) Drop reads below the 50 bases long. Trimmed sequence reads were mapped to human genome (Reference database: hg37decv0.1) by KneadData (v0.10.0) to remove reads originated from the host. Pair-end two reads were concatenated together.Example 2: Methods for determining human gut bacteria composition and differential bacterial species between UC patients and normal individuals

[0599] Analysis of the Bacterial Microbiome

[0600] Profiling of the composition of bacterial communities was performed on metagenomic trimmed reads via MetaPhlAn3 (v3.0.13) . Mapping reads to clade-specific markers gene and annotation of species pangenomes was done through Bowtie2 (v2.2.4.2) . The output table contained bacterial species and its relative abundance in different levels, from kingdom to strain level. The resulting data were analyzed in R v3.6.1 using ggpubr (v0.2) and phyloseq (v1.24.2) . Human gut bacteria composition and the selected bacterial species were compared between UC patients and normal individuals via Microbiome Multivariable Associations with Linear Models (MaAsLin2) .

[0601] Machine Learning Model

[0602] MaAsLin2 was used for the identification of discriminative features. Random forest (RF) was chosen to build UC patients versus normal controls prediction model using the selected fecal microbes because of its superior performance for classification with binary features. Random Forest is one of the approaches in metagenomic data analysis to build prediction models. As a widely used ensemble learning algorithm, Random Forest consists of a series of classification and regression trees (CARTs) to form a strong classifier. A subset of data randomly sampled from the original dataset with replacement is known as bootstrap sampling, applying to build the trees. When the training dataset for the current tree is drawn by the bootstrap method,  observations are left out from the overall dataset. With infinite N, there are 36.8%data not occurred in the training samples called out-of-bag (OOB) observations, which would not be used for constructing the trees. In addition, extra randomness introduced to the random forest as each decision tree splits nodes based on a random subset of features selected from the overall features. The features with the least Gini (Gini are used to evaluate the purity of the node) would be utilized to split the nodes in each iteration to generate the trees. With different subsets of data and features, the algorithm is able to train different trees and obtain the final classification by averaging or majority voting the result from the tree models.

[0603] A total of 205 UC patients and 118 normal controls were included as the discovery cohort for modeling. The selected species which determined by MaAsLin2 were imported for the random forest model construction. The performance of the model was evaluated using 5-fold cross-validation. The performance of the model was evaluated in terms of binary classifiers with Area Under the Curve (AUC) in Receiver Operating Characteristic (ROC) curves. The parameters for model construction were tuned to achieve the best accuracy and kappa. These analysis were done using R packages MaAsLin2 v2_1.4.0, randomForest v4.6-14 and pROC v1.18.2.

[0604] Droplet Digital PCR (ddPCR)

[0605] Primers and probe design for ten selected bacterial species

[0606] The specific gene sequence of each bacterium was downloaded from the MetaPhlAn3 database, and the specificity of each bacterium’s gene sequence was confirmed by performing the BLAST program using a public database such as GenBank. Based on the gene sequence, primers and probes were designed in Primer3Plus (https:  / / www. primer3plus. com / index. html) by evaluating their Tm value, GC content, and possible secondary structure. Primer and probe sequences for the 16s rDNA internal control were the same as in reported studies. Primers and probes are synthesized in BGI Bio-Solutions HongKong Co., Limited.

[0607] Detection of selected bacterial species using droplet digital PCR (ddPCR)

[0608] Droplet digital PCR reactions were designed to detect the selected bacterial species (listed in Table 2.2 and Table 2.3) . The first reaction was designed to detect Blautia hansenii (probe conc. 180 nM) , Clostridium leptum (probe conc. 400 nM) , Actinomyces sp. oral taxon 181 (probe conc. 180 nM) , and Gemella morbillorum (probe conc. 450 nM) . The second reaction was designed to detect Clostridium spiroforme (probe conc. 120 nM) , Fusicatenibacter saccharivorans (probe conc. 480 nM) , Odoribacter splanchnicus (probe conc. 200 nM) , and Gemmiger formicilis (probe conc. 450 nM) . The third reaction was designed to detect Bilophila wadsworthia (probe conc. 120 nM) , Ruminococcus torques (probe conc. 480 nM) and 16s reference gene (probe conc. 200 nM) . The sequences of primers and probes for detecting certain identified bacterial species (including the ten selected bacterial species described herein) and the reference gene are listed in Table 2.1. The ddPCR mixture consisted of 10 μL ddPCR Supermix for Probes (No dUTP, Bio-rad Cat No. 1863024) , primers (900 nM) , probes (concentration was listed herein) , 2 μL DNA (diluted to 5 ng / μL in the reaction 1 and 2, diluted to 0.05 ng / μL in the reaction 3) , and nuclease-free water (to 20 μL) . The 20 μL ddPCR mixture and 70 μL Droplet Generation Oil for Probes (Bio-rad Cat No. 1863005) were then loaded into the DG8TM Cartridges (Bio-rad Cat No. 1864008) . DG8TM Gaskets (Bio-rad Cat No. 1863009) was hooked over the cartridge holder. The droplets for each sample would be generated by the QX200 Droplet Generator (Bio-rad) , and then transferred into the ddPCRTM 96-Well Plates (Bio-rad Cat No. 12001925) . After covering the plate with foil seal (Bio-rad Cat No. 1814040) and sealing in PX1 PCR Plate Sealer (Bio-Rad) , the 96-well plates were run on Bio-Rad T100 PCR System. The PCR program was: (1) initial denaturation at 95 ℃ for 10 min; (2) 40 cycles of denaturation at 94 ℃ for 30 s, and annealing and extension at 59 ℃ for 1 min; (3) enzyme deactivation at 98 ℃ for 10 min. After PCR, the fluorescence of droplet was detected in the QX200 Droplet Reader (Bio-rad) and the data would be analyzed by the QuantaSoftTM Analysis Pro (v 1.0.596) . The concentration of each bacterial species was determined and then normalized by the concentration of 16s reference gene (Normalized abundance of target = Concentration of the target  / Concentration of the 16s) .Table 2.1: Nucleotide sequences of primers and probes for the bacterial species. Nucleic acid sequences of PCR products amplified by primers in Table 2.1 for the bacterial species

[0609] The nucleic acid sequences of PCR products amplified by primers in Table 2.1 ( “amplified fragments” ) for the selected bacterial species and the nucleic acid sequences of genes where the amplified fragments are located ( “marker genes” ) are listed below.

[0610] Bilophila wadsworthia

[0611] Amplified fragment: Region from 202065 to 202184 of Bilophila wadsworthia ATCC 49260 T370DRAFT_scaffold00002.2_C, whole genome shotgun sequence (GenBank: JNJP01000002.1)

[0612]

[0613] Marker gene: T370_RS0102035 (1044 nt)

[0614] Gene product: uroporphyrinogen decarboxylase family protein

[0615] Nucleic acid sequence:

[0616]

[0617] Clostridium leptum

[0618] Amplified fragment: Region from 59277 to 59356 of [Clostridium] leptum DSM 753 3c, whole genome shotgun sequence (GenBank: NOXF01000003.1)

[0619] Nucleic acid sequence:

[0620]

[0621] Marker gene: CH238_05410 (2664 nt)

[0622] Nucleic acid sequence:

[0623]

[0624] Fusicatenibacter saccharivorans

[0625] Amplified fragment: Region from 75654 to 75747 of Fusicatenibacter saccharivorans strain 2789STDY5834923 genome assembly, contig: SCcontig000005, whole genome shotgun sequence (GenBank: CZBB01000005.1)

[0626] Nucleic acid sequence:

[0627]

[0628] Marker gene: dacB_3 (1308 nt)

[0629] Gene product: D-alanyl-D-alanine carboxypeptidase dacB precursor

[0630] Nucleic acid sequence:

[0631]

[0632] Gemmiger formicilis

[0633] Amplified fragment: Region from 1148 to 1265 of Gemmiger formicilis strain ATCC 27749 genome assembly, contig: EI45DRAFT_scaffold00074.74, whole genome shotgun sequence (GenBank: FUYF01000074.1)

[0634] Nucleic acid sequence:

[0635]

[0636] Marker gene: SAMN02745178_02911 (1362 nt)

[0637] Nucleic acid sequence:

[0638]

[0639] Odoribacter splanchnicus

[0640] Amplified fragment: Region from 2137841 to 2137944 of Odoribacter splanchnicus strain NCTC10825 genome assembly, chromosome: 1 (GenBank: LT906459.1)

[0641] Nucleic acid sequence:

[0642]

[0643] Marker gene: SAMEA44545918_01834 (834 nt)

[0644] Gene product: putative lipoprotein

[0645] Nucleic acid sequence:

[0646]

[0647] Ruminococcus torques

[0648] Amplified fragment: Region from 2028515 to 2028624 of Ruminococcus torques L2-14 draft genome (GenBank: FP929055.1)

[0649] Nucleic acid sequence:

[0650]

[0651] Marker gene: RTO_19540 (927 nt)

[0652] Gene product: tRNA pseudouridine synthase B

[0653] Nucleic acid sequence:

[0654]

[0655] Eubacterium sp. CAG: 274

[0656] Amplified fragment: Region from 7396 to 7500 of Eubacterium sp. CAG: 274 WGS project CBEX01000000 data, contig, whole genome shotgun sequence (GenBank: CBEX010000022.1)

[0657] Nucleic acid sequence:

[0658]

[0659] Marker gene: BN582_00070 (861 nt)

[0660] Gene product: prolipoprotein diacylglyceryl transferase

[0661] Nucleic acid sequence: same as the nucleic acid sequence for marker gene BN582_00070 as described in Embodiment 1.

[0662] Lawsonibacter asaccharolyticus

[0663] Amplified fragment: Region from 627394 to 627504 of Lawsonibacter asaccharolyticus strain OA10 chromosome, complete genome (GenBank: CP091870.1)

[0664] Nucleic acid sequence:

[0665]

[0666] Marker gene: L9O85_02835 (633nt)

[0667] Gene product: TetR / AcrR family transcriptional regulator

[0668] Nucleic acid sequence: same as the nucleic acid sequence for marker gene L9O85_02835 as described in Embodiment 1.

[0669] Oscillibacter sp. CAG: 241

[0670] Amplified fragment: Region from 1167 to 1283 of Oscillibacter sp. CAG: 241 WGS project CBDI01000000 data, contig, whole genome shotgun sequence (GenBank: CBDI010000003.1)

[0671] Nucleic acid sequence:

[0672]

[0673] Marker gene: BN557_01301 (1506 nt)

[0674] Nucleic acid sequence: same as the nucleic acid sequence for marker gene BN557_01301 as described in Embodiment 1.

[0675] Asaccharobacter_celatus

[0676] Amplified fragment: Region from 1205228 to 1205347 of Adlercreutzia equolifaciens subsp. celatus JCM 14811 DNA, complete genome (GenBank: AP024470.1)

[0677] Nucleic acid sequence:

[0678]

[0679] Marker gene: ADLECEL_09850 (1215 nt)

[0680] Gene product: acyl-CoA dehydrogenase

[0681] Nucleic acid sequence:

[0682]

[0683] Butyricimonas virosa

[0684] Amplified fragment: Region from 4511686 to 4511794 of Butyricimonas virosa strain DSM 23226 chromosome, complete genome (GenBank: CP102269.1)

[0685] Nucleic acid sequence:

[0686]

[0687] Marker gene: NQ494_18680 (1347 nt)

[0688] Gene product: TlpA family protein disulfide reductase

[0689] Nucleic acid sequence:

[0690]

[0691] Clostridium sp. CAG: 58

[0692] Amplified fragment: Region from 14734 to 14849 of Clostridium sp. CAG: 58 WGS project CBFK01000000 data, contig, whole genome shotgun sequence (GenBank: CBFK010000011.1)

[0693] Nucleic acid sequence:

[0694]

[0695] Marker gene: BN719_00886 (483 nt)

[0696] Nucleic acid sequence:

[0697]

[0698] Collinsella stercoris

[0699] Amplified fragment: Region from 39110 to 39203 of Collinsella stercoris strain DSM 13279 chromosome, complete genome (GenBank: CP102276.1)

[0700] Nucleic acid sequence:

[0701]

[0702] Marker gene: NQ498_00160 (1653 nt)

[0703] Gene product: C69 family dipeptidase

[0704] Nucleic acid sequence:

[0705]

[0706] Phascolarctobacterium faecium

[0707] Amplified fragment: Region from 2275359 to 2275458 of Phascolarctobacterium faecium JCM 30894 DNA, complete genome (GenBank: AP019004.1)

[0708] Nucleic acid sequence:

[0709]

[0710] Marker gene: corC_2 (1341 nt)

[0711] Gene product: Magnesium and cobalt efflux protein CorC

[0712] Nucleic acid sequence:

[0713]

[0714] Blautia hansenii

[0715] Amplified fragment: Region from 333953 to 334071 of Blautia hansenii DSM 20583 chromosome, complete genome (GenBank: CP022413.2)

[0716] Nucleic acid sequence:

[0717]

[0718] Marker gene: CGC63_01640 (3174 nt)

[0719] Gene product: cell wall-binding protein

[0720] Nucleic acid sequence:

[0721]

[0722] Clostridium spiroforme (Thomasclavelia spiroformis)

[0723] Amplified fragment: Region from 652607 to 652721 of Thomasclavelia spiroformis DSM 1552 chromosome, complete genome (GenBank: CP102275.1)

[0724] Nucleic acid sequence:

[0725]

[0726] Marker gene: NQ543_02855 (1245 nt)

[0727] Gene product: HD domain-containing protein

[0728] Nucleic acid sequence:

[0729]

[0730] Gemella morbillorum

[0731] Amplified fragment: Region from 1033945 to 1034042 of Gemella morbillorum strain NCTC11323 genome assembly, chromosome: 1 (GenBank: LS483440.1)

[0732] Nucleic acid sequence:

[0733]

[0734] Marker gene: rpoB (3819 nt)

[0735] Gene product: DNA-directed RNA polymerase subunit beta

[0736] Nucleic acid sequence:

[0737]

[0738] Actinomyces sp. oral taxon 181

[0739] Amplified fragment: Region from 458 to 544 of Actinomyces sp. oral taxon 181 str. F0379 A_sporaltaxon181-1.0_Cont73.5, whole genome shotgun sequence (GenBank: AMEW01000023.1)

[0740] Nucleic acid sequence:

[0741]

[0742] Marker gene: HMPREF9061_01114 (545 nt)

[0743] Nucleic acid sequence: same as the nucleic acid sequence for marker gene HMPREF9061_01114 as described in Embodiment 1.Example 3: Determination of PerformanceGut bacterial profile is different between UC patients and normal individuals.

[0744] Gut bacterial profile is different between UC patients and normal individuals

[0745] Now referring to FIG. 13A, a diagram of the top bacterial species associated with ulcerative colitis (UC) identified in the study is shown. The lollipop plot in the left panel shows the coefficient of each species with disease calculated by MaAsLin2 with age and gender adjusted. The box in the middle panel indicates the phylum of each species. The bar plot in the right panel demonstrated the proportion of each species present in UC and healthy groups.

[0746] As shown in FIG. 13A, with MaAsLin2 analysis, 48 bacterial species are found to be negatively correlated with UC, i.e., UC-depleted species, namely Collinsella aerofaciens, Clostridium leptum, Ruminococcus torques, Asaccharobacter celatus, Gemmiger formicilis, Fusicatenibacter saccharivorans, Alistipes putredinis, Dorea longicatena, Ruminococcus bromii, Odoribacter splanchnicus, Lachnospira pectinoschiza, Coprococcus comes, Adlercreutzia equolifaciens, Roseburia inulinivorans, Blautia obeum, Eubacterium rectale, Dorea formicigenerans, Alistipes shahii, Bacteroides stercoris, Parabacteroides merdae, Oscillibacter sp. CAG: 241, Akkermansia muciniphila, Alistipes finegoldii, Eubacterium hallii, Oscillibacter sp. 57_20, Roseburia hominis, Phascolarctobacterium faecium, Collinsella stercoris, Bifidobacterium adolescentis, Roseburia intestinalis, Butyricimonas virosa, Alistipes indistinctus, Eubacterium sp. CAG: 38, Anaerostipes hadrus, Bacteroides caccae, Clostridium sp. CAG: 242, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Coprococcus catus, Bacteroides massiliensis, Roseburia faecis, Clostridium sp. CAG: 58, Ruthenibacterium lactatiformans, Bacteroides plebeius, Bilophila wadsworthia, Firmicutes bacterium CAG: 145, Eubacterium ramulus, and Enterorhabdus caecimuris.

[0747] Fifteen bacterial species are found to be positively correlated with UC, i.e., UC-enriched species, namely Gemella morbillorum, Veillonella infantium, Haemophilus sp. HMSC71H05, Actinomyces sp. oral taxon 181, Blautia producta, Lactobacillus mucosae, Enterococcus avium, Veillonella atypica, Eubacterium sulci, Blautia hansenii, Rothia mucilaginosa, Clostridium spiroforme, Tyzzerella nexilis, Veillonella parvula, and Bacteroides fragilis.

[0748] Among the 48 UC-depleted species as described herein, 14 bacterial species are identified as the potential bacterial markers for UC diagnosis, found to be negatively correlated with UC, i.e., UC-depleted species, namely Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, and Ruminococcus torques. The mean relative abundance in normal individuals (%) of the UC-depleted species are shown in Table 2.2.

[0749] Among the 15 UC-enriched species as described herein, four species are also identified as the potential bacterial markers for UC diagnosis, found to be positively correlated with UC, i.e., UC-enriched species, namely Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii. The mean relative abundance in normal individuals (%) of the UC-enriched species are shown in Table 2.3.

[0750] As shown in FIG. 13B, the associations between UC, gender, age and the relative abundance of the 14 UC-depleted bacterial species and the 4 UC-enriched species calculated by general linear model MaAsLin2 are shown. Significant associations (FDR < 0.05) were marked with a plus for positive correlations and a minus for negative correlations, respectively. False discovery rate (FDR) was computed by Benjamini–Hochberg correction. These bacteria can be used to determine the risk of or diagnose ulcerative colitis in a subject.

[0751] As such, bacteria listed in Table 2.2 and Table 2.3 can be used in different combinations to build an assessment model to determine the risk of or diagnose ulcerative colitis in a subject, and whether microbiome restoration therapy or supplementation is required. The relative abundance can be determined by qPCR or ddPCR using a panel of primers, or by metagenomics sequencing to determine the risk of or diagnose ulcerative colitis in the subject.

[0752] Bacteria listed in Table 2.2 can be administered to UC patients for relieving symptoms of UC.Table 2.2: Bacterial Species Enriched in Normal Individuals Compared to UC patients (UC-depleted species) . Table 2.3: Bacterial Species Enriched in UC patients Compared to Normal Individuals (UC-enriched species)

[0753] Performance of machine learning model based on single bacterial marker or different bacterial markers combination using metagenomic data.

[0754] Single bacterial marker or different bacterial markers combination were used in the machine learning model to evaluate the performance of UC diagnosis. The model performance ranged from 0.344 to 0.799 with single bacterial biomarker (No. 1-18 in Table 2.4) . To enhance the model capability for disease diagnosis, different bacterial markers combination was tested, with the AUC ranging from 0.7 to 0.903 in the discovery cohort (No. 19-132 in Table 2.4) . The combination of ten bacterial biomarkers achieved the best diagnostic performance with AUC of 0.903 (No. 19 in Table 2.4) .Table 2.4: Performance of machine learning model based on different bacterial markers combination for Risk Prediction of UC using metagenomic data.

[0755] Performance of machine learning model based on the ten selected bacterial biomarkers using metagenomic data.

[0756] [Corrected under Rule 26, 14.03.2025]Now referring to FIGs. 14A-14B, which show the relative abundance (%) of 18 bacterial species markers (as listed in Table 2.2 and Table 2.3) determined by metagenomic sequencing in UC patients and normal individuals in Hong Kong, China (HK, China) discovery cohort (FIG. 14A) and Hong Kong, China (HK, China) validation cohort (FIG. 14B) , respectively.

[0757] [Corrected under Rule 26, 14.03.2025]To obtain the model with best performance in different population, different combinations of bacterial biomarkers were evaluated. Ten selected bacterial species biomarkers (also referred to as “biomarkers” ) were finally used in the machine learning model, including Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, Blautia hansenii (as listed in Table 2.5) . As shown in FIGs. 14A-14B, the relative abundance of these 10 biomarkers is consistent in HK, China discovery cohort and HK, China validation cohort.

[0758] [Corrected under Rule 26, 14.03.2025]Now referring to FIG. 15A, which shows the receiver operating characteristic (ROC) curve and the area under the curve (AUC) of the machine learning model. AUC of random forest model using the 10 biomarkers determined by metagenomic sequencing in the training set, test set and the Hong Kong, China (HK, China) validation cohort.

[0759] [Corrected under Rule 26, 14.03.2025]The final model using these 10 markers has an Area Under the Curve (AUC) in Receiver Operating Characteristic (ROC) curve of 0.9025 (95%CI: 0.8431-0.9618; sensitivity of 88.06%, specificity of 80.95%with optimized threshold of 0.564) in HK, China discovery cohort (labeled as ‘Testset’ as shown in FIG. 15A) . In HK, China validation cohort, the AUC is 0.8075 with these 10 markers (labeled as ‘Validationcohort’ as shown in FIG. 15A) .

[0760] [Corrected under Rule 26, 14.03.2025]Now referring to FIG. 15B, which shows the receiver operating characteristic (ROC) curve and the area under the curve (AUC) values of the machine learning model (random forest model) using 10 markers determined by metagenomic sequencing in three public datasets including subjects from the United States, Netherlands and China. In the three public datasets, the AUCs of diagnostic model in United States, Netherlands and China are 0.8549, 0.8706 and 0.8213 respectively.

[0761] The results indicate that the machine learning model using the 10 biomarkers is accurate in predicting the risk of ulcerative colitis in a subject.Table 2.5: Bacterial Species Included in the Machine Learning Model for Risk Prediction of UC

[0762] Determination of ulcerative colitis risk using different combinations of bacterial biomarkers

[0763] To determine the risk of or diagnose ulcerative colitis in a subject, the following steps are carried out:(1) obtaining a set of training data by determine the relative abundance of species selected from Table 2.2 or Table 2.3 in a cohort of normal individuals and ulcerative colitis patients;(2) determining the relative abundance of these species in the subject whose risk of ulcerative colitis is to be determined;(3) comparing the relative abundance of these species in the subject with the training data using random forest model; and(4) generating decision trees by random forest from the training data. The relative abundances will be run down the decision trees and generate a risk score. If more than 50%trees in the model consider the subjects have ulcerative colitis, the subject being tested is deemed to be at an increased risk for ulcerative colitis. If less than 50%trees in the model consider the subject as normal individual, the subject being tested is deemed to not have an increased risk for ulcerative colitis.

[0764] Bacterial species determined by droplet digital PCR (ddPCR)

[0765] Now referring to FIGs. 16A-16C. Using the specific primers and probes listed in the Table 2.1, the level of 6 species that are negatively correlated with UC and 4 species that are positively correlated with UC (Table 2.5) was determined by ddPCR in a subgroup of discovery cohort (UC = 205, Controls = 84) .

[0766] Performance of machine learning model based on single bacterial marker or different bacterial markers combination using ddPCR data.

[0767] Single bacterial marker or different bacterial markers combination were used in the machine learning model to evaluate the performance of UC diagnosis. The model performance ranged from 0.403 to 0.789 with single bacterial biomarker (No. 1-10 in Table 2.6) . To enhance the model capability for disease diagnosis, different bacterial markers combination was tested, with the AUC ranging from 0.614 to 0.881 (No. 11-136 in Table 2.6) . The combination of ten selected bacterial biomarkers achieved the best diagnostic performance with AUC of 0.881 (No. 11 in Table 2.6)Table 2.6: Performance of machine learning model based on different bacterial markers combination for Risk Prediction of UC using ddPCR data.

[0768] Performance of Machine Learning Model using ddPCR data

[0769] Ten selected bacterial species biomarkers were used in the machine learning model, including Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, Blautia hansenii (Table 2.5) . The subgroup of discovery cohort (UC = 205, Controls = 84) was randomly divided into training set (UC = 165, Controls = 62) and test set (UC = 40, Controls = 22) .

[0770] FIG. 17 shows the receiver operating characteristic (ROC) curve and the area under the curve (AUC) of the machine learning model. AUC of random forest model using the 10 markers determined by droplet digital PCR. The final model using these 10 markers achieved an AUC of 0.9521 in training set and 0.8807 in the test set (95%CI: 0.7885-0.9729; sensitivity of 85.00%, specificity of 81.82%with optimized threshold of 0.658) .

[0771] UC-depleted bacterial species biomarkers contributed to the beneficial pathways which depleted in UC patients

[0772] FIGs. 18A-18D are stacked bar plots comparing the relative abundance (%) of different bacterial species (i.e, Actinomyces sp. oral taxon 181, Bilophila wadsworthia, Blautia hansenii, Clostridium leptum, Clostridium spiroforme, Fusicatenibacter saccharivorans, Gemella morbillorum, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, and others) in the functional pathways of L-arginine biosynthesis I (via L-ornithine) , L-valine biosynthesis, L-ornithine biosynthesis I and adenosine ribonucleotides de novo biosynthesis, respectively, between healthy controls and UC patients. Each stacked bar plot indicates the contribution of bacteria species biomarkers and other bacteria in each functional pathway.

[0773] As shown in FIGs. 18A-18D, the abundance of functional pathways including Amino acid biosynthesis (L-arginine, L-ornithine, L-valine biosynthesis) and Nucleoside and Nucleotide Biosynthesis (adenosine ribonucleotides de novo biosynthesis) were decresed in the UC patients. These pathways have higher abundance in healthy controls. The higher abundance of these pathways in healthy controls were mainly contributed by the UC-depleted species, such as C leptum, F. saccharivorans, G formicilis, and R. torques. As such, by providing an effective amount of C leptum, F. saccharivorans, G formicilis, and / or R. torques to a UC patient, the abundance of these functional pathways can be increased, which is beneficial in the treatment of UC.Example 4: Risk Prediction for UC on Test Subjects

[0774] In this example, to determine the risk score of ulcerative colitis in a subject, the following steps were carried out:(1) measuring the relative abundance of 10 selected bacterial species biomarkers as shown in Table 2.5 in the subject by metagenomics sequencing or ddPCR;(2) importing the relative abundance data into the pre-trained machine learning model (random forest model) to generate the disease risk score (probability of disease) for the subject;(3) If a subject's risk score (probability of disease) is higher than the optimized cutoff value of 0.564 for metagenomics data or the optimized cutoff value of 0.658 for ddPCR data, the subject is regarded as suffering or at risk of ulcerative colitis. Conversely, if the risk score (probability of disease) is not higher than the optimized cutoff value, the individual is considered not to be suffering from or at risk of ulcerative colitis.

[0775] Table 2.7 shows the risk prediction (probability of disease) for UC on 109 subjects (including known UC patients and controls, as shown in “original group” ) using the machine learning model based on metagenomic data. Table 2.8 shows the risk prediction (probability of disease) for UC on 62 subjects (including known UC patients and controls, as shown in “original group” ) using the machine learning model based on ddPCR data.

[0776] The “original group” column indicates the condition evaluated by health professions by other means (e.g., colonoscopy) in each subject, where “UC” refers to a subject diagnosed with “ulcerative colitis” and “Controls” refers to a subject without ulcerative colitis The risk score (probability of disease or risk prediction) was calculated after analyzing the sample from the subjects according to the disclosed method as described in the preceding examples. If the risk score was higher than the optimized cutoff value, the subject was labeled as “UC” in the “test results” column. Conversely, if the risk score was not higher than the optimized cutoff value, the subject was labeled as “Controls” in the “test results” column.

[0777] As seen in Table 2.7, 34 out of 42 subjects in the original group of “Controls” were correctly identified as “Controls” in the test results using the machine learning model based on metagenomic data, and 59 out of 67 subjects in the original group of “UC” were correctly identified as “UC” in the test results using the machine learning model based on metagenomic data. As seen in Table 2.8, 18 out of 22 subjects in the original group of “Controls” were correctly identified as “Controls” in the test results using the machine learning model based on ddPCR data, and 34 out of 40 subjects in the original group of “UC” were correctly identified as “UC” in the test results using the machine learning model based on ddPCR data.

[0778] The sensitivity is 88.06%and specificity is 80.95%in the machine learning models based on metagenomic data. The sensitivity is 85.00%and specificity is 81.82%in the machine learning models based on ddPCR data.

[0779] These results demonstrate that the machine learning models based on metagenomic data and ddPCR data are both highly accurate in predicting the risk of UC in a subject.Table 2.7: Risk Score (probability of disease) for UC on subjects using the Machine Learning Model based on metagenomic data. Table 2.8: Risk Score (probability of disease) for UC on subjects using the Machine Learning Model based on ddPCR data. Example 5: Risk Prediction for UC patient with different disease activityDisease risk scores generated by machine learning model utilizing ten selected UC bacterial markers can reflect the metabolic dysregulations in patients with UC

[0780] Metabolic dysfunctional score for each subject was determined by calculating the median Bray-Curtis dissimilarity to the reference control group based on the metabolic functional pathways. Now referring to FIG. 19, which shows a plot showing the correlation between functional dysbiosis scores and probability of disease generated by the machine learning model based on the 10 selected UC bacterial species biomarkers according to Table 2.5. The correlation coefficient R and p value were given by Spearman correlation (*, p<0.05) . As shown in FIG. 19, the disease risk scores generated by the machine learning model using the 10 selected UC bacterial markers positively correlated with the dysfunctional scores. This indicates that the machine learning model can reflect the metabolic dysregulations in patients with UC.Diagnostic models based on the ten selected UC bacterial markers can diagnose UC patients in remission stage

[0781] UC is characterized by the alternation of periods of disease activity including flares and remissions. The accuracy and stability of the 10 selected bacterial markers (listed in Table 2.5) in UC patients with different disease activity were evaluated. Using the Mayo score, UC patients were categorized to those with active disease (Mayo>2) and inactive disease (Mayo≤2) . FIGs. 20A-20C showed that the six UC depleted bacterial markers Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Ruminococcus torques, Odoribacter splanchnicus, Bilophila wadsworthia were also decreased in inactive UC patients compared to healthy controls, while the four UC enriched bacterial markers Gemella morbillorum, Blautia hansenii, Actinomyces sp. oral taxon 181, Clostridium spiroforme were increased in inactive UC patients compared to healthy controls. FIG. 20D showed that the disease risk scores generated by the machine learning model using the 10 selected UC bacterial markers showed no significant difference between UC patients with inactive (N = 110) and active (N = 12) status. FIG. 20E showed that the model achieved an AUC of 0.885 in classifying UC patients in remission from controls. This indicates that the machine learning model can be used to determine the risk of ulcerative colitis in an individual regardless of disease activity.

[0782] Among the 10 selected bacterial markers, C. leptum, F. saccharivorans, O. splanchnicus and G formicili showed lower levels in active UC patients compared to inactive UC patients, whereas Actinomyces sp. oral taxon 181 and C. spiroforme exhibited higher levels in active UC patients than in inactive UC patients. As such, these species are useful as bacterial markers for monitoring disease activity in UC patients. For example, a machine learning model can be trained in a cohort of UC patients with active and inactive disease utilizing a panel of markers comprising these species in order to be used as a biomarker for disease monitoring and for predicting disease flare.Example 6: Differentiating UC from other diseasesDiagnostic models based on the ten selected UC bacterial markers can differentiate UC from other gastrointestinal (GI) and non-GI diseases

[0783] In light of shared microbiota alterations across various diseases, it is important to verify disease specificity for the identified bacteria markers, thereby ensuring a low false positive rate for IBD diagnosis. For this purpose, several non-IBD disease datasets were assessed, consisting of subjects with gastrointestinal diseases (n=439) including colorectal cancer (CRC, n=160) , colorectal adenomas (CA, n=162) , irritable bowel syndrome (diarrhea subtype, IBS-D, n=117) and non-gastrointestinal diseases (n=291) including obesity (Body mass index>28; n=148) and cardiovascular disease (CVD; n=143) .

[0784] Now referring to FIGs. 21A-21B. FIG. 21A is a plot showing ROC curves of classifying UC patients from patients with other GI diseases, including CA (n=162) , IBS-D (n=117) , and CRC (n=160) . As shown in FIG. 21A, the UC diagnostic model / machine learning model (random forest model) using the 10 selected bacterial markers determined by metagenomic sequencing can differentiate UC subjects from patients with other GI diseases with an AUROCs of 0.8893, 0.8649, 0.7575 for colorectal adenomas (CA) , irritable bowel syndrome (diarrhea subtype, IBS-D) and colorectal cancer (CRC) , respectively. FIG. 21B is the comparison of probability of disease (risk score) generated by the model based on the 10 selected UC bacterial species biomarkers for UC patients and patients with other GI diseases. P values were calculated using the Wilcoxon rank-sum test (p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001; ns, no significance) . As shown in FIG. 21B, the probability of disease (risk score) for patients with other GI diseases is significantly lower than the probability of disease for UC patients, showing that the machine learning model is accurate in discriminating UC patients from patients with other GI diseases. FIG. 21C is a plot showing the ROC curve of classifying UC patients from patients with non-IBD GI diseases, where the UC diagnostic model differentiated UC subjects from patients with other GI diseases with an average AUROC of 0.8348.

[0785] Now referring to FIGs. 22A-22B. FIG. 22A is a plot showing ROC curves of classifying UC patients from patients with other non-GI diseases, including obesity (n=148) and CVD (n=143) , and the comparison of their probability of disease generated by model based on the 10 selected UC bacterial species biomarkers. As shown in FIG. 22A, the same model can also differentiate UC subjects from patients with other non-GI diseases, including obesity (n=148) and cardiovascular disease (CVD) (n = 143) , with AUROCs of 0.8436 and 0.8996 for obesity and CVD, respectively. FIG. 22B is a plot showing the comparison of probability of disease (risk score) generated by model based on 10 selected UC bacterial species biomarkers among UC and other non-GI diseases. P values were calculated using the Wilcoxon rank-sum test (p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001; ns, no significance) . As shown in FIG. 22B, the probability of disease (risk score) for patients with other non-GI diseases is significantly lower than the probability of disease for UC patients, showing that the machine learning model is accurate in discriminating UC patients from patients with other non-GI diseases.

[0786] FIG. 22C is a plot showing the ROC curve of classifying UC patients from patients with all other non-IBD diseases (GI and non-GI diseases) (n = 730) using the machine learning model based on the 10 selected bacterial species biomarkers. As shown in FIG. 22C, the machine learning model can differentiate UC subjects from patients with all other non-IBD diseases with an AUROC of 0.8493.

[0787] The results of the model based on the ten selected UC bacterial species biomarkers in differentiating UC from GI diseases, and other non-GI disease from UC are summarized in Table 2.9.Table 2.9: Performance of UC diagnostic model in differentiating UC from other gastrointestinal (GI) and non-GI diseases

[0788] These results suggest that the diagnostic model encompassing the 10 selected UC-associated biomarkers exhibited satisfactory performance in differentiating UC from all other non-IBD diseases, i.e., GI diseases and / or other non-GI diseases.

[0789] Among the ten selected bacterial markers, R. torques is depleted in UC patients compared with other diseases, while C. spiroforme is enriched in UC patients compared with other diseases.Example 7: Method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual

[0790] Now referring to FIG. 23, which shows an example of a method 200 for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual with the following steps involved:

[0791] Step 210: determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species includes one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

[0792] Step 220: comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;

[0793] Step 230: determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value;

[0794] Step 240: if the individual is determined to have an increased risk for or is suffering from UC, optionally treating the individual.EMBODIMENT 3

[0795] Despite recent progress in the understanding of the association between the gut microbiome and inflammatory bowel disease (IBD) , the role of microbiome biomarkers in IBD diagnosis remains underexplored. This study resulted in a microbiome-based diagnostic test for IBD. Utilizing metagenomic data from 5,979 fecal samples with and without IBD from different geographies and ethnicities, microbiota alterations in IBD were identified and ten and nine bacterial species were selected to construct diagnostic models for ulcerative colitis (UC) and Crohn's disease (CD) , respectively. These diagnostic models achieved area under the curves (AUCs) greater than 0.90 for distinguishing IBD patients from controls in the discovery cohort. In the transethnic validation cohorts consisting of IBD patients (UC: n=817, CD: n=1,065) and non-IBD subjects (n=3,600) , the UC and CD diagnostic models maintained satisfactory performance with AUCs of 0.78 and 0.72, respectively. A multiplex droplet digital polymerase chain reaction targeting selected IBD-associated bacterial species was further developed, and the m-ddPCR-based models showed numerically higher performance than fecal calprotectin in discriminating UC and CD from controls (UC: 0.74 vs 0.61; CD: 0.78 vs 0.56) . In this example, universal IBD-associated bacterial species shared across cohorts were identified and the potential applicability of a multi-bacteria biomarker panel as a non-invasive tool for IBD diagnosis was demonstrated.

[0796] Altered gut microbial composition and metabolic pathways have been shown in patients with IBD. However, the role of microbiome biomarkers in IBD diagnosis remains underexplored.

[0797] In this example, a microbiome-based diagnostic test for IBD was developed. Comprehensive analyses of metagenomic datasets were performed to assess the predictability of selected bacterial species for IBD diagnosis, diagnostic models using disease-specific species were constructed and a multiplex droplet digital PCR (m-ddPCR) -based multi-bacteria biomarker panel for IBD diagnosis was developed, as shown in FIG. 24A.

[0798] FIG. 24A-24C are diagrams and charts showing the overview of the study workflow and comparison of fecal microbiome in patients with UC, CD, and controls. Some elements were created with BioRender. com. A total of 5979 samples, including 1884 samples from in-house sequencing datasets and 4095 samples from public datasets, was included in this study. Discovery cohort includes 174 CD patients, 205 UC patients and 118 controls. Independent IBD cohorts includes 139 UC, 190 CD, and 328 Controls. Public datasets include 678 UC, 875 CD, and 1699 Controls. Non-IBD cohorts includes 146 IBS, 230 CA, 372 CRC, 318 obesity, 143 CVD, and 364 corresponding Controls. FIG. 24B are violin plots showing the Shannon index and observed species of fecal microbiome in patients with UC (N=205) , CD (N=174) , and controls (N=118) . Data were shown in boxplots as the median (centre line) , 25th and 75th percentiles (box limits) , and 5th and 95th percentiles (whiskers) . P values were calculated using the two-sided Wilcoxon rank-sum test. FIG. 24C is a principal Coordinates Analysis (PCoA) plot showing the varied microbial composition among groups (174 CD patients, 205 UC patients and 118 controls) . Data were shown in boxplots as the median (centre line) , 25th and 75th percentiles (box limits) , and 5th and 95th percentiles (whiskers) . P values of beta diversity based on Bray-Curtis distance were calculated with PERMANOVA by 999 permutations (Df=2, R2=0.02219, F=5.606, P=0.001) . FIG. 24C’ is a chart of multivariate analysis showing the amount of explained variance and the respective P value determined by PERMANOVA based on Bray-Curtis dissimilarity at species level.

[0799] FIG. 24D is a stacked bar chart showing the relative abundance of the six most abundant phyla in patients with UC (N=205) , CD (N=174) , and controls (N=118) . “Others” represented the phyla that were not shown in the figure. CD: Crohn’s disease; UC: Ulcerative colitis; CA, Colorectal adenoma; CRC, Colorectal cancer; IBS, Irritable bowel syndrome; CVD, Cardiovascular disease. *, p<0.05; **, p<0.01; ***, p<0.001. FIG. 24D’ is a comparison chart of the relative abundance of six phyla among patients with UC, CD, and healthy controls. Boxplots represent the minimum, Q1, median, Q3 and maximum. P values were calculated using the two-sided Wilcoxon rank-sum test. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001, ns no significance. CD: Crohn’s disease; UC: Ulcerative colitis.Results

[0800] Characterization of gut microbiome diversity and taxonomy in IBD

[0801] [Corrected under Rule 26, 14.03.2025]Fecal metagenomics data from 4, 406 samples from 13 IBD cohorts across eight countries and regions were analyzed to identify and gut microbial biomarkers for IBD diagnosis were validated. Specifically, in-house sequencing data from Hong Kong, China as a discovery cohort (Table 3.1) were utilized, and three additional independent in-house cohorts from Hong Kong, China and Australia, as well as nine public datasets from the United States, the Netherlands, China, Spain, Denmark, and United Kingdom as validation cohorts (Table 3.2) were included.

[0802] In the discovery cohort, a total of 1, 175 taxa (3 kingdoms, 14 phyla, 25 classes, 40 orders, 85 families, 226 genera, and 788 species) were identified. At the species level, a total of 674, 637, and 506 bacterial species were identified in UC, CD, and control groups, respectively. Decreased microbial diversity (median: UC 2.73, CD 2.71, controls 3.08; P<0.001) and richness (median: UC 82, CD 86.5, controls 95; P<0.001) were found in patients with UC and CD compared with controls, but there was no statistically significant difference between UC and CD, as show in FIG. 24B. On principal coordinates analysis (PCoA) based on Bray-Curtis distances, the gut microbiome of patients with UC and CD clustered separately from that of controls, and UC patients exhibited a greater distance from controls than CD patients, as shown in FIG. 24C. The presence of IBD accounted for 2.22%of the microbiome variance (P<0.001) , whilst age and gender contributed 0.28% (P=0.107) and 0.29% (P=0.073) , respectively. Factors explaining microbiota variance are shown in FIG. 24C’. There were significant differences in gut microbial communities at the phylum level between patients with IBD and controls characterized by a reduction in Firmicutes and an enrichment of Proteobacteria in IBD. CD patients had lower levels of Bacteroidetes compared with UC and controls (both P<0.001) (FIG. 24D, and FIG. 24D’) . Patients with IBD were found to have harbored reduced microbial diversity compared with controls, and patients with CD and UC showed distinct differences in their gut microbial composition.

[0803] Identification of gut microbiome signatures in UC and CD

[0804] General linear models as implemented in MaAsLin2 was next used to identify differentially abundant bacterial species in UC and CD after filtering out low prevalent species (<10%) followed by adjustment for age and gender. FIGs. 25C-25H are charts showing the differential bacterial species and dysbiosis of metabolic pathways in UC and CD patients compared with controls. FIG. 13A and 1A are charts showing the top bacterial species associated with UC and CD, respectively. Lollipop plot in the left panel showed the coefficient of each species with disease calculated by MaAsLin2 with age and gender adjusted. The box in the middle panel indicates the phylum of each species. The bar plot in the right panel demonstrated the proportion of each species present in UC, CD, and controls groups. FIG. 25C is a chart showing the relative abundance of ten bacterial species biomarkers in UC (N=205) and control group (N=118) , and FIG. 25D is a chart showing the relative nine bacterial species biomarkers in CD (N=174) and control group (N=118) determined by metagenomics. Data were shown in boxplots as the median (centre line) , 25th and 75th percentiles (box limits) , and 5th and 95th percentiles (whiskers) . P values were calculated using the two-sided Wilcoxon rank-sum test. FIGs. 25E-25F are charts showing the performance of model with ten UC and nine CD bacterial species biomarkers for classifying UC and CD patients, respectively, with controls in discovery cohort. Shaded areas of the ROC curves represent the 95%confidence interval of the AUC for the test set. FIGs. 25G-25H are charts showing the Shapley Additive Explanations (SHAP) values of UC and CD bacterial species biomarkers, respectively, for each sample. Each point represents the SHAP value of each biomarker for each sample. The distribution of the points indicates the impact of each biomarker on the model output. The color represents the relative abundance of the biomarkers (yellow high, purple low) . It was found that 126 and 161 species were differentially abundant in UC and CD, respectively, compared with controls (FDR<0.25) . Amongst these, 15 species, including Bacteroides fragilis, Veillonella parvula, Tyzzerella nexilis, Clostridium spiroforme, Rothia mucilaginosa, Blautia hansenii were enriched in UC (FDR<=0.1, coefficient>=0.1) , whereas 48 species including Collinsella aerofaciens, Clostridium leptum, Ruminococcus torques, Asaccharobacter celatus, Gemmiger formicilis, Fusicatenibacter saccharivorans were depleted in the gut of patients with UC (FDR<0.05, coefficient<=-0.45) , as shown in FIG. 13A. In CD, 58 bacterial species were either enriched or depleted compared with controls (FDR<0.001 and |coefficient|>=0.5) . In particular, certain bacterial species with proposed anti-inflammatory properties, including Eubacterium hallii, Blautia obeum, Roseburia inulinivorans, and Eubacterium rectale, were depleted in CD, as shown in FIG. 1A. In addition, E. coli and some Streptococcus species were enriched in the gut of patients with CD but not UC. Altogether, these data highlighted the presence of disease-specific bacterial species in UC and CD.

[0805] Development of metagenomics-based diagnostic models for UC and CD diagnosis

[0806] A five-fold cross-validation with all discriminative bacterial species were next performed to construct diagnostic models. Stable classification performances were achieved by utilizing seven and eight bacterial features in UC and CD, respectively, as shown in FIG. 25K and FIG. 25L. As shown in FIG. 25K, a total of 125 species features were used in the UC diagnostic model. The vertical dotted line in x=7 represented the minimum number of features to maintain a relatively stable performance of the model (horizontal dotted line, AUC= 0.8937) . As shown in FIG. 25L, a total of 161 species features were used in the CD diagnostic model. The vertical dotted line in x=8 represented the minimum number of features to maintain relatively stable performance of the model (horizontal dotted line, AUC= 0.9096) . After accounting for functional properties of selected bacteria and optimal distribution of enriched and depleted bacterial species, ten bacterial species were selected as biomarkers for UC (4 enriched: Gemella morbillorum, Blautia hansenii, Actinomyces sp. oral taxon 181, Clostridium spiroforme; 6 depleted: Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Ruminococcus torques, Odoribacter splanchnicus, Bilophila wadsworthia) and nine bacterial species as biomarkers for CD (3 enriched: Bacteroides fragilis, Escherichia coli, Actinomyces sp. oral taxon 181; 6 depleted: Roseburia inulinivorans, Blautia obeum, Lawsonibacter asaccharolyticus, Roseburia intestinalis, Dorea formicigenerans, Eubacterium sp. CAG: 274) (FIG. 25C and FIG. 25D) . Amongst them, Actinomyces sp. oral taxon 181, and C. spiroforme were reported for the first time to be enriched in UC whereas Actinomyces sp. oral taxon 181, L. asaccharolyticus and Eubacterium sp. CAG: 274 were novel bacterial species associated with CD. Using Random Forest algorithm, the ten selected bacterial species discriminated patients with UC from controls with an AUC of 0.95 (95%CI: 0.92-0.98) in the training set and 0.90 in the test set (95%CI: 0.84-0.96; sensitivity 88.06%, specificity 80.95%) . F. saccharivorans, C. leptum, and G. formicilis were the top three key discriminatory bacterial species in the UC model (FIG. 25E and FIG. 25G) . In CD, nine bacterial species showed an AUC of 0.95 (95%CI: 0.92-0.98) and 0.94 (95%CI: 0.89-0.98; sensitivity 88.33%, specificity 89.47%) in discriminating CD from controls in the training set and test set, respectively. B. obeum, L. asaccharolyticus, R. inulinivorans (depleted in CD) and Actinomyces sp. oral taxon 181 and E. coli (enriched in CD) were the top-ranking bacteria in our model (FIG. 25F and FIG. 25H) .

[0807] Functional and metabolic pathways of bacterial species in UC and CD

[0808] Apart from altered microbial composition and taxonomy, it was found that metabolic functions were dysregulated in UC and CD. Using MaAsLin2 comparison analysis, 545 metabolic pathways were identified, consisting of 244 and 315 differential pathways in UC and CD, respectively, compared with controls (FDR<0.05) .

[0809] Those pathways involved in amine and polyamine degradation, and fatty acid and lipid biosynthesis were significantly enriched in UC and CD patients compared with controls. UC-and CD-enriched bacterial species biomarkers were positively correlated with disease-enriched metabolic pathways, while UC-and CD-depleted bacterial species biomarkers showed negative correlations (FIGs. 26A-26D) . FIGs. 26A-26D are charts showing the differential functional pathways between UC / CD patients and healthy controls, and their correlation with bacterial species biomarkers. Differential functional pathways determined by MaAsLin2 with age and gender adjusted. The correlation coefficient and P value between ten UC or nine CD bacterial species biomarkers and differential functional pathways were given by Spearman correlation. *, p<0.05; **, p<0.01; ***, p<0.001.

[0810] The findings from stratified analyses revealed that pathways belonging to amino acid biosynthesis (L-arginine, L-ornithine, and L-valine biosynthesis) were mainly contributed by bacterial species depleted in UC patients, including C leptum, F. saccharivorans, G formicilis, and R. torques. Depletion of these key bacterial species in UC was also associated with a significant decrease in the abundance of their respective functional pathways (FIGs. 18A-18D for UC) . Similarly, bacterial species, including B. obeum, R. inulinivorans, D. formicigenerans, Eubacterium sp. CAG: 274, and R. intestinalis also contributed to pathways of amino acid biosynthesis (L-tryptophan biosynthesis) , carbohydrate degradation (starch degradation) , cofactor, carrier and vitamin biosynthesis (Thiamine phosphate formation from purithiamine and exythiamine) in controls. In CD, there was a shift in main contributors of these pathways characterized by a predominance of E. coli instead of a diversified bacteria profile (FIGs. 6A-6D for CD) . Variations in metabolic pathways of the bacteria suggested that altered microbiota-mediated metabolic capabilities may be essential for IBD development.

[0811] To examine the role of bacterial biomarkers in mediating metabolic functions, a dysfunctional score for each subject was developed by calculating the median Bray-Curtis dissimilarity to controls using metabolic pathway profiles. Distribution and level of dysfunctional scores in patients with UC and CD were different from that of controls, implying dysfunctional changes in IBD (FIGs. 26E-26F) . The probability of disease determined by diagnostic models showed a positive correlation with dysfunctional scores in UC and CD implying that the bacterial biomarkers could reflect the metabolic dysregulations in patients with IBD (FIG. 19 and FIG. 7) .

[0812] Diagnostic performance of multi-bacteria biomarker panel in relation to host inflammation

[0813] The accuracy and stability of the multi-bacteria biomarker panel in IBD with different disease activities was next evaluated. Using the Mayo score and Crohn’s disease activity index (CDAI) , patients were categorized into active (UC: Mayo>2; CD: CDAI>150) and inactive disease (UC: Mayo≤2; CD: CDAI≤150) . There was differential abundance of CD-and UC-associated bacterial species in patients with inactive IBD compared with controls (6 decreased, 4 increased in UC; 6 decreased, 3 increased in CD) . Moreover, levels of some bacterial species changed with disease activity. FIGs. 27A-27F are charts showing the Relative abundance of bacterial species biomarkers in healthy controls and patients at inactive and active status. Shaded areas of the ROC curves represent the 95%confidence interval of the AUCROC for each cohort. Boxplots represent the minimum, Q1, median, Q3 and maximum. The gray diamond represents the mean value. P values were given by the two-sided Wilcoxon rank sum test. *p<0.05, **p<0.01, ***p<0.001, NS no significance. In this example, the relative abundance of six UC-depleted species were lower in patients with active UC compared with inactive UC, as shown in FIGs. 27A and 27B; and R. inulinivorans, B. obeum, L. asaccharolyticus, D. formicigenerans and Eubacterium sp. CAG: 274 were lower in patients with active CD compared with inactive CD, whereas some disease-enriched bacteria, such as G. morbillorum, were higher in active UC, and B. fragilis, E. coli, and Actinomyces sp. oral taxon 181 were higher in active CD, as shown in FIGs. 27D and 27E, implying these bacterial species may be involved in disease activity and severity. Diagnostic models were next utilized to calculate a probability value for disease risk and found that disease probability showed no significant difference between inactive and active IBD. Our model was able to distinguish inactive UC and CD patients from controls with AUC of 0.89 and 0.84, respectively, as shown in FIGs. 27C and FIG. 27F. These data suggested that the bacterial biomarkers may not just be a consequence of inflammation but also contribute to, or may reflect underlying disease pathogenesis.

[0814] Validation of multi-bacteria biomarker panel diagnostic model in independent cohorts

[0815] [Corrected under Rule 26, 14.03.2025]Next, two independent datasets from Hong Kong, China (139 UC, 139 controls; 92 CD, 108 controls) and Australia (98 CD, 81 controls) and three public datasets from the United States (53 UC, 68 CD, 34 controls) , Netherlands (23 UC, 20 CD, 22 controls) and China (25 UC, 15 controls; 48 CD, 54 controls) were analyzed.

[0816] [Corrected under Rule 26, 14.03.2025]FIGs. 28A-28C are charts showing the abundance and prevalence of bacterial species biomarkers and the performance of diagnostic models in cohorts from different ethnicities and regions. FIG. 28A and FIG. 28B are charts showing the signature of bacterial species biomarkers for UC and CD diagnosis in patients and healthy individuals of discovery cohort, validation cohort, and three downloaded public datasets. The abundance of species was normalized to log2 fold change (log2FC) relative to the mean of control samples. P values were calculated using the two-sided Wilcoxon rank-sum test. P values were then converted to -log10 (P-value) after using Benjamini–Hochberg correction to control for multiple testing. *p<0.05, **p<0.01, ***p<0.001. Prevalence indicates the proportion of bacterial presence in UC, CD, and healthy group of each cohort. FIG. 28C are charts showing the probability of disease calculated by the random forest model between UC / CD patients and healthy controls in Hong Kong, China discovery cohort, validation cohort from Hong Kong, China and Australia, and public datasets (USA, Netherland, and China) . Boxplots represent the minimum, Q1, median, Q3 and maximum. P values were calculated using the two-sided Wilcoxon rank-sum test. FIG. 28D is a chart showing the correlation among the ten UC bacterial species markers. UC-depleted bacteria were labelled with green color while the UC-enriched ones were labeled with yellow color. FIG. 28E is a chart showing the correlation among the nine CD bacterial species markers. CD-depleted bacteria were labelled with green color while the CD-enriched ones were labeled with orange color. Grids in red indicated positive correlation, while grids in blue indicated negative correlation. The correlation coefficient and P value were given by Spearman correlation.

[0817] The abundance and prevalence of the bacterial biomarkers in the independent cohorts and public datasets were consistent with those reported in the discovery cohort (FIGs. 28A and 28B) .

[0818] [Corrected under Rule 26, 14.03.2025]FIGs. 29A-29M are charts showing the performance of model with bacterial species biomarkers to discriminate patients with UC or CD from controls in independent cohorts and public datasets. FIG. 29A is a chart showing the performance of model with ten UC selected bacterial species biomarkers for classifying UC patients with controls in Hong Kong, China validation cohort. FIGs. 29B-29C are charts showing the performance of model with nine CD selected bacterial species biomarkers for classifying CD patients with controls in Hong Kong, China and Australia validation cohorts. FIGs. 29B-29C are charts showing the performance of model with the selected bacterial species biomarkers for classifying UC or CD patients with controls in the three downloaded public datasets. FIG. 29D is a chart showing the performance of model with the selected bacterial species biomarkers for classification of UC patients with controls in the three downloaded public datasets.

[0819]

[0820] FIG. 29E is a chart showing the performance of model with the selected bacterial species biomarkers for classification of CD patients with controls in the three downloaded public datasets. FIGs. 29F-29G are charts showing the associations between disease group, geography, ethnicity and the relative abundance of bacterial species biomarkers were calculated by MaAsLin2 in all IBD cohorts. Positive associations were colored by red, while negative associations were colored by blue. Significant associations (FDR<0.05) were marked with a plus for positive associations and a minus for negative associations, respectively. False discovery rate (FDR) was computed by Benjamini–Hochberg correction. FIG. 29H is a chart showing the Performance of model with the selected bacterial species biomarkers for classifying UC patients (n=817) with controls (n=1746) in all UC validation cohorts. FIG. 29I Performance of model with the selected bacterial species biomarkers for classifying CD patients (n=1065) with controls (n=1873) in all CD validation cohorts. FIG. 29J is a chart showing the model performance in distinguishing treated and  UC patients from controls in two downloaded public datasets. FIG. 29K is a chart showing the model performance in distinguishing treated and  CD patients from controls in two downloaded public datasets. FIG. 29L is a chart showing the model performance in distinguishing UC patients from controls compared with fecal calprotectin test in two downloaded public datasets. FIG. 29M is a chart showing the model performance in distinguishing CD patients from controls compared with fecal calprotectin test in two downloaded public datasets. Shaded areas of the ROC curves represent the 95%confidence interval of the AUC for each cohort.

[0821] [Corrected under Rule 26, 14.03.2025]In this example, the diagnostic models as disclosed herein showed desirable performances in classifying IBD from controls with AUCs of 0.81 (95%CI: 0.76-0.86) for UC, as shown in FIG. 29A, and 0.83 (95%CI: 0.77-0.89) for CD in Hong Kong, China cohorts, as shown in FIG. 29B, and 0.73 (95%CI: 0.65-0.80) for CD in Australian cohort, as shown in FIG. 29C. Using datasets from United States, Netherlands, and China, the diagnostic model achieved AUCs of 0.85, 0.87, and 0.82, respectively, for UC diagnosis (FIG. 29D, FIG. 28C) , and AUCs of 0.89, 0.86, and 0.97, respectively, for CD diagnosis (FIG. 29E, FIG. 28C) . Bacteria ecological network analysis showed co-occurring correlations between depleted bacterial species in IBD, and stable co-excluding correlations among disease-depleted and disease-enriched species in almost all cohorts (FIGs. 28D-28E) . To further validate the diagnostic models, additional metagenomic datasets from the United States, Netherlands, Spain, Denmark and United Kingdom (577 UC, 739 CD, 1574 controls) were further utilized. By integrating data from all IBD cohorts and adjusting for geography and ethnicity, the multi-bacteria biomarker panel remained significantly different between IBD and controls (FIGs. 29F-29G) . The UC model maintained an overall AUC of 0.82, and the CD model achieved an AUC of 0.76 in all validation cohorts (FIGs. 29H-29I) . Altogether, these results highlighted the robustness of diagnostic performance of our multi-bacteria biomarker panel across different regions and ethnicities.

[0822] Given that drugs used to induce and maintain disease remission in IBD can alter diversity and composition of the gut microbiome, the models were tested to determine whether they would be affected by treatment. Using metagenomic data from two public datasets (United States and Netherlands) whereby treatment data were available for IBD cases, the UC model was found to discriminate   (n=14) and treated-UC patients (n=62) from controls (n=56) with AUCs of 0.74 and 0.89 respectively, whilst the CD model discriminated   (n=20) and treated-CD patients (n=66) from controls (n=56) with AUCs of 0.89 and 0.88, respectively, suggesting that the performance of the multi-bacteria biomarker panel is unlikely to be influenced by treatment (FIGs. 29J-29K) .

[0823] To compare the models with a commonly used IBD screening test-fecal calprotectin, data from two public datasets (the United States and the Netherlands) were used whereby fecal calprotectin data were available (46 UC, 65 CD, 42 controls) . It was found that the diagnostic models based on bacteria biomarkers had a numerically higher AUC than fecal calprotectin for the diagnosis of UC (AUC 0.85 vs 0.81) and CD (AUC 0.87 vs 0.79) (FIGs. 29L-29M) . The multi-bacteria biomarker panel also showed a higher sensitivity (72%vs 54%for CD; 67%vs 57%for UC) and specificity (95%vs 86%for CD; 88%vs 86%for UC) than fecal calprotectin.Specificity of IBD diagnostic models based on multi-bacteria biomarker panel

[0824] [Corrected under Rule 26, 14.03.2025]In light of shared microbiota alterations across various diseases, the study in this example verifies disease specificity for bacterial biomarkers. Several in-house non-IBD disease metagenomic datasets from Hong Kong, China were assessed, which included subjects with various gastrointestinal diseases (n=439) including colorectal cancer (CRC, n=160) , colorectal adenomas (CA, n=162) , irritable bowel syndrome (diarrhea subtype, IBS-D, n=117) and non-gastrointestinal diseases (n=291) including obesity (n=148) and cardiovascular disease (CVD; n=143) .

[0825] [Corrected under Rule 26, 14.03.2025]FIGs. 30A-30B are charts showing the relative abundance of bacterial species biomarkers in UC / CD and other non-IBD disease groups in Hong Kong, China cohort. FIG. 30A are charts showing the relative abundance of 10 UC bacterial species biomarkers in UC and other non-IBD disease group. FIG. 30B are charts showing the relative abundance of 9 CD bacterial species biomarkers in CD and other non-IBD disease group. Boxplots represent the minimum, Q1, median, Q3 and maximum. The gray diamond represents the mean value. P values were calculated using the two-sided Wilcoxon rank-sum test. *p<0.05, **p<0.01, ***p<0.001, NS no significance. CD: Crohn’s disease; UC: Ulcerative colitis; IBS-D, Irritable bowel syndrome (diarrhea subtype) ; CA, Colorectal adenomas; CRC, Colorectal cancer; CVD, Cardiovascular disease.

[0826] Among the UC bacterial biomarkers, a depletion in R. torques and enriched C. spiroforme were unique to UC patients compared with all other non-IBD diseases (FIG. 30A) . Among the CD bacterial biomarkers, a depletion of R. inulinivorans and B. obeum, and an increase of B. fragilis and E. coli was specifically associated with CD (FIG. 30B) .

[0827] FIGs. 31A-31E shows the performance of model with bacterial species biomarkers to discriminate patients with UC or CD from other subjects with and without gastrointestinal disorders in international cohorts. FIG. 31A is a chart showing the composition of international multi-disease datasets from different countries and regions. FIG. 31B is a chart showing the comparison of their probability of disease generated by UC model based on ten UC bacterial species biomarkers in controls (n=2391) , CVD (n =143) , obesity (n=318) , CA (n=230) , CRC (n=372) , IBS-D (n=146) and UC patients (n =817) . Data were shown in boxplots as the median (centre line) , 25th and 75th percentiles (box limits) , and 5th and 95th percentiles (whiskers) . FIG. 31C is a chart showing the performance of UC model in classifying UC patients (n = 817) from other non-IBD subjects (n = 3600) . FIG. 31D is a chart showing the comparison of their probability of disease generated by CD model based on nine CD bacterial species biomarkers in controls (n=2391) , CVD (n = 143) , obesity (n=318) , CA (n=230) , CRC (n=372) , IBS-D (n=146) and CD patients (n = 1065) . Data were shown in boxplots as the median (centre line) , 25th and 75th percentiles (box limits) , and 5th and 95th percentiles (whiskers) . FIG. 31E is a chart showing the performance of CD model in classifying CD patients (n = 1065) from other non-IBD subjects (n = 3600) . Boxplots represent the minimum, Q1, median, Q3 and maximum. P values were calculated using the two-sided Wilcoxon rank-sum test. ****, p<0.0001. Shaded areas of the ROC curves represent the 95%confidence interval of the AUC for each cohort. CD: Crohn’s disease; UC: Ulcerative colitis; IBS-D, Irritable bowel syndrome (diarrhea subtype) ; CA, Colorectal adenomas; CRC, Colorectal cancer; CVD, Cardiovascular disease.

[0828] To validate the models in an internationally diverse non-IBD cohort, 843 extra metagenomic data of non-IBD cohorts (212 CRC, 68 CA, 29 IBS-D, 170 obesity, and 364 controls) from Austria, France, Germany, Japan, the United States, and Denmark (FIG. 31A) were included. The probability of disease generated by the models showed significant differences between IBD and non-IBD subjects (FIG. 31B, FIG. 31D) . The UC diagnostic model discriminated patients with UC from non-IBD subjects with an AUC of 0.78 (FIG. 31C) , whereas the CD diagnostic model distinguished CD from non-IBD subjects with an AUC of 0.72 (FIG. 31E) . Overall, these results demonstrate that our multi-bacteria panel was specific to UC and CD.

[0829] Development of a general IBD model based on the multi-bacteria biomarker panel

[0830] Since there is an unmet need for universal biomarkers to differentiate IBD from non-IBD subjects, an IBD model using a total 18 bacterial species identified for UC and CD (UC: 10 species; CD: 9 species; 1 species overlapping in UC and CD) was developed.

[0831] FIGs. 32A-32C are charts showing the performance of general IBD model in classifying IBD from and non-IBD subjects. FIG. 32A is a chart showing the ROC of general IBD model in classifying IBD from controls and non-IBD in test set, IBD validation cohort, and non-IBD cohort. FIG. 32B is a chart showing the prototypical standards for reporting diagnostic accuracy studies (STARD) diagram reporting the flow of participants in independent international IBD cohort (IBD=1882, Controls=2027) . FIG. 32C is a chart showing the comparison of diagnostic performance of general IBD model and fecal calprotectin in classifying IBD from and IBS subjects. Shaded areas of the ROC curves represent the 95%confidence interval of the AUC for each cohort.

[0832] The discriminative power of the IBD model achieved an AUC of 0.91 (95%CI: 0.85-0.97) with a sensitivity of 92%and specificity of 75%in the discovery cohort, and an AUC of 0.81 (95%CI: 0.80-0.82) with a sensitivity of 78%, and specificity of 70%in the validation cohort (controls=2, 027, IBD=1, 882) . In the international multi-disease cohort (IBD=1, 882; non-IBD=3, 600) , which included subjects with GI and non-GI diseases, the model could differentiate IBD from non-IBD with an AUC of 0.77 (95%CI: 0.75-0.78) (FIG. 32A and FIG. 32B) . In a pilot cohort, direct comparison of the multi-bacteria biomarker panel with fecal calprotectin in samples from aforementioned in-house cohorts (36 UC, 36 CD, and 36 IBS patients) was performed. The multi-bacteria biomarker panel (AUC=0.91; sensitivity: 79%, specificity: 92%) showed numerically higher performance than fecal calprotectin (AUC=0.86; sensitivity: 68%, specificity: 89%) in distinguishing patients with IBD from IBS (FIG. 32C) .

[0833] Developing multiplex droplet digital PCR (m-ddPCR) -based multi-bacteria biomarker panel

[0834] To translate the metagenome-derived multi-bacteria biomarker panel into a simple and affordable clinical tool, a m-ddPCR-based method was developed to quantify selected bacterial species in fecal samples.

[0835] FIGs. 33A-33C show the panel design of multiplex droplet digital PCR and correlation between the abundance of bacterial species biomarkers determined by metagenomics and multiplex droplet digital PCR method. FIG. 33A are charts showing the panel design of multiplex droplet digital PCR for UC and CD bacterial species markers. FIG. 33B and FIG. 33C are charts showing the correlation between the abundance of the ten UC bacterial species biomarkers and nine CD bacterial species biomarkers, respectively, which were determined by metagenomics and multiplex droplet digital PCR method. The correlation coefficient and P value were given by Spearman correlation.

[0836] Three reactions were designed to measure abundance of bacterial species and to ensure there was no cross reaction among the primers and probes of targeted species (FIG. 33A) . The abundance of bacterial species was quantified by m-ddPCR and correlated them with abundance of species generated from metagenomics in UC (205 UC, 84 controls) and CD (172 CD, 86 controls) . Quantification by metagenomic sequencing and m-ddPCR showed strong correlations (Spearman r=0.34-0.73 for UC biomarkers, and Spearman r=0.49-0.94 for CD biomarkers) , indicating that both measurements were reliable and consistent (FIG. 33B and FIG. 33C) .

[0837] [Corrected under Rule 26, 14.03.2025]FIGs. 34A-34H are related to the bacterial species biomarkers in patients and healthy individuals determined by multiplex droplet digital PCR. FIG. 34A (and FIGs. 16A-16C) are charts showing the relative abundance of ten bacterial species biomarkers in UC and healthy control group in discovery cohort (205 UC; 84 controls) . Boxplots represent the minimum, Q1, median, Q3 and maximum. The gray diamond represents the mean value. FIG. 34B is a chart showing the diagnostic performance to discriminate patients with UC from healthy control with ten bacterial species biomarkers determined by m-ddPCR in discovery cohort (testset, N=62) and Hong Kong, China cohort (N=108) . FIG. 34C (and FIGs. 4A-4C) are charts showing the relative abundance of nine bacterial species biomarkers in CD and healthy control group in discovery cohort (172 CD; 86 controls) . Boxplots represent the minimum, Q1, median, Q3 and maximum. The gray diamond represents the mean value. FIG. 34D is a chart showing the diagnostic performance to discriminate patients with CD from healthy control with nine bacterial species biomarkers determined by m-ddPCR in discovery cohort (testset, N=66) , Hong Kong, China cohort (N=153) and Australia cohort (N=177) . FIG. 34E are charts showing the diagnostic performance of fecal calprotectin test and UC model with ten bacterial species biomarkers determined by m-ddPCR in Canada cohort (100 UC, 53 Controls) and Taiwan, China cohort (40 UC, 40 Controls) . FIG. 34F are charts showing the diagnostic performance of fecal calprotectin test and CD model with ten bacterial species biomarkers determined by m-ddPCR in Canada cohort (100 CD, 53 Controls) and Taiwan, China cohort (40 CD, 40 Controls) . FIG. 34G are charts showing the comparison of the probability of disease calculated by UC / CD model using m-ddPCR data and fecal calprotectin results between UC / CD patients at inactive and active status, and healthy controls in Canada and Taiwan, China cohort. FIG. 34H is a chart showing the diagnostic performance of fecal calprotectin test and UC model with ten bacterial species biomarkers determined by m-ddPCR in distinguishing inactive UC patients (N=81) and healthy controls (N=93) . Shaded areas of the ROC curves represent the 95%confidence interval of the AUCROC for each cohort. P values were calculated using the two-sided Wilcoxon rank-sum test. *p<0.05, **p<0.01, ***p<0.001, ****p<0.001, NS no significance. CD: Crohn’s disease; UC: Ulcerative colitis; m-ddPCR, multiplex droplet digital PCR.

[0838] FIG. 35A is a chart showing the difference of probability of disease (POD) calculated by metagenomics-based model and m-ddPCR-based model in UC.

[0839] FIG. 35B is a chart showing the difference of probability of disease (POD) calculated by metagenomics-based model and m-ddPCR-based model in CD.

[0840] From the m-ddPCR results, significant differences in the six depleted bacterial species and one enriched bacterial species were found in UC compared with controls, while the other three enriched bacterial species showed an increasing trend in UC patients compared with controls (FIG. 34A) . Random Forest diagnostic model constructed using m-ddPCR data yielded an AUC of 0.88 (sensitivity 85.0%; specificity 81.8%) for UC diagnosis in the discovery cohort (FIG. 34B) . In CD, significant differences in abundance of six depleted and three enriched bacterial species in CD was also identified when compared with controls (FIG. 34C) . By using m-ddPCR data of the multi-bacteria biomarker panel, a Random Forest diagnostic model was constructed, which showed an AUC of 0.87 for CD (sensitivity 90.2%; specificity 76.0%) (FIG. 34D) . Furthermore, the UC model achieved AUC of 0.89 (FIG. 34B) , while the CD model achieved AUC of 0.75 and 0.73 (FIG. 34D) in the independent validation cohorts. The probability of disease (POD) values derived from the metagenomic model and m-ddPCR model were compared, and it was found that differences of POD values between the two models were -0.03 (95%CI: -0.05, -0.01) in UC and -0.07 (95%CI: -0.08, -0.05) in CD (FIG. 35A and FIG. 35B) , indicating that findings from m-ddPCR were consistent with that of metagenomics.

[0841] [Corrected under Rule 26, 14.03.2025]To compare the diagnostic performance of the multi-bacteria biomarker panel with that of fecal calprotectin, both tests were performed on fecal samples from two independent cohorts from Canada (100 UC, 100 CD, 53 controls) and Taiwan, China (40 UC, 40 CD, 40 controls) with analysis blinded relative to each test. It was demonstrated that the multi-bacteria biomarker panel showed better performance than that of fecal calprotectin in UC (FIG. 34E) , Canada: AUC=0.74 vs 0.63; Taiwan, China: 0.79 vs 0.57) and performed slightly better than or comparable to fecal calprotectin in CD model (FIG. 34F) , Canada: AUC=0.77 vs 0.75; Taiwan, China: 0.71 vs 0.71) . In the subgroup analysis, the CD multi-bacteria biomarker panel could differentiate active and inactive CD patients from controls. In addition, the UC multi-bacteria biomarker panel showed higher performance than fecal calprotectin in discriminating inactive UC from controls (AUC=0.78 vs 0.56) (FIG. 34G and FIG. 34H) .

[0842] Discussion

[0843] In the present study, IBD-associated gut microbiome and its ability to distinguish IBD from non-IBD subjects were comprehensively assessed. Through extensive and rigorous validation, whereby data were generated from eight countries and regions and data used for training were separated from that for the testing, disease-specific bacteria species were identified and a non-invasive microbiome-based tool for IBD diagnosis was developed. In particular, metagenomics-based model trained on selected bacteria species from multiple studies maintained an AUC of 0.81 (95%CI: 0.80-0.82) in distinguishing patients with IBD from controls, which is above the threshold (AUC=0.80) that is generally considered clinically useful. The multi-bacteria biomarker panel also showed numerically higher diagnostic performance than fecal calprotectin, a standard non-invasive clinical test for inflammation commonly used in IBD.

[0844] To date, microbial identification and analysis using mass spectrometry, 16S amplicon and metagenomics sequencing face challenges including high cost, complex operations and interpretation procedures. Targeted detection technologies, such as fluorescent quantitative PCR, nucleic acid hybridization, fluorescent probe labeling, and digital PCR, have been applied in pathogen detection, environmental monitoring, and liquid biopsy, but there is limited research applying these techniques for disease diagnosis. Herein, existing IBD microbiome research was taken a step further by translating metagenomic-generated data to bacteria detection based on m-ddPCR, which is more user-friendly and less operator dependent. Specific primers and probes for bacterial species identified from metagenomics were developed and m-ddPCR assays for quantification were designed. M-ddPCR-based results replicated performance from metagenomics with moderately higher accuracy than fecal calprotectin. Importantly, the potential cost of m-ddPCR is substantially lower and the turnaround time is more rapid than that of metagenomic-based tools.

[0845] The metagenomic findings in this example showed enrichments of E. coli and B. fragilis in the gut of patients with CD. Specifically, adherent-invasive E. coli (AIEC) was present in more than half of the CD patients and has been linked to mucosal dysbiosis and functional alteration, associated with disease activity and endoscopic recurrence after surgery, and B. fragilis may induce intestinal inflammation through toxin production. In addition, a novel oral bacterium, Actinomyces sp. oral taxon 181, was discovered, which was significantly enriched in stool samples of patients with CD and UC. It is possible that these resident oral bacteria translocate to the gastrointestinal tract through the bloodstream or the digestive system and colonize and induce inflammation in the gut by activating intestinal immune system. An increased abundance of G. morbillorum was also observed in patients with active UC than those in remission, suggesting its potential role in the inflammatory process.

[0846] The underlying mechanisms were explored to further understand the role of selected bacteria in IBD pathogenesis. The functional microbiome has emerged as a prerequisite for host phenotype and physiology and increasing efforts have been made to link functional traits and mechanisms of organisms to their environments to predict survival and community structure. In this study, functional metabolic perturbations were observed in IBD patients, and the probability of disease obtained from our multi-bacteria biomarker panel could reflect these metabolic dysregulations. Notably, some of the bacterial species drove the functional alterations. For example, reduction in bacterial species with putative anti-inflammatory properties may lead to a reduced capacity for fermentation of dietary fiber and / or production of short-chain fatty acid (SCFA) , whereas deficiency of SCFA is often connected with impaired intestinal mucosal barrier function and induction of intestinal inflammation. The data in this example also showed that several bacteria depleted in IBD were major contributors in amino acid biosynthesis pathway, and the impairment of these pathways may impact intestinal tissue repair and immune regulation in IBD.

[0847] Although dietary data were not specifically collected, the consistent diagnostic performance achieved using data from transethnic public datasets whereby dietary habits vary suggests that the bacterial biomarkers are unlikely to be influenced by diet. The performance of the biomarker panels remained unaffected regardless of different medications that IBD patients were receiving. In this study, the distribution of disease and controls was balanced which may not reflect the true lower prevalence of IBD in the real-life.

[0848] In conclusion, altered gut microbiome signatures and metabolic pathways associated with UC and CD were uncovered in this example. The targeted ddPCR-based quantification of bacterial species consistent with metagenomics data from different populations serves as the foundation for diagnostic assays that are sufficiently robust, sensitive, and cost-effective for clinical application. The identification of reproducible bacterial biomarkers for IBD helps enable the design of non-invasive diagnostic tools for more precise and personalized approaches in IBD detection and management.

[0849] Methods

[0850] Subject recruitment

[0851] [Corrected under Rule 26, 14.03.2025]Metagenomic profiling in two IBD cohorts consisting of patients with CD and UC, and non-IBD controls was performed. The first cohort consisted of 344 patients with UC, 266 patients with CD and 365 age and sex-matched controls from the Prince of Wales Hospital and other hospitals in Hong Kong, China which formed the basis for the biomarker discovery (205 UC, 174 CD, and 118 controls) (Table 3.1) and validation (139 UC and 139 controls; 92 CD and 108 controls) (Table 3.2) . A second group of patients was recruited from St Vincent's Hospital, Melbourne, Australia was used as a validation cohort (98 CD and 81 controls) (Table 3.2) . Patients with CD and UC were diagnosed according to standard criteria of endoscopy, radiology, and histology. Crohn’s Disease Activity Index (CDAI≤150, inactive; >150, active) scores in patients with CD and Mayo scores (≤2, inactive; >2, active) in patients with UC were collected. Individuals with no existing gut disorders such as inflammatory bowel diseases, cancer, advanced adenoma, irritable bowel syndrome, or other GI symptoms were recruited as controls. In addition, five in-house metagenomic cohorts of non-IBD patients with other gastrointestinal disorders, including colorectal adenomas (CA, n=162) , colorectal cancer (CRC, n=160) , irritable bowel syndrome (diarrhea subtype, IBS-D, n=117) , and non-IBD patients with non-gastrointestinal disorders, including obesity (Body mass index>28; n=148) and cardiovascular disease (CVD, n=143) , were also included for validation (Table 3.3) . Subjects with CRC and CA were diagnosed by colonoscopy and confirmed on histology examinations. Subjects with IBS were diagnosed according to the ROME III criteria, and endoscopy and enteroscopy were performed to exclude other GI disorders such as IBD, coeliac disease, parasite infestations, or other organic disorders. Subjects with CVD were recruited from the public as part of a survey of cardiovascular health in the Hong Kong, China general population. Subjects underwent carotid ultrasounds to measure intima-media thickness (IMT) of the common, internal, and external carotid arteries (CCA, ICA and ECA, respectively) and carotid bulbs and subjects that had ≥50%stenosis in a single or multiple vessels were regarded as having the risk of CVD.

[0852] [Corrected under Rule 26, 14.03.2025]Two independent international cohorts from Canada (100 UC; 100 CD; 53 controls) and Taiwan, China (40 UC; 40 CD; 40controls) were recruited from several centers in Canada and from the Taiwan, China University Hospital, respectively. Patients with CD and UC were diagnosed according to standard criteria of endoscopy, radiology, and histology. Crohn’s Disease Activity Index (CDAI≤150, inactive; >150, active) scores or Harvey-Bradshaw Index (HBI≤4, inactive; >4, active) in patients with CD and Mayo scores (≤2, inactive; >2, active) in patients with UC were collected. Individuals with no existing gut disorders such as inflammatory bowel diseases, cancer, advanced adenoma, irritable bowel syndrome, or other GI symptoms were recruited as controls.

[0853] Sample collection

[0854] [Corrected under Rule 26, 14.03.2025]All participants were required to provide at least one spoonful of stool sample using a stool collection tube provided by the investigator in advance. After collection, the stool samples were divided into 2 ml tubes and promptly transferred to a -80℃ ultra-low temperature freezer for storage until further processing. Aliquot tubes will be used for different tests to avoid repeated freezing and thawing. Samples from other centers were also processed using the same procedure and shipped to Hong Kong, China at low temperatures using dry ice ...

Claims

1.A method for determining the risk of, diagnosing, preventing, or treating Inflammatory Bowel Disease (IBD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from IBD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from IBD, optionally treating the individual;wherein the IBD is an active or inactive IBD.2.A method for determining the risk of, diagnosing, preventing, or treating inflammatory bowel disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:(a) determining the relative abundance of a first set of bacterial species, optionally the level of fecal calprotectin, and optionally the relative abundance of a second set of bacterial species in a stool sample from the individual, wherein the first set of bacterial species and the second set of bacterial species each independently comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;(b) comparing the relative abundance of each bacterial species in the first set of bacterial species from the individual with the relative abundance of each bacterial species in the first set of bacterial species in a first reference data set and optionally comparing the level of fecal calprotectin from the individual with the level of fecal calprotectin in the first reference data set using a first machine learning model to generate a first risk score;(c) determining the individual as having increased risk for or is suffering from IBD if the first risk score is higher than a first cutoff value;(d) if the individual is determined to have an increased risk for or is suffering from IBD, optionally comparing the relative abundance of each bacterial species in the second set of bacterial species from the individual with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score;(e) optionally determining the individual as having increased risk for or is suffering from CD if the second risk score is higher than a second cutoff value, or determining the individual as having increased risk for or is suffering from UC if the risk score is lower than or equal to the second cutoff value; and(f) if the individual is determined to have an increased risk for or is suffering from IBD in step (c) , or if steps (d) and (e) are performed and if the individual is determined to have an increased risk for or is suffering from UC or CD, optionally treating the individual.3.The method of claim 2, wherein the first set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.4.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Odoribacter splanchnicus and Gemella morbillorum; Odoribacter splanchnicus and Clostridium leptum; Odoribacter splanchnicus and Blautia hansenii; Odoribacter splanchnicus and Clostridium spiroforme; Odoribacter splanchnicus and Escherichia coli; Odoribacter splanchnicus and Actinomyces sp. oral taxon 181; Gemella morbillorum and Clostridium leptum; Gemella morbillorum and Ruminococcus obeum (Blautia obeum) ; Gemella morbillorum and Blautia hansenii; Gemella morbillorum and Clostridium spiroforme; and Gemella morbillorum and Escherichia coli.5.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises three or more bacterial species, wherein the three or more bacterial species comprises a group of three bacterial species selected from the group consisting of: Odoribacter splanchnicus, Gemella morbillorum and Clostridium leptum; Odoribacter splanchnicus, Gemella morbillorum and Blautia hansenii; Odoribacter splanchnicus, Gemella morbillorum and Clostridium spiroforme; Gemella morbillorum, Clostridium leptum and Ruminococcus obeum (Blautia obeum) ; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Clostridium spiroforme; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Escherichia coli; Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Clostridium spiroforme; and Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Escherichia coli.6.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises two or more bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Ruminococcus obeum (Blautia obeum) , and Blautia hansenii.7.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, and Ruminococcus obeum (Blautia obeum) .8.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Ruminococcus obeum (Blautia obeum) , and Bilophila wadsworthia.9.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Ruminococcus obeum (Blautia obeum) , and Clostridium spiroforme.10.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises at least three bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Ruminococcus obeum (Blautia obeum) , and Clostridium spiroforme.11.The method of claim 2 or claim 3, wherein the first set of bacterial species comprises at least three bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Escherichia coli, and Clostridium spiroforme.12.The method of any one of claims 2-11, wherein the second set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Ruminococcus obeum (Blautia obeum) and Gemmiger formicilis; Ruminococcus obeum (Blautia obeum) and Roseburia intestinalis; Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Ruminococcus obeum (Blautia obeum) and Roseburia inulinivorans; Ruminococcus obeum (Blautia obeum) and Gemella morbillorum; Gemmiger formicilis and Blautia hansenii; Gemmiger formicilis and Roseburia intestinalis; and Gemmiger formicilis and Gemella morbillorum.13.The method of any one of claims 2-11, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, and Bacteroides fragilis.14.The method of any one of claims 2-11, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, and Escherichia coli.15.The method of any one of claims 2-11, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, Escherichia coli, Clostridium leptum, and Odoribacter splanchnicus.16.The method of any one of claims 2-15, wherein the first set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii, and the second set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.17.The method of any one of the preceding claims, wherein the first set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii, and the second set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.18.The method of any one of the preceding claims, whereinthe first reference data set is obtained by determining relative abundances of each bacterial species in the first set of bacterial species and optionally levels of fecal calprotectin in stool samples from a first group of subjects comprising patients diagnosed with IBD and individuals without IBD in a reference cohort; andthe second reference data set is obtained by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects, wherein the second group of subjects comprises patients diagnosed with UC and patients diagnosed with CD.19.The method of claim 18, wherein the subjects in the second group of subjects are selected from the patients diagnosed with IBD in the first group of subjects.20.The method of any one of the preceding claims, wherein the first machine learning model is constructed byi) obtaining the first reference data set by determining relative abundances of each bacterial species in the first set of bacterial species and optionally levels of fecal calprotectin in the stool samples from a first group of subjects comprising patients diagnosed with IBD and individuals without IBD in a reference cohort;ii) dividing the first reference data set into a first training data set and a first test data set;iii) dividing the first training data set into a first training subset and a first validation subset;iv) training a first candidate model with the first training subset and evaluating the performance of the first candidate model using the first validation subset;v) optionally tuning the first candidate model by adjusting hyperparameters in the first candidate model;vi) repeating step (iv) and optionally step (v) to produce a set of first candidate models, and selecting an optimal first candidate model among the set of first candidate models;vii) evaluating the performance of the optimal first candidate model using the first test data set;viii) optionally training a first final model with the first reference data set using a set of hyperparameters corresponding to the optimal first candidate model to produce the first final model; andix) using the optimal first candidate model of step (vii) or if step (viii) is used to train the first final model, using the first final model of step (viii) as the first machine learning model for the comparing step in step (b) to generate the first risk score.21.The method of any one of the preceding claims, wherein the second machine learning model is constructed byi. obtaining the second reference data set by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects comprising patients diagnosed with UC and patients diagnosed with CD;ii. dividing the second reference data set into a second training data set and a second test data set;iii. dividing the second training data set into a second training subset and a second validation subset;iv. training a second candidate model with the second training subset and evaluating the performance of the second candidate model using the second validation subset;v. optionally tuning the second candidate model by adjusting hyperparameters in the second candidate model;vi. repeating step (iv) and optionally step (v) to produce a set of second candidate models, and selecting an optimal second candidate model among the set of second candidate models;vii. evaluating the performance of the optimal second candidate model using the second test data set;viii. optionally training a second final model with the second reference data set using a set of hyperparameters corresponding to the optimal second candidate model to produce the second final model; andix. using the optimal second candidate model of step (vii) or if step (viii) is used to train the second final model, using the second final model of step (viii) as the second machine learning model for the comparing step in step (d) to generate the second risk score.22.The method of any one of the preceding claims, wherein the first machine learning model is a first random forest model, and wherein the first risk score in step (b) is generated by the following steps:(b1) generating an ensemble of decision trees by the first random forest model using the relative abundance of the first set of bacterial species and optionally levels of fecal calprotectin in the first reference data set; and(b2) running the relative abundance of the first set of bacterial species and optionally the level of fecal calprotectin from the individual along the ensemble of decision trees to generate the first risk score.23.The method of any one of the preceding claims, wherein the second machine learning model is a second random forest model, and wherein the second risk score in step (d) is generated by the following steps:(d1) generating an ensemble of decision trees by the second random forest model using the relative abundance of the second set of bacterial species in the second reference data set; and(d2) running the relative abundance of the second set of bacterial species from the individual along the ensemble of decision trees to generate the second risk score.24.The method of claim 20, wherein the optimal first candidate model is the first candidate model with the highest performance among the set of first candidate models.25.The method of claim 21, wherein the optimal second candidate model is the second candidate model with the highest performance among the set of second candidate models.26.The method of any one of claims 20-25, wherein the performance of the first candidate model and / or the second candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the first candidate model and / or the second candidate model.27.The method of any one of the preceding claims, wherein the first cutoff value and / or the second cutoff value are in the range of 0.4 to 0.75.28.The method of any one of the preceding claims, wherein the first cutoff value is determined by the following steps:(c1) calculating a Youden’s index using a first risk score data obtained from the first reference data set; and(c2) determining the first cutoff value based on the Youden’s index.29.The method of any one of the preceding claims, wherein the second cutoff value is determined by the following steps:(e1) calculating a Youden’s index using a second risk score data obtained from the second reference data set; and(e2) determining the second cutoff value based on the Youden’s index.30.The method of any one of the preceding claims, wherein the IBD is an active IBD or an inactive IBD, the UC is an active UC or an inactive UC, and the CD is an active CD or an inactive CD.31.The method of any one of the preceding claims, wherein the method identifies the individual’s risk for IBD over non-IBD diseases, wherein the non-IBD diseases comprise colorectal cancer, colorectal adenomas, irritable bowel syndrome (IBS) , obesity, cardiovascular disease, acid reflux, colon polyp, functional dyspepsia, and gastrointestinal stromal tumor (GIST) .32.The method of any one of the preceding claims, wherein the method identifies an individual as having an increased risk for or suffering from inactive IBD over irritable bowel syndrome (IBS) .33.The method of any one of the preceding claims, wherein the relative abundance of the first set of bacterial species and the relative abundance of the second set of bacterial species are determined by metagenomic sequencing and / or polymerase chain reaction (PCR) .34.The method of claim 33, wherein the PCR is digital droplet PCR (ddPCR) or quantitative PCR (qPCR) performed by using one or more sets of primers and optionally one or more fluorophore-quencher probes, wherein each set of primers is used for amplifying a target nucleic acid sequence belonging to a bacterial species in the first set of bacterial species or the second set of bacterial species to produce an amplification product, and each fluorophore-quencher probe is used for detecting the amplification product of each set of primers.35.The method of claim 34, wherein the target nucleic acid sequence amplified by each set of primers comprises a portion of a nucleotide sequence selected from the group consisting of SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 85, SEQ ID NO: 89, SEQ ID NO: 92, SEQ ID NO: 137, SEQ ID NO: 139, SEQ ID NO: 141, SEQ ID NO: 143, SEQ ID NO: 145, SEQ ID NO: 147, SEQ ID NO: 160, SEQ ID NO: 162, and SEQ ID NO: 164.36.The method of claim 34, wherein the target nucleic acid sequence amplified by each set of primers comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 67, SEQ ID NO: 71, SEQ ID NO: 74, SEQ ID NO: 136, SEQ ID NO: 138, SEQ ID NO: 140, SEQ ID NO: 142, SEQ ID NO: 144, SEQ ID NO: 146, SEQ ID NO: 159, SEQ ID NO: 161, and SEQ ID NO: 163.37.The method of claim 34, wherein the one or more sets of primers comprise a forward primer of SEQ ID NO: 1 and a reverse primer of SEQ ID NO: 2; a forward primer of SEQ ID NO: 3 and a reverse primer of SEQ ID NO: 4; a forward primer of SEQ ID No: 5 and a reverse primer of SEQ ID NO: 6; a forward primer of SEQ ID NO: 9 and a reverse primer of SEQ ID NO: 10; a forward primer of SEQ ID NO: 11 and a reverse primer of SEQ ID NO: 12; a forward primer of SEQ ID NO: 13 and a reverse primer of SEQ ID NO: 14; a forward primer of SEQ ID NO: 19 and a reverse primer of SEQ ID NO: 20; a forward primer of SEQ ID NO: 33 and a reverse primer of SEQ ID NO: 34; a forward primer of SEQ ID NO: 94 and a reverse primer of SEQ ID NO: 95; a forward primer of SEQ ID NO: 96 and a reverse primer of SEQ ID NO: 97; a forward primer of SEQ ID NO: 98 and a reverse primer of SEQ ID NO: 99; a forward primer of SEQ ID NO: 100 and a reverse primer of SEQ ID NO: 101; a forward primer of SEQ ID NO: 102 and a reverse primer of SEQ ID NO: 103; a forward primer of SEQ ID NO: 104 and a reverse primer of SEQ ID NO: 105; a forward primer of SEQ ID NO: 106 and a reverse primer of SEQ ID NO: 107; a forward primer of SEQ ID NO: 118 and a reverse primer of SEQ ID NO: 119; and / or a forward primer of SEQ ID NO: 120 and a reverse primer of SEQ ID NO: 121; and each fluorophore-quencher probe comprises a probe sequence selected from the group consisting of SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 48, SEQ ID NO: 52, SEQ ID NO: 55, SEQ ID NO: 122, SEQ ID NO: 123, SEQ ID NO: 124, SEQ ID NO: 125, SEQ ID NO: 126, SEQ ID NO: 127, SEQ ID NO: 128, SEQ ID NO: 134, and SEQ ID NO: 135.38.The method of any one of the preceding claims, wherein if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, treating the individual with a therapeutic treatment, fecal microbiota transplantation (FMT) , intestinal microbiota transplantation (IMT) , lifestyle and diet modifications, or surgical intervention.39.The method of claim 38, wherein the therapeutic treatment is selected from the group consisting of a 5-Aminosalicylate (5-ASA) , a steroid, a thiopurine, a TNF inhibitor, an anti-integrin, and an anti-interleukin (IL) 12 / 23.40.The method of claim 39, wherein the 5-Aminosalicylate is mesalazine or sulphasalazine; wherein the steroid is budesonide or prednisolone; wherein the thiopurine is azathioprine, mercaptopurine or methotrexate; wherein the TNF inhibitor is infliximab, adalimumab or certolizumab pegol; wherein the anti-integrin is vedolizumab; and wherein the anti-interleukin (IL) 12 / 23 is ustekinumab.41.The method of any one of the preceding claims, wherein if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, treating the individual with an appropriate IBD treatment, UC treatment or CD treatment, respectively.42.The method of claim 41, wherein the UC treatment is selected from the group consisting of 5-Aminosalicylate (5-ASA) , tofacinib, mirikuzumab, and guselkumab.43.The method of claim 41, wherein the CD treatment is natalizumab.44.The method of claim 41, wherein the UC treatment is a composition comprising one or more bacterial species selected from the group consisting of Fusicatenibacter saccharivorans, Clostridium leptum, Gemmiger formicilis, Ruminococcus torques, Odoribacter splanchnicus, Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, and Bilophila wadsworthia and optionally a pharmaceutically acceptable carrier.45.The method of claim 41, wherein the CD treatment is a composition comprising one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans, Dorea formicigenerans, Eubacterium eligens, Roseburia intestinalis, Eubacterium ventriosum, Anaerostipes hadrus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, and Romboutsia ilealis and optionally a pharmaceutically acceptable carrier.46.The method of any one of the preceding claims, wherein the first machine learning model and / or the second machine learning model is a generalized linear model.47.The method of any one of the preceding claims, wherein the individual is an individual without an IBD treatment, an IBD individual, or an individual exposed to an IBD treatment.48.A method for diagnosing, preventing, or treating an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:a) determining if the individual has increased risk for or is suffering from IBD;b) if the individual is determined to have an increased risk for or is suffering from IBD, determining the relative abundance of a second set of bacterial species in a stool sample from the individual, wherein the second set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;c) comparing the relative abundance of each bacterial species in the second set of bacterial species from the individual with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score;d) determining the individual as having increased risk for or is suffering from CD if the second risk score is higher than a second cutoff value, or determining the individual as having increased risk for or is suffering from UC if the risk score is lower than or equal to the second cutoff value; ande) if the individual is determined to have an increased risk for or is suffering from UC or CD, optionally treating the individual.49.The method of claim 48, wherein the determining step in step (a) further comprises the following steps:(a1) determining the relative abundance of a first set of bacterial species and optionally the level of fecal calprotectin in the stool sample from the individual, wherein the first set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;(a2) comparing the relative abundance of each bacterial species in the first set of bacterial species from the individual with the relative abundance of each bacterial species in the first set of bacterial species in a first reference data set and optionally comparing the level of fecal calprotectin from the individual with the level of fecal calprotectin in the first reference data set using a first machine learning model to generate a first risk score;(a3) determining the individual as having increased risk for or is suffering from IBD if the first risk score is higher than a first cutoff value.50.The method of claim 48 or claim 49, wherein the second set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Ruminococcus obeum (Blautia obeum) and Gemmiger formicilis; Ruminococcus obeum (Blautia obeum) and Roseburia intestinalis; Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Ruminococcus obeum (Blautia obeum) and Roseburia inulinivorans; Ruminococcus obeum (Blautia obeum) and Gemella morbillorum; Gemmiger formicilis and Blautia hansenii; Gemmiger formicilis and Roseburia intestinalis; and Gemmiger formicilis and Gemella morbillorum.51.The method of claim 48 or claim 49, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, and Bacteroides fragilis.52.The method of claim 48 or claim 49, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, and Escherichia coli.53.The method of claim 48 or claim 49, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, Escherichia coli, Clostridium leptum, and Odoribacter splanchnicus.54.The method of any one of the preceding claims, wherein the second set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.55.The method of any one of the preceding claims, wherein the second set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.56.The method of any one of the preceding claims, whereinthe second reference data set is obtained by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects, wherein the second group of subjects comprises patients diagnosed with UC and patients diagnosed with CD.57.The method of claim 49, wherein the first reference data set is obtained by determining relative abundances of each bacterial species in the first set of bacterial species and optionally levels of fecal calprotectin in stool samples from a first group of subjects comprising patients diagnosed with IBD and individuals without IBD in a reference cohort; and wherein the subjects in the second group of subjects are selected from the patients diagnosed with IBD in the first group of subjects.58.The method of any one of the preceding claims, wherein the second machine learning model is constructed byi. obtaining the second reference data set by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects comprising patients diagnosed with UC and patients diagnosed with CD;ii. dividing the second reference data set into a second training data set and a second test data set;iii. dividing the second training data set into a second training subset and a second validation subset;iv. training a second candidate model with the second training subset and evaluating the performance of the second candidate model using the second validation subset;v. optionally tuning the second candidate model by adjusting hyperparameters in the second candidate model;vi. repeating step (iv) and optionally step (v) to produce a set of second candidate models, and selecting an optimal second candidate model among the set of second candidate models;vii. evaluating the performance of the optimal second candidate model using the second test data set;viii. optionally training a second final model with the second reference data set using a set of hyperparameters corresponding to the optimal second candidate model to produce the second final model; andix. using the optimal second candidate model of step (vii) or if step (viii) is used to train the second final model, using the second final model of step (viii) as the second machine learning model for the comparing step in step (c) to generate the second risk score.59.The method of any one of the preceding claims, wherein the second machine learning model is a second random forest model, and wherein the second risk score in step (c) is generated by the following steps:(c1) generating an ensemble of decision trees by the second random forest model using the relative abundance of the second set of bacterial species in the second reference data set; and(c2) running the relative abundance of the second set of bacterial species from the individual along the ensemble of decision trees to generate the second risk score.60.The method of claim 58, wherein the optimal second candidate model is the second candidate model with the highest performance among the set of second candidate models.61.The method of any one of claims 58-60, wherein the performance of the second candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the second candidate model.62.The method of any one of the preceding claims, wherein the second cutoff value is in the range of 0.4 to 0.75.63.The method of any one of the preceding claims, wherein the second cutoff value is determined by the following steps:(e1) calculating a Youden’s index using a second risk score data obtained from the second reference data set; and(e2) determining the second cutoff value based on the Youden’s index.64.The method of any one of the preceding claims, wherein the IBD is an active IBD or an inactive IBD, the UC is an active UC or an inactive UC, and the CD is an active CD or an inactive CD.65.The method of any one of the preceding claims, wherein the relative abundance of the first set of bacterial species and / or the second set of bacterial species is determined by metagenomic sequencing and / or polymerase chain reaction (PCR) .66.The method of claim 65, wherein the PCR is digital droplet PCR (ddPCR) or quantitative PCR (qPCR) performed by using one or more sets of primers and optionally one or more fluorophore-quencher probes, wherein each set of primers is used for amplifying a target nucleic acid sequence belonging to a bacterial species in the first set of bacterial species or the second set of bacterial species to produce an amplification product, and each fluorophore-quencher probe is used for detecting the amplification product of each set of primers.67.The method of claim 66, wherein the target nucleic acid sequence amplified by each set of primers comprises a portion of a nucleotide sequence selected from the group consisting of SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 85, SEQ ID NO: 89, SEQ ID NO: 92, SEQ ID NO: 137, SEQ ID NO: 139, SEQ ID NO: 141, SEQ ID NO: 143, SEQ ID NO: 145, SEQ ID NO: 147, SEQ ID NO: 160, SEQ ID NO: 162, and SEQ ID NO: 164.68.The method of claim 66, wherein the target nucleic acid sequence amplified by each set of primers comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 67, SEQ ID NO: 71, SEQ ID NO: 74, SEQ ID NO: 136, SEQ ID NO: 138, SEQ ID NO: 140, SEQ ID NO: 142, SEQ ID NO: 144, SEQ ID NO: 146, SEQ ID NO: 159, SEQ ID NO: 161, and SEQ ID NO: 163.69.The method of claim 66, wherein the one or more sets of primers comprise a forward primer of SEQ ID NO: 1 and a reverse primer of SEQ ID NO: 2; a forward primer of SEQ ID NO: 3 and a reverse primer of SEQ ID NO: 4; a forward primer of SEQ ID No: 5 and a reverse primer of SEQ ID NO: 6; a forward primer of SEQ ID NO: 9 and a reverse primer of SEQ ID NO: 10; a forward primer of SEQ ID NO: 11 and a reverse primer of SEQ ID NO: 12; a forward primer of SEQ ID NO: 13 and a reverse primer of SEQ ID NO: 14; a forward primer of SEQ ID NO: 19 and a reverse primer of SEQ ID NO: 20; a forward primer of SEQ ID NO: 33 and a reverse primer of SEQ ID NO: 34; a forward primer of SEQ ID NO: 94 and a reverse primer of SEQ ID NO: 95; a forward primer of SEQ ID NO: 96 and a reverse primer of SEQ ID NO: 97; a forward primer of SEQ ID NO: 98 and a reverse primer of SEQ ID NO: 99; a forward primer of SEQ ID NO: 100 and a reverse primer of SEQ ID NO: 101; a forward primer of SEQ ID NO: 102 and a reverse primer of SEQ ID NO: 103; a forward primer of SEQ ID NO: 104 and a reverse primer of SEQ ID NO: 105; a forward primer of SEQ ID NO: 106 and a reverse primer of SEQ ID NO: 107; a forward primer of SEQ ID NO: 118 and a reverse primer of SEQ ID NO: 119; and / or a forward primer of SEQ ID NO: 120 and a reverse primer of SEQ ID NO: 121; and each fluorophore-quencher probe comprises a probe sequence selected from the group consisting of SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 48, SEQ ID NO: 52, SEQ ID NO: 55, SEQ ID NO: 122, SEQ ID NO: 123, SEQ ID NO: 124, SEQ ID NO: 125, SEQ ID NO: 126, SEQ ID NO: 127, SEQ ID NO: 128, SEQ ID NO: 134, and SEQ ID NO: 135.70.The method of any one of the preceding claims, wherein if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, treating the individual with a therapeutic treatment, fecal microbiota transplantation (FMT) , intestinal microbiota transplantation (IMT) , lifestyle and diet modifications, or surgical intervention.71.The method of claim 70, wherein the therapeutic treatment is selected from the group consisting of a 5-Aminosalicylate (5-ASA) , a steroid, a thiopurine, a TNF inhibitor, an anti-integrin, and an anti-interleukin (IL) 12 / 23.72.The method of claim 71, wherein the 5-Aminosalicylate is mesalazine or sulphasalazine; wherein the steroid is budesonide or prednisolone; wherein the thiopurine is azathioprine, mercaptopurine or methotrexate; wherein the TNF inhibitor is infliximab, adalimumab or certolizumab pegol; wherein the anti-integrin is vedolizumab; and wherein the anti-interleukin (IL) 12 / 23 is ustekinumab.73.The method of any one of the preceding claims, wherein if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, treating the individual with an appropriate IBD treatment, UC treatment or CD treatment, respectively.74.The method of claim 73, wherein the UC treatment is selected from the group consisting of 5-Aminosalicylate (5-ASA) , tofacinib, mirikuzumab, and guselkumab.75.The method of claim 73, wherein the CD treatment is natalizumab.76.The method of claim 73, wherein the UC treatment is a composition comprising one or more bacterial species selected from the group consisting of Fusicatenibacter saccharivorans, Clostridium leptum, Gemmiger formicilis, Ruminococcus torques, Odoribacter splanchnicus, Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, and Bilophila wadsworthia and optionally a pharmaceutically acceptable carrier.77.The method of claim 73, wherein the CD treatment is a composition comprising one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, Roseburia inulinivorans, Dorea formicigenerans, Eubacterium eligens, Roseburia intestinalis, Eubacterium ventriosum, Anaerostipes hadrus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, and Romboutsia ilealis and optionally a pharmaceutically acceptable carrier.78.The method of any one of the preceding claims, wherein the first machine learning model and / or the second machine learning model is a generalized linear model.79.A method for identifying a stool sample with an altered Inflammatory Bowel Disease’s (IBD) microbiome and / or an altered IBD subtype microbiome, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:a) determining the relative abundance of a first set of bacterial species, optionally the level of fecal calprotectin, and optionally the relative abundance of a second set of bacterial species in the stool sample, wherein the first set of bacterial species and the second set of bacterial species each independently comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the first set of bacterial species from the stool sample with the relative abundance of each bacterial species in the first set of bacterial species in a first reference data set and optionally comparing the level of fecal calprotectin from the stool sample with the level of fecal calprotectin in a the first reference data set using a first machine learning model to generate a risk score; andc) determining the stool sample as having an altered IBD microbiome if the first risk score is higher than a first cutoff value;d) if the stool sample is determined to have the altered IBD microbiome, optionally comparing the relative abundance of each bacterial species in the second set of bacterial species from the stool sample with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score; ande) optionally determining the stool sample as having an altered CD microbiome if the second risk score is higher than a second cutoff value, or determining the stool sample as having an altered UC microbiome if the risk score is lower than or equal to the second cutoff value.80.The method of claim 79, wherein the first set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.81.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Odoribacter splanchnicus and Gemella morbillorum; Odoribacter splanchnicus and Clostridium leptum; Odoribacter splanchnicus and Blautia hansenii; Odoribacter splanchnicus and Clostridium spiroforme; Odoribacter splanchnicus and Escherichia coli; Odoribacter splanchnicus and Actinomyces sp. oral taxon 181; Gemella morbillorum and Clostridium leptum; Gemella morbillorum and Ruminococcus obeum (Blautia obeum) ; Gemella morbillorum and Blautia hansenii; Gemella morbillorum and Clostridium spiroforme; and Gemella morbillorum and Escherichia coli.82.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises three or more bacterial species, wherein the three or more bacterial species comprises a group of three bacterial species selected from the group consisting of: Odoribacter splanchnicus, Gemella morbillorum and Clostridium leptum; Odoribacter splanchnicus, Gemella morbillorum and Blautia hansenii; Odoribacter splanchnicus, Gemella morbillorum and Clostridium spiroforme; Gemella morbillorum, Clostridium leptum and Ruminococcus obeum (Blautia obeum) ; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Clostridium spiroforme; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Escherichia coli; Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Clostridium spiroforme; and Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Escherichia coli.83.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises two or more bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Ruminococcus obeum (Blautia obeum) , and Blautia hansenii.84.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, and Ruminococcus obeum (Blautia obeum) .85.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Ruminococcus obeum (Blautia obeum) , and Bilophila wadsworthia.86.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Ruminococcus obeum (Blautia obeum) , and Clostridium spiroforme.87.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises at least three bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Ruminococcus obeum (Blautia obeum) , and Clostridium spiroforme.88.The method of claim 79 or claim 80, wherein the first set of bacterial species comprises at least three bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Escherichia coli, and Clostridium spiroforme.89.The method of any one of claims 79-88, wherein the second set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Ruminococcus obeum (Blautia obeum) and Gemmiger formicilis; Ruminococcus obeum (Blautia obeum) and Roseburia intestinalis; Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Ruminococcus obeum (Blautia obeum) and Roseburia inulinivorans; Ruminococcus obeum (Blautia obeum) and Gemella morbillorum; Gemmiger formicilis and Blautia hansenii; Gemmiger formicilis and Roseburia intestinalis; and Gemmiger formicilis and Gemella morbillorum.90.The method of any one of claims 79-88, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, and Bacteroides fragilis.91.The method of any one of claims 79-88, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, and Escherichia coli.92.The method of any one of claims 79-88, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, Escherichia coli, Clostridium leptum, and Odoribacter splanchnicus.93.The method of any one of the preceding claims, wherein the first set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii, and the second set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.94.The method of any one of the preceding claims, wherein the first set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii, and the second set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.95.The method of any one of the preceding claims, wherein2. the first reference data set is obtained by determining relative abundances of each bacterial species in the first set of bacterial species and optionally levels of fecal calprotectin in stool samples from a first group of subjects comprising patients diagnosed with IBD and individuals without IBD in a reference cohort; and3. the second reference data set is obtained by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects, wherein the second group of subjects comprises patients diagnosed with UC and patients diagnosed with CD.96.The method of claim 95, wherein the subjects in the second group of subjects are selected from the patients diagnosed with IBD in the first group of subjects.97.The method of any one of the preceding claims, wherein the first machine learning model is constructed byi) obtaining the first reference data set by determining relative abundances of each bacterial species in the first set of bacterial species and optionally levels of fecal calprotectin in the stool samples from a first group of subjects comprising patients diagnosed with IBD and individuals without IBD in a reference cohort;ii) dividing the first reference data set into a first training data set and a first test data set;iii) dividing the first training data set into a first training subset and a first validation subset;iv) training a first candidate model with the first training subset and evaluating the performance of the first candidate model using the first validation subset;v) optionally tuning the first candidate model by adjusting hyperparameters in the first candidate model;vi) repeating step (iv) and optionally step (v) to produce a set of first candidate models, and selecting an optimal first candidate model among the set of first candidate models;vii) evaluating the performance of the optimal first candidate model using the first test data set;viii) optionally training a first final model with the first reference data set using a set of hyperparameters corresponding to the optimal first candidate model to produce the first final model; andix) using the optimal first candidate model of step (vii) or if step (viii) is used to train the first final model, using the first final model of step (viii) as the first machine learning model for the comparing step in step (b) to generate the first risk score.98.The method of any one of the preceding claims, wherein the second machine learning model is constructed byi. obtaining the second reference data set by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects comprising patients diagnosed with UC and patients diagnosed with CD;ii. dividing the second reference data set into a second training data set and a second test data set;iii. dividing the second training data set into a second training subset and a second validation subset;iv. training a second candidate model with the second training subset and evaluating the performance of the second candidate model using the second validation subset;v. optionally tuning the second candidate model by adjusting hyperparameters in the second candidate model;vi. repeating step (iv) and optionally step (v) to produce a set of second candidate models, and selecting an optimal second candidate model among the set of second candidate models;vii. evaluating the performance of the optimal second candidate model using the second test data set;viii. optionally training a second final model with the second reference data set using a set of hyperparameters corresponding to the optimal second candidate model to produce the second final model; andix. using the optimal second candidate model of step (vii) or if step (viii) is used to train the second final model, using the second final model of step (viii) as the second machine learning model for the comparing step in step (d) to generate the second risk score.99.The method of any one of the preceding claims, wherein the first machine learning model is a first random forest model, and wherein the first risk score in step (b) is generated by the following steps:(b1) generating an ensemble of decision trees by the first random forest model using the relative abundance of the first set of bacterial species and optionally levels of fecal calprotectin in the first reference data set; and(b2) running the relative abundance of the first set of bacterial species and optionally the level of fecal calprotectin from the stool sample along the ensemble of decision trees to generate the first risk score.100.The method of any one of the preceding claims, wherein the second machine learning model is a second random forest model, and wherein the second risk score in step (d) is generated by the following steps:(d1) generating an ensemble of decision trees by the second random forest model using the relative abundance of the second set of bacterial species in the second reference data set; and(d2) running the relative abundance of the second set of bacterial species from the stool sample along the ensemble of decision trees to generate the second risk score.101.The method of claim 97, wherein the optimal first candidate model is the first candidate model with the highest performance among the set of first candidate models.102.The method of claim 98, wherein the optimal second candidate model is the second candidate model with the highest performance among the set of second candidate models.103.The method of any one of the preceding claims, wherein the performance of the first candidate model and / or the second candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the first candidate model and / or the second candidate model.104.The method of any one of the preceding claims, wherein the first cutoff value and / or the second cutoff value are in the range of 0.4 to 0.75.105.The method of any one of the preceding claims, wherein the first cutoff value is determined by the following steps:(c1) calculating a Youden’s index using a first risk score data obtained from the first reference data set; and(c2) determining the first cutoff value based on the Youden’s index.106.The method of any one of the preceding claims, wherein the second cutoff value is determined by the following steps:(e1) calculating a Youden’s index using a second risk score data obtained from the second reference data set; and(e2) determining the second cutoff value based on the Youden’s index.107.The method of any one of the preceding claims, wherein the IBD is an active IBD or an inactive IBD, the UC is an active UC or an inactive UC, and the CD is an active CD or an inactive CD.108.The method of any one of the preceding claims, wherein the relative abundance of the first set of bacterial species and the relative abundance of the second set of bacterial species are determined by metagenomic sequencing and / or PCR.109.The method of claim 108, wherein the PCR is digital droplet PCR (ddPCR) or quantitative PCR (qPCR) performed by using one or more sets of primers and optionally one or more fluorophore-quencher probes, wherein each set of primers is used for amplifying a target nucleic acid sequence belonging to a bacterial species in the first set of bacterial species or the second set of bacterial species to produce an amplification product, and each fluorophore-quencher probe is used for detecting the amplification product of each set of primers.110.The method of claim 109, wherein the target nucleic acid sequence amplified by each set of primers comprises a portion of a nucleotide sequence selected from the group consisting of SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 85, SEQ ID NO: 89, SEQ ID NO: 92, SEQ ID NO: 137, SEQ ID NO: 139, SEQ ID NO: 141, SEQ ID NO: 143, SEQ ID NO: 145, SEQ ID NO: 147, SEQ ID NO: 160, SEQ ID NO: 162, and SEQ ID NO: 164.111.The method of claim 109, wherein the target nucleic acid sequence amplified by each set of primers comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 67, SEQ ID NO: 71, SEQ ID NO: 74, SEQ ID NO: 136, SEQ ID NO: 138, SEQ ID NO: 140, SEQ ID NO: 142, SEQ ID NO: 144, SEQ ID NO: 146, SEQ ID NO: 159, SEQ ID NO: 161, and SEQ ID NO: 163.112.The method of claim 109, wherein the one or more sets of primers comprise a forward primer of SEQ ID NO: 1 and a reverse primer of SEQ ID NO: 2; a forward primer of SEQ ID NO: 3 and a reverse primer of SEQ ID NO: 4; a forward primer of SEQ ID No: 5 and a reverse primer of SEQ ID NO: 6; a forward primer of SEQ ID NO: 9 and a reverse primer of SEQ ID NO: 10; a forward primer of SEQ ID NO: 11 and a reverse primer of SEQ ID NO: 12; a forward primer of SEQ ID NO: 13 and a reverse primer of SEQ ID NO: 14; a forward primer of SEQ ID NO: 19 and a reverse primer of SEQ ID NO: 20; a forward primer of SEQ ID NO: 33 and a reverse primer of SEQ ID NO: 34; a forward primer of SEQ ID NO: 94 and a reverse primer of SEQ ID NO: 95; a forward primer of SEQ ID NO: 96 and a reverse primer of SEQ ID NO: 97; a forward primer of SEQ ID NO: 98 and a reverse primer of SEQ ID NO: 99; a forward primer of SEQ ID NO: 100 and a reverse primer of SEQ ID NO: 101; a forward primer of SEQ ID NO: 102 and a reverse primer of SEQ ID NO: 103; a forward primer of SEQ ID NO: 104 and a reverse primer of SEQ ID NO: 105; a forward primer of SEQ ID NO: 106 and a reverse primer of SEQ ID NO: 107; a forward primer of SEQ ID NO: 118 and a reverse primer of SEQ ID NO: 119; and / or a forward primer of SEQ ID NO: 120 and a reverse primer of SEQ ID NO: 121; and each fluorophore-quencher probe comprises a probe sequence selected from the group consisting of SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 48, SEQ ID NO: 52, SEQ ID NO: 55, SEQ ID NO: 122, SEQ ID NO: 123, SEQ ID NO: 124, SEQ ID NO: 125, SEQ ID NO: 126, SEQ ID NO: 127, SEQ ID NO: 128, SEQ ID NO: 134, and SEQ ID NO: 135.113.A method for identifying a stool sample with an altered IBD subtype microbiome, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said method comprising the steps of:a) determining if the stool sample has an altered IBD microbiome;b) if the stool sample is determined to have an altered IBD microbiome, determining the relative abundance of a second set of bacterial species in a stool sample from the individual, wherein the second set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;c) comparing the relative abundance of each bacterial species in the second set of bacterial species from the stool sample with the relative abundance of each bacterial species in the second set of bacterial species in a second reference data set using a second machine learning model to generate a second risk score; andd) determining the stool sample as having an altered CD microbiome if the second risk score is higher than a second cutoff value, or determining the stool sample as having an altered UC microbiome if the risk score is lower than or equal to the second cutoff value.114.The method of claim 113, wherein the determining step in step (a) further comprises the following steps:(a1) determining the relative abundance of a first set of bacterial species and optionally the level of fecal calprotectin in the stool sample, wherein the first set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;(a2) comparing the relative abundance of each bacterial species in the first set of bacterial species from the stool sample with the relative abundance of each bacterial species in the first set of bacterial species in a first reference data set and optionally comparing the level of fecal calprotectin from the stool sample with the level of fecal calprotectin in the first reference data set using a first machine learning model to generate a first risk score;(a3) determining the stool sample as having an altered IBD microbiome if the first risk score is higher than a first cutoff value.115.The method of claim 113 or claim 114, wherein the second set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Ruminococcus obeum (Blautia obeum) and Gemmiger formicilis; Ruminococcus obeum (Blautia obeum) and Roseburia intestinalis; Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Ruminococcus obeum (Blautia obeum) and Roseburia inulinivorans; Ruminococcus obeum (Blautia obeum) and Gemella morbillorum; Gemmiger formicilis and Blautia hansenii; Gemmiger formicilis and Roseburia intestinalis; and Gemmiger formicilis and Gemella morbillorum.116.The method of claim 113 or claim 114, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, and Bacteroides fragilis.117.The method of claim 113 or claim 114, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, and Escherichia coli.118.The method of claim 113 or claim 114, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, Escherichia coli, Clostridium leptum, and Odoribacter splanchnicus.119.The method of any one of the preceding claims, wherein the second set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.120.The method of any one of the preceding claims, wherein the second set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.121.The method of any one of the preceding claims, whereinthe second reference data set is obtained by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects, wherein the second group of subjects comprises patients diagnosed with UC and patients diagnosed with CD.122.The method of claim 114, wherein the first reference data set is obtained by determining relative abundances of each bacterial species in the first set of bacterial species and optionally levels of fecal calprotectin in stool samples from a first group of subjects comprising patients diagnosed with IBD and individuals without IBD in a reference cohort; and wherein the subjects in the second group of subjects are selected from the patients diagnosed with IBD in the first group of subjects.123.The method of any one of the preceding claims, wherein the second machine learning model is constructed byi. obtaining the second reference data set by determining relative abundances of each bacterial species in the second set of bacterial species in the stool samples from a second group of subjects comprising patients diagnosed with UC and patients diagnosed with CD;ii. dividing the second reference data set into a second training data set and a second test data set;iii. dividing the second training data set into a second training subset and a second validation subset;iv. training a second candidate model with the second training subset and evaluating the performance of the second candidate model using the second validation subset;v. optionally tuning the second candidate model by adjusting hyperparameters in the second candidate model;vi. repeating step (iv) and optionally step (v) to produce a set of second candidate models, and selecting an optimal second candidate model among the set of second candidate models;vii. evaluating the performance of the optimal second candidate model using the second test data set;viii. optionally training a second final model with the second reference data set using a set of hyperparameters corresponding to the optimal second candidate model to produce the second final model; andix. using the optimal second candidate model of step (vii) or if step (viii) is used to train the second final model, using the second final model of step (viii) as the second machine learning model for the comparing step in step (d) to generate the second risk score.124.The method of any one of the preceding claims, wherein the second machine learning model is a second random forest model, and wherein the second risk score in step (c) is generated by the following steps:(c1) generating an ensemble of decision trees by the second random forest model using the relative abundance of the second set of bacterial species in the second reference data set; and(c2) running the relative abundance of the second set of bacterial species from the individual along the ensemble of decision trees to generate the second risk score.125.The method of claim 123, wherein the optimal second candidate model is the second candidate model with the highest performance among the set of second candidate models.126.The method of any one of claims 123-125, wherein the performance of the second candidate model is evaluated by Area Under the Curve (AUC) in Receiving Operating Characteristic (ROC) curve of the second candidate model.127.The method of any one of the preceding claims, wherein the second cutoff value is in the range of 0.4 to 0.75.128.The method of any one of the preceding claims, wherein the second cutoff value is determined by the following steps:(e1) calculating a Youden’s index using a second risk score data obtained from the second reference data set; and(e2) determining the second cutoff value based on the Youden’s index.129.The method of any one of the preceding claims, wherein the IBD is an active IBD or an inactive IBD, the UC is an active UC or an inactive UC, and the CD is an active CD or an inactive CD.130.The method of any one of the preceding claims, wherein the relative abundance of the first set of bacterial species and / or the second set of bacterial species is determined by metagenomic sequencing and / or PCR.131.The method of claim 130, wherein the PCR is digital droplet PCR (ddPCR) or quantitative PCR (qPCR) performed by using one or more sets of primers and optionally one or more fluorophore-quencher probes, wherein each set of primers is used for amplifying a target nucleic acid sequence belonging to a bacterial species in the first set of bacterial species or the second set of bacterial species to produce an amplification product, and each fluorophore-quencher probe is used for detecting the amplification product of each set of primers.132.The method of claim 131, wherein the target nucleic acid sequence amplified by each set of primers comprises a portion of a nucleotide sequence selected from the group consisting of SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 85, SEQ ID NO: 89, SEQ ID NO: 92, SEQ ID NO: 137, SEQ ID NO: 139, SEQ ID NO: 141, SEQ ID NO: 143, SEQ ID NO: 145, SEQ ID NO: 147, SEQ ID NO: 160, SEQ ID NO: 162, and SEQ ID NO: 164.133.The method of claim 131, wherein the target nucleic acid sequence amplified by each set of primers comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 67, SEQ ID NO: 71, SEQ ID NO: 74, SEQ ID NO: 136, SEQ ID NO: 138, SEQ ID NO: 140, SEQ ID NO: 142, SEQ ID NO: 144, SEQ ID NO: 146, SEQ ID NO: 159, SEQ ID NO: 161, and SEQ ID NO: 163.134.The method of claim 131, wherein the one or more sets of primers comprise a forward primer of SEQ ID NO: 1 and a reverse primer of SEQ ID NO: 2; a forward primer of SEQ ID NO: 3 and a reverse primer of SEQ ID NO: 4; a forward primer of SEQ ID No: 5 and a reverse primer of SEQ ID NO: 6; a forward primer of SEQ ID NO: 9 and a reverse primer of SEQ ID NO: 10; a forward primer of SEQ ID NO: 11 and a reverse primer of SEQ ID NO: 12; a forward primer of SEQ ID NO: 13 and a reverse primer of SEQ ID NO: 14; a forward primer of SEQ ID NO: 19 and a reverse primer of SEQ ID NO: 20; a forward primer of SEQ ID NO: 33 and a reverse primer of SEQ ID NO: 34; a forward primer of SEQ ID NO: 94 and a reverse primer of SEQ ID NO: 95; a forward primer of SEQ ID NO: 96 and a reverse primer of SEQ ID NO: 97; a forward primer of SEQ ID NO: 98 and a reverse primer of SEQ ID NO: 99; a forward primer of SEQ ID NO: 100 and a reverse primer of SEQ ID NO: 101; a forward primer of SEQ ID NO: 102 and a reverse primer of SEQ ID NO: 103; a forward primer of SEQ ID NO: 104 and a reverse primer of SEQ ID NO: 105; a forward primer of SEQ ID NO: 106 and a reverse primer of SEQ ID NO: 107; a forward primer of SEQ ID NO: 118 and a reverse primer of SEQ ID NO: 119; and / or a forward primer of SEQ ID NO: 120 and a reverse primer of SEQ ID NO: 121; and each fluorophore-quencher probe comprises a probe sequence selected from the group consisting of SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 48, SEQ ID NO: 52, SEQ ID NO: 55, SEQ ID NO: 122, SEQ ID NO: 123, SEQ ID NO: 124, SEQ ID NO: 125, SEQ ID NO: 126, SEQ ID NO: 127, SEQ ID NO: 128, SEQ ID NO: 134, and SEQ ID NO: 135.135.A kit for determining the risk of or diagnosing Inflammatory Bowel Disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , said kit comprising a reagent for detecting a first set of bacterial species and optionally a reagent for detecting a second set of bacterial species, wherein the first set of bacterial species and the second set of bacterial species each independent comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.136.The kit of claim 135, wherein the first set of bacterial species comprises one or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.137.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Odoribacter splanchnicus and Gemella morbillorum; Odoribacter splanchnicus and Clostridium leptum; Odoribacter splanchnicus and Blautia hansenii; Odoribacter splanchnicus and Clostridium spiroforme; Odoribacter splanchnicus and Escherichia coli; Odoribacter splanchnicus and Actinomyces sp. oral taxon 181; Gemella morbillorum and Clostridium leptum; Gemella morbillorum and Ruminococcus obeum (Blautia obeum) ; Gemella morbillorum and Blautia hansenii; Gemella morbillorum and Clostridium spiroforme; and Gemella morbillorum and Escherichia coli.138.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises three or more bacterial species, wherein the three or more bacterial species comprises a group of three bacterial species selected from the group consisting of: Odoribacter splanchnicus, Gemella morbillorum and Clostridium leptum; Odoribacter splanchnicus, Gemella morbillorum and Blautia hansenii; Odoribacter splanchnicus, Gemella morbillorum and Clostridium spiroforme; Gemella morbillorum, Clostridium leptum and Ruminococcus obeum (Blautia obeum) ; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Clostridium spiroforme; Gemella morbillorum, Ruminococcus obeum (Blautia obeum) and Escherichia coli; Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Clostridium spiroforme; and Clostridium leptum, Ruminococcus obeum (Blautia obeum) and Escherichia coli.139.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises two or more bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Ruminococcus obeum (Blautia obeum) , and Blautia hansenii.140.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, and Ruminococcus obeum (Blautia obeum) .141.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Ruminococcus obeum (Blautia obeum) , and Bilophila wadsworthia.142.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises at least one bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Ruminococcus obeum (Blautia obeum) , and Clostridium spiroforme.143.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises at least three bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Ruminococcus obeum (Blautia obeum) , and Clostridium spiroforme.144.The kit of claim 135 or claim 136, wherein the first set of bacterial species comprises at least three bacterial species selected from the group consisting of Odoribacter splanchnicus, Gemella morbillorum, Clostridium leptum, Blautia hansenii, Escherichia coli, and Clostridium spiroforme.145.The kit of any one of claims 135-144, wherein the second set of bacterial species comprises two or more bacterial species, wherein the two or more bacterial species comprises a group of two bacterial species selected from the group consisting of: Ruminococcus obeum (Blautia obeum) and Gemmiger formicilis; Ruminococcus obeum (Blautia obeum) and Roseburia intestinalis; Ruminococcus obeum (Blautia obeum) and Blautia hansenii; Ruminococcus obeum (Blautia obeum) and Roseburia inulinivorans; Ruminococcus obeum (Blautia obeum) and Gemella morbillorum; Gemmiger formicilis and Blautia hansenii; Gemmiger formicilis and Roseburia intestinalis; and Gemmiger formicilis and Gemella morbillorum.146.The kit of any one of claims 135-144, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, and Bacteroides fragilis.147.The kit of any one of claims 135-144, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, and Escherichia coli.148.The kit of any one of claims 135-144, wherein the second set of bacterial species comprises at least one bacterial species selected from the group consisting of Ruminococcus obeum (Blautia obeum) , Gemmiger formicilis, Roseburia intestinalis, Blautia hansenii, Gemella morbillorum, Lawsonibacter asaccharolyticus, Bacteroides fragilis, Escherichia coli, Clostridium leptum, and Odoribacter splanchnicus.149.The kit of any one of claims 135-144, wherein the first set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii, and the second set of bacterial species comprises four or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.150.The kit of any one of the preceding claims, wherein the first set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii, and the second set of bacterial species consists essentially of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.151.The kit of any one of the preceding claims, wherein the reagent comprises one or more sets of primers and optionally one or more fluorophore-quencher probes, wherein each set of primers is used for amplifying a target nucleic acid sequence belonging to a bacterial species in the first set of bacterial species or the second set of bacterial species via PCR to produce an amplification product, and each fluorophore-quencher probe is used for detecting the amplification product of each set of primers.152.The kit of claim 151, wherein the target nucleic acid sequence amplified by each set of primers comprises a portion of a nucleotide sequence selected from the group consisting of SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 85, SEQ ID NO: 89, SEQ ID NO: 92, SEQ ID NO: 137, SEQ ID NO: 139, SEQ ID NO: 141, SEQ ID NO: 143, SEQ ID NO: 145, SEQ ID NO: 147, SEQ ID NO: 160, SEQ ID NO: 162, and SEQ ID NO: 164.153.The kit of claim 151, wherein the target nucleic acid sequence amplified by each set of primers comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 63, SEQ ID NO: 64, SEQ ID NO: 67, SEQ ID NO: 71, SEQ ID NO: 74, SEQ ID NO: 136, SEQ ID NO: 138, SEQ ID NO: 140, SEQ ID NO: 142, SEQ ID NO: 144, SEQ ID NO: 146, SEQ ID NO: 159, SEQ ID NO: 161, and SEQ ID NO: 163.154.The kit of claim 151, wherein the one or more sets of primers comprise a forward primer of SEQ ID NO: 1 and a reverse primer of SEQ ID NO: 2; a forward primer of SEQ ID NO: 3 and a reverse primer of SEQ ID NO: 4; a forward primer of SEQ ID No: 5 and a reverse primer of SEQ ID NO: 6; a forward primer of SEQ ID NO: 9 and a reverse primer of SEQ ID NO: 10; a forward primer of SEQ ID NO: 11 and a reverse primer of SEQ ID NO: 12; a forward primer of SEQ ID NO: 13 and a reverse primer of SEQ ID NO: 14; a forward primer of SEQ ID NO: 19 and a reverse primer of SEQ ID NO: 20; a forward primer of SEQ ID NO: 33 and a reverse primer of SEQ ID NO: 34; a forward primer of SEQ ID NO: 94 and a reverse primer of SEQ ID NO: 95; a forward primer of SEQ ID NO: 96 and a reverse primer of SEQ ID NO: 97; a forward primer of SEQ ID NO: 98 and a reverse primer of SEQ ID NO: 99; a forward primer of SEQ ID NO: 100 and a reverse primer of SEQ ID NO: 101; a forward primer of SEQ ID NO: 102 and a reverse primer of SEQ ID NO: 103; a forward primer of SEQ ID NO: 104 and a reverse primer of SEQ ID NO: 105; a forward primer of SEQ ID NO: 106 and a reverse primer of SEQ ID NO: 107; a forward primer of SEQ ID NO: 118 and a reverse primer of SEQ ID NO: 119; and / or a forward primer of SEQ ID NO: 120 and a reverse primer of SEQ ID NO: 121; and each fluorophore-quencher probe comprises a probe sequence selected from the group consisting of SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 48, SEQ ID NO: 52, SEQ ID NO: 55, SEQ ID NO: 122, SEQ ID NO: 123, SEQ ID NO: 124, SEQ ID NO: 125, SEQ ID NO: 126, SEQ ID NO: 127, SEQ ID NO: 128, SEQ ID NO: 134, and SEQ ID NO: 135.155.A computer program product for determining the risk of or diagnosing Inflammatory Bowel Disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of claims 2-47.156.A computer program product for determining the risk of, diagnosing, or suggesting treatment for Inflammatory Bowel Disease (IBD) and / or an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of claims 2-47, and optionally further configured to enable the execution of the following step:if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, suggesting an appropriate IBD treatment, UC treatment or CD treatment, respectively to the individual.157.A computer program product for determining the risk of or diagnosing an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of claims 48-78.158.A computer program product for determining the risk of, diagnosing, or suggesting treatment for an IBD subtype in an individual, wherein the IBD subtype is Ulcerative Colitis (UC) or Crohn’s disease (CD) , wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of the method according to any one of claims 48-78, and optionally further configured to enable the execution of the following step:if the individual is determined to have an increased risk for or is suffering from IBD, UC or CD, suggesting an appropriate IBD treatment, UC treatment or CD treatment, respectively to the individual.159.The method of any one of claims 79-112, wherein the first machine learning model and / or the second machine learning model is a generalized linear model.160.The method of any one of claims 113-134, wherein the first machine learning model and / or the second machine learning model is a generalized linear model.161.A method for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from CD, optionally treating the individual.162.A method for identifying a stool sample with an altered Crohn’s Disease microbiome comprising the steps of:a) determining the relative abundance of a set of bacterial species in the stool sample, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the stool sample as having an altered CD microbiome if the risk score is higher than a cutoff value.163.A method of increasing amino acid biosynthesis in a cell, wherein the amino acid is L-arginine, L-ornithine, or L-valine, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.164.A method of increasing carbohydrate degradation in a cell, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.165.A method of increasing cofactor, carrier, and vitamin biosynthesis in a cell, wherein the vitamin is thiamine phosphate, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, and Roseburia intestinalis to the cell.166.A method of increasing amino acid biosynthesis in a cell, wherein the amino acid is L-tryptophan, comprising providing an effective amount of Ruminococcus obeum (B. obeum) , Roseburia inulinivorans, Dorea formicigenerans, Eubacterium sp. CAG: 274, or Roseburia intestinalis to the cell.167.A composition comprising Lawsonibacter asaccharolyticus and Eubacterium sp. CAG: 274.168.A composition comprising Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, and Roseburia inulinivorans.169.A kit for determining the risk of or diagnosing Crohn’s Disease (CD) in an individual, comprising a reagent for detecting a set of bacterial species, wherein the set of bacterial species is selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum, Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3.170.A computer program product for determining the risk of or diagnosing Crohn’s Disease (CD) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value.171.A method for determining the risk of, diagnosing, preventing, or treating Crohn’s Disease (CD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Gemmiger formicilis, Eubacterium hallii, Blautia obeum, Roseburia inulinivorans, Fusicatenibacter saccharivorans, Eubacterium rectale, Roseburia faecis, Asaccharobacter celatus, Collinsella aerofaciens, Faecalibacterium prausnitzii, Anaerostipes hadrus, Lachnospira pectinoschiza, Ruminococcus torques, Clostridium leptum, Parabacteroides merdae, Ruminococcus bromii, Roseburia intestinalis, Adlercreutzia equolifaciens, Alistipes putredinis, Eubacterium sp. CAG: 38, Roseburia hominis, Agathobaculum butyriciproducens, Dorea formicigenerans, Bacteroides stercoris, Odoribacter splanchnicus, Oscillibacter sp. 57_20, Coprococcus catus, Bacteroides vulgatus, Dorea longicatena, Alistipes shahii, Lawsonibacter asaccharolyticus, Firmicutes bacterium CAG: 83, Oscillibacter sp. CAG: 241, Eubacterium ramulus, Alistipes finegoldii, Bacteroides caccae, Akkermansia muciniphila, Coprococcus comes, Butyricimonas virosa, Alistipes indistinctus, Bacteroides uniformis, Bacteroides coprocola, Eubacterium sp. CAG: 274, Desulfovibrio piger, Bacteroides massiliensis, Clostridium sp. CAG: 299, Eubacterium eligens, Romboutsia ilealis, Eubacterium ventriosum, Anaerostipes hadrus, Lactobacillus rogosae, Ruminococcus obeum, Gemella haemolysans, Streptococcus mitis, Actinomyces odontolyticus, Actinomyces sp. S6-Spd3, Actinomyces graevenitzii, Lactobacillus mucosae, Streptococcus parasanguinis, Rothia mucilaginosa, Actinomyces sp. oral taxon 181, Veillonella parvula, Escherichia coli, Bacteroides fragilis, and Eubacterium sulci;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from CD, optionally treating the individual.172.A method for monitoring disease activity of Crohn’s Disease (CD) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Ruminococcus obeum (B. obeum) , Lawsonibacter asaccharolyticus, and Escherichia coli;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased disease activity if the risk score is higher than a cutoff value.173.A method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from UC, optionally treating the individual.174.A method for identifying a stool sample with an altered ulcerative colitis microbiome comprising the steps of:a) determining the relative abundance of a set of bacterial species in the stool sample, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the stool sample as having an altered UC microbiome if the risk score is higher than a cutoff value.175.A method of increasing nucleoside and nucleotide biosynthesis in a cell, comprising providing an effective amount of one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, and Ruminococcus torques to the cell.176.A composition comprising Fusicatenibacter saccharivorans, Clostridium leptum, and Gemmiger formicilis.177.A kit for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, comprising a reagent for detecting a set of bacterial species, wherein the set of bacterial species is selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.178.A computer program product for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value.179.A method for determining the risk of, diagnosing, preventing, or treating ulcerative colitis (UC) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Collinsella aerofaciens, Clostridium leptum, Ruminococcus torques, Asaccharobacter celatus, Gemmiger formicilis, Fusicatenibacter saccharivorans, Alistipes putredinis, Dorea longicatena, Ruminococcus bromii, Odoribacter splanchnicus, Lachnospira pectinoschiza, Coprococcus comes, Adlercreutzia equolifaciens, Roseburia inulinivorans, Blautia obeum, Eubacterium rectale, Dorea formicigenerans, Alistipes shahii, Bacteroides stercoris, Parabacteroides merdae, Oscillibacter sp. CAG: 241, Akkermansia muciniphila, Alistipes finegoldii, Eubacterium hallii, Oscillibacter sp. 57_20, Roseburia hominis, Phascolarctobacterium faecium, Collinsella stercoris, Bifidobacterium adolescentis, Roseburia intestinalis, Butyricimonas virosa, Alistipes indistinctus, Eubacterium sp. CAG: 38, Anaerostipes hadrus, Bacteroides caccae, Clostridium sp. CAG: 242, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Coprococcus catus, Bacteroides massiliensis, Roseburia faecis, Clostridium sp. CAG: 58, Ruthenibacterium lactatiformans, Bacteroides plebeius, Bilophila wadsworthia, Firmicutes bacterium CAG: 145, Eubacterium ramulus, Enterorhabdus caecimuris, Gemella morbillorum, Veillonella infantium, Haemophilus sp. HMSC71H05, Actinomyces sp. oral taxon 181, Blautia producta, Lactobacillus mucosae, Enterococcus avium, Veillonella atypica, Eubacterium sulci, Blautia hansenii, Rothia mucilaginosa, Clostridium spiroforme, Tyzzerella nexilis, Veillonella parvula, and Bacteroides fragilis;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from UC, optionally treating the individual.180.A method for monitoring disease activity of ulcerative colitis (UC) in an individual, comprising the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Clostridium leptum, Fusicatenibacter saccharivorans, Odoribacter splanchnicus Gemmiger formicili, Actinomyces sp. oral taxon 181 and Clostridium spiroforme;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score; andc) determining the individual as having increased disease activity if the risk score is higher than a cutoff value.181.A computer program product for determining the risk of, diagnosing, or suggesting treatment for Crohn’s Disease (CD) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Dorea formicigenerans, Eubacterium eligens, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (B. obeum) , Eubacterium ventriosum, Anaerostipes hadrus, Lawsonibacter asaccharolyticus, Oscillibacter sp. CAG: 241, Oscillibacter sp. 57_20, Lactobacillus rogosae, Eubacterium sp. CAG: 274, Romboutsia ilealis, Bacteroides fragilis, Escherichia coli, Eubacterium sulci, Actinomyces sp. oral taxon 181, and Actinomyces sp. S6-Spd3;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from CD if the risk score is higher than a cutoff value; andd) if the individual is determined to have an increased risk for or is suffering from CD, optionally suggesting an appropriate CD treatment to the individual.182.A computer program product for determining the risk of or diagnosing ulcerative colitis (UC) in an individual, wherein the computer program product comprises a computer readable medium encoded with computer executable code, wherein the computer executable code is configured to enable the execution of the steps of:a) determining the relative abundance of a set of bacterial species in a stool sample from the individual, wherein the set of bacterial species comprises one or more bacterial species selected from the group consisting of Phascolarctobacterium faecium, Asaccharobacter celatus, Collinsella stercoris, Oscillibacter sp. CAG: 241, Lawsonibacter asaccharolyticus, Butyricimonas virosa, Clostridium sp. CAG: 58, Eubacterium sp. CAG: 274, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Actinomyces sp. oral taxon 181, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii;b) comparing the relative abundance of each bacterial species in the set of bacterial species from the individual with the relative abundance of each bacterial species in the set of bacterial species in a reference data set to generate a risk score;c) determining the individual as having increased risk for or is suffering from UC if the risk score is higher than a cutoff value; andif the individual is determined to have an increased risk for or is suffering from UC, optionally suggesting an appropriate UC treatment to the individual.183.The method of any one of claims 2-47, wherein the first set of bacterial species and the second set of bacterial species each independently comprises two or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.184.The method of any one of claims 48-78, wherein the first set of bacterial species and the second set of bacterial species each independently comprises two or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.185.The method of any one of claims 79-112, wherein the first set of bacterial species and the second set of bacterial species each independently comprises two or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.186.The method of any one of claims 113-160, wherein the first set of bacterial species and the second set of bacterial species each independently comprises two or more bacterial species selected from the group consisting of Actinomyces sp. oral taxon 181, Bacteroides fragilis, Escherichia coli, Lawsonibacter asaccharolyticus, Eubacterium sp. CAG: 274, Roseburia inulinivorans, Roseburia intestinalis, Ruminococcus obeum (Blautia obeum) , Dorea formicigenerans, Bilophila wadsworthia, Clostridium leptum, Fusicatenibacter saccharivorans, Gemmiger formicilis, Odoribacter splanchnicus, Ruminococcus torques, Clostridium spiroforme, Gemella morbillorum, and Blautia hansenii.

Citation Information

Patent Citations

  • Biomarkers for rheumatoid arthritis and usage therof

    CN106795479A

  • Methods of determining colorectal cancer status in an individual

    US20200041510A1