Multi-omics diagnostic models for inflammatory bowel disease and their biomarker screening methods, applications, and reagent kits
By constructing a multi-omics diagnostic model, utilizing machine learning algorithms and recursive feature elimination methods, and combining gut microbial biomarkers, bacterial functional genes, and metabolite information, the lack of a gold standard in the diagnosis of inflammatory bowel disease has been addressed, achieving highly accurate non-invasive prediction and diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2026-03-10
AI Technical Summary
Current technologies lack a gold standard for the diagnosis of inflammatory bowel disease, leading to significant differences in validation characteristics among different populations and reducing the diagnostic reference value of gut microbiota and metabolites.
By utilizing machine learning algorithms and recursive feature elimination methods, and combining gut microbial biomarkers, bacterial functional genes, and metabolite information, a multi-omics diagnostic model was constructed. The gut microbial biomarkers, bacterial functional genes, and characteristic metabolites were trained using fecal metagenomics and metabolomics data to establish a non-invasive predictive model.
It achieves highly accurate (area under ROC curve 0.92–0.98) prediction of inflammatory bowel disease, provides a non-invasive diagnostic tool, assists in differentiating inflammatory bowel disease, and improves the reliability and accessibility of diagnosis.
Smart Images

Figure CN117253548B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a multi-omics diagnostic model for inflammatory bowel disease, as well as its biomarker screening method, application, and kit. Background Technology
[0002] Recent research has found that alterations in the gut microbiota and metabolites are associated with human health and disease, including inflammatory bowel disease (IBD). IBD is a chronic inflammatory disease that affects the gastrointestinal tract and includes two main forms: Crohn's disease (CD) and ulcerative colitis (UC). Millions of people worldwide are affected by IBD, and its incidence is shifting from developed to developing countries, highlighting the importance of early diagnosis.
[0003] Previous studies have reported some characteristics of the gut microbiota and metabolites, but the differences between different studies make it challenging to validate these characteristics in different populations, thereby reducing the diagnostic reference value of the microbiota and metabolome in IBD. Summary of the Invention
[0004] The purpose of this invention is to provide a multi-omics diagnostic model for inflammatory bowel disease, as well as its biomarker screening method, application, and reagent kit.
[0005] To address the above problems, this invention provides a method for screening biomarkers for inflammatory bowel disease, characterized by comprising:
[0006] Using machine learning algorithms, a recursive feature elimination method was employed, and the abundance information of gut microbial biomarkers was used to construct a model for screening gut microbial biomarkers; using machine learning algorithms, a recursive feature elimination method was employed, and the functional gene information of gut bacteria was used to construct a model for screening functional genes of gut bacteria; using machine learning algorithms, a recursive feature elimination method was employed, and the content information of metabolites in the gut was used to construct a model for screening characteristic metabolites.
[0007] Metagenomic and metabolomics data cohorts were obtained from internal sequencing data and publicly available data on "inflammatory bowel disease" and "gut microbiota". Some samples only have corresponding metagenomic data in the metagenomic data cohort; some samples only have corresponding metabolomics data in the metabolomics data cohort; and some samples have corresponding metagenomic data in both the metagenomic and metabolomics data cohorts.
[0008] Based on the metagenomic data queue, the model for screening gut microbiota biomarkers is trained to obtain gut microbiota biomarkers;
[0009] Based on the metagenomic data queue, the model for screening the functional genes of intestinal bacteria is trained to obtain the functional genes of intestinal bacteria.
[0010] Based on the metabolomics data queue, the model for screening characteristic metabolites is trained to obtain the characteristic metabolites.
[0011] Furthermore, in the above method, the gut microbiota markers are 31 in number, including:
[0012] 1.Bifidobacterium_adolescentis;
[0013] 2. Asaccharobacter_celatus;
[0014] 3. Gordonibacter_pamelaeae;
[0015] 4. Bacteroides_plebeius;
[0016] 5.Bacteroides_xylanisolvens;
[0017] 6.Barnesiella_intestinihominis;
[0018] 7. Clostridium sp_CAG_58;
[0019] 8.Intestinimonas_butyriciproducens;
[0020] 9.Lawsonibacter_asaccharolyticus;
[0021] 10. Eubacterium_eligens;
[0022] 11. Eubacterium_hallii;
[0023] 12.Eubacterium_sp_CAG_274;
[0024] 13. Eubacterium_ventriosum;
[0025] 14. Ruminococcus_torques;
[0026] 15. Coprococcus_comes;
[0027] 16.Dorea_longicatena;
[0028] 17.Fusicatenibacter_saccharivorans;
[0029] 18.Lachnospira_pectinoschiza;
[0030] 19. Roseburia_inulinivorans;
[0031] 20.Oscillibacter_sp_57_20;
[0032] 21.Faecalibacterium_prausnitzii;
[0033] 22. Gemmiger_formicilis;
[0034] 23. Clostridium leptum;
[0035] 24. Eubacterium_siraeum;
[0036] 25. Ruminococcus_bromii;
[0037] 26. Holdemania_filiformis;
[0038] 27. Firmicutes_bacterium_CAG_110;
[0039] 28. Firmicutes_bacterium_CAG_83;
[0040] 29.Phascolarctobacterium_faecium;
[0041] 30. Bilophila wadsworthia;
[0042] 31.Akkermansia_muciniphila.
[0043] Furthermore, in the above method, the intestinal bacteria have 25 functional genes, including: 1. K00180;
[0044] 2.K00343;
[0045] 3.K00394;
[0046] 4.K01846;
[0047] 5.K01873;
[0048] 6.K01921;
[0049] 7.K02071;
[0050] 8.K02536;
[0051] 9.K03284;
[0052] 10.K03709;
[0053] 11.K04564;
[0054] 12.K05305;
[0055] 13.K05341;
[0056] 14.K06142;
[0057] 15.K06334;
[0058] 16.K06975;
[0059] 17.K07099;
[0060] 18.K08999;
[0061] 19.K11534;
[0062] 20.K12511;
[0063] 21.K14445;
[0064] 22.K18704;
[0065] 23.K19068;
[0066] 24.K19506;
[0067] 25.K20627.
[0068] Furthermore, in the above method, the characteristic metabolites are 13 in number, including:
[0069] 1. HMDB0000008;
[0070] 2.HMDB0000062;
[0071] 3. HMDB0000139;
[0072] 4. HMDB0000148;
[0073] 5. HMDB0000191;
[0074] 6. HMDB0000192;
[0075] 7.HMDB0000214;
[0076] 8.HMDB0000517;
[0077] 9. HMDB0000535;
[0078] 10.HMDB0000619;
[0079] 11.HMDB0000641;
[0080] 12.HMDB0000929;
[0081] 13.HMDB0002226.
[0082] Furthermore, in the above method, the model for screening gut microbiota biomarkers is trained based on the metagenomic data queue to obtain gut microbiota biomarkers, including:
[0083] Step S31: Divide the metagenomic data queue into: a first training set, a first validation set, and a first test set; the data in the first training set, the first validation set, and the first test set do not overlap.
[0084] Step S32: Based on the first training set and the first validation set, perform cross-validation between queues and leave-one-out validation, and adjust the parameters of the model for screening gut microbiota markers based on the validation results;
[0085] Step S33: Based on the first test set, the model for screening gut microbiota markers is tested. If the test results meet the requirements, a non-inflammatory bowel disease cohort is obtained. Based on the current model for screening gut microbiota markers and its parameters, the false positive rate of the gut microbiota markers is verified. If the false positive rate of the gut microbiota markers meets the requirements, the corresponding gut microbiota markers are obtained based on the parameters of the current model for screening gut microbiota markers. If the false positive rate of the gut microbiota markers does not meet the requirements, steps S32 to S33 are repeated until the false positive rate of the gut microbiota markers meets the requirements. If the test results do not meet the requirements, steps S32 to S33 are repeated until the false positive rate of the gut microbiota markers meets the requirements.
[0086] Furthermore, in the above method, the model for screening functional genes of intestinal bacteria is trained based on a metagenomic data queue to obtain functional genes of intestinal bacteria, including:
[0087] Step S41: Based on the first training set and the first validation set, perform cross-validation between queues and leave-one-out validation, and adjust the model for screening functional genes of intestinal bacteria based on the validation results;
[0088] Step S42: Based on the first test set, the model for screening functional genes of intestinal bacteria is tested. If the test results meet the requirements, a non-inflammatory bowel disease cohort is obtained. Based on the current model for screening functional genes of intestinal bacteria and its parameters, the false positive rate of functional genes of intestinal bacteria in the non-inflammatory bowel disease cohort is verified. If the false positive rate of functional genes of intestinal bacteria meets the requirements, the corresponding functional genes of intestinal bacteria are obtained based on the parameters of the current model for screening functional genes of intestinal bacteria. If the false positive rate of functional genes of intestinal bacteria does not meet the requirements, steps S41 to S42 are repeated until the false positive rate of functional genes of intestinal bacteria meets the requirements. If the test results do not meet the requirements, steps S41 to S42 are repeated until the false positive rate of functional genes of intestinal bacteria meets the requirements.
[0089] Furthermore, in the above method, based on the metabolomics data queue, the model for screening characteristic metabolites is trained to obtain characteristic metabolites, including:
[0090] Step S51: Divide the metabolomics data queue into: a second training set, a second validation set, and a second test set;
[0091] Step S52: Based on the second training set and the second validation set, perform cross-validation between queues and leave-one-out-of-work validation, and adjust the parameters of the model for screening feature metabolites based on the validation results;
[0092] Step S53: Based on the second test set, the model for screening characteristic metabolites is tested. If the test results meet the requirements, a non-inflammatory bowel disease cohort is obtained. Based on the current model for screening characteristic metabolites and its parameters, the false positive rate of characteristic metabolites in the non-inflammatory bowel disease cohort is verified. If the false positive rate of characteristic metabolites meets the requirements, the corresponding gut microbiota markers are obtained based on the parameters of the current model for screening characteristic metabolites. If the false positive rate of characteristic metabolites does not meet the requirements, steps S52 to S53 are repeated until the false positive rate of characteristic metabolites meets the requirements. If the test results do not meet the requirements, steps S52 to S53 are repeated until the false positive rate of characteristic metabolites meets the requirements.
[0093] According to another aspect of the present invention, a multi-omics diagnostic model for inflammatory bowel disease is also provided, wherein the characteristic metabolites are 13, including:
[0094] The multi-omics diagnostic model for inflammatory bowel disease utilizes metagenomic data queues and metabolomics data queues as datasets, wherein the datasets are divided into: a third training set, a third validation set, and a third test set;
[0095] The multi-omics diagnostic model for inflammatory bowel disease is constructed based on the abundance information of gut microbial biomarkers, the functional gene information of gut bacteria, and the content information of characteristic metabolites.
[0096] The inflammatory bowel disease multi-omics diagnostic model is trained and validated based on the third training set and the third validation set, using cross-validation between cohorts and retention-of-samples verification.
[0097] The multi-omics diagnostic model for inflammatory bowel disease was independently validated based on the third test set.
[0098] According to another aspect of the present invention, an application of an inflammatory bowel disease marker is also provided, which is used in the preparation of a diagnostic reagent for inflammatory bowel disease, wherein the inflammatory bowel disease marker in the diagnostic reagent includes: intestinal microbial markers, functional genes of intestinal bacteria, and characteristic metabolites.
[0099] According to another aspect of the present invention, an inflammatory bowel disease (IBD) detection kit is also provided, characterized in that the IBD biomarkers in the detection kit include: gut microbiota biomarkers, functional genes of gut bacteria, and characteristic metabolites.
[0100] Compared to existing technologies, there is currently no gold standard for the diagnosis of inflammatory bowel disease (IBD), leading to many unidentifiable suspected cases in clinical practice. This invention utilizes information on the abundance of gut bacteria, the functional genes of gut bacteria, and the content of intestinal metabolites to establish a predictive model for IBD. Furthermore, this invention integrates these two omics approaches and three information dimensions for joint prediction, establishing a predictive model with impressive predictive capabilities (area under the ROC curve 0.92–0.98). The method and model of this invention can assist in the diagnosis and differentiation of IBD. Moreover, this predictive model is a non-invasive, non-surgical diagnostic model, facilitating its clinical application. Attached Figure Description
[0101] Figure 1 This is a flowchart of a method for screening biomarkers for inflammatory bowel disease according to an embodiment of the present invention. Detailed Implementation
[0102] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0103] like Figure 1 As shown, a method for screening biomarkers for inflammatory bowel disease is characterized by comprising:
[0104] Step S1: Using machine learning algorithms, a recursive feature elimination method is used, and the abundance information of gut microbial biomarkers is used to construct a model for screening gut microbial biomarkers; using machine learning algorithms, a recursive feature elimination method is used, and the functional gene information of gut bacteria is used to construct a model for screening functional genes of gut bacteria; using machine learning algorithms, a recursive feature elimination method is used, and the content information of metabolites in the gut is used to construct a model for screening characteristic metabolites.
[0105] Step S2: Obtain metagenomic data queues and metabolomics data queues from internal sequencing and publicly available data on "inflammatory bowel disease" and "gut microbiota," respectively. Some samples have corresponding metagenomic data only in the metagenomic data queue; some samples have corresponding metabolomics data only in the metabolomics data queue; and some samples have corresponding metagenomic data in both the metagenomic data queue and the metabolomics data queue.
[0106] Step S3: Based on the metagenomic data queue, train the model for screening gut microbiota biomarkers to obtain gut microbiota biomarkers;
[0107] Here, the metagenomic data queue can be divided into: a first training set, a first validation set, and a first test set;
[0108] Step S4: Based on the metagenomic data queue, train the model for screening the functional genes of intestinal bacteria to obtain the functional genes of intestinal bacteria.
[0109] Step S5: Based on the metabolomics data queue, train the model for screening characteristic metabolites to obtain characteristic metabolites.
[0110] Inflammatory bowel disease (IBD) is a serious chronic gastrointestinal disease with an insidious onset. Establishing mathematical prediction models for IBD by detecting certain biological markers in subject samples such as blood, feces, and saliva is an effective and feasible method for predicting IBD, facilitating prevention and early warning efforts in large-scale populations.
[0111] Previous studies have confirmed that the development and progression of inflammatory bowel disease are closely related to the gut microbiota. This invention studies the gut microbiota and its metabolites of subjects using two fecal detection methods (fecal metagenomic analysis and fecal metabolomics analysis): fecal metagenomic analysis allows us to obtain information on the types of bacteria in the subject's gut and the proportion (abundance) of each type, as well as the functional gene information of the bacteria; fecal metabolomics allows us to obtain the content of various metabolites in the subject's gut.
[0112] Fecal shotgun metagenomics offers a powerful tool for identifying disease-associated species and understanding the synergistic metabolism between the host and microbiome at a higher taxonomic resolution, while metabolomics reveals changes in gut metabolites, which act as messengers in communication between the gut microbiome and the host. The combination of metagenomics and metabolomics provides a promising approach for understanding the development of IBD and related changes in the gut environment, and offers a non-invasive biomarker for IBD.
[0113] This invention utilizes fecal metagenomic sequencing to powerfully identify disease-associated species and gain deeper insights into the co-metabolic relationships between the host and the microbiome. Simultaneously, fecal metabolomics reveals changes in gut metabolites, serving as a medium for communication between the gut microbiome and the host. Combining metagenomics and metabolomics provides a promising approach for understanding the development of IBD and related changes in the gut environment, and offers a non-invasive IBD biomarker.
[0114] Currently, there is no gold standard for diagnosing inflammatory bowel disease (IBD), leading to many unidentifiable suspected cases in clinical practice. This invention utilizes information on the abundance of gut bacteria, the functional genes of gut bacteria, and the content of intestinal metabolites to establish a predictive model for IBD. Furthermore, this invention integrates these two omics approaches and three information dimensions for joint prediction, establishing a predictive model with impressive predictive capabilities (area under the ROC curve 0.92–0.98). The method and model of this invention can assist in the diagnosis and differentiation of IBD.
[0115] Here, the metabolomics data queue can be divided into: the second training set, the second validation set, and the second test set;
[0116] In one embodiment of the inflammatory bowel disease marker screening method of the present invention, the intestinal microbial markers are 31 in number, including:
[0117] 1.Bifidobacterium_adolescentis;
[0118] 2. Asaccharobacter_celatus;
[0119] 3. Gordonibacter_pamelaeae;
[0120] 4. Bacteroides_plebeius;
[0121] 5.Bacteroides_xylanisolvens;
[0122] 6.Barnesiella_intestinihominis;
[0123] 7.Clostridium_sp_CAG_58;
[0124] 8. Intestinimonas_butyriciproducens;
[0125] 9.Lawsonibacter_asaccharolyticus;
[0126] 10. Eubacterium_eligens;
[0127] 11. Eubacterium_hallii;
[0128] 12. Eubacterium_sp_CAG_274;
[0129] 13. Eubacterium_ventriosum;
[0130] 14.Ruminococcus_torques;
[0131] 15. Coprococcus_comes;
[0132] 16. Dorea_longicatena;
[0133] 17.Fusicatenibacter_saccharivorans;
[0134] 18.Lachnospira_pectinoschiza;
[0135] 19. Roseburia_inulinivorans;
[0136] 20.Oscillibacter_sp_57_20;
[0137] 21. Faecalibacterium_prausnitzii;
[0138] 22.Gemmiger_formicilis;
[0139] 23.Clostridium_leptum;
[0140] 24. Eubacterium_siraeum;
[0141] 25.Ruminococcus_bromii;
[0142] 26. Holdemania_filiformis;
[0143] 27. Firmicutes_bacterium_CAG_110;
[0144] 28. Firmicutes_bacterium_CAG_83;
[0145] 29.Phascolarctobacterium_faecium;
[0146] 30. Bilophila wadsworthia;
[0147] 31.Akkermansia_muciniphila.
[0148] In one embodiment of the inflammatory bowel disease biomarker screening method of the present invention, the intestinal bacteria contain 25 functional genes, including:
[0149] 1.K00180;
[0150] 2.K00343;
[0151] 3.K00394;
[0152] 4.K01846;
[0153] 5.K01873;
[0154] 6.K01921;
[0155] 7.K02071;
[0156] 8.K02536;
[0157] 9.K03284;
[0158] 10.K03709;
[0159] 11.K04564;
[0160] 12.K05305;
[0161] 13.K05341;
[0162] 14.K06142;
[0163] 15.K06334;
[0164] 16.K06975;
[0165] 17.K07099;
[0166] 18.K08999;
[0167] 19.K11534;
[0168] 20.K12511;
[0169] 21.K14445;
[0170] 22.K18704;
[0171] 23.K19068;
[0172] 24.K19506;
[0173] 25.K20627.
[0174] In one embodiment of the inflammatory bowel disease biomarker screening method of the present invention, the characteristic metabolites are 13, including:
[0175] 1. HMDB0000008;
[0176] 2.HMDB0000062;
[0177] 3. HMDB0000139;
[0178] 4. HMDB0000148;
[0179] 5. HMDB0000191;
[0180] 6. HMDB0000192;
[0181] 7.HMDB0000214;
[0182] 8.HMDB0000517;
[0183] 9. HMDB0000535;
[0184] 10.HMDB0000619;
[0185] 11.HMDB0000641;
[0186] 12.HMDB0000929;
[0187] 13.HMDB0002226.
[0188] In one embodiment of the inflammatory bowel disease biomarker screening method of the present invention, step S3, based on the metagenomic data queue, trains the model for screening gut microbiota biomarkers to obtain gut microbiota biomarkers, including:
[0189] Step S31: Divide the metagenomic data queue into: a first training set, a first validation set, and a first test set;
[0190] Here, the data in the first training set, the first validation set, and the first test set do not overlap.
[0191] Step S32: Based on the first training set and the first validation set, perform cross-validation between queues and leave-one-out validation, and adjust the parameters of the model for screening gut microbiota markers based on the validation results;
[0192] Step S33: Based on the first test set, the model for screening gut microbiota markers is tested. If the test results meet the requirements, a non-inflammatory bowel disease cohort is obtained. Based on the current model for screening gut microbiota markers and its parameters, the false positive rate of the gut microbiota markers is verified. If the false positive rate of the gut microbiota markers meets the requirements, the corresponding gut microbiota markers are obtained based on the parameters of the current model for screening gut microbiota markers. If the false positive rate of the gut microbiota markers does not meet the requirements, steps S32 to S33 are repeated until the false positive rate of the gut microbiota markers meets the requirements. If the test results do not meet the requirements, steps S32 to S33 are repeated until the false positive rate of the gut microbiota markers meets the requirements.
[0193] Here, the method for constructing a metagenomic bacterial abundance model to predict the occurrence of inflammatory bowel disease can be achieved by using internal sequencing and publicly available metagenomic cohorts of "inflammatory bowel disease" and "gut microbiota" to obtain a dataset. For the dataset, feature screening is performed on the included microbial markers related to "inflammatory bowel disease". Using machine learning algorithms, a recursive feature elimination method is used to screen out the characteristic microorganisms for constructing the model, resulting in 31 gut microbial markers.
[0194] Specifically, the stability and reproducibility of the model can be evaluated by cross-validation between cohorts and leave-one-out validation across all studies in the training set; independent validation can be performed in the independent validation cohort; and the false positive rate of biomarkers can be validated in the non-inflammatory bowel disease cohort.
[0195] Step S4: Based on the metagenomic data queue, train the model for screening functional genes of gut bacteria to obtain functional genes of gut bacteria, including:
[0196] Step S41: Based on the first training set and the first validation set, perform cross-validation between queues and leave-one-out validation, and adjust the model for screening functional genes of intestinal bacteria based on the validation results;
[0197] Step S42: Based on the first test set, the model for screening functional genes of intestinal bacteria is tested. If the test results meet the requirements, a non-inflammatory bowel disease cohort is obtained. Based on the current model for screening functional genes of intestinal bacteria and its parameters, the false positive rate of functional genes of intestinal bacteria in the non-inflammatory bowel disease cohort is verified. If the false positive rate of functional genes of intestinal bacteria meets the requirements, the corresponding functional genes of intestinal bacteria are obtained based on the parameters of the current model for screening functional genes of intestinal bacteria. If the false positive rate of functional genes of intestinal bacteria does not meet the requirements, steps S41 to S42 are repeated until the false positive rate of functional genes of intestinal bacteria meets the requirements. If the test results do not meet the requirements, steps S41 to S42 are repeated until the false positive rate of functional genes of intestinal bacteria meets the requirements.
[0198] Specifically, the method for constructing a model for predicting the occurrence of inflammatory bowel disease (IBD) based on the abundance of metagenomic bacterial functional genes can utilize internal sequencing and publicly published studies on "inflammatory bowel disease" and "gut microbiota" as datasets. For the datasets, feature screening is performed on the functional genes of the gut bacteria included that are associated with "inflammatory bowel disease". Using machine learning algorithms, a recursive feature elimination method is used to screen out the functional genes of the gut bacteria used to construct the model, resulting in 25 functional genes of gut bacteria.
[0199] Specifically, cross-validation and leave-one-out validation can be performed across all studies in the training set to assess the model's stability and reproducibility; independent validation can be performed in the independent validation cohort. The false positive rate of biomarkers can also be validated in the non-inflammatory bowel disease cohort.
[0200] Step S5: Based on the metabolomics data queue, train the model for screening characteristic metabolites to obtain characteristic metabolites, including:
[0201] Step S51: Divide the metabolomics data queue into: a second training set, a second validation set, and a second test set;
[0202] Here, the data in the second training set, the second validation set, and the second test set do not overlap with each other;
[0203] Step S52: Based on the second training set and the second validation set, perform cross-validation between queues and leave-one-out-of-work validation, and adjust the parameters of the model for screening feature metabolites based on the validation results;
[0204] Step S53: Based on the second test set, the model for screening characteristic metabolites is tested. If the test results meet the requirements, a non-inflammatory bowel disease cohort is obtained. Based on the current model for screening characteristic metabolites and its parameters, the false positive rate of characteristic metabolites in the non-inflammatory bowel disease cohort is verified. If the false positive rate of characteristic metabolites meets the requirements, the corresponding gut microbiota markers are obtained based on the parameters of the current model for screening characteristic metabolites. If the false positive rate of characteristic metabolites does not meet the requirements, steps S52 to S53 are repeated until the false positive rate of characteristic metabolites meets the requirements. If the test results do not meet the requirements, steps S52 to S53 are repeated until the false positive rate of characteristic metabolites meets the requirements.
[0205] Here, the metabolomics data queue can be divided into: a second training set, a second validation set, and a second test set.
[0206] The method for constructing a metabolomics-based model for predicting the occurrence of inflammatory bowel disease (IBD) can utilize internal sequencing data and published studies on IBD and fecal metabolomics as datasets. For these datasets, feature screening is performed on the fecal metabolites associated with IBD. Using machine learning algorithms, a recursive feature elimination method is employed to select the feature metabolites for model construction, resulting in 13 feature metabolites.
[0207] Specifically, cross-validation between cohorts and retain-one-sample validation are performed across all studies in the training set to assess the model's stability and reproducibility; independent validation can also be performed in the independent validation cohort. Furthermore, the false positive rate of biomarkers can be validated in the non-inflammatory bowel disease cohort.
[0208] According to another aspect of the present invention, a multi-omics diagnostic model for inflammatory bowel disease is also provided, comprising:
[0209] The multi-omics diagnostic model for inflammatory bowel disease utilizes metagenomic and metabolomics data queues as datasets, wherein the datasets are divided into a third training set, a third validation set, and a third test set; the data in the third training set, the third validation set, and the third test set do not overlap with each other;
[0210] The multi-omics diagnostic model for inflammatory bowel disease is constructed based on the abundance information of gut microbial biomarkers, the functional gene information of gut bacteria, and the content information of characteristic metabolites.
[0211] The multi-omics diagnostic model for inflammatory bowel disease is trained and validated based on the third training set and the third validation set, using cross-validation between cohorts and retention-of-samples validation to assess the model's stability and reproducibility.
[0212] The multi-omics diagnostic model for inflammatory bowel disease was independently validated based on the third test set.
[0213] Specifically, this embodiment can use internal sequencing and publicly published research on "inflammatory bowel disease" and "intestinal fecal metagenomics and metabolomics" as datasets. Based on the datasets, a multi-omics prediction model for inflammatory bowel disease can be constructed using the aforementioned 31 intestinal microbial biomarkers, 25 functional genes and characteristic metabolites of intestinal bacteria.
[0214] Preferably, the dataset is divided into a third training set, a third validation set, and a third test set; cross-validation between cohorts and retention-of-samples validation can be performed in all studies of the training set to evaluate the stability and reproducibility of the model; and independent validation can be performed in the independent validation cohort.
[0215] Preferably, the modeling method of the present invention may be, for example, a random forest model with a sample size (dataset) of: nine metagenomic cohorts (1363 cases) and four metabolomics cohorts (398 cases);
[0216] This invention can reveal the relationship between the gut microbiota and its functions and gut metabolites through multinational, large-scale cohorts, multi-omics characterization, standardized sampling and analysis, and model systems. Cross-cohort comprehensive analysis (CCIA) promises to address these challenges by comparing multiple metagenomic case-control studies to assess the robustness of disease-microbiome associations. The goal of CCIA is to discover consistent associations across different cohorts, thereby minimizing the influence of biological or technological confounding factors. CCIA has demonstrated strong performance in previous studies, highlighting its potential as an important tool across various research fields.
[0217] This invention comprehensively analyzed nine metagenomic cohorts (1363 cases) and four metabolomics cohorts (398 cases) of IBD patients from different countries or regions using CCIA. Our aim was to discover novel gut bacteria, metabolites, and associated Kyoto Genome Encyclopedia (KEGG) selected (KO) genes that contribute to the progression of IBD in different cohorts, and to establish diagnostic models using disease-specific biomarkers from different IBD cohorts.
[0218] Model validation method: 10x cross-validation, with 9 copies used as the training set to train the model and the remaining copy used as the validation set to evaluate the model performance.
[0219] • A predictive model built using abundance information of gut bacteria: including 31 differentially expressed bacteria.
[0220] 1.Bifidobacterium_adolescentis
[0221] 2.Asaccharobacter_celatus
[0222] 3.Gordonibacter_pamelaeae
[0223] 4.Bacteroides_plebeius
[0224] 5.Bacteroides_xylanisolvens
[0225] 6.Barnesiella_intestinehominis
[0226] 7.Clostridium_sp_CAG_58
[0227] 8.Intestinimonas_butyriciproducens
[0228] 9. Lawsonibacter_asaccharolyticus
[0229] 10.Eubacterium_eligens
[0230] 11.Eubacterium_hallii
[0231] 12.Eubacterium_sp_CAG_274
[0232] 13.Eubacterium_ventriosum
[0233] 14.Ruminococcus_torques
[0234] 15.Coprococcus_comes
[0235] 16.Dorea_longicatena
[0236] 17. Fusicatenibacter_saccharivorans
[0237] 18.Lachnospira_pectinoschiza
[0238] 19.Roseburia_inulinivorans
[0239] 20.Oscillibacter_sp_57_20
[0240] 21.Faecalibacterium_prausnitzii
[0241] 22. Gemmiger_formicilis
[0242] 23. Clostridium leptum
[0243] 24.Eubacterium_siraeum
[0244] 25. Ruminococcus_bromii
[0245] 26. Holdemania_filiformis
[0246] 27.Firmicutes_bacterium_CAG_110
[0247] 28.Firmicutes_bacterium_CAG_83
[0248] 29.Phascolarctobacterium_faecium
[0249] 30. Bilophila_wadsworthia
[0250] 31.Akkermansia_muciniphila
[0251] • A predictive model built using functional gene information from gut bacteria: This model includes 25 differentially expressed bacterial functional genes.
[0252] 1.K00180
[0253] 2.K00343
[0254] 3.K00394
[0255] 4.K01846
[0256] 5.K01873
[0257] 6.K01921
[0258] 7.K02071
[0259] 8.K02536
[0260] 9.K03284
[0261] 10.K03709
[0262] 11.K04564
[0263] 12.K05305
[0264] 13.K05341
[0265] 14.K06142
[0266] 15.K06334
[0267] 16.K06975
[0268] 17.K07099
[0269] 18.K08999
[0270] 19.K11534
[0271] 20.K12511
[0272] 21.K14445
[0273] 22.K18704
[0274] 23.K19068
[0275] 24.K19506
[0276] 25.K20627
[0277] • A predictive model built using information on the content of metabolites in the gut: This model includes 11 differentially expressed metabolites.
[0278] 1.HMDB0000008
[0279] 2.HMDB0000062
[0280] 3.HMDB0000139
[0281] 4.HMDB0000148
[0282] 5.HMDB0000191
[0283] 6.HMDB0000192
[0284] 7.HMDB0000214
[0285] 8.HMDB0000517
[0286] 9.HMDB0000535
[0287] 10.HMDB0000619
[0288] 11.HMDB0000641
[0289] 12.HMDB0000929
[0290] 13.HMDB0002226
[0291] Currently, there is no gold standard for the diagnosis of inflammatory bowel disease, and many suspected cases that cannot be distinguished in clinical practice will appear. The method and model of this invention can also assist in the diagnosis and differentiation of inflammatory bowel disease.
[0292] According to another aspect of the present invention, an application of an inflammatory bowel disease marker is also provided, which is used in the preparation of a diagnostic reagent for inflammatory bowel disease, wherein the inflammatory bowel disease marker in the diagnostic reagent includes: intestinal microbial markers, functional genes of intestinal bacteria, and characteristic metabolites.
[0293] According to another aspect of the present invention, an inflammatory bowel disease (IBD) detection kit is also provided, wherein the IBD biomarkers in the detection kit include: gut microbiota biomarkers, functional genes of gut bacteria, and characteristic metabolites.
[0294] In detail, firstly, the workflow for cross-cohort integrated analysis of fecal metagenomics and metabolomics in IBD of this invention is as follows:
[0295] This invention employs a multi-omics approach integrating fecal metagenomics and metabolomics to investigate alterations in the gut microbiota of individuals with IBD. This study included nine metagenomic cohorts (N=1363 cases) from four different regions or countries. These cohorts were divided into six discovery cohorts and three validation cohorts. In addition, we included four metabolomics cohorts (N=398 cases), with two external cohorts using untargeted metabolomics and two internal cohorts using targeted metabolomics.
[0296] To ensure consistency in bioinformatics analysis, we used MetaPhlan3 for classification analysis and HUMAnN3 for functional analysis to reprocess all raw sequencing data. Furthermore, by annotating metabolite names with uniform IDs using the Human Metabolome Database (HMDB), we identified 79 metabolites common to the four cohorts. These metabolites will be used for cross-cohort analysis of the metabolomics data.
[0297] Furthermore, our goal was to reveal the variation patterns of the gut microbiota and its metabolites through comprehensive statistical analysis, and then utilize machine learning techniques to diagnose IBD. First, we excluded samples collected repeatedly from external cohorts. Then, using a series of differential analyses and feature selection, we identified 31 species, 25 functional genes, and 13 metabolites effective for diagnosing IBD patients. Subsequently, we selected four cohorts (N=391 cases) containing metagenomic and metabolomics data for comprehensive analysis, thereby establishing the most accurate diagnostic model.
[0298] Second, this invention uses a cross-cohort model to demonstrate changes in the gut microbiota of IBD patients, identifying bacterial biomarkers for diagnosing IBD at the species level:
[0299] To identify potential microbial biomarkers for diagnosing IBD, we employed a method from a previous study to analyze the composition of microbial species. With an FDR less than 0.0001, 74 microorganisms with significantly different abundances in the gut microbiome were identified using CCIA. Subsequently, we utilized a machine learning approach (Random Forest, RF) to diagnose IBD. To improve model accuracy and interpretability and minimize the impact of redundant and irrelevant features, we employed Iterative Feature Elimination (IFE) for feature selection. Therefore, we rigorously selected 31 characteristic species from the 74 differentially abundant species for modeling analysis; these species primarily belong to the Fungi phylum. We initially built a Random Forest model using 10-fold cross-validation with the 31 characteristic species from six cohorts. This model demonstrated strong IBD detection capability across all cohorts, with AUROCs ranging from 0.66 to 0.95. However, when analyzing the classifier's performance across the six cohorts, we found that the results for the IBDMDB cohort were significantly worse than the other five cohorts. This difference in results may be due to the fact that patients, although from the same country, came from different centers. This variability and heterogeneity of the data may contribute to the reduced accuracy of the classifier. Furthermore, to assess the transferability and geographic diversity of identified features used to diagnose IBD, we performed inter-cohort transfer analysis and LOCO analysis using established methods. For AUROC, the average performance of cohort-to-cohort transfer analysis for species-level models ranged from 0.79 to 0.86, with most values hovering around 0.8. Our analysis showed that the performance of LOCO analysis ranged from 0.71 to 0.94.
[0300] Furthermore, to validate the accuracy and transferability of our model in independent cohorts, we included three independent IBD metagenomic cohorts. Specifically, the mean AUROC was 0.70 for the HallAB 2017 cohort, 0.90 for the FranzosaEA2019B cohort, and 0.89 for the Pudong cohort. However, the AUROC in LOCO analysis showed a slight improvement: 0.72 for the HallAB 2017 cohort, 0.96 for the FranzosaEA2019B cohort, and 0.91 for the Pudong cohort. If we consider the HallAB 2017 cohort as an outlier, then the independent validation of our model yields an AUROC of approximately 0.90.
[0301] Previous research has revealed that changes in the microbiome may be associated with various diseases, highlighting the importance of identifying disease-specific microbiome signatures. Next, we investigated the false positive rate (FPR) of the metagenomic classifier by analyzing the metagenomics of patients with gastrointestinal (GI) diseases (such as adenomas and colorectal cancer (CRC)) and non-GI diseases (such as type 2 diabetes (T2D)). Therefore, we used the species-specific LOCO classification model, achieving FPRs of 0.08 and 0.13 for the CRC dataset, respectively, after calibration. We also found relatively low FPRs for other disease datasets, with 0.15 for adenomas and 0.11 for T2D. These results demonstrate that our model exhibits excellent disease specificity. In conclusion, our findings indicate that our model possesses excellent specificity and can accurately identify disease-specific microbiome signatures.
[0302] Third, this invention identifies IBD diagnostic markers by demonstrating changes in the orthogonal genome (KO) of the Kyoto Genome Encyclopedia (KEGG) in different IBD cohorts:
[0303] Metagenomic functional analysis plays a crucial role in understanding the complex interactions between the gut microbiome and human health. In our study, we first used the KEGG orthogonal database to annotate gene families obtained from metagenomic analysis as KO genes, resulting in 9270 KO genes. We employed a low abundance screening step, ultimately obtaining 3732 KO genes, and then performed abundance differential analysis using the same methods as before. To avoid model overfitting, we employed a rigorous FDR method (CCIA analysis FDR < 1 × 10⁻¹²), thereby identifying 162 KO genes differentially expressed between normal individuals and IBD patients.
[0304] Furthermore, we sought to evaluate the potential role of KO genes in IBD diagnosis. We used the IFE method for feature selection, identifying 25 KO genes from 162 as features for random forest modeling. After 10-fold cross-validation, we observed that, except for the HeQ 2017 cohort which showed excellent diagnostic performance (AUROC: 0.98), the AUROC values of other cohorts were lower compared to the bacterial species model. Inter-cohort transfer analysis showed that the mean AUROC values for all cohorts ranged from 0.74 to 0.81. Consistent with the bacterial species model, LOCO analysis showed that the diagnostic value of KO genes was slightly higher than that of the inter-cohort transfer analysis. Overall, although the diagnostic performance of the KO gene model was slightly lower than that of the bacterial species model, its diagnostic sensitivity remained acceptable.
[0305] Furthermore, to validate the diagnostic potential of the 25 KO genes, we applied them to the aforementioned independent cohorts. In the cohort-to-cohort transfer model, the AUROC ranged from 0.71 to 0.96, while in the LOCO classification model, the AUROC ranged from 0.61 to 0.87. We also tested the FPR by analyzing other disease datasets, including adenoma, CRC, and T2D. These KO genes also exhibited relatively low FPRs in the non-diabetic cohort: 0.04 for adenoma, 0.13 for CRC, and 0.07 for T2D, indicating excellent disease specificity.
[0306] We also used the EggNOG homology classification method to diagnose IBD, but the accuracy of IBD detection was slightly lower compared to using the KO model. The AUROC values for inter-cohort transfer validation ranged from 0.63 to 0.80, and the AUROC values for LOCO validation ranged from 0.65 to 0.92. Considering the accuracy and interpretability of the results, we decided to use only the annotation results from the KO database in subsequent analyses.
[0307] Fourth, this invention demonstrates the metabolomic changes in different IBD populations and uses characteristic metabolites to diagnose IBD:
[0308] We are intrigued by the intricate interactions between the gut microbiota and host co-metabolism, and therefore sought to further explore the profile of fecal metabolite changes through metabolomics. Using targeted metabolomics, we delved into the molecular complexity of fecal samples, meticulously assessing the overall differences in fecal metabolites between IBD and healthy samples. To investigate metabolic differences between IBD patients and controls, we used PCoA analysis and the OPLS-DA model, revealing significant differences in the metabolomics composition between the two groups. Our results indicate that no specific set of metabolites is superior to the collection of all metabolites in distinguishing between IBD patients and healthy individuals. However, amino acid compounds exhibited the strongest discriminative power.
[0309] To further understand the differences in metabolites between IBD patients and healthy controls, we performed differential analysis, identifying 78 metabolites. Our next goal was to investigate whether a specific set of metabolites could serve as an accurate diagnostic tool for IBD, regardless of whether it was a targeted or non-targeted dataset obtained through the CCIA method. We then integrated four metabolomics studies, identifying 79 metabolites using HMDB ID, which were prevalent across all four cohorts. Subsequently, we performed univariate differential analysis (FDR < 0.0001) and OPLS-DA analysis (VIP score > 1), ultimately identifying 32 candidate metabolites. To further refine the screening, we used the IFE method to narrow down the metabolite library to 13. To account for differences in metabolite detection methods and numerical units between internal and external cohorts, we restricted cross-validation to cohorts with the same units. In the analysis of the domestic cohort, our model achieved an AUROC of 0.945 for 10-fold cross-validation, while LOOCV performed slightly lower at 0.937, and the AUROC for the independent validation cohort was 0.867. In the US-Netherlands cohort, the AUROC values for 10-fold cross-validation and LOOCV were similar, both exceeding 0.9, while the AUROC for the independent validation cohort was 0.841, demonstrating good performance. Based on these results, it can be inferred that metabolomics has greater potential for disease diagnosis compared to metagenomics and is expected to become a diagnostic biomarker in the future.
[0310] To further validate the disease specificity of our characteristic metabolites, we included four non-biodiversity disease metabolomics cohorts: one adenoma cohort, two CRC cohorts, and one T1D cohort. Differential analysis showed that most of the 13 metabolites we used to diagnose IBD did not show significant differences across these cohorts (FDR > 0.05). These findings confirm the disease specificity of our characteristic metabolites in IBD.
[0311] Fifth, this invention integrates multi-omics features to diagnose IBD in different cohorts:
[0312] We previously identified three panels consisting of 31 species, 25 KO genes, and 13 metabolites that accurately distinguished IBD patients from healthy controls. To investigate whether integrating multiple data sources could improve diagnostic accuracy, we further examined the interactions between the gut microbiota and its metabolites. We first combined species and KO genes to differentiate IBD, achieving satisfactory diagnostic results in 10-fold cross-validation and LOOCV. The AUROC value exceeded 0.97 in the domestic cohort, while the AUROC value in the international cohort increased to over 0.9 (compared to using a single species or KO gene). The AUROC values for these panels also exceeded 0.9 in independent validation. Subsequently, we combined species and metabolites, and metabolites and KO genes, finding that their diagnostic performance was generally higher than 0.9 (AUROC). In particular, the combination of species and metabolites performed best in the combined panels, with diagnostic performance superior to individual panels.
[0313] After achieving good diagnostic performance with the aforementioned model, we further explored whether combining all panels could improve its diagnostic performance. We were pleasantly surprised to find that combining features significantly improved the diagnostic performance of the random forest model. In the domestic cohort, the AUROC value for 10-fold cross-validation was 0.98, while the AUROC value for independent validation was 0.96. In the US-Netherlands cohort, the AUROC value for 10-fold cross-validation was 0.93, and the AUROC value for independent validation was 0.92. Our results indicate that combining species, KO genes, and metabolites can significantly improve the diagnostic performance of our model in fecal metagenomics and metabolomics analysis.
[0314] Sixth, identify multiple sets of biological markers used to distinguish IBD subtypes:
[0315] Since the treatment strategies for UC and CD differ significantly, our next goal was to identify a subset of biomarkers from the aforementioned multi-omics panel that could distinguish between these two subtypes of IBD. Using the IFE method, we selected 12 features from the multi-omics panel. RF modeling showed that the selected 12 biomarkers effectively distinguished between UC and CD. In both internal and external cohorts, the AUROC value for 10-fold cross-validation and LOOCV was approximately 0.8, while the AUROC value for independent validation was higher than 0.7. This subpanel can help clinicians further classify the disease after a diagnosis of IBD.
[0316] Seventh, differential abundance analysis was used to identify gut microbial species and functional genes:
[0317] The significance of differential abundance (DA) between groups was assessed using the "coin" package in R and the blocking Wilcoxon test. The test was performed separately for each species or gene, and the data were partitioned by cohort to control for confounding effects from varying cohort composition. Within each cohort, permutations were performed to obtain a conditional empty distribution to account for variations in block size and composition. The p-values were adjusted using the FDR method to account for multiple hypothesis testing. Furthermore, the generalized fold change (gFC) method was used to calculate the degree of difference between the control and IBD samples.
[0318] Eighth, identification of differentially expressed metabolites related to IBD:
[0319] We identified differentially metabolites based on two criteria:
[0320] (1) Using a nonparametric univariate method (Wilcoxon rank-sum test), the false discovery rate (FDR) was <0.0001. The FDR correction for the p-value of each metabolite was performed using a two-tailed Mann-Whitney U test.
[0321] (2) The significance of predictor variables (VIP scores) greater than 1 was determined using the OPLS-DA model. OPLS-DA, short for Orthogonal Projections to Latent Structures Discriminant Analysis, is a multivariate statistical method used to analyze multivariate data. The significance of the projected variables is a measure of how much a metabolite contributes to the separation between the two groups being compared. The OPLS-DA model validation utilized the permutation test in the "ropls" R software package.
[0322] Eighth, preprocessing of metabolomics data:
[0323] Given the significant differences in metabolomics techniques, processing methods, and outputs across different studies, we preprocessed the metabolomics dataset to facilitate subsequent cross-cohort analysis. To this end, we used the MetaboAnalyst (5.0) compound ID conversion procedure to normalize metabolite names in both internal and external cohorts to a common HMDB ID, identifying 79 metabolites common across all four cohorts. Subsequently, we performed a log2 transformation on the metabolite values and then converted them to z-scores.
[0324] Ninth, elimination of recursive features:
[0325] To improve the reliability and robustness of the model while reducing its size and complexity, we employed the Recursive Feature Elimination (IFE) feature selection method in Python. First, we performed differential feature analysis to identify latent features. Then, we trained a Random Forest (RF) model using the scikit-learn package and performed stratified 10-fold cross-validation to distinguish between IBD and the normal control group. We used stratified 10-fold cross-validation to reasonably distribute the training and testing datasets. Next, we used the Recursive Feature Elimination (IFE) step to improve the performance of subsequent RF models. Finally, we selected the most important features from the best-performing model (the one with the highest AUROC value) as the final features for modeling.
[0326] Tenth, multi-omics statistical modeling workflow and model evaluation:
[0327] Given the excellent performance of the Random Forest model (an ensemble machine learning approach) in microbial data classification, our machine learning model also adopts this model. First, we used a 10-fold cross-validation technique commonly used in machine learning and statistical analysis in the cohorts. This involves dividing the available data into 10 equal parts, training the model on 9 of them, and evaluating its performance on the remaining parts. In inter-cohort transfer validation, we train the classifier in one cohort and then test it on all other cohorts. In "Outlier-Outlier (LOCO)" validation, we use the data from one cohort as the outer validation set and then train the model on the remaining data from all other cohorts. Then, we use the same nested cross-validation procedure as in inter-cohort transfer validation. Through these methods, we can evaluate the generality of the metagenomic classifier and its good performance on multiple cohorts of data. Single-cohort cross-validation (LOOCV) involves removing a sample (or observation) from the dataset and then training the model on the remaining data. We then use the removed sample as the validation dataset to evaluate the model's performance. We repeat this process for each sample in the dataset, ensuring that each sample is used as a validation dataset once. Data preprocessing, model building, and model evaluation were performed using the following R packages: SIAMCAT (v.1.14.0), caret (v.6.0.90), randomForest (v.4.7.1.1), pROC (v.1.18.0), and ROCR (v.1.0.11).
[0328] Eleventh, independent validation of external metagenomic cohorts:
[0329] To ensure the reliability of metagenomic signatures as diagnostic biomarkers for IBD, we validated our findings using three independent datasets from the United States and China. Following the same procedure used to build models in the discovery cohort, we performed inter-cohort and LOCO analyses to assess the strength and consistency of the identified biomarkers.
[0330] Validate the specificity of microbial biomarkers in non-IBD cohorts.
[0331] To minimize the risk of misdiagnosing IBD, we assessed the specificity of metagenomic markers by analyzing the AUROC values of models constructed using the most effective feature sets. Our analysis included patients with non-IBD diseases such as colorectal cancer (60 and 65 controls, from PRJEB27928, WirbelJ 2018 cohort), type 2 diabetes (45 and 39 controls, from PRJEB1786, KarlssonFH 2013 cohort), and adenoma (47 and 61 controls, from PRJEB7774, FengQ 2015 cohort).
[0332] Twelfth, validate the specificity of metabolic biomarkers in non-IBD cohorts.
[0333] To validate the disease specificity of the selected characteristic metabolites, we integrated four metabolomics cohorts, including an adenoma cohort (KIM ADENOMAS2020, N=204), two CRC cohorts (KIM ADENOMAS2020, N=138 and YACHIDA CRC 2019, N=347), and one T1D cohort (KOSTIC INFANTSDIABETES2015, N=103). All data were obtained from the study by Muller, E. et al.
[0334] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0335] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0336] Obviously, those skilled in the art can make various modifications and variations to the invention without departing from the spirit and scope of the invention. Therefore, if these modifications and variations fall within the scope of the claims of the invention and their equivalents, the invention is also intended to include these modifications and variations.
Claims
1. A method of screening for inflammatory bowel disease markers, characterized by, Comprise: A model for screening intestinal microbial markers is constructed using a machine learning algorithm, a recursive feature elimination method, and the abundance information of intestinal microbial markers; a model for screening functional genes of intestinal bacteria is constructed using a machine learning algorithm, a recursive feature elimination method, and the functional gene information of intestinal bacteria; and a model for screening characteristic metabolites is constructed using a machine learning algorithm, a recursive feature elimination method, and the content information of intestinal metabolites; Macro-genomic data queues and metabolome data queues are obtained respectively, wherein some sample individuals have corresponding macro-genomic data only in the macro-genomic data queues; some sample individuals have corresponding metabolome data only in the metabolome data queues; and some sample individuals have both corresponding macro-genomic data in the macro-genomic data queues and corresponding metabolome data in the metabolome data queues; The model for screening intestinal microbial markers is trained based on the macro-genomic data queues to obtain intestinal microbial markers; The model for screening functional genes of intestinal bacteria is trained based on the macro-genomic data queues to obtain functional genes of intestinal bacteria; The model for screening characteristic metabolites is trained based on the metabolome data queues to obtain characteristic metabolites; The intestinal microbial markers are 31, consisting of:
1. Bifidobacterium adolescentis; 2. Asaccharobacter celatus; 3. Gordonibacter pamelaeae; 4. Bacteroides plebeius; 5. Bacteroides xylanisolvens; 6. Barnesiella intestinihominis; 7. Clostridium sp CAG 58; 8. Intestinimonas butyriciproducens; 9. Lawsonibacter asaccharolyticus; 10. Eubacterium eligens; 11. Eubacterium hallii; 12. Eubacterium sp CAG 274; 13. Eubacterium ventriosum; 14. Ruminococcus torques; 15. Coprococcus comes; 16. Dorea longicatena; 17. Fusicatenibacter saccharivorans; 18. Lachnospira pectinoschiza; 19. Roseburia inulinivorans; 20. Oscillibacter sp 57 20; 21. Faecalibacterium prausnitzii; 22. Gemmiger_formicilis; 23. Clostridium_leptum; 24. Eubacterium_siraeum; 25. Ruminococcus_bromii; 26. Holdemania_filiformis; 27. Firmicutes_bacterium_CAG_110; 28. Firmicutes_bacterium_CAG_83; 29. Phascolarctobacterium_faecium; 30. Bilophila_wadsworthia; 31. Akkermansia_muciniphila; The functional genes of the intestinal bacteria are 25, consisting of: 1.K00180; 2.K00343; 3.K00394; 4.K01846; 5.K01873; 6.K01921; 7.K02071; 8.K02536; 9.K03284; 10.K03709; 11.K04564; 12.K05305; 13.K05341; 14.K06142; 15.K06334; 16.K06975; 17.K07099; 18.K08999; 19.K11534; 20.K12511; 21.K14445; 22.K18704; 23.K19068; 24.K19506; 25.K20627; The characteristic metabolites are 13, consisting of:
1. HMDB0000008; 2. HMDB0000062; 3. HMDB0000139; 4. HMDB0000148; 5. HMDB0000191; 6. HMDB0000192; 7. HMDB0000214; 8. HMDB0000517; 9. HMDB0000535; 10. HMDB0000619; 11. HMDB0000641; 12. HMDB0000929; 13. HMDB0002226.
2. The inflammatory bowel disease marker screening method according to claim 1, wherein Based on the macro gene data queue, the model for screening the intestinal microbial marker is trained to obtain an intestinal microbial marker, comprising: Step S31, the macro gene data queue is divided into: a first training set, a first validation set and a first test set; The data between the first training set, the first validation set and the first test set do not overlap with each other; Step S32, based on the first training set and the first validation set, cross-validation and leave-one-out method validation between queues are carried out, and the parameters of the model for screening the intestinal microbial marker are adjusted based on the validation result; Step S33, based on the first test set, the model for screening the intestinal microbial marker is tested, if the test result meets the requirements, the non-inflammatory intestinal disease queue is obtained, the false positive rate of the intestinal microbial marker is verified based on the current model for screening the intestinal microbial marker and its parameters, if the false positive rate of the intestinal microbial marker meets the requirements, the corresponding intestinal microbial marker is obtained based on the parameters of the current model for screening the intestinal microbial marker; If the false positive rate of the intestinal microbial marker does not meet the requirements, steps S32-S33 are circularly executed until the false positive rate of the intestinal microbial marker meets the requirements; If the test result does not meet the requirements, steps S32-S33 are circularly executed until the false positive rate of the intestinal microbial marker meets the requirements.
3. The inflammatory bowel disease marker screening method according to claim 1, wherein Based on the macro gene data queue, the model for screening the intestinal bacterial functional gene is trained to obtain the intestinal bacterial functional gene, comprising: Step S41, based on the first training set and the first validation set, cross-validation between queues and leave-one-out study method verification are performed, and the model of the functional gene of the screened intestinal bacteria is adjusted based on the verification result; Step S42, based on the first test set, the model of the functional gene of the screened intestinal bacteria is tested, if the test result meets the requirement, a non-inflammatory bowel disease queue is obtained, based on the current model of the functional gene of the screened intestinal bacteria and its parameters, the false positive rate of the functional gene of the intestinal bacteria in the non-inflammatory bowel disease queue is verified, if the false positive rate of the functional gene of the intestinal bacteria meets the requirement, the corresponding functional gene of the intestinal bacteria is obtained based on the parameters of the current model of the functional gene of the screened intestinal bacteria; if the false positive rate of the functional gene of the intestinal bacteria does not meet the requirement, steps S41-S42 are executed in a loop until the false positive rate of the functional gene of the intestinal bacteria meets the requirement; if the test result does not meet the requirement, steps S41-S42 are executed in a loop until the false positive rate of the functional gene of the intestinal bacteria meets the requirement.
4. The inflammatory bowel disease marker screening method according to claim 1, wherein Based on the metabolome data queue, the model of the screened feature metabolite is trained to obtain the feature metabolite, comprising: Step S51, the metabolome data queue is divided into: a second training set, a second validation set and a second test set; the data between the second training set, the second validation set and the second test set do not overlap with each other; Step S52, based on the second training set and the second validation set, cross-validation between queues and leave-one-out study method verification are performed, and the parameters of the model of the screened feature metabolite are adjusted based on the verification result; Step S53, based on the second test set, the model of the screened feature metabolite is tested, if the test result meets the requirement, a non-inflammatory bowel disease queue is obtained, based on the current model of the screened feature metabolite and its parameters, the false positive rate of the feature metabolite in the non-inflammatory bowel disease queue is verified, if the false positive rate of the feature metabolite meets the requirement, the corresponding intestinal microbial marker is obtained based on the parameters of the current model of the screened feature metabolite; if the false positive rate of the feature metabolite does not meet the requirement, steps S52-S53 are executed in a loop until the false positive rate of the feature metabolite meets the requirement; if the test result does not meet the requirement, steps S52-S53 are executed in a loop until the false positive rate of the feature metabolite meets the requirement.
5. A method for constructing an inflammatory bowel disease multi-omics diagnosis model, characterized by, Comprising: The inflammatory bowel disease multi-omics diagnosis model uses the macrogene data queue and the metabolome data queue as the data set, wherein the data set is divided into: a third training set, a third validation set and a third test set; the data between the third training set, the third validation set and the third test set do not overlap with each other; The inflammatory bowel disease multi-omics diagnosis model is constructed based on the abundance information of the intestinal microbial marker, the functional gene information of the intestinal bacteria and the content information of the feature metabolite; The inflammatory bowel disease multi-omics diagnosis model is trained and verified based on the third training set and the third validation set, cross-validation between queues and leave-one-out study method verification are performed. The inflammatory bowel disease multi-omics diagnosis model is independently verified based on the third test set; The intestinal microorganism markers are 31, consisting of:
1. Bifidobacterium adolescentis; 2. Asaccharobacter celatus; 3. Gordonibacter pamelaeae; 4. Bacteroides plebeius; 5. Bacteroides xylanisolvens; 6. Barnesiella intestinihominis; 7. Clostridium sp CAG 58; 8. Intestinimonas butyriciproducens; 9. Lawsonibacter asaccharolyticus; 10. Eubacterium eligens; 11. Eubacterium hallii; 12. Eubacterium sp CAG 274; 13. Eubacterium ventriosum; 14. Ruminococcus torques; 15. Coprococcus comes; 16. Dorea longicatena; 17. Fusicatenibacter saccharivorans; 18. Lachnospira pectinoschiza; 19. Roseburia inulinivorans; 20. Oscillibacter sp 57 20; 21. Faecalibacterium prausnitzii; 22. Gemmiger formicilis; 23. Clostridium leptum; 24. Eubacterium siraeum; 25. Ruminococcus bromii; 26. Holdemania filiformis; 27. Firmicutes bacterium CAG 110; 28. Firmicutes bacterium CAG 83; 29. Phascolarctobacterium faecium; 30. Bilophila wadsworthia; 31. Akkermansia muciniphila; The functional genes of the intestinal bacteria are 25, consisting of: 1.K00180; 2.K00343; 3.K00394; 4.K01846; 5.K01873; 6.K01921; 7.K02071; 8.K02536; 9.K03284; 10.K03709; 11.K04564; 12.K05305; 13.K05341; 14.K06142; 15.K06334; 16.K06975; 17.K07099; 18.K08999; 19.K11534; 20.K12511; 21.K14445; 22.K18704; 23.K19068; 24.K19506; 25.K20627; The characteristic metabolites are 13, consisting of:
1. HMDB0000008; 2. HMDB0000062; 3. HMDB0000139; 4. HMDB0000148; 5. HMDB0000191; 6. HMMD0000192; 7. HMMD0000214; 8. HMMD0000517; 9. HMMD0000535; 10. HMMD0000619; 11. HMMD0000641; 12. HMMD0000929; 13. HMMD0002226.
6. The use of inflammatory bowel disease markers in the preparation of a detection reagent for inflammatory bowel disease, characterized by: the use in the preparation of a detection reagent for inflammatory bowel disease, the inflammatory bowel disease markers in the detection reagent, comprising: intestinal microbial markers, intestinal bacterial functional genes and characteristic metabolites; The intestinal microbial markers are 31, consisting of:
1. Bifidobacterium adolescentis; 2. Asaccharobacter celatus; 3. Gordonibacter pamelaeae; 4. Bacteroides plebeius; 5. Bacteroides xylanisolvens; 6. Barnesiella intestinihominis; 7. Clostridium sp CAG 58; 8. Intestinimonas butyriciproducens; 9. Lawsonibacter asaccharolyticus; 10. Eubacterium eligens; 11. Eubacterium hallii; 12. Eubacterium sp CAG 274; 13. Eubacterium ventriosum; 14. Ruminococcus torques; 15. Coprococcus comes; 16. Dorea longicatena; 17. Fusicatenibacter saccharivorans; 18. Lachnospira pectinoschiza; 19. Roseburia inulinivorans; 20. Oscillibacter sp 57_20; 21. Faecalibacterium prausnitzii; 22. Gemmiger formicilis; 23. Clostridium leptum; 24. Eubacterium siraeum; 25. Ruminococcus bromii; 26. Holdemania filiformis; 27. Firmicutes bacterium CAG 110; 28. Firmicutes bacterium CAG 83; 29. Phascolarctobacterium faecium; 30. Bilophila wadsworthia; 31. Akkermansia muciniphila; The functional genes of the intestinal bacteria are 25, consisting of: 1.K00180; 2.K00343; 3.K00394; 4.K01846; 5.K01873; 6.K01921; 7.K02071; 8.K02536; 9.K03284; 10.K03709; 11.K04564; 12.K05305; 13.K05341; 14.K06142; 15.K06334; 16.K06975; 17.K07099; 18.K08999; 19.K11534; 20.K12511; 21.K14445; 22.K18704; 23.K19068; 24.K19506; 25.K20627; The characteristic metabolites are 13, consisting of:
1. HMDB0000008; 2. HMDB0000062; 3. HMDB0000139; 4. HMDB0000148; 5. HMDB0000191; 6. HMDB0000192; 7. HMDB0000214; 8. HMDB0000517; 9. HMDB0000535; 10. HMDB0000619; 11. HMDB0000641; 12. HMDB0000929; HMDB0002226.
7. An inflammatory bowel disease detection kit, characterized in that the inflammatory bowel disease markers in the detection kit include intestinal microbial markers, functional genes of intestinal bacteria and characteristic metabolites. The intestinal microbial markers are 31, consisting of:
1. Bifidobacterium adolescentis; 2. Asaccharobacter celatus; 3. Gordonibacter pamelaeae; 4. Bacteroides plebeius; 5. Bacteroides xylanisolvens; 6. Barnesiella intestinihominis; 7. Clostridium sp. CAG_58; 8. Intestinimonas butyriciproducens; 9. Lawsonibacter asaccharolyticus; 10. Eubacterium eligens; 11. Eubacterium hallii; 12. Eubacterium sp. CAG_274; 13. Eubacterium ventriosum; 14. Ruminococcus torques; 15. Coprococcus comes; 16. Dorea longicatena; 17. Fusicatenibacter saccharivorans; 18. Lachnospira pectinoschiza; 19. Roseburia inulinivorans; 20. Oscillibacter sp. 57_20; 21. Faecalibacterium prausnitzii; 22. Gemmiger formicilis; 23. Clostridium leptum; 24. Eubacterium siraeum; 25. Ruminococcus bromii; 26. Holdemania filiformis; 27. Firmicutes bacterium CAG 110; 28. Firmicutes bacterium CAG 83; 29. Phascolarctobacterium faecium; 30. Bilophila wadsworthia; 31. Akkermansia muciniphila; The functional genes of the gut bacteria are 25 consisting of: 1.K00180; 2.K00343; 3.K00394; 4.K01846; 5.K01873; 6.K01921; 7.K02071; 8.K02536; 9.K03284; 10.K03709; 11.K04564; 12.K05305; 13.K05341; 14.K06142; 15.K06334; 16.K06975; 17.K07099; 18.K08999; 19.K11534; 20.K12511; 21.K14445; 22.K18704; 23.K19068; 24.K19506; 25.K20627; The characteristic metabolites are 13 consisting of:
1. HMDB0000008; 2. HMDB0000062; 3. HMDB0000139; 4. HMDB0000148; 5. HMDB0000191; 6. HMDB0000192; 7. HMDB0000214; 8. HMDB0000517; 9. HMDB0000535; 10. HMDB0000619; 11. HMDB0000641; 12. HMDB0000929; 13. HMDB0002226.
Citation Information
Patent Citations
Detection method for biomarkers, for accurately intervening in hyperlipidemia, of radix astragali powder based on transcription science
CN110343758A
Intestinal bacteria and faeces metabolite capable of serving as type 1 diabetes biomarker and application of intestinal bacteria and faeces metabolite
CN114317671A
Bladder cancer metabolism marker screening method and system based on deep learning
CN114997303A