Computational methods for mapping peptides using metagenomic data
Metagenomic sequencing and mass spectrometry-based methods create refined microbial protein databases for precise gut proteome analysis, addressing the limitations of current genomic approaches and enabling detailed disease diagnosis and dietary impact assessment.
Patent Information
- Application Number
- PCT/IL2025/050564
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Current methods for characterizing the human intestinal proteome and microbiome interactions are fragmented and lack high-resolution assignment, particularly in understanding the impact of dietary changes on gut health and disease, while existing genomic approaches fail to distinguish between dead and live microbes and rely on subjective dietary assessments.
A method involving metagenomic sequencing to construct microbial protein databases, refined by species and coverage thresholds, combined with mass spectrometry and database searches, to identify and quantify microbial, host, and dietary protein sources in samples, enabling detailed profiling and disease diagnosis.
Provides a comprehensive and objective assessment of gut microbiome and host interactions, allowing for precise disease diagnosis and dietary compliance analysis, uncovering novel biomarkers and therapeutic targets for conditions like IBD.
Smart Images

Figure IL2025050564_08012026_PF_FP_ABST
Abstract
Description
[0001] COMPUTATIONAL METHODS FOR MAPPING PEPTIDES USING METAGENOMIC
[0002] DATA
[0003] RELATED APPLICATION / S
[0004] This application claims the benefit of priority of Israel Patent Application No. 314089 filed on 2 July 2024, the contents of which are incorporated herein by reference in their entirety.
[0005] FIELD AND BACKGROUND OF THE INVENTION
[0006] The present invention, in some embodiments thereof, relates to methods of constructing protein microbial databases for identification of proteins in a sample by proteomic analysis. The microbial databases have been specifically tailored according to metagenomics sequencing data of the sample and may optionally be combined with comprehensive and refined host and dietary databases.
[0007] The human gastrointestinal tract (GIT) and its commensal microbiome constitute a metabolically- and immune-active interface that impedes the colonization of pathogens, while concurrently executing nutrient digestion and absorption. Dietary signals constitute key modulators of host gut mucosal surface, microbiome configuration, and their intricate yet poorly understood interaction networks. Perturbation of the delicate tri-partite host-microbiome-diet equilibrium in individuals carrying genetic risk traits results in an array of medical human disorders ranging from inflammatory bowel disease (IBD), Celiac Disease, malnutrition, cardiometabolic disease, cancer and even neurodegeneration, through complex mechanisms that include systemic immune modulation and metabolite influx. However, simultaneous functional characterization of the host and microbiome under shifting dietary exposures remains fragmented, non-generalizable and indirect, thereby challenging the elucidation of their interaction networks and impacts on human disease. For example, human microbiome characterization mainly relies on genomic DNA characterization (using 16S rDNA analysis or shotgun metagenomic sequencing), which, at best, provides a snapshot of potential functionality while disregarding dead-live distinction or direct functional outputs. Inferring microbial function from genomic readouts assumes that all detected microbes express all of their proteins constitutively; However, microbial protein biosynthesis is a cost-effective regulated process, in which only selected sets of proteins required under certain conditions are translated at a given context.
[0008] Many biological effector functions of both the host and its microbiome are performed at protein level. Previous investigations pertaining to assessment of the human intestinal proteome have predominantly focused on either host or bacterial protein characterization in identifying bacterial proteins annotated at a family or genus level of resolution. The human host gut mucosal secretory protein responses to altering conditions, diets and microbial (bacteria, virus, archaea, fungi) signals, and their impacts on human intestinal disease and associated gut microbiome alterations (‘dysbiosis’) remain poorly understood. Consequently, despite the presence of a complex signature of hundreds of proteins and peptides differentially secreted into the gut lumen during health and disease states only one combination of two inter-connected anti-microbial peptides, S100A8 and S100A9, termed calprotectin, is utilized in clinical assessment of IBD severity and features a moderate sensitivity and specificity. Collectively, the intestinal protein landscape and its high-resolution assignment to host, microbial species, their cross-regulatory networks; and shifting landscapes upon exposure to environmental, genomic or immune perturbations remain elusive to date.
[0009] Likewise, dietary content, timing, and long-term consumption patterns are considered central modifiers of host gut mucosal and microbiome function across the GIT. However, quantification of dietary intake currently relies on subjective questionnaires or food diaries that are prone to recall bias, altered compliance, and insufficient resolution. Objective methodologies that would non-invasively assess the global landscape of food exposure and absorption at a single nutrient level are currently missing. Moreover, unbiased assessment of niche- specific impacts of dietary exposure on host and microbiome functions, and host and microbiome effects on food digestion and absorption would constitute invaluable tools for deciphering their cumulative effects on holobiont health and disease risk.
[0010] Additional background art includes US 20130338932A1.
[0011] SUMMARY OF THE INVENTION
[0012] According to an aspect of the present invention, there is provided a method of generating a microbial protein database comprising:
[0013] (a) sequencing at least one microbiome sample of at least one host subject species to obtain metagenomic data;
[0014] (b) aligning genomic sequences of the metagenomic data to reference genomes to obtain a database of microbial species that are potentially present in the at least one microbiome sample, and their relative abundance;
[0015] (c) refining the database by:
[0016] (i) selecting those microbes having a species abundance above a predetermined level; and (ii) selecting those microbes whose genomes have a % coverage of greater than a first predetermined % to obtain an initially refined database;
[0017] (d) obtaining protein sequences of the microbes of the initially refined database and align to the metagenomic data; and
[0018] (e) selecting those proteins having a % sequence alignment coverage greater than a second predetermined % to obtain a protein database which contains microbial proteins potentially produced by microbes in the at least one microbiome sample.
[0019] According to embodiments of the invention, the at least one microbiome sample is selected from the group consisting of a fecal sample, a gastrointestinal sample, a saliva sample, a blood sample, a urine sample, a tissue sample, a tumor sample and cerebrospinal fluid.
[0020] According to embodiments of the invention, the at least one host subject species is a mammalian subject species.
[0021] According to embodiments of the invention, the mammalian subject species has a disease or disorder.
[0022] According to embodiments of the invention, the at least one microbiome sample comprises at least two microbiome samples each derived from a different mammalian subject of the same species and having the same disease or disorder.
[0023] According to embodiments of the invention, the at least one mammalian subject species is healthy.
[0024] According to embodiments of the invention, the at least one microbiome sample comprises at least two microbiome samples each derived from a different mammalian subject of the same species and each being healthy.
[0025] According to embodiments of the invention, the at least one microbiome sample comprises at least two microbiome samples, a first sample of the at least two microbiome samples derived from a mammalian species having a disease or disorder and a second sample of the at least two microbiome samples derived from a mammalian species being healthy.
[0026] According to embodiments of the invention, the disease or disorder is a gastrointestinal disease.
[0027] According to embodiments of the invention, the disease is inflammatory bowel disease (IBD).
[0028] According to embodiments of the invention, the disease is an allergic disease.
[0029] According to embodiments of the invention, the allergic disease is Celiac disease.
[0030] According to embodiments of the invention, the method further comprises adding protein sequences of the host subject to the refined protein database. According to embodiments of the invention, the method further comprises adding protein sequences of proteins in a diet of the host subject species to the refined protein database.
[0031] According to another aspect of the invention, there is provided a method of profiling a microbiome sample of a host subject comprising:
[0032] (a) generating a protein database according to the method described herein;
[0033] (b) performing a mass spectrometry (MS) analysis on a protein extract of the microbiome sample; and
[0034] (c) performing a database search to determine a taxonomical origin of peptides generated by the MS analysis, using the database generated in step (a).
[0035] According to embodiments of the invention, the taxonomical origin is a bacterial, viral, fungal, parasite or archaeal origin.
[0036] According to embodiments of the invention, the microbiome sample is selected from the group consisting of a fecal sample, a gastrointestinal sample, a saliva sample, a blood sample, a urine sample, a tissue sample, a tumor sample and cerebrospinal fluid.
[0037] According to embodiments of the invention, the method further comprises adding protein sequences of the host subject species to the refined protein database following step (a) and prior to step (b).
[0038] According to embodiments of the invention, the method further comprises adding protein sequences of proteins in a diet of the host subject species to the refined protein database following step (a) and prior to step (b).
[0039] According to embodiments of the invention, the method further comprises determining an origin of the peptides, wherein the origin is selected from the group consisting of a microbial origin, a food origin and a host subject origin.
[0040] According to an aspect of the invention, there is provided a method of obtaining an indication for a subject suspected of having a disease comprising:
[0041] (a) performing a mass spectrometry (MS) analysis on a protein extract of a microbiome sample of the subject to identify sequences of peptides present in the protein extract;
[0042] (b) performing a database search to identify the protein source of the peptides and quantify an abundance of the protein source, the database being generated according to the method described herein; and
[0043] (c) comparing the abundance of the proteins identified in step (b) to an abundance in a protein extract of a microbiome sample of a healthy control subject, wherein a change in the abundance is indicative of the disease. According to embodiments of the invention, the method further comprises classifying a microbial, host or food origin of the peptides identified in step (b).
[0044] According to an aspect of the invention, there is provided a method of analyzing whether a subject is complying with a dietary intervention comprising:
[0045] (a) performing a mass spectrometry (MS) analysis on a protein extract of a gastrointestinal microbiome sample of the subject to identify sequences of peptides present in the protein extract; and
[0046] (b) performing a database search to identify the protein source of the peptides and quantify an abundance of the protein source, the database being generated according to the method described herein;
[0047] (c) classifying the protein source to a microbial, host or diet origin;
[0048] (d) analyzing the abundance of identified proteins that correspond to the proteins of the diet origin; wherein an amount or presence of the identified protein which is associated with the dietary intervention in the protein extract is indicative as to whether the subject is complying with the dietary intervention.
[0049] According to still another aspect of the invention, there is provided a method of analyzing gut barrier function of a subject comprising:
[0050] (a) performing a mass spectrometry (MS) analysis on a protein extract of a blood sample of the subject to identify sequences of peptides present in the protein extract; and
[0051] (b) performing a database search to identify the protein source of the peptides and quantify an abundance of the protein source, the database being generated according to the method described herein;
[0052] (c) classifying the proteins to a microbial, host or diet origin;
[0053] (d) analyzing the abundance of identified proteins that correspond to the proteins of the subject species; wherein an increased amount or presence of the identified protein compared to control values is indicative as to whether the subject has an impaired gut barrier function.
[0054] According to yet another aspect there is provided a method of analyzing small intestinal absorptive / digestive function comprising:
[0055] (a) performing a mass spectrometry (MS) analysis on a protein extract of a gastrointestinal sample of the subject to identify sequences of peptides present in the protein extract; and (b) performing a database search to identify the protein source of the peptides and quantify an abundance of the protein source, the database being generated according to the method described herein;
[0056] (c) classifying the proteins to a microbial, host or diet origin; and
[0057] (d) analyzing the abundance of identified proteins that correspond to the proteins of the subject species; wherein an increased amount or presence of the identified dietary protein compared to control values is indicative as to whether the subject has an impaired small intestinal absorptive / digestive function.
[0058] According to yet another aspect there is provided a computer readable medium comprising computer readable instructions stored thereon that when run on a computer perform a method as described herein.
[0059] According to yet another aspect there is provided a system, for producing a protein database for a microbiome sample comprising: one or more processors which executes the method described herein.
[0060] According to yet another aspect there is provided a method of diagnosing an inflammatory bowel disease (IBD) of a subject comprising measuring in a gut microbiome sample of the subject an amount of at least one bacteria selected from the group consisting of Alistipes putredinis, Dialister succinatiphilus, Oscillibacter sp. ER4, Romboutsia sp. DR1, Gemmiger formicilis, Faecalibacterium prausnitzii, Fusicatenibacter saccharivorans, Ruminococcus callidus and Bacteroides coprocola, wherein a decrease in the amount of the at least one bacteria as compared to an amount in a gut microbiome sample of a healthy subject is indicative of the subject having IBD.
[0061] According to yet another aspect there is provided a method of diagnosing Crohn’s disease (CD) of a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 1, wherein an amount of the proteins is indicative of the subject having CD.
[0062] According to yet another aspect there is provided a method of diagnosing ulcerative colitis (UC) of a subject comprising measuring in a gut microbiome sample of the subject for an amount of a combination of proteins set forth in a row of Table 2, wherein an amount of the proteins is indicative of the subject having UC.
[0063] According to yet another aspect there is provided a method of distinguishing between UC and CD in a subject comprising measuring in a gut microbiome sample of the subject for an amount of a combination of proteins set forth in a row of Table 3, wherein an amount of the proteins is indicative as to whether the subject has UC or CD.
[0064] According to yet another aspect there is provided a method of determining the severity of CD in a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 4, wherein an amount of the proteins is indicative as to the severity of the CD.
[0065] According to yet another aspect there is provided a method of determining the severity of UC in a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 5, wherein an amount of the proteins is indicative as to the severity of the UC.
[0066] According to embodiments of the invention, the method further comprises treating the subject according to the results of the analyzing.
[0067] According to embodiments of the invention, the measuring is effected on the protein level.
[0068] According to embodiments of the invention, the measuring is effected on the RNA level.
[0069] According to yet another aspect there is provided a method of treating IBD in a subject in need thereof comprising administering to the subject a therapeutically effective amount of at least one bacteria selected from the group consisting of Alistipes putredinis, Dialister succinatiphilus, Oscillibacter sp. ER4, Romboutsia sp. DR1, Gemmiger formicilis, Faecalibacterium prausnitzii, Fusicatenibacter saccharivorans, Ruminococcus callidus and Bacteroides coprocola, thereby treating the IBD.
[0070] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting.
[0071] BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWING(S)
[0072] Some embodiments of the invention are herein described, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is stressed that the particulars shown are by way of example and for purposes of illustrative discussion of embodiments of the invention. In this regard, the description taken with the drawings makes apparent to those skilled in the art how embodiments of the invention may be practiced.
[0073] In the drawings:
[0074] FIGs. 1A-N. In silico and in vitro validation of IPHOMED. (A-C) E. coli containing the plasmid pTriEx-Manf was incubated in vitro with either Isopropyl 0-d-l- thiogalactopyranoside (IPTG, red, n=6) to induce Lac operon controlled MANF expression or vehicle (PBS, white, n=6). After incubation, the culture was centrifuged, and bacterial protein expression was characterized in the bacterial pellets by the IPHOMED pipeline. (A) Experimental scheme. (B) The fold change in Manf gene copy number and protein abundance between a culture incubated with IPTG or vehicle. Bars represent the mean, and whiskers represent standard error. (C) Differential protein abundance analysis between a culture incubated with IPTG or vehicle. Each dot denotes a protein produced by E. coli. differentially abundant proteins are colored red. Statistical comparisons were performed using protein-wise linear models combined with empirical Bayes statistics, with FDR correction by Benjamini-Hochberg. (D-G) A single CFU of E. coli was inoculated into four different growth media (BHI, MHB, TSB, LB), and incubated in either aerobic or anaerobic conditions to generate eight growth conditions in total. Following incubation all cultures were centrifuged and bacterial pellets were analyzed using IPHOMED. Four biological repetitions were performed. (D) Experimental scheme. (E) Principal component analysis (PCA) of cultures bacterial proteome. Circles denote aerobic cultures, squares denote anaerobic cultures, and colors represent growth media. Each data point denotes a single culture. Significance by PERMANOVA. (F) The top 10 enriched BioCyc pathways in aerobic E.coli cultures compared with anaerobic E.coli cultures. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (G) Biocyc pathway enrichment analysis of differentially abundant E. coli proteins between TSB and LB broths in aerobic conditions. Only top 10 pathways are presented. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (H-J) IPHOMED analysis of in vitro bacterial cultures containing between one and six species in various combinations. 14 species that were not used in the culture were advertently included in the bacterial protein database to assess IPHOMED’ s false protein identification. (H) Experimental scheme. (I) Relative bacterial protein count as quantified by IPHOMED. Colors represent the bacterial strains (one of six combination strains or 14 strains not cultured in the experiment) and rectangle legends at the bottom denote which strains were truly present in each culture. This experiment was repeated twice with each small column within every bacterial combination representing a biological replicate. (J) Proteomic data was searched against a metagenomics data-driven protein database (IPHOMED, 65,249 proteins) or the IGC (Integrated Gene Catalog, 11,445,354 proteins, left). The middle color legend denotes which strain or strain combination comprised that culture. Bar plots represent the numbers of high-quality proteins (more than three unique peptides were identified) assigned to each bacterial mixture in each culture. (K-M) Comparative analysis between shotgun metagenomics (MG, purple) and metaproteomics (MP, red) was performed on stool samples from a cohort of naive SPF mice. (K) Gene- or proteinlevel alpha diversity, estimated by Shannon Index in all stool samples. Mann-Whitney test. (L) Bray-Curtis dissimilarity between samples by MG (taxon, gene, pathway levels) or MP (protein, pathway levels). Data is summarized using Tukey method. (M) Normalized, scaled, bacterial gene and protein abundances grouped by bacterial species. Each data value depicts the mean normalized gene / protein abundance across all naive SPF mice. Only top bacterial species with the highest number of gene / proteins are included. Full list of detected species. Unpaired multiple Mann-Whitney with Benjamini, Krieger, & Yekutieli FDR correction. (N) Percentage of bacterial proteins with an extracellular location for both predictable (protein normalized abundance is predicted by gene abundance, red) and unpredictable (protein normalized abundance not predicted by gene abundance, white) proteins. Extracellular / intracellular location defined by UniProt. Mann- Whitney test. Data is summarized using Tukey method. BHI, Brain Heart Infusion; DB, database; FC, fold change; FDR, false discovery rate; IPTG, Isopropyl P-d- 1 -thiogalactopyranoside; EB, Euria Broth; LC-MS / MS, liquid chromatography-mass spectrometry; MG, metagenomics; MP, metaproteomics; MHB, Mueller Hinton Broth; NB, nutrient broth; PBS, phosphate-buffered saline; PC, principal component; TSB, tryptic soy broth; act., activity; biosyn, biosynthesis; metab, metabolism.
[0075] FIGS. 2A-O. In-silico metagenomics and proteomics. (A-F) Ten different synthetic metagenomic samples, each comprised of metagenomics reads from 40 commonly detected gut microbiome species, were generated in silica. The refined bacterial metagenomic-based protein catalog was built using IPHOMED, comparing its performance with other methodologies (Methods). (A) Experimental design. (B) Circos plot representing the bacterial species detected with Kraken-Bracken (IPHOMED, red) or Metaphlan (pink). Species highlighted in bold represent the selected 40 bacterial species. In the center, species sensitivity of Kraken-Bracken (IPHOMED) or Metaphlan. (C) Distribution of the percentage of genome coverage by metagenomic reads after aligning all reads to the bacterial genomes of detected species using Kraken / Bracken (IPHOMED). (D) Species sensitivity and specificity of IPHOMED after selecting species with a genome coverage higher than 75%. (E) Number of bacterial proteins detected by IPHOMED or by genome assembly. Dot line represents the number of simulated proteins present in the selected 40 bacterial species. (F) Percentage of proteins assigned to the simulated 40 bacterial species, not included species or higher taxonomical resolution by IPHOMED or genome assembly. (G-K) Synthetic LC-MS / MS data of proteins from the combination of 1, 2 or 3 bacterial species (Bacteroides vulgatus, Faecalibacterium prausnitzii and Bacteroides uniformis) was generated and tested using IPHOMED methodology or a generic database (db) containing a total of 40 bacterial species (including the 3 species used for the simulation). (G) Experimental design. (H) Number of bacterial proteins detected with the generic database (db) and IPHOMED, compared to the number of proteins included in the synthetic LC- MS / MS data. (I) Sensitivity of the generic db and IPHOMED, for the detection of the proteins included in the synthetic LC-MS / MS data. (J) Specificity of the generic db and IPHOMED, for the detection of the proteins included in the synthetic LC-MS / MS data. (K) Number of proteins produced by species not included in the synthetic LC-MS / MS data and detected by the generic db and IPHOMED. (L-O) In silica simulation of IPHOMED analysis of naive SPF mice (n=l l) stool samples using protein databases artificially supplemented with random bacterial proteins that were not detected by metagenomic sequencing. LC-MS / MS output was searched against a metagenomically-defined protein database (herein MG, pooled data from a cohort of wild-type C57BL / 6 SPF mice, n=l l), or the same database supplemented with additional 1 to 5 million random bacterial proteins not detected by metagenomic sequencing. (L) The number of detected bacterial proteins at species (red) or higher taxonomic level (white), (M) the number of identified bacterial species, present (red) or absent (white) in metagenomics data and (N) the duration of computational processing (database search) runtime, as a function of the bacterial protein database size. (O) Taxonomical resolution of bacterial proteins comparing IPHOMED to database search against a general protein database (IGC, hereby termed ‘database’)23. Values represent the percentage of proteins assigned to each taxonomical level. Abund., abundance; db, database; Cov., Coverage; gen, genome; IPH, IPHOMED; MG, metagenomics; RDB, refined database; sp., species; k, kingdom; p, phylum; c, class; o, order; f, family; g, genus; s, species.
[0076] FIGs. 3A-M. The intestinal host proteomic landscape across different microbial colonization states in mice. (A-D) Two groups of GF mice were mono-colonized with either mouse segmented filamentous bacteria (mSFB, n=10) or rat SFB (rSFB, n=9) by oral gavage. A third group of mice that were gavaged with vehicle (PBS, n=9) were used as controls. After 14 days of colonization, terminal ileal samples were harvested and analyzed by IPHOMED. (A) Experimental scheme. (B) PCA of host proteins in mucosal ileal samples obtained from mice colonized with an attaching commensal or controls. Significance by PERMANOVA. (C) Differential host protein abundance analysis in mucosal Ileal samples obtained from mice colonized with mSFB or controls. Dashed line: q value = 0.05. Each datapoint denotes a host protein, differentially abundant proteins are colored blue. Statistical comparisons were performed using protein- wise linear models combined with empirical Bayes statistics, with FDR correction by Benjamini-Hochberg. (D) Reactome pathway enrichment analysis of significantly enriched host proteins in mSFB -colonized mice compared to PBS-treated controls. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (E-H) GF mice were treated with bi-daily intraperitoneal (IP) injections of recombinant IE-18 or vehicle (PBS). After five days of treatment, colonic samples were harvested and analyzed by IPHOMED. (E) Experimental scheme. (F) PCA of host proteomic analysis of colonic samples obtained from IL- 18 treated mice and controls. Significance by PERMANOVA. (G) Differential host protein abundance analysis in colonic samples obtained from GF mice treated with IL- 18 compared with controls. Dashed line: q value = 0.05. Each data point denotes a host protein. Differentially abundant proteins are shown in blue. Statistical comparisons were performed using protein-wise linear models combined with empirical Bayes statistics, with FDR correction by Benjamini- Hochberg. (H) GO pathway enrichment analysis (biological Process) of differentially abundant host proteins between IL-18-treated mice and controls. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (LL) SPF mice (n=12) were infected with Citrobacter rodentium by oral gavage. Distal colonic mucosal samples were collected before infection (dayO), during the peak of infection (day7) and at the recovery phase (day 14) and processed by IPHOMED and targeted (PRM assay) proteomics. (I) Experimental scheme. (J) PCA of bacterial proteomic profiles in colonic mucosal samples obtained from mice infected with C. rodentium at days 0, 7, and 14. Each data point denotes the bacterial proteome in one colon sample. Significance by PERMANOVA. (K) Differential bacterial protein abundance analysis in colonic samples obtained from mice infected with C. rodentium at day 7 compared with baseline (day 0). Dashed line: q value = 0.05. Each data point denotes a bacterial protein. Differentially abundant proteins are shown in empty circles (day 0) or filled red circles (day 7). Statistical comparisons were performed using protein-wise linear models combined with empirical Bayes statistics, with FDR correction by Benjamini-Hochberg. (L) Heatmap of differentially abundant C. rodentium proteins at day 0, day 7, and day 14 in the colonic mucosa. Values represent normalized abundances (z-score). (M) GO pathway enrichment analysis (biological process) of differentially abundant host proteins between C. rodent ium-inicctcd mice at day 7 and at baseline (day 0). Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. FC, fold change; GF, germ-free; mSFB, mouse segmented filamentous bacteria; ns, not significant; PBS, phosphate-buffered saline; PC, principal component; rSFB, rat segmented filamentous bacteria; SPF, specific pathogen-free. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0077] FIGs. 4A-I. Dietary influence on host and bacterial proteomes across the gastrointestinal tract. SPF mice were fed with either NC (n=5) or HFD (n=5) for one week. Luminal samples were collected from the stomach, duodenum, jejunum, ileum, cecum, and colon, and were subsequently analyzed by IPHOMED. (A) PCA of the host proteome quantified in all luminal samples. Each triangle data point denotes a sample from NC-fed mice, and each circle data point denotes a sample from HFD fed mice. Color denotes intestinal region. Significance by PERMANOVA (Table SI). (B) Number of differentially abundant host proteins detected in each region between mice fed with either NC or HFD. (C) Top 10 significantly enriched host Reactome pathways in the differentially abundant host proteins detected in the stomach between NC and HFD diets. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (D-E) Alluvial graph depicting the percentage distribution of bacterial proteins produced by each bacterial species throughout the gastrointestinal tract (GI) in mice fed with (D) NC or (E) HFD. (F) PCA of the bacterial proteome quantified in all luminal samples. Significance by PERMANOVA (Table SI). (G) Carbohydrate- Active Enzymes (CAZymes) enriched among bacterial proteins detected in the stomach that were significantly more abundant in mice fed with NC compared to HFD. Significance by hypergeometric test and Benjamini- Hochberg correction for multiple testing. (H) KEGG orthology (KO) terms enriched among bacterial proteins detected in the stomach that were significantly more abundant in mice fed with HFD compared to NC. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (I) Abundance of bacterial proteins associated with K00626, produced by Lachnospiraceae bacterium COE1 (upper panel) and Mucispirillum schaedleri (lower panel). Protein abundance was normalized using variance-stabilizing transformation (vst). Diff., differential; HFD, high-fat diet; NC, normal chow; PC, principal component; prot., protein; vst, variance- stabilizing transformation. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0078] FIGs. 5A-L. The intestinal host proteomic landscape in mouse during intestinal inflammation. SPF mice (n=l 1) were treated with 2% dextran sodium sulfate (DSS) added to the drinking water for seven days followed by the resumption of regular water. Mice were monitored for 31 days by weighing and longitudinal stool collection. (A) Experimental scheme. (B-C) PC As of shotgun metagenomics (B) and untargeted bacterial proteomics (C) in stool samples at days 0, 3, 5, 16 and 31. Each data point denotes a sample. Significance by PERMANOVA. (D-F) Differentially abundant bacterial proteins at day 0 compared with day 3 (D) day 16, (E) day 31 (F). Proteins were grouped by species-level taxa. A ‘+’ sign indicates taxa that exhibited differential abundance in a metagenomic analysis. (G) PCA of untargeted host proteomics of stool samples at days 0, 3, 5, 16 and 31. Each data point denotes a sample. Significance by PERMANOVA. (H) Number of bacterial (red) and host (blue) differentially abundant proteins at days 3, 5, 16, and 31 compared with baseline. (I) The bacterial (red) and host (blue) proteins with the highest predictive value in a random-forest classifier on proteomic analysis of stool samples at day 3 compared with baseline. Only the five most predictive proteins for host and bacteria are shown. Producing taxa is in parenthesis (J) Heatmap depicting protein abundance of differentially abundant inflammatory host proteins between DSS-treated mice throughout the experiment. Values represent the z-score of the mean normalized abundance per day. (K) GO pathway enrichment analysis (Biological process) of differentially abundant host proteins between DSS-treated mice at day 16 and day 0. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (L) Host (blue) and bacterial (red) proteins significantly correlated with body weight change (adjusted p-value < 0.05 & R2>0.35). Color gradient represents the coefficient of determination (R2). A ‘+’ sign indicates a positive correlation, and a sign indicates a negative correlation. Linear mixed effects models were used to correlate the protein abundance with the body weight change, including the individual as a random effect. DSS, dextran sulfate sodium; FC, fold change; PC, principal component; SPF, specific pathogen-free. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0079] FIGs. 6A-L. Host and bacterial activity in mouse during gut inflammation. (A-L) Microbiome transfer experiments from SPF mice donors to GF mice recipients where either recipients (A-H) or donors (I-L) were pre-treated or not with DSS prior to the transfer. (A-E) Treatment-naive (n=3) or DSS-treated (3 days of treatment, n=4) GF mice recipients were colonized with stool microbiome from treatment-naive SPF mice (Methods). Two days after colonization stool samples were collected from recipients and processed by IPHOMED. (A) Experimental scheme. (B) PCA of bacterial proteomic profiles of stool samples obtained from naive (I, pink) or DSS-treated (II, red) GF mice recipients colonized with naive SPF microbiome. Each data point denotes a stool sample. Significance by PERMANOVA. (C) Differential abundance analysis of bacterial proteins between naive (I, pink) or DSS-treated (II, red) GF mice colonized with naive microbiome. Labeled proteins are triosephosphate isomerases (TPI, KO 1803). Each data point denotes a bacterial protein. Dashed line denotes q=0.05. Differentially abundant proteins are colored red. Statistical comparisons were performed using protein-wise linear models combined with empirical Bayes statistics, with FDR correction by Benjamini-Hochberg. (D) Differentially abundant bacterial proteins between naive (I, pink) or DSS-treated (II, red) GF mice colonized with naive microbiome, grouped by species. (E) KEGG orthology (KO) terms enriched among differential bacterial proteins between naive (I) or DSS- treated (II) GF mice colonized with naive microbiome. Significance by hypergeometric test and Benjamini-Hochberg correction for multiple testing. (F-H) Treatment-naive (I, n=5) or DSS- treated (II, n=4) GF mice were colonized with stool microbiome from treatment-naive C57BL / 6 SPF mice (Methods). Four days after colonization, stool samples were collected from recipients and processed by IPHOMED. (F) Experimental scheme. (G) Differential abundance analysis of bacterial proteins between naive (I, pink) or DS S -treated (II, red) GF mice colonized with naive microbiome. Each data point denotes a bacterial protein. Dashed line denotes q=0.05. Differentially abundant proteins are colored red. Labeled proteins are top three bacterial proteins shared with DSS mouse model (Day 16, Figure 5E). Statistical comparisons were performed using protein-wise linear models combined with empirical Bayes statistics, with FDR correction by Benjamini-Hochberg. (H) Differentially abundant bacterial proteins between naive (I, pink) or DSS-treated (II, red) GF mice colonized with naive microbiome, grouped by species. (I-L) Naive GF mice were colonized with microbiome from treatment-naive (III, n=3) or DSS-treated (IV, n=4) SPF mice. Two days after colonization stool samples were collected and processed by IPHOMED. (I) Experimental scheme. (J) PCA of host proteomic profiles in stool samples obtained from naive GF mice colonized with microbiomes from naive (III, light blue) or DSS- treated (IV, blue) SPF mice. Each data point denotes a stool sample. Significance by PERMANOVA. (K) Differential abundance analysis of host proteins between naive GF mice colonized with microbiome from naive (III, light blue) or DSS-treated (IV, blue) SPF mice. Statistical comparisons were performed using protein-wise linear models combined with empirical Bayes statistics, with FDR correction by Benjamini-Hochberg. (L) Reactome pathways enriched among differential host proteins between naive GF mice colonized with microbiome from naive or DSS-treated SPF mice. Significance by hypergeometric test and Benjamini- Hochberg correction for multiple testing. DSS, dextran sulfate sodium; FC, fold change; GF, germ-free; PC, principal component; KO, KEGG Orthology; *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0080] FIGs. 7A-L. The intestinal dietary proteomic landscape of mice across different dietary regimens and diseases. (A) Proteomic analysis of food pellets of four different commercial rodent diets. Percentage of the dietary proteomics intensity associated with each dietary genus in each diet. (B) Differential protein abundance analysis between small intestine and colonic luminal samples from GF mice, colonized with Escherichia coli, that were fed with NC or HFD for five days (n=4 per condition). Each data point denotes a dietary protein. Dashed line denotes q=0.05. (C) SPF mice were fed with a casein-based diet (circles) or an amino acid-only diet (AA, rectangles) for two weeks (n=8 per condition). Each data point denotes the Alpha-Si -casein intensity detected in stool sample from each mouse by IPHOMED. Line represents the mean, and whiskers represent standard error. Significance by Mann-Whitney test. (D-G) Naive SPF mice were fed with either NC or HFD (n=5 per group) for two weeks. Luminal samples were collected from the stomach, duodenum, jejunum, ileum, cecum, and colon, and were subsequently processed using IPHOMED. (D) PCA of the dietary protein abundance profiles in all luminal samples. Colors represent the GI region of the luminal samples, and the shape represents the type of diet. Each data point denotes a luminal sample. Significance by PERMANOVA. (E) Percentage of detected peptides in the stomach associated with either a fulltrypsin or semi-trypsin profile in both NC- and HFD-fed mice. Significance by unpaired t test. (F) Distribution of dietary protein coverages (percentage of protein sequence covered by detected peptides) detected in luminal samples along the upper GI in SPF mice fed with either NC or HFD. Bars represent the mean, and whiskers represent standard error. Statistics by RM one-way ANOVA, multiple test correction by Holm-Sfdak. (G) Percentage of the dietary intensity originating from each detected dietary protein presented for each luminal sample from HFD-fed mice. (H-L) Colitis mouse model induced by administration of 2% DSS in drinking water for seven days and followed up for 31 days (n=l l mice as in Figure 5A). (H) PCA of the dietary protein abundance profiles during the colitis progression. Each dot represents one stool sample and colors represent the day of the experiment. Significance by PERMANOVA. (I) Mean percentage of the proteomics intensity assigned to each dietary genus in all samples across time. Data is summarized using the Tukey method. Statistics by mixed model (mouse as random effect), multiple test correction by Holm-Sfdak. (J) Total dietary proteomics intensity detected at each time point by the developed proteomics methodology. Each dot represents a stool sample. Bars represent the mean, and whiskers represent standard error. Statistics by mixed model (mouse as random effect), multiple test correction by Holm-Sfdak. (K) Correlation between total dietary proteomics intensity and S100A8 levels. Statistics by simple linear regression model. (L) Correlation between total dietary proteomics intensity and S100A9 levels. Statistics by simple linear regression model. AA, amino acid; DSS, dextran sodium sulfate; FC, fold change; FDR, false discovery rate; HFD, high-fat diet; HPD, high-protein diet; LPD, low-protein diet; NC, normal chow; PC, principal component; PERMANOVA, Permutational multivariate analysis of variance; SI, small intestine; SPF, specific pathogen free. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0081] FIGs. 8A-J. A schematic overview of proteomic host-microbiome-diet pipeline. (A) DNA sequencing. Each microbiome sample is sequenced by shotgun metagenomics. (B) Experimental database refinement. Bacterial protein database: iterative integration of speciesand gene-level metagenomic data from an entire pooled cohort is performed to define a geneprotein reference catalog, containing all bacterial proteins potentially produced by the microbes detected by sequencing in this set of samples (Methods). Dietary protein database: All proteins potentially present in dietary components are clustered together and protein markers of dietary genera are identified. Host protein database: a catalog of host proteins (human, mouse, etc.) is generated. The refined experimental database combines the bacterial, dietary, and host databases and constitutes a representation of all three major protein sources in the mammalian gastrointestinal tract - bacteria, host, and diet. (C) Mass-spectrometry. All samples undergo protein extraction, tryptic digestion, and liquid chromatography and tandem mass spectrometry (LC-MS / MS) to measure mass-to-charge ratio (m / z) of charged particles (ions). (D) Database search. Identification and quantification of bacterial, dietary and host proteins in all samples label-free quantitative (LFQ) proteomic analysis of the measured LC-MS / MS data, using a database search against the experimental refined database. (E) Statistical analysis. Differential abundance analysis and functional enrichment analysis of the quantified proteins, integrated at different taxonomical levels. Integrative analysis can be performed to identify correlations between bacterial, host and dietary proteins. IPHOMED methodology enables (F) the detection of novel host, bacterial and dietary biomarkers for diagnosis or disease severity, (G) provides an objective tool for diet quantification and dietary compliance assessment, and (H) uncovers novel potential mechanisms involved in disease onset and progression, thus enabling the discovery of potential therapeutic targets. (I-J) Protein characterization in stool samples from healthy, treatment-naive WT SPF mice (n=17). Total number (I) and percentage (J) of detected host (blue), bacterial (red), and dietary (yellow) proteins. Blue, red, and yellow colors represent host, bacterial, and dietary proteins throughout the manuscript. Cov., coverage; gen., genus; H. sapiens, Homo sapiens', LC-MS / MS, liquid chromatography tandem mass spectrometry; M. musculiis, Mas musculiis; path, pathway; sp., species.
[0082] FIGs. 9A-F. The bacterial and host proteomic landscape across the healthy human gastrointestinal tract. (A-D) Stool and luminal samples from the stomach, duodenum, jejunum, terminal ileum, cecum, and descending colon were obtained by esophagogastroduodenoscopy and colonoscopy in treatment-naive (n=l 1) and broad-spectrum ABX-treated (n=10) healthy adult individuals (Methods). Subjects from both groups were age- and gender-matched. (A) Mean relative bacterial protein abundance of microaerophilic (dark red), anaerobic (red), aerobic (pink), and facultative anaerobic (light pink) bacteria across the GIT tract in treatment-naive and ABX-treated individuals. Colors indicate species, shared species between adjacent sites of the GIT are shown as connections between bars. (B) Mean total number of different bacterial proteins detected in each GIT region and observed in treatment-naive and ABX-treated individuals. Bars represent the mean, and whiskers represent standard error. Kruskal-Wallis with FDR correction by two-stage linear step-up procedure of Benjamini, Krieger and Yekutieli. (C) Total number of proteins associated with resistance to different types of antibiotics (antibiotic resistant proteins, ARPs) in each GIT region and observed in treatment-naive and ABX-treated individuals. Colors indicate GIT region. (D) Total number of different host proteins detected in each GIT region in treatment-naive and ABX-treated individuals. Bars represent the mean, and whiskers represent standard error. Kruskal-Wallis with FDR correction by two-stage linear step- up procedure of Benjamini, Krieger and Yekutieli. (E-F) Principal component analysis (PCA) of the human proteome across the GIT in (E) Treatment-naive (p=0.001) and (F) ABX-treated (p=0.001) individuals. Each data point represents a sample. ABX, antibiotics; ARPs, antibiotic resistance proteins; ns, non-significant; PC, Principal Component; Stom, Stomach; Duo, Duodenum; Jej, Jejunum; He, Ileum; Cec, Cecum; Col, Colon; Stoo, Stool; FDR, False Discovery Rate; GIT, Gastrointestinal Tract. **, Q<0.01.
[0083] FIGs. 10A-I. The host and bacterial intestinal proteomic landscape in pediatric Crohn’s disease. (A-I) We profiled the metagenomics and metaproteomics landscape in stool samples from a cohort of Israeli subjects with pediatric Crohn’s disease (ISR-pCD, n=26) and age- and gender-matched healthy controls (ISR-HC, n=28). (A) Experimental scheme. (B) PCA of host proteome in stool samples from subjects with ISR-pCD (blue) and controls (white). Each data point represents a stool sample. (C) Differential host protein abundance analysis between subjects with ISR-pCD and ISR-HC. Each data point represents a host protein. Dashed line denotes q=0.05. (D) Gene Ontology (Molecular function terms) enrichment analysis of differentially abundant host proteins between ISR-pCD and ISR-HC. The top 10 most enriched pathways are shown. The dotted vertical line denotes the threshold for statistical significance (q-value = 0.05). (E) PCA of the bacterial proteome in stool samples from subjects with ISR-pCD (red) and controls (white). Each data point represents a stool sample. (F) Differential bacterial protein abundance analysis between ISR-pCD patients and healthy individuals. Each data point represents a bacterial protein. Dashed line denotes q=0.05. (G) For each bacterial species, number of proteins significantly enriched in healthy individuals or ISR- pCD patients. All species that were found to be significantly enriched or depleted by a hypergeometric test and FDR correction by Benjamini-Hochberg are shown. A ‘+’ sign indicates that the species was also significantly enriched by a metagenomic differential abundance analysis. White bars depict bacterial proteins that are enriched in healthy controls, and red bars depicts bacterial proteins that are enriched in ISR-pCD. (H) The top 10 enriched bacterial BioCyc pathways in differentially abundant proteins between ISR-pCD patients and healthy controls. The dotted vertical line denotes the threshold for statistical significance (q-value = 0.05). (I) At the top, heatmap depicting differentially abundant host proteins between ISR-pCD patients and healthy controls. Values represent the normalized abundance (z-score) across all individuals. At the bottom, a correlation analysis between host and bacterial proteins. For each host protein, only the bacterial protein with the highest correlation (R2) is depicted. pCD, pediatric Crohn’s disease; HC, Healthy Controls; FC, Fold Change; PC, Principal Component; FDR, False Discovery Rate. *, Q<0.05, **, Q<0.01, ****, Q<0.0001.
[0084] FIGs. 11A-J. The host and bacterial intestinal proteomic landscape in Ulcerative colitis. (A-J) We profiled the metagenomics and metaproteomics landscape in stool samples from a cohort of German subject with Ulcerative colitis (GER-UC, n=50) and age- and gender-matched healthy controls (GER-HC, n=50). (A) Experimental design. (B) PCA of host proteome in stool samples from subjects with GER-UC (blue) and controls (white). Each data point represents a stool sample. (C) Differential host protein abundance analysis between subjects with GER-UC and GER-HC. Each data point represents a host protein. Dashed line denotes q=0.05. (D) Gene Ontology (Molecular function terms) enrichment analysis of differentially abundant host proteins between GER-UC and GER-HC. The top 10 enriched pathways are shown. (E) PCA of the bacterial proteome in stool samples from subjects with GER-UC (red) and controls (white). Each data point represents a stool sample. (F) Differential bacterial protein abundance analysis between GER-UC patients and healthy individuals. Each data point represents a bacterial protein. Dashed line denotes q=0.05. (G) For each bacterial species, number of proteins significantly differentially abundant in either healthy individuals or GER-UC patients. All species that were found to be significantly enriched or depleted by a hypergeometric test and FDR correction by Benjamini-Hochberg are shown. A ‘+’ sign indicates that the species was also significantly enriched by a metagenomic differential abundance analysis. White bars depict bacterial proteins that are enriched in healthy controls, and red bars depicts bacterial proteins that are enriched in GER-UC. (H) Total number of bacterial proteins significantly enriched in healthy individuals or IBD patients in both IBD cohorts. Statistics were calculated with hypergeometric test and FDR correction by Benjamini-Hochberg. White bars depict bacterial proteins that are enriched in healthy controls. Red bars depict bacterial proteins that are enriched in subjects with IBD. (I) Accuracy of the diagnostic classification (AUROC) of the ISR-pCD and GER-UC cohorts by both metaproteomics (protein abundances) and metagenomics (gene abundances). Random-forest classification evaluated in repeated 5-fold cross-validation iterations (10 repetitions). Data represent the median of each iteration. Ordinary one-way ANOVA and Sidak correction. (J) At the top, heatmap depicting differentially abundant host proteins between GER- UC patients and healthy controls. Values represent the normalized abundance (z-score) across all individuals. At the bottom, a correlation analysis between host and bacterial proteins. For each host protein, only the bacterial protein with the highest correlation (R2) is depicted. UC, Ulcerative colitis; CD, Crohn’s disease; IBD, Inflammatory Bowel Disease; ISR, Israel; GER, Germany; FC, Fold Change; HC, Healthy Controls; AUROC, Area Under the Receiver Operator Curve; PC, Principal Component; FDR, False Discovery Rate. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0085] FIGs. 12A-K. Assessing host-microbiome interaction at protein-level by fecal microbiome transfer (FMT) from donor subjects with UC and matched healthy controls into germ-free (GF) mice recipients. (A) Experimental scheme. Stool microbiomes from subjects with GER-UC (n=3) and GER-HC (n=3) were transferred to GF mice (Methods). The microbiome of each human donor was used to inoculate 6-8 GF mice recipients. The metagenomic and proteomic landscapes were analyzed in stool obtained from mice recipients after two weeks of colonization. (B-C) PC A of recipients’ stool samples two weeks following colonization, based on (B) bacterial metagenome, (C) bacterial proteome. Each data point represents a recipient’s stool sample. Shapes denote individual donor. White datapoints denote samples from recipients of a healthy microbiome (I) and colored datapoints denote samples from recipients of a GER-UC microbiome (II). (D) Total number of bacterial proteins significantly enriched in mice colonized with a GER-HC (I) or GER-UC (II) microbiome and not significantly different at gene level. (E) Bacterial proteins produced by Bacteroides vulgatus significantly more abundant in recipients of a human GER-UC microbiome. Values represent the normalized protein abundances (z-score) across all donors (average of all samples per donor). (F) Differentially abundant bacterial proteins in recipients of a GER-UC microbiome and shared with human GER-UC microbiome (Figure 11G). (G) PCA of the host proteome from recipients’ stool samples two weeks following colonization. (H) Differential protein abundance analysis in host protein between mice colonized with GER-UC-associated microbiome (II) and mice colonized with GER-HC-associated microbiome (I). Statistical comparisons were performed using linear mixed-effect models, including sample donors as random effects. Dotted line depicts the threshold for statistical significance. (I) The area under the receiver operator curve (AUROC) in a random-forest diagnostic classifier, used to classify donors’ IBD-status, based on either bacterial genes (gene-level metagenomic data), bacterial proteins (protein-level proteomic data), or host proteins (protein-level proteomic data). Random-forest classification was evaluated in five-fold cross-validation iterations: samples from the same donor were only included in the training or the test groups. Statistics by Kruskal-Wallis test with Dunn’s correction for multiple comparisons. (J) Differential bacterial protein abundance analysis between mice colonized with UC-associated microbiome and mice colonized with GER-HC-associated microbiome. Statistical comparisons were performed using linear mixed-effect models, including sample donors as random effects. Each data point represents a bacterial protein. Dashed line denotes q=0.05. (K) Biocyc pathway enrichment analysis of differentially abundant bacterial proteins between both groups. Dashed line denotes q=0.05. UC, Ulcerative Colitis; HC, Healthy Control; GF, Germ- Free; FC, Fold Change; AUC, Area Under the Curve; ROC, Receiver Operator Curve; PC, Principal Component; FDR, False Discovery Rate. *, Q<0.05.
[0086] FIGs. 13A-H. Identifying non-invasive biomarkers for IBD diagnosis and activity. (A-B) IPHOMED analysis of stool samples from two cohorts of the Human Microbiome Project with Ulcerative colitis (UC) and Crohn’s disease (CD) patients, Massachusetts General Hospital (HMP-MGH, UC=12, CD=34, HC=52) and the Cincinnati Hospital (HMP-Cin, UC=29, CD=48, HC=45). (A) Venn diagrams depicting the number of detected host (blue) and bacterial (red) proteins in the HMP-MGH and HMP-Cin cohorts. Overlap denotes proteins that were detected in both cohorts. (B) The area under the receiver operator curve (AUROC) in a random-forest classifier for UC or CD diagnosis, combining both host and bacterial proteins in the HMP-MGH and HMP-Cin cohorts. Random-forest classification was evaluated in repeated five-fold cross- validation iterations (10 repetitions). (C) Top 50 discriminating proteins for CD diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for CD diagnosis. Combinations of CD discriminating proteins were tested in the HMP-MGH, HMP-Cin and Israeli pediatric CD cohorts (ISR-pCD=26, HC=28). Heatmap values of the AUROC of a random-forest classifier for CD diagnosis in the two HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Eollipop plot depicts the performance in the ISR-pCD cohort. (D) Top 50 discriminating proteins for UC diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for UC diagnosis. Combinations of UC discriminating proteins were tested in the HMP-MGH, HMP-Cin and GER-UC cohorts (UC=50, HC=50). Heatmap of AUROC values of a random-forest classifier for UC diagnosis in the HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Lollipop plot depicts the performance in the GER-UC cohort. (E) Random-forest classifier distinguishing between UC and CD, using both host and bacterial proteins in the HMP-MGH and HMP-Cin cohorts. Top 50 discriminating proteins were randomly combined in sets of 1 or 3 proteins to test the accuracy for distinguishing between CD and UC. Combinations of discriminating proteins were tested in the HMP-MGH, HMP-Cin and HMP-Ced cohort (CD=60, UC=54). Heatmap values the AUROC of a randomforest classifier for IBD subtype in the HMP-MGH and HMP-Cin cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot depicts the performance in the HMP-Ced. (F) Top 50 discriminating proteins for CD diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for CD severity prediction in Mills-CD cohort (Mills et al, CD=119). Support Vector Machine (SVM) for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R-squared values of the regression model. (G) Top 50 discriminating proteins for UC diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for UC severity prediction in the Mills-UC cohort (Mills et al, UC=63). Support Vector Machine (SVM) for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R- squared values of the regression model. (H) Bacterial species producing proteins associated with KOO133 (aspartate- semialdehyde dehydrogenase) across all cohorts and IBD subtypes. MGH, Massachusetts General Hospital; Cin, Cincinnati; UC, Ulcerative Colitis; CD, Crohn’s disease; AUC, Area Under the Curve; ROC receiver operator curve; KO, KEGG Orthology; GER, Germany; ISR, Israel. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0087] FIGs. 14A-G. A dietary intervention of food item enrichment in humans, to assess proteomics and metagenomics sensitivity in detecting dietary signals. Six healthy subjects were enrolled in an interventional dietary trial of controlled consumption of four plant-based foods: bananas (yellow), apples (red), cucumbers (green), and peanuts (brown). Throughout the study, banana and apple consumption was only allowed during a four-days enrichment period (Days 0- 3) and was instructed to account for a daily minimum of 300 grams of bananas and 400 grams of apples. Cucumber and peanut consumptions were only allowed during another separate four- days enrichment period (days 11-14) and was instructed to account for a daily minimum of 350 grams of cucumbers and 100 grams of peanuts. Stool samples were obtained at prespecified time points and analyzed by proteomics and metagenomics (Methods). (A) Experimental scheme. (B- C) The mean signal intensity in (B) metaproteomics (MP) and (C) metagenomics (MG) of all four food genera in stool samples throughout the study. (D-E) Area under the curve (AUC) of the mean signal intensity as a function of time curves for each dietary genus as measured by (D) MP and (E) MG. (F) Most abundant and recurrently consumed dietary genera. Bars represent the number of participants that consumed the dietary genus during the intervention. (G) 15 healthy human participants were enrolled in a published clinical study and reported all consumed food items . Metagenomics and proteomics data from stool samples were downloaded and analyzed IPHOMED. Heatmap showing the sensitivity of our methodology for the detection of the reported food items. Yellow color represents proteomic signal detected with IPHOMED. Thick line in the square represents reported food items and thin line in the square represents not reported food items. Vertical bar plots represent the sensitivity of the methodology for the detection of each reported dietary genus across all participants. Horizontal bar plots represent the sensitivity of the methodology for the detection of all dietary foods reported by a participant. AUC, area under the curve; MP, metaproteomics; MG, metagenomics; ns, not significant. **, Q<0.01, ***, Q<0.001.
[0088] FIGs. 15A-L. The gastrointestinal dietary proteomic profiles of healthy humans and subjects with IBD. (A) Heatmap showing the percentage of proteomics signal assigned to every dietary genus for each human cohort. In total, 6 different cohorts were included: 4 datasets from the Human Microbiome Project (collected by Cedars-Sinai Medical Center, Cincinnati Children’s Hospital and Massachusetts General Hospital) and the German-UC (GER-UC) and Israeli-CD (ISR-pCD) cohorts included in this study. (B-D) Comparison of dietary habits of German and Israeli populations. (B) Percentage of proteomics signal assigned to Sus (pork) genus. Bars represent the mean, and whiskers represent standard error. Mann- Whitney test (C) Percentage of proteomics signal assigned to individual Sus (pork) proteins in the German (white) and Israeli (yellow) individuals. Bars represent the mean, and whiskers represent standard error. Multiple Mann- Whitney test with FDR correction by Benjamini and Hochberg (D) Percentage of proteomics signal assigned to Gallus (chicken) genus. Bars represent the mean, and whiskers represent standard error. Mann-Whitney test (E-G) Pediatric subjects with CD (pCD) were enrolled in a clinical trial involving commencement of exclusive enteral nutrition (EEN) to treat an inflammatory flare upon diagnosis. Stool samples from seven pCD patients were collected before (yellow, regular diet) and after (white) commencing on EEN and processed by IPHOMED. (E) The number of detected dietary proteins in stools from all pCD patients before (regular diet) and after commencing EEN. Paired T-test. (F) Total dietary signal quantified in all stool samples before (regular diet) and after commencing EEN. Paired T-test. (G) Correlation between number of dietary proteins detected and the calprotectin (S100A9 subunit) levels in stool samples. Pearson correlation coefficient. (H) Stool samples from a human cohort of celiac patients before (n=9, yellow) and during gluten-free diet (n=5, white) were collected and processed for proteomics. Values represent gluten protein abundance in all individuals. Unpaired t-test. (I-L) Dietary protein quantification in stool samples from IBD patients and healthy individuals in two independent cohorts from the Human Microbiome Project: Cincinnati Children’s Hospital (HMP-Cin) and Massachusetts General Hospital (HMP-MGH). (I) The pattern of intestinal involvement (location of bowel inflammation) of Crohn’s disease in the HMP-Cin and HMP-MGH cohorts by the Montreal classification system for GIT location (LI, ileal; L2, colonic; L3, ileocolonic; +L4, concomitant upper GI disease). (J-K) Total dietary protein abundance quantified in IBD patients (yellow) and healthy individuals (white) from the (J) HMP-MGH cohort and (K) HMP-Cin cohort. Line represents the median. Mann-Whitney test. (L) Total dietary protein abundance quantified in IBD patients stratified by the affected GIT location in the Cincinnati cohort. Line represents the median. Mann-Whitney test. GER, Germany; ISR, Israel; MGH, Massachusetts General Hospital; ped, pediatric; GI, gastrointestinal; IBD, Inflammatory Bowel Disease; UC, Ulcerative colitis; pCD, pediatric Crohn’s disease; HC, Healthy Control; EEN, Exclusive Enteral Nutrition; ns, not significant. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0089] FIGs. 16A-C. Accuracy of host and bacterial proteome for IBD diagnosis. IPHOMED analysis of stool samples from two cohorts of the Human Microbiome Project (HMP) with Ulcerative colitis (UC) and Crohn’s disease (CD) patients, Massachusetts General Hospital (HMP-MGH, UC=12, CD=34, HC=52) and the Cincinnati Hospital (HMP-Cin, UC=29, CD=48, HC=45). (A-B) Top 50 discriminating features and their importance (average of both cohorts) for (A) CD and (B) UC diagnosis. (C) The area under the receiver operator curve (AUROC) in a random-forest classifier for UC or CD diagnosis, including host (blue) or bacterial (red) proteins in the HMP-MGH and HMP-Cin cohorts. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). MGH, Massachusetts General Hospital; Cin, Cincinnati; UC, Ulcerative Colitis; CD, Crohn’s disease; AUC, area under the curve; ROC receiver operator curve.
[0090] FIGs. 17A-E. Identifying non-invasive biomarkers for IBD diagnosis and activity. (A-B) IPHOMED analysis of stool samples from two cohorts of the Human Microbiome Project (HMP) with Ulcerative colitis (UC) and Crohn’s disease (CD) patients, Massachusetts General Hospital (HMP-MGH, UC=12, CD=34, HC=52) and the Cincinnati Hospital (HMP-Cin, UC=29, CD=48, HC=45). Random-forest classifier for UC or CD diagnosis, was performed combining host proteins and bacterial proteins collapsed into KO terms in the HMP-MGH and HMP-Cin cohorts. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). (A) Top 50 discriminating proteins for CD diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for CD diagnosis. Combinations of CD discriminating proteins were tested in the HMP-MGH, HMP-Cin and the Israeli pediatric CD cohort (ISR-pCD=26, HC=28). Heatmap values the AUROC of a random-forest classifier for CD diagnosis in the HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Lollipop plot depicts the performance in the ISR-pCD cohort. (B) Top 50 discriminating proteins for UC diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for UC diagnosis. Combinations of UC discriminating proteins were tested in the HMP-MGH, HMP-Cin and German UC cohorts (GER-UC, UC=50, HC=50). Heatmap values the AUROC of a random-forest classifier for UC diagnosis in the HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Lollipop plot depicts the performance in the GER-UC cohort. (C) Random-forest classifier distinguishing between UC and CD, combining host proteins and bacterial proteins collapsed into KO terms in the HMP-MGH and HMP-Cin cohorts. Top 50 discriminating proteins were randomly combined in sets of 1 or 2 proteins to test the accuracy for distinguishing between CD and UC. Combinations of discriminating proteins were tested in the HMP-MGH, HMP-Cin and Cedars (HMP-Ced, CD=60, UC=54) cohorts. Heatmap values the AUROC of a random- forest classifier for IBD subtype in the HMP-MGH and HMP-Cin cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot depicts the performance in the HMP-Ced cohort. (D) Top 50 discriminating proteins for CD diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for CD severity prediction in the Mills- CD cohort15(CD=119). Support Vector Machine (SVM) for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R-squared values of the regression model. (E) Top 50 discriminating proteins for UC diagnosis were randomly combined in sets of 1 or 2 proteins to test the accuracy for UC severity prediction in the Mills-UC cohort (UC=63). Support Vector Machine (SVM) for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R-squared values of the regression model. MGH, Massachusetts General Hospital; ISR, Israel; GER, Germany; AUC, area under the curve; ROC receiver operator curve; UC, Ulcerative colitis; CD, Crohn’s disease; HC, Healthy Control. *, Q<0.05, **, Q<0.01, ***, Q<0.001, ****, Q<0.0001.
[0091] FIGs. 18A-E. Identifying non-invasive biomarkers for IBD diagnosis and activity. IPHOMED analysis of stool samples from two cohorts of the Human Microbiome Project (HMP) with Ulcerative colitis (UC) and Crohn’s disease (CD) patients, Massachusetts General Hospital (HMP-MGH, UC=12, CD=34, HC=52) and the Cincinnati Hospital (HMP-Cin, UC=29, CD=48, HC=45). Random-forest classifier for UC or CD diagnosis, was performed combining host proteins and bacterial proteins in the HMP-MGH and HMP-Cin cohorts. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). (A) Top 50 discriminating proteins for UC diagnosis were randomly combined in sets of 3 proteins to test the accuracy for UC diagnosis. Combinations of UC discriminating proteins were tested in the HMP-MGH, HMP-Cin and GER-UC cohorts (UC=50, HC=50). Heatmap values the AUROC of a random-forest classifier for UC diagnosis in the HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross- validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Lollipop plot depicts the performance in the GER-UC cohort. (B) Top 50 discriminating proteins for UC diagnosis were randomly combined in sets of 3 proteins to test the accuracy for UC severity prediction in the Mills-UC cohort15(UC=63). Support Vector Machine (SVM) for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R-squared values of the regression model. (C) Top 50 discriminating proteins for CD diagnosis were randomly combined in sets of 3 proteins to test the accuracy for CD diagnosis. Combinations of CD discriminating proteins were tested in the HMP-MGH, HMP-Cin and Israeli pediatric CD cohorts (ISR-pCD=26, HC=28). Heatmap values the AUROC of a random-forest classifier for CD diagnosis in the HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross- validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Lollipop plot depicts the performance in the ISR-pCD cohorts. (D) Top 50 discriminating proteins for CD diagnosis were randomly combined in sets of 3 proteins to test the accuracy for CD severity prediction in the Mills-CD cohort15(CD=119). SVM for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R-squared values of the regression model. (E) Random-forest classifier distinguishing between UC and CD, using both host and bacterial proteins in the HMP-MGH and HMP-Cin cohorts. Top 50 discriminating proteins were randomly combined in sets of 3 proteins to test the accuracy for distinguishing between CD and UC. Combinations of discriminating proteins were tested in the HMP-MGH, HMP-Cin and Cedars cohort (HMP-Ced, CD=60, UC=54). Heatmap values the AUROC of a random-forest classifier for IBD subtype in the HMP-MGH and HMP- Cin cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot depicts the performance in the HMP-Ced cohort. AUC, area under the curve; ROC receiver operator curve; UC, Ulcerative colitis; CD, Crohn’s disease, HC, Healthy Control.
[0092] FIGs. 19A-F. Identifying non-invasive biomarkers for IBD diagnosis and activity. Related to Figure 5 and Figure 7. IPHOMED analysis of stool samples from two cohorts of the Human Microbiome Project (HMP) with Ulcerative colitis (UC) and Crohn’s disease (CD) patients, Massachusetts General Hospital (HMP-MGH, UC=12, CD=34, HC=52) and the Cincinnati Hospital (HMP-Cin, UC=29, CD=48, HC=45). Random-forest classifier for UC or CD diagnosis, was performed combining host proteins and bacterial proteins collapsed into KO terms in the HMP-MGH and HMP-Cin cohorts. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). (A) Top 50 discriminating proteins for UC diagnosis (Figure 18B) were randomly combined in sets of 3 proteins to test the accuracy for UC diagnosis. Combinations of UC discriminating proteins were tested in the HMP-MGH, HMP-Cin and German UC cohorts (GER-UC, UC=50, HC=50). Heatmap values the AUROC of a random-forest classifier for UC diagnosis in the HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Lollipop plot depicts the performance in these GER-UC cohort. (B) Top 50 discriminating proteins for UC diagnosis were randomly combined in sets of 3 proteins to test the accuracy for UC severity prediction in the Mills-UC cohort15(UC=63). Support Vector Machine (SVM) for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R-squared values of the regression model. (C) Top 50 discriminating proteins for CD diagnosis (Figure 18 A) were randomly combined in sets of 3 proteins to test the accuracy for CD diagnosis. Combinations of CD discriminating proteins were tested in the HMP-MGH, HMP-Cin and Israeli pediatric CD cohorts (ISR-pCD, pCD=26, HC=28). Heatmap values the AUROC of a random-forest classifier for CD diagnosis in the HMP cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross- validation iterations (10 repetitions). Calprotectin performance is included at the bottom. Lollipop plot depicts the performance in the ISR-pCD cohorts. (D) Top 50 discriminating proteins for CD diagnosis were randomly combined in sets of 3 proteins to test the accuracy for CD severity prediction in in Mills-CD cohort (CD=119). SVM for regression was evaluated by repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot represents the median of R-squared values of the regression model. (E) Random-forest classifier distinguishing between UC and CD, using both host and bacterial proteins (collapsed into KO terms) in the HMP-MGH and HMP-Cin cohorts. Top 50 discriminating proteins were randomly combined in sets of 3 proteins to test the accuracy for distinguishing between UCD and UC. Combinations of discriminating proteins were tested in the HMP-MGH, HMP-Cin and Cedars cohort (HMP-Ced, CD=60, UC=54). Heatmap values the AUROC of a random-forest classifier for IBD subtype in the HMP-MGH and HMP-Cin cohorts for each combination of proteins. Random-forest classification was evaluated in repeated five-fold cross-validation iterations (10 repetitions). Lollipop plot depicts the performance in the HMP-Ced cohort. (F) Dietary proteomics signal was quantified at genus level. In total, 6 different cohorts were included: 4 datasets from the Human Microbiome Project (HMP, collected by Cedars-Sinai Medical Center, Cincinnati Children’s Hospital and Massachusetts General Hospital) and the German UC (GER-UC) and Israeli pediatric CD (ISR-pCD) cohorts included in this study. Percentage of proteomics signal assigned to Triticum (wheat) genus in Israeli (n=54) and German (n=100) populations. Bars represent the mean, and whiskers represent standard error. AUC, area under the curve; ROC receiver operator curve. GER, Germany; ISR, Israel; UC, Ulcerative colitis; CD, Crohn’s disease; HC, Healthy Control; SVM, Support Vector Machine; MGH, Massachusetts General Hospital; ns, nonsignificant.
[0093] FIG. 20. SPF mice were fed with a casein-based diet (circles) or an amino acid-only diet (AA, rectangles) for two weeks (n=8 per condition). IPHOMED quantification of total dietary intensity detected in stool sample from each mouse. Lines represent the mean. Significance by Mann-Whitney test.
[0094] FIG. 21. NC-fed mice were administered a protein shake, comprising proteins sourced from pea (Pisum) and rice (Oryza) (n=14), or control (n=6). Values represent relative dietary intensity of pea and rice.
[0095] FIG. 22. SPF mice were fed with either NC (n=5) or HFD (n=5) for one week. Luminal samples were collected from the stomach, duodenum, jejunum, ileum, cecum, and colon, and were subsequently analyzed by IPHOMED. Percentage of detected peptides in the stomach associated with either a pepsin-trypsin profile in both NC- and HFD-fed mice. Significance by unpaired t test.
[0096] FIG. 23. Longitudinal analysis of host- microbiome interaction at protein level during dextran sodium sulfate (DSS) colitis using IPHOMED. SPF mice (n=l 1) were treated with 2% DSS added to the drinking water for seven days followed by the resumption of regular water. Mice were monitored for 31 days by weighing and longitudinal stool collection.
[0097] FIGs. 24A-B. Dietary protein quantification in stool samples from CD patients and healthy individuals in two independent cohorts from the Human Microbiome Project: Cincinnati Children’s Hospital (HMP-Cin) and Massachusetts General Hospital (HMP-MGH). Patients were classified according to the intestinal involvement (location of bowel inflammation) of Crohn’s disease in the HMP-Cin and HMP-MGH cohorts by the Montreal classification system for GIT location (LI, ileal; L2, colonic; L3, ileocolonic; +L4, concomitant upper GI disease). (A) Total dietary protein abundance quantified in CD patients (yellow) and healthy individuals (white) from the HMP-Cin cohort. Line represents the median. Mann-Whitney test. (B) Total dietary protein abundance quantified in CD patients stratified by the affected GIT location in the HMP-Cin cohort. Line represents the median. Mann-Whitney test.
[0098] DESCRIPTION OF SPECIFIC EMBODIMENTS OF THE INVENTION
[0099] The present invention, in some embodiments thereof, relates to methods of constructing protein microbial databases for identification of proteins in a sample by proteomic analysis. The microbial databases have been specifically tailored according to metagenomics sequencing data of the sample and may optionally be combined with comprehensive and refined host and dietary databases.
[0100] Before explaining at least one embodiment of the invention in detail, it is to be understood that the invention is not necessarily limited in its application to the details set forth in the following description or exemplified by the Examples. The invention is capable of other embodiments or of being practiced or carried out in various ways.
[0101] The present inventors have now devised a high throughput metagenomic-metaproteomic pipeline that uniquely utilizes a tailored, per-experiment protein-coding sequence database to enable a comprehensive, parallel, high-resolution quantification of proteins in a sample. The database is restricted to proteins which are potentially present in the sample and can include at least one of the following sources of proteins: microbial (e.g. bacteria, virus, fungi, archaea), host and dietary proteins. The present inventors demonstrate that the database features superior performance and enhanced taxonomical resolution in functionally characterizing microbiome community structure compared to other metagenomic and proteomic approaches (see for example Figures 2A-K).
[0102] Thus, according to a first aspect of the invention there is provided a method of generating a microbial protein database comprising:
[0103] (a) sequencing at least one microbiome sample of at least one host subject species to obtain metagenomic data;
[0104] (b) aligning genomic sequences of said metagenomic data to reference genomes to obtain a database of microbial species that are potentially present in the at least one microbiome sample, and their relative abundance; (c) refining the database by:
[0105] (i) selecting those microbes having a species abundance above a predetermined level; and
[0106] (ii) selecting those microbes whose genomes have a % coverage of greater than a first predetermined % to obtain an initially refined database;
[0107] (d) obtaining protein sequences of said microbes of the initially refined database and align to said metagenomic data; and
[0108] (e) selecting those proteins having a % sequence alignment coverage greater than a second predetermined % to obtain a protein database which contains microbial proteins potentially produced by microbes in the at least one microbiome sample.
[0109] As used herein, the term “microbes” (also referred to herein as microorganisms) refers to microscopic organisms found as part of the microbiome of a region of another, larger multicellular, organism. Examples of microorganisms which may be identified using the method of the invention include bacteria (such as gram-positive bacteria, gram-positive bacterial spores, gram-negative bacteria, gram-negative bacterial spores), fungus (including fungal spores), protozoa, viruses (including bacteriophages), viroids and archaea. Microorganisms identified by the methods of the invention are live microorganisms obtained from a biological sample (e.g. a microbiome sample) of a subject.
[0110] As used herein, the term “microbiome” refers to the microorganisms of a particular region of a host subject. For example, the gut microbiome refers to the community of microorganisms in the gut.
[0111] A microbiome sample is a biological sample retrieved from a mammal (e.g. organism, organ, tissue etc.) comprising microbes of that source.
[0112] According to a particular embodiment, the microbiome sample is a human microbiome sample.
[0113] The microbiome sample may be derived from a healthy subject or a subject having a disease or disorder.
[0114] A “gut microbiome sample” in the context of the present invention is a biological sample comprising microbes of the gut (e.g. gastrointestinal tract).
[0115] According to a particular embodiment, the sample is a fecal sample.
[0116] Additional microbiomes contemplated by the present inventors include but are not limited to oral microbiome, cutaneous microbiome, vaginal microbiome, rumen microbiome. According to one embodiment the microbiome sample (e.g. fecal sample) is frozen and / or lyophilized prior to analysis. According to another embodiment, the sample may be subjected to solid phase extraction methods.
[0117] The term “metagenomics data” refers to data retrieved by metagenomics sequencing. The data may include nucleic acid sequence data and / or amino acid sequence data.
[0118] As used herein, the term “metagenomic sequencing” refers to a process wherein nucleic acids of a sample are subjected to nucleic acid sequencing (see, for example, Sharpton, 2014 Front Plant Sci. 5:209; Wang et al., 2015 World J Gastroenterol 21:803-813; Quince et al., 2017 Nat Biotechnol. 35:833-844; Kumar et al., 2017, Virus Res. 239: 172- 179). Metagenomic sequencing can be achieved using any method known in the art such as by next-generation sequencing (NGS). In an embodiment, the Illumina HiSeq X Ten System is used for metagenomics sequencing.
[0119] As the skilled person will appreciate, the portion analyzed will need to be of sufficient length to classify the taxon source of the nucleic acid gene.
[0120] Once sequencing has been performed the metagenomics data is analyzed.
[0121] In an embodiment, metagenomic analysis following metagenomic sequencing typically includes identification and quantification of genomes of microorganisms in a sample guided by genome-assembly or using a database reference. Regarding microbial abundance quantification using database reference, two different types of methodologies can be used: i) a taxonomically organised k-mer based sequence database can be generated from reference genome sequences of multiple microbes. Metagenomic reads are then assigned based on sequence alignment (e.g. identity) to determine sample species composition (e.g. Kraken described by Wood and Salzberg (2014) Genome Biology and genomebiologydotbiomedcentraldotcom / articles / 10dotl l86 / gb- 2014-15-3-r46 and Forster et al. (2019) Nat. Biotechnol. 37: 186-192); ii) profiling the composition of microbial communities using unique clade- specific marker genes identified from microbial genomes, such as Metaphlan (Duy Tin Truong et al, Genome Research, 2017), PhymmBL (Brady et al, Nature Methods, 2011).
[0122] As used herein, “reference genome” refers to a genetic sequence for a particular organism with which other sequenced genomes can be compared. Microbial genome references can be downloaded generic databases such as NCBI and Biocyc.
[0123] In an example, sequence alignment (i.e. comparison, identity) can be performed using BLAST, Megablast, BLAT and SSAHA.
[0124] In general, "sequence identity" refers to an exact nucleotide-to-nucleotide or amino acid- to-amino acid correspondence of two polynucleotides or polypeptide sequences, respectively. Typically, techniques for determining sequence identity include determining die nucleotide sequence of a polynucleotide and / or determining the amino acid sequence encoded thereby, and comparing diese sequences to a second nucleotide or amino acid sequence. Two or more sequences (polynucleotide or amino acid) can be compared by determining their "percent identity.” The percent identity of two sequences, whether nucleic acid or amino acid sequences, is the number of exact matches between two aligned sequences divided by the length of the shorter sequences and multiplied by 100. Percent identity may also be determined, for example, by comparing sequence information using the advanced BLAST computer program, including version 2.2.9, available from the National Institutes of Health. The BLAST program is based on the alignment method of Karlin and Altschul, Proc. Natl. Acad. Sci. USA 87:2264-2268 (1990) and as discussed in Altschul, et al., J. Mol. Biol. 215:403-410 (1990); Karlin And Altschul, Proc. Natl. Acad. Sci. LISA 90:5873-5877 (1993); and Altschul et ah, Nucleic Acids Res. 25:3389- 3402 (1997). Briefly, the BLAST program defines identity as the number of identical aligned symbols (i.e., nucleotides or amino acids), divided by the total number of symbols in the shorter of the two sequences. The program may be used to determine percent identity over the entire length of the proteins being compared. Default parameters are provided to optimize searches with short query sequences in, for example, with the blastp program. The program also allows use of an SEG filter to mask-off segments of the query sequences as determined by the SEG program of Wootton and Federhen. Computers and Chemistry 17: 149-163 (1993). Ranges of desired degrees of sequence identity are approximately 80% to 100% and integer values therebetween. In general, an exact match indicates 100% identity over the length of the shortest of the sequences being compared (or over the length of both sequences, if identical).
[0125] Microbes are selected having a species abundance above a predetermined level - e.g. the top 85 % of species are included in the updated database, the top 90 % of species are included in the updated database, the top 95 % of species are included in the updated database or the top 97 % of species are included in the updated database. It will be appreciated that the predetermined level can be determined according to user requirements and the predetermined levels provided herein should not be construed as limiting. A database (e.g. catalogue, dataset) of microbial species is then generated which includes the microbial species that have selected based on the previous filtering step.
[0126] The database may comprise a catalogue of microbial species / strains derived from a single microbiome sample or a plurality of microbiome samples (e.g. a number of microbiome samples in a cohort). In one embodiment, these microbiome samples used to generate the database are of healthy subjects, or patients. In one embodiment, the plurality of microbiome samples which are used to generate the database are derived from subjects having an identical disease (e.g. inflammatory bowel disease, cancer, diabetes, allergic disease such as Celiac disease etc.) and their matched healthy controls.
[0127] Metagenomics reads are then mapped to the generated genomic database using algorithms such as Minimap2 (Li et al, Bioinformatics, 2018), bwa-mem (Li et al, Bioinformatics, 2010) and bowtie2 (Langmead et al, Genome Biology, 2009).
[0128] In one embodiment, a coverage score is determined. The coverage score can be determined as a cumulative score representing a sum of the counts of aligned reads at each genomic position (e.g., base position) along the reference sequence. In another embodiment, the coverage score can be determined using a sum of genomic positions that have at least one read aligned to that position. The coverage score may be a percent of the genome that is “covered” by at least one aligned read, and thus the sum of genomic positions that have at least one read aligned to that position can be divided by the total number of positions in the database for that reference genome.
[0129] Coverage scores may be computed for all the microbial species included in the previously generated database. Microbes are selected whose genomes have a % coverage greater than a first predetermined % (e.g. only include those microbes that have 50 % coverage, 60 % coverage, 70 % coverage, 80 % coverage), as determined by the previous metagenomic alignment. It will be appreciated that the % coverage can be determined according to user requirements and the % coverages provided herein should not be construed as limiting.
[0130] A second refined protein database is generated by downloading the protein sequences encoded in the selected microbial species from the previous filtering step. These amino acid sequences can be downloaded from generic databases such as Biocyc, NCBI, Uniprot.
[0131] Similarly, metagenomic reads are then aligned to the second refined protein database using algorithms like DIAMOND (Buchfink et al, Nature Methods, 2014) and Blastx. Proteins having a predetermined % coverage are selected - e.g. (e.g. only include those microbes that have 50 % coverage, 60 % coverage, 70 % coverage, 80 % coverage). It will be appreciated that the % coverage can be determined according to user requirements and the % coverages provided herein should not be construed as limiting. This set of proteins represents the catalog of microbial proteins that are included in the global protein database.
[0132] According to another embodiment, metagenomic analysis following metagenomic sequencing typically includes identification and quantification of genomes of microorganisms in a sample guided by genome-assembly. Metagenomic reads may be directly assembled to generate metagenome assembled genomes (MAGs) with species composition and proportion determined through prevalence of these sequences within the complete sample (see, for example, Almeida (2019) Nature 568:499-504). Protein sequences can be detected in those MAGs by algorithms such as blastp, DIAMOND to search for NCBI’s COG database or PRODIGAL (Hyatt, BMC Bioinformatics, 2010). This set of proteins represents the catalog of microbial proteins that are included in the global protein database.
[0133] By carrying out one of the two above mentioned actions (genome-assembly or database reference), a microbial protein database may be generated which has an increased likelihood to only contain proteins that are produced by microbes in the sample (as determined by metagenomic data analysis).
[0134] According to user requirements, additional protein sequences may be added to the database generated using the methods described herein. The protein sequences are typically those that are likely to be present in the microbiome sample. Thus, for example, if a gastrointestinal (e.g. fecal) microbiome sample is being analyzed, the database may further include host proteins (e.g. human proteins) and / or dietary proteins.
[0135] In one embodiment, the additional protein sequences are of the host species from where the microbiome samples were taken. For example, in one embodiment, human protein and dietary protein sequences are added to the database. Host protein sequences may be found for example UniProt, NCBI Protein Database, Ensembl and Human Protein Atlas. Dietary protein sequences can be also downloaded from UniProt and NCBI Protein Database. Refinements can be performed to the dietary database. Proteins from dietary species included in curated dietary databases, such as FooDB (wwwdotfoodbdotca / ), can be downloaded from UniProt and NCBI Protein Database. Using protein sequence clustering algorithms, such as MMseqs2 (Mirdita et al, Bioinformatics, 2021), proteins with a minimum sequence identity of and alignment coverage are selected. This reduced set of proteins (dietary marker proteins) represents the catalog of dietary proteins that are included in the global protein database
[0136] Embodiments of the invention use a multi-omic approach to help increase the accuracy of proteomic analysis. Methods of the invention use metagenomic data to assist the generation of target protein list and / or to improve protein identification using peptide-to-protein mapping. Methods of the invention use a generalized methodology that is based on metagenomic level data and proteome level data, regardless of what techniques are used to obtain these data. Therefore, methods of the invention may be used with a variety of technologies.
[0137] According to another aspect of the invention there is provided a method of profiling a microbiome sample of a host subject comprising:
[0138] (a) generating a protein database according to the methods described herein above; (b) performing a mass spectrometry (MS) analysis on a protein extract of said microbiome sample; and
[0139] (c) performing a database search to identify peptides and protein of origin, using the generated protein database in (a).
[0140] In one embodiment, the profiling is taxonomical profiling and refers to classifying a taxon from where a peptide of the sample is derived.
[0141] The terms "taxon" (plural "taxa"), “taxonomic group," and "taxonomic unit" are used interchangeably to refer to a group of one or more organisms that comprises a node in a clustering tree. The level of a cluster is determined by its hierarchical order. In one embodiment, a taxon is a group tentatively assumed to be a valid taxon for purposes of phylogenetic analysis. In another embodiment, a taxon is any of the extant taxonomic units under study. In yet another embodiment, a taxon is given a name and a rank. For example, a taxon can represent a domain, a sub-domain, a kingdom, a sub-kingdom, a phylum, a sub- phylum, a class, a sub-class, an order, a sub-order, a family, a subfamily, a genus, a subgenus, or a species. In some embodiments, taxa can represent one or more organisms from the kingdoms eubacteria, protista, or fungi at any level of a hierarch al order.
[0142] In one embodiment, a microbiome sample is obtained from a host subject. A fraction of the microbiome sample is manipulated to obtain a DNA sample thereof and used to generate the reference protein database, as described herein above. Another fraction of the microbiome sample is manipulated to obtain a protein sample (i.e. protein extract). The protein sample is then analyzed by mass spectrometry (MS) analysis.
[0143] Methods of performing mass spectrometry (MS) analysis are described herein below.
[0144] The mass spectrometry analysis step is carried out on a paired metagenomic sample (wherein one sample is a DNA sample and the other sample is a protein sample).
[0145] In one embodiment, a fecal sample is profiled. The database includes, as described above, amino acid sequence of microbial proteins of the sample and optionally includes amino acid sequences of host proteins (e.g. human proteins) and / or dietary proteins. It will be appreciated that the method may be used to identify a microbial origin, a food origin and / or a host subject origin.
[0146] According to another aspect of the present invention there is provided a method of obtaining an indication for a subject suspected of having a disease comprising:
[0147] (a) performing a mass spectrometry (MS) analysis on a protein extract of a microbiome sample of the subject to obtain sequences of peptides present in said protein extract; (b) performing a database search to identify both the protein source of the peptides and to quantify an abundance of the protein source, said database being generated according to the methods described herein; and
[0148] (c) comparing the abundance of the proteins identified in step (b) to an abundance in a protein extract of a microbiome sample of a healthy control subject, wherein a change in said abundance is indicative of the disease.
[0149] As above, the mass spectrometry analysis step is carried out on a paired metagenomic sample.
[0150] The indication may be used to rule in or rule out whether the subject has the disease or is likely to have the disease. In some embodiments, the indication may be a classification of the disease, a progression of the disease, a response to a therapy, a forecast concerning the outcome of the disease, and / or determining prospects of recovery from the disease.
[0151] The indication is determined based on abundance of at least one, two, three, four, five, six, seven, eight, nine, ten or more proteins as determined using proteomic analysis and compared to control proteomic data (e.g. derived from a healthy subject, or derived from a subject where it is known the subject does not suffer from the disease).
[0152] A change in abundance may be an upregulation or a downregulation (e.g. in a statistically significant manner).
[0153] In one embodiment, the amount of protein is up / down-regulated by at least 1.5 fold, at least 2 fold, at least 3 fold, at least 4 fold or more compared to the amount of the protein in the control sample.
[0154] It will be appreciated that as well as analyzing protein abundance, the methods described herein can be used to analyze abundance of bacteria present in the sample. The abundance of at least one, two, three, four, five, six, seven, eight, nine, ten or more bacteria identified using the methods described herein compared to an abundance of those bacteria in a protein extract of a microbiome sample of a healthy control subject is indicative of the disease.
[0155] According to still another aspect of the invention there is provided a method of analyzing whether a subject is complying with a dietary intervention using the same methodology described above. An amount or presence of said identified protein which is associated with the dietary intervention in the protein sample is indicative as to whether the subject is complying with the dietary intervention.
[0156] In one embodiment, the amount of at least one, two, three, four, five, six, seven, eight, nine, ten or more identified proteins is used to determine whether the subject is complying with the dietary invention. For example, if the dietary intervention is provision of a particular food type, proteins associated with that food type may be identified in order to determine compliance. If the dietary intervention is absence of a particular food type, if proteins associated with that food type are identified, it is indicative that the subject is not complying with the dietary intervention.
[0157] In another embodiment, a total amount of dietary proteins is determined. This data provides information concerning human nutritional absorption. An amount of total dietary proteins above a predetermined level (e.g. significantly higher than that in a healthy subject) is indicative that the subject is suffering from a disease / disorder which is related to nutritional absorption (e.g. inflammatory bowel disease). In one embodiment, the total amount of dietary protein is used to determine upper intestinal protein malabsorption of a subject.
[0158] Examples of diseases which may be treated by dietary intervention include, but are not limited to obesity, heart disease, cancer, colitis, Crohn’s disease, Celiac disease, food allergies. According to the dietary intervention, a person of skill in the art is able to determine which proteins or types of proteins are predictive of compliance to the dietary intervention. Thus, for example, a reduction of fecal dietary proteins in a fecal sample of a subject suffering from IBD and being on an anti-inflammatory dietary intervention of exclusive enteral nutrition (EEN) is indicative that the subject is complying with the intervention. Thus, for example, a reduction of fecal dietary gluten proteins in a fecal sample of a subject suffering from Celiac disease and being on a non-gluten dietary intervention is indicative that the subject is complying with the intervention.
[0159] Using the annotated protein database generated according to the pipeline described herein, the present inventors were able to identify numerous bacterial protein markers (by mass spectrometry) which were up or down regulated in a disease state (e.g. inflammatory bowel disease) as compared to in a non-disease state (see for example Figures 16A-B). Furthermore, numerous bacteria could also be identified (both on the species level and the functional (KO) level) which can potentially serve as makers for the disease state (see for example Figures 10G and 11G).
[0160] In addition, numerous host proteins were also uncovered which could serve as markers for a disease state (Figures 16A-B).
[0161] Combinations of proteins identified by supervised machine learning were shown to have a higher degree of accuracy for ruling in IBD as compared with the clinically used biomarker calprotectin (comprised of S100A8 and S100A9) - see for example Figures 13C-G.
[0162] Thus, according to another aspect of the invention there is provided a method of diagnosing Crohn’s disease (CD) of a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 1, wherein an amount of said proteins is indicative of the subject having CD.
[0163] According to another aspect of the present invention there is provided a method of diagnosing ulcerative colitis (UC) of a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 2, wherein an amount of said proteins is indicative of the subject having UC.
[0164] According to another aspect of the present invention there is provided a method of distinguishing between UC and CD in a subject comprising measuring in a gut microbiome sample of the subject for an amount of a combination of proteins set forth in a row of Table 3, wherein an amount of said proteins is indicative as to whether the subject has UC or CD.
[0165] According to another aspect of the present invention there is provided a method of determining the severity of CD in a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 4, wherein an amount of said proteins is indicative as to the severity of the CD.
[0166] According to another aspect of the present invention there is provided a method of determining the severity of UC in a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 5, wherein an amount of said proteins is indicative as to the severity of the UC.
[0167] For any of the aspects disclosed herein, the term “measuring” or “measurement,” or alternatively “detecting” or “detection,” means assessing the presence, absence, quantity or amount (which can be an effective amount) of the determinant (e.g. protein) within a clinical or subject-derived sample, including the derivation of qualitative or quantitative concentration levels of such determinants.
[0168] In one embodiment, the diagnosing comprises ruling in or ruling out whether the subject has the disease or is likely to have the inflammatory bowel disease (e.g. ulcerative colitis or Crohn's disease). In some embodiments, the diagnosing comprises classifying the inflammatory bowel disease or a symptom thereof, determining a severity of the inflammatory bowel disease, monitoring inflammatory bowel disease progression, forecasting an outcome of inflammatory bowel disease and / or determining prospects of recovery from the inflammatory bowel disease.
[0169] The microbiome sample may be taken from a subject known or suspected of having a inflammatory bowel disease for which a definitive positive or negative diagnosis is not available via clinical tests, The sample may be taken from a subject experiencing unexplained symptoms, such as fatigue, nausea, weight loss, aches and pains, weakness. The sample may be taken from a subject having explained symptoms (e.g. abdominal pain, diarrhea, rectal bleeding, and fever). The sample may be taken from a subject at risk of developing IBD e.g. as a result from, e.g., a family history of the disease, a genotype which predisposes to the disease, or phenotypic symptoms which predispose to the disease.
[0170] According to a particular embodiment, the subject is a human. The subject can be male or female. The subject may be an adult (e.g. older than 18, 21, or 22 years or a child (e.g. younger than 18, 21 or 22 years). In another embodiment, the subject is an adolescent (between 12 and 21 years), an infant (29 days to less than 2 years of age).
[0171] The proteins may be analyzed using any method known in the art. In one embodiment, the protein is analyzed on the protein level using methods including Western, Enzyme linked immunosorbent assay (ELISA), Fluorescence activated cell sorting (FACS) automated immunoassay, lateral flow immunoassay, immunohistochemical analysis etc.
[0172] In another embodiment, the proteins are measured on the RNA level using methods including for example Northern blot analysis and PCR.
[0173] The polypeptide determinants can be detected in any suitable manner, but are typically detected by contacting a sample from the subject with an antibody, which binds the determinant and then detecting the presence or absence of a reaction product. The antibody may be monoclonal, polyclonal, chimeric, or a fragment of the foregoing, as discussed in detail above, and the step of detecting the reaction product may be carried out with any suitable immunoassay. The sample from the subject is typically a biological sample as described above, and may be the same sample of biological sample used to conduct the method described above.
[0174] In one embodiment, the antibody which specifically binds the determinant is attached (either directly or indirectly) to a signal producing label, including but not limited to a radioactive label, an enzymatic label, a hapten, a reporter dye or a fluorescent label.
[0175] Immunoassays carried out in accordance with some embodiments of the present invention may be homogeneous assays or heterogeneous assays. In a homogeneous assay the immunological reaction usually involves the specific antibody (e.g., anti- determinant antibody), a labeled analyte, and the sample of interest. The signal arising from the label is modified, directly or indirectly, upon the binding of the antibody to the labeled analyte. Both the immunological reaction and detection of the extent thereof can be carried out in a homogeneous solution. Immunochemical labels, which may be employed, include free radicals, radioisotopes, fluorescent dyes, enzymes, bacteriophages, or coenzymes.
[0176] In a heterogeneous assay approach, the reagents are usually the sample, the antibody, and means for producing a detectable signal. Samples as described above may be used. The antibody can be immobilized on a support, such as a bead (such as protein A and protein G agarose beads), plate or slide, and contacted with the specimen suspected of containing the antigen in a liquid phase. The support is then separated from the liquid phase and either the support phase or the liquid phase is examined for a detectable signal employing means for producing such signal. The signal is related to the presence of the analyte in the sample.
[0177] Means for producing a detectable signal include the use of radioactive labels, fluorescent labels, or enzyme labels. For example, if the antigen to be detected contains a second binding site, an antibody which binds to that site can be conjugated to a detectable group and added to the liquid phase reaction solution before the separation step. The presence of the detectable group on the solid support indicates the presence of the antigen in the test sample. Examples of suitable immunoassays are oligonucleotides, immunoblotting, immunofluorescence methods, immunoprecipitation, chemiluminescence methods, electrochemiluminescence (ECL) or enzyme-linked immunoassays.
[0178] Those skilled in the art will be familiar with numerous specific immunoassay formats and variations thereof which may be useful for carrying out the method disclosed herein. See generally E. Maggio, Enzyme-Immunoassay, (1980) (CRC Press, Inc., Boca Raton, Fla.); see also U.S. Pat. No. 4,727,022 to Skold et al., titled “Methods for Modulating Ligand-Receptor Interactions and their Application,” U.S. Pat. No. 4,659,678 to Forrest et al., titled “Immunoassay of Antigens,” U.S. Pat. No. 4,376,110 to David et al., titled “Immunometric Assays Using Monoclonal Antibodies,” U.S. Pat. No. 4,275,149 to Litman et al., titled “Macromolecular Environment Control in Specific Receptor Assays,” U.S. Pat. No. 4,233,402 to Maggio et al., titled “Reagents and Method Employing Channeling,” and U.S. Pat. No. 4,230,767 to Boguslaski et al., titled “Heterogeneous Specific Binding Assay Employing a Coenzyme as Label.” The determinant can also be detected with antibodies using flow cytometry. Those skilled in the art will be familiar with flow cytometric techniques which may be useful in carrying out the methods disclosed herein. These include, without limitation, Cytokine Bead Array (Becton Dickinson) and Luminex technology.
[0179] Antibodies can be conjugated to a solid support suitable for a diagnostic assay (e.g., beads such as protein A or protein G agarose, microspheres, plates, slides or wells formed from materials such as latex or polystyrene) in accordance with known techniques, such as passive binding. Antibodies as described herein may likewise be conjugated to detectable labels or groups such as radiolabels (e.g.,35S,125I,131I), enzyme labels (e.g., horseradish peroxidase, alkaline phosphatase), and fluorescent labels (e.g., fluorescein, Alexa, green fluorescent protein, rhodamine) in accordance with known techniques. Antibodies can also be useful for detecting post-translational modifications of determinant proteins, polypeptides, mutations, and polymorphisms, such as tyrosine phosphorylation, threonine phosphorylation, serine phosphorylation, glycosylation (e.g., O- GlcNAc). Such antibodies specifically detect the phosphorylated amino acids in a protein or proteins of interest, and can be used in immunoblotting, immunofluorescence, and ELISA assays described herein. These antibodies are well-known to those skilled in the art, and commercially available. Post-translational modifications can also be determined using metastable ions in reflector matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDL TOF) (Wirth U. and Muller D. 2002).
[0180] For determinant-proteins, polypeptides, mutations, and polymorphisms known to have enzymatic activity, the activities can be determined in vitro using enzyme assays known in the art. Such assays include, without limitation, kinase assays, phosphatase assays, reductase assays, among many others. Modulation of the kinetics of enzyme activities can be determined by measuring the rate constant KM using known algorithms, such as the Hill plot, Michaelis-Menten equation, linear regression plots such as Lineweaver-Burk analysis, and Scatchard plot.
[0181] In particular embodiments, the antibodies of the present invention are monoclonal antibodies.
[0182] Suitable sources for antibodies for the detection of determinants include commercially available sources such as, for example, Abazyme, Abnova, AssayPro, Affinity Biologicals, AntibodyShop, Aviva bioscience, Biogenesis, Biosense Laboratories, Calbiochem, Cell Sciences, Chemicon International, Chemokine, Clontech, Cytolab, DAKO, Diagnostic BioSystems, eBioscience, Endocrine Technologies, Enzo Biochem, Eurogentec, Fusion Antibodies, Genesis Biotech, GloboZymes, Haematologic Technologies, Immunodetect, Immunodiagnostik, Immunometrics, Immunostar, Immunovision, Biogenex, Invitrogen, Jackson ImmunoResearch Laboratory, KMI Diagnostics, Koma Biotech, LabFrontier Life Science Institute, Lee Laboratories, Lifescreen, Maine Biotechnology Services, Mediclone, MicroPharm Ltd., ModiQuest, Molecular Innovations, Molecular Probes, Neoclone, Neuromics, New England Biolabs, Novocastra, Novus Biologicals, Oncogene Research Products, Orbigen, Oxford Biotechnology, Panvera, PerkinElmer Life Sciences, Pharmingen, Phoenix Pharmaceuticals, Pierce Chemical Company, Polymun Scientific, Polysiences, Inc., Promega Corporation, Proteogenix, Protos Immunoresearch, QED Biosciences, Inc., R&D Systems, Repligen, Research Diagnostics, Roboscreen, Santa Cruz Biotechnology, Seikagaku America, Serological Corporation, Serotec, SigmaAldrich, StemCell Technologies, Synaptic Systems GmbH, Technopharm, Terra Nova Biotechnology, TiterMax, Trillium Diagnostics, Upstate Biotechnology, US Biological, Vector Laboratories, Wako Pure Chemical Industries, and
[0183] Zeptometrix. However, the skilled artisan can routinely make antibodies, against any of the polypeptide determinants described herein.
[0184] Polyclonal antibodies for measuring polypeptides include without limitation antibodies that were produced from sera by active immunization of one or more of the following: Rabbit, Goat, Sheep, Chicken, Duck, Guinea Pig, Mouse, Donkey, Camel, Rat and Horse.
[0185] Examples of additional detection agents, include without limitation: scFv, dsFv, Fab, sVH, F(ab')2, Cyclic peptides, Haptamers, A single-domain antibody, Fab fragments, Singlechain variable fragments, Affibody molecules, Affilins, Nanofitins, Anticalins, Avimers, DARPins, Kunitz domains, Fynomers and Monobody.
[0186] The presence of a label can be detected by inspection, or a detector which monitors a particular probe or probe combination is used to detect the detection reagent label. Typical detectors include spectrophotometers, phototubes and photodiodes, microscopes, scintillation counters, cameras, film and the like, as well as combinations thereof. Those skilled in the art will be familiar with numerous suitable detectors that widely available from a variety of commercial sources and may be useful for carrying out the method disclosed herein. Commonly, an optical image of a substrate comprising bound labeling moieties is digitized for subsequent computer analysis. See generally The Immunoassay Handbook [The Immunoassay Handbook. Third Edition. 2005].
[0187] In another embodiment, the proteins are measured by mass spectrometry.
[0188] Samples may be processed or purified io obtain preparations that are suitable for analysis by mass spectrometry. Such purification will usually include chromatography, such as liquid chromatography, and may also often involve an additional purification procedure that is performed prior to chromatography. Various procedures may be used for this purpose depending on the type of sample or the type of chromatography. Examples include filtration, centrifugation, combinations thereof and the like.
[0189] Various methods have been described involving the use of HPLC for sample clean-up prior to mass spectrometry analysis. One of skill in the art may select HPLC instruments and columns that are suitable for use in the methods. The chromatographic column typically includes a medium (i.e., a packing material) to facilitate separation of chemical moieties (i.e., fractionation). The medium may include minute particles. The particles include a bonded surface that interacts with the various chemical moieties to facilitate separation of the chemical moieties. One suitable bonded surface is a hydrophobic bonded surface such as an alkyl bonded surface. Alkyl bonded surfaces may include C-4, C-8, or C-18 bonded alkyl groups. ’The chromatographic column includes an inlet port for receiving a sample and an outlet port for discharging an effluent that includes the fractionated sample.
[0190] In certain embodiments, the proteins are purified by applying a sample to a column under conditions where the molecule of interest is reversibly retained by the column packing material, while one or more other materials are not retained. In these embodiments, a first mobile phase condition can be employed where protein of interest is retained by the column and a second mobile phase condition can subsequently be employed to remove retained material from the column, once the non-retained materials are washed through. Alternatively, a protein may be purified by applying a sample to a column under mobile phase conditions where the analyte of interest elutes at a differential rate in comparison to one or more other materials, Such procedures may enrich the amount of one or more proteins of interest relative to one or more other components of the sample.
[0191] Numerous column packings are available for chromatographic separation of samples and selection of an appropriate separation protocol is an empirical process that depends on the sample characteristics, analyte of interest, presence of interfering substances and their characteristics, etc. Commercially available HPLC columns include, but are not limited to, polar-, ion exchange (both cation and anion), hydrophobic interaction, phenyl, C-2, C-8, C-18, and polar coating on porous polymer columns.
[0192] During chromatography, the separation of materials is effected by variables such as choice of eluent (also known as a “mobile phase”), choice of gradient elution and the gradient conditions, temperature, etc.
[0193] In various embodiments, the proteins may then be ionized by any method known to the skilled artisan. Mass spectrometry is performed using a mass spectrometer, which includes an ion source for ionizing the fractionated sample and creating charged molecules for further analysis.
[0194] Ionization sources used in various MS techniques include, but are not limited to, electron ionization, chemical ionization, electrospray ionization (ESI), photon ionization, atmospheric pressure chemical ionization (APCI), photoionization, atmospheric pressure photoionization (APPI), fast atom bombardment (FAB ( / liquid secondary' ionization (LSIMS), matrix assisted laser desorption ionization (MALDI), field ionization, field desorption, thermospray / plasmaspray ionization, surface enhanced laser desorption ionization (SELDI), inductively coupled plasma (ICP) and particle beam ionization. The skilled artisan will understand that the choice of ionization method may be determined based on the analyte to be measured, type of sample, the type of detector, the choice of positive versus negative mode, etc. After the sample has been ionized, the positively charged ions thereby created may be analyzed to determine m / z. Suitable analyzers for determining m / z include quadrupole analyzers, ion trap analyzers, and time-of-flight analyzers. The ions may be detected using one of several detection modes. For example, only selected ions may be detected using a selective ion monitoring mode (SIM), or alternatively, multiple ions may be detected using a scanning mode, e.g., multiple reaction monitoring (MRM) or selected reaction monitoring (SRM). In preferred embodiments, ions are detected using SRM.
[0195] One may enhance the resolution of the MS technique by entploying “tandem mass spectrometry,” or “MS / MS.” In this technique, a precursor ion (also called a parent ion) generated from a molecule of interest can be filtered in an MS instrument, and the precursor ion subsequently fragmented to yield one or more fragment ions (also called daughter ions or product ions) that are then analyzed in a second MS procedure. By careful selection of precursor ions, only ions produced by certain analytes are passed to the fragmentation chamber, where collision with atoms of an inert gas produce the fragment ions. Because both the precursor and fragment ions are produced in a reproducible fashion under a given set of ionization / fragmentation conditions, the MS / MS technique may provide an extremely powerful analytical tool. For example, the combination of filtration / fragmentation may be used to eliminate interfering substances, and may be particularly useful in complex samples, such as biological samples,
[0196] Additionally, recent advances in technology, such as matrix-assisted laser desorption ionization coupled with time-of-flight analyzers (“MALDI-TOF’) permit the analysis of analytes at femtomole levels in very short ion pulses. Mass spectrometers that combine time-of-flight analyzers with tandem MS are also well known to the artisan. Additionally, multiple mass spectrometry steps may be combined in methods known as “MS / MS”. Various other combinations may be employed, such as MS / MS / TOF, MALDIZMS / MS / TOF, or SELDI / MS / MS / TOF mass spectrometry.
[0197] The mass spectrometer typically provides the user with an ion scan: that is, the relative abundance of each ion with a particular m / z over a given range (e.g,, 400 to 1600 amu), The results of an analyte assay, that is, a mass spectrum, may be related to the amount of the analyte in the original sample by numerous methods known in the art. For example, given that sampling and analysis parameters are carefully controlled, the relative abundance of a given ion may be compared to a table (e.g. the database described herein) that converts that relative abundance to an absolute amount of the original molecule. Alternatively, molecular standards may be run with the samples and a standard curve constructed based on ions generated from those standards. Using such a standard curve, the relative abundance of a given ion may be converted into an absolute amount of the original molecule. In certain preferred embodiments, an internal standard is used to generate a standard curve for calculating the quantity of one of the proteins described herein. Methods of generating and using such standard curves are well known in the art and one of ordinary skill is capable of selecting an appropriate internal standard. Numerous other methods for relating the amount of an ion to the amount of the original molecule will be well known to those of ordinary skill in the art.
[0198] One or more steps of the methods may be performed using automated machines. In certain embodiments, one or more purification steps are performed on-line, and more preferably all of the LC purification and mass spectrometry steps may be performed in an on-line fashion.
[0199] In one embodiment, at least one of the proteins recited in Figure 16A or 16B is measured in order to diagnose IBD (e.g. UC or CD).
[0200] A change (e.g. an increase or decrease) in the amount of at least one of the proteins described herein, above or below a predetermined level, is indicative of a subject having IBD. In one embodiment, the predetermined level is the amount of (or a function of the amount of) the protein in a control sample derived from one or more subjects who do not have the disorder (i.e., healthy individuals) or who do not have IBD. In a further embodiment, such subjects are monitored and / or periodically retested for a diagnostically relevant period of time (“longitudinal studies”) following such test to verify continued absence of IBD. Such period of time may be one month, two months, two to five months, five months, five to ten months, ten month, or ten or more months from the initial testing date for determination of the reference value. Furthermore, retrospective measurement of protein levels in properly banked historical subject samples may be used in establishing these reference values, thus shortening the study time required.
[0201] A reference value can also comprise the amounts of protein markers derived from subjects who show an improvement as a result of treatments and / or therapies for the IBD. A reference value can also comprise the amounts of protein markers derived from subjects who have confirmed IBD by known techniques.
[0202] In one embodiment, at least two of the proteins recited in Figure 16A or 16B is measured in order to diagnose IBD.
[0203] In one embodiment, at least three of the proteins recited in Figure 16A or 16B is measured in order to diagnose IBD.
[0204] In one embodiment, at least four of the proteins recited in Figure 16A or 16B is measured in order to diagnose IBD. In one embodiment, at least five of the proteins recited in Figure 16A or 16B is measured in order to diagnose IBD.
[0205] According to a specific embodiment, at least two of the proteins appearing in the same row of Table 1 are measured for ruling in Crohn’s disease Table 1
[0206] According to a specific embodiment, at least two of the proteins appearing in the same row of Table 2 are measured for ruling in Ulcerative colitis.
[0207] Table 2
[0208] According to a specific embodiment, at least two of the proteins appearing in the same row of Table 3 are measured for distinguishing between CD and UC. Table 3
[0209] According to a specific embodiment, at least two of the proteins appearing in the same row of Table 4 are measured for determining severity of CD.
[0210] Table 4
[0211] According to a specific embodiment, at least two of the proteins appearing in the same row of Table 5 are measured for determining severity of UC.
[0212] Table 5 Preferably the combinations which are us to diagnose / classify the IBD do not exceed 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, or 2 proteins. In another embodiment, no more than 40 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 30 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 20 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 10 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 9 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 8 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 7 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 6 protein markers are analyzed in a single test / analysis, for the diagnosis / classification. In another embodiment, no more than 5 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 4 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 3 protein markers are analyzed in a single test / analysis for the diagnosis / classification. In another embodiment, no more than 2 protein markers are analyzed in a single test / analysis for the diagnosis / classification.
[0213] On the basis of the classification of the disease, clinical decisions may be made.
[0214] According to some embodiments of the invention, the method further comprises informing the subject of results of the diagnosis / classification.
[0215] As used herein the phrase “informing the subject” refers to advising the subject that based on the diagnosis the subject should seek a suitable treatment regimen.
[0216] Once the diagnosis is determined, the results can be recorded in the subject’s medical file, which may assist in selecting a treatment regimen and / or determining prognosis of the subject (e.g.. colonoscopy, watchful waiting, referral to a gastroenterologist, etc.).
[0217] Once the classifications are made, additional tests may be made in order to corroborate the result or to further classify the disease.
[0218] In some embodiments, the methods disclosed herein are useful in monitoring the treatment of IBD. For example, in some embodiments, the methods may be performed immediately before, during and / or after a treatment (e.g. a. dietary treatment) to monitor treatment success. In some embodiments, the methods are performed at intervals on disease free patients to insure treatment success. Kits
[0219] Some aspects of the invention also include a protein marker-detection reagent such as combinations of internal standards (for mass spec analysis) or combinations of antibodies corresponding to particular protein markers packaged together in the form of a kit. The kit may contain in separate containers internal standards, control formulations (positive and / or negative). The kits of this aspect of the present invention may comprise additional components that aid in the detection of the determinants such as columns, buffers etc.
[0220] Exemplary combinations of internal standards and antibodies are provided herein above in Tables 1-5.
[0221] According to another embodiment of the invention there is provided a method of diagnosing an inflammatory bowel disease (IBD) of a subject comprising measuring in a gut microbiome sample of the subject an amount of at least one bacteria selected from the group consisting of Alistipes putredinis, Dialister succinatiphilus, Oscillibacter sp. ER4, Romboutsia sp. DR1, Gemmiger formicilis, Faecalibacterium prausnitzii, Fusicatenibacter saccharivorans, Ruminococcus callidus and Bacteroides coprocola, wherein a decrease in the amount of said at least one bacteria as compared to an amount in a gut microbiome sample of a healthy subject is indicative of the subject having IBD.
[0222] Measuring a level or presence of a microbe may be affected by analyzing for the presence of microbial component or a microbial by-product. Thus, for example the level or presence of a microbe may be affected by measuring the level of a DNA sequence. In some embodiments, the level or presence of a microbe may be affected by measuring 16S rRNA gene sequences or 18S rRNA gene sequences. In other embodiments, the level or presence of a microbe may be affected by measuring RNA transcripts. In still other embodiments the level or presence of a microbe may be affected by measuring proteins. In still other embodiments, the level or presence of a microbe may be affected by measuring metabolites.
[0223] Quantifying Microbial Levels:
[0224] It will be appreciated that determining the abundance of microbes may be affected by taking into account any feature of the microbiome. Thus, the abundance of microbes may be affected by taking into account the abundance at different phylogenetic levels; at the level of gene abundance; gene metabolic pathway abundances; sub-species strain identification; SNPs and insertions and deletions in specific bacterial regions; growth rates of bacteria, the diversity of the microbes of the microbiome, as further described herein below.
[0225] In some embodiments, determining a level or set of levels of one or more types of microbes or components or products thereof comprises determining a level or set of levels of one or more DNA sequences. In some embodiments, one or more DNA sequences comprises any DNA sequence that can be used to differentiate between different microbial types. In certain embodiments, one or more DNA sequences comprises 16S rRNA gene sequences. In certain embodiments, one or more DNA sequences comprises 18S rRNA gene sequences. In some embodiments, 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, 100, 1,000, 5,000 or more sequences are amplified.
[0226] 16S and 18S rRNA gene sequences encode small subunit components of prokaryotic and eukaryotic ribosomes respectively. rRNA genes are particularly useful in distinguishing between types of microbes because, although sequences of these genes differs between microbial species, the genes have highly conserved regions for primer binding. This specificity between conserved primer binding regions allows the rRNA genes of many different types of microbes to be amplified with a single set of primers and then to be distinguished by amplified sequences.
[0227] In some embodiments, a microbiota sample (e.g. fecal sample) is directly assayed for a level or set of levels of one or more DNA sequences. In some embodiments, DNA is isolated from a microbiota sample and isolated DNA is assayed for a level or set of levels of one or more DNA sequences. Methods of isolating microbial DNA are well known in the art. Examples include but are not limited to phenol-chloroform extraction and a wide variety of commercially available kits, including QIAamp DNA Stool Mini Kit (Qiagen, Valencia, Calif.).
[0228] In some embodiments, a level or set of levels of one or more DNA sequences is determined by amplifying DNA sequences using PCR (e.g., standard PCR, semi-quantitative, or quantitative PCR). In some embodiments, a level or set of levels of one or more DNA sequences is determined by amplifying DNA sequences using quantitative PCR. These and other basic DNA amplification procedures are well known to practitioners in the art and are described in Ausebel et al. (Ausubel F M, Brent R, Kingston R E, Moore D, Seidman J G, Smith J A, Struhl K (eds). 1998. Current Protocols in Molecular Biology. Wiley: New York).
[0229] In some embodiments, DNA sequences are amplified using primers specific for one or more sequence that differentiate(s) individual microbial types from other, different microbial types. In some embodiments, 16S rRNA gene sequences or fragments thereof are amplified using primers specific for 16S rRNA gene sequences. In some embodiments, 18S DNA sequences are amplified using primers specific for 18S DNA sequences.
[0230] In some embodiments, a level or set of levels of one or more 16S rRNA gene sequences is determined using phylochip technology. Use of phylochips is well known in the art and is described in Hazen et al. ("Deep-sea oil plume enriches indigenous oil-degrading bacteria." Science, 330, 204-208, 2010), the entirety of which is incorporated by reference. Briefly, 16S rRNA genes sequences are amplified and labeled from DNA extracted from a microbiota sample. Amplified DNA is then hybridized to an array containing probes for microbial 16S rRNA genes. Level of binding to each probe is then quantified providing a sample level of microbial type corresponding to 16S rRNA gene sequence probed. In some embodiments, phylochip analysis is performed by a commercial vendor. Examples include but are not limited to Second Genome Inc. (San Francisco, Calif.).
[0231] In some embodiments, determining a level or set of levels of one or more types of microbes comprises determining a level or set of levels of one or more microbial RNA molecules (e.g., transcripts). Methods of quantifying levels of RNA transcripts are well known in the art and include but are not limited to northern analysis, semi-quantitative reverse transcriptase PCR, quantitative reverse transcriptase PCR, and microarray analysis.
[0232] In some embodiments, determining a level or set of levels of one or more types of microbes comprises determining a level or set of levels of one or more microbial polypeptides. Methods of quantifying polypeptide levels are well known in the art and include but are not limited to Western analysis and mass spectrometry.
[0233] As mentioned herein above, as well as (or instead of) analyzing the level of microbes, the present invention also contemplates analyzing the level of microbial products.
[0234] Examples of microbial products include, but are not limited to mRNAs, polypeptides, carbohydrates and metabolites.
[0235] In some embodiments, the presence, level, and / or activity of metabolites of at least ten types of microbes are measured. In other embodiments, the presence, level, and / or activity of metabolites of between 2 and 100 types of microbes are measured.
[0236] Performance and Accuracy Measures of the Invention.
[0237] The performance and thus absolute and relative clinical usefulness of the invention may be assessed in multiple ways as noted above. Amongst the various assessments of performance, some aspects of the invention are intended to provide accuracy in clinical diagnosis and prognosis. The accuracy of a diagnostic or prognostic test, assay, or method concerns the ability of the test, assay, or method to rule in whether subjects having a colon cell proliferative disorder is based on whether the subjects have, a “significant alteration” (e.g., clinically significant and diagnostically significant) in the levels of a determinant. By “effective amount” it is meant that the measurement of an appropriate number of determinants (which may be one or more) to produce a “significant alteration” (e.g. level of expression or activity of a determinant) that is different than the predetermined cut-off point (or threshold value) for that determinant (s) and therefore indicates that the subject has a colon cell proliferative disorder for which the determinant (s) is an indication. The difference in the level of determinant is preferably statistically significant. As noted below, and without any limitation of the invention, achieving statistical significance, and thus the preferred analytical, diagnostic, and clinical accuracy, may require that combinations of several determinants be used together in panels and combined with mathematical algorithms in order to achieve a statistically significant determinant index.
[0238] In the categorical diagnosis of a disease state, changing the cut point or threshold value of a test (or assay) usually changes the sensitivity and specificity, but in a qualitatively inverse relationship. Therefore, in assessing the accuracy and usefulness of a proposed medical test, assay, or method for assessing a subject’s condition, one should always take both sensitivity and specificity into account and be mindful of what the cut point is at which the sensitivity and specificity are being reported because sensitivity and specificity may vary significantly over the range of cut points. One way to achieve this is by using the Matthews correlation coefficient (MCC) metric, which depends upon both sensitivity and specificity. Use of statistics such as area under the ROC curve (AUC), encompassing all potential cut point values, is preferred for most categorical risk measures when using some aspects of the invention, while for continuous risk measures, statistics of goodness-of-fit and calibration to observed results or other gold standards, are preferred.
[0239] By predetermined level of predictability it is meant that the method provides an acceptable level of clinical or diagnostic accuracy. Using such statistics, an “acceptable degree of diagnostic accuracy”, is herein defined as a test or assay (such as the test used in some aspects of the invention for determining the clinically significant presence of determinants, which thereby indicates the presence a colon cell proliferative disorder) in which the AUC (area under the ROC curve for the test or assay) is at least 0.60, desirably at least 0.65, more desirably at least 0.70, preferably at least 0.75, more preferably at least 0.80, and most preferably at least 0.85.
[0240] By a “very high degree of diagnostic accuracy”, it is meant a test or assay in which the AUC (area under the ROC curve for the test or assay) is at least 0.75, 0.80, desirably at least 0.85, more desirably at least 0.875, preferably at least 0.90, more preferably at least 0.925, and most preferably at least 0.95.
[0241] Alternatively, the methods predict the presence or absence of a colon cell proliferative disorder or response to therapy with at least 75% total accuracy, more preferably 80%, 85%, 90%, 95%, 97%, 98%, 99% or greater total accuracy.
[0242] Alternatively, the methods predict the presence or absence of colon cell proliferative disorder or response to therapy with an MCC larger than 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or 1.0. In general, alternative methods of determining diagnostic accuracy are commonly used for continuous measures, when a disease category has not yet been clearly defined by the relevant medical societies and practice of medicine, where thresholds for therapeutic use are not yet established, or where there is no existing gold standard for diagnosis of the pre-disease. For continuous measures of risk, measures of diagnostic accuracy for a calculated index are typically based on curve fit and calibration between the predicted continuous value and the actual observed values (or a historical index calculated value) and utilize measures such as R squared, Hosmer- Lemeshow P-value statistics and confidence intervals. It is not unusual for predicted values using such algorithms to be reported including a confidence interval (usually 90% or 95% CI) based on a historical observed cohort’s predictions, as in the test for risk of future breast cancer recurrence commercialized by Genomic Health, Inc. (Redwood City, California).
[0243] In general, by defining the degree of diagnostic accuracy, i.e., cut points on a ROC curve, defining an acceptable AUC value, and determining the acceptable ranges in relative concentration of what constitutes an effective amount of the determinants of the invention allows for one of skill in the art to use the determinants to identify, diagnose, or prognose subjects with a pre-determined level of predictability and performance.
[0244] Furthermore, other unlisted biomarkers will be very highly correlated with the determinants (for the purpose of this application, any two variables will be considered to be “very highly correlated” when they have a Coefficient of Determination (A2) of 0.5 or greater). Some aspects of the present invention encompass such functional and statistical equivalents to the aforementioned determinants. Furthermore, the statistical utility of such additional determinants is substantially dependent on the cross-correlation between multiple biomarkers and any new biomarkers will often be required to operate within a panel in order to elaborate the meaning of the underlying biology.
[0245] According to still another embodiment of the invention there if provided a method of treating IBD in a subject in need thereof comprising administering to the subject a therapeutically effective amount of at least one bacteria selected from the group consisting of Alistipes putredinis, Dialister succinatiphilus, Oscillibacter sp. ER4, Romboutsia sp. DR1, Gemmiger formicilis, Faecalibacterium prausnitzii, Fusicatenibacter saccharivorans, Ruminococcus callidus and Bacteroides coprocola, thereby treating the IBD.
[0246] The therapeutic bacteria described herein (also referred to as probiotic microorganisms) may be provided in any suitable form, for example in a powdered dry form. In addition, the probiotic microorganism may have undergone processing in order for it to increase its survival. For example, the microorganism may be coated or encapsulated in a polysaccharide, fat, starch, protein or in a sugar matrix. Standard encapsulation techniques known in the art can be used. For example, techniques discussed in U.S. Pat. No. 6,190,591, which is hereby incorporated by reference in its entirety, may be used.
[0247] According to a particular embodiment, the probiotic microorganism composition is formulated in a food product, functional food or nutraceutical.
[0248] In some embodiments, a food product, functional food or nutraceutical is or comprises a dairy product. In some embodiments, a dairy product is or comprises a yogurt product. In some embodiments, a dairy product is or comprises a milk product. In some embodiments, a dairy product is or comprises a cheese product. In some embodiments, a food product, functional food or nutraceutical is or comprises a juice or other product derived from fruit. In some embodiments, a food product, functional food or nutraceutical is or comprises a product derived from vegetables. In some embodiments, a food product, functional food or nutraceutical is or comprises a grain product, including but not limited to cereal, crackers, bread, and / or oatmeal. In some embodiments, a food product, functional food or nutraceutical is or comprises a rice product. In some embodiments, a food product, functional food or nutraceutical is or comprises a meat product.
[0249] Prior to administration, the subject may be pretreated with an agent which reduces the number of naturally occurring microbes in the microbiome (e.g. by antibiotic treatment). According to a particular embodiment, the treatment significantly eliminates the naturally occurring gut microflora by at least 20 %, 30 % 40 %, 50 %, 60 %, 70 %, 80 % or even 90 %.
[0250] In some embodiments, administering comprises any means of administering an effective (e.g., therapeutically effective) or otherwise desirable amount of a composition to an individual. In some embodiments, administering a composition comprises administration by any route, including for example parenteral and non-parenteral routes of administration. Parenteral routes include, e.g., intraarterial, intracerebroventricular, intracranial, intramuscular, intraperitoneal, intrapleural, intraportal, intraspinal, intrathecal, intravenous, subcutaneous, or other routes of injection. Non-parenteral routes include, e.g., buccal, nasal, ocular, oral, pulmonary, rectal, transdermal, or vaginal. Administration may also be by continuous infusion, local administration, sustained release from implants (gels, membranes or the like), and / or intravenous injection.
[0251] Particular doses or amounts to be administered in accordance with the present invention may vary, for example, depending on the nature and / or extent of the desired outcome, on particulars of route and / or timing of administration, and / or on one or more characteristics (e.g., weight, age, personal history, genetic characteristic, lifestyle parameter, severity of IBD and / or level of risk of IBD, etc., or combinations thereof). Such doses or amounts can be determined by those of ordinary skill. In some embodiments, an appropriate dose or amount is determined in accordance with standard clinical techniques. Alternatively or additionally, in some embodiments, an appropriate dose or amount is determined through use of one or more in vitro or in vivo assays to help identify desirable or optimal dosage ranges or amounts to be administered.
[0252] In some particular embodiments, appropriate doses or amounts to be administered may be extrapolated from dose-response curves derived from in vitro or animal model test systems. The effective dose or amount to be administered for a particular individual can be varied (e.g., increased or decreased) over time, depending on the needs of the individual. In some embodiments, where bacteria are administered, an appropriate dosage comprises at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more bacterial cells. In some embodiments, the present invention encompasses the recognition that greater benefit may be achieved by providing numbers of bacterial cells greater than about 1000 or more (e.g., than about 1500, 2000, 2500, 3000, 35000, 4000, 4500, 5000, 5500, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 75,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, IxlO6, 2xl06, 3 xlO6, 4 xlO6, 5 xlO6, 6 xlO6, 7 xlO6, 8 xlO6, 9 xlO6, 1 xlO7, 1 xlO8, 1 xlO9, 1 xlO10, 1 xlO11, 1 xlO12, 1 xlO13or more bacteria.
[0253] In some embodiments, the term “about” refers to ± 20% or ± 10%.
[0254] The terms "comprises", "comprising", "includes", "including", “having” and their conjugates mean "including but not limited to".
[0255] The term “consisting of’ means “including and limited to”.
[0256] The term "consisting essentially of" means that the composition, method or structure may include additional ingredients, steps and / or parts, but only if the additional ingredients, steps and / or parts do not materially alter the basic and novel characteristics of the claimed composition, method or structure.
[0257] As used herein, the singular form "a", "an" and "the" include plural references unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" may include a plurality of compounds, including mixtures thereof.
[0258] Throughout this application, various embodiments of this invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0259] Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integral) within the indicated range. The phrases “ranging / ranges between” a first indicate number and a second indicate number and “ranging / ranges from” a first indicate number “to” a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numerals therebetween.
[0260] As used herein the term "method" refers to manners, means, techniques and procedures for accomplishing a given task including, but not limited to, those manners, means, techniques and procedures either known to, or readily developed from known manners, means, techniques and procedures by practitioners of the chemical, pharmacological, biological, biochemical and medical arts.
[0261] As used herein, the term “treating” includes abrogating, substantially inhibiting, slowing or reversing the progression of a condition, substantially ameliorating clinical or aesthetical symptoms of a condition or substantially preventing the appearance of clinical or aesthetical symptoms of a condition.
[0262] The embodiments described herein may find applicability in any computing or processing environment. Various embodiments may be implemented in hardware, software, or a combination of hardware and software. Embodiments may be implemented using one or more computer programs on programmable computers and / or on a computing or communication device connected to a network, that each includes a processor, and a storage medium (e.g., a remote storage server) readable by the processor (including volatile and non-volatile memory and / or storage elements) without limitation.
[0263] Any software programs used in any embodiments may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system and / or user devices. Also, the programs can be implemented in assembly or machine language. The language may be a compiled, interpreted, or scripted language without limitation. Computer programs may be stored on any suitable storage medium or device (e.g., CD-ROM, hard disk, or magnetic diskette, optical or other drive) that is readable by a general or special purpose programmable machine for configuring and operating the computer when the storage medium or device is read by the computer to execute software instructions. Computer resources, such as processing and storage, may be configured in a single device or spread among several devices in the same location or distributed in remote locations, interconnected by a suitable network, as is well known in the art, without limitation.
[0264] Databases used in various embodiments may be any type of information storage and access system commercially available, publicly available or built for the purposes as described in this disclosure. Databases may be stored on a single server or distributed in a cloud-operating environment. Databases may be locally stored or connected via a network to user devices, which may include any form of computing device including mobile devices.
[0265] The databases used in various embodiments may be accessed by a variety of user interfaces either commercially available or built for the purposes described in this disclosure. The user interfaces may be built for use on any computing device as is known in the art, together with security measures to assure only authorized users may search for, view, edit or create data records.
[0266] The embodiments described herein may also use machine intelligence without limitation, specifically to include aspects of artificial intelligence, particularly for the correlation functions between candidate microbes (microbes yet to be characterized with a recipe) and available recipes together with reports on extent of success of the recipe in fostering growth. However, machine intelligence may be used for any aspect of this disclosure.
[0267] It will be appreciated that an exemplary computing system is merely illustrative of a computing environment in which the herein described systems and methods may operate, and therefore does not limit the implementation of the described systems and methods in possibly different computing environments that may have different components and configurations. In other words, the inventive concepts described herein may be implemented in various computing environments using various components and configurations. Moreover, those of skill in the art will appreciate that the herein described apparatuses, engines, devices, systems and methods are susceptible to various modifications and alternative constructions. There is no intention to limit the scope of the invention to the specific constructions described herein. Also, it should be apparent that the embodiments disclosed herein are not limited to a specific architecture or programming language. Rather, the systems and methods described herein are intended to cover all modifications, alternative constructions, and equivalents falling within the scope and spirit of the disclosure, any appended claims, and any equivalents thereto.
[0268] While several inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means for performing the function or obtaining the results and advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment of the invention. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.
[0269] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.
[0270] EXAMPLES
[0271] Reference is now made to the following examples, which together with the above descriptions illustrate some embodiments of the invention in a non limiting fashion.
[0272] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques.
[0273] EXAMPLE 1
[0274] MATERIALS AND METHODS
[0275] Assessing the induction of lac operon in vitro. E. coli (BL21(DE3), Thermo Scientific, EC0114) was grown in LB (1% glucose, 50ug / mL Kanamycin) overnight at 37 °C 180 rounds per minute (RPM). Induction of the lac operon was achieved by supplementing the culture with Isopropyl B-D-l -thiogalactopyranoside (IPTG, 2.5pm) and incubating for three hours at 37 °C 180 RPM. Control cultures were supplemented with vehicle (sterile PBS- / ). Lastly, cultures were centrifuged for 8 minutes at 3,200g, the supernatants were discarded, and pellets were frozen at - 80°C until their processing for shotgun metagenomic sequencing and liquid-chromatography tandem mass spectrometry (LC-MS / MS). This experiment included two biological replicates (using two different CFUs) with three technical replicates per each biological replicate.
[0276] Proteomic profiling of E. coli in different growth conditions in vitro. A single CFU of E. coli K12 (ATCC 700926) was used to prepare a primary culture in BHI, (brain heart infusion) broth and incubated overnight at 37 °C. This culture was diluted 1:200 (vol: vol) into secondary cultures using either BHI, LB (Luria Broth), TSB (Tryptic Soy Broth), or MHB (Mueller Hinton broth) and incubated overnight at 37 °C in either aerobic or anaerobic conditions using an anaerobic chamber (Coy Laboratory Products, 75% N2, 20% CO2, 5% H2). All growth media were purchased from BD. For anaerobic conditions, all culture broth media were introduced to the anaerobic chamber at least 24 hours prior to use, to allow for proper oxygen reduction. Lastly, cultures were centrifuged for 8 minutes at 3,200g, supernatant was discarded, and pellets were frozen at -80°C until their processing for shotgun metagenomic sequencing and LC- MS / MS. This experiment was repeated independently three times to generate three biological replicates.
[0277] Proteomic profiling of bacterial cultures for taxonomic annotation in vitro. Secondary cultures of Staphylococcus aureus (ATCC 12600), Klebsiella oxytoca (ATCC 13182), Micrococcus luteus (DSM 20030), E. coli K12 (ATCC 700926), Enterococcus faecalis (ATCC 29212), and Bacillus subtilis strain 168 (ATCC 23857) were grown aerobically at 37 °C. The same broth (MHB) was used for culturing all six species. Optical density (OD) was equalized to 1 in all cultures, and different isovolumic culture mixtures were created. Lastly, cultures were centrifuged for 8 minutes at 3,200g, supernatant was discarded, and pellets were frozen at -80 °C until their processing for shotgun metagenomic sequencing and LC-MS / MS. This experiment was repeated twice to generate two biological replicates. The database used as a reference to search and quantify bacterial proteins in the LC-MS / MS output included the above- mentioned six species and other 14 bacterial species commonly found in the gut microbiome: Klebsiella pneumoniae, Bacteroides thetaiotaomicron, Bacteroides uniformis, Bifidobacterium adolescentis, Faecalibacterium prausnitzii, Eubacterium rectale, Ruminococcus torques, Veillonella atypica, Prevotella copri, Akkermansia muciniphila, Roseburia faecis, Coprococcus comes, Bacteroides dorei, Lactococcus lactis. These 14 non-present taxa were included in the database to simulate real-world conditions and enable the possibility of erroneous identification of proteins not truly present in the samples.
[0278] Mice. Specific-pathogen-free (SPF) C57BL / 6 mice were purchased from Harlan Envigo and acclimatized to the animal facility environment for two weeks before commencing experiments. C57BL / 6 germ-free (GF) mice were born and kept in the GF facility of the Weizmann Institute of Science and routinely monitored for sterility. All mice were maintained under a strict 12-h light-dark cycle, with lights on at 6 a.m. and off at 6 p.m. Eight-week-old male wild-type mice were used in all experiments unless otherwise specified.
[0279] Segmental filamentous bacteria (SFB) mono-colonization in GF mice. Fresh frozen cecal contents from mSFB -monocolonized or rSFB -monocolonized mice. GF mice were inoculated by oral gavage with 200pl of bacterial solution of mSFB (n=10) or rSFB (n=9) cultured overnight, containing approximately 109colony-forming units (CFU). Littermates GF mice orally inoculated with vehicle (sterile PBS / _, n=9) were used as negative controls. After two weeks of colonization, mice were euthanized, and terminal ileal mucosal samples were harvested and divided into three aliquots. Fresh stool samples were collected and snap-frozen at -
[0280] 80 °C until their processing for metagenomic sequencing, LC-MS / MS, and bulk-RNA sequencing.
[0281] Interleukin- 18 (IL- 18) supplementation in GF mice. GF mice were intraperitoneally injected with recombinant IL-18 (MBL, Cat#B004-2, powder dissolved in sterile PBS / _, n=10) for five days at a dose of Ipg / mouse / inj ection, twice daily. Littermate GF mice injected with vehicle (sterile PBS / _, n=10) only were used as negative controls. Exogenous microbial contamination was minimized by housing mice in sterilized iso-cages, and handling mice and IL- 18 solution in a designated germ- free biosafety cabinet. After five days distal colonic mucosal samples were harvested and snap frozen at 80°C until their processing for metagenomic sequencing and LC-MS / MS.
[0282] Citrobacter rodentium infection in SPF mice. A kanamycin-resistant C. rodentium strain, DBS 100 (ICC180), was used for infection. SPF mice (n=12) were infected by oral gavage with 200pl of bacterial solution cultured overnight, containing approximately 109colonyforming units (CFU). Both the distal colonic mucosal and stool samples were collected from mice without infection (day 0), during the peak of infection (day 7), and at the infection recovery phase (day 14) and snap frozen at 80°C until their processing for metagenomic sequencing and LC-MS / MS. The pathogen load was quantified by CFU counting of stool CFU. In detail, stool samples were weighed, homogenized in sterile PBS / _, serially diluted in PBS / _, and plated on LB kanamycin plates. After incubating overnight at 37 °C, bacterial CFUs were counted and normalized to the stool weight.
[0283] Proteomic profiling of E. coli in different growth conditions in vivo. GF mice were given either autoclaved NC (n=4) or a sterile HFD (D18110605ii, 60% kcal from fat, Research Diets, n=4) for two days prior to inoculation. Mono-colonization was performed by oral gavage of 200pl culture broth containing 108CFUs of E. coli K12 (ATCC 700926). Mice were euthanized five days after colonization and small intestinal and colonic luminal samples were obtained and snap frozen at -80°C until their processing for shotgun metagenomic sequencing and LC-MS / MS.
[0284] Proteomic profiling of luminal samples across the murine gastrointestinal tract under different dietary conditions. SPF mice were fed with either NC (n=5) or HFD (n=5) for one week. Subsequently, mice were euthanized by CO2 asphyxiation, and six different segments of the digestive tract were exposed and harvested: stomach, duodenum, jejunum, ileum, cecum and colon. For each section, the luminal contents were collected snap-frozen at -80 °C until their processing for metagenomic sequencing and LC-MS / MS. Dextran sulfate sodium (DSS)-induced colitis. SPF mice (n=l 1) were treated with 2% (weight / volume) DSS (molecular weight, 36,000-50,000 Da; MP Biomedicals, CAT. no 9011- 18-1) added to the drinking water for seven days followed by the resumption of regular water. Disease severity was monitored by mice weighing. Fresh stool samples were collected before DSS treatment (day 0, baseline), at the early inflammatory phase (day 3, day 5), late inflammatory phase (day 16), and the end of the weight recovery phase (day 31) and snap-frozen at -80 °C until their processing for metagenomic sequencing and LC-MS / MS.
[0285] Transfer of non-treated SPF microbiome to GF mice with or without colonic inflammation. GF recipient mice (n=4) were treated with 2.5% (weight / volume) DSS (molecular weight, 36,000-50,000 Da; MP Biomedicals, CAT 9011-18-1) added to the drinking water for three or five days. Another group of littermate GF recipient mice (n=3) was given regular drinking water as a control. A single sample of pooled, filtered, and homogenized stool from treatment-naive donor SPF mice (n=5) was transferred into these auto-inflamed or control GF mice. Fresh stool pellets were collected from mice donors, and quickly transferred to an anaerobic chamber for further processing. Within the anaerobic chamber pellets were resuspended in sterile PBS / _at a ratio of 5 pellets per ImL of fluid, homogenized using a vortex, and spun down before supernatant was collected for gavage. Fresh stool samples were collected from recipient mice at baseline, three days after DSS treatment, and two or four days after microbiome colonization. Fresh stool samples were collected and snap-frozen at -80 °C until their processing for metagenomic sequencing and LC-MS / MS.
[0286] Transfer of SPF microbiome from inflamed or non-inflamed intestine to nontreated GF mice. Similarly, SPF donor mice (n=4) were treated with 2.5% (weight / volume) DSS added to the drinking water for three or five days. A group of littermate SPF donor mice (n=3) was given regular drinking water as a control. GF recipient mice were colonized with microbiomes from either DSS-treated (n=4) or control SPF mice (n=3). Stool microbiome transfer was performed as described above. Fresh stool samples were collected at baseline and three days after DSS treatment from donor mice, and two days after microbiome colonization in recipient mice. Fresh stool samples were collected and snap-frozen at -80 °C until their processing for metagenomic sequencing and LC-MS / MS.
[0287] Mice exclusive enteral diet. SPF mice were maintained for two weeks on an amino acidbased diet (Teklad 200727, n=8) or a control diet containing whole casein (Teklad 200726, n=7). Fresh stool samples were collected and snap-frozen at -80 °C until their processing for metagenomic sequencing and LC-MS / MS. Analysis of rodent’s diets. Pellets from NC (Teklad 2018), HFD (Research diets D12492i and D18110605ii), high-protein diet (Teklad 90016) and low-protein diet (Teklad 91352) were processed for LC-MS / MS, directly homogenized in lysis buffer containing 5% SDS in 50 mM Tris-HCl pH 7.4.
[0288] RNA purification. Total RNA from intestinal mucosal samples (terminal ileum, distal colon) was purified using RNAeasy kit (QIAGEN, 74104) according to the manufacturer’s instructions.
[0289] RNA-sequencing. Ribosomal RNA was selectively depleted by RNaseH (New England Biolabs, M0297) according to a modified version of a published method105. Specifically, a pool of 50 bp DNA oligos (25 nM, IDT) that is complementary to murine rRNA18S and 28S, was resuspended in 75 ml of 10 mM Tris pH 8.0. Total RNA (1000 ng in 10 ml H2O) was mixed with an equal amount of rRNA oligo pool. The RNA was added to 2 pl diluted oligo pool and 3 pl 5x rRNA hybridization buffer (0.5M Tris-HCl, 1 M NaCl, titrated with HC1 to pH 7.4). Samples were incubated at 95 °C for 2 min, then the temperature was slowly reduced (0.1 C / s) to 37 °C. RNaseH enzyme mix (2pl of 10U RNaseH, 2 pL lOx RNaseH buffer, ImL H2O, total 5 pl mix) was prepared 5 min before the end of the hybridization and preheated to 37 °C. The enzyme mix was added to the samples when they reached 37 °C and the mix was incubated at this temperature for 30 min. Samples were purified with 2.2x SPRI beads (Ampure XP, Beckmann Coulter) according to the manufacturer’s instructions. Residual oligos were removed with DNase treatment (ThermoFisher Scientific, AM2238) by incubation with 5 pl DNase reaction mix (1 pl Turbo DNase, 2.5 pl Turbo DNase lOx buffer, 1.5 pl H2O) that was incubated at 37 °C for 30 min. Samples were again purified with 2.2x SPRI beads and suspended in 3.6 pl priming mix (0.3 ml random primers of New England Biolab, E7420, 3.3 pl H2O). Samples were subsequently primed at 65 °C for 5 min. Samples were then transferred to ice and 2 pl of the first strand mix was added (1 pl 5x first strand buffer, NEB E7420; 0.125 pl RNase inhibitor, NEB E7420; 0.25 pl ProtoScript II reverse transcriptase, NEB E7420; and 0.625 pl of 0.2 pg / ml Actinomycin D, Sigma, A1410). The first strand synthesis and all subsequent library preparation steps were performed using NEBNext Ultra Directional RNA Library Prep Kit for Illumina (NEB, E7420) according to the manufacturer’s instructions (all reaction volumes reduced to a quarter). Libraries that passed quality control were loaded with a concentration of 2 pM on 75- cycle high output flow cells (Illumina, FC-404-2005) and sequenced on a NextSeq 500 (Illumina) with the following cycle distribution: 8 bp index 1, 8 bp index 2, 75 bp read 1.
[0290] RNA-sequencing analysis. We used Illumina’s bcl2fastq script to convert the raw files to fastq files. Fastq files were quality filtered using fastp (v0.20.0) with default parameter. We aligned the reads to the murine reference transcriptome with Gencode annotation (GRCm38.p6, Release M24) and quantified gene expression using STAR (v2.7.3a). Differential expression analysis was carried out with the DESeq2 package with a model for paired samples ( ~subject+ condition). Gene Ontology analysis was carried out with g:Profiler with default settings.
[0291] Microbiome DNA purification. Intestinal mucosal (ileum, distal colon), luminal or fecal DNA was extracted and purified using a Purelink Microbiome DNA purification kit (Invitrogen).
[0292] Shotgun metagenomics sequencing and processing. For shotgun sequencing, Illumina libraries were prepared using a Nextera DNA Sample Prep kit (Illumina, 20034198), according to the manufacturer’s protocol, and sequenced on the Illumina Novaseq platform with a read length of 122bp. Illumina’s bcl2fastq script was implemented to generate the fastq files. QC trimming was carried out using Fastp, and host sequences were removed with Bowtie2 (2.4.1) and human genome reference hg37. The cleared fastq files were subsampled using Seqtk (1.3- rl l4). We carried out the taxonomic assignment of bacterial DNA relying on exact the alignment of k-mers with Kraken2 (2.0.9) against the Genome Taxonomy Database (https: / / gkb.ecogenomk.org / , version 95). To improve the accuracy of species-level classification, we applied Bayesian re-estimation of bacterial abundance with Bracken (v2.5.3). Alpha and beta diversity, based on Bray-Curtis dissimilarities, were analyzed using Vegan R- package (https: / / CRAN.R-project.org / package=vegan). Differential species abundance analysis was performed using DESeq2. Functional annotation was implemented using protein alignment with DIAMOND, thereby, only the first hit was considered, and an e value <0.0001 was accepted. Principal coordinate analysis was performed based on Bray-Curtis dissimilarity using the ape R package. Permutational multivariate analysis of variance (PERMANOVA), on Bray- Curtis dissimilarities, was used to test for differences in the taxonomical and functional profiles.
[0293] Mass spectrometry-based proteomics sample preparation. Samples were subjected to in-solution tryptic digestion using suspension trapping (S-trap) as previously described115. Briefly, mucosal tissues were first homogenized in cold sterile PBS- / and then supplemented with equal volume of 10% SDS in 50 mM Tris-HCl pH 7.4 (5% SDS final concentration). Mouse and human stool samples were directly homogenized in lysis buffer containing 5% SDS in 50 mM Tris-HCl pH 7.4. Lysates were then incubated at 96 °C for 5 min, followed by six cycles of 30 sec of sonication (Bioruptor Pico, Diagenode, USA). Protein concentration was measured using a BCA assay (Thermo Scientific, USA). An amount of 50 ug total protein was reduced with 5mM dithiothreitol and alkylated with lOmM iodoacetamide in the dark. Each sample was loaded onto S-Trap microcolumns (Protifi, USA) according to the manufacturer’s instructions. After loading, samples were washed with 90: 10% methanol / 50 mM ammonium bicarbonate. Samples were then digested with trypsin (1:50 trypsin / protein) for 1.5 h at 47 °C. The digested peptides were eluted using 50 mM ammonium bicarbonate. Trypsin was added to this fraction and incubated overnight at 37 °C. Two more elutions were made using 0.2% formic acid and 0.2% formic acid in 50% acetonitrile. The three elutions were pooled together and vacuum- centrifuged to dryness. Stool samples were subjected to an additional cleaning step using solidphase extraction (Oasis HBL, Waters, MA, USA) according to manufacturer instructions. For samples to be analyzed by parallel reaction monitoring (PRM), equal amount of peptide mix containing 21 stable isotope-labeled (SIL) synthetic peptides was spiked into the digests. Samples were kept at -80°C until further analysis.
[0294] Liquid chromatography. LC / MS-MS grade solvents were used for all chromatographic steps. Dry-digested samples were dissolved in 97:3% PhO / acetonitrile + 0.1% formic acid. Each sample was loaded and analyzed using split-less nano-Ultra Performance Liquid Chromatography (nanoUPLC, 10 kpsi nanoAcquity; Waters, Milford, MA, USA). The mobile phase was: A) H2O + 0.1% formic acid and B) acetonitrile + 0.1% formic acid. Desalting of the samples was performed online using a Symmetry C18 reversed-phase trapping column (180 pm internal diameter, 20 mm length, 5 pm particle size; Waters). The peptides were then separated using an HSS T3 nano-column (75 pm internal diameter, 250 mm length, 1.8 pm particle size; Waters) at 0.3 5pL / min. For label-free quantitative (LFQ) proteomic analysis, peptides were eluted from the column into the mass spectrometer using the following gradient: 4% to 25% B in 155 min, 25% to 90% B in 5 min, maintained at 90% for 5 min, and then back to initial conditions. For PRM analysis the gradient was: 4% to 30%B in 97 min, 30% to 90%B in 5 min, maintained at 90% for 5 min and then back to initial conditions.
[0295] Mass Spectrometry. The nanoUPLC was coupled online through a nanoESI emitter (10pm tip; New Objective; Wobum, MA, USA) to a quadrupole orbitrap mass spectrometer (Q Exactive HFX or Exploris 480, Thermo Scientific) using a Flexion nanospray apparatus (Proxeon). For LFQ analysis, data were acquired in data-dependent acquisition (DDA) mode, using a ToplO method. MSI resolution was set to 120,000 (at 200 m / z), mass range of 375-1650 m / z, AGC of 3e6, and maximum injection time was set to 60msec. MS2 resolution was set to 15,000, quadrupole isolation 1.7 m / z, AGC of le5, dynamic exclusion of 45 sec, and maximum injection time of 60msec. For PRM analysis, data was acquired in PRM mode with scheduled monitoring of 148 native (unlabeled) peptides and 21 SIL synthetic peptides, corresponding to 68 proteins (supplementary table 2). MSI resolution was set to 120,000 (at 200m / z), mass range of 375-1650m / z, AGC of le6 and maximum injection time was set to 60msec. MS2 resolution was set to 30,000, quadrupole isolation 1.7m / z, AGC of 2e5, and maximum injection time of 100msec.
[0296] Targeted proteomics. In complementing the discovery approach and its inherent technical limitations, antimicrobial peptides (AMPs) and AMP-candidate list was utilized in a targeted AMP-focused proteomic approach (‘targeted approach’) enabling increased sensitivity and high confidence in AMP identification and quantification in the complex stool and GI mucosal sample matrix. Specifically, 148 tryptic peptides corresponding to 68 known or putative AMPs were measured in a custom designed parallel reaction monitoring (PRM) pipeline. Stable isotope labeled (SIL) peptides were synthesized (JPT, Berlin, Germany) for a subset of 21 of these peptides, spiked into the stool samples, and monitored along with their unlabeled native counterparts in the same experiment, thus serving as internal standards for validation of the respective AMP identity.
[0297] Constructing a refined catalog of bacterial proteins using per-experiment metagenomic data. To assemble a cohort- specific proteomic database the microbiome samples from the entire cohort are sequenced by shotgun metagenomics as outlined above. Reads from all FASTQ files are then pooled together and classified to the species taxonomic level using Kraken2 (Wood, D. E., Lu, J. & Langmead, B. Improved metagenomic analysis with Kraken 2. Genome Biol. 20, 257 (2019)) and the reference Genome Taxonomy Database as outlined above. Based on taxonomical classification, the relative abundance of all detected taxa is calculated after additional unclassified read distribution using Bracken (Lu, J., Breitwieser, F. P., Thielen, P. & Salzberg, S. L. Bracken: estimating species abundance in metagenomics data. PeerJ Comput. Sci. 3, el04 (2017). After the completion of metagenomic sequencing a stepwise filtration process involving the following three criteria is applied (Figure 8A-B): (1) Genomes of species belonging to the 95thpercentile most abundant taxa are included in the database. (2) All reads are then realigned to the genomes of all remaining species using Minimap2 Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094-3100 (2018), and the genomes of bacterial taxa with more than 70% of covered genome are included in the database. (3) All metagenomic reads are then realigned using the DIAMOND algorithm Buchfink, B., Xie, C. & Huson, D. H. Fast and sensitive protein alignment using DIAMOND. Nat. Methods 12, 59-60 (2014) (blastx option) to the amino acid sequence of the proteins produced by the remaining species, and downloaded from the BioCyc Caspi, R. et al. BioCyc: A genomic and metabolic web portal with multiple omics analytical tools. FASEB J. 33, (2019) database. Proteins with more than 75% of covered amino acid sequences were selected and constituted the final metagenomic -based protein catalog. The aim of this filtration process is to include only protein coding sequences that are likely to be present in a given experiment.
[0298] Constructing a refined catalog of dietary protein markers. To generate this dietary catalog, 310 species present in diet were selected. Proteins produced by the genera containing these species were downloaded from UniProt (UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51, D523-D531 (2023). Using the MMseqs2 algorithm Steinegger, M. & Sdding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35, 1026-1028 (2017)., proteins with a minimum sequence identity of 0.6 and alignment coverage of 0.8.
[0299] Constructing the final refined catalog of proteins. The refined database is generated by combining the above-described bacterial and dietary catalogs with all host proteins (i.e., Homo sapiens or Mus musculus) present in UniProt database UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51, D523-D531 (2023).
[0300] Mass spectrometry-based proteomics analysis. The refined database is used as a reference to search and quantify host, bacterial and dietary proteins in the detected signal by LC-MS / MS using Metamorpheus (v0.0.319). Detected dietary proteins are subsequently verified by realigning the assigned unique peptides to the NCBI’s non-redundant protein sequences (July 2022) (blastp -word_size 2 -matrix PAM30 -threshold 11 -comp_based_stats 0 -outfmt 7 -evalue 200000 -gapopen 9 -gapextend 1 -num_alignments 100 -window_size 40). Only proteins where, at least, one unique peptide is assigned with the highest score to the expected dietary genus were selected. Log 10 transformation and quantile normalization were performed on mass spectrometry intensities using R packages DEP (v 1.18.0) and limma. Principal component analysis (PCA) was carried out based on log-transformed intensity. Permutational multivariate analysis of variance (PERMANOVA) was performed for pairwise comparisons on a distance matrix between group levels with corrections for multiple testing using the R package pairwiseAdonis (v0.4, httpsVgidmb.com / pmanmezarbizii / pakwiseAdoms). Differential protein expression was tested by applying protein-wise linear models combined with empirical Bayes statistics, from DEP package. For dietary proteomics analysis, PCA was computed with both protein and genus abundances, while differential abundance analysis was applied using Mann- Whitney test and Benjamini-Hochberg correction for multiple testing. Bacterial functional analysis. Metabolic pathway annotations for all bacterial proteins included in the refined database are downloaded from BioCyc. Functional enrichment analysis of differentially abundant bacterial proteins was performed with the enricherQ function (R package clusterProfiler, v4.4.4) (Wu, T. et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innovation (Camb) 2, 100141 (2021), applying hypergeometric test and Benjamini-Hochberg correction for multiple testing. Host functional analysis. Functional enrichment analysis of differential abundant host proteins was computed using the gost function of the gprofiler2 R package Raudvere, U. et al. g: Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic Acids Res. 47, W191-W198 (2019).
[0301] PRM data was processed with the Skyline algorithm (MacLean et al., 2010). Extracted ion chromatograms for all relevant peptide precursors and fragments were imported and manually curated for removing interfering signals and determining peak boundaries. The signal from SIL peptides was used to validate assignment of peak identity of the corresponding light (native) peptides. Peak assignment validation of native peptides without a corresponding SIL peptide was done by constructing a peptide identification library in Skyline using the above LFQ analysis results and comparing to PRM data. Peptide assignments of PRM fragment clusters which did not have a high confidence match to the relevant peptide in the library (dotp score<0.80) in any sample were filtered out. Total fragment area was exported from Skyline to Microsoft Excel, normalized to total ion current and log-transformed for statistical analysis.
[0302] In silico metagenomics. 40 common gut bacterial species were randomly selected, and both their genomes and proteomes were downloaded from Biocyc. Metagenomics reads (10 samples, 10 million reads per sample) were in silico simulated using InSilicoSeq124and the downloaded 40 genomic sequences (NovaSeq error models and log-normal distribution of the 40 bacterial species). Then, we constructed the refined bacterial metagenomic-based protein catalog using the above-described methodology (IPHOMED) or Metaphlan125for species abundance quantification (Figures 2A-O).
[0303] In silico proteomics. Proteins from the top 3 most abundant species in the gut microbiome were downloaded from Biocyc. Combinations of proteins from 1, 2 or 3 different bacterial species (Bacteroides vulgatus, Faecalibacterium prausnitzii and Bacteroides uniformis) were generated. Proteins were subsampled to 60% to simulate biological conditions. Spectral libraries were generated with Oktoberfest126(trypsin digestion, spectronaut output format and default parameters). Then, we created synthetic LC-MS / MS data using Synthedia127(one sample per group, centroided data). Synthetic mzML files were processed using the above-mentioned methodology, performing database search for the refined database or a generic database of bacterial proteins from the top 40 most abundant species in the gut microbiome (Figures 2A-O).
[0304] Comparison between IPHOMED and genome assembly-based methodology. Paired- end metagenomics data and LC-MS / MS were downloaded from ENA repository (PRJNA473429) and ProteomeXchange Consortium via the PRIDE (PXD020695). For in silico proteomics, synthetic paired-end metagenomics reads were used. Genome assembly was performed using Megahit (metagenome configuration, —presets meta-large). Proteins present in contigs were detected using eggNOG mapper (metagenome configuration, -itype metagenome). Taxonomical assignment of contigs (and their proteins) was performed using Kraken2. Database search was performed against the detected catalog of proteins using Metamorpheus.
[0305] Comparison between shotgun metagenomics and bacterial proteomics. Abundance quantification of bacterial genes (shotgun metagenomics) and bacterial proteins (metaproteomics) was carried out using the refined catalog of bacterial genes / proteins for both methodologies. To compare the distribution of gene / protein abundances, both quantifications were combined and transformed using quantile normalization. Alpha diversity was measured based on the Shannon Index.
[0306] Biomarkers of early colitis. Random forest analysis was performed using the randomForest function from the randomForest R package (v4.7, five-fold cross-validation iterations, ntree=1500) and applied to the mouse samples from day 0 (baseline) and day 3 (early colitis) from the DSS-induced colitis mouse model.
[0307] Quantification and statistical analysis. All statistical analyses were performed using R (3.6.3) and GraphPad Prism software (GraphPad Software, Inc.). For principal component analysis, we used permutational analysis of variance (PERMANOVA). For differential protein abundance analysis, we used protein- wise linear models combined with empirical Bayes statistics and FDR correction by Benjamini-Hochberg. For pathway enrichment analysis we used a hypergeometric test and FDR correction by Benjamini-Hochberg.
[0308] RESULTS
[0309] IPHOMED-mediated functional microbiome characterization features a high species-level specificity and sensitivity. IPHOMED is based on a refined and comprehensive, per-experiment generation of a protein reference catalog that enables the identification and quantification of the majority of proteins comprising the intestinal metaproteome, namely proteins originating from diet, host, and its bacterial microbiome, as overviewed in Figure 8A-H (Methods). Collectively, in stool samples obtained from cohorts of healthy treatment-naive wildtype (WT) C57BL / 6 mice (n=17) consuming normal chow (NC) and housed under specific pathogen-free (SPF) conditions, IPHOMED detected 630.9+13.8 host proteins, 2655.7+137.6 bacterial proteins, and 23.4+0.29 dietary proteins, respectively (Figure 81). Bacterial, host, and dietary proteins accounted for 79.83+0.64%, 19.44+0.62%, and 0.72+0.02% of the total detected stool proteome, respectively (Figure 8 J). We began our assessment of functional IPHOMED-mediated characterization of microbiome community structure by testing the sensitivity of our pipeline in a well-controlled in vitro context. We first examined whether IPHOMED is capable of detecting subtle changes in protein abundance that occur downstream of genetic code and are thus not amenable for detection by metagenomics. Such translational changes constitute a central motivation to adopt a proteomic rather than a genetic approach, especially since microorganisms rapidly adapt to changing environments and evade the immune system by expressing specific clusters of genes (i.e., operons58). To that end, we cultured Escherichia coli (E coli) that can express the mouse gene Mesencephalic Astrocyte Derived Neurotrophic Factor (Manf) under a Lac operon, followed by supplementation with the lac operon inducer isopropyl P-d-1 -thiogalactopyranoside (IPTG) or with vehicle control (phosphate buffered saline, PBS, Figure 1A). Expectedly, metagenomic sequencing of lac-operon-activated or non-activated E. coli detected no changes in MANF abundance at the gene-level. In contrast, IPHOMED identified a higher abundance of MANF protein upon IPTG supplementation (Figure IB), coupled with a significantly altered abundance of multiple other E. coli proteins, including purT and purC, which are involved in purine biosynthetic processes, thereby indicating a shift towards an altered rate of cellular division and energetic equilibrium (Figure 1C).
[0310] IPHOMED was also effective in detecting the dynamics of bacterial activity upon environmental modulation. In vitro culturing of a singular colony forming unit (CFU) E. coli strain in four different culture broths in aerobic and anaerobic conditions (Figure ID) demonstrated substantially different global IPHOMED profiles (Figure IE), including an expected enrichment of pathways associated with aerobic respiration (i.e., the Krebs Cycle) observed in all aerobically cultured conditions (Figure IF). Within the same oxygenic conditions, each one of the four-growth media activated distinct E. coli metabolic pathways, likely owing to a different nutrient profile. As an example, differential abundance analysis between tryptic soy broth (TSB) and Luria broth (LB) showed a significant change in the abundance of multiple proteins in TSB associated with tryptophan biosynthetic process, consistent with the higher abundance of tryptophan in TSB compared to LB media (Figure 1G). Together, these results highlight the effectiveness of IPHOMED in detecting functional changes that would otherwise remain undetected through shotgun metagenomic analysis.
[0311] To evaluate the taxonomic specificity of IPHOMED in correctly attributing proteins to particular bacterial species within a sample, we conducted a proteomic analysis of bacterial mixtures grown in the same broth medium in vitro. These were comprised of one to six bacterial strains in various combinations (Figure 1H), including E. coli, Micrococcus luteus, Bacillus subtilis, Staphylococcus aureus, Klebsiella oxytoca, and Enterococcus faecalis. To examine the potential for erroneous detection and quantification of proteins produced by species absent in a given combination, we constructed a metagenomic -based bacterial protein database comprising all the protein sequences encoded by these bacterial species, coupled with protein sequences from fourteen other species commonly found in the gut microbiome, such as Bacteroides thetaiotaomicron, Faecalibacterium prausnitzii, Akkermansia muciniphila and Eubacterium rectale (Methods). Indeed, IPHOMED analysis of individual cultures demonstrated a high level of specificity, with fewer than 10% of proteins being assigned to non-present species (Figure II). Moreover, a higher accuracy was observed when analyzing combinations containing multiple bacterial species. Importantly, IPHOMED identified a higher number of high- confidence proteins with enhanced taxonomical resolution, compared to proteomics analysis using nonredundant reference catalogs of proteins (Li et al Nat. Biotechnol. 32, 834-841 (2014)) (Figure 1J). These findings support a high IPHOMED accuracy in identifying bacterial proteins with species-level resolution.
[0312] We next compared the in vivo IPHOMED fecal microbiome signal to that generated by shotgun metagenomic sequencing. To this end, we generated a refined experimental catalog of bacterial protein sequences from fecal samples of treatment-naive SPF mice. Gene and protein abundances were then quantified by either metagenomics or IPHOMED, respectively, using the same refined database as a reference. Expectedly, considering that some bacterial genes are not constitutively expressed, alpha diversity (i.e., features richness within an individual sample) was greater in the metagenomic analysis compared with the IPHOMED analysis (Figure IK), with IPHOMED detecting only proteins corresponding to genes of relatively higher abundance. IPHOMED readouts (protein and pathway) were less variable among samples than corresponding genomic readouts (gene and pathway; Figure IL), in line with observations demonstrating that functional microbiome readouts feature less inter- individual variability compared with species-level community structure. In multiple bacterial taxa, the abundance distribution of genes and their corresponding protein was dissimilar to each other (Figure IM). Even at an individual bacterium level (for example, Lactobacillus murinus), multiple proteins featured significantly different abundances compared to their corresponding genes across samples. This observation suggested that only some bacterial proteins (whose abundances at gene and protein level are similar across individuals) can be inferred or 'predicted' from genomic data. Interestingly, such ‘unpredictable’ proteins were more likely to be extracellular (Figure IN), suggesting that genomic approaches may be especially suboptimal in investigating bacterial functions related to the secreted protein landscape. To validate our metagenomic-driven bacterial protein database refinement strategy, we next utilized an in silico approach. To this end, ten synthetic metagenomic communities, comprised of the 40 random bacterial species commonly found in the gut, were generated using their representative bacterial genomes and proteomes (Figure 2A, Methods). Bacterial abundance quantification in the synthetic metagenomics data was computed using either IPHOMED or MetaPhlAn60. IPHOMED correctly identified and included 100% of the included bacterial species in its refined database, while MetaPhlAn failed to include 10% of species in the artificial in silico "sample” (Figure 2B). After realigning metagenomic reads to the bacterial genomes of the detected species (Figure 2C), IPHOMED achieved a specificity of 100%, while maintaining 100% sensitivity (Figure 2D). Importantly, IPHOMED identified more than 90% of bacterial protein sequences after discarding protein sequences with low coverage (i.e., less than 75% coverage, Methods Figure 2E). In comparison, genome assembly methodology identified 14.5% more proteins than the simulated ones and with lower taxonomic resolution, compared to IPHOMED (Figure 2E-F).
[0313] To further exemplify how IPHOMED's database refinement strategy improves the analytical performance we generated three artificial proteomic datasets comprising one, two, and three bacterial species using their respective proteomes. To simulate real-world conditions, only 60% of the proteins were modeled (Figure 2G). We then searched these mock datasets using the metagenomically-driven database (IPHOMED) or a generic database containing proteins from 40 different bacterial species (representing a non-refined generic database, Methods). IPHOMED analysis revealed sensitivity and specificity values higher than 97% and 87%, respectively (Figure 2H-K). Collectively, this in silico simulation revealed that the inclusion of bacterial proteins, that lie beyond the potential proteomic framework of a given experiment (e.g., bacterial proteins that are unlikely to be truly present as their corresponding genes are absent from a given experimental metagenomic set), negatively affects sensitivity and specificity. In contrast, defining the bacterial reference genome and hence potential bacterial proteome by shotgun metagenomics sequencing on a per-experiment basis significantly enhances the analytical performance of bacterial proteomics.
[0314] To experimentally validate these findings using real-life data, we used IPHOMED to analyze stool samples of 11 SPF mice. In total, IPHOMED identified 5,446 bacterial proteins, 97% of which were successfully assigned to bacterial taxa at species-level resolution (Figure 2L). We then supplemented the refined experimental protein database with varying numbers of random bacterial protein sequences from species that were not detected by metagenomics (ranging from 1 to 5 million, Methods), and iteratively reanalyzed proteomic readouts from the same experiment using these artificially ‘inflated’ databases. Indeed, this random addition significantly impaired the output’s sensitivity and specificity, in a dose-dependent manner (Figure 2L) and was associated with the detection of hundreds of proteins spuriously assigned to bacterial species that were not identified by metagenomics (Figure 2M). Importantly, utilization of IPHOMED’s refined database (RDB) also resulted in significantly shorter computational processing time for the proteomic database search (Figure 2N). Moreover, in the context of complex microbial compositions, IPHOMED revealed a markedly higher taxonomical resolution of the detected bacterial proteins up to the species level compared to proteomics analysis using reference catalogs of proteins and genome-assembly-based methodologies (Figure 20). Functional assessment of host secretory responses to commensal and pathogenic colonization. We next utilized IPHOMED to characterize the in vivo host secretome landscape. We began by assessing the host protein- secretion response to mono-colonization with segmented filamentous bacteria (SFB), commensals that preferentially colonize (in a species- specific manner) the ileal mucosa of mice and induce the differentiation of Th 17 cells, production of secretory immunoglobulin A and secretion of antimicrobial peptides (AMPs). To assess the dependency of the host secretory proteomic response in commensal colonization, we monocolonized germ-free (GF) mice with SFB indigenous to either mice (mSFB) or rats (rSFB), or with a sterile vehicle (PBS) as a negative control, by a single oral gavage (Figure 3A). Of note, only mSFB are capable of colonizing the terminal ileum of mice through adherence to ileal epithelial cells, as confirmed by metagenomic sequencing. Indeed, only mice who were colonized by an attaching-capable commensal (mSFB) exhibited a global secretome shift using IPHOME pipeline (Figure 3B). These findings were validated by a targeted approach using parallel reaction monitoring (PRM) (targeting 68 known / putative AMPs, Methods). Interestingly, differentially abundant proteins between PBS- and mSFB-colonized mice were significantly associated with immune responses and antimicrobial activity (Figure 3C-D). This activation of host immune response as a consequence of commensal attachment was also recapitulated at the transcriptomic level, in RNA-sequencing of terminal ileum samples. Comparison between host transcriptomic and proteomic profiles revealed multiple host proteins that were differentially abundant at protein level, but not at the transcript level. For example, units of the antimicrobial protein calprotectin (S100A8 and S100A9), which are often used as a sole severity marker for intestinal inflammation, were found to be upregulated upon mSFB colonization by proteomics but not transcriptomics. Induction of a host intestinal secretory response can be also indirectly induced by commensals, through mucosal activation of innate immune signaling pathways. For example, gut commensal-mediated interleukin- 18 (IE- 18) production may lead to downstream expression of multiple AMPs. Indeed, intraperitoneal injection of recombinant IL- 18 to GF mice (Figure 3E, Methods) substantially impacted the colonic secreted host IPHOMED profile, as compared to vehicle control (PBS, Figure 3F-G), including enrichment of pathways related to defense response to symbionts, antigen presentation and antimicrobial activity (Figure 3G-H).
[0315] We next utilized IPHOMED to assess the host-microbiome interface shaping during GI pathogenic gastrointestinal infection, in which intestinal homeostasis is disrupted in a dynamic process that ranges from the initial pathogen invasion to infectious resolution by host elicitation of a protective mucosal immune response. To this aim, we orally infected SPF mice with 109CFUs of Citrobacter rodent ium, an attaching and effacing pathogen which simulates human enteropathogenic and enterohaemorrhagic E. coli (EHEC) infection (Methods). Colonic mucosal samples were collected at baseline (day 0), peak of infection (day 7) and the recovery phase (day 14), assessed by IPHOMED and validated by PRM (Methods, Figure 31). Upon infection, IPHOMED detected significant proteomic colonic mucosal dysbiosis that persisted throughout the duration of infection (Figure 3J). Further assessing the colonic bacterial proteome, C. rodentium, and particularly four of its bacterial proteins, Adk, AhpC, GAPDH, and GroES, were detected by IPHOMED in the colonic mucosa at day 7 post-infection and eliminated during infectious resolution at day 14 (Figure 3K-L). Assessing the host proteome, C. rodentium- induced colitis led to a profound host-secreted response, detected by IPHOMED, at day 7, that persisted after pathogen elimination from the colonic mucosa at day 14. Functional analysis of differentially abundant host proteins revealed significant defense responses to other organisms, cytokine production, and stress responses (Figure 3M). These pathogen-driven immune responses, including an expanded AMP signature, were validated by PRM, and included previously identified effector proteins such as neutrophil elastase (ELANE)63, and novel effector candidates such as Cystatin-B (CSTB). Collectively, IPHOMED detected hard-wired, regionspecific functional host proteomic programs that were specifically modulated by introduction of commensals or pathogens into the intestinal ecosystem.
[0316] Decoding of niche-specific dietary regulation of functional trans-kingdom interactions. We next assessed whether IPHOMED could sensitively and simultaneously assess host-microbiome responses to environmental perturbations inflicted on this trans-kingdom ecosystem, such as ones induced by nutritional alterations. Indeed, dietary interventions were previously suggested to impact host gut metabolic responses, but these were mainly studied as ‘cherry picking’ snapshots quantified at the gene expression level, while concomitant reciprocal changes induced on commensal function remained elusive. We began by assessing the in vivo, niche- specific, dietary-induced impact on host and microbial dynamics in mono-colonized mice. To this aim, we monocolonized either normal chow (NC)- or high fat diet (HFD)-fed germ-free (GF) mice with E. coli by oral gavage and collected small intestinal and colonic luminal samples after five days for IPHOMED analysis. Indeed, exposure of monocolonized mice to different diets induced substantial niche-dependent IPHOMED alterations in both the host and the colonizing bacterium. Diet- induced modulation of secreted intestinal host proteins was more notable in the small intestine compared with the colon, including an HFD-induced upregulation of multiple metabolic pathways related to amino acids, fatty acids and the Krebs Cycle. Likewise, E. coil's cellular function was influenced by diet, intriguingly also mostly in the jejunum and was hallmarked by HFD-induced upregulation of jejunal E. co / z-ascribed proteins involved in bacterial lipid metabolism (YecR, aes, dmsB). This harmonized biogeographydependent alteration of both host and E. coli function likely stems from commensal exposure and dietary absorption of most dietary nutrients occurring in the proximal small intestine.
[0317] We next assessed the niche- specific dietary impact on the host secretome landscape in a full microbiome colonization setting. To this aim, we applied IPHOMED on luminal samples collected from different gastrointestinal (GI) tract regions (stomach, duodenum, jejunum, ileum, cecum, and colon) of SPF mice fed with either NC (n=5) or HFD (n=5) for one week (Methods). The global host-secreted proteomic profiles were significantly different across different GI regions, and between the different diets (Figure 4A). The highest number of secreted host proteins were detected by IPHOMED in the duodenum and jejunum. However, differential abundance analysis revealed no significant host secretome changes in these regions between NC- and HFD-fed mice. In contrast, profound alterations in the secreted host protein IPHOMED landscape were observed in the stomach, as well as in the ileum, cecum, and colon upon HFD consumption (Figure 4B). These included a strong activation of host pathways related to lipid metabolism in HFD-fed mice (Figure 4C). Simultaneous bacterial IPHOMED profiles significantly differed in each GI region and upon changes in diet (Figure 4D-F). The highest numbers of bacterial proteins were detected in the stomach, cecum, and colon in both NC- and HFD-fed mice. Differential abundance analysis identified notable changes in microbiome activity in the stomach, cecum and colon. Of note, the IPHOMED-based stomach microbiome featured a greater similarity to that of the distal colon, likely reflecting murine coprophagy, as this phenomenon was not recapitulated in humans. Interestingly, IPHOMED, but not metagenomics, discovered multiple bacterial proteins to be significantly differentially abundant in the stomachs of either NC-fed or HFD-fed mice. For example, enrichment of multiple gastric bacterial carbohydrate-active enzymes (CAZymes) during NC consumption (Figure 4G), was consistent with higher concentrations of multiple complex carbohydrates in NC compared to HFD. In contrast, Lachnospiraceae bacterium C0E1- and Mucispirillum schaedleri- produced acetyl-CoA C-acetyltransferases (K00626), which participates in fatty acid degradation, were expanded in the stomachs of HFD-fed mice (Figure 4H-I). Interestingly, these results suggest that coprophagy in mice, and the associated gastric microbiome functions it uniquely entails, likely contributes to distinct gastric capacities of diet-dependent metabolic food degradation (Figure 4G-I). Collectively, IPHOMED recognized unique and highly harmonized transkingdom, host and microbiome functional changes occurring in response to dietary manipulation. Such coordinated niche-specific, reactive holobiont functions enable a high- resolution study of host-microbiome network contributions to physiology and disease risk.
[0318] Functional elucidation of host-commensal cross-regulation during intestinal autoinflammation. Given the above functional trans-kingdom host and microbiome insights, we next harnessed IPHOMED in disentangling more complex host-microbiome interactions likely occurring in multi-factorial diseases, such as IBD. To this end, we utilized the murine dextran sulfate sodium (DSS) colitis model, characterized by an acute development of colonic dysbiosis and dysregulation of mucosal immune responses, resembling some key features of human IBD. We induced acute colitis through oral administration of 2% DSS to SPF mice for seven days, followed by regular drinking water supplementation (Methods), and applied IPHOMED on fecal samples collected longitudinally before DSS exposure (day 0), at the early ‘pre-clinical’ inflammatory phases (days 3 and 5), advanced inflammatory and tissue destruction phase (day 16), and the recovery phase (day 31, Figure 5A). As early as 3 days after initiation of DSS supplementation, the fecal microbiome, as assessed by either metagenome or IPHOMED, has significantly changed compared to baseline and continued to change throughout disease course (Figure 5B-F). Already at this early inflammatory timepoint (day 3), IPHOMED was able to detect significant proteomic shifts in multiple bacterial species, including three species (Dorea sp. 5-2, Lachnospiraceae 28-4, and Lachnospiraceae 10-1 ), marked by increased presence of 92, 90 and 73 species-specific bacterial proteins, respectively (Figure 5D) whose taxonomical abundances remained unaltered by metagenomics. These early colitis-driven proteomic shifts attest to the advantage of IPHOMED in identifying early functional changes at a species-level resolution, which cannot be detected by conventional metagenomics analysis.
[0319] In parallel, IPHOMED quantified the secreted protein host response along disease course. These were already notable on day 3 and peaked at days 16 and 31 following induction of DSS- colitis (Figure 5G). This host response was slower to develop compared to dysbiosis, which peaked at day 3 (Figure 5H). Accordingly, the most discriminating proteins for the prediction of early colitis (day 3) were bacterial proteins produced by Butyricimonas virosa, Muricabulum intestinale, Akkermansia muciniphila and Parabacteroid.es goldsteinii (Figure 51). In contrast, at day 16, the host exhibited a pronounced inflammatory signal and antimicrobial activity (Figure 5J-K), which were confirmed by PRM. Interestingly, even at late recovery stage, host inflammatory and antimicrobial activity remained elevated (Figure 5J, 5G). These findings collectively demonstrate that alterations in functional microbiome composition and activity in this model precede the host’s inflammatory response.
[0320] Importantly, longitudinal samples enabled to identify novel biomarkers of severity in this IBD model, complementing conventional methods that rely on fecal calprotectin and circulating C -reactive protein (CRP). Using body weight change as an objective indicator of colitis severity, we uncovered host and bacterial proteins linked with an auto-inflammation flareup (Figure 5L). For example, serum amyloid P component (Apes), which increased with colitis severity, has been previously proposed as a serum marker for disease onset and severity. Similarly, the Rho- related GTP-binding protein (Rhoc), a Rho GTPase subfamily member, was positively correlated with body weight change. Of note, Rho GTPases play crucial roles in IBD-associated cellular events that are emerging as potential IBD treatment targets (Figure 5L). Importantly, multiple bacterial proteins, mainly from Lachnospiraceae and Clostridium species, were reduced during severity peaks. Building upon these concepts in human IBD may enable the identification of bacterial and host protein markers complementing calprotectin in assessing disease development and severity.
[0321] Intriguingly, IPHOMED could be harnessed to distinguish between the earliest changes induced by the inflamed gut mucosa on the indigenous microbiome, and the reciprocal changes induced by the dysbiotic auto-inflamed microbiome on the naive mucosa. Understanding these two closely intertwined early functional aberrations may enable to decode the initiating sequence set of events that occur before the occurrence of irreversible tissue damage and full-blown IBD, while uncovering new early microbial and host diagnostic and therapeutic targets. First, we dissected the impact of an inflamed axenic host intestinal environment on a naive microbiome (Figure 6A). To this end, we pre-treated GF recipient mice with DSS prior to colonizing them with stool microbiome from treatment-naive SPF mice donors (Figure 6A, Methods). Vehicle- treated littermate GF mice served as controls. DSS consumption elicited an inflammatory intestinal proteomic response in GF mice recipients within three days. This ‘primary’, microbiome-independent host response was characterized by numerous differentially abundant mouse proteins that were predominantly enriched in immune-related pathways. These autoinflamed or control GF recipient mice were then inoculated with a single stool microbiome obtained from treatment-naive SPF donor mice, enabling an IPHOMED assessment of the earliest impacts induced by an inflamed host gut on previously naive fecal microbial function (Figure 6A). Indeed, colonization of a naive foreign microbiome in a sterile inflamed gut environment resulted in marked bacterial functional shift developing as soon as two days following transfer, as compared to microbiome colonizing healthy non-inflamed colons (Figure 6B). Specifically, the abundance of 264 bacterial proteins changed significantly between the two groups (Figure 6C). For example, proteins produced by organisms like Dorea sp. 5-2 increased in abundance, while those from Parabacteroid.es goldsteinii and Parabacteroides bacterium YL27 exhibited a notable decrease in abundance (Figure 6D). Enrichment analysis identified several bacterial KEGG orthology (KO) terms, such as triosephosphate isomerase (TPI, KO18O3) and glyceraldehyde 3-phosphate dehydrogenase (K00134), that were significantly overrepresented among the differentially abundant proteins in inflamed GF mice colonized with naive SPF microbiome (Figure 6C, 6E). Interestingly, TPI has been suggested to counter, in some bacterial species, the host's inflammatory response. A similar naive microbiome transfer into GF recipients exposed to DSS for a longer time (5 days) induced comparable microbial shifts. Importantly, a longer naive commensal exposure (of 4 days) to a previously sterile host inflammatory milieu (Figure 6F) unveiled a substantial microbial inactivation (reduced overall protein abundance) (Figure 6G-H), mirroring similar bacterial inactivation noted in DSS-treated SPF mice (Figure 5D-F).
[0322] We next disentangled the reciprocal impacts induced by microbiomes obtained from colitis-induced mice on the naive GF host intestine (Figure 61). To this aim, we induced acute colitis in SPF mice via a three-day oral administration of 2.5% DSS in drinking water. A comparable group of vehicle-treated littermate SPF mice served as controls. DSS supplementation induced a significant shift in microbiome activity in DSS-treated donor mice compared with control donor mice, as reflected by altered IPHOMED-quantified fecal bacterial signatures. Notably, various bacterial proteins produced by Lachnospiraceae species were more prevalent in DSS-supplemented mice, while proteins from P. goldsteinii, P. bacterium YL27 or Eubacterium plexicaudatum were more abundant in control mice. Fecal microbiomes from DSS- treated or control SPF donor mice were then transferred into two groups of matched littermate GF mice, to assess the early impacts of the transferred inflamed microbial configurations on the naive recipient gut secretome (Figure 61). Indeed, the colitis-associated microbiomes induced a significant alteration in the naive recipient fecal host IPHOMED profile within two days of colonization (Figure 6J), manifesting as an enrichment of several host pathways, predominantly those associated with antimicrobial peptide production and immune responses (Figure 6K-L). These included the pro-inflammatory markers in IBD, S100A8 and S1009 (constituting subunits of calprotectin), together with other antimicrobial proteins such as Lyz2 and bone marrow proteoglycan (Prg2, Figure 6K). Prg2 features cytotoxic and microbicidal activities and was previously suggested as a potential biomarker for active human Crohn’s disease (CD). To corroborate these findings in an enhanced dysbiotic setting, we transferred fecal samples from mice exposed to DSS for a longer period (5 days) into naive germ-free mice. Interestingly, as was noted in DSS-treated SPF mice and the above reciprocal transfer experiments (Figure 6F- H), naive microbiomes exposed to an enhanced inflammatory exposure (in donor mice) featured a marked inactivation of multiple bacterial activities. Recipient naive GF mice colonized with such ‘inflamed’ microbiomes developed a similar shift of the host response to the one noted with a shorter gut mucosal exposure to an ‘inflamed’ microbiome (Figure 6I-L), characterized by activation of immune-related proteins and AMPs, such as PRG2, S100A8 and S100A9 (Figure 6K-L). Overall, IPHOMED enabled a simultaneous in-depth characterization of the shifting functional host and microbiome ‘reactome’ that characterizes different stages and clinical manifestations of intestinal auto-inflammation. It enabled to disentangle host and microbiome inter-modulation, while differentiating primary host auto-inflammatory changes from those induced by the reactive microbiome.
[0323] IPHOMED-mediated dietary and digestive quantification. The above simultaneous IPHOMED assessment of the murine host and bacterial protein landscapes enabled the quantification of approximately 3,000 proteins produced by the host or the microbiome. We hypothesized that some of the remaining unidentified proteomic signal may be comprised of dietary proteins originating in plants and animals, thereby presenting an opportunity to integrate into each IPHOMED analysis a quantitative assessment of dietary exposure and possibly other features of digestion and absorption. To this end, we used standardized dietary questionnaires72to compile a list of edible plant and animal taxa, totaling 310 widely consumed species encompassing vegetables, fruits, legumes, grains, seeds, meats, seafoods, and dairy products. Together, the species included in the database constitute widely consumed food items that are commonly represented in dietary practices. Protein markers for each dietary genus (213 genera) were selected for the dietary database as a reference for the search and quantification of dietary proteins. Importantly, after proteomics database search, only proteins featuring at least one peptide uniquely assigned to an expected dietary genus were included (Methods).
[0324] To test the sensitivity and specificity of the IPHOMED dietary assessment capacity, we processed chow pellets of four different rodent diets, in which the genera of all dietary proteins are comprehensively characterized: NC (Teklad 2018) is comprised of proteins from soybean (Glycine), com (Zea) and wheat (Triticum); HFD (Research diets D12492i) consists of a combination of proteins from soybean (Glycine) and casein; the remaining two diets, high and low protein diets (Teklad 90016 and 91352, respectively), almost exclusively contain casein. Importantly, IPHOMED accurately identified the genus abundance of all dietary elements present in these rodent diets (Figure 7A). To test IPHOMED-based dietary assessment in the in vivo setting, we fed SPF mice with either NC (Teklad 2018) or HFD (Research diets D12492i) for one week (n=4 per condition), a timeframe that falls short of inducing significant weight changes. Distinct small intestinal and colonic dietary IPHOMED signatures reflected dietary intake, including a significantly higher abundance of casein in HFD-fed mice, and an increased signal of corn (Zea) in NC-fed mice, as expected from the known composition of these two diets (Figure 7B). To assess the specificity of the pipeline, we supplemented 7-week-old SPF mice with either a casein-only diet or an elemental amino acid-only diet. Through IPHOMED analysis, we exclusively detected the presence of alpha-S 1 casein protein in mice receiving the casein-only diet (Figure 7C), while mice consuming an elemental amino-acid-only diet featured a complete lack of IPHOMED-based dietary signals (Figure 7C).
[0325] The dietary IPHOMED signal may also enable quantification of some aspects of dietary digestion and absorption, as assessed in luminal samples obtained from six different GI locations of NC- or HFD-fed mice (Figure 7D). Interestingly, dietary IPHOMED intensity in HFD-fed mice was markedly pronounced in the stomach, while in NC-fed mice it was also pronounced in the cecum and colon. To elucidate these regional differences in the context of protein digestion, we analyzed the peptide digestion profiles in the stomach under both dietary regimens. In this analysis, we took advantage of the fact that gastric protein digestion is mainly mediated by pepsin, while downstream intestinal digestion, and subsequent ex vivo IPHOMED digestion (Methods), are mediated by trypsin. As such, a higher rate of extra-gastric small intestinal digestion (followed by ex vivo IPHOMED digestion of non-indigenously hydrolyzed proteins) would result in higher proportions of fully trypsinized dietary proteins, while a higher rate of pepsin-induced gastric digestion would result in higher proportions of semi-trypsinized dietary proteins. Indeed, in IPHOMED-processed fecal samples of HFD-fed mice, we observed a significantly greater percentage of peptides exhibiting a semi-trypsinized profile (Figure 7E), suggesting that HFD-based protein is increasingly digested in the stomach and rapidly absorbed thereafter. In line with these digestion profiles, elevated protein coverage (percentage of the protein sequence covered by detected peptides) was detected in the stomach of both NC-fed and HFD-fed mice, which diminished in the upper GI distally to the stomach, especially in mice fed with HFD (Figure 7F). These findings were in line with the increased gastric presence of host- secreted (Figure 4B-C) and microbial (Figure 4G-I) lipid-metabolizing proteins in HFD- consuming mice. Additional digestive patterns could be identified on a per-protein basis along the GI tract. For example, in HFD-fed mice (Figure 7G), we observed a pronounced gastric alpha- SI -casein intensity, which diminished thereafter, likely due to gastric digestion and proximal small intestinal absorption of this protein. Other spatial abundance patterns, assessed in more complex dietary regimens, may help uncover individualized digestive patterns, and merit further studies.
[0326] IPHOMED could also assess temporally shifting fecal dietary signals, likely reflecting upstream digestive alterations. To exemplify this concept, we assessed the IPHOMED dietary signal in fecal samples at different time points after induction of DSS-induced colitis, in which all mice were fed the same diet (i.e., NC, comprised mostly of soybean, wheat, and corn; (Figure 5A, Figure 7A). Interestingly, fecal IPHOMED dietary profiles significantly shifted during colitis (Figure 7H), including an increase in the abundance of fecal wheat (Triticum) proteins, a concomitant decreased abundance of soybean (Glycine) proteins, and stably abundant com- related proteins (Figure 71). Importantly, the total fecal abundance of dietary proteins increased over time (Figure 7J), despite a lack of shift in the dietary content, and was positively correlated with the peptides S100A8 and S100A9, comprising the inflammatory biomarker calprotectin (Figure 7K-L). Collectively, these findings suggest that gut inflammation-induced protein malabsorption / maldigestion may be proxied by discrete dietary protein fecal accumulation patterns, which, in turn, correlate with inflammatory severity and mucosal tissue damage. In all, non-invasive IPHOMED quantification of such nutrient- and host-specific malabsorption patterns may enable the rational personalization of nutritional interventions to the individual, and merit further studies.
[0327] EXAMPLE 2
[0328] MATERIALS AND METHODS
[0329] Human study design
[0330] Israeli pediatric Crohn’s disease cohort (ISR-pCD cohort). This study was approved by the Schneider Children’s Hospital Institutional Review Board (IRB approval number 0722-19- RMC). In total, 54 children were recruited, including 26 pediatric Crohn’s disease (ISR-pCD) patients whose diagnoses were based on standard endoscopic, radiographical, and histological criteria, and 28 healthy controls (ISR-HC). Participants from these two arms were matched by age and gender controlling for bias. All subjects fulfilled the following inclusion criteria: males and females aged 6-18. Exclusion criteria included: (i) chronic treatment with any drug upon enrollment for healthy controls; (ii) the use of systemic antibiotics, probiotics, or proton pump inhibitors three months prior to enrollment; (iii) a diagnosis of type 1 or type 2 diabetes; (iv) any chronic disease (other than IBD for the first arm); (v) any psychiatric disorder; (vi) alcohol or substance abuse; (vii) gut-related surgery, including bariatric surgery; (viii) pregnancy, breastfeeding or fertility treatments for female participants; (ix) morbid obesity (BMI > 95thpercentile for their age and gender). After signing the informed consent, participants enrolled in both arms collected stool samples using the same protocol: a single stool sample was collected at home and then frozen in a home freezer (-20°C) for up to seven days and then brought on site in a cold cooler. Samples were then frozen and stored at -80°C before protein was extracted for proteomics and DNA was extracted for shotgun metagenomic sequencing, as described below.
[0331] German Ulcerative colitis cohort (GER-UC cohort). In total, 100 stool samples were retrieved from a biobank including 50 UC (GER-UC) patients and 50 age- and gender- matched healthy controls (GER-HC). Patients’ diagnoses were based on standard endoscopic, radiographical, and histological criteria. All subjects fulfilled the following inclusion criteria: males and females aged 18-70. Exclusion criteria included: (i) chronic treatment with any drug upon enrollment for healthy controls; (ii) the use of systemic antibiotics, probiotics, or proton pump inhibitors three months prior to enrollment; (iii) a diagnosis of type 1 or type 2 diabetes; (iv) any chronic disease (other than IBD for the first arm); (v) any psychiatric disorder; (vi) alcohol or substance abuse; (vii) gut-related surgery, including bariatric surgery; (viii) pregnancy, breastfeeding or fertility treatments for female participants; (ix) morbid obesity (BMI > 95thpercentile for their age and gender). After signing the informed consent, participants enrolled in both arms collected stool samples using the same protocol: a single stool sample was collected at home into a stool collection tube (Sarstedt cat. Number NC0705093) and sent back by mail to the Institute of Clinical Molecular Biology. Samples were then aliquoted, frozen, and stored at -80°C before protein was extracted for proteomics and DNA was extracted for shotgun metagenomic sequencing, as described below.
[0332] Interventional dietary study (Israel). This short-term interventional dietary study in healthy participants consisted of two short periods in which increased intake (i.e., “enrichment”) of a pair of food items was implemented. The first pair of food items was apples and bananas (pair A), and the second pair of food items was peanuts and cucumbers (pair B). The study began with a weeklong “run-in” period in which participants were not allowed to consume food items belonging to either pair A or pair B. Then a four-days long enrichment of pair A was ensued in which each participant needed to consume a minimum of three apples (approximately 400 grams) and two bananas (approximately 300 gram) daily. During pair A enrichment, participants were allowed to consume anything they wished, in addition to the prescribed bananas and apples, except for food items belonging to pair B. This four-days enrichment of pair A was followed by a “washout” period of seven days in which participants were allowed to consume anything they desired except for apples, bananas, cucumbers, and peanuts. Subsequently, a second four-days enrichment period of pair B was ensued, during which each participant needed to consume at least four cucumbers (approximately 350 grams) and 100 grams of peanuts daily. During the second enrichment period, participants were allowed to consume any other food item in addition to cucumbers and peanuts except for apples and bananas. All subjects fulfilled the following inclusion criteria: Females and males, aged 18-70, with normal BMI (18.5-25). Exclusion criteria: (i) known food allergies to the tested food items; (ii) chronic constipation (less than one bowel movement every two days); (iii) chronic treatment with any drug upon enrolment; (iv) the use of systemic antibiotics, probiotics, or proton pump inhibitors three months prior to enrollment; (v) diagnosis of type 1 or type 2 diabetes; (vi) any chronic disease at the discretion of the study team- such as active or recent cancer, inflammatory bowel disease, systemic autoimmunity; (vii) any psychiatric disorders; (viii) alcohol or substance abuse; (ix) gut-related surgery, including bariatric surgery, but not including previous appendectomy; (x) pregnancy, breastfeeding, or fertility treatments six months prior to enrollment for female participants. Stool samples were collected at home every other day using spoon tubes and stored temporarily and no longer than seven days in a home freezer at -20°C until their transfer to -80°C. Protein was extracted for proteomics and DNA was extracted for shotgun metagenomic sequencing, as described below.
[0333] Celiac cohort (Israel). All subjects satisfied the following inclusion criteria: children aged 3-18 years with suspected Celiac disease, that were referred for a duodenal biopsy due to positive serology or diagnosed with celiac without biopsy using the ESPGHAN guidelines. All subjects were consuming a gluten-containing diet at the time of enrollment. Parents provided informed consent. Exclusion criteria: (i) children on a gluten free diet; (ii) children with suspected or known inflammatory bowel disease or other gastrointestinal diseases that might affect the gut microbiome; (iii) children who have been treated with antibiotics two months prior to endoscopy; (iv) children who underwent gastrointestinal surgeries. A total of 408 subjects with newly diagnosed Celiac disease were enrolled between 2017 and 2023. Some participants provided a stool sample at baseline before commencing on a gluten-free diet, and some participants provided a sample at routine medical follow-up after 1-2 years of gluten-free diet. Samples were collected at home using a DNA-pre serving buffe-containing tube (Genotek Omni-gene tube catalog No.23- 410-000). Samples were then delivered to the Weizmann Institute of Science, aliquoted, frozen, and stored at -80°C before protein was extracted for proteomics and DNA was extracted for shotgun metagenomic sequencing, as described below.
[0334] EEN cohort (Israel). All subjects satisfied the following inclusion criteria: children aged 6-18 years with suspected CD. Parents provided informed consent. Exclusion criteria: (i) any known autoimmune or rheumatologic disease; (ii) chronic gastrointestinal disorder (e.g. celiac disease, eosinophilic esophagitis, collagenous gastritis, autoimmune gastritis, etc.); (iii) past or present history of malignancy; (iv) primary immunodeficiency; (v) Pregnancy in the last six months, breastfeeding, (vi) Morbid obesity (BMI > 95thpercentile for their age and gender); (vii) following dietary regimen / dietitian consultation / participation in another study; (viii) chronic treatment with any oral / systemic immunosuppressive or anti-inflammatory drugs (e.g. steroids, 5- amino salicylic acid, immunomodulators, biologies, etc.). Patients receiving these drugs as inhalers / creams / ointments were be excluded from the study; (ix) use of systemic antibiotics, probiotics or NSAIDs during 30 days prior to enrollment; (x) gut-related surgery, including bariatric surgery. A total of 109 subjects with new-onset CD were enrolled between 2020 and 2024. Participants provided the first stool sample at baseline before any medical or dietary intervention, and then after one week of consuming 100% exclusive enteral nutrition (EEN). Samples were collected at home using a DNA-preserving buffer-containing tube (Genotek Omnigene tube catalog No.23-410-000). Samples were then delivered to the Weizmann Institute of Science, aliquoted, frozen and stored at -80°C before protein was extracted for proteomics and DNA was extracted for shotgun metagenomic sequencing, as described below.
[0335] Fecal microbiome transfers
[0336] Eight weeks old C57BL / 6 wild-type GF mice were routinely monitored for sterility. All mice were maintained under a strict 12-h light-dark cycle, with lights on at 6 am and off at 6 pm.. Stools from a UC patient or a HC were thawed, resuspended in sterile PBS, and homogenized inside an anaerobic chamber. Solids were allowed to sink for five minutes, and the supernatant was collected into a sterile anaerobic Hungate tube and transferred to the animal facility. GF mice were inoculated with either 200pL of the UC or healthy control supernatant by oral gavage and placed in a sterilized isocage. 6-8 mice were used as recipients per each donor. Following two weeks of colonization, mice were sacrificed, and stool samples were collected for proteomics and metagenomics. Colon samples were preserved in 4% PFA and H&E-stained slides were prepared. Histopathologic evaluation was performed by a certified veterinarian specialized in pathology who remained blinded with regard to specimen group. Blood samples were obtained by cardiac puncture, centrifuged for 5 minutes at 3,200g at 4°C, and supernatant was frozen in -80°C until they were analyzed by Multiplexed ELISA as described below. This experiment was repeated three independent times with three sets of matched donor pairs - one UC donor and one HC donor per each repetition.
[0337] Multiplexed ELISA
[0338] The V-PLEX Mouse Cytokine 19-Plex Kit (CatNo K15255D-1, Meso Scale Diagnostics, Rockville, MD, US) was utilized. Multiplex MSD assays were performed as per manufacturer instructions. The kit is composed of two MULTISPOT® 96-well panels and utilizes sandwich immunoassay technology: the cytokine panel was coated with IL-9, MCP-1, IL-33, IL-27p28 / IL- 30, IL-15, IL-17A / F, MIP-la, IP-10, and MIP-2 capture antibodies and the pro-inflammatory panel coated with INF-g, IL-1B, IL-2, IL-4, IL-5, IL-6, KC / GRO, IL-10, IL-12p70, and TNF-a. Following washing with washing buffer, MSD standards or samples diluted 1: 1 in diluent were added to the wells. All samples and standards were run in duplicate. After 2 h incubation and washing, SULFO-TAG™ labeled detection antibody solution was added and the plates incubated a further 2 h. Following washing, MSD® Read Buffer T was added and plates were read immediately using a MESO® QuickPlex SQ 120. The standard curve was established by fitting the signals from the standard using a 4-parameter logistic model. Concentrations of samples (pg / ml) were determined from the electrochemiluminescence signals by back-fitting to the standard curve and multiplied by the dilution factor.
[0339] RNA sequencing and analysis
[0340] Purification. Total RNA from intestinal mucosal samples (terminal ileum, distal colon) was extracted & purified using an RNAeasy kit (QIAGEN, 74104) according to manufacturer’s instructions.
[0341] RNA-sequencing. Ribosomal RNA was selectively depleted by RNAseH (New England Biolabs, M0297). Specifically, a pool of 50bp DNA oligos (25 nM, IDT) that is complementary to murine rRNA18S and 28S, was resuspended in 75 pL of 10 mM Tris pH 8.0. Total RNA (lOOOng in 10 pL H2O) was mixed with an equal amount of rRNA oligo pool. The RNA was added to 2pL diluted oligo pool and 3pL 5x rRNA hybridization buffer (0.5 M Tris-HCl, IM NaCl, titrated with HC1 to pH 7.4). Samples were incubated at 95 °C for 2 min, then the temperature was slowly reduced (0.1 C / s) to 37 °C. RNaseH enzyme mix (2 pL of 10U RNaseH, 2 pL lOx RNaseH buffer, IpL H2O, total 5pL mix) was prepared 5 min before the end of the hybridization and preheated to 37 °C. The enzyme mix was added to the samples when they reached 37 °C and they were incubated at this temperature for 30 min. Samples were purified with 2.2x SPRI beads (Ampure XP, Beckmann Coulter) according to the manufacturers’ instructions. Residual oligos were removed with DNase treatment (ThermoFisher Scientific, AM2238) by incubation with 5pL DNase reaction mix (IpL Turbo DNase, 2.5 pL Turbo DNase lOx buffer, 1.5 pL H2O) that was incubated at 37 °C for 30 min. Samples were again purified with 2.2x SPRI beads and suspended in 3.6 pL priming mix (0.3 pL random primers of New England Biolab, E7420, 3.3pl H2O). Samples were subsequently primed at 65 °C for 5 min. Samples were then transferred to ice and 2 pL of the first strand mix was added (1 pL 5x first strand buffer, NEB E7420; 0.125 pl RNase inhibitor, NEB E7420; 0.25 pL ProtoScript II reverse transcriptase, NEB E7420; and 0.625 pL of 0.2 pg / mL Actinomycin D, Sigma, A1410). The first strand synthesis and all subsequent library preparation steps were performed using NEBNext Ultra Directional RNA Library Prep Kit for Illumina (NEB, E7420) according to the manufacturers’ instructions (all reaction volumes reduced to a quarter). Libraries that passed quality control were loaded with a concentration of 2 pM on 75 cycle high output flow cells (Illumina, 20024906) and sequenced on a NextSeq 500 (Illumina) with the following cycle distribution: 8 bp index 1, 8 bp index 2, 75 bp read 1.
[0342] RNA-sequencing analysis. We used Illumina’s bcl2fastq script to convert the raw files to fastq files. Fastq files were quality filtered using Fastp (v0.20.0) with default parameters. We aligned the reads to the murine reference transcriptome with Gencode annotation (GRCm38.p6, Release M24) and quantified gene expression using STAR (v2.7.3a). Differential expression analysis was carried out with the DESeq2 package with a model for paired samples ( -subject + condition ). Gene Ontology analysis was carried out with g:Profiler with default settings.
[0343] Microbiome sequencing and analysis
[0344] DNA purification. DNA from intestinal mucosa (ileum, distal colon) or stool was extracted and purified using a Purelink Microbiome DNA purification kit (Invitrogen).
[0345] Shotgun metagenomics sequencing and processing. For shotgun sequencing, Illumina libraries were prepared using a Nextera DNA Sample Prep kit (Illumina, FC-121-1031), according to the manufacturer’s protocol, and sequenced on the Illumina NextSeq or Novaseq platform with a read length of 80bp or 125bp respectively. Illumina’s bcl2fastq script was implemented to generate the fastq files. QC trimming was carried out using Fastp, and host sequences were removed with Bowtie2 (2.4.1) and human genome reference hg37. The cleared fastq files were subsampled, by dataset, using Seqtk. We carried out the taxonomic assignment of bacterial DNA relying on exact alignment of k-mers with Kraken2 (2.0.9) against the Genome Taxonomy Database (https: / / gtdb.ecogenomic.org / , version 95). To improve the accuracy of species-level classification, we applied Bayesian re-estimation of bacterial abundance with Bracken (v2.5.3). Alpha and beta diversity, based on Bray-Curtis dissimilarities, were analyzed using Vegan R- package (https: / / CRAN.R-project.org / package=vegan). Differential species abundance analysis was performed using DESeq2. Functional annotation was implemented using protein alignment with DIAMOND, thereby, only the first hit was considered, and an e value <0.0001 was accepted. Principal component analysis was performed based on Bray-Curtis dissimilarity using ape R package. Permutational multivariate analysis of variance (PERMANOVA), on Bray-Curtis dissimilarities, was used to test for differences in the taxonomical, functional profiles.
[0346] Mass spectrometry-based proteomics
[0347] Sample preparation. Samples were subjected to in-solution tryptic digestion using suspension trapping (S-trap) as previously described114. Briefly, mucosal tissues were first homogenized in cold sterile PBS / _. After centrifugation, the clear supernatant containing soluble proteins was supplemented with lysis buffer to a final concentration of 5% SDS in 50mM Tris-HCl pH 7.4. Mouse and human stool samples were directly homogenized in lysis buffer containing 5% SDS in 50mM Tris-HCl pH 7.4. Lysates were then incubated at 96°C for 5 min, followed by six cycles of 30 sec of sonication (Bioruptor Pico, Diagenode, USA). Protein concentration was measured using the BCA assay (Thermo Scientific, USA). An amount of 50pg total protein was reduced with 5mM dithiothreitol and alkylated with lOmM iodoacetamide in the dark. Each sample was loaded onto S-Trap microcolumns (Protifi, USA) according to the manufacturer’s instructions. After loading, samples were washed with 90: 10% methanol / 50mM ammonium bicarbonate. Samples were then digested with trypsin (1:50 trypsin / protein) for 1.5 h at 47°C. The digested peptides were eluted using 50mM ammonium bicarbonate. Trypsin was added to this fraction and incubated overnight at 37°C. Two more elutions were made using 0.2% formic acid and 0.2% formic acid in 50% acetonitrile. The three elutions were pooled together and vacuum- centrifuged to dryness. Stool samples were subjected to an additional cleaning step using solidphase extraction (Oasis HBL, Waters, MA, USA) according to manufacturer instructions.
[0348] Liquid chromatography. ULC / MS grade solvents were used for all chromatographic steps. Dry digested samples were dissolved in 97:3% FLO / acctonitrilc + 0.1% formic acid. Each sample was loaded and analyzed using split-less nano-Ultra Performance Liquid Chromatography (10 kpsi nanoAcquity; Waters, Milford, MA, USA). The mobile phase was: A) H2O + 0.1% formic acid and B) acetonitrile + 0.1% formic acid. Desalting of the samples was performed online using a Symmetry C18 reversed-phase trapping column (180pm internal diameter, 20mm length, 5pm particle size; Waters). The peptides were then separated using an HSS T3 nano-column (75pm internal diameter, 250mm length, 1.8pm particle size; Waters) at 0.35pL / min. For label-free quantitative (LFQ) proteomic analysis, peptides were eluted from the column into the mass spectrometer using the following gradient: 4% to 25%B in 155 min, 25% to 90%B in 5 min, maintained at 90% for 5 min, and then back to initial conditions.
[0349] Mass Spectrometry. The nanoUPLC was coupled online through a nanoESI emitter (10pm tip; New Objective; Woburn, MA, USA) to a quadrupole orbitrap mass spectrometer (Q Exactive HFX or HF, Thermo Scientific) using a Flexion nanospray apparatus (Proxeon). For EFQ analysis, data were acquired in data-dependent acquisition (DDA) mode, using a Top 10 method. MSI resolution was set to 120,000 (at 200m / z), mass range of 375-1650m / z, AGC of 3e6, and maximum injection time was set to 60msec. MS2 resolution was set to 15,000, quadrupole isolation 1.7m / z, AGC of le5, dynamic exclusion of 45sec, and maximum injection time of 60msec.
[0350] Proteomics analysis. The refined database is used as a reference to search and quantify host, bacterial and dietary proteins in the detected signal by EC-MS / MS using Metamorpheus (v0.0.319), as described in Example 1, herein above. Detected dietary proteins were subsequently verified by realigning the assigned unique peptides to the NCBI’s non-redundant protein sequences (July 2022) (blastp -word_size 2 -matrix PAM30 -threshold 11 -comp_based_stats 0 -outfmt 7 - evalue 200000 -gapopen 9 -gapextend 1 -num_alignments 100 -window_size 40). Only proteins where, at least, one unique peptide was assigned, with the highest score to the expected dietary genus, were selected. Log 10 transformation and quantile normalization were performed on mass spectrometry intensities using R packages DEP (v 1.18.0) and limma after dividing proteins into host, bacterial and dietary proteins. Principal component analysis (PCA) was carried out based on log-transformed intensity. Permutational multivariate analysis of variance (PERMANOVA) was performed for pairwise comparisons on a distance matrix between group levels with corrections for multiple testing using the R package pairwiseAdonis (v0.4). Differential protein expression was tested by applying protein- wise linear models combined with empirical Bayes statistics, from DEP package. For dietary proteomics analysis, PCA was computed with both protein and genus abundances, while differential abundance analysis was applied using Mann-Whitney test and Benjamini-Hochberg correction for multiple testing.
[0351] Bacterial functional analysis. Metabolic pathway annotations for all bacterial proteins included in the refined database were downloaded from Biocyc. Functional enrichment analysis of differentially abundant bacterial proteins was performed with the enricherQ function (R package clusterProfiler, v4.4.4), applying hypergeometric test and Benjamini-Hochberg correction for multiple testing.
[0352] Host functional analysis. Functional enrichment analysis of differentially abundant host proteins was computed using the gost function of the gprofiler2 R package (vO.2.1).
[0353] Comparison between shotgun metagenomics and bacterial proteomics. Abundance quantification of bacterial genes (shotgun metagenomics) and bacterial proteins (metaproteomics) was carried out using the refined catalog of bacterial genes / proteins for both methodologies. To compare the distribution of gene / protein abundances, both quantifications were combined and transformed using quantile normalization. Alpha diversity was measured based on Shannon Index. Random-forest analysis was performed for both methodologies using randomForest function from the randomForest R package (v4.7, repeated five-fold cross-validation iterations, 10 repetitions) and applied to the GER and ISR IBD cohorts.
[0354] Correlation between microbiome composition / function and host response. From a cohort of eight healthy individuals, RNA-seq analysis from stomach, duodenum, jejunum, terminal ileum, and descending colon biopsies was correlated with proteomics profiles from luminal samples across the gastrointestinal tract and stool samples.
[0355] Host-host correlations'. 1) For each individual and GIT region, host protein abundances, measured in the luminal sample, were correlated with the abundances of the corresponding host genes quantified in the paired gastrointestinal biopsy. 2) For each individual and GIT region, host protein abundances, measured in the luminal sample of the GIT region, were correlated with the abundance of the same host proteins quantified in the stool sample.
[0356] Microbiome-microbiome correlation: for each individual and GIT region, bacterial protein abundances, measured in the luminal sample of the GIT region, were correlated with the same bacterial protein abundances quantified in the stool sample. For all these comparisons, gene abundances were normalized applying the variance-stabilizing transformation (VST) from the DESeq2 R package (v 1.36.0); protein abundances were transformed using quantile normalization. Correlations were performed using general linear regression models. Global correlation values were computed as the median R2across all individuals.
[0357] Bacterial-host protein correlations'. 1) For each GIT region, all bacterial protein abundances, measured in all luminal samples of that GIT region, were correlated with the abundance of all host genes quantified in the paired gastrointestinal biopsies. 2) For each GIT region, all bacterial protein abundances, measured in all stool samples, were correlated with the abundance of all host genes quantified in the paired gastrointestinal biopsies. Correlations were performed using general linear regression models. For each species, p-values were corrected for multiple testing by Benjamini-Hochberg. For each bacterial species, functional enrichment analysis was performed for all host genes that were significantly correlated with, at least, one protein produced by this bacterial species: we applied the enricher() function (R package clusterProfiler, v4.4.4), based on the hypergeometric test and Benjamini-Hochberg correction for multiple testing.
[0358] Biomarker analysis'.. Random-forest classification, evaluated in repeated five-fold cross- validation iterations (10 repetitions), was developed to identify discriminating proteins, including host and bacterial proteins, for UC or CD diagnosis. Top discriminating features were randomly combined in sets of 1, 2 or 3 proteins and tested in different cohorts for diagnosis or severity score prediction. Performance was compared to calprotectin, the gold standard for IBD diagnosis and severity. Diagnostic accuracy was analyzed using Random-forest classification, while severity score prediction was evaluated using Support Vector Machine for regression (repeated five-fold cross-validation iterations). HMP64, Mills cohort, as well as our German and Israeli cohorts, were processed and analyzed, as described above, by combining the randomForest (v4.7) and el071 (vl.7-11) R packages. Similar biomarker analysis was performed after collapsing bacterial protein intensity in KO terms, by summing of the intensity of all bacterial proteins associated with a KO term.
[0359] Datasets
[0360] Shotgun metagenomics and proteomics data were downloaded from the IBD multi-omics database, including the 3 different cohorts from the Cincinnati Children's Hospital, Cedars-Sinai Medical Center and Massachusetts General Hospital5. Similarly, two IBD datasets collected by Mills et al. Nat Microbiol 7, 262-276 (2022) were downloaded and reanalyzed for biomarker analysis: proteomics data from the MassIVE repository (MSV000082094, MSV000086509), and metagenomics from the European Nucleotide Archive (https: / / www.ebi.ac.uk / ena, PRJEB42151). Furthermore, shotgun metagenomics and proteomics data of intestinal and stool samples from 15 healthy individuals were downloaded from the European Nucleotide Archive (PRJEB42151) and Proteomics Identifications Database, PXD038906), respectively. Finally, a different dataset of shotgun metagenomics data from 12 treatment-naive and 7 broad- spectrum antibiotics-treated individuals was downloaded from the European Nucleotide Archive (PRJEB28097).
[0361] Quantification and statistical analysis.
[0362] All statistical analyses were performed using R (3.6.3) and GraphPad Prism software (GraphPad Software 10.0.2, Inc.). For principal component analysis we used permutational analysis of variance (PERMANOVA). For differential protein abundance analysis, we used protein- wise linear models combined with empirical Bayes statistics and FDR correction by Benjamini-Hochberg. For pathway enrichment analysis we used a hypergeometric test and FDR correction by Benjamini-Hochberg.
[0363] RESULTS
[0364] Adaptation of IPHOMED for human use. IPHOMED was calibrated, validated, and tested at the in silica, in vitro, and in vivo animal model setting, using a variety of dietary, colonization, infection and auto-inflammatory conditions [as described in Example 1]. In stool samples from healthy humans (n=76), IPHOMED detected 362.8+5.3 host proteins, 3746.7+46.3 bacterial proteins, and 57.6+1.6 dietary proteins, respectively, collectively accounting for 89.84+0.12%, 8.75+0.13%, and 1.39+0.04% of the total gastrointestinal metaproteome, respectively.
[0365] We began our adaptation of IPHOMED into human use, by performing an in silica simulation to analyze the sensitivity and specificity of our methodology. Using pre-defined criteria [as described in Example 1], we constructed a refined bacterial protein database using metagenomic data derived exclusively from stool samples of 100 individuals (Germany, Methods), as to include only bacterial proteins that are likely to be specifically present in this cohort. Mass spectra obtained from liquid chromatography-mass spectrometry (LC-MS / MS) were then searched against the refined cohort- specific database to quantify (i) the number of detected bacterial proteins, (ii) the number of detected bacterial species and (iii) the duration of database search runtime. To quantitate the impacts of using a generic, non-refined, bacterial protein database, we in silica adulterated the refined custom-built protein database with millions of random bacterial proteins downloaded from the BioCyc database. As with mice [see Example 1], utilization of a generic human microbiome proteomic dataset resulted in a reduced detection of bacterial proteins and species-level taxonomic resolution, coupled with markedly longer processing time. Together, these results indicate that experimentally defining the potential bacterial metaproteome by use of per-experiment metagenomic bacterial protein-coding data as reference, significantly improves the sensitivity, specificity, and computational efficiency of the analysis.
[0366] Shotgun metagenomic sequencing is the most commonly used modality for assessment of microbiome functional potential, by the assumption that protein abundance (and thus function) can be inferred from gene abundance. However, as in eukaryotic cells, the existence of transcriptional, post-transcriptional and post-translational regulatory mechanisms in prokaryotic cells often results in de-coupling of gene and protein abundances. To compare metagenomic and IPHOMED-based functional assessment, we analyzed stool samples of healthy individuals from two independent and age-diverse human cohorts (Israel, n=28, and Germany, n=50, Methods). Expectedly, considering that not all bacterial genes are constitutively expressed, microbiome diversity was significantly higher by at gene-level as compared to protein-level. Generally, genes featuring a lower mean abundance were less likely to be detected by metaproteomics, indicating that proteomics is less sensitive in detecting the products of low-abundance genes. A greater metagenomic interindividual variability (at species, gene, and pathway level) was noted compared to IPHOMED, likely resulting from the presence of numerous genes at the DNA level that are not actively expressed. Importantly, the majority of species exhibited protein abundances that were independent of gene abundance distribution. Even within a single species level, some proteins displayed similar abundance distributions across individuals as their corresponding genes, while others featured discordance to metagenomic abundance, suggesting a regulated expression pattern. Interestingly, genomically-unpredictable proteins (i.e., proteins having different abundance distribution to their corresponding gene) were significantly more likely to have extracellular location and function. Since such secreted proteins and peptides are often involved in bacteria- bacteria and trans-kingdom signaling, their assessment by IPHOMED may be advantageous in studying host-microbiome interactions.
[0367] We next compared IPHOMED’ s customized per-experiment protein-coding reference genome assembly to a previously used approached of de novo genome assembly, that does not utilize reference genomes (assembled based on pre-existing taxonomic data. To this aim, we analyzed a dataset obtained from Human Microbiome Project (HMP) with shotgun metagenomics and LC-MS / MS data of individuals with Ulcerative Colitis (UC), Crohn’s disease (CD), and matched healthy controls (HC; n=12, n=34, n=52, respectively). We utilized the same paired-end- sequenced metagenomic data to construct two refined proteomic databases, by either performing de novo genome assembly or the IPHOMED per-experiment pipeline (Methods). In total, the refined de novo genome assembly-based and IPHOMED-based bacterial protein database comprised 549,130 and 284,197 bacterial proteins, respectively. Use of genome assembly enabled the detection of 5,459 bacterial proteins, of which 2,794 (51.18%) were proteins assigned to species and the rest assigned to taxa at a lower taxonomic resolution (i.e., genus or family). In comparison, use of IPHOMED enabled the detection of 12,316 bacterial proteins, of which 11,580 (94.02%) were assigned to species level. At the genomic level, IBD patients often feature a dysbiotic microbiome composition as compared to healthy controls. To test which approach better captures the differences between IBD patients and healthy controls, we trained a random-forest classifier on either data generated from de novo genome assembly or IPHOMED. Indeed, IPHOMED-generated custom reference database yielded an improved diagnostic accuracy of IBD (UC and CD combined, while unraveling a markedly increased number of differentially abundant bacterial proteins and a higher number of contributing species for nearly all identified bacterial pathway. For example, only 10 bacterial species contributed to a gram-negative bacterial secretion systems (KEGG pathway ko03070) according to the genome-assembly approach, compared with 33 species using IPHOMED. A similar comparison of the accuracy of bacterial protein quantification using another published dataset of IBD patients (n=184) and HC (n=21) utilizing another generic reference database approach, exhibited an improved performance of IPHOMED, which successfully identified a greater number of differentially abundant proteins between healthy individuals and IBD patients. These comparative experiments suggest that IPHOMED enables a superior performance and enhanced taxonomical resolution in functionally characterizing microbiome community structure.
[0368] IPHOMED-based functional characterization of the naive and antibiotic-perturbed human gastrointestinal tract. We next utilized IPHOMED to characterize the functional hostmicrobiome landscape along the human gut by assessing a healthy cohort of treatment-naive (n=12) and bro ad- spectrum antibiotics -treated (ABX, n=7, metronidazole 500 mg tri-daily and ciprofloxacin 500 mg bi-daily) adult volunteers (Methods). All participants underwent a standard bowel preparation (20 hours of clear liquid diet prior to the examination, oral consumption of sodium picosulfate-based laxative, and two fleet enemas), followed by an upper and lower endoscopy to obtain luminal and mucosal samples from the stomach, duodenum, jejunum, terminal ileum, cecum, descending colon, and stool (Methods). Overall, the IPHOMED species composition of the gut microbiome in treatment-naive or antibiotics -treated healthy adults showcased distinct microbiome activity across the GIT (Figure 9A). Examples of representative niche- specific commensals identified by protein configuration included Helicobacter pylori in the stomach; Eubacterium ramulus, and Ruminococcus gnavus in the jejunum; Bifidobacterium adolescentis in the terminal ileum; and Prevotella buccae, Paraprevotella clara, and Dialister succinatiphilus in stool. For example, IPHOMED detected 39 gastric proteins produced by Helicobacter pylori, most of which were detected exclusively in one individual, who expressed the cag pathogenicity island (cag PAI) protein, a well-established virulence factor, implicated in 90% of all chronic active gastritis, peptic ulcer disease, and gastric cancer cases. Importantly, the same individual also expressed an oxygen-insensitive NAD(P)H nitroreductase, which has been associated with resistance to nitroimidazole antibiotics, underscoring IPHOMED’ s potential clinical utility.
[0369] The number of detected bacterial proteins (Figure 9B) and species distally increased along the gut in treatment-naive individuals, in agreement with previous metagenomic studies demonstrating a bacterial density gradient along the GIT. Compared to the small intestinal IPHOMED profile, the colonic gut bacterial metaproteome was more similar to the stool bacterial metaproteome. However, even the distal colonic luminal IPHOMED microbiome profile remained distinct from that of stool, indicating that, alike microbiome composition, functional colonic community structure could only partially be inferred from stool samples. Coprophagy-associated gastric-colonic similarity noted in mice was absent in humans, suggesting that human fecal-oral transition had substantially less impact on gastric microbial community structure than in caged mice. Expectedly, proteins from aerobic species, such as Streptococcus salivarus and Streptococcus thermophilus, were mostly detected in proximal parts of the GIT. Such IPHOMED- based niche- specific functional commensal configurations were distinct in multiple features from those identified by shotgun metagenomic sequencing. For example, 53 proteins expressed by S. salivarius, including multiple enzymes involved in carbohydrate uptake mechanisms, such as phosphoenolpyruvate-protein phosphotransferase, phosphocarrier protein HPr, and the amino acid ABC transporter ATP-binding protein, were significantly more abundant in stomach samples compared to stool using IPHOMED but not shotgun metagenomic sequencing. These transporters endow Streptococci with metabolic flexibility, allowing them to exploit alternative energy sources, when glucose, their preferred substrate, is unavailable.
[0370] Some characteristics of the IPHOMED-based functional microbiome configuration in naive humans closely resembled those observed in naive mice (see Example 1). For example, the enzymes specialized in synthesizing and degrading complex carbohydrates (CAZymes) such as the GH13 family of enzymes and the CBM48 carbohydrate-binding modules, which possess glycogenbinding functions and are appended to GH13 modules to break down starch, were highly abundant in the colon and cecum of both species. Differences between the two host species were noted in other CAZymes, such as GH20, which is implicated in the breakdown of animal glycans, and was significantly more abundant in the human lower GI tract and, likely as a result of inclusion of such components in the human diet. In contrast, the enzyme GH73, mirroring the activity of virulence- associated peptidoglycan hydrolases in breaking down peptidoglycans during host cell invasion, was significantly enriched throughout the murine gastrointestinal tract compared to the human GI tract.
[0371] Broad-spectrum antibiotic treatment resulted in a substantial functional dysbiosis (Figure 9A), marked by an expansion of facultative anaerobes, with a higher proportion of proteins produced by Escherichia coli and Enterococcus species noted in the duodenum, terminal ileum, cecum, colon, and stool. Unlike treatment-naive humans, ABX-treated individuals featured a relatively constant number of proteins (Figure 9B) and species along their lower GI tract, reflecting their dysbiosis-associated loss of GIT microbial gradient. In naive individuals, bacterial proteins associated with antibiotic resistance were detected across all lower GI tract sections and stool samples, conferring resistance to a diverse spectrum of antibiotics (Figure 9C). Following antibiotic treatment, marked individualized differences in this resistance profile were observed. Notably, most IPHOMED-detected Antibiotics-resistant proteins (ARP) were undetectable in fecal samples, suggesting that stool analysis may not be a reliable approach the determine the presence of antibiotic-resistant strains in individuals treated with antibiotics.
[0372] The landscape of host proteins along the GIT was likewise niche- specific. Unlike bacterial proteins, the number of host proteins was not significantly affected by GIT region or by ABX exposure (Figure 9D). The global host proteome signature along the GIT featured three main clusters that partly overlap the embryonal division to foregut, midgut, and hindgut: (i) the stomach, (ii) the duodenum and jejunum, (iii) the ileum, large intestine, and stool (Figure 9E). For example, essential gastric proteins, including hydrogen-potassium ATPase, crucial for gastric acid secretion were enriched in the stomach. Additionally, calpastatin and destrin, reduced in gastric tumors and associated with tumor suppressive activities, were more abundant in gastric samples. Similarly, eight proteins, all involved in protein metabolism, were significantly increased in abundance in the duodenum and jejunum. Among them, GTP cyclohydrolase I feedback regulatory protein (GCHFR) was exclusively detected in the upper GI tract, participating in processes like antigenreceptor signaling and the methyl-transferase network. Notably, this protein is downregulated in the inflamed duodenum of patients with Common Variable Immunodeficiency (CVID). Likewise in the large intestine, we detected 39 host protein markers, including multiple transporters facilitating the absorption of short-chain fatty acids, sugar phosphates, sulfates or carboxylates in the cecum and colon.
[0373] Interestingly, broad-spectrum antibiotic treatment had no effect on this triple-cluster configuration (Figure 9F), while the profile of host protein markers identified in naive individuals was recapitulated in volunteers subjected to antibiotic treatment. Regional differential abundance analysis unveiled several host proteins consistently affected by antibiotic treatment across various GI regions. Notably, the expression of the non-ATPase regulatory subunits 2 and 3 of the 26S proteasome (PSMD2 and PSMD3) showed a significant reduction in lower GI samples following antibiotic treatment. Dysfunction of the ubiquitin-proteasome system has been linked to IBD, colorectal cancer, and gastrointestinal infections.
[0374] We next compared the human host gut function assessed by mucosal RNA-sequencings to that assessed in corresponding luminal samples by IPHOMED. In a global analysis of host proteome (Figure 9E), host transcriptome, bacterial metagenome, and bacterial proteome landscapes along the GIT, small intestinal samples clustered separately from large intestinal samples. Interestingly, host proteome signatures of the terminal ileum clustered together with cecal, colonic, and fecal samples (Figure 9E), while they clustered with small intestinal samples in host transcriptome. In concordance with previous reports, we observed only a weak transcriptprotein correlation across the human GIT. Abundance of host proteins in stool samples featured a stronger correlation with the abundance of host proteins in luminal samples from the large intestine, compared to luminal samples from the small intestine. In contrast, bacterial proteins across the GIT exhibited a limited correlation with stool bacterial proteins. Collectively, these results demonstrate that the niche- specific host gut mucosal function inferred by gene expression at RNA level, is markedly distinct from that directly assessed by IPHOMED at protein level. Interestingly, multiple host epithelial transcripts, involved in signal transduction and immune function, significantly correlated with regional IPHOMED bacterial protein abundances, including for example proteins produced by Clostridium symbiosum in the duodenum and Dorea formicigenerans in the cecum. In addition, several host epithelial transcripts significantly correlated with the abundance of stool bacterial proteins, and, in particular, Faecalibacterium prausnitzii strongly correlated with host transcriptome in the ileum.
[0375] IPHOMED is superior to metagenomics in identifying differential host and microbiome features in human inflammatory bowel disease (IBD). We next utilized IPHOMED to uncover new aspects of host-microbiome interplays in IBD. We began by assessing the fecal IPHOMED profile in a cohort of Israeli subjects with pediatric Crohn’s disease (ISR- pCD, n=26) and age- and gender-matched healthy controls (ISR-HC, n=28, Figure 10A). Fecal ISR-pCD host IPHOMED signatures were markedly distinct from the ISR-HC group (Figure 10B). Elevated disease-associated host proteins included calprotectin subunits (S100A8, S100A9), cathepsin G, carcinoembryonic antigen-related cell adhesion molecule 8 (CEACAM8), myeloperoxidase (MPO), and multiple immunoglobulins (Figure IOC), with host pathway enrichment demonstrating that seven out of the top ten enriched pathways were associated with immunity, particularly neutrophil function and the defense response to bacteria and fungi (Figure 10D). Likewise, the bacterial proteomic profile in stool samples of subjects with ISR-pCD significantly differed from healthy individuals (Figure 10E). Differential abundance analysis of bacterial proteins revealed multiple proteins that were significantly altered in ISR-pCD patients (Figure 10F). Grouping differentially abundant bacterial proteins into their respective producing species, revealed that the bacterium that produces the highest number of differentially abundant proteins in ISR-pCD, Alistipes putredinis, was not detected by traditional differential taxonomical abundance analysis with shotgun metagenomics data, highlighting the additive value of proteinbased approaches in identifying disease-associated bacterial candidates (Figure 10G). Pathway enrichment analysis revealing multiple bacterial metabolic pathways that were altered in pCD patients, including those related to biosynthesis of short-chain fatty acids (SCFAs) and branched- chain amino acids (BCAAs) (Figure 10H). Both SCFAs and BCAAs serve important immune functions including lymphocyte growth and activity and maintenance of gut barrier function. Accurate quantification of bacterial and host proteins facilitated the discovery of associations between host protein markers indicative of ISR-pCD and bacterial proteins (Figure 101). Interestingly, several essential bacterial proteins produced by Eubacterium rectale, Eubacterium halii, and Faecalibacterium prausnitzii, recognized for their depletion in IBD patients, exhibited a negative correlation with host proteins implicated in immune-inflammatory responses, such as
[0376] CTSG, PRTN3, S100A8, and S100A9.
[0377] We next analyzed the IPHOMED landscape in a cohort of German adults with Ulcerative colitis (GER-UC, n=50) as compared to age- and gender-matched GER-HC (n=50) (Figure 11A). Similar to subjects with ISR-pCD, the host proteome of adult GER-UC patients featured a marked global shift (Figure 11B), with multiple proteins related to mucosal immune response being differentially abundant, including calprotectin subunits (S100A8, S100A9), mucins, CEACAM6, CEACAM7, CEACAM8, proteinase 3, and multiple immunoglobulins (Figure 11C). The most enriched host pathways in the fecal proteome of subjects with GER-UC included immune functions associated with neutrophil activity and defense response to fungi (Figure 11D). While a considerable share of differentially abundant proteins in both ISR-pCD and GER-UC were immune related, particularly with respect to neutrophil function, some of the IBD-associated differentially abundant proteins were observed only in one of the disease cohorts. This could be explained by disease-related disparity, differences in age-related IBD manifestations, extent and involved regions of inflammation, samples size, and demographics, among other factors. Similar to ISR-pCD, the bacterial proteome of GER-UC patients was also globally perturbed compared to healthy individuals (Figure HE). Several hundred bacterial proteins were differentially abundant in diseased adults (Figure HF). Bacterial functional pathway enrichment analysis identified several pathways related to carbohydrate metabolism and fermentation to be significantly altered between the microbiome of GER-UC patients and that of GER-HC, a perturbation likely reflecting a response to environmental stress in an inflammatory environment. Grouping the differentially abundant bacterial proteins into their respective producing bacterial species revealed several differentially abundant bacterial species, such as Faecalibacterium prausnitzii, Fusicatenibacter saccharivorans, Ruminococcus callidus, Alistipes putredinis, and Bacteroides coprocola, which were not discovered by shotgun metagenomics (Figure 11G). Interestingly, the vast majority of IBD-associated differentially abundant bacterial proteins in both ISR-pCD pediatric (Figure 10G) and GER-UC adult (Figure 11G) populations were significantly downregulated as compared to their respective HC (Figure 11H). The expansion of intestinal antimicrobial peptides noted in IBD patients, coupled with under- abundance of bacterial proteins across multiple taxa in two independent cohorts, potentially implicates a host innate immune suppressive effect on the commensal microbiome, in contrast with pathobiont expansion pursued by most IBD studies.
[0378] To substantiate the apparent superiority of IPHOMED over metagenomics in detecting IBD-associated microbiome alterations (Figure 10G, 11G), we compared the diagnosis predictive capacity of IPHOMED versus traditional metagenomic analysis. Indeed, in both ISR-pCD and adult GER-UC cohorts, random-forest diagnostic prediction demonstrated a significantly superior performance of IPHOMED (based on protein abundance), as compared to metagenomics (based on gene abundance, Figure 111). Interestingly, a correlation abundance analysis between host and bacterial proteins uncovered a significant positive association between host antimicrobial peptides, including S100A8, S100A9, LTF, PRTN3 and MPO, and the bacterial iron-sulfur cluster assembly enzyme (IscU) produced by Escherichia coli (Figure 11J). Notably, iron-sulfur clusters play a pivotal role as cofactors for diverse transcriptional regulators in bacteria, including prominent mammalian pathogens like E. coli. The adaptability of iron-sulfur clusters to factors such as iron availability, oxygen tension, and reactive oxygen and nitrogen species enables bacteria to rapidly adjust to varying environmental conditions, thereby modulating the expression of critical bacterial genes associated with virulence. These regulatory mechanisms may contribute to the pathogenesis of conditions such as IBD.
[0379] Next, we employed IPHOMED to explore the functional cross-talk between host and microbiome in the IBD setting. To this aim, we conducted fecal microbiome transfers from healthy individuals or from GER-UC patients into naive germ-free (GF) mice (Methods, Figure 12A). Analysis of fecal microbiome in GF mice recipients by both metagenomics (Figure 12B) and IPHOMED (Figure 12C) exhibited significant differences recipients of UC-associated microbiome and recipients of healthy microbiomes. IPHOMED signals included hundreds of bacterial proteins produced by multiple species that were not differentially abundant by shotgun metagenomics, and multiple bacterial proteins that were not significantly different at gene level (Figure 12D). In particular, we observed a consistent reduced intensity of proteins produced by Bacteroides and Parabacteroides species in GF recipients of GER-UC-associated microbiomes (Figure 12D). In contrast, Bacteroides xylanisolvens exhibited an over-expression of several bacterial proteins, such as Rag / Sus family nutrient uptake outer membrane proteins (Figure 12E), previously noted to be elevated in GER-UC patients, and induce an increased local IgG response in patients with periodontitis Interestingly, the noted depletion of multiple bacterial proteins, originating from Faecalibacterium prausnitzii, Akkermansia muciniphila, and Blautia wexlerae species in recipient GF recipients of UC-associated microbiome, mirrored the functional microbiome suppression observed in human GER-UC donors (Figure 12F).
[0380] Importantly, the host proteome of recipient mice displayed a striking clustering in response to colonization with microbiomes from GER-UC patients versus GER-HC (Figure 12G), driven by 115 differentially abundant host proteins (Figure 12H), implying a regulatory role of gut bacteria in shaping the intestinal host proteome. To compare the extent by which recipient mouse metagenomic microbiome data, IPHOMED microbiome data, and IPHOMED host data predicted the human donor's diagnosis, we trained a random-forest classifier based on each of the three datasets and quantified their classification accuracy (Methods). Bacterial IPHOMED analysis provided the highest classification accuracy, as compared to bacterial genomics and host IPHOMED analysis (Figure 121). Interestingly, enrichment analysis of differentially abundant recipient host proteins primarily identified pathways related to metabolic rather than immune functions, with a total of 729 bacterial proteins exhibiting differential abundance in recipient mice of UC versus healthy microbiomes (Figure 12J). Similar to the microbiomes of the human donors, an enrichment analysis of bacterial proteins in recipient mice revealed alterations in multiple metabolic pathways related to fermentation and carbohydrate metabolism (Figure 12K). Clinically, naive GF mice colonized with GER-UC-associated microbiome did not exhibit any histopathologic evidence of colitis, or significant differences in blood levels of inflammatory biomarkers, suggesting that, in the absence of underlying host IBD disease susceptibility factors, the microbiome is likely not able to solely elicit an overt auto-inflammatory response and associated tissue damage. Collectively, these data suggest that the IBD-perturbed gut microbiome may impact host gut metabolic function, rather than merely the expected immune function.
[0381] IPHOMED enables identification of diagnostic and disease-severity biomarkers in human IBD. Given these results, we next utilized IPHOMED for identification and quantification of fecal bacterial and host proteins that could be potentially usable as IBD biomarkers, while enhancing the performance of the clinically used biomarker calprotectin (comprised of S100A8 and S100A9). To this aim, we conducted a feature selection analysis by training a random-forest classifier on proteomic data from six distinct IBD cohorts, aiming to identify specific combinations of bacterial and host proteins that exhibited significant potentials of outperforming calprotectin in diagnostic accuracy and prediction of disease severity (Methods). We began by utilizing publicly available shotgun metagenomic and proteomic data from two Human Microbiome Project (HMP) cohorts, a Massachusetts General Hospital cohort (HMP-MGH, UC=12, CD=34, HC=52) and a Cincinnati Medical Center cohort (HMP-Cin, UC=29, CD=48, HC=45). In these datasets, IPHOMED analysis identified a total of 1,028 and 770 fecal disease-associated host proteins and 13,491 and 10,631 disease-associated bacterial proteins in the HMP-MGH and HMP-Cin cohorts, respectively (Figure 13A). Training cross-validated random-forest classifiers with the aim of feature selection revealed a high accuracy in diagnosing individuals as either being HC or suffering from IBD, with a median area under the curve (AUC) exceeding 0.9 in both cohorts (Figure 13B, repeated five-fold cross-validation random-forest). Interestingly, the majority of highly IBD-predictive human proteins (ELANE, LTF, PRTN3, CTSG, APCS, MPO, CEACAM8, PGLYRP1, SA1008-9) were not CD- or UC-specific, whereas highly IBD-predictive bacterial proteins were mostly specific to either CD or UC. Notably, the most discriminative proteins predicting a CD diagnosis consisted of combinations of host and bacterial proteins (Figure 16A), whereas those best predicting UC diagnosis predominantly consisted of host proteins. To investigate the respective contributions of the global host and bacterial proteomes to IBD diagnosis, we conducted a random-forest analysis separately for host and bacterial proteins. Remarkably, both sets of proteins exhibited a similarly high diagnostic performance, emphasizing the potential of combining host and bacterial proteins for enhanced diagnostic accuracy in IBD.
[0382] To corroborate and further refine the individual diagnostic accuracy capabilities of IPHOMED’s top 50 discriminating features, we utilized the data of the GER-UC cohort, the ISR- pCD cohort and another HMP cohort, the Cedars-Sinai cohort (HMP-Ced, CD=60, UC=54). In short, we generated random combinations of 1 or 2 discriminatory proteins and assessed their performance in diagnostic accuracy in comparison to calprotectin by repeated five-fold cross- validation random-forest models independently in each cohort. Notably, we identified lactotransferrin (LTF, also known as lactoferrin) in multiple combinations with other host proteins or bacterial proteins, to feature superior performance to calprotectin in CD diagnosis, in agreement with previous studies (Figure 13C). Fecal lactoferrin, a major biomarker of intestinal mucosal neutrophils, increases proportionally with neutrophil degranulation during inflammation and has been suggested as a diagnostic tool to monitor intestinal inflammation. Interestingly, the combination of lactoferrin and flagellin (G1H02-140), produced by Roseburia hominis, exhibited one of the highest accuracies (Figure 13C). This bacterial protein, decreased in IBD patients, is associated with the expansion of regulatory T cells via a Toll-like receptor 5 (TLR5)-dependent mechanism, crucial for host defense and IBD protection. In UC, we identified several pairs of proteins, combining bacterial and host proteins, that outperformed calprotectin for diagnosis (Figure 13D). Notably, most of these combinations included bacterial proteins produced by Alistipes putredinis, significantly depleted in UC patients (Figure 11G, 13D). Several of these pairs included two distinct host proteins, LTF and myeloblastin (PRTN3). In particular, the myeloblastin, a serine protease primarily produced by neutrophils, was observed to exhibit significantly higher levels in colonic biopsies from UC patients compared to control samples.
[0383] To further refine our biomarker search, considering the functional redundancy of microbiome proteins, wherein multiple proteins from different bacterial species perform similar activities, we merged the abundances of all bacterial proteins based on KEGG orthology (KO) and conducted a similar biomarker analysis (Methods). Interestingly, we identified one combination, including LTF and the bacterial functional entity K01006 (ppdK), a central glycolytic enzyme, that outperformed calprotectin for diagnosis of both UC and CD (Figure 17A-B). Some IBD patients present with intermediate colitis, an IBD subtype in which the distinction between CD and UC cannot be made, despite the fact that such distinction has important implications on medical management. Identifying a biomarker with specificity to just CD or UC can facilitate treatment decisions and improve clinical outcomes. Using a similar IPHOMED-based biomarker analysis, including bacterial proteins or KO terms, we detected several combinations enhancing differentiation between adult CD and UC, including one pair of proteins, haptoglobin (HP) and type 1 glyeraldehyde-3 phosphate dehydrogenase (RS04585, Dysgonomonas mossii) that features an average area under the curve of 0.78 (Figure 13E). Indeed, ELISA-based quantification of haptoglobin, a protein with immunomodulatory and antimicrobial properties that binds to hemoglobin, revealed that it is significantly more abundant in UC patients compared to those with CD.
[0384] To test disease severity prediction capabilities of IPHOMED’s top 50 discriminating features for CD and UC, we used the data from a previously published IBD cohort by Mills et al Nat Microbiol 7, 262-276 (2022) (UC=63, CD=119), which included colonoscopy-based severity scores (UC and CD Endoscopic Index of Severity; UCEIS and CDEIS). Similarly, we assessed the performance of the above-mentioned protein pairs for severity prediction, in comparison to calprotectin, by support vector regression analysis independently in each cohort (repeated five-fold cross-validations, Figure 13F-G, 17D-E). Among them, the combination of Annexin A3 (ANXA3) and the bacterial KOO133 (aspartate-semialdehyde dehydrogenase) was correlated with the Ulcerative Colitis Endoscopic Index of Severity (UCEIS) with a coefficient of determination (R2) of 0.63 (Figure 13G). We next asked which bacterial taxa are capable of producing KOO133 in our multi-cohort human sample set. Eubacterium rectale and Bacteroides vulgatus contribute to this functional activity across all cohorts, with Akkermansia muciniphila exclusively detected in the Mills cohort, and Bacteroides fragilis found in the GER-UC and ISR-pCD cohorts (Figure 13H). Repeat analyses using combinations of three different proteins produced by the host or the microbiome demonstrated similar performances to that achieved by two protein combinations (Figure 18A-E, FIG. 19A-F), except for a combination of haptoglobin (HP) and two bacterial proteins produced by Bacteroides stercoris and Clostridium clostridioforme that exhibited superior performance in distinguishing between UC and CD (AUC=0.827, Figure 13E, 18E).
[0385] Fecal IPHOMED assessment of human dietary exposure and absorption patterns in health and IBD. We next assessed whether uncharacterized proteomic signal could be annotated to dietary proteins, using a dietary protein reference database developed and validated in animal models (see Example 1). To validate the accuracy of IPHOMED-based dietary assessment in humans, we began by enrolling six healthy volunteers (Methods) into an interventional dietary trial involving a controlled consumption of two pairs of plant-based foods: bananas (a minimum of 300 grams) and apples (a minimum of 400 grams), or cucumbers (a minimum of 350 grams) and peanuts (a minimum of 100 grams), allowed only during separate four-days enrichment periods, in addition to participant’s regular dietary regimens (Figure 14A). A fecal IPHOMED analysis accurately reflected the dietary exposure, by exclusively detecting the four dietary elements in their respective consumption periods (Figure 14B, 14D). In contrast, mapping fecal DNA reads to the dietary taxonomic database in a metagenomic analysis of the same samples featured low and nonspecific dietary compound detection capacity (Figure 14C, 14E), indicating that genomic readouts are insufficient in providing dietary insights with fecal samples. In addition to the test foods, multiple consumed food items could be detected in stool, reflecting varying nutritional exposures experienced by participants during the documented time frame (Figure 14F). To assess the sensitivity of IPHOMED-based nutritional assessment, we analyzed a cohort of healthy human participants (n=15) recruited as part of a published clinical study that reported all food items consumed during a three-day follow-up IPHOMED-based dietary assessment of stool samples identified protein evidence of reportedly consumed food items in more than 85% of instances (Figure 14G).
[0386] We next used IPHOMED to profile the global dietary landscape in several human cohorts (n=547 in total), which revealed a substantial variability in nutritional composition, likely reflecting diverse geographical and cultural dietary habits between cohorts (Figure 15A). Indeed, IPHOMED-based nutritional assessment could accurately capture population-based food exposure differences, driven by distinct cultural differences in nutritional consumption habits. In an IPHOMED-based comparison of dietary exposure of a cohort from Israel (n=54) and Germany (n=100), we equally detected traces of wheat-containing foods in fecal samples from inhabitants of Germany and Israel; However, only samples from Germans featured pork-related meat products, while samples from Israelis (both Jews and Muslims) did not feature pork-related IPHOMED signals (Figure 15B). This difference is likely explained by the fact that Jews and Muslims (the vast majority of Israelis) generally refrain from such products due to religious restrictions. Importantly, this striking difference stemmed from higher abundance of six different pork-related proteins (Figure 15C), highlighting the sensitivity and specificity of IPHOMED in robustly quantifying dietary exposure. Furthermore, the absence of pork-related protein exposure among Israeli individuals was partially ‘compensated’ by a higher exposure of chicken-derived proteins, as compared to the German counterparts (Figure 15D).
[0387] Further evidence for the utility of IPHOMED in assessing disease-related nutritional changes and their host ramifications could be seen in pediatric subjects with IBD, in which the first-line anti-inflammatory treatment often consists of dietary intervention of exclusive enteral nutrition (EEN). Fecal IPHOMED nutritional assessment of samples from pediatric patients (n=7) diagnosed with new-onset CD before and following the commencement of EEN (Methods), detected a marked reduction of fecal dietary proteins (Figure 15E-F), reflecting the adherence to the amino acid- and dipeptide-based EEN (both not detectable by IPHOMED). Importantly, a significant positive correlation was noted between the number of dietary proteins detected in stools of pediatric CD patients and the levels of fecal inflammatory markers (i.e., S100A9, Figure 15G), suggesting that the level of adherence to EEN is directly associated with its intestinal antiinflammatory activity. Likewise, IPHOMED accurately captured the dietary changes occurring during initiation of a gluten-free diet in children diagnosed with Celiac disease, while accurately detecting one child not adhering to a gluten-free diet despite medical recommendation (Figure 15H). Overall, these results suggest that IPHOMED can provide insights into population and individual dietary habits and can be utilized in assessing compliance to human dietary interventions.
[0388] Finally, and given our findings in animal models (Figure 7) we assessed the capacity of IPHOMED to quantify some features of human nutritional absorption, a critical intestinal function impacting host metabolic and immune homeostasis, which remains clinically elusive by currently- employed non-invasive clinical sampling. Absorptive disruption may occur during small intestinal inflammation and tissue damage occurring in CD enteritis, but not UC that exclusively involves the large intestine. We therefore performed a dietary IPHOMED analysis of two IBD HMP cohorts, the HMP-Cin cohort that features IBD patients with an upper GI involvement, and the HMP-MGH cohort which only features patients with lower intestinal involvement according to endoscopic classification (Figure 151). Indeed, in the HMP-Cin cohort, but not the HMP-MGH cohort (Figure 15J-K), a significant difference in the total dietary protein intensity was noted between IBD patients and controls, potentially indicating an upper intestinal protein malabsorption leading to fecal dietary protein accumulation. In agreement, within the HMP-Cin cohort, IBD patients who were reported to suffer from upper intestinal involvement were notable for a significantly higher total abundance of IPHOMED-based fecal dietary proteins compared to patients shown by endoscopy to only feature ileocolonic involvement (Figure 15L).
[0389] EXAMPLE 3
[0390] Corroboration of results (from Examples 1 and 2)
[0391] A new analysis of total dietary signal between casein and AA diet was performed. Figure 20 illustrates that IPHOME exclusively detected the presence of alpha-S 1 casein protein in mice receiving casein-only diet, featured a complete loss of signal in mice consuming an elemental diet.”
[0392] An additional mouse experiment was performed, combining normal diet and proteinshake gavage. We exclusively detected pea (Pisum) and rice (Oryza) in mice administered a protein supplement comprised from these two genera. As illustrated in Figure 21, IPHOMED exclusively detected pea (Pisum) and rice (Oryza) in mice administered a protein supplement comprised from these two genera.
[0393] The percentage of pepsin-trypsin peptides in murine gastric samples between HFD vs NC. Was analyzed. In IPHOMED-processed gastric samples of HFD-fed mice, we observed a significantly lower percentage of peptides exhibiting a fully-trypsinized profile, coupled with a higher percentage of peptides featuring a pepsin-trypsin profile (Figure 22), suggesting that HFD-based protein is increasingly digested in the stomach and rapidly absorbed thereafter.
[0394] A new targeted AMP analysis was carried out, identifying a dynamic interaction between AMPs and bacterial composition and function in a DSS colitis mouse model. Figure 23 illustrates that host AMPs, measured by targeted proteomics, significantly correlated with bacterial species abundances. Heatmap depicts the linear correlation coefficient. Importantly, the dynamic changes in the abundance of individual AMPs were tightly correlated with the abundance of specific bacterial species, as quantified by MIM. For example, the decreased abundance of Regl, BPI fold-containing family A member 2 (Bpifa2) and lEC-produced AMPs (Ang4, alpha-defensins), coupled with the increased abundance of polymorphonuclear- neutrophils -produced AMPs (Ecn2, Etf, Ngp, Mpo, Camp) and other AMPs (Apes, Ponl, Mbl, Ctsb, Vtn, Dmbtl), were predictive of the loss of commensal bacterial species and the blooming of colitis-associated species (Figure 23). Further correlations were identified between discrete host AMPs and bacterial function (for example, in Eachnospiraceae and Muribaculaceae, as quantified by their respective protein production dynamics.
[0395] A reanalysis of dietary signal in HMP IBD cohorts (MGH and Cincinnati) was carried out, comparing CD patients and healthy controls, showing impaired digestion / absorption functions in CD patients. To display this capacity, we performed a dietary fecal MIM analysis of two similarly processed IBD HMP cohorts5, the Cincinnati Medical Center cohort (HMP-Cin, UC=29, CD=48, HC=45)5, that features CD patients with an upper GIT involvement, and the HMP-MGH cohort which only features CD patients with lower intestinal involvement according to endoscopic classification. Indeed, in the HMP-Cin cohort, but not the HMP-MGH cohort (Figure 24A), a significant enhancement of the total fecal dietary protein intensity was noted in IBD patients compared to healthy controls, potentially indicating an upper intestinal small intestinal malabsorption leading to fecal dietary protein accumulation. In agreement, even within the HMP-Cin cohort, CD patients who were reported to suffer from upper intestinal involvement were notable for a significantly higher total abundance of MIM-based fecal dietary proteins compared to patients shown by endoscopy to only feature ileocolonic involvement (Figure 24B).
[0396] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims. It is the intent of the applicant(s) that all publications, patents and patent applications referred to in this specification are to be incorporated in their entirety by reference into the specification, as if each individual publication, patent or patent application was specifically and individually noted when referenced that it is to be incorporated herein by reference. In addition, citation or identification of any reference in this application shall not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not be construed as necessarily limiting. In addition, any priority document(s) of this application is / are hereby incorporated herein by reference in its / their entirety.
Claims
WHAT IS CLAIMED IS:
1. A method of generating a microbial protein database comprising:(a) sequencing at least one microbiome sample of at least one host subject species to obtain metagenomic data;(b) aligning genomic sequences of said metagenomic data to reference genomes to obtain a database of microbial species that are potentially present in the at least one microbiome sample, and their relative abundance;(c) refining the database by:(i) selecting those microbes having a species abundance above a predetermined level; and(ii) selecting those microbes whose genomes have a % coverage of greater than a first predetermined % to obtain an initially refined database;(d) obtaining protein sequences of said microbes of the initially refined database and align to said metagenomic data; and(e) selecting those proteins having a % sequence alignment coverage greater than a second predetermined % to obtain a protein database which contains microbial proteins potentially produced by microbes in the at least one microbiome sample.
2. The method of claim 1, wherein said at least one microbiome sample is selected from the group consisting of a fecal sample, a gastrointestinal sample, a saliva sample, a blood sample, a urine sample, a tissue sample, a tumor sample and cerebrospinal fluid.
3. The method of claim 1, wherein said at least one host subject species is a mammalian subject species.
4. The method of claim 3, wherein said mammalian subject species has a disease or disorder.
5. The method of claims 3 or 4, wherein said at least one microbiome sample comprises at least two microbiome samples each derived from a different mammalian subject of the same species and having the same disease or disorder.
6. The method of claim 3, wherein said at least one mammalian subject species is healthy.
7. The method of claim 6, wherein said at least one microbiome sample comprises at least two microbiome samples each derived from a different mammalian subject of the same species and each being healthy.
8. The method of claims 1 or 2, wherein said at least one microbiome sample comprises at least two microbiome samples, a first sample of said at least two microbiome samples derived from a mammalian species having a disease or disorder and a second sample of said at least two microbiome samples derived from a mammalian species being healthy.
9. The method of any one of claims 4, 5, or 8, wherein said disease or disorder is a gastrointestinal disease.
10. The method of claim 9, wherein said disease is inflammatory bowel disease (IBD).
11. The method of claim 9, wherein said disease is an allergic disease.
12. The method of claim 11, wherein said allergic disease is Celiac disease.
13. The method of any one of claims 1-12, further comprising adding protein sequences of said host subject to said refined protein database.
14. The method of any one of claims 1-13, further comprising adding protein sequences of proteins in a diet of said host subject species to said refined protein database.
15. A method of profiling a microbiome sample of a host subject comprising:(a) generating a protein database according to any one of claims 1-12;(b) performing a mass spectrometry (MS) analysis on a protein extract of said microbiome sample; and(c) performing a database search to determine a taxonomical origin of peptides generated by said MS analysis, using the database generated in step (a).
16. The method of claim 15, wherein said taxonomical origin is a bacterial, viral, fungal, parasite or archaeal origin.
17. The method of claim 15, wherein said microbiome sample is selected from the group consisting of a fecal sample, a gastrointestinal sample, a saliva sample, a blood sample, a urine sample, a tissue sample, a tumor sample and cerebrospinal fluid.
18. The method of any one of claims 15-17, further comprising adding protein sequences of said host subject species to said refined protein database following step (a) and prior to step (b).
19. The method of any one of claims 15-18, further comprising adding protein sequences of proteins in a diet of said host subject species to said refined protein database following step (a) and prior to step (b).
20. The method of claims 18 or 19, further comprising determine an origin of the peptides, wherein the origin is selected from the group consisting of a microbial origin, a food origin and a host subject origin.
21. A method of obtaining an indication for a subject suspected of having a disease comprising:(a) performing a mass spectrometry (MS) analysis on a protein extract of a microbiome sample of the subject to identify sequences of peptides present in said protein extract;(b) performing a database search to identify the protein source of said peptides and quantify an abundance of said protein source, said database being generated according to the method of any one of claims 1-14; and(c) comparing said abundance of said proteins identified in step (b) to an abundance in a protein extract of a microbiome sample of a healthy control subject, wherein a change in said abundance is indicative of the disease.
22. The method of claim 21, further comprising classifying a microbial, host or food origin of said peptides identified in step (b).
23. A method of analyzing whether a subject is complying with a dietary intervention comprising:(a) performing a mass spectrometry (MS) analysis on a protein extract of a gastrointestinal microbiome sample of the subject to identify sequences of peptides present in said protein extract; and(b) performing a database search to identify the protein source of said peptides and quantify an abundance of said protein source, said database being generated according to the method of any one of claims 1-14;(c) classifying the protein source to a microbial, host or diet origin;(d) analyzing said abundance of identified proteins that correspond to said proteins of said diet origin; wherein an amount or presence of said identified protein which is associated with the dietary intervention in said protein extract is indicative as to whether the subject is complying with the dietary intervention.
24. A method of analyzing gut barrier function of a subject comprising:(a) performing a mass spectrometry (MS) analysis on a protein extract of a blood sample of the subject to identify sequences of peptides present in said protein extract; and(b) performing a database search to identify the protein source of said peptides and quantify an abundance of said protein source, said database being generated according to the method of any one of claims 1-14;(c) classifying the proteins to a microbial, host or diet origin;(d) analyzing said abundance of identified proteins that correspond to said proteins of said subject species; wherein an increased amount or presence of said identified protein compared to control values is indicative as to whether the subject has an impaired gut barrier function.
25. A method of analyzing small intestinal absorptive / digestive function comprising:(a) performing a mass spectrometry (MS) analysis on a protein extract of a gastrointestinal sample of the subject to identify sequences of peptides present in said protein extract; and(b) performing a database search to identify the protein source of said peptides and quantify an abundance of said protein source, said database being generated according to the method of any one of claims 1-14;(c) classifying the proteins to a microbial, host or diet origin; and(d) analyzing said abundance of identified proteins that correspond to said proteins of said subject species; wherein an increased amount or presence of said identified dietary protein compared to control values is indicative as to whether the subject has an impaired small intestinal absorptive / digestive function.
26. A computer readable medium comprising computer readable instructions stored thereon that when run on a computer perform a method according to any one of claims 1 to 25.
27. A system, for producing a protein database for a microbiome sample comprising: one or more processors which executes the method of any one of claims 1-25.
28. A method of diagnosing an inflammatory bowel disease (IBD) of a subject comprising measuring in a gut microbiome sample of the subject an amount of at least one bacteria selected from the group consisting of Alistipes putredinis, Dialister succinatiphilus, Oscillibacter sp. ER4, Romboutsia sp. DR1, Gemmiger formicilis, Faecalibacterium prausnitzii, Fusicatenibacter saccharivorans, Ruminococcus callidus and Bacteroides coprocola, wherein a decrease in the amount of said at least one bacteria as compared to an amount in a gut microbiome sample of a healthy subject is indicative of the subject having IBD.
29. A method of diagnosing Crohn’s disease (CD) of a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 1, wherein an amount of said proteins is indicative of the subject having CD.
30. A method of diagnosing ulcerative colitis (UC) of a subject comprising measuring in a gut microbiome sample of the subject for an amount of a combination of proteins set forth in a row of Table 2, wherein an amount of said proteins is indicative of the subject having UC.
31. A method of distinguishing between UC and CD in a subject comprising measuring in a gut microbiome sample of the subject for an amount of a combination of proteins set forth in a row of Table 3, wherein an amount of said proteins is indicative as to whether the subject has UC or CD.
32. A method of determining the severity of CD in a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 4, wherein an amount of said proteins is indicative as to the severity of the CD.
33. A method of determining the severity of UC in a subject comprising measuring in a gut microbiome sample of the subject an amount of a combination of proteins set forth in a row of Table 5, wherein an amount of said proteins is indicative as to the severity of the UC.
34. The method of any one of claims 28-33, further comprising treating the subject according to the results of the analyzing.
35. The method of any one of claims 28-33, wherein said measuring is effected on the protein level.
36. The method of any one of claims 28-33, wherein said measuring is effected on the RNA level.
37. A method of treating IBD in a subject in need thereof comprising administering to the subject a therapeutically effective amount of at least one bacteria selected from the group consisting of Alistipes putredinis, Dialister succinatiphilus, Oscillibacter sp. ER4, Romboutsia sp. DR1, Gemmiger formicilis, Faecalibacterium prausnitzii, Fusicatenibacter saccharivorans, Ruminococcus callidus and Bacteroides coprocola, thereby treating the IBD.
Citation Information
Patent Citations
Computational method for mapping peptides to proteins using sequencing data
US20130338932A1
Heat sealable laminated propylene polymer packaging material
US4230767A
Reagents and method employing channeling
US4233402A
Macromolecular environment control in specific receptor assays
US4275149A
Immunometric assays using monoclonal antibodies
US4376110A