Compositions and methods for improving microbial fitness in the mammalian gut

A recombinant bacterial expression vector with YigZ and TrhP genes enhances GI tract colonization, addressing the challenge of identifying colonization factors for diverse bacteria, improving gut health and therapeutic applications.

WO2026096887A1PCT designated stage Publication Date: 2026-05-07THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK
Filing Date
2025-10-31
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing technologies have limited understanding of how bacteria colonize the mammalian gastrointestinal tract and lack effective methods to systematically identify and utilize colonization factors for therapeutic applications, particularly for diverse and genetically intractable species.

Method used

A recombinant bacterial expression vector containing genes encoding gastrointestinal tract colonization factors, such as IMPACT family member YigZ and tRNA hydroxylation protein P (TrhP), is introduced to enhance bacterial colonization in the GI tract, using a computational approach to identify these factors across diverse organisms.

Benefits of technology

Enhances bacterial colonization in the GI tract, improving gut health and therapeutic efficacy for disorders like inflammatory bowel disease, and allows for the development of probiotic compositions to modulate the gut microbiome effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000067_0001
    Figure IMGF000067_0001
  • Figure IMGF000070_0001
    Figure IMGF000070_0001
  • Figure IMGF000075_0001
    Figure IMGF000075_0001
Patent Text Reader

Abstract

The present disclosure provides for systematic method of identifying genes that enhance bacterial fitness in the mammalian gastrointestinal (GI) tract. The identified genes can then be used to generate recombinant bacteria which may improve the health of a mammal, such as a human, or may be used to combat a gastrointestinal disorder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] COMPOSITIONS AND METHODS FOR IMPROVING MICROBIAL FITNESS

[0002] IN THE MAMMALIAN GUT

[0003] CROSS REFERENCE TO RELATED APPLICATION

[0004] This application claims priority to U.S. Provisional Application No. 63 / 714,313 filed on October 31, 2024, which is incorporated herein by reference in its entirety.

[0005] GOVERNMENT LICENSE RIGHTS

[0006] This invention was made with government support under AI077562 awarded by the National Institutes of Health. The government has certain rights in the invention.

[0007] FIELD OF THE INVENTION

[0008] The present invention relates to methods and compositions for improving the health of a mammal by modulating the population of bacteria in the mammalian gastrointestinal (GI) tract. In particular, the present invention relates to promoting the growth of beneficial gut bacteria.

[0009] BACKGROUND

[0010] Microbial organisms inhabit a vast variety of habitats. The molecular mechanisms by which individual bacteria achieve high fitness in such diverse and specific environments remain largely unexplored. The mammalian gastrointestinal (GI) tract is a complex and dynamic environment colonized by a diverse array of microbes. The mechanisms by which bacteria colonize the GI tract remain largely unknown. Understanding the genetic factors that influence the colonization of commensal organisms is crucial not only for elucidating the mechanisms underlying gut microbiome dysbiosis in disease states11 14, but also for developing effective live biotherapeutic products (LBPs) with robust colonization capabilities to achieve desired clinical efficacy15 16. However, studies systematically identifying colonization factors have been limited to only a few species17-26, mostly within the Bacteroides genus. Expanding these efforts to include a wider range of commensals, particularly those that are unendurable or genetically intractable, presents a significant challenge.

[0011] Bacteria that colonize the human gastrointestinal (GI) tract have a wide variety of effects on their hosts. Core features exist in the microbial communities that colonize the human body (Rooks et al., Gut microbiota, metabolites and host immunity. Nat Rev Immunol. 2016; 16(6): pp. 341-52). In recent years, research has shown a clear link between gut microbiome composition and overall health, making gut microbiome optimization a therapeutic goal for various diseases.

[0012] Developing novel and effective therapeutics for the gut connected systems will require an understanding of gut bacteria colonization at the level of individual microbial genes (Bradley et al., Phylogeny-corrected identification of microbial gene families relevant to human gut colonization. PLoS Comput Biol. 2018; 14(8): el006242.) Metagenomic sequencing helps reveal interactions between microbial metabolism and host.

[0013] As we aspire to engineer microbes and microbial ecosystems to benefit human health1 10and environmental sustainability, we need next-generation approaches to determine the genetic basis of niche colonization systematically and rapidly across the biosphere.

[0014] Computational identification of microbial genes associated with specific habitat colonization offers a potentially powerful approach for predicting colonization factors. However, this approach has been largely limited to narrow analyses of closely related species27,28. With the increasing metagenomic sampling of diverse habitats29-34and accumulation of metagenome-assembled genomes (MAGs), it is now finally feasible to perform comparative genomic analyses across the tree-of-life to infer potential common colonization factors among diverse organisms. Despite this progress, such analyses remain challenging due to the extremely high taxonomic31,32and genomic diversity35and the large fraction of proteins of unknown function36.

[0015] SUMMARY

[0016] The present disclosure provides for a recombinant bacterial expression vector. The recombinant bacterial expression vector may comprise a gene encoding a gastrointestinal (GI) tract colonization factor. The colonization factor may comprise IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

[0017] The present disclosure also provides for a method of enhancing gastrointestinal (GI) tract colonization of a bacterium in a mammal. The method may comprise introducing a recombinant bacterial expression vector into the bacterium. The bacterial expression vector may comprise a gene encoding a gastrointestinal (GI) tract colonization factor. The colonization factor may comprise IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

[0018] The present disclosure provides for a method of enhancing gastrointestinal (GI) tract colonization of a bacterial cell in a mammal. The method may comprise introducing a gene encoding a gastrointestinal (GI) tract colonization factor into the bacterial cell. The colonization factor may comprise IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

[0019] YigZ may comprise (or consist essentially of, or consist of) an amino acid sequence at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%. at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, at least or about 98%, at least or about 99%, or about 100%, identical to the amino acid sequence set forth in SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, or SEQ ID NO:5.

[0020] TrhP may comprise (or consist essentially of, or consist of) an amino acid sequence at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, at least or about 98%, at least or about 99%, or about 100%, identical to the amino acid sequence set forth in SEQ ID NO:6.

[0021] The colonization factor may comprise YigZ (GenBank accession number: WP_001295262), YigZ (GenBank accession number: XKS61981), YigZ (UniProt accession number: A0A833CEL3), YigZ (UniProt accession number: A0A430FQ87), YigZ (UniProt accession number: A0A174QGE9), TrhP (GenBank accession number: NP_416585), or combinations thereof.

[0022] The gene may be under the control of a constitutive promoter, or an inducible promoter.

[0023] The colonization factor may be wild-type or a mutant.

[0024] The present disclosure provides for a recombinant bacterium comprising the bacterial expression vector.

[0025] The present disclosure provides for a recombinant bacterium comprising a bacterial expression vector. The bacterial expression vector may comprise a gene encoding a gastrointestinal (GI) tract colonization factor. The colonization factor may comprise IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

[0026] The present disclosure also provides for a recombinant bacterium comprising a gene encoding a gastrointestinal (GI) tract colonization factor. The colonization factor may comprise IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

[0027] The colonization factor may be endogenous to the bacterium, where the bacterium overexpresses the colonization factor.

[0028] The colonization factor may be heterologous or exogenous to the bacterium.

[0029] The gene may be integrated into a bacterial chromosome or may be episomal. The (recombinant) bacterium may belong to phyla Bacteroidetes, Firmicutes, Proteobacteria, or Actinobacteria.

[0030] The (recombinant) bacterium may belong to genera Bacteroides, Clostridium, Escherichia, Bacillus, Lactobacillus, or Bifidobacterium.

[0031] The (recombinant) bacterium may be Escherichia coli, Lactobacillus gasseri, Bifidobacterium dolichotidis, Bacteroides thetaiotaomicron, or Salmonella enterica.

[0032] The (recombinant) bacterium may be Bacillus subtilis, Bifidobacterium longum, E. coli Nissle 1917, Lactobacillus plantarum, Lactobacillus lactis, Lactobacillus casei, Lactobacillus reuteri, Lactobacillus gasseri, Streptococcus typhimurium, Bacteroides thetaiotaomicron, Bacteroides fragilis, Bacteroides ovatus, or Akkermansia muciniphila.

[0033] The (recombinant) bacterium may be Escherichia coli (E. coli). Bacillus subtilis (B. subtilis), Lactobacillus plantarum (L. plantarum), Lactobacillus reuteri (L. reuteri), Lactobacillus rhamnosus (L. rhamnosus), Bifidobacterium longum (B. longum), Bacteroides thetaiotaomicron, Clostridium butyricum, Bacteroides fragilis, Bacteroides melaninogenicus, Bacteroides oralis, Bacteroides amylophilus, Clostridium butyricum, Clostridium perfringens, Clostridium tetani, or Clostridium septicum.

[0034] The (recombinant) bacterium may be from a natural gut microbiota. The natural gut microbiota may be from the gut of a healthy mammal, or from the gut of a mammal with inflamed GI tract. The mammal with inflamed GI tract may have an inflammatory bowel disease (IBD) or irritable bowel syndrome (IBS). The IBD may be ulcerative colitis or Crohn's disease. The healthy mammal, or the mammal with inflamed GI tract, may be a human subject.

[0035] The (recombinant) bacterium may be a spore-forming bacterium.

[0036] The (recombinant) bacterium may be a human commensal bacterium.

[0037] The present disclosure provides for a pharmaceutical composition comprising the recombinant bacterial expression vector.

[0038] The present disclosure provides for a pharmaceutical composition or a probiotic composition comprising the recombinant bacterium.

[0039] The pharmaceutical composition may be formulated for oral administration.

[0040] The pharmaceutical composition may be a food composition, a beverage composition, or a feedstuff composition. The pharmaceutical composition may be a dairy product.

[0041] The pharmaceutical composition or probiotic composition may be formulated for rectal administration. Also encompassed by the present disclosure is a method of treating a disorder in a subject.

[0042] The present disclosure provides for a method for treating a subject in need of a probiotic treatment, a method of regulating the immune system of a subject, a method of improving digestion in a subject, a method of maintaining or improving health of a subject, and / or a method of maintaining or improving health of a gastrointestinal tract of a subject.

[0043] The present method may comprise administering the pharmaceutical composition or probiotic composition to the subject. The method may comprise administering an effective amount of the pharmaceutical composition or probiotic composition to the subject.

[0044] The method may comprise administering the recombinant bacterium to the subject. The method may comprise administering an effective amount of the recombinant bacterium to the subject.

[0045] Administration of the pharmaceutical composition or probiotic composition to the subject may result in colonization of the gastrointestinal tract of the subject by the recombinant bacterium.

[0046] Administration of the recombinant bacterium to the subject may result in colonization of the gastrointestinal tract of the subject by the recombinant bacterium.

[0047] Colonization of the gastrointestinal tract of the subject by the recombinant bacterium may be for at least or about 1 day, at least or about 2 days, at least or about 3 days, at least or about 4 days, at least or about 5 days, at least or about 6 days, at least or about 7 days, at least or about 8 days, at least or about 9 days, at least or about 10 days, at least or about 11 days, at least or about 12 days, at least or about 13 days, at least or about 2 weeks, at least or about 3 weeks, at least or about 4 weeks, at least or about 1 month, at least or about 2 months, or longer.

[0048] The disorder may be an infection or an infectious disease, an autoimmune disease, an allergic disease or cancer.

[0049] The autoimmune disease may comprise organ transplant rejection, inflammatory bowel disease (IBD), ulcerative colitis, Crohn's disease, sprue, rheumatoid arthritis, Type I diabetes, graft versus host disease, or multiple sclerosis.

[0050] The infection may be a bacterial infection or viral infection.

[0051] The infection or infectious disease may be associated with infectious pathogens comprising, Salmonella, Shigella, Clostridium difficile, Mycobacterium, protozoa, filarial nematodes, Schistosoma, Toxoplasma, Leishmania, hepatitis C virus (HCV), hepatitis B virus (HBV). and / or herpes simplex viruses.

[0052] The subject may be a mammal. The mammal may be a primate, a bovine, an equine, a canine or a feline. The subject may be a human.

[0053] The present disclosure provides for a method of identifying colonization factors for bacteria to colonize the mammalian gastrointestinal (GI) tract. The method may comprise:

[0054] (a) collecting mouse gastrointestinal (GI) metagenome-assembled genomes (MAGs), human GI MAGs, and human isolates;

[0055] (b) aligning protein sequences from host-associated dataset against protein sequences from both host-associated dataset and environmental datasets to create a phylogenetic profile, denoting homolog presence across genomes;

[0056] (c) analyzing genotype-phenotype association on each of the mouse GI MAGs, human GI MAGs and human isolates against their respective matched environmental genomes, by ranking proteins based on their mutual information Z-score (MI-Z) and correlating habitat phenotype with their phylogenetic profile; and

[0057] (d) identifying colonization factors (CF) by (1) combining top proteins from the genotypephenotype association analyses of the mouse GI MAGs, human GI MAGs and human isolates, and (2) consolidating the top proteins into homologous groups based on protein sequence similarity.

[0058] BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figures 1A-1F. Identification of microbial colonization factors across mouse and human gut microbes. (A) Top panel shows the schematic of computational pipeline. See Figures 8A-8F and Methods for details. Bottom panel shows steps for dataset- specific (steps 1 and 2) and combined (steps 3 and 4) analyses. (B) Distribution of mutual information z-scores (MI-Z) of the 8.15, 2.72, 2.58 million proteins from mouse gut MAGs, human gut MAGs, and human microbial isolates, showing their associations with gut colonization phenotype. For each dataset, top 0.2% proteins with highest MI-Z were selected for further analysis with text indicate the number of proteins selected. Dashed lines indicate MI-Z cutoff. (C) The number of CFs that are significantly associated with colonization in three datasets. See Methods on how significant associations of CFs were determined. (D) Prevalence of 63 CFs (highly associated with colonization in mouse or human gut MAGs) across habitats. Panel D compares CF prevalence in environmental MAGs vs mammalian gut MAGs. Gray dots are all 16985 protein families from a random simulation of 27,000 proteins random non-colonization associated protein families (refer to Figure 9F for details). (E) Correlation of mutual information z-scores of 63 CFs in mouse and human gut MAGs (R: Pearson correlation coefficient). (F) Network of 79 CFs showing homologous connections (edge: e-value < IE-10).

[0060] Figures 2A-2I. De novo inferred coinherited colonization modules capture colonization- related biological processes. (A) Phylogenetic profiles (left) and their pairwise coinheritance strength (right) of 79 CFs across the 9,472 genomes (genomes with no CFs are not shown). For phylogenetic profile, gray / white indicate presence / absence of homologs across genomes. The left annotation bars show the CF-dataset relationship (as categorized in Figure 1C). The top annotation bars show the genomes’ metadata, including habitat and type (MAG or isolates). The CMs detailed in panels D-H are highlighted. (B) CM identification workflow. See Methods for details. (C) Number of CMs in each category. ‘Homology’ indicates that all member CFs are homologous to each other as defined in Figure IF. Genomic context (operon, non-operon) is defined based on genomes of three representative species: E. coli K.12, Bacteroides thetaiotaomicron VPI5482, and Clostridium difficile S-0253. (D-H) Example of CMs in representative species. (D-E) operonic CMs: CM1 and CM6; (F-G) non-operonic CMs: CM11 and CM 12. (H) an example of operonic CM with duplication (CM 18). (I) Annotation status of the 63 MAG-associated CFs based on external databases. Refer to Methods and Figure 10D for criteria of annotation status. From top to bottom: whether protein is annotated using: BioCyc, UniRef50 database, BioCyc or UniRef50 database, GO biological process (UniRef50 GO BP) or molecular function (UniRef50 GO MF). Gray bars represent unannotated genes per annotation database.

[0061] Figures 3A-3E. YigZ and TrhP are required for E. coli MP13 colonization. (A) The rank of 79 CFs based on their MI z-score in human gut MAGs. The top three CFs (CF1, CF2_4, CF2_7) correspond to anaerobic RNR pathway. The tRNA modification factors include CF15, 22, 7, 24_29 and 24_46, and other translation-related CFs include CF9, 56 and 69. (B) Differential habitat and host preferences across E. coli strains. (C) Protein sequence comparisons of YigZ, TrhP, TcdA and Ybak across four E. coli strains. The position of amino acid substitutions relative to MG1655 is highlighted with numbers. Shading indicates no substitutions. (D) Study design for evaluating the impact on colonization fitness following deletion of four genes of interest through in vivo competition of MP7 wildtype versus MP 13 wildtype or mutants. (E) LoglO-transformed normalized ratios of wild-type MP13 or its mutants to MP7. determined by colony counting of mCherry and GFP colonies. See Methods for details. Samples of post-gavage day > 12 include intestinal content directly collected from the cecum, colon, and small intestine during mice sacrifice. P-values are calculated using one-sided students’ t-tests: * p<0.05, ** p< 0.01, *** p<0.001.

[0062] Figure 4A-4D: YigZ overexpression is sufficient to enhance gut colonization of E. coli MG1655. (A) Diagram of the genomic barcoding system for simultaneous in vivo fitness profiling of multiple E. coli strains. To introduce the genomic barcode, a dsDNA containing a random 20bp sequence was generated through PCR. This dsDNA was chromosomally integrated into the intergenic locus between rprA and ydiL, which is conserved among multiple E. coli strains, using the lambda-red system94. See Methods and Figures 12A-12B for details. (B) Study design of the gain-of-function (GOF) mouse experiment: barcoded MG1655 strains, each with a unique genomic barcode and GOF plasmid, are co-administered to mice pre-treated with streptomycin. Stool samples are then collected for metagenomic DNA analysis to assess barcode frequencies and selective plating to estimate absolute colonization levels. (C) Design of the variable gene cassette in the GOF plasmids, featuring genes with their native promoters cloned from the E. coli MP 13 genome. Control plasmid contains no additional sequences in this region. The constant region containing the replication origin SC 101 and a chloramphenicol resistance gene are not shown. (D) Normalized frequencies of different strains on post-gavage day 2 (left) and day 3 (right). Statistics were done using Student’s t-tests. p-value: * <0.05, ** <0.01. See Methods for details.

[0063] Figures 5A-5C. Taxonomic distribution and genomic coding pattern of CMs. (A) Distribution of the 35 MAG-associated CMs across major gut microbial classes. Only classes with more than 5 species are shown. The bar plot on the right shows the number of unique species per class. A total of 5,149 host-associated microbes (both MAGs and isolates) are used to calculate species CM profile. (Refer to Methods and Figure 13A for details). TrhP- and YigZ-containing CMs are highlighted with black boxes on the x-axis. (B-C) Total number of unique CMs per genome across genomes collected from different habitats combined (B) or by microbial classes (C). Text reflects median number for each group. P- values are calculated by Wilcoxon rank-sum test with the aquatic and terrestrial genomes combined as the reference group: *<0.01, ** <2.2xl0‘16. The Phylum labels are shown above in panel C.

[0064] Figures 6A-6B. Within-species stability of CMs. (A) Within- species frequencies of CMs. The tile color indicates the within-species frequency. Only species with more than 10 genomes were included in this analysis. The phylogenetic tree is derived from the GTDB bacterial taxonomy tree (r95) and collapsed at the genus level (see Methods for details). TrhP- and YigZ-containing CMs are highlighted with black boxes. (B) Histogram of CM within-species frequencies in C. difficile (n=794 genomes) and E. coli (n=3,533 genomes). Dashed lines on the left and right show CM within-species frequency at 10% and 90%, respectively. See Methods for details.

[0065] Figures 7A-7D. Natural YigZ allelic variation contributes to the colonization differences among E. coli strains. (A) The sequence variation of YigZ residues (entropy) among the 3,112 E. coli isolates collected from the NCBI Pathogen Detection Project (PDP). Entropy values were computed for each residue position in the aligned YigZ sequences, with YigZKi2 as the reference, via the entropy function from the infotheo R package. (B) Structure of YigZKn (PDB: 1 vi7)81with the two frequently variable sites highlighted. (C) In vivo frequency™™ of barcoded MG1655 strains harboring a control plasmid (1stfrom left) or YigZ GOF cassette (PyigZMPi-YigZMPi: 2ndfrom left, PyigZKi2-YigZMPi: 3rdfrom left, PyigZMPi-YigZKi2i 4thfrom left, or PyigZKi2-YigZKi2: 5thfrom left) at post-gavage day 3 (n=12 mice). Each dot reflects the mean frequencynorm of the indicated strains in one mouse (see Methods for frequencynorm calculation). Fold changes relative to the control, indicated at the bottom, were calculated using the median frequencynorm values across all mice. Statistical significance was determined using Wilcoxon rank-sum test with Benjamini-Hochberg multiple comparisons adjustment; p-values: **** <0.0001, ***<0.001, **<0.01, * <0.05. (D) Frequency distribution of YigZ variants based on the two non- synonymous mutations (M25L and H146S) among human-associated and environmentally-associated E. coli isolates. The variants were categorized into YigZMGi655, single mutations (M25L, H146S), and double mutations (M25L+H146S, YigZmpi3). The frequencies for each variant were calculated separately for the human and environmental-associated groups and indicated with text. *p=0.006, by one-sided two-sample equal proportion test.

[0066] Figures 8A-8F. Data selection, computational pipeline and discovery of colonization factors, related to Figure 1. (A) Criteria and selection of 9,475 genomes for genotype-habitat association. Notably, the data selection is done separately for mouse GI MAGs, human GI MAGs and human isolates. *Human isolates include genomes annotated as associated with the gastrointestinal tract and those that are undetermined. See Methods for details. (B-C) The geographical distribution (B) and taxonomic diversity (C) of the 9,475 genomes selected. For genomes without longitude and latitude information, geographical distribution was inferred using country or region information. (D) Detailed computational pipeline. Proteins from genomes were collected and subjected to semi all-against-all alignment (See Methods for detailed information and parameter selection) using Diamond default '-fast' mode. The alignment results were converted to protein PPs. To assess protein-habitat association, the mutual information (MI) and Spearman correlation between protein PPs and the binary habitat vector were calculated. Multiplying MI with the sign of corresponding Spearman correlation coefficient yielded signed MI. The signed MI values were converted to MI z-score (MI-Z) using frequency-matched MI null distributions (See Methods). Proteins with the highest MI-Z (top-associated proteins) were selected and consolidated into protein families based on amino acid sequence similarity as candidate colonization factors (CFs). To assess the candidate CFs more accurately, the PPs of CF were regenerated by aligning these top-associated proteins against the whole proteome under the Diamond '-very-sensitive' mode and candidate CFs with refined MI z-scores above 95% of dataset were kept as colonization factors (See Methods), which are subject for colonization modules (CM) identification based on their coinheritance strength. (E) Hierarchical assignment of 27,096 top-associated proteins identified from three datasets to protein families visualized by sanky plot. The proteins were assigned to protein families using the graph clustering algorithm Markov Cluster (MCL)95with three inflation values (1=1.4, 2.5, and 3), yielding the formation of 71, 79, and 82 protein families respectively. Displayed are the distribution and flow of proteins across three levels of protein family classification. Levels 1 (1=1.4) and 3 (1=3) were used to define colonization factors (CFs). CFs with names that include an underscore indicate they are further divided into subgroups at level 3. (F) The distribution of protein families identified from randomly selecting same number of proteins (n=27,096) 1,000 times. An inflation value of 1=1.4 was used.

[0067] Figures 9A-9J. Reevaluation of CF significance, related to Figure 1. (A) Comparison of the number of obtained protein hits under DIAMOND'S '-fast' versus '-very-sensitive' mode. This analysis is conducted separately for each dataset, as shown in Figure 1 A bottom panel, with each dot representing one top-habitat-associated protein from the corresponding dataset. (B) Methods for summarizing protein-level MI z-scores at the CF level. (C) Correlation of CF MI z-scores using the human MAGs versus the mouse MAGs (left) or the human isolate dataset (right). Dashed lines indicate significance cutoffs (95th percentile MI z-score for all proteins) for the corresponding datasets. Pearson correlation was utilized. (See Methods for details) (D-E) Assessing Pipeline Stability with Varying Alignment Coverage Parameters. (D) Venn diagram showing the overlap of top colonization-associated proteins from human GI MAGs identified with '-query-cover set to 66 (n=6,l 11), 75 (n=6,445), or 80 (n=6,448). This figure reflects an intermediate analysis during pipeline development. (E) Comparison of homologous relationships between proteins identified using non-default parameter ('— query-cover' = 75 or 80) versus the default f -query -cover' = 66). Light gray dots represent the 31 ,746 top proteins identified with the default parameter across human GI MAGs, mouse GI MAGs, and human isolates. Dark gray dots highlight 3.808 top proteins from human GI MAGs identified exclusively with the nondefault parameters, indicating they mostly expand existing CFs, suggesting that the CFs acquired with default parameters are robust to varying alignment coverage. (F-G). Generating non-habitat associated protein families as controls protein families. (F) Process for selecting non-habitat- associated proteins. A total of 27,096 proteins were randomly selected from human MAGs and grouped into protein families using MCL (1=1.4). Protein families with |MI z-score|<50 were considered non- habitat associated. (G) Illustration of the cutoff employed (-50 and 50) among the total MI z-score distribution. (H) Same as Figure ID, with gray dots on the right side of the graph indicating all protein families (n = 16,985) from a random sampling of 20,703 proteins as controls. (I- J) Case study of homologous CFs: CF2_4 and CF2_7. (I) Sequence alignment results of proteins from CF2_4 and CF2_7 against the longest protein from two CFs jgi_bl0634|p953, gem42569|pl388) as the reference protein. Gray darkness represents the percentage of identity for the aligned protein regions, based on DIAMOND aligner. Proteins in CF2_4 exhibit higher sequence identity to the reference protein from CF2_4 (jgi_bl0634|p953, left) compared to the reference protein from CF2_7 (gem42569|pl388, right) and vice versa. (J) PPs (selective mode) of CFs 2_4 and 2_7, demonstrating their distinct taxonomic distribution. The order of 9,475 genomes were based on hierarchical clustering (hclusf) using the selective PPs of the complete set of 79 CFs.

[0068] Figures 10A-10G. Discovery and functional annotation of colonization modules (CM), related to Figure 2. (A) Hierarchical clustering dendrogram of CFs. based on their PPs in permissive mode. Vertical long lines indicating attempted height for R's 'entree' function ( / ?). The final CM discovery in the present study was done using / ?=0.4. (B-C) Example CMs in three representative species. Panels focus on CM2 (B) and CM7 (C). Representative species are E. coli K12, Bacteroides thetaiotaomicron VPI5482, and Clostridium difficile S-0253. (D) Criteria for determining the annotation status of CFs based on the UniRef50 and BioCyc databases. Briefly, all top associated proteins are mapped to both databases to find the best hit. which is then summarized at the CF level, resulting in a CF-level reference protein and percentage of member proteins. If a CF has no hit or if the best hit has no associated protein names in the database, it is classified as unannotated. GO terms were collected through the UniRef50 reference protein and consolidated using the ReViGO webserver (See Methods). For each CF, the consensus GO terms associated with more than 1% of the member proteins are retained. (E) Annotation status of all 79 CFs where count numbers over bars correspond to the number of CFs with annotations. (F-G) The association of CFs with Gene Ontology (GO) terms. Panels focus on GO terms of biological processes (F) and molecular functions (G). GO terms were collected through UniRef50 as described in Figure 10D. Tile colors indicate the percentage of proteins whose best UniRef50 hit is associated with the indicated GO terms. CFs are annotated on the x-axis and highlighted in text labels within each tile. See Methods for detailed information.

[0069] Figures 11A-11E. Functional validation of translation-Associated CFs and in vitro / in vivo phenotypic analyses, related to Figure 3. (A-B) Expanded information regarding translation- related CFs (A) and selection of four Genes for in vivo functional validation (B). Prevalence among E. coli was calculated using 3,533 complete E. coli genome assemblies from the NCBI Pathogen Detection Project (Methods). (C) The ratio between MP13 and MP7 wildtype or mutant in the inoculum. Inoculum was generated by mixing MP7 and the corresponding MP13 strain equally. Inoculum was diluted and a lOOmL volume was plated on selective plate (LB + tetracycline). The number GFP (MP 13) or mcherry (MP7) fluorescing colonies was counted to determine the ratio (see Methods for details). Two to three plates of technical replicates were used for each group. (D) Growth curves of MP wildtype and mutant strains in Luria-Bertani (LB) broth. Doubling time is estimated using Growthcurver R package96. For each strain, a frozen stock was streaked on an LB plate and two single colonies were picked into LB medium for overnight growth. This overnight culture was then diluted 1:200 into a 96- well plate, with two technical replicates per culture. The ODeoo was measured by a plate reader every 10 minutes. (E) Representative image showing small colony phenotype of MP13 Dy baK in mouse 2-1 after in vivo evolution. Mouse fecal-PBS slurry was grown on selective plate (tetracycline) overnight to induce fluorescence: MP13 (GFP), MP7 (mCherry). Left shows heterogeneous colony size of MP13 DybaK observed in mouse 2-1 at post-gavage day 3. This mouse was excluded from analysis and Figure 3E. Right panel shows MP7 and MP 13 wildtypes recovered from control mouse for visual comparison.

[0070] Figures 12A-12F. Design and implementation of genomic barcoding system in multiple E. coli strains, related to Figures 4 and 7. (A) Schematic representation of the barcoding system. Step 1: The conserved intergenic barcode locus (between rprA an ydiL) was identified by systematic analysis of genomes from three E. coli strains (MG1655, Nissle, MP1). Step 2: A PCR product containing a 20bp random barcode and kanR, flanked by homologous arms targeting the rprA and ydiL site, was generated using the pKD4 plasmid and primers (M42, M43). Step 3: The PCR product was integrated into the genomes of E. coli using the lambda-red recombination system. See Methods for details. (B) Successful insertion of the barcode cassette in three E. coli strains MG1655, MP, and Nissle, confirmed by PCR (M28 and W1583). Sanger sequencing confirmed 10 random barcode sequences in the 10 barcoded MG1655 strains shown in panel B. (C) Tracking the absolute abundance of barcoded MG1655 strains in mouse stools. The absolute abundance of each strain was estimated by multiplying its frequency (determined by barcode sequencing) by the total colony-forming units (CFU) counted from spot plating 5 pL of fecal-PBS slurry on selective plates (kanamycin and chloramphenicol). Each panel displays data from one of the 8 mice with successful colonization. (D-E) The frequency of the six bcMG1655 strains in the initial inoculum (D) or after 6 hours of in vitro competition (E). The inoculum was generated by pooling six individually growing strain (early exponential phase) equally based on their ODeoo (see Methods for details). A total of 30ml of the inoculum was inoculated into 50ml LB for in vitro competition for ~ 6 hours (final ODeoo: 0.92). Barcode frequency of each strain was determined (see Methods for details) and shown. Dashed line indicate 17%, which is the hypothetical ratio if the six strains are evenly represented. (F) Sequence comparison of promoter of pepQ-yigZ- trkH-hemG operon encoded in the genomes of E. coli MG1655, MP1, or Nissle (EcN1918), related to Figure 7. Mutations relative to the MG1655 promoter are highlighted.

[0071] Figures 13A-13E, Distribution of CFs / CMs within and across diverse Species, related to Figures 5 and 6. (A) Procedure to determine species-level CFs and CMs presence, related to Figure 5. (B) The relationship between microbial taxonomy distance and CM profile dissimilarity, related to Figure 5. The pairwise CM dissimilarities are calculated using the Euclidian distance between genomes binary CM PP (selective mode). The points reflect the mean CM dissimilarities for all inter- and intra-taxa genome pairs. (C) Histogram showing the distribution of CMs within-species frequency across bacterial species, related to Figure 6. Species with more than 10 genomes are shown. Both MAGs and isolates are included in this analysis. (D-E) Histogram showing the distribution of within-species frequency of 79 CFs across various species with more than 10 genomes available (D) and two representative species (E), related to Figure 6. The presence of CF was determined based on the CF PP in selective mode.

[0072] Figure 14A shows an endogenously induced colonization-enhancing vector. Figure 14B shows an arabinose inducible colonization-enhancing vector. Figure 14C shows a colonization-enhancing vector with a kill switch (suicidal control system). Figure 14D shows an example of a colonizationenhancing vector with a kill switch (suicidal control system).

[0073] DETAILED DESCRIPTION

[0074] The present disclosure provides for a powerful and systematic method of identifying colonization factors or fitness genes (that encode colonization factors) which enhance bacterial fitness in the mammalian gastrointestinal (GI) tract. The identified colonization factors or fitness genes can be used to generate recombinant bacteria to be included in a probiotic composition. The composition may improve the health of a mammal, such as a human, or may be used to combat a gastrointestinal disorder.

[0075] The present method can advance the general understanding of microbiome dynamics by elucidating genetic components of gut bacteria in a streamlined manner. Additionally, the present method allows for the optimization of bacterial colony composition in the gut and thus improves therapeutic approaches to gut health.

[0076] In certain embodiments, the present method uses a computational approach to identify bacterial genes and genetic combinations that enable and enhance colonization in the mammalian GI tract.

[0077] The colonization factor or fitness gene may enhance bacterial fitness in the mammalian GI tract. In certain embodiments, the overexpression of a GI tract colonization factor increases the colonization of a weakly colonizing E. coli strain in the mammalian GI tract by more than a factor of 100.

[0078] The present compositions and methods may be used as research tools for studying bacteria genetics, or as therapeutics targeting the gut microbiome. The present compositions and methods may be used to enhance nutrient absorption or for pathogenesis disruption.

[0079] The present composition can modify the GI tract colonization efficiency of bacterial strains to benefit human / animal health.

[0080] The present disclosure provides for a method of enhancing gastrointestinal (GI) tract colonization of a bacterium in a mammal. The method may comprise introducing a recombinant bacterial expression vector into the bacterium, where the bacterial expression vector comprises a gene encoding a gastrointestinal (GI) tract colonization factor.

[0081] Fitness genes or colonization factors that improve gut colonization can lead to enrichment of the bacterial abundance in the mammalian GI tract (e.g., across the microbial population).

[0082] Also encompassed by the present disclosure are recombinant bacteria (e.g., in a probiotic composition) engineered with genes that can enhance bacterial colonization / fitness in the mammalian GI tract. The abundance (or colonization) of the bacterium overexpressing or expressing the colonization factor may be at least or about 1.5 fold, at least or about 2 fold, at least or about 2.5 fold, at least or about 3 fold, at least or about 3.5 fold, at least or about 4 fold, at least or about 4.5 fold, at least or about 5 fold, at least or about 5.5 fold, at least or about 6 fold, at least or about 6.5 fold, at least or about 7 fold, at least or about 8 fold, at least or about 9 fold, at least or about 10 fold, at least or about 15 fold, at least or about 20 fold, at least or about 25 fold, at least or about 30 fold, at least or about 35 fold, at least or about 40 fold, at least or about 45 fold, at least or about 50 fold, at least or about 55 fold, at least or about 60 fold, at least or about 65 fold, at least or about 70 fold, at least or about 75 fold, at least or about 80 fold, at least or about 85 fold, at least or about 90 fold, at least or about 95 fold, at least or about 100 fold, at least or about 120 fold, at least or about 150 fold, at least or about 175 fold, at least or about 200 fold, at least or about 220 fold, at least or about 250 fold, at least or about 275 fold, at least or about 300 fold, at least or about 320 fold, at least or about 350 fold, at least or about 375 fold, or at least or about 400 fold, of the abundance of the bacteria without the colonization factor, about 12 hours, about 24 hours, about 36 hours, about 48 hours, about 60 hours, about 72 hours, about 1 day, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 1 week, about 2 weeks, about 3 weeks, or about 4 weeks after the administration of the present composition by a subject (e.g., a mammal), compared to the bacterium which does not overexpress the colonization factor or which does not express the colonization factor.

[0083] The present disclosure provides for a method of treating or preventing a disorder (e.g.. an infection such as a bacterial infection or viral infection), or decreasing antibiotic resistance of gastrointestinal microbiome, in a subject, the method comprising administering the present recombinant bacterium or composition to the subject.

[0084] The bacterial infection may be an antibiotic -resistant infection, such as a vancomycin-resistant infection.

[0085] The present method may further comprise administering one or more antibiotics to the subject.

[0086] Classes of antibiotics may include: amikacin, aminoglycosides, carbapenems, cephalosporins (classified based on generation), glycopeptides, fluoroquinolones, lincosamides. macrolides, monobactams, nitroimidazoles, penicillins, penicillins with beta-lactamase inhibitors, polymixin, sulfabased, rifamycin, tetracyclines, and combinations thereof.

[0087] The present disclosure provides for a method of identifying a colonization factor, or a bacterial fitness gene, for bacteria to colonize the mammalian gastrointestinal (GI) tract. The method may contain the following steps:

[0088] (a) collecting mouse gastrointestinal (GI) metagenome-assembled genomes (MAGs), human GI MAGs, and human isolates;

[0089] (b) aligning protein sequences from host-associated dataset against protein sequences from both host-associated dataset and environmental datasets to create a phylogenetic profile, denoting homolog presence across genomes;

[0090] (c) analyzing genotype-phenotype association on each of the mouse GI MAGs, human GI MAGs and human isolates against their respective matched environmental genomes, by ranking proteins based on their mutual information Z-score (MI-Z) and correlating habitat phenotype with their phylogenetic profile; and

[0091] (d) identifying colonization factors (CF) by (1) combining top proteins from the genotypephenotype association analyses of the mouse GI MAGs, human GI MAGs and human isolates, and (2) consolidating the top proteins into homologous groups based on protein sequence similarity.

[0092] A metagenome-assembled genome is a collection of the genomic DNAs of a mixture of organisms, such as a mixture of microbes (e.g., a mammalian GI tract microbiota).

[0093] Colonization factors

[0094] The present disclosure provides for a recombinant bacterial expression vector which comprises a gene encoding a gastrointestinal (GI) tract colonization factor.

[0095] The colonization factors that can enhance bacterial fitness in the mammalian GI tract may include IMPACT family protein YigZ; TrhP (tRNA wobble base hydroxylation protein, or tRNA hydroxylation protein P); TcdA (tRNA threonylcarbamoyladenosine dehydratase); YbaK (Cys- tRNAPro / Cys-tRNACys deacylase); GTP-binding protein YihA / YsxC; anaerobic ribonucleoside-triphosphate reductase (nrdD-nrdG)', epoxyqueuosine reductase; quorum sensing molecule AI-2; galactokinase; glucosamine-6-phosphate deaminase; FprA family A-type flavoprotein; Cof-type HAD-IIB family hydrolase; tRNA modifying enzymes, or proteins involved in tRNA processing and translation. The colonization factors may include Pfs and LuxS that catalyze the two- step reactions for the biosynthesis of quorum sensing molecule autoinducer-2 and L-homocysteine from S-adenosylhomocysteine (SAH).

[0096] The fitness gene (encoding the colonization factor) engineered into the recombinant bacteria may be wild-type or be mutated.

[0097] The colonization factor may comprise IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

[0098] The colonization factor may comprise an amino acid sequence at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, at least or about 98%, at least or about 99%, or about 100%, identical to the amino acid sequence of YigZ (GenBank accession number: WP_001295262), YigZ (GenBank accession number: XKS61981), YigZ (UniProt accession number: A0A833CEL3), YigZ (UniProt accession number: A0A430FQ87), YigZ (UniProt accession number: A0A174QGE9), TrhP (GenBank accession number: NP_416585), or combinations thereof.

[0099] Amino acid sequence of YigZ (E. coli Nissle 1917) (GenBank accession number: WP_001295262): MESWLIPAAPVTVVEEIKKSRFITLLAHTDGVEAAKAFVESVRAEHPDARHHCVAWVAGAP DDSQQLGFSDDGEPAGTAGKPMLAQLMGSGVGEITAVVVRYYGGILLGTGGLVKAYGGGV NQALRQLTTQRKTPLTEYTLQCEYSQLTGIEALLGQCDGKIINSDYQAFVLLRVALPAAKVA EFSAKLADFSRGSLQLLAIEE (SEQ ID NO:1)

[0100] Amino acid sequence of YigZ (E. coli MG1655) (GenBank accession number: XKS61981): MESWLIPAAPVTVVEEIKKSRFITMLAHTDGVEAAKAFVESVRAEHPDARHHCVAWVAGA PDDSQQLGFSDDGEPAGTAGKPMLAQLMGSGVGEITAVVVRYYGGILLGTGGLVKAYGGG VNQALRQLTTQRKTPLTEYTLQCEYHQLTGIEALLGQCDGKIINSDYQAFVLLRVALPAAKV AEFSAKEADFSRGSEQEEAIEE (SEQ ID NO:2)

[0101] Amino acid sequence of YigZ (Lactobacillus gasseri) (UniProt accession number: A0A833CEL3): MSSKQLNYLTISKAGQHELIIKKSKFICSLARTKTVEEAQEFIEQISKKYHDATHNTYAY TLGLNDNQVKASDNGEPSGTAGIPELKALQLMKLKNVTAVVTRYFGGIKLGAGGLIRAYS NSVTEAAQNIGVVKCVMQQRIQFSIPYNRIDEINHYLEENRISIANQEYTTNVTIQIYLD LDQIQKVEDDLINLLSGKVEFNKLDQRFNEIPVTDFNFHEQ (SEQ ID NOG)

[0102] Amino acid sequence of YigZ (Bifidobacterium dolichotidis) (UniProt accession number: A0A430FQ87): MRTLLNPPEEPAHDSFVEKKSEFIGDACHVESFEDALAFVQSIRDQHPKARHVAWAVVCT DENGNASERMSDDGEPSGTAGKPILEVLRMNELTNVAVTVTRYFGGILLGSGGLTRAYST GASIAVKAAQQAQIVPCSAYHTTIEYTQLGQMQRLLQQMDGEQRDAEFTDRVSLTAVVPS DRAQVFEDQVRESFNATVSLEPAGTVMRNVTA (SEQ ID NO:4)

[0103] Amino acid sequence of YigZ (Bacteroides thetaiotaomicron) (UniProt accession number: A0A174QGE9):

[0104] MTAEDTYKTIVEPSEGIYTEKRSKFIAIALPVRTLDEIKAHLETYQKKYYDARHVCYAYM LGAARKDFRANDNGEPSGTAGKPILGQINSNELTDILIIVVRYFGGIKLGTSGLIVAYKA AAAEAISAATIIEKTVDEEVTVMFEYPFMNDIMRIVKEEEPEILSQSYDMDCSMTLRIRR SMMPKLRARLEKVETARILDEE (SEQ ID NO:5)

[0105] Amino acid sequence of TrhP (E. coli MG1655) (GenBank accession number: NP_416585): MFKPELLSPAGTLKNMRYAFAYGADAVYAGQPRYSLRVRNNEFNHENLQLGINEAHALG KKFYVVVNIAPHNAKLKTFIRDLKPVVEMGPDALIMSDPGLIMLVREHFPEMPIHLSVQAN AVNWATVKFWQQMGLTRVILSRELSLEEIEEIRNQVPDMEIEIFVHGALCMAYSGRCLLSG YINKRDPNQGTCTNACRWEYNVQEGKEDDVGNIVHKYEPIPVQNVEPTLGIGAPTDKVFM IEEAQRPGEYMTAFEDEHGTYIMNSKDLRAIAHVERLTKMGVHSLKIEGRTKSFYYCART AQVYRKAIDDAAAGKPFDTSLLETLEGLAHRGYTEGFLRRHTHDDYQNYEYGYSVSDRQ QFVGEFTGERKGDLAAVAVKNKFSVGDSLELMTPQGNINFTLEHMENAKGEAMPIAPGD GYTVWLPVPQDLELNYALLMRNFSGETTRNPHGK (SEQ ID NO:6)

[0106] YigZ (E. coli MG1655) may be encoded by the yigZ gene (GenBank Gene ID: 948334). TrhP (E. coli MG1655) may be encoded by the trhP gene (GenBank Gene ID: ID: 946609).

[0107] The colonization factor may comprise (or consist essentially of, or consist of) an amino acid sequence at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%. at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%. at least or about 96%, at least or about 97%, at least or about 98%. at least or about 99%. or about 100% identical to the amino acid sequence set forth in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NOG, SEQ ID NO:4, SEQ ID NOG, SEQ ID NOG, or combinations thereof. The colonization factor may have a UniProt accession number selected from below or combinations thereof, or may comprise an amino acid sequence at least or about 70%, at least or about 75%, at least or about 80%, at least or about 81%, at least or about 82%, at least or about 83%, at least or about 84%, at least or about 85%, at least or about 86%, at least or about 87%, at least or about 88%, at least or about 89%, at least or about 90%, at least or about 91%, at least or about 92%, at least or about 93%, at least or about 94%, at least or about 95%, at least or about 96%, at least or about 97%, at least or about 98%, at least or about 99%, or about 100%, identical to the amino acid sequence of the below protein / polypeptide having any of the below UniProt accession numbers or combinations thereof:

[0108] R7J371, A0A7J0AC44, A0A0P0ENB9, A0A564UCE3, G5HK78, H1B8B3, A0A174K1F7, A0A7X4ZNW0, A0A354I9L9, R7EVK6, A0A1C5YI91, A0A0S2W493, A0A174ICL4, A0A173U8U2, A0A7U9NAG2, A0A174JJZ2, A0A0P0M2V2, A0A174FRF4, A0A143XCX0, UPI00189C5A1D, A0A108TDB1, A0A7X4ZRL5, A0A7J0A665, A0A4S1ZDI9, A0A174HR35, A0A7X7V4L9, A0A412FJY3, A0A0F0CFC1, A0A356F6Q9, A0A173ZD17, A0A140DTE8, A0A031WA66, A0A174J055, A0A395YJF8, R5VEM5, A0A641W383, A0A1C7PAS9, A0A318ETP5, A0A1E3AZD4, A0A2X3ACV5, A0A174GSU5, F4BSJ3, C0FP32, UPI00195D25C8. A0A4Q1ZYB3, A0A174ER39, A0A1C7ID27, A0A380LKG9, A0A4P6M049, A0A173RYD9, A0A174X360, C9LE70, A0A0F5IIR0, A0A174B8Z5, A0A0U5JSG3, A0A1C6FHC6, F2JJX4, A0A1V5GNT7. A0A1Y4TEC4. A0A031W9I7, A0A136WIS3, A0A143XCB5. A0A174ALP9, A0A396MQ18, A0A174E0S0, R9J1E0, UPI001D065EA4, R9LUK8, UPI000369FF66, A0A2N3PJI3, A0A3A8ZQN3, A0A353WYG2, A0A3N2N6Y7, A0A174PZR6, A0A174RP02, A0A1I3NDR1, A0A4P6M4T5, A0A4R8MG33, A0A7U9N8N5, R5W2U5, R6IPS9, R9JB32, UPI00137ADAE7, A0A095XIG9, A0A174RXG0, A0A1C6L4K8, A0A1E3ACD8, A0A1T4QIV2, A0A316LKZ9, A0A355X5L4, A0A6N2YTV4, A0A7V6G645, A0A173REZ9. A0A1C6K7F1, A0A369LIQ3, B0MEQ5, J9URN2, V2XXU7, A0A0U5NM54, A0A174AJC8, A0A174YYP8, A0A 1C6BBH9, A0A1C6JGW0, A0A1S9CDX4, A0A2X1ZTG1, A0A3C0WVA4, A0A6N3H1U0, A0A848BZT7, F9Z4S4, H5XWS1. R6N6G3, R7ARV0, UPI00143ABEA5, A0A143XBW3, A0A174K574.

[0109] A0A1C6B0G7, A0A1E8F017, A0A1Y4DBM0, A0A369BAM4, R5BBL4, R9I4B1, A0A143WXK7, A0A1C5ZZJ7, A0A1G3L0H5, A0A3C0TED7, A0A6N7RQE3, C3WGD0, Q838B2, UPI000A485BAE, A0A143WVZ1, A0A173R5Y2, A0A174ALB9, A0A1Y4EZ58. A0A353F0E9, A0A7C6FVH7, E5Y9Q6, G2IEC8, H0UJB2, V2YFY9, A0A011WUG5, A0A081C665, A0A1Q6RFJ1, A0A1V6H433, A0A1V6JQH1. A0A267MBY6, A0A2U1B9Y8, A0A353SAJ4, A0A355JBS9, A0A369XWV6, A0A377GXQ0, A0A4R2TBT9, A0A7V6KFA2, A0A7W4EVU5, G6C5K2, R6UN61, V2YEB6, A0A031WFB7, A0A0G2Z7I5, A0A133N3Y8. A0A161YGZ7, A0A174MRW0, A0A1E3KMS5, A0A1H6S4C3, A0A1I3QRW8, A0A1K.1NJB5, A0A1Q6N425, A0A1V5HLH7, A0A1V5L0K8, A0A239TGI8, A0A316MPY6, A0A318KHB4, A0A348ANS2, A0A373N3A5, A0A379CFK5, A0A380YA74. A0A3D4HY82, C0EC62, E8LH97, F9MW62, N2AI63, R7K5G1, U5MVB5, A0A0C7QI87, A0A0D0RBG6, A0A0F0C7R7, A0A0F0CDB4, A0A0F0CHD4, A0A0M2NLM0, A0A0P8Y7I6, A0A0Q1A8N1, A0A0R1TU86, A0A0U5NNI2, A0A136WCF7, A0A140DRI5, A0A154BVV3, A0A174IQK3, A0A174LA83, A0A194AJT6. A0A1A6AR74, A0A1C5KYE6, A0A1C5Y773, A0A1C6D8G5, A0A1C7ICS6, A0A1E7S2J5, A0A1I5TRM5, A0A1M6P7S0, A0A1R1EMG2, A0A1V4SKZ2, A0A1V5I8V8, A0A1Y4LDV9, A0A267MP81, A0A2K9J6M3, A0A316N2V6, A0A351VBM8, A0A371IKK4, A0A377JTF2, A0A381J873, A0A3B8XUC4, A0A3D3FJA2, A0A3D5UMA5, A0A412I6K0, A0A423UNG6, A0A496NQ74, A0A564S1K7. A0A6N3EY58, A0A7C6CJ66, A0A7M1Q0X0, A0A7T7XJZ0, A0A7U9MVL2, A0A7U9X5W9, A0A7W8D0L8, A0A7X4Z6T7, A0A7X5SWF8, A0A7X8UEK2, A0A829Z7N1, A0A876WFW7, A0A8G2FM70, A4BA73, B0N261, B6W8U6, B7C8B8, D1AMC7, D1PPL8. D2Z7B4, F4LUF4. F7V2V8, J4TLA7, R5A2M2, R5E4R7, R5XEA0, R9KEV0, S0FNJ5, S2ZME5, T2RE82, T4VMF5, UPI000400ECFF, UPI0008361932, UPI000C07E220, UPI000CFB010C, UPI00166BEBDC, UPI0018ABDAC5, UPI001A9B3845, V7I4J9 A0A174BQG7, A0A7U9SUR8. A0A7U9NJ06, A0A1C5MK14, A0A1U7NM68. A0A1Y1VHI4, A0A1C5YLZ3, A0A564SB80, AOAOFOCKCO, A0A6I2UGF6, A0A174V4I5, UPI0018AA151A, V2XWV5. A0A4P6M1Q0, A0A7X8EVT8, A0A1Y4H7N5, UPI00048A3F03, A0A1C6CQ03, A0A490LK34, A0A174FFS7, A0A1C6KH04, R6DJ08, A0A430JTX7, A0A830UBE1, A0A843CUG4, D4MJJ9, A0A7U9XBG2, G5IFB6, A0A0C7QI87, A0A1M7L608, A0A1U7NG52, A0A354HFY0, A0A7X7H793, D1AI58, Q7WZ38, R7AET9. UPI000C7B89FD, UPI0013EC40F3 A0A174KHG9, A0A316NCC3, D4MU96, A0A349ZRP4, B0NBS9, T2F2X3, A0A0P0G3K1, E6LS26, A0A355V1Z1, A0A1C6H6R9, A0A3D1RG23, A0A4Q0J115, E0S4M6, A0A3N2M1J6, B0MDN5. A0A0P0FBD1, D4CJ61. A0A1V5GFC9. A0A1D3TT77, A7AXJ8, A0A3A9EKJ3.

[0110] A0A1I7IRM9, A0A4S2EZN0, A0A369P297, A0A6M8IYD9, A0A843F9S8, B0N507, A0A3D6B142, A4MVI8, A0A174RZA7, A0A316LF09, A0A268FDP3, A0A7H2BCV3, A0A7X6NVS6, A0A0W7TMN0, A0A110A6R3, A0A1B3SKA0, A0A0P0L0X5. A0A0Z9C1T3, A0A3B9DXK6. A0A4V2DZ87, A0A7X6YLT6, A0A1V4I3W1, A0A1C5R2Y1, A0A255TDS2, A0A3B8SNJ6, A0A0F3H8W1, A0A1V4IY11, A0A7X7ADV9, C4GDQ9, G4KP44, A0A133Q095, A0A1C0BWG0, A0A352WKQ3, A0A143ZW43, A0A1C5RNL1, A0A2D3D7H2, A0A800BRH3, C2CFF7, A0A0U4JTC1, A0A139Q679, A0A174SMH1, A0A2P1G460, A0A6N3EZ01, UPI00148BCC14, A0A1H6ZGW3, A0A2L2WVH1, A0A354MVS0, A0A3C1MED6, A0A7W4HGV2, A7VPC5, R5NBE2, A0A239WAZ2, A0A2W5SL79, A0A7X8LA18, J9FVY0, A0A349YTN3, A0A3B0YRC0, M5RBX1, Q9L645, A0A1H9SJS0, A0A2V2F3W7, A0A848CD97, E0E2Z6, E3EAV3, UPI000BE3B5E4, A0A1M6JEB5, A0A1V5RWE4, A0A369LW75, A0A415CGA8, A0A847DNH4, A8RAW6, C9KJA3, E8R0V9, A0A357ZFH3, A0A4Z0YC11, A0A6I2UWI2, A0A7C1Z031, A0A7C6SHT1, A0A7X0HRR7. A0A0A2FX21, A0A1C6HBF6, A0A1F2PI34, A0A1H4CIE4, A0A1J5EPH8, A0A2N2E481, A0A2P9IUF8, A0A3B9VDZ1, UPI001ADDBCFC, UPI001C10C82B, A0A136CQR0, A0A151AME4, A0A5C7R4F9, A0A7V6H5B8, A0A7X4ZRQ5, D4S5Y0, Q6KG76, UPI0009941DAD, W1Q542, A0A0J6GZL5, A0A143X6V3, A0A1S8M773, A0A380WW17, A0A415ZYS1, A0A432IGX4, A0A6N8BFL8, A0A7V6GPB0, C8VXR7, C9RK15, R5VC77, U7UV53, UPI001ABBCF2F, A0A0A7I4B0, A0A0F9TL10, A0A0H3XJ05, A0A0N7GFM7, A0A1Q9K2R9, A0A1Q9NKF2, A0A2K4ZM26, A0A2U3PMI1, A0A317UCF0, A0A373KPI0, A0A374DUW8, A0A3D3J0A7, A0A6N7S8D2, E8N5G3, F0GUH6, K7QLT1, N2A9S1, R5PMB8, R6TIS0, UPI001C105BEF, W4BEQ8, A0A017RUR2, A0A090ZGC8. A0A0K0HHI2, A0A0Q1AFS0, A0A0R3JS94, A0A1H8GG86, A0A1L8CSG0, A0A2A2W5C0, A0A2K9J1X6, A0A2N5ZGB5, A0A2Z2KNZ7, A0A2Z6IEP5. A0A315ZC49, A0A350H6U5, A0A377FU17, A0A379CI75. A0A3D4JCQ7, A0A3D5NNR9. A0A3Q8S9B2, A0A417E0I9, A0A4Y6E9C1, A0A538Q940. A0A5B9N4K6, A0A6L5XE67, A0A6M1X8X4, A0A6N3AXM4, A0A7C6GBG4, A0A7C6J0G8, A0A7C6TWS2, A0A7C6UUL4, A0A7X9J1B4, A0A7X9P4Z8, A0A7Z0PFW8, A0A829I8C8, A0A843FG79, A0A850GIK3, A0A8B0KJH6, A0A8F6HP22, B0TGD7, C6JR79, E6TY96, 11TK00, K0AZU1, Q7Y465, R5SSS2, UPI00042226B1, UPI0006890BB3, UPIOOOD 144308, UPI00187B65BE A0A348AAH6, A0A1C3SH66, A0A2S6DJR3, A0A291BEM8, A0A2V2EF74, A0A2C2UGV9, A0A4U9HVT3, A0A 173SZT6, A0A2N6T3F0, A0A358NE44, A0A5A9DK39, A0A5R9BYZ6, A0A0L0W7I7, A0A0U1QSB7, A0A1H8ZEG9, A0A1S6A782, A0A443TFA0, P94499, A0A090SAT6, A0A2S7N0G5, R6BT40, A0A380N1Y4, A0A6I2GKH1, C2KJR0, UPI0003637298, J8Q4I1, A0A1M4N937, A0A1T0A3G4, A0A4U3BQT5, N2BPH6, UPI00097C473B, UPI0011A65857, A0A3C0WZU3, A0A4U2MFS0, UPI000364C2D3, A0A099S916, A0A0A7FXG8, A0A0C1TZ53, A0A0J8D5I9, A0A0X1U7A5, A0A1E8VXT1, A0A1G6N8L8, A0A239ZK46, A0A2H9T588, A0A2L0PKC6, A0A378PJ67, A0A430A7D3, A0A891X7W2, P19072, R7HFL8, UPI0018A289DB, V2Y0C6, A0A0H3BZ54, A0A1E8F0H7, A0A1E9HMF8, A0A1G8P0K0, A0A1H9DJ91, A0A2I1M9H4, A0A2I1MTJ0, A0A2S0JH38, A0A2Z3TU73, A0A374JAC6, A0A3A9K1Y9, A0A3N0A8B9, A0A3R6QZA6, A0A417LVY0, A0A4R3KZS6, A0A6G7ZB43, A0A6H1X1I1, A0A7C6BLR9, A0A7G9WB86, A0A7Z0LG50, A7I2J6, C0AYW0, C9LCC7, D9PQ36, E4KPX3, F9DWM8, J5HU31, P25185, P54104, Q49XK9, Q8REN9, S3L2G2, UPI0003724782, UPI0006D7C77D, UPI0008265F0E. UPI000BF44005. UPI001AE2A758, UPI001CEF57D6, UPI001D0AA327,

[0111] R9N9C7, A0A1N7HW37, A0A3C1XHB5, A0A077KLP6, A0A1E7X9T3, P07464, A0A174SH54, A0A1K1MVR0, A0A1G9KS56, A0A143XT69, A0A1B8TDH5, A0A3C1ED19, A0A3N2M6W7, A0A0K9GDR5, A0A3Q8STH2, A0A1C5M2A8, A0A412JEF9, A0A3D0FNK5, R7PNE8, A0A174MVC5, A0A0F5EWL6. K0XQB7, R7FAK4, A0A173Z9E8, A0A329TM50, A0A3D0SPE2, W6S6K7, A0A0L6Z791, A0A174X2U9, A0A3D0JTH7, A0A7U9SUE0, F3UZP9, Q5HCZ5, R7K8V0,

[0112] A0A3C1H2Y2, Q2RKK1, A0A7H2BCQ4, A0A160T6Y3, A6LC43, A0A836P8D3, D3IDS6, A0A3M1HW86, A0A1C2BQ37, A0A350IZD5, A1A190, I7KL61, A0A425VY26, A0A2E2UQN8, A0A4S3B619, A0A316SQU1, A0A1F5ZQC3, A0A2S4DWN9, A0A662EF28, A0A0J6WZH1, A0A1H4BS97, Q8AAN7, Q9ZEE3, A0A137SVQE A0A381IC96, A0A7S3K6S0, A0A0C2UYJ1, A0A0X8H0P7, A0A1C6BB31, A0A1F2WEH2, A0A1F8NHT0, A0A4D9CRK1, A0A6N8UDL2, A0A7S1TFB5, A0A7S2I9T6, A0A7S3ACC3, A0A8B0IHZ1, E0NU32, Q7X402, R5VHV8. R6HNA7, UPI001CF84A77.

[0113] U1J8Z6, P58824, Q897F7, A0A3D1V7W2, A0A117MAN5, A0A094IM03, A0A4R0XPR1, A0A7S1KQI3, A0A7S3XLE8, A0A292IH60, A0A7S3P4P8, B8I6R1, Q03PZ8, A0A2T2UHZ0, A0A2V2EPF6, A0A074ZZK9, A0A7Z9XCW1, A0A090I2W1, A0A1S2F2J6, A0A6C0CQJ4, B1AJ02, B2GE56, P54264, W6A8Y9,

[0114] P76403, A0A2V2GKT7, D9S976. A0A352BH46, A0A377J6A5. A0A078KQN0, B0PFG8, A0A3A9AXE9, A0A380LJY8, A0A1D7YRR4, A0A1R4EEB9, A0A0A2WLK1, A0A1F3IN85, A0A4U9RF74, 032034, A0A1Q6T658, B9Y9W6, E0RY86, P44700, A9BH48, A0A3D1YX96, D4C991, H1D2E2, A0A081CAV5, A0A1V5S4X0, A0A2E9QFM3, A0A350J789, A0A7G9YHI9, A0A7X5HX49, F2JR01, H3K5Z7, R6XN27, R9A9R0, R9K2L3, UPI0004BCF0CF

[0115] Q8ZSU3, A0A0P0N5H1, A0A3L7YM55, Q9JYC3, A0A4U9XNK0, G3C9G6, AOLMHO, A0A7C4QQSE A0A7C7RN80. Q97ST4. A0A7C4L0Y8. D5SNX5, A0A117LIM7, A0A347ZV19. A0A7Y2CTM3, A0A842W0T8, D8FA22,

[0116] A0A174MZH2, A0A1Q9YIV3, A0A416MG87, C4IDA4, R9KRQ9, C0C2F7, A0A7U9RQE7, A0A174A5D9, A0A1U7NKP0, A0A349HY12, A0A829ZD01, A0A380ZBN9, A0A7X4ZNG9, A0A7X8C8H4, R9MT60, UPI0015E07926, UPI001D17865B, A0A174HTT2, A0A1C6BVF8, A0A292QPU0, A0A2N2EJY6, A0A4Q4IIP0, A0A6L4X207, A0A8B2DYT6, R6U0M6, A0A173T793, A0A1Q6UGF2, A0A327J068, A0A327JAR0, A0A355TXD3, A0A3E2UTD2, A0A6N2ZQR3, A0A7X2MY39, R9I231. R9J018, R9J633, S5NS31. UPI000C776BB6, A0A099WFL2, A0A0F5IQ87, A0A0F5IVR9, A0A0H5SUX1, A0A117LNR0, A0A174SM12, A0A1B2ISP4, A0A1C5N5C4, A0A1C6C6M5, A0A1I0CVQ9, A0A1S8TDG4, A0A1V8X9D6, A0A1W1YWB3, A0A1Y4UFP4, A0A2U1S9L5, A0A350XUY4, A0A373IIJ1. A0A374LJB7, A0A377KL52, A0A3C0NMG5, A0A3E4Q6C2, A0A4Q2BM76, A0A4R3ZBR3, A0A4U7JFX2, A0A4U8TPK1, A0A4V3RS61, A0A4Y9FIM0, A0A564W5M6, A0A6F9YL06, A0A6G2CDX6, A0A6I3NKN2, A0A6L9H4K4, A0A6L9HVW7, A0A6S6SW32, B9YAA3, F7K5P2, N2BMI5, Q5LIR2, R6CW81, R6D8Y6, R9LRH3, UPI0011A28367, UPI0013EAB5AD, UPI00157131BE, UPI001914BC77, UPI001C2C2572, W1SES6, A0A140DV18, Q8A5W2, A0A1Y4J174, Q8Y4R7, A0A1H4C1J5, UPI0004AFA52A, P13375, D2MNP3, R5M0U5, R7EZB9, A0A7C6MQA6, A0A7X5EA30, P13376, 18UFL9, Q0TN51, A0A3D4TVJ1, A0A3D8HUE5, A0A7U9MS03, Q6KH90, Q9X1A5, A0A174AMK8, A0A1C5WED5, D2MN49, C7HTI9, A0A7T8ENS5, A0A376CKI0, A0A7V8D756, A0A7Z2PFM0, A0A5B9N4A3, A0A5R1NB56, P28903, B0N506, A0A7X9H2S4, A0A660MRJ5, A0A090ZJN6, A0A349YSS0. UPI00126036EA. I1TE01. R6BF34, A0A349YTN4, A7VS55, M4QTE4, A0A315Z9R3, UPI0016525174, A0A351YVL7, A0A347WM99, B9YEC6, P07071, A0A1X7MUV9, A0A6G6XUQ6, C6JR80, A0A4Y7RWQ5, Q8RGI2, A0A7Y3SVD6, 11TJZ7, A0A0R2LKS1, A0A2C9CYL3, A0A556UAG4, A0A5Q2N6N9, R5YG62, A0A357IRR3, A0A6N8CS93, A0A1H0T954, A0A1Q6EL13, A0A2I7REQ5, A0A373KMX7, A0A4Q1KXP3, B1C8M3, D3ELL1.

[0117] R5D8G5, A0A417BT84, A0A352HR48, B 1C8M3, B0N506, A0A5C1QD25, A0A4S2AT59, B9YEC6, A0A1Y4CZP4, A0A1H0T954, R6Q4S7, A0A4P7AHN5, A0A6N9P1P3, D3ELL1, A0A374LFH6, A0A4U3A5B2, A7VS55. A0A7U9S0M1. A0A0J1IJK3. A0A316N9U3, A0A143WYE6, A0A1Q6LDQ2, A0A2K4ZM04, A0A353T5Q2, A0A378X6B2, A0A5Q2N6N9, A0A7X6XAN2. A0A7X8LAE9, 10GNZ7, G4KS62, C0W2X6, A0A0A1U415, A0A0S8DS84, Q6NEC9, Q99Z89, A0A7X8ZTC9, A0A174IVT2, A0A3C0WVW3, F2I7I8. K9EFC1, Q898P3, A0A2H0WR67, A0A2M7XJI5, A0A3M1HDJ5, A0A7J4HY98, A0A7X6YMI4,

[0118] A0A3A8ZEZ9, A0A143XY64, A0A140PR14, UPI0009FA75B3, A0A174AZC8, A0A6N2ZDX6, A0A7X8CHH4, R7CFG2, R9L0N7, A0A173XCM5, A0A3D3WBJ7, A0A1C6DVL7, A0A3D0LFF9, A0A3S7JAE3, A0A1D3TP12, A0A2W4I0N0, A0A410P3I2, E2NTX3, N2AFX1, A0A174HJK8, A0A2G6G4C0, A0A379C5U8, A0A525D4R9, A0A7C3TS31, V7I876,

[0119] A0A6G7Z924, A0A378HHT0, A0A1V6LNB9, A0A2X3Y1X0, 17L9Z8, A0A0X3ASS1, UPI0003FDFE14, A0A0R1K5L5, A0A553GP26, A0A239SQX0, A0A3B9ICU4, A0A239SRE4, A0A174PSX4, 17LDN2, A0A0D5YL19, A0A162QBU3, A0A389M574, F8LJ64, A0A023CV3E A0A089KBM9, A0A0E4H637, A0A0S3KFC2, A0A0X8FME6, A0A116M2G2, A0A139NT79, A0A139QJY0, A0A143WXX7, A0A173M779, A0A1C5UIG6, A0A1D3TT58, A0A239U0B1, A0A291R0S1, A0A2R7ZZ11, A0A2X0SS26, A0A318L4L3, A0A3D4VRY0, A0A3L9DVA7, A0A3Q8SQL0, A0A5S4TU39, A0A6M0H0N5, A0A7X6S4D0, A0A7X6VGX2, A0A8B0GIH1, C2EGL1, E7FNS9, 17LSG7, S6FB19, UPI0005D20D96, UPI000838C6EB, UPI00083D9DF4, UPI000C82BDAD,

[0120] A0A061NAR0, C3WCH3, F3Y8H4, G5H650,

[0121] UPI001BE9A614, A0A0L0WF80,

[0122] A0A1W9SBU9, A0A2H0W361, A0A0M1J156, A0A7V0XFP8, R5VP73, A0A3D5A121, A0A485KXE4, A0A523CEN9, A0A7C4KR76, A0A7U6GE24, A0A191HX73, A0A1V5K442, C0QRT0, A0A1Y4NHB4, A0A388TFQ6, A0A3C0X233, A0A7C5EPK5. C9L892, D4CKS9. V1CS35,

[0123] A0A532UVW8, R5VP73, A0A7U6GE24, A0A521IHG5, A0A7V6J2B0, R6XUQ6, A0A081KAJ6, A0A2D6N9S3, A0A7C3FVJ3, A0A097ST36, A0A0K9NB09, A0A174A5U4, A0A1C5KRL7, A0A363T710, A0A3A9BU90, A0A4V2Q833, A0A537LT94, A0A6I3SJ06, A0A6N7WZQ4, B9CK55, F3ZX75, J1H3N9,

[0124] A0A1Q6M4T8, A0A5Q3Q963, A0A0D8ID86, L8TKX4, G4L5Y9, A0A1 Y4HKZ8, A0A389M7M0, A0A7C6FS10, A0A0Q9YTW3, A0A2T5G8C3, A0A101I8Q7, A0A1C7GRR2, A0A507EQA8, A0A644W5Q4, A0A6F8SPP3, D3E3W2. K0B170, R5KPL5. R6H9K8, A0A2K8M3R6, A0A653UAG4, A0A0B9A5K4, R5HAT7,

[0125] B0NGS1, A0A127EG47, A0A227LGX0, A0A078KV03, Q58156, A0A2X2YCM9, A0A1V4STL7, A0A1S8P1A2, A0A259UCD7, A0A3N0IA80. Q97D83, A0A419G0S7. A0A645BTX3, UPI0002473F4D, A1HNP3, A0A3B8W5D1, A0A1Y4TLR2, UPI001A1E898C, A0A379G972, A0A4V2WS86, A0A1C5LFB9, A0A522YUG2, A0A7C6TN16, R6M2Z4, A0A2A7MIT7, R9IWK6, B1C894, A0A3B8R1J9, A0A0S6VUH2, A0A355V563, A0A174PP73, A0A1C6HQ07, G1WR51, A0A6N2RD53, R5AVW0,

[0126] A0A1Z5KPM5, Q5HKH9, A0A2P0VNR5, A0A1Y1WTI3, A0A6J4X6K0, A0A0B7I7V9, A0A2P6TKX7, A0A847JPY4, 032797,

[0127] A0A0P0EXS4, A0A108TCB3, R5XUE8. S0IZV8, A0A088F4C0, A0A1V6ACZ0, A0A7X5CAP1, R6JRQ6, UPI000411E3F1, U5F9C3,

[0128] A0A174IC81. P13243, A0A1V2BS88, A0A5Q8BTQ4, B0NJ34, A0A174A4Y2, P0A4SE A0A0S7BSQ4, A0A1C5VVD7, A0A1S6IR80, A0A352PLZ2, A0A380BUT5, E0RVV1, P77704,

[0129] A0A1C5YG70, D0WED4, P37877, A0A5K1IMC7, Q9WYB1, A0A7C6MK02, Q726S6, A0A086ZJV6, R7IE24, A0A348N4I8, B8J374,

[0130] A0A1F8V5W6, A0A1Q6Q6T9, F9MV73, A0A7U6QJ67, A0A173M5A2, A0A174DTW8, A0A2S1TZ13, A0A354HHK3, E0S1T0, UPI00058E9F2F, A0A139DRU9, A0A173UI18, A0A1S8TC65, A0A1V5XG99, A0A3D1WSX1, A0A3D4L3G1, A0A6N2WXD2, A0A7C6HMD4, B2A6I5, D7UWX0, F9MPP7, H1HT16, R6GNT9,

[0131] A0A7U9RPS2, F0KGN8, A0A6D2CJK9, A0A1H8AU05, A0A7J0BC12, R9IXX0, R9ME70, C0ZE92,

[0132] A0A1E3A1K3, A0A7J0AJ87. A0A4S2ALZ3, U2EUZ7, A0A7U9S067. A0A4Q0U7J3, A0A6L9H405, WOESXO, A0A7J0ACZ6, A0A7U6KJQ5, A0A1C7GNV3, A0A0F0CCI8, A0A4Q0IGV5, A0A0F0CE14, B0NDL4, A0A3N2MBA8, A0A4S2FAA4, A0A174KQM7, A0A1C7FZ58, 15AWA3, A0A175ADG3, A0A829Z9L4, A0A0X8V908, A0A3A8ZYH3, AOAOFOCALO, A0A352HNC1, A0A373Q250, R7GAL1, A0A4Z0XXW9, A0A0D8IXL5, A0A3A9FSG0, A0A6N9PAJ8, A0A4S2HN22, A0A1C5ZVY2, A0A0D8IXJ1, A0A4S2F7N0, D5HGB0, V2Q216, R7KWQ4, A0A140DVY8, A0A1T5SVJ4, A0A316TB64, A0A4P6M8Q0, A0A6L5XCA3, A0A7U9RCQ7, A0A7W5UL62, D4J8P8, R5L1W9, R5L822, R5Y1I1, X8K446, A0A143X6Q7, A0A1C6A7G6, A0A1I3S9M6. A0A3D4S7Z7, A0A1E3A6R7, A0A1E7S263. A0A3A9IYC6, A0A3C1EBP5, A0A417HB94, A0A0G3WKJ4, A0A1G4WXN3, A0A416CWT7, R6QSE6, A0A1E2ZZN6, A0A413WSN0, A0A658JQL1, B7AN89, A0A316P9E3, A0A3B8XUN4, A0A415Y2C3, A0A4U7NG26, A0A6N9PAA4, A0A7U9NE60, R5CLA2, R7BT26. A0A1G5E3P6, A0A2V2FZX1, A0A4Q7PNW2, A0A4S2F728, A0A644XY53, A0A6N9PGJ2, A0A6N9QGT3, A0A7U9MP58, D8IAW3, H6LIE7, R5NGV4, A0A1C6GBB5, A0A1E3AXN3, A0A1K1MN39, A0A1M6NHH1, A0A316SSU6, A0A3A6JVI6, A0A3D4XC58, A0A4P6LSE7, A0A7X8FLL8, A7VSE5, D4LK59, A0A0F0C8X1, A0A0F0CD04, A0A0F0CHV8, A0A161LGE2, A0A1C5M7A7, A0A1E3ACE7, A0A1G9G8R9, A0A1I0FB86, A0A351ASF0, A0A354IFM5, A0A3C1UWR7, A0A3D4RPF3, A0A3E2VY93, A0A417PK68, A0A4R2LDW2, A0A7U9MMR5, A0A832KRV9, D4J6V1. F0Z3K3, J1H099, R5GJR3, UPI00195C8267, D4LD72, A0A7X1J1B1, A0A5C4LV03, A0A261F3R6, UPI00030BD737, A0A1C6RLF3, A0A1I2AKH9, A0A6A7USL4, A0A425YG93, R6SUK2, A0A1H3CZU9, A0A7U9SRV8, A0A174FI06, A0A4U9RJJ4, A0A4V2WS86, A0A0Q3WUZ2, A0A0R2DT51. M5DYF5, A0A2A8H8C2, A0A448FCF1, A0A3A1YHJ8, A0A1E7K433, F9N3R1, A0A1Y1WQM9, A0A1G5WW98, A0A1V5GRE4, A5I1C7, B1WQP0, A0A1Y3UF83, A0A416CYY0, UPI0012AC47BC, A0A1U7M7A6, A0A224AJQ7, C4FSL7, A0A6N3B692, A0A7X7BTK8, C2BI81, A0A088T200, A0A1V4M5V6, A0A352A2Y4, A0A3C1U6V4, A0A524JPI6, A0A7X4I9A9, A0A3B8RGF0, A0A0R2C8N8, A0A7X2ST07, A0A6N2R088, A0A0R2CZZ3, Q1LTN6, A0A1M4YHN4, A0A4S2EV47, A0A847Z758. X8ITW3, R5PWZ6, R5WHK4. A0A1C5SZL7, UPI0015F5FAD0. A0A1C6D2K4. A0A2M9H975, A0A316PGU1, A0A8A8TM14, A0A223B341, A0A3D1VD87, R5K9M6, A0A2V2G4F8, A0A173RVW9, A0A6N4TLH6, A0A1H3ZVR7, G4Q3A2, A0A1V5GFJ3, R6GMB7, A0A174LLY2, A0A174T347, B9Y2P3, A0A0Z8CVX1, A0A174DKN9, A0A1C6IX29, A0A1Q6KTF4, A0A1Q9YE56, A0A7C4ECP2, A6BJ03, A9KQ64, A0A173S9Y5, A0A1G5D6S6, F9MUZ5, R6IYH5. R9LND1, A0A143WY22, A0A174KF34, A0A7X6UQU2, C7H221, A0A1Y3TUF5, A0A1Y4LH93, A0A4Q2BUE4, A0A7X2PDD0, G4Q6Z3, A0A174S7D9, A0A1C5KNL0, A0A1C5UN08, A0A1Y4AK24, A0A7C6J6A6, A0A848B5M8, A0A1C5WF37, A0A1H3J7G6, A0A2T0B6R5, A0A3A8ZM13, A0A653AU13. A0A7Y0HS47. A7B571, R7G603, R7HQ80, A0A174DYD6, A0A174ZP48, A0A1Y4TX12, A0A3R6JK96, C3X2S8, R7F8C8, A0A0M6WI15, A0A0X8JE99, A0A174ATN3, A0A1I0CIX2, A0A1M6NH40, A0A1T4MLC3, A0A2P8ENV4, A0A396QIE1, A0A3C0NYZ8, A0A3N0IBB2. A0A644WXK0, U7UL44.

[0133] A0A1H5VGF2, A0A1Q6IZD4, A0A3D5TED5, A0A3R6I177, A8SS54, C5EIA7, D9R881, G5GHP6, UPI0009F42FF7, UPI001CF928E1, A0A143YYD1, A0A1B0ZK90, A0A1C6KDM6, A0A2Y9BM17, A0A3D2CKF9, A0A3D2W8Q3, A0A7U9RRR2, A0A7X8FEJ6, A7VQ73, E0RVN6, F0RRV1, G4D198, A0A0F9ZQI4, A0A0R1RL38, A0A0R2IBS3, A0A174F4G5, A0A1V4J0K0, A0A2S0KME9, A0A351EH46, A0A352WS53, A0A3B8U228, A0A3C1LXT8, A0A3S4T0W6, A0A847QC24, A5TW09, R5R4M8, R9IKU5, U2AMV4, UPI0004037F19, UPI001D09F3EF, A0A143X8H2, A0A173ZXZ9, A0A174R8V6. A0A1C6BID7, A0A1H9EDB7, A0A1Q9YGK6, A0A1S8S5Q4, A0A1V4I597, A0A2V2EPC6, A0A3A9S8J1, A0A3C1HVK5, A0A4V0Z7Q0, A0A6F8SNM7, A0A7C6I7T9, A0A7X8ZNH8, A0A8A5CFU0, D2RJD4, E2ZC28, E7MQ42, F3B1X5. F5RJR0, H1D108, H3NHH0, R6N1I7. UPI000288B4EB, UPI00196A1B54, UPI001A9AE198, UPI001C120363, A0A0F0CHE8, A0A174G415, A0A1C5KGB3, A0A1C5MA39, A0A1C5U3V8, A0A1C6JAP2, A0A1G9ME12, A0A1I0BLR5, A0A1T4VV61, A0A2S0L1F3, A0A2U1AYM9, A0A351XKU1, A0A356N8K6, A0A357SXJ4, A0A358RJX3, A0A3C1XFI0, A0A3D1WT45, A0A3D5KZ62, A0A417NFG8, A0A6N3CBT7, A0A7C5MM84, A0A7T7XNL7, A0A7X6X6U2, B0MHB2, E6MIE1, E6UE67, F1T4J8, K0B2K1, R6LGY2. R7G2R8, R9M0N7, UPI000A865B1D, UPI001031648F, V2Y237, A0A0R1SG05, A0A133XQL6, A0A173VPB0, A0A173WJU6, A0A173YCB5, A0A174A1W8, A0A174AG14, A0A174AGV6, A0A174HZ90, A0A174K8Y5, A0A174NNZ7, A0A174PM42, A0A174Q5P8, A0A174WVF2, A0A174ZNX7, A0A1C5XD32, A0A1C6JY27, A0A1H6GNC4, A0A1H7LLI0, A0A1I2BMU6, A0A1I2DLR8, A0A1M7ULS8, A0A1T5D5W6, A0A1V2YC10, A0A1Y4MCY7, A0A2A7MC33, A0A2G3E781, A0A2J8B673. A0A2V2DXF9, A0A2X4NCY2, A0A353DLS8, A0A354ARF5. A0A356XDQ6. A0A374RXR6, A0A378HJ73, A0A397S760, A0A3B8U1I2, A0A3B9VQT6, A0A3C1UV95, A0A3D2TS10, A0A3F3S4S2, A0A3G9J2T0, A0A416E9D2, A0A4R3YJP4, A0A511B0F2, A0A661Z1H0, A0A6N7SDZ2, A0A7C6ET68, A0A7C6TPY6, A0A7J0AAX8, A0A7U9MQD8, A0A7X8IGK7, A0A7X8W0Y0, A0A7X8W573, A0A7X8ZMS3, C5EG99, C7N7E1, F4GIY2, F8HYG3, F8N5K3. F9PP43, H7EJV5, Q73P68, R5VLV0, R6BLP5, R6BST6, R6XVQ6, R7BWL0, R7D2F8, R7EPH5, S6A2A0, UPT000975B333, UPT00141307CC, UPI00195D750F, D4J7H3, D4CE20, A0A3M0ZMS7, E0NME1.

[0134] A0A174EZL1,

[0135] R5LLQ6, C6L9L8. A0A5K1IXN3, K0B2C3, R7AL67,

[0136] A0A371JJ75,

[0137] A0A357T7R5,

[0138] A0A7C7AHX5, A0A108TD57, R9LQN7, A0A3E2XMG7, A0A347ZVR3, A0A0D8IYH6, A0A2V1IP85. A0A3D4N2K1, A0A0B2XK64, J1H898. A0A031WF92, A0A3A0C0W1, A0A7Y0L6N2, A0A1V5L8Y2, A0A356AXE9, A0A356L4S7, UPI0001E2F33A, UPI000CE1ADCA, A0A1E3L7I4, A0A351TJ30, A0A3G9K558, A0A7V1EGP1, X5EL59, A0A7C3LZR8, F7V3E8, UPI000C832B7E, A0A1A6L183, A0A173Y1Q1, A0A351G3C8, A0A4U2LAJ7. A0A662ZJ39, U7UJR7, A0A2E9LCU3, A0A7W0QFY6, C8NG38, J6IJ91, A0A0P6WYP3, A0A367CF17, A0A645BZC9, A0A7R8PBU9, D2MMJ3, F0FUH9, Q6LNC4, UPI000420B92D, UPI000A3CF3E9, A0A088F4Q2, A0A0R2DUM6, A0A1Q4UGU0, A0A534YSX4, A0A6L5Y1Y4, UPI0018EE9827, A0A0M9FKS3, A0A101H0G6, A0A359DVS8, A0A380H9B1, A0A381DKK5, A0A3D0RS09, A0A3G9J482. A0A3Q9BJL3, A9WKN0. D4MVR0, D6TJA7, S2YTA3, W1TTV6, A0A081QQY2, A0A0H2YS73, A0A1I0YLR7, A0A2I0MYM1, A0A3D0FZP1, A0A448TS85, A0A6L7U3P3, A0A7J7AIP2, P36922, UPI001B832DAE, AOAOHOYFQO, A0A143XYZ2, A0A149SKF8, A0A149SSW1, A0A1C3SDN2, A0A1Q3QQM8, A0A1T4KEF9, A0A1W9XFY3, A0A239XB75, A0A351Y950, A0A354XTC4, A0A3M1J9U0, A0A523BTL3, A0A6A5L529, A0A6I7P4M0, A0A6M6E2V0, A0A7C5JUM7, A0A7C9QW30, A0A7X6TWS2, A0A8A7APD4, A4SS95, C4GGT3, H6LAM1. S2VUY0, UPI001C969E3B.

[0139] A0A096LI29,

[0140] B0NJI8, R9K8Y2, G9RRR2, A0A087CFV8, F3QMX8, A9KI52, A0A1M6GDC5, A0A498CKA8, A0A645CA69, D7UVY3, A0A2D5B0I5, A0A4Z0Y857, A0A174M4M0, A0A1G9H7B5, A0A1V5TYA3, A0A645E328, A0A1M6AW46, A0A261F2Z7, A0A2N0UUC6, A0A3B9JIM7, A0A351X9U8, A0A564U6U7, R5NNJ4, V2Y0V3, A0A1C6FFC6, J9FDR1, R5IXW0, A0A0M9D4W4, A0A3D5XSE6, A0A6H9KXH0, A0A173R2Y4, A0A1C6J568, A0A1K1LIY6, A0A1V5GN19, A0A1V6AJU5, A0A7X8VT26, G2HD78, A0A239U2J3, A0A316LU79, A0A387B8U9, A0A4P7L4J6, C9XTB4, D7CYC6, 17LG63, P27862, R6RSH3, UPI0006BB6171. A0A087BTI7, A0A099Y9Q9, A0A173X960, A0A1C6FHP0, A0A1E3L521, A0A1I9YMI2, A0A1S1V856, A0A1S6QJP2, A0A240ASF0, A0A2N1R2V4, A0A2N2EHU7, A0A2P6TF13, A0A2X0WV36, A0A316PJ25, A0A4Z0D3X5, A0A645H2G6. A0A6L4ZPT4, A0A7T7ALX4, A0A7X6HWU9, D6GQG7, R6QFC9, R7BS39, U4TMS4. Bacterial expression vectors

[0141] A bacterial expression vector may refer to any recombinant nucleic acid (DNA or RNA) molecule that is used for expression of a gene (e.g., a fitness gene encoding a colonization factor) in a bacterial cell.

[0142] The bacterial expression vector may be capable of autonomous replication in a host bacterial cell into which the vector is introduced. For example, the bacterial expression vector may contain an origin of replication for enabling replication of the expression vector in the host bacterial cells. Nonlimiting examples of E. coli replication origin include SC 101. colEl, pBR, and R6K.

[0143] The bacterial expression vector may be integrated into the genome of a host cell upon introduction into the host cell, and are replicated along with the host genome.

[0144] The bacterial expression vector may contain at least one cloning site (e.g., a multiple cloning site (MSC)) to enable cloning of a gene (e.g., a fitness gene) under the control of a promoter. The cloning site may be a sequence of several unique restriction endonuclease sites.

[0145] The bacterial expression vector may contain a selectable marker gene. A selectable marker gene may encode a protein that (a) confers resistance to antibiotics or other toxins, e.g., ampicillin, neomycin, methotrexate, or tetracycline, (b) complements auxotrophic deficiencies, and / or (c) supplies critical nutrients not available from complex media, e.g., the gene encoding D-alanine racemase for Bacilli.

[0146] In certain embodiments, the bacterial expression vector contains one or more of the following regulatory sequences, e.g., one or more promoters (which can be used by the bacterial cell for expression of the gene (e.g., a fitness gene)), one or more ribosome binding sites, one or more termination signals, and the like.

[0147] Non-limiting examples of vectors include plasmids, cosmids, fosmids, viral vectors, phage lambda, bacteriophage Pl, Pl artificial chromosomes (PACs), and bacterial artificial chromosomes (BACs). Exemplary vectors include GMVlc, pCCIFOS, pWE15, pFOS1 , p!ndigoBAC536, pWEB, pSMART, pET, pSE-420, pUC18, pUC19, pUC118, pUC119, pBR322 and its derivatives, Lambda ZAP, pHOS2, pUC and its derivatives, pBluescript and its derivatives, M13. pTOPO-XL, and pCF430.

[0148] Vectors that may be used with the present compositions and methods include pBbS6C-RFP (Addgene, 35293), the Bifidobacterium-Escherichia coli shuttle vector series (e.g., pKO403, www.addgene.org / 174725 / , H. Altaib et al., Bifidobacterium-Escherichia coli Shuttle Vector Series pKO403, with Temperature-Sensitive Replication Origin for Gene Knockout in Bifidobacterium, Microbiol. Resour. Announc. 11, e0088421 (2022)), and pTRK892 for gene expression in Lactobacillus gasseri (www.addgene.org / 71803 / ; Mimee et al., Programming a Human Commensal Bacterium, Bacteroides thetaiotaomicron, to Sense and Respond to Stimuli in the Murine Gut Microbiota, Cell Syst. 2, 214 (2016)). Bacteroides species (e.g., pSIEl which can achieve targeted genetic manipulation of diverse wild-type Bacteroides species from the human gut) and B. thetaiotaomicron may be chromosomally edited through existing systems such as the ones in Jones et al., Sequence and characterization of shuttle vectors for molecular cloning in Porphyromonas, Bacteroides and related bacteria, Mol. Oral. Microbiol. 35, 181-191 (2020), and Lim et al., Engineered Regulatory Systems Modulate Gene Expression of Human Commensals in the Gut, Cell, 169, 547- 558:e515 (2017).

[0149] As used herein, the term “recombinant” or “engineered” when used with reference to, e.g., a vector, a cell, a nucleic acid, or a protein, indicates that the vector, cell, nucleic acid, or protein, has been modified by the introduction of a heterologous nucleic acid or protein, or the alteration of a native nucleic acid or protein, or that the material is derived from a cell so modified. Thus, for example, a recombinant cell contains or expresses genes that are not found within the native (non-recombinant) form of the cell, or express native genes that are otherwise abnormally expressed (e.g., overexpressed, under-expressed, or not expressed at all).

[0150] An endogenous nucleic acid sequence (or the encoded protein product of that sequence) in a cell is deemed recombinant or engineered if a heterologous sequence is placed adjacent to the endogenous nucleic acid sequence, such that the expression of this endogenous nucleic acid sequence is altered. In this context, a heterologous sequence is a sequence that is not naturally adjacent to the endogenous nucleic acid sequence, whether or not the heterologous sequence is itself endogenous to the organism (e.g., originating from the same organism or progeny thereof) or exogenous (e.g., originating from a different organism or progeny thereof). By way of example, a promoter sequence can be substituted for the native promoter of a gene in the cell of an organism, such that this gene has an altered expression pattern. This gene would be recombinant or engineered because it is separated from at least some of the sequences that naturally flank it. A vector or nucleic acid is also considered “recombinant” if it contains any modifications that do not naturally occur in the corresponding vector or nucleic acid. For instance, an endogenous gene or coding sequence is considered recombinant or engineered if it contains an insertion, deletion or a point mutation introduced artificially, e.g., by human intervention. A recombinant vector or nucleic acid also includes a vector or nucleic acid integrated into a host cell chromosome at a heterologous site, and a vector or nucleic acid present as an episome in a cell. A nucleic acid sequence, a gene or a protein can be “exogenous” or “heterologous” which means that it is foreign to the cell into which the nucleic acid sequence, gene or protein is being introduced, or that the nucleic acid sequence or gene is homologous to a nucleic acid sequence or gene in the cell but in a position within the host cell nucleic acid in which the sequence or gene is ordinarily not found. A heterologous nucleic acid sequence or protein may be any nucleic acid sequence or protein placed at a location where it does not normally occur. A heterologous nucleic acid sequence or protein may comprise a nucleic acid sequence or protein that does not naturally occur in a cell, or it may comprise the nucleic acid sequence or protein naturally found in the cell, but placed at a non- normally occurring location in the cell.

[0151] In some embodiments, the heterologous nucleic acid sequence or protein is not an endogenous nucleic acid sequence or protein. In certain embodiments, the heterologous nucleic acid sequence or protein is an endogenous nucleic acid sequence or protein that is derived from a different cell. In certain embodiments, the heterologous nucleic acid sequence is a nucleic acid sequence or protein that occurs naturally in a cell but is then relocated to another site where it does not naturally occur, rendering it a heterologous nucleic acid sequence or protein at that new site.

[0152] Bacteria

[0153] The recombinant bacterial expression vector can be introduced into a bacterium by any suitable method. The bacterium may be transformed by a suitable method, including, but not limited to, electroporation, heat shock, calcium phosphate precipitation, biolistic transformation, and sonic transformation. Hanahan et al.. “Plasmid Transformation of Escherichia coli and Other Bacteria,” Meth. EnzymoL, 204:63-113 (1991), which is hereby incorporated by reference in its entirety. When the vector is a viral vector, the vector is introduced into a bacterium through transduction.

[0154] A recombinant (or engineered) bacterium or bacterial cell as a recipient or host cell includes a bacterial cell that comprises a heterologous nucleic acid, or expresses a peptide or protein encoded by a heterologous nucleic acid. A host cell can contain genes that are not found within the native (nonrecombinant) form of the cell, genes found in the native form of the cell where the genes are modified and re-introduced into the cell by artificial means, or a nucleic acid endogenous to the cell that has been artificially modified without removing the nucleic acid from the cell.

[0155] The recipient / host bacterium may be any suitable bacterium. In one embodiment, the recipient bacterium has low in vivo fitness or may be non-adapted in vivo. In another embodiment, the recipient bacterium is a commensal bacterium. In still another embodiment, the recipient bacterium is a probiotic bacterium.

[0156] The bacterium may be a bacterium that exists naturally in the GI tract. The bacteria may be a mixture of bacteria, from, e.g., a metagenomic source. For example, the bacterium may be from a natural gut microbiota of the gut of a healthy mammal, or from the gut of an unhealthy mammal (e.g., with a gastrointestinal disorder). The unhealthy mammal may have a condition influenced by the GI tract microbiota. For example, the unhealthy mammal may have inflamed GI tract, such as an inflammatory bowel disease (IBD), including ulcerative colitis or Crohn's disease. The unhealthy mammal may have collagenous colitis, lymphocytic colitis, diversion colitis, Behcet’s disease, indeterminate colitis, irritable bowel syndrome (IBS, or spastic colon), mucous colitis, microscopic colitis, antibiotic-associated colitis, constipation, diverticulosis, polyposis coli or colonic polyps.

[0157] The bacterium may be Gram-positive or Gram-negative.

[0158] The bacterium may be from any of the following phyla: Fusobacteriota, Bacillota, and Spirochaetota, Bacteroidetes, Firmicutes, Actinobacteria, Proteobacteria, etc.

[0159] The bacterium may be from any of the following genera: Bacteroides, Clostridium, Bifidobacterium, Lactobacillales (lactic acid bacteria or LAB), Lactobacillus, Lactococcus, Enterococcus, Streptococcus, Klebsiella, Escherichia, Enterobacter, Peptostreptococcus, Peptococcus, Bacillus, Propionibacteria, Ruminococcus, Gemmiger, Desulfomonas, etc.

[0160] The bacterium may be from any of the following: Bacteroidia, and Gammaproteobacteria, Actinomycetia, Coriobacteriia, Negativicutes, Campylobacteria, Alphaproteobacteria, Verrucomicrobia, Desulfovibrionia, Spirochaetia, Fusobacteriia, etc.

[0161] The bacterium may be Bacteroides thetaiotaomicron, Clostridium butyricum, Clostridium bolteae, Anaerotruncus colihominis , Sellinimonas intestinalis, Clostridium symbiosum, Blautia producta, Dorea longicatena, Clostridium innocuum, Flavonifractor plantd' . Bacteroides fragilis, Bacteroides melaninogenicus, Bacteroides oralis, Enterococcus faecalis, Escherichia coli, Bifidobacterium bifidum, Staphylococcus aureus, Clostridium perfringens, Proteus mirabilis, Clostridium tetani, Clostridium septicum, Pseudomonas aeruginosa, Salmonella enteritidis, Bifidobacterium longum, Bifidobacterium lactis, Bifidobacterium animalis, Bifidobacterium breve, Bifidobacterium infantis, Lactobacillus lactis, Lactobacillus gasseri, Lactobacillus acidophilus, Lactobacillus casei, Lactobacillus salivarius, Lactococcus lactis, Lactobacillus reuteri, Lactobacillus rhamnosus, Lactobacillus paracasei, Lactobacillus johnsonii, Lactobacillus plantarum, Lactobacillus salivarius, and Enterococcus faecium. U.S. Patent No. 8,591,880.

[0162] In certain embodiments, the bacterium is of the strains E. coli K.-12, E. coli MG1655, E. coli HS, or E. coli Nissle 1917.

[0163] Non-limiting examples of the bacterium also include: Bacillus subtilis, Bacillus coagulans, B. lentus. Bacillus licheniformis, B. mesentericus, B. pumilus, B. natto, Bacteroides thetaiotaomicron, Bacteroides fragilis, Bacteroides ovatus, C. scindens, Akkermansia muciniphila, Bacteroides amylophilus, Bac. capillosus, Bac. ruminocola, Bac. suis, Bifidobacterium adolescentis, B. animalis, B. breve, B. pseudolongum, B. thermophilum, Enterococcus cremoris, E. diacetylactis, E. intermedins, E. lactis, E. muntdi, E. thermophilus, Kluyveromyces fragilis, L. alimentarius, L. amylovorus, L. crispatus, L. brevis, L. curvatus, L. cellobiosus, Lactobacillus delbrueckii, L. farciminis, L. fermentum, L. gasseri, L. helveticus, L. sakei, L. salivarius, Leuconostoc mesenteroides, Pediococcus damnosus, Pediococcus acidilactici, P. pentosaceus, Propionibacteriumfreudenreichii, Prop, shermanii, Staphylococcus carnosus, Staph, xylosus, Streptococcus typhimurium, Streptococcus infantarius, Strep, salivarius, Streptococcus thermophiles, Strep. Lactis, and E. mundtii.

[0164] The recombinant bacterium may be introduced to the GI tract of one or more mammals. The mammal may be gnotobiotic. The mammal may be germ-free. The mammal may have an already established microbiota.

[0165] The bacteria carrying the bacterial fitness gene may be en terally administered to a mammal. For example, the recipient bacteria can be introduced by gavage to a mouse.

[0166] Beneficial genes that improve gut colonization leads to a measurable enrichment of their relative abundance across the microbial population. As used herein, the term “enrich” refers to an increase in abundance (or percentage or concentration) of a particular group of bacteria.

[0167] Due to the colonization factor or fitness gene, the abundance of the bacterium overexpressing or expressing the colonization factor at a time point (e.g., on day 3, 4, 5, 6. 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30; after 1 , 2, 3 or 4 weeks; after 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months; or after 1, 2, 3, 4, 5 years or longer, counting from the delivery of the recombinant bacterium or composition to the subject) may be at least or about 1.2 fold, at least or about 1.4 fold, at least or about 1.5 fold, at least or about 1.8 fold, at least or about 2 fold, at least or about 3 fold, at least or about 4 fold, at least or about 5 fold, at least or about 6 fold, at least or about 7 fold, at least or about 8 fold, at least or about 10 fold, at least or about 15 fold, at least or about 20 fold, at least or about 25 fold, at least or about 30 fold, at least or about 35 fold, at least or about 40 fold, at least or about 50 fold, at least or about 60 fold, at least or about 70 fold, at least or about 80 fold, at least or about 90 fold, at least or about 100 fold, at least or about 120 fold, at least or about 150 fold, at least or about 175 fold, at least or about 200 fold, at least or about 220 fold, at least or about 250 fold, at least or about 270 fold, at least or about 300 fold, at least or about 320 fold, at least or about 350 fold, at least or about 370 fold, at least or about 400 fold, at least or about 500 fold, at least or about 600 fold, at least or about 700 fold, at least or about 800 fold, at least or about 900 fold, at least or about 1,000 fold, at least or about 1,100 fold, at least or about 1,200 fold, at least or about 1,300 fold, at least or about 1,400 fold, at least or about 1,500 fold, at least or about 1,600 fold, at least or about 1,700 fold, at least or about 1,800 fold, at least or about 1.900 fold, or at least or about 2,000 fold, compared to abundance of the bacterium which does not overexpress the colonization factor or which does not express the colonization factor, after delivery to the subject at the corresponding time point.

[0168] Functional relevance of enriched genes may be assessed by metabolic pathway analysis using the KEGG (Kyoto Encyclopedia of Genes and Genomes) and COG databases.

[0169] The composition and abundance of the established microbiota can be studied by sequencing the 16S ribosomal RNA (or 16S rRNA) gene. 16S rRNA is a component of the 30S small subunit of prokaryotic ribosomes.

[0170] The composition and abundance of the established microbiota may be determined by screening bacterial 16S rRNA genes using PCR.

[0171] In one embodiment, changes in a mammalian gut bacterial populations are assessed by fluorescent in situ hybridization (FISH) with 16S rRNA probes. These 16S rRNA probes, specific for predominant classes of the gut microflora (bacteroides, bifidobacteria, Clostridia, and lactobacilli / enterococci), are tagged with fluorescent markers. For example, the probes can include Bifl64 (Langendijk et al., Appl. Environ. Microbiol., 61: 3069-3075 (1995)), Bac303 (Manz, Microbiology, 142: 1097-1106 (1996)), Hisl50 (Franks, Appl. Environment. Microbiol., 64: 3336- 3345 (1998)), and Labl58 (Harmsen et al., Microbial Ecology Health Disease, 11: 3-12 (1996)). The nucleic acid stain 4'6-diamidino-2-phenylindole (DAPI) may be used for total bacterial counts. Fermentation samples are diluted and fixed in paraformaldehyde. These cells are then washed and resuspended. The cell suspension is then added to the hybridization mixture and filtered. Hybridization is carried out at appropriate temperatures for the probes. Subsequently, the hybridization mix is vacuum filtered and the filter mounted on a microscope slide and examined using fluorescence microscopy, such that the bacterial groups could be enumerated (Ryecroft et al.. J. Appl. Microbiol., 91: 878 (2001)). U.S. Patent No. 8,313,789. Recombinant bacteria engineered with genes that enhance fitness

[0172] The present disclosure also involves recombinant bacteria engineered with genes that can enhance bacterial fitness in the mammalian GI tract (“fitness genes”). The gene can be exogenous or endogenous. When the gene is endogenous, it can be overexpressed (i.e., having a generally higher expression than the gene in its natural form) and / or constitutively expressed. For example, the overexpressed fitness gene (encoding a colonization factor) can express its encoded protein at a level at least or about 1.2 fold, at least or about 1.4 fold, at least or about 1.5 fold, at least or about 1.8 fold, at least or about 2 fold, at least or about 3 fold, at least or about 4 fold, at least or about 5 fold, at least or about 6 fold, at least or about 7 fold, at least or about 8 fold, at least or about 9 fold, at least or about 10 fold, at least or about 15 fold, at least or about 20 fold, at least or about 25 fold, at least or about 30 fold, at least or about 35 fold, at least or about 40 fold, at least or about 50 fold, at least or about 60 fold, at least or about 70 fold, at least or about 80 fold, at least or about 90 fold, at least or about 100 fold, at least or about 200 fold, of the protein expression level of the gene in its natural (or endogenous) form.

[0173] The gene may be integrated into the bacterial chromosome or may be episomal.

[0174] The gene may be codon-optimized.

[0175] Expression of the gene requires that appropriate signals be provided in the vectors, and which include various regulatory elements, such as promoters that drive expression of the genes of interest in host cells.

[0176] The terms “under the control”, “under transcriptional control”, “operatively positioned”, and “operatively linked” mean that a promoter is in a correct functional location and / or orientation in relation to a nucleic acid sequence or a gene to control transcriptional initiation and / or expression of that sequence or gene.

[0177] A promoter may be one naturally associated with a gene or sequence, as may be obtained by isolating the 5' non-coding sequences located upstream of the coding segment and / or exon. Such a promoter can be referred to as “endogenous”. Alternatively, certain advantages will be gained by positioning the gene under the control of a recombinant or heterologous promoter, which refers to a promoter that is not normally associated with a gene in its natural environment. Such promoters may include promoters of other genes, and promoters isolated from any other prokaryotic, viral, or eukaryotic cell, and promoters or enhancers not naturally occurring, i.e., containing different elements of different transcriptional regulatory regions, and / or mutations that alter expression.

[0178] The promoters employed may be constitutive, inducible, and / or useful under the appropriate conditions to direct high-level expression of the introduced DNA segment. The promoter may be exogenous (heterologous) or endogenous.

[0179] A constitutive promoter is an unregulated promoter that allows for continual transcription of the gene under the promoter’s control. Non-limiting examples of constitutive promoters include the constitutive E. coli o70promoter, constitutive E. coli ospromoter, constitutive E. coli o32promoter, constitutive E. coli o54promoter, constitutive E. coli o70promoter, constitutive B. subtilis <JApromoter, constitutive B. subtilis oBpromoter, T7 promoter, and SP6 promoter. A list of constitutive bacterial promoters may be found in the database of Registry of Standard Biological Parts. They are active in all circumstances in the cell. In one embodiment, the constitutive promoter is pL.

[0180] Non-limiting examples of the promoters include the T7 phage promoter, lac promoter, trp promoter, recA promoter, ribosomal RNA promoter, a hybrid trp-lacUV5 (tac) promoter, the PR and PL promoters of coliphage lambda and others, including but not limited, to lacUV5, ompF, Ipp, and the like.

[0181] The gene may be under the control of an inducible promoter. The transcriptional activity of these promoters may be induced by chemical or physical factors. Chemically regulated inducible promoters may include promoters whose transcriptional activity is regulated by the presence or absence of oxygen, a metabolite, alcohol, tetracycline, steroids, metal and other compounds. Physically regulated inducible promoters, including promoters whose transcriptional activity is regulated by the presence or absence of heat, low or high temperatures, acid, base, or light. In one embodiment, the inducible promoter is pH- sensitive (pH inducible).

[0182] The inducer for the inducible promoter may be located in the biological tissue or environmental medium to which the composition is administered or targeted, or is to be administered or targeted. For example, the inducer for the inducible promoter may be located in the mammalian GI tract.

[0183] The pH level of a particular biological tissue can affect the inducibility of the pH inducible promoter. See. for example, Boron, et al., Medical Physiology: A Cellular and Molecular Approach. Elsevier / Saunders. (2004), ISBN 1-4160-2328-3. which is incorporated herein by reference.

[0184] Examples of acid inducible promoters include, but are not limited to, P170, Pl, P3, baiAl, baiA3, lipF promoter. FiFo-ATPase promoter, gadC, gad D, glutamate decarboxylase promoter, etc. See, for example, Cotter and Hill, Microbiol, and Mol. Biol. Rev. vol. 67, no. 3, pp. 429-453 (2003); Hagenbeek, et al., Plant Phys., vol. 123, pp. 1553-1560 (2000); Madsen, et al., Abstract, Mol. Microbiol, vol. 56, no. 3, pp. 735-746 (2005); U.S. Pat. No. 6,242,194; Richter, et al.. Abstract, Gene, vol. 395, no. 1-2, pp. 22-28 (2007), Mallonee, et al., J. Bacteriol., vol. 172, no. 12, pp. 7011-7019 (1990); each of which is incorporated herein by reference. U.S. Patent No. 8,852,916.

[0185] Some non-limiting examples of promoters induced by a change in temperature include P2, P7, and PhS. See, for example, Taylor, et al, Cell, Abstract, vol. 38, no. 2, pp. 371-381 (1984); U.S. Pat. No. 6,852,511, Wang, et al., Biochem. and Biophys. Res. Commun. Abstract, vol. 358, no. 4, pp. 1148-1153 (2007), U.S. Pat. No. 7,462,708, each of which is incorporated herein by reference.

[0186] In an embodiment, the acid inducible promoter is inducible at a pH of about 0.0, about 0.5, about 1.0, about 1.5, about 2.0, about 2.5, about 3.0, about 3.5, about 4.0, about 4.5, about 5.0, about 5.5, about 6.0. about 6.5, about 6.6, about 6.7, about 6.8, about 6.9. or any value therebetween or less.

[0187] In an embodiment, the base inducible promoter is inducible at a pH of about 7.1, about 7.5, about 8.0, about 8.5, about 9.0, about 9.5, about 10.0, about 10.5, about 11.0, about 11.5, about 12.0, about 12.5, about 13.0, about 13.5, about 14.0, or any value therebetween or greater.

[0188] Examples of inducers that can induce the activity of the inducible promoters also include, but are not limited to, radiation, temperature change, alcohol, antibiotic, steroid, metal, salicylic acid, ethylene, benzothiadiazole, or other compound. In an embodiment, the at least one inducer includes at least one of arabinose, lactose, maltose, sucrose, glucose, xylose, galactose, rhamnose, fructose, melibiose, starch, inunlin, lipopolysaccharide, arsenic, cadmium, chromium, temperature, light, antibiotic, oxygen level, xylan, nisin, L-arabinose, allolactose, D-glucose, D-xylose, D-galactose, ampicillin, tetracycline, penicillin, pristinamycin, retinoic acid, or interferon. Other examples of inducers include, but are not limited to, at least a portion of one of an organic or inorganic small molecule, clathrate or caged compound, protocell, coacervate, microsphere, Janus particle, proteinoid, laminate, helical rod, liposome, macroscopic tube, niosome. sphingosome, vesicular tube, vesicle, unilamellar vesicle, multilamellar vesicle, multivesicular vesicle, lipid layer, lipid bilayer, micelle, organelle, nucleic acid, peptide, polypeptide, protein, glycopeptide, glycolipid, lipoprotein, lipopolysaccharide, sphingolipid, glycosphingolipid, glycoprotein, peptidoglycan, lipid, carbohydrate, metalloprotein, proteoglycan, chromosome, nucleus, acid, buffer, protic solvent, aprotic solvent, nitric oxide, vitamin, mineral, nitrous oxide, nitric oxide synthase, amino acid, micelle, polymer, copolymer, monomer, prepolymer, cell receptor, adhesion molecule, cytokine, chemokine, immunoglobulin, antibody, antigen, extracellular matrix, cell ligand, zwitterionic material, cationic material, oligonucleotide, nanotube, piloxymer, transfersome, gas, element, contaminant, radioactive particle, radiation, hormone, virus, quantum dot. temperature change, thermal energy, or contrast agent. See. for example, Theys, et al., Abstract, Curr. Gene Ther. vol. 3, no. 3 pp. 207-221 (2003), which is incorporated herein by reference. Kill Switch

[0189] The present recombinant bacterial expression vector or bacterium may comprise a suicide system or a kill switch (see, e.g., U.S. Patent No. 9,487,764 incorporated herein by reference in its entirety). The kill switch is intended to actively kill engineered microbes in the absence of external stimuli, or in response to external stimuli.

[0190] Kill switches can be designed such that a toxin is produced in response to an environmental condition or external signal (e.g., the bacteria are killed in response to an external cue) or, alternatively designed such that a toxin is produced once an environmental condition no longer exists or an external signal is ceased. The kill switch may be triggered by the absence of a particular factor in the environment that induces the production of toxic molecules within the microbe that cause cell death. The kill switch may be triggered by a particular factor in the environment that induces the production of toxic molecules within the microbe that cause cell death.

[0191] The kill switch can be activated by the absence of a single environmental factor, or by the absence of several activators, to induce cell death. The kill switch can be activated by a single environmental factor or may require several activators to induce cell death.

[0192] Bacteria engineered for in vivo administration to treat a disease or disorder may be programmed to die at a specific time after the delivery and expression of a bacterial fitness gene, or after the subject has experienced the therapeutic effect. Specifically, it may be useful to prevent long-term colonization of subjects by the microorganism, spread of the microorganism outside the area of interest (for example, outside the gut) within the subject, or spread of the microorganism outside of the subject into the environment (for example, spread to the environment through the stool of the subject).

[0193] For example, the recombinant bacteria may comprise a toxin gene that is under the control of an inducible promoter (e.g., the toxin gene is expressed in response to an environmental condition(s) and / or signal(s) (such as arabinose)).

[0194] Examples of such toxins that can be used in kill switches include, but are not limited to, bacteriocins. lysins. and other molecules that cause cell death by lysing cell membranes, degrading cellular DNA, or other mechanisms. Such toxins can be used individually or in combination. The switches that control their production can be based on, for example, transcriptional activation, translation (riboregulators), or DNA recombination (recombinase-based switches).

[0195] In one embodiment, an anti-toxin inhibits the activity of the toxin, thereby delaying death of the genetically engineered bacterium. In one embodiment, the genetically engineered bacterium is killed by the bacterial toxin when the bacterial fitness gene encoding the anti-toxin is no longer expressed when the exogenous environmental condition is no longer present.

[0196] In one embodiment, the genetically engineered bacterium is killed by the bacterial toxin. In one embodiment, the genetically engineered bacterium further expresses a heterologous gene encoding an anti-toxin in response to the exogenous environmental condition. In one embodiment, the anti-toxin inhibits the activity of the toxin when the exogenous environmental condition is present, thereby delaying death of the genetically engineered bacterium. In one embodiment, the genetically engineered bacterium is killed by the bacterial toxin when the heterologous gene encoding the anti-toxin is no longer expressed when the exogenous environmental condition is no longer present.

[0197] In one embodiment, a toxin is produced in the presence of an environmental factor or signal. In another embodiment, a toxin may be repressed (or not produced) in the presence of an environmental factor and then produced once the environmental condition or external signal is no longer present.

[0198] In certain embodiments, the disclosure provides recombinant bacterial cells which express one or more heterologous genes (e.g., a bacterial fitness gene, an anti-toxin gene, etc.) upon sensing arabinose or other sugar in the exogenous environment. In this aspect, the recombinant bacterial cells may contain the araC gene, which encodes the AraC transcription factor, as well as one or more genes under the control of the araBAD promoter. In the absence of arabinose, the AraC transcription factor adopts a conformation that represses transcription of genes under the control of the araBAD promoter. In the presence of arabinose, the AraC transcription factor undergoes a conformational change that allows it to bind to and activate the AraBAD promoter, which induces expression of the desired gene.

[0199] Thus, in some embodiments in which one or more heterologous gene(s) are expressed upon sensing arabinose in the exogenous environment, the one or more heterologous genes are directly or indirectly under the control of the araBAD promoter. In some embodiments, the expressed heterologous gene is selected from one or more of the following: a heterologous therapeutic gene (e.g., a fitness gene), a heterologous gene encoding an antitoxin, a heterologous gene encoding a repressor protein or polypeptide, for example, a TetR repressor, a heterologous gene encoding an essential protein not found in the bacterial cell, and / or a heterologous encoding a regulatory protein or polypeptide.

[0200] Arabinose inducible promoters are known in the art, including Para, ParaB, ParaC, and ParaBAD. In one embodiment, the arabinose inducible promoter is from E. coli. hr some embodiments, the ParaC promoter and the ParaBAD promoter operate as a bidirectional promoter, with the ParaBAD promoter controlling expression of a heterologous gene(s) in one direction, and the ParaC (in close proximity to, and on the opposite strand from the ParaBAD promoter), controlling expression of a heterologous gene(s) in the other direction. In the presence of arabinose, transcription of both heterologous genes from both promoters is induced. However, in the absence of arabinose, transcription of both heterologous genes from both promoters is not induced.

[0201] In one exemplary embodiment of the disclosure, the engineered bacteria of the present disclosure contains a kill-switch having at least the following sequences: a ParaBAD promoter operably linked to a heterologous gene encoding a Tetracycline Repressor Protein (TetR), a ParaC promoter operably linked to a heterologous gene encoding AraC transcription factor, and a heterologous gene encoding a bacterial toxin operably linked to a promoter which is repressed by the Tetracycline Repressor Protein (PTetR). In the presence of arabinose, the AraC transcription factor activates the ParaBAD promoter, which activates transcription of the TetR protein which, in turn, represses transcription of the toxin. In the absence of arabinose, however, AraC suppresses transcription from the ParaBAD promoter and no TetR protein is expressed. In this case, expression of the heterologous toxin gene is activated, and the toxin is expressed. The toxin builds up in the recombinant bacterial cell, and the recombinant bacterial cell is killed. In one embodiment, the AraC gene encoding the AraC transcription factor is under the control of a constitutive promoter and is therefore constitutively expressed.

[0202] In one embodiment of the disclosure, the recombinant bacterial cell further comprises an antitoxin under the control of a constitutive promoter. In this situation, in the presence of arabinose, the toxin is not expressed due to repression by TetR protein, and the antitoxin protein builds-up in the cell. However, in the absence of arabinose, TetR protein is not expressed, and expression of the toxin is induced. The toxin begins to build up within the recombinant bacterial cell. The recombinant bacterial cell is no longer viable once the toxin protein is present at either equal or greater amounts than that of the anti-toxin protein in the cell, and the recombinant bacterial cell will be killed by the toxin.

[0203] In another embodiment of the disclosure, the recombinant bacterial cell further comprises an antitoxin under the control of the ParaBAD promoter. In this situation, in the presence of arabinose, TetR and the anti-toxin are expressed, the anti-toxin builds up in the cell, and the toxin is not expressed due to repression by TetR protein. However, in the absence of arabinose, both the TetR protein and the anti-toxin are not expressed, and expression of the toxin is induced. The toxin begins to build up within the recombinant bacterial cell. The recombinant bacterial cell is no longer viable once the toxin protein is expressed, and the recombinant bacterial cell will be killed by the toxin.

[0204] The bacterial toxin may be a lysin, Hok, Fst, TisB, LdrD, Kid, SymE, MazF, FImA, lbs, XCV2162, dinJ, CcdB, ParE, YafO, Zeta, hicB, relB, yhaV, yoeB, chpBK, hipA, microcin B, microcin B17, microcin C, microcin C7-051, microcin J25, microcin ColV, microcin 24, microcin L, microcin D93, microcin L, microcin E492, microcin H47, microcin 147, microcin M, colicin A, colicin El, colicin K, colicin N, colicin U, colicin B, colicin la, colicin lb, colicin 5, colicinlO, colicin S4, colicin Y, colicin E2, colicin E7. colicin E8, colicin E9, colicin E3, colicin E4, colicin E6; colicin E5, colicin D, colicin M, and cloacin DF13, a biologically active fragment thereof, or combinations thereof.

[0205] The anti-toxin may be an anti-lysin, Sok, RNAII, IstR, RdID, Kis, SymR, MazE, FImB, Sib, ptaRNAl, yafQ, CcdA, ParD, yafN. Epsilon, HicA, relE, prlF, yefM, chpBI, hipB, MccE, MccECTD, MccF, Cai, ImmEl, Cki, Cni, Cui, Cbi, lia, Imm, Cfi, ImlO, Csi, Cyi, Im2, Im7, Im8, Im9, Im3, Im4, ImmE6, cloacin immunity protein (Cim), ImmE5, ImmD, and Cmi, a biologically active fragment thereof, or combinations thereof.

[0206] In one embodiment, the bacterial toxin is bactericidal to the recombinant bacterium. In one embodiment, the bacterial toxin is bacteriostatic to the recombinant bacterium.

[0207] In certain embodiment, the fitness gene (encoding the colonization factor) and an anti-toxin gene are under the control of one or more inducible promoters, and the recombinant bacterial expression vector further comprises a kill switch (e.g., toxin- antitoxin system): when an inducer, or an inducing condition, is present, both the colonization factor and antitoxin (e.g., MazE) are produced, to enhance colonization and to neutralize a constitutively expressed toxin (e.g., MazF); when the inducer, or inducing condition, is no longer provided or no longer present, the cell will be killed by the toxin.

[0208] Conditions to be treated

[0209] The present vectors, bacteria, compositions and methods may be used for the treatment and / or prophylaxis of a disorder.

[0210] The method may comprise administering an effective amount of the present bacteria or composition to the subject.

[0211] The method may comprise administering a therapeutically effective amount of the present bacteria or composition to the subject. A therapeutically effective amount is an amount effective to ameliorate one or more symptoms of the disorder.

[0212] The method may comprise administering a prophylactically effective amount of the present bacteria or composition to the subject. A prophylactically effective amount is an amount which prevents the onset of one or more symptoms of the disorder. The present compositions and methods may be used for the treatment and / or prophylaxis of a disorder associated with the presence in the gastrointestinal tract of a mammalian host of abnormal (or an abnormal distribution of) microbiota.

[0213] Such disorders include, but are not limited to, the following conditions: gastrointestinal disorders including irritable bowel syndrome (IBS, or spastic colon), and intestinal inflammation, functional bowel disease (FBD), including constipation predominant FBD, pain predominant FBD, upper abdominal FBD, non-ulcer dyspepsia (NUD), gastroesophageal reflux, inflammatory bowel disease including Crohn's disease, ulcerative colitis, indeterminate colitis, collagenous colitis, microscopic colitis, Clostridium difficile infection, pseudomembranous colitis, mucous colitis, antibiotic associated colitis, idiopathic or simple constipation, diverticular disease, AIDS enteropathy, small bowel bacterial overgrowth, coeliac disease, polyposis coli, colonic polyps, chronic idiopathic pseudo obstructive syndrome; chronic gut infections with specific pathogens including bacteria, viruses, fungi and protozoa (e.g., Clostridium difficile infection (CDI)); viral gastrointestinal disorders, including viral gastroenteritis, Norwalk viral gastroenteritis, rotavirus gastroenteritis, AIDS related gastroenteritis; liver disorders such as primary biliary cirrhosis, primary sclerosing cholangitis, fatty liver or cryptogenic cirrhosis; rheumatic disorders such as rheumatoid arthritis, non-rheumatoid arthritis, non-rheumatoid factor positive arthritis, ankylosing spondylitis, Lyme disease, and Reiter's syndrome; immune mediated disorders such as glomerulonephritis, hemolytic uremic syndrome, type 1 diabetes mellitus, type 2 diabetes mellitus, mixed cryoglobulinemia, polyarteritis, familial Mediterranean fever, amyloidosis, scleroderma, systemic lupus erythematosus, and Behcet’s syndrome; autoimmune disorders including systemic lupus, idiopathic thrombocytopenic purpura, Sjogren's syndrome, hemolytic uremic syndrome or scleroderma; neurological syndromes such as chronic fatigue syndrome, migraine, multiple sclerosis, amyotrophic lateral sclerosis, myasthenia gravis, Gillain-Barre syndrome, Parkinson's disease, Alzheimer's disease, Chronic Inflammatory Demyelinating Polyneuropathy, and other degenerative disorders; psychiatric disorders including chronic depression, schizophrenia, psychotic disorders, manic depressive illness; regressive disorders including Asperger's syndrome, Rett syndrome, autism, attention deficit hyperactivity disorder (ADHD), and attention deficit disorder (ADD); sudden infant death syndrome (SIDS), anorexia nervosa; and dermatological conditions such as, chronic urticaria, acne, dermatitis herpetiformis and vasculitis disorders. U.S. Patent Publication No. 20140234260. Ozdemir et al., Synthetic Biology and Engineered Live Biotherapeutics: Toward Increasing System Complexity, Cell Syst. 2018, 7( 1):5- 16. Kim et al., Systems and synthetic biology-driven engineering of live bacterial therapeutics, Front Bioeng. Biotechnol. 2023; 11:1267378. Li et al., Function of Akkermansia muciniphila in type 2 diabetes and related diseases, Front Microbiol. 2023; 14:1172400.

[0214] Other metabolic disorders that can be treated or prevented by the present compositions and methods include obesity, insulin resistance, hyperglycemia, hepatic steatosis, and small intestinal bacterial overgrowth (SIBO). U.S. Patent No. 8,110,177.

[0215] The present compositions and methods may be used for the treatment and / or prophylaxis of a disorder, such as cancer, hypercholesterolemia, colitis, diabetes, hyperlipidemia, a liver disease, Lyme disease, mucosal injury, obesity, tetanus, a cardiovascular disease, and cognitive impairment.

[0216] The present compositions and methods may be used for the treatment and / or prophylaxis of an infection, caused by, or associated with, E. coli, enterotoxigenic E. coli, Helicobacter pylori, S. enteritidis, S. typhimurium, Streptococcus, Vibrio cholerae, and / or HIV (human immunodeficiency virus).

[0217] The present compositions and methods may be used to treat an infection caused by, or associated with, antibiotics resistant pathogens such as Clostridium difficile or vancomycin-resistant Enterococcus.

[0218] The present compositions and methods may be used for immune modulation to treat an autoimmune disease.

[0219] The present compositions and methods may be used for immune modulation to enhance the efficacy of cancer immune therapy.

[0220] The present compositions and methods may be used for drug delivery or detoxification (e.g.. to treat rare diseases such as hyperoxaluria).

[0221] The present disclosure provides for compositions and methods for the treatment or prophylaxis of bacterial infections or viral infections.

[0222] For prophylaxis, the present composition can be administered to a subject in order to prevent the onset of one or more symptoms of a bacterial infection or viral infection. In one embodiment, the subject can be asymptomatic. The subject may have been, or have not been, exposed to the bacterium or vims. A prophylactically effective amount of the agent or composition is administered to such a subject. A prophylactically effective amount is an amount which prevents the onset of one or more symptoms of the bacterial infection or viral infection.

[0223] The present composition can be administered to a subject to treat a bacterial infection or viral infection. In one embodiment, the subject is symptomatic. In another embodiment, the subject can be asymptomatic.

[0224] The bacterial infections may be a nosocomial infection, and / or an opportunistic infection.

[0225] The bacterial infections may be a respiratory tract infection, a pulmonary tract infection, respiratory pneumonia, a urinary tract infection, a blood infection, an ear infection, an eye infection, a central nervous system infection, a surgical site wound infection, bacteremia, a gastrointestinal tract infection, a bone infection, a joint infection, a skin infection, a bum infection, a wound infection, dental plaque, gingivitis, chronic sinusitis, endocarditis, or combinations thereof. The infection may be of the pulmonary tract and may be pneumonia.

[0226] The subject may have cystic fibrosis, and / or primary ciliary dyskinesia. The subject may be immunocompromised or immunosuppressed. The subject may be undergoing, or has undergone, surgery, implantation of a medical device, and / or a dental procedure. For example, the medical device can be a catheter, a joint prosthesis, a prosthetic cardiac valve, a ventilator, a stent, an intrauterine device, or combinations thereof. The treatment may be therapeutic or prophylactic. In certain embodiments, the present compositions and methods are used prophylactically when the subject is undergoing surgery, a dental procedure or implantation of a medical device.

[0227] The present compositions and methods may be used for the treatment and / or prophylaxis of a disorder associated with the presence in the gastrointestinal tract of a mammalian host of abnormal microbiota (or an abnormal distribution of microbiota). The method comprises administering an effective amount of the present composition.

[0228] The present compositions and methods may be used to treat or prevent pathogen colonization. The pathogen colonization may comprise gastrointestinal pathogen colonization, such as multidrugresistant (MDR) bacterial colonization. The pathogen colonization may comprise colonization by antibiotic-resistant pathogens. The pathogen colonization may comprise colonization by vancomycin- resistant Enterobacteriaceae (VRE), beta-lactamase (ESBL) producing Gram-negative bacteria, Klebsiella pneumonia carbapenemase (KPC)-producing bacteria, methicillin-resistant Staphylococcus aureus (MRSA), or combinations thereof. The pathogen colonization may comprise colonization by Enterobacteriaceae, Staphylococcus, Pseudomonas, or combinations thereof.

[0229] The present compositions and methods may be used for delivery of beneficial metabolites such as short-chain fatty acids (e.g., for colonization resistance to C. difficile').

[0230] The present compositions and methods may be administered to a subject who is at high risks for infection, such as a transplant recipient (e.g., a bone marrow transplant recipient, a solid organ transplant recipient, a tissue transplant recipient, etc.), or other patients who are immunosuppressed or immunocompromised.

[0231] The present recombinant bacteria and compositions may additionally confer benefits to a subject. These additional benefits are generally known to those skilled in the art and may include managing lactose intolerance, prevention of colon cancer, lowering cholesterol, lowering blood pressure, improving immune function and preventing infections, reducing inflammation and / or improving mineral absorption.

[0232] Bacterial Infections

[0233] The present compositions and methods may be used to treat, or treat prophylactically, bacterial infection. The bacterial infection may be caused by, or associated with. Gram-negative or Grampositive bacteria. For example, the bacterial infection may be caused by, or associated with, bacteria from one or more of the families Clostridium, Pseudomonas, Escherichia, Klebsiella, Enterococcus, Enterobacter, Serratia, Morganella, Yersinia, Salmonella, Proteus, Pasteurella, Haemophilus, Citrobacter, Burkholderia, Brucella, Moraxella, Mycobacterium, Streptococcus or Staphylococcus. Particular examples include Clostridium, Pseudomonas, Escherichia, Klebsiella, Enterococcus, Enterobacter, Streptococcus and Staphylococcus. The bacterial infection may be caused by, or associated with, one or more bacteria selected from Moraxella catarrhalis, Brucella abortus, Burkholderia cepacia, Citrobacter species, Escherichia coli, Haemophilus Pneumonia, Klebsiella Pneumonia, Pasteurella multocida, Proteus mirabilis, Salmonella typhimurium, Clostridium difficile, Yersinia enterocolitica Mycobacterium tuberculosis, Staphylococcus aureus, group B streptococci, Streptococcus Pneumonia, and Streptococcus pyogenes, e.g., from E. coli and K. pneumoniae.

[0234] For example, the bacterial infection may be caused by, or associated with, gram-negative bacteria including, but not limited to, Pseudomonas (including, but not limited to Pseudomonas aeruginosa). Burkholderia cepaci, C. violaceum, V harveyi, Neisseria gonorrhoeae, Neisseria meningitidis, Bordetell pertussis, Haemophilus influenzae, Legionella pneumophila, Brucella, Francisella, Xanthomonas, Agrobacterium, enteric bacteria, such as Escherichia coli and its relatives, the members of the family Enterobacteriaceae, such as Salmonella and Shigella, Proteus, and Yersinia peslis. U.S. Patent No. 9,751,851.

[0235] Gram-negative bacteria that can be inhibited by the present compositions include, but are not limited to, Pseudomonas (including, but not limited to Pseudomonas aeruginosa), Burkholderia cepaci, C. violaceum, V harveyi, Neisseria gonorrhoeae, Neisseria meningitidis, Bordetell pertussis, Haemophilus influenzae, Legionella pneumophila, Brucella, Francisella, Xanthomonas, Agrobacterium, enteric bacteria, such as Escherichia coli and its relatives, the members of the family Enterobacteriaceae, such as Salmonella and Shigella, Proteus, and Yersinia pestis.

[0236] The present compositions and methods can be used to treat, or treat prophylactically, infections of the pulmonary tract, urinary tract, bums, and wounds, caused by, or associated with, gram negative bacteria such as P. aeruginosa. The present compositions and methods can be used to treat, or treat prophylactically, catheter-associated infections, blood infections, middle ear infections, formation of dental plaque, gingivitis, chronic sinusitis, endocarditis, coating of contact lenses, and infections associated with implanted devices (e.g., catheters, joint prostheses, prosthetic cardiac valves and intrauterine devices), caused by, or associated with, gram negative bacteria such as P. aeruginosa. The present compositions and methods can be used to treat, or treat prophylactically, infections of the central nervous system, gastrointestinal tract, bones, joints, ears and eyes, caused by, or associated with, gram negative bacteria such as P. aeruginosa.

[0237] The present compositions and methods can be used to treat, or treat prophylactically, inhibit, and / or ameliorate infections including opportunistic infections and / or antibiotic -resistant bacterial infections caused by gram negative bacteria. Examples of such opportunistic infections, include, but are not limited to P. aeruginosa, or poly-microbial infections of P. aeruginosa with, for example, Staphylococcus aureus or Burkholderia cepacia.

[0238] The present compositions and methods can be used to treat, or treat prophylactically, burns and / or other traumatic wounds as well as common or uncommon infections. Examples of such wounds and infection disorders include, but are not limited to puncture wounds, radial keratotomy, ecthyma gangrenosum, osteomyelitis, external otitis, and / or dermatitis.

[0239] In one embodiment, the present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate pulmonary infections. In one embodiment, the present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate pneumonia. Pneumonia can be caused by colonization of medical devices, such as ventilator- associated pneumonia, and other nosocomial pneumonia. In one embodiment, the present compositions and methods can be used to treat, treat prophy tactically, prevent, and / or ameliorate lung infections, such as pneumonia, in cystic fibrosis patients. In one embodiment, the present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate an infection caused by, or associated with, gram negative bacteria (such as by P. aeruginosa) in cystic fibrosis patients.

[0240] The present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate septic shock. The present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate septic shock in neutropenic, immunocompromised, and / or immunosuppressed patients or patients infected with antibiotic resistant bacteria, such as, for example, antibiotic resistant P. aeruginosa.

[0241] The present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate urinary tract or pelvic infections. The present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate gastrointestinal infections, such as necrotizing enterocolitis, often seen in premature infants and / or neutropenic cancer patients.

[0242] The present compositions and methods can be used to treat, treat prophylactically, prevent, and / or ameliorate urinary dysentery (for example, dysentery caused by bacillary dysentery), food poisoning and / or gastroenteritis (for example, caused by Salmonella enterica), typhoid fever (for example, caused by Salmonella typhi), whooping cough (or pertussis) as is caused by Bordetella pertussis, Legionnaires' pneumonia, caused by Legionella pneumophila, sexually transmitted diseases, such as gonorrhea, caused by Neisseria gonorrhoeae, or meningitis, caused by, for example, Neisseria meningitidis or Haemophilus influenzae, brucellosis which is caused by brucellae, and more specifically, Brucella abortus.

[0243] The present compositions and methods may be used to attenuate bacterial virulence.

[0244] In one embodiment, the present compositions are administered to a subject who is free of bacterial disease. Administration may be in advance of an anticipated health-related procedure known to increase susceptibility to gram-negative bacteria (e.g., P. aeruginosa) pathogenicity, for example, in advance of a surgical procedure, including dental procedures, procedures involving implants, and / or insertion of catheters or other devices.

[0245] The present composition may be administered alone or in combination with other compounds, such as an antibiotic, an antimicrobial agent, a probiotic, and / or an anti-inflammatory agent. In certain embodiments, the present composition may be administered in combination with one or more antibiotics. Combinations may be administered either concomitantly, e.g., as an admixture, separately but simultaneously or concurrently; or sequentially. This includes presentations in which the combined agents are administered together as a therapeutic mixture, and also procedures in which the combined agents are administered separately but simultaneously. Administration "in combination" further includes the separate administration of one of the compounds or agents given first, followed by the second.

[0246] This may be achieved by administering a composition that includes both agents (e.g., an antibiotic and the present bacteria, or an antimicrobial agent and the present bacteria, or by administering two compositions, at the same time or within a short time period, wherein one composition comprises an antibiotic, and the other composition includes the present bacteria.

[0247] Probiotic Compositions and Pharmaceutical Compositions

[0248] The present disclosure provides for a probiotic composition or pharmaceutical composition comprising the present recombinant bacteria. The recombinant bacteria may express one or more of the colonization factors.

[0249] The present disclosure also provides for a pharmaceutical composition comprising the present recombinant bacterial expression vector. The recombinant bacterial expression vector may comprise the bacterial fitness gene encoding one or more of the colonization factors.

[0250] Probiotics are microorganisms, or processed compositions of microorganisms which beneficially affect a host. Salminen et al., Probiotics: how should they be defined, Trends Food Sci. Technol. 1999: 10 107-10. U.S. Patent No. 8.216,563.

[0251] The present probiotic composition may be administered enterally, such as oral, sublingual and rectal administration.

[0252] The present probiotic composition can be a food composition, a beverage composition, a pharmaceutical composition, or a feedstuff composition.

[0253] The present probiotic composition may comprise a liquid culture. The probiotic composition may be lyophilized, pulverized and powdered. As a powder it can be provided in a palatable form for reconstitution for drinking or for reconstitution as a food additive. The composition can be provided as a powder for sale in combination with a food or drink. The food or drink may be a dairy-based product or a soy-based product. The invention therefore also includes a food or food supplement containing the present composition. Typical food products that may be prepared in the framework of the present invention may be milk-powder based products; instant drinks: ready-to-drink formulations: nutritional powders; milk-based products, such as yogurt or ice cream; cereal products; beverages such as water, coffee, malt drinks; culinary products and soups. The composition can be combined with other adjuvants such as antacids to dampen bacterial inactivation in the stomach. Acid secretion in the stomach could also be pharmacologically suppressed using H2-antagonists or proton pump inhibitors. Typically, the H2-antagonist is ranitidine. Typically the proton pump inhibitor is omeprazole.

[0254] The present composition may further contain one or more of the following: earners, protective hydrocolloids (such as gums, proteins, modified starches), binders, film forming agents, encapsulating agents / materials, wall / shell materials, matrix compounds, coatings, emulsifiers, surface active agents, solubilizing agents (oils, fats, waxes, lecithins etc.), adsorbents, fillers, co-compounds, dispersing agents, wetting agents, processing aids (solvents), flowing agents, taste masking agents, weighting agents, jellifying agents, gel forming agents, antioxidants and antimicrobials.

[0255] The present composition may comprise a source of protein. Any suitable dietary protein may be used, for example animal proteins (such as milk proteins, meat proteins and egg proteins); vegetable proteins (such as soy protein, wheat protein, rice protein, and pea protein); mixtures of free amino acids; or combinations thereof. The proteins may be intact, hydrolyzed, partially hydrolyzed or a mixture thereof.

[0256] The composition may also contain a source of carbohydrates and a source of fat. A source of carbohydrate may be added to the composition. Any suitable carbohydrate may be used, for example sucrose, lactose, glucose, fructose, com syrup solids, maltodextrins, and mixtures thereof.

[0257] The pharmaceutical compositions can be. e.g., in a solid, semi-solid, or liquid formulation. Compositions can also take the form of tablets, pills, capsules, semisolids, powders, sustained release formulations, emulsions, suspensions, or any other appropriate compositions.

[0258] The present composition may be in the form of: an enema composition which can be reconstituted with an appropriate diluent; enteric-coated capsules or microcapsules; powder for reconstitution with an appropriate diluent for naso-enteric infusion, naso-duodenal infusion or colonoscopic infusion; powder for reconstitution with appropriate diluent, flavoring and gastric acid suppression agent for oral ingestion; or powder for reconstitution with food or drink. U.S. Patent Publication No. 20140234260.

[0259] The composition may also contain conventional pharmaceutical additives and adjuvants, excipients and diluents, including, but not limited to, water, gelatine of any origin, vegetable gums, ligninsulfonate, talc, sugars, starch, gum arabic, vegetable oils, polyalkylene glycols, flavouring agents, preservatives, stabilizers, emulsifying agents, buffers, lubricants, colorants, wetting agents, fillers, and the like. In all cases, such further components will be selected having regard to their suitability for the intended recipient. U.S. Patent Publication No. 8,741,622.

[0260] The present compositions may be used in vitro or administered to a subject. The administration may be oral, topical, intranasal, or any other suitable route as described herein.

[0261] Oral dosage forms may be tablets, capsules, bars, sachets, granules, syrups and aqueous or oily suspensions. Tablets may be formed form a mixture of the active compounds with fillers, for example calcium phosphate; disintegrating agents, for example maize starch, lubricating agents, for example magnesium stearate; binders, for example microcrystalline cellulose or polyvinylpyrrolidone and other optional ingredients known in the art to permit tableting the mixture by known methods. Similarly, capsules, for example hard or soft gelatin capsules, may be prepared by known methods. The contents of the capsule may be formulated using known methods so as to give sustained release of the bacteria or vectors. Other dosage forms for oral administration include, for example, aqueous suspensions in an aqueous medium in the presence of a non-toxic suspending agent such as sodium carboxymethylcellulose, and oily suspensions in a suitable vegetable oil, for example arachis oil. The bacteria or vectors may be formulated into granules with or without additional excipients. The granules may be ingested directly by the patient or they may be added to a suitable liquid carrier (e.g. water) before ingestion. The granules may contain disintegrants, e.g. an effervescent pair formed from an acid and a carbonate or bicarbonate salt to facilitate dispersion in the liquid medium. U.S. Patent No. 8,263.662.

[0262] Additional compositions include formulations in sustained or controlled delivery, such as using liposome or micelle carriers, bioerodible microparticles or porous beads and depot injections. The pharmaceutical composition can be prepared in single unit dosage forms.

[0263] Appropriate frequency of administration can be determined by one of skill in the art and can be administered once or several times per day (e.g., twice, three, four or five times daily). The present compositions may also be administered once each day or once every other day. The compositions may also be given twice weekly, weekly, monthly, or semi-annually. In the case of acute administration, treatment is typically carried out for periods of hours or days, while chronic treatment can be carried out for weeks, months, or even years.

[0264] A sufficient dose of the recombinant bacteria is usually consumed per day in order to achieve successful colonization. The daily dose of probiotics in the composition will depend on the particular person or animal to be treated. Important factors to be considered include age, body weight, sex and health condition. Daily doses generally range from about 102to about 1014cfu (colony forming units), from about 102to about 1012cfu, from about 104to about 1012cfu, from about 106to about IO10cfu, from about 106to about 1014cfu, about 107to about 1013cfu, about IO10to about 1014cfu, about 1011to about 1013cfu, about l-4xl012cfu, or from about 107to about 109cfu per day. U.S. Patent No. 8,021,656.

[0265] The dosage of the present recombinant bacteria can be adjusted by those skilled in the art to the designated purpose.

[0266] Appropriate frequency of administration can be determined by one of skill in the art and can be administered once or several times per day (e.g., twice, three, four or five times daily). The compositions of the invention may also be administered once each day or once every other day. The compositions may also be given twice weekly, weekly, monthly, or semi-annually. U.S. Patent No. 8,501,686.

[0267] Different treatment regimens may be used. In some embodiments, a daily dosage of the present composition is administered once, twice, three times, or four times a day, for at least or about 2 days, at least or about 3 days, at least or about 4 days, at least or about 5 days, at least or about 6 days, at least or about 7 days, at least or about 8 days, at least or about 9 days, at least or about 10 days, at least or about 11 days, at least or about 12 days, at least or about 13 days, at least or about 2 weeks, at least or about 3 weeks, at least or about 4 weeks, at least or about 1 month, at least or about 2 months, at least or about 3 months, at least or about 4 months, at least or about 5 months, at least or about 6 months, at least or about 7 months, at least or about 8 months, at least or about 9 months, at least or about 10 months, at least or about 11 months, at least or about 1 year, or longer. Depending on the stage and severity of the condition, a shorter treatment time (e.g., up to five days) may be employed along with a high dosage, or a longer treatment time (e.g., ten or more days, or weeks, or a month, or longer) may be employed along with a low dosage. In some embodiments, a once- or twice-daily dosage is administered every other day.

[0268] Prebiotics

[0269] The present composition may be administered to a subject in combination with a prebiotic. The composition may further contain at least one prebiotic.

[0270] The term "prebiotic" may refer to an agent that increases the number and / or activity of one or more desired bacteria. A prebiotic may be an agent that allows specific changes both in the composition and / or activity in the gastrointestinal microbiota that may (or may not) confer benefits upon the host. Prebiotic may mean food substances intended to promote the growth of probiotic bacteria in the intestines. Prebiotics may be dietary fibers. Dietary fibers may be selected from the group consisting of fructo-oligosaccharides, galacto-oligosaccharides, xylo-oligosaccharides, isomalto- saccharides, soya oligosaccharides, pyrodextrins, transgalactosylated oligosaccharides, lactulose, beta-glucan, insulin, raffinose, stachyose. They furthermore may contribute to the present methods by improving gastrointestinal health and by increasing satiety. The prebiotic may be selected from the group consisting of oligosaccharides and optionally contains fructose, galactose, mannose, soy and / or inulin; and / or dietary fibers.

[0271] In some embodiments, a prebiotic can be a comestible food or beverage or ingredient thereof.

[0272] Prebiotics include, but are not limited to, oligosaccharides and optionally contain fructose, galactose, mannose, soy and / or inulin; and / or dietary fibers. Dietary fibers include, but are not limited to, inulin, inulin-type fructans, fructo-oligosaccharides, oligofructose, galacto-oligosaccharides, xylooligosaccharides, isomalto-saccharides, soya oligosaccharides, pyrodextrins, transgalactosylated oligosaccharides, N-acetylglucosamine, N-acetylgalactosamine, glucose, other five- and six-carbon sugars (such as arabinose, maltose, lactose, sucrose, cellobiose, etc.), amino acids, alcohols, sugar alcohols, resistant starch, lactulose, beta-glucan, raffinose, stachyose, raffinose family oligosaccharides (RFO), and mixtures thereof. See, e.g., Ramirez-Farias et al.. Br J Nutr (2008) 4:1-10; Pool-Zobel and Sauer, J Nutr (2007), 137:2580S-2584S.

[0273] Prebiotics may include complex carbohydrates, amino acids, peptides, minerals, or other essential nutritional components for the survival of the bacterial composition. Prebiotics include, but are not limited to, amino acids, biotin, fructooligosaccharide, galactooligosaccharides, hemicelluloses (e.g., arabinoxylan, xylan, xyloglucan, and glucomannan), inulin, chitin, lactulose, mannan oligosaccharides, oligofructose-enriched inulin, gums (e.g., guar gum, gum arabic and carregenaan), oligofructose, oligodextrose, tagatose, resistant maltodextrins (e.g., resistant starch), transgalactooligosaccharide, pectins (e.g., xylogalactouronan, citrus pectin, apple pectin, and rhamnogalacturonan-I), dietary fibers (e.g., soy fiber, sugarbeet fiber, pea fiber, corn bran, and oat fiber) and xylooligosaccharides. Non-limiting examples of prebiotics include a monomer or polymer of arabinoxylan, xylose, soluble fiber dextran, soluble com fiber, polydextrose, lactose, N-acetyl- lactosamine, glucose, and combinations thereof. Non-limiting examples of prebiotics also include a monomer or polymer of galactose, fructose, rhamnose, mannose, uronic acids, 3'-fucosyllactose, 3'- sialylactose. 6'-sialyllactose, lacto-N-neotetraose, 2'-2'-fucosyllactose, and combinations thereof. Nonlimiting examples of prebiotics include a monosaccharide selected from the group consisting of arabinose, fructose, fucose, galactose, glucose, mannose, D-xylose, xylitol, ribose, and combinations thereof. In one embodiment, the prebiotic comprises a disaccharide selected from the group consisting of xylobiose, sucrose, maltose, lactose, lactulose, trehalose, cellobiose, and combinations thereof. In another embodiment, the prebiotic comprises a polysaccharide, wherein the polysaccharide is xylooligosaccharide. In one embodiment, the prebiotic comprises a sugar selected from the group consisting of arabinose, fructose, fucose, lactose, galactose, glucose, mannose, D-xylose, xylitol, ribose, xylobiose, sucrose, maltose, lactose, lactulose, trehalose, cellobiose, xylooligosaccharide, and combinations thereof. In one embodiment, the sugar is xylose.

[0274] Subjects, which may be treated, include all animals which may benefit from the present compositions and methods. Such subjects include mammals, preferably humans (infants, children, adolescents and / or adults), but can also be an animal such as dogs and cats, farm animals such as cows, pigs, sheep, horses, goats and the like, and laboratory animals (e.g., rats, mice, guinea pigs, and the like).

[0275] The term “about” can mean a range of ±10% of a given value.

[0276] As used herein, the term "bacteria" encompasses both prokaryotic organisms and archaea present in mammalian microbiota.

[0277] The terms "intestinal microbiota", "gut flora", and "gastrointestinal microbiota" are used interchangeably to refer to bacteria in the digestive tract.

[0278] The term "Eubacteria" refers to all bacteria and excludes archaea. In mammals. >90% of all colonic bacteria are in the phyla Firmicutes or Bacteroidetes (Ley et al., Nat Rev Microbiol 2008; 6:776-88).

[0279] The following are examples of the present invention and are not to be construed as limiting.

[0280] EXAMPLES

[0281] Example 1 Conserved Genetic Basis for Microbial Colonization of the Gut

[0282] Despite the fundamental importance of gut microbes, the genetic basis of their colonization remains largely unexplored. Here, by applying cross-species genotype-habitat association at the tree- of-life scale, we identify conserved microbial gene modules associated with gut colonization. Across thousands of species, we discovered 79 taxonomically diverse putative colonization factors organized into operonic and non-operonic modules. They include previously characterized colonization pathways such as autoinducer-2 biosynthesis and novel processes including tRNA-modification and translation. In vivo functional validation revealed YigZ (IMPACT family) and TrhP (tRNA hydroxylation protein- P) are required for E. coli intestinal colonization. Overexpressing YigZ alone is sufficient to enhance colonization of the poorly-colonizing MG1655 E. coli by >100 fold. Moreover, natural allelic variations in YigZ impact inter-strain colonization efficiency. Our findings highlight the power of large-scale comparative genomics in revealing the genetic basis of microbial adaptations. These broadly conserved colonization factors may prove critical for understanding GI dysbiosis and developing therapeutics.

[0283] Here, we present a robust computational framework that overcomes these challenges and enables the systematic analysis of -3.700 microbial species selected from an initial pool of 280K genomes (Figure 1 A, Figure 8 A). We present the initial set of putative gut colonization factors and their modular organization, encompassing both known and previously undescribed gut colonization processes. In vivo functional validation in mice demonstrates that this strategy reveals previously undescribed, bona fide colonization factors of large effect. Furthermore, we show that presence and absence of these factors are stable within various species, making them challenging to identify through comparative genomic analysis limited to closely related species. Collectively, our findings highlight the power of large-scale cross-species association studies in uncovering novel genetic basis for niche colonization.

[0284] Results

[0285] Tree-of-life Scale Genotype-Habitat Association Reveals Conserved Protein Families as Putative Colonization Factors

[0286] We hypothesized that diverse bacterial taxa utilize common, conserved colonization mechanisms, and that the genetic basis of these mechanisms could be inferred by the computational identification of genes specific to mammalian (human and murine) gut microbes compared to those from environmental habitats. To enable this analysis, we carefully selected 9,475 high-quality genomes from an initial pool of 208K genomes with the goal of minimizing genomic redundancy, while maintaining taxonomic and geographical diversity (Figure 1 A step 1 , Figure 8A-8C). These genomes include both culture-dependent microbial isolates and culture-independent metagenome-assembled genomes (MAGs), enabling us to study a broad range of species regardless of their cultivability status across the tree-of-life.

[0287] Accurately identifying homologs across distantly related organisms poses a significant challenge in cross-species genomic comparison at this scale. To address this, instead of clustering proteins based solely on amino acid identity, we employed a computationally intensive, alignment-based method with a customized computational module optimized for speed and accuracy (see Methods for details). This enabled exhaustive identification of homologs across the genomes analyzed on practical time scales (Figure 1 A, step 2; see Methods for details). Proteins from three datasets (mouse MAGs, human MAGs, and human isolates) underwent an all-against-all alignment, which allowed us to generate a reliable binary phylogenetic profile (PP) that reflects the presence or absence of the protein’s homologs across all genomes (Figure 1A step 2, Figure 8D). We then utilized mutual information (MI)37,38to identify proteins whose distribution across genomes (PP) is highly informative of their binary habitat colonization: mammalian gut vs other habitats (Figure 1A, step 2). MI values were converted to MI z- scores using frequency-matched null distributions (Figure 8D). Notably, this database-independent approach allows for assessing the association of every protein with gut colonization, including those of unknown function.

[0288] To this end, we identified 16,307, 5,436, and 5,353 proteins as the top gut colonization-associated proteins (top 0.2%, MI z-scores > 1548, 1087 and 139) from three datasets (mouse MAGs, human MAGs, and human isolates), respectively (Figure IB). Grouping these proteins (n=27,096) based on sequence similarity (Figure 1A, step 3) using hierarchical network clustering yielded 82 protein families (Figure 8E, Methods). This indicates that only a small set of functions are enriched in these high-scoring proteins, as evidenced by simulations in which random sampling of equivalent number of proteins produced an average of 17K protein families (Figure 8F). Re-evaluation of these 82 protein families using a more precise alignment mode yielded 79 families that were significantly associated with at least one dataset (human isolates, human GI MAGs, and mouse GI MAGs) (Methods, Figures 9A-9C), providing a final set of 79 putative Colonization Factors (CFs) (Table 1).

[0289] Taxonomically Diverse Colonization Factors Largely Overlap Between Human and Mouse Gut Microbes

[0290] Notably, 37 of the 79 CFs were common across all datasets (Figure 1C), highlighting potentially common microbial adaptation mechanisms for mouse and human gut colonization that are independent of cultivability. Nine CFs are unique to the human microbes (Figure 1C). Some CFs are potentially culture-dependent with 16 and 13 CFs specific to isolates or MAGs, respectively (Figure 1C). These CFs are robust to variations in the critical alignment parameter: ‘ query protein coverage ’ (Figures 9D- 9E).

[0291] The CFs are highly prevalent in gut microbes, some of which are encoded by more than 90% of gut MAGs analyzed (Figure ID). The non-gut-associated proteins (gray dots in Figure ID, Figures 9F- 9G) and randomly selected protein families (gray dots in Figure 9H) are either rare or prevalent in a nonhabitat distinguishable manner. The mutual information z-scores of CFs in the human and mouse MAG datasets are significantly correlated (Figure IE). Most CFs are orthogonal with no detectable sequence homology (Figure IF). A small number of CFs share homology but are sufficiently divergent to be classified as separate CFs, a distinction captured by our hierarchical homolog consolidation approach (Figure 8E, see Methods). For instance, CFs 2_4 and 2_7, which are from the same high-ranking protein family, show divergence in amino acid sequence (Figure 91) and taxonomic distributions (Figure 9J). We, thus, find that putative CFs are conserved protein families consistently associated with gut colonization regardless of host species (mouse and human) and technical variations (genome recovery method and pipeline alignment parameters), underscoring the statistical robustness of their inference.

[0292] De novo Identification of Colonization Modules Reveals Operonic and Non-Operonic Functional Links Among Colonization Factors

[0293] To better understand the higher-level organization of CFs, we next sought to group them into functionally coherent modules. We previously demonstrated that genes contributing to a phenotype can be organized into modules based on their patterns of cross-species coinheritance (Figure 1A, step 4), revealing functionally coherent pathways or protein complexes38. Consistent with this notion, we observed that cross-genomic occurrences of 79 CFs, as indicated by their PPs, exhibited a highly modular structure (Figure 2A, left). We therefore calculated the pairwise co-inheritance strength among the 79 CFs (Figure 2A, right, Methods), based on which we grouped the CFs into 47 colonization modules (CM), comprising 23 multi-CF CMs (denoted by CM#) and 24 single-CF CMs (Figure 2B, Table 1, and Figure 10A, Methods).

[0294] Investigating the genomic context of multi-CF CMs in three representative species (E. coli, B. thetaiotaomicron, and C. difficile) revealed that CMs are associated with both operonic and non-operonic structures (Figure 2C). For instance, operonic CM1 corresponds to the nrdD-nrdG (anaerobic ribonucleoside-triphosphate reductase), crucial for DNA synthesis (Figure 2D), and CM6 corresponds to an uncharacterized operon that is only present in C. difficile (Figure 2E). The majority of CMs are non-operonic (Figure 2C). Some of them nevertheless capture functionally linked CFs, as shown by CMl l(Figure 2F), which consists of genes encoding Pfs and LuxS that catalyze the two-step reactions for the biosynthesis of quorum sensing molecule autoinducer-2 and L-homocysteine from S- adenosylhomocysteine (SAH)39’40. Non-operonic CMs also contain high-scoring CFs with uncharacterized relationships, such as CM 12 (Figure 2G) that consists of the prevalent IMPACT family protein YigZ (CF9)41and the GTP-binding protein YihA / YsxC (CF69)42’45.

[0295] Some operonic CMs exhibit further gene duplications (Figure 2C) in a species-specific manner (Figure 2H, Figures 1 OB- IOC). For example, CM 18 corresponds to acetate (pta-ackA) and butyrate (ptb- buk) fermentation operons (Figure 2H). E. coli encodes pta-ackA and a second copy of phosphotransacetylase (eutD) (Figure 2H) that is evolutionarily related to pta, which can fulfill the same role as pta under some conditions despite differential regulation and enzymatic efficiency46,47. C. difficile encodes two ptb-buk operons with two additional buk homologs (Figure 2H). CFs and CMs, thus, correspond to broad protein families, and species- specific duplications may indicate their differential utilization / dependency on specific modes of colonization. Some of the CFs correspond to protein families with no functional annotation (Figure 21, Figures 10D-10E). In summary, these multi-CF CMs captured known operonic and non-operonic functional relationships and predicted novel links among CFs.

[0296] Colonization Modules Reveal Known and Novel colonization mechanisms

[0297] To characterize the biological processes enriched in CMs, we functionally annotated them using the Biocyc and UniRef databases (Table 1, Figures 10D-10E, Methods). The CMs include genes functioning in metabolic processes, such as nucleoside catabolism and cofactor metabolism (CM1, CM3, CF46, CF37) (Figure 10F). Other functions involve transporters and ion binding proteins (CM5, CM7, CM14, F54, CM21, CM22, CF49_65, CF0_6), quorum sensing (CM11) and transcriptional regulators (CF38, CF7) (Figure 10G). Many of the CFs have metal ion- or ATP-binding functions (Figure 10G). However, a subset of CMs. including those with the highest association signals (CM3, CF20), are proteins of unknown functions (Figure 21, Figure 10E) or proteins with conserved domains of unknown function (CF21: DUF1846, CF64: DUF2179). In summary, these CMs revealed a compact set of processes including known mechanisms of gut colonization, such as metabolic niche17,21,48, quorum sensing49and sugar PTS systems50,51. However, many of the CMs had not been previously associated with gut colonization, including several genes of unknown function.

[0298] Translation-Related Factors are Among the Top-Scoring CMs

[0299] Intriguingly, the top-scoring CMs include several factors related to tRNA processing and translation (Figure 3A, Figure 10F, Figure 11 A), which, unlike those associated with anaerobic growth (CM1, anaerobic RNR) (Figure 3 A), represent functions not previously linked to gut colonization. Specifically. CM2 includes two tRNA modifying enzymes (CF15 and 22, Figure 11A). CM12 encompasses IMPACT family proteins (CF9) with evidence of genetic interactions with proteins involved in translation41and GTP-binding protein YihA / YsxC (CF69) that directly binds to rRNA affecting ribosome assembly42-45(Figure 11 A). CF7 corresponds to YbaK (Cys-tRNAproand Cys- tRNAcys deacylase) (Figure 11 A). CF24_29 and CF24_46 correspond to epoxyqueuosine reductase (Figure 11 A). The recurrence of these translation-related functions among top CMs motivated us to experimentally determine their role in murine gut colonization.

[0300] TrhP and YigZ are Bona Fide Colonization Factors in Murine Gut Colonization

[0301] E. coli strains exhibit a wide spectrum of colonization efficiencies (Figure 3B). At the two extremes, the MG 1655 (KI 2) laboratory strain is a poor colonizer52 54, while the MP1 strain shows highly robust colonization55, offering the opportunity to evaluate essentiality and sufficiency of CFs in murine gut colonization (Figure 3B). Among all translation-related CFs (Figure 11 A), we focused on the four genes (yigZ, tcdA, trhP, ybaK that satisfied three key criteria (TcdA: tRNA threonylcarbamoyladenosine dehydratase; YbaK: Cys-tRNAPro / Cys-tRNACys deacylase) (Figure 1 IB): (1) significant across all three analyses (human GI MAGs, mouse GI MAGs, and human isolates); (2) part of the E. coli core genome; and (3) non-essential in vitro, allowing for the generation of knockouts. All four genes are present in the genomes of the E. coli strains under study (Figure 3B). However, sequence comparisons revealed amino-acid substitutions in YigZ, TrhP, TcdA, and YbaK in all gut-colonizing strains compared to MG 1655 (Figure 3C). Therefore, we set out to test the contribution of the four CFs in the colonization of E. coli MP1 and MG 1655 strains.

[0302] The E. coli MP1 strain was isolated from mice and exhibits high colonization capacity55. Two florescence-tagged derivatives of MP1 exist: MP7 (mCherry) and MP13 (GFP)55. We generated single knockouts of each of the four genes in the MP13 background for in vivo competition with MP7 as the control strain (Figure 3D). To facilitate colonization, the mice were pretreated with streptomycin for 72 hours, and then recovered for 24 hours as previously described55. The mice were single housed to ensure strict independence across biological samples. Each MP13 strain was grown to the early exponential phase, combined in equal proportions with MP7, and administered to mice via oral gavage (Figure 3D, Figure 11C).

[0303] Colonization of AyigZ and AtrhP strains was significantly diminished in all mice compared to MP7 (Figure 3E). Importantly, AyigZ and AtrhP strains showed no detectable growth defects in vitro (Figure 11D). AtcdA and AybaK showed much higher inter-individual variation (Figure 3E). Heterogeneity in colony morphology was observed in AybaK cells recovered from feces after day 3 postgavage, suggesting that a subpopulation of this strain may have acquired compensatory mutations (Figure 1 IE). In summary, our findings reveal that yigZ and trhP are required for intestinal colonization of E. coli MP1 strain.

[0304] Over-expression of yigZ is Sufficient to Enhance Colonization of E. coli K12 (MG1655)

[0305] The domesticated “WT” E. coli strain K12 MG1655 has been cultured and passaged in lab for nearly a century, and exhibits poor colonization in mice52and humans53,54. Sequence comparisons revealed amino-acid substitutions in YigZ, TrhP, TcdA, and YbaK, in all gut-colonizing strains compared to MG1655 (Figure 3C). We, therefore, wondered whether these CF allelic variants in MG1655 contribute to its reduced colonization potential, and if so, whether complementing these proteins from a gut-colonizing strain (MP1) could enhance MG1655’s colonization capacity.

[0306] To improve quantification sensitivity and throughput, we developed an intergenic genomic barcoding system (Figure 4A, Figures 12A-12B), which allowed us to evaluate the in vivo fitness of multiple strains, in parallel, in a single animal (Figure 4B). We generated 6 barcoded MG1655 mutants and transformed each with a gain-of-function (GOF) plasmid (Figure 4B-C). These GOF plasmids contained three components: (z) pSClOl replication origin, which ensures a tight control of the plasmid copy number (~5 copies per cell)56,57; (z’z) a GOF gene cassette (Figure 4C), each with one of the four CFs, or all combined, cloned from E. coli MP13; (z'zz) the CmR gene that confers chloramphenicol resistance. A plasmid with an empty GOF cassette served as control (Figure 4C). All CFs are driven by their endogenous promoters cloned from MP 13.

[0307] Remarkably, overexpressing yigZmv\ in MG1655 enhanced colonization by 114-fold at day 2 and 302-fold at day 3 relative to control (Figure 4D). The strain with the plasmid containing all four genes exhibited 29- and 76-fold enhanced colonization at day 2 and 3, respectively (Figure 4D). Overexpression of tcd.Ampv slightly reduced MG1655 colonization (Figure 4D). A significant reduction in MG 1655 colonization was observed across all mice at day 7 (Figure 12C). The yzgZnipi3 overexpressing strain was outcompeted in vitro (Figures 12D-12E), suggesting that the fitness advantage of yigZmpi3 was specific to the in vivo environment. These results demonstrate that modest gene copy increase (~5-fold) of yzgZmpi3 can significantly enhance early intestinal colonization of E. coli. Altogether, our loss-of-function and gain-of-function experiments revealed that two out of the four translation-related CFs tested are bona fide E. coli colonization factors.

[0308] Gut Adaptation is Associated with Greater Diversity of CMs across Diverse Phyla

[0309] Building on our experimental validation, we set out to investigate the distribution and evolution of CMs among diverse microbial species. CMs and CFs are widespread across major gut microbial taxa (Figure 5A, Figure 13A). TrhP- and YigZ-containing CMs are among the most widespread ones among the gut microbes, with the exception that YigZ is absent in Verrucomicrobia (Figure 5A, highlighted with black boxes on the X-axis). As expected, genomes of gut microbes, both MAGs and isolates, contain more CMs than genomes from environmental sources (Figure 5B). Within each taxonomic class, gut microbes encode a larger set of CMs compared to taxonomically similar counterparts (Figure 5C). This trend persists even within Alpha- and Gamma-proteobacteria, which are more commonly found in environmental habitats (Figure 5C). The species similarity in terms of the presence of CMs and taxonomic distance follows the distance-decay relationship (Figure 13B), resembling a similar pattern of species habitat distribution shown previously58.

[0310] CF Distribution and Conservation across the Tree of Life

[0311] The 9K genomes selected for genotype-habitat association were limited to high-quality genomes with well-defined habitat annotations. To achieve a more comprehensive characterization of CFs across a truly representative tree of life, we expanded our analysis to include the complete set of 113,103 representative genomes from the GTDB database59,60using MMseqs profile search61,62(Methods). We observed that CFs display varying degrees of cross-phylum prevalence and within-phylum conservation. The phyla with the highest average number of CFs include Fusobacteriota, Bacillota, and Spirochaetota. Notably, these CF-rich phyla are distributed across diverse branches of the tree rather than being confined to closely related clades, suggesting that the roles of CFs in species habitat adaptation may have been conserved across long evolutionary time scales. Furthermore, CF gene set completeness is lower than that of GTDB single-copy marker genes across all phyla, suggesting greater species divergence and variability of CFs across bacterial lineages.

[0312] To quantitatively assess the degree to which each CF is associated with phylogeny, we calculated the phylogenetic signal of each CF in the top 5 largest genera across four major classes (Figure 5A). We observed a complex, yet intriguing, pattern that a CF may exhibit drastically varying phylogenetic patterns, with statistically significant phylogenetic signal in one genus, but absent in another. For instance. CF10 (branched-chain amino acid transporter BmQ) shows a strong phylogenetic signal in the genus Flavobacterium within the Bacteroidia class (delta z-score: 287) but weak signals in other genera (delta z-score: -0.15 to 11). Similarly, CF43 exhibits a strong signal in Paraburkhol.de ria from Gammaproteobacteria (delta z-score: 400) but is weakly correlated in others (delta z-score: <2). In contrast, the experimentally validated CF YigZ is prevalent across all studied genera and shows no significant phylogenetic association. These findings suggest that, depending on the taxonomic context, CFs may be acquired through various mechanisms, such as vertical selection within specific lineages or horizontal transfer, and at different points in evolutionary time. In summary, while CFs are broadly distributed across the tree of life, they exhibit distinct patterns in taxonomic range, conservation, and phylogenetic signal.

[0313] Bimodal Distribution of Consistently Present or Absent CMs Across Diverse Species

[0314] To assess the conservation of CMs at the species level, we calculated the frequency of each CM across all genomes for a set of diverse species (Figure 6A). The within- species frequencies of CMs exhibited a striking bimodal distribution (Figure 6 A, Figure 13C), dividing the CMs into two groups- consistently present or absent in most genomes from the same species. Although different species encode different subsets of CMs (Figure 6A), this bimodal distribution of within-species CM frequency persists for every species analyzed (Figure 13C). An expanded analysis with 794 and 3,533 high-quality C. difficile and E. coli genome assemblies further substantiates this pattern (Figure 6B). Such bimodal within-species frequency is also observed at the CF level (Figures 13D-13E). In summary, CFs / CMs consistently display a bimodal distribution, being either predominantly present or largely absent. Consequently, these CFs will not emerge as differentially present when only comparing closely related genomes, underscoring the necessity of performing genotype-habitat association at the tree-of-life scale.

[0315] Natural Allelic Variations in yigZ Lead to Differential Colonization Capacity

[0316] Despite the conserved presence of CMs within species, some CMs exhibit nonsynonymous coding (Figure 3C) and copy number variations across strains. We thus sought to investigate whether these within-species variations can lead to variability in a strain’s gut colonization phenotype, using YigZ as a test case. An in-depth analysis of 2,948 naturally occurring E. coli isolates revealed that YigZ harbors two frequent amino acid substitutions at the 25th and 146th residues (Figure 7A). The two residues, located separately in the ancient N-terminal IMPACT domain63-68or the C-terminal EF-G-like domain (Figure 7B), also differentiate YigZKi2and YigZMPi (M25L and H146S) from the domesticated K12 strain and host-associated MP1 strain, respectively (Figure 3C). To test whether yigZKi and yigZMPi confer differential colonization capacity, we cloned the CDS and promoter region of yigZ separately from both MG1655 (K12) and MP1 strains (Figure 3C, Figure 12F) and overexpressed them, in combinations, in MG 1655 for in vivo competition. We found that PyigZmpi-YigZMPi confers approximately a 4.5-fold higher colonization advantage than PyigZKi2-YigZKi2 at post-gavage day 3 (Figure 7C). Replacing either the promoter or CDS to the MG1655 version reduced the gain-of-function effects (Figure 7C). More broadly, at the population level, YigZMPi-coding strains were more frequently observed among the 1,135 human-associated E. coli isolates relative to those from the environment (Figure 7D). These findings, based on laboratory experiments and ecological observations, highlight that sequence-level evolution of CMs can further drive differential habitat adaptation across closely related strains from the same species.

[0317] Discussion

[0318] We hypothesized that common conserved mechanisms of gut colonization are preferentially enriched in microbes that have adapted to the gut environment relative to those that live in distinct (aquatic and terrestrial) habitats. We, thus, developed a computational framework for microbial genotype-habitat association at the tree-of-life scale and applied it to identify putative genetic factors for mammalian gut colonization. The scale of our study was enabled by the critical efforts to survey diverse microbial habitats29,69'72and methodologies developed for assembling genomes directly from metagenomes73,74. Utilizing both isolate genomes and MAGs, we identified putative colonization factors independent of organisms’ cultivability. Examining millions of proteins across thousands of unique species from diverse habitats, we identified 79 broadly conserved CFs that strongly associate with human and / or mouse GI colonization. These CFs were organized into co-inherited CMs, reflecting putative key biological processes that may causally contribute to mammalian gut colonization across diverse commensal gut bacteria.

[0319] The identified CMs correspond to distinct biological processes, some of which align with previously known mechanisms of gut colonization, including metabolic niche17,21,48, quorum sensing49and sugar phosphotransferase systems50,51. The Quorum sensing molecule AI-2 (CM11) has been shown to mediate murine gut colonization by E. colt15and multi- species biofilm formation76. It is also suggested to be a universal communication molecule among microbes due to its broad prevalence40. Our discovery of these processes as CMs highlights a potentially broader roles for these functions in the colonization of other commensal bacteria. Although our experimental study is only conducted in E. coli, some CFs have been suggested to be essential for colonization in B. thetaiotaomicron in a transposon mutagenesis study17, including CM19 (BT0370, Galactokinase; BT4127: glucosamine-6-phosphate deaminase), CM9 (BT4126, FprA family A-type flavoprotein), and CM10 (BT4699, Cof-type HAD-IIB family hydrolase).

[0320] Although CMs were initially identified in a binary-mode analysis, we observed more nuanced variations, including gene duplications (Figures 2C and 2H) and sequence variations (Figure 3C) both within and across species. In the case study of the gene yigZ, experimental and ecological evidence suggest that these finer- scale variations can impact colonization capacity.

[0321] Our work revealed several translation-related CFs, among which we showed that TrhP and YigZ are essential for E. coli colonization, resonating with recent work that demonstrated gut colonization by Bacteroides requires alternative protein synthesis factors77,78. These data together suggest a broader role for alternative translational control in the colonization of commensal bacteria in murine and human guts. Remarkably, a ~5-fold increase in the copy number of yigZmpi3 led to a 100-300 fold increase in colonization of E. coli MG1655. This finding suggests a potential avenue for using CFs to enhance the fitness of gut commensal strains, particularly for live biotherapeutic products (LBPs), whose colonization success directly impacts clinical efficacy, as reported by a recent clinical trial15.

[0322] YigZ, encoded by 91% of the mouse and human gut microbes analyzed (Figure ID), is a 204 amino-acid protein in E. coli. Previous work has implicated yigZ in early biofilm formation by E. coli BW2511379. YigZ contains two conserved domains. Intriguingly, its N-terminus domain, IMPACT, is conserved across kingdoms, with homologs in yeast (Yihl)65and human (IMPACT)65involved in regulating the critical GCN2-eIF2 axis63-68. YigZ interacts with multiple translation-related proteins41, and knockout of yigZ increases translation readthrough41. YigZ physically interacts with the large ribosomal subunit assembly factor BipA80. Structurally, YigZ C-terminal domain resembles Elongation Factor G (EF-G)81.

[0323] Other studies have explored similar questions using different computational approaches. A previous study used clustering of microbial genes to demonstrate that most species-level genes are habitat-specific35. Another study82developed an elegant phylogenetic-correction based method to associate genes with species gut prevalence within the four major gut phyla (Bacteroidetes, Firmicutes, Proteobacteria, or Actinobacteria). Another study employed a clustering-based method to group proteins into clusters, identifying those associated with plant colonization across nine phyla83. Other studies have used comparative genomics to study colonization, but narrowly focused on the same or closely related species18. We approached this question differently by searching for common conserved genes specifically present in genomes recovered from mammalian gut, relative to other distinct habitats. To do this, we employed the gold-standard BLASTP alignment to reliably identify remote protein homologs across diverse species, and rank each protein’s association with habitat quantitatively against a frequency-matched null distribution. Our pipeline, thus, enables genotype-habitat association at the tree- of-life scale. Critically, we bypass the computationally intensive task of first clustering proteins into orthogroup at a large scale - a typical bottleneck that could affect the precision and statistical power of associ •ati-on 84.

[0324] Microbial phylogeny is an active area of research; factors such as tree construction methods and marker gene selection introduce uncertainty to a true tree-of-life phylogenetic scale. Additionally, recent studies suggest that gut microbial phylogeny does not always mirror their phenotypes85. Therefore, we chose not to include phylogenetic correction in the initial steps of protein-habitat association. Indeed, the expanded analysis reveals that CFs and CMs are widely distributed across diverse gut phyla and classes, and show varying phylogenetic signal within distinct phylogenetic contexts. The high within- species conservation of colonization factors underscores the power of a cross-genomic approach covering broad phylogeny, as their discovery would have been challenging if performed at the intraspecies level alone86.

[0325] In conclusion, our cross-genomic genotype-habitat association approach has uncovered both known and previously undescribed mechanisms for microbial gut colonization. Our successful experimental validation of 2 of 4 translation-related factors in E. coli and the large impact of YigZ suggest the broader validity of our findings. As we show for YigZ, expression of CFs can strongly modulate colonization of individual species, providing a powerful strategy for microbiome manipulation for both scientific investigation and therapeutic intervention.

[0326] Data and code availability

[0327] The computational tool is available at tavazoielab.c2b2.columbia.edu / GHA. The processed phylogenetic profile data and downstream analysis code can be downloaded from Zenodo (doi.org / 10.5281 / zenodo.12516934). The sequencing results can be accessed through SRA under PRJNA1126302.

[0328] METHODS

[0329] EXPERIMENTAL MODEL AND STUDY PARTICIPANT DETAILS

[0330] Mouse model for intestinal colonization

[0331] Six to eight-week-old male C57BL / 6 mice were purchased from Jackson Laboratory and kept in (specific pathogen-free) SPF mouse facility. Before gavage, they were fed streptomycin and glucose in their drinking water, both at a concentration of 5 mg / ml, for 72 h as previously described55. After a ~24h recovery with no antibiotic supplementation, the mice were orally gavaged with bacteria inoculum. Mice were single-housed to prevent coprophagia so that each mouse was a strictly independent biological replicate (IACUC: #21054-H). METHOD DETAILS

[0332] Genomes data curation

[0333] Data download. The metagenome-assembled genomes (MAGs) were collected from GEM29and MGBC70. The isolate genomes were collected from JGI IMG71,72. Only genomes designated with 'high-quality' labels (>90% complete with <5% contamination)100were included for downstream selection. For isolates, to honor JGI data usage policy, we only included genomes that are either (i) publicly released ('Is Published'="Yes") or (ii) JGI sequenced genomes ('Sequencing Center'="JGI") released after 2018-11-1 with release days more than 365 days at the time of analysis ('release.after.2018'="Yes") and ('release.days'>365). The 9,475 genomes represent a substantial portion of GTDB taxonomy (based on v207), covering 52.7% (78 out of 148) of bacterial phyla and 46.2% (197 out of 425) of bacterial classes, as well as 40% (8 out of 20) of archaeal phyla and 43.4% (23 out of 53) of archaeal classes. Statistics are generated in alignment with GTDB conventions; for instance, / ? Firmicutes_A, p Firmicutes_B, and p Firmicutes_C are all counted under p Firmicutes.

[0334] Reducing taxonomic redundancy. To preserve taxonomic diversity with minimal genomic redundancy, a cutoff was set on the maximum number of genomes per species to be included. The Genomic Taxonomy Database (GTDB)59 97 101taxonomy associated with each genome was acquired from the original data source29,70-72. Genomes were sorted by total size, and the top N from each species were selected for each habitat to include as many accessory genes as possible. Positive genomes (human and mouse gut microbes) and negative genomes (environmental microbes) were selected separately, with N set to 10 or 5, respectively. Such separate selection across datasets aims to: 1) balance the total number of genomes for the positive and negative classes to achieve statistical power, and 2) ensure inclusion of genomes from the same species with distinct habitats labels.

[0335] Protein alignment

[0336] The proteomes associated with MAGs and isolates were downloaded from the original sources and subjected to all-against-all alignments. Given the significant computational demands of this step, we opted to use proteins from phenotype-positive genomes (gut microbes) as the query proteins to align against the protein database constructed with the complete set of proteins from positive and negative genomes. This decision assumes that proteins important for gut colonization are predominantly enriched in gut microbes. By doing so, we convert the all-against-all alignment into a semi all-against-all approach, effectively reducing the task's magnitude by half. The pipeline allows users to choose between DIAMOND and MMseqs261.

[0337] For MMseqs2, the parameters below are recommended

[0338] 'Speed setting: “-s 4,

[0339] Coverage: “-c 0.5”,

[0340] Exhaustive search: “-max-seqs 10000”,

[0341] Evalue cutoff: “-e IE- 10”'

[0342] In the current study, alignment was performed using the protein aligner DIAMOND (version 2.O.9)102 103tp unless otherwise noted, with the pipeline parameters:

[0343] ' ALIGNMENT M AX= 10000',

[0344] ' ALIGNMENT_E VALUE- 1 e- 1 O' ,

[0345] ' ALIGNMENT_QUERY_COVERAGE=66' ,

[0346] 'ALIGNMENT_SUBJECT_COVERAGE=50', which specifies four important aspects: increase the confidence of a hit being a functionally homologous protein to the query protein, we force the aligned region to cover the majority of both the subject and query proteins by setting DIAMOND '-query-cover=66' and ' -subject-cover=50' . This selection aims to reduce the chance of acquiring a protein hit that only partially matches the query protein, for example by just one domain of a multi-domain protein. We also show that most of the CFs are robust to changing this parameter as they are consistently identified when '-query-cover was set to alternative values

[0347] (Figures 9D-9E). value: The maximum E ■value to report a protein hit is set to le- 10 with '-evalue=le-10' . This cutoff has been shown to be successful in previous studies with 202 genomes for various bacterial phenotypes such as chemotaxis and sporulation38 104.

[0348] Maximum targets The default DIAMOND only reports 25 targets per query. To ensure we find all homologs for each query protein to generate an accurate phylogenetic profile, we set the maximum alignment hits to 10,000 by '-max-target-seqs -10000' , where 10,000 is greater than the number of genomes we have for the three analyses. We then retrospectively confirmed the actual number of hits per query was smaller than 10,000 to ensure this parameter indeed enables exhaustive search through all the proteomes for every protein. DIAMOND mode: The initial alignment was performed under DIAMOND '-very-fast' mode, based on which we generated phylogenetic profiles for each protein used for protein-habitat association. For top habitat-associated proteins identified, their homologs are searched again under DIAMOND '-very- sensitive' mode (See the “Reassessing the association of CF with three datasets” section below). To ensure the highest detection sensitivity of CF homologs in representative species (A. coli, B. theta, C. difficile'), ' -ultrasensitive' was used.

[0349] Phylogenetic profiles (PP)

[0350] Protein-level PP. The protein-level PP is a binary vector of 0 or 1 to reflect the absence or presence, respectively, of DIAMOND hits across genomes for a particular protein. This is done using the diamond_to_pp function within the association.py script as part of the pipeline.

[0351] CF-level PP. The protein-level PP is summarized to the CF-level PP under three distinct modes, which are used in different scenarios. (1) Permissive mode. A CF is considered present in a genome if any member protein has a hit for that genome. Permissive PPs were used for protein-habitat association and CF-CF coinheritance analysis (module discovery). (2) Selective mode. Recognizing that a protein hit can be simultaneously assigned to multiple homologous CFs, we, thus assign a protein uniquely to the most related CF, whose member protein shows largest alignment bitscore. If two CFs have equally highest bitscores, the protein is assigned to the larger CF (with more member proteins). Selective mode is used when investigating the distribution of CFs and CMs (Figures 5A-5C. 6A-6B, 13C, 13D) and for calculating CM genomic distance (Figure 13B). For example, the PPs of CF2_4 and 2_7 are similar under the permissive mode, but distinct under the selective mode (Figure 9J). (3) Copy number mode. For a particular genome, the total number of proteins assigned to a CF under the selective mode was summarized and used as the copy number of CFs in that genome.

[0352] Protein-habitat association

[0353] Signed Mutual Information (MI). Given the habitat vector (H) and a protein phylogenetic profile (PP) H G {0 (non-gut), l(gut)}wPP G {0 (absent), l(present)]Nwhere N is the total number of genomes. The mutual information M / 38and Spearman's rank correlation coefficient (p) between PP and H are calculated to determine the signed MI:

[0354] Signed MI = sign(p(H, PP)) • MI(H; PP) MI z-score. To convert a protein’s signed MI into MI z-score, we generate a frequency-matched null distribution. Specifically, for a given protein with frequency M, its NULL distribution NULLMis generated by shuffling a PP vector 1000 times:

[0355] NULLM= (s1,s2, ... , sloOo) where each element sMi= sign(p(H, PP*)) • MI(H; PP*) and PP* = shuffle(PPM) and J i PPi =

[0356] M. From this null distribution, the mean ( ) and standard deviation (cr) are calculated and signed MI is converted into MI z-score:

[0357] „ Siqned MI - u

[0358] MI z-score = — - -

[0359] In practice, instead of generating a new null distribution for every protein, we pre-generated N — 1 null vectors that correspond to all possible protein frequencies: M e {1,2, ... , N - 1}. When calculating MI z-score for each protein, the o and / / of the corresponding null vector were retrieved and used.

[0360] Procedures for enhancing pipeline speed

[0361] Genome preprocessing. We assigned new designations to input proteins using ' genome\protein' . facilitating efficient genome index retrieval to convert protein alignment results to PP vectors for millions of proteins.

[0362] Combining genomes for alignments. Since DIAMOND is optimized for large input files of more than one million proteins103, we amalgamated every 200 proteomes into one query dataset to utilize the full potential of DIAMOND. This strategic decision played a crucial role, allowing us to complete a single run of analyses within a feasible timeframe for pipeline optimization.

[0363] Colonization factor (CF) discovery

[0364] Grouping top colonization-associated proteins into CFs. The top-colonization associated protein were grouped into protein families (same as CFs) using the Markov Cluster Algorithm (MCL)95based on their sequence homology. First, the pairwise alignment e-value for the top-associated proteins aligned against each other were extracted and transformed with mcxload under the following transformation parameters: ' -stream-neg-loglO' and ' -stream-tf ceil(200)' , whose output is a .abc file subject for MCL clustering. Then, MCL was performed with inflation (I) values = 1.4, 2.5 or 3, yielding 71, 79, and 82 protein families. Level 1 (1=1.4) and Level 3 (1=3) MCL results were used for CF assignment, resulting in 82 candidate CFs. CFs with names that include an underscore (X_Yi, X_Y ) indicate that the level 1 protein family (X) is further divided into subfamilies (Yi and Y2) at level 3, therefore automatically reflecting homologous relationships. The CF classification is quite consistent across different iterations even though MCL is based on the simulation of stochastic flows in graphs. The significance of these 82 candidate CFs was reassessed as described below.

[0365] Ml z-score based on more accurate ali The initial protein-habitat associations were conducted using protein PP generated under DIAMOND '-fast' mode, which offers desired computational speed for the semi-all-against-all alignment step (Figure 1A step 2) but may compromise accuracy102 103. To evaluate the association of the 81 putative CFs with the gut colonization phenotype more accurately, we performed protein alignment again under the DIAMOND '-very -sensitive mode for the 27,096 top- associated proteins and recalculated their Ml z-scores across three datasets (human GI MAGs, mouse GI MAGs, human isolates). The protein-level MI z-scores were subsequently summarized to CF-level MI z-scores by taking the maximum MI z-score among corresponding member proteins (Figure 9B), referred to as the refined MI z-scores.

[0366] Reassessing the association of CFs with three datasets. Based on the refined Ml z-scores, we reassessed the significant association of 82 candidate CFs with the gut colonization phenotype across three datasets: human GI MAGs, mouse GI MAGs, and human isolate. A CF is considered significant for a dataset if its refined MI z-scores (Figure 9B) exceed the 95th percentile of the original MI z- scores for all proteins from that dataset (human GI MAGs: >382.22, mouse GI MAGs: >579.93. human isolates: >44.20). This reassessment revealed that 79 out of the 82 CFs are significantly associated with at least one dataset and are included for downstream analysis (Figure 1C). Notably, the original MI z-scores distribution is skewed towards positive values (Figure IB), especially for MAGs, making our significance threshold more stringent. Nevertheless, we decided to focus on only the initial set of highly significant CFs in this study.

[0367] Colonization module ( CM) discovery

[0368] Coinheritance To quantify the degree of CF coinheritance, we calculate the frequency of consensus clustering between any CF pairs among the 79 CFs using the customized function

[0369] ‘ get_coclusterjreq’ as previously described38. Specifically, 1,000 K- means clustering analyses were performed using the permissive PP of the 79 CFs, from which the frequency of two CFs being in the same cluster was recorded and referred to as their coinheritance strengths. Various K values were used to capture clusters at different similarity levels. The current CM result is generated using the parameters as follows: 'k_min = 0.1', 'k_max = 0.8', 'k_step = 0.1', 'k_times = 1,000'. reflecting K values are assigned from 0.1 to 0.8 times the total number of CFs and a total of 1,000 independent K- means runs. CM identification To group co-inherited CFs into CMs, the inverse of CF pairwise coinheritance (1 - coinheritance) was calculated and stored in a 79 x 79 square matrix, which is used as the distance matrix for hierarchical clustering (hclust in R, default parameters). The hclust output tree was cut with various heights ranging from 0.4 to 0.9, among which 0.4 was used for final CM identification as it successfully identified groups of operonic CFs into CMs for multiple operons.

[0370] Functional annotation of CFs and CMs

[0371] CF functional annotation To functionally annotate the CFs. we exploited two reference databases of protein functions: (1) BioCyc tier 1 database (Table I)103was used as the primary resource; (2)

[0372] UniRef50 was used as supporting evidence, particularly to annotate the proteins with no hits in Biocyc. Furthermore, three well-studied species: Escherichia coli K12 MG1655 (from BioCyc), Bacteroides thetaiotaomicron VPI5482 (GCA_014131755.1), and Clostridium difficile S-0253

[0373] (GCF_018885085.1) were used as representative reference species. To convert protein-level annotation to CF-level, for the 27,096 top-associated proteins, the most similar reference protein from each source was identified using DIAMOND under " -very-sensitive" (database-based annotation) or '- -ultrasensitive" (representative species-based annotation) mode. The best hits of all member proteins from a given CF are summarized, resulting in CF-level reference proteins associated with a percentage of member proteins. If a CF has no hit or the most similar protein name is blank, it is classified as unannotated.

[0374] Gene (GO). The Gene Ontology (GO) terms associated with CFs were acquired by collecting GO terms annotations of corresponding UniRef50 protein hits via the UniprotR106package in R. The GO terms are consolidated using the ReViGO107webserver ('Medium (0.7)'). Then, for each

[0375] CF, the consensus GO terms associated with >1% member proteins are retained.

[0376] Calculating inter- and intra-taxa distance

[0377] To investigate the relationship between the CM presence and taxonomy distance among the gut microbes, we calculate the CM dissimilarity between any genome pair among the mammalian gut microbes (n=5148) via Euclidean distance. A CM is deemed present if any member CF is present under the selective mode. The inter- and intra-taxa distances were collected using customized functions (' get_intra_dist" and ' get_inter_dist') at each taxonomic level. When calculating intra-taxa distance for a taxonomic level, two filtering steps were performed: 1) genomes with no or unclassified taxonomic labels at that particular taxonomic level were excluded; 2) taxa with only one genome were excluded from intra-taxa distance calculation.

[0378] Deduplication of NCBI pdp E. coll genomes

[0379] E. coli .faa files were analyzed using Mash108with all parameters set to default, except for the '-a' option, which specifies the use of the amino acid alphabet. Genomes with 100% identity, collected by the same collector, at the same time, and from the same location, were considered duplicate genomes. A total of 259 genomes were assigned to 95 duplicate clusters, with each cluster represented by a single genome for downstream analysis.

[0380] Searching CFs against the GTDB database

[0381] The latest GTDB (v220) database59'60,97was downloaded via MMseqs261by 'mmseqs databases GTDB gtdb'. The GTDB amino acid database was generated using 'mmseqs tar2db' with the following parameters '--tar- include faa.gz$'. To search CFs against the GTDB database, the 27K top-colonization associated proteins were first clustered using 'MMseqs2 cluster' with the parameter below '-e 1.000E- 10', and then the centroid of each cluster was used to query the GTDB amino acid database with parameters '-e IE-10, -c 0.5, -exhaustive-search -s 4.0 -num-iterations 2'. The 120 single-copy marker genel of the v220 GTDB representative genomes are downloaded from the GTDB website.

[0382] Quantifying phylogenetic signal associated with each CF

[0383] We selected four major classes Clostridia, Bacilli, Bacteroidia, and Gammaproteobacteria) from three distinct phyla, as shown in Figure 5A, to quantify the degree to which each CF is associated with microbial phylogeny. For each genus, the phylogenetic tree of all GTDB representative genomes and their CF profiles (based on MMseqs2 results) were extracted, and their associations were quantified using two metrics: delta98and lambda99. Specifically, delta was calculated using the delta function98with the following parameters: 'sim = 100, thin = 10, burn = 10, lambdaO = 0.1 , se = -0.5'. A null distribution of delta from 100 random simulations, frequency-matched with 10 frequency windows per genus, was generated to compute -values and delta z-scores. Lambda, along with its corresponding -values were obtained using the 'phylosig' function from the Phytools109package. Since lambda is designed for continuous traits, it is considered a secondary metric in the current study. If a CF is present in all representative genomes, the calculation for that genus is skipped, and NA values are assigned.

[0384] Culture conditions E. coli strains were routinely grown at 37°C with shaking. 1 mL of culture was grown in 14 mL round-bottom culture tubes shaking at 300 rpm. E. coli was grown in LB Miller Broth (DF0446-07-5, Fisher Scientific) unless otherwise noted. LB Miller plates were used for growth on solid medium. For plasmid maintenance, plates or broth were supplemented with 25 g / mL chloramphenicol, 50pg / mL kanamycin, 50pg / mL carbenicillin or 15pg / mL tetracycline. Cells were routinely pelleted by centrifugation at 5,000xg for 5-10 minutes.

[0385] Bacterial strains

[0386] MP 13 deletion strains. Deletions were generated by lambda-red mediated homologous recombination94.

[0387] The recombination templates were generated through PCR from pKD4 with the following primers: M3 and M4 (ybaK), M5 and M6 (yigZ), Mi l and M12 (irhP). M13 and M14 tcdA). The chromosomal kanR was removed with plasmid pCP2094, which was subsequently lost after non-selective overnight growth at 42°C. All strains were confirmed to have lost pCP20 and pKD46 with no remaining antibiotic resistance.

[0388] Barcoded MG 1655 strains Primers M42 (containing a 20bp stretch of Ns as barcode) and M43 were used to amplify the PCR product from pKD4 for lambda-red mediated recombination (see section Genomic barcode system for details). The six barcoded bcMGl-6 strains were transformed with GOF plasmids pML10-15, respectively. All strains were confirmed to have lost pKD46 with no remaining ampicillin resistance.

[0389] Plasmid cloning

[0390] NEB HiFi assembly (E2621L, NEB) was used to generate the GOF plasmids. The GOF plasmid backbone was amplified from pBbS6C-RFP (Addgene, 35293) with primers Ml 16 / MI 17. The other insert was generated with PCR performed using previously generated plasmids (pMLl - pML4) as templates, which contain GOF gene cassettes amplified from MP13 genomic DNA. The inserts for pMLlO, pMLl l, pML12 and pML13 were PCR products generated from plasmid pMLl with primers M 120 / M 121 (ybaKp-ybt / k'), plasmid pML2 and primers Ml 22 / M 123 (yigZp-yz.gZ), plasmid pML3 and primers M124 / M125 (trhPp-frW5), plasmid pML4 and primers M126 / M127 (tcdAp- 7A), respectively.

[0391] The GOF plasmid pML14 with four genes combined was constructed using HiFi assembly (E2621L, NEB) of 4 PCR products amplified using pMLlO with primers M134 / M135, pMLl l with primers M136 / M137, pML12 with primers M140 / M141, and pML13 with primers M138 / M139. The empty control plasmid pML15 was constructed using linear T4 ligase (NEB, M0202L) of PCR product generated from plasmid pBbS6C-RFP (Addgene, 35293) and primers Ml 16 / MI 17. The final GOF plasmids were confirmed using either Sanger sequencing of the insert region or long-read sequencing of the whole plasmid for pML14 (Genewiz, Plasmid-EZ).

[0392] Mouse model for intestinal colonization. Six to eight-week-old male C57BL / 6 mice were purchased from Jackson Laboratory and kept in SPF mouse facility. Before gavage, they were fed streptomycin and glucose in their drinking water, both at a concentration of 5 mg / ml, for 72 h as previously described35. After a ~24h recovery with no antibiotic supplementation, the mice were orally gavaged with bacteria inoculum. Mice were single-housed to prevent coprophagia so that each mouse was a strictly independent biological replicate.

[0393] Bacterial inoculum

[0394] Escherichia coli cells in the exponential (log) phase, with an optical density (ODeoo) of 0.4 to 0.6, were prepared for use as an inoculum. The E. coli strains were initially revived from frozen stocks and cultured overnight on LB agar plates containing the appropriate antibiotics. On the day prior to the experiment (Day -1), a single bacterial colony was selected and cultured in LB broth with the appropriate antibiotics for an overnight incubation.

[0395] Inoculum for loss-of-function experiment. On the day of the experiment (Day 0), the overnight cultures of MP7 and MP13 strains were diluted at a 1:200 ratio and grown in LB Lennox (Sigma-Aldrich, L7275- 500TAB) for 2-3 hours until the ODeoo reached ~0.4. Subsequently, the bacterial cells were pelleted by centrifugation at 5,000 rpm for 10 minutes and resuspended in cold phosphate-buffered saline (PBS). Each mouse was administered 109bacterial cells suspended in 100 pL of PBS, encompassing MP7 and MP 13 (wildtype or mutants) equally.

[0396] Inoculum for gain-of-function experiment. On the day of the experiment (Day 0), the overnight culture MG strains were diluted at a 1:200 ratio and grown in LB with kanamycin and chloramphenicol for 2-3 hours until ODeoo reached -0.4. The actual ODeoo ranged from 0.41-0.49 across the 6 strains. Subsequently, the bacterial cells were pelleted by centrifugation at 5000 rpm for 10 minutes and resuspended in cold PBS. Each mouse was administered a total of 6 x 109cells suspended in 100 pL of PBS. encompassing 6 barcoded strains harboring different GOF plasmids evenly.

[0397] In vivo fitness of E. coli MP strain

[0398] Fecal sample collection and plating. Fecal pellets were collected from each mouse, weighed, and then homogenized in PBS using pre-filled soft tissue homogenizing beads (1.4 mm Ceramic Beads. Omni). These homogenized samples were serially diluted and plated on LB agar plates containing tetracycline (15 pg / mL), which activates the fluorescence cassette. For each mouse, a 100 pL fecal-PBS slurry was plated on LB plates with tetracycline (to induce fluorescence) in triplicates for different dilutions.

[0399] After overnight incubation, plates displaying a countable range of colony-forming units (CFUs) were imaged for colony counting. The fecal-PBS slurry was kept in the 4°C refrigerator to allow for the possibility of plating a higher dilution the next day if the initial dilution did not yield sufficient colonies for counting. After this, the fecal-PBS slurry was moved to -80°C for long-term storage. and colony counting. Plates were imaged using a Bio-Rad ChemiDoc system with multi-view using two channels Alexa Fluor488(GFP) and Alexa Fluor546(mCherry). These images were subsequently processed for the quantification of green (MP13) and red (MP7) fluorescent colonies using a customized CellProfiler pipeline. Briefly, for each plate, the Alexa Fluor488and Alexa Fluor546images were rescaled, and putative colonies were identified using the IdentifyPrimaryObjects function using the ' O . u' thresholding method and 'adaptive' thresholding strategy (see pipeline for other parameters). The candidate colony objects were filtered using shape ('pixel units' = 1-25) and size

[0400] ('Eccentricity' < 0.8). For each colony, the GFP and mCherry intensity was quantified from the Alexa Fluor488and Alexa Fluor546images. Plating colonies’ GFP and mCherry on a scatterplot, we identified the cutoff line dividing two major populations (Y = 1.3X - 0.1). Therefore, if a colony has

[0401] GFP > 1.3 mCherry — 0.1, it is classified as GFP; otherwise, it is classified as mCherry. The colonies that are classified as GFP or mCherry, as well as colonies removed during the filtering step, were summarized and combined into three separate masked images for manual checking to determine if corrections need to be made or if an image needs to be manually counted.

[0402] Calculating MP 13 over MP7 ratio. The ratio of MP 13 wildtype or mutants to MP7 was calculated for each mouse with 2-3 replicate plates. The ratio was then normalized to the corresponding ratio at the inoculum:

[0403] Ratio normalized as described in55. For visualization in figure 3E, the samples with zero colonies for GFP (MP13) or RFP (MP7) were represented with 10‘4(MP13 not detected) or 103(MP13 took over), respectively, to avoid negative infinity during log 10 transformation or errors due to division by zero. Additionally, one mouse (2-1) that received ybaK mutants displayed the characteristic small colony morphology of ybaK, and eventually, AybaK became dominant (see Figure 1 IE). This sample was excluded from the data analysis in Figure 3E.

[0404] Genomic barcode system

[0405] To increase the throughput of in vivo colonization profiling, we developed a barcoding system that allows us to determine the frequency of multiple strains via next-generation sequencing (NGS). We chose a genomic location for barcode insertion based on two criteria: 1) an intergenic region with no functional elements such as promoters or terminators, to avoid impacting gene expression; 2) conservation among multiple E. coli strains, allowing for efficient use in inter-strain competition in the future. We conducted a systematic analysis of the genomes of three E. coli strains (MG1655, Nissle, MP1) and identified one locus located 29 bp downstream from the small regulatory RNA gene RprA and 90 bp upstream from the DUF1870 domain-containing protein YdiL. This locus was manually examined to exclude any known functional motifs, such as DNA-binding motifs or terminators. A PCR product containing a 20 bp random barcode and kanR was generated using the pKD4 plasmid and primers M42 and M43 and was chromosomally integrated into E. coli using lambda-red-mediated homologous recombination. Colonies with successful barcode integration were verified by PCR using primers M28 and W1043. The barcode sequence was determined through Sanger sequencing. The procedure for recovering the barcode from mouse fecal samples is described in the next section.

[0406] Evaluation of in vivo fitness of MG1655 strain

[0407] Absolute colonization levels. Fecal pellets from mice colonized with the MG1655 strain were collected and processed in the same way as in the previous mouse experiment with MP strains. Serial dilution of fecal-PBS slurry was dot-plated (5 pL) on LB agar plates with kanamycin (for the genomic barcode) and chloramphenicol (for GOF plasmid maintenance) for overnight incubation to determine the absolute abundance of E. coli by counting CFUs. The same plating was done on LB plates containing only kanamycin to screen for plasmid loss. No significant plasmid loss was observed across all samples, as they exhibit colonies at the same dilution. The only exception was one mouse (B4) at one timepoint (day 2), but it did not show significant plasmid loss at the latter timepoint, thus was still included in the analysis.

[0408] Fecal metagenomic DNA extraction. The fecal slurry suspended in PBS was subsequently stored at - 80°C until it was ready for sample preparation for barcode sequencing (barcode-seq). Only samples that exhibited colonies based on plating data were included for DNA extraction and barcode-seq. For DNA extraction, 100 pL of the fecal slurry was processed using the Qiagen PowerSoil kit (Qiagen, 47014). The extracted DNA was then eluted in either 30 or 50 pL of molecular-grade water. Barcode- . To determine the frequency of barcoded MG 1655 strains in the mouse fecal metagenome

[0409] (Figure 4D and Figure 7C), 1 L of the metagenomic DNA was used to amplify the 20 nt barcode for sequencing (barcode-seq). The initial barcode-seq library preparation involved two rounds of PCR: first, amplification using custom primers M132 / M133, followed by commercial NEB i5 / i7 primers from the NEBNext kit (NEB, E7600S), yielding a final PCR product of 204 bp. For a subset of samples, we observed non-specific amplification of a shorter product at high frequency. Therefore, we revised the barcode-seq protocol to include three rounds of PCRs, using primers M154 and M155 (first round), M156 and M157 (second round), and NEB i5 and i7 (third round), yielding a final product of 509 bp. The 3-step protocol efficiently depleted non-specific product formation and was adopted as the default.

[0410] PCRs were performed using a Thermal Cycler (ProFlex, Applied Biosystems™) with Q5 (NEB, #M0491). For the 2-step barcode-seq protocol, the first PCR was performed with the following settings: initial denaturation: 98°C 30 s; 8 or 16 cycles (8 or 16 were used for pure E. coli DNA from inoculum or fecal metagenome, respectively): 98°C 10 s, 58°C 20 s, 72°C 30 s; 72°C 2 min. The PCR product was cleaned with the Zymo-Clean-5 kit (Zymo, D4014) and eluted in 30 pL of molecular- grade water. A total of 8 p L was used as the template for the second PCR in a 25 pL reaction. Before performing the second PCR, 5 pL from the mixed reaction was taken to run qPCR (QuantStudio5, Applied Biosystems) to determine the number of cycles for PCR to ensure the amplification is within the exponential range, while the rest of the reaction remained on ice or in a 4°C refrigerator. PCR2 was performed as follows: 98°C 30s; 16-19 cycles of (depending on qPCR results): 98°C 10 s, 58°C 20 s, 72°C 45 s; 72°C 2 min. PCR product was cleaned using AMPure XP beads (Beckman Coulter, 1.4x) or Axyprep (Axygen™, 1.4x).

[0411] For the 3-step barcode-seq protocol, 1.5 pL of metagenome DNA was used as template for a 35 pL reaction in the first PCR that ran with the following settings: initial denaturation: 98°C 30 s; 8 cycles: 98°C 10 s, 65°C 30 s, 72°C 30 s; 72°C 2 min. The PCR product was purified with Axyprep beads (1.2X, Axygen™) and eluted in 20 pL, of which 2.5 pL was used as template for PCR2. PCR2 was performed in 25 pL reaction with the same settings as PCR1. PCR2 products were cleaned with Axyprep beads (1.2X) and eluted in 20 pL, of which 10 pL was used as template for PCR3. PCR3 was run in a 25 pL reaction. Before performing PCR3, 5 pL from the reaction was taken to run qPCR to determine the number of cycles for each sample to ensure the amplification is within the exponential range, while the rest of the reaction remained on 4°C. PCR3 was performed with the following settings: initial denaturation: 98°C for 30 s; 5-19 cycles (depending on qPCR results): 98°C for 10 s, 65°C for 45 s, 72°C for 30 s; 72°C 2 min. The PCR was cleaned using Axyprep beads (0.8X).

[0412] The library concentration was measured using the Agilent Bioanalyzer High Sensitivity DNA Kit (5067- 4626, Agilent). Libraries were sequenced for 75 cycles using the NextSeq 500 / 550 High

[0413] Output Kit v2.5 (20024906, Illumina) or 150 cycles using commercial vendor (Genewiz).

[0414] Barcode frequency determination. For samples processed with the 2-step PCR protocol, R1 fastq file was used; otherwise, R2 was used. To determine the occurrences of the six barcodes in each sample, we supplied a fasta file containing a whitelist of six barcodes and query each fastq file using seqkit110by the command 'seqkit locate -degenerate -pattern-file barcodelist.fa' . The output was analyzed using 'cut -f2\sort\uniq -c to determine the count of sequences containing each barcode (count;). The minimum barcode depth (total count of barcodes) was 50K per sample in the current study. The frequency of each barcode was calculated by freqL= — -!— where n = 6 (Figure 4) or 9 (Figure

[0415] 7). If a strain barcode is not detected, a pseudo-count of 1 is used and its frequency is defined as 1

[0416] - , where sequencing depth refers to the total number of reads in that sample. Sequencing depth

[0417] The fitness of each MG strains is determined by the frequency of corresponding barcode normalized against the control strain (harbors the empty plasmid) based on their frequencies in the

[0418] , .. .rfreq< ,rbaseline freq < baseline inoculum: J freq nti, n noorrmmaalliizzeedd — - f actor - f, where ’ factor1= - baseline freqcon-t—ro l.

[0419] Table 1. Colonization factor (CF) information (related to Figures 1 and 2). Information includes CF z-scores across datasets, CF-dataset associations, CF-CM assignments, and functional annotations based on the BioCyc database.

[0420] CF Dataset significance: significance across three datasets (human GI MAGs, mouse GI MAGs and human isolates). Multi-CF CM: CM assignment. CM type: CM type based on member CF homologous relationship and genomic context (Fig. 2C). Biocyc annotation: Most similar protein from the on Biocyc database. Biocyc annotation percentage: % of member proteins that mapped to the annotation proten. MI z-score: CF-level refined MI z-scores based on DIAMOND '-very-sensitive'. CF size: number of proteins belonging to this CF among the top 27K colonization-associated proteins. Homologous CF: Homologous CF based on Fig. ID. Table 2. Key Resources

[0421]

[0422] References

[0423] 1. Devlin et al. (2016). Modulation of a Circulating Uremic Solute via Rational Genetic Manipulation of the Gut Microbiota. Cell Host Microbe 20, 709-715. 2. Guo et al. (2019). Depletion of microbiome-derived molecules in the host using Clostridium genetics. Science 366, eaavl282. 10.1126 / science.aavl282.

[0424] 3. Sun et al. (2020). Bifidobacterium alters the gut microbiota and modulates the functional metabolism of T regulatory cells in the context of immune checkpoint blockade. Proceedings of the National Academy of Sciences 117, 27509-27515.

[0425] 4. Chen et al. (2023). Engineered skin bacteria induce antitumor T cell responses against melanoma. Science 380, 203-210. 10.1126 / science.abp9563.

[0426] 5. Din et al. (2016). Synchronized cycles of bacterial lysis for in vivo delivery. Nature 536, 81-85. 10.1038 / naturel8930.

[0427] 6. Gurbatri, C.R., Arpaia, N., and Danino, T. (2022). Engineering bacteria as interactive cancer therapies. Science 378, 858-864. 10.1126 / science.add9667.

[0428] 7. Savage et al. (2023). Chemokines expressed by engineered bacteria recruit and orchestrate antitumor immunity. Science advances 9, eadc9436. 10.1126 / sciadv.adc9436.

[0429] 8. Hahn et al. (2023). Bacterial therapies at the interface of synthetic biology and nanomedicine. Nature Reviews Bioengineering. 10.1038 / s44222-023-00119-4.

[0430] 9. Vincent et al. (2023). Probiotic-guided CAR-T cells for solid tumor targeting. Science 382, 211- 218.

[0431] 10. Lawson et al. (2019). Common principles and best practices for engineering microbiomes. Nature Reviews Microbiology 17, 725-741.

[0432] 11. Brandl et al. (2008). Vancomycin-resistant enterococci exploit antibiotic-induced innate immune deficits. Nature 455, 804-807.

[0433] 12. Kim et al. (2019). Microbiota-derived lantibiotic restores resistance against vancomycin- resistant Enterococcus. Nature 572, 665-669.

[0434] 13. Shan, Y., Lee, M., and Chang, E.B. (2022). The Gut Microbiome and Inflammatory Bowel Diseases. Annu Rev Med 73, 455-468. 10.1146 / annurev-med-042320-021020.

[0435] 14. Caruso, R., Lo, B.C., and Nunez, G. (2020). Host-microbiota interactions in inflammatory bowel disease. Nature Reviews Immunology 20, 411-426. 10.1038 / s41577-019-0268-7.

[0436] 15. Louie et al. (2023). VE303, a Defined Bacterial Consortium, for Prevention of Recurrent Clostridioides difficile Infection: A Randomized Clinical Trial. JAMA 329, 1356-1366.

[0437] 16. Dsouza et al. (2022). Colonization of the live biotherapeutic product VE303 and modulation of the microbiota and metabolites in healthy volunteers. Cell Host Microbe 30, 583-598, e588.

[0438] 17. Goodman et al. (2009). Identifying genetic determinants needed to establish a human gut symbiont in its habitat. Cell Host Microbe 6, 279-289.

[0439] 18. Wu et al. (2015). Genetic determinants of in vivo fitness and diet responsiveness in multiple human gut Bacteroides. Science 350, aac5992. 10.1126 / science.aac5992.

[0440] 19. Liu et al. (2008). Regulation of surface architecture by symbiotic bacteria mediates host colonization. Proc Natl Acad Sci USA 105, 3951-3956.

[0441] 20. Donaldson et al. (2020). Spatially distinct physiology of Bacteroides fragilis within the proximal colon of gnotobiotic mice. Nat Microbiol 5, 746-756.

[0442] 21. Shepherd et al. (2018). An exclusive metabolic niche enables strain engraftment in the gut microbiota. Nature 557, 434-438. 10.1038 / s41586-018-0092-4.

[0443] 22. Lee et al. (2013). Bacterial colonization factors control specificity and stability of the gut microbiota. Nature 501, 426-429. 10.1038 / nature 12447.

[0444] 23. Jung et al. (2019). Genome-Wide Screening for Enteric Colonization Factors in Carbapenem- Resistant ST258 Klebsiella pneumoniae. mBio 10. 10.1128 / mBio.02663-18.

[0445] 24. Fu et al. (2013). Tn-Seq analysis of Vibrio cholerae intestinal colonization reveals a role for T6SS-mediated antibacterial activity in the host. Cell Host Microbe 14, 652-663.

[0446] 25. Kennedy et al. (2023). Dynamic genetic adaptation of Bacteroides thetaiotaomicron during murine gut colonization. Cell Rep 42, 113009. 10.1016 / j.celrep.2023.113009.

[0447] 26. Shiver et al. (2023). A mutant fitness compendium in Bifidobacteria reveals molecular determinants of colonization and host-microbe interactions. bioRxiv. 10.1101 / 2023.08.29.555234.

[0448] 27. Guerrero-Egido et al. (2024). bacLIFE: a user-friendly computational workflow for genome analysis and prediction of lifestyle-associated genes in bacteria. Nat Commun 75, 2072.

[0449] 28. Xiao et al. (2021). Mining genome traits that determine the different gut colonization potential of Lactobacillus and Bifidobacterium species. Microb Genom 7. 10.1099 / mgen.0.000581.

[0450] 29. Nayfach et al. (2021). A genomic catalog of Earth’s microbiomes. Nature Biotechnology 39, 499-509. 10.1038 / s41587-020-0718-6.

[0451] 30. Tully et al. (2018). The reconstruction of 2,631 draft metagenome- assembled genomes from the global oceans. Scientific Data 5, 170203.

[0452] 31. Huttenhower et al. (2012). Structure, function and diversity of the healthy human microbiome. Nature 486, 207-214. 10.1038 / nature 11234.

[0453] 32. Thompson G., et al. (2017). A communal catalogue reveals Earth’s multiscale microbial diversity. Nature 557, 457-463. 10.1038 / nature24621.

[0454] 33. Nayfach et al. (2019). New insights from uncultivated genomes of the global human gut microbiome. Nature 568, 505-510. 10.1038 / s41586-019-1058-x.

[0455] 34. Hug et al. (2016). A new view of the tree of life. Nat Microbiol 1, 16048. 10.1038 / nmicrobiol.2016.48.

[0456] 35. Coelho et al. (2022). Towards the biogeography of prokaryotic genes. Nature 601, 252-256.

[0457] 36. Rodriguez et al. (2023). Functional and evolutionary significance of unknown genes from uncultivated taxa. Nature. 10.1038 / s41586-023-06955-z.

[0458] 37. Jim, K., Parmar, K., Singh, M., and Tavazoie, S. (2004). A cross-genomic approach for systematic mapping of phenotypic traits to genes. Genome Res 14, 109-115. 10.1101 / gr.1586704.

[0459] 38. Slonim, N., Elemento, O., and Tavazoie, S. (2006). Ab initio genotype-phenotype association reveals intrinsic modularity in genetic networks. Mol Syst Biol 2, 2006 0005. 10.1038 / msb4100047.

[0460] 39. Surette et al. (1999). Quorum sensing in Escherichia coli, Salmonella typhimurium, and Vibrio harveyi: a new family of genes responsible for autoinducer production. Proc Natl Acad Sci U S A 96, 1639-1644. 10.1073 / pnas.96.4.1639.

[0461] 40. Vendeville et al. (2005). Making 'sense' of metabolism: autoinducer-2, LuxS and pathogenic bacteria. Nat Rev Microbiol 3, 383-396. 10.1038 / nrmicrol l46.

[0462] 41. Gagarinova et al. (2016). Systematic Genetic Screens Reveal the Dynamic Global Functional Organization of the Bacterial Translation Machinery. Cell Rep 17, 904-916.

[0463] 42. Wicker- Planquart et al. (2008). Interactions of an essential Bacillus subtilis GTPase, YsxC, with ribosomes. J Bacteriol 190, 681-690. 10.1128 / JB.01193-07.

[0464] 43. Schaefer et al. (2006). Multiple GTPases participate in the assembly of the large ribosomal subunit in Bacillus subtilis. J Bacteriol 788, 8252-8258. 10.1128 / JB.01213-06.

[0465] 44. Ni et al. (2016). YphC and YsxC GTPases assist the maturation of the central protuberance, GTPase associated region and functional core of the 50S ribosomal subunit. Nucleic Acids Res 44, 8442- 8455. 10.1093 / nar / gkw678.

[0466] 45. Cooper et al. (2009). YsxC, an essential protein in Staphylococcus aureus crucial for ribosome assembly / stability. BMC Microbiol 9, 266. 10.1186 / 1471-2180-9-266.

[0467] 46. Bologna et al. (2010). Characterization of Escherichia coli EutD: a phosphotransacetylase of the ethanolamine operon. J Microbiol 48, 629-636. 10.1007 / sl2275-010-0091-0.

[0468] 47. Brinsmade et al. (2004). The eutD gene of Salmonella enterica encodes a protein with phosphotransacetylase enzyme activity. J Bacteriol 186, 1890- 1892. 10.1128 / JB.186.6.1890-1892.2004.

[0469] 48. Wang et al. (2023). Strain dropouts reveal interactions that govern the metabolic output of the gut microbiome. Cell 186, 2839-2852.e2821. 10.1016 / j.cell.2023.05.037. 49. Pena-Diaz et al. (2023). Quorum sensing modulates bacterial virulence and colonization dynamics of the gastrointestinal pathogen Citrobacter rodentium. Gut Microbes 75, 2267189.

[0470] 50. Yang et al. (2022). Within-host evolution of a gut pathobiont facilitates liver translocation. Nature 607, 563-570. 10.1038 / s41586-022-04949-x.

[0471] 51. Hullahalli et al. (2021). Pathogen clonal expansion underlies multiorgan dissemination and organ- specific outcomes during murine systemic infection. Elife 10. 10.7554 / eLife.70910.

[0472] 52. Russell et al. (2022). Intestinal transgene delivery with native E. coli chassis allows persistent physiological changes. Cell 785. 3263-3277, e3215.

[0473] 53. Anderson, E.S. (1975). Viability of, and transfer of a plasmid from, E. coli K12 in human intestine. Nature 255, 502-504.

[0474] 54. Smith, H.W. (1975). Survival of orally administered E. coli K12 in alimentary tract of man. Nature 255, 500-502.

[0475] 55. Lasaro et al. (2014). Escherichia coli isolate for studying colonization of the mouse intestine and its application to two-component signaling knockouts. J Bacteriol 196, 1723-1732.

[0476] 56. Furuno et al. (2000). Negative control of plasmid pSClOl replication by increased concentrations of both initiator protein and iterons. J Gen Appl Microbiol 46, 29-37. 10.2323 / jgam.46.29.

[0477] 57. Thompson et al. (2018). Isolation and characterization of novel mutations in the pSClOl origin that increase copy number. Sci Rep 8, 1590. 10.1038 / s41598-018-20016-w.

[0478] 58. von Mering et al. (2007). Quantitative phylogenetic assessment of microbial communities in diverse environments. Science 315, 1126-1130. 10.1126 / science.1133420.

[0479] 59. Parks et al. (2022). GTDB: an ongoing census of bacterial and archaeal diversity through a phylogenetically consistent, rank normalized and complete genome-based taxonomy. Nucleic Acids Res 50, D785-D794. 10.1093 / nar / gkab776.

[0480] 60. Zhou et al. (2020). GTDB: an integrated resource for glycosyltransferase sequences and annotations. Database (Oxford) 2020. 10.1093 / database / baaa047.

[0481] 61. Steinegger, M., and Soding, J. (2017). MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol 35, 1026-1028. 10.1038 / nbt.3988.

[0482] 62. Hauser et al. (2016). MMseqs software suite for fast and deep clustering and searching of large protein sequence sets. Bioinformatics 32, 1323-1330.

[0483] 63. Harjes et al. (2021). Experimentally based structural model of Yihl provides insight into its function in controlling the key translational regulator Gcn2. FEBS Lett 595, 324-340.

[0484] 64. Silva et al. (2015). The Gcn2 Regulator Yihl Interacts with the Cyclin Dependent Kinase Cdc28 and Promotes Cell Cycle Progression through G2 / M in Budding Yeast. PLoS One 10, e0131070.

[0485] 65. Cambiaghi et al. (2014). Evolutionarily conserved IMPACT impairs various stress responses that require GCN1 for activating the eIF2 kinase GCN2. Biochem Biophys Res Commun 443, 592-597.

[0486] 66. Waller, T., Lee, S.J., and Sattlegger, E. (2012). Evidence that Yihl resides in a complex with ribosomes. FEBS J 279, 1761-1776. 10.1111 / j.1742-4658.2012.08553.X.

[0487] 67. Sattlegger et al. (2011). Genl and actin binding to Yihl: implications for activation of the eIF2 kinase GCN2. J Biol Chem 286, 10341-10355.

[0488] 68. Sattlegger et al. (2004). YIH1 is an actin-binding protein that inhibits protein kinase GCN2 and impairs general amino acid control when overexpressed. J Biol Chem 279, 29952-29962.

[0489] 69. Sunagawa et al. (2015). Structure and function of the global ocean microbiome. Science 348, 1261359. 10.1126 / science.l261359.

[0490] 70. Beresford-Jones et al. (2022). The Mouse Gastrointestinal Bacteria Catalogue enables translation between the mouse and human gut microbiotas via functional mapping. Cell Host Microbe 30, 124-138 el28. 10.1016 / j.chom.2021.12.003.

[0491] 71. Mukherjee et al., (2020). Genomes OnLine Database (GOLD) v.8: overview and updates. Nucleic Acids Research 49, D723-D733.

[0492] 72. Chen et al. (2021). The IMG / M data management and analysis system v.6.0: new tools and advanced capabilities. Nucleic Acids Res 49, D751-D763.

[0493] 73. Nayfach et al. (2021). CheckV assesses the quality and completeness of metagenome- assembled viral genomes. Nat Biotechnol 39, 578-585. 10.1038 / s41587-020-00774-7.

[0494] 74. Tyson et al. (2004). Community structure and metabolism through reconstruction of microbial genomes from the environment. Nature 428, 37-43. 10.1038 / nature02340.

[0495] 75. Laganenka et al. (2023). Chemotaxis and autoinducer-2 signalling mediate colonization and contribute to co-existence of Escherichia coli strains in the murine gut. Nat Microbiol 8, 204-217.

[0496] 76. Laganenka et al. (2018). Autoinducer 2-Dependent Escherichia coli Biofilm Formation Is Enhanced in a Dual-Species Coculture. Applied and Environmental Microbiology 84, e02638-02617.

[0497] 77. Townsend et al. (2020). A Master Regulator of Bacteroides thetaiotaomicron Gut Colonization Controls Carbohydrate Utilization and an Alternative Protein Synthesis Factor. mBio 11. 10.1128 / mBio.03221-19.

[0498] 78. Han et al. (2023). Gut colonization by Bacteroides requires translation by an EF-G paralog lacking GTPase activity. EMBO J. 42, el 12372. 10.15252 / embj.2022112372.

[0499] 79. Holden et al. (2021). Massively parallel transposon mutagenesis identifies temporally essential genes for biofilm formation in Escherichia coli. Microb Genom 7. 10.1099 / mgen.0.000673.

[0500] 80. Butland et al. (2005). Interaction network containing conserved and essential protein complexes in Escherichia coli. Nature 433, 531-537. 10.1038 / nature03239.

[0501] 81. Park et al. (2004). Crystal structure of YIGZ, a conserved hypothetical protein from Escherichia coli kl2 with a novel fold. Proteins 55, 775-777. 10.1002 / prot.20087.

[0502] 82. Bradley et al. (2018). Phylogeny-corrected identification of microbial gene families relevant to human gut colonization. PLOS Computational Biology 14, el006242. 10.1371 / journal.pcbi.1006242.

[0503] 83. Levy et al. (2018). Genomic features of bacterial adaptation toplants. Nature Genetics 50, 138- 150. 10.1038 / S41588-017-0012-9.

[0504] 84. Dutilh et al. (2013). Explaining microbial phenotypes on a genomic scale: GWAS for microbes. Brief Funct Genomics 12, 366-380. 10.1093 / bfgp / elt008.

[0505] 85. Han et al. (2021). A metabolomics pipeline for the mechanistic interrogation of the gut microbiome. Nature 595, 415-420. 10.1038 / s41586-021-03707-9.

[0506] 86. VanEvery et al. (2023). Microbiome epidemiology and association studies in human health. Nat Rev Genet 24, 109-124. 10.1038 / s41576-022-00529-x.

[0507] 87. Hutson, M. (2023). Foldseek gives AlphaFold protein database a rapid search tool. Nature. 10.1038 / d41586-023-02205-4.

[0508] 88. Kempen et al. (2023). Fast and accurate protein structure search with Foldseek. Nat Biotechnol. 10.1038 / S41587-023-01773-0.

[0509] 89. Barrio-Hernandez et al. (2023). Clustering predicted structures at the scale of the known protein universe. Nature 622, 637-645. 10.1038 / s41586-023-06510-w.

[0510] 90. Carini et al. (2016). Relic DNA is abundant in soil and obscures estimates of soil microbial diversity. Nature Microbiology 2, 16242. 10.1038 / nmicrobiol.2016.242.

[0511] 91. Zhernakova et al. (2024). Host genetic regulation of human gut microbial structural variation. Nature 625, 813-821. 10.1038 / s41586-023-06893-w.

[0512] 92. Goodrich et al. (2016). Cross-species comparisons of host genetic associations with the microbiome. Science 352, 532-535. 10.1126 / science.aad9379.

[0513] 93. Jiang et al. (2020). Comprehensive Genome-wide Perturbations via CRISPR Adaptation Reveal Complex Genetics of Antibiotic Sensitivity. Cell 180, 1002-1017 el031. 10.1016 / j.cell.2020.02.007.

[0514] 94. Datsenko et al. (2000). One-step inactivation of chromosomal genes in Escherichia coli K-12 using PCR products. Proc Natl Acad Sci U S A 97, 6640-6645. 10.1073 / pnas.120163297.

[0515] 95. van Dongen, S., and Abreu-Goodger, C. (2012). Using MCL to extract clusters from networks. Methods Mol Biol 804, 281-295. 10.1007 / 978-l-61779-361-5_15.

[0516] 96. Sprouffske, K., and Wagner, A. (2016). Growthcurver: an R package for obtaining interpretable metrics from microbial growth curves. BMC Bioinformatics 17, 172. 10.1186 / S12859-016-1016-7.

[0517] 97. Parks et al. (2018). A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat Biotechnol 36, 996-1004. 10.1038 / nbt.4229.

[0518] 98. Borges et al. (2018). Measuring phylogenetic signal between categorical traits and phylogenies. Bioinformatics 35, 1862-1869. 10.1093 / bioinformatics / bty800.

[0519] 99. Pagel, M. (1999). Inferring the historical patterns of biological evolution. Nature 401, 877-884. 10.1038 / 44766.

[0520] 100. Parks et al. (2015). CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res 25, 1043-1055. 10.1101 / gr.186072.114.

[0521] 101. Chaumeil et al. (2019). GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics 36, 1925-1927. 10.1093 / bioinformatics / btz848.

[0522] 102. Buchfink, B., Xie, C., and Huson, D.H. (2015). Fast and sensitive protein alignment using DIAMOND. Nat Methods 12, 59-60. 10.1038 / nmeth.3176.

[0523] 103. Buchfink, B., Reuter, K„ and Drost, H.-G. (2021). Sensitive protein alignments at tree-of-life scale using DIAMOND. Nature Methods 18, 366-368. 10.1038 / s41592-021-01101-x.

[0524] 104. Girgis, H.S., Liu, Y., Ryu, W.S., and Tavazoie, S. (2007). A comprehensive genetic characterization of bacterial motility. PLoS Genet 3, 1644-1660. 10.1371 / joumal.pgen.0030154.

[0525] 105. Karp et al. (2019). The BioCyc collection of microbial genomes and metabolic pathways. Brief Bioinform 20, 1085-1093. 10.1093 / bib / bbx085.

[0526] 106. Soudy et al. (2020). UniprotR: Retrieving and visualizing protein sequence and functional information from Universal Protein Resource (UniProt knowledgebase). J Proteomics 213, 103613. 10.1016 / j.jprot.2019.103613.

[0527] 107. Supek, F., Bosnjak, M., Skunca, N., and Smuc, T. (201 1). REVIGO summarizes and visualizes long lists of gene ontology terms. PLoS One 6, e21800. 10.1371 / journal.pone.0021800.

[0528] 108. Ondov et al. (2016). Mash: fast genome and metagenome distance estimation using MinHash. Genome Biol 17, 132. 10.1186 / sl3059-016-0997-x.

[0529] 109. Revell, L.J. (2024). phytools 2.0: an updated R ecosystem for phylogenetic comparative methods (and other things). PeerJ 12, el6505. 10.7717 / peerj.16505.

[0530] 110. Shen, W., Le, S., Li, Y., and Hu, F. (2016). SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA / Q File Manipulation. PLoS One 11, e0163962. 10.1371 / journal.pone.0163962. 111. Van Dongen, S. (2008). Graph clustering via a discrete uncoupling process. SIAM Journal on Matrix Analysis and Applications 30, 121-141.

[0531] 112. Enright, A.J., Van Dongen, S., and Ouzounis, C.A. (2002). An efficient algorithm for large-scale detection of protein families. Nucleic Acids Res 30, 1575-1584. 10.1093 / nar / 30.7.1575.

[0532] 113. Carpenter et al. (2006). CellProfiler: image analysis software for identifying and quantifying cell phenotypes. Genome Biol 7, R100. 10.1186 / gb-2006-7-10-rl00.

[0533] 114. Stirling et al. (2021). CellProfiler 4: improvements in speed, utility and usability. BMC Bioinformatics 22, 433.

[0534] 115. Duong et al. Construction of vectors for inducible and constitutive gene expression in Lactobacillus. Microb Biotechnol 4, 357-367 (2011).

[0535] 116. Wortelboer et al. Fecal microbiota transplantation beyond Clostridioides difficile infections. EBioMedicine 44, 716-729 (2019).

[0536] 117. Pierog et al., Fecal microbiota transplantation in children with recurrent Clostridium difficile infection. The Pediatric infectious disease journal 33, 1198-1200 (2014).

[0537] 118. Routy et al., Gut microbiome influences efficacy of PD-l-based immunotherapy against epithelial tumors. Science 359. 91-97 (2018).

[0538] 119. Gopalakrishnan et al., Gut microbiome modulates response to anti-PD-1 immunotherapy in melanoma patients. Science 359, 97-103 (2018).

[0539] 120. Dubin et al., Intestinal microbiome analyses identify melanoma patients at risk for checkpointblockade-induced colitis. Nat Commun 7, 10391 (2016).

[0540] Example 2

[0541] A set of plasmids containing the identified colonization factor YigZ is described here. The yigZ gene may be driven by its native promoter in the targets LBP (Figure 14A), or a universal inducible promoter activated by arabinose (Figure 14B). The expression system may be combined with a “suicide” toxin-antitoxin system to achieve a scenario of dual colonize-or-die effect (Figures 14C and 14D): when inducer arabinose is added, both YigZ and antitoxin (MazE) are produced to enhance colonization but also to neutralize a constitutively expressed toxin (MazF); when arabinose is no longer provided, the cell will be killed by the toxin, achieving a suicide effect to enable precise control of LBP colonization.

[0542] The yigZ and toxin / antitoxin genes may be codon-optimized for the target LBP host and inducible promoters will be optimized for induction in the target LBP. Vectors that may be used with the present compositions and methods include the Bifidobaclerium-Escherichia coll shuttle vector series (e.g., pKO403, www.addgetie.org / 174725 / , H. Altaib et al., Bifidobacterium-Escherichia coli Shuttle Vector Series pKO403, with Temperature- Sensitive Replication Origin for Gene Knockout in Bifidobacterium, Microbiol. Resour. Announc. 11, e0088421 (2022)), and pTRK892 for gene expression in Lactobacillus gasseri (www. addgene. org / 71803 / ; Mimee et al., Programming a Human Commensal Bacterium, Bacteroides thetaiotaomicron, to Sense and Respond to Stimuli in the Murine Gut Microbiota, Cell Syst. 2, 214 (2016)). Bacteroides species (e.g., pSIEl which can achieve targeted genetic manipulation of diverse wild-type Bacteroides species from the human gut) and B. thetaiotaomicron may be chromosomally edited through existing systems such as the ones in Jones et al., Sequence and characterization of shuttle vectors for molecular cloning in Porphyromonas, Bacteroides and related bacteria, Mol. Oral. Microbiol. 35, 181-191 (2020), and Lim et al., Engineered Regulatory Systems Modulate Gene Expression of Human Commensals in the Gut, Cell, 169, 547-558;e515 (2017).

[0543] The scope of the present invention is not limited by what has been specifically shown and described hereinabove. Those skilled in the art will recognize that there are suitable alternatives to the depicted examples of materials, configurations, constructions and dimensions. Numerous references, including patents and various publications, are cited and discussed in the description of this invention. The citation and discussion of such references is provided merely to clarify the description of the present invention and is not an admission that any reference is prior art to the invention described herein. All references cited and discussed in this specification are incorporated herein by reference in their entirety. Variations, modifications and other implementations of what is described herein will occur to those of ordinary skill in the art without departing from the spirit and scope of the invention. While certain embodiments of the present invention have been shown and described, it will be obvious to those skilled in the art that changes and modifications may be made without departing from the spirit and scope of the invention. The matter set forth in the foregoing description and accompanying drawings is offered by way of illustration only and not as a limitation. SEQUENCES:

[0544] Amino acid sequence of YigZ (E. coli Nissle 1917):

[0545] (GenBank accession number: WP_001295262)

[0546] MESWLIPAAPVTVVEEIKKSRFITLLAHTDGVEAAKAFVESVRAEHPDARHHCVAWVAGAP

[0547] DDSQQLGFSDDGEPAGTAGKPMLAQLMGSGVGEITAVVVRYYGGILLGTGGLVKAYGGGV NQALRQLTTQRKTPLTEYTLQCEYSQLTGIEALLGQCDGKIINSDYQAFVLLRVALPAAKVA EFSAKLADFSRGSLQLLAIEE (SEQ ID NO:1)

[0548] Amino acid sequence of YigZ (E. coli MG1655):

[0549] (GenBank accession number: XKS61981)

[0550] MESWLIPAAPVTVVEEIKKSRFITMLAHTDGVEAAKAFVESVRAEHPDARHHCVAWVAGA

[0551] PDDSQQLGFSDDGEPAGTAGKPMLAQLMGSGVGEITAVVVRYYGGILLGTGGLVKAYGGG

[0552] VNQALRQLTTQRKTPLTEYTLQCEYHQLTGIEALLGQCDGKIINSDYQAFVLLRVALPAAKV AEFSAKLADFSRGSLQLLAIEE (SEQ ID NO:2)

[0553] Amino acid sequence of YigZ (Lactobacillus gasseri):

[0554] (UniProt accession number: A0A833CEL3)

[0555] MSSKQLNYLTISKAGQHELIIKKSKFICSLARTKTVEEAQEFIEQISKKYHDATHNTYAY TLGLNDNQVKASDNGEPSGTAGIPELKALQLMKLKNVTAVVTRYFGGIKLGAGGLIRAYS NSVTEAAQNIGVVKCVMQQRIQFSIPYNRIDEINHYLEENRISIANQEYTTNVTIQIYLD LDQIQKVEDDLINLLSGKVEFNKLDQRFNEIPVTDFNFHEQ (SEQ ID NO:3)

[0556] Amino acid sequence of YigZ (Bifidobacterium dolichotidis):

[0557] (UniProt accession number: A0A430FQ87)

[0558] MRTLENPPEEPAHDSFVEKKSEFIGDACHVESFEDAEAFVQSIRDQHPKARHVAWAVVCT

[0559] DENGNASERMSDDGEPSGTAGKPILEVERMNEETNVAVTVTRYFGGIEEGSGGLTRAYST

[0560] GASIAVKAAQQAQIVPCSAYHTTIEYTQLGQMQRLEQQMDGEQRDAEFTDRVSETAVVPS DRAQVFEDQVRESFNATVSEEPAGTVMRNVTA (SEQ ID NO:4)

[0561] Amino acid sequence of YigZ (Bacteroides thetaiotaomicrori):

[0562] (UniProt accession number: A0A174QGE9)

[0563] MTAEDTYKTIVEPSEGIYTEKRSKFIAIALPVRTLDEIKAHLETYQKKYYDARHVCYAYM LGAARKDFRANDNGEPSGTAGKPILGQINSNELTDILIIVVRYFGGIKLGTSGLIVAYKA

[0564] AAAEAISAATIIEKTVDEEVTVMFEYPFMNDIMRIVKEEEPEILSQSYDMDCSMTLRIRR SMMPKLRARLEKVETARILDEE (SEQ ID N0:5)

[0565] Amino acid sequence of TrhP (E. coli MG1655):

[0566] (GenBank accession number: NP_416585)

[0567] MFKPELLSPAGTLKNMRYAFAYGADAVYAGQPRYSLRVRNNEFNHENLQLGINEAHALG

[0568] KKFYVVVNIAPHNAKLKTFIRDLKPVVEMGPDALIMSDPGLIMLVREHFPEMPIHLSVQAN

[0569] AVNWATVKFWQQMGLTRVILSRELSLEEIEEIRNQVPDMEIEIFVHGALCMAYSGRCLLSG

[0570] YINKRDPNQGTCTNACRWEYNVQEGKEDDVGNIVHKYEPIPVQNVEPTLGIGAPTDKVFM

[0571] IEEAQRPGEYMTAFEDEHGTYIMNSKDLRAIAHVERLTKMGVHSLKIEGRTKSFYYCART

[0572] AQVYRKAIDDAAAGKPFDTSLLETLEGLAHRGYTEGFLRRHTHDDYQNYEYGYSVSDRQ

[0573] QFVGEFTGERKGDLAAVAVKNKFSVGDSLELMTPQGNINFTLEHMENAKGEAMPIAPGD

[0574] GYTVWLPVPQDLELNYALLMRNFSGETTRNPHGK (SEQ ID NO:6)

Claims

What is claimed is:

1. A method of enhancing gastrointestinal (GI) tract colonization of a bacterium in a mammal, the method comprising introducing a recombinant bacterial expression vector into the bacterium, wherein the bacterial expression vector comprises a gene encoding a gastrointestinal (GI) tract colonization factor, wherein the colonization factor comprises IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

2. A method of treating a disorder in a subject, the method comprising administering a recombinant bacterium to the subject, wherein the recombinant bacterium comprises a bacterial expression vector, wherein the bacterial expression vector comprises a gene encoding a gastrointestinal (GI) tract colonization factor, wherein the colonization factor is IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

3. The method of claim 2, wherein the disorder is an infectious disease, an autoimmune disease, an allergic disease or cancer.

4. The method of claim 3, wherein the autoimmune disease comprises organ transplant rejection, inflammatory bowel disease (IBD), ulcerative colitis, Crohn's disease, sprue, rheumatoid arthritis, Type I diabetes, graft versus host disease, or multiple sclerosis.

5. The method of claim 3, wherein the infectious disease is associated with infectious pathogens comprising, Salmonella, Shigella, Clostridium difficile, Mycobacterium, protozoa, filarial nematodes, Schistosoma, Toxoplasma, Leishmania, hepatitis C virus (HCV), hepatitis B virus (HBV), and / or herpes simplex viruses.

6. The method of any of claims 2-5, wherein the subject is a mammal.

7. The method of claim 6, wherein the mammal is a primate, a bovine, an equine, a canine or a feline.

8. The method of any of claims 2-7, wherein the subject is a human.

9. The method of any of claims 1-8, wherein YigZ comprises an amino acid sequence at least 90% identical to the amino acid sequence set forth in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, or SEQ ID NO:5.

10. The method of any of claims 1-8, wherein TrhP comprises an amino acid sequence at least 90% identical to the amino acid sequence set forth in SEQ ID NO:6.

11. The method of any of claims 1-8, wherein the colonization factor is selected from the group consisting of: YigZ (GenBank accession number: WP_001295262), YigZ (GenBank accession number: XKS61981), YigZ (UniProt accession number: A0A833CEL3), YigZ (UniProt accession number: A0A430FQ87), YigZ (UniProt accession number: A0A174QGE9), TrhP (GenBank accession number: NP_416585), or combinations thereof.

12. The method of any of claims 1-11, wherein the bacterium belongs to phyla Bacteroidetes, Firmicutes, Proteobacteria, or Actinobacteria.

13. The method of any of claims 1-11, wherein the bacterium belongs to genera Bacteroides, Clostridium, Escherichia, Bacillus, Lactobacillus, or Bifidobacterium.

14. The method of any of claims 1-11, wherein the bacterium is Escherichia coli, Lactobacillus gasseri. Bifidobacterium dolichotidis, Bacteroides thetaiotaomicron, or Salmonella enterica.

15. A recombinant bacterial expression vector, comprising a gene encoding a gastrointestinal (GI) tract colonization factor, wherein the colonization factor comprises IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

16. The bacterial expression vector of claim 15, wherein YigZ comprises an amino acid sequence at least 90% identical to the amino acid sequence set forth in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NOG, SEQ ID NO:

4. or SEQ ID NOG.

17. The bacterial expression vector of claim 15, wherein TrhP comprises an amino acid sequence at least 90% identical to the amino acid sequence set forth in SEQ ID NO:6.

18. The bacterial expression vector of claim 15, wherein the colonization factor is selected from the group consisting of: YigZ (GenBank accession number: WP_001295262), YigZ (GenBank accession number: XKS61981), YigZ (UniProt accession number: A0A833CEL3), YigZ (UniProt accession number: A0A430FQ87), YigZ (UniProt accession number: A0A174QGE9), TrhP (GenBank accession number: NP_416585), or combinations thereof.

19. The bacterial expression vector of any of claims 15-18, wherein the gene is under the control of a constitutive promoter.

20. The bacterial expression vector of any of claims 15-18, wherein the gene is under the control of an inducible promoter.

21. The bacterial expression vector of any of claims 15-18, wherein the colonization factor is wildtype or a mutant.

22. A recombinant bacterium, comprising the bacterial expression vector of any of claims 15-21.

23. A recombinant bacterium, comprising a bacterial expression vector, wherein the bacterial expression vector comprises a gene encoding a gastrointestinal (GI) tract colonization factor, wherein the colonization factor is IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

24. A recombinant bacterium, comprising a gene encoding a gastrointestinal (GI) tract colonization factor, wherein the colonization factor is IMPACT family member YigZ, tRNA hydroxylation protein P (TrhP), or a combination thereof.

25. The recombinant bacterium of any of claims 22-24, wherein the colonization factor is endogenous to the bacterium, and wherein the bacterium overexpresses the colonization factor.

26. The recombinant bacterium of any of claims 22-24, wherein the colonization factor is heterologous to the bacterium.

27. The recombinant bacterium of any of claims 22-26, wherein the gene is integrated into a bacterial chromosome or is episomal.

28. The recombinant bacterium of any of claims 22-27, belonging to phyla Bacteroidetes, Firmicutes, Proteobacteria. or Actinobacteria.

29. The recombinant bacterium of any of claims 22-27, belonging to genera Bacteroides, Clostridium, Escherichia, Bacillus, Lactobacillus, or Bifidobacterium.

30. The recombinant bacterium of any of claims 22-27, wherein the bacterium is Escherichia coll, Lactobacillus gasseri, Bifidobacterium dolichotidis, Bacteroides thetaiotaomicron, or Salmonella enterica.

31. The recombinant bacterium of any of claims 22-27, wherein the recombinant bacterium is Bacillus subtilis, Bifidobacterium longum, E. coli Nissle 1917, Lactobacillus plantarum, Lactobacillus lactis, Lactobacillus casei, Lactobacillus reuteri, Lactobacillus gasseri, Streptococcus typhimurium, Bacteroides thetaiotaomicron, Bacteroides fragilis, Bacteroides ovatus, or Akkermansia muciniphila.

32. The recombinant bacterium of any of claims 22-27, wherein the recombinant bacterium is Escherichia coli (E. coli), Bacillus subtilis (B. subtilis), Lactobacillus plantarum (L. plantarum), Lactobacillus reuteri (L. reuteri), Lactobacillus rhamnosus (L. rhamnosus), Bifidobacterium longum (B. longum), Bacteroides thetaiotaomicron, Clostridium butyricum, Bacteroides fragilis, Bacteroides melaninogenicus, Bacteroides oralis, Bacteroides amylophilus, Clostridium butyricum, Clostridium perfringens, Clostridium tetani, or Clostridium septicum.

33. The recombinant bacterium of any of claims 22-27, wherein the bacterium is from a natural gut microbiota.

34. The recombinant bacterium of claim 33, wherein the natural gut microbiota is from the gut of a healthy mammal, or from the gut of a mammal with inflamed GI tract.

35. The recombinant bacterium of claim 34, wherein the mammal with inflamed GI tract has an inflammatory bowel disease (IBD) or irritable bowel syndrome (IBS).

36. The recombinant bacterium of claim 35, wherein the IBD is ulcerative colitis or Crohn's disease.

37. The recombinant bacterium of claim 34, wherein the healthy mammal, or the mammal with inflamed GI tract, is a human subject.

38. A pharmaceutical composition, comprising the recombinant bacterium of any of claims 22-37.

39. The pharmaceutical composition of claim 38, wherein the pharmaceutical composition is formulated for oral administration.

40. The pharmaceutical composition of claims 38 or 39, wherein the composition is a food composition, a beverage composition, or a feedstuff composition.

41. The pharmaceutical composition of any of claims 38-40, wherein the composition is a dairy product.

42. The pharmaceutical composition of claim 38, wherein the pharmaceutical composition is formulated for rectal administration.

43. A method of identifying colonization factors for bacteria to colonize the mammalian gastrointestinal (GI) tract, the method comprising:(a) collecting mouse gastrointestinal (GI) metagenome-assembled genomes (MAGs), human GI MAGs, and human isolates;(b) aligning protein sequences from host-associated dataset against protein sequences from both host-associated dataset and environmental datasets to create a phylogenetic profile, denoting homolog presence across genomes;(c) analyzing genotype-phenotype association on each of the mouse GI MAGs, human GI MAGs and human isolates against their respective matched environmental genomes, by ranking proteins based on their mutual information Z-score (MI-Z) and correlating habitat phenotype with their phylogenetic profile; and(d) identifying colonization factors (CF) by (1) combining top proteins from the genotypephenotype association analyses of the mouse GI MAGs, human GI MAGs and human isolates, and (2) consolidating the top proteins into homologous groups based on protein sequence similarity.