iMODULON GENE CLUSTERS AND METHODS TO IDENTIFY AND USE SAME

By employing computational analysis and adaptive laboratory evolution, iModulons facilitate efficient transfer of cellular functions across species, addressing the inefficiencies of traditional methods and optimizing host compatibility.

WO2025171062A1PCT designated stage Publication Date: 2025-08-14RGT UNIV OF CALIFORNIA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/014671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2025-02-05
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing methods for identifying and transferring cellular functions across species in synthetic biology are inefficient and require extensive experimental efforts, such as genome-wide knockout studies and Tn-Seq, which are slow and labor-intensive.

Method used

The use of computational tools and large transcriptomic datasets to identify independently modulated gene sets (iModulons) that represent specific cellular functions, combined with adaptive laboratory evolution, allows for the transfer of these functions into alternate hosts.

Benefits of technology

This approach enables efficient and accurate cross-species transfer of cellular functions by identifying the genetic basis for traits, optimizing host compatibility, and enhancing functional expression through automated methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025014671_14082025_PF_FP_ABST
    Figure US2025014671_14082025_PF_FP_ABST
Patent Text Reader

Abstract

Applicant applied machine learning applied to large compendia of transcriptomic data to enable the decomposition of bacterial transcriptomes to identify independently modulated sets of genes, iModulons, that represent specific cellular functions. The identification of iModulons enables accurate identification of genes necessary and sufficient for cross-species transfer of cellular functions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Att Dkt: 114198-5960 iMODULON GENE CLUSTERS AND METHODS TO IDENTIFY AND USE SAME CROSS-REFERENCE TO RELATED PATENT APPLICATIONS This application claims priority under 35 U.S.C. § 119(e) of U.S. Provisional Application No.: 63 / 550,406, filed February 6, 2024, the content of which is incorporated herein by reference in its entirety. BACKGROUND Synthetic biology aims to engineer new, or modify existing, cellular functions by designing, editing, and assembling the underlying genetic material. To achieve this goal, knowledge-based designs have been traditionally used that identify genes of known molecular functions to piece together the targeted cellular function. However, identifying all the necessary genetic components is not trivial and requires large-scale analysis such as genome-wide knockout studies and Tn-Seq1,2. This approach also calls for serial steps of genetic transformation with separate testing of each individual gene of the targeted function3slowing the development of new therapies. For example, development of the biosynthetic pathway of the antimalarial drug artemisinin demanded extensive research efforts and engineering4–6. This disclosure provides new methods that simultaneously query multiple genes involved with targeted cellular functions. SUMMARY OF THE DISCLOSURE To circumvent the difficulties with present approaches, computational tools have been developed to supplement experimental procedures and aid design7–9. More importantly, recent advances in new data analytics applied to large transcriptomic datasets has enabled Applicant’s alternative approach. Using the methods as disclosed herein, Applicant can now identify sets of independently modulated genes (called iModulons) that constitute a particular cellular function10. This advancement opens up the possibility to identify targeted functions in a particular strain and transfer them into alternate hosts. In one aspect, successful cross-species transfer of desired functions would thus require: i) the identification of the full genetic basis for the trait, ii) use of recombinant or DNA synthesis methods to capture these genes into a plasmid, iii) transfer of the plasmid into the target host, and iv) making any needed changes to the new host that are critical to accommodate the transferred function. All these Att Dkt: 114198-5960 capabilities now exist, with ii) and iii) representing known approaches, while i) and iv) require novel transcriptomic analysis and the use of automated adaptive laboratory evolution. As shown herein, the disclosed methods can be applied to find or identify source regulatory signals in bacterial transcriptomes10–13. iModulons are fundamental units of bacterial transcriptomes that have been found to represent the genetic basis for various cellular functions and are associated with particular transcriptional regulator(s)14–16. Many identified iModulons include genes that are distantly located on the genome, genes of unknown functions, or contain accessory genes that augment the targeted cellular function. iModulons thus represent an advanced scale of synthetic biology to transfer naturally evolved traits across species. Also as disclosed herein, Applicant’s methods use iModulons to create new cellular functions in a new host. Applicant shows that transferring iModulons is superior to using operons or single genes identified by genome annotation algorithms and rational design approaches9,17,18. Applicant also provides methods that use adaptive laboratory evolution (ALE) to enable the host to optimally use the new function under selection pressure. Thus, as disclosed herein, Applicant demonstrates that cross-species iModulon transfer is a versatile tool for synthetic biology. In one aspect using this method, Applicant provides compositions comprising or consisting essentially of, or consisting of an iModulon gene cluster wherein the gene cluster is selected from the group of the polynucleotides of: SEQ ID Nos: 1-4 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 5-13 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 14-18 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 19-25 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 26-66 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 67-76 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 26-43 or an equivalent of each thereof; or the polynucleotides of SEQ ID Nos: 53-54, 57-61, and 64-65, or an equivalent of each thereof. This disclosure also provides an isolated host cell or population of cells comprising or consisting essentially of a gene cluster, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 1-4. In other embodiments, an isolated host cell or a population of cells, comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 5-13. In other embodiments, an isolated host Att Dkt: 114198-5960 cell or population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 14-18. In other embodiments, an isolated host cell or a population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 19-25. In other embodiments, an isolated host cell or a population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 26- 66. In other embodiments, an isolated host cell or a population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 67-76. In other embodiments, an isolated host cell or a population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 26-43. In other embodiments, an isolated host cell or a population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 53-54, 57-61, and 64-65. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another aspect, a method for preparing an isolated cell or population of cells with an exogenous biochemical pathway is provided herein, the method comprising, or consisting essentially of, or yet further consisting of introducing an iModulon gene cluster selected from the group of the polynucleotides of: SEQ ID Nos: 1-4 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 5-13 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 14-18 or an equivalent of each thereof; the polynucleotides of SEQ ID Nos: 19- 25 or an equivalent of each thereof, the polynucleotides of SEQ ID Nos: 26-66 or an equivalent of each thereof, the polynucleotides of SEQ ID Nos: 67-76 or an equivalent of each thereof, the polynucleotides of SEQ ID Nos: 26-43 or an equivalent of each thereof, or the polynucleotides of SEQ ID Nos: 53-54, 57-61, and 64-65 or an equivalent of each thereof into the host cell or population of cells such that the polynucleotides of the gene cluster is expressed in the host cell or population of cells, thereby providing the host cell or population of cells with the exogenous biochemical pathway. The host cell can be a prokaryotic host cell or eukaryotic host cell. Also provided herein is a method for preparing an isolated cell or population of cells with exogenously obtained vanillate transport and catabolic pathway of Pseudomonas putida, comprising, or consisting essentially of, or yet further consisting of introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 into the host cell or the population of cells, wherein the host cell or population of cells expresses the polynucleotides of the gene Att Dkt: 114198-5960 cluster and gains the vanillate transport and catabolic pathway of Pseudomonas putida. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another embodiment, a method for preparing an isolated cell or population of cells with exogenously obtained malonate transport and catabolic pathway of Pseudomonas aeruginosa is provided herein, comprising, or consisting essentially of, or yet further consisting of introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 5-13 into the host cell or the population of cells, wherein the host cell expresses the polynucleotides of the gene cluster gains the malonate transport and catabolic pathway of Pseudomonas aeruginosa. The host cell can be a prokaryotic host cell or eukaryotic host cell. Further provided herein is a method for preparing an isolated cell or population of cells with exogenously obtained 2,3-butanediol catabolic pathway of Pseudomonas putida, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 14-18 into the host cell or population of cells, wherein the host cell or the population of cells expresses the gene cluster and gains the 2,3-butanediol catabolic pathway of Pseudomonas putida. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained ampicillin resistance property of Pseudomonas aeruginosa is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 19-25, wherein the host cell or population of cells expresses the gene cluster gains the ampicillin resistance property of Pseudomonas aeruginosa. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained protocatechuate transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43, wherein the host cell or population of cells expresses the gene cluster gains the protocatechuate transport and catabolic pathway of Pseudomonas putida. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained hydroxycinnamates transport and catabolic pathway of Att Dkt: 114198-5960 Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 and 26-66, wherein the host cell or population of cells expresses the gene cluster gains the hydroxycinnamates transport and catabolic pathway of Pseudomonas putida. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained benzoate transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43 and 67-76, wherein the host cell or population of cells expresses the gene cluster gains the benzoate transport and catabolic pathway of Pseudomonas putida. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained vanillate transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 and 26-43, wherein the host cell or population of cells expresses the gene cluster gains the vanillate transport and catabolic pathway of Pseudomonas putida. The host cell can be a prokaryotic host cell or eukaryotic host cell. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained hydroxycinnamates transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43, 53-54, 57-61, and 64-65, wherein the host cell or population of cells expresses the gene cluster gains the hydroxycinnamates transport and catabolic pathway of Pseudomonas putida. The host cell can be a prokaryotic host cell or eukaryotic host cell. This disclosure also provides compositions comprising an iModulon gene cluster as disclosed herein, and vectors or carriers, for use in the methods as disclosed herein. In embodiment, a method comprises, or consists essentially of, or consists of, generating transcriptomic data of one or more organisms under a plurality of conditions, wherein the transcriptomic data comprises expression levels for a set of genes under each of Att Dkt: 114198-5960 the plurality of conditions; determining one or more iModulons from the transcriptomic data by applying independent component analysis on the transcriptomic data, the one or more iModulons each comprising a plurality of genes comprising at least two genes that are expressed together across the plurality of different conditions and comprising a first at least one gene with a known function and a second at least one gene with an unknown function; for each iModulon of the one or more iModulons: creating a strain of a phenotypic expression from the plurality of genes of the iModulon by refactoring the plurality of genes of the iModulon into a gene transfer vehicle; and sequencing the strain of phenotypic expression and the gene transfer vehicle to determine mutations for the strain of phenotypic expression. In some embodiments, the method further comprises applying an adaptive laboratory evolution algorithm to the strain of the phenotypic expression to create an optimized strain of phenotypic expression for performing a function. In some embodiments, refactoring the plurality of genes of the iModulon into the gene transfer vehicle comprises refactoring the plurality of genes of the iModulon into a bacterial artificial chromosome. In some embodiments, the method further comprises creating an expression profile for each of the plurality of conditions, each expression profile comprising a vector indicating an expression level of each gene of the plurality of genes in the condition of the expression profile, wherein determining the one or more iModulons of the organism comprises applying independent component analysis on the one or more expression profiles. In some embodiments, determining the one or more iModulons of the organism comprises determining a potential iModulon includes only a single gene with a gene expression exceeding a threshold; and discarding the potential iModulon responsive to the determination that the potential iModulon only includes a single gene. In some embodiments, the plurality of conditions comprise a pH level adjustment, a change in ions, and an added mixture of other cells. In some embodiments, the method further comprises mapping the known function of the first at least one gene onto the iModulon. In some embodiments, the method further comprises generating a record identifying the mutations of the iModulon; and creating a second strain of the phenotypic expression based on the sequenced mutations. Further provided herein is a method comprising, or consisting essentially of, or consisting of determining one or more iModulons of an organism or cell by applying Att Dkt: 114198-5960 independent component analysis on gene expression data of the organism or cell, the one or more iModulons comprising one or more genes of an unidentified function; creating a strain of a phenotypic expression from the one or more genes of the unidentified function by refactoring the one or more genes of the unidentified function into a gene transfer vehicle; optionally applying an adaptive laboratory evolution algorithm to the strain of the phenotypic expression to create an optimized strain of phenotypic expression for performing a function; and sequencing the optimized strain of phenotypic expression and the gene transfer vehicle to determine mutations for the optimized strain of phenotypic expression. BRIEF DESCRIPTION OF THE DRAWINGS FIGS. 1A-1E: Cross-species transfer of the Pseudomonas putida VanR iModulon into E. coli. (FIG. 1A) Vanillate transport and conversion in P. putida. OM, outer membrane. CM, cytoplasmic membrane. (FIG.1B) iModulon weights of genes in P. putida. Four genes (vanA, vanB, vanK, galP-IV) with high weighting constitute the VanR iModulon. Gray lines indicate thresholds of determining iModulon membership. Gray circles identify genes not in the iModulon. (FIG. 1C) Graphical representation of vanR locus on the P. putida chromosome. (FIG. 1D) The VanR iModulon was refactored in a single operon under the control of trc promoter (PTrc) resulting in the pVanR_iM plasmid. Shades show genetic rearrangement for cloning purpose. (FIG. 1E) Vanillate (VA) conversion of E. coli carrying empty or pVanR_iM plasmid into proto-catechuate (PCA). Circles, diamonds, and triangles indicate cell density, VA, and PCA levels of the culture, respectively. Measurements from E. coli carrying empty or pVanR_iM plasmid are represented by hollow or filled symbols, respectively. Data are presented as mean values + / − SD. Error bars indicate SD of three replicate cultures. Source data are provided as a Source Data file. FIGS 2A-2D: E. coli carrying the Pseudomonas aeruginosa AmpC iModulon confers better ampicillin resistance than cells expressing beta-lactamase alone. (FIG. 2A) iModulon weights of genes in P. aeruginosa. Seven genes constitute the AmpC iModulon. Gray lines indicate thresholds of determining iModulon membership. Gray circles identify genes not in the iModulon. (FIG. 2B) Refactoring the P. aeruginosa AmpC iModulon on bacterial artificial chromosome (BAC). Genes are expressed with the trc promoter (PTrc). Shades show genetic rearrangement for cloning purpose. (FIG. 2C) Dose- kill curves of P. aeruginosa and E. coli carrying empty BAC, BAC_ampC, or BAC_AmpC_iM. Data are presented as mean values + / − SD. Error bars indicate SD of biological replicates (n=3). Dose-kill curves show lower cell density over time in order from Att Dkt: 114198-5960 0 µg / ml ampicillin (top row) to 8192 µg / ml ampicillin (bottom row). Note that the range of ampicillin concentration (Amp) is different, due to the huge difference in ampicillin tolerance. (FIG. 2D) Cell density of cultures treated with different ampicillin concentrations after 10 hr of incubation. Arrows indicate the minimum inhibitory concentration (MIC). Data are presented as mean values + / − SD. Error bars indicate SD of biological replicates (n=3). Source data are provided as a Source Data file. FIGS. 3A-3E: Cross-species transfer of 2,3-butanediol (2,3-BDO) utilization iModulon of Pseudomonas putida in E. coli. (FIG. 3A) A pathway responsible for 2,3- BDO utilization. (FIG. 3B) Scatter plot shows weights of genes in P. putida to AcoR iModulon. Gray lines indicate thresholds of determining iModulon membership. Five genes constitute the AcoR iModulon. Gray circles identify genes not in the iModulon. Black circles are three neighboring genes. (FIG. 3C) Genomic structure of the AcoR iModulon. Shade shows predicted operonic structure. Genes in the iModulon are acoX, acoA, acoB, acoC, and bdhA. Arrows indicate three different plasmid constructs for cross-species transfer. (FIG. 3D) 2,3-BDO degradation by P. putida. Formation of acetoin was negligible. Shaded boxes represent 2,3-BDO and acetoin in the culture medium. Circles show cell density. Dots indicate individual data points. Data are presented as mean values + / − SD. Error bars indicate SD of the three biological replicates. (FIG. 3E) 2,3-BDO and acetoin degradation by E. coli carrying empty plasmid or one of the three constructs. 2,3-BDO was added at the start of the culture and the remaining amount and acetoin formation was measured. Shaded boxes represent 2,3-BDO and acetoin in the culture medium. Circles show cell density. Dots indicate individual data points. Data are presented as mean values + / − SD. Error bars indicate SD of the three biological replicates. FIGS. 4A-4H: Adaptive laboratory evolution improves functionality of Pseudomonas aeruginosa MdcR iModulon in E. coli. (FIG. 4A) Malonate catabolic pathway in P. aeruginosa. (FIG. 4B) Genetic structure of the malonate catabolic operon of P. aeruginosa cloned in a heterologous expression plasmid, pMdcR_iM. Genes other than PA0207 and mdcR constitute MdcR iModulon. (FIG. 4C) Malonate utilization of P. aeruginosa, E. coli carrying empty plasmid, and pMdcR iM. Cells were incubated for up to 72 hr in M9 malonate (2 g / l) media. Circles and diamonds show cell density and malonate concentration in culture, respectively. Shaded lines represent P. aeruginosa, E. coli carrying empty plasmid, and pMdcR iM plasmid, respectively. Data are presented as mean values + / − SD. Error bars indicate SD of three replicated cultures. (FIG. 4D) Growth rates of E. coli Att Dkt: 114198-5960 carrying the MdcR iModulon over the course of evolution. Dashed lines are moving averages of three individual ALE lineages. Growth rates for each ALE lineage are shaded differently. (FIG. 4E) Malonate utilization and growth of three evolved populations. Circles represent cell density, with the solid circles being extracellular malonate concentration. Measurement for each ALE lineage are shaded differently. Data are presented as mean values of two replicated cultures. (FIG. 4F) Growth rates of clones isolated from malonate- evolved populations in M9 malonate medium. Data are presented as mean values of two replicated cultures. Dots show individual data points. Strain names are given as AX.IY. X is the ALE lineage number and Y is an arbitrary identifying number for the clonal isolate from the same ALE lineage. (FIG. 4G) Adaptive mutations in the ALE endpoint clones that are not present in the parental strain. fs, frameshift mutation. Heatmap shows allele frequencies. (FIG. 4H) Plasmid to chromosome copy number ratio (P / C ratio) and expression level of mdcA of unevolved parent strain and evolved clones. Shaded boxes represent P / C ratio and mdcA expression level, respectively. Data are presented as mean values two biologically replicated cultures. Dots and are individual data points each of which is composed of two technical duplicates. FIGS. 5A-5D: Additional Pseudomonas putida aromatic compound degradation pathway. (FIG. 5A) Hydroxycinnamates (HCA), vanillate (VanR), protocaechuic acid (PCA), and benzoate (BenR) sub-iModulons in Pseudomonas putida is shown. (FIG. 5B) Protocatechuic acid (PCA) utilization of P. aeruginosa, E. coli carrying empty plasmid, and PCA sub-iM is shown. Cells were incubated for up to 30 hr in PCA (2 g / l) media. Empty and filled circles show cell density and PCA concentration in culture, respectively. Shaded lines represent P. aeruginosa, E. coli carrying empty plasmid, and PCA sub-iM plasmid, respectively. (FIG. 5C) Vanillic acid (VA) utilization of P. aeruginosa, E. coli carrying PCA sub-iM, and PCA sub-iM + VanR sub-iM is shown. Cells were incubated for up to 72 hr in VA (2 g / l) media. Empty and filled circles show cell density and VA concentration in culture, respectively. Shaded lines represent P. aeruginosa, E. coli carrying PCA sub-iM, and PCA sub-iM + VanR sub-iM, respectively. (FIG. 5D) Hydroxycinnamates utilization (HCA), tested with 40coumaric acid (4CA), of P. aeruginosa, E. coli carrying PCA sub-iM, and PCA sub-iM + HCA sub-iM is shown. Cells were incubated for up to 54 hr in 4AC (2 g / l) media. Empty and filled circles show cell density and 4AC concentration in culture, respectively. Shaded lines represent P. aeruginosa, E. coli carrying PCA sub-iM plasmid, and PCA sub-iM Att Dkt: 114198-5960 + HCA sub-iM, respectively. Data are presented as mean values + / − SD. Error bars indicate SD of three replicated cultures. FIG. 6: is an illustration of an example system for detecting, in accordance with implementations. FIG. 7: is an illustration of an example method for detecting and sequencing iModulons, in accordance with implementations. FIG. 8: is a block diagram illustrating an architecture for a computer system that can be employed to implement elements of the systems and methods described and illustrated herein, including, for example, the system depicted in FIG. 6 and the method depicted in FIG. 7. FIGS. 9A – 9E: Reconfiguration and repair of branched-chain amino acids (BCAA) biosynthetic iModulons in E. coli K-12. (FIG. 9A) Scatter plots show iModulon weights of genes contained in isoleucine and leucine biosynthetic iModulons. Horizontal gray lines indicate thresholds for determining iModulon membership. (FIG. 9B) The BCAA biosynthetic pathway in E. coli MG1655. ilvG carries a frameshift mutation inactivating the gene. Enzymes from Isoleucine and Leucine (i.e., IIvC, LeuABCD)_ iModulons are shown. MOB, 3-methyl-2-oxobutanoate. MOP, 4-methyl-2-oxopentanoate. HS, homoserine. (FIG. 9C) Activities of isoleucine and leucine iModulons in M9 defined medium (dark shade) or LB medium (light shade). The PRECISE 1K dataset is centered on M9 thus the iModulon activities are near zero under these conditions, but deactivated in LB. Box limits, whiskers, and center lines indicate 1st and 3rd quartiles, 10 and 90 percentiles, and median of the distribution, respectively. (FIG. 9D) Genomic landscape of genes in the BCAA iModulons in the MG1655 strain and their reconfigured structure (RECON). Transcription start sites (TSSs) and transcription termination sites (TTSs) guide the border of genetic rearrangement. Genes from Isoleucine and Leucine iModulons are colored red and blue, respectively. Shades indicate genetic rearrangements. (FIG. 9E) Growth rates of E. coli MG1655 (gray bars), BCAA biosynthesis knock-out (ΔBCAA, orange bars), and BCAA reconfigured (RECON, blue bars) strains in M9 glucose, M9 glucose supplemented with all three BCAA (ILV), isoleucine (I), leucine (L), or valine (V). Data are presented as mean values + / − SD. Error bars indicate SD of five biological replicates. Circles show individual data points. Att Dkt: 114198-5960 FIG. 10: Activity of AmpC iModulon in response to antibiotics treatment. Data are presented as mean values + / − SD. Error bars indicate SD of the sample set (when n>2). Circles represent individual transcriptome samples. FIG. 11: Ampicillin disc diffusion assay of E. coli carrying empty plasmid, P. aeruginosa ampC, or AmpC iModulon. Numbers indicate concentration of ampicillin added to the paper disk. Disks with inhibitory halos are marked with boxes. FIG. 12: MdcR iModulon weights of genes in P. aeruginosa. Nine genes constitute the MdcR iModulon. Gray circles below the weight threshold identify genes not in the iModulon. Gray lines indicate iModulon weight threshold. DETAILED DESCRIPTION OF THE DISCLOSURE Definitions Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All nucleotide sequences provided herein are presented in the 5′ to 3′ direction. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, particular, non-limiting exemplary methods, devices, and materials are now described. All technical and patent publications cited herein are incorporated herein by reference in their entirety. In some respect, the reference to a technical publication is identified by an author name and date, or an Arabic number. The full bibliographic citation for these technical publications is found immediately preceding the claims. Nothing herein is to be construed as an admission that the disclosure is not entitled to antedate such disclosure by virtue of prior disclosure. The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of cell culture, tissue culture, immunology, molecular biology, microbiology, cell biology and recombinant DNA, which are within the skill of the art. See, e.g., Green and Sambrook eds. (2012) Molecular Cloning: A Laboratory Manual, 4thedition; the series Ausubel et al. eds. (2015) Current Protocols in Molecular Biology; the series Methods in Enzymology (Academic Press, Inc., N.Y.); MacPherson et al. (2015) PCR 1: A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; McPherson et al. (2006) PCR: The Basics (Garland Science); Harlow and Lane eds. (1999) Antibodies, A Laboratory Manual; Greenfield ed. (2014) Antibodies, A Laboratory Manual; Freshney (2010) Culture of Animal Cells: A Manual of Basic Technique, Att Dkt: 114198-5960 6thedition; Gait ed. (1984) Oligonucleotide Synthesis; U.S. Pat. No. 4,683,195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Herdewijn ed. (2005) Oligonucleotide Synthesis: Methods and Applications; Hames and Higgins eds. (1984) Transcription and Translation; Buzdin and Lukyanov ed. (2007) Nucleic Acids Hybridization: Modern Applications; Immobilized Cells and Enzymes (IRL Press (1986)); Grandi ed. (2007) In Vitro Transcription and Translation Protocols, 2ndedition; Guisan ed. (2006) Immobilization of Enzymes and Cells; Perbal (1988) A Practical Guide to Molecular Cloning, 2ndedition; Miller and Calos eds, (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London); Lundblad and Macdonald eds. (2010) Handbook of Biochemistry and Molecular Biology, 4thedition; and Herzenberg et al. eds (1996) Weir's Handbook of Experimental Immunology, 5thedition. All numerical designations, e.g., pH, temperature, time, concentration, and molecular weight, including ranges, are approximations which are varied (+) or (−) by increments of 1.0 or 0.1, as appropriate or alternatively by a variation of + / − 15%, or alternatively 10% or alternatively 5% or alternatively 2%. It is to be understood, although not always explicitly stated, that all numerical designations are preceded by the term “about”. It also is to be understood, although not always explicitly stated, that the reagents described herein are merely exemplary and that equivalents of such are known in the art. As used in the specification and claims, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a polypeptide” includes a plurality of polypeptides, including mixtures thereof. As used herein, the term “comprising” is intended to mean that the compositions and methods include the recited elements, but do not exclude others. “Consisting essentially of” when used to define compositions and methods, shall mean excluding other elements of any essential significance to the combination for the intended use. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers, such as phosphate buffered saline, preservatives, and the like. “Consisting of” shall mean excluding more than trace elements of other ingredients and substantial method steps for administering the compositions disclosed herein. Embodiments defined by each of these transition terms are within the scope of this disclosure. Att Dkt: 114198-5960 As used herein, the term “optional” or “optionally” means that the subsequently described circumstance may or may not occur, so that the description includes instances where the circumstance occurs and instances where it does not. As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”). As used herein, the term “about” is used to indicate that a value includes the standard deviation of error for the device or method being employed to determine the value. The term “about” when used before a numerical designation, e.g., temperature, time, amount, and concentration, including range, indicates approximations which may vary by (+) or (–) 15%, 10%, 5%, 3%, 2%, or 1 %. The term “substantially” or “essentially” means nearly totally or completely, for instance, 95% or greater of some given quantity. In some embodiments, “substantially” or “essentially” means 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. As used herein, the term “host cell” refers not only to the particular subject cell but to the progeny or potential progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term as used herein. The host cell can be a prokaryotic or a eukaryotic cell. As used herein, the term “substantially homogeneous population of cells” refers to a plurality of cells that are of the nearly totally or completely the same kind, e.g., at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99%, or 100% phenotypically similar or containing the exogenous iModulon as disclosed herein. As used herein, the term “exogenous biochemical pathway” refers to a biochemical pathway, components of biochemical pathways, and the like that originates from a taxonomy differing from the host cell. In some embodiments, origination refers to the original source of nucleotide sequences. As used herein, the term “adaptive laboratory evolution” refers to the process of culturing cells in artificial conditions to improve one or more cellular functions through the imitation of natural evolution. Adaptive laboratory evolution is a technique well known to those skilled in the art. Att Dkt: 114198-5960 In some embodiments, a first sequence (nucleic acid sequence or amino acid) is compared to a second sequence, and the identity percentage between the two sequences can be calculated. In further embodiments, the first sequence can be referred to herein as an equivalent and the second sequence can be referred to herein as a reference sequence. In yet further embodiments, the identity percentage is calculated based on the full-length sequence of the first sequence. In other embodiments, the identity percentage is calculated based on the full-length sequence of the second sequence. The term “protein”, “peptide” and “polypeptide” are used interchangeably and in their broadest sense to refer to a compound of two or more subunit amino acids, amino acid analogs or peptidomimetics. The subunits may be linked by peptide bonds. In another embodiment, the subunit may be linked by other bonds, e.g., ester, ether, etc. A protein or peptide must contain at least two amino acids and no limitation is placed on the maximum number of amino acids which may comprise a protein's or peptide's sequence. As used herein the term “amino acid” refers to either natural and / or unnatural or synthetic amino acids, including glycine and both the D and L optical isomers, amino acid analogs and peptidomimetics. The terms “polynucleotide” and “oligonucleotide” are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides or analogs thereof. Polynucleotides can have any three-dimensional structure and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: a gene or gene fragment (for example, a probe, primer, EST or SAGE tag), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, RNAi, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes and primers. A polynucleotide can comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the polynucleotide. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide can be further modified after polymerization, such as by conjugation with a labeling component. The term also refers to both double- and single-stranded molecules. Unless otherwise specified or required, any embodiment disclosed herein that is a polynucleotide encompasses both the double-stranded form and each of two complementary single-stranded forms known or predicted to make up the double-stranded form. Att Dkt: 114198-5960 A polynucleotide is composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); thymine (T); and uracil (U) for thymine when the polynucleotide is RNA. Thus, the term “polynucleotide sequence” is the alphabetical representation of a polynucleotide molecule. This alphabetical representation can be input into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching. The term “isolated” or “recombinant” as used herein with respect to nucleic acids, such as DNA or RNA, refers to molecules separated from other DNAs or RNAs, respectively that are present in the natural source of the macromolecule as well as polypeptides. The term “isolated or recombinant nucleic acid” is meant to include nucleic acid fragments which are not naturally occurring as fragments and would not be found in the natural state. The term “isolated” is also used herein to refer to polynucleotides, polypeptides and proteins that are isolated from other cellular proteins and is meant to encompass both purified and recombinant polypeptides. In other embodiments, the term “isolated or recombinant” means separated from constituents, cellular and otherwise, in which the cell, tissue, polynucleotide, peptide, polypeptide, protein, antibody or fragment(s) thereof, which are normally associated in nature. For example, an isolated cell is a cell that is separated from tissue or cells of dissimilar phenotype or genotype. An isolated polynucleotide is separated from the 3′ and 5′ contiguous nucleotides with which it is normally associated in its native or natural environment, e.g., on the chromosome. As is apparent to those of skill in the art, a non-naturally occurring polynucleotide, peptide, polypeptide, protein, antibody or fragment(s) thereof, does not require “isolation” to distinguish it from its naturally occurring counterpart. It is to be inferred without explicit recitation and unless otherwise intended, that when the present disclosure relates to a polypeptide, protein, polynucleotide or antibody, an equivalent or a biologically equivalent of such is intended within the scope of this disclosure. As used herein, the term “biological equivalent thereof” is intended to be synonymous with “equivalent thereof” when referring to a reference protein, antibody, fragment, polypeptide or nucleic acid, intends those having minimal homology while still maintaining desired structure or functionality. Unless specifically recited herein, it is contemplated that any polynucleotide, polypeptide or protein mentioned herein also includes equivalents thereof. In one aspect, an equivalent polynucleotide is one that hybridizes under stringent conditions to the polynucleotide or complement of the polynucleotide as described herein for use in the described methods. In another aspect, an equivalent antibody or antigen binding polypeptide Att Dkt: 114198-5960 intends one that binds with at least 70%, or alternatively at least 75%, or alternatively at least 80%, or alternatively at least 85%, or alternatively at least 90%, or alternatively at least 95% affinity or higher affinity to a reference antibody or antigen binding fragment. In another aspect, the equivalent thereof competes with the binding of the antibody or antigen binding fragment to its antigen tinder a competitive ELISA assay. In another aspect, an equivalent intends at least about 80% homology or identity and alternatively, at least about 85%, or alternatively at least about 90%, or alternatively at least about 95%, or alternatively 98% percent homology or identity and exhibits substantially equivalent biological activity to the reference protein, polypeptide or nucleic acid. A polynucleotide or polynucleotide region (or a polypeptide or polypeptide region) having a certain percentage (for example, 80%, 85%, 90%, or 95%) of “sequence identity” to another sequence means that, when aligned, that percentage of bases (or amino acids) are the same in comparing the two sequences. The alignment and the percent homology or sequence identity can be determined using software programs known in the art, for example those described in Current Protocols in Molecular Biology (Ausubel et al., eds. 1987) Supplement 30, section 7.7.18, Table 7.7.1. In certain embodiments, default parameters are used for alignment. A non-limiting exemplary alignment program is BLAST, using default parameters. In particular, exemplary programs include BLASTN and BLASTP, using the following default parameters: Genetic code=standard; filter=none; strand=both; cutoff=60; expect=10; Matrix=BLOSUM62; Descriptions=50 sequences; sort by=HIGH SCORE; Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Details of these programs can be found at the following Internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST. Sequence identity and percent identity were determined by incorporating them into clustalW (available at the web address:align.genome.jp, last accessed on Mar. 7, 2011. “Homology” or “identity” or “similarity” refers to sequence similarity between two peptides or between two nucleic acid molecules. Homology can be determined by comparing a position in each sequence which may be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same base or amino acid, then the molecules are homologous at that position. A degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. An “unrelated” or “non-homologous” sequence shares less than 40% identity, or alternatively less than 25% identity, with one of the sequences of the present disclosure. Att Dkt: 114198-5960 “Homology” or “identity” or “similarity” can also refer to two nucleic acid molecules that hybridize under stringent conditions. “Hybridization” refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The hydrogen bonding may occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex may comprise two strands forming a duplex structure, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute a step in a more extensive process, such as the initiation of a PCR reaction, or the enzymatic cleavage of a polynucleotide by a ribozyme. Examples of stringent hybridization conditions include: incubation temperatures of about 25° C. to about 37° C.; hybridization buffer concentrations of about 6×SSC to about 10×SSC; formamide concentrations of about 0% to about 25%; and wash solutions from about 4×SSC to about 8×SSC. Examples of moderate hybridization conditions include: incubation temperatures of about 40° C. to about 50° C.; buffer concentrations of about 9×SSC to about 2×SSC; formamide concentrations of about 30% to about 50%; and wash solutions of about 5×SSC to about 2×SSC. Examples of high stringency conditions include: incubation temperatures of about 55° C. to about 68° C.; buffer concentrations of about 1×SSC to about 0.1×SSC; formamide concentrations of about 55% to about 75%; and wash solutions of about 1×SSC, 0.1×SSC, or deionized water. In general, hybridization incubation times are from 5 minutes to 24 hours, with 1, 2, or more washing steps, and wash incubation times are about 1, 2, or 15 minutes. SSC is 0.15 M NaCl and 15 mM citrate buffer. It is understood that equivalents of SSC using other buffer systems can be employed. As used herein, “expression” refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently being translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in an eukaryotic cell. The term “encode” as it is applied to polynucleotides refers to a polynucleotide which is said to “encode” a polypeptide if, in its native state or when manipulated by methods well known to those skilled in the art, it can be transcribed and / or translated to produce the mRNA for the polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom. Att Dkt: 114198-5960 A “composition” is intended to mean a combination of polynucleotides, active agent or another compound or composition, inert (for example, a detectable agent or label) or active, such as an adjuvant. A “pharmaceutical composition” is intended to include the combination of an active agent with a carrier, inert or active, making the composition suitable for diagnostic or therapeutic use in vitro, in vivo or ex vivo. “Pharmaceutically acceptable carriers” refers to any diluents, excipients, or carriers that may be used in the compositions disclosed herein. Pharmaceutically acceptable carriers include ion exchangers, alumina, aluminum stearate, lecithin, serum proteins, such as human serum albumin, buffer substances, such as phosphates, glycine, sorbic acid, potassium sorbate, partial glyceride mixtures of saturated vegetable fatty acids, water, salts or electrolytes, such as protamine sulfate, disodium hydrogen phosphate, potassium hydrogen phosphate, sodium chloride, zinc salts, colloidal silica, magnesium trisilicate, polyvinyl pyrrolidone, cellulose-based substances, polyethylene glycol, sodium carboxymethylcellulose, polyacrylates, waxes, polyethylene-polyoxypropylene-block polymers, polyethylene glycol, microspheres, microparticles, or nanoparticles (comprising e.g., biodegradable polymers such as Poly(Lactic Acid-co-Glycolic Acid)), and wool fat. Suitable pharmaceutical carriers are described in Remington's Pharmaceutical Sciences, Mack Publishing Company, a standard reference text in this field. They may be selected with respect to the intended form of administration, that is, oral tablets, capsules, elixirs, syrups and the like, and consistent with conventional pharmaceutical practices. “Liposomes” are microscopic vesicles consisting of concentric lipid bilayers. Structurally, liposomes range in size and shape from long tubes to spheres, with dimensions from a few hundred Angstroms to fractions of a millimeter. Vesicle-forming lipids are selected to achieve a specified degree of fluidity or rigidity of the final complex providing the lipid composition of the outer layer. These are neutral (cholesterol) or bipolar and include phospholipids, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE), phosphatidylinositol (PI), and sphingomyelin (SM) and other types of bipolar lipids including but not limited to dioleoylphosphatidylethanolamine (DOPE), with a hydrocarbon chain length in the range of 14-22, and saturated or with one or more double C═C bonds. Examples of lipids capable of producing a stable liposome, alone, or in combination with other lipid components are phospholipids, such as hydrogenated soy phosphatidylcholine (HSPC), lecithin, phosphatidylethanolamine, lysolecithin, lysophosphatidylethanol-amine, Att Dkt: 114198-5960 phosphatidylserine, phosphatidylinositol, sphingomyelin, cephalin, cardiolipin, phosphatidic acid, cerebrosides, distearoylphosphatidylethan-olamine (DSPE), dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), palmitoyloteoylphosphatidylcholine (POPC), palmitoyloleoylphosphatidylethanolamine (POPE) and dioleoylphosphatidylethanolamine 4-(N-maleimido-triethyl)cyclohexane-1- carboxylate (DOPE-mal). Additional non-phosphorous containing lipids that can become incorporated into liposomes include stearylamine, dodecylamine, hexadecylamine, isopropyl myristate, triethanolamine-lauryl sulfate, alkyl-aryl sulfate, acetyl palmitate, glycerol ricinoleate, hexadecyl stereate, amphoteric acrylic polymers, polyethyloxylated fatty acid amides, and the cationic lipids mentioned above (DDAB, DODAC, DMRIE, DMTAP, DOGS, DOTAP (DOTMA), DOSPA, DPTAP, DSTAP, DC-Chol). Negatively charged lipids include phosphatidic acid (PA), dipalmitoylphosphatidylglycerol (DPPG), dioteoylphosphatidylglycerol and (DOPG), dicetylphosphate that are able to form vesicles. Typically, liposomes can be divided into three categories based on their overall size and the nature of the lamellar structure. The three classifications, as developed by the New York Academy Sciences Meeting, “Liposomes and Their Use in Biology and Medicine,” December 1977, are multi-lamellar vesicles (MLVs), small uni-lamellar vesicles (SUVs) and large uni-lamellar vesicles (LUVs). The polynucleotides can be encapsulated in such for administration in accordance with the methods described herein. A “micelle” is an aggregate of surfactant molecules dispersed in a liquid colloid. A typical micelle in aqueous solution forms an aggregate with the hydrophilic “head” regions in contact with surrounding solvent, sequestering the hydrophobic tail regions in the micelle center. This type of micelle is known as a normal phase micelle (oil-in-water micelle). Inverse micelles have the head groups at the center with the tails extending out (water-in-oil micelle). Micelles can be used to attach a polynucleotide, polypeptide, antibody or composition described herein to facilitate efficient delivery to the target cell or tissue. Also included as a micelles are lipid nanoparticles. A “gene delivery vehicle” is defined as any molecule that can carry inserted polynucleotides into a host cell. Examples of gene delivery vehicles are liposomes, micelles biocompatible polymers, including natural polymers and synthetic polymers; lipoproteins; polypeptides; polysaccharides; lipopolysaccharides; artificial viral envelopes; metal particles; and bacteria, or viruses, such as baculovirus, adenovirus and retrovirus, bacteriophage, cosmid, plasmid, fungal vectors and other recombination vehicles typically used in the art Att Dkt: 114198-5960 which have been described for expression in a variety of eukaryotic and prokaryotic hosts, and may be used for gene therapy as well as for simple protein expression. A polynucleotide disclosed herein can be delivered to a cell or tissue using a gene delivery vehicle. “Gene delivery,” “gene transfer,” “transducing,” and the like as used herein, are terms referring to the introduction of an exogenous polynucleotide (sometimes referred to as a “transgene”) into a host cell, irrespective of the method used for the introduction. Such methods include a variety of well-known techniques such as vector-mediated gene transfer (by, e.g., viral infection / transfection, or various other protein-based or lipid-based gene delivery complexes) as well as techniques facilitating the delivery of “naked” polynucleotides (such as electroporation, “gene gun” delivery and various other techniques used for the introduction of polynucleotides). The introduced polynucleotide may be stably or transiently maintained in the host cell. Stable maintenance typically requires that the introduced polynucleotide either contains an origin of replication compatible with the host cell or integrates into a replicon of the host cell such as an extrachromosomal replicon (e.g., a plasmid) or a nuclear or mitochondrial chromosome. A number of vectors are known to be capable of mediating transfer of genes to mammalian cells, as is known in the art and described herein. A “plasmid” is an extra-chromosomal DNA molecule separate from the chromosomal DNA which is capable of replicating independently of the chromosomal DNA. In many cases, it is circular and double-stranded. Plasmids provide a mechanism for horizontal gene transfer within a population of microbes and typically provide a selective advantage under a given environmental state. Plasmids may carry genes that provide resistance to naturally occurring antibiotics in a competitive environmental niche, or alternatively the proteins produced may act as toxins under similar circumstances. “Plasmids” used in genetic engineering are called “plasmid vectors”. Many plasmids are commercially available for such uses. The gene to be replicated is inserted into copies of a plasmid containing genes that make cells resistant to particular antibiotics and a multiple cloning site (MCS, or polylinker), which is a short region containing several commonly used restriction sites allowing the easy insertion of DNA fragments at this location. Another major use of plasmids is to make large amounts of proteins. In this case, researchers grow bacteria containing a plasmid harboring the gene of interest. Just as the bacterium produces proteins to confer its antibiotic resistance, it can also be induced to produce large amounts of proteins Att Dkt: 114198-5960 from the inserted gene. This is a cheap and easy way of mass-producing a gene or the protein it then codes for. A “yeast artificial chromosome” or “YAC” refers to a vector used to clone large DNA fragments (larger than 100 kb and up to 3000 kb). It is an artificially constructed chromosome and contains the telomeric, centromeric, and replication origin sequences needed for replication and preservation in yeast cells. Built using an initial circular plasmid, they are linearized by using restriction enzymes, and then DNA ligase can add a sequence or gene of interest within the linear molecule by the use of cohesive ends. Yeast expression vectors, such as YACs, YIps (yeast integrating plasmid), and YEps (yeast episomal plasmid), are extremely useful as one can get eukaryotic protein products with posttranslational modifications as yeasts are themselves eukaryotic cells, however YACs have been found to be more unstable than BACs, producing chimeric effects. A “viral vector” is defined as a recombinantly produced virus or viral particle that comprises a polynucleotide to be delivered into a host cell, either in vivo, ex vivo or in vitro. Examples of viral vectors include retroviral vectors, adenovirus vectors, adeno-associated virus vectors, alphavirus vectors and the like. Infectious tobacco mosaic virus (TMV)-based vectors can be used to manufacturer proteins and have been reported to express Griffithsin in tobacco leaves (O'Keefe et al. (2009) Proc. Nat. Acad. Sci. USA 106(15):6099-6104). Alphavirus vectors, such as Semliki Forest virus-based vectors and Sindbis virus-based vectors, have also been developed for use in gene therapy and immunotherapy. See, Schlesinger & Dubensky (1999) Curr. Opin. Biotechnol. 5:434-439 and Ying et al. (1999) Nat. Med. 5(7):823-827. In aspects where gene transfer is mediated by a retroviral vector, a vector construct refers to the polynucleotide comprising the retroviral genome or part thereof, and a therapeutic gene. Further details as to modern methods of vectors for use in gene transfer may be found in, for example, Kotterman et al. (2015) Viral Vectors for Gene Therapy: Translational and Clinical Outlook Annual Review of Biomedical Engineering 17. As used herein, “retroviral mediated gene transfer” or “retroviral transduction” carries the same meaning and refers to the process by which a gene or nucleic acid sequences are stably transferred into the host cell by virtue of the virus entering the cell and integrating its genome into the host cell genome. The virus can enter the host cell via its normal mechanism of infection or be modified such that it binds to a different host cell surface receptor or ligand to enter the cell. As used herein, retroviral vector refers to a viral particle capable of introducing exogenous nucleic acid into a cell through a viral or viral-like entry mechanism. Att Dkt: 114198-5960 Retroviruses carry their genetic information in the form of RNA; however, once the virus infects a cell, the RNA is reverse-transcribed into the DNA form which integrates into the genomic DNA of the infected cell. The integrated DNA form is called a provirus. In aspects where gene transfer is mediated by a DNA viral vector, such as an adenovirus (Ad) or adeno-associated virus (AAV), a vector construct refers to the polynucleotide comprising the viral genome or part thereof, and a transgene. Adenoviruses (Ads) are a relatively well characterized, homogenous group of viruses, including over 50 serotypes. See, e.g., PCT International Pat. Application Publication No. WO 95 / 27071. Ads do not require integration into the host cell genome. Recombinant Ad derived vectors, particularly those that reduce the potential for recombination and generation of wild-type virus, have also been constructed. See, PCT International Pat. Application Publication Nos. WO 95 / 00655 and WO 95 / 11984, Wild-type AAV has high infectivity and specificity integrating into the host cell's genome. See, Hermonat & Muzyczka (1984) Proc. Natl. Acad. Sci. USA 81:6466-6470 and Lebkowski et al. (1988) Mol. Cell. Biol. 8:3988-3996. Vectors that contain both a promoter and a cloning site into which a polynucleotide can be operatively linked are well known in the art. Such vectors are capable of transcribing RNA in vitro or in vivo, and are commercially available from sources such as Stratagene (La Jolla, Calif.) and Promega Biotech (Madison, Wis.). In order to optimize expression and / or in vitro transcription, it may be necessary to remove, add or alter 5′ and / or 3′ untranslated portions of the clones to eliminate extra, potential inappropriate alternative translation initiation codons or other sequences that may interfere with or reduce expression, either at the level of transcription or translation. Alternatively, consensus ribosome binding sites can be inserted immediately 5′ of the start codon to enhance expression. As used herein, the term “expression” refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently being translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell. The expression level of a gene may be determined by measuring the amount of mRNA or protein in a cell or tissue sample. In one aspect, the expression level of a gene from one sample may be directly compared to the expression level of that gene from a control or reference sample. Att Dkt: 114198-5960 Gene delivery vehicles also include DNA / liposome complexes, micelles and targeted viral protein-DNA complexes. Liposomes that also comprise a targeting antibody or fragment thereof can be used in the methods disclosed herein. In addition to the delivery of polynucleotides to a cell or cell population, direct introduction of the proteins described herein to the cell or cell population can be done by the non-limiting technique of protein transfection, alternatively culturing conditions that can enhance the expression and / or promote the activity of the proteins disclosed herein are other non-limiting techniques. As used herein, the term “label” intends a directly or indirectly detectable compound or composition that is conjugated directly or indirectly to the composition to be detected, e.g., N-terminal histidine tags (N-His), magnetically active isotopes, e.g.,115Sn,117Sn and119Sn, a non-radioactive isotopes such as13C and15N, polynucleotide or protein such as an antibody so as to generate a “labeled” composition. The term also includes sequences conjugated to a polynucleotide that will provide a signal upon expression of the inserted sequences, such as green fluorescent protein (GFP) and the like. While the term “label” generally intends compositions covalently attached to the composition to be detected, in one aspect it specifically excludes naturally occurring nucleosides and amino acids that are known to fluoresce under certain conditions (e.g., temperature, pH, etc.) when positioned within the polynucleotide or protein in its native environment and generally any natural fluorescence that may be present in the composition to be detected. The label may be detectable by itself (e.g., radioisotope labels or fluorescent labels) or, in the case of an enzymatic label, may catalyze chemical alteration of a substrate compound or composition which is detectable. The labels can be suitable for small scale detection or more suitable for high-throughput screening. As such, suitable labels include, but are not limited to magnetically active isotopes, non-radioactive isotopes, radioisotopes, fluorochromes, chemiluminescent compounds, dyes, and proteins, including enzymes. The label may be simply detected or it may be quantified. A response that is simply detected generally comprises a response whose existence merely is confirmed, whereas a response that is quantified generally comprises a response having a quantifiable (e.g., numerically reportable) value such as an intensity, polarization, and / or other property. In luminescence or fluorescence assays, the detectable response may be generated directly using a luminophore or fluorophore associated with an assay component actually involved in binding, or indirectly using a luminophore or fluorophore associated with another (e.g., reporter or indicator) component. Examples of luminescent labels that produce signals include, but are not limited to bioluminescence and Att Dkt: 114198-5960 chemiluminescence. Detectable luminescence response generally comprises a change in, or an occurrence of a luminescence signal. Suitable methods and luminophores for luminescently labeling assay components are known in the art and described for example in Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6thed). Examples of luminescent probes include, but are not limited to, aequorin and luciferases. Examples of suitable fluorescent labels include, but are not limited to, fluorescein, rhodamine, tetramethylrhodamine, eosin, erythrosin, coumarin, methyl-coumarins, pyrene, Malacite green, stilbene, Lucifer Yellow, Cascade Blue™, and Texas Red. Other suitable optical dyes are described in the Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6thed.). In another aspect, the fluorescent label is functionalized to facilitate covalent attachment to a cellular component present in or on the surface of the cell or tissue such as a cell surface marker. Suitable functional groups, including, but not are limited to, isothiocyanate groups, amino groups, haloacetyl groups, maleimides, succinimidyl esters, and sulfonyl halides, all of which may be used to attach the fluorescent label to a second molecule. The choice of the functional group of the fluorescent label will depend on the site of attachment to either a linker, the agent, the marker, or the second labeling agent. “Eukaryotic cells” comprise all of the life kingdoms except monera. They can be easily distinguished through a membrane-bound nucleus. Animals, plants, fungi, and protists are eukaryotes or organisms whose cells are organized into complex structures by internal membranes and a cytoskeleton. The most characteristic membrane-bound structure is the nucleus. Unless specifically recited, the term “host” includes a eukaryotic host, including, for example, yeast, higher plant, insect and mammalian cells. Non-limiting examples of eukaryotic cells or hosts include simian, bovine, porcine, murine, rat, avian, reptilian and human. “Prokaryotic cells” that usually lack a nucleus or any other membrane-bound organelles and are divided into two domains, bacteria and archaea. In addition to chromosomal DNA, these cells can also contain genetic information in a circular loop called on episome. Bacterial cells are very small, roughly the size of an animal mitochondrion (about 1-2 μm in diameter and 10 μm long). Prokaryotic cells feature three major shapes: rod shaped, spherical, and spiral. Instead of going through elaborate replication processes like Att Dkt: 114198-5960 eukaryotes, bacterial cells divide by binary fission. Examples include but are not limited to Bacillus bacteria, E. coli bacterium, and Salmonella bacterium. As used herein, “solid phase support” or “solid support”, used interchangeably, is not limited to a specific type of support. Rather a large number of supports are available and are known to one of ordinary skill in the art. Solid phase supports include silica gels, resins, derivatized plastic films, glass beads, cotton, plastic beads, alumina gels. As used herein, “solid support” also includes synthetic antigen-presenting matrices, cells, and liposomes. A suitable solid phase support may be selected on the basis of desired end use and suitability for various protocols. For example, for peptide synthesis, solid phase support may refer to resins such as polystyrene (e.g., PAM-resin obtained from Bachem Inc., Peninsula Laboratories, etc.), POLYHIPE® resin (obtained from Aminotech, Canada), polyamide resin (obtained from Peninsula Laboratories), polystyrene resin grafted with polyethylene glycol (TentaGel®, Rapp Polymere, Tubingen, Germany) or polydimethylacrylamide resin (obtained from Milligen / Biosearch, Calif.). An example of a solid phase support include glass, polystyrene, polypropylene, polyethylene, dextran, nylon, amylases, natural and modified celluloses, polyacrylamides, gabbros, and magnetite. The nature of the carrier can be either soluble to some extent or insoluble. The support material may have virtually any possible structural configuration so long as the coupled molecule is capable of binding to a polynucleotide, polypeptide or antibody. Thus, the support configuration may be spherical, as in a bead, or cylindrical, as in the inside surface of a test tube, or the external surface of a rod. Alternatively, the surface may be flat such as a sheet, test strip, etc. or alternatively polystyrene beads. Those skilled in the art will know many other suitable carriers for binding antibody or antigen, or will be able to ascertain the same by use of routine experimentation. Ampicillin resistance denotes the ability of bacteria to withstand the effects of ampicillin, an antibiotic that inhibits bacterial cell wall synthesis. This resistance can be quantified by measuring the minimum inhibitory concentration (MIC). The MIC is determined by treating cells with varying concentrations of ampicillin and identifying the minimum concentration at which ampicillin completely inhibits the growth of the bacterium. “Independent Component Analysis” is a statistical and computational technique used in machine learning to separate a multivariate signal into its independent non-Gaussian components. The goal of ICA is to find a linear transformation of the data such that the Att Dkt: 114198-5960 transformed data is as close to being statistically independent as possible. The heart of ICA lies in the principle of statistical independence. ICA identify components within mixed signals that are statistically independent of each other. As used herein, an “iModulon” a fundamental unit of bacterial transcriptomes that have been found to represent the genetic basis for various cellular functions and are associated with particular transcriptional regulator(s). “Adaptive laboratory evolution” is an evolutionary engineering approach in artificial conditions that improves organisms through the imitation of natural evolution. See, for example Hirasawa T, Maeda T. Adaptive Laboratory Evolution of Microorganisms: Methodology and Application for Bioproduction. Microorganisms.2022 Dec 29;11(1):92. doi: 10.3390 / microorganisms11010092. PMID: 36677384; PMCID: PMC9864036 for examples of various ALE methodologies. As used herein, upregulated or downregulated expression of an iModulon gene cluster refers to increased or decreased RNA expression of the iModulon gene cluster or one or more elements thereof, respectively. As used herein, a gene of unidentified function refers to a gene having a function that has not been characterized in the preexisting art. Therefore, a skilled person would not be able to ascertain the function of a gene of unidentified function based on the preexisting art. As used herein, a substrate (e.g., substrate compatibility) refers to a material that a cell grows on. For example, the substrate may comprise a matrix of molecules used to support the growth of the cell, including nutrients to support the growth of the cell. In some embodiments, the substrate comprises a gelling agent, e.g., agar. Modes For Carrying Out the Disclosure Applicant applied ICA to find source regulatory signals in bacterial transcriptomes10–13. iModulons are fundamental units of bacterial transcriptomes that have been found to represent the genetic basis for various cellular functions and are associated with particular transcriptional regulator(s)14–16. Many identified iModulons include genes that are distantly located on the genome, genes of unknown functions, or contain accessory genes that augment the targeted cellular function. iModulons thus represent a new scale of synthetic biology to transfer naturally evolved traits across species. Att Dkt: 114198-5960 Here, iModulons were used to create new cellular functions in a new host. As shown herein, transferring iModulons is superior to using operons or single genes identified by genome annotation algorithms and rational design approaches9,17,18. In some cases, adaptive laboratory evolution (ALE) is used to enable the host to optimally use the new function under selection pressure. Thus as shown herein, cross-species iModulon transfer is a new and versatile tool for synthetic biology. Specifically, sets of transcriptionally independently modulated genes (called iModulons) as a unit of engineering that constitute a particular cellular function, were used. The iModulons were identified by independent component analysis (ICA) of large transcriptomic datasets that are publicly available. Applicant captured and refactored members (genes) of an iModulon into a single transcription unit (operon) carried by a heterologous plasmid DNA. This was sufficient to recreate function of the iModulon captured after transformation of the plasmid into E. coli. The iModulon harbors all the necessary genes supporting the function, thus transferring iModulon immediately makes the recipient cell to have the function. Engineering of the host cell or iModulon genes is not required, providing a simple and rapid method to engineer specific cellular traits in microorganisms. Overall, this method of iModulon transfer followed by adaptive laboratory evolution can achieve at least a technical solution of engineering cellular traits across multiple bacterial species. iModulons, Vectors and Isolated Host Cells Using Applicant’s methods, provided herein are compositions comprising or consisting essentially of an isolated iModulon gene cluster, wherein the gene cluster is selected from the group of polynucleotides comprising, or consisting essentially of, or consisting of: SEQ ID Nos: 1-4 or an equivalent of each thereof; SEQ ID Nos: 5-13 or an equivalent of each thereof; SEQ ID Nos: 14-18 or an equivalent of each thereof; SEQ ID Nos: 19-25 or an equivalent of each thereof; SEQ ID Nos: 26-66 or an equivalent of each thereof SEQ ID Nos: 67-76 or an equivalent of each thereof; SEQ ID Nos: 26-43 or an equivalent of each thereof; or SEQ ID Nos: 53-54, 57-61, and 64-65, or an equivalent of each thereof, that are optionally detectably labeled. The polynucleotides can be DNA, RNA or mRNA. Further provided are the complements of the polynucleotides. Also provided are the protein encoded by the isolated iModulon gene cluster as identified herein. Applicant also provides isolated polynucleotides encoding one or more of an uncharacterized gene as identified in Table 3, e.g., one or more of SEQ ID NOs: 4, 14, 20-25, Att Dkt: 114198-5960 36, 37, 39-43, 45-47, 51, 54, 57, 58, 63, 64, and / or 66, or equivalents of each thereof. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. The polynucleotides can be detectably labeled. The polynucleotides can be operatively linked to regulatory elements for the replication and / or expression of the polynucleotides, in vitro or in vivo. Non-limiting examples are described herein, e.g., promoters, enhancers, operons, etc. for expression in prokaryotic or eukaryotic cells. The polynucleotides can be incorporated into vectors, e.g., viral vectors or plasmids. These can be combined with a carrier, e.g, a pharmaceutically acceptable carrier. Also within the composition can be stabilizers and / or preservatives to aid in transport or storage of the polynucleotides, e.g., glycerol. When the polynucleotide is inserted into a suitable host cell, e.g., a prokaryotic or a eukaryotic cell and the host cell replicates, the proteins or polypeptides can be recombinantly produced. Suitable host cells will depend on the vector and can include mammalian cells, animal cells, human cells, simian cells, insect cells, yeast cells, and bacterial cells constructed using known methods. See Sambrook, et al. (1989) supra. In addition to the use of viral vector and plasmids for insertion of exogenous polynucleotides into cells, the polynucleotides can be inserted into the host cell by methods known in the art such as transformation for bacterial cells; transfection using calcium phosphate precipitation for mammalian cells; or DEAE-dextran; electroporation; or microinjection. See, Sambrook et al. (1989) supra, for methodology. Thus, this disclosure also provides a host cell, e.g. a mammalian cell, an animal cell (rat or mouse), a human cell, or a prokaryotic cell such as a bacterial cell, containing polynucleotides encoding a gene cluster. Also provided are the polypeptides encoded by the polynucleotides, e.g., one or more of SEQ ID NOs: 4, 14, 20-25, 36, 37, 39-43, 45-47, 51, 54, 57, 58, 63, 64, and / or 66. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or consisting essentially of, or yet further consisting of SEQ ID Nos: 1-4 (i.e., the VanR iModulon for vanillate conversion into PCA, see Table 3) or an equivalent of each thereof, that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or Att Dkt: 114198-5960 consisting essentially of, or yet further consisting of SEQ ID Nos: 5-13 (i.e., the MdcR iModulon for malonate catabolism, see Table 3) or an equivalent of each thereof; that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or consisting essentially of, or yet further consisting of SEQ ID Nos: 14-18 (i.e., the AcoR iModulon for 2,3-BDO cataboism, see Table 3) or an equivalent of each thereof; that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or consisting essentially of, or yet further consisting of SEQ ID Nos: 19-25 (i.e., the AmpC iModulon for ampicillin resistance, see Table 3) or an equivalent of each thereof; that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or consisting essentially of, or yet further consisting of SEQ ID Nos: 26-66 (i.e., the PcaR iModulon for hydroxycinnamates catabolism, see Table 3) or an equivalent of each thereof; that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or consisting essentially of, or yet further consisting of SEQ ID Nos: 67-76 (i.e., the BenR iModulon for benzoate catabolism, see Table 3) or an equivalent of each thereof; that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or consisting essentially of, or yet further consisting of SEQ ID Nos: 26-43 (i.e., the PCA sub- iModulon for vanillate transport and catabolism, see Table 3) or an equivalent of each Att Dkt: 114198-5960 thereof; that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. In one aspect, provided is a composition comprising an isolated iModulon gene cluster comprising or consisting essentially of a group of polynucleotides comprising, or consisting essentially of, or yet further consisting of SEQ ID Nos: 53-54, 57-61, 64-65 (i.e., the HCA sub-iModulon for hydroxycinnamates transport and catabolism, see Table 3) or an equivalent of each thereof; that are optionally detectably labeled. The polynucleotides can be DNA, RNA, or mRNA. Further provided are the complements of the polynucleotides. Vectors comprising the polynucleotides of the gene clusters are further provided, examples of which are known in the art and briefly described herein. Non-limiting examples include plasmids or viral vectors. In one aspect where more than one polynucleotide is to be expressed as a single unit, the polynucleotides can be contained within a polycistronic vector, e.g., under the control of an operon. One of skill in the art can make such polynucleotides using the information provided herein and knowledge of those of skill in the art. See Goodman and Kay (1999) J. Biological Chem. 274(52):37004-37011 and Kamashev and Rouviere-Yaniv (2000) EMBO J. 19(23):6527-6535. The disclosure further provides the polynucleotides operatively linked to a promoter of RNA transcription, as well as other regulatory sequences for replication and / or transient or stable expression of the DNA or RNA. As used herein, the term “operatively linked” means positioned in such a manner that the promoter will direct transcription of RNA off the DNA molecule. Examples of such promoters are SP6, T4 and T7. In certain embodiments, cell-specific promoters are used for cell-specific expression of the inserted polynucleotide. Vectors which contain a promoter or a promoter / enhancer, with termination codons and selectable marker sequences, as well as a cloning site into which an inserted piece of DNA can be operatively linked to that promoter are known in the art and commercially available. For general methodology and cloning strategies, see Gene Expression Technology (Goeddel ed., Academic Press, Inc. (1991)) and references cited therein and Vectors: Essential Data Series (Gacesa and Ramji, eds., John Wiley & Sons, N.Y. (1994)) which contains maps, functional properties, commercial suppliers and a reference to GenEMBL accession numbers for various suitable vectors. Expression vectors containing the polynucleotides are useful to obtain host vector systems to produce proteins and polypeptides. It is implied that these expression vectors must Att Dkt: 114198-5960 be replicable in the host organisms either as episomes or as an integral part of the chromosomal DNA. Non-limiting examples of suitable expression vectors include plasmids, yeast vectors, viral vectors and liposomes. Adenoviral vectors are particularly useful for introducing genes into tissues in vivo because of their high levels of expression and efficient transformation of cells both in vitro and in vivo. When a polynucleotide is inserted into a suitable host cell, e.g., a prokaryotic or a eukaryotic cell and the host cell replicates, the proteins or polypeptides can be recombinantly produced. Suitable host cells will depend on the vector and can include mammalian cells, animal cells, human cells, simian cells, insect cells, yeast cells, and bacterial cells constructed using known methods. See Sambrook, et al. (1989) supra. In addition to the use of viral vector and plasmids for insertion of exogenous polynucleotides into cells, the polynucleotides can be inserted into the host cell by methods known in the art such as transformation for bacterial cells; transfection using calcium phosphate precipitation for mammalian cells; or DEAE-dextran; electroporation; or microinjection. See, Sambrook et al. (1989) supra, for methodology. Thus, this disclosure also provides a host cell, e.g. a mammalian cell, an animal cell (rat or mouse), a human cell, or a prokaryotic cell such as a bacterial cell, containing polynucleotides encoding a gene cluster. This disclosure also provides genetically modified host cells that contain and / or express the polynucleotides of this disclosure. The genetically modified cells can be produced by insertion of upstream regulatory sequences such as promoters or gene activators (see, U.S. Patent No. 5,733,761), using methods as described herein or known in the art. The polynucleotides can be conjugated to a detectable marker, e.g., an enzymatic label or a radioisotope for detection of nucleic acid and / or expression of the gene in a cell. A wide variety of appropriate detectable markers are known in the art, including fluorescent, radioactive, enzymatic or other ligands, such as avidin / biotin, which are capable of giving a detectable signal. In one aspect, one will likely desire to employ a fluorescent label or an enzyme tag, such as urease, alkaline phosphatase or peroxidase, instead of radioactive or other environmentally undesirable reagents. In the case of enzyme tags, calorimetric indicator substrates can be employed to provide a means visible to the human eye or spectrophotometrically, to identify specific hybridization with complementary nucleic acid-containing samples. Thus, this disclosure further provides a method for detecting a single-stranded polynucleotide or its complement, by contacting target single-stranded polynucleotide with a labeled, single-stranded polynucleotide (a probe) which is a portion of Att Dkt: 114198-5960 the polynucleotide of this disclosure under conditions permitting hybridization (preferably moderately stringent hybridization conditions) of complementary single-stranded polynucleotides, or more preferably, under highly stringent hybridization conditions. Hybridized polynucleotide pairs are separated from un-hybridized, single-stranded polynucleotides. The hybridized polynucleotide pairs are detected using methods known to those of skill in the art and set forth, for example, in Sambrook et al. (1989) supra. The polynucleotide embodied in this disclosure can be obtained using chemical synthesis, recombinant cloning methods, PCR, or any combination thereof. Methods of chemical polynucleotide synthesis are known in the art and need not be described in detail herein. One of skill in the art can use the sequence data provided herein to obtain a desired polynucleotide by employing a DNA synthesizer or ordering from a commercial service. The polynucleotides of this disclosure can be isolated or replicated using PCR. PCR technology is the subject matter of U.S. Patent Nos. 4,683,195; 4,800,159; 4,754,065; and 4,683,202 and described in PCR: The Polymerase Chain Reaction (Mullis et al. eds., Birkhauser Press, Boston (1994)) or MacPherson et al. (1991) and (1995) supra, and references cited therein. Alternatively, one of skill in the art can use the sequences provided herein and a commercial DNA synthesizer to replicate the DNA. Accordingly, this disclosure also provides a process for obtaining the polynucleotides of this disclosure by providing the linear sequence of the polynucleotide, nucleotides, appropriate primer molecules, chemicals such as enzymes and instructions for their replication and chemically replicating or linking the nucleotides in the proper orientation to obtain the polynucleotides. In a separate embodiment, these polynucleotides are further isolated. Still further, one of skill in the art can insert the polynucleotide into a suitable replication vector and insert the vector into a suitable host cell (prokaryotic or eukaryotic) for replication and amplification. The DNA so amplified can be isolated from the cell by methods known to those of skill in the art. A process for obtaining polynucleotides by this method is further provided herein as well as the polynucleotides so obtained. RNA can be obtained by first inserting a DNA polynucleotide into a suitable host cell. The DNA can be delivered by any appropriate method, e.g., by the use of an appropriate gene delivery vehicle (e.g., liposome, plasmid or vector) or by electroporation. When the cell replicates and the DNA is transcribed into RNA; the RNA can then be isolated using methods known to those of skill in the art, for example, as set forth in Sambrook et al. (1989) supra. For instance, mRNA can be isolated using various lytic enzymes or chemical solutions Att Dkt: 114198-5960 according to the procedures set forth in Sambrook et al. (1989) supra, or extracted by nucleic-acid-binding resins following the accompanying instructions provided by manufactures. Polynucleotides exhibiting sequence complementarity or homology to a polynucleotide of this disclosure are useful as hybridization probes or as an equivalent of the specific polynucleotides identified herein. Since the full coding sequence of the transcript is known, any portion of this sequence or homologous sequences, can be used in the methods of this disclosure. It is known in the art that a “perfectly matched” probe is not needed for a specific hybridization. Minor changes in probe sequence achieved by substitution, deletion or insertion of a small number of bases do not affect the hybridization specificity. In general, as much as 20% base-pair mismatch (when optimally aligned) can be tolerated. Preferably, a probe useful for detecting the aforementioned mRNA is at least about 80% identical to the homologous region. More preferably, the probe is 85% identical to the corresponding gene sequence after alignment of the homologous region; even more preferably, it exhibits 90% identity. These probes can be used in radioassays (e.g. Southern and Northern blot analysis) to detect, prognose, diagnose or monitor various cells or tissues containing these cells. The probes also can be attached to a solid support or an array such as a chip for use in high throughput screening assays for the detection of expression of the gene corresponding a polynucleotide of this disclosure. Accordingly, this disclosure also provides a probe comprising or corresponding to a polynucleotide of this disclosure, or its equivalent, or its complement, or a fragment thereof, attached to a solid support for use in high throughput screens. The total size of fragment, as well as the size of the complementary stretches, will depend on the intended use or application of the particular nucleic acid segment. Smaller fragments will generally find use in hybridization embodiments, wherein the length of the complementary region may be varied, such as between at least 5 to 10 to about 100 nucleotides, or even full length according to the complementary sequences one wishes to detect. Nucleotide probes having complementary sequences over stretches greater than 5 to 10 nucleotides in length are generally preferred, so as to increase stability and selectivity of Att Dkt: 114198-5960 the hybrid, and thereby improving the specificity of particular hybrid molecules obtained. More preferably, one can design polynucleotides having gene-complementary stretches of 10 or more or more than 50 nucleotides in length, or even longer where desired. Such fragments may be readily prepared by, for example, directly synthesizing the fragment by chemical means, by application of nucleic acid reproduction technology, such as the PCR technology with two priming oligonucleotides as described in U.S. Patent No. 4,603,102 or by introducing selected sequences into recombinant vectors for recombinant production. In one aspect, a probe is about 50-75 or more alternatively, 50-100, nucleotides in length. Examples of probes are provided herein. The polynucleotides of the present disclosure can serve as primers for the detection of genes or gene transcripts that are expressed in cells described herein. In this context, amplification means any method employing a primer-dependent polymerase capable of replicating a target sequence with reasonable fidelity. Amplification may be carried out by natural or recombinant DNA-polymerases such as T7 DNA polymerase, Klenow fragment of E. coli DNA polymerase, and reverse transcriptase. For illustration purposes only, a primer is the same length as that identified for probes. One method to amplify polynucleotides is PCR and kits for PCR amplification are commercially available. After amplification, the resulting DNA fragments can be detected by any appropriate method known in the art, e.g., by agarose gel electrophoresis followed by visualization with ethidium bromide staining and ultraviolet illumination. Methods for administering an effective amount of a gene delivery vector or vehicle to a cell have been developed and are known to those skilled in the art and described herein. Methods for detecting gene expression in a cell are known in the art and include techniques such as in hybridization to DNA microarrays, in situ hybridization, PCR, RNase protection assays and Northern blot analysis and functional assays as described herein. Such methods are useful to detect and quantify expression of the gene in a cell. Alternatively, expression of the polypeptide can be detected by various methods. In particular it is useful to prepare polyclonal or monoclonal antibodies that are specifically reactive with the target polypeptide. Such antibodies are useful for visualizing cells that express the polypeptide using techniques such as immunohistology, ELISA, and Western blotting. These techniques can be used to determine expression level of the expressed polynucleotide. Att Dkt: 114198-5960 Also provided are the polypeptides encoded by the gene clusters as described herein, that can be optionally detectably labeled. This disclosure also provides an isolated host cell or population of cells comprising or consisting essentially of a gene cluster, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 1-4. In other embodiments, provided herein is an isolated host cell or population of cells, comprising a gene cluster as provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 5-13. In other embodiments, an isolated host cell or population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 14-18. In other embodiments, an isolated host cell or population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 19-25. In other embodiments, an isolated host cell or population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 26- 66. In other embodiments, an isolated host cell or population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 67-76. In other embodiments, an isolated host cell or population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 26-43. In other embodiments, an isolated host cell or population of cells comprising a gene cluster is provided herein, wherein the gene cluster comprises the polynucleotides of SEQ ID Nos: 53-54, 57-61, and 64-65. In one aspect, the population of cells are substantially homogenous or a clonal population of the cells. The cells can be eukaryotic or prokaryotic, examples of such include E. coli, P. putida, Corynebacterium glutamicum, and genetically minimized species such as Vibrio natriegens and JCVI_Syn3A (see https: / / elifesciences.org / articles / 36842, incorporated herein by reference). The cells are commercially available from vendors. Compositions Compositions are further provided. The compositions comprise a carrier and one or more of a gene cluster of the disclosure, the polypeptides encoded by the gene clusters, vectors of the disclosure, an isolated host cell or population of cells of the disclosure. The carriers can be one or more of a solid support, a carrier or a pharmaceutically acceptable carrier. In one aspect, the compositions further comprise a preservative and / or a cryopreservative, e.g., glycerol, or medium containing 10% glycerol, or medium containing Att Dkt: 114198-5960 10% dimethylsulfoxide, or 50% cell-conditioned medium with 50% fresh medium with 10% glycerol or 10% DMSO. Methods of Preparation In another aspect, a method for preparing an isolated cell or population of cells with an exogenous biochemical pathway is provided herein, comprising, or consisting essentially of, or yet further consisting of introducing an iModulon gene cluster selected from the group of: SEQ ID Nos: 1-4 or an equivalent of each thereof; SEQ ID Nos: 5-13 or an equivalent of each thereof; SEQ ID Nos: 14-18 or an equivalent of each thereof; SEQ ID Nos: 19-25 or an equivalent of each thereof; SEQ ID Nos: 26-43 or an equivalent of each thereof; SEQ ID Nos: 1-4 and 26-66 or an equivalent of each thereof; SEQ ID Nos: 26-43 and 67-76 or an equivalent of each thereof; SEQ ID Nos: 1-4 and 26-43 or an equivalent of each thereof; or SEQ ID Nos: 26-43, 53-54, 57-61, and 64-65 or an equivalent of each thereof, into the host cell or population of cells such that the gene cluster is expressed in the host cell or population of cells, thereby providing the host cell or population of cells with the exogenous biochemical pathway. The cells can be eukaryotic or prokaryotic, examples of such are provided herein. In a further aspect, the method further comprises optimizing the biochemical pathway through adaptive laboratory evolution. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. Also provided herein is a method for preparing an isolated cell or population of cells with exogenously obtained vanillate transport and catabolic pathway of Pseudomonas putida , comprising, or consisting essentially of, or yet further consisting of introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 into the host cell or the population of cells, wherein the host cell or population of cells expresses the polynucleotides of the gene cluster and gains the vanillate transport and catabolic pathway of Pseudomonas putida. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. In another embodiment, method for preparing an isolated cell or population of cells with exogenously obtained malonate transport and catabolic pathway of Pseudomonas aeruginosa is provided herein, comprising, or consisting essentially of, or yet further consisting of introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 5-13 into the host cell or the population of cells, wherein the host cell expresses the Att Dkt: 114198-5960 polynucleotides of the gene cluster gains the malonate transport and catabolic pathway of Pseudomonas aeruginosa. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. Further provided herein is a method for preparing an isolated cell or population of cells with exogenously obtained 2,3-butanediol catabolic pathway of Pseudomonas putida, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 14-18 into the host cell or population of cells, wherein the host cell or the population of cells expresses the gene cluster and gains the 2,3-butanediol catabolic pathway of Pseudomonas putida. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained ampicillin resistance property of Pseudomonas aeruginosa is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 19-25, wherein the host cell or population of cells expresses the gene cluster gains the ampicillin resistance property of Pseudomonas aeruginosa. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained protocatechuate transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43, wherein the host cell or population of cells expresses the gene cluster gains the protocatechuate transport and catabolic pathway of Pseudomonas putida. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained hydroxycinnamates transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the Att Dkt: 114198-5960 polynucleotides of SEQ ID Nos: 1-4 and 26-66, wherein the host cell or population of cells expresses the gene cluster gains the hydroxycinnamates transport and catabolic pathway of Pseudomonas putida. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained benzoate transport and catabolic pathway of Pseudomonas putuda is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43 and 67-76, wherein the host cell or population of cells expresses the gene cluster gains the benzoate transport and catabolic pathway of Pseudomonas putida. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained vanillate transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 and 26-43, wherein the host cell or population of cells expresses the gene cluster gains the vanillate transport and catabolic pathway of Pseudomonas putida. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. In another embodiment, a method for preparing an isolated cell or population of cells with an exogenously obtained hydroxycinnamates transport and catabolic pathway of Pseudomonas putida is provided herein, comprising, or consisting essentially of, or yet further consisting of, introducing a gene cluster of a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43, 53-54, 57-61, and 64-65, wherein the host cell or population of cells expresses the gene cluster gains the hydroxycinnamates transport and catabolic pathway of Pseudomonas putida. In another aspect, the method further comprises culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. Methods to culture the cells are known in the art and described herein. Methods for Identifying iModulons Att Dkt: 114198-5960 Also provided herein is a method comprising, or consisting essentially of, or consisting of determining one or more iModulons of an organism or cell by applying independent component analysis on gene expression data of the organism or cell, the one or more iModulons comprising one or more genes of an unidentified function; creating a strain of a phenotypic expression from the one or more genes of the unidentified function by refactoring the one or more genes of the unidentified function into a gene transfer vehicle; optionally applying an adaptive laboratory evolution algorithm to the strain of the phenotypic expression to create an optimized strain of phenotypic expression for performing a function; and sequencing the optimized strain of phenotypic expression and the gene transfer vehicle to determine mutations for the optimized strain of phenotypic expression. In one aspect, the refactoring of the one or more genes of the unidentified function into the gene transfer vehicle comprises, or consists essentially of, or consisting of refactoring the one or more genes of the unidentified function into a bacterial artificial chromosome. In another aspect, the determining of the one or more iModulons of the organism or cell comprises: creating an expression profile of genes of the organism; and determining the one or more iModulons of the organism or cell by applying independent component analysis on the expression profile of the genes of the organism or cell. The following experimental examples are provided to illustrate the clauses and aspects of this disclosure. Examples Example 1 MATERIAL AND METHODS Bacterial strains and culture conditions E. coli strain K-12 substrain MG1655 is used as a recipient of iModulon. Cells were grown in LB medium (Novagen, 71753) or M9 defined medium (47.75 mM Na2HPO4, 22.04 mM KH2PO4, 8.56 mM NaCl, 18.70 mM NH4Cl, 2 mM MgSO4, 0.1 mM CaCl2, and trace elements). Trace elements were prepared in 2000× concentrated solution (100 mM FeCl3, 9.54 mM ZnCl2, 8.41 mM CoCl2, 8.27 mM Na2MoO4, 0.75 mM CaCl2, 0.91 mM CuCl2, and 0.5 mM H3BO3in 3.7% (w / w) hydrochloric acid solution). Branched-chain amino acids were supplemented with a final combined concentration of 0.15 mM. To construct the ΔBCAA strain (MG1655 ΔilvLGMEDAYC::ilvY, ΔleuLABCD, thrL-119::sacB-cat), sequential lambda-recombinations were performed51. First, ilvLGMEDAYC was replaced with a Att Dkt: 114198-5960 kanamycin resistance cassette flanked by two flippase recognition sequences. The cassette was removed by flippase-mediated excision and a DNA cassette containing ilvY-pheS*-nptII was introduced as replacing the flippase recombination scar. Then pheS*-nptII dual- selectable marker was again removed by the lambda recombination. A successful recombinant was selected by p-chloro-phenylalanine counter selection52on LB-agar plate containing 2 mM p-chloro-phenylalanine. The genomic locus containing leuLABCD was removed by introducing the kanamycin cassette followed by flippase recombination. Lastly, a sacB-cat dual selection marker was inserted upstream of thrL serving as a landing pad for REXER. Yeast was grown in YPDA (10 g / l yeast extract, 20 g / l peptone, 30 mg / l adenine sulfate, and 20 g / l dextrose) or synthetic His dropout medium (1.7 g / l yeast nitrogen base without amino acid and ammonium sulfate, 1.9 g / l yeast synthetic drop-out medium supplements without histidine, 5 g / l ammonium sulfate, 20 g / l agar, and 20 g / l dextrose) at 30°C. MdcR and AmpC iModulons were sourced from Pseudomonas aeruginosa strain K273316. AcoR and VanR iModulons were amplified from Pseudomonas putida strain KT2440. Fermentation experiment was done in 15 ml liquid culture contained in a 30-ml test tube. Cells were incubated at 37°C on a heat-block and stirred at 600 rpm with a cross- shaped magnetic stir bar. To induce heterologous iModulons, (IPTG) was added to cultures to 0.1 or 1 mM of final concentration. All reagents were purchased from Sigma-Aldrich unless stated otherwise. Investigation of target iModulons and transfer design Target iModulons for transfer were selected from a catalog of iModulons available through iModulonDB (https: / / imodulondb.org / )12. Initially, iModulons of P. aeruginosa and P. putida were inspected by their predicted functional annotations. iModulons associated with functions that are uncharacterized, already present in E. coli (e.g., translation), or challenging to experimentally screen (e.g., quorum sensing) were excluded. Subsequently, the functions of iModulons were further investigated by a comprehensive examination of individual gene functions using gene annotation databases (Biocyc18and Pseudomonas Genome Database53) and a literature search. The functional annotation of an iModulon was revised in this step if needed. For instance, the P. aeruginosa MdcR iModulon was initially annotated as malonate biosynthetic process, but Applicant corrected its function to malonate catabolism as it encodes the malonate catabolic enzyme complex21. Finally, iModulon activities across diverse experimental conditions were examined using the built-in graph function of iModulonDB in the individual iModulon page, which can also be obtained by Att Dkt: 114198-5960 downloading two tables, namely Experimental Conditions and iModulon Activity. In this step, Applicant took a rational approach to test if changes in iModulon activity under specific conditions were consistent with established biological knowledge. For example, activities of BCAA biosynthetic iModulons and AmpC iModulon are induced in the absence of BCAA and the presence of beta-lactam in the media (FIG. 9C and FIG. 10), affirming the functional annotations. The strategy for iModulon transfer was established by addressing the genetic context of iModulons. Firstly, gene members of iModulons were extracted by selecting genes with their absolute iModulon weight exceeding the threshold (FIG. 1B, FIG. 2A, FIG. 3B, FIG. 9A, and FIG. 12) (Gene Table in an iModulon page). Next, the genetic organization (location, orientation, and operonic structure) was assessed through Biocyc18and Pseudomonas Genome Database53. If genes were found in a single operon, the entire operon was directly amplified and cloned. If genes were located in multiple genomic loci, they were refactored into an operon, while preserving the genetic order to ensure an optimal expression level as demonstrated previously25,26. DNA assembly and cloning for capturing iModulon The bacterial artificial chromosome (BAC) encoding BCAA biosynthetic iModulon was constructed by assembling PCR amplified pCAP_BAC (Addgene plasmid #120229)54fragments, nptII promoter, pheS-nptII, thr, ilvG, ilvMEDA, ilvC, and leu loci using transformation-associated recombination (TAR) cloning55in Saccharomyces cerevisiae strain VL6-48 (ATCC MYA-3666; MATα his3-Δ200 trp1-Δ1 ura3−52 lys2 ade2−1 met14 cir0). To perform TAR cloning, yeast VL6-48 was grown in 50 ml YPDA broth at 30 °C until A600of 0.4. Cells were collected by centrifuging at 3000 ×g for 5 min. Then the cell pellet was washed with 10 ml of buffer 1 (100 mM LiAc, 10 mM Tris-HCl (pH8.0), and 1 mM EDTA (pH8.0)) and resuspended with 0.5 ml of buffer 1. Cell suspension (100 µl) was combined with 100 ng each of the DNA fragments, 100 µg of denatured salmon sperm DNA (Sigma- Aldrich, D9156), and 600 µl of buffer 2 (100 mM LiAc, 10 mM Tris-HCl (pH8.0), 1 mM EDTA (pH8.0), 40% (w / v) PEG-3350 (Sigma-Aldrich, P4338)). The mixture was mixed thoroughly and incubated at 30 °C for 30 min. Then, the mixture was incubated at 42 °C for 15 min after addition of 70 µl of DMSO. Cell was harvested by centrifugation at 16000 ×g for 1min after incubation on ice for 1 min. Liquid supernatant was discarded and cell pellet was resuspended with 1 ml of YPDA broth. After incubation at 30 °C for 2 hr, cell was harvested again with centrifugation at 16000 ×g for 1min and resuspended with 100 µl of Att Dkt: 114198-5960 sterile water. Successful clones were screened on solid synthetic His dropout medium. Spacer array targeting both the landing pad and the BAC was constructed by annealing of two primers RX4_1_F and RX_4_1_R followed by primer extension using RX4_2_F and RX4_2_R. The spacer array was cloned in a pUC plasmid containing crRNA leader sequence. MdcR and AcoR iModulons were PCR amplified from P. aeruginosa and P. putida genomic DNA. DNA fragments had 15-20 nt homology to pTrcHis2A (Invitrogen, V36520) plasmid and cloned into the linearized plasmid backbone using In-Fusion Cloning Kit (Takara Bio, 638948). Two converging operons encoding VanR iModulons were PCR amplified separately from P. putida genomic DNA and cloned into pTrcHis2A plasmid backbone in three-fragments ligation reaction using In-Fusion Cloning Kit resulting in pVanR_iM. AmpC iModulon was PCR amplified as 6 different fragments from P. aeruginosa genomic DNA. Together with the Trc promoter fragment amplified from pTrcHis2A plasmid and pCAP_BAC, entire DNA fragments were assembled using TAR cloning. Assembled BAC was extracted using Gentra Puregene Yeast / Bact. Kit (Qiagen, 158567). Plasmids and BACs were electro-transformed to E. coli NEB10β (New England Biolabs, C3020K) for propagation and MG1655 for downstream experiment. Primer sequences for PCR amplification were summarized in Table 2. All the plasmids were sequence confirmed by whole-plasmid sequencing (Primordium Labs). REXER The refactored BCAA biosynthetic iModulon was introduced into genome using the replicon excision enhanced recombination (REXER) method56,57with minor modification. First, BAC_BCAA_Biosynthetic_iM plasmid containing BCAA biosynthetic iModulons was introduced using electroporation to E. coli ΔBCAA strain carrying pREDCas9 plasmid and sacB-cat dual-selection landing pad upstream of thrL gene. Recipient cell was grown in terrific broth (TB) containing 25 µg / ml chloramphenicol, 50 µg / ml kanamycin, 100 µg / ml carbenicillin, and 30 mM L-arabinose at 37°C to cell density (A600) of 0.8. Two micrograms of PCR-amplified linear guide-RNA construct, targeting both the landing pad and the BAC, were introduced using electroporation and cells were screened on NaCl-free LB-agar plate (10 g / l tryptone, 5 g / l yeast extract, and 1.5% agar) containing 100 µg / ml carbenicillin, 50 µg / ml kanamycin, and 10% sucrose. Successful recombinants were further screened by chloramphenicol sensitivity and PCR genotyping. HPLC Att Dkt: 114198-5960 Culture supernatant was collected by filtering 200 µl of the crude culture through 96- Well PVDF Filtration Plate (0.2 µm; Agilent, 203980-100). Exo-metabolites were analyzed from 10 µl of the sample using 1260 Infinity II HPLC System (Agilent) equipped with Multisampler (G7167A), HIP Degasser (G4225A), Binary Pump (G1312C), and Refractive Index Detector (G1362A), Thermostatted Column Compartment (G1316A), and Aminex HPX-87H HPLC Column (300 × 7.8 mm; BioRad, 1250140). Five millimolar sulfuric acid was used as a mobile phase and the detector temperature was maintained at 30°C. Vanillate, proto-catechuate, acetoin, 2,3-butanediol were detected at column temperature of 65°C with mobile phase flow rate of 0.6 ml / min. For detection of malonate column temperature was maintained at 45°C. Quantitative PCR To measure the plasmid to chromosome ratio and the expression level of mdcR iModulon construct, cells were grown in 15 ml M9 malonate medium and sampled at the mid-log phase. Total DNA was extracted from 1 ml of the culture using Quick-DNA Miniprep Kit (Zymo Research, D3024) as instructed by the manufacturer. 100 ng of DNA extract was subject to quantitative PCR in a 20 µl reaction containing AccuPower PCR 2× Master Mix (Bioneer, K-2018), 10 µM each of primers, SYBR Green I Nucleic Acid Gel Stain (Invitrogen, S7563). Fluorescence signals were monitored by CFX Duet Real-Time PCR System (RioRad, 12016265). Amount of beta-lactamase (bla) gene on the plasmid was compared to the alaA gene on chromosome to estimate the relative number of the plasmid to chromosome. PCR efficiencies were measured from three 2-fold dilution series of each gene. Plasmid to chromosome ratio (P / C ratio) is calculated by where EalaA and Ebla are the PCR efficiencies of alaA gene and bla gene, respectively. Cq⋅alaA and Cq⋅bla are the quantification of cycles of respective genes. To estimate expression level of MdcR iModulon, relative expression level of mdcA to 16S rRNA was measured. RNA was extracted from 14 ml of the culture using Quick-RNA Fungal / Bacterial Miniprep Kit (Zymo Research, R2014) as instructed by the manufacturer. Residual DNA was removed by incubating 2 µg of RNA extract at 37°C for 30 min in a 50 µl reaction containing 2 U of RNase-free DNase I (New England Biolabs, M0303) followed by Att Dkt: 114198-5960 a purification using RNA Clean & Concentrator Kit (Zymo Research, R1013). cDNA was synthesized from 300 ng of the DNA-depleted RNA sample using SuperScript II First-Strand Synthesis System (Invitrogen, 11904018) as instructed by the manufacturer. 1 µl of cDNA synthesis reaction was subject to quantitative PCR in a 20 µl reaction containing AccuPower PCR 2× Master Mix , 10 µM each of primers, SYBR Green I Nucleic Acid Gel Stain. Relative amount of mdcA transcript was quantified by the ΔΔCq method using rrsA transcript as a reference, after adjusting PCR efficiencies that were measured from three 2-fold dilution series of each gene. All the primers were designed by Primer-BLAST58to have no predicted cross-reactivity. Primer sequences are summarized in Table 2. Antibiotics sensitivity assay For disc diffusion assay, overnight grown E. coli culture was diluted to 5×108cells / ml and 500 µl of the diluted culture was spread on a LB-agar plate. Thirty microliters of ampicillin solutions with different concentrations were dropped on sterilized filter paper discs (9 mm diameter; Sigma-Aldrich, 1703932). After drying for 15 min, the discs were placed on the LB-agar plate and the plate was incubated at 37°C overnight. For dose-killing assay, overnight grown E. coli culture was diluted to 5×108cells / ml in LB medium containing appropriate concentration of ampicillin and 100 µl of the diluted culture was transferred to a 96-well microplate. Cells were incubated at 37°C with agitation in Infinite 200 Pro microplate reader (Tecan) and A600was monitored every 15 min up to 10 hr. Adaptive Laboratory Evolution E. coli K-12 MG1655 carrying pMdcR_iM plasmid was evolved via serial propagation of 150 µl into 15 ml M9 malonate (2 g / l) minimal medium containing 100 µg / ml carbenicillin in three biologically replicated cultures. Cultures were incubated at 37°C, aerated by magnetic stirring, and reinoculated at A600 of 0.3 (Tecan Sunrise microplate reader; equivalent to an A600 of 0.750 on a conventional spectrophotometer with a path-length of 10 mm) using an automated system. After 40 days of propagation, evolved cultures were subjected to malonate utilization assay and three clones were isolated from each culture. DNA sequencing and analysis Genomic DNA was isolated using Quick-DNA Miniprep Kit as instructed by the manufacturer. Whole-genome DNA-seq libraries were generated with a NEBNext Ultra II DNA Library Prep Kit for Illumina (New England Biolabs, E7645) and run on an Illumina NovaSeq X Plus with 100 cycles pair-ended recipe. The sequencing results were processed Att Dkt: 114198-5960 with the in-house pipeline59that incorporates BreSeq pipeline to identify mutations60. E. coli K-12 MG1655 genome sequence (National Center for Biotechnology Information accession no. NC_000913.3) was used as a reference sequence. Data Availability The genome resequencing data generated in this study have been deposited in the EMBL Nucleotide Sequence Database (ENA) under accession code PRJEB69310 [https: / / www.ebi.ac.uk / ena / browser / view / PRJEB69310]. The iModulon activity and gene membership data used in this study are available in the iModulonDB [https: / / imodulondb.org / ]. Predicted operonic structures were from Biocyc database [https: / / www.biocyc.org / ]. Genome sequence and genetic organization information were assessed through Pseudomonas Genome Database [https: / / www.pseudomonas.com / ]. All other data are available in the article, Supplementary items, or Source Data file. Source data are provided with this paper. Code Availability The mutation detection from the DNA sequencing results was performed with the ALEdb (v1.1.0) pipeline [https: / / www.aledb.org / ] that incorporates BreSeq pipeline59,60. RESULTS Cross-species transfer of Pseudomonas iModulons into E. coli: To initiate the project and prior to implementing cross-species iModulon transfer, Applicant refactored a known cellular function within the original host as a proof of concept. Successful homologous refactoring and complementation of E. coli’s branched-chain amino acid (BCAA) metabolism was achieved (FIG. 9) to demonstrate identification, reconstruction, and transfer of genetic constituent of a biological function based on iModulon (i, ii, and iii). This motivated Applicant to investigate the potential for transferring biological functions across species. Among the available species with iModulon structures in iModulonDB12, Pseudomonas is well-known for their versatile metabolism to degrade and utilize diverse compounds, including aromatics19–21. First, Applicant chose to reconstruct and transfer a simple bioconversion process from Pseudomonas putida15to E. coli in order to examine iModulon’s capability to rapidly identify genes associated with specific function. The VanR iModulon that is responsible for vanillate (VA) transport and conversion into proto-catechuate (PCA) was chosen for Applicant’s first cross-species iModulon transfer (FIG. 1A). It comprises three genes with annotated functions, vanA, vanB, vanK, and Att Dkt: 114198-5960 predicted porin-like galP-IV (FIG. 1B) in two converging operons (FIG. 1C). Notably, the iModulon exactly matches with the genes for the vanillate transport and metabolism22,23. Four genes, vanA, vanB, galP, and vanK are functionally annotated to encode for vanillate O- demethylase oxidoreductase complex, outer-membrane porin, and a major facilitator superfamily transporter, respectively22. Although the function of outer membrane OprD- domain containing galP-IV has never been addressed, it is hypothesized that it facilitates the diffusion of the ligand through the outer membrane23,24. Since the mechanism of VanR regulation has not been established, the four genes constituting the VanR iModulon were cloned and heterologously expressed under the control of IPTG-inducible Trc promoter on a plasmid, pVanR_iM (FIG. 1D). When refactoring iModulons for heterologous expression, Applicant tried to preserve native genetic arrangement, for VanR and following iModulons if possible, to ensure optimal expression levels of the gene members as demonstrated elsewhere25,26. E. coli carrying pVanR_iM converted VA into PCA up to 15.34 mg / l passively diffused to the supernatant27,28during 48 hours of fermentation in M9 glucose (4 g / l) medium supplemented with 100 mg / l VA, while the negative control carrying empty plasmid did not metabolize any VA (FIG. 1E). This first cross-species iModulon transplantation illustrates the rapid identification of enzymes required for biotransformation by ICA. Furthermore, iModulon engraftment provided a rapid way to biochemically verify a predicted pathway in a heterologous host. Auxiliary genes may be needed for optimal function of cross-species transferred iModulons: Next Applicant chose to transfer an ampicillin resistance function of Pseudomonas aeruginosa to E. coli. P. aeruginosa displays beta-lactam resistance with endogenous beta-lactamase, AmpC, and has an iModulon involved in the inducible ampicillin resistance16. Activity levels of the AmpC iModulon are highly induced against beta-lactam challenge, but not under other antibiotic treatments (FIG. 10). In the previous iModulon engraftment examples, genes comprising an iModulon matched with the predicted genes necessary for building the desired function. However, identifying all the genes necessary to build a biological function may not be trivial given previous characterization efforts. Many iModulons contain genes whose functions are unknown or are seemingly unrelated to the overall function being transferred. The AmpC iModulon comprises class C beta-lactamase encoded by the ampC gene29that serves as a core for the functionality and six lesser characterized auxiliary genes, carO Att Dkt: 114198-5960 (PA0320), creD (PA0465), PA0466, PA0467, PA4111, and PA4112 (FIG. 2A). The seven iModulon genes are distributed across three genomic loci separated by over 4 Mb. P. aeruginosa readily becomes resistant to ampicillin by transcriptional activation of ampC30. However, it is not known if the resistance trait is carried by this single gene. To examine if this resistance function is transferable across species, the constituent genes were refactored into a single operon (FIG. 2B). In addition, Applicant constructed a plasmid that contained beta-lactamase alone to address any involvement of auxiliary factors into the function. Ampicillin disc diffusion assay revealed that E. coli carrying the AmpC iModulon or ampC gene were resistant to ampicillin, while E. coli carrying empty plasmid were not (FIG. 11). The source of AmpC iModulon, P. aeruginosa, showed ampicillin resistance with the minimum inhibitory concentration (MIC) of 2048 µg / ml (FIG. 2C). The MIC of ampicillin for laboratory E. coli strain MG1655 with empty plasmid was 16 µg / ml, which is comparable to previous reports31,32(FIG. 2C and FIG. 2D). E. coli strain with the P. aeruginosa beta- lactamase showed a dramatic increase in ampicillin resistance with an MIC of 1024 µg / ml, while it was lower than that of the original host FIG. 2D). Strikingly, E. coli harboring the entire AmpC iModulon, six auxiliary genes in addition to ampC, had an MIC of 4096 µg / ml, which was four times higher than that with ampC alone (FIG. 2D). Although little is known about the molecular function of auxiliary genes, they were required to completely replicate the ampicillin resistance characteristics of P. aeruginosa. Previous reports have shown a decrease in beta-lactam resistance of the inner membrane protein creD knockout mutant of P. aeruginosa33and growth enhancement of E. coli by endogenous creD overproduction (shares 37.4% sequence identity; BLOSUM62)34. Although the function of CreD is still elusive, reports indicate its relevance in biofilm development in P. aeruginosa35and envelope integrity in Stenotrophomonas maltophilia36. Additionally, calcium-regulated oligonucleotide / oligosaccharide binding (OB)-fold protein CarO has been reported to be related to susceptibility to various stresses in bacteria37. Also, it shares similarity with Salmonella enterica stress-related protein VisP (38% sequence identity) which binds to peptidoglycan and inhibits the lipid A modifying enzyme LpxO38. Since lipid A is an anchor of lipopolysaccharide to the outer membrane and affects the properties of the outer membrane, expression of carO might be beneficial for cells to maintain structural integrity under cell wall deficient conditions induced by beta-lactam39. Engrafting Pseudomonas iModulons to E. coli highlighted critical properties of iModulon gene membership. Harnessing only core genes for transferred cellular function Att Dkt: 114198-5960 may not be sufficient as auxiliary genes may be needed to reconstruct an optimal function. Full iModulon gene membership helps to recreate the targeted cellular function, even without a complete understanding of the molecular function of all the genes involved. Complete iModulon gene membership is needed for successful cross-species transfer: As illustrated by the AmpC case, Applicant further investigated iModulon-based transfer of cellular traits and compared it to the alternative conventional methods. The 2,3- butanediol (2,3-BDO) iModulon was chosen to examine the role of iModulon genes of unknown functions. 2,3-BDO is a byproduct of bacterial fermentation processes that can be produced by a variety of microorganisms, including Pseudomonas species40–42. In Pseudomonas , 2,3-BDO can serve as a carbon and energy source and is degraded by enzymes in the 2,3-BDO catabolic pathway42. This catabolic pathway involves the conversion of 2,3-BDO into acetoin, which is further converted into acetaldehyde and acetyl- CoA by butanediol dehydrogenase and acetoin dehydrogenase, respectively (FIG. 3A). Applicant transferred the 2,3-BDO iModulon of P. putida (called the AcoR iModulon15) to E. coli. The AcoR iModulon comprises acoABC (encoding acetoin dehydrogenase complex), bdhA (encoding 2,3-BDO dehydrogenase), and a gene acoX (FIG. 3B). AcoX encodes for a protein of unknown function and co-exists with acetoin utilizing genes in various bacteria41,43. Operon prediction also suggests that the transcriptional unit contains acoX and two other hypothetical proteins (PP_0550 and PP_0551) in addition to characterized metabolic enzymes, acoABC-bdhA (FIG. 3C)18,44. To examine which genes are required for recreating the 2,3-BDO catabolic pathway, Applicant built three different plasmid based on (1) operonic structure (Op353; acoXABC- bdhA-PP_0551-PP_0550), (2) iModulon structure (acoXABC-bdhA), and (3) four genes encoding enzymes predicted to be sufficient for converting 2,3-BDO into acetaldehyde and acetyl-CoA based on current gene annotations (pathway; acoABC-bdhA) (FIG. 3C). 2,3- BDO dehydrogenase activities of the source organism and E. coli strains carrying the three plasmids individually were examined during 96 hours of batch cultivation in LB medium supplemented with 2 g / l of 2,3-BDO. The original strain, P. putida KT2440, showed 2,3- BDO utilization with negligible level of acetoin (FIG. 3D). The negative control, E. coli MG1655 carrying an empty plasmid converted 0.77 g / l of 2,3-BDO into acetoin, possibly due to endogenous promiscuous alcohol dehydrogenase activity (FIG. 3E). On the other hand, the plasmids based on pathway, operonic structure, and iModulon showed higher conversion of 2,3-BDO with amounts of 1.36, 1.75, and 1.96 g / l, respectively (FIG. 3E). Att Dkt: 114198-5960 Interestingly, the strains showed varying levels of acetoin dehydrogenase activity. First, all the 2,3-BDO consumed by the negative control resulted in roughly the equimolar amount of acetoin; not surprising since there is no acetoin dehydrogenase introduced. The strain carrying the functional gene annotation-based pathway plasmid did not further convert acetoin into downstream products, even though it contained genes encoding for the acetoin dehydrogenase complex. Second, strains with the full operon or AcoR iModulon not only consumed more than 1.7 g / l of 2,3-BDO but there was only a small amount of acetoin left in the medium, indicating conversion of acetoin by acetoin dehydrogenase. The difference between annotation-based and iModulon-based plasmid is the presence of acoX (FIG. 3C), a gene encoding a predicted small molecule kinase that has been reported to have no acetoin, NAD, or pyruvate kinase activity45. However, acoX was critical for acetoin dehydrogenase activity. Although the acoX product has no known function in acetoin metabolism, it is conserved and colocalizes on the genome with the acetoin dehydrogenase in several acetoin- utilizing bacteria from multiple phyla, such as P. aeruginosa (76% sequence identity) and Clostridium magnum (32% sequence identity)42. However, there is no significant match of AcoX from BLASTP search on other acetoin-utilizing bacteria such as Bacillus subtilis, Klebsiella pneumoniae, and Pelobacter carbinolicus. Therefore, the requirement of AcoX in acetoin metabolism is species-specific and could not be determined by analyzing the genome sequence context. When the iModulon and operonic constructs were compared, the iModulon construct performed better than the operonic construct for 2,3-BDO degradation (FIG. 3E). Two additional genes in the operonic construct encode the predicted membrane occupation and recognition nexus (MORN) domain-containing peptidase and a NAD(P)-binding oxidoreductase, whose relation with 2,3-BDO metabolism is unknown. These two genes were irrelevant for function. Instead, expression of the hypothetical proteins reduced 2,3- BDO degradation, possibly by imposing unnecessary transcriptional burden on the cell. The iModulon gene membership provided information on the necessary genes to support a 2,3- BDO catabolic process that would not have been found using only functional gene annotation. This example illustrates the unique advantages of using iModulon structure for cross-species transfer of the full genetic basis for a desired integrated function. ALE optimizes the functionality of catabolic iModulons: Applicant chose the MdcR iModulon from P. aeruginosa16to transfer into E. coli that, again, comprises genes Att Dkt: 114198-5960 identical to a reported set for the malonate transport and utilization. The MdcR iModulon comprises seven subunits of malonate decarboxylase complex21and two putative membrane proteins, MadL-MadM (FIG. 4A and FIG. 12). Although the function of these membrane proteins have not been elucidated in P. aeruginosa, MadL and MadM have 71 and 81% of sequence homology to malonate transporters in Malonomonas rubra46, respectively, suggesting a potential malonate uptake function. These genes are encoded in a single operon on the P. aeruginosa genome, and thus the entire operon was subjected to cross-species transfer. The operon was cloned and heterologously expressed under the control of a Trc promoter on a plasmid, named pMdcR_iM (FIG. 4B). Malonate is a non-native nutrient for E. coli, thus it is expected that a strain with the pMdcR_iM alone would then enable growth in M9 malonate medium as the breakdown product of the pathway, acetate, can support growth47. Applicant experimented with varying levels of expression using different concentrations of the inducer (IPTG) to activate the MdcR iModulon. E. coli could slowly utilize (doubling time of 11.2±0.6 hr; over the course of 72 hours of fermentation in M9 malonate medium) malonate as a carbon source only at weak expression level (FIG. 4C). In contrast to complete utilization of malonate by P. aeruginosa within 12 hours of fermentation (FIG. 4C), the observed slow utilization by E. coli suggests a potential metabolic imbalance in E. coli, perturbed by and unable to accommodate the malonate pathway. Therefore, Applicant implemented adaptive laboratory evolution to allow E. coli to rebalance and optimize its metabolism with malonate as a substrate. The E. coli strain carrying the pMdcR_iM was grown in M9 malonate medium and evolved using serial passaging that imposes growth rate selection pressure (FIG. 4D) on an automated ALEbot48. After 21 passages, populations showed faster growth with a short lag phase compared to their ancestor (FIG. 4E). The evolved populations fully consumed malonate within 40 hrs. Subsequently, three clones were isolated from each replicate evolved population and they all displayed a faster growth rate than the ancestor (FIG. 4F). To understand the genetic bases of improved growth, Applicant resequenced the genome of the evolved clones (Table 1). All the evolved clones carried mutations on DNA polymerase I, encoded by polA, which is required for plasmid maintenance (FIG. 4G)49. Previous studies reported a change of plasmid copy number induced by polA mutation50. Quantitative measurement of plasmid copy number indicated a reduction of plasmid copy number, which led to a reduction of MdcR iModulon expression (FIG. 4H). Thus, the initial Att Dkt: 114198-5960 metabolic failure was likely due the sub-optimal expression of the MdcR iModulon, which could be optimized by ALE. Engraftment of the MdcR iModulon, in addition to three other iModulons, demonstrated cross-species iModulon transfer as a rapid way of creating new functionality in bacteria with minimal engineering. Applicant found that the overall behavior of the iModulon interferes with the host factors that require modifications to optimally support the system. This optimization could be rapidly achieved by ALE that identified few genetic changes in the host, while the transferred genes acquired no adaptive mutations. Additional Pseudomonas putida aromatic compound degradation pathway: Lastly, Applicant divided the PcaR iModulon, comprising 45 genes of known and unknown function related to aromatic compound degradation in P. putida, into smaller modules called sub-iModulons (FIG. 5A). See Lim et. al., Metabolic Engineering 72:297-310 (2022) for additional information about the PcaR iModulon, which is hereby incorporated by reference in its entirety. The hydroxycinnamates (HCA) sub-iModulon, comprising SEQ ID Nos: 1-4, 53-54, 57-61, and 64-65; the VanR sub-iModulon, comprising SEQ ID Nos: 1-4; and the protocatechuic acid (PCA) sub-iModulon, comprising SEQ ID Nos: 26-43 from P. putida were transferred into E. coli E. coli could slowly utilize PCA as a carbon source (FIG. 5B). In contrast, E. coli expressing the PCA sub-iModulon utilized PCA similar to P. putida (FIG. 5B). E. coli could slowly utilize vanillic acid (VA) as a carbon source (FIG. 5C). In contrast, E. coli expressing the PCA and VanR sub-iModulons utilized VA similar to P. putida (FIG. 5C). E. coli could slowly utilize 4-coumaric acid (4CA) as a carbon source (FIG. 5D). In contrast, E. coli expressing the PCA and HCA sub-iModulons utilized 4CA significantly better than E. coli (FIG.5D). Engraftment of the PCA, VanR, and HCA sub-iModulons demonstrated cross-species iModulon transfer as a rapid way of creating new functionality in bacteria with minimal engineering. Engraftment of the BenR sub-iModulon, comprising SEQ ID Nos: 67-76, although not demonstrated, would additionally impart new benzoate metabolism functionality to bacteria. DISCUSSION Historically synthetic biology has built multi-genic functions in a serial manner. Individual genes are introduced to the host and their function is assessed. iModulons, Att Dkt: 114198-5960 obtained through big data analysis of transcriptome compendia, describe sets of co-expressed genes that constitute independent cellular functions, suggesting that multigenic traits can be captured and transferred. Here Applicant demonstrate that this is possible through cross- species transfer of cellular functions from Pseudomonas species into E. coli. Applicant successfully transfer three metabolic traits and an antimicrobial resistance trait between these Gram-negative species. iModulons lead to identification of genes that are not functionally annotated but essential for recreating a cellular function in a new host. Thus, these genes gain a network level functional annotation, as opposed to classical molecular level functional annotation. Applicant also find that iModulon genes may have an auxiliary function, i.e., they enhance the cellular functions being transferred. Finally, Applicant find that laboratory evolution can improve the transferred cellular function in the new host. Sequencing laboratory evolved strains can identify host factors that enable or enhance the cellular function encoded on the transferred construct. Taken together, these factors lead to an advancement of synthetic biology, including identification of all genes needed to constitute a cellular function and revealing host factors that need modification to optimize the engineered strain. This acceleration and higher predictability in strain design and construction should accelerate the development of synthetic biology and its deployment for practical purposes such as biomanufacturing. Example 2 Repairing and reconfiguring broken branched-chain amino acid synthetic iModulons in E. coli K-12 Genes that are members of iModulons exist in trans-locations and transfer of all the genes on an iModulon requires refactoring into a contiguous piece of DNA. We refactored branched-chain amino acid (BCAA) biosynthetic iModulon in its native host. BCAAs are essential amino acids that play a critical role in protein synthesis and energy metabolism. The genes that constitute these biosynthetic functions are found in two iModulons in E. coli K-12 strains(Lamoureux, C. R. et al., Escherichia coli. Nucleic Acids Res., 2023). They can synthesize all three BCAAs — leucine, isoleucine, and valine — using intermediates from central carbon metabolism, pyruvate and oxaloacetate. The biosynthesis involves a series of enzymatic reactions, including the use of acetohydroxyacid synthase (AHAS) complexes. E. coli has three different AHAS isozymes, each with unique properties and product inhibition Att Dkt: 114198-5960 sensitivities (Blatt, J. M. et al., Biochem. Biophys. Res. Commun. 48, 1972; Vinogradov, V. et al., Biochim. Biophys. Acta 1760, 2006). This ensures the proper balance of production and utilization of the three BCAAs. However, the ilvG gene in E. coli K-12, which encodes a catalytic subunit of valine-insensitive AHAS II in conjunction with a regulatory subunit ilvM, has a frameshift mutation that renders the enzyme non-functional (Lawther, R. P. et al., Proc. Natl. Acad. Sci. U. S. A. 78, 1981). This mutation, combined with the inhibitory effects of valine on the other two AHAS isozymes, results in the failure of BCAA biosynthesis in the presence of excess valine (Choe, D. et al., Nat. Commun. 10, 2019). In addition, the deficiency induces repeated isoleucine starvation in fermentation (Andersen, D. C. et al., Biotechnol. Bioeng. 75, 2001). During efforts to correct this genetic deficiency in the strain, Applicant explored the potential of iModulon-based genome refactoring, with the goal of gaining insights that could be applied to the transfer of iModulons whose genes are in a trans- configuration on the genome. We are particularly interested in whether the presence of all the iModulon genes is sufficient and if the genomic organization, such as location, direction, and order of genes, is important for BCAA biosynthesis. The iModulon structure of two closely-related E. coli K-12 strains MG1655 and BW25113, determined from more than 400 unique experimental conditions, shows that E. coli K-12 strains have two different iModulons dedicated to BCAA synthesis (FIG. 9A) (Lamoureux, C. R. et al., Escherichia coli. Nucleic Acids Res., 2023). These iModulons encode multiple genes to synthesize BCAA from pyruvate or homoserine and complement each other due to the multifunctionality of some of the genes (FIG. 9B). The two iModulons contain a full repertoire of genes required for the BCAA biosynthesis. Thus, we included the two iModulons in the reconfiguration of the complete BCAA biosynthetic pathways. The reconfiguration involves moving iModulon members from four different genomic loci to one locus altogether and aligning their orientation into the same direction. To set the boundary of transfer to preserve regulatory elements in the iModulon, we utilized transcription start sites (TSSs) and transcription termination sites (TTSs) represented by Bitomic information (Choe, D. et al., Nat. Commun. 10, 2019), (Lamoureux, C. R. et al., Nucleic Acids Res. 48, 2020). Outermost TSSs and TTSs that define the largest transcriptional unit were chosen for boundaries and DNA fragments containing the transcriptional unit were cloned and assembled (FIG. 9D). During the cloning, a single nucleotide insertion was introduced to fix the frameshift mutation in ilvG. After removal of endogenous copies of all iModulon genes (ΔBCAA knockout strain) from the chromosome, the assembled iModulon construct was Att Dkt: 114198-5960 integrated into the chromosome at thr locus using replicon excision for enhanced genome engineering through programmed recombination (REXER) (Wang, K. et al., Nature 539, 2016), (Robertson, W. E. et al., Nat. Protoc. 16, 2021). (FIG. 9D). In an M9 glucose defined medium, the reconfigured strain (RECON) showed growth rates comparable to its parental strain, while the knock-out strain (ΔBCAA) failed to grow (FIG. 9E). The knock-out strain was only able to grow with supplementation of all three BCAAs (FIG. 9E). Wild-type MG1655 and the reconfigured strain behaved similarly with isoleucine and leucine supplementation (FIG. 9E). With valine supplementation, the wild- type strain showed dramatic growth retardation due to the aforementioned valine toxicity (FIG. 9E). The reconfigured strain showed no growth inhibition with external valine, indicating rewiring of BCAA biosynthesis. This reconfiguration demonstrates that the order, orientation, and location of genes on the genome had negligible effects on functionality if the refactored system had the complete set of genes. It suggests transferability of iModulons based on the transcriptional framework decoded by ICA(Lamoureux, C. R. et al., Escherichia coli. Nucleic Acids Res., 2023). Metabolic failure induced by MdcR iModulon The E. coli strain carrying the pMdcR_iM was unable to grow in M9 malonate medium when the governing promoter (IPTG-inducible Trc promoter) was induced. It is likely due to the inhibition of endogenous enzymes — succinate:quinone oxidoreductase (Maklashina, E. et al., Escherichia coli. Arch. Biochem. Biophys. 369, 1999) and isocitrate lyase11(Hoyt, J. C. et al., Biochim. Biophys. Acta 966, 1988). constituting TCA and glyoxylate cycle, respectively — by malonate. Suboptimal over expression of the iModulon would produce the malonate transporter MdcL and MdcM, increasing intracellular malonate levels, leading to a metabolic failure. We demonstrated that heterologous expression of iModulon could be rapidly optimized by ALE. During ALE, polA mutations occurred and the copy number of the plasmid was reduced, resulting in 4 to 50-fold decrease in mdcA expression. Besides, there was no mutation on the genes composing MdcR iModulon, which again indicates that iModulons are units of biological traits that require no extensive engineering at transfer. iModulon Identification and Cell Function Translation Many approaches to strain engineering and metabolic optimization face significant technical challenges in identifying unknown gene functions within organisms. The Att Dkt: 114198-5960 complexity of cellular metabolism, combined with the large number of uncharacterized genes (e.g., genes with an unknown function) and their interactions, makes it difficult to rationally design strains with desired phenotypes. Additionally, given the large number of uncharacterized genes, it can take a long time and a substantial amount of processing resources to identify groups of genes that operate together to perform a function, particularly given the large amount of noise in the data (e.g., expression data) regarding each individual gene. The systems and methods described herein overcome these technical hurdles by using independent component analysis (ICA) to identify and isolate iModulons from gene expression data. Doing so can facilitate the identification of previously unknown functional gene groups and transfer of the iModulons into an organism. For instance, a computer can use ICA to determine new iModulons and then transfer the new iModulons into cells, bacterial artificial chromosomes (BACs), or other gene transfer vehicles of an organism. The transfer can cause the cells or bacteria to mutate and / or have a new function or functionality (e.g., the functionality of the transferred iModulon). The method can further involve filtering out low-expression genes and focusing on strongly coregulated gene sets to lead to a more reliable identification of functionally related genes and / or provide the functionality to BACs of an organism. This approach can detect coordinated gene expression patterns that may not be evident through conventional genomic identification methods. By using ICA to identify and isolate iModulons, the system and method can identify regulatory relationships of iModulons without prior knowledge of each gene function and identify functional associations between genes with a known and unknown function. Using ICA for iModulon identification for transfer into organisms has multiple technical advantages over other methods (e.g., pattern matching). For example, using ICA can substantially reduce the processing resources and latency required to identify iModulons and / or respond to requests for iModulons compared with systems that use other methods, such as pattern matching between groups of genes, particularly given that there are millions of permutations or combinations of genes that may or may not be involve in metabolic processes. For example, instead of analyzing thousands or millions of individual gene correlations, the ICA can facilitate the compression of expression data into a smaller set of meaningful regulatory components. This can make the computational process more efficient and help identify the most significant gene relationships in the large dataset that includes different characteristics of genetic data. Att Dkt: 114198-5960 Optionally, the method can involve using adaptive laboratory evolution (ALE) to allow for the systematic optimization of strain performance under selective pressure. For instance, the method can involve the iterative process of measuring fitness and selecting improved variants, which can facilitate the discovery of beneficial mutations that might. This approach, combined with whole genome sequencing, can provide an understanding of intended and unintended genetic changes that contribute to improved phenotypic expression. The method can further provide technical benefits for strain engineering. For example, the method can involve sequencing both the optimized strain and the gene transfer vehicle to facilitate the identification of both on-target and off-target mutations that contribute to a desired phenotype. The ability to distinguish between mutations within the transferred genes and those in the host genome can provide information for subsequent strain design iterations. The analysis can allow for the incorporation of beneficial mutations into future strain designs, creating a feedback loop that can continuously improve the engineering process and leads to increasingly optimized strains. For example, FIG. 6 is an illustration of an example system 600 for iModulon detection, in accordance with implementations. In brief overview, the system 600 can include a data processing system 602, a remote data source 618, and a client device 620. The client device 620 can be or include a computing device, a server, a mobile device, a table, a personal computer, or any other type of computing device. The data processing system 602 can receive or retrieve genetic data of multiple organisms from the remote data source. The data processing system 602 can apply independent component analysis on the genetic data to identify one or more iModulons. Each iModulon can include one or more genes with a known function and one or more genes with an unknown function. The data processing system 600 can transmit a message to the client device 620 identifying each of the one or more iModulons (e.g., identifying the genes included in each iModulon). The system 600 may include more, fewer, or different components than shown in FIG. 6. The remote data source 618 can be or include one or more data sources that store genetic data of individual or different organisms and / or external data. The remote data source 618 can be a database, a computer, a laptop, a server, or any other type of device that can store data. For example, the remote data source 618 can be or include a database (e.g., a relational or graph database) that stores data indicating levels of expression different groups of genes exhibited under different conditions (e.g., culture conditions). Examples of the conditions can include varying the pH level, changing the ions, changing substrates, and / or Att Dkt: 114198-5960 putting other cells in the media containing the genes. Each level can be represented as numerical, alphabetical, and / or alphanumerical value. For each condition, the remote data source 618 can store a vector identifying the level of expression of individual genes cultured under the condition. The remote data source 618 can store such vectors for any number of groups of genes for an individual condition. The remote source 618 can generate and / or store such vectors using genes from a single organism or from multiple organisms. In some cases, the remote data source 618 can store the expression data separately and not in vector form. In some cases, the remote data source 618 can store data or metadata about individual genes. For example, the remote data source 618 can store data indicating the functionality (e.g., molecular functionality and / or genetic functionality) of individual genes and / or indications of whether the genes have a known function. The remote data source 618 can store any amount and / or type of data regarding genes. In one example, the remote data source 618 can include one or more data sources (e.g., as one or more computing devices at the same or different physical locations) for gene expression data in transcriptomic datasets. The remote data source 618 can store databases, such as the Gene Expression Omnibus (GE), ArrayExpress, and / or the Sequence Read Archive (SR). The different databases can store gene expression data and / or can include RNA sequencing data showing the number of reads mapped to each gene, or microarray data with hybridization intensity values, along with metadata about experimental conditions, sample characteristics, and / or technical parameters. The datasets may additionally or instead contain quality metrics, alignment statistics, and / or information about splice variants or alternative transcripts, depending on the experimental platform and processing pipeline used. The data can be generated from individual experiments or from multiple experiments that involve measuring the expression levels of genes in one or more conditions. The data processing system 602 may comprise one or more processors that are configured to determine iModulons. The data processing system 602 may comprise a network interface, a processor, and / or memory. The data processing system 602 may communicate with the remote data source 618 and / or the client device 620 via the network interface, which may be or include an antenna or other network device that enables communication across a network and / or with other devices. The processor may be or include an ASIC, one or more FPGAs, a DSP, circuits containing one or more processing components, circuitry for supporting a microprocessor, a group of processing components, or other suitable electronic processing components. In some embodiments, the processor may Att Dkt: 114198-5960 execute computer code or modules (e.g., executable code, object code, source code, script code, machine code, etc.) stored in memory to facilitate the activities described herein. The memory may be any volatile or non-volatile computer-readable storage medium capable of storing data or computer code. The memory may include a data collector 604, a profile generator 606, an iModulon generator 608, and / or a data repository 612, in some embodiments. The components 604-612 may operate to identify iModulons containing different sets of genes that correspond to different functions. In doing so, the components 604-612 can identify groups of genes that each include at least one gene with a known function (e.g., known genetic function or molecular function) and at least one gene with an unknown function (e.g., unknown genetic function or molecular function). For example, the data collector 604 may comprise programmable instructions that, upon execution, cause the data processing system 602 to communicate with the remote data source 618, the client device 620, and / or any other computing devices. The data collector 604 may be or include an application programming interface (API) that facilitates communication between the data processing system 602 and other computing devices. The communicator 604 may communicate with the remote data source 618, the client device, and / or any other computing device across a network 601 (e.g., a wired or wireless network). The data collector 604 can establish connections with the remote data source 618 and / or the client device 620. The data collector 604 can establish the connections over the network 601. To do so, the data collector 604 can communicate with the remote data source 618 and / or the client device 620 across the network 601. In one example, the data collector 604 can transmit syn packets to the remote data source 618 and establish the connection using a TLS handshaking protocol. The data collector 604 can use any handshaking protocol to establish a connection with the remote data source 618. The data collector 604 can collect gene expression data from the remote data source 618. For example, the data collector 604 can send requests or otherwise crawl the remote data source 618 for genetic expression levels (e.g., a relative expression level, such as a fold- change relative to a control condition, or an absolute expression level, such as a copy number per cell) under different conditions. In response, the data remote data source 618 can transmit the expression levels under the different conditions to the data processing system 602. Att Dkt: 114198-5960 The profile generator 606 may comprise programmable instructions that, upon execution, cause the data processing system 602 to generate transcriptomic data comprising expression profiles for groups of genes when placed under different conditions. For example, the profile generator 606 can identify the expression levels received from the remote data source 618. The profile generator 606 can identify the expression levels of the genes under the different conditions and group the expression levels together. Examples of the conditions can include conditions relating to varying levels of humidity, CO2 levels, light / dark cycles, agitations speed, nitrogen source (e.g., amino acids, ammonium), carbon source (e.g., glucose, glycerol), oxygen levels, pH levels, temperature, introductions of different cells, etc. In grouping the expression levels of the genes, for example, the profile generator 606 can generate a separate vector for each condition from the expression levels of the genes that were cultured under the condition. Each vector can include a numerical value at different index values that represent an expression level of an individual gene under the condition for which the profile generator 606 is generating the vector. Each vector can be a gene expression profile for a given condition. In some cases, instead of aggregating the gene expression level data into separate vectors, the profile generator 606 can receive such vectors from the remote data source 618 when the remote data source 618 provides the gene expression data to the data processing system 602. The profile generator 606 can store the vectors in the data repository 612 as expression profiles 614. Accordingly, the profile generator 606 can generate and / or store vectors with different numerical values that represent the expression levels of genes. The iModulon generator 608 may comprise programmable instructions that, upon execution, cause the data processing system 602 to identify or generate groups of genes of iModulons. The iModulon generator 608 can generate the iModulons using independent component analysis (ICA). In doing so, the iModulon generator 608 can generate or identify individual iModulons that each include at least one gene with a known function and at least one gene with an unknown function. Such iModulons can be transferred into bacteria or cells of organisms to give the bacteria or cells a new functionality. For example, to identify the iModulons, the iModulon generator 608 can apply ICA to the expression profile vectors. The iModulon generator 608 can do so using the FastICA algorithm, in some cases. For example, the iModulon generator 608 can apply ICA to the expression profile vectors by first organizing the vectors into a data matrix X, where each Att Dkt: 114198-5960 row represents a condition and each column represents a gene, or vice versa. This matrix can contain the expression levels for all genes across all experimental conditions. In some cases, prior to generating the matrix, the iModulon generator 608 can normalize the values. The iModulon generator 608 can do so, for example, by performing a z-score normalization and / or log2 transformation to the expression values. By doing so, the iModulon generator 608 can reduce biases from technical variations across experiments and / or make data comparable across different experiments and conditions. The iModulon generator 608 can decompose the matrix X into two matrices: a mixing matrix A and a source matrix S. The iModulon generator 608 can decompose the matrix X such that X is approximately equal to AS. The mixing matrix A can contain the independent components that represent distinct regulatory patterns, while S can contain the weights describing how each gene contributes to these patterns. This decomposition can identify statistically independent patterns of gene expression variation across conditions. During decomposition, the iModulon generator 608 can maximize the statistical independence between the components. The iModulon generator 608 can do so using techniques such as maximizing non-Gaussianity or minimizing mutual information. The resulting independent components can represent distinct regulatory modules, with each component corresponding to a potential iModulon (e.g., a candidate iModulon). The iModulon generator 608 can identify genes that significantly contribute to the different respective components. The iModulon generator 608 can do so, for example, by identifying the weight values in the matrix S and comparing the weight values to a threshold. The iModulon generator 608 can determine which genes are associated with each component based on the genes corresponding to weights exceeding the threshold. The iModulon generator 608 can group genes with weights exceeding this threshold as potential members of an iModulon, such as by assigning a common or identical identifier to each gene of the respective groups. The iModulon generator 608 can perform a filtering technique to remove or discard potential iModulons from memory or otherwise consideration as being an iModulon. The iModulon generator 608 can do so, for example, for each potential iModulon, by analyzing data or metadata for each gene grouped into the iModulon. The iModulon generator 608 can identify data indicating whether the gene has a known function or not. The iModulon generator 608 may identify such data by retrieving or querying the remote data source 618 for Att Dkt: 114198-5960 such data. The iModulon generator 608 can determine whether at least one gene of the iModulon has a known function and at least one gene of the iModulon has an unknown function (e.g., has an unknown function or has no known function). The iModulon generator 608 can determine whether this condition is true for each potential iModulon and discard (e.g., remove from memory or otherwise from consideration as an iModulon) any potential iModulon for which the iModulon generator 608 determines the condition is not true. In another example, the iModulon generator 608 may remove any potential iModulons that only include a single gene or a number of genes below a threshold. For example, the iModulon can instantiate or generate a counter for each iModulon. The iModulon generator 608 can increment each counter for each gene in the iModulon. The iModulon generator 608 can compare the counts of the counters to a threshold (e.g., 1.5). Responsive to determining the number of genes in a potential iModulon exceeds the threshold, the iModulon generator 608 can remove or discard the potential iModulon. The iModulon generator 608 can periodically use such filters, and any number of other filters, to remove or discard potential iModulons. In some cases, the iModulon generator 608 can validate the statistical significance of each potential iModulon after the filtering using various metrics. For example, the iModulon generator 608 can calculate internal consistency metrics including mean pairwise correlation scores, explained variance ratios, and internal clustering coefficients for gene expression patterns within each potential iModulon. The iModulon generator 608 can compare such metrics for each potential iModulon to one or more thresholds (e.g., a threshold specific to the metric of the comparison or the same threshold for each metric) to determine whether the threshold is satisfied (e.g., exceeded or beneath, depending on the metric and / or the configuration of the iModulon generator 608). In one example, the iModulon generator 608 can determine whether each potential iModulon has a mutual information score less than 0.1, a cross-correlation coefficient less than 0.15, and / or a principal angle greater than 60 degrees. The iModulon generator 608 can remove or discard any potential iModulons where at least one of these criteria is not met. In another example, the iModulon generator 608 can perform a robustness validation using bootstrap analysis to remove or discard any potential iModulons with a minimum 95% confidence threshold and / or a cross-validation stability score below 0.8. Doing so can ensure reproducibility of the identified iModulons. The iModulon generator 608 can identify or output the filtered and / or validated iModulons. Each filtered and / or validated iModulon can be represented as a set of genes with Att Dkt: 114198-5960 their corresponding weights from the ICA decomposition. The weights can indicate the strength and direction of each gene's contribution to the iModulon's function, providing insight into the relative importance of both known and unknown genes within each functional module. The iModulon generator 608 can store the validated iModulons and / or the corresponding in the data repository as iModulons 616. In some cases, the data processing system can store an association between the genes of the iModulon and / or any other data or metadata about the genes and / or iModulon in memory. In doing so, the data processing system can map (e.g., assign) the function (e.g., molecular function or genetic function) of the genes with the known function of the iModulon to the grouped genes of the iModulon. In some embodiments, in one example, the data processing system can determine (e.g., using a mapping table of genetic functions to metabolic functions) a metabolic function for the iModulon based on the genetic functions of the genes with the known genetic functions of the iModulon and assign the metabolic function to the iModulon. The data processing system can repeat this process for any number of potential iModulons. In some cases, the data processing system 602 can transmit the identified iModulons 616 and / or any data (e.g., the functionality of the genes of the iModulon) to the client device 620. FIG. 7 is an illustration of an example method 700 for iModulon identification and cell functionality modification, in accordance with implementations. One or more, or all, of the operations of the method 700 can be performed by one or more systems or components depicted in FIG. 6 including, for example, the data processing system 602. Performance of the method 700 can facilitate the transfer of genes of an unknown function into gene transfer vehicles of an organism to change the functionality of the gene transfer vehicles or otherwise provide the gene transfer vehicles with new functionality. At operation 702, one or more iModulons can be determined. iModulons can be or include groups of genes that each include at least one gene with an unknown function (e.g., an unknown cellular function or unknown molecular function) and at least one gene with a known function (e.g., a known cellular function or a known molecular function). Each iModulon can correspond with a particular function (e.g., metabolic function) to which the different genes of the iModulon contribute. The data processing system can determine the iModulons using independent component analysis techniques. For example, the data processing system can generate transcriptomic data of one or more organisms including individual expression profiles (e.g., vectors with expression levels at different index values) for different conditions. The data processing system can organize Att Dkt: 114198-5960 the expression profiles into a matrix X, where different rows represent different conditions and different columns represent different genes, or vice versa. The data processing system can decompose the matrix into a mixing matrix and a source matrix. The mixing matrix can contain independent components representing different regulatory patterns, while the source matrix can contain weights indicating an impact of genes’ contribution to the patterns. The data processing system can perform the decomposition to maximize the statistical independence between components, such as by maximizing non-Gaussianity or minimizing mutual information, with each component representing a potential iModulon. For example, the data processing system can decompose the matrix X into a mixing matrix A and a source matrix S. The matrix X can be or include an m x n matrix with m rows in which m corresponds to experimental conditions, n corresponds to genes, and each entry Xij is an expression level of gene j in condition i. Matrix A can be or include an m x k matrix in which k is the number of independent components, each column can be a regulatory pattern across conditions, and each entry Aij is a strength of component j in condition i. Matrix S can be or include a k x n matrix in which each row corresponds to weights for genes in a component and each entry Sijis a contribution of gene j to component i. The data processing system can identify significant genes by comparing the weight values in the source matrix to a threshold. The data processing system can group genes that exceed the threshold (e.g., exceed the threshold for a particular regulatory pattern or component) into potential iModulons. The data processing system can then apply filtering techniques to the potential iModulons, removing potential iModulons that don't contain genes with both known and unknown functions and / or potential iModulons with genes below a threshold number. The data processing system can validate the remaining iModulons using metrics such as mean pairwise correlation scores, explained variance ratios, and internal clustering coefficients. The data processing system can compare the metrics to one or more thresholds. The data processing system can remove or discard any potential iModulons with at least one metric that does not satisfy a threshold based on the comparisons. The data processing system can determine, identify, or output the remaining iModulons after the filtering and / or the validation. The output iModulons can be or include sets of genes with their corresponding weights from the ICA decomposition. The weights can indicate the strength and / or direction of each gene's contribution to the iModulon's function. The validated iModulons can be stored in memory. The data processing system can transmit the output iModulons to a client device automatically or in response to a request Att Dkt: 114198-5960 from the client device. In some cases, the data processing system can determine the iModulons in response to a request from the client device and automatically transmit the iModulons to the client device in response to determining the iModulons. In a non-limiting example, the data processing system can analyze gene expression data from E. coli grown under various nutrient conditions. For instance, the data processing system can identify or generate expression profiles from cells grown in media with and without leucine. The data processing system can create or generate a matrix X where each row represents a different growth condition (e.g., rich media, minimal media, stress conditions, etc.) and each column represents individual genes, such as leuA, which is a known leucine synthesis gene, and yibT, which has an unknown function or does not have a known function. For instance, one or more studies can include of 100 experimental conditions across 1000 genes. The data processing system can collect expression data from these studies and generate the matrix X to include 100 rows and 1000 columns. The data processing system can decompose the matrix X into a mixing matrix A and a source matrix S. The matrix X can be or include an m x n matrix with m rows in which m corresponds to experimental conditions, n corresponds to genes, and each entry Xijis an expression level of gene j in condition i. A can be or include an m x k matrix in which k is the number of independent components, each column can be a regulatory pattern across conditions, and each entry Aijis a strength of component j in condition i. Matrix S can be or include a k x n matrix in which each row corresponds to weights for genes in a component and each entry Sij is a contribution of gene j to component i. The data processing system can identify one or more patterns from the decomposed matrices. For instance, the matrix X may include profiles for the following three conditions: growth in leucine-rich media, growth in leucine-depleted media, and growth under heat stress. The data processing system can identify a component from the matrix A and / or the matrix S where leuA has a weight of .85 (e.g., a strong positive above a threshold of .7), yibT has a weight of .78 (e.g., a strong positive above the threshold of .7), and hspB has a weight of .12 (e.g., a weak correlation below the threshold of .7). Based on the weights’ relationship with the threshold of .7 (e.g., whether the weights are above or below the threshold of .7), the data processing system can determine leuA and yibT are a part of the same potential iModulon (e.g., an iModulon involved in leucine metabolism), while hspB may belong to a different potential iModulon. Att Dkt: 114198-5960 The decomposition can reveal that leuA and yibT are strongly co-expressed under leucine-related conditions, while hspB (heat shock protein) shows minimal correlation. The data processing system can use these weights to establish that leuA and yibT likely belong to the same iModulon involved in leucine metabolism, while hspB may belong to a different regulatory network. The data processing system can validate the potential iModulon by confirming the iModulon has at least a threshold number of genes with a known function (e.g., leuA) and at least a threshold number of genes with an unknown function or not a known function (e.g., yibT). The data processing system can additionally verify the statistical significance through correlation scores. If the genes show strong correlation (greater than 0.8) and meet other statistical thresholds (e.g., a mutual information score less than 0.1, a cross-correlation coefficient less than 0.15, and / or a principal angle greater than 60 degrees), the data processing system can determine the potential iModulon is an iModulon. Responsive to determining the potential iModulon is an iModulon, the data processing system can generate a notification or message including or identifying the iModulon (e.g., identifying each of the genes in the iModulon and / or any other data about the iModulon, such as the function of the known genes of the iModulon). In some cases, the data processing system can include weights or expression weights of the genes of the iModulon in the message. In some cases, the data processing system can include the known functionality of the genes in the message, which can include assigning (e.g., mapping) the known functions (e.g., molecular functions and / or cellular function) of the genes of the iModulon to the iModulon. In some cases, the data processing system can store an association between the genes of the iModulon and / or any other data or metadata about the genes and / or iModulon in memory. In doing so, the data processing system can map (e.g., assign) the function of the genes to the known function of the iModulon. The data processing system can repeat this process for any number of potential iModulons. The data processing system can include each iModulon that passed the filtering and verification steps in a message or in separate messages and then transmit the message or messages to client devices of one or more researchers or individuals. At operation 704, a strain of phenotypic expression can be made for the one or more iModulons by refactoring the one or more iModulons comprising one or more genes of unidentified function and one or more genes of a known function into a gene transfer vehicle. The gene transfer vehicle can be or include circularized DNA, a plasmid, a viral Att Dkt: 114198-5960 vector, a bacterial expression vector pCAP_BAC, a BAC, etc. For instance, individuals accessing the client devices can view the message containing the identification of the iModulon (e.g., identifications of the genes of the iModulon), and select genes comprising the iModulon for refactoring into the gene transfer vehicle. If selected genes are found in a single operon, the entire operon can be directly amplified and cloned into the gene transfer vehicle. If the selected genes are located in multiple genomic loci, they can be refactored into an operon, while preserving the genetic order to ensure an optimal expression level, and the refactored operon can be amplified and cloned into the gene transfer vehicle. In some cases, an artificial chromosome encoding an iModulon can be constructed by assembling amplified DNA fragments comprising the iModulon and a promoter. In one example, a bacterial artificial chromosome (BAC) encoding an iModulon can be constructed by assembling PCR amplified pCAP_BAC (Addgene plasmid #120229) fragments, nptII promoter, pheS-nptII, thr, ilvG, ilvMEDA, ilvC, and leu loci using transformation-associated recombination (TAR) cloning in Saccharomyces cerevisiae strain VL6-48 (ATCC MYA- 3666; MATα his3-Δ200 trp1-Δ1 ura3−52 lys2 ade2−1 met14 cir0). In some cases, successful clones can be screened in or on a medium. In one example, successful clones can be screened on solid synthetic His dropout medium. In some cases, a synthetic DNA fragment can be integrated into a recipient cell genome to allow for precise gene-insertion. In one example, a spacer array targeting both the landing pad and the BAC can be constructed by annealing of two primers RX4_1_F and RX_4_1_R followed by primer extension using RX4_2_F and RX4_2_R. In one example, the spacer array can be cloned in a pUC plasmid containing crRNA leader sequence. In some cases, iModulons can be PCR amplified from bacterial genomic DNA. In one example, DNA fragments have 15-20 nt homology to pTrcHis2A (Invitrogen, V36520) plasmid and can be cloned into the linearized plasmid backbone using In-Fusion Cloning Kit (Takara Bio, 638948). In some cases, iModulons can be amplified as different fragments from bacterial genomic DNA. In one example, assembled BAC can be extracted using Gentra Puregene Yeast / Bact. Kit (Qiagen, 158567). In some cases, a refactored iModulon can be introduced into a genome using the replicon excision enhanced recombination (REXER) method with or without modifications. In some cases, the iModulon can be introduced into a bacteria by transformation. In one example, a plasmid containing an iModulon can be introduced using electroporation to E. coli ΔBCAA strain carrying pREDCas9 plasmid and sacB-cat dual-selection landing pad upstream of thrL gene. In some cases, a recipient cell can be screened for incorporation of Att Dkt: 114198-5960 the iModulon with antibiotics. In one example, a recipient cell can be grown in terrific broth (TB) containing 25 µg / ml chloramphenicol, 50 µg / ml kanamycin, 100 µg / ml carbenicillin, and 30 mM L-arabinose at 37°C to cell density (A600) of 0.8. In some cases, a recipient cell can be screened for genomic incorporation of the iModulon with antibiotics. In one example, two micrograms of PCR-amplified linear guide-RNA construct, targeting both the landing pad and the BAC, can be introduced using electroporation and cells can be screened on NaCl- free LB-agar plate (10 g / l tryptone, 5 g / l yeast extract, and 1.5% agar) containing 100 µg / ml carbenicillin, 50 µg / ml kanamycin, and 10% sucrose. In one example, successful recombinants can be further screened by chloramphenicol sensitivity and PCR genotyping. At operation 706, optionally, an adaptive laboratory evolution algorithm can be applied to the strain of the phenotypic expression to create an optimized strain of phenotypic expression for performing a function. In some cases, recipient cells are evolved via serial propagation in a medium comprising a metabolite. In one example, E. coli K-12 MG1655 carrying pMdcR_iM plasmid can be evolved via serial propagation of 150 µl into 15 ml M9 malonate (2 g / l) minimal medium containing 100 µg / ml carbenicillin in three biologically replicated cultures. In one example, cultures can be incubated at 37°C, aerated by magnetic stirring, and reinoculated at A600 of 0.3 (Tecan Sunrise microplate reader; equivalent to an A600 of 0.750 on a conventional spectrophotometer with a path-length of 10 mm) using an automated system. In some cases, recipient cells are screened for improved utilization of the metabolite. In one example, after 40 days of propagation, evolved cultures can be subjected to malonate utilization assay to isolate clones from a culture. At operation 708, Genomic DNA of the optimized strain of phenotypic expression can be sequenced to determine mutations for the optimized strain of phenotypic expression. Genomic DNA can be isolated and sequenced using any suitable technique known in the art. In one example, genomic DNA can be isolated using Quick-DNA Miniprep Kit as instructed by the manufacturer. In one example, whole-genome DNA-seq libraries can be generated with a NEBNext Ultra II DNA Library Prep Kit for Illumina (New England Biolabs, E7645) and run on an Illumina NovaSeq X Plus with 100 cycles pair-ended recipe. The sequencing data can be aligned to a reference genome. In one example, the sequencing results can be processed with the pipeline that incorporates BreSeq pipeline to identify mutations. In one example, the E. coli K-12 MG1655 genome sequence (National Center for Biotechnology Information accession no. NC_000913.3) can be used as a reference sequence. In some embodiments, mutations are identified as nucleotides of the sequencing data that differ from Att Dkt: 114198-5960 nucleotides of the aligned reference genome using any suitable technique known in the art. The data processing system can generate a record (e.g., a table, data structure, user interface, entry in a table, etc.) identifying the mutations and identifications of each gene and / or a position within the sequencing data of the optimized strain of phenotypic expression. The record can be used to create a second cell strain of optimized phenotypic expression. For example, the record can be identified and used to generate circular DNA, a plasmid, or a viral vector comprising nucleotides of (i) the sequencing data and / or (ii) mutations identified the record. In some embodiments, the circularized DNA, plasmid, or viral vector is integrated into the genome of the second cell (e.g., bacteria) by transformation or infection, thereby creating a second cell strain. In some embodiments, the cell type of the second cell strain is a different type than the cell of the (first) optimized strain of phenotypic expression. In other embodiments, the cell type of the second cell strain is the same type as the cell of the (first) optimized strain of phenotypic expression. In some embodiments, the cell of the second cell strain may be a strain comprising an iModulon. For example, an E. coli strain comprising an iModulon and / or other genomic feature may be transformed with a plasmid comprising nucleotides of the record to create a new second cell strain of optimized phenotypic expression. In a non-limiting example, the data processing system can generate transcriptomic data by subjecting E. coli to 50 distinct growth conditions, including variations in temperature (25°C to 42°C), nutrient availability (minimal to rich media), and environmental stressors (pH 5.5 to 8.5). Using RNA sequencing, data processing system can measure expression levels for 4000 genes across all conditions, creating an expression matrix that captures diverse cellular responses. Through ICA, the data processing system can decompose the expression matrix into its constituent components, identifying an iModulon containing both characterized genes (e.g., leuA and leuB, known for leucine biosynthesis) and uncharacterized genes (e.g., yibT and yggX). For example, these genes demonstrate coordinated expression patterns across multiple conditions, particularly showing strong correlation during amino acid starvation responses. To create a phenotypic strain, the data processing system can design a gene construct that incorporates identified genes (e.g., leuA, leuB, yibT, yggX) along with their native promoter sequences and standardized ribosome binding sites. This construct can be Att Dkt: 114198-5960 assembled into a plasmid, e.g., the pCAP_BAC vector, using transformation-associated recombination, for example, in Saccharomyces cerevisiae, followed by transformation into E. coli strain ΔBCAA. The transformed strain can undergo optimization through successive rounds of growth selection in minimal media lacking a metabolite, e.g., leucine. The data processing system monitors growth rates and metabolite production, selecting colonies showing enhanced growth characteristics for further analysis. The data processing system can compare sequencing data of both the optimized strain and the pCAP_BAC construct. This reveals key mutations, for example a two-fold increase in yibT promoter strength and a single nucleotide variant in the yggX coding sequence that correlates with improved strain performance. These mutations provide insight into the functional roles of the previously uncharacterized genes, e.g., within the leucine biosynthesis pathway. The sequencing data and / or the mutations can later be used to create a new strain to transfer into a gene transfer vehicle of an organism to provide the organism with the functionality (e.g., genetic functionality) of the strain. Advantageously, by implementing the systems and methods described herein, the data processing system can systematically identify and characterize previously unknown gene functions through a combination transcriptome analysis and strain engineering. The method implements independent component analysis (ICA) to detect statistically independent patterns of gene co-expression across diverse conditions, facilitating the identification of functional relationships between characterized and uncharacterized genes that may be obscured in traditional correlation-based analyses. Using ICA instead of other approaches can facilitate parsing complex, high-dimensional transcriptomic data to reveal subtle but biologically meaningful associations between genes, providing a framework for identifying the functions of previously uncharacterized genes. Furthermore, the method provides a robust pathway to strain engineering and optimization. By refactoring the identified genes into a gene transfer vehicle and creating phenotypic strains, the direct testing of the predicted functional relationships can be facilitated. The subsequent sequencing and mutation analysis of optimized strains can identify the molecular roles of the uncharacterized genes. This integrated approach can accelerate the process of gene function identification of the mutations that enhance strain performance. Att Dkt: 114198-5960 FIG. 8 is a block diagram of an example computer system 800. The computer system or computing device 800 can include or be used to implement the system 600 or their components such as the data processing system 602. The computing system 800 includes a bus 805 or other communication component for communicating information and a processor 810 or processing circuit coupled to the bus 805 for processing information. The computing system 800 can also include one or more processors 810 or processing circuits coupled to the bus for processing information. The computing system 800 also includes main memory 815, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 805 for storing information, and instructions to be executed by the processor 810. The main memory 815 can be or include the data repository 612. The main memory 815 can also be used for storing position information, temporary variables, or other intermediate information during execution of instructions by the processor 810. The computing system 800 may further include a read only memory (ROM) 820 or other static storage device coupled to the bus 805 for storing static information and instructions for the processor 810. A storage device 825, such as a solid state device, magnetic disk or optical disk, can be coupled to the bus 805 to persistently store information and instructions. The storage device 825 can include or be part of the data repository 612. The computing system 800 may be coupled via the bus 805 to a display 835, such as a liquid crystal display, or active matrix display, for displaying information to a user. An input device 830, such as a keyboard including alphanumeric and other keys, may be coupled to the bus 805 for communicating information and command selections to the processor 810. The input device 830 can include a touch screen display 835. The input device 830 can also include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 810 and for controlling cursor movement on the display 835. The display 835 can be part of the data processing system 602, the client device 620 or other component of FIG. 6, for example. The processes, systems and methods described herein can be implemented by the computing system 800 in response to the processor 810 executing an arrangement of instructions contained in main memory 815. Such instructions can be read into main memory 815 from another computer-readable medium, such as the storage device 825. Execution of the arrangement of instructions contained in main memory 815 causes the computing system 800 to perform the illustrative processes described herein. One or more processors in a multi- processing arrangement may also be employed to execute the instructions contained in main Att Dkt: 114198-5960 memory 815. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software. Although an example computing system has been described in FIG. 8, the subject matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more circuits of computer program instructions, encoded on one or more computer storage media for execution by, or to control the operation of, data processing apparatuses. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. While a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate components or media (e.g., multiple CDs, disks, or other storage devices). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources. The terms “data processing system” “computing device” “component” or “data processing apparatus” encompass various apparatuses, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC Att Dkt: 114198-5960 (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures. For example, the data collector 604, the profile generator 606, and the iModulon generator 608 data and other data processing system 602 components can include or share one or more data processing apparatuses, systems, computing devices, or processors. A computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can correspond to a file in a file system. A computer program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network. The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs (e.g., components of the data processing system 602) to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatuses can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Att Dkt: 114198-5960 The subject matter described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or a combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks). The computing system such as the system 600 can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network (e.g., the network 601). The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client- server relationship to each other. In some implementations, a server transmits data (e.g., data packets representing a digital component) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server (e.g., received by the data processing system 602 from the client device 602). While operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. Actions described herein can be performed in a different order. The separation of various system components does not require separation in all implementations, and the described program components can be included in a single hardware or software product. For example, the data collector 604, the profile generator 606, and the iModulon generator 608 can be a single component, app, or program, or a logic device having one or more processing circuits, or part of one or more servers of the data processing system 602. Having now described some illustrative implementations, it is apparent that the foregoing is illustrative and not limiting, having been provided by way of example. In Att Dkt: 114198-5960 particular, although many of the examples presented herein involve specific combinations of method acts or system elements, those acts and those elements may be combined in other ways to accomplish the same objectives. Acts, elements and features discussed in connection with one implementation are not intended to be excluded from a similar role in other implementations or implementations. The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including” “comprising” “having” “containing” “involving” “characterized by” “characterized in that” and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components. Any references to implementations or elements or acts of the systems and methods herein referred to in the singular may also embrace implementations including a plurality of these elements, and any references in plural to any implementation or element or act herein may also embrace implementations including only a single element. References in the singular or plural form are not intended to limit the presently disclosed systems or methods, their components, acts, or elements to single or plural configurations. References to any act or element being based on any information, act or element may include implementations where the act or element is based at least in part on any information, act, or element. Any implementation disclosed herein may be combined with any other implementation or embodiment, and references to “an implementation,” “some implementations,” “one implementation” or the like are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation or embodiment. Such terms as used herein are not necessarily all referring to the same implementation. Any implementation may be combined with any other implementation, inclusively or exclusively, in any manner consistent with the aspects and implementations disclosed herein. References to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all of the described terms. References Att Dkt: 114198-5960 to at least one of a conjunctive list of terms may be construed as an inclusive OR to indicate any of a single, more than one, and all of the described terms. For example, a reference to “at least one of ‘A’ and ‘B’” can include only ‘A’, only ‘B’, as well as both ‘A’ and ‘B’. Such references used in conjunction with “comprising” or other open terminology can include additional items. Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included to increase the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any claim elements. The systems and methods described herein may be embodied in other specific forms without departing from the characteristics thereof. The foregoing implementations are illustrative rather than limiting of the described systems and methods. Scope of the systems and methods described herein is thus indicated by the appended claims, rather than the foregoing description, and changes that come within the meaning and range of equivalency of the claims are embraced therein. Equivalents It is to be understood that while the disclosure has been described in conjunction with the above embodiments, that the foregoing description and examples are intended to illustrate and not limit the scope of the disclosure. Other aspects, advantages and modifications within the scope of the disclosure will be apparent to those skilled in the art to which the disclosure pertains. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All nucleotide sequences provided herein are presented in the 5′ to 3′ direction. The embodiments illustratively described herein may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, for example, the terms “comprising”, “including,” containing”, etc. shall be read expansively and without limitation. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the disclosure. Att Dkt: 114198-5960 Thus, it should be understood that although the present disclosure has been specifically disclosed by specific embodiments and optional features, modification, improvement and variation of the embodiments therein herein disclosed may be resorted to by those skilled in the art, and that such modifications, improvements and variations are considered to be within the scope of this disclosure. The materials, methods, and examples provided here are representative of particular embodiments, are exemplary, and are not intended as limitations on the scope of the disclosure. The scope of the disclosure has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the disclosure. This includes the generic description with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein. In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that embodiments of the disclosure may also thereby be described in terms of any individual member or subgroup of members of the Markush group. All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety, to the same extent as if each were incorporated by reference individually. In case of conflict, the present specification, including definitions, will control. Sequence Listing SEQ ID Nos: 1-76 vanA – SEQ ID NO: 1 ATGTACCCCAAAAACACCTGGTACGTCGCCTGCACCCCCGATGAGATCGCCACCA AACCCCTGGGCCGGCAGATCTGCGGGGAAAAAATCGTGTTCTACCGCGCCCGCG AGAACCAAGTAGCCGCCGTCGAGGACTTCTGCCCGCACCGCGGCGCACCGTTGT CGTTGGGCTATGTCGAGGACGGCAACCTGGTGTGCGGCTACCACGGCCTGGTGAT GGGTTGCGACGGCAAGACCGTGTCGATGCCGGGCCAACGGGTGCGTGGCTTCCC CTGCAACAAGACCTTTGCGGCCGTCGAGCGCTATGGCTTCATCTGGGTCTGGCCC GGTGACCAGGCGCAGGCCGACCCGGCGCTGATTCCGCATCTGGAATGGGCGGTG AGTGATGAGTGGGCCTACGGCGGCGGGCTGTTCCACATCGGTTGCGACTACCGCC TGATGATCGACAACCTCATGGACCTCACCCATGAAACCTATGTGCACGCCTCCAG CATCGGCCAGAAGGAGATCGACGAGGCACCGCCGGTCACCACCGTCACCGGCGA CGAAGTGGTCACCGCCCGGCACATGGAAAACATCATGGCGCCACCGTTCTGGCG Att Dkt: 114198-5960 CATGGCCTTGCGTGGCAATGGCCTGGCCGACGATGTACCAGTGGACCGCTGGCA GATCTGCCGTTTCACCCCACCTAGCCATGTGCTGATCGAAGTGGGTGTAGCGCAT GCCGGCAAGGGCGGCTACCACGCCGAGGCACAGCATAAGGCGTCGAGCATCGTG GTCGACTTCATCACCCCTGAGAGCGATACCTCTATCTGGTACTTCTGGGGCATGG CGCGCAACTTCGCTGCGCACGACCAGACCCTGACCGACAACATTCGTGAGGGCC AGGGCAAGATTTTCAGCGAAGACCTGGAAATGCTCGAACGCCAGCAGCAGAACC TGCTGGCCCACCCCGAGCGCAACTTGCTGAAGCTGAATATCGACGCCGGCGGCG TGCAGTCACGCAAAGTGCTGGAGCGGATCATCGCCCAAGAGCGTGCGCCGCAGC CGCAACTGATCGCCACCAGCGCCAACCCTGCCTGA vanB – SEQ ID NO: 2 ATGATCGATGCCGTAGTGGTATCCCGTAACGATGAAGCGCAGGGTATCTGCAGCT TCGAGCTGGCCGCGGCAGATGGCAGCCTGCTGCCGGCGTTCAGCGCCGGCGCCC ATATCGACGTGCACCTGCCCGACGGGCTGGTGCGCCAGTATTCGCTGTGCAACCA CCCCGAAGAACGCCATCGCTATCTGATTGGCGTACTCAACGACCCGGCTTCGCGG GGCGGTTCTCGTAGCCTGCACGAACAGGTGCAAGCCGGTGCCCGGCTGCGTATC AGTGCGCCGCGCAACCTGTTCCCGCTGGCCGAGGGTGCGCAGCGCAGTTTGCTGT TTGCTGGCGGTATCGGCATTACCCCAATCCTGTGCATGGCCGAGCAGCTGTCCGA CAGCGGCCAGGCCTTCGAGCTGCACTACTGTGCCCGCTCCAGCGAGCGTGCGGC GTTTGTCGAGCGGATCCGCAGCGCGCCGTTCGCTGATCGGCTGTTCGTGCATTTT GACGAGCAGCCGGAAACGGCGCTGGACATCGCCCAGGTGCTGGGCAACCCGCAA GATGATGTGCACCTGTATGTATGCGGGCCCGGCGGGTTCATGCAGCATGTGCTGG ACAGCGCGAAGGGGCTGGGCTGGCAGGAGGCCAACCTGCACCGCGAGTACTTCG CCGCAGCACCGGTGGATGCCAGCAACGATGGCAGTTTCGCGGTGCAGGTGGGCA GCACGGGACAGGTGTTCGAGGTGCCAGCCGACCGGACCGTGGTGCAGGTGCTGG AAGAGAATGGTATCGAGATCGCCATGTCGTGCGAGCAGGGTATTTGCGGCACCT GCCTGACACGCGTGCTGCAGGGCACACCGGACCATCGCGATCTGTTTCTCACCGA AGAGGAACAGGCCCTGAACGATCAGTTCACGCCCTGCTGCTCGCGCTCGAAGAC GCCGCTGCTGGTGCTGGACATCTGA vanK– SEQ ID NO: 3 ATGAATACCGATCTCAAGCAACGTCTGGACAACGACCCGATGGGTCGCTTTCAGT GCCTGGCCATTGGTATTTGCATCATCCTCAACATGATCGATGGCTTCGACGTGCT GGTCATGGCCTTCACCGCCGCCTCGGTGTCTGCCGAGTGGGGCCTGAATGGGGCG CAGGTTGGGCTTCTGCTCAGTGCCGGCTTGTTCGGTATGGCGGCCGGGTCATTGT TCATCGCGCCGTGGGCTGACCGCTTTGGCCGACGTCCGCTGATCCTGCTGTGTCT GGCGCTGTCCGGTATCGGCATGCTGCTCTCTGCACTGAGCCAGAGCCCGCTGCAA CTGGCGCTGCTGCGCGGCCTGACCGGGCTGGGCATCGGTGGCATCCTGGCCAGC AGCAATGTGATTGCCAGTGAATATGCCAGCGAACGCTGGCGCGGCTTGGCGGTG AGCCTGCAGTCCACTGGCTATGCGCTGGGTGCGACCCTTGGCGGTTTGCTGGCGG TATGGCTGCTGGGGCATTGGGGCTGGCGATCGGTATTTCTGTTTGGCGGCGTCGT CACCGTACTGGTGATCCCGCTGGTACTGCTGTGGTTGCCAGAGTCGCTGGACTTT CTGCTGGCACGCCGGCCGGCCAATGCTCTGGCCCGGGTCAACCGCCTGGCCGTGC Att Dkt: 114198-5960 GCCTTGGCCAGCCGGCATTGGCGCAACTACCGGCTGTTGCCGCAAGGGCAGAGG GTGCCAAAACCGGGTTCAGGCAGCTGCTGGCGCCGGCCATGCGCCGCACCACGC TGGTGATCTGGTTGCTGTTCTTCCTGGTCATGTTCGGCTTCTACTTCGTGATGAGC TGGACGCCGAAACTGTTGGTGGCCGCCGGCTTGTCGGCCCAGCAAGGCATCACC GGCGGCGTGCTGTTGAGCGTGGGCGGCATCGTTGGCGCGGCGCTGATCGGCGGC CTGGCTTCGCGCTGGCCACTGACTCGCGTGCTGGCGTTGTTCATGCTCATCACCG CCGCGCTGCTGGTGCTGTTCGTGGCCAATGGTGCCTCGGTCACTGCGGCGCTGGC GTTGGGCCTGCTGATCGGGCTGTTTTCCAACGGCTGCGTGGCTGGGCTGTATGCC TTGTCGCCGGTGGTCTATGACGCATCGGTACGCGCCACCGGTGTGGGTTGGGGGA TCGGCATCGGCCGCATGGGCGCGATTCTGTCGCCGACCGTGGCTGGCGTGCTGCT CGATGGTGGCTGGCAACCCCTGCATCTGTATGGCGTGTTCGCCAGCGTGTTCGTG CTTGCCGCAGGTTGCCTGCTGCTGTTGCGCCCAACGGCCAAGCCTGCGTCTGCCT TGCCTGCTGATGCACTGAGCCACTGA galP-IV– SEQ ID NO: 4 ATGATGGACACAACTTGCACGCTGGCCCGTTCGGTCCTGGCTGTTTCGCTGGCGT TTGGCCTTGGTTTGCCGGCCTGGGCGCAAGAGGGGGGCTTCTTGGAAGATGCCAA GGGTGGGCTGACATTGCGCAACTACTACTTCAACCGCGATTTCCGCGACCCGGGG GCCGCCAAGGGCAAGGTTGAAGAGTGGGCCCAGGGTTTCATCTTCAAGTTCAGC TCTGGCTACACCCCGGGGCTGATCGGTTTTGGCCTGGATGGTATCGCCCTGTTCG GCGTGAAGCTGGACAGCGGCGGCGGTACCAGCGGCTCAGAACTGTTGCCGGTGC ACGATGACGGGCGAGCGGCCGACGACTACGGTCGGGCCGGGCTGGCGGCGAAG CTGCGTATGTCCGCGACCGAGCTGAAGGTGGGGGAGTTGTTGCCGGATATTCCGC TGTTGCGTTACGACGATGGCCGCCTGCTGCCGCAGACCTTCCGTGGTGCCATGCT CGACTCGCGGGAAATCGAAGGCCTGAGCCTGCAGGCCGGCCAGTACCGCGAGGT GAGCCTGCGCCACTCTTCGGACATGCAGGACCTATCTGCCTGGGCCGCGCCCGGG GTCACCTCGGACGGCTTCAACTACGCCGGCGCCGAGTACCGCTTCAACCAGCAGC GAACACTGATCGGCGCCTGGCACTCGCAATTGGAGGACATCTACCAGCAAAGCT ACTTCAACCTGCTGCACAGTCAGCGGGTAGGGGAGTGGACGTTGGGGGCCAACC TTGGCTACTTCATCGACCGGGATGACGGACAGGCGCGCATCGGCGACATCCAGA GCCGCACCGGTTACGCTTTGCTGTCAGCTTCGCACAGTGGCCACACGCTGTACCT GGGGTTGCAGAAGGTCAGTGGCGACAGTCAGTGGATGTCGGTTTACGGCAGCAG TGGGCGAACCCTGGGCAACGACATGTTCAACGGTAACTTCAGTAATGCTGACGA GCGTTCGTGGCAAGTGCGTTACGACTACAACTTTGCCGCGCTCGGTGTGCCAGGG CTGCTGGCCATGGTGCGGTACGGGCATGGCGAGAACGCCACCACGGCTGCGGGC AGCAATGGCAAGGAGTGGGAGCGCGATGCCGAGGTGAACTACACCTTCCAGAGC GGGGCTTTGAAGAACCTCAACATCCGGCTGAACAACGCGACCAACCGGCGCAGC Att Dkt: 114198-5960 TTCAACAGTGATCTCGACCAGACGCGCCTGATCGTCAGCTATCCGCTGACCTTGT AG mdcA– SEQ ID NO: 5 ATGACGACACCGATCTCTCCCCCGCCGCAGTGGTCGCGGCGCCGCCAGGAAAAG CAGCGGCGCCTCGAGCGCGTGCGCGGCCTGGCCGACGGCGCGGTGCTGCCGCGC GAAGGCCTGGTCGCCGCGCTGGAGGCGCTGATCGCGCCGGGCGATCGGGTGGTG CTGGAAGGCAACAACCAGAAGCAGGCGGATTTCCTCTCCCGTTCGCTGGCGCGG GTCGATCCCGGCAAGCTGCACGACCTGCACATGATCATGCCCAGCGTCGGCCGCC CGGAGCATCTCGACCTGTTCGAACTGGGCATCGCGCGCAAGCTCGACTTCTCCTT CTCCGGCCCGCAGAGCTTGCGCATCGGCCAGTTGCTCGAAGACGGCCTGCTGGA GATCGGCGCGATCCACACCTACATCGAGCTGTACGCGCGGCTGGTGGTCGACCTG ATCCCCAACGTGGCGCTGGTCGCCGGTTTCGTCGCCGACCGCGAAGGCAACGTGT ACACCGGGCCGAGCACGGAAGACACCCCGGCGCTGGTGGAGCCCACCGCGTTCA GCGACGGCATCGTCATCGTCCAGGTCAACCGTATCGTCGACGACCCGCGCGACCT GCCGCGAGTGGATATCCCGGCGTCCTGGGTGGACTTCGTGGTCGAGGCCGACCA GCCGTTCTACATCGAGCCGCTGTTCACCCGCGACCCGCGCCACATCAAGCCGGTC CACGTGCTGATGGCGATGATGGCGATCCGCGGCATCTACCAGCGGCACAACGTG CAGTCGCTGAACCACGGCATCGGCTTCAACACCGCGGCCATCGAGCTGATCCTGC CGACCTACGGCGAGTCCCTCGGCCTGAAGGGCAAGATCTGCCGACACTGGACCC TCAACCCGCACCCGACCCTTATCCCCGCCATCGAGAGCGGCTGGGTCGAGAGCGT GCACTGCTTCGGCACCGAGCTGGGCATGGAGGGCTACATCGCCCAGCGTCCCGA CGTGTTCTTCACCGGTCGCGACGGCTCGCTGCGTTCCAATCGGATGTTCTGCCAA CTGGCAGGACAATACGCCGTCGACCTGTTCATCGGCGCTACCCTCCAGGTCGACG GCGACGGCCATTCCTCCACCGTCACCCGCGGACGCCTGGCCGGTTTCGGCGGCGC CCCGAACATGGGCCACGACCCGCGCGGCCGGCGCCACTCGACGCCGGCCTGGCT GGACATGCGGGGCGAACCCGAGGCCTTGCTGGAGCGCGGCAGGAAGCTGGTGGT GCAGATGGTCGAGACCTTCCAGGACGGCGGCAAGCCGACCTTCGTCGAGCGCCT CGACGCGCTGGAGGTGGCGCGCCAGACCGGCATGCCGCTGGCGCCGGTGATGAT CTATGGCGACGACGTCACCCATGTGCTCACCGAGGAAGGCATCGCCTACCTGTAC AAGGCGCGCTCCCTGGAGGAACGTCAGGCGATGATCGCGGCGGTGGCGGGGATC AGTCCGATCGGCCTGCGCCACGATCCGCGCGAGACCCAGCGGATGCGCCGCGAA GGGCTGATCGCGCTGCCCGAGGACCTCGGCATCCGCCGCACCGACGCCTCGCGC GAGCTGCTGGCGGCGAAGAGCATCGCCGAGCTGGTCGAGTGGTCCGGCGGGCTG TACCAGCCGCCGGCGCGTTTCAGGAGCTGGTGA mdcB -– SEQ ID NO: 6 ATGAACGCCATCGCGAACCTCGCGGCGACGCCGGGCGCCGACCTCGGCGAATGC CTGGCCGACCTGGCGGTGGACGCCCTCATCGACGAAGCCGAACTGTCGCCCAAG CCGGCCCTGGTGGACCGCCGCGGCAATGGCGCGCATGCCGACCTGCACCTGGGC CTGATGCAGGCCTCGGCGCTGAGCCTGTGGCCGTGTTTCAAGGAAATGGCCGAC GCCGCACAGCGGCATGGGCGCATCGATGCGCGCCTGCGCGGCGTCCTTGGCCAA Att Dkt: 114198-5960 CTGGGGCGCGACGGCGAGGCGGCGATGCTGCGTACCACCGAAGGCGTGAACACC CATCGCGGAGCGATCTGGGCGCTCGGCCTGCTGGTCGCGGCCGCCGCTCTCGAAC CGCGCCGCACGCAGGCCGGCGAAGTGGCCGCGCGCGCAGGACGCATCGCCCTGC TCGACGATCCGGCGGCCGCCATCGGCGACAGCCACGGCGAGCGCGTCCGTCGGC GCTACGGCGTCGGCGGCGCCCGCGAGGAGGCGCGTCTCTGCTTCCCCCGCGCAGT GCGCCACGGCCTGCCGCAGCTATGGCGCAGCCGCGAAGGCGGCGCTGGCGAGCA GAACGCCAGGCTCGATGCGCTGCTGGCGATCATGAGCGTCCTCGACGATACCTGC GTGCTGCACCGCGCCGGCCGCGTCGGCCTCGCCGCGATGCAGGATGGCGCCCGC GCGGTGCTCGCCGCCGGCGGCAGCGCCAGTCTCGCCGGGCGTCGCCGCCTGTGC GAGCTGGACCGCCGCCTGCTGGCGCTGAACGCTTCGCCCGGCGGCGCCGCCGAC CTGCTGGCCGCCTGTCTTTTCCTCGATCGCCTGCCCGCCGTGTCCGGTGGGTGGGC TGGGAGTCTCTGA mdcC – SEQ ID NO: 7 ATGGAAACCCTGACCTTCGAATTTCCCGCCGGCGCTCCGGCGCGCGGTCGCGCCC TGGCCGGCTGTGTCGGCTCCGGCGACCTGGAGGTGCTGCTCGAGCCGGCCGCCG GAGGCGCGCTGAGCATCCAGGTGGTGACCTCGGTAAACGGCAGCGGCCCGCGCT GGCAACAACTCTTCGCGCGGGTCTTCGCGGCGTCCACGGCGCCGGCCGCGTCGAT CCGCATCCACGATTTCGGCGCCACTCCCGGGGTGGTCCGCCTGCGCCTGGAGCAG GCACTGGAGGAGGCTGGCCATGACTGA mdcD– SEQ ID NO: 8 ATGACTGACGTCGCCCGCCTGCTGGCGCTGCGCAGCTTCACCGAACTGGGAGCGC GCCAACGCGCCCGTGCCCTGCTTGACGCGGGCAGCTTCCGCGAGTTGCTCGATCC CTTCGCCGGCGTGCAGTCGCCCTGGCTGGAACGACAGGGCATCGTGCCCCAGGC CGACGATGGGGTAGTGGTCGCCCGTGGCCTCCTCGACGGCCAGCCGGCGGTGCT CGCCGCCATCGAAGGTGCTTTCCAGGGCGGCAGCCTTGGCGAAGTCAGCGGAGC GAAGATCGCCGGCGCGCTGGAGCTGGCCGCCGAGGACAATCGCAACGGCGTTCC GACCCGGGCCCTGCTGTTGCTGGAAACCGGTGGCGTGCGCCTGCAGGAAGCCAA CCTCGGACTGGCGGCGATCGCCGAGATCCAGGCGGCCATCGTCGATCTGCAGCG CTACCAGCCGGTGGTGGCGGTGATTGCCGGGCCGGTCGGCTGCTTCGGCGGCATG TCCATCGCCGCCGGGCTGTGCAGCTACGTGCTGGTGACCCGCGAGGCGCGCCTCG GCTTGAACGGTCCGCAGGTGATCGAACAGGAGGCCGGCATCGCTGAATACGACT CGCGCGACCGGCCGTTCATCTGGAGCCTCACCGGCGGCGAGCAGCGCTTCGCCA GCGGTCTCGCCGACGCCTACCTGGCCGACGACCTCGATGAGGTACGCACCTCGGT ACTCGCGTATTTCGCCAAGGGCCTGCCGGCACGGCCGCGCTGCCGCCGCGCCGA GGATTACCTGCGGCGGCTGGGCGACCTCGATACCGCCGAGCAACCGGATGCCGC CGGCGTGCGCCGGCTGTACCAGGGGCTGGGCCAAGGAGATGCAACATGA mdcE– SEQ ID NO: 9 ATGAGCCAACCGTTCGCTTCCCGTGGCCTGGCCTGGTTCCAGGCGCTGGCCGGTA GCCTCGCGCCGCGTCCCGGCGATCCTGCTTCGTTGCGCGTCGCCGATGCCGAGCT Att Dkt: 114198-5960 GGACGGCTACCCGGTGCGTTTCCTGGCCGTCGTGCCCGATCCCGACAATCCGTTC CCGCGCGCTCGCCAGGGCGAGGTCGGGCTGCTCGAGGGCTGGGGCCTGGCCGCC GCCGTCGACGAGGCGCTGGAGGCCGACCGCGAGGCGCCACGCAAGCGCGCGCTG CTGGCCATCGTCGACGTGCCGAGCCAGGCCTACGGCCGCCGCGAGGAGGCCCTG GGTATCCACCAGGCGCTGGCCGGGGCGGTGGATGCCTACGCCCGCGCCCGACTC GCCGGGCATCCGCTGATCGGCCTGCTGGTGGGCAAGGCGATGTCCGGCGCGTTCC TCGCCCACGGCTACCAGGCCAACCGCCTGATCGCTTTGCACGATCCGGGGGTGAT GGTCCACGCCATGGGCAAGGCCGCCGCCGCGCGGATCACCCTGCGCAGCGTCGA GGAACTGGAGGCGTTGGCGGCGAAGGTGCCGCCGATGGCCTACGACATCGACAG CTACGCGTCGCTCGGCCTGCTCTGGAGGACCCTTCCGGTGGAGACCGTCGAGGTG CCGTCGACGGCGGACCTGGTGCGGGTCCGCACCTGCCTGGGCGAGGCCCTGGCC GATATCCTCGGCGGCCCGCGCGACCTCGGCGGCCGCCTCGGCGCGGCCAACCGC GAGGCGTCGGCGCGGGTGCGCCGGCTGCTGCGCGAGCAGTGGTGA mdcG– SEQ ID NO: 10 GTGGTCATCGCGAACAGCTCACAGCGGGGCTCCAGCCGGCGACGGCTCGCTCCT GATCCGATGCGTCGCAGGGCAGGGCCGGAACCCCATGACTTGCTCCTGGGCATG GCGCCGGCGCATTTGCCGGAGAATGCTCCGGCCTGGGTCTGCGCGGCACTGAAG GCCGGCTGGCCGGTGGTGGTGCGCCGTGCGCCGGCAGATCCCGCGCGGGTGGCC ATCGGCGTGCGCGGCCTGGCGCGGGAGCAGCGCTGGGCCGGCTGGATGCCGTTG GCGGCAATCACCCGGCGCTTGCGCCCGGAAGCGCTGGGCCAGCGCGAGCCGGCG ATGCTGCGCGATCTGCCGGCCTGGCGTGCCTTGCGCGACCTGCGTGCGCCGCTGA ACGCGCTGGGACTGGCCTGGGGCGTGACCGGCGGCGCCGGTTTCGAGCTGGCCA GCGGGGTCGCCGTGCTGCATCCGGACAGCGACCTCGACCTGCTGCTGCGCACCCC GCGCCCGTTCCCGCGCGATGACGCGCTGCGCCTGCTGCAGTGCTTCGAGCAGTGT CCCTGCCGGATCGACCTGCAACTGCAGACCCCGGCGGGCGGCGTGGCCCTGCGC GAATGGGCCGAGGGTCGGCCCCGGGTGCTGGCGAAAGGCAGCGAGGCGCCGCTG TTGCTGGAAGACCCCTGGCGGATCGCGGAGGTCGAGGCATGA mdcH– SEQ ID NO: 11 ATGAGCGTGCTGTTCGCCTATCCCGGCCAGGGCGCCCAACGTCCCGGCATGCTCG CGGCCTTGCCCGACGAACCGCCGGTGCGCGCCTGCCTGGAGCAGGCGGCGGACT GTCTCGGCCAGGCGCCGGCCGAACTGGAGAGCGCCGAGGCCTTGCGCGGCACCC GCGCGGTGCAGTTGTGCCTGCTGATCGCCGGGGTCGCCGCCAGCCGCCTGCTGGA GACGCGCGGGCATCGGCCGGGGCTGGTCGCCGGGCTGTCCATCGGCGCCTATCC GGCGGCGGTGGTAGCCGGCGCGCTGGACTTCGACGACGCGCTGCGTCTGGTGGC GCTGCGTGGCGAACTGATGCAGGCCGCCTGGCCCGAGGGCTACGGGATGAGCGC CATCCTCGGTCTCGAACAGGCGCAACTGGAAGCGCTGATCCTGGCGGTGCGGCG CGAGCACCCGCCGCTCTACCTGGCCAACGTCAACGCCGAGCGGCAACTGGTGGT GGCCGGCAGCGAAGCGGCTCTGGCCGCCCTCGCCGAACGCGCCCGTGCCGCCGG CGCCAGCGCGGCGAAGCGCCTGGCGGTGAGCGTACCGTCGCACTGTGCGCTGCT CGACGAACCGGCCGCGCGCCTGGCCGAAGCCTTCGCCGGCATCCGCCTGCATCG GCCGCGCGTGCCCTACCTGAGCAGCAGCCGGGCGCGCCTGGTCGCCGAACCGGC Att Dkt: 114198-5960 CGCGCTGGCCGACGATCTCGCCGGCAACATGGCGCGCCGCGTCGAGTGGCTGGC GACATTGCGCAGCGCCTACGAGCGCGGTGCCCGCCTGCATCTCGAGCTGCCGCCG GGACGGGTCCTCAGCGGTCTGGCGCGGCCTCTGTTCGGCTGCGCCACGCCGGCCT TCGAAGGCTCCCGGGCGGATACCCTGGACGCGCTGTTGCGCGAGGAGGAAAAGA GAACCCGATAA madL– SEQ ID NO: 12 ATGATTATCTACGGTGTCGCACTGCTGTCTCTGTGTACCCTCGCCGGGCTGTTCAT CGGCGAGATCCTTGGCGTGCTGCTCGGCGTGCCGGCCAACGTCGGTGGCGTCGGC ATCGCCATGCTGCTGCTGATCTTCGTCGGCAGCTACCTGGGCAAGCGTGGCCTGC TCTCGGCGAAGACCGAGCAGGGCGTGGAGTTCTGGAGCGCGGTCTACATTCCCA TCGTGGTCGCCATGGCCGCCCAGCAGAATGTCCTCGGCGCGCTGAGCGGCGGGC CGGCCGCTATCCTCGCCGGCATCGCCGCGGTCGCCGTGGCCTTCGCCCTGGTCCC GCTGCTCGACCGCTTCGGCCGGAAGAAAGCCGAAGGCGCCCGGCCATTGAAACC GCGCACCAGCGTGCAGGGGTGA madM– SEQ ID NO: 13 ATGTTCGAGTCCCTGAACAAAGTCATCAATGGCTACGGCCTGATCAGCGGCTTCG CCATCATCGGCGTGACCATGGCGGTGTCCTACTGGTTGTCGGCGCGCCTCACCCG CGGGCGCCTGCACGGCTCGGCGATCGCGATCTTCCTCGGCCTGGTGCTGTCCTAC GTCGGCGGCGTCGCCACCGGCGGGCAGAAAGGATTGGTGGACGTGCCCCTGATG TCCGGTATCGGCCTGCTGGGTGGCGCCATGCTGCGCGACTTCGCCATCGTCGCCA CGGCCTTCGGGGTGAATGTCGAGGAACTGCGTCGGGCAGGGGTTTCCGGGGTGG TGTCGCTGTTCCTCGGGGTCGGTTCGTCGTTCGTCGCCGGGGTCGCGGTGGCCAT GGCGTTCGGCTACACCGATGCGGTCAGCCTGACCACCATCGGCGCCGGCGCGGT CACCTACATCGTCGGACCGGTGACCGGCGCCGCCCTGGGGGCCAGCTCCGAGGT GATGGCGCTGAGCATCGCCGCCGGGCTGGTGAAGGCGATCCTGGTGATGGTCCT GACGCCTTTCGTCGCGCCCTACATCGGCCTGAACAATCCACGTACGGCGGTGATC TTCGGTGGTCTCATGGGCACCTCCAGTGGCGTCGCCGGCGGCCTCGCCGCCACCG ACCCGAAGCTGGTGCCCTACGGTTGCCTGACCGCGGCGTTCTATACCGCCATCGG TTGTCTGCTGGGGCCGTCGCTGCTGTATTTCAGCGTGCGCGGCCTGGTTGGGTGA acoX– SEQ ID NO: 14 ATGCCGCTACCCGCCCCGATTGTCGGGATCATTGCCAACCCTGCCTCTGGCCGCG ACCTGCGCCGCCTGACCGCCAATGCCGGGCTCTATTCGAGCACCGACAAGGCCTC GGCTATCCAGCGCCTGCTGGCCGCTTTCGGTGCCACCGGCATTGGCCAGGTGTTG TTGCCCAGCGACATGACCGGCATCGCTGCCGCCGTGCTCAAGGCCAGCCAGGGC CCGGGGGCCCGTGACCAGCACTGGCCGGCCATCGAAATCCTCGACCTGCCACTG AGCCAGACCGTGGCCGATACCCGCCTGGCTACCCGGCACATGGTCCAGCGCGGC GTGGCGATGATCGCCGTGCTGGGCGGTGATGGCACCCACAAGGCGGTCGCCGCC GAGGCCGGTGACGTGCCGCTGCTGACCCTGTCCACCGGCACCAACAACGCCTTCC CCGAACTTCGCGAAGCCACCAGCGCGGGCCTGGCCGGTGGCCTGCTTGCCAGTG Att Dkt: 114198-5960 GCCGGGTACCGGCCAGCATCGGCTTGCGCCGCAACAAACGCTTACTGGTGAACG TGCCTGAACAACAGCTGGCCGAGTGGGCCTTGGTCGAGGTGGCCGTGTCGCCGC AGCGCTTCATCGGTGCCCGTGCACTCAGCCGCAGCGAGGACCTCTGCGAAGTCTT CGCCTGCTTCGCCGAGCCCCACGCCATCGGCCTGTCGGCCCTGTGCGGTCTGTGG TGCCCGGTGTCACGGCAAGATCCGCACGGCGCCTGGGTACGCCTGGCCCCCGAC GCCGAACAGGCACTGTTGGCGCCGTTGGCGCCCGGGCTGCTGCAAGCCTGCGGC ATCACTGCCTGCGGTCCGCTGACACCCGGCGTCGCCCAGCGCCTGAGCCTGACCA GCGGCACCTTGGCGCTGGATGGCGAACGCGAAATCGAATTTGCCGAGCATGACA CGCCCACCATCACCCTCGACCACCAAGGCCCACTCAGCGTGGACGTCGAGGCCG TGCTGGCGCATGCCGCCCGCCACCACCTGCTGGCCGTCCCGCGCGGCCACCGGCT GCACCCTGCAAACCCGTGTTGA acoA– SEQ ID NO: 15 ATGTCCAATCAACTCAGTACTGAACAACTGCTGCATGCCTATGAAGTCATGCGGA CCATCCGCGCCTTCGAAGAACGCCTGCATGTGGAGTTCGCCACCGGCGAAATTCC CGGCTTCGTCCACCTGTATGCCGGCGAAGAAGCCTCTGCCGCCGGGGTCATGGCC CACCTGCGCGACAGCGACTGCATCGCCTCGACCCACCGCGGCCATGGCCACTGC ATCGCCAAAGGCGTGGACGTGTACGGCATGATGGCCGAGATCTACGGCAAGAAG ACCGGCGTGTGCGGTGGCAAAGGCGGCTCGATGCACATTGCCGACCTGGAAAAG GGCATGCTCGGTGCCAACGGTATCGTCGGCGCCGGCGCACCACTGGTGGCCGGG GCGGCGCTGGCGGCCAAGCTCAAGGGCCGCGATGACGTGTCGGTAGCCTTCTTC GGCGATGGGGCATCCAACGAGGGCGCCGTATTCGAGGCCATGAACATGGCGTCG ATCATGAACCTGCCCTGCATCTTCGTGGCCGAGAACAACGGCTACGCCGAGGCC ACTGCCTCCAACTGGTCCGTAGCCTGCGACCACATTGCCGACCGCGCTGCCGGCT TCGGCATGCCAGGCGTCACCATCGACGGCTTTGACTTCTTTGCCGTATACGAGGC CGCCGGCGCTGCCATCGAGCGCGCCCGCTCCGGGCAAGGGCCGTCGCTGATCGA GGTCAAGCTCAGCCGTTACTACGGCCATTTCGAAGGCGATGCGCAAACCTACCGC GCCCCTGATGAAGTGAAGAACCTGCGCGAATCCCGCGACTGCCTCATGCAATTTC GCAACAAGACCACCCGCGCCGGCCTGCTCACGGCCGAGCAACTGGACGCCATCG ACGCCCGCGTCGAGGACCTGATCGAAGACGCCGTTCGCCGCGCCAAATCCGATC CCAAGCCACAGCCCGCCGACCTGCTCACCGACGTCTACGTCGCCTACCCCTGA acoB– SEQ ID NO: 16 ATGGCAAGAAAGATCAGCTACCAGCAGGCCATCAACGAGGCCCTGGCCCAGGAA ATGCGCCGCGACAGCACAGTATTCATCATTGGCGAAGACGTCGCCGGCGGTGCC GGTGCGCCCGGCGAGGATGACGCCTGGGGTGGCGTGTTGGGCGTGACCAAAGGC CTGTATCACCAATTCCCCGGCCGCGTGCTTGACGCGCCGCTGTCGGAAATCGGCT ACGTCGGCGCGGCCGTCGGTGCTGCCACCCAGGGCCTGCGCCCGGTGTGCGAAC TGATGTTCGTCGACTTCGCCGGCTGCTGCCTGGACCAGATCCTCAACCAGGCCGC CAAGTTCCGCTACATGTTCGGCGGCAAGGCAGTGACCCCGCTGGTGATGCGCACC ATGTATGGCGCCGGCCTGCGCGCGGCTGCCCAGCACTCGCAGATGCTCACCTCGC TATGGACGCACATCCCCGGCCTGAAGGTGGTGTGCCCATCCTCGCCTTACGATGC CAAGGGCCTGTTGATCCAGGCGATCCGCGACAACGACCCGGTGATCTTCTGCGA Att Dkt: 114198-5960 ACACAAACTGCTGTACAGCATGCAGGGCGAGGTGCCGGAAGAGGTGTACACCGT ACCGTTCGGCGAAGCCAACTTCTTGCGCGATGGCGACGACGTGACCCTCGTTACC TACGGGCGCATGGTGCATGTGGCGCTGGAGGCCGCCAACAACCTGGCCCGCCAG GGCATCGACTGCGAAGTGCTGGACCTGCGCACTACCAGCCCGCTGGACGAGGAC AGCATTCTGGAAAGTGTGGAGAAGACCGGCCGCCTGGTGGTGATCGATGAAGCC AACCCGCGCTGCTCGATGGCCACCGACATCAGCGCGCTGGTGGCACAGAAGGCT TTTGGCGCACTCAAGGGCCCGATCGAAATGGTCACCGCGCCGCATACCCCGGTGC CGTTCTCCGACGCCCTGGAAGACCTGTACATCCCTGACGCGGCGAAGATCGAAG CCGCCGTGCGCAAGGTGATCGAAGCCGCAAGGAGTGCCGCATGA acoC– SEQ ID NO: 17 ATGAGCCAGATCCATACCCTGACTATGCCCAAGTGGGGCCTGTCGATGACCGAG GGCCGGGTGGACACCTGGCTCAAGCAGGAAGGTGACGAAATCAGCAAGGGCGA CGAGGTGCTGGACGTCGAGACCGACAAGATCAGCAGCAGCGTCGAAGCGCCTTT CAGCGGCGTGCTGCGCCGCCAGGTAGCCAAGCCGGATGAAACCCTGCCCGTCGG CGCACTGCTGGCGGTGGTAGTCGAAGGCGAAGCAGAGGAGGCCGAGATCGATGC CGTGGTACAGCGCTTCCAGGCCGAGTTCGTCGCCGAAGGTGGCGCCGACCAAAG CCAGGGGCCGGCCCCTCAGAAGGCGGAAGTAGGCGGGCGCCTGCTGCGCTGGTT CGAGCTGGGTGAAGGCGGTACACCGCTGGTGCTGGTGCACGGGTTTGGTGGTGA CCTCAACAACTGGCTGTTCAACCACCCGGCGCTGGCTGCCGAGCGGCGGGTCATC GCCCTCGACCTGCCGGGCCACGGCGAGTCGGCCAAGGCCCTGCAGCGGGGCGAC CTGGATGAGTTGAGCGAAACCGTGCTCGCCCTGCTCGACCACCTGGACATCGCCA AGGCCCACCTGGCCGGCCATTCCATGGGCGGAGCCGTCAGCCTCAACGTGGCGC GCCTGGCCCCGCAGCGGGTGGCCAGCCTGAGCCTGGTGGCCAGCGCCGGCCTGG GCGAAGCGATCAACGGGCAGTACTTGCAAGGCTTCGTCACTGCGGCCAACCGCA ATGCACTCAAACCGCAGATGGTGCAGTTGTTTGCCGACCCGGCGCTGGTCACCCG GCAAATGCTGGAGGACATGCTCAAGTTCAAGCGCCTGGAAGGTGTGGACCAGGC CCTGCAGCAACTGGCAGGGGCGCTCGCCGACGGCGACCGGCAGCGCCATGACTT GCGCGGGGTCCTTGGCAATCACCCGGCACTGGTGGTGTGGGGTGGCAAGGATGC GATCATCCCGGCCAGCCATGCCGAGGGCCTGGAAGCCGAAGTACAGGTGCTGCC AGAGGCCGGCCACATGGTGCAGATGGAGGCGGCCGAACAGGTCAACCAGCAACT GCTCGCATTCCTGCGCAAGCACTAA bdhA– SEQ ID NO: 18 ATGCGCGCGGCCGTCTGGCATGGCCGCCACGATATTCGTGTCGAACAGGTACCTT TGCCGGCCGACCCTGCGCCGGGCTGGGTGCAGATCAAGGTGGACTGGTGCGGCA TCTGCGGCTCCGACCTGCACGAATATGTTGCCGGCCCGGTGTTCATCCCGGTAGA GGCCCCGCACCCGCTGACCGGCATTCAGGGCCAGTGCATCCTCGGCCACGAATTC TGCGGCCACATCGCCAAGCTTGGCGAAGGCGTGGAAGGCTATGCCGTAGGCGAC CCGGTGGCGGCAGACGCGTGCCAGCATTGTGGTACCTGCTATTACTGCACCCATG GCCTGTACAACATCTGCGAACGCCTGGCGTTCACCGGCCTGATGAACAACGGTGC CTTCGCCGAGCTGGTCAACGTGCCCGCCAACCTGCTCTACCGGCTGCCGCAGGGC TTCCCTGCCGAAGCCGGGGCACTGATCGAGCCGCTGGCGGTGGGTATGCACGCG Att Dkt: 114198-5960 GTGAAAAAGGCCGGCAGCCTGCTTGGGCAAACCGTTGTAGTGGTTGGGGCCGGC ACCATCGGCCTGTGCACCATCATGTGCGCCAAGGCTGCAGGTGCGGCACAGGTC ATCGCCCTTGAGATGTCCTCTGCGCGCAAAGCCAAGGCCAAGGAAGCGGGCGCC AACGTGGTGCTGGACCCCAGCCAGTGCGATGCCCTGGCGGAAATCCGCGCACTG ACTGCTGGGCTGGGCGCCGATGTGAGTTTTGAGTGCATCGGCAACAAACATACG GCCAAGCTGGCCATCGACACCATCCGCAAAGCAGGCAAGTGCGTGCTGGTGGGT ATTTTCGAAGAGCCCAGCGAGTTCAACTTCTTCGAGCTGGTGTCCACCGAGAAGC AAGTGCTGGGGGCGTTGGCGTACAACGGCGAGTTTGCTGACGTGATTGCCTTCAT TGCTGATGGTCGGCTGGATATTCGCCCGCTGGTAACCGGCCGGATCGGATTGGAG CAGATTGTCGAGCTGGGCTTCGAGGAACTGGTGAACAACAAAGAGGAGAACGTG AAGATCATCGTTTCACCAGGTGTGCGCTAATAA ampC– SEQ ID NO: 19 ATGCGCGATACCAGATTCCCCTGCCTGTGCGGCATCGCCGCTTCCACACTGCTGT TCGCCACCACCCCGGCCATTGCCGGCGAGGCCCCGGCGGATCGCCTGAAGGCAC TGGTCGACGCCGCCGTACAACCGGTGATGAAGGCCAATGACATTCCGGGCCTGG CCGTAGCCATCAGCCTGAAAGGAGAACCGCATTACTTCAGCTATGGGCTGGCCTC GAAAGAGGACGGCCGCCGGGTGACGCCGGAGACCCTGTTCGAGATCGGCTCGGT GAGCAAGACCTTCACCGCCACCCTCGCCGGCTATGCCCTGACCCAGGACAAGAT GCGCCTCGACGACCGCGCCAGCCAGCACTGGCCGGCACTGCAGGGCAGCCGCTT CGACGGCATCAGCCTGCTCGACCTCGCGACCTATACCGCCGGCGGCTTGCCGCTG CAGTTCCCCGACTCGGTGCAGAAGGACCAGGCACAGATCCGCGACTACTACCGC CAGTGGCAGCCGACCTACGCGCCGGGCAGCCAGCGCCTCTATTCCAACCCGAGC ATCGGCCTGTTCGGCTATCTCGCCGCGCGCAGCCTGGGCCAGCCGTTCGAACGGC TCATGGAGCAGCAAGTGTTCCCGGCACTGGGCCTCGAACAGACCCACCTCGACG TGCCCGAGGCGGCGCTGGCGCAGTACGCCCAGGGCTATGGCAAGGACGACCGCC CGCTACGGGTCGGTCCCGGCCCGCTGGATGCCGAAGGCTACGGGGTGAAGACCA GCGCGGCCGACCTGCTGCGCTTCGTCGATGCCAACCTGCATCCGGAGCGCCTGGA CAGGCCCTGGGCGCAGGCGCTCGATGCCACCCATCGCGGTTACTACAAGGTCGG CGACATGACCCAGGGCCTGGGCTGGGAAGCCTACGACTGGCCGATCTCCCTGAA GCGCCTGCAGGCCGGCAACTCGACGCCGATGGCGCTGCAACCGCACAGGATCGC CAGGCTGCCCGCGCCACAGGCGCTGGAGGGCCAGCGCCTGCTGAACAAGACCGG TTCCACCAACGGCTTCGGCGCCTACGTGGCGTTCGTCCCGGGCCGCGACCTGGGC CTGGTGATCCTGGCCAACCGCAACTATCCCAATGCCGAGCGGGTGAAGATCGCCT ACGCCATCCTCAGCGGCCTGGAGCAGCAGGGCAAGGTGCCGCTGAAGCGCTGA PA4111– SEQ ID NO: 20 ATGAGCGCTTCCCCGCCCCTCCGCCAGCCCTGCCCGCCCGGCGCCTGCGTCTGCG AACGCGAGCGCCTGGAGGCGCCCGGCGCGGACCGTCGCATCCTCCTCCTGACCC GCCAGGAAGAGCAGCGCCTGGCCGCTCGCCTGGAAGCCCTGCGCAGCCTGGAAG ACCTGGAACACCTGCTGCGGCGCATGGAGGAACAACTGGGCATCCGCCTGCGGA TCGCCCCGGCCTTCGGCGAGGTGCGCAGCATGCGCGGCATCCGCATGCGCTTCGA GGAGCAGCCGGGGCTCTGCCGCAAGACCCGCCAGGCGATTCCCGCCGCCATCCG Att Dkt: 114198-5960 CCGCGGCCTGGAGAAGCGCCCGGAAGTCGCCTACGCCCTGCTCAACGCCCATGA CCTGCTGCGCGACGCCTGA PA4112– SEQ ID NO: 21 GTGAACGTACAACAGAGCAAGTTCAGGCTGAGCCGCTGGGCGGGCGCGACACTG GCGCTCGGACTGCTGCTGTCCGGGCTGGGCGCCTGGGGGCTGGCGCTGGTCAAC GAGCAACAGGCGCGGATGGCGCTGGAACGCGAGGCGGAACTGCTGGCCGAGGC GGTGACCCGCCGCGTGGAACTCTACCAGTATGGCTTGCGCGGGGTGCGCGGTGC CTTGCTGACGGCGGGCGAGGTGCACATCGACCGCGAACTGTTCCGCCGCTACAG CCTGACCCGCGACATCGACCGCGAGTTTCCCGGCGCCCGCGGCTTCGGCTTCATA CGCCGGGTGGCGGCCGCCGACGAGGCCGGGTTCCTGCGCCAGGCGCGGGCCGAC GGCCAGCCGGAGTTCCGCATCCAGCAACTGACTCCGCACGATGGCGAGCGCTAC GTCATCCAGTACATCGAGCCGGTGGCGCGCAACGGCCAGGCGTTGGGCCTGGAC ATCGCCTCGGAAGCCAATCGCCGGGAGGCCGCCCGCGCCGCCCTGGAGACCGGG CAGGTGCGCCTGACGGGGCCGATCACCCTGGTCCAGGCCAGCGGCTTGCGCCAG CAATCGTTCCTCATCCTGATGCCGATCTACCGCAGCGGCATCACTCCACCGCCGG GGCCCCAGCGCGAACTCGAAGGCTTCGGCTGGAGCTATGCGCCGCTGCTGACCG GCGAGGTGCTGGCCGGGTTGCCCATCGACAACGCGGCGATCCACCTCGAACTCA GCGACGTCACCGGCGACGGCGCTGCCGTGCCCTTCTTCGTCAACGGCGCCGCAGC GCCGGCTCAACGGCTGTTCGGCCACATGCTGCGGCGGGAGATCTACGGTCGGCA CTGGCAGATGGCGTTCAGCGCACTGCCGCTGTTCGTCCAGCGCCTGCACCAGCCT TCGCCGCGCATCCTGTTCCTGGCCGGCAGCCTGGTCAGCCTGTTGCTGGCGGCGC TGGTCAACGCCAGCGCCCTCGGCCGCCTGCGCCGGCGTCGGGAGGCGGCGACCC AGGCGCGGCTGGCGGCGATCGTCGGCAATTCGGCGGACGGCATCATCGGCGTCG GCCTCGACGGCGTCATCAGCGACTGGAATCGCGGGGCCGAGGCGTTGTTCGGCT ACCGCGAGGAGCAGGCGGTCGGGCGGCGGGTGGTGGACCTGCTGGTGCCGCCAT CCAAGGAGAACGAGGAACTCGACATCCTCGCCCGCATCGCCCGTAAAGAACAGG TGGTCAGCTTCGATACCGTGCGCCGGCACCAGGATGGCCACCTGCTCGACGTGGC GGTGACCGTCTCGCCGATCCTCGGTCCGGACGGCGGCGTGGTCGGCGCCTCGAA GACCGTACGCGACATCTCCGCGAAGAAGGCCGCCGAGGCCCGCATCCGCGAACT CAATACCGGCCTGGAGCACCAGATCGCCGAGCGTACCGCCGAACTGCGTCGGTT GAACGTACTGCTCGGCAGCGTGCTGCAGGCGGCGTCCGAGGTGTCGATCATCAC CCTCGACCGCGATGGCATCATCAGCGGCTTCAATCTCGGTGCCGAGCGGATGCTC GGCTACCGCGCCGACGAGGTGGTCGGCAAGGCCTCGCCGACCCTGCTGCACAGC GAGCGGGAGTTGCTCGCCCGCGGCCAGGAGCTGGGCGAGGAGGGTTTCCGCGTG CTGGTGGCGCGGGCCGAACAGGAGGGCGCCGAGACCCGCGAATGGACCTATCTG CGCAAGGACGGCTCGCCGCTCAGCGTGACGCTTTCGGTTACCGTCCTGCGCGACG AGGCTGGGGCGATCAACGGCTACCTGGGCATCGCCGTCGATGTCACCGAATGGC GCAACGCCGAGCGCGAGATGGCCGCGGCGCGCGACCAGTTGCAAATGGCCGCCG ACGTGGCGCGCCTGGGCATCTGGCGCTGGAACCTGGCCGACGACAGCCTGCAGT GGAACGAGCGCATGTGCGAGATGTATGGCCAGCCGCTGGCGTTGCGCGACGGCG GCCTGGTCTACGAGCATTGGCGCAGCCGCCTGCATCCGGAAGACCTCGAGCGCA CCGAGGCCAGCCTGCGCGCGGCGGTGGAGGGGCGCGGCAACTACGATGTGATCT Att Dkt: 114198-5960 TCCGCGTGGTGCTGCCCGACGGCGGCATCCGCTTCATCCAGGCCGGCGCGCAGGT CGAGCGCGACGCCGACGGCAACCCGCTGCAGGTCACCGGGATCAATATCGATAT CACCAGCGACCAGCAACTGCAGGCGCGCCTGCGCGAGGCCAAGGAGCAGGCGG ACGCCGCCAGCGCGGCGAAGTCGTCGTTCCTGGCGAACATGAGCCACGAGATCC GCACGCCGATGAATGCCGTGCTGGGCATGCTGCAACTGGCCCGGCAGACCGATC TCAACGAACGCCAGCGCGACTACCTGGACAAGGCCAGCTCGGCGGCAACCTCGT TGCTCGGCCTGCTCAACGACATCCTCGACTACTCGAAGATCGAGGCCGGCAAACT GGTCCTGGAAATGCTCCCCTTCGAGCTGGAGCCGCTGATGCAGGACCTGGCCGTG GTGCTCTCCGGCAACCAGGGCGACAAGGACGTCGAGGTGATCTTCGACATCGAT CCCGAGCTGCCCAGCGCGGTGGTCGGCGACCGCCTGCGCCTGCAGCAGATCCTC ATCAACCTGGCCGGCAACGCACTGAAGTTCACCGCCCGCGGGCACGTCCTGGTG AGCTTGCGGCGTCTGGCGCACGACGCCCACCTGGTGCGCCTGCGGGTGCTGGTGG CGGACACCGGCATCGGCATCAGCGCGGAACAGCAGCAGCGCATCTTCGAAGGTT TCACCCAGGCCGAGGCGTCCACCTCGCGACGCTTCGGCGGCACCGGCCTCGGCCT GTTCATATGCAAACGCCTGGTCGACCTGATGGGCGGTGAATTGCGGGTCGAGAG CGCGCCGGGCAGCGGTAGCCGGTTCTGGTTCGACCTGGACCTCGACGTCGCCCAC GACCAACCGCTGCGTGCGGCCTGCCCCGGGGCCGGCGAACCGCTGCGGCTGCTG GTGGCGGACGACAACCTGGTGGCCGGCGAGCTGCTCGAACGCACGGTCGGCGCC CTCGGCTGGCGCGCCGACTGCGTCGGCAGCGGCAGCGAGGCGGTGGCGCGGGTC CAGGCGGCGATGGCCGAGGGCCGTCGCTACGACGTGGTGCTGATGGATTGGCGC ATGCCCGATCTCGACGGCCTCAGCGCGGCGCAACTGATCCGCCAGTTGCAGGGC GACCTGCCGCCGCCCATGGTGATCATGATCACCGCCTATGGCCGCGAGGTGCTGG CCGATGCGCGCGACCACAGCGCGCCGCCCTTCGTCGATTTTCTCACCAAGCCGGT CACGCCCAAGCAACTGGCCGACTCGGTGCTGCATGCGTTGCACGGCGAACAGGG CGCCCCGGCCAACCCGCCGCCGCGCCCGGTGGAGCGCACCCAGCGGTTGCGGGG CGTGCGTCTGCTGGTAGTGGAGGACAACGCGCTGAACCGCCAGGTCGCCGCCGA GCTGTTGAGCAGCGAGGGCGCCCGGGTGGCCCTCGCCGACGGCGGGCTGGCGGG CGTTCAGCAGGTGCTGGAGGCGAGCGTGCCGTTCGACGCGGTGCTGATGGACAT GCAGATGCCGGATATCGATGGTCTCGAGGCGACCCGGAGAATCCGCGCCGACGG GCGCTTCGCCGGGCTGCCGATCCTGGCCATGACCGCCAACGCCTCGCTGGCCGAC CGCGAAGCCTGCCTGGCGGCGGGCATGAACGACCATGTGGCGAAGCCGATCGAC AAGGAGCGCCTGGTGCTTTGCCTGCTCGGTCATCTCGGCAGGAGCGGCGCCAGG GGCGCGCCGGCAACGGCAGCGGACGCGGGCGAACTGGTCGAGGCGCGGGGCGA TATCGTCGGGCGTTTCGGCGGCAGCCTGGAGCTGATCGTCCAGGTGTTGCGGCGC TTCGTTCCGGACATGCAGGACCTCTTCGCCCAGCTCGAACGCCAGCTCGGCGAGG GCGATGTCCAGGGCAGCGCCGCGACCCTGCATACCATCAAGGGCAGCGCTTCCA CCGTTGGCGCCAGCGCCCTGGCCGGCCGTGCCAGCGAGCTGGAGCAGGCCCTGC GCCGCGCCGATCCGCTGCGCGGCATGGAAATCCTCGCCGGCATCCGCCTCGACG AACTGCGCACCCTGTGCGATGCCTGCCTGCTCCGCCTGCAGGCGATGTTCGGCGA CGTGCAGGCGGAGTCGGCGAGCTGA creD– SEQ ID NO: 22 Att Dkt: 114198-5960 ATGAACCGCACGCTGGCCTACAAGCTCGGCGCCATCGCCCTACTCATCCTGCTAT TGCTGATTCCCCTGTTGATGATCGACGGCCTGATCCGCGATCGCCAGGAGGTTCG TGACGGCGTGCTGCAGGATATCGCCCGCAGCTCCAGCTACAGCCAGCGCCTCACC GGCCCGCTGCTGGTGGTGCCGTACCGCAGGACGGTGCGCGAGTGGGTCACCGAG GAGAAGACCGAGAAGCGCGTGTTGCAGGAGCGCGAGCAGCGCGGCGAGCTGTA CTTCCTGCCCGACCGCTTCGCCCTCGACGGGCAGATGCGCACCGAGCTGCGCTAC CGTGGCATCTACCAGGCGCGCCTGTACCACGTCGACAACAAGGTTGCCGGCTACT TCCAGGTGCCTGCGCACTATGGCATCGAAGAAGACCTGGAGGACTACCAGTTCG AAACGCCGTTCCTGGCGATGGGCATCAGCGACATCCGCGGCATCGAGAATGCCC TGCGCCTGTCCCTCAACGGTGCCAGCGTCGACTTCGAGCCGGGTAGCCGGGTCGG CCTGCTCGGCAGCGGTGTGCACGCACCGCTGCAAGGCGTCGACGGGCGCCAGGC GCAGCGCCTGGAGTACGCCTTCGACCTGTCGTTGCTCGGCAGCGAGCGGCTGGAC ATCGTGCCGGTGGGCCGCGACAGCCAGGTCATCCTGAAGGCCGACTGGCCGCAC CCGAGCTTCGGCGGCGAGTTCCTGCCCAGCGAGCGGGAGATCACCGCCCAGGGC TTCACCGCGCGCTGGCAGACCAGCTTCTTCGCCACCAACATGGAAGAGGCGCTGC GCAGTTGCGTGGAGGAGCAGCGTTGCGACGGCTTCCAGGCGCGCGCCTTCGGCG TCGGCCTGGTCGATCCGGTGGACCAGTACCTGAAGGCCGACCGCGCGATCAAGT ACGCGCTGCTGTTCATCACCCTGACCTTCGCCGGCTTCTTCCTCTTCGAGGTGCTC AAGCGCCTGGCGGTACACCCGGTCCAGTACGCGCTGGTCGGCCTGGCGCTGGCG TTCTTCTACCTGCTGCTGCTGTCGCTGGCCGAGCACGTCGGCTTCGAGCTGGCCTA CCTGGTCTCCGCCGGAGCCTGTGTCGGCCTGATCGGCTTCTACCTGTGCTTCGTCC TGCGCAGCGTGGCCCGAGGGCTGGGCTTCAGCGTCGGCCTGGCAGGGCTGTACG GCCTGCTCTACGGCTTGCTGAGCGCGGAGGACTACGCCTTGCTGATGGGCTCGCT GCTGCTGTTCGCGGTGCTTGCCGCGGTGATGGTGCTGACCCGCCGCCTGGACTGG TACGGGGTCGGCCGCAAGGACTTCCCGGCGGTACCCGCCAGGGCCTGA PA0466– SEQ ID NO: 23 ATGCGCCTGTTCTTCGAGTTCACCCTCTGTGCCGCCCTCGGCCTGCTGGTCAGCGG CGTGATCCTGATTTTTCTGCTGTTCTTCGGCAAGTTCGGCCTGATCGAGTACTGGA TCGCCAGCGGCAAGCCCCTGGCCCACGCCATGCTGGTGCTGATGCCGGACGGATT CTGGGAAGGCCTCAGCGGCATCGAGGGCGCCGCCGGCAATCCGCACGTGCGCTC CTTCCTGGAGCTTTGCATGGGGCTCGGACAGAGCGCGCTGTTGCTGGGCCTGCTG CTGTTCCGCCTGGTGTGCTGGAAATGA PA0467– SEQ ID NO: 24 ATGAGCGCTACCCATACGCTGTTCTACGCCGCCGCCTCGCCCTTCGTGCGCAAGG TCCTGGTCCTGCTCCACGAAACCGGCCAACGCGAGCGCGTGGCGCTCGAGGAAG TCACGCCGACGCCGGTCGCGCCCATCCGGCAGCTCAACGCCAGCAACCCGGCCG GCAAGATTCCCGCCCTGCGCCTGCCGGACGGCCAGGTGCTGCACGACAGCCGGG TGATCTGCGACTACTTCGACCAGCAGCACGTCGGCGAACCGCTGATCCCCCGCGA AGGCAGCGCCCGCTGGCGACGCCTGACCATCGCCTCGCTGGCGGACGCCGTCCTC GACGCCGCGGTGCTGTCCCGCTACGAAACCTTCGTCCGCCCCGAGGAGAAGCGC TGGGACACCTGGCTCGAAGCACAGCGGGAAAAGATCGGGCGGTCCCTGGCGTGG Att Dkt: 114198-5960 CTGGAGGGCGACTGCATCGCCGAACTGCAGGCGCGCTTCGACATCGCCGCCATC GGCGTCGCCTGCGCCCTGGGCTATCTCGACCTGCGCCAGCCGGAGTGGGACTGGC GCGGCCGCTACCCGCGACTCGCCGCCTGGTTCGCCGAAGTCAGCCAGCGCCCGTC GATGCAGGCCACCCGCGCCTGA carO– SEQ ID NO: 25 ATGAAACTTCGTCACCTGCCCCTCATCGCTGCCATCGGCCTGTTCTCCACCGTCAC CCTGGCCGCCGGCTATACCGGCCCGGGCGCTACCCCGACCACCACCACGGTGAA GGCCGCGCTGGAAGCCGCCGACGACACCCCGGTGGTCCTCCAGGGCACCATCGT CAAGCGCATCAAGGGCGACATCTACGAGTTCCGCGATGCCACCGGCAGCATGAA GGTGGAGATCGACGACGAAGACTTCCCGCCGATGGAAATCAACGACAAGACCCG GGTCAAGCTGACCGGCGAAGTCGACCGCGACCTGGTCGGCCGCGAGATCGACGT CGAGTTCGTCGAAGTGATCAAGTAA pcaF-I-SEQ ID NO: 26 ATGCACGACGTATTCATCTGTGACGCCATCCGTACCCCGATCGGCCGCTTCGGCG GCGCCCTGGCCAGCGTGCGGGCCGACGACCTGGCCGCCGTGCCGCTGAAGGCGC TGATCGAGCGCAACCCTGGCGTGCAGTGGGACCAGGTAGACGAAGTGTTCTTCG GCTGCGCCAACCAGGCCGGTGAAGACAACCGCAACGTGGCCCGCATGGCACTGC TGCTGGCCGGCCTGCCGGAAAGCATCCCGGGCGTCACCCTGAACCGTCTGTGCGC GTCGGGCATGGATGCCGTCGGCACCGCGTTCCGCGCCATCGCCAGCGGCGAGAT GGAGCTGGTGATTGCCGGTGGCGTCGAGTCGATGTCGCGCGCCCCGTTCGTCATG GGCAAGGCTGAAAGCGCCTATTCGCGCAACATGAAGCTGGAAGACACCACCATT GGCTGGCGTTTCATCAACCCGCTGATGAAGAGCCAGTACGGTGTGGATTCCATGC CGGAAACCGCCGACAACGTGGCCGACGACTATCAGGTTTCGCGTGCTGATCAGG ACGCTTTCGCCCTGCGCAGCCAGCAGAAGGCTGCCGCTGCGCAGGCTGCCGGCTT CTTTGCCGAAGAAATCGTGCCGGTGCGTATCGCTCACAAGAAGGGCGAAATCAT CGTCGAACGTGACGAACACCTGCGCCCGGAAACCACGCTGGAGGCGCTGACCAA GCTCAAACCGGTCAACGGCCCGGACAAGACGGTCACCGCCGGCAACGCCTCGGG CGTGAACGACGGTGCTGCGGCGATGATCCTGGCCTCGGCCGCAGCGGTGAAGAA ACACGGCCTGACTCCGCGTGCCCGCGTTCTGGGCATGGCCAGCGGCGGCGTTGCG CCACGTGTCATGGGCATTGGCCCGGTGCCGGCGGTGCGCAAACTGACCGAGCGT CTGGGGATAGCGGTAAGTGATTTCGACGTGATCGAGCTTAACGAAGCGTTTGCCA GCCAAGGCCTGGCGGTGCTGCGTGAGCTGGGTGTGGCTGACGATGCGCCCCAGG TAAACCCTAATGGCGGTGCCATTGCCCTGGGCCACCCCCTGGGCATGAGCGGTGC ACGCCTGGTACTGACTGCGTTGCACCAGCTGGAGAAGAGTGGCGGTCGCAAGGG CCTGGCGACCATGTGTGTGGGTGTCGGCCAAGGTCTGGCGTTGGCCATCGAGCGG GTTTGA pcaB-SEQ ID NO: 27 ATGAGCAACCAACTGTTCGACGCCTATTTCACCGCGCCGGCCATGCGCGAGATTT TCTCCGACCGAGGCCGCCTGCAGGGCATGCTGGATTTCGAAGCCGCGCTTGCCCG Att Dkt: 114198-5960 AGCCGAAGCCTCTGCCGGTTTGGTCCCGCACAGCGCGGTAGCGGCCATCGAGGC GGCATGCCAGGCCGAGCGCTATGACGTTGGCGCGCTGGCCAATGCCATCGCCAC CGCGGGCAACTCGGCCATTCCGCTGGTGAAAGCGTTGGGCAAGGTGATCGCCAC CGGCGTGCCAGAGGCTGAGCGCTATGTGCACCTTGGGGCCACCAGCCAGGATGC GATGGATACCGGTCTGGTTCTGCAGCTGCGCGATGCCCTCGATTTGATCGAGGCC GACCTCGGCAAGCTGGCCGATACCCTGTCGCAGCAGGCCTTGAAGCACGCCGAT ACGCCCTTGGTGGGTCGTACCTGGTTGCAACACGCCACCCCGGTGACCCTGGGCA TGAAACTGGCCGGTGTACTGGGTGCTTTGACCCGCCACCGTCAGCGCCTGCAGGA ACTGCGCCCGCGCCTTCTGGTCCTGCAGTTCGGCGGTGCCTCGGGCAGCCTGGCG GCGCTGGGCAGCAAGGCGATGCCGGTGGCCGAAGCGCTGGCCGAACAGCTCAAG CTGACCCTGCCCGAGCAGCCCTGGCACACCCAGCGCGACCGCCTGGTGGAGTTTG CCTCGGTATTGGGCCTGGTTGCCGGCAGCCTGGGCAAGTTCGGCCGTGATATCAG CTTGCTGATGCAAACCGAGGCGGGGGAGGTGTTTGAGCCTTCTGCGCCGGGCAA GGGTGGTTCTTCGACCATGCCACACAAGCGCAACCCGGTGGGTGCCGCCGTGTTG ATCGGTGCCGCGACCCGCGTGCCGGGCCTGCTGTCGACGCTGTTCGCAGCCATGC CTCAGGAGCACGAACGCAGCCTGGGCCTATGGCATGCCGAGTGGGAAACCCTGC CGGATATCTGCTGCCTGGTCTCTGGCGCCCTGCGCCAGGCTCAAGTGATTGCCGA GGGCATGGAGGTGGATGCCGCGCGCATGCGCCGTAACCTCGACCTGACCCAAGG CCTGGTGCTGGCCGAAGCGGTGAGCATCGTCCTCGCCCAGCGTCTGGGTCGCGAC CGTGCCCACCACCTGCTGGAACAATGCTGCCAACGCGCGGTGGCCGAACAGCGG CACCTGCGTGCCGTGCTGGGTGACGAGCCGCAGGTCAGCGCCGAGCTGTCTGGC GAAGAACTCGATCGCCTGCTCGACCCTGCCCATTACCTGGGCCAGGCCCGCGTCT GGGTGGCGCGCGCCGTGTCCGAACATCAACGTTTCACTGCCTGA pcaD-SEQ ID NO: 28 GTGGCGCACTTGCAACTGGCCGATGGCGTTTTGAATTACCAGATCGATGGCCCGG ATGACGCCCCGGTGCTGGTCCTGTCCAACTCGCTGGGTACCGACCTGGGCATGTG GGACACCCAGATTCCGCTCTGGAGTCAGCACTTCCGGGTGCTGCGCTATGACACC CGTGGTCACGGCGCATCGCTGGTCACTGAAGGCCCTTACAGCATCGAACAGCTG GGCCGCGACGTGCTGGCCCTGCTCGATGGCCTGGACATTCAAAAGGCTCACTTCG TCGGCCTGTCGATGGGCGGCCTGATCGGCCAGTGGCTGGGTATCCATGCAGGTGA GCGCCTGCACAGCCTGACCCTGTGCAACACGGCCGCCAAGATCGCCAATGACGA GGTGTGGAACACCCGTATCGACACGGTACTCAAAGGCGGCCAGCAGGCCATGGT CGACCTGCGCGATGCCTCCATCGCCCGCTGGTTCACCCCGGGCTTTGCCCAGGCG CAGGCGGAGCAGGCCCAGCGTATCTGCCAGATGCTGGCGCAAACCAGCCCGCAA GGCTACGCAGGCAACTGTGCAGCGGTACGTGACGCTGATTATCGTGAGCAACTG GGCCGCATCCAGGTGCCTGCGCTGATCGTTGCCGGTACCCAAGACGTGGTTACCA CCCCTGAGCATGGCCGCTTCATGCAGGCCGGTATCCAAGGTGCCGAGTACGTCGA CTTCCCGGCGGCGCACCTGTCCAATGTCGAGATTGGCGAGGCCTTCAGCCGCCGC GTGCTCGATTTCCTGCTGGCTCACTGA pcaC-SEQ ID NO: 29 Att Dkt: 114198-5960 ATGGACGAGAAACAACGTTACGACGCTGGCATGCAAGTGCGCCGCGCAGTGCTG GGTGATGCCCACGTGGACCGCAGCCTGGAGAAGCTCAACGACTTCAATGGCGAG TTCCAGGAAATGATCACCCGCCACGCCTGGGGTGACATCTGGACCCGCCCGGGG CTGCCGCGCCATACCCGCAGCCTGATCACCATCGCCATGCTGATTGGCATGAACC GCAACGACGAGCTGAAGCTGCACCTGCGTGCGGCGGCCAACAATGGCGTGACCC GCGACGAGATCAAGGAAGTGCTGATGCAGAGCGCGATCTACTGCGGCATTCCGG CGGCCAATGCCACGTTCCACCTGGCTGAGTCGGTGTGGGATGAACTTGGCGTAGA GTCTCGCCAGTAA pcaI-SEQ ID NO: 30 TTGATCAATAAAACGTACGAGTCCATCGCCAGCGCGGTGGAAGGGATTACCGAC GGTTCGACCATCATGGTCGGTGGCTTCGGCACGGCTGGCATGCCGTCCGAGCTGA TCGATGGCCTCATTGCCACCGGTGCCCGCGACCTGACCATCATCAGCAACAACGC CGGCAACGGCGAGATCGGCCTGGCCGCCCTGCTCATGGCAGGCAGCGTGCGCAA GGTGGTCTGCTCGTTCCCGCGCCAGTCCGACTCCTACGTGTTCGACGAACTGTAC CGCGCCGGCAAGATCGAGCTGGAAGTGGTCCCGCAGGGCAACCTGGCCGAGCGT ATCCGCGCCGCAGGCTCCGGCATTGGTGCGTTCTTCTCGCCAACCGGCTACGGCA CCCTGCTGGCCGAGGGCAAGGAAACCCGTGAGATCGATGGCCGCATGTACGTGC TGGAAATGCCGCTGCACGCCGACTTCGCACTGATCAAGGCGCACAAGGGTGACC GTTGGGGCAACCTGACCTACCGCAAGGCCGCCCGCAACTTCGGCCCGATCATGG CCATGGCTGCCAAGACCGCCATCGCCCAGGTCGACCAGGTCGTCGAACTCGGTG AACTGGACCCGGAACACATCATCACCCCGGGTATCTTCGTCCAGCGCGTGGTCGC CGTCACCGGTGCTGCCGCTTCTTCGATTGCCAAAGCTGTCTGA pcaJ-SEQ ID NO: 31 ATGACCATCACCAAAAAGCTCTCCCGCACCGAGATGGCCCAACGCGTGGCCGCA GACATCCAGGAAGGCGCGTACGTAAACCTGGGCATCGGCGCACCGACCCTGGTG GCCAACTACCTGGGCGACAAGGAAGTGTTCCTGCACAGCGAGAACGGCCTGCTG GGCATGGGCCCAAGCCCTGCGCCGGGCGAGGAAGACGATGACCTGATCAACGCC GGCAAGCAGCACGTCACCCTGCTGACCGGTGGTGCCTTCTTCCACCATGCCGATT CGTTCTCGATGATGCGTGGCGGCCACCTGGACATCGCTGTACTGGGCGCCTTCCA GGTGTCGGTCAAGGGCGACCTGGCCAACTGGCACACGGGTGCCGAAGGCTCGAT CCCGGCCGTAGGCGGTGCAATGGACCTGGCCACCGGCGCCCGCCAGGTGTTCGT GATGATGGACCACCTGACCAAGACCGGCGAAAGCAAGCTGGTGCCCGAGTGCAC CTACCCGCTGACCGGTATCGCTTGCGTCAGCCGCATCTACACCGACCTGGCCGTA CTGGAAGTGACACCTGAAGGGCTGAAAGTGGTCGAAATCTGCGCGGACATCGAC TTTGACGAGCTGCAGAAACTCAGTGGCGTGCCGCTGATCAAGTGA pcaH-SEQ ID NO: 32 ATGCCCGCCCAGGACAACAGCCGCTTCGTGATCCGTGATCGCAACTGGCACCCTA AAGCCCTTACGCCTGACTACAAGACCTCCGTTGCCCGCTCGCCGCGCCAGGCACT GGTCAGCATTCCGCAGTCGATCAGCGAAACCACTGGTCCGGACTTTTCCCATCTG Att Dkt: 114198-5960 GGCTTCGGCGCCCACGACCATGACCTGCTGCTGAACTTCAATAACGGTGGCCTGC CCATTGGCGAGCGCATCATCGTCGCCGGCCGTGTCGTCGACCAGTACGGCAAGCC TGTGCCGAACACTTTGGTGGAGATGTGGCAAGCCAACGCCGGCGGCCGCTATCG CCACAAGAACGATCGCTACCTGGCGCCCCTGGACCCGAACTTCGGTGGTGTTGGG CGGTGTCTGACCGACCGTGACGGCTATTACAGCTTCCGCACCATCAAGCCGGGCC CGTACCCATGGCGCAACGGCCCGAACGACTGGCGCCCGGCGCATATCCACTTCG CCATCAGCGGCCCATCGATCGCCACCAAGCTGATCACCCAGTTGTACTTCGAAGG TGACCCGCTGATCCCGATGTGCCCGATCGTCAAGTCGATCGCCAACCCGCAAGCC GTGCAGCAGTTGATCGCCAAGCTCGACATGAGCAACGCCAACCCGATGGACTGC CTGGCCTACCGCTTTGACATCGTGCTGCGCGGCCAGCGCAAGACCCACTTCGAAA ACTGCTGA pcaG-SEQ ID NO: 33 ATGCCAATCGAACTGCTGCCGGAAACCCCTTCGCAGACTGCCGGCCCCTACGTGC ACATCGGCCTGGCCCTGGAAGCCGCCGGCAACCCGACCCGCGACCAGGAAATCT GGAACTGCCTGGCCAAGCCAGACGCCCCGGGCGAGCACATTCTGCTGATCGGCC ACGTATATGACGGAAACGGCCACCTGGTGCGCGACTCGTTCCTGGAAGTGTGGC AGGCCGACGCCAACGGTGAGTACCAGGATGCCTACAACCTGGAAAACGCCTTCA ACAGCTTTGGCCGCACGGCTACCACCTTCGATGCCGGTGAGTGGACGCTGCAAAC GGTCAAGCCGGGTGTGGTGAACAACGCTGCTGGCGTGCCGATGGCGCCGCACAT CAACATCAGCCTGTTTGCCCGTGGCATCAACATCCACCTGCACACGCGCCTGTAT TTCGATGATGAGGCCCAGGCCAATGCCAAGTGCCCGGTGCTCAACCTGATCGAG CAGCCGCAGCGGCGTGAAACCTTGATTGCCAAGCGTTGCGAAGTGGATGGGAAG ACGGCGTACCGCTTTGATATCCGCATTCAGGGGGAAGGGGAGACCGTCTTCTTCG ACTTCTGA pcaK-SEQ ID NO: 34 ATGAACCAAGCGCAAACCAACGTCGGCAAAAGCCTCGACGTTCAGTCATTCATC AATCAGCAGCCGCTGTCCCGCTATCAGTGGCGGGTCGTGCTGCTGTGCTTCCTGA TCGTCTTCCTCGACGGCCTGGACACCGCTGCCATGGGGTTCATCGCCCCTGCACT GTCGCAAGAGTGGGGGATCGACCGCGCCAGCCTTGGGCCGGTCATGAGTGCCGC GTTGATCGGCATGGTGTTCGGTGCCTTGGGCTCCGGGCCGCTGGCTGACCGTTTC GGCCGCAAAGGCGTGCTGGTGGGGGCGGTGCTGGTGTTCGGTGGCTTCAGCCTG GCCTCGGCCTATGCCACCAACGTCGACCAGTTGCTGGTACTGCGCTTCCTGACTG GCCTGGGCCTGGGTGCAGGCATGCCCAACGCTACCACGCTGTTGTCCGAATACAC CCCGGAGCGCCTCAAGTCGCTGCTGGTGACCAGCATGTTCTGTGGCTTCAACCTG GGCATGGCGGGTGGTGGCTTCATTTCCGCCAAGATGATCCCGGCCTACGGCTGGC ACAGCCTGCTGGTCATCGGTGGCGTGCTGCCGTTGCTGCTGGCGCTGGTGCTGAT GATCTGGCTGCCGGAGTCGGCGCGGTTCCTGGTGGTGCGCAACCGCGGCACCGA CAAGGTGCGCAAGACCTTGTCGCCAATCGCGCCGCAAGTGGTGGCCGAGGCAGG CAGTTTCAGCGTGCCTGAGCAAAAGGCGGTGGCTGCGCGCAACGTGTTCGCGGT GATCTTCTCTGGCACCTATGGCTTGGGCACCGTGTTGCTGTGGCTGACCTATTTCA TGGGCCTGGTGATCGTCTACCTCCTGACCAGCTGGCTGCCAACCCTGATGCGTGA Att Dkt: 114198-5960 CAGCGGCGCGAGCATGGAGCAGGCGGCGTTCATCGGCGCGCTGTTCCAGTTTGG CGGCGTGCTAAGTGCAGTCGGTGTGGGTTGGGCCATGGACCGTTTCAATCCGCAC AAGGTAATTGGCATTTTCTATCTGCTGGCCGGGGTGTTCGCCTACGCAGTGGGGC AGAGCCTGGGCAACATCACTCTGCTAGCCACGCTGGTGCTCGTCGCTGGTATGTG CGTGAACGGTGCGCAGTCGGCAATGCCATCGCTGGCGGCGCGCTTCTATCCCACC CAGGGCCGGGCCACAGGCGTATCGTGGATGCTGGGTATTGGCCGCTTTGGTGCG ATTCTCGGCGCCTGGAGTGGTGCAACGCTGCTGGGCCTGGGCTGGAGCTTCGAGC AGGTGCTGACGGCGTTGCTGGTACCGGCGGCGCTGGCGACCGTGGGAGTGGTCG TCAAGGGGCTGGTCAGCCACGCCGACGCAACCTGA azoR2-SEQ ID NO: 35 ATGTCCCGCGTACTGATCATCGAAAGCAGCGCCCGCCAGCAGGATTCCGTTTCCC GTCAGCTGACCAAGGACTTCATCCAGCAATGGCAGGCCGCCCACCCGGCCGATC AGATCACTGTGCGCGACCTGGCAGTAAGCCCGGTTCCCCACCTGGATGCGAACCT GCTGGGCGGCTGGATGAAACCCGAAGAACAGCGCAGTGCCGCCGAACTGGAGGC CCTGGCCCGTTCCAACGAACTGACGGATGAGTTGTTGGCTGCCGACGTACTGGTG ATGGCTGCGCCCATGTACAACTTCACCATCCCCAGCACCCTGAAAGCCTGGCTGG ACCATGTGCTGCGTGCCGGCATCACCTTCAAATACACCCCGACCGGCCCGCAAGG CCTGCTGACTGGCAAGCGCGCCATCGTCCTGACTGCCCGCGGCGGCATCCATGCG GGTGCCAGCAGTGACCATCAGGAACCGTACCTGCGCCAGGTCATGGCTTTCATCG GCATCCACGATGTCGACTTCATCCATGCCGAAGGCCTGAACATGAGCGGCGAGTT CCACGAGAAGGGCGTCAACCAGGCCAAGGCCAAGCTGGCAGCGGTGGCCTGA PP_3671-SEQ ID NO: 36 ATGCAAAGGCTCAGCACCCGCACCGGCCTGGACCTGCCCGCCATCGGCCTGGGC ACCTGGCCAATGACCGGCAGTGAATGCACCCAGGCCGTGCGCCAGGCACTGGAC GTGGGTTACCGGCACATCGACACCGCCACCGCCTACGAAAACGAAGCAGCCGTT GGCCAGGCGCTGCGCGACAGCGACGTGCCACGCGAGCAGATCCACCTGACCACC AAGGTCTGGTGGGACCGCCTGGAGCCTAAAGCCATGCGTCAGTCGCTGGAAGAC AGCCTGCGCGCGCTGGGTACCGAACAGGTCGACCTGTTCCACATCCACTGGCCAG GCACCGACTGGGACCTGGCCCGCAGCATCGACACCCTGGTAGCCCTGCGTGACG AAGGCAAGGCGCGCCACATTGGCGTGGCCAACTTCCCATTGGGGTTACTGCGCC AGGTGGTGGAAACCCTGGGCGCGCCGTTGTCGGCAATTCAGGTGGAATACCATG TTCTGCTCAGCCAGCAACCGCTACTCGACTTCGCCCGCGGCCATGACCTGTTGCT CACCGCCTACACCCCGCTGGCACGCGGCCAGGCAGCGGCGCAGCCGGTGATCCA GGCGATCGCCCGCAAGCATGGCGTGCTGCCCAGCCAGGTGGCGCTGAAATGGCT GCTGGACCAGGACGGCGTGGCGGCGATCCCCAAAGCCAGCAGCCGCGAAAACCA GCTGGCCAACCTGGCCGCGTTGACCGTGCCGCTGGATGACGAAGACCGCACAGC CATTGCCGGGCTGCCGAAAGACCAGCGCGTGGTCAGCCCGCCGTTTGCCCCGGA CTGGAACAGCTGA PP_1399-SEQ ID NO: 37 Att Dkt: 114198-5960 ATGAAAAACAAACTGATCCTTACCCTGGCCCTCTCGGTTTTCGCAGCTGGCGCCT TCGCCGAAGATGGCTTCGACCGGACCGACGCTCACACCTTTGCCGCAGTTCAAAC CCAGTCGCCCCGCTACGCTGAAGACGGCTTCGACCGCACCGGTGGCCAACGCTTC GCCGAAGACGGCTTCGACCGCACCGGCGGCCAGCGCTTCGCCGAAGACGGCTTC GACCGCACCGGTGGCCAACGCTTCGCTGAAGACGGCTTCGACCGCACCGGTGGC CAACGCTTCGCTGAAGACGGCTTCGACCGCACCGGTGGCCAACGCTTCGCTGAA GACGGCTTCGACCGCACCGGTGGCCAACGCTTCGCTGAAGACGGCTTCGACCGC ACCAACGCGCACCGCATCAGCTGA ttgR-SEQ ID NO: 38 ATGGTCCGTCGAACCAAAGAAGAAGCCCAGGAAACCCGTGCCCAGATCATCGAG GCGGCGGAAAGGGCCTTCTACAAGCGCGGGGTTGCGCGAACCACCCTGGCCGAC ATCGCCGAACTGGCGGGTGTGACGCGGGGGGCGATCTACTGGCACTTCAACAAC AAGGCCGAGCTGGTGCAGGCGTTGCTCGACAGCCTGCACGAGACCCATGACCAT TTGGCGCGCGCGAGCGAAAGCGAGGATGAACTCGACCCGCTTGGCTGCATGCGC AAGCTGCTGTTGCAAGTGTTTAACGAGCTGGTGCTCGATGCCCGAACCCGTCGTA TCAATGAAATCCTGCATCACAAGTGCGAGTTCACCGATGACATGTGTGAAATTCG CCAGCAGCGCCAGAGTGCGGTGCTGGATTGCCACAAGGGCATCACCCTGGCGCT GGCCAATGCCGTACGCCGGGGCCAGTTGCCTGGCGAGCTGGATGTCGAGCGCGC GGCGGTGGCCATGTTTGCCTATGTCGATGGCCTGATCGGGCGCTGGCTATTGCTG CCTGACAGTGTCGATCTGTTGGGTGATGTCGAAAAATGGGTCGACACCGGGCTG GATATGCTGCGCTTGAGCCCGGCTCTGCGCAAATGA galP-II-SEQ ID NO: 39 ATGAGCATTTCGCTACCCGCACGTCAGCTGCTGCCTGGCCTGTTGGCCATGTCCT GCGCCCTCCCCGCATTCGCCGCCAACGAAGGTGGTTTCATCGAGGACGCCAAGG CCACCCTCAACCTGCGCAACTTCTACATCAACCGCAACTTCGTCGACCCGGCCCA CCCACAGGGCAAGGCCGAGGAATGGACACAAAGCTTCATCCTCGATGCCCGTTC CGGTTTCACCCAGGGCACGGTCGGTTTCGGCGTGGACGTACTGGGCCTGTACTCG GTCAAGCTCGACGGCGGCAAAGGCACCACCAACACTCACCTGCTGCCCGTGCAT GGCGACGGCCGCCCGGCCGATGACTTTGGCCGGCTGGGTGTGGCGCTGAAGGCC AAGCTGTCCGAAACCGAGCTGAAAGTCGGTGAATGGATGCCGGTACTGCCAATC CTGCGTTCGGACGATGGCCGCTCACTGCCGCAAACATTCCGCGGTGGCCAGCTGA CCTCCAGGGAAATCGCCGGCCTGACCCTGTACGCCGGCCAGTTCCGCGGCAACA GCCCGCGCAACGATGCCAGCATGGAAGACATGTCGCTCAACGGTAAAGCCGCCT TCACCTCCGACCGTTTCAACTTCGGCGGCGGTGAGTACACCTTCAACGACAAGCG CACCATGATCGGCCTGTGGGATGCCCAGCTCAAGGACATCTACCGCCAGCAATA CCTCAACCTGACGCACAGCCAGCCGCTGGGCGACTGGACGCTGGGCGCCAACCT GGGCTACTTCATCGGCAAGGAAGATGGCGCCGAACGGGCCGGCAAGCTGGACAA CCGCACCGCCTCCGCCTTGCTTTCGGCGCGCTACCAAGGCCACACCTTTTACGTT GGCCTGCAGAAGGTCAGCGGCGACGATGCCTGGATGCGGGTCAATGGCACCAGT GGCGGTACCCTGGCCAACGACAGCTACAATTCAAGCTTCGACAATGCCAAGGAG CGTTCCTGGCAGGTTCGCCATGACTTCAACTTCGCCACCGTCGGCGTACCCGGGC Att Dkt: 114198-5960 TGACCCTGATGAACCGCTATATCAAGGGTGACAACGTCACCGTGGGCAGCGTGG ACGATGGCAAGGAATGGGCGCGGGAAACAGAGCTGGCCTATGTGGTGCAGTCGG GCAGCTTCAAGGACCTGTCGCTGAAATGGCGCAACTCGACCATGCGTCGCGACTT CAGCACCAACGCGTTCGATGAGAACCGCTTGATTGTGAGTTATCCGCTGAATCTT CTGTGA PP_1864-SEQ ID NO: 40 TTGCAGCGTCCGGCGCGCTGGCTGGAGTTGTACCGTCAGCGTCAGGAACTGGCCA GTCTCAGTGATGCAACTTTGCGCGACTTCGGGTTGAGTCGGGCGGATATCCAGCA AGAGGCCGAGCGGCATTTCTGGGATGATCCACTGAGGAAGTGA PP_2222-SEQ ID NO: 41 ATGATGAATAGCCGAGTCGATGTGGCTGTGATGATAGGGTCGGGTGTGCCGGCC ACCTTGCAGGCGCTGGGCAAGCGTATTTGCTGGGTAGTGCTGGTCAATGGCGAGC GTCGCGGCACCGCCTTCGCCACGCGCGAAGAAGCGCTTGAGTGCCAGGCTGCCT GGCAGCAGCAGCTGAAGAAAGGGCTGGCTGCCTGA PP_3494-SEQ ID NO: 42 ATGAAACGCTCGATCGCCGTCCTTGCCATTACCGCTGCCTTCGCCTCGTTCGGTGC TTTTGCCGATGCAGGCAATAACCAGCCAGTCAGCAGCCAATATGAATACGGCAT GCCGCTGGACGTGGCCAAGGTCATCTCCATCACGCCAGCCAGCAACGCCGCTGA CTGCCAAGTGGGCACCGCGCACATGGTGTACGTCGACCATCAGGGCCAGAAGCG CGAAATCGACTACCGGGAAATGGGCAACTGCTCGCAGCAGTGA PP_5458-SEQ ID NO: 43 ATGTCCAACAGTATGGGTATTGCCAGCGTCTTCGTCCTGTCTTCCCTGTTGTTGTC GCCCCTGGCCATGGCCGAAGAATCCCCGGCATTCGTTGCCCGAAATGCAGAGAT CGCGGCTGTCCATCAAGAAGCCCGTGACGCTGTAGCCGCAAACAAGGCCGATAC CGTAAAAGATGCGAAACAGACAGCTTCCAAAGTCAGCGTCGATACCAGTGCCGA CAGCTGA PP_0175-SEQ ID NO: 44 ATGTCCCTTGATTTCCTGCGCCTGCAAGTCAGCAGTGGCCTGGTAGTCGCTGCCC GCCACTGGCGGCGCATCTGTCATACCGCCTTGACCGGCTACGGCATTTCCGAGGC CTGCGCTGCGCCTTTGCTGATGATCGTGCGCCTGGGCGACGGCGTGCACCAGGTG GCTGTGGCCCAGGCCGCCGGCCTGGAAAGCCCGTCGCTGGTACGCCTGCTTGACC AACTGTGCAAAGCTGGCCTGGTGTGCCGCAGCGAAGACCCACTGGATCGCCGCG CCAAGGCACTGCGCCTGACTGCCGAAGGCCGTGCCCTGGCCGAAGCCATTGAAG CCGAACTGGTGCGCTTGCGCCGCGATGTGCTCGGCGGTATCGATCAGGCCGACCT GGAAGCTGCCTTGCGTGTGCTTCGTGCCTTCGAGGCTGCCGGCCTGGGCAGCGCA GGCGGGGTGGCATGA Att Dkt: 114198-5960 PP_0176-SEQ ID NO: 45 ATGAACGGTTTCTTCAGCTCTGTCCCTCCGGCCCGTGACTGGTTCTACGGCGTGC GCACCTTCGGCGCGTCCATGATTGCCCTGTACATTGCCCTGCTCATGCAACTGCC GCGTCCCTACTGGGCCATGGCGACGGTGTACATCGTCTCCAGCCCCTTCGTCGGC CCGACCAGCTCCAAGGCGTTGTACCGTGCCGTCGGCACGTTGCTCGGCGCTGGCG GGGCAATCTTCCTGGTGCCGCCGCTGGTGCAGTCACCGTTGCTGCTGAGCATCGC CGTTGCCCTGTGGACCGGCACCTTGCTGTTCCTTTCGCTGAACCTGCGCACGGCC AACAACTACGTGCTGATGCTGGCCGGTTATACCCTGCCGATGATTGCCTTGGCCG TGGTGGATAATCCGCTGGCAGTGTTCGATGTGGCTTCTTCTCGTGCGCAGGAAAT CTGTTTGGGAATTGTCTGCGCGGCAGTCGTCGGGGCTATCTTCTGGCCACGCCGG CTGGCGCCGGTGGTAGTCGGTGCCACCGGAAACTGGTTCAACGAGGCGATTCGC TACAGCGACACCTACCTGGCCCGCGATGCCAGCGCTGACAAGGTCGGCGGTATG CGCGGGGCGATGGTCACTACCTTCAACTCGCTGGAGCTGATGATCGGTCAGCTCG CCCACGAAGGTGCCGGCCCGCATACCCTGAAGAACGCCCGCGAATTGCGTGGGC GGATGATCCATCTGCTGCCGGTGATCGACGCCCTCGACGACGCACTGGTAGCCCT GGAAGGCCGCGCCCCGGCCCAGTTCGCCCAGCTTCAACCCGTGCTGGAAGCCGC TCGTGAATGGCTCAAAGGCACCGCCGACAGCGCTTCGGTAGCGCACTGGGCCAC GTTGCACGAGCAGATCGGCCGCCTGCAGCCCGCTACAGCCGCCCTCGACCAGCG CGCCGAGCTGCTGTTGTCCAACGCCCTGTACCGCCTGACCGAATGGGCCGACCTG TGGCAGGACTGCTGCACCCTGCAGCACGCCCTGCGTACCGATGATGCCAAGCCTT GGCGTGCGGTGTATCGCCACTGGCGGCTGGGCCGCCTGACAGCGTTTTTCGATCG TGGCCTGATGCTCTACTCGGTAGCCTCCACGGTTCTGGCGATCGTGGTCGCCTGC GGCCTGTGGATCGGCCTGGGCTGGAACGACGGCGCCAGTGCGGTAATCCTCGCG GCAGTTGCGTGCAGCTTCTTCGCCGCCATGGACGACCCGGCACCGCAGATCTACC GGTTCTTTTTCTGGACGCTGATGTCGGTCATCTTCTCCAGTCTGTACCTGTTTTTG GTACTGCCCAACCTGCACGACTTCCCGATGCTGGTGCTGGCCTTTGCCGTGCCGT TCATCTGCATAGGCACCCTGACCGTACAGCCGCGCTTCTACCTGGGCACCTTGCT GACCATCGTCAACACGTCGACCTTCATCAGCATCCAGGGTGCCTACGACGCCGAC TTCTTTACCTTCCTCAACTCCAACCTTGCCGGCCCCGCGGGCCTTCTGTTCGCCTT TATCTGGACCCTGGTGATGCGCCCGTTCGGCGTGGAACTGGCGGCCAAGCGCATG ACCCGCTTCGCCTGGCGCGACATCGTCGAAATGACCGAGCCTGCGACATTGGCCG AACACCGCCAGGTTGGCGTACAGATGCTCGACCGCCTGATGCAGCACCTGCCAC GTCTGTCGCAAACCGGCCAAGACAGTGGTGTGGCACTGCGTGACCTGCGCGTAG GGTTGAACCTGCTCGACCTGCTGGCGTACATGCCGCGCGCCGGCCAGCAGGCTCG CGAGCGCCTGAACACTGTGGTCGAGGAAGTCGGTGCGCACTACGCCGCCTGCCT GCGTGCTGGCGAACGCCTGCATGCCCCGGCCGCACTGCTGCGCAACATGGAACG CGCTCGCCTGGCGCTGAACCTGGATGAGCTGTACGAGCGCGGCGATGCCCGCAC CCACCTGCTGCATGCCCTGAGTGGCCTGCGCCTGGCGCTACTGCCGGGGGTGGAG GTGATGCTCGAACCCGCCGAACAACCGCAACTGCCCCCCGGGCTCGACGGAGCG CCCTTGTGA PP_0178-SEQ ID NO: 46 Att Dkt: 114198-5960 ATGAAAAAACCTTTGCTGACCTTGGGCCGTGTGGTCCTGACCTTGCTGGTAGTGA CCTTCGCCGCCGTGCTCGTGTGGCAGATGGTGGTCTACTACATGTTTGCCCCCTG GACCCGAGACGGCCATATCCGCGCTGACGTGATCCAGATCGCCCCCGATGTGTCG GGATTGATCCAGAAGGTCGAGGTGCGCGACAACCAGACCGTCAAACGTGGCGAC GTGCTGTTCACCATCGACCAGGACCGCTTCACCCTGGCCCTGCGTCAAGCCAAGG CAACCCTTGGCGAGCGCCAGGAGACCCTGGCCCAGGCTTCCCGCGAAGCGCAGC GCAACCGCAAGCTGGGCAACCTGGTGGCCGCCGAGCAACTGGAAGAAAGCCAGT CCCGCGAGGCCCGTGCCCGTTCGGCTGTCAGCGAGGCGCAAGTGGCTGTCGATA CTGCCCAGCTCAACCTTGACCGCTCGGTGGTGCGCAGCCCGGTAGACGGCTACCT CAACGACCGCGCGCCGCGTAACCACGAGTTCGTCACTGCTGGCCGCCCGGTGCTT TCGGTGGTCGACAGTGCCTCCTACCACGTCGACGGCTACTTCGAGGAAACCAAGC TGGGGGGTATCCATATTGGTGACGCCGTGGATATCCGCGTGATGGGCGACAACA CCCGCCTGCGTGGCCATGTGCAGAGCTTCGCCGCTGGCATCGAAGACCGCGACC GCAGCAGTGGCGCCAACCTGCTGCCCAACGTCAACCCGGCGTTCAGCTGGGTAC GCCTGGCCCAGCGCATTCCGGTGCGCATCGCCTTTGACGAAGTGCCGCAAGACTT CCGCATGATCGCCGGGCGTACTGCGACGGTGTCGATCATCGAGGGCCAGCGCCC ATGA PP_0487-SEQ ID NO: 47 ATGGACATCGCTATGGAAATGTTGGGTTCTTACTGGGCGGAGCACTGGTTGCTGG CAGTCCTCATACTTGGCGCGATCGCCTTCGTTGCGGGCTTCGTCGACAGCATTGC CGGTGGTGGAGGTTTGTTTCTGGTGCCGGGTTTTCTGCTGGTGGGCATGCCGCCC CAGGTGGCACTGGGCCAGGAAAAACTGGTCAGTACCCTCGGCACGCTGGCGGCT ATCCGCAACTTCCTGGCCAACAGCAAGATGGTCTGGCAGGTAGCGTTGGTCGGG GTGCCGTTTTCGCTGCTAGGCGCCTACCTGGGGGCTCACCTGATCGTGTCGATCT CGCAGGAAACTGTGGGCAAGATCATCCTGGCACTGATCCCGCTCGGTATTCTCAT TTTCCTTACCCCCAAGGACCGCCCGGTCGAGGAGCGCGAACTGTCATCGCGCATG CTGTTCACCGTGGTCCCGCTGACCTGCCTGGCCATTGGCTTCTATGACGGCTTCTT CGGCCCCGGTACCGGCAGCATGTTCATCATCGCTTTCCATTACCTGCTGCGCATG GACCTGGTCTCCAGTTCGGCCAACTCCAAAACCTTCAACTTCGCTTCCAACATCG GTGCTTTGGTGGCGTTCGTCAGCGCAGGCAAGGTCGTCTACCTGCTGGCATTGCC ACTTGTGGCCTGCAACATCCTCGGTAACCACCTGGGCAGTTCCCTGGCCTTGCGT AAAGGCAACGACGTGGTGCGCAAGGCGTTGGTGTTCTCGATGCTCTGCCTGTTCA CCAGCCTGGGCGTGAAGTACCTGACCTGA pcaT-SEQ ID NO: 48 ATGACCTCAACCTATTACACCGGCGAAGAACGCAGCAAGCGCATCTTCGCTATTG TCGGTGCCTCATCAGGCAACCTGGTGGAATGGTTCGACTTCTACGTCTACGCCTT CTGCGCGATCTACTTCGCCCCGGCCTTCTTCCCCTCCGACGACCCCACCGTGCAG CTGTTGAACACGGCGGGTGTATTCGCCGCGGGCTTTTTGATGCGCCCCATTGGTG GCTGGATCTTCGGCCGCCTCGCCGACCGCCATGGTCGCAAGAATTCCCTGATGAT CTCGGTGCTGATGATGTGCTTCGGCTCCCTGATGATCGCCTGCCTGCCCACCTATG CCAGCATCGGCACCTGGGCCCCGGCGCTGCTATTGCTGGCCCGCCTGATCCAGGG Att Dkt: 114198-5960 GCTGTCGGTAGGGGGCGAATACGGCACCACGGCCACCTACATGAGTGAAGTAGC GCTGCGTGGGCAGCGCGGCTTCTTTGCTTCGTTCCAGTACGTCACCCTGATCGGC GGCCAGTTGCTGGCAGTCCTGGTCGTGGTGATCCTGCAACAGCTGCTGACTGAAG ATGAACTGCGGGCCTGGGGCTGGCGCATTCCCTTCGTGGTGGGCGCCATCGCCGC CCTGATCTCGCTGATGCTGCGCCGTTCGCTGCACGAGACCAGTAGCGCCGAAGCC CGCAATGACAAGGATGCAGGTTCGATCGGAGGTTTGTTCCGCAATCACGCAGCG GCGTTCATCACGGTACTGGGCTACACCGCGGGTGGCTCGCTGATTTTCTATACTTT CACCACCTATATGCAGAAGTACCTGGTCAACACCGCGGGCATGACCGCCAAGAA CGCCAGCTACGTGATGACCGGGGCGTTGTTCCTGTTCATGGTGGTGCAGCCGTTC TTCGGCATGCTGTCCGACCGTATCGGCCGGCGCAATTCGATGCTGCTGTTCGGCG GCCTCGGTACCCTGTGCACCGTGCCGCTGCTGATGGCGCTGAAAACCGTGACCAG CCCGATCATGGCCTTCGTGCTGATCAGCCTGGCCCTGTGTATCGTGAGTTTCTACA CCTCGATCAGCGGTCTGGTGAAGGCCGAGATGTTCCCGCCGCAGGTGCGTGCACT GGGTGTTGGCCTGGCCTACGCGGTGGCCAACGCAGCATTCGGCGGTTCGGCCGA GTATGTGGCCCTGGGCCTGAAAACCCTGGGGATGGAAAACACTTTCTATTGGTAC GTGACGGCGATGATGGCGATTGCCTTCCTGTTCAGCCTGCGCCTGCCGAAGCAGG CGGCGTACCTGCACCATGATGATTAA PP_2008-SEQ ID NO: 49 GTGAGGTGCCCGCATTCCGTGATCACGCCATCAGGGGACCTTTCCATGGCCGCCG ACCGCTACCCGCACCTGCTTGCTCCGCTGGACCTGGGCTTTACCACCTTGCGCAA CCGCACCCTGATGGGCTCGATGCACACCGGCCTCGAAGAGCGCCCCGGCGGCTT CGAGCGCATGGCAGCTTACTTTGCCGAGCGCGCCCGGGGCGGCGTTGGCCTGAT GGTCACGGGCGGCATTGCGCCCAATGATGAAGGCGGGGTGTATTCCGGTGCGGC AAAGCTCAGCACCGAGGAAGAGGCCGACAAGCACCGCATCGTCACCGAGGCGGT GCACGCTGCCGGTGGCAAGATCTGCCTGCAGATACTGCATGCCGGGCGCTACGC CTACAGCCCACGGCAGGTGGCACCTAGCGCGATCCAGGCGCCGATCAACCCGTT CAAGCCCAAAGAGCTGGATGAGGCGGGCATCGAGAAGCAGATCGCCGACTTCGT CAATTGTGCCGTGCTGGCTCAGCGTGCCGGTTACGACGGCGTCGAAATCATGGGT TCGGAAGGCTACTTCATCAACCAGTTCCTGGCCGCCCACACCAACCACCGCACCG ACCGCTGGGGCGGCAGTTATGAAAACCGCATGCGCCTGGCAGTGGAAATCGTCA GCCGGGTGCGTGGCGCGGTAGGGCCGAACTTCATCATCATCTTCCGCCTGTCGAT GCTCGACCTGGTCGAGGGTGGCAGCACCTGGGACGAGATCGAGCTGCTGGCCAA GGCCATCGAGCAGGCCGGCGCGACCTTGATCAACACCGGAATCGGTTGGCACGA GGCGCGTATTCCGACCATCGCCACCAAAGTGCCGCGTGCGGCCTTCAGCAAAGTC ACCGCCAAGTTGCGCGGCGTGGTGAGCATTCCGCTGATCACCACCAACCGCATCA ACACCCCGGAAGTGGCCGAGGCAGTGCTGGCCGAGGGCGATGCGGACATGGTCT CGATGGCGCGACCGTTTCTCGCCGACCCGGACTTCGTCAACAAGGCCGCTGCCGG TCGTGCGGATGAAATCAACACCTGCATCGGCTGCAACCAGGCCTGCCTGGACCAT ACCTTCGGCGGCAAGCTGACCAGTTGCCTGGTCAACCCGCGGGCCTGCCACGAG ACCGAACTCAACTACTTGCCTGTACGTACGGTGAAACGCATTGCCGTGGTCGGCG CCGGCCCGGCTGGCCTGGCGGCGGCCACCGTGGCGGCCGAGCGGGGCCACGCGG TGACCCTGTTTGACGCCGCCAGCGAAATCGGTGGCCAGTTCAACGTGGCCAAGC Att Dkt: 114198-5960 GGGTGCCGGGCAAGGAAGAATTCTTCGAAACGCTGCGTTACTTCCGCAACAAGG TCAAAAGCACGGGCGTCGACCTGCGCCTGAATACCCGCGTGGATGTGCAGGCAC TGGTGGGCGGCGGCTTTGATGAAGTCATCCTGGCTACCGGCATCGCCCCGCGTAC CCCGGACATCGCGGGCGTGGAGCATGCCAAGGTGCTCAGCTACCTGGACGTGCT GCTCGAGCGCAAGCCGGTGGGCAAGTCGGTGGCCGTGATTGGCGCGGGAGGTAT CGGCTTCGATGTGTCCGAGTACCTGGTGCATCAGGGCGTGGCCACCAGTCAGGAC CGAGCGGCATTCTGGAAAGAGTGGGGCATCGATACCCATCTTCAGGCGCGAGGT GGTGTGGCCGGGATCAAGGCCGAGCCGCATGCTCCGGCGCGGCAGGTGTACCTG TTGCAGCGCAAGAAATCCAAGGTGGGCGACGGGCTCGGCAAGACGACCGGCTGG ATTCACCGCACCGGGTTGAAGAACAAGGGGGTGCAGATGCTCAACAGTGTCGAG TATCTGGGTATCGACGATGCCGGCCTGCACATTCGTGTGGACGGCGGCGAGCCCC AGGTGCTGGCGGTGGATAACGTGGTGATCTGTGCCGGGCAAGATCCGCTGCGCG AGCTGCAGGAAGGGCTGGTGGCGGCGGGGCAGTCGGTGCACCTGATCGGCGGCG CGGATGTGGCGGCCGAGCTGGATGCCAAGCGGGCGATCAACCAAGGCTCGCGGT TGGCGGCTGAGCTCTGA PP_2406-SEQ ID NO: 50 ATGAGCCAGCAAGCCATCCTCGCCGGCCTGATCGGCCGCGGCATCCAACTGTCAC GTACCCCGGCCCTGCACGAACACGAAGGCGACGCCCAGGCCCTGCGCTACCTGT ACCGGCTGATCGACGCCGACCAGTTGCAACTGGACGACAGCGCCCTGCCCGGCC TGCTCGAAGCCGCGCAACATACCGGTTTCACCGGGCTGAACATTACCTACCCGTT CAAGCAGGCGATCCTGCCGTTGCTCGACGAGCTGTCCGACGAGGCCCGTGGTATC GGCGCGGTAAACACCGTGGTGCTCAAGGACGGCAAGCGCGTCGGCCACAACACC GATTGCCTGGGCTTTGCCGAAGGCTTGCGCCGTGGCCTGCCCGACGTTGCACGGC GCCAGGTGGTGCAGATGGGTGCCGGTGGCGCTGGCTCGGCCGTGGCCCATGCCC TGCTGGGTGAAGGGGTAGAACGGCTGGTGCTGTTCGAAGTGGATGCGACCCGTG CGCAGGCGTTGGTGGACAACCTGAACACCCATTTCGGCGCGGAGCGCGCCGTGC TTGGCACCGACCTGGCCACAGCCTTGGCCGAAGCGGACGGGCTGGTCAACACCA CGCCGGTGGGCATGGCCAAGCTACCGGGCACGCCACTGCCGGTGGAGTTGTTGC ATCCCCGCCTGTGGGTGGCAGAGATCATCTACTTCCCGCTGGAGACCGAACTGCT GCGCGCGGCCCGGGCACTGGGCTGCCGCACGCTGGATGGCAGCAACATGGCGGT GTTCCAGGCGGTGAAGGCATTCGAGCTGTTCAGCGGGCGGCAGGCGGATGCGGC GCGGATGCAGGCGCACTTTGCCAGTTTCACCTGA PP_2604-SEQ ID NO: 51 ATGAACCACTCCAAGATCCCCCCCGCCGCGCCCCGGATGTCCCAGGGCTCGATCG GCGACAAGCTCCGCGGCGCCTTCGGCGTCGGCAAGACCCGTTGGGGCATGCTCG CCCTGGTGTTCTTCGCCACCACCCTGAACTACATCGACCGCGCCGCCCTTGGCAT CATGCAGCCCGTGCTGGCCAAGGAAATGAGCTGGACGGCGATGGACTACGCCAA CATCAATTTCTGGTTCCAGGTCGGCTATGCCATCGGCTTTTTGCTGCAAGGCCGG CTGATCGACAAGGTCGGAGTCAAACGGGCCTTCTTCTTCGCTGTACTGCTGTGGA GCCTGGCCACGGGTGCCCACGGCCTGGCCACCTCGGCTGCAGGTTTCATGGTCTG CCGCTTCATCCTCGGCCTGACCGAGGCCGCCAACTACCCGGCCTGCGTCAAGACC Att Dkt: 114198-5960 GTGCGCCTGTGGTTCCCGGCGGGTGAGCGGGCGATTGCAACGGGGTTGTTCAAC GCCGGTACCAATGTTGGCGCCATGGTCACCCCCGCGCTGTTGCCGCTGATCCTGG CTGTATGGGGCTGGCAGGCTGCGTTCATCGCCATGGGCAGCCTGGGCCTGGTGTG GGTCGTGTTTTGGGTGCTGAAATACTACAACCCGGAAGACCACCCCCGCGTGCGT CAGAGCGAGGTGCAGTACATCCAGCAGGATGACGAGCCCGAGCCGACCCGCGTG CCGTTCAGCCGCATCCTGCGCCTGCGTGGCACCTGGGCTTTTGCAGTGGCCTATG CGATTACCGCGCCGGTGTTCTGGTTCTACCTGTACTGGCTGCCGCCGTTCCTGAAC CAGCAATACAGCCTGGGCATCAACGTGACCCAGATGGGCATCCCGTTGATCATC ATCTGGCTGACGGCAGACTTCGGCAGTATCGGGGGAGGCATTCTGTCCTCGTGGC TGATCGGCCGTGGCGTGCGTGCCACAACAGCGCGCTTGGTGTCCATGCTGATCTT TGCCTGCACCATGCTCAGCGTCATCTTTGCTGCCAACGCCAGTGGACTGTGGGTC GCGGTGGCGGCCATCTCGATCGCGGTAGGCGCGCACCAGGCGTGGACGGCGAAT ATCTGGAGCCTGGTGATGGACTACACACCCAAGCATCTGGTGAGCACCGTATTCG GCTTCGGCAGCATGTGTGCGGCAATTGGCGGGATGTTCATGACGCAGATTGTCGG CAGCGTGCTGACCGCTACCAACAACAACTATGCCGTGCTGTTTACCATGATCCCG GCCATGTACTTCCTGGCGTTGGTGTGGATGTATTTCATGGCACCGCGCAAGGTCG AGAGCGCCTGA PP_2643-SEQ ID NO: 52 GTGGTGCCAACAAGGAGTACTGCCCGCATGCTTGCAAACCTGAAAATTCGCACC GGAATGTTCTGGGTGTTGTCGCTGTTCAGCCTGACCCTGCTGTTTTCGACCGCCAG TGCCTGGTGGGCTGCCTTGGGTAGCGACCAGCAGATCACTGAACTGGACCAGAC TGCGCACCAGTCGGACCGTTTGAACAATGCGTTGCTGATGGCGATTCGTTCCAGC GCCAACGTGTCTTCGGGTTTTATCGAGCAGTTGGGTGGCCATGACGAGAGTGCAG GCAAACGCATGGCGCTGTCGGTGGAGCTGAACAACAAGAGCCAGGCGCTGGTCG ACGAATTTGTCGAAAACGCCCGCGAACCCGCCCTTCGCGGGTTGGCAACCGAGC TGCAGGCTACCTTTGCCGAATACGCCAAGGCGGTGGCTGGGCAGCGTGAAGCGA CCCGCCAGCGTTCGCTGGAGCAGTATTTCAAGGTCAACAGCGATGCCGGTAATGC CATGGGGCGGTTGCAAACCCTGCGCCAGCAACTGGTGACCACACTGAGCGAACG TGGCCAGCAAATCATGCTCGAGTCCGACCGGCGTCTGGCCCGGGCACAATTGCTG AGCCTGTGCCTGCTGGGCGTGACCGTGGTGCTGGCGGTGCTGTGCTGGGCCTTCA TTGCCCAGCGTGTGTTGCACCCGCTGCGTGAAGCCGGTGGGCATTTCCGGCGCAT TGCCAGTGGCGACCTGAGTGTGCCGGTGCAAGGGCAGGGTAACAACGAAATCGG CCAACTGTTCCATGAGCTCCAGCGCATGCAGCAGAGCCAGCGCGACACTCTGGG GCAGATCAACAACTGTGCCCGGCAGCTGGACGCCGCCGCCACGGCGTTGAACGC CGTCACCGAGGAAAGCGCCAACAACCTGCGTCAGCAGGGGCAGGAGCTGGAGC AGGCCGCCACTGCCGTGACCGAAATGACCACGGCAGTGGAGGAAGTTGCGCGCA ACGCGATCACCACCTCGCAAACCACCAGTGAATCCAACCAGCTGGCCGCGCAAA GCCGCCGGCAAGTCAGCGAAAACATCGACGGCACCGAAGCCATGACCCGCGAAA TCCAGACCAGCAGTGCGCACCTGCAGCAGTTGGTGGGGCAGGTGCGGGACATCG GCAAGGTGCTGGAAGTGATCCGCTCGGTGTCCGAGCAAACCAACCTGCTGGCGC TCAATGCTGCCATCGAGGCGGCCCGTGCCGGTGAGGCCGGGCGCGGCTTTGCGG TGGTGGCGGATGAGGTGCGCACACTGGCCTACCGCACGCAACAGTCGACCCAGG Att Dkt: 114198-5960 AAATCGAGCAGATGATCGGCAGCGTGCAGGCGGGGACCGAAGCTGCGGTGGCTT CCATGCAGGCCAGCACCAACCGCGCCCAGTCGACACTGGACGTGACCCTGGCCT CGGGGCAAGTGCTGGAGGGCATCTACAGTGCGATCGGCGAAATCAACGAGCGCA ATCTGGTCATTGCCAGTGCGGCCGAAGAGCAGGCTCAGGTGGCGCGGGAAGTAG ACCGCAACTTGCTGAACATCCGCGAACTGTCCAACCATTCCGCTGCGGGTGCGCA GCAGACCAGCGAGGCGAGCAAGGCGTTGTCGGGGCTGGTGGGGGAGATGACGG CGTTGGTGGGGCGGTTCAAGGTTTGA mhpT-SEQ ID NO: 53 ATGCATCATTCATGCGCGTCATCGGGCAAGGCCGTGCTCACCATCGGCTTGTGTT TTCTCGTTGCCTTGCTGGAAGGGCTGGACCTGCAGGCCACCGGCATCGCCGCGCC GCACATGGCCAAGGCATTCAACCTAACCCCTGCCATGCTCGGTTGGGTATTCAGT GCCGGGTTGCTGGGGTTGCTGCCCGGCGCACTGATCGGCGGTTGGCTGGCAGACC GCTTTGGCCGCAAAGCGATCCTGATCGTGGCGGTGCTGCTGTTTGGCGGCTTCTC GCTAGGGACCGCCCATGCCCAGACCTACGATTCGCTGTTGATCGCGCGGCTGATG ACAGGGCTCGGGCTGGGGGCGGCGTTGCCGATCCTGATCGCCCTGGCCAGCGAG GCGGCGCCTGAGCGTCTGCGGAGTACGGCGGTCAGCCTGACTTACTGTGGTGTGC CGCTGGGCGGGGCGGTTGCGTCGCTGATCGGCATGGCCGGAGTGGGCGACGGTT GGCGTACGGTGTTCTACGTCGGCGGCATCGCACCGATCGTGATCGCCTTCGTGCT GATGATCTGGCTGAAAGAGTCCCAGGCATTCCGTGCCCAAGGCGTTGCCAAGGC TGGCAGCGAAGGTGTGCTTGCCCAACTGTTCGGGCCGCAGCAGGCAAGCCGCAC GCTGTTGCTGTGGGTGGCATGTTTCTTCACCCTGACAGTGTTGTACATGTTGCTCA ACTGGCTGCCGTCGCTGTTGATCGGCCAAGGCTTCAGCCGGCCGCAGGCGGGCG CCGTGCAGATTCTGTTCAACCTGGGTGGGGCGGCTGGATCGTTCCTGACCGGTCG GATGATGGACCGCGGGTTTGCCGGGCGGGCGGTATTGATCGCCTACGCCGGCAT GCTGGCGTCGTTGGCGGGGCTGGGGCTGTCGAGCAGTTTCGGGCTGATGTTGCTG GCCGGGTTCACTGCAGGGTATTGCGCCATTGGTGGCCAGTTGGTGCTGTACGCGC TGGCGCCGACGCTGTATTCGACCCAGGTGCGCGCAACGGGCTTGGGCGCTTCGGT GGCGGTGGGGCGCCTGGGTTCCATGGCCGGTCCTCTGGCGGCCGGGCAGATCCTT GCAGCAGGTGCAGGTGCGGGGGGGCTGCTGATGGCCGCGGCTCCAGGCCTGGTG CTCGCGGCGTTGGCGGCACGTGCGCTACTGAAACACCGAACCCGCACGCAACGC CCGGAGCCGGGTGAGGCTGCAGCCCAGGACCCTGCATAG PP_3350-SEQ ID NO: 54 ATGATCAAGCCGTTGAACGGGGTATGCCTGCCCCTCATGGTGCTTTCGCTGTGCC ATTCCGCCTACGCGGCCAGCCCGCTTGAAGAGAGCCGGCCGGGGAGGCCAGGGG CCCCTCGCTGGATCGAGGACTACCGCTTTCTCGACGACCCCGCCAAGGCCACTGC CCCTTTCGATGGCCTGCGCTACCATCGGCTGTCAGAGTCGGCCTGGCTGCAACTC GGGGCTGAGGCCCGCTACCGGGCCGATGCTTCGGACAAACCCTTCTTTGGCCTGC GCGGGCTGAACGACGATTCCTATCTGCAGCAACGGCTGCAGGCCCATGCCGACC TGCACCTGTTCGACGATGCCGTGCGTACCTTCGTCCAGGTCGAGAATACCCGTGC CTTTGGCAAGGACTTGTACTCGCCCAATGATGAAAGCCGCAATGAAGTGCGACA GGCCTTCGTCGATTTCAACCACGATTTCGCCGCCGGGCGCTATACCACACGCGTA Att Dkt: 114198-5960 GGTCGCCAGGAGATGGGCTTCGGCGACCAGGTGTTGGTCACCTACCGCGACGCG CCGAACATCCGCCTGAGCTTCGACGGCGTGCGCGCCAGCCTCAACCTCAAGGAC GGACGCAAGCTGGATGCATTCGCTGTGCGCCCGCTGAAAACCGGTGAAGACAGT TGGGACGATGGCAGCAACAACAACATCAAATTCTACGGCCTCTACGGCACCTTG CCGCTGTCGGCGTCGTGGAACATGGACCTGCTAGCCTTTGGCCTGGAAACCGACG ACCGCACCCTGGCCGGCGAGTCGGGTGACGAGCAGCGCTACACTTTCGGCACGC GGCTGTTCGGCCGCCATCAGGCACTGGACTGGAGCTGGAACCTCGTCGGGCAGA CCGGCCACCTGGGCAACGCGTCGATCCGCGCCTGGGCGCTGTCGAGCGACAGCG GCTTCACCTTCACCCACCCCTGGCAACCGCGCCTGGCCATGCGCATCGATGCCGC AAGCGGCGACAGTGACCTGGGCGACGGAAAAGTCGGCACCTTCGACCCGCTTTA CCCACGCAATGGTGTCTATGGCGAAGCCAGCCTCACCACCCTGAGCAACATCATC GTGGTGGGGCCGACCTTCGGCTTCTCGCCCTGGCGCACGCTGCGCATCGAGCCCG GCATCTTCGAAGTGTGGAAACAGCGTGAGGAAGACGGCGTGTACATGCCCGGCA TGAGCATGCTGGCCAACACCCGCGGCACGGGTCGCCATGTCGGTACCATCTACCG CGCCAGCACGCGCTGGCTTGCCACTCCCAACCTGACCCTGGACCTTGACCTGAAG TACTACGACGTGGGCACGGCCATCAAGGAGGCTGGCGGCGAAGATTCGTCGTTT GTCTCAGTACGCGCCACGTTCCGCCTTTGA atsA-SEQ ID NO: 55 ATGAATTGCCTCCACCCCCTTGCCGCCCTCGGCGTGCTCATGCTCGGTGCCAGCC ACGCCATGGCCGCACCAACCGTTAAACAACCAAACATTCTGCTGGTGGTCGCTGA TGACCTCGGTTTCTCTGACCTGTCGGCCTTGGGCAGCGAAATCCAGACGCCCAAC CTGGACAAACTGTTGGCCGACGGCCGCCTGATGCTCGATTTCCACGCCGCGCCCA GTTGCTCACCGACGCGGGCGATGCTGATGTCCGGCAACGACCCACACCTGGTCG GGCTGGGCATGATGTCGGAGACGCGCAAGCGCCTGCCGGCCGGCGCAGAGGTAC GCGCCGGTTACGAAGGCGTGCTCGGTGCCAATGTCGCGACCCTCGCCGAACTGCT GCATGACGCCGGCTACCGCACCTACATCGCCGGCAAATGGCACCTGGGCATGCA ACCGGCCCAACAGCCACACAACCGCGGCTTCGACCAGGCCTACGTGCTGCTCGA AGGCGGCGCGGCACACTTCAAGCAGGCCAACATGGACCTGCAGTCGAACTACAG CGCCACCTATCGCCACAACGGCGAGCCGATCGCGCTGCCCGACGACTTCTATTCC ACCACCTGGTACACCGACCGCCTGATCGACGCGATCGACCAAGGCAAGGCCAGC AACAAACCGTTCTTCGCCTTTGCCGCCTACACCGCGCCGCACTGGCCGCTGCAGG CGCCGGATGCGTACCTGGCGCAGTACCGAGGCCGCTACGACCAAGGCTACGAAG TCATCGCCCACGAACGCCTGAAGCGCCAACGCGAAGCCGGCCTGGTGCCCGCCG ACTTCTCTCTGCAAGCACAACTGGAAGGTGTGCCGAGCTGGGCATCGCTGAGCG ATGAACAAAAACGCCAGTCCGCACGCACCATGGAGGTGTATGCGGCCATGGTCA CGGCGCTGGACGCCGAGGTGGGGCGCCTGGTGGAACACCTACGCAAGACCGGCG AGCTGGAAAACACGATGATCCTGTTCATGTCCGACAACGGTCCGGAGAACGCCA CCCGTGTCGACCCCGAGTGGATCAAGGCAAACTTCGATAACCGCCTGGACAACT ACGGCCGCCGCGATTCGTTCCTGATGATCGGCTCGGCCTGGGCCCAGGTCAGCGC CCTGCCCAGCCGTCGCTACAAGATGACCACCTACCAGGGCGGCATCCGCGTACC GGCCTTCGTCTACTACCCGGCCACCATCAAGGCCGGCCGCAGCGAAGAGCTGGC CAGCGTCCGCGACATCATGCCGACCCTGCTGCAACTGGCCGGTGCGCCGATGCCC Att Dkt: 114198-5960 GGCGTCAAGTACCGCGACCGCGATATCGTGCCGCCGCAAGGCACATCGGCGCTG AACTACCTGCAAGGCAAGGACAGCCAGATCCACGCCCATGACGAAGCCTTGGGG TGGGAAATCAACGGCAGCGCCTCGTTGCGCAAGGGCCCGCTGAAGCTGGTGTGG GATGCCACCGACAGCAAGCCGAGCTGGCACTTGTATGACCTGCGCAAGGACCCT GGCGAGCACCACGACATCGGCGCAGGGCACCCCGAGCAGGTGGCGCTGATGCTG CGCGACTGGAAGGCTTATGCCCAACGCGGCAACATCGCGATCGATGCCGACGGC CGTCCGGCCAGCCCGATCCTGGTCAAGCCGCAATGA PP_3353-SEQ ID NO: 56 ATGATCGCCCGGGCAGCCGGCTGCGCAGCACTCGCCACCGGCCTGCTCGCCGCC GCCTATGCCGTGTGGCCGCAAGCCCCGGCGCCGGCGCTGCTGTGTACGGGCTACA GCGGCCTGCCCGCGCAGCGCCAAGGCCCTTGGCGAGACGGCATGGTGCACGTGC CGGGCGGCGAGTTCAGCTTTGGTTCAAGCCGCTTTTACGACGAAGAAGGCCCGCC TCACCCCGCCAAGGTGTCCGGCTTCTGGATTGACGTGCATCCGGTCACCAACGCC CAGTTCGCGCGCTTCGTCAAGGCCACGGGGTATGTCACCCATGCCGAGCGCGGTA CCCGTGTCGAGGACGACCCTGCCCTGCCCGACGCGCTGCGGATACCGGGTGCGA TGGTGTTTCATCAGGGTGCGGACGTGCTCGGCCCCGGCTGGCAGTTCGTGCCCGG CGCCAACTGGCGACACCCGCAAGGGCCGGGCAGCAGCCTGGCCGGGCTGGACAA CCATCCGGTGGTGCAGATCGCCCTGGAAGATGCCCAGGCCTATGCCCGCTGGGC AGGCCGCGAACTGCCCAGCGAGGCGCAGCTGGAATACGCCATGCGCGGCGGCCT GACCGATGCCGACTTCAGCTGGGGTACCACCGAGCAGCCCAAGGGCAAGCTCAT GGCCAATACCTGGCAGGGTCAGTTCCCTTATCGCAATGCGGCGAAGGATGGTTTT ACCGGTACATCGCCCGTGGGTTGCTTCCCGGCCAACGGCTTTGGCCTGTTCGATG CCGGCGGCAATGTCTGGGAGCTGACTCGCACGGGCTATCGGCCAGGCCATGACG CACAGCGCGACGCCAAGCTCGACCCCTCAGGCCCGGCCCTGAGTGACAGCTTCG ACCCGGCAGACCCCGGCGTGCCGGTGGCGGTAATCAAAGGCGGCTCGCACCTGT GTTCGGCGGACCGCTGCATGCGCTACCGCCCCTCGGCACGCCAGCCGCAGCCGGT GTTCATGACGACCTCGCACGTGGGTTTCAGAACGATTCGGCAATGA PP_3354-SEQ ID NO: 57 ATGACCTGGAAGGCGCCCCTGCGCGACATGCAGTTCGTCCTGCAACACTGGCTGC GGGCCGATCAGGCCTGGGCGCAATTACCGGCATATGCCGACGTCGACCTCGACC TGGCGGCCCAGGTACTGGAAGAGGCGGCCCGTTTCTGCGAGCAGGTGCTGGCCC CTTTGAACAGTTCGGGCGATCGCCAGGGCTGCCAGATGGAGGGCGGGGAGGTAC GCACCCCGGACGGATTCCCGGCGGCTTATCGCGCCTATGTCGACGGTGGCTGGCC GTCGCTGGCATGCGCCCCGGCGCTTGGCGGCCAAGGGCTGCCGTTAGTGCTGGA GGCGGCCTTGCAAGAGATGCTGTTCGGCAGCAACCACGGCTGGGCCATGTACAC CGGCATCGCCCACGGTGCCTACCTGTGCCTGAAAGCCCATGCATCCGCCGAGCTG CAGGCGCGCTATCTGCCGGGCATCGTCAGTGGTGCCACCCTGCCTACCATGTGCC TGACCGAACCCCAGGCCGGCAGCGATCTCAGCCTGTTGCGCTGCCGCGCCGTGCC CCTGGGCGACGGTGGCGACGGTGGCTACGCGGTCAGTGGCAACAAGCTGTTCAT CTCTGGCGGCGAACACGACCTGACCCGCGACATACTCCACCTGGTGTTGGCCCGG TTGCCGGATGCGCCGCCCGGCAACCGCGGGCTCTCGCTGTTGCTGGTGCCGAAAT Att Dkt: 114198-5960 GGCTGGAGGATGGCCAGCGCAATACACTGTACTGCGACGGCCTGGAACATAAGC TGGGCATCCGGGGCAGCGCCACCTGTGCACTACGCCTGGAGGGGGCCCGGGGCT GGTTGATAGGCCAGCCCCATGGCGGCCTGGCGGCGATGTTCGTGATGATGAATTC CGCGCGTCTGCATGTTGGCCTGCAAGGGCTGGGGCATGCTGAGGCCGCCTGGCA ACAGGCCCGCGACTATGCCCTGGAGCGGCGCCAGATGAGTGCCCCGGGGCGGGC CAAGGGGGGCGCCCAGGGCGCCGACCCGATCCACCTGCACCCGGCCATGCGTCG GGTCTTGCTGGAGTTGCGGGTGCGTACCGAGGGGATGCGCGCGCTGGCCTACTG GACGGCCCAGTGGCTGGATATCGCCGACGCCGCCGTCGAGCCTGCACAACGGGC GAAGGCTGCGACATTGGCGGCGCTGTTGACCCCCATCGTCAAGGCGGCCTTCACC GAGCAGGGGTTCAGCCTGGCAAGCAAGGCCTTGCAGGTGTTGGGTGGCTATGGC TACACCCAGGAGTTTGCCATCGAGCAGACCTTGCGCGACAGCCGCATCGCGATG ATCTACGAAGGCACCAACGAAATACAGGCCAACGACCTGATGCTGCGCAAGGTA TTGGGCGATGGGGGGGAAGGTTTCGCCCTGCTGTTGAGTGAGTTGCGTACCGAG GCCGCTGCGGACGCTGGCGCCGCCGCCGCCGCGTTGGGCAGTGCGTGCGACACG CTGGAGGCAGCGCTGGCCAAGGTGCTGTCGATGACCGCGGGCGATCCCGAATTT CCGTACCGCGCCGCAGGCGACTTCCTGCAGTTGTGCGCAACGGTGCTGCAGGCCT TCGCCTGGGCCCGCAGCGCGCGCTGCATCCAGGCGTTGCCAGTGGGGGCGGCGC AGCGCCGGGAAAAGCAGGAAAGCGTTGGGTTCTTCCTCGGCTACCTGCTGCCCG ACCTCGACCGGCAGGTGGCGGCCATCGACGCGGCCGCCACGCCGTTGCCATTCAT TGCCGAATCGTTCTGA PP_3355-SEQ ID NO: 58 ATGGTGCGCACGCCCTGGGTCGACCTGGGTGGCGCATTGGCCGCGATCTCGCCCA TCGACCTGGGCATTCAGGCCGGCCGCGCGGTGCTGGCACGTGCAGCTGCCGCCCC TCAGGCGGTAGACAGTGTGCTCGCCGGCAGCATGGCTCAGGCCAGCTTCGACGC CTACATGCTGCCACGCCATGTCGGCCTCTACTGTGGCGTGCCGCAGGCGGTACCG GCACTGGCGGTGCAGCGCATCTGTGGCACGGGCCTTGAACTGTTGCGTCAGGCCG GTGAACAGCTGCGCAGTGGCGCCCGCCAGGTGCTGTGCGTGGGCGCGGAGTCGA TGTCGCGCAATCCGATCGCGGCCTATGAACACCGGGGCGGTTTTCGCCTGGGGGC GCCGGTCGGGTTCAAGGACTTTCTCTGGGAAGCCCTGTACGACCCGGCTGCCGGC GTGGACATGATCGGTACTGCGGAAAACCTGGCCCGTGCCTATGGCCTGGCGCGT GAAACGGTGGACGCGTGGGCCCTGCGCAGCCATCAGCGGGCCCTGCAAGCCCAG GTGCAGGGCTGGTTCGACGAGGAAATCGTCAGCGTCACCGCGCAGGCACTGGAG GTAGAAGGCTGCCAGCCACGCGGCATTGCATTGCCGCGTGGCGTGAGTGAGGTC AGCCAGGACAGCCACCCACGGCCTACCGACGCGGCCGCACTGGCCCGACTGCGC CCGGTGCACGCCGGTGGCGTGCAGACGGCTGGCAACAGTTGCGCGGTGGTCGAT GGCGCGGCGGCAGCCTTGGTCACGCGCTACTCCGATTGCACTCATCCGCCTTTGG CGCGCTTGCTGATGGCCACGGCTGTTGGCGTGCCGCCCGGCCTCATGGGCATCGG GCCGGCGCCGGCCATTACCTTGCTGCTCGAGCGCAGCGGCCTGCGCCTGGAGCA GATCGACCGCCTGGAAATCAACGAGGCCCAGGCCGCCCAAGTGCTGGCGGTGGC CCAGGCGCTCGAACTCGATGTGGACAAGCTCAATGTCCATGGCGGCGCTATCGCC CTTGGTCACCCATTGGCCGCTACAGGTCTGCGCCTGGTCCACACCTTGGCCCGTC Att Dkt: 114198-5960 AATTGCGCGACGGCAACTTGCGCTATGGCATCGCAGCGGCATGCATCGGCGGTG GCCAAGGCATGGCGCTGCTGATCGAAAACCCTCATTTCGTTGCCTGA fcs-SEQ ID NO: 59 GTGAATAACGAAGCCCGCTCAGGGTCGACCGACCCTGGCCAACGTCCGCGCTAC CGCCAGGTGGCCATCGGGCATCCCCAGGTGCAGGTCAGTCACGTCGACGACGTG CTGCGCATGCAACCTGTCGAGCCACTGGCGCCGCTGCCGGCGCGCCTGCTCGAGC GCCTGGTGCATTGGGCCCAGGTGCGCCCGGACACCACTTTCATCGCGGCACGCCA GGCAGACGGTGCCTGGCGTTCGATCAGCTACGTGCAGATGCTCGCCGATGTGCGC ACCATCGCCGCCAACTTGCTAGGACTGGGCCTCAGTGCCGAGCGCCCGCTGGCGC TGCTTTCCGGCAACGACATCGAACACCTGCAAATCGCCCTCGGCGCCATGTATGC CGGTATTGCCTATTGCCCGGTGTCGCCGGCCTACGCGCTGTTGTCGCAAGACTTC GCCAAGTTGCGCCATGTCTGCGAGGTGCTCACCCCCGGAGTGGTCTTCGTCAGCG ACAGCCAGCCGTTCCAGCGCGCCTTCGAGGCGGTGCTGGACGATTCGGTCGGCGT GATCAGCGTGCGTGGCCAGGTCGCAGGTCGCCCCCATATAAGCTTCGACAGCCTG TTGCAACCGGGTGACCTGGCGGCGGCCGATGCGGCTTTCGCCGCCACCGGGCCG GACACCATCGCCAAATTCCTCTTCACCTCGGGCTCGACCAAGCTGCCCAAGGCGG TGATCACCACCCAGCGCATGCTGTGCGCCAATCAGCAGATGCTTCTGCAGACTTT TCCGACGTTCGCCGAGGAGCCGCCGGTGCTGGTGGACTGGCTGCCGTGGAACCA CACGTTCGGCGGTAGCCACAACCTCGGCATCGTGCTTTACAACGGGGGCAGTTTC TACCTGGACGCCGGCAAGCCGACCCCGCAAGGCTTCGCCGAGACCTTGCGCAAT CTGCGCGAGATTTCCCCCACGGCCTACCTCACCGTACCCAAGGGCTGGGAGGAA CTGGTCAAGGCACTGGAGCAGGACCCCGCGCTACGCGAGGTGTTCTTTGCCCGCA TCAAGCTGTTCTTCTTTGCCGCCGCAGGCCTGTCGCAAAGCGTCTGGGACCGGCT GGACCGCATTGCCGAGCAACACTGTGGCGAACGCATCCGCATGATGGCCGGCCT TGGCATGACCGAAGCCTCGCCATCGTGCACCTTCACCACCGGGCCTTTGTCGATG GCCGGCTATGTCGGGCTGCCGGCACCTGGCTGCGAAGTGAAGCTGGTGCCGGTG GGCGACAAGCTCGAGGCGCGCTTCCGTGGCCCGCATATCATGCCGGGCTACTGG CGCTCGCCGCAGCAGACCGCCGAGGCGTTCGACGAGGAGGGCTTCTACTGTTCG GGCGACGCGTTGAAGCTGGCCGATGCCAGGCAGCCCGAGCTTGGCCTGATGTTC GATGGCCGTATCGCTGAGGACTTCAAACTTTCGTCCGGGGTATTCGTCAGTGTCG GGCCGCTGCGCAACCGCGCAGTGCTGGAGGGCTCGCCTTACGTACAGGACATCG TGGTCACCGCGCCGGACCGTGAATGCCTGGGCCTGCTGGTGTTCCCGCGTCTGCC CGAGTGTCGGCGCCTGGCCGGGCTGGCAGAGGATGCCAGCGATGCGCGGGTGCT GGCCAACGACACCGTGCGCAGTTGGTTCGCTGACTGGCTGGAGCGCTTGAACCG CGATGCCCAAGGCAACGCCAGCCGTATCGAATGGCTGTCGCTGCTGGCCGAGCC GCCGTCGATCGACGCCGGTGAAATCACCGACAAGGGCTCGATCAATCAGCGCGC CGTGCTGCAGCGGCGCGCCGCTCAGGTCGAGGCGCTGTACCGTGGCGAAGACCC CGACGCATTGCACGCCAAGGTGCGGCCTTGA vdh-SEQ ID NO: 60 ATGTTGCAGGTGCCTTTGCTGATTGGCGGGCAGTCGCGCCCCGCCAGCGATGGAC GAACCTTCGAGCGCTGTAACCCGGTGACTGGCGAGGTGGTGTCGCAGGCTGCCG Att Dkt: 114198-5960 CCGCCACACTGGCCGATGCCGATGCCGCGGTGGCTGCTGCCAGCGCGGCGTTTCC GGCCTGGGCCGCCCTGGCACCGGGCGAGCGGCGCAGCCGCTTGCTGGCAGGCGC TGATCTGTTGCAGGCGAGGGCCGCCGAGTTCATCGCCGCCGCCGGTGAAACCGG GGCCATGGCCAACTGGTATGGCTTCAACGTGAAGTTGGCCGCCAACATGCTGCGC GAGGCTGCAGCCATGACCACGCAGATCACCGGTGAAGTGATCCCCTCGGACGTT CCCGGCAGCTTCGCAATGGCCCTGCGCGCGCCCTGCGGCGTGGTGTTGGGCATCG CACCGTGGAACGCCCCGGTGATACTGGCCACGCGTGCCATTGCCATGCCGCTGGC CTGCGGCAACACCGTGGTGCTCAAGGCCTCGGAGCTGAGCCCGGCGGTCCATCG GCTGATCGGCCAGGTGCTGCACGATGCAGGCATCGGCGACGGCGTGGTCAATGT CATCAGCAATGCGCCGCAGGATGCCCCCGCCATCGTCGAGCGGCTGATCGCCAA CCCTGCGGTACGCCGGGTCAACTTCACCGGTTCGACGCACGTCGGGCGCATCGTC GGCGAACTGGCGGCCCGCCATCTCAAGCCGGCCCTGCTCGAACTGGGCGGCAAG GCACCTTTGCTGGTGCTCGACGATGCCGACCTGGACGCCACGGTCGAAGCGGCG GCCTTCGGTGCCTACTTCAACCAGGGGCAGATCTGCATGTCCACCGAGCGCCTTG TGGTGGACAGCTGTATTGCCGACGCTTTCGTCGACAAGCTGGCGGTGAAGATCGC CGGGCTGCGTGCAGGTGATCCGCAAGCCAGCACCTCGGTGCTCGGCTCGCTGGTC AGCGCAGCGGCCGGCGAGCGCATCAAGGCACTGATCGACGATGCCGTGGCCAAG GGCGCGCGCCTGGTCAGCGGCGGCCAGCTGGAAGGCAGCATCCTGCAACCGACC TTGCTCGACAACGTCGATGCCAGCATGCGCCTGTACCGCGAGGAGTCCTTCGGCC CGGTGGCGGTGGTACTGCGCGCCGAAGGCGACGAAGCCTTGCTGCAGCTGGCCA ACGACTCGGAGTTCGGTCTGTCATCGGCCATTTTCAGCCGCGACACCAGCCGCGC CCTGGCCTTGGCCCAACGGGTGGAGTCGGGTATCTGCCATATCAACGGCCCGACC GTTCACGATGAAGCGCAGATGCCGTTTGGCGGGGTCAAGTCCAGCGGCTATGGC AGCTTCGGCAGCCGCACGGCCATCGATCAGTTCACCCAGTTGCGCTGGGTCACCC TCCAGCACGGCCCGCGTCACTATCCCATCTAG PP_3358-SEQ ID NO: 61 ATGAGCAAATACGAAGGCCGCTGGACCACCGTGAAGGTCGAACTGGAAGCGGGC ATCGCCTGGGTGACCCTCAATCGCCCGGAAAAACGCAATGCCATGAGCCCCACC CTGAACCGGGAAATGGTCGACGTGCTGGAAACCCTTGAGCAGGACGCTGACGCT GGCGTGCTGGTATTGACCGGTGCCGGCGAGTCCTGGACCGCCGGCATGGACCTG AAGGAGTACTTCCGCGAGGTGGACGCCGGCCCGGAAATCCTCCAGGAAAAGATT CGTCGCGAAGCCTCGCAATGGCAATGGAAGTTGCTGCGTCTGTATGCCAAACCG ACCATCGCCATGGTCAACGGCTGGTGCTTCGGCGGCGGCTTCAGCCCACTGGTGG CATGCGACCTGGCGATCTGCGCCAACGAAGCGACCTTCGGCCTGTCGGAAATCA ACTGGGGCATCCCGCCTGGTAACCTGGTCAGCAAGGCCATGGCCGATACCGTTG GCCATCGTCAGTCGCTGTACTACATCATGACCGGCAAGACCTTCGATGGTCGCAA GGCTGCCGAGATGGGCCTGGTGAACGACAGTGTGCCGCTGGCCGAGCTGCGTGA AACCACCCGCGAGTTGGCGCTGAACCTGCTGGAAAAGAACCCGGTGGTGCTGCG TGCCGCGAAGAATGGCTTCAAGCGTTGCCGCGAGCTGACCTGGGAACAGAACGA GGACTACCTCTACGCCAAGCTCGACCAGTCGCGCCTGCTGGACACTACCGGCGGC CGCGAGCAGGGCATGAAGCAGTTCCTCGACGACAAGAGCATCAAGCCAGGCCTG CAGGCCTACAAGCGCTGA Att Dkt: 114198-5960 PP_3359-SEQ ID NO: 62 ATGGCTAGGTCTGCCCGTAGTACCGACGACGCTTGCGTGGCTGCTCCTGTGGGAG AAGGGGTGCTTGAAGACTTGATCGGCTACGCCTTGCGACGCGCGCAATTGAAGC TGTTTCAGAACCTTATTGCCCGGCTCTCGGCCCATGACCTGCGCCCGGCCCAATTT TCCGCCCTGGCGATCATCGACCAGAACCCCGGGCTGATGCAGGCCGACCTGGCG CGTGCGTTGGCAATCGACCCCCCGCAAGTCGTGCCAATGCTGAACAAACTGGAA GAGCGCGCGCTGGCCGTGCGCGTGCGGTGCAAACCGGACAAGCGCTCGTATGGG ATTTTCCTGAGCAAATCGGGCGAGGCCCTGTTGAAGGAGTTGAAGCACATCGCC GCCGACAGCGATCACCAGGCGACATCCAACCTCTCGGATGACGAAAGGACTGAA CTGTTGAGGTTATTGAAGAAAATCTACCGGGACTGA PP_3360-SEQ ID NO: 63 ATGAACGACCGCATCATCGCGTTTTTCACCGAAACCACCGTATTCGGCATCTCCA TTCTCAATCTGCTGGTGGCGCTGACAGTCGCCACCTTGGTCTTCCTGGTCGCCCGT GCCGCCATCAGTTTCCTGATGCGCCGGGTACGGCGCTGGTCCGAACACGAAGGC CCACTGAGCCAGATAATGGCCAAGGTGCTTTCCGGCACCAGCAACTTTCTGTTGT TGCTGGCCTCCTTGCTGGTGGGCCTGAGCATGCTCGACCTGCCCGAGCGCTGGCT GACCAGGGTCAGCAGCCTGTGGTTCGTGGTCGCGGCCTTGCAGATTGGCCTGTGG GCCAACCGCGCCATCGGCCTGGGCCTGAGCCGTTACTTCGCCCGCCACCGTACCG ACGGCCTGAACCAGGGCAGCGCACTGGCAACCCTTTCAGCCTGGGGCGCGCGGG TCCTGTTGTGGTCAGTCGTAGTGCTGGCGATGCTGTCGAACCTGGGGGTGAACAT CACCGCCTTCGTCGCCAGCCTGGGCGTGGGCGGCATCGCAGTGGCGCTGGCTGTG CAGAATATCCTTGGCGACCTGTTCGCTTCGTTGTCGATCGCCGTGGACAAGCCGT TCGAGATCGGTGACTTCATCGTCATCGGGCCCCTGGCGGGCACCGTTGAGCACGT CGGCCTCAAAACCACACGCATCCGCAGCCTCGGCGGCGAACAGATCGTCATGGC CAACGCCAGCATGATCAGCAGCACCATCCAGAATTACAAGCGCCTGCAAGAGCG TCGCATCGTGTTCGAGTTCGGGCTTTCCTACGACACGCCCACCGAGGCCGTGAAG AAAGCGCCAGCCATCGTAGAAGACGCCATCAAGGCCCAGGAGCAGGCGCGTTTC GATCGCGCTCACCTGCGCGGCTTCGGCAAAGAGGCGCTGGAGTTCGAGTGTGTCT ACATCGTCCAGGACCCAGGCTACAACCTTTACATGGACATTCAGCAGGCGATCA ACTTCAGGCTGCTGGAGCGGTTCTCGCAAATCGGCGCCAAGTTCGCCGTGCCGGT GCGCGCGATCAAGGTCACGGCATTGCCAGAGGCGTCCGGCCTGCATGGCCAGGC TCGCGAAGGCCTGAGCATCAGGGGCGCTTGA PP_3536-SEQ ID NO: 64 GTGGAGTGTCGATTGTCTTTTAACATGGAACGTAGAGAGGGCACCACGATGAGC GAACTCCCCACCACAAGTTTTTTGATCTTCCATGATCCACACGCGCCTTACCTGGT CGAGACACTTGATCCGAAGGAAGTGGCGCAGCTTGAAGCCGACCATGAGAAATC CAAAATAGACGACATTGAACAACTGAAAGCGCAATTTTTACAGATGGAACGGGA AAAAATGCTCCCCAAGGAGGATCAGGAGGCGCTTTACAACGCCAGACACGCCAA AGCAGATCTTACAGACCCTGTGAGTGAGTGGGGCCCGTGGGGCGAGTTGAGGAA TGTATCTATTTCAAGTGTGCCTCATGACTTGGAGCCTGAGACACCAGAGTCTCCA Att Dkt: 114198-5960 GCTGAGGTGGAGGCGCTTCAAATCAAGATCGACGCTCTTGAGCGACAGGCCCGT CAATACGCGTTAGATCCAGCGATAGAATCTCGTTTGGACGAGGTCGTTCAAGCCG TCCAGGCCCTCAGCGCCGATGAAGATGATGCTCAGAAATACATTAGATATACTTT CCCCCGCGATGATGAAGCGTACTGGCGAAACTACGCTTTGTTGGTGCACAAAATT GTAAGTGCTAAATTTGAACTGCCGGATTTTCGGTATTCTCCCTCCTTAAGTCCCGA CCGCTTCGTGTTAATCAGGCGCGACGGTACTTTGGGTATCAATATGGCCGATGAA CTGAAAGCCCAGCCTTTACTAAAGTTCGAAGTGGTCGACGGTTGGTCCGCAGATG AAAGGTGGGGGGGGTATCTTCGCGTCTCAAGCTCAACTGAAAAGGGCAAGTACC TCTACACGAGAGACACATTGGATGAGGGCGCGTTTTTTAAGAATGACAGTCTAG ATAACGATGACAGTGGGTGGGTAGAACTGACTAGCTCTGAAGGGGCCAAAATCA CGTTTTGCAAAGTGGGAAACCATTACGAAATCTGGCAGAGCCTCAAGTCCGAAA TCTTCGGGCCCACGACCCGGTGGTTGACCGTGGTTGACGGCCGGCTGAAATTTGT AATGGAGGGATCCTCAGCTAAGTGGAACATTGTGTCTACAGACGTAGACTAA pobA-SEQ ID NO: 65 ATGAAAACTCAGGTTGCAATTATTGGTGCAGGTCCGTCTGGCCTGCTGCTGGGCC AGCTGCTGCACAAGGCCGGTATCGATAACATCATCGTCGAACGCCAGACTGCCG AGTACGTACTAGGCCGCATCCGCGCCGGGGTGCTAGAGCAAGGCACGGTCGACC TGCTGCGCGAGGCTGGCGTGGCCGAGCGCATGGACCGTGAAGGCCTGGTGCACG AGGGGGTTGAACTGCTGGTTGGCGGGCGCCGCCAGCGTCTGGATCTCAAAGCCC TGACCGGCGGCAAGACGGTGATGGTCTACGGCCAGACCGAAGTCACCCGTGACC TGATGCAGGCCCGCGAAGCCAGTGGTGCGCCGATCATTTATTCAGCCGCCAACGT TCAGCCGCATGAATTGAAAGGCGAGAAGCCCTACCTGACGTTCGAAAAGGATGG CCGGGTGCAGCGGATTGACTGCGACTATATCGCCGGCTGCGACGGCTTCCACGGT ATCTCGCGGCAGAGCATCCCGGAGGGCGTGCTGAAACAGTATGAGCGGGTTTAC CCGTTTGGCTGGCTGGGCCTGCTGTCGGACACACCGCCAGTCAATCACGAGTTGA TCTACGCCCACCATGAGCGCGGTTTCGCGTTGTGTAGCCAACGCTCGCAAACACG CAGCCGCTACTACCTGCAGGTACCTTTGCAGGATCGGGTCGAGGAGTGGTCTGAC GAGCGTTTCTGGGACGAACTGAAAGCCCGTCTGCCCGCCGAGGTGGCGGCGGAC CTGGTCACAGGCCCCGCGTTGGAAAAAAGTATTGCGCCGCTGCGTAGCCTGGTG GTCGAACCCATGCAGTATGGTCACCTGTTCCTGGTGGGGGACGCGGCGCACATCG TCCCCCCTACGGGTGCCAAAGGCCTTAACCTGGCGGCCTCCGACGTCAACTACCT GTACCGCATTCTGGTCAAGGTGTACCACGAAGGGCGCGTCGACCTGCTTGCGCAA TACTCGCCGCTGGCACTGCGCCGCGTGTGGAAGGGCGAGCGCTTCAGCTGGTTCA TGACCCAACTGCTGCATGACTTCGGTAGCCACAAGGACGCCTGGGACCAGAAGA TGCAGGAAGCTGACCGCGAGTACTTCCTGACCTCGCCGGCGGGCCTGGTGAACA TTGCCGAGAACTATGTGGGGCTGCCGTTCGAGGAAGTTGCCTGA PP_4557-SEQ ID NO: 66 ATGAAACCATTACTGATTCTTGCATTGGTCGGTATGTCCTCGTTCGCCCTCGCCGA TGAGGCCAAACCCGTCTCCAGCGAGCCTGTGGCTCAGCAGTACGACTACTCGATG AACCTGGACATCAAGCGGGTGATCAACCTGTCGACCATCCCCAATGTCTGCGAA Att Dkt: 114198-5960 GTGGTACCTGCGACCATGACCTATGAAGACCATCAGGGCCAGGTGCACACCATT CAGTACCGGGCCATGGGTGAAGGTTGCCAGCAGGGTTGA benA-SEQ ID NO: 67 ATGTCCCTGGGATTCGACTACCTCAATGCCATGCTCGAGGACGACCGTGAAAAG GGCATCTACCGCTGCAAGCGGGAGATGTTCACCGACCCACGGCTGTTCGACCTGG AGATGAAACACATCTTCGAGGGCAACTGGATCTACCTGGCCCACGAAAGCCAGA TCCCCGAGAAGAATGACTTCCTCACCCTGACCATGGGGCGCCAGCCGATTTTCAT CGCGCGCAACAAGGACGGTGAGCTCAATGCCTTCCTCAACGCCTGCAGCCACCG CGGTGCCATGCTGTGCCGGCACAAGCGCGGCAACCGTTCCAGCTATACCTGCCCG TTCCACGGCTGGACGTTCAACAACAGCGGCAAACTGCTGAAAGTGAAAGACCCC AGCAACGCCGGCTACCCCGACAGCTTCAACTGCGACGGCTCCCACGACCTGACC AAGGTGGCACGCTTCGAGTCGTACCGGGGCTTCTTGTTCGGCAGCCTGAACGCCG ATGTGAAGCCGCTGGTCGAGCACCTGGGCGAGTCGGCGAAGATCATCGACATGA TCGTCGATCAGTCGCCTGAAGGCCTGGAAGTGCTGCGCGGTGCCAGCTCGTACAT CTACGAAGGCAACTGGAAGCTCACCGCCGAAAACGGCGCTGACGGCTACCACGT AAGCTCCGTGCACTGGAACTACGCCGCCACCCAGAACCAGCGCAAGCAGCGCGA GGCGGGTGACGAGATCAAGACCATGAGTGCTGGCGCCTGGGCCAAGCAGGGCGG CGGTTTCTATTCCTTCGACCACGGCCACCTGCTGCTGTGGACCCGTTGGGCCAAC CCGGAAGACCGCCCGGCCTATGAGCGCCGCGACCAACTGGCCGCCGATTTCGGC CAGGCACGTGCCGACTGGATGATCGAAAACTCGCGCAACCTGTGCCTGTACCCC AACGTGTACCTGATGGACCAGTTCAGCTCGCAGATCCGCGTTGCCCGGCCGATTT CGGTGAACAAGACCGAAATCACCATCTACTGCATCGCGCCGAAAGGCGAGAGCG CCGATGCTCGTGCCAAGCGTATTCGCCAGTACGAAGACTTCTTCAACGTCAGCGG CATGGCCACCCCGGACGACCTGGAAGAATTCCGCTCGTGCCAGACCGGCTATGG CGGCGGCACTGGCTGGAACGACATGTCCCGTGGCGCGAAACACTGGGTCGAAGG CGCCGATGAGGCGGCCAAGGAGATTGAACTCGAACCCCTGCTGTCGGGCGTGCG CACCGAGGACGAAGGCCTGTTCGTGCTGCAGCACAAGTACTGGCAGGACACCAT GATCCAGGCCCTCAAGGACGAACAGCAGCTGATCCCCGTGGAGGCCGTGCAATG A benB-SEQ ID NO: 68 ATGAGCCTGTATGACACCGTGCGCGACTTCCTGTACCGCGAAGCGCGCTACCTGG ACGATGCCCAGTGGGACCAGTGGCTGGAACTGTACGCCAGCGATGCCAGCTTCT GGATGCCGAGCTGGGACGATGACGACACCCTCACCGAAGACCCACAAAGCGAAA TCTCGCTGATCTGGTACGGCAACCGTGGCGGCCTGGAAGACCGTGTGTTCCGCAT CAAGACCGAGCGCTCCAGTGCGACCGTGCCCGACACCCGCACCTCGCACAACAT CAGCAACATCGAGATCGTCGAGCAGGGCGAGGACAACTGCCAGGTGCGCTTCAA CTGGCACACCCTGAGCTTCCGCTACAAGACCACCGACAGTTACTTCGGCACCAGC TTCTACACCCTCGACCTGCGCGGCGAGCAGCCGCTGATCAAGGCCAAGAAGGTG GTGCTGAAGAACGACTACGTCCGCCAGGTCATCGACATCTACCACATCTGA benC-SEQ ID NO: 69 Att Dkt: 114198-5960 ATGAGCTACCAGATCGCACTGAATTTCGAAGACGGGGTTACCCGTTTCATTGAGG CCACCGGTCAGGAAACCGTGGCCGACGCCGCTTACCGCCAAGGCATGAACATTC CGCTGGACTGCCGCGACGGCGCTTGCGGGACCTGCAAGTGCAAGGCTGAGTCCG GTCGTTATGACCTTGGCGACAACTTCATCGAAGACGCCTTGAGCGAAGACGAAA TCGCCGAAGGTTACGTGCTGACGTGCCAGATGCGCGCCGAAAGCGACTGCGTCA TCCGCATCCCGGCTTCGTCGCAGCTGTGCAAGACCGAGCAGGCCACTTTCGAGGC CGCCATCAGCGATGTGCGCCAACTGTCGGTCAGCACCATTGCCCTGTCGATCAAG GGTGAGGCCCTGAGCCGTCTGGCCTTCCTCCCGGGGCAGTACGTCAACCTGAAGG TTCCGGGTAGCGAGCAGAGCCGGGCCTACTCGTTCAGCTCGCTGCAGAAGGACG GCGAAGTCAGCTTCCTGATCCGCAACGTACCGGGTGGGCTGATGAGCAGCTTCCT GACCAACCTTGCCAAAGCTGGCGACAGCCTGAGCCTGGCCGGGCCGTTGGGCAG CTTCTACCTGCGGCCGATCCAGCGCCCGCTGTTGCTGCTGGCCGGTGGTACCGGG CTGGCACCGTTCACCGCGATGCTGGAGAAGATCGCCGAGCAGGGCAGCGAGCAC CCGCTGCACCTGATCTACGGCGTGACCAACGACTTTGATCTGGTCGAACTCGACC GCCTGCAAGCGCTGGCGGCACGTATCCCCAACTTCACCTACAGCGCCTGCGTCGC CAACCCGGACAGCCAGTACCCGCAGAAGGGCTATGTCACCCAGCACATCGAGCC ACGCCACCTCAATGACGGCGATGTGGATGTGTACCTGTGCGGCCCGCCACCCATG GTGGAAGCGGTCAGCCAGTACGTGCGTGAGCAGGGCATTACCCCGGCCAACTTC TACTACGAGAAGTTTGCGGCTGCGGCCTGA benD-SEQ ID NO: 70 ATGAACAACAGATTCCAAGGCAAGGTTGCCCTGGTCACCGGTGCCGCCCAAGGC ATCGGCCGGGGCGTGTGCTGGCGGCTGAAGGCTGAAGGCGCGCAAGTGGTGGCG GTAGACCGTTCCGAGCTGGTCCACGAACTGGCTGGCGAGGGCATGCTCACCCTCA CCGCCGACCTGGAACAACACGCCGATTGCGCCCGGGTCATGGCCAGCGCGGTGG AAACCCATGGCCGCCTGGACATCCTGGTCAACAACGTGGGCGGCACCATCTGGG CCAAGCCCTTCGAGCACTACGAAGTCGAGCAGATCGAGGCCGAGGTGCGCCGCT CGCTGTTCCCCACCCTGTGGTGCTGCCACGCCGCGCTGCCGTACATGCTCGAGCG TGGCAGTGGCGCCATCGTCAATGTCTCGTCCATCGCTACCCGCGGCGTCAACCGT GTGCCTTACGGCGCAGCCAAGGGTGGCGTCAACGCGCTCACTGCCTGCCTGGCCT TCGAAACCGCCGGCCGGGGCATCCGCGTCAACGCCACTGCGCCCGGCGGCACCG AAGCACCGCCCCGGCGCGTACCGCGCAACAGTGCAGAGCAAACCGAGCAGGAA AAGCGCTGGTACCAGCAGATCGTCGACCAGACCCTCGACAGCAGCCTGATGCAC CGCTACGGCAGCATCGACGAGCAGGTGGGGGCCATTCTGTTCCTCGCGTCCGACG AAGCCTCGTACATCACCGGCGTGACCCTGCCGGTAGGGGGCGGCGACCTGGGCT GA benK-SEQ ID NO: 71 ATGCGAACCCTTGATGTACACCCGATCATCGACAATGCGCGCTTTACCCCGTTTC ACTGGATGGTCATGGCCTGGTGCGGCCTGCTGCTGATTTTCGACGGCTATGACCT GTTCATCTACGGTGTGGTACTGCCGGTCATCATGAAAGAGTGGGGCCTGACCCCG TTGCAGGCTGGTGCACTGGGCAGCTATGCGCTGTTCGGCATGATGTTCGGCGCCC TGGCCTTCGGCAGCCTGGCCGACCGCATCGGGCGCAAGAAGGGCATTGCCATCT Att Dkt: 114198-5960 GTTTCGCCTTGTTCTCGGGGGCGACCATCCTCAATGGCTTTGCCAGCAACCCGAG CGAGTTTGGCATCTACCGCTTCATCGCCGGCCTGGGCTGTGGCGGCCTGATGCCC AACGCCGTGGCACTGATGAACGAATACGCACCCAAGCGCCTGCGCAGCACGCTG GTGGCGATCATGTTCAGTGGTTATTCGCTGGGCGGCATGCTGTCGGCAGGTGTCG GCATCTTCATGCTGCCGCGTTTTGGCTGGGAGTCGATGTTCTTCGCCGCAGCGGT GCCACTGCTGCTGTTACCGGTGATTCTCTACTACCTGCCTGAATCCATCGGCTTCC TGGTACGCCAAGGCCGTATCGAGGAAGCACGCACATTGCTCAAACGGCTGGACC CGGACTGCGATGTCCAGGCCGATGATGTGCTTCAGGCGGCCGACCGCAAAGGCA GCGGCGCTTCGGTGCTGGAACTGTTCCGCAATGGCCTGGCGATTCGCACCCTGGC GCTGTGGGTTGCGTTCTTCTGCTGCCTGCTGATGGTCTACGCCCTGAGCTCCTGGC TGCCAAAGCTGATGGCCAACGCCGGCTACAGCCTGGGCTCAAGCCTGTCATTCCT GCTGGCACTGAACTTTGGCGGCATGGCCGGGGCGATCCTCGGCGGCTGGCTGGG TGACCGCTACAACCTGGTCAAGGTGAAGGTGGCGTTCTTCATTGCCGCCGCACTG TCGATCAGCCTGCTGGGGGTAAACAGCCCGATGCCGGTGCTGTACCTGCTGATCT TCATTGCCGGCGCCACCACCATTGGTACACAGATTCTGCTGTACGCCGGCGCTGC ACAGATGTACGGCCTGTCGGTGCGCTCTACCGGCCTGGGCTGGGCCTCGGGCATC GGCCGTAATGGCGCCATCGTCGGCCCGCTGCTGGGCGGTGCGCTGATGGGCATC AACCTGCCACTGCAACTCAATTTCATCGCCTTTGCCATCCCTGGCGCCATCGCTGC GCTGGCCATGGCCGTGCACCTGGCCAGTGGTCGGCGCCACGCCCAGACCCTGGC GGCGCAGGCCTGA catA-II-SEQ ID NO: 72 ATGACCGTGAACATTTCCCATACTGCCGAGGTACAGCAGTTCTTCGAGCAGGCCG CAGGCTTTTGTAATGCGGCCGGCAACCCACGCCTCAAACGCATCGTGCAGCGCCT GCTGCAGGATACCGCGCGGCTGATCGAAGACCTGGACATCAGCGAAGACGAGTT CTGGCACGCCGTCGATTACCTCAACCGCCTGGGCGGTCGCGGCGAAGCCGGGTT GCTGGTGGCGGGGCTGGGCATCGAACACTTCCTCGACCTGCTGCAGGATGCCAA GGACCAGGAGGCAGGGCGCGTTGGCGGCACCCCACGCACCATCGAAGGCCCGTT GTACGTGGCTGGCGCACCGATTGCCCAAGGTGAAGTGCGCATGGACGACGGCAG CGAGGAGGGCGTGGCCACGGTGATGTTCCTGGAAGGCCAGGTGCTGGACCCGCA CGGACGCCCGCTGCCGGGTGCCACGGTCGACCTGTGGCATGCCAATACCCGTGGT ACCTACTCGTTCTTCGACCAAAGCCAGTCGGCGTACAACCTGCGTCGGCGCATCG TTACCGATGCCCAGGGGCGCTACCGCGCGCGCTCCATCGTGCCATCGGGCTATGG CTGCGACCCGCAGGGGCCAACCCAGGAATGCCTGGACCTGCTGGGCCGTCATGG CCAGCGCCCGGCGCACGTGCACTTCTTTATCTCGGCCCCAGGGTACCGGCACCTG ACCACGCAGATAAACCTGTCGGGGGACAAGTACCTGTGGGATGACTTTGCCTAT GCCACACGGGATGGGCTGGTCGGGGAGGTGGTGTTCGTCGAAGGGCCGGATGGT CGGCATGCCGAGCTGAAGTTCGACTTCCAGTTGCAGCAGGCCCAGGGCGGTGCC GATGAGCAGCGCAGCGGGCGGCCGCGAGCTTTGCAGGAGGCCTGA benE-II-SEQ ID NO: 73 ATGCCCGCGAAGCAGCCAACACCCAAACAAGAGCCCGCCACAAAGGCCGGCCGC TGCCTTGCGAGAAGCTTCATGAAAACGCTGATCAAGGACTGCTCTGTGTCGGCCA Att Dkt: 114198-5960 TCGTCGCCGGCAGCATTGCCACGACGATCTCCTACGCCGGCCCGTTGGTGATCAT TTTCCACGCCGCCGAAGCGGCCGGCCTGTCGCACCAGCAGTTGTCATCCTGGGTC TGGGCCGTGTCCATGGGCAGTGCGTTGCTCGGTGCGCTGCTCAGCCTGCGCTACC GCGTGCCGGTAGTAATCGCCTGGTCCATCCCAGGGTCGGCCCTGTTGGTCACCGC CTTGCCGCAACTGGGCCTGGAGCAGGCCGTGGGCGCCTACCTGGTGGCCAACCT GGTCCTGCTGCTGATCGGCATCAGCGGGGCCTTCGACCGCATCATCGCGCGACTG CCGGGGTCTATCGCTGCCGGTATGCAGGCCGGGATCCTGTTCAGCTTTGGCATCG AGGTGTTCCGCGCCTTGCCGGTGCAACCGCTGCTGGTACTGGCGATGTTCGTCAC CTACGTGCTGATGCGTCGGGTGCAGCCGCGCTATGCAGTGGCTGCGGTACTGGGC GTGGGCATGCTGGTGACCGTGGCCAGCGGTACGTTGCGCAGCGAGGCACTGGTG CTGGAACTGGCCTCACCCCAGTGGATCACCCCCCAGTTCAGCCTGTCGGCGGTAT TCAGCCTGGCGCTGCCGATGGTGCTGGTGGCGCTGACCGGGCAATTCATGCCGGG TATGGCGGTACTGCGCAACGCTGGCTACCCCACGCCGGCCAGCCCGCTGATCAGC GCCAGTGCCCTGATCGGAGTGCTACTGGCGCCGTTCGGTTGCCACGGGCTGAACC TGGCGGCGGTTACCGCCAGCCTGTGCACCGGCCATGAAGCCCATGAAAACCCGC AACGGCGCTATGTAGCGGCGGTGGCTGGCAGTGTGTTGTACCTGCTACTGGGCAT TGCCGGGGCCACCTTGATGTCGCTGTTCGCCGCCTTTCCCGCCGTACTGATCGCTG CCCTGGCCGGCTTGGCCCTGTATGGCGCCATCAGCGAAGCCCTGGTCCGCAGCCT GGCAGAACCGGCCGAGCGCGACGCCGGGCTGTTCACCTTCCTGGTCACCGCTTCG GGGGTGGCTTTCCTGGGCCTGTCGGCAGCGTTCTGGGGGTTGCTGTTCGGCCTGC TGGCGCACGCGCTGCTGAGGCTGCGCCAGCCACGTGCCCGGGAGGTGCTGCGCC GTACCCCATGA catB-SEQ ID NO: 74 ATGACAAGCGTGCTGATTGAACACATAGATGCAATTATCGTCGATCTCCCGACCA TTCGCCCGCACAAGCTGGCGATGCACACCATGCAGCAGCAGACCCTGGTGGTATT GCGACTGCGCTGCAGCGATGGCGTGGAAGGCATCGGTGAAGCCACCACCATCGG TGGCCTGGCGTATGGCTACGAAAGCCCCGAAGGGATCAAGGCCAACATCGACGC GTACCTCGCCCCAGCGTTGATTGGCCTGCCGGCAGACAACATCAATGCCGCCATG CTCAAGCTGGACAAGCTGGCCAAGGGCAACACCTTCGCCAAGTCCGGCATCGAA AGCGCCTTGCTCGACGCCCAGGGCAAACGCCTGGGCCTGCCGGTCAGCGAACTG CTGGGTGGCCGCGTGCGTGACAGCCTGGAAGTGGCCTGGACCCTGGCCAGCGGC GACACCGCCCGCGACATCGCCGAAGCACAGCACATGCTGGACATTCGCCGGCAC CGCGTGTTCAAGCTGAAAATCGGCGCCAACCCGGTGGCGCAGGACCTCAAGCAC GTGGTCGCGATCAAGCGCGAGCTGGGTGACAGCGCCAGCGTGCGGGTCGACGTC AACCAGTACTGGGACGAGTCCCAGGCCATCCGCGCCTGCCAGGTATTGGGCGAC AACGGCATCGACCTGATCGAGCAGCCGATTTCGCGCATCAACCGCGCTGGCCAG GTGCGCCTGAACCAGCGCAGTCCGGCTCCGATCATGGCCGATGAGTCGATCGAA AGCGTCGAGGACGCCTTCAGCCTGGCCGCCGACGGCGCCGCCAGCATCTTCGCCC TGAAAATCGCCAAGAATGGTGGCCCGCGCGCGGTTCTGCGCACTGCACAGATCG CCGAGGCCGCTGGCATCGCCTTGTACGGCGGCACCATGCTCGAAGGTTCGATCGG CACCCTGGCTTCGGCTCATGCATTCCTCACCCTGCGCCAGCTCACCTGGGGTACA GAGCTGTTCGGGCCGCTGCTGCTGACCGAGGAGATCGTCAACGAGCCGCCGCAA Att Dkt: 114198-5960 TACCGCGACTTCCAGCTGCACATCCCCCACACCCCAGGCCTGGGCCTGACGTTGG ACGAACAGCGCCT GGCGCGCTTCGCCCGTCGCTGA catC-SEQ ID NO: 75 ATGTTGTTCCACGTGAAGATGACCGTGAAGCTGCCGGTCGACATGGACCCGGCC AAGGCCGCCCAGCTCAAGGCCGACGAAAAGGAACTGGCCCAGCGCCTGCAGCGC GAAGGCATCTGGCGTCACCTGTGGCGCATTGCCGGGCATTACGCCAACTACAGC GTGTTCGATGTGCCCAGCGTCGAGGCATTGCATGACACGCTGATGCAGCTGCCGC TGTTCCCGTACATGGATATCGAGGTCGACGGCCTGTGTCGGCATCCCTCGTCTAT TCACAGCGACGATCGCTGA catA-I-SEQ ID NO: 76 ATGACCGTGAAAATTTCCCACACTGCCGACATTCAAGCCTTCTTCAACCGGGTAG CTGGCCTGGACCATGCCGAAGGAAACCCGCGCTTCAAGCAGATCATTCTGCGCGT GCTGCAAGACACCGCCCGCCTGATCGAAGACCTGGAGATTACCGAGGACGAGTT CTGGCACGCCGTCGACTACCTCAACCGCCTGGGCGGCCGTAACGAGGCAGGCCT GCTGGCTGCTGGCCTGGGTATCGAGCACTTCCTCGACCTGCTGCAGGATGCCAAG GATGCCGAAGCCGGCCTTGGCGGCGGCACCCCGCGCACCATCGAAGGCCCGTTG TACGTTGCCGGGGCGCCGCTGGCCCAGGGCGAAGCGCGCATGGACGACGGCACT GACCCAGGCGTGGTGATGTTCCTTCAGGGCCAGGTGTTCGATGCCGACGGCAAG CCGTTGGCCGGTGCCACCGTCGACCTGTGGCACGCCAATACCCAGGGCACCTATT CGTACTTCGATTCGACCCAGTCCGAGTTCAACCTGCGTCGGCGTATCATCACCGA TGCCGAGGGCCGCTACCGCGCGCGCTCGATCGTGCCGTCCGGGTATGGCTGCGAC CCGCAGGGCCCAACCCAGGAATGCCTGGACCTGCTCGGCCGCCACGGCCAGCGC CCGGCGCACGTGCACTTCTTCATCTCGGCACCGGGGCACCGCCACCTGACCACGC AGATCAACTTTGCTGGCGACAAGTACCTGTGGGACGACTTTGCCTATGCCACCCG CGACGGGCTGATCGGCGAACTGCGTTTTGTCGAGGATGCGGCGGCGGCGCGCGA CCGCGGTGTGCAAGGCGAGCGCTTTGCCGAGCTGTCATTCGACTTCCGCTTGCAG GGTGCCAAGTCGCCTGACGCCGAGGCGCGAAGCCATCGGCCGCGGGCGTTGCAG GAGGGCTGA TABLES Table 1. Mutations occurred during the malonate ALE. Arrows indicate genetic orientation. Strain names are given as AX.IY. X is the ALE lineage number. Y is an arbitrary identifying number for the clonal isolate from the same ALE lineage. First five mutations are genetic variations observed in multiple isolates of the MG165512,13. Att Dkt: 114198-5960 Att Dkt: 114198-5960

[0002] Atty. Dkt. No.: 114198-5960 Table 2. Primers of the present disclosure. Experiment Primer name Primer sequence (5' to 3') TTTTTATGAAGAAATTATGGAGAAAAATGACAGGGAAAAAGGAG AC TT AA AC GA TT A AG GA CT A AT AA AA AG TT AA TT CT CA GC Atty. Dkt. No.: 114198-5960 selection ilv CAAAAATGCAGCGGACAAAGGATGAACTACGAGGAAGGGAACA landing pad Y_PK_HF ACATTCAAAAAATGCCTGATAGCGCTT il GC C TA GA T AC G GC GC TC A AC CC TT AC C GC GT TA Atty. Dkt. No.: 114198-5960 vanK_F GTGCTGGACATCTGAGCCTGGCCGGCAGGTG vanK_R TTGTTCTCCACGTTCTCAGTGGCTCAGTGCATCAG C CG AC AC TA CG CA G C TT CC G C TT A TT A G C TT AA Atty. Dkt. No.: 114198-5960 BAC_AmpC_iM_2 CAATTCGCGCTAACTCACATTAATTGCGTTGCGCCCACTATTTAT _R ACCATGGGA GC G TT AG Table 3. iModulons and Sub-iModulons of the present disclosure. Sub SEQ ID Gene Gene Product Status Atty. Dkt. No.: 114198-5960 SEQ ID 8 mdcD characterized protein N / ASEQ ID 9 mdcE characterized protein N / A Atty. Dkt. No.: 114198-5960 SEQ ID 30 pcaI characterized protein PCASEQ ID 31 pcaJ characterized protein PCA Atty. Dkt. No.: 114198-5960 SEQ ID 58 PP_3355 uncharacterized protein HCASEQ ID 59 fcs characterized protein HCA

[0003] Atty. Dkt. No.: 114198-5960 References 1. Campos, M. et al. Genomewide phenotypic analysis of growth, cell morphogenesis, and cell cycle events in Escherichia coli. Mol. Syst. Biol.14, e7573 (2018). 2. Griffin, J. E. et al. High-resolution phenotypic profiling defines genes essential for mycobacterial growth and cholesterol catabolism. PLoS Pathog.7, e1002251 (2011). 3. Tang, Y. & Gomer, R. H. An improved shotgun antisense method for mutagenesis and gene identification. Biotechniques 68, 163–165 (2020). 4. Bertea, C. M. et al. Identification of intermediates and enzymes involved in the early steps of artemisinin biosynthesis in Artemisia annua. Planta Med.71, 40–47 (2005). 5. Dietrich, J. A. et al. A novel semi-biosynthetic route for artemisinin production using engineered substrate-promiscuous P450(BM3). ACS Chem. Biol.4, 261–267 (2009). 6. Paddon, C. J. & Keasling, J. D. Semi-synthetic artemisinin: a model for the use of synthetic biology in pharmaceutical development. Nat. Rev. Microbiol.12, 355–367 (2014). 7. Blin, K. et al. antiSMASH 7.0: new and improved predictions for detection, regulation, chemical structures and visualisation. Nucleic Acids Res.51, W46–W50 (2023). 8. Paley, S. M. & Karp, P. D. Evaluation of computational metabolic-pathway predictions for Helicobacter pylori. Bioinformatics 18, 715–724 (2002). 9. Nielsen, A. A. K. et al. Genetic circuit design automation. Science 352, aac7341 (2016). 10. Sastry, A. V. et al. The Escherichia coli transcriptome mostly consists of independently regulated modules. Nat. Commun.10, 5536 (2019). 11. Lamoureux, C. R. et al. A multi-scale expression and regulation knowledge base for Escherichia coli. Nucleic Acids Res.51, 10176–10193 (2023). 12. Rychel, K. et al. iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning. Nucleic Acids Res.49, D112–D120 (2021). 13. Saelens, W., Cannoodt, R. & Saeys, Y. A comprehensive evaluation of module detection methods for gene expression data. Nat. Commun.9, 1090 (2018). Atty. Dkt. No.: 114198-5960 14. Shin, J., Rychel, K. & Palsson, B. O. Systems biology of competency in Vibrio natriegens is revealed by applying novel data analytics to the transcriptome. Cell Rep.42, 112619 (2023). 15. Lim, H. G. et al. Machine-learning from Pseudomonas putida KT2440 transcriptomes reveals its transcriptional regulatory network. Metab. Eng.72, 297–310 (2022). 16. Rajput, A. et al. Advanced transcriptomic analysis reveals the role of efflux pumps and media composition in antibiotic responses of Pseudomonas aeruginosa. Nucleic Acids Res.50, 9675–9688 (2022). 17. Seemann, T. Prokka: rapid prokaryotic genome annotation. Bioinformatics 30, 2068– 2069 (2014). 18. Karp, P. D. et al. The BioCyc collection of microbial genomes and metabolic pathways. Brief. Bioinform.20, 1085–1093 (2019). 19. Molina-Henares, M. A. et al. Functional analysis of aromatic biosynthetic pathways in Pseudomonas putida KT2440. Microb. Biotechnol.2, 91–100 (2009). 20. Assinder, S. J. & Williams, P. A. The TOL plasmids: determinants of the catabolism of toluene and the xylenes. Adv. Microb. Physiol.31, 1–69 (1990). 21. Maderbocus, R. et al. Crystal structure of a Pseudomonas malonate decarboxylase holoenzyme hetero-tetramer. Nat. Commun.8, 160 (2017). 22. Wada, A. et al. Characterization of aromatic acid / proton symporters in Pseudomonas putida KT2440 toward efficient microbial conversion of lignin-related aromatics. Metab. Eng. 64, 167–179 (2021). 23. Brunel, F. & Davison, J. Cloning and sequencing of Pseudomonas genes encoding vanillate demethylase. J. Bacteriol.170, 4924–4930 (1988). 24. Yoshihara, E., Yoneyama, H., Ono, T. & Nakae, T. Identification of the catalytic triad of the protein D2 protease in Pseudomonas aeruginosa. Biochem. Biophys. Res. Commun.247, 142–145 (1998). 25. Nagy-Staron, A. et al. Local genetic context shapes the function of a gene regulatory network. Elife 10, (2021). Atty. Dkt. No.: 114198-5960 26. Boulas, I. et al. Assessing in vivo the impact of gene context on transcription through DNA supercoiling. Nucleic Acids Res.51, 9509–9521 (2023). 27. Vermaas, J. V. et al. Passive membrane transport of lignin-related compounds. Proc. Natl. Acad. Sci. U.S.A.116, 23117–23123 (2019). 28. Pugh, S., McKenna, R., Osman, M., Thompson, B. & Nielsen, D. R. Rational engineering of a novel pathway for producing the aromatic compounds p-hydroxybenzoate, protocatechuate, and catechol in Escherichia coli. Process Biochem.49, 1843–1850 (2014). 29. Naas, T. et al. Beta-lactamase database (BLDB) - structure and function. J. Enzyme Inhib. Med. Chem.32, 917–919 (2017). 30. Bagge, N. et al. Constitutive high expression of chromosomal beta-lactamase in Pseudomonas aeruginosa caused by a new insertion sequence (IS1669) located in ampD. Antimicrob. Agents Chemother.46, 3406–3411 (2002). 31. Ibacache-Quiroga, C., Oliveros, J. C., Couce, A. & Blázquez, J. Parallel Evolution of High-Level Aminoglycoside Resistance in Under Low and High Mutation Supply Rates. Front. Microbiol.9, 427 (2018). 32. Tashiro, Y., Eida, H., Ishii, S., Futamata, H. & Okabe, S. Generation of Small Colony Variants in Biofilms by Escherichia coli Harboring a Conjugative F Plasmid. Microbes Environ. 32, 40–46 (2017). 33. Moya, B. et al. β-Lactam Resistance Response Triggered by Inactivation of a Nonessential Penicillin-Binding Protein. PLoS Pathog.5, e1000353 (2009). 34. Takebayashi, Y., Taylor, E. S., Whitwam, S. & Avison, M. B. CreC Sensor Kinase Activation Enhances Growth of Escherichia coli in the Presence of Cephalosporins and Carbapenems. Antimicrob. Agents Chemother.63, (2019). 35. Zamorano, L. et al. The Pseudomonas aeruginosa CreBC two-component system plays a major role in the response to β-lactams, fitness, biofilm growth, and global regulation. Antimicrob. Agents Chemother.58, 5084–5095 (2014). 36. Huang, H.-H. et al. Expression and Functions of CreD, an Inner Membrane Protein in Stenotrophomonas maltophilia. PLoS One 10, e0145009 (2015). Atty. Dkt. No.: 114198-5960 37. Fukushima, K., Kumar, S. D. & Suzuki, S. YgiW homologous gene from Pseudomonas aeruginosa 25W is responsible for tributyltin resistance. J. Gen. Appl. Microbiol.58, 283–289 (2012). 38. Moreira, C. G. et al. Virulence and stress-related periplasmic protein (VisP) in bacterial / host associations. Proc. Natl. Acad. Sci. U.S.A.110, 1470–1475 (2013). 39. Murtha, A. N. et al. High-level carbapenem tolerance requires antibiotic-induced outer membrane modifications. PLoS Pathog.18, e1010307 (2022). 40. Ali, N. O., Bignon, J., Rapoport, G. & Debarbouille, M. Regulation of the acetoin catabolic pathway is controlled by sigma L in Bacillus subtilis. J. Bacteriol.183, 2497–2504 (2001). 41. Liu, Q. et al.2,3-Butanediol catabolism in Pseudomonas aeruginosa PAO1. Environ. Microbiol.20, 3927–3940 (2018). 42. Liu, Y. et al. Dehydrogenation Mechanism of Three Stereoisomers of Butane-2,3-Diol in Pseudomonas putida KT2440. Front Bioeng Biotechnol 9, 728767 (2021). 43. Priefert, H. et al. Identification and molecular characterization of the Alcaligenes eutrophus H16 aco operon genes involved in acetoin catabolism. J. Bacteriol.173, 4056–4071 (1991). 44. Romero, P. R. & Karp, P. D. Using functional and organizational information to improve genome-wide computational prediction of transcription units on pathway-genome databases. Bioinformatics 20, 709–717 (2004). 45. Xiao, Z. J., Huang, Y. L., Zhu, X. K. & Qiao, S. L. Functional study of AcoX, an unknown protein involved in acetoin catabolism. Adv. Mat. Res.393-395, 776–779 (2011). 46. Schaffitzel, C., Berg, M., Dimroth, P. & Pos, K. M. Identification of an Na+-dependent malonate transporter of Malonomonas rubra and its dependence on two separate genes. J. Bacteriol.180, 2689–2693 (1998). 47. Kumari, S., Tishel, R., Eisenbach, M. & Wolfe, A. J. Cloning, characterization, and functional expression of acs, the gene which encodes acetyl coenzyme A synthetase in Escherichia coli. J. Bacteriol.177, 2878–2886 (1995). Atty. Dkt. No.: 114198-5960 48. LaCroix, R. A. et al. Use of adaptive laboratory evolution to discover key mutations enabling rapid growth of Escherichia coli K-12 MG1655 on glucose minimal medium. Appl. Environ. Microbiol.81, 17–30 (2015). 49. Kogoma, T. & Maldonado, R. R. DNA polymerase I in constitutive stable DNA replication in Escherichia coli. J. Bacteriol.179, 2109–2115 (1997). 50. Yang, Y. L. & Polisky, B. Suppression of ColE1 high-copy-number mutants by mutations in the polA gene of Escherichia coli. J. Bacteriol.175, 428–437 (1993). 51. Datsenko, K. A. & Wanner, B. L. One-step inactivation of chromosomal genes in Escherichia coli K-12 using PCR products. Proc. Natl. Acad. Sci. U.S.A.97, 6640–6645 (2000). 52. Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514–518 (2019). 53. Winsor, G. L. et al. Enhanced annotations and features for comparing thousands of Pseudomonas genomes in the Pseudomonas genome database. Nucleic Acids Res.44, D646–53 (2016). 54. Bauman, K. D. et al. Refactoring the Cryptic Streptophenazine Biosynthetic Gene Cluster Unites Phenazine, Polyketide, and Nonribosomal Peptide Biochemistry. Cell Chem Biol 26, 724– 736.e7 (2019). 55. Gietz, R. D. Yeast transformation by the LiAc / SS carrier DNA / PEG method. Methods Mol. Biol.1205, 1–12 (2014). 56. Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59–64 (2016). 57. Robertson, W. E. et al. Creating custom synthetic genomes in Escherichia coli with REXER and GENESIS. Nat. Protoc.16, 2345–2380 (2021). 58. Ye, J. et al. Primer-BLAST: a tool to design target-specific primers for polymerase chain reaction. BMC Bioinformatics 13, 134 (2012). 59. Phaneuf, P. V., Gosting, D., Palsson, B. O. & Feist, A. M. ALEdb 1.0: a database of mutations from adaptive laboratory evolution experimentation. Nucleic Acids Res.47, D1164– D1171 (2018). Atty. Dkt. No.: 114198-5960 60. Deatherage, D. E. & Barrick, J. E. Identification of mutations in laboratory-evolved microbes from next-generation sequencing data using breseq. Methods Mol. Biol.1151, 165–188 (2014).

Claims

Atty. Dkt. No.: 114198-5960 WHAT IS CLAIMED IS:

1. A method comprising: generating transcriptomic data of one or more organisms under a plurality of conditions, wherein the transcriptomic data comprises expression levels for a set of genes under each of the plurality of conditions; determining one or more iModulons from the transcriptomic data by applying independent component analysis on the transcriptomic data, the one or more iModulons each comprising a plurality of genes comprising at least two genes that are expressed together across the plurality of different conditions and comprising a first at least one gene with a known function and a second at least one gene with an unknown function; for each iModulon of the one or more iModulons: creating a strain of a phenotypic expression from the plurality of genes of the iModulon by refactoring the plurality of genes of the iModulon into a gene transfer vehicle; and sequencing the strain of phenotypic expression and the gene transfer vehicle to determine mutations for the strain of phenotypic expression.

2. The method of claim 1, further comprising: applying an adaptive laboratory evolution algorithm to the strain of the phenotypic expression to create an optimized strain of phenotypic expression for performing a function.

3. The method of claim 1, wherein refactoring the plurality of genes of the iModulon into the gene transfer vehicle comprises refactoring the plurality of genes of the iModulon into a bacterial artificial chromosome.

4. The method of claim 1, further comprising: creating an expression profile for each of the plurality of conditions, each expression profile comprising a vector indicating an expression level of each gene of the plurality of genes in the condition of the expression profile,Atty. Dkt. No.: 114198-5960 wherein determining the one or more iModulons of the organism comprises applying independent component analysis on the one or more expression profiles.

5. The method of claim 4, wherein determining the one or more iModulons from the transcriptomic data by applying independent component analysis on the transcriptomic data comprises: generating, from the expression profiles, a matrix comprising a plurality of columns each corresponding to a different condition of the plurality of conditions and a plurality of rows each corresponding to a different gene of the set of genes; and decomposing the matrix.

6. The method of claim 1, wherein determining the one or more iModulons of the organism comprises: determining a potential iModulon includes only a single gene with a gene expression exceeding a threshold; and discarding the potential iModulon responsive to the determination that the potential iModulon only includes a single gene.

7. The method of claim 1, wherein the plurality of conditions comprise a pH level adjustment, a change in ions, and an added mixture of other cells.

8. The method of claim 1, further comprising: mapping the known function of the first at least one gene onto the iModulon.

9. The method of claim 1, further comprising: generating a record identifying the mutations of the iModulon; and creating a second strain of the phenotypic expression based on the sequenced mutations.

10. An iModulon gene cluster selected from the group of polynucleotides of: a. SEQ ID Nos: 1-4 or an equivalent of each thereof; b. SEQ ID Nos: 5-13 or an equivalent of each thereof;Atty. Dkt. No.: 114198-5960 c. SEQ ID Nos: 14-18 or an equivalent of each thereof; d. SEQ ID Nos: 1925 or an equivalent of each thereof; e. SEQ ID Nos: 26-66 or an equivalent of each thereof; f. SEQ ID Nos: 67-76 or an equivalent of each thereof; g. SEQ ID Nos: 26-43 or an equivalent of each thereof; or h. SEQ ID Nos: 53-54, 57-61, and 64-65 or an equivalent of each thereof.

11. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 1-4 or an equivalent of each thereof.

12. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 5-13 or an equivalent of each thereof.

13. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 14-18 or an equivalent of each thereof.

14. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 19-25 or an equivalent of each thereof.

15. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 26-66 or an equivalent of each thereof.

16. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 67-76 or an equivalent of each thereof.

17. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 26-43 or an equivalent of each thereof.

18. An isolated host cell, comprising an iModulon gene cluster of claim 10, wherein the gene cluster comprises SEQ ID Nos: 53-54, 57-61, and 64-65 or an equivalent of each thereof.

19. The isolated host cell of any one of claims 11-18, wherein the isolated cell is a prokaryotic cell or a eukaryotic cell.Atty. Dkt. No.: 114198-5960 20. The isolated host cell of claim 19, wherein the prokaryotic cell is a bacterial cell, optionally selected from an E. coli, P. putida, Corynebacterium glutamicum, Vibrio natriegens, and genetically minimized species such as JCVI_Syn3A.

21. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 1-4.

22. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 5-13.

23. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 14-18.

24. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 19-25.

25. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 26-66.

26. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 67-76.

27. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 26-43.

28. The isolated host cell of claim 20, wherein the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 53-54, 57-61, and 64-65.

29. A substantially homogeneous population of cells of any one of claims 11 to 28.

30. A composition comprising the isolated host cell of any one of claims 11 to 28, and a carrier, and optionally a preservative or cryoprotective agent.

31. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 1-4.Atty. Dkt. No.: 114198-5960 32. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 5-13.

33. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 14-18.

34. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 19-25.

35. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 26-66.

36. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 67-76.

37. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 26-43.

38. The composition of claim 10 or the isolated host cell of any one of claims 11-28, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity to SEQ ID Nos: 53-54, 57-61, and 64- 65.

39. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster of the composition of claim 10 into a host cell lacking the biochemical pathway.Atty. Dkt. No.: 114198-5960 40. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 into the host cell, wherein the host cell gains the vanillate transport and catabolic pathway of Pseudomonas putida.

41. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 5-13 into the host cell, wherein the host cell gains the malonate transport and catabolic pathway of Pseudomonas aeruginosa.

42. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 14-18 into the host cell, wherein the host cell gains the 2,3-butanediol catabolic pathway of Pseudomonas putida.

43. A method for preparing an isolated cell or population of cells with an exogenous biochemical property, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 19-25, wherein the host cell gains the ampicillin resistance property of Pseudomonas aeruginosa.

44. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43 into the host cell, wherein the host cell gains the protocatechuate transport and catabolic pathway of Pseudomonas putida.

45. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 and 26-66 into the host cell, wherein the host cell gains the hydroxycinnamates transport and catabolic pathway of Pseudomonas putida.

46. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides ofAtty. Dkt. No.: 114198-5960 SEQ ID Nos: 26-43 and 67-76 into the host cell, wherein the host cell gains the benzoate transport and catabolic pathway of Pseudomonas putida.

47. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-4 and 26-43 into the host cell, wherein the host cell gains the vanillate transport and catabolic pathway of Pseudomonas putida.

48. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 26-43, 53-54, 57-61, and 64-65 into the host cell, wherein the host cell gains the hydroxycinnamates transport and catabolic pathway of Pseudomonas putida.

49. The method of any one of claims 39 to 48, further comprising optimizing the biochemical pathway through adaptive laboratory evolution.

50. The method of any one of claims 39 to 49, further comprising culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture.

Citation Information

Patent Citations

  • Methods and systems for genomic analysis

    US20160019341A1

  • Performance enhancing genetic variants of e. coli

    US20170198276A1

  • Bacterial artificial chromosomes

    US20190111125A1