Systems and methods for formulating growth media

WO2025171112A3PCT designated stage Publication Date: 2025-10-09RGT UNIV OF CALIFORNIA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/014764
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2025-02-06
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Media formulation for microbial growth is traditionally empirical, lacking a detailed genetic and molecular basis, which hinders the understanding of bacterial nutrition and impacts the development of effective growth media.

Method used

Utilizing transcriptome data and advanced analytics to identify iModulon gene clusters that influence cellular stress response and metabolism, enabling the formulation of media based on foundational biological principles, and iteratively modifying media formulations to optimize growth conditions.

Benefits of technology

This approach significantly reduces time and costs in developing growth-optimized media, enhances compatibility with specific cellular processes, and facilitates customized media formulations for improved microbial growth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025014764_09102025_PF_FP_ABST
    Figure US2025014764_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for optimizing a media formulation or substrate for cellular growth, comprising growing a first plurality of cells in a media formulation, on a substrate, or on a substrate in a media formulation; isolating RNA from the plurality of cells; detecting based on the upregulation or downregulation of an iModulon gene cluster comprising a metabolic, trace element, or stress- related iModulon; and iteratively modifying the media formulation or the substrate to downregulate or upregulate the iModulon gene cluster, growing a second population of cells in the modified media formulation or on the modified substrate, isolating RNA from the second plurality of cells, and detecting optimized expression of the iModulon gene cluster.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYSTEMS AND METHODS FOR FORMULATING GROWTH MEDIA BACKGROUND OF THE DISCLOSURE This application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Provisional Application No.: 63 / 550,406, filed February 6, 2024, and U.S. Provisional Patent Application No.: 63 / 555,817, filed February 20, 2024, the content of each of which is incorporated herein by reference in its entirety. BACKGROUND OF THE DISCLOSURE Media formulation is fundamental to all of microbiology. For a microorganism to grow, all of its nutritional requirements need to be met. Media formulation has been achieved through empirical means and once a media is found to be supportive of the growth of a particular microbe, it gets reused by a large number of laboratories and gains a legacy status. The nutrient environment is known to influence many cellular responses1,2. They influence the state of the transcriptional regulatory network (TRN), thus influencing metabolic shifts3-5, stringent response activation6, pathogenicity changes7, antibiotic resistance8, and even interspecies interactions9. Despite its fundamental importance, media formulation is still empirical and lacks a detailed genetic and molecular basis. However, the advent of genome-scale science, with available full genome wide profiling methods allow the determination of the full molecular state underlying an observed phenotypic state. In particular, since 2013, after purification mRNA methods were established for bacteria10,11, a large number of transcriptomic profiles have become available in the public domain12,13. SUMMARY OF THE DISCLOSURE Formulating growth media, a fundamental challenge in microbiology, has traditionally depended on empirical methods. This disclosure leverages extensive transcriptome data and advanced analytics to transform this process. This approach enables a deep mechanistic look into how bacteria fundamentally respond to various nutrients, moving beyond relying on simple growth rate measurements. Both stress-free and stress-inducing effects of nutrients on bacteria are deciphered, including low nutrition levels, metal ion exposure, and stressful conditions. This strategy not only broadens the understanding of bacterial nutrition but also assists in developing more effective growth media. The findings hold the potential for expanding the comprehension of the foundations of bacterial nutrition, with wide-ranging applications from basic science to industrial processes and environmental studies. Understanding the diverse nutritional needs of bacteria is foundational in microbial research and biotechnology. However, defining these requirements at a fundamental level for specific strains remains challenging, directly impacting media formulation strategies. In one aspect as described herein, knowledge-enriched data analytics are applied to a compendia of transcriptomics data to define 64 sets of genes that are independently modulated in Vibrio natriegens. This strategy resulted in; i) identification of novel transporter systems for diverse substrates, ii) detailing the influence of trace elements on metabolism and growth, and iii) a broad assessment of nutrient-induced stress responses, including osmotic stress, low glycolytic flux, proteostasis, and altered protein expression. The linking of acetate-associated regulon to low-glycolytic flux led Applicant to modify this strategy to disclose a technical solution for improved growth under nutrient-limited conditions. By going beyond traditional growth rate metrics, this disclosure not only advances the understanding of bacterial nutrition but also facilitates the formulation of media based on foundational biological principles and genome-wide cellular responses. The methods disclosed herein yield: 1) significant savings in time and costs for developing growth-optimized media formulations comparted to traditional trial and error approaches, 2) enablement of the optimization of specific cellular processes through customized media formulations, and 3) genetically modified cells with improved compatibility with existing media formulations. In one aspect, Applicant has identified 64 iModulon gene clusters that affect cellular stress response and growth. In another aspect, Applicant has developed a method for determining the RNA expression profile of an iModulon gene cluster from a media formulation or a substrate for cellular growth, the method comprising, or consisting essentially of, or yet further consisting of determining one or more iModulons of a first organism in a cell culture medium by applying independent component analysis (ICA) on gene expression data of the cultured first organism, the one or more iModulons comprising one or more genes of an unidentified function; inserting the one or more iModulons into a second organism to create a strain of a phenotypic expression from the one or more iModulons; determining the RNA expression profile of the iModulon gene cluster in the second organism. In some embodiments, the method for determining the RNA expression profile of an iModulon gene cluster from a media formulation or a substrate for cellular growth further comprises, or further consists essentially of, or yet further consists of iteratively modifying the cell culture media to downregulate or upregulate the iModulon gene cluster, growing a third population of cells in the modified media formulation or on the modified substrate, determining the RNA expression from the third cell culture, and detecting the RNA expression profile of the iModulon gene cluster. In some embodiments, the iModulon comprises, or consists essentially of, or yet further consists of an iModulon selected from a metabolic iModulon, a trace element iModulon, a stress-related iModulon, or an Acetate iModulon gene cluster. In some embodiments, the method for determining the RNA expression profile of an iModulon gene cluster from a media formulation or a substrate for cellular growth further comprises, or further consists essentially of, or yet further consists of preparing an isolated cell or population of cells with optimized media or substrate compatibility, comprising genetically modifying the isolated cell or population of cells to upregulate or downregulate the iModulon gene cluster identified in claim 1, optionally wherein the iModulon gene cluster is the Acetate iModulon gene cluster. In other embodiments, the isolated cell is the wild-type phenotype of the cell. In yet another aspect, Applicant has developed a method for optimizing media formulations or substrates, the method comprising, or consisting essentially of, or yet further consisting of growing a first population of cells in a media formulation, on a substrate, or on a substrate in a media formulation; isolating RNA from the plurality of cells; detecting by the RNA expression profile upregulation or downregulation of an iModulon gene cluster comprising a metabolic, trace element, or stress-related iModulon; and iteratively modifying the media formulation or the substrate to downregulate or upregulate the iModulon gene cluster, growing a second population of cells in the modified media formulation or on the modified substrate, isolating RNA from the second plurality of cells, and detecting optimized expression of the iModulon gene cluster. Further provided herein, Applicant has developed a method for preparing an isolated cell or population of cells with optimized media or substrate compatibility, the method comprising, or consisting essentially of, or yet further consisting of genetically modifying the isolated cell or population of cells to upregulate or downregulate the iModulon gene cluster, optionally wherein the iModulon gene cluster is the Acetate iModulon gene cluster. In some embodiments, the stress-related iModulon comprises, or consists essentially of, or yet further consists of the OxyR, RpoE, Low-osmolarity, Macrolide-efflux, PhrR, CSP, or Ectoine iModulon gene cluster. In other embodiments, the trace element iModulon comprises, or consists essentially of, or yet further consists of the Zur, TPP, CueR, Fur-1, Fur-2, Enterobactin, or thiosulfate iModulon gene cluster. In other embodiments, the metabolic iModulon comprises, or consists essentially of, or yet further consists of the NarP-1, TonB, FNR, CytochromeO, NarP-2, GABA, Alr, HutC, MetJ, LiuR, ArgR, Cysteine, Prophage-1, Prophage-2, β-ox, Flagellar-1, Flagellar-2, Chaperone, Biofilm, Ribosomal proteins, TfoX, Qst, UC-1, HapR, Cellobiose, TRAP, Formate, TreR, Arabinose, AGL, NagC, GntR, Maltose, FruR, Glycerol, TTT, Acetate, GalR, Rhamnose, Null-1, Null-2, Null-3, UC-2, UC-3, UC-4, UC-5, UC-6, UC-7, UC-8, or UC-9 iModulon gene cluster. In some embodiments, the media formulation comprises, or consists essentially of, or yet further consists of acetate, ethanol, cellobiose, galactose, glycogen, or GlcN. In yet a further aspect, a method for optimizing media composition for microbial growth includes obtaining transcriptomic data from a plurality of microbial samples grown in media containing different conditions, the different conditions each corresponding to a type and concentration of a nutrient supplement added to the media; determining one or more iModulons by applying independent component analysis to the transcriptomic data, each iModulon comprising at least one gene with a known function and at least one gene with an unknown function; calculating a nutrient score for each condition based on a weighted sum of iModulon activity levels of the one or more iModulons, wherein the weights are determined based on iModulon activity levels and iModulon growth rate in the media containing the nutrient supplements of the conditions; selecting a condition identifying a type and concentration of the nutrient supplement based at least on the condition corresponding to a positive nutrient score; and supplementing a growth medium with the type and concentration of the nutrient supplement of the condition. In some embodiment, the method includes or consists essentially of, or yet further consists of categorizing the iModulons into functional groups comprising catabolism, stress responses, and translation. In some embodiment, calculating the nutrient score for each condition includes or consists essentially of, or yet further consists of identifying each iModulon categorized into the stress responses group; and calculating the nutrient score based only on iModulon activity by the identified iModulons categorized into the stress responses group. In some embodiment, calculating the nutrient score includes or consists essentially of, or yet further consists of assigning positive weights to growth-promoting iModulons; and assigning negative weights to stress-related iModulons. In some embodiment, the nutrient supplement is selected from the group consisting of L-methionine, L-cysteine, and D-malic acid. In some embodiment, the method further includes or consists essentially of, or yet further consists of ranking each of the conditions based on the nutrient scores calculated for the conditions, wherein selecting the condition is based at least on the rankings of the conditions. In some embodiments, the method includes or consists essentially of, or yet further consists of generating a matrix based on expression profiles for individual conditions in the transcriptomic data; applying independent component analysis to the matrix to decompose the matrix into at least a mixing matrix containing a different column for each iModulon and a different row for each condition, wherein values in a column for an iModulon indicate an activity level of the iModulon in the conditions; determining the weight for each nutrient supplement based at least on the values in the column for the iModulon indicating the activity level of the iModulon in the conditions. In some embodiments, generating the matrix includes or consists essentially of, or yet further consists of, for each condition identifying an expression level of each gene grown under the condition; and generating a vector with each identified expression level for the condition. Further provided herein is an iModulon gene cluster comprising SEQ ID Nos: 1-6 and optionally, SEQ ID Nos: 7-8, or an equivalent of each thereof. In some embodiments, an isolated host cell comprises the iModulon gene cluster. In some embodiments, the isolated cell is a prokaryotic cell or a eukaryotic cell. In some embodiments, the prokaryotic cell is a bacterial cell, optionally selected from an E. coli, P. putida, Corynebacterium glutamicum, Vibrio natriegens and genetically minimized species such as JCVI_Syn3A. In some embodiments, the bacterial cell is an E. coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 1-6 and optionally, SEQ ID Nos: 7-8. In some embodiments, the disclosure provides a substantially homogenous population of cells comprising the iModulon. In some embodiments, a composition comprises the isolated host cell, a population of cells, and a carrier, and optionally a preservative or cryoprotective agent. In some embodiments, the composition or the isolated host cell comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID Nos: 1-6 and optionally, SEQ ID Nos: 7-8. In another aspect, Applicant has developed a method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster of the composition into a host cell lacking the biochemical pathway. In yet another aspect, Applicant has developed a method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 1-6 and optionally, SEQ ID Nos: 7-8into the host cell, wherein the host cell gains the acetate metabolic pathway of Pseudomonas putida. In some embodiments, the methods further comprise optimizing the biochemical pathway through adaptive laboratory evolution. In some embodiments, the methods further comprise culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. BRIEF DESCRIPTION OF THE DRAWINGS FIGS. 1A – 1B: Determining bacterial nutrition through deep transcriptome analysis. (FIG. 1A) Depicts the conceptual framework of the methods of this disclosure. A matrix of detailed stress and nutrient transcriptomic responses can be evaluated for a large number of nutrient perturbations. The individual responses are characterized through the condition-dependent activity level of independently modulated sets of genes, called iModulons. (FIG. 1B) A comprehensive map of the 46 iModulons identified using the natPRECISE148 compendium is shown. Each iModulon number is represented with gene names enclosed in a gray box, indicating their association with a transcriptional regulator or cellular function. The shades of gray represent the genes contained in each iModulon and are darker when the genes in each iModulon overlap (e.g., Enterobactin iM). For detailed information on each iModulon number, refer to Table 4 and the iModulonDB website: https: / / www.imodulondb.org / 16. iModulons exclude for clarity include three categorized as Null, nine as Uncharacterized, two as Prophage, as well as the specific iModulons Biofilm, FNR, HapR, and PhrR. FIGS. 2A – 2E: Transcriptomic responses of V. natriegens to commonly used substrates. (FIG. 2A) The activity of 15 identified transcriptomic responses (rows) to various carbon sources (columns) is shown. The conservation of these iModulons was determined by comparing them with iModulons from other species (FIG. 8D). Arrows indicate the anticipated pairing of a substrate with its primary iModulon counterpart. TTT and TRAP iModulons were previously unknown. (FIG. 2B) Relationship between the activity levels of the AGL and GalR iModulons, representing galactose and glycogen catabolism, respectively, is shown. The Pearson's correlation coefficient (R) and the corresponding P-values are shown. (FIG. 2C) iModulon gene composition, described through gene weightings for the TTT, TRAP, and GntT iModulons and their genome position is shown. The horizontal dashed lines indicate a cut-off threshold for gene weights. OM, outer membrane; IM, inner membrane. (FIG. 2D) A comparison between the TTT and TRAP iModulons within other TRAP genes in GntR iModulon is shown. Connections between homologous genes are gray, with percentage sequence identity noted on the connecting lines. (FIG. 2E) An evaluation of the function of TTT iModulon, TRAP iModulon, and unknown transporters is shown. To determine their role in growth, all genes within the TTT and TRAP iModulons, as well as selected gene members of the GntR, GalR, and Rhamnose iModulons, were deleted as detailed in FIG. 9D. Each growth condition utilized a 15 mM concentration of the respective carbon source in M9Na minimal medium. The specific growth rates of these strains under various carbon conditions were measured, and the functions of deleted genes were inferred from the average relative growth rate (Relative GR) compared to the wild-type (WT) strain. FIGS. 3A – 3D: Trace-element-related iModulon activity changes and cell growth variations across different substrates. (FIG. 3A) Activity of seven element homeostasis iModulons (rows) under various nutrient conditions (columns) is shown. (FIG. 3B) Activity correlation among the seven iModulons is shown. (FIG. 3C) Growth profiles in various carbon conditions with ZnCl2 treatment is shown. 8 carbon sources known to induce the Zur iModulon was used. A 15 mM concentration of the respective carbon source in M9Na minimal media was used. Data is presented as mean ± SD from four biologically independent samples. (FIG. 3D) Growth profiles in various carbon conditions with ZnCl2 treatment is shown. 6 carbon sources that do not induce Zur iModulon was used. A 15 mM concentration of the respective carbon source in M9Na minimal media was used. Data is presented as mean ± SD from four biologically independent samples. FIGS. 4A – 4G: Nutrients affect activity levels of Ribosome and nine stress- related iModulons, and impact proteostasis. (FIG. 4A) The activity of nine stress-related and Ribosome iModulons (rows) under various substrates (columns) is shown. iModulon activities > 10 in rich and minimal media are summarized in column and row plots associated with the heat map. Nutrients are classified as stress-free (SF) or stress-inducing (SI) based on the stress-iModulon activity that they induce. (FIG. 4B) The specific growth rate comparison between SF and SI substrates is shown. Different shades represent data from each substrate. WT V. natriegens were grown in 96-well plates, with OD600 measured using a Tecan Infinite 200Pro microplate reader. Each condition involved a 15 mM concentration of the respective carbon source in M9Na minimal media. (FIG. 4C) The activity correlation among the ten iModulons used for nutrient assessment is shown. (FIG. 4D) The correlation between Chaperone and Ribosome iModulon activities is shown. (FIG. 4E) Fluorescence and growth profiles in various nutrient conditions is shown. V. natriegens with Ptac-GFP plasmid were cultured in black / clear 96-well plates. GFP fluorescence and OD600were measured. Cellular GFP signals (GFP / OD600) were normalized to peak values. Upper and lower dashed lines indicate the maximum values for ribose and glucose samples, respectively. Data are mean ± SD from four biologically independent samples. (FIG. 4F) A comparison of relative growth rate, relative GFP rates, maximum OD600, and maximum cellular GFP signals (GFP / OD600) in various carbon conditions is shown. Each nutrient condition was compared to the glucose sample for relative growth and relative GFP rates. A dashed line represents the level of the glucose 1 mM IPTG sample. Data are presented as mean ± SD from four biological replicates. The statistical significance for each carbon condition, in comparison to the glucose sample with the same IPTG concentration, was determined using Student’s t-test (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001), except for the relative growth rate graph. (FIG. 4G) Fluorescence and growth profile in LBv2 and BHIN Media is shown. Vn carrying Ptac-GFP plasmid grown in black / clear 96-well plates. GFP signal and OD600measured. Cellular GFP signals (GFP / OD600) were normalized to peak values. Mean ± SD from three independent samples. FIGS. 5A – 5F: Functions of the Acetate iModulon under low-glycolytic flux conditions. (FIG. 5A) A graphical summary depicting the function of the Acetate iModulon in central carbon metabolism at low glycolytic flux states is shown. The left dashed lines indicate acetate-induced inhibition of glycolysis and TCA cycle enzymes44. (FIGS. 5B) An analysis of acetate supplementation in stress-inducing (SI) and stress-free (SF) carbon conditions is shown. Specific growth rates were measured following the modification in SI and SF conditions, compared to control conditions (no acetate or WT strain). Statistical significance in SI and SF environments was evaluated using the two-tailed Mann-Whitney test (***P < 0.001; ****P < 0.0001). Shading distinguishes specific carbon conditions, and detailed growth rate changes for each carbon source are provided in (FIGS. 11C – 11D). (FIG. 5C), An analysis of Acetate iModulon gene deletion effects in stress-inducing (SI) and stress-free (SF) carbon conditions is shown. Specific growth rates were measured following this modification in SI and SF conditions, compared to control conditions (no acetate or WT strain). Statistical significance in SI and SF environments was evaluated using the two-tailed Mann-Whitney test (***P < 0.001; ****P < 0.0001). Shading distinguishes specific carbon conditions, and detailed growth rate changes for each carbon source are provided in (FIGS. 11C – 11D). (FIG. 5D) Correlation between Acetate and GalR iModulons’ activities is shown. Pearson's R (R) and P-value (P) are displayed. (FIG. 5E) TTT and TRAP iModulons' activity under SI and SF carbon conditions is shown. Statistical significance was assessed using the two-tailed Wilcoxon matched-pairs signed rank test (***P < 0.001). (FIG. 5F) Interrelations among Acetate, TRAP, TTT, and GntR iModulons is shown. Pearson's R (R) and P-value (P) are shown. FIGS. 6A – 6E: Strain engineering with the Acetate iModulon. (FIG. 6A) A graphical overview of enhancing Acetate iModulon function with native regulation is shown. AceiMBoost is activated in low glycolytic flux states to enhance the phenotype. (FIG. 6B) A schematic of the Acetate iModulon boosting circuit (top) and its genomic integration (bottom) is shown. AcetateiM-Boost modules (Boost-A and Boost-B) are integrated following the native acetate iModulon gene (RS14410). For simplicity, all constructs except araC and araC’s RBS are depicted on the same strand as araC, though actual integration is in the opposite direction. Boost-A contains two operons of the Acetate iModulon, while Boost-B includes the acsA gene with a terminator. (FIG. 6C) Growth profiles of control and Boost strains in glucose (top) and cellobiose (bottom) conditions is shown. Control strains express ampicillin and spectinomycin resistance genes under the same promoter as the two Boost strains, integrated into the arabinose catabolism and endA genes, respectively. (FIG. 6D) Changes in growth rate and lag-time due to the Acetate iModulon boosting module is shown. Two Boost strains and a control strain were cultured in 15 mM carbon source in M9Na minimal media, with OD600measured using a Tecan Infinite 200Pro microplate reader. Relative growth rates compared to the control is shown. Data are mean ± SD from three to four independent samples. Growth rate statistical significance was determined using Student’s t-test (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001). Lag-time significance was assessed with the two-tailed Mann-Whitney test (*P < 0.05). (FIG. 6E) Changes in growth rate and lag-time due to the Acetate iModulon boosting module is shown. Two Boost strains and a control strain were cultured in 15 mM carbon source in M9Na minimal media, with OD600 measured using a Tecan Infinite 200Pro microplate reader. Lag times are compared to the control is shown. Data are mean ± SD from three to four independent samples. Growth rate statistical significance was determined using Student’s t-test (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001). Lag-time significance was assessed with the two-tailed Mann-Whitney test (*P < 0.05). FIGS. 7A – 7H: Development and analysis of the natPRECISE148 database in Vibrio natriegens. (FIG. 7A) A comparison of the number of single nutrient perturbation samples in both complex and defined media against existing iModulon data in other species92is shown. (FIG. 7B) The database configuration comparison between the previous natPRECISE10494and the updated natPRECISE148 is shown. Note that the formate condition newly incorporated in the natPRECISE148 dataset diverges from the minimal M9Na media standard, as detailed elsewhere96. (FIG. 7C) The process for generating natPRECISE148, highlighting the quality control steps is shown. (FIG. 7D) The range of correlations (represented by darker bars) between different samples indicating the presence of diverse expression states is shown. Lighter bars show the correlation between replicates, with the median Pearson’s R value for replicates being 0.98. (FIG. 7E) The classification of independent components across various dimensions is shown. (FIG. 7F) A cumulative variance plot demonstrating the variance explained versus the number of independent components (i.e., iModulons) or principal components in both natPRECISE104 and natPRECISE148 is shown. (FIG. 7G) A scatter plot of iModulon and regulon recall is shown. (FIG. 7H) A treemap of the 64 iModulons in natPRECISE148 is shown. Each rectangle's size corresponds to the proportion of variance explained by that iModulon. FIGS. 8A – 8E: iModulon analysis and cross-species comparison in natPRECISE148. (FIG. 8A) Breakdown of natPRECISE104 and natPRECISE148 iModulons by their annotation category: ‘Regulatory’ for significant enrichment of one or more known regulators; ‘Functional’ for significant enrichment of Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment or functional annotation (e.g., ribosomal proteins, flagellar) without significant regulator enrichment, or indicating potential new regulons; and ‘Uncharacterized’ for unclassified iModulons. Eleven detailed iModulon functions were further characterized for functional and regulatory iModulons. (FIG. 8B) A comparison of ICA results of natPRECISE104 and natPRECISE148 is shown. Each iModulon is shade-coded by function. iModulons with correlated gene coefficients (Pearson’s R > 0.25) are linked by arrows, with widths and shades of arrows indicating the strength of correlation. (FIG. 8C) A carbon utilization iModulon matrix displaying the presence and absence of identified carbon-utilization iModulons for each specific carbon source catabolism is shown. The matrix compares the existing iModulon data across different species, highlighting the cross-species similarities and differences. (FIG. 8D) An iModulon cross-species comparison in natPRECISE148 is shown. Heatmap presents a comparison of all iModulons in natPRECISE148 with iModulon data from other species. A heatmap indicates the iModulons whose gene coefficients are correlated across different species, showcasing the similarities and differences in iModulon profiles. (FIG. 8E) A heatmap showing the conservation of iModulon function categories across various species, as represented in their respective iModulon data, is shown. The heatmap highlights the degree of conservation for each category. The left panel shows the number of unique iModulons in each function category specific to natPRECISE148, providing insights into the distinct functional aspects of these iModulons. FIGS. 9A – 9H: Carbon utilization and trace element-related iModulons and their roles. (FIG. 9A) An analysis of activity correlation among 15 carbon utilization iModulons is shown. (FIG. 9B) The relationship between the AGL iModulon activity and the expression of seven putative transcription factors found in GalR iModulon is shown. High correlation suggests that a transcription factor may regulate AGL iModulon. Dashed lines represent the best fit for each graph. Pearson's R (R) and corresponding P-value (P) are shown for each graph. (FIG. 9C) Genotyping by PCR is shown. Deletion of the target gene was confirmed by PCR using primers specific for the knockout strain (“Δstrain-name"_Val_For / Rev; see Table 3). Filled and unfilled triangles indicate the expected amplicon sizes for wild-type (WT) and mutant strains, respectively. (FIG. 9D) Specific growth rates and maximum OD600 (Max OD600) in various carbon conditions treated with ZnCl2is shown. 8 carbon sources known to induce Zur iModulon were used. The specific growth rates under different conditions were measured, and the effect of ZnCl2 treatment was inferred from the relative growth rate (Relative GR) compared to the untreated sample for each carbon source. Data are presented as means ± SEM from four biological replicates. Statistical significance was assessed using Student’s t- test (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001). (FIG. 9E) An illustration depicting Zn2+ interactions with enzymes based on iModulon and growth profile data is shown. This figure highlights enzymes known to be dependent on Zn2+ (nagA), those inhibited by Zn2+ (AgL, fruk), and enzymes potentially activated by Zn2+ ions (CAA, gpsA). (FIG. 9F) Specific growth rates and maximum OD600 (Max OD600) in various carbon conditions treated with ZnCl2 is shown. 6 carbon sources that do not induce Zur iModulon were used. The specific growth rates under different conditions were measured, and the effect of ZnCl2treatment was inferred from the relative growth rate (Relative GR) compared to the untreated sample for each carbon source. Data are presented as means ± SEM from four biological replicates. Statistical significance was assessed using Student’s t- test (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001). (FIG. 9G) iModulon gene weights for the Zur iModulon is shown. Genes without a specific name are either uncharacterized or have putative functions, as indicated by their Clusters of Orthologous Groups (COGs). (FIG. 9H) A comparison between proteins associated with zinc-related iModulons across species using the BLASTp algorithm is shown. Genes with an E-value lower than e−10are shaded according to the key. FIGS. 10A – 10J: Functional roles of stress-related iModulons. (FIG. 10A) iModulon gene weights, and activity across different salt concentrations is shown. The top panels show gene weights, and bottom panels depict iModulon activity for the Low-osmolarity iModulon. An orange dashed line marks the NaCl concentration calculated reversely from Ectoine activity in acetate samples, determined through curve- fitting analysis. (FIG. 10B) iModulon gene weights, and activity across different salt concentrations is shown. The top panels show gene weights, and bottom panels depict iModulon activity for the Ectoine iModulon. An orange dashed line marks the NaCl concentration calculated reversely from Ectoine activity in acetate samples, determined through curve-fitting analysis. (FIG. 10C) a Validation of Low-osmolarity iModulon (Low-Os. iM) and Ectoine iModulon functionality is shown. Wild-type (WT) and knockout (KO) strains (FIG. 9C) were cultured under M9woNa (M9 medium without NaCl) containing 0.5% wt / vol glucose. The statistical significance was evaluated using Student’s t-test, with significance levels marked as *P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001. (FIG. 10D) iModulon gene weights for the RpoE iModulon (top) and the relationship between the RpoE iModulon activity and the expression of rpoE transcription factor (bottom) is shown. (FIG. 10E) Gene weights and activity for Ribosome iModulons is shown. The left panels display gene weights, while the right panels illustrate iModulon activity across all tested conditions. Gene functions are additionally categorized and highlighted using Clusters of Orthologous Groups (COGs), each represented by a different shade. (FIG. 10F) Gene weights and activity for Chaperone iModulons is shown. The left panels display gene weights, while the right panels illustrate iModulon activity across all tested conditions. Gene functions are additionally categorized and highlighted using Clusters of Orthologous Groups (COGs), each represented by a different shade. (FIG. 10G) iModulon activity of RpoE, Ribosome, and Chaperone iModulons across different salt concentrations is shown. (FIG. 10H) A schematic illustration of rhamnose carbon uptake inducing low salt stress, based on TTT and TRAP iModulons is shown. (FIG. 10I) Gene weights and activity for Prophage-1 (top) and Prophage-2 (bottom) iModulons is shown. The left panels display gene weights, while the right panels show iModulon activity under tested conditions. (FIG. 10J) A comparison of growth profile and growth rate in LBv2 and BHIN media is shown. WT V. natriegens cells were grown in 96-well plates with shaking, and OD600 was measured using a Tecan Infinite 200 Pro microplate reader. Data presented mean ± SD from three independent samples. The statistical significance was assessed using Student’s t-test FIGS. 11A – 11G: Stress-related iModulons functional roles across carbon conditions. (FIG. 11A) iModulon genes weights of Acetate iModulon is shown. (FIG. 11B) iModulon activity Acetate iModulon across with cell growth phage in different samples in LBv2 media is shown. (FIG. 11C) An evaluation of the impact of acetate supplementation on the utilization of various carbon sources is shown. Statistical significance was determined using Student’s t-test (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001). (FIG. 11D) An evaluation of the impact of Acetate iModulon genes on the utilization of various carbon sources is shown. Statistical significance was determined using Student’s t-test (*P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001). (FIG. 11E) Exploring the activity relationship between the Acetate and BkdR iModulons in the putidaPRECISE321 database91is shown. (FIG. 11F) A schematic diagram illustrating the utilization of stress-induced (SI) or poor carbon sources by cells is shown. The diagram depicts the activation of TTT, TRAP, and Acetate iModulons as a response mechanism to support cellular growth under these conditions is shown. (FIG. 11G) Genotypic verification of the Boost strains using PCR is shown. Integration of the Acetate iM-Boost modules (A and B) was confirmed by PCR, targeting the left, internal, and right junctions with strain-specific primers listed in Tables 1 – 2. The triangles indicate the expected sizes of PCR products for the Boost strains. The sequence of each PCR product was further confirmed by Sanger sequencing. Lanes #1 to #4 represent different bacterial clones, with clones #4 of Boost-A and Boost-B selected for further experiments. FIG. 12 is an illustration of an example system for detecting iModulons and selecting nutrient supplements, in accordance with implementations. FIG. 13 is an illustration of an example method for nutrient selection and media supplementation, in accordance with implementations. FIG. 14 is a block diagram illustrating an architecture for a computer system that can be employed to implement elements of the systems and methods described and illustrated herein, including, for example, the system depicted in FIG. 12 and the method depicted in FIG. 13. FIGS. 15A - C: An example sequence for decoding nutrient responses using a knowledge-enriched transcriptomic approach. (FIG. 15A) An overview of the nutrient perturbation dataset constructions is shown. A total of 346 newly generated samples were integrated with 535 previously generated PRECISE-1K database samples to create the PRECISE-NP881 database. These samples include diverse nutrient types, categorized as carbon sources (e.g., sugars, sugar acids, sugar alcohols), nitrogen sources (e.g., amino acids, nucleobases), and supplements (e.g., peptides). (FIG. 15B) Identification and categorization of iModulons (iMs) is shown. iModulons were grouped into functional categories, including catabolism, stress, translation, and metal ion regulation, based on gene memberships. (FIG. 15C) A visualization of responses of nutrient supplementations is shown. A t- SNE plot shows the clustering of samples based on transcriptomic profiles. Samples represent nutrient supplementation conditions, other samples of PRECISE-1K, and a star marker for the glucose (carbon) + NH4Cl (nitrogen) baseline condition. FIG. 16 is an illustration of an overview of a nutrient perturbation dataset compared with other transcriptomic datasets. The overview shows a comparison of unique conditions among datasets. PRECISE-NP881 includes 117 unique nutrient supplementation conditions introduced in this study, surpassing other datasets, which typically include fewer unique conditions. FIGS. 17A and 17B: Two plots illustrating transcriptional trade-offs and interactions between RpoS, translation, and ppGpp iModulons in response to nutrient perturbations. (FIG. 17A) A scatter plot showing the fear-greed relationship between RpoS iModulon (iM) and translation iModulon activities is shown. Each point represents a transcriptomic sample, with yellow points highlighting nutrient supplementation (NS) conditions and gray points representing all conditions. Specific amino acid additions (+Thr, +Gth, +Cys) associated with high stress responses are annotated. (FIG. 17B) A three-dimensional plot depicting the interplay among RpoS iModulon, translation iModulon, and ppGpp iModulon activities is shown. Samples are shade-coded by nutrient perturbation types: carbon replacements (NR-C), nitrogen replacements (NR-N), and nutrient supplementation (NS). The plot reveals distinct clustering patterns and interactions between stress (RpoS and ppGpp) and translation responses under various nutrient conditions. FIGS. 18A - C: An example sequence for generating nutrient scores and relationships with growth rates a stress responses. (FIG. 18A) A calculation of the nutrient score is shown. The nutrient score is computed as a weighted sum of stress-related iModulon (iM) activities (Ai), where n represents the total number of stress-related iModulons (n = 30), and Wiis the weight assigned to each iModulon. (FIG. 18B) The weight (Wi) corresponds to the correlation coefficient between the iModulon activity and growth rate across all nutrient conditions. Positive weights indicate a positive correlation with growth, while negative weights indicate a negative correlation. (FIG. 18C) A relationship between the nutrient score and growth rate is shown. A significant positive correlation (R = 0.73) is observed, suggesting that higher nutrient scores correspond to higher growth rates. FIGS. 19A - B: Example charts ranking nutrient conditions based on nutrient scores for carbon and nitrogen replacements. (FIG. 19A) Nutrient scores for nutrient replacement for carbon sources (NR-Carbon) are shown. Each point represents a specific carbon source, shaded according to its chemical category. Nutrient scores are ranked from lowest (poor) to highest (good), reflecting the capacity of each nutrient to support growth. The bar above the plot provides a visual summary of nutrient categories, corresponding to the shaded points. (FIG. 19B) Nutrient scores for nutrient replacement for nitrogen sources are shown. Similar to panel (a), each point represents a nitrogen source, categorized and ranked by nutrient score. The summary bar above reflects the distribution of nutrient categories. FIG. 20 includes a chart ranking nutrient supplements based on nutrient scores, in accordance with implementations. DETAILED DESCRIPTION OF THE DISCLOSURE Definitions Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All nucleotide sequences provided herein are presented in the 5′ to 3′ direction. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, particular, non-limiting exemplary methods, devices, and materials are now described. All technical and patent publications cited herein are incorporated herein by reference in their entirety. In some respects, this disclosure contains author name and date or an Arabic number to refer to a citation, the complete bibliographic information for which can be found immediately preceding the claims. Nothing herein is to be construed as an admission that the disclosure is not entitled to antedate such disclosure by virtue of prior disclosure. The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of cell culture, tissue culture, immunology, molecular biology, microbiology, cell biology and recombinant DNA, which are within the skill of the art. See, e.g., Green and Sambrook eds. (2012) Molecular Cloning: A Laboratory Manual, 4thedition; the series Ausubel et al. eds. (2015) Current Protocols in Molecular Biology; the series Methods in Enzymology (Academic Press, Inc., N.Y.); MacPherson et al. (2015) PCR 1: A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; McPherson et al. (2006) PCR: The Basics (Garland Science); Harlow and Lane eds. (1999) Antibodies, A Laboratory Manual; Greenfield ed. (2014) Antibodies, A Laboratory Manual; Freshney (2010) Culture of Animal Cells: A Manual of Basic Technique, 6thedition; Gait ed. (1984) Oligonucleotide Synthesis; U.S. Pat. No. 4,683,195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Herdewijn ed. (2005) Oligonucleotide Synthesis: Methods and Applications; Hames and Higgins eds. (1984) Transcription and Translation; Buzdin and Lukyanov ed. (2007) Nucleic Acids Hybridization: Modern Applications; Immobilized Cells and Enzymes (IRL Press (1986)); Grandi ed. (2007) In Vitro Transcription and Translation Protocols, 2ndedition; Guisan ed. (2006) Immobilization of Enzymes and Cells; Perbal (1988) A Practical Guide to Molecular Cloning, 2ndedition; Miller and Calos eds, (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London); Lundblad and Macdonald eds. (2010) Handbook of Biochemistry and Molecular Biology, 4thedition; and Herzenberg et al. eds (1996) Weir's Handbook of Experimental Immunology, 5thedition. All numerical designations, e.g., pH, temperature, time, concentration, and molecular weight, including ranges, are approximations which are varied (+) or (−) by increments of 1.0 or 0.1, as appropriate or alternatively by a variation of + / − 15%, or alternatively 10% or alternatively 5% or alternatively 2%. It is to be understood, although not always explicitly stated, that all numerical designations are preceded by the term “about”. It also is to be understood, although not always explicitly stated, that the reagents described herein are merely exemplary and that equivalents of such are known in the art. As used in the specification and claims, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a polypeptide” includes a plurality of polypeptides, including mixtures thereof. As used herein, the term “comprising” is intended to mean that the compositions and methods include the recited elements, but do not exclude others. “Consisting essentially of” when used to define compositions and methods, shall mean excluding other elements of any essential significance to the combination for the intended use. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers, such as phosphate buffered saline, preservatives, and the like. “Consisting of” shall mean excluding more than trace elements of other ingredients and substantial method steps for administering the compositions disclosed herein. Embodiments defined by each of these transition terms are within the scope of this disclosure. As used herein, the term “optional” or “optionally” means that the subsequently described circumstance may or may not occur, so that the description includes instances where the circumstance occurs and instances where it does not. As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”). As used herein, the term “about” is used to indicate that a value includes the standard deviation of error for the device or method being employed to determine the value. The term “about” when used before a numerical designation, e.g., temperature, time, amount, and concentration, including range, indicates approximations which may vary by (+) or (–) 15%, 10%, 5%, 3%, 2%, or 1 %. The term “substantially” or “essentially” means nearly totally or completely, for instance, 95% or greater of some given quantity. In some embodiments, “substantially” or “essentially” means 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. As used herein, the term “host cell” refers not only to the particular subject cell but to the progeny or potential progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term as used herein. The host cell can be a prokaryotic or a eukaryotic cell. As used herein, the term “substantially homogeneous population of cells” refers to a plurality of cells that are of the nearly totally or completely the same kind, e.g., at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99%, or 100% phenotypically similar or containing the exogenous iModulon as disclosed herein. As used herein, the term “media formulation” and “substrate” refers to a liquid or surface having specific amounts of carbon, nitrogen, micronutrients such as vitamins and trace elements, and other physical characteristics, such as pH, supportive of cell viability and growth. As used herein, the term “modified” or “modifying” refers to changing the quantity one or more carbon, nitrogen, micronutrient, or other physical characteristic of a media formulation or substrate. As used herein, the term “optimized media or substrate compatibility” refers to a modified media formulation of substrate wherein at least one iModulon gene cluster is determined to be maximally or minimally expressed. In some embodiments, a first sequence (nucleic acid sequence or amino acid) is compared to a second sequence, and the identity percentage between the two sequences can be calculated. In further embodiments, the first sequence can be referred to herein as an equivalent and the second sequence can be referred to herein as a reference sequence. In yet further embodiments, the identity percentage is calculated based on the full-length sequence of the first sequence. In other embodiments, the identity percentage is calculated based on the full-length sequence of the second sequence. The term “protein”, “peptide” and “polypeptide” are used interchangeably and in their broadest sense to refer to a compound of two or more subunit amino acids, amino acid analogs or peptidomimetics. The subunits may be linked by peptide bonds. In another embodiment, the subunit may be linked by other bonds, e.g., ester, ether, etc. A protein or peptide must contain at least two amino acids and no limitation is placed on the maximum number of amino acids which may comprise a protein's or peptide's sequence. As used herein the term “amino acid” refers to either natural and / or unnatural or synthetic amino acids, including glycine and both the D and L optical isomers, amino acid analogs and peptidomimetics. The terms “polynucleotide” and “oligonucleotide” are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides or analogs thereof. Polynucleotides can have any three-dimensional structure and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: a gene or gene fragment (for example, a probe, primer, EST or SAGE tag), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, RNAi, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes and primers. A polynucleotide can comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the polynucleotide. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide can be further modified after polymerization, such as by conjugation with a labeling component. The term also refers to both double- and single-stranded molecules. Unless otherwise specified or required, any embodiment disclosed herein that is a polynucleotide encompasses both the double-stranded form and each of two complementary single-stranded forms known or predicted to make up the double-stranded form. A polynucleotide is composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); thymine (T); and uracil (U) for thymine when the polynucleotide is RNA. Thus, the term “polynucleotide sequence” is the alphabetical representation of a polynucleotide molecule. This alphabetical representation can be input into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching. The term “isolated” or “recombinant” as used herein with respect to nucleic acids, such as DNA or RNA, refers to molecules separated from other DNAs or RNAs, respectively that are present in the natural source of the macromolecule as well as polypeptides. The term “isolated or recombinant nucleic acid” is meant to include nucleic acid fragments which are not naturally occurring as fragments and would not be found in the natural state. The term “isolated” is also used herein to refer to polynucleotides, polypeptides and proteins that are isolated from other cellular proteins and is meant to encompass both purified and recombinant polypeptides. In other embodiments, the term “isolated or recombinant” means separated from constituents, cellular and otherwise, in which the cell, tissue, polynucleotide, peptide, polypeptide, protein, antibody or fragment(s) thereof, which are normally associated in nature. For example, an isolated cell is a cell that is separated from tissue or cells of dissimilar phenotype or genotype. An isolated polynucleotide is separated from the 3′ and 5′ contiguous nucleotides with which it is normally associated in its native or natural environment, e.g., on the chromosome. As is apparent to those of skill in the art, a non-naturally occurring polynucleotide, peptide, polypeptide, protein, antibody or fragment(s) thereof, does not require “isolation” to distinguish it from its naturally occurring counterpart. It is to be inferred without explicit recitation and unless otherwise intended, that when the present disclosure relates to a polypeptide, protein, polynucleotide or antibody, an equivalent or a biologically equivalent of such is intended within the scope of this disclosure. As used herein, the term “biological equivalent thereof” is intended to be synonymous with “equivalent thereof” when referring to a reference protein, antibody, fragment, polypeptide or nucleic acid, intends those having minimal homology while still maintaining desired structure or functionality. Unless specifically recited herein, it is contemplated that any polynucleotide, polypeptide or protein mentioned herein also includes equivalents thereof. In one aspect, an equivalent polynucleotide is one that hybridizes under stringent conditions to the polynucleotide or complement of the polynucleotide as described herein for use in the described methods. In another aspect, an equivalent antibody or antigen binding polypeptide intends one that binds with at least 70%, or alternatively at least 75%, or alternatively at least 80%, or alternatively at least 85%, or alternatively at least 90%, or alternatively at least 95% affinity or higher affinity to a reference antibody or antigen binding fragment. In another aspect, the equivalent thereof competes with the binding of the antibody or antigen binding fragment to its antigen tinder a competitive ELISA assay. In another aspect, an equivalent intends at least about 80% homology or identity and alternatively, at least about 85%, or alternatively at least about 90%, or alternatively at least about 95%, or alternatively 98% percent homology or identity and exhibits substantially equivalent biological activity to the reference protein, polypeptide or nucleic acid. A polynucleotide or polynucleotide region (or a polypeptide or polypeptide region) having a certain percentage (for example, 80%, 85%, 90%, or 95%) of “sequence identity” to another sequence means that, when aligned, that percentage of bases (or amino acids) are the same in comparing the two sequences. The alignment and the percent homology or sequence identity can be determined using software programs known in the art, for example those described in Current Protocols in Molecular Biology (Ausubel et al., eds. 1987) Supplement 30, section 7.7.18, Table 7.7.1. In certain embodiments, default parameters are used for alignment. A non-limiting exemplary alignment program is BLAST, using default parameters. In particular, exemplary programs include BLASTN and BLASTP, using the following default parameters: Genetic code=standard; filter=none; strand=both; cutoff=60; expect=10; Matrix=BLOSUM62; Descriptions=50 sequences; sort by=HIGH SCORE; Databases=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Details of these programs can be found at the following Internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST. Sequence identity and percent identity were determined by incorporating them into clustalW (available at the web address:align.genome.jp, last accessed on Mar. 7, 2011. “Homology” or “identity” or “similarity” refers to sequence similarity between two peptides or between two nucleic acid molecules. Homology can be determined by comparing a position in each sequence which may be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same base or amino acid, then the molecules are homologous at that position. A degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. An “unrelated” or “non-homologous” sequence shares less than 40% identity, or alternatively less than 25% identity, with one of the sequences of the present disclosure. “Homology” or “identity” or “similarity” can also refer to two nucleic acid molecules that hybridize under stringent conditions. “Hybridization” refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The hydrogen bonding may occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex may comprise two strands forming a duplex structure, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute a step in a more extensive process, such as the initiation of a PCR reaction, or the enzymatic cleavage of a polynucleotide by a ribozyme. Examples of stringent hybridization conditions include: incubation temperatures of about 25° C. to about 37° C.; hybridization buffer concentrations of about 6×SSC to about 10×SSC; formamide concentrations of about 0% to about 25%; and wash solutions from about 4×SSC to about 8×SSC. Examples of moderate hybridization conditions include: incubation temperatures of about 40° C. to about 50° C.; buffer concentrations of about 9×SSC to about 2×SSC; formamide concentrations of about 30% to about 50%; and wash solutions of about 5×SSC to about 2×SSC. Examples of high stringency conditions include: incubation temperatures of about 55° C. to about 68° C.; buffer concentrations of about 1×SSC to about 0.1×SSC; formamide concentrations of about 55% to about 75%; and wash solutions of about 1×SSC, 0.1×SSC, or deionized water. In general, hybridization incubation times are from 5 minutes to 24 hours, with 1, 2, or more washing steps, and wash incubation times are about 1, 2, or 15 minutes. SSC is 0.15 M NaCl and 15 mM citrate buffer. It is understood that equivalents of SSC using other buffer systems can be employed. As used herein, “expression” refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently being translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in an eukaryotic cell. The term “encode” as it is applied to polynucleotides refers to a polynucleotide which is said to “encode” a polypeptide if, in its native state or when manipulated by methods well known to those skilled in the art, it can be transcribed and / or translated to produce the mRNA for the polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom. “Liposomes” are microscopic vesicles consisting of concentric lipid bilayers. Structurally, liposomes range in size and shape from long tubes to spheres, with dimensions from a few hundred Angstroms to fractions of a millimeter. Vesicle-forming lipids are selected to achieve a specified degree of fluidity or rigidity of the final complex providing the lipid composition of the outer layer. These are neutral (cholesterol) or bipolar and include phospholipids, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE), phosphatidylinositol (PI), and sphingomyelin (SM) and other types of bipolar lipids including but not limited to dioleoylphosphatidylethanolamine (DOPE), with a hydrocarbon chain length in the range of 14-22, and saturated or with one or more double C═C bonds. Examples of lipids capable of producing a stable liposome, alone, or in combination with other lipid components are phospholipids, such as hydrogenated soy phosphatidylcholine (HSPC), lecithin, phosphatidylethanolamine, lysolecithin, lysophosphatidylethanol-amine, phosphatidylserine, phosphatidylinositol, sphingomyelin, cephalin, cardiolipin, phosphatidic acid, cerebrosides, distearoylphosphatidylethan-olamine (DSPE), dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), palmitoyloteoylphosphatidylcholine (POPC), palmitoyloleoylphosphatidylethanolamine (POPE) and dioleoylphosphatidylethanolamine 4-(N-maleimido-triethyl)cyclohexane-1- carboxylate (DOPE-mal). Additional non-phosphorous containing lipids that can become incorporated into liposomes include stearylamine, dodecylamine, hexadecylamine, isopropyl myristate, triethanolamine-lauryl sulfate, alkyl-aryl sulfate, acetyl palmitate, glycerol ricinoleate, hexadecyl stereate, amphoteric acrylic polymers, polyethyloxylated fatty acid amides, and the cationic lipids mentioned above (DDAB, DODAC, DMRIE, DMTAP, DOGS, DOTAP (DOTMA), DOSPA, DPTAP, DSTAP, DC-Chol). Negatively charged lipids include phosphatidic acid (PA), dipalmitoylphosphatidylglycerol (DPPG), dioteoylphosphatidylglycerol and (DOPG), dicetylphosphate that are able to form vesicles. Typically, liposomes can be divided into three categories based on their overall size and the nature of the lamellar structure. The three classifications, as developed by the New York Academy Sciences Meeting, “Liposomes and Their Use in Biology and Medicine,” December 1977, are multi-lamellar vesicles (MLVs), small uni-lamellar vesicles (SUVs) and large uni-lamellar vesicles (LUVs). The polynucleotides can be encapsulated in such for administration in accordance with the methods described herein. A “micelle” is an aggregate of surfactant molecules dispersed in a liquid colloid. A typical micelle in aqueous solution forms an aggregate with the hydrophilic “head” regions in contact with surrounding solvent, sequestering the hydrophobic tail regions in the micelle center. This type of micelle is known as a normal phase micelle (oil-in-water micelle). Inverse micelles have the head groups at the center with the tails extending out (water-in-oil micelle). Micelles can be used to attach a polynucleotide, polypeptide, antibody or composition described herein to facilitate efficient delivery to the target cell or tissue. Also included as a micelles are lipid nanoparticles. A “gene delivery vehicle” is defined as any molecule that can carry inserted polynucleotides into a host cell. Examples of gene delivery vehicles are liposomes, micelles biocompatible polymers, including natural polymers and synthetic polymers; lipoproteins; polypeptides; polysaccharides; lipopolysaccharides; artificial viral envelopes; metal particles; and bacteria, or viruses, such as baculovirus, adenovirus and retrovirus, bacteriophage, cosmid, plasmid, fungal vectors and other recombination vehicles typically used in the art which have been described for expression in a variety of eukaryotic and prokaryotic hosts, and may be used for gene therapy as well as for simple protein expression. A polynucleotide disclosed herein can be delivered to a cell or tissue using a gene delivery vehicle. “Gene delivery,” “gene transfer,” “transducing,” and the like as used herein, are terms referring to the introduction of an exogenous polynucleotide (sometimes referred to as a “transgene”) into a host cell, irrespective of the method used for the introduction. Such methods include a variety of well-known techniques such as vector-mediated gene transfer (by, e.g., viral infection / transfection, or various other protein-based or lipid-based gene delivery complexes) as well as techniques facilitating the delivery of “naked” polynucleotides (such as electroporation, “gene gun” delivery and various other techniques used for the introduction of polynucleotides). The introduced polynucleotide may be stably or transiently maintained in the host cell. Stable maintenance typically requires that the introduced polynucleotide either contains an origin of replication compatible with the host cell or integrates into a replicon of the host cell such as an extrachromosomal replicon (e.g., a plasmid) or a nuclear or mitochondrial chromosome. A number of vectors are known to be capable of mediating transfer of genes to mammalian cells, as is known in the art and described herein. A “plasmid” is an extra-chromosomal DNA molecule separate from the chromosomal DNA which is capable of replicating independently of the chromosomal DNA. In many cases, it is circular and double-stranded. Plasmids provide a mechanism for horizontal gene transfer within a population of microbes and typically provide a selective advantage under a given environmental state. Plasmids may carry genes that provide resistance to naturally occurring antibiotics in a competitive environmental niche, or alternatively the proteins produced may act as toxins under similar circumstances. “Plasmids” used in genetic engineering are called “plasmid vectors”. Many plasmids are commercially available for such uses. The gene to be replicated is inserted into copies of a plasmid containing genes that make cells resistant to particular antibiotics and a multiple cloning site (MCS, or polylinker), which is a short region containing several commonly used restriction sites allowing the easy insertion of DNA fragments at this location. Another major use of plasmids is to make large amounts of proteins. In this case, researchers grow bacteria containing a plasmid harboring the gene of interest. Just as the bacterium produces proteins to confer its antibiotic resistance, it can also be induced to produce large amounts of proteins from the inserted gene. This is a cheap and easy way of mass-producing a gene or the protein it then codes for. A “yeast artificial chromosome” or “YAC” refers to a vector used to clone large DNA fragments (larger than 100 kb and up to 3000 kb). It is an artificially constructed chromosome and contains the telomeric, centromeric, and replication origin sequences needed for replication and preservation in yeast cells. Built using an initial circular plasmid, they are linearized by using restriction enzymes, and then DNA ligase can add a sequence or gene of interest within the linear molecule by the use of cohesive ends. Yeast expression vectors, such as YACs, YIps (yeast integrating plasmid), and YEps (yeast episomal plasmid), are extremely useful as one can get eukaryotic protein products with posttranslational modifications as yeasts are themselves eukaryotic cells, however YACs have been found to be more unstable than BACs, producing chimeric effects. A “viral vector” is defined as a recombinantly produced virus or viral particle that comprises a polynucleotide to be delivered into a host cell, either in vivo, ex vivo or in vitro. Examples of viral vectors include retroviral vectors, adenovirus vectors, adeno-associated virus vectors, alphavirus vectors and the like. Infectious tobacco mosaic virus (TMV)-based vectors can be used to manufacturer proteins and have been reported to express Griffithsin in tobacco leaves (O'Keefe et al. (2009) Proc. Nat. Acad. Sci. USA 106(15):6099-6104). Alphavirus vectors, such as Semliki Forest virus-based vectors and Sindbis virus-based vectors, have also been developed for use in gene therapy and immunotherapy. See, Schlesinger & Dubensky (1999) Curr. Opin. Biotechnol. 5:434-439 and Ying et al. (1999) Nat. Med. 5(7):823-827. In aspects where gene transfer is mediated by a retroviral vector, a vector construct refers to the polynucleotide comprising the retroviral genome or part thereof, and a therapeutic gene. Further details as to modern methods of vectors for use in gene transfer may be found in, for example, Kotterman et al. (2015) Viral Vectors for Gene Therapy: Translational and Clinical Outlook Annual Review of Biomedical Engineering 17. As used herein, “retroviral mediated gene transfer” or “retroviral transduction” carries the same meaning and refers to the process by which a gene or nucleic acid sequences are stably transferred into the host cell by virtue of the virus entering the cell and integrating its genome into the host cell genome. The virus can enter the host cell via its normal mechanism of infection or be modified such that it binds to a different host cell surface receptor or ligand to enter the cell. As used herein, retroviral vector refers to a viral particle capable of introducing exogenous nucleic acid into a cell through a viral or viral-like entry mechanism. Retroviruses carry their genetic information in the form of RNA; however, once the virus infects a cell, the RNA is reverse-transcribed into the DNA form which integrates into the genomic DNA of the infected cell. The integrated DNA form is called a provirus. In aspects where gene transfer is mediated by a DNA viral vector, such as an adenovirus (Ad) or adeno-associated virus (AAV), a vector construct refers to the polynucleotide comprising the viral genome or part thereof, and a transgene. Adenoviruses (Ads) are a relatively well characterized, homogenous group of viruses, including over 50 serotypes. See, e.g., PCT International Pat. Application Publication No. WO 95 / 27071. Ads do not require integration into the host cell genome. Recombinant Ad derived vectors, particularly those that reduce the potential for recombination and generation of wild-type virus, have also been constructed. See, PCT International Pat. Application Publication Nos. WO 95 / 00655 and WO 95 / 11984, Wild-type AAV has high infectivity and specificity integrating into the host cell's genome. See, Hermonat & Muzyczka (1984) Proc. Natl. Acad. Sci. USA 81:6466-6470 and Lebkowski et al. (1988) Mol. Cell. Biol. 8:3988-3996. Vectors that contain both a promoter and a cloning site into which a polynucleotide can be operatively linked are well known in the art. Such vectors are capable of transcribing RNA in vitro or in vivo, and are commercially available from sources such as Stratagene (La Jolla, Calif.) and Promega Biotech (Madison, Wis.). In order to optimize expression and / or in vitro transcription, it may be necessary to remove, add or alter 5′ and / or 3′ untranslated portions of the clones to eliminate extra, potential inappropriate alternative translation initiation codons or other sequences that may interfere with or reduce expression, either at the level of transcription or translation. Alternatively, consensus ribosome binding sites can be inserted immediately 5′ of the start codon to enhance expression. As used herein, the term “expression” refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently being translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell. The expression level of a gene may be determined by measuring the amount of mRNA or protein in a cell or tissue sample. In one aspect, the expression level of a gene from one sample may be directly compared to the expression level of that gene from a control or reference sample. Gene delivery vehicles also include DNA / liposome complexes, micelles and targeted viral protein-DNA complexes. Liposomes that also comprise a targeting antibody or fragment thereof can be used in the methods disclosed herein. In addition to the delivery of polynucleotides to a cell or cell population, direct introduction of the proteins described herein to the cell or cell population can be done by the non-limiting technique of protein transfection, alternatively culturing conditions that can enhance the expression and / or promote the activity of the proteins disclosed herein are other non-limiting techniques. “Eukaryotic cells” comprise all of the life kingdoms except monera. They can be easily distinguished through a membrane-bound nucleus. Animals, plants, fungi, and protists are eukaryotes or organisms whose cells are organized into complex structures by internal membranes and a cytoskeleton. The most characteristic membrane-bound structure is the nucleus. Unless specifically recited, the term “host” includes a eukaryotic host, including, for example, yeast, higher plant, insect and mammalian cells. Non-limiting examples of eukaryotic cells or hosts include simian, bovine, porcine, murine, rat, avian, reptilian and human. “Prokaryotic cells” that usually lack a nucleus or any other membrane-bound organelles and are divided into two domains, bacteria and archaea. In addition to chromosomal DNA, these cells can also contain genetic information in a circular loop called on episome. Bacterial cells are very small, roughly the size of an animal mitochondrion (about 1-2 μm in diameter and 10 μm long). Prokaryotic cells feature three major shapes: rod shaped, spherical, and spiral. Instead of going through elaborate replication processes like eukaryotes, bacterial cells divide by binary fission. Examples include but are not limited to Bacillus bacteria, E. coli bacterium, and Salmonella bacterium. As used herein, “solid phase support” or “solid support”, used interchangeably, is not limited to a specific type of support. Rather a large number of supports are available and are known to one of ordinary skill in the art. Solid phase supports include silica gels, resins, derivatized plastic films, glass beads, cotton, plastic beads, alumina gels. As used herein, “solid support” also includes synthetic antigen-presenting matrices, cells, and liposomes. A suitable solid phase support may be selected on the basis of desired end use and suitability for various protocols. For example, for peptide synthesis, solid phase support may refer to resins such as polystyrene (e.g., PAM-resin obtained from Bachem Inc., Peninsula Laboratories, etc.), POLYHIPE® resin (obtained from Aminotech, Canada), polyamide resin (obtained from Peninsula Laboratories), polystyrene resin grafted with polyethylene glycol (TentaGel®, Rapp Polymere, Tubingen, Germany) or polydimethylacrylamide resin (obtained from Milligen / Biosearch, Calif.). An example of a solid phase support include glass, polystyrene, polypropylene, polyethylene, dextran, nylon, amylases, natural and modified celluloses, polyacrylamides, gabbros, and magnetite. The nature of the carrier can be either soluble to some extent or insoluble. The support material may have virtually any possible structural configuration so long as the coupled molecule is capable of binding to a polynucleotide, polypeptide or antibody. Thus, the support configuration may be spherical, as in a bead, or cylindrical, as in the inside surface of a test tube, or the external surface of a rod. Alternatively, the surface may be flat such as a sheet, test strip, etc. or alternatively polystyrene beads. Those skilled in the art will know many other suitable carriers for binding antibody or antigen, or will be able to ascertain the same by use of routine experimentation. Ampicillin resistance denotes the ability of bacteria to withstand the effects of ampicillin, an antibiotic that inhibits bacterial cell wall synthesis. This resistance can be quantified by measuring the minimum inhibitory concentration (MIC). The MIC is determined by treating cells with varying concentrations of ampicillin and identifying the minimum concentration at which ampicillin completely inhibits the growth of the bacterium. As used herein, “RNA expression profile” or “gene expression data” refers to measurements of mRNA levels, showing the pattern of genes express by a cell at the transcription level. mRNA is a single-stranded molecule of RNA that corresponds to the genetic sequence of a gene and is read by a ribosome in the process of synthesizing a protein. RNA expression profiles and gene expression data can be measured using standard laboratory techniques, including PCR, using primers and techniques disclosed herein, or by other techniques known to one of ordinary skill in the art. As used herein, “iModulon gene cluster” refers to a group of genes that that is independently modulated, wherein the group of genes is controlled by the same or related transcription regulators. In some embodiments, an iModulon comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes, accounting for a part or all of a biological process, such as Acetate metabolism. As used herein, “booster module” refers to genetically modifying a host cell’s genome to artificially include a genetic circuit comprising additional transcription regulators. In some embodiments, the transcription regulator araC is position after the original Acetate iModulon gene cluster to repress the transcription regulator lacl, resulting in amplified Acetate iModulon activity specifically under poor carbon source conditions. A booster module may comprise, for example, SEQ ID No. 83 and / or SEQ ID No. 84, or an equivalent of each thereof. As used herein, “determining one or more iModulons” refers to measuring whether an iModulon gene cluster’s RNA expression level is upregulated or downregulated compared to a control. As used herein, upregulated or downregulated expression of an iModulon gene cluster refers to increased or decreased RNA expression of the iModulon gene cluster and / or one or more members of the cluster, respectively. As used herein, a gene of unidentified function refers to a gene having a function that has not been characterized in the preexisting art. Therefore, a skilled person would not be able to ascertain the function of a gene of unidentified function based on the preexisting art. As used herein, a substrate (e.g., substrate compatibility) refers to a material that a cell grows on. For example, the substrate may comprise a matrix of molecules used to support the growth of the cell, including nutrients to support the growth of the cell. In some embodiments, the substrate comprises a gelling agent, e.g., agar. As used herein, “strain of phenotypic expression” refers to a strain genetically engineered using techniques disclosed herein or know to those skilled in the art to overexpress or underexpress an iModulon gene cluster, the iModulon gene cluster accounting for a part or all of a biological process, such as Acetate metabolism. As used herein, “metabolic iModulon” refers to an iModulon gene cluster accounting for a part or all of a carbon-based metabolic biological process, such as Acetate metabolism. An acetate iModulon may comprise, for example, SEQ ID Nos: 77-82 and optionally SEQ ID Nos: 83-84 or an equivalent of each thereof. As used herein, “trace element iModulon” refers to an iModulon gene cluster accounting for a part or all of a cell’s trace element homeostasis, such as Zn2+transport. Trace elements are present in relatively small quantities and primarily function as enzyme cofactors. As used herein, “stress-related iModulon” refers to an iModulon gene cluster accounting for a part or all of a cell’s response to stressors, such as temperature or osmolarity. As used herein, “TTT” refers to transporters and tripartite tricarboxylate transporters, “TPP” refers to thiamine pyrophosphate, “CSP” refers to cold shock proteins, FNR refers to fumarate and nitrate reduction, TRAP refers to tripartite ATP-independent periplasmic transporters, GABA refers to gamma-aminobutyric acid, AGL refers to alpha-1.6 glucosidase, Null and UC refers to unknown function, and GlcN refers to N- acetylglucosamine. As used herein, “iteratively modifying the cell culture media” refers to repetitively modifying the media or substrate and assaying iModulon expression until the iModulon expression is maximized, minimized, or optimized. Modes For Carrying Out the Disclosure Despite its fundamental importance, media formulation is still empirical and lacks a detailed genetic and molecular basis. However, the advent of genome-scale science, with available full genome wide profiling methods allow the determination of the full molecular state underlying an observed phenotypic state. In particular, since 2013, after purification mRNA methods were established for bacteria10,11, a large number of transcriptomic profiles have become available in the public domain12,13. Applicant has applied ICA to find source regulatory signals in bacterial transcriptomes. iModulons are fundamental units of bacterial transcriptomes that have been found to represent the genetic basis for various cellular functions and are associated with particular transcriptional regulator(s). Many identified iModulons include genes that are distantly located on the genome, genes of unknown functions, or contain accessory genes that augment the targeted cellular function. iModulons thus represent a new scale of synthetic biology to transfer naturally evolved traits across species. Here, Applicant provides methods comprising the use of transcriptome analysis of iModulons in response to media and substrate nutrient perturbations to identify 64 responsive iModulons. The methods can further comprise modification of media and substrate formulations to correct, or optimize, the iModulon transcriptional signature in response to nutrient perturbations. Further, Applicant upregulates or downregulates the expression of iModulons to optimize the cells’ response to media or substrate formulations having nutrient perturbations. Methods of Preparing Host Cells with Optimized Media Compatibility This disclosure also provides methods for creating genetically modified host cells that differentially express the iModulons of this disclosure. Optionally, the Acetate iModulon gene cluster is upregulated or downregulated with a booster module. The genetically modified cells can be produced by insertion of upstream regulatory sequences such as promoters or gene activators (see, U.S. Patent No. 5,733,761), using methods as described herein or known in the art. The polynucleotides of the present disclosure can serve as primers for the detection of genes or gene transcripts that are expressed in cells described herein. In this context, amplification means any method employing a primer-dependent polymerase capable of replicating a target sequence with reasonable fidelity. Amplification may be carried out by natural or recombinant DNA-polymerases such as T7 DNA polymerase, Klenow fragment of E. coli DNA polymerase, and reverse transcriptase. For illustration purposes only, a primer is the same length as that identified for probes. One method to amplify polynucleotides is PCR and kits for PCR amplification are commercially available. After amplification, the resulting DNA fragments can be detected by any appropriate method known in the art, e.g., by agarose gel electrophoresis followed by visualization with ethidium bromide staining and ultraviolet illumination. Methods for administering an effective amount of a gene delivery vector or vehicle to a cell have been developed and are known to those skilled in the art and described herein. Methods for detecting gene expression in a cell are known in the art and include techniques such as in hybridization to DNA microarrays, in situ hybridization, PCR, RNase protection assays and Northern blot analysis and functional assays as described herein. Such methods are useful to detect and quantify expression of the gene in a cell. Alternatively, expression of the polypeptide can be detected by various methods. In particular it is useful to prepare polyclonal or monoclonal antibodies that are specifically reactive with the target polypeptide. Such antibodies are useful for visualizing cells that express the polypeptide using techniques such as immunohistology, ELISA, and Western blotting. These techniques can be used to determine expression level of the expressed polynucleotide. In one aspect, the population of cells are substantially homogenous or a clonal population of the cells. The cells can be eukaryotic or prokaryotic, examples of such include E. coli, P. putida, Corynebacterium glutamicum, and genetically minimized species such as Vibrio natriegens and JCVI_Syn3A (see https: / / elifesciences.org / articles / 36842, incorporated herein by reference). The cells are commercially available from vendors. Methods of Preparing Optimized Media Formulations In another aspect, a method for preparing optimized media formulations is provided herein, comprising, or consisting essentially of, or yet further consisting of growing a first population of cells in a media formulation, on a substrate, or on a substrate in a media formulation; isolating RNA from the plurality of cells; detecting by an RNA expression profile upregulation or downregulation of an iModulon gene cluster comprising a metabolic, trace element, or stress-related iModulon; and iteratively modifying the media formulation or the substrate to downregulate or upregulate the iModulon gene cluster, growing a second population of cells in the modified media formulation or on the modified substrate, isolating RNA from the second plurality of cells, and detecting optimized expression of the iModulon gene cluster, thereby informing the next iteration of media formulation until iModulon expression is optimized. The cells for the media or substrate formulation can be eukaryotic or prokaryotic, examples of such are provided herein. In another aspect, the media formulation comprises, or consists essentially of, or yet further consists of acetate, ethanol, cellobiose, galactose, glycogen, or GlcN, thereby improving cell growth in media formulations having stress-free or stress-inducing carbon sources. Methods to culture the cells are known in the art and described herein. The following experimental examples are provided to illustrate the clauses and aspects of this disclosure. Experimental The following experimental examples are provided to illustrate the clauses and aspects of this disclosure. Example 1 As transcriptomic compendia have grown, source signal extraction algorithms have been applied to identify sets of independently modulated sets of genes (called iModulons). Independent component analysis (ICA), a machine learning method, applied to transcriptomic compendia for various bacteria has proved to be particularly effective14for identifying quantitative and strain-specific TRNs15,16. Since iModulons are big data analogs of regulons, mapping of known regulatory and molecular biology information has knowledge-enriched the ICA signals, leading to deep understanding of the modularization of TRNs and determination of their activity state16,17. iModulons can measure the activity states of 100s of cellular functions. These functions include metabolism, proteostasis, various stresses, two-component systems18, antibiotic response19,20, adaptations to stresses21and activation of latent phages22. The activity states of iModulons are a direct measure of the functional state of the TRN, and what the cell is sensing and responding to in a particular environment. For many organisms, enough transcriptomic data has become available to allow for the development of knowledge- enriched data analytics allowing for fundamental understanding of genome-wide cellular responses16. As an example of this method, disclosed herein is the generation of a transcriptomic compendium for Vibrio natriegens (Vn) under a variety of nutrient conditions. The measurement of the activity state of all iModulons under these conditions reveals genome- wide responses to different media compositions. The responses show the intricate relationship between nutrients and the TRN, where many responses were known, but most were not. These TRN responses surprisingly are highly informative about what cells sense and how they respond. Furthermore, they allow the formulation of media based on first biological principles and genome-wide cellular responses. Modularization of the transcriptome to evaluate cellular responses to nutrients To investigate Vn's cellular responses to media composition changes, iModulons related to carbon utilization, trace elements, stress responses, and uncharacterized iModulons were focused on to establish a comprehensive framework for evaluating the effects of various nutrients on cellular processes (FIG. 1A). The previous natPRECISE104 database23was expanded to natPRECISE148 by adding samples from use of diverse substrates and stress conditions, and analyzed using an ICA pipeline24, resulting in 64 iModulons that enhance insights into bacterial responses to nutritional and stress factors (FIG. 7). Each iModulon was assigned to one of 11 functional groups, providing a systems-level perspective (FIG. 1B and FIG. 7H). Significant overlap was observed when comparing the 64 iModulons disclosed herein with 45 known iModulons23, mapping previously characterized regulatory signals to the new findings (FIG. 8A, FIG. 8B). This overlap suggests that the TRN structure remains robust to adding new data15. Despite its relatively small size compared to transcriptomic compendium for other bacterial species (FIG. 7A), natPRECISE148 effectively captures unique catabolic iModulons (FIG. 8C, FIG. 8D, FIG. 8E). The updated natPRECISE148 compendium reveals an increase in most iModulon categories, including eight new catabolic processes, four stress responses, four element homeostasis, two amino acid biosynthesis, two nitrogen and energy responses, and two natural competence, demonstrating comprehensive coverage of responses to changes in media composition (FIG. 8A). Using these 64 iModulons, this disclosure covers a spectrum of transcriptional responses to diverse nutrient conditions (FIG. 1A). This disclosure covers various aspects of Vn nutrition, including nutrient uptake and catabolism, the involvement of metal ions as cofactors, their influence on protein synthesis and cell growth, and an array of specific stress responses. These stress responses include osmolarity and oxidative stress, challenges related to low glycolytic flux, and the activation of prophage genes. Condition-dependent iModulon activity levels allow the decoding intricate nutrient-cell response interactions and associated regulatory mechanisms. Transcriptomic responses to substrates reveal TTT- and TRAP-related functions 15 substrate utilization iModulons were analyzed in Vn across 21 substrates (FIG. 2A). Most iModulons were activated under their specific growth condition (the highlighted boxes with a thick border, FIG. 2A). Additionally, multiple substrates activated unexpected iModulons. Notably, the GalR and AGL iModulons, related to galactose and glycogen catabolism, showed a strong correlation (Pearson’s R= 0.62), indicating potential co- regulation (FIG. 2B and FIG. 9A) Transcription factors highly correlated with these iModulons (Pearson’s R < 0.68) could be potential AGL iModulon regulators (FIG. 9B). These experiments identified the TTT (Transporters and Tripartite Tricarboxylate Transporters) and TRAP (Tripartite ATP-Independent Periplasmic) iModulons. These systems, transporting small organic molecules via ion-electrochemical gradients25,26, exhibit varied substrate specificities25,27,28. Despite their distinct functional annotations, they share a common configuration of substrate binding along with small and large subunits (FIG. 2C, FIG. 2D). The GntR iModulon also contains other TRAP genes with varying amino acid similarities (FIG. 2D). The TTT and TRAP iModulons are highly activated under growth on more than eight substrates, suggesting their pleiotropic role in the uptake of a broad range of substrates (FIG. 2A). The TTT and TRAP iModulons, unique to Vn (FIG. 8D), were further investigated alongside uncharacterized transporters in the GalR and Rhamnose iModulons. Comparing knockout (KO) and WT (wild-type) strains’ growth rate revealed their roles in carbon utilization (FIG. 2E and FIG. 9C). KO strains of TTT and TRAP iModulons showed significantly reduced growth rates under three and five carbon conditions, respectively. The KO strains of GntR and TRAP iModulons demonstrated distinct nutrient impacts, excluding fumarate, highlighting the substrate specificity differences between these two TRAP genes (FIG. 2E). Deletion of uncharacterized transporter genes (RS09130–RS09145) in the GalR iModulon led to a significant decrease in growth rate (< 0.44-fold; P-value < 0.00015) under galactose and cellobiose, further emphasizing the role of GalR iModulon in cellobiose utilization. Notably, rhamnose utilization highly depended on the uncharacterized transporter (RS17440) in the Rhamnose iModulon and on both TTT and TRAP iModulons. Thus, iModularization of transcriptomic responses showed expected primary catabolic responses, revealed new pleiotropic uptake mechanisms, and showed substrate- specific activation of several cellular processes. These substrate-specific responses were further delineated in the next section. iModulon activities reveal interactions of trace elements with substrates The influence of trace elements in the medium on iModulon activity was examined using various substrates in M9Na media, formulated without any supplementary trace element solutions. This approach revealed the significance of specific trace elements with specific substrates. Among seven iModulons associated with trace elements homeostasis (rows in FIG. 3A), the Zur iModulon, associated with Zn2+transport, showed notable activation on many substrates, such as fructose and GlcNAc. Further analysis revealed correlations between the seven iModulons (FIG. 3B). The Zur iModulon, uncorrelated with the other six, suggests an independent regulatory mechanism. The Enterobactin iModulon shows a strong correlation (Pearson’s R = 0.80) with the Fur-1 iModulon, as expected from the overlap in their gene membership (FIG. 1B). Additionally, the Fur-1 iModulon exhibits a correlation (Pearson’s R = 0.53) with the Thiosulfate iModulon activity. Without wishing to be bound to any particular theory, this suggests that Fur regulation might also influence thiosulfate uptake, linking iron homeostasis and sulfur metabolism. Following the supplementation of ZnCl2(0, 1, and 3 µM as a typical concentration of trace element solution29into the M9Na media with both substrates that activate the Zur iModulon and those that do not. This approach revealed Zur iModulon activation’s relationship with zinc ion levels, clarifying Zn2+regulation under various growth conditions (FIG. 3C, FIG. 3D). This result indicates Zur iModulon’s key role in Zn+2regulation under these growth conditions. The growth responses to zinc supplementation were categorized as enhanced, neutral, and inhibited. Surprisingly, significant improvements in growth rate (P- value < 0.0195, excepting for trehalose) and maximum OD600 (P-value < 0.041) were observed in GlcNAc, glycerol, sucrose, and trehalose conditions (FIG. 3C and FIG. 9D). This enhancement indicates a Zn2+dependency in pathways, such as GlcNAc catabolism involving the zinc-dependent NagA enzyme30, contrasting with the non-zinc-dependent NagB enzyme, showing neutral growth response. Similar Zn2dependencies in glycerol31and trehalose32catabolic enzymes in other species suggest a comparable requirement for Vn (FIG. 9E). Conversely, Zn2+addition inhibited growth in fructose, glycogen, and fumarate conditions, evidenced by a decrease in growth rate (fructose, P-value < 2.1 × 10–5) or maximum OD600values (glycogen and fumarate, P-value < 3.4 × 10–4) (FIG. 3C and FIG. 9F). This inhibition suggests that Zn2+negatively affected enzymes such as fructose kinase33and glycogen debranching pullulanase34,35. Furthermore, the identification of a putative inorganic ion transporter within the Zur iModulon (FIG. 9G), a feature not observed in Zinc- related iModulons in other species (FIG. 9H), suggests a dual functionality in both zinc uptake and export, similar to that observed in Salmonella36. In the Pseudomonas putida iModulon database16,37, a deactivation of Zur iModulon under the fructose indicates a potential Zn2+-mediated regulatory requirement for the FruR iModulon's activities33. Thus, iModulon analysis related to trace elements highlights the critical influence of trace elements in affecting iModulon activity and impacting bacterial growth across different substrates. Trace elemental composition of media may thus need alteration depending on the substrate for optimal and stress free growth. iModulon activity suggests an impact of nutrients on membrane and prophage- related stresses Nine stress-related iModulons were employed to categorize substrates as stress-free (SF) or stress-inducing (SI) (FIG. 4A). Generally, SI substrates, except for glycogen, succinate, and ribose, showed lower growth rates than SF (average 0.6-fold, P-value = 0.0006) (FIG. 4B). Focusing on membrane-related stress, two unique osmolarity-related iModulons were identified in natPRECISE148: the Low-osmolarity and the Ectoine iModulons, responsible for putrescine and ectoine synthesis, respectively (FIG. 10A, FIG. 10B). The Low-osmolarity iModulon, enhancing membrane stability, activates under low salt conditions, whereas the Ectoine iModulon, an osmoprotectant, activates under high salt conditions (iModulon activity > 30). KO strains for these iModulons, subjected to varied salt concentrations, revealed their crucial role in osmotic stress adaptation, as indicated by changes in growth rates (FIG. 10C). A range of salt concentrations was created by adding varying amounts of NaCl to the medium. The experiment was conducted in 96-well plates with shaking, and OD600of these cultures was measured using a Tecan Infinite 200Pro microplate reader. The relative growth rates of the strains under these salt conditions were calculated based on their specific growth rates compared to the 200 mM NaCl concentration condition. In addition, RpoE, Ribosome, and Chaperone iModulons responded to large changes in salt concentration (FIGS. 10D – 10G). Acetate and rhamnose notably induced these iModulons (FIG. 4A) due to the use of sodium acetate raising sodium levels in the medium (250 mM of M9Na + 120 mM of Na+= 370 mM; FIG. 10B) and rhamnose causing low salt stress through sodium ion-gradient dependent uptake via TTT and TRAP iModulons (FIG. 10H). The RpoE iModulon’s response to extracytoplasmic stress in glycogen or ethanol conditions aligns with other species’ findings38. Furthermore, different substrates significantly influence the activation of Prophage- 1 and Prophage-2 iModulons (FIG. 4A and FIG. 10I). This observation suggests that specific nutrients can elicit prophage activation, a process that can be both stressful and resource- intensive for bacteria39,40. These results reveal intricate relationships between substrates and stresses that they inflict on the host. Chaperone activation highlights protein expression and stress response dependence on media composition The activities of the Chaperone iModulon was examined and it was observed that cellular stress responses, including low salt conditions, stimulate its activity (FIG. 4C, FIG. 10G). The correlation between Chaperone and Ribosome iModulon activities (FIG. 4D) indicated that certain conditions, such as rhamnose or ribose, might 1) challenge protein folding41, or 2) necessitate an increase in ribosome concentration, subsequently influencing the rate of protein synthesis in response to the intracellular nutritional state42,43. To assess how nutrients affect proteostasis, GFP protein expression was measured across various substrates in M9Na media (FIGS. 4E – 4F). Ribose showed a distinct pattern (FIG. 4E), with rhamnose activating Chaperone and Ribosome iModulons, this appeared to be an amplification due to the low-salt stress associated with rhamnose uptake (FIG. 4D and FIGS. 10E – 10G). Ribose, while not enhancing growth as rapidly as glucose (0.65-fold, P- value = 1.1 × 10–9), showed the highest levels of relative fluorescence units (RFU; 1.9-fold, P-value = 7.7 × 10–6) and RFU normalized by OD600 (1.2-fold, P-value = 0.027), indicating a significant increase in protein production within cells. Interestingly, ribose exhibited robust protein expression even without inducers (2.3-fold, P-value = 2.8 × 10–5, compared to glucose) (FIG. 4E), suggesting that it may activate a stress response that enhances translation more effectively than other carbons. Therefore, choosing substrates activating Ribosome and Chaperone iModulons might enhance protein production. Additionally, analysis revealed that the complex media BHIN (Brain Heart Infusion + 1.5% wt / vol NaCl) triggered several stress iModulons (FIG. 4A). This led to lower growth rates (0.5-fold, P-value = 1.5 × FIG. 10J), slower rates of RFU increase (0.7-fold, P- value = 0.048), and lower maximum RFU / OD600 levels (0.41-fold, P-value = 0.0005) versus LBv2 medium (FIG. 4G). These findings suggest that BHIN might not be ideal for Vn cultivation. Thus, substrates can induce specific stresses, with substrate choice impacting translation and protein production. This shows that substrate selection is crucial in bacterial cultivation and strain development, influencing growth and protein synthesis efficiency. Nutrients affect Acetate iModulon activity that is associated with glycolytic flux Striking behavior of the Acetate iModulon was observed across various substrates (FIG. 2A and FIG. 4A), prompting a deeper analysis. The Acetate iModulon, involved in both carbon utilization and stress response, displayed patterns indicative of its involvement in low glycolytic flux states44. This finding indicates the Acetate iModulon (FIG. 11A) could be a biomarker for low flux glycolytic states (FIG. 5A). In Escherichia coli, acetate metabolism involves phosphotransacetylase (Pta) and acetate kinase (Ack), with the acetate CoA-synthetase (AcsA) gene playing a pivotal role45(FIG. 5A). AcsA crucially converts acetate into acetyl-CoA, essential for the acetate switch process. Similarly, the Acetate iModulon in Vn includes the Acs enzyme, the acetate symporter (ActP), acyl-CoA synthetase, methylisocitrate lyase, and eight additional proteins with functions to be elucidated. Bacteria typically excrete acetate when utilizing standard carbon sources, like sugars, then re-uptake it when depleted, particularly in the stationary phase46(FIG. 11B). SF substrates that promote high glycolytic flux tend to lead to acetate secretion, as acetate inhibits both the glycolytic pathway and the TCA cycle45. Conversely, poor carbon sources inducing low glycolytic flux activate the Acetate iModulon, facilitating acetate reuse or limiting secretion44. For instance, E. coli growing on glycerol, a poor carbon source, employs a "carbon source foraging strategy", avoiding acetate production47. To further investigate this effect, 5 mM acetate44was added to various substrates in M9Na. This addition generally reduced growth rate (0.69–0.88-fold, P-value < 0.0001) in SF carbon samples, except for glucosamine, but improved growth rate under SI conditions (1.15–5.51-fold, P-value = 0.0002) (FIG. 5B and FIG. 11C). Rhamnose and ribose, not inducing the Acetate iModulon, exhibited expected growth patterns. KO studies of the Acetate iModulon gene (RS14410–RS14420) indicated improved growth rate in SF carbons (1.11–1.28-fold, P-value < 0.0001, except for gluconate and glucosamine) but reduced growth rate in SI conditions (0–0.89-fold, P-value < 0.0001, except for ribose and succinate) (FIG. 5C and FIG. 11D). Notably, both acetate and, unexpectedly, galactose significantly depend on the Acetate iModulon, underscoring its importance (FIG. 5D). To assess the effect of acetate on carbon source utilization in FIGS. 11C – 11D, WT strains were cultured with and without acetate (5 mM) in 96-well plates with shaking. Additionally, the functional roles of Acetate iModulon genes were examined by deleting genes (RS14410–RS14420) (FIG. 9C). Both WT and KO strains were cultured under identical conditions. OD600measurements were conducted using a Tecan Infinite 200 Pro microplate reader. Each experimental setup involved a 15 mM concentration of the respective carbon source in M9Na minimal media. The specific growth rates of these strains across different carbon conditions were measured. The functions of these strains were inferred from the relative growth rate compared to either the acetate-untreated sample or the WT strain. Data are presented as mean ± SEM from three biologically independent samples. Despite the limited availability of single nutrient perturbation samples (FIG. 7A), consistent patterns in the Acetate iModulon have been identified across species in the iModulonDB16. In the Pseudomonas putida iModulon database, certain carbon sources such as ferulate, citrate, coumarate, fructose, and serine, and specific genomic modifications48were found to activate the starvation-related BkdR iModulon37along with the Acetate iModulon. The simultaneous activation suggests a significant correlation (Pearsons’ R = 0.57, P < 0.001) between BkdR and Acetate iModulons (FIG. 11E). Similarly, in E. coli, the Acetate iModulon is activated by carbon sources such as fructose, acetate, and glycerol, highlighting their limited role as general-use carbon sources without inducing acetate Furthermore, a new interaction between the TTT and TRAP iModulons and SI substrates was found. These iModulons were significantly activated under SI conditions (average 21.1-fold for TTT iModulon and 381.8-fold for TRAP iModulon) (FIG. 5E). Consequently, there is a correlation between Acetate iModulon activity and the TTT and TRAP iModulons (FIG. 5F). Comparisons revealed that the TRAP iModulon activation is growth phage-independent, unlike the TTT iModulon, and uncorrelate with the GntR iModulon’s different TRAP genes (FIG. 5F). Overall, these findings provide insights into cells adaptation to poor carbon sources via the TTT / TRAP and Acetate iModulons (FIG. 11F). In summary, the Acetate iModulon's response to various nutrients revealed its importance as an indicator of low glycolytic flux, providing insights into bacterial metabolism and potential strategies for optimizing substrate utilization. Enhancing growth on poor carbon sources via Acetate iModulon boosting The potential use of the Acetate iModulon for strain engineering was evaluated, aiming to elevate its activity without disrupting its natural regulation. This strategy aims to boost growth under most SI carbon sources (FIG. 6A), keeping the Acetate iModulon inactive under SF carbon conditions45,51. Implementing this strategy, a ‘boost module’ genetic circuit was integrated adjacent to the original Acetate iModulon genes in the genome (FIG. 6B, FIG. 11G). This circuit included araC positioned after the original Acetate iModulon genes, tasked with repressing the lacI gene. This lacI, in turn, represses an additional set of Acetate iModulon genes within the boost module. Such a design amplifies the Acetate iModulon activity specifically under poor carbon source conditions. Two strains were engineered, Boost-A and Boost-B, by integrating this booster module after the original Acetate iModulon genes. Boost-A included acsA, RS14450, and acsP, whereas Boost-B contained only the key enzyme acsA. Experiments revealed minimal growth differences between the control and the two booster strains under SF conditions, except for Boost-B in glucosamine and mannitol conditions. However, under SI carbon sources, the engineered strains exhibited significant growth enhancements, as evidenced by increased growth rates (1.1–2.4-fold, P-value < 0.034) and shorter lag times (0.47–0.81-fold, P-value < 0.005) in the Boost-A strain (FIGS. 6C – 6E), except for ribose which did not induce the Acetate iModulon. Thus, activating the Acetate iModulon can substantially improve the utilization of poor carbon sources, demonstrating the practical application of iModulon information from nutrient change experiments for optimization of cellular functions. Experimental Discussion Modern genome-scale methods can be deployed to address optimal medium formulation by deploying new modularization methods of the bacterial transcriptome, determining the activation state of all identifiable independently modulated cellular processes as a function of media composition, accessing catabolism, trace elements use, proteostasis, phage activation, and stress responses3-5. Analyzing catabolic iModulon activity revealed pathways activated by different nutrients, showing new iModulons linked to substrate uptake. Notable are the pleiotropic TTT and TRAP iModulons25,26, essential for metabolizing seven substrates like rhamnose, ribose, and fumarate, underscoring their broad substrate utility in Vn52. New regulatory relationships and gene functions in catabolism are disclosed herein, such as interactions within specific substrates such as cellobiose, galactose, and glycogen, especially regarding the GalR and AGL iModulons activities. Additionally, interconnections among Acetate, TTT, and TRAP iModulons under conditions inducing low-glycolytic flux were observed, enhancing the understanding of Vn’s metabolic adaptability. iModulon analysis, mainly focusing on the Zur iModulon, illuminated the complex interplay between trace elements like Zn2+and various substrates. This analysis highlights the intricate regulatory mechanisms that control trace elements such as Zn2+, essential for bacterial metabolism and growth. Fine-tuning trace elements, as evidenced by the Zur iModulon activities, is shown to impact bacterial physiology substantially. This balancing includes influencing physiological activities29, optimizing large-scale cultivation processes53, regulating the expression of heterologous proteins54,55, affecting anaerobic metabolism56, and enhancing biochemical production57. Stress-response iModulons analysis uncovers the diverse physiological challenges of different nutrients, moving beyond the traditional focus on carbon starvation58,59and the stringent response6,60. Stress responses were identified related to protein-folding, oxidative, extracytoplasmic, and osmotic stress. This broadened understanding illuminates how bacteria adapt to nutritional changes. Notably, stresses in commonly used BHIN media for Vn studies61-63, such as oxidative and extracytoplasmic stress, are linked to reduced growth rates and protein synthesis. The high- and low-osmolarity stresses due to sodium acetate and rhamnose usage illustrates the influence of carbon source types and their specific uptake mechanisms, particularly involving TTT and TRAP iModulons. These results reveal intricate relationships between substrates and the stresses they inflict on the host. Furthermore, activating the Acetate iModulon under certain conditions reveals insights into bacterial metabolic adjustments for low glycolytic flux44. This disclosure highlights the utility of iModulon analysis in designing optimal media composition and guiding strain engineering. By analyzing iModulon activities, nutrients tailored to specific experimental objectives can be selected. A notable example is ribose, which stands out for its potential to boost protein expression in Vn. Ribose is pivotal in maintaining intracellular redox balance and promoting efficient amino acid biosynthesis64,65. Its unique ability to activate Ribosome and Chaperone iModulons—without inducing the osmolarity stress responses observed with rhamnose—suggests a potential to increase heterologous protein production. Stress-free growth conditions are crucial for understanding bacterial pathogenicity66, metabolism67, and antibiotic resistance68. Also, the detailed insights provided by gene membership and iModulon activity dynamics across various conditions are invaluable for microbial engineering, such as using the Acetate iModulon to boost growth on poor carbon sources. Materials and Methods Experiment No. 1 Bacterial strains and growth conditions Bacterial strains, both utilized and generated for this study, are listed in Table 1. This study used Vibrio natriegens ATCC 14048 (Vn) as the wild-type strain. Routine cultivation of Vn was conducted at 30°C in LBv2 media (25 g / L LB Miller broth, 200 mM NaCl, 4.2 mM KCl, and 23.14 mM MgCl2), agitated at 180 RPM, or on LBv2 agar plates (LBv2 plus 1.5% wt / vol agar). Brain Heart Infusion broth with added NaCl (BHIN: 37 g / L, 1.5% wt / vol) served as a comparative medium to LBv2. For the construction of 44 RNA-Seq data, Vn strains were cultivated at 30°C with agitation at 180 RPM. Unless stated otherwise, chemical reagents used for cell culture were sourced from Sigma-Aldrich (Burlington, MA). For additional RNA-Seq carbon samples, a range of carbon sources were added into M9Na medium (M9 minimal medium enriched with NaCl [42 mM Na2HPO4, 22 mM KH2PO4, 258.5 mM NaCl, 18.6 mM NH4Cl, 2 mM MgSO4, 0.1 mM CaCl2]) at a uniform concentration of 1.0% wt / vol, including sodium acetate, cellobiose, glycogen, and rhamnose, consistent with previous data23. Oxidative stress conditions in M9Na were induced by adding H2O2(0 mM, 0.1 mM, 0.2 mM, and 0.4 mM) to WT and VN-ALE-1 strains, the latter being an evolved strain from adaptive laboratory evolution experiments under oxidative stress69. Low- and high-temperature stress conditions were set in M9Na at 25°C and 40°C, respectively. Osmolarity stress conditions were established by varying NaCl concentrations (0, 200, 500, and 800 mM) in M9woNa (M9 medium without NaCl [42 mM Na2HPO4, 22 mM KH2PO4, 18.6 mM NH4Cl, 2 mM MgSO4, 0.1 mM CaCl2]). When an antibiotic selection of Vn was required, the antibiotics were used at specified concentrations: 50 μg / mL carbenicillin (Carb), 10 μg / mL chloramphenicol (Cm), or 360 μg / mL spectinomycin (Spec). For plasmid cloning, NEB® 10-beta Competent Escherichia coli (New England BioLabs, Ipswich, MA) was cultivated aerobically at 37°C in LB Miller broth (LB, #71753-6) with shaking at 180 RPM or on LB agar (LB with 1.5% wt / vol agar). When E. coli harbored a plasmid, appropriate antibiotics were used: 100 μg / mL ampicillin / carbenicillin, 25 μg / mL chloramphenicol, or 120 μg / mL spectinomycin. RNA extraction and transcriptomic data generation 44 new RNA-Seq datasets were produced pertinent to carbon and stress-related conditions. Transcriptomic data were derived from RNA isolated under 22 distinct conditions, encompassing a range of carbon sources (acetate, cellobiose, glycogen, and rhamnose), temperature stress (25°C and 40°C), oxidative stress (0, 0.1, 0.2, and 0.4 mM H2O2), and salt stress (0, 200, 500, and 800 mM NaCl). Each condition was replicated biologically, as detailed in FIG. 7B. Applicant followed the protocol outlined by the previous study(23) for RNA extraction and library preparation. Briefly, samples were centrifuged for 10 min at 5,000 × g at 4°C, and the supernatant was removed. RNA was isolated from the harvested cells using the Quick-RNA Fungal / Bacterial Microprep Kit (Zymo Research, Irvine, CA), according to the manufacturer's guidelines. As previously described, ribosomal RNA and genomic DNA contaminants were eliminated from 1 μg total RNA using the RiboRid method (23, 70). rRNA depletion was verified via the 4150 TapeStation System (Agilent, Santa Clara, CA) with High Sensitivity RNA ScreenTape. The rRNA-depleted RNA was then converted into libraries using the KAPA RNA HyperPrep kit (Roche, Basel, Switzerland) following the manufacturer's protocols. Library quality was assessed with the 4150 TapeStation System (Agilent) using D1000 ScreenTape, and quantification was done with the Qubit 2.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA) using the Qubit dsDNA HS Assay Kit. Subsequently, the libraries were combined and sequenced using a 100 bp single-end protocol on the Illumina NovaSeq 6000 platform at the UC San Diego IGM Genomics Center. Compilation of natPRECISE148 dataset In this disclosure, in addition to the 44 RNA-Seq datasets that were generated, 8 RNA-Seq samples from the NCBI Sequence Read Archive were also included, accessed before March 1, 2023, using the fasterq-dump software from https: / / github.com / ncbi / sra- tools. This expanded the collection to a total of 52 RNA-Seq samples. For data processing and quality control prior to Independent Component Analysis (ICA), the procedures outlined in the Modulome workflow were followed, detailed at https: / / github.com / avsastry / modulome-workflow24. Initially, raw read trimming was performed using Trim Galore, available at https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / , and FastQC, found at https: / / www.bioinformatics.babraham.ac.uk / projects / fastqc / . The quality reads were then aligned to the Vn reference genome (GCA_001456255.1)71using Bowtie72. The generated SAM files were converted into BAM files using the sam2bam function of Samtools, available at http: / / www.htslib.org / . Gene read counts in each library were computed with RSeQC73and FeatureCounts74. All quality control metrics were compiled using MultiQC at https: / / multiqc.info / 75. To maintain high data quality, datasets that failed to meet any of the following FASTQC criteria were excluded: per_base_sequence_quality, per_sequence_quality_scores, per_base_n_content, and adapter_content. Samples with fewer than 400,000 reads mapped to coding sequences were also discarded. To minimize technical variability, samples were further filtered out based on three conditions: (a) deviation from the general expression pattern, as determined by hierarchical clustering, (b) weak correlation within biological replicates (R2 less than 0.90), and (c) absence of a biological replicate, as detailed in FIGS. 7B – 7C. After thorough quality control, our final natPRECISE148 (natrigens Precision RNA-seq Expression Compendium for Independent Signal Exploration) contained 148 high-quality expression profiles: 104 from the previous natPRECISE104 database16, 23, 42 generated in this study, and 2 expression profiles76derived from public databases. The read counts were then normalized and presented as log2-transformed Transcripts per Million (log-TPM). Independent component analysis (ICA) The bioinformatics pipeline detailed in previous studies15,23,24was used. ICA was utilized to decompose the transcriptomic data matrix (X, 4515 genes by 148 conditions) into two components (M and A, representing iModulons and their activities, respectively). Independent components (ICs) through 100 iterations were calculated with the FastICA algorithm77and Scikit-Learn78and then clustered these ICs using Scikit-Learn's DBSCAN79to determine robust ICs. The optimal dimension for ICs was ascertained iteratively80, testing dimensions from 10 to 150 in steps of 10 as per15. The dimension of 140 was chosen based on the consistency between the number of robust components and the final components at this dimension (FIG. 7E). Consequently, an M matrix was derived with 64 robust iModulons and an A matrix detailing their activities across conditions. In the M matrix, each iModulon's gene weighting was determined, though most were insignificant. Significant iModulon genes were identified by setting optimal thresholds based on the D'Agostino K2 test in Scikit-Learn, as delineated previously15,23,24. The process involved iteratively removing genes with the highest absolute weight until the remaining genes approximated a normal distribution (D'Agostino K2 test statistic < 500). The genes and weights removed at this point were deemed significant, setting the iModulon thresholds. Functional characterization of iModulons The functional characterization of iModulons was conducted as described previously23. Briefly, the Pymodulon tool (https: / / github.com / SBRG / pymodulon) was utilized to examine 64 identified iModulons. The initial step involved augmenting the transcriptional regulatory network (TRN) to ascertain the functions of these iModulons, utilizing 931 TF-gene interactions previously identified23. The transcription regulators for each iModulon were inferred using Fisher's Exact Test, applying a Benjamini-Hochberg correction to control the false discovery rate (FDR) below 10–5. iModulons with significant overlap with the TRN were named according to the associated transcription factors (TFs). Furthermore, iModulon functions were deduced by gene annotation against the Kyoto Encyclopedia of Genes and Genomes (KEGG)81and the Cluster of Orthologous Groups (COG) databases, utilizing the EggNOG mapper82. KEGG modules with a statistically significant Benjamini-Hochberg corrected FDR of less than 10–2(Fisher's Exact test) were noted. Uniprot IDs were obtained through the Uniprot ID mapper83and sourced operon information from BioCyc84. Additionally, Gene Ontology (GO) annotations were acquired from AmiGO285. Each iModulon was then named based on its significantly enriched functional traits. Optical density and fluorescence measurement For flask culture experiments, the optical density (OD600) at 600 nm of bacterial cultures was determined using the BioMate™ 3S Spectrophotometer (Thermo Fisher Scientific, Waltham, MA). Additionally, the optical density for bacterial cultures was measured at 600 nm using the Infinite 200 Pro Plate Reader (Tecan, Männedorf, Switzerland). The measurements were conducted in Flat Bottom 96-well plates, with each well containing 100 µl of culture. For experiments involving Green Fluorescent Protein (GFP), fluorescence and / or OD600were assessed using the same plate reader. This was done at 485 / 515 nm excitation / emission wavelengths, utilizing Bio-One CELLSTAR μClear™ 96- well plates (Greiner, Kremsmünster, Austria). The plates were incubated at 30°C with orbital shaking. Measurements of OD600or fluorescence signals were recorded at 10 to 15-minute intervals. Relative fluorescence units (RFU) were adjusted by deducting the corresponding blanks, specifically the medium both with and without molecules. For the analysis of growth dynamics and GFP expression, parameters such as lag-time (λ), specific growth rate (μ), and GFP rate were calculated through linear regression. This analysis was facilitated by the QurvE web tool86. In some instances, the data was normalized on relative growth rates and GFP rates to make it simpler to compare with the control sample. Vector construction Oligonucleotides employed in this study are detailed in Tables 2 – 3. Oligonucleotides were sourced from Integrated DNA Technologies (Coralville, IA). PrimeSTAR GXL DNA Polymerase (Takara Bio, Shiga, Japan) and Q5 High-Fidelity DNA Polymerase (New England Biolabs) were used for high-fidelity PCR amplifications and genetic analysis. PCR products and plasmids from E. coli were purified using the DNA Clean & Concentrator (Zymo Research) and the Monarch Plasmid DNA Miniprep Kit (New England Biolabs), respectively. Plasmid was constructed using the NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs). To facilitate gene deletion, knockout (KO) plasmids were created to generate transforming DNA (tDNA), in line with methods detailed previously23. Homology arms (HAs) of about 2 Kb adjacent to each target region were amplified from Vn genomic DNA, employing specific primers listed in Table 3. These HAs, together with the antibiotic resistance cassettes (CarbRor SpecR), were integrated into the pACYC184 DNA fragment using the NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs), resulting in ptDNA-“Δstrain name” plasmids (Table 2). The sequence of each tDNA plasmid was confirmed by Sanger sequencing at Eton Bioscience (San Diego, CA). Table 1: Bacterial Strains Table 2: Plasmids Table 3: Primers Electroporation Electrocompetent Vn cells were prepared following a procedure as previously described23. Vn cells were cultured overnight in LBv2 medium at 30°C with agitation at 180 RPM. The cultures were centrifuged at 5,000 × g for 5 minutes at 4°C, washed in LBv2, and inoculated into fresh LBv2 medium at a 1:100 dilution. The culture was grown until an OD600of 0.4 was reached. Afterward, cells were centrifuged and washed twice in cold 1M sorbitol. The final pellet was resuspended in 250 μL of cold 1M sorbitol and divided into 50 μL aliquots. For electroporation, a mixture of 100–200 ng of DNA and the electrocompetent cells was transferred to a 0.2 mm Gene Pulser cuvette (Bio-Rad Laboratories) and electroporated using the Bio-Rad Gene Pulser at 800 V, 25 μF, and 1000 Ω. The cells were then recovered in 1 mL of LBv2 media for 1 hour at 30°C with shaking, plated on LBv2 agar plates with antibiotics, and incubated for at least 12 hours at room temperature or 8 hours at 30°C. Similarly, BAC DNA was electroporated into NEB 10-beta electrocompetent E. coli cells per manufacturer's instructions. The cells were recovered in 1 mL of SOC media for 1 hour at 37°C with shaking, then plated on LB agar plates with 12.5 μg / mL chloramphenicol and incubated for at least 8 hours at 37°C. DNA assembly in Saccharomyces cerevisiae The pcBAC15a shuttle vector for DNA assembly in yeast and protein expression in Vn, was developed based on a design previously described87. This vector includes a Bacterial Artificial Chromosome (BAC), S. cerevisiae replication centromere CEN, p15A origin, chloramphenicol resistance, HIS3 marker, and oriT, with construction primers listed in Table 3. For GFP expression, three vector variants were constructed: pcBAC-Ptac-GFP, pcBAC- Ptet-GFP, and pcBAC-PBAD-GFP, utilizing primers also listed in Table 3. On these vectors, GFP expression can be induced by IPTG, anhydrotetracycline, or arabinose. Additionally, Applicant developed two plasmids, pcBAC-tetR-Boost-GFP and pcBAC-araC-Boost-GFP, to investigate the function of booster genetic circuits in Vn (Table 2). The former plasmid, pcBAC-tetR-Boost-GFP, exhibits GFP expression inhibition in the presence of anhydrotetracycline, but activation in the absence of an inducer or under IPTG conditions. In contrast, the latter, pcBAC-araC-Boost-GFP, shows GFP expression inhibition with arabinose and activation in the absence of an inducer or with IPTG. Following initial experimentation, pcBAC-araC-Boost-GFP was chosen to advance the development of the Acetate Boost Module. To integrate the Acetate Boost Module in the Vn genome, two shuttle vectors containing the modules was assembled. DNA fragments with 60–70 bp homologies, as detailed in Table 2, were transformed and assembled in Saccharomyces cerevisiae VL6-48 using the LiAc / SS carrier DNA / PEG method88. Yeast clones harboring accurately assembled BACs were identified through colony PCR targeting the junctions of the constructs. The BACs were extracted using the Gentra Puregene Yeast / Bact. Kit (Qiagen) from validated yeast clones and electroporated into E. coli NEB 10-beta cells. Post-purification in E. coli, the sequence of BACs was further verified by the Whole Plasmid Sequencing service in Plasmidsaurus (Eugene, OR). Genome editing via natural transformation To delete genes or insert genetic modules into the genome precisely, the transforming DNA (tDNA) was prepared for natural competence, as previously described23. For gene deletion, tDNA was PCR-amplified using the "Δstrain name"_Left_For / Right_Rev primer pair and corresponding knockout (KO) plasmids as templates, as detailed in Tables 2- 3. The PCR products were treated with DpnI enzyme (New England Biolabs) and purified using the DNA Clean & Concentrator kit (Zymo Research). To insert the Acetate iModulon Boost Module into the genome, tDNA was amplified with Ace2_tDNA_F and Ace2_tDNA_R primer sets, using pcAceiMBoost-A-REV and pcAceiMBoost-B-REV, respectively. The amplified tDNA was further purified from agarose gels using the Zymoclean Gel DNA Recovery Kit (Zymo Research). Natural transformation assays followed the method previously described23. For gene deletion experiments, the Vn-PTrc- TfoX strain (wild-type with pTrc-tfoX plasmid) was used, while ΔArabinose-PTrc-TfoX (ΔArabinose strain with pTrc-tfoX plasmid) was used for constructing AceiMBoost strains. To activate natural competence, these strains carrying the pTrc-tfoX plasmid (inducing tfoX expression under the Trc promoter with isopropyl ß-D-1-thiogalactopyranoside (IPTG)) were grown overnight in LBv2 media supplemented with 10 μg / mL Cm and 100 μM IPTG. The cell culture's OD600was adjusted to 4.0, and 5 μL of this culture was transferred to 350 μL of competence buffer (28 g / L Instant Ocean Sea Salt) with 500 μM IPTG. 50 ng of tDNA was added to the cells and mixed gently. The mixtures were incubated at 30°C for 4 hours without shaking, then recovered in 1 mL of fresh LBv2 for 2 hours at 30°C. The cells were then plated on selective agar plates (LBv2 with 50 μg / mL Carb or 360 μg / mL Spec). Gene deletion (FIG. 9C) and target DNA integration (FIG. 11G) were verified by PCR using genomic DNA and primer set "Δstrain-name"_Val_For / Rev or AceiMBoost_Left / Right_Val_For / Rev (for left and right junction) and BoostSm_For / BoostPBAD_For (for internal region), as listed in Table 3. Table 4: Summary of iModulons. Metadata and iModulon data of natPRECISE148. This table encompasses both metadata and intricate iModulon details pertinent to natPRECISE148. It incorporates comprehensive information regarding the iModulons identified within this study. Sheet “Metadata”: This section contains the metadata of natPRECISE148, providing essential background and contextual details. Sheet “Summary”: Offers a concise overview of the 64 iModulons uncovered in this study. Sheet "X matrix": Displays the expression data (X matrix), where rows represent genes and columns correspond to samples. The values in this matrix are log-transformed TPM (Transcripts Per Million) figures. Sheet "M matrix": iModulon Structure (M matrix): This sheet elaborates on the iModulon structure (M matrix), which encompasses groups of genes that are independently modulated, as revealed through ICA. It catalogs the iModulons, arranging them in rows with genes in columns. The entries signify the influence of each gene within its respective iModulon. The top row indicates the thresholds applied to each iModulon. Sheet "A matrix": iModulon Activity (A matrix): Presents the varying activities of iModulons depending on the conditions. It aligns iModulons in rows and different conditions in columns, with the activity values reflecting the relative operation of an iModulon in a specific condition compared to a baseline condition. Sheet "TRN": Details the reconstructed transcriptional regulation network (TRN) architecture of Vn. Sheet "Annotation": A gene annotation table important for the characterization of the iModulons.

[0002]  781.3035-2898-6194 Expanding the natPRECISE transcriptomic compendium to encompass diverse nutrient and stress conditions To address the limited availability of single nutrient perturbation samples in the existing iModulon database92(FIG. 7A), the existing transcriptomic database was strategically expanded to include a wider range of nutrient and stress conditions. The existing natPRECISE104 transcriptomic database was expanded for Vibrio natriegens (Vn)94– which previously held the largest collection of single nutrient perturbation samples in minimal M9Na media (FIG. 7B) – to include a wider array of substrates and stress conditions. The expanded database, named natPRECISE148, adds various substrates, including sodium acetate, cellobiose, glycogen, and rhamnose, each at a uniform concentration of 1.0% wt / vol. The selection of these substrates for Vn growth in the M9Na minimal medium was guided by previous studies90,95. Several new stress conditions were added to the compendium. Oxidative stress was induced with varying concentrations of H2O2 in both wild-type (WT) and VN-ALE-1 strains, the latter being an evolved strain adapted to oxidative stress89. Temperature stress was generated at 25°C and 40°C, while osmolarity stress was induced by altering NaCl concentrations. For general samples not requiring specific conditions like competence, cells were harvested during the mid-exponential phase at specified optical densities (OD600). The resulting natPRECISE148 dataset encompasses 148 RNA-Seq samples (FIG. 7C), covering 20 single nutrient perturbation conditions and nine unique stress conditions. This transcriptome compendium exhibits high consistency (Pearson’s R = 0.98 between replicates), providing a robust foundation for our nutritional analytics. An ICA pipeline93was then run on this compendium, leading to the identification of 64 iModulons (FIG. 7E) that contain 1133 genes (25% of all coding sequences in Vn). These iModulons account for 73% of the variation in the expression compendium (FIG. 7F). Although principal component analysis (PCA) is effective at capturing variance within a dataset, it does not offer straightforward biological insights. However, independent component analysis (ICA) is adept at isolating consistent and reproducible regulatory signals that are rich in genetic, biochemical, and biological contexts. Therefore, the 73% of variation explained by iModulons is directly linked to biological significance. Of these 64 iModulons, 22 were categorized as 'regulatory,' showing statistical overlap with known regulons and alignment with those in other Vibrio species (FIG. 8A). Further classification was achieved based on iModulon recall and regulon recall metrics (FIG. 7G). To determine the relationship between iModulons and these known regulons, two metrics were employed: iModulon Recall (MR) and Regulon Recall (RR). Regulon Recall is calculated by dividing the number of genes common to both an iModulon and its corresponding regulon by the total number of genes in that regulon. Conversely, iModulon Recall is determined by dividing the number of shared genes by the total number of genes in the iModulon. This plot illustrates the relationship between iModulon Recall (MR) and Regulon Recall (RR) across four categories of the 64 iModulons: well-matched (MR and RR ≥ 0.6), regulon subset (MR ≥ 0.6 and RR < 0.6), regulon discovery (MR < 0.6 and RR ≥ 0.6), and poorly matched (both MR and RR < 0.6). Circle sizes on the plot represent the number of genes in each iModulon, while shades categorize the general type of iModulon. Note that these categories are approximate and sensitive to the threshold applied. An additional 31 iModulons were identified as 'functional,' determined through Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment or functional annotations. The remaining 11 were labeled as 'uncharacterized,' indicating unknown gene functions and transcriptional regulation. Each iModulon was also assigned to one of 11 functional groups, providing a systems-level perspective (FIG. 1B and FIG. 7H). Experiment No. 2 Nutrient Supplement and Concentration Selection and Use Many approaches to optimizing cell growth in media relies on properly providing nutrient supplements to the media during cell culturing. However, given the diverse range of genes and reactions that genes have to different nutrient supplements, it can be difficult to identify or determine nutrient supplements that have a growth supplementing process across a wide array of genes. These problems can be compounded because different nutrient supplements may improve cell growth for some genes and reduce or inhibit cell growth for other genes. Many methods to address this problem involve using trial and error to identify nutrient supplements that work in different situations. However, such methods can take an exorbitant amount of time and processing resources given the large number of genes and differing reactions to nutrient supplements by the different genes. Moreover, such methods struggle to identify the functional roles of uncharacterized genes and their relationships to growth conditions. Accordingly, there is a need for an improved method of optimizing nutrient supplements to add to media to improve microbial growth. The systems and methods described herein overcome these technical hurdles by using independent component analysis (ICA) and iModulons. For example, a computer can obtain transcriptomic data from microbial samples that were cultured in media. The transcriptomic data may include data indicating genetic expressions in different conditions that are different only with regards to the type and / or concentration of nutrient supplements that are added for the culturing. The computer can use independent component analysis (ICA) on the transcriptomic data to determine or identify one or more iModulons containing at least one characterized gene and at least one uncharacterized gene. By using ICA to identify and isolate iModulons, the system and method can identify regulatory relationships of iModulons without prior knowledge of each gene function and identify functional associations between genes with a known and unknown function. Using ICA for iModulon identification has multiple technical advantages over other methods (e.g., pattern matching). For example, using ICA can substantially reduce the processing resources and latency required to identify iModulons and / or respond to requests for iModulons compared with systems that use other methods, such as pattern matching between groups of genes, particularly given that there are millions of permutations or combinations of genes that may or may not be involved in metabolic processes. For example, instead of analyzing thousands or millions of individual gene correlations, ICA can facilitate the compression of expression data into a smaller set of meaningful regulatory components. This can make the computational process more efficient and help identify the most significant gene relationships in the large dataset that includes different characteristics of genetic data. The computer can calculate a nutrient score for each of the conditions. The computer can do so using the outputs of the ICA analysis. For example, when determining the iModulons, the computer can generate a matrix (e.g., a mixing matrix) that contains columns that correspond to different iModulons and rows that correspond to the conditions. The values at the intersections of the matrix can indicate activity levels of the iModulons when cultured in the different conditions. The computer can determine a nutrient score for each condition by aggregating the activity levels of the iModulons under the condition. In some cases, the computer can determine the nutrient scores as weighted sums of iModulon activity levels in which the activity levels of the iModulons are weighted based on the growth rate of the iModulons across conditions. In doing so, the computer can determine a relationship between each of the conditions (e.g., type of nutrient supplement and concentration of nutrient supplement) in which a higher positive nutrient score indicates a higher growth rate and a negative nutrient score indicates a reduced or inhibited growth rate (e.g., the magnitude of the negative nutrient score can indicate or correspond to the amount in which growth is reduced or inhibited when the nutrient supplement is added). The computer can select a condition based on the nutrient score of the condition. For example, the computer can rank the conditions based on the nutrient scores determined for the conditions. The computer can compare the rankings and identify the highest ranked condition. The computer can select the identified condition and generate a recommendation identifying the type of nutrient supplement and concentration of the condition. The computer can transmit the recommendation to a client device. A user can use the recommended type and concentration of the nutrient supplement to supplement a growth medium for cell growth. For example, the user can view the recommended type and concentration of the nutrient supplement on a user interface of a display of the computer. The user can retrieve the type and concentration of the nutrient supplement and supplement a cell culturing container with the type and concentration of the nutrient supplement. Operating in this manner can substantially improve the cell growth process and do so with relatively low computing latency and processing resources by using and leveraging ICA. For example, FIG. 12 is an illustration of an example system 1200 for nutrient supplement selection, in accordance with implementations. In brief overview, the system 1200 can include a data processing system 1202, a remote data source 1218, and a client device 1220. The client device 1220 can be or include a computing device, a server, a mobile device, a table, a personal computer, or any other type of computing device. The data processing system 1202 can receive or retrieve genetic transcriptomic data from the remote data source 1218. The data processing system 1202 can receive additional transcriptomic data from the client device 122 that was generated using nutrient supplementation data that includes cell expression levels for cells that were grown in media in which varying types of nutrient supplements and / or varying amounts of the nutrient supplements were added. The data processing system can combine the two sets of transcriptomic data to create an aggregated transcriptomic data set. The data processing system 1202 can apply independent component analysis on the aggregated transcriptomic data set or at least the additional transcriptomic data of the aggregated transcriptomic data set to identify one or more iModulons. Each iModulon can include one or more genes with a known function and one or more genes with an unknown function. The data processing system can calculate a nutrient score for the different conditions of the nutrient supplementation data. The data processing system 1202 can select a condition (e.g., a type and concentration of a nutrient supplement) based at least on the condition corresponding to a positive nutrient score. The data processing system 1200 can transmit a message to the client device 1220 identifying the selected condition. The system 1200 may include more, fewer, or different components than shown in FIG. 12. The system 1200 may be used to perform any of the process described herein and is not limited to the operations described with respect to FIG. 12. The remote data source 1218 can be or include one or more data sources that store genetic data of individual or different organisms and / or external data. The remote data source 1218 can be a database, a computer, a laptop, a server, or any other type of device that can store data. For example, the remote data source 1218 can be or include a database (e.g., a relational or graph database) that stores a list of previously determined iModulons (e.g., the PRECISE-1K dataset). The remote data source 1218 can store data indicating levels of expression different groups of genes exhibited under different conditions (e.g., culture conditions). Examples of the conditions can include varying the pH level, changing the ions, changing substrates, and / or putting other cells in the media containing the genes. Each level can be represented as numerical, alphabetical, and / or alphanumerical value. For each condition, the remote data source 1218 can store a vector identifying the level of expression of individual genes cultured under the condition. The remote data source 1218 can store such vectors for any number of groups of genes for an individual condition. The remote source 1218 can generate and / or store such vectors using genes from a single organism or from multiple organisms. In some cases, the remote data source 1218 can store the expression data separately and not in vector form. In some cases, the remote data source 1218 can store data or metadata about individual genes. For example, the remote data source 1218 can store data indicating the functionality (e.g., molecular functionality and / or genetic functionality) of individual genes and / or indications of whether the genes have a known function. The remote data source 1218 can store any amount and / or type of data regarding genes. In one example, the remote data source 1218 can include one or more data sources (e.g., as one or more computing devices at the same or different physical locations) for gene expression data in transcriptomic datasets. The remote data source 1218 can store databases, such as the Gene Expression Omnibus (GE), ArrayExpress, and / or the Sequence Read Archive (SR). The different databases can store gene expression data and / or can include RNA sequencing data showing the number of reads mapped to each gene, or microarray data with hybridization intensity values, along with metadata about experimental conditions, sample characteristics, and / or technical parameters. The datasets may additionally or instead contain quality metrics, alignment statistics, and / or information about splice variants or alternative transcripts, depending on the experimental platform and processing pipeline used. The data can be generated from individual experiments or from multiple experiments that involve measuring the expression levels of genes in one or more conditions. The data processing system 1202 may comprise one or more processors that are configured to determine iModulons. The data processing system 1202 may comprise a network interface, a processor, and / or memory. The data processing system 1202 may communicate with the remote data source 1218 and / or the client device 1220 via the network interface, which may be or include an antenna or other network device that enables communication across a network and / or with other devices. The processor may be or include an ASIC, one or more FPGAs, a DSP, circuits containing one or more processing components, circuitry for supporting a microprocessor, a group of processing components, or other suitable electronic processing components. In some embodiments, the processor may execute computer code or modules (e.g., executable code, object code, source code, script code, machine code, etc.) stored in memory to facilitate the activities described herein. The memory may be any volatile or non-volatile computer-readable storage medium capable of storing data or computer code. The memory may include a data collector 1204, an iModulon generator 1206, a score generator 1208, a nutrient selector 1210, and / or a data repository 1212, in some embodiments. The components 1204-1212 may operate to identify iModulons containing different sets of genes that correspond to different functions. The iModulons may have been cultured in media containing varying or different nutrient supplements. The components 1204-1212 can calculate nutrient scores for different nutrient supplements based on stress- related activity by the iModulons. The components 1204-1212 can select a condition corresponding to a type and concentration of a nutrient supplement based on the nutrient score for the condition. In doing so, the components 1204-1212 can operate to identify a nutrient supplement and / or concentration that optimally grows the iModulons that can be used for future cell culturing. For example, the data collector 1204 may comprise programmable instructions that, upon execution, cause the data processing system 1202 to communicate with the remote data source 1218, the client device 1220, and / or any other computing devices. The data collector 1204 may be or include an application programming interface (API) that facilitates communication between the data processing system 1202 and other computing devices. The communicator 1204 may communicate with the remote data source 1218, the client device, and / or any other computing device across a network 1201 (e.g., a wired or wireless network). The data collector 1204 can establish connections with the remote data source 1218 and / or the client device 1220. The data collector 1204 can establish the connections over the network 1201. To do so, the data collector 1204 can communicate with the remote data source 1218 and / or the client device 1220 across the network 1201. In one example, the data collector 1204 can transmit syn packets to the remote data source 1218 and establish the connection using a TLS handshaking protocol. The data collector 1204 can use any handshaking protocol to establish a connection with the remote data source 1218. The data collector 1204 can collect a list of previously determined iModulons from the remote data source 1218. For example, the data collector 1204 can send a request or otherwise crawl the remote data source 1218 for a dataset identifying a list of iModulons (e.g., the PRECISE-1K dataset). In response, the data remote data source 1218 can transmit the list of iModulons to the data processing system 1202. The data collector 1204 can collect or request any type of data or metadata regarding genes from the remote data source 1218. Additionally, the data collector 1204 can collect gene expression data from the client device 1220. For example, one or more individuals may conduct experiments evaluating the gene expressions of genes that were cultured in separate media. Each medium may have been supplemented with a different nutrient supplement and / or a different concentration of a nutrient supplement (e.g., correspond to a different condition). The media may have been divided or placed into separate containers (e.g., each container can correspond to a different experiment) each containing one of 63 unique nutrient supplements and / or varying concentrations of such nutrient supplements (e.g., 2-10mM). The medium may have been M9 glucose + NH4Cl. The media may be the same or different between each container. Each nutrient supplement may have been added to the media at varying concentrations between the containers (e.g., 2mM of a first nutrient supplement may have been added to a first container, 3mM of the first nutrient supplement may have been added to a second container, 3mM of the first nutrient supplement may have been added to a third container, and so forth). Accordingly, each nutrient supplement may correspond to multiple containers. One or more nutrient supplements may have been added to each container under each condition. The individuals (e.g., experimenters) may culture bacteria or cells in the separate containers. For example, bacteria or cells of a particular strain (e.g., E. coli MG1655a) may be added or grown in each of the containers. The containers can be incubated, such as for aeration. Varying nutrient supplements and / or varying amounts or concentrations of nutrient supplements can be added to the containers. The bacteria or cells in the containers can grow in the conditions (e.g., the type of nutrient supplement added and the amount or concentration of the nutrient supplement) over time until reaching a state or phase (e.g., mid-log) at which the bacteria or cells can be harvested to capture their transcriptional state. The individuals may extract or determine transcriptomic data from each of the containers. For example, the individuals may harvest the cells from the different containers, such as during the mid-log phase of growth. The individuals may use the client device 120 to generate transcriptomic data for the extracted cells. Examples of transcriptomic data that can be generated can include gene expression measurements (e.g., raw read counts per gene, gene expression values, etc.), pathway-level data (e.g., expression patterns of metabolic pathways, nutrient transport system activation, stress response pathway activity, energy metabolism gene clusters, cell envelope modification pathways, etc.), regulatory network information (e.g., transcription factor activity profiles, operon structure and regulation, etc.), technical and quality metrics (e.g., RNA quality scores, sequencing depth per sample, etc.), time-course data (e.g., expression changes over time after supplementation, adaptation responses, sequential gene activation patterns, temporal regulatory network dynamics), and / or condition- specific transcriptional features (e.g., novel transcripts, alternative splice variants, antisense transcription, etc.). The computer or the client device 120 can label the individual genes with the corresponding transcriptomic data determined for the genes. The client device 120 can transmit the generated transcriptomic data to the data processing system 102, which can be received by the data collector 1204. The iModulon generator 1206 may comprise programmable instructions that, upon execution, cause the data processing system 1202 to generate or identify iModulons from transcriptomic data. The iModulon generator 1206 can identify iModulons using the transcriptomic data the iModulon generator 1206 received from the client device 120 generated based on conditions relating to nutrient supplement type and / or nutrient supplement concentration. The iModulon generator 1206 can identify the expression levels received from the client device 1220. The iModulon generator 1206 can identify the expression levels of the genes under the different conditions (e.g., as grown with the different types of nutrient supplements and / or different concentrations of nutrient supplements) and group the expression levels together. In grouping the expression levels of the genes, for example, the iModulon generator 1206 can generate a separate vector for each condition from the expression levels of the genes that were cultured under the condition. Each vector can include a numerical value at different index values that represents an expression level of an individual gene under the condition for which the iModulon generator 1206 is generating the vector. Each vector can be a gene expression profile for a given condition. In some cases, instead of aggregating the gene expression level data into separate vectors, the iModulon generator 1206 can receive such vectors from the client device 1220 when the client device 1220 provides the gene expression data to the data processing system 1202. The iModulon generator 1206 can store the vectors in the data repository 1212 as expression profiles 1214. Accordingly, the iModulon generator 1206 can generate and / or store vectors with different numerical values that represent the expression levels of genes. The iModulon generator 1206 can use the expression profiles 1214 to identify or determine iModulons. The iModulon generator 1206 can identify or determine the iModulons using independent component analysis (ICA). In doing so, the iModulon generator 1206 can identify individual iModulons that each include at least one gene with a known function and at least one gene with an unknown function. Such iModulons can be transferred into bacteria or cells of organisms to give the bacteria or cells a new functionality, in some cases. For example, to identify the iModulons, the iModulon generator 1206 can apply ICA to the expression profile vectors. The iModulon generator 1206 can do so using the FastICA algorithm, in some cases. For instance, the iModulon generator 1206 can apply ICA to the expression profile vectors by first organizing the vectors into a data matrix X, where each row represents a condition and each column represents a gene, or vice versa. This matrix can contain the expression levels for all genes across all experimental conditions. In some cases, prior to generating the matrix X, the iModulon generator 1206 can normalize the values. The iModulon generator 1206 can do so, for example, by performing a z-score normalization and / or log2 transformation to the expression values. By doing so, the iModulon generator 1206 can reduce biases from technical variations across experiments and / or make data comparable across different experiments and conditions. The iModulon generator 1206 can decompose the matrix X into two matrices: a mixing matrix A and a source matrix S. The iModulon generator 1206 can decompose the matrix X such that X is approximately equal to AS. The mixing matrix A can contain the independent components that represent distinct regulatory patterns, while S can contain the weights describing how each gene contributes to these patterns. This decomposition can identify statistically independent patterns of gene expression variation across conditions. During decomposition, the iModulon generator 1206 can maximize the statistical independence between the components. The iModulon generator 1206 can do so using techniques such as maximizing non-Gaussianity or minimizing mutual information. The resulting independent components can represent distinct regulatory modules, with each component corresponding to a potential iModulon (e.g., a candidate iModulon). The iModulon generator 1206 can identify genes that significantly contribute to the different respective components. The iModulon generator 1206 can do so, for example, by identifying the weight values in the matrix S and comparing the weight values to a threshold. The iModulon generator 1206 can determine which genes substantially contribute to each component based on the genes corresponding to weights exceeding the threshold. The iModulon generator 1206 can group genes with weights exceeding this threshold as potential members of an iModulon, such as by assigning a common or identical identifier to each gene of the respective groups. The iModulon generator 1206 can perform a filtering technique to remove or discard potential iModulons from memory or otherwise consideration as being an iModulon. The iModulon generator 1206 can do so, for example, for each potential iModulon, by analyzing data or metadata for each gene grouped into the iModulon. The iModulon generator 1206 can identify data indicating whether the gene has a known function or not. The iModulon generator 1206 may identify such data by retrieving or querying the remote data source 1218 for the data. The iModulon generator 1206 can determine whether at least one gene of the iModulon has a known function and at least one gene of the iModulon has an unknown function (e.g., has an unknown function or has no known function). The iModulon generator 1206 can determine whether this condition is true for each potential iModulon and discard (e.g., remove from memory or otherwise from consideration as an iModulon) any potential iModulon for which the iModulon generator 1206 determines the condition is not true. In another example, the iModulon generator 1206 may remove any potential iModulons that only include a single gene or a number of genes below a threshold. For example, the iModulon can instantiate or generate a counter for each iModulon. The iModulon generator 1206 can increment each counter for each gene in the iModulon. The iModulon generator 1206 can compare the counts of the counters to a threshold (e.g., 1.5). Responsive to determining the number of genes in a potential iModulon exceeds the threshold, the iModulon generator 1206 can remove or discard the potential iModulon. The iModulon generator 1206 can periodically use such filters, and any number of other filters, to remove or discard potential iModulons. In some cases, the iModulon generator 1206 can validate the statistical significance of each potential iModulon after the filtering using various metrics. For example, the iModulon generator 1206 can calculate internal consistency metrics including mean pairwise correlation scores, explained variance ratios, and internal clustering coefficients for gene expression patterns within each potential iModulon. The iModulon generator 1206 can compare such metrics for each potential iModulon to one or more thresholds (e.g., a threshold specific to the metric of the comparison or the same threshold for each metric) to determine whether the threshold is satisfied (e.g., exceeded or beneath, depending on the metric and / or the configuration of the iModulon generator 1206). In one example, the iModulon generator 1206 can determine whether each potential iModulon has a mutual information score less than 0.1, a cross-correlation coefficient less than 0.15, and / or a principal angle greater than 120 degrees. The iModulon generator 1206 can remove or discard any potential iModulons where at least one of these criteria is not met. In another example, the iModulon generator 1206 can perform a robustness validation using bootstrap analysis to remove or discard any potential iModulons with a minimum 95% confidence threshold and / or a cross-validation stability score below 0.8. Doing so can ensure reproducibility of the identified iModulons. The iModulon generator 1206 can identify or output the filtered and / or validated iModulons. Each filtered and / or validated iModulon can be represented as a set of genes with their corresponding weights from the ICA decomposition. The weights can indicate the strength and direction of each gene's contribution to the iModulon's function, providing insight into the relative importance of both known and unknown genes within each functional module. The iModulon generator 1206 can store the validated iModulons and / or the corresponding in the data repository 1212 as iModulons 1216. In some cases, the iModulon generator 1206 can store an association between the genes of the iModulon and / or any other data or metadata about the genes and / or iModulon in memory. In doing so, the iModulon generator 1206 can map (e.g., assign) the function (e.g., molecular function or genetic function) of the genes with the known function of the iModulon to the grouped genes of the iModulon. In some embodiments, in one example, the data processing system can determine (e.g., using a mapping table of genetic functions to metabolic functions) a metabolic function for the iModulon based on the genetic functions of the genes with the known genetic functions of the iModulon and assign the metabolic function to the iModulon. The data processing system can repeat this process for any number of potential iModulons. In some cases, the iModulon generator 1206 can transmit the identified iModulons 1216 and / or any data (e.g., the functionality of the genes of the iModulon) to the client device 1220. The iModulon generator 1206 can add any identified iModulons to the iModulon dataset (e.g., PRECISE-1K) stored in and / or retrieved from the remote data source, such as to create a new dataset (e.g., PRECISE-NP881). The iModulon generator 1206 can categorize the iModulons into functional groups. For example, the iModulon generator 1206 can categorize the iModulons into catabolism, stress responses, and translation. The iModulons categorized into the catabolism group can indicate that the genes of such iModulons are involved in breaking down nutrients. The iModulons grouped into the stress responses group can indicate that the genes of such iModulons are activated during various environmental stresses. The iModulons grouped into the translation group can indicate that the genes of such iModulons are involved in protein synthesis. To categorize the iModulons into these groups, the iModulon generator 1206 can identify the known functions of the genes within the iModulons and / or the mapping of the known functions to the iModulons. The known functions may each correspond to or be mapped (e.g., in a stored table) to one of the categories (e.g., identifications of the categories in memory). The iModulon generator 1206 can categorize the iModulons into functional groups based on the mappings of the known functions of the genes within the iModulons to identifications of the respective categories. The iModulon generator 1206 can assign labels to each of the iModulons indicating the determined categories of the respective iModulons. The score generator 1208 may comprise programmable instructions that, upon execution, cause the data processing system 1202 to generate scores (e.g., nutrient scores) for the different conditions. The score generator 1208 can do so using the mixing matrix A. For example, the mixing matrix A may include separate columns representing the different iModulons and rows representing the conditions, or vice versa. The intersections between the columns and rows can indicate a level of activity (e.g., strength of a regulatory pattern) in the respective conditions. The levels of activity can each be zero or a positive or negative numerical value that indicates the magnitude of the level of activity of the iModulons in the different conditions. A positive value can indicate activation, and a negative value can indicate repression. The score generator 1208 may identify iModulons that were categorized into the stress responses category by the iModulon generator 1206. The score generator 1208 can identify the columns or rows of the identified iModulons. The score generator 1208 can determine a weight for each of the iModulons. The score generator 1208 can do so based on correlation coefficients between iModulon activity and growth rate across all conditions indicating the contribution of each iModulon to growth. To do so, the score generator 1208 can identify the levels of activity of the iModulons in the different conditions from the mixing matrix A. The score generator 1208 can determine the weight for each of the iModulons (e.g., the iModulons categorized into the stress category or each of the determined iModulons) using the Pearson correlation coefficient, R, between activity vectors across each of N conditions (e.g., samples extracted from the conditions) (aj1, aj2, ..., ajN), wherein aij is the activity level of iModulon j under condition I, and a growth rate vector for the same N conditions (μ1, μ2, ..., μN). The score generator 1208 can identify the growth rates of the conditions from the transcriptomic data provided by the client device 1220. The score generator 1208 can apply the following formula to the vectors to determine the weight for each iModulon j: wj = R = Σ((aji - āj)(μi - μ)̄) / square root(Σ(aji - āj)² × Σ(μi - μ)̄²), where i is the condition. In some cases, the score generator 1208 can normalize the weights, such as to ensure comparable scales across iModulons. The score generator 1208 can normalize each weight by dividing the value of the weight by a sum of the absolute value of each of the weights. In some cases, when assigning the weights, growth-promoting iModulons, such as those associated with protein biosynthesis, received positive weights, while stress-related iModulons linked to resource allocation for stress management (e.g., CRP, NtrC, RpoS) were assigned negative weights. The score generator 1208 can determine a score (e.g., a nutrient score) for each nutrient condition (e.g., each combination or permutation of nutrient type and / or concentration). The score generator 1208 can do so, for example, using the following formula: Nutrient scorei= ΣWj*Aijwhere i is a particular condition, j is a particular iModulon, Wjis the weight for the iModulon j, and Aij is the iModulon activity level in the mixing matrix A for the iModulon j for condition i. The nutrient score can indicate how nutrient perturbations influence the balance between growth and stress responses. The score generator 1208 can assign the scores to the different nutrients by storing a value indicating the nutrient score for each nutrient condition in memory. The score generator 1208 can determine and / or assign positive and negative nutrient scores to the conditions. A positive nutrient score can indicate a growth-promoting condition and that growth-related iModulons may be more active and stress responses iModulons may be less active when grown under the condition. A negative nutrient score can indicate stress responses dominate and that stress-related iModulons may be more active and growth- promoting processes may be reduced when grown under the condition. The nutrient selector 1210 may comprise programmable instructions that, upon execution, cause the data processing system 1202 to select nutrient supplements or conditions corresponding to the nutrient supplements and the concentration of the nutrient supplements. The nutrient selector 1210 can select the conditions based on the nutrient scores determined for the conditions. For example, the nutrient selector 1210 can rank the conditions in descending or ascending order based on the nutrient scores determined for the conditions. The nutrient selector 1210 can select one or a defined number of the highest ranked conditions based on the ranking. In another example, the nutrient selector 1210 can filter out any conditions associated with a negative nutrient score and then perform the ranking. By doing so, the nutrient selector 1210 can ensure that the nutrient selector 1210 does not select a condition with a negative nutrient score that may reduce or harm growth-promoting processes. The nutrient selector 1210 can transmit an indication of the selected condition (e.g., an identification of the nutrient supplement and / or concentration of the nutrient supplement) to the client device 102. For example, the data processing system 102 can initiate the process of selecting a condition based on a nutrient score of the condition in response to a request from the client device 102 (e.g., a request containing the transcriptomic data of the nutrient- based conditions). The data processing system 102 can select the condition based on the nutrient score for the condition in response to the request. The data processing system can transmit an identification of the selected condition back to the client device 102 in a message. The identification can be a recommendation to use the nutrient type and concentration of the condition for culturing in the future to improve the cell culturing process. In some cases, the individual at the client device 1220 can use the recommended condition in another cell growth operation. For example, the client device 1220 can present the selected condition on a user interface. The individual can identify the selected condition and use the selected condition to supplement a growth medium (e.g., the same type of medium from which the transcriptome data was generated, such as M9 glucose + NH4Cl). The individual can culture cells in the growth medium under the selected condition. Using the selected condition for the cell culturing can promote growth and minimize stress on the cells. In some cases, the individual and the data processing system 102 can operate to improve the growth media by iteratively determining types and types of nutrient supplements to add to the same growth media to improve the growth media (e.g., improve the capability of the growth media to facilitate cell growth. For example, the data processing system 102 and determine iModulons grown in a particular growth media and first nutrient scores for conditions in which the iModulons were grown as described above. Based on the first nutrient scores, the data processing system can select a type and concentration of a nutrient supplement to add to the cell growth media to improve the growth stimulation properties of the growth media. Subsequently, the data processing system and individual can repeat the same process using the growth media including the added selected nutrient type and concentration by adding new cells (e.g., a second population or plurality of cells) to the updated growth media. The data processing system can determine one or more iModulons from the new cells that were cultured in the updated growth media. The data processing system can determine second nutrient scores for conditions in which the new cells were cultured in the updated growth media and select a second condition indicating a second type of nutrient and concentration of the second type of nutrient. The individual can add the second type of nutrient in the selected concentration to another growth media with the initially selected nutrient supplement such that the growth media has the two selected types of nutrient supplements with the initial growth media. The individual and data processing system can iteratively repeat this process any number of times to improve the growth media for cell growth. The data processing system can stop the iterative process responsive to a criterion being satisfied. For example, the data processing system can determine a criterion is satisfied if the nutrient scores that the data processing system generates for an iteration are all below a defined threshold (e.g., below zero or below another threshold). Responsive to such a determination, the data processing system can generate an alert to the client device 120 indicating to stop the iterative process. In some cases, instead of stopping the process, the individual may attempt to continue improving the growth media by repeating the process with another set of nutrient supplements. The data processing system and individual can continually improve the growth media using different nutrient supplements to improve growth rates of cells in the growth media. FIG. 13 is an illustration of an example method 1300 for nutrient selection and media supplementation, in accordance with implementations. One or more, or all, of the operations of the method 1300 can be performed by one or more systems or components depicted in FIG. 12 including, for example, the data processing system 1202. Performance of the method 1300 can facilitate the selection and supplementation of media with nutrient supplements to improve cell growth. At operation 1302, transcriptomic data can be obtained. The transcriptomic data can include gene expression data and / or metadata regarding the known functionality or known lack of functionality in genes. The transcriptomic data can be created or generated when a user cultures microbial samples (e.g., cells from microorganisms or cells from other types of organisms) in different containers of media. The media may be the same or different between the containers, but the user may supplement the different containers with different types or concentrations of nutrient supplements in different conditions. The user or a computer can extract transcriptomic data from the cultured cells of each condition. The user can transmit the transcriptomic data to the data processing system. At operation 1304, the data processing system can determine one or more iModulons from the transcriptomic data. iModulons can be or include groups of genes that each include at least one gene with an unknown function (e.g., an unknown cellular function or unknown molecular function) and at least one gene with a known function (e.g., a known cellular function or a known molecular function). Each iModulon can correspond with a particular function (e.g., metabolic function) to which the different genes of the iModulon contribute. The data processing system can determine the iModulons using independent component analysis (ICA) techniques. For example, the data processing system can identify transcriptomic data generated from culturing the genes in the different conditions of varying types of nutrients and / or concentrations of the nutrients. The transcriptomic data can include individual expression profiles (e.g., vectors with expression levels at different index values) for different conditions. The data processing system can organize the expression profiles into a matrix X, where different rows represent different conditions and different columns represent different genes, or vice versa. The data processing system can decompose the matrix into a mixing matrix and a source matrix. The mixing matrix can contain independent components representing different regulatory patterns (e.g., activity of iModulons in the different conditions), while the source matrix can contain weights indicating an impact of genes’ contribution to the patterns. The data processing system can perform the decomposition to maximize the statistical independence between components, such as by maximizing non- Gaussianity or minimizing mutual information, with each component representing a potential iModulon. For example, the data processing system can decompose the matrix X into a mixing matrix A and a source matrix S. The matrix X can be or include an m x n matrix with m rows in which m corresponds to experimental conditions, n corresponds to genes, and each entry Xijis an expression level of gene j in condition i. Matrix A can be or include an m x k matrix in which k is the number of independent components, each column can be a regulatory pattern across conditions, and each entry Aij is a strength of component j in condition i. Matrix S can be or include a k x n matrix in which each row corresponds to weights for genes in a component and each entry Sijis a contribution of gene j to component i. The data processing system can identify significant genes by comparing the weight values in the source matrix to a threshold. The data processing system can group genes that exceed the threshold (e.g., exceed the threshold for a particular regulatory pattern or component) into potential iModulons. The data processing system can then apply filtering techniques to the potential iModulons, removing potential iModulons that don't contain genes with both known and unknown functions and / or potential iModulons with genes below a threshold number. The data processing system can validate the remaining iModulons using metrics such as mean pairwise correlation scores, explained variance ratios, and internal clustering coefficients. The data processing system can compare the metrics to one or more thresholds. The data processing system can remove or discard any potential iModulons with at least one metric that does not satisfy a threshold based on the comparisons. The data processing system can determine, identify, or output the remaining iModulons after the filtering and / or the validation. The output iModulons can be or include sets of genes with their corresponding weights from the ICA decomposition. The weights can indicate the strength and / or direction of each gene's contribution to the iModulon's function. The validated iModulons can be stored in memory. The data processing system can transmit the output iModulons to a client device automatically or in response to a request from the client device. In some cases, the data processing system can determine the iModulons in response to a request from the client device and automatically transmit the iModulons to the client device in response to determining the iModulons. In a non-limiting example, the data processing system can analyze gene expression data from E. coli grown under various nutrient conditions. For instance, the data processing system can identify or generate expression profiles from cells grown in media with and without leucine. The data processing system can create or generate a matrix X where each row represents a different growth condition (e.g., rich media, minimal media, stress conditions, etc.) and each column represents individual genes, such as leuA, which is a known leucine synthesis gene, and yibT, which has an unknown function or does not have a known function. For instance, one or more studies can include of 100 experimental conditions across 1000 genes. The data processing system can collect expression data from these studies and generate the matrix X to include 100 rows and 1000 columns. The data processing system can decompose the matrix X into a mixing matrix A and a source matrix S. The matrix X can be or include an m x n matrix with m rows in which m corresponds to experimental conditions, n corresponds to genes, and each entry Xij is an expression level of gene j in condition i. A can be or include an m x k matrix in which k is the number of independent components, each column can be a regulatory pattern across conditions, and each entry Aijis a strength of component j in condition i. Matrix S can be or include a k x n matrix in which each row corresponds to weights for genes in a component and each entry Sij is a contribution of gene j to component i. The data processing system can identify one or more patterns from the decomposed matrices. For instance, the matrix X may include profiles for the following three conditions: growth in leucine-rich media, growth in leucine-depleted media, and growth under heat stress. The data processing system can identify a component from the matrix A and / or the matrix S where leuA has a weight of .85 (e.g., a strong positive above a threshold of .7), yibT has a weight of .78 (e.g., a strong positive above the threshold of .7), and hspB has a weight of .12 (e.g., a weak correlation below the threshold of .7). Based on the weights’ relationship with the threshold of .7 (e.g., whether the weights are above or below the threshold of .7), the data processing system can determine leuA and yibT are a part of the same potential iModulon (e.g., an iModulon involved in leucine metabolism), while hspB may belong to a different potential iModulon. The decomposition can reveal that leuA and yibT are strongly co-expressed under leucine-related conditions, while hspB (heat shock protein) shows minimal correlation. The data processing system can use these weights to establish that leuA and yibT likely belong to the same iModulon involved in leucine metabolism, while hspB may belong to a different regulatory network. The data processing system can validate the potential iModulon by confirming the iModulon has at least a threshold number of genes with a known function (e.g., leuA) and at least a threshold number of genes with an unknown function or not a known function (e.g., yibT). The data processing system can additionally verify the statistical significance through correlation scores. If the genes show strong correlation (greater than 0.8) and meet other statistical thresholds (e.g., a mutual information score less than 0.1, a cross-correlation coefficient less than 0.15, and / or a principal angle greater than 60 degrees), the data processing system can determine the potential iModulon is an iModulon. At operation 1306, the data processing system can calculate a nutrient score for each condition based on which the data processing system determined the one or more iModulons. The data processing system can do so using the mixing matrix A, which was calculated during the ICA process. For example, the mixing matrix A may include separate columns representing the different iModulons and rows representing the conditions, or vice versa. The intersections between the columns and rows can indicate a level of activity (e.g., strength of a regulatory pattern) in the respective conditions. The levels of activity can each be zero or a positive or negative numerical value that indicates the magnitude of the level of activity of the iModulons in the different conditions. The data processing system can determine a weight for each of the iModulons. The data processing system can do so based on correlation coefficients between iModulon activity and growth rate across all conditions indicating the contribution of each iModulon to growth. To do so, the data processing system can identify the levels of activity of the iModulons in the different conditions from the mixing matrix A. The data processing system can determine the weight for each of the iModulons (e.g., the iModulons categorized into the stress category or each of the determined iModulons) using the Pearson correlation coefficient, R, between activity vectors across each of N conditions (e.g., samples extracted from the conditions) (aj1, aj2, ..., ajN), wherein aij is the activity level of iModulon j under condition i, and a growth rate vector for the same N conditions (μ1, μ2, ..., μN). The data processing system can identify the growth rates of the conditions from the transcriptomic data. The data processing system can apply the following formula to the vectors to determine the weight for each iModulon j: wj = R = Σ((aji - āj)(μi - μ)̄) / square root(Σ(aji - āj)² × Σ(μi - μ)̄²), where i is the condition. In some cases, the data processing system can normalize the weights, such as to ensure comparable scales across iModulons. The data processing system can normalize each weight by dividing the value of the weight by a sum of the absolute value of each of the weights. In some cases, when assigning the weights, growth-promoting iModulons, such as those associated with protein biosynthesis, may receive positive weights, while stress-related iModulons linked to resource allocation for stress management (e.g., CRP, NtrC, RpoS) may be assigned negative weights. The data processing system can determine a score (e.g., a nutrient score) for each condition (e.g., each combination or permutation of nutrient type and / or concentration). The data processing system can do so, for example, using the following formula: Nutrient scorei= ΣWj*Aijwhere i is a particular condition, j is a particular iModulon, Wjis the weight for the iModulon j, and Aij is the iModulon activity level in the mixing matrix A for the iModulon j for condition i. The nutrient scores can indicate how nutrient perturbations influence the balance between growth and stress responses. The data processing system can assign the scores to the different nutrients by storing a value indicating the nutrient score for each nutrient condition in memory. At operation 1308, the data processing system can select a condition. The condition can identify a type and concentration of a nutrient supplement. The data processing system can select the condition based on the nutrient score for the condition satisfying a criterion. For example, the data processing system can identify the nutrient score that is generated for each of the conditions. The data processing system can rank the conditions in order according to the nutrient scores generated for the conditions with the highest nutrient score ranked first by comparing the nutrient scores. The data processing system can select the condition based on the nutrient score for the condition being the highest or otherwise based on the condition having the highest ranking. At operation 1310, a growth medium can be supplemented with the nutrient supplement type and concentration of the selected condition. For example, a growth medium (e.g., the same growth medium that was used to generate the transcriptomic data) can be prepared. The type of nutrient and concentration of the selected condition can be inserted into the growth medium. Cells of microorganisms can be inoculated into the supplemented growth medium. Because the growth medium was supplemented with the concentration of the selected condition, the cells inoculated into the supplemented growth medium may grow at a faster rate than if the growth medium was not supplemented or an untested supplemental type of nutrient was just randomly selected and input into the growth medium. In some cases, the selection process can avoid introduction of supplemental nutrients that inhibit or reduce growth of cells in the medium. In a non-limiting example, the data processing system can select nutrient supplements and / or concentrations for microbial E. coli growth optimization. For instance, RNA-seq data may be collected from 50 different media conditions supplemented with varying zinc concentrations (0-500 μM ZnSO4). The RNA-seq data can be or include gene expression values of the genes being cultured under the different conditions. The RNA-seq can be uploaded or transmitted to the data processing system. The data processing system can receive the RNA-seq data and apply independent component analysis to the RNA-seq data to identify at least two iModulons: iModulon-Zn1 containing known zinc transport genes (znuABC) alongside uncharacterized genes yjiL and ykgM, and iModulon-Zn2 comprising the zinc-dependent transcription factor Zur and the unknown gene yrdH. The data processing system can calculate the nutrient scores for the different conditions based on the activity levels of the two iModulons under each of the conditions. For instance, the data processing system can calculate a weight for each iModulon as a function of the growth rate of the iModulon across conditions. For instance, the data processing system can determine that the znuABC iModulon demonstrates a strong positive correlation coefficient of 0.85, while the Zur iModulon shows a negative correlation of -0.72. For each condition, the data processing system can use the weight to determine a weighted sum of cell activity levels of the iModulons to determine a nutrient score for the condition. The data processing system can select a nutrient type and concentration based on the nutrient scores for the conditions. For example, the data processing system can compare the scores and select 150 μM ZnSO4 as the optimal concentration based on 150 μM ZnSO4 corresponding to the highest positive nutrient score of 0.92. The data processing system can transmit an identification of the selection to a user device of a researcher or organization with a recommendation to use the selected nutrient supplement and concentration in a medium for future cell growth. The researcher or organization can view the recommendation and subsequently supplement a medium for cell culturing with the selected nutrient supplement of 150 μM ZnSO4. Advantageously, by implementing the systems and methods described herein, the data processing system can automatically identify nutrient supplements to optimize cell growth in a growth media. The data processing system can select the nutrient supplement to optimize the growth of cells with genes with an unknown function by using independent component analysis (ICA) to detect statistically independent patterns of gene co-expression across diverse conditions and using the intermediate layer of the analysis (e.g., the mixing matrix generated in the analysis) to score conditions for different nutrient supplements and concentrations of the nutrient supplements. The data processing system can select a nutrient supplement and concentration to improve cell growth of iModulons and a user or individual can supplement a medium with the selected nutrient supplement and concentration accordingly. Using ICA and scoring in this framework instead of other approaches can facilitate parsing complex, high-dimensional transcriptomic data to reveal subtle but biologically meaningful associations between genetic reactions to nutrient supplements and cell growth, providing a framework for identifying the optimal nutrient supplement and concentration for cell growth of groups of genes containing uncharacterized genes. A study was performed exploring how the addition of 63 diverse nutrients at varying concentrations (2-10 mM) to a standard growth medium (M9 glucose + NH₄Cl) influences cellular responses, particularly stress mitigation and growth promotion. Using an iModulon- based framework, the resulting dataset was systematically analyzed to quantify the impact of specific nutrients on media formulation. This example demonstrates how nutrient supplementation can minimize intracellular stress while enhancing microbial growth and illustrates the practical application of iModulon insights in refining media composition. Furthermore, this approach complements previous work on trace element supplementation, highlighting the versatility of iModulon-guided strategies for media optimization. To demonstrate that nutrient supplementation, informed by transcriptomic analysis and iModulon activities, can systematically optimize media composition, enhance microbial growth, and minimize cellular stress responses. Once one supplementation has proven successful, then Applicant repeated the process with another nutrient addition. Thus, the subject matter of this disclosure can be practiced in an iterative fashion. To investigate the effects of nutrient supplementation on E. coli growth, Applicant expanded the PRECISE-1K dataset (Lamoureux et al. 2023) to create PRECISE-NP881, which includes 252 new transcriptomic samples derived from experiments involving 63 unique nutrients at varying concentrations (2–10 mM) in a standard medium (M9 glucose + NH₄Cl) (FIG. 15A). This dataset, focusing exclusively on E. coli MG1655, represents one of the most extensive collections of nutrient supplementation data to date (FIG. 16). Using Independent Component Analysis (ICA) (Sastry et al. 2019), Applicant decomposed the transcriptome into 137 iModulons, which are categorized into three primary functional groups: (1) catabolism, (2) stress responses, and (3) translation (FIG. 15b). Nutrient-specific transcriptional activities were further decoded through iModulon activity patterns (FIG. 15C, C). Key functional modules focusing on various stress responses and translation exhibited varying activities across nutrient conditions, providing insights into how microbes adapt to different environmental inputs. This modularization enabled a systematic evaluation of transcriptional changes induced by nutrient supplementation. To assess the impact of nutrient supplementation on global transcriptional states, Applicant performed t-SNE analysis on the transcriptomic data (FIG. 15C). Nutrient supplementation samples formed distinct clusters, clearly separating from non-supplemented controls, indicating that specific supplements induce unique transcriptional patterns. This separation in transcriptional states underscores the complex regulatory strategies employed by microorganisms to adapt to nutrient changes, particularly the need to balance growth and stress responses. A fundamental challenge for microorganisms is balancing resource allocation between growth and stress responses to maintain optimal fitness. This dynamic trade-off, known as the fear vs. greed trade-off, involves prioritizing either growth-related genes (“greedy” genes, such as ribosomal proteins) or stress-response genes (“fearful” genes) (Dalldorf et al. 2024). This balance is regulated by competing sigma factors (RpoD and RpoS), the stringent response regulator ppGpp, and other regulatory networks. In E. coli, the fear vs. greed trade-off has been demonstrated using iModulons, data-driven regulons representing independently modulated gene sets (Dalldorf et al. 2024). The strong negative correlation observed between the RpoS (fear) and Translation (greed) iModulons epitomizes this trade-off. To explore the dynamics of the fear vs. greed trade-off under nutrient supplementation conditions, Applicant analyzed the activities of the RpoS and Translation iModulons across various conditions (FIG. 17A). Nutrient supplementation (NS) samples were characterized by high Translation iModulon activity and low RpoS activity, reflecting a growth-prioritized state. In contrast, conditions with elevated RpoS activity showed reduced translation, indicative of a stress-responsive state. Specific supplements, such as threonine (+Thr), glutathione (+Gth), and cysteine (+Cys), were strongly associated with increased RpoS activity and decreased translation, highlighting their role in inducing stress responses. Marginal histograms further reinforced the distinct distributions of iModulon activities between supplemented and control samples. Further analysis revealed that nutrient supplementation generally suppresses stress- related iModulons while promoting growth-related iModulons, including those associated with ribosomal proteins. To deepen Applicant’s understanding, Applicant incorporated the ppGpp iModulon, a key regulator of the stringent response (FIG. 17B). Nutrient perturbations were categorized into three types: NR-C (carbon replacements), NR-N (nitrogen replacements), and NS (nutrient supplements). Clustering patterns showed that nutrient supplementation samples occupied regions of high translation and low stress activity, while carbon replacements and nutrient supplements conditions displayed varying degrees of ppGpp and RpoS activation, reflecting their stress-inducing effects. These findings demonstrate that nutrient supplementation tends to shift transcriptional priorities toward growth, suppressing stress-related pathways. In contrast, nutrient replacement conditions with diverse nutrients often provoke stress responses, particularly under nitrogen or carbon limitations. This transcriptional trade-off underscores the adaptive strategies employed by E. coli to optimize resource allocation in diverse environmental conditions. To evaluate the effects of nutrient supplementation quantitatively, Applicant generated a nutrient score, a weighted sum of 30 stress-related iModulon activities(^^^^^^^^^^^^^^^^ ^^^^^^^^^^ ൌ ∑^^ୀ^ ൈ ^^^ ). The weights (^^^), derived from the correlationcoefficients between iModulon activity (^^^) and growth rate (µ) across all conditions, indicate the contribution of each iModulon to growth. Growth-promoting iModulons, such as those associated with protein biosynthesis, received positive weights, while stress-related iModulons linked to resource allocation for stress management (e.g., CRP, NtrC, RpoS) were assigned negative weights (FIG. 18A, 18B). This metric provides a robust framework for assessing how nutrient perturbations influence the balance between growth and stress responses. Using the nutrient score, Applicant first systematically ranked 40 carbon sources and 22 nitrogen sources in M9 media based on their ability to promote growth while minimizing stress (FIGS. 19A, 19B). High-ranking carbon sources, such as D-glucose, L-arabinose, and Glc-6-P, demonstrated strong alignment with central metabolic pathways, facilitating efficient energy production and minimal stress responses (FIG. 19A). In contrast, low- ranking substrates, including L-proline, alpha-ketoglutarate, and L-asparagine, were associated with metabolic inefficiencies and heightened stress. A similar trend was observed for nitrogen sources. Top-ranked supplements such as NH₄Cl, adenine, and D-alanine supported robust growth with minimal stress induction (FIG.19B), whereas poor-performing sources like L-proline, L-ornithine, and guanosine induced significant stress responses, as indicated by their negative nutrient scores. These rankings highlight the nutrient score’s ability to distinguish between effective and suboptimal nutrient sources. Further analysis revealed a strong correlation between nutrient score and observed growth rate (Pearson’s R = 0.73, P = 2.73 × 10⁻²³), validating the metric’s predictive power (FIG. 18C). This correlation underscores the utility of the nutrient score in decoding complex cellular responses and optimizing medium composition to enhance E. coli growth. By integrating iModulon activities into this framework, Applicant refined media formulations to promote growth and minimize stress, providing a quantitative foundation for understanding how microbial cells allocate transcriptional resources in diverse nutrient environments. Applicant applied the nutrient score framework to evaluate 126 nutrient supplementation conditions in M9 glucose + NH₄Cl media. This analysis revealed a clear gradient of nutrient quality, with key supplements such as L-methionine (2 mM, 10 mM), L- cysteine (2 mM), and D-malic acid (2 mM) achieving the highest nutrient scores (FIG. 20). These supplements promoted growth by enhancing Translation iModulon activity while suppressing stress-related iModulons, effectively shifting the cellular state toward a growth- prioritized mode. Conversely, low-performing supplements, such as inosine (10 mM), thymidine (10 mM), and adenine (2 mM), exhibited negative nutrient scores, reflecting heightened stress responses and reduced growth potential. Non-supplemented conditions also fell into this low-performance category, further emphasizing the critical role of nutrient selection in media optimization. The nutrient-driven transcriptional adaptations observed in this analysis highlight the delicate balance between growth and stress responses, exemplified by the fear vs. greed trade-off. High-scoring supplements, such as L-methionine, effectively reduced RpoS activity while enhancing ribosomal protein synthesis, leading to substantial growth improvements. Specifically, L-methionine supplementation (2 mM) resulted in an average increase of approximately 19.2 in Translation iModulon activity and a decrease of about 22.1 in RpoS iModulon activity, compared to non-supplemented standard controls (both of which had baseline activities of 0). Therefore, the ranking of supplements based on nutrient scores provides actionable insights for media design. High-scoring supplements enhance metabolic flux through growth-promoting pathways while minimizing stress responses, making them ideal candidates for optimized formulations. On the other hand, low-scoring conditions often introduce metabolic bottlenecks or activate stress pathways, thereby impeding growth. Overall, this analysis highlights the power of nutrient score-based evaluations to refine media composition systematically. By prioritizing high-scoring supplements like L- methionine and D-malic acid, researchers can design formulations that not only support robust growth but also mitigate cellular stress. These findings underscore the practical utility of transcriptomic insights in improving microbial performance across diverse applications. These findings demonstrate the power of machine-learning-based transcriptomic analysis in guiding media optimization. The nutrient score provides a quantitative framework for evaluating nutrient quality, enabling systematic and iterative refinement of media composition. Key takeaways include: - iModulon-guided insights: iModulon analysis offers a mechanistic understanding (based on first biological principles) of how specific nutrients influence cellular physiology, revealing pathways that can be targeted for optimization. - Predictive metrics: The nutrient score correlates strongly with growth rate (Pearson’s R = 0.73), validating its utility as a predictive metric for media formulation. - Practical applications: High-scoring nutrients, such as L-methionine and D- malic acid, represent promising candidates for improving industrial and research-scale microbial cultivation. - Range of strains: The method can be applied to wild type strains to optimize their growth. The methods can also be applied to strains that have been pre- engineered to produce a product, so that growth and productivity are both considered. The method can also be applied to strain after they have been optimized using adaptive laboratory evolution. The methods can be applied to prokaryotic and eukaryotic organisms. This example introduces a systematic approach to media optimization through nutrient supplementation, guided by transcriptomic analysis and iModulon activities. The nutrient score provides a robust, quantitative metric for evaluating nutrient quality, offering actionable insights into media composition. By balancing growth promotion with stress mitigation, this method establishes a scalable framework for advancing microbial performance in both research and industrial contexts. The wild-type strain Escherichia coli str. K-12 MG1655 (E. coli) was used throughout this study. E. coli was routinely grown in LB medium (#71753) or M9 medium. The M9 medium contained 47.75 mM Na2HPO4, 22.04 mM KH2PO4, 8.56 mM NaCl, 18.70 mM NH4Cl, 2 mM MgSO4, 0.1 mM CaCl2 and trace elements. Trace elements were prepared as a 2000 × concentrated solution, consisting of 100 mM FeCl3, 9.54 mM ZnCl2, 8.41 mM CoCl2, 8.27 mM Na2MoO4, 0.75 mM CaCl2, 0.91 mM CuCl2 and 0.5 mM H3BO3 in a 3.7% (w / w) hydrochloric acid solution. For additional RNA-Seq samples of nutrient supplementations, a range of nutrients were added into M9 medium at two concentrations of 2 mM and 10 mM, including D-Glucose-6-Phosphate, D-Ribose, Maltose, D-Melibiose, L- Arabinose, Maltotriose, L-Rhamnose, a-D-Lactose, L-Lyxose, D-Galactose, D-Trehalose, D- Mannose, L-Fucose, D-Saccharic acid, D-Gluconic acid, D-Galacturonic acid, Mucic acid, D- Glucuronic acid, D-Sorbitol, D-Mannitol, Dulcitol, m-Tartaric acid, D-Malic acid, DL-Malic acid, Pyruvic acid, a-Ketoglutaric acid, Methylpyruvate, b-Methyl-D-Galactoside, D- Glucosamine, N-Acetyl-D-Glucosamine, N-Acetyl-Neuraminic acid, Histidine, Isoleucine, Leucine, g-Amino-N-Butyric acid, Lysine, Methionine, Phenylalanine, Proline, Serine, Threonine, Tryptophan, Alanine, Tyrosine, Valine, Arginine, Asparagine, Aspartic acid, Cysteine, Glutamic acid, Glycine, Glutamine, Putrescine, 2-Deoxyadenosine, Adenosine, Inosine, Uridine, Cytidine, Thymidine, Adenine, and Cytosine. LB was supplemented with 1.5% wt / vol agar for solid media. Unless stated otherwise, chemical reagents for cell culture were primarily obtained from Sigma-Aldrich (Burlington, MA). In flask culture experiments, OD600 of bacterial cultures was measured at 600 nm using the BioMate™ 3S Spectrophotometer (Thermo Fisher Scientific, Waltham, MA). For high-throughput assessments, the Infinite 200 Pro Plate Reader (Tecan, Männedorf, Switzerland) was used for OD600 measurements in Flat Bottom 96-well plates, each containing 100 μl of culture. Plates were incubated at 37°C with orbital shaking, with readings taken at 10 to 15-minute intervals. Growth dynamics involved calculating lag-time (λ) and specific growth rate (μ) via linear regression, utilizing the QurvE web tool (Wirth et al. 2023). The Applicant produced a total of 252 new RNA-Seq datasets pertinent to nutrient supplementation conditions. The Applicant obtained transcriptomic data from two biological replicates. The RNA extraction and library preparation procedures were in accordance with the protocol established in our prior study (Shin, Rychel, and Palsson 2023). Briefly, samples were centrifuged for 10 min at 5000 × g at 4 °C, and the supernatant was removed. RNA was isolated from the harvested cells using the Quick-RNA Fungal / Bacterial Microprep Kit (Zymo Research, Irvine, CA), according to the manufacturer's guidelines. As previously described, ribosomal RNA and genomic DNA contaminants were eliminated from 1 μg total RNA using the RiboRid method (Choe et al. 2021). rRNA depletion was verified via the 4150 TapeStation System (Agilent, Santa Clara, CA) with High Sensitivity RNA ScreenTape. The rRNA-depleted RNA was then converted into libraries using the KAPA RNA HyperPrep kit (Roche, Basel, Switzerland) following the manufacturer's protocols. Library quality was assessed with the 4150 TapeStation System (Agilent) using D1000 ScreenTape, and quantification was done with the Qubit 2.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA) using the Qubit dsDNA HS Assay Kit. Subsequently, the libraries were combined and sequenced using a 100 bp single-end protocol on the Element AVITI System Sequencing platform at the Genomics Core Facility in Scripps Research. For data processing and initial RNA-Seq quality control, raw reads were trimmed with Trim Galore (https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / ) and assessed with FastQC (https: / / www.bioinformatics.babraham.ac.uk / projects / fastqc / ). Aligned reads to the E. coli reference genome (GCF_000005845.2) were achieved using Bowtie (Langmead et al. 2009), with SAM to BAM file conversion via Samtools (http: / / www.htslib.org / ). Gene read counts were determined using RSeQC (Wang, Wang, and Li 2012) and FeatureCounts (Liao, Smyth, and Shi 2014), with all quality control metrics aggregated by MultiQC (https: / / multiqc.info / ) (Ewels et al. 2016). The Applicant ensured RNA-Seq data quality by confirming compliance with FASTQC standards, including per_base_sequence_quality, per_sequence_quality_scores, per_base_n_content, and adapter_content. A high correlation was observed within biological replicates (R2> 0.94) in this RNA-Seq data. The read counts were then normalized and presented as log2-transformed Transcripts per Million (logTPM) for downstream analysis. To identify differentially expressed genes (DEGs), raw read counts were normalized using the Bioconductor package DESeq2 (Love, Huber, and Anders 2014). Significance was determined with an adjusted P- value (Padj) < 0.05. For the iM-based engineering strategy implementation, the Applicant utilized iM data generated through Independent Component Analysis (ICA) based on the PRECISE-1K iModulonDB (detailed at https: / / imodulondb.org / ) (Shin, Zielinski, and Palsson 2024). iM analysis is able to be performed using the (i) iModulonMiner (https: / / github.com / sbrg / imodulonminer) (Sastry et al. 2024), (ii) Pymodulon package (https: / / github.com / SBRG / pymodulon), and iModulonDB (https: / / imodulondb.org / ) (Rychel et al. 2021). Following methodologies described in prior studies, the Applicant modeled differences in iM activities across biological replicates using a log-normal distribution. The approach involved comparing mean activity differences across all iMs to compute P-values. These P-values were adjusted using the Benjamini-Hochberg method to address multiple hypothesis testing. iMs demonstrating an FDR change of less than 0.05 under specific conditions were identified as statistically significant. FIG. 14 is a block diagram of an example computer system 1400. The computer system or computing device 1400 can include or be used to implement the system 1200 or their components such as the data processing system 1202. The computing system 1400 includes a bus 1405 or other communication component for communicating information and a processor 1410 or processing circuit coupled to the bus 1405 for processing information. The computing system 1400 can also include one or more processors 1410 or processing circuits coupled to the bus for processing information. The computing system 1400 also includes main memory 1415, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 1405 for storing information, and instructions to be executed by the processor 1410. The main memory 1415 can be or include the data repository 1212. The main memory 1415 can also be used for storing position information, temporary variables, or other intermediate information during execution of instructions by the processor 1410. The computing system 1400 may further include a read only memory (ROM) 1420 or other static storage device coupled to the bus 1405 for storing static information and instructions for the processor 1410. A storage device 1425, such as a solid state device, magnetic disk or optical disk, can be coupled to the bus 1405 to persistently store information and instructions. The storage device 1425 can include or be part of the data repository 1212. The computing system 1400 may be coupled via the bus 1405 to a display 1435, such as a liquid crystal display, or active matrix display, for displaying information to a user. An input device 1430, such as a keyboard including alphanumeric and other keys, may be coupled to the bus 1405 for communicating information and command selections to the processor 1410. The input device 1430 can include a touch screen display 1435. The input device 1430 can also include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 1410 and for controlling cursor movement on the display 1435. The display 1435 can be part of the data processing system 1202, the client device 1220 or other component of FIG. 12, for example. The processes, systems and methods described herein can be implemented by the computing system 1400 in response to the processor 1410 executing an arrangement of instructions contained in main memory 1415. Such instructions can be read into main memory 1415 from another computer-readable medium, such as the storage device 1425. Execution of the arrangement of instructions contained in main memory 1415 causes the computing system 1400 to perform the illustrative processes described herein. One or more processors in a multi-processing arrangement may also be employed to execute the instructions contained in main memory 1415. Hard-wired circuitry can be used in place of or in combination with software instructions together with the systems and methods described herein. Systems and methods described herein are not limited to any specific combination of hardware circuitry and software. Although an example computing system has been described in FIG. 14, the subject matter including the operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware. Equivalents It is to be understood that while the disclosure has been described in conjunction with the above embodiments, that the foregoing description and examples are intended to illustrate and not limit the scope of the disclosure. Other aspects, advantages and modifications within the scope of the disclosure will be apparent to those skilled in the art to which the disclosure pertains. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All nucleotide sequences provided herein are presented in the 5′ to 3′ direction. The embodiments illustratively described herein may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, for example, the terms “comprising”, “including,” containing”, etc. shall be read expansively and without limitation. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the disclosure. Thus, it should be understood that although the present disclosure has been specifically disclosed by specific embodiments and optional features, modification, improvement and variation of the embodiments therein herein disclosed may be resorted to by those skilled in the art, and that such modifications, improvements and variations are considered to be within the scope of this disclosure. The materials, methods, and examples provided here are representative of particular embodiments, are exemplary, and are not intended as limitations on the scope of the disclosure. The scope of the disclosure has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the disclosure. This includes the generic description with a proviso or negative limitation removing any subject matter from the genus, regardless of whether or not the excised material is specifically recited herein. In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that embodiments of the disclosure may also thereby be described in terms of any individual member or subgroup of members of the Markush group. All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety, to the same extent as if each were incorporated by reference individually. In case of conflict, the present specification, including definitions, will control. Sequence Listing RS23895 (uncharacterized protein) SEQ ID NO: 1 ATGGTAGAAACGAACGACTATCAAATGGCTATCGATTTACTACGTTGTCACCTAG GTATCTCAGAGGACGAGGCAAAGCAACAACTTGGTATTTCGACTGATGATCATA TGGCTAATCGTATTGCAGAAACGCAACACGCATTGATGGGTCTCGGTAACGAAA AATAA Acetate—CoA ligase gene (acs; characterized protein) SEQ ID NO: 2 ATGAGCGAAGCCCACGTTTATCCGGTAAAAGAAAACATTAAAACTCATACACAC GCGGATAATGAAACTTACCTAGCCATGTATCAGCAGTCGGTAACCGACCCAGAG GGCTTCTGGAACGAGCACGGCAAAATCGTTGATTGGATTAAACCTTTCACCCAAG TAAAAAGCACCTCTTTCGACACGGGTCACGTCGACATCCGCTGGTTTGAAGACGG CACACTTAACGTTTCAGCAAACTGTATTGACCGCCATCTAGCAGAACACGGCGAC GACGTAGCAATAATCTGGGAAGGCGATGACCCTGCAGACGATAAAACGCTGACG TTCAATGAACTGCACAAAGAAGTATGTAAATTCTCAAACGCATTGAAAGATCAA GGCGTACGTAAAGGCGACGTAGTTTGTCTCTACATGCCAATGGTACCTGAAGCTG CAATCGCGATGCTGGCTTGTACCCGCATCGGCGCAGTCCACACTGTGGTATTCGG CGGTTTCTCACCAGAAGCACTTTCTGGCCGTATCATTGACTCAGACGCTAAAGTT GTCATCACCGCAGACGAAGGCGTTCGTGGCGGACGTGCGGTTCCACTGAAAAAG AATGTTGATGAAGCACTGACTAACCCAGACGTGAAAACCATCAGCAAAGTGTTG GTTCTTAAACGTACTGGTGGCGATGTTGAATGGCATGATCACCGTGATGTTTGGT GGCATGAAGCAACCGCAAGTGTTTCTGACGTTTGCCCACCAGAAGAGATGAAAG CCGAAGATCCACTTTTCATCCTTTACACCTCAGGCTCTACGGGTAAACCTAAAGG CGTACTGCATACCACTGGTGGCTACCTTGTTTATGCCGCAATGACATTTAAATAC GTCTTCGACTACCAGCCGGGCGAAACCTTCTGGTGTACTGCTGACGTGGGCTGGA TTACTGGTCACACGTACCTTATCTACGGTCCACTAGCGAACGGCGCTAAAACCAT CTTGTTTGAGGGTGTGCCAAACTACCCAAGCACAAGCCGTATGAGCGAAGTGGT TGATAAGCATCAAGTGAACATCCTTTACACTGCGCCAACTGCGATTCGTGCGCTA ATGGCAAAAGGTAATGAAGCGGTTGAAGGCACGTCTCGTACAAGCCTTCGCATC ATGGGTTCGGTAGGTGAGCCAATCAACCCAGAAGCATGGGAGTGGTACTACAAG ACCATCGGTAACGAAAATTCACCGATTGTCGATACTTGGTGGCAAACTGAAACA GGCGGCATCTTGATTGCTCCACTACCAGGCGCTACAGATCTAAAACCAGGTTCCG CGACCCGTCCATTCTTCGGCGTACAACCTGCGCTTGTCGATAACATGGGTAACAT TATTGAAGGTGCAGCTGAAGGCAACCTTGTGATTCTTGATTCGTGGCCAGGTCAG ATGCGTACGGTCTATGGTGACCATGAACGCTTCGAACAAACTTACTTCTCAACCT TCAAGGGCATGTACTTCACCAGTGATGGCGCTCGTCGTGACGAAGATGGTTACTA CTGGATCACAGGCCGTGTGGATGACGTATTGAACGTTTCTGGACACCGTATGGGT ACCGCAGAGATTGAATCTGCCCTAGTCGCGCACCACAAGATTGCAGAAGCAGCC ATTGTAGGTATCCCGCACGACATCAAAGGTCAGGCTATCTATGCTTACATCACGC TAAACGATGGTGAGTTCCCTTCAGCGGAACTGCACAAAGAAGTTAAAGACTGGG TGCGCAAAGAGATCGGTCCAATTGCAACACCAGATGTACTGCACTGGACGGATT CCCTACCGAAGACTCGCTCTGGTAAGATCATGCGTCGAATTCTGCGTAAGATTGC AACTGGCGATACGAGCAACCTGGGTGACACGTCAACGCTTGCCGACCCTAGCGT TGTAGACAAACTTATCGCTGAAAAAGCTGAGCTGGCGTAA 3’-5’ exonuclease gene (RS14415; characterized protein) SEQ ID NO: 3 ATGAATTGGCTACAACGAAAATATTGGCACTATAAACTGAAAGGCTCGCCTTATC AGTCGCTATTTTGTGCACCAGATAAAACCGAGCTGGTTTCTCTCGACTGCGAAAC CACCAGCCTAGACCCGAATAGAGCAGAGCTGGTGACCATCGCAGCCACTAAAAT CATTGATAACCGCATTATCACCAGCCAACCATTCGAAGTACATTTGCGAGCGCCA CAATCTCTCGACTCCGGCTCGGTAAAAATCCATAAAATTCGCCATCAGGATTTAG TCGACGGCATCAGCGAAAAGGATGCGCTACTAAAATTAATAGACTTTATCGGCA ATCGTCCGCTCGTCGGCTATCACATCCGTTACGACAAAAAAATCTTAGACCTCGC CTGTCAAAGACAACTGGGATTCCCTCTGCCCAACCCTCTTATTGAAGTAAGCCAG ATTTACCACGACAAGCTGGAGCGACATTTACCGAACGCCTATTTCGACTTAAGTC TGGAGGCGATTTGCAAGCACCTTGAGCTGCCAATTCAAGATAAGCACGACGCTTT GCAAGATGCGATATCCGCAGCATTGGTATTTGTTCGTTTAACCAAAGGCGATTTA CCGAGTCTCAGCGTGCCATACACATAA DUF294 nucleotidyltransferase-like domain-containing gene (RS14420; uncharacterized protein) SEQ ID NO: 4 ATGCCTGATAAGTTTAATATGCAATCTCCCCCGTTCGATCGCCTGACCAGCAAGC AGCAAGCGCGACTGCGTTCGTCTCTCGACGTGGCATATTATCGAACTCGAGATGT GATTTTATCTTGTGGGCAAACCAACCCTCATTTGCATATTTTGATCAAAGGTGCC GTGGAAGAGCGCTCCAAAGACCAGAGCGAAGTTTTCGCTCACTATGCCAATGAC GATATGTTTGATGTCCGTTCTTTGTTTGAAGAGAGTGTGCGTCATCAATATGTAGC GCTTGAGGATACGCTGTCCTATTTGCTGCCAAAAGAAGTGTTTCTTGAACTCTAT AACGAGAATGGACAATTCGCAGCGTACTTCGATAACAACCTATCTAAACGACAG GCATTAATCGAAGCCGCGCAGCAACAGCAAAATCTGGCCGAGTTTATTCTGACT AAAGTCGATAAGAGCATTTACCATCCTCCGATGATTCTTAAGCCCGACATGCCTA TCAATGAGGTCACACGAACCTTAAAAGAAAACGGCATCGACGCCGCATTAGTTG AGCTCAACCCAGATGATGCACGCCTCGAAAAATGGCCTTCAGCGCATCCGTACG CCATTGTCACGCGAACAAACATGCTCCATTCCGTGATGTTAGATAACTGCGCTGT CGATACACCGGTGGGCGAGATCGCCACCTTCCCGGTTTACCACGTTGATGAAGGC GACTTCCTGTTTAATGCCATGATTACGATGACGCGCCACAGGATGAAGCGCCTGA TGGTCTGCGATGGTAATCAAGCCATTGGCATGTTGGATATGACTCAGATCCTCAG TGCATTCTCCACTCATTCACACGTGCTCACCCTCAGCATCGCCAGAGCCAGCAGT GTGGAAGAACTCGCTCTGGCGTCCAACCGACAACGCCAGTTGGTAGAGAGTCTG GTCAGAAACGGCATCCGCACCCGATTTATCATGGAATTGATATCTGCCGTTAATG AACAAATTATCGAGAAAGCGTTTGAATTGGTTGTGCCACCAGCACTCCACGACC ACTGCTGCCTAATTGTTTTAGGCTCGGAAGGCCGAGGTGAGCAGATCCTCAAGAC CGACCAAGATAACGCGCTGATCATTAAAGACGGTTTAGAGTGGTCGCACAGTGA ACAGGTGATGCAATCATTCACCCACACCCTGCAGCAACTGGGTTATCCGCTCTGC CCTGGCAAAGTGATGGTCAACAATCCTAAGTGGGTAAACTCTCAATCAGAGTGG ATTCGAACCCTTGATAATTGGATTTCAAAAGCGCAGCCAGAACAAATCATGGAC TTAGCCATCTTCAGTGATGCTCAAGCCGTTGCCGGCAATCGAGAATTACTCACTC CCGTCGCCAACCATCTCAGAGACACAATGAAAGATCGCATGTTGATCCTGTCCGA CTTTACCCGGCCAGCGTTGCAGTTTTCTGTGCCTCTGACTCTGTTTGGTAACGTAA AAAGTGCCAAAGATGGCTTAGATATCAAACGCGGCGGTATTTTCCCGATAGTGC ACGGCATTAGAACTCTGTCTCTCGAATATGCCATTGAAGAGAAAAATACCTTTGC CCGCATAGAAGCCCTGAGAAACAAGCGAATTTTAGAACCAGAAACGGCGGATAA CTTAAACGAAGCACTCAAGTTGTTCTTTAAGCTTCGGCTCAATCAGCAACTTAAC CAACAAGAGGCACAGAATAATATCGACCTCAAACAACTCGACCGAACAGAGCGC GACTTGCTGCGTCATAGCCTGCATGTCGTGAAAAAGTTTAAGCAGTTTTTAGGTT TCCACTATCAGATCCGTGATTAA Cation acetate symporter gene (RS14445; characterized protein) SEQ ID NO: 5 ATGGATTTGAAAACTATTACCTACCTGGTTGTTGGCGCAACCTTCATCCTGTACA TCGGGATTGCGATTTGGGCGCGCGCAGGCTCAACCAAAGAGTTCTATGTTGCGG GTGGCGGTGTAAACCCAATCGCGAACGGTATGGCGACAGCAGCAGACTGGATGT CAGCGGCGTCGTTTATTTCGATGGCGGGTCTGATTGCCTTCATGGGATACGGTGG TTCTGTATTCCTAATGGGTTGGACTGGCGGTTACGTTCTGCTTGCTCTGCTTCTTG CTCCTTACCTACGTAAGTTCGGTAAATTTACCGTGCCAGAGTTTGTCGGTGAGCG TTTCTACTCTAAGACAGCGCGTATCGTAGCGGTTGTGTGTCTGATCATCGCATCA GTAACATACGTTATCGGTCAGATGAAAGGCGTAGGTGTTGCATTCGGTCGTTTCT TAGAAGTAGATTACGCAACAGGTCTACTGATTGGTATGTGTATCGTATTTATGTA CGCAGTACTAGGCGGCATGAAAGGCATCACATACACGCAGATTGCTCAATATTG TGTACTCATTCTTGCTTACACTATCCCAGCAGTATTTATCTCTCTGCAATTGACTG GTAACCCAATTCCGCAAATTGGCCTTGGTAGTACCATGGCGGGAACCGATGTCTA TCTACTCGATCGGCTCGATCAGGTGGTGACCGAACTCGGCTTTAGTGAATACACC ACTCAAGTTCGTGGTGATACACTAAACATGTTTGCTTACACCATGTCACTGATGA TCGGTACGGCTGGTCTGCCACACGTAATCATCCGTTTCTTTACGGTACCTAAGGT TCGTGACGCACGTACATCAGCAGGTTGGGCACTAATCTTCATTGCTATTCTTTAC ACAACAGCACCTGCAGTATCAGCAATGGCGCGCCTAAACCTAATGAACACAGTT AACCCGGCGCCTGGCCAACACCTAGCTTACGACGAGCGTCCTGCTTGGTTTAAAA ACTGGGAAAAAACAGGCCTGCTAGGTTTCGATGATAAGAACGGTGATGGCAACA TTCAGTACACTTCTGATGCGGCTACTAACGAGCTAAAAGTTGATAACGACATCAT GGTACTGGCTAACCCAGAGATTGCTAAGCTACCAAACTGGGTAATCGCATTGGTT GCGGCAGGTGGTCTGGCAGCAGCACTATCAACCGCTGCGGGTCTACTACTGGCG ATTTCGTCCGCGATATCTCACGACTTAATCAAAGGGGTCATTAACCCGAATATCT CCGAGAAGAAAGAGCTCCTTGCCAGCCGGATTTCCATGGCAGTGGCAATTGCGG TAGCAGGCTACTTAGGTCTACACCCACCGGGCTTCGCAGCAGGTACCGTTGCACT GGCCTTTGGTCTGGCAGCGTCCTCTATCTTCCCTGCATTGATGATGGGTATCTTCA GTAAGAGCATCAACAAAGAAGGCGCTATCGCAGGTATGATCGCAGGTATCAGTA TTACACTGTTCTACGTATTCCAACACAAAGGTATCCTATTCATCGCGGATTGGAC TTACCTGGAAAGCTGGGGCAGCAACTGGTTCCTAGGCATTGAGCCAAACGCGTTT GGTGCAATTGGTGCGCTATTCAACTTCATCGTGGCGTTCGCAGTATCGAGAGTAA CAGCAGAAACACCACAAGAAGTGAAAGATCTGGTTGAGCACGTACGTGTACCAG TAGGTGCGGGCCAAGCTGTCGATCACTAA DUF4212 domain-containing gene (RS14450; uncharacterized protein) SEQ ID NO: 6 ATGGCGTTTGAAAGCGAAGAAAGAGCCAAAGCCTACTGGGACAAGAACGTGAA GCTAATGATTGGCCTAATGGTCATCTGGTTCTTAGTGTCTTTCGGCTGCGGCATTT TATTTGTCGATGTACTAAATCAATTTCAACTAGGTGGATACAAACTTGGTTTCTG GTTTGCTCAACAAGGCTCCATCTACGCCTTCCTGGGTATTATTTTCTACTACGCGT GGAAGATGCGCAAAATAGACCGTGAATTCGGCGTGGATGAGTAA Acetate iModulon Booster A SEQ ID NO: 7 ttcaaaagcgcagccagaacaaatcatggacttagccatcttcagtgatgctcaagccgttgccggcaatcgagaattactcactcccg tcgccaaccatctcagagacacaatgaaagatcgcatgttgatcctgtccgactttacccggccagcgttgcagttttctgtgcctctgac tctgtttggtaacgtaaaaagtgccaaagatggcttagatatcaaacgcggcggtattttcccgatagtgcacggcattagaactctgtct ctcgaatatgccattgaagagaaaaatacctttgcccgcatagaagccctgagaaacaagcgaattttagaaccagaaacggcggata acttaaacgaagcactcaagttgttctttaagcttcggctcaatcagcaacttaaccaacaagaggcacagaataatatcgacctcaaac aactcgaccgaacagagcgcgacttgctgcgtcatagcctgcatgtcgtgaaaaagtttaagcagtttttaggtttccactatcagatccg tgattaagcccggtaagttgtaaaagaggtatctgggatgaattggctacaacgaaaatattggcactataaactgaaaggctcgcctta tcagtcgctattttgtgcaccagataaaaccgagctggtttctctcgactgcgaaaccaccagcctagacccgaatagagcagagctgg tgaccatcgcagccactaaaatcattgataaccgcattatcaccagccaaccattcgaagtacatttgcgagcgccacaatctctcgact ccggctcggtaaaaatccataaaattcgccatcaggatttagtcgacggcatcagcgaaaaggatgcgctactaaaattaatagacttta tcggcaatcgtccgctcgtcggctatcacatccgttacgacaaaaaaatcttagacctcgcctgtcaaagacaactgggattccctctgc ccaaccctcttattgaagtaagccagatttaccacgacaagctggagcgacatttaccgaacgcctatttcgacttaagtctggaggcga tttgcaagcaccttgagctgccaattcaagataagcacgacgctttgcaagatgcgatatccgcagcattggtatttgttcgtttaaccaaa ggcgatttaccgagtctcagcgtgccatacacataattcctgttgacctggctcccataattacttttatttcaaccctttttcggtagatccct ctctatcatcctaaagtctaattgtggaaatacgcgctgtgactgaaagttattagcacaaggtcagtcttgcgacgaccaactgtttaaaa caaatatccgttctattctttatgttttagaaaatcataatgttgatagagcgggtaaatacaacagcccgacaaagtaaagattcaaagtct tatgtagattgttgtaaggccgagaaataaaggcacaaggagagaagccatgagcgaagcccacgtttatccggtaaaagaaaacatt aaaactcatacacacgcggataatgaaacttacctagccatgtatcagcagtcggtaaccgacccagagggcttctggaacgagcac ggcaaaatcgttgattggattaaacctttcacccaagtaaaaagcacctctttcgacacgggtcacgtcgacatccgctggtttgaagac ggcacacttaacgtttcagcaaactgtattgaccgccatctagcagaacacggcgacgacgtagcaataatctgggaaggcgatgac cctgcagacgataaaacgctgacgttcaatgaactgcacaaagaagtatgtaaattctcaaacgcattgaaagatcaaggcgtacgtaa aggcgacgtagtttgtctctacatgccaatggtacctgaagctgcaatcgcgatgctggcttgtacccgcatcggcgcagtccacactg tggtattcggcggtttctcaccagaagcactttctggccgtatcattgactcagacgctaaagttgtcatcaccgcagacgaaggcgttc gtggcggacgtgcggttccactgaaaaagaatgttgatgaagcactgactaacccagacgtgaaaaccatcagcaaagtgttggttctt aaacgtactggtggcgatgttgaatggcatgatcaccgtgatgtttggtggcatgaagcaaccgcaagtgtttctgacgtttgcccacca gaagagatgaaagccgaagatccacttttcatcctttacacctcaggctctacgggtaaacctaaaggcgtactgcataccactggtgg ctaccttgtttatgccgcaatgacatttaaatacgtcttcgactaccagccgggcgaaaccttctggtgtactgctgacgtgggctggatta ctggtcacacgtaccttatctacggtccactagcgaacggcgctaaaaccatcttgtttgagggtgtgccaaactacccaagcacaagc cgtatgagcgaagtggttgataagcatcaagtgaacatcctttacactgcgccaactgcgattcgtgcgctaatggcaaaaggtaatga agcggttgaaggcacgtctcgtacaagccttcgcatcatgggttcggtaggtgagccaatcaacccagaagcatgggagtggtactac aagaccatcggtaacgaaaattcaccgattgtcgatacttggtggcaaactgaaacaggcggcatcttgattgctccactaccaggcgc tacagatctaaaaccaggttccgcgacccgtccattcttcggcgtacaacctgcgcttgtcgataacatgggtaacattattgaaggtgc agctgaaggcaaccttgtgattcttgattcgtggccaggtcagatgcgtacggtctatggtgaccatgaacgcttcgaacaaacttacttc tcaaccttcaagggcatgtacttcaccagtgatggcgctcgtcgtgacgaagatggttactactggatcacaggccgtgtggatgacgt attgaacgtttctggacaccgtatgggtaccgcagagattgaatctgccctagtcgcgcaccacaagattgcagaagcagccattgtag gtatcccgcacgacatcaaaggtcaggctatctatgcttacatcacgctaaacgatggtgagttcccttcagcggaactgcacaaagaa gttaaagactgggtgcgcaaagagatcggtccaattgcaacaccagatgtactgcactggacggattccctaccgaagactcgctctg gtaagatcatgcgtcgaattctgcgtaagattgcaactggcgatacgagcaacctgggtgacacgtcaacgcttgccgaccctagcgtt gtagacaaacttatcgctgaaaaagctgagctggcgtaagctaactagccaAAAGAGGAGAAAATGGCTGAAG CGCAAAATGatcCCCTGCTGCCGGGATACTCGTTTAATGCCCATCTGGTGGCGGGT TTAACGCCGATTGAGGCCAACGGTTATCTCGATTTTTTTATCGACCGACCGCTGG GAATGAAAGGTTATATTCTCAATCTCACCATTCGCGGTCAGGGGGTGGTGAAAA ATCAGGGACGAGAATTTGTTTGCCGACCGGGTGATATTTTGCTGTTCCCGCCAGG AGAGATTCATCACTACGGTCGTCATCCGGAGGCTCGCGAATGGTATCACCAGTG GGTTTACTTTCGTCCGCGCGCCTACTGGCATGAATGGCTTAACTGGCCGTCAATA TTTGCCAATACGGGGTTCTTTCGCCCGGATGAAGCGCACCAGCCGCATTTCAGCG ACCTGTTTGGGCAAATCATTAACGCCGGGCAAGGGGAAGGGCGCTATTCGGAGC TGCTGGCGATAAATCTGCTTGAGCAATTGTTACTGCGGCGCATGGAAGCGATTAA CGAGTCGCTCCATCCACCGATGGATAATCGGGTACGCGAGGCTTGTCAGTACATC AGCGATCACCTGGCAGACAGCAATTTTGATATCGCCAGCGTCGCACAGCATGTTT GCTTGTCGCCGTCGCGTCTGTCACATCTTTTCCGCCAGCAGTTAGGGATTAGCGT CTTAAGCTGGCGCGAGGACCAACGTATCAGCCAGGCGAAGCTGCTTTTGAGCAC CACCCGGATGCCTATCGCCACCGTCGGTCGCAATGTTGGTTTTGACGATCAACTC TATTTCTCGCGGGTATTTAAAAAATGCACCGGGGCCAGCCCGAGCGAGTTCCGTG CCGGTTGTGAAGAAAAAGTGAATGATGTAGCCGTCAAGTTGTCATAACTACGCT CTGGCTGCTTAGgggtaaccaggcatcaaataaaacgaaaggctcagtcgaaagactgggcctttcgttttatctgttgtttgt cggtgaacgctctctactagagtcacactggctcaccttcgggtgggcctttctgCGTTTATACGCTAGTAAGCGA AAAGAAACCAATTGTCCATATTGCATCAGACATTGCCGTCACTGCGTCTTTTACT GGCTCTTCTCGCTAACCAAACCGGTAACCCCGCTTATTAAAAGCATTCTGTAACA AAGCGGGACCAAAGCCATGACAAAAACGCGTAACAAAAGTGTCTATAATCACGG CAGAAAAGTCCACATTGATTATTTGCACGGCGTCACACTTTGCTATGCCATAGCA TTTTTATCCATAAGATTAGCGGATCCTACCTGACGCTTTTTATcgcaaCTCTCTACTgt ttctcCATtactagagaaagaggagaaaATGAAACCAGTAACGTTATACGATGtcgcagagtatgccggtg tctcttatatgaccgtttcccgcgtggtgaaccaggccagccacgtttctgcgaaaacgcgggaaaaagtggaagcggcgatggtgg agctgaattacattcccaaccgcgtggcacaacaactggcgggcaaacagtcgttgctgattggcgttgccacctccagtctggccct gcacgcgccgtcgcaaattgtcgcggcgattaaatctcgcgccgatcaactgggtgccagcgtggtggtgtcgatggtagaacgaag cggcgtcgaagcctgtaaagcggcggtgcacaatcttctcgcgcaacgcgtcagtgggctgatcattaactatccgctggatgaccag gatgccattgctgtggaagctgcctgcactaatgttccggcgttatttcttgatgtctctgaccagacacccatcaacagtattatttactcc catgaggacggtacgcgactgggcgtggagcatctggtcgcattgggtcaccagcaaatcgcgctgttagcgggcccattaagttctg tctcggcgcgtctgcgtctggctggctggcataaatatctcactcgcaatcaaattcagccgatagcggaacgggaaggcgactggag tgccatgtccggttttcaacaaaccatgcaaatgctgaatgagggcatcgttcccactgcgatgctggttgccaacgatcagatggcgct gggcgcaatgcgcgccattaccgagtccgggctgcgcgttggtgcggatatctcggtagtgggatacgacgataccgaagatagctc atgttatatcccgccgttaaccaccatcaaacaggattttcgcctgctggggcaaaccagcgtggaccgcttgctgcaactctctcagg gccaggcggtgaagggcaatcagctgttgccagtctcactggtgaaaagaaaaaccaccctggcgcccaatacgcaaaccgcctct ccccgcgcgttggccgattcattaatgcagctggcacgacaggtttcccgactggaaagcgggcagtgataaggatcctaattggtaa cgaatcagacaattgacggctcgagggagtagcatagggtttgcagaatccctgcttcgtccatttgacaggcacattatgcatcgatg ataagctgtcaaacatgagcagatcctctacgccggacgcatcgtggccggcatcaccggcgccacaggtgcggttgctggcgccta tatcgccgacatcaccgatggggaagatcgggctcgccacttcgggctcatgagcaaatattttatctgagctccttttgttatcaataaa aaaggccccccgatttgggagGCCTTCAATAATTGGAGGGGGGTCTGACGCTCAGTAGcGGC GGCGAAACCCCGCCCTGTCAGGGGCGGGGTTTGCGGCGTTAAGCgatccccctatgcaagg gtttattgttttctaaaatctgattaccaattagaatgaatatttcccaaatattaaataataaaacaaaaaaattgaaaaaagtgtttccaccat tttttcaatttttttataatttttttaatctgttatttaaatagtttatagttaaatttacattttcattagtccattcaatattctctccaagataactacg aactgctaacaaaattctctccctatgttctaatggagaagattcagccactgcatttcccgcaatatcttttggtatgattttacccgtgtcc atagttaaaatcatacggcataaagttaatatagagttggtttcatcatcctgataattatctattaattcctctgacgaatccataatggctctt ctcacatcagaaaatggaatatcaggtagtaattcctctaagtcataatttccgtatattcttttattttttcgttttgcttggtaaagcattatggt taaatctgaatttaattccttctgaggaatgtatccttgttcataaagctcttgtaaccattctccataaataaattcttgtttgggaggatgattc cacggtaccatttcttgctgaataataattgttaattcaatatatcgtaagttgcttttatctcctattttttttgaaataggtctaattttttgtataa gtatttctttactttgatctgtcaatggttcagatacgacgactaaaaagtcaagatcactatttggttttagtccactctcaactcctgatcca aacatgtaagtaccaataaggttattttttaaatgtttccgaagtatttttttcactttattaatttgttcgtatgtattcaaatatatcctcctcatT TTCTCCTCTTTCTCTAGTAAGTAGCTAGCACTATACCTAGGActGAGCTAGCCGTC AACTCCCTGCGTTTATACGCTAGTAAGCgaaaaaaaaccccgccctgtcaggggcggggtttttttttttaaG CACGCTGCTGACAACTTTTACGCCAGCTCAGCTTTTTCagcgataagtttgtctacaacgctagggt cggcaagcgttgacgtgtcacccaggttgctcgtatcgccagttgcaatcttacgcagaattcgacgcatgatcttaccagagcgagtct tcggtagggaatccgtccagtgcagtacatctggtgttgcaattggaccgatctctttgcgcacccagtctttaacttctttgtgcagttccg ctgaagggaactcaccatcgtttagcgtgatgtaagcatagatagcctgacctttgatgtcgtgcgggatacctacaatggctgcttctgc aatcttgtggtgcgcgactagggcagattcaatctctgcggtacccatacggtgtccagaaacgttcaatacgtcatccacacggcctgt gatccagtagtaaccatcttcgtcacgacgagcgccatcactggtgaagtacatgcccttgaaggttgagaagtaagtttgttcgaagc gttcatggtcaccatagaccgtacgcatctgacctggccacgaatcaagaatcacaaggttgccttcagctgcaccttcaataatgttac ccatgttatcgacaagcgcaggttgtacgccgaagaatggacgggtcgcggaacctggttttagatctgtagcgcctggtagtggagc aatcaagatgccgcctgtttcagtttgccaccaagtatcgacaatcggtgaattttcgttaccgatggtcttgtagtaccactcccatgcttc tgggttgattggctcacctaccgaacccatgatgcgaaggcttgtacgagacgtgccttcaaccgcttcattaccttttgccattagcgca cgaatcgcagttggcgcagtgtaaaggatgttcacttgatgcttatcaaccacttcgctcatacggcttgtgcttgggtagtttggcacac cctcaaacaagatggttttagcgccgttcgctagtggaccgtagataaggtacgtgtgaccagtaatccagcccacgtcagcagtacac cagaaggtttcgcccggctggtagtcgaagacgtatttaaatgtcattgcggcataaacaaggtagccaccagtggtatgcagtacgcc tttaggtttacccgtagagcctgaggtgtaaaggatgaaaagtggatcttcggctttcatctcttctggtgggcaaacgtcagaaacactt gcggttgcttcatgccaccaaacatcacggtgatcatgccattcaacatcgccaccagtacgtttaagaaccaacactttgctgatggttt tcacgtctgggttagtcagtgcttcatcaacattctttttcagtggaaccgcacgtccgccacgaacgccttcgtctgcggtgatgacaac tttagcgtctgagtcaatgatacggccagaaagtgcttctggtgagaaaccgccgaataccacagtgtggactgcgccgatgcgggta caagccagcatcgcgattgcagcttcaggtaccattggcatgtagagacaaactacgtcgcctttacgtacgccttgatctttcaatgcgt ttgagaatttacatacttctttgtgcagttcattgaacgtcagcgttttatcgtctgcagggtcatcgccttcccagattattgctacgtcgtcg ccgtgttctgctagatggcggtcaatacagtttgctgaaacgttaagtgtgccgtcttcaaaccagcggatgtcgacgtgacccgtgtcg aaagaggtgctttttacttgggtgaaaggtttaatccaatcaacgattttgccgtgctcgttccagaagccctctgggtcggttaccgactg ctgatacatggctaggtaagtttcattatccgcgtgtgtatgagttttaatgttttcttttaccgGATAAACGTGGGCTTCGC TCATTTTCTCCTCTTTCTCTAGTAAATTGtgagcgctcacaattccacacattatacgagccgatgattaattG TCAACACAGCCGAATGGaaaaccgccactttgatggcggtttttttgctttgggacgcaagttagccacacccttgatgt ggtcattaggttatttttgttaacttggttgcagaaagaacatgagctattgggagattagaccagtaaattgctgctaattgctatgaattag ctataatcttggtcgattttttaaattttgctcggttttaacaaaaaaatcgtgttgaattatcgtatatttgggaaaataaacgttccactcacg ataagggaagatagcaacataatgtctgcaaagtcacggattcttgttctaaacggaccaaaccttaacctgttaggtctgcgcgaaccg acacactatggtaataacaccttagcacagattgtgaacacgctaacggagcaagctcacaacgcaggggtggaattagaacaccta caatcaaatcgtgagtacgaactgattgaagccattcatgccgcacatggcaagattgatttcatcatcatcaatccagctgctttcactca taccagtgtagcactgcgagatgcattacttggcgtcgccatcccatttatcgaagtgcacctatcaaacgtgcacgcacgtgagccgtt tcgccatcactcatacctgtcagataaggcagaaggggtgatttgcggtctaggcgcacaaggttatgaatttgctctgtctgctgtaatt aacaaacttcagacaaagtaacccttacactctgcaatcctacggggttgtcttaatcacaagataaaagagaaagaaaagatggatatt cgtaaaatcaaaaagctaatcgagttagttgaagagtctggcattgctgagctagaaatttcagaaggtgaagaatcagtacgcatcagt cgtcacggcactgcagctgcgccagcaccagtacactacgcagcagctccagtagcagcacctgcgccagtagcagcggctccag tagcagaagcaccagcagctgaagctccagcagtacctgcgggtcaccaagttctttctccaatggttggtactttttaccgttctccaag ccctgactcgaaagcattcgtagaagtaggccaaaaagttagcgctggcgatactctatgtatcgttgaagcaatgaagatgatgaacc aaatcgaagcggacaaatctggcgttgttacagctatccttgttgaagacggccaaccagttgaattcgaccaacctctagttgtaatcg aataagcgaggcttgccttatgctagacaaagtagtcatcgcgaaccgaggtgaaatcgcacttcgtattcttcgcgcttgtaaagagtt aggtattaagaccgttgctgttcactcaacagctgaccgcgatcttaagcacgtactacttgcagatgaaaccgtatgtatcggtcctgct cgtggtatcgacagctacctaaacatcccacgtatcatttctgcagcagaagtaacgggtgcggttgctatccaccctggttacggcttct tgtcagaaaacgcagacttcgcagagcaagttgagcgtagcggcttcatctttgttggtcctaaagcagaaactatccgcttgatgggc gacaaagtatcagcgatcaacgcaatgaagaaagctggcgttccttgtgtaccaggttctgacggtccattagacaatgatgaagataa aaacaaagcgtttgctaaacgcattggttacccagtcatcatcaaagcgtcaggcggcggcggcggtcgtggtatgcgtgttgtacgtt ctgaaaaagaactagtacaagctattgctatgacacgtgcagaagcgaaagcagcgttcaacaacgac Acetate iModulon Booster B SEQ ID NO: 8 ttcaaaagcgcagccagaacaaatcatggacttagccatcttcagtgatgctcaagccgttgccggcaatcgagaattactcactcccg tcgccaaccatctcagagacacaatgaaagatcgcatgttgatcctgtccgactttacccggccagcgttgcagttttctgtgcctctgac tctgtttggtaacgtaaaaagtgccaaagatggcttagatatcaaacgcggcggtattttcccgatagtgcacggcattagaactctgtct ctcgaatatgccattgaagagaaaaatacctttgcccgcatagaagccctgagaaacaagcgaattttagaaccagaaacggcggata acttaaacgaagcactcaagttgttctttaagcttcggctcaatcagcaacttaaccaacaagaggcacagaataatatcgacctcaaac aactcgaccgaacagagcgcgacttgctgcgtcatagcctgcatgtcgtgaaaaagtttaagcagtttttaggtttccactatcagatccg tgattaagcccggtaagttgtaaaagaggtatctgggatgaattggctacaacgaaaatattggcactataaactgaaaggctcgcctta tcagtcgctattttgtgcaccagataaaaccgagctggtttctctcgactgcgaaaccaccagcctagacccgaatagagcagagctgg tgaccatcgcagccactaaaatcattgataaccgcattatcaccagccaaccattcgaagtacatttgcgagcgccacaatctctcgact ccggctcggtaaaaatccataaaattcgccatcaggatttagtcgacggcatcagcgaaaaggatgcgctactaaaattaatagacttta tcggcaatcgtccgctcgtcggctatcacatccgttacgacaaaaaaatcttagacctcgcctgtcaaagacaactgggattccctctgc ccaaccctcttattgaagtaagccagatttaccacgacaagctggagcgacatttaccgaacgcctatttcgacttaagtctggaggcga tttgcaagcaccttgagctgccaattcaagataagcacgacgctttgcaagatgcgatatccgcagcattggtatttgttcgtttaaccaaa ggcgatttaccgagtctcagcgtgccatacacataattcctgttgacctggctcccataattacttttatttcaaccctttttcggtagatccct ctctatcatcctaaagtctaattgtggaaatacgcgctgtgactgaaagttattagcacaaggtcagtcttgcgacgaccaactgtttaaaa caaatatccgttctattctttatgttttagaaaatcataatgttgatagagcgggtaaatacaacagcccgacaaagtaaagattcaaagtct tatgtagattgttgtaaggccgagaaataaaGGCACAAGGAGAGAAGCCatgagcgaagcccacgtttatccggtaa aagaaaacattaaaactcatacacacgcggataatgaaacttacctagccatgtatcagcagtcggtaaccgacccagagggcttctg gaacgagcacggcaaaatcgttgattggattaaacctttcacccaagtaaaaagcacctctttcgacacgggtcacgtcgacatccgct ggtttgaagacggcacacttaacgtttcagcaaactgtattgaccgccatctagcagaacacggcgacgacgtagcaataatctggga aggcgatgaccctgcagacgataaaacgctgacgttcaatgaactgcacaaagaagtatgtaaattctcaaacgcattgaaagatcaa ggcgtacgtaaaggcgacgtagtttgtctctacatgccaatggtacctgaagctgcaatcgcgatgctggcttgtacccgcatcggcgc agtccacactgtggtattcggcggtttctcaccagaagcactttctggccgtatcattgactcagacgctaaagttgtcatcaccgcagac gaaggcgttcgtggcggacgtgcggttccactgaaaaagaatgttgatgaagcactgactaacccagacgtgaaaaccatcagcaaa gtgttggttcttaaacgtactggtggcgatgttgaatggcatgatcaccgtgatgtttggtggcatgaagcaaccgcaagtgtttctgacg tttgcccaccagaagagatgaaagccgaagatccacttttcatcctttacacctcaggctctacgggtaaacctaaaggcgtactgcata ccactggtggctaccttgtttatgccgcaatgacatttaaatacgtcttcgactaccagccgggcgaaaccttctggtgtactgctgacgt gggctggattactggtcacacgtaccttatctacggtccactagcgaacggcgctaaaaccatcttgtttgagggtgtgccaaactaccc aagcacaagccgtatgagcgaagtggttgataagcatcaagtgaacatcctttacactgcgccaactgcgattcgtgcgctaatggcaa aaggtaatgaagcggttgaaggcacgtctcgtacaagccttcgcatcatgggttcggtaggtgagccaatcaacccagaagcatggg agtggtactacaagaccatcggtaacgaaaattcaccgattgtcgatacttggtggcaaactgaaacaggcggcatcttgattgctcca ctaccaggcgctacagatctaaaaccaggttccgcgacccgtccattcttcggcgtacaacctgcgcttgtcgataacatgggtaacatt attgaaggtgcagctgaaggcaaccttgtgattcttgattcgtggccaggtcagatgcgtacggtctatggtgaccatgaacgcttcgaa caaacttacttctcaaccttcaagggcatgtacttcaccagtgatggcgctcgtcgtgacgaagatggttactactggatcacaggccgt gtggatgacgtattgaacgtttctggacaccgtatgggtaccgcagagattgaatctgccctagtcgcgcaccacaagattgcagaagc agccattgtaggtatcccgcacgacatcaaaggtcaggctatctatgcttacatcacgctaaacgatggtgagttcccttcagcggaact gcacaaagaagttaaagactgggtgcgcaaagagatcggtccaattgcaacaccagatgtactgcactggacggattccctaccgaa gactcgctctggtaagatcatgcgtcgaattctgcgtaagattgcaactggcgatacgagcaacctgggtgacacgtcaacgcttgccg accctagcgttgtagacaaacttatcgctgaaaaagctgagctgGCGTAAGCTAACTAGCCAAAAGAGGAG AAAATGGCTGAAGCGCAAAATGatcCCCTGCTGCCGGGATACTCGTTTAATGCCCA TCTGGTGGCGGGTTTAACGCCGATTGAGGCCAACGGTTATCTCGATTTTTTTATCG ACCGACCGCTGGGAATGAAAGGTTATATTCTCAATCTCACCATTCGCGGTCAGGG GGTGGTGAAAAATCAGGGACGAGAATTTGTTTGCCGACCGGGTGATATTTTGCTG TTCCCGCCAGGAGAGATTCATCACTACGGTCGTCATCCGGAGGCTCGCGAATGGT ATCACCAGTGGGTTTACTTTCGTCCGCGCGCCTACTGGCATGAATGGCTTAACTG GCCGTCAATATTTGCCAATACGGGGTTCTTTCGCCCGGATGAAGCGCACCAGCCG CATTTCAGCGACCTGTTTGGGCAAATCATTAACGCCGGGCAAGGGGAAGGGCGC TATTCGGAGCTGCTGGCGATAAATCTGCTTGAGCAATTGTTACTGCGGCGCATGG AAGCGATTAACGAGTCGCTCCATCCACCGATGGATAATCGGGTACGCGAGGCTT GTCAGTACATCAGCGATCACCTGGCAGACAGCAATTTTGATATCGCCAGCGTCGC ACAGCATGTTTGCTTGTCGCCGTCGCGTCTGTCACATCTTTTCCGCCAGCAGTTAG GGATTAGCGTCTTAAGCTGGCGCGAGGACCAACGTATCAGCCAGGCGAAGCTGC TTTTGAGCACCACCCGGATGCCTATCGCCACCGTCGGTCGCAATGTTGGTTTTGA CGATCAACTCTATTTCTCGCGGGTATTTAAAAAATGCACCGGGGCCAGCCCGAGC GAGTTCCGTGCCGGTTGTGAAGAAAAAGTGAATGATGTAGCCGTCAAGTTGTCAT AACTACGCTCTGGCTGCTTAGgggtaaccaggcatcaaataaaacgaaaggctcagtcgaaagactgggcctttc gttttatctgttgtttgtcggtgaacgctctctactagagtcacactggctcaccttcgggtgggcctttctgCGTTTATACGCT AGTAAGCGAAAAGAAACCAATTGTCCATATTGCATCAGACATTGCCGTCACTGC GTCTTTTACTGGCTCTTCTCGCTAACCAAACCGGTAACCCCGCTTATTAAAAGCA TTCTGTAACAAAGCGGGACCAAAGCCATGACAAAAACGCGTAACAAAAGTGTCT ATAATCACGGCAGAAAAGTCCACATTGATTATTTGCACGGCGTCACACTTTGCTA TGCCATAGCATTTTTATCCATAAGATTAGCGGATCCTACCTGACGCTTTTTATcgca aCTCTCTACTgtttctcCATtactagagaaagaggagaaaATGAAACCAGTAACGTTATACGATGtc gcagagtatgccggtgtctcttatatgaccgtttcccgcgtggtgaaccaggccagccacgtttctgcgaaaacgcgggaaaaagtgg aagcggcgatggtggagctgaattacattcccaaccgcgtggcacaacaactggcgggcaaacagtcgttgctgattggcgttgcca cctccagtctggccctgcacgcgccgtcgcaaattgtcgcggcgattaaatctcgcgccgatcaactgggtgccagcgtggtggtgtc gatggtagaacgaagcggcgtcgaagcctgtaaagcggcggtgcacaatcttctcgcgcaacgcgtcagtgggctgatcattaactat ccgctggatgaccaggatgccattgctgtggaagctgcctgcactaatgttccggcgttatttcttgatgtctctgaccagacacccatca acagtattatttactcccatgaggacggtacgcgactgggcgtggagcatctggtcgcattgggtcaccagcaaatcgcgctgttagcg ggcccattaagttctgtctcggcgcgtctgcgtctggctggctggcataaatatctcactcgcaatcaaattcagccgatagcggaacg ggaaggcgactggagtgccatgtccggttttcaacaaaccatgcaaatgctgaatgagggcatcgttcccactgcgatgctggttgcc aacgatcagatggcgctgggcgcaatgcgcgccattaccgagtccgggctgcgcgttggtgcggatatctcggtagtgggatacgac gataccgaagatagctcatgttatatcccgccgttaaccaccatcaaacaggattttcgcctgctggggcaaaccagcgtggaccgctt gctgcaactctctcagggccaggcggtgaagggcaatcagctgttgccagtctcactggtgaaaagaaaaaccaccctggcgccca atacgcaaaccgcctctccccgcgcgttggccgattcattaatgcagctggcacgacaggtttcccgactggaaagcgggcagtgata aggatcctaattggtaacgaatcagacaattgacggctcgagggagtagcatagggtttgcagaatccctgcttcgtccatttgacagg cacattatgcatcgatgataagctgtcaaacatgagcagatcctctacgccggacgcatcgtggccggcatcaccggcgccacaggt gcggttgctggcgcctatatcgccgacatcaccgatggggaagatcgggctcgccacttcgggctcatgagcaaatattttatctgagc tccttttgttatcaataaaaaaggccccccgatttgggagGCCTTCAATAATTGGAGGGGGGTCTGACGCT CAGTAGcGGCGGCGAAACCCCGCCCTGTCAGGGGCGGGGTTTGCGGCGTTAAGCg atccccctatgcaagggtttattgttttctaaaatctgattaccaattagaatgaatatttcccaaatattaaataataaaacaaaaaaattgaa aaaagtgtttccaccattttttcaatttttttataatttttttaatctgttatttaaatagtttatagttaaatttacattttcattagtccattcaatattct ctccaagataactacgaactgctaacaaaattctctccctatgttctaatggagaagattcagccactgcatttcccgcaatatcttttggta tgattttacccgtgtccatagttaaaatcatacggcataaagttaatatagagttggtttcatcatcctgataattatctattaattcctctgacg aatccataatggctcttctcacatcagaaaatggaatatcaggtagtaattcctctaagtcataatttccgtatattcttttattttttcgttttgct tggtaaagcattatggttaaatctgaatttaattccttctgaggaatgtatccttgttcataaagctcttgtaaccattctccataaataaattctt gtttgggaggatgattccacggtaccatttcttgctgaataataattgttaattcaatatatcgtaagttgcttttatctcctattttttttgaaata ggtctaattttttgtataagtatttctttactttgatctgtcaatggttcagatacgacgactaaaaagtcaagatcactatttggttttagtccac tctcaactcctgatccaaacatgtaagtaccaataaggttattttttaaatgtttccgaagtatttttttcactttattaatttgttcgtatgtattca aatatatcctcctcatTTTCTCCTCTTTCTCTAGTAAGTAGCTAGCACTATACCTAGGActGA GCTAGCCGTCAACTCCCTGCGTTTATACGCTAGTAAGCgaaaaaaaaccccgccctgtcagggg cggggtttttttttttaaGCACGCTGCTGACAACTTTTAGTGATCGACAGCTTGGcccgcacctactgg tacacgtacgtgctcaaccagatctttcacttcttgtggtgtttctgctgttactctcgatactgcgaacgccacgatgaagttgaatagcgc accaattgcaccaaacgcgtttggctcaatgcctaggaaccagttgctgccccagctttccaggtaagtccaatccgcgatgaatagga tacctttgtgttggaatacgtagaacagtgtaatactgatacctgcgatcatacctgcgatagcgccttctttgttgatgctcttactgaagat acccatcatcaatgcagggaagatagaggacgctgccagaccaaaggccagtgcaacggtacctgctgcgaagcccggtgggtgt agacctaagtagcctgctaccgcaattgccactgccatggaaatccggctggcaaggagctctttcttctcggagatattcgggttaatg acccctttgattaagtcgtgagatatcgcggacgaaatcgccagtagtagacccgcagcggttgatagtgctgctgccagaccacctg ccgcaaccaatgcgattacccagtttggtagcttagcaatctctgggttagccagtaccatgatgtcgttatcaacttttagctcgttagtag ccgcatcagaagtgtactgaatgttgccatcaccgttcttatcatcgaaacctagcaggcctgttttttcccagtttttaaaccaagcagga cgctcgtcgtaagctaggtgttggccaggcgccgggttaactgtgttcattaggtttaggcgcgccattgctgatactgcaggtgctgtt gtgtaaagaatagcaatgaagattagtgcccaacctgctgatgtacgtgcgtcacgaaccttaggtaccgtaaagaaacggatgattac gtgtggcagaccagccgtaccgatcatcagtgacatggtgtaagcaaacatgtttagtgtatcaccacgaacttgagtggtgtattcact aaagccgagttcggtcaccacctgatcgagccgatcgagtagatagacatcggttcccgccatggtactaccaaggccaatttgcgga attgggttaccagtcaattgcagagagataaatactgctgggatagtgtaagcaagaatgagtacacaatattgagcaatctgcgtgtat gtgatgcctttcatgccgcctagtactgcgtacataaatacgatacacataccaatcagtagacctgttgcgtaatctacttctaagaaacg accgaatgcaacacctacgcctttcatctgaccgataacgtatgttactgatgcgatgatcagacacacaaccgctacgatacgcgctgt cttagagtagaaacgctcaccgacaaactctggcacggtaaatttaccgaacttacgtaggtaaggagcaagaagcagagcaagcag aacgtaaccgccagtccaacccattaggaatacagaaccaccgtatcccatgaaggcaatcagacccgccatcgaaataaacgacg ccgctgacatccagtctgctgctgtcgccataccgttcgcgattgggtttacaccgccacccgcaacatagaactctttggttgggcctg cgcgcgcccaaatcgcaatcccgatgtacaggatgaaggttgcgccaacaaccaggtaggtaatagttttcaaatccatctgagtgact ccttactcatccacgccgaattcacggtctattttgcgcatcttccacgcgtagtagaaaataatacccaggaaggcgtagatggagcct tgttgagcaaaccagaaaccaagtttgtatccacctagttgaaattgatttagtacatcgacaaataaaatgccgcagccgaaagacact aagaaccagatgaccattaggccaatcattagcttcacgttcttgtcccagtaggctttggctctttcttcgctttcaaacgcCATTGC CGTCTCCTTACGATTTACGCCAGCTCAGCTTTTTCagcgataagtttgtctacaacgctagggtcggca agcgttgacgtgtcacccaggttgctcgtatcgccagttgcaatcttacgcagaattcgacgcatgatcttaccagagcgagtcttcggt agggaatccgtccagtgcagtacatctggtgttgcaattggaccgatctctttgcgcacccagtctttaacttctttgtgcagttccgctga agggaactcaccatcgtttagcgtgatgtaagcatagatagcctgacctttgatgtcgtgcgggatacctacaatggctgcttctgcaatc ttgtggtgcgcgactagggcagattcaatctctgcggtacccatacggtgtccagaaacgttcaatacgtcatccacacggcctgtgatc cagtagtaaccatcttcgtcacgacgagcgccatcactggtgaagtacatgcccttgaaggttgagaagtaagtttgttcgaagcgttca tggtcaccatagaccgtacgcatctgacctggccacgaatcaagaatcacaaggttgccttcagctgcaccttcaataatgttacccatg ttatcgacaagcgcaggttgtacgccgaagaatggacgggtcgcggaacctggttttagatctgtagcgcctggtagtggagcaatca agatgccgcctgtttcagtttgccaccaagtatcgacaatcggtgaattttcgttaccgatggtcttgtagtaccactcccatgcttctgggt tgattggctcacctaccgaacccatgatgcgaaggcttgtacgagacgtgccttcaaccgcttcattaccttttgccattagcgcacgaat cgcagttggcgcagtgtaaaggatgttcacttgatgcttatcaaccacttcgctcatacggcttgtgcttgggtagtttggcacaccctca aacaagatggttttagcgccgttcgctagtggaccgtagataaggtacgtgtgaccagtaatccagcccacgtcagcagtacaccaga aggtttcgcccggctggtagtcgaagacgtatttaaatgtcattgcggcataaacaaggtagccaccagtggtatgcagtacgcctttag gtttacccgtagagcctgaggtgtaaaggatgaaaagtggatcttcggctttcatctcttctggtgggcaaacgtcagaaacacttgcgg ttgcttcatgccaccaaacatcacggtgatcatgccattcaacatcgccaccagtacgtttaagaaccaacactttgctgatggttttcacg tctgggttagtcagtgcttcatcaacattctttttcagtggaaccgcacgtccgccacgaacgccttcgtctgcggtgatgacaactttagc gtctgagtcaatgatacggccagaaagtgcttctggtgagaaaccgccgaataccacagtgtggactgcgccgatgcgggtacaagc cagcatcgcgattgcagcttcaggtaccattggcatgtagagacaaactacgtcgcctttacgtacgccttgatctttcaatgcgtttgag aatttacatacttctttgtgcagttcattgaacgtcagcgttttatcgtctgcagggtcatcgccttcccagattattgctacgtcgtcgccgt gttctgctagatggcggtcaatacagtttgctgaaacgttaagtgtgccgtcttcaaaccagcggatgtcgacgtgacccgtgtcgaaag aggtgctttttacttgggtgaaaggtttaatccaatcaacgattttgccgtgctcgttccagaagccctctgggtcggttaccgactgctga tacatggctaggtaagtttcattatccgcgtgtgtatgagttttaatgttttcttttaccgGATAAACGTGGGCTTCGCTC ATTTTCTCCTCTTTCTCTAGTAAATTGtgagcgctcacaattccacacattatacgagccgatgattaattGTC AACACAGCCGAATGGAAAACCGCCACTTTGATGgcggtttttttgctttgggacgcaagttagccacac ccttgatgtggtcattaggttatttttgttaacttggttgcagaaagaacatgagctattgggagattagaccagtaaattgctgctaattgct atgaattagctataatcttggtcgattttttaaattttgctcggttttaacaaaaaaatcgtgttgaattatcgtatatttgggaaaataaacgttc cactcacgataagggaagatagcaacataatgtctgcaaagtcacggattcttgttctaaacggaccaaaccttaacctgttaggtctgc gcgaaccgacacactatggtaataacaccttagcacagattgtgaacacgctaacggagcaagctcacaacgcaggggtggaattag aacacctacaatcaaatcgtgagtacgaactgattgaagccattcatgccgcacatggcaagattgatttcatcatcatcaatccagctgc tttcactcataccagtgtagcactgcgagatgcattacttggcgtcgccatcccatttatcgaagtgcacctatcaaacgtgcacgcacgt gagccgtttcgccatcactcatacctgtcagataaggcagaaggggtgatttgcggtctaggcgcacaaggttatgaatttgctctgtct gctgtaattaacaaacttcagacaaagtaacccttacactctgcaatcctacggggttgtcttaatcacaagataaaagagaaagaaaag atggatattcgtaaaatcaaaaagctaatcgagttagttgaagagtctggcattgctgagctagaaatttcagaaggtgaagaatcagta cgcatcagtcgtcacggcactgcagctgcgccagcaccagtacactacgcagcagctccagtagcagcacctgcgccagtagcagc ggctccagtagcagaagcaccagcagctgaagctccagcagtacctgcgggtcaccaagttctttctccaatggttggtactttttaccg ttctccaagccctgactcgaaagcattcgtagaagtaggccaaaaagttagcgctggcgatactctatgtatcgttgaagcaatgaagat gatgaaccaaatcgaagcggacaaatctggcgttgttacagctatccttgttgaagacggccaaccagttgaattcgaccaacctctagt tgtaatcgaataagcgaggcttgccttatgctagacaaagtagtcatcgcgaaccgaggtgaaatcgcacttcgtattcttcgcgcttgta aagagttaggtattaagaccgttgctgttcactcaacagctgaccgcgatcttaagcacgtactacttgcagatgaaaccgtatgtatcgg tcctgctcgtggtatcgacagctacctaaacatcccacgtatcatttctgcagcagaagtaacgggtgcggttgctatccaccctggttac ggcttcttgtcagaaaacgcagacttcgcagagcaagttgagcgtagcggcttcatctttgttggtcctaaagcagaaactatccgcttg atgggcgacaaagtatcagcgatcaacgcaatgaagaaagctggcgttccttgtgtaccaggttctgacggtccattagacaatgatga agataaaaacaaagcgtttgctaaacgcattggttacccagtcatcatcaaagcgtcaggcggcggcggcggtcgtggtatgcgtgtt gtacgttctgaaaaagaactagtacaagctattgctatgacacgtgcagaagcgaaagcagcgttcaacaacgac

[0003] References For Experiment No.1: 1. J. D. Wang, P. A. Levin, Metabolism, cell growth and the bacterial cell cycle. Nat. Rev. Microbiol.7, 822–827 (2009). 2. D. B. Roszak, R. R. Colwell, Survival strategies of bacteria in the natural environment. Microbiol. Rev.51, 365–379 (1987). 3. D. W. Erickson, et al., A global resource allocation strategy governs growth transition kinetics of Escherichia coli. Nature 551, 119–123 (2017). 4. M. Basan, et al., A universal trade-off between growth and lag in fluctuating environments. Nature 584, 470–474 (2020). 5. R. Balakrishnan, R. T. de Silva, T. Hwa, J. Cremer, Suboptimal resource allocation in changing environments constrains response and growth in bacteria. Mol. Syst. Biol.17, e10597 (2021). 6. S. E. Irving, N. R. Choudhury, R. M. Corrigan, The stringent response and physiological roles of (pp)pGpp in bacteria. Nat. Rev. Microbiol.19, 256–271 (2021). 7. W. Eisenreich, T. Dandekar, J. Heesemann, W. Goebel, Carbon metabolism of intracellular bacterial pathogens and possible links to virulence. Nat. Rev. Microbiol.8, 401–412 (2010). 8. J. L. Martínez, F. Rojo, Metabolic regulation of antibiotic resistance. FEMS Microbiol. Rev.35, 768–789 (2011). 9. D. Bajic, A. Sanchez, The ecology and evolution of microbial metabolic strategies. Curr. Opin. Biotechnol.62, 123–128 (2020). 10. N. J. Croucher, N. R. Thomson, Studying bacterial transcriptomes using RNA-seq. Curr. Opin. Microbiol.13, 619–624 (2010). 11. N. J. Croucher, et al., A simple method for directional transcriptome sequencing using Illumina technology. Nucleic Acids Res.37, e148 (2009). 12. T. Barrett, et al., NCBI GEO: archive for functional genomics data sets--10 years on. Nucleic Acids Res.39, D1005–10 (2011). 13. R. Leinonen, H. Sugawara, M. Shumway, International Nucleotide Sequence Database Collaboration, The sequence read archive. Nucleic Acids Res.39, D19–21 (2011). 136 4916-8982-5303.1 14. W. Saelens, R. Cannoodt, Y. Saeys, A comprehensive evaluation of module detection methods for gene expression data. Nat. Commun.9, 1–12 (2018). 15. A. V. Sastry, et al., The Escherichia coli transcriptome mostly consists of independently regulated modules. Nat. Commun.10, 1–14 (2019). 16. K. Rychel, et al., iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning. Nucleic Acids Res.49, D112–D120 (2021). 17. C. R. Lamoureux, et al., A multi-scale expression and regulation knowledge base for Escherichia coli. Nucleic Acids Res.51, 10176–10193 (2023). 18. K. S. Choudhary, et al., Elucidation of Regulatory Modes for Five Two-Component Systems in Escherichia coli Reveals Novel Relationships. mSystems 5 (2020). 19. A. Rajput, et al., Advanced transcriptomic analysis reveals the role of efflux pumps and media composition in antibiotic responses of Pseudomonas aeruginosa. Nucleic Acids Res.50, 9675–9688 (2022). 20. A. V. Sastry, et al., Machine Learning of Bacterial Transcriptomes Reveals Responses Underlying Differential Antibiotic Susceptibility. mSphere 6, e0044321 (2021). 21. K. Rychel, et al., Laboratory evolution, transcriptomics, and modeling reveal mechanisms of paraquat tolerance. Cell Rep.42, 113105 (2023). 22. S. Poudel, et al., Revealing 29 sets of independently modulated genes in Staphylococcus aureus, their regulators, and role in key physiological response. Proceedings of the National Academy of Sciences 117, 17228–17239 (2020). 23. J. Shin, K. Rychel, B. O. Palsson, Systems biology of competency in Vibrio natriegens is revealed by applying novel data analytics to the transcriptome. Cell Rep.42, 112619 (2023). 24. A. V. Sastry, et al., Mining all publicly available expression data to compute dynamic microbial transcriptional regulatory networks. bioRxiv, 2021.07.01.450581 (2021). 25. L. T. Rosa, M. E. Bianconi, G. H. Thomas, D. J. Kelly, Tripartite ATP-Independent Periplasmic (TRAP) Transporters and Tripartite Tricarboxylate Transporters (TTT): From Uptake to Pathogenicity. Front. Cell. Infect. Microbiol.8, 33 (2018). 26. C. Mulligan, M. Fischer, G. H. Thomas, Tripartite ATP-independent periplasmic (TRAP) transporters in bacteria and archaea. FEMS Microbiol. Rev.35, 68–86 (2011). 137 4916-8982-5303.1 27. J. Herrou, et al., Structure-based mechanism of ligand binding for periplasmic solute- binding protein of the Bug family. J. Mol. Biol.373, 954–964 (2007). 28. P. Rucktooa, et al., Crystal structures of two Bordetella pertussis periplasmic receptors contribute to defining a novel pyroglutamic acid binding DctP subfamily. J. Mol. Biol.370, 93– 106 (2007). 29. Y. Soma, et al., Trace impurities in sodium phosphate influences the physiological activity of Escherichia coli in M9 minimal medium. Sci. Rep.13, 17396 (2023). 30. F. M. Ferreira, et al., Structural analysis of N-acetylglucosamine-6-phosphate deacetylase apoenzyme from Escherichia coli. J. Mol. Biol.359, 308–321 (2006). 31. S. N. Ruzheinikov, et al., Glycerol dehydrogenase. structure, specificity, and mechanism of a family III polyol dehydrogenase. Structure 9, 789–802 (2001). 32. J. Schothorst, G. V. Zeebroeck, J. M. Thevelein, Identification of Ftr1 and Zrt1 as iron and zinc micronutrient transceptors for activation of the PKA pathway in Saccharomyces cerevisiae. Microb. Cell Fact.4, 74–89 (2017). 33. B. Nocek, et al., Structural studies of ROK fructokinase YdhR from Bacillus subtilis: insights into substrate binding and fructose specificity. J. Mol. Biol.406, 325–342 (2011). 34. K. Wangpaiboon, T. Charoenwongpaiboon, M. Klaewkla, R. A. Field, P. Panpetch, Cassava pullulanase and its synergistic debranching action with isoamylase 3 in starch catabolism. Front. Plant Sci.14, 1114215 (2023). 35. W. Wei, J. Ma, S.-Q. Chen, X.-H. Cai, D.-Z. Wei, A novel cold-adapted type I pullulanase of Paenibacillus polymyxa Nws-pp2: in vivo functional expression and biochemical characterization of glucans hydrolyzates analysis. BMC Biotechnol.15, 96 (2015). 36. D. Osman, et al., Fine control of metal concentrations is necessary for cells to discern zinc from cobalt. Nat. Commun.8, 1884 (2017). 37. H. G. Lim, et al., Machine-learning from Pseudomonas putida KT2440 transcriptomes reveals its transcriptional regulatory network. Metab. Eng.72, 297–310 (2022). 38. C. E. Harty, et al., Ethanol Stimulates Trehalose Production through a SpoT-DksA-AlgU- Dependent Pathway in Pseudomonas aeruginosa. J. Bacteriol.201 (2019). 138 4916-8982-5303.1 39. N. Derdouri, N. Ginet, Y. Denis, M. Ansaldi, A. Battesti, The prophage-encoded transcriptional regulator AppY has pleiotropic effects on E. coli physiology. PLoS Genet.19, e1010672 (2023). 40. J. N. Carey, et al., Phage integration alters the respiratory strategy of its host. Elife 8 (2019). 41. W. Voth, U. Jakob, Stress-Activated Chaperones: A First Line of Defense. Trends Biochem. Sci.42, 899–913 (2017). 42. D. Hidalgo, C. A. Martínez-Ortiz, B. O. Palsson, J. I. Jiménez, J. Utrilla, Regulatory perturbations of ribosome allocation in bacteria reshape the growth proteome with a trade-off in adaptation capacity. iScience 25, 103879 (2022). 43. E. Bosdriesz, D. Molenaar, B. Teusink, F. J. Bruggeman, How fast-growing bacteria robustly tune their ribosome concentration to approximate growth-rate maximization. FEBS J. 282, 2029–2044 (2015). 44. P. Millard, T. Gosselin-Monplaisir, S. Uttenweiler-Joseph, B. Enjalbert, Acetate is a beneficial nutrient for E. coli at low glycolytic flux. EMBO J.42, e113079 (2023). 45. P. Millard, B. Enjalbert, S. Uttenweiler-Joseph, J.-C. Portais, F. Létisse, Control and regulation of acetate overflow in Escherichia coli. Elife 10 (2021). 46. Wolfe Alan J., The Acetate Switch. Microbiol. Mol. Biol. Rev.69, 12–50 (2005). 47. K. Martínez-Gómez, et al., New insights into Escherichia coli metabolism: carbon scavenging, acetate metabolism and carbon recycling responses during growth on glycerol. Microb. Cell Fact.11, 46 (2012). 48. G. J. Bentley, et al., Engineering glucose metabolism for enhanced muconic acid production in Pseudomonas putida KT2440. Metab. Eng.59, 64–75 (2020). 49. D. E. Chang, S. Shin, J. S. Rhee, J. G. Pan, Acetate metabolism in a pta mutant of Escherichia coli W3110: importance of maintaining acetyl coenzyme A flux for growth and survival. J. Bacteriol.181, 6656–6663 (1999). 50. W. R. Farmer, J. C. Liao, Reduction of aerobic acetate production by Escherichia coli. Appl. Environ. Microbiol.63, 3205–3210 (1997). 139 4916-8982-5303.1 51. S. Pinhal, D. Ropers, J. Geiselmann, H. de Jong, Acetate Metabolism and the Inhibition of Bacterial Growth by Acetate. J. Bacteriol.201 (2019). 52. F. Thoma, B. Blombach, Metabolic engineering of Vibrio natriegens. Essays Biochem. 65, 381–392 (2021). 53. J. Soini, K. Ukkonen, P. Neubauer, High cell density media for Escherichia coli are generally designed for aerobic cultivations - consequences for large-scale bioprocesses and shake flask cultures. Microb. Cell Fact.7, 26 (2008). 54. T. Shang, C. M. Fang, C. E. Ong, Y. Pan, Heterologous Expression of Recombinant Human Cytochrome P450 (CYP) in Escherichia coli: N-Terminal Modification, Expression, Isolation, Purification, and Reconstitution. BioTech (Basel) 12 (2023). 55. G. L. Rosano, E. A. Ceccarelli, Recombinant protein expression in Escherichia coli: advances and challenges. Front. Microbiol.5, 172 (2014). 56. T. C. Stadtman, Selenoproteins--tracing the role of a trace element in protein function. PLoS Biol.3, e421 (2005). 57. P. R. Nimbalkar, et al., Role of Trace Elements as Cofactor: An Efficient Strategy toward Enhanced Biobutanol Production. ACS Sustain Chem Eng 6, 9304–9313 (2018). 58. C. You, et al., Coordination of bacterial proteome with metabolism by cyclic AMP signalling. Nature 500, 301–306 (2013). 59. A. Kolb, S. Busby, H. Buc, S. Garges, S. Adhya, Transcriptional regulation by cAMP and its receptor protein. Annu. Rev. Biochem.62, 749–795 (1993). 60. M. Zhu, X. Dai, Stringent response ensures the timely adaptation of bacterial growth to nutrient downshift. Nat. Commun.14, 467 (2023). 61. J. Xu, et al., Vibrio natriegens as a pET-Compatible Expression Host Complementary to Escherichia coli. Front. Microbiol.12, 627181 (2021). 62. R. G. Eagon, Pseudomonas natriegens, a marine bacterium with a generation time of less than 10 minutes. J. Bacteriol.83, 736–737 (1962). 63. E. Hoffart, et al., High Substrate Uptake Rates Empower Vibrio natriegens as Production Host for Industrial Biotechnology. Appl. Environ. Microbiol.83, e01614–17 (2017). 140 4916-8982-5303.1 64. Y.-F. Sui, et al., Engineering cofactor metabolism for improved protein and glucoamylase production in Aspergillus niger. Microb. Cell Fact.19, 198 (2020). 65. B. Moritz, K. Striegel, A. A. de Graaf, H. Sahm, Changes of pentose phosphate pathway flux in vivo in Corynebacterium glutamicum during leucine-limited batch cultivation as determined from intracellular metabolite concentration measurements. Metab. Eng.4, 295–305 (2002). 66. F. C. Fang, E. R. Frawley, T. Tapscott, A. Vázquez-Torres, Bacterial Stress Responses during Host Infection. Cell Host Microbe 20, 133–143 (2016). 67. S. Jozefczuk, et al., Metabolomic and transcriptomic stress response of Escherichia coli. Mol. Syst. Biol.6, 364 (2010). 68. J. Dawan, J. Ahn, Bacterial Stress Responses as Potential Targets in Overcoming Antibiotic Resistance. Microorganisms 10 (2022). 69. A. Anand, et al., OxyR Is a Convergent Target for Mutations Acquired during Adaptation to Oxidative Stress-Prone Metabolic States. Mol. Biol. Evol.37, 660–667 (2020). 70. D. Choe, et al., RiboRid: A low cost, advanced, and ultra-efficient method to remove ribosomal RNA for bacterial transcriptomics. PLoS Genet.17, e1009821 (2021). 71. H. H. Lee, et al., Functional genomics of the rapidly replicating bacterium Vibrio natriegens by CRISPRi. Nat Microbiol 4, 1105–1113 (2019). 72. B. Langmead, C. Trapnell, M. Pop, S. L. Salzberg, Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol.10, R25 (2009). 73. L. Wang, S. Wang, W. Li, RSeQC: quality control of RNA-seq experiments. Bioinformatics 28, 2184–2185 (2012). 74. Y. Liao, G. K. Smyth, W. Shi, featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30, 923–930 (2014). 75. P. Ewels, M. Magnusson, S. Lundin, M. Käller, MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics 32, 3047–3048 (2016). 76. J. Tian, et al., Discovery and remodeling of Vibrio natriegens as a microbial platform for efficient formic acid biorefinery. Nat. Commun.14, 7758 (2023). 141 4916-8982-5303.1 77. A. Hyvärinen, Fast and robust fixed-point algorithms for independent component analysis. IEEE Trans. Neural Netw.10, 626–634 (1999). 78. F. Pedregosa, et al., Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res.12, 2825–2830 (2011). 79. Ester, Kriegel, Sander, Xu, A density-based algorithm for discovering clusters in large spatial databases with noise. KDD (1996). 80. J. L. McConn, C. R. Lamoureux, S. Poudel, B. O. Palsson, A. V. Sastry, Optimal dimensionality selection for independent component analysis of transcriptomic data. BMC Bioinformatics 22, 584 (2021). 81. M. Kanehisa, M. Furumichi, M. Tanabe, Y. Sato, K. Morishima, KEGG: new perspectives on genomes, pathways, diseases and drugs. Nucleic Acids Res.45, D353–D361 (2017). 82. J. Huerta-Cepas, et al., eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res.47, D309–D314 (2019). 83. UniProt Consortium, UniProt: the universal protein knowledgebase in 2021. Nucleic Acids Res.49, D480–D489 (2021). 84. P. D. Karp, et al., The BioCyc collection of microbial genomes and metabolic pathways. Brief. Bioinform.20, 1085–1093 (2019). 85. Gene Ontology Consortium, The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Res.49, D325–D334 (2021). 86. N. T. Wirth, J. Funk, S. Donati, P. I. Nikel, QurvE: user-friendly software for the analysis of biological growth and fluorescence data. Nat. Protoc.18, 2401–2403 (2023). 87. B. Jana, K. Keppel, D. Salomon, Engineering a customizable antibacterial T6SS-based platform in Vibrio natriegens. EMBO Rep.22, e53681 (2021). 88. R. D. Gietz, R. H. Schiestl, High-efficiency yeast transformation using the LiAc / SS carrier DNA / PEG method. Nat. Protoc.2, 31–34 (2007). 89. , A., Chen, K., Catoiu, E., Sastry, A. V., Olson, C. A., Sandberg, T. E., Seif, Y., Xu, S., Szubin, R., Yang, L., Feist, A. M., & Palsson, B. O. (2020). OxyR Is a Convergent Target for 142 4916-8982-5303.1 Mutations Acquired during Adaptation to Oxidative Stress-Prone Metabolic States. Molecular Biology and Evolution, 37(3), 660–667. https: / / doi.org / 10.1093 / molbev / msz251 90. Erian, A. M., Freitag, P., Gibisch, M., & Pflugl, S. (2020). High rate 2,3-butanediol production with Vibrio natriegens. Bioresource Technology Reports, 10, 100408. https: / / doi.org / 10.1016 / j.biteb.2020.100408 91. Lim, H. G., Rychel, K., Sastry, A. V., Bentley, G. J., Mueller, J., Schindel, H. S., Larsen, P. E., Laible, P. D., Guss, A. M., Niu, W., Johnson, C. W., Beckham, G. T., Feist, A. M., & Palsson, B. O. (2022). Machine-learning from Pseudomonas putida KT2440 transcriptomes reveals its transcriptional regulatory network. Metabolic Engineering, 72, 297–310. https: / / doi.org / 10.1016 / j.ymben.2022.04.004 92. Rychel, K., Decker, K., Sastry, A. V., Phaneuf, P. V., Poudel, S., & Palsson, B. O. (2021). iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning. Nucleic Acids Research, 49(D1), D112–D120. https: / / doi.org / 10.1093 / nar / gkaa810 93. Sastry, A. V., Poudel, S., Rychel, K., Yoo, R., Lamoureux, C. R., Chauhan, S., Haiman, Z. B., Al Bulushi, T., Seif, Y., & Palsson, B. O. (2021). Mining all publicly available expression data to compute dynamic microbial transcriptional regulatory networks. In bioRxiv (p. 2021.07.01.450581). https: / / doi.org / 10.1101 / 2021.07.01.450581 94. Shin, J., Rychel, K., & Palsson, B. O. (2023). Systems biology of competency in Vibrio natriegens is revealed by applying novel data analytics to the transcriptome. Cell Reports, 42(6), 112619. https: / / doi.org / 10.1016 / j.celrep.2023.112619 95. Thoma, F., & Blombach, B. (2021). Metabolic engineering of Vibrio natriegens. Essays in Biochemistry, 65(2), 381–392. https: / / doi.org / 10.1042 / EBC20200135 96. Tian, J., Deng, W., Zhang, Z., Xu, J., Yang, G., Zhao, G., Yang, S., Jiang, W., & Gu, Y. (2023). Discovery and remodeling of Vibrio natriegens as a microbial platform for efficient formic acid biorefinery. Nature Communications, 14(1), 7758. https: / / doi.org / 10.1038 / s41467- 023- 43631-2 143 4916-8982-5303.1 References for Experiment No.2 Choe, Donghui, Richard Szubin, Saugat Poudel, Anand Sastry, Yoseb Song, Yongjae Lee, Suhyung Cho, Bernhard Palsson, and Byung-Kwan Cho.2021. “RiboRid: A low cost, advanced, and ultra-efficient method to remove ribosomal RNA for bacterial transcriptomics.” PLoS Genetics 17 (9): e1009821. Dalldorf, Christopher, Kevin Rychel, Richard Szubin, Ying Hefner, Arjun Patel, Daniel C. Zielinski, and Bernhard O. Palsson.2024. “The Hallmarks of a Tradeoff in Transcriptomes That Balances Stress and Growth Functions.” MSystems 9 (7): e0030524. Ewels, Philip, Måns Magnusson, Sverker Lundin, and Max Käller.2016. “MultiQC: summarize analysis results for multiple tools and samples in a single report.” Bioinformatics 32 (19): 3047–48. Lamoureux, Cameron R., Katherine T. Decker, Anand V. Sastry, Kevin Rychel, Ye Gao, John Luke McConn, Daniel C. Zielinski, and Bernhard O. Palsson.2023. “A Multi-Scale Expression and Regulation Knowledge Base for Escherichia Coli.” Nucleic Acids Research 51 (19): 10176–93. Langmead, Ben, Cole Trapnell, Mihai Pop, and Steven L. Salzberg.2009. “Ultrafast and Memory-Efficient Alignment of Short DNA Sequences to the Human Genome.” Genome Biology 10 (3): R25. Liao, Yang, Gordon K. Smyth, and Wei Shi.2014. “FeatureCounts: An Efficient General Purpose Program for Assigning Sequence Reads to Genomic Features.” Bioinformatics 30 (7): 923–30. Love, Michael I., Wolfgang Huber, and Simon Anders.2014. “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2.” Genome Biology 15 (12): 550. Rychel, Kevin, Katherine Decker, Anand V. Sastry, Patrick V. Phaneuf, Saugat Poudel, and Bernhard O. Palsson.2021. “iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning.” Nucleic Acids Research 49 (D1): D112–20. Sastry, Anand V., Ye Gao, Richard Szubin, Ying Hefner, Sibei Xu, Donghyuk Kim, Kumari Sonal Choudhary, Laurence Yang, Zachary A. King, and Bernhard O. Palsson.2019. 144 4916-8982-5303.1 “The Escherichia coli transcriptome mostly consists of independently regulated modules.” Nature Communications 10 (1): 5536. Sastry, Anand V., Yuan Yuan, Saugat Poudel, Kevin Rychel, Reo Yoo, Cameron R. Lamoureux, Gaoyuan Li, et al.2024. “IModulonMiner and PyModulon: Software for Unsupervised Mining of Gene Expression Compendia.” PLoS Computational Biology 20 (10): e1012546. Shin, Jongoh, Kevin Rychel, and Bernhard O. Palsson.2023. “Systems Biology of Competency in Vibrio Natriegens Is Revealed by Applying Novel Data Analytics to the Transcriptome.” Cell Reports 42 (6): 112619. Shin, Jongoh, Daniel C. Zielinski, and Bernhard O. Palsson.2024. “Modulating Bacterial Function Utilizing A Knowledge Base of Transcriptional Regulatory Modules.” Nucleic Acids Research, August. https: / / doi.org / 10.1093 / nar / gkae742. Wang, Liguo, Shengqin Wang, and Wei Li.2012. “RSeQC: Quality Control of RNA-Seq Experiments.” Bioinformatics 28 (16): 2184–85. Wirth, Nicolas T., Jonathan Funk, Stefano Donati, and Pablo I. Nikel.2023. “QurvE: User- Friendly Software for the Analysis of Biological Growth and Fluorescence Data.” Nature Protocols 18 (8): 2401–3. 145 4916-8982-5303.1

Claims

WHAT IS CLAIMED IS:

1. A method comprising: growing a first plurality of cells in a media formulation, on a substrate, or on a substrate in a media formulation; isolating RNA from the plurality of cells; detecting by an RNA expression profile upregulation or downregulation of an iModulon gene cluster; iteratively modifying a media formulation or a substrate to downregulate or upregulate the iModulon gene cluster; growing a second population of cells in the modified media formulation or on the modified substrate; isolating RNA from the second plurality of cells; and detecting optimized expression of the iModulon gene cluster from the second plurality of cells.

2. The method of claim 1, wherein the iModulon gene cluster is selected from the group of a metabolic iModulon gene cluster, a trace element iModulon gene cluster, a stress-related iModulon gene cluster or an Acetate iModulon gene cluster.

3. A method for preparing an isolated cell or population of cells with optimized media or substrate compatibility, comprising genetically modifying an isolated cell or a population of cells to upregulate or downregulate an iModulon gene cluster, optionally wherein the iModulon gene cluster is selected from the group of a metabolic iModulon gene cluster, a trace element iModulon gene cluster, a stress-related iModulon gene cluster or an Acetate iModulon gene cluster.

4. The method of any one of claims 1-3, wherein the iModulon gene cluster is a stress- related iModulon gene cluster that comprises the OxyR, RpoE, Low-osmolarity, Macrolide- efflux, PhrR, CSP, or Ectoine iModulon gene cluster. 146 4916-8982-5303.

15. The method of any one of claims 1-3, wherein the iModulon gene cluster is a trace element iModulon gene cluster that comprises a Zur, TPP, CueR, Fur-1, Fur-2, Enterobactin, or thiosulfate iModulon gene cluster.

6. The method of any one of claims 1-3, wherein the iModulon gene cluster is a metabolic iModulon gene cluster that comprises the NarP-1, TonB, FNR, CytochromeO, NarP-2, GABA, Alr, HutC, MetJ, LiuR, ArgR, Cysteine, Prophage-1, Prophage-2, β-ox, Flagellar-1, Flagellar-2, Chaperone, Biofilm, Ribosomal proteins, TfoX, Qst, UC-1, HapR, Cellobiose, TRAP, Formate, TreR, Arabinose, AGL, NagC, GntR, Maltose, FruR, Glycerol, TTT, Acetate, GalR, Rhamnose, Null-1, Null-2, Null-3, UC-2, UC-3, UC-4, UC-5, UC-6, UC-7, UC-8, or UC-9 iModulon gene cluster.

7. The method of any one of claims 1-3, wherein the media formulation comprises acetate, ethanol, cellobiose, galactose, glycogen, or GlcN.

8. The method of any one of claims 1-7, wherein the isolated cell or population of cells is a prokaryotic cell or a eukaryotic cell or a population of each thereof.

9. The method of claim 8, wherein the prokaryotic cell is a bacterial cell, optionally selected from an E. coli, P. Putida, Corynebacterium glutamicum, or Vibrio natriegens.

10. A method for optimizing media composition for microbial growth, comprising: obtaining transcriptomic data from a plurality of microbial samples grown in media containing different conditions, the different conditions each corresponding to a type and concentration of a nutrient supplement added to the media; determining one or more iModulons by applying independent component analysis to the transcriptomic data, each iModulon comprising at least one gene with a known function and at least one gene with an unknown function; 147 4916-8982-5303.1calculating a nutrient score for each condition based on a weighted sum of iModulon activity levels of the one or more iModulons, wherein the weights are determined based on iModulon activity levels and iModulon growth rate in the media containing the nutrient supplements of the conditions; selecting a condition identifying a type and concentration of the nutrient supplement of the condition based at least on the condition corresponding to a positive nutrient score; and supplementing a growth medium with the type and concentration of the nutrient supplement of the condition.

11. The method of claim 10, wherein the growth medium is of the same type as the media in which the plurality of microbial samples were grown, the method further comprising: obtaining additional transcriptomic data from microbial samples grown in the supplemented growth medium; calculating a second nutrient score for additional nutrient supplements; and selecting a second nutrient supplement based on the second nutrient score.

12. The method of claim 11, further comprising: supplementing a second growth medium with the type and concentration of the nutrient supplement of the condition and the second nutrient supplement.

13. The method of claim 10, further comprising: categorizing the iModulons into functional groups comprising catabolism, stress responses, and translation.

14. The method of claim 13, wherein calculating the nutrient score for each condition comprises: identifying each iModulon categorized into the stress responses group; and calculating the nutrient score based only on iModulon activity by the identified iModulons categorized into the stress responses group. 148 4916-8982-5303.

115. The method of claim 10, wherein calculating the nutrient score comprises: assigning positive weights to growth-promoting iModulons; and assigning negative weights to stress-related iModulons.

16. The method of claim 10, wherein the nutrient supplement is selected from the group consisting of L-methionine, L-cysteine, and D-malic acid.

17. The method of claim 10, further comprising: ranking each of the conditions based on the nutrient scores calculated for the conditions, wherein selecting the condition is based at least on the rankings of the conditions.

18. The method of claim 10, comprising: generating a matrix based on expression profiles for individual conditions in the transcriptomic data; applying independent component analysis to the matrix to decompose the matrix into at least a mixing matrix containing a different column for each iModulon and a different row for each condition, wherein values in a column for an iModulon indicate an activity level of the iModulon in the conditions; and determining the weight for each nutrient supplement based at least on the values in the column for the iModulon indicating the activity level of the iModulon in the conditions.

19. The method of claim 18, wherein generating the matrix comprises, for each condition: identifying an expression level of each gene grown under the condition; and generating a vector with each identified expression level for the condition.

20. A method for determining an RNA expression profile of an iModulon gene cluster from a media formulation or a substrate for cellular growth, the method comprising: 149 4916-8982-5303.1determining one or more iModulons of a first organism in a cell culture medium by applying independent component analysis (ICA) on gene expression data of the cultured first organism, the one or more iModulons comprising one or more genes of an unidentified function; inserting the one or more iModulons into a second organism to create a strain of a phenotypic expression from the one or more iModulons; determining the RNA expression profile of the iModulon gene cluster in the second organism.

21. The method of claim 20, further comprising iteratively modifying the cell culture medium to downregulate or upregulate the iModulon gene cluster, growing a third population of cells in the modified media formulation or on the modified substrate, determining the RNA expression from the third cell culture, and detecting RNA expression profile of the iModulon gene cluster in the third cell culture.

22. The method of claim 20 or 21, wherein the one or more iModulons are selected from a metabolic iModulon gene cluster, a trace element iModulon gene cluster, a stress-related iModulon gene cluster or an Acetate iModulon gene cluster.

23. The method of any one of claims 20-22, further comprising preparing an isolated cell or population of cells with optimized media or substrate compatibility, comprising genetically modifying the isolated cell or population of cells to upregulate or downregulate the iModulon gene cluster identified in claim 17, optionally wherein the iModulon gene cluster is an Acetate iModulon gene cluster.

24. The method of claim 22 or 23, wherein the stress-related iModulon comprises a OxyR, RpoE, Low-osmolarity, Macrolide-efflux, PhrR, CSP, or Ectoine iModulon gene cluster.

25. The method of claim 22 or 23, wherein the trace element iModulon comprises a Zur, TPP, CueR, Fur-1, Fur-2, Enterobactin, or thiosulfate iModulon gene cluster. 150 4916-8982-5303.

126. The method of claim 22 or 23, wherein the metabolic iModulon comprises a NarP-1, TonB, FNR, CytochromeO, NarP-2, GABA, Alr, HutC, MetJ, LiuR, ArgR, Cysteine, Prophage- 1, Prophage-2, β-ox, Flagellar-1, Flagellar-2, Chaperone, Biofilm, Ribosomal proteins, TfoX, Qst, UC-1, HapR, Cellobiose, TRAP, Formate, TreR, Arabinose, AGL, NagC, GntR, Maltose, FruR, Glycerol, TTT, Acetate, GalR, Rhamnose, Null-1, Null-2, Null-3, UC-2, UC-3, UC-4, UC- 5, UC-6, UC-7, UC-8, or UC-9 iModulon gene cluster.

27. The method of any one of claims 20-26, wherein the media formulation comprises acetate, ethanol, cellobiose, galactose, glycogen, or GlcN.

28. The method of any one of claims 20-27, wherein the isolated cell or population of cells is a prokaryotic cell or a eukaryotic cell.

29. The method of claim 28, wherein the prokaryotic cell is a bacterial cell, optionally selected from an E. coli, P. Putida, Corynebacterium glutamicum, or Vibrio natriegens.

30. An iModulon gene cluster comprising SEQ ID Nos: 1-6 and optionally, SEQ ID Nos: 7- 8, or an equivalent of each thereof.

31. An isolated host cell, comprising the iModulon gene cluster of claim 30.

32. The isolated host cell of claim 31, wherein the isolated cell is a prokaryotic cell or a eukaryotic cell.

33. The isolated host cell of claim 32, wherein the prokaryotic cell is a bacterial cell, optionally selected from an E. coli, P. putida, Corynebacterium glutamicum, Vibrio natriegens and genetically minimized species such as JCVI_Syn3A. 151 4916-8982-5303.

134. The isolated host cell of claim 33, wherein the bacterial cell is an E. Coli cell and the gene cluster comprises the nucleic acid sequences of SEQ ID Nos: 1-6 and optionally, SEQ ID Nos: 7-8.

35. A substantially homogenous population of cells of any one of claims 31-34.

36. A composition comprising the isolated host cell of any one of claims 31-34, and a carrier, and optionally a preservative or cryoprotective agent.

37. The composition of claim 30 or the isolated host cell of any one of claims 31-34, wherein the equivalent thereof comprises a nucleic acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID Nos: 1-6 and optionally, SEQ ID Nos: 7-8.

38. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster of the composition of claim 30 into a host cell lacking the biochemical pathway.

39. A method for preparing an isolated cell or population of cells with an exogenous biochemical pathway, comprising introducing a gene cluster comprising the polynucleotides of SEQ ID Nos: 77-82 and optionally, SEQ ID Nos: 83-84 into the host cell, wherein the host cell gains the acetate metabolic pathway of Pseudomonas putida.

40. The method of any one of claims 38-39, further comprising optimizing the biochemical pathway through adaptive laboratory evolution.

41. The method of any one of claims 38-40, further comprising culturing the isolated cell or population of cells and isolating the product of the biochemical pathway from the cell culture. 152 4916-8982-5303.1

Citation Information

Patent Citations

  • Cell culture medium

    US20220154137A1

  • Genomic and proteomic approaches for the development of cell culture medium

    WO2004101808A2

  • Genetic algorithm and imodulon based optimization of media formulation for quality, titer, strain, and process improvement biologics

    WO2024030344A1