Performance prediction of fundamental transcriptional programs

WO2024163909A3PCT designated stage expired Publication Date: 2025-06-12GEORGIA TECH RES CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/014269
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-02
Filing Date
2024-02-02
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The complexity of synthetic genetic circuits in biological systems makes predictive design challenging due to non-modular biological circuit components and limitations in DNA-level composability, requiring time-consuming trial and error, and existing metrics like RPU for promoter strength are inaccurate due to local genetic context influences.

Method used

Development of a quantitatively predictive method for designing compressed genetic circuits using cellobiose-responsive anti-repressors and the relative expression unit (REU) metric to accurately quantify expression levels, enabling orthogonal 3-input transcriptional programs and reducing design complexity.

Benefits of technology

The method allows for accurate prediction of transcriptional program performances, reducing design complexity and improving the accuracy of genetic circuit designs, enabling efficient control of metabolic pathways and operons, with an average error of less than 1.4-fold in fundamental circuit predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024014269_12062025_PF_FP_ABST
    Figure US2024014269_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides constructs and cell compositions for reprogramming cells. The present disclosure also provides methods using constructs and cell compositions to modify and / or monitor cells. The present disclosure provides transcriptional programming, and modeling thereof, to three inputs that can provide provide up to 256 logical operations, e.g., to form a Turing complete and scalable decision-making platform technology for biocomputing and biological intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

PERFORMANCE PREDICTION OF FUNDAMENTALTRANSCRIPTIONAL PROGRAMSStatement of Government Support

[0001] This invention was made with government support under Grant Nos. GCR 1934836; MCB 2123855; MCB 1921061; and CBET 1804639 awarded by the National Science Foundation. The government has certain rights in the invention.Related Application

[0002] The international PCT application claims priority to, and the benefit of, U.S. Provisional Patent Application No. 63 / 482,992, filed February 2, 2023, which is incorporated by reference herein in its entirety.Field

[0003] The present disclosure relates to nucleic acid constructs and uses thereof.Background

[0004] Biological computation, at its core, is the ability to engineer and develop systems capable of converting information (inputs) into a programmable gene expression (output(s)). Gene regulation in biological systems can be viewed as a molecular computer. That is, gene expression can be modeled as on-off states of Boolean (digital) logic, which can integrate multiple digital inputs into a desired output. Currently, living cells can be programmed with genetic parts such as promoters, transcription factors, and metabolic genes to encode logical operations that integrate environmental and cellular signals. Synthetic genetic logic gates have been engineered, including those capable of accomplishing Boolean functions (e.g., AND, OR, and NOT functions), which have been employed for pharmaceutical and biotechnological applications. Moreover, combinations of such gates can be used to construct biological analogs of more advanced electronic circuits including switches, logic, memory, pulse generators, and oscillators. Although logic in synthetic gene networks can be accomplished either at the transcriptional or translational levels, the former is more commonly employed in the development of synthetic gene networks via the use of transcription factors (TFs) to activate or repress genes of interest. Broadly speaking, TFs are DNA-binding proteins capable of blocking (or recruiting) RNA polymerase activity at the site of genetic promoters, and these functions can be combined in modular ways to engineer synthetic gene networks. For the most part, early bacterial gene circuits were based on a coreset of repressors, namely, TetR, LacI, and bacteriophage λ cl, which have been extensively studied.

[0005] As circuit complexity increases in synthetic biology in developing new biological parts, devices, and systems or to redesign existing systems, there is a benefit to modeling synthetic biology.Summary

[0006] The present disclosure provides constructs and cell compositions for reprogramming cells. The present disclosure also provides methods using constructs and cell compositions to modify and / or monitor cells.

[0007] The present disclosure expands transcriptional programming from logical operations having two inputs that can provide up to 16 logical operations to three inputs that can provide provide up to 256 logical operations, e.g., to form a Turing complete and scalable decision- making platform technology for biocomputing and biological intelligence. The integrated biological circuit can increase biological computing capacity while minimizing metabolic burden.

[0008] Additionally, exemplary systems and methods are disclosed employing modeling of transcriptional programming that leverages systems of engineered transcription factors to impart decision-making (e.g., Boolean logic) in chassis cells to predict the performance of transcriptional programs. As the number of components used to construct decision-making systems in transcriptional programming is rapidly increasing, an exhaustive experimental evaluation of iterations of biological circuits can be impractical. The exemplary systems and methods can be employed as a predictive tool to guide and accelerate the design of transcriptional programs.

[0009] In some embodiments, the exemplary system employs the development and experimental characterization of a large collection of network capable single-INPUT logical operations - i.e., engineered BUFFER (repressor) and engineered NOT (antirepressor) logical operations. Using the single-INPUT data and developed metrology, the system can model and predict the performances of all fundamental two-INPUT compressed logical operations (i.e., compressed AND gates, and compressed NOR gates). In addition, the system can model and predict the performance of compressed mixed phenotype logical operations (A NIMPLY B gates, and complementary B NIMPLY A gates). A study was conducted and the results demonstrated that single-INPUT data is sufficient to accurately predict both the qualitative and quantitative performance of a complex circuit. The exemplary system can be employed for the predictive design of transcriptional programs of greater complexity.

[0010] In one aspect, a construct is disclosed comprising a plurality of nucleic acid sequences encoding a first group of one or more regulatory core domains; a second group of one or more regulatory core domains; one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose- responsive anti-repressor; and one or more DNA operator elements, wherein the one or more DNA operator elements are each specifically recognized by one of the DNA binding domains.

[0011] In some embodiments, the cellobiose-responsive anti-repressor comprises cellobiose- responsive anti-repressor EA1, EA2, EA3, or a variant thereof.

[0012] In some embodiments, the first group of the one or more regulatory core domains, the second group of the one or more regulatory core domains, and a third group of one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose-responsive anti-repressor, to form a three-input transcription program.

[0013] In some embodiments, the first group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0014] In some embodiments, the second group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0015] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises a second repressor.

[0016] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises an anti-repressor.

[0017] In some embodiments, the first group of one or more regulatory core domains comprises at first anti-repressor, and the second group of one or more regulatory core domains comprises a second anti-repressor.

[0018] In one aspect, disclosed herein is a method of modifying a gastrointestinal tract microbiome in a subject, comprising administering to the subject an effective amount of a cell comprising a construct comprising a plurality of nucleic acid sequences encoding a first group of one or more regulatory core domains, a second group of one or more regulatory core domains formed of a cellobiose-responsive anti-repressor, one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA bindingdomains, and one or more DNA operator elements, wherein the one or more DNA operator elements are each specifically recognized by one of the DNA binding domains.

[0019] In some embodiments, the cellobiose-responsive anti-repressor comprises cellobiose- responsive anti-repressor EA1, EA2, EA3, or a variant thereof.

[0020] In some embodiments, the first group of the one or more regulatory core domains, the second group of the one or more regulatory core domains, and a third group of one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose-responsive anti-repressor, to form a three-input transcription program.

[0021] In some embodiments, the first group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0022] In some embodiments, the second group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0023] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises a second repressor.

[0024] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises an anti-repressor.

[0025] In some embodiments, the first group of one or more regulatory core domains comprises a first anti-repressor, and the second group of one or more regulatory core domains comprises a second anti-repressor.

[0026] In some embodiments, the method of modifying a gastrointestinal tract microbiome in a subject, comprising administering to the subject an effective amount of the cell of any preceding aspect, wherein the cell comprises the construct of any preceding aspect.

[0027] In one aspect, disclosed herein is a method of modifying a gastrointestinal tract microbiome in a subject, comprising administering to the subject an effective amount of a cell comprising a construct comprising a plurality of nucleic acid sequences encoding a first group of one or more regulatory core domains, a second group of one or more regulatory core domains formed of a cellobiose-responsive anti-repressor, one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA binding domains, and one or more DNA operator elements, wherein the one or more DNA operator elements are each specifically recognized by one of the DNA binding domains.

[0028] In some embodiments, the cellobiose-responsive anti-repressor comprises cellobiose- responsive anti-repressor EA1, EA2, EA3, or a variant thereof.

[0029] In some embodiments, the first group of the one or more regulatory core domains, the second group of the one or more regulatory core domains, and a third group of one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose-responsive anti-repressor, to form a three-input transcription program.

[0030] In some embodiments, the first group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0031] In some embodiments, the second group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0032] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises a second repressor.

[0033] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises an anti-repressor.

[0034] In some embodiments, the first group of one or more regulatory core domains comprises at first anti-repressor, and the second group of one or more regulatory core domains comprises a second anti-repressor.

[0035] In some embodiments, the method of modifying a gastrointestinal tract microbiome in a subject, comprising administering to the subject an effective amount of the cell of any preceding aspect, wherein the cell comprises the construct of any preceding aspect.

[0036] In one aspect, a method is disclosed to predict transcriptional programming of gene expression, the method comprising: providing a gene model having network repressors and network anti-repressors, wherein the network repressors and network anti-repressors are designated for binding to a plurality of candidate DNA binding function (e.g., YQR | O1, HQN | Ottg, or GKR | Ogac); applying one or more of the plurality of candidate DNA binding function as a first 2-node single-in single-out (SISO) networks of transcription factors binned via the alternate DNA binding function; and determining and outputting a predicted value for at least one of: relative expression units (REU), fold induction, fold anti-induction, fluorescence in an absence of inducer (e.g., a), maximum fluorescence relative to a basal expression of an expression state (e.g., σ), a traceability score, or a combination thereof.

[0037] In one aspect, a method is disclosed comprising: providing a gene model having network repressors and network anti-repressors, wherein the network repressors and network anti-repressors are designated for binding to a plurality of candidate DNA binding function (e.g., YQR | O1, HQN | Otg, or GKR | Ogac); applying one or more of the plurality of candidate DNA binding function as a 3 -node networks of transcription factors binned via the alternate DNA binding function; and determining logical operations for the 3 -node networks of transcription factors, wherein the logical operations are made accessible to biocomputation operation.

[0038] In some embodiments, the method further includes applying one or more of the plurality of candidate DNA binding function as a second 2-node single-in single-out (SISO) networks of transcription factors binned via the alternate DNA binding function following the application of the one or more of the plurality of candidate DNA binding function as the first 2-node SISO networks to provide a combination of the first 2-node SISO networks and the second 2-node SISO networks, wherein the combination of the first 2-node SISO networks and the second 2-node SISO networks is equivalent to a multiple input single output (MISO) network.

[0039] In some embodiments, the outputted predicted value is for the combination.

[0040] In some embodiments, the plurality of candidate DNA binding function is applied via two or more of: a 1 -INPUT BUFFER logic operation (e.g., single repressor), a 1 -INPUT NOT logic gate operation (e.g., having a single anti-repressor), a 2-INPUT AND logic gate operation (e.g., having a repressor pair), a 2-INPUT NOR logic gate operation (e.g., having an antirepressor pair), a 2-INPUT A NIMPLY B logic gate operation (e.g., having a repressor / anti-repressor pair), a 2-INPUT B NIMPLY A logic gate operation (e.g., having an anti-repressor / repressor pair), a 2-INPUT XNOR logic gate operation, and a 2-INPUT NAND logic gate operation.

[0041] In some embodiments, the 2-INPUT AND logic gate operation comprises:wherein ε is a value for fluorescence in the absence of inducer, σ is a value constant for a maximum fluorescence relative to basal expression of an OFF-state, is a coarse-grained Hill function (e.g.,having a value of 0 or 1), and AA(I) is a coarse-grained antithetical Hill-function for anti- repression (e.g., where 0 INPUT corresponds to the ON-state).

[0042] In some embodiments, the 2-INPUT NOR logic gate operation comprises:

[0043] In some embodiments, the 2-INPUT A NIMPLY B logic gate operation comprises:

[0044] In some embodiments, the 2-INPUT B NIMPLY A logic gate operation comprises:

[0045] In some embodiments, the predicted value for fold induction is determined by: Ω+(1) = σɅ+(l) + ε.

[0046] In some embodiments, the gene model comprises: a lactose repressor (LacI) topology having (i) a regulatory core domain (RCD) and (ii) a DNA binding domain (DBD).

[0047] In some embodiments, each 2-node SISO network comprises (i) a single transcription factor expressed on the pLacI plasmid (Novagen) (e.g., having pl 5a origin (copy number 20-30 / cell)) and (ii) a super folder fluorescent protein (GFP) reporter (e.g., expressed on the pZS*22-sfGFP plasmid containing a pSClOl origin).

[0048] In some embodiments, the REU is a measure of fluorescence of a reporter fusion to a nucleic acid sequence (e.g., first 25 AA) of a gene of interest.Brief Description of the Drawings

[0049] The components in the drawings are not necessarily to scale relative to each other. Like reference, numerals designate corresponding parts throughout the several views.

[0050] Figs, la-lf shows an example of circuit compression is shown. Fig. 1 shows a NOR gate that can constructed using inversion (left) or by using anti-repressors (middle). The logical truth table is shown to the right. Fig. lb shows a workflow for engineering a cellobiose anti-repressor. Fig. 1c shows a transcriptional regulation of GFP expression by the TFs in Fig. lb. Fig. Id shows the putative transcriptional programming design space for repressor / anti-repressor pairs is shown. IPTG is known to inhibit fucose-responsive TFs. Fig. le shows reduced design space for transcriptional programming based on orthogonality.

[0051] Fig. 2 shows the validation of anti-CelRs equipped with alternate DBDs.

[0052] Fig. 3a shows an expression of GFP versus LacLGFP fusion reporters from five synthetic promoters is shown. Fig. 3b shows an expression of LacLGFP versus LacI-mKate from five synthetic promoters is shown. Fig. 3c shows an RBS library applied to three promoters does not show modularity. Fig. 3d shows the genetic schematic of an REU standard curve for Nanoluc is shown. Fig. 3e shows the measured REU as a function of AHL concentration is shown for Fig. 3d. Fig. 3f shows activity of full-length Nanoluc as a function of REU is shown to correlate linearly. Fig. 3g shows the genetic schematic of an REU standard curve for LacI is shown. Fig. 3h shows the measured REU as a function ofAHL concentration is shown for Fig. 3g. Fig. 2i shows REU(LacI) correlates linearly with REU(Nanoluc) over varying transcription levels.

[0053] Figs. 4a - 41 shows a genetic schematic for the transcription factor titration circuit and associated operation. Fig. 4a shows the genetic schematic for the transcription factor titration circuit. Fig. 4b shows an example dataset generated from the measurement of RbsR(YQR) in the circuit from Fig. 4a. REUin corresponds to the REU value at which the TF is expressed, and REUout corresponds to the measured fluorescence of the LacLGFP reporter. Fig. 4c shows the performance metric for Fig. 4b. Fig. 4d shows the genetic context of an RBS library designed for modulating the constitutive expression of TFs. Fig. 4e shows data corresponding to select RBSs from Fig. 3d. Fig. 34 shows the genetic schematic for BUFFER / NOT gates. Figs. 4g - 4i shows BUFFER gate performances are shown compared to the predicted expression levels based on REU modeling. Figs. 4j - 41 shows NOT gate performances compared to predicted expression levels based on REU modeling.

[0054] Fig. 5 shows an investigation of the impact of GOI length on LacLGFP fusion reporters.

[0055] Fig. 6 shows the non-modularity of RBS libraries applied to different promoters.

[0056] Fig. 7 shows non-modularity of ribozymes.

[0057] Fig. 8a shows the workflow for predictive transcriptional program design is shown for an A IMPLY B gate. Fig. 8b shows the wiring diagram for a 3-input consensus program is shown (left) along with the measured performance (right).

[0058] Figa. 9a - 9c show titration curves for all LacI and 1(A) (Fig. 9a), all RbsR and R(A) (Fig. 9b), and all CelR and E(A) (Fig. 9c).

[0059] Fig. 10 shows transcription factor orthogonality.

[0060] Fig. 11 shows additional examples of 2-input circuits.

[0061] Fig. 12a shows the biosynthetic pathway for lycopene production. Fig. 12b shows the genetic schematic for measuring REU of each Crt gene. Fig. 12c shows the dose-response functions of the three circuits from Fig. 12b. The vertical dashed lines indicate the induced concentration that yields an REU of 100. Fig. 12d shows the genetic schematic for the fully refactored synthase pathway. Fig. 12e shows the measured lycopene titer from the circuit in Fig. 12d. Fig. 12f The genetic schematic for measuring REU of the Crt genes in an operon. Fig. 512 shows the measured REU values for the three constructs from Fig. 12f. Fig. 512 shows the genetic schematic for the functional operon.

[0062] Fig. 13 shows the measured lycopene titer from the circuit along with the titer achieved from expressing the circuit in Fig. 12d at the same levels.

[0063] Fig. 14 shows a comparison of DNA sequence-based models of expression.

[0064] Fig. 15a- 15g shows modular components used in a design space. Fig. 15a shows the performance card of a repressor (X+) and the abstraction of metrics to a logical BUFFER operation. Fig. 15b shows the performance card of an anti-repressor (XA) and the abstraction of metrics to a logical NOT operation. Fig. 15c shows the design space overview in which each of the 5 X+or XARCDs can be paired with 1 of 8 ADRs and directed to 1 of 2 operator positions (OPs), resulting in a putative design space of 80 BUFFER and 80 NOT operations. Figs. 15d - 15g shows example genetic architectures for a PROXIMAL architecture with an operator position downstream of the promoter in which the transcription factor interferes with RNA polymerase’s ability to transcribe DNA (Fig. 15d), a CORE architecture featuring an operator intercalated between the -35 and -10 hexamers of the synthetic trc promoter in E. coli in which the transcription factor competes with RNA polymerase binding to DNA (Fig. 15e). Figs. 15f — 15g shows a two-input architecture. Fig. 15f shows PROXIMAL SE-PA architecture, as shown in Fig. 15d, with two transcription factors directed to the operator. Fig. 15g shows CORE SE-PA architecture as shown in Fig. 15e, with two transcription factors directed to the DNA operator.

[0065] Figs. 16a - 16b show example combinatorial sets of SE-PA AND gates. Fig. 16a shows an illustration of non-synonymous repressor pairs combined with 8 ADRs yielding 80 putative PROXIMAL SE-PA AND gates. Repressors classified as non-operational (Fig. 26) are shown faded, and incompatible repressor pairs (Fig. 28) are highlighted in red. Consideration of nonoperational pairs results in a reduced space of 61 PROXIMAL SE-PA AND gates. Fig. 16b shows CORE SE-PA architecture AND gates. Elimination of non- operational and incompatible repressors results in 72 CORE SE-PA AND gates.

[0066] Figs. 17a - 17b show SE-PA AND operation and NOR operation predictive models using BUFFER SISO and NOT SISO parameters. Fig. 17a shows an AND gate logic modeled using a quadratic function of IX and IY, which control the repressor state functions A J and Ay. Each term has a coefficient α0, α1, α2, or α3, which are estimated as functions of BUFFER gate parameters εx, εY, σX, and σY (also see Fig. 15a). Functions for parameters a0, al, a2, and a3 are derived using four assumptions corresponding to each INPUT condition. (B) NOR gate logic is modeled analogous to AND logic, however, with a pair of NOT gates parameterized with anti-repressor state functions (also see Fig. IB). Given thatAy and Ay functions capture the ON-OFF state inversion from the repressor to anti-repressorphenotype, α0, α1, α2 , and α3 parameters are estimated with the same functions for both AND and NOR models.

[0067] Figs. 18a - 18b show results showing the correlation between predicted and measured OUTPUT of 133 SE-PA AND gates. Under-predictions and over-predictions fall above and below the theoretical value of 1 (red line), respectively. Fig. 18a shows correlation results for 61 PROXIMAL SE-PA AND gates across the 4 INPUT conditions. INPUTS A and B correspond to repressors X+and Y+’ respectively) and can be inferred from each BUFFER pair depicted in Fig. 16b. Fig. 18b shows the correlation between predicted and measured OUTPUT of 72 CORE SE-PA AND gates. INPUTS A and B correspond to repressors X+and Y+and can be inferred from each BUFFER pair depicted in Fig. 16b.

[0068] Figs. 19a - 19b shows a combinatorial set of 131 SE-PA NOR gates. Fig. 19b shows an illustration of non-synonymous anti-repressor pairs combined with 8 ADRs yielding 80 putative PROXIMAL SE-PA NOR gates. Anti-repressors classified as nonoperational (see Fig. 26) are shown faded, and incompatible anti-repressor pairs (see Fig. 21b) are highlighted in red. These non-operational pairs result in a reduced space of 60 proximal SE-PA NOR gates. Fig. 19b shows CORE SE-PA architecture NOR gates. Elimination of non-operational and incompatible anti-repressors results in 71 CORE SE-PA.

[0069] Figs. 20a - 20b show results showing the correlation between predicted and measured OUTPUT of 131 SE-PA NOR gates. Under-predictions and over-predictions fall above and below the theoretical value of 1 (red line), respectively. Fig. 20a shows correlation results for 60 PROXIMAL SE-PA NOR gates across the 4 INPUT conditions. INPUTS A and B correspond to anti-repressors XAand YA, respectively, and can be inferred from each NOT pair depicted in Fig. 5A. Fig. 20b shows the correlation between predicted and measured OUTPUT of 71 CORE SE-PA NOR gates. INPUTS A and B correspond to repressors X+and Y+and can be inferred from each NOT pair depicted in Fig. 19b.

[0070] Figs. 21a - 211 show results for 12 SE-PA NIMPLY logic gates at the CORE operator position. Signal INPUTs (IPTG, Ribose, Fucose, and Fructose) were selected based on the ability to perform both BUFFER and NOT logic (i.e., induce repressors and anti-repressors). This corresponds to 6 A and B INPUT pairs, which cover the full combinatorial space for NIMPLY logic. Figs. 21a - 21f show an A NIMPLY B logic employing a repressorwhich responds to INPUT A and antirepressor which responds to INPUT B. Figs. 21g- 211 show complimentary A NIMPLY B logic utilizing an anti-repressor and repressor

[0071] Figs. 22a - 22b show NIMPLY predictive models using BUFFER and NOT gate parameters. Fig. 22a shows an A NIMPLY B gate logic is modeled using a quadratic function of lx and IY, which controls the repressor state function and anti-repressor statefunction . Each term has a coefficient α0, α1, α2, or α3, which are estimated as functions of BUFFER and NOT gate parameters εx, εY, σX, and σY. Functions for parameters ao, ai, a2, and as are derived using a set of four assumptions corresponding to each INPUT condition. Fig. 22b shows a B NIMPLY A gate logic modeled analogous to the A NIMPLY B logic but with an anti-repressor state function and repressor state function

[0072] Figs. 23a - 231 shows results for 6 SERI AND operations and 6 SERI NOR operations. Figs. 23a - 23f shows AND logic gates employing a repressordirected to a cognate PROXIMAL operator (top input), and second repressor directed to a cognateCORE operator (bottom input). Results for OUTPUT prediction using SE-PA SISO parameters, prediction using SERI SISO parameters, and measured OUTPUT are shown on the right. Figs. 9g - 91 show NOR logic gates employing antirepressors

[0073] Figs. 24 (part 1) - 24 (part 5) show PROXIMAL BUFFER and NOT Gate Performance Cards. Each card displays experimental ON and OFF state OUTPUT values, INPUT signal type, DNA operator (ADR) type, and system performance metrics. Card outline color depicts the phenotype of each operation, consistent with Fig. 26. Fig. 24 (Part 1) shows LacIperformance cards, Fig. 24 (Part 2) shows RbsR ( performancecards, Fig. 24 (Part 3) shows CelRperformance cards, Fig. 24 (Part 4) shows GalR performance cards, and Fig. 24 (Part 5) shows FruR performance cards.

[0074] Figs. 24 (part 6) - 24 (part 10) show PROXIMAL NOT gate performance cards. Each card is analogous to those in Parts 1-5 but includes respective metrics for NOT gates. Fig. 24 (Part 6) shows Anti-Laci performance cards, Fig. 24 (Part 7) shows Anti-RbsRperformance cards, Fig. 24 (Part 8) shows PurRperformance cards, Fig. 24 (Part 9) shows Anti-GalS performance cards, and Fig. 24 (Part 10) shows Anti-FruRperformance cards.

[0075] Fig. 25 shows CORE BUFFER and NOT Gate Performance Cards. Each card displays experimental ON and OFF state OUTPUT values, INPUT signal type, DNA operator (ADR) type, and system performance metrics. Card outline color depicts the phenotype of each operation, consistent with Figure S3. (S2 - Part 1) LacI (I+ADR) performance cards, (S2 - Part 2) RbsR (R+ADR) performance cards, (S2 - Part 3) CelR (E+ADR) performance cards, (S2 - Part 4) GalR (G+ADR) performance cards, and (S2 -Part 5) FruR (F+ADR)performance cards. CORE NOT gate performance cards. Each card is analogous to those in Parts 1-5 but includes respective metrics for NOT gates. (S2 - Part 6) Anti -LacI (IA ADR) performance cards, (S2 - Part 7) Anti-RbsR (RAADR) performance cards, (S2 - Part 8) PurR (PAADR) performance cards, (S2 - Part 9) Anti-Gal S (SAADR) performance cards, and (S2 - Part 10) Anti-FruR (FA ADR) performance cards.

[0076] Fig. 26 shows operational and non-operational SISO logic gates. Operational gates consist of either (A) repressor - i.e., BUFFER logic - or (B) anti-repressori.e., NOT logic - phenotypes. Classification of non-operational gates as either (C) super- repressor or (D) nonfunctional phenotypes.

[0077] Figs. 27a - 27e show example genetic architectures. Fig. 27a shows a PROXIMAL architecture with an operator position downstream of the promoter. Transcription factor blocks RNA polymerase from transcribing DNA to regulate expression. Fig. 27b shows a CORE architecture featuring an operator intercalated between the -35 and -10 hexamers of the synthetic trc promoter in E. coli. Transcription factor competes with RNA polymerase to bind DNA to regulate output expression. Figs. 27c - 27e show two-input architectures. Fig. 27c shows a PROXIMAL SE-PA architecture as shown in Fig. 27a, with two transcription factors directed to the operator. Fig. 27d shows a CORE SE-PA architecture, as shown in Fig. 27b, with two transcription factors directed to the operator. Fig. 27e shows SERI architecture featuring a CORE operator and a second (non-synonymous) PROXIMAL operator.

[0078] Figs. 28a - 28b show compatible and incompatible AND gate components. Fig. 28b shows two compatible BUFFER operations constitute an AND gate when the OFF-state OUTPUT of either repressor is lower than the ON-state OUTPUT of the other. Fig. 28b shows a pair of BUFFER operations are incompatible when the inequalities shown in Fig. 28a are not met. Incompatible pairs are unlikely to produce a functional AND gate in that relative ON-state OUTPUT cannot be achieved across four input conditions.

[0079] Figs. 29a - 29b show histograms of prediction error. Fig. 29a shows error, defined as the ratio of measured to predicted OUTPUT, for all 133 SE-PA AND gates across all four INPUT conditions (the error is equivalent to values given in plots illustrated in Figs. 18a - 18b). Values below 1 are overpredictions, and values above 1 are underpredictions. Blue bars indicate < 2-fold error, and red bars indicate > 2-fold error in either direction. Fig. 29b shows histograms of prediction error for 131 NOR gates given in Figs. 20a - 20b.

[0080] Figs. 30a - 30b shows compatible and incompatible NOR gate components. Fig. 30a shows two compatible NOT operations that constitute a NOR gate when the OFF-state OUTPUT of either anti-repressor is lower than the ON-state OUTPUT of the other anti- repressor. Fig. 30b shows a pair of NOT operations is incompatible when the inequalities shown in Fig. 30a are not met. Incompatible pairs are unlikely to produce a functional NOR gate because relative ON-state OUTPUT cannot be achieved across four input conditions.

[0081] Figs. 3 la - 3 Id show an PROXIMAL and CORE SE-PA NIMPLY Logic. Fig. 31a shows an example A NIMPLY B logic gate comprising an BUFFER and NOToperation both directed to the O1PROXIMAL SE-PA genetic architecture. Fig. 31b shows a complementary B NIMPLY A operation utilizing anBUFFER operation. Figs. 31c - 3 Id show CORE SE-PA NIMPLY logic. Fig. 31c shows an example A NIMPLY B logic gate shown in Fig. 3 la at the CORE operator position. Fig. 3 Id shows a complimentary B NIMPLY A operation shown in Fig. 3 lb also directed to the CORE position. All NIMPLY gates respond to the same two INPUTs; however, variations of transcription factor phenotypes and DNA operator position yield differences in performance.

[0082] Figs. 32a - 321 show results for PROXIMAL SE-PA NIMPLY Logic, which is analogous to Figs. 21 - 211 but at the proximal operator position. Figs. 32a - 32f show an X NIMPLY Y logic employing a repressorwhich responds to INPUT A and anti- repressor which responds to INPUT B. Figs. 32g - 321 show complimentary a B NIMPLY A logic utilizing an anti-repressorand repressor

[0083] Figs. 33a - 331 show results for Insulated SERI AND Gates and NOR Gates. Specifically, results for 6 insulated SERI AND operations and 6 insulated SERI NOR operations (analogous gates to those in Figs. 23a - 231, with the addition of the genetic insulator RiboJIO). Figs. 33a - 33f show AND logic gates employing a repressordirected to a cognate PROXIMAL operator (top input), and second repressor directedto a cognate CORE operator (bottom input). Results for OUTPUT prediction using SERI SISO parameters and measured OUTPUT are shown on the right. Figs. 33g - 331 show insulated NOR logic gates employing anti-repressors via the SERI geneticarchitecture.Detailed Specification

[0084] The following description of the disclosure is provided as an enabling teaching of the disclosure in its best, currently known embodiment s). To this end, those skilled in the relevant art will recognize and appreciate that many changes can be made to the variousembodiments of the invention described herein, while still obtaining the beneficial results of the present disclosure. It will also be apparent that some of the desired benefits of the present disclosure can be obtained by selecting some of the features of the present disclosure without utilizing other features. Accordingly, those who work in the art will recognize that many modifications and adaptations to the present disclosure are possible and can even be desirable in certain circumstances and are a part of the present disclosure. Thus, the following description is provided as illustrative of the principles of the present disclosure and not in limitation thereof.

[0085] Reference will now be made in detail to the embodiments of the invention, examples of which are illustrated in the drawings and the examples. This invention may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.

[0086] Terminology

[0087] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs. The term “comprising” and variations thereof as used herein is used synonymously with the term “including” and variations thereof and are open, non-limiting terms. Although the terms “comprising” and “including” have been used herein to describe various embodiments, the terms “consisting essentially of’ and “consisting of’ can be used in place of “comprising” and “including” to provide for more specific embodiments and are also disclosed. As used in this disclosure and in the appended claims, the singular forms “a”, “an”, “the”, include plural referents unless the context clearly dictates otherwise.

[0088] The following definitions are provided for the full understanding of terms used in this specification.

[0089] The terms "about" and "approximately" are defined as being “close to” as understood by one of ordinary skill in the art. In one non-limiting embodiment the terms are defined to be within 10%. In another non-limiting embodiment, the terms are defined to be within 5%. In still another non-limiting embodiment, the terms are defined to be within 1%.

[0090] As used herein, the terms "may," "optionally," and "may optionally" are used interchangeably and are meant to include cases in which the condition occurs as well as cases in which the condition does not occur. Thus, for example, the statement that a formulation "may include an excipient" is meant to include cases in which the formulation includes an excipient as well as cases in which the formulation does not include an excipient.

[0091] “Composition” refers to any agent that has a beneficial biological effect. Beneficial biological effects include both therapeutic effects, e.g., treatment of a disorder or other undesirable physiological condition, and prophylactic effects, e.g., prevention of a disorder or other undesirable physiological condition. The terms also encompass pharmaceutically acceptable, pharmacologically active derivatives of beneficial agents specifically mentioned herein, including, but not limited to, a vector, polynucleotide, cells, salts, esters, amides, proagents, active metabolites, isomers, fragments, analogs, and the like. When the term “composition” is used, then, or when a particular composition is specifically identified, it is to be understood that the term includes the composition per se as well as pharmaceutically acceptable, pharmacologically active vector, polynucleotide, salts, esters, amides, proagents, conjugates, active metabolites, isomers, fragments, analogs, etc.

[0092] "Comprising" is intended to mean that the compositions, methods, etc. include the recited elements, but do not exclude others. "Consisting essentially of' when used to define compositions and methods, shall mean including the recited elements, but excluding other elements of any essential significance to the combination. Thus, a composition consisting essentially of the elements as defined herein would not exclude trace contaminants from the isolation and purification method and pharmaceutically acceptable carriers, such as phosphate buffered saline, preservatives, and the like. "Consisting of' shall mean excluding more than trace elements of other ingredients and substantial method steps for administering the compositions provided and / or claimed in this disclosure. Embodiments defined by each of these transition terms are within the scope of this disclosure.

[0093] An "increase" can refer to any change that results in a greater amount of a symptom, disease, composition, condition, or activity. An increase can be any individual, median, or average increase in a condition, symptom, activity, composition in a statistically significant amount. Thus, the increase can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100%, or more, increase so long as the increase is statistically significant.

[0094] A "decrease" can refer to any change that results in a smaller amount of a symptom, disease, composition, condition, or activity. A substance is also understood to decrease the genetic output of a gene when the genetic output of the gene product with the substance is less relative to the output of the gene product without the substance. Also, for example, a decrease can be a change in the symptoms of a disorder such that the symptoms are less than previously observed. A decrease can be any individual, median, or average decrease in a condition, symptom, activity, composition in a statistically significant amount. Thus, thedecrease can be a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% decrease so long as the decrease is statistically significant.

[0095] " Inhibit," "inhibiting," and "inhibition" mean to decrease an activity, response, condition, disease, or other biological parameter. This can include but is not limited to the complete ablation of the activity, response, condition, or disease. This may also include, for example, a 10% reduction in the activity, response, condition, or disease as compared to the native or control level. Thus, the reduction can be a 10, 20, 30, 40, 50, 60, 70, 80, 90, 100%, or any amount of reduction in between as compared to native or control levels.

[0096] By “reduce” or other forms of the word, such as “reducing” or “reduction,” is meant lowering of an event or characteristic. It is understood that this is typically in relation to some standard or expected value, in other words, it is relative, but that it is not always necessary for the standard or relative value to be referred to.

[0097] By “prevent” or other forms of the word, such as “preventing” or “prevention,” is meant to stop a particular event or characteristic, to stabilize or delay the development or progression of a particular event or characteristic, or to minimize the chances that a particular event or characteristic will occur. Prevent does not require comparison to a control as it is typically more absolute than, for example, reduce. As used herein, something could be reduced but not prevented, but something that is reduced could also be prevented. Likewise, something could be prevented but not reduced, but something that is prevented could also be reduced. It is understood that where reduce or prevent are used, unless specifically indicated otherwise, the use of the other word is also expressly disclosed.

[0098] The term “subject” refers to any individual who is the target of administration or treatment. The subject can be a vertebrate, for example, a mammal. In one aspect, the subject can be human, non-human primate, bovine, equine, porcine, canine, or feline. The subject can also be a guinea pig, rat, hamster, rabbit, mouse, or mole. Thus, the subject can be a human or veterinary patient. The term “patient” refers to a subject under the treatment of a clinician, e.g., a physician.

[0099] A “promoter,” as used herein, refers to a sequence in DNA that mediates the initiation of transcription by an RNApolymerase. Transcriptional promoters may comprise one ormore of a number of different sequence elements as follows: 1) sequence elements present at the site of transcription initiation; 2) sequence elements present upstream of the transcription initiation site and; 3) sequence elements down- stream of the transcription initiation site. The individual sequence elements function as sites on theDNA, where RNA polymerases and transcription factors that facilitate positioning of RNA polymerases on the DNA bind.

[0100] A “transcription factor” refers to a sequence-specific DNA-binding protein that controls the rate of transcription of genetic information from DNA to messenger RNA, by binding to a specific DNA sequence.

[0101] As used herein, a “transcription terminator” or a “terminator” refers to a segment of a nucleic acid sequence that marks the end of gene in genomic DNA during the transcription process, or gene expression. This sequence mediates or signals the end of transcription by providing signaling nucleotides in newly synthesized RNA transcripts that trigger an RNA polymerase to release the DNA and newly synthesized RNA.

[0102] The word “vector” refers to any vehicle that carries a polynucleotide into a cell for the expression of the polynucleotide in the cell. The vector may be, for example, a plasmid, a virus, a phage particle, or a nanoparticle. A “bacterial plasmid” is a small extrachromosomal DNA molecule that can be incorporated into another cell that is physically separated from the chromosomal DNA and is easily replicated. Once transformed into a suitable host, the vector may replicate and function independently of the host genome, or may, in some instances, integrate into the genome itself. In some embodiments, the vector is a DNA construct containing a DNA sequence which is operably linked to a suitable control sequence capable of effecting the expression of the DNA in a suitable host cell. Such control sequences can include a promoter to effect transcription, an optional operator sequence to control such transcription, a sequence encoding suitable mRNA ribosome binding sites, and sequences that control the termination of transcription and translation.

[0103] The term “administer,” “administering,” or derivatives thereof refer to delivering a composition, substance, inhibitor, or medication to a subject or object by one or more the following routes: oral, topical, intravenous, subcutaneous, transcutaneous, transdermal, intramuscular, intra-joint, parenteral, intra-arteriole, intradermal, intraventricular, intracranial, intraperitoneal, intralesional, intranasal, rectal, vaginal, by inhalation or via an implanted reservoir. The term “parenteral” includes subcutaneous, intravenous, intramuscular, intra- articular, intra-synovial, intrastemal, intrathecal, intrahepatic, intralesional, and intracranial injections or infusion techniques.

[0104] Generally, “host” refers to an organism or cell into which a heterologous component (polynucleotide, polypeptide, other molecule, cell) has been introduced. As used herein, a “host cell” refers to an in vivo or in vitro eukaryotic cell, prokaryotic cell (e.g., bacterial or archaeal cell), or cell from a multicellular organism (e.g., a cell line) cultured as a unicellularentity, into which a heterologous polynucleotide or polypeptide has been introduced. In some embodiments, the cell is selected from the group consisting of: an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic single-cell organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, in an invertebrate cell, a vertebrate cell, a fish cell, a frog cell, a bird cell, an insect cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell. In some cases, the cell is in vitro. In some cases, the cell is in vivo.

[0105] An "effective amount" is an amount sufficient to affect beneficial or desired results. An effective amount can be administered in one or more administrations, applications or dosages.

[0106] “Effective amount” encompasses, without limitation, an amount that can ameliorate, reverse, mitigate, prevent, or diagnose a symptom or sign of a medical condition or disorder (e.g., HIV-1 infection). Unless dictated otherwise, explicitly or by context, an “effective amount” is not limited to a minimal amount sufficient to ameliorate a condition. The severity of a disease or disorder, as well as the ability of a treatment to prevent, treat, or mitigate the disease or disorder, can be measured, without implying any limitation, by a biomarker or by a clinical parameter.

[0107] The term “microbiota” refers to the range of microorganisms that may be commensal, symbiotic, or pathogenic found in and on all multicellular organisms, including plants and animals. These include bacteria, archaea, protists, fungi, and viruses and have been found to be crucial for the immunologic, hormonal, and metabolic homeostasis of the host.

[0108] As used herein, “monitoring” refers to the actions of observing and checking the progress or quality of a treatment or procedure over a period of time. Herein, “monitoring” refers to the actions of observing and checking for changes to the GI tract microbiome following the administration of a cell comprising a construct to (re)program to transcriptional regulation of the microbiome.

[0109] A “nucleotide” is a compound consisting of a nucleoside, which consists of a nitrogenous base and a 5-carbon sugar, linked to a phosphate group forming the basic structural unit of nucleic acids, such as DNA or RNA. The four types of nucleotides are adenine (A), cytosine (C), guanine (G), and thymine (T), each of which are bound together by a phosphodiester bond to form a nucleic acid molecule.

[0110] A “nucleic acid” is a chemical compound that serves as the primary information- carrying molecules in cells and make up the cellular genetic material. Nucleic acids comprise nucleotides, which are the monomers made of a 5-carbon sugar (usually ribose ordeoxyribose), a phosphate group, and a nitrogenous base. A nucleic acid can also be a deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA).[OHl] The terms “percent identity” and “% identity,” as applied to polynucleotide sequences, refer to the percentage of residue matches between at least two polynucleotide sequences aligned using a standardized algorithm. Such an algorithm may insert, in a standardized and reproducible way, gaps in the sequences being compared in order to optimize alignment between two sequences and, therefore, achieve a more meaningful comparison of the two sequences. Percent identity for a nucleic acid sequence may be determined as understood in the art. (See, e.g., U.S. Pat. No. 7,396,664, which is incorporated herein by reference in its entirety). A suite of commonly used and freely available sequence comparison algorithms is provided by the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST) (Altschul, S. F. et al. (1990) J. Mol. Biol. 215:403 410), which is available from several sources, including the NCBI, Bethesda, Md., at its website. The BLAST software suite includes various sequence analysis programs, including “blastn,” that is used to align a known polynucleotide sequence with other polynucleotide sequences from a variety of databases. Also available is a tool called “BLAST 2 Sequences” which is used for direct pairwise comparison of two nucleotide sequences. “BLAST 2 Sequences” can be accessed and used interactively at the NCBI website. The “BLAST 2 Sequences” tool can be used for both blastn and blastp (discussed above).

[0112] Percent identity may be measured over the length of an entire defined polynucleotide sequence or may be measured over a shorter length, for example, over the length of a fragment taken from a larger, defined sequence, for instance, a fragment of at least 20, at least 30, at least 40, at least 50, at least 70, at least 100, or at least 200 contiguous nucleotides. Such lengths are exemplary only, and it is understood that any fragment length may be used to describe a length over which percentage identity may be measured.

[0113] A “full length” polynucleotide sequence is one containing at least a translation initiation codon (e.g., methionine) followed by an open reading frame and a translation termination codon. A “full length” polynucleotide sequence encodes a “full length” polypeptide sequence.

[0114] A “variant,” “mutant,” or “derivative” of a particular nucleic acid sequence may be defined as a nucleic acid sequence having at least 50% sequence identity to the particular nucleic acid sequence over a certain length of one of the nucleic acid sequences using blastn with the “BLAST 2 Sequences” tool available at the National Center for Biotechnology Information's website. (See Tatiana A. Tatusova, Thomas L. Madden (1999), “Blast 2sequences — a new tool for comparing protein and nucleotide sequences”, FEMS Microbiol Lett. 174:247-250). In some embodiments a variant polynucleotide may show, for example, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length relative to a reference polynucleotide.

[0115] As used herein, “upstream” refers to the relative position of a genetic sequence, either DNA or RNA. Upstream relates to the 5’ to 3’ direction relative to the start site of transcription, wherein upstream is usually closer to the 5’ end of a genetic sequence.

[0116] As used herein, “downstream” refers to the relative position of a genetic sequence, either DNA or RNA. Downstream relates to the 5’ to 3’ direction relative the start site of transcription, wherein downstream is usually closer to the 3’ end of a genetic sequence.

[0117] “ Gene” includes a nucleic acid fragment that expresses a functional molecule such as, but not limited to, a specific protein, including regulatory sequences preceding (5’ noncoding sequences) and following (3’ non-coding sequences) the coding sequence. “Native gene” refers to a gene as found in its natural endogenous location with its own regulatory sequences.

[0118] The terms “knock-out,” “gene knock-out,” and “genetic knock-out” are used interchangeably herein. A knock-out represents a DNA sequence of a cell that has been rendered partially or completely inoperative by targeting with a Cas protein; for example, a DNA sequence prior to knock-out could have encoded an amino acid sequence or could have had a regulatory function (e.g., promoter).

[0119] The terms “knock-in,” “gene knock-in,” “gene insertion,” and “genetic knock-in” are used interchangeably herein. A knock-in represents the replacement or insertion of a DNA sequence at a specific DNA sequence in cell by targeting with a Cas protein (for example, by homologous recombination (HR), wherein a suitable donor DNA polynucleotide is also used) examples of knock-ins are a specific insertion of a heterologous amino acid coding sequence in a coding region of a gene, or a specific insertion of a transcriptional regulatory element in a genetic locus.

[0120] By “domain” it means a contiguous stretch of nucleotides (that can be RNA, DNA, and / or RNA-DNA-combination sequence) or amino acids.

[0121] An “enhancer” is a DNA sequence that can stimulate promoter activity and may be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue-specificity of a promoter. Promoters may be derived in their entirety from a native gene or be composed of different elements derived from different promoters found in natureand / or comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different tissues or cell types, at different stages of development, or in response to different environmental conditions. It is further recognized that since, in most cases, the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of some variation may have identical promoter activity.

[0122] Nucleic acid constructs and Cell compositions

[0123] The present disclosure provides transcriptional programming using logical operations having three or more inputs that can provide up to 256 logical operations, e.g., to form a Turing complete and scalable decision-making platform technology for biocomputing and biological intelligence.

[0124] In one aspect, a construct is disclosed comprising a plurality of nucleic acid sequences encoding a first group of one or more regulatory core domains; a second group of one or more regulatory core domains; one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose- responsive anti-repressor; and one or more DNA operator elements, wherein the one or more DNA operator elements are each specifically recognized by one of the DNA binding domains.

[0125] In some embodiments, the cellobiose-responsive anti-repressor comprises cellobiose- responsive anti-repressor EA1, EA2, EA3, or a variant thereof.

[0126] In some embodiments, the first group of the one or more regulatory core domains, the second group of the one or more regulatory core domains, and a third group of one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose-responsive anti-repressor, to form a three-input transcription program.

[0127] In some embodiments, the first group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0128] In some embodiments, the second group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0129] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises a second repressor.

[0130] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises an anti-repressor.

[0131] In some embodiments, the first group of one or more regulatory core domains comprises at first anti-repressor, and the second group of one or more regulatory core domains comprises a second anti-repressor.

[0132] In some embodiments, the first group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof. In some embodiments, the first group of one or more regulatory core domains comprises one, two, three, four, five or more repressors or one, two, three, four, five or more anti-repressor, or a combination thereof. In some embodiments, the first group of one or more regulatory core domains comprises at least two repressors, at least two anti-repressors, or a combination thereof.

[0133] In some embodiments, the second group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof. In some embodiments, the second group of one or more regulatory core domains comprises one, two, three, four, five or more repressors or one, two, three, four, five or more anti-repressor, or a combination thereof. In some embodiments, the second group of one or more regulatory core domains comprises at least two repressor or at least two anti-repressors, or a combination thereof.

[0134] In some embodiments, the first group of one or more regulatory core domains is specifically recognized by a first agent. In some embodiments, the first agent is isopropyl-P- D- 1 -thiogalactopyranoside.

[0135] In some embodiments, the second group of one or more regulatory core domains is specifically recognized by a second agent. In some embodiments, the second agent is D- ribose.

[0136] In some embodiments, the first and second groups of the one or more regulatory core domains are linked to a same DNA binding domain. In some embodiments, the first and second groups of the one or more regulatory core domains are linked to different DNA binding domains.

[0137] In some embodiments, the construct comprises a plurality of nucleic acid sequences encoding a first group of two regulatory core domains, a second group of two regulatory core domains, and three DNA binding domains, wherein the first group of the regulatory core domains and the second group of the regulatory core domains are each linked to one of thethree DNA binding domains, and three DNA operator elements that are each specifically recognized by one of the three DNA binding domains.

[0138] In some embodiments, the construct further comprises a nucleic acid sequence encoding a reporter including, but not limited to green fluorescent protein (GFP), yellow fluorescent protein (YFP), blue fluorescent protein (BFP), cyane fluorescent protein (CFP), monomeric red fluorescent protein (mRFP), Discosoma striata (DsRed), mCherry, mOrange, tdTomato, mSTrawberry, mPlum, photoactivatable GFP (PA-GFP), Venus, Kaede, monomeric kusabira orange (mKO), Dronpa, enhanced CFP (ECFP), Emerald, Cyan fluorescent protein for energy transfer (CyPet), super CFP (SCFP), Cerulean, photoswitchable CFP (PS-CFP2), photoactivatable RFP1 (PA-RFP1), photoactivatable mCherry (PA-mCherry), monomeric teal fluorescent protein (mTFP1), Eos fluorescent protein (EosFP), Dendra, TagBFP, TagRFP, enhanced YFP (EYFP), luciferase, Topaz, Citrine, yellow fluorescent protein for energy transfer (YPet), super YFP (SYFP), enhanced GFP (EGFP), Superfolder GFP, T-Sapphire, Fucci, mK02, mOrange2, m Apple, Sirius, Azurite, EBFP, and / or EBFP2.

[0139] In some embodiments, the construct is coupled to a nucleic acid sequence encoding components of a CRISPR gene editing system.

[0140] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR- associated system (CRISPR / -Cas9) is a popular tool for genome editing. However, use of CRISPR-Cas9 as a programmable genome editing tool is hindered by off-target DNA cleavage (Cong et al., 2013; Doudna, 2020; Fu et al., 2013; Jinek et al., 2013), and the underlying mechanisms by which Cas9 recognizes mismatches are poorly understood (Kim et al., 2019; Liu et al., 2020; Slaymaker and Gaudelli, 2021). Although Cas9 variants with greater discrimination against mismatches have been designed (Chen et al., 2017; Kleinstiver et al., 2016; Slaymaker et al., 2016), these suffer from significantly reduced on-target DNA cleavage rates (Kim et al., 2020; Liu et al., 2020).

[0141] In some embodiments, the construct further comprises a nucleic acid sequence encoding a dead Cas9 endonuclease (dCas9) and a single guide RNA (sgRNA).

[0142] The dCas9, also known as an endonuclease deficient Cas, is a variant form of the parent Cas9, whose endonuclease activity is removed by mutating the endonuclease domains. It should be understood however that dCas9 may still possess binding activity to guide RNA and targeted DNA strands.

[0143] Disclosed herein is an isolated Cas9 variant or a fragment. By “variant” or “fragment” is meant a functional fragment or functional variant of a native Cas protein, or a protein thatshares at least 30%, between 30% and 35%, at least 35%, between 35% and 40%, at least 40%, between 40% and 45%, at least 45%, between 45% and 50%, at least 50%, 50%, between 50% and 55%, at least 55%, between 55% and 60%, at least 60%, between 60% and 65%, at least 65%, between 65% and 70%, at least 70%, between 70% and 75%, at least 75%, between 75% and 80%, at least 80%, between 80% and 85%, at least 85%, between 85% and 90%, at least 90%, between 90% and 95%, at least 95%, between 95% and 96%, at least 96%, between 96% and 97%, at least 97%, between 97% and 98%, at least 98%, between 98% and 99%, or at least 99% sequence identity to a parent Cas9 polypeptide. It is noted that “parent” and “native” are referred to alternatively herein and have the same meaning, which is the naturally occurring Cas9 on which the variant or fragment thereof is based.

[0144] The terms “single guide RNA” and “sgRNA” are used interchangeably herein and relate to a synthetic fusion of two RNA molecules, a crRNA (CRISPR RNA) comprising a variable targeting domain (linked to a tracr mate sequence that hybridizes to a tracrRNA), fused to a tracrRNA (trans-activating CRISPR RNA). The single guide RNA can comprise a crRNA or crRNA fragment and a tracrRNA or tracrRNA fragment of the CRISPR / Cas system that can form a complex with a Cas endonuclease, wherein said guide RNA / Cas endonuclease complex can direct the Cas endonuclease to a DNA target site, enabling the Cas endonuclease to recognize, optionally bind to, and optionally nick or cleave (introduce a single or double-strand break) the DNA target site.

[0145] It should be understood that the construct can be introduced and / or integrated into the cell by techniques commonly known in the art, including, but not limited to, the method of transformation. "Transformation" of a cellular organism with DNA means introducing DNA into an organism so that at least a portion of the DNA is replicable, either as an extrachromosomal element or by chromosomal integration. The term "transformed" refers to a cell in which DNA was introduced. The cell is termed "host cell," and it may be either prokaryotic or eukaryotic. Typical prokaryotic host cells include various strains of E. coli. Typical eukaryotic host cells are mammalian, such as gastrointestinal cells of human origin. The introduced DNA sequence may be from the same species as the host cell or a different species from the host cell, or it may be a hybrid DNA sequence containing some foreign and some homologous DNA.

[0146] Methods

[0147] The present disclosure also provides methods of using nucleic acid constructs and / or cell compositions to modify and / or monitor a GI microbiome.

[0148] In one aspect, disclosed herein is a method of modifying a gastrointestinal tract microbiome in a subject, comprising administering to the subject an effective amount of a cell comprising a construct comprising a plurality of nucleic acid sequences encoding a first group of one or more regulatory core domains, a second group of one or more regulatory core domains formed of a cellobiose-responsive anti-repressor, one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA binding domains, and one or more DNA operator elements, wherein the one or more DNA operator elements are each specifically recognized by one of the DNA binding domains.

[0149] In some embodiments, the cellobiose-responsive anti-repressor comprises cellobiose- responsive anti-repressor EA1, EA2, EA3, or a variant thereof.

[0150] In some embodiments, the first group of the regulatory core domains, the second group of the regulatory core domains, and a third group of one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose- responsive anti-repressor, to form a three-input transcription program.

[0151] In some embodiments, the first group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0152] In some embodiments, the second group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

[0153] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises a second repressor.

[0154] In some embodiments, the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises an anti-repressor.

[0155] In some embodiments, the first group of one or more regulatory core domains comprises at first anti-repressor, and the second group of one or more regulatory core domains comprises a second anti-repressor.

[0156] In some embodiments, the method of modifying a gastrointestinal tract microbiome in a subject comprises administering to the subject an effective amount of the cell of any preceding aspect, wherein the cell comprises the construct of any preceding aspect.

[0157] As used herein, “modifying a gastrointestinal tract microbiome” refers to transcriptionally increasing or decreasing functions, cell numbers, or combinations thereof in a host organism, such as humans, to promote or revert the host GI tract to a normalfunctioning state. The method of modifying a GI tract microbiome” also refers to transcriptionally increasing or decreasing functions, cell numbers, gene expression, or combinations thereof in a host organism to facilitate the understanding of disease pathogeneses associated with the GI tract and further understanding bacterial populations within the GI tract microbiome.

[0158] In one aspect, a method is disclosed to predict transcriptional programming of gene expression, the method comprising: providing a gene model having network repressors and network anti-repressors, wherein the network repressors and network anti-repressors are designated for binding to a plurality of candidate DNA binding function (e.g., YQR | O1, HQN | Ottg, or GKR | Ogac); applying one or more of the plurality of candidate DNA binding function as a first 2-node single-in single-out (SISO) networks of transcription factors binned via the alternate DNA binding function; and determining and outputting a predicted value for at least one of: relative expression units (REU), fold induction, fold anti-induction, fluorescence in an absence of inducer (e.g., a), maximum fluorescence relative to a basal expression of an expression state (e.g., o), a traceability score, or a combination thereof.

[0159] In one aspect, a method is disclosed comprising: providing a gene model having network repressors and network anti-repressors, wherein the network repressors and network anti-repressors are designated for binding to a plurality of candidate DNA binding functions (e.g., YQR | O1, HQN | Otg, or GKR | Ogac); applying one or more of the plurality of candidate DNA binding function as a 3 -node network of transcription factors binned via the alternate DNA binding function; and determining logical operations for the 3 -node networks of transcription factors, wherein the logical operations are made accessible to biocomputation operation.

[0160] In some embodiments, the method further includes applying one or more of the plurality of candidate DNA binding functions as a second 2-node single-in single-out (SISO) networks of transcription factors binned via the alternate DNA binding function following the application of the one or more of the plurality of candidate DNA binding function as the first 2-node SISO networks to provide a combination of the first 2-node SISO networks and the second 2-node SISO networks, wherein the combination of the first 2-node SISO networks and the second 2-node SISO networks is equivalent to a multiple input single output (MISO) network.

[0161] In some embodiments, the outputted predicted value is for the combination.

[0162] In some embodiments, the plurality of candidate DNA binding function is applied via two or more of: a 1 -INPUT BUFFER logic operation (e.g., single repressor), a 1 -INPUT NOT logic gate operation (e.g., having a single anti-repressor), a 2-INPUT AND logic gate operation (e.g., having a repressor pair), a 2-INPUT NOR logic gate operation (e.g., having an antirepressor pair), a 2-INPUT A NIMPLY B logic gate operation (e.g., having a repressor / anti-repressor pair), a 2-INPUT B NIMPLY A logic gate operation (e.g., having an anti-repressor / repressor pair), a 2-INPUT XNOR logic gate operation, and a 2-INPUT NAND logic gate operation.

[0163] In some embodiments, the 2-INPUT AND logic gate operation comprises:wherein ε is a value for fluorescence in the absence of inducer, σ is a value constant for a maximum fluorescence relative to basal expression of an OFF-state, A+(I) is a coarse-grained Hill function (e.g., having a value of 0 or 1), and AA(I) is a coarse-grained antithetical Hill-function for anti- repression (e.g., where 0 INPUT corresponds to the ON-state).

[0164] In some embodiments, the 2-INPUT NOR logic gate operation comprises:

[0165] In some embodiments, the 2-INPUT A NIMPLY B logic gate operation comprises:

[0166] In some embodiments, the 2-INPUT B NIMPLY A logic gate operation comprises:

[0167] In some embodiments, the predicted value for fold induction is determined by: Ω+(1) = oA+(l) + £.

[0168] In some embodiments, the gene model comprises a lactose repressor (LacI) topology having (i) a regulatory core domain (RCD) and (ii) a DNA binding domain (DBD).

[0169] In some embodiments, each 2-node SISO network comprises (i) a single transcription factor expressed on the pLacI plasmid (Novagen) (e.g., having pl 5a origin (copy number 20-30 / cell)) and (ii) a super folder fluorescent protein (GFP) reporter (e.g., expressed on the pZS*22-sfGFP plasmid containing a pSClOl origin).

[0170] In some embodiments, the REU is a measure of fluorescence of a reporter fusion to a nucleic acid sequence (e.g., first 25 AA) of a gene of interest.

[0171] In one aspect, disclosed herein is a method of monitoring a gastrointestinal tract microbiome in a subject, comprising administering to the subject an effective amount of a cell comprising a construct comprising a plurality of nucleic acid sequences encoding a firstgroup of one or more regulatory core domains, a second group of one or more regulatory core domains, one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA binding domains, and one or more DNA operator elements, wherein the one or more DNA operator elements are each specifically recognized by one of the DNA binding domains.

[0172] In some embodiments, the method of monitoring a gastrointestinal tract microbiome in a subject, comprising administering to the subject an effective amount of the cell of any preceding aspect, wherein the cell comprises the construct of any preceding aspect.

[0173] As used herein, “monitoring a gastrointestinal tract microbiome” refers to the processes of observing and / or routinely checking the increases or decreases in functions, cell numbers, or combinations thereof caused by transcriptionally (re)programming a host microbiome. It should be understood that the process of monitoring can be performed as often or as sparingly necessary to observe a desired effect. In some embodiments, the host can be monitored every day, every 2 days, every 3 days, every 4 days, every 5 days, every 6 days, every 7 days, or more. In some embodiments, the host can be monitored every week, every 2 weeks, every 3 weeks, every 4 weeks, or more. In some embodiments, the host can be monitored every month, every 2 months, every 3 months, every 4 months, every 5 months, every 6 months, every 7 months, every 8 months, every 9 months, every 10 months, every 11 months, every 12 months, or more. In some embodiments, the host can be monitored every year, every 2 years, every 3 years, every 4 years, every 5 years, or more.

[0174] In some embodiments, the host can be monitored 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37,38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62,63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87,88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or more times.

[0175] In one aspect, disclosed herein is a method of treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject an effective amount of a cell comprising a construct comprising a plurality of nucleic acid sequences encoding a first group of one or more regulatory core domains, a second group of one or more regulatory core domains, one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA binding domains, and one or more DNA operator elements, wherein the one or more DNA operator elements are eachspecifically recognized by one of the DNA binding domains, and wherein the construct transcriptionally (re)programs a bacterial population within the subject’s GI tract to improve the host’s health.

[0176] In one aspect, disclosed herein is a method of treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject an effective amount of the cell comprising the construct of any preceding aspect, wherein the construct transcriptionally (re)programs a bacterial population within the subject’s GI tract to improve the host’s health.

[0177] In some embodiments, the method (re)programs the bacterial population into a therapeutic bacteria. In some embodiments, the bacterial population comprises a Bacteroides species including, but not limited to B. thetaiotaomicron (Bt), B. fragilis (Bf), B. vulgatus (Bv), B. ovatus (Bo), or B. uniformis (Bu).

[0178] In some embodiments, the disease or disorder includes, but are not limited to a cancer, a gastrointestinal disease, a congenital disease or disorder, an infectious disease, or combinations thereof.

[0179] In some embodiments, the cancer includes, but is not limited to acoustic neuroma, adenocarcinoma, adrenal gland cancer, anal cancer, angiosarcoma (e.g., lymphangiosarcoma, lymphangioendotheliosarcoma, hemangiosarcoma), appendix cancer, benign monoclonal gammopathy, biliary cancer (e.g., cholangiocarcinoma), bladder cancer, breast cancer (e.g., adenocarcinoma of the breast, papillary carcinoma of the breast, mammary cancer, medullary carcinoma of the breast), bronchus cancer, carcinoid tumor, cervical cancer (e.g., cervical adenocarcinoma), choriocarcinoma, chordoma, craniopharyngioma, colorectal cancer (e.g., colon cancer, rectal cancer, colorectal adenocarcinoma), epithelial carcinoma, ependymoma, endotheliosarcoma (e.g., Kaposi's sarcoma, multiple idiopathic hemorrhagic sarcoma), endometrial cancer (e.g., uterine cancer, uterine sarcoma), esophageal cancer (e.g., adenocarcinoma of the esophagus, Barrett's adenocarinoma), Ewing's sarcoma, familiar hypereosinophilia, gall bladder cancer, gastric cancer (e.g., stomach adenocarcinoma), gastrointestinal stromal tumor (GIST), oral cancer (e.g., oral squamous cell carcinoma (OSCC), throat cancer (e.g., laryngeal cancer, pharyngeal cancer, nasopharyngeal cancer, oropharyngeal cancer)), a one or more leukemias and / or lymphomas known in the art, multiple myeloma (MM)), heavy chain disease (e.g., alpha chain disease, gamma chain disease, mu chain disease), hemangioblastoma, inflammatory myofibroblastic tumors, immunocytic amyloidosis, kidney cancer (e.g., nephroblastoma a.k.a. Wilms' tumor, renal cell carcinoma), liver cancer (e.g., hepatocellular cancer (HCC), malignant hepatoma), lungcancer (e.g., bronchogenic carcinoma, small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC), adenocarcinoma of the lung), leiomyosarcoma (LMS), mastocytosis (e.g., systemic mastocytosis), myelodysplastic syndrome (MDS), mesothelioma, myeloproliferative disorder (MPD) (e.g., polycythemia Vera (PV), essential thrombocytosis (ET), agnogenic myeloid metaplasia (AMM) a.k.a. myelofibrosis (MF), chronic idiopathic myelofibrosis, osteosarcoma, ovarian cancer (e.g., cystadenocarcinoma, ovarian embryonal carcinoma, ovarian adenocarcinoma), papillary adenocarcinoma, pancreatic cancer (e.g., pancreatic adenocarcinoma, intraductal papillary mucinous neoplasm (IPMN), Islet cell tumors), penile cancer (e.g., Paget's disease of the penis and scrotum), pinealoma, prostate cancer (e.g., prostate adenocarcinoma), rectal cancer, rhabdomyosarcoma, salivary gland cancer, skin cancer (e.g., squamous cell carcinoma (SCC), keratoacanthoma (KA), melanoma, basal cell carcinoma (BCC)), small bowel cancer (e.g., appendix cancer), sebaceous gland carcinoma, sweat gland carcinoma, synovioma, testicular cancer (e.g., seminoma, testicular embryonal carcinoma), thyroid cancer (e.g., papillary carcinoma of the thyroid, papillary thyroid carcinoma (PTC), medullary thyroid cancer), urethral cancer, vaginal cancer and vulvar cancer (e.g., Paget's disease of the vulva).

[0180] In some embodiments, the gastrointestinal disease includes, but is not limited to heartburn, irritable bowel syndrome, lactose intolerance, gallstones, cholecystitis, cholangitis, anal fissure, hemorrhoids, proctitis, colon polyps, infective colitis, ulcerative colitis, ischemic colitis, Crohn’s disease, radiation colitis, celiac disease, diarrhea (chronic or acute), constipation (chronic or acute), diverticulosis, diverticulitis, acid reflux (gastroesophageal reflux (GER) or gastroesophageal reflux disease (GERD)), Hirschsprung disease, abdominal adhesions, achalasia, acute hepatic porphyria (AHP), anal fistulas, bowel incontinence, centrally mediated abdominal pain syndrome (CAPS), clostridioides difficile infection, cyclic vomiting syndrome (CVS), dyspepsia, eosinophilic gastroenteritis, globus, inflammatory bowel disease, malabsorption, scleroderma, volvulus, and other gastrointestinal diseases.

[0181] In some embodiments, the congenital disease or disorder includes, but is not limited to amniotic band syndrome, Angelman syndrome, Barth syndrome, chromosomal abnormalities (including, but not limited to abnormalities to chromosome 9, 10, 16, 18, 20, 21, 22, X chromosome, and Y chromosome), congenital adrenal hyperplasia, congenital hyperinsulinism, congenital sucrase-isomaltase deficiency (CSID), cystic fibrosis, De Lange syndrome, fetal alcohol syndrome, first arch syndrome, gestational diabetes, Haemophilia, heterochromia, Jacobsen syndrome, Katz syndrome, Klinefelter syndrome, Kabuki syndrome, Kyphosis, Larsen syndrome, Laurence-Moon syndrome, macrocephaly, Marfan syndrome,microcephaly, Nager’s syndrome, neonatal jaundice, neurofibromatosis, Noonan syndrome, Pallister-Killian syndrome, Pierre Robin syndrome, Poland syndrome, Prader-Willi syndrome, Rett syndrome, sickle cell disease, Smith-Lemli-Optiz syndrome, spina bifida, congenital syphilis, teratoma, Treacher Collins syndrome, Turner syndrome, Umbilical hernia, Usher syndrome, Waardenburg syndrome, Werner syndrome, Wolf-Hirschhorn syndrome, Wolff-Parkinson-White syndrome, and other congenital diseases or disorders.

[0182] In some embodiments, the infectious disease includes, but is not limited to common cold, influenza ( including, but not limited to human, bovine, avian, porcine, and simian strains of influenza), measles, acquired immune deficiency syndrome / human immunodeficiency virus (AIDS / HIV), anthrax, botulism, cholera, Campylobacter infections, chickenpox, chlamydia infections, cryptosporidosis, dengue fever, diphtheria, hemorrhagic fevers, Escherichia coli (E. coll) infections, ehrlichiosis, gonorrhea, hand-foot-mouth disease, hepatitis A, hepatitis B, hepatitis C, legionellosis, leprosy, leptospirosis, listeriosis, malaria, meningitis, meningococcal disease, mumps, pertussis, polio, pneumococcal disease, paralytic shellfish poisoning, rabies, rocky mountain spotted fever, rubella, salmonella, shigellosis, small pox, syphilis, tetanus, trichinosis (trichinellosis), tuberculosis (TB), typhoid fever, typhus, west nile virus, yellow fever, yersiniosis, and zika.

[0183] In some embodiments, the cell of any preceding aspect or the construct of any preceding aspect is administered in combination with a therapeutic agent. In some embodiments, the therapeutic agent includes, but is not limited to an antibiotic, a probiotic, an anti-inflammatory compound, a vitamin, a mineral, or combinations thereof.

[0184] In some embodiments, the antibiotic includes, but is not limited to penicillins (including, but not limited to amoxicillin, clavulanate and amoxicillin, ampicillin, dicloxacillin, oxacillin, and penicillin V potassium), tetracyclines (including, but not limited to demeclocycline, doxycycline, eravacycline, minocycline, omadacycline, sarecycline, and tetracycline), cephalosporins (cefaclor, cefadroxil, cefdinir, cephalexin, cefprozil, cefepime, cefiderocol, cefotaxime, cefotetan, ceftaroline, cefazidme, ceftriaxone, and cefuroxime), quinolones (also referred to as fluoroquinolones include, but are not limited to ciprofloxacin, delafloxacin, levofloxacin, moxifloxacin, and gemifloxacin), lincomycins (including clindamycin and lincomycin), macrolides (including, but not limited to azithromycin, clarithromycin, erythromycin, and fidaxomicin (ketolide)), sulfonamides (including sulfamethoxazole and trimethoprim, and sulfasalazine), glycopeptides (including, but not limited to dalbavancin, oritavancin, telavancin, and vancomycin), aminoglycosides (including, but not limited to gentamicin, tobramycin, and amikacin), carbapenems(including, but not limited to imipenem and cilastatin, meropenem, and ertapenem), and topical antibiotics (including, but not limited to neomycin, bacitracin, polymyxin B, and praxomine) used alone or in combination.

[0185] In some embodiments, the probiotic comprises a food or supplement comprising a beneficial bacterial species including, but not limited to Bifidobacteria animalis, Bifidobacteria breve, Bifidobacteria bifidum, Bifidobacteria lactis, Bifidobacteria longum, Lactobcillus acidophilus, Lactobacillus reuteri, Lacticaseibacillus rhamnosus, Lacticaseibacillus casei, Lactiplantibacillus plantarum, Ligilactobacillus salivarius, Limosilactobacillus fermentum, Lactobacillus paracasei, Lactobacillus gasseri, Lactobacillus acidophilus, Saccharomyces boulardii, Limosilactobacillus reuteri, Bacillus coagulans, or Streptococcus thermophilus alone or in combination.

[0186] In some embodiments, the anti-inflammatory compound includes, but is not limited to, a non-steroidal anti-inflammatory compound including, but is not limited to, aspirin, ibuprofen, ketoprofen, naproxen, steroids, glucocorticoids (including, but not limited to betamethasone, budesonide, dexamethasone, hydrocortisone, hydrocortisone acetate, methylprednisolone, prednisolone, prednisone, and triamcinolone), methotrexate, sulfasalazine, lefunomide, anti-Tumor Necrosis Factor (TNF) medications, cyclophosphamide, and mycophenolate used alone or in combination.

[0187] In some embodiments, the vitamin or mineral includes, but is not limited to, vitamin D, magnesium, vitamin K, vitamin A, riboflavin, vitamin B 12, thiamine, zinc, vitamin B6, biotin, vitamin C, folic acid, vitamin B3, calcium, iron, or derivatives thereof, given alone or in combination.

[0188] In some embodiments, the cell of any preceding aspect or the construct of any preceding aspect is administered in combination with a lifestyle change including, but not limited to, dietary changes, exercise, physical therapy, or combinations thereof.

[0189] In one aspect, disclosed herein is a nucleic acid construct or cell of any preceding aspect and a pharmaceutically acceptable carrier selected from an excipient, a diluent, a salt, a buffer, a stabilizer, a lipid, an emulsion, and a nanoparticle. One or more active agents (e.g., the nucleic acid construct) can be administered in the “native” form, if desired, in the form of salts, esters, amides, prodrugs, a derivative that is pharmacologically suitable, or within a transformed cell. Salts, esters, amides, prodrugs, and other derivatives of the active agents can be prepared using standards procedures known to those skilled in the art of synthetic organic chemistry and described, for example, by March ( \ 992 Advanced Organic Chemistry; Reactions, Mechanisms, and Structure, 4thEd. N.Y. Wiley-Interscience.

[0190] The cell comprising the construct or the native construct may be administered in such amounts, time, and route deemed necessary in order to achieve the desired result. The exact amount of the cell comprising the construct or the native construct will vary from subject to subject, depending on the species, age, and general condition of the subject, the severity of the disease or disorder, the particular composition, its mode of administration, its mode of activity, and the like. The cell comprising the construct or the native construct is preferably formulated in dosage unit form for ease of administration and uniformity of dosage. It will be understood, however, that the total daily usage of the cell comprising the construct or the native construct will be decided by the attending physician within the scope of sound medical judgment. The specific therapeutically effective dose level for any particular subject will depend upon a variety of factors including the disease or disorder being treated and the severity of the disease or disorder; the activity of the cell comprising the construct or the native construct employed; the specific cell comprising the construct or the native construct employed; the age, body weight, general health, sex and diet of the patient; the time of administration, route of administration, and rate of excretion of the specific cell comprising the construct or the native construct employed; the duration of the treatment; drugs used in combination or coincidental with the specific cell comprising the construct or the native construct employed; and like factors well known in the medical arts.

[0191] The cell comprising the construct or the native construct may be administered by any route deemed appropriate to achieve the desired effect. In some embodiments, the cell comprising the construct or the native construct is administered via a variety of routes, including oral, intravenous, intramuscular, intra-arterial, intramedullary, intrathecal, subcutaneous, intraventricular, transdermal, intradermal, rectal, intravaginal, intraperitoneal, mucosal, nasal, buccal, enteral, sublingual; by intratracheal instillation, or bronchial instillation. In general, the most appropriate route of administration will depend upon a variety of factors, including the nature of the cell comprising the construct or the native construct (e.g., its stability in the environment of the gastrointestinal tract), the condition of the subject (e.g., whether the subject is able to tolerate the chosen route of administration), etc.

[0192] The exact amount of the cell comprising the construct or the native construct required to achieve a therapeutically or prophylactically effective amount will vary from subject to subject, depending on species, age, and general condition of a subject, severity of the side effects, identity of the particular compound(s), mode of administration, and the like. The amount to be administered to, for example, a child or an adolescent can be determined by amedical practitioner or person skilled in the art and can be lower or the same as that administered to an adult.

[0193] A number of embodiments of the disclosure have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the invention. Accordingly, other embodiments are within the scope of the following claims.

[0194] By way of non-limiting illustration, examples of certain embodiments of the present disclosure are given below.

[0195] EXAMPLES

[0196] The following examples are set forth below to illustrate the compositions, devices, methods, and results according to the disclosed subject matter. These examples are not intended to be inclusive of all aspects of the subject matter disclosed herein, but rather to illustrate representative methods and results. These examples are not intended to exclude equivalents and variations of the present invention which are apparent to one skilled in the art.

[0197] Introduction

[0198] Synthetic genetic circuits enable the reprogramming of cells to perform novel functions and have advanced the study of natural biological processes with molecular precision. Early efforts in genetic circuit design yielded biological devices such as the toggle switch [F] and repressilator [2’], proving that synthetic biology could be used to emulate natural cellular programs. While the complexity of synthetic genetic circuits has greatly increased in the following years, successful designs have typically required time-consuming and laborious trial and error, despite efforts to develop methods for the predictive design of circuit behaviors. The major challenge of the field is that biological circuit components (e.g., transcription and translation machinery) are not strictly modular like their electronic counterparts, which is further exacerbated by limitations in the composability of genetic elements at the DNA level [3’], [4’]. As a result, the de novo design and quantitative modeling of even the simplest prokaryotic gene circuits have not been achieved using thermodynamic and bottom-up modeling. Empirical modeling and coarse-grained approaches have thus become the preferred methods for designing genetic circuits.

[0199] Control of transcriptional networks has proven to be a reliable technique for genetic circuit design in prokaryotic systems [6’]-[l 1’]. Robust transcriptional control can be achieved using transcription factors (TFs) and CRISPR-Cas-based approaches for activation (CRISPRa) [12’] or inhibition (CRISPRi) [13’] of transcription. Synthetic, engineered TFshave shown to function well as simple OFF-ON gene switches [14’], while natural TFs have been repurposed as NOT gates to allow for layered genetic circuit design [15’]. Combining these systems has led to the development of genetic circuit design automation techniques [7’] that are generally qualitatively accurate but lack precise quantitative prediction.Transcriptional programming (T-Pro) using modular, synthetic TFs has proven to be a powerful and complementary approach for circuit design, particularly due to a reduction in design complexity achieved through circuit compression [I F], [16’], [17’].

[0200] It has been previously demonstrated that combining protein and genetic engineering strategies allows for the development of a scalable genetic circuit platform referred to as T- Pro [11’], [17’], [18’]. By engineering allosteric TFs from the LacI / GalR family, the present disclosure has developed a suite of over 100 transcriptional regulators that can be networked using synthetic promoters. Central to this technique is the use of modular DNA-binding domains (DBDs) and engineered anti-repressors (single-protein NOT gates) that collectively enable circuit compression [16’] (Fig. la). The T-Pro technique for qualitative circuit design allows for the engineering of cellular programs with significantly reduced complexity compared to the state of the art [7’].

[0201] Here, the present disclosure significantly expands the T-Pro technology through the development of a quantitatively predictive method for designing compressed genetic circuits. A cellobiose-responsive anti-repressor has been engineered to enable the design of orthogonal 3-input transcriptional programs with previously developed IPTG- and D-ribose- responsive TFs.

[0202] There are inaccuracies in the frequently used relative promoter unit (RPU) concept and develop a genetic context-specific metric (the relative expression unit; REU) for quantifying expression levels with increased accuracy. By using the REU metric to characterize the cooperativity of T-Pro regulators, the present disclosure provides numerous genetic circuits with unprecedented accuracy in their predicted performances. REU modeling can be applied to metabolic engineering and operons in order to predictively control flux through a toxic biosynthetic pathway.

[0203] Results

[0204] Example #1 - Transcriptional programming with engineered cellobiose anti- repressors

[0205] Figs, la - le shows the engineering of a cellobiose anti -repressor to expand the transcriptional programming design space. Fig. la shows an example of circuit compression. Fig. lb shows the engineering of a corresponding anti-celR. Fig. 1c shows the performanceof variants with lacRBS regulating GFP. Fig. Id shows the global design space for transcriptional programming. Fig. le shows the constrained design space based on orthogonality.

[0206] Many diverse transcriptional programs have been constructed using engineered TFs, facilitating Boolean, analog, and sequential circuit behaviors [11’], [19’], [20’]. However, complementary repressor / anti-repressor TF pairs that respond to a common input signal are necessary to create compressed circuits [16’], [18’] (Fig. la). Four repressor / anti-repressor pairs have been previously developed using a protein engineering workflow [17’], [18’], [21’]. However, some engineered TFs exhibited either a low dynamic range or a response to multiple inducers, the latter feature enabling the creation of bandpass and bandstop programs [17’]. In order to expand the capacity of orthogonal transcriptional programming, an additional repressor / anti-repressor pair was developed, specifically by engineering a cellobiose anti-repressor from the chimeric CelR repressor equipped with a synthetic DBD [11’], [19’].

[0207] Here, the present disclosure attempted to directly evolve an anti-repressor using error- prone polymerase chain reaction (EP-PCR) on CCIRTAN. After generating a library with ~108variants, the desired phenotype after screening with FACS was not identifiable.

[0208] A previously developed workflow for engineering anti-repressors was then re- explored that employs the generation of a super-repressor mutant, which can then be evolved using EP-PCR [18’], [21’] (Fig. 1b). It was contemplated that the success with the technique would be predicated on the manipulation of allosteric communication within the protein [22’]. Site saturation mutagenesis was performed on residue L75 (L78 in wild-type CelR) based on a previous report implicating this position in the allosteric communication of CelR [23’]. Mutant L75H displayed the desired super-repressor phenotype (Fig. Ib-lc), and this template was then used to perform EP-PCR. Screening the new library with FACS yielded three unique anti-repressors ( EA1, EA2, EA3, provided as SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 5), highlighting the importance of the initial super-repressor mutation in evolving the anti-repressor phenotype (Figs. Ib-lc). The anti-CelRs were validated to be equiped with four additional DNA-binding domains, DBDs, and maintain the anti-repressor phenotype (Fig. 2). Specifically, Fig. 2 shows validation of anti-CelRs (shown as Anti-CelR 1, 2, 3) equipped with alternate DBDs. With this result, the transcriptional programming design space was expanded to include five repressor / anti-repressor pairs that can be equipped with seven unique DBDs (Fig. Id). Moving forward, a constrained design space was developed based on orthogonality of ligand inputs and DBDs (Fig. le).

[0209] Genetic context dictates the non-modularity of circuit components. Synthetic genetic circuits can be modeled from first principles by accounting for the theoretical interactions between protein, DNA, and RNA components involved in the circuit. However, the absolute quantification of these macromolecules in living cells is difficult to achieve, requiring specialized techniques that are typically low throughput and costly to implement. This poses a challenge for those designing synthetic genetic circuits where it is necessary to precisely control the levels of individual biomolecules that often have complex interactions with one another. As a result, the most widely adopted technique for gene circuit design is the use of fluorescent protein reporters (e.g. , green fluorescent protein; GFP) as a proxy for promoter strength. This led to the development of the RPU concept, where GFP fluorescence is assumed to correlate with transcriptional flux from a promoter [24’]. Transcriptional circuits can thus be modeled as the regulation of RNA polymerase (RNAP) flux from promoter inputs and promoter outputs, typically achieved through protein-based regulators of transcription e.g., TFs and CRISPRi), and accounting for the interactions between regulators, nucleic acids, RNAP, and ribosomes.

[0210] One major limitation of the RPU concept is the phenomenon of local genetic context influencing the behavior of a circuit “part” (including, but not limited to, promoters, ribozymes, RBSs, genes, and terminators). For instance, an RBS may be considered “strong” when paired with a particular promoter and gene of interest (GOI), but the same RBS may appear to be “weak” if the promoter or GOI sequence changes, potentially due to alternative folding of the mRNA or the introduction of unintended promoters, for example. Similarly, the use of hammerhead ribozyme insulators[4’] has become common practice as they are intended to “normalize” mRNA structures by removing extraneous 5’ UTR sequences after transcription. However, we have found that the use of sequence-distinct ribozymes that have the same biophysical function can impact the actual expression level of a GOI from a given promoter owing to the local DNA (and resulting mRNA) sequences that arise during the composition of the expression cassette.

[0211] To illustrate how the RPU metric provides a misguided measurement of promoter activity, five repressible promoters were developed, and each was assigned a unique ribozyme insulator (Fig. le). Each promoter has one of five synthetic operators (O1, tta, gta, ttg, or agg) inserted between the -35 and -10 hexamers (core position) and an identical operator inserted downstream of the transcription start site (TSS, proximal position). We measured the constitutive expression of a GFP reporter gene paired with a putatively strongRBS compared to a chimeric reporter gene consisting of a fusion of the first 25 amino acids (AAs) of LacI to GFP paired with the cognate lac RBS (Fig. 3a).

[0212] Figa. 3a - 3d shows modeling of genetic context is critical for the prediction of gene expression. Fig. 3a shows GFP reporter inaccurately predicts expression level via RPU “promoter strength” not consistent when RBS and 5’ region of gene vary. Fig. 3b shows the leader sequence of gene is most important for expression level; REU concept generalizes to different FPs. Fig. 3c provides a demonstration of the non-modularity of RBS even with fixed promoter, ribozyme, and gene. Fig. 3d provides an example of REU nanoluc. Fig. 3e shows the standard curve. Fig. 3f shows that the REU correlates with real protein expression levels. Fig. 3g shows the REU lac circuit. Fig. 3h shows the standard curve. Fig. 3i shows that the REUlac and REU nano correlate from the same promoter over different transcript levels, bringing attention to SI fig comparing 5 promoters with constitutive LG vs. NG expression (the ratio is not conserved).

[0213] If expression patterns were conserved between the two GFP reporters, a consistent expression ratio can be expected to be observed from the promoters when scaled to the respective reporter genes. However, it is found that the expression patterns varied not only in magnitude depending on the reporter gene but also in relation to the other promoters. It is posited that this is largely due to differences in mRNA processing and folding when the RBS and initial region of the transcript are changed between the genetic constructs. It is also possible that the difference in DNA sequence between the constructs can affect the relative promoter activity through the introduction or removal of cryptic promoters [25’].

[0214] There is further evidence that the difference in expression patterns is specifically dictated by the leader sequence of mRNA because replacing GFP in the LacI fusion construct with mKate, a red fluorescent protein, yields a conserved expression pattern across all five promoters (Fig. 3b). The fluorescence of LacLGFP fusion reporters were characterized ranging from 0-24 AA of the LacI gene to determine the point of stabilized expression. Protein expression varied significantly until AA 16, after which a stable pattern of expression was achieved (Fig. 5). Furthermore, a de novo RBS library for the O1promoter was characterized to achieve a series of different expression levels for our LacLGFP reporter and applied the same library to the other four promoters to examine the modularity of the RBS sequence alone. The relative RBS “strengths” were not conserved across promoters, emphasizing a lack of modularity (Fig. 3c, also see Fig. 6). Finally, unique ribozyme sequences were demonstrated also affect expression levels dramatically even when the core promoter and RBS sequences are unchanged (Fig. 7). With these results in mind, it wasproposed that the RPU be replaced by a new metric for relative expression units (REU) which can be defined as the measured fluorescence of a GFP fusion to the first 25 AA of a GOI, specifically in the context of a promoter-ribozyme-RBS combination. In principle, the REU could be converted to an absolute measurement through fluorescent protein calibration curves [26’].

[0215] To confirm that the REU provides an accurate measurement of the true expression level of a GOI, an REU standard curve was generated for Nanoluc [27’] luciferase using a synthetic LuxR-activated promoter (pLuxm) that produces a graded response to 3OC6 homoserine lactone (AHL) (Fig. 3d-3e). The Nanoluc-GFP fusion reporter was replaced with the full Nanoluc gene and measured the catalytic activity via a luminescence assay. The REU fluorescence and Nanoluc luminescence showed near perfect correlation (R2= 0.9958), confirming that the REU provides an accurate measurement of gene expression (Fig. 3f).Next, an REU standard curve was generated for our set of chimeric TFs by measuring the fluorescence of the LacI-GFP reporter expressed from the pLux circuit (Fig. 3g-3h). Because the TFs used for T-Pro share the LacI DBD (up to AA 16), we posited that this single standard curve would be applicable to all TFs. A standard Hill function was fitted to these data to convert AHL concentration to a pLuxREU standard curve (Equation 1).(Eq. 1)

[0216] In Equation 1, β is a basal expression level, y is the span of the Hill function, x is the AHL concentration, x50is the EC50 AHL value, and n is the Hill slope. REUluxcan be defined as the REUlux= relative expression units

[0217] Genetic context is shown with this experiment, as the Nanoluc and LacI REU standard curves showed unique expression patterns, with the only experimental difference being the first 25 AA of the reporter genes. When compared with the Nanoluc and LacI REU standard curves, it was observed that there is a strong linear correlation, indicating that the expression levels scale over a range of transcription rates (Fig. 3i).

[0218] Predictive design of fundamental transcriptional programs. Figs. 4a-41 show validation of the accuracy of fundamental gene circuit predictions. The REU metric was applied to the predictive design of transcriptional programs by systematically characterizing the regulatory performance of 30 TFs (Fig. le) over a wide range of expression levels. Transcriptional programming requires precise control over the expression levels of TFs fromboth constitutive and regulated promoters, as underexpression leads to weak transcriptional regulation, while overexpression can limit the dynamic range of a given TF. To this end, a titration circuit was developed that uses the pLuxmpromoter to vary the expression level of a TF that, in turn, regulates its cognate synthetic promoter (Fig. 4a). The LacI REU standard curve allows the conversion of the AHL dose to an input REU value (Fig. 3g-3h). The circuit was characterized with and without the cognate inducer of the TF to generate Hill functions (Equations 2-3) describing the performance of the regulator (Fig. 4b). A performance metric (Equation 4) was then calculated to determine the input REU level that yields optimal TF regulation (Fig. 4c). Equations 2 and 3 are the Hill functions for the TF titration circuit.(Eq. 3)

[0219] In Equations 2 and 3, REU0N / 0FFis the relative expression units of the GFP reporter. The ON / OFF subscript denotes the ON or OFF state of the GFP reporters based on + / - inducter for anti-repressors. REUlux 50is the EC50 REUiux value. All 30 TFs were characterized using the titration circuit, revealing unique inducibility patterns for the different regulators (Fig. 4b, also see Fig 9). Fig. 9 shows TF titration curves.

[0220] From this dataset, it was posited that fundamental BUFFER and NOT gate circuits could be modeled by constitutively expressing TFs at their optimal REU values based on a performance metric φ in Equation 4. φ is the effective dynamic range of a TF in terms of span.(Eq. 4)

[0221] To achieve this, experiments characterizing a context-specific RBS library were performed to provide a wide range of expression levels encompassing the span of the LacI REU standard curve (Fig. 4d-4e). Specific RBS sequences were then assigned to each TF to approximate their optimal performance REU values when regulating their cognate promoter(Fig. 4f). These 30 fundamental transcriptional programs were then built and tested to validate the predictive model (Fig. 4g-41). The data showed highly accurate predictions from the model, with an average error of <1.4-fold (-40%). The results confirmed that the REU metric can be used for the quantitative prediction of expression levels in fundamental circuits involving two genes. While the chosen DBDs are putatively orthogonal based on previous studies, we reevaluated the DBD-operator affinities of the 30 TFs (120 potential non-cognate interactions) for our current system. As expected, the majority of noncognate DBD-promoter pairings showed insignificant interactions, but rare pairs did exhibit binding which is important to consider in the design of more complex circuits (Fig. 10). Fig. 10 shows TF orthogonality.

[0222] Example #2 - Multi-input transcriptional programs

[0223] Figs. 8a and 8b show the predictive design of complex circuits. Fig. 8a shows an example workflow for predictive circuit design. Fig. 8b shows example 3 input XNOR gate. The program can be employed for 2 input gates, 3 input gates, 4 input gates, 5 input gates, among others.

[0224] Based on the success of modeling single-input transcriptional programs, the model was applied to the design of the 10 nontrivial 2-input programs (i.e., AND, NOR, NIMPLY, NAND, OR, IMPLY, XOR, and XNOR). The workflow for this design process is illustrated in Fig. 8a. A user first provides a truth table mapping inducer input states to output gene expression states. A generic circuit topology can then be assigned based on transcriptional programming rules (*expand in Supplement, hopefully with lookup table generated by ISYE lab). Multiple iterations of the circuit can then be simulated in silico by assigning different combinations of promoters, RBSs, and DBDs to the different regulators. (A circuit score can be calculated as in Cello to determine the optimal circuit architecture.) The model first predicts the outputs of promoters regulated by constitutively expressed TFs as described above. These output REU values are then linearly scaled based on RBS libraries characterized for each inducible promoter (Fig. 3c and Fig. 6). Linear scaling of REU is predicated on the assumption that transcription rates do not change when different RBSs or GOIs are combined with a promoter-ribozyme pair (see Equations 5-10 for a derivation of this assumption).

[0225] Modeling Gene Expression can be expressed as:(Eq 5)

[0226] In Equation 5, [mi is the concentration of mRNA I, αm iis the transcription initiation rate, [Gi is the concentration of gene, and Yi is the degradation rate of mRNA i.(Eq 6)

[0227] In Equation 6, [Pi is the concentration of protein i, αP iis the translation initiation rate, is the dilution rate of protein I, and t is the doubling time of strain.

[0229] In Equation 10, Jiis the lumped parameter for promoter flux. The model assumptions include (i) the promoter flux corresponds to production of full-length transcripts, (ii) an average mRNA degradation rate can be assumed (0.00407 1 / s) from literature, (iii) the transcription initiation rate not being influenced by changing the RBS sequence or GOI, and is constant as long as a promoter-ribozyme combination is maintained, and (iv) the protein dilution rate being dominated by cell division. This allows protein expression levels to be linearly scaled from a promoter by changing translation initiation rate.

[0230] For promoters regulating a TF, their output REU is scaled to match the ideal REU of the TF being expressed through the RBS assignment. The ON / OFF states of the final reporter gene can then be calculated based on the corresponding TF titration curve. In the instance where multiple TFs regulate a single promoter simultaneously, the predicted output corresponds to the lower of the two TF response functions (Equation 11).

[0231] For SEPA Regulation:

[0232] In Equation 11, subscript “IC” indicates a particular induction condition (i.e., inducer combination). Subscript “TFn” indicates the REU corresponding to a particular TF regulating a promoter, with “n+1” corresponding to a unique TF co-regulating the same promoter. Simply, for a SEPA regulation scheme, the REU level of an induction condition will equal the minimum of the set of REUs of that promoter when regulated by the single TFs.

[0233] The workflow was executed for the design of the remaining 2-input programs and observed strong predictive performance with an average error of only -40% (Fig. 8a and Fig. 11). Fig. 11 shows additional examples of 2-input circuits.

[0234] An exemplar 3 -input program (3 -input XNOR, or Consensus) was then chosen that required 45 parts* to construct via Cello programming [7’] to test the accuracy of our transcriptional programming model. Notably, the design of this program using transcriptional programming only requires 20 parts* (Fig. 8b) (*definition of part used by Voigt is different from those used herein). After constructing and testing this program, no state was greater than -1.5-fold off from the prediction, which is substantially more accurate than the Cello design, where most states were > 10-fold off from the prediction. This program employed the simultaneous prediction of the expression levels of seven interacting proteins.

[0235] e. Figs. 12a - 12i show expression of operons can be modeled for metabolic and strain engineering. Table 1 shows the operons of Fig. 12a.Table 1

[0236] A useful application of T-Pro is the control of metabolic pathways for biomanufacturing. The lycopene biosynthetic pathway is provided as an example to demonstrate how the REU metric can be used to inform the design of a multistep synthesis. The production of lycopene in E. coli can be achieved by expressing three heterologous genes (crtE, crtB, and crtl)28(Fig. 12a), which are often organized into an operon [29’]-[31Additionally, overexpression of the native E. coli proteins Dxs and Idi has been shown to improve lycopene titer by increasing precursor concentrations [30’], [32’], [33’]. To control lycopene production, a synthetic crtEBI operon [34’] was initially placed under the control of R+YQR but experienced issues with toxicity during cloning and assaying of the circuit.

[0237] Attempts to clone the CrtEBI operon led to spontaneous mutations and DNA deletions or insertions which we attributed to toxicity caused by high expression levels. This was mitigated by adding a transcriptional regulator to minimize expression during cloning steps. Successfully cloned circuits then underwent spontaneous mutation during induction, validated by DNA sequencing post-induction. These issues were fixed only by lowering the expression levels for all Crt genes.

[0238] While others have also reported challenges with toxicity when expressing this pathway

[0035] -

[0037] , it was reasoned that the REU metric could be used to characterize the production of each gene and then rationally engineer the expression levels using fusion reporters. To achieve this, each gene was placed under the control of a different T-Pro regulator and determined the dose-response functions for each expression cassette using GFP fusion reporters (Fig. 12b-12c). A refactored circuit was then assembled to express the full- length genes, along with the Dxs and Idi proteins (Fig. 12d). To express each crt gene at a low REU value of 100 (noting that, in the system, the maximum REU measured was -50,000), the appropriate inducer concentrations (xx uM D-Ribose, xx uM cellobiose, and xx uM IPTG) was provided. This allowed a lycopene titer of 350 ng / ml to be generated in batch culture and mitigated the issues with genetic instability of the circuit (Fig. 12e).

[0239] It was then hypothesized that the REU metric could be applied to operons as well as single expression cassettes, which would be valuable for the predictive design of multi-gene pathways. It was posited that three key factors govern the expression levels of genes in a polycistronic mRNA and would be important to model. First, the length of the transcript is known to be important, with upstream genes potentially experiencing a higher copy number compared to downstream genes due to RNAP generating a range of transcript lengths. Second, ribosome reinitiation has been shown to be a significant contributor to the coupling of gene expression in synthetic operons [38’]. Third, as with single expression cassettes, the local mRNA structure around an internal RBS should be critical in governing the translation initiation rate of the associated gene.

[0240] With these factors in mind, three unique GFP fusion constructs were designed to quantify the REU of each gene in an operon context, with the goal of matching the 100 REU level used above for each gene (Fig. 12f). Each GFP fusion was designed to retain all DNAupstream of the GOI, including the first 25 AA of the GOI as before. At the outset, libraries for each of the three RBSs were generated and screened for variants with an REU of -100. The final expression levels were close to the target, with REUs of 100, 100, and 50 for CrtE, B, and I, respectively (Fig. 12g). A functional operon was assembled under control of R+YQR and measured the lycopene titer to be 360 ng / ml, which was close to the titer achieved with the refactored system expressed at the same levels (Fig. 12h-12i). The consistency in lycopene production between the refactored circuit and operon demonstrated that our REU metric can be used to quantitatively measure gene expression in polycistronic mRNAs.

[0241] Discussion

[0242] The applications of synthetic biology are anticipated to have major impacts ranging from improved therapeutics to mitigation of the current climate crisis [39’]-[45’]. However, the field has not realized its full potential due to the challenges involved with the design of genetic circuits — the fundamental method used to reprogram cells. The instant study developed the most accurate genetic circuit design platform to date, enabled by the use of combined protein and genetic engineering strategies along with the key adoption of the REU metric. While the accuracy of the REU metric was proven in the context of the present disclosure, there are potential limitations, such as the use of 25 AA fusion proteins to model particularly long genes or multimeric proteins that have long maturation times. Additionally, if the native full-length gene has unintended internal promoters, RBS, or terminators, this could reduce the accuracy of the fusion reporter. Fortunately, these are well-known phenomena that can be solved by strategic codon optimization of a GOI.

[0243] The present disclosure demonstrated the power of using empirical modeling to accurately predict the expression levels of diverse proteins (DNA-binding proteins and multiple enzymes). There have been significant efforts to develop DNA sequence-based models for the design and control of transcription [46’] - [49’] and translation [50’], [51’], but they have not progressed to the point where their accuracy surpasses experimental characterization of unique constructs (i.e., library screening). To illustrate this, the most recent versions of the RBS

[0050] and Promoter Calculators

[0049] were used to predict the expression level of 75 constructs tested in this study. The data correlated poorly with the model prediction, indicating that the instant operation for experimental characterization is superior to sequence-based model predictions of gene expression (Fig. 13).

[0244] Fig. 14 shows a comparison of DNA sequence-based models of expression

[0245] Method

[0246] The development and optimization of the transcriptional programming platform based on the use of LacI / GalR TFs as a general strategy can be used to apply combined protein and genetic engineering to other classes of regulators. For example, TetR family TFs have also been shown to be amenable to alteration of the DBD [52’] and ligand-binding region [53’], as well as being engineered to have the anti-repressor phenotype [54’]. Rational design can be combined with deep-learning approaches to streamline the process of protein engineering in order to control the ligand and DNA interactions of synthetic TFs. Such an approach would democratize the field of genetic circuit design through a standardized workflow, greatly expanding the potential for programming advanced behaviors into cells.

[0247] Bacterial strains and media. E. coli strains used were NEB® 10-beta (for cloning) and 3.320 (JacZ13(Oc) lacI22 λ- el4- relAl spoTl thiE1, Yale CGSC #5237) (for assays). E. coli were routinely cultured aerobically in LB Miller medium (Fisher BP9723) at 37°C (unless otherwise specified), in M9 minimal medium (MM) (MM contains 3 g / L KH2PO4, 0.5 g / L NaCl, 6.78 g / L Na2HPO4, 1 g / L NH4C1, 0.1 mM CaCh, 2 mM MgSO4, 1 mM thiamine hydrochloride, 0.4% D-glucose, and 0.2% casamino acids), or on LB Miller agar (Fisher BP1425). Antibiotics for plasmid selection were used at the following concentrations: carbenicillin (Goldbio C-103-25)- 100 pg / ml; chloramphenicol (Goldbio C-105-25)- 25 pg / ml; kanamycin (Goldbio K-120-25)- 35 pg / ml.

[0248] Chemical inducers. The following chemicals were used as inducers: Isopropyl-beta- D-thiogalactoside (IPTG, Goldbio 12481C); D-ribose (D-rib, Alfa Aesar Al 7894); Cellobiose (cello, Acros Organics 108461000); 3-Oxohexanoyl-homoserine lactone (AHL, Sigma K3007). Unless otherwise specified, the final concentrations used for each inducer were: 10 mM IPTG; 10 mM D-rib; 10 mM cello; 0.1 nM-10 μM 3OC6 AHL.

[0249] Cloning and plasmid construction. Plasmids were created using Golden Gate assembly [55’], inverse PCR followed by blunt-end ligation, or Gibson cloning [56’]. Q5 polymerase (NEB M0491L) was used for PCR. T4 DNA ligase (NEB M0202L), BsmBI-v2 (R0739L), and BsaI-HFv2 (NEB R3733L) were used for Golden Gate cloning. NEBuilder HiFi DNA Assembly Master Mix (NEB E2621X) was used for Gibson cloning. All DNA primers were synthesized by Eurofins Genomics. The DNA sequences of all constructs were verified by Sanger sequencing or whole plasmid sequencing (Eurofins Genomics).

[0250] Fluorescence assay. Cells were transformed with the plasmid(s) harboring a circuit of interest and selected on LB agar with the appropriate antibiotics for plasmid maintenance. After overnight incubation, individual colonies were used to inoculate 200 μl LB cultures (with appropriate antibiotics), which were grown in a flat-bottom 96-well plate (Corning3370) sealed with a Breathe Easier membrane (Electron Microscopy Sciences 70536-20). After overnight growth (-16-20 hours) in a Thermo Scientific MaxQ 4000 shaker at 300 rpm, cells were diluted 1 :200 into fresh LB medium and grown for an additional 8 hours. Cells were then diluted 1 :200 into MM with the appropriate inducer(s) and grown for 12 hours. After this final growth period, 100 pl of each culture was transferred to a black-walled, clear- bottom 96-well plate (#). OD600 absorbance and GFP fluorescence were then measured using an M2e Spectramax spectrophotometer (485 / 510 nm excitation / emission).

[0251] Example #3 - Performance Prediction of Fundamental Transcriptional Programs

[0252] Introduction

[0253] Significant efforts have been devoted to engineering logical (decision-making) responses within a variety of chassis cells as a general proof-of-concept [l]-[9], and for a variety of logic-based applications - e.g., biosensing [7],

[0010] ,

[0011] , biological clocks

[0012] -

[0015] , oscillators

[0012] ,

[0016] -

[0019] , controllers [5],

[0020] -

[0022] , and therapeutics

[0023] -

[0028] , An emerging technology in biotic decision-making is transcriptional programming

[0028] -

[0030] ,

[0254] Transcriptional programming makes use of fundamental logic principles by assigning an inducer molecule as the INPUT, and by assigning a coupled regulated reading frame (coding or non-coding) as the OUTPUT. The operating constraints for said biotic programs are predicated on digitizing the INPUT to 0 or 1, where an INPUT 1 is achieved via the maintenance of saturating concentrations of the cognate inducer molecule - typically 10 mM. Digitizing the INPUT facilitates a constant level of OUTPUT - e.g., the amount of green fluorescent protein (GFP) is present at a steady state.

[0255] Figs. 15a -15b show modular components used in a design space and method thereof. The fundamental 1 -INPUT logical operations in transcriptional programming are: i) BUFFER gates regulated via engineered repressors (Fig. 15a), and ii) NOT gates regulated via engineered anti -repressors (Fig. 15b). Notably, antirepressors are an important and unique feature of transcriptional programming in that said transcription factors enable circuit compression. That is, the anti-repressor eliminates the need for the inversion of a repressor function to achieve the said logical operation. Anti-repression versus Inversion is provided below.

[0256] Another important feature of transcriptional programing is the ability to direct two or more engineered transcription factors to a single DNA operator element - enabling the systematic construction of 2-INPUT logical operations, see Fig. 15f, 15g. The design workflow for the 2-input operation is provided below..

[0257] Anti-repression versus Inversion. Inversion is a process in which a single repressor is expressed on one layer and is directed to interact with a cognate DNA element to reject an output located on a second layer, and can be regarded as a NOT operation - notably, Cello circuits are constructed via the said inversion process 1. In contrast, anti-repressors reduce the NOT operation to a single layer and single promoter, and the reduction in components (e.g., promoters) is defined as circuit compression (Fig. 15b and Fig. 26).

[0258] Design Workflow for 2-INPUT Operations. To construct single layer 2-INPUT operations from 1 -INPUT operations requires the use of engineered transcription factors and engineered cognate genetic architectures, see Fig. 15. Engineered transcription factors were developed via modular design, which enables the development of synonymous DNA binding functions for two transcription factors that process two different INPUTS. In turn, coupled DNA functions between two engineered transcription factors can be directed via a SE-PA or SERI genetic architecture (Fig. 27) to facilitate the construction of a 2-INPUT operation.

[0259] The engineered transcription factors used in transcriptional programming were developed via modular design (Fig. 15c). The engineered transcription factors via modular design are provided below. Briefly, the design template is based on the lactose repressor (LacI) topology, which can be decomposed into two functional regions: i) a regulatory core domain (RCD), and ii) a DNA binding domain (DBD).

[0260] Engineered Transcription Factors via Modular Design. The design template LacI may be a part of a large family of proteins that share a topology and putative mechanism of action 2. The LacI / GalR protein family is made up of over 1,000 homologues. Moreover, the LacI / GalR transcription regulatory proteins mediate responses to a wide range of environmental and metabolic changes. Structurally, the general LacI / GalR topology can be defined by two fundamental domains - i.e., (i) a regulatory core domain and (ii) a DNA binding domain. Accordingly, we can regard this collection of paralogues as a putative design space - when carefully decomposed - positing that said functional domains can be mixed and matched to form new allosteric transcription factors.

[0261] Given that LacI belongs to a large family of homologous transcription factors with similar topology that can process different INPUT ligands and bind to different DNA operators

[0031] ,

[0032] , a putative design space can be gleaned. Accordingly, several groups have demonstrated that functional chimera can be constructed based on said engineering principles

[0029] ,

[0033] -

[0037] ,

[0262] Wilson et al. in a collection of studies, posited and demonstrated that modular design could be applied to engineered domains - that is, alternate (engineered) DNA bindingfunctions

[0029] ,

[0035] , alternate (engineered) allosteric communication

[0030] ,

[0035] ,

[0037] ,

[0038] , and alternate (engineered) ligand binding functions

[0036] ,

[0037] - to create a system of transcription factors (repressors and anti-repressors) that are network capable. To date, using this collection of engineered repressors and antirepressors, transcriptional programs are designed and constructed intuitively - though with an apparent rule set. Descriptions of transcription programming are provided below.

[0263] Transcriptional Programming. Transcriptional programming is predicated on a definitive bottom-up combinational rule set. Single-input single-output operations (BUFFER and NOT) represent the fundamental binaries, that can be systematically combined to create all proper two-input single-output operations. Complex circuit development via transcriptional programming (e.g., OR, NAND, A IMPLY B, B IMPLY A, XOR, and XNOR) involve feeding forward information [3] may be performed.

[0264] As a result, transcriptional programs often require iterative tuning and re-design. While intuitive program design and construction have proven to be effective, the said approach is time-consuming and expensive. What is needed now is a means to predictively design transcriptional programs - i.e., in terms of qualitative outcomes and quantitative performances. To accomplish the aforementioned, the ins / tant study leveraged and built upon the model introduced by Zong et al.

[0039] to predict Multipl e-INPUT Single-OUTPUT (MISO) logical operations, from Single-INPUT Single-OUTPUT (SISO) data, without requiring parameter fitting from MISO experiments. Namely, the study systematically designed, built, and tested a large collection of BUFFER SISO and NOT SISO with corresponding metrology for said fundamental logical operations. In turn, the study leveraged the standardized SISO data to design, build, and test the corresponding set of MISO logical operations (via transcriptional programming) allowed at a single operator-promoter position - that is, forming AND, NOR, A NIMPLY B, and B NIMPLY A operations.

[0265] Finally, the study showed that simple (coarse-grained) models can qualitatively design and quantitatively predict the fundamental performances of MISO logical operations from SISO data only - establishing the foundation for the predictive design of transcriptional programs.

[0266] Experimental Results and Examples

[0267] The study developed and evaluated the exemplary system. Figs. 16a - 16b show example combinatorial set of SE-PA AND gates conducted in the study.

[0268] MATERIALS AND METHODS

[0269] Example: Cloning BUFFER and NOT plasmids. Each SISO system comprised (1) a single transcription factor expressed on the pLacI plasmid (Novagen), which contained the pl5a origin (copy number 20-30 / cell), and (2) a super folder green fluorescent protein (GFP) reporter expressed on the pZS*22-sfGFP plasmid which contains the pSClOl origin (copy number 3-5 / cell). Chloramphenicol and kanamycin resistance genes were used as selection markers for transcription factor and reporter plasmids, respectively. Transcription factor and reporter plasmids were taken from previous works (Rondon et al., Groseclose et al.) and when necessary, ADR or operator variants were cloned using site-directed mutagenesis PCR (Phusion DNA Polymerase, NEB) with custom primers (Eurofins Genomics) followed by kinase, ligase and Dpnl reactions (KLD enzyme mix, NEB). The reactions were transformed into chemically competent DH5a cells (huA2 A(argF-lacZ)U169 phoA glnV44 cp80A(lacZ)M15 gyrA96 recAl relAl endAl thi-1 hsdR17; New England Biolabs) and plated on LB agar with appropriate antibiotic. A transformant was cultured overnight and mini- prepped (Omega Bio-Tek) to yield each plasmid, and the sequence was confirmed with DNA sequencing (Eurofins Genomics). A LacSTOP control plasmid, which contained a LacI gene with ochre mutations at codons 2 and 3, was also cloned using this site-directed mutagenesis protocol.

[0270] Example: Cloning AND, NOR, andNIMPLY transcription factor plasmids. Transcription factor plasmids used in MISO systems were identical to those in SISO systems, except contained two independently driven transcription factor genes. AND, NOR, and NIMPLY transcription factor plasmids were cloned using a mix-and-match Golden Gate Assembly method. Transcription factor inserts were PCR amplified (Q5 DNA Polymerase, NEB) using BUFFER and NOT plasmids as templates, gel extracted (Qiagen), and desired pairs were matched and assembled with BsmBI-v2 and T4 DNA ligase (BsmBI-v2 Golden Gate Assembly Kit, NEB). The resulting plasmids were transformed and isolated according to exemplary methods.

[0271] Example: GFP microwell plate assay. For each logic gate, the transcription factor plasmid contains either a single repressor (BUFFER), single anti-repressor (NOT), repressor pair (AND), antirepressor pair (NOR), or repressor / anti-repressor pair (NIMPLY). Transcription factor and corresponding reporter plasmids were double transformed into homemade chemically competent 3.32 E. coli cells (Genotype lacZ13(Oc), lacI22, LAM-, el4-, relAl, spoTl, and thiEl, Yale CGSC #5237) and transformants were precultured for 6 hours in LB media with chloramphenicol (25 pg / mL, VWR Life Sciences) and kanamycin (35 pg / mL, VWR Life Sciences) antibiotics.

[0272] Precultures were then diluted in sextuplicate into glucose (100 mM, Fisher Scientific) M9 minimal media supplemented with 0.2% (w / v) casamino acids (VWR Life Sciences), 1 mM thiamine HC1 (Alfa Aesar), antibiotics, and respective inducers, and grown in a flat bottom 96-well microplate (Costar) for 16 hours (37 °C, 300 rpm). Microwell plates were sealed with Breathe-Easy membranes (Diversified Biotech) to prevent evaporation. Inducer concentrations used are as follows: isopropyl-β-D-thiogalactoside (IPTG; 10 mM, reduced to 1 mM for IPTG-fucose gates), D-ribose (10 mM), cellobiose (10 mM), D-fucose (10 mM), fructose (10 mM), and adenine (1 mM). Optical density (OD600) and GFP fluorescence ( λ!# = 485 nm, λ!# = 510 nm) were measured with a Spectramax M2e plate reader (Molecular devices).

[0273] The LacSTOP control plasmid was also assayed with each reporter construct to determine the maximum expression level of each genetic architecture. Measurements were corrected by subtracting values of blank media from sample values, and fluorescence values were normalized to optical density in Microsoft Excel (Microsoft).

[0274] Data analysis and model predictions. All fluorescence data was normalized to a global maximum of 75,000 relative fluorescence units (RFU). A two-tailed t-test was used to determine statistical significance between ON and OFF states for each BUFFER and NOT gate (significance level = 0.001). Gates with a p-value > 0.001 were regarded as not significant and classified as either non-functional (X-) or super-repressor (XS) phenotypes. Compatibility tests and OUTPUT predictions for AND, NOR, and NIMPLY gates were performed in Microsoft Excel (Microsoft) using the appropriate models and experimental BUFFER and NOT data (see Supplemental Data).

[0275] First, σ and ε values were calculated using normalized ON and OFF state OUTPUTS. Model parameters α0, α1, α2 and α3 were then evaluated using σ and ε values and plugged into the respective model equation for each 2-INPUT gate. Prediction values and experimental data were plotted using GraphPad (Prism) for correlation analysis. Prediction error was calculated for each INPUT condition across all gates as the ratio of measured OUTPUT to predicted OUTPUT.

[0276] Insulated SE-PA and SERI logic gates. For each proximal AND and NOR RCD pair, the ADR variant with the largest prediction error (determined as the magnitude of fold change, averaged across all four INPUT conditions) was selected for the insulated genetic architecture case study. Insulated reporters were cloned using site-directed mutagenesis PCRas described previously using a template (Oaggcore Otgproximal RiboJIO GFP reporter, provided by Groseclose et al.).

[0277] Transcription factor and insulated reporter plasmids were double transformed, and both SISO (BUFFER or NOT) and MISO (AND or NOR) operations were constructed and assayed as described previously. For SERI gates, core operators were inserted with site- directed mutagenesis PCR (using SE-PA reporters as templates), and both SISO and MISO operations were constructed and assayed. Data analysis and model predictions for insulated SE-PA and SERI gates are described in the methods above.

[0278] RESULTS

[0279] Design, Metrology, and Modeling for Single -INPUT Single-OUTPUT (SISO) Logical Operations. In previous studies, a metrology was established for SISO XfDRBUFFER

[0029] and SISO XfDRNOT

[0030] gate performance - which was extended and further developed in the instant study. For a given engineered transcription factor X (or Y) defines the regulatory core domain (RCD), the superscript “+” defines the repressor phenotype (Fig. 15a), the superscript “A” defines the anti-repressor phenotype (Fig. 15b), and the subscript defines the alternate (engineered) DNA recognition (ADR) function. The putative design space for said engineered transcription factors is given in Fig. 15c.

[0280] Briefly, given a transcription factor and cognate operator DNA element regulating a green florescent protein (GFP) OUTPUT the performance metrics of a BUFFER gate can be given by the: (i) fold induction, (ii) repression strength, (iii) and two-part traceability score - i.e., induction units (IU) and repression units (RU) - relative to a reference system (see Fig.1 A). Description of Metrics for engineered XfDRRepressors are provided below. Likewise, for a NOT gate the performance can be reported by similar metrics, however: (i) fold anti- induction replaces fold induction to reflect the change in phenotype, and two-part traceability score is modified accordingly - i.e., reporting anti -induction units (AIU) - relative to the same reference system (see Fig. 15B and description of metrics for engineered XfDRrepressors ).

[0281] In addition to reporting the metrology for a given SISO, the study modeled the induction profile of a given SISO BUFFER operation via a coarse-grained binding function defined per Equation 12. Ω+(1) = σΛ+(l) + ε (Eq. 12)

[0282] In Equation 12, σ is a constant representing the maximum fluorescence relative to basal expression of the OFF-state, A+(I) is a coarse-grained Hill function that can assume avalue of 0 or 1, and & represents fluorescence in the absence of inducer; that is, the OFF-state (see Fig. 15a). Given that the transition region cannot maintain a setpoint, intermediate INPUT concentrations are excluded, analogous to the naive Hill model reported by Zong et al.

[0039] - as only the steady-state (binary) performance of interest of a given open-loop operation.

[0283] Likewise, to model the performance of a given SISO NOT gate the study used an analogous coarse-grained binding function - though for antirepression - defined per Equation 13. ΩA(1) = σɅA(1) + ε (Eq. 13)

[0284] In Equation 13, σ is a constant representing the maximum fluorescence, minus ligand, relative to basal expression of the OFF-state, AA(I) is a coarse-grained antithetical Hill- function for anti-repression where 0 INPUT corresponds to the ON-state, and 1 INPUT corresponds to the OFF-state, and a represents fluorescence in the presents of inducer; that is, the OFF-state (see Fig. 15b).

[0285] The study designed, built, and tested 80 BUFFER operations and 80 NOT operations congruent with the design space given in Fig. 15c to provide 40 systems at the PROXIMAL position (see Fig. 15d and Fig. 24), and 40 systems at the CORE position (see Fig. 15e and Fig. 25) for each putative logical operation. In addition, the study performed metrological analysis on and modeling of said transcription factors (see description of metrics for engineered engineered repressors and Supplementary Data Set 1 at Milner et al.“Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108). For the BUFFER operations the design space consisted of 5 nonsynonymous regulatory core domains, 8 alternate DNA binding operations, and 2 operator positions (see Fig. 15c). Similarly, the design space for the purported NOT operations was composed of 5 anti-RCDs (4 of which were antithetical to a given X+), with complete overlap with respect to the given alternate DNA binding functions and cognate DNA operators. Out of the 40 transcription factors tested at the PROXIMAL position,35 (-87%) resulted in objective (qualitative) BUFFER logic gating - i.e., having statistically significant differences between the ON-state (with ligand) and OFF-state (without ligand) based on a student T-test. Whereas 38 (-95%) out of the 40 transcription factors testedat the CORE position resulted in BUFFER logic (see Supplementary Data Set 1 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108). Together this resulted in 73 (out of 80 - or -91%)functional BUFFER SISO control systems. In contrast, 36 (out of 40) PROXIMAL and 40 (out of 40) CORE XAADR anti-repressor transcription factors resulted in objective NOT logic gate performance - for a total of 76 (-95%) operational NOT SISO (see Supplementary Data Set 1 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108). The study posited that any differences observed in performance between the PROXIMAL and CORE operator positions for a given (equivalent) logical operation can be attributed to the variation in binding competition between the two sites (Figs. 15d and 15e). In general, for a given logical operation, the CORE position had fewer nonoperational gates relative to the PROXIMAL position. Non- operational gates can be classified by two additional phenotypes: (i) super-repressor (Xs), or (ii) non-functional (X‘) - see Fig. 26. Notably, the majority of the nonoperational SISO gates were classified as nonfunctional (X‘). Finally, the overlap in DNA binding functions for said BUFFER and NOT operations can facilitate networked cooperation between SISO - i.e., when sets of transcription factors are directed to a single DNA operator element. The aforesaid networking capability can enable the bottom -up construction of Multiple-INPUT Single-OUTPUT (MISO) logical operations - illustrated in the following sections.

[0286] Metrics for Engineered Repressors. The Fraction of Maximum Output (F.M.O.)= [GFP / OD600] / [Max LacSTOP value**], where (i) F.M.O. Repression is the system minus ligand, and (ii) F.M.O. Induction is the system plus ligand. Repression Strength = F.M.O. normalized output minus ligand, and Fold Induction = FI, such that FI = (F.M. O.Induction) / (F.M. O. Repression). Part 1 of the traceability score is given in terms of the induced state (where IU = Induction Units) such that the IU traceability scores were calculated as IU Traceability Score = IU Reference Score == 1. Part 2 of the traceability score is given in terms of therepressed state (where RU = Repression Units) such that the RU traceability scores were calculated as RU Traceability Score =Reference Score = refers to MaxLacSTOP value = 75,000 relative fluorescence units (rfu), OD600 normalized.

[0287] Metrics for Engineered Anti-repressors. The Fraction of Maximum Output(F.M.O.) = [GFP / OD600] / [Max LacSTOP value**], where (i) F.M.O. anti-repression is thesystem minus ligand, and (ii) F.M.O. anti-induction is the system plus ligand. Fold Anti- induction = FAI, such that FAI = (F.M. O. Anti-repression) / (F.M. O. Anti-induction). Part 1 of the traceability score is given in terms of the anti-repressed state (where AIU = Anti- Induction Units) such that the AIU traceability scores were calculated: AIU Traceability Part 2 of the traceability score is given in terms of therepressed state (where RU = Repression Units) such that the RU traceability scores for an anti-repressor were calculated as: RU Traceability Score =Max LacSTOP value = 75,000 relative fluorescence units (rfu), OD600 normalized.

[0288] Design Rules for MISO AND Logical Gate Construction from BUFFER SISO Data. The construction of an AND (MISO) logical gate via transcriptional programming can use either a (i) series (SERI) (Fig. 27) or (ii) series-parallel (SE-PA) (Fig. 15f and 15g) genetic architecture. Here the study focused on the construction of 2-INPUT AND gates using the SE-PA iteration, as this particular design simplifies the accounting of independent transcription factor operator interactions - as both transcription factors are directed to the same DNA element. To identify putative sets of BUFFER logical operations that can be paired (via SE-PA DNA operators) to form objective 2-INPUT AND logical gates, the study initially used a two-step decision process informed by the SISO data alone. Namely, first, the study identified all BUFFER SISO logical operations with measurable dynamic ranges (i.e., statistical differences between the ON and OFF states), and in the second tier of the decision process, the study evaluated compatibility between two networked transcription factors. When sufficient inequality (i.e., MISO compatibility) exists between the ON-state and OFF- state of SE-PA networked BUFFER gates, we posited that an objective 2-INPUT AND logic gate can be constructed - see example in Fig. 28a. In this illustration, SISO data for I+YQR was compared to SISO data from R+YQR - where X = I or R, ADR = YQR, and + = repressor phenotype. Here, the OFF-state of I+YQR has a lower threshold relative to the ON-state of R+YQR - likewise for the complementary ON and OFF states. Accordingly, the two BUFFER operations can be regarded as being compatible with respect to forming a 2-INPUT, SE-PAdirected AND logic gate - i.e., when directed to a cognate operator at a fixed position. In contrast, given two functional BUFFER gates (e.g., E+YQR and F+YQR), potential incompatibility can arise when a pair of logical operations do not have sufficient inequality (i.e., MISO incompatibility) between the ON-state of one transcription factor i.e., F+YQR), relative to the OFF-state of the complementary (networked) transcription factor (i.e., E+YQR) - see Fig. 28b.

[0289] Unmitigated, the pairwise (2-INPUT) network space for AND gate construction is represented by 80 operations at the PROXIMAL position and 80 operations at the CORE position. However, with the initial constraints imposed by the number of functional BUFFER SISO (i.e., 35 PROXIMAL, 38 CORE), the putative network space is reduced to 62 PROXIMAL AND gates (Fig. 16a) and 72 CORE AND gates (Fig. 16b) - without factoring in putative incompatibilities. Including said incompatibilities, the putative networked space is further reduced by one - i.e., to 61 PROXIMAL AND gates and 72 CORE AND gates - resulting in a total of 133 2-INPUT logical operations that are purportedly functional (Figs. 16a and 16b).

[0290] Building, Testing, Modeling AND Logic Gates. After designing said AND logic gates we built and tested the complete set of gates with predicted function - i.e., 62 PROXIMAL AND gates (Fig. 16a) and 72 CORE AND gates (Fig. 16b). Briefly, each AND gate was experimentally tested in the same E. coli chassis cell as the SISO systems and regulated the same GFP OUTPUT. All of the AND gates were functional with qualitative (objective) performances congruent with the said logical operation (see Supplementary Data Set 2 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108).

[0291] To further validate our qualitative predictions, the study constructed several systems that were predicted to be non-operational (see Supplementary Data Set 3 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108). In general, both sets of data affirmed our qualitative prediction. Namely, on average, the nonoperational gates did not result in objective logic gating (or had poor qualitative performance - i.e., had dynamic ranges < 2) when experimentally tested.

[0292] To better interpret and predict the qualitative performance of our 2-INPUT AND logic gates, the study constructed a coarse-grained model per Equation 14.(Eq. 13)

[0293] In Equation 13, QAND is the OUTPUT expression, is the Hill state function ofrepressor X+, is the Hill state function of repressor Y+, lx is the inducer state of X+(either 0 or 1), IY is the inducer state of Y+(either 0 or 1), and a0, a1, a2, and α3 are parameters determined from the SISO gates by a set of four equations (Fig. 17a).

[0294] Qualitatively, aois the minimum OUTPUT of the gate (or overall leakiness, i.e., EX or εY), ai is the OUTPUT increase (from the baseline ao) in response to lx, a2 is the OUTPUT increase (from the baseline ao) in response to IY, and as is the OUTPUT increase from the maximum OFF state to the ON state. Further description of the AND logic model is provided below..

[0295] In general, the AND gate model predicted the quantitative performances of experimental outcomes with a high degree of accuracy - with a mean error (measured output / predicted output) of 1.256, see Figs. 18a - 18b and Fig. 29a.

[0296] Only -12% of the values had a 2-fold or greater difference relative to the predicted value - i.e., the model could accurately predict the qualitative and quantitative performance of measured values in the context of AND logic in -88% of cases. Interestingly, PROXIMAL AND logic gates (Fig. 4A) had a greater degree of spread - i.e., for a given data set per individual operation - relative to the CORE AND logic gates (Fig. 18b). The study attributed this difference to the presence of a variable ‘5-UTR (untranslated region) in the PROXIMAL systems, which can variably affect ribosome binding and thus the apparent rate of translation.

[0297] AND logic model. To better interpret and predict the qualitative performance of our 2- INPUT AND logic gates, the study constructed a coarse-grained model per Equation 13. In Equation 13, a0, a1, a2, and a3 are parameters determined as ao = min(εx, εY), a0 + a1 = εY, a0 + a2 = εx, and a0 + a1 + a2 + as = min (σx + εX, σY + εY). Qualitatively, ao is the minimum OUTPUT of the gate (or overall leakiness, i.e., sx or εY), ai is the OUTPUT increase (from the baseline ao) in response to lx, a2 is the OUTPUT increase (from the baseline ao) in response to IY, and as is the OUTPUT increase from the maximum OFF state to the ON state.

[0298] Equations for a0, a1, a2, and a3 are derived from solving Equation 13 using three distinct assumptions. First, it is assumed that when neither lx or IY is present, the TF with the lowest SISO OFF state OUTPUT controls DAND. This is represented as Ω and (0,0) = mand ao = min(εx, εY). Second, it is assumed that when either lx or IY is present, the TF in the OFF-state controls Ω and be either (i) Ω and ( 1,0) = (0); a0 + a1(1) + a2(0) += σx(0) + εx; a0 + a2 = εx.

[0299] Third, it is assumed that when both lx and IY are present, DANDis given by the TF with the lowest ON state OUTPUT includes: DAND(1, 1) = (ctx(l) + εx, σY(1) + εY); a0 +a1 + a2 + a3 = min(σx + εx, σY + εY).

[0300] As a summary of the outcomes, the model could accurately predict the qualitative performance of all measured values in the context of the AND model in -88% of cases - i.e., only -12% of the values had a 2-fold or greater difference relative to the predicted value. An error in AND condition 1 (-,-) correlates to the OFF states of both TFx and TFY, an error in AND condition 2 (-,+) correlates to the OFF state of TFY, an error in AND condition 3 (+,-) correlates to the OFF state of TFx, and error in AND condition 4 (+,+) correlates to the ON states of both TFX and TFY, also see Fig. 17a.

[0301] Design Rules for MISO NOR Logical Gate Construction from NOT SISO Data.Similar to the workflow for identifying functional AND gates, a similar process can be used to identify putative 2-INPUT NOR logical gates - paired via SE-PA operators. Namely, the initial selection (design) criteria employed the identification of said NOT SISO operations with statistically significant differences between the ON-state (without ligand) and the OFF- state (with ligand) - i.e., adequate dynamic ranges for a set of network capable antirepressors.

[0302] The second hierarchical design criteria require sufficient inequality between complementary ON and OFF states (Fig. 30a). In other words - with the design goal of forming a 2-INPUT NOR logical operation - incompatibility between two NOT gates occurs when said SISO does not have sufficient distinction between the ON-state of one operation (PAYQR) relative to the OFF-state of the complementary operation (IA(9)YQR) - see Fig. 30b. The unrestricted, 2-INPUT network space for NOR gate construction is represented by 80 operations at the PROXIMAL position and 80 operations at the CORE position.

[0303] However, when accounting for the non-functional SISO logic gates at the PROXIMAL position, the network space is reduced to 64 putative NOR gates (Fig. 19a). Moreover, including the 4 incompatible sets the PROXIMAL network space is reduced to 60 putative NOR gates. In contrast, the CORE position did not contain any non-functional NOT gates. However, 9 putative incompatible NOR sets were predicted at the CORE position - resulting in 71 putative NOR gates (Fig. 19b).

[0304] The study built and experimentally tested all putative NOR gates - 60 putative PROXIMAL NOR gates, and 71 putative CORE NOR gates. In addition, the study constructed a model to better interpret and predict the qualitative performance of our 2- INPUT NOR logic gates given the corresponding SISO data per Equation 14.

[0305] In Equation 14, ΩNOR is the OUTPUT expression, Ay is the Hill state function of anti- repressor XA, is the Hill state function of anti-repressor YA, lx is the inducer state of XA(either 0 or 1), IY is the inducer state of YA(either 0 or 1), and a0, a1, a2, and a3 are parameters determined by the set of four equations described previously - also see Fig. 17b. Further description of the NOR logic model is provided below.

[0306] Qualitatively, all of the predicted NOR gates were functional - i.e., validated by experiment (see Fig. 20a - 20b). Quantitatively, the model accurately predicted the experimental values in -89% of cases, with a mean error of 1.28 (Fig. 29b). Moreover, select non-operational data affirmed our expectations - i.e., on average, objective (qualitative) NOR gating was not observed or had poor performance (see Supplementary Data Set 3 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108). Congruent with the observation made for differences in performance at the PROXIMAL versus CORE positions for the AND gates, the tested NOR gates had similar differences in data spread between the two operator positions.Namely, PROXIMAL NOR gates (Fig. 20a) had a greater degree of spread in the standard deviation per data point relative to CORE NOR gates (Fig. 20b) - which we again attributed to a variable ‘5-UTR in the PROXIMAL operations.

[0307] NOR logic model. For 2-INPUT NOR logic gates, the study modified the model shown in Equation 13 to include anti -repressor state functions pertaining to NOT SISO logic. The model for NOR is provided Equation 14. In Equation 14, a0, a1, a2, and a3 are parameters determined as ao = min(εx, εY), a0 + a1 = εY, a0 + a2 = ex, and a0 + a1 + a2 + as = min (σx + εx, σY + εY). Equations for a0, a1, a2, and a3 are derived from solving Equation 14 using similar assumptions described in the AND model but with antithetical input conditions (due to the antithetical phenotype of the anti-repressors from repressors). First, it is assumed that when neither lx or IY is present, the TF with the lowest SISO OFF state OUTPUT controls (ΩNOR. This is represented= min (σx(0) + εx, σY(0) + εY); and a0 = min(sx, εY). Second, it is assumed that when either lx or IY is present, the TF in the OFF-state controls ( ΩNOR. includes (i) (ΩNOR(1,0) =; a0 + a1(0) + a2(l) + a3(0)(l) = σx(0) + εx;a0 + a2 ex. or (ii) (ΩNOR(0,1) =(1); a0 + a1(l) + a2(0) + a3(l)(0) = σY(0) + εY;a0 + a1 = εX. Third, it is assumed that whenneither lx and IY are present, the TF with the lowest ON state OUTPUT includes (ctx(l) + EX,σY( 1 ) + εY); a0 + a1 + a2 + a3 = min(σX + εX, σY + εY).

[0308] Building, Testing, Modeling Nonimplication Logic Gates. Based on the study, any network of two transcription factors with divergent phenotypes (i.e., repressor and anti- repressor) can be performed with a shared (SEPA) DNA operator. This class of simple mixed networks objectively can provide, e.g., an A NIMPLY B logical operation, see Figs. 31a - 31d.

[0309] Likewise, using the complementary set of transcription factors (i.e., R+YQR, and IAYQR) the complementary logical operation B NIMPLY A can be generated, see Figs. 31a - 31d . In the illustrations, the study qualitatively predicted that the repressor I+YQR can be paired with the anti-repressor RAYQR, and this operation will only produce an OUTPUT when the INPUT signal that corresponds to the repressor is present, see Figs. 21a - 211 and Figs. 32a - 321..

[0310] As with previous 2-INPUT operations, the study can model nonimplication logic gates to better interpret quantitative performances. Namely, for 2-INPUT A NIMPLY B logic gates, the study modified the model shown in Equation 13 to include one repressor and one anti-repressor state function pertaining to both BUFFER and NOT SISO logic. The model for A NIMPLY B is provided in Equation 15 A.(Eq. 15 A)

[0311] In Equation 15 A, Ω A NIMPLY B is the OUTPUT expression, is the Hill state functionof repressor X+, is the Hill state function of anti-repressor YA, lx is the inducer state of X+(either 0 or 1), IY is the inducer state of YA(either 0 or 1), and a0, a1, a2, and a3 are parameters determined by the set of four equations described previously - also see Figs. 22a - 22b. Further description of the A NIMPLY B logic model is provided below..

[0312] The model for 2-INPUT B NIMPLY A logic gates follows that described above for A NIMPLY B gates, with the modification that TFs X and Y phenotypes are switched so that this system contains an anti-repressor XAand repressor Y+. The model is provided as Equation 15B.(Eq. 15B)

[0313] In Equation 15B, Ω B NIMPLY A is the OUTPUT expression, s the Hill state functionof anti-repressor XA, is the Hill state function of repressor Y+, lx is the inducer state ofXA (either 0 or 1), IY is the inducer state of Y+(either 0 or 1), and ao, ai, a2, and as are parameters determined by the set of four equations described - also see Fig. 22b. Further description of the B NIMPLY A logic model is provided below.

[0314] Given that total combinatorial space for said nonimplication logic gates is represented by 160 operations per operator position (i.e., a total of 320 logical operations), we selected 12 exemplars (i.e., 24 when considering CORE and PROXIMAL operator positions) to illustrate gate construction and to test our model’s accuracy given the corresponding SISO data (see Fig. 21a - 211, Figs. 32a - 321, and Supplementary Data Set 2 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108). Qualitatively, all of the tested nonimplication logic gates performed as expected.

[0315] Moreover, the model accurately predicted the qualitative performance in all cases and quantitatively predicted the performance of said nonimplication gates in -86% of the tested cases.

[0316] A NIMPLY B logic model. For 2-INPUT B NIMPLY A logic gates, the model shown in Equation 13 can be modified to include one repressor and one anti -repressor state function pertaining to both BUFFER and NOT SISO logic. The model for X NIMPLY Y is provided in Equation 15 A. In Equation 15 A, a0, a1, a2, and a3 are parameters determined as ao = min(sx, εY); a0 + a1 = εY;a0 + a2 = εx;and a0 + a1 + a2 + as = min (σx + εx, σY + εY).Equations for a0, a1, a2, and a3 are derived from solving Equation 15A using similar assumptions described in the AND and NOR model but with conditions reflecting the phenotype of each TF. First, it is assumed that when only IY is present, the TF with the lowest SISO OFF state OUTPUT controls Ω A NIMPLY B can be represented as Ω A NIMPLY B ( 0,1) =ao = min(sx, εY). Second, it is assumed that when neither lx or IY is present, the repressor X+controlsσx(0) + εX; a0 + a2 = εX. Similarly, when both lx or IY are present, the anti-repressor YAcontrols+ εY; a0 + a1 = εY. Third, it is assumed that when only lx is present, NIMPLYYcan be given by the TF with the lowest ON state OUTPUT where Ω A NIMPLY B(1,0) =(ox(l) + ex, σY(1) + εY); a0+ a1 + a2 + a3 = min(σX + εX, σY + εY).

[0317] B NIMPLY A logic model. The model for 2-INPUT B NIMPLY A logic gates follows that described above for A NIMPLY B gates, with the modification that TFs X and Y phenotypes are switched, so that this system contains an anti-repressor XAand repressor Y+. The model is provided in Equation 15B. In Equation 15B, a0, a1, a2, and a3 are parameters determined as ao = min(εx, εY), a0 + a1 = εY, a0 + a2 = ex, and a0 + a1 + a2 + as = min (σx + εx, σY + εY). Equations for a0, a1, a2, and a3 are derived from solving Equation 15B using similar assumptions described in the AND model but with input conditions reflecting the phenotypes of each TF. First, it is assumed that when only lx is present, the TF with the lowest SISO OFF state OUTPUT controlsNIMPLY Acanbe represented as Ω B NIMPLY A(1,0) =ex, σY(0) + εY); and ao = min(εx, εY). Second, it is assumed that when neither lx or IY is present, the repressor Y+controlsσY(0) + εY; a0 + a1 = εY. Simarly, when both lx or IY are present, the anti-repressor XAcontrolsεx; a0+ a2 = ex. Third, it is assumed that when only IY is present, Ω B NIMPLY A can be given by the TF with the lowest ON state OUTPUT where Ω B NIMPLY A(0,1) = ex, σY(1) + BY); a0 +ai + a2 + as = min(σX + εX, σY + εY).

[0318] Discussion. As circuit complexity increases in synthetic biology, there is a growing need to predict the performance of a desired complex system (i.e., qualitatively and quantitatively) prior to construction. Ideally, this would begin with modeled interaction data - which would involve comprehensive predictions of protein-ligand interactions, protein-DNA interactions, and allosteric communication in terms of binding energetics. While such predictive capabilities are not practical, the extrapolation of simple SISO data to predict the performance of MISO logical operations has shown great promise

[0039] ,

[0319] The instant study provides for the predicting of circuit performance for simple (single promoter) transcriptional programs that can potentially scale to more complex operations that involve feeding forward information.

[0320] While the ability to predict 2-INPUT circuit performance from SISO data was remarkably accurate, in some cases the system intimated that properties of the circuit wereresponsible for increased variability in the performance of a given state of a circuit. Namely, the instant study posited that the variable 5’-UTR in PROXIMAL circuits contributed to decreased accuracy in the predictions - as the CORE circuits had better correlations with said predictions (Figs. 18a - 18b and Figs. 20a - 20b). To test this assertion, the instant study inserted a genetic insulator upstream of the putative UTR in a subset of PROXIMAL AND circuits and NOR circuits - i.e., gates with the largest prediction error for each 2-INPUT combinatorial pair - followed by a re-test of circuit performance (see Supplementary Data Set 4 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108).

[0321] With the addition of the genetic insulator, the study observed a ~3-fold and ~6-fold improvement in the accuracy of the prediction of experimental results relative to the model for AND gates and NOR gates, respectively. An important feature of transcriptional programming is the ability to use network repressors and anti-repressors to build multiple INPUT operations that are compressed

[0030] , The instant study accomplished this with the SE-PA architecture to form simple 2-node networks of transcription factors binned via the alternate DNA binding function - e.g., YQR | O1, HQN | Otg, or GKR | Ogac. This SE-PA network form resulted in 7 orthogonal (binned) DNA binding networks, with inter-bin communication facilitated via the INPUT signals. The description is provided in the supplemental discussion below. For 7 nonsynonymous ADR and 5 RCDs each with the capacity to interact with one of 5 non-synonymous INPUTS, the result was a putative network space of 70 operations (i.e., restricted to said AND gates and NOR gates), with 103signal coupled operations. When considering mixed unit operations (i.e., said A NIMPLY B gates, and B NIMPLY A gates) the network space is represented by 350 putative operations, with signal coupling three times larger than AND gate (or NOR gate) coupled operations.

[0322] In the context of network development, the DNA binding network in transcriptional programming can be expanded via the SERI architecture. Namely, in a given SERI genetic architecture, two nonsynonymous DNA operators can be paired in tandem - i.e., one located at the CORE position and the other at the PROXIMAL position (Fig. 27e). When extrapolated based on the engineered transcription factors and cognate DNA operators used in this study, the putative SERI networked DNA space results in 103operations, with a signal coupling on the order of 105. The description is provided in the supplemental discussion below.

[0323] To determine if SERI circuits are amenable to modeling (i.e., MISO predictions from SISO data) the instant study leveraged the workflows that we established for SE-PA circuits.Given the enormous combinatorial space for SERI circuits the study opted to demonstrate predictive capacity using a small sample set - i.e., 6 AND gates (figs. 23a - 231) + 6 NOR gates (Figs. 23g - 231).

[0324] Congruent with the workflows in the previous study, first the isntant study collected SISO data for each transcription factor - however in this case operating in the context of a given SERI operator-promoter (opposed to SE-PA). The rationale for re-collecting SISO data is evidenced in the differences in performances between SE-PA and SERI 1 -INPUT operations for equivalent transcription factors and cognate DNA interactions. In turn, we built, tested, and modeled the corresponding SERI AND gates (see Supplementary Data Set 5 at Milner et al. “Performance Prediction of Fundamental Transcriptional Programs,” ACS Synthetic Biology 12, no. 4 (2023): 1094-1108). In all cases, the experimental data and model predictions were in good agreement. Likewise, the experimental data and corresponding model of NOR gates were in good agreement; however, many SERI NOR circuits had divergent performance relative to synonymous SE-PA circuits. For example, RA(3)Ksr and IA(6)HTK, in principle, should form an objective NOR gate based on the SE-PA SISO data. However, in the context of the SERI architecture, said operations are incompatible and, as predicted, result in a non-functional operation (Fig. 23k).

[0325] Moreover, the inclusion of a genetic insulator does not improve circuit fidelity, indicating that properties and functions that precede translation (e.g., changes in transcription factor DNA interactions and possibly changes in promoter strength) impact the performance of the SERI circuit, see Figs. 33a - 331.

[0326] Supplemental Discussion. The number of possible SE-PA logical BUFFER or NOT SISO operations can be calculated from simple combinations of the selection of an ADR (1 of 5 either repressor [AND] or antirepressor [NOT]) and DBD (1 of 8). This combination can be placed in either the PROXIMAL or CORE position leading to:= 40 SISO operations per position or 80 in total for either the BUFFER or NOT. It should be noted that two of the ADRs O1and OSYMrecognize the same DNA binding domain and so are synonymous, which leads to an effective design space of 70 SISO operations. For the SE-PA operation of AND, two ADR (2 of 5) can be selected but only one DBD to which each ADR will be coupled. Again, this SE-PA AND operator can be placed in either the PROXIMAL or CORE position:x 2 = 160 = possible SE-PA AND gates, or 140 non-synonymous combinations. A similar argument holds for the NOR gates where we again select the samenumbers of components but from the set of DBDs with anti-repressor behavior, leading to 160 possible NOR gates.

[0327] For the SE-PA operation of NIMPLY, 1 ADR can be selected from the set of repressors and 1 ADR from the set of anti-repressors, and one DBD. In general, this would lead to possible nonsynonymous NIMPLY designs per position for atotal of 350 possible designs. However, four of the five signals are the same for both the repressor and anti-repressor (Cellobiose and Adenine being the two different ones, respectively). This means that except for these two, when the signal is selected for the first operator, the second one can only be selected from the remaining four. This leads to (5+4x4)x8 or (5+4x4)x7 NIMPLY designs for each position, 168 or 147, respectively.

[0328] For the SERI architecture for the two-input one output AND gate two non- synonymous ADRs can be selected and placed uniquely in the PROXIMAL and CORE positions. Two DBD’s can be selected to be coupled to these ADRs: For theNIMPLY logical operation the combinations are similar with one repressor and one anti- repressor being selected from each set and each one coupled to anBoth of these designs for SERI architecture can be doubled in number because the ordering of the two ADR’s in the PROXIMAL and CORE positions can be switched. However, for this specific system, the same argument about the signal overlap applies. Hence the number of unique combinations of the DBD’s is reduced to 21 leading to logicaloperations that can be doubled in number by changing the order of the ADR’s in the PROXIMAL and CORE positions. Tables 2 and 3 provides a summary.Table 2Table 3

[0329] It will be apparent to those skilled in the art that various modifications and variations can be made in the present disclosure without departing from the scope or spirit of the invention. Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the methods disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the invention being indicated by the following claims.

[0330] Example Analysis System

[0331] Computer-executable instructions, such as program modules, being executed by a computer may be used. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. In its most basic configuration, the controller includes at least one processing unit and memory. Depending on the exact configuration and type of computing device, memory may be volatile (such as random-access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two. The controller of Fig. 12 may have additional features / functionality. For example, the computing device may include additional storage (removable and / or non-removable), including, but not limited to, magnetic or optical disks or tape. Such additional storage may include removable storage and / or non-removable storage.

[0332] It should be understood that the various techniques described herein may be implemented in connection with hardware components or software components or, where appropriate, with a combination of both. Illustrative types of hardware components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc. The methods and apparatus of the presently disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, or any other machine-readable storage medium where, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the presently disclosed subject matter.

[0333] The following patents, applications and publications as listed below and throughout this document, are hereby incorporated by reference in their entirety herein.Reference List #1[ 1´] Gardner, T.S., Cantor, C.R. & Collins, J. J. Construction of a genetic toggle switch in Escherichia coli. Nature 403, 339--2 (2000).[2'] Elowitz, M.B. & Leibler, S. A synthetic oscillatory network of transcriptional regulators. Nature 403, 335-8 (2000).[3'] Kosuri, S. et al. Composability of regulatory sequences controlling transcription and translation in Escherichia coli. Proc Natl Acad Sci USA 110, 14024-9 (2013).[4’] Lou, C., Stanton, B., Chen, Y.J., Munsky, B. & Voigt, C.A. Ribozyme-based insulator parts buffer synthetic circuits from genetic context. Nat Biotechnol 30, 1137-42 (2012).[5’] SimSek, E., Yao, Y., Lee, D. & You, L. Toward predictive engineering of gene circuits. Trends Biotechnol 41, 760-768 (2023).[6’] Moon, T.S., Lou, C., Tamsir, A., Stanton, B.C. & Voigt, C.A. Genetic programs constructed from layered logic gates in single cells. Nature 491, 249-53 (2012).[7’] Nielsen, A.A. et al. Genetic circuit design automation. Science 352, aac7341 (2016).[8’] Daniel, R., Rubens, J.R., Sarpeshkar, R. & Lu, T.K. Synthetic analog computation in living cells. Nature 497, 619-23 (2013).[9’] Cox, R.S., 3rd, Surette, M.G. & Elowitz, M.B. Programming gene expression with combinatorial promoters. Mol Sy st Biol 3, 145 (2007).[10’] Sexton, J.T. & Tabor, J.J. Multiplexing cell-cell communication. Mol SystBiol 16, e9618 (2020).[I L] Rondon, R.E., Groseclose, T.M., Short, A.E. & Wilson, C.J. Transcriptional programming using engineered systems of transcription factors and genetic architectures. Nat Commun 10, 4784 (2019).[12’] Dong, C., Fontana, J., Patel, A., Carothers, J.M. & Zalatan, J.G. Synthetic CRISPR-Cas gene activators for transcriptional reprogramming in bacteria. Nat Commun 9, 2489 (2018).[13’] Qi, L.S. et al. Repurposing CRISPR as an RNA-guided platform for sequence- specific control of gene expression. Cell 152, 1173-83 (2013).[14’] Meyer, A. J., Segall-Shapiro, T.H., Glassey, E., Zhang, J. & Voigt, C.A. Escherichia coli "Marionette" strains with 12 highly optimized small-molecule sensors. Nat Chem Biol 15, 196-204 (2019).[15’] Stanton, B.C. et al. Genomic mining of prokaryotic repressors for orthogonal logic gates. Nat Chem Biol 10, 99-105 (2014).[16’] Huang, B.D., Groseclose, T.M. & Wilson, C.J. Transcriptional programming in a Bacteroides consortium. Nat Commun 13, 3901 (2022).[17’] Groseclose, T.M., Hersey, A.N., Huang, B.D., Realff, M.J. & Wilson, C.J. Biological signal processing filters via engineering allosteric transcription factors. Proc Natl Acad Set USA 118, e2111450118 (2021).[18’] Groseclose, T.M., Rondon, R.E., Herde, Z.D., Aldrete, C.A. & Wilson, C.J. Engineered systems of inducible anti -repressors for the next generation of biological programming. Nat Commun 11, 4440 (2020).[19’] Chan, C.T., Lee, J.W., Cameron, D.E., Bashor, C.J. & Collins, J.J. 'Deadmar and 'Passcode' microbial kill switches for bacterial containment. Nat Chem Biol 12, 82-6 (2016).[20’] Shis, D.L., Hussain, F., Meinhardt, S., Swint-Kruse, L. & Bennett, M.R. Modular, multi-input transcriptional logic gating with orthogonal LacI / GalR. family chimeras. ACS Synth Biol 3, 645-51 (2014).[21’] Richards, D.H., Meyer, S. & Wilson, C.J. Fourteen Ways to Reroute Cooperative Communication in the Lactose Repressor: Engineering Regulatory Proteins with Alternate Repressive Functions. ACS Synth Biol 6, 6-12 (2017).[22’] Herde, Z.D. et al. Engineering allosteric communication. Curr Opin Struct Biol 63, 115-122 (2020).[23’] Fu, Y. et al. Structural and functional analyses of the cellulase transcription regulator CelR. FEBS Lett 592, 2776-2785 (2018).[24’] Kelly, J.R. et al. Measuring the activity of BioBrick promoters using an in vivo reference standard. J Biol Eng 3, 4 (2009).[25’] Espah Borujeni, A., Zhang, J., Doosthosseini, H., Nielsen, A.A.K. & Voigt, C.A. Genetic circuit characterization by inferring RNA polymerase movement and ribosome usage. Nat Commun 11, 5001 (2020).[26’] Csibra, E. & Stan, G.B. Absolute protein quantification using fluorescence measurements with FPCountR. Nat Commun 13, 6600 (2022).[27’] Hall, M.P. et al. Engineered luciferase reporter from a deep sea shrimp utilizing a novel imidazopyrazinone substrate. ACS Chem Biol 7, 1848-57 (2012).[28’] Misawa, N. et al. Elucidation of the Erwinia uredovora carotenoid biosynthetic pathway by functional analysis of gene products expressed in Escherichia coli. J Bacteriol 172, 6704-12 (1990).[29’] Cunningham, F.X., Jr., Sun, Z., Chamovitz, D., Hirschberg, J. & Gantt, E. Molecular structure and enzymatic function of lycopene cyclase from the cyanobacterium Synechococcus sp strain PCC7942. Plant Cell 6, 1107-21 (1994).[30’] Kim, S.W. & Keasling, J.D. Metabolic engineering of the nonmevalonate isopentenyl diphosphate synthesis pathway in Escherichia coli enhances lycopene production. Biotechnol Bioeng 72, 408-15 (2001).[31’] Sun, T. et al. Production of lycopene by metabolically-engineered Escherichia coli. Biotechnol Lett 36, 1515-22 (2014).[32’] Kajiwara, S., Fraser, P.D., Kondo, K. & Misawa, N. Expression of an exogenous isopentenyl diphosphate isomerase gene enhances isoprenoid biosynthesis in Escherichia coli. Biochem J 324 ( Pt 2), 421-6 (1997).[33’] Matthews, P.D. & Wurtzel, E.T. Metabolic engineering of carotenoid accumulation in Escherichia coli by modulation of the isoprenoid precursor pool with expression of deoxyxylulose phosphate synthase. Appl Microbiol Biotechnol 53, 396-400 (2000).[34’] McNerney, M.P. & Styczynski, M.P. Precise control of lycopene production to enable a fast-responding, minimal-equipment biosensor. Metab Eng 43, 46-53 (2017).[35’] Miguez, A.M., McNerney, M.P. & Styczynski, M.P. Metabolomics Analysis of the Toxic Effects of the Production of Lycopene and Its Precursors. Front Microbiol 9, 760 (2018).[36’] Yoon, K.W., Doo, E.H., Kim, S.W. & Park, J.B. In situ recovery of lycopene during biosynthesis with recombinant Escherichia coli. J Biotechnol 135, 291-4 (2008).[37’] Albermann, C. High versus low level expression of the lycopene biosynthesis genes from Pantoea ananatis in Escherichia coli. Biotechnol Lett 33, 313-9 (2011).[38’] Levin-Karp, A. et al. Quantifying translational coupling in E. coli synthetic operons using RBS modulation and fluorescent reporters. ACS Synth Biol 2, 327-36 (2013).[39’] Higashikuni, Y., Chen, W.C. & Lu, T.K. Advancing therapeutic applications of synthetic gene circuits. Curr Opin Biotechnol 47, 133-141 (2017).[40’] Hong, M., Clubb, J.D. & Chen, Y.Y. Engineering CAR-T Cells for Next- Generation Cancer Therapy. Cancer Cell 38, 473-488 (2020).[41’] Cubillos-Ruiz, A. et al. Engineering living therapeutics with synthetic biology. Nat Rev Drug Discov (2021).[42’] Gonzalez, L.M., Mukhitov, N. & Voigt, C.A. Resilient living materials built by printing bacterial spores. Nat Chem Biol 16, 126-133 (2020).[43’] McCarty, N.S. & Ledesma-Amaro, R. Synthetic Biology Tools to Engineer Microbial Communities for Biotechnology. Trends Biotechnol 37, 181-197 (2019).[44’] Khalil, A.S. & Collins, J. J. Synthetic biology: applications come of age. Nat Rev Genet 11, 367-79 (2010).[45’] Voigt, C.A. Synthetic biology 2020-2030: six commercially-available products that are changing our world. Nat Commun 11, 6379 (2020).[46’] Zrimec, J. et al. Controlling gene expression with deep generative design of regulatory DNA. Nat Commun 13, 5099 (2022).[47’] Zhang, P. et al. Deep flanking sequence engineering for efficient promoter design using DeepSEED. Nat Commun 14, 6309 (2023).[48’] Van Brempt, M. et al. Predictive design of sigma factor-specific promoters. Nat Commun 11, 5822 (2020).[49’] LaFleur, T.L., Hossain, A. & Salis, H.M. Automated model-predictive design of synthetic promoters to control transcriptional profiles in bacteria. Nat Commun 13, 5159 (2022).[50’] Salis, H.M., Mirsky, E.A. & Voigt, C.A. Automated design of synthetic ribosome binding sites to control protein expression. Nat Biotechnol 27, 946-50 (2009).[51’] Seo, S.W. et al. Predictive design of mRNA translation initiation region to control prokaryotic translation efficiency. Metab Eng 15, 67-74 (2013).[52’] Krueger, M., Scholz, O., Wisshak, S. & Hillen, W. Engineered Tet repressors with recognition specificity for the tetO-4C5G operator variant. Gene 404, 93-100 (2007).[53’] Dimas, R.P. et al. Engineering DNA recognition and allosteric response properties of TetR family proteins by using a module-swapping strategy. Nucleic Acids Res 47, 8913-8925 (2019).[54’] Kamionka, A., Bogdanska-Urbaniak, J., Scholz, O. & Hillen, W. Two mutations in the tetracycline repressor change the inducer anhydrotetracycline to a corepressor. Nucleic Acids Res 32, 842-7 (2004).[55’] Engler, C., Kandzia, R. & Marillonnet, S. A one pot, one step, precision cloning method with high throughput capability. PLoS One 3, e3647 (2008).[56’] Gibson, D.G. et al. Enzymatic assembly of DNA molecules up to several hundred kilobases. Nat Methods 6, 343-5 (2009).Reference List #2[1] Yokobayashi, Y., Weiss, R. and Arnold, F.H. (2002) Directed evolution of a geneticcircuit. Proc Natl Acad Sci USA, 99, 16587-16591.[2] Ellis, T., Wang, X. and Collins, J.J. (2009) Diversity-based, model-guided construction of synthetic gene networks with predicted functions. Nat Biotechnol, 27, 465- 471.[3] Siuti, P., Yazbek, J. and Lu, T.K. (2013) Synthetic circuits integrating logic and memory in living cells. Nat Biotechnol, 31, 448-+.[4] Stanton, B.C., Nielsen, A.A.K., Tamsir, A., Clancy, K., Peterson, T. and Voigt, C.A. (2014) Genomic mining of prokaryotic repressors for orthogonal logic gates. Nat Chem Biol, 10, 99-105.[5] Nielsen, A.A., Der, B.S., Shin, J., Vaidyanathan, P., Paralanov, V., Strychalski, E.A., Ross, D., Densmore, D. and Voigt, C.A. (2016) Genetic circuit design automation. Science, 352, aac7341.[6] Guiziou, S., Sauveplane, V., Chang, H.J., Clerte, C., Declerck, N., Jules, M. and Bonnet, J. (2016) A part toolbox to tune genetic expression in Bacillus subtilis.Nucleic Acids Res, 44, 7495-7508.[7] Meyer, A. J., Segall-Shapiro, T.H., Glassey, E., Zhang, J. and Voigt, C.A. (2019) Escherichia coli "Marionette" strains with 12 highly optimized smallmolecule sensors. Nat Chem Biol, 15, 196-+.[8] Chen, Y., Zhang, S.Y., Young, E.M., Jones, T.S., Densmore, D. and Voigt, C.A. (2020) Genetic circuit design automation for yeast. Nat Microbiol, 5, 1349-+.[9] Taketani, M., Zhang, J., Zhang, S., Triassi, A.J., Huang, Y.J., Griffith, L.G. and Voigt,C.A. (2020) Genetic circuit design automation for the gut resident species Bacteroides thetaiotaomicron. Nat Biotechnol, 38, 962-969.

[0010] Brandsen, B.M., Mattheisen, J.M., Noel, T. and Fields, S. (2018) A Biosensor Strategy for E-coli Based on Ligand-Dependent Stabilization. Acs Synth Biol, 7, 1990-1999.

[0011] Koch, M., Pandi, A., Borkowski, O., Batista, A.C. and Faulon, J.L. (2019) Custom-made transcriptional biosensors for metabolic engineering. Curr Opin Biotech, 59, 78-84.

[0012] Elowitz, M.B. and Leibler, S. (2000) A synthetic oscillatory network of transcriptional regulators. Nature, 403, 335-338.

[0013] Danino, T., Mondragon-Palomino, O., Tsimring, L. and Hasty, J. (2010) A synchronized quorum of genetic clocks. Nature, 463, 326-330.

[0014] Hussain, F., Gupta, C., Hirning, A.J., Ott, W., Matthews, K.S., Josie, K. and Bennett,M.R. (2014) Engineered temperature compensation in a synthetic genetic clock. P Natl Acad Sci USA, 111, 972-977.

[0015] Pattanayak, G.K., Lambert, G., Bernat, K. and Rust, M.J. (2015) Controlling the Cyanobacterial Clock by Synthetically Rewiring Metabolism. Cell Rep, 13, 2362-2367.

[0016] Atkinson, M.R., Savageau, M.A., Myers, J.T. and Ninfa, A. J. (2003) Development of genetic circuitry exhibiting toggle switch or oscillatory behavior in Escherichia coli. Cell, 113, 597-607.

[0017] Gardner, T.S., Cantor, C.R. and Collins, J.J. (2000) Construction of a genetic toggle switch in Escherichia coli. Nature, 403, 339-342.

[0018] Stricker, J., Cookson, S., Bennett, M.R., Mather, W.H., Tsimring, L.S. and Hasty, J. (2008) A fast, robust and tunable synthetic gene oscillator. Nature, 456, 516-U539.

[0019] Tigges, M., Marquez-Lago, T.T., Stelling, J. and Fussenegger, M. (2009) A tunable synthetic mammalian oscillator. Nature, 451, 309-312.

[0020] Chan, C.T.Y., Lee, J.W., Cameron, D.E., Bashor, C.J. and Collins, J.J. (2016) 'Deadman' and 'Passcode' microbial kill switches for bacterial containment. Nat Chem Biol, 12, 82-+.

[0021] Del Vecchio, D., Abdallah, H., Qian, Y.L. and Collins, J.J. (2017) A Blueprint for aSynthetic Genetic Feedback Controller to Reprogram Cell Fate. Cell Sy st, 4, 109-+.

[0022] Milias-Argeitis, A., Rullan, M., Aoki, S.K., Buchmann, P. and Khammash, M. (2016) Automated optogenetic feedback control for precise and robust regulation of gene expression and cell growth. NatCommun, 7.

[0023] Auslander, D., Auslander, S., Charpin-El Hamri, G., Sedlmayer, F., Muller, M., Frey, O., Hierlemann, A., Stelling, J. and Fussenegger, M. (2014) A Synthetic Multifunctional Mammalian pH Sensor and CO2 Transgene-Control Device. Mol Cell, 55, 397-408.

[0024] Mohamad H. Abedi, M.S.Y., David R. Mittelstein, Avinoam Bar-Zion, Margaret Swift, Audrey Lee-Gosselin, Mikhail G. Shapiro. (2021) Acoustic Remote Controlof B acteri al Immunotherapy .

[0025] Smole, A., Lainscek, D., Bezeljak, U., Horvat, S. and Jerala, R. (2017) A Synthetic Mammalian Therapeutic Gene Circuit for Sensing and Suppressing Inflammation. Mol Ther, 25, 102-119.

[0026] Ye, H.F., Charpin-El Hamri, G., Zwicky, K., Christen, M., Folcher, M. and Fussenegger, M. (2013) Pharmaceutically controlled designer circuit for the treatment of the metabolic syndrome. P Natl Acad Sci USA, 110, 141-146.

[0027] Ye, H.F., Daoud-El Baba, M., Peng, R.W. and Fussenegger, M. (2011) A SyntheticOptogenetic Transcription Device Enhances Blood-Glucose Homeostasis in Mice. Science, 332, 1565-1568.

[0028] Huang, B.D., Groseclose, T.M. and Wilson, C.J. (2022) Transcriptional programming in a Bacteroides consortium. Nat Commun, 13, 3901.

[0029] Rondon, R.E., Groseclose, T.M., Short, A.E. and Wilson, C.J. (2019) Transcriptional programming using engineered systems of transcription factors and genetic architectures. Nat Commun, 10.

[0030] Groseclose, T.M., Rondon, R.E., Herde, Z.D., Aldrete, C.A. and Wilson, C.J. (2020) Engineered systems of inducible anti-repressors for the next generation of biological programming. Nat Commun, 11.

[0031] Weickert, M.J. and Adhya, S. (1992) A family of bacterial regulators homologous to Gal and Lac repressors. J Biol Chem, 267, 15869-15874.

[0032] Swint-Kruse, L. and Matthews, K.S. (2009) Allostery in the LacI / GalR family: variations on a theme. Curr Opin Microbiol, 12, 129-137.

[0033] Tungtur, S., Egan, S.M. and Swint-Kruse, L. (2007) Functional consequences of exchanging domains between LacI and PurR are mediated by the intervening linker sequence. Proteins, 68, 375-388.

[0034] Shis, D.L., Hussain, F., Meinhardt, S., Swint-Kruse, L. and Bennett, M.R. (2014) Modular, Multi-Input Transcriptional Logic Gating with Orthogonal Lacl / GaIR Family Chimeras. Acs Synth Biol, 3, 645-651.

[0035] Rondon, R.E. and Wilson, C.J. (2019) Engineering a New Class of Anti-Laci Transcription Factors with Alternate DNA Recognition. Acs Synth Biol, 8, 307-317.

[0036] Rondon, R. and Wilson, C.J. (2021) Engineering Alternate Ligand Recognitionin the PurR Topology: A System of Novel Caffeine Biosensing Transcriptional Antirepressors. Acs Synth Biol, 10, 552-565.

[0037] Groseclose, T.M., Hersey, A.N., Huang, B.D., Realff, M.J. and Wilson, C.J. (2021), Biological signal processing filters via engineering allosteric transcription factors. Proc Natl Acad Sci USA, 118.

[0038] Richards, D.H., Meyer, S. and Wilson, C.J. (2017) Fourteen Ways to Reroute Cooperative Communication in the Lactose Repressor: EngineeringRegulatory Proteins with Alternate Repressive Functions. Acs Synth Biol, 6, 6-12.

[0039] Zong, D.M., Cinar, S., Shis, D.L., Josie, K., Ott, W. and Bennett, M.R. (2018) Predicting Transcriptional Output of Synthetic Multi-input Promoters. Acs Synth Biol, 7, 1834-1843.Bold nucleotides / amino acids represent super-repressor mutation.Underlined nucleotides / amino acids represent anti-repressor mutation.Italicized nucleotides / amino acids represent appended sequences.

Claims

What is claimed is:

1. A computerized method to predict transcriptional programming of gene expression, the method comprising: providing a gene model having network repressors and network anti-repressors, wherein the network repressors and network anti-repressors are designated for binding to a plurality of candidate DNA binding function; applying one or more of the plurality of candidate DNA binding function as a first 2- node single-in single-out (SISO) networks of transcription factors binned via the alternate DNA binding function; and determining and outputting a predicted value for at least one of: relative expression units, fold induction, fold anti-induction, fluorescence in an absence of inducer, maximum fluorescence relative to a basal expression of an expression state, a traceability score, or a combination thereof.

2. The computerized method of claim 1, further comprising: applying one or more of the plurality of candidate DNA binding function as a second 2-node single-in single-out (SISO) networks of transcription factors binned via the alternate DNA binding function following the application of the one or more of the plurality of candidate DNA binding function as the first 2-node SISO networks to provide a combination of the first 2-node SISO networks and the second 2-node SISO networks, wherein the combination of the first 2-node SISO networks and the second 2-node SISO networks is equivalent to a multiple input single output (MISO) network.

3. The method of claim 2, wherein the outputted predicted value is for the combination.

4. The method of any one of claims 1-3, wherein the plurality of candidate DNA binding function is applied via two or more of: a 1 -INPUT BUFFER logic operation, a 1 -INPUT NOT logic gate operation,a 2-INPUT AND logic gate operation, a 2-INPUT NOR logic gate operation, a 2-INPUT A NIMPLY B logic gate operation, a 2-INPUT B NIMPLY A logic gate operation, a 2-INPUT XNOR logic gate operation, and a 2-INPUT NAND logic gate operation.

5. The method of claim 4, wherein the 2-INPUT AND logic gate operation comprises:wherein ε is a value for fluorescence in the absence of inducer, σ is a value constant for a maximum fluorescence relative to basal expression of an OFF-state, A+(I) is a coarse-grained Hill function (e.g., having a value of 0 or 1), and AA(I) is a coarse-grained antithetical Hill- function for anti-repression.

6. The method of claim 4, wherein the 2-INPUT NOR logic gate operation comprises:

7. The method of claim 4, wherein the 2-INPUT A NIMPLY B logic gate operation comprises:

8. The method of claim 4, wherein the 2-INPUT B NIMPLY A logic gate operation comprises:

9. The method of any one of claims 1-8, wherein the predicted value for fold induction is determined by: Ω+(1) = σA+(l) + ε10. The method of any one of claims 1-9, wherein the gene model comprises: a lactose repressor (LacI) topology having (i) a regulatory core domain (RCD) and (ii) a DNA binding domain (DBD).

11. The method of any one of claims 1-10, wherein each 2-node SISO network comprises (i) a single transcription factor expressed on the pLacI plasmid and (ii) a super folder fluorescent protein (GFP) reporter.

12. The method of claim 1, wherein the REU is a measure of fluorescence of a reporter fusion to a nucleic acid sequence of a gene of interest.

13. A system comprising: a processor; and a memory having instructions stored thereon, wherein execution of the instructions by the processor causes the processor to perform any one of the computerized methods of any one of claims 1-12.

14. A non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by a processor causes the processor to perform any of the computerized method of any one of claims 1-12.

15. A construct comprising a plurality of nucleic acid sequences encoding a first group of one or more regulatory core domains; a second group of one or more regulatory core domains; one or more DNA binding domains, wherein the first group of the one or more regulatory core domains and the second group of the one or more regulatory core domains are each linked to one of the DNA binding domains formed of a cellobiose-responsive anti- repressor; and one or more DNA operator elements, wherein the one or more DNA operator elements are each specifically recognized by one of the DNA binding domains.

16. The construct of claim 15, wherein a nucleic acid sequence encoding the cellobiose- responsive anti-repressor comprises SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 5, or a variant thereof.

17. The construct of claim 15 or 16, wherein the first group of the one or more regulatory core domains, the second group of the one or more regulatory core domains, and a third group of one or more regulatory core domains are each linked to one of the DNA bindingdomains formed of a cellobiose-responsive anti-repressor, to form a three-input transcription program.

18. The construct of any one of claims 15-17, wherein the first group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

19. The construct of any one of claims 15-17, wherein the second group of one or more regulatory core domains comprises at least one repressor, at least one anti-repressor, or a combination thereof.

20. The construct of any one of claims 15-17, wherein the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises a second repressor.

21. The construct of any one of claims 15-17, wherein the first group of one or more regulatory core domains comprises a first repressor, and the second group of one or more regulatory core domains comprises an anti-repressor.

22. The construct of any one of claims 15-17, wherein the first group of one or more regulatory core domains comprises a first anti-repressor, and the second group of one or more regulatory core domains comprises a second anti-repressor.

23. A method comprising: providing a gene model having network repressors and network anti-repressors, wherein the network repressors and network anti-repressors are designated for binding to a plurality of candidate DNA binding functions; applying one or more of the plurality of candidate DNA binding functions as a 3 -node network of transcription factors binned via the alternate DNA binding function; and determining logical operations for the 3 -node networks of transcription factors, wherein the logical operations are made accessible to biocomputation operation.

Citation Information

Patent Citations

  • Superior bioinformatics process for identifying at risk subject populations

    US20160117439A1