Synthetic promoters for directed evolution

EP4724579A2Pending Publication Date: 2026-04-15UNIV OF UTAH RES FOUND
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Current methods for improving nucleic acid and protein sequences, such as directed evolution, face inefficiencies, high costs, and limited applicability, necessitating the development of more effective approaches for generating variant proteins.

Method used

The method involves a library of synthetic promoters to drive the continuous evolution of transgenes through viral infection, where host cells are cultured with viruses containing a transgene, and the expression of viral genes is controlled by these promoters, allowing for the selection and optimization of promoters that activate in response to specific conditions, thereby enhancing protein evolution.

Benefits of technology

This approach enables the rapid identification and utilization of optimal synthetic promoters, leading to improved protein activity and increased efficiency in generating variant proteins with enhanced properties, such as increased viral packaging and RNA production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000028_0001
    Figure IMGF000028_0001
  • Figure IMGF000028_0002
    Figure IMGF000028_0002
  • Figure IMGF000029_0001
    Figure IMGF000029_0001
Patent Text Reader

Abstract

Provided herein are compositions, systems, and methods for directed evolution. Further provided herein are compositions, systems, and methods comprising use of synthetic promotor libraries for directed evolution.
Need to check novelty before this filing date? Find Prior Art

Description

SYNTHETIC PROMOTERS FOR DIRECTED EVOLUTIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 506,499, filed on June 6, 2023, which is incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under GM 146247 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND

[0003] Improving the properties and function of nucleic acid and protein sequences is applicable to fields such as medicine, diagnostics, and biotechnology. However, many approaches (e.g., directed evolution) may suffer from low efficiency, high cost, or limited applicability. Therefore, there is a need for improved methods of directed evolution.BRIEF SUMMARY

[0004] Provided herein are methods, and systems for generating variant proteins through continuous directed evolution of a transgene, the method comprising: selecting a polynucleotide from a library of polynucleotides; culturing a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle and wherein expression of the at least one viral gene is under control of the synthetic promoter from the library of polynucleotides; and infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene. In some embodiments, library of polynucleotides comprises at least 50 synthetic promoters. In some embodiments, each of the polynucleotides from the library of polynucleotides are 100-300 bases in length. In some embodiments, each of the polynucleotides are at most 300 bases in length. In some embodiments, each of the polynucleotides from the library of polynucleotides are at least 300 bases in length. In some embodiments, each of the polynucleotides from the library of polynucleotides comprises a helical turn. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprise one or more copies of a transcription factor binding motif (TFBM) comprising a DNAsequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site. In some embodiments, each of the one or more copies of TFBM are separated by one or more nucleotides. In some embodiments, each of the polynucleotides from the library of polynucleotides comprises at least four TFBMs. In some embodiments, each of the polynucleotides from the library of polynucleotides comprises four TFBMs. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprise a spacer before, between, and / or after the TFBMs. In some embodiments, the spacer comprises two random sequence spacer sets. In some embodiments, each of the polynucleotides from the library of polynucleotides comprises at least 3 spacers. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprise a minimal promoter, a transcription start site, and / or a DNA sequence capable of initiating transcription. In some embodiments, the TFBM is at least 3 helix positions relative to the minimal promoter. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprises at least 3 minimal promoters. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprise a 5 ’ UTR comprising a detectable nucleotide barcode. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprise a nucleotide sequence encoding a detectable polypeptide. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprise a 3 ’ UTR comprising a detectable nucleotide barcode. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprises in 5’ to 3’ direction: one or more copies of transcription factor binding motif (TFBM) comprising a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site; a minimal promoter, a transcription start site, and / or a DNA sequence capable of initiating transcription; a 5 ’ UTR comprising a detectable nucleotide barcode; a nucleotide sequence encoding a detectable polypeptide; and a 3 ’ UTR comprising a detectable nucleotide barcode. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprises a mammalian transcription response element (TRE). In some embodiments, each of the TREs independently comprises one or more copies of transcription factor binding motif (TFBM) comprising a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site. In some embodiments, each of the TREs independently comprises a spacer before, between, and / or after the TFBMs. In some embodiments, each of the polynucleotides from the library of polynucleotides independently comprises a synthetic promoter. In some embodiments, each of the synthetic promoters independently comprises theTRE and a minimal promoter. In some embodiments, the synthetic promoter is not constitutively active or inactive in the host cell or the virus. In some embodiments, each of the TFBM of the one or more copies of TFBM is comprises a different sequence from the other TFBMs. In some embodiments, the library does not comprise a synthetic promoter activated by viral infection. In some embodiments, the 5’ UTR nucleotide barcode or 3’ UTR nucleotide barcode are unique to the composition of the synthetic promoter. In some embodiments, the TRE comprises a nuclear receptor response element, a type 1 nuclear receptor response element, a type 2 nuclear receptor response element, a nuclear factor of activated T-cells response element (NFAT-RE), a cAMP response element (CRE), a B recognition element, a hypoxia-responsive element, a serum response element (SRE), a serum response factor response element (SRF-RE), a metal- responsive element (MRE), a IFN-stimulated response element (ISRE), a calcium-response element, an antioxidant response element (ARE), a p53 response element, a sterol regulatory element, a polycomb response element (PRE), a rev response element (RRE), and / or a wnt response element (WRE). In some embodiments, the selection comprises: transfecting a population of host cells with the library of polynucleotides; infecting the population of host cells with a population of viruses, wherein the population of viruses comprise a transgene; treating the host cells with an agent capable of activating the transgene; harvesting the host cells; isolating RNA from the harvested host cells; sequencing the isolated RNA; quantifying the RNA and identifying the synthetic promoter based on the RNA sequencing; and selecting the synthetic promoter, wherein the synthetic promoter has a higher number of sequence reads compared to a control. In some embodiments, the RNA comprises a unique detectable nucleotide barcode. In some embodiments, the RNA comprises a reporter. In some embodiments, the synthetic promoter is not constitutively active or inactive in the absence of the agent. In some embodiments, the method comprises: (d) incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene. In some embodiments, the method comprises: (e) generating a second population of viruses expressed by the first population of the host cells. In some embodiments, the method comprises: (f) infecting a second population of host cells with the second population of viruses. In some embodiments, steps (a)-(f) are repeated at least one, two, or five times. In some embodiments, steps (a)-(f) are repeated at least once. In some embodiments, steps (a)-(f) are repeated at least two times. In some embodiments, steps (a)-(f) are repeated at least five times. In some embodiments, the method comprises: (g) amplifying RNA from the second population of viruses; (h) sequencing the RNA; and (i) quantifying the RNA and / or barcode counts. In some embodiments, the transgene is about 500, 1000, or 2000 nucleotides in length. In some embodiments, the transgene is at least about 500 nucleotides in length. In some embodiments, the transgene is at least about1000 nucleotides in length. In some embodiments, the transgene is at least about 2000 nucleotides in length. In some embodiments, the second population of viruses comprise at least one mutated transgene of interest. In some embodiments, the repeating of steps (a)-(f) results in at least one mutation in the transgene of interest. In some embodiments, the mutated transgene of interest comprises at least one nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene of interest comprises at least two nucleic acid substitutions, insertions, or deletions compared to the transgene. In some embodiments, the mutated transgene of interest comprises at most about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene. In some embodiments, the mutated transgene of interest comprises about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene. In some embodiments, the mutated transgene of interest comprises at least about 99% sequence identity to the transgene. In some embodiments, the mutated transgene of interest encodes for a mutated protein comprising at least one amino acid substitution, insertion, or deletion compared to a protein encoded by the transgene. In some embodiments, the mutated transgene of interest comprises same, increased, or decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene. In some embodiments, the mutated transgene of interest comprising improved activity results in a higher number of infectious viral particles. In some embodiments, the mutated transgene of interest comprising improved activity has a higher RNA count than the transgene or other mutated transgenes. In some embodiments, a ratio of quantified RNA to DNA of the transgene is at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, a ratio of quantified RNA to DNA of the transgene is at most about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, a ratio of quantified RNA to DNA of the transgene is about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal. In some embodiments, the agent activates the transgene. In some embodiments, the activity of the transgene is a signaling event. In some embodiments, the activity of the transgene produces an activity- coupled transcriptional response. In some embodiments, the activity-coupled transcriptional response binds to the synthetic promoter. In some embodiments, binding of the activity-coupled transcriptional response with the synthetic promoter results in transcription of the at least oneviral gene for packaging a virus into an infectious viral particle. In some embodiments, the transgene encodes for a protein. In some embodiments, the protein is activated by the agent and produces an activity-coupled transcriptional response wherein the activity-coupled transcriptional response is an endogenous transgene -mediated signaling pathway. In some embodiments, the protein is a G protein-coupled receptor (GPCR). In some embodiments, the agent activates the GPCR and produces a GPCR-mediated signaling response. In some embodiments, the GPCR-mediated signaling response binds to SRE or CRE. In some embodiments, the virus is a mutagenic virus. In some embodiments, the virus has a mutation rate of about 10"5- 10"3mutations per base replicated. In some embodiments, the virus has a mutation rate of about 1 O’3mutations per base replicated. In some embodiments, the virus has a mutation rate of about 1 O’5mutations per base replicated. In some embodiments, the virus comprises an error prone replicase. In some embodiments, the virus is a DNA virus, an RNA virus, or a reverse transcribing virus. In some embodiments, the virus is an RNA virus. In some embodiments, the RNA virus is a positive-strand RNA virus. In some embodiments, the virus is a Sindbis virus. In some embodiments, the host cells are mammalian cells. In some embodiments, the mammalian cells are BHK-21 cells, HEK-293 cells, HEK293T cells, or CHO cells. Provided herein are compositions composition comprising: a library of polynucleotides, wherein the library comprises a plurality of synthetic promoters; and a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle and wherein expression of the at least one viral gene is under control of at least one synthetic promoter from the library of polynucleotides. In some embodiments, the composition further comprises a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene.

[0005] Provided herein are systems configured to perform the steps of one or more methods described herein.INCORPORATION BY REFERENCE

[0006] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will beobtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:

[0008] FIGs. 1A-1B depicts a plot of a dose-response curve measuring the response of GPR40 at various concentrations of Compound 1 (FIG. 1A) or Compound 2 (FIG. IB). The FIG. 1A x-axis is labeled Log[Compound 1] from -10 to -4 at 2 unit intervals; the y-axis is labeled Fluc / RLuc (ratio firefly to Renilla luciferase activity) from 0.0 to 2.0 at 0.5 unit intervals. Data points for experiments with GPR40 are shown as circles; empty vector controls are shown as squares. FIG. IB x-axis is labeled Log[Compound 1] from -10 to -4 at 2 unit intervals; the y-axis is labeled Fluc / RLuc (ratio firefly to Renilla luciferase activity) from 0.0 to 0.8 at 0.2 unit intervals. Data points for experiments with GPR40 are shown as circles; empty vector controls are shown as squares.

[0009] FIGs. 2A-2B depict flow cytometry plots measuring the response in the presence or absence Compound 1 (FIG. 2A) or Compound 2 (FIG. 2B) to compare promoters from a library of promoters. The x-axis is labeled log2(Fold Change) from -6 to 6 at 3 unit intervals. The y-axis is labeled Density.

[0010] FIGs. 3A-3B depict plots of dose-response curves measuring the response of GPR40 at various concentrations of Compound 1 (FIG. 3A) or Compound 2 (FIG. 3B) using a promoter selected from the library of promoters. The FIG. 3A x-axis is labeled Log[Compound 1] starting at “ND” and labeled from -8 to -4 at 2 unit intervals; the y-axis is labeled Fluc / RLuc from 0 to 8 at 2 unit intervals. Legend: GPR40+pGL4.33R (light grey shaded circles); Empty Vector+pGL4.33R (unshaded light grey circles); GPR40+SRE-minCMV (shaded black circles); and Empty Vector+SRE-minCMV (unshaded black circles). The FIG. 3B x-axis is labeled Log[Compound 2] starting at “ND” and labeled from -11 to -6 at 1 unit intervals; the y-axis is labeled Fluc / RLuc from 0 to 25 at 5 unit intervals. Legend: GPR40+pGL4.29R (black shaded circles); Empty Vector (EV) +pGL4.29R (unshaded black circles); GPR40+CRE-minCMV (shaded grey triangles); and Empty Vector (EV)+CRE-minCMV (unshaded triangles).

[0011] FIG. 4 depicts a RT-qPCR plot measuring the quantification cycle over multiple rounds of a VEGAS continuous evolution campaign. Experimental conditions are shown on the x-axis; the y-axis is labeled cq (quantification cycles) from 50 to 0 at 10 unit intervals. Legend: RT (modAmbrose, shaded circles); no RT (modAmbrose, unshaded circles); RT (modVEGAS, shaded squares); and no RT (modVEGAS, unshaded squares). Evolution of the tTA (tetracycline-controlled transactivator) was used as a positive control.

[0012] FIG. 5 shows a summary of Sanger sequencing results of identified mutants in each of the rounds aligned with the GPR40 construct. Round 1 (24 clones): 22 / 24 align to GPR40,1 / 24 failed); Round 2 (24 clones): 2 / 24 align to GPR40, 14 / 24 align to empty backbone; 8 / 24 failed); Round 3 (14 clones): 2 / 12 align to empty backbone; 10 / 12 failed).

[0013] FIGs. 6A-6D depict plots of dose-response curves measuring the response of wildtype of VEGAS mutated GPR40 protein at various concentrations of Compound 1 or Compound 2 with the promoters selected from the library of promoters. FIG. 6A depicts GPR40+Compound 1+SRE; the x-axis is labeled Log[Compound 1] starting at “ND” and labeled from -9 to -4 at 1 unit intervals; the y-axis is labeled Fluc / RLuc from 0 to 10 at 2 unit intervals.. FIG. 6B depicts GPR40+Compound 1+CRE; the x-axis is labeled Log[Compound 1] starting at “ND” and labeled from -9 to -4 at 1 unit intervals; the y-axis is labeled Fluc / RLuc from 0.0 to 2.5 at 0.5 unit intervals.. FIG. 6C depicts GPR40+Compound 2+SRE; the x-axis is labeled Log[Compound 2] starting at “ND” and labeled from -11 to -7 at 1 unit intervals; the y- axis is labeled Fluc / RLuc from 0 to 15 at 5 unit intervals. FIG. 6D depicts GPR40+Compound 2+CRE; the x-axis is labeled Log[Compound 2] starting at “ND” and labeled from -11 to -6 at 1 unit intervals; the y-axis is labeled Fluc / RLuc from 0 to 40 at 10 unit intervals. Legend: WT (black shaded circles); F191S (grey shaded squares); C136R (black shaded triangles); EV (empty vector, grey shaded circles).DETAILED DESCRIPTION OF THE INVENTION

[0014] Provided herein are directed evolution platforms. The VEGAS directed evolution platform (English et al., Cell 2019 Jul 25;178(3):748-761.el7) delivers transgenes to cells via viral infection and evolves these genes through random mutagenesis to produce mutant versions of the transgene. In the VEGAS system, a viral packaging gene is coupled to the activity of said transgene via promoter comprising a transcription response element (TRE). This efficiently couples the signaling pathways and other biological processes upstream of or directly in line with transcription. Mutagenesis of the transgene may result in more or less activity of the transgene, thus resulting in more or less expression of the viral packaging gene. If the transgene accumulates additional advantageous mutations, more virus comprising said mutagenized transgene will be packaged. Transcriptional activation can be encoded on synthetic, exogenous, or endogenous DNA promoters that ultimately induce the host cell to produce RNA that encodes the Sindbis virus nucleocapsid and glycosylation coat proteins (together, Sindbis Structural Genome).

[0015] Thus, a major component of the success of the VEGAS platform is in the selection of promoters activated by cellular events that produce minimal transcriptional activity in the absence of the desired evolved product from the virus and high expression when the correct activation conditions are met by the transgene. Herein is described a method to rapidly identifysubsets of synthetic promoters suitable for performing VEGAS directed evolution on any transgenic target in any cell line. In summary, the approach provides a large scale promoter library to cells in a massively parallel reporter assay format to select said suitable promoters.

[0016] Further provided herein are libraries of synthetic promoters. The composition of the library can be any combination of synthetic, endogenous, exogenous, or other transcription regulatory sequences that drive transcription. Cells are treated with Sindbis virus lacking an encoded transgene or a functional transgene such as a receptor, transcription factor, peptide, RNA ribozyme, or any other functional transgene for which transcriptional activity coupling is to be evolved. The transgene is transduced with an agent such as a ligand, reagent, cofactor, or other stimulating agent, as necessary. Such an agent would also be provided to the empty viral and inert transgene cellular samples as a background control. Activity of all promoter sequences in the library are quantified through barcode reads, reporter assay, or RNA / DNA signal measurement to identify a transcription rate ratio between DNA content in the cells and observed RNA output. Those promoters which are activated by viral infection without transgene or virus expressing inert transgenic material are discarded from the library and the remaining constructs which are responsive to transgene or transgene and agent as necessary, are maintained and the library is recreated as a “minimal screening library” which can be used to quantify responses in the directed evolution platform across ligand dose, time, or other stimulations or cellular conditions to optimize the selection process. Promoters identified as having a large response in the presence of the infected transgene and agent and a minimal response in the absence of agent can be selected for use in the VEGAS system.Definitions

[0017] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which these inventions belong.

[0018] Throughout this disclosure, numerical features are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiments. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range to the tenth of the unit of the lower limit unless the context clearly dictates otherwise. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual values within that range, for example, 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges mayindependently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention, unless the context clearly dictates otherwise.

[0019] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of any embodiment. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.

[0020] Unless specifically stated or obvious from context, as used herein, the term “about” in reference to a number or range of numbers is understood to mean the stated number and numbers + / - 10% thereof, or 10% below the lower listed limit and 10% above the higher listed limit for the values listed for a range.

[0021] As used herein, the term “unique molecular identifier (UMI)” refers to a unique nucleic acid sequence that is attached to each of a plurality of nucleic acid molecules. When incorporated into a nucleic acid molecule, an UMI in some instances is used to correct for subsequent amplification bias by directly counting UMIs that are sequenced after amplification. The design, incorporation and application of UMIs is described, for example, in Int. Pat. Appl. Pub. No. WO 2012 / 142213, Islam et al. Nat. Methods (2014) 11 : 163-166, Kivioja, T. et al. Nat. Methods (2012) 9: 72-74, Brenner et al. (2000) PNAS 97(4), 1665, and Hollas and Schuler, (2003) Conference: 3rd International Workshop on Algorithms in Bioinformatics, Volume: 2812.

[0022] As used herein, the term "barcode" refers to a nucleic acid tag that can be used to identify a sample or source of the nucleic acid material. Thus, where nucleic acid samples are derived from multiple sources, the nucleic acids in each nucleic acid sample are in some instances tagged with different nucleic acid tags such that the source of the sample can be identified. Barcodes, also commonly referred to indexes, tags, and the like, are well known to those of skill in the art. Any suitable barcode or set of barcodes can be used. See, e.g., nonlimiting examples provided in U.S. Pat. No. 8,053,192 and Int. Pat. Appl. Pub. No. W02005 / 068656.Directed Evolution Platforms

[0023] Provided herein are directed evolution platforms comprising systems and methods. In some instances, a directed evolution platform provided herein comprises a viral replication system inside of a transfected cell. Viral replication systems in some embodiments comprise a transgene (e.g., the target sequence to be evolved), a synthetic promoter driven by the activity or presence of the transgene (or transgene product, such as a protein) which induces expression of a viral replication protein. Cells producing transgene or transgene products with the selected properties activate the viral replication protein, leading to increased infectivity.

[0024] Provided herein are methods for continuous evolution of a transgene, comprising selecting an optimal promoter, transfecting mammalian host cells with the optimal promoter and at least one viral gene required for viral packaging, and infecting the host cells with virus comprising a transgene. In some embodiments, virus expressed by the infected host cells comprises a mutated transgene. In some embodiments, the method further comprising infecting a new population of host cells with virus expressed by the infected host cells. In some embodiments, the method is repeated for multiple cycles of infecting host cells and collecting the resulting expressed virus comprising uniquely mutated transgenes to infect a new population of host cells.

[0025] The agent activating the transgene is coupled to an endogenous activity response transcriptional element in order for the agent to activate the promoter. In some embodiments, the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal. In some embodiments, the agent activates the transgene. In some embodiments, the activity of the transgene is a signaling event. In some embodiments, the activity of the transgene produces an activity- coupled transcriptional response. In some embodiments, the activity-coupled transcriptional response binds to the promoter. In some embodiments, binding of the activity-coupled transcriptional response with the synthetic promoter results in transcription of the at least one viral gene for packaging a virus into an infectious viral particle. Exemplary agents include but are not limited to: GPR40 agonists, Fetal bovine serum, Forskolin, ATP / GTP, AICAR, Dexamethasone, CdCh, ZnSCE, LiCh, Deferoxamine, Thapsigargin, Serotonin, Dopamine, Thrombin, Epinephrine, cis-epoxysuccinate, (R)-zn3573.

[0026] The transgene is a gene that is not endogenously expressed by the virus. In some embodiments, the transgene comprises DNA. In some embodiments, the transgene encode for a protein. In some embodiments, the protein is expressed in order to be activated by the agent. In some embodiments, the protein is activated by the agent and produces an activity-coupled transcriptional response wherein the activity-coupled transcriptional response is an endogenous transgene-mediated signaling pathway. In some embodiments, the protein is a G protein-coupledreceptor (GPCR). In some embodiments, the protein is GPR40. In some embodiments, the agent activates the GPCR and produces a GPCR-mediated signaling response. In some embodiments, the GPCR-mediated signaling response binds to SRE or CRE. In some instances the transgene comprises an antibody or antigen binding fragment, such as a monospecific Fab2, bispecific Fab2, trispecific Fab3, monovalent IgG, scFv, bispecific diabody, trispecific triabody, scFv-Fc, sdAb, minibody, nanobody, IgNAR, V-NAR, hcIgG, VHH, or peptibody. Transgenes in some instances also comprise transcription factors.

[0027] Through subsequent round of the continuous evolution method, the transgene accumulates random mutations caused by errors in the mutagenic virus replication process. Therefore, the virus should be a mutagenic virus in order to randomly generate the mutations. In some embodiments, the virus is a mutagenic virus. In some embodiments, the virus has a mutation rate of about 10“9-10-1, 10“8-10“2, 10“7-10“3, 10“6-10“3, 10“5-10“3, lO^-lO-4, or 10“4-10“3mutations per base replicated. In some embodiments, the virus has a mutation rate of at least about 10“9- 10-1, 10“8-10“2, 10“7-10“3, 10“6-10“3, 10“5-10“3, 1 O’5- 1 O’4, or I ’4- 1 ’3mutations per base replicated. In some embodiments, the virus comprises an error prone replicase. In some embodiments, the viral mutations are caused by mutagenic nucleoside analogs. In some embodiments, the virus is a DNA virus, an RNA virus, or a reverse transcribing virus. In some embodiments, the virus is an RNA virus. In some embodiments, the virus is a Sindbis virus.

[0028] In some embodiments, the host cells are mammalian cells. In some embodiments, the mammalian cells are BHK-21 cells, HEK-293 cells, HEK293T cells, or CHO cells.

[0029] Through subsequent round of the continuous evolution method, the transgene accumulates random mutations caused by errors in the mutagenic virus replication process. In some embodiments, the method further comprising sequencing virus expressed by the infected host cells. In some embodiments, the method comprises identifying mutations in the mutated transgene. In some embodiments, the mutated transgene comprises at least one, two, three, four, five, six, seven, eight, nine, or ten nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene comprises at least one nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene comprises one nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene of interest comprises at most about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, or 40% sequence identity to the transgene. In some embodiments, the mutated transgene of interest comprises at least about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, or 40% sequence identity to the transgene. In some embodiments, the mutatedtransgene of interest comprises about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, or 40% sequence identity to the transgene. In some embodiments, the mutated transgene of interest encodes for a mutated protein comprising at least one, two, three, four, five, six, seven, eight, nine, or ten amino acid substitution, insertion, or deletion compared to a protein encoded by the transgene.

[0030] Mutations in the transgene are randomly generated by the mutagenic virus, but may end up having an effect on the activity of the transgene. In some embodiments, the mutated transgene of interest comprises same, increased, or decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises at least about 1 %, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises at least about 1 %, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene.

[0031] As the activity of the transgene improves, there is an increased activity coupled transcription response. The increase in activity coupled transcription response results in more expression of the viral gene required for viral packaging controlled by the promoter activated by the activity coupled transcription response. Thus, as activity of the transgene improves, there is more viral packaging. This evolutionary pressure results in more virus comprising the mutated transgene with greater activity. This method can also be used in the opposite direction to apply selective pressure for less activity in a transgene. In some embodiments, the mutated transgene of interest comprising improved activity results in a higher number of infectious viral particles. In some embodiments, the mutated transgene of interest comprising improved activity has a higher RNA count than the transgene or other mutated transgenes. In some embodiments, the ratio of quantified RNA to DNA of the transgene is at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. In some embodiments, the ratio of quantified RNA to DNA of the transgene is at most about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. In some embodiments, the ratio of quantified RNA to DNA ofthe transgene is about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25.

[0032] Polynucleotide libraries

[0033] Libraries of polynucleotides may be used with a directed evolution system provided herein. In some instances, identifying a polynucleotide library comprises synthetic promoters. Polynucleotide libraries in some instances comprise at least 10, 20, 50, 100, 200, 500, 700, 1000, 2000, 5000, 7000, or at least 10,000 unique sequences. Polynucleotide libraries in some instances comprise no more than 10, 20, 50, 100, 200, 500, 700, 1000, 2000, 5000, 7000, or no more than 10,000 unique sequences. Polynucleotide libraries in some instances comprise about 10-10,000, 10-50,000, 10-5,000, 10-1000, 10-500, 10-100, 50-10,000, 50-50,000, 50-5,000, 50- 1000, 50-500, 1000-10,000, 1000-50,000, 5,000-50,000, or 10,000-50,000 unique sequences. Polynucleotides in a library in some instances are about 25-1000, 50-1000, 100-1000, 200-500, 200-300, 500-1000, 500-750, or 750-1000 bases in length. Polynucleotides in a library in some instances are about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 350, 400, 450, 500, or about 1000 bases in length.

[0034] Each of the polynucleotides in the library of polynucleotides comprises a unique sequence. In some instances, the polynucleotide is part of a biological circuit which is activated in response to a stimulus (e.g., a transgene protein and / or agent). In some embodiments, each of the polynucleotides comprise one or more copies of a transcription factor binding motif (TFBM). In some embodiments, each of the polynucleotides comprise one or more copies of a unique transcription factor binding motif (TFBM). In some embodiments, the TFBM comprises a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site. In some embodiments, each of the polynucleotides comprises at least one, two, three, four, or five TFBMs. In some embodiments, each of the polynucleotides comprises one, two, three, four, or five TFBMs. In some embodiments, each of the polynucleotides comprises four TFBMs. In some embodiments, each of the one or more copies of TFBM are separated by one or more nucleotides. In some embodiments, each of the polynucleotides independently comprise a spacer before, between, and after the TFBMs. In some embodiments, each of the polynucleotides independently comprise a spacer before, between, or after the TFBMs. In some embodiments, each of the polynucleotides comprises at least 1, 2, 3, 4, 5, or 6 spacers. In some embodiments, each of the polynucleotides comprises 1, 2, 3, 4, 5, or 6 spacers. In some embodiments, each of the polynucleotides comprises 3 spacers. In some embodiments, the spacer comprises two random sequence spacer sets. In some embodiments, each of the polynucleotides independently comprise a minimal promoter, a transcription start site, and / or a DNA sequence capable of initiating transcription. Insome embodiments, each of the polynucleotides comprises a helical turn (“rotation”). In some embodiments, the TFBM is at least 3 helix positions relative to the minimal promoter. In some embodiments, each of the polynucleotides independently comprises at least 3 minimal promoters. In some embodiments the minimal promotor comprises minCMV, minSV40, miniTK (thymidine kinase), minProm (pGL4 plasmid), or other minimal promoter. In some embodiments, each of the polynucleotides independently comprise a 5 ’ UTR comprising a detectable nucleotide barcode. In some embodiments, each of the polynucleotides independently comprise a 3 ’ UTR comprising a detectable nucleotide barcode. In some embodiments, each of the polynucleotides independently comprise a nucleotide sequence encoding a detectable polypeptide. In some embodiments, each of the polynucleotides independently comprises in the 5 ’ to 3 ’ direction: one or more copies of transcription factor binding motif (TFBM) comprising a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site; a minimal promoter, a transcription start site, and / or a DNA sequence capable of initiating transcription; a 5 ’ UTR comprising a detectable nucleotide barcode; a nucleotide sequence encoding a detectable polypeptide; and a 3’ UTR comprising a detectable nucleotide barcode. In some embodiments, each of the polynucleotides independently comprises a synthetic promoter comprising in the 5’ to 3’ direction: a first TFBM, two random sequence spacer sets; a second TFBM; two random sequence spacer sets; a third TFBM; two random sequence spacer sets; a fourth TFBM; three helix positions relative to a minimal promoter; and a minimal promoter. In some instances, all TFBMs comprise the same sequence. In some embodiments, the synthetic promoter drives expression of a reporter. In some instances the reporter comprises Luc2. In some instances the reporter is flanked by a barcode. In some instances, polynucleotides do not comprise a spacer set. A polynucleotide in some instances comprises a promoter and one or more sequences (such as a barcode or other sequence).

[0035] In some embodiments, the polynucleotide may comprise a transcription response element (TRE) that comprises at least one TFBM and at least one spacer. In some embodiments, each of the polynucleotides independently comprises a mammalian transcription response element (TRE). In some embodiments, each of the TREs independently comprises one or more copies of transcription factor binding motif (TFBM) comprising a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site. In some embodiments, each of the TREs independently comprises a spacer before, between, and / or after the TFBMs. In some embodiments, the TRE comprises a nuclear receptor response element, a type 1 nuclear receptor response element, a type 2 nuclear receptor response element, a nuclear factor of activated T-cells response element (NFAT-RE), a cAMP response element (CRE), a B recognition element, a hypoxia-responsive element, a serumresponse element (SRE), a serum response factor response element (SRF-RE), a metal- responsive element (MRE), a IFN-stimulated response element (ISRE), a calcium-response element, an antioxidant response element (ARE), a p53 response element, a sterol regulatory element, a polycomb response element (PRE), a rev response element (RRE), and / or a wnt response element (WRE). TREs in some instances comprise any sequence which is responsive to the presence of an agent or mutated transgene protein.

[0036] In some embodiments, the polynucleotide may comprise a unique synthetic promoter that comprises at least one TFBM, at least one space, and at least one minimal promoter. In some embodiments, the synthetic promoter comprises a TRE and at least one minimal promoter. In some embodiments, each of the polynucleotides independently comprises a synthetic promoter. In some embodiments, each of the synthetic promoters independently comprises the TRE and a minimal promoter. In some embodiments, each of the TFBM of the one or more copies of TFBM is comprises a different sequence from the other TFBMs. In some embodiments, the 5 ’ UTR nucleotide barcode or 3 ’ UTR nucleotide barcode are unique to the composition of the synthetic promoter. In some embodiments, the synthetic promoter is not constitutively active or inactive in the host cell or the virus. In some embodiments, the library does not comprise a synthetic promoter activated by viral infection.

[0037] Massively Parallel Reporter Assay

[0038] A promoter for use in a VEGAS continuous directed evolution campaign optimally produces a large response when activated by the transgene. Selecting individual promoters to determine the best promoters in a large library would be time and resource intensive. To select a promoter for VEGAS, a parallel screen can be performed in a VEGAS-like system to optimally select a responsive promoter. This massively parallel reporter assay (MPRA) is similar to VEGAS as it involves infecting a host cell with a virus and generating a response in a transgene to an agent, but the output is not a mutated transgene or a packaged virus particle. Instead of encoding for a virus particle, the polynucleotide comprises a gene encoding for a report and / or a barcode. Therefore, the response can be measured by either the reporter in an assay or barcode reads instead of virus particle output. This is a simplified VEGAS system, but mimics the system enough that promoters can be high throughput tested to determine the best responders in a VEGAS campaign.

[0039] In some embodiments, selecting a polynucleotide, which each comprise a unique synthetic promoter, comprises transfecting host cells with the library of polynucleotides, infecting the host cells with a population of virus comprises a transgene, treating with an agent, and quantifying the response to the agent by a reporter assay or quantifying barcode reads. In some embodiments, the selection comprises: transfecting a population of host cells with thelibrary of polynucleotides; infecting the population of host cells with a population of viruses, wherein the population of viruses comprise a transgene; treating the host cells with an agent capable of activating the transgene; harvesting the host cells; isolating RNA from the harvested host cells; sequencing the isolated RNA; quantifying the RNA and identifying the synthetic promoter based on the RNA sequencing; and selecting the synthetic promoter, wherein the synthetic promoter has a higher number of sequence reads compared to a control. In some embodiments, the RNA comprises a unique detectable nucleotide barcode. In some embodiments, the RNA comprises a reporter. In some embodiments, the RNA comprises a unique detectable nucleotide barcode and a reporter. In some embodiments, the synthetic promoter is not constitutively active or inactive in the absence of the agent. In some embodiments, the selecting comprises determining the synthetic promoters to have the highest response in the presence of agent. In some embodiments, the selecting comprises choosing a higher responding synthetic promoter. In some embodiments, the selecting comprising using the chosen synthetic promoter in a VEGAS continuous directed evolution campaign.

[0040] A reporter system may be optimized to minimize background signal (e.g., in the absence of an agent or activated transgene protein). In some instances, a reporter system generates an observable signal that is at least 10%, 25%, 50%, 100%, 200%, 500%, 10 times, 20 times, 50 times, or at least 100 times higher in the presence of an agent or activated transgene protein. In some instances, a reporter system generates an observable signal that is at least 10%, 25%, 50%, 100%, 200%, 500%, 10 times, 20 times, 50 times, or at least 100 times higher in the presence of an agent or activated transgene protein when compared to an empty vector. In some instances, a reporter system generates an observable signal that is at least 10%- 1000%, 25- 1000%, 50%-1000%, 100%- 100%, 200%-1000%, 500%-100%, 5-10 times, 5-20 times, or 10-50 times higher in the presence of an agent or activated transgene protein when compared to an empty vector.

[0041] For evolution of the transgene, mutagenesis may be used to evolve new transgene variants. In some instances, the presence of the transgene is configured to activate a synthetic promoter from a promotor library. In some instances, viruses comprising a high mutation rate are used to induce mutations in the transgene. In some instances, the transgene encodes for a protein. In some instances, the protein is activated by the agent and produces an activity-coupled transcriptional response wherein the activity-coupled transcriptional response is an endogenous transgene-mediated signaling pathway. In some instances, the protein is a membrane -bound protein. In some instances, the protein is a G protein-coupled receptor (GPCR). In some instances. In some instances, the agent activates a GPCR and produces a GPCR-mediatedsignaling response. In some instances, the GPCR-mediated signaling response binds to an SRE (serum-response element) or CRE (cAMP-response element).

[0042] Agents

[0043] In some instances, an agent is added to activate an endogenous pathway in the cell which activates the synthetic promoter from a promotor library. In some instances, an agent is added to activate the same endogenous pathway in the cell as the transgene. For example, an agent that is an agonist for a transgene protein is added. In some instances, an agent mimics the activated state of a transgene product. As the amount of the agent is decreased, viral replication will only occur if the transgene itself has evolved into a more active state. In some instances, the agent inhibits activation of a synthetic promoter from a promotor library. The agent in some instances is titrated to control the selection pressure on a viral system. The agent in some instances is titrated to increase the selection pressure on a viral system to evolve more active transgenes. In some instances, the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal. In some instances, the agent activates the transgene. In some instances, the activity of the transgene is a signaling event. In some instances, the activity of the transgene produces an activity-coupled transcriptional response. In some instances, the activity-coupled transcriptional response binds to the synthetic promoter. In some instances, binding of the activity-coupled transcriptional response with the synthetic promoter results in transcription of the at least one viral gene for packaging a virus into an infectious viral particle.

[0044] In some embodiments, the agent comprises Compound 1 or Compound 2. In some embodiments, Compound 1 comprises fasiglifam (TAK-875) (CAS No.: 1000413-72-8). Fasiglifam is a potent, selective and orally bioavailable GPR40 agonist with EC50 of 72 nM. In some embodiments, Compound 2 comprises Compound A (see: Rives et al. Mol Pharmacol, (2018), 93(6), 581-591).

[0045] Methods and evolution systems provided herein may provide for rapid generation of mutations in the transgene. In some instances, a system provided herein has a mutation rate of at least 10"5, 10"4, 10"3, IO-2, or at least 10"1mutations per base per hour. In some instances, a system provided herein has a mutation rate of about 10“5-10-1, 10“5-10“2, lO^-lO-1, 10“4-10“2, 10"3- 10'1, 1 O'3- 10“2, or about 10“2- 10'1mutations per base per hour. In some instances, a system provided herein has a mutation rate of about 10"5, 10"4, 10"3, 10"2, or about 10"1mutations per base per hour.

[0046] Viruses

[0047] Viruses may be used with the systems and methods for directed evolution provided herein. In order to produce variants, a virus which promotes mutations (e.g., a mutagenic virus)in the transgene may be used. In some instances, the virus comprises an error prone replicase. In some instances, the virus is a DNA virus, an RNA virus, or a reverse transcribing virus. In some instances, the RNA virus is a positive-strand RNA virus. In some instances the positive-strand virus includes but is not limited to Bymoviruses, comoviruses, nepoviruses, nodaviruses, picornaviruses, potyviruses, sobemoviruses luteoviruses Carmoviruses, dianthoviruses, flaviviruses, pestiviruses, statoviruses, tombusviruses, single-stranded RNA bacteriophages, hepatitis C virus Alphaviruses, carlaviruses, fiiroviruses, hordeiviruses, potexviruses, rubiviruses, tobraviruses, tricomaviruses, tymoviruses, apple chlorotic leaf spot virus, beet yellows virus and hepatitis E vims. In some instances, the vims is an alphavirus. In some instances, the alphavirus is a rubi-like, tobamo-like, or tymo-like viruses. In some instances, the vims is a Sindbis virus. In some instances, a vims is altered to decrease its replication fidelity. In some instances, the vims lacks proof-reading polymerases or other mechanisms which increase mutation rate. In some instances, the vims has a mutation rate of about 10"5- 10"1, 10"5- 10"2, 10"4- 1 O'1, 1 O'4- 1 O'2, 1 O'3- 1 O'1, 1 O'3- 1 O'2, or about 1 O'2- 1 O'1, mutations per base replicated. In some instances, the vims has a mutation rate of about 10"5, 10"4, 10"3, 10"2, or about 10"1mutations per base replicated. In some instances, the vims has a mutation rate of at least 10"5, 10"4, 10"3, 10"2, or at least 10"1mutations per base replicated.

[0048] Cells

[0049] Directed evolution systems provided herein may be used inside various cell types. In some instances, the cell comprises an animal, plant, fungal, archaeal, or bacterial cell. In some instances, the cell comprises a mammalian cell. In some instances, the mammalian cell is derived from a human, pig, cow, hamster, or other mammal. In some instances, the mammalian cell does not comprise a human cell. In some instances, a cell is selected such that it can be infected by a virus described herein. In some instances, the mammalian cell is genetically modified, for instance to incorporate machinery for directed evolution. In some instances, the host cell and virus are modified for dependency / complementarity or safety.

[0050] Methods

[0051] Provided herein are methods for directed evolution. In some instances, a directed evolution platform is configured to continuously perform the steps of (1) mutagenesis, (2) expression, and (3) screening. In some instances, mutagenic outcomes of interest are linked to cellular activation of a transcriptional element, such as a synthetic promotor. The synthetic promoter in some instances drives a measurable output, such as viral replication and / or a reporter assay. A method for continuous evolution of a transgene in some instances comprises one or more steps of selecting a synthetic promoter from a library of polynucleotides; culturing a first population of host cells; and infecting the first population of host cells with a firstpopulation of viruses. In some embodiments, the method for directed evolution comprises, selecting a synthetic promoter from a library of polynucleotides, culturing a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle and wherein expression of the at least one viral gene is under control of the synthetic promoter from the library of polynucleotides, and infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene. In some embodiments, the selecting comprises a method described in the specification.

[0052] In some embodiments, the first population of host cells comprises at least one viral gene for packaging a virus into an infectious viral particle comprising a transgene and a subset of viral genes. In some embodiments, the host cell is transfected with a plasmid. In some embodiments, the plasmid encodes for at least a synthetic promoter and a viral gene for packaging a virus comprises a viral capsid gene under control of said synthetic promoter. In some embodiments, a viral gene for packaging a virus comprises one or more of a viral capsid gene, E3, E2, and E2 of a virus (e.g., Sindbis virus). In some instances, expression of the at least one viral gene is under control of the synthetic promoter from the library of polynucleotides. The subset of viral genes in some instances comprises viral packaging genes. The subset of viral genes in some instances comprises one or more of nSPl, nSP2, nSP3, and nSP4 of a virus (e.g., Sindbis virus).

[0053] In some embodiments, infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene. In some embodiments, the subset of viral genes comprise non- structural proteins, capsid, two envelope proteins. In some embodiments, the subset of viral genes comprise four non-structural proteins, capsid, two envelope proteins. In some embodiments, the subset of viral genes comprise non-structural proteins. In some embodiments, the subset of viral genes comprise capsid. In some embodiments, the subset of viral genes comprise envelope proteins. In some embodiments, the subset of viral genes does not comprise a non-structural protein, the capsid, or an envelope protein. In some embodiments, the subset of viral genes does not comprise the viral gene transfected into the host cell. In some embodiments, the virus cannot be packaged without expression of the viral gene transfected into the host cell.

[0054] In some embodiments, the method further comprises incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene. In some embodiments, the method further comprises generating a second population of viruses expressed by the first population of the host cells. In some embodiments, the method further comprises infecting a second population of host cells with the second population of viruses. In some embodiments, the method comprises incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene, generating a second population of viruses expressed by the first population of the host cells, and generating a second population of viruses expressed by the first population of the host cells. In some embodiments, the method is repeated at least repeated at least once, twice, three times, four times, five times, six times, seven times, eight times, nine times, or ten times. In some embodiments, the method is repeated at least once, twice, three times, four times, five times, six times, seven times, eight times, nine times, or ten times.

[0055] The infection and subsequent expression of the virus results in mutated transgenes. To analyze the results, the virus is sequenced to detect mutations in the transgene. In some embodiments, the method comprises amplifying RNA from the second population of viruses. In some embodiments, the method comprises sequencing the RNA. In some embodiments, the method comprises quantifying the RNA and / or barcode counts. In some embodiments, the method comprises amplifying RNA from the second population of viruses, sequencing the RNA, and quantifying the RNA and / or barcode counts.

[0056] Additional evolution methods may be utilized with the systems and methods described herein. In some instances, the systems and methods are configured for evolution in mammalian cells. In some instances, one or more essential viral genes are removed and placed under control of a synthetic promoter from a polynucleotide library.

[0057] For example, synthetic promoter libraries may be used to select appropriate response elements for use in host-dependent propagation of viral like particles (VLV). Mutant transgenes with the desired properties activate response elements. These response elements in turn regulate VSVG (Indiana vesiculovirus G coat protein), which allows viruses harboring the mutant transgene to package and infect additional cells. An exemplary VLV system (PROTEUS) is described in Cole et al., bioRxiv 2024.04.20.590384.

[0058] An engineered adenovirus may also be employed, wherein the adenovirus has genes El, AdPol, AdProt, and E3 deleted. These genes are instead complemented by the host cell, including a constitutively active error prone polymerase (e.g., AdPol) to induce mutations in the transgene. Mutated transgenes are configured to activate adenoviral protease (AdProt) needed for viral replication. The synthetic promoter libraries described herein may be used to driveexpression of AdProt in response to the transgene. An exemplary adenovirus system (mPACE) is described in Berman et al. (J Am Chem Soc. 2018 December 26; 140(51): 18093-18103).

[0059] Multiple cycles of evolution may be used to evolve a transgene. The continuous evolution system in some embodiments, is executed for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or more than 30 days. In some embodiments, the continuous evolution system is executed for no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or no more than 30 days. In some embodiments, the continuous evolution system is executed for 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or no more than 30 days. In some embodiments, the continuous evolution system is executed for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or 30 cycles. In some embodiments, the continuous evolution system is executed for no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or 30 cycles. In some embodiments, the continuous evolution system is executed for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or 30 cycles.

[0060] After optimization, final promoters with optimal RNA / DNA ratios are chosen for the directed evolution process. Optimal RNA / DNA ratios in some instances are at least about 0.1, 0.2, 0.3 0.4, 0.5, 0.7, 0.9, 1, 1.5, 2, 2.5, or 3. Optimal RNA / DNA ratios in some instances are at most about 0.1, 0.2, 0.3 0.4, 0.5, 0.7, 0.9, 1, 1.5, 2, 2.5, or 3. Optimal RNA / DNA ratios in some instances are about 0.1, 0.2, 0.3 0.4, 0.5, 0.7, 0.9, 1, 1.5, 2, 2.5, or 3.NUMBERED EMBODIMENTS

[0061] Provided herein are numbered embodiments 1-19. Embodiment 1. A library comprising mammalian promoter sequences of at least 50 polynucleotides that differ in nucleotide sequence from each other, wherein each of the at least 50 polynucleotides independently comprises, in 5' to 3' direction: (a) one or more copies of transcription factor binding motifs (TFBMs), broadly defined as any DNA sequence wherein a polypeptide, protein, or other agent capable of inducing transcription can be targeted for binding upstream of a transcriptional start site. Optionally wherein each of the one or more copies of the TFBM are separated from its adjacent one(s) by one or more nucleotides; (b) optionally an additional spacer between TFBMs or after the TFBMs; (c) a minimal promoter, transcription start site, or other DNA sequence capable of initiating transcription; (d) optionally a 5 ’ UTR containing a detectable nucleotide barcode; (e) optionally a nucleotide sequence encoding a detectable polypeptide; and (f) optionally a 3' UTR containing a detectable nucleotide barcode; wherein the promoter is not constitutively active or inactive in a cell or virus-infected cell; wherein the TFBM composition in each of the at least 50 polynucleotide promoters, or the barcode identity, is different from the TFBM in each of the others; wherein the library does not comprise a polynucleotide comprising a synthetic promoter that is activated by virus infection of a cell; andwherein the 5 ’ or 3 ’ UTR nucleotide barcodes are unique to the composition of the synthetic promoter. Embodiment 2. The method of embodiment 1, wherein the library comprises promoters >50 polynucleotides. Embodiment 3. The method of any of the preceding embodiments, wherein each polynucleotide is present in a separate construct. Embodiment 4. The method of any of the preceding embodiments, wherein the detectable polypeptide is luciferase. Embodiment 5. A method of preparing a library of at least 100 polynucleotides, the method comprising: (a) obtaining a first library of polynucleotides, wherein each of the polynucleotides of the first library independently comprises, in 5' to 3' direction: (i) two or more copies of a transcription factor binding motif (TFBM), optionally wherein each of the two or more copies of the TFBM is separated from its adjacent one(s) by one or more nucleotides; (ii) optionally an additional spacer; (iii) a promoter; (iv) optionally a 5 ’ UTR containing a detectable nucleotide barcode; (v) a nucleotide sequence encoding a detectable polypeptide; and (vi) a 3' UTR containing a detectable nucleotide barcode; wherein the encoded detectable polypeptide in each of the polynucleotides of the first library is the same; and (b) selecting from the first library at least 100 polynucleotides, wherein the TFBM in each of the selected at least 100 polynucleotides: is less than 30 nucleotides in length, is not constitutively active in a cell, is not constitutively inactive in a cell, and is not activated by virus infection of a cell; wherein the library of at least 100 polynucleotides does not comprise a polynucleotide comprising a synthetic promoter that is activated by virus infection of a cell. Embodiment 6. A method of identifying a promoter, the method comprising: (a) transfecting a first plurality of mammalian cells with the library of embodiment 1 ; (b) treating the transfected first plurality of mammalian cells of step (a) with an agent(s); (c) culturing the first plurality of mammalian cells of step (b); (d) harvesting the cultured first plurality of mammalian cells of step (c), (e) isolating RNA from the harvested first plurality of mammalian cells of step (d); (f) sequencing the isolated promoter library barcoded RNA from step (e); and (g) identifying a library promoter that is upregulated compared to a sequenced RNA control, wherein the RNA control comprises RNA obtained from a second plurality of mammalian cells transfected with the library of embodiment 1 and cultured in the absence of one or more agent(s); and (h) The promoter is not constitutively active in the plurality of mammalian cells in the absence of any agent(s); (i) Is not constitutively inactive in the plurality of mammalian cells after addition of any agent(s). Embodiment 7. A polynucleotide comprising, in 5' to 3' direction: (a) two or more copies of a TFBM, optionally wherein each of the two or more of the first MTRE is separated from its adjacent one(s) by one or more nucleotides; (b) optionally a first linker; (c) a first promoter; (d) a first nucleotide sequence encoding each structural protein of an RNA virus; (e) a first 3' UTR; wherein the MTRE is not constitutively active, is not constitutively inactive, is not activated by virus infection, and is nota previously disclosed or known synthetic promoter comprising these elements. Embodiment 8. A method of performing directed evolution of a transgene in an RNA virus, the method comprising: (a) culturing a first plurality of mammalian cells that contain a first polynucleotide comprising, in 5' to 3' direction, a promoter as identified in embodiment 6, 5' to the inframe DNA nucleotide sequences encoding the polypeptide structural genome components of an RNA virus; (b) infecting the first plurality of mammalian cells of step (a) with a transgenic RNA virus, wherein at least the transgenic RNA virus comprises an RNA encoding the transgene, and wherein the RNA virus does not comprise an RNA encoding all the structural proteins of the RNA virus in the event that the composition is changed in such a way that the virus does encode a component of the structural genome; (c) contacting the first plurality of mammalian cells of step (b) with any agent(s); (d) culturing the cells of step (c); and (e) generating infectious RNA viruses, wherein the generating of infectious RNA viruses is increased by directed evolution of the transgene to activate the promoter defined in 7. Embodiment 9. A method for the selection of mammalian promoters for application in directed evolution. Embodiment 10. The method of embodiment 9, wherein mammalian promoters are selected for use for directed evolution in which a virus is used as a vector for genes of interest to be evolved. Embodiment 11. The method of embodiment 9, wherein mammalian promoters are selected for use for directed evolution in which an Alphavirus is used as a vector for genes of interest to be evolved. Embodiment 12. The method of embodiment 9, wherein mammalian promoters are selected for use for directed evolution in which a Sindbis virus is used as a vector for genes of interest to be evolved. Embodiment 13. The method of embodiment 9, wherein activation of the promoter must be below a certain threshold prior to viral infection and subsequent transgene expression and above a specific threshold when the transgene is expressed or propagated in the target cell either with or without a co-factor agent for activation. Embodiment 14. The method of embodiment 9, wherein viral infection absent the affiliated transgene does not cause the promoter to activate above the threshold specified in embodiment 13. Embodiment 15. The method of embodiment 9, wherein the identified promoter in embodiment 13 and embodiment 14 is applied for directed evolution selection. Embodiment 16. The method of embodiment 9, wherein the identified synthetic promoter in embodiment 13 and embodiment 14 is applied to express the structural genome for a virus in a directed evolution paradigm. Embodiment 17. The method of embodiment 9, wherein the identified promoter in embodiment 13 and embodiment 14 is applied to express the structural genome for an Alphavirus in a directed evolution paradigm. Embodiment 18. The method of embodiment 9, wherein the identified promoter in embodiment 13 and embodiment 14 is applied to express the structural genome for a Sindbisvirus in a directed evolution paradigm. Embodiment 19. One or more of the peptides, nucleic acid sequences, compositions, methods, or libraries described herein.EXAMPLES

[0062] The following examples are set forth to illustrate more clearly the principle and practice of embodiments disclosed herein to those skilled in the art and are not to be construed as limiting the scope of any claimed embodiments. Unless otherwise stated, all parts and percentages are on a weight basis.Example 1: Directed Evolution with TRE Library

[0063] A directed evolution campaign was used to identity transgene mutants of activated GPR40 using the VEGAS platform and a TRE library. VEGAS directed evolution steps were performed according to the general methods of English et al., Cell 2019 Jul 25;178(3):748- 76 Eel 7. The goal of the VEGAS campaign was to evolve a constitutively active mutant of GPR40.

[0064] Before beginning the VEGAS campaign on GPR40, SRE and CRE promoters were tested using a Transcriptional Response Element (TRE)-Luciferase assay and commercially available SRE and CRE transcriptional response elements. Both SRE and CRE displayed an expected response upon application of increasing amounts of Compound 1 (fasiglifam or TAK- 875, PubChemID: 24857286, CAS No.: 1000413-72-8) or Compound 2 (“Compound A”, see Rives et al. Molecular Pharmacology June 1, 2018, 93(6), 581-591; PubChemID: 134465785); however, when compared to a previously successfully evolved system the transcriptional response of CRE and SRE with GPR40 was relatively low (FIG. 1A and FIG. IB). The signaling circuit between the transgene and promoter used in the VEGAS platform is important for the success of the evolution campaign, therefore a promoter with a more robust response was optimized.

[0065] Identification of Promoters from a TRE library. To generate a biological circuit / promoter system for use with directed evolution, a TRE library was first evaluated for its ability to drive a reporter in response to GPR40 activation. BHK-21 cells were maintained in MEMalpha supplemented with 5% fetal bovine serum, 10% tryptose phosphate broth, and IX penicillin / streptomycin (‘growth media’). For TRE-MPRA experiments, 5 x 106cells were plated on 15 cm treated tissue culture dishes in growth media. The following day, cells were transfected with 10 pg of the TRE library (Zahm et al. bioRxiv 2023.05.11.539703) comprising a massively parallel reporter assay library containing 6184 synthetic promoters, each less than 250 bp in length using TransIT-2020 (Mirus Bio) according to manufacturer’s instructions. Synthetic promotors in the library were assembled from combinations of TFBMs, two spacer sets, three different rotations, and three minimal promoters (minCMV, minPromega, and minTK). Thegeneral structure of an exemplary synthetic promoter of the library such as that of IRF9-1-1753, 8, set2, minCMV, comprised: AACGAAACCGAAACTCAACGAAACCGAAACTAACGAAACCGAAACTTAACGAAAC CGAAACTTCCC7NCGG7NCC (SEQ ID NO: 15); wherein the underlined regions are TFBMs, the bold regions are spacer sets, and italicized regions represent the rotation region. IRF9-1-1753 comprised 8 rotations, contained spacer set 2, and further included the minCMV minimal promoter at the 3 ’ end: (GTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAG ATC) (SEQ ID NO: 30).

[0066] After six hours, the media was replaced with growth media lacking fetal bovine serum (‘serum- free media’) and cells were cultured overnight at 37 °C. The next day, cells were treated with Sindbis virus packaged either with an empty pTSin4-derived ssRNA or with Sindbis virus packed with pTSin4-GPR40-derived ssRNA. Cells were treated with a GPR40 agonist (Compound 1 or Compound 2).

[0067] RNA fractions from cell pellets were isolated using QIAshredder homogenizers and the AllPrep DNA / RNA Mini Kit (Qiagen). Immediately after homogenization, a pool of four synthesized spike-in RNAs (2.5 fM each, 10 fM total) was added to each 600 pL sample (Table S4). Spike-in RNAs serve to identify any samples with poor barcode recovery. RNA fractions were eluted in 30 pL nuclease-free water. Following isolation, the RNA fraction was treated with TURBO DNase (Invitrogen) to remove any plasmid DNA carryover. DNA removal from RNA fractions was confirmed by RT-PCR using the SuperScript IV One-Step RT-PCR System (Invitrogen) with and without addition of SuperScript IV RT Mix, followed by agarose gel electrophoresis. RT-PCR was performed according to manufacturer’s instructions with the following conditions: 55 °C for 10 minutes, 98 °C for 2 minutes, 30 cycles of PCR (98 °C for 10 seconds, 60 °C for 10 seconds, 72 °C for 8 seconds), and 72 °C for 5 minutes.

[0068] Dual-indexed amplicons from RNA samples were generated using the SuperScript IV One-Step RT-PCR System (Invitrogen). 0.5- 1.0 pL of RNA is used as the template in 20 pL reactions.

[0069] PCR was performed according to manufacturer’s instructions with the following conditions: 98 °C for 30 seconds, 16 cycles of PCR (98 °C for 5 seconds, 60 °C for 10 seconds, 72 °C for 5 seconds), and 72 °C for 5 minutes. Products were separated by agarose gel electrophoresis and library amplicons are extracted using the QIAquick Gel Extraction Kit (Qiagen). Amplicons were quantified with the KAPA Library Quantification Kit for Illumina Platforms (KAPA Biosystems) in 384-well format using a CFX Opus 384-Well Real-Time System (Bio-Rad). Libraries prepared from RNA samples were pooled at equimolarconcentrations and combined with plasmid DNA input amplicons such that plasmid amplicons represent approximately 6-8% of the pool. RNA and plasmid DNA libraries were sequenced using a 50 cycle SP flow cell on a NovaSeq 6000 (Illumina) with custom sequencing primers.

[0070] Raw barcode counts from TRE-MPRA sample demultiplexed fastq files were processed using a custom pipeline that relied upon MPRAnalyze, a statistical framework developed for use with MPRA data sets. The input to the pipeline was the raw barcode counts of each sample and the outputs were estimations of transcription rates for each synthetic promoter in each treatment condition and the identification of differential promoter activity between samples. This pipeline was used to compare cells infected with empty (transgene- free) Sindbis virus to cells infected with Sindbis virus containing a transgenic cargo. Cargo-specific synthetic promoters identified by the custom pipeline could then be assessed in an orthogonal dualluciferase assay to further validate the identified promoter.

[0071] FIGs. 1A-1B depict plots of a dose-response curve measuring the ratio firefly to Renilla luciferase activity (Fluc / Rluc) at various concentrations of Compound 1 (FIG. 1A) or Compound 2 (FIG. IB). The promoter used in this experiment was a SRE promoter, pGL4.33R, in FIG. 1A and a CRE promoter, pGL4.29R, in FIG. IB. The SRE promoter had a maximal detected Fluc / Rluc ratio of about 1.5 for the protein GPR40. The CRE promoter had a maximal detected Fluc / Rluc ratio of about 0.6 for the protein GPR40. The SRE and CRE displayed an expected response upon application of increasing amounts of their respective compounds; however, when compared to a previously successfully evolved systems, the transcriptional response of SRE and CRE with GPR40 was relatively low. The signaling circuit between the transgene and promoter used in the VEGAS platform is important for the success of the evolution campaign, therefore a promoter with a more robust response should be selected.

[0072] FIGs. 2A and 2B depict flow cytometry plots measuring the response to Compound 1 (FIG. 2A) or Compound 2 (FIG. 2B). The leftmost unshaded area plots show response with vehicle only and the rightmost shaded area plots show treatment with compound. The results show the top 10 identified promoters that had the largest fold change response in the presence of the compounds. These hits identified from the library could be used in the system as described above to improve the response and gain a greater signal to noise ratio. The Promega SRE promoter (Promega_SRE-2048, 8, none, minCMV) and the Promega CRE promoter (Promega_CRE-2002, 1, none, minCMV) were identified as potentially improving the response for treatment with Compound 1 and Compound 2, respectively. Alternatively, the promoters could be swapped given their robust responses with either Compound 1 or Compound 2. Sequences of the top 10 identified TFBM motifs and TREs are shown in Tables 1 and 2, respectively. Minimal promoter sequences are shown in Table 3. Sequences of the syntheticpromoters selected from the library, Promega_SRE-2048, 8, none, minCMV (SEQ ID NO: 33) and Promega CRE promoter Promega_CRE-2002, 1, none, minCMV (SEQ ID NO: 34), are shown in Table 4.Table 1Table 2*Minimal promoter sequences not shown.Table 3Table 4*minCMV promoter sequence is underlined.

[0073] FIGs. 3A and 3B depict plots of dose-response curves measuring the ratio firefly to Renilla luciferase activity (Fluc / Rluc) at various concentrations of Compound 1 (FIG. 3A) or Compound 2 (FIG. 3B). The promoter construct used for treatment with Compound 1 was Promega_SRE-2048 (8 rotations, no spacers, minCMV), as identified described above, and had a maximal detected Fluc / Rluc ratio of about 6 for the protein GPR40. The promoter construct used for treatment with Compound 2 was Promega_CRE-2002 (1 rotation, no spacer set, minCMV), as identified described above, and had a maximal detected Fluc / Rluc ratio of about 20 for the protein GPR40. Both TRE-MPRA-identified promoters showed significantly increased response to their respective agonist when compared to the previous tested promoters.These results demonstrate that the method for selecting the promoter using the library system was validated by showing the increased signal to noise ratio. Overall, these results demonstrate the efficacy of the TRE library selection for use in the VEGAS system.

[0074] Use of an identified synthetic promoter with VEGAS. The cargo-specific synthetic promoter identified using the above protocol (Promega_CRE-2002, 1, none, minCMV) was then utilized in a VEGAS campaign. Host cells were transfected with the plasmid (expressing the Sindbis structural genome (SSG) gated by Promega-CRE-2002_minCMV activation) comprising the identified promoter and a gene encoding for at least one viral protein required for packaging. Initial packaging of virus (round 0, or RO) was completed by electroporating approximately equimolar amounts of RNA expressing the following: (1) viral replication machinery and the transgene of interest (pTSin4-GPR40), (2) capsid protein (pCapSin) and (3) glycoprotein (pGlySin). This campaign included three rounds of virus generation (Rl, R2, R3), each with decreasing amounts of applied compound in an attempt to provide selective pressure for the generation of a constitutively active mutant. Each round of harvested viral supernatant was concentrated through 20% sucrose by applying all available viral supernatant (~16 mL) to a layer of 8 mL of 20% sucrose, brought to volume with sucrose gradient buffer (50mM Tris-HCl pH 7, 100 mM NaCl, 0.5mM EDTA), and spun in an Ultracentrifuge at 30,000 RPM for 90 minutes at 4°C. Pelleted virus was resuspended in 400 pL of sucrose gradient buffer and the entire volume was used for RNA isolation using the MagMAX Viral Isolation Kit.

[0075] Isolated RNA was amplified via RT-PCR using a 26S-binding forward primer and a pool of 3 ’UTR-binding reverse primers to amplify the GPR40 transgene region. A pool of reverse primers was used to try and mitigate the potential issue of the 3’UTR mutating throughout the evolution rounds, leading to loss of primer binding sites. For each round, RT- PCR was run with and without reverse transcriptase (RTase) to account for potential DNA contamination. Using primer / probe sets that bind and amplify two distinct regions of the pTSin4 RNA in an attempt to mitigate the potential issue of evolving primer binding sites, RT-qPCR was run on the isolated RNA from each round.

[0076] RT-PCR products were cleaned up using the Qiagen PCR Clean-Up kit, then assembled into digested pcDNA3.1 backbone using NEBuilder HiFi DNA Assembly. Colony forming units (CFU) were compared between the no insert control and later rounds, and suggested the presence of amplifiable RNA material. Single colonies from each round were picked, miniprepped, and sequenced using sanger sequencing. Samples were further processed with unique molecular identifier (UMI) labels for Nanopore sequencing. Of those mutations identified, two were of interest: F 191 S and C 136R.

[0077] To test the function of the mutants, F191S GPR40 and C136R GPR40 were synthesized and run on the TRE-Luciferase assay with both SRE and CRE identified from the TRE-MPRA platform.

[0078] FIG. 4 shows a RT-qPCR plot measuring the quantification cycle over multiple rounds of a VEGAS continuous evolution campaign. RT-qPCR was performed using a 10 pL reaction with 1 pl RNA input) w / sucrose gradient pipetted viral RNA. Briefly, 8 mL 20% sucrose + 29 mL viral supernatant (~16mL virus brought up to volume with sucrose gradient buffer) were spun at 30k rpm at 4°C for 90 min, media aspirated off, the “pellet” resuspended in 400 pL sucrose gradient buffer, and the RNA from pelleted virus isolated using MagMAX Kit (400 pL input and 50 pL final resuspension volume). Using the promoter as selected described above, the VEGAS system was applied to GPR40 for continuous evolution of the protein to have greater activity in the presence of Compound 2. As shown on the leftmost data point, in Ro there was identified a large pool of Sindbis virus. Collected virus populations were used in subsequent rounds of Ri, R2, and R3. The virus populations showed diminishing virus counts over the subsequent rounds. Population samples were collected after each round and sequenced. These results show that the virus can be used in multiple rounds of VEGAS. Legend: GPR40 R0: virus collected after Round 0 (Ro); GPR40 Rl : virus collected after Ri; GPR40 R2: virus collected after R2; GPR40 R3: virus collected after R3; water: negative control.

[0079] FIG. 5 shows a summary of Sanger sequencing results of identified mutants in each of the rounds. After Round 1, 24 unique clones were identified, 22 of which comprised GPR40 mutants shown. After Round 2, 24 unique clones were identified, and 2 aligned to GPR40. After Round 3, 24 unique clones were identified, and 0 aligned to GPR40. Of those mutations identified, two were of functional interest: Mutations identified from sequencing were at amino acid positions F 191 , G238, G254, G280, and R281. Follow on sequencing identified the mutation Cl 36R. The identified mutant GPR40 F191S and C136R were further validated in follow-on experiments.

[0080] FIGs. 6A-6D depict plots of dose-response curves measuring the ratio firefly to Renilla luciferase activity (Fluc / Rluc) at various concentrations of Compound 1 or Compound 2 with either the CRE or SRE promoters. The data show the response of GPR40 wildtype compared to two identified mutants from the campaign as described above (F191S and C136R) or vehicle control. FIG. 6A shows the dose response curve of GPR40 wildtype and mutants with Compound 1 and the library identified CRE promoter. FIG. 6B shows the dose response curve of GPR40 wildtype and mutants with Compound 1 and the library identified SRE promoter. FIG. 6C shows the dose response curve of GPR40 wildtype and mutants with Compound 2 and the library identified CRE promoter. FIG. 6D shows the dose response curve of GPR40wildtype and mutants with Compound 2 and the library identified SRE promoter. These results demonstrate that the VEGAS system could be used in a continuous evolution system to mutate a protein, such as GPR40.

[0081] Overall, the data shows a workflow from using a library of promoters in a mammalian cell system to identify a promoter that demonstrates a high signal to noise response for a given protein of interest and then using said promoter in a VEGAS continuous evolution campaign to mutate said protein.Example 2: Adenovirus-based directed evolution using a TRE Library

[0082] A directed evolution campaign is performed following the general methods described in Berman et al. (J Am Chem Soc. 2018 December 26; 140(51): 18093-18103), with modification: a TRE library described herein is first evaluated to identify a promoter that is driven by a target property of a transgene. For directed evolution, the identified promoter is then used to drive expression of adenoviral protease (AdProt).

[0083] The examples described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

WHAT IS CLAIMED IS:

1. A method for continuous evolution of a transgene, comprising:(a) selecting a polynucleotide from a library of polynucleotides;(b) culturing a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle and wherein expression of the at least one viral gene is under control of the synthetic promoter from the library of polynucleotides; and(c) infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene.

2. The method of claim 1, wherein library of polynucleotides comprises at least 50 synthetic promoters.

3. The method of claim 1, wherein each of the polynucleotides from the library of polynucleotides are 100-300 bases in length.

4. The method of claim 1, wherein each of the polynucleotides from the library of polynucleotides are at most 300 bases in length.

5. The method of claim 1, wherein each of the polynucleotides from the library of polynucleotides are at least 300 bases in length.

6. The method of claim 1, wherein each of the polynucleotides from the library of polynucleotides comprises a helical turn.

7. The method of any one of claims 1-6, wherein each of the polynucleotides from the library of polynucleotides independently comprise one or more copies of a transcription factor binding motif (TFBM) comprising a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site.

8. The method of any one of claims 1-7, wherein each of the one or more copies of TFBM are separated by one or more nucleotides.

9. The method of any one of claims 1-8, wherein each of the polynucleotides from the library of polynucleotides comprises at least four TFBMs.

10. The method of any one of claims 1-8, wherein each of the polynucleotides from the library of polynucleotides comprises four TFBMs.

11. The method of any one of claims 1-10, wherein each of the polynucleotides from the library of polynucleotides independently comprise a spacer before, between, and / or after the TFBMs.

12. The method of any one of claims 1-11, wherein the spacer comprises two random sequence spacer sets.

13. The method of any one of claims 1-12, wherein each of the polynucleotides from the library of polynucleotides comprises at least 3 spacers.

14. The method of any one of claims 1-13, wherein each of the polynucleotides from the library of polynucleotides independently comprise a minimal promoter, a transcription start site, and / or a DNA sequence capable of initiating transcription.

15. The method of any one of claims 1-14, wherein the TFBM is at least 3 helix positions relative to the minimal promoter.

16. The method of any one of claims 1-15, wherein each of the polynucleotides from the library of polynucleotides independently comprises at least 3 minimal promoters.

17. The method of any one of claims 1-16, wherein each of the polynucleotides from the library of polynucleotides independently comprise a 5 ’ UTR comprising a detectable nucleotide barcode.

18. The method of any one of claims 1-17, wherein each of the polynucleotides from the library of polynucleotides independently comprise a nucleotide sequence encoding a detectable polypeptide.

19. The method of any one of claims 1-18, wherein each of the polynucleotides from the library of polynucleotides independently comprise a 3 ’ UTR comprising a detectable nucleotide barcode.

20. The method of any one of claims 1-19, wherein each of the polynucleotides from the library of polynucleotides independently comprises in 5’ to 3’ direction:(a) one or more copies of transcription factor binding motif (TFBM) comprising a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site;(b) a minimal promoter, a transcription start site, and / or a DNA sequence capable of initiating transcription;(c) a 5’ UTR comprising a detectable nucleotide barcode;(d) a nucleotide sequence encoding a detectable polypeptide; and(e) a 3 ’ UTR comprising a detectable nucleotide barcode.

21. The method of any one of claims 1 -20, wherein each of the polynucleotides from the library of polynucleotides independently comprises a mammalian transcription response element (TRE).

22. The method of any one of claims 1-21, wherein each of the TREs independently comprises one or more copies of transcription factor binding motif (TFBM) comprising a DNA sequence wherein a polynucleotide, protein, or other agent of inducing transcription can be bind upstream of a transcriptional start site.

23. The method of any one of claims 1-22, wherein each of the TREs independently comprises a spacer before, between, and / or after the TFBMs.

24. The method of any one of claims 1-23, wherein each of the polynucleotides from the library of polynucleotides independently comprises a synthetic promoter.

25. The method of any one of claims 1-24, wherein each of the synthetic promoters independently comprises the TRE and a minimal promoter.

26. The method of any one of claims 1-25, wherein the synthetic promoter is not constitutively active or inactive in the host cell or the virus.

27. The method of any one of claims 1-26, wherein each of the TFBM of the one or more copies of TFBM is comprises a different sequence from the other TFBMs.

28. The method of any one of claims 1-27, wherein the library does not comprise a synthetic promoter activated by viral infection.

29. The method of any one of claims 1-28, wherein the 5’ UTR nucleotide barcode or 3’ UTR nucleotide barcode are unique to the composition of the synthetic promoter.

30. The method of any one of claims 1-29, wherein the TRE comprises a nuclear receptor response element, a type 1 nuclear receptor response element, a type 2 nuclear receptor response element, a nuclear factor of activated T-cells response element (NFAT-RE), a cAMP response element (CRE), a B recognition element, a hypoxia-responsive element, a serum response element (SRE), a serum response factor response element (SRF-RE), a metal-responsive element (MRE), a IFN-stimulated response element (ISRE), a calcium- response element, an antioxidant response element (ARE), a p53 response element, a sterol regulatory element, a polycomb response element (PRE), a rev response element (RRE), and / or a wnt response element (WRE).

31. The method of any one of claims 1-30, wherein the selection comprises:(a) transfecting a population of host cells with the library of polynucleotides;(b) infecting the population of host cells with a population of viruses, wherein the population of viruses comprise a transgene;(c) treating the host cells with an agent capable of activating the transgene;(d) harvesting the host cells;(e) isolating RNA from the harvested host cells;(f) sequencing the isolated RNA;(g) quantifying the RNA and identifying the synthetic promoter based on the RNA sequencing; and(h) selecting the synthetic promoter, wherein the synthetic promoter has a higher number of sequence reads compared to a control.

32. The method of any one of claims 1-31, wherein the RNA comprises a unique detectable nucleotide barcode.

33. The method of any one of claims 1-32, wherein the RNA comprises a reporter.

34. The method of any one of claims 1-33, wherein the synthetic promoter is not constitutively active or inactive in the absence of the agent.

35. The method of any one of claims 1-34, wherein the method comprises:(d) incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene.

36. The method of any one of claims 1-35, wherein the method comprises:(e) generating a second population of viruses expressed by the first population of the host cells.

37. The method of any one of claims 1-36, wherein the method comprises:(f) infecting a second population of host cells with the second population of viruses.

38. The method of any one of claims 1-37, wherein steps (a)-(f) are repeated at least one, two, or five times.

39. The method of any one of claims 1-37, wherein steps (a)-(f) are repeated at least once.

40. The method of any one of claims 1-37, wherein steps (a)-(f) are repeated at least two times.

41. The method of any one of claims 1-37, wherein steps (a)-(f) are repeated at least five times.

42. The method of any one of claims 1-41, wherein the method comprises:(g) amplifying RNA from the second population of viruses;(h) sequencing the RNA; and(i) quantifying the RNA and / or barcode counts.

43. The method of any one of claims 1-42, wherein the transgene is about 500, 1000, or 2000 nucleotides in length.

44. The method of any one of claims 1-42, wherein the transgene is at least about 500 nucleotides in length.

45. The method of any one of claims 1-42, wherein the transgene is at least about 1000 nucleotides in length.

46. The method of any one of claims 1-42, wherein the transgene is at least about 2000 nucleotides in length.

47. The method of any one of claims 1-46, wherein the second population of viruses comprise at least one mutated transgene of interest.

48. The method of any one of claims 1-47, wherein the repeating of steps (a)-(f) results in at least one mutation in the transgene of interest.

49. The method of any one of claims 1-48, wherein the mutated transgene of interest comprises at least one nucleic acid substitution, insertion, or deletion compared to the transgene.

50. The method of any one of claims 1-49, wherein the mutated transgene of interest comprises at least two nucleic acid substitutions, insertions, or deletions compared to the transgene.

51. The method of any one of claims 1 -50, wherein the mutated transgene of interest comprises at most about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene.

52. The method of any one of claims 1-50, wherein the mutated transgene of interest comprises about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene.

53. The method of any one of claims 1-50, wherein the mutated transgene of interest comprises at least about 99% sequence identity to the transgene.

54. The method of any one of claims 1-53, wherein the mutated transgene of interest encodes for a mutated protein comprising at least one amino acid substitution, insertion, or deletion compared to a protein encoded by the transgene.

55. The method of any one of claims 1-54, wherein the mutated transgene of interest comprises same, increased, or decreased activity of the transgene.

56. The method of any one of claims 1-55, wherein the mutated transgene of interest comprises increased activity of the transgene.

57. The method of any one of claims 1-55, wherein the mutated transgene of interest comprises decreased activity of the transgene.

58. The method of any one of claims 1-55, wherein the mutated transgene of interest comprises at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene.

59. The method of any one of claims 1-58, wherein the mutated transgene of interest comprising improved activity results in a higher number of infectious viral particles.

60. The method of any one of claims 1-59, wherein the mutated transgene of interest comprising improved activity has a higher RNA count than the transgene or other mutated transgenes.

61. The method of any one of claims 1 -60, wherein a ratio of quantified RNA to DNA of the transgene is at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8,9, or 10.

62. The method of any one of claims 1-60, wherein a ratio of quantified RNA to DNA of the transgene is at most about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

63. The method of any one of claims 1-60, wherein a ratio of quantified RNA to DNA of the transgene is about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or10.

64. The method of any one of claims 1-63, wherein the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal.

65. The method of any one of claims 1-64, wherein the agent activates the transgene.

66. The method of any one of claims 1-65, wherein the activity of the transgene is a signaling event.

67. The method of any one of claims 1-66, wherein the activity of the transgene produces an activity-coupled transcriptional response.

68. The method of any one of claims 1-67, wherein the activity-coupled transcriptional response binds to the synthetic promoter.

69. The method of any one of claims 1-68, wherein binding of the activity-coupled transcriptional response with the synthetic promoter results in transcription of the at least one viral gene for packaging a virus into an infectious viral particle.

70. The method of any one of claims 1-69, wherein the transgene encodes for a protein.

71. The method of any one of claims 1-70, wherein the protein is activated by the agent and produces an activity-coupled transcriptional response wherein the activity-coupled transcriptional response is an endogenous transgene -mediated signaling pathway.

72. The method of any one of claims 1-71, wherein the protein is a G protein-coupled receptor (GPCR).

73. The method of any one of claims 1-72, wherein the agent activates the GPCR and produces a GPCR-mediated signaling response.

74. The method of any one of claims 1-73, wherein the GPCR-mediated signaling response binds to SRE or CRE.

75. The method of any one of claims 1-74, wherein the virus is a mutagenic virus.

76. The method of any one of claims 1-75, wherein the virus has a mutation rate of about 10" 5- 1 O’3mutations per base replicated.

77. The method of any one of claims 1-75, wherein the virus has a mutation rate of about 10" 3 mutations per base replicated.

78. The method of any one of claims 1-75, wherein the virus has a mutation rate of about 10" 5 mutations per base replicated.

79. The method of any one of claims 1-78, wherein the virus comprises an error prone replicase.

80. The method of any one of claims 1-79, wherein the virus is a DNA virus, an RNA virus, or a reverse transcribing virus.

81. The method of any one of claims 1-80, wherein the virus is an RNA virus.

82. The method of claim 81, wherein the RNA virus is a positive-strand RNA virus.

83. The method of any one of claims 1-82, wherein the virus is a Sindbis virus.

84. The method of any one of claims 1-83, wherein the host cells are mammalian cells.

85. The method of any one of claims 1-84, wherein the mammalian cells are BHK-21 cells, HEK-293 cells, HEK293T cells, or CHO cells.

86. A system configured to perform the steps of any one of claims 1-85.

87. A composition comprising: a library of polynucleotides, wherein the library comprises a plurality of synthetic promoters; and a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle and wherein expression of the at least one viral gene is under control of at least one synthetic promoter from the library of polynucleotides.

88. The composition of claim 87, wherein the composition further comprises a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene.