Unique molecular index sequencing for genetic mutations
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2026-04-15
AI Technical Summary
Current methods for directed evolution, such as those used in nucleic acid and protein sequence improvement, face inefficiencies, high costs, and limited applicability, necessitating the development of more effective approaches for generating and analyzing genetic mutations.
The method involves continuous evolution of a transgene using a viral system where host cells are infected with viruses carrying a transgene and viral genes, with the transgene's activity triggering viral packaging, and unique molecular identifiers (UMIs) are used to label and sequence oligonucleotides, allowing for the detection of mutational frequencies and identification of mutated transgenes with improved activity.
This approach enables the rapid and accurate generation and analysis of mutated transgenes, enhancing the efficiency of directed evolution by identifying variants with increased or decreased activity, thereby improving viral particle production and gene function.
Smart Images

Figure IMGF000020_0001 
Figure IMGF000024_0001 
Figure IMGF000024_0002
Abstract
Description
UNIQUE MOLECULAR INDEX SEQUENCING FOR GENETIC MUTATIONSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 506,505, filed on June 6, 2023, which is incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under GM146247 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND
[0003] Improving the properties and function of nucleic acid and protein sequences is applicable to fields such as medicine, diagnostics, and biotechnology. However, many approaches (e.g., directed evolution) may suffer from low efficiency, high cost, or limited applicability. Therefore, there is a need for improved methods of directed evolution.BRIEF SUMMARY
[0004] Provided herein are methods for continuous evolution of a transgene, compnsing: culturing a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle, wherein expression of the at least one viral gene is under control of a promoter; infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene, and wherein packaging the virus into an infectious viral particle results in a second population of viruses; and labeling oligonucleotides of the second population of viruses expressed by the first population of host cells with at least one barcode. In some embodiments, the oligonucleotides comprise RNA, DNA, or cDNA. In some embodiments, the at least one barcodes are uniquely identifiable. In some embodiments, the at least one barcodes comprises at least one unique molecular identifier (UMI). In some embodiments, the labeling oligonucleotides comprises labeling with at least about IxlO3, IxlO4, IxlO5, IxlO6, 3x106, IxlO7, or IxlO8barcodes. In some embodiments, each barcode comprises at least 24 nucleotides. In some embodiments, each barcode comprises at most 24 nucleotides. In some embodiments,each barcode comprises about 24 nucleotides. In some embodiments, each barcode comprise at least three YR nucleotide repeats. In some embodiments, each barcode comprise each of the three YR nucleotide repeats are separated by at least three nucleotides. In some embodiments, each barcode is restricted to at least about 10% Y nucleotides. In some embodiments, each barcode is restricted to at least about 10% R nucleotides. In some embodiments, each barcode comprises the oligonucleotide sequence NNNYRNNNYRNNNYRNNN. In some embodiments, each barcode comprises the oligonucleotide sequence VHBDVHBDVHBDBDHVBDHVBDHV (SEQ ID NO: 16). In some embodiments, each barcode comprises a Hamming distance of at least 1, 2, 3, 4, or 5 relative to other barcodes. In some embodiments, each barcode comprises a Hamming distance of at least 2 relative other UMIs. In some embodiments, each barcode comprises a Levenshtein distance of at least 1, 2, 3, 4, or 5 relative to other barcodes. In some embodiments, each barcode comprises a Levenshtein distance of at least 2 relative to other barcodes. In some embodiments, the RNA of the second population of viruses is attached with the barcode. In some embodiments, the RNA of the second population of viruses is amplified. In some embodiments, the amplified RNA is sequenced to generate a plurality of reads. In some embodiments, the plurality of reads are organized to distinguish between amplification errors and single nucleotide polymorphisms present in the amplified RNA. In some embodiments, the RNA of the second population of viruses is reverse transcribed into a library of cDNA. In some embodiments, the library of cDNA is attached with the barcode. In some embodiments, the library of cDNA is amplified. In some embodiments, the amplified library of cDNA is sequenced to generate a plurality of reads. In some embodiments, the plurality of reads are organized to distinguish between amplification errors and single nucleotide polymorphisms present in the amplified library of cDNA. In some embodiments, the barcode is attached by primer-based addition, PCR-based addition, ligation-based addition, fragmentation and end repair, or adapter ligation. In some embodiments, the method detects mutational frequencies of at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%. In some embodiments, the method detects mutational frequencies of at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%. In some embodiments, the method detects mutational frequencies of about 7%. In some embodiments, the sequencing is third-generation sequencing, next generation sequencing, shortread sequencing, long-read sequencing, single molecule real time sequencing (SMRT), or nanopore sequencing. In some embodiments, the sequencing is long-read sequencing. In some embodiments, the method comprises incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene. In someembodiments, the method comprises: incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene. In some embodiments, the method comprises generating a second population of viruses expressed by the first population of the host cells. In some embodiments, the method comprises infecting a second population of host cells with the second population of viruses. In some embodiments, steps (a)-(b) are repeated at least once, twice, three times, four times, five times, six times, seven times, eight times, nine times, or ten times. In some embodiments, steps (a)-(b) are repeated at least once. In some embodiments, steps (a)-(b) are repeated at least five times. In some embodiments, steps (a)-(b) are repeated at least six times. In some embodiments, the method comprises: amplifying RNA from the second population of viruses; sequencing the RNA; and quantifying the RNA and barcode counts. In some embodiments, the transgene is about 500, 1000, 2000 nucleotides in length. In some embodiments, the transgene is at least about 500 nucleotides in length. In some embodiments, the transgene is at least about 1000 nucleotides in length. In some embodiments, the transgene is at least about 2000 nucleotides in length. In some embodiments, the transgene is at least about 500 nucleotides in length. In some embodiments, the transgene is at least about 1000 nucleotides in length. In some embodiments, the transgene is at least about 2000 nucleotides in length. In some embodiments, the transgene is about 500 nucleotides in length. In some embodiments, the transgene is about 1000 nucleotides in length. In some embodiments, the transgene is about 2000 nucleotides in length. In some embodiments, the second population of viruses comprise at least one mutated transgene of interest. In some embodiments, the repeating of steps (a)-(b) results in at least one mutation in the transgene of interest. In some embodiments, the mutated transgene of interest comprises at least one nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene of interest comprises at least two nucleic acid substitutions, insertions, or deletions compared to the transgene. In some embodiments, the mutated transgene of interest comprises at most about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene. In some embodiments, the mutated transgene of interest comprises about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene. In some embodiments, the mutated transgene of interest comprises at least about 99% sequence identity to the transgene. In some embodiments, the mutated transgene of interest encodes for a mutated protein comprising at least one amino acid substitution, insertion, or deletion compared to a protein encoded by the transgene. In some embodiments, the mutated transgene of interest comprises same, increased, or decreased activity of the transgene. In someembodiments, the mutated transgene of interest comprises increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene. In some embodiments, the mutated trans gene of interest comprises at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises about 1%, 2%, 3%, 5%. 10%. 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene. In some embodiments, the mutated trans gene of interest comprises at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprising better activity results in a higher number of infectious viral particles. In some embodiments, the mutated transgene of interest comprising better activity has a higher RNA count than the transgene or other mutated transgenes. In some embodiments, a ratio of quantified RNA to DNA of the transgene is at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, a ratio of RNA to DNA of the transgene of the first population of host cells is at most about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, a ratio of RNA to DNA of the transgene of the first population of host cells is about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9. or 10. In some embodiments, the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal. In some embodiments, the agent activates the transgene. In some embodiments, the activity of the transgene is a signaling event. In some embodiments, the activity of the transgene produces an activity-coupled transcriptional response. In some embodiments, the activity-coupled transcriptional response binds to the synthetic promoter. In some embodiments, binding of the activity-coupled transcriptional response with the synthetic promoter results in transcription of the at least one viral gene for packaging a virus into an infectious viral particle. In some embodiments, the transgene encodes for a protein. In some embodiments, the protein is activated by the agent and produces an activity-coupled transcriptional response wherein the activity- coupled transcriptional response is an endogenous transgene-mediated signaling pathway. In some embodiments, the protein is a G protein-coupled receptor (GPCR). In some embodiments, the agent activates the GPCR and produces a GPCR-mediated signaling response. In some embodiments, the GPCR-mediated signaling response binds to the SRE or the CRE. In someembodiments, the virus is a mutagenic virus. In some embodiments, the virus has a mutation rate of about 10'5-l O'3mutations per base replicated. In some embodiments, the virus has a mutation rate of about IO3mutations per base replicated. In some embodiments, the virus has a mutation rate of about 10'5mutations per base replicated. In some embodiments, the virus has a mutation rate of at least about 10'3mutations per base replicated. In some embodiments, the vims has a mutation rate of at most about 10'3mutations per base replicated. In some embodiments, the vims comprises an error prone replicase. In some embodiments, the virus is a DNA vims, an RNA vims, or a reverse transcribing vims. In some embodiments, the virus is an RNA vims. In some embodiments, the RNA vims is a positive-strand RNA vims. In some embodiments, the vims is a Sindbis virus.
[0005] Provided herein are systems for generating variant proteins through continuous directed evolution of a transgene, the system comprising: a system configured to perform the steps of any one of the previous embodiments. Provided herein are one or more of the peptides, nucleic acid sequences, compositions, methods, or libraries described herein.INCORPORATION BY REFERENCE
[0006] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0008] FIGs. 1A-1B depict plots of dose-response curves measuring the response of GPR40 at various concentrations of Compound 1 (FIG. 1A) or Compound 2 (FIG. IB) using a promoter selected from the library of promoters. The FIG. 1A x-axis is labeled Log[Compound 1] starting at “ND” and labeled from -8 to -4 at 2 unit intervals; the y-axis is labeled Fluc / RLuc from 0 to 8 at 2 unit intervals. Legend: GPR40+pGL4.33R (light grey shaded circles); Empty Vector+pGL4.33R (unshaded light grey circles); GPR40+SRE-minCMV (shaded black circles); and Empty Vector+SRE-minCMV (unshaded black circles). The FIG. IB x-axis is labeled Log[Compound 2] starting at “ND” and labeled from -11 to -6 at 1 unit intervals; the y-axis islabeled Fluc / RLuc from 0 to 25 at 5 unit intervals. Legend: GPR40+pGL4.29R (black shaded circles); Empty Vector (EV) +pGL4.29R (unshaded black circles); GPR40+CRE-minCMV (shaded grey triangles); and Empty Vector (EV)+CRE-minCMV (unshaded triangles).
[0009] FIG. 2 depicts a RT-qPCR plot measuring the quantification cycle over multiple rounds of a VEGAS continuous evolution campaign. Experimental conditions are shown on the x-axis; the y-axis is labeled cq (quantification cycles) from 50 to 0 at 10 unit intervals. Legend: RT (modAmbrose, shaded circles); no RT (modAmbrose, unshaded circles); RT (modVEGAS, shaded squares); and no RT (modVEGAS, unshaded squares). Evolution of the tTA (tetracycline-controlled transactivator) was used as a positive control.
[0010] FIG. 3 shows a summary of Sanger sequencing results of identified mutants in each of the rounds aligned with the GPR40 construct. Round 1 (24 clones): 22 / 24 align to GPR40, 1 / 24 failed); Round 2 (24 clones): 2 / 24 align to GPR40, 14 / 24 align to empty backbone; 8 / 24 failed); Round 3 (14 clones): 2 / 12 align to empty backbone; 10 / 12 failed).
[0011] FIG. 4A shows a polyacrylamide gel electrophoresis (PAGE) gel of PCR amplified and UMI tagged samples from the virus populations. Lanes from left to right show a ladder, initial virus population, Ro vims population, Ri virus population, R2 virus population, R3 virus population, and library of pooled virus for sequencing. FIG. 4B depicts a bar plot of nanopore sequencing results with UMIs. Sequencing results for different rounds of evolution are shown on the x-axis, from left to right: IVT (control), Round 0, Round 1, Round 2, and Round 3. The y- axis is labeled # consensus seq from 0 to 2000 at 1000 unit intervals. The shaded portions of teach bar represent specific regions sequenced, including GPR40, mScarlet, helper, and undetermined.
[0012] FIGs. 5A-5D depict plots of dose-response curves measuring the response of wildtype of VEGAS mutated GPR40 protein at various concentrations of Compound 1 or Compound 2 with the promoters selected from the library of promoters. FIG. 5A depicts GPR40+Compound 1+SRE; the x-axis is labeled Log[Compound 1] starting at “ND” and labeled from -9 to -4 at 1 unit intervals; the y-axis is labeled Fluc / RLuc from 0 to 10 at 2 unit intervals. FIG. 5B depicts GPR40+Compound 1+CRE; the x-axis is labeled LogfCompound 1] starting at “ND” and labeled from -9 to -4 at 1 unit intervals: the y-axis is labeled Fluc / RLuc from 0.0 to 2.5 at 0.5 unit intervals. FIG. 5C depicts GPR40+Compound 2+SRE; the x-axis is labeled Log[Compound 2] starting at “ND” and labeled from -11 to -7 at 1 unit intervals; the y- axis is labeled Fluc / RLuc from 0 to 15 at 5 unit intervals. FIG. 5D depicts GPR40+Compound 2+CRE; the x-axis is labeled LogfCompound 2] starting at “ND” and labeled from -11 to -6 at 1unit intervals; the y-axis is labeled Fluc / RLuc from 0 to 40 at 10 unit intervals. Legend: WT (black shaded circles); F191S (grey shaded squares); C136R (black shaded triangles); EV (empty vector, grey shaded circles).DETAILED DESCRIPTION OF THE INVENTION
[0013] Provided herein are directed evolution platforms. The VEGAS directed evolution platform (English et al., Cell 2019 Jul 25; 178(3):748-761.el 7) delivers transgenes to cells via viral infection and evolves these genes through random mutagenesis to produce mutant versions of the transgene. In the VEGAS system, a viral packaging gene is coupled to the activity of said transgene via promoter comprising a transcription response element (TRE). This efficiently couples the signaling pathways and other biological processes upstream of or directly in line with transcription. Mutagenesis of the transgene may result in more or less activity of the transgene, thus resulting in more or less expression of the viral packaging gene. If the transgene accumulates additional advantageous mutations, more virus comprising said mutagemzed transgene will be packaged. Transcriptional activation can be encoded on synthetic, exogenous, or endogenous DNA promoters that ultimately induce the host cell to produce RNA that encodes the Sindbis virus nucleocapsid and glycosylation coat proteins (together, Sindbis Structural Genome).
[0014] Sindbis vims is a single positive stranded RNA alphavirus that replicates in human cells with a high error rate. It is used in the context of directed evolution and directed molecular evolution to generate gene variants with selected functions. Isolating and accurately sequencing the many variants generated within a population of viral genomes is the explicit purpose of this patent. The general approach involves using specific degenerate DNA primers targeting conserved regions of the Sindbis virus genome, or chemically conjugated nucleotide sequences to the Sindbis genome, to generate barcoded amplicons of the Sindbis genome. These amplicons are then repeatedly amplified using a “landing site” primer to generate many copies of the barcoded amplicons. Using the methods provided herein, the pool of sequences is analyzed via Nanopore sequencing and aligned using software that pairs the barcodes together and generates a consensus sequence of the data contained between the two barcode sequences for each measured DNA molecule. The final consensus produces an essentially error-free complete identity of a single genome of the Sindbis virus. This is performed at scale to generate thousands of consensus sequences to evaluate the single molecule sequence composition of a pool of Sindbis viruses simultaneously.Definitions
[0015] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art to which these inventions belong.
[0016] Throughout this disclosure, numerical features are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiments. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range to the tenth of the unit of the low er limit unless the context clearly dictates otherwise. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual values within that range, for example, 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention, unless the context clearly dictates otherwise.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of any embodiment. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0018] Unless specifically stated or obvious from context, as used herein, the term “about” in reference to a number or range of numbers is understood to mean the stated number and numbers + / - 10% thereof, or 10% below the lower listed limit and 10% above the higher listed limit for the values listed for a range.Directed Evolution Platforms
[0019] Provided herein are directed evolution platforms comprising systems and methods. In some instances, a directed evolution platform provided herein comprises a viral replication system inside of a transfected cell. Viral replication systems in some embodiments comprise a transgene (e.g., the target sequence to be evolved), a synthetic promoter driven by the activity or presence of the transgene (or transgene product, such as a protein) which induces expression of a viral replication protein. Cells producing transgene or transgene products with the selected properties activate the viral replication protein, leading to increased infectivity.
[0020] Provided herein are methods for continuous evolution of a transgene, compnsing selecting an optimal promoter, transfecting mammalian host cells with the optimal promoter and at least one viral gene required for viral packaging, and infecting the host cells with virus comprising a transgene. In some embodiments, virus expressed by the infected host cells comprises a mutated transgene. In some embodiments, the method further comprising infecting a new population of host cells with virus expressed by the infected host cells. In some embodiments, the method is repeated for multiple cycles of infecting host cells and collecting the resulting expressed virus comprising uniquely mutated transgenes to infect a new population of host cells.
[0021] In some embodiments, the agent activating the transgene is coupled to an endogenous activity response transcriptional element in order for the agent to activate the promoter. In some embodiments, the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal. In some embodiments, the agent activates the transgene. In some embodiments, the activity of the transgene is a signaling event. In some embodiments, the activity of the transgene produces an activity-coupled transcriptional response. In some embodiments, the activity-coupled transcriptional response binds to the promoter. In some embodiments, binding of the activity- coupled transcriptional response with the synthetic promoter results in transcription of the at least one viral gene for packaging a virus into an infectious viral particle. Exemplary agents include but are not limited to: GPR40 agonists, Fetal bovine serum, Forskohn, ATP / GTP, AICAR, Dexamethasone, CdCh, ZnSO4, LiCE, Deferoxamine, Thapsigargin, Serotonin, Dopamine, Thrombin, Epinephrine, cis-epoxysuccinate, (R)-zn3573.
[0022] The transgene is a gene that is not endogenously expressed by the virus. In some embodiments, the transgene comprises DNA. In some embodiments, the transgene encode for a protein. In some embodiments, the protein is expressed in order to be activated by the agent. In some embodiments, the protein is activated by the agent and produces an activity -coupledtranscriptional response wherein the activity -coupled transcriptional response is an endogenous transgene-mediated signaling pathway. In some embodiments, the protein is a G protein-coupled receptor (GPCR). In some embodiments, the protein is GPR40. In some embodiments, the agent activates the GPCR and produces a GPCR-mediated signaling response. In some embodiments, the GPCR-mediated signaling response binds to SRE or CRE. In some instances the transgene comprises an antibody or antigen binding fragment, such as a monospecific Fab2, bispecific Fab2, trispecific Fab3, monovalent IgG, scFv, bispecific diabody, trispecific triabody, scFv-Fc, sdAb, minibody, nanobody, IgNAR, V-NAR, hdgG, VHH, or peptibody. Transgenes in some instances also comprise transcription factors.
[0023] Through subsequent round of the continuous evolution method, the transgene accumulates random mutations caused by errors in the mutagenic virus replication process. Therefore, the virus should be a mutagenic virus in order to randomly generate the mutations. In some embodiments, the virus is a mutagenic virus. In some embodiments, the virus has a mutation rate of about 10'9-10 , 10'8-10’2, 10'7-10'3, 10'6-10'3, 10'5-10'3, 10'5-104, or 10’4-10’3mutations per base replicated. In some embodiments, the virus has a mutation rate of at least about 10'9-104, 10'8-10'2, 10'7-10'3, 10'6-10'3, 1O'5-1O'3, 10'5-10'4, or 10'4-10'3mutations per base replicated. In some embodiments, the virus comprises an error prone replicase. In some embodiments, the viral mutations are caused by mutagenic nucleoside analogs. In some embodiments, the virus is a DNA virus, an RNA virus, or a reverse transcribing virus. In some embodiments, the virus is an RNA virus. In some embodiments, the virus is a Sindbis virus.
[0024] In some embodiments, the host cells are mammalian cells. In some embodiments, the mammalian cells are BHK-21 cells, HEK-293 cells, HEK293T cells, or CHO cells.
[0025] Through subsequent round of the continuous evolution method, the transgene accumulates random mutations caused by errors in the mutagenic virus replication process. In some embodiments, the method further comprising sequencing virus expressed by the infected host cells. In some embodiments, the method comprises identifying mutations in the mutated transgene. In some embodiments, the mutated transgene comprises at least one, two, three, four, five, six, seven, eight, nine, or ten nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene comprises at least one nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene comprises one nucleic acid substitution, insertion, or deletion compared to the transgene. In some embodiments, the mutated transgene of interest comprises at most about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 80%,75%, 70%, 65%, 60%, 55%, 50%, 45%, or 40% sequence identity to the transgene. In some embodiments, the mutated transgene of interest comprises at least about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, or 40% sequence identity to the transgene. In some embodiments, the mutated transgene of interest comprises about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, or 40% sequence identity to the transgene. In some embodiments, the mutated transgene of interest encodes for a mutated protein comprising at least one, two, three, four, five, six, seven, eight, nine, or ten amino acid substitution, insertion, or deletion compared to a protein encoded by the transgene.
[0026] Mutations in the transgene are randomly generated by the mutagenic virus, but may end up having an effect on the activity of the transgene. In some embodiments, the mutated transgene of interest comprises same, increased, or decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises decreased activity of the transgene. In some embodiments, the mutated transgene of interest comprises about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene. In some embodiments, the mutated trans gene of interest comprises at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene. In some embodiments, the mutated transgene of interest comprises about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene. In some embodiments, the mutated trans gene of interest comprises at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene.
[0027] As the activity of the transgene improves, there is an increased activity coupled transcription response. The increase in activity coupled transcription response results in more expression of the viral gene required for viral packaging controlled by the promoter activated by the activity coupled transcription response. Thus, as activity of the transgene improves, there is more viral packaging. This evolutionary pressure results in more virus comprising the mutated transgene with greater activity. This method can also be used in the opposite direction to apply selective pressure for less activity in a transgene. In some embodiments, the mutated transgene of interest comprising improved activity results in a higher number of infectious viral particles. In some embodiments, the mutated transgene of interest comprising improved activity has a higher RNA count than the transgene or other mutated transgenes. In some embodiments, theratio of quantified RNA to DNA of the transgene is at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. In some embodiments, the ratio of RNA to DNA of the transgene of the first population of host cells is at most about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. In some embodiments, the ratio of RNA to DNA of the transgene of the first population of host cells is about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12. 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25.
[0028] In some embodiments, the promoter comprises at least one transcription factor binding motif (TFBM), at least one space, and at least one minimal promoter. In some embodiments, the synthetic promoter comprises a transcriptional response element (TRE) and at least one minimal promoter. In some embodiments, each of the polynucleotides independently comprises a synthetic promoter. In some embodiments, each of the synthetic promoters independently comprises the TRE and a minimal promoter. In some embodiments, each of the TFBM of the one or more copies of TFBM is comprises a different sequence from the other TFBMs. In some embodiments, the 5’ UTR nucleotide barcode or 3’ UTR nucleotide barcode are unique to the composition of the synthetic promoter. In some embodiments, the synthetic promoter is not constitutively active or inactive in the host cell or the vims. In some embodiments, the library does not comprise a synthetic promoter activated by viral infection.
[0029] A reporter system may be optimized to minimize background signal (e.g., in the absence of an agent or activated transgene protein). In some instances, a reporter system generates an observable signal that is at least 10%, 25%, 50%, 100%, 200%, 500%, 10 times, 20 times, 50 times, or at least 100 times higher in the presence of an agent or activated transgene protein. In some instances, a reporter system generates an observable signal that is at least 10%, 25%, 50%, 100%, 200%, 500%, 10 times, 20 times, 50 times, or at least 100 times higher in the presence of an agent or activated transgene protein when compared to an empty vector. In some instances, a reporter system generates an observable signal that is at least 10%- 1000%, 25- 1000%, 50%-1000%, 100%-100%, 200%-1000%, 500%-100%, 5-10 times, 5-20 times, or 10-50 times higher in the presence of an agent or activated transgene protein when compared to an empty vector.
[0030] For evolution of the transgene, mutagenesis may be used to evolve new transgene variants. In some instances, viruses comprising a high mutation rate are used to induce mutations in the transgene. In some instances, the transgene encodes for a protein. In someinstances, the protein is activated by the agent and produces an activity-coupled transcriptional response wherein the activity-coupled transcriptional response is an endogenous transgene- mediated signaling pathway. In some instances, the protein is a membrane-bound protein. In some instances, the protein is a G protein-coupled receptor (GPCR). In some instances. In some instances, the agent activates a GPCR and produces a GPCR-mediated signaling response. In some instances, the GPCR-mediated signaling response binds to an SRE (serum-response element) or CRE (cAMP -response element).
[0031] Agents
[0032] In some instances, an agent is added to activate an endogenous pathway in the cell which activates the synthetic promoter from a promotor 1 ibrary. In some instances, an agent is added to activate the same endogenous pathway in the cell as the transgene. For example, an agent that is an agonist for a transgene protein is added. In some instances, an agent mimics the activated state of a transgene product. As the amount of the agent is decreased, viral replication will only occur if the transgene itself has evolved into a more active state. In some instances, the agent inhibits activation of a synthetic promoter from a promotor library. The agent in some instances is titrated to control the selection pressure on a viral system. The agent in some instances is titrated to increase the selection pressure on a viral system to evolve more active transgenes. In some instances, the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal. In some instances, the agent activates the transgene. In some instances, the activity of the transgene is a signaling event. In some instances, the activity of the transgene produces an activity-coupled transcriptional response. In some instances, the activity-coupled transcriptional response binds to the synthetic promoter. In some instances, binding of the activity-coupled transcriptional response with the synthetic promoter results in transcription of the at least one viral gene for packaging a virus into an infectious viral particle.
[0033] In some embodiments, the agent comprises Compound 1 or Compound 2. In some embodiments, Compound 1 comprises fasiglifam (TAK-875, PubChemID: 24857286, CAS No.: 1000413-72-8). Fasiglifam is a potent, selective and orally bioavailable GPR40 agonist with EC50 of 72 nM. In some embodiments, Compound 2 comprises “Compound A” (see: Rives et al. Mol Pharmacol, (2018), 93(6), 581-591, PubChemID: 134465785).
[0034] Methods and evolution systems provided herein may provide for rapid generation of mutations in the transgene. In some instances, a system provided herein has a mutation rate of at least 10'5, 10'4, 10'3, 10'2, or at least IO'1mutations per base per hour. In some instances, asystem provided herein has a mutation rate of about 10’5-10 , 10'5-10'2, 10'4-10 , 10’4-10’2, 10'3- I0’1, 1 O’3- 1 O’2, or about lO’ O’1mutations per base per hour. In some instances, a system provided herein has a mutation rate of about 10'5, 10'4, 10'3, 10'2, or about IO mutations per base per hour.
[0035] Viruses
[0036] Viruses may be used with the systems and methods for directed evolution provided herein. In order to produce variants, a virus which promotes mutations (e.g., a mutagenic virus) in the transgene may be used. In some instances, the virus comprises an error prone replicase. In some instances, the virus is a DNA virus, an RNA virus, or a reverse transcribing virus. In some instances, the RNA virus is a positive-strand RNA virus. In some instances the positive-strand virus includes but is not limited to Bymoviruses, comoviruses, nepoviruses, nodaviruses, picomaviruses, potyviruses, sobemoviruses luteoviruses Carmoviruses, dianthoviruses, flaviviruses, pestiviruses, statoviruses, tombusviruses, single-stranded RNA bacteriophages, hepatitis C virus Alphaviruses, carlaviruses, furoviruses, hordeiviruses, potexviruses, rubiviruses, tobraviruses, tncomaviruses, tymoviruses, apple chlorotic leaf spot virus, beet yellows virus and hepatitis E virus. In some instances, the virus is an alphavirus. In some instances, the alphavirus is a rubi-like, tobamo-like, or tymo-like viruses. In some instances, the vims is a Sindbis vims. In some instances, a vims is altered to decrease its replication fidelity. In some instances, the virus lacks proof-reading polymerases or other mechanisms which increase mutation rate. In some instances, the virus has a mutation rate of about 10'5-l O'1, 10’3-10’2, 10'4- 10’1, IO - 10’2, lO’MO’1, 10’3-l O’2, or about 10’2-l O’1, mutations per base replicated. In some instances, the vims has a mutation rate of about 10'5, 104, 10'3, 10'2, or about 10 mutations per base replicated. In some instances, the virus has a mutation rate of at least 10'5, 104, 10'3, 10'2, or at least 104mutations per base replicated.
[0037] Cells
[0038] Directed evolution systems provided herein may be used inside various cell types. In some instances, the cell comprises an animal, plant, fungal, archaeal, or bacterial cell. In some instances, the cell comprises a mammalian cell. In some instances, the mammalian cell is derived from a human, pig, cow, hamster, or other mammal. In some instances, the mammalian cell does not comprise a human cell. In some instances, a cell is selected such that it can be infected by a vims described herein. In some instances, the mammalian cell is genetically modified, for instance to incorporate machinery for directed evolution. In some instances, the host cell and vims are modified for dependency / complementarity or safety.
[0039] Methods
[0040] Provided herein are methods for directed evolution. In some instances, a directed evolution platform is configured to continuously perform the steps of (1) mutagenesis, (2) expression, and (3) screening. In some instances, mutagenic outcomes of interest are linked to cellular activation of a transcriptional element, such as a synthetic promotor. The synthetic promoter in some instances drives a measurable output, such as viral replication and / or a reporter assay. A method for continuous evolution of a transgene in some instances comprises one or more steps of selecting a synthetic promoter from a library of polynucleotides; culturing a first population of host cells; and infecting the first population of host cells with a first population of viruses. In some embodiments, the method for directed evolution comprises, selecting a synthetic promoter from a library of polynucleotides, culturing a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle and wherein expression of the at least one viral gene is under control of the synthetic promoter from the library of polynucleotides, and infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the at least one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene. In some embodiments, the selecting comprises a method described in the specification.
[0041] In some embodiments, the first population of host cells comprises at least one viral gene for packaging a virus into an infectious viral particle comprising a transgene and a subset of viral genes. In some embodiments, the host cell is transfected with a plasmid. In some embodiments, the plasmid encodes for at least a synthetic promoter and a viral gene for packaging a virus comprises a viral capsid gene under control of said synthetic promoter. In some embodiments, a viral gene for packaging a virus comprises one or more of a viral capsid gene, E3, E2, and E2 of a virus (e.g. , Sindbis virus). In some instances, expression of the at least one viral gene is under control of the synthetic promoter from the library of polynucleotides. The subset of viral genes in some instances comprises viral packaging genes. The subset of viral genes in some instances comprises one or more of nSPl, nSP2, nSP3, and nSP4 of a virus (e.g., Sindbis virus).
[0042] In some embodiments, infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, and wherein the atleast one viral gene for packaging the virus into infectious viral particles is expressed in response to activity of the transgene. In some embodiments, the subset of viral genes comprise non-structural proteins, capsid, two envelope proteins. In some embodiments, the subset of viral genes comprise four non-structural proteins, capsid, two envelope proteins. In some embodiments, the subset of viral genes comprise non-structural proteins. In some embodiments, the subset of viral genes comprise capsid. In some embodiments, the subset of viral genes comprise envelope proteins. In some embodiments, the subset of viral genes does not comprise a non-structural protein, the capsid, or an envelope protein. In some embodiments, the subset of viral genes does not comprise the viral gene transfected into the host cell. In some embodiments, the virus cannot be packaged without expression of the viral gene transfected into the host cell.
[0043] In some embodiments, the method further comprises incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene. In some embodiments, the method further comprises generating a second population of viruses expressed by the first population of the host cells. In some embodiments, the method further comprises infecting a second population of host cells with the second population of viruses. In some embodiments, the method comprises incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene, generating a second population of viruses expressed by the first population of the host cells, and generating a second population of viruses expressed by the first population of the host cells. In some embodiments, the method is repeated at least repeated at least once, twice, three times, four times, five times, six times, seven times, eight times, nine times, or ten times. In some embodiments, the method is repeated at least once, twice, three times, four times, five times, six times, seven times, eight times, nine times, or ten times.
[0044] The infection and subsequent expression of the virus results in mutated transgenes. To analyze the results, the virus is sequenced to detect mutations in the transgene. In some embodiments, the method comprises amplifying RNA from the second population of viruses. In some embodiments, the method comprises sequencing the RNA. In some embodiments, the method comprises quantifying the RNA and / or barcode counts. In some embodiments, the method comprises amplifying RNA from the second population of viruses, sequencing the RNA, and quantify ing the RNA and / or barcode counts.
[0045] Additional evolution methods may be utilized with the systems and methods described herein. In some instances, the systems and methods are configured for evolution inmammalian cells. In some instances, one or more essential viral genes are removed and placed under control of a synthetic promoter from a polynucleotide library.
[0046] For example, synthetic promoter libraries may be used to select appropriate response elements for use in host-dependent propagation of viral like particles (VLV). Mutant transgenes with the desired properties activate response elements. These response elements in turn regulate VSVG (Indiana vesiculovirus G coat protein), which allows viruses harboring the mutant transgene to package and infect additional cells. An exemplary VLV system (PROTEUS) is described in Cole et al., bioRxiv 2024.04.20.590384 which is hereby incorporated by reference for its teaching of a VLV system.
[0047] In some embodiments, an engineered adenovirus may also be employed, wherein the adenovirus has genes El, AdPol, AdProt, and E3 deleted. These genes are instead complemented by the host cell, including a constitutively active error prone polymerase (e.g., AdPol) to induce mutations in the transgene. Mutated transgenes are configured to activate adenoviral protease (AdProt) needed for viral replication. The synthetic promoter libraries described herein may be used to drive expression of AdProt in response to the transgene. An exemplary adenovirus system (mPACE) is described in Berman et al. (J Am Chem Soc. 2018 December 26; 140(51): 18093-18103) which is hereby incorporated by reference for its teaching of an adenovirus system.
[0048] Multiple cycles of evolution may be used to evolve a transgene. The continuous evolution system in some embodiments, is executed for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or more than 30 days. In some embodiments, the continuous evolution system is executed for no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or no more than 30 days. In some embodiments, the continuous evolution system is executed for 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or no more than 30 days. In some embodiments, the continuous evolution system is executed for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or 30 cycles. In some embodiments, the continuous evolution system is executed for no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or 30 cycles. In some embodiments, the continuous evolution system is executed for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, or 30 cycles.
[0049] After optimization, final promoters with optimal RNA / DNA ratios can be chosen for the directed evolution process. Optimal RNA / DNA ratios in some instances are at least about 0.1, 0.2, 0.3 0.4, 0.5, 0.7, 0.9, 1, 1.5, 2, 2.5, or 3. Optimal RNA / DNA ratios in some instances are at most about 0.1, 0.2, 0.3 0.4, 0.5, 0.7, 0.9, 1, 1.5, 2, 2.5, or 3. Optimal RNA / DNA ratios in some instances are about 0.1, 0.2, 0.3 0.4, 0.5, 0.7, 0.9, 1, 1.5, 2, 2.5, or 3.Barcodes
[0050] A component of the VEGAS platform is the identification of mutations in the transgene. After each round of VEGAS, vims populations can be sequenced to identify mutations in the transgene. RNA, DNA, or cDNA from the vims populations can be labeled with a unique barcode and amplified prior to sequencing. Sequencing allows for high resolution reads, enabling accurate detection of tme variants. A true variant will be present in every amplified product originating from the original sequence as identified by aligning all products with a barcode. Each sequence amplified will have a different random barcode that will indicate that the amplified product originated from that sequence. Background caused by the fidelity of the amplification process can be eliminated because tme variants will be present in all amplified products and background representing random error will only be present in single amplification products. The use of barcodes in this method will help distinguish between tme mutations generated during VEGAS and errors generated during the amplification and sequencing process.
[0051] As used herein, the term “unique molecular identifier (UMI)” refers to a unique nucleic acid sequence that is attached to each of a plurality of nucleic acid molecules. When incorporated into a nucleic acid molecule, an UMI in some instances is used to correct for subsequent amplification bias by directly counting UMIs that are sequenced after amplification. The design, incorporation and application of UMIs is described, for example, in Int. Pat. Appl. Pub. No. WO 2012 / 142213, Islam et al. Nat. Methods (2014) 11: 163-166, Kivioja, T. et al. Nat. Methods (2012) 9: 72-74, Brenner et al. (2000) PNAS 97(4), 1665, and Hollas and Schuler, (2003) Conference: 3rd International Workshop on Algorithms in Bioinformatics, Volume: 2812 which is hereby incorporated by reference for its teaching of unique molecular indentifiers.
[0052] As used herein, the term "barcode" refers to a nucleic acid tag that can be used to identify a sample or source of the nucleic acid material. Thus, where nucleic acid samples are derived from multiple sources, the nucleic acids in each nucleic acid sample are in some instances tagged with different nucleic acid tags such that the source of the sample can be identified. Barcodes, also commonly referred to indexes, tags, and the like, are well known to those of skill in the art. Any suitable barcode or set of barcodes can be used. See, e.g., nonlimiting examples provided in U.S. Pat. No. 8,053,192 and Int. Pat. Appl. Pub. No. W02005 / 068656. In preferred embodiments, sequencing is performed using unique molecular identifiers (UMI). The term “unique molecular identifiers” (UMI) as used herein refers to a sequencing linker or a subtype of nucleic acid barcode used in a method that uses molecular tagsto detect and quantify unique amplified products. The UMI may also be used to determine the number of transcripts that gave rise to an amplified product.
[0053] In certain embodiments, each barcode is uniquely identifiable. In certain embodiments, each barcode comprises a unique molecular identifier (UMI). In certain embodiments, the barcode comprises at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In certain embodiments, each barcode comprises at most 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In certain embodiments, each barcode comprises at least one, two, three, four, or five YR nucleotide repeats. In certain embodiments, each barcode comprises one, two, three, four, or five YR nucleotide repeats. In certain embodiments, each YR repeat is separated by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides. In certain embodiments, each YR repeat is separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides. In certain embodiments, the barcode is restricted to at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% Y nucleotides. In certain embodiments, the barcode is restricted to about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% Y nucleotides. In certain embodiments, the barcode is restricted to at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% R nucleotides. In certain embodiments, the barcode is restricted to about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% R nucleotides. In some embodiments each barcode comprises the oligonucleotide sequence NNNN, NNNNNN,NNYRNN, NNNYRNNN, NNNYRNNNYRNNN, NNNYRNNNYRNNNYRNNN, NNNYRNNNYRWNYRNNNYRNNN. In certain embodiments, each barcode comprises a Hamming distance of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other barcodes. In certain embodiments, each barcode comprises a Hamming distance of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other barcodes. In certain embodiments, each barcode comprises a Hamming distance of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other UMI. In certain embodiments, each barcode comprises a Hamming distance of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other UMI.In certain embodiments, each barcode comprises a Levenshtein distance of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other barcodes. In certain embodiments, each barcode comprises a Levenshtein distance of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other barcodes. In certain embodiments, each barcode comprises a Levenshtein distance of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other UMI. In certain embodiments, each barcode comprises a Levenshtein distance of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 from other UMI.
[0054] In certain embodiments, the barcode is attached to viral oligonucleotides. In certain embodiments, the barcode is added to the 5’ end of viral oligonucleotides. In certain embodiments, the barcode is added to the 3’ end of viral oligonucleotides. In certain embodiments, the barcode is attached by primer-based addition, PCR-based addition, ligationbased addition, fragmentation and end repair, or adapter ligation. In certain embodiments, the viral oligonucleotides are RNA. In certain embodiments, the RNA is amplified. In certain embodiments, the RNA from the second population of viruses is attached with the barcode. In certain embodiments, the amplified RNA is sequenced to generate a plurality of reads. In certain embodiments, the plurality of reads are organized to distinguish between amplification errors and single nucleotide polymorphisms present in the amplified RNA. In certain embodiments, the viral oligonucleotides are cDNA. In certain embodiments, the RNA from the second population of viruses is reverse transcribed into a library of cDNA. In certain embodiments, the library of cDNA is attached with the barcode. In certain embodiments, the library of cDNA is amplified. In certain embodiments, the amplified library of cDNA is sequenced to generate a plurality of reads. In certain embodiments, the plurality of reads are organized to distinguish between amplification errors and single nucleotide polymorphisms present in the amplified library of cDNA.
[0055] Viral oligonucleotides can be optionally labeled with multiple barcodes in combinatorial fashion (for example, using multiple barcodes bound to one or more specific binding agents that specifically recognizing the target molecule), thus greatly expanding the number of unique identifiers possible within a particular barcode pool. In certain embodiments, barcodes are added to a growing barcode concatemer attached to a target molecule, for example, one at a time. In other embodiments, multiple barcodes are assembled prior to attachment to a target molecule.
[0056] Sequencing allows for high resolution reads, enabling accurate detection of true variants. A true variant will be present in every amplified product originating from the original sequence as identified by aligning all products with a UMI. Each sequence amplified will have adifferent random UMI that will indicate that the amplified product originated from that sequence. Background caused by the fidelity of the amplification process can be eliminated because true variants will be present in all amplified products and background representing random error will only be present in single amplification products (See e.g., Islam S. et al., 2014. Nature Methods No: 11, 163-166). Labeled oligonucleotides can be amplified by methods known in the art, such as polymerase chain reaction (PCR). For example, the barcode can contain universal primer recognition sequences that can be bound by a PCR primer for PCR amplification and subsequent high-throughput sequencing. In certain embodiments, the barcode includes or is linked to sequencing adapters (for example, universal primer recognition sequences) such that the barcode and sequencing adapter elements are both coupled to the oligonucleotides. In particular examples, the sequence of the barcode is amplified, for example using PCR. In some embodiments, a barcode further comprises a sequencing adaptor. In some embodiments, a barcode further comprises universal priming sites. A barcode (or a concatemer thereof), an oligonucleotide (for example, a DNA, cDNA, or RNA molecule), and / or a nucleic acid encoding a transgene may be optionally sequenced by any method known in the art, for example, methods of high-throughput sequencing, also known as next generation sequencing or deep sequencing. A oligonucleotide labeled with a barcode can be sequenced with the barcode to produce a single read and / or contig containing the sequence, or portions thereof, of both the oligonucleotide and the barcode. Exemplary next generation sequencing technologies include, for example, Illumina sequencing, Ion Torrent sequencing, 454 sequencing, SOLiD sequencing, and nanopore sequencing amongst others. In certain embodiments, the sequencing is third- generation sequencing, next generation sequencing, short-read sequencing, long-read sequencing, single molecule real time sequencing (SMRT), or nanopore sequencing.
[0057] In certain embodiments, the method detects mutational frequencies of at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99%. In certain embodiments, the method detects mutational frequencies of about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99%.EXAMPLES
[0058] The following examples are set forth to illustrate more clearly the principle and practice of embodiments disclosed herein to those skilled in the art and are not to be construedas limiting the scope of any claimed embodiments. Unless otherwise stated, all parts and percentages are on a weight basis.Example 1: Directed Evolution using UMIs
[0059] This is a method for the amplification of viral genomic material for application in a Unique Molecular Index sequencing pipeline using long read DNA sequencing technology. Sindbis virus has been used to illustrate the method (general method described in English et al., Cell 2019 Jul 25;178(3):748-761.el7.).
[0060] VEGAS was performed on tTA through activation of tetO7-driven Sindbis structural genome expression without selection. Viral supernatant was harvested from each round of evolution and clarified by centrifugation at 4°C for 20 minutes at 3000 x g. Clarified viral supernatant was then precipitated through 20% sucrose and the total volume of precipitated virus was used as input for isolating viral RNA using MagMAX Viral RNA Isolation Kit (Applied Biosystems cat no. AM1939). Viral RNA was aliquoted in single-use volumes and stored at -80 degrees Celsius. cDNA copies of target sequences within the viral RNA were made using the SuperScript IV One-Step RT-PCR system (Invitrogen cat no. 12594100) with target-specific primers. These reactions were cleaned up using a 0.5X ratio of KAPA Pure Beads (Roche cat no. 7983271001), resuspended in lOmM Tris-Cl and DNA concentration was measured with either a spectrophotometer or fluorometer. Unique molecular index (UMI) labeling occurred using 1 pg of cleaned-up cDNA as input and 2 cycles of PCR using Platinum SuperFi II Green PCR Master Mix (Invitrogen cat no. 12369010) with primers containing (1) a sequence specific to the 3’ and 5’ ends of the cDNA target region, (2) a degenerate primer sequence and (3) a universal “landing site” for the subsequent round of PCR amplification (see Table 1). UMI labeled cDNA was cleaned up using a 0.5X ratio of KAPA Pure Beads with I pg of salmon DNA and resuspended in lOmM Tris-Cl. The cleaned-up, UMI-labeled cDNA was then used as input for a full amplification PCR with 35 cycles using Platinum SuperFi II Green PCR Master Mix and primers that bind to the universal landing sites mentioned above. The fully amplified products were again cleaned using a 0.5X ratio of KAPA Pure Beads and measured with a spectrophotometer or fluorometer. UMI-labeled amplicons were then barcoded and prepared for sequencing following Oxford Nanopore Technology’s Native Barcoding Kit 24 V14 protocol using their Native Barcoding Kit 24 V14 library prep kit (ONT cat no. SQK-NBD114.24). 20 finol of library was loaded onto a MinlON Flow Cell (R10.4. 1) and sequenced for ~18h with super- accurate basecalling using ConSeqUMI.
[0061] Sample primers used in this example are shown in Table 1 and Table 2.Table 1: Primers to UMI tag and amplify transgene regionTable 2: Primers to UMI tag and amplify CSE region
[0062] UMI Processing with ConSeqUMI
[0063] The pools of virus were collected and sequenced as described above. Mutant GPR40 sequences were identified in each pooled virus population. GPR40 mutants were then validated in an activity assay. Basecalling was performed using the ConSeqUMI online tool (GitHub).
[0064] In addition to sequencing files in fastq format, ConSeqUMI uses a text file with four lines, representing the four adapters used in the wet lab portion of the protocol. These adapters represent the template-specific and the amplification primer binding portions of the UMI primers or the template-specific sequence and backbone sequences immediately adjacent to the UMI in the pUMI plasmid library. ConSeqUMI utilizes the cutadapt software package to identify these adapters and extract UMI sequences and target sequences from input reads in fastq format. The first and last two hundred base pairs of each nanopore read are initially isolated for adapter identification, and are labeled the top and bottom of the read, respectively. The bottom of the read is converted to its reverse complement to reflect the second set of adapters on thecomplementary strand of the original sequence. Using cutadapt, the front and back linked adapters are used to extract a UMI from the 5’ and 3’ sections of each read. ConSeqUMI recognizes that reverse strands also appear in nanopore data and thus, if cutadapt fails to identify the appropriate linked adapters, the process is repeated on the reverse complement of the entire strand. In this way, the final collection of UMI and target sequences are all representatives of the forward strand of each read.
[0065] As an optional setting, ConSeqUMI allows for user-specified UMI lengths to aid in adapter identification. During experimentation, some reads were not sequenced to completion. Without being bound by theory, this may have been from possibly truncating the front adapter of either linked adapter pair. Provided the program knows the expected length of the UMI, an identified adapter internal to the UMI can independently be used to extract the UMI while identifying the start or end of the target sequence. ConSeqUMI allows users to provide the expected UMI length, in which case the front adapters are considered optional, thereby retrieving any prematurely terminated reads containing internal adapters and additional nucleotides equal to or greater than the specified UMI length.
[0066] Once adapters were identified and UMIs and target sequences extracted, ConSeqUMI uses the starcode software package to cluster UMIs based on Levenshtein distance similarity. Starcode offers a significant reduction in time and memory complexity compared to other clustering software for short sequences. The starcode ‘seq-id’ setting attaches lists of original indices to each cluster, retaining memory linking reads to the cluster categorization. UMI pairs are found by running starcode with this setting separately on UMIs extracted from the top and bottom of each read and finding the intersection of indices between their clusters.
[0067] UMI chimera identification and removal
[0068] Once UMI pairs are identified and matched to reads, target sequences are then clustered together for downstream consensus sequence generation. In the final output of the UMI processing step, sequences are saved one fastq file per cluster in a bins folder in the output directory. Furthermore, data files regarding read rejection, chimera analysis and starcode output are saved to a data analysis folder in the output directory for further investigation.
[0069] De novo consensus sequence generation
[0070] The pairwise algorithm uses a dual-method approach to generating consensus sequences from bins of target sequences. In the first stage of the pairwise method, a preliminary consensus sequence is generated using a custom frameshift algorithm. After buffering the ends of the binned sequences with spaces to match the maximum (central vector) length, thealgorithm extracts a ten base pair window from the front of each sequence. The most common sequence is established as the front end of the consensus sequence. The algorithm determines the most common eleventh base pair for all sequences that contain the initial ten base pairs. Each base pair is then added sequentially, searching for the base pair immediately following the last ten identified base pairs of the consensus sequences. These ten base pairs are searched within a tw enty base pair window that shifts for every identified base pair. In some instances, the method results in skipping or duplicating base pairs in homopolymeric regions, but generates a respectable reference sequence. This method is also used for providing a reference to the medaka consensus algorithm.
[0071] In the second stage of the pairwise method, the preliminary consensus sequence generated in the first stage is adjusted to more accurately reflect binned sequence commonalities. A pairwise alignment between the preliminary consensus sequence and each binned sequence is performed, where individual discrepancies are saved and collected. The most common discrepancy between all binned sequences is then applied to the consensus sequence. This process is repeated so long as each added discrepancy increases the average alignment score. When an added discrepancy decreases the average, the discrepancy is ignored and the prior consensus sequence is considered the true consensus sequence.
[0072] Provided a directory of fastq files wherein each file contains reads expected to coincide with a single consensus, such as the bin directory output of the UMI processing step, ConSeqUMI generates a single fasta file containing the consensus sequence for each fastq file. This file is saved in a newly created directory to separate it from other processes.
[0073] By default, ConSeqUMI only generates consensus sequences for clusters with 50 or more reads. However, this benchmark can be adjusted by the end user, as the sequencing error rate can differ significantly between sequencing protocols and runs. To help researchers establish a benchmark level appropriate for their own data, ConSeqUMI includes a bootstrapped analysis workflow^ in addition to the UMI processing and consensus generation workflows. This setting extracts a set of binned reads, such as an output fastq file of the UMI processing step, and iteratively and randomly extracts subclusters incrementing at a user-defined size. A consensus sequence is generated from each subcluster, and subsampling at each size occurs a user-defined number of times. While both the subsampling step size and iterations of each step size can be set by the end user, the default values are 10 and 100, respectively. The Levenshtein distance between consensus sequences generated for each subcluster and the one generated from the full cluster are then collected and returned in a comma-delimited file for benchmark cutoff analysis.Example 2: GPR40 campaign using UMIs
[0074] A directed evolution campaign was used to identity transgene mutants of activated GPR40 using the VEGAS platform and a TRE library. VEGAS directed evolution steps were performed according to the general methods of Example 1. The goal of the VEGAS campaign was to evolve a constitutively active mutant of GPR40 using Compound 1 (fasiglifam, TAK- 875, PubChemID: 24857286, CAS No.: 1000413-72-8) or Compound 2 (“Compound A”, see Rives et al., PubChemID: 134465785). Promoters (SEQ ID NO: 14 and 15) were identified by screening of a TRE library (Zahm et al. bioRxiv 2023.05. 11.539703) and are shown in Table 3Table 3^Promoter Key: Library ID#, rotations, spacer set, minimal promoter)
[0075] FIGs. 1A and IB depict plots of dose-response curves measuring the ratio firefly to Renilla luciferase activity (Fluc / Rluc) at various concentrations of Compound 1 (FIG. 1A) or Compound 2 (FIG. IB). The promoter construct used for treatment with Compound 1 was Promega_SRE-2048 (8 rotations, no spacers, minCMV), as identified described above, and had a maximal detected Fluc / Rluc ratio of about 6 for the protein GPR40. The promoter construct used for treatment with Compound 2 was Promega_CRE-2002 (1 rotation, no spacer set,minCMV), as identified described above, and had a maximal detected Fluc / Rluc ratio of about 20 for the protein GPR40. These results demonstrate that the method for selecting the promoter using the library system was validated by showing the increased signal to noise ratio. Overall, these results demonstrate the efficacy of the TRE library selection for use in the VEGAS system.
[0076] The cargo-specific synthetic promoter identified using the above protocol (Promega_CRE-2002, 1, none, minCMV) was then utilized in a VEGAS campaign. Host cells were transfected with the plasmid comprising the identified promoter and a gene encoding for at least one viral protein required for packaging. Initial packaging of virus (round 0, or RO) was completed by electroporating approximately equimolar amounts of RNA expressing the following: (1) viral replication machinery and the transgene of interest (pTSin4-GPR40), (2) capsid protein (pCapSin) and (3) glycoprotein (pGlySin). This campaign included three rounds of virus generation (Rl, R2, R3), each with decreasing amounts of applied compound in an attempt to provide selective pressure for the generation of a constitutively active mutant. Each round of harvested viral supernatant was concentrated through 20% sucrose by applying all available viral supernatant (~16 mL) to a layer of 8 mL of 20% sucrose, brought to volume with sucrose gradient buffer (50mM Tris-HCl pH 7, 100 mM NaCl, 0.5mM EDTA), and spun in an Ultracentrifuge at 30,000 RPM for 90 minutes at 4°C. Pelleted virus was resuspended in 400 LL L of sucrose gradient buffer and the entire volume was used for RNA isolation using the MagMAX Viral Isolation Kit.
[0077] Isolated RNA was amplified via RT-PCR using a 26S-binding forward primer and a pool of 3’UTR-binding reverse primers to amplify the GPR40 transgene region. A pool of reverse primers was used to try and mitigate the potential issue of the 3’UTR mutating throughout the evolution rounds, leading to loss of primer binding sites. For each round, RT- PCR was run with and without reverse transcriptase (RTase) to account for potential DNA contamination. Using primer / probe sets that bind and amplify two distinct regions of the pTSin4 RNA in an attempt to mitigate the potential issue of evolving primer binding sites, RT-qPCR was run on the isolated RNA from each round.
[0078] RT-PCR products were cleaned up using the Qiagen PCR Clean-Up kit, then assembled into digested pcDNA3. 1 backbone using NEBuilder HiFi DNA Assembly. Colony forming units (CFU) were compared between the no insert control and later rounds, and suggested the presence of amplifiable RNA material. Single colonies from each round were picked, mmiprepped, and sequenced using Sanger sequencing.
[0079] Samples were further processed with unique molecular identifier (UMI) labels. Entire transformation plates were harvested, UMI tagged, and library prepped for sequencing on the Nanopore MinlON sequencer. In addition to packaged virus (RO) and the three evolution rounds (R1-R3), in vitro transcribed (IVT) RNA was also prepared and sequenced. IVT RNA was used for initial packaging and served as a reference sequence for any observed sequence mutations present in later rounds. A rough approximation of the general contents of the sequencing run for all rounds and IVT material was generated (FIG. 4B). A significant number of reads aligned to GPR40 for most rounds, which was an improvement to the Sanger sequencing results.
[0080] FIG. 2 shows a RT-qPCR plot measuring the quantification cycle over multiple rounds of a VEGAS continuous evolution campaign. RT-qPCR was performed using a 10 pL reaction with 1 pl RNA input) w / sucrose gradient pipetted viral RNA. Briefly, 8 mL 20% sucrose + 29 mL viral supernatant (~16mL virus brought up to volume with sucrose gradient buffer) were spun at 30k rpm at 4°C for 90 min, media aspirated off, the “pellet” resuspended in 400 pL sucrose gradient buffer, and the RNA from pelleted virus isolated using MagMAX Kit (400 pL input and 50 pL final resuspension volume). Using the promoter as selected described above, the VEGAS system was applied to GPR40 for continuous evolution of the protein to have greater activity in the presence of Compound 2. As shown on the leftmost data point, in Ro there was identified a large pool of Sindbis virus. Collected virus populations were used in subsequent rounds of Ri, R2, and R3. The virus populations showed diminishing virus counts over the subsequent rounds. Population samples were collected after each round and sequenced. These results show that the virus can be used in multiple rounds of VEGAS. Legend: GPR40 R0: virus collected after Round 0 (Ro); GPR40 Rl: virus collected after Ri; GPR40 R2: virus collected after R2; GPR40 R3: virus collected after R3; water: negative control.
[0081] FIG. 3 shows a summary of Sanger sequencing results of identified mutants in each of the rounds. After Round 1, 24 unique clones were identified, 22 of which comprised GPR40 mutants shown. After Round 2, 24 unique clones were identified, and 2 aligned to GPR40. After Round 3, 24 unique clones were identified, and 0 aligned to GPR40. The results show that using VEGAS, selective pressure can be used to generate mutants in the protein of interest. Mutations identified from sequencing were at amino acid positions F 191 , G238, G254, G280, and R281. The identified mutant GPR40 Fl 91 S was further validated in follow-on experiments.
[0082] FIG. 4A shows a polyacrylamide gel electrophoresis (PAGE) gel of PCR amplified and UMI tagged samples from the virus populations. From right to left, the lanes comprise a ladder, initial virus population, Ro virus population, Ri virus population, R2 virus population, R3virus population, and library of pooled virus for sequencing. The virus populations and pooled virus were PCR amplified and UMI tagged material. 1 pg of viral samples were prepared were amplified with forward and reverse primers for two cycles. The PCR amplified and UMI tagged material was then run on the gel. The sequencing library was prepped with 200 frnol input, assuming 1.5 kB in length) and 20 frnol loaded on to a nanopore sequencer. About equal amounts of viral material was observed in initial, Ro, and Ri lanes. There was a decrease in the amount of material in the Ri and R2 lanes which aligned with the results seen in FIG. 2. In FIG. 4B, each column (left to right) represents first IVT (original input used for viral packaging as a control) and each subsequent round. The height of the column shows the number of consensus seq that come out for each round. Each bar shading pattern represent the portion of sequences that aligned to various reference sequences. A summary of the consensus sequences is shown in Table 4. UMI labeling identified additional mutant GPR40 sequences. In addition to the mutations identified by Sanger sequencing, more mutations were identified from UMI labeling and nanopore sequencing. Five mutations, C165R, and nonsense mutations A98X, L140X, and two L190X (a 1 base pair deletion and a 2 base pair deletion) were identified. These results demonstrate that UMIs increase identification of mutations that arise during a VEGAS campaign.Table 4N.D. : Not determined
[0083] FIGs. 5A-5D depict plots of dose-response curves measuring the ratio firefly to Renilla luciferase activity (Fluc / Rluc) at various concentrations of Compound 1 or Compound 2 with either the CRE or SRE promoters. The data show the response of GPR40 wildtype compared to two identified mutants from the campaign as described above (F191S and C136R) or vehicle control. FIG. 5A shows the dose response curve of GPR40 wildt pe and mutants withCompound 1 and the library identified CRE promoter. FIG. 5B shows the dose response curve of GPR40 wildtype and mutants with Compound 1 and the library identified SRE promoter. FIG. 5C shows the dose response curve of GPR40 wildtype and mutants with Compound 2 and the library identified CRE promoter. FIG. 5D shows the dose response curve of GPR40 wildtype and mutants with Compound 2 and the library identified SRE promoter. These results demonstrate that the VEGAS system could be used in a continuous evolution system to mutate a protein, such as GPR40.
[0084] Overall, the data shows results for a VEGAS continuous directed evolution campaign to mutate a transgene, GPR40, and combining with UMI methods to identify unique mutations generated in VEGAS.Example 3: Additional UMI designs
[0085] The general methods of Examples 1-2 were followed with modification: UMI barcodes were selected from Table 5, which provides both the general UMI library design (SEQ ID NO.: 16) and exemplary sequences in the library (SEQ ID NOS. : 17-18).Table 5Key: B = C or G or T; D = A or G or T; H = A or C or T; and V =A or C or G.Example 4: Adenovirus-based directed evolution with UMI barcoding
[0086] A directed evolution campaign is performed following the general methods described in Berman et al. (J Am Chem Soc. 2018 December 26; 140(51): 18093-18103), with modification: following at least one round of evolution, mutations in transgene nucleic acids are identified using UMI barcoding as described above.
[0087] The examples described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
WHAT IS CLAIMED IS:
1. A method for continuous evolution of a transgene, comprising:(a) culturing a first population of host cells comprising at least one viral gene for packaging a virus into an infectious viral particle, wherein expression of the at least one viral gene is under control of a promoter;(b) infecting the first population of host cells with a first population of viruses comprising a transgene and a subset of viral genes, wherein the first population of viruses is capable of infecting the first population of host cells, wherein the at least one viral gene for packaging the virus into an infectious viral particle is expressed in response to activity of the transgene, and wherein packaging the virus into an infectious viral particle results in a second population of viruses; and(c) labeling oligonucleotides of the second population of viruses expressed by the first population of host cells with at least one barcode.
2. The method of claim 1, wherein the oligonucleotides of the second population of viruses comprise RNA, DNA, or cDNA.
3. The method of claims 1 or 2, wherein the at least one barcode is uniquely identifiable.
4. The method of any one of claims 1-3, wherein the at least one barcodes comprises at least one unique molecular identifier (UMI).
5. The method of any one of claims 1-4, wherein the step of labeling oligonucleotides comprises labeling with at least about IxlO3, IxlO4, IxlO5, IxlO6, 3xl06, IxlO7, or 1x108barcodes.
6. The method of any one of claims 1-5, wherein each barcode comprises at least 24 nucleotides.
7. The method of any one of claims 1-5, wherein each barcode comprises at most 24 nucleotides.
8. The method of any one of claims 1-5, wherein each barcode comprises about 24 nucleotides.
9. The method of any one of claims 1-8, wherein each barcode comprise at least three YR nucleotide repeats.
10. The method of claim9, wherein each barcode comprise each of the three YR nucleotide repeats are separated by at least three nucleotides.
11. The method of any one of claims 1-10, wherein each barcode is restricted to at least about 10% Y nucleotides.
12. The method of any one of claims 1-11, wherein each barcode is restricted to at least about 10% R nucleotides.
13. The method of any one of claims 1-12, wherein each barcode comprises the oligonucleotide sequence NNNYRNNNYRNNNYRNNN.
14. The method of any one of claims 1-12, wherein each barcode comprises the oligonucleotide sequence VHBDVHBDVHBDBDHVBDHVBDHV (SEQ ID NO: 16).
15. The method of any one of claims 1-13, wherein each barcode comprises a Hamming distance of at least 1, 2, 3, 4, or 5 from other barcode.
16. The method of any one of claims 1-15, wherein each barcode comprises a Hamming distance of at least 2 relative to other UMIs.
17. The method of any one of claims 1-16, wherein each barcode comprises a Levenshtein distance of at least 1, 2, 3, 4, or 5 from other barcode.
18. The method of any one of claims 1-17, wherein each barcode comprises a Levenshtein distance of at least 2 relative to other barcodes.
19. The method of any one of claims 2-18, wherein the RNA of the second population of viruses is attached with the barcode.
20. The method of any one of claims 2-19, wherein the RNA of the second population of viruses is amplified.
21. The method of claim 20, wherein the amplified RNA is sequenced to generate a plurality of reads.
22. The method of claim 21, wherein the plurality of reads are organized to distinguish between amplification errors and single nucleotide polymorphisms present in the amplified RNA.
23. The method of any one of claims 2-18, wherein the RNA of the second population of viruses is reverse transcribed into a library of cDNA.
24. The method of claim 23, wherein the library of cDNA is attached with the barcode.
25. The method of any one of claims 23-24, wherein the library of cDNA is amplified.
26. The method of claim 25, wherein the amplified library of cDNA is sequenced to generate a plurality of reads.
27. The method of claim 26, wherein the plurality of reads are organized to distinguish between amplification errors and single nucleotide polymorphisms present in the amplified library of cDNA.
28. The method of any one of claims 1-27, wherein the barcode is attached by primer-based addition, PCR-based addition, ligation-based addition, fragmentation and end repair, or adapter ligation.
29. The method of any one of claims 1-28, wherein the method detects mutational frequencies of at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%.
30. The method of any one of claims 1-28, wherein the method detects mutational frequencies of at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%.
31. The method of any one of claims 1-28, wherein the method detects mutational frequencies of about 7%.
32. The method of any one of claims 21-31, wherein the sequencing is third-generation sequencing, next generation sequencing, short-read sequencing, long-read sequencing, single molecule real time sequencing (SMRT), or nanopore sequencing.
33. The method of any one of claims 21-32, wherein the sequencing is long-read sequencing.
34. The method of any one of claims 1-33, wherein the method comprises incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene.
35. The method of any one of claims 1-34, wherein the method comprises: incubating the first population of host cells infected with the first population of virus with an agent capable of activating the transgene.
36. The method of any one of claims 1-35, wherein the method comprises infecting a second population of host cells with the second population of viruses.
37. The method of any one of claims 1-36, wherein steps (a)-(b) are repeated at least once, twice, three times, four times, five times, six times, seven times, eight times, nine times, or ten times.
38. The method of any one of claims 1-36, wherein steps (a)-(b) are repeated at least once.
39. The method of any one of claims 1-36, wherein steps (a)-(b) are repeated at least five times.
40. The method of any one of claims 1-36, wherein steps (a)-(b) are repeated at least six times.
41. The method of any one of claims 1-40, wherein the method further comprises:(g) amplifying RNA from the second population of viruses;(h) sequencing the RNA; and(i) quantifying the RNA and barcode counts.
42. The method of any one of claims 1-41, wherein the transgene is about 500, 1000, 2000 nucleotides in length.
43. The method of any one of claims 1-41, wherein the transgene is at least about 500 nucleotides in length.
44. The method of any one of claims 1-41, wherein the transgene is at least about 1000 nucleotides in length.
45. The method of any one of claims 1-41, wherein the transgene is at least about 2000 nucleotides in length.
46. The method of any one of claims 1-41, wherein the transgene is about 500 nucleotides in length.
47. The method of any one of claims 1-41, wherein the transgene is about 1000 nucleotides in length.
48. The method of any one of claims 1-41, wherein the transgene is about 2000 nucleotides in length.
49. The method of any one of claims 1-48, wherein the second population of viruses comprise at least one mutated transgene of interest.
50. The method of any one of claims 1-49, wherein repeating of steps (a)-(f) results in at least one mutation in the transgene of interest.
51. The method of any one of claims 1-50, wherein the mutated transgene of interest comprises at least one nucleic acid substitution, insertion, or deletion compared to the transgene.
52. The method of any one of claims 1-51, wherein the mutated transgene of interest comprises at least two nucleic acid substitutions, insertions, or deletions compared to the transgene.
53. The method of any one of claims 1-52, wherein the mutated transgene of interest comprises at most about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene.
54. The method of any one of claims 1-52, wherein the mutated transgene of interest comprises about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the transgene.
55. The method of any one of claims 1-52, wherein the mutated transgene of interest comprises at least about 99% sequence identity to the transgene.
56. The method of any one of claims 1-55, wherein the mutated transgene of interest encodes for a mutated protein comprising at least one amino acid substitution, insertion, or deletion compared to a protein encoded by the transgene.
57. The method of any one of claims 1-56, wherein the mutated transgene of interest comprises same, increased, or decreased activity of the transgene.
58. The method of any one of claims 1-56, wherein the mutated transgene of interest comprises increased activity of the transgene.
59. The method of any one of claims 1-56, wherein the mutated transgene of interest comprises decreased activity of the transgene.
60. The method of any one of claims 1-56, wherein the mutated transgene of interest comprises about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the trans gene.
61. The method of any one of claims 1-56, wherein the mutated transgene of interest comprises at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% increased activity of the transgene.
62. The method of any one of claims 1-60, wherein the mutated transgene of interest compnses about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene.
63. The method of any one of claims 1-56, wherein the mutated transgene of interest compnses at least about 1%, 2%, 3%, 5%, 10%, 25%, 50%, 100%, 200%, 250%, 500%, or 1000% decreased activity of the transgene.
64. The method of any one of claims 1-63, wherein the mutated transgene of interest comprising better activity results in a higher number of infectious viral particles.
65. The method of any one of claims 1-64, wherein the mutated transgene of interest comprising better activity has a higher RNA count than the transgene or other mutated transgenes.
66. The method of any one of claims 1-65, wherein a ratio of RNA to DNA of the transgene of the first population of host cells is at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
67. The method of any one of claims 1-65, wherein a ratio of RNA to DNA of the transgene of the first population of host cells is at most about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9. or 10.
68. The method of any one of claims 1-65, wherein a ratio of quantified RNA to DNA of the transgene is about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10.
69. The method of any one of claims 34-68, wherein the agent is a peptide, a protein, an antibody, a small molecule, a hormone, light, a vitamin, a neurotransmitter, an intermediary metabolite, a nucleic acid, or a metal.
70. The method of any one of claims 34-69, wherein the agent activates the transgene.
71. The method of any one of claims 1-70, wherein the activity of the trans gene is a signaling event.
72. The method of any one of claims 1-71, wherein the activity of the trans gene produces an activity-coupled transcriptional response.
73. The method of any one of claims 1-72, wherein the activity-coupled transcriptional response binds to the synthetic promoter.
74. The method of any one of claims 1-73, wherein binding of the activity-coupled transcriptional response with the synthetic promoter results in transcription of the at least one viral gene for packaging a virus into an infectious viral particle.
75. The method of any one of claims 1-74, wherein the transgene encodes for a protein.
76. The method of any one of claims 1-75, wherein the protein is activated by the agent and produces an activity-coupled transcriptional response wherein the activity-coupled transcriptional response is an endogenous transgene-mediated signaling pathway.
77. The method of any one of claims 1-76, wherein the protein is a G protein-coupled receptor (GPCR).
78. The method of any one of claims 1-77, wherein the agent activates the GPCR and produces a GPCR-mediated signaling response.
79. The method of any one of claims 1-78, wherein the GPCR-mediated signaling response binds to the SRE or the CRE.
80. The method of any one of claims 1-79, wherein the virus is a mutagenic virus.
81. The method of any one of claims 1-80, wherein the virus has a mutation rate of about 10" 5-10-3mutations per base replicated.
82. The method of any one of claims 1-80, wherein the virus has a mutation rate of about 10’ 3 mutations per base replicated.
83. The method of any one of claims 1-80, wherein the virus has a mutation rate of about 10'5mutations per base replicated.
84. The method of any one of claims 1-83, wherein the virus has a mutation rate of at least about 10'3mutations per base replicated.
85. The method of any one of claims 1-84, wherein the virus has a mutation rate of at most about 10'3mutations per base replicated.
86. The method of any one of claims 1-85, wherein the virus comprises an error prone replicase.
87. The method of any one of claims 1-86, wherein the virus is a DNA vims, an RNA vims, or a reverse transcribing virus.
88. The method of any one of claims 1-87, wherein the vims is an RNA vims.
89. The method of claim 87, wherein the RNA vims is a positive-strand RNA virus.
90. The method of any one of claims 1-89, wherein the vims is a Sindbis vims.
91. A system configured to perform the steps of any one of claims 1-90.