Compositions and methods for nucleic acid replication
Patent Information
- Application Number
- PCT/SG2025/050162
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2025-03-07
- Publication Date
- 2025-10-02
AI Technical Summary
Current methods for in vivo continuous directed evolution in E. coli, such as Phage-Assisted Continuous Evolution (PACE) and Targeted Artificial DNA Replisome (TADR), are limited by the size of genes that can be evolved, replication stability, error rate, and the need for specialized bioreactors, making it difficult to evolve larger genes and multiple genes efficiently.
A T7-based orthogonal replication system using a set of proteins including a DNA polymerase, RNA polymerase, single-strand binding protein, helicase-primase, and linearising proteins, capable of replicating DNA with a T7 origin of replication, which includes a protelomerase to maintain DNA in a linear form, and an error-prone DNA polymerase to increase mutagenesis rates.
The system enables efficient replication and mutagenesis of larger genes and multi-gene constructs up to 40 kb, overcoming size and error rate limitations, and allows for accelerated molecular evolution in E. coli without affecting host DNA.
Abstract
Description
[0001] COMPOSITIONS AND METHODS FOR NUCLEIC ACID REPLICATION
[0002] Technical field
[0003] The present invention relates, in general terms, to genetic engineering and more particularly to compositions and methods for replicating and mutagenising DNA containing a T7 origin of replication.
[0004] Background
[0005] Orthogonal replication systems allow the transfer of genes and genetic parts between species and arc useful in genetic engineering and vector production. Such systems can also be used in combination with mutagenic polymerases for accelerated continuous evolution of genes or proteins. They have been developed for this purpose in yeast and Bacillus bacteria.
[0006] E. coll is the most popular host for molecular biology and biotechnology. However, there are currently no orthogonal replication systems for E. coll. Current methods for in vivo continuous directed evolution in E. coli rely on phage replication, such as the Phage- Assisted Continuous Evolution (PACE) platform. PACE suffers from various disadvantages, including: (i) the size of the gene that can be evolved is limited to 5000 bp to be able to fit into the M13 phage; (ii) only fast-acting biological processes can be evolved, as the M13 phage only replicates stably in live E. coli for a short while, and the biological process selected for must take place before phage replication ceases; (iii) there is an intrinsic limit on rate of mutagenesis, since mutagenesis relies on error-prone DNA replication machinery that is also required for host cell replication, and the error rate is therefore limited to what is viable for the E. coli genome; and (iv) there is a need for specialised biorcactors to maintain an optimal balance of M13 phage concentration and E. coli cell density to enable evolution to occur.
[0007] Continuous evolution in E. coli is also possible using a Targeted Artificial DNA Replisome (TADR) which is a synthetic replisome containing a DNA nickase, a helicase and an error- prone DNA polymerase. TADR can achieve high replication error rates, but only over a short section of DNA (up to about 1600 bp) — the rest of the DNA replicates at a wild-type error rate. Evolution of larger genes and multiple genes is therefore impossible. There is thus a need for an orthogonal DNA replication system which can improve the rate and efficiency of in vivo continuous directed evolution for discovery of novel protein and nucleic acid variants, and which can be readily adapted for use in a variety of host cells.
[0008] It would be desirable to overcome or alleviate at least one of the above-described problems, or at least to provide a useful alternative.
[0009] Summary
[0010] Disclosed here is a method for replicating DNA, comprising contacting a DNA molecule with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication (T7 ori) and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, and wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori.
[0011] Disclosed here is a method for DNA mutagenesis, comprising contacting a DNA molecule with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori, and wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate.
[0012] Disclosed herein is method for evolving a target nucleic acid sequence, the method comprising contacting a DNA molecule comprising the target sequence with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori, and wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate.
[0013] Disclosed here is a kit for DNA replication, comprising: (a) a set of proteins comprising a DNAP, an RNAP, a SSBP, a HP, and one or more linearising proteins, wherein the set of proteins is capable of replicating a DNA molecule at a T7 ori in the DNA molecule; (b) a nucleic acid molecule encoding the set of proteins of (a); or (c) a host cell comprising the nucleic acid molecule of (b).
[0014] Disclosed herein is a kit for DNA mutagenesis, comprising: (a) a set of proteins comprising a DNAP, an RNAP, a SSBP, a HP, and one or more linearising proteins, wherein the set of proteins is capable of replicating a DNA molecule at a T7 ori in the DNA molecule, and wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate; or (b) a nucleic acid molecule encoding the set of proteins of (a); or (c) a host cell comprising the nucleic acid molecule of (b).
[0015] Disclosed herein is a DNA molecule comprising a T7 origin of replication (T7 ori) and one or more nucleic acid sequences capable of binding to a linearising protein, wherein the one or more nucleic acid sequences capable of binding to the linearising protein comprises a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 20-30
[0016] Brief description of the drawings
[0017] Embodiments of the present invention will now be described, by way of non-limiting example, with reference to the drawings in which:
[0018] FIG. 1 shows T7 repli some-dependent replication of plasmid with T7 origin in a cell-free system. (A) E. coli cell-free extract containing four T7 proteins (DNAP, RNAP, HH, SSB) was added to a plasmid carrying the T7 replication origin sequence. After incubation, the amount of T7ori plasmid was quantified by qPCR. (B) Expression cassettes used to express the four T7 proteins. (C) T7ori plasmid replicates in presence of all four T7 proteins but not in absence of any one or all of them. FIG. 2 shows that linear plasmid conformation resolves T7 origin plasmid separation in vivo. (A) Schematic showing maintenance of linear conformation of T7ori plasmid containing TelRL sequence in presence of protelomerase TelN. Replication of the linear conformation allows plasmids to be resolved, leading to normal cell division, while replication of the circular plasmid results in catenanes that prohibit plasmid separation. (B) The T7ori plasmid containing TelRL is efficiently transformed into E. coli expressing TelN and four T7 proteins (DNAP, RNAP, HP, SSB). In the absence of any one of the five proteins, no colonies were observed after transformation. (C) Conformation of the T7ori- plasmid by PCR. Primer 1 gives a band at 750 bp when T7ori plasmid is present, and Primer 2 gives a band at 1100 bp when T7ori-plasmid is present. Primer 3 amplifies across the TclLR site and would only give a band at 850 bp when the plasmid not cleaved by TclLR and remains circular, while it will not give a band if TelN cleaves the TelLR and the plasmid is linear. (D) The copy number of the T7ori plasmid containing TelRL was determined by qPCR together with common E. coli plasmid origins.
[0019] FIG. 3 shows that T7ORep selectively increases on-target mutation rates for mutagenesis.
[0020] (A) Reversion assay with an ampicillin resistance gene with an internal stop codon. After propagation for 10-30 generations and normalisation to total cell count, the frequency of reversion was scored. The same experiment was conducted in E. coli expressing an error- prone T7 DNAP and the T7 replisome (RNAP, HH, and SSB), as well as TelN. No significant change in reversion was observed. Finally, the ampicillin gene was carried on our T7ori plasmid in the E. coli with epDNAP, leading to a 23-fold increase in reversion rate.
[0021] (B) Data for (A).
[0022] FIG. 4 shows the average GFP intensity of the 5 best clones (of 100,000 events) obtained from directed evolution of dEGFP using the T7ORep system. dEGFP was cloned into a T7ori plasmid and propagated for 3 days in E. coli expressing the T7 replisome with an error- prone DNAP. By Fluorescence- Activated Cell Sorting (FACS), the top 1% most fluorescent cells were isolated, and the 5 most fluorescent ones were characterised by flow cytometry.
[0023] Detailed description There has been limited success in replicating plasmids containing the T7 origin of replication (T7 ori) in vivo in E. coli. Possible bottlenecks include poor folding of the replisome proteins, T7 replisome inactivation by endogenous E. coli proteins and targeted DNA degradation. The inventors found that four proteins in the T7 replisome are sufficient for in vitro replication of the T7 ori: a DNA polymerase (DNAP), an RNA polymerase (RNAP), a helicase-primase (HP), and a single-strand binding protein (SSBP). Furthermore, it was found that plasmid resolution is the bottleneck that prevents successful replication of circular T7 ori plasmids in E. coli hosts. By introducing a system that linearises circular DNA and maintains the replicated DNA in linear form, the inventors showed that plasmids with a T7 ori can be replicated in E. coli. In one example, the addition of a protelomerase recognition sequence (TelRL) to a T7 ori plasmid, in combination with the use of a protelomerase during DNA replication, enabled linearisation, replication and maintenance of a T7 ori plasmid in E. coli. Advantageously, the T7-based system can be used for DNA replication in a host cell orthogonal to the host replication system. The system can also be used for cell-free replication of DNA containing a T7 ori in vitro.
[0024] By using an error-prone T7 DNA polymerase, the inventors further showed that the T7- based replication system can increase mutagenesis of genes in a T7 ori plasmid, while other plasmids arc replicated at a natural error rate in E. coli. Because the present system replicates the T7 ori plasmid only (and not host endogenous DNA), there is no intrinsic limit on the error rate, hence molecular evolution can be accelerated. The present replication system can also accommodate much larger genes and even multi-gene constructs such as operons, multigene pathways and multi-gene protein complexes. The size of the construct may be up to 40 kb (the size of T7 phage genome). Directed evolution using the present orthogonal replication system is thus advantageous over existing systems for evolution such as PACE.
[0025] Accordingly, this disclosure provides methods for replicating and mutagenising DNA with a T7 origin of replication (T7ori) using a replisome comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, and wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori. Also provided arc compositions for carrying out the methods. Disclosed herein is a method for replicating DNA, comprising contacting a DNA molecule with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication (T7 ori) and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, and wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori.
[0026] Disclosed herein is a method for DNA mutagenesis, comprising contacting a DNA molecule with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori, and wherein one or more of the proteins in the set of proteins is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate.
[0027] Disclosed herein is a method for evolving a target nucleic acid sequence, the method comprising contacting a DNA molecule comprising the target sequence with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori, and wherein one or more of the proteins in the set of proteins is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate.
[0028] General definitions
[0029] As used herein, a “linearising protein” refers to a protein that is capable of binding to and / or processing a double-stranded DNA molecule to maintain it in a linear conformation, but excludes proteins that maintain or change the topological state of DNA, such as helicases, topoisomerases, gyrascs, and the like. The protein may keep the DNA in an open linear form (i.e., with free ends) or in a closed linear form (i.e., with ends ligated). A linearising protein generally recognises and binds to a cognate nucleic acid sequence on a DNA molecule. Such a sequence may be within the DNA molecule or at the ends of the DNA molecule. The protein may bind to the DNA covalently or non-covalently. The protein may bind as a monomer, multimer, or as part of a protein complex. The linearising protein may process the DNA within or proximal to its cognate nucleic acid sequence. For example, the protein may introduce single- or double-strand breaks, add or remove nucleotides, or ligate DNA within or proximal to its cognate nucleic acid sequence. Alternatively or additionally, the linearising protein may stabilise the ends of open linear DNA to prevent the ends from circularising or being ligated. A linearising protein may be involved in DNA replication, for example, to maintain replicated DNA in a linear conformation.
[0030] The term “rcplisomc” refers to the set of proteins capable of carrying out DNA replication. A replisome generally includes a protein with duplex separation functionality (e.g., a helicase), a protein capable of preventing re-annealing of the separated duplex (e.g., a singlestrand binding protein), and a protein with polymerase activity (e.g., a DNA or RNA polymerase). A replisome may also include unwinding proteins (e.g., gyrases or topoisomerases), primases (for generating a RNA primer), nickases, clamp factors, exonucleases, ligases, protelomerases and / or terminal proteins. A “T7 replisome” refers a set of proteins capable of replicating a nucleic acid molecule containing a T7 origin of replication.
[0031] As used herein, an “origin of replication” refers to a specific DNA sequence in the genome of prokaryotes, eukaryotes, viruses and plasmids where DNA replication is initiated. The origin of replication serves as the starting point for assembly of the replisome. A “T7 origin of replication” or “T7 ori” refers to an origin of replication derived from the genome of the T7 phage.
[0032] A “protelomerase” or “telomere resolvase” herein is any polypeptide capable of cleaving and rejoining DNA comprising a protelomerase cognate nucleic acid sequence in order to produce a covalently closed linear- DNA molecule. Thus, a protelomerase has DNA cleavage and ligation functions. A typical substrate for pro telomerase is circular double stranded DNA. If this DNA contains a protelomerase target site, the enzyme can cut the DNA at this site and ligate the ends to create a linear double stranded covalently closed DNA molecule. A “terminal protein” herein is any polypeptide capable of covalent association with a cognate nucleic acid sequence at an end of a nucleic acid molecule, typically the 5’ end. Terminal proteins function to initiate replication of linear DNA (such as linear viral genomes or linear plasmids), and may also function to stabilise the DNA ends, for example, by influencing the conformation of the DNA ends or preventing them from being recognised as damaged DNA.
[0033] As used herein, the term “fidelity” refers to the accuracy of DNA or RNA polymerisation by a template-dependent DNA or RNA polymerase. The fidelity of a polymerase is typically measured by the error rate, i.e., the frequency of incorporating an inaccurate nucleotide (i.e., a nucleotide that is not incorporated in a template-dependent manner). The error rate or mutation rate of a polymerase may be expressed as the number of inaccurate nucleotides per base per replication. This is typically denoted as a fraction, such as 10’6errors per base. Fidelity is the inverse of the error rate, i.e., a lower error rate indicates higher fidelity.
[0034] As used herein, an “error-prone polymerase” refers to a DNA or RNA polymerase that exhibits an increased error rate (i.e., increased misincorporation of nucleotides) during nucleic acid synthesis compared to high-fidelity polymerases. The error-prone polymerase may be naturally-occurring or an engineered polymerase. The increased error rate may result from, for example, reduced proofreading (3’^-5’ exonuclease) activity or relaxed basepairing stringency.
[0035] As used herein, a “nucleoside deaminase” is an enzyme that catalyses the deamination of aminated nucleosides, and includes enzymes that deaminate either or both ribo- and deoxyribonucleosides. Deamination reactions occur in the nucleobase moiety, including cytosine, 5-mcthylcytosinc, guanine and adenine nucleosides, which arc transformed into their corresponding nucleoside analogues containing, respectively, uracil, thymine, xanthine and hypoxanthine as the nucleobases. Adenosine deaminases deaminate (deoxy)adenosine to give (deoxy)inosine. Cytidine deaminases deaminate (deoxy)cytidines to give (deoxy )uridine.
[0036] As used herein, the term “wild-type” and “parent’ ’ are used interchangeably to refer to a gene or gene product (c.g., protein) that has the characteristics (c.g., sequence) of that gene or gene product isolated from a naturally occurring source. In contrast, the term “mutant” or “variant” refers to a gene or gene product that displays modifications in sequence when compared to the wild-type gene or gene product. Mutant genes or gene products may be naturally-occurring or synthetic (i.e., containing altered sequences that do not occur in nature).
[0037] The term “construct” generally refers to recombinant nucleic acid, generally recombinant DNA, that has been generated for the purpose of the expression of a specific nucleic acid sequence (i.e., an “expression construct”), or is to be used in the construction of other recombinant nucleotide sequences (i.e., a “cloning construct”). The construct may be contained within a vector. Two or more constructs can be contained within a single nucleic acid molecule, such as a single vector, or can be contained within two or more separate nucleic acid molecules, such as two or more separate vectors.
[0038] An “expression construct” generally includes at least an element operably linked to a nucleotide sequence of interest to direct expression of the nucleic acid sequence in a host cell. Such elements may include control elements such as a promoter that is operably linked to direct transcription of the nucleic acid sequence of interest, and often include a polyadenylation sequence as well. For the practice of the present invention, conventional compositions and methods for preparing and using constructs and host cells arc well known to one skilled in the art, see for example, Molecular Cloning: A Laboratory Manual, 3rd edition Volumes 1, 2, and 3. J. F. Sambrook, D. W. Russell, and N. Irwin, Cold Spring Harbor Laboratory Press, 2000.
[0039] As used herein “operably linked” is the association of two or more nucleic acid fragments in a construct such that the function of one is controlled by the other, for example DNA encoding a protein associated with DNA encoding a promoter.
[0040] By “control element” or “control sequence” is meant nucleic acid sequences (e.g., DNA) necessary for expression of an operably linked coding sequence in a particular host cell. The control sequences that are suitable for prokaryotic cells for example, include a promoter, and optionally a cis-acting sequence such as an operator sequence and a ribosome binding site. Control sequences that are suitable for eukaryotic cells include transcriptional control sequences such as promoters, polyadcnylation signals, transcriptional enhancers, translational control sequences such as translational enhancers and internal ribosome binding sites (IRES), nucleic acid sequences that modulate mRNA stability, as well as targeting sequences that target a product encoded by a transcribed polynucleotide to an intracellular compartment within a cell or to the extracellular environment. Promoters suitable for use with expression constructs or vectors of the present disclosure include but are not limited to the phage lambda PL promoter, the E. coli lac, phoA and tac promoters, and the SV40 early and late promoters.
[0041] By “vector” is meant a nucleic acid molecule, suitably a DNA molecule derived, for example, from a plasmid, bacteriophage, yeast or virus, into which a polynucleotide can be inserted or cloned. A vector may contain one or more unique restriction sites or multiple cloning sites and can be capable of autonomous replication in a defined host cell including a target cell or tissue or a progenitor cell or tissue thereof, or be integrated with the genome of the defined host such that the cloned sequence is reproducible. Accordingly, the vector can be an autonomously replicating vector, i.e., a vector that exists as an extra-chromosomal entity, the replication of which is independent of chromosomal replication, e.g., a linear or closed circular plasmid, an extra-chromosomal element, a mini-chromosome, or an artificial chromosome. The vector can contain any means for assuring self-replication. Alternatively, the vector can be one which, when introduced into the host cell, is integrated into the genome and replicated together with the chromosomc(s) into which it has been integrated. A vector system may comprise a single vector or plasmid, two or more vectors or plasmids which together contain the total DNA to be introduced into the genome of the host cell, or a transposon. The choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector may also include a selection marker such as an antibiotic resistance gene that can be used for selection of suitable transformants. Examples of such resistance genes are well known to those of skill in the art.
[0042] The terms “host cell”, “host cell line” and “host cell culture” are used interchangeably and refer to cells into which exogenous nucleic acid has been introduced, including the progeny of such cells. Host cells include “transformants” and “transformed cells”, which include the primary transformed cell and progeny derived therefrom without regard to the number of passages. Progeny may not be completely identical in nucleic acid content to a parent cell but may contain mutations. Mutant progeny that have the same function or biological activity as screened or selected for in the originally transformed cell are included herein. Host cells include cultured cells, e.g., mammalian cultured cells, such as CHO cells, BHK cells, NSO cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells or hybridoma cells, yeast cells, insect cells, and plant cells, to name only a few, but also cells comprised within a transgenic animal, transgenic plant or cultured plant or animal tissue.
[0043] The term “sequence identity” as used herein refers to the extent that sequences are identical on a nucleotide-by-nucleotide basis or an amino acid-by-amino acid basis over a window of comparison. Thus, a “percentage of sequence identity” is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical nucleic acid base (e.g., A, T, C, G and I) or the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity. Sequence identity may be measured using sequence analysis software, for example, BLAST, ClustalW, ClustalOmega, MUSCLE, TCoffee or ProbCons).
[0044] Sequence variations may arise from conservative amino acid substitutions. A “conservative amino acid substitution” is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art, which can be generally sub-classified as follows:
[0045] Table A. Amino acid sub-classification.
[0046] Conservative amino acid substitution also includes groupings based on side chains. For example, a group of amino acids having aliphatic side chains is glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains is serine and threonine; a group of amino acids having amide-containing side chains is asparagine and glutamine; a group of amino acids having aromatic side chains is phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains is lysine, arginine, and histidine; and a group of amino acids having sulphur-containing side chains is cysteine and methionine. For example, it is reasonable to expect that replacement of a valine with a methionine, an aspartate with a glutamate, a threonine with a serine, or a similar replacement of an amino acid with a structurally related amino acid will not have a major effect on the properties of the resulting mutant polypeptide.
[0047] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (or).
[0048] As used in this application, the singular form “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “an agent” includes a plurality of agents, including mixtures thereof.
[0049] Throughout this specification and the claims which follow, unless the context requires otherwise, the word “comprise”, and variations such as “comprises” and “comprising”, will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.
[0050] Throughout this specification and the claims which follow, unless the context requires otherwise, the phrase “consisting essentially of’, and variations such as “consists essentially of’ will be understood to indicate that the recited element(s) is / are essential i.e. necessary elements of the invention. The phrase allows for the presence of other non-recited elements which do not materially affect the characteristics of the invention but excludes additional unspecified elements which would affect the basic and novel characteristics of the method defined.
[0051] T7 ori
[0052] Compositions and methods herein may be used to replicate DNA with a T7 origin of replication (T7 ori). The DNA molecule may be a linear' or circular DNA molecule. In certain embodiments, the DNA molecule comprises additional origins of replication orthogonal to the T7 ori that allow the DNA to be replicated by the endogenous replication system of a host cell.
[0053] In some embodiments, the T7 ori comprises a nucleic acid sequence having at least 70% sequence identity (such as about 71 %, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to a nucleic acid sequence set forth in SEQ ID NO: 1-6.
[0054] Table 1. Exemplary T7ori sequences.
[0055] Linearising proteins
[0056] Linearising proteins for use in methods and compositions herein may be capable of maintaining the replicated DNA in open linear form or closed linear form.
[0057] In one embodiment, the one or more linearising proteins is a protelomerase (also known as a telomere resolvase). Protelomerases maintain the replicated DNA in a closed linear conformation, which is advantageous for preventing attack by enzymes such as exonucleases, preventing the DNA from integrating with other DNA molecules present (such as genomic DNA), and preventing concatamerisation of DNA molecules. Examples of protclomcrascs include but arc not limited to E. coli phage N15 TclN, Klebsiella oxytoca phage cpKO2 TelK, Yersinia phage PY54 TelY, Vibrio phage VP882 gp54, Vibrio parahaemolyticus phage VP58.5 gp40, Halomonas phage cpHAP-1 gp34, Agrobacterium tumefaciens C58 Tel A, and Borrelia burgdorferi ResT.
[0058] In some embodiments, the protelomerase comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 7-14.
[0059] In one embodiment, the one or more linearising proteins is a terminal protein (TP). Terminal proteins maintain the replicated DNA in an open linear conformation. Examples of terminal proteins include but are not limited to B. subtilis phage q>29 terminal protein, Bacillus phage GA-1 terminal protein, Streptococcus phage Cp-1 terminal protein, enterobacteria phage PRD1 terminal protein, and human adenovirus C terminal protein. In some embodiments, the terminal protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 15-19.
[0060] In one embodiment, the linearising protein is a protelomerase comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 7 (N15 phage TclN). In one embodiment, the linearising protein is a protelomerase comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 8. In one embodiment, the linearising protein is a protelomerase comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 9. In one embodiment, the linearising protein is a protelomerase comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 10. In one embodiment, the linearising protein is a protclomcrasc comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 11. In one embodiment, the linearising protein is a protelomerase comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 12. In one embodiment, the linearising protein is a protelomerase comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 13. In one embodiment, the linearising protein is a protelomerase comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 14.
[0061] In one embodiment, the linearising protein is a terminal protein comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71 %, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 15. In one embodiment, the linearising protein is a terminal protein comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 16. In one embodiment, the linearising protein is a terminal protein comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 17. In one embodiment, the linearising protein is a terminal protein comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 18. In one embodiment, the linearising protein is a terminal protein comprising or consisting of an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 19. The protelomerase or terminal protein recognises and binds to a nucleic acid sequence in the DNA to maintain the DNA in linear form. The binding sequence may comprise an inverted repeat sequence, i.e., a double- stranded DNA sequence having two-fold rotational symmetry, also described herein as a palindromic sequence. The length of the inverted repeat differs depending on the protelomerase or terminal protein. The palindrome or inverted repeat may be perfect or imperfect. For example, the protelomerase TelN from N15 phage recognises a specific nucleotide sequence termed TelRL, which is a slightly imperfect inverted palindromic structure comprising two halves, TelR and TelL, flanking a 22 base pair perfect inverted repeat TelO.
[0062] In some embodiments, the one or more sequences capable of binding to the linearising protein comprises a nucleic acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to a nucleic acid sequence set forth in SEQ ID NO: 20-30.
[0063] In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 20. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 21. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71 %, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 22. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 23. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%), 94%;, 95%;, 96%, 97%;, 98%), or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 24. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 25. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 26. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%), 94%;, 95%;, 96%, 97%;, 98%), or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 27. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70%> sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 28. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 29. In one embodiment, the one or more sequences capable of binding to the linearising protein comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%), 94%;, 95%;, 96%, 97%;, 98%), or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 30.
[0064] Table 2. Exemplary linearising proteins
[0065] Table 3. Exemplary sequences recognised by linearising proteins
[0066] Replisome proteins
[0067] The set of replisome proteins may comprise any DNA polymerase (DNAP), RNA polymerase (RNAP), single-strand binding protein (SSBP) and helicase-primase (HP) capable of replicating the T7 origin of replication (T7ori).
[0068] In one embodiment, the DNA polymerase in the replisome is a T7 DNAP, or a homologue of a T7 DNAP. In one embodiment, the T7 DNAP comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%), 97%., 98%), or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 31.
[0069] In one embodiment, the RNA polymerase in the replisome is a T7 RNAP, or a homologue of a T7 RNAP. In one embodiment, the T7 RNAP comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 32.
[0070] In one embodiment, the single strand binding protein in the replisome is a T7 SSBP, or a homologue of a T7 SSBP. In one embodiment, the T7 SSBP comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 33.
[0071] In one embodiment, the helicase-primase in the replisome is a T7 HP, or a homologue of a T7 HP. In one embodiment, the T7 HP comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence in SEQ ID NO: 34.
[0072] Homologues of T7 DNAP, RNAP, SSBP and HP are identifiable by a skilled person by comparison of amino acid sequences, e.g., manually or by using homology-based search algorithms such as those commonly known and referred to as BLAST, Smith- Waterman, Needleman-Wunsch, MMseqs2, and DIAMOND. A local sequence alignment program, e.g. BLAST, can be used to search a database of sequences to find similar sequences, and the summary expectation value (E-value) used to measure the sequence base similarity.
[0073] In some embodiments, the set of proteins comprises one or more proteins selected from a T7 DNAP, a T7 RNAP, a T7 SSBP and a T7 HP.
[0074] In one embodiment, the set of proteins comprises at least two proteins selected from a T7 DNA polymerase, a T7 RNA polymerase, a T7 single-strand binding protein, and a T7 helicase-primase. In one embodiment, the at least two proteins comprises a T7 DNAP and a T7 RNAP.
[0075] In one embodiment, the set of proteins comprises a T7 DNAP, a T7 RNAP, a T7 SSBP, a T7 HP.
[0076] Table 4. Exemplary DNAPs, RNAPs, SSBPs and HPs.
[0077] One or more of the proteins in the set of proteins may be a mutant protein. Thus, for example, the DNAP, RNAP, SSBP, HP and / or the one or more linearising proteins may be a mutant protein. The mutant protein may be a naturally occurring mutant or an engineered protein. The mutant protein may have improved activity or an additional or alternative functionality compared to a wild-type protein. For instance, a mutant DNAP or HP may exhibit improved processivity during DNA replication. A mutant DNAP may be an error-prone DNAP having a replication error rate that is above the replication error rate of a wild-type DNAP. One or more of the proteins in the set of proteins may be a fusion protein. Thus, for example, the DNAP, RNAP, SSBP, HP and / or the one or more linearising proteins may be a fusion protein. The fusion protein may comprise a sequence (such as a peptide, protein fragment or protein domain) which is not present in the wild-type DNAP, RNAP, SSBP, HP or linearising protein. In some embodiments, the fusion protein comprises sequences derived from two or more different proteins. The two or more different proteins may have different functionalities.
[0078] Mutagenic polymerases
[0079] The set of rcplisomc proteins may comprise a mutagenic polymerase to increase the error rate of DNA replication. This allows DNA mutagenesis during DNA replication and may be advantageous in continuous directed evolution systems.
[0080] The mutagenic polymerase may be a DNAP with an increased replication error rate compared to a natural replication error rate (such as the error rate of a corresponding wildtype parent polymerase). In some embodiments, the mutagenic polymerase is a mutant and / or fusion DNAP.
[0081] In one embodiment, the mutant and / or fusion DNAP is an error-prone polymerase. The error-prone DNAP may be a naturally-occurring or engineered DNAP. In some embodiments, the error-prone DNAP has reduced or abolished proofreading activity, for example, the DNAP may be absent an exonuclease domain for proofreading, or the exonuclease domain may contain one or more inactivating mutations. In some embodiments, the error-prone DNAP has reduced substrate selectivity and / or base-pairing stringency, for example, through the introduction of mutations in the substrate recognition domain. In some embodiments, the error-prone DNAP is capable of trans-lesion synthesis, i.e., the DNAP is able to bypass lesions or distortions in the DNA template, such as the presence of thymine dimers, oxidised or alkylated bases, or abasic sites.
[0082] The error-prone polymerase may be capable of increasing the DNA replication error rate by about 20%, about 50%, about 100%, about 2-fold, about 5-fold, about 10-fold, about 20- fold, about 50-fold, about 100-fold, about 200-fold, about 500-fold, about 1000-fold, about 104-fold, about 105-fold, about 106-fold, about 107-fold, about 108-fold, or more than 108- fold above a natural replication error rate. The natural replication error rate may be the error rate of a corresponding wild-type DNAP, the error rate of a DNAP endogenous to the host cell, or the background mutation rate of the host cell. For example, the background mutation rate of spontaneous base-pair substitutions in E. coli is reported to be about IO"10mutations per base per generation.
[0083] In some embodiments, the error-prone DNAP is characterised by a mutation rate that is at least about 102-fold higher than the background mutation rate of the host cell. In some embodiments, the error-prone DNAP is characterised by a mutation rate that is about 102- to about 108-fold higher than the background mutation rate of the host cell. A skilled person may select an error-prone DNAP with an error rate that produces a desired rate of mutagenesis in a target DNA sequence.
[0084] Assays to determine fidelity or error rate of DNA replication are known in the art. One exemplar}' assay is the Luria-Delbriick fluctuation assay. In the fluctuation assay, a premature stop codon is added to the coding sequence of a selectable marker, e.g., an antibiotic resistance gene. Replication using an error-prone polymerase allows reversion of the stop codon to a sense codon, enabling growth on the antibiotic. By counting the number of viable colonies and applying statistical methods, such as the Poisson distribution, the mutation rate can be estimated. The fluctuation assay is suitable for detecting low mutation rates given the high sensitivity of the assay. Another exemplar}' assay is the LacZ inactivation assay. In this assay, the polymerase variant is used to replicate a LacZ gene on a plasmid. A higher mutation rate increases the rate of introducing deleterious mutations to the LacZ gene, thereby inactivating the LacZ gene. Growth of bacteria on agar containing X-gal can determine rate of inactivation by blue-white colony counting (with blue colonics having functional LacZ and white colonies having inactivated LacZ). This assay is generally less sensitive and usually used when the mutation rate is high.
[0085] In one embodiment, the error-prone DNAP is an engineered T7 DNAP. In one embodiment, the error-prone DNAP comprises an amino acid sequence having at least 70% sequence identity (such as about 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence set forth in SEQ ID NO: 39. In one embodiment, the error-prone DNAP comprises a nucleoside deaminase domain. The coupling of a nucleoside deaminase to the DNAP allows the introduction of random basepair changes during DNA replication, which can increase the error rate of replication. For example, adenosine deaminases can convert A-T base pairs to G-C, and cytidine deaminases can convert C-G base pairs to T-A. The increase in the error rate provided by deaminase conjugation may further contribute to the mutagenicity of an error-prone DNAP. The nucleoside deaminase domain may be conjugated to the N- or C-terminus of the DNAP, optionally via a linker.
[0086] In some embodiments, the nucleoside deaminase is an adenosine or a cytidine deaminase. The adenosine deaminase or cytidine deaminase may be capable of deaminating deoxynucleosides, or both nucleosides and deoxynucleosides.
[0087] In some embodiments, the adenosine deaminase is TadA-7.10, TadA-8e, or a variant thereof. TadA-7.10 is described in Gaudelli, N. M. et al. Nature 551, 464-471 (2017), and TadA-8 is described in Richter, M. F. et al. Nat. Biotechnol. 38, 883-891 (2020). Both references are incorporated by reference in their entirety herein. In some embodiments, the cytidine deaminase is AID, APOBEC1, APOBEC3A, APOBEC3G, PmCDAl, or a variant thereof.
[0088] Variants contemplated herein include natural and engineered enzymes with improved DNA- binding affinity, catalytic activity, substrate range, overall mutagenic activity, reduced off- target activity (i.e., deamination activity that is not in conjunction with polymerase activity), or smaller size. Exemplary TadA-8e variants are described in Neugebauer, M. E. et al., Nat. Biotechnol. 41, 673-685 (2023).
[0089] In one embodiment, the nucleoside deaminase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 35-38. The base deaminase may comprise a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a nucleic acid sequence set forth in SEQ ID NO:35-38.
[0090] In some embodiments, the mutant and / fusion polymerase is the only DNAP present in the set of replisome proteins. In some embodiments, the replisome comprises both a wild-type and a mutant / fusion DNAP. In yet other embodiments, the replisome may comprise two or more mutant / fusion DNAPs, such as DNAPs with different replication error rates. The use of multiple DNAPs with different replication fidelities or error rates may allow tuning of the rate of mutagenesis during DNA replication. Where multiple DNAPs are used, expression of the DNAPs may be inducible under different conditions to control which DNAP is active at a given time.
[0091] TadA-8e
[0092] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRV1GEGWNRA1GLHDP TAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRN SKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKA QSSIN (SEQ ID NO: 35)
[0093] PmCDA!
[0094] MMTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYA VNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYN QELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKI FIQSSHNQLNENRWLEKTLKRAEKWRSELSIMIQVKILHTTKSPAVS (SEQ ID NO: 36)
[0095] APOBEC1
[0096] MGSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTS QNTNKHVEVNF1EKFTTERYFCPNTRCS1TWFLSWSPCGECSRA1TEFLSRYPHVTL FIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWP RYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWA TGLK (SEQ ID NO: 37)
[0097] AID
[0098] MGDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNK NGCHVELLFLRY1SDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSL RIF ARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKA WEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 38) Exemplary error-prone T7 DNAP
[0099] MIVSDIE AN ALLES VTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALEAEVARGGLI VFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVLSRLIHSNLKDTDMGLLRS GKLPGKRLGSHALEAWGYRLGEMKGEYKDDFKRMLEEQGEEYVDGMEWWNFN EEMMDYNVQDVVVTKALLEKLLSDKHYFPPEIDFTDVGYTTFWSESLEAVDIEHR AAWLLAKQERNGFPFDTKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMF CHPRTGKPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVE HVVFNPSSRDH1QKKLQEAGWVPTKYTDKGAPVVDDEVLEGVRVDDPEKQAA1D LIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPNGAVTGRATHAFPNLA Q1PGVRSPYGEQCRAAFGAEHHLDG1TGKPWVQAG1DASGLELRCLAHFMARFDN GEYAHEILNGDIHTKNQIAAELPTRDNAKTFIYGFLYGAGDEKIGQIVGAGKERGK
[0100] ELKKKFLENTPAIAALRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHV RSPHAALNTLLQSAGALTCKLWTTKTEEMLVEKGLKHGWDGDFAYMAWVHDETQ VGCRTEEIAQVVIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH (SEQ ID NO: 39)
[0101] Methods for DNA replication and mutagenesis
[0102] Methods herein may be used to replicate linear or circular DNA molecules.
[0103] In some embodiments, DNA replication occurs in vitro in the absence of cells. The set of proteins may be provided in a cell lysate (e.g., lysate from a cell expressing the set of replisome proteins). Alternatively, the purified or substantially pure proteins may be used for in vitro DNA replication.
[0104] In some embodiments, DNA replication occurs in vivo in a host cell. A variety of host cell types can be used, including prokaryotic hosts, eukaryotic hosts, and cell lines. Eukaryotic hosts may be particularly useful for evolving genes, proteins, and pathways that cannot be expressed or performed effectively in prokaryotes.
[0105] In some embodiments, the host cell is a prokaryotic cell. In one embodiment, the prokaryotic cell is a bacterial cell. In one embodiment, the bacterial cell is an E. coli cell. Suitable E. coli hosts include, without limitation, the K-12, MG 1655, BL21, AD494, Origami, HMS174, BLR, HMS174, Tuner, Rosetta, Lemo21, NiCo2, T7 Express, Shuffle Express, C41, C43, and ml 5 pREP4 strains, and derivatives thereof, such as DE3 prophage-lyosogenised strains. These E. coll strains are widely available commercially.
[0106] In some embodiments, the host cell is a eukaryotic cell. In one embodiment, the eukaryotic cell is a fungal cell, such as a yeast cell. In one embodiment, the eukaryotic cell is an insect cell. In one embodiment, the eukaryotic cell is a plant cell. In one embodiment, the eukaryotic cell is an algal cell, such as a microalgal cell. In one embodiment, the eukaryotic cell is a mammalian cell.
[0107] In some embodiments, the host cell expresses a thioredoxin. Thioredoxin acts an accessory protein for T7 DNAP. The thioredoxin may be, for example, E. coli thioredoxin, or a homologous protein. In one embodiment, the thiorcdoxin comprises or consists of an amino acid sequence having at least 80% sequence identity (such as about 80%, 81 %>, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or greater sequence identity) to an amino acid sequence set forth in SEQ ID NO: 40.
[0108] One or more components of the replisome may be encoded by a nucleic acid molecule (e.g., a vector, plasmid or expression construct) in the host cell. In some embodiments, the nucleic acid molecule is an cxtrachromosomal clement that can be replicated by the endogenous DNA replication system of the host cell. In other embodiments, one or more components of the replisome is encoded by the DNA molecule to be replicated, or by a DNA molecule that is compatible with the T7 replisome. In some embodiments, one or more components of the replisome is encoded in the host cell genome. For example, the RNAP may be a T7 RNAP that is encoded by a E3 lysogen in an E. coli host. Introduction of nucleic acids into host cells and integration of nucleic acids into host cell genomes may be performed using methods widely known in the art.
[0109] In some embodiments, the set of replisome proteins is capable of replicating the DNA molecule orthogonal to an endogenous DNA replication system in the host cell. An orthogonal DNA replication system is less capable of replicating host-endogenous DNA compared to the host native DNA replication system. This reduced ability means that the orthogonal replication system can be used to mutate DNA that is compatible with the orthogonal system, without significantly increasing the mutation rate of the host cell genome. Thus, methods herein may be used for DNA mutagenesis and for evolving target nucleic acid sequences (e.g., protein coding sequences) in TV-compatible DNA molecules.
[0110] Target nucleic acid sequences for mutagenesis or directed evolution may be sequences that affect the fitness of the host cell during culture, and thus can be selected for. The target sequence may, for example, encode multi-gene metabolic pathways, enzymes or protein complexes required for cell metabolism, growth and / or proliferation. Preferably the target genes or pathways produce activities or products which occur or accumulate in the cell or in the culture medium over a period of time. For example, it may be anticipated that mutations in the target sequence can increase the fitness of the host during culture. Thus, the culture conditions may be such as to enable only sub-optimal growth of the host when the target sequence is not mutated. Mutation of the target sequence may lead to an optimisation or acquisition of a particular function (or conversely a reduction or loss of a function) in the host cell that increases host fitness under the culture conditions (i.e., the culture conditions impose a selection pressure for the evolution of target sequences beneficial to the fitness of the host). The host cells may be cultured and selected iteratively, for example, using increasingly stringent culture conditions that select for the most desirable mutations.
[0111] Alternatively, the target sequence may be a regulatory sequence or a reporter gene, and host cells may be selected based on an increase or decrease in regulatory or reporter function following mutagenic DNA replication.
[0112] There is no particular limitation to the size of the target sequence that can be replicated and evolved. In certain embodiments, the target sequence may be as large as about 40 kb, which is the approximate size of the T7 phage genome. Methods herein are particularly well- suited for evolving large genes, such as those encoding large enzymes, and multi-gene constructs (encoding, for example, metabolic pathways or protein complexes) where the individual gene components need to evolve in tandem.
[0113] Exemplary thioredoxin from E. coli
[0114] MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTV AKLN1DQNPGTAPKYG1RG1PTLLLFKNGEVAATKVGALSKGQLKEFLDANLA (SEQ ID NO: 40) Vectors and kits
[0115] This disclosure also provides vectors and kits comprising or encoding the set of T7 replisome proteins.
[0116] Disclosed herein is a DNA molecule comprising a T7 origin of replication (T7 ori) and one or more nucleic acid sequences capable of binding to a linearising protein. The linearising protein binding sequence may comprise a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 20-30.
[0117] The DNA molecule may be a lineal' or circular DNA molecule. In some embodiments, the DNA molecule is a vector, such as a plasmid. The DNA molecule may comprise additional origins of replication that are compatible with the endogenous replication system of a suitable host cell. Furthermore, the DNA molecule may comprise a marker gene or a resistance gene for selection of host cells containing the DNA molecule.
[0118] Disclosed herein is a kit for DNA replication, comprising: (a) a set of proteins comprising a DNAP, an RNAP, a SSBP, a HP, and one or more linearising proteins, wherein the set of proteins is capable of replicating a DNA molecule at a T7 ori in the DNA molecule; (b) a nucleic acid molecule encoding the set of proteins of (a); or (c) a host cell comprising the nucleic acid molecule of (b). The nucleic acid molecule may be genomic DNA in the host cell or an extrachromosomal entity, such as a plasmid or vector. The host cell may be prokaryotic or eukaryotic. The host cell may be a cell line.
[0119] Also disclosed herein is a kit for DNA mutagenesis, comprising: (a) a set of proteins comprising a DNAP, an RNAP, a SSBP, a HP, and one or more linearising proteins, wherein the set of proteins is capable of replicating a DNA molecule at a T7 ori in the DNA molecule, and wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate; (b) a nucleic acid molecule encoding the set of proteins of (a); or (c) a host cell comprising the nucleic acid molecule of (b). The nucleic acid molecule may be genomic DNA in the host cell or an extrachromosomal entity, such as a plasmid. In some embodiments, the kit comprises a DNA molecule containing a T7 origin of replication (T7 ori) and one or more nucleic acid sequences capable of binding to a linearising protein.
[0120] The kit may further contain reagents and buffers, such as nucleotides and divalent metal ions, required by one or more of the replisome proteins for DNA replication.
[0121] The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms pail of the common general knowledge in the field of endeavour to which this specification relates.
[0122] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. It is to be understood that the invention includes all such variations and modifications, which fall within the spirit and scope. The invention also includes all of the steps, features, compositions and compounds referred to or indicated in this specification, individually or collectively, and any and all combinations of any two or more of said steps or features.
[0123] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0124] Certain embodiments of the invention will now be described with reference to the following examples which arc intended for the purpose of illustration only and arc not intended to limit the scope of the generality hereinbefore described.
[0125] EXAMPLES
[0126] Example 1: Orthogonal T7-based replication system in E. coli
[0127] T7 replisome-dependent plasmid replication in a cell-free extract of E. coli It was verified that the four T7 proteins are necessary and sufficient for the replication of a T7ori plasmid in an E. coll cell-free extract (Fig. 1). The cell-free extracts contained all the cellular machinery of the host but without membrane compartmentalisation. Such extracts are increasingly employed in molecular and synthetic biology, as they provide a unique opportunity to take advantage of cellular processes in their native environment, without the need to keep cells alive. The use of cell-free extracts also bypass the need for replicated plasmids to resolve (separate) to allow cell division under selection for the plasmid.
[0128] T7 DNA polymerase (DNAP), RNA polymerase (RNAP), helicase-primase (HP), and single strand-binding protein (SSB) were expressed and, after incubation, the amount of T7ori plasmid was quantified by qPCR (Fig. 1A, B). Successful replication of the T7ori plasmid was evident from a significant increase in plasmid concentration in presence of all four T7 replisome proteins (Fig. 1C). However, when all or any one of the proteins was omitted, no replication was detected. Instead, the plasmid concentrations dropped, presumably due to degradation by DNases in the cell-free extract. Hence, it was concluded that T7 DNAP, RNAP, HH, and SSB are necessary and sufficient for replication of T7ori plasmids in E. coli.
[0129] Linear conformation enables a T7ori plasmid to resolve upon replication in vivo
[0130] In E. coli, the genome and most plasmids are maintained in a circular conformation, resulting in interlocked catenanes or multimer formation upon replication, that are subsequently resolved by endogenous DNA resolution systems to allow separation of replicated plasmids for segregation into daughter cells. In contrast, the bacteriophage T7 genome is maintained in a linear conformation, and it is therefore quite possible that there is no endogenous mechanism to resolve catenanes that are formed through T7 replication (Fig. 2A). To explore if an inability to resolve a circular plasmid conformation is the bottleneck for replication of T7ori plasmids, the only linear plasmid system known to be maintained E. coli (the phage- derived protelomerase TelN and the 56 bp TelN recognition DNA sequence, TelRL) was investigated. Upon binding to TelRL, TelN performs a double-strand cut and seals the ends together to create two hairpin ends.
[0131] The TelRL sequence was cloned into the T7ori plasmid to allow linearisation (Fig. 2A). To facilitate cloning and handling of this plasmid, the conditional origin oriR6K that exclusively replicates in a pir-expressing strain was also included. After cloning and propagation in a pir+ host (the pir protein is strictly required for replication of oriR6K), the plasmid was transformed into pir- cells expressing TelN and all or some of the T7 replisome components (Fig. 2B). In the presence of all replisome components, the transformation led to the formation of thousands of colonies, while absence of any one component entirely abolished the formation of viable colonies. Further, PCR was employed to confirm that the T7ori plasmid is circular in the pir-i- cloning strain where TelN is absent, while it adopts a linear conformation in the E. coli strain with TelN and the T7 replisome (Fig. 2C). These results demonstrate i) that replication of a T7ori plasmid is possible in E. coli if the plasmid maintains a lineal' conformation, and ii) that all components of the anticipated T7 replisome (and TelN in the present setup) are necessary to replicate and maintain the plasmid. It further supports the hypothesis that replication of circular T7ori plasmids fails due to unresolved catenane formation. Finally, the copy-number of the T7ori plasmid in E. coli was determined by qPCR (Fig 2D). With an average copy-number of 7.96 ± 2.55, its abundance is comparable to that of plasmids carrying the common pl5a origin.
[0132] An orthogonal plasmid system in E. coli
[0133] With this first demonstration of a functional T7-derived plasmid replication system in E. coli, it was next asked if the plasmid replicated orthogonally in E. coli. A reversion assay was employed, where an ampicillin resistance gene with an internal stop codon (ampTAA) was evolved and subsequently selected for to determine the rate of stop codon reversion. The ampTAA gene was placed on the T7ori plasmid and a p!5a ori plasmid due to their comparable copy number. The plasmids were propagated in E. coli for 10-30 generations in absence of ampicillin, then plated under ampicillin selection, and the emerging colonies were scored (Fig. 3). Next, the experiment was repeated in a strain expressing an error-prone mutant of the T7 DNAP (epDNAP), the rest of the T7 replisome (RNAP, HH, and SSB) and TelN. The number of resistant colonies obtained did not differ between the experiments, implying that the T7 replisome, specifically the T7 DNAP, did not engage in replication of the pl5a plasmid. Finally, the ampicillin resistance gene with stop codon was installed in the T7ori plasmid and the above experiment was repeated (Fig. 3). This led to a 23-fold increase in stop codon reversion compared to the pl5a plasmids, demonstrating i) that the T7ori plasmid is replicated by the T7 DNAP, and ii) that this plasmid can be propagated with significantly higher error-rate than other plasmids in E. coli. The new orthogonal plasmid system was named T7ORep. Example 2: Employing T7ORep for evolution of fluorescence intensity
[0134] As proof of concept, the orthogonal replication system was used to evolve the fluorescent protein dEGFP, a variant of GFP that is dimmer than sfGFP. The dEGFP gene was cloned into the T7ori plasmid and propagated for 3 days in an E. coli strain expressing the T7 replisome with an error-prone DNAP, as well as TelN. By Fluorescence-Activated Cell Sorting (FACS), the top 1% most fluorescent cells were isolated and plated. Of the emerging colonies, the 5 most fluorescent ones were characterised by flow cytometry (Fig. 4A). Three of the five clones had improved fluorescence of more than 2-fold, with the most fluorescent clone achieving an increase of 2.6-fold over the unevolved dEGFP (Fig. 4B).
[0135] Other applications
[0136] The orthogonal replication system may be used for continuous directed evolution of proteins whose function can be coupled to the survival of the cell. This may be used, e.g., to evolve enzymes (or a pathway of multiple enzymes, or protein complexes with multiple subunits) that are able to produce a compound of interest. By tying the product of the enzymatic reaction or pathway to cell survival (e.g., by using a genetic circuit and a biosensor that detects the enzymatic product), the enzyme or pathway can be set up to continuously evolve towards improved productivity. This enables discovery, engineering and optimisation of novel enzyme variants and metabolic pathways that may have commercial value. One potential application is the evolution of non-ribosomal peptide synthetases (NRPSs), which have the ability to produce short modified peptides with pharmaceutical relevance. NRPSs are large enzymes and do not function if divided into separate units. Orthogonal replication with mutagenesis can accommodate large constructs and thus may be used for evolution of NRPSs.
[0137] The system may also be applied for in vitro DNA replication or production. Host factors may be added to further increase the efficiency of replication. In vitro DNA production is relevant to pharmaceutical companies working on gene therapy and RNA vaccines.
[0138] Further, the present system can serve as a platform for establishing and developing other orthogonal bioprocesses, e.g., extending the genetic code by incorporating unnatural bases during replication. An engineered DNAP that can accommodate a mix of canonical and non- canonical bases may be used for such an application, without affecting replication of the host cell genome or host cell viability.
[0139] Moreover, the biorthogonal replication system can be employed to provide biosafety and biosecurity by ensuring that an exogenous DNA is only propagated in a target cell and not in other strains upon release.
[0140] It will be appreciated that many further modifications and permutations of various aspects of the described embodiments are possible. Accordingly, the described aspects are intended to embrace all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims.
Claims
CLAIMS1. A method for replicating DNA, comprising contacting a DNA molecule with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a hclicasc-primasc (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication (T7 ori) and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, and wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori.
2. The method of claim 1, wherein the set of proteins comprises one or more proteins selected from a T7 DNAP, a T7 RNAP, a T7 SSBP and a T7 HP.
3. The method of claim 2, wherein the set of proteins comprises a T7 DNAP, a T7 RNAP, a T7 SSBP, and a T7 HP.
4. The method of any one of claims 1 to 3, wherein the one or more linearising proteins is capable of maintaining the replicated DNA in open or closed linear form.
5. The method of claim 4, wherein the one or more linearising proteins is a protelomerase or a terminal protein.
6. The method of claim 5, wherein the protelomerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence in SEQ ID NO: 7-14.
7. The method of claim 6, wherein the protelomerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 7.
8. The method of claim 5, wherein the terminal protein comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence in SEQ ID NO: 15- 19.
9. The method of any one of claims 1 to 8, wherein the one or more nucleic acid sequencescapable of binding to the linearising protein comprises a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 20- 30.
10. The method of claim 9, wherein the one or more nucleic acid sequences capable of binding to the linearising protein comprises a nucleic acid sequence having at least 80%> sequence identity to a nucleic acid sequence set forth in SEQ ID NO 20.
11. The method of any one of claims 1 to 10, wherein the DNA molecule is a linear or circular DNA molecule.
12. The method of any one of claims 1 to 11, wherein the T7 ori comprises a nucleic acid sequence having at least 70% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 1-6.
13. The method of any one of claims 1 to 12, wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate.
14. The method of claim 13, wherein the DNAP comprises a nucleoside deaminase domain.
15. The method of claim 14, wherein the nucleoside deaminase domain comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 35-38.
16. The method of any one of claims 13 to 15, wherein the DNAP is an error-prone polymerase.
17. The method of claim 16, wherein the error-prone DNA polymerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 39.
18. The method of any one of claims 1 to 17, wherein the contacting occurs in vivo in a host cell.
19. The method of claim 18, wherein the host cell is a prokaryotic cell.
20. The method of claim 19, wherein the prokaryotic cell is a bacterial cell.
21. The method of claim 20, wherein the bacterial cell is an E. coli cell.
22. The method of claim 18, wherein the host cell is a eukaryotic cell.
23. The method of claim 22, wherein the eukaryotic cell is a yeast or mammalian cell.
24. The method of any one of claims 18 to 23, wherein the host cell expresses a thioredoxin.
25. The method of any one of claims 1 to 17, wherein the contacting occurs in vitro in the absence of cells.
26. The method of any one of claims 18 to 24, wherein the set of proteins replicates the DNA molecule orthogonal to an endogenous DNA replication system in the host cell.
27. The method of any one of claims 1 to 26, wherein the set of proteins is encoded by the DNA molecule to be replicated.
28. A method for DNA mutagenesis, comprising contacting a DNA molecule with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand binding protein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori, and wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate.
29. A method for evolving a target nucleic acid sequence, the method comprising contacting a DNA molecule comprising the target sequence with a set of proteins comprising a DNA polymerase (DNAP), an RNA polymerase (RNAP), a single-strand bindingprotein (SSBP), a helicase-primase (HP), and one or more linearising proteins, wherein the DNA molecule comprises a T7 origin of replication and one or more nucleic acid sequences that is capable of binding to the one or more linearising proteins, wherein the set of proteins is capable of replicating the DNA molecule at the T7 ori, and wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate.
30. A kit for DNA replication, comprising: a) a set of proteins comprising a DNAP, an RNAP, a SSBP, a HP, and one or more linearising proteins, wherein the set of proteins is capable of replicating a DNA molecule at a T7 ori in the DNA molecule; b) a nucleic acid molecule encoding the set of proteins of (a); or c) a host cell comprising the nucleic acid molecule of (b).
31. A kit for DNA mutagenesis, comprising: a) a set of proteins comprising a DNAP, an RNAP, a SSBP, a HP, and one or more linearising proteins, wherein the set of proteins is capable of replicating a DNA molecule at a T7 ori in the DNA molecule, and wherein the DNAP is a mutant and / or fusion protein capable of increasing the DNA replication error rate above a natural replication error rate; or b) a nucleic acid molecule encoding the set of proteins of (a); or c) a host cell comprising the nucleic acid molecule of (b).
32. A DNA molecule comprising a T7 origin of replication (T7 ori) and one or more nucleic acid sequences capable of binding to a linearising protein, wherein the one or more nucleic acid sequences capable of binding to the linearising protein comprises a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence set forth in SEQ ID NO: 20-30.