E. coli orthogonal DNA replication systems

Engineered T7 DNA polymerases and replisomes in E. coli provide high mutation rates and targeted mutagenesis, overcoming limitations of existing systems by achieving efficient and accelerated directed evolution.

WO2026015343A1PCT designated stage Publication Date: 2026-01-15THE SCRIPPS RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/036214
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2025-07-02
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing orthogonal replication systems in E. coli are limited by low mutation rates and high off-target mutations, making directed evolution inefficient and laborious, while systems compatible with yeast are not suitable for E. coli's faster generation times and higher cell densities.

Method used

Development of engineered T7 DNA polymerases with specific mutations and T7 replisomes for E. coli, enabling high mutation rates and targeted mutagenesis without off-target effects, using vectors derived from pBR322, pSC101, or CloDF13, and incorporating T7 single-stranded DNA binding protein, RNA polymerase, and primase-helicase.

Benefits of technology

The engineered T7 replisomes achieve a mutation rate 500-fold higher than previous systems, allowing accelerated evolution and efficient transformation of libraries, suitable for evolving enzymes, antibodies, and metabolic pathways in E. coli.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000003_0001
    Figure IMGF000003_0001
  • Figure IMGF000014_0001
    Figure IMGF000014_0001
  • Figure IMGF000014_0002
    Figure IMGF000014_0002
Patent Text Reader

Abstract

The present invention provides engineered T7 replisomes that contain one or more modified protein components, and E. coli based orthogonal DNA replication systems that contain such engineered T7 replisomes. Relative to known orthogonal replication systems, the engineered DNA replication systems of the invention are capable of evolving target polynucleotide sequence with enhanced mutation rates.
Need to check novelty before this filing date? Find Prior Art

Description

Atty Docket: TSRI 2250.1PC E. coli Orthogonal DNA Replication Systems CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The subject patent application claims the benefit of priority to U.S. Provisional Patent Application Number 63 / 670,151 (filed July 12, 2024; now pending). The full disclosure of the priority application is incorporated herein by reference in its entirety and for all purposes. STATEMENT OF GOVERNMENT SUPPORT

[0002] This invention was made with government support under GM062159 awarded by the National Institutes of Health. The government has certain rights in the invention. SEQUENCE LISTING

[0003] This application includes by incorporation of reference a sequence listing in the XML format, 2250_1PC_Sequence Listing, which was created on June 13, 2025 and contains 16 KB in content. BACKGROUND OF THE INVENTION

[0004] Low organismal mutation rates (~10-10substitutions per base pair, spb) render mere passaging of cells under selective pressure an inefficient strategy for the diversification of genes on laboratory timescales.(1, 2)Historically, directed evolution addressed this limitation by mutagenesis in vitro (e.g., error-prone PCR) followed by transformation of a microbial host with the diversified libraries for selection of phenotypic fitness.(3–5)However, such in vitro mutagenesis renders directed evolution time consuming and laborious and limits both the depth (rounds of mutagenesis) and scale (number of parallel experiments) of these experiments. While alternative randomization strategies (e.g., DNA shuffling) have improved the efficiency of this process, the removal of ex vivo diversification altogether by increasing in vivo mutagenicity enables continuous evolution at an accelerated pace.(6)

[0005] Early efforts to increase in vivo mutation rates relied on chemical mutagens, engineered host DNA polymerases, and impairing of host DNA repair machinery.(7–9)However, the mutation rates that can be achieved according to these strategies are fundamentally limited by the large number of essential host genes that must be maintained faithfully to ensure survival. Efforts to target mutagenesis to a defined locus or gene using translational fusions of deaminases to nCas9 or T7 RNA polymerase, as well as of nCas9 to error-prone DNA polymerases have been investigated. These approaches are still subject to elevated levels of genomic off-target mutations outside of the specified gene of interest, which can lead to competitive means of selection escape and reduced host fitness.(10–12)These limitations can be mitigated by using an orthogonal replication system where a highly error-prone DNA polymerase maintains a dedicated episome.Here, mutation rates are not subject to the limitations imposed by the error-threshold for genome maintenance, and mutagenesis is constricted to a defined replicon thus minimizing selection escape via off-target mutations.

[0006] The first orthogonal replication system, OrthoRep, was developed by minimizing extranuclear plasmids of yeast that are replicated by an orthogonal cytosolic protein-primed DNA polymerase.(13)Subsequent high-throughput engineering of this DNA polymerase yielded clones with 100,000-fold increased mutation rates of 10-5spb in vivo.(12)The prowess of this highly mutagenic system has been demonstrated by evolving enzymes, antibodies, and entire metabolic pathways.(16–18)However, in the context of directed evolution, the majority of genetic tools are not directly compatible with yeast, but have been developed for Escherichia coli (E. coli), the workhorse organism of synthetic biology.(19)More importantly, an orthogonal replication system in E. coli will benefit from faster generation times (20–30 min) and higher cell densities (109–1010ml-1) compared to yeast (~1.5–2.5 h; 107–108cells ml-1), thus enabling significantly accelerated evolution experiments at the same mutation rate (~500-fold more mutations for a given time and culture volume).(20)Recently, a bacterial replication system EcORep based on the protein-primed DNA polymerase of the lytic bacteriophage PDR1 has been developed for E. coli.(15)However, the mutation rates that can stably be maintained in the disclosed system (~2.0 × 10-7spb) were only modestly enhanced over genomic levels, effectively mutating only 0.02% of a hypothetical 1 kb gene each generation.(20)

[0007] There is still a strong need for highly mutagenic orthogonal replication systems with high transformation efficiencies in E. coli for continuous hypermutation and accelerated evolution. The present invention addresses this and other unmet needs in the art.SUMMARY OF THE INVENTION

[0008] In one aspect, the invention provides engineered T7 DNA polymerases with enhanced mutation rate relative to the wildtype T7 polymerase. These engineered T7 DNA polymerases contain a deletion of residues K118-R145in the exonuclease domain, and a substitution at residue N520, with the amino acid numbering being based on wildtype T7 DNA polymerase (Gp5) sequence under UniProt ID P00581 (Accession No. CAA24412.1; SEQ ID NO:1). In some embodiments, residue N520 can be replaced with a small hydrophobic residue or a small polar residue. In some of these embodiments, the replacing amino acid residues can be small hydrophobic residue M, I, L, or C, or small polar residue T, N, D, or E. In some embodiments, the engineered T7 DNA polymerase can additionally contain one or more substitutions at residues P560, V443, Y24, and Q585. In various embodiments, residue P560 is substituted with residue V, M, I, L, G, A, or K; residue V443 is substituted with residue K, S, T, N, Q, R, H, or C; residue Y24 is substituted with residue F or W; and residue Q585 is substituted with residue R, T, R. E, or C. In some embodiments, the engineered T7 DNA polymerase contains one or more mutations that are selected from the group consisting of P560V, V443K, Y24F, and Q585R. In some embodiments, the engineered T7 DNA polymerase contains N520M, P560V, V443K, Y24F, and Q585R substitutions. Some engineered T7 DNA polymerases contain an amino acid sequence as set forth in any one of SEQ ID NOs:3-7. In a related aspect, the invention provides polynucleotide sequences that encode any of the engineered T7 DNA polymerases described herein.

[0009] In another aspect, the invention provides engineered T7 replisomes. The engineered T7 replisomes typically contain one or more vectors that express (1) a mutant T7 DNA polymerase (Gp5) containing a substitution at residue N520 (e.g., N520M substitution) and a deletion of residues K118-R145in the exonuclease domain, (2) a T7 single-stranded DNA binding protein (Gp2.5) or variant thereof, (3) a T7 RNA polymerase (Gp1) or variant thereof, and (4) a T7 primase-helicase (Gp4) or variant thereof, with the amino acid numbering of T7 DNA polymerase being based on wildtype T7 DNA polymerase (Gp5) sequence under UniProt ID P00581 (Accession No. CAA24412.1; SEQ ID NO:1). In some engineered T7 replisomes of the invention, the one or more vectors each contain an origin of replication that replicates in E. coli. In some embodiments, the one or more vectors are derived from pBR322, pSC101, orCloDF13. In some engineered T7 replisomes of the invention, the mutant T7 DNA polymerase further contains one or more substitutions at residues P560, V443, Y24, and Q585 (e.g., P560V, V443K, Y24F, and Q585R). In some of these embodiments, the mutant T7 DNA polymerase additionally contains P560V, V443K, Y24F, and Q585R substitutions. In some engineered T7 replisomes, the mutant T7 DNA polymerase contains the amino acid sequence as set forth in any one of SEQ ID NOs:3-7. In some embodiments, the T7 primase-helicase is expressed from a Gp4 polynucleotide sequence that does not express Gp4B. In some engineered T7 replisomes, the T7 RNA polymerase (Gp1) is fused to a T7 lysozyme (Gp3.5) mutant that is deficient in hydrolase activity. In some of these embodiments, the T7 lysozyme mutant contains a C131S substitution, with the amino acid numbering of T7 lysozyme being based on wildtype T7 lysozyme sequence under UniProt ID P00806 (SEQ ID NO:8).

[0010] In still another aspect, the invention provides engineered E. coli orthogonal DNA replication systems, which are also termed mutagenic T7-ORACLE systems herein. The E. coli orthogonal DNA replication systems of the invention contain an E. coli cell into which is introduced an engineered T7 replisome. The engineered T7 replisome contains one or more vectors that express (1) a mutant T7 DNA polymerase (Gp5) containing a substitution at residue N520 (e.g., N520M) and a deletion of residues K118-R145in the exonuclease domain, (2) a T7 single-stranded DNA binding protein (Gp2.5) or variant thereof, (3) a T7 RNA polymerase (Gp1) or variant thereof, and (4) a T7 primase-helicase (Gp4) or variant thereof; with the amino acid numbering of T7 DNA polymerase being based on wildtype T7 DNA polymerase (Gp5) sequence under UniProt ID P00581 (Accession No. CAA24412.1; SEQ ID NO:1). In some E. coli orthogonal DNA replication systems, the one or more vectors are derived from pBR322, pSC101, or CloDF13. In some E. coli orthogonal DNA replication systems, the mutant T7 DNA polymerase further contains one or more substitutions at residues P560, V443, Y24, and Q585 (e.g., P560V, V443K, Y24F, and Q585R). In some of these embodiments, the mutant T7 DNA polymerase additionally contains P560V, V443K, Y24F, and Q585R substitutions. In some E. coli orthogonal DNA replication systems of the invention, the employed mutant T7 DNA polymerase contains the amino acid sequence as set forth in any one of SEQ ID NOs:3-7. In some embodiments, the T7 primase-helicase is expressed from a Gp4 polynucleotide sequence that does not express Gp4B. In some embodiments, the T7 RNA polymerase (Gp1) is fused to a T7 lysozyme (Gp3.5) mutant that is deficient in hydrolase activity. In some of theseembodiments, the T7 lysozyme mutant contains a C131S substitution, with the amino acid numbering of T7 lysozyme being based on wildtype T7 lysozyme sequence under UniProt ID P00806 (Accession No. Q38567; SEQ ID NO:8).

[0011] In another aspect, the invention provides methods for evolving a target polypeptide. The methods entail (1) transforming an E. coli orthogonal replication system described herein with a vector containing a T7 origin of replication and harboring a polynucleotide that encodes the target polypeptide, (2) culturing the transformed E. coli replication system under appropriate conditions, and (3) selecting one or more E. coli cells from the cultured E. coli replication system with an enhanced property or a desired phenotype. In some embodiments, the T7 origin of replication in the vector is φOR. In various embodiments, the target polypeptide to be evolved is an enzyme, a hormone, a therapeutic protein, or an antibody molecule. Some methods of the invention can involve one or more additional manipulations including, e.g., isolating the vector from the cultured E. coli cell, expressing the evolved polynucleotide, and purifying the evolved target polypeptide expressed from the evolved polynucleotide.

[0012] A further understanding of the nature and advantages of the present invention may be realized by reference to the remaining portions of the specification and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1. Establishing an orthogonal replication system in E. coli based on the bacteriophage T7 replisome. (A) T7 RNA polymerase initiates replication by priming the replisome at the T7 origin of replication. T7 lysozyme stimulates initiation of replication. (B) T7 DNA polymerase is primed by the orthogonal transcript and T7 helicase / primase is recruited to the origin, leading to (C) concerted leading- and lagging-strand synthesis. T7 ssDNA-binding protein stabilizes the replication loop and E. coli thioredoxin serves as a sliding clamp to increase T7 DNA polymerase processivity. (D) Transformation of a plasmid with a selection marker and a T7 origin of replication is only replicated by E. coli strains that (E) provide the T7 replisome proteins in trans. (F) T7 origin plasmids are stably maintained, and transformed strains show robust growth with a doubling time of 28 min and a final OD600 of 7.2. (G) The Δ28 mutant of T7 DNA polymerase displays an elevated mutation rate of 3.31 × 10-8spb on the origin replicon (pOR-1), without increasing genomic mutation rate (3.6 ×10-10spb).

[0014] Figure 2. Directed evolution of a mutagenic T7 DNA polymerase. (A) Structure of T7 DNA polymerase (white piped helices) bound to DNA and its thioredoxin cofactor. Domains of the polymerase (3’-5’-proofreading exonuclease domain, fingers domain, palm domain, and thioredoxin-binding site) are labeled. (B) Mutations in T7 DNA polymerase that decrease base pairing fidelity (Y24F, V443K, N520M, P560V, Q585R) are highlighted (ball and stick representation). (C) Mutation rates of T7 DNA polymerase mutants as determined by fluctuation analysis. Reversion rates are converted to mutation rates by fitting the data to the Luria-Delbruck function using the Ma-Sandri-Sarkar Maximum Likelihood Estimator. (D) Comparison of the mutagenesis by T7 DNA polymerase mutant G on the orthogonal replicon (pOR-3) to the genomic mutation rate shows five orders of magnitude increase in targeted mutagenesis and no off-target activity.

[0015] Figure 3. Continuous evolution of TEM-1 β-Lactamase. (A) TEM-1 β- Lactamase is continuously mutagenized by the mutagenic T7 replisome (T7rep-7, quintuple mutant). (B) Selection for expanded substrate scopes is achieved by passaging cells (96 replicates) in the presence of increasing amounts of cephalosporin and monobactam antibiotics. (C) Chemical structures of carbenicillin, cephalosporin β- lactams (cefotaxime, cefepime, ceftazidime), and the monobactam aztreonam. (D) Sanger sequencing for 96 cefotaxime resistance evolution experiments after 1000-fold increase in antibiotic concentration. (E) Structure of TEM-1 β-lactamase with mutations that multiple replicates converge upon and the catalytic S70 residue shown.

[0016] Figure 4. Time course of β-lactamase evolution experiments. The time course for evolution of (A) cefotaxime and (B) ceftazidime resistance was interrogated by deep-sequencing of all 96 replicates after each passage (once per day)..

[0017] Figure 5. Continuous evolution of TEM-1 β-lactamase saturation libraries. Evolutionary trajectories on a hypothetical fitness landscape for (A) evolution from a single orthologstarting point, evolution from individual members of a site-saturation library multiple starting points in in (B) isolation, and (C) in competition. (D) Comparison of the active-site mutations in the WT enzyme, the TEM-1 evolution campaign starting from the WT enzyme, as well as the evolution campaigns from isolated members of the site-saturation library, and from competing members of the site-saturation library. (E) Sanger sequencing for 96 cefotaxime resistance evolutionexperiments after starting from 96 individual colonies of a 5-site saturation library of TEM-1 (E104X, Y105X, G238X, E240X, and R244X) after 1000-fold increase in antibiotic concentration. (F) Sanger sequencing for 96 cefotaxime resistance evolution experiments after starting from 96 cultures comprising the entire TEM-1 library (E104X, Y105X, G238X, E240X, and R244X) after 1000-fold increase in antibiotic concentration.

[0018] Figure 6. Time course of β-lactamase evolution experiments. The time course for evolution of cefepime resistance by deep-sequencing of 96-replicates after each passage (once per day).

[0019] Figure 7. Time course of β-lactamase evolution experiments. The time course for evolution of aztreonam resistance was interrogated by deep-sequencing of 96-replicates after each passage (once per day). DETAILED DESCRIPTION OF THE INVENTION I. Overview

[0020] The invention provides engineered T7 replisomes that contain modified protein components, and E. coli based orthogonal DNA replication systems that contain such engineered T7 replisomes. Relative to known orthogonal replication systems, the engineered DNA replication systems of the invention are capable of evolving target polynucleotide sequences with enhanced mutation rates. The invention is derived in part from studies undertaken by the inventors to develop an orthogonal T7 replisome for continuous hypermutation and accelerated evolution in E. coli. As detailed herein, the inventors established an orthogonal replication system in E. coli based on the replisome of bacteriophage T7. The exemplified orthogonal replication system demonstrated significant technological advantages over other E. coli based DNA replication systems known in the art. Containing an engineered T7 DNA polymerase obtained by directed evolution for increased mutation rates, optionally other modified components of T7 replisome, the exemplified orthogonal replication system can be continuously passaged at a mutation rate of 1.7 × 10-5spb. This rate is two orders of magnitude higher than that of the best reported orthogonal replication system in E. coli to date (i.e., EcORep). The combination of the high mutagenesis rate, the fast growth rate of E. coli, as well as the ease with which both the E. coli and the circular replicon plasmid can be integrated into standard molecular biology workflows differentiates theexemplified orthogonal replication system from the prior art systems such as OrthoRep and EcORep.

[0021] In comparison to PACE, another known continuous evolution system in E. coli, the orthogonal replisome systems described herein benefit from simplicity and the lack of specialized equipment that is required to run evolution campaigns, as well as from the ability to autonomously run continuous evolution campaigns in a large number of replicates.(48, 49) The helicase-dependent concerted leading- and lagging-strand synthesis by the engineered T7 replisomes of the invention, unlike helicase-independent replication modules of known orthogonal replication systems, enables high transformation efficiencies of circular plasmids (2.4 × 1010cfu μg-1) and thus the integration of pre-diversified libraries (e.g., site-saturation libraries) and continuous mutagenesis into a single workflow. This will be particularly beneficial for evolving fundamentally new as opposed to improved activities, as well as for investigating vast fitness landscapes through divergent, as opposed to convergent evolution campaigns.

[0022] The invention accordingly provides novel modified protein components of T7 replisome (e.g., T7 DNA polymerase), engineered T7 replisomes containing the novel components, E. coli based orthogonal replication systems containing the engineered replisomes (aka “mutagenic T7-ORACLE”), as well as methods of using the orthogonal replisome systems for evolving target proteins or polypeptides of interest. It is noted that this invention is not limited to the particular methodology, protocols, and reagents described as these may vary. The novel compositions of the invention and related methods can all be generated or performed in accordance with the procedures exemplified herein or routinely practiced methods well known in the art. Unless otherwise indicated, the practice of the present invention can employ conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry and immunology, which are within the skill of the art. Such techniques are explained fully in the literature. For example, exemplary methods are described in the following references, e.g., Sambrook, et al., Molecular Cloning: A Laboratory Manual (3rdEdition, 2000); Brent et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (ringbou ed., 2003); DNA Cloning: A Practical Approach, vol. I & II (D. Glover, ed.); Oligonucleotide Synthesis: Methods and Applications (P. Herdewijn, ed., 2004); Farrell, R., RNA Methodologies: A Laboratory Guide for Isolation and Characterization (3rdEdition 2005); Nucleic Acid Hybridization: Modern Applications (Buzdin and Lukyanov, eds., 2009); Transcriptionand Translation (B. Hames & S. Higgins, eds., 1984); Current Protocols in Protein Science (CPPS) (John E. Coligan, et. al., ed., John Wiley and Sons, Inc.); Current Protocols in Cell Biology (CPCB) (Juan S. Bonifacino et. al. ed., John Wiley and Sons, Inc.); Animal Cell Culture (R. Freshney, ed., 1986); and Freshney, R.I. (2005) Culture of Animal Cells, a Manual of Basic Technique, 5thEd. Hoboken NJ, John Wiley & Sons.

[0023] The following sections provide additional guidance for practicing the compositions and methods of the present invention. II. Definitions

[0024] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. The following references provide one of skill with a general definition of many of the terms used in this invention: Academic Press Dictionary of Science and Technology, Morris (Ed.), Academic Press (1sted., 1992); Oxford Dictionary of Biochemistry and Molecular Biology, Smith et al. (Eds.), Oxford University Press (revised ed., 2000); Encyclopaedic Dictionary of Chemistry, Kumar (Ed.), Anmol Publications Pvt. Ltd. (2002); Dictionary of Microbiology and Molecular Biology, Singleton et al. (Eds.), John Wiley & Sons (3rded., 2002); Dictionary of Chemistry, Hunt (Ed.), Routledge (1sted., 1999); Dictionary of Pharmaceutical Medicine, Nahler (Ed.), Springer-Verlag Telos (1994); Dictionary of Organic Chemistry, Kumar and Anandand (Eds.), Anmol Publications Pvt. Ltd. (2002); and A Dictionary of Biology (Oxford Paperback Reference), Martin and Hine (Eds.), Oxford University Press (4thed., 2000). In addition, the following definitions are provided to assist the reader in the practice of the invention.

[0025] The singular terms "a," "an," and "the" include plural referents unless the context clearly indicates otherwise. Similarly, the word "or" is intended to include "and" unless the context clearly indicates otherwise.

[0026] As used herein, the term "amino acid" of a peptide refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid.

[0027] As used herein the term "comprising" or "comprises" is used in reference to compositions, methods, and respective component(s) thereof, that are essential to the invention, yet open to the inclusion of unspecified elements, whether essential or not.

[0028] As used herein the term "consisting essentially of" refers to those elements required for a given embodiment. The term permits the presence of elements that do not materially affect the basic and novel or functional characteristic(s) of that embodiment of the invention.

[0029] The term "consisting of" refers to compositions, methods, and respective components thereof as described herein, which are exclusive of any element not recited in that description of the embodiment.

[0030] The term "conservatively modified variant" applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, conservatively modified variants refers to those nucleic acids which encode identical or essentially identical amino acid sequences, or where the nucleic acid does not encode an amino acid sequence, to essentially identical sequences. Because of the degeneracy of the genetic code, a large number of functionally identical nucleic acids encode any given protein. For instance, the codons GCA, GCC, GCG and GCU all encode the amino acid alanine. Thus, at every position where an alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are “silent variations,” which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid. One of skill will recognize that each codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine, and TGG, which is ordinarily the only codon for tryptophan) can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid that encodes a polypeptide is implicit in each described sequence.

[0031] For polypeptide sequences, “conservatively modified variants” refer to a variant which has conservative amino acid substitutions, amino acid residues replaced with other amino acid residue having a side chain with a similar charge. Families of amino acid residues having side chains with similar charges have been defined in the art. These families include amino acids with basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine) and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine).

[0032] The term “engineered cell” or “recombinant host cell” (or simply “host cell”) refers to a cell into which a recombinant expression vector has been introduced. It should be understood that such terms are intended to refer not only to the particular subject cell but to the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term “host cell” as used herein.

[0033] The term "isolated" means a reference molecule (e.g., a polypeptide or a polynucleotide) is removed from its natural surroundings. However, some of the components found with it may continue to be with an "isolated" molecule. Thus, an “isolated polypeptide” is not as it appears in nature but may be substantially less than 100% pure protein.

[0034] The terms “identical” or percent “identity,” in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same. Two sequences are "substantially identical" if two sequences have a specified percentage of amino acid residues or nucleotides that are the same (i.e., 60% identity, optionally 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity over a specified region, or, when not specified, over the entire sequence), when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. Optionally, the identity exists over a region that is at least about 50 nucleotides (or 10 amino acids) in length, or more preferably over a region that is 100 to 500 or 1000 or more nucleotides (or 20, 50, 200 or more amino acids) in length.

[0035] Methods of alignment of sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman, Adv. Appl. Math.2:482c, 1970; by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol.48:443, 1970; by the search for similarity method of Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988; by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, Madison, WI); or by manual alignment and visual inspection (see, e.g., Brent et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (ringbou ed., 2003)). Two examples of algorithms that are suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., Nuc. Acids Res.25:3389-3402, 1977; and Altschul et al., J. Mol. Biol.215:403-410, 1990, respectively.

[0036] Other than percentage of sequence identity noted above, another indication that two nucleic acid sequences or polypeptides are substantially identical is that the polypeptide encoded by the first nucleic acid is immunologically cross reactive with the antibodies raised against the polypeptide encoded by the second nucleic acid, as described below. Thus, a polypeptide is typically substantially identical to a second polypeptide, for example, where the two peptides differ only by conservative substitutions. Another indication that two nucleic acid sequences are substantially identical is that the two molecules or their complements hybridize to each other under stringent conditions, as described below. Yet another indication that two nucleic acid sequences are substantially identical is that the same primers can be used to amplify the sequence.

[0037] The term “operably linked” refers to a functional relationship between two or more polynucleotide (e.g., DNA) segments. Typically, it refers to the functional relationship of a transcriptional regulatory sequence to a transcribed sequence. For example, a promoter or enhancer sequence is operably linked to a coding sequence if it stimulates or modulates the transcription of the coding sequence in an appropriate host cell or other expression system. Generally, promoter transcriptional regulatory sequences that are operably linked to a transcribed sequence are physically contiguous to the transcribed sequence, i.e., they are cis-acting. However, some transcriptional regulatory sequences, such as enhancers, need not be physically contiguous or located in close proximity to the coding sequences whose transcription they enhance.

[0038] An origin of replication is a sequence motif of DNA at which replication is initiated on a chromosome, plasmid or virus. The T7 phage genome replication starts from the primary origins, φ1.1A and φ1.1B, but replication can also be initiated at secondary origins (e.g., φOR, φ13 and φ6.5) if the primary origins are deleted. Early in vitro studies demonstrated the initiation of DNA synthesis in plasmids containing a T7 primary origin using purified T7 phage enzymes. An essential element of every T7 origin of replication (T7 ori) is a promoter dependent on T7 RNA polymerase.

[0040] “Regulatory sequences” are nucleotide sequences located upstream (5' non- coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, translation leader sequences, introns, and polyadenylation signal sequences. They include natural and synthetic sequences as well as sequences that may be a combination of synthetic and natural sequences. As is noted above, the term “suitable regulatory sequences” is not limited to promoters. However, some suitable regulatory sequences useful in the present invention will include, but are not limited to constitutive promoters, development-specific promoters, regulatable promoters and viral promoters. “5' non-coding sequence” refers to a nucleotide sequence located 5' (upstream) to the coding sequence. It is present in the fully processed mRNA upstream of the initiation codon and may affect processing of the primary transcript to mRNA, mRNA stability or translation efficiency.

[0041] “Promoter” refers to a nucleotide sequence, usually upstream (5') to its coding sequence, which directs and / or controls the expression of the coding sequence by providing the recognition for RNA polymerase and other factors required for proper transcription. “Promoter” includes a minimal promoter that is a short DNA sequence comprised of a TATA- box and other sequences that serve to specify the site of transcription initiation, to which regulatory elements are added for control of expression. “Promoter” also refers to a nucleotide sequence that includes a minimalpromoter plus regulatory elements that is capable of controlling the expression of a coding sequence or functional RNA. This type of promoter sequence consists of proximal and more distal upstream elements, the latter elements often referred to as enhancers. Accordingly, an “enhancer” is a DNA sequence that can stimulate promoter activity and may be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of a promoter. It is capable of operating in both orientations (normal or flipped), and is capable of functioning even when moved either upstream or downstream from the promoter. Both enhancers and other upstream promoter elements bind sequence-specific DNA-binding proteins that mediate their effects. Promoters may be derived in their entirety from a native gene, or be composed of different elements derived from different promoters found in nature, or even be comprised of synthetic DNA segments. A promoter may also contain DNA sequences that are involved in the binding of protein factors that control the effectiveness of transcription initiation in response to physiological or developmental conditions.

[0042] The terms "open reading frame" and "ORF" refer to the amino acid sequence encoded between translation initiation and termination codons of a coding sequence. The terms "initiation codon" and "termination codon" refer to a unit of three adjacent nucleotides ('codon') in a coding sequence that specifies initiation and chain termination, respectively, of protein synthesis (mRNA translation).

[0043] The term “transformation” refers to the transfer of a nucleic acid fragment into the genome of a host cell, resulting in genetically stable inheritance. A “host cell” is a cell that has been transformed, or is capable of transformation, by an exogenous nucleic acid molecule. Host cells containing the transformed nucleic acid fragments are referred to as “transgenic” cells. Transformed,” “transduced,” “transgenic” and “recombinant” refer to a host cell into which a heterologous nucleic acid molecule has been introduced. Unless otherwise noted, the term “transformation” is used herein to refer to delivery of DNA into prokaryotic (e.g., E. coli) cells, and the term “transduction” is used to refer to infecting cells with viral particles. The term "untransformed" refers to normal cells that have not been through the transformation process.

[0044] The term “vector” is intended to refer to a polynucleotide molecule capable of transporting another polynucleotide to which it has been linked. One type of vector is a “plasmid”, which refers to a circular double stranded DNA loop into which additionalDNA segments may be ligated. Another type of vector is a viral vector, wherein additional DNA segments may be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as “recombinant expression vectors” (or simply, “expression vectors”). III. Engineered T7 DNA polymerases and T7 replisomes

[0045] DNA replication is carried out by the replisome, a large protein complex that includes DNA helicase, DNA polymerase (DNAP), and other proteins. The bacteriophage T7 replisome has been identified to be a simple and efficient model system for the study of DNA replication, as only four proteins are required for basic phage replication, yet it mimics more complex systems. Briefly, T7 DNAP, a 1:1 complex of gene 5 protein (Gp5) and its processivity factor Escherichia coli thioredoxin (trx), is responsible for nucleotide polymerization. The protein product of T7 gene 4 (Gp4) provides helicase and primase activities. The helicase activity, required for unwinding the parental DNA strand, resides in the C-terminal half of the protein, while the primase activity, required to initiate lagging-strand DNA synthesis by synthesizing RNA primers, is located in the N-terminal of the protein. T7 helicase is a hexameric motor and couples the hydrolysis of nucleoside triphosphate to translocate along single- stranded DNA (ssDNA) and unwind double-stranded DNA (dsDNA). Finally, the product of gene 2.5 of bacteriophage T7 (Gp2.5), a single-stranded DNA-binding (SSB) protein, coats ssDNA to remove its secondary structure.

[0046] The invention provides engineered T7 replisomes that enable orthogonal replication of target DNA sequences in E. coli with enhanced mutation rates. Unless otherwise noted, an engineered T7 replisome of the invention encompasses a composition or combination that contains (1) the essential T7 proteins for DNA replication, DNA polymerase (Gp5), DNA primase-helicase (Gp4), and ssDNA-binding protein (Gp2.5), (2) polynucleotides encoding these proteins, or (3) expression vectors containing the polynucleotides. Typically, the T7 DNA polymerase in the T7 replisomes of the invention is an engineered Gp5 protein described herein thatdemonstrated enhanced mutation rates relative to the wildtype Gp5 protein. The other components in the replisomes can be either wildtype proteins or modified variants of the wildtype proteins as described herein.

[0047] Other than the essential components, the engineered T7 replisome can contain additional components that can promote orthogonal DNA replication by the replisome in an E. coli host cell. It is also noted that, while a thioredoxin protein may also be included in the engineered T7 replisomes of the invention, this essential protein for T7 DNA replication is usually provided endogenously by an E. coli host cell. As detailed below, one or more of the protein components of the engineered T7 replisomes are modified molecules of the respective wildtype proteins. In various embodiments, the different protein components of the engineered T7 replisome can be expressed from one, two, three or more expression vectors. Typically, the expression vectors contain a replication origin and transcription regulatory elements (e.g., a promoter) that function in E. coli cells. For example, the vectors can be based on plasmids pBR322, pSC101 or CloDF13, as exemplified herein.

[0048] To enable orthogonal DNA replication with enhanced mutation rates in an E. coli host, the engineered T7 replisomes typically contain an engineered T7 DNA polymerase, which is essential to the T7 replisome mediated DNA replication. A T7 DNA polymerase variant (Δ28-T7 DNA Pol; SEQ ID NO:2) which contains deletion of residues K118-R145(SEQ ID NO:11) was known to lack the 3’-5’ exonuclease activity, and is capable of replicating DNA sequences with an improved mutation rate over the wildtype enzyme. See, e.g., Tabor et al., J. Biol. Chem.1989, 264:6447-58; and Joneja et al., Anal. Biochem.2011, 414:58-69. Typically, the engineered T7 DNA polymerase of the invention contains the Δ28 deletion and also a substitution of residue N520 (e.g., with a small hydrophobic amino acid residue or a small polar amino acid residue). In various embodiments, the replacing residue for N520 can be hydrophobic residue M, I, L, or C, or polar residue T, N, D, or E. As a specific exemplification, residue N520 in the engineered T7 DNA polymerase of the invention is replaced with a Met residue (SEQ ID NO:3). As demonstrated herein, T7 DNA polymerase containing the Δ28 deletion and the substitution at residue N520 results in a further substantial improvement of mutation rate over the Δ28 T7 DNA polymerase (see, e.g., Fig.2, C).

[0049] In some embodiments, the engineered T7 DNA polymerase can additionally contain one or more specific amino acid substitutions at residues P560, V443, Y24 and Q585. In various embodiments, residue P560 can be substituted with small hydrophobicresidue V, M, I, L, G, or A, as well as K; residue V443 can be substituted with positively charged residue K, R or H, uncharged residue S, T, N, or Q, as well as residue C; residue Y24 can be substituted with aromatic residue F or W; and residue Q585 can be substituted with residue Q, T, R, E or C. Exemplary T7 DNA polymerases containing mutations at these amino acid residues include mutants containing one or more of substitutions P560V, V443K, Y24F, and Q585R. It is noted that the amino acid numbering is based on the sequence of the full length wildtype T7 DNA polymerase under UniProt ID P00581 (Accession No. CAA24412.1; SEQ ID NO:1). As demonstrated herein, these mutants were identified via directed evolution of the T7 DNA polymerase for increased mutation rates. In some embodiments, the engineered T7 DNA polymerase contains the Δ28 deletion, the N520M substitution, and the P560V substitution (SEQ ID NO:4). In some embodiments, the engineered T7 DNA polymerase contains the Δ28 deletion, the N520M substitution, the P560V substitution, and the V443K substitution (SEQ ID NO:5). In some embodiments, the engineered T7 DNA polymerase contains the Δ28 deletion, the N520M substitution, the P560V substitution, the V443K substitution, and the Y24F substitution (SEQ ID NO:6). In still some embodiments, the engineered T7 DNA polymerase contains the Δ28 deletion, the N520M substitution, the P560V substitution, the V443K substitution, the Y24F substitution, and the Q585R substitution (SEQ ID NO:7). In various embodiments, engineered T7 DNA polymerases of the invention can contain an amino acid sequence that is substantially identical (e.g., at least 90%, 95%, 98%, or 99% identical) to, or a conservatively modified variant of, one of the engineered T7 DNA polymerases specifically exemplified herein (e.g., SEQ ID NOs:3-7).

[0050] Wildtype T7 DNA Pol sequence (SEQ ID NO:1) MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALEAEVARG GLIVFHNGHKYDVPALTKLAKLQLNREFHLPRENCIDTLVLSRLIHSNLKDTDM GLLRSGKLPGKRFGSHALEAWGYRLGEMKGEYKDDFKRMLEEQGEEYVDGM EWWNFNEEMMDYNVQDVVVTKALLEKLLSDKHYFPPEIDFTDVGYTTFWSES LEAVDIEHRAAWLLAKQERNGFPFDTKAIEELYVELAARRSELLRKLTETFGS WYQPKGGTEMFCHPRTGKPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELD TREYVAGAPYTPVEHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPVVDDEV LEGVRVDDPEKQAAIDLIKEYLMIQKRIGQSAEGDKAWLRYVAEDGKIHGSVN PNGAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWVQAGI DASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPTRDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAALRESIQQTLVESSQWVA GEQQVKWKRRWIKGLDGRKVHVRSPHAALNTLLQSAGALICKLWIIKTEEML VEKGLKHGWDGDFAYMAWVHDEIQVGCRTEEIAQVVIETAQEAMRWVGDH WNFRCLLDTEGKMGPNWAICH

[0051] Δ28-T7 DNA Pol sequence (SEQ ID NO:2): containing deletion of residues 118-145. Additional residues mutated in the engineered T7 polymerase variants (SEQ ID NOs:3-7), Y24, V443, N520, P560, and Q585, are highlighted. MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALEAEVARGG LIVFHNGHKYDVPALTKLAKLQLNREFHLPRENCIDTLVLSRLIHSNLKDTDMG LLRSGKLPG-----------------------------------------------------------MLEEQGEEYVDGME WWNFNEEMMDYNVQDVVVTKALLEKLLSDKHYFPPEIDFTDVGYTTFWSESL EAVDIEHRAAWLLAKQERNGFPFDTKAIEELYVELAARRSELLRKLTETFGSW YQPKGGTEMFCHPRTGKPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTR EYVAGAPYTPVEHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPVVDDEVLE GVRVDDPEKQAAIDLIKEYLMIQKRIGQSAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWVQAGIDA SGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPTRDNAKTFIYGFLY GAGDEKIGQIVGAGKERGKELKKKFLENTPAIAALRESIQQTLVESSQWVAGE QQVKWKRRWIKGLDGRKVHVRSPHAALNTLLQSAGALICKLWIIKTEEMLVE KGLKHGWDGDFAYMAWVHDEIQVGCRTEEIAQVVIETAQEAMRWVGDHWN FRCLLDTEGKMGPNWAICH

[0052] In addition to the novel T7 DNA polymerase variant, the engineered T7 replisomes of the invention can optionally contain one or more other modified components of proteins involved in T7 DNA replication. In some embodiments, the engineered replisome utilizes a modified Gp4 protein. Bacteriophage T7 expresses two forms of gene 4 protein (Gp4). The 63-kDa full-length Gp4 (Gp4A) contains both the helicase and primase domains. T7 phage also express a 56-kDa truncated Gp4 (Gp4B) lacking the zinc binding domain of the primase; the protein has helicase activity but no DNA-dependent primase activity. At the N terminus of the Gp4A protein is comprised of the primase domain, including the zinc binding domain (ZBD) and RNA polymerase domain (RPD), connected together by a flexible linker. Following the primase domain is another linker of 26 residues and a helicase domain at the C terminus. Due to the truncation of the first 63 amino acids in 4B, it contains only the RBD and the helicase domain. During T7 infection of E. coli, equal amounts of 4A and 4B are expressed.Since the Gp4B protein is not for replication of T7 phage, the Gp4 protein to be included in the replisomes of the invention can contain only the Gp4A protein. As exemplified herein, this can be achieved via removing the internal start codon of the open reading frame for Gp4B translation in the Gp4 coding sequence in the vectors expressing components of the replisome, e.g., a M64L substitution in the modified Gp4 protein described in the Examples below. As explained below, this serves to eliminate toxic effects that have been observed of Gp4 primase / helicase due to internal promoter in Gp4 coding sequence for translation of Gp4B.

[0053] In some embodiments, a modified T7 lysozyme (Gp3.5) can be included in the engineered T7 replisomes. T7 lysozyme plays an important role in the switch between viral transcription and genome replication. It interacts with and inhibits the viral RNA polymerase that becomes unable to produce additional late transcripts. This lysozyme-polymerase complex in turn plays an active role in viral genome replication and packaging. To maximumly stimulate DNA replication, the T7 lysozyme is modified to eliminate its hydrolase activity while maintaining its function in inhibition of transcription. In some embodiments, this can be realized by introducing a mutation in the hydrolase domain of T7 lysozyme. As exemplification, a modified T7 lysozyme (SEQ ID NO:9) containing a C131S substitution in the wildtype T7 lysozyme (Gp3.5) sequence is found to be catalytically inactive. The wildtype T7 lysozyme sequence is provided under UniProt ID P00806 (Accession No. Q38567; SEQ ID NO:8). To achieve optimal effect on DNA replication, the modified T7 lysozyme can be translationally fused to the T7 RNA polymerase (Gp1). In some embodiments, the T7 RNA polymerase can be fused to the C-terminus of the modified T7 lysozyme. To facilitate interaction of the two fused proteins, a peptide spacer or linker motif can be employed to connect the modified T7 lysozyme to the N-terminus of the T7 RNA polymerase. An example of such a linker motif, TACSTSSTFDRLLSDCSHSGFVATRSTAT (SEQ ID NO:10), was used in the fusion protein exemplified in the Examples herein.

[0054] Wildtype T7 lysozyme (Gp3.5) sequence (SEQ ID NO:8): residue mutated in the engineered T7 lysozyme (SEQ ID NO:9) is highlighted. MARVQFKQRESTDAIFVHCSATKPSQNVGVREIRQWHKEQGWLDVGYHFIIKR DGTVEAGRDEMAVGSHAKGYNHNSIGVCLVGGIDDKGKFDANFTPAQMQSL RSLLVTLLAKYEGAVLRAHHEVAPKACPSFDLKRWWEKNELVTSDRGIV. E. coli based orthogonal DNA replication systems

[0055] The invention provides E. coli orthogonal DNA replication system (“E. coli based orthogonal DNA replication systems” or “engineered E. coli cells for orthogonal DNA replication” ) that contain the engineered T7 replisomes described herein. The engineered T7 replisomes present in the E. coli host can be in the form of expressed proteins, polynucleotides encoding the component proteins of the replisome, or expression vectors containing the polynucleotide sequences. In some preferred embodiments, the orthogonal DNA replication systems of the invention contain an E. coli cell into which is transformed one or more vectors expressing an engineered replisome described herein. They are able to orthogonally replicate a target polynucleotide of interest, which can be separately introduced into the E. coli cell, e.g., via a vector containing a T7 origin of replication as exemplified herein.

[0056] In general, any E. coli cell can be employed to produce the orthogonal DNA replication systems of the invention, as long as it is compatible with expression of the vectors encoding the protein components of the engineered T7 replisome. In preferred embodiments, the Escherichia coli thioredoxin (trx) protein required for the function of the T7 replisome is exogenously expressed by the cell. In various embodiments, the employed E. coli cell can be from any of the routinely used E. coli strains in academic or industrial settings for expressing exogeneous proteins. These commonly used E. coli strains include, e.g., NEB 10-beta, AG1, AB1157, B2155, BHB2688, BL21, BL21[pT7POL26], BL21(AI), BL21(DE3), BL21(DE3) pLysS, BL21(DE3) pLysE, BLR, BNN93, BNN97, BW25113, BW26434 (CGSC Strain # 7658), BW313, C600, C600 hflA150 (Y1073, BNN102), CSH50, D1210, NEB dam– / dcm–, DB3.1, DC10B, DH1, DH5α, DH5αF, DH5αF(λ), DH5αpir, DH5αpir116 variant, DH10B, DH12S, DM1, DP50, E. cloni(r) 5alpha, E. cloni(r) 10G, E. cloni(r) 10GF', EPI300, ER2738, ER2566, ER2267, FHK12, GM31, GM124, GM2198, H12R8a, HB94, HB101, HMS174(DE3), High-Control(tm) BL21(DE3), High-Control(tm) 10G, IJ1126, IJ1127, IM01B, IM08B, IM30B, IM93B, JK242, JM83, JM101, JM103, JM105, JM106, JM107, JM108, JM109, JM109(DE3), JM110, JM2.300, JTK165, K123000, K12ΔH1Δtrp, LE392, LK111, M15, M5219, Mach1, MC1061, MC1061(λ), MC1061Rif, MC4100, MFDpir, MG1655, MG1655 seqA-eYFP, MG1655 seqA- mEOS3.2, MG1655 seqA-PAmCherry, OmniMAX2, OverExpress(tm)C41(DE3), OverExpress(tm)C41(DE3)pLysS, OverExpress(tm)C43(DE3), OverExpress(tm)C43(DE3)pLysS, Rosetta™(DE3)pLysS, Rosetta-gami(DE3)pLysS,RR1, RV308, S26, S26R1d, S26R1e,SG4044, SG4121, SM10(λpir), SOLR, SS320 (Lucigen), STBL2, STBL4, SURE, SURE2, TG1, TOP10, Top10F', Turbo, W3110, W3110 (λ857S7), WK6mutS(λ), WM3064, XL1-Blue, XL1-Blue MRF', XL2-Blue, XL2-Blue MRF', XL1-Red, XL10-Gold, and XL10-Gold KanR.

[0057] The essential characteristics of these E. coli strains, including their genotypes and appropriate culturing conditions, have been described in the literatures. See, e.g., Bachmann BJ: Derivation and genotypes of some mutant derivatives of Escherichia coli K-12. Escherichia coli and Salmonella typhimurium. Cellular and Molecular Biology (Eds. F C Neidhardt et al.). Washington, D.C., American Society for Microbiology 1987, 2:1190-1219; Hohn B. In vitro packaging of lambda and cosmid DNA. Methods Enzymol.1979;68:299-309; Studier et al., 2009, J. Mol. Biol. 394(4), 653; Studier et al., Methods Enzymol.1990;185:60-89; Kunkel et al., Proc Natl Acad Sci U S A.1985:488-9; Appleyard,et al., Genetics 1954, 39: 440; Hanahan, 1983, J. Mol. Biol.166, 577; Blomfeld et al., J. Bact.173: 5298-5307, 1991; Bernard and Couturier, J Mol Biol.1992, 226(3):735-45; Monk et al., mBio.2012; 3(2):e00277- 11; Meselson and Yuan, Nature 1968; 217:1110-4; Grant et al., 1990; Proc. Natl. Acad. Sci. USA 87: 4645-4649; Sitaras et al., Plasmid.2011; 65(3):232-8; Sitaras et al., Plasmid 2011; 65:232-8; Lin et al., BioTechniques 1992; 12: 718; Leder et al., Science 1977; 196(4286):175-7; Kolmar et al., EMBO J.1995; 14:3895-904; Marinus et al., Mol Gen Genet. 1983; 192(1-2): 288-9; Remaut and Fiers, J. Mol. Biol.1972; 71(2):243-61; Boyer et al., J Mol Biol.1969; 41: 459-72; Monk et al., mBio.2015; 6(3):e00308-15; Parker et al., Mol Gen Genet.1980;180(2):275-81; Messing et al., Nucleic Acids Res.1981; 9: 309; Yanisch-Perron et al., Gene 1985; 33: 103; Mairhofer et al., J Biotechnol.2010; 146:130-7; Brenner et al., Proc. Natl. Acad. Sci. USA 2007; 104: 17300-4; Kittleson et al., J. Biol. Eng.2011; 5: 10; Loomis et al., J. Mol. Biol.1967; 23: 487-94; Bernard et al., Gene 1979; 5: 59-76; Remaut et al., Gene 1981; 15: 81-93; Zabeau et al., EMBO J. 1982; 1:1217-24; Casadaban et al., J. Mol. Biol.1980; 138: 179-207; Mertens et al., Gene 1995; 164: 9-15; Casadaban et al., J. Mol. Biol.1980; 138: 179-207; Blattner et al., Science 1997; 277: 1453-62; Mika et al., Faraday Discuss.2015; 184: 425-50; Garen et al., J. Mol. Biol.1965; 4: 167-78; Gottesman et al., J. Bacteriol.1981; 148: 265-73; Keidel et al., Eur. J. Biochem.1992; 204: 1141-8; and Durfee et al., J. Bacteriol.2008; 190: 2597-606.

[0058] E. coli strains suitable for the invention may be obtained from the various sources as described in the literature, or biological material depositories such as theAmerican Type Culture Collection (ATCC). Some of these strains are also available from commercial suppliers, which include New England Biolabs (NEB), Stratagene, Invitrogen, Lucigen, Epicentre, and Qiagen. Some orthogonal T7 replisomes of the invention are based on E. coli strain BW25113 as exemplified herein. Genotype of E. coli strain BW25113 is known, which is F- LAM- rrnB3 DElacZ4787 hsdR514 DE(araBAD)567 DE(rhaBAD)568 rph-1. The complete genome sequence of strain BW25113 is described in Grenier et al., Genome Announc.2014 Sep-Oct; 2(5): e01038-14, with a GenBank accession number CP009273. This E. coli strain has been well characterized, and is available commercially available, e.g., from vendors such as Horizon Discovery (Cambridge, UK).

[0059] Some orthogonal replication systems of the invention contain one or more transformed vectors expressing the component proteins of the engineered T7 replisome, including an engineered T7 DNA polymerase with the ∆28 deletion and the N520 substitution (e.g., N520M), a T7 lysozyme and T7 RNA polymerase fusion, T7 primase-helicase, and the ssDNA-binding protein. In some embodiments, the employed T7 lysozyme in a lysozyme-T7 RNA polymerase fusion containing a hydrolase deficient mutant lysozyme (e.g., C131S mutant) as exemplified herein. In some embodiments, the employed T7 primase-helicase is expressed from a modified sequence that eliminates the internal start codon for Gp4B (helicase) translation. As a result, only the primase-helicase fusion protein is expressed from the Gp4 sequence. Examples of E. coli orthogonal replication systems include T7Rep-1, T7Rep-2, T7Rep-3, T7Rep-4, T7Rep-5, T7Rep-6, and T7Rep-7exemplified herein.

[0060] In some E. coli orthogonal DNA replication systems of the invention, the T7 replisome contains one or both of the modified Gp4 / Gp2.5 fusion and the modified Gp3.5 / Gp1 fusion, plus a wildtype T7 DNA polymerase (e.g., T7Rep-1) or a T7 DNA polymerase containing the ∆28 deletion (e.g., T7Rep-2). In some systems, the T7 replisome contains an engineered T7 DNA polymerase having the ∆28 deletion and a N520 substitution such as N520M (e.g., T7Rep-3). In some systems, the engineered T7 DNA polymerase of the replisome additionally contains one or more mutations beyond the ∆28 deletion and the N520 substitution relative to the wildtype T7 DNA polymerase. In some of these systems, the engineered T7 DNA polymerase of the replisome additionally contains a P560V substitution (e.g., T7Rep-4). In some systems, the engineered T7 DNA polymerase of the replisome additionally contains P560V and V443K substitutions (e.g., T7Rep-5). In still some other systems, the engineered T7 DNApolymerase of the replisome additionally contains P560V, V443K and Y24F substitutions (e.g., T7Rep-6). In still some other systems, the engineered T7 DNA polymerase of the replisome additionally contains P560V, V443K, Y24F and Q585R substitutions (e.g., T7Rep-7).

[0061] Transformation of vectors encoding the engineered T7 replisome into the E. coli cell, culturing and maintenance of transformed strains, and expression of the protein components can all be readily performed in accordance with standard techniques that are routinely practiced in the art and / or the specific protocols that are exemplified herein. For example, the calcium chloride method is a suitable method for transforming an E. coli host cell with expression vectors bearing a compatible origin of replication, e.g., vectors containing a pMB1 ori such as pBR322. Alternatively, electroporation may be used to transform the vectors into the E. coil cell. The transformed cells are selected by growth on an antibiotic, e.g., tetracycline (tet) or ampicillin (amp), to which they are rendered resistant due to the presence of tet and / or amp resistance genes on the vector. See, e.g., Sambrook, et al., Molecular Cloning: A Laboratory Manual (3rdEdition, 2000); Brent et al., Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (ringbou ed., 2003); and Neumann et al., EMBO J. 1982; 1:841-5. V. Polynucleotides, vectors and host cells

[0062] The invention additionally provides polynucleotides (DNA or RNA) which encode the various components of the engineered T7 replisomes described herein. They include, without limitation, messenger RNA (mRNA), DNA / RNA hybrids, or synthetic nucleic acids. The nucleic acids of the invention may be single-stranded, or partially or completely double-stranded (duplex). Duplex nucleic acids may be homoduplex or heteroduplex. Vectors containing the polynucleotides and engineered host cells harboring the vectors are also provided in the invention. The vectors of the invention are not subject to any particular limitation, and may be, for example, bacteriophages, plasmids, cosmids or phagemids. Examples of recombinant bacteriophage or phagemid vectors include that based on a filamentous phage such as M13. Plasmid vectors include those based on plasmids from, e.g., E. coli (e.g., pBR322, pBR325, pUC118 and pUC119), plasmids from Bacillus subtilis (e.g., pUB110 and pTP5), and plasmids from yeasts (e.g., YEp13, YEp24 and YCp50). The vectors can also include animal viruses such as retroviruses, vaccinia viruses and insect viruses (e.g., baculoviruses).

[0063] In some preferred embodiments, the vectors of the invention are expression constructs that can express one or more of the T7 replisome protein components in an E. coli host cell. These expression constructs are based on vectors containing a replication origin (ori) and transcription regulatory elements (e.g., a promoter) that function in E. coli cells. Examples of such base vectors include the ColE1 family of plasmids which contain the pMB1 ori, e.g., pBR22, pUC, pET and pMB1. In the expression vectors, the polynucleotide encoding the replisome components is typically operably linked to the transcription regulatory elements including the promoter. In various embodiments, the promoter in the expression vectors can be, e.g., the tetracycline promoter, the Trp promoter, the T7 promoter, the lac promoter, the recA promoter, the λ promoter and the lpp promoter. In addition to the promoter sequence, the expression vectors can contain, if desired, an enhancer, a splicing signal, a poly(A) addition signal, a ribosome binding sequence (SD sequence), a selective marker and the like. Examples of selective markers include the tetracycline resistance gene, the carbencillin resistance gene, the dihydrofolate reductase gene, the ampicillin resistance gene and the neomycin resistance gene. The expression vectors of the invention may additionally include a polynucleotide having a nucleotide sequence encoding an amino acid sequence for enhancing translation and / or a polynucleotide having a nucleotide sequence encoding a peptide sequence for purification. For example, the expression vectors can employ a translational enhancer element (TEE) sequence (see, e.g., Batten et al., FEBS Lett.580:2591-7, 2006). Specific examples of expression vectors of the invention include vectors that express an engineered T7 DNA polymerase (e.g., pC0260, pC0261, pC0262, pC0263, pC0264, pC0265 and pC02660), a vector expressing a modified Gp4 / Gp2.5 fusion (pCD0280), and a vector expressing a mutant Gp3.5 / Gp1 fusion (pCD0275).

[0064] Host cells of the invention are genetically engineered (transduced, transformed or transfected) with the recombinant vectors or expression constructs disclosed herein for production of the T7 replisome protein components. The host cells to which the vectors are introduced can be any of a variety of host cells well known in the art, e.g., bacteria (e.g., E. coli), yeast cell, or animal cells such as CHO, COS or 293 cells. The host cell for production or expression of a construct of the invention are preferably an E. coli cell as exemplified herein. In some of these embodiments, the host cells are the above-described E. coli based orthogonal DNA replication systems which harbor a complete T7 replisome. They also encompass cells, including E. coli cells andother cells, that contain one or more vectors that encode less than the complete T7 replisome as noted above. These include, e.g., a vector encoding just the engineered T7 DNA polymerase, the mutated T7 primase-helicase, or the modified T7 lysozyme. In other embodiments, the host cells of the invention can also be non-E. coli prokaryotic cells, higher eukaryotic cells such as mammalian cells, or lower eukaryotic cells such as a yeast cell. The selection of an appropriate host is within the scope of those skilled in the art and also exemplified in the Examples herein. In addition to E. coli cells, other representative examples of appropriate host cells of the invention include, but limited to, bacterial cells such as Streptomyces, Salmonella tvphimurium; fungal cells such as yeast; insect cells such as Drosophila S2 and Spodoptera Sf9; animal cells such as CHO, COS or 293 cells; plant cells; or any suitable cell already adapted to in vitro propagation or so established de novo.

[0065] Introduction of the vector or expression construct into the host cell can be effected by a variety of methods with which those skilled in the art will be familiar. These include, but not limited to, transformation with calcium chloride (heat shock), transformation via electroporation, calcium phosphate transfection, or DEAE-Dextran mediated transfection. See, e.g., Brent et al., supra. Expression and, if desired, purification, of an encoded T7 replisome component in a transformed or transfected host cell can be carried out in accordance with any of the routinely practiced methods in the art, e.g., Sambrook et al., supra; and Brent et al., supra. For example, to generate the E. coli host cells of the invention, a vector which harbors and expresses a T7 replisome component can be introduced into a suitable host, e.g., E. coli BW25113 cell as exemplified herein. The engineered host cells can be cultured in conventional nutrient media modified as appropriate for activating promoters, selecting transformants or amplifying particular polynucleotide sequences such as the sequences encoding a component of the engineered T7 replisomes. The culture conditions for particular host cells selected for expression, such as temperature, pH and the like, will be readily apparent to the ordinarily skilled artisan. VI. Evolving target polypeptides of interest

[0066] The invention further provides methods of evolving proteins or polypeptides of interest, using the E. coli based orthogonal replication systems of the invention. In these methods, a vector harboring a polynucleotide sequence encoding the target polypeptide of interest is introduced into an E. coli orthogonal replication system (e.g.,which already contain an engineered T7 replisome of the invention, e.g., the exemplified T7Rep-7 system. Alternatively, the vector harboring the target polynucleotide sequence can be co-transformed into the base E. coli strain (e.g., BW25113 cells) with constructs expressing the components of the engineered T7 replisome. Preferably, the vector is a circular plasmid that can be stably replicated in the E. coli host cell. Typically, to ensure replication of the target polynucleotide sequence by the T7 replisome, the vector harboring the target polynucleotide sequence vector contains a T7 origin of replication, e.g., the φOR ori. As demonstrated herein, transformation efficiency of such circular replicon plasmids into the E. coli replication systems of the invention was found to be several orders of magnitude higher than transformation efficiency achieved with linear plasmids into the orthogonal replication systems known in the art (e.g., OrthoRep and EcORep).

[0067] Once the vector encoding the target polynucleotide sequence has been introduced into the E. coli orthogonal replication system of the invention, the cells are cultured in accordance with the procedures described herein and / or standard E. coli culturing conditions well known in the art. The culturing condition should enable the engineered T7 replisome to replicate the target polynucleotide sequence with enhanced mutation rates. As exemplified herein, the cells can be cultured at 37 °C in a 2YT media with appropriate maintenance antibiotics that correspond to antibiotic resistance markers that are present in the vectors transformed into the cells. At various time points (e.g., 4 hrs, 8 hrs, 12 hrs, 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or longer), aliquots of cells are removed from the culture and subject to further analysis of the evolved target polynucleotide, e.g., to identify an enhanced property of the target polypeptide or a desired phenotype of the cell that is indicative of an evolved property of the target polypeptide. For example, an evolved target sequence encoding antibiotic resistance can be analyzed by selecting cells from the aliquoted cells (e.g., on LB agar plates) in the presence of increasing concentrations of the corresponding antibiotic, as exemplified herein for evolving a β-lactamase encoding polynucleotide sequence. Screening and selection of evolved sequences encoding other functional polypeptides (e.g., enzymes or antibodies) can be similarly performed. As described below, subsequent analysis of evolved target polynucleotides encoding these other polypeptides in the cultured cells can all be performed with many known assays that are suitable for monitoring the specific functional or structural property to be evolved in the target polypeptide.

[0068] Target polypeptides or proteins of interest that can be evolved with the methods and kits of the invention include any natural or non-natural proteins known in the art. These include, e.g., any polypeptides with medical or industrial applications such as therapeutic proteins, antibodies and enzymes. In some embodiments, the polypeptide or protein of interest is one that encodes a therapeutic protein. Examples of therapeutic proteins include factor VIII, factor IX, β-globin, interleukins, interferons, cytokines, insulin, erythropoietin, tissue plasminogen activator (tPA), urokinase, streptokinase, granulocyte colony stimulating factor (G-CSF)), and thrombopoietin (TPO). In some embodiments, the polypeptide or protein of interest is a therapeutic antibody. Suitable therapeutic antibodies include, e.g., Trastuzumab, Rituximab, Daclizumab, Adalimumab, Pembrolizumab, Atezolizumab, Avelumab, Glofitamab, Donanemab, Lecanemab, Tislelizumab, Nirsevimab, Tremelimumab, Tezepelumab, Tisotumab vedotin, Bimekizumab, Evinacumab, Naxitamab, Inebilizumab, Teprotumumab, Crizanlizumab, Romosozumab, Emapalumab, Elranatamab, and Pozelimab. In some embodiments, the polypeptide or protein of interest is an enzyme. Examples of suitable enzymes include amylases, lipases, proteases, oxidoreductases, transferases, hydrolases, lyases, isomerases, ligases, and translocases.

[0069] Other than therapeutic protein, enzymes and antibodies, various other types of target polypeptides of interest can also be similarly evolved with the methods disclosed herein. These include receptor proteins (e.g., Fc-receptor or its fragments), cytokines, hormones, toxins, enzyme inhibitors, DNA-binding domains (e.g., zinc fingers), protein ligands or any other peptides or proteins (e.g., protein A or protein L). Some specific examples of such target molecules include, e.g., tumor necrosis factor α (TNFα), vascular endothelial growth factor (VEGF), transforming growth factor (TGF), fibroblast growth factor (FF), platelet derived growth factor (PDGF), angiotensin II, IL- 2, IL-10, insulin-like growth factor, insulin receptor, MHC proteins (e.g. class I MHC and class II MHC protein), CD3 receptor, cytokine receptors (e.g., interleukin receptors), G-protein coupled receptors, chemokine receptors, Maurotoxin, Agitoxin, and transcription factor Zif268. Selection and screening of evolved mutants of any of these target molecules can be readily performed in accordance with methods exemplified herein or well known in the art.

[0070] In some other embodiments, suitable polypeptides of interest are specific membrane proteins or other intracellular proteins. Examples of membrane proteins include, but are not necessarily limited to adrenergic receptors, serotonin receptors,low-density lipoprotein receptor, CD-18, sarcoglycans (which are deficient in muscular dystrophy), etc. Examples of intracellular proteins include fumarylacetoacetate hydrolase (FAH), antiviral proteins (e.g., proteins that can provide for inhibition of viral replication or selective killing of infected cells), and structural protein such as collagens.

[0071] Depending on the specific target polypeptide of interest, screening and / or selection of evolved molecules with desired activities or characteristics can be performed with any known assays or methods that have been used to measure and monitor one or more specific biological or pharmaceutical properties of the target polypeptide. In some embodiments, the selection can be based on an enhanced antibiotic resistance, as exemplified herein for evolving a β-Lactamase encoding sequence for resistance to beta-lactam antibiotics such as aztreonam, cefotaxime, ceftazidime, and cefepime. In some embodiments, the selection can be based on an enhanced binding affinity for a binding partner. This includes binding of a target antibody sequence to a cognate antigen, binding of a ligand to a cognate receptor molecule, or binding of an enzyme to a substrate. In some embodiments, the selection can be based on an improved catalytic activity, e.g., digestion of a substrate by a target enzyme. As a specific exemplification, the selection can be based on monitoring survival of a cell expressing a target polynucleotide sequence. For example, this can be performed by linking expression of the target polypeptide (e.g., a transcription regulator) to transcription of an essential gene or formation of an essential metabolite. In another exemplified embodiment, selection of an evolved binding partner (e.g., an antibody) can be performed by fluorescence or magnetism activated cell sorting. This involves display of the protein / antibody of interest on the cell surface, followed by binding of a fluorophore or magnetic species. In a further exemplified embodiment, selection of an evolved enzyme molecule can be performed by linking the enzyme function to a fluorescent readout in a high throughput format. For example, this can be achieved via combining the orthogonal replication system with microfluidic sorters, e.g., fluorescence activated droplet sorting (FADS). EXAMPLES

[0072] The following examples are provided to further illustrate the invention but not to limit its scope.Example 1 Establishing an orthogonal replication system based on the bacteriophage T7 replisome

[0073] While no natural orthogonal replisome-replicon pairs exist for E. coli, there are lytic coliphages that encode for their own replisome.(21)Thus, to establish an orthogonal replisome-replicon pair, we capitalized on the vast body of literature detailing the replisome of the lytic bacteriophage T7 (Fig.1).(22–24)The replication of the linear double stranded genome of T7 phage is initiated by priming of its replisome by T7 RNA polymerase (Gp1) at a T7 origin of replication (Fig.1, A).(25) T7 origins of replication rely on the orthogonality of the T7 RNA polymerase / T7 promoter pair to E. coli polymerase / promoter pairs for specificity in recruiting the replisome (Fig.1, B).(26)The orthogonal T7 transcript primes the T7 DNA polymerase (Gp5), which further recruits the T7 helicase (Gp4A) and helicase / primase (Gp4A) fusion protein to the transcription bubble. The host factor thioredoxin (TrxA) binds to the T7 DNA polymerase and serves as a β-sliding clamp. Replisome assembly initiates concerted leading and lagging strand synthesis, which further requires the presence of the T7 single-stranded DNA binding protein (Gp2.5) for stabilizing the lagging strand replication loop (Fig.1, C).(25)

[0074] To generate an orthogonal T7 replication system in E. coli, we generated a strain based on E. coli BW25113 (parent strain of Keio collection) in which all replisome genes are supplied in trans using plasmid-based expression systems (Table 1).(27)As described in the literature, we observed severe toxicity for expression of Gp4B (helicase), but less toxicity for Gp4A (helicase-primase fusion). Since Gp4B is not essential for replication of T7 phage, we removed the internal start codon of its open reading frame via a M64L substitution. Expression and functionality of the exogenous proteins was confirmed by infection of the replisome strain with four T7 phages, each with one of the essential replisome genes knocked out. While replication of each knock-out phage proved successful, initial attempts at transforming the replisome strain with a circular plasmid comprising both the primary T7 origin of replication proved unsuccessful.(30)

[0075] We reasoned that additional T7 genes might be required to establish an orthogonal replication system, but not strictly required for phage propagation.(31)While not strictly essential to replication, phages with a T7 lysozyme knockout (Gp3.5) are known to display severe replication deficiencies.(32)T7 lysozyme is a bifunctionalenzyme that displays both cysteine hydrolase activity on N-acetylmuramoyl-L-alanine amide bonds of peptidoglycan (major constituent of the bacterial cell wall), as well as inhibition of transcription by T7 RNA polymerase.(33, 34)The concerted expression of T7 lysozyme with the other replication proteins (Class II genes of bacteriophage T7) implies a regulatory function, in which its accumulation seems to signify a switch from transcription to replication in the phage life cycle.(35)Mechanistically, this can be rationalized based on the mechanism of inhibition by T7 lysozyme, which stabilizes T7 RNA polymerase in its binding conformation and prevents the switch to the elongation conformation required for gene transcription.(36–39)This leads to abortive cycling around the RNA polymerase binding site and to the formation of short RNA ‘primers’ in lieu of the transcription of full genes. The constant opening of the DNA double helix by the transcription bubble enables priming of the T7 DNA polymerase and stimulates assembly of the T7 helicase / primase at the origin.

[0076] To maximally benefit from this regulatory effect, we introduced catalytically inactive lysozyme (C131S, hydrolase deficient) as a translational fusion to the T7 RNA polymerase.(40)Based on literature precedent, we changed the origin of replication from the primary origin of replication to a secondary origin φOR for which replication stimulation by T7 lysozyme was reported to be highest (Fig.1, A).(35)Finally, we chose a previously described exonuclease deficient clone of T7 DNA polymerase (Δ28, T7 Sequenase 2.0) for initial studies as it had been utilized preferentially in reports on replisome-based in vitro replication assays.(41–43)Implementation of these changes to the replisome strain enabled transformation of a plasmid carrying a T7 origin of replication (pOR-1; Table 1), maintaining stable replication for more than 2 weeks (Fig. 1, D and E). The strain shows robust growth (Tdouble = 29 min) and reaches high cell densities (OD600: 7.2; ~7 × 109cells ml-1) in liquid culture (Fig.1, E). Mutation rates of the initial replisome strain were determined by fluctuation analysis, and showed increased mutagenesis of the T7 origin plasmid (2.81 × 10-7spb per generation) without affecting the genomic mutation rate of the host (3.66 × 10-10spb; Table 4). The copy number of the T7 origin plasmid was determined to be 8.5 based on quantitative PCR, which translates to an adjusted mutation rate of 3.31 × 10-8spb (Fig.2, G; and Table 3). Finally, the transformation efficiency of the circular replicon plasmid into the replisome strain was determined to be 2.4 × 1010cfu μg-1, orders of magnitude higher than the transformation efficiency for linear plasmids of alternative orthogonal replication systems (i.e., OrthoRep and EcORep).(12, 15)Example 2 Engineering T7 DNA Polymerase Mutants with Increased Mutation Rate

[0077] The in vivo mutation rate of bacteria is controlled by three distinct mechanisms that collectively account for the high fidelity (10-9to 10-10spb) of DNA replication. Specifically, base selection in the T7 DNA polymerase active site accounts for a fidelity of ~10-5spb, 3’-5’ exonucleolytic proofreading further increases fidelity by ~10-2spb, and host mismatch repair is responsible for the remaining ~10-3spb of fidelity. In our efforts to devise a mutagenic orthogonal replisome, we began with a T7 DNA polymerase mutant for which the 3’-5’ proofreading exonuclease activity was removed using a previously described 28 amino acid mutant (“Δ28”; Fig.2, A). To approximate the extent to which this deletion contributed to mutagenesis in vivo, we carried out fluctuation analysis on both the WT polymerase (1.4 × 10-9spb), as well as on the Δ28 mutant (3.31 × 10-8spb), yielding an increase in mutation rate of 25-fold (Fig.2, B). This is in agreement with the magnitude of reported increases in mutation rate in vitro (WT: 2.2 × 10-6spb, Δ285.4 × 10-5spb, 24.5-fold). Note, the fidelity is ~3 orders of magnitude higher in vivo due to proof-reading by host mismatch repair.(44)

[0078] Since DNA mismatch repair cannot be altered without compromising host fitness, we next targeted the fidelity of base selection in the polymerase active site to further increase mutagenicity. Initial efforts focused on the identification of sites that contribute to changes in fidelity of homologous polymerases (e.g., T. aquaticus polA, E. coli polA, and S. cerevisiae mip1), and homology mapping of the corresponding residues onto T7 DNA polymerase. On the basis of previous reports, we targeted 11 sites of T7 DNA polymerase and generated single site-saturation libraries at each position (i.e., R429X, V443X, R444X, L479X, E480X, N520X, A521X, K522X, T523X, F524X, Y530X, P560X, and N611X). We next transformed a replisome strain deficient in T7 DNA polymerase (T7Rep-ΔGp5, Table 2) with each site-saturation library and carried out reversion assays in which the polymerase libraries were challenged to revert a premature ochre stop codon in the TEM-1 β-lactamase gene (E26*ochre) on the T7 origin plasmid (pOR-3, Table 1), thus restoring antibiotic resistance. After demonstrating survival on the maintenance antibiotic to select for functional DNA polymerases, a reversion selection for mutagenicity on carbenicillin was performed. Surviving clones were sequenced using Illumina amplicon sequencing. Clones that were found to be enriched in the reversion assays were individually subjected to a fluctuation analysis which identified three optimal mutations (N520M, P560V, andV443K, Fig.2, B) that we mapped onto the Δ28 mutant. The mutations at the three identified sites are hypothesized to target distinct mechanisms for increasing the mutation rate. N520M is a mutation in the O-helix, which is in direct contact with the incoming nucleotide in the active site. Mutations at this site likely decrease geometric constraints that contribute to increased base pairing selectivity. P560V is located in a conformationally flexible region of the fingers domain that contacts the O-helix. Mutations in these ‘hinge’ loops that connect the alpha helices of this domain are known to affect the kinetics of conformational changes upon binding of the correct / incorrect nucleotides, thus impacting fidelity. Finally, V443K is a mutation at a site that is in direct contact with the DNA double helix at the replication fork. Such mutations likely decrease base pairing constraints by increasing the flexibility at the replication fork, which minimizes strain due to pairing of mismatched bases. The identified mutations increased the mutation rate of exonuclease deficient T7 DNAP to 6.75 × 10-7spb for the best single mutant, 4.26 × 10-6spb for the best double mutant, and 1.04 × 10-5spb for the best triple mutant (Fig.2, C).

[0079] Next, we constructed a random mutagenesis library by error-prone PCR on top of the polymerase triple mutant and subjected it to the same antibiotic reversion assay. Sequencing to identify codon positions at which mutations were enriched yielded an additional 7 sites for which individual single-site saturation libraries were constructed and screened (Y24X, P84X, 118X, R171X, 433X, L437X, Q585X, and 678X). The two most beneficial mutations identified were Y24F and Q585R (Fig.2, B). Notably, the first site is located in the exonuclease domain and was found to reverse growth deficiencies observed for the single-, double- and triple mutants (Table 5) but did not show a change in mutation rate (1.1 × 10-5spb). The last mutation was found in the Palm domain of the polymerase and yielded a mutation rate of 1.73 × 10-5spb for the quintuple mutant (Fig.2, C; Table 3).

[0080] The mutational spectrum of the quintuple mutant was investigated by Illumina sequencing using the NovaseqX platform (2 replicates of 600 million paired end reads each) and analyzed using a custom analysis pipeline. The observed mutation rate across the two replicates was calculated as 1.93 × 10-5spb, in good agreement with the results from fluctuation analysis. Heatmaps of the mutational spectrum were generated by calculating the mutational frequency of each nucleotide substitution across the two replicate data sets. A fairly unbiased mutational spectrum for both transition mutations and transversion mutations was observed.

[0081] Orthogonality to the genome by fluctuation analysis using rifampicin resistance assays was confirmed for the final mutant, which displayed a genomic mutation rate of 3.9 × 10-10spb (Fig.2, D; Table 4). Stability of the quintuple mutant was further confirmed by passaging the strain for two weeks, after which sequencing of the quintuple mutant of the replisome strain showed no changes in sequence. We further determined that plasmid sizes of up to 13kb in size can be maintained (pOR-6) with comparable mutation rates (1.53 × 10-5spb).

[0082] In addition to the top mutants identified for substitutions at each of the noted positions (e.g., N520M at N520, P560V at P560, V443K at V443, Y24F at Y24, and Q585R at Q585), T7 DNA polymerase mutants with other substituting residues at these positions have also been observed and confirmed via fluctuation analysis to increase mutation rates. These include (1) mutants with N520 substituted with small hydrophobic residue I, L, or C, or small polar residue T, N, D, or E; (2) mutants with P560 substituted with small hydrophobic residue M, I, L, G, or A, as well as K; (3) mutants with V443 substituted with uncharged residue S, T, N, or Q, or positively charged residue R or H, as well as C; (4) mutants with Y24 substituted with aromatic residue W; and (5) mutants with Q585 substituted with residue T, R, E or C. Example 3 Continuous Evolution of TEM-1 β-Lactamase

[0083] Next, to demonstrate the utility of the mutagenic T7 orthogonal replisome assisted continuous laboratory evolution (T7-ORACLE) system for evolving protein function, we applied it to the continuous evolution of TEM-1 β-lactamase for an expanded substrate scope and improved activity against monobactam (aztreonam) and cephalosporin antibiotics (cefotaxime, ceftazidime, and cefepime, Fig.3, A-C). To this end, the most mutagenic (T7rep-7; quintuple mutant) replisome strain was transformed with a T7 origin plasmid encoding chloramphenicol resistance (CamR) as a maintenance antibiotic marker, as well as the blaTEM-1β-lactamase gene (Fig.3, A). Initial plating of the strain on antibiotic concentrations that approximate the MIC of TEM-1 against the respective antibiotics of interest (0.1 μg ml-1for aztreonam, cefotaxime, and cefepime; 0.4 μg ml-1for ceftazidime) served as the starting point for the evolution experiment. For each experiment 96 colonies were picked and the cells were passaged under increasing concentrations of antibiotic for 6 days with an overall increase in concentration of ≥1000-fold (100 μg ml-1for aztreonam, cefotaxime, cefepime; 500 μg ml-1for ceftazidime; Fig.3, B). At the end of the experiment, the coding sequences of the β-lactamase genes, as well as the pAmpR promoter region were amplified by PCR and analyzed by sequencing (Fig.3, D).

[0084] Convergence of the mutations across the 96 replicates was observed for all antibiotics, with each of the enriched mutations having previously been described in clinical isolates of extended spectrum β-lactamases and / or in laboratory evolution experiments.(45, 46)Both the nature, and the combination of mutations differ for the four distinct β-lactam antibiotics. In addition, the entire 96-well plate for each evolutionexperiment was pooled after each passage and subjected to nanopore amplicon sequencing to elucidate the order and establish a time-course of when mutations occur (Fig.4, A and B; and Figs.6 and 7). For instance, the evolution of cefotaxime resistance (Fig.4, A) begins to enrich sequences comprising the G238S substitution from the first passage, in agreement with the substitution being associated with increased hydrolysis of cefotaxime. After passage 3, only clones that acquire a secondary mutation, E104K, survive. E104K is generally associated with mutations G238S or R164H / S (see below) where in combination it drastically increases resistance to cefotaxime, or ceftazidime and aztreonam, respectively.(45, 47) Starting at passage 4, mutations that increase the strength of the weak promoter pAmpR are observed. Finally, over passages 4, 5, and 6 one of four accessory mutations are observed for the majority of cultures (M182T, L49M, T265M, or G92D), all of which have been found as secondary mutations in both clinical isolates and evolution experiments.(45) These mutations are more distant from the active site and have no direct effect on catalytic activity. M182T increases thermodynamic stability and helps prevent aggregation. Similarly, G92D and T265M were identified in evolution campaigns that select for stabilizing and compensatory mutations.(45) Finally, the role of L49M is unknown, but the fact that it is observed frequently and specifically in the context of cefotaxime resistance supports an adaptive effect.(45)

[0085] In the ceftazidime resistance time course experiment, mutations in the promoter region are observed from the outset (Fig.5, B). This can be attributed to the higher initial activity against this β-lactam, making increased expression levels a viable solution for improved fitness. Additionally, mutations R164H / S (increases accessibility of β-lactams with larger side-chains to the active site) are enriched from the outset, along with the aforementioned E104K secondary mutation. After passage 3, additional mutations (i.e., M182T, E240K, and D179G) are observed. E240K is a secondary mutation that is commonly observed in combination with R164S / H, where it significantly increases activity against ceftazidime and aztreonam.(45, 47) D179G has been observed in evolution campaigns using ceftazidime as a selective agent.(45) Over the last three passages, the percentage of D179G mutations is steadily decreasing, indicating that the mutation is being outcompeted at higher antibiotic concentrations. It is interesting to note that this one-week experiment perfectly recapitulates the vast bodyof work on extended-spectrum β-lactamases collected over decades, highlighting the power of the T7 evolution system in interrogating complex evolutionary fitness landscapes.

[0086] A major advantage of T7-ORACLE compared to other continuous evolution systems is that maintenance of a circular replicon enables high transformation efficiencies and integration of continuous hypermutation with pre-diversified libraries (e.g., site-saturation mutagenesis and gene shuffling). Under constant selective pressure, laboratory evolution tends to follow a path of least resistance across a hypothetical fitness landscape. This can be illustrated for the evolution of TEM-1 for activity against cefotaxime, where replicate evolution experiments broadly converge to the same solution (Fig.3, D-E). Conceptually, this can be described as evolution trajectories on a hypothetical fitness landscape being trapped in local fitness maxima as opposed to sampling the entire landscape (Fig.5, A–C). To overcome this limitation, we generated a site saturation library of TEM-1 (E104X, Y105X, G238X, E240X, and R244X), covering the five residues that line the enzyme’s active site. Transformation of the library plasmid pOR-5 (Table 1) into the T7-ORACLE strain yielded full library coverage (3.4 × 107-member library).96 clonal starting points were picked after initial selection (0.1 μg ml-1cefotaxime) and passaged under increasing selection stringency. After 5 days, the replicate evolution experiments displayed vast diversity in their active site residues and differed from the solution for evolution of wild-type TEM-1 (Fig.5, D-E). Notably, while E104K was one of the most important mutations for cefotaxime resistance in wild-type TEM-1, 16 distinct solutions were observed at this site in the evolution of the site-saturation library (Fig.5, D). Similarly, 14 distinct solutions were observed at site 238, seven for site 105, and seven for site 240. No substitutions for R244 were identified, indicating that the residue is essential for activity. Secondary mutations related to stability and expression levels followed similar trends as for the WT enzyme evolution (e.g., pAmpR, M182T, L49M, T264M, and G92D).

[0087] A second evolution experiment was carried out where, instead of 96 individual clones, the entire library was used as a starting point in each of the replicates. In this case the active site mutations converged again, with >80% sequence identity at four of the five sites (E104F, Y105W, G238S, and R244R) and divergence only being observed at site 240 (Fig.5, D and F). In addition, starting from the entire libraryallowed increasing the cefotaxime concentration even further (500 μg ml-1) without any loss in viability, supporting the fact that evolution under these conditions can identify global maxima on the protein’s fitness landscape. This is further corroborated by the fact that the active site residues that converged in the experiment starting from the site- saturation library, when tested clonally, were found to be more active than the clones that converged from the evolution starting from WT TEM-1. Example 4 Materials and methods

[0088] Some specific techniques and protocols for practicing methods of the invention are exemplified below.

[0089] Culture conditions and reagents. All E. coli cultures were grown in 2YT media unless otherwise specified. Cloning was carried out in E. coli DH5a. All replisome strains were constructed in E. coli BW25113 (F- LAM-::rrnB3, ΔlacZ4787, hsdR514, Δ(araBAD)567, Δ(rhaBAD)568 rph-1), the Keio collection parent strain.(50) Antibiotics were used at the following concentrations unless specified otherwise: Kanamycin (50 µg / ml), streptomycin (100 µg / ml), gentamycin (20 µg / ml), carbenicillin (100 µg / ml), chloramphenicol (25 µg / ml), and rifampicin (50 µg / ml).

[0090] DNA cloning. Plasmids used in this study are listed in Table 1. E. coli strain NEB5a (NEB) was used for all DNA cloning steps. All primers used in this study were purchased from Integrated DNA Technologies, USA. All enzymes for PCR and cloning were obtained from NEB. Plasmids were assembled by Gibson assembly. To clone the T7 phage replisome constructs (Table 1), DNA fragments encoding the open-reading frames of Gp1, Gp2.5, Gp4, Gp3.5, and Gp5 were purchased from IDT.

[0091] Growth curves. Growth curves of the replisome strains (Table 2) transformed with pOR-3 (Table 1) were measured in a 96-well plate format using the BioTek LogPhase 600 microbiology reader. Cultures of 150 μl in volume were grown at 37 °C at a shaking speed of 800 (a.u.). Terrific broth was used as the culture media and was supplemented with antibiotics kanamycin (22.5 µg / ml), streptomycin (90 µg / ml), gentamycin (18 µg / ml), and chloramphenicol (25 µg / ml). Optical density was recorded every 10 minutes and samples were measured in 12 replicates for each experiment.

[0092] Transformation efficiency determination. The transformation efficiency is defined as the theoretical number of colony-forming units (cfu’s) produced by transforming 1 µg of plasmid DNA into a given volume of competent cells. In practice, this efficiency was calculated by transforming 100 pg of purified plasmid under idealized conditions. Electrocompetent cells of strain T7rep-2were generated by back- diluting an overnight culture 1000-fold, followed by growing the culture (500 ml) to an optical density (OD600) of 0.4 at 37°C.4 cycles of centrifugation (7 min, 3000×g, 16°C) and media exchange with a full culture volume of 10% glycerol in water were used to remove remaining salt from the media. The final OD600of the competent cells was adjusted to ~200.50 µl of competent cells were mixed with 100 pg of purified DNA of pOR-2 (Table 1) and electroporated using pre-chilled electroporation cuvettes. Cultures were recovered in 250µl of SOC media for 1 hour, and diluted 10-fold before plating 30 µl on pre-warmed plates (37°C), supplemented with kanamycin (45 µg / ml), streptomycin (90 µg / ml), gentamycin (18 µg / ml), and carbenicillin (100 µg / ml). The effective transformation efficiency was calculated by normalizing the number of cfu’s to a hypothetical DNA concentration of 1 µg.

[0093] Site-saturation mutagenesis. NNK site-saturation libraries were generated by amplifying target sequences using primers for which one codon is randomized (NNK codon), followed by plasmid ligation using Gibson assembly. The cloned libraries were transformed into NEB® 10-beta electrocompetent E. coli cells and plasmid DNA was isolated and purified. Codons that were randomized in the coding sequence of T7 DNA polymerase comprise R429, V443, R444, L479, E480, N520, A521, K522, T523, F524, Y530, P560, N611 in plasmid pCD0261 (Table 1; Δ28 exonuclease deficient mutant), and Y24, P84, M146X, R171, A433, L437, Q585, and R678 in the coding sequence of T7 DNA polymerase of plasmid pCD0264 (Table 1; triple mutant). The numbering of all residues is based off of the wild type T7 DNA polymerase sequence.

[0094] Reversion assays of single site-saturation library of T7 DNAP. Purified DNA of the NNK single site-saturation T7 DNA polymerase library plasmid (see above) was co-transformed with plasmid pOR-3 (Table 1) by electroporation into strain T7Rep-ΔGp5 (Table 2) and plated on LB agar supplemented with maintenance antibiotics (kanamycin (45 µg / ml), streptomycin (90 µg / ml), gentamycin (18 µg / ml), and chloramphenicol (25 µg / mL). Cells were scraped and resuspended using Dulbecco'sPhosphate Buffered Saline (DPBS) and 1ml to 10 ml of OD600 = 1 re-plated onto LB agar plates of appropriate size supplemented with maintenance antibiotics, as well as selection antibiotic (carbenicillin, 100 µg / ml). Survival of clones is contingent upon reversion of a premature ochre stop codon in the TEM-1 β-Lactamase gene on pOR-3 (i.e., statistical selection for mutagenic T7 DNA polymerase clones).

[0095] Illumina amplicon sequencing. Cells of carbenicillin resistant clones from the single site-saturation libraries (see above; plated on selective media) were scraped using Dulbecco's Phosphate Buffered Saline (DPBS) buffer and their plasmid DNA was purified using the ZymoPURE II Plasmid Midiprep kit. A 500 base pair sequence comprising the single-site NNK library in the T7 DNA polymerase gene was amplified from the isolated DNA (1 µg of template for 50 µl PCR reaction) using primers functionalized with adaptors for Illumina sequencing. The product was PCR purified using the Monarch® PCR & DNA Cleanup Kit (NEB), and the DNA quantified using the qubit DNA fluorometry method. Samples were analyzed using Azenta Genewiz Amplicon EZ sequencing (50,000 reads) and the codon distribution analyzed for enrichment to identify mutagenic mutants.

[0096] Error-prone PCR library reversion assay. Purified DNA of the error-prone library plasmid (see above) was co-transformed with plasmid pOR-3 (Table 1) by electroporation into strain T7Rep-ΔGp5(Table 2) and plated on LB agar supplemented with maintenance antibiotics (kanamycin (45 µg / ml), streptomycin (90 µg / ml), gentamycin (18 µg / ml), and chloramphenicol (25 µg / mL). Cells were scraped and resuspended using Dulbecco's Phosphate Buffered Saline (DPBS) and 1ml to 10 ml of OD6001 re-plated onto LB agar plates of appropriate size supplemented with maintenance antibiotics kanamycin, as well as selection antibiotic (carbenicillin, 100 µg / ml). Survival of clones is contingent upon reversion of a premature ochre stop codon in the TEM-1 β-Lactamase gene on pOR-3 (i.e., statistical selection for mutagenic T7 DNA polymerase clones).

[0097] Small-scale pOR-3 fluctuation analysis. Strain T7Rep-ΔGp5was co- transformed with pOR-3 and promising T7 DNAP clones in either pCD0261 (identified from single site-saturation library reversion assays) or pCD0264 (identified from error- prone PCR library reversion assay). The transformation was plated on LB agar supplemented with maintenance antibiotics kanamycin (45 µg / ml), streptomycin (90µg / ml), gentamycin (18 µg / ml), and chloramphenicol (25 µg / ml). Single colonies were picked to grow out 8 replicates overnight in a 96-well plate (150μl volume, 37 °C, shaking speed of 800 a.u.). The cultures were diluted 105-fold and grown until they reach stationary phase. Volumes of 0.5 – 100μl were plated onto selective LB agar media supplemented with maintenance antibiotics and selection antibiotic carbenicillin (100 µg / ml) and the number of revertants (cfu’s) was quantified. Dilution plates were used to determine the absolute number of cells plated (no selection antibiotic) to calculate mutation rates. Luria-Delbrück fluctuation analysis was carried out by fitting the obtained data using the MSS maximum likelihood method in the Fluctuation AnaLysis CalculatOR (FalCOR) online tool.(51)

[0098] Large-scale pOR-3 fluctuation analysis. Strains T7Rep-1, T7Rep-2, T7Rep-3, T7Rep-4, T7Rep-5, and T7Rep-6, were transformed with pOR-3 and plated on LB agar media supplemented with maintenance antibiotics kanamycin (45 µg / ml), streptomycin (90 µg / ml), gentamycin (18 µg / ml), and chloramphenicol (25 µg / ml). Single colonies were picked to grow out 24 replicates overnight in a 96-well plate (150μl volume, 37 °C, shaking speed of 800 [a.u.]). The cultures were back-diluted 105-fold and grown until stationary phase. Volumes of 0.5 – 100μl (depending on the mutation rate) were plated onto selective LB agar media supplemented maintenance antibiotics and selection antibiotic carbenicillin (100 µg / ml) and the number of revertants (cfu’s) counted. Dilution plates were used to determine the absolute number of cells plated (no selection antibiotic) to calculate mutation rates. Luria-Delbrück fluctuation analysis was carried out by fitting the data using the MSS maximum likelihood method in the Fluctuation AnaLysis CalculatOR (FalCOR) tool.(2) To calculate the mutation rate per base pair (spb), the mutation rate was normalized by the copy number of the plasmid (see below) and the number of mutations that yield functional TEM-1 derivatives (7 / 9 possible mutations of the ochre codon; divide observed mutation rate by 7 / 3).

[0099] Genomic fluctuation analysis. T7Rep-2 and T7Rep-6 were grown out from glycerol stocks overnight in a 96-well plate (45 µg / ml kanamycin, 90 µg / ml streptomycin, 18 µg / ml gentamycin). The culture was back-diluted 104-fold and 12 replicates grown out to stationary phase in a 96-well plate format (150μl volume, 37 °C, shaking speed of 800 [a.u.].100 μl of culture were plated onto selective LB agar media supplemented with maintenance antibiotics kanamycin (45 µg / ml), streptomycin (90µg / ml), gentamycin (18 µg / ml), and selection antibiotic rifampicin (50 µg / ml) and the number of revertants (cfu’s) counted. Dilution plates were used to determine the absolute number of cells plated (no selection antibiotic) to calculate mutation rates. Luria-Delbrück fluctuation analysis was carried out by fitting the data using the MSS maximum likelihood method in the Fluctuation AnaLysis CalculatOR (FalCOR) tool.(51) To calculate the mutation rate per base pair (s.p.b.), the mutation rate was normalized by the number of mutations in the rpoB gene that that impart rifampicin resistance (77 known point mutations, divide observed mutation rate by 77 / 3).(52)

[0100] pOR-3 copy number determination. The copy number of pOR-3 in the different replisome strains (T7Rep-1 – T7Rep-7) was evaluated by qPCR using the Power SYBR Green PCR Master Mix (Thermo Fisher, Massachusetts, United States) on the QuantStudio 7 Flex Real-Time PCR System, 384-well (Thermo Fisher, Massachusetts, United States). Calibrator plasmid pSR-1 was constructed by cloning E. coli gene dxs into pUC19 using Gibson assembly to enable absolute quantitation. dxs and blaTEM-1 were amplified from 10-fold serial dilutions of pSR-1 and a calibration curve was constructed. For each replisome strain (T7Rep-1– T7Rep-7), 100 µL of culture was grown to mid log phase in 2YT media supplemented with kanamycin (45 µg / ml), streptomycin (90 µg / ml), gentamycin (18 µg / ml), and chloramphenicol (25 µg / ml). The cultures were boiled for ten minutes and spun down for 10 min (21.000 × g). The lysate was serial- diluted in 10-fold increments, which were used as template for qPCR. Amplification of blaTEM-1 of pOR-3 and dxs of the chromosomal DNA was used to determine the absolute copy number of T7 origin plasmid in the replisome strains by the ΔΔCT method, with dxs used to normalize the plasmid DNA amount to the genomic copy number of 1.(53)

[0101] β-Lactamase nomenclature. The numbering of residues in TEM-1 β- lactamase was done according to a standardized nomenclature for class A β- lactamases.(54)

[0102] β-Lactamase evolution experiments. Strain T7Rep-7transformed with pOR-4 was grown out overnight in 2YT media supplemented with maintenance antibiotics (45 µg / ml kanamycin, 90 µg / ml streptomycin, 18 µg / ml gentamycin, 25 µg / ml chloramphenicol.100 µl of liquid culture was plated on LB agar plates supplemented with both maintenance antibiotics and selection antibiotic (cefotaxime, ceftatzidime,cefepime, and aztreonam) at a concentration approximating the respective minimum inhibitory concentration of the four β-lactams (0.1 µg / ml, 0.4 µg / ml, 0.1 µg / ml, and 0.1 µg / ml, respectively).96 colonies for each selection were picked and grown out in 96 deep-well plates (1.2 ml culture volume). The concentration of the selection antibiotic was continuously increased 1000-fold in 5 increments. Cultures were back-diluted up to two times per day (once stationary phase is reached, 100-fold dilutions) and the antibiotic concentration increased once per day (for cefotaxime / cefepime / aztreonam: 0.1 µg / ml, 0.5 µg / ml, 2.5 µg / ml, 10 µg / ml, 50 µg / ml, 100 µg / ml; for ceftazidime: 0.4 µg / ml, 2.0 µg / ml, 10 µg / ml, 50 µg / ml, 250 µg / ml, 500 µg / ml).

[0103] g

[0104] g

[0105] g

[0106]

[0107] β-Lactamase evolution Sanger sequencing. At the end of each β-Lactamase evolution experiment the entire 96-well plate was spun down, each individual culture was resuspended in 1 ml of molecular biology grade water, and the entire TEM-1 β- Lactamase coding sequence, as well as the pAmpR promoter were amplified by PCR. The 96 samples were analyzed by Sanger sequencing of the unpurified PCR product.

[0108] β-Lactamase evolution nanopore amplicon sequencing. After each day of the β-Lactamase evolution experiment, 100µl of each culture of the 96-well plate were combined, 5ml of the mixed culture was spun down, the plasmid DNA was isolated, and the TEM-1 β-Lactamase gene was amplified by PCR using 1 ng of isolated DNA as template. The samples were analyzed by Primordium nanopore amplicon sequencing. Table 1. Some plasmids used in the work Name Specifier Origin ofSelectable replicationGene(s) of interestmarker pCD0280 pOR1 pBR322 PBAD::Gp4; PGlnS::Gp2.5 KanRpC0260 pOR2 pSC101 PpolI::Gp5 GenRpCD0261 pOR2 pSC101 PpolI::Gp5(Δ28) GenRpCD0262 pOR2 pSC101 PpolI::Gp5(Δ28, N520M) GenRpCD0263 pOR2 pSC101 PpolI::Gp5(Δ28, N520M, P560V) GenRpCD0264 pOR2 pSC101 PpolI::Gp5(Δ28, N520M, P560V, V443K) GenRpCD0265 pOR2 pSC101 PpolI::Gp5(Δ28, N520M, P560V, V443K, Y24F) GenRT7Rep-1BW25113 pCD0280, pCD0260, pCD0275 KanR, GenR, SmRT7Rep-2BW25113 pCD0280, pCD0261, pCD0275 KanR, GenR, SmRT7Rep-3BW25113 pCD0280, pCD0262, pCD0275 KanR, GenR, SmRT7Rep-4BW25113 pCD0280, pCD0263, pCD0275 KanR, GenR, SmRT7Rep-5BW25113 pCD0280, pCD0264, pCD0275 KanR, GenR, SmRT7Rep-6BW25113 pCD0280, pCD0265, pCD0275 KanR, GenR, SmRTable 3. Mutations rates on orthogonal T7 replicon. Mutation Upper 95% Lower 9Strain T7 ori5% PlasmidT7 DNAP MutationsRate CI [spb]CI [spb] Copy[spb]NumberT7Rep-1pOR-3 Wild type 2.8×10-98.0×10-107.0×10-101.5 T7Rep-2pOR-3 Δ28 3.3×10-89.4×10-98.5×10-912.5 T7Rep-3pOR-3 Δ28, N520M 6.8×10-71.5×10-81.4×10-85.4T7Rep-4pOR-3 Δ28, N520M, P560V 4.3×10-65.5×10-75.2×10-73.8 T7Rep-5pOR-3 Δ28, N520M, P560V, R443K 1.0×10-51.6×10-61.5×10-64.6 T7Rep-6pOR-3 Δ28, N520M, P560V, R443K, Y24F 1.1×10-51.5×10-61.5×10-65.0 T7Rep-7pOR-3 Δ28, N520M, P560V, R443K, Y24F, Q585R 1.7×10-92.6×10-102.5×10-104.1 Table 4. Genomic mutation rates. Strain T7 DNAP MutationsMutation RateUpper 95% CI Lower 95% CI [spb][spb] [spb] T7Rep-2Δ28 3.6×10-101.6×10-101.4×10-10T7Rep-7Δ28, N520M, P560V, R443K, Y24F, Q585R 3.9×10-101.7×10-101.6×10-10Table 5. Maximum growth rate and doubling time of replisome strains. Strain T7 ori Plasmid Max. rate [ΔOD / min]T7Rep-1pOR-3 2.34×10-229.7 T7Rep-2pOR-3 2.46×10-228.2 T7Rep-3pOR-3 2.11×10-232.9 T7Rep-4pOR-3 2.37×10-229.3 T7Rep-5pOR-3 2.17×10-232.0 T7Rep-6pOR-3 2.45×10-228.3 T7Rep-6pOR-3 2.36×10-229.4 Table 6. Some strain used in TEM-1 β-lactamase evolution experiments. MutationStrainT7 oriT7 DNAP MuMutations / (cell× Plasmid tationsRate[min] Copy Number [spb] kb×24h) Δ28, N520M, T7Rep-7pOR-4 P560V, R443K, 1.72×10-528.3 4.1 3.59 Y24F, Q585R

[0109] Some additional cited references 1. J. W. Drake, A constant rate of spontaneous mutation in DNA-based microbes. Proc. Natl. Acad. Sci.88, 7160–7164 (1991). 2. C. K. Biebricher, M. Eigen, What is a quasispecies? Quasispecies concept Implic. Virol., 1–31 (2006). 3. M. S. Packer, D. R. Liu, Methods for the directed evolution of proteins. Nat. Rev. Genet.16, 379–394 (2015).4. F. H. Arnold, Design by directed evolution. Acc. Chem. Res.31, 125–131 (1998). 5. A. J. Simon, S. d’Oelsnitz, A. D. Ellington, Synthetic evolution. Nat. Biotechnol. 37, 730–743 (2019). 6. W. P. C. Stemmer, Rapid evolution of a protein in vitro by DNA shuffling. Nature 370, 389–391 (1994). 7. D. M. Hillis, J. J. Bull, M. E. White, M. R. Badgett, I. J. Molineux, Experimental phylogenetics: Generation of a known phylogeny. Science 255, 589–592 (1992). 8. B. C. Dickinson, M. S. Packer, A. H. Badran, D. R. Liu, A system for the continuous directed evolution of proteases rapidly reveals drug-resistance mutations. Nat. Commun.5 (2014). 9. M. Camps, J. Naukkarinen, B. P. Johnson, L. A. Loeb, Targeted gene evolution in Escherichia coli using a highly error-prone DNA polymerase I. Proc. Natl. Acad. Sci. U. S. A.100, 9727–9732 (2003). 10. J. J. Bull, R. Sanjuan, C. O. Wilke, Theory of lethal mutagenesis for viruses. J. Virol.81, 2930–2939 (2007). 11. A. J. Herr, M. Ogawa, N. A. Lawrence, L. N. Williams, J. M. Eggington, M. Singh, R. A. Smith, B. D. Preston, Mutator suppression and escape from replication error–induced extinction in yeast. PLoS Genet.7, e1002282 (2011). 12. A. Ravikumar, G. A. Arzumanyan, M. K. A. Obadi, A. A. Javanpour, C. C. Liu, Scalable, continuous evolution of genes at mutation rates above genomic error thresholds. Cell 175, 1946-1957 (2018). 13. A. Ravikumar, A. Arrieta, C. C. Liu, An orthogonal DNA replication system in yeast. Nat. Chem. Biol.10, 175–177 (2014). 14. R. Tian, R. Zhao, H. Guo, K. Yan, C. Wang, C. Lu, X. Lv, J. Li, L. Liu, G. Du, Engineered bacterial orthogonal DNA replication system for continuous evolution. Nat. Chem. Biol.19, 1504–1512 (2023). 15. R. Tian, F. B. H. Rehm, D. Czernecki, Y. Gu, J. F. Zürcher, K. C. Liu, J. W. Chin, Establishing a synthetic orthogonal replication system enables accelerated evolution in E. coli. Science 383, 421–426 (2024). 16. G. Rix, E. J. Watkins-Dulaney, P. J. Almhjell, C. E. Boville, F. H. Arnold, C. C. Liu, Scalable continuous evolution for the generation of diverse enzyme variants encompassing promiscuous activities. Nat. Commun.11, 5644 (2020). 17. A. Wellner, C. McMahon, M. S. A. Gilman, J. R. Clements, S. Clark, K. M. Nguyen, M. H. Ho, V. J. Hu, J.-E. Shin, J. Feldman, Rapid generation of potent antibodies by autonomous hypermutation in yeast. Nat. Chem. Biol.17, 1057– 1064 (2021). 18. E. D. Jensen, F. Ambri, M. B. Bendtsen, A. A. Javanpour, C. C. Liu, M. K. Jensen, J. D. Keasling, Integrating continuous hypermutation with high‐ throughput screening for optimization of cis, cis‐muconic acid production in yeast. Microb. Biotechnol.14, 2617–2626 (2021). 19. J. E. Cronan, Escherichia coli as an experimental organism. eLS (2014).20. R. L. Williams, C. C. Liu, Accelerated evolution of chosen genes. Science 383, 372–373 (2024). 21. C. Weigel, H. Seitz, Bacteriophage replication modules. FEMS Microbiol. Rev. 30, 321–381 (2006). 22. C. C. Richardson, Bacteriophage T7: minimal requirements for the replication of a duplex DNA molecule. Cell 33, 315–317 (1983). 23. A. W. Kulczyk, C. C. Richardson, The replication system of bacteriophage T7. Enzym.39, 89–136 (2016). 24. Y. Gao, Y. Cui, T. Fox, S. Lin, H. Wang, N. de Val, Z. H. Zhou, W. Yang, Structures and operating principles of the replisome. Science 363, eaav7003 (2019). 25. S. J. Lee, C. C. Richardson, Choreography of bacteriophage T7 DNA replication. Curr. Opin. Chem. Biol.15, 580–586 (2011). 26. S. Tabor, C. C. Richardson, A bacteriophage T7 RNA polymerase / promoter system for controlled exclusive expression of specific genes. Proc. Natl. Acad. Sci.82, 1074–1078 (1985). 27. T. Baba, T. Ara, M. Hasegawa, Y. Takai, Y. Okumura, M. Baba, K. A. Datsenko, M. Tomita, B. L. Wanner, H. Mori, Construction of Escherichia coli K‐12 in‐frame, single‐gene knockout mutants: The Keio collection. Mol. Syst. Biol.2, 8–2006 (2006). 28. H. Zhang, S.-J. Lee, A. W. Kulczyk, B. Zhu, C. C. Richardson, Heterohexamer of 56-and 63-kDa gene 4 helicase-primase of bacteriophage T7 in DNA replication. J. Biol. Chem.287, 34273–34287 (2012). 29. W. Ma, A. Phan, R. Walsh, K. Ye, Building an orthogonal replication system for performing directed evolution in Escherichia coli: A strategic review and a summary of the initial steps in cloning bacteriophage T7 Gp4 primase / helicase. JEMI, 1–8 (2015). 30. K. Becker, A. Meyer, T. M. Roberts, S. Panke, Plasmid replication based on the T7 origin of replication requires a T7 RNAP variant and inactivation of ribonuclease H. Nucleic Acids Res.49, 8189–8198 (2021). 31. H. Saito, S. Tabor, F. Tamanoi, C. C. Richardson, Nucleotide sequence of the primary origin of bacteriophage T7 DNA replication: Relationship to adjacent genes and regulatory elements. Proc. Natl. Acad. Sci. U. S. A.77, 3917–3921 (1980). 32. W. T. McAllister, H.-L. Wu, Regulation of transcription of the late genes of bacteriophage T7. Proc. Natl. Acad. Sci.75, 804–808 (1978). 33. B. A. Moffatt, F. W. Studier, T7 lysozyme inhibits transcription by T7 RNA polymerase. Cell 49, 221–227 (1987). 34. X. Cheng, X. Zhang, J. W. Pflugrath, F. W. Studier, The structure of bacteriophage T7 lysozyme, a zinc amidase and an inhibitor of T7 RNA polymerase. Proc. Natl. Acad. Sci.91, 4034–4038 (1994). 35. X. Zhang, F. W. Studier, Multiple roles of T7 RNA polymerase and T7 lysozymeduring bacteriophage T7 infection. J. Mol. Biol.340, 707–730 (2004). 36. X. Zhang, F. W. Studier, Mechanism of inhibition of bacteriophage T7 RNA polymerase by T7 lysozyme. J. Mol. Biol.269, 10–27 (1997). 37. D. Jeruzalmi, T. A. Steitz, Structure of T7 RNA polymerase complexed to the transcriptional inhibitor T7 lysozyme. EMBO J.17, 4101–4113 (1998). 38. A. Kumar, S. S. Patel, Inhibition of T7 RNA polymerase: Transcription initiation and transition from initiation to elongation are inhibited by T7 lysozyme via a ternary complex with RNA polymerase and promoter DNA. Biochemistry 36, 13954–13962 (1997). 39. N. M. Stano, S. S. Patel, T7 lysozyme represses T7 RNA polymerase transcription by destabilizing the open complex during initiation. J. Biol. Chem. 279, 16136–16143 (2004). 40. B. C. Dickinson, M. S. Packer, A. H. Badran, D. R. Liu, A system for the continuous directed evolution of proteases rapidly reveals drug-resistance mutations. Nat. Commun.5, 5352 (2014). 41. S. Tabor, C. C. Richardson, DNA sequence analysis with a modified bacteriophage T7 DNA polymerase. Proc. Natl. Acad. Sci.84, 4767–4771 (1987). 42. S. Doublié, T. Ellenberger, The mechanism of action of T7 DNA polymerase. Curr. Opin. Struct. Biol.8, 704–712 (1998). 43. B. Zhu, Bacteriophage T7 DNA polymerase–sequenase. Front. Microbiol.5, 89476 (2014). 44. T. A. Kunkel, S. S. Patel, K. A. Johnson, Error-prone replication of repeated DNA sequences by T7 DNA polymerase in the absence of its processivity subunit. Proc. Natl. Acad. Sci.91, 6830–6834 (1994). 45. M. L. M. Salverda, J. A. G. M. De Visser, M. Barlow, Natural evolution of TEM-1 β-lactamase: Experimental reconstruction and clinical relevance. FEMS Microbiol. Rev.34, 1015–1036 (2010). 46. R. Fernandes, P. Amador, C. Oliveira, C. Prudêncio, Molecular characterization of ESBL-producing Enterobacteriaceae in northern Portugal. Sci. World J.2014 (2014). 47. T. Palzkill, Structural and mechanistic basis for extended-spectrum drug- resistance mutations in altering the specificity of TEM, CTX-M, and KPC β- lactamases. Front. Mol. Biosci.5, 16 (2018). 48. S. M. Miller, T. Wang, D. R. Liu, Phage-assisted continuous and non-continuous evolution. Nat. Protoc.15, 4101–4127 (2020). 49. K. M. Esvelt, J. C. Carlson, D. R. Liu, A system for the continuous directed evolution of biomolecules. Nature 472, 499–503 (2011). 50. T. Baba, T. Ara, M. Hasegawa, Y. Takai, Y. Okumura, M. Baba, K. A. Datsenko, M. Tomita, B. L. Wanner, H. Mori, Construction of Escherichia coli K‐12 in‐frame, single‐gene knockout mutants: the Keio collection. Mol. Syst. Biol.2, 8–2006 (2006).51. B. M. Hall, C.-X. Ma, P. Liang, K. K. Singh, Fluctuation AnaLysis CalculatOR: a web tool for the determination of mutation rate using Luria–Delbrück fluctuation analysis. Bioinformatics 25, 1564–1565 (2009). 52. A. H. Badran, D. R. Liu, Development of potent in vivo mutagenesis plasmids with broad mutational spectra. Nat. Commun.6, 1–10 (2015). 53. C. Lee, J. Kim, S. G. Shin, S. Hwang, Absolute and relative QPCR quantification of plasmid copy number in Escherichia coli. J. Biotechnol.123, 273–280 (2006). 54. R. P. Ambler, A. F. Coulson, J.-M. Frère, J.-M. Ghuysen, B. Joris, M. Forsman, R. C. Levesque, G. Tiraby, S. G. Waley, A standard numbering scheme for the class A beta-lactamases. Biochem. J.276, 269 (1991). ***

[0110] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be readily apparent to one of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.

[0111] All publications, databases, sequence accession numbers, patents, and patent applications cited in this specification are herein incorporated by reference as if each was specifically and individually indicated to be incorporated by reference.

Claims

WHAT IS CLAIMED IS:

1. An engineered T7 replisome, comprising one or more vectors that express (1) a mutant T7 DNA polymerase (Gp5) containing a substitution at residue N520 and a deletion of residues K118-R145in the exonuclease domain, (2) a T7 single- stranded DNA binding protein (Gp2.5) or variant thereof, (3) a T7 RNA polymerase (Gp1) or variant thereof, and (4) a T7 primase-helicase (Gp4) or variant thereof; wherein amino acid numbering of T7 DNA polymerase is based on wildtype T7 DNA polymerase (Gp5) sequence under UniProt ID P00581 (Accession No. CAA24412.1; SEQ ID NO:1).

2. The engineered T7 replisome of claim 1, wherein residue N520 is substituted with a small hydrophobic residue M, I, L, or C, or a small polar residue T, N, D, or E.

3. The engineered T7 replisome of claim 1, wherein the one or more vectors each comprises an origin of replication that replicates in E. coli.

4. The engineered T7 replisome of claim 1, wherein the one or more vectors are derived from pBR322, pSC101, or CloDF13.

5. The engineered T7 replisome of claim 1, wherein the mutant T7 DNA polymerase comprises N520M substitution.

6. The engineered T7 replisome of claim 1, wherein the mutant T7 DNA polymerase further comprises P560V, V443K, Y24F, and Q585R substitutions.

7. The engineered T7 replisome of claim 1, wherein the mutant T7 DNA polymerase comprises the amino acid sequence as set forth in any one of SEQ ID NOs:3-7.

8. The engineered T7 replisome of claim 1, wherein the T7 primase- helicase is expressed from a Gp4 polynucleotide sequence that does not express Gp4B.

9. The engineered T7 replisome of claim 1, wherein the T7 RNA polymerase (Gp1) is fused to a T7 lysozyme (Gp3.5) mutant that is deficient in hydrolase activity.

10. The engineered T7 replisome of claim 9, wherein the T7 lysozyme mutant comprises a C131S substitution, and wherein amino acid numbering of T7 lysozyme is based on wildtype T7 lysozyme sequence under UniProt ID P00806 (SEQ ID NO:8).

11. An E. coli orthogonal DNA replication system, comprising an E. coli cell into which is introduced the engineered T7 replisome of claim 1.

12. The orthogonal DNA replication system of claim 11, wherein residue N520 is substituted with a small hydrophobic residue M, I, L, or C, or a small polar residue T, N, D, or E.

13. The orthogonal DNA replication system of claim 11, wherein the one or more vectors are derived from pBR322, pSC101, or CloDF13.

14. The orthogonal DNA replication system of claim 11, wherein the mutant T7 DNA polymerase comprises a N520M substitution.

15. The orthogonal DNA replication system of claim 11, wherein the mutant T7 DNA polymerase further comprises P560V, V443K, Y24F, and Q585R substitutions.

16. The orthogonal DNA replication system of claim 11, wherein the mutant T7 DNA polymerase comprises the amino acid sequence as set forth in any one of SEQ ID NOs:3-7.

17. The orthogonal DNA replication system of claim 11, wherein the T7 primase-helicase is expressed from a Gp4 polynucleotide sequence that does not express Gp4B.

18. The orthogonal DNA replication system of claim 11, wherein the T7 RNA polymerase (Gp1) is fused to a T7 lysozyme (Gp3.5) mutant that is deficient in hydrolase activity.

19. The orthogonal DNA replication system of claim 18, wherein the T7 lysozyme mutant comprises a C131S substitution, and wherein amino acid numbering of T7 lysozyme is based on wildtype T7 lysozyme sequence under UniProt ID P00806 (Accession No. Q38567; SEQ ID NO:8).

20. A method for evolving a target polypeptide, comprising (1) transforming the E. coli orthogonal replication system of claim 11 with a vector harboring a polynucleotide that encodes the target polypeptide, wherein the vector contains a T7 origin of replication, (2) culturing the transformed E. coli replication system under appropriate conditions, and (3) selecting one or more E. coli cells from the cultured E. coli replication system with an enhanced property or a desired phenotype; thereby evolving the target polypeptide.

21. The method of claim 20, wherein the T7 origin of replication is φOR.

22. The method of claim 20, wherein the target polypeptide is an enzyme, a hormone, a therapeutic protein, or an antibody molecule.

23. The method of claim 20, further comprising isolating the vector from the cultured E. coli cell, expressing the evolved polynucleotide, and purifying the evolved target polypeptide expressed from the evolved polynucleotide.

24. An engineered T7 DNA polymerase with enhanced mutation rate relative to the wildtype T7 polymerase, comprising a deletion of residues K118-R145in the exonuclease domain, and a substitution at residue N520; wherein the amino acid numbering is based on wildtype T7 DNA polymerase (Gp5) sequence under UniProt ID P00581 (Accession No. CAA24412.1; SEQ ID NO:1).

25. The engineered T7 DNA polymerase of claim 24, wherein residue N520 is substituted with a hydrophobic residue or a polar residue.

26. The engineered T7 DNA polymerase of claim 25, wherein the hydrophobic residue is M, I, L, or C, and the polar residue is T, N, D, or E.

27. The engineered T7 DNA polymerase of claim 24, further comprising one or more mutations at residues P560, V443, Y24, and Q585.

28. The engineered T7 DNA polymerase of claim 27, wherein residue P560 is substituted with residue V, M, I, L, G, A, or K; residue V443 is substituted with residue K, S, T, N, Q, R, H, or C; residue Y24 is substituted with residue F or W; and residue Q585 is substituted with residue R, T, R. E, or C.

29. The engineered T7 DNA polymerase of claim 28, wherein the one or more mutations at residues P560, V443, Y24, and Q585 are selected from the group consisting of P560V, V443K, Y24F, and Q585R.

30. The engineered T7 DNA polymerase of claim 28, comprising N520M, P560V, V443K, Y24F, and Q585R substitutions.

31. The engineered T7 DNA polymerase of claim 24, comprising the amino acid sequence as set forth in any one of SEQ ID NOs:3-7.

32. A polynucleotide encoding the engineered T7 DNA polymerase of claim 24.

Citation Information

Patent Citations

  • Use of DNA polymerases as exoribonucleases

    EP1916312A1