Compositions and methods for directed evolution

The LySE system addresses scaling and control issues in directed evolution by using error-prone DNA polymerases and phage vectors to efficiently evolve large gene constructs and metabolic pathways with controlled mutagenesis and selection cycles, achieving high mutation rates and reduced off-target mutations.

WO2026111651A1PCT designated stage Publication Date: 2026-05-28NATIONAL UNIVERSITY OF SINGAPORE

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NATIONAL UNIVERSITY OF SINGAPORE
Filing Date
2025-11-19
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing directed evolution platforms face challenges in scaling, time consumption, and limited exploration of sequence space due to trade-offs between improving gene construct function and phage fitness, with conventional methods being cumbersome and continuous methods lacking control over mutagenesis.

Method used

A phage-assisted directed evolution system (LySE) that uses a propagation-defective phage vector to introduce an error-prone DNA polymerase in a host cell for mutagenesis, replicating and packaging gene constructs into phage particles, followed by transduction into a selection host for controlled evolution cycles.

Benefits of technology

Enables accelerated evolution of large gene constructs and metabolic pathways with reduced off-target mutations, achieving high mutation rates and selective pressure on the gene of interest, overcoming limitations of conventional and continuous evolution systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050734_28052026_PF_FP_ABST
    Figure SG2025050734_28052026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates, in general terms, to directed evolution, and more specifically to compositions and methods for phage-assisted directed evolution in microbes. In one embodiment, there is provided a method for directed molecular evolution, the method comprising: a) introducing a propagation-defective phage vector into a first host cell (mutagenic host) competent to propagate the phage vector, wherein the first host cell comprises a gene construct of interest (GOI) to be evolved; and wherein the phage vector allows for (i) expression of an error-prone DNA polymerase in the first host cell for mutagenesis, and (ii) replication and packaging of the GOI into infectious phage particles; b) incubating the first host cell under conditions for replication and mutagenesis of the GOI and release of phage particles comprising mutated GOIs; c) infecting a second host cell (selection host) with phage particles from step b); and d) selecting for a desirable function of the GOI in the second host cell.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] COMPOSITIONS AND METHODS FOR DIRECTED EVOLUTION Technical field

[0002] The present invention relates, in general terms, to directed evolution, and more specifically to compositions and methods for phage-assisted directed evolution in microbes.

[0003] Background

[0004] Directed evolution is an indispensable bioengineering tool that enables the discovery of novel protein variants. Classical directed evolution platforms (such as phage or yeast display) are based on discrete cycles of mutagenesis, selection, and amplification. These approaches provide control over the evolution process, allowing sequence analysis and adjustment of selection parameters between cycles to prevent the accumulation of off-target mutations. However, they can be time-consuming to iterate, difficult to scale, and usually do not fully explore the sequence space of a gene construct ofinterest (GOI).

[0005] Phage-assisted continuous evolution (PACE) and related continuous evolution systems offer a different approach by selectively mutagenising target GOIs inside cells. In PACE, the GOI is fused to a phage gene that is essential for phage survival and replication, and selection is based on the ability of gene variants to improve phage propagation. PACE enables continuous cycling between mutagenesis and selection with minimal external input, thus allowing accelerated evolution timelines and high throughput workflows. However, because the GOI must be incorporated into the phage genome, its size is limited by phage’s packaging capacity. Furthermore, as evolution proceeds, there may be trade-offs between improving the GOI’s desired function and maintaining the phage’s own fitness and replication capacity. These trade-offs can limit both the speed and the extent of evolutionary progress.

[0006] There is thus a need for an improved directed evolution platform that combines the controllability of conventional display approaches with the speed and throughput advantages of continuous evolution platforms. It would be desirable to overcome or alleviate at least one of the above-described problems, or at least to provide a useful alternative. Summary

[0007] Disclosed herein is a method for directed molecular evolution, the method comprising: a) introducing a propagation-defective phage vector into a first host cell (mutagenic host) competent to propagate the phage vector, wherein the first host cell comprises a gene construct of interest (GOI) to be evolved; and wherein the phage vector allows for (i) expression of an error-prone DNA polymerase in the first host cell for mutagenesis, and (ii) replication and packaging of the GOI into infectious phage particles;

[0008] b) incubating the first host cell under conditions for replication and mutagenesis of the GOI and release of phage particles comprising mutated GOls;

[0009] c) infecting a second host cell (selection host) with phage particles from step b); and d) selecting for a desirable function of the GOI in the second host cell.

[0010] Disclosed herein is an error-prone DNA polymerase that is distinguished from a wild-type DNA polymerase by i) at least one amino acid substitution at a position corresponding to position 479, 523 or 560 of SEQ ID NO: 1, and / or ii) a conjugated nucleoside deaminase.

[0011] Disclosed herein is an error-prone DNA polymerase comprising an amino acid sequence having at least 70% sequence identity to an amino acid sequence in SEQ ID NO: 1, wherein the error-prone polymerase comprises at least one of the following:

[0012] a) a N at a position corresponding to position 479 of SEQ ID NO: 1;

[0013] b) a R at a position corresponding to position 523 of SEQ ID NO: 1; and

[0014] c) a H at a position corresponding to position 560 of SEQ ID NO: 1; and / or

[0015] d) a conjugated nucleoside deaminase.

[0016] Disclosed herein is a polynucleotide comprising a nucleic acid sequence encoding an error-prone DNA polymerase as defined herein.

[0017] Disclosed herein is an expression vector comprising a polynucleotide as defined herein.

[0018] Disclosed herein is a host cell comprising a polynucleotide or an expression vector as defined herein.

[0019] Disclosed herein is a method of producing an error-prone DNA polymerase as defined herein, comprising culturing a host cell as defined herein under conditions suitable for expressing the error-prone DNA polymerase.

[0020] Disclosed herein is a kit for performing a method as defined herein, the kit comprising an error-prone DNA polymerase or an expression vector as defined herein.

[0021] Brief description of the drawings

[0022] Embodiments of the present invention will now be described, by way of non-limiting example, with reference to the drawings in which:

[0023] Figure 1 is a schematic showing the evolution of gene constructs through cycles of lysis and transduction using LySE. (a) Controlled replication and phage packaging, (b) The LySE cycle, (c) Temporal comparison of LySE versus conventional directed evolution methods, (d) LySE is a hybrid system between discrete and continuous evolution. Transduction replaces transformation to enable seamless transition between mutagenesis and selection, a characteristic of continuous evolution. Cheater mutations are a problem in PACE and orthogonal replication systems because the GOI replicates with the host. LySE creates discrete checkpoints to prevent cheater accumulation by refreshing the host in each cycle.

[0024] Figure 2 shows the engineering and characterisation of hypermutagenic T7 DNA polymerase variants, (a) Mutational frequencies of engineered T7 DNAP variants measured by a fluctuation assay using a phagemid-encoded chloramphenicol resistance gene containing an internal stop codon. Wild-type T7 DNAP (wt), mutant T7 DNAPs (vl-v4), and wild-type T7 DNAP-deaminase fusions (v5-v6). Data represent mean ± SD from three independent experiments, (b) Snapshots from molecular dynamics simulations of T7 DNAP (PDB: 1T7P): wild-type (top) and T523R variant mutated in silico (bottom) with A-G nucleotide mismatch. Hydrogens not displayed. Distances measured in Angstrom, (c) Mutational frequencies of advanced T7 DNAP variants determined through a lacZ inactivation assay using a phagemid-encoded lacZ reporter. Mutations were quantified by screening for white or light blue colonies indicative of loss-of-function mutations. Mutant T7 DNAPs (v2-v4), mutant T7 DNAP-deaminase fusions (v7-v8). Data represent mean ± SD from three independent experiments, (d) AlphaFold3-predicted structure of variant v8, highlighting key engineered modifications: vl (S399T) - thumb domain relaxation affecting template minor groove interactions; v2 (T523R) - fingers domain modification proximal to active site; v4 (D5A, E7A, Y64C, F120L) - exonuclease domain inactivation; v8 - fusion with TadA-8e adenosine deaminase enabling concurrent DNA deamination during replication.

[0025] Figure 3 shows continuous lytic cycling of phagemids by multiplicity tuning, (a) Change in optical density of E. coli cultures expressing T7 DNAP wild type (wt) and variants (v2.4-v8) when infected with T7ADNAP over 3 hours at varying multiplicities of infection (MOI: 0.01-10). Error bars represent mean ± SD (n = 3). (b) Lysis kinetics of E. coli cultures expressing wild-type (wt, dashed lines) or hypermutagenic (v8, solid lines) T7 DNAP at different MOIs (0.01-10). (c) Schematic representation of proposed multiplicity tuning mechanism: After infection (step 1), wild type T7 DNAP enables efficient phage replication at low MOI (step Ila), while error-prone variant v8 requires higher MOI due to increased phage inactivation during replication (step lib), (d) Quantification of phagcmid packaging efficiency across T7 DNAP variants, demonstrating maintained library diversity despite reduced transduction rates in error-prone variants. Data shown as mean ± SD (n = 3). (e) Representative 20-hour time course of a complete LySE cycle, showing distinct phases of bacterial growth and phage-mediated lysis initiated by simple mixing of phage lysates and cell cultures.

[0026] Figure 4 shows engineering of a hybrid deaminase with a broad mutagenesis spectrum, (a) Biochemical mechanism of nucleobase modifications by deaminases. Fusion of error-prone T7 DNAP with the adenosine deaminase TadA-8e (v8) catalyses adenosine-to-inosine deamination resulting in A: T— > G: C transitions. Installation of the dual adenine-cytosine deaminase TadDE (v9) enables both adenosine and cytidine deamination, facilitating A: T— > G: C and C: G— > T: A transitions, (b) Substitution frequencies and ratios of A: T— > G: C and C: G— > T: A transitions for v8 and v9 variants derived from NGS data (Illumina sequencing, mean coverage >14,000x per base pair), (c) Distribution of substitution frequencies across a 39-kb BAC-phagemid (mean coverage = 14,098x per base pair). Mutational spectra were calculated for every 10,000 bp. (d) Mutational spectra of LySE v9.

[0027] Figure 5 shows phagemid-targeted evolution of genes and gene clusters, (a) Comparative workflows of LySE versus adaptive laboratory evolution (ALE). LySE confines mutations to the phagemid through phage-mediated cycling, while ALE allows accumulation of genome-wide mutations during serial passaging, (b) Tigecycline resistance profiles of an evolved pool after five generations of LySE (E5 LySE), compared to the culture before evolution (El LySE). A representative single clone from E5 (E5_2 LySE) is also shown, (c) The same phagemid and tetA evolved by ALE. Resistance profile of pool after five passages (E5 ALE). The acquired resistance was almost entirely lost when transferring the phagemid to a fresh host cell (E5T ALE), indicating genome-dependent rather than phagemid-encoded adaptation, (d) Structure of TetA and identified mutations in the evolved pool by Sanger sequencing of 32 clones, (e) Metabolic pathway for ethylene glycol (EG) assimilation in E. coli. (f) Linearised representation of the phagemid with the EG assimilation pathway for evolution. Identified mutations from 8 clones are labelled in red, with corresponding clone number (EGA1-8). (g) Comparative workflows for semi-relaxing evolution of EG assimilation pathway using LySE versus ALE. Both approaches employed progressive selection from 1 g / L glucose + 8 g / L EG (El) to 0 g / L glucose + 12 g / L EG (E5). The LySE protocol included cell recovery in ampicillin and kanamycin to maintain phagemid and accessory plasmid, followed by washing and selection in minimal medium. Similarly, the ALE protocol comprised initial growth, washing, minimal medium selection, antibiotic recovery to maintain phagemid and accessory plasmid, and final washing before the next selection cycle. Growth curves for each selection cycle are shown, (h) Growth curves of eight isolated evolved phagemids (EGA1-8), and wild type phagemid (WT) in M9 media with 10 g / L EG and no glucose. For panel (g) and (h), data are the mean ± standard deviation from three (n = 3) independent biological replicates

[0028] Figure 6 shows T7 phagemid packaging during lysis, (a) Schematic of the T7 lytic cycle, (b) Biocontainment of phage T7ADNAP. (c) Plaque assay of T7ADNAP infecting E. coli cells with or without AP carrying T7 DNAP. For cells expressing T7 DNAP, the bacteria lawn was completed cleared upon infection with T7ADNAP of 5 x 104PFU or higher. No plaques were observed on cells not expressing T7 DNAP. (d) Lysis kinetics of T7ADNAP infecting E. coli cells with or without AP carrying T7 DNAP. MOI = 0.1. e. Transduction of a phagemid into new host cells by phage T7ADNAP. E. coli cells containing the AP and phagemid were lysed by addition of phage T7ADNAP. The lysate was washed with chloroform, transduced 1:100 into fresh E. coli cells and directly spotted with dilution on LB agar with 50 pg / mL kanamycin to select for cells that received the phagemid. Three independent replicates are shown.

[0029] Figure 7 shows molecular dynamics simulations of T7 DNAP crystal structure, (a) Root Mean Square Deviation (RMSD) values of dGTP incorrectly paired with adenine in wild type T7 DNAP (blue) and T523R mutant (orange) in triplicate. dGTP shows increased movement with R523 compared to T523 throughout the simulation, (b) RMSD values of dGTP incorrectly paired with guanine in wild type T7 DNAP (blue) and T523R mutant (orange) in triplicate. dGTP shows slightly increased movement with R523 compared to T523, but this is less pronounced than with other mispairings, (c) RMSD values of dGTP incorrectly paired with thymine in wild type T7 DNAP (blue) and T523R mutant (orange) in triplicate. dGTP shows increased movement with R523 compared to T523 throughout the simulation, (d) RMSD values of dGTP correctly paired with cytosine in wild type T7 DNAP (blue) and T523R mutant (orange) in triplicate. Correctly paired dGTP shows notably low RMSD values in the T523R mutant, while RMSD increases in wild type compared to other conditions. This contrast correlates with a conformational change in the protein when matching nucleotides align in the mutated structure. R523 may interact strongly only with mismatched nucleotides, potentially contributing to increased error rates, (e) Snapshot of wild type T7 DNAP with threonine in position 523 (grey) with correct pairing between incoming dGTP (green) and template cytosine (cyan). The a-helix with T523 maintains the confirmation observed for base mispairings in wild type and T523R T7 DNAP. (f) Snapshot of T7 DNAP with T523R mutation (grey) and correct base pairing between incoming dGTP (green) and template cytosine (cyan). Correct base pairing prevents R523 from positioning between incoming and template nucleotides, resulting in an entirely different helix conformation.

[0030] Figure 8 shows transduction efficiency and growth rates for T7 DNAP variants, (a) Quantification of phagemid packaging efficiency across wild type (wt) and engineered (vl-8) T7 DNAP variants by selection with kanamycin after transduction. Data shown as mean ± SD (n = 3). (b) Growth kinetics of E. coli strains harboring no T7 DNAP (Empty), wild type T7 DNAP (wt), or the hypermutagenic T7 DNAP variant v8.

[0031] Figure 9 shows LySE replication, packaging and transduction of a 39 kb BAC-phagemid. (a) Overall base substitution frequencies for T7 DNAP variants v8 and v9 derived from Illumina next-generation sequencing data (mean coverage > 14,000x per base pair), (b) Read length distribution from nanopore sequencing of transductants after one generation of LySE with 39 kb BAC-phagcmid. E. coli cells containing accessory plasmid (WT T7 DNAP) and BAC-phagemid were lysed by addition of phage T7ADNAP. The lysate was washed with chloroform, transduced 1: 100 into fresh E. coli cells and recovered overnight in LB medium with 50 pg / mL kanamycin to select for the BAC-phagemid. Phagemids were extracted by miniprep and sequenced by nanopore.

[0032] Figure 10 shows tetA mutants obtained from LySE evolution for tigecycline resistance, (a) Screening of 32 clones from LySE E5 evolution of tetA gene for tigecycline resistance, (b) Sanger sequencing of the tetA gene in the 32 evolved clones. Convergence of sequences in the promoter region (pCAT) as well as in the coding sequence (V145A, G283S). (c) Normalized tetA expression levels in E5 LySE compared to El pools quantified by RT-qPCR (n = 3; mean ± SD).

[0033] Figure 11 shows LySE evolution of ethylene glycol (EG) assimilation pathway and selection with only EG. Shown are the growth rates of cells containing EG assimilation pathway phagemid after each round of LySE evolution in E. colt using a semi-relaxed selection protocol. Control is E. coli with no phagemid. Cultures were grown in M9 minimal medium with 10 g / L EG. Data are the mean ± standard deviation from three (n = 3) independent biological replicates. The best replicate from each round was picked for subsequent rounds of LySE.

[0034] Figure 12 shows endpoint biomass (t = 48 hours) of evolved ethylene glycol (EG) assimilation pathway phagemid clones after five generations of LySE. Cultures were grown in M9 minimal medium with 10 g / L EG. Five clones exhibited accelerated growth on EG compared to the wild type (Starting from best: EGA5, EGA3, EGA7, EGA8, EGA4). Data are the mean ± standard deviation from three (n = 3) independent biological replicates. Significance was tested via two-tailed unpaired two-sample t-test. * p < 0.01; ** p < 0.001; *** p < 0.0001.

[0035] Figure 13 shows a map of an exemplary T7 phagemid.

[0036] Figure 14 shows a map of the T7 phagemid (pS J78) used to evolve the tetA gene.

[0037] Figure 15 shows a map of the T7 phagemid (pAN29) used to evolve the ethylene glycol (EG) pathway.

[0038] Detailed description

[0039] The inventors have developed a phage-assisted directed evolution system that allows the evolution of large gene constructs and metabolic pathways in microbes, and a natural cycling between mutagenesis and selection. Also termed LySE (LYtic Selection and Evolution) in this disclosure, the system leverages the lytic cycle of a lytic bacteriophage to replicate a gene construct of interest (GO I) on a phagemid, and an error-prone phage DNA polymerase to introduce mutations during replication. The lytic phage cycle naturally packages the phagemid into phage particles which are then transduced into fresh host cells for selection. Desired functions are selected for by coupling gene expression to host fitness. The cycle is repeated by initiating a new round of phage infection. LySE uniquely combines continuous evolution with discrete evolution cycles, enabling accelerated evolution of large gene constructs while maintaining stringent control over mutational trajectories to reduce off-target mutations.

[0040] In an exemplary, non-limiting embodiment, a LySE system based on the T7 bacteriophage was developed. The T7 phage was rendered replication-deficient by deleting the T7 DNA polymerase from the phage genome. Complementation of a host E. coli cell with an error-prone T7 DNA polymerase (carried on a plasmid) allows phage replication and mutagenesis. Expression of the error-prone polymerase is placed under inducible control to keep basal mutation levels low. A T7 phagemid was also engineered to shuttle gene constructs of interest (GOI) between the phage and the E. coli host.

[0041] In the first stage of the directed evolution process, E. coli hosts carrying a T7 phagemid with the GOT are infected with the engineered T7 phage, which induces production of the error-prone DNA polymerase. The polymerase replicates and mutagenises both the T7 genome and the phagemid, generating a pool of phages at the end of the lytic cycle containing either mutated T7 genomes or GOI variants. The phage particles are used to infect a new batch of host cells in the second stage. Because error-prone replication naturally gives rise to a population of phages with reduced multiplicity of infection (MOI), this second round of infection mostly leads to transduction of the host cells rather than lysis, which generates a cell library of GOI variants that can be used for selection and screening. Another cycle of mutagenesis and selection is initiated by infecting the host cells with T7 phage at a high multiplicity of infection enough to induce lysis.

[0042] By cycling between high MOI phage exposure (infection and lysis) and low MOI phage exposure (transduction), the LySE method allows closed-loop cycling of mutagenesis and selection. Furthermore, because mutagenesis and evolution of the GOI is decoupled from phage and host cell fitness, much higher rates of mutagenesis can be achieved, exceeding the genomic error threshold that is tolerable for phage viability.

[0043] As the phage DNA polymerase (DNAP) acts orthogonally to the host DNA polymerase, the phage DNAP can be engineered to achieve extremely high error rates without affecting host viability. Furthermore, the implementation of controlled error-prone replication and multiplicity tuning in LySE redirects pressure away from maintaining phage genome fidelity, enabling the use of mutation rates that exceed error thresholds for phage viability. The inventors engineered hypermutagenic T7 DN A polymerases that can achieve a mutation rate exceeding 10-4substitutions per base, which is ~106times higher than the genomic mutation rate of E. coli, and also exceeds the T7 phage error threshold. The mutational spectrum was further optimised by fusing a deaminase (such as adenine or cytosine deaminase) to the DNA polymerase to introduce random base changes during replication.

[0044] Using the LySE method and the engineered error-prone DNA polymerase, a 25-fold increase in tigecycline resistance could be achieved in an E. coli host in just 5 cycles. LySE also enables selection of slow-manifesting metabolic functions by coupling large gene cluster expression to host fitness. This was demonstrated by evolving a pathway that allows E. coli to utilise ethylene glycol (the monomer of PET) as its sole carbon source.

[0045] Unlike the PACE system where the GO1 is inserted into the phage genome (which fundamentally limits the size of the GOI that can be evolved), the LySE approach employs a phagemid to carry' the GOT, thus leveraging the full capacity of the phage capsid. For example, with a T7 phagemid and T7 phage (packaging capacity -40 kb), genetic constructs of up to -39 kb can be evolved. The inventors evolved a 29 kb gene cluster using T7 LySE, which exceeds the size limit of PACE and many other continuous evolution systems.

[0046] In directed evolution systems where mutagenesis and selection occur inside the same host cell, off-target mutations (also known as hitchhiker mutations) and cheater mutations can accumulate. Off-target mutations are genetic changes that arise outside the GOT (e.g., in the host genome, on accessory plasmids, or within the selection circuitry) but are nonetheless carried forward during the evolution process. Accumulation of such mutations over successive host generations can alter host physiology or selection pathways, thus complicating phenotypic attribution. Certain off-target mutations may bypass selection pressures altogether, for example by falsely activating a reporter in selection assays or by conferring a growth or replication advantage independent of the GOI or selection mechanism. Such cheater mutations can outcompete desirable GOT variants and become fixed in the cell population, ultimately ending the evolution process prematurely. Directed evolution systems that rely on lysogenic phages (e.g., the Pl phage) can introduce additional complications, as these phages can initiate genetic recombination events, leading to the inadvertent transfer of host genomic sequences into phagemids or phage particles. This can make it harder to resolve genetic changes that have occurred in the GOI. The LySE system overcomes these limitations by employing a lytic phage (thus reducing the risk of recombination-driven artifacts) and by using different hosts for mutagenesis and selection. Phage-induced cell lysis completely eliminates the host culture after each selection cycle, effectively removing all off-target mutations and ensuring that only the phagemid is carried forward for selection. The use of a separate selection host ensures that the phenotype evaluated during selection can be attributed to changes in the GOI rather than hitchhiker mutations that may have accumulated during replication. The LySE system thus provides a discrete checkpoint for maintaining the mutation trajectory not unlike in classical directed evolution systems, with the added flexibility of control over the rate of evolution and the size of the evolved construct.

[0047] Another strength of LySE lies in its modularity and compatibility with other mutagens. For example, more targeted mutagenic agents (e.g., CRISPR-guided DNAPs) can be used in the LySE platform to further increase error rates in specific regions of the phagemid. Other global mutagens such as the mutagenesis plasmid from PACE, mutator host strains like XL-1 red, or even chemical mutagens and UV irradiation, may also be used in the LySE method. If further genetic diversity is required to initiate a LySE evolution campaign, in vitro diversification methods such as error-prone PCR and DNA shuffling can also be used to generate a starting variant library.

[0048] Furthermore, LySE allows the use of multiple hosts and the exploration of multiple selection schemes in a single evolution campaign without changing the phage vector. In growth-coupled metabolic engineering, this facilitates rapid switching between different auxotrophic biosensor strains, where each may offer a different level of selection stringency suited to specific target genes or clusters. Although LySE may have a longer cycle time than PACE (depending on the selection process), LySE uniquely accommodates slow-manifesting phenotypes.

[0049] The LySE process is also easy to implement, with the first stage involving only the mixing of phage lysates and cell cultures, and is readily scalable and adaptable for high-throughput formats with automated liquid handling. Furthermore, due to the short T7 lytic cycle, at least one round of replication, transduction, recovery and selection can be completed within 24 hours. This contrasts with display-based directed evolution systems where a single round of mutagenesis and panning may take a week or longer. Accordingly, this disclosure provides compositions and methods for directed molecular evolution in bacteria based on the lytic phage cycle and error-prone phage DNA polymerases.

[0050] Disclosed herein is a method for directed molecular evolution, the method comprising: a) introducing a propagation-defective phage vector into a first host cell (mutagenic host) competent to propagate the phage vector,

[0051] wherein the first host cell comprises a gene construct of interest (GOI) to be evolved; and wherein the phage vector allows for (i) expression of an error-prone DNA polymerase in the first host cell for mutagenesis, and (ii) replication and packaging of the GOI into infectious phage particles;

[0052] b) incubating the first host cell under conditions for replication and mutagenesis of the GOI and release of phage particles comprising mutated GOIs;

[0053] c) infecting a second host cell (selection host) with phage particles from step b); and d) selecting for a desirable function of the GOI in the second host cell.

[0054] Also disclosed herein is an error-prone DNA polymerase that is distinguished from a wildtype DNA polymerase by i) at least one amino acid substitution at a position corresponding to position 479, 523 or 560 of SEQ ID NO: 1, and / or ii) a conjugated nucleoside deaminase.

[0055] Also disclosed herein is a kit for performing a method as defined herein, the kit comprising an error-prone DNA polymerase as defined herein, or an expression vector encoding the error-prone DNA polymerase.

[0056] The compositions and methods herein can be adapted for use in a variety of host cells because the critical components (such as the gene construct of interest and the error-prone polymerase) can be encoded in transformable vectors, and mutagenesis is not contingent on any specific genetic defect in the host cell. The compositions and methods can introduce a broad spectrum of mutations to, and be used to evolve, any target gene, prokaryotic or eukaryotic, with a screenable phenotype.

[0057] General definitions The terms “nucleic acid” and “polynucleotide” are used interchangeably herein to refer to a polymer of nucleotides, which can be mRNA, RNA, cRNA, cDNA or DNA. The term typically refers to polymeric form of nucleotides of at least 10 bases in length, either ribonucleotides or deoxynucleotides or a modified form of either type of nucleotide. The term includes single and double stranded forms of DNA and RNA.

[0058] The terms “polypeptide”, “peptide” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Typically, a polypeptide will be at least three amino acids long. These terms do not exclude modifications, for example, glycosylations, acetylations, phosphorylations, and the like. Soluble forms of the subject proteinaceous molecules are particularly useful.

[0059] As used herein, the term “gene construct of interest” refers to a nucleic acid sequence which is the subject of a directed evolution process. A GOI may comprise a single gene, multiple discrete genes, or a gene cluster such as an operon containing functionally related genes. The term encompasses coding sequences that encode polypeptides and functional RNA molecules, as well as non-coding sequences (such as regulatory sequences) that are operably linked to such coding sequences, including but not limited to promoters, enhancers, ribosome binding sites, terminators, and untranslated regions (UTRs).

[0060] The term “evolution” or “molecular evolution” refers to a process that results in the production of new nucleic acids and polypeptides that retain at least some of the structural features and / or functional activity of the parent nucleic acids or polypeptides from which they have developed. In some embodiments, the evolved nucleic acids or polypeptides have new or enhanced activity compared with the parent. In some embodiments, the evolved nucleic acids or polypeptides have reduced activity compared with the parent.

[0061] The term “directed evolution” or “directed molecular evolution” refers to a process for producing new nucleic acids and polypeptides with desired properties using iterative rounds of mutagenesis to generate new nucleic acids and polypeptides, followed by selection or screening to identify variants with the desired properties.

[0062] As used herein, the term “fidelity” refers to the accuracy of DNA or RNA polymerisation by a template-dependent DNA or RNA polymerase. The fidelity of a polymerase is typically measured by the error rate, i.e., the frequency of incorporating an inaccurate nucleotide (i.e., a nucleotide that is not incorporated in a template-dependent manner). The error rate or mutation rate of a polymerase may be expressed as the number of inaccurate nucleotides per base per replication cycle. This is typically denoted as a fraction, such as 10-6errors per base per replication cycle, with higher values indicating lower accuracy. Fidelity is the inverse of the error rate, i.e., a lower error rate indicates higher fidelity.

[0063] As used herein, an “error-prone polymerase” refers to a DNA or RNA polymerase that exhibits an increased error rate (i.e., increased frequency of misincorporation of nucleotides) during nucleic acid synthesis compared to high-fidelity polymerases. The error-prone polymerase may be naturally-occurring or engineered. The increased error rate may result from, for example, reduced proofreading (3’— >5’ exonuclease) activity or relaxed basepairing stringency.

[0064] As used herein, a “nucleoside deaminase” is an enzyme that catalyses the deamination of aminated nucleosides, and includes enzymes that deaminate either or both ribo- and deoxyribonucleosides. Deamination reactions occur in the nucleobase moiety, including cytosine, 5-methylcytosine, guanine and adenine nucleosides, which are transformed into their corresponding nucleoside analogues containing, respectively, uracil, thymine, xanthine and hypoxanthine as the nucleobases. Adenosine deaminases deaminate (deoxy)adenosine to give (deoxy)inosine. Cytidine deaminases deaminate (deoxy)cytidines to give (deoxy )uridine.

[0065] As used herein, the terms “wild-type” and “parent” are used interchangeably to refer to a gene or gene product (e.g.. protein) that has the characteristics (e.g.. sequence) of that gene or gene product isolated from a naturally occurring source. In contrast, the term “mutant” or “variant” refers to a gene or gene product that displays modifications in sequence when compared to the wild-type gene or gene product. Mutant genes or gene products may be naturally-occurring or synthetic (i.e., containing altered sequences that do not occur in nature).

[0066] As used herein, “a position corresponding to” or recitation that amino acid positions “correspond to” amino acid positions in a disclosed sequence, such as set forth in the sequence listing, refers to amino acid positions identified upon alignment with the disclosed sequence to maximise identity using a standard alignment algorithm or software (such as the BLAST, ClustalW, ClustalOmega, MUSCLE, TCoffee or ProbCons programmes). By aligning the sequences, one skilled in the art can identify corresponding residues.

[0067] As used herein, the term “sequence identity” refers to the extent that sequences are identical on an amino acid-by-amino acid basis over a window of comparison. Thus, a “percentage of sequence identity” is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Arg, Asn, Asp, Cys, Gin, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr and Val) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.c., the window size), and multiplying the result by 100 to yield the percentage of sequence identity. Sequence identity may be measured using sequence analysis software (for example, BLAST, ClustalW, ClustalOmega, MUSCLE, TCoffee or ProbCons programmes).

[0068] Sequence variations may arise from conservative amino acid substitutions. A “conservative amino acid substitution” is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art, which can be generally sub-classified as follows:

[0069] Table A. Amino acid sub-classification

[0070] Sub-classes Amino acids

[0071] Acidic Aspartic acid (D), Glutamic acid (E)

[0072] Basic Noncyclic: Arginine (R), Lysine (K); Cyclic: Histidine (H) Charged Aspartic acid (D), Glutamic acid (E), Arginine (R), Lysine (K),

[0073] Histidine (H)

[0074] Small Glycine (G), Serine (S), Alanine (A), Threonine (T), Proline (P) Polar / neutral Asparagine (N), Histidine (H), Glutamine (Q), Cysteine (C),

[0075] Serine (S), Threonine (T)

[0076] Polar / large Asparagine (N), Glutamine (Q)

[0077] Non-polar Tyrosine (Y), Valine (V), Isoleucine (I), Leucine (L),

[0078] Methionine (M), Phenylalanine (F), Tryptophan (W)

[0079]

[0080] Aromatic Tryptophan (W), Tyrosine (Y), Phenylalanine (F) Residues that influence Glycine (G) and Proline (P)

[0081] chain orientation

[0082]

[0083] Conservative amino acid substitution also includes groupings based on side chains. For example, a group of amino acids having aliphatic side chains is glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains is serine and threonine; a group of amino acids having amide-containing side chains is asparagine and glutamine; a group of amino acids having aromatic side chains is phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains is lysine, arginine, and histidine; and a group of amino acids having sulphur-containing side chains is cysteine and methionine. For example, it is reasonable to expect that replacement of a valine with a methionine, an aspartate with a glutamate, a threonine with a serine, or a similar' replacement of an amino acid with a structurally related amino acid will not have a major effect on the properties of the resulting mutant polypeptide.

[0084] Whether an amino acid change results in a functional polypeptide can readily be determined by assaying its activity. Conservative substitutions are shown in the table below under the heading of exemplary substitutions. Amino acid substitutions falling within the scope of the invention, are, in general, accomplished by selecting substitutions that do not differ significantly in their effect on maintaining (a) the structure of the peptide backbone in the area of the substitution, (b) the charge or hydrophobicity of the molecule at the target site, or (c) the bulk of the side chain. After the substitutions are introduced, the variants are screened for biological activity.

[0085] Table B, Exemplary amino acid substitutions

[0086] Original Residue Exemplary Substitutions

[0087] Ala Val, Leu, Ile

[0088] Arg Lys, Gin, Asn

[0089] Asn Gin, His, Lys, Arg

[0090] Asp Glu

[0091] Cys Ser

[0092] Gin Asn, His, Lys,

[0093]

[0094] Glu Asp, Lys

[0095] Gly Pro

[0096] His Asn, Gin, Lys, Arg

[0097] Ile Leu, Val, Met, Ala, Phe

[0098] Leu Ile, Val, Met, Ala, Phe

[0099] Lys Arg, Gin, Asn

[0100] Met Leu, Ile, Phe

[0101] Phe Leu, Val, Ile, Ala

[0102] Pro Gly

[0103] Ser Thr

[0104] Thr Ser

[0105] Trp Tyr

[0106] Tyr Trp, Phe, Thr, Ser

[0107] Val Ile, Leu, Met, Phe, Ala

[0108]

[0109] The term “construct” generally refers to recombinant nucleic acid, generally recombinant DNA, that has been generated for the purpose of the expression of a specific nucleic acid sequence (i.e., an “expression construct”), or is to be used in the construction of other recombinant nucleotide sequences (i.e., a “cloning construct”). The construct may be contained within a vector. Two or more constructs can be contained within a single nucleic acid molecule, such as a single vector, or can be contained within two or more separate nucleic acid molecules, such as two or more separate vectors.

[0110] An “expression construct” generally includes at least an element operably linked to a nucleic acid sequence of interest to direct expression of the nucleic acid sequence in a host cell. Such elements may include control elements such as a promoter that is operably linked to direct transcription of the nucleic acid sequence of interest, and often include a polyadenylation sequence as well. Conventional compositions and methods for preparing and using constructs and host cells arc well known to one skilled in the art, see for example, Molecular Cloning: A Laboratory Manual, 3rd edition Volumes 1, 2, and 3. J. F. Sambrook, D. W. Russell, and N. Irwin, Cold Spring Harbor Laboratory Press, 2000. As used herein “operably linked” is the association of two or more nucleic acid sequences in a construct such that the function of one is controlled by the other, for example DNA encoding a protein associated with DNA encoding a promoter.

[0111] By “control element” or “control sequence” is meant nucleic acid sequences (e.g., DNA) necessary for expression of an operably linked coding sequence in a particular host cell. The control sequences that are suitable for prokaryotic cells for example, include a promoter, and optionally a cis-acting sequence such as an operator sequence and a ribosome binding site. Control sequences that are suitable for eukaryotic cells include transcriptional control sequences such as promoters, polyadenylation signals, transcriptional enhancers, translational control sequences such as translational enhancers and internal ribosome binding sites (IRES), nucleic acid sequences that modulate mRNA stability, as well as targeting sequences that target a product encoded by a transcribed polynucleotide to an intracellular compartment within a cell or to the extracellular environment. Promoters suitable for use with expression constructs or vectors of the present disclosure include but are not limited to the phage lambda PL promoter, the E. coll lac, phoA and tac promoters, and the SV40 early and late promoters.

[0112] By “vector” is meant a nucleic acid molecule, suitably a DNA molecule derived, for example, from a plasmid, bacteriophage, yeast or virus, used to introduce heterologous nucleic acids into cells for either expression or replication thereof. The vector can be an autonomously replicating vector, i.e., a vector that exists as an extra-chromosomal entity, the replication of which is independent of chromosomal replication, e.g., a linear or closed circular plasmid, a phagemid, an extra-chromosomal element, a mini-chromosome, or an artificial chromosome. The vector can contain any means for assuring self-replication. Alternatively, the vector can be one which, when introduced into the host cell, is integrated into the genome and replicated together with the chromosome(s) into which it has been integrated. A vector system may comprise a single vector or plasmid, two or more vectors or plasmids which together contain the total DNA to be introduced into the genome of the host cell, or a transposon. The choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector may also include a selection marker such as an antibiotic resistance gene that can be used for selection of suitable transformants. Examples of such resistance genes arc well known to those of skill in the art. An “expression vector” generally includes at least an element operably linked to a nucleotide sequence of interest to direct expression of the nucleic acid sequence in a host cell. An expression vector may contain one or more unique restriction sites or multiple cloning sites and can be capable of autonomous replication in a host cell, or be integrated with the genome of the host such that the cloned sequence is reproducible.

[0113] The term “phage vector”, as used herein, refers to a nucleic acid comprising a phage genome that, when introduced into a suitable host cell, can be replicated and packaged into infectious phage particles (i.e., phage particles capable of transferring the phage genome into another host cell). The phage genome can be single- or double- stranded RNA or DNA, in either linear' or circular form. Phages and phage vectors are well known to those of skill in the art. Non-limiting examples of phages that arc useful for carrying out the methods provided herein are I (Lysogen), T2, T4, T5, T6, T7, X174, Felix 01, SP6, 029, SP01, SP10, AP50, ICP1, PVP40, and VP882. The term phage vector extends to vectors comprising truncated or partial phage genomes, and phage genomes in which certain phage genes are inactivated.

[0114] The term “phage particle”, as used herein, refers to a phage genome, for example, a DNA or RNA genome, that is associated with a coat of a phage protein or proteins (forming a capsid). A phage particle may also contain a tail (which facilitates binding to the bacterial surface and genome injection), tail fibres or spikes (for recognising and attaching to specific receptors on a bacterial host), a base plate, or a lipid envelope.

[0115] The term “infectious phage particle”, as used herein, refers to a phage particle able to transport the phage genome it comprises into a suitable host cell. Particles unable to accomplish this are referred to as a non-infectious phage particles. Phage particles may be rendered non-infectious, for example, if the capsid is unstable, or one or more of the phage tail, tail fibre or tail spike is lacking or non-functional.

[0116] A “propagation-defective phage vector” herein is a phage vector in which a gene encoding a protein essential for the generation of infectious phage particles is deleted or inactivated. In suitable host cells, however, such as host cells comprising the functional phage gene under the control of a conditional promoter, the propagation-defective phage vector can replicate and generate infectious phage particles. As used herein, the term “multiplicity of infection” or “MOI” refers to the ratio of phage vectors to host cells used during infection or transduction of host cells. For example, if 108vectors are used to transduce 106host cells, the multiplicity of infection is 100. The term encompasses introduction of a vector into a host by any method including natural infection, lipofection, microinjection, and electroporation.

[0117] As used herein, the term “phagemid” refers to a vector that derives from both a plasmid and a bacteriophage genome. A phagemid herein comprises a plasmid-derived origin of replication, a phage-derived origin of replication and a phage packaging signal, and is thus capable of propagating as a plasmid, and also be packaged into phage particles. The term phagemid is not limited to vectors having an Fl origin of replication. Similarly to a plasmid, a phagemid can be used to clone DNA fragments and be introduced into a bacterial host by a range of techniques (transformation, electroporation). Infection of a bacterial host containing a phagemid with a phage providing the necessary viral components can enable phagemid replication and packaging into phage particles. A “T7 phagemid” refers to a phagemid containing an origin of replication and a phage packaging signal derived from T7 phage.

[0118] The terms “host cell”, “host cell line” and “host cell culture” arc used interchangeably and refer to cells that can host a vector or a nucleic acid construct, including the progeny of such cells. Progeny may not be completely identical in nucleic acid content to a parent cell or may contain mutations. Mutant progeny that have the same function or biological activity as screened or selected for in the originally transformed cell are included herein.

[0119] In some embodiments, a host cell is one that can host a phage vector useful for an evolution process as provided herein. A cell can host a phage vector if it supports expression of genes from the phage vector, replication of the phage genome, and / or the generation of phage particles. For example, if the phage vector is a T7 phage, then a suitable host cell would be any cell that can support the wild-type T7 phage life cycle. Suitable host cells for phage vectors useful in evolution processes are known to those of skill in the ait, and the invention is not limited in this respect.

[0120] In other embodiments, a host cell is one that is suitable for protein expression from an expression construct or expression vector. Host cells for protein expression include, without limitation, prokaryotic hosts such as Escherichia coli, Bacillus subtilis, Pseudomonas species, Streptomyces species, and Lactobacillus species; fungal hosts such as Saccharomyces cerevisiae, Komagataella phaffii (previously known as Pichia pastoris), Hansenula polymorpha, Kluyveromyces lactis, Aspergillus species, and Trichoderma species; insect hosts including Spodoptera frugiperda (Sf9 or Sf21), Trichoplusia ni (High Five™), Drosophila S2, and Aedes albopictus (C6 / 36); and mammalian hosts, such as CHO, HEK, BHK, COS, HeLa, NIH / 3T3, SP2 / 0, YO myeloma, P3X63 myeloma, or hybridoma cells.

[0121] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (or).

[0122] As used in this application, the singular form “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “an agent” includes a plurality of agents, including mixtures thereof.

[0123] Throughout this specification and the claims which follow, unless the context requires otherwise, the word “comprise”, and variations such as “comprises” and “comprising”, will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.

[0124] Throughout this specification and the claims which follow, unless the context requires otherwise, the phrase "consisting essentially of", and variations such as "consists essentially of' will be understood to indicate that the recited element(s) is / are essential i.e. necessary elements of the invention. The phrase allows for the presence of other non-rccitcd elements which do not materially affect the characteristics of the invention but excludes additional unspecified elements which would affect the basic and novel characteristics of the method defined.

[0125] LySE (LYtic Selection and Evolution) platform

[0126] Disclosed herein is a method for directed molecular evolution, the method comprising: a) introducing a propagation-defective phage vector into a first host cell (mutagenic host) competent to propagate the phage vector,

[0127] wherein the first host cell comprises a gene construct of interest (GOI) to be evolved; and wherein the phage vector allows for (i) expression of an error-prone DNA polymerase in the first host cell for mutagenesis, and (ii) replication and packaging of the GOI into infectious phage particles;

[0128] b) incubating the first host cell under conditions for replication and mutagenesis of the GOI and release of phage particles comprising mutated GOls;

[0129] c) infecting a second host cell (selection host) with phage particles from step b); and d) selecting for a desirable function of the GOI in the second host cell.

[0130] Phage vectors

[0131] The phage vector may be made propagation-defective by deleting or inactivating one or more essential genes required for phage replication, assembly and / or infection in the phage genome. This ensures bio-containment of the phage until it infects a suitable host cell containing a functional copy of the phage gene(s) required for phage propagation. The phage propagation gene(s) may be complemented in the host cell, for example, by insertion into the host cell genome, or more preferably in a vector such as a plasmid. Gene complementation using a vector allows the methods herein to be carried out in any suitable host cell without host cell engineering.

[0132] Expression of the phage propagation gene(s) in the host cell may be conditional upon phage infection of the host cell. For example, the phage propagation gene may be operably linked to a promoter which is responsive to phage entry or to a signal generated by the phage vector. This has the advantage of allowing phage propagation to occur autonomously without further input following phage infection. Alternatively, the promoter may be an inducible promoter so that expression of the phage propagation gene may be induced a later stage in the methods herein, such as after a mutagenesis step. Controlling expression of the phage propagation gene in this way can provide for a delay between phage entry and production of phage particles so that mutagenesis of the GOI can occur in the host cell. In some embodiments, the phage vector is derived from a lytic phage genome. The use of a lytic phage allows phage propagation to occur autonomously without the need for a helper phage or exogenous induction, as would be required for temperate phages.

[0133] The phage may be a DNA or RNA phage, and may comprise a linear or circular genome. In some embodiments, the phage has a genome that is at least about 10 kb in size, such as at least about 15 kb, at least about 20 kb, at least about 25 kb, at least about 30 kb, at least about 35 kb, at least about 40 kb, at least about 45 kb, at least about 50 kb, at least about 55 kb, at least about 60 kb, at least about 65 kb, at least about 70 kb, at least about 75 kb, at least about 80 kb, at least about 85 kb, at least about 90 kb, at least about 95 kb, at least about 100 kb, at least about 105 kb, at least about 110 kb, at least about 115 kb, or at least about 120 kb. Larger genomes allow for packaging and evolution of larger gene constructs and multi-gene pathways.

[0134] In some embodiments, the phage vector is derived from a lytic phage capable of infecting Escherichia coli, Bacillus spp., Salmonella spp., or Vibrio natriegens.

[0135] In some embodiments, the lytic phage is an E. coli phage, which includes but is not limited to T2, T4, T5, T6 and T7 phage. In one embodiment, the phage vector is derived from T7 phage.

[0136] In some embodiments, the essential propagation gene that is deleted or inactivated in the phage vector is the phage DNA polymerase gene. Thus, in some embodiments, the present disclosure also provides engineered phages, for example, engineered lytic phages, with a deleted or non-functional DNA polymerase gene.

[0137] Host cells and mutagens

[0138] The host cells may be any bacteria capable of propagating the phage vector. In some embodiments, the host cell is selected from Escherichia coli. Bacillus spp., Salmonella spp., and Vibrio natriegens. In one embodiment, the host cell is an E. coli cell.

[0139] The first (mutagenic) and second (selection) host cells may be different bacteria and / or different engineered strains. Each host may be engineered to perform their respective functions in the LySE workflow, thus allowing independent optimisation of mutagenic pressure and selection stringency.

[0140] In some embodiments, the first host cell is engineered to be mutagenic. In non-limiting examples, such a host may be engineered to express an error-prone DNA polymerase, carry inducible mutagenic cassettes, reduce mismatch repair, and / or to be more tolerant of mutations or an elevated mutation load.

[0141] In some embodiments, the second host cell is engineered for phenotypic screening and / or selection of a desirable function of the GOI. For example, the second host cell may be engineered to report or to couple its fitness to a desired property of an expressed gene product. In one example, the second host cell is engineered as an auxotroph whose growth depends on an activity of the GOI. In another example, the second host cell may contain biosensor circuitry that converts the molecular activity of a GOI into a selectable or screenable signal.

[0142] In some embodiments, the phage vector comprises a deleted or inactivated DNA polymerase, which is complemented in the host cells. Thus, in some embodiments, the first and / or second host cell (preferably both first and second host cells) comprise a gene for a phage DNA polymerase capable of propagating the phage vector. Some embodiments of the methods herein may comprise a step of introducing the gene for the phage DNA polymerase into the first and / or second host cell prior to phage infection. The gene for the DNA polymerase may be inserted into the host cell genome, or comprised in a vector such as a plasmid.

[0143] In one embodiment, the phage DNA polymerase in the first and / or second host cell is a T7 DNA polymerase. The T7 DNA polymerase may comprise an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence in SEQ ID NO: 1.

[0144] In preferred embodiments, the first and / or second host cell comprise a gene for an error-prone phage DNA polymerase capable of propagating the phage vector. Advantageously, such an arrangement allows the error-prone polymerase to function both in mutagenesis and in phage propagation, thus removing the need for an additional mutagenesis agent in the host cell. In one embodiment, both the first and second host cells comprise a gene for an error-prone phage DNA polymerase.

[0145] The error-prone DNA polymerase may be a polymerase as provided in this disclosure, such as an engineered error-prone DNA polymerase described below.

[0146] When the error rate of the polymerase is very high, deleterious mutations can accumulate in the phage genome and produce non-viable phages. To ensure replication and packaging of the vector containing the GOI and completion of the lytic cycle, a higher multiplicity of infection (MOI) may be used in step a) than in step c). In this way, DNA polymerases with error rates that would exceed the genomic error threshold tolerable for phage viability can be used to provide a high rate of mutagenesis. Embodiments of the methods herein may comprise varying the MOI in step a) according to the error rate of the error-prone DNA polymerase.

[0147] In one embodiment, the MOI used in step a) is at least 1, such as an MOI of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, or more than 10. In one embodiment, an MOI of between about 2 to about 10 is used in step a).

[0148] In some embodiments, the gene for the error-prone DNA polymerase is provided in a vector (e.g., a mutagenesis vector). Methods herein may comprise introducing such a vector into the first and / or second host cell prior to step a). In one embodiment, the mutagenesis vector is a plasmid.

[0149] In some embodiments, the gene for the DNA polymerase or error-prone DNA polymerase is operably linked to a promoter which is responsive to a signal generated by the phage vector. In some embodiments, the promoter is responsive to a phage RNA polymerase produced upon entry of the phage vector into the host cell. By way of non-limiting example, the gene for a T7 DNA polymerase or error-prone T7 DNA polymerase may be operably linked to a T7 promoter in the host cell. Infection of the host by a T7 phage delivers the T7 RNA polymerase gene, which drives production of the DNA polymerase (or error-prone polymerase) by binding to the T7 promoter. In some embodiments, the host cell comprises two or more genes encoding different error-prone DNA polymerases. For example, the host cell may comprise a first error-prone polymerase comprising a conjugated adenosine deaminase (as described herein) and a second error-prone polymerase comprising a conjugated cytidine deaminase (as described herein) to achieve a balanced mutational spectrum during random mutagenesis. The expression of each error-prone polymerase may be conditional on a signal provided by the phage vector.

[0150] In some embodiments, the first host cell is engineered to express one or more other mutagens in addition to the error-prone DNA polymerase. For example, the first host cell may contain an inducible mutagenic cassette, a PACE mutagenesis plasmid, or a targeted mutagenic agent such as a CRISPR-guided DNA polymerase. The expression of these mutagens may be made conditional upon phage entry, or inducible upon the addition of an external inducer.

[0151] In other embodiments, the first host cell may be engineered to be deficient in DNA repair genes (such as mutS, mutD, mutT, mutL, mutH, uvrA, uvrB, uvrC, UNG or recA), or to express an inhibitor for DNA repair genes to reduce mismatch repair capacity and increase the spontaneous mutation rate.

[0152] In some embodiments, the first host cell is deficient in a uracil-DNA glycosylase (UNG). UNG removes uracil from DNA molecules and may reverse mutations introduced by a cytidine deaminase. It may thus be beneficial to remove or inactivate one or more UNG enzymes in the host cell to maximise mutagenesis. Alternatively, the first host cell may express a uracil glycosylase inhibitor (UGI) to inhibit the endogenous UNG. Thus, in some embodiments, the first host cell may comprise a polynucleotide comprising a nucleic acid sequence encoding a uracil glycosylase inhibitor (UGI). The gene for the UGI may be operably linked to the same promoter controlling the error-prone DNA polymerase (e.g., a T7 promoter), so that both the error-prone polymerase and the UGI are only expressed upon phage entry.

[0153] GQIs

[0154] Any desired gene construct that produces a phenotype that can be selected or screened may be targeted for directed evolution in accordance with the presently disclosed methods. Non- limiting examples include genes encoding regulatory RNAs; enzymes; binding proteins (e.g., receptors, antigen-binding molecules); regulatory proteins; antibodies; functional peptides; nutrition utilisation pathways; stress response pathways; multi-enzyme pathways; and metabolic pathways, including biosynthetic and biodegradation pathways. In some embodiments, the first host cell comprises two or more GOIs for directed evolution.

[0155] The expanded capacity of LySE broadens potential applications for continuous evolution. For example, it enables evolution of anabolic pathways for small molecule synthesis, catabolic pathways for waste assimilation or carbon capture, and protein complexes, many of which exceed 10 kb when including regulatory elements.

[0156] Phagemids

[0157] In some embodiments, the GOI is comprised in a selection vector capable of being packaged into the phage particles. For example, the selection vector may comprise a phage origin of replication for replication by the error-prone phage DNA polymerase, and phage packaging signals for packaging into phage particles. The selection vector may additionally comprise a host origin of replication and selection markers for maintenance in the host cell.

[0158] In one embodiment, the selection vector is a phagemid. In one embodiment, the phagemid is a T7 phagemid. In an alternative embodiment, the selection vector is a bacterial artificial chromosome. The phagemid or bacterial artificial chromosome may be derived from a E. coli plasmid.

[0159] The error-prone polymerase replicates the phage vector and the GOI, introducing mutations into both in the process. Replicated mutant phage vectors and mutant GOIs are packaged into individual infectious phage particles, generating a mixed population comprising progeny phages and a phage library of mutant GOIs. Production of phage particles induces host cell lysis to release the phage particles.

[0160] The completion of cell lysis may be assessed, for example, by monitoring culture turbidity (optical density, or OD), by timed incubation based on known lysis kinetics of the phagehost system being used, or by phage quantification (e.g., using a plaque assay or amplification-based methods). For example, lysis may be considered complete when the culture turbidity has dropped and stabilised (as determined by OD measurements of the culture), or when phage titres (determined via plaque assay or qPCR) plateau. A few drops of chloroform may be added to the cultures to ensure complete lysis of host cells without significantly affecting the phage particles.

[0161] The phage particles may be purified prior to infecting the second host cell. In preferred embodiments, the lysate is used directly for infection. For example, a population of the second host cells may be added to the culture of the first host cells following lysis of the first host cells.

[0162] In embodiments where an error-prone polymerase is used for replication, progeny phages will generally be less fit than the parental phage vector, and phage propagation will occur to a less extent in the second host cell, i.e., the multiplicity of infection (MOI) is generally lower for step c) than for step a). Advantageously, this allows for transduction of the second host cell (rather than phage propagation and lysis) to generate a cell library for selection and screening.

[0163] In one embodiment, the MOI used in step c) is not more than 1, such as an MOI of about 1, about 0.5, about 0.1, about 0.05, about 0.01, or less than 0.01. In one embodiment, an MOI of between about 0.05 to about 1 is used in step a).

[0164] Variant screening and selection

[0165] In some embodiments, selection in step d) comprises coupling the function of the GOI to the fitness of the second host cell, and selecting host cells with a fitness above a predetermined threshold. For example, the GOI may be part of a metabolic pathway that is required for the growth of the bacterial host under selective conditions, and selection may involve culturing the bacterial host using selective media, and screening for viable colonies or colony size indicating the presence of GOI variants with improved bioactivity. In other examples, the GOI may rescue a conditional lethal phenotype; enhance resistance to antibiotics, toxins, antimicrobial compounds, or environmental stressors; activate a growth-promoting pathway; or relieve a conditional growth inhibition mechanism. By tying the product of the gene to cell survival (e.g., by rendering the cell dependent on the product; either through auxotrophy), the directed evolution platform can be set up to continuously evolve target proteins or pathways for improved properties. Advantageously, the use of a lytic phage and an error-prone phage DNA polymerase allows methods herein to be performed largely autonomously and without external input, such as the addition of inducer molecules or mutagenic agents, or purification steps.

[0166] In some embodiments, selection in step d) comprises coupling the function of the GOI to the generation of a detectable signal in the second host cell, and selecting host cells with a signal above or below a predetermined threshold. For example, the GO1 or its gene product(s) may interact with, modulate or activate a reporter system such as a fluorescent protein, enzyme, colorimetric marker, or other quantifiable readout.

[0167] Multiple modes of selection may be performed, either sequentially or in combination, to increase the stringency of the evolution process. In some embodiments, two or more engineered host strains each optimised for a different selection mechanism are used. For example, a first strain may couple GOI activity to host fitness, while a second strain may couple the same activity to a reporter-based signal, enabling orthogonal confirmation of function. Alternatively, the host strains may allow selection for different protein activities in a gene cluster or metabolic pathway.

[0168] Selection may involve isolating, enriching and / or preferentially amplifying host cells exhibiting a desirable phenotype.

[0169] In some embodiments, the method further comprises repeating steps a) to d) by introducing a propagation-defective phage vector to the selected cell from step d). The phage vector may be a phage vector which has not been replicated by the error-prone DNA polymerase. Introduction of the phage vector may be at a higher MOI than in step d) (e.g., at an MOI of more than 1) to ensure completion of the lytic cycle and generation of a new phagemid library of GOI variants. The steps may be repeated as necessary to obtain gene products with desired functions or properties.

[0170] In some embodiments, methods herein are high-throughput and are performed in multiple chambers or vessels concurrently. Methods herein may also be automated, and may comprise, for example, the use of automated liquid handlers for liquid, reagent and / or culture transfer.

[0171] This disclosure also provides evolved products of the present methodology, including evolved genes and gene products thereof.

[0172] Kits for directed evolution

[0173] This disclosure further provides kits comprising reagents, vectors, cells, and / or apparatus for carrying out the methods provided herein.

[0174] Disclosed herein is kit for directed molecular evolution, comprising an error-prone DNA polymerase as defined herein, or an expression vector encoding such. The kit may also comprise: a selection vector, such as a phagemid or bacterial artificial chromosome, for cloning in the gene construct of interest; a propagation-defective phage vector or phage particle; and / or a host cell amenable to phage infection and capable of producing infectious phage particles.

[0175] In certain embodiments, the phage vector is a T2, T4, T5, T6 or T7 phage. The phage vector may be deficient in a phage DNA polymerase. In certain embodiments, the selection vector is a phagemid that can be packaged in a T2, T4, T5, T6 or T7 phage. The phagemid may be derived from an E. coli plasmid and thus be replicable in an E. coli host. In certain embodiments, the host cell is E. coli. The E. coli host may contain a vector encoding the error-prone DNA polymerase. In one embodiment, the E. coli host is deficient in a uracil-DNA glycosylase (UNG).

[0176] In one embodiment, the kit is for performing a method of directed evolution as provided herein.

[0177] In one embodiment, the kit comprises an expression vector comprising a gene encoding an error-prone T7 DNA polymerase. The gene for the polymerase may be operably linked to a T7 promoter. The expression vector may further comprise a gene encoding a uracil glycosylase inhibitor (UGI). The kit may further comprise one or more of: a T7 phagemid, a T7 phage vector deficient in the T7 DNA polymerase, and / or an E. coli host. In one embodiment, the E. coli host is deficient in a uracil-DNA glycosylase (UNG).

[0178] Error-prone DNA polymerases

[0179] This disclosure also provides error-prone DNA polymerases that are engineered to allow an increased frequency of random mutagenesis during DNA replication. The error-prone DNA polymerases contain mutations in the catalytic fingers domain that reduce replication fidelity, and / or contain a conjugated nucleoside deaminase to induce base switching. The error-prone DNA polymerases may be used in the directed evolution methods described in this disclosure any application where mutagenesis is required, such as in.

[0180] In some embodiments, the error-prone DNA polymerase is engineered from a wild-type phage DNA polymerase. In some embodiments, the wild-type phage polymerase is endogenous to a phage that infects E. coli, Bacillus spp., Salmonella spp. or Vibrio natriegens.

[0181] In one embodiment, the wild-type polymerase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence set forth in SEQ ID NO: 1 (T7 phage DNA polymerase). In one embodiment, the wild-type polymerase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence set forth in SEQ ID NO: 2 (T2 phage DNA polymerase). In one embodiment, the wild-type polymerase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence set forth in SEQ ID NO: 3 (T4 phage DNA polymerase). In one embodiment, the wild-type polymerase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%;, about 98%;, about 99%, or 100%; sequence identity) to an amino acid sequence set forth in SEQ ID NO: 4 (T5 phage DNA polymerase). In one embodiment, the wild-type polymerase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence set forth in SEQ ID NO: 5 (T6 phage DNA polymerase).

[0182] In one embodiment, the wild-type polymerase comprises or consists of an amino acid sequence set forth in SEQ ID NO: 1.

[0183] In some embodiments, the error-prone DNA polymerase is distinguished from the wild-type DNA polymerase by at least one amino acid substitution at a position corresponding to position 523, 479 or 560 of SEQ ID NO: 1.

[0184] In some embodiments, the error-prone DNA polymerase comprises one or more amino acid substitutions at a position corresponding to position 523, 479 or 560 of SEQ ID NO: 1, and further comprises one or more amino acid substitutions at a position corresponding to position 5, 7, 64, 120, 399, 429, 443, 444, 480, 520, 521, 522, 524, 530 or 611 of SEQ ID NO: 1.

[0185] In some embodiments, the error-prone DNA polymerase comprises one or more amino acid substitutions at a position corresponding to position 523, 479 or 560 of SEQ ID NO: 1, and further comprises one or more amino acid substitutions at a position corresponding to position 5, 7, 64, 120 or 399 of SEQ ID NO: 1.

[0186] The amino acid substitution may be a substitution to any one of the canonical amino acids, i.e., a substitution to A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W or Y.

[0187] In some embodiments, the error-prone DNA polymerase comprises an amino acid substitution at positions corresponding to one of the following sets of positions of SEQ ID NO: 1:

[0188] (a) 523;

[0189] (b) 479 and 523;

[0190] (c) 523 and 560; (d) 479, 523 and 560;

[0191] (e) 64, 120, 399 and 479;

[0192] (f) 64, 120, 399 and 523;

[0193] (g) 64, 120, 399 and 560;

[0194] (h) 64, 120, 399, 479 and 523;

[0195] (i) 64, 120, 399, 479 and 560;

[0196] (j) 64, 120, 399, 523 and 560;

[0197] (k) 64, 120, 399, 479, 523 and 560;

[0198] (l) 5, 7, 64, 120, 399 and 523;

[0199] (m) 5, 7, 64, 120, 399, 479 and 523;

[0200] (n) 5, 7, 64, 120, 399, 523 and 560; or

[0201] (o) 5, 7, 64, 120, 399, 479, 523 and 560.

[0202] In some embodiments, the error-prone DNA polymerase comprises an amino acid substitution at positions corresponding to one of the following sets of positions of SEQ ID NO: 1:

[0203] (a) 64, 120, 399 and 479;

[0204] (b) 64, 120, 399 and 523;

[0205] (c) 64, 120, 399 and 560;

[0206] (d) 64, 120, 399, 479 and 523;

[0207] (e) 64, 120, 399, 479 and 560;

[0208] (f) 64, 120, 399, 523 and 560;

[0209] (g) 64, 120, 399, 479, 523 and 560;

[0210] (h) 5, 7, 64, 120, 399 and 523;

[0211] (i) 5, 7, 64, 120, 399, 479 and 523;

[0212] (j) 5, 7, 64, 120, 399, 523 and 560; or

[0213] (k) 5, 7, 64, 120, 399, 479, 523 and 560.

[0214] In some embodiments, the error-prone DNA polymerase comprises one or more amino acid substitutions selected from the following:

[0215] (a) a substitution to A at a position corresponding to position 5 of SEQ ID NO: 1 (b) a substitution to A at a position corresponding to position 7 of SEQ ID NO: 1 (c) a substitution to C at a position corresponding to position 64 of SEQ ID NO: 1; (d) a substitution to L at a position corresponding to position 120 of SEQ ID NO: 1; (e) a substitution to T at a position corresponding to position 399 of SEQ ID NO: 1; (f) a substitution to N at a position corresponding to position 479 of SEQ ID NO: 1; (g) a substitution to R at a position corresponding to position 523 of SEQ ID NO: 1; and (h) a substitution to H at a position corresponding to position 560 of SEQ ID NO: 1;

[0216] Disclosed herein is an error-prone DNA polymerase comprising an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence in SEQ ID NO: 1, wherein the error-prone polymerase comprises at least one of the following:

[0217] a) a N at a position corresponding to position 479 of SEQ ID NO: 1;

[0218] b) a R at a position corresponding to position 523 of SEQ ID NO: 1;

[0219] c) a H at a position corresponding to position 560 of SEQ ID NO: 1; and / or d) a conjugated nucleoside deaminase.

[0220] In some embodiments, the error-prone DNA polymerase further comprises at least one of the following:

[0221] a) a A at a position corresponding to position 5 of SEQ ID NO: 1;

[0222] b) a A at a position corresponding to position 7 of SEQ ID NO: 1;

[0223] c) a C at a position corresponding to position 64 of SEQ ID NO: 1;

[0224] d) a L at a position corresponding to position 120 of SEQ ID NO: 1; and / or

[0225] e) a T at a position corresponding to position 399 of SEQ ID NO: 1.

[0226] In some embodiments, the error-prone DNA polymerase of this disclosure comprises a nucleoside deaminase. The coupling of a nucleoside deaminase to the DNA polymerase allows the introduction of random base-pair changes during DNA replication, and may increase the error rate of the DNA polymerase. For example, adenine deaminases can convert A-T base pairs to G-C, and cytosine deaminases can convert C-G base pairs to T-A. The increase in the error rate provided by deaminase conjugation may further contribute to the mutagenicity of the polymerase. In some embodiments, the nucleoside deaminase has adenine deaminase and / or cytosine deaminase activity. The nucleoside deaminase may be capable of deaminating nucleosides, deoxy nucleosides, or both nucleosides and deoxynucleosides.

[0227] In some embodiments, the adenine deaminase is TadA-7.10, TadA-8e, or a variant thereof. TadA-7.10 is described in Gaudelli, N. M. et al. Nature 551, 464-471 (2017), and TadA-8e is described in Richter, M. F. et al. Nat. Biotechnol. 38, 883-891 (2020). Both references are incorporated by reference in their entirety herein. Variants of TadA-7.10 and TadA-8 contemplated herein include natural and engineered enzymes with improved DNA-binding affinity, catalytic activity, substrate range (e.g., ability to deaminate cytidine), overall mutagenic activity, reduced off-target activity (i.e., deamination activity that is not in conjunction with polymerase activity), or smaller size. Exemplary variants are described in Neugebauer, M. E. et al., Nat. Biotechnol. 41, 673-685 (2023).

[0228] In some embodiments, the cytosine deaminase is AID, APOBEC1, APOBEC3A, APOBEC3G, PmCDAl, or a variant thereof. Variants contemplated herein include natural and engineered enzymes with improved DNA-binding affinity, catalytic activity, substrate range (e.g., ability to deaminate adenosine), overall mutagenic activity, reduced off-target activity (i.e., deamination activity that is not in conjunction with polymerase activity), or smaller size.

[0229] In one embodiment, the nucleotide deaminase has both adenine deaminase and cytosine deaminase activity. A non-limiting example of a nucleotide deaminase with adenine and cytosine deaminase activity is TadDE (SEQ ID NO: 29).

[0230] In one embodiment, the nucleoside deaminase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence in SEQ ID NO: 27-29.

[0231] The nucleoside deaminase may be conjugated to the N- or C-terminus of the DNA polymerase, optionally via a linker. In preferred embodiments, the nucleoside deaminase is conjugated to the N-terminus of the DNA polymerase. N-terminal conjugation can reduce the impact of the additional domain on polymerase activity.

[0232] In some embodiments, the nucleoside deaminase is conjugated to the DNA polymerase via a polypeptide linker. In some embodiments, the polypeptide linker comprises at least 24 amino acids. In some embodiments, the polypeptide linker comprises between 24 and 33 amino acids, such as 24, 25, 26, 27, 28, 29, 30, 31, 32, or 33 amino acids. The inventors have found that shorter linkers can increase target mutagenesis. A minimum of 24 amino acids is optimal to maintain polymerase functionality after deaminase conjugation. Non-limiting examples of polypeptide linker sequences are provide as SEQ ID NO: 30-32.

[0233] In some embodiments, the error-prone DNA polymerase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence set forth in SEQ ID NO: 7 to 26. In some embodiments, the error-prone DNA polymerase comprises or consists of an amino acid sequence set forth in SEQ ID NO: 7 to 26.

[0234] In some embodiments, the error-prone polymerase comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%>, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%>, about 99%, or 100% sequence identity) to an amino acid sequence set forth in SEQ ID NO: 24 or SEQ ID NO: 26. In some embodiments, the error-prone polymerase comprises or consists of an amino acid sequence set forth in SEQ ID NO: 24, or SEQ ID NO: 26.

[0235] The modifications of this disclosure, such as amino acid substitutions and deaminase conjugation, can improve the error rate (or reduce the fidelity) of the DNA polymerase.

[0236] For example, the error-prone DNA polymerase of this disclosure may exhibit a higher rate of misincorporation of nucleotides during DNA polymerisation compared to a corresponding wild-type DNA polymerase. The increase in the error rate may be an increase of at least about 10%, for example, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%>, about 70%, about 80%, about 90%, about 100%, or about 2-fold, about 5- fold, about 10-fold, about 50-fold, about 100-fold, about 500-fold, about 1000-fold, about 5000-fold, about 10,000-fold, about 50,000-fold, about 100,000-fold, about 500,000-fold, or greater than 500,000-fold when compared to the error rate of a corresponding wild-type DNA polymerase.

[0237] Assays to determine fidelity or error rate are known in the art. One exemplary assay is the Luria-Delbrück fluctuation assay. In the fluctuation assay, a premature stop codon is added to the coding sequence of a selectable marker, e.g., an antibiotic resistance gene. Replication using an error-prone polymerase allows reversion of the stop codon to a sense codon, enabling growth on the antibiotic. By counting the number of viable colonies and applying statistical methods, such as the Poisson distribution, the mutation rate can be estimated. The fluctuation assay is suitable for detecting low mutation rates given the high sensitivity of the assay. Another exemplary assay is the LacZ inactivation assay. In this assay, the polymerase variant is used to replicate a LacZ gene on a plasmid. A higher mutation rate increases the rate of introducing deleterious mutations to the LacZ gene, thereby inactivating the LacZ gene. Growth of bacteria on agar containing X-gal can determine rate of inactivation by blue-white colony counting (with blue colonies having functional LacZ and white colonies having inactivated LacZ). This assay is generally less sensitive and usually used when the mutation rate is high.

[0238] Disclosed herein is a polynucleotide comprising a nucleic acid sequence encoding an error-prone DNA polymerase as defined herein. In one embodiment, the polynucleotide also comprises a nucleic acid sequence encoding a uracil glycosylase inhibitor (UGI). In one embodiment, the UGI comprises an amino acid sequence having at least 70% sequence identity (such as about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100% sequence identity) to an amino acid sequence in SEQ ID NO: 33.

[0239] Disclosed herein is an expression vector comprising a polynucleotide as defined herein. In one embodiment, the nucleic acid sequence encoding the error-prone DNA polymerase in the expression vector is operably linked to an inducible promoter. The inducible promoter may be responsive to a phage RNA polymerase. In one embodiment, the expression vector comprises a gene for the error-prone DNA polymerase operably linked to a T7 promoter. Disclosed herein is a host cell comprising a polynucleotide or expression vector as defined herein. The host cell may be any expression host (e.g., a prokaryotic or eukaryotic cell) capable of expressing a polypeptide encoded by the polynucleotide or vector.

[0240] Those skilled in the field of molecular biology will understand that any of a wide variety of expression systems may be used to provide the error-prone DNA polymerase defined herein. The DNA polymerase may be produced in a prokaryotic host (e.g., E. coli) or in a eukaryotic host (e.g., Saccharomyces cerevisiae, insect cells, e.g., Sf21 cells, or mammalian cells, e.g., NIH 3T3, HeLa, COS cells). Such cells are available from a wide range of sources (e.g., the American Type Culture Collection, Rockland, MD). Non-limiting examples of insect cells are, Spodoptera frugiperda (Sf) cells, e.g., Sf9, Sf21, Trichoplusia ni cells, e.g.. High Five cells, and Drosophila S2 cells. Examples of fungi (including yeast) host cells are S. cerevisiae, Kluyveromyces lactis (K. lactis). species of Candida including C. albicans and C. glabrata, Aspergillus nidulans, Schizosaccharomyces pombe (S. pombe), Komagataella pastoris (previously known as Pichia pastoris), and Yarrowia lipolytica. Examples of mammalian cells are COS cells, baby hamster kidney cells, mouse L cells, LNCaP cells, Chinese hamster ovary (CHO) cells, human embryonic kidney (HEK) cells, African green monkey cells, CV1 cells, HeLa cells, MDCK cells, Vero and Hep-2 cells. Xenopus laevis oocytes, or other cells of amphibian origin, may also be used. Prokaryotic host cells include bacterial cells, for example, E. coli, B. subtilis, and mycobacteria.

[0241] Disclosed herein is a method of producing an error-prone DNA polymerase as disclosed herein, the method comprising culturing a host cell as defined herein under conditions suitable for expressing the DNA polymerase.

[0242] Methods to grow cells that produce the DNA polymerase of this disclosure include, but are not limited to, batch, batch-fed, continuous and perfusion cell culture techniques. Typically, cell culture is performed under sterile, controlled temperature and atmospheric conditions. A bioreactor is a chamber used to culture cells in which environmental conditions such as temperature, atmosphere, agitation and / or pH can be monitored. The bioreactor can be a stainless steel chamber or a pre-sterilised plastic bag (e.g., Cellbag. RTM., Wave Biotech, Bridgewater, N. J.). The bioreactor may be dimensioned for cultures of about 10 L to about 50,000 L. In some embodiments, the method further comprises recovering the DNA polymerase from the cell culture. Proteins may be isolated from the cell and / or media fraction using methods that preserve their integrity, such as by gradient centrifugation, e.g., caesium chloride, sucrose and iodixanol, as well as standard purification techniques including, e.g., ion exchange, gel filtration or affinity chromatography.

[0243] Table 1. Exemplary wild-type DNA polymerases and error-prone variants (linkers in bold) SEQ DNA Amino acid sequence

[0244] ID polymerase

[0245] NO. (DNAP)

[0246] 1 Wild-type MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE (WT) T7 AEVARGGLIVFHNGHKYDVPALTKLAKLQLNREFHLPRENCIDTLVL DNAP SRLIHSNLKDTDMGLLRSGKLPGKRFGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVVVTKALLEKLLSD KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAATDLTKEYLMTQKRIGQSAEGDKAWLRYVAEDGKTHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALTCKLWT TKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0247] 2 Wild-type MKEFYISIETVGNNIIERYIDENGKERTREVEYLPTMFRHCKEESKY (WT) T2 KDIYGKNCAPQKFPSMKDARDWMKRMEDIGLEALGMNDFKLAYISDT DNAP YGSEIVYDRKFVRVANCDTEVTGDKFPDPMKAEYETDAITHYDSIDD RFYVFDLLNSMYGSVSKWDAKLAAKLDCEGGDEVPQEILDRVIYMPF DNERDMLMEYINLWEQKRPAIFTGWNIEGFDVPYIMNRVKMILGERS MKRFSPIGRVKSKLIQNMYGSKEIYSIDGVSILDYLDLYKKFAFTNL PSFSLESVAQHETKKGKLPYDGPINKLRETNHQRYISYNIIDVESVQ AIDKIRGFIDLVLSMSYYAKMPFSGVMSPIKTWDAIIFNSLKGEHKV IPQQGSHVKQSFPGAFVFEPKPIARRYIMSFDLTSLYPSIIRQVNIS PETIRGQFKVHPIHEYIAGTAPKPSDEYSCSPNGWMYDKHQEGIIPK EIAKVFFQRKDWKKKMFAEEMNAEAIKKIIMKGAGSCSTKPEVERYV

[0248]

[0249] KFTDDFLNELSNYTESVLNSLTEECEKAATLANTNQLNRKTLTNSLY GALGNIHFRYYDLRNATAITIFGQVGIQWIARKINEYLNKVCGTNDE DFIAAGDTDSVYVCVDKVIEKVGLDRFKEQNDLVEFMNQFGKKKMEP MIDVAYRELCDYMNNREHLMHMDREAISCPPLGSKGVGGFWKAKKRY ALNVYDMEDKRFAEPHLKIMGMETQQSSTPKAVQEALEESIRRILQE GEESVQEYYKNFEKEYRQLDYKVIAEVKTANDIAKYDDKGWPGFKCP FHIRGVLTYRRAVSGLGVAPILDGNKVMVLPLREGNPFGDKCIAWPS GTELPKEIRSDVLSWIDYSTLFQKSFVKPLAGMCESAGMDYEEKASL DFLFG

[0250] Wild-type MKEFYTSIETVGNNTVERYIDENGKERTREVEYLPTMFRHCKEESKY (WT) T4 KDIYGKNCAPQKFPSMKDARDWMKRMEDIGLEALGMNDFKLAYISDT DNAP YGSEIVYDRKFVRVANCDIEVTGDKFPDPMKAEYEIDAITHYDSIDD RFYVFDLLNSMYGSVSKWDAKLAAKLDCEGGDEVPQEILDRVIYMPF DNERDMLMEYINLWEQKRPAIFTGWNIEGFDVPYIMNRVKMILGERS MKRFSPIGRVKSKLIQNMYGSKEIYSIDGVSILDYLDLYKKFAFTNL PSFSLESVAQHETKKGKLPYDGPINKLRETNHQRYISYNIIDVESVQ AIDKIRGFIDLVLSMSYYAKMPFSGVMSPIKTWDAIIFNSLKGEHKV IPQQGSHVKQSFPGAFVFEPKPIARRYIMSFDLTSLYPSIIRQVNIS PETIRGQFKVHPIHEYIAGTAPKPSDEYSCSPNGWMYDKHQEGIIPK EIAKVFFQRKDWKKKMFAEEMNAEAIKKIIMKGAGSCSTKPEVERYV KFSDDFLNELSNYTESVLNSLIEECEKAATLANTNQLNRKILINSLY GALGNIHFRYYDLRNATAITIFGQVGIQWIARKINEYLNKVCGTNDE DFIAAGDTDSVYVCVDKVIEKVGLDRFKEQNDLVEFMNQFGKKKMEP MIDVAYRELCDYMNNREHLMHMDREAISCPPLGSKGVGGFWKAKKRY ALNVYDMEDKRFAEPHLKIMGMETQQSSTPKAVQEALEESIRRILQE GEESVQEYYKNFEKEYRQLDYKVIAEVKTANDIAKYDDKGWPGFKCP FHIRGVLTYRRAVSGLGVAPILDGNKVMVLPLREGNPFGDKCIAWPS GTELPKEIRSDVLSWIDHSTLFQKSFVKPLAGMCESAGMDYEEKASL DFLFG

[0251] Wild-type MKIAVVDKALNNTRYDKHFQLYGEEVDVFHMCNEKLSGRLLKKHITI (WT) T5 GTPENPFDPNDYDFVILVGAEPFLYFAGKKGIGDYTGKRVEYNGYAN DNAP WIASISPAQLHFKPEMKPVFDATVENIHDIINGREKIAKAGDYRPIT DPDEAEEYIKMVYNMVIGPVAFDSETSALYCRDGYLLGVSISHQEYQ GVYIDSDCLTEVAVYYLQKILDSENHTIVFHNLKFDMHFYKYHLGLT FDKAHKERRLHDTMLQHYVLDERRGTHGLKSLAMKYTDMGDYDFELD KFKDDYCKAHKIKKEDFTYDLIPFDIMWPYAAKDTDATIRLHNFFLP KIEKNEKLCSLYYDVLMPGCVFLQRVEDRGVPISIDRLKEAQYQLTH

[0252]

[0253] NLNKAREKLYTYPEVKQLEQDQNEAFNPNSVKQLRVLLFDYVGLTPT GKLTDTGADSTDAEALNELATQHPIAKTLLEIRKLTKLISTYVEKIL LSIDADGCIRTGFHEHMTTSGRLSSSGKLNLQQLPRDESIIKGCVVA PPGYRVIAWDLTTAEVYYAAVLSGDRNMQQVFINMRNEPDKYPDFHS NIAHMVFKLQCEPRDVKKLFPALRQAAKAITFGILYGSGPAKVAHSV NEALLEQAAKTGEPFVECTVADAKEYIETYFGQFPQLKRWIDKCHDQ IKNHGFIYSHFGRKRRLHNIHSEDRGVQGEEIRSGFNAIIQSASSDS LLLGAVDADNE 11 SLGLEQEMKIVMLVHD SVVAIVREDL IDQYNE IL IRNIQKDRGISIPGCPIGIDSDSEAGGSRDYSCGKMKKQHPSIACID DDEYTRYVKGVLLDAEFEYKKLAAMDKEHPDHSKYKDDKFIAVCKDL DNVKRILGA

[0254] Wild-type MKEFYISIETVGNNIVERYIDENGKERTREVEYLPTMFRHCKEESKY (WT) T6 KDIYGKNCAPQKFPSMKDARDWMKRMEDIGLEALGMNDFKLAYISDT DNAP YGSEIVYDRKFVRVANCDIEVTGDKFPDPMKAEYEIDAITHYDSIDD RFYVFDLLNSMYGSVSKWDAKLAAKLDCEGGDEVPQEILDRVIYMPF DNERDMLMEYINLWEQKRPAIFTGWNIEGFDVPYIMNRVKMVLGERS MKRFSPIGRVKSKLIQNMYGSKEIYSIDGVSILDYLDLYKKFAFTNL PSFSLESVAQHETKKGKLPYDGPINKLRETNHQRYISYNIIDVESVQ AIDKIRGFIDLVLSMSYYAKMPFSGVMSPIKTWDAIIFNSLKGEHKV IPQQGSHVKQSFPGAFVFEPKPIARRYIMSFDLTSLYPSIIRQVNIS PETIRGQFKVHPIHEYIAGTAPKPSDEYSCSPNGWMYDKHQEGIIPK EIAKVFFQRKDWKKKMFAEEMNAEAIKKIIMKGAGSCSTKPEVERYV KFNDDFLNELSNYTESVLNSLIEECEKAATLANTNQLNRKILINSLY GALGNIHFRYYDLRNATAITIFGQVGIQWIARKINEYLNKVCGTNGE DFIAAGDTDSVYVCVDKVIEKVGLDRFKEQNDLVEFMNQFGKKKMEP MIDVAYRELCDYMNNREHLMHMDREAISCPPLGSKGVGGFWKAKKRY ALNVYDMEDKRFAEPHLKIMGMETQQSSTPKAVQEALEESIRRILQE GEESVQEYYKNFEKEYRQLDYKVIAEVKTANDIAKYDDKGWPGFKCP FHIRGVLTYRRAVSGLGVAPILDGNKVMVLPLREGNPFGDKCIAWPS GTELPKEIRSDVLSWIDYSTLFQKSFVKPLAGMCESAGMDYEEKASL DFLFG

[0255] T7 DNAP vl MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE (Y64C, AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL F120L, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVVVTKALLEKLLSD

[0256] S399T)

[0257] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0258]

[0259] LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0260] T7 DNAP MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v2.2 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0261] F120L,

[0262] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0263] S399T,

[0264] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0265] L479N) LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0266] T7 DNAP MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v2.4 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0267] F120L,

[0268] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0269] S399T,

[0270] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0271] T523R) LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN

[0272]

[0273] TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0274] T7 DNAP MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v2.5 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0275] F120L,

[0276] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0277] S399T,

[0278] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0279] P560H) LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTHAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0280] T7 DNAP MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v2.7 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0281] F120L,

[0282] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0283] S399T,

[0284] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0285] L479N, LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV T523R) EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAATDLTKEYLMTQKRIGQTAEGDKAWLRYVAEDGKTHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0286] T7 DNAP MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v2.8 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0287] F120L,

[0288] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0289] S399T,

[0290]

[0291] L479N, TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP P560H) LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTHAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQVVIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0292] T7 DNAP MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v2.11 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVVVTKALLEKLLSD

[0293] F120L,

[0294] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0295] S399T,

[0296] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0297] T523R, LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV P560H) EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTHAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0298] T7 DNAP MIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v2.13 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0299] F120L,

[0300] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0301] S399T,

[0302] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0303] L479N,

[0304] LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV

[0305] T523R, EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD P560H) PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTHAIAA

[0306]

[0307] LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0308] T7 DNAP v3 MIVSAIAANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE (D5A, E7A, AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVVVTKALLEKLLSD

[0309] F120L,

[0310] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0311] S399T,

[0312] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0313] L479N, LPKYPRIKTPKVGGTFKKPKNKAQREGREPCELDTREYVAGAPYTPV T523R, EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD P560H) PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0314] T7 DNAP v4 MIVSAIAANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE (D5A, E7A, AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL Y64C, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0315] F120L,

[0316] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0317] S399T,

[0318] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0319] T523R) LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0320] T7 DNAP MIVSAIAANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v4.2 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (D5A, E7A, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0321] Y64C,

[0322]

[0323] F120L, KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD S399T, TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP L479N, LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDD

[0324] T523R)

[0325] PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0326] T7 DNAP MIVSAIAANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALE v4.3 AEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVL (D5A, E7A, SRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKD DFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSD

[0327] Y64C,

[0328] KHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFD

[0329] F120L,

[0330] TKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKP

[0331] S399T, LPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPV L479N, EHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPVVDDEVLEGVRVDD T523R, PEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPN P560H) GAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWV QAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPT RDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTHAIAA LRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALN TLLQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQV GCRTEEIAQVVIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0332] T7 DNAP MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNR v5.3 Al GLHDP TAHAE IMALRQGGLVMQNYRLIDATLYVTFEP CVMCAGAM (TadA-8e- I H S R I GRWF GVRNS KRGAAG S LMNVLNY P GMNHRVE I T E G I LAD E C AAL L C DF YRMP RQ VFNAQKKAQ S S I N SGGGSETPGTSE S ATPE SGGS

[0333] 24aa-WT T7

[0334] IKGIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDA DNAP)

[0335] LEAEVARGGLIVFHNGHKYDVPALTKLAKLQLNREFHLPRENCIDTL VLSRLIHSNLKDTDMGLLRSGKLPGKRFGSHALEAWGYRLGEMKGEY KDDFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVWTKALLEKLL SDKHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFP FDTKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTG KPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYT

[0336]

[0337] PVEHVVFNPSSRDHTQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRV DDPEKQAAIDLIKEYLMIQKRIGQSAEGDKAWLRYVAEDGKIHGSVN PNGAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKP WVQAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAEL PTRDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAI AALRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAA LNTLLQSAGALICKLWIIKTEEMLVEKGLKHGWDGDFAYMAWVHDEI QVGCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAIC H

[0338] T7 DNAP MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVTGEGWNR v5.5 Al GLHDP TAHAE IMALRQGGLVMQNYRLIDATLYVTFEP CVMCAGAM (TadA-8e- I H S R I GRWF GVRNS KRGAAG S LMNVLNY P GMNHRVE I T E G I LAD E C AALLCDFYRMPRQVFNAQKKAQSSINSGSETPGTSESATPESGGSDY

[0339] 33aa-WT T7

[0340] KDDDDKGSLIKGIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRP DNAP)

[0341] SDFGAYLDALEAEVARGGLIVFHNGHKYDVPALTKLAKLQLNREFHL PRENCIDTLVLSRLIHSNLKDTDMGLLRSGKLPGKRFGSHALEAWGY RLGEMKGEYKDDFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVW TKALLEKLLSDKHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLL AKQERNGFPFDTKAIEELYVELAARRSELLRKLTETFGSWYQPKGGT EMFCHPRTGKPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTR EYVAGAPYTPVEHWFNPSSRDHIQKKLQEAGWVPTKYTDKGAPVVD DEVLEGVRVDDPEKQAAIDLIKEYLMIQKRIGQSAEGDKAWLRYVAE DGKIHGSVNPNGAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEH HLDGITGKPWVQAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIH TKNQIAAELPTRDNAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKK KFLENTPAIAALRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRK VHVRSPHAALNTLLQSAGALICKLWIIKTEEMLVEKGLKHGWDGDFA YMAWVHDEIQVGCRTEEIAQVVIETAQEAMRWVGDHWNFRCLLDTEG KMGPNWAICH

[0342] T7 DNAP MMTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERR v6.1 ACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFT INWYSS (PmCDAl- WSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWN LRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKW

[0343] 24aa-WT T7

[0344] RSELSIMIQVKILHTTKSPAVSSGGGSETPGTSESATPESGGSIKGI DNAP)

[0345] VSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALEAE VARGGLIVFHNGHKYDVPALTKLAKLQLNREFHLPRENCIDTLVLSR LIHSNLKDTDMGLLRSGKLPGKRFGSHALEAWGYRLGEMKGEYKDDF

[0346]

[0347] KRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSDKH YFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFDTK AIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKPLP KYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEH VVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPVVDDEVLEGVRVDDPE KQAAIDLIKEYLMIQKRIGQSAEGDKAWLRYVAEDGKIHGSVNPNGA VTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWVQA GIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPTRD NAKTFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAALR ESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALNTL LQSAGALICKLWI IKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQVGC RTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0348] T7 DNAP v7 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNR (TadA-8e- Al GLHDP TAHAE IMALRQGGLVMQNYRLIDATLYVTFEP CVMCAGAM 24aa-Y64C, I H S R I GRWF G VRNS KRGAAG S LMNVLNY P GMNHRVE I T E G I LAD E C AAL L C DF YRMP RQVFNAQKKAQ S S I N SGGGSETPGTSE S ATPE SGGS

[0349] F120L,

[0350] IKGIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDA

[0351] S399T,

[0352] LEAEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTL

[0353] T523R) VLSRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEY KDDFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVWTKALLEKLL SDKHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFP FDTKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTG KPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYT PVEHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRV DDPEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVN PNGAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKP WVQAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAEL PTRDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAI AALRESTQQTLVESSQWVAGEQQVKWKRRWTKGLDGRKVHVRSPHAA LNTLLQSAGALICKLWIIKTEEMLVEKGLKHGWDGDFAYMAWVHDEI QVGCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAIC H

[0354] T7 DNAP MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNR v7.2 Al GLHDP TAHAE IMALRQGGLVMQNYRLIDATLYVTFEP CVMCAGAM (TadA-8e- I H S R I GRWF G VRNS KRGAAG S LMNVLNY P GMNHRVE I T E G I LAD E C AAL L C DF YRMP RQVFNAQKKAQ S S I N SGGGSETPGTSE S ATPE SGGS

[0355] 24aa-Y64C,

[0356] IKGIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDA

[0357] F120L,

[0358]

[0359] S399T, LEAEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTL L479N, VLSRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEY T523R) KDDFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVWTKALLEKLL SDKHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFP FDTKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTG KPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYT PVEHWFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRV DDPEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVN PNGAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKP WVQAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAEL PTRDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAI AALRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAA LNTLLQSAGALICKLWIIKTEEMLVEKGLKHGWDGDFAYMAWVHDEI QVGCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAIC H

[0360] T7 DNAP MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNR v7.3 AT GLHDP TAHAE IMALRQGGLVMQNYRLIDATLYVTFEP CVMCAGAM (TadA-8e- I H S R I GRWF GVRNS KRGAAG S LMNVLNY P GMNHRVE I T E G I LAD E C AAL L C DF YRMP RQ VFNAQKKAQ S S I N SGGGSETPGTSE S ATPE SGGS

[0361] 24aa-Y64C,

[0362] IKGIVSDIEANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDA

[0363] F120L,

[0364] LEAEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTL

[0365] S399T, VLSRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEY L479N, KDDFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVWTKALLEKLL T523R, SDKHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFP P560H) FDTKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTG KPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYT PVEHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRV DDPEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVN PNGAVTGRATHAFPNLAQTPGVRSPYGEQCRAAFGAEHHLDGTTGKP WVQAGIDASGNELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAEL PTRDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTHAI AALRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAA LNTLLQSAGALICKLWIIKTEEMLVEKGLKHGWDGDFAYMAWVHDEI QVGCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAIC H

[0366] T7 DNAP v8 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNR Al GLHDP TAHAE IMALRQGGLVMQNYRLIDATLYVTFEP CVMCAGAM

[0367]

[0368] (TadA-8e- I H S R I GRWF GVRNS KRGAAG S LMNVLNY P GMNHRVE I T E G I LAD E C 24aa-D5A, AAL L C DF YRMP RQVFNAQKKAQ S S I N SGGGSETPGTSE S ATPE SGGS E7A, Y64C, IKGIVSAIAANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDA LEAEVARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTL

[0369] F120L,

[0370] VLSRLIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEY

[0371] S399T,

[0372] KDDFKRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDVWTKALLEKLL

[0373] T523R) SDKHYFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFP FDTKAIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTG KPLPKYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYT PVEHVVFNPSSRDHIQKKLQEAGWVPTKYTDKGAPVVDDEVLEGVRV DDPEKQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVN PNGAVTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKP WVQAGIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAEL PTRDNAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAI AALRESIQQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAA LNTLLQSAGALICKLWIIKTEEMLVEKGLKHGWDGDFAYMAWVHDEI QVGCRTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAIC H

[0374] T7 DNAP MMTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERR v8.2 ACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFT INWYSS (PmCDA-8e- WSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWN LRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKW

[0375] 24aa-D5A,

[0376] REELS IMIQVKILHTTKSPAVSSGGGSETPGTSESATPESGGSIKGI

[0377] E7A, Y64C,

[0378] VSAIAANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALEAE

[0379] F120L,

[0380] VARGGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVLSR

[0381] S399T, LIHSNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKDDF T523R) KRMLEEQGEEYVDGMEWWNFNEEMMDYNVQDWVTKALLEKLLSDKH YFPPEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFDTK AIEELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKPLP KYPRIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEH WFNPSSRDHIQKKLQEAGWVPTKYTDKGAPWDDEVLEGVRVDDPE KQAAIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPNGA VTGRATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWVQA GIDASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPTRD NAKRFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAALR ESTQQTLVESSQWVAGEQQVKWKRRWTKGLDGRKVHVRSPHAALNTL

[0382]

[0383] LQSAGALICKLWT IKTEEMLVEKGLKHGWDGDFAYMAWVHDETQVGC RTEEIAQWIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0384] 26 T7 DNAP v9 MSE VE F SHE YWMRHALT LAKRARDE GE AP VGAVLVLNNRVI GE GWNR (TadDE- RIGLHDPTAHAEIMALRQGGLVMQNSRLIDATLYVTFEPCVMCAGAM 24aa-D5A, TNSRIGRWFGVRNSKRGAAGSLMNVLNYPGMNVTGLADECAALCDF YRMPRQVFNAQKKAQSSINSGGGSETPGTSESATPESGGSIKGIVSA

[0385] E7A, Y64C,

[0386] IAANALLESVTKFHCGVIYDYSTAEYVSYRPSDFGAYLDALEAEVAR

[0387] F120L,

[0388] GGLIVFHNGHKCDVPALTKLAKLQLNREFHLPRENCIDTLVLSRLIH

[0389] S399T, SNLKDTDMGLLRSGKLPGKRLGSHALEAWGYRLGEMKGEYKDDFKRM T523R) LEEQGEEYVDGMEWWNFNEEMMDYNVQDVWTKALLEKLLSDKHYFP PEIDFTDVGYTTFWSESLEAVDIEHRAAWLLAKQERNGFPFDTKAIE ELYVELAARRSELLRKLTETFGSWYQPKGGTEMFCHPRTGKPLPKYP RIKTPKVGGIFKKPKNKAQREGREPCELDTREYVAGAPYTPVEHVVF NPSSRDHIQKKLQEAGWVPTKYTDKGAPVVDDEVLEGVRVDDPEKQA AIDLIKEYLMIQKRIGQTAEGDKAWLRYVAEDGKIHGSVNPNGAVTG RATHAFPNLAQIPGVRSPYGEQCRAAFGAEHHLDGITGKPWVQAGID ASGLELRCLAHFMARFDNGEYAHEILNGDIHTKNQIAAELPTRDNAK RFIYGFLYGAGDEKIGQIVGAGKERGKELKKKFLENTPAIAALRESI QQTLVESSQWVAGEQQVKWKRRWIKGLDGRKVHVRSPHAALNTLLQS AGALICKLWIIKTEEMLVEKGLKHGWDGDFAYMAWVHDEIQVGCRTE EIAQVVIETAQEAMRWVGDHWNFRCLLDTEGKMGPNWAICH

[0390]

[0391] PmCDAl MMTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQS GTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTL KIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEK TLKRAEKWRSELSIMIQVKILHTTKSPAVS (SEQ ID NO: 27)

[0392] TadA-8e MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEI MALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRWFGVRNSKRGAAGSLMNV LNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 28)

[0393] TadDE MSEVEFSHEYWMRHALTLAKRARDEGEAPVGAVLVLNNRVIGEGWNRRIGLHDPTAHAEI MALRQGGLVMQNSRLIDATLYVTFEPCVMCAGAMINSRIGRWFGVRNSKRGAAGSLMNV LNYPGMNVIGLADECAALCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 29)

[0394] Exemplary 24-amino acid linker

[0395] SGGGSETPGTSESATPESGGSIKG (SEQ ID NO: 30)

[0396] Exemplary 29-amino acid linker

[0397] SGGSSGGSSGGGSETPGTSESATPESGGS (SEQ ID NO: 31)

[0398] Exemplary 33 -amino acid linker

[0399] SGSETPGTSESATPESGGSDYKDDDDKGSLIKG (SEQ ID NO: 32)

[0400] Exemplary uracil glycosylase inhibitor (UGI)

[0401] MTNLSDI IEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTS DAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 33)

[0402] The reference in this specification to any prior publication (or information derived from it), or to any matter which is29 known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.

[0403] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. It is to be understood that the invention includes all such variations and modifications, which fall within the spirit and scope. The invention also includes all of the steps, features, compositions and compounds referred to or indicated in this specification, individually or collectively, and any and all combinations of any two or more of said steps or features.

[0404] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this invention belongs. Certain embodiments of the invention will now be described with reference to the following examples which are intended for the purpose of illustration only and are not intended to limit the scope of the generality hereinbefore described.

[0405] EXAMPLES

[0406] Methods

[0407] General methods

[0408] E. coli MG1655 was used as the host strain for LySE. Phage T7 was obtained from ATCC (BAA-1025-B2). Cloning of all plasmids was carried out using MDS42 cells (Scarab Genomics, E-6265-05K). Plasmids were constructed with NEBuilder HiFi DNA Assembly (New England Biolabs) unless otherwise stated. Native E. coli and T7 genes were amplified by PCR directly from genomic DNA. Other genes were synthesised as gBlocks Gene Fragments (Integrated DNA Technologies), unless otherwise stated. PCR reactions were performed using PrimeSTAR GXL DNA Polymerase (Takara) for cloning and using Rapid Taq DNA polymerase (Vazyme) for genotyping. All oligos were synthesised by Integrated DNA Technologies. Sanger sequencing was performed by 1st BASE, Singapore. Nanopore sequencing was performed by Plasmidsaurus, CA, U. S. A. Sample preparation for sequencing was done according to each companies’ protocol.

[0409] Culture media

[0410] E. coli was grown in LB medium or on LB agar (Bio Basic Asia Pacific, Singapore) added kanamycin (50 pg / mL), ampicillin (lOO pg / mL), streptomycin (lOO pg / mL), chloramphenicol (20 pg / mL), tetracycline (10 pg / mL), or hygromycin (200 pg / mL) where appropriate. Ethylene glycol (EG) selection and growth experiments were carried out without antibiotics in standard M9 minimal media (50 mM Na2HPO4, 20 mM KH2PO4, 1 mM NaCl, 20 mM NH4C1, 2 mM MgSO4and 100 pM CaCh, 134 pM EDTA, 13 pM FeC13-6H2O, 6.2 pM ZnCh, 0.76 pM CuC12-2H2O, 0.42 pM CoC12-2H2O, 1.62 pM H3BO3, 0.081 µM MnCl₂·4H₂O). Carbon sources were used as indicated in the text.

[0411] Electrocompetent cell prepar ation

[0412] E. coli cells were inoculated in 10 mL LB and grown overnight at 37 °C and shaking at 225 rpm, with appropriate antibiotics if applicable. The next morning, cells were diluted 1:50 fold in the same media to a final volume of 500 mL, grown to OD₆₀₀ = 0.5, and incubated on ice for 20 min. The cells were spun down at 4000 g for 10 min at 4 °C. The supernatant was discarded, and the cells resuspended in chilled 50 mL sterile dH₂O. The cells were further washed once again with sterile dH₂O followed by chilled 16% (w / v) glycerol. The cells were then resuspended in 1 mL chilled 16% glycerol and used immediately for transformation by electroporation, or flash frozen and stored at -80°C.

[0413] Scarless gene deletion by homologous recombination

[0414] Plasmid pPBGOl was created for arabinose-inducible lambdaRED recombination by cloning genes gam, beta, and exo from E. coli C321. AA.opt (Addgene #87359) into a CloDF13 vector (Addgene #69669). E. coli MG1655, J” / K43Rwas used for all gene deletions. It was created by PCR amplification of rpsLK43R from E. coli DHIOb, which wc electroporated into competent and induced MG1655 cells harbouring pPBGOl. After recovery, the cells were plated on LB agar with streptomycin to select for successful recombination. We further created a double-selection cassette of negative selection gene rpsL (from MG1655) and positive selection gene hygR (from Addgene #104405) combined by overhang extension PCR. For gene deletions, we amplified the rpsL-hygR cassette by PCR using primers containing homology to the regions flanking the target gene to be deleted. We electroporated the the PCR product into competent and induced MG1655, JM / K43Rharbouring pPBGOl and selected for genomic integration with hygromycin. Successful gene deletions were verified by genotyping. To achieve scarless gene deletion, the rpsL-hygR cassette was subsequently removed by a second round of recombineering, by electroporating 10 pg of a 90-base oligo covering the genomic homology regions targeting the lagging strand. Loss of the rpsL-hygR cassette was selected for with streptomycin and verified by genotyping and Sanger sequencing.

[0415] Phage plaque assay

[0416] Phage from clarified phage lysate (or from rebooting in cell-free TXTL extract, see below) were serially diluted ten-fold in LB medium. Overnight cultures of E. coli host strains were prepared, diluted 4-fold in LB, and incubated with shaking for 1 hour at 37°C. For each 100 mm petri dish, 100 pL of cells were mixed with 1 pL of serially diluted phage and incubated for 5 mins at 37°C. The phage-cell mixtures were then mixed with 4 mL pre-warmed molten 0.35% top LB agar and immediately poured uniformly on to 100 mm petri dishes containing 15 mL of solidified 1.5% bottom LB agar. The top agar was left to solidify for 1 hour, then incubated for 4 hours at 37°C for plaque formation.

[0417] Phage titre determination

[0418] Overnight cultures of E. coli were diluted and embedded in top agar as above, but without phage. Plates were dried for 1 hour. Each phage 'as serially diluted and 5 pL were spotted onto the top agar and left to dry for 15 min. The plates were incubated at 37°C, facing down, for 4 h. The number of plaques was counted at each dilution to determine plaque forming units (PFU). The MOI was determined by taking the ratio of PFU of phage added to the colony forming units (CFU) of cells used for phage infection. We determined empirically that on average, approximate phage lysates are 1 xlO1’1PFU / mL and diluted overnight cell cultures are 1×1011CFU / mL.

[0419] T7ADNAP phage assembly and rebooting in TXTL

[0420] T7 genomic DNA excluding T7 DNAP was PCR amplified in five fragments of 10.0 kb (Fl), 4.4 kb (F2), 3.5 kb (F4), 10.0 kb (F5) and 10.0 kb (F6), with 25-30 bp overlapping sequences. Fragment F3 containing the trxA gene was amplified from the E. coli genome, with overlapping sequences with F2 and F4 to replace T7 DNAP. Purified PCR fragments were mixed, including 300 ng of F3, and 100 ng of each of the other fragments. They were then assembled by Gibson assembly using NEBuilder HiFi DNA Assembly (New England Biolabs). The assembly mix, together with 50 ng of T7 genomic DNA, was mixed with myTXTL Linear DNA Expression Kit (Arbor Biosciences) according to the manufacturer’s protocol and incubated overnight at 37°C to generate a mixture of rebooted wild type T7 and T7ADNAP phage. To select for the mutant phage, a plaque assay was performed by infecting MG1655AtrxA containing pSJ55, which expresses wild type T7 DNAP. Wild type T7 not expressing trxA would not propagate in this strain. Only T7ADNAP trxA would be able to propagate in this strain and form plaques. We confirmed successful generation of T7ADNAP trxA by genotyping and Sanger sequencing the genomic region where T7 DNAP was deleted from.

[0421] Phage infection kinetic assay

[0422] Infection kinetics were carried out in 96 well plates in BioTek Synergy HTX Microplate Reader (Agilent). Overnight cultures of E. coli host strains were prepared, diluted by half w'ith LB medium, and 100 pL added to each well. In each well, 20 uL of different serial phage dilutions were added to tune MOI. Each condition was replicated in three different wells. A lid was added to the 96 well plates to reduce evaporation during acquisition. The microplate reader was set to 37°C with continuous orbital shaking at 300 rpm. ODeoo was measured every 3 mins and monitored for at least 2 hours.

[0423] Yeast assembly

[0424] Large phagemids were assembled by transformation-associated recombination in yeast as previously described (Robertson, W. E. et al. Nature Protocols 16, 2345-2380 (2021)).. S'. cerevisiae BY4741 (MATa his3Al leu2A0 metl5A0 ura3A0) was obtained from ATCC (#201388). The cell-wall digestion step was modified slightly: We only used 1 pL zymolyase, and we measured OD660 every 10 min after 30 min zymolyase digestion. Digested cells were resuspended by swirling and inverting the tubes.

[0425] Phagemid assembly was assessed by genotyping of the resulting yeast colonies, amplifying 350-1000 bp across the intersections of assembled fragments. Positive yeast colonies were cultured, phagemids were isolated with Zymoprep Yeast Plasmid Miniprep Kit (Zymo Research, USA) and electroporated into £'. coll MDS42 with selection. Individual colonies were genotyped again with the same primers, and phagemids were isolated from positive clones using Monarch Plasmid Miniprcp Kit and verified by Nanoporc sequencing.

[0426] General LySE strain preparation

[0427] All LySE experiments were conducted with E. coli MG1655. Cells were first transformed with accessory plasmids (APs) containing different T7 DNAP variants by electroporation as described above and selected with ampicillin. Cells containing the AP are then transformed a second time with the phagemid (PM) by electroporation and selected with both ampicillin and kanamycin. Genotyping was performed at each step to confirm successful transformation.

[0428] LySE cycle

[0429] To propagate T7ADNAP phage, E. coli MG1655 with pSJ55 was grown overnight in LB medium with ampicillin, then diluted with equal volumes of LB the next day. T7ADNAP phage lysates were mixed with the diluted cells at a volume ratio of 1: 1000 (approximately MOI = 0.1) and incubated at 37°C with shaking for 2 hours until there was no further reduction in ODeoo. The phage lysates were then washed once with equal volumes of chloroform to remove residual cells and debris.

[0430] The LySE cycle begins with an overnight culture of MG1655 cells containing AP and PM. Overnight culture cells were diluted with equal volumes of LB the next day. To 100 pL of diluted cells, 20 µL of phage is added (approximately MOI = 10; High MOI) and incubated at 37°C with shaking for 2 hours until there was no further reduction in OD₆₀₀. The phage lysates were washed once with equal volumes of chloroform. Next, to 1 mL overnight culture of MG1655 cells containing only AP diluted with equal volumes of LB, we mixed 10 pL phage lysates containing phagemids and incubated at 37°C with shaking for 1 hour for complete transduction of phagemids (approximately MOI = 1; Low MOI). Phagemid packaging efficiency was determined by serial dilution of the transduced cells, followed by spot plating on LB agar with kanamycin to determine CFU of transduced phagemid. To continue the LySE cycle, transduced cells were diluted in selection media and grown to confluency. The exact protocol for selection and culture recover}' is specific to each evolution experiment, but should include addition of ampicillin to maintain the AP. The next LySE cycle is continued by adding phage at high MOI to lyse cultured cells.

[0431] Fluctuation assay

[0432] We performed Luria-Delbriick fluctuation analysis to quantify the mutation rate per generation of LySE for each T7 DNAP variant. We cloned pSJ51, a phagemid constitutively expressing a chloramphenicol resistance gene (CmR) containing a premature stop codon (Q38TAG). MG1655 containing pSJ51 and an AP encoding the T7 DNAP variant to be tested was grown overnight, diluted and lysed as per standard LySE protocols. After one generation of phagemid replication and transduction, transduced cells were serially diluted ten-fold and spotted on LB agar with kanamycin to quantify total transduction CFU, and on LB agar with chloramphenicol to quantify stop codon reversion CFU respectively. After overnight incubation at 37°C, CFUs from three independent replicates were counted. Mutation per generation per base pair p (s.p.b.) was calculated using p (s.p.b.) = m / (R x C). Mutation frequency (m) was calculated based on the ratio of cells grown on chloramphenicol to that grown on kanamycin. For the parameter R, which is the number of distinct mutation sites that make the resistance gene effective, we found that 8 / 9 possible single base substitutions (which yield sense codons) can result in chloramphenicol resistance. So, R = 8 / 3. C is the gene copy number. Since T7 packages concatemeric phagemid DNA via a head-full mechanism, multiple CmR copies are expected per phage, while only a single reverted TAG is sufficient to confer chloramphenicol resistance. We therefore determined C by taking the fraction of phagemid to T7 genome size, so C = 39937 / 3469 = 11.5.

[0433] LacZ inactivation assay

[0434] LacZ inactivation assay was performed to quantify the mutation rates of more error-prone T7 DNAP variants. We performed scarless deletion of lacIZYA from E coli MG1655 as described above. We cloned pSJ78, a phagemid constitutively expressing wild type lacZ. MG1655∆lacIZYA containing pSJ78 and an AP encoding the T7 DNAP variant to be tested was grown overnight, diluted and lysed as per standard LySE protocols. After one generation of phagemid replication and transduction, transduced cells were serially diluted ten-fold and spotted on LB agar with kanamycin and 200 pg / mL X-Gal (Thermo Scientific). After overnight incubation at 37°C, CFUs from three independent replicates were counted. The fraction of white or light blue colonies (lacZ- phenotype) was counted as a function of all colonies (blue+light blue+white) and used as a measure of mutation frequency for the lacZ cassette.

[0435] Molecular dynamics simulation

[0436] We used the crystal structure published by Doublie et al. (PDB: 1T7P) as the basis for our molecular dynamics simulations of the wild type and T523R mutant T7 DNAP. 1T7P contains a growing DNA strand terminated with a dideoxy cytosine nucleotide, and the incoming nucleotide is dideoxyguanosine triphosphate (ddGTP). To represent the real biomolecules as closely as possible, we manually added 3’-hydroxyl groups to the chainterminating nucleotide of the growing DNA strand and ddGTP. A wild type and mutant T523R T7 DNAP variants were created in silico, and subsets of these containing DNA substitutions of the leading cytosine nucleotide on the template strand were implemented to study the effect T523R has on base mispairing. The structure is dissolved in water under physiological conditions using the solution builder on CHARMM-GUI. We placed each protein structure in a 130A x 130A x 130A simulation box with water containing 215 / 216 sodium ions and 183 chloride ions (150 mM) to balance protein charges at pH 7.0. For each condition, we ran a 5000 step steepest descent energy minimisation. This was followed by an NVT equilibration with a simulation time of 125 ps (125,000 steps) at 303.15 K. The subsequent NPT production simulation is 10 ns (5,000,000 steps) at 303.15 K. All simulations used the CHARMM36m force field and were run using CUDA supported GROMACS, version 2023.3, on high performance computing facilities (National University of Singapore HPC).

[0437] Genomic fluctuation assay

[0438] E. coli MG1655 was grown from glycerol stocks overnight in LB. To determine total CFU, the cells were serially diluted ten-fold and spotted on LB agar. To determine frequency of rifampicin resistance, 2 mL of cells were spun down and the cell pellet plated onto selective LB agar with 50 pg / mL rifampicin. After overnight incubation at 37°C, CFUs from three independent replicates were counted. Mutation frequency was calculated based on the ratio of cells grown on rifampicin to that grown on LB. To calculate substitutions per base pair (s.p.b.), the mutation rate was normalized by the number of mutations in the rpoB gene that impart rifampicin resistance (77 known point mutations, divide observed mutation rate by 77 / 3)31. We determined E. coli genomic mutation rates to be 2.39 ± 1.10 × 10-10s.p.b, comparable to those previously reported.

[0439] Illumina NGS and data analysis

[0440] We cloned pSJ77, a 39 kb BAC-phagemid by yeast assembly. E. coli MG1655 containing pSJ77 and an AP encoding either T7 DNAP variant v8 or v9 was grown overnight, diluted and lysed as per standard LySE protocols. After one generation of phagemid replication and transduction into MG1655 with no plasmids, transduced cells were diluted 10-fold in LB with kanamycin and recovered overnight. The recovered mutated BAC-phagemid library was purified by ZymoPURE Plasmid Miniprep Kit (Zymo Research, USA). Illumina Next-Generation Sequencing (NGS) of the BAC-phagemid library was performed by Bio Basic Asia Pacific. Next generation sequencing library preparations were constructed following the manufacturer’s protocol. For each sample, 200 pg DNA was randomly fragmented by Covaris to an average size of 300-350 bp. The fragments were treated with End Prep Enzyme Mix for end repairing, 5’ Phosphorylation and 3’ adenylated, to add adaptors to both ends. Size selection of Adaptor-ligated DNA was then performed by DNA Cleanup beads. Each sample was then amplified by PCR for 8 cycles using P5 and P7 primers, with both primers carrying sequences that can anneal with flowcell to perform bridge PCR and P7 primer carrying a six-base index allowing for multiplexing. The PCR products were cleaned up and validated using an Agilent 2100 Bioanalyzer. The qualified libraries were sequenced pair end PE150 on the Illumina Novaseq System. Fastp (v0.23.0) was used for quality control and preprocessing, including removal of adaptor sequences, PCR primers, reads with more than 14 N bases, and reads with less than 40% bases above Q20. The cleaned data was then mapped to the reference genome using the Sentieon pipeline (v202112.02). A custom python script using the pysam (v0.23.0) module was used to align the NGS reads with Q score >30 to the wild type sequence and count the nucleotide positions from which the experimental sample deviates from the wild type sequence. General mismatch rates and A:T>G:C and C:G>T:A mutation rates per 5 bp were calculated and plotted. Overall mutational spectra, and for every 10,000 bp, were calculated. We yielded an average of > 14,000 reads per position for each of the sequenced samples.

[0441] LySE evolution of tetA

[0442] Wc cloned pSJ78, a phagemid constitutively expressing the tetA tetracycline efflux gene. E. coli MG1655 containing pSJ78 and pSJ139 (T7 DNAP v9) was grown overnight, diluted and lysed as per standard LySE protocols. After one generation of phagemid replication and transduction into MG1655 with pSJ139, transduced cells were diluted 10-fold with LB with ampicillin, kanamycin, and 0.1 pg / mL tigecycline. The cells were incubated overnight at 37°C with shaking at 225 rpm. The next day, the cells were diluted with equal volumes of LB, and T7ADNAP phages were added to lyse the culture, starting another round of LySE. Evolution cycles were repeated another four times for a total of five evolution cycles, using the best growing cultures from the previous cycle as the starting material for the next cycle. Tigecycline concentrations were incrementally increased from LySE El to E5, with concentrations 0.1 pg / mL, 0.25 pg / mL, 0.5 pg / mL, 0.75 pg / mL and 1 pg / mL respectively. After the 5thLySE cycle, the best growing culture was streaked on LB agar with kanamycin and 1 pg / mL tigecycline. Individual colonies were picked and inoculated separately in 1 mL LB with kanamycin in 24-well plates and grown overnight. 32 evolved clones were spotted on LB agar with increasing concentrations of tigecycline to test resistance. The 32 clones were PCR amplified for the tetA cassette using PrimeSTAR GXL DNA Polymerase and the amplicon was sequenced by Sanger sequencing.

[0443] For ALE, E. coli MG1655 containing pSJ78 and pSJ55 (WT T7 DNAP) was grown overnight in LB with kanamycin and ampicillin, then diluted 10-fold with LB with kanamycin, ampicillin, and 0.1 pg / mL tigecycline. The cells were incubated overnight at 37 °C with shaking at 225 rpm. The next day, the cells were diluted another 10-fold to continue ALE. A total of 5 passages were performed, each time with increasing concentrations of tigecycline identical to the LySE schedule. Cells at the 5thpassage were spotted on LB agar with increasing concentrations of tigecycline to test resistance. The cells were also lysed by addition of T7ADNAP, the lysate washed with chloroform, and phagemids transduced to fresh MG1655 with no plasmids. The transduced cells were diluted 10-fold with LB with kanamycin, then grown overnight. The recovered E5T ALE cells were then spotted on LB agar with tigecycline to test resistance after transduction.

[0444] Quantitative real-time PCR (qPCR)

[0445] Total RNA was first isolated from the E. coli cells. Cells were grown overnight and diluted 1:50 with LB and appropriate antibiotics and cultured until ODeoo = 0.5. Then, 500 pL of cells were transferred into an Eppendorf tube, spun down, and the pellet dried by dabbing on a paper towel. The pellet was resuspended in 100 pL of Tris-EDTA buffer pH 8.5 with 15 mg / mL lysozyme, vortex for 10 s and incubated at room temperature for 5 mins with shaking. To the mixture, 400 pL of TRK Lysis buffer (Omega Biotek) with 4 pL 2-mercaptoethanol was added and total RNA was extracted immediately with E. Z. N. A RNA isolation kit (Omega Biotek). Purified RNA was quantified with NanoDrop (Thermo Fisher Scientific). One microgram of RNA was converted to cDNA using the GoScript™ Reverse Transcriptase (Promega). qPCR was performed with GoTaq qPCR (Promega) using the CFX Opus 96 Real-Time PCR System (Bio Rad). Fold changes were normalized to 16s rRNA and are based on relative expression values calculated using the 2−ΔΔCT method.

[0446] LySE evolution of EG assimilation pathway

[0447] The gox0313 gene (UniProt: Q5FU50) was synthesised by GentleGen, China. We cloned pAN29, a phagemid containing metabolic pathway genes for complete assimilation of EG. E. coli MG1655 containing pAN29 and pSJ139 (T7 DNAP v9) was grown overnight, diluted and lysed as per standard LySE protocols. After one generation of phagemid replication and transduction into MG1655 with pSJ139, transduced cells were diluted 10-fold with LB with ampicillin and kanamycin and incubated overnight at 37°C with shaking at 225 rpm. The next day, 5 mL of recovered cells were pelleted by centrifugation at 3900 x g for 3 mins, washed three times with M9 medium, and then resuspended in M9 medium with EG, with or without glucose at concentrations as indicated in the text until the OD600 is approximately 0.2. No antibiotics were used for selection. To each well in a 24-well plate, 1 mL of the cell suspension was added and cultured overnight at 37°C with shaking at 300 rpm in a ThermoShaker PST-60HL-4 Microplate Reader (BioSan, Latvia). Every hour, ODeoo was measured to monitor cell growth. Evolution cycles were repeated another four times for a total of five evolution cycles, using the best growing cultures from the previous cycle as the starting material for the next cycle. After the 5thLySE cycle, the best growing culture was streaked on M9 minimal agar plates supplemented with 2 g / L of glucose and 10 g / L of EG. Individual colonies were picked and inoculated separately in 1 mL LB with kanamycin in 24-well plates and grown overnight. The following day, the cultures were washed with M9 media at 3900 x g for 1 min. Subsequently, the cells were inoculated in M9 medium supplemented with 10 g / L EG in 12 mL culture tubes for 6 hours to reach OD600=0.1. 200 µL of each culture was transferred into a sterile 96-well microplate as triplicates and incubated overnight with continuous orbital shaking at 600 rpm. Growth rates were determined by measuring ODeoo at 12, 18. 24, 36, 42, and 48 hours post-inoculation using the microplate reader. Eight clones were PCR amplified for the phagemid using PrimeSTAR GXL DNA Polymerase and the amplicon was sequenced by nanopore.

[0448] For ALE, 5 mL E. coli MG1655 containing pAN29 was grown overnight in LB with kanamycin, washed three times with M9 medium, and then resuspended in M9 medium with EG, with or without glucose at concentrations as indicated in the text until the OD600 is approximately 0.2. To each well in a 24-well plate, 1 mL of the cell suspension was added and cultured overnight. The next day, 1 µL of the cells were added to 10 mL of LB with kanamycin and grown overnight before 5 mL was taken and washed, following the same washing, resuspension procedure as above for the next round of ALE. Evolution cycles were repeated another four times for a total of five evolution cycles, using the best growing cultures from the previous cycle as the starting material for the next cycle.

[0449] Example 1: The LYtic Selection and Evolution system (LYSE)

[0450] Bacteriophage T7 is a lytic phage that infects E. coli, replicates its 40 kb genome, and lyses the host cell within 17 minutes, releasing approximately 180 progeny phages (Fig. 6a). We exploited the ability of the T7 phage for rapid multiplication and distribution of large genetic material to develop a system for continuous, accelerated evolution of large Gene Clusters of Interest (GCO1). First, we engineered a T7 phage variant lacking T7 DNAP, by in vitro assembly of the complete genome except gp5, and rebooting the phage in a cell-free extract from E. coli. This phage, T7ADNAP, efficiently propagates only in hosts that carry an accessory plasmid (AP) expressing the T7 DNAP under a T7 promoter (Fig. 6c, d). The absence of phage propagation without accessory plasmid demonstrates the strict biocontainment of the system.

[0451] Phage T7 can replicate, package, and transduce phagemids (PM); circular plasmids containing the T7 origin of replication, a T7 terminal repeat, and a host origin of choice. Importantly, the phagemid is replicated by the host replication machinery during cell cycle, and only by the phage replication machinery during phage infection. We constructed a phagemid (PM) with a pl5a host origin (Fig. 13) and demonstrated that it was efficiently packaged and transduced by phage T7ADNAP in cells containing accessory plasmid (AP) expressing wildtype or error-prone T7 DNAP (Fig. la, 6e). T7 DNAP expression from the accessory plasmid (AP) is tightly regulated under a T7 promoter during the cell cycle. Upon T7 phage infection, the T7 RNA polymerase (T7 RNAP) induces expression of error-prone T7 DNAP, which replicates the phagemid (PM) for subsequent packaging and transduction.

[0452] When equipped with an error-prone T7 DNAP variant, this system will enable cyclic evolution, where the phagemid undergoes error-prone replication and packaging during a lytic phase (high Multiplicity Of Infection, MOI), followed by transduction to fresh hosts for growth and selection during a cellular phase (low MOI) (Fig. lb). At high multiplicity of infection (MOI), the system enters a lytic phase where the gene construct of interest (GOI), carried on a phagemid, undergoes error-prone replication by error-prone T7 DNAP. Mutated variants are packaged into T7 phage particles and released. During subsequent transduction at low MOI, these mutated phagemids are introduced into new host cells, enabling fitness-based selection through cell growth. The evolved gene cluster pool is then cycled back to the lytic phase by reintroducing phage T7ADNAP.

[0453] During the cellular phase, T7 DNAP expression will cease and host DNAPIII maintains the phagemid with high fidelity. The selection of improved GCOI variants is achieved by linking phagemid-encoded functions to host cell fitness during the cellular phase. The short lysis time and large burst size of T7 facilitate quick turnover of large genetic pools, thus reducing the time for directed evolution from days to hours (Fig. 1c). The workflow requires only simple mixing of phage lysates and cell cultures without specialized equipment, making LySE readily portable and accessible to most laboratories regardless of experience. Utilising T7 as an efficient gene shuttle between cells, LySE creates a hybrid system between continuous and discrete evolution by seamlessly transitioning between mutagenesis and selection while preserving discrete checkpoints (Fig. Id).

[0454] Example 2: Engineering hypermutagenic T7 DNA polymerase variants

[0455] Several error-prone T7 DNAP variants were engineered and tested to drive targeted mutagenesis of phagemid-encoded GCOIs (see Table 1 for the variants). The design strategy incorporated complementary error-inducing mechanisms for additive effects. We began with a previously characterised error-prone variant, T7 DNAP vl, containing mutations in both the thumb domain (S399T) and exonuclease domain (Y64C / F120L). S399T has been proposed to increase error rates by disturbing the nucleic acid binding cleft, while Y64C and F120L likely reduce proof-reading activity. This variant exhibited a 13.5-fold increased error rate compared to wild type T7 DNAP (Fig. 2a), in line with previous reports.

[0456] Mutation rates were quantified by fluctuation analysis using a chloramphenicol resistance gene (CmR) containing a premature stop codon (Q38TAG). After one generation of phagemid replication and transduction, we measured the frequency of cells acquiring chloramphenicol resistance through point mutations converting the TAG to sense codons. We calculated apparent mutation rates in substitutions per base pair per generation (s.p.b) by correcting for multiple phagemid copies per phage particle.

[0457] To further enhance mutagenesis, we employed homology-guided design based on fidelity-reduced E. coll DNAP I variants. We identified several potential mutations in the fingers domain of T7 DNAP (L479N / H506Y / T523R / P560H), of which some are positioned near the polymerase active site. Through systematic experimental testing, we determined T523R to significantly increase error rates, and when combined with vl (v2.4: Y64C / F120L / S399T / T523R) achieved error rates 157-fold higher than wild-type (wt). Molecular dynamics simulations of the T7 DNAP crystal structure revealed that T523R significantly alters nucleotide positioning in the active site (Fig 2b, 7). An incoming triphosphate nucleotide is pulled deeper into the active site in presence of the arginine substitution. Thus, we propose that the arginine stabilises incorrect base-pairing.

[0458] Complete inactivation of the exonuclease function through D5A / E7A mutations further reduced replication fidelity. Incorporating D5A / E7A into T7 DNAP v2.4 generated variant v4 with a 2-fold increased error rate (Fig 2a). Despite successfully engineering multiple error-prone variants, we observed a ceiling at 4.27 x 10-5s.p.b in the fluctuation assay, whereby increased mutation rates were accompanied by declining phagemid transduction efficiency (Fig 8a). We hypothesised that higher mutation rates might be obscured if inactivating mutations elsewhere in the CmR gene counteracted TAG reversion effects.

[0459] We therefore employed a ZacZ-inactivation assay capable of detecting higher mutation frequencies (Fig.2c). We encoded lacZ on the phagemid, yielding blue colonies in presence of X-Gal. At high error rates, lacZ disruption produces white or light blue colonies. Using variants v2.4 and v4 as benchmarks, we observed 0.86% and 2.96% lacZ inactivation rates respectively — a 3.4-fold difference consistent with the fluctuation assay results.

[0460] In parallel we explored error-prone replication by fusing T7 DNAP to DNA deaminases for concurrent deamination during replication. We fused the adenosine deaminase TadA-8e to the N-terminus of wild type T7 DNAP with linkers of varying lengths (v5.1-5.5, Table 1).

[0461] A 24-residue linker (v5.3) resulted in the highest error rate of 1.79 x 10-5, comparable to the rate of v2.4 (Fig 2a). The cytosine deaminase PmCDAl fusion (v6.1) proved less effective than TadA-8e. A uracil glycosylase inhibitor (UGI) was co-expressed with the polymerasedeaminase fusion protein to inactivate uracil-DNA glycosylase (UNG), which removes uracil nucleotides introduced by the cytosine deamination activity of PmCDAl. The UGI gene was added downstream of the gene for the polymerase fusion in the same operon (controlled by the T7 promoter).

[0462] We then combined the two complementary approaches by fusing TadA-8e to the most error-prone T7 DNAP variants with mutations in the thumb, fingers and exonuclease domain (v2.4) as well as exonuclease inactivation (v4), creating v7 and v8, respectively. These fusion variants resulted in 10-fold and 40-fold increased lacZ inactivation compared to the mutant variants (Fig. 2c). T7 DNAP v8 achieved error rates of 9.40 x 10-4s.p.b — nearly one substitution per one thousand base pairs, and 4 million times higher than the determined genomic mutation rate in E. coli (2.39 ± 1.10 x 10-10s.p.b). Variant v8 had negligible impact on host growth kinetics (Fig.8b), demonstrating tight control of error-prone replication. The final engineered T7 DNAP v8 comprises a hypermutagenic T7 DNAP suitable for mutagenizing and diversifying phagemid libraries for continuous evolution with LySE (Fig 2d). We next sought to broaden the mutagenesis spectrum of the error-prone T7 DNAP to introduce all transition mutations at comparable frequencies. Variant v8 contains TadA-8e, which catalyses adenosine-to-inosine deamination, resulting in A: T— > G: C transitions during DNA replication (Fig. 4a). Illumina next-generation sequencing (NGS) confirmed that transition mutation frequencies were skewed toward A: T— > G: C transitions at a ratio of 1:0.8 (Fig. 4b). To address this bias, we installed a recently engineered dual adenine-cytosine deaminase, TadDE (v9), capable of performing both adenosine and cytidine deamination, facilitating A: T— > G: C and C: G— > T: A transitions, respectively (Fig. 4a). This modification improved the transition ratio to 1:0.92, demonstrating that v9 introduces all transition mutations at nearly equivalent frequencies (Fig. 4b) while maintaining a high overall error rate (Fig. 9a).

[0463] To probe the maximum capacity of LySE to evolve large GCOIs, we assembled a 39 kb T7 phagemid with a bacterial artificial chromosome (BAC) origin by homologous recombination in S. cerevisiae. After performing one generation of LySE with this BAC-phagemid, nanopore sequencing of transductants confirmed that the intact vector was successfully transferred into fresh host cells (Fig. 9b). NGS analysis revealed a uniform distribution of mutations and base transitions throughout the entire 39 kb BAC -phagemid (Fig. 4c). LySE exhibited a broad mutagenesis spectrum with all transitions occurring at similar frequencies, while maintaining the biologically relevant preference for transitions over transversions (Fig.4d). T aken together, LySE presents a significant improvement over existing continuous evolution tools by enabling the potential evolution of large gene clusters up to 40 kb, equivalent to the size of the complete T7 genome.

[0464] Example 3: Tuning of multiplicity of infection (MOI) for controlled lysis

[0465] We assessed how the error-prone T7 DNAP variants affected T7 phage replication. E. coli cultures harbouring T7 DNAP variants were infected with T7 ADNAP at varying MOIs while we monitored cell density over time (Fig. 3a). Wild type T7 DNAP caused efficient cell lysis independent of the initial MOI, due to rapid phage propagation. However, as the error rate of replication increased (v2.4-v8), the resulting lysis became highly MOI-dependent. In cultures with T7 DNAP v8, the cell density increased during infection at low MOIs, while it decreased at high MOI (Fig. 3b). These findings demonstrate that lysis by T7 phage — a strictly lytic phage in nature — can be tuned by multiplicity of infection under error-prone replication.

[0466] During faithful replication with wild type T7 DNAP, the phage propagates and lyses the cells exponentially, independent of MOIs (Fig, 3c). Conversely, error-prone variants like v8 compromise replication fidelity, resulting in unsustainable propagation and MOI-dependent lysis. We verified that this compromised propagation still permits substantial phagemid packaging and transduction even with highly error-prone DNAP v8 (Fig. 3d).

[0467] The ability to tune the degree of culture lysis by adjusting T7 phage multiplicity enables controlled and distinct phases of the LySE cycle. At high MOI, mutated phagemids are packaged into virions and released through efficient cell lysis (Fig. 3e). Upon lowering the MOT by infecting a fresh culture, phagemid libraries get transduced to fresh cells without widespread lysis. This library of transductants can be grown under selective conditions to enrich for improved GCOI performance before initiating another round of lysis. These distinct LySE phases allow adjustment of growth and selection duration, facilitating evolution of slow-manifesting phenotypes such as low-rate metabolic pathways or slow-folding protein complexes.

[0468] Example 4: Accelerated evolution of genes and gene clusters

[0469] To demonstrate the utility of LySE for accelerated gene evolution, we first evolved the tetracycline resistance gene lelA to confer resistance to tigecycline. We expressed lelA from a phagemid and subjected it to 5 generations of LySE with progressively increasing tigecycline concentrations (Fig.5a, 14). For comparison, we performed adaptive laboratory evolution (ALE) under identical selection conditions. Following LySE evolution, we obtained cells tolerating 2.5 pg / mL of tigecycline, a 25-fold increase over the wild type tolerance of 0.1 pg / mL (E5 LySE, Fig.5b). In contrast, cells evolved by ALE tolerated only up to 1 pg / mL (E5 ALE). Furthermore, when the ALE-evolved phagemid was transferred to fresh cells, the acquired resistance was lost (E5T ALE, Fig. 5b). This indicates that the evolved resistance was caused by genomic mutations rather than mutations in the target phagemid, and illustrates that ALE is highly sensitive to off-target and cheater mutations. LySE overcomes these issues by refreshing the host in each cycle, thus only allowing evolution through on-target mutations of the phagemid. From evolved clones of E5 LySE, we identified four point mutations in let A (Fig.5d, 10a, b), and a mutation in the promoter that increased tetA expression approximately 200-fold (Fig.

[0470] 10c). This demonstrates the ability of LySE to simultaneously evolve both regulatory and coding regions, which is particularly useful for optimising complex phenotypes.

[0471] We next explored the capability to evolve a complex process — a multigene cluster constituting a complete metabolic pathway for assimilation of ethylene glycol (EG), a monomer derived from PET degradation. This pathway was chosen due to its significant implications for plastic recycling and its potential contribution towards developing a sustainable circular' economy for plastic waste management.

[0472] We constructed the pathway starting with EG reduction to glycol aldehyde using Gluconobacter oxydans alcohol dehydrogenase (Gox0313). selected for its superior performance and usage of only NAD+ as cofactor (Fig. 5e). Glycolaldehyde is further reduced to glycolate and then to glyoxylate by endogenous aldehyde dehydrogenase (aldA) and glycolate oxidase complex ( glcDEF) from E. coli. The resulting glyoxylate enters the glycerate pathway for biomass and energy production (Fig. 5e). All five genes were cloned into a T7 phagemid, creating a 9,715 bp plasmid (Fig. 5f, 15). The phagemid was subjected to 5 generations of evolution by LySE or ALE under the same selection regime. We implemented a semi-relaxed selection protocol with gradual glucose withdrawal from the culture medium, resulting in improved normalised growth rates for LySE-evolved cells compared to ALE-evolved cells (Fig. 5g, 11).

[0473] Sequencing of clones evolved by LySE revealed mutations in the T7 ori / pac site, the pl5a ori, near the aldA promoter, and nonsynonymous mutations in three of the five genes (Fig.

[0474] 5f). Five clones exhibited accelerated growth on EG compared to wild type, featuring a novel Y94F substitution in Gox0313 (EGA3), and I312F substitution in glcD (EGA7), with improved endpoint biomass by 50.9% and 46.1 % respectively after just 5 generations (Fig.

[0475] 5h, 12). These experiments demonstrate LySE as a powerful directed evolution platform that enables rapid, targeted genetic optimization across both single genes and complex gene clusters. Example 5: Discussion

[0476] LySE is a robust T7 phage-based system that bridges the fundamental trade-off between controllability and speed in directed evolution. By leveraging T7 as an efficient gene shuttle between cells, LySE transfers DNA up to 40 kb without transformation losses while performing hypermutagenesis on target gene clusters. The system preserves discrete checkpoints for mutagenesis and selection yet enables seamless transitions between phases, uniquely combining continuous evolution with controlled discrete cycles.

[0477] Unlike previous phage-based continuous evolution methods such as PACE and T7AE (T7 assisted evolution), which insert GOIs directly into the phage genome (thereby limiting the capacity to less than 5 kb), the LySE approach employs a T7 phagemid with T7 origin and packaging signals to direct the phage to package the GCOI as-is into its capsid. This strategy leverages the full capacity of the T7 capsid to enable evolution of 40 kb constructs, of which 29 kb of GCOI was tested in this w'ork (Fig. 9b). This substantially exceeds the capabilities of existing phage-assisted and in vivo continuous evolution methods. The expanded capacity of LySE significantly broadens potential applications for continuous evolution. It enables work with anabolic pathways for small molecule synthesis, catabolic pathways for waste assimilation (Fig. 5) or carbon capture, and evolution of protein complexes, many of which exceed 10 kb when including regulatory elements. This versatility eliminates size constraints that have previously limited directed evolution approaches.

[0478] While orthogonal replication systems have gained traction for handling larger constructs of up to a 16.5 kb replicon, they remain vulnerable to off- target mutations that can accumulate as genomic hitchhikers, complicating phenotypic attribution. These systems also risk enabling cheater mutations that bypass selection pressure, especially in biosensor-based evolution. These vulnerabilities stem from the fundamental mechanism of orthogonal replication, where the GOI replicates with the host, allowing off-target mutations to be carried over to subsequent generations. Similarly, Pl phage-based methods like IDE, despite accommodating large inserts of up to 36 kb, are known to transfer flanking genomic regions after integration, which presents challenges for precise directed evolution.

[0479] LySE overcomes these limitations through its lytic cycle. T7 phage-mediated cell lysis completely eliminates the E. coli culture after each selection cycle, effectively removing all off-target mutations. This process mirrors the discrete checkpoints of classical directed evolution but operates with the speed and throughput of continuous systems. There remains a risk of off-target mutations within the phagemid itself, such as alterations to the E. coli origin that could increase plasmid copy number instead of evolving the GCOI (Fig.5f). This can be addressed through stringent selection pressures or biosensor-based approaches, particularly when evolving enzyme specificity for synthetic metabolism.

[0480] Furthermore, as an orthogonal DNA polymerase, T7 DNAP can be engineered to achieve extremely high error rates without affecting cell viability. Our implementation of controlled error-prone replication and multiplicity tuning redirects pressure away from maintaining T7 genome fidelity, enabling us to exceed standard T7 phage error thresholds for the first time (Fig. 2a), achieving an estimated 9.40 x 10-4s.p.b – one of the highest error rates in continuous evolution.

[0481] Unlike other phage-assisted evolution methods, LySE employs intracellular selection rather than viral fitness coupling. This enables direct coupling of GCOI function to cellular fitness, more appropriate for applications ultimately deployed in cellular contexts, such as strain or therapeutic cell engineering, while also enabling selections based on physical cell properties through techniques like cell sorting.

[0482] The LySE workflow (lysis-transduction-recovery-selection) requires only mixing phage lysates with cell cultures, making it accessible to users regardless of experience or access to equipment. Though manually performed here, the process is amenable to automation via liquid handling systems for increased throughput. The present results establish a generalisable method for accelerated evolution of whole gene clusters, providing researchers with a robust tool for engineering large metabolic pathways and protein complexes without the complications of off-target mutations or mutations that circumvent selection.

[0483] The LYSE platform has advantages over other directed evolution systems:

[0484] 1. Phage-assisted continuous evolution (PACE): The PACE system uses the M13 phage, which has a limited pay load size of less than 8 kb. The M13 phage also undergoes a lysogenic cycle, not killing the host during infection. This means that it relies on the error rate of the host cell, which cannot go beyond the extinction threshold, limiting the error rate that can be used for mutagenesis. PACE also requires the GOI to be part of the M13 phage genome, which limits the size of the payload and requires additional phage assembly steps. Due to limited error rates, PACE is expected to be slower than LYSE and cannot evolve large gene constructs.

[0485] 2. Inducible Directed Evolution (IDR) (Al’Abri et al. Nucleic Acids Res. 2022 Jun 10;50(10):e58): IDR uses the Pl phage and a Pl phagemid to carry the GOI. Similar to M13, the P1 phage is a temperate phage, and IDR relies on the error rate of the host cell for mutagenesis, limiting the error rate which can be introduced. Due to limited error rates, IDR is expected to be slower than the LYSE system.

[0486] 3. T7 assisted evolution (T7AE) (Serrano et al. Sci Rep. 14, Article number: 2377 (2024)):

[0487] The T7AE system uses the T7 phage to carry the GOI on the T7 genome, which limits the size of the payload to the size of the T7 genome. The GOI is linked directly to the fitness of the T7 phage, thereby limiting the use of T7AE. Further, incorporating the GOI into the T7 genome is more cumbersome as it requires recombineering with selection, thereby increasing the handling time and limiting the number of variants that can be introduced.

[0488] It will be appreciated that many further modifications and permutations of various aspects of the described embodiments are possible. Accordingly, the described aspects are intended to embrace all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims.

Claims

CLAIMS1. A method for directed molecular evolution, the method comprising:a) introducing a propagation-defective phage vector into a first host cell (mutagenic host) competent to propagate the phage vector,wherein the first host cell comprises a gene construct of interest (GOI) to be evolved;and wherein the phage vector allows for (i) expression of an error-prone DNA polymerase in the first host cell for mutagenesis, and (ii) replication and packaging of the GOI into infectious phage particles;b) incubating the first host cell under conditions for replication and mutagenesis of the GOI and release of phage particles comprising mutated GOIs;c) infecting a second host cell (selection host) with phage particles from step b); and d) selecting for a desirable function of the GOI in the second host cell.

2. The method of claim 1, wherein the phage vector is deficient in a DNA polymerase required for phage propagation, and wherein the error-prone DNA polymerase allows for propagation of the phage vector.

3. The method of claim 1 or 2, wherein the phage vector is derived from a lytic phage.

4. The method of claim 3, wherein the lytic phage is T2, T4, T5, T6 or T7 phage.

5. The method of claim 4, wherein the lytic phage is T7 phage.

6. The method of any one of claims 1 to 5, wherein the multiplicity of infection (MOI) is higher in step a) than in step c).

7. The method of any one of claims 1 to 6, wherein the error-prone DNA polymerase is distinguished from a wild-type DNA polymerase by i) at least one amino acid substitution at a position corresponding to position 523, 479 or 560 in SEQ ID NO: 1, and / or ii) a conjugated nucleoside deaminase.

8. The method of claim 7, wherein the error-prone DNA polymerase further comprises one or more amino acid substitutions at a position corresponding to position 5, 7, 64, 120, 399, 429, 443, 444, 480, 520, 521, 522, 524, 530 or 611 in SEQ ID NO: 1.

9. The method of claim 7 or 8, wherein the error-prone DNA polymerase comprises an amino acid substitution at positions corresponding to one of the following sets of positions in SEQ ID NO: 1:(a) 523;(b) 479 and 523;(c) 523 and 560;(d) 479, 523 and 560;(c) 64, 120, 399 and 479;(f) 64, 120, 399 and 523;(g) 64, 120, 399 and 560;(h) 64, 120, 399, 479 and 523;(i) 64, 120, 399, 479 and 560;(j) 64, 120, 399, 523 and 560;(k) 64, 120, 399, 479, 523 and 560;(l) 5, 7, 64, 120, 399 and 523;(m) 5, 7, 64, 120, 399, 479 and 523;(n) 5, 7, 64, 120, 399, 523 and 560; and(o) 5, 7, 64, 120, 399, 479, 523 and 560.

10. The method of any one of claims 7 to 9, wherein the error-prone DNA polymerase comprises one or more amino acid substitutions selected from the following:(a) a substitution to A at a position corresponding to position 5 of SEQ ID NO: 1; (b) a substitution to A at a position corresponding to position 7 of SEQ ID NO: 1; (c) a substitution to C at a position corresponding to position 64 of SEQ ID NO: 1; (d) a substitution to L at a position corresponding to position 120 of SEQ ID NO: 1; (e) a substitution to T at a position corresponding to position 399 of SEQ ID NO: 1; (f) a substitution to N at a position corresponding to position 479 of SEQ ID NO: 1; (g) a substitution to R at a position corresponding to position 523 of SEQ ID NO: 1;and(h) a substitution to H at a position corresponding to position 560 of SEQ ID NO: 1;11. The method of any one of claims 7 to 10, wherein the wild-type DNA polymerase is a T7 DNA polymerase comprising an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 1.

12. The method of any one of claims 7 to 11, wherein the conjugated nucleoside deaminase has adenine deaminase and / or cytosine deaminase activity.

13. The method of claim 12, wherein the nucleoside deaminase is TadA-8e, PmCDAl or TadDE, or a variant thereof.

14. The method of any one of claims 7 to 13, wherein the nucleoside deaminase is conjugated to the N-terminus of the DNA polymerase.

15. The method of any one of claims 7 to 14, wherein the nucleoside deaminase is conjugated to the DNA polymerase via a polypeptide linker.

16. The method of claim 15, wherein the polypeptide linker comprises at least 24 amino acids.

17. The method of any one of claims 7 to 16, wherein the error-prone polymerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 7 to 26.

18. The method of claim 17, wherein the error-prone polymerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 24 or SEQ ID NO: 26.

19. The method of any one of claims 1 to 18, wherein the first and / or second host cell comprises a vector comprising a gene encoding the error-prone polymerase.

20. The method of any one of claim 19, wherein the gene for the error-prone DNA polymerase is operably linked to a promoter responsive to a signal from the phage vector.

21. The method of any one of claims 1 to 20, wherein the GOI is comprised in a selection vector capable of being packaged into the phage particles.

22. The method of claim 21, wherein the selection vector is a phagemid.

23. The method of any one of claims 1 to 22, wherein the first and / or second host cell is Escherichia coli, Bacillus spp., Salmonella spp., or Vibrio natriegens.

24. The method of any one of claims 1 to 23, wherein selection in step d) comprises coupling the function of the GOI to the fitness of the second host cell, and selecting host cells with a fitness above a predetermined threshold.

25. The method of any one of claims 1 to 23, wherein selection in step d) comprises coupling the function of the GOI to the generation of a detectable signal in the second host cell, and selecting host cells with a signal above or below a predetermined threshold.

26. The method of any one of claims 1 to 25, wherein the first host cell is deficient in a uracil-DNA glycosylase (UNG).

27. The method of any one of claims 1 to 26, wherein the first host cell comprises a polynucleotide comprising a nucleic acid sequence encoding a uracil glycosylase inhibitor (UGI).

28. The method of any one of claims 1 to 27, wherein the method further comprises repeating steps a) to d) by introducing a propagation-defective phage vector to a selected cell from step d).

29. An error-prone DNA polymerase that is distinguished from a wild-type DNA polymerase by i) at least one amino acid substitution at a position corresponding to position 523, 479 or 560 of SEQ ID NO: 1, and / or ii) a conjugated nucleoside deaminase.

30. The error-prone DNA polymerase of claim 29, wherein the polymerase further comprises one or more amino acid substitutions at a position corresponding to position 5, 7, 64, 120, 399, 429, 443, 444, 480, 520, 521, 522, 524, 530 or 611 of SEQ ID NO: 1.

31. The error-prone DNA polymerase of claim 29 or 30, wherein the polymerase comprises an amino acid substitution at positions corresponding to one of the following sets of positions of SEQ ID NO: 1:(a) 523;(b) 479 and 523;(c) 523 and 560;(d) 479, 523 and 560;(e) 64, 120, 399 and 479;(f) 64, 120, 399 and 523;(g) 64, 120, 399 and 560;(h) 64, 120, 399, 479 and 523;(i) 64, 120, 399, 479 and 560;(j) 64, 120, 399, 523 and 560;(k) 64, 120, 399, 479, 523 and 560;(l) 5, 7, 64, 120, 399 and 523;(m) 5, 7, 64, 120, 399, 479 and 523;(n) 5, 7, 64, 120, 399, 523 and 560; or(o) 5, 7, 64, 120, 399, 479, 523 and 560.

32. The error-prone DNA polymerase of any one of claims 29 to 31, wherein the polymerase comprises one or more amino acid substitutions selected from the following:(a) a substitution to A at a position corresponding to position 5 of SEQ ID NO: 1 (b) a substitution to A at a position corresponding to position 7 of SEQ ID NO: 1 (c) a substitution to C at a position corresponding to position 64 of SEQ ID NO: 1; (d) a substitution to L at a position corresponding to position 120 of SEQ ID NO: 1; (e) a substitution to T at a position corresponding to position 399 of SEQ ID NO: 1; (f) a substitution to N at a position corresponding to position 479 of SEQ ID NO: 1; (g) a substitution to R at a position corresponding to position 523 of SEQ ID NO: 1;and(h) a substitution to H at a position corresponding to position 560 of SEQ ID NO: 1;33. The error-prone DNA polymerase of any one of claims 29 to 32, wherein the wild-type DNA polymerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 1.

34. The error-prone DNA polymerase of any one of claims 29 to 33, wherein the conjugated nucleoside deaminase has adenine deaminase and / or cytosine deaminase activity.

35. The error-prone DNA polymerase of claim 34, wherein the nucleoside deaminase is TadA-8e, PmCDAl, or TadDE, or a variant thereof.

36. The error-prone DNA polymerase of any one of claims 29 to 35, wherein the nucleoside deaminase is conjugated to the N-terminus of the DNA polymerase.

37. The error-prone DNA polymerase of any one of claims 29 to 36, wherein the nucleoside deaminase is conjugated to the DNA polymerase via a polypeptide linker.

38. The error-prone DNA polymerase of claim 37, wherein the polypeptide linker comprises at least 24 amino acids.

39. The error-prone DNA polymerase of any one of claims 29 to 38, wherein the polymerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 7 to 26.

40. The error-prone DNA polymerase of claim 39, wherein the error-prone polymerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 24 or SEQ ID NO: 26.

41. An error-prone DNA polymerase comprising an amino acid sequence having at least 70% sequence identity to an amino acid sequence in SEQ ID NO: 1, wherein the error- prone polymerase comprises at least one of the following:a) a N at a position corresponding to position 479 of SEQ ID NO: 1;b) a R at a position corresponding to position 523 of SEQ ID NO: 1;c) a H at a position corresponding to position 560 of SEQ ID NO: 1; and / or d) a conjugated nucleoside deaminase.

42. The error-prone DNA polymerase of claim 41, wherein the DNA polymerase further comprises at least one of the following:a) a A at a position corresponding to position 5 of SEQ ID NO: 1;b) a A at a position corresponding to position 7 of SEQ ID NO: 1;c) a C at a position corresponding to position 64 of SEQ ID NO: 1;d) a L at a position corresponding to position 120 of SEQ ID NO: 1; and / or e) a T at a position corresponding to position 399 of SEQ ID NO: 1.

43. The error-prone DNA polymerase of claim 41 or 42, wherein the DNA polymerase comprises an amino acid sequence having at least 70% sequence identity to an amino acid sequence set forth in SEQ ID NO: 7 to 26.

44. A polynucleotide comprising a nucleic acid sequence encoding an error-prone DNA polymerase of any one of claims 29 to 43.

45. The polynucleotide of claim 44, wherein the polynucleotide comprises a nucleic acid sequence encoding a uracil glycosylase inhibitor (UGI).

46. An expression vector comprising a polynucleotide of claim 44 or 45.

47. A host cell comprising a polynucleotide of claim 44 or 45 or an expression vector of claim 46.

48. A method of producing an error-prone DNA polymerase of any one of claims 29 to 43, comprising culturing a host cell of claim 47 under conditions suitable for expressing the error-prone DNA polymerase.

49. A kit for performing a method according to any one of claims 1 to 28, the kit comprising an error-prone DNA polymerase of any one of claims 29 to 43, or an expression vector of claim 46.

50. The kit of claim 49, further comprising a phagemid.

51. The kit of claim 49 or 50, further comprising a propagation-defective phage vector.

52. The kit of any one of claims 49 to 51, further comprising a host cell for the phage vector.