Error-prone dnaps for orthogonal DNA replication
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2026-04-08
AI Technical Summary
Current orthogonal DNA replication systems, such as OrthoRep, have mutation rates that are too low to facilitate rapid evolution of genes in laboratory timescales, especially under purifying selection or without strong directional selection, limiting the ability to generate extensive genetic diversity and functional adaptation.
Development of DNA polymerases with increased error rates and activities, specifically orthogonal DNAPs that replicate plasmids at mutation rates exceeding 10^-4 substitutions per base, allowing for rapid and extensive evolution of encoded genes by introducing amino acid substitutions at strategic positions, thereby enhancing mutation rates and spectra.
These high-error-rate DNAPs enable rapid adaptation and divergence of encoded proteins, achieving significant genetic diversity and functional evolution within laboratory timescales, comparable to natural evolutionary processes, and allowing for the exploration of evolutionary constraints and selective forces.
Smart Images

Figure US2024031672_05122024_PF_FP_ABST
Abstract
Description
ERROR-PRONE DNAPS FOR ORTHOGONAL DNA REPLICATION
[0001] This application claims benefit of United States provisional patent application number 63 / 504,905, filed May 30, 2023, the entire contents of which are incorporated by reference into this application. REFERENCE TO A SEQUENCE LISTING
[0002] The content of the XML file of the sequence listing named “UCI017_seq”, which is 189 kb in size, created on May 29, 2024, and electronically submitted herewith the application, is incorporated herein by reference in its entirety. ACKNOWLEDGEMENT OF GOVERNMENT SUPPORT
[0003] This invention was made with government support under Grant Numbers 1DP2GM119163 and 1R35GM136297, awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND
[0004] Over the history of life, evolution has carried out a large-scale experiment exploring how gene sequences change under the constraints of prevailing or shifting structural and functional demands. The results of this natural experiment, embedded within the patterns of diversity across extant gene sequences, are of fundamental value to almost all areas of life sciences. For example, sequence conservation within a gene family is used to identify functionally critical residues, covariation among positions in an RNA or protein is used to deduce structural contacts and sectors of connectivity, differences in amino acid composition reveal environmental preferences (e.g., temperature or subcellular localization), and differences in the conserved physicochemical properties across regions of a protein reflect driving forces behind folding. Natural diversity across homologs also serves as a shared biomolecular engineering resource that can be mined for desired activities or recombined to access new functions. Additionally, machine learning (ML) models have proven incredibly effective at extracting meaningful representations of biomolecular structure and function from the extensive diversity within and across gene families, as exemplified by the ability of ML models to predict functional effects of mutations, design functional sequences, and predict protein structures. However, the natural evolution of highly diverse gene sequences under the constraints of selective forces — or conversely, the imprinting of selective forces and design principles into the statistics of sequence diversity — takes a long time at the slow rates of mutation in cellular and multicellular organisms. For example, reaching the 11% median divergence separating essential mouse and human genes took ~96 million years.Moreover, generating extensive collections of diverged sequences required complex histories of geographical isolation and speciation to maintain variation across populations in the face of within-population selective sweeps or genetic drift. Is it possible to compress long and vast gene evolution processes into laboratory experiments? Doing so would allow us to systematically detect novel structural and functional constraints governing biology, engineer custom biomolecules, create rich new sources of genetic variation for neofunctionalization or ML, and prospectively study the mechanisms and principles by which histories of selective forces become embedded into the patterns of sequence diversity.
[0005] We and others have endeavored towards this goal, including through the development of scalable accelerated continuous evolution systems such as our orthogonal DNA replication (OrthoRep) system in Saccharomyces cerevisiae. OrthoRep cells have an additional DNA replication system comprising an orthogonal DNA polymerase (DNAP) / plasmid pair wherein the orthogonal DNAP (TP-DNAP1) durably replicates the orthogonal plasmid (p1) but not the host genome; likewise, host DNAPs replicate the host genome but not p1 (Fig.1A). Through this architecture, OrthoRep supports the sustainable coexistence of two independent mutation rates in the same cell: a low mutation rate of 10-10substitutions per base (s.p.b.) for the host genome and a high mutation rate of 10-5s.p.b. forp1. OrthoRep’s 10-5s.p.b. mutation rate exceeds the error thresholds of the large hostgenome but not of the small p1 plasmid, allowing us to drive the rapid, continuous evolution of p1- encoded chosen genes as cells autonomously propagate. However, 10-5s.p.b. is stillnot high enough to observe extensive evolution on laboratory timescales in the general case where evolution occurs both with and without positive selection. In the specific case of evolution under strong positive selection, sufficiently large population sizes can directly compensate for moderate mutation rates by increasing the beneficial mutation supply on which selection “pulls”; in this case, OrthoRep systems have successfully evolved enzymes, biosynthetic pathways, biosensors, drug targets, and antibodies through long adaptive mutational pathways. Yet in the general case that includes when purifying selection is dominant or when selection is absent, both highly relevant in the generation of natural diversity and the ability to escape local fitness optima, mutation becomes the main force pushing sequence change. Without the pull of positive selection, 10-5s.p.b. OrthoRepsystems would take 100 generations (8-12 days for the yeast host of OrthoRep) just to observe an average of 1 new mutation in a typical 1 kb gene.
[0006] There remains a need for improved DNAPs that allow for the continuous evolution of genes at maximal speeds in vivo, with improved overall error rate, transversion rate, and activity.SUMMARY
[0007] Described herein are DNA polymerases (DNAPs) with increased error rates and activities. These DNAPs are ideally suited for OrthoRep due to the superior mutation rates and mutation spectrum they achieve. These new DNAPs allow OrthoRep-driven biomolecular evolution experiments to occur with much greater speed than before. Described herein are orthogonal DNAPs that durably replicate p1 at mutation ratesexceeding 10-4s.p.b. and which generate a remarkable level of divergence.
[0008] As demonstration of the power of these new orthogonal DNAPs for OrthoRep, small yeast populations encoding an enzyme on the orthogonal DNA replication system were serially passaged under selection for two months (540 generations), which resulted in the enzyme’s rapid adaptation and divergence into thousands of unique functional orthologs separated by an average pairwise distance of ~35 amino acids (10% divergence) and a maximum of ~60 amino acids (17% divergence). Applications for which OrthoRep has been used, including enzyme and antibody engineering, and future applications can be upgraded by these new DNAPs such that evolution of desired functions will be both faster and better, reaching superior outcomes or functions previously inaccessible.
[0009] In some embodiments, the error-prone DNA polymerase (DNAP) comprises an amino acid sequence having at least 90% identity with SEQ ID NO: 1 and at least three amino acid substitutions relative to SEQ ID NO: 1, wherein the at least three amino acid substitutions are at positions selected from E266, N282, I327, N449, L474, E488, Q598, K635, P680, F702, N713, K753, I761, I777, T828, I863, L900, and F965, and wherein theDNAP has a mutation rate greater than 10-6substitutions per base. In some embodiments,the DNAP has a mutation rate of at least 10-4substitutions per base. In some embodiments,the measured rate of mutation for any individual mutation type (i.e. A / T / G / C → A / T / G / C) isabove 10-7substitutions per base. In some embodiments, the measured rate of transversionmutations is greater than 4.89 x 10-7. In some embodiments, the measured rate of transversion mutations is greater than 6 x 10-6. In some embodiments, the measured rate of transversion mutations is greater than 3 x 10-5. In some embodiments, the measured rate oftransversion mutations is 1.89 x 10-6to 3.18 x 10-5. In some embodiments, the measuredrate of transition mutations is greater than 1.32 x 10-5. In some embodiments, the measuredrate of transition mutations is 1.41 x 10-5to 1.38 x 10-4.
[0010] In some embodiments, the DNAP comprises one or more of the amino acid substitutions shown in Table 1. In some embodiments, the at least three substitutions comprise (a) P680T; (b) I777K, I177T, or I777S; and (c) L900S. In some embodiments, theat least three substitutions comprise K635R, K753R, and F965Y. In some embodiments, the at least three substitutions comprise L474S and E488G.
[0011] In some embodiments, the DNAP comprises an amino acid sequence having at least 95% identity with an amino acid sequence selected from SEQ ID NOs: 3-18, and wherein the at least three substitutions comprise P680T. In some embodiments, the DNAP consists of an amino acid sequence selected from SEQ ID NOs: 3-18. In some embodiments, the DNAP comprises an amino acid sequence that varies from the sequences described herein due to truncation, insertions, deletions, and / or N- or C-terminal tags. Those skilled in the art understand that the numbering of amino acid positions may change with such modifications, but alignment with SEQ ID NO: 1 can be employed to identify substitutions corresponding to the positions identified as E266, N282, I327, N449, L474, E488, Q598, K635, P680, F702, N713, K753, I761, I777, T828, I863, L900, and F965 of SEQ ID NO: 1.
[0012] Also included is a nucleic acid molecule encoding the DNAP described herein. In some embodiments, the nucleic acid molecule further comprises a promoter. In some embodiments, the promotor sequence is selected from pSAC6, pPSP2, and pSLD3. In addition, included is a yeast host cell comprising a p1 plasmid and DNAP as described herein, and one or more p2 components for orthogonal replication of the p1 plasmid.
[0013] Described herein is a method of engineering a protein having a desired characteristic. In some embodiments, the method comprises subjecting a yeast host cell containing a p1 plasmid encoding the protein, and a DNAP as described herein, to error prone orthogonal replication. The method further comprises selecting yeast cells expressing the protein having the desired characteristic.
[0014] Further described is a kit for use in implementing methods of the disclosure. In some embodiments, the kit comprises reagents for integration of a gene of interest onto a p1 plasmid, and a DNAP as described herein. The kit optionally further comprises one or more reagents or devices for transforming a yeast cell therewith. In some embodiments, the kit further comprises a p1 plasmid packaged together with a yeast host cell comprising one or more p2 components for orthogonal replication of the p1 plasmid. In some embodiments, the yeast host cell is packaged together with one or more reagents or devices for culturing and / or transforming the yeast host cell. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIGS.1A-1F. Engineering orthogonal DNA polymerases for increased mutation rates. (1A) Architecture of the OrthoRep system. A DNA polymerase (TP-DNAP1) that exclusively replicates a specific cytoplasmically localized plasmid via protein primedreplication at a high error rate enables in vivo targeted mutagenesis without mutagenizing genomic DNA. (1B) Schematic for a directed evolution approach to engineer TP- DNAP1’s mutation rates and mutation spectrum incorporating both a direct selection for rare transversion mutations as well as high accuracy mutation rate measurement using a mutation accumulation and high throughput sequencing (HTS) assay. (1C) Mutations identified in TP-DNAP1 variants presented in this study. Shading and single letter amino acid code are used for only the first TP-DNAP1 in which a mutation is identified. Shading only is shown for all other instances of a mutation. (1D-1F) Mutation rate measurements via mutation accumulation for a series of TP-DNAP1 directed evolution intermediates showing either overall mutation rates as boxplots (1D), mean mutation rates for individual substitution types as heatmaps (1E), or mutation rates for individual substitution types for three individual TP-DNAP1 variants as boxplots (1F). Points are representative of individual biological replicates, each representing 3-4 timepoints with >50 sequences each. Boxplots and heatmaps are representative of n=4 biological replicates. Box plot central line, boxes, and whiskers represent the median, interquartile range, and minimum / maximum values, respectively.
[0016] FIGS.2A-2E. Massively parallel continuous diversification and evolution of TrpB. (2A) Schematic for OrthoRep-driven evolution of the tryptophan synthase β-subunit from Thermotoga maritima (TrpB) for standalone function in yeast. TrpB was integrated onto the p1 plasmid in a yeast strain lacking the native yeast tryptophan synthase gene (TRP5).96 independent cultures of the resulting strain were passaged mostly under selective pressure for Trp production using exogenously supplied indole over ~540 generations. DNA from fifteen timepoints throughout the evolution campaign was harvested and sequenced using HTS. (2B) Selection pressure for TrpB function is applied by lowering or eliminating exogenously supplied Trp and lowering exogenously supplied indole over time. The schedule of selection pressure imposed throughout extensive evolution is plotted. Timepoints at which cultures were harvested and sequenced using HTS are indicated. Selection periods were characterized as ‘no selection’, ‘mostly positive selection’, and ‘mostly purifying selection’ based on the Trp and indole amounts supplied and the progress of evolution. See Table 2 for a description of media conditions and selection pressure derivation. (2C-2E) Distribution of amino acid hamming distances for both pairwise comparisons and comparisons to the wild-type sequence, at the first and last timepoint (2D and 2E respectively) and as pairwise cumulative distributions for all timepoints (2C).
[0017] FIGS.3A-3C. Revealed effects of selective constraints. (3A) AlphaFold structure of Thermotoga maritima TrpB with different regions colored according to their structural role. First shell, β-β interface, and β-α interface residues are designated as such if they are withinthe αββα heterotetramer holoenzyme, respectively. Alignment to Pyrococcus furiosus TrpB crystal structures (PDB codes 5E0K and 5DW3) were used to determine distances from substrate and cofactor, α-subunit, and β-subunit. Mean solvent accessible surface area (SASA) was used to categorize all remaining residues as either surface (SASA≥0.2) or buried (SASA<0.2). (3B) Heatmap of mutations among OrthoRep-evolved TrpB sequences applied to the AlphaFold structure. (3C) Distributions of mutations among OrthoRep-evolved TrpB sequences within the 6 structural regions compared to a null model composed of a simulated dataset of sequences mutagenized in silico according to the mutation rates and preferences measured for TP-DNAP1 BadBoy2 until the number of synonymous mutations in the simulated sequences and real sequences were equivalent. (3D) Heatmap of mutation frequency for all mutations among the 20 most frequently mutated positions in the timepoint corresponding to either 50 or 70 generations. Frequencies for simulated sequences were subtracted to account for bias due to BadBoy2’s mutation preference and wt TrpB sequence content. Isoelectric points of the wild type TrpB, the TrpA-TrpB holoenzyme ortholog from Saccharomyces cerevisiae (Trp5) and an N-terminally- truncated Trp5 homologous to TrpB (Trp5-ΔN) are shown for comparison. (3E) Violin plots of isoelectric points for all OrthoRep- evolved and simulated TrpB sequences, split by timepoint. Points and black bars denote the means and interquartile range for all sequences within each timepoint.
[0018] FIGS.4A-4C. Lineage barcodes reveal covarying residues. (4A) Schematic of computational processing used to reduce phylogenetic contingency of residue covariation. (4B) Residue covariation plot for the most frequently covarying residues among all timepoints in the TrpB evolution dataset for 93 lineages downsampled to 100 sequences per lineage. (4C-4D) Heatmaps of most frequently mutated residues among sequences containing mutations A20V or A20T, downsampled to 100 sequences and chosen from specific timepoints. Heatmaps of all positions for early (generations 70 and 90) and late (generations 480 and 540) timepoints are overlaid onto a TrpB AlphaFold structure with the 20 most frequently mutated positions shown as spheres (4C) or shown as mutation-specific heatmaps of the most frequently mutated 20 positions in late timepoints (4D).
[0019] FIGS.5A-5G. Pooled measurement and TransceptEVE prediction of TrpB variant fitness. (5A) Schematic of pooled TrpB fitness assay using HTS. (5B) Spot plating growth assay of control sequences included in the pooled fitness assay. (5C) Hexbin plot of replicate concordance among pairs of replicates under growth conditions with Trp (no selection), without Trp and with 400 uM indole (weak selection) or without Trp and with 25 uM indole (strong selection) for highly functional sequences (enrichment score > -5) (5D) Distribution of mean enrichment scores (average of n=2 biological replicates) among thethree selection conditions. (5E and 5F) Hexbin plot of TranceptEVE score vs. measured mean enrichment score with strong selection for either all enrichment scores (5E) or enrichment scores for highly active sequences (5F). The percentage of all sequences that fall in the upper or lower quartile of score predictions and are classified as either high or low function (enrichment score greater than or less than -5, respectively) are shown in the respective sections of the plot in (5E). (5G) Hexbin plot of TransceptEVE score vs. number of nonsynonymous substitutions for all sequences with a measured strong selection mean enrichment score. r, Pearson correlation.
[0020] FIGS.6A-6D. Validation of mutation accumulation and comparison of p1 maintenance by legacy TP-DNAP1 variants. (6A-6D) Yeast strains encoding p1-leu2*-URA3 were transformed with plasmids encoding wild-type (wt) TP-DNAP1, TP-DNAP1-4-2, or TP- DNAP1-KS and passaged for 130 generations under selection for URA3 to allow for accumulation of mutations in leu2*. DNA was isolated from these samples at four timepoints throughout the experiment. Gel electrophoresis of these DNA samples (6A) revealed the resurgence of a DNA band corresponding in length to wt p1 (~9 kb) in the TP-DNAP1-4-2 sample, but not others. Timepoint DNA isolates were used for PCR amplification, high throughput sequencing, and mutation analysis of a ~450 bp region demonstrated that TP- DNAP1-KS maintained both a consistent mutation rate (6B, 6C) and monotonic diversification throughout the experiment (6D). Two replicates were isolated for the first timepoint for gel electrophoresis, but only one of these was isolated for subsequent timepoints and used for mutation accumulation. Mutations per base for TP-DNAP1 wt is shown rescaled in (6B) to highlight the poor linear fit. n.d., not detected.
[0021] FIGS.7A-7B. Mutation spectrum of legacy TP-DNAP1s. (7A-7B) Heatmaps of log transformed individual substitution rates for all 12 substitution types for TP-DNAP1-4-2 (7A) and TP-DNAP1-KS (7B) from the first two timepoints of mutation accumulation in HTS dataset I (~60 generations, see Fig.6).
[0022] FIG.8. Amino acid accessibility by mutation type. Accessibility of any codon for each amino acid from each of the 64 possible codons by one (left column), two (middle column), or three (last column) nucleotide mutations. Key (labeled “accessibility”) indicates markings to show whether the codon can reach the corresponding amino acid with only transversions (diagonal cross-hatched), with only transitions (leftward hatching), either with only transversions or with only transitions (straight cross-hatched), only with both transversions and transitions (rightward hatching), or is inaccessible (white) with at most the indicated number of mutable positions. Multiple mutations to the same position within the codon were not considered.
[0023] FIGS.9A-9C. Validation of transversion-specific selection and epPCR 1. (9A) Sanger sequencing of the ura3-K93N (ura3*) or trp5-K384* (trp5*) transversion mutations in two distinct clonal populations of yeast harboring the p1-ura3*-trp5* plasmid following selection in -uracil or -tryptophan media, respectively. In all cases, the only visible peaks were those corresponding to the original sequence or the desired transversion mutation.100% fixation of the expected mutation was not enforced due to the multicopy nature of p1. (9B) Plating assay demonstrating that p1-encoded auxotrophy markers URA3 and TRP5 are rendered inactive by the K93N and K384* mutations to catalytic lysines. (9C) URA3 revertant colony counts for either TP-DNAP1 epPCR 1 libraries or the parent TP-DNAP1-KS transformed into OR-Y488. TP-DNAP1-KS was used as the template sequence for error prone PCR mutagenesis of one of two regions of the polymerase expressed under one of two promoters for a total of four 103– 104member TP-DNAP1 libraries which were then transformed into the OR-Y488 strain harboring p1-ura3*-trp5*. Two biological replicates of each of these libraries were subject to selection for URA3 alongside mock libraries encoding only the parental TP-DNAP1 sequence and resulting revertants were quantified by titering and replating.
[0024] FIG.10. Polymerase replacement transformation. A strain encoding a “landing pad” DNA sequence that includes the wild type (wt) TP-DNAP1 and an I-SceI cut site at the CAN1 locus was co-transformed with both a TP-DNAP1 variant (typically in library format) and a transient I-SceI endonuclease expression cassette. Integration of NatMX was selected for using nourseothricin. Retention of the landing pad was prevented by counterselection of CAN1 using L-canavanine.
[0025] FIG.11. Flowchart of the programmatic steps performed by the mutation analysis for parallel laboratory evolution (Maple) pipeline. To facilitate rapid exploration and analysis of high throughput sequencing datasets, we built an end-to-end data processing pipeline that takes as input a sequencing dataset and a minimal set of additional user inputs and carries out the necessary steps to produce commonly desired visualizations as well as the data that supports those visualizations. This includes generating high accuracy consensus sequences from multiple reads of a sequence via rolling circle amplification (RCA) or unique molecular identifiers (UMI), alignment-based demultiplexing to separate and label sequences derived from different samples, and mutation analysis to generate human-readable .csv outputs that are further analyzed and visualized by Maple or can be viewed and analyzed by the user. Maple also includes an interactive dashboard that allows for user interaction with visualizations, such as selection of sequences that cluster together when plotted by the output of the dimensionality reduction tool PaCMAP.
[0026] FIGS.12A-12C. Identification of TP-DNAP1-TKS (epPCR 1). (12A) Fluctuation assay results for ~90 colonies isolated from OR-Y488 transformed with either the epPCR 1 library (diagonally hatched) or the parent TP-DNAP1-KS (white / horizontal hatch lines), then subject to URA3 selection. Note that per base mutation rate cannot be calculated by fluctuation analysis without copy number measurement; only per-cell mutation frequency is shown. Twelve library members (leftward hatching) were chosen for Sanger sequencing, using both mutation frequency and lineage (to minimize duplicate TP-DNAP1 variants) as criteria. Dotted line, mean mutation frequency for the three parent controls. (12B) Mutation rate measurements for TP-DNAP1-KS and five unique TP-DNAP1 variants isolated from this directed evolution round. Following selection for URA3 and either fluctuation analysis (using trp5) or trp5 selection, individual clones were grown in -uracil growth medium for mutation accumulation on trp5. Two timepoints separated by 100 generations of growth were used as template for high-throughput sequencing. Mutation rates are shown as a box plot and points for individual replicates where n > 1 biological replicate, or one line where n = 1. (12C) Cumulative hamming distance distribution for all high throughput sequencing samples at the end of mutation accumulation.
[0027] FIGS.13A-13C. Identification of Trixy and SgtKis (epPCR 2). (13A) Mutation rate measurements for all substitutions from HTS dataset 3 by mutation accumulation. Mutation rates are shown as a box plot and points for individual replicates where n > 1 biological replicate, or one line where n = 1. (13B-C) Heatmap representation of individual mutation rates for Trixy (13B) and SgtKis (13C) as measured in HTS dataset 3.
[0028] FIGS.14A-14D. Mutation accumulation and HTS analysis of TP-DNAP1 variants from manual recombination. (14A) Illustration of manually constructed TP-DNAP1 variants derived from mutation subsets of SgtKis transplanted onto Trixy. (14B-14D) HTS analysis from an 80-generation mutation accumulation on a ~1 kb region of p1, highlighting the cumulative pairwise hamming distance for a subset of polymerase variants tested (14B), the overall mutation rate measurements for all polymerases tested (14C), and the effect of mutation subset 2 on the mutation spectrum by log transformed heatmap of individual mutation rates (14D). Mutation rates in (14C) are plotted as points for rates measured for each replicate and box plots of summary statistics for all biological replicates (n ≥ 2).
[0029] FIGS.15A-15B. Single timepoint HTS analysis of epPCR 3. (15A-15B) An error prone PCR library (using TP-DNAP1-Trixy as the template) was subject to either simultaneous or sequential selection for ura3* and trp5* reversion and a ~2 kb region of p1 was PCR amplified and subject to HTS. Analysis of mean nucleotide mutations per sequence (15A) and unique transversions per sequence (total unique transversion mutations / number of sequences analyzed) (15B) was used to nominate mutations for the followinground of directed evolution. Nonsynonymous amino acid substitutions in addition to those found in Trixy are listed. Horizontal hatching, no nonsynonymous TP-DNAP1 mutations in addition to TP-DNAP1-Trixy mutations identified. Widely spaced diagonal hatching, multiple isolates contained the same polymerase. Narrowly spaced, isolate contained a unique mutation combination. Rightward diagonal hatching, contains a mutation to residue 777.
[0030] FIG.16. Relationship between p1 length and mutation rate. Mutation rate measurements for a subset of TP-DNAP1 variants and the length of recombinant p1 used to generate the indicated mutation accumulation dataset are shown. TP-DNAP1-Trixy, whose mutation rates show the most obvious relationship with p1 length, is highlighted with diagonal hatching. Data are identical to those shown in Figs.1C, 13, and 14 (left to right).
[0031] FIG.17. The effect of mutations to residue 777 on transition mutation rates. Mutation rate measurements from HTS dataset 6 for all TP-DNAP1 variants assayed for all four transition mutations, highlighting variants 3B, 3C, which differ from Trixy only by a single mutation (K777T / K777S, respectively) and BB-3B and BB-3C, which differ from BadBoy3 by only a single mutation (K777T / K777S, respectively).
[0032] FIGS.18A-18B. Control of p1 copy number via TP-DNAP1 expression level. TP- DNAP1 variants were placed under control of one of 3 promoter sequences of varying strength as determined by median protein abundance data. (18A) and p1 copy number in the resulting strains was determined by qPCR (18B) Points represent data for one technical replicate, thick horizontal lines represent average measurement for one biological replicate.
[0033] FIGS.19A-19C. The effect of selection on mutation rate. (19A) Illustration of the p1 subjected to mutation accumulation in HTS dataset 6. Cells were passaged in +uracil / - leucine synthetic complete media, enforcing functional selection for LEU2, but not ura3 or mScarlett-I (mScar). (19B-19C) Comparison of mutation rates (HTS dataset 6) for regions of similar length (~1 kb), either under selection (LEU2) or not under selection (ura3*, mScar) showing substitution rates for all substitution types as box plots and points (n=4 biological replicates) for a subset of TP-DNAP1 variants (19B) or as a scatter plot showing the relationship between mutation rates in all pairs of the three regions (19C) (one point per each n=1 individual polymerase variant replicate for all 17 TP-DNAP1 variants assayed). Dashed line depicts the x=y trendline in (19C).
[0034] FIGS.20A-20D. A rolling circle amplification (RCA) method for high accuracy long read nanopore sequencing. (20A) Illustration of the steps involved in the RCA method. Note that RCA primers must bind internally to the inner primers to prevent amplification of primer dimers. Primer sequences correspond to SEQ ID NOs: 76-79, respectively. (20B-20D) An example RCA and nanopore sequencing results using a Flongle flow cell for an amplicon of3 kb in length, showing gel electrophoresis of the RCA product (20B), the distribution of read lengths (20C), and the distribution of the number of repeats (estimated as the ratio of read length to subread length) vs. subread length (20D). Number of repeats and subread length were used for quality filtering prior to downstream analysis, the results of which are shownin 20D. Note that the bands corresponding to 3+ repeats were gel extracted for sequencing, resulting in the minimal amplicons with 2 or fewer repeats observed in 20C.
[0035] FIGS.21A-21E. Diversity of evolved TrpB variants. (21A) 2-dimensional representation for all unique genotypes identified using PaCMAP for dimensionality reduction, with the timepoint from in which each genotype was identified represented with shading intensity. (21B and 21C) 2D representations as in 21E, with shading intensity used to indicate the most frequently observed combinations of mutations among the six most frequently mutated positions in the final timepoint. (21D) Heatmap of all mutations to the 20 most commonly mutated positions throughout the entire experiment. (21E) Heatmaps of all mutations to the 20 most commonly mutated positions in 3 representative timepoints, with the mutation frequencies for the equivalent timepoint in the simulated sequences subtracted as background.
[0036] FIGS.22A-22E. Comparison of substitution distributions for evolved TrpB sequences and projected distributions. (22A-22C) Violin plots of total nucleotide substitutions (22A), synonymous substitutions (22B), or nonsynonymous substitutions (22C) per sequence in each timepoint of TrpB evolution (darker shapes) compared to projected distributions corresponding to sequences generated by a computational bulk mutation process using mutation rates and preferences of BadBoy2 (lighter shapes), with the extent of mutagenesis determined by the estimated number of generations. Note that synonymous substitution counts match the projected model, suggesting that synonymous substitutions are neutral. This contrasts with nonsynonymous substitution counts, suggesting they are not neutral. Points and black bars denote the means and interquartile range for all sequences within each timepoint.
[0037] FIG.23. Nonsynonymous mutation distributions for simulated and evolved sequences. Violin plots of nonsynonymous mutations per sequence in each timepoint of TrpB evolution compared to that of sequences generated by a computational mutation simulation process using mutation rates and preferences of BadBoy2, with the extent of mutagenesis determined by generating an equivalent number of synonymous mutations as that observed in evolved sequences. Points and black bars denote the means and interquartile range for all sequences within each timepoint.
[0038] FIG.24. Effect of long-term mutagenesis on hydrophobicity. Violin plots of Kyte- Doolittle hydrophobicity index in each timepoint of TrpB evolution for both evolved and simulated sequences. Points and black bars denote the means and interquartile range for all sequences within each timepoint. Hydrophobicity indices of the wild type TrpB, the TrpA- TrpB holoenzyme ortholog from Saccharomyces cerevisiae (Trp5) and an N-terminally- truncated Trp5 (Trp5-ΔN), homologous to TrpB, are shown for comparison.
[0039] FIG.25. Distribution of mesophile mutation fraction for both evolved and simulated TrpB sequences. Mesophile mutation fraction was calculated for each sequence as the number of mesophile mutations divided by the total mutations. A set of 17 amino acid replacements (e.g. Proline to Serine) identified by Haney et al. were designated as mesophile mutations. Only sequences from the final timepoint (540 generations) were included. DETAILED DESCRIPTION
[0040] Described herein are improved OrthoRep systems with a mutation rate up to 1.7 × 10-4s.p.b., corresponding to a new mutation in a typical 1 kb gene once every <10generations in the absence of any selection. This intensified mutational force on chosen genes, operating at >1 million times the mutation rate of the host genome, allows one to mimic extended periods of natural gene evolution on laboratory timescales in the general case. Shown herein is how, in ~3 months of laboratory passaging (totaling <15 hours of researcher intervention) of 96 independent populations, a conditionally essential gene encoded on p1 diverges to an extent where the median distance separating pairs of evolved sequences is 35 amino acids, with thousands of unique pairs separated by >60 amino acids. This corresponds to an amino acid divergence of ~9% to >15%, exceeding the median 11% distance between mouse and human orthologous genes. By analyzing the rich collection of diverged sequences throughout their laboratory evolutionary history, we uncover hidden forces shaping and constraining sequence change, such as a preference for negative net charge, supporting a proposed mechanism by which proteins avoid large-scale indiscriminate clustering in the crowded environment inside cells. We also extract examples of allosteric network remodeling through the cooccurrence of mutations across distinct clades and temperature optimization through amino acid content change. Overall, the invention provides an approach to systematically reveal the evolutionary constraints and selective forces that genes experience and delivers an upgraded OrthoRep system for broad application.Definitions
[0041] All scientific and technical terms used in this application have meanings commonly used in the art unless otherwise specified. As used in this application, the following words or phrases have the meanings specified.
[0042] As used herein, “parental sequence” refers to an initial sequence that is subjected to mutagenesis and selection. In the context of mutagenesis by OrthoRep , the parental sequence refers to the sequence of the gene of interest provided on a p1 integration plasmid or the protein it encodes that is to be artificially evolved to have one or more desired characteristics. Although one or more sequences on the p1 integration plasmid that are provided for effecting orthogonal replication, surface display, selection, and / or detection may also be artificially evolved by way of being integrated on the p1 expression plasmid, such a sequence is not considered part of the parental sequence unless mutations in the sequence caused by OrthoRep will be specifically selected over its original starting sequence. In the context of mutagenesis by ex vivo error-prone PCR, the parental sequence refers to the sequence that is used as an error-prone PCR template for a particular cycle of engineering.
[0043] As used herein, a “p1 plasmid” refers to a plasmid capable of orthogonal replication in yeast cells. P1 plasmids comprise recognition elements, which minimally include p1- specific terminal proteins (TPs) and terminal inverted repeats, that are needed for replication of a gene of interest by a TP-DNAP1. The terms “p1” and “P1” may be used interchangeably.
[0044] As used herein, a “p1 integration plasmid” refers to a circular or linear plasmid that is used to insert a gene of interest into a p1 plasmid of a yeast cell by homologous recombination after transducing the yeast cell therewith.
[0045] As used herein, a “p1 expression plasmid” refers to the p1 plasmids of a yeast cell that have been modified to express a given parental sequence and copies thereof resulting from one or more OrthoRep rounds.
[0046] As used herein, “p2 components” refers to the components encoded on naturally occurring p2 plasmids and derivatives thereof that are needed for orthogonal replication of p1 plasmids. One or more of the p2 components need not be encoded on a p2 plasmid, but may instead be encoded in the yeast host cell’s nuclear DNA or in another plasmid (including p1 expression plasmids) found in the yeast host cell. The terms “p2” and “P2” may be used interchangeably.
[0047] As used herein, a “desired characteristic” refers to a structure or function that one desires a given protein to obtain that it does not already possess. Such desiredcharacteristics include: affinity; selectivity; agonism; antagonism; inhibition; irreversible binding; enhancement; a different affinity, avidity, and / or specificity for a target the protein is already capable of binding; an ability to bind a new target; an ability to catalyze a given reaction it is already capable of catalyzing but with a different efficiency and / or under different reaction conditions; an ability to catalyze a new reaction that gives a new product or the same reaction product it already produces but by way of a different synthetic pathway; a change in its resistance or susceptibility to a given condition, e.g., heat, moisture, a given pH, a given chemical or other biomolecule (e.g., protease), degradation, agglutination; a change in a structural domain, a structural motif, a protein fold, and / or supersecondary structure; and the like.
[0048] As used herein, an “affinity reagent” refers to a compound (e.g., an antibody or fragment thereof, a receptor, an enzyme, etc.) that specifically binds a given target (e.g., a compound or composition, a protein, a nucleic acid molecule, etc.), or vice versa. For example, an affinity reagent may an enzyme that binds with a protein substrate or the affinity reagent may be the protein substrate that binds with the enzyme.
[0049] As used herein, a given percentage of “sequence identity” refers to the percentage of nucleotides or amino acid residues that are the same between sequences, when compared and optimally aligned for maximum correspondence over a given comparison window, as measured by visual inspection or by a sequence comparison algorithm in the art, such as the BLAST algorithm, which is described in Altschul et al., (1990) J Mol Biol 215:403-410. Software for performing BLAST (e.g., BLASTP and BLASTN) analyses is publicly available through the National Center for Biotechnology Information (ncbi.nlm.nih.gov). The comparison window can exist over a given portion, e.g., a functional domain, or an arbitrarily selection a given number of contiguous nucleotides or amino acid residues of one or both sequences. Alternatively, the comparison window can exist over the full length of the sequences being compared. For purposes herein, where a given comparison window (e.g., over 80% of the given sequence) is not provided, the recited sequence identity is over 100% of the given sequence. Additionally, for the percentages of sequence identity of the proteins provided herein, the percentages are determined using BLASTP 2.8.0+, scoring matrix BLOSUM62, and the default parameters available at blast.ncbi.nlm.nih.gov / Blast.cgi. See also Altschul, et al., (1997) Nucleic Acids Res 25:3389-3402; and Altschul, et al., (2005) FEBS J 272:5101-5109.
[0110] Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith & Waterman, Adv Appl Math 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J Mol Biol 48:443 (1970), by the search for similarity method of Pearson & Lipman, PNAS USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT,FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by visual inspection.
[0050] As used herein, the terms “protein”, “polypeptide” and “peptide” are used interchangeably to refer to two or more amino acids linked together. Groups or strings of amino acid abbreviations are used to represent peptides. Except when specifically indicated, peptides are indicated with the N-terminus on the left and the sequence is written from the N-terminus to the C-terminus.
[0051] Polypeptides may be made using methods known in the art including chemical synthesis, biosynthesis or in vitro synthesis using recombinant DNA methods, and solid phase synthesis. See, e.g., Kelly & Winkler (1990) Genetic Engineering Principles and Methods, vol.12, J. K. Setlow ed., Plenum Press, NY, pp.1-19; Merrifield (1964) J Amer Chem Soc 85:2149; Houghten (1985) PNAS USA 82:5131-5135; and Stewart & Young (1984) Solid Phase Peptide Synthesis, 2ed. Pierce, Rockford, IL, which are herein incorporated by reference. Polypeptides may be purified using protein purification techniques known in the art such as reverse phase high-performance liquid chromatography (HPLC), ion-exchange or immunoaffinity chromatography, filtration or size exclusion, or electrophoresis. See, e.g., Olsnes and Pihl (1973) Biochem.12(16):3121-3126; and Scopes (1982) Protein Purification, Springer-Verlag, NY, which are herein incorporated by reference. Alternatively, the polypeptides may be made by recombinant DNA techniques known in the art.
[0052] As used herein, “antibody” refers to naturally occurring and synthetic immunoglobulin molecules and immunologically active portions thereof (i.e., molecules that contain an antigen binding site that specifically bind the molecule to which antibody is directed against, such as minibodies and nanobodies). As such, the term antibody encompasses not only whole antibody molecules, but also antibody multimers and antibody fragments as well as variants (including derivatives) of antibodies, antibody multimers and antibody fragments. Examples of molecules which are described by the term “antibody” herein include: single chain Fvs (scFvs), Fab fragments, Fab’ fragments, F(ab’)2, disulfide linked Fvs (sdFvs), Fvs, and fragments comprising or alternatively consisting of, either a VL or a VH domain.
[0053] As used herein, a compound (e.g., receptor or antibody) “specifically binds” a given target (e.g., ligand or epitope) if it reacts or associates more frequently, more rapidly, with greater duration, and / or with greater binding affinity with the given target than it does with a given alternative, and / or indiscriminate binding that gives rise to nonspecific binding and / or background binding. As used herein, “non-specific binding” and “background binding” refer to an interaction that is not dependent on the presence of a specific structure (e.g., a givenepitope). An example of a compound that specifically binds a given target is an antibody that binds its target antigen with greater affinity, avidity, more readily, and / or with greater duration than it does to other compounds. As used herein, an “epitope” is the part of a molecule that is recognized by an antibody. Epitopes may be linear epitopes or three-dimensional epitopes. As used herein, the terms “linear epitope” and “sequential epitope” are used interchangeably to refer to a primary structure of an antigen, e.g., a linear sequence of consecutive amino acid residues, that is recognized by an antibody. As used herein, the terms “three-dimensional epitope” and “conformational epitope” are used interchangeably to refer a three-dimensional structure that is recognized by an antibody, e.g., a plurality of non- linear amino acid residues that together form an epitope when a protein is folded.
[0054] As used herein, “binding affinity” refers to the propensity of a compound to associate with (or alternatively dissociate from) a given target and may be expressed in terms of itsdissociation constant, Kd. In some embodiments, the antibodies have a Kd of 10-5or less,10-6or less, preferably 10-7or less, more preferably 10-8or less, even more preferably 10-9or less, and most preferably 10-10or less, to their given target. Binding affinity can bedetermined using methods in the art, such as equilibrium dialysis, equilibrium binding, gel filtration, immunoassays, surface plasmon resonance, and spectroscopy using experimental conditions that exemplify the conditions under which the compound and the given target may come into contact and / or interact. Dissociation constants may be used determine the binding affinity of a compound for a given target relative to a specified alternative. Alternatively, methods in the art, e.g., immunoassays, in vivo or in vitro assays for functional activity, etc., may be used to determine the binding affinity of the compound for the given target relative to the specified alternative.
[0055] “Nucleotide sequence” refers to a heteropolymer of deoxyribonucleotides, ribonucleotides, or peptide-nucleic acid sequences that may be assembled from smaller fragments, isolated from larger fragments, or chemically synthesized de novo or partially synthesized by combining shorter oligonucleotide linkers, or from a series of oligonucleotides, to provide a sequence which is capable of expressing the encoded protein.
[0056] As used herein, “a” or “an” means at least one, unless clearly indicated otherwise. DNA Polymerases (DNAPs)
[0057] Described herein are DNA polymerases (DNAPs) with increased error rates and activities. These DNAPs are ideally suited for OrthoRep due to the superior mutation rates and mutation spectrum they achieve. These new DNAPs allow OrthoRep-driven biomolecular evolution experiments to occur with much greater speed than before.Described herein are orthogonal DNAPs that durably replicate p1 at mutation ratesexceeding 10-4s.p.b. and which generate a remarkable level of divergence.
[0058] In some embodiments, the error-prone DNA polymerase (DNAP) comprises an amino acid sequence having at least 90% identity with SEQ ID NO: 1 and at least three amino acid substitutions relative to SEQ ID NO: 1, wherein the at least three amino acid substitutions are at positions selected from E266, N282, I327, N449, L474, E488, Q598, K635, P680, F702, N713, K753, I761, I777, T828, I863, L900, and F965, and wherein theDNAP has a mutation rate greater than 10-6substitutions per base. In some embodiments,the DNAP has a mutation rate of at least 10-4substitutions per base. In some embodiments,the measured rate of mutation for any individual mutation type (i.e. A / T / G / C → A / T / G / C) isabove 10-7substitutions per base. In some embodiments, the measured rate of transversionmutations is greater than 4.89 x 10-7. In some embodiments, the measured rate of transversion mutations is greater than 6 x 10-6. In some embodiments, the measured rate of transversion mutations is greater than 3 x 10-5. In some embodiments, the measured rate oftransversion mutations is 1.89 x 10-6to 3.18 x 10-5. In some embodiments, the measuredrate of transition mutations is greater than 1.32 x 10-5. In some embodiments, the measuredrate of transition mutations is 1.41 x 10-5to 1.38 x 10-4.
[0059] In some embodiments, the DNAP comprises one or more of the amino acid substitutions shown in Table 1. In some embodiments, the at least three substitutions comprise (a) P680T; (b) I777K, I177T, or I777S; and (c) L900S. In some embodiments, the at least three substitutions comprise K635R, K753R, and F965Y. In some embodiments, the at least three substitutions comprise L474S and E488G.
[0060] Table 1: TP-DNAP1 mutation rates measured in HTS dataset 6
[0061] Table 1 (cont’d): median substitution rate (s.p.b., n=4)
[0062] Table 1 (cont’d): median substitution rate (s.p.b., n=4)
[0063] Table 1 (cont’d): median substitution rate (s.p.b., n=4)
[0064] Table 1 (cont’d): mutations
[0065] Table 1 (cont’d): mutations*n.d., not detected
[0066] In some embodiments, the DNAP comprises an amino acid sequence having at least 95% identity with an amino acid sequence selected from SEQ ID NOs: 3-18, and wherein the at least three substitutions comprise P680T. In some embodiments, the DNAP consists of an amino acid sequence selected from SEQ ID NOs: 3-18. In some embodiments, the DNAP comprises an amino acid sequence that varies from the sequences described herein due to truncation, insertions, deletions, and / or N- or C-terminal tags. Those skilled in the art understand that the numbering of amino acid positions may change with such modifications, but alignment with SEQ ID NO: 1 can be employed to identify substitutions corresponding to the positions identified as E266, N282, I327, N449, L474, E488, Q598, K635, P680, F702, N713, K753, I761, I777, T828, I863, L900, and F965 of SEQ ID NO: 1.
[0067] Also included is a nucleic acid molecule encoding the DNAP described herein. In some embodiments, the nucleic acid molecule further comprises a promoter. Promoters can be selected in accordance with desired TP-DNAP1 expression levels to modify p1 copy number, as shown in Fig.18. In some embodiments, the promotor sequence is selected from pSAC6, pPSP2, and pSLD3. In addition, included is a yeast host cell comprising a p1 plasmid and DNAP as described herein, and one or more p2 components. The p2 components can include RNA polymerases and other transcriptional machinery for expression of genes encoded on p1 and p2, capping enzymes and other machinery for translation of genes encoded on p1 and p2, and replication machinery for the replication of p1 and p2. P2 components can be interpreted as accessories to the orthogonal replication of the p1 plasmid by DNAPs described herein. Methods
[0068] As demonstration of the power of these new orthogonal DNAPs for OrthoRep, small yeast populations encoding an enzyme on the orthogonal DNA replication system were serially passaged under selection for two months (540 generations), which resulted in the enzyme’s rapid adaptation and divergence into thousands of unique functional orthologs separated by an average pairwise distance of ~35 amino acids (10% divergence) and a maximum of ~60 amino acids (17% divergence). Applications for which OrthoRep has been used, including enzyme and antibody engineering, and future applications can be upgraded by these new DNAPs such that evolution of desired functions will be both faster and better, reaching superior outcomes or functions previously inaccessible.
[0069] Described herein is a method of engineering a protein having a desired characteristic. In some embodiments, the method comprises subjecting a yeast host cell containing a p1 plasmid encoding the protein, and a DNAP as described herein, to error prone orthogonal replication. The method further comprises selecting yeast cells expressing the protein having the desired characteristic.
[0070] In some embodiments, different TP-DNAP1 expression levels are employed to modify p1 copy number, as illustrated in Fig.18, which demonstrates that altering the promoter driving the expression of TP-DNAP1 can influence p1 copy number.
[0071] In some embodiments, the engineered protein is an enzyme. In some embodiments, the engineered protein is an antibody, binding portion thereof, or other protein capable of selectively binding a target. In some embodiments, the engineered protein is a biosensor, capable of sensing small molecules or macromolecules and transducing a response. In some embodiments, the engineered protein is a gene editor or gene therapy agent. In someembodiments, the engineered protein may be multiple proteins or enzymes comprising a complex or metabolic pathway. Kits
[0072] Further described is a kit for use in implementing methods of the disclosure. In some embodiments, the kit comprises reagents for integration of a gene of interest onto a p1 plasmid, and a DNAP as described herein. The gene of interest typically encodes a protein to be engineered using the methods described herein. The kit optionally further comprises one or more reagents or devices for transforming a yeast cell therewith. In some embodiments, the kit further comprises a p1 plasmid packaged together with a yeast host cell comprising one or more p2 components for orthogonal replication of the p1 plasmid. In some embodiments, the yeast host cell is packaged together with one or more reagents or devices for culturing and / or transforming the yeast host cell. EXAMPLES
[0073] The following examples are presented to illustrate the present invention and to assist one of ordinary skill in making and using the same. The examples are not intended in any way to otherwise limit the scope of the invention. Example 1: Motivation and general strategy for orthogonal DNAP engineering
[0074] When a gene evolves under prevailing and shifting functional demands, those demands become embedded into the resulting diversity of homologous sequences as patterns of conservation and change. We have long looked for conservation in these patterns to extract features responsible for gene function and have also reconstructed phylogenetic trees from these patterns to infer how genes changed to satisfy new demands, offering lessons in biomolecular design. Yet for the most part, only natural evolution, through its millions of years of operation, has diverged genes far enough to yield sets of homologous sequences that contain abundant statistical information on the details of gene function. Compressing long-term gene evolution to laboratory timespans at scale would break Nature’s monopoly on the production of extensive gene diversity, enabling the systematic detection of novel structural and functional demands governing biology, the engineering of custom biomolecules, and the prospective study of evolutionary mechanisms and principles by which extant gene diversity was generated in the first place.
[0075] Orthogonal DNA replication (OrthoRep) is a genetic architecture for continuously hypermutating user-defined genes in vivo, but the mutation rates of legacy systems are too low to condense long gene evolutionary trajectories onto laboratory timescales when strong directional selection is absent. We engineered OrthoRep systems in yeast to have mutationrates exceeding 10-4 substitutions per base: a new mutation in a typical 1 kb gene is sampled once every <10 generations. This intensified mutational force acting on chosen genes should enable their extensive laboratory evolution towards nature-like levels of diversity under all types of selection.
[0076] We encoded a maladapted gene, TrpB, onto OrthoRep and continuously evolved it to gain and maintain the ability to synthesize tryptophan. During ~540 generations in 96 independent populations (~3 months of laboratory passaging and <15 hours of researcher intervention), TrpB diverged extensively such that the median distance separating pairs of evolved sequences reached 35 amino acids (~9%), with thousands of unique pairs separated by >60 amino acids (~15%). For comparison, the median distance between mouse and human orthologous genes is 11%. The high fitness of extensively diverged TrpB variants were not predictable by advanced machine learning models trained on natural variation. Analyzing the rich collection of resultant TrpB sequences – referenced against a precise null model of sequence change simulated from detailed measurements on the mutation rates and preferences of our new OrthoRep systems – revealed both known and unexpected factors influencing TrpB’s function and evolution at high resolution. We identified varying degrees of conservation consistent with structural models, pinpointed regions and cooccurring networks of mutation responsible for functional adaptation, found statistical signatures of thermoadaptation, and detected a clear trend toward net negative charge in support of a proposed mechanism by which proteins avoid large-scale indiscriminate clustering.
[0077] OrthoRep’s newfound ability to condense long adaptive and neutral gene evolutionary processes into accessible laboratory experiments should drive the evolutionary engineering of new biomolecules, the broader mapping of fitness landscapes, and the statistically-powered extraction of novel biological forces governing the functions of genes and biomolecules. Example 2: Materials and Methods
[0078] DNA plasmid construction
[0079] Plasmids used in this study are listed in Table 4, along with sources for DNA templates. Complete maps for these plasmids are made available for download on github at github.com / liusynevolab / OrthoRep_Rix_2023. All DNA templates for PCR were derived from previous studies or gBlocks (IDT). All primers were synthesized by IDT. All relevant primer pairs are listed in Table 5. Amplicons for construction of clonal plasmids were generated using Q5 Hot Start High-Fidelity DNA Polymerase (NEB). All non-library plasmids wereconstructed using Gibson Assembly and transformed into chemically competent E. coli strain TOP10 (ThermoFisher). Clonal plasmids were sequence verified by either Sanger sequencing (Azenta) or whole plasmid sequencing (Primordium).
[0080] DNA library construction
[0081] Amplicons for TP-DNAP1 libraries were generated using error prone PCR with GeneMorph II (Agilent) according to manufacturer instructions, aiming for ~3-5 nucleotide substitutions per sequence. Amplicons for all other libraries were generated with Q5 Hot Start High-Fidelity DNA Polymerase (NEB). For epPCR 1, the resulting PCR product was assembled into plasmids using Gibson assembly in 20 μL reaction volumes. For all other libraries, resulting PCR products were assembled into plasmids with Golden Gate assembly with T4 DNA ligase and BsaI-HF v2 or PaqCI (all NEB) in a 40 μL reaction volume. Gibson reactions were run at 50 °C for 1 hour. Golden gate reactions were run isothermally at 37 °C for 1 hour and heat inactivated at 65 °C for 10 minutes. Reactions were purified with AMPure XP beads (Beckman), typically with a 0.9:1 bead:sample ratio according to manufacturer instructions. Libraries were transformed into high-competency electrocompetent E. coli TOP10 cells (ThermoFisher).
[0082] Yeast strains, media, transformations, and DNA extraction
[0083] All yeast strains used in this study and their provenance are listed in Table 6. Yeast were grown in liquid or on plates at 30 °C in synthetic complete (SC) growth medium (20 g / L dextrose, 6.7 g / L yeast nitrogen base w / ammonium sulfate w / o amino acids (US Biological), appropriate nutrient drop-out mix (US Biological), as directed) or MSG SC growth medium (20 g / L dextrose, 1.72 g / L yeast nitrogen base w / o ammonium sulfate w / o amino acids (US Biological), appropriate nutrient drop-out mix (US Biological), as directed, 1 g / L L-Glutamic acid monosodium salt hydrate (ThermoFisher)) minus nutrients (referred to as -X where X is either the single letter amino acid code for an amino acid nutrient, or U for uracil) required for appropriate auxotrophy selection(s). Where selection for MET15 was required, cells were propagated in media lacking both methionine and cysteine.500 μL liquid yeast cultures in 96-well deep well plates were incubated with shaking at 750 rpm. All other liquid yeast cultures were incubated with shaking at 200 rpm.
[0084] Yeast transformations, including p1 integrations and polymerase replacement integration, were performed as previously described (38). For all integration transformations, plasmid DNA was linearized prior to transformation using either ScaI-HF or EcoRI-HF (both NEB) for p1 or genomic integrations, respectively. Due to its repetitive nature, deletion of FLO1 was performed by a URA3 knock-in knock-out method (70) (see Table 4, pFLO1-KO). Genetic deletions for TRP5 and MET15 were performed as previously described (71), usingspacer sequences TTTGAGCCTGATCCCACTAG (SEQ ID NO.74) and GCTAAGAAGTATCTATCTAA (SEQ ID NO.75), respectively. When isolating individual clones from genetic deletion and integration transformations, colonies were restreaked onto media agar plates of the same formulation to ensure isolation of only cells that have the desired genetic change.
[0085] All p1 plasmid sequences were generated by first generating a strain harboring a ‘landing pad’ p1 via integration and then integrating over this landing pad to generate the desired p1 construct. To enable construction of the landing pad strain, the wt TP-DNAP1 was integrated at the CAN1 locus using pGR475. A sequence encoding a non-functional partial LEU2 sequence lacking the N terminus was then integrated over the wt p1 using pGR420 to generate the landing pad p1. The wt p1, which encodes the TP-DNAP1, was then cured out via 3-41:1000 passages. The p1 plasmid(s) encoding the desired sequence were then generated via integration using cassette(s) that include a LEU2 sequence lacking the C terminus (e.g. pGR438). The overlap between the LEU2 on the landing pad and the new integration cassette reconstitutes full length LEU2 only when integration occurs on p1, reducing the likelihood of genomic integration.
[0086] The polymerase replacement integration transformation was performed by first digesting 0.5-2 μg of the polymerase replacement plasmid or library with EcoRI-HF in a 25 μL reaction per 1x transformation followed by directly transforming this digestion reaction into a yeast strain encoding the CAN1-WT-TP-DNAP1 landing pad (all polymerase libraries were transformed into OR-Y488). Library transformations were carried out at 20-40x scale. Transformed yeast were plated onto solid MSG SC -LR or -MCR media w / 100 mg / L nourseothricin (for positive selection of integration) and 200 mg / L l-canavanine (a toxic l- arginine analog for counterselection of cells that fail to perform polymerase replacement and remove the arginine permease CAN1). Leu or Met / Cys dropout was used to maintain selection for p1-encoded LEU2 or MET15, respectively, while Arg dropout was used to improve l-canavanine selection.
[0087] Extraction of genomic DNA (gDNA) and p1 / p2 plasmids was performed as previously described for 1.5 mL yeast culture volumes (38). This procedure was used for all experiments except for DNA extracted for use in HTS dataset 6, which was instead performed in 96-well format. In brief, a 96-well block of 500 μL of saturated yeast cultures was centrifuged (2500 × G, 5 min), supernatant was discarded, pellets were resuspended in 1 mL 0.9% NaCl, this resuspension was again centrifuged (2500 × G, 5 min), and the supernatant was discarded. The resulting pellet was resuspended in 250 μL Zymolyase solution (0.9 M D-Sorbitol (Sigma Aldrich), 0.1 M Ethylenediaminetetraacetic acid (EDTA, Sigma Aldrich), 10 U / mL Zymolyase (US Biological)) and incubated with shaking (37 °C, 200RPM). The 96-well block was then centrifuged (2500 × G, 5 min), supernatant was discarded, and pellets were resuspended in 280.5 μL proteinase K solution (250 μL TE (50 mM Tris-HCl (pH 7.5), 20mM EDTA), 25 μL 10% sodium dodecyl sulfate (SDS, Sigma Aldrich), 5.5 μL proteinase K stock solution (10 mg / mL proteinase K (ThermoFisher)). The 96-well block was then incubated at 65 °C for 30 min, combined with 75 μL 5M potassium acetate (ThermoFisher), and incubated on ice for 30 min. The 96-well block was centrifuged at 12,000 × g for 10 min, the resulting supernatant was combined and mixed with 2 volumes buffer PB (5 M Guanidine hydrochloride (ThermoFisher), 30% isopropanol, 70% water), and this mixture was applied to a 96 well DNA-binding plate (Epoch Life Science) on a vacuum manifold. Flow through was discarded, columns were washed with PE buffer (10 mM Tris- HCl (ThermoFisher), 80% ethanol, 20% water, pH 7.5), centrifuged and dried, and 60 μL water was applied to columns for elution by centrifugation (2500 × G, 5 min).
[0088] High throughput sequencing
[0089] All high throughput sequencing datasets are listed in Table 7, along with the method used to construct them. All PCRs for high throughput sequencing were performed with Platinum SuperFi II DNA Polymerase (ThermoFisher). For short read paired end sequencing, both low yield (AmpliconEZ, Azenta) and high yield (HiSeq paired end 150, Novogene) were performed directly on PCR products generated in either one or two rounds of PCR using primers that each included an adapter sequence, a 6- or 7-nucleotide barcode, or both.
[0090] For in-house long read high throughput sequencing, we used the Oxford Nanopore Technologies nanopore sequencing platform. Due to the lower accuracy of nanopore sequencing, we employed modified versions of previously described methods for the construction of DNA libraries that yield multiple reads of the same original DNA molecule, allowing for computational reconstruction of high accuracy sequences (53, 72). We refer to the first of these as in vivo downsampled unique molecular identifier (UMI) PCR, which was used for most nanopore sequencing in this study (Table 7). This involved first a 2-cycle “UMI tagging” reaction, in which primers (e.g., primer pair 2) were used to append both UMIs and universal DNA sequences (for further amplification) to both ends of the target sequence with the following components: ~1-50 ng of purified yeast or E. coli miniprep 5 μL 2x SuperFi II master mix 1 μM each primer water to 10 μLand with these thermocycler conditions: 98 °C for 30 sec 98 °C for 10 sec 65 °C for 1 sec 60 °C for 45 sec w / ramp down from 65 to 60 @ 0.2 °C per second (this should result in 25 seconds of ramp down time and 20 seconds of hold time) 72 °C extension, 1 min / kb Go to step 2 (1x) 72 °C for 2 min.
[0091] Next, UMI primers were removed using ExoSAP-IT (ThermoFisher) according to manufacturer instructions.10 μL of the resulting reaction was then used as template in a 25 μL Platinum SuperFi II PCR with 1 mM MgCl2 supplemented, using primers (e.g., primer pair 3) that included BsaI or PaqCI sites and 7 nucleotide barcodes in forward / reverse combinations that were unique for each sample. The resulting uniquely barcoded PCR products were then combined, purified using AMPure XP beads, used in a library Golden Gate reaction with an E. coli vector, then transformed into high competency E. coli as described above. Plasmid pGR554 (Table 4) was designed for this purpose and contains CcdB and sfGFP, which both are replaced with the insert during Golden Gate assembly, as well as NotI and SbfI sites, strategically placed to enable separation of the desired library insert from the backbone prior to sequencing. CcdB and sfGFP provided counterselection and visualization of colonies resulting from undigested vector. (Cloning into any E. coli vector will suffice however, so long as unique library members are associated with a unique relatively short (20-50 bp) sequence.) Resulting colonies each contained many copies of a unique plasmid species encoding a UMI-tagged library member. To obtain good coverage of each sequence and UMI with nanopore sequencing, the resulting library was downsampled by only harvesting ~20-fold fewer colonies than the expected number of reads, amounting to 100 – 200 thousand colonies for a standard MinION flow cell. Plasmid DNA from this library was then miniprepped, digested (for example with NotI-HF or NotI-HF / SbfI-HF (both NEB) if using plasmid pGR554), and gel extracted prior to sequencing.
[0092] Construction of the library for HTS dataset 8a was performed using a similar approach, but with a yeast expression vector, and with UMIs and library members inserted into the vector in two distinct steps so that UMIs are present at a single location in the plasmid and would not need to be immediately adjacent to library members, mitigating any potential effects of UMIs on expression. In brief, libraries of UMIs generated via PCR usingprimer pair 8 were cloned into plasmid pUMI by Golden Gate assembly with BsmBI-v2 and E. coli electrotransformation to generate an intermediate UMI library. Different variants of pUMIs with unique, known 7-nucleotide barcodes were used for each library to allow multiplexing. These libraries were then used as the vector into which evolved TrpB library members, PCR amplified using primer pair 9, were cloned via standard restriction cloning, with inserts digested with PaqCI and vectors digested with BsaI-HFv2. Isothermal ligation (1 hour 37 °C, 20 mins 65 °C) was then performed with T4 ligase in 40 μL reaction volumes, AMPure bead purified, and transformed into high efficiency E. coli. This was carried out for 6 unique libraries: downsampled in yeast with selection, downsampled in yeast without selection, and not downsampled in yeast, each for both timepoints (generation 350 and generation 540) chosen from TrpB evolution. The intermediate UMI libraries were generated at a library size of >100-fold larger than the desired final library size to minimize the chance that distinct library members would be tagged with identical UMIs.
[0093] The second method used for generating libraries for high accuracy nanopore sequencing was adapted from methods described in Volden et al. (72), Oliynyk and Church (73), and Zhang and Tanner (74). It involved circularization of the target sequence and use of this circularized product as template for rolling circle amplification (RCA) using strand displacing DNA polymerases (Fig.20). In brief, UMI-tagged PCR products were generated as described above, albeit with complementary Type IIS cut sites (BsaI or PaqCI) on the forward and reverse primers used during PCR amplification such that the two ends of the amplicon ligated to each other during a Golden Gate assembly reaction, forming a circular product. Due to the large amount of template DNA required for rolling circle amplification, a large amount of amplicon was used in the circularization reaction, typically 1-2 pmols in 100 μL Golden Gate assembly reactions. Oliynyk and Church (73) provide a useful discussion of relevant considerations for such circularization reactions. An isothermal Golden Gate assembly reaction was then performed to circularize the amplicon library.
[0094] Following Golden Gate assembly, uncircularized DNA was digested with lambda exonuclease (NEB), exonuclease I (NEB), and exonuclease III (NEB) in a 10:10:1 ratio, which was added directly to the Golden Gate reaction in a 1:10 exonuclease mixture to sample ratio. This reaction was incubated at 37 °C for 45 min, then 80 °C for 15 min. The reaction was then AMPure bead purified with a 0.7:1 bead:sample ratio, eluting in a maximum of 10 μL of water, and the resulting circularized library was used in a rolling circle amplification reaction with the following components combined on ice: 4 μL 10X NEB buffer 4 (NEB) 4.8 μL 10 mM dNTPs2.64 μL 5 U / μL Bsu DNA Polymerase, Large Fragment (NEB) 1.6 μL 10 μg / μLT4 gene 32 protein (NEB) 0.2 – 1 μg circularized DNA library RCA primers, 2 μM each (must bind internal to first primer set, e.g. primer pair 7) Water to 40 μL.
[0095] This reaction was incubated at 37 °C for 3 hours. SDS-containing loading dye was added directly to the reaction, and the entire sample was run on an agarose gel. Bands corresponding to 3x–6x concatemers were gel extracted, and purified DNA was used for nanopore sequencing. Unlike RCA using Phi29, this method enabled size selection of specific repeat numbers and did not require a ‘debranching’ step but required large amounts of input DNA and in our experience suffered from high sensitivity to DNA contamination. Use of Phi29 with random hexamers, followed by debranching, is therefore a reasonable alternative.
[0096] Following construction of in vivo downsampled UMI PCR or RCA libraries, nanopore library preparation and sequencing was performed using the most up-to-date ligation sequencing kit (e.g. LSK-114) and flow cell (e.g. R10.4.1), following manufacturer instructions, with two exceptions. First, ½ volumes (but unaltered DNA input) for end prep and ligation reactions were used and second, FFPE Repair Mix was not used during end prep reactions.
[0097] All relevant high throughput sequencing datasets are made available on the NCBI sequence read archive (SRA), accession number PRJNA1050257.
[0098] Error rate measurement by mutation accumulation
[0099] Following a polymerase replacement integration transformation, colonies were picked into liquid media of the same media formulation as that used for selection after transformation, then grown to saturation. The resulting saturated culture was miniprepped to serve as the 0th passage, p0. The culture was then also propagated in the appropriate growth medium for maintenance of the orthogonal plasmid, typically SC -L, using a dilution factor d (typically 128, 256, or 512) that is consistent throughout the experiment. At least one additional saturated culture from this time course was miniprepped to serve as the tth passage, pt. We used only two passages / timepoints in all mutation accumulation error rate measurement experiments except that which produced HTS dataset 6. The number of generations, g, that separated p0 from pt was used to calculate the mutation rate and can be approximated as^௧ൌ ^ ൈ ^^^ଶ^Ǥ
[0100] We note that this approximation assumes no cell death and equivalent saturation at each passage. Cultures were manually propagated several times until a total number of generations of at least 50 was reached. Miniprepped yeast DNA for each timepoint was used as template DNA for high throughput sequencing, and custom scripts, organized within the Maple pipeline, were used to calculate mutation rates.
[0101] Mutation rates were calculated as the rate of accumulation of a mutation type (all substitutions, individual substitutions, insertions, or deletions), j. We denote μjas this rate calculated for mutation type j. High accuracy HTS was used to first obtain cj,t, the total counts of mutation j (e.g., A to T) among all sequences in passage t. For substitution mutations, to account for the influence of variable A / T / G / C content in the sequence being analyzed, this count is normalized to obtainthe expected count for an idealized sequence with a 1:1:1:1 A:T:G:C ratio. For a substitution at a particular nucleotide where the nucleotide occurs w times within a reference sequence of length L, nj,tis calculated as
[0102] The total normalized count of all substitution mutations at each timepoint is then calculated as the sum of all nj,tfor all twelve substitution types. However, for insertions and deletions, is not normalized in this way, and is instead equivalent to cj,t, the total number of nucleotides inserted or deleted among all sequences for that timepoint.
[0103] To obtain the per-nucleotide, per-generation mutation rate μj, we used the total number of sequences analyzed for each timepoint, st, to calculate the per-nucleotide frequency of mutation j for each timepoint. We then use linear regression on these normalized per-nucleotide frequencies according to
[0104] where μjand bjare the slope and intercept, respectively, of the best fit line for all t timepoints. When the number of generations between the 0thtimepoint and the initiation of mutagenesis (typically polymerase replacement) is accurately estimated and no mutations fully fixed within the population prior to the first timepoint, bjshould be close to 0. Regardless, we do not report bj, as it has no bearing on μjacross experiments. We report μjas the per-base per-generation rate of accumulation of mutation type j. Mutation tabulation was performed by the script mutation_analysis.py and all other operations related to mutation rate calculation were performed by the script plot_mutation_rates.py, both of whichare contained within the Maple pipeline. All reported mutation rates were calculated using mutation tabulation within a sequence region that was not under functional selection.
[0105] TP-DNAP1 library selection and screening
[0106] Error prone TP-DNAP1 libraries were cloned in E. coli. Mutagenesis was validated by Sanger sequencing of 8-12 individual clones. See below for a summary of these libraries:
[0107] Following TP-DNAP1 replacement library construction, resulting purified plasmid DNA was used for a polymerase replacement integration transformation. Prior to transformation, OR-Y488 was grown up in SC -L + 1 mg / L 5-fluoroorotic acid (US Biological) for counterselection against cells that had reverted the inactivating mutation in ura3* by chance. Following library transformation, plating on media selecting for cells that had replaced the wild type TP-DNAP1 with library variants, and 48 hour incubation, colonies were harvested in bulk and immediately plated onto SC -LU plates. These plates were then incubated for 48 hours, and resulting colonies were either picked into liquid media for mutation rate or frequency characterization (by mutation accumulation or fluctuation assays, respectively) or were harvested in bulk and immediately plated onto SC -LUW media. After colony formation, colonies were either picked into liquid media for mutation rate characterization by mutation accumulation or were harvested in bulk, miniprepped for gDNA isolation, and used as template for a non-mutagenic PCR to generate amplicons for Golden Gate assembly into pGR554, which was then retransformed into OR-Y488 to repeat the selection and perform mutation rate screening.
[0108] The fluctuation test was performed as follows. Following transformation of the epPCR 1 TP-DNAP1 library into OR-Y488 and selection for ura3* reversion on solid media,individual colonies were picked from this plate, inoculated into 500 μL SC -LU media in a 96 well block, and grown to saturation. Cultures were then passaged 1:10,000 into 200 μL SC - LU, 12 replicates per each individual colony, and grown to saturation. Cultures were centrifuged, washed with 0.9% NaCl, and pellets were resuspended in 35 μL 0.9% NaCl.10 μL of each resuspension was then plated onto SC -LUW plates. A subset of cultures were titered and plated on SC -LU plates to estimate population size. Plated cells were allowed to grow for 4 days, and revertants on each spot were counted. Counts were used to estimate the m value using the FALCOR online web tool (lianglab.brocku.ca / FALCOR / ) and the Ma- Sandri-Sarkar Maximum Likelihood Estimator. Mutation frequency was calculated from this m value as previously described (42), using a target size of 1 (only one mutation is capable of restoring Trp5 activity). Copy number was not considered and therefore per base substitution rate was not calculated using this method.
[0109] Copy number measurement
[0110] Quantitative PCR (qPCR) was used to determine p1 copy number. gDNA was isolated from samples according to the 1.5 mL DNA extraction protocol referred to above. qPCR reactions were performed in 10 μL volumes using the Sybr Powerup Master Mix (ThermoFisher), with 1-5 ng of gDNA template. Reactions were run at ‘standard speed’ on a ThermoFisher Quantstudio 6.
[0111] The reactions measured the amplification of genomic GAL1, and p1-encoded LEU2 using primer pairs 16 and 17, respectively (Table 5), yielding cycle threshold (Ct) values for each sample. A standard curve was generated correlating Ctwith DNA quantity using reactions containing known quantities of plasmids encoding GAL1 and LEU2. Standard curves were used to validate the assumption that the two primer pairs had the same amplification efficiency. Relative copy numbers of GAL1 and LEU2 in each sample were determined from the standard curves, and absolute p1 copy numbers were normalized to genomic copy numbers to obtain absolute per-cell p1 copy number by dividing LEU2 relative copy numbers by GAL1 relative copy numbers for each sample.
[0112] TrpB evolution
[0113] Plasmid pGR595 (TP-DNAP1 BadBoy2) was first transformed into yeast strain OR- Y484 according to the polymerase replacement procedure described above. The resulting strain (OR-Y538) was then transformed with plasmid pGR438 (TrpB with lineage barcodes), following the p1 integration procedure described above, plating on SC -L media. ~400 resulting colonies were harvested together and passaged into 512 μL SC -L media in all wells of a 96-well block, grown to saturation, then again passaged 1:1024 into SC -L media. DNA extracted from these resulting cultures served as passage / timepoint 0, which weapproximate to be ~50 generations from TrpB p1 integration. These cultures were also passaged 1:1024 (0.5 μL into 512 μL) for all passages in the experiment into growth media and with timepoints taken as described in Table 2. DNA extraction for timepoints was performed by combining all 96 saturated cultures for a specific timepoint and extracting DNA from the pooled cultures according to the 1.5 mL DNA extraction protocol referred to above.
[0114] Selection pressures shown in Table 2 and Fig.2B for visualization purposes were derived from the concentrations of Trp and indole in the growth media used for each passage. These two components are inversely related to the selection pressure supplied by each component, scaled to range from 0 to 0.5 such that the maximum concentration used for both Trp and indole would yield a selection pressure of 0, media without Trp and with the maximum concentration of indole would yield a selection pressure of 0.5, and media without either component would yield a selection pressure of 1. The three selection periods ‘no selection’, ‘mostly positive’ and ‘mostly purifying’ were characterized as such based on the selection pressure used during that period. The initial ‘no selection’ period only included the minimum selection pressure (0). The subsequent ‘mostly positive’ period proceeded until the populations were exposed to the highest selection pressure for the first time. Subsequently, populations did not need to further adapt to more stringent selection conditions, so we refer to the following period as ‘mostly purifying’.
[0115] To downsample evolved populations for cloning UMI-tagged TrpB libraries, yeast cultures from the passages corresponding to generation 340 and 510 were inoculated into SC -L media from glycerol stock and grown to saturation. Saturated cultures were then combined, and multiple serial dilutions of both cultures were plated onto SC -L media. Plates derived from generation 340 and 510 with ~3700 and ~1700 colonies respectively were harvested. These cultures were passaged 1:1000 into 2 mL of either SC -LW + 400 μM indole (Sigma-Aldrich) (selective) or SC -L (nonselective) growth media and allowed to grow to saturation. This process was repeated twice more for each of the two media. All three of the resulting saturated cultures for each selective and nonselective conditions were pooled and DNA was extracted to serve as template for cloning the TrpB library fitness assay library, in addition to DNA extracted from generation 340 and 510 without downsampling. Note that this downsampling performed in yeast preceded the downsampling in E. coli necessary for proper sequencing coverage.
[0116] Pooled TrpB fitness assay
[0117] UMI-tagged evolved TrpB libraries and equivalent plasmids expressing two control TrpB sequences (TmTriple and TrpB-003-1-A) were transformed into yeast strain OR-Y260 and plated onto -LH media, resulting in 40-fold coverage of the ~120 thousand memberlibrary. Colonies were harvested and the library was spiked with each of the two control TrpB-expressing yeast strains at a 1:1000 control:library ratio. DNA was extracted from the resulting library to serve as the 0thtimepoint. This library was also passaged 1:100 into 50 mL of either SC -L (nonselective), SC -LW + 400 μM indole (weakly selective), or SC -LW + 25 μM indole (strongly selective) and grown to saturation. This passaging and growth was repeated for each of the three growth conditions five times for a total of six passages. Of these, DNA was extracted from passages 1, 2, 5, and 6 to serve as additional timepoints. DNA from all timepoints were used as templates for PCR amplification of only the UMI and barcode regions for high throughput sequencing.
[0118] Enrichment scores were calculated following the procedure described in Rubin et al. (75), albeit with two modifications to account for disparate sequencing coverage over multiple timepoints. First, counts of each UMI at each timepoint were normalized to the average count of all UMIs that persisted throughout all timepoints prior to log transformation and weighted linear regression. Second, regression weights were multiplied by this average count. All enrichment calculations were performed by the Maple pipeline within the script enrichment.py.
[0119] High-throughput sequencing analysis
[0120] High accuracy consensus sequence generation, alignment, demultiplexing, genotype identification, mutation rate analysis, and basic dataset visualization were performed by version v0.10.4 of the Maple pipeline, which uses the Snakemake workflow management library, and is available on github.com / gordonrix / maple. Parameters and settings for Maple analyses for each dataset and all other code used for analysis are available on Github (github.com / liusynevolab / OrthoRep_Rix_2023).
[0121] TrpB fitness prediction with TranceptEVE (66)
[0122] A multiple sequence alignment (MSA) of natural TrpB subunits was created using 5 iterations of jackhmmer (76) to query the UniRef100 database with a bitscore of 0.9. Columns with more than 20% gaps were ignored, and we used a theta parameter of 0.8 to downweight sequences with more than 80% sequence homology, as described in Hopf et al. (77). We trained 4 separate EVE models using this MSA and used them to calculate the log probabilities of the mutated sequences. These scores were ensembled with the Tranception log probabilities for each sequence by taking a weighted average of 60% Tranception score and 40% EVE score.
[0123] Null hypothesis sequence simulation and analysis
[0124] A dataset of TrpB variants that contained random mutations representative of biases due to the relevant mutation preferences and wild type sequence, but that were not subject to selective forces, was generated through a simple simulation. First, the number of synonymous mutations within the ORF of each unique evolved TrpB sequence was approximated as the number of nucleotide mutations minus the number of nonsynonymous mutations. For each real unique genotype, a corresponding simulated genotype was generated by starting from the wild type TrpB sequence used in the evolution experiment and stochastically sampling nucleotide mutations with probabilities determined by the mutation rates of the same polymerase used for TrpB evolution (BadBoy2) until the same number of synonymous mutations as the evolved genotype was reached. Additional information such as the timepoint from which the real sequence was identified and the count of the genotype was also replicated for the corresponding simulated sequence. To account for some minor strand-dependent mutational biases, mutation rates and spectrum calculated for a sequence in the same position and orientation (relative to the LEU2 gene) as TrpB were used to generate the simulated sequences.
[0125] To validate our assumption that synonymous mutations were minimally subject to selective forces and could therefore be used as a neutral molecular clock in the simulation of mutation accumulation, we computationally generated mutagenized sequences using a bulk mutation process. Mutagenesis was carried out as above, but the number of synonymous mutations present in the evolved sequences was not used to determine the extent of mutagenesis. Instead, the projected average number of mutations per sequence was estimated using the overall nucleotide substitution rate of BadBoy2 and the estimated number of generations at each timepoint, and 10,000 sequences per timepoint were mutagenized to match this average of mutations per sequence. The resulting distribution of mutations per sequence in these sequences is shown in Fig.22 as ‘projected’. All other figures referring to in silico mutagenized sequences refer to the simulation described in the previous paragraph.
[0126] Simulated and evolved sequences were analyzed identically. Isoelectric point and hydrophobicity index (gravy) were both calculated using the ProtParam module in BioPython (78). Mesophilic adaptation mutations were selected from Haney et al. (13) as the inverse of the 17 mesophile to thermophile “Replacements most biased in number” with P<0.005. Alternative sets of amino acid replacements were not evaluated.
[0127] Computational lineage downsampling
[0128] Lineage barcodes were identified using the demultiplexing feature within the Maple pipeline. The 100 most frequently observed lineage barcodes only from the first timepoint ofTrpB evolution were used for all lineage analyses. Lineage barcodes for all remaining timepoints were assigned from this list of 100 barcodes, allowing for a nucleotide hamming distance of 1. For global covariation analysis, all barcodes appearing in at least 100 sequences were identified, and 100 sequences from each lineage were randomly extracted for analysis of mutation frequencies. Covariation with specific mutations was performed similarly, except that only sequences that contained the specific mutation and were identified from specified timepoints were considered prior to randomly extracting sequences for analysis of mutation frequencies. To ensure random sampling did not bias results, this process was performed multiple times with virtually identical results. Example 3: Results and Discussion
[0129] OrthoRep engineering
[0130] The current state-of-the-art OrthoRep system uses TP-DNAP1-4-2 as the error-prone orthogonal DNAP. Besides its suboptimal error rate of 10-5s.p.b., TP-DNAP1-4-2 also haslow replicative activity (Fig.6) and exhibits a heavily transition-biased mutation spectrum (Fig.7), suppressing the impact of point mutations on amino acid sequence during protein evolution (Fig.8). We carried out a directed evolution campaign on TP-DNAP1 to increase OrthoRep’s overall error rate, transversion rate, and activity. A selection strain, OR-Y488, was engineered to contain a p1 plasmid (p1-ura3*-trp5*) encoding two auxotrophic marker genes, ura3 and trp5, each specifically disabled via an active site missense mutation whose sole option for functional reversion is a transversion (Fig.1B and Fig.9). TP-DNAP1s with the highest transversion rates should restore URA3 and TRP5 most frequently, resulting in their enrichment from genomically-integrated TP-DNAP1 libraries (see Fig.10) when OR- Y488 is grown in the absence of exogenous uracil or tryptophan. Selection was designed to occur in two sequential stages, first for URA3 restoration and then for TRP5 restoration, to suppress the enrichment of low error rate TP-DNAP1 variants in revertants that stochastically emerge from the long tail of the Luria-Delbrück distribution (50).
[0131] To precisely guide the TP-DNAP1 directed evolution campaign, we developed a 30 mutation accumulation assay (51) for p1 and coupled it with high throughput sequencing (HTS) of p1 amplicons using the Oxford Nanopore Technologies (ONT) platform (52), allowing us to accurately determine the rate for any individual type of mutation (Fig.1B). In this assay, a strain containing p1 replicated by a given TP-DNAP1 variant is grown for a set number of generations. A region of p1 not under selection is sequenced at two or more timepoints using unique molecular identifiers (UMIs) for error correction (53, 54), and the rate of change in the number of mutations per position is calculated individually for all typesof mutations to fully describe the overall mutation rate and mutation preferences of the TP- DNAP1. To facilitate rapid characterization, we developed a custom analysis pipeline that could carry out most of the analysis steps autonomously (Fig.11). This pipeline, Mutation Analysis for Parallel Laboratory Evolution or Maple, performs consensus sequence generation, demultiplexing, mutation identification, mutation rate analysis, and many other operations to generate a collection of visualizations and data tables that accelerate analysis of mutation-rich sequencing datasets while minimizing user input.
[0132] With the genetic selection, HTS-based mutation accumulation measurement pipeline, and Maple in place, we carried out the TP-DNAP1 directed evolution campaign over five rounds (Fig.1B), yielding a collection of TP-DNAP1 variants (Fig.1C and Table 1) with broad mutational spectra and error rates up to and exceeding 10-4s.p.b. (Figs. 1D-F). Thefull course of the directed evolution campaigns is described in Figs.12-18. Here, we emphasize three key outcomes. First, because our mutation rate measurement pipeline obtained complete mutation rate and preference data for our evolved TP-DNAP1 variants, the use of OrthoRep in driving continuous gene evolution experiments comes with the ability to generate accurate null models of sequence change in the absence of selective forces against which evolutionary trajectories and outcomes can be compared. Second, we saw evidence suggestive of nearing gene error thresholds at OrthoRep’s new 10-4s.p.b. mutationrates. We found that mutation frequencies were consistently higher in regions of p1 not under selection compared to regions encoding genes under selection (Figs.19A-B), and that the ratio between the two was highest for the lowest fidelity TP-DNAP1 variants (Fig.19C). This implies that the 10-4s.p.b. mutation rate of TP-DNAP1 variants BadBoy1, BadBoy2, andBadBoy3 (Fig.1D) quickly degrades the function of genes when purifying selection is absent. Furthermore, an extremely error-prone 1.7 × 10-4s.p.b. TP-DNAP1 variant weisolated (BB-5k) did not durably maintain p1 in two of four biological replicates under selection over ~120 generations of mutation accumulation, possibly because BB-5k exerted an excessive mutational load on the selection marker used to maintain p1 leading to mutational meltdown. Third, in conjunction with past efforts, the TP-DNAP1s obtained here complete a set of OrthoRep systems evenly spanning a range of ~5 orders-of-magnitude, from ~10-9s.p.b., similar to the mutation rate of modern cellular genomes, up to ~10-4s.p.b.,which is likely in the regime where the error thresholds of individual genes reside, above which gene mutational meltdown occurs, and near which maximal gene adaptation rates can be reached (55,56). Thus, the overall ecosystem of OrthoRep systems should not only allow for the rapid continuous evolution of chosen genes in vivo, but also investigations on the detailed role of mutation rates and error thresholds in molecular evolution for which theory is abundant but experiment is sparse.
[0133] Extensive divergence of a conditionally essential gene on laboratory timescales
[0134] Our new OrthoRep systems should be capable of driving rapid evolution of chosen genes regardless of the type of selection imposed. We encoded the β-subunit of Thermotoga maritima’s tryptophan synthase (TrpB) onto p1 in a tryptophan auxotroph wherep1 is exclusively replicated by BadBoy2 at a mutation rate of 1.4 × 10-4s.p.b. (Fig.2A).TrpB condenses indole with serine to yield tryptophan (Trp), but T. maritima TrpB is maladapted for this standalone reaction, since it normally functions in complex with TrpA (57,58). Therefore, cells grown in the absence of Trp need to evolve improved TrpB activity to propagate, allowing TrpB to serve as the subject of an extended evolution experiment that included a range of selection pressures (Fig.2B and Table 2).
[0135] Table 2: TrpB selection schedule*reported generations correspond to the culture after it has reached saturation
[0136] We designed our evolution experiment to prioritize sequence divergence and diversity to maximize the amount of evolutionary information that could later be extracted. The evolution experiment was therefore run for ~540 generations (~3 months) at a scale of 96 independent replicate 500 μL cultures.1:1024 (10 generation) transfers into fresh growth medium were made every one or two days, depending on cell density, following a passaging schedule that included all types of selection pressures (Fig.2B): an initial period of evolution without selection (‘no selection’), then a period of adaptation when strong selection pressure was applied and functional improvements in TrpB’s function were observed (‘mostly positive selection’), followed by a long period characterized by mostly purifying selection where adapted TrpBs were pressured to maintain the fitness they evolved (‘mostly purifying selection’) (Fig.2B). The ‘mostly purifying’ period also included some brief episodes of relaxed or removed selection pressure, which we introduced with the intention of promoting sequence divergence. We collected DNA from cells at 15 timepoints throughout the ~540- generation evolution experiment (Fig.2B) and used a rolling circle amplification-based sequencing strategy in conjunction with HTS and Maple to analyze the TrpB sequences sampled from these timepoints (See Fig.20).
[0137] Overall, we observed a monotonic increase in both the average number of mutations and diversity (as measured by pairwise hamming distances) in TrpB throughout all phases of evolution (Figs.2C-E), resulting in a large number of distinct evolutionary outcomes (Fig. 21). The rate at which mutations accumulated in the population was highest in the ‘no selection’ period (~0.15 amino acid changes per generation) followed by the ‘mostly positive selection’ and ‘mostly purifying selection’ periods (~0.039 and ~0.021 amino acid changes per generation, respectively) (Table 3). Notably, the ‘mostly positive selection’ period did not have the highest rate of mutation accumulation even though there was substantial adaptation in producing TrpBs that supported cell growth in the absence of Trp and presence of moderate concentrations of indole. This suggests that BadBoy2’s error rate was high enough to consistently “saturate” positive selection with an overabundance of beneficial mutations in TrpB, predicting the broader power of upgraded OrthoRep mutation rates in biomolecular evolution applications. Notably, mutation accumulation in the ‘mostly purifyingselection’ period was appreciable yet substantially slower than in the ‘no selection’ and ‘mostly positive selection’ periods. This suggests that BadBoy2’s error rate was high enough to constantly test the constraints of structure and function, signaling the general potential of OrthoRep in uncovering biological forces governing how genes and biomolecules operate. At the end of the evolution experiment, each TrpB sequence had an average of 20.6 amino acid and 44.5 nucleotide mutations (Table 3). In the last two timepoints, over 1800 sequences (~5%) had accumulated more than 30 amino acid changes from the ancestral 398 amino acid wt TrpB. The distribution of pairwise hamming distances showed that sequences had substantially diverged from each other (Figs.2C-D) such that in the final timepoint, 24% of sequences were separated from each other by 40 amino acids or more, including more than 4000 sequence pairs differing by a pairwise amino acid hamming distance of at least 60 (Fig.2E, inset). This level of sequence divergence (>15%) approximates that between human and mouse essential gene orthologs (35). Indeed, our experiment demonstrates that we can compress extensive and complex gene evolution processes into laboratory timescales, resulting in the generation of highly diverse sequence families shaped by varied selection conditions over long mutational pathways.
[0138] Table 3: Summary mutation statistics for TrpB evolution
[0139] Table 3 (cont’d)
[0140] General structural and functional constraints
[0141] ~500,000 sequences of TrpB with an average of 13.1 amino acid replacements each were captured over the evolution experiment, and over 90% of those sequences were unique (Table 3). With such a diverse evolutionary dataset, patterns of conservation should contain structural and functional constraints defining TrpB. To test this notion, we used an AlphaFold structure (33) and knowledge from previous studies on TrpB (60, 61) to first categorize each residue in TrpB according to its general structural or functional role, as outlined in Fig.3A. We then asked whether different categories showed different levels of conservation. We immediately noticed a congruence between relative conservation and buried residues, revealing the well-known importance of a buried hydrophobic core in protein folding (Fig.3B). We also noticed that residues within 5 Å of TrpB’s active site were highly conserved. Additionally, there was a relative abundance of amino acid replacements at certain positions in the COMM domain, suggesting that it was a target of adaptation.
[0142] To improve our resolution in such observations, we generated a simulated dataset of TrpB mutants where each sequence accumulated mutations from the exact encoding of wt TrpB using BadBoy2’s fully described mutation biases. For each simulated sequence, mutation accumulation was stopped once it matched the number of synonymous mutations of a corresponding sequence from the real dataset. The simulated dataset serves as the null model where patterns in evolved TrpB sequences are simply a reflection of the mutation preferences of BadBoy2 and codon usage of wt TrpB. Under the assumption that synonymous mutations have no effect on fitness – we confirm this assumption for the experiment at hand in Fig.22 – the nonsynonymous differences between the real dataset and the simulated dataset contain the influence of selective forces. An excess of nonsynonymous changes in the real dataset compared to the simulated dataset is therefore an indication of positive selection, while the opposite signifies purifying selection. As shown in Fig.3C, at the generation 70 timepoint, the real dataset has a paucity of nonsynonymous mutations per sequence in the active site region and buried residues and an excess of nonsynonymous mutations per sequence in the COMM domain. Generation 70 is during the first phase of positive selection for TrpB’s operation as a standalone enzyme capable of generating tryptophan (Fig.2B), so this timepoint is most likely to reveal signatures of adaptation. (This is supported by Fig.23, which shows the overall mutation accumulation dynamics; generation 70 is the timepoint with the greatest excess of nonsynonymous mutations relative to the simulated dataset among the timepoints for which selection for TrpB function had been applied.) At generation 70, we find that positive selection had enriched mutations to buried and COMM domain residues while purifying selection had already removed changes to the active site region. Our detection of the COMM domain as a focus ofadaptation can be explained since TrpB is a well-studied enzyme. The COMM domain mediates allosteric activation of TrpB by TrpA (57). As our evolution experiment required TrpB to operate in a standalone manner without TrpA, the remodeling of allosteric networks through the COMM domain was an expected means to adaptation in line with previous studies of engineered TrpB standalone activity (60, 61). Had TrpB not been well-studied, the detection of this critical region for adaptation from the evolutionary information could have suggested such an explanation.
[0143] We also considered the final timepoint of the evolution experiment in detail. By generation 540, the influence of purifying selection had dominated, as evidenced by the paucity of mutations in the real data compared to the simulated (Fig.3C) or, similarly, the consistently low rate of nonsynonymous versus synonymous mutation accumulation after generation 90 (Figs.22-23). Although purifying selection constrained all regions of TrpB, some were clearly more constrained than others. For example, the active site region had almost no mutations (maximally 1 or 2 but mostly 0) and deviated from the simulated mutant distribution more than all other regions. Buried residues also had substantially fewer nonsynonymous mutations than the simulated dataset. In contrast, the effect of purifying selection was less pronounced on surface residues, reflecting the relative tolerance of protein surfaces to mutation. Surprisingly, this also applied to the newly solvent-exposed β-α interface region. In the absence of the α-subunit, this region should be more solvent- exposed than in TrpB’s native context. The fact that there was little noticeable difference in the effects of selection on the β-α and β-β interfaces suggests that solvent-exposure of this region had minimal impact on TrpB fitness.
[0144] Isoelectric point evolution for intracellular compatibility
[0145] We examined the 20 residues that were mutated in greatest excess across real evolved sequences relative to simulated sequences in the first two timepoints (Fig.3D). Despite the entire wt TrpB protein containing only 19 arginine residues in total (5%), arginine constituted 8 of these 20 most frequently mutated residues by generation 50. This came as a surprise because this enrichment occurred even before selection for TrpB function was imposed (Fig.2B). This led us to hypothesize charge optimization as a driving selective force, because charge could influence not only TrpB function itself, but also the cellular environment within which TrpB operated. To evaluate this hypothesis, we calculated the isoelectric point (pI) of sequences throughout the evolution experiment and examined its distribution over time (Fig.3E). Indeed, we found that the pI of sequences was significantly lower at the end of the experiment than early in the experiment (p<0.0001, Mann-Whitney U test). Comparison of pI change against simulation corroborates that this effect was driven by positive selection. A similar analysis of hydrophobicity revealed a modest decrease inhydrophobicity over time for simulated sequences that was mitigated by selection in the real data (Fig.23). The change in hydrophobicity for the real sequences throughout the experiment was less pronounced than the change in pI, however, suggesting that charge optimization, and not polarity in general, was the dominant selective force. Notably, the majority of the shift in the pI distribution occurred in the latter half of the experiment (Fig.3E), highlighting the importance of sustained rapid mutagenesis over long periods of evolution to embed such presumably subtle selective forces into the data.
[0146] TrpB’s pI evolved to be comfortably below the typical yeast cytosolic pH of 6.8 to 7.2 (62), which is consistent with the notion that intracellular proteins (49, 63) prefer to be negatively charged to minimize large-scale clustering with RNAs and other proteins. One possible mechanism by which this preference could have driven the observed adaptation is through its influence on TrpB function itself, for example by increasing the diffusivity of the enzyme (64). Another mechanism by which a preference for negative charge in TrpB could have been adaptive is by lessening its perturbation on other entities in the cell, for example by preventing spurious association or aggregation that would disturb the function of the proteome (49). Our data does not exclude either mechanism but suggests that the latter mechanism is present. In generations 0- 50 of the evolution experiment, TrpB was not under selection for function as excess Trp was supplied to the growth media. Indeed, HTS at the end of 50 generations detected the presence of many stop codons, which were mostly eliminated soon after positive selection for TrpB adaptation had been imposed (for example, HTS at the end of generation 70) (Fig.3D). Yet at the end of 50 generations, pIs were significantly lower than that of the null model (p<0.0001). Our observation of charge optimization even when there was no selection for the enzyme’s function demonstrates that the intracellular context can impose constraints on the physicochemical properties of proteins independent of its primary molecular function. It also highlights the value of evolving proteins in vivo where subtle constraints dictating intracellular compatibility can both be revealed and included in the evolutionary optimization of protein function.
[0147] Thermoadaptation
[0148] Given that our parental T. maritima TrpB was from a thermophile but needed to evolve standalone activity in a mesophile, we looked for statistical evidence of thermoadaptation in our evolved sequences. Haney et al. studied the patterns of amino acid replacements between natural orthologous proteins in mesophilic versus thermophilic organisms and found 17 amino acid replacements that distinguished the mesophilic variants from the thermophilic variants at homologous positions in multiple sequence alignments with high confidence (13). When we evaluated the frequency of these 17 amino acid replacements among all mutations in our evolution experiment’s outcomes, we found thatreplacements in the mesophilic direction were enriched (Fig.25). As before, this illustrates the ability of extensive gene evolution to reveal selective forces through the evolutionary information embedded into the resulting diversity.
[0149] Networks of coupled mutations
[0150] In addition to general selective forces, we investigated finer patterns in the outcomes of TrpB evolution with the expectation that these may reveal coevolving networks of amino acids responsible for adaptation. At the beginning of our evolution experiment, we had included short barcodes adjacent to the TrpB sequence integrated onto p1. This allowed us to 1) isolate the largest clades for analysis, since these should correspond to the fittest sequences, and 2) reduce the contribution of phylogeny by computationally limiting the number of sequences analyzed per clade (Fig.4A). The latter increases the signature of mutations independently discovered across multiple clades, favoring the detection of mutations whose cooccurrence was functionally significant. Specifically, we considered clades whose barcode had at least 100 associated sequences — there were 93 such clades — and randomly downsampled to 100 sequences for clades whose members exceeded this number.
[0151] This analysis revealed that some sets of mutations frequently cooccur (Fig.4B). In one particularly striking example, the A20V mutation was found to cooccur at a frequency above ~0.7 with a set of 5 other mutations (A118V, F122S, V245A, I271V, T292S) in 8 distinct lineages in later timepoints, implying a strong relationship among these mutations (Figs.4C-D). This set of mutations, as well as other sets identified, are spread throughout the structure of TrpB, in line with recent studies demonstrating the prevalence of structurally distributed allosterically activating mutations in TrpB and other proteins (6, 60, 61, 65). Among these five mutations is the T292S mutation, which was previously identified in a TrpB directed evolution campaign as highly activating, alone conferring a >5-fold increase in the standalone catalytic efficiency of TrpB (20). Intriguingly, the association among these five mutations was sensitive to their specific identities. For example, the A20T mutation was individually present at a higher frequency than A20V among all sequences (Fig.21D) but was not associated with any specific other mutation at a frequency above ~0.4 (Fig.4D). The overall picture from this analysis suggests the existence of heavily enriched individual mutations that are broadly activating across many sequence contexts and individual mutations that are only activating in combination with other mutations, implying long range epistatic interactions.
[0152] Fitness of TrpBs and their predictability
[0153] Accurately modeling the fitness landscapes of proteins is a major goal of ML. However, it is known that ML models are biased towards favoring sequences that are more similar to the natural sequences on which they were trained (66), and it remains unclear to what extent they can predict the function of sequences that are many mutations away from these natural sequences. It is also unclear whether ML can model how mutant sequences will perform on new functions that deviate from natural functions, in service of bioengineering goals such as enzyme and antibody engineering. To provide insight into these questions, we tested whether an advanced ML model could predict the fitness of our evolutionary outcomes.
[0154] To gain high-resolution fitness information on evolved TrpBs, we first profiled evolved variants in a high-throughput enrichment assay (Fig.5A). We cloned a library of ~100,000 TrpB sequences isolated from our evolution experiment into a standard yeast plasmid that would not be subject to hypermutation by OrthoRep, transformed this library into yeast, applied selection for TrpB function, and tracked the enrichment or depletion of individual variants via HTS. Included in this library were two previously engineered control TrpB variants known to be either highly functional (TrpB-003-A) or nearly nonfunctional (TmTriple) in the context of yeast Trp production growth complementation (Fig.5B)(38). We evaluated Trp production by members of this library using three distinct growth conditions: Trp- supplemented media (no selection), media lacking Trp with a high concentration of indole (400 μM, weak selection), and media lacking Trp with a low concentration of indole (25 μM, strong selection). We tracked the abundance of library members in replicate yeast transformations of the same library over 4 timepoints taken at the beginning and end of 6 passages to obtain fitness scores. We found that fitness scores above a threshold (enrichment score > -5) were well correlated among replicates for both weak and strong selection conditions (Pearson correlation = 0.79 and 0.84, respectively), but not for nonselective conditions where Trp was present (Fig.5C), confirming the reliability of the assay for highly functional TrpB variants. Thousands of multi-mutation sequences at least asfunctional as the previously engineered high-fitness TrpB-003-1-A, with a kcat / KMof 1.4 ×-1 -1 105M s (38), were identified (Fig.5D). We also observed that a large fraction of sequences had low activity (Fig.5D), which likely owes to the multi-copy nature of p1 (Fig. 18) that creates a delay in the action of purifying selection on recently generated mutants hitchhiking with functional TrpBs in the same cell.
[0155] We then asked whether a state-of-the-art ML model called TranceptEVE (67), which ensembles an autoregressive LLM (Tranception) trained across protein families with a variational autoencoder (EVE) trained on a specific family of proteins (in this case, TrpBs), could predict the measured fitness scores of our lab-evolved TrpBs (Figs.5E-F). Whilesequences that were predicted to have low fitness did exhibit little or no function in our enrichment assay, we found essentially no correlation between the predicted scores and the real enrichment scores of high function TrpBs (Fig.5F, strong selection mean enrichment score >-5, Pearson correlation 0.033, p<0.0001). For example, the highest predicted score was assigned to the nearly nonfunctional TmTriple variant. In contrast and as expected, we found that predicted scores exhibited a much stronger and negative correlation with the number of nonsynonymous amino acid mutations (Pearson correlation -0.765, p<0.0001, Fig.5F). Our interpretation is that the evolution of TrpB for standalone function, rarely found in nature, combined with the extensiveness of sequence divergence from wt TrpB have brought our evolved sequences into regions of the fitness landscape that are out-of- distribution of natural sequences. It has been shown that large ML models can generate highly functional artificial sequences that are more dissimilar to natural sequences than our TrpBs are to wt TrpB; it has also been shown that these models can nominate artificial sequences containing evolutionarily plausible mutations that improve non-natural functions (68). Yet these ML successes do not preclude the possibility that ML-generated sequences nonetheless miss important regions of underlying fitness landscapes. The ability for scalable continuous evolution to enter and explore functional regions of fitness landscapes that ML models do not, and vice-versa, highlights both open challenges in ML and the potential value of combining these two types of approaches going forward.
[0156] Directed evolution of TP-DNAP1
[0157] We carried out our TP-DNAP1 directed evolution campaign over five rounds. In round 1, we started from an epPCR library generated from TP-DNAP1-KS. TP-DNAP1-KS is a relative of TP-DNAP1-4-2 with a mutation rate near 10-5s.p.b. and greater activity than TP- DNAP1-4-2 (42) (Fig.6). We reasoned that its higher activity would confer mutational robustness, increasing the fraction of active library members when used as the parental sequence. After integration of the epPCR library into OR-Y488 and enrichment of TP- DNAP1s that could restore URA3 via a transversion mutation (Fig. S4C), we isolated ~90 colonies and screened for their ability to restore TRP5 through the second selection stage, which was done in multiple replicates per clone following a fluctuation analysis format (42) to obtain estimates on phenotypic mutation rate. We isolated several clones whose phenotypic mutation rates were up to 10-fold elevated (Fig. S7A) and evaluated their genotypic mutation rates in detail using our mutation accumulation assay and Maple. Among these clones, TP- DNAP1-TKS, with the mutation P680T, had the highest per base mutation rate (Fig.12B- 12C). In round 2, selection applied to an epPCR library generated from TP-DNAP1-TKS resulted in the enrichment of several clones whose full mutation rates and spectra were then determined (Fig.13). These several clones turned out to represent two unique TP-DNAP1variants. The two variants satisfied the requirements of selection via distinct strategies. One variant, TP-DNAP1-SgtKis, contained 5 nonsynonymous mutations in addition to those in the parent TP-DNAP1-TKS and demonstrated an altered mutation spectrum favoring transversions but only a minimal apparent increase in overall mutation rate. Another variant, TP-DNAP1-Trixy contained three nonsynonymous mutations and had the highest overall mutation rate measured, but only marginal changes to the mutation spectrum.
[0158] Since TP-DNAP1-SgtKis had an increased transversion rate and TP-DNAP1-Trixy had an increased overall substitution rate, we reasoned that their combination could yield orthogonal DNAPs with both high overall and transversion rates. We cloned seven new TP- DNAP1 variants where a subset of mutations from TP-DNAP1-SgtKis were added to TP- DNAP1-Trixy and obtained their fully described mutation rates (SgtKis / Trixy recombination round, Fig.14). Remarkably, all TP-DNAP1 variants that included the mutations L474S and E488G from TP-DNAP1-SgtKis exhibited a dramatic elevation in their overall mutation rate, in each case bringing the per base rate to ~10-4s.p.b. (Fig.1C, Fig.14C), 1-million-fold higher than the yeast genomic mutation rate (42, 56). Furthermore, the broad mutation spectrum of TP-DNAP1-SgtKis was preserved in these variants (Fig.1D-E, Fig.14D), which we named BadBoy1, BadBoy2, and BadBoy3 (Table 1) to recognize their poor fidelity.
[0159] Our directed evolution campaign yielded two additional notable TP-DNAP1 variants resulting from combining mutations enriched from an epPCR library derived from TP- DNAP1-Trixy (epPCR 3 round, Fig.15) with those in BadBoy3 (BadBoy3 + epPCR 3 round, Table 1). One of the resulting TP-DNAP1s, named BB-5k, exhibited a further increase in mutation rate, to 1.7 × 10-4s.p.b. (Fig.1D), representing ~1 mutation every time a 5 kb recombinant p1 plasmid is replicated. However, BB-5k did not durably maintain p1-ura3*- trp5* in two of four biological replicates over the ~120 generations of mutation accumulation tested, possibly because it exerts an excessive mutational load on the LEU2 marker used to maintain p1-ura3*-trp5*. The other DNAP of potential value, BB-Tv, has a relatively low mutation rate (1.6 × 10-5s.p.b.), like that of our previous OrthoRep systems. Yet unlike previous systems, BB-Tv demonstrated a near-ideal mutation spectrum, with transversions accounting for 43% of all mutations, up from only 2.5% for TP-DNAP1-KS (Table 1, Fig.1D). BB-Tv should therefore be useful in continuous evolution experiments involving larger targets that have lower error thresholds.
[0160] Detailed characterization of mutation accumulation across our several TP-DNAP1 variants revealed interesting trends. For example, we found that the mutation rate of TP- DNAP1-Trixy was correlated with p1 length while TP-DNAP1-SgtKis did not exhibit this relationship, suggesting an interplay between mutation rate and the number of bases replicated that is dependent on mutation mechanism (Fig.16). For BadBoy1, BadBoy2, andBadBoy3, mutation rates were largely independent of p1 length (Fig.16). Characterization of the rate of insertions and deletions revealed that our engineering efforts had a larger impact on insertion rates, with ~50-fold and 2-fold increases in the rates of insertions and deletions, respectively, by BadBoy2 over the parental TP-DNAP1-KS (Table 1). Finally, examination of substitution-type-specific mutation rates among our TP-DNAP1 variants showed that mutation of position 777 to either Ser or Thr yields a large drop (~10-fold) in the A:T→G:C transition mutation rate while having little effect on the reverse G:C→A:T rate (Fig.17), demonstrating that even similar nucleotide substitution types can be generated by independent mechanisms. These subtle observations that come through our precise and rigorous mutation rate measurement pipeline should aid the future engineering of OrthoRep and other continuous evolution systems (57).
[0161] Previous engineering efforts demonstrated that p1 copy number can vary as a function of the TP-DNAP1 replicating it (82). We therefore measured the copy number of p1 when replicated by the wt TP-DNAP1 variant or two representative engineered variants from the current study, BadBoy2 and BadBoy3 (Fig.18). When expressed under the control of a relatively strong constitutive promoter, all three variants supported a similar p1 copy number, ranging from ~100-140. However, weaker expression of both BadBoy2 and BadBoy3 resulted in reduced copy number, down to a copy number of 30 with the weakest promoter tested. Expression of the orthogonal DNAP therefore provides a way to tune copy number, which may prove useful in controlling the outcomes of continuous evolution experiments, since copy number can alter the supply of mutations within a cell and change the rate of fixation of new mutants (83).
[0162] Table 4: Plasmids used in Rix et al.2023
[0163] Table 4 (cont’d)
[0164] Table 5: Primer pairs used in Rix et al.2023
[0165] Table 5 (cont’d)*X denotes a nucleotide with a known sequence used for multiplexing (barcode); N, Y, and R denote mixed nucleotides within a region used as a UMI
[0166] Table 6: Yeast strains used in Rix et al.2023
[0167] Table 6 (cont’d)
[0168] Table 7: High throughput sequencing datasets generated in Rix et al.2023
[0169] Table 7 (cont’d)
[0170] Discussion
[0171] In this work, we have engineered OrthoRep’s mutation rate to reach >10-4s.p.b. whilealso reducing OrthoRep’s bias against transversion mutations to maximize the exploration of sequence space. The durable action of these new mutation rates and preferences on a chosen gene, TrpB, enabled a new modality for continuous evolution where mutations quickly and continuously accumulated on laboratory timescales even through periods of no selection or predominantly purifying selection. In the absence of selection, TrpB sequences swiftly diffused into new regions of sequence space; under strong positive selection, new functional adaptations in TrpB rapidly emerged; and during mostly purifying selection, TrpB sequences quickly sampled the space bounded by the constraints of structure and function, thereby revealing those constraints. At the end of our ~540 generation evolution experiment on TrpB, we obtained thousands of unique sequences, pairs of which were separated by an average ~35 amino acids (~9% divergence) including many >60 amino acids apart (>15% divergence). The amount of evolutionary information recorded into such extensive diversity allowed us to infer both known and unknown mechanisms shaping TrpB’s evolution, including the focusing of adaptation on the COMM domain, the importance of certain allosterically linked positions on the function of TrpB, and the reduction of TrpB’s isoelectric point to yield negatively charged variants even when selection for TrpB’s enzymatic function was absent. The evolutionary information from the experiment also revealed structural and functional constraints acting on TrpB, including conservation of positions near the active site and in buried regions. Additionally, evolved outcomes included many highly active TrpBs that exceeded the in vivo fitness of previously evolved and engineered TrpBs. These variants were distinct from what could be predicted by ML models trained on natural proteins, indicating the discovery of high fitness regions in the fitness landscape of TrpB that are out- of-distribution of natural variation.
[0172] While TrpB was used to demonstrate the new capabilities of OrthoRep in both the evolutionary improvement of a gene’s function and the extensive evolutionary recording of selective forces into sequence diversity, our experiments should easily extend beyond TrpB. Our experiments should also be capable of scaling beyond 96 replicate lines to support the generation and maintenance of greater diversity and / or comparative evolution acrossdifferent selection schedules. Such efforts can also be assisted by automated liquid handling systems (69, 70) that we did not leverage in the current study but may offer additional scale in the future. Overall, the practicality of condensing long adaptive and neutral gene evolutionary processes into laboratory experiments here realized should find broad applications in the evolutionary engineering of biomolecules, the finer and broader mapping of sequence-function relationships, revealing novel biological constraints that shape evolution, and understanding how genes evolve — from their own points of view.
[0173] References
[0174] D. M. Blow, et al. Nature 221, 337–340 (1969).
[0175] 2. G. Casari, et al. Nat Struct Biol 2, 171–178 (1995).
[0176] 3. J. A. Capra, M. Singh, Bioinformatics 23, 1875–1882 (2007).
[0177] 4. B. Reva, et al. Nucleic Acids Res 39, 37–43 (2011).
[0178] 5. N. Halabi, et al. Cell 138, 774–786 (2009).
[0179] 6. J. W. McCormick, et al. Elife 10, 1–38 (2021).
[0180] 7. F. Morcos, et al. Proc Natl Acad Sci U S A 108 (2011).
[0181] 8. D. S. Marks, et al. PLoS One 6 (2011).
[0182] 9. S. W. Lockless, R. Ranganathan, Science 286, 295–299 (1999).
[0183] 10. E. Rivas, et al. Nat Methods 14, 45–48 (2016).
[0184] I. N. Shindyalov, et al. Protein Engineering, Design and Selection 7, 349–358 (1994).
[0185] D. Altschuh, et al. J Mol Biol 193, 693–707 (1987).
[0186] P. J. Haney, et al. Proc Natl Acad Sci U S A 96, 3578–3583 (1999).
[0187] 14. G. Gianese, et al. Protein Eng 14, 141–148 (2001).
[0188] 15. I. N. Berezovsky, E. I. Shakhnovich, Proc Natl Acad Sci U S A 102, 12742–12747 (2005).
[0189] 16. E. M. Marcotte, et al. Proc Natl Acad Sci U S A 97, 12115–12120 (2000).
[0190] 17. E. Shakhnovich, et al. Nature 379, 96–98 (1996).
[0191] 18. R. V. Wolfenden, et al. Science 206, 575–577 (1979).
[0192] D. Collias, et al. Nat Commun 12, 1–12 (2021).
[0193] 20. J. Murciano-Calles, et al. Angewandte Chemie - International Edition 55, 11577–11581 (2016).
[0194] 21. F. Baier, et al. Elife 8, 1–20 (2019).
[0195] G. Gasiunas, et al. Nat Commun 11 (2020).
[0196] 23. M. H. Medema, et al. Nat Rev Genet 22, 553–571 (2021).
[0197] 24. D. L. Trudeau, et al. Curr Opin Chem Biol 17, 902–909 (2013).
[0198] 25. A. Crameri, et al. Nature 391, 288–291 (1998).
[0199] 26. A. J. Riesselman, et al. Nat Methods 15, 816–822 (2018).
[0200] Frazer, et al. Nature 599, 91–95 (2021).
[0201] 28. A. Rives, et al. Proc Natl Acad Sci U S A 118 (2021).
[0202] 29. J. E. Shin, et al. Nat Commun 12, 1–11 (2021).
[0203] 30. D. Bryant, et al. Nat Biotechnol 39, 691–696 (2021).
[0204] 31. A. Madani, et al. Nat Biotechnol, doi: 10.1038 / s41587-022-01618-2 (2023).
[0205] 32. T. A. Hopf, et al. Cell 149, 1607–1621 (2012).
[0206] 33. J. Jumper, et al. Nature 596, 583–589 (2021).
[0207] 34. K. Tunyasuvunakool, et al. Nature 596, 590–596 (2021).
[0208] 35. W. Makałowski, et al. Proc Natl Acad Sci U S A 95, 9407–9412 (1998).
[0209] 36. M. Nei, et al. Proc Natl Acad Sci U S A 98, 2497–2502 (2001).
[0210] 37. M. A. Stiffler, et al. Cell Syst 10, 15-24.e5 (2020).
[0211] 38. G. Rix, et al. Nat Commun 11, 1–11 (2020).
[0212] 39. M. S. Morrison, et al. Nat Chem Biol 16, 610–619 (2020).
[0213] 40. R. S. Molina, et al. Nature Reviews Methods Primers 2 (2022).
[0214] 41. A. Ravikumar, et al. Nat Chem Biol 10, 175–177 (2014).
[0215] 42. A. Ravikumar, et al. Cell 175, 1946-1957.e13 (2018).
[0216] 43. Z. Zhong, et al. ACS Synth Biol 7, 2930–2934 (2018).
[0217] 44. J. D. García-García, et al. Plant Physiol, doi: 10.1093 / plphys / kiab500 (2021).
[0218] 45. E. D. Jensen, et al. Microb Biotechnol 14, 2617–2626 (2021).
[0219] 46. A. A. Javanpour, et al. ACS Synth Biol 10, 2705–2714 (2021).
[0220] 47. A. Wellner, et al. Nat Chem Biol 17, 1057–1064 (2021).
[0221] 48. E. P. Harvey, et al. Nat Commun 13 (2022).
[0222] 49. E. Vallina Estrada, et al. Curr Opin Struct Biol 81 (2023).
[0223] 50. S. E. Luria, M. Delbrück, Mutations of Bacteria From Virus Sensitivity To Virus Resistance. Genetics 28, 491–511 (1943).
[0224] 51. P. L. Foster, Methods Enzymol 409, 195–213 (2006).
[0225] 52. M. Jain, et al. Genome Biol 17, 1–11 (2016).
[0226] 53. P. J. Zurek, et al. Nat Commun 11, 1–10 (2020).
[0227] 54. S. M. Karst, et al. Nat Methods 18, 165–169 (2021).
[0228] 55. H. A. Orr, The Rate of Adaptation in Asexuals.968, 961–968 (2000).
[0229] 56. P. J. Gerrish, et al. J R Soc Interface 10 (2013).
[0230] 57. M. F. Dunn, Arch Biochem Biophys 519, 154–166 (2012).
[0231] 58. A. R. Buller, et al. Proc Natl Acad Sci U S A 112, 14599–14604 (2015).
[0232] 59. D. S. Goodsell, et al. Structure 27, 1716-1720.e1 (2019).
[0233] 60. A. R. Buller, et al. J Am Chem Soc 140, 7256–7266 (2018).
[0234] 61. M. A. Maria-Solano et al. J Am Chem Soc 141, 13049–13056 (2019).
[0235] 62. R. Orij, et al. Microbiology (N Y) 155, 268–278 (2009).
[0236] 63. H. Wennerström, et al. Proc Natl Acad Sci U S A 117, 10113–10121 (2020).
[0237] 64. L. Xiang, R. et al. Nano Lett 23, 1711–1716 (2023).
[0238] 65. M. Leander, et al. Proc Natl Acad Sci U S A 117, 25445–25454 (2020).
[0239] 66. A. Shaw, et al, Removing bias in sequence models of protein fitness. (2023).
[0240] 67. P. Notin, et al. bioRxiv, 2022.12.07.519495 (2022).
[0241] 68. B. L. Hie, et al. Nat Biotechnol, doi: 10.1038 / s41587-023-01763-2 (2023).
[0242] 69. Z. Zhong, et al. ACS Synth Biol 9, 1270–1276 (2020).
[0243] 70. E. A. DeBenedictis, et al. Nat Methods 19, 55–64 (2022).
[0244] 71. E. Alani, et al. Genetics 116, 541–545 (1987).
[0245] 72. O. W. Ryan, J. H. D. Cate, “Multiplex engineering of industrial yeast genomes using CRISPRm” in Methods in Enzymology (Elsevier Inc., ed.1, 2014)vol.546, pp.473–489.
[0246] 73. R. Volden, et al. Proc Natl Acad Sci U S A 115, 9726–9731 (2018).
[0247] 74. R. T. Oliynyk, G. M. Church, Commun Biol 5, 1–10 (2022).
[0248] 75. Y. Zhang, N. A. Tanner, Sci Rep 7, 1–9 (2017).
[0249] 76. A. F. Rubin, et al. Genome Biol 18, 1–15 (2017).
[0250] 77. L. S. Johnson, et al. BMC Bioinform 11, 431. BMC Bioinformatics 11, 431 (2010).
[0251] 78. T. A. Hopf, et al. Nat Biotechnol 35, 128–135 (2017).
[0252] 79. P. J. A. Cock, et al. Bioinformatics 25, 1422–1423 (2009).
[0253] 80. G. I. Lang, et al. Genetics 178, 67–82 (2008).
[0254] 81. R. Tian, et al. Nat Chem Biol, doi: 10.1038 / s41589-023-01387-2 (2023).
[0255] 82. R. Tian, et al. Science 383, 421–426 (2024).
[0256] 83. A. Ravikumar, et al. Cell 175, 1946-1957.e13 (2018).
[0257] 84. A. Garoña, et al. Mol Biol Evol 38, 5610–5624 (2021).
[0258] 85. M. E. Lee, et al. ACS Synth Biol 4, 975–986 (2015).
[0259] 86. B. Ho, et al. Cell Syst 6, 192-205.e3 (2018).
[0260] Throughout this application various publications are referenced. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to describe more fully the state of the art to which this invention pertains. Also incorporated by reference herein, in its entirety, is U.S. Patent Publication Number US- 20220195442-A1, published June 3, 2022.
[0261] Those skilled in the art will appreciate that the conceptions and specific embodiments disclosed in the foregoing description may be readily utilized as a basis for modifying or designing other embodiments for carrying out the same purposes of the present invention. Those skilled in the art will also appreciate that such equivalent embodiments do not depart from the spirit and scope of the invention as set forth in the appended claims.
Claims
What is claimed is:
1. An error-prone DNA polymerase (DNAP) comprising an amino acid sequence having at least 90% identity with SEQ ID NO: 1 and at least three amino acid substitutions relative to SEQ ID NO: 1, wherein the at least three amino acid substitutions are at positions selected from E266, N282, I327, N449, L474, E488, Q598, K635, P680, F702, N713, K753, I761, I777, T828, I863, L900, and F965, and wherein the DNAP has a mutation rate greater than 10-6substitutions per base.
2. The DNAP of claim 1, wherein the at least three substitutions comprise P680T; I777K, I177T, or I777S; and L900S.
3. The DNAP of claim 1 or 2, wherein the at least three substitutions comprise K635R, K753R, and F965Y.
4. The DNAP of claim 1, 2, or 3, wherein the at least three substitutions comprise L474S and E488G.
5. The DNAP of claim 1, comprising an amino acid sequence having at least 95% identity with an amino acid sequence selected from SEQ ID NOs: 3-18, and wherein the at least three substitutions comprise P680T.
6. A nucleic acid molecule encoding the DNAP of any of the preceding claims.
7. The nucleic acid molecule of claim 6, further comprising a promotor.
8. The nucleic acid molecule of claim 7, wherein the promotor is pSAC6, pPSP2, and / or pSLD3.
9. A yeast host cell comprising a p1 plasmid and DNAP of any one of claims 1-5, and one or more p2 components for orthogonal replication of the p1 plasmid.
10. A method of engineering a protein having a desired characteristic, which comprises: a. subjecting a yeast host cell containing a p1 plasmid encoding the protein, and a DNAP of any one of claims 1-5, to error prone orthogonal replication; and b. selecting yeast cells expressing the protein having the desired characteristic.
11. A kit comprising: a. reagents for integration of a gene of interest onto a p1 plasmid; b. a DNAP of any one of claims 1-5; and c. one or more reagents or devices for transforming a yeast cell therewith.
12. The kit of claim 11, further comprising a p1 plasmid packaged together with a yeast host cell comprising one or more p2 components for orthogonal replication of the p1 plasmid.
13. The kit of claim 12, wherein the yeast host cell is packaged together with one or more reagents or devices for culturing and / or transforming the yeast host cell.