Systems for the continuous, rationally directed evolution of biomolecules
Patent Information
- Application Number
- PCT/US2026/021079
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-03-26
- Publication Date
- 2026-10-01
Smart Images

Figure US2026021079_01102026_PF_FP_ABST
Abstract
Description
Aty. Dkt. No. 125141.04976 MGH2024-443SYSTEMS FOR THE CONTINUOUS, RATIONALLY DIRECTED EVOLUTION OF BIOMOLECULESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Application No.63 / 778,073 filed on March 26, 2025, the content of which is incorporated by reference in its entirety.BACKGROUND
[0002] With the advent of protein language models and artificial intelligence, the ability to predict potential fitness-improving protein variants has greatly outpaced the ability to screen protein variants. At present, existing high-throughput protein screening methods have centered on the use of ‘display’ technologies, such as phage, yeast, and mammalian display. These approaches typically involve the generation of large vector libraries via ex vivo molecular cloning, in which select codons within a gene of interest encoded on a vector are diversified by PCR with degenerate oligonucleotide primers (for example NNN or NNK). Such libraries are then transformed into recipient cells and the encoded protein variants are screened for desirable properties, such as enhanced antigen binding or specificity in the case of antibodies or soluble T cell receptors (TCRs). While each method offers distinct advantages and drawbacks, all share the fundamental limitation of being highly labor-intensive and time-consuming when performed iteratively. Strategies that employ full-saturation mutagenesis and continuous selection of protein variants with fitnessimproving mutations are desired.SUMMARY
[0003] In an aspect, provided herein is a construct comprising sequences encoding: a diversity generating retroelement (DGR) accessory variability determinant (Avd) protein; a DGR reverse transcriptase (RT); a single-stranded DNA annealing protein (SSAP); and at least one non-coding RNA; wherein the sequences encoding the Avd, the RT, the SSAP, and the non-coding RNA are operably linked to one or more promoters; wherein the non-coding RNA comprises a target homology region (THR) that comprises homology to a target nucleic acid; and wherein at least one nucleotide in the THR is a non-homologous adenine.Aty. Dkt. No. 125141.04976 MGH2024-443
[0004] In another aspect, provided herein is a plasmid comprising the construct described herein.
[0005] In another aspect, provided herein is a kit comprising: one or more plasmids comprising: a diversity generating retroelement (DGR) accessory variability determinant (Avd) protein; a DGR reverse transcriptase (RT); a single-stranded DNA annealing protein (SSAP); and at least one noncoding RNA; wherein each of the sequences encoding the Avd, the RT, the SSAP, and the noncoding RNA are operably linked to one or more promoters;wherein the non-coding RNA comprises a target homology region (THR) that comprises homology to a target nucleic acid; and wherein at least one nucleotide in the THR is a non-homologous adenine.
[0006] In another aspect, provided herein is a kit comprising the plasmid described herein; and a phagemid vector or a phage vector encoding the target nucleic acid.
[0007] In another aspect, provided herein is a cell comprising the construct or plasmid described herein.
[0008] In another aspect, provided herein is a phage or phagemid vector comprising the construct or plasmid described herein.
[0009] In another aspect, provided herein is a kit comprising the cell described herein; and a helper cell, wherein the helper cell comprises in its genome all genes necessary for propagation of a phage vector or phagemid vector except genes encoded on the phage vector or phagemid vector or on an accessory plasmid.
[0010] In another aspect, provided herein is an E. coli cell comprising at least one of a Red gene deletion; an sbcB gene deletion; and a Q576A mutation in DNA gyrase.
[0011] In another aspect, provided herein is a method for generating and identifying functional mutants of a starting biomolecule, the method comprising: a) back-diluting and growing an initial fresh culture of host cells to mid log phase, wherein the host cells are transformed with: a helper plasmid comprising all genes required for phage propagation except gVI and gill; an accessory plasmid comprising gVI operably linked to a synthetic gene circuit, wherein when the synthetic gene circuit is activated, the gVI is transcribed; and the plasmid described herein; b) infecting the host cells with a selection phagemid vector comprising gill, wherein the selection phagemid vector comprises the target nucleic acid; wherein the target nucleic acid encodes the biomolecule; wherein activity of the biomolecule activates the synthetic gene circuit; c) incubating the host cells under conditions that allow for production of an infectious phagemid, wherein infectious phagemidAty. Dkt. No. 125141.04976 MGH2024-443encoding variants of the biomolecule that activate the synthetic gene circuit to a greater extent than the starting biomolecule will receive a propagation advantage; d) harvesting phagemid from the culture and discarding host cells from the culture; e) back-diluting and growing to mid-log phase a subsequent fresh culture of host cells; and infecting the host cells with the harvested phagemid from step d); f) repeating steps b) through e); and g) isolating phagemid encoding variants of the starting biomolecule that inactivate the synthetic gene circuit to a greater extent than the starting biomolecule.
[0012] In another aspect, provided herein is a method for generating and identifying functional mutants of a starting first biomolecule, the method comprising: a) back-diluting and growing to mid-log phase an initial fresh culture of host cells, wherein the host cells are transformed with: a helper plasmid encoding all genes required for phage propagation except gVI and gill; a first accessory plasmid encoding gVI operably linked to a synthetic gene circuit, wherein when the synthetic gene circuit is activated, the gVI is transcribed; and a second accessory plasmid encoding at least one additionalbiomolecule that interacts with the first biomolecule; wherein the interaction of the first and at least one additional biomolecules activates the synthetic gene circuit; and the plasmid described herein; b) infecting the host cells with a selection phagemid comprising gill, wherein the phagemid comprises the target nucleic acid; wherein the target nucleic acid encodes the first biomolecule; c) incubating the host cells under conditions that allow for production of an infectious phagemid, wherein infectious phagemid encoding variants of the starting first biomolecule that activate the synthetic gene circuit to a greater extent than the starting first biomolecule will receive a propagation advantage; d) harvesting phagemid from the culture and discarding host cells; e) back-diluting and growing to mid-log phase a subsequent fresh culture of host cells; and infecting the host cells with the harvested phagemid from step d); f) repeating steps b) through e); and g) isolating phagemid encoding variants of the first biomolecule that activate the synthetic gene circuit to a greater extent that the starting first biomolecule.
[0013] In another aspect, provided herein is a method for generating and identifying a mutant of a biomolecule endogenously expressed in a host cell or a mutant in a component of the genetic circuitry that regulates expression of the biomolecule, the method comprising: a) transforming a population of the host cell with the construct described herein, wherein the construct is operably linked to an inducible promoter, wherein the target nucleic acid encodes the endogenously expressed biomolecule or the component of the genetic circuitry that regulates expression of theAty. Dkt. No. 125141.04976 MGH2024-443biomolecule; b) contacting the host cell with an inducing agent that activates the inducible promoter; c) incubating the cell for between about 0 and about 24 hours under conditions that permit the function of the construct; d) selecting host cells containing mutant biomolecules or mutant components of the genetic circuitry that confer an improvement of interest; and e) extracting and sequencing gDNA from the host cells selected in d).
[0014] In another aspect, provided herein is a system for directed evolution of a protein of interest in a bacterial periplasm, the system comprising: (a) a phagemid vector comprising a nucleic acid encoding the protein of interest; and gill; (b) an accessory helper plasmid comprising: (i) a helper plasmid construct comprising all genes required for phage assembly except gill and gVI; (ii) a target antigen construct comprising a gene encoding a fusion protein, the fusion protein comprising a transmembrane protein and a target antigen; wherein the target antigen binds the protein of interest; and wherein when the fusion protein is expressed, the target antigen is anchored in the membrane and exposed to the periplasm; (iii) a conditional propagation construct comprising a conditional promoter operably linked to gVI; wherein the conditional promoter comprises an operator to which the transmembrane protein is capable of binding; and wherein when the transmembrane protein binds to the operator, gVI is expressed; and (c) an engineered bacterial cell comprising: disruption or deletion of a gene encoding the transmembrane protein; and a CadBA operon that is naturally operably linked to the conditional promoter; wherein the conditional promoter is replaced with a heterologous constitutive promoter; wherein interaction of the protein of interest with the target antigen in the periplasm activates the conditional propagation construct and enables propagation of the phage.
[0015] In another aspect, provided herein is an engineered bacterial cell for periplasmic evolution of proteins encoded on a phagemid, the cell comprising: deletion or inactivation of cadC; and a heterologous constitutive promoter operably linked to cadBA.
[0016] In another aspect, provided herein is a method for generating a library of mutants of a biomolecule, the method comprising performing a directed evolution technique using the plasmid described herein; wherein the target nucleic acid encodes the biomolecule.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] FIGS. 1A-1B. Components of the BPP-1 DGR generate adenine-diversified cDNA libraries in E. coli s2060. (A) Schematic of the split intron reporter assay used to verify activityAty. Dkt. No. 125141.04976 MGH2024-443of BPP-1 DGR proteins in E.coli s2060. An Intron Plasmid (TP, purple) encoding BPP-1 DGR proteins bRT (navy) and Avd (teal) as well as a DGR-RNA with a self-splicing td intron (grey) are transformed into E.coli s2060 (tan). Cells are grown to OD=0.3 and induced with arabinose leading to expression of the DGR proteins from the Pbad promoter. The td intron self splices leaving a DGR-RNA with the 5’UTR (orange), TR region (light navy), and 3’UTR (green) ready to be reverse transcribed by the bRT-Avd complex. Reverse transcription of the TR region proceeds in a manner highly biased towards A to N mutations, creating a library of cDNA in the cytoplasm divergent at positions corresponding to adenines in the TR region of the DGR-RNA-encoding gene on the IP. The self-splicing intron creates a length differential between PCRs off the cDNA and off the IP, allowing visualization and extraction of cDNAs for NGS analysis (B). Visualization of amplicons off the cDNA (300bp band) and off the IP (750bp band).
[0018] FIGS.2A-2C. Mutagenic analysis of cDNA libraries produced in E.coli s2060 by DGR split intron reporter assay. (A) A cDNA library produced 5 hours post arabinose induction was sequenced in an Illumina Miseq with 100,000 reads / sample and aligned against the TR sequence as a reference using the CRISPREsso2 analysis suite. A bell-curve distribution of substitutions relative to the TR was identified, with a mean of 12 substitutions per read. (B) Total mutagenic spectra of cDNA libraries described in (A). The mutagenic distribution across all template cytosines, guanines, uracils and adenines was separately measured and averaged. The overwhelming majority of mutations in cDNA libraries are at positions corresponding to an adenine base in the reference TR. (C) Mutagenic analysis of cDNA libraries described in (A) at each position corresponding to an adenine base in the reference TR. The reverse-transcribed base at each position is provided: incorporation of thymine (T) opposite a template adenine (A) is reported as a correct incorporation (no change), while cytosine, adenine, and guanine incorporations opposite template adenines are reported as mutations.
[0019] FIG. 3. Allelic mapping of cDNA libraries produced in E.coli s2060 by DGR split intron reporter assay. Allelic mapping was performed on every position in the cDNA library with a corresponding AAC (N) codon in the TR. AAC can be diversified into 15 of the 20 canonical amino acids by bRT; all 20 amino acids are shown together with stop codons. Data generated by Crispresso2 analysis of 100,000 reads of a cDNA library harvested five hours post IP induction.
[0020] FIGS. 4A-4B. Mutagenic rates and spectra measured in DGR split intron reporter assays are not impacted by pre-digestion of the IP or the polymerase used for NGS libraryAty. Dkt. No. 125141.04976 MGH2024-443preparation. (A) Mutagenic spectra measured on template adenines for the same cDNA library harvested 5 hours post arabinose induction with either Nebnext or Q5 polymerase used to prepare the NGS library. For each condition, extracted DNA was either subject to digestion with Xbal, which cleaves in the IP in the intron sequence, or no digestion. Digestion is intended to reduce the possibility of template switching during NGS library preparation, which could alter the mutagenic spectra measured in the final result. Data generated by Crispresso2 analysis of 100,000 reads of a cDNA library harvested five hours post IP induction. (B) Gel image of the libraries analyzed in (A). A clear band off the IP is present in the no digest condition (750bp) which is absent or reduced in the digestion condition, demonstrating that Xbal digestion was successful.
[0021] FIG. 5. A ‘reticulase’ is a novel class of prokaryotic retrotransposon that combines components of the BPP-1 DGR and single-stranded annealing proteins to enable hypermutation of select gene regions in bacterial chromosomes with single base-pair precision. Schematic of the reticulase mechanism of action on an coli chromosomal locus. The Reticuase Plasmid (RP, green) encoding four reticulase proteins under the control of the arabinose-inducible Pbad promoter (Avd, teal), (bRT, navy), (ssap, lime), (MutL *, pink) as well as a regRNAregRNA expressed off a constitutive Pbad promoter is transformed intoE.coli s2064 (tan). Cells are grown to OD=0.3 and induced with arabinose, leading to the expression of the reticulase proteins off the RP. The regRNAregRNA is transcribed, compositing the DGR-RNA 5’UTR (orange) a ‘Target Homology Region’ (THR, light navy), and DGR-RNA 3’UTR (green). Codons or gene regions in the target genome selected for mutagenesis are replaced with AAA in the THR. bRT and Avd reverse-transcribe the regRNAregRNA in a manner that is highly biased towards A to N mutations, producing a cDNA library that is highly degenerate at user-specified positions. Ssap (lime) binds cDNA and anneals it in place of an Okazaki fragment on the lagging strand of a replication fork during chromosomal replication. MutL* is a dominant negative mismatch repair protein that prevents reversion of mismatched bases long enough for a subsequent round of genome replication to result in a stable double stranded edit. The end result is the full saturation (NNN) of the user-specified codon in the bacterial genome.
[0022] FIGS. 6A-6D. Iterative improvements made to the reticulase, host strain, and induction conditions to improve reticulase editing efficiency on E coli chromosomal loci. (A) Comparison of reticulase editing efficiency in the s2064 strain (s2060 ARecJ AsbcB) and the s2065 strain (s2060 ARecJ A sbcB DnaG Q576A) for a discrete C>T edit in rpob nucleotide 1591 with anAty. Dkt. No. 125141.04976 MGH2024-44380bp THR. Data represent the average of four biological replicates, induced with 1 OmM arabinose and harvested 20 hours post induction. (B) Impact of RP copy number on editing efficiency of either a discrete OT edit in rpob nucleotide 1591 with an 80bp THR or a discrete OT edit in gyrA nucleotide 248 with an 80bp THR. Data represent the average of four biological replicates in s2064, induced with lOmM arabinose and harvested 20 hours post induction. (C) Comparison of arabinose induction concentration on reticulase editing efficiency with a discrete C>T edit in rpob nucleotide 1591 with an 40bp THR. Data represent the average of 3 biological replicates in s2064 harvested 20 hours post induction. (D) Comparison of arabinose induction concentration on reticulase editing efficiency with either a discrete C>T edit in rpob nucleotide 1591 with an 80bp THR or a discrete C>T edit in gyrA nucleotide 248 with an 80bp THR. Data represent the average of 4 biological replicates in s2064 harvested 20 hours post induction.
[0023] FIGS. 7A-7C. The spG4C mutation in the 3’UTR of the regRNAregRNA does not increase editing efficiency on E. coli chromosomal loci. (A) Schematic of the DGR-RNA associating with bRT (transparent navy) and Avd (transparent teal). The spG4 nucleotide sits at the base of the stem loop that lays on top of the Avd pentamer and makes extensive hydrogen bonding contact with Avd (B) Comparison of the WT and spG4C regRNAregRNAs with a discrete C>T edit in rpob nucleotide 1591 with an 80bp THR. Data represent the average of 6 biological replicates in s2064 harvested 20 hours post induction. (C) Comparison of the WT and spG4C regRNAregRNAs with a discrete C>T edit in gyrA nucleotide 248 with an 80bp THR. Data represent the average of 6 biological replicates in s2064 harvested 20 hours post induction.
[0024] FIGS.8A-8B. Truncation of the regRNAregRNA THR enables discretionary targeted mutagenesis of select gene regions through homology shielding. (A) Comparison of reticulase editing efficiencies for a series of THRs of varying length making a discrete C>T edit in rpob nucleotide 1591. Data represent an average of 3 biological replicates in s2064 induced with ImM arabinose and harvested 20 hours post-induction. (B) Heat map displaying the substitution frequency in edited reads identified in (a). The nucleotides around the edited nucleotide (rpob T 1591, highlighted in red) are shown on the X axis, and the corresponding THR lengths are shown on the Y axis. The sequences contained in each THR are boxed in black.
[0025] FIGS. 9A-9B. Reticulases enable site saturation mutagenesis of user-selected codons in bacterial genomes with single base pair precision. (A) Substitution frequency around rpob codon S531 targeted by a 36bp THR with triple adenine (AAA) substituting the endogenous codonAty. Dkt. No. 125141.04976 MGH2024-443(TCC) present in the target locus. Data represent the average of three biological replicates of s2064 induced with ImM arabinose and harvested 20 hours post-induction. (B) Allelic mapping of amino acids substituted at position S531 in the experiment described in (a). Data represent mean of three biological replicates of s2064 induced with ImM arabinose and harvested 20 hours post-induction. An average of 500,000 reads per sample were analyzed with Crispresso2.
[0026] FIGS. 10A-10C. Reticulases enable rational evolution of user-selected codons in a bacterial genome with single base-pair precision. (A) Schematic of rational evolution assay used to identify novel mutations conferring resistance to the antibiotic rifamycin at position 531 in E. coli gene rpoB. E.coli s2064 (tan) are transformed with the RP (green) and grown up to OD=0.3 and induced with arabinose. The rpob codon S531 (TCC) in the E.coli chromosome (grey) is the target of a 36bp THR with a triple adenine (AAA) at the position corresponding to position 531. The RP generates a cDNA library with NNN at the position corresponding to 531, which is integrated into the bacterial genome to create a population of bacteria which are highly degenerate at this position. 20 hours post-induction, cells are plated on 2xYT plates either with or without 25ng / uL rifamycin. Colony counting allows estimation of the fraction of E.coli that have become rifamycin resistant. Sanger sequencing of the rpob locus of single rifamycin-resistant colonies allows identification of novel rifamycin resistance mutations at position 531. (B) Fraction of rifamycin-resistant colony forming units (cfu) from the experiment described in (A) induced with either 12.5mM arabinose or lOOmM glucose. The reticulase proteins Avd, bRT, ssap, and MutL* are under the control of the arabinose-inducible Pbad promoter, which can be suppressed through the addition of glucose to the media. Data from one biological replicate shown. (C) Representative traces generated by sanger sequencing of the rpob locus from rifamycin-resistant colonies described in (B). Data from one biological replicate shown.
[0027] FIGS. 11A-11C. Reticulases enable target genomes to transcend large evolutionary valleys and access unusual fitness-improving genotypes. (A) 24 colonies from the experiment described in FIG. 10A were sanger sequenced, and the amino acid substitutions made at position 531 are tabulated and graphed. Novel resistance mutations S531K (TCOAAA / AAG), S531R (TCOAGA / AGG), S53 II (TCOATT) were identified which have not been previously reported, as well as S531G (TCOGGG / GGA), S531Q (TCOCAG), and S531L (TCOTTA) which have previously been reported to confer rifamycin resistance. The ability of the RP to introduce a wide range of triple substitutions scarlessly in frame enables the discovery of novel fitness-improvingAty. Dkt. No. 125141.04976 MGH2024-443genotypes that would be unlikely to arise via random mutagenesis scanning. (B) the experiment performed in (A) was repeated at 190 colonies were sanger sequenced, and the amino acid substitutions made at position S531 were tabulated as in (A). (C) Crystal structure of rpob (blue) in complex with rifamycin (green). Serine S531 (red) points directly into the rifamycin binding pocket in rpob (PDB 5UAC).
[0028] FIGS. 12A-12B. Division of the RegRNARegRNA in the 3’UTR between sp20-sp21 will enable multiplexed editing of distinct genomic regions. (A) Schematic of the split regRNAregRNA in association with the bRT (transparent navy) - Avd (transparent teal) complex. The regRNAregRNA is split between nucleotides sp20 and sp21 in the 3’UTR, which form the end of a hairpin loop in the WT DGR-RNA. Extensive sequence homology between the two RNAs enables assembly into the same overall secondary structure as the WT DGR-RNA. The 5’UTR (orange) has roughly 20bp homology to the 3’UTRb region (lime), while the 3’UTRa (green) shares roughly 30bp homology to the 3’UTRb. This secondary structure facilitates cDNA priming off the 2’ OH group of nucleotide sp56 in the 3’UTRb (as in the WT DGR-RNA), enabling reverse transcription of the THR region present in the split RNA with the 3’UTRa portion (hereafter referred to as ‘regRNAregRNA A’). These RNAs can be expressed off distinct genomic loci and may be expressed polycistronic with the reticulase proteins. (B) Schematic of multiplexed cDNA library generation using a split reticulase array. Multiple regRNAregRNA A’s can be strung together, with each THR targeting a distinct genomic locus. A concatenation of the 3’UTRa and 5’UTR form an iterative repeat sequence of 57bp between each THR. In our best split construct (Architecture 3, FIG. 13) the regRNAregRNA A is expressed with the arabinose-inducible Pbad promotor polycistronic with and upstream of the reticulase protein-encoding genes, while the 3’UTRb is expressed off a separate locus on the RP with the strong constitutive promoter ProD.
[0029] FIGS. 13A-13B. Schematic of reticulase architectures tested using both full RNA and split-RNA approaches. (A) Schematic of reticulase architectures tested using both the full RNA and split-RNA approaches. Architecture 1 represents the standard architecture used in all experiments in this manuscript unless otherwise noted: Pbad driving Avd (teal), bRT (blue), ssap (yellow), and MutL* (pink), hereafter referred to as the ‘reticulase operon,’ and the strong constitutive promoter ProD driving the full regRNAregRNA comprising the DGR-RNA 5’UTR (orange), THR (navy), and full DGR-RNA 3’UTR (region upstream and including sp20 colored lime, region including sp21 and downstream colored green). Architecture 2: the fullAty. Dkt. No. 125141.04976 MGH2024-443regRNAregRNA is expressed polycistronic with the reticulase operon off Pbad. Architecture 3: The regRNAregRNA A is expressed polycistronic with the reticulase operon off Pbad, and the 3’UTRb is expressed off a separate locus on the RP with ProD. Architecture 4: The 3’UTRb is expressed polycistronic with the reticulase operon off Pbad, and the regRNAregRNA A is expressed off a separate locus on the RP with ProD. (B) Reticulase editing efficiency for an 80bp THR encoding a discrete OT edit in E. coll chromosomal gene gyrA at nucleotide 248 for each of the architectures described in (A). Data represent eight biological replicates of s2064 cells transformed with each of the reticulases, induced with lOmM arabinose, and harvested 20 hours post induction. The low copy sclOl origin of replication was used for these reticulases, as we previously identified this ori produced the highest editing on gyrA. A non-targeting (nt) Architecture 1 reticulase diversifying rpob was used as a negative control.
[0030] FIG. 14. Reticulases enable targeted hypermutation of genes encoded on M13 phagemid. Schematic of reticulase-mediated M13 phagemid editing. A gene of interest (red) is encoded on the genome of a phagemid (bright green) downstream of its two origins of replication (grey). s2061 host cells (tan) transformed with a helper plasmid (HP, orange) and a reticulase plasmid (RP, dark green) are grown to OD=0.3 and induced with arabinose and infected with phagemid at l*104cfu ml;1unless otherwise noted. The helper plasmid expresses all genes required for phage replication and packaging except for gill, which is encoded on the phagemid. gV (pink) on the HP is responsible for regulating the duration of rolling circle amplification and Fibonacci expansion during each infection cycle of the phagemid (see FIG. 21). The RP continually produces cDNA libraries set to diversify corresponding positions on the gene of interest encoded on the phagemid once the latter infects the cell. After 18 hours, phagemid are harvested from the supernatant, titered and subjected to NGS for sequence analysis.
[0031] FIGS. 15A-15B. Reticulases edit M13 phagemids maintained as resident plasmids in E.coli s2060 and derived strains. (A) Schematic of phagemid plasmid propagation assay used to determine impact of E.coli genetic background and ColEl ori-dependent replication on reticulase editing efficiency. A selection phagemid (SP, red) contains both an Ml 3 phage origin of replication and a ColEl origin of replication, an antibiotic resistance marker (CarbR, teal), a gene of interest which is the target of the reticulase (red arrow), and Ml 3 phage gill, which encodes the minor coat protein pill (not shown). In the absence of a helper plasmid (HP), which provides M13 genes gll, gV, and gX required for facilitating rolling circle amplification initiating from the M13Aty. Dkt. No. 125141.04976 MGH2024-443phage origin of replication, only the ColEl plasmid origin of replication is active. Cells transformed with the SP and grown in the presence of Carbenicillin will maintain the SP as a stable resident plasmid with a copy number of approximately 20 copies / cell. s2060 cells or a derivative strain (tan) are transformed with the SP and RP (green), grown to OD=0.3 and induced with arabinose. The RP makes a discrete AT>CG edit in the transcription factor encoding gene clopt with a 68bp THR. RP with a medium copy number origin of replication (CloDF13, ~40 copies / cell) was used. After 18 hours cells are lysed and subject to NGS for analysis. (B) Reticulase editing efficiency on resident plasmid SP following induction with ImM arabinose. s2060 cells (WT), s2062 (s2060 AsbcB). s2061 (s2060 ARecJ), or s2064 (s2060 AsbcB ARecJ) were tested in parallel. Data are average of three biological replicates.
[0032] FIGS. 16A-16D. Reticulases edit M13 phagemids as phage propagating in s2061 cell lines containing a resident helper plasmid (HP). (A) Initial data for experiment described in FIG. 15. s2061 cell lines transformed with HP and RP are grown to OD=0.3 and infected with phagemid at a 1*104 cfu mL-1 and induced with lOmM arabinose. The RP makes a discrete AT>CG edit in the transcription factor encoding gene clopt with a 68bp THR. RP with either a medium copy number origin of replication (CloDF13, ~40 copies / cell) or a low copy number of replication (sclOl, ~5 copies / cell) were tested. 18 hours post-induction phage are harvested from the supernatant and subject to NGS for analysis. Nontargeting controls for each RP (with THRs targeting rpob) are shown for comparison. Data represent average of six biological replicates. (B) titering data for experiment described in (A). (C) Titer and editing efficiency data from (A) and (B) for the high copy number RP are graphed on an X / Y plot. Simple linear regression analysis generates an R squared value of .1857, with a P value of .1610, demonstrating no statistically significant correlation between editing efficiency and final titer for the high copy RP conditions tested.(D) Titer and editing efficiency data from (A) and (B) for the low copy number RP are graphed on an X / Y plot. Simple linear regression analysis generates an R squared value of .2410, with a P value of .1051, demonstrating no statistically significant correlation between editing efficiency and final titerfor the low copy RP conditions tested.
[0033] FIG. 17. Expression of the reticulase components ssap and MutL* , not bRT and Avd, mildly depresses titers of M13 phagemid propagated in s2061. Schematic of M13 phagemid multi-passage experiment to determine impact of respective reticulase components on Ml 3 phagemid titer. s2061 cells (tan) are transformed with a helper plasmid (HP, orange) lacking gVIAty. Dkt. No. 125141.04976 MGH2024-443(the gene encoding Ml 3 protein pVT), an accessory plasmid (AP,blue) encoding a synthetic circuit in which expression of gVI is contingent on the binding of the lambda phage transcriptional operator cl opt, and a nontargeting reticulase plasmid (RP, green) with a regRNAregRNA targeting mcherry or Mtd. Cells were transformed with either a full reticulase, a reticulase encoding only the BPP-1 DGR proteins bRT and Avd, or no reticulase. A phagemid (light green) encoding a clopt lambda phage transcriptional operator gene (red) infects the s2061 cells transformed with HP, AP, and RP (hereafter referred to as host cells) grown to OD=0.3 at a starting titer of l*106cfu ml;1. Host cells are induced with ImM arabinose to activate the RP at the time of infection. The clopt transcriptional operator (red) is expressed off the phagemid upon infection, leading to activation of the synthetic circuit and expression of pVI, which is required for phage propagation. After 18 hours, phage are harvested from the supernatant and back diluted l*10-3into a new batch of host cells grown to OD=0.3 and induced with ImM arabinose. This process of serial dilutions is repeated for a total of five serial passages, resulting in a total dilution factor of the starting phage stock of l*10‘12.
[0034] FIGS. 18A-18C. Expression of the reticulase components ssap and MutL*, not bRT and Avd, mildly depresses titers of M13 phagemid propagated in s2061. (A) Phagemid titers from the experiment described in (A). Each condition refers to a different host cell line with either a full reticulase expressing all reticulase proteins (Avd, bRT, ssap, and MutL*), a reticulase expressing only the BPP-1 DGR proteins Avd and bRT, or host cells with no reticulase. Experiment performed as a single biological replicate. (B) Titer differential of data shown in (B) between phagemid propagated in host cells with the full nontargeting reticulase vs the nontargeting reticulase with only BPP-1 DGR proteins Avd and bRT. (C) Titer differential of data shown in (B) between phagemid propagated in host cells with the nontargeting reticulase with only BPP-1 DGR proteins Avd and bRT vs no reticulase.
[0035] FIGS. 19A-19C. Reticulases can efficiently edit genes encoded on M13 phagemid irrespective of starting infection titer. (A) 2061 cell lines were transformed with HP and RP, grown to OD=0.3, induced with lOmM arabinose, and infected with starting phagemid titer ranging from l*102cfu mL'1to l*109cfu mL'1in one-log increments. The RP makes a discrete AT>CG edit in the transcription factor gene clopt encoded on the phagemid with a 68bp THR. RP with a medium copy number origin of replication (CloDF13, ~40 copies / cell) was used. A nontargeting reticulase set to diversify rpoB was used as a negative control. Phagemid were harvestedAty. Dkt. No. 125141.04976 MGH2024-44318 hours post infection and subject to NGS analysis and titered. Data represent the average of six biological replicates. (B) Final titering data from the targeting reticulase condition of experiment described in (A). Singlicate data shown. (C) Final titering data from the nontargeting reticulase condition of experiment described in (A). Singlicate data shown.
[0036] FIGS. 20A-20B. Mathematical modeling of maximum reticuase editing efficiency theoretically achievable on an M13 phagemid population expanding through either rolling circle amplification / Fibonacci expansion or leading lagging exponential expansion. (A) Equations characterizing expansion of M13 phagemid population through a Fibonacci sequence (dependent on M13 origin of replication) or exponential sequence (ColEl origin of replication). To calculate maximum edited portion of the population, equation models a single edit made at the first point in the replication cycle of the genome when the replaced strand is being polymerized (see cartoons in FIGS. 21-23). ‘Reticulate fibonasion’ is used to refer to editing in which a cDNA is integrated into the genome during rolling circle amplification / Fibonacci expansion (dependent on Ml 3 phage origin of replication). ’Recombineering’ is used to describe edits made via cDNA annealing in place of an Okazaki fragment on the lagging strand of the replication fork (dependent on ColEl origin of replication), n refers to number of replication cycles, or ‘generation number.’ Fnis total population size at generation n. Enis edited population size at generation n. Fraction edited (En / .Fn) is multiplied by 100 to generate edited fraction (%) reported in (B). (B) Edited fraction (%) of the population for each of the equations described in (A), graphed to generation (n) 20.
[0037] FIG. 21. Schematic and population parameters for first four generations of -strand replacement reticulate fibonasion. -strand replacement refers to the integration of a cDNA at the target locus in place of the -strand during reticulate fibonasion (cDNA shares GCT homology with the +strand and GCT identity with the -strand.) Here, first edit can be made at generation n=l.
[0038] FIG. 22. Schematic and population parameters for first four generations of +strand replacement reticulate fibonasion. +strand replacement refers to the integration of a cDNA at the target locus in place of the +strand during reticulate fibonasion (cDNA shares GCT homology with the - strand and GCT identity with the + strand.) Here, first edit can be made at generation n=2.
[0039] FIG.23. Schematic and population parameters for first four generations of edit made through Okazaki fragment annealing (recombineering). Here, first edit can be made atAty. Dkt. No. 125141.04976 MGH2024-443generation n=2.
[0040] FIGS. 24A-24C. Optimization of the M13 phagemid architecture enables high-efficiency reticulase editing (invention of the M13 tracemid). (A) Schematic of the phagemid architectures tested. The M13 origin of replication and the ColEl origin of replication (box arrows) can be placed in four arrangements relative to the reticulase target gene (red arrow). The polymorphisms identified for the M13+ ColEl+ architecture (labeled vl, v2, and v3) are labeled above the schematic (we have defined this dual origin orientation as an Ml 3 tracemid). (B) Reticulase editing efficiency on phagemids for each of the six architectures described in (A). s2061 cells transformed with the HP and RP were grown to OD=0.3 and induced with lOmM arabinose and infected with each phagemid at l*104cfu mL’1. Phage were harvested 18 hours post infection. The RP makes a discrete AT>CG edit in the transcription factor encoding gene clopt with a 68bp THR. RP with a medium copy number origin of replication (CloDF13, ~40 copies / cell) was used. Data represent mean of five biological replicates. (C) Titer data for experiment described in (B). Singlacate titering data shown; boxed datapoints in (B) represent titered replicates shown in (C).
[0041] FIGS.25A-25B. Recombineering with the ColEl plasmid origin of replication on M13 tracemid with the (M13+ C0IEI+) architecture is a minor but statistically significant contributor to total reticulase editing efficiency. (A) Reticulase editing efficiency on phagemids with either the (vl M13+ C0IEI+) architecture (which we have defined as an M13 tracemid) or the (M13+ ColEl-) architecture. s2061 cells transformed with the HP and RP were grown to OD=0.3 and induced with lOmM arabinose and infected with each phagemid at l*104cfu mL'1. Phage were harvested 18 hours post infection. The RP makes a discrete AT>CG edit in the transcription factor encoding gene clopt with a 68bp THR. RP with a medium copy number origin of replication (CloDF13, ~40 copies / cell) was used. Data represent 42 biological replicates. (B) Results of unpaired T test on data generated in experiment described in (A).
[0042] FIG. 26. Schematic summarizing proposed reticulase editing mechanism on M13 tracemid with the (vl M13+ C0IEI+) architecture. Recombineering (cDNA replacement of an Okazaki fragment on the lagging strand of a replication fork initiated by the ColEl origin of replication) is a minor contributor (<20%) to total reticulase editing efficiency on this vector. -strand replacement reticulate fibonation (cDNA replacement of the -strand during -strand synthesis originating from the M13 origin of replication) is the major contributor (>80%) to total reticulase editing efficiency on this vector. cDNA (purple bar), ssap (green), origins of replicationAty. Dkt. No. 125141.04976 MGH2024-443(grey).
[0043] FIGS.27A-27C. Fine-tuning the M13 life cycle to increase reticulase editing efficiency and final library titer. (A) Schematic of the role of pV (purple) in the natural life cycle of M13 phage. pV is the protein product of Ml 3 gV and regulates the duration of Fibonacci expansion during the Ml 3 life cycle. pV is a single-strand DNA-binding protein that associates with the +strand, (green) and prepares it for packaging and cell export. In early infection, pV concentration in the cell is low and +strands are free to undergo -strand (red) synthesis and subsequent rolling circle amplification / Fibonacci expansion. In late infection, pV reaches a high concentration in the cell at which point all +strands are bound by hundreds of copies of pV, arresting RCA and preparing the +strands for packaging. The spacer sequence upstream of gV on the helper plasmid (HP, orange) controls the expression level of pV in the cell and can be used to fine tune the duration of the Fibonacci expansion phase of the Ml 3 phage life cycle. (B) Impact of pV concentration on reticulase editing of Ml 3 tracemid. A series of previously characterized spacer sequences upstream of gV on the helper plasmid were tested, arranged by decreasing level of pV expression. WT (18k-0) is included for comparison. s2061 cells transformed with various HP (differing by gV spacer sequence) and RP were grown to OD=0.3 and induced with lOmM arabinose and infected with phagemid at l*104cfu mL'1. Phage were harvested 18 hours post infection. The RP makes a discrete AT>CG edit in the transcription factor encoding gene clopt with a 68bp THR. RP with a medium copy number origin of replication (CloDF13, ~40 copies / cell) was used. Data represent average of five biological replicates. A non-targeting reticulase diversifying rpob together with the WT HP is included as a negative control. (C) Titer data for both WT and 18k-62 conditions described in (B). Data represent mean of 6 biological replicates.
[0044] FIGS. 28A-28C. Reticulase optimization is a codon optimization scheme to enhance hypermutation specificity in genes encoded on extrachromosomal targets. (A) Tabulation of substitutions made by the reticulase optimization codon optimizer batch replace function. Codon frequencies attained from Genscript (https: / / www.genscript.com / tools / codon-frequency-table). (B) Example reticulase optimization of the region E28-V71 in the lambda phage transcriptional operator clopt. Adenines in both the original sequence and reticulase optimized sequence are shown in red. Sequence content comparison tabulated below sequences. (C) Reticulase optimized sequence region of clopt with dual stop codon inserted in place of Y38 and E39 (TAC GAG> TAA TAG). A 68bp reticulase THR programmed to revert this edit is shown below (maroon).Aty. Dkt. No. 125141.04976 MGH2024-443
[0045] FIGS.29A-29D. Reticulases enable rapid generation of antibody libraries encoded on M13 tracemid. (A) Crystal structure of Trastuzumab in complex with Human Epidermal growth factor 2 (Her2) (PDB:1N8Z). Her2 antigen colored blue, light chain colored red, heavy chain colored salmon. The residue H91 in the CDR3 loop of the heavy chain makes direct contact with antigen and is shown as space-filling spheres. (B) Titering of a multi- passage hypermutation of Trastuzumab H91 encoded on a (vl M13+ ColEl+) tracemid backbone. s2061 cells transformed with the HP and RP were grown to OD=0.3 and induced with lOmM arabinose and infected with phagemid at l*104cfu m ’1. Phage were harvested 18 hours post infection and back-diluted 10’3to infect a new batch of host cells for a subsequent round of hypermutation. The RP encodes a 73bp THR set to diversify codon 91 in Trastuzumab with full saturation (AAA). RP with a medium copy number origin of replication (CloDF13, ~40 copies / cell) was used. Trastuzumab was encoded on the phagemid in scFv format. Data represent average of six biological replicates. (C) Percentage of phagemid after each hypermutation round described in (B) with codons other than CAT (the endogenous codon used to encode histidine) at Trastuzumab position 91. Data represent average of six biological replicates. (D) Incidence of phagemid encoding each amino acid substitution at Trastuzumab position 91 in the total phagemid population after each passage described in (B) and (C). Passage 1 and 2 data generated by allelic mapping of 80,000 reads / sample, passage 3 data generated by allelic mapping of 100,000 reads / sample. Data represent average of six biological replicates.
[0046] FIGS.30A-30B. Reticulases enable full codon saturation mutagenesis of select codons in antibody CDR loops encoded on M13 tracemid. (A) Substitution frequency across the light chain CDR3 loop region of Trastuzumab for passage three phage from experiment described in FIG. 24. The 73bp THR used to diversify position 91 is shown in red. CDR3 loop residues are bolded. Substitution frequency of phage passaged in the targeting condition shown in purple. Substitution frequency of phage passaged in the non-targeting condition (Reticulase set to diversify rpob) shown in blue. Data represent average of six biological replicates, sequenced with 100,000 reads / sample. (B) Incidence of each of each of the 63 possible codon substitutions at Trastuzumab position 91 (all codons but CAT) in the overall phagemid population for the experiment described in (A). Data represent average of six biological replicates, sequenced with 100,000 reads / sample.
[0047] FIG. 31. Targeted Reticulase-Assisted Continuous Evolution (TRACE) combinesAty. Dkt. No. 125141.04976 MGH2024-443reticulate fibonasion and conditional M13 propagation circuits to enable the continuous, rationally directed evolution of biomolecules. Schematic of multi-passage TRACE experiment to evolve desirable variants of an arbitrary protein of interest. Selection tracemid (ST, light green) are cloned to encode a gene of interest (GOI, red) which expresses a protein of interest (POI, red). s2061 cells (tan) are transformed with a helper plasmid (HP, orange) lacking gVI (the gene encoding M13 protein pVI), an accessory plasmid (AP, blue) encoding a conditional M13 phage propagation circuit in which expression of gVI is contingent on some desirable activity of the POI, and a reticulase plasmid (RP) which diversifies user-specified positions in the GOI encoded on the ST. The s2061 cells transformed with the HP, AP, and RP (hereafter described as host cells) are grown to OD=0.3 infected with SP and induced with lOmM arabinose. ST that encode more fit POI that can more strongly activate the synthetic circuit on the AP are rewarded with pVI, which stabilizes the tail end of the phage coat and is essential for phage propagation. Those phage which have less fit POI that are unable to activate the synthetic circuit and are penalized and fail to propagate to the same extent as those phage with fitter POI variants. ST from each passage are back-diluted and used to infect fresh host cells to start a subsequent passage. Over multiple serial passages, those ST encoding the most fit POI variants come to dominate the population.
[0048] FIG.32. Schematic of typical TRACE experiment workflow. A variety of methods may be used to nominate positions in a POI that may improve its fitness for a desired phenotype. Host cells are prepared with an RP (green) encoding a reticulase array programmed to saturate the chosen positions in the GOI (red) encoded on the ST for the duration of the TRACE. The ST each express variant POI (red) divergent at select positions. The synthetic circuit on the AP (blue) continuously selects for SP encoding POI variants that exhibit a desired phenotype; a vast array of synthetic circuits can be used to select for a wide variety of biomolecules with desirable properties. As ST encoding POI variants with favorable gain-of-function mutations at one of the saturated sites come to dominate the population, they serve as the template for RP- mediated saturation at a different site. This process continues, allowing ST to accumulate beneficial gain-of-function mutations at each of the saturated positions. By the end of the TRACE, a full Iterative Saturation Mutagenesis (ISM) campaign has been executed in vivo; the ST encoding the fittest POI variant fully dominates the population and can be readily isolated by phage titering.
[0049] FIGS. 33A-33D. Conditional M13 phage propagation circuits select for reticulase-introduced gain-of-function mutations in a gene of interest encoded on a selection phagemid.Aty. Dkt. No. 125141.04976 MGH2024-443(A) Schematic of the proof-of-concept selection circuit used in this experiment. A selection phagemid (SP, red) encoding a copy of the lambda phage transcriptional operator clopt was cloned with two stop codons inserted in place of Y38 and E39; the gene region around these stop codons was reticulase optimized (see FIG. 23C). A reticulase plasmid (RP, green) encoding a 68bp THR was programmed to revert these stop codons with a discrete AT>CG edit. Edited SP are able to express clopt, which binds to the PR / PRM promoter on the accessory plasmid (AP), leading to expression of Ml 3 pVI. Reticulase-edited phage are expected to dominate the population over multiple serial passages due to the propagation advantage provided by the abundant pVI expressed off the AP in the presence of functional clopt. (B) Titer data for experiment described in (A). s2061 cells (hereafter referred to as host cells) were transformed with the helper plasmid AgVI (HP), the AP, and either a targeting RP or a non-targeting RP set to diversify mCherry. Host cells were grown to OD=0.3, infected with SP at a starting titer of l*106cfu mL'1and induced with ImM arabinose. After 18-22 hours, phage were back diluted l*10'3and used to infect a fresh batch of host cells. This process was repeated for a total of five serial passages. This experiment was performed prior to the discovery of the optimized vl M13+ColEl+ tracemid backbone, so the standard pLITMUS* M13- ColEl+ backbone was used. This experiment was performed at 30°C (unlike all other experiments in this manuscript performed at 37°C), because the clopt transcriptional operator is temperature sensitive and non-active at 37°C. Data from one biological replicate. (C) Percent of phagemid population with stop codons reverted in clopt after each serial passage described in (B). NGS data from 30,000 reads / sample. Data from one biological replicate. (D) Representative full-phagemid Zeroprep sequencing of a single SP colony isolated by phage titering at passage 3. Starting SP genome (reference, top) aligned to the evolved SP genome (Isolated phagemid, middle). Consensus (bottom) between sequences shows dark line at positions with sequence divergence. Magnification of edited area shown, along with the portion of the RP THR (red) which made the discrete AT>CG edit.
[0050] FIG. 34 illustrates the step-by-step reticulase-mediated mutagenesis process for genomes undergoing discontinuous lagging strand / Okazaki fragment genomic replication.
[0051] FIG. 35 illustrates the step-by-step reticulase-mediated mutagenesis process for genomes undergoing rolling circle amplification.
[0052] FIG.36 illustrates a practical example of the reticulase-mediated mutagenesis process.
[0053] FIG.37 illustrates the use of a bioreactor and phagemid for the example shown in FIG. 35.Aty. Dkt. No. 125141.04976 MGH2024-443
[0054] FIGS. 38A-38I: Continuous selection for affinity and solubility-enhancing mutations in a Trastuzumab scFv. (A) Conditional propagation circuit for selection of affinity-enhanced Trastuzumab scFv. scFvs are exported to the periplasmic space, which supports antibody folding and disulfide bond formation; scFvs homodimerize by their GCN4 linker. scFv binding to the target epitope triggers dimerization of CadC transcription factor and expression of essential M13 phage propagation gVI, providing a propagation advantage to ST encoding affinity-enhanced scFv variants. (B) Titers for library generation / affinity maturation experiment with circuit shown in (A); 18-hour passages. Four phage libraries were saturated at position H91 and evolved independently. (C) Dominant scFv genotypes resulting from each evolution experiment. H91Y is a mutation known to increase scFv affinity for this epitope from 166nM to 66nM. H91 Y / T94S is a novel dual mutation that has not been previously reported. (D) Titers of three independent second-round library generation / affinity maturation experiments with circuit shown in (A); 18 hour passages. The H91 Y variant was used as a template and ST were saturated at specified positions in the CDR2 heavy chain. (E) Genotypes emerging from evolution described in (D). (F) AlphaFold 3 predicted structure of Trastuzumab scFv in complex with the H98 peptide. Positions hypermutated in (E) highlighted. (G) Titers for library generation / affinity maturation experiment with circuit shown in (A); 18 hour passages. The signal sequence (SS) has been interrupted by a trans-splicing intein, reducing the quantity of scFv exported to the periplasmic space. Six phage libraries were saturated at position A34 and evolved independently. (H) Genotypes for the experiment described in (G) at passage 4 and passage 8. (I) Alphafold 3 predicted structure of Trastuzumab scFv in complex with the H98 peptide. Substitutions made in replicate three of (H) specified.
[0055] FIGS.39A-39E. RT phylogeny of Diversity Generating Retroelements. (A) (Figure and legend adapted from: Paul, B. G. & Eren, A. M. Eco-evolutionary significance of domesticated retroelements in microbial genomes. Mobile DNA 13, 6 (2022)) Maximum-likelihood phylogenetic tree of RT representatives aligned with ANMV-1 and DUSEL4 Nanoarchaeota sequences. Green branches correspond to bacterial and bacteria-derived RTs (from chromosomes, plasmids, mitochondria, chloroplasts and bacteriophage), red branches indicate archaeal and archaeal virus RTs, and black branches represent RTs from eukaryotes and their viruses. Retroelement clades and key representatives are labelled as follows: DGRs, diversity-generating retroelements; DIRS, Dictyostelium retrotransposons; GemV, geminiviridae; G2L, group-II intron-like (G2L are numbered according to Simon and Zimmerly 24)); Hpdn, hepadnaviruses;Aty. Dkt. No. 125141.04976 MGH2024-443LTR, long terminal repeat retroelements; NPV, nucleopolyhedraviruses; non-LTR, non-long terminal repeat retroelements; RtV, retroviridae; unk, unknown or unclassified. The scale shows substitutions per site. For clarity, bootstrap values are not shown for the full RT tree. (B) (Figure and legend adapted from: Paul, B. G. & Eren, A. M. Eco-evolutionary significance of domesticated retroelements in microbial genomes. Mobile DNA 13, 6 (2022)). Expanded subtree view of DGR RT representatives. NCBI accession codes are given for representatives in the subtree, but previously described bacterial DRs are explicitly named. The representative for Bordetella phage BPP is labelled ‘BPP’. Colored circles at internal nodes indicate branch support. (C) (Figure and legend adapted from: Roux, S. etal. Ecology and molecular targets of hypermutation in the global microbiome. Nat Commun 12, 3076 (2021)). Phylogeny of DGR and non-DGR reverse transcriptases (RT). RT protein sequences were first grouped into “RT clusters”, and a representative for each cluster was included in the tree building process. Branches are colored according to the type of RT in the corresponding cluster. All nodes with support <50% were collapsed. From inside to outside, the outer rings display the consensus genome type, taxonomic classification, and biome of each RT cluster. CPR: Candidate Phyla Radiation. DP ANN Diapherotrites, Parvarchaeota, Aenigmarchaeota, Nanoarchaeota, Nanohaloarchaeota, FCB Flavobacteria, Fibrobacteres, Chlorobi, Bacteroides, PVC Planctomycetes, Verrucomicrobia, Chlamydiae, Aq aquatic, Te Terrestrial, En Engineered, H-a Host-associated. NA corresponds to cases for which the feature could not be estimated. (D) (Figure and legend adapted from: Roux, S. etal. Ecology and molecular targets of hypermutation in the global microbiome. Nat Commun 12, 3076 (2021)). Prevalence and sequence characteristics of the most abundant DGR target protein clusters (PCs). The 24 PCs listed here represent >92% of all identified DGR targets. Distribution across DGR clades and relative position of the VR region within the target sequence shown. The boxplot lower and upper hinges correspond to the first and third quartiles, respectively, and the whiskers extend no further than ±1.5 times the interquartile range. (E) (Figure and legend adapted from: Paul, B. G. & Eren, A. M. Eco-evolutionary significance of domesticated retroelements in microbial genomes. Mobile DNA 13, 6 (2022)). Conserved and putative regulatory features of Nanoarchaeota DGRs. IMH sites (IMH and IMH*) are shown as yellow boxes, and the trinucleotide-loop hairpin is given in an expanded view at right. Dark grey arrows indicate ORFs between RT and TP whose amino-acid sequences have comparable isoelectric point and molecular weight to accessory variability determinant (Avd; pl 'A 9±1; Mw % 10±5).Aty. Dkt. No. 125141.04976 MGH2024-443
[0056] FIGS. 40A-40B. Comparative Architecture of the Bordetella hinzii and Bordetella pertussis DGRs. (A) Suspected diversity generating retroelement cassette in Bordetella hinzii strain FY01 (gb|CP049736.1|: 1093595-1096644). Open reading frames of Major trophism determinant (Mtd), Accessory Variability Determinant (Avd) and Reverse Transcriptase (bRT) annotated. Regions suspected to constitute homing hairpin (red) as well as RNA structural features (blue) annotated. Regions with clear conservation to RNA structural features in pertussis DGR shown in beige. IMH: Initiation of Mutagenic Homing. VR: Variable region. TR: Template Region. UTR: Untranslated Region. (B) Architecture of the well-characterized BPP-1 Bordetella pertussis DGR. Same coloring as in A.
[0057] FIGS. 41A-41D. Constructing a novel reticulase out of Bordetella hinzii DGR components. (A) Comparative editing efficiencies of Bordetella hinzii and Bordetella pertussis reticulases. Reticulase constructs were constructed with either hinzii or pertussis Avd-bRT proteins, hinzii or pertussis regRNAs and with or without the 22bp 5’GCT spacer between the 5’UTR and the THR. A 73bp THR reverting a stop codon in Trastuzumab CDR3L was used. Tracemid infected at a final titer of le6 cfu / mL and induced with 50uM arabinose. Cells harvested 20 hours post-induction and subject to NGS. Data average of eight biological replicates. (B) Substitution rates on Tracemid edited by the highest-performing hinzii reticulase in a. 73bp THR shown. (C) Substitution rates on Tracemid edited by the highest-performing pertussis reticulase in a. 73bp THR shown. (D) Substitution rates on Tracemid in negative control condition with a bRT knockout. 73bp THR shown.
[0058] FIGS. 42A-42B. Overview of conditional propagation circuit for selection of affinity- enhanced Trastuzumab scFv variants, as described by Morrison et al.2021 for phage-based vectors. (A) The periplasmic space supports antibody folding and disulfide bond formation. A selection phage (SP) encodes an Trastuzumab scFv that is exported to the periplasmic space by an N-terminal PhoA signal sequence (SS), which is cleaved off after passing through the SEC translocon. ScFvs homodimerize by their C-terminal GCN4 linker. An accessory plasmid 2 (AP2) expresses a target epitope (H98) tethered to the transmembrane transcription factor CadC. scFv binding to the target epitope triggers dimerization of CadC transcription factor in the cytoplasm. Dimerized CadC binds to the Cadi and Cad2 operators in the PcadBA promoter on the accessory plasmid 1 (API), triggering expression of essential M13 phage propagation gene gill. scFv variants with enhanced binding to the target antigen activate this circuit to a greater degree,Aty. Dkt. No. 125141.04976 MGH2024-443providing a propagation advantage to SP encoding affinity-enhanced scFv variants. (B) Genetic architecture of circuit shown in a. Origins of replication shown in grey, resistance genes shown in light green, phage genes shown in light red, gill shown in dark blue, gVI shown in dark purple, signal sequence shown in light purple, CadC cytoplasmic DNA binding domain shown in dark green, target antigen shown in orange.
[0059] FIGS. 43A-43B. Overview of conditional propagation circuit for dual selection of affinity and solubility enhanced Trastuzumab scFv variants, as described by Morrison et al.2021 for phage-based vectors. (A) A trans splicing intein from Nostoc puiictiforme (Npu) is inserted between amino acids 8 and 9 of the PhoA signal sequence. The PhoA(l-8)-NpuN protein is expressed off the AP2, while the NpuC-PhoA(9-21) is encoded on the SP upstream of the scFv. The Npu intein must splice prior to scFv periplasmic export, decreasing scFv quantity in the periplasmic space and placing a selection pressure on scFv solubility. The remainder of the circuit remains the same as described in FIG. 42. (B) Genetic architecture of circuit shown in a. Origins of replication shown in grey, resistance genes shown in light green, phage genes shown in light red, gill shown in dark blue, gVI shown in dark purple, signal sequence shown in light purple, Npu segments shown in light green, CadC cytoplasmic DNA binding domain shown in dark green, target antigen shown in orange.
[0060] FIGS. 44A-44C. Summary of genetic and biological components for periplasmic evolution of scFvs encoded on phage vectors, as described by Morrison et al. 2021. (A) Selection phages (SP) for periplasmic scFv evolution from FIG. 42 and 43. Origins of replication shown in grey, resistance genes shown in light green, phage genes shown in light red, gill shown in dark blue, gVI shown in dark purple, signal sequence shown in light purple, Npu segments shown in light green, CadC cytoplasmic DNA binding domain shown in dark green, target antigen shown in orange. B) Accessory plasmids (AP) for periplasmic scFv evolution from FIG. 42 and 43. The API may be used with either AP2. colored as in a. (C) E. coli host strains s536 / sl367 used for periplasmic scFv evolution, derived by deletion / replacement of the CadCBA operon with a kanamycin resistance marker in PACE host strains sl030 / s2060, respectively.
[0061] FIGS. 45A-45E. Comparative summary of genetic components for selection of novel cl transcription factor variants with activity on orthogonal PRM lambda phage promoters, as described by Brodel et al.2016 for phagemid vectors. Vectors colored as in FIG. 42 and 43. (A) Wild-type M13 phage with all relevant promoters and genes labeled for reference. (B) GenericAty. Dkt. No. 125141.04976 MGH2024-443Selection Phage (SP) commonly used in PACE experiments. The SP has only two differences from a WT M13 phage: 1) a gene of interest (GOI) replaces gill, 2) a synthetic promoter (PgVI synth) is used to express gVI. (C) Helper plasmids for the generation of Selection Phagemid (PM). Helper plasmid 134351 (HP casette gll-glV Aglll) provides gVI indiscriminately and is used to generate / amplify the titer of all phagemid in a population. Helper plasmid 80840 (HP casette gll-glV Aglll AgVI) does not express gVI and is used in concert with an accessory plasmid (AP) to facilitate phagemid production during continuous selection. Significant homology tracts (>=14bp) between the Helper plasmid (HP) and Selection Phagemid (PM) capable of facilitating HP -PM recombination highlighted in green; recombination sites share numbering such that site 1 on HP shares perfect homology with site 1 on the PM. (D) Selection phagemid for selection of novel cl transcription factor variants. (E) Accessory plasmid with gVI expression under the control of the PRM lambda phage promoter. This AP selects for phagemid encoding cl transcription factors with activity on the PRM lambda phage promoter variant. Note: The commercially available TGI cell line was used by Brodel et al to facilitate phagemid selection. These authors did not perform any engineering or modification of this strain.
[0062] FIGS. 46A-46E. Summary of genetic and biological components for periplasmic evolution of scFvs encoded on phagemid vectors (this work). All novel, non-obvious innovations made to genetic and biological components described by Morrison et al. or Brodel et al. are numbered and highlighted in red. Numbering explained in main test. Color scheme same as in previous figures. (A) Helper plasmid for the generation of selection tracemid (ST). Derived through modification of helper plasmid Addgene 134351 (HP cassette gll-glV Aglll) provides gVI indiscriminately and is used to generate / amplify the titer of all ST in a population. (B) Selection tracemid (ST) used for periplasmic scFv evolution. Top: ST for scFv affinity circuit (depicted in FIG. 42A). Bottom: ST for scFv dual affinity-solubility circuit (depicted in FIG. 43A). Derived through extensive modification of Addgene 80852. (C) Accessory helper plasmid (AHP) and accessory plasmid 2 (AP2) for periplasmic scFv affinity circuit. Analogous to API and AP2 (respectively) in FIG. 42. The helper plasmid cassette in AHP is derived from Addgene 80840 (HP cassete gll-glV Aglll AgVI), which does not express gVI. (D) Single-vector AHP constructs for periplasmic scFv affinity circuit (top) and periplasmic scFv dual affinity-solubility circuit (bottom). The helper plasmid cassette is the same as in c. (E)E. coli TRACE host strains s!037 / s2070 used for periplasmic scFv evolution, derived by chromosomal deletion / replacement of CadC andAty. Dkt. No. 125141.04976 MGH2024-443PcadBA with a gentamycin resistance marker in TRACE host strains s1035 / s2061, respectively. The weak constitutive promoter ProA is used to drive cadBA expression in these strains.
[0063] FIG.47. Illustration of the CadCBA operon in E. coli. Cartoon described in main text.In response to low periplasmic pH (<6.6), the Cad operon replaces periplasmic lysine (pKa=10.5, more acidic) with periplasmic cadaverine (pKa=10.25, more basic), restoring pH homeostasis. This process also consumes an intracellular proton, helping to maintain the proton motive force (PMF) across the inner membrane.
[0064] FIGS.48A-48F. Constitutive cadBA chromosomal expression rescues high-efficiency transformation of helper plasmids in TRACE host strains with CadC / PcadBA knockouts.Plate images of demonstrating representative transformation efficiencies of either (A, B, C) a control plasmid (sclOl, spec) or (D, E, F) a full helper plasmid (sclOl, spec, M13 gll-gVI Aglll). Cells on plates a and b have WT CadCBA locus. Cells on plates B and E are ACadCBA. Cells on plates C and F are ACacC and APcadBA, with the weak constitutive promoter driving cadBA expression (Innovation 15). A gentamycin resistance marker was left in the genome of ACadCBA / ACacC, APcadBA, ProA-cadBA strains to ensure no cross-contamination could occur with other strains in the lab. Images taken 20 hours post-electroporation.
[0065] FIGS. 49A-49B. TRACE strain s2070 (ACacC, APcadBA, ProA-cadBA) supports comparable ST production to WT TRACE strain s2061. (A) Cartoon of protocol to produce data in (B). Cells grown to OD=0.4, ST infected at titer of l*104cfu m ’1, phage harvested and titered after 18 hours. (B) Production of phage in TRACE host strains s2061 or s2070. The pMY0045 HP has the medium copy pl5a origin of reapplication; the pMY0168 HP has the low copy sclOl origin of replication (Innovation 1). Data represent average of 4 biological replicates.
[0066] FIGS. 50A-50D. Chromosomal deletion of CadC is essential for periplasmic scFv circuit performance in TRACE host strain sl037. (A) Cartoon of periplasmic scFv circuit as adapted for phagemid-based vectors (ST). Abbreviations as described in previous figures. (B) Diagramed components for the circuit depicted in A. (C) Change in titer following one passage for either an ST encoding a WT Trastuzumab scFv (160nM binder for H98 peptide), or an affinity matured Trastuzumab scFv (66nM binder for H98 peptide). WT TRACE host strain (si 035) or ACadC TRACE host strain (si 037) were infected with each ST at le7; ST were harvested 18 hours post infection (Pl) and titered to measure the change relative to the starting titer (P0). The change in titer (P1 / P0) for the WT ST is compared to the change in titer (P1 / P0) for the H91Y ST to assessAty. Dkt. No. 125141.04976 MGH2024-443the performance of the circuit. Data represent average of eight biological replicates. (D) Log circuit activation for experiment described in C calculated via Log[H91Y(Pl / P0)]-Log[WT(Pl / P0)] for either the WT or ACadC TRACE host strain.
[0067] FIGS. 51A-51C. PcadBA-gVI antisense orientation relative to the HP cassette is essential to achieve CadC-dimerization induced gVI expression on the AHP. (A) Cartoon of constitutive CadC dimerization circuit. The AP2 expresses CadC with the arabinose-inducible Pbad promoter tethered to either the homodimerizing GCN4 leucine zipper or a dimerizationincompetent GCN4* mutant (7P;14P); comparison with GCN4* provides a control fortrue CadC-dimerization dependent binding as opposed to high CadC expression or cellular metabolic changes. The AHP encodes gVI under the PcadBA promoter. An irrelevant ST encoding a transcription factor was used to measure the degree of gVI expression off the AHP in each condition. (B) Diagramed components for the AHP in either the sense or antisense orientations. The s7002 strain contains the antisense AHP (Innovation 11), the s7004 strain contains the sense AHP. (C) ST produced in each condition for experiment described in a and b. Data generated from two biological replicates.
[0068] FIGS.52A-52D. Split-Vector AHP-AP2 approach for the selection of functional scFv variants. (A) Cartoon of split vector AHP-AP2 circuit, analogous to that described by Morrison et al in FIG. 42. (B) Diagramed components for the AHP, AP2 and ST for circuit depicted in a. The ProB or ProC promoters on the AP2 are highlighted. Negative control AP2 not shown. The scFv variants (WT or non-functional WT with stop codons inserted into CDR3 region) labeled. (C) Change in titer following one passage for each condition. ST was infected at a titer of 5e6 and harvested / titered 18 hours post-infection. Change in titer calculated as in FIG. 50. (D) Log circuit activation for experiment described in (C) calculated via Log[H91Y(Pl / P0)]-Log[WT(Pl / P0)] for each condition. All conditions average of three biological replicates except no AP2, which is average of two biological replicates.
[0069] FIGS. 53A-53B. Altering gVI translational level does not alter periplasmic scFv circuit dynamic range. (A) Identical circuit to that described in FIG. 52 with altering rbs used to initiate translation of gVI on the AHP1. Change in titer following one passage for either an ST encoding a WT Trastuzumab scFv (160nM binder for H98 peptide), or an affinity matured Trastuzumab scFv (66nM binder for H98 peptide). ST infected at titer of 5e6 and harvested / titered 18 hours post infection. Change in titer calculated as in FIG. 50. (B) Log circuit activation forAty. Dkt. No. 125141.04976 MGH2024-443experiment described in a calculated via Log[H91 Y(Pl / P0)]-Log[WT(Pl / P0)] for each condition. All conditions average of three biological replicates.
[0070] FIGS. 54A-54E. Antisense orientation of the ProC-CadC-H98 cassette relative to the PcadBA cassette doubles the dynamic range of the single-vector AHP. (A) Diagram of antisense (Innovationl3) and sense single vector accessory plasmids. Identical circuit and protocol to that described in FIG. 50. rbs driving CadC expression shown in red box. (B) Change in titer following one passage for either an ST encoding a WT Trastuzumab scFv (160nM binder for H98 peptide, pHMY0392), or an affinity -matured Trastuzumab scFv (66nM binder for H98 peptide, pHMY0436). ST infected at titer of le7 and harvested / titered 18 hours post infection. Change in titer calculated as in FIG. 50. rbs with strength 0.2 relative translational units used to drive CadC. (C) Log circuit activation for experiment described in (B) calculated via Log[H91Y(Pl / P0)]-Log[WT(Pl / P0)] for each condition. All conditions averaged across four biological replicates. (D) Change in phage tiers resulting from titration of rbs strength used to drive CadC; identical experiment to that described in b; average of 3 biological replicates. (E) Log circuit activation for experiment described in D.
[0071] FIGS. 55A-55D. Further improvements to the antisense orientation single-vector AHP. (A) Change in titer following one passage for either an ST encoding a WT Trastuzumab scFv (160nM binder for H98 peptide), or an affinity -matured Trastuzumab scFv (66nM binder for H98 peptide). ST was infected at a titer of le7 and harvested / titered 18 hours post-infection. Change in titer calculated as in FIG. 50. rbs with strength 0.13 relative translational units used to drive CadC. Comparison of an AHP with either the full 600bp PcadBA promoter or a minimal 289bp PcadBA promoter (Innovation 12). (B) Log circuit activation for experiment described in B calculated via Log[H91Y(Pl / P0)]-Log[WT(Pl / P0)] for each condition. All conditions average of four biological replicates. (C) Change in ST tiers on AHP with minimal 289bp PcadBA promoter and Bba_B1006 terminator downstream of CadC swapped out for lambda phage terminator; identical experiment to that described in A; average of eight biological replicates. (D) Log circuit activation for experiment described in C.
[0072] FIGS. 56A-56D. Consolidation of all dual affinity-solubility circuit components onto a single AHP vector. (A) Cartoon of dual affinity- solubility circuit with the single-vector AHP, analogous to that described by Morrison et al in FIG. 43. (B) Diagramed components for the AHP and ST for circuit depicted in A. Polycistronic expression of SS(l-8)-NpuN and CadC-H98Aty. Dkt. No. 125141.04976 MGH2024-443construct with the ProC promoter enables the elimination of an operon and straightforward consolidation of all components onto exiting AHP architecture (Innovation 14). (C) Change in titer following one passage for each condition. ST infected at titer of le7 and harvested / titered 18 hours post infection. Change in titer calculated as in FIG. 50. pHMY0414 is ST with WT scFv. pHMY0415 is ST with A34D;Y49S scFv, reported by Morrison et al. to have a strong selective advantage on this circuit. (D) Log circuit activation for experiment described in c calculated via Log[A34D;Y49S(Pl / P0)]-Log[WT(Pl / P0)] for each condition. All conditions average of three biological replicates.
[0073] FIGS. 57A-57C. High-stringency phagemid circuits rapidly select for ST-HP recombinants. (A) ST with either WT Trastuzumab scFv (pHMY0206), WT scFv with no PhoA signal sequence (pHMY0293) or a 1:100 heterogeneous population of these two vectors (1:100 206:293) were passaged through the AHP1 / AP2 ProB circuit described in FIG. 52. The no PhoA signal sequence phage should never be able to activate the periplasmic circuit and act as a negative control. (B) Full phagemid sequencing of negative control ST isolated at Passage 2. ST 1 is the expected 5125bp vector with no recombination. ST 2 is a 12,302bp ST-HP recombinant formed through recombination during ST preparation. The 69bp gill homology region shared by ST and HP is highlighted in a red box on both vectors. (C) Sequence view of recombination region between the ST and HP.
[0074] FIGS. 58A-58E. Homology regions between Addgene vectors 80852 (Phagemid) and 134351 (helper plasmid).
[0075] FIGS. 59A-59C. Development of recombination-resistant ST-HP pairings for high-stringency continuous selection on ST vectors. (A) Diagram of ST-HP vectors. Innovations made to either helper plasmid (Addgene 134351) or phagemid backbone (Addgene 80852) highlighted in red, same numbering as in main test and FIG. 46. Regions with homology that required removal on the other vector labeled in lilac with the same number as corresponding innovation. For instance, innovation 5 on the phagemid (homology removal labeled in red) was necessary due to presence of homologous sequence (5 lilac) on the HP. Each innovation described in main text. ST with the dual M13+ ColEl+ orientation (innovation 6) were not tested in the data on this figure; instead, the first-generation ST backbone (ColEl- M13+, STI) was tested, thus STI is labeled as “relevant to innovation 6.” (B) Fold titer expansion after 20 hours for phagemid / ST infecting si 035 with an RP and either a standard HP (low copy sclOl origin innovation 1 only) orAty. Dkt. No. 125141.04976 MGH2024-443a homology removed HP (innovations 1 ,2,3). Phagemid / ST tested either standard (Addgene 80852 backbone) or innovation 5 (Promoter gll removal), either with M13- ColeEl+ (standard phagemid orientation, PM) or ColEl-M13+ (first-generation ST orientation, STI (relevant to innovation 6)). Four biological replicates titered. (C) Reticulase editing efficiency for experiment described in b; eight biological replicates sequenced. Additional homology removal (innovations 4 and 9) did not discernibly change either titer expansion or editing efficiency from those of the ST in the data shown here.
[0076] FIGS. 60A-60C. Optimization of gl rbs sequence on the AHP HP cassette to remove ST-AHP homology in the gill region. (A) Diagram of AHP, AP2 and ST relevant to this figure. The region upstream of gl (10, red box) has 69bp homology to gill on the ST (10, lilac) and could theoretically support ST-AHP recombination. An AP2 with a rbs of 0.13 relative translational units was used in this figure. (B) ST encoding either the WT or H91 Y Trastuzumab variants were used to test the dynamic range of the circuit as previously described in FIG. 52. The s3701 strain contains the standard AHP with no homology removal. The s3704 strain has an AHP the gill homology removed with the native gill rbs being used to drive gl expression (innovation 10). The s3707 strain has an AHP with the gill homology region removed and the native gVI rbs used to drive gl expression. Data represent three biological replicates. (C) Log circuit activation for experiment described in b, calculated as in FIG. 52.
[0077] FIGS. 61A-61C. Recombination between the ST and AP2 in a homologous promoter region enables a novel cheating strategy. (A) Diagram of AHP, AP2 and ST relevant to this experiment. The pHMY0320 ST encodes the WT Trastuzumab scFv. The pHMY0321 ST encodes the WT scFv with no signal sequence and serves as a negative control. The homologous ProB and Pro3 promoter regions on the AP2 and ST are boxed in green. An rbs with relative translational strength 0.2 was used to express CadC-H98. (B) Titering of a three-passage experiment with the ST and AHP / AP2 described in a. ST harvested after an 18-hour passage were used to infect the subsequent passage. Each graph represents a single biological replicate. (C) Full-ST sequencing of ST isolated at passage 3 replicate 2 in experiment described in (B). Sequence close-up (top) of homology region highlighted in red on full-phagemid map. 93bp homology region overscored in red.
[0078] FIGS. 62A-62C. scFv expression on the ST with the promoter PgVI-synth yields a larger circuit activation than with the Pro3 promoter and removes AP2 / ST homology. (A)Aty. Dkt. No. 125141.04976 MGH2024-443Diagram of ST tested in this figure. The circuit used is the same as in FTG. 52. An rbs with relative translational strength of 0.048 was used to express gVI on the AHP1; the Pro3 promoter and an rbs with relative translational strength of 0.2 was used to express CadC-H98 on the AP2. A WT and Mature (H91Y) variant were tested for each ST to assess the relative circuit activation produced by each promoter. (B) Change in titer for experiment described in a. Data average of three biological replicates. (C) Log circuit activation for experiment described in A and B. The Psynth-gVI promoter was used for all future experiments (Innovation 7).
[0079] FIGS. 63A-63C. Mock selection of affinity-matured Trastuzumab variants encoded on ST vectors. (A) Cartoon of mock-selection experiment. A heterogenous population of ST is prepared in which the H91 Y scFv variant is mixed in 1 : 1000 with ST encoding the WT scFv. The population is passaged and re-infected into fresh host cells. ST populations are titered and sequenced at each passage to determine genotypic makeup of the population. The AHP / AP2 circuit from FIG. 52 was used; a rbs with relative translational strength of 0.13 was used to drive expression of CadC-H98 on the AP2. (B) Titers of ST for experiment described in a. Data from three biological replicates. (C) Representative sequencing of ST isolated at passage 3 (replicate 2 sequencing shown).
[0080] FIGS. 64A-64D. Discovery and engineering of evolution-resistant PhoA signal sequences for regulated scFv export. (A) ST titers for a homogeneous population of ST encoding only the WT scFv for the experiment described in FIG. 63. (B) Sequencing of signal sequence mutations in ST at passage three leading to enhanced circuit activation. (C) Incorporation of evolved signal sequence variants into ST. WT SS-scFv construct expressed with an rbs of relative translational strength of 0.2; P17L or T19I SS-scFv constructs expressed with a rbs of relative translational strength of 0.01. Circuit activation measured with WT (parent) and H91Y (mature) scFv variants on same circuit as in FIG. 52. The ProC promoter and an rbs with relative translational strength of 0.13 was used to express CadC-H98 on the AP2. ST titered 18 hours postinfection. Data from three biological replicates. (D) Log circuit activation for experiment described in C. The T19I mutation signal sequence was used in all future experiments (Innovation 8).
[0081] FIGS. 65A-65C. Optimal single-strand annealing protein activity balances reticulase editing efficiency with efficient M13 Tracemid titer expansion. (A) Diagram of constructs used. A helper plasmid (HP) expresses all genes required for M13 phage replication except gill. The selection tracemid (ST) encodes the target of the reticulase downstream of the two origins ofAty. Dkt. No. 125141.04976 MGH2024-443replication, a CarbR gene and gill. The reticulase plasmid (RP) encodes the reticulase operon under the arabinose-inducible Pbad Promoter consisting of BPP-1 DGR proteins Avd and bRT, the single-stranded annealing protein (ssap) and dominant negative repair protein MutL*. The reticulase evolutionary guide RNA (regRNA) is expressed constitutively and encodes the target homology region (THR not shown) programmed to revert a stop codon in the gene of interest on the ST. In this experiment, an ST encoding either a clopt transcription factor or Trastuzumab scFv are used are the target. (B) Editing efficiency for a panel of eight reticulases with different ssap tested on the two target ST described in a. Cells grown till OD 0.4 and induced with arabinose and infected with l*10A4 cfu / mL phage; phage harvested after 20 hours. A CspRecT reticulase with a non-targeting regRNA is used as a negative control. Data from eight biological replicates. (C) ST titers for experiment described in a and b. Dotted line represents starting infection titer. Data from eight biological replicates.
[0082] FIGS. 66A-66B. Confirmatory repetition of CspRecT and EcTRecT conditions for experiment described in FIG. 65. (A) Editing efficiency for both the CspRecT and EcTRecT reticuases targeting the clopt transcription factor-encoding ST. Identical experiment to that described in FIG. 42. Data from eight biological replicates. (B) ST titering for experiment described in a. Dotted line represents starting infection titer. Data from eight biological replicatesDETAILED DESCRIPTION
[0083] Described herein are the methodology and molecular components used to perform the continuous rationally-directed evolution of biomolecules in bacterial cells using a ‘reticulase’ - a synthetic ribonucleoprotein complex. A reticulase facilitates the continuous hypermutation of select genes in bacteria at any gene region, codon, or nucleotide chosen by the user with single base pair precision. This complex can scarlessly introduce a broad range of single, double, and triple substitutions to select codons of an open reading frame, generating immense genetic diversity at user-selected positions of a protein-encoding gene. Reticulases may also be used to diversify select positions of non-coding genes (such as those of ncRNAs), as well as ribosomal binding sites, promotors, and other regulatory elements. Importantly, these hypermutated positions can be chosen by the researcher on a rational basis as those most likely to induce a desirable phenotype or expression profile (as informed by structural, bioinformatic, or other analyses). When used concomitantly with a method applying selection pressure for a desired phenotype (such as aAty. Dkt. No. 125141.04976 MGH2024-443conditional M13 phage propagation circuit), this invention enables continuous rationally-directed evolution of biomolecules in bacterial cells.
[0084] A second envisioned purpose for the reticulase is its use for basic science and biomedical research. The ability of reticulases to facilitate the transcendence of large evolutionary fitness valleys through triple substitutions in select codons enables the generation of genotypes that would be unlikely to arise through natural means or through traditional laboratory methods of random mutagenesis. For instance, positions in an E. coli gene that confer resistance to an antibiotic could be hypermutated by site-saturation mutagenesis to identify novel antibiotic resistance mutations unlikely to arise in E. coli itself due to specific codon usage). Thus, from a broader perspective, this technology could be utilized to rapidly diversify a wide array of biomolecules to uncover the functional roles of key amino acid residues. Thus, the reticulase may also serve as a powerful basic research tool.
[0085] Constructs
[0086] In a first aspect, provided herein is a construct comprising sequences encoding: a diversity generating retroelement (DGR) accessory variability determinant (Avd) protein; a DGR reverse transcriptase (RT); a single-stranded DNA annealing protein (SSAP); and at least one non-coding RNA; wherein the non-coding RNA comprises a target homology region (THR) that comprises homology to a target nucleic acid; and wherein at least one nucleotide in the THR is a non-homologous adenine.
[0087] Diversity generating retroelements (DGRs) are a family of retroelements that were first found in Bordetella pertussis (Bordetella pertussis positive-phase phage 1 (BPP-1)), and have since been found in various bacteria, Archaea, bacteriophages and Archaeal viruses. DGRs benefit their host by continuously hypermutating amino acid positions within the C-terminal domains of specific target proteins. DGR-mediated hypermutation is facilitated by ‘mutagenic retrohomoing,’ in which the accessory variability determinant protein and error-prone reverse transcriptase of the DGR form a complex which generates mutagenized cDNA (containing substantial A to N mutations) using a template RNA transcribed from the DGR template region (TR). This cDNA is then integrated into a variable region (VR) that forms part of the C-terminal region of the target protein’s open reading frame. The VR and TR share near-complete sequence homology, differing only at positions corresponding to adenine bases in the TR. The BPP-1 DGR is currently the most well-characterized DGR. Accordingly, in embodiments, the Avd and RT are derived from BPP-1.Aty. Dkt. No. 125141.04976 MGH2024-443However, other DGR Avd and RT proteins may be used. For instance, the inventors identified a novel DGR in the genome of Bordetella hinzii (species within the Bordetella genus distinct from pertussis) and demonstrated the Avd, RT and DGR-RNA components of this DGR could also be used as part of a reticulase. Additionally, as demonstrated in Example 2, the inventors demonstrate DGR RTs belong to a distinct evolutionary clade relative to other retroelement RTs (e.g., Group II intron maturases, retrons, etc.) and may be characterized by their biochemistry to a greater extent than by their structure. For example, the Bordetella hinzii and Bordetella pertussis DGRs do not share significant sequence homology (about 60% identity) but are able to provide similar function in the reticulase. Therefore, the DGR RT may be characterized by one or more features including, but not limited to, a loose active site (low binding affinity for dNTPs relative to related RTs such as retron or group-II intron maturase RTs), low processivity (can reverse transcribe up to about 300 base pairs, whereas other retroelement RTs can reverse transcribe much longer sequences), and a hydrophobic template site (higher incidence of hydrophobic amino acids surrounding the RNA template base relative to relative to related RTs such as retron or group-II intron maturase RTs). The DGR RT may comprise a protein or RNA sequence that groups within a DGR clade in a maximum-likelihood phylogenetic analysis, as opposed to another retroelement clade. The RNA sequence may be a 3’ UTR sequence. In embodiments, the RT is encoded by a sequence comprising SEQ ID NO: 4 or 6, or a sequence having at least 50% identity to thereto. The RT may have between 50-100% identity to SEQ ID NO: 4 or 6, or any percent identity or range in between. The RT may have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, at least 90%, or at least 95% identity to SEQ ID NO: 4 or 6.
[0088] The Avd works in complex with the RT to mediate mutagenic retrohoming. In embodiments, the Avd is encoded by a sequence having at least 90% identity to SEQ ID NO: 3 or 5.
[0089] The non-coding RNA component (referred to herein as a ‘reticulase evolutionary guide RNA’ or ‘regRNA’) is a short ribonucleotide strand that is composed of the 5’ and 3’ untranslated regions (UTRs) of a core diversity generating retroelement RNA which flank a target homology region (THR) that has sequence homology to a target gene of interest. The core DGR RNA refers to a DGR RNA in which the UTRs are truncated segments of the natural DGR RNA UTRs. These truncated segments are functional for producing cDNA together with Avd and RT. The core DGR RNA comprises 294 nts encompassing only the terminal 20 bp of the 5’ UTR and the first 157 ntAty. Dkt. No. 125141.04976 MGH2024-443of the 3’ UTR. This core DGR RNA is sufficient to produce cDNA, but does not facilitate cDNA retrohoming or integration in vivo. For purposes of the reticulase described herein, only the cDNA production function is needed.
[0090] To select positions for continuous rational evolution within the target gene, nucleotides are replaced with adenine (A) bases in the THR of the regRNA. This is due to the unique error rate activity of the DGR-RT, which averages 0.5%, 0.3%, and 1.6% when reverse transcribing uracil (U), cytosine (C), and guanine (G) bases respectively, but has an error rate of 51.6% when reverse transcribing adenine (A) bases. The non-coding RNA may be between about 200 and about 500 nucleotides long, or any length or range in between. In exemplary embodiments, the non-coding RNA is about 200 nucleotides. The THR may comprise between about 20 and about 400 nucleotides long, or any length or range in between. In exemplary embodiments, the THR is between about 20 and about 100 nucleotides long.
[0091] As used herein, the term “homology” refers to two nucleic acid sequences which share significant sequence identity with one other. For example, the non-coding regRNA is an RNA molecule that shares significant sequence identity in its THR with a portion of a gene of interest, with the exception of selected nucleotides within the THR sequence that are replaced with an adenine.
[0092] The target nucleic acid may encode a protein or an RNA that performs a biological function (e g. a guide RNA). The target nucleic acid may also encode a regulatory sequence (e.g. a ribosomal binding site) that controls the expression of a biomolecule encoded in the RNA.
[0093] Following selection of positions within a target gene, the mechanism by which a reticulase facilitates continuous hypermutation in cells can be divided into two steps. First, error-prone reverse transcription of the THR in the regRNA is facilitated by DGR-RT and DGR-Avd, both of which are necessary and sufficient for this process. Reverse transcription of this region is cis primed off a 2’OH group in the 3’ untranslated region of the regRNA and occurs independently of the presence of a target locus. Due to the selective infidelity of the DGR-RT for adenine (A) bases, these A nucleotides in the THR undergo an average mutagenic rate (relative to the sense strand) of A>A (no change, 48.4%), A>T (22.2%), A>G (17.8%) and A>C (11.6%). Thus, the DGR-RT can be conceived of as an A>N editor. The DGR-RT faithfully synthesizes a reverse complemented cDNA product in the cytoplasm of the cell with largely conserved sequence at all U, C and G bases in the regRNA THR, and fairly random nucleotides at all positions corresponding to an adenineAty. Dkt. No. 125141.04976 MGH2024-443base. Therefore, to enable hypermutation of select DNA nucleotides, codons or gene regions, a researcher can simply replace the corresponding positions within the regRNA THR with adenine bases. For instance, a serine codon TCC present in the target locus can be replaced with AAA at the corresponding position of the regRNA THR, resulting in the synthesis of a reverse complimented cDNA library with essentially NNN at this position. The synthesis of such rationally diversified cDNA libraries occurs in the bacterial cytoplasm when the reticulase components (DGR-RT, DGR-Avd and regRNA) are actively expressed.
[0094] In the second step of reticulase-mediated hypermutation, cDNAs produced in the first step are bound by the reticulase SSAP protein and annealed into the target locus via one of two distinct genomic integration pathways (A or B). Which pathway is used is dependent on the replication mechanism employed by the genome in which the target locus is situated. Pathway A is used to integrate cDNA into genomes undergoing discontinuous DNA polymerization involving the synthesis of Okazaki fragments at a replication fork. This replication mechanism applies to all bacterial and archaeal chromosomes, as well as some plasmids which contain an origin of replication (such as ColEl) that stimulate discontinuous DNA polymerization involving the synthesis of Okazaki fragments at a replication fork. Pathway B is used to integrate cDNA into genomes undergoing rolling circle amplification (RCA), in which one of the two strands in a double-stranded circular genome (often designated the - strand) serves as the template for continuous polymerization of a novel complementary (+) strand while displacing the existing + strand; the new double-stranded + - genome may undergo a second round of RCA, while the displaced + strand will undergo a round of continuously polymerized - strand synthesis to become a complete double-stranded + - genome that can begin undergoing RCA. This replication mechanism is utilized by certain phages (such as Ml 3 phage) as well as some plasmids.
[0095] In pathway A (relevant to genomes undergoing discontinuous polymerization of Okazaki fragments at a replication fork), cDNAs produced in the first step are bound by the reticulase SSAP protein and annealed to transiently single-stranded regions of the target locus formed during discontinuous DNA polymerization of the lagging strand. The cDNAs produced in the first step are designed to be complementary to the lagging strand of the target locus, enabling their integration through the same DNA repair pathway that is utilized for the integration of Okazaki fragments. In pathway B (relevant to genomes undergoing RCA) cDNAs produced in the first step are bound by the reticulase SSAP protein and annealed to the single-stranded + strand displacedAty. Dkt. No. 125141.04976 MGH2024-443during RCA. The cDNAs produced during the first step are designed to be complementary to the + strand, such that they replace a segment of the - strand during - strand synthesis and can serve as a template for subsequent rounds of RCA once a complete + - genome is formed. For both integration pathways A and B, the result of the integration event is a target locus with conventional Watson-Crick pairing at all nucleotide positions corresponding to U, C and G bases in the regRNA THR, but potential mispairing at positions corresponding to an adenine (A) base in the regRNA THR. A subsequent round of DNA replication using the cDNA-derived strand as a template results in a stable double-stranded edit. This integration event is entirely scarless - with the exception of potential substitutions at adenine (A) positions in the regRNA THR - and therefore does not introduce insertions, deletions or other disruptions to target locus sequences required for integration. As such, reticulases are capable of performing scarless rationally-directed hypermutation anywhere in or outside the open reading frame of a gene. Moreover, because the sequence of the target locus is largely conserved post-edit at all positions corresponding to U, C and G bases in the regRNA THR, reticulases are capable of continuous, successive rounds of rationally-directed hypermutation of the same gene over multiple generations of genome replication.
[0096] For synthetic genomes which use both genomic replication mechanisms (such as M13 phagemid with M13 origin of replication and ColEl origin replication) both pathway A and B may be used to integrate cDNA into the target locus in an orthogonal manner to increase efficiency.
[0097] The steps of reticulase-mediated hypermutation for Pathway A are illustrated in FIG. 34. The steps of reticulase-mediated hypermutation for Pathway B are illustrated in FIG. 35.
[0098] SSAPs are proteins that promote reannealing of complementary single-stranded DNA strands, particularly during homologous recombination and double-strand break repair. In exemplary embodiments, the single-stranded DNA annealing protein is CspRecT (from Collinsella stercoris) or E. coli Rac prophage SSAP (EcRecT). However, any suitable SSAP may be used. For instance, the inventors have demonstrated reticulase-mediated phagemid editing with E. coli phage lambda bet (Red Beta), E. coli Rac prophage SSAP (EcRecT), Bacteriophage P22 essential recombination function (P22-Erf), Lactococcus phage ul36 Sak (ul36 Sak), and Staphylococcus phage phi 11 sak4 (phi 11 sak4). In embodiments, the SSAP is encoded by a sequence having at least 90% identity to SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, or 14.
[0099] In the construct, the sequences encoding the Avd, the RT, the SSAP, and the non-codingAty. Dkt. No. 125141.04976 MGH2024-443RNA are operably linked to one or more promoters. The genes encoding the three proteins and RNA component of the reticulase can be expressed in a variety of ways: either in cis (together) or in trans (separately) from one another, off either a bacterial chromosome or extrachromosomal construct (such as a plasmid). Any combination of reticulase components (some, all, or none) can be placed under the control of an inducible external stimulus (such as the arabinose inducible PBAD promoter), while the remaining components may be expressed constitutively. Accordingly, in embodiments, at least one of the sequences encoding the Avd, the RT, or the SSAP is operably linked to an inducible promoter, such as a PBAD promoter. In embodiments, the sequence encoding the regRNA is operably linked to a constitutive promoter, such as the ProD promoter or the Pro3 promoter (from pDAS112). Any suitable inducible and constitutive promoters may be used. Inducible promoters include, but are not limited to, the lactose-inducible lac operon promoter (Phc) and the M13 pIV-inducible phage shock promoter (PPSP). Constitutive promoters include, but are not limited, to the ProC, ProB, or ProA promoters.
[0100] In embodiments, the construct comprises more than one non-coding RNA, and each THR comprises homology to a different nucleic acid. The construct may comprise two, three, four, five, six, etc. non-coding RNAs. Each non-coding RNA may be between about 200 and about 500 nucleotides in length. Each THR may be between about 30 and about 100 nucleotides in length.
[0101] The construct may further comprise a spacer sequence between the sequence encoding the 5’ UTR and the sequence encoding the non-coding RNA or the THR sequence. The spacer sequence may be a 22bp GCT spacer sequence comprising SEQ ID NO: 89.
[0102] The construct may further comprise a 3’ spacer between the THR and the first nucleotide reverse transcribed ((CGGGGCGCGCGGCGTCTG) where the G on the far right is the first nucleotide reverse transcribed.
[0103] As described in the Examples, to reduce or avoid collateral mutagenesis in the target locus, the THR may be selected to target a region of the target locus that has no adenines within about 10 base pairs of the 5’ end and about 10 base pairs of the 3’ end. In other embodiments, the THR may be selected to target a region of the target locus that does have adenines present within about 10 base pairs of the 5’ end and about 10 base pairs of the 3’ end. In this latter embodiment, those adenines in the target locus corresponding to adenines present in the terminal regions of the THR (within about 10 base pairs of the THR 5’ or 3’ end) will exhibit a substantially lower mutagenic rate than those adenines in the target locus corresponding to adenines present in the central regionAty. Dkt. No. 125141.04976 MGH2024-443of the THR (more than about 10 base pairs from either end). This difference in mutagenic rate is due to the preference of the SSAP to selectively integrate cDNAs which exhibit about 10 base pairs of perfect homology to the target locus at their 5’ and 3’ ends. Put another way, those randomly generated cDNAs that do not have mutated adenines in their 5’ and 3’ terminal regions will be integrated into the target locus at a substantially higher rate than those randomly generated cDNAs that do have mutated adenines in their 5’ and 3’ terminal regions. Thus, collateral mutagenesis of some adenine bases in the target locus may be suppressed by using a THR for which the corresponding THR adenines fall within about 10 base pairs of the 5’ or 3’ end of the THR. This strategy is referred to as ‘homology shielding.’
[0104] In order to minimize the length of the non-coding RNA, the 5’ UTR and 3’ UTR flanking the THR may be minimal. In exemplary embodiments, a minimal 5’ UTR is 20 nucleotides in length, and a minimal 3’ UTR is 157 nucleotides in length. For bordetella pertussis DGR:, the minimal 5' UTR is AAGGGCAGGCTGGGAAATAA; and the minimal 3' UTR is TGCCCATCACCTTCTTGCATGGCTCTGCCAACGCTACGGCTTGGCGGGCTGGCCTTTC CTCAATAGGTGGTCAGCCGGTTCTGTCCTGCTTCGGCGAACACGTTACACGGTTCGG CAAAACGTCGATTACTGAAAATGGAAAGGCGGGGCCGACTTC.
[0105] In embodiments, the 3’ UTR is split in half and placed within the construct at two distinct sites that allow for expression of more than one THR from a single RNA without repeating the full 157 nucleotide 3’ UTR after each THR. For example, the 3’ UTR may be divided into a 3’ UTRa segment and a 3’ UTRb segment; wherein the construct comprises a first site and a second site; wherein the construct further comprises a second THR; and wherein at the first site the non-coding RNA comprises, from 5’ to 3’: the 5’ UTR, the first THR, the 3’ UTRa segment, the 5’ UTR, the second THR, and the 3’ UTRa segment; wherein at the second site, the non-coding RNA comprises the 3 ’UTRb and, when expressed, complexes with the 3 ’UTRa segment adjacent to each THR to support reverse transcription of each THR. The second THR may comprise homology to a second target nucleic acid sequence; wherein at least one nucleotide in the second THR in a non-homologous adenine.
[0106] The construct may further comprise a sequence encoding a dominant-negative mismatch repair protein. The protein will reduce or prevent correction of mismatched bases on the target nucleic acid during the mutagenesis process. In exemplary embodiments, the mismatch repair protein is MutL E32K. However, any suitable dominant negative mismatch repair protein may beAty. Dkt. No. 125141.04976 MGH2024-443used. In other embodiments, the reticulase may be used in a strain which has been engineered to lack a functional mismatch repair system (for . coli, such a phenotype may be attained through deletion or disruption of the MutS gene). In such embodiments, expression of the dominantnegative mismatch repair protein MutL E32K as part of the reticulase is mechanistically redundant and unnecessary.
[0107] As described in the Examples, in order to achieve an approximate 5:1 stoichiometry of the Avd and the RT in the cell, the construct may encode an Avd ribosomal binding site (RBS) of SEQ ID NO: 17 and an RT RBS of SEQ ID NO: 18.
[0108] As used herein, the term “construct” refers to a recombinant polynucleotide, i.e., a polynucleotide that was formed artificially by combining at least two polynucleotide components from different sources (natural or synthetic). As used herein, the term “promoter” refers to a DNA sequence that regulates the transcription of a polynucleotide. Typically, a promoter is a regulatory region that is capable of binding RNA polymerase and initiating transcription of a downstream sequence. However, a promoter may be located at the 5’ end, within a coding region, or within an intron of a gene that it regulates. Promoters may be derived in their entirety from a native gene, may be composed of elements derived from multiple regulatory sequences found in nature, or may comprise synthetic DNA segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different cell types, at different stages of development, or in response to different environmental conditions. A promoter is “operably linked” to a polynucleotide if the promoter is connected to the polynucleotide such that it may affect transcription of the polynucleotide.
[0109] The promoter used in the constructs described herein may be a heterologous promoter (i.e., a promoter that is not naturally associated with the encoded protein), an endogenous promoter (i.e., a promoter that is naturally associated with the encoded protein), or a synthetic promoter that is artificially designed to function in a desired manner in a particular host cell.
[0110] The terms “polynucleotide,” “polynucleotide sequence,” “nucleic acid,” and “nucleic acid sequence” refer to a nucleotide, oligonucleotide, polynucleotide (which terms may be used interchangeably), or any fragment thereof. A “polynucleotide” may refer to a polydeoxyribonucleotide (containing 2-deoxy-D-ribose), a polyribonucleotide (containing D-ribose), and to any other type of polynucleotide that is an N glycoside of a purine or pyrimidine base. There is no intended distinction in length between the terms “nucleic acid”, “oligonucleotide”Aty. Dkt. No. 125141.04976 MGH2024-443and “polynucleotide”, and these terms will be used interchangeably. These terms refer only to the primary structure of the molecule. Thus, these terms include double- and single-stranded DNA, as well as double- and single-stranded RNA. For use in the present methods, an oligonucleotide also can comprise nucleotide analogs in which the base, sugar, or phosphate backbone is modified as well as non-purine or non-pyrimidine nucleotide analogs. As used herein, the term “nucleic acid” or “polynucleotide” refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA) and DNA / RNA hybrids. Polynucleotides may be single-stranded or double-stranded. Nucleic acids include, but are not limited to: pre-messenger RNA (pre-mRNA), messenger RNA (mRNA), RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, synthetic RNA, genomic RNA (geRNA), guide RNA, tracRNA, crRNA, sgRNA, plus strand RNA (RNA(+)), minus strand RNA (RNA(-)), synthetic RNA, genomic DNA (gDNA), PCR amplified DNA, complementary DNA (cDNA), synthetic DNA, or recombinant DNA. These phrases also refer to nucleic acids of genomic, natural, or synthetic origin.
[0111] A “recombinant nucleic acid” is a sequence that is not naturally occurring or has a sequence that is made by an artificial combination of two or more otherwise separated segments of sequence. This artificial combination is often accomplished by chemical synthesis or, more commonly, by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques known in the art. The term recombinant includes nucleic acids that have been altered solely by addition, substitution, or deletion of a portion of the nucleic acid. Frequently, a recombinant nucleic acid may include a nucleic acid sequence operably linked to a promoter sequence. Such a recombinant nucleic acid may be part of a vector that is used, for example, to transform a cell.
[0112] As used herein, the terms “complementary” or “complementarity” are used in reference to “polynucleotides” and “oligonucleotides” related by the base-pairing rules as typically seen in the forward strand and reverse strand of a DNA molecule. For example, the sequence “5'-C-A-G-T,” is complementary to the sequence “5 -G-T-C-A.” Complementarity can be “partial” or “total.” “Reverse complementary” refers to a sequence that is complementary to the reverse of another sequence. For example, the sequence “5'-C-A-G-T,” is reverse complementary to the sequence “5'-A-C-T-G.” “Partial” complementarity is where one or more nucleic acid bases is not matched according to the base pairing rules. “Total” or “complete” complementarity between nucleic acids is where each and every nucleic acid base is matched with another base under the base pairingAty. Dkt. No. 125141.04976 MGH2024-443rules.
[0113] A “gene” as used herein refers to a nucleic acid region, also referred to as a transcribed region, which expresses a polynucleotide, such as an RNA. The transcribed polynucleotide can have a sequence encoding a polypeptide, such as a functional protein, which can be translated into the encoded polypeptide when placed under the control of an appropriate regulatory region. A gene may comprise several operably linked fragments, such as a promoter, a 5' leader sequence, a coding sequence and a 3' non-translated sequence, such as a poly adenylation site. A chimeric or recombinant gene is a gene not normally found in nature, such as a gene in which, for example, the promoter is not associated in nature with part or all of the transcribed DNA region. “Expression of a gene” refers to the process wherein a gene is transcribed into an RNA and / or translated into a functional protein.
[0114] The terms “protein”, “polypeptide”, and “peptide” refer to a polymer of amino acids. Typically, a “polypeptide” or “protein” is defined as a longer polymer of amino acids, of a length typically of greater than 50, 60, 70, 80, 90, or 100 amino acids. A “peptide” is defined as a short polymer of amino acids, of a length typically of 50, 40, 30, 20 or less amino acids.
[0115] Plasmids and Phagemids
[0116] In a second aspect, provided herein is a plasmid comprising any of the constructs described herein. The plasmid encoding the components of the reticulase may be optimized to provide a high editing efficiency. In exemplary embodiments, the plasmid comprises a sclOl origin of replication (low copy number). In other exemplary embodiments, the plasmid comprises a CloDF13 origin of replication (high copy number). In embodiments, the plasmid comprises an RK2 origin of replication.
[0117] The components of the reticulase may be provided on one or more plasmids. For example, a first plasmid may comprise a construct comprising sequences encoding the Avd, the RT, and the SSAP; and a second plasmid may comprise the at least one non-coding RNA.
[0118] The term “plasmid” as used herein refers to a small, circular, extrachromosomal DNA molecule that is physically separated from chromosomal DNA in a cell, and can replicate independently.
[0119] Also provided herein are phagemid vectors and phage vectors encoding the target nucleic acids or genes of interest to be mutated by the reticulase. The term “phage vector” as used herein refers to a vector comprising the genes encoding a bacteriophage engineered to carry foreign DNAAty. Dkt. No. 125141.04976 MGH2024-443into a host cell for molecular cloning, gene expression, or phage display. A phagemid vector is a vector that has both bacteriophage and plasmid properties, including an origin of replication derived from bacteriophage as well as a plasmid origin of replication. In exemplary embodiments, the phagemid vector comprises an Ml 3 phage origin of replication. However, other phagemids may be used including, but not limited to fl and fd filamentous phagemids. In exemplary embodiments, the phagemid comprises a pColEl plasmid origin of replication. In other embodiments, the phagemid comprises a pl 5a origin of replication. The phagemid may also encode Ml 3 gill and a resistance marker. In embodiments, the resistance marker is a carbenicillin resistance marker (CarbR). In embodiments, the phage origin of replication and the plasmid origin of replication both initiate replication in the direction of the gene of interest encoded on the phagemid to be mutated by the reticulase. Phagemid with this orientation of the phage and plasmid origins of replication are sometimes referred to as ‘tracemid’ in this work.
[0120] Also provided herein is a helper plasmid. A helper plasmid comprises genes (g) encoding the proteins (p) necessary for replication and packaging of the phagemid DNA into phage particles. The helper plasmid may comprise all of the genes necessary for propagation of a phage vector or a phagemid vector except such genes encoded on the phage vector or phagemid vector or on an accessory plasmid. “Propagation” refers to assembly of structural proteins of the phage (e.g. capsid) and lytic proteins for breaking down the host cell wall, amplification of the phage, and lytic infection of the host cell.
[0121] As described in the Examples, in order to increase phage titer and / or editing efficiency of the reticulase, the helper plasmid may comprise a spacer sequence upstream of a gV start codon, as described in Lee et al. Optimizing protein V untranslated region sequence in M13 phage for increased production of single-stranded DNA for origami. Nucleic Acids Res. 49, 6596-6603 (2021). The spacer sequence should be sufficient to reduce protein V expression, prolonging the Fibonacci expansion phase of the M13 life cycle. In embodiments, the helper plasmid encodes all genes required for phage propagation except gVI and gill. In other embodiments, the helper plasmid encodes all genes required for phage propagation except gill.
[0122] Also provided herein are accessory plasmids. In an embodiment, an accessory plasmid encoding gVI operably linked to a synthetic gene circuit that acts as a sensor for a desired activity of a biomolecule of interest encoded by the target nucleic acid is used in the methods described herein. In some embodiments, the accessory plasmid may encode a secondary or tertiaryAty. Dkt. No. 125141.04976 MGH2024-443biomolecule that interacts with the biomolecule of interest encoded on the phagemid and constitutes a component of the synthetic gene circuit used in the methods described herein. In other embodiments, a second accessory plasmid may be used that encodes a secondary or tertiary biomolecule that interacts with the biomolecule of interest encoded on the phagemid and constitutes a component of the synthetic gene circuit used in the methods described herein.
[0123] In embodiments, the helper plasmid and accessory plasmid are combined onto the same vector to create an “accessory helper plasmid” (AHP). The AHP may comprise all genes necessary for propagation of a phage vector or a phagemid vector except such genes encoded on the phage vector or phagemid vector itself. The AHP constitutively expresses all genes required for phage propagation except gVI and gill and encodes all or part of the components of the synthetic gene circuit which results in the transcription of gVI in response to a target nucleic acid or biomolecule of interest encoded on the phagemid performing a desired function.
[0124] Cells and Phage
[0125] In a third aspect, provided herein is a cell comprising any of the plasmids, phage vectors, or phagemid vectors described herein. In preferred embodiments the cell is a bacterial cell. The cell may be an . coli cell. In exemplary embodiments, the cell is anE. coli s!030 or s2060-derived strain. The cell may be a TGI, EcBS2, EcNR2, EcNR5, or MG1655 E.coli strain. The E. coli genes encoding exonucleases Red and sbcB may be knocked out for improved editing efficiency. The cell may be engineered with a Q576A mutation in DNA gyrase (DnaG) for improved editing efficiency.
[0126] In embodiments the cell is a helper cell, in which the cell comprises in its genome all genes necessary for propagation of a phage vector or a phagemid vector except such genes encoded on the phage vector, phagemid vector, or an accessory plasmid. The helper cell may be used in lieu of a helper plasmid in any of the kits and methods described herein.
[0127] In a fourth aspect, provided herein is an E. coli cell comprising at least one of a RecJ gene deletion; an sbcB gene deletion; and a Q576A mutation in DnaG. TheE coli cell may be an s!030 or s2060 strain.
[0128] In a fifth aspect, provided herein is a phage or phagemid vector comprising any of the constructs or plasmids described herein.
[0129] Kits
[0130] In a sixth aspect, provided herein is a kit comprising any of the phagemid vectors encodingAty. Dkt. No. 125141.04976 MGH2024-443the target nucleic acid and plasmids encoding the reticulase described herein. The kit may further comprise a helper plasmid. In exemplary embodiments, the phagemid is an M13 phagemid, and the helper plasmid is an M13KO7-derived helper plasmid. However, any suitable phagemid / helper plasmid pairings may be used. The M13KO7 helper phage cassete from the M13KO7 plasmid may be inserted into the backbone for sclOl (creating a low copy helper plasmid). Other helper plasmids include, but are not limited to, R408, VCSM13, hyperphage, R408d3, KM13m, and CM13. The kit may further comprise a host cell, such as E. coli cell comprising the mutations conferring enhanced editing efficiency. The kit may further comprise a bioreactor configured for continuous evolution, as described further below. The kit may further comprise any reagents and instructions necessary for using the reticulase-targeted phagemids for saturated mutagenesis and directed evolution.
[0131] The cells, phages, phagemids, plasmids, vectors, etc. in the kits described herein may each be provided in separate containers.
[0132] A bioreactor may comprise a cell culture vessel comprising a population of the phagemid and a population of the host cell; an inflow connected to a chemostat that maintains fresh host cells in mid-log phase (e.g. OD 600 = 0.3-0.4); an outflow connected to a waste container; and a controller that controls the inflow and outflow rates of fresh media and waste through the bioreactor.
[0133] Methods
[0134] In a seventh aspect, provided herein is a method for generating and identifying functional mutants of a starting biomolecule, the method comprising: a) back-diluting and growing an initial fresh culture of host cells to mid log phase, wherein the host cells are transformed with: a helper plasmid comprising all genes required for phage propagation except gVI and gill; an accessory plasmid comprising gVI operably linked to a synthetic gene circuit, wherein when the synthetic gene circuit is activated, the gVI is transcribed; and the plasmid encoding any of the reticulase constructs described herein; b) infecting the host cells with a selection phagemid vector comprising gill, wherein the selection phagemid vector comprises the target nucleic acid; wherein the target nucleic acid encodes the biomolecule; wherein activity of the biomolecule activates the synthetic gene circuit; c) incubating the host cells under conditions that allow for production of an infectious phagemid , wherein infectious phagemid encoding variants of the biomolecule that activate the synthetic gene circuit to a greater extent than the starting biomolecule will receive a propagationAty. Dkt. No. 125141.04976 MGH2024-443advantage; d) harvesting phagemid from the culture and discarding host cells from the culture; e) back-diluting and growing to mid-log phase a subsequent fresh culture of host cells; and infecting the host cells with the harvested phagemid from step d); f) repeating steps b) through e); and g) isolating phagemid encoding variants of the starting biomolecule that inactivate the synthetic gene circuit to a greater extent than the starting biomolecule.
[0135] When a variant provides a propagation advantage, this will increase its frequency in the population of cells in the culture. The phage can be further sequenced to identify mutations that permit an improved ability of the biomolecule to perform a desired activity of interest.
[0136] Alternatively, the method may comprise a) back-diluting and growing an initial fresh culture of host cells to mid log phase, wherein the host cells are transformed with: a helper plasmid encoding all genes required for phage propagation except gVI and gill; an accessory plasmid encoding gVI operably linked to a synthetic gene circuit which acts as a sensor for a desired activity of a biomolecule of interest; a second accessory plasmid which operates as an integral part of the aforementioned synthetic circuit, such as by encoding a biomolecule that acts as a sensor for the desired activity of the biomolecule of interest encoded on the phagemid and facilitates transcription of gVI off the first accessory plasmid; and a reticulase plasmid comprising any of the reticulase constructs described herein; b) infecting the host cells with a selection phagemid; wherein the phagemid encodes a biomolecule of interest and gill; wherein the target nucleic acid of the reticulase is the gene encoding the biomolecule of interest encoded on the phagemid; and wherein a desired activity of the biomolecule results in the activation of the synthetic gene circuit and transcription of gVI from the first accessory plasmid; c) incubating the host cells under conditions that allow for production of an infectious phagemid, wherein infectious phagemid encoding variants of the biomolecule that activate the synthetic gene circuit to a greater extent than the starting biomolecule will receive a propagation advantage; d) harvesting phagemid from the culture and discarding host cells from the culture; e) back-diluting and growing to mid-log phase a subsequent fresh culture of host cells; and infecting the host cells with the harvested phagemid from step d); f) repeating steps b) through e); and g) isolating phagemid encoding variants of the starting biomolecule that inactivate the synthetic gene circuit to a greater extent than the starting biomolecule. The phagemid can be further sequenced to identify mutations that permit an improved ability of the biomolecule to perform a desired activity of interest. A schematic of the disclosed method is provided at FIG. 36.Aty. Dkt. No. 125141.04976 MGH2024-443
[0137] When performing the method in a 96 well plate, the cells are spun down and phage are harvested from the supernatant. VI is required for them to get out of the cell and into the supernatant, therefore any phagemid in the supernatant are those that expressed gVI. Those that didn't express gVI are stuck in the cell and thereby are sequestered in the pellet (thus centrifugation is the key method distinguishing the phagemid encoding non / less functional GOI and those with more valuable / high activity mutations). Once the phagemid is harvested, a fresh population of host cells is infected with the phagemid. Thus, all 'viable' phagemid encoding the gene of interest with higher-activity mutations are distinguished by centrifugation.
[0138] The host cell may be infected with between about IxlO2cfu ml / 1and about IxlO9cfu ml;1of the selection phagemid. In exemplary embodiments, the host cell is infected with about IxlO6cfu mL1of the selection phagemid.
[0139] These phage-based methods adopt similar principles as the phage-assisted continuous evolution (PACE) method, which employs a mutagenesis plasmid to randomly introduce genetic variation into a gene of interest and introduces selective pressures to identify useful mutants. The methods described herein (e.g. TRACE method) are much improved over PACE, primarily due to the use of the reticulase. With the reticulase full saturation mutagenesis of the gene of interest is achieved, allowing for all twenty amino acids to be encoded at any given position. In contrast, PACE relies on random mutagenesis and therefore only about 5-7 amino acids interconversions can be accessed at any given position (via single nucleotide changes).
[0140] Some added features of the constructs and mutations to the host cells described herein are important for the TRACE method functionality. For example, the host cell may be an A’. coli cell comprising at least one of a Red gene deletion; an sbcB gene deletion; and a Q576A mutation in DnaG, as described above.
[0141] In embodiments, the method is performed in a bioreactor in which the first and subsequent populations of host cells are continuously flowed into the culture and wherein old host cells and phagemid are continuously flowed out of the culture at a constant flow rate. The process is illustrated at FIG. 37. The flow rate may be between 0.5 and 3 volumes per hour, or any rate or range in between. In exemplary embodiments, the flow rate is about 0.5 volumes per hour. In other exemplary embodiments, the flow rate is about 3 volumes per hour. The flow rate of the bioreactor impacts the selection pressure of the system. For example, a higher flow rate selects for a biomolecule that leads to stronger activation of the synthetic circuit. In embodiments in which theAty. Dkt. No. 125141.04976 MGH2024-443desired activity of the biomolecule is the binding of an antigen, increasing flow rate in a bioreactor would increase the selection pressure for biomolecules that bind the antigen with high affinity.
[0142] The term “culture” as used herein refers to a population of cells grown under controlled conditions in vitro. The cells described herein may be cultured in a liquid, nutrient rich broth or cell culture medium. The term “lagoon” may be used herein to refer to a culture in a fixed volume vessel or chamber. The bioreactor described herein may comprise a lagoon, where the bioreactor is configured to constantly cycle fresh media and host cells into, and used media, old host cells and phagemid encoding biomolecules that fail to activate the synthetic circuit above the level required for steady phagemid propagation at the given bioreactor flow rate set by the researcher, out of the lagoon.
[0143] In other embodiments, the method is performed manually in a cartridge or deep well plate, e g. a 96-well plate.
[0144] The construct may comprise an inducible promoter, such as the PBAD promoter, as described herein. In such embodiments, the method may further comprise contacting the host cell with an inducing agent, such as arabinose. “Contacting” a cell refers to adding the inducing agent to the culture.
[0145] In an eighth aspect, provided herein is a method for generating and identifying a mutant of a biomolecule endogenously expressed in a host cell or a mutant in the genetic circuitry (promotors, ribosomal binding sites, operator sites, terminators) that regulates expression of a biomolecule endogenously expressed in a host cell, the method comprising: a) transforming the host cell with any of the constructs or plasmids described herein, wherein the construct is operably linked to an inducible promoter, wherein the target nucleic acid encodes the endogenously expressed biomolecule or a component of the genetic circuitry that regulates expression of the exogenously expressed biomolecule; b) contacting the host cell with an inducing agent that activates the inducible promoter; c) allowance of a period of incubation of between 0-24 hours under conditions that permit the function of any of the plasmids or constructs described herein; d) a method for screening or selecting members of the host population containing mutations in the endogenous biomolecule or its genetic regulatory circuitry that confers an improvement of interest (e.g. improved strain fitness under certain conditions, a gain of function mutation or improved activity of the biomolecule of interest, increased production of a desirable metabolite, or any other biomolecular activity of interest to the researcher); e) extracting and sequencing gDNA from theAty. Dkt. No. 125141.04976 MGH2024-443host cell if it passes the screening or selection step outlined in d); and f) repeating the above steps a-e if applicable. The host cell may be an E.coli cell comprising any of the mutations described herein. The construct may comprise an inducible promoter, such as the PBAD promoter, as described herein, and the method may further comprise contacting the host cell with an inducing agent, such as arabinose.
[0146] In a ninth aspect, provided herein is a method for generating and screening a library of mutants of a biomolecule using any of the reticulase plasmids described herein. The method comprising a) using any of the reticulase plasmids described herein to generate a library of mutants of a biomolecule encoded on a plasmid, virus, phage or any vector known in the art and b) performing a screening / directed evolution technique. The screening / directed evolution technique may be selected from phagemid display, phage display, ribosome display, yeast display, mammalian display, or any other screening / directed evolution technique known in the art. The method may further comprise screening the library of mutants for those having an activity of interest. An activity of interest may be an increased or decreased level of the typical activity of the biomolecule. An activity of interest may be a different activity than the typical activity of the biomolecule. The activity of interest may be improved or reduced stability solubility, activity, substrate specificity, etc.
[0147] Systems for Directed Evolution of Proteins in Periplasmic Space
[0148] In a tenth aspect, provided herein is a system for directed evolution of a protein of interest in a bacterial periplasm, the system comprising: (a) a phagemid vector comprising a nucleic acid encoding the protein of interest; and gill; (b) an accessory helper plasmid comprising: (i) a helper gene construct comprising all genes required for phage assembly except gill and gVI; (ii) a target antigen construct comprising a gene encoding a fusion protein, the fusion protein comprising a transmembrane protein and a target antigen; wherein the target antigen binds the protein of interest; and wherein when the fusion protein is expressed, the target antigen is anchored to the membrane and exposed to the periplasm; (iii) a conditional propagation construct comprising a conditional promoter operably linked to gVI; wherein the conditional promoter comprises an operator to which the transmembrane protein is capable of binding; and wherein when the transmembrane protein binds to the operator, gVI is expressed; and (c) an engineered bacterial cell comprising: disruption or deletion of a gene encoding the transmembrane protein; and a CadBA operon that is naturally operably linked to the conditional promoter; wherein the conditional promoter is replaced with aAty. Dkt. No. 125141.04976 MGH2024-443heterologous constitutive promoter; wherein interaction of the protein of interest with the target antigen in the periplasm activates the conditional propagation construct and enables propagation of the phagemid.
[0149] In alternative embodiments, the target antigen construct may comprise two genes: 1) a gene encoding a fusion protein, the fusion protein comprising a transmembrane protein and a universal donor protein that is anchored in the membrane and exposed to the periplasm when expressed and 2) a gene encoding a second fusion protein, the second fusion protein comprising a target antigen and a universal acceptor protein which are exported into the periplasm when expressed; wherein when expressed the donor protein and acceptor protein domains associate and non-covalently anchor the target antigen to the membrane; wherein the target antigen binds the protein of interest.
[0150] A “fusion protein” is an engineered hybrid protein created by joining two or more distinct genes or protein domains, enabling them to be expressed as a single polypeptide.
[0151] The term “engineered” refers to any artificial manipulation of a cell, a polynucleotide, a polypeptide, etc. that results in a detectable change in the naturally occurring (or ‘native’) organism or molecule, such as changes in expression of genes in a cell or changes in a sequence of a polynucleotide or polypeptide. The engineered bacterial cell may comprise a deletion or disruption of a gene encoding the native transmembrane protein. The gene encoding the native transmembrane protein may comprise a mutation that disrupts expression of the protein. In exemplary embodiments, the gene encoding the transmembrane protein is deleted and replaced with an antibiotic resistance cassette.
[0152] A “transmembrane protein” is a protein that spans the entirety of a cell membrane.
[0153] As used herein, the term “target antigen” refers to a peptide, protein, or protein fragment expressed or positioned in the periplasmic space that can bind the protein of interest in a system.
[0154] As used herein, the terms “universal donor protein” and “universal acceptor protein” refer to protein domains that specifically associate with one another and may be fused to other proteins to non-covalently associate those other proteins or bring them within proximity to one another. In exemplary embodiments, the universal donor protein may be HA4, and the universal acceptor protein may be SH2. The HA4 sequence (101 amino acid protein that associates 1:1 with SH2) is GGCAGCTCTGTGAGTAGCGTTCCGACCAAACTGGAAGTGGTTGCAGCAACCCCGAC GAGCCTGCTGATTTCTTGGGATGCCCCGATGTCTAGTAGCTCTGTGTATTACTATCGT ATCACCTACGGTGAAACGGGCGGTAACAGCCCGGTGCAGGAATTTACGGTTCCGTAAty. Dkt. No. 125141.04976 MGH2024-443TAGTAGCTCTACCGCGACGATTAGTGGCCTGAGCCCGGGTGTGGATTACACCATCAC GGTTTATGCATGGGGCGAAGATAGCGCGGGTTACATGTTCATGTATTCTCCGATTAG TATCAATTATCGTACCTGCTGA. The SH2 sequence (100 amino acid protein that associates 1:1 with HA4) is AGTCTGGAAAAACACAGCTGGTATCATGGCCCTGTGAGCCGTAACGCGGCCGAATA CCTGCTGAGCTCTGGCATTAATGGTTCTTTTCTGGTTCGTGAAAGTGAAAGTAGCCC GGGCCAGCGCAGCATTTCTCTGCGTTATGAAGGTCGCGTGTATCACTACCGTATCAA CACCGCCAGCGATGGCAAACTGTACGTTTCTAGTGAATCTCGCTTCAATACCCTGGC AGAACTGGTGCATCACCATAGCACGGTTGCGGATGGTCTGATCACCACGCTGCATTA TCCGGCGCCGAAACGCTAATAG.
[0155] “Periplasm” or “periplasmic space” is a cell compartment located between the inner cytoplasmic membrane and the outer membrane of Gram-negative bacteria.
[0156] The term “conditional promoter” as used herein refers to a promoter or promoter region that is activated when a protein of interest in the system is able to bind the target antigen, which can then impact the conformation of the transmembrane protein such that the transmembrane protein can bind to the conditional promoter or an operator within the conditional promoter. In exemplary embodiments, when the protein of interest binds the target antigen, the fusion protein binds to the operator to induce transcription from the conditional promoter. In embodiments, the fusion protein dimerizes and binds to the operator to induce transcription from the conditional promoter. In exemplary embodiments, the transmembrane protein is CadC, and the conditional promoter is a CadBA promoter. The CadC protein may be a truncated protein comprising the cytosolic DNA-binding region of the protein (CadCi-155).
[0157] In embodiments, the accessory helper plasmid comprises SEQ ID NO: 113 or 127, or a sequence having at least 80% identity thereto, wherein all of the Innovations are preserved. The helper plasmid may have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, at least 90%, or at least 95% identity to SEQ ID NO: 113 or 127.
[0158] In embodiments, the phagemid vector comprises SEQ ID NO: 121 or 130, or a sequence having at least 80% identity thereto, wherein all of the Innovations are preserved. The phagemid vector may have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, at least 90%, or at least 95% identity to SEQ ID NO: 113 or 127.
[0159] A “heterologous promoter” is a promoter sequence that is not naturally linked to a gene orAty. Dkt. No. 125141.04976 MGH2024-443nucleic acid sequence.
[0160] The system may comprise any of the mutations (“Innovations”) described in Example 3.
[0161] In an eleventh aspect, provided herein is an engineered bacterial cell for periplasmic evolution of proteins encoded on a phagemid, the cell comprising: deletion or inactivation of cadC; and a heterologous constitutive promoter operably linked to cadBA. The cadC may be deleted and replaced with another sequence. In exemplary embodiments, the cadC gene is replaced with a gentamycin resistance cassette. The heterologous constitutive promoter may be a PproAor a Pj23iso promoter.
[0162] In exemplary embodiments, the engineered bacterial cell is ax E.coli cell. The E. coli cell may be an s2060 or s!030 strain. The E. coli cell may be any of the cells described herein.
[0163] Miscellaneous
[0164] Unless otherwise specified or indicated by context, the terms “a”, “an”, and “the” mean “one or more.”
[0165] As used herein, “about,” “approximately,” “substantially,” and “significantly” will be understood by persons of ordinary skill in the art and will vary to some extent on the context in which they are used. If there are uses of these terms which are not clear to persons of ordinary skill in the art given the context in which they are used, “about” and “approximately” will mean plus or minus <10% of the particular term and “substantially” and “significantly” will mean plus or minus >10% of the particular term.
[0166] As used herein, the terms “include” and “including” have the same meaning as the terms “comprise” and “comprising” in that these latter terms are “open” transitional terms that do not limit claims only to the recited elements succeeding these transitional terms. The term “consisting of,” while encompassed by the term “comprising,” should be interpreted as a “closed” transitional term that limits claims only to the recited elements succeeding this transitional term. The term “consisting essentially of,” while encompassed by the term “comprising,” should be interpreted as a “partially closed” transitional term which permits additional elements succeeding this transitional term, but only if those additional elements do not materially affect the basic and novel characteristics of the claim. Embodiments recited as “including,” “comprising,” or “having” certain elements are also contemplated as “consisting essentially of’ and “consisting of’ those certain elements.
[0167] Recitation of ranges of values herein are merely intended to serve as a shorthand methodAty. Dkt. No. 125141.04976 MGH2024-443of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. For example, if a concentration range is stated as 1% to 50%, it is intended that values such as 2% to 40%, 10% to 30%, or 1% to 3%, etc., are expressly enumerated in this specification. These are only examples of what is specifically intended, and all possible combinations of numerical values between and including the lowest value and the highest value enumerated are to be considered to be expressly stated in this disclosure. Use of the word “about” to describe a particular recited amount or range of amounts is meant to indicate that values very near to the recited amount are included in that amount, such as values that could or naturally would be accounted for due to manufacturing tolerances, instrument and human error in forming measurements, and the like. All percentages referring to amounts are by weight unless indicated otherwise.
[0168] In those instances where a convention analogous to “at least one of A, B and C, etc.” is used, in general such a construction is intended in the sense of one having ordinary skill in the art would understand the convention (e.g., “a system having at least one of A, B and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description or figures, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”
[0169] “Substantial identity” of amino acid sequences means that a polynucleotide or polypeptide comprises a sequence that has at least 85% sequence identity to a reference sequence (SEQ ID NO) using a sequence alignment program; preferably BLAST using standard parameters. A preferred percent identity of polynucleotides and polypeptides can be any integer from 85% to 100%. A preferred percent identity may be 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to a reference sequence.
[0170] No admission is made that any reference, including any non-patent or patent document cited in this specification, constitutes prior art. In particular, it will be understood that, unless otherwise stated, reference to any document herein does not constitute an admission that any of these documents forms part of the common general knowledge in the art in the United States or inAty. Dkt. No. 125141.04976 MGH2024-443any other country. Any discussion of the references states what their authors assert, and the applicant reserves the right to challenge the accuracy and pertinence of any of the documents cited herein. All references cited herein are fully incorporated by reference, unless explicitly indicated otherwise. The present disclosure shall control in the event there are any disparities between any definitions and / or description found in the cited references.
[0171] All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[0172] Preferred aspects of this invention are described herein, including the best mode known to the inventors for carrying out the invention. Variations of those preferred aspects may become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors expect a person having ordinary skill in the art to employ such variations as appropriate, and the inventors intend for the invention to be practiced otherwise than as specifically described herein. Accordingly, this invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the invention unless otherwise indicated herein or otherwise clearly contradicted by context.
[0173] EXAMPLES
[0174] The following Examples are illustrative and should not be interpreted to limit the scope of the claimed subject matter.
[0175] Example 1. A System for the Continuous Rationally Directed Evolution of Biomolecules
[0176] Introduction
[0177] With the advent of protein language models and artificial intelligence, our ability to predict potential fitness-improving protein variants has greatly outpaced our ability to screen protein variants.1,2At present, existing high-throughput protein screening methods have centered on the use of ‘display’ technologies, such as phage, yeast, and mammalian display.3These approaches typically involve the generation of large vector libraries via ex vivo molecular cloning, in whichAty. Dkt. No. 125141.04976 MGH2024-443select codons within a gene of interest encoded on a vector are diversified by PCR with degenerate oligonucleotide primers (for example NNN or NNK).4Such libraries are then transformed into recipient cells and the encoded protein variants are screened for desirable properties, such as enhanced antigen binding or specificity in the case of antibodies or soluble T cell receptors (TCRs). While each method offers distinct advantages and drawbacks, all share the fundamental limitation of being highly labor-intensive and time-consuming when performed iteratively.
[0178] In each display method, the maximum number of variants that can be effectively screened is rapidly exceeded when even a moderate number of positions in a gene of interest are selected for diversification. Phage display can screen a maximum library size of approximately 10l ,-1012, while yeast display and mammalian display can screen 109- 1010, and 1 CP-106, respectively.3Thus, for libraries involving full codon saturation (substitution of all 19 amino acids at a given position), phage display is limited to the saturation screening of nine positions (209= 5.12*10n), yeast display seven positions (207= 1.28*109), and mammalian display just four positions (204= 1.60* 105). These limitations have necessitated the use of strategies broadly referred to as ‘Iterative Saturation Mutagenesis’ (ISM) for the screening of libraries that exceed the capacity of a chosen method4’6_ In ISM, a subset of selected positions are saturated and screened, and a selected variant from the first round is used as the template vector for a subsequent round of screening at a different subset of positions. In this iterative approach the total number of variants screened is additive rather than multiplicative, drastically decreasing the total number of potential variants that can be screened. For instance, a selected variant from a first-round nine-positional phage display saturation library can be used as the template vector for a subsequent round of saturation screening at a different set of nine positions. A manageable S.I ^IO11+ 5.12*10n= 1.02* 1012variants would be screened in total, as opposed to the unmanageable 2018= 2.62* 1023which would be screened in a single-round library. Even in cases where the desired combinatorial diversity could theoretically be achieved in a single round library, ISM is still often utilized because it enables the construction of significantly smaller, more manageable libraries and because it maximizes the chance of accumulating multiple beneficial mutations on the same biomolecule through successive rounds of screening and validation.6
[0179] While ISM may overcome the issue of library size, it also introduces many significant drawbacks. Each single round of library construction, screening, and clonal isolation generally occurs over the course of one to two months for display methods, although the specifics of theAty. Dkt. No. 125141.04976 MGH2024-443biomolecule screened, library size, and method used can either reduce or extend this timeframe. The need to perform library screening across multiple, iterative rounds can add years to development timelines for certain therapeutically-valuable biologies - such as soluble TCRs -which require extensive affinity and stabilization maturation that cannot be achieved in a singleround library screen.7Moreover, the order in which positions in a protein are chosen for diversification in an ISM campaign (the ISM scheme) can profoundly impact the fitness of the final product. Epistatic interactions between positional variants selected across different rounds of ISM can provide or prohibit access to local fitness maxima.8For instance, a comprehensive study seeking to alter the enantioselectivity of an epoxide hydrolase by saturating five positions in the protein demonstrated that less than half of the 120 ISM schemes yielded the fittest protein product.9The possibility of selecting a suboptimal ISM scheme can necessitate the pursuit of distinct ISM schemes in parallel, naturally increasing labor and decreasing throughput. Strategies that shortcut ISM by simply combining fitness-improving variants identified in different libraries on a single molecule (a strategy known as ‘additivity’) have long been known to be unreliable due to epistatic clashes between positional variants identified in isolation.10
[0180] In contrast to display technologies, continuous directed evolution technologies aim to streamline all parts of a directed evolution campaign in vivo with minimal researcher intervention. Such platforms include Viral Evolution of Genetically Actuating Sequences (VEGAS) in human cell culture, Orthorep in yeast and Phage- Assisted Continuous Evolution (PACE) in E. coli.n’13In each of these systems, a high rate of random mutagenesis in the gene of interest is achieved in vivo, and a method for the continuous selection of library variants is applied to amplify the most fit library variants from the population while either penalizing or removing unfit variants.14Because the full process of mutagenesis and selection occurs in vivo, mutations iteratively accumulate in the most fit library variants, further improving their fitness. While these platforms have been used to evolve a broad range of biomolecules with significantly shorter turnaround times and less labor input than their respective ‘discontinuous’ display analogs, their reliance on indiscriminate random mutagenesis precludes their ability to generate and screen protein variants divergent at select positions. Moreover, although these technologies have achieved high error rates and broad mutagenic spectra, none are capable of full codon saturation in vivo. Each of these systems is generally able to introduce at most two substitutions in a codon in a single generation of genome replication and are highly biased towards transitions.11,14’15For reference, of the 380 possibleAty. Dkt. No. 125141.04976 MGH2024-443amino acid interconversions, 150 (39%) are possible with single nucleotide substitutions, 166 (44%) require double substitutions, and 64 (17%) require triple substitutions; more than a third (34.5%) of transition mutations do not change the encoded amino acid, as opposed to 14.2% of transversions.16’17Codon degeneracy also creates an uneven topology of fitness valleys that makes ‘random’ mutagenesis follow a predictable hierarchy that is heavily skewed towards certain interconversions over others.3,18,19For instance, R (CGG vs AGG vs AGA) to W (TGG) requires either a C>T transition, an A>T transversion, or an A>T / A>G dual transversion / transition. Thus, the probability that an R>W interconversion will occur in a random mutagenesis scheme is therefore heavily dependent on the R codon used. Unsurprisingly, the inability of current in vivo mutagenesis methods to generate all amino acid interconversions from all starting codons can gate local fitness maxima behind insurmountable evolutionary fitness valleys. Such gating has periodically necessitated a need for ex vivo cloning and undermined the inherent benefit of continuous evolution.18Random mutagenesis in protein regions not determinative of a desired phenotype can also enable false positives via ‘cheating’ of selection conditions, as recently reported for an scFv antibody that homodimerized within a linker region far from the antigen binding site to short-circuit a PACE campaign.19In light of these drawbacks, it is not surprising that a recent review described a method for the continuous, targeted mutagenesis of a gene of interest in vivo as the unrealized ‘Holy Grail’ of directed evolution.20
[0181] Thus, we argue that a method for the continuous, full-saturation mutagenesis of select codons within a gene of interest in vivo would be a major development for modern directed evolution. When combined with a method for the continuous selection of functional library variants, such a method would effectively automate ISM - enabling beneficial mutations at user-selected positions in a gene of interest to accumulate continuously without researcher intervention. If all designated positions could be saturated and selected for in a multiplexed manner over the course of a campaign, such a method would effectively explore all possible ISM schemes that could be taken from a set of chosen positions.
[0182] With this in mind, we describe the invention of Targeted Reticulase-Assisted Continuous Evolution (TRACE). This system is the first continuous directed evolution platform enabling full saturation mutagenesis of user-selected codons in a gene of interest with single base-pair precision. We combine components of a Diversity-Generating Retroelement (DGR) and single-strand DNA annealing proteins (ssap) to create a synthetic prokaryotic retrotransposon we term a ‘reticulase.’Aty. Dkt. No. 125141.04976 MGH2024-443The reticulase utilizes an error-prone reverse transcriptase with a >40% error rate and broad, transversion-prone mutagenic spectra exclusively on template adenines (A), enabling the synthesis of cDNA libraries with NNN at any position in the template RNA with a corresponding AAA. Short cDNAs (27-80bp) directed by GCT homology are then integrated into a gene of interest by ssap, pasting over the endogenous bases regardless of the codon present to enable full saturation of any user-specified codon with single base pair precision. We achieve editing efficiencies of 3-8% across two genomic loci in E. coli and report the first instance of in vivo rational evolution in a bacterial genome, discovering three novel mutations that provide resistance to the antibiotic rifamycin at the codon 531 in rpoB. Next, we engineer the life cycle of bacteriophage M13 to accommodate high-efficiency (>10%) reticulase editing on Ml 3 phagemid through a novel gene editing pathway we refer to as ‘reticulate fibonasion.’ Using this method, we generate libraries of more than 1 billion reticulase-edited Ml 3 phagemid from a single ImL culture of E. coli in an 18-hour period. To underscore the utility of our method to rapidly generate libraries of clinically translatable targets, we use a reticulase to introduce all 19 amino acid substitutions at a chosen position in the CDR3 loop of an FDA-approved antibody implicated in the treatment of type II metastatic breast cancer. Finally, we combine our phagemid gene-saturation method with conditional M13 phage propagation circuits to enable the continuous selection of library members with desirable fitness-improving mutations from a diverse population of more than 108phage. Library members with desirable fitness-improving mutations in a transcription factor encoding gene are selected by a synthetic circuit and dominate the population after just five days, enabling straightforward isolation of evolved clones via phage plaquing without the need for tedious biopanning or phage ELISA. Thus, an entire experimental workflow of rational library generation, variant screening, and clonal isolation is achieved in less than a week: more than ten times faster and significantly less labor-intensive than traditional display methods. By combining rationally targeted saturation mutagenesis with continuous selection, TRACE represents the first system for the continuous rationally directed evolution of biomolecules.
[0183] Results
[0184] Functional transplantation of diversity generating retroelement components to E. coli s2060. Diversity generating retroelements (DGRs) are a unique family of prokaryotic retrotransposons with evolutionary relationships to retrons and group II introns that generate continuous hyper-directed sequence variation of the protein-encoding genes they target.23,21DGRsAty. Dkt. No. 125141.04976 MGH2024-443utilize a unique mechanism of site-specific retrotransposition in which sequence variants are inserted into a flexible coding scaffold avoiding non-specific variation in conserved genomic regions.22The necessary components of a DGR are typically localized in a single genomic locus that spans 5-10kb, however, the synteny and organization of these components can vary.23Mechanistically, diversification is facilitated by a reverse transcriptase that acts on a non-coding RNA transcribed from a template repeat (TR) region in the DGR locus. The template repeat region is almost identical to a variable region (VR) often found in a nearby gene, which encodes the DGR-variable protein. The reverse transcription of the RNA intermediate into cDNA is carried out by the error-prone RT which favors A to N mutations with a mutation rate exceeding 50%.24The mutated cDNA then replaces the VR, which often encodes a sequence that corresponds to the flexible residues in a ligand-binding structural domain belonging to the C-type lectin protein family.25,26The TR is invariable, and transposition occurs in a unidirectional manner from TR to VR, generating continuous sequence diversity at positions in the VR that contain adenine at the corresponding positions in the TR. The variation created by DGRs reaches at least IO20possible sequences in a single gene - with metagenomic analyses identifying variable proteins with IO30possible sequence variants - a scale of diversity outstripping that generated by vertebrate somatic hypermutation by more than 15 orders of magnitude25,27DGRs have been identified in bacteria, archaea and their viruses from across diverse ecosystems, and are predicted to account for more than 10% of total genomic diversity in some organisms.28
[0185] To our knowledge no DGR or DGR components have ever been used for any biotechnological, industrial, or commercial purpose. However, we reasoned that components of a DGR could form the lynchpin of a next-generation continuous directed evolution platform. As the subject of our study, we selected a DGR from Bordetella Positive Phase bacteriophage 1 (BPP-1), the first DGR discovered and the one that has been the most extensively characterized. In nature, BPP-1 infects Bordetella, the bacterial pathogen that causes whooping cough in humans7,29The BPP-1 DGR diversifies 12 amino acids in the tail fiber tip protein (Mtd) that bind the Bordetella host receptor, dynamically altering the tropism of this phage.22Surface plasticity m ' Bordetella cell surface proteins during its infectious cycle creates a natural selection pressure for such DGR-mediated continuous tropism switching in the phage tail fiber. This DGR composes two proteins: the BPP-1 DGR reverse transcriptase (bRT), and an Accessory Variability Determinant protein (Avd) that associates with the bRT through more than 800A buried surface area to stabilize theAty. Dkt. No. 125141.04976 MGH2024-443DGR-RNA in its active site.30The 580bp DGR-RNA which forms the template for reverse transcription is expressed polycistronic with these proteins: its extensive 296bp 5’ UTR is the Avd coding region, and its 150bp 3’UTR comprises a 140bp spacer region (sp) upstream of the bRT start codon and the first 6bp of the bRT ORF.31
[0186] We hypothesized that components of the BPP-1 DGR could be functionally transplanted to E. coli and reengineered to facilitate programmable hypermutation of select gene regions in bacterial genomes or M13 phagemid. To our knowledge, no components of a diversity generating retroelement have been functionally transplanted to an exogenous organism or placed under the control of a tunable external stimulus. Thus, prior to construction of a gene-editing complex that incorporated DGR components, we intended to verify that the BPP-1 DGR bRT, Avd, and DGR-RNA could produce adenine-diversified cDNA libraries in the cytoplasm of E. coli. Because the aim of our work was to create a phage-based continuous directed evolution platform, we performed all experiments in E. coli s2060 - the host strain for phage-assisted continuous evolution (PACE) - or strains derived from s2060 through further genomic engineering.32,33
[0187] The arabinose-inducible promoter Pbad was used to express the BPP-1 DGR proteins Avd and bRT off an Intron Plasmid (IP) (FIG. 1A). To achieve an approximate 5:1 stoichiometry of these proteins in the cell - which have been demonstrated to associate as an Avd pentamer - bRT monomer - we used the ribosomal binding sites SD8 for Avd and sd8 for bRT.30These RBS’ have been reported to exhibit 1 to 0.2 relative arbitrary translational units, respectively.34Previous in vitro data has demonstrated that only the terminal 20 base pairs of the DGR-RNA 5’UTR (Avd 368-387) are required to enable cDNA production off the TR template, while a longer 157bp stretch of DGR-RNA 3’UTR (IMH* 5-21 + spl-140) is required.31We expressed this ‘core DGR-RNA’ with the minimal functional 5’ and 3’ UTRs flanking the native TR sequence from the IP using the strong constitutive promoter ProD.33A 429bp self-splicing intron from E. coli bacteriophage T4 flanked by Sall sites was inserted between nucleotides TR84 and TR85 in the DGR-RNA.36This location in the TR was chosen because it had previously been shown to tolerate an intron insertion in a retrohoming assay in Bordetella31Following transcription, the td intron is expected to self-splice, generating an RNA with a scar of two 18 bp exons between TR84-85. This RNA can then serve as the template for bRT-Avd mediated reverse transcription. NGS primers landing in the TR region upstream and downstream of the splice sites are expected to generate either a 750bp band (if amplifying off the IP template) or a 300bp band (if amplifying off cDNAAty. Dkt. No. 125141.04976 MGH2024-443generated through reverse transcription of a post-spliced RNA). Thus, this assay provides both a) an efficient means for the verification of bRT-Avd activity in E. coll (through band visualization on an agarose gel) and b) an efficient means for selective extraction / sequencing of cDNA generated in E. coli without contamination from amplicons generated by polymerization off the template IP (through agarose gel size exclusion).
[0188] s2060 cells transformed with the IP were grown to mid-log phase (OD=0.3) and induced with arabinose. Upon cell lysis, DNA extraction and amplification, a distinct 300bp band could be visualized at 2hr, 5hr, and 24hr post induction with arabinose (FIG. IB). NGS of the 300bp band revealed a bell curve distribution of substitutions with a median of 12 substitutions per read (FIG.2A). Importantly, these substitutions were made almost exclusively at positions corresponding to adenines in the template DGR-RNA (FIG. 2B). We observed an average error rate of 40.8% on template adenines, with a broad mutagenic spectrum corresponding to 16.0% A>T, 15.0% A>G, and 9.7% A>C. The error rate ranged widely but within the same order of magnitude, from a low of 18.3% on TR35 to a high of 69.0% on TR62 (FIG. 2C). Error rates observed for the other three template bases were markedly low: 1.3% on template cytosine, 0.3% on template guanine, and 1.4% on template uracil. We also mapped the frequency of alleles reverse transcribed across each AAC motif present in the TR and observed that bRT can diversify AAC into 15 of the 20 canonical amino acids and not a stop codon (FIG. 3). We measured a broad spectrum of the 15 expected amino acids across each AAC in the TR, as well as a non-insignificant level (1.8% of reads) of a 16th amino acid, Lysine (AAG / AAA) at TR39-41 resulting from mild mutagenesis of this AAC codon’s cytosine. To our knowledge these data represent the first report of the allelic frequency of a DGR-RT produced cDNA library.
[0189] The mutagenic rate on template adenines we observed was slightly lower than has been reported in vitro (40.8% reported here as opposed to 51%).38To rule out the possibility of amplification off of the IP and template switching during NGS library preparation - which could result in a lower measured mutagenic rate - we performed pre-digestion of the IP backbone and library preparation with different polymerases (FIGS. 4A and 4B). These preparatory steps had negligible impact on the observed error rate or mutagenic spectra, and template switching was ruled out. It is possible that the presence of the exogenous 36bp sequence of spliced td exons between TR84-85 had a mild depressing impact on the overall mutagenic rate. We observed among the lowest rates of adenine mutagenesis on TR83 and TR84 (21.8% and 24.8% respectively)Aty. Dkt. No. 125141.04976 MGH2024-443immediately upstream of this exon region, while the mutagenic rate at these positions measured in vitro with no exons present has been reported as closer to the average rate.38The precise mechanism that determines mutagenic rates in DGR RTs is currently unknown, and it is conceivable that the presence of certain exogenous sequences in the template could have a minor depressing or enhancing impact on mutagenesis. Despite the slightly lower observed mutagenic rate observed in our assay, the relative mutagenic profile observed on each TR adenine, both in terms of total error rate and the spectra of substituted bases was almost identical to what has been reported on this template in vitro. These observations are consistent with and further support the present consensus that adenine- specific hypermutation in DGR variable cassettes is attributable to the intrinsic biophysical properties of the DGR-RT and is not dependent upon host DNA-repair mechanisms such as cytosine deamination or uracil excision.
[0190] Collectively, these data demonstrate that components of the BPP-1 DGR can be functionally transplanted to E. coli and placed under the control of a tunable external stimulus. We observed no fitness penalty associated with expression of bRT-Avd in E. coli (the proteins appear to be entirely non-toxic). To date, measurement of the bRT error rate and mutagenic spectra has relied on protein purification and in vitro characterization. We anticipate that this intron reporter assay may be of use for more high-throughput experiments for the characterization of DGR RTs, such as through the generation and characterization of large libraries of IP with different bRT point mutants or DGR-RNA templates with varying intergenic motifs.
[0191] A ‘reticulase’ is a novel class of synthetic retrotransposon combining components of the BPP-1 DGR with single-stranded DNA annealing proteins. In nature, DGRs integrate cDNA into their VR through retrohoming to an intergenic DNA cruciform structure immediately downstream of the variable protein ORF.37A similar mechanism of genomic integration is shared by group II intron maturases, the closest evolutionary relative of DGRs.39The reliance of DGRs on such non-coding intergenic motifs constrains all DGR-mediated hypermutation in the global microbiome to the extreme C-terminus of target genes and prevents multiplexed hypermutation of different regions within the same primary sequence. Indeed, a recent metagenomic analysis of DGRs from across bacterial, archaeal, phage, and archaeal virus clades revealed DGR-variable proteins are universally predicted to adopt the C-type lectin fold, in which the DGR-targeted, solvent-exposed amino acids that determine antigen specificity are arranged in a single loop at the C-terminus of the protein.28This natural integration mechanism of DGRs fatally undermines theirAty. Dkt. No. 125141.04976 MGH2024-443potential utility as tools for the directed evolution of desirable biomolecules. Notably, proteins of the immunoglobulin fold (encompassing antibodies, T-cell receptors, and nanobodies) engage antigen with amino acids arranged in three to six CDR loops centrally located in the protein primary sequence - a set of genomic targets that would be effectively impossible to diversify with a natural DGR.
[0192] We hypothesized that the native mechanism of DGR retrohoming could be uncoupled from cDNA synthesis and replaced with an entirely novel mechanism of genomic integration unconstrained by reliance on intergenic homing motifs. To our knowledge DGRs are the only class of retrotransposon in nature in which cDNA synthesis is cis primed off a 2’OH group in the 3’UTR of an RNA template. Many prominent classes of retrotransposons (non-LTR retrotransposons, LINEs, polyA retrotransposons) trans prime cDNA synthesis off a 3 ’OH group generated by a nick in the target locus.40’43This unique ability of DGRs to cis prime reverse transcription off an RNA template allows cDNA synthesis to be mechanistically decoupled from genomic integration (demonstrated in the IP assays above) in a manner that is not possible for trans primed retrotransposons. The other prominent class of retroelements which cis prime cDNA off a 2’OH in a template RNA are retrons, a class of bacterial retroelements which play a role in adaptive antiphage defense.44,43Although retron cDNA is not naturally integrated into a host genome, retrons have recently been used together with single-strand DNA annealing proteins (ssap) to create ‘recombitrons:’ synthetic retrotransposons that integrate retron-generated cDNA in place of an Okazaki fragment on the lagging strand of a target locus replication fork.46’49
[0193] We hypothesized that a similar approach could be taken to integrate bRT-Avd-generated cDNA into a target locus without reliance on any DGR intergenic homing motifs (FIG. 5). To test this hypothesis, a new plasmid was constructed in which an arabinose-inducible Pbad promoter was used to control the expression of bRT and Avd as well as a ssap and a dominant-negative mismatch repair protein MutL E32K (MutL*).50The constitutive ProD promoter was used to drive the expression of an RNA consisting of the minimal functional 5’ and 3’UTRs of the core DGR-RNA flanking a ‘Target Homology Region’ (THR) with homology to a selected target locus. Upon arabinose induction, Avd-bRT cis primes cDNA off the RNA 3’ UTR and generates a cDNA library by error-prone reverse-transcription of the THR template. All template adenines in the THR are expected to be selectively hypermutated in the cDNA library. Adenine substitutions can be made in the THR at positions corresponding to nucleotides in the target locus chosen forAty. Dkt. No. 125141.04976 MGH2024-443hypermutation. For example, the lysine codon AAA can be used to generate a library of cDNA with NNN at the chosen position, regardless of the native adenine content of the target codon. Alternatively, discrete edits can be made by substitution of cytosine, guanine and uracil at chosen positions in the THR due to the relative high fidelity of Avd-bRT on these bases. cDNA generated by Avd-bRT is bound by the ssap and annealed in place of an Okazaki fragment on the target locus during genome replication. Sequence conservation between the target locus and cDNA at all cytosines, guanines and uracils enables scarless homology-directed annealing of the cDNA into the lagging strand. The dominant-negative mismatch-repair protein MutL* prevents reversion of the mismatched bases on the edited strand until a subsequent round of genome replication results in a stable double-stranded edit. This novel method for the integration of bRT-Avd-generated cDNA is entirely non-reliant on intergenic homing motifs, and therefore could be used to hypermutate any nucleotide, codon, or gene region irrespective of its position within or outside an ORF. This pathway is also RecA-independent, as the s2060 cell line is ARecA.32
[0194] The term ‘reticulate evolution’ has been used to describe processes of symbiosis, symbiogenesis, and horizontal gene transfer in which populations evolve through the continuous interchange of genetic elements.51‘Reticulate’ comes from the latin reticulatus, (Tittle net’) referring to the bifurcating pattern of crossings and mergings which epitomize networks of ancestral descent in populations undergoing continuous horizontal gene transfer. Because our RNA-protein complex diversifies target genomes through the introduction of horizontally-transferred cDNA, we named our synthetic retrotransposon a ‘reticulase.’ As noted, a reticulase is not a DGR, but a novel class of synthetic retrotransposon with far greater targeting versatility due to its flexible mechanism of genomic integration.
[0195] We first tested our reticulase with an 80bp THR making a rpob S531F edit (nucleotide C1591>T). We found that a Reticulase with the ssap from Collinsella stercoris phage (CspRecT) outperformed one with the ssap red beta from E. coli lambda phage, while the expression of MutL* further increased editing.52Our initial efficiency for this construct was .01% (data not shown). To increase the editing efficiency of our system, we generated a new cell line, s2064, in which the primary 5 ’-3’ E. coli exonuclease RecJ and the primary 3 ’-5’ exonuclease sbcB were knocked out from s2060. These exonucleases degrade linear single-stranded DNA in the cytoplasm, and their removal has been previously shown to improve the editing efficiency of recombineering-based genetic engineering platforms.50We further engineered this strain with a Q576A mutation in DNAAty. Dkt. No. 125141.04976 MGH2024-443gyrase (DnaG) (SEQ ID NO: 64), which further improved editing efficiency (FIG. 6A)?3This new strain was labeled s2065. DnaG is responsible for synthesizing RNA primers on the lagging strand of the replication fork, and the Q576A mutation has previously been shown to disrupt the interaction between DnaG and helicase DnaH, leading to less frequent RNA primer synthesis, longer Okazaki fragments and an increased recombineering efficiency. Next, we tested a series of arabinose induction conditions (FIGS. 6C and 6D) and found that 5 and 10 mM arabinose provided the highest editing efficiency. We also varied the copy number of the reticulase by swapping out the CloDF13 origin of replication (~40 copies cell) with the sclOl ori (-5 copies cell).54A consensus has emerged in the recombitron gene editing field that the low copy ori sclOl provides the highest editing efficiency. However, we found that the CloDF13 origin provided higher efficiency on the rpoB locus and the sclOl origin provided higher efficiency on the gyrA locus (FIG. 6B).46,50Finally, we tested a substitution of spG4C in the 3’UTR of the regRNA, which has been reported to increase the production of cDNA in vitro by as much as 30%.30This modification provided a greater variability in editing efficiency, but did not improve efficiency to a statistically significant extent (FIG. 7). Collectively, these iterative improvements allowed us to reliably achieve between 3%-8% editing across two genomic loci for an 80bp THR. In addition, to our knowledge, this work represents the first instance of the complete removal and replacement of the integration mechanism of a natural retrotransposon.
[0196] Reticulase homology shielding enables discretionary targeted mutagenesis of select gene regions. Over the course of our experiments, we observed a high degree of collateral mutagenesis on endogenous adenines present in edited target loci. For instance, rpoB K527 (codon AAA) exhibited the expected -40% error rate on each adenine when using an 80bp THR, consistent with the intrinsic mutagenic spectra of the bRT (FIG. 8B). We wanted to create a system in which collateral mutagenesis of endogenous adenines could be effectively mitigated, enabling discretionary saturation of select codons in the target locus. It has previously been reported that ssap such as red beta and CspRecT require -10 bp of perfect homology between the donor DNA and the target locus to facilitate integration.49We hypothesized that undesired collateral mutagenesis of endogenous adenines could be prevented by strategic truncation of the THR that placed endogenous adenines within this ~10bp ssap homology overlap region. Importantly, such ‘homology shielding’ is not intended to prevent the generation of cDNA library members with diversified template adenines. Rather, this method is intended to provide a highly selectiveAty. Dkt. No. 125141.04976 MGH2024-443preference for the integration of library members which have perfect homology (i.e. unmutated bases) within the ~10bp ssap homology overlap region. By placing the intended adenine substitutions more than lObp from either end of the THR, we theorized that select codons in the target locus could be hypermutated while mitigating collateral adenine mutagenesis.
[0197] We tested this hypothesis by making a progression of THR truncations from 80bp to 27bp around the rpoB S531F (C1591>T) edit (FIGS. 8A and 8B). The editing efficiency dropped as the guide length decreased, consistent with similar findings in the retron recombineering field. We also observed that the collateral mutagenesis in the target locus was reduced as the THR was truncated. Specifically, template adenines within ~9-10bp of the ends of the THR exhibited a significantly reduced mutagenic rate in comparison to those adenines centrally located in the THR, consistent with our expected model of ‘homology shielding.’ Most notably, the adenines of the K527 codon dropped from ~40% mutated in the 80bp THR (23bp from the THR 5’ end) to ~10% mutated in the 60bp THR (8bp from the THR 5’ end) to <1% mutated in all other THRs (extreme 5’ end of the THR). Collectively, these results demonstrate that homology shielding is a valid method for minimizing collateral mutagenesis of endogenous adenines in a target locus.
[0198] Reticulases enable full saturation mutagenesis and rational evolution of select codons in bacterial chromosomes with single base-pair precision. To test the ability of the reticulase to facilitate rationally directed evolution of a bacterial genome, we selected rpoB S531 for full saturation mutagenesis. Several amino acid substitutions in rpoB S531 can convey resistance to the antibiotic rifamycin, and mutations in the homologous codon are routinely identified in isolates of multidrug-resistant Mycobacterium Tuberculosis (rnTb).55-’6We used a 36bp THR to enable narrow precision hypermutation of S531 via homology shielding, and the position corresponding to S531 was replaced with AAA in the THR. Deep sequencing (500,000 reads / sample) of s2064 cells induced with this construct revealed exceedingly precise targeted hypermutation of codon 531 (FIG. 9A). Approximately 0.5% of cells possessed amino acid substitutions at this position, consistent with the lower efficiency expected of shorter THRs. Allelic mapping identified reads corresponding to 19 of the 20 canonical amino acids (all except histidine) as well as stop codons (FIG. 9B). 15 of the 20 amino acids had an incidence of >~0.5% of substituted reads, with Lysine making up approximately half of all amino acids substituted. Trace levels (in some cases single reads) of Aspartic Acid, Methionine, Tryptophan, and Cysteine were also detected at an incidence of <0.5% total substituted reads.Aty. Dkt. No. 125141.04976 MGH2024-443
[0199] To demonstrate that reticulase-introduced edits could convey a selective fitness advantage to a protein / host genome, we plated our edited cells 20 hours post-reticulase induction on +-rifamycin plates and colony counted (FIG. 10A). Approximately 0.2% of colony forming units were rifamycin resistant, or 1 / 500 cells (FIG. 10B). Sanger sequencing of 24 individual colonies identified six amino acid substitutions which conveyed resistance at rpoB 531: Lysine, Arginine, Glycine, Isoleucine, Glutamate, and Leucine (FIGS. 10B, 10C and HA). We repeated this experiment and sanger sequenced 190 colonies, which revealed the same set of mutations, minus Glycine and with Asparagine (FIG. 1 IB). To our knowledge this is the first report of the S53 IK, S531R, or S53 II mutations as conferring rifamycin resistance. One study reported that S53 IN did not confer rifamycin resistance, however in our hands it did at 25ug / mL rifamycin.57Analysis of the rpob crystal structure demonstrates that serine 531 faces directly into the rifamycin binding pocket (FIG. 11C). Thus, the finding that substitutions of larger bulky groups at this position (K, R, I, N) conferred resistance to rifamycin - presumably by displacing the molecule - had a strong structural basis. The S531Q, S531G, and S531L mutations have previously been reported as conferring rifamycin resistance.58,59The identification of S531L was especially notable because this mutation is the single most common mutation identified in rifamycin-resistant isolates of multidrug resistant mTb but is rarely identified in rifamycin- resistant isolates of E. coli. This discrepancy has been attributed to idiosyncrasies of the respective Serine codons used. In E. coli the S531L mutation requires a double substitution (either TCOTTG or TTA) while in mTb the homologous S>L mutation requires only a single transition (TCG>TTG).5?Thus, the ability of the reticulaseto execute this targeted double substitution (TCOTTA) underscores its ability to access favorable genotypes which may be gated behind large evolutionary fitness valleys. We would underscore that none of these mutations were directly programmed in the THR. Rather, a cDNA library with all possible mutations (NNN) was generated from a monoclonal plasmid with AAA at the corresponding position, integrated into the genome at the targeted locus, and those individuals with mutations that conveyed rifamycin resistance were selected from the population in response to selection pressure. Because this codon was selected on a rational basis (we knew that rifamycin resistance mutations have been previously identified at this position), we describe this experiment as the first instance of fully in vivo rationally directed protein evolution.
[0200] We argue these data represent a significant advance in the field of in vivo continuous evolution. Current methods for the full saturation of select codons in bacterial genomes rely on theAty. Dkt. No. 125141.04976 MGH2024-443electroporation of large libraries of oligonucleotides in strains prepared with a recombineering plasmid such as pORTMAGE.60While such methods can be automated through repeated rounds of electroporation and competent cell preparation - as in Multiplex Automated Gene Editing (MAGE) - such automated methods require large apparatuses which are costly to acquire and impractical to operate for most small to moderately-sized laboratories.61Moreover, the fundamentally discontinuous nature of such methods (requiring repeated rounds of electrocompetent cell preparation) precludes their use for the fully in vivo continuous targeted mutagenesis of dynamic populations, such as phage replicating under the control of a conditional propagation circuit (as in PACE). Existing fully in vivo targeted mutagenesis technologies, such as EvolvR (a Cas9-directed error-prone DNA Poll) and REGES (an error-prone retron Recombitron) can raise the local mutagenesis rate within a semi-defined window (18bp and 50bp, respectively)49,62However, these technologies do not enable single base-pair precision hypermutation of select codons, are highly biased towards transitions, and have not been demonstrated to enable full codon saturation. Thus, the ability of a monoclonal reticulase plasmid to saturate a user-selected position in a bacterial genome with single base pair precision in a 20-hour period following a single induction represents a significant advance over existing methods.
[0201] Engineering a split regRNA to enable multiplexed targeting. The ability of our proposed system to automate ISM in vivo is contingent on the ability of the RP to perform multiplexed saturation mutagenesis of distinct genomic regions. Although the construction of a single RP with multiple full regRNA regions (ProD-RNA coding region-terminator) is possible, such a construct would necessitate the assembly of multiple long iterative repeats (each 400bp total) which would be difficult to clone with conventional methods. Thus, we sought to create a method to enable the polycistronic expression of multiple cDNA libraries from a single RNA with small (<60bp) iterative repeats that would be straightforward to clone with conventional methods.
[0202] Any truncation of the minimal 5’ and 3’ UTRs of the core DGR-RNA has been reported to significantly reduce cDNA production in vitro.31While the minimal 5’UTR is already very short (20bp), we hypothesized that the longer 160bp 3’UTR could be further engineered. The 3’UTR region sp3-33 forms a hairpin loop that sits on top of the Avd pentamer. The nucleotides spl9-22 (ACGG) form the stem loop of this hairpin. We hypothesized that the 3’UTR could be divided between nucleotides sp20 and sp21 to create two RNAs: one RNA with the 5’ UTR, the variable THR sequence and the first half of the 3’UTR (what we designate as the 3’UTRa); and a secondAty. Dkt. No. 125141.04976 MGH2024-443RNA with the second half of the 3’ UTR (what we designate as the 3’ UTRb) (FIGS. 12A and 12B). These RNAs could be expressed off separate loci and associate via their hairpin homology to create the same overall secondary structure that enables cDNA production. Importantly, the RNA containing the THR would be significantly shorter than the full-length RNA. When put end to end, multiple RNAs containing the THR segments would have an iterative repeat of the 3’UTRa and the 5’UTR (37bp + 20bp = 57bp total). When staggered, 60bp annealed oligonucleotides could easily contain such a sequence, together with short 4bp overhangs to enable rapid assembly via golden gate cloning.47Thus, this split-RNA approach could enable the straightforward construction of ‘reticulase arrays’ capable of polycistronic production of multiple cDNA libraries targeting distinct genomic loci. The invariable 3’ UTRb would be expressed off a separate locus on the RP and would never need to be altered (FIG. 12B).
[0203] We tested a series of architectures expressing these RNA components off different regions of the reticulase (FIG. 13 A). The construct that performed the best in our panel was one in which the THR segment with 3’UTRa was expressed polycistronic with Avd, hR7 CspRecT, and MutL* and the 3 ’UTRb was expressed off ProD from a separate locus (Architecture 3) (FIG. 13B). This construct only slightly underperformed the standard full regRNA construct (Architecture 1), however not to a statistically significant extent. We were pleased to find that the regRNA could be cut in half with negligible detriment to editing efficiency. This finding was especially striking because single base-pair changes in DGR-RNA UTRs have been reported to nearly abolish cDNA production in vitro, and DGR-RNAs from across distinct taxa share near-identical UTR sequences, underscoring the extreme degree of evolutionary conservation and functional importance of these sequences.30Collectively, these data validate our hypothesis that the regRNA could be split between sp20-21 and still support robust reticulase editing.
[0204] ‘Reticulate Fibonasion’ is a novel gene editing method for the targeted hypermutation of phage genomes undergoing rolling circle amplification and Fibonacci expansion. Due to limitations in selection methods, the variety of biomolecules that can be evolved through E. coli chromosomal editing are strictly limited to those that can be linked to prokaryotic cell fitness, such as antibiotic resistance genes or those that confer resistance to oxidative stress. While invaluable for certain industrial applications, such limitations preclude the development of valuable biologies such as antibodies. In contrast, a diverse array of selection schemes exist for the selection of library members encoded on Ml 3 phage.20In addition to phage display, which provides an efficientAty. Dkt. No. 125141.04976 MGH2024-443selection scheme for essentially any antigen engager, conditional Ml 3 phage propagation circuits can theoretically select for any biomolecular activity (protein cleavage, DNA editing, antigen engagement, small molecule binding) so long as that activity can be linked to the expression of a gene in E. co!in
[0205] To this end, we sought to engineer a system that enabled efficient reticulase editing of genes encoded on Ml 3 phagemid. Ml 3 phagemid are bioengineered Ml 3 phage / plasmid hybrids that typically contain a phage packaging signal and origin replication, a plasmid origin of replication, an antibiotic resistance marker, a gene of interest and a phage gene (typically gill). As a starting point for our system, we elected to use a pLITMUS* Ml 3 phagemid together with an M13KO7 helper plasmid (HP), which provides the genes required for phage packaging not encoded on the phagemid. This phagemid / helper plasmid pairing (and modified derivatives) is one of the most commonly used for antibody and soluble TCR phage display.4This phagemid / helper plasmid pairing has also been used to facilitate Phagemid-Assisted Continuous Evolution (PACEmid), a continuous directed evolution method analogous to PACE.63-67
[0206] To generate hypermutated libraries of Ml 3 phagemid, host cells are transformed with the HP and the RP, grown to OD=0.3 and infected with phagemid / induced with arabinose (FIG. 14). Phage are infected at a low titer ( IO4cfu mL’1) and over the course of 18 hours, phagemid replicate, are packaged, and are released into solution to infect other host cells. Typically, the phage titer is boosted to between l*109-l*10ncfu mL'1during this period, an increase of 5-7 orders of magnitude. The RP targets a gene encoded on the phagemid, which sits downstream of its two origins of replication. As a resident plasmid in the host cell, the RP continuously produces hypermutated cDNA and integrates it into the phagemid genome once the latter infects the cell. The F’ (or F factor, which encodes the F pilus required for M13 phage infection) encodes a tetracycline resistance marker and cells are grown in the presence of tetracycline to ensure continuous phagemid infection and production. After 18 hours, phagemid are titered and subject to NGS. Allelic frequencies determined by NGS can be used together with titer data to estimate the number of phages in the population that possess a certain genotype. Both NGS-determined editing rates and titer data are provided for each experiment.
[0207] Our initial attempts to produce phagemid in the s2064 and s2065 cell lines yielded inconsistent results, with cells displaying signs of significant metabolic burden: greatly reduced growth rates, low electroporation efficiencies and periodic loss of antibiotic resistance, indicativeAty. Dkt. No. 125141.04976 MGH2024-443of plasmid loss. We discovered that the transformation of the HP into cells with a AsbcB genotype was necessary and sufficient to produce this metabolic burden phenotype. Both standard s2060 cells and cells with only the RecJ modification (s2061) exhibited no fitness penalty when transformed with HP and constantly produced robust phagemid titers (data not shown). sbcB is the primary 3 ’ to 5 ’ exonuclease in E. coli and is implicated in several DNA damage repair pathways.68To our knowledge, this work is the first indication that sbcB may be an essential gene implicated in the mitigation of genomic stress associated with M13 phage infection. To determine whether the AsbcB genotype was needed for efficient phagemid editing, we performed an experiment in which both the RP and the phagemid were maintained as resident plasmids without the presence of HP. In this experiment, cells were grown in the presence of Carbenicillin, (CarbR is the corresponding marker on the phagemid), grown to OD=0.3 and induced with arabinose to enable reticulase editing. The transcription factor gene clopt encoded on the phagemid was the target of a discrete AT>CG edit encoded on a 68bp THR (FIG. 15). We found that the total editing efficiency in this experiment was approximately one log lower than for chromosomal loci (0.3% in the s2064 cell line as opposed to the 3-8% for 80bp THRs in the chromosome). Such lower editing could be due to the higher copy number of the phagemid (more genomes need to be edited in the population to achieve the same aggregate editing efficiency) or to the mechanics of phagemid genome replication. The ColEl plasmid ori on the pLITMUS* phagemid stimulates the PirA pathway of gene replication, which involves the synthesis of far shorter Okazaki fragments (250bp) than the 1000-2000bp Okazaki fragments found in the E. coli chromosome.69Shorter Okazaki fragments can lead to more competition on the lagging strand between DnaG-synthesized RNA primers and reticulase-produced cDNA, leading to lower editing efficiency (as seen for chromosomal editing in FIG. 6A). Nonetheless, the resident plasmid editing efficiency we observed in s2061 cell line ARecJ) only mildly underperformed that in the s2064 cell line (AsbcB ARecJ), and not to a statistically significant extent. Thus, we proceeded with the s2061 (ARecJ) cell line for the remainder of our work.
[0208] Our initial experiments editing the phagemid as a phage in the s2061 cell line (together with HP) yielded significantly higher editing efficiencies than when editing this same construct as a resident plasmid (FIG. 16A). We observed -2% editing efficiency when using a medium copy number RP (CloDF13 ori, ~40 / copies cell), and -1.2% editing efficiency when using a low copy number RP (sclOl ori, ~5 / copies cell). Titering revealed more than 1010phagemid produced in anAty. Dkt. No. 125141.04976 MGH2024-44318-hour period (FIG. 16B). We observed no statistically significant correlation between editing efficiency and titer expansion across four conditions each with six biological replicates (FIGS.16C and 16D). This finding was crucial, as the utility of the reticulase would be undermined if its activity precluded efficient library expansion.
[0209] We next sought to understand the inherent impact of various reticulase components on phagemid titer expansion irrespective of whether phage was being specifically targeted (FIG. 17). We tested phagemid titer expansion over multiple generations in host cells with either a full nontargeting reticulase with all reticulase proteins (Avd, bRT, CspRecT, and MutL*), a non-targeting reticulase with only Avd and bRT proteins (ACspRecT, and AMutL*) or no reticulase (FIG. 18A). We found that the expression of CspRecT and MutL* (but not Avd and bRT) resulted in a consistent depression of phage titer expansion by approximately 1.2-2.2 orders of magnitude (FIGS. 18B and 18C). We suspect that this drop in titer was due to the activity of CspRecT, which indiscriminately binds single-strand DNA and could plausibly interfere with the ability of a singlestranded DNA phage to replicate. Nonetheless, the ability of our system to consistently produce more than l*109cfu mb'1phagemid underscored that this reduction in titer did not undermine its ability to still generate large libraries. Moreover, the ability of phagemid to persist over multiple serial passages (each serial passage involving a l*10'3dilution, for a total dilution factor of 1*10’12by the end of this experiment) underscored that the mild depression in titer caused by the full reticulase did not preclude its use for multi-passage or automated dilution continuous evolution experiments, as in phage-assisted noncontinuous evolution (PANCE) or PACE, respectively. Finally, we sought to understand the impact of starting phagemid titer on reticulase editing efficiency. Phage replicating under the control of a conditional propagation circuit (as in PACE) constitute a dynamic population with titers ranging from l*103cfu mb'1to l*108cfu mL'1.32Thus, although of less relevance to the generation of phage display libraries, it was important to validate that the RP could consistently edit phage at a variety of titers. We infected cells with a serial dilution of phage ranging from l*102cfu mb'1to l*109cfu mL'1(FIG. 19). Reticulase editing efficiency was effectively unchanged over this titer range, underscoring its translatability as a tool for the continuous hypermutation of dynamic phage populations. We found that the 1 * 104cfu mL'1starting titer produced the highest final titer, while infection at l*108cfu mL'1and l*109cfu mL'1resulted in negligible expansion.
[0210] The observation that reticulase editing efficiency was significantly higher when theAty. Dkt. No. 125141.04976 MGH2024-443phagemid was propagating as a phage than when the phagemid was propagating as a resident plasmid suggested to us that a secondary gene-editing pathway - one only active when the phagemid was propagating as a phage - might be at play. In nature, M13 phage replicates its genome using a method of rolling circle amplification (RCA) that does not involve discontinuous polymerization or Okazaki fragment synthesis (FIG. 20).70In this pathway, the invading singlestranded M13 genome (the +strand) serves as the template for the synthesis of the complementary (-) strand. The -strand is then used as the template for RCA, in which the existing +strand is displaced from the -strand via the synthesis of a new +strand. This new +strand can then serve as the template for a subsequent round of -strand synthesis, while the + / - double stranded genome can perform an additional round of RCA. As this process continues, the expansion of the population can be modeled as a Fibonacci sequence, in which the total number of genomes at any given generation (Fn) is equal to the sum of the populations from the previous two generations (Fn = Fn-i + F11-2). In contrast, the ColEl plasmid origin amplifies its resident genome using a replication fork and leading / lagging strand synthesis as in the E. colt chromosome. Populations replicating through leading / lagging strand synthesis (E. coli genomes and ColEl plasmid) can be modeled as a simple exponential (Fn= 2n). The M13 RCA pathway is facilitated by phage proteins pll, pV and pX, the gene products of gll, gV and gX, which in our system are encoded on HP71,72Because the HP was absent when we propagated the phagemid as a resident plasmid, RCA could not have been active during this experiment. In contrast, RCA is active during phage propagation (by definition) due to the presence of the HP. Thus, we were able to atribute this enhanced editing efficiency during phage propagation to the ability of cDNA to integrate into the phagemid while it was undergoing RCA from the M13 origin of replication.
[0211] To our knowledge, these data represent the first instance of the targeted in vivo editing of a genome undergoing RCA. In contrast to ‘recombineering’ which is used to refer to the annealing of single-stranded DNAs in place of an Okazaki fragment on the lagging strand of a replication fork, we refer to this novel gene-editing pathway as Reticulate Fibonacci Expansion, or ‘Reticulate Fibonasion.’ While we demonstrate reticulate fibonasion on Ml 3 phage / phagemid, we underscore that all of the phage most commonly used for phage display (M13, fl, and fd) undergo RCA / Fibonacci Expansion.72-73Thus, our method represents a highly generalizable approach for the rapid modification of phage genomes used in a wide variety of industrial and commercial settings.Aty. Dkt. No. 125141.04976 MGH2024-443
[0212] Optimization of the M13 phagemid architecture enables high efficiency reticulase editing. We hypothesized that a synergistic arrangement between the two origins of replication on the phagemid could significantly improve the efficiency of reticulate fibonasion. Typical phagemids are constructed with the M13 and ColEl ori facing in opposite directions from one another (M13- ColEl+). We made a series of phagemid constructs in which the two ori were pointing in each of the four possible arrangements relative to the gene of interest (FIG. 24A). The M13+ C0IEI+ phagemid failed numerous cloning attempts, including full gene block synthesis. It is possible that such an architecture results in a phagemid with high copy number that is metabolically taxing and difficult to clone. However, we were able to identify M13+ C0IEI+ clones with minor polymorphisms in the origins of replication that were stably maintained and did not appear metabolically taxing when transformed into our cloning strain; we labeled these vl (M13 ori A469G), v2 (M13 ori G487A), and v3 (ColEl ori G191T). A recent PACEmid paper reported a similar case of metabolic burden for pLITMUS* phagemids that were alleviated through ori polymorphisms that reduced copy number.63We suspect that the polymorphisms identified here may play a similar role in mitigating metabolic burden by reducing phagemid copy number. The M13- ColEl- orientation did not provide difficulty cloning.
[0213] We tested a total of six constructs and found that the orientation of the M13 ori in the + direction (towards the gene being edited) universally enhanced editing to nearly 10% (FIG. 24B). We found that the v2 and v3 M13+ C0IEI+ constructs resulted in a lower final titer (8*108cfu mL’1), while all other four constructs yielded the higher titer 1.2*1O10cfu m '1(FIG. 24C). These results represented a truly impressive quantity of reticulase-edited phage: for the vl M13+ C0IEI+ phagemid, more than 1 billion (1.2*109) edited phage were produced from a single mL of culture in an 18-hour period. For the v2 and v3 M13+ C0IEI+ constructs, it is conceivable that the polymorphisms which lower metabolic burden (presumably via copy number reduction) also impaired replication and result in a lower titer. The finding that the Ml 3+ orientation provided the best editing suggested that cDNA was effective at replacing the -strand at the target locus during -strand synthesis. Because the -strand serves as the template for RCA, it is not surprising that selective replacement of this strand - as opposed to the +strand - resulted in higher editing efficiency. Mathematical modeling of Ml 3 populations undergoing Fibonacci expansion suggests that the maximal editing efficiency that should be achievable if the -strand was replaced by a reticulase-produced cDNA is 61.08%; if the +strand was replaced, this number drops to 23.61%Aty. Dkt. No. 125141.04976 MGH2024-443(FIGS. 20-22). Thus, our experimental finding that -strand annealing was the most efficient means of editing agrees with the predicted outcome. For the ColEl ori (amplifying exponentially), the maximal editing efficiency that should be achievable through recombineering is 25% (FIGS. 20 and 23). However, due to the finding from our plasmid propagation assay that indicated recombineering with the ColEl ori was highly inefficient (0.25% in s2061), the finding in our phage propagation editing assay that the orientation of the ColEl ori was essentially a non-factor for editing efficiency was not surprising. To determine whether the ColEl ori had any statistically significant contribution to editing, we tested 42 replicates of M13+ ColEl+ vl and M13+ ColEl-phagemids (FIG. 25A). Analysis of these data with an unpaired t test indicated that the difference in editing efficiency (7.551% vs 6.262%) was significant, with a P value of 0.0409 (FIG. 25B). Thus, we postulate that recombineering with the ColEl ori is a minor contributor to phagemid editing while Reticulate Fibinasion (-strand replacement and RCA from the M13 ori) is the major contributor to phagemid editing (FIG. 26). Collectively these results demonstrate the invention of a novel gene-editing method for the efficient reticulase-mediated editing of M13 phagemid in vivo. Because the vl M13+ C0IEI+ backbone performed the best of our panel, we proceeded with this backbone for all subsequent experiments; phagemid with this backbone (M13+ C0IEI+ ori orientation) we have defined / refer to as Ml 3 ‘tracemid.’
[0214] Fine-tuning the M13 life cycle to increase reticulase editing efficiency and final library titer. We hypothesized that the life cycle of M13 phage could be fine-tuned to enhance reticulase editing efficiency. pV is the protein product of M13 gV and regulates the duration of Fibonacci expansion during the Ml 3 life cycle (FIG. 27A). pV is a single-strand DNA binding protein that associates with the single-strand M13 genome (the +strand) and prepares it for packaging and cell export. In early infection, pV concentration in the cell is low and +strands are free to undergo -strand synthesis and subsequent RCA. In late infection, pV reaches a high concentration in the cell at which point all +strands are bound by hundreds of copies of pV, arresting RCA and preparing the +strands for packaging. Thus, pV concentration in the cell is the critical determinant for the duration and extent of Ml 3 Fibonacci expansion. We hypothesized that because our gene-editing method was contingent on editing phage during Fibonacci expansion, extending this phase of the Ml 3 life cycle would provide a larger time window during each infection when the RP could introduce an edit, leading to higher editing efficiency. At the same time, additional cycles of Fibonacci expansion could also lead to a higher titer, helping to compensate for the reduction inAty. Dkt. No. 125141.04976 MGH2024-443titer induced by CspRecT / MutL*, and resulting in a larger final library size. A previous study published a series of spacer sequences upstream of the gV start codon which reduced the expression of pV in the cell by 10-20% arbitrary translational units.74These spacers were reported to increase the final titer by as much as 5-fold. We tested HP with each of these spacer sequences upstream of gV (FIG. 27B). As expected, as the concentration of pV was reduced in the cell, the editing efficiency of the RP increased from approximately 10% to 12%. The spacer reported to reduce the pV concentration the most (80% of WT) increased titer with respect to the WT spacer by more than an order of magnitude (FIG. 27C). These data confirmed that the life cycle of Ml 3 phage could be fine-tuned to increase reticulase editing efficiency and final library titer.
[0215] To further improve the editing efficiency on tracemid, we tested a panel of eight singlestranded annealing proteins (ssap) within the reticulase (FIG. 65A). We found that CspRecT, EcTRecT, and P22 erf all produced >1% editing across two targets, while the other members of the panel gave low to undetectable editing efficiency (FIG. 65B). Intriguingly, we found that although EcTRecT produced comparable editing efficiency to CspRecT, this ssap appeared to significantly suppress titer expansion by two to three orders of magnitude relative to CspRecT (FIG. 65C). To confirm this finding, we repeated the CspRecT and EcTRecT conditions on one of the tracemid and generated similar results (FIGS. 66A-66B). It is conceivable that a ssap with high affinity for both linear DNA (such as reticulase-produced cDNA) and circular DNA (such as Ml 3 genome) could suppress Ml 3 titer expansion while also enabling high efficiency editing (the phenotype seen for EcTRecT). In contrast, a ssap that required exposed 375’ DNA ends on one of its substrates for high efficiency annealing could conceivably produce high editing efficiency without disrupting the replication / life cycle of a circular DNA genome (the phenotype seen for CspRecT). Collectively these data demonstrate that optimal single strand annealing protein activity balances reticulase editing efficiency with efficient M13 Tracemid titer expansion. Because high-titer highly edited populations are desirable, we proceeded to use CspRecT as the reticulase ssap for all further experiments.
[0216] Reticulase optimization is a codon optimization scheme to enhance hypermutation specificity in genes encoded on extrachromosomal targets. In nature, the non-diversified amino acid codons of DGR variable regions have evolved to become highly de-adenated (lacking the use of adenine). The depletion of adenine content outside of variable codons allows DGRs to hyperfocus their diversity on codons which encode solvent-exposed, antigen engaging amino acidAty. Dkt. No. 125141.04976 MGH2024-443residues while avoiding hypermutation of the scaffold residues that support them.28We hypothesized that a similar ‘de-adenation’ codon optimization approach could be taken for target gene regions encoded on Ml 3 phagemid to enable increased precision hypermutation of select codons.
[0217] Our intended ‘reticulase optimization’ codon optimization scheme had three primary objectives. First, wherever possible, reduce the adenine content of the coding strand. This objective entailed the substitution of codons with higher adenine content for ones with lower or no adenine content. Three of the canonical amino acids (C, W, F) can only be encoded with GCT codons and are therefore always protected from hypermutation. Seven of the 20 canonical amino acids (R, S, A, V, L, G, P) contain an adenine in the wobble position of their codon (the third nucleotide) that can simply be substituted for one of the other three nucleotides without a change in the encoded amino acid. Thus, half of the canonical amino acids can be encoded without the use of adenine in their codon; not coincidentally, these are the 10 amino acids that overwhelmingly form the invariable structural scaffold regions of DGR variable proteins. Of the other 10 amino acids, all but two (N and K) can be encoded by codons containing only one adenine. Thus, for members of this set that contain codons with two adenines, it is possible to substitute a codon which contains only one. Of those two amino acids which must have two adenines present in their codon, N (AAC / AAT) cannot be substituted, while the K codon AAA can be substituted for AAG.
[0218] The second objective of our codon optimization scheme was to avoid drastically increasing the GC content of optimized sequences where possible. While not intrinsically detrimental to gene expression, gene regions with high GC content can be difficult to clone or order as gene block sequences. Thus, for codons chosen to be substituted for de-adenated counterparts (objective one) substitutions of A for T were made where possible to avoid increases in GC content. The third objective of our codon optimization scheme was to generally avoid the substitution of codons with a lower codon frequency than those replaced. It bears underscoring that translation initiation, not elongation, is the limiting factor of protein expression. The substitution of low- abundance codons (for instance CGA, an Arginine codon with 7% use frequency in E.coli) for high-abundance codons (for instance CGT, an Arginine codon with 36% use frequency in E.coli) does not inherently increase the expression level of a given protein. Encoding of a protein with a long stretch of low-frequency codons can lead to reduced expression and is best avoided where possible.
[0219] Taking these factors into account, we wrote a simple batch replace function which makesAty. Dkt. No. 125141.04976 MGH2024-443a series of find and replace codon substitutions to any input nucleic acid sequence. A total of 17 of the 61 possible amino-acid coding codons are replaced anywhere they appear in the input sequence, impacting a total of 12 encoded amino acids (FIG. 28A). Of these substitutions, eight result in the substitution of a higher-frequency codon (Afrequency in E. coli > 4%), six result in the substitution of a codon with a negligible frequency differential (Afrequency in E. coli<+-4° / o) and three result in the substitution of a lower-frequency codon (Afrequency in E.coli >- 4%). Eight of these codon substitutions increase the codon GC content while the other seven do not. An example is provided in FIG. 28b for the reticulase optimization of a target region in the transcription factor encoding gene clopt. The adenine content of the sequence is reduced from 26.5% to 17.4%, a removal of more than a third (34.3%) of adenines present. The total GC content of the sequence is raised by 9.1%.
[0220] We underscore that such reticulase optimization is not required for the functioning of our system, nor is it required for precision hypermutation of select codons. As shown, the high-precision chromosomal hypermutation of rpob S531 was achieved exclusively through homology shielding, despite this sequence not being reticulase-optimized and having an AAA Lysine codon present in the template overlap region. Many therapeutically relevant targets we intend to evolve with our system, such as antibodies and soluble TCRs, are mammalian in origin and are ordered as codon-optimized gene blocks for prokaryotic expression. As such, these sequences are already synthesized exogenously and can employ any additional codon optimization schemes desired by the researcher. During our research, we have ordered numerous reticulase-optimized sequences and have never encountered synthesis errors with commercial vendors. Moreover, reticulase-optimized genes appear fully functional in assays dependent on protein expression and activity (FIG. 33).
[0221] Reticulases enable rapid generation of antibody libraries encoded on M13 tracemid.To demonstrate the ability of our system to generate libraries of clinically translatable biomolecules, we used a reticulase to hypermutate Trastuzumab, an FDA-approved antibody used to treat type II metastatic breast cancer. Trastuzumab binds to Human Epidermal Growth Factor-2 (HER2) presented on the surface of HER2 positive breast cancer cells.73Only -20% of HER2 positive breast cancer cells present high surface levels of this antigen, while the remaining cells present low levels. We reasoned that higher-affinity variants of Trastuzumab could prove useful for the clearance of low-level HER2 positive breast cancer cells. Thus, we sought to generate aAty. Dkt. No. 125141.04976 MGH2024-443library of Trastuzumab variants encoded on Ml 3 tracemid which were rationally diversified at positions most likely to improve affinity for HER2. Analysis of the crystal structure of Trastuzumab bound to HER2 revealed that numerous residues in the light chain CDR3 loop made direct contact with this antigen (FIG. 29A). We chose to diversify Histidine 91 in this region as a proof of principle. A geneblock encoding Trastuzumab in scFv format was ordered with the light chain CDR3 loop region reticulase-optimized and was cloned into our high efficiency (vl M13+ ColEl+) tracemid backbone. A 73bp THR was selected to diversify H91 with full codon saturation (AAA).
[0222] To generate a rich library, we diversified Trastuzumab over a series of three serial passages in which the edited phage from of each passage were diluted and used to seed the subsequent passage (FIG. 29B). By passage three the total sum of all insertions and deletions in the sequence remained less than 1% (0.91%+-0.13% as opposed to 0.63%+-0.12% in a non- targeting control, data not shown). Because we were only interested in perfect in-frame modifications, only reads with no frameshift modifications were analyzed. With each serial passage, the proportion of tracemid that contained a codon other than CAT (the endogenous codon used to encode Histidine) increased from 6% to 11% to 18% (FIG. 29C). We performed allelic mapping at this position over each successive generation (FIG. 29D). In passage 1, we identified reads for 17 of the 19 possible amino acid substitutions introduced at position 91 at a read density of 80,000 reads / sample over six biological replicates. By passage two all 19 amino acid interconversions were detected. We underscore that the absence of two amino acids (F and C) in passage one is not inherently indicative of their absence from the population, but more likely indicative of their occurrence at an incidence below the level of detection for 80,000 reads / sample. Thus, for passage three we increased our read density to 100,000 reads / sample. We detected all amino acids at position 91 at an incidence in the phage population at or exceeding ~4.2* 10’5. All but M, W and C were present at an incidence of greater than ~l*10'4. While these incidences may appear low, we highlight that a conditional Ml 3 phage propagation circuit (the selection method used in PACE) can readily select for favorable mutations present in a phage population at an incidence of approximately ~l*10'8.32Thus, even those amino acids with the lowest probability of being generated (such as W, only encoded by one codon that the reticulase must access through a triple AAA > TGG substitution during cDNA synthesis and then paste perfectly in frame over a CAT codon during reticulate fibonasion) were present in the population at an excess of more than 4000-fold than would beAty. Dkt. No. 125141.04976 MGH2024-443required for selection in PACE. Similarly, the final titer of the passage three population was 4 17*108cfu mL’1, meaning that even the lowest frequency amino acid (M, incidence of 4.19*10’ ’) was encoded in-frame by approximately 17000 tracemid in the population. Such an aggregate total is far in excess of that required for selection via conventional biopanning as in phage display.76Of the 64 possible codons, only one was not detected in the targeting dataset after passage three: CCC (one of the four Proline codons) (FIG. 30B). This finding was consistent with the fact that the A>C transversion is the lowest frequency bRT generated mutation on template adenines.
[0223] To underscore that these substitutions were statistically significant from control, we performed allelic mapping of all codons introduced at position 91 (FIGS. 30A and 30B). A nontargeting control was used for comparison in which an RP was set to diversify rpoB. Because the RP encodes MutL*, which can increase the rate of random mutations by approximately 30-fold, this non-targeting control allowed us to determine whether the mutations mapped were truly attributable to cDNA integration or to random background mutagenesis.60It is expected that single nucleotide changes in the CAT H91 codon (for instance CAT>CAG) could arise over multiple passages with MutL* expression. Alternatively, incorrect base calls in the Illumina MiSeq could also generate reads registered as single-nucleotide polymorphisms. However these data were generated on a run with 85% Q30 score, which meets Illumina quality control standards for the Micro kit used.77As expected, essentially all non-CAT codons in the non-targeting passage three control (0.2% of reads as opposed to 18% in the passage three targeting condition) could be introduced via single-nucleotide polymorphisms in CAT. These data allowed us to identify which substitutions in the targeting condition could be attributed to cDNA integration. For instance, while the introduction of proline codon CCT in the targeting control was not statistically significant compared to control, the introduction of proline codon CCA was. Indeed, for each amino acid at least one codon was present at a level that was statistically significant from control. These data allowed us to confidently conclude that the RP was able to introduce all possible amino acid substitutions to a chosen position on a therapeutically relevant target encoded on an M13 tracemid. Moreover, all substitutions were introduced at an incidence level high enough to enable selection via a variety of commonly used selection methods.
[0224] This phagemid / tracemid hypermutation method has numerous potential commercial and industrial applications. We highlight that the construction of phage display libraries withAty. Dkt. No. 125141.04976 MGH2024-443degenerate oligonucleotide primers and ex vivo cloning can take more than two weeks, and must be quality controlled prior to phage production to ensure adequate diversity.65,78End-to-end our method can be performed in under a week, and only involves the cloning of a monoclonal RP, which is trivial and highly cost-effective as discussed. Moreover, our method generates libraries directly on phage, eliminating the time required for phage production following ex vivo cloning. Although here we only report the saturation of a single codon (and are currently in the process of demonstrating multiplexing), we underscore that the rapid generation of single-codon saturation libraries is of immense utility to identify valuable gain of function mutations in therapeutically relevant biomolecules that are inaccessible via conventional continuous evolution methods.18
[0225] Targeted reticulase-assisted continuous evolution (TRACE) combines reticulate fibonasion and conditional M13 phage propagation circuits to enable the continuous rationally directed evolution of biomolecules. While phage display is an effective method for the screening of large phage / phagemid libraries, this method requires tedious rounds of biopanning and phage ELISA which can take weeks and are highly labor intensive.76In contrast, conditional Ml 3 phage propagation circuits can enable the selection and isolation of phagemid encoding functional clones in under a week with minimal researcher intervention. A variety of directed evolution methods have employed this selection method, including PACE, PACEmid, Phage-Assisted Continuous Selection (PACS), as well as various derivatives of these methods.13,32,67,79,80Herein we describe the specifics of this selection scheme as it applies to our system, although variations do exist (FIG. 31). In our selection scheme, a selection tracemid (ST) is cloned encoding a gene of interest (GOI) that expresses a protein of interest (POI). The resident helper plasmid (HP) in the host strain encodes all genes required for Ml 3 phage propagation except gVI, which encodes pVI. pVI stabilizes the tail region of the phage coat and is required for release of the phage from the host cell.64ST without pVI are non-infectious and are actively selected against in the population. To construct a conditional propagation circuit, gVI is encoded on an accessory plasmid (AP) in which the expression of pVI is controlled by a promoter that is linked to the activity of POI encoded on the ST. Those ST that encode POI variants that can activate this synthetic circuit are rewarded with pVI and are able to escape from the cell and continue to propagate. Performing multiple passages with fresh host cells and ST from the previous passage allows the overall proportion of phage in the population encoding POI variants with gain-of-function mutations to increase. Over multiple generations, those ST encoding POI with favorable gain-of-functionAty. Dkt. No. 125141.04976 MGH2024-443mutations come to dominate the population. Thus, this method affords a rapid means of selection for POI variants with desirable gain-of-function mutations. This selection paradigm has been used to evolve a wide variety of biomolecules in E. coli, including Cas9 variants,81,82proteases,83‘83transcription factors,67polymerases,13, 86single chain antibody fragments,19,87cytosine base editors,88adenine base editors,89tRNA synthetases,90agricultural toxins,91biosynthetic pathways,92TALENs,33serine integrases,93and molecular glues.94Biomolecular properties selected for such molecules have included enhanced catalytic activity, antigen specificity, antigen affinity, stability and solubility.32,87Although here we report the use of multiple serial passages to facilitate continuous selection (taking course over 3-5 passages over 3-5 days), we note that this method can also be performed in a bioreactor in which new host cells are continuously flowed into a chamber (the lagoon) which contains tracemid, while old host cells are deposited into a waste container. Such a continuous flow bioreactor enables even less researcher input than serial passaging, while also allowing stringency modulation via adjustment of flow rate. We underscore that the underlying selection scheme used between either serial passaging or a bioreactor (i.e. the use of a conditional M13 phage propagation circuit) is fundamentally the same, and our method could readily be adapted to be used in a continuous flow bioreactor.
[0226] In a typical TRACE workflow, a variety of methods (PLMs, crystal structure analysis, mutational scans / screens, artificial intelligence) may be used to rationally nominate positions in a POI that are the most likely to have an impact on fitness (FIG. 32). An RP is cloned with an array or regRNAs programmed to saturate the chosen positions in the GOI encoded on the ST. Over the course of the TRACE, the RP saturates these positions on ST via reticulate fibonasion. A conditional Ml 3 phage propagation circuit corresponding to the activity of the POI is used to continuously select for variant POI with desirable gain-of-function mutations. As an ST with a favorable gain-of-function mutation at one of the saturated sites comes to dominate the population, it acts as the recipient template for saturation at another site. This continuous accumulation of favorable gain-of-function mutations in the GOI effectively automates ISM in vivo with minimal researcher intervention. Moreover, because phage with favorable mutations dominate the population by the end of the TRACE, simple phage titering and Sanger sequencing allows isolation and identification of the selected genotype without any need for biopanning or tedious phage ELISA. Thus, TRACE enables the full process of ISM (iterative rounds of rational library generation and screening, followed by clonal isolation) to be performed in under a week withAty. Dkt. No. 125141.04976 MGH2024-443minimal researcher intervention. We underscore that a wide variety of Ml 3 phage selection circuits have already been optimized and published in the literature and could readily be adapted for use in our system.
[0227] In this work we wanted to demonstrate the validity of our method with a simple proof-of concept selection circuit (FIG. 33A). We adapted a circuit that has been used to select variants of the dual transcriptional activator / repressor cl from E. coli bacteriophage lambda. In this circuit, an optimized variant of cl (denoted clopt) is encoded on the phagemid (at the time this experiment was performed we had not yet discovered the high-efficiency tracemid backbone). This transcription factor binds to operator 1 and 2 in the lambda phage bi-directional Promoter PRPRM and promotes transcription in the forward direction while repressing transcription in the opposite direction.64For our purposes, we obliterated the third operator site in this promoter which regulates transcription in the opposite direction, turning this system into a simple, unidirectional ON circuit for clopt binding. The clopt gene region encoding the amino acids E28-K43 was reticulase optimized, and a dual stop codon was inserted in place of Y38 and E39 (TAC GAG replaced with TAA TAG) (FIG.28C). An RP was cloned in which the THR contained a discrete edit which would revert these stop codons to the YE amino acids typically present. Those phagemid which are reticulase edited should therefore be able to express clopt, allowing activation of the circuit and continuous selection for RP-edited phage. Thus, this simple proof of concept experiment allows us to demonstrate that a conditional Ml 3 phage propagation circuit can select for gain of function mutations in a gene of interest encoded on an Ml 3 phagemid which were introduced by an RP. We used a non-targeting RP with a THR set to diversify mCherry as a negative control in this experiment. We note that the Promoter PRPRM used in this experiment is known to be leaky: namely, a low level of expression of pVI off this promoter is expected irrespective of the activity of clopt. This circuit still provides selection for phage encoding clopt but does not penalize non-clopt encoding phage to such an extent they are wholly unable to propagate.64
[0228] Our host s2061 strain was transformed with HPAgVI, the AP, and either the targeting or non-targeting RP. Cells were grown to OD=03, induced with lOmM arabinose and infected with phagemid at a starting titer of PIO6cfu mL’1. Each passage was run for 18-22 hours, and the resulting phagemid were back diluted 1 / 1000 to start the next passage in fresh host cells grown to OD=0.3. By passage 3, the phagemid titer in the host strain containing the targeting reticulase (that which reverted the stop codon) shot up approximately three orders of magnitude more than theAty. Dkt. No. 125141.04976 MGH2024-443phagemid tier in the cells with the non-targeting RP (FIG. 33B). NGS of the population at each passage of this campaign demonstrates that the population of phagemid with the stop reverted went from only 2% in passage 1 to more than 70% by passage three and nearly 100% by passage 5 (FIG.33C). Thus, NGS sequencing demonstrates that these RP-edited phage did indeed come to dominate the population. Finally, to demonstrate the ease of clonal isolation, we performed fullplasmid Zeroprep sequencing on one of the isolated colonies from our titering plates (FIG. 33D). A representative clone from passage three is shown mapped against the reference phagemid sequence which was used to start the evolution campaign. This sequencing revealed the reversion of the stop codons (as was programmed in the THR) concomitant with an A to G mutation in the TAC Tyrosine codon, changing it to TGC (Cystine), This adenine-specific mutation is the hallmark of bRT mediated mutagenesis. Notably, we did not observe mutations outside of the diversity accommodation region of the THR used to make this mutation on the GOI, or outside the GOI in the SP backbone. Such targeted mutagenesis stands in stark contrast to that which is observed in a typical PACE experiment when using a random mutagenesis plasmid (MP), which increases the global mutagenesis rate within the phage GOI, the phage backbone, and the host cells. Indeed, for a typical PACE with the mutagenesis plasmid MP6, the majority of mutations in the SP are expected to be made outside the GOI entirely.89We highlight that random mutagenesis of gene regions outside those which are likely to determine fitness for a desired phenotype can lead to ‘cheating’ in a selection circuit, as noted above. Thus, these data validate not only the feasibility of TRACE but also underscore its technical advantages over methods that rely on random mutagenesis (PACE).
[0229] Conclusion
[0230] In this manuscript we have described a system for the continuous rationally directed evolution of biomolecules we refer to as TRACE (Targeted Reticulase-Assisted Continuous Evolution). This method involves four primary innovations: 1) the use of a DGR-RT to generate short single stranded DNAs in the cytoplasm of a prokaryotic cell with high degeneracy at select positions 2) a method for the continuous integration of such single stranded DNAs into an E.coli genome in vivo 3) the invention of a novel gene-editing method (reticulate fibonasion) for the integration of such single stranded DNAs specifically into the genome of an Ml 3 phagemid undergoing Rolling circle amplification as a means of targeted library generation and 4) the combination of this phagemid gene editing method with conditional M13 phage propagationAty. Dkt. No. 125141.04976 MGH2024-443circuits to enable continuous in vivo library generation and selection, followed by clonal isolation by phagemid titering. We argue that both the constituent components of this invention as well as their synergistic combination has numerous academic, industrial and commercial applications for the rapid directed evolution of diverse biomolecules with desirable biophysical properties.
[0231] References1. Jiang, K. et al. Rapid in silico directed evolution by a protein language model with EVOLVEpro. Science 0, eadr6006 (2024).2. Hie, B. L. et al. Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol. 1-9 (2023) doi:10.1038 / s41587-023-01763-2.3. Robinson, R. A., McMurran, C., McCully, M. L. & Cole, D. K. Engineering soluble T-cell receptors for therapy. FEBSJ. 288, 6159-6173 (2021).4. Li, Y. et al. Directed evolution of human T-cell receptors with picomolar affinities by phage display. Nat. Biotechnol. 23, 349-354 (2005).5. Reetz, M. T. & Carballeira, J. D. Iterative saturation mutagenesis (ISM) for rapid directed evolution of functional enzymes. Nat. Protoc. 2, 891-903 (2007).6. Acevedo-Rocha, C. G., Hoebenreich, S. & Reetz, M. T. Iterative Saturation Mutagenesis: A Powerful Approach to Engineer Proteins by Systematically Simulating Darwinian Evolution, in Directed Evolution Library Creation: Methods and Protocols (eds. Gillam, E. M. J., Copp, J. N. & Ackerley, D.) 103-128 (Springer, New York, NY, 2014). doi:10.1007 / 978-l-4939-1053-3_7.7. Dolgin, E. First soluble TCR therapy opens ‘new universe’ of cancer targets. Nat. Biotechnol. 40, 441-444 (2022).8. Gumulya, Y., Sanchis, J. & Reetz, M. T. Many Pathways in Laboratory Evolution Can Lead to Improved Enzymes: How to Escape from Local Minima. ChemBioChem 13, 1060-1066 (2012).9. Reetz, M. T. & Sanchis, J. Constructing and Analyzing the Fitness Landscape of an Experimental Evolutionary Process. ChemBioChem 9, 2260-2267 (2008).10. Yang, W. P. et al. CDR walking mutagenesis for the affinity maturation of a potent human anti- HIV-1 antibody into the picomolar range. J. Mol. Biol. 254, 392-403 (1995).11. English, J. G. et al. VEGAS as a Platform for Facile Directed Evolution in Mammalian Cells. Cell 178, 748-761. el7 (2019).12. Wellner, A. et al. Rapid generation of potent antibodies by autonomous hypermutation inAty. Dkt. No. 125141.04976 MGH2024-443yeast. Nat. Chem. Biol. 17, 1057-1064 (2021).13. Esvelt, K. M., Carlson, J. C. & Liu, D. R. A system for the continuous directed evolution of biomolecules. Nature 472, 499-503 (2011).14. Molina, R. S. et al. In vivo hypermutation and continuous evolution. Nat. Rev. Methods Primer 2, 1-22 (2022).15. Badran, A. H. & Liu, D. R. Development of potent in vivo mutagenesis plasmids with broad mutational spectra. Nat. Commun. 6, 8425 (2015).16. Wong, T. S., Roccatano, D. & Schwaneberg, U. Are transversion mutations better? A Mutagenesis Assistant Program analysis on P450 BM-3 heme domain. Biotechnol. J. 2, 133- 142 (2007).17. Wong, T. S., Wong, T. S., Roccatano*, D. & Schwaneberg, U. Challenges of the genetic code for exploring sequence space in directed protein evolution. Biocatal. Biotransformation 25, 229-241 (2007).18. Zhang, E., Neugebauer, M. E., Krasnow, N. A. & Liu, D. R. Phage-assisted evolution of highly active cytosine base editors with enhanced selectivity and minimal sequence context preference. Nat. Commun. 15, 1697 (2024).19. Morrison, M. S., Wang, T., Raguram, A., Hemez, C. & Liu, D. R. Disulfide-compatible phage-assisted continuous evolution in the periplasmic space. Nat. Commun. 12, 5959 (2021). 20. Vidal, L. S., Isalan, M., Heap, J. T. & Ledesma-Amaro, R. A primer to directed evolution: current methodologies and future directions. RSC Chem. Biol. 4, 271-291 (2023).21. Macadangdang, B. R., Makanani, S. K. & Miller, J. F. Accelerated Evolution by Diversity-Generating Retroelements. Anna. Rev. Microbiol. 76, 389-411 (2022).22. Guo, H., Arambula, L., Ghosh, P. & Miller, J. F. Diversity-generating Retroelements in Phage and Bacterial Genomes, in Mobile DNA III 1237-1252 (John Wiley & Sons, Ltd, 2015). doiilO.l 128 / 9781555819217. ch53.23. Wu, L. et al. Diversity-generating retroelements: natural variation, classification and evolution inferred from a large-scale genomic survey. Nucleic Acids Res. 46, 11-24 (2018). 24. Naorem, S. S. et al. DGR mutagenic transposition occurs via hypermutagenic reverse transcription primed by nicked template RNA. Proc. Natl. Acad. Sci. 114, E10187-E10195 (2017).25. Miller, J. L. et al. Selective Ligand Recognition by a Diversity-Generating Retroelement Variable Protein. PLOS Biol. 6, el 31 (2008).Aty. Dkt. No. 125141.04976 MGH2024-44326. McMahon, S. A. et al. The C-type lectin fold as an evolutionary solution for massive sequence variation. Nat. Struct. Mol. Biol. 12, 886-892 (2005).27. Litman, G. W., Rast, J. P. & Fugmann, S. D. The origins of vertebrate adaptive immunity. Nat. Rev. Immunol. 10, 543-553 (2010).28. Roux, S. et al. Ecology and molecular targets of hypermutation in the global microbiome. Nat. Commun. 12, 3076 (2021).29. Liu, M. et al. Reverse Transcriptase-Mediated Tropism Switching in Bordetella Bacteriophage. Science 295, 2091-2094 (2002).30. Handa, S., Biswas, T., Chakraborty, J., Paul, B. G. & Ghosh, P. Structural Requirements for Reverse Transcription by a Diversity-Generating Retroelement. http: / / bioixiv.org / lookup / doi / 10.1101 / 2023.10.23.563531 (2023)31. Handa, S. et al. Template-assisted synthesis of adenine-mutagenized cDNA by a retroelement protein complex. Nucleic Acids Res. 46, 9711-9725 (2018).32. Miller, S. M., Wang, T. & Liu, D. R. Phage-assisted continuous and non-continuous evolution. Nat. Protoc. 15, 4101-4127 (2020).33. Hubbard, B. P. et al. Continuous directed evolution of DNA-binding proteins to improve TALEN specificity. Nat. Methods 12, 939-942 (2015).34. Ringquist, S. et al. Translation initiation in Escherichia coli: sequences within the ribosome- binding site. Mol. Microbiol. 6, 1219-1229 (1992).35. Davis, J. H., Rubin, A. J. & Sauer, R. T. Design, construction and characterization of a set of insulated bacterial promoters. Nucleic Acids Res. 39, 1131-1141 (2011).36. Sandegren, L. & Sjoberg, B.-M. Self-Splicing of the Bacteriophage T4 Group I Introns Requires Efficient Translation of the Pre-mRNA In Vivo and Correlates with the Growth State of the Infected Bacterium. J. Bacterial. 189, 980-990 (2007).37. Guo, H. et al. Target site recognition by a diversity -generating retroelement. PLoS Genet.7, el002414 (2011).38. Handa, S., Reyna, A., Wiryaman, T. & Ghosh, P. Determinants of adenine-mutagenesis in di versi ty -generating retroelements. Nucleic Acids Res. 49, 1033-1045 (2020).39. Robart, A. R., Seo, W. & Zimmerly, S. Insertion of group II intron retroelements after intrinsic transcriptional terminators. Proc. Natl. Acad. Sci. U. S. A. 104, 6620-6625 (2007). 40. Luan, D. D., Korman, M. H., lakubczak, J. L. & Eickbush, T. H. Reverse transcription ofAty. Dkt. No. 125141.04976 MGH2024-443R2Bm RNA is primed by a nick at the chromosomal target site: a mechanism for non-LTR retrotransposition. Cell'll, 595-605 (1993).41. Eickbush, D. G., Luan, D. D. & Eickbush, T. H. Integration ofBombyx mori R2 Sequences into the 28S Ribosomal RNA Genes of Drosophila melanogaster. Mol. Cell. Biol. 20, 213-223 (2000).42. Han, J. S. Non-long terminal repeat (non-LTR) retrotransposons: mechanisms, recent developments, and unanswered questions. Mob. DNA 1, 15 (2010).43. Cost, G. J., Feng, Q., Jacquier, A. & Boeke, J. D. Human LI element target-primed reverse transcription in vitro. EMBO J. 21, 5899-5910 (2002).44. Millman, A. et al. Bacterial Retrons Function In Anti-Phage Defense. Cell 183, 1551- 1561. el2 (2020).45. Bobonis, J. et al. Phage proteins block and trigger retron toxin / antitoxin systems.2020.06.22.160242 Preprint at https: / / doi.org / 10.1101 / 2020.06.22.160242 (2020).46. Ellington, A. J. & Reisch, C. R. Efficient and iterative retron-mediated in vivo recombineering in Escherichia coli. Synth. Biol. 7, ysac007 (2022).47. Gonzalez-Delgado, A., Lopez, S. C., Rojas-Montero, M., Fishman, C. B. & Shipman, S. L. Simultaneous multi-site editing of individual genomes using retron arrays. bioRxiv 2023.07.17.549397 (2023) doi: 10.1101 / 2023.07.17.549397.48. Khan, A. G. et al. An Experimental Census of Retrons for DNA Production and Genome Editing. htp: / / biorxiv.org / lookup / doi / 10.! 101 / 2024.01.25.577267 (2024) doi: 10.1101 / 2024.01.25.577267.49. Liu, W. et al. Retron-mediated multiplex genome editing and continuous evolution in Escherichia coli. Nucleic Acids Res. 51, 8293-8307 (2023).50. Schubert, M. G. et al. High-throughput functional variant screens via in vivo production of single-stranded DNA. Proc. Natl. Acad. Sci. 118, e2018181118 (2021).51. Gontier, N. Reticulate Evolution Everywhere, in Reticulate Evolution: Symbiogenesis, Lateral Gene Transfer, Hybridization and Infectious Heredity (ed. Gontier, N.) 1-40 (Springer International Publishing, Cham, 2015). doi: 10.1007 / 978-3-319-16345-1 1.52. Wannier, T. M. et al. Improved bacterial recombineering by parallelized protein discovery. Proc. Natl. Acad. Sci. U. S. A. 117, 13689-13698 (2020).53. Egbert, R. G. et al. A versatile platform strain for high-fidelity multiplex genome editing.Aty. Dkt. No. 125141.04976 MGH2024-443Nucleic Acids Res. 47, 3244-3256 (2019).54. Morgan, K. Plasmids 101: Origin of Replication, https: / / blog.addgene.org / plasmid-101-origin- of-replication.55. Goldstein, B. P. Resistance to rifampicin: a review. J. Antibiot. (Tokyo) 67, 625-630 (2014).56. Lin, Y.-H., Tai, C.-H., Li, C.-R., Lin, C.-F. & Shi, Z.-Y. Resistance profiles and rpoB gene mutations of Mycobacterium tuberculosis isolates in Taiwan. J. Microbiol. Immunol. Infect. Wei Mian Yu Gan Ran Za Zhi 46, 266-270 (2013).57. Ning, Q. et al. Predicting Rifampicin Resistance Mutations in Bacterial RNA Polymerase Subunit Beta Based on Machine Learning Algorithms. Preprint at https: / / doi.org / 10.21203 / rs.3.rs- 127050 / vl (2020).58. Jing, W. et al. Rifabutin Resistance Associated with Double Mutations in rpoB Gene in Mycobacterium tuberculosis Isolates. Front. Microbiol. 8, 1768 (2017).59. Li, M.-C. et al. rpoB Mutations are Associated with Variable Levels of Rifampin and Rifabutin Resistance in Mycobacterium tuberculosis. Infect. Drug Resist. 15, 6853-6861 (2022).60. Nyerges, A. et al. A highly precise and portable genome engineering method allows comparison of mutational effects across bacterial species. Proc. Natl. Acad. Sci. 113, 2502- 2507 (2016).61. Wang, H. H. et al. Programming cells by multiplex genome engineering and accelerated evolution. Nature 460, 894-898 (2009).62. Halperin, S. O. et al. CRISPR-guided DNA polymerases enable diversification of all nucleotides in a tunable window. Nature 560, 248-252 (2018).63. Davenport, B. I., Tica, J. & Isalan, M. Reducing metabolic burden in the PACEmid evolver system by remastering high-copy phagemid vectors. Eng. Biol. 6, 50-61 (2022).64. Brodel, A. K., Jaramillo, A. & Isalan, M. Engineering orthogonal dual transcription factors for multi-input synthetic promoters. Nat. Commun. 7 , 13858 (2016).65. Brodel, A. K., Jaramillo, A. & Isalan, M. Intracellular directed evolution of proteins from combinatorial libraries based on conditional phage replication. Nat. Protoc. 12, 1830-1843 (2017).66. Brodel, A. K., Isalan, M. & Jaramillo, A. Engineering of biomolecules by bacteriophage directed evolution. Curr. Opin. Biotechnol. 51, 32-38 (2018).67. Brodel, A. K., Rodrigues, R., Jaramillo, A. & Isalan, M. Accelerated evolution of a minimalAty. Dkt. No. 125141.04976 MGH2024-44363- amino acid dual transcription factor. Sci. Adv. 6, eaba2728 (2020).68. Naom, I. S., Morton, S. J., Leach, D. R. & Lloyd, R. G. Molecular organization of sbcC, a gene that affects genetic recombination and the viability of DNA palindromes in Escherichia coli K- 2. Nucleic Acids Res. 17, 8033-8045 (1989).69. Allen, J. M. et al. Roles of DNA polymerase I in leading and lagging-strand replication defined by a high-resolution mutation footprint of ColEl plasmid replication. Nucleic Acids Res.39, 7020-7033 (2011).70. Rakonjac, J., Bennett, N. J., Spagnuolo, J., Gagic, D. & Russel, M. Filamentous bacteriophage: biology, phage display and nanotechnology applications. Curr. Issues Mol. Biol.13, 51-76 (2011).71. Suzuki, T. & Kamiya, H. Easily-controllable, helper phage-free single- stranded phagemid production system. Genes Environ. 44, 25 (2022).72. Smeal, S. W., Schmitt, M. A., Pereira, R. R., Prasad, A. & Fisk, J. D. Simulation of the M13 life cycle I: Assembly of a genetically-structured deterministic chemical kinetic simulation. Virology 500, 259-274 (2017).73. Liu, Q. et al. Chapter 105 - Virus assembly, in Molecular Medical Microbiology (Third Edition) (eds. Tang, Y.-W. et al.) 2131-2175 (Academic Press, 2024). doi:10.1016 / B978-0-12-818619- 0.00162-3.74. Lee, B.-Y., Lee, J., Ahn, D. J., Lee, S. & Oh, M.-K. Optimizing protein V untranslated region sequence in Ml 3 phage for increased production of single- stranded DNA for origami. Nucleic Acids Res. 49, 6596-6603 (2021).75. Molina, M. A. et al. Trastuzumab (herceptin), a humanized anti-Her2 receptor monoclonal antibody, inhibits basal and activated Her2 ectodomain cleavage in breast cancer cells. Cancer Res. 61, 4744-4749 (2001).76. Barbas, C. F. Phage Display: A Laboratory Manual . (CSHL Press, 2001).77. Does my sequencing run look good? | Illumina Knowledge. https: / / knowledge.illumina.com / instrumentation / general / instrumentation-general-reference_material-list / 000001922 (2024).78. Kille, S. etal. Reducing Codon Redundancy and Screening Effort of Combinatorial Protein Libraries Created by Saturation Mutagenesis. ACS Synth. Biol. 2, 83-92 (2013).79. Jones, K. A., Snodgrass, H. M., Belsare, K., Dickinson, B. C. & Lewis, J. C. Phage-Aty. Dkt. No. 125141.04976 MGH2024-443Assisted Continuous Evolution and Selection of Enzymes for Chemical Synthesis. ACS Cent. Set.7, 1581-1590 (2021).80. Zinkus-Boltz, J., DeValk, C. & Dickinson, B. C. A Phage- Assisted Continuous Selection Approach for Deep Mutational Scanning of Protein-Protein Interactions. ACS Chem. Biol. 14, 2757-2767 (2019).81. Miller, S. M. et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs. Nat. Biotechnol. 38, 471-481 (2020).82. Hu, J. H. et al. Evolved Cas9 variants with broad PAM compatibility and high DNA specificity. Nature 556, 57-63 (2018).83. Blum, T. R. et al. Phage-assisted evolution of botulinum neurotoxin proteases with reprogrammed specificity. Science 371, 803-810 (2021).84. Packer, M. S., Rees, H. A. & Liu, D. R. Phage-assisted continuous evolution of proteases with altered substrate specificity. Nat. Commun. 8, 956 (2017).85. Dickinson, B. C., Packer, M. S., Badran, A. H. & Liu, D. R. A system for the continuous directed evolution of proteases rapidly reveals drug-resistance mutations. Nat. Commun. 5, 5352 (2014).86. Carlson, J. C., Badran, A. H., Guggiana-Nilo, D. A. & Liu, D. R. Negative selection and stringency modulation in phage-assisted continuous evolution. Nat. Chem. Biol. 10, 216-222 (2014).87. Wang, T., Badran, A. H., Huang, T. P. & Liu, D. R. Continuous directed evolution of proteins with improved soluble expression. Nat. Chem. Biol. 14, 972-980 (2018).88. Thuronyi, B. W. et al. Continuous evolution of base editors with expanded target compatibility and improved activity. Nat. Biotechnol. 37, 1070-1079 (2019).89. Richter, M. F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol. 38, 883-891 (2020).90. Bryson, D. I. et al. Continuous directed evolution of aminoacyl-tRNA synthetases. Nat. Chem. Biol. 13, 1253-1260 (2017).91. Badran, A. H. et al. Continuous evolution of Bacillus thuringiensis toxins overcomes insect resistance. Nature 533, 58-63 (2016).92. Johnston, C. W ., Badran, A. H. & Collins, J. J. Continuous bioactivity-dependent evolution of an antibiotic biosynthetic pathway. Nat. Commun. 11, 4202 (2020).Aty. Dkt. No. 125141.04976 MGH2024-44393. Pandey, S. et al. Efficient site-specific integration of large genes in mammalian cells via continuously evolved recombinases and prime editing. Nat. Biomed. Eng. 1-18 (2024) doi : 10.1038 / s41551 -024-01227- 1.94. Mercer, J. A. M. et al. Continuous evolution of compact protein degradation tags regulated by selective molecular glues. Science 383, eadk4422 (2024).
[0232] Materials and Methods
[0233] Preparation and transformation of electrocompetent cells
[0234] To prepare electrocompetent cells of s2060 or sl030 derived strains, overnight cultures were grown from single colonies or glycerol stocks in 2xYT media (United States Biologicals) supplemented with appropriate antibiotics at 37C with shaking at 250rpm. Cells were back diluted 50-fold into 25mL 2xYT media supplemented with appropriate antibiotics and grown at 37C with shaking at 250 rpm until reaching OD600= 0.3-0.4, then pelleted by centrifugation at 4000 g for 10 minutes at 4C. Following centrifugation, the supernatant was discarded, and cells were resuspended in 2mL ice-cold 10% glycerol and pelleted at 4000 g for 10 minutes at 4C. This wash step was repeated for a total of three times maintaining cells on ice before spins before cells were finally resuspended in ImL ice-cold 10% glycerol. 50uL of electrocompetent cells were aliquoted into ImL culture tubes. Aliquoted cells were either electroporated immediately (see below) or flash-frozen for 15 seconds in liquid nitrogen for indefinite storage at -80C.
[0235] To transform electrocompetent s2060 or si 030 derived strains, 50uL electrocompetent cells were thawed on ice (if flash frozen and maintained at 80C) for 15 minutes prior to mixing with up to lOOng plasmid DNA. Cells + plasmid were mixed via gentile pipetting and left to equilibrate for 10 minutes before transfer to pre-chilled cuvettes (Bio-Rad). Cells were electroporated per manufacturer's instruction and recovered for 1 hour at 37C in 500mL SOC media (new England biolabs) with shaking at 250 rpm before being spun down at 10,000g for 1 minute. Cells were resuspended in 200uL SOC media and streaked onto 2xTY media +1% agar plates with appropriate antibiotics prior to incubation at 37C for 12-24 hours.
[0236] Genomic engineering of E. coli strains
[0237] All experiments were carried out in E. coli strains derived from s2060. This strain was engineered via a) Lambda Red recombineering1for large chromosomal deletions or b) recombitron-mediated in vivo recombineering for minor chromosomal point substitutions.2For large chromosomal deletions (RecJ and sbcB), s2060 cells were transformed with the pDK46Aty. Dkt. No. 125141.04976 MGH2024-443plasmid as described above and individual colonies were grown overnight in 2xYT media supplemented with Streptomycin, Tetracycline, and Carbenicillin at 30C. Cells were back diluted 1:50 in 2xYT media supplemented with antibiotics and lOmM arabinose, grown to OD=0.3-0.4, and made electrocompetent as described above. Primers MY0270 and MY0271 with 5’ homology to regions flanking Red locus in the E. coli genome were used to amplify the Kanamycin resistance cassette on pKD13. Primers MY0272 and MY0273 with 5’ homology to regions flanking the sbcB locus in the E. coli genome were used to amplify the Kanamycin resistance cassette on pKD13. The PCR product was gel-purified and transformed into 500uL of electrocompetent cells + pDK46 and recovered overnight at 37C in 6mL of SOC media with shaking at 250rpm before plating on 2xYT +1% agar + Streptomycin, Tetracycline, Kanamycin and incubation at 37C for 20 hours.
[0238] Insertion of the Kanamycin resistance cassette was confirmed via colony PCR with primers MY0278 and MY0279 (RecJ) or primers MY0282 and MY0283 (sbcB). To remove the Kanamycin resistance cassette, successful colonies were grown overnight in 2xYT + Streptomycin, Tetracycline, Kanamycin at 37C with shaking at 250rpm and made electrocompetent as described above. Cells were then transformed with the pCP20 plasmid3 and plated overnight on 2xTY+ 1% agarose plates supplemented with Streptomycin, Tetracycline, and Carbenicillin at 30C. Single colonies were grown overnight in 2xTY media supplemented with Streptomycin, Tetracycline, and Carbenicillin at 30C and shaking at 250rpm, then confluent cultures were heat shocked at 43C for 15 minutes to induce expression of the FLP recombinase. Cells were then plated at 37C on 2xYT+ 1% agarose plates supplemented with Streptomycin and Tetracycline to cure the pCP20 plasmid. Removal of the Kanamycin resistance cassette was confirmed via PCR as described above and cells were confirmed to be Kanamycin non-resistant and Carbenicillin non-resistant. Successful colonies were preserved as glycerol stocks and used for experiments or subsequent rounds of genomic engineering.
[0239] For minor point mutations (DNAG Q576A), a retron recombitron2was cloned with an 80bp cDNA template region encoding the DNAG Q576A mutation together with a silent mutation to assist in evasion of MMR (acgcatggttaagcaacgaagaacgcctggagctctggacattaaacGCggaActggcgaaaaagtgatttaacggcttaagtgccg ) (SEQ ID NO: 64). Electrocompetent cells were prepared as described above and transformed with the Chloramphenicol-resistant recombitron plasmid. Cells were outgrown in 500mL SOC forAty. Dkt. No. 125141.04976 MGH2024-443one hour at 37C with shaking at 250 rpm before plating on 2xYT+ 1% agarose plates supplemented with Streptomycin, Tetracycline, and Chloramphenicol and lOOmM glucose followed by incubation at 37C. A single colony was picked and grown overnight in 2xYT media supplemented with antibiotics and lOOmM glucose at 37C with shaking at 250rpm. Cells were back-diluted 1:50 in 2xYT media supplemented with antibiotics and grown at 37C with shaking at 250rpm to OD=0.3-0.4 before induction with lOmM arabinose. Cells were grown for 18 hours at 37C with shaking at 250rpm, then plated on 2xYT + 1% agarose plates supplemented with Streptomycin and Tetracycline. The DNAG Q576A mutation was confirmed via colony PCR with primers MY0417 and MY0418. Successful colonies were streaked out on 2xYT + 1% agarose plates supplemented with Streptomycin and Tetracycline over 2-3 days, then streaked out on both 2xYT + 1% agarose plates supplemented with Streptomycin and Tetracycline or 2xYT + 1% agarose plates supplemented with Chloramphenicol to confirm loss of the recombitron plasmid. Successful Chloramphenicol-non-resistant colonies were preserved as glycerol stocks and used for experiments or subsequent rounds of genomic engineering.
[0240] Intron plasmid assays
[0241] Single colonies of S2060 cells transformed with the intron plasmid (IP) were grown overnight in 2xYT + maintenance antibiotics at 37C with shaking at 250rpm and back diluted the following day 1 :50 in Davis Rich Media (DRM). Cells were grown at 37C with shaking at 250rpm until reaching OD=0.3-0.4, then induced with lOmM arabinose. After the allotted induction time (1, 2, 5 or 20 hours), cells were spun down and lysed via incubation at 95C for 10 minutes. NGS primers MY0029 and MY0032 were used to amplify cDNA (NGS round 1, 18 cycles) before barcoding with universal barcode primers (NGS round 2, 22 cycles). 300bp cDNA bands were visualized on an agarose gel as in FIG. IB, purified and loaded on a Miseq (Illumina) with a read density of approximately 100,000 reads / sample.
[0242] E. coli chromosomal editing and rifamycin resistance assays
[0243] s2060-derived strains were transformed with a reticulase plasmid (RP) and plated on 2xYT media 1% agarose plates supplemented with maintenance antibiotics and lOOmM glucose. Maintenance of cells containing the RP on glucose is essential to attenuate metabolic burden or random mutagenesis associated with leaky expression of the ssap and MutL* off the RP. Single colonies were picked into 2xYT media + maintenance antibiotics and lOOmM glucose and grown overnight at 37C with shaking at 250rpm. Cells were then back-diluted the following day 1:50 inAty. Dkt. No. 125141.04976 MGH2024-443Davis Rich Media (DRM)4+ maintenance antibiotics (no glucose), grown at 37C with shaking at 250rpm until reaching OD=0.3-0.4, then induced with arabinose and incubated for 20 hours at 37C with shaking at 250rpm. For NGS sequencing of edited loci, cells were spun down and lysed as described above. For NGS sequencing of the rpoB locus, primers MY0371 and MY0372 were used. For NGS sequencing of the gyrA locus, primers MY0519 and MY0520 were used. For rifamycin resistance assays, cells post-20 hour induction were serial diluted (1:10 dilutions) in 2xYT media. 30uL of each cell dilution was plated on either 2xYT +1% agarose plates or 2xYT +1% agarose plates supplemented with 25ug / mL rifamycin, then left to outgrow for 18 hours at 37C. To calculate the % cfu resistant, cells were counted on each of the plates and the cfu / mL on the + rifamycin plates was divided by the cfu / mL on the - rifamycin plates. Sanger sequencing of rifamycin-resistant colonies was performed via colony PCR with primers MY0280 and MY0281.
[0244] Phage production and titering
[0245] To produce phagemid particles from phagemid plasmids, s2060-derived strains were transformed with the helper plasmid pMY0045 and made electrocompetent as described above to create an electrocompetent ‘helper strain’. 50-100ng of phagemid plasmid was transformed into the helper strain, followed by outgrowth in 500mL SOC for one hour at 37C with shaking at 250rpm. This 500mL culture was used to directly seed 7mL 2xYT media supplemented with streptomycin, tetracycline, kanamycin (resistance marker of the helper plasmid) and carbenicillin (resistance marker of the phagemid plasmid), before incubation for 20 hours at 37C with shaking at 250rpm. The following day, the helper strain culture was spun down and phagemid were harvested from the supernatant and sterile filtered using a 20pm CA syringe filter (Celltreat).
[0246] To titer phage, a series of 1 : 100 serial dilutions of phagemid harvested from the supernatant as described above were prepared in 2xYT media. Typically, four serial dilutions of 1:100 (10A-2), 1:10,000 (10A-4), 1:1,000,000 (10A-6), and 1:100,000,000 (10A-8), were prepared. An overnight culture of s2060 was back-diluted 1:50 in 2xYT media supplemented with maintenance antibiotics (streptomycin and tetracycline) and grown at 37C with shaking at 250rpm until reaching OD=0.3-0.4. 90uL portions of this s2060 cell culture were then each infected with lOuL of serial phagemid dilution, resulting in an additional 1:10 dilution of phagemid. The infected cell culture was incubated at 37C for one hour with no shaking before 30uL of the phage / cell mix was plated onto a pre-warmed 2xYT +1% agarose 24 well plate supplemented with Carbenicillin. Plates were incubated for 18 hours at 37C before titer counting the following day.Aty. Dkt. No. 125141.04976 MGH2024-443
[0247] To calculate phagemid titer, Carbenicillin-resistant colonies on the titering plates described above were counted at the dilution factor at which single colonies could be readily distinguished; typically, no more than 100 colonies would be counted at a given dilution factor. For the protocol described above, the titer was determined using the following calculation:, 1 cfu mL1= colonies counted * — - - - - — - - - - - * 333dilution factor at which colonies were counted
[0248] The set multiplier 333 accounts for the 1:10 dilution of lOuL phage in 90uL cells (an additional lOx multiplier), and the correction to estimate cfu / mL; 30uL of phage / cell mix plated, 30uL= 1 / (33. ) ImL (an additional 33.3x multiplier, rounded down): 10*33.3= 333.
[0249] Where appropriate, titering data was used to prepare stocks of phage at normalized working concentrations (typically l*10A6 cfu / mL or l*10A8 cfu / mL). Normalized stocks were aliquoted into 1.5mL eppendorf tubes and stored indefinitely at -20C.
[0250] The set multiplier 333 accounts for the 1:10 dilution of lOuL phagemid in 90uL cells (an additional lOx multiplier), and the correction to estimate cfu / mL; 30uL of phagemid / cell mix plated, 30uL= 1 / (33.3) ImL (an additional 33.3x multiplier, rounded down): 10*33.3= 333.
[0251] Where appropriate, titering data was used to prepare stocks of phage at normalized working concentrations (typically l*10A6 cfu / mL or l*10A8 cfu / mL). Normalized stocks were aliquoted into 1.5mL eppendorf tubes and stored indefinitely at -20C.
[0252] Reticulase-mediated phagemid editing
[0253] To edit phagemid populations using the reticulase, single colonies of a helper strain (s2060 derived cell line containing a helper plasmid (HP)) transformed with a reticulase plasmid (RP) were grown in 2xYT media + maintenance antibiotics and lOOmM glucose overnight at 37C with shaking at 250rpm. Cells were back-diluted the following day 1:50 in Davis Rich Media (DRM) + maintenance antibiotics (no glucose), grown at 37C with shaking at 250rpm until reaching OD=0.3-0.4, then induced with arabinose and infected with phagemid; unless otherwise noted, lOmM (1% IM) arabinose was used for RP induction, and phage were infected via a 1:100 fold dilution of a l*10A6 cfu / mL stock in the cell culture, resulting in a final starting tier of l*10A4 cfu / mL. Cells were incubated for 10 minutes at 37C without shaking to allow infection, then incubated for 18-20 hours at 37C with shaking at 250rpm. Cells were then spun down and phagemid was harvested from the supernatant. Harvested phagemid could then be titered as described above. Where appropriate, harvested phagemid could be used to infect a subsequentAty. Dkt. No. 125141.04976 MGH2024-443passage of s2060-derived cells transformed with the HP and RP, increasing the population of edited phage by editing over multiple serial passages. Typically, phage from the previous passage was diluted 1:1000 in the fresh batch of s2060-derived cells to start the subsequent passage. For NGS of phage populations, 20uL of phagemid harvest was transferred to a 96 well plate and heat lysed via incubation at 95C for 10 minutes. This lysate was then used as the template for NGS. Primers MY0452 and MY0453 were used for NGS of the clopt gene, and primers MY0729 and MY0730 were used for NGS of the Trastuzumab light chain CDR3 region. Total edited phage for any given sample was calculated by multiplying the fraction of edited phage by the corresponding phage titer.
[0254] Targeted Reticulase-Assisted Continuous Evolution (TRACE) of clopt
[0255] S2061 cells (s2060 ARecJ) transformed with the helper plasmid (HP), accessory plasmid (AP) and reticulase plasmid (RP) were plated on 2xYT media + 1% agarose plates supplemented with maintenance antibiotics and lOOmM glucose. Single colonies of these TRACE host cells were grown in 2xYT media + maintenance antibiotics and lOOmM glucose overnight at 37C with shaking at 250rpm. Cells were back-diluted the following day 1:50 in Davis Rich Media (DRM) + maintenance antibiotics (no glucose), grown at 30C with shaking at 250rpm until reaching OD=0.3-0.4, then induced with arabinose and infected with phagemid; lOmM (1% IM) arabinose was used for RP induction, and phagemid were infected via a 1:100 fold dilution of a l*10A8 cfu / mL stock into the cell culture, resulting in a final starting tier of l*10A6 cfu / mL. Cells were incubated for 10 minutes at 30C without shaking to allow infection, then incubated for 18-20 hours at 30C with shaking at 250rpm. Phagemid were harvested from the supernatant and used to infect a subsequent batch of TRACE host cells grown to OD=0.3-0.4 as described above. Each passage was initiated with a 1:1000 dilution of phagemid from the previous passage in a fresh batch of TRACE host cells. The evolution of clopt was performed at 30C because the activity of this protein is known to be thermoregulated: clopt is folded at 30C and denatured at 37C. For a typical protein with Tm>37C, TRACE should be performed at 37C to maximize M13 phage infectivity. The harvested phage of each passage was titered and subject to NGS as described above. Single Carbenicillin-resistant colonies grown on titer plates were sanger sequenced with primers MY0373 and MY0374 to determine genotypes of individual phage selected by the conditional M13 phage propagation circuit. For full phagemid sequencing, Carbenicillin-resistant colonies from the titer plates were picked into 5mL 2xYT supplemented with Carbenicillin and grown overnight at 37CAtty. Dkt. No. 125141.04976 MGH2024-443with shaking at 250rpm. The following day cell cultures were miniprepped (Qiagen) and sent for full-plasmid sequencing.
[0256] Table 1. Relevant PrimersMY0270 AAAGAATTCCTCGACGAACACCAAA F primer for pKD13 knockout vector for AAATGACCAGCGGTAAATAATTCGC RecJ locus ATTCCGGGGATCCGTCGACC MY0271 AAGAGCGGGATTGTACCCAATCCAC R primer for pKD13 knockout vector for GCTCTTTTTTATAGAGAAGATGACGT RecJ locus GTAGGCTGGAGCTGCTTCG MY0272 CAAATTGTGGCGCTAAAGCTGATTA F primer for pKD13 knockout vector for GCACGGTGATATTTGATACTCTGGCA SbcB locus TTCCGGGGATCCGTCGACC MY0273 CCTGACTCAACATTGTCCTCCGCCGT R primer for pKD13 knockout vector for ACCAGCGGCGGAGGCTTCAAATTAT SbcB locus GTAGGCTGGAGCTGCTTCG MY0278 AATGGCACACTTGTTCCGGGTT RecJ F sequencing primerMY0279 GGTCTGATTTCTTTTATTGAGCTAGT RecJ R sequencing primerC MY0282 TGTGACTACCCTCTCATTTTTATCTG sbcB F sequencing primerAC MY0283 AAGAAGGACATGCCATCACCGA sbcB R sequencing primer MY0417 CTGTCGATGTGGGACGATATAG F primer dnaG sequencng primer MY0418 CAACGATAATTACGAGGGCGTAC R primer dnaG sequencng primer MY0280 ATCCGTTCCGTTGGCGAAATGG rpoB F sequencing primer MY0281 CAGACAGGTAGTGAATTTCGTCAG rpoB R sequencing primer MY0373 GGTAGAAGTTTGCGACGTTTTAGCA pMY0057 / 58 F sequencing primer MY0374 GGTGAGAACATCCCTGCCTGAA pMY0057 / 58 R sequencing primer MY0029 ACACTCTTTCCCTACACGACGCTCTT F sequencing primer 1 BPP1 DGR 5prime CCGATCTNNNNCGCTGCTGCGCTATT homologyCGG MY0032 TGGAGTTCAGACGTGTGCTCTTCCGA R sequencing primer 1 BPP1 DGR GC TCT cagacgccgcgcgc terminusMY0371 ACACTCTTTCCCTACACGACGCTCTT rpoB F alternate sequencing primer CCGATCTNNNNGAGTTCTTCGGTTCC AGCCAG MY0372 TGGAGTTCAGACGTGTGCTCTTCCGA rpoB R alternate sequencing primerTCTgagtcgggtgtacgtctcgaaAtty. Dkt. No. 125141.04976 MGH2024-443MY0519 ACACTCTTTCCCTACACGACGCTCTT gyrA NGS F primer CCGATCTNNNNGCAATGACTGGAAC AAAGCCTATAAA MY0520 TGGAGTTCAGACGTGTGCTCTTCCGA gyrA NGS R primer correct TCTCGAAGTTACCCTGACCGTCT MY0452 ACACTCTTTCCCTACACGACGCTCTT clopt NGS F primer correct CCGATCTNNNNGACGCACGTCGCCT TAAAGCAATTT MY0453 TGGAGTTCAGACGTGTGCTCTTCCGA clopt NGS R primer correct TCTCAACGCTAACCTTGAGAATCTTG GC MY0729 ACACTCTTTCCCTACACGACGCTCTT Trastuzumab NGS F primer CCGATCTNNNNACTTCACGCTTACGA TTTCAAGCC MY0730 TGGAGTTCAGACGTGTGCTCTTCCGA Trastuzumab NGS R primerTCTAGACCCTCCGCCGCCTGAA
[0257] Materials and Methods References1. Datsenko, K. A. & Wanner, B. L. One-step inactivation of chromosomal genes in Escherichia coli K-12 using PCR products. Proc. Natl. Acad. Set. U. S. A. 97, 6640-6645 (2000).2. Schubert, M. G. et al. High-throughput functional variant screens via in vivo production of single-stranded DNA. Proc. Natl. Acad. Set. 118, e2018181118 (2021).3. Cherepanov, P. P. & Wackemagel, W. Gene disruption in Escherichia coli. TcR and KmR cassettes with the option of Flp-catalyzed excision of the antibiotic-resistance determinant. Gene 158, 9-14 (1995).4. Miller, S. M., Wang, T. & Liu, D. R. Phage-assisted continuous and non-continuous evolution. Nat. Protoc. 15, 4101-4127 (2020).
[0258] Example 2. Construction of Novel Reticulases from Metagenomic-Mined Diversity Generating Retroelement Components
[0259] I: Diversity Generating Retroelements in Natural Continuous Evolution
[0260] Diversity-Generating Retroelements (DGRs) are a unique family of retroelements with evolutionary relationships to retrons and group II introns that generate continuous, hyper-directed sequence variation in the genes they target.5DGRs utilize a unique mechanism of site-specific retrotransposition in which sequence variants are inserted into a flexible coding scaffold avoiding non-specific variation in conserved genomic regions.25The necessary components of a DGR areAty. Dkt. No. 125141.04976 MGH2024-443typically localized in a single genomic locus that spans 5-1 Okb, however, the synteny and organization of these components can vary.14Mechanistically, diversification is facilitated by a reverse transcriptase that acts on a non-coding RNA transcribed from a template repeat (TR) region in the DGR locus. The template repeat region is almost identical to a variable region (VR) often found in a nearby gene, which encodes the DGR-variable protein (VP). The reverse transcription of the RNA intermediate into cDNA is carried out by the error-prone RT which favors A to N mutations.19The mutated cDNA then replaces the VR, which often encodes a sequence that corresponds to the flexible residues in a ligand-binding structural domain belonging to the C-type lectin protein family.101The TR is invariable, and transposition occurs in a unidirectional manner from TR to VR, generating continuous sequence diversity at positions in the VR that contain adenine at the corresponding positions in the TR.
[0261] The variation created by DGRs reaches at least IO20possible sequences, a scale of diversity far outstripping that generated by somatic hypermutation in vertebrate adaptive immune systems-which reach around 1014possible sequences in antibodies.36Importantly, this 106greater magnitude of diversity generated by DGRs is achieved through the hypermutation of 103shorter DNA sequences than those sequences diversified by vertebrate adaptive immune systems.36 41 70A DGR protein with IO20possible sequence variations has been structurally characterized,106and a DGR protein with IO30possible sequence variants - eight fold more variants than the number of stars in the visible universe - was recently identified in metagenomic analyses.11DGRs are currently the only known source of immense protein-encoding sequence variation in prokaryotes, and produce the greatest level of protein sequence plasticity in the natural world.32 34 35 2
[0262] The first DGR variable protein was discovered in 2002 in the Bordetella Positive Phase bacteriophage 1 (BPP-1), which infects the Bordetella bacterial pathogen that causes whooping cough in humans.108 110The BPP-1 DGR diversifies 12 amino acids in the tail fiber tip protein (Mtd) that bind the Bordetella host receptor; this DGR is theoretically capable of generating nearly 10 trillion protein variants of Mtd differing only in these 12 amino acid positions.106 25Surface plasticity in Bordetella cell surface proteins during the infectious cycle of Bordetella creates a natural advantage for this DGR-mediated continuous evolution of the phage tail fiber. The Bordetella adhesin pertactin is expressed on the cell surface during the in vivo / pathogenic Bvg+phase of Bordetella and can serve as the receptor for Mtd; during the ex vivo Bvg phase, cellsurface expression of pertactin is lost.110 101DGR-mediated tropism switching enables the BPP-1Aty. Dkt. No. 125141.04976 MGH2024-443DGR to produce Mtd variants capable of binding cell-surface receptors expressed by the Bvg Bordetella, and thereby continue its infectious cycle.110 101 70The genomic locus that enabled the BPP-1 to adaptively alter its tropism was termed a ‘tropism switching cassette’ in the initial publication that reported its discovery. While this term is still used to describe this genetic locus in BPP-1 specifically, the term ‘Diversity Generating Retroelement’ has become the universal term used to describe this class of genetic elements; DGR refers to the entire cassette and associated genes, not only the reverse transcriptase that facilitates the mutagenic cDNA synthesis. The mechanism whereby a DGR facilitates mutagenesis of a target gene through mutagenic cDNA synthesis and subsequent integration into a target locus is termed ‘mutagenic retrohoming.’ This phrase does not refer to the capacity of a virus to ‘hone in’ on a novel receptor on the surface of the cell using a DGR.
[0263] Shortly after their initial discovery in phage, cellular DGRs were identified and characterized in bacterial pathogens Legionella pneumophila and Treponema denticola, in these bacteria, mutagenic diversification is targeted at cell surface proteins believed to be implicated in cell-cell attachment.8692The discovery of DGRs in the context of viral-cell or cell-cell interactions led them to be initially categorized as genomic elements primarily involved in host pathogenesis or symbiosis. However, the critical components of DGRs have been found in metagenomic studies across many lineages of prokaryotic life, which may imply a broader range of utility than the diversification of ligand-binding domains associated with cellular attachment.146388 11For instance, lab isolates of cyanobacteria Nodularia spumigena CCY9414 and Trichodesmium erythraeum IMS101 appear to have high expression levels of the DGR RNA intermediate when subjected to light and oxidative stress, suggesting that mutagenic diversification may play a role in environmental adaptation.82 78 85
[0264] A recent bioinformatics review identified six major evolutionary lineages of DGR-RTs across all prokaryotic domains of life and viruses-three lineages primarily associated with viruses-that share a common ancestor; this review suggested that DGRs may be responsible for >10% of amino acid changes in some organisms. Out of 36,611 genes suspected to be DGR target variable proteins in publicly available genomes and metagenomes, all were multi-domain proteins with the variable region in the C-terminus; this finding was consistent with what appears to be the universal mechanism of DGR function involving C-terminal retrohoming mediated by indispensable VR intergenic sequences downstream of the retrotransposon insertion site. Intriguingly, while N-termAty. Dkt. No. 125141.04976 MGH2024-443domains of variable proteins exhibit a variety of folds, the C-term domains were systematically associated with members of the C-type lectin fold family, and the invariable residues within these C-type lectin folds were highly conserved across variable proteins from all domains of life and viruses. This finding indicates that the C-type lectin fold family is uniquely optimized as a scaffold capable of accommodating immense amino acid variation in solvent-exposed residues that can be appended to myriad protein architectures with different functions.63
[0265] II: Components of the BPP-1 PGR and Their Phylogenic Relationships to Other Retroelements
[0266] ID vivo evidence has shown that no BPP-1 phage products other than those encoded within the BPP-1 DGR are required for BPP-1 DGR function. All other proteins and RNA elements presumed to be required for BPP-1 DGR mutagenic retrotransposition (such as DNA ligases, polymerases, etc.) must originate from the Bordetella host itself and have yet to be identified.94The BPP-1 DGR encodes three proteins: the BPP-1 tail-fiber (Mtd), the BPP-1 reverse transcriptase (bRT), and the accessory variability determinant protein (Avd). To date the only the BPP-I DGR has been mechanistically and biochemically studied, and much of our understanding of DGR biology stems from this model system. While many DGR components appear generally conserved, one must keep in mind there may be undiscovered variation (additional accessory proteins, RNA components etc.) that differentiate some natural DGRs from this BPP-1 model system.
[0267] III: Structure and Association of bRT and Avd
[0268] DGR-RTs belong to their own distinct clade of RTs that are evolutionarily most closely related to the reverse transcriptase domain of group II intron maturases.9089The bRT structure was solved in complex with Avd and the DGR-RNA to 3.1 A resolution. Its structure closely resembles the reverse transcriptase domain of group II intron maturases and does not contain any accessory RNAse domain.76 75 79 72 2The bRT exhibits an N-terminal extension subdomain (aa 1-59) which it shares in common with non-Long Terminal Repeat (non-LTR) reverse transcriptases, followed by canonical finger (aa 60-102), palm (aa 103-272) and thumb (aa 273-328) subdomains that are related to those of the HIV-1 RT.23 2The central active site motif of the bRT is YADD, which has been replaced with SMAA to create an inactive mutant for negative controls.94Sequence comparison to the HIV-RT was able to identify D138 as a third active-site aspartic acid (together with D215 and D216 noted above in YADD), and a D138Q mutation has been used to create aAty. Dkt. No. 125141.04976 MGH2024-443bRT mutant with one-fifth WT catalytic activity.110Overall low catalytic efficiency and especially low catalytic efficiency on template adenines are intrinsic properties of the bRT and are responsible for the hallmark A-specific mutational profile of the BPP-1 DGR.65
[0269] The Avd protein is 128 amino acids, and its structure was solved in isolation from bRT to 2.69A resolution. Avd forms a pentameric, highly positively charged barrel.87Hydrophobic interprotomer interfaces stabilize the quaternary structure of the Avd pentamer. Many DGRs are predicted to encode similar Avd-like proteins that are also predicted to form pentameric barrellike assemblages; sequence alignment indicates these proteins have a shared relationship with S23 ribosomal and S23 ribosomal-like proteins.108The structure of the bRT-Avd complex confirms that the 3’UTR of the DGR-RNA sits on top of the narrow end of the Avd pore and makes myriad specific base pair- amino acids contacts with the Avd pentamer. The structure of Avd free and bound to bRT is essentially identical (1.4A rmsd, 456 Ca), which may suggest Avd pentamer formation precedes association with bRT.2The bRT N-terminal subdomain (aa 1-53) makes extensive contact with two of the five Avd protamers, creating a total of 2300A buried surface area of mainly electrostatic interface. One of the Avd pentamers nestles its N-terminal 10 amino acids into a crevice in the four-helix bundle of the bRT N-terminal subdomain; in the other four Avd pentamers this N terminal subdomain is disordered.2The entire bRT-Avd complex is required for bRT-mediated reverse transcription of DGR-RNA templates. In vitro reverse transcription of DGR-RNAs using exogenous primers still requires Avd, indicating that Avd has a functional interaction with bRT in addition to its interaction with 3’ spacer-region nucleotides Sp4-30 in the 3’UTR.2Because Avd is positioned some 40A from the bRT active site, it is not believed to play a direct role in the chemistry of reverse transcription.2To our knowledge, no retron - the other primary class of retroelements that cis-prime off a 2’ OH group in their template RNA - has been discovered with dependance on an Avd-like protein. Thus, this Avd-dependance appears to differentiate DGRs-RTs from their close mechanistic relatives.
[0270] II.2: Structure of the DGR-RNA
[0271] The BPP-1 DGR-RNA is 580nt, containing an extensive 5’ 297nt UTR from avd, the central 134nt Template Region (TR), and 150nt 3’ UTR consisting of a 144nt spacer (sp) between TR and brt and the beginning 6nt of brt. The 297nt 5’ UTR of the DGR-RNA is essential for near WT retrohoming ability in the WT DGR. Truncations of this 5’ UTR, especially below 115nt result in 500-1400-fold lower retrohoming activity.87However, almost the entirety of this 5’ UTR is notAty. Dkt. No. 125141.04976 MGH2024-443required for cDNA synthesis. A truncated version of the 5’ UTR with only the final 20 nt of avd has been shown to produce 7-fold higher quantities of full-length cDNA in vitro than the full-length DGR-RNA. These results indicate that mutagenic cDNA synthesis and retrohoming can be decoupled by the removal of the first 277 nt of the 5’ UTR, producing a truncated RNA (referred to in the literature as the 294 nt ‘core DGR-RNA’) that can facilitate high-efficiency synthesis of adenine-mutagenized cDNA but not its homing to the endogenous VR locus. The reticulase uses the truncated 5’ and 3’UTRs of this ‘core DGR-RNA’. Complementarity between bases Avd 369-379 in the 5’UTR (part of the 20bp 5’UTR present in the core DGR-RNA) and Sp 59-70 in the 3’ UTR create a steric block that arrests reverse transcription just after the final base pairs of the TR. This RNA-RNA homoduplex thereby prevents reverse transcription into the 5’UTR (cPRT stop).2
[0272] The 3’ TR-Z>r / spacer region is essential for cDNA synthesis either from the cis-primed DGR-RNA template itself or from an exogenous primer. This region provides an essential binding site for Avd (sp4-30) as well as the nucleotide (Sp56A) from which cis-priming is initiated.7470In the synthesis-competent / retrohoming incompetent ‘core DGR-RNA,’ a minor truncation of lOnt at the 3’ end of the spacer region excludes the final 4nt of sp and the 6nt of brt, as a result, the first 140nt of the sp region are included in the core DGR-RNA. Sequence homology between 5’-TR118-125 and Sp49-56 (except for TR120C and sp54U) position the Sp56A priming nucleophile for reverse transcription. Following the Avd-binding stem loop, base pairs Sp59-70 bind the 5’UTR and assist in terminating reverse transcription, while Sp71-78 form a stem- tetraloop structure of the highly abundant UNCG class. This stem tetraloop appears to assist in creating a physical block to polymerization together with the Sp71-78 region; deletion of the stem-tetraloop results in reverse transcription at a level of 2% WT, indicating that this motif may also play a role in the proper positioning of the bRT rather than only in terminating reverse transcription. A large single-stranded portion of Sp (Sp82-94) wraps around the thumb ring of bRT, while a second large hairpin structure capped by a heptaloop (Sp94-140) makes extensive contact with bRT. Disruption of bases in any of these bRT -binding regions of Sp significantly decreases but does not abolish reverse transcriptase activity in vitro. All of these RNA secondary structural elements in the 3’ UTRs of the DGR-RNA appear highly conserved across DGRs identified in metagenomic studies with clear Avd homologs. It has been proposed that these large stem-loops serve as ‘staples’ which direct the initial association of bRT-Avd with the DGR-RNA and serve to maintain this association through multiple rounds of polymerization. Such a staple mechanism could prove essential becauseAty. Dkt. No. 125141.04976 MGH2024-443the bRT and other DGR-RTs lack the a loop found in group II intron RTs which promotes processive polymerization.2
[0273] III: The Mechanism of the Mutagenic Retrohoming Is Believed to Be Broadly Conserved Across DGRs
[0274] Because the DGRs of bacteria, archaea, and their respective viruses appear structurally similar and have related RTs and signatures of adenine-specific diversification, it has been hypothesized that the fundamental mechanism of mutagenic retrohoming epitomized by the BPP-1 DGR is broadly conserved among DGRs across the prokaryotic domains of life.73 74At present, the mechanism of the BPP-1 DGR is by far the best characterized in the literature. The present review has divided the BPP-1 DGR mechanism into three parts: Initiation of single stranded cDNA synthesis, cDNA synthesis / adenine mutagenesis, and homing / integration into VR with appropriate sub-divisions. While only the initiation and single- stranded cDNA synthesis steps are relevant to the function of the reticulase in TRACE, the full mechanism is provided here.
[0275] IIL1: Initiation of ss cDNA synthesis
[0276] In several characterized retrotransposons, reverse transcription is facilitated by a nick in the target locus that creates a 3’ DNA end to serve as the priming nucleophile for cDNA synthesis. This mechanism-referred to as Target Primed Reverse Transcription (or ‘TPRT’)- couples cDNA synthesis with retrohoming by requiring the localization of the RNA / RT complex to a 3’ DNA end in the target locus. TPRT has been characterized in the R2 element in Bombyx mori (B2Bm) and several non-Long Terminal Repeat retrotransposons (also referred to as non-LTR retrotransposons, LINEs, polyA retrotransposons, and target primed retrotransposons).62 103 97 112 125Unlike in these systems, bRT cDNA synthesis does not require the presence of the target VR for initiation and is instead cis-primed by the DGR-RNA itself (a priming mechanism similar to that implemented by bacterial retrons)74Reverse transcription to produce TR-cDNA has been performed in vitro and requires only the bRT, the DGR-RNA, and the Avd protein.110 7487In vitro evidence indicates that the 2’-OH of Sp56A provides the priming nucleophile from which the first cDNA base is added by bRT, resulting in branched RNA-cDNA molecule linked by a 2’-5’ phosphodiester bond. A synthetic DGR-RNA molecule with deoxyadenosine at Sp56A abolished cDNA transcription activity relative to a synthetic full RNA control, demonstrating the requirement for a 2’ -OH at this position for priming. This mechanism of internal priming off of a 2’OH has been characterized in retrons, in which a specific internal Guanosine (as opposed to adenine) in an RNA primes cDNAAty. Dkt. No. 125141.04976 MGH2024-443synthesis by the retron reverse transcriptase, creating a branched RNA-cDNA molecule linked by a 2’-5’ phosphodiester bond.128 121 81 126Importantly, this cis-priming mechanism is distinct from the TPRT mechanism implemented by group II introns, which are predicted to share the closest evolutionary relationship to DGRs.119 118 108 107An exogenous primer can be used to initiate reverse transcription off of a non-DGR-RNA template, however only short 5-35nt cDNAs are synthesized, indicating that template-priming and processive polymerization are properties specific to the DGR RNA. These short reverse transcripts initiated with exogenous primers still display adeninespecific mutagenesis, indicating that adenine mutagenesis is an intrinsic property of the bRT-Avd complex that operates independently of the mechanism of priming or the RNA template.2
[0277] The ability of the BPP-1 DGR-RNA to cis-prime is suspected to be indicative of a mechanism broadly employed in DGRs to expedite their horizontal gene transfer between bacteria and phages by consolidating the components required for functionality. Such consolidation would enable the DGR cassette to transfer horizontally as a single, functional genetic element. For instance, there is evidence of extensive horizontal gene transfer of a unique clade of DGRs in cyanobacteria suspected to play a role in diversifying signaling kinases.9The requirement of the BPP-1 DGR-RNA for cv.s-priming any reverse transcript longer than 35nt may also play a role in protecting the host genome from DNA damage by greatly restricting the chance of off-target reverse transcription. Genome damage due to retrotransposons such as Ty elements and non-LTR retrotransposons has been observed in which aberrant reverse transcription of cellular mRNAs by these retroelements creates cDNAs that are erroneously incorporated into the genome as pseudogenes through DNA repair pathways.126 116 105
[0278] IIL2: Single-Stranded cDNA Synthesis and Termination
[0279] The substitution rates of bRT are surprisingly even among mispaired bases, with a slight A to T transversion bias: A T (22.2%), A — G (17.8%) and A— ► C (11.6%). The bRT error rates on the other three template bases are comparatively quite low: 0.5 ± 0.2% for T— > N, 0.3 ± 0.2% for C — > N, and 1.6 ± 1.2% for G —> N respectively.65The hallmark motif of DGR biology is the codon AAY (most often AAC). Due to its profile of adenine-specific mutagenesis, bRT can diversify this codon into those encoding for 15 of the 20 canonical amino acids - representing the full gamut of size and chemical properties - and not a stop codon. A total of 10 amino acids can be encoded without adenine present in their codon: Phe, Leu, Vai, Ser, Pro, Ala, Cys, Trp, Arg, Gly. These residues are protected from diversification by bRT and compose the invariable scaffoldAty. Dkt. No. 125141.04976 MGH2024-443that supports the solvent-exposed variable amino acids that confer Mtd’s trophism.
[0280] The mechanistic basis of adenine-specific mutagenesis by bRT-Avd has been shown in in vitro biochemical assays to be a function of low catalytic efficiency of base incorporation by the bRT-Avd complex across template adenines. The catalytic efficiency (kcat / Kni) of the bRT-Avd complex is remarkably low across all template bases and lowest on adenine. Notably, the catalytic efficiency for misincorporated bases across a template adenine is only slightly lower than for correctly incorporated bases at the same position.65Because a hallmark of low fidelity polymerases is generally low catalytic efficiency, this underlying mechanism of adenine mutagenesis is in keeping with characterized related systems.113The error rate of around 50% for the bRT on TR adenine residues is more than 13,000 times the error rate of the HIV type 1 (HIV-1) RT a notoriously low-fidelity RT- which exhibits a misincorporation rate of 1.4-3.0*10"5per cycle per base.122 100
[0281] The C6 position of the purine ring is a critical determinant of the misincorporation rate on a template base itself. The presence of an amine group at C6 (as in adenine) or the absence of any group at this position (as in purine) neither increases nor decreases the misincorporation rate. However, the introduction of a carboxyl group at this site (as in guanine) markedly decreases the misincorporation rate.65The C6 substituent on the template adenine base is not directly contacted by bRT. At present, it is thought that bRT’s A-specific mutagenesis is attributable to the differences in the dipole moments of the template bases. Adenine exhibits a low dipole moment of p 2.5, while guanine, cytosine and uracil have far higher dipole moments of p 6.8, 7.0 and 4.7 respectively. A high dipole moment is expected to exhibit a high energetic cost at the largely hydrophobic template site in bRT, which could be compensated by hydrogen bonding to the incoming respective dNTP to stabilize the bRT active site to promote chemistry; mispaired bases would fail to offset this energetic penalty and dissociate before chemistry could occur. In contrast, only a template adenine (with a low dipole moment) would intrinsically reduce the energetic cost associated with high template site hydrophobicity, allowing mispaired dNTPs to occupy the active site long enough for chemistry to occur.2An adenine nucleobase analog with a methyl group attached to the C6 carbon (which increases overall base hydrophobicity and further diminishes the dipole moment) resulted in a greater misincorporation frequency as measured in vitro, further substantiating this hypothesis.63Because DGR-RTs form a conserved evolutionary clade and all appear to display A-specific mutagenesis (as opposed to T, G or C specific mutagenesis), thisAty. Dkt. No. 125141.04976 MGH2024-443‘template dipole moment dependent infidelity’ is believed to be the universally conserved biochemical mechanism underlying DGR hypermutation. Importantly, our lab has created variants of the bRT with increased hydrophobicity in their template site through rational substitution of key active site residues; these variants display enhanced A>N mutagenesis, as well as nearly 500% increased mutagenesis on template uracil (the base with the second-lowest dipole moment after adenine). These results demonstrate that mechanistic insights can be leveraged to engineer DGR components with desirable biochemical properties.
[0282] Within the 134bp TR region itself, the first 22bp between the 5’UTR and the first AAC motif we refer to as the 5’GCT spacer; at the 3’ end of the TR between the final AAC motif and the ‘Initiation of Mutagenic Homing’* sequence (IMH* discussed below) lies a 14bp stretch of GC repeats we refer to as the 3’GC spacer. While these 5’ and 3’ spacers play roles in the integration of cDNA into the VR (discussed below), our lab has also discovered they play a critical role in cDNA synthesis itself. We have demonstrated that the editing efficiency of a Reticulases with a given THR can vary widely depending on the spacers used; as the most extreme example of this trend, a 2bp 3’ (G / C) spacer enabled 15% editing efficiency on a CDR3 loop locus while a 4bp 3’ (G / C) spacer enabled >50% editing on the same locus (forthcoming data packet). These trends do not appear to generalize between loci, with different length 5’ and 3’ spacers being optimal for each THR. Due to their placement flanking the TR region, we believe these spacers play a role in the initiation and termination of reverse transcription in a template sequencedependent manner; it is likely that they impact the quantity of cDNA produced by the RT, although we have yet to demonstrate this hypothesis biochemically.
[0283] III.3: Retrohoming and Integration of cDNA into the VR Locus
[0284] Unlike cDNA initiation, synthesis, and adenine-specific diversification, cDNA retrohoming and integration for the BPP-1 DGR have not been reconstituted in vitro and are in the process of being mechanistically solved. In vivo evidence strongly indicates that the antisense cDNA produced off the template DGR-RNA is integrated directly into the VR locus rather than serving as a template for HDR. Retrohoming is entirely RecA independent in Bordetella, substantiating a non-HDR mechanism.103
[0285] III.3a: The 3 ’ Intergenic VR Homing Hairpin
[0286] A broadly conserved 3’ intergenic motif found across DGRs consists of an inverted repeat sequence predicted to form a hairpin structure in ssDNA and a cruciform structure in dsDNA.Aty. Dkt. No. 125141.04976 MGH2024-443These hairpin / cruciform motifs are predicted to form 7-1 Obp stems and are found at varying distances downstream of predicted IMH* sequences in many different DGRs. The stems of these hairpins are often GC-rich and do not exhibit sequence conservation between DGRs, however, the loops of these hairpins are generally 4 nucleotides and predominantly exhibit the conserved motif 5’-GRNA.94 108The widely-conserved nature of this 3’ intergenic sequence in DGRs has been taken as evidence that the mechanism VR homing is broadly conserved among this unique family of retroelements. The dependency of DGR homing on the presence of such non-coding 3’ intergenic sequences has also been proposed to explain the observation that out of some 36,000 DGR analyzed in publicly available genomes, all diversified multi-domain proteins with the variable region at the very end of the C-terminus (placing the variable region proximate to this hairpin / cruciform motif).63Importantly, the group IIC introns (a subtype of the group II introns to which DGRs are closely related) are known to insert downstream of factor-independent transcriptional terminators consisting of similar GC-rich stem-loop structures; the dependency of these group IIC introns on such stem-loop translational-termination structures has been proposed help avoid insertion into critical coding sequences and facilitate their spread horizontally through bacterial genomes that broadly exhibit stem-loop transcriptional terminators downstream of their111 104 115
[0287] For the BPP-1 DGR, in vivo homing is highly dependent on a 35bp sequence downstream of VR that contains 8bp inverted repeats that are predicted to form the hairpin / cruciform structure typical of phage DGRs. Deletions of this 35bp hairpin / cruciform sequence abolish homing in vivo, deletions downstream of this hairpin / cruciform sequence in VR (3’delta 58) are well tolerated, however further downstream deletions (3’ delta 68) also effectively abolish homing. Perturbations of the stem length and especially insertions or sequence changes in the 4-nucleotide loop are highly detrimental to homing efficiency, indicating that the sequence is generally optimized to the BPP- 1 DGR. In vitro restriction assays indicate that harpins form on each strand of DNA in this motif (forming a full cruciform secondary structure) and that this cruciform secondary structure requires negatively supercoiled DNA to form. The polarity of genome replication does not impact homing efficiency and that the cruciform structure functions in an orientation-independent manner.94Because the 5’ end of the DGR-RNA is required for efficient homing and is likely wrapped around back onto bRT-Avd due to the predicted structure of the bRT-Avd-DGR-RNA ribonucleoprotein complex, it is plausible that a protein interaction between a sequence at the 5’ end of the DGR-Aty. Dkt. No. 125141.04976 MGH2024-443RNA and the intergenic hairpin / cruciform sequence downstream of the VR recruits the bRT-Avd-DGR-RNA ribonucleoprotein complex to the VR in a manner that positions the cis-primed cDNA for 3’ strand invasion.70It has also been proposed that this hairpin structure creates a strand bias and / or helps direct a host endonuclease for antisense strand nicking within the (G / C)i4 region to enable 3’ strand invasion.94
[0288] III.3b: cDNA Strand Displacement and Integration in VR.
[0289] Strand invasion and integration of antisense cDNA at the 3’ end of the VR occurs within the (G / C)i4 motif between positions TR107 and TRI 12; heterogeneity in marker transfer studies at position TR109 indicates that the replacement of the native antisense DNA strand with the antisense cDNA does not occur consistently at one position in this region.103The final nucleotides of the TR (TRI 14-134) encode the 21nt ‘Initiation of Mutagenic Homing*’ (IMH*) sequence, which is almost identical to the IMH sequence in the VR downstream of the VR (G / C)14 motif, differing only by 5nt. Homing and integration into the 3’ end of the VR locus requires the (G / C)i4 motif in TR / VR and IMH* / IMH, however no sequences upstream of the (G / C)i4 motif are required for 3’ integration. Removal of all sequences upstream of the (G / C)i4 motif still allows for cDNA integration at the 3 ’ end of VR. cDNA strand invasion is facilitated by either a nick in the antisense strand or a double-strand break within the (G / C)i4 motif. The endonuclease responsible for this cleavage activity could be either a protein or catalytic RNA and is currently unidentified, as bRT does not appear to possess a DNA endonuclease domain and has not been shown to have DNA endonuclease activity in vitro ,103The IMH-IMH* sequences preserve the unidirectionality of transposition in the DGR cassette and prevents deadenization of the TR; again, the IMH and IMH* sequences are located in the VR and TR, respectively, and differ by only 5nt. The IMH sequence in VR is indispensable for retrohoming, and mutations of this sequence drastically decrease homing efficiencies in in vivo PCR-based homing assays. Intriguingly, the simple substitution of 5nt to replace IMH* in the TR with a copy of IMH is sufficient to stimulate cDNA integration into the TR; this difference of 5bp between IMH* and IMH is therefore considered critical to maintaining the unidirectionality of transposition in the DGR and preventing deadeninization of the TR.108
[0290] Integration of cDNA at the 5’ end of the VR locus does not require specific intergenic sequences and only requires short segments of sequence homology (4-13nt) between the VR and corresponding 3’ end of the antisense cDNA. Existing data are consistent with a model of strandAty. Dkt. No. 125141.04976 MGH2024-443displacement of the native antisense strand by the transposed cDNA following cDNA integration at the 3’ end of VR. Integration at the 5’ end of the VR locus is scarless and requires no upstream intergenic sequences. Thus, the novel DNA sequence is inserted within the open reading frame of Mtd this scarless insertion allows expression of the BPP-1 tail fiber with diversified, solvent-exposed C-terminal receptor binding amino acids following DNA ligation and resolution of singlestranded adenine substitutions to double-stranded DNA. Importantly, the VR downstream intergenic sequences ((G / C)i4 region, IMH, and the DNA hairpin / cruciform) are not disrupted by the retrotransposition reaction, allowing the continuous retrotransposition of novel sequences into the VR; this capacity for continuous evolution is both a hallmark of DGR function and is considered to be critical to their function in dynamic environmen...
Claims
1. Aty. Dkt. No. 125141.04976 MGH2024-443CLAIMSWe claim:
1. A construct comprising sequences encoding:a diversity generating retroelement (DGR) accessory variability determinant (Avd) protein;a DGR reverse transcriptase (RT);a single-stranded DNA annealing protein (SSAP); andat least one non-coding RNA;wherein the sequences encoding the Avd, the RT, the SSAP, and the non-coding RNA are operably linked to one or more promoters;wherein the non-coding RNA comprises a target homology region (THR) that comprises homology to a target nucleic acid; andwherein at least one nucleotide in the THR is a non-homologous adenine.
2. The construct of claim 1, wherein at least one of the sequences encoding the Avd, the RT, or the SSAP is operably linked to an inducible promoter.
3. The construct of claim 1, wherein the inducible promoter is an arabinose-inducible promoter (PBAD).
4. The construct of claim 1, wherein the sequence encoding the non-coding RNA is operably linked to a constitutive promoter.
5. The construct of claim 4, wherein the constitutive promoter is ProD or Pro3.
6. The construct of claim 1, wherein the THR is flanked by a sequence encoding a minimal 5’ UTR and a sequence encoding a minimal 3’ UTR.
7. The construct of claim 1, wherein the construct comprises more than one non-coding RNA; and wherein each THR comprises homology to a different target nucleic acid.Aty. Dkt. No. 125141.04976 MGH2024-4438. The construct of claim 1 , wherein the construct further comprises a sequence encoding a dominant-negative mismatch repair protein.
9. The construct of claim 8, wherein the mismatch repair protein is MutL E32K.
10. The construct of claim 1, wherein the THR is between about 30 and about 100 nucleotides in length.
11. The construct of claim 1, wherein the non-coding RNA is between about 200 and about 500 nucleotides in length.
12. The construct of claim 11, wherein the non-coding RNA is about 200 nucleotides in length.
13. The construct of claim 1, wherein the RT comprises a protein or RNA sequence that groups within a DGR clade in a maximum-likelihood phylogenetic analysis.
14. The construct of claim 13, wherein the RNA sequence is a 3’ UTR sequence.
15. The construct of claim 1, wherein the RT is characterized by at least one of a loose active site, low processivity, and a hydrophobic template site.
16. The construct of claim 1, wherein the RT is derived from a Bordetella species.
17. The construct of claim 16, wherein the Bordetella species is Bordetella pertussis or Bordetella hinzii.
18. The construct of claim 1, wherein the RT is encoded by a sequence comprising SEQ ID NO: 4, SEQ ID NO: 6, or a sequence having at least 50% identity thereto.Aty. Dkt. No. 125141.04976 MGH2024-44319. The construct of claim 1 , wherein the Avd is derived from Bordetella pertussis or Bordetella hinzii.
20. The construct of claim 1, further comprising a spacer sequence between the sequence encoding the 5’ UTR and the sequence encoding the non-coding RNA.
21. The construct of claim 20, wherein the spacer sequence comprises SEQ ID NO: 89.
22. The construct of claim 1, wherein the SSAP is CspRecT, EcRecT, P22 Erf, ul36 Sak,or red beta.
23. The construct of claim 1, wherein the construct further encodes an Avd ribosomal binding site (RBS) of SEQ ID NO: 17; and an RT RBS of SEQ ID NO: 18.
24. The construct of claim 1, wherein the sequence encoding the Avd has at least 90% identity to SEQ ID NO: 3 or 5.
25. The construct of claim 1, wherein the sequence encoding the SSAP has at least 90% identity to SEQ ID NO: 7, 8, 9, 10, 11, 12, 13, or 14.
26. A plasmid comprising the construct of claim 1.
27. The plasmid of claim 26, wherein the plasmid comprises a sclOl origin of replication, a CloDF13 origin of replication, or aRK2 origin of replication.
28. A kit comprising:one or more plasmids comprising:a diversity generating retroelement (DGR) accessory variability determinant (Avd) protein;a DGR reverse transcriptase (RT);a single-stranded DNA annealing protein (SSAP); andAty. Dkt. No. 125141.04976 MGH2024-443at least one non-coding RNA;wherein each of the sequences encoding the Avd, the RT, the SSAP, and the non-coding RNA are operably linked to one or more promoters;wherein the non-coding RNA comprises a target homology region (THR) that comprises homology to a target nucleic acid; andwherein at least one nucleotide in the THR is a non-homologous adenine.
29. The kit of claim 28, comprising a first plasmid comprising:the construct comprising sequences encoding:the diversity generating retroelement (DGR) accessory variability determinant (Avd) protein;the DGR reverse transcriptase (RT); andthe single-stranded DNA annealing protein (SSAP)wherein the sequences encoding the Avd, the RT, and the SSAP; anda second plasmid comprising:the at least one non-coding RNA operably linked to a promoter.
30. The kit of claim 29, further comprising a phagemid vector or a phage vector encoding the target nucleic acid.
31. A kit comprising the plasmid of claim 26; and a phagemid vector or a phage vector encoding the target nucleic acid.
32. The kit of claim 30, wherein the phagemid vector or phage vector comprises an M13 phage origin of replication.
33. The kit of claim 30, wherein the phagemid vector comprises a pColEl plasmid origin of replication or a pl 5a plasmid origin of replication.
34. The kit of claim 33, wherein the phage origin of replication and the plasmid origin of replication both initiate replication in the direction of the target nucleic acid.Aty. Dkt. No. 125141.04976 MGH2024-44335. The kit of claim 28, further comprising a helper plasmid.
36. The kit of claim 35, wherein when the phagemid vector or phage vector comprises an M13 phage origin of replication, the helper plasmid is an M13KO7-derived helper plasmid.
37. The kit of claim 35, wherein the helper plasmid comprises a spacer sequence upstream of a gV start codon.
38. The kit of claim 35, wherein the helper plasmid comprises all genes necessary for propagation of a phage vector or a phagemid vector except genes encoded on the phage vector or phagemid vector or on an accessory plasmid.
39. The kit of claim 35, wherein the helper plasmid comprises all genes necessary for propagation of the phagemid vector or phage vector except gVI and gill.
40. The kit of claim 28, further comprising at least one accessory plasmid comprising gVI operably linked to a promoter, wherein the promoter activity is increased by activity of a protein of interest encoded by the target nucleic acid.
41. The kit of claim 28, further comprising a bioreactor configured for continuous evolution.
42. A cell comprising the construct of claim 1 or the plasmid of claim 26.
43. The cell of claim 42, further comprising a phagemid vector or a phage vector encoding the target nucleic acid.
44. The cell of claim 43, wherein the phagemid vector or phage vector comprises an M13 phage origin of replication.
45. The cell of claim 43, wherein the phagemid vector comprises a pColEl plasmid origin of replication or a pl 5a origin of replication.Aty. Dkt. No. 125141.04976 MGH2024-44346. The cell of claim 44, wherein the phage origin of replication and the plasmid origin of replication both initiate replication in the direction of the target nucleic acid.
47. The cell of claim 42, further comprising a helper plasmid.
48. The cell of claim 47, wherein when the phagemid vector or phage vector comprises an Ml 3 phage origin of replication, the helper plasmid is an M13KO7-derived helper plasmid.
49. The cell of claim 47, wherein the helper plasmid comprises a spacer sequence upstream of a gV start codon.
50. The cell of claim 47, wherein the helper plasmid comprises all genes necessary for propagation of a phage vector or a phagemid vector except genes encoded on the phage vector or phagemid vector or on an accessory plasmid.
51. The cell of claim 47, wherein the helper plasmid comprises all genes necessary for propagation of the phagemid vector except gVI and gill.
52. The cell of claim 42, further comprising at least one accessory plasmid.
53. The cell of claim 42, wherein the cell is an E. coli cell.
54. The cell of claim 53, wherein the E. coli cell is an s2060, sl030, or TGI strain.
55. The cell of claim 42, wherein an endogenous gene encoding Red is knocked out.
56. The cell of claim 55, wherein an endogenous gene encoding sbcB is knocked out.
57. The cell of claim 42, wherein the cell comprises a Q576A mutation in DNA gyrase.Aty. Dkt. No. 125141.04976 MGH2024-44358. A phage or phagemid vector comprising the construct of claim 1 or the plasmid of claim 26.
59. A kit comprising the cell of claim 42; and a helper cell, wherein the helper cell comprises in its genome all genes necessary for propagation of a phage vector or phagemid vector except genes encoded on the phage vector or phagemid vector or on an accessory plasmid.
60. An E. coli cell comprising at least one of a RecJ gene deletion; an sbcB gene deletion; and a Q576A mutation in DNA gyrase.
61. A method for generating and identifying functional mutants of a starting biomolecule, the method comprising:a) back-diluting and growing an initial fresh culture of host cells to mid log phase, wherein the host cells are transformed with: a helper plasmid comprising all genes required for phage propagation except gVI and gill; an accessory plasmid comprising gVI operably linked to a synthetic gene circuit, wherein when the synthetic gene circuit is activated, the gVI is transcribed; and the plasmid of claim 26;b) infecting the host cells with a selection phagemid vector comprising gill, wherein the selection phagemid vector comprises the target nucleic acid; wherein the target nucleic acid encodes the biomolecule; wherein activity of the biomolecule activates the synthetic gene circuit;c) incubating the host cells under conditions that allow for production of an infectious phagemid, wherein infectious phagemid encoding variants of the biomolecule that activate the synthetic gene circuit to a greater extent than the starting biomolecule will receive a propagation advantage;d) harvesting phagemid from the culture and discarding host cells from the culture; e) back-diluting and growing to mid-log phase a subsequent fresh culture of host cells; and infecting the host cells with the harvested phagemid from step d);f) repeating steps b) through e); andg) isolating phagemid encoding variants of the starting biomolecule that inactivate the synthetic gene circuit to a greater extent than the starting biomolecule.Aty. Dkt. No. 125141.04976 MGH2024-44362. A method for generating and identifying functional mutants of a starting first biomolecule, the method comprising:a) back-diluting and growing to mid-log phase an initial fresh culture of host cells, wherein the host cells are transformed with: a helper plasmid encoding all genes required for phage propagation except gVI and gill; a first accessory plasmid encoding gVI operably linked to a synthetic gene circuit, wherein when the synthetic gene circuit is activated, the gVI is transcribed; and a second accessory plasmid encoding at least one additionalbiomolecule that interacts with the first biomolecule; wherein the interaction of the first and at least one additional biomolecules activates the synthetic gene circuit; and the plasmid of claim 26;b) infecting the host cells with a selection phagemid comprising gill, wherein the phagemid comprises the target nucleic acid; wherein the target nucleic acid encodes the first biomolecule;c) incubating the host cells under conditions that allow for production of an infectious phagemid, wherein infectious phagemid encoding variants of the starting first biomolecule that activate the synthetic gene circuit to a greater extent than the starting first biomolecule will receive a propagation advantage;d) harvesting phagemid from the culture and discarding host cells;e) back-diluting and growing to mid-log phase a subsequent fresh culture of host cells; and infecting the host cells with the harvested phagemid from step d);f) repeating steps b) through e); andg) isolating phagemid encoding variants of the first biomolecule that activate the synthetic gene circuit to a greater extent that the starting first biomolecule.
63. The method of claim 61, wherein the host cell is an E. coll cell.
64. The method of claim 63, wherein the E. coli cell is an s2060 or an si 030 strain.
65. The method of claim 63, wherein the E. coli cell comprises at least one of a Red gene deletion; an sbcB gene deletion; and a Q576A mutation in DNAgyrase.Aty. Dkt. No. 125141.04976 MGH2024-44366. The method of claim 61, wherein the construct comprises an inducible promoter; and wherein step b) comprises contacting the host cell with an inducing agent that activates the inducible promoter.
67. The method of claim 66, wherein the inducible promoter is a PBAD promoter; and wherein the inducing agent is arabinose.
68. The method of claim 61, wherein the host cell is infected with between about IxlO2cfu ml;1and about IxlO9cfu mL’1of the selection phagemid.
69. The method of claim 68, wherein the host cell is infected with about IxlO6cfu mL1of the selection phagemid.
70. The method of claim 61, wherein the method is performed in a bioreactor in which the initial and subsequent populations of host cells are continuously flowed into the culture and wherein old host cells and phagemid encoding biomolecules that do not yield activation of the synthetic gene circuit above the level required for steady phagemid propagation at a bioreactor flow rate set by a researcher are continuously flowed out of the culture at a constant flow rate.
71. The method of claim 70, wherein the constant flow rate is between about 0.5 and about 3 volumes / hour.
72. The method of claim 71, wherein the constant flow rate is about 0.5 volumes / hour.
73. The method of claim 71, wherein the constant flow rate is about 3 volumes / hour.
74. The method of claim 61, further comprising h) sequencing the phagemid isolated in g).
75. A method for generating and identifying a mutant of a biomolecule endogenously expressed in a host cell or a mutant in a component of the genetic circuitry that regulates expression of the biomolecule, the method comprising:Aty. Dkt. No. 125141.04976 MGH2024-443a) transforming a population of the host cell with the construct of any one of claims 1-25, wherein the construct is operably linked to an inducible promoter, wherein the target nucleic acid encodes the endogenously expressed biomolecule or the component of the genetic circuitry that regulates expression of the biomolecule;b) contacting the host cell with an inducing agent that activates the inducible promoter; c) incubating the cell for between about 0 and about 24 hours under conditions that permit the function of the construct;d) selecting host cells containing mutant biomolecules or mutant components of the genetic circuitry that confer an improvement of interest; ande) extracting and sequencing gDNA from the host cells selected in d).
76. The method of claim 75, wherein the cell is an E. coli cell.
77. The method of claim 76, wherein the E. coli cell is an s2060, sl030, EcNR2, EcNR5, or MG1655 strain.
78. The method of claim 76, wherein the E. coli cell comprises at least one of a RecJ gene deletion; an sbcB gene deletion; and a Q576A mutation in DNAgyrase.
79. The method of claim 75, wherein the inducible promoter is a PBAD promoter; and wherein the inducing agent is arabinose.
80. A system for directed evolution of a protein of interest in a bacterial periplasm, the system comprising:(a) a phagemid vector comprising a nucleic acid encoding the protein of interest; and gill;(b) an accessory helper plasmid comprising:(i) a helper plasmid construct comprising all genes required for phage assembly except gill and gVI;(ii) a target antigen construct comprising a gene encoding a fusion protein, the fusion protein comprising a transmembrane protein and a target antigen; wherein the target antigen bindsAty. Dkt. No. 125141.04976 MGH2024-443the protein of interest; and wherein when the fusion protein is expressed, the target antigen is anchored in the membrane and exposed to the periplasm;(iii) a conditional propagation construct comprising a conditional promoter operably linked to gVI; wherein the conditional promoter comprises an operator to which the transmembrane protein is capable of binding; and wherein when the transmembrane protein binds to the operator, gVI is expressed; and(c) an engineered bacterial cell comprising: disruption or deletion of a gene encoding the transmembrane protein; and a CadBA operon that is naturally operably linked to the conditional promoter; wherein the conditional promoter is replaced with a heterologous constitutive promoter;wherein interaction of the protein of interest with the target antigen in the periplasm activates the conditional propagation construct and enables propagation of the phage.
81. The system of claim 80, wherein the transmembrane protein is CadC; and wherein the conditional promoter is a CadBA promoter.
82. The system of claim 80, wherein the CadC protein is a cytosolic DNA-binding region of CadC (CadCi-155).
83. The system of claim 80, wherein the nucleic acid encoding the protein of interest on the phagemid or the fusion protein further comprises a signal peptide for periplasmic export.
84. The system of claim 83, wherein the signal peptide is a PhoA signal peptide.
85. The system of claim 84, where in the PhoA signal peptide comprises a T19I mutation relative to SEQ ID NO: 93.
86. The system of claim 80, wherein the phagemid vector does not comprise SEQ ID NO: 116; wherein the accessory helper plasmid does not comprise SEQ ID NO: 132; and wherein the accessory helper plasmid does not comprise SEQ ID NO: 90.Aty. Dkt. No. 125141.04976 MGH2024-44387. The system of claim 80, wherein the phagemid vector does not comprise SEQ ID NOs: 91 and 92.
88. The system of claim 80, wherein the gene encoding the protein of interest is operably linked to a Psynth-gvi promoter comprising SEQ ID NO: 117.
89. The system of claim 80, wherein the accessory helper plasmid comprises an sclOl origin of replication.
90. The system of claim 80, wherein the phagemid vector comprises a M13+ ColEl+ origin of replication.
91. The system of claim 80, wherein the accessory helper plasmid comprises a minimal gill rbs upstream of a gl gene on the accessory helper plasmid that does not comprise SEQ ID NO: 118.
92. The system of claim 80, wherein the helper plasmid construct and the conditional propagation construct are oriented towards each other.
93. The system of claim 80, wherein the conditional promoter is a PcadBA promoter.
94. The system of claim 93, wherein the PcadBA promoter consists of SEQ ID NO: 119.
95. The system of claim 80, wherein the target antigen construct is in the antisense orientation relative to the conditional propagation construct.
96. The system of claim 80, wherein the accessory helper plasmid further comprises a polynucleotide encoding amino acids 1-9 of the PhoA signal peptide and an N-terminal half of Nostoc piinctiforme intein (NpuN), and comprising SEQ ID NO: 120, upstream of the target antigen construct; and wherein the phagemid vector further comprises amino acids 9-21 of theAty. Dkt. No. 125141.04976 MGH2024-443PhoA signal peptide and a C-terminal half of the Nostoc punctiforme intein (NpuC), and comprising SEQ ID NO: 131, upstream of the nucleic acid encoding the protein of interest.
97. The system of claim 80, wherein the accessory helper plasmid comprises SEQ ID NO: 113 or 127.
98. The system of claim 80, wherein the phagemid vector comprises SEQ ID NO: 121 or 130.
99. The system of claim 80, wherein all possible recombination sites between the phagemid vector and the accessory helper plasmid are disrupted.
100. The system of claim 80, further comprising the plasmid of claim 26.
101. The system of claim 80, wherein the protein of interest is an antibody or an antibody fragment.
102. The system of claim 101, wherein the antibody fragment is an scFv.
103. An engineered bacterial cell for periplasmic evolution of proteins encoded on a phagemid, the cell comprising:deletion or inactivation of cadC; anda heterologous constitutive promoter operably linked to cadBA.
104. The engineered cell of claim 103, wherein the heterologous constitutive promoter is a Ppro or a Pi23i50 promoter.
105. The engineered bacterial cell of claim 103, wherein the cell is an E. coli cell.
106. The engineered bacterial cell of claim 105, wherein the E. coli cell is an s2060 or sl030 strain.Atty. Dkt. No. 125141.04976 MGH2024-443107. The engineered bacterial cell of claim 105, wherein the E. coli cell comprises at least one of a RecJ gene deletion; an sbcB gene deletion; and a Q576A mutation in DNA yrase.
108. A method for generating a library of mutants of a biomolecule, the method comprising performing a directed evolution technique using the plasmid of claim 26; wherein the target nucleic acid encodes the biomolecule.
109. The method of claim 108, wherein the directed evolution technique is selected from phagemid display, phage display, ribosome display, yeast display, and mammalian display.
110. The method of claim 108, further comprising screening the library of mutants for mutants having an activity of interest.