Evolved and engineered prime editors with improved editing efficiency

EP4716740A1Pending Publication Date: 2026-04-01THE BROAD INST INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-23
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Current prime editing systems face challenges in improving reverse transcriptase and Cas9 variants for enhanced editing efficiency, particularly in mammalian cells, with previous attempts showing limited success in protein engineering and lower efficiencies compared to the highly engineered M-MLV RT and Cas9 variants.

Method used

The development of phage-assisted continuous evolution (PACE) and protein engineering to generate new reverse transcriptase and Cas9 variants, such as PE6a-PE6g, which offer improved editing efficiencies and are compatible with dual AAV delivery systems, enabling longer edits and twin prime editing in vivo.

Benefits of technology

The new variants demonstrate significant improvements in prime editing efficiency, with average 12- to 183-fold enhancements for 38- to 42-bp edits, achieving 62% targeted installation of loxP sequence in mouse cortex cells, and show reduced indel frequencies compared to previous state-of-the-art systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024030786_28112024_PF_FP_ABST
    Figure US2024030786_28112024_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides evolved and engineered reverse transcriptase variants and Cas9 variants with improved properties (e.g., improved editing efficiency when used in the context of a prime editor). Fusion proteins, including for example prime editors, comprising the reverse transcriptase variants and Cas9 variants described herein are also provided by the present disclosure. The present disclosure also provides polynucleotides encoding the reverse transcriptase variants, Cas9 variants, and prime editors provided herein, as well as vectors comprising such polynucleotides. Pharmaceutical compositions and cells comprising the reverse transcriptase variants, Cas9 variants, and prime editors described herein are also provided by the present disclosure. The present disclosure also provides methods and uses involving the reverse transcriptase variants, Cas9 variants, and prime editors described herein.
Need to check novelty before this filing date? Find Prior Art

Description

EVOLVED AND ENGINEERED PRIME EDITORS WITH IMPROVED EDITING EFFICIENCY RELATED APPLICATIONS

[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application, U.S.S.N.63 / 503,892, filed May 23, 2023; U.S. Provisional Application, U.S.S.N.63 / 506,026, filed June 2, 2023; U.S. Provisional Application, U.S.S.N.63 / 510,078, filed June 23, 2023; U.S. Provisional Application, U.S.S.N.63 / 596,006, filed November 3, 2023; and U.S. Provisional Application, U.S.S.N.63 / 508,616, filed June 16, 2023, each of which is incorporated herein by reference. REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0002] The contents of the electronic sequence listing (B119570180WO00-SEQ-TNG.xml; Size: 226,657 bytes; and Date of Creation: May 15, 2024) is incorporated herein by reference in its entirety. GOVERNMENT SUPPORT

[0003] This invention was made with government support under Grant Nos. UG3AI150551, U01AI142756, R35GM118062, and RM1HG009490, awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND OF THE INVENTION

[0004] The ability to install precise, targeted changes into the genomes of living cells and organisms has advanced the understanding of biological systems and may provide one-time treatments for genetic diseases. Prime editing (PE) is a versatile gene editing technology capable of installing any base substitution, insertion, or deletion without generating DSBs1. Since >95% of pathogenic substitutions, insertions, deletions, or combinations thereof are ≤50 bp in length2, PE raises the possibility of correcting a large fraction of known disease- causing mutations. PE requires a prime editing guide RNA (pegRNA) and a prime editor protein, which comprises a programmable nickase (typically S. pyogenes Cas9 H840A nickase) and a reverse transcriptase (RT). The first-generation prime editor (PE1) used the wild-type Moloney murine leukemia virus (M-MLV) RT, while subsequent prime editors (PE2-PE5) use an engineered pentamutant M-MLV RT (FIG.1A)1,3. The pegRNA contains a core guide RNA scaffold that binds the programmable nickase, a spacer that specifies the B1195.70180WO00 12418099.1target site, a primer binding site (PBS) that is complementary to the target DNA, and a reverse transcriptase template (RTT) that encodes the desired edit. To install an edit, the prime editor•pegRNA complex pairs with one strand of the target genomic locus and nicks the opposite strand to generate an exposed 3′ end of nicked genomic DNA, which binds to the complementary PBS of the pegRNA. The RT engages the resulting primer-template complex and initiates reverse transcription of the RTT, generating a 3′ DNA flap containing the desired edit. The newly synthesized 3′ flap is incorporated into the genome by cellular DNA repair pathways, replacing the original DNA sequence and leading to the permanent installation of the desired edit1. In the PE3 and PE5 systems, an additional sgRNA is used to nick the non-edited DNA strand, improving editing efficiencies by biasing cellular mismatch repair to favor replacement of the non-edited strand (FIG.1A)1,3.

[0005] Since their development, PE systems have been improved by stabilizing or circularizing pegRNAs4–6, varying the prime editor architecture3,4, and manipulating or evading cellular mismatch repair to favor desired editing outcomes3,9. Twin prime editing (twinPE) and related methods have also been developed that use two pegRNAs to install edited sequence on both DNA strands, replacing the original genomic sequence between the two prime editing nicks with larger (>100-bp) programmable insertions and deletions10-13,15,16. Prime editing and twinPE have also been used to install site-specific recombinase landing sites, enabling recombinase-mediated gene-sized (>5,000 bp) targeted insertions or inversions10. Prime editing, twin prime editing, and prime editors are further described, e.g., in International Patent Application No. PCT / US2020 / 023721, filed March 19, 2020, which published as WO 2020 / 191239; International Patent Application No. PCT / US2021 / 031439, filed May 7, 2021, which published as WO 2021 / 226558; and International Patent Application No. PCT / 2021 / 052097, filed September 24, 2021, which published as WO 2022 / 067130; the contents of each of which is incorporated by reference herein.

[0006] Despite this progress, the reverse transcriptase at the heart of prime editors has proven challenging to improve through protein engineering. Many of the prime editing systems reported to date, including the current PE4max and PE5max systems, use the engineered M- MLV RT in PE2. The five M-MLV RT mutations in PE2 were identified over several decades of in vitro screening for improved RT variants18–21, followed by screening of many combinations of M-MLV RT mutants that optimize prime editing efficiencies1. While these mutations are critical to the efficiency of prime editing, few analogous mutations have been described for other RTs that have been tested in prime editing experiments. Prime editor proteins that use non-M-MLV RTs in principle could offer important benefits, including B1195.70180WO00 12418099.1smaller size that could facilitate in vivo prime editor delivery, mRNA production, or ribonucleoprotein (RNP) preparation. Different RT enzymes may also improve properties of PE such as editing efficiency, suitability for longer or shorter prime edits, or compatibility with installing sequences of different composition, just as different deaminases have provided a diverse collection of base editors that greatly increase the likelihood of finding one ideally suited to a particular application22. Despite these potential benefits, all previously reported prime editors that do not use the engineered M-MLV RT in PE2 have shown substantially lower prime editing efficiencies than PE2 for most target sequences, even after extensive protein engineering4,17,24. Further improvement of the highly engineered M-MLV RT in PE2 has also proven difficult, as all reported variants of this RT have also yielded little or no improvements in prime editing efficiency in mammalian cells17,24. Similarly, although it has been shown that Cas9 mutations known to improve nuclease performance can also increase prime editing efficiency3, mutants of Cas9 identified specifically to improve prime editing have not yet been reported. Accordingly, additional RT and Cas9 variants evolved and / or engineered with the purpose of improving prime editing efficiency would advance the art. SUMMARY OF THE INVENTION

[0007] As presented herein, a phage-assisted continuous evolution (PACE)26selection for prime editing was developed, and PE PACE and protein engineering was used to generate new polymerase (e.g., reverse transcriptase or “RT”) and Cas9 variants that enhance prime editing efficiency and in-vivo deliverability. First, natural RTs were screened from a wide variety of organisms, and it was found that most exhibited negligible prime editing activity in mammalian cells. Two weakly active RTs, those from the Escherichia coli Ec48 retron27and from the Schizosaccharomyces pombe Tf1 retrotransposon28, were evolved to create next- generation prime editors (PE6a and PE6b) that are 516-810 bp smaller than PE2 while offering mammalian prime editing efficiencies comparable to or higher than those of PE2 for many target sites and types of edits. It was discovered that the reduced RT processivity of PEmax∆RNaseH (i.e., PEmax comprising an MMLV reverse transcriptase with a truncation of the C-terminal RNaseH domain), the commonly used prime editor variant used in dual- AAV delivery systems4,23,29–31, causes it to underperform at long edits with a high degree of secondary structure. To generate dual AAV-compatible RTs that can install longer edits or edits that require RT templates with a high degree of secondary structure, PE PACE and protein engineering were used to generate PE6c and PE6d. These RT mutants offer large benefits in editing efficiency compared to PEmax∆RNaseH for edits that require structured B1195.70180WO00 12418099.1pegRNA RT templates. The PE6a-PE6d RTs also offer improvements in editing efficiency and fewer indel frequencies over full-length PEmax. Finally, PE PACE was used to evolve the Cas9 nickase domain of prime editors to create PE6e-PE6g, which further improve prime editing efficiencies. The improved RT and Cas9 nickase domains of PE6 variants can be combined with each other, as well as with mismatch repair evasion strategies3, epegRNAs5, and the PEmax architecture3, to offer cumulative benefits in a variety of contexts, including in patient-derived fibroblasts and primary human T cells. Finally, it is demonstrated that PE6c and PE6d are uniquely enabling for performing long prime edits and twinPE in vivo. After dual-AAV delivery of PE6 systems, on average 12- to 183-fold improvement in prime editing efficiency was achieved compared with previous state-of-the-art systems for installation of 38- to 42-bp edits in the mouse cortex, yielding 62% targeted installation of the loxP sequence among transduced cells in the mouse cortex.

[0008] The improved prime editors PE6a-PE6g described herein comprise the following amino acid substitutions relative to particular wild-type reverse transcriptase or Cas9 proteins:B1195.70180WO00 12418099.1

[0009] In some embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6a (or a reverse transcriptase at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the reverse transcriptase of PE6a) and a nucleic acid-programmable DNA-binding protein (napDNAbp) (e.g., a Cas9 protein). In some embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6b (or a reverse transcriptase at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the reverse transcriptase of PE6b) and a napDNAbp (e.g., a Cas9 protein). In some embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6c (or a reverse transcriptase at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the reverse transcriptase of PE6c) and a napDNAbp (e.g., a Cas9 protein). In some embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6d (or a reverse transcriptase at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the reverse transcriptase of PE6d) and a napDNAbp (e.g., a Cas9 protein and a napDNAbp (e.g., a Cas9 protein).

[0010] In some embodiments, the present disclosure provides prime editors comprising the Cas9 protein of PE6e (or a Cas9 protein at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the Cas9 protein of PE6e) and a polymerase (e.g., a reverse transcriptase). In some embodiments, the present disclosure provides prime editors comprising the Cas9 protein of PE6f (or a Cas9 protein at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the Cas9 protein of PE6f) and a polymerase (e.g., a reverse transcriptase). In some embodiments, the present disclosure provides prime editors comprising the Cas9 protein of PE6g (or a Cas9 protein at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the Cas9 protein of PE6g) and a polymerase (e.g., a reverse transcriptase). B1195.70180WO00 12418099.1

[0011] In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6a and the Cas9 protein of PE6e (PE6a-e). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6a and the Cas9 protein of PE6f (PE6a-f). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6a and the Cas9 protein of PE6g (PE6a-g). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6b and the Cas9 protein of PE6e (PE6b-e). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6b and the Cas9 protein of PE6f (PE6b-f). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6b and the Cas9 protein of PE6g (PE6b-g). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6c and the Cas9 protein of PE6e (PE6c-e). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6c and the Cas9 protein of PE6f (PE6c-f). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6c and the Cas9 protein of PE6g (PE6c-g). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6d and the Cas9 protein of PE6e (PE6d-e). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6d and the Cas9 protein of PE6f (PE6d-f). In certain embodiments, the present disclosure provides prime editors comprising the reverse transcriptase of PE6d and the Cas9 protein of PE6g (PE6d-g).

[0012] In one aspect, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1 (Tf1 reverse transcriptase), wherein the reverse transcriptase variant comprises amino acid substitutions at positions 70, 72, 87, 102, 106, 118, 128, 158, 269, 363, 413, and 492 relative to SEQ ID NO: 1, or corresponding substitutions in a homologous sequence. In some embodiments, the reverse transcriptase variant further comprises amino acid substitutions at positions 188, 260, 297, and 288 relative to SEQ ID NO: 1.

[0013] In another aspect, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30 (MMLV reverse transcriptase), wherein the reverse transcriptase variant comprises amino acid substitutions at positions 128 and 200 relative to SEQ ID NO: 30, or corresponding substitutions in a homologous B1195.70180WO00 12418099.1sequence. In some embodiments, the reverse transcriptase variant further comprises amino acid substitutions at positions 223, 306, 313, and 330 relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence. In some embodiments, the reverse transcriptase variant comprises a truncation of the RNaseH domain of SEQ ID NO: 30 (e.g., a truncation at D497 in SEQ ID NO: 30).

[0014] In another aspect, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30 (MMLV reverse transcriptase), wherein the reverse transcriptase variant comprises the amino acid substitutions T128N and V223M; T128N and V223Y; T128F and V223M; or D200C and V223M relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence.

[0015] In another aspect, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30 (MMLV reverse transcriptase), wherein the reverse transcriptase variant comprises amino acid substitutions at positions 128, 129, 196, 200, and 223 relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence.

[0016] In another aspect, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2 (Streptococcus pyogenes Cas9 nickase), wherein the Cas9 variant comprises amino acid substitutions at positions 775 and 918 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

[0017] In another aspect, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2 (Streptococcus pyogenes Cas9 nickase), wherein the Cas9 variant comprises amino acid substitutions at positions 99, 471, 632, 645, and 721 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

[0018] In another aspect, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2 (Streptococcus pyogenes Cas9 nickase), wherein the Cas9 variant comprises amino acid substitutions at positions 99, 471, and 632 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

[0019] In another aspect, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least B1195.70180WO00 12418099.199% sequence identity with SEQ ID NO: 2 (Streptococcus pyogenes Cas9 nickase), wherein the Cas9 variant comprises amino acid substitutions at positions 471 and 918 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

[0020] In another aspect, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2 (Streptococcus pyogenes Cas9 nickase), wherein the Cas9 variant comprises amino acid substitutions at positions 753 and 1151 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

[0021] In another aspect, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2 (Streptococcus pyogenes Cas9 nickase), wherein the Cas9 variant comprises one or more amino acid substitutions at positions selected from the group consisting of 260, 298, 395, 769, 778, 1014, 1034, 1100, 1106, 1138, 1152, and 1320 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

[0022] In another aspect, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2 (Streptococcus pyogenes Cas9 nickase), wherein the Cas9 variant comprises amino acid substitutions at positions 23 and 754 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

[0023] In another aspect, the present disclosure provides prime editors comprising (i) any of the reverse transcriptase variants provided herein, and (ii) a napDNAbp, for example, a Cas9 protein (e.g., a Cas9 nickase, or any of the Cas9 variants provided herein (which may also be Cas9 nickases), or a Cas9 nuclease or nuclease-inactivated Cas9 (dCas9)).

[0024] In another aspect, the present disclosure provides prime editors comprising (i) any of the Cas9 variants provided herein, and (ii) a polymerase (e.g., a reverse transcriptase, such as any of the reverse transcriptase variants provided herein). In some embodiments, the reverse transcriptase comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 7 (Ec48 reverse transcriptase), wherein the reverse transcriptase comprises amino acid substitutions at positions 60, 87, 165, 243, 267, 279, 318, and 343 relative to SEQ ID NO: 7, or corresponding positions in a homologous sequence.

[0025] In some embodiments, the reverse transcriptase variants provided herein comprise the amino acid sequence of any one of SEQ ID NOs: 25-27 or 50, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or B1195.70180WO00 12418099.1at least 99% identical to the amino acid sequence of any one of SEQ ID NOs: 25-27 or 50. In some embodiments, the Cas9 variants provided herein comprise the amino acid sequence of any one of SEQ ID NOs: 28, 48, or 49, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of any one of SEQ ID NOs: 28, 48, or 49. In some embodiments, the prime editors provided herein comprise a Cas9 variant comprising the amino acid sequence of any one of SEQ ID NOs: 28, 48, or 49, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of any one of SEQ ID NOs: 28, 48, or 49, and a reverse transcriptase variant comprising the amino acid sequence of any one of SEQ ID NOs: 25-27 or 50, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of any one of SEQ ID NOs: 25-27 or 50.

[0026] In another aspect, the present disclosure provides fusion proteins comprising any of the Cas9 variants provided herein and an effector domain. In certain embodiments, the effector domain comprises nuclease activity, nickase activity, recombinase activity, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity, transcriptional repression activity, or polymerase activity.

[0027] In another aspect, the present disclosure provides complexes comprising any of the prime editors or other fusion proteins provided herein and a prime editing guide RNA (pegRNA).

[0028] In some aspects, the present disclosure provides polynucleotides encoding any of the reverse transcriptase variants, Cas9 variants, fusion proteins, or prime editors provided herein. In another aspect, the present disclosure provides vectors comprising any of the polynucleotides provided herein.

[0029] In another aspect, the present disclosure provides adeno-associated virus (AAV) particles comprising any of the reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, complexes, polynucleotides, and / or vectors provided herein.

[0030] In another aspect, the present disclosure provides cells comprising any of the reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, complexes, polynucleotides, vectors, and / or AAV particles provided herein. B1195.70180WO00 12418099.1

[0031] In another aspect, the present disclosure provides pharmaceutical compositions comprising any of the reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, complexes, polynucleotides, vectors, AAV particles, and / or cells provided herein.

[0032] In another aspect, the present disclosure provides methods for editing a nucleic acid molecule by prime editing comprising contacting a nucleic acid molecule with any of the prime editors or complexes provided herein. In certain embodiments, the edit in the nucleic acid molecule comprises one or more nucleotide insertions, one or more nucleotide substitutions, one or more nucleotide deletions, or a combination thereof. In certain embodiments, the method is a method of twin prime editing (also known as dual flap prime editing). In some embodiments, the present disclosure provides methods of using the prime editors, complexes, polynucleotides, or vectors provided herein in veterinary uses. In some embodiments, the present disclosure provides methods of using the prime editors, complexes, polynucleotides, or vectors provided herein in agricultural uses.

[0033] In another aspect, the present disclosure provides kits comprising any of the reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, complexes, polynucleotides, vectors, AAV particles, and / or cells provided herein.

[0034] In another aspect, the present disclosure provides for the use of any of the reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, complexes, polynucleotides, vectors, AAV particles, and / or cells provided herein in the manufacture of a medicament.

[0035] In another aspect, the present disclosure provides for the use of any of the reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, complexes, polynucleotides, vectors, AAV particles, and / or cells provided herein in medicine.

[0036] In some aspects, the present disclosure provides systems for phage-assisted continuous and non-continuous evolution (PACE and PANCE) of prime editors. In certain embodiments, the present disclosure provides systems comprising: i) a first polynucleotide encoding a pegRNA and the gIII gene; ii) a second polynucleotide encoding a Cas9 protein fused to an N-intein; iii) a third polynucleotide encoding an RNA polymerase; iv) a fourth polynucleotide encoding proteins capable of mutagenizing a phage, optionally wherein the fourth polynucleotide comprises the MP6 plasmid; and v) a fifth polynucleotide encoding a reverse transcriptase fused to a C-intein. In certain embodiments, the present disclosure provides systems comprising i) a first polynucleotide encoding a pegRNA and the gIII gene; ii) a second polynucleotide encoding a prime editor; iii) a third polynucleotide encoding an B1195.70180WO00 12418099.1RNA polymerase; and iv) a fourth polynucleotide encoding proteins capable of mutagenizing a phage, optionally wherein the fourth polynucleotide comprises the MP6 plasmid.

[0037] It should be appreciated that the foregoing concepts, and additional concepts discussed below, may be arranged in any suitable combination, as the present disclosure is not limited in this respect. Further, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non- limiting embodiments when considered in conjunction with the accompanying Figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The following Figures form part of the present specification and are included to further demonstrate certain aspects of the present disclosure, which can be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.

[0039] FIGs.1A-1J: Identification and engineering of reverse transcriptase enzymes into new prime editor candidates. FIG.1A shows an overview of prime editing using the PE1, PE2, and PE3 systems. All three systems use a prime editor protein comprising SpCas9(H840A) nickase fused to a reverse transcriptase (RT) enzyme. The PE1 system uses the RT from the Moloney murine leukemia virus (M-MLV), while the PE2 system uses an engineered pentamutant variant of the M-MLV RT with D200N, L603W, T306K, W313F, and T330P. An additional single guide RNA (sgRNA) is used in the PE3 system to nick the non-edited strand. PBS = primer binding site. RT template = reverse transcriptase template. FIG.1B shows phylogenetic classification of all RTs tested for prime editing herein (circles). Enzymes that exhibit activity in the PE system (dark gray circles) belong to four different RT classes. FIG.1C shows 20 different RT enzymes other than the M-MLV RT exhibit activity in the prime editing system at endogenous sites in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. Throughout all figures, prime editing efficiencies shown reflect the frequency of the intended prime editing outcome with no indels or other changes at the target site. FIG.1D shows comparison of wild type (WT) Tf1 RT, PE2∆RNaseH (i.e., comprising a truncation of the C-terminal RNaseH domain of the MMLV reverse transcriptase, e.g., between amino acids D497 and I498 in an MMLV reverse transcriptase of SEQ ID NO: 30), and PE2 at three longer, complex PE (HEK3) or twinPE (CCR5 and IDS) edits in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.1E shows a comparison of prime editors containing engineered retroviral RT variants with their WT counterparts in HEK293T B1195.70180WO00 12418099.1cells. rdPERV = porcine endogenous retrovirus RT D200N, T306K, W313F, E330P, L603W. rdAVIRE = avian reticuloendotheliosis virus RT D200N, T306K, W313F, G330P, L603W. rdKORV = koala retrovirus RT D198N, T304K, W311F, E328P, L600W. rdWMSV = woolly monkey sarcoma virus RT D198N, T304K, W311F, E328P, L600W. All values from n=3 independent replicates are shown. Horizontal bars show the mean value. FIG.1F shows residues mutated to improve editing of the Tf1 RT prime editor correspond to V188, R118, L258, M281 and V286 (red) in Ty3 RT (blue). V188 and R118 are in close proximity to the RNA (green) substrate and correspond to K118 and S188 in Tf1, respectively. L258, M281 and V286 are near the DNA (yellow) substrate and correspond to I260, S297 and R288 in Tf1, respectively. FIG.1G shows that rationally designed Tf1 pentamutant variant (rdTf1) shows improvements in editing over its WT counterpart in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. All edits are PE edits, except the AAVS1 site, which is twinPE. rdTf1 = Tf1 RT K118R, S188K, I260L, S297Q, R288Q. FIG.1H shows that rationally designed Ec48 triple mutant variant (rdEc48) shows improvements in editing over its WT counterpart for five edits in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. rdEc48 = Ec48 RT R315K, L182N, T189N. FIG.1I shows a comparison of prime editors containing engineered RT variants with PE2 in HEK293T cells. All values from n=3 independent replicates are shown. Horizontal bars show the mean value. All edits using single-flap prime editing, except the AAVS1 site, which uses twinPE. FIG.1J shows a comparison of rdTf1 with PE2 and its WT counterpart at three longer, complex PE (HEK3), or twinPE (CCR5 and IDS) edits in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values.

[0040] FIGs.2A-2K: Development and validation of a prime editing PACE selection. FIG. 2A is a schematic of PE PACE selection circuit. Upon infection of host E. coli cells by selection phage (SP, blue), the NpuN intein and NpuC intein (pink) mediate protein splicing to reconstitute the two halves of the PE2 prime editor (purple and pink). The prime editor then engages a pegRNA (dark green) and corrects a frameshift in T7 RNAP (orange) via prime editing. Functional T7 RNAP then transcribes gIII (light green), which enables propagation of the SP. FIG.2B shows evaluation of the v1 PE PACE circuit. Phage replication levels from overnight propagation of empty phage (red), NpuC-PE2-RT phage (purple), and T7-RNAP phage (green) on host cells harboring the PE PACE circuit before pegRNA optimization. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.2C shows a screen of pegRNAs for the v1 PE PACE circuit. B1195.70180WO00 12418099.1Overnight propagation values of empty phage (red), NpuC-PE2-RT phage (purple), and T7- RNAP phage (green) are shown. Each point reflects the mean value of n=3 independent biological replicates for a different pegRNA. Individual replicates are shown in FIG.9C. FIG.2D shows overnight propagation of empty phage (red), NpuC-PE1-RT phage (light purple), NpuC-PE2-RT phage (dark purple), and T7-RNAP phage (green) in the v1 pegRNA- optimized circuit. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.2E shows PANCE titers for the evolution of NpuC-PE1-RT phage. Grey shading indicates a passage of evolutionary drift, in which phage were supplied gIII in the absence of selection to allow free mutagenic replication. Titers of four replicate lagoons are shown. FIG.2F shows a mutation table for RT clones that enriched during PANCE of NpuC-PE1-RT phage. Four clones from each lagoon (L1-L4, with clones ordered by lagoon) were sequenced. Light purple denotes a conserved mutation. Dark purple denotes a conserved mutation that was also present in the previously engineered RT in PE21. FIG.2G shows a schematic of the PE PACE selection for evolution of the whole prime editor, including the Cas9 domain. The P1 plasmid (green) and P3 plasmid (orange) are identical to those used in FIG.2A. FIG.2H shows a PANCE experiment to compare the outcome of selection on v1 (requiring a 1-bp insertion) and v2 (requiring a 20-bp insertion) selection circuits. Whole- editor phage were divided into 16 separate lagoons, with 8 lagoons evolved on the v1 circuit (yellow) and 8 lagoons evolved on the v2 circuit (blue). After 31 passages, clones from each selection were Sanger sequenced, and the resulting mutations were compared to generate FIGs.2D-2F. From top to bottom, the sequences are SEQ ID NOs: 134-137. FIG.2I shows violin plots showing the number of mutations per clone for the M-MLV domain of whole- editor phage evolved with either the v1 (yellow) or v2 (blue) circuit. Data are shown as individual values, with one dot representing one sequenced phage. The mean value is shown as a dotted line. FIG.2J shows predicted positions of mutated residues in M-MLV from v1 (yellow) or v2 (blue) PANCE. The structure is from the highly homologous XMRV (PDB: 4HKQ). FIG.2K shows overnight propagation of pools of wild-type RT and evolved RT phage on their cognate or noncognate host-cell selection strains. Phage were from PANCE on the v1 circuit (yellow bars), from PANCE on the v2 circuit (blue bars), or wild-type-PE2 phage (grey bars). Propagation was then measured in the v1 circuit (left) or the v2 circuit (right). Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values.

[0041] FIGs.3A-3H: Phage-assisted evolution of compact RTs for prime editing. FIG.3A shows a summary of evolution campaigns for NpuC-RT phage encoding Gs RT, Ec48 RT, or B1195.70180WO00 12418099.1Tf1 RT. Shading indicates which selection circuit (v1 in yellow, v2 in blue, and v3 in purple) was used. Whether a given evolution was PANCE or PACE is specified: the number in parentheses after a PANCE or PACE label specifies how many passages of PANCE were performed (p) or how many hours of PACE (h) were performed. Arrowheads indicate that an evolution was stopped and increased in stringency without mammalian characterization; mutants characterized in mammalian cells are denoted with a dot and labeled. Finally, evolutions that used extra manipulations to increase stringency are labeled in pink, reflecting either a change in the PBS or a change in the expression of the target T7 RNAP gene. FIG. 3B shows the position of residues in the Gs RT close to the DNA•RNA substrate that were mutated following evolution mapped onto the structure of the Gs RT (PDB: 6AR1). Residues mutated following PANCE in the v2 circuit are red, residues mutated following PANCE and PACE in the v1 circuit are blue, the DNA substrate is green, and the RNA substrate is yellow. FIG.3C shows predicted positions of residues in the Ec48 RT close to the DNA•RNA substrate (E60, E279, and K318) that were mutated after PANCE in the v1 and v2 circuit. Residues are mapped onto the AlphaFold predicted structure of the Ec48 overlayed with the substrate of the XMRV RT (PDB: 4HKQ). Residue mutated following PANCE in the v1 circuit is blue, residues mutated following PANCE in the v2 circuit are red, the DNA substrate is green, and the RNA substrate is yellow. FIG.3D shows predicted positions of several conserved residues in the Tf1 RT that were mutated after PANCE in the v1, v2, and v3 circuit. Residues are mapped onto the AlphaFold predicted structure of the Tf1 RT overlayed with the substrate of the Ty3 RT (PDB: 4OL8). Residues S492, K413, I128, and K118 are all predicted to be close to the substrate while residues P70, G72, M102, and K106 decorate the surface of the enzyme that may be important for its interaction with the RTT of the pegRNA. Residues mutated following PANCE in the v1 circuit are blue, residues mutated following PANCE in the v2 circuit are red, and residues mutated following PANCE in the v3 circuit are in orange. The DNA substrate is green, and the RNA substrate is yellow. FIG.3E shows prime editing using prime editors containing wild-type (grey) Gs, Ec48, and Tf1 RTs, evolved Gs-RT (evoGs, green), evolved Ec48 RT (evoEc48, blue), and evolved Tf1 RT (evoTf1, yellow) in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. Throughout all figures, prime editing efficiencies shown reflect the frequency of the intended prime editing outcome with no indels at the target site. FIG.3F shows a comparison of prime editors in the optimized PEmax architecture containing either engineered pentamutant Marathon RT (Marathon penta, red), evoEc48 (blue), or evoTf1 (yellow) with PEmax (gray) in HEK293T cells. Bars reflect the mean of B1195.70180WO00 12418099.1n=3 independent replicates. Dots show individual replicate values. FIG.3G shows prime editing in primary human T-cells at commonly edited test loci. Bars reflect the mean of n=4 independent replicates. Dots show individual replicate values. Indel-free editing is shown in blue or pink, and indels are shown in grey. FIG.3H shows correction of the HEXA 1278insTATC mutation that causes Tay-Sachs disease in a HEK293T cell line model previously engineered to harbor the mutation (left) and in patient-derived fibroblasts (right). Bars reflect the mean of n=3 independent replicates for the HEK293T cell in model. Bars reflect n=2 independent replicates for the patient-derived fibroblasts. Dots show individual replicate values.

[0042] FIGs.4A-4J: Evolved prime editor preferences and summary of RT evolution campaigns described herein. FIG.4A shows a summary of evolution and engineering campaigns used to generate PE6c and PE6d. FIG.4B shows conserved mutations from M- MLV RT evolution. The structure of XMRV RT (PDIB 4HKQ), which is highly homologous to M-MLV, shows PACE-evolved residues (blue) lie close to the enzyme active site (dark grey) and DNA / RNA duplex substrate (pink / purple). An incoming dNTP is shown in yellow. Below, pink lines indicate locations in the M-MLV RT at which PACE-evolved mutations truncated the protein. FIG.4C shows fold-change in editing efficiency relative to PEmax for PEmax∆RNaseH, PE6c, and PE6d in HEK293T cells. Individual replicates are plotted, with n=3 biological replicates per edit. FIG.4D shows editing efficiencies of PEmax∆RNaseH and PE6d at the HEK3 +1 loxP insertion edit (pink) and the HEK3 +1 FLAG insertion edit (orange) in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. The NUPACK-predicted structures of the RTT and PBS extensions for each edit are shown. FIG.4E shows results of a TdT assay on the HEK3 +1 loxP insertion edit in HEK293T cells. The y-axis indicates the percentage of total RT products of a given length, and the x-axis represents the length of the product in base pairs. PEmax∆RNaseH is shown in grey, and PE6d is shown in blue. The lines are mean values from n=3 biological replicates. The pink box indicates DNA bases templated by the structured portions of the pegRNA. FIG.4F shows editing efficiencies of PEmax∆RNaseH (grey) and PE6d (blue) at an example engineered hairpin edit and its corresponding unpinned control in HEK293T cells. The sequence of the RTT is shown (from top to bottom, SEQ ID NOs: 138 and 139), with point mutations in the unpinned control shown in red. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. The NUPACK-predicted structures of the RTT and PBS extensions for each edit is shown. FIG. B1195.70180WO00 12418099.14G shows the relationship between pegRNA RTT / PBS secondary structure and PE6d improvements. The y-axis reflects the fold-improvement of PE6d over PEmax∆RNaseH. The x-axis is the absolute value of the free energy of pegRNA folding as measured by NUPACK. Each dot represents one edit in HEK293T cells that was calculated from the mean values from n=3 biological replicates. See FIG.11D for individual editing values and edit identities. FIG.4H shows a comparison of evolved and engineered RTs to PEmax∆RNaseH at typical twinPE edits in HEK293T cells. Bars reflect the mean indel-free editing efficiency of n=3 independent replicates. Dots show individual replicate values. Solid bars indicate editing efficiency. Striped bars indicate indels. FIG.4I shows twinPE-mediated insertion of the 38- bp attB sequence into the Rosa26 locus in N2a cells. Indel-free editing is shown in yellow, and indels are shown in grey. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.4J shows PE-mediated insertion of a 42-bp sequence containing loxP into the Dnmt1 locus in N2a cells. Indel-free editing is shown in yellow, and indels are shown in grey. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values.

[0043] FIGs.5A-5H: Characterization of PE6 variants compared with PEmax. FIG.5A shows prime editing efficiencies of PE6c, PE6d, and PEmax at challenging twinPE edits in HEK293T cells. Bars reflect the mean indel-free editing efficiency of n=3 independent replicates. Dots show individual replicate values. FIG.5B shows edit to indel ratios of PE6c, PE6d, and PEmax at sites shown in 5A in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.5C shows twin prime editing in primary human T-cells at the CCR5 safe harbor locus. Indel-free editing is shown in red, and indels are shown in grey. Bars reflect the mean of n=4 independent replicates. Dots show individual replicate values. FIG.5D shows edit to indel ratios of PE6b and PEmax∆RNaseH normalized to that of PEmax in HEK293T cells. Individual replicates are plotted, with n=3 biological replicates per edit. Lines reflect the mean across all edits and replicates. Individual editing efficiencies and indel levels are shown in FIGs.12D-12E. FIG. 5E shows edit to indel ratios of prime editors at endogenous HEK293T sites. The editor with the highest edit:indel ratio was picked and plotted side-by-side with PEmax for each specific edit. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. Individual editing efficiencies and indel levels are shown in FIGs.12D-12E. FIG.5F shows prime editing efficiencies of PE6b and PE6c normalized to the editing efficiency of PEmax at 77 edits that install a pathogenic allele into endogenous sites in HEK293T cells. No B1195.70180WO00 12418099.1nicking gRNA was used and MLH1dn plasmid was simultaneously transfected with prime editor plasmid for all conditions. All values from n=3 replicates are shown. Lines reflect the mean across all edits and replicates. Prime editing efficiencies for edits where PE6b or PE6c outperformed PEmax by more than 1.5-fold are shown on the right. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. Prime editing efficiencies used are the frequency of the intended prime editing outcome with no indels or other changes at the target site. FIG.5G shows correction of pathogenic mutations implicated in Crigler- Najjar Syndrome, Bloom Syndrome, and Pompe disease in HEK293T cell models using PEmax, PE6b, and PE6c. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.5H shows correction of mutations implicated in Crigler- Najjar Syndrome (UGT1A1) and Bloom Syndrome (RECQL3) in patient-derived fibroblast using PE6c and PEmax. Bars reflect the mean of n=3 independent replicates for treated samples and n=1-3 replicates of an untreated control for editing (red) and indels (gray). Dots show individual replicate values.

[0044] FIGs.6A-6G: Evolution and engineering of improved Cas9 domains for prime editing, and summary of PE6 use. FIG.6A shows a summary of evolution campaigns for whole PE2 phage. Shading indicates which circuit (v1 in yellow, v2 in blue, and v3 in purple) an evolution was performed in. Green shading indicates reversion analysis. Whether a given evolution was PANCE or PACE is specified: the number in parentheses after a PANCE or PACE label specifies how many passages of PANCE (p) were performed or how many hours of PACE (h) were performed. Arrowheads indicate that an evolution was stopped and increased in stringency without mammalian characterization; mutants characterized in mammalian cells are denoted with a dot and labeled. Finally, evolutions that utilized extra manipulations to increase stringency are labeled in pink, reflecting either a change in the PBS or a change in the expression of the target T7 RNAP gene. FIG.6B shows an evaluation of PACE-evolved clones in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. EvoCas9-1 through evoCas9-4 were isolated from low-stringency evolution. EvoCas9-5 and evoCas9-6 were isolated from high- stringency evolution. FIG.6C shows an assessment of individual Cas9 mutations on prime editing efficiency at two test sites. The y-axis shows editing efficiency at the Pcsk9 +3 C to G and +6 G to C edit in N2a cells. The x-axis shows editing efficiency for the RNF2 +5 G to T edit in HEK293T cells. Mutants incorporated into final Cas9 variants are shown in green. Mutants previously shown to, or structurally predicted to, decrease Cas9 binding are shown in maroon. PEmax∆RNaseH is shown in orange FIG.6D shows a comparison of combined B1195.70180WO00 12418099.1Cas9 mutants to PEmax∆RNaseH in HEK293T cells and N2a cells. Editing efficiencies of variants are normalized to the editing efficiency generated by PEmax∆RNaseH. Individual replicates are plotted, with n=3 biological replicates per edit. FIG.6E shows a comparison of PEmax, PE6a, and PE6a / e at two sites in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.6F shows a comparison of PEmax∆RNaseH, PE6c, and PE6g in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.6G shows a decision tree for selecting a PE6 variant.

[0045] FIGs.7A-7C: PE6 variants enable new classes of in vivo prime editing. FIG.7A is a schematic showing a dual-AAV delivery system for twinPE (v3em twinPE-AAV). In the N- terminal AAV, production of the N-terminal portion of Cas9 (yellow) fused to an N-terminal Npu split intein (orange) is regulated by the Cbh promoter (green) and the SV40 late polA signal (tan). In the C-terminal AAV, the C-terminal Npu split intein (dark green) is fused to the remainder of the prime editor (Cas9, yellow and RT, purple). The SV40 late polyA signal (tan), two epegRNAs (light and dark blue, AAV ITRs (black) are also shown. FIG.7B shows an injection route and twinPE editing efficiency of PEmax∆RNaseH and PE6d viruses in the for the twinPE-mediated insertion of a 38-bp attB sequence at murine Rosa26 in the mouse cortex. N- and C- terminal twinPE viruses are administered via ICV injection (4x1010vg total) along with a GFP-KASH virus. Editing efficiencies (light and dark blue) and indel frequencies (black and grey) are shown to the right. Bars reflect the mean of n=3-4 mice. Dots show individual mice. FIG.7C shows an injection route and PE editing efficiency of PEmax∆RNaseH and PE6d viruses for the installation of a 42-bp insertion containing loxP at the Dnmt1 locus in the mouse cortex. (Left) The C-terminal virus is modified to include one epegRNA and one nicking sgRNA to encode a PE edit as opposed to a twinPE edit. (Right) Editing efficiencies (light / dark pink) and indel rates (black / grey). Bars reflect the mean of n=3 mice. Dots show individual mice.

[0046] FIGs.8A-8J: Characterization and engineering of reverse transcriptase enzymes for prime editing, related to FIGs.1A-1J. FIG.8A show that native small RT enzymes demonstrate poor activity in the prime editing system (HEK293T cells, HEK3 +5 G to T edit). RT enzymes engineered in FIGs.1A-1J are highlighted in green, and the WT M-MLV RT used in the PE1 system is highlighted in black. All other enzymes are in red. Dots reflect the mean of n=3 independent replicates. FIG.8B shows an overview of twinPE. The prime editor protein (grey and blue) uses two pegRNAs (dark blue and teal) to target opposite B1195.70180WO00 12418099.1strands of DNA. The prime editor generates two 3′ flaps (red) that are complementary to each other. After these newly synthesized 3′ flaps anneal and the original DNA sequence in the 5′ flaps is degraded, the edited sequence in the flaps is permanently installed at the target DNA site. FIG.8C shows incorporation of each of the five mutations analogous to those in PE2 improves the activity of four retroviral RT enzymes in HEK293T cells. PERV = porcine endogenous retrovirus RT, AVIRE = avian reticuloendotheliosis virus RT, KORV = koala retrovirus RT and WMSV = woolly monkey sarcoma virus RT. Combining all five mutations together (Penta) further improves the activity of each enzyme. All values from n=3 independent replicates are shown. Horizontal bars show the mean value. FIG.8D shows structure-guided rational engineering of the Tf1 RT identifies five mutations that improve prime editing in HEK293T cells. The solved structure of the Tf1 RT homolog, Ty3 RT, was used to predict mutations that could increase contacts of the RT with its DNA-RNA substrate (PDB: 4OL8). All values from n=3 independent replicates are shown. Horizontal bars show the mean value across all sites and replicates. FIG.8E shows combining all mutations identified from structure-guided rational engineering improves the activity of the Tf1 RT prime editor in HEK293T cells. The final rationally designed Tf1 variant (rdTf1) is a combination of five mutations: K118R, S188K, I260L, R288Q, and S297Q. All values from n=3 independent replicates are shown. Horizontal bars show the mean value. FIG.8F shows an AlphaFold-predicted structure of the Ec48 RT enzyme. FIG.8G shows that aligning the AlphaFold-predicted structure of the Ec48 RT (blue) with the RT from xenotropic murine leukemia virus-related virus (XMRV, PDB 4HKQ, yellow), a close relative of the M-MLV RT, suggests that the residue analogous to the D200 residue in M-MLV RT is the T189 residue in Ec48 RT. FIG.8H shows structure-guided rational engineering of the Ec48 RT identifies six mutations that improve prime editing. An AlphaFold-generated predicted structure of the Ec48 RT was overlayed with the structure of the RT from the xenotropic murine leukemia virus-related virus (XMRV) (PDB: 4HKQ) to perform structure-guided mutagenesis. All values from n=3 independent replicates are shown. Horizontal bars show the mean value. FIG.8I shows the positions of residues (red) proximal to the substrate that were mutated to improve the activity of the Ec48 RT prime editor. Residues are mapped onto the predicted AlphaFold structure of the Ec48 RT aligned with the solved substrate of the XMRV RT (PDB: 4HKQ). L182 and T385 are proximal to the DNA substrate (green), R315 and K307 are proximal to the RNA substrate (yellow) and R378 is proximal to both the DNA and RNA rate. FIG.8J shows that combining the top three mutations identified from structure-guided engineering improves the activity of the Ec48 RT prime editor in HEK293T B1195.70180WO00 12418099.1cells. The final rationally designed Ec48 RT variant (rdEc48) contains three mutations: L182N, T189N, and R315K. All values from n=3 independent replicates are shown. Horizontal bars show the mean value.

[0047] FIGs.9A-9F: Design and validation of a PE PACE circuit, related to FIGs.2A-2K. FIG.9A shows a summary of phage-assisted continuous evolution (PACE). Host E. coli (grey) harboring relevant selection circuit plasmids (green, pink, and orange) and the mutagenesis plasmid (MP, black) continuously flow into a fixed-volume lagoon (left). Addition of arabinose induces expression of mutagenic genes on the MP. Selection phage (blue) harboring an NpuC-RT transgene (purple) infect the E. coli and are mutagenized. If a mutagenized RT is inactive (red, bottom / right), then prime editing does not trigger gIII expression and pIII production, and phage are not able to propagate. These phage encoding inactive RTs are washed out of the lagoon by continuous flow. If a mutagenized RT is active (green, center), then prime editing leads to pIII production, and phage encoding that RT can propagate faster than the rate at which they are diluted out of the lagoon. FIG.9B shows a summary of phage-assisted non-continuous evolution (PANCE). The same principles shown in FIG.9A are used in PANCE, except periodic discrete dilution steps instead of continuous flow is used to dilute selection cultures. Mid-log-phase cultures of selection E. coli are infected with phage, and arabinose is added to induce mutagenesis (left). After an overnight incubation, cultures are centrifuged to pellet bacteria and allow isolation of propagating phage from the supernatant (middle). A small volume of supernatant (typically a 1:50 dilution factor) is used to infect a fresh lagoon of mid-log selection strains (right). This process is iterated until phage titers stabilize (i.e., when overnight phage propagation is equal to or greater than the dilution factor). FIG.9C shows the effect of pegRNA optimization on PE2 phage propagation. Overnight propagation of empty phage (native control, red), PE2 phage (purple), and T7 RNAP phage (positive control, green) in strains harboring pegRNAs of different PBS and RTT lengths. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. This data was used to generate FIG.2C. FIG.9D shows a luciferase assay to screen pegRNAs for the v2 PE PACE circuit. Selection strains encoding luxAB transcriptionally coupled to gIII were infected with either empty phage (red) or PE2 phage (purple).4 h after infection, OD600-normalized luminescence was measured as a proxy for circuit activation. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. Strains in which PE2 phage outperformed empty phage were used for v2 evolutions. FIG.9E shows overnight propagation of pools of wild-type RT and evolved RT phage on their cognate or noncognate host-cell selection strains. Additional B1195.70180WO00 12418099.1evolved pools of phage are shown here beyond those provided in FIG.2K. Phage were from PANCE on the v1 circuit (yellow bars), from PANCE on the v2 circuit (blue bars), or wild- type-PE2 phage (grey bars). Propagation was then measured in the v1 circuit (left) or the v2 circuit (right). Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.9F shows the design of v3 circuit and improvements compared to v1 and v2 designs. A long insertion edit (20-bp insertion edit with a 60-bp RTT) was used to select for high-processivity, high-activity prime editors. Unlike v1 and v2 circuits, the v3 pegRNA (grey) targets the noncoding strand of T7 RNAP; this shortens the time between prime editing and wild type T7 RNAP production. In addition to the 20-bp insertion (green) needed to restore the frame of T7 RNAP, the v3 pegRNA also encodes silent PAM edits (maroon) and a seed edit (blue) that prevents subsequent binding and nicking of the edited sequence. From top to bottom, the sequences are SEQ ID NOs: 140 and 141.

[0048] FIGs.10A-10F: Evolution and characterization of compact RTs for prime editing, related to FIGs.3A-3H. FIG.10A shows overnight propagation of phage encoding dead M- MLV RT (red), Gs (blue), or PE2 (purple) RTs in the NpuC-RT phage architecture in the pegRNA-optimized v1 PE PACE circuit. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.10B shows phage titers during PANCE of NpuC- Gs-RT phage. Grey shading indicates a passage of evolutionary drift, in which phage were supplied gIII in the absence of selection to allow free mutagenic replication. Titers of four replicate lagoons are shown. FIG.10C shows PACE of NpuC-Gs-RT phage. The left y-axis and pink and blue lines show the SP titer of three different replicate lagoons at various timepoints. The right y-axis and dotted grey line show the flow rate in volumes per hour. FIG.10D shows indel frequencies for prime editors in the optimized PEmax architecture containing either engineered pentamutant Marathon RT (Marathon penta, red), evoEc48 (blue), or evoTf1 (yellow) with PEmax (gray) in HEK293T cells. Editing frequencies corresponding to this data is in FIG.3F. Bars reflect the mean of three independent replicates. Dots show individual replicate values. FIG.10E shows performance of PE6a and PE6b in the presence and absence of epegRNAs in HEK293T cells. All values from n=3 independent replicates are shown. Horizontal bars show the mean value. FIG.10F shows a comparison of PE6a, PE6b, and PEmax at three longer, complex edits in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values.

[0049] FIGs.11A-11G: Development and characterization of highly processive, dual AAV- compatible RTs. FIG.11A shows editing efficiencies of prime editors containing single M- MLV mutants in HEK293T cells. Prime editing efficiencies used are the frequency of the B1195.70180WO00 12418099.1intended prime editing outcome with no indels or other changes at the target site. Lines reflect the mean of n=2 independent replicates per edit. Dots show individual replicate values. FIG.11B shows an overview of the terminal deoxynucleotidyl transferase (TdT) assay for sequencing newly reverse-transcribed DNA flaps that have not been incorporated into the genome. Shortly after treatment with a prime editor and pegRNA, cells are lysed, and DNA is purified. A terminal transferase enzyme (yellow) adds a polyG sequence to all DNA 3′ ends. PCR amplification for high-throughput DNA sequencing is performed using a locus- specific forward primer and a polyC reverse primer. FIG.11C shows results of a TdT assay on the HEK3 +1 FLAG insertion edit in HEK293T cells. The y-axis indicates the percentage of total RT products of a given length, and the x-axis represents the length of the product in base pairs. PEmax∆RNaseH is shown in grey, and PE6d is shown in blue. The lines are mean values from n=3 biological replicates. FIG.11D shows editing efficiencies of PE6b-d, PEmax, and PEmax∆RNaseH for edits engineered to contain varying levels of secondary structure. “UC” indicates an unpinned control for a corresponding hairpin edit. These values were used to generate the free energy vs fold improvement plot in FIG.4G. All edits are in HEK293T cells. Individual replicates are shown, with n=3 replicates per condition. FIG.11E shows editing efficiencies (left) and indel rates (right) of PE6d and PEmax∆RNaseH for a series of prime edits that use short unstructured pegRNAs in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. PEmax∆RNaseH is shown on the right for each edit site, and PE6d is shown on the left for each edit site. FIG. 11F shows results of a TdT assay on the RNF2 +5 G to T edit in HEK293T cells. Note that the x-axis differs from other TdT plots shown herein: instead of RTT-templated bases correctly installed, it quantifies the number of sgRNA scaffold-templated bases aberrantly installed (for example, x=1 indicates the addition of one extra scaffold-templated base). The y-axis indicates the percentage of edit-containing flaps that have a given number of scaffold- templated bases. For each prime editor, the line reflects the mean of n=3 independent replicates. Pie charts indicate the percentages of edit-containing flaps that either have ≤2 bp (solid color) or >2 bp (striped) of scaffold-templated bases. Data shown are the mean of three independent biological replicates. FIG.11G shows unique molecular identifier (UMI) analysis of prime editing efficiencies for twinPE edits in N2a cells (left) and HEK293T cells (middle, right). UMI protocol was applied to remove PCR bias, and trends agree with the data shown in FIGs.4A-4J. Bars reflect the mean of n=3 independent replicates. Dots show B1195.70180WO00 12418099.1individual replicate values. For each instance of “Edit” and “Indels,” PEmax∆RNaseH is shown on the left, PE6c is shown in the middle, and PE6d is shown on the right.

[0050] FIGs.12A-12J: Comparison of PE6 variants with PEmax, related to FIGs.5A-5H. FIG.12A shows prime editing efficiencies of the best performing PE6 variant (either PE6c or PE6d) normalized to the editing efficiency of PEmax at sites tested in FIG.5A. All values from n=3 independent replicates are shown. Editing was performed in HEK293T cells. The horizontal bar shows the mean value. FIG.12B shows indel frequencies of PEmax, PE6c, and PE6d at edits tested in FIG.5A. This data was used for FIG.5B. Bars reflect the mean of three independent replicates. Editing was performed in HEK293T cells. Dots show individual replicate values. FIG.12C shows screening PE6 variants for insertion of attB into the CCR5 locus in primary human T cells. Bars reflect the mean of n=4 independent replicates for editing (red) and indels (grey). Dots show individual replicate values. FIG.12D shows absolute prime editing efficiencies of PE6 variants, PEmax∆RNaseH, and PEmax in HEK293T cells used to plot data for FIGs.5D-5E. Prime editing efficiencies used are the frequency of the intended prime editing outcome with no indels or other changes at the target site. Bars reflect the mean of three independent replicates. Dots show individual replicate values. FIG.12E shows indel frequencies of PE6 variants, PEmax∆RNaseH, and PEmax in HEK293T cells used to plot data for FIGs.5D-5E. Bars reflect the mean of three independent replicates. Dots show individual replicate values. FIG.12F shows a percentage of sequencing reads containing a pegRNA scaffold insertion after prime editing using PE6 variants, PEmax∆RNaseH, and PEmax in HEK293T cells. These reads contribute to the total indel frequency. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.12G shows prime editing efficiencies for edits where PE6b or PE6c outperformed PEmax using a nicking gRNA. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. Prime editing efficiencies used are the frequency of the intended prime editing outcome with no indels or other changes at the target site in HEK293T cells. FIG.12H shows indel frequencies of PE6 variant and PEmax at sites shown in FIG.5F in HEK293T cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.12I shows a correction of mutation implicated in Pompe disease in patient-derived fibroblast using PE6c and PEmax. Bars reflect the mean of n=3 independent replicates for editing (red) and indels (grey). Dots show individual replicate values. FIG.12J shows a distribution of editing outcomes after correction of the pathogenic mutation implicated in Pompe disease in patient-derived fibroblasts using PE6c. The patient B1195.70180WO00 12418099.1was heterozygous. Indel genotypes are shown. From top to bottom, the sequences are SEQ ID NOs: 142, 143, 144, 144, 142, 144, 31, 143, and 144.

[0051] FIGs.13A-13F: Evolution and engineering of Cas9 mutants for PE, related to FIGs. 6A-6G. FIG.13A shows a representative PACE campaign for the v1 circuit. Different colored lines represent different replicate lagoons. PACE experiments with less than four lagoons shown experienced cheating (activity-independent phage propagation likely from rare gene III recombination onto the SP) or washout (complete loss of viable phage) for one or more lagoons. Top graphs represent the phage titer over a PACE experiment. Bottom graphs show the flow rate at the corresponding time. FIG.13B shows a reversion analysis of EvoCas9-4 in HEK293T cells. Editing efficiency was normalized to the values obtained using PE2. Data are shown as individual data points for n=3 biological replicates and as the grand mean across the four sites tested. FIG.13C shows a structural analysis of mutations that harm mammalian prime editing activity. (Left) Structure (PDB: 4UN3) of wild-type Sp Cas9 (grey) bound to its guide RNA (purple) and DNA substrate (yellow / orange). Residue K1151 is shown in dark pink. (Right) Structure (PDB: 4OO8) of wild-type Sp Cas9 (grey) bound to its guide RNA (purple) and DNA substrate (orange). Wild-type residues K1003, K1014, and A1034 are shown in dark pink. FIG.13D shows a circuit for examining editing- independent effects of the prime editor on the PE PACE circuit. E. coli harbor a corrected T7 RNAP, luxAB gene under the T7 promoter, and the pegRNA used during selection. Prime editor variants are introduced on a plasmid under the control of an arabinose-inducible promoter. After induction, OD-normalized luminescence for n=3 biological replicates were used to measure circuit turn on. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. FIG.13E shows prime editing efficiencies N2a cells (left, Ctnnb1 through Pcks9) and HEK293T cells (right, CXCR4 through RNF2) used to generate the fold changes reported in FIG.6D. Individual replicates are plotted, with n=3 biological replicates per edit. FIG.13F show the structure (PDB: 4UN3) of Cas9 (grey) bound to its sgRNA (purple). Residue H721, which is mutated to Tyr in evolutions, is shown in green sticks. Dotted lines denote predicted polar contacts between H721 and other atoms.

[0052] FIGs.14A-14E: In vivo prime editing with PE6c and PE6d delivered via dual AAV, related to FIGs.7A-7C. FIG.14A shows an analysis of truncated PE6c variants. Editing (yellow) and indels (grey) are shown for the installation of an attB sequence at the murine Rosa26 locus in N2a cells. Bars reflect the mean of n=3 independent replicates. Dots show individual replicate values. The number below each variant indicates the number of DNA bases that have been deleted from the C-terminal end of the Tf1 gene. FIG.14B shows B1195.70180WO00 12418099.1representative flow plots for the isolation of unsorted and sorted nuclei from mouse cortices. Left: scatter plot of all events, gate A set to collect nuclei. Middle: selection of single-nuclei droplets in Gate B, Right: FITC signal was used to collect unsorted cells (Gate C) and transduced, GFP-positive cells (Gate D). FIG.14C shows twinPE editing efficiency of PEmax∆RNaseH and PE6c viruses in the mouse cortex. N- and C- terminal twinPE viruses are administered via ICV injection (4x1010vg total) along with a GFP-KASH virus. Editing efficiencies (light and dark blue) and indel (black / grey) rates are shown to the right. Bars reflect the mean of n=3-4 mice. Dots show individual mice. FIG.14D shows injection route and PE editing (Dnmt1 loxP insertion) efficiency of PEmax∆RNaseH and PE6d viruses at a low viral dose (2 x1010vg total) in the mouse cortex. (Left) The C-terminal virus is modified to include one epegRNA and one nicking sgRNA to encode a PE edit as opposed to a twinPE edit. (Right) Editing efficiencies (light / dark pink) and indel rates (black / grey). Bars reflect the mean of n=3 mice. Dots show individual mice. FIG.14E shows off-target editing from AAV-treated and untreated mice. Bars reflect the mean of n=3 mice. Dots show individual mice. PE6d bulk (light pink) and transduced (dark pink) values were either less than 0.1% on average or were not statistically significant from untreated controls (light grey). For both ns notes, p = 0.08. Analyses were performed with an unpaired t test with Welch correction. The y-axis indicates off-target editing and indels summed (see methods for calculation). OT6 failed to amplify. All treated samples are from the high AAV dose condition.

[0053] FIGs.15A-15B: Mutation tables from v1 PE PA(N)CE Gs evolution, related to FIGs. 3A-3H. Mutations in clones emerging from evolution. Lagoon 2 cheated during PACE and was not sequenced. Silent mutations omitted for clarity.

[0054] FIG.16: Mutation table from v1 PE PANCE Tf1 evolution, related to FIGs.3A-3H. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0055] FIG.17: Mutation table from v1 PE PANCE Ec48 evolution, related to FIGs.3A-3H. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0056] FIG.18: Mutation table from v1 PE PANCE Vc95 evolution, related to FIGs.3A-3H. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0057] FIGs.19A-19J: Mutation tables from v1 whole editor evolution, related to FIGs.4A- 4G. Mutations in clones emerging from evolution. Blue shading indicates an amino acid change, and orange shading indicates a truncating mutation, either a stop codon (*) or a frameshift (FS). Silent mutations omitted for clarity. B1195.70180WO00 12418099.1

[0058] FIG.20: Mutation table from v1 and v2 comparative PE PANCE whole editor evolution, related to FIGs.4A-4G. Mutations in clones emerging from evolution. Mutations from v1 are shaded in yellow, and mutations from v2 evolution are shaded in blue. Silent mutations omitted for clarity. Frameshift mutations are not shown.

[0059] FIG.21: Mutation table from v2 PE PANCE Gs evolution, related to FIGs.4A-4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0060] FIG.22: Mutation table from v2 PE PANCE Ec48 evolution, related to FIGs.4A-4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0061] FIG.23: Mutation table from v2 high stringency PE PANCE Ec48 evolution, related to FIGs.4A-4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0062] FIG 24: Mutation table from v2 PE PANCE Tf1 evolution, related to FIGs.4A-4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0063] FIG.25: Mutation table from v3 PE PANCE Tf1 evolution, related to FIGs.4A-4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0064] FIG.26: Mutation table from v1 PE PACE PE2 RT evolution, related to FIGs.4A- 4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0065] FIG.27: Mutation table from v2 PE PANCE PE2 RT evolution, related to FIGs.4A- 4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0066] FIGs.28A-28C: Mutation tables from v3 PE PANCE PE2 RT evolution, related to FIGs.4A-4G. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0067] FIGs.29A-29F: Mutation tables from v1-v3 whole editor evolutions, related to FIGs. 6A-6F. Mutations in clones emerging from evolution. Silent mutations omitted for clarity.

[0068] FIG.30 shows neonatal cerebroventricular (P0 ICV) injections of dual AAV delivering PEmax∆RNaseH or PE6 variants to the CNS of 12 mice (n=4 for each of three groups).

[0069] FIG.31 shows in vivo liver editing after ICV injection. PE6 editors substantially enhance liver editing (30% editing in the liver after ICV injection). DEFINITIONS

[0070] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms B1195.70180WO00 12418099.1used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise. Adeno-Associated Virus (AAV)

[0071] An “adeno-associated virus” or “AAV” is a virus that infects humans and some other primate species. The wild-type AAV genome is a single-stranded deoxyribonucleic acid (ssDNA), either positive- or negative-sensed. The genome comprises two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs): rep and cap between the ITRs. The rep ORF comprises four overlapping genes encoding Rep proteins required for the AAV life cycle. The cap ORF comprises overlapping genes encoding capsid proteins: VP1, VP2, and VP3, which interact together to form the viral capsid. VP1, VP2, and VP3 are translated from one mRNA transcript, which can be spliced in two different manners: either a longer or shorter intron can be excised resulting in the formation of two isoforms of mRNAs: a ~2.3 kb- and a ~2.6 kb-long mRNA isoform. The capsid forms a supramolecular assembly of approximately 60 individual capsid protein subunits into a non-enveloped, T-1 icosahedral lattice capable of protecting the AAV genome. The mature capsid is composed of VP1, VP2, and VP3 (molecular masses of approximately 87, 73, and 62 kDa respectively) in a ratio of about 1:1:10.

[0072] Recombinant AAV (rAAV) particles may comprise a nucleic acid vector (e.g., a recombinant genome), which may comprise at a minimum: (a) one or more heterologous nucleic acid regions comprising a sequence encoding a protein or polypeptide of interest (e.g., a split prime editor) or an RNA of interest (e.g., a gRNA), or one or more nucleic acid regions comprising a sequence encoding a Rep protein; and (b) one or more regions comprising inverted terminal repeat (ITR) sequences (e.g., wild-type ITR sequences or engineered ITR sequences) flanking the one or more nucleic acid regions (e.g., heterologous nucleic acid regions). In some embodiments, the nucleic acid vector is between 4 kb and 5 kb in size (e.g., 4.2 to 4.7 kb in size). In some embodiments, the nucleic acid vector further comprises a region encoding a Rep protein. In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, a double- stranded nucleic acid vector may be, for example, a self-complimentary vector that contains a B1195.70180WO00 12418099.1region of the nucleic acid vector that is complementary to another region of the nucleic acid vector, initiating the formation of the double-strandedness of the nucleic acid vector.

[0073] In some embodiments, an AAV is used to deliver any of the reverse transcriptase variants, Cas9 variants, fusion proteins, prime editors, and / or polynucleotides or vectors encoding the same. Cas9

[0074] The term “Cas9” or “Cas9 nuclease” refers to an RNA-guided nuclease comprising a Cas9 domain, or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or the gRNA binding domain of Cas9). A “Cas9 domain,” as used herein, is a protein fragment comprising an active or fully or partly inactive cleavage domain of Cas9 and / or the gRNA binding domain of Cas9. A “Cas9 protein” is a full length Cas9 protein. A Cas9 nuclease is also referred to sometimes as a casn1 nuclease or a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves a linear or circular dsDNA target complementary to the spacer. The strand in the target DNA not complementary to crRNA is first cut endonucleolytically, then trimmed 3′-5′ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gRNA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the contents of which are incorporated herein by reference. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self versus non-self. Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. B1195.70180WO00 12418099.198:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816- 821(2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, a Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0075] A nuclease-inactivated Cas9 domain may interchangeably be referred to as a “dCas9” protein (for nuclease-“dead” Cas9). Methods for generating a Cas9 domain (or a fragment thereof) having an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell.28;152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, whereas the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science.337:816-821(2012); Qi et al., Cell.28;152(5):1173-83 (2013)). In some embodiments, proteins comprising fragments of a Cas9 protein are provided. For example, in some embodiments, a protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9, or fragments thereof, are referred to as “Cas9 variants.” A Cas9 variant shares homology to Cas9, or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% B1195.70180WO00 12418099.1identical, or at least about 99.9% identical to wild type Cas9 (e.g., SpCas9 of SEQ ID NO: 6). In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild type Cas9 (e.g., SpCas9 of SEQ ID NO: 6). In some embodiments, the Cas9 variant comprises a fragment of SEQ ID NO: 6 Cas9 (e.g., a gRNA binding domain or a DNA-cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild type Cas9 (e.g., SpCas9 of SEQ ID NO: 6). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild type Cas9 (e.g., SpCas9 of SEQ ID NO: 6).

[0076] In some embodiments, a Cas9 protein comprises any of the amino acid substitutions described herein. In certain embodiments, a Cas9 protein comprises the amino acid substitutions K775R and K918A relative to wild type Streptococcus pyogenes Cas9 or relative to Streptococcus pyogenes Cas9 nickase (SEQ ID NO: 2). In certain embodiments, a Cas9 protein comprises the amino acid substitutions H99R, E471K, I632V, D645N, R654C, and H721Y relative to wild type Streptococcus pyogenes Cas9 or relative to Streptococcus pyogenes Cas9 nickase (SEQ ID NO: 2). In certain embodiments, a Cas9 protein comprises the amino acid substitutions H99R, E471K, I632V, D645N, H721Y, and K918A relative to wild type Streptococcus pyogenes Cas9 or relative to Streptococcus pyogenes Cas9 nickase (SEQ ID NO: 2). CRISPR

[0077] CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of prior infections by a virus that have invaded the prokaryote. The snippets of DNA are used by the prokaryotic cell to detect and destroy DNA from subsequent attacks by similar viruses and effectively compose, along with an array of CRISPR- associated proteins (including Cas9 and homologs thereof) and CRISPR-associated RNA, a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of pre-crRNA requires a trans-encoded small RNA B1195.70180WO00 12418099.1(tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves a linear or circular dsDNA target complementary to the RNA. Specifically, the DNA strand in the target that is not complementary to crRNA is first cut endonucleolytically, then trimmed 3′-5′ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gRNA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species – the guide RNA. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the entire contents of which is hereby incorporated by reference. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self versus non-self. CRISPR biology, as well as Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A.98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816- 821(2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.

[0078] In general, a “CRISPR system” refers collectively to transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated (“Cas”) genes, including sequences encoding a Cas gene, a tracr (trans-activating CRISPR) sequence (e.g., tracrRNA or an active partial tracrRNA), a tracr mate sequence (encompassing a “direct B1195.70180WO00 12418099.1repeat” and a tracrRNA-processed partial direct repeat in the context of an endogenous CRISPR system), a guide sequence (also referred to as a “spacer”), or other sequences and transcripts from a CRISPR locus. The tracrRNA of the system is complementary (fully or partially) to the tracr mate sequence present on the guide RNA. Edit strand and non-edit strand

[0079] The terms “edit strand” and “non-edit strand” are terms that may be used when describing the mechanism of a prime editing system on a double-stranded DNA substrate. The “edit strand” refers to the strand of DNA that is nicked by the prime editor complex to form a 3ʹ end, which is then extended as a newly synthesized single stranded DNA (also referred to herein as the newly synthesized 3′ DNA flap), which comprises a desired edit and ultimately displaces and replaces the single strand region of DNA just downstream of the nick, thereby installing the 3ʹ DNA flap containing the desired edit downstream of the nick on the “edit strand.” In some embodiments, the newly synthesized 3′ DNA flap comprising the nucleotide edit is paired in a heteroduplex with the non-edit strand that does not comprise the nucleotide edit, thereby creating a mismatch. In some embodiments, the mismatch is recognized by DNA repair machinery, and / or replication machinery, e.g., an endogenous DNA repair machinery. In some embodiments, through DNA repair, the intended nucleotide edit is incorporated into both strands of the target double-stranded DNA substrate. The application may also refer to the “edit strand” as the “protospacer strand” or the “PAM strand” since these elements are present in that strand. The “edit strand” may also be called the “non-target strand” since the edit strand is not the strand that becomes annealed to the spacer of the pegRNA molecule, but rather is the complement of the strand that is annealed by the spacer of the pegRNA. The “non-edit” strand is not directly edited by the PE system. Rather, the desired edit created by the PE system in the 3ʹ DNA flap is incorporated into the “non-edited strand” through DNA replication and / or repair. In some embodiments, the “non- edit strand” is the strand that anneals to the spacer of the pegRNA, and thus is also called the “target strand.” Fusion protein

[0080] The term “fusion protein” as used herein refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C- terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid-programmable DNA-binding domain (e.g., the gRNA binding domain of Cas9 B1195.70180WO00 12418099.1that directs the binding of the protein to a target site) and a reverse transcriptase (i.e., a prime editor). In some embodiments, a fusion protein comprises any of the reverse transcriptase variants provided herein fused to at least one other domain (e.g., a Cas9 protein (such as a Cas9 variant provided herein), an NLS, or any other domain disclosed herein). In some embodiments, a fusion protein comprises any of the Cas9 variants provided herein fused to at least one other domain (e.g., a reverse transcriptase (such as a reverse transcriptase variant provided herein), an NLS, or any other domain disclosed herein). In certain embodiments, a fusion protein comprises any of the reverse transcriptase variants provided herein and any of the Cas9 variants provided herein. Any of the fusion proteins provided herein may be produced by any method known in the art. For example, the prime editor fusion proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), which is incorporated herein by reference. Genetic Elements of AAV Particle Vectors

[0081] Nucleic acids of the present disclosure (e.g., nucleic acids delivered by an AAV particle as described herein) may include one or more genetic elements. A “genetic element” refers to a particular nucleotide sequence that has a role in nucleic acid expression (e.g., promoter, enhancer, terminator) or encodes a discrete product of an engineered nucleic acid (e.g., a nucleotide sequence encoding a guide RNA and / or a protein).

[0082] A “promoter” refers to a control region of a nucleic acid sequence at which initiation and rate of transcription of the remainder of a nucleic acid sequence are controlled. A promoter may also contain sub-regions at which regulatory proteins and molecules may bind, such as RNA polymerase and other transcription factors. Promoters may be constitutive, inducible, activatable, repressible, tissue-specific, or any combination thereof. A promoter drives expression or drives transcription of the nucleic acid sequence that it regulates. Herein, a promoter is considered to be “operably linked” when it is in a correct functional location and orientation in relation to a nucleic acid sequence it regulates to control (“drive”) transcriptional initiation and / or expression of that sequence.

[0083] A promoter may be one naturally associated with a gene or sequence, as may be obtained by isolating the 5' non-coding sequences located upstream of the coding segment of a given gene or sequence. Such a promoter is referred to as an “endogenous promoter.” In B1195.70180WO00 12418099.1some embodiments, a coding nucleic acid sequence may be positioned under the control of a recombinant or heterologous promoter, which refers to a promoter that is not normally associated with the encoded sequence in its natural environment. Such promoters may include promoters of other genes; promoters isolated from any other cell; and synthetic promoters or enhancers that are not “naturally occurring” such as, for example, those that contain different elements of different transcriptional regulatory regions and / or mutations that alter expression through methods of genetic engineering that are known in the art. In addition to producing nucleic acid sequences of promoters and enhancers synthetically, sequences may be produced using recombinant cloning and / or nucleic acid amplification technology, including polymerase chain reaction (PCR).

[0084] In some embodiments, promoters used in accordance with the present disclosure are “inducible promoters,” which are promoters that are characterized by regulating (e.g., initiating or activating) transcriptional activity when in the presence of, influenced by or contacted by an inducer signal. An inducer signal may be endogenous or a normally exogenous condition (e.g., light), compound (e.g., chemical or non-chemical compound), or protein that contacts an inducible promoter in such a way as to be active in regulating transcriptional activity from the inducible promoter. Thus, a “signal that regulates transcription” of a nucleic acid refers to an inducer signal that acts on an inducible promoter. A signal that regulates transcription may activate or inactivate transcription, depending on the regulatory system used. Activation of transcription may involve directly acting on a promoter to drive transcription or indirectly acting on a promoter by inactivation of a repressor that is preventing the promoter from driving transcription. Conversely, deactivation of transcription may involve directly acting on a promoter to prevent transcription or indirectly acting on a promoter by activating a repressor that then acts on the promoter.

[0085] A “transcriptional terminator” is a nucleic acid sequence that causes transcription to stop. A transcriptional terminator may be unidirectional or bidirectional. It is comprised of a DNA sequence involved in specific termination of an RNA transcript by an RNA polymerase. A transcriptional terminator sequence prevents transcriptional activation of downstream nucleic acid sequences by upstream promoters. A transcriptional terminator may be necessary in vivo to achieve desirable expression levels or to avoid transcription of certain sequences. A transcriptional terminator is considered to be “operably linked to” a nucleotide sequence when it is able to terminate the transcription of the sequence it is linked to.

[0086] The most commonly used type of terminator is a forward terminator. When placed downstream of a nucleic acid sequence that is usually transcribed, a forward transcriptional B1195.70180WO00 12418099.1terminator will cause transcription to abort. In some embodiments, bidirectional transcriptional terminators are provided, which usually cause transcription to terminate on both the forward and reverse strand. In some embodiments, reverse transcriptional terminators are provided, which usually terminate transcription on the reverse strand only.

[0087] In prokaryotic systems, terminators usually fall into two categories (1) rho- independent terminators and (2) rho-dependent terminators. Rho-independent terminators are generally composed of a palindromic sequence that forms a stem loop rich in G-C base pairs followed by several T bases. Without wishing to be bound by theory, the conventional model of transcriptional termination is that the stem loop causes RNA polymerase to pause, and transcription of the poly-A tail causes the RNA:DNA duplex to unwind and dissociate from RNA polymerase.

[0088] In eukaryotic systems, the terminator region may comprise specific DNA sequences that permit site-specific cleavage of the new transcript so as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3' end of the transcript. RNA molecules modified with this polyA tail appear to more stable and are translated more efficiently. Thus, in some embodiments involving eukaryotes, a terminator may comprise a signal for the cleavage of the RNA. In some embodiments, the terminator signal promotes polyadenylation of the message. The terminator and / or polyadenylation site elements may serve to enhance output nucleic acid levels and / or to minimize read through between nucleic acids.

[0089] Terminators for use in accordance with the present disclosure include any terminator of transcription described herein or known to one of ordinary skill in the art. Examples of terminators include, without limitation, the termination sequences of genes such as, for example, the bovine growth hormone terminator, and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA, and arcA terminator. In some embodiments, the termination signal may be a sequence that cannot be transcribed or translated, such as those resulting from a sequence truncation. Linker

[0090] The term “linker,” as used herein, refers to a molecule linking two other molecules or moieties. The linker can be an amino acid sequence in the case of a peptide linker joining two domains of a fusion protein. For example, a napDNAbp (e.g., Cas9) can be fused to a reverse transcriptase by an amino acid linker sequence. The linker can also be a nucleotide sequence in the case of joining two nucleotide sequences together (e.g., in a gRNA). For example, in B1195.70180WO00 12418099.1the instant case, the traditional guide RNA is linked via a spacer or linker nucleotide sequence to the RNA extension of a prime editing guide RNA which may comprise an RT template sequence and an RT primer binding site. In other embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5- 200 amino acids in length, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. napDNAbp

[0091] As used herein, the term “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which Cas9 is an example, refers to a protein that uses RNA:DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic-acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to a complementary sequence.

[0092] Without being bound by theory, the binding mechanism of a napDNAbp–guide RNA complex, in general, includes the step of forming an R-loop whereby the napDNAbp induces the unwinding of a double-strand DNA target, thereby separating the strands in the region bound by the napDNAbp. The guide RNA protospacer then hybridizes to the “target strand.” This displaces a “non-target strand” that is complementary to the target strand, which forms the single strand region of the R-loop. In some embodiments, the napDNAbp includes one or more nuclease activities, which then cut the DNA, leaving various types of lesions. For example, the napDNAbp may comprise a nuclease activity that cuts the non-target strand at a first location, and / or cuts the target strand at a second location. Depending on the nuclease activity, the target DNA can be cut to form a “double-stranded break” whereby both strands are cut. In other embodiments, the target DNA can be cut at only a single site, i.e., the DNA is “nicked” on one strand. Exemplary napDNAbp with different nuclease activities include “Cas9 nickase” (“nCas9”) and a deactivated Cas9 having no nuclease activities (“dead Cas9” or “dCas9”). Exemplary sequences for these and other napDNAbp are provided herein. In some embodiments, a napDNAbp has nickase activity in a RuvC domain and / or an HNH domain. In some embodiments, a napDNAbp has nickase activity in a RuvC domain or an HNH domain. B1195.70180WO00 12418099.1Nickase

[0093] As used herein, a “nickase” refers to a napDNAbp (e.g., a Cas protein) which is capable of cleaving only one of the two complementary strands of a double-stranded target DNA sequence, thereby generating a nick in that strand. In some embodiments, the nickase cleaves a non-target strand of a double stranded target DNA sequence. In some embodiments, the nickase comprises an amino acid sequence with one or more mutations in a catalytic domain of a canonical napDNAbp (e.g., a Cas protein), wherein the one or more mutations reduces or abolishes nuclease activity of the catalytic domain. In some embodiments, the nickase is a Cas9 that comprises one or more mutations in a RuvC-like domain relative to a wild type Cas9 sequence or to an equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the nickase is a Cas9 that comprises one or more mutations in an HNH-like domain relative to a wild type Cas9 sequence or to an equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the nickase is a Cas9 that comprises an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 relative to a canonical SpCas9 sequence or to an equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the nickase is a Cas9 that comprises an H840A, N854A, and / or N863A mutation relative to a canonical SpCas9 sequence, or to an equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the term “Cas9 nickase” refers to a Cas9 with one of the two nuclease domains inactivated. This enzyme is capable of cleaving only one strand of a target DNA. In some embodiments, the nickase is a Cas protein that is not a Cas9 nickase.

[0094] In some embodiments, the napDNAbp of the prime editing complex comprises an endonuclease having nucleic acid programmable DNA binding ability. In some embodiments, the napDNAbp comprises an active endonuclease capable of cleaving both strands of a double stranded target DNA. In some embodiments, the napDNAbp is a nuclease active endonuclease, e.g., a nuclease active Cas protein, that can cleave both strands of a double stranded target DNA by generating a nick on each strand. For example, a nuclease active Cas protein can generate a cleavage (a nick) on each strand of a double stranded target DNA. In some embodiments, the two nicks on both strands are staggered nicks, for example, generated by a napDNAbp comprising a Cas12a or Cas12b1. In some embodiments, the two nicks on both strands are at the same genomic position, for example, generated by a napDNAbp comprising a nuclease active Cas9. In some embodiments, the napDNAbp comprises an endonuclease that is a nickase. For example, in some embodiments, the napDNAbp comprises an endonuclease comprising one or more mutations that reduce nuclease activity of B1195.70180WO00 12418099.1the endonuclease, rendering it a nickase. In some embodiments, the napDNAbp comprises an inactive endonuclease, for example, in some embodiments, the napDNAbp comprises an endonuclease comprising one or more mutations that abolish the nuclease activity. In various embodiments, the napDNAbp is a Cas9 protein or variant thereof. The napDNAbp can also be a nuclease active Cas9, a nuclease inactive Cas9 (dCas9), or a Cas9 nickase (nCas9). In a preferred embodiment, the napDNAbp is Cas9 nickase (nCas9) that nicks only a single strand. In other embodiments, the napDNAbp can be selected from the group consisting of: Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas12b2, Cas13a, Cas12c, Cas12d, Cas12e, Cas12h, Cas12i, Cas12g, Cas12f (Cas14), Cas12f1, Cas12j (CasΦ), and Argonaute and optionally has a nickase activity such that only one strand is cut. In some embodiments, the napDNAbp is selected from Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas12b2, Cas13a, Cas12c, Cas12d, Cas12e, Cas12h, Cas12i, Cas12g, Cas12f (Cas14), Cas12f1, Cas12j (CasΦ), and Argonaute and optionally has a nickase activity such that one DNA strand is cut preferentially to the other DNA strand. Nuclear localization sequence (NLS)

[0095] The term “nuclear localization sequence” or “NLS” refers to an amino acid sequence that promotes import of a protein into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and would be apparent to the skilled artisan. For example, NLS sequences are described in Plank et al., international PCT application, PCT / EP2000 / 011690, filed November 23, 2000, published as WO 2001 / 038547 on May 31, 2001, the contents of which are incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In some embodiments, an NLS is included in a fusion protein (e.g., in a prime editor as described herein). In certain embodiments, an NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 94), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 99), KRTADGSEFESPKKKRKV (SEQ ID NO: 97), KRTADGSEFEPKKKRKV (SEQ ID NO: 106), NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 107), PAAKRVKLD (SEQ ID NO: 98), RQRRNELKRSF (SEQ ID NO: 108), or NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 109). Nucleic acid

[0096] The term “nucleic acid,” as used herein, refers to a polymer of nucleotides. The polymer may include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl B1195.70180WO00 12418099.1adenosine, 5-methylcytidine, C5 bromouridine, C5 fluorouridine, C5 iodouridine, C5 propynyl uridine, C5 propynyl cytidine, C5 methylcytidine, 7-deazaadenosine, 7- deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, 4-acetylcytidine, 5- (carboxyhydroxymethyl)uridine, dihydrouridine, methylpseudouridine, 1-methyl adenosine, 1-methyl guanosine, N6-methyl adenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, 2′-O-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioates and 5ʹ N phosphoramidite linkages). In some embodiments, a nucleic acid is a pegRNA or an epegRNA. In some embodiments, a nucleic acid is a target nucleic acid to be editing, e.g., in a genome. PEgRNA

[0097] As used herein, the terms “prime editing guide RNA” or “PEgRNA” or “pegRNA” or “extended guide RNA” refer to a specialized form of a guide RNA that has been modified to include one or more additional sequences for implementing the prime editing as described herein. As described herein, the prime editing guide RNAs comprise one or more “extended regions,” also referred to herein as “extension arms,” of nucleic acid sequence. The extended regions may comprise, but are not limited to, single-stranded RNA or DNA. Further, the extended regions may occur at the 3′ end of a traditional guide RNA. In other arrangements, the extended regions may occur at the 5′ end of a traditional guide RNA. In still other arrangements, the extended region may occur at an intramolecular region of the traditional guide RNA, for example, in the gRNA core region which associates and / or binds to the napDNAbp. The extended region comprises a “DNA synthesis template” or “reverse transcriptase template” that encodes (by the polymerase / reverse transcriptase of the prime editor) a single-stranded DNA which, in turn, has been designed to be (a) homologous with the endogenous target DNA to be edited, and (b) which comprises at least one desired nucleotide change (e.g., a transition, a transversion, a deletion, or an insertion) to be introduced or integrated into the endogenous target DNA. The extended region may also comprise other functional sequence elements, such as, but not limited to, a “primer binding site” and a “linker” sequence, or other structural elements, such as, but not limited to, aptamers, stem loops, hairpins, toe-loops (e.g., a 3′ toeloop), or an RNA-protein recruitment domain (e.g., MS2 hairpin). As used herein, the “primer binding site” comprises a sequence that hybridizes to a single-strand DNA sequence having a 3′ end generated from the nicked DNA of the R-loop. B1195.70180WO00 12418099.1

[0098] In certain embodiments, the pegRNAs have a 3ʹ extension arm, a spacer, and a gRNA core. The 3ʹ extension arm further comprises in the 5ʹ to 3ʹ direction a DNA synthesis template, a primer binding site, and a linker. The DNA synthesis template may also be referred to more broadly as the “DNA synthesis template” where the polymerase of a prime editor described herein is not an RT, but another type of polymerase.

[0099] In certain other embodiments, the pegRNAs have a 5ʹ extension arm, a spacer, and a gRNA core. The 5ʹ extension further comprises in the 5ʹ to 3ʹ direction a DNA synthesis template, a primer binding site, and a linker. The DNA synthesis template may also be referred to more broadly as the “DNA synthesis template” where the polymerase of a prime editor described herein is not an RT, but another type of polymerase.

[0100] In still other embodiments, the pegRNAs have in the 5ʹ to 3ʹ direction a spacer, a gRNA core, and an extension arm. The extension arm is at the 3ʹ end of the pegRNA. The extension arm further comprises in the 5ʹ to 3ʹ direction a homology arm, an edit template, and a primer binding site. The extension arm may also comprise an optional modifier region at the 3ʹ and 5ʹ ends, which may be the same sequences or different sequences. In addition, the 3ʹ end of the pegRNA may comprise a transcriptional terminator sequence. These sequence elements of the pegRNAs are further described and defined herein.

[0101] In still other embodiments, the pegRNAs have in the 5ʹ to 3ʹ direction an extension arm, a spacer, and a gRNA core. The extension arm is at the 5ʹ end of the pegRNA. The extension arm further comprises in the 3ʹ to 5ʹ direction a primer binding site, an edit template, and a homology arm. The extension arm may also comprise an optional modifier region at the 3ʹ and 5ʹ ends, which may be the same sequences or different sequences. The pegRNAs may also comprise a transcriptional terminator sequence at the 3ʹ end. These sequence elements of the pegRNAs are further described and defined herein.

[0102] In some embodiments, the spacer sequence of the pegRNA is about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, or about 25 nucleotides in length. In certain embodiments, the spacer sequence of the pegRNA is about 20 nucleotides in length. In some embodiments, the prime binding site is about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, or about 17 nucleotides in length. In some embodiments, the homology arm of the pegRNA is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 nucleotides in length. In some embodiments, the DNA synthesis template is from about 5 to about 58 nucleotides in length, about 10 to about B1195.70180WO00 12418099.116 nucleotides in length, or about 12 to about 17 nucleotides in length. In certain embodiments, the DNA synthesis template is less than 15 nucleotides in length.

[0103] In some embodiments, a pegRNA is an “engineered pegRNA” (“epegRNA”). Relative to a pegRNA, an epegRNA comprises an additional structured motif, for example, attached to its 3′ end. Such additional structured motifs may stabilize the pegRNA or otherwise prevent it from being degraded. Suitable structured motifs include, but are not limited to, toe-loops, hairpins, stem-loops, pseudoknots, aptamers, G-quadruplexes, tRNAs, riboswitches, and ribozymes. In some embodiments, a 3′ structured motif comprises evopreq1.

[0104] pegRNAs are further described, e.g., in International Patent Application No. PCT / US2020 / 023721, filed March 19, 2020, which published as WO 2020 / 191239; International Patent Application No. PCT / US2021 / 031439, filed May 7, 2021, which published as WO 2021 / 226558; International Patent Application No. PCT / 2021 / 052097, filed September 24, 2021, which published as WO 2022 / 067130; International Patent Application No. PCT / US2022 / 012054, filed January 11, 2022, which published as WO 2022 / 150790; International Patent Application No. PCT / US2022 / 078655, filed October 25, 2022, which published as WO 2023 / 076898; and International Patent Application No. PCT / US2022 / 074628, filed August 5, 2022, which published as WO 2023 / 015309; the contents of each of which is incorporated by reference herein. PE1

[0105] As used herein, “PE1” refers to a prime editing composition comprising 1) a fusion protein comprising a Cas9 protein variant Cas9(H840A) and a wild type MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(wt)] -NLS and 2) a desired PEgRNA, wherein the fusion protein (referred to as the PE1 protein) has the amino acid sequence of SEQ ID NO: 3, which is shown as follows. MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLG NTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMA KVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDST DKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPI NASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNF DLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNT EITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYID GGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGE LHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETI B1195.70180WO00 12418099.1TPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKV KYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEIS GVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFA NRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKV VDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQIL KEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKD DSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVI TLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVY GDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIET NGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKL IARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSS FEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELA LPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVIL ADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYT STKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESS GGSSGGSSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPL KATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQ DLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAATSEL DCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPT PKTPRQLREFLGTAGFCRLWIPGFAEMAAPLYPLTKTGTLFNWGPDQQKAYQEIKQALLT APALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRM VAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQF GPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEG QRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHI HGEIYRRRGLLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQ AARKAAITETPDTSTLLIENSSPSGGSKRTADGSEFEPKKKRKV (SEQ ID NO: 3) KEY: NUCLEAR LOCALIZATION SEQUENCE (NLS) TOP:(SEQ ID NO: 95), BOTTOM: (SEQ ID NO: 96) CAS9(H840A) (SEQ ID NO: 10) 33-AMINO ACID LINKER (SEQ ID NO: 80) M-MLV reverse transcriptase (SEQ ID NO: 30). B1195.70180WO00 12418099.1PE2

[0106] As used herein, “PE2” refers to a prime editing composition comprising 1) a fusion protein comprising a Cas9 protein variant Cas9(H840A) and a variant MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]- [MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)] -NLS and 2) a desired PEgRNA, wherein the fusion protein (referred to as the PE2 protein) has the amino acid sequence of SEQ ID NO: 4, which is shown as follows: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLG NTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMA KVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDST DKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPI NASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNF DLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNT EITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYID GGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGE LHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETI TPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKV KYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEIS GVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFA NRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKV VDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQIL KEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKD DSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVI TLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVY GDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIET NGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKL IARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSS FEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELA LPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVIL ADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYT STKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESS GGSSGGSSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPL KATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQ DLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSEL DCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPT PKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLT APALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRM VAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQF GPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEG QRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHI HGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQ AARKAAITETPDTSTLLIENSSPSGGSKRTADGSEFEPKKKRKV (SEQ ID NO: 4) KEY: B1195.70180WO00 12418099.1NUCLEAR LOCALIZATION SEQUENCE (NLS) TOP:(SEQ ID NO: 95), BOTTOM: (SEQ ID NO: 96) CAS9(H840A) (SEQ ID NO: 10) 33-AMINO ACID LINKER (SEQ ID NO: 80) M-MLV reverse transcriptase (SEQ ID NO: 29). PE3

[0107] As used herein, “PE3” refers to a prime editing composition comprising a PE2 prime editor and further comprising a second-strand nicking guide RNA that complexes with PE2 and introduces a nick in the non-edit DNA strand in order to induce preferential replacement of the edit strand. PE3b

[0108] As used herein, “PE3b” refers to a prime editing composition comprising PE2 and further comprising a second-strand nicking guide RNA that complexes with PE2 and introduces a nick in the non-edit DNA strand, wherein the second-strand nicking guide RNA is designed for temporal control such that the second strand nick is not introduced until after the installation of the desired edit. This is achieved by designing the second strand nicking guide RNA with a spacer sequence that comprises complementarity to, and only hybridizes with, the edited strand after installation of the desired nucleotide edit(s), but not the endogenous target DNA sequence. Using this strategy, mismatches between the nicking guide RNA spacer and the unedited target DNA should disfavor nicking by the sgRNA until after the editing event on the PAM strand takes place. PE4

[0109] As used herein, “PE4” refers to a prime editing composition comprising a PE2 and further comprising an MLH1 dominant negative protein variant (i.e., wild-type MLH1 with amino acids 754-756 truncated, which may be referred to herein as “MLH1 Δ754-756” or “MLH1dn”). The MLH1 dominant negative protein variant may be expressed in trans in some embodiments. In some embodiments, a PE4 system comprises a fusion protein comprising a PE2 protein and an MLH1 dominant negative protein joined via an optional linker. PE5 and PE5b

[0110] As used herein, “PE5” refers to a prime editing composition comprising a PE3 prime editor and further comprising an MLH1 dominant negative protein variant (i.e., wild-type MLH1 with amino acids 754-756 truncated, which may be referred to as “MLH1 Δ754-756” or “MLH1dn”). The MLH1 dominant negative variant may be expressed in trans in some embodiments. In some embodiments, a PE5 system comprises a fusion protein comprising a B1195.70180WO00 12418099.1PE2 protein and an MLH1 dominant negative protein joined via an optional linker. “PE5b” refers to a prime editing composition comprising a PE3 and an MLH1 dominant negative protein, wherein the second-strand nicking guide RNA is designed for temporal control such that the second strand nick is not introduced until after the installation of the desired edit. This is achieved by designing the second strand nicking guide RNA with a spacer sequence that comprise complementarity to, and hybridize with, only the edited strand after installation of the desired nucleotide edit(s), but not the endogenous target DNA sequence. PE6

[0111] The term “PE6” refers to a suite of next-generation prime editors described herein (PE6a, PE6b, PE6c, PE6d, PE6e, PE6f, and PE6g) comprising improved reverse transcriptase and / or Cas9 variants. The improved reverse transcriptase and Cas9 domains of the PE6 variants can also be combined with each other to offer cumulative benefits. For example, a PE6 prime editor comprising an improved reverse transcriptase variant of PE6a and an improved Cas9 variant of PE6e is referred to herein as the prime editor “PE6a-e” (or “PE6e- a”). Any possible combination of PE6 prime editors is contemplated by the present disclosure including, for example, PE6a-e, PE6a-f, PE6a-g, PE6b-e, PE6b-f, PE6b-g, PE6c-e, PE6c-f, PE6c-g, PE6d-e, PE6d-f, and PE6d-g.

[0112] Each of the PE6 prime editors comprise a Cas9 domain, e.g., a Cas9 variant, and a reverse transcriptase domain, e.g., a reverse transcriptase variant. PE6a comprises a reverse transcriptase variant comprising the amino acid substitutions E60K, K87E, E165D, D243N, R267I, E279K, K318E, and K343N relative to an Ec48 reverse transcriptase (SEQ ID NO: 7). PE6b comprises a reverse transcriptase variant comprising the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N relative to a Tf1 reverse transcriptase (SEQ ID NO: 1). PE6c comprises a reverse transcriptase variant comprising the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to a Tf1 reverse transcriptase (SEQ ID NO: 1). PE6d comprises a reverse transcriptase variant comprising the amino acid substitutions T128N, D200C, and V223Y (and the substitutions T306K, W313F, and T330P used in the MMLV reverse transcriptase of PE2 and PEmax) relative to a MMLV reverse transcriptase (SEQ ID NO: 30) with a truncation of the C-terminal RNaseH domain (e.g., between D497 and I498 of SEQ ID NO: 30). PE6e comprises a Cas9 variant comprising the amino acid substitutions K775R and K918A relative to wild type Streptococcus pyogenes Cas9 or Streptococcus pyogenes Cas9 nickase (SEQ ID NO: 2). PE6f comprises a Cas9 variant comprising the amino acid B1195.70180WO00 12418099.1substitutions H99R, E471K, I632V, D645N, H721Y, and K918A relative to wild type Streptococcus pyogenes Cas9 or Streptococcus pyogenes Cas9 nickase (SEQ ID NO: 2). PE6g comprises a Cas9 variant comprising the amino acid substitutions H99R, E471K, I632V, D645N, R654C, and H721Y relative to wild type Streptococcus pyogenes Cas9 or Streptococcus pyogenes Cas9 nickase (SEQ ID NO: 2). The Cas9 domain and the reverse transcriptase variant of a PE6 prime editor described herein can be covalently linked or associated, for example, directly fused or connected to each other via a linker peptide to form a fusion protein. Alternatively, the Cas9 domain and the reverse transcriptase may be provided in trans, i.e., not covalently connected. Any of the PE6 prime editor fusion proteins provided herein may also comprise the architecture of the prime editor fusion proteins described herein or known in the art, for example, a PE2 protein architecture or a PEmax protein architecture. Components, sequences, and corresponding architecture of an exemplary PEmax protein is provided below. In some embodiments, any of the PE6 prime editors provided herein may further comprise additional amino acid mutations, e.g., any of those included in PEmax as provided below.

[0113] In some embodiments, a PE6 protein comprises a reverse transcriptase variant disclosed herein and a Cas9 protein that recognizes a non-canonical PAM sequence (e.g., a Cas9 protein of SEQ ID NO: 133, or at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 133). For example, the prime editor “PE6b-NRCH” comprises the reverse transcriptase of PE6b (SEQ ID NO: 25) and the NRCH-Cas9 protein of SEQ ID NO: 133. PE6b-NRCH comprises the amino acid sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNT DRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDS FFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLI YLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAIL SARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDT YDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMD GTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNREKIEKI LTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDK NLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTN RKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENED ILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGWGRLSRKLINGIRDB1195.70180WO00 12418099.1KQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAG SPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEE GIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVP QSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDN LTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVIT LKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDY KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEI VWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNSDKLIARKKDWDP KKYGGFNSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAK GYKEVKKDLIIKLPKYSLFELENGRKRMLASAGVLQKGNELALPSKYVNFLYLASHY EKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTINRKQYNTTKEVLDATLIRQSITGLYET RIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSISSSKHTLSQMNKV SNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVK MQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDYRPLNKYVKPNVYP LPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGIST APAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIINQ AKCEFHQSQVKFIGYHISEKGLTPCQENIDKVLQWKQPKNRKELRQFLGSVNYLRKF IPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLETD VSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHWRH YLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIAD ALSRIVDETEPIPKDNEDNSINFVNQISIKRTADGSEFESPKKKRKVPAAKRVKLD (SEQ ID NO: 155).

[0114] The prime editor “PE6c-NRCH” comprises the reverse transcriptase of PE6c (SEQ ID NO: 26) and the NRCH-Cas9 protein of SEQ ID NO: 133. PE6c-NRCH comprises the amino acid sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNT DRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDS FFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLI YLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAIL SARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDT YDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMD GTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNREKIEKI B1195.70180WO00 12418099.1LTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDK NLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTN RKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENED ILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGWGRLSRKLINGIRD KQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAG SPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEE GIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVP QSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDN LTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVIT LKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDY KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEI VWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNSDKLIARKKDWDP KKYGGFNSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAK GYKEVKKDLIIKLPKYSLFELENGRKRMLASAGVLQKGNELALPSKYVNFLYLASHY EKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTINRKQYNTTKEVLDATLIRQSITGLYET RIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSISSSKHTLSQMNKV SNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQENYRLPIRNYPLTPVK MQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDYRPLNKYVKPNVYP LPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPRGVFEYLVMPYGIK TAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKDVLQKLKNANLIIN QAKCEFHQSQVKFLGYHISEKGLTPCQENIDKVLQWKQPKNQKELRQFLGQVNYLR KFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPVLRHFDFSKKILLE TDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSDKEMLAIIKSLEHW RHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFNFEINYRPGSANHIA DALSRIVDETEPIPKDNEDNSINFVNQISIKRTADGSEFESPKKKRKVPAAKRVKLD (SEQ ID NO: 156).

[0115] The prime editor “PE6d-NRCH” comprises the reverse transcriptase of PE6d (SEQ ID NO: 27) and the Cas9 variant as set forth in SEQ ID NO: 133. In some embodiments, a prime editor fusion protein comprises a non-truncated version of the MMLV reverse transcriptase variant having amino acid substitutions T128N, D200C, V223Y, T306K, W313F and T330P relative to SEQ ID NO: 30, and the NRCH-Cas9 protein of SEQ ID NO: 133. PE6d-NRCH comprises the amino acid sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNT B1195.70180WO00 12418099.1DRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDS FFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLI YLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAIL SARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDT YDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMD GTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNREKIEKI LTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDK NLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTN RKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENED ILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGWGRLSRKLINGIRD KQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAG SPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEE GIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVP QSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDN LTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVIT LKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDY KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEI VWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNSDKLIARKKDWDP KKYGGFNSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAK GYKEVKKDLIIKLPKYSLFELENGRKRMLASAGVLQKGNELALPSKYVNFLYLASHY EKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTINRKQYNTTKEVLDATLIRQSITGLYET RIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETS KEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEAR LGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNV PNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWT RLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALL QTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQ LREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPA LGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLR MVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLD TDRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVPAAK RVKLD (SEQ ID NO: 157). B1195.70180WO00 12418099.1

[0116] An exemplary prime editor fusion protein comprising a non-truncated version of the MMLV reverse transcriptase variant having amino acid substitutions T128N, D200C, V223Y, T306K, W313F, T330P, and L603W relative to SEQ ID NO: 30 and the NRCH- Cas9 protein of SEQ ID NO: 133 (which has a full-length MMLV reverse transcriptase domain relative to SEQ ID NO: 30 as described herein) can have the amino acid sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNT DRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDS FFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLI YLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAIL SARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDT YDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMD GTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNREKIEKI LTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDK NLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTN RKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENED ILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGWGRLSRKLINGIRD KQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAG SPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEE GIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVP QSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDN LTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVIT LKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDY KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEI VWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNSDKLIARKKDWDP KKYGGFNSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAK GYKEVKKDLIIKLPKYSLFELENGRKRMLASAGVLQKGNELALPSKYVNFLYLASHY EKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTINRKQYNTTKEVLDATLIRQSITGLYET RIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETS KEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEAR LGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNV PNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWT RLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLB1195.70180WO00 12418099.1QTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQ LREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPA LGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLR MVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLD TDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYT DGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLN VYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGH QKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSPSGGSKRTADGSEFESPKK KRKVPAAKRVKLD (SEQ ID NO: 158).

[0117] In some embodiments, a PE6d prime editor comprises a fusion protein that comprises a Cas9 having amino acid substitutions R221K, N394K, and relative to SEQ ID NO: 2 (i.e., R221K, N394K, and H840A substitutions relative to a wildtype Cas9) and a MMLV-RT having amino acid substitutions T128N, D200C, V223Y, T306K, W313F, and T330P, and a C terminal truncation between D497 and I498, relative to SEQ ID NO: 30. In some embodiments, a PE6d prime editor comprises a fusion protein comprising the reverse transcriptase variant as set forth in SEQ ID NO: 27 and the Cas9 variant as set forth in SEQ ID NO: 11. In some embodiments, a PE6d prime editor comprises a fusion protein having the following configuration (the “PEmax architecture”): [bipartite NLS]- [Cas9(R221K)(N394K)(H840A)]-[linker]- [MMLV_RT(T128N)(D200C)(V223Y)(T306K)(T330P)(D497 / I498 C-term truncation) )]- [bipartite NLS]-[NLS]. In some embodiments, a PE6d prime editor comprises a fusion protein having the following sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLG NTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMA KVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDST DKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPI NASGVDAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNF DLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNT EITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYID GGASQEEFYKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGE LHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETI TPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKV KYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEIS GVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFA NRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKV VDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQIL KEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKD DSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK B1195.70180WO00 12418099.1AERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVI TLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVY GDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIET NGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKL IARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSS FEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELA LPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVIL ADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYT STKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKR KVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLII PLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRP VQDLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFE WRDPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAAT SELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMG QPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQ ALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPP CLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTD RVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKR VKLD (SEQ ID NO: 159) KEY: BIPARTITE SV40 NUCLEAR LOCALIZATION SEQUENCE (NLS) TOP: (SEQ ID NO: 95), CAS9(R221K N394K H840A) (SEQ ID NO: 11) SGGSx2-BIPARTITE SV40NLS-SGGSx2 LINKER (SEQ ID NO: 79) M-MLV reverse transcriptase(T128N D200C V223Y T306K W313F T330PD497 / I498truncation ) (SEQ ID NO: 27) Other linker sequence (SEQ ID NO: 82) BIPARTITE SV40NLS (SEQ ID NO: 97) Other linker sequence (GSG) c-Myc NLS (SEQ ID NO: 98)

[0118] In some embodiments, a PE6d prime editor comprises a fusion protein that comprises a Cas9 comprising an amino acid sequence as set forth in SEQ ID NO: 10 (i.e., H840A substitutions relative to a wildtype Cas9) and a MMLV-RT having amino acid substitutions T128N, D200C, V223Y, T306K, W313F, and T330P, and a C terminal truncation between D497 and I498, relative to SEQ ID NO: 30. In some embodiments, a PE6d prime editor comprises a fusion protein comprising the reverse transcriptase as set forth in SEQ ID NO: 27 and the Cas9 variant as set forth in SEQ ID NO: 10. In some embodiments, a PE6d prime editor comprises a fusion protein having the following configuration (the “PE2 architecture”): B1195.70180WO00 12418099.1[bipartite NLS]-[Cas9 (H840A)]-[linker]- [MMLV_RT(T128N)(D200C)(V223Y)(T306K)(T330P)(D497 / I498 C-term truncation) )]- [NLS]. In some embodiments, a PE6d prime editor comprises a fusion protein having the following sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLG NTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMA KVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDST DKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPI NASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNF DLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNT EITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYID GGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGE LHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETI TPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKV KYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEIS GVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFA NRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKV VDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQIL KEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKD DSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVI TLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVY GDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIET NGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKL IARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSS FEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELA LPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVIL ADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYT STKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESS GGSSGGSSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPL KATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQ DLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSEL DCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPT PKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLT APALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRM VAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQF GPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFEPKKKRKV (SEQ ID NO: 160) KEY: NUCLEAR LOCALIZATION SEQUENCE (NLS) TOP:(SEQ ID NO: 95), BOTTOM: (SEQ ID NO: 96) CAS9(H840A) (SEQ ID NO: 10) 33-AMINO ACID LINKER (SEQ ID NO: 80)

[0119] M-MLV reverse transcriptase (SEQ ID NO: 27). B1195.70180WO00 12418099.1PE7

[0120] The term “PE7” refers to the PE6 prime editors plus a second strand nicking guide RNA. For example, “PE7a” refers to the PE6a prime editor as provided herein, plus a second strand nicking guide RNA. PEmax

[0121] As used herein, “PEmax” refers to a prime editing composition comprising 1) a fusion protein comprising a Cas9 protein variant Cas9(R221K N394K H840A) and a variant MMLV RT having the following structure: [bipartite NLS]-[Cas9(R221K)(N394K)(H840A)]- [linker]-[MMLV_RT(D200N)(T306K)(W313F)(T330P)(L603W)]-[bipartite NLS]-[NLS] and 2) a desired PEgRNA, wherein the fusion protein (referred to as the PEmax protein) has the amino acid sequence of SEQ ID NO: 5, which is shown as follows: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLG NTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMA KVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDST DKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPI NASGVDAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNF DLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNT EITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYID GGASQEEFYKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGE LHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETI TPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKV KYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEIS GVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFA NRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKV VDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQIL KEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKD DSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTK AERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVI TLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVY GDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIET NGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKL IARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSS FEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELA LPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVIL ADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYT STKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKR KVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQA PLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGT NDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPT SQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQ YVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEG QRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTL FNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWR RPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALV B1195.70180WO00 12418099.1KQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILA EAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGT SAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKN KDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLL IENSSPSGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLD (SEQ ID NO: 5) KEY: BIPARTITE SV40 NUCLEAR LOCALIZATION SEQUENCE (NLS) TOP: (SEQ ID NO: 95), CAS9(R221K N394K H840A) (SEQ ID NO: 11) SGGSx2-BIPARTITE SV40NLS-SGGSx2 LINKER (SEQ ID NO: 79) M-MLV reverse transcriptase(D200N T306K W313F T330P L603W) (SEQ ID NO: 29) Other linker sequence (SEQ ID NO: 82) BIPARTITE SV40NLS (SEQ ID NO: 97) Other linker sequence (GSG) c-Myc NLS (SEQ ID NO: 98) PACE

[0122] The term “phage-assisted continuous evolution (PACE),” as used herein, refers to continuous evolution that employs phage as viral vectors. The general concept of PACE technology has been described, for example, in International PCT Application, PCT / US2009 / 056194, filed September 8, 2009, published as WO 2010 / 028347 on March 11, 2010; International PCT Application, PCT / US2011 / 066747, filed December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; U.S. Application, U.S. Patent No. 9,023,594, issued May 5, 2015, International PCT Application, PCT / US2015 / 012022, filed January 20, 2015, published as WO 2015 / 134121 on September 11, 2015, and International PCT Application, PCT / US2016 / 027795, filed April 15, 2016, published as WO 2016 / 168631 on October 20, 2016, the entire contents of each of which are incorporated herein by reference. PANCE

[0123] Phage-assisted non-continuous evolution (PANCE),” as used herein, refers to non- continuous evolution that employs phage as viral vectors. PANCE is a simplified technique for rapid in vivo directed evolution using serial flask transfers of evolving “selection phage” (SP), which contain a gene of interest to be evolved, across fresh E. coli host cells, thereby allowing genes inside the host E. coli to be held constant while genes contained in the SP continuously evolve. Serial flask transfers have long served as a widely-accessible approach for laboratory evolution of microbes, and, more recently, analogous approaches have been B1195.70180WO00 12418099.1developed for bacteriophage evolution. The PANCE system features lower stringency than the PACE system. Polymerase

[0124] As used herein, the term “polymerase” refers to an enzyme that synthesizes a nucleotide strand and that may be used in connection with the prime editor delivery systems described herein. The polymerase can be a “template-dependent” polymerase (i.e., a polymerase that synthesizes a nucleotide strand based on the order of nucleotide bases of a template strand). The polymerase can also be a “template-independent” polymerase (i.e., a polymerase that synthesizes a nucleotide strand without the requirement of a template strand). A polymerase may also be further categorized as a “DNA polymerase” or an “RNA polymerase.” In various embodiments, the prime editor system comprises a DNA polymerase. In various embodiments, the DNA polymerase can be a “DNA-dependent DNA polymerase” (i.e., whereby the template molecule is a strand of DNA). In such cases, the DNA template molecule can be a pegRNA, wherein the extension arm comprises a strand of DNA. In such cases, the pegRNA may be referred to as a chimeric or hybrid pegRNA that comprises an RNA portion (i.e., the guide RNA components, including the spacer and the gRNA core) and a DNA portion (i.e., the extension arm). In various other embodiments, the DNA polymerase can be an “RNA-dependent DNA polymerase” (i.e., whereby the template molecule is a strand of RNA). In such cases, the pegRNA is RNA, i.e., including an RNA extension. The term “polymerase” may also refer to an enzyme that catalyzes the polymerization of nucleotides (i.e., the polymerase activity). Generally, the enzyme will initiate synthesis at the 3′-end of a primer annealed to a polynucleotide template sequence (e.g., such as a primer sequence annealed to the primer binding site of a pegRNA) and will proceed toward the 5′ end of the template strand. A “DNA polymerase” catalyzes the polymerization of deoxynucleotides. As used herein in reference to a DNA polymerase, the term DNA polymerase includes a “functional fragment thereof.” A “functional fragment thereof” refers to any portion of a wild-type or mutant DNA polymerase that encompasses less than the entire amino acid sequence of the polymerase and which retains the ability, under at least one set of conditions, to catalyze the polymerization of a polynucleotide. Such a functional fragment may exist as a separate entity, or it may be a constituent of a larger polypeptide, such as a fusion protein. Prime editing

[0125] As used herein, the term “prime editing” refers to an approach for gene editing using napDNAbps, a polymerase (e.g., a reverse transcriptase), and specialized guide RNAs that B1195.70180WO00 12418099.1include a primer binding site and a DNA synthesis template for encoding desired new genetic information (or deleting genetic information) that is then incorporated into a target DNA sequence. Prime editing is described in Anzalone, A. V. et al., Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149–157 (2019), which is incorporated herein by reference. See also International PCT Application, PCT / US2020 / 023721, filed March 19, 2020, and published as WO 2020 / 191239, which is incorporated herein by reference.

[0126] Prime editing represents a platform for genome editing that is a versatile and precise method to directly write new genetic information into a specified DNA site using a nucleic acid programmable DNA binding protein (“napDNAbp”) working in association with a polymerase (i.e., in the form of a fusion protein or otherwise provided in trans with the napDNAbp), wherein the prime editing system is programmed with a prime editing (PE) guide RNA (“pegRNA”) that both specifies the target site and templates the synthesis of the desired edit in the form of a replacement DNA strand by way of an extension (either DNA or RNA) engineered onto a guide RNA (e.g., at the 5ʹ or 3ʹ end, or at an internal portion of a guide RNA). The replacement strand containing the desired edit (e.g., a single nucleobase substitution) shares the same sequence as the endogenous strand (or is homologous to it) immediately downstream of the nick site of the target site to be edited (with the exception that it includes the desired edit). Through DNA repair and / or replication machinery, the endogenous strand downstream of the nick site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing may be thought of as a “search-and-replace” genome editing technology since the prime editors, as described herein, not only search and locate the desired target site to be edited, but at the same time, encode a replacement strand containing a desired edit that is installed in place of the corresponding target site endogenous DNA strand. The prime editors of the present disclosure relate, in part, to the discovery that the mechanism of target-primed reverse transcription (TPRT) or “prime editing” can be leveraged or adapted for conducting precision CRISPR / Cas-based genome editing with high efficiency and genetic flexibility. TPRT is naturally used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial Group II introns. Cas protein-reverse transcriptase fusions or related systems are used to target a specific DNA sequence with a guide RNA, generate a single strand nick at the target site, and use the nicked DNA as a primer for reverse transcription of an engineered DNA synthesis template that is integrated with the guide RNA. However, while the concept begins with prime editors that use reverse transcriptase as the DNA polymerase component,B1195.70180WO00 12418099.1the prime editors described herein are not limited to reverse transcriptases but may include the use of virtually any DNA polymerase. Indeed, while the application throughout may refer to prime editors with “reverse transcriptases,” it is set forth here that reverse transcriptases are only one type of DNA polymerase that may work with prime editing. Thus, wherever the specification mentions a “reverse transcriptase,” the person having ordinary skill in the art should appreciate that any suitable DNA polymerase may be used in place of the reverse transcriptase. Thus, in one aspect, the prime editors may comprise Cas9 (or an equivalent napDNAbp), which is programmed to target a DNA sequence by associating it with a specialized guide RNA (i.e., pegRNA) containing a spacer sequence that anneals to a complementary sequence (the complementary sequence to an endogenous protospacer sequence) in the target DNA. The pegRNA also contains new genetic information in the form of an extension that encodes a replacement strand of DNA containing a desired nucleotide change which is used to replace a corresponding endogenous DNA strand at the target site. To transfer information from the pegRNA to the target DNA, the mechanism of prime editing involves nicking the target site in one strand of the DNA to expose a 3′-hydroxyl group. The exposed 3′-hydroxyl group can then be used to prime the DNA polymerization of the edit- encoding extension on pegRNA directly into the target site. In various embodiments, the extension—which provides the template for polymerization of the replacement strand containing the edit—can be formed from RNA or DNA. In the case of an RNA extension, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (such as a reverse transcriptase). In the case of a DNA extension, the polymerase of the prime editor may be a DNA-dependent DNA polymerase. The newly synthesized strand (i.e., the replacement DNA strand containing the desired nucleotide edit) that is formed by the prime editor would be homologous to the genomic target sequence (i.e., have the same sequence as), except for the inclusion of one or more desired nucleotide changes (e.g., a single nucleotide substitution, a deletion, or an insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA may also be referred to as a single strand DNA flap, which would compete for hybridization with the complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. Resolution of the hybridized intermediate (also referred to as a heteroduplex, comprising the single strand DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand with the exception of mismatches at positions where desired nucleotide edits are installed in the edit strand) can include removal of the resulting displaced flap of endogenous DNA (e.g., with a 5ʹ end DNA flap endonuclease, FEN1), ligation of the synthesized single B1195.70180WO00 12418099.1strand DNA flap to the target DNA, and assimilation of the desired nucleotide changes as a result of cellular DNA repair and / or replication processes.

[0127] In various embodiments, prime editing operates by contacting a target DNA molecule (for which a change in the nucleotide sequence is desired to be introduced) with a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editing guide RNA (pegRNA). In various embodiments, the prime editing guide RNA (pegRNA) comprises an extension at the 3′ or 5′ end of the guide RNA, or at an intramolecular location in the guide RNA, and encodes the desired nucleotide change (e.g., single nucleotide substitution, insertion, or deletion). First, the napDNAbp / extended gRNA complex contacts the DNA molecule, and the extended gRNA guides the napDNAbp to bind to a target locus. Next, a nick in one of the strands of DNA of the target locus is introduced (e.g., by a nuclease or chemical agent), thereby creating an available 3′ end in one of the strands of the target locus. In certain embodiments, the nick is created in the strand of DNA that corresponds to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence, i.e., the “non-target strand.” The nick, however, could be introduced in either of the strands. That is, the nick could be introduced into the R-loop “target strand” (i.e., the strand hybridized to the protospacer of the extended gRNA) or the “non-target strand” (i.e., the strand forming the single-stranded portion of the R-loop and which is complementary to the target strand). In the next step, the 3′ end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA in order to prime reverse transcription (i.e., “target-primed RT”). In certain embodiments, the 3′ end DNA strand hybridizes to a specific RT priming sequence on the extended portion of the guide RNA, i.e., the “reverse transcriptase priming sequence” or “primer binding site” on the pegRNA. In the next step, a reverse transcriptase (or other suitable DNA polymerase) is introduced that synthesizes a single strand of DNA from the 3′ end of the primed site towards the 5′ end of the prime editing guide RNA. The DNA polymerase (e.g., reverse transcriptase) can be fused to the napDNAbp or alternatively can be provided in trans to the napDNAbp. This forms a single-strand DNA flap comprising the desired nucleotide change (e.g., the single base change, insertion, or deletion, or a combination thereof) and that is otherwise homologous to the endogenous DNA at or adjacent to the nick site. In the next step, the napDNAbp and guide RNA are released. The final two steps relate to the resolution of the single strand DNA flap such that the desired nucleotide change becomes incorporated into the target locus. This process can be driven towards the desired product formation by removing the corresponding 5′ endogenous DNA flap that forms once the 3′ single strand DNA flap invades and hybridizes to the endogenous B1195.70180WO00 12418099.1DNA sequence. Without being bound by theory, the cell’s endogenous DNA repair and replication processes resolve the mismatched DNA to incorporate the nucleotide change(s) to form the desired altered product. The process can also be driven towards product formation with “second strand nicking.” This process may introduce at least one or more of the following genetic changes: transversions, transitions, deletions, and insertions. Prime editor

[0128] The term “prime editor” refers to the polypeptide or polypeptide components involved in prime editing as described herein. In some embodiments, a prime editor comprises a fusion construct comprising a napDNAbp (e.g., Cas9 nickase, and / or any of the Cas9 variants provided herein) and a reverse transcriptase (e.g., any of the reverse transcriptase variants provided herein). In some embodiments, a prime editor is capable of carrying out prime editing on a target nucleotide sequence in the presence of a pegRNA (or “extended guide RNA”). In some embodiments, a prime editor comprises a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase provided in trans, i.e., the napDNAbp and the reverse transcriptase are not fused. The in trans napDNAbp and the reverse transcriptase may be tethered via a non-peptide linkage, e.g., an MS2 RNA-protein binding RNA sequence and a MS2 coat protein fused to either the napDNAbp or the reverse transcriptase, or may be unlinked to each other and simply recruited by the pegRNA. In some embodiments, a prime editor composition, system, or complex provided herein comprises a fusion protein or a fusion protein complexed with a pegRNA, and / or further complexed with a second-strand nicking sgRNA. In some embodiments, the prime editor system may also refer to the complex comprising a fusion protein (reverse transcriptase fused to a napDNAbp), a pegRNA, and a regular guide RNA capable of directing the second-site nicking step of the non-edited strand as described herein. Protein, peptide, and polypeptide

[0129] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide may also be a B1195.70180WO00 12418099.1single molecule or may be a multi-molecular complex. A protein, peptide, or polypeptide may be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the contents of which are incorporated herein by reference. Protospacer

[0130] As used herein, the term “protospacer” refers to the sequence (e.g., of ~20 bp) in DNA adjacent to the PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA (except that a protospacer contains Thymine and the spacer sequence contains Uracil). The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand thereof, i.e., the “target strand” versus the “non-target strand” of the target DNA sequence). In some embodiments, in order for a Cas nickase component of a prime editor to function, it also requires a specific protospacer adjacent motif (PAM) that varies depending on the Cas protein component itself, e.g., the type of Cas protein and the bacterial species from which it is derived. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes a PAM sequence of NGG that is directly downstream of the protospacer sequence in the genomic DNA, on the non-target strand. Protospacer adjacent motif (PAM)

[0131] As used herein, the term “protospacer adjacent motif” or “PAM” refers to a DNA sequence (e.g., an approximately 2-6 nucleotide sequence) that is an important targeting component of a Cas nuclease, e.g., a Cas9. For example, in some embodiments for a Cas9 nuclease, the PAM sequence is on either strand and is downstream in the 5ʹ to 3ʹ direction of the Cas9 cut site. The canonical PAM sequence (i.e., the PAM sequence that is associated with the Cas9 nuclease of Streptococcus pyogenes or SpCas9) is 5ʹ-NGG-3ʹ, wherein “N” is any nucleobase followed by two guanine (“G”) nucleobases. In some embodiments, SpCas9 can also recognize additional non-canonical PAMs (e.g., NAG and NGA).

[0132] Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., B1195.70180WO00 12418099.1SpCas9, may be modified to alter the PAM specificity of the nuclease such that the nuclease recognizes an alternative PAM sequence. Reverse transcriptase

[0133] The term “reverse transcriptase” describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize a DNA transcript from an RNA template. Historically, reverse transcriptase has been used primarily to transcribe mRNA into cDNA, which can then be cloned into a vector for further manipulation. Avian myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme has 5ʹ-3ʹ RNA-directed DNA polymerase activity, 5ʹ-3ʹ DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5ʹ and 3ʹ ribonuclease specific for the RNA strand for RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Errors in transcription cannot be corrected by reverse transcriptase because known viral reverse transcriptases lack the 3ʹ-5ʹ exonuclease activity necessary for proofreading (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNaseH activity has been presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase that is used extensively in molecular biology is reverse transcriptase originating from Moloney murine leukemia virus (M-MLV or “MMLV”). See, e.g., Gerard, G. R., DNA 5:271-279 (1986) and Kotewicz, M. L., et al., Gene 35:249-258 (1985). M-MLV reverse transcriptase substantially lacking in RNase H activity has also been described. See, e.g., U.S. Pat. No.5,244,797. The invention contemplates the use of any such reverse transcriptases, or variants or mutants thereof.

[0134] In some embodiments, the prime editors provided herein comprise MMLV RT, or a variant or fragment of MMLV RT. In some embodiments, the prime editors provided herein comprise Ec48 RT, or a variant or fragment of Ec48 RT. In some embodiments, the prime editors provided herein comprise Tf1 RT, or a variant or fragment of Tf1 RT.

[0135] In certain embodiments, a reverse transcriptase comprises the amino acid substitutions E60K, K87E, E165D, D243N, R267I, E279K, K318E, and K343N relative to an Ec48 reverse transcriptase (SEQ ID NO: 7). In certain embodiments, a reverse transcriptase comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N relative to a Tf1 reverse transcriptase (SEQ ID NO: 1). In certain embodiments, a reverse transcriptase comprises the amino acid B1195.70180WO00 12418099.1substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to a Tf1 reverse transcriptase (SEQ ID NO: 1). In certain embodiments, a reverse transcriptase comprises the amino acid substitutions T128N, D200C, and V223Y relative to an MMLV reverse transcriptase (SEQ ID NO: 30) with a truncation of the C-terminal RNaseH domain. Reverse transcription

[0136] As used herein, the term “reverse transcription” indicates the capability of an enzyme to synthesize a DNA strand (that is, complementary DNA or cDNA) using RNA as a template. In some embodiments, the reverse transcription can be “error-prone reverse transcription,” which refers to the properties of certain reverse transcriptase enzymes that are error-prone in their DNA polymerization activity. Spacer sequence

[0137] As used herein, the term “spacer sequence” in connection with a guide RNA or a pegRNA refers to the portion of the guide RNA or pegRNA of about 20 nucleotides that contains a nucleotide sequence that shares the same sequence as the protospacer sequence in the target DNA sequence. The spacer sequence anneals to the complement of the protospacer sequence to form a ssRNA / ssDNA hybrid structure at the target site and a corresponding R loop ssDNA structure of the endogenous DNA strand. Substitution

[0001] The term “substitution,” as used herein, refers to replacement of a residue within a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or a deletion or insertion of one or more residues within a sequence. The term “mutation” may also be used throughout the present disclosure to refer to a substitution (i.e., a “nucleic acid mutation” or an “amino acid mutation”). Substitutions are typically described herein by identifying the original residue followed by the position of the residue within the sequence and the identity of the newly mutated / substituted residue. Various methods for making the amino acid substitutions provided herein are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, a substitution is in a reverse transcriptase, e.g., an MMLV reverse transcriptase, an Ec48 reverse transcriptase, or a Tf1 reverse transcriptase. In some embodiments, a substitution is in a Cas9 protein, e.g., an SpCas9 protein. B1195.70180WO00 12418099.1Variant

[0138] As used herein, the term “variant” should be taken to mean the exhibition of qualities that have a pattern that deviates from what occurs in nature. The term “variant” encompasses homologous proteins having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 975, at least 98%, or at least 99% identity with a reference sequence and having the same or substantially the same functional activity or activities as the reference sequence. The term also encompasses mutants, truncations, or domains of a reference sequence that display the same or substantially the same functional activity or activities as the reference sequence.

[0139] In some embodiments, a variant comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more amino acid substitutions relative to a wild type sequence. In some embodiments, a variant is a reverse transcriptase variant. In certain embodiments, a reverse transcriptase variant comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more amino acid substitutions relative to a wild type reverse transcriptase sequence (e.g., wild type MMLV reverse transcriptase, wild type Ec48 reverse transcriptase, or wild type Tf1 reverse transcriptase). In certain embodiments, a Cas9 variant comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more amino acid substitutions relative to a wild type Cas9 sequence (e.g., wild type SpCas9) or Cas9 nickase (e.g., SpCas9 nickase). Vector

[0140] The term “vector,” as used herein, refers to a nucleic acid that can be modified to encode a gene of interest and that is able to enter a host cell, mutate, and replicate within the host cell, and then transfer a replicated form of the vector into another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phage, and conjugative plasmids. Additional suitable vectors will be apparent to those of skill in the art based on the instant disclosure. Wild type

[0141] As used herein the term “wild type” or “WT” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene, or characteristic as it occurs in nature as distinguished from mutant or variant forms. B1195.70180WO00 12418099.1DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS

[0142] The present disclosure describes the use of directed evolution and protein engineering to generate new reverse transcriptase and Cas9 variants that enhance editing efficiency when used in the context of prime editors. In particular, next-generation prime editors that are approximately 500-800 bp smaller than PE2 (PE6a and PE6b), while offering mammalian prime editing efficiencies comparable to or higher than those of PE2, were developed. Additionally, highly active and processive prime editors that use either the M-MLV RT or Tf1 RT (PE6c and PE6d) were also developed. These evolved and engineered reverse transcriptases offer substantial improvements over previously used prime editors, e.g., increased editing efficiency for longer edits. Evolved variants of the Cas9 nickase domain of prime editors were also created (PE6e-PE6g), further improving prime editing efficiencies.

[0143] Thus, the present disclosure provides evolved and engineered reverse transcriptase variants and Cas9 variants with improved properties (e.g., improved editing efficiency when used in the context of a prime editor). Fusion proteins, including for example prime editors, comprising the reverse transcriptase variants and Cas9 variants described herein are also provided by the present disclosure. The present disclosure also provides polynucleotides encoding the reverse transcriptase variants, Cas9 variants, fusion proteins, and prime editors provided herein, as well as vectors comprising such polynucleotides. Pharmaceutical compositions, AAVs and cells comprising the reverse transcriptase variants, Cas9 variants, and prime editors (and / or polynucleotides or vectors encoding the same) described herein are also provided by the present disclosure. The present disclosure also provides methods and uses involving the reverse transcriptase variants, Cas9 variants, and prime editors described herein. PE6 Prime Editors, Cas9 Variants, Reverse Transcriptase Variants, and Fusion Proteins

[0144] Some aspects of the present disclosure provide evolved and / or engineered reverse transcriptases and Cas9 proteins, and prime editors comprising the same, with various improved properties (e.g., smaller size to increase delivery efficiency (for example, using AAVs), improved prime editing efficiency (for example, for edits that require structured pegRNA RT templates), decreased frequency of indels, etc.). In some embodiments, pegRNA structure and folding prediction, including free energy of the folding of pegRNA components, e.g., the RT template or the extension arm, can be measured by NUPACK free energy prediction as described in Zadeh, J.N. et al., (2011). NUPACK: Analysis and design of B1195.70180WO00 12418099.1nucleic acid systems. J. Comput. Chem.32, 170–173, which is incorporated herein by reference.

[0145] The variants provided by the present disclosure include variants of Escherichia coli Ec48 reverse transcriptase, Schizosaccharomyces pombe Tf1 reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, and Streptococcus pyogenes Cas9, as well as variants comprising the amino acid substitutions disclosed herein at corresponding positions in homologous proteins.

[0146] In one aspect, the present disclosure provides reverse transcriptase variants comprising various amino acid substitutions relative to the amino acid sequence of Schizosaccharomyces pombe Tf1 reverse transcriptase, which is provided below: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLPPGKMQAMNDEINQGLKSGIIRESKAINACPVMFVPKKEGTLRMVVD YKPLNKYVKPNIYPLPLIEQLLAKIQGSTIFTKLDLKSAYHLIRVRKGDEHKLAFRCPR GVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFIGYHISEKGFTPCQENIDKVLQWKQPKNRKE LRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPV LRHFDFSKKILLETDASDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSD KEMLAIIKSLKHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFN FEINYRPGSANHIADALSRIVDETEPIPKDSEDNSINFVNQISI (SEQ ID NO: 1).

[0147] In some embodiments, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1, wherein the reverse transcriptase variant comprises amino acid substitutions at positions 70, 72, 87, 102, 106, 118, 128, 158, 269, 363, 413, and 492 relative to SEQ ID NO: 1, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 70 is a P70X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 70 is a P70T substitution. In some embodiments, the amino acid substitution at position 72 is a G72X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 72 is a G72V substitution. In some embodiments, the amino acid substitution at position 87 is an S87X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 87 is an S87G substitution. In some embodiments, the amino acid substitution at position 102 is an M102X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the B1195.70180WO00 12418099.1amino acid substitution at position 102 is an M102I substitution. In some embodiments, the amino acid substitution at position 106 is a K106X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 106 is a K106R substitution. In some embodiments, the amino acid substitution at position 118 is a K118X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 118 is a K118R substitution. In some embodiments, the amino acid substitution at position 128 is an I128X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 128 is an I128V substitution. In some embodiments, the amino acid substitution at position 158 is an L158X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 158 is an L158Q substitution. In some embodiments, the amino acid substitution at position 269 is an F269X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 269 is an F269L substitution. In some embodiments, the amino acid substitution at position 363 is an A363X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 363 is an A363V substitution. In some embodiments, the amino acid substitution at position 413 is a K413X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 413 is a K413E substitution. In some embodiments, the amino acid substitution at position 492 is an S492X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 492 is an S492N substitution. In certain embodiments, the reverse transcriptase variant comprises the substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N relative to SEQ ID NO: 1. In some embodiments, the reverse transcriptase variants further comprise amino acid substitutions at positions 188, 260, 297, and 288 relative to SEQ ID NO: 1. In some embodiments, the amino acid substitution at position 188 is an S188X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 188 is an S188K substitution. In some embodiments, the amino acid substitution at position 260 is an I260X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 260 is an I260L substitution. In some embodiments, the amino acid substitution at position 297 is an S297X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 297 is an S297Q substitution. In some embodiments, the amino acid substitution at position 288 is an R288X B1195.70180WO00 12418099.1substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 288 is an R288Q substitution. In certain embodiments, the reverse transcriptase variant further comprises the substitutions S188K, I260L, S297Q, and R288Q relative to SEQ ID NO: 1.

[0148] In another aspect, the present disclosure provides reverse transcriptase variants comprising various amino acid substitutions relative to the amino acid sequence of MMLV reverse transcriptase, which is provided below: TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTP VSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLR EVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAAT SELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKE TVMGQPTPKTPRQLREFLGTAGFCRLWIPGFAEMAAPLYPLTKTGTLFNWGPDQQK AYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKK LDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLS NARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDL TDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIA LTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRRRGLLTSEGKEIKNKDEILALLK ALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP (SEQ ID NO: 30).

[0149] In some embodiments, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises amino acid substitutions at positions 128 and 200 relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence. In some embodiments, the substitution at position 128 is a T128X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the substitution at position 128 is a T128N substitution. In some embodiments, the substitution at position 200 is a D200X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the substitution at position 200 is a D200C substitution. In certain embodiments, the reverse transcriptase variant comprises the amino acid substitutions T128N and D200C. In some embodiments, the reverse transcriptase variant further comprises an amino acid substitution at position 223 relative to SEQ ID NO: 30. In some embodiments, the amino acid substitution at position 223 is a V223X substitution, wherein X is any amino acid other than wild type. In B1195.70180WO00 12418099.1certain embodiments, the amino acid substitution at position 223 is a V223Y substitution. In some embodiments, the reverse transcriptase variant further comprises amino acid substitutions from the MMLV reverse transcriptase used in PE2 and PEmax (e.g., the amino acid substitutions T306K, W313F, and T330P). In certain embodiments, the reverse transcriptase variant comprises a truncation of all or part of the C-terminal RNaseH domain of MMLV reverse transcriptase (e.g., a truncation at amino acid position 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, or 510 of SEQ ID NO: 30). In certain embodiments, the reverse transcriptase variant comprises a truncation between positions D497 and I498 of SEQ ID NO: 30.

[0150] In some embodiments, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises the amino acid substitutions T128N and V223M relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence. In some embodiments, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises the amino acid substitutions T128N and V223Y relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence. In some embodiments, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises the amino acid substitutions T128F and V223M relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence. In some embodiments, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises the amino acid substitutions D200C and V223M relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence.

[0151] In some embodiments, the present disclosure provides reverse transcriptase variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises amino acid substitutions at positions 128, 129, 196, 200, and 223 relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence. In B1195.70180WO00 12418099.1some embodiments, the amino acid substitution at position 128 is a T128X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 128 is a T128N substitution. In some embodiments, the amino acid substitution at position 129 is a V129X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 129 is a V129A substitution. In certain embodiments, the amino acid substitution at position 129 is a V129G substitution. In some embodiments, the amino acid substitution at position 196 is a P196X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 196 is a P196S substitution. In certain embodiments, the amino acid substitution at position 196 is a P196T substitution. In certain embodiments, the amino acid substitution at position 196 is a P196F substitution. In some embodiments, the amino acid substitution at position 200 is an N200X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 200 is an N200S substitution. In certain embodiments, the amino acid substitution at position 200 is an N200Y substitution. In some embodiments, the amino acid substitution at position 223 is a V223X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 223 is a V223A substitution. In certain embodiments, the amino acid substitution at position 223 is a V223M substitution. In certain embodiments, the amino acid substitution at position 223 is a V223L substitution. In certain embodiments, the amino acid substitution at position 223 is a V223E substitution.

[0152] In some embodiments, any of the reverse transcriptase variants provided herein further comprise amino acid substitutions from the MMLV reverse transcriptase used in PE2 and PEmax (e.g., any of the amino acid substitutions D200N, T306K, W313F, T330P, and L603W). In certain embodiments, any of the reverse transcriptase variants provided herein comprise a truncation of all or part of the C-terminal RNaseH domain of MMLV reverse transcriptase (e.g., a truncation at amino acid position 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500, 501, 502, 503, 504, 505, 506, 507, 508, 509, or 510 of SEQ ID NO: 30). In certain embodiments, any of the reverse transcriptase variants provided herein comprise a truncation between positions D497 and I498 of SEQ ID NO: 30.

[0153] In some embodiments, the reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO: 25 (the RT domain of “PE6b”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 25: B1195.70180WO00 12418099.1ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFIGYHISEKGLTPCQENIDKVLQWKQPKNRKE LRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPV LRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSD KEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFN FEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (SEQ ID NO: 25).

[0154] In some embodiments, a PE6b prime editor comprises the amino acid sequence: MKRTADGSEFESPKKKRKV[CAS9]SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGG SISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGISTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVKD VLQKLKNANLIINQAKCEFHQSQVKFIGYHISEKGLTPCQENIDKVLQWKQPKNRKE LRQFLGSVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSPPV LRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVSD KEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDFN FEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISIKRTADGSEFESPKKKR KVPAAKRVKLD (SEQ ID NOs: 146, 147), wherein [CAS9] comprises any Cas9 protein (e.g., any of the Cas9 variants disclosed herein).

[0155] In some embodiments, the reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO: 26 (the RT domain of “PE6c”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 26: ISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVK DVLQKLKNANLIINQAKCEFHQSQVKFLGYHISEKGLTPCQENIDKVLQWKQPKNQ KELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSP PVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVS B1195.70180WO00 12418099.1DKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDF NFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISI (SEQ ID NO: 26).

[0156] In some embodiments, a PE6c prime editor comprises the amino acid sequence: MKRTADGSEFESPKKKRKV[CAS9]SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGG SISSSKHTLSQMNKVSNIVKEPELPDIYKEFKDITADTNTEKLPKPIKGLEFEVELTQEN YRLPIRNYPLTPVKMQAMNDEINQGLKGGIIRESKAINACPVIFVPRKEGTLRMVVDY RPLNKYVKPNVYPLPLIEQLLAKIQGSTIFTKLDLKSAYHQIRVRKGDEHKLAFRCPR GVFEYLVMPYGIKTAPAHFQYFINTILGEAKESHVVCYMDDILIHSKSESEHVKHVK DVLQKLKNANLIINQAKCEFHQSQVKFLGYHISEKGLTPCQENIDKVLQWKQPKNQ KELRQFLGQVNYLRKFIPKTSQLTHPLNKLLKKDVRWKWTPTQTQAIENIKQCLVSP PVLRHFDFSKKILLETDVSDVAVGAVLSQKHDDDKYYPVGYYSAKMSKAQLNYSVS DKEMLAIIKSLEHWRHYLESTIEPFKILTDHRNLIGRITNESEPENKRLARWQLFLQDF NFEINYRPGSANHIADALSRIVDETEPIPKDNEDNSINFVNQISIKRTADGSEFESPKKK RKVPAAKRVKLD (SEQ ID NOs: 146, 148), wherein [CAS9] comprises any Cas9 protein (e.g., any of the Cas9 variants disclosed herein).

[0157] In some embodiments, the reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO: 27 (the RT domain of “PE6d”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 27: TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTP VSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLR EVNKRVEDIHPNVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWR DPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAAT SELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKE TVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKA YQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKL DPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD (SEQ ID NO: 27).

[0158] In some embodiments, a PE6d prime editor comprises the amino acid sequence: [optional NLS]-[CAS9]-[SEQ ID NO:27]-[optional NLS] wherein CAS9 comprises any Cas9 protein (e.g., any of the Cas9 variants disclosed herein), and wherein “optional NLS” comprises one or more nuclear localization signals described herein or known in the art, and wherein each of the “]-[” independently comprises a optional peptide linker described herein or known in the art. B1195.70180WO00 12418099.1

[0159] In some embodiments, the N-terminal NLS of the PE6d prime editor comprises a bipartite SV40 NLS as set forth in SEQ ID NO: 95.

[0160] In some embodiments, the C-terminal NLS of the PE6d prime editor comprises a bipartite SV40 NLS as set forth in SEQ ID NO: 97. In some embodiments, the C-myc NLS of the PE6d prime editor comprises a bipartite SV40 NLS as set forth in SEQ ID NO: 98. In some embodiments, the C-terminal NLS of the PE6d prime editor comprises the sequence SGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLD.

[0161] In some embodiments, the C-terminal NLS of the PE6d prime editor comprises a bipartite SV40 NLS as set forth in SEQ ID NO: 96.

[0162] In some embodiments, the peptide linker connecting the Cas9 and the MMLV-RT variant of a PE6d fusion protein comprises SEQ ID NO: 80.

[0163] In some embodiments, the peptide linker connecting the Cas9 and the MMLV-RT variant of a PE6d fusion protein comprises SEQ ID NO: 79.

[0164] In some embodiments, a PE6d prime editor comprises the amino acid sequence: MKRTADGSEFESPKKKRKV[CAS9]SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGG STLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATST PVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDL REVNKRVEDIHPNVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEW RDPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAA TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARK ETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQK AYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKK LDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLS NARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDKRTADGSEFESP KKKRKVPAAKRVKLD (SEQ ID NOs: 146, 149), wherein [CAS9] comprises any Cas9 protein (e.g., any of the Cas9 variants disclosed herein).

[0165] In another aspect, the present disclosure provides Cas9 variants comprising various amino acid substitutions relative to the amino acid sequence of Streptococcus pyogenes Cas9 nickase (H840A), which is provided below: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHE RHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEG DLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLP GEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYA B1195.70180WO00 12418099.1DLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPE KYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQ RTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRF AWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFT VYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECF DSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEE RLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFAN RNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDEL VKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENT QLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAG FIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFY KVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEI GKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVL SMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVL VVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYS LFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLF VEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 2).

[0166] In some embodiments, the Cas9 comprises the amino acid sequence as set forth in SEQ ID NO: 10.

[0167] In some embodiments, the Cas9 comprises the amino acid sequence as set forth in SEQ ID NO: 11.

[0168] In some embodiments, the Cas9 comprises the amino acid sequence as set forth in SEQ ID NO: 133.

[0169] In some embodiments, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 775 and 918 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 775 is a K775X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 775 is a K775R substitution. In some embodiments, the amino acid substitution at position 918 is a K918X substitution, B1195.70180WO00 12418099.1wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 918 is a K918A substitution. In certain embodiments, the Cas9 variant comprises K775R and K918A substitutions relative to SEQ ID NO: 2.

[0170] In some embodiments, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 99, 471, 632, 645, and 721 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 99 is an H99X substitution, wherein X is any amino acid. In certain embodiments, the amino acid substitution at position 99 is an H99R substitution. In some embodiments, the amino acid substitution at position 471 is an E471X substitution, wherein X is any amino acid. In certain embodiments, the amino acid substitution at position 471 is an E471K substitution. In some embodiments, the amino acid substitution at position 632 is an I632X substitution, wherein X is any amino acid. In certain embodiments, the amino acid substitution at position 632 is an I632V substitution. In some embodiments, the amino acid substitution at position 645 is a D645X substitution, wherein X is any amino acid. In certain embodiments, the amino acid substitution at position 645 is a D645N substitution. In some embodiments, the amino acid substitution at position 721 is an H721X substitution, wherein X is any amino acid. In certain embodiments, the amino acid substitution at position 721 is an H721Y substitution. In some embodiments, the Cas9 variant further comprises an amino acid substitution at position 654 relative to SEQ ID NO: 2. In some embodiments, the amino acid substitution at position 654 is an R654X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 654 is an R654C substitution. In some embodiments, the Cas9 variant further comprises an amino acid substitution at position 918 relative to SEQ ID NO: 2. In some embodiments, the amino acid substitution at position 918 is a K918X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 918 is a K918A substitution.

[0171] In some embodiments, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 99, 471, and 632 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 99 is an H99X substitution, wherein X is any amino acid other than wild type. In B1195.70180WO00 12418099.1certain embodiments, the amino acid substitution at position 99 is an H99R substitution. In some embodiments, the amino acid substitution at position 471 is an E471X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 471 is an E471K substitution. In some embodiments, the amino acid substitution at position 632 is an I632X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 632 is an I632V substitution. In certain embodiments, the Cas9 variant comprises the amino acid substitutions H99R, E471K, and I632V relative to SEQ ID NO: 2. In some embodiments, the Cas9 variant further comprises an amino acid substitution at position 721 relative to SEQ ID NO: 2. In some embodiments, the amino acid substitution at position 721 is an H721X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 721 is an H721K substitution.

[0172] In some embodiments, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 471 and 918 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 471 is an E471X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 471 is an E471K substitution. In some embodiments, the amino acid substitution at position 918 is a K918X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 918 is a K918A substitution. In certain embodiments, the Cas9 variant comprises the amino acid substitutions E471K and K918A relative to SEQ ID NO: 2.

[0173] In some embodiments, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 753 and 1151 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 753 is an R753X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 753 is an R753G substitution. In some embodiments, the amino acid substitution at position 1151 is a K1151X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1151 is a K1151E substitution. In certain embodiments, the Cas9B1195.70180WO00 12418099.1variant comprises the amino acid substitutions R753G and K1151E relative to SEQ ID NO: 2.

[0174] In some embodiments, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten amino acid substitutions at positions selected from the group consisting of 260, 298, 395, 769, 778, 1014, 1034, 1100, 1106, 1138, 1152, and 1320 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 260 is an E260X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 260 is an E260K substitution. In some embodiments, the amino acid substitution at position 298 is a D298X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 298 is a D298N substitution. In some embodiments, the amino acid substitution at position 395 is an R395X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 395 is an R395C substitution. In some embodiments, the amino acid substitution at position 769 is a T769X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 769 is a T769P substitution. In some embodiments, the amino acid substitution at position 778 is an R778X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 778 is an R778Q substitution. In some embodiments, the amino acid substitution at position 1014 is a K1014X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1014 is a K1014E substitution. In some embodiments, the amino acid substitution at position 1034 is an A1034X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1034 is an A1034E substitution. In some embodiments, the amino acid substitution at position 1100 is a V1100X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1100 is a V1100I substitution. In some embodiments, the amino acid substitution at position 1106 is an S1106X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1106 is an S1106F substitution. In some embodiments, the amino acid substitution at position 1138 is a T1138X substitution, wherein X is any amino acid other than wild type. In certain B1195.70180WO00 12418099.1embodiments, the amino acid substitution at position 1138 is a T1138A substitution. In some embodiments, the amino acid substitution at position 1152 is a G1152X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1152 is a G1152E substitution. In some embodiments, the amino acid substitution at position 1320 is an A1320X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1320 is an A1320T substitution. In certain embodiments, the Cas9 variant comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more amino acid substitutions E260K, D298N, R395C, T769P, R778Q, K1014E, A1034E, V1100I, S1106F, T1138A, G1152E, and A1320T. In certain embodiments, the Cas9 variant comprises the amino acid substitutions E260K, D298N, R395C, T769P, R778Q, K1014E, A1034E, V1100I, S1106F, T1138A, G1152E, and A1320T. In some embodiments, the Cas9 variant further comprises one or more additional amino acid substitutions at positions selected from the group consisting of 102, 753, 804, and 1003 relative to SEQ ID NO: 2. In some embodiments, the amino acid substitution at position 102 is an E102X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 102 is an E102K substitution. In some embodiments, the amino acid substitution at position 753 is an R753X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 753 is an R753G substitution. In some embodiments, the amino acid substitution at position 804 is a T804X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 804 is a T804A substitution. In some embodiments, the amino acid substitution at position 1003 is a K1003X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 1003 is a K1003R substitution. In some embodiments, the Cas9 variant comprises amino acid substitutions at positions 102, 395, 753, 778, and 1100; 753, 769, 1034, and 1320; 298, 753, 1034, and 1138; 102, 260, 395, 753, 778, 804, 1003, 1100, 1106, and 1152; or 102, 260, 395, 753, 778, 804, 1003, 1014, 1100, 1106, and 1152; relative to SEQ ID NO: 2. In certain embodiments, the Cas9 variant comprises amino acid substitutions at the positions: E102K, R395C, R753G, R778Q, and V1100I; R753G, T769P, A1034E, and A1320T; D298N, R753G, A1034E, and T1138A; E102K, E260K, R395C, R753C, R778Q, T804A, K1003R, V1100I, S1106F, and G1152E; or E102K, E260K, R395C, R753G, R778Q, T804A, K1003R, K1014E, V1100I, S1106F, and G1152E; relative to SEQ ID NO: 2. B1195.70180WO00 12418099.1

[0175] In some embodiments, the present disclosure provides Cas9 variants having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 23 and 754 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence. In some embodiments, the amino acid substitution at position 23 is a D23X substitution, wherein X is any amino acid other than wild type (i.e., D). In certain embodiments, the amino acid substitution at position 23 is a D23G substitution. In some embodiments, the amino acid substitution at position 754 is an H754X substitution, wherein X is any amino acid other than wild type (i.e., H). In certain embodiments, the amino acid substitution at position 754 is an H754R substitution.

[0176] In certain embodiments, the Cas9 variant comprises the amino acid sequence of SEQ ID NO: 28 (the Cas9 domain of “PE6e”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 28: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHE RHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEG DLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLP GEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYA DLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPE KYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQ RTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRF AWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFT VYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECF DSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEE RLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFAN RNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDEL VKVMGRHKPENIVIEMARENQTTQKGQRNSRERMKRIEEGIKELGSQILKEHPVENT QLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAG FIARQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFY KVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEI GKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVL SMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVL B1195.70180WO00 12418099.1VVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYS LFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLF VEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 28).

[0177] In some embodiments, a PE6e prime editor comprises the amino acid sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNT DRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDS FFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLI YLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAIL SARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDT YDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEH HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDG TEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKIL TFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKN LPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNR KVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDI LEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRD KQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAG SPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQRNSRERMKRIEE GIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVP QSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDN LTKAERGGLSELDKAGFIARQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVIT LKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDY KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEI VWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDP KKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAK GYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHY EKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETR IDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS[REVERSE TRANSCRIPTASE]KRTADGSEFESPKKKRKVPAAKRVKLD (SEQ ID NOs: 150, 151), wherein [REVERSE TRANSCRIPTASE] comprises any reverse transcriptase (e.g., any of the reverse transcriptase variants disclosed herein). B1195.70180WO00 12418099.1

[0178] In certain embodiments, the Cas9 variant comprises the amino acid sequence of SEQ ID NO: 48 (the Cas9 domain of “PE6f”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 48: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFRRLEESFLVEEDKKHE RHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEG DLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLP GEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYA DLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPE KYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQ RTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRF AWMTRKSEKTITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFT VYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECF DSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMVEE RLKTYAHLFDNKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFAN RNFMQLIHDDSLTFKEDIQKAQVSGQGDSLYEHIANLAGSPAIKKGILQTVKVVDEL VKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENT QLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAG FIARQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFY KVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEI GKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVL SMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVL VVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYS LFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLF VEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 48).

[0179] In some embodiments, a PE6f prime editor comprises the amino acid sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNT DRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDS FFRRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLI YLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAIL B1195.70180WO00 12418099.1SARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDT YDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEH HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDG TEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKIL TFRIPYYVGPLARGNSRFAWMTRKSEKTITPWNFEEVVDKGASAQSFIERMTNFDKN LPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNR KVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDI LEDIVLTLTLFEDREMVEERLKTYAHLFDNKVMKQLKRRRYTGWGRLSRKLINGIR DKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLYEHIANLA GSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIE EGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIV PQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFD NLTKAERGGLSELDKAGFIARQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKV ITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGD YKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDW DPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEA KGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASH YEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHR DKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYET RIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS[REVERSE TRANSCRIPTASE]KRTADGSEFESPKKKRKVPAAKRVKLD (SEQ ID NOs: 152, 151), wherein [REVERSE TRANSCRIPTASE] comprises any reverse transcriptase (e.g., any of the reverse transcriptase variants disclosed herein).

[0180] In certain embodiments, the Cas9 variant comprises the amino acid sequence of SEQ ID NO: 49 (the Cas9 domain of “PE6g”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 49: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFRRLEESFLVEEDKKHE RHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEG DLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLP GEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYA DLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPE B1195.70180WO00 12418099.1KYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQ RTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRF AWMTRKSEKTITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFT VYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECF DSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMVEE RLKTYAHLFDNKVMKQLKRCRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFAN RNFMQLIHDDSLTFKEDIQKAQVSGQGDSLYEHIANLAGSPAIKKGILQTVKVVDEL VKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENT QLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAG FIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFY KVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEI GKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVL SMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVL VVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYS LFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLF VEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 49).

[0181] In some embodiments, a PE6g prime editor comprises the amino acid sequence: MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNT DRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDS FFRRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLI YLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAIL SARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDT YDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEH HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDG TEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKIL TFRIPYYVGPLARGNSRFAWMTRKSEKTITPWNFEEVVDKGASAQSFIERMTNFDKN LPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNR KVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDI LEDIVLTLTLFEDREMVEERLKTYAHLFDNKVMKQLKRCRYTGWGRLSRKLINGIR DKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLYEHIANLA GSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIE B1195.70180WO00 12418099.1EGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIV PQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFD NLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKV ITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGD YKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDW DPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEA KGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASH YEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHR DKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYET RIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS[REVERSE TRANSCRIPTASE]KRTADGSEFESPKKKRKVPAAKRVKLD (SEQ ID NOs: 153, 151), wherein [REVERSE TRANSCRIPTASE] comprises any reverse transcriptase (e.g., any of the reverse transcriptase variants disclosed herein).

[0182] In certain embodiments, the Cas9 variant comprises the amino acid sequence of SEQ ID NO: 145, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 145: MDKKYSIGLDIGTNSVGWAVITGEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGE TAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHE RHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEG DLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLP GEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYA DLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPE KYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQ RTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRF AWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFT VYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECF DSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEE RLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFAN RNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDEL VKVMGRRKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENT QLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAG B1195.70180WO00 12418099.1FIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFY KVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEI GKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVL SMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVL VVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYS LFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLF VEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 145).

[0183] In some embodiments, any of the PE6 prime editors provided herein may comprise the architecture of PEmax. In some embodiments, any of the PE6 prime editors provided herein may comprise one or more additional amino acid substitutions (e.g., one, two, three, four, five, six, seven, eight, nine, or ten or more additional amino acid substitutions) relative to a wild type amino acid sequence, for example, any of the amino acid substitutions included in the Cas9 protein of PEmax or the MMLV reverse transcriptase of PEmax.

[0184] In some aspects, the present disclosure provides fusion proteins comprising any of the Cas9 variants provided herein and an effector domain. In certain embodiments, the effector domain comprises nuclease activity, nickase activity, recombinase activity, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity, transcriptional repression activity, or polymerase activity.

[0185] It should be appreciated that any of the amino acid mutations described herein, (e.g., E60K) from a first amino acid residue (e.g., E) to a second amino acid residue (e.g., K) may also include mutations from the first amino acid residue to an amino acid residue that is similar to (e.g., conserved) the second amino acid residue. For example, mutation of an amino acid with a hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan) may be a mutation to a second amino acid with a different hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan). For example, a mutation of an alanine to a threonine may also be a mutation from an alanine to an amino acid that is similar in size and chemical properties to a threonine, for example, serine. As another example, mutation of an amino acid with a positively charged side chain (e.g., arginine, histidine, or lysine) may be a mutation to a second amino acid with a different positively charged side chain (e.g., arginine, histidine, or lysine). As another example, mutation of an amino acid with a polarB1195.70180WO00 12418099.1side chain (e.g., serine, threonine, asparagine, or glutamine) may be a mutation to a second amino acid with a different polar side chain (e.g., serine, threonine, asparagine, or glutamine). Additional similar amino acid pairs include, but are not limited to, the following: phenylalanine and tyrosine; asparagine and glutamine; methionine and cysteine; aspartic acid and glutamic acid; and arginine and lysine.

[0186] The skilled artisan would recognize that such conservative amino acid substitutions will likely have minor effects on protein structure and are likely to be well tolerated without compromising function. In some embodiments, any of the amino acid mutations provided herein from one amino acid to a threonine may be an amino acid mutation to a serine. In some embodiments, any of the amino acid mutations provided herein from one amino acid to an arginine may be an amino acid mutation to a lysine. In some embodiments, any of the amino acid mutations provided herein from one amino acid to an isoleucine may be an amino acid mutation to an alanine, valine, methionine, or leucine. In some embodiments, any of the amino acid mutations provided herein from one amino acid to a lysine may be an amino acid mutation to an arginine. In some embodiments, any of the amino acid mutations provided herein from one amino acid to an aspartic acid may be an amino acid mutation to a glutamic acid or asparagine. In some embodiments, any of the amino acid mutations provided herein from one amino acid to a valine may be an amino acid mutation to an alanine, isoleucine, methionine, or leucine. In some embodiments, any of the amino acid mutations provided herein from one amino acid to a glycine may be an amino acid mutation to an alanine. It should be appreciated, however, that additional conserved amino acid residues would be recognized by the skilled artisan, and any of the amino acid mutations to other conserved amino acid residues are also within the scope of this disclosure.

[0187] In some aspects, the present disclosure provides reverse transcriptase variants comprising mutations corresponding to any of the mutations disclosed herein, or any combination thereof, at a homologous position in another reverse transcriptase. Examples of additional reverse transcriptases include, but are not limited to, the following:B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1

[0188] Additional reverse transcriptases are known in the art and will be readily apparent to those of skill in the art.

[0189] In some aspects, the present disclosure provides Cas9 variants comprising mutations corresponding to any of the mutations disclosed herein, or any combination thereof, at a homologous position in another Cas9 protein. Examples of additional Cas9 proteins include, but are not limited to, the following:B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1B1195.70180WO00 12418099.1

[0190] Additional Cas9 proteins are known in the art and will be readily apparent to those of skill in the art.

[0191] In some aspects, the present disclosure provides prime editors comprising any of reverse transcriptase variants and / or any of the Cas9 variants described herein. In some embodiments, the present disclosure provides prime editors comprising any of the reverse transcriptase variants provided herein and a napDNAbp. In certain embodiments, the napDNAbp comprises a Cas9 protein (e.g., a Cas9 nickase). In some embodiments, the Cas9 B1195.70180WO00 12418099.1protein comprises any of the Cas9 proteins of SEQ ID NOs: 2, 6, 8, 9, 12-24, or 133, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 2, 6, 8, 9, 12-24, or 133. In certain embodiments, the Cas9 protein comprises the amino acid sequence of SEQ ID NO: 133, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 133. In certain embodiments, the napDNAbp comprises one of the Cas9 variants disclosed herein. In some embodiments, the present disclosure provides prime editors comprising any of the Cas9 variants provided herein and a polymerase. In some embodiments, the polymerase is a reverse transcriptase. In certain embodiments, the reverse transcriptase is one of the reverse transcriptase variants provided herein.

[0192] In some embodiments, the reverse transcriptase variant used in a prime editor comprises various amino acid substitutions relative to the amino acid sequence of Escherichia coli Ec48 reverse transcriptase (SEQ ID NO: 7), which is provided below: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTL KEIPKIDGSKRIVYSLHPKMRLLQSRINKRIFKELVVFPSFLFGSVPSKNDVLNSNVKR DYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALEYLVDICTKDD FVVQGALTSSYIATLCLFAVEGDVVRRAQRKGLVYTRLVDDITVSSKISNYDFSQMQ SHIERMLSEHDLPINKHKTKIFHCSSEPIKVHGLRVDYDSPRLPSDEVKRIRASIHNLK LLAAKNNTKTSVAYRKEFNRCMGRVNKLGRVGHEKYESFKKQLQAIKPMPSKRDV AVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASL KPL (SEQ ID NO: 7)

[0193] In some embodiments, the reverse transcriptase variant used in a prime editor comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 7, wherein the reverse transcriptase comprises amino acid substitutions at positions 60, 87, 165, 243, 267, 279, 318, and 343 relative to SEQ ID NO: 7. In some embodiments, the amino acid substitution at position 60 is an E60X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 60 is an E60K substitution. In some embodiments, the amino acid substitution at position 87 is a K87X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 87 is a K87E substitution. In some embodiments, the amino acid substitution at position 165 is an E165X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 165 is an B1195.70180WO00 12418099.1E165D substitution. In some embodiments, the amino acid substitution at position 243 is a D243X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 243 is a D243N substitution. In some embodiments, the amino acid substitution at position 267 is an R267X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 267 is an R267I substitution. In some embodiments, the amino acid substitution at position 279 is an E279X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 279 is an E279K substitution. In some embodiments, the amino acid substitution at position 318 is a K318X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 318 is a K318E substitution. In some embodiments, the amino acid substitution at position 343 is a K343X substitution, wherein X is any amino acid other than wild type. In certain embodiments, the amino acid substitution at position 343 is a K343N substitution. In certain embodiments, the reverse transcriptase variant comprises the amino acid substitutions E60K, K87E, E165D, D243N, R267I, E279K, K318E, and K343N relative to SEQ ID NO: 7.

[0194] In certain embodiments, the reverse transcriptase variant comprises the amino acid sequence of SEQ ID NO: 50 (the RT domain of “PE6a”), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 50: GRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYTL KKIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVKR DYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALDYLVDICTKDD FVVQGALTSSYIATLCLFAVEGDVVRRAQRKGLVYTRLVDDITVSSKISNYDFSQMQ SHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLKL LAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSNRDVA VIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASLK PL (SEQ ID NO: 50)

[0195] In some embodiments, a PE6a prime editor comprises the amino acid sequence: MKRTADGSEFESPKKKRKV[CAS9]SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGG SGRPYVTLNLNGMFMDKFKPYSKSNAPITTLEKLSKALSISVEELKAIAELSLDEKYT LKKIPKIDGSKRIVYSLHPKMRLLQSRINERIFKELVVFPSFLFGSVPSKNDVLNSNVK RDYVSCAKAHCGAKTVLKVDISNFFDNIHRDLVRSVFEEILHIKDEALDYLVDICTKD DFVVQGALTSSYIATLCLFAVEGDVVRRAQRKGLVYTRLVDDITVSSKISNYDFSQMB1195.70180WO00 12418099.1QSHIERMLSEHNLPINKHKTKIFHCSSEPIKVHGLIVDYDSPRLPSDKVKRIRASIHNLK LLAAKNNTKTSVAYRKEFNRCMGRVNELGRVGHEKYESFKKQLQAIKPMPSNRDV AVIDAAIKSLELSYSKGNQNKHWYKRKYDLTRYKMIILTRSESFKEKLECFKSRLASL KPLKRTADGSEFESPKKKRKVPAAKRVKLD (SEQ ID NOs: 146, 154), wherein [CAS9] comprises any Cas9 protein (e.g., any of the Cas9 variants described herein).

[0196] The prime editors provided herein may, in some embodiments, comprise both a reverse transcriptase variant provided herein and a Cas9 variant provided herein, or reverse transcriptase variants and Cas9 variants at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any of those provided herein. For example, the present disclosure contemplates prime editors comprising the reverse transcriptase variant of SEQ ID NO: 50 (PE6a) and the Cas9 variant of SEQ ID NO: 28 (PE6e), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 50 and SEQ ID NO: 28. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 50 (PE6a) and the Cas9 variant of SEQ ID NO: 48 (PE6f), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 50 and SEQ ID NO: 48. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 50 (PE6a) and the Cas9 variant of SEQ ID NO: 49 (PE6g), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 50 and SEQ ID NO: 49. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 25 (PE6b) and the Cas9 variant of SEQ ID NO: 28 (PE6e), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 25 and SEQ ID NO: 28. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 25 (PE6b) and the Cas9 variant of SEQ ID NO: 48 (PE6f), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 25 and SEQ ID NO: 48. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 25 (PE6b) and the Cas9 variant of SEQ ID NO: 49 (PE6g), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 25 and SEQ ID NO: 49. In some embodiments, a prime editor comprises the B1195.70180WO00 12418099.1reverse transcriptase variant of SEQ ID NO: 26 (PE6c) and the Cas9 variant of SEQ ID NO: 28 (PE6e), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 26 and SEQ ID NO: 28. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 26 (PE6c) and the Cas9 variant of SEQ ID NO: 48 (PE6f), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 26 and SEQ ID NO: 48. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 26 (PE6c) and the Cas9 variant of SEQ ID NO: 49 (PE6g), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 26 and SEQ ID NO: 49. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 27 (PE6d) and the Cas9 variant of SEQ ID NO: 28 (PE6e), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 27 and SEQ ID NO: 28. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 27 (PE6d) and the Cas9 variant of SEQ ID NO: 48 (PE6f), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 27 and SEQ ID NO: 48. In some embodiments, a prime editor comprises the reverse transcriptase variant of SEQ ID NO: 27 (PE6d) and the Cas9 variant of SEQ ID NO: 49 (PE6g), or a reverse transcriptase variant and Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 27 and SEQ ID NO: 49. Nuclear localization sequences (NLS)

[0197] In various embodiments, the prime editors described herein may comprise one or more nuclear localization sequences (NLS), which help promote translocation of a protein into the cell nucleus. Such sequences are well-known in the art and can include the following examples:B1195.70180WO00 12418099.1

[0198] The NLS examples above are non-limiting. The prime editors disclosed herein may comprise any known NLS sequence, including any of those described in Cokol et al., “Finding nuclear localization signals,” EMBO Rep., 2000, 1(5): 411-415 and Freitas et al., “Mechanisms and Signals for the Nuclear Import of Proteins,” Current Genomics, 2009, 10(8): 550-7, each of which are incorporated herein by reference.

[0199] In various embodiments, the prime editors described herein further comprise one or more (and preferably at least two) nuclear localization sequences. In certain embodiments, the prime editors comprise at least two NLSs. In embodiments with at least two NLSs, the NLSs can be the same NLSs or they can be different NLSs. In some embodiments, one or more of the NLSs are bipartite NLSs (“bpNLS”). In certain embodiments, the prime editors comprise two bipartite NLSs. In some embodiments, the prime editors comprise more than two bipartite NLSs.

[0200] The location of the NLS fusion can be at the N-terminus, the C-terminus, or within a sequence of a prime editor (e.g., inserted between the encoded napDNAbp component (e.g., Cas9) and a polymerase domain (e.g., a reverse transcriptase).

[0201] The NLSs may be any known NLS sequence in the art. The NLSs may also be any future-discovered NLSs for nuclear localization. The NLSs also may be any naturally- occurring NLS, or any non-naturally occurring NLS (e.g., an NLS with one or more desired mutations).

[0202] The term “nuclear localization sequence” or “NLS” refers to an amino acid sequence that promotes import of a protein into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and would be apparent to the skilled artisan. For example, NLS sequences are described in Plank et al., International PCT application PCT / EP2000 / 011690, filed November 23, 2000, published as WO / 2001 / 038547 B1195.70180WO00 12418099.1on May 31, 2001, the contents of which are incorporated herein by reference. In some embodiments, an NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 94), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 99), KRTADGSEFESPKKKRKV (SEQ ID NO: 97), or KRTADGSEFEPKKKRKV (SEQ ID NO: 106). In other embodiments, an NLS comprises the amino acid sequences NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 107), PAAKRVKLD (SEQ ID NO: 98), RQRRNELKRSF (SEQ ID NO: 108), or NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 109).

[0203] In one aspect of the disclosure, a prime editor described herein may be modified with one or more nuclear localization sequences (NLS), preferably at least two NLSs. In certain embodiments, the prime editors are modified with two or more NLSs. The disclosure contemplates the use of any nuclear localization sequence known in the art at the time of the disclosure, or any nuclear localization sequence that is identified or otherwise made available in the state of the art after the time of the instant filing. A representative nuclear localization sequence is a peptide sequence that directs the protein to the nucleus of the cell in which the sequence is expressed. A nuclear localization signal is predominantly basic, can be positioned almost anywhere in a protein's amino acid sequence, generally comprises a short sequence of four amino acids (Autieri & Agrawal, (1998) J. Biol. Chem.273: 14731-37, incorporated herein by reference) to eight amino acids, and is typically rich in lysine and arginine residues (Magin et al., (2000) Virology 274: 11-16, incorporated herein by reference). Nuclear localization sequences often comprise proline residues. A variety of nuclear localization sequences have been identified and have been used to effect transport of biological molecules from the cytoplasm to the nucleus of a cell. See, e.g., Tinland et al., (1992) Proc. Natl. Acad. Sci. U.S.A.89:7442-46; Moede et al., (1999) FEBS Lett.461:229-34, which is incorporated herein by reference. Translocation is currently thought to involve nuclear pore proteins.

[0204] Most NLSs can be classified in three general groups: (i) a monopartite NLS exemplified by the SV40 large T antigen NLS (PKKKRKV (SEQ ID NO: 94)); (ii) a bipartite motif consisting of two basic domains separated by a variable number of spacer amino acids and exemplified by the Xenopus nucleoplasmin NLS (KRXXXXXXXXXXKKKL (SEQ ID NO: 110)); and (iii) noncanonical sequences such as M9 of the hnRNP Al protein, the influenza virus nucleoprotein NLS, and the yeast Gal4 protein NLS (Robbins, J. et al., Cell 1991, 64(3), 615-623).

[0205] Nuclear localization sequences appear at various points in the amino acid sequences of proteins. NLS have been identified at the N-terminus, the C-terminus, and in the central B1195.70180WO00 12418099.1region of proteins. Thus, the disclosure provides prime editors that may be modified with one or more NLSs at the C-terminus and / or the N-terminus, as well as at internal regions of the prime editor. The residues of a longer sequence that do not function as component NLS residues should be selected so as not to interfere, for example, tonically or sterically, with the nuclear localization signal itself. Therefore, although there are no strict limits on the composition of an NLS-comprising sequence, in practice, such a sequence can be functionally limited in length and composition.

[0206] The present disclosure contemplates any suitable means by which to modify a prime editor to include one or more NLSs. In one aspect, the prime editors may be engineered to express a prime editor that is translationally fused at its N-terminus or its C-terminus (or both) to one or more NLSs, i.e., to form a prime editor-NLS fusion construct. In other embodiments, a prime editor-encoding nucleotide sequence may be genetically modified to incorporate a reading frame that encodes one or more NLSs in an internal region of the encoded prime editor. In addition, the NLSs may include various amino acid linkers or spacer regions encoded between the prime editor and the N-terminally, C-terminally, or internally attached NLS amino acid sequence, e.g., and in the central region of proteins. Thus, the present disclosure also provides for nucleotide constructs, vectors, and host cells for expressing fusion proteins that comprise a prime editor and one or more NLSs, among other components.

[0207] The prime editors described herein may also comprise nuclear localization sequences that are linked to a prime editor through one or more linkers, e.g., a polymeric, amino acid, nucleic acid, polysaccharide, chemical, or nucleic acid linker element. The linkers within the contemplated scope of the disclosure are not intended to have any limitations and can be any suitable type of molecule (e.g., polymer, amino acid, polysaccharide, nucleic acid, lipid, or any synthetic chemical linker domain) and can be joined to the prime editor by any suitable strategy that effectuates forming a bond (e.g., covalent linkage, hydrogen bonding) between the prime editor and the one or more NLSs.

[0208] In some embodiments, the prime editors provided herein comprise an NLS comprising the amino acid sequence of SEQ ID NO: 95, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 95. In some embodiments, the prime editors provided herein comprise an NLS comprising the amino acid sequence of SEQ ID NO: 97, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least B1195.70180WO00 12418099.198%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 97. In some embodiments, the prime editors provided herein comprise an NLS comprising the amino acid sequence of SEQ ID NO: 98, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 98. In certain embodiments, the prime editors provided herein comprise a first NLS comprising the amino acid sequence of SEQ ID NO: 95, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 95, a second NLS comprising the amino acid sequence of SEQ ID NO: 97, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 97, and a third NLS comprising the amino acid sequence of SEQ ID NO: 98, or an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 98. Linkers

[0209] In various embodiments, the napDNAbp and the reverse transcriptase of the prime editors provided herein may be provided in trans or otherwise not fused to one another. In other embodiments, the prime editors provided herein comprise a napDNAbp and a reverse transcriptase fused to one another, for example, via one or more linkers. As defined above, the term “linker,” as used herein, refers to a chemical group or a molecule linking two molecules or moieties, e.g., a binding domain and a cleavage domain of a nuclease. In some embodiments, a linker joins a gRNA binding domain of an RNA-programmable nuclease and a polymerase (e.g., a reverse transcriptase). In some embodiments, a linker joins a Cas9 protein and a reverse transcriptase (e.g., any of the Cas9 variants provided herein and / or any of the reverse transcriptase variants provided herein). Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. B1195.70180WO00 12418099.1

[0210] The linker may be as simple as a covalent bond, or it may be a polymeric linker many atoms in length. In certain embodiments, the linker is a polypeptide, or amino acid-based. In other embodiments, the linker is not peptide-like. In certain embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In certain embodiments, the linker is a carbon-nitrogen bond of an amide linkage. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched, aliphatic or heteroaliphatic linker. In certain embodiments, the linker is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In other embodiments, the linker comprises amino acids. In certain embodiments, the linker comprises a peptide. In certain embodiments, the linker comprises an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker may include functionalized moieties to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0211] In some other embodiments, the linker comprises the amino acid sequence (GGGGS)n (SEQ ID NO: 84), (G)n(SEQ ID NO: 85), (EAAAK)n(SEQ ID NO: 86), (GGS)n(SEQ ID NO: 87), (SGGS)n (SEQ ID NO: 81), (XP)n (SEQ ID NO: 88), or any combination thereof, wherein n is independently an integer between 1 and 30, and wherein X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS)n(SEQ ID NO: 87), wherein n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 89). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 90). In some embodiments, the linker comprises the amino acid sequence SGGSGGSGGS (SEQ ID NO: 91). In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 82). In other embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSS B1195.70180WO00 12418099.1GGS (SEQ ID NO: 83, 60AA). In some embodiments, the linker comprises the amino acid sequence GGS, GGSGGS (SEQ ID NO: 92), GGSGGSGGS (SEQ ID NO: 93), SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 80), SGSETPGTSESATPES (SEQ ID NO: 89), or SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSS GGS (SEQ ID NO: 83).

[0212] In certain embodiments, linkers may be used to link any of the peptides or peptide domains or moieties of the invention (e.g., a napDNAbp linked or fused to a reverse transcriptase domain, and / or a napDNAbp linked to one or more NLS). Any of the domains of the prime editors described herein may also be connected to one another through any of the presently described linkers. PegRNAs

[0213] The prime editing systems and methods described herein contemplate the use of any suitable pegRNAs, e.g., to introduce recombinase recognition sites into a target DNA sequence, such as a genome, using prime editing. PEgRNA architecture

[0214] In some embodiments, an extended guide RNA, or pegRNA, used in the prime editing systems and methods disclosed herein includes a spacer sequence (e.g., a ~20 nt spacer sequence) and a gRNA core region, which binds with the napDNAbp. In some embodiments, the pegRNA includes an extended RNA segment, i.e., an extension arm, at the 5′ end, i.e., a 5′ extension. In some embodiments, the 5′ extension includes a DNA synthesis template sequence, a primer binding site, and an optional 5-20 nucleotide linker sequence. The RT primer binding site hybridizes to the free 3ʹ end that is formed after a nick is formed in the non-target strand of the R-loop, thereby priming reverse transcriptase for DNA polymerization in the 5′-3′ direction.

[0215] In another embodiment, an extended guide RNA (i.e., a pegRNA) used in the prime editing systems and methods provided herein includes a spacer sequence (e.g., a ~20 nt spacer sequence) and a gRNA core, which binds with the napDNAbp. In some embodiments, the pegRNA includes an extended RNA segment, i.e., an extension arm, at the 3′ end, i.e., a 3′ extension. In some embodiments, the 3′ extension includes a DNA synthesis template sequence, and a reverse transcription primer binding site. The RT primer binding site hybridizes to the free 3′ end that is formed after a nick is formed in the non-target strand of B1195.70180WO00 12418099.1the R-loop, thereby priming reverse transcriptase for DNA polymerization in the 5′-3′ direction.

[0216] In another embodiment, an extended guide RNA (i.e., a pegRNA) used in the prime editing systems and methods provided herein includes a spacer sequence (e.g., a ~20 nt spacer sequence) and a gRNA core, which binds with the napDNAbp. In some embodiments, the pegRNA includes an extended RNA segment, i.e., an extension arm, at an intermolecular position within the gRNA core, i.e., an intramolecular extension. In some embodiments, the intramolecular extension includes a DNA synthesis template sequence, and a reverse transcription primer binding site. The RT primer binding site hybridizes to the free 3′ end that is formed after a nick is formed in the non-target strand of the R-loop, thereby priming reverse transcriptase for DNA polymerization in the 5′-3′ direction.

[0217] In one embodiment, the position of the intermolecular RNA extension is not in the spacer sequence of the guide RNA. In another embodiment, the position of the intermolecular RNA extension is in the gRNA core. In still another embodiment, the position of the intermolecular RNA extension is anywhere within the guide RNA molecule except within the spacer sequence, or at a position which disrupts the spacer sequence. In one embodiment, the intermolecular RNA extension is inserted downstream from the 3′ end of the spacer sequence. In another embodiment, the intermolecular RNA extension is inserted at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides downstream of the 3′ end of the spacer sequence.

[0218] In other embodiments, the intermolecular RNA extension is inserted into the gRNA core, which refers to the portion of a traditional guide RNA corresponding or comprising the tracrRNA, which binds and / or interacts with the napDNAbp, e.g., a Cas9 protein or equivalent thereof (i.e., a different napDNAbp). Preferably the insertion of the intermolecular RNA extension does not disrupt or minimally disrupts the interaction between the tracrRNA portion and the napDNAbp.

[0219] The length of the RNA extension (which includes at least the RT template and primer binding site) can be any useful length. In various embodiments, the RNA extension is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 B1195.70180WO00 12418099.1nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0220] The RT template sequence can also be any suitable length. For example, the RT template sequence can be at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0221] In still other embodiments, the reverse transcription primer binding site sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0222] In other embodiments, the optional linker or spacer sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at B1195.70180WO00 12418099.1least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0223] The RT template sequence, in certain embodiments, encodes a single-stranded DNA molecule which is homologous to the non-target strand (and thus, complementary to the corresponding site of the target strand) but includes one or more nucleotide changes, e.g., for introducing a recombinase recognition sequence into a target DNA molecule. The one or more nucleotide changes may include one or more single-base nucleotide changes, one or more deletions, and / or one or more insertions.

[0224] The synthesized single-stranded DNA product of the RT template sequence is homologous to the non-target strand except that it contains one or more nucleotide changes. The single-stranded DNA product of the RT template sequence hybridizes in equilibrium with the complementary target strand sequence, thereby displacing the homologous endogenous target strand sequence. The displaced endogenous strand may be referred to in some embodiments as a 5′ endogenous DNA flap species. This 5′ endogenous DNA flap species can be removed by a 5′ flap endonuclease (e.g., FEN1) and the single-stranded DNA product, now hybridized to the endogenous target strand, may be ligated, thereby creating a mismatch between the endogenous sequence and the newly synthesized strand. The mismatch may be resolved by the cell’s innate DNA repair and / or replication processes.

[0225] In various embodiments, the nucleotide sequence of the RT template sequence corresponds to the nucleotide sequence of the non-target strand that becomes displaced as the 5′ flap species and that overlaps with the site to be edited.

[0226] In various embodiments of the extended guide RNAs, the DNA synthesis template sequence may encode a single-strand DNA flap that is complementary to an endogenous DNA sequence adjacent to a nick site, wherein the single-strand DNA flap comprises a desired nucleotide change. The single-stranded DNA flap may displace an endogenous single-strand DNA at the nick site. The displaced endogenous single-strand DNA at the nick site can have a 5′ end and form an endogenous flap, which can be excised by the cell. In various embodiments, excision of the 5′ end endogenous flap can help drive product formation since removing the 5′ end endogenous flap encourages hybridization of the single- strand 3′ DNA flap to the corresponding complementary DNA strand, and the incorporation or assimilation of the desired nucleotide change carried by the single-strand 3′ DNA flap into the target DNA.

[0227] The terms “cleavage site,” “nick site,” and “cut site” as used interchangeably herein in the context of prime editing, refer to a specific position in between two nucleotides or two B1195.70180WO00 12418099.1base pairs in the double-stranded target DNA sequence. In some embodiments, the position of a nick site is determined relative to the position of a specific PAM sequence. In some embodiments, the nick site is the particular position where a nick will occur when the double stranded target DNA is contacted with a napDNAbp, e.g., a nickase such as a Cas nickase, that recognizes a specific PAM sequence. For each PEgRNA described herein, a nick site (e.g., the “first nick site” when referred to in the context of PE3, PE5 and similar approaches), is characteristic of the particular napDNAbp to which the gRNA core of the PEgRNA associates with, and is characteristic of the particular PAM required for recognition and function of the napDNAbp. For example, for a PEgRNA that comprises a gRNA core that associates with a SpCas9, the nick site in the phosphodiester bond between bases three (“-3” position relative to the position 1 of the PAM sequence) and four (“-4” position relative to position 1 of the PAM sequence).

[0228] In some embodiments, a nick site is in a target strand of the double-stranded target DNA sequence. In some embodiments, a nick site is in a non-target strand of the double- stranded target DNA sequence. In some embodiments, the nick site is in a protospacer sequence. In some embodiments, the nick site is adjacent to a protospacer sequence. In some embodiments, a nick site is downstream of a region, e.g., on a non-target strand, that is complementary to a primer binding site of a PEgRNA. In some embodiments, a nick site is downstream of a region, e.g., on a non-target strand, that binds to a primer binding site of a PEgRNA. In some embodiments, a nick site is immediately downstream of a region, e.g., on a non-target strand, that is complementary to a primer binding site of a PEgRNA. In some embodiments, the nick site is upstream of a specific PAM sequence on the non-target strand of the double stranded target DNA, wherein the PAM sequence is specific for recognition by a napDNAbp that associates with the gRNA core of a PEgRNA. In some embodiments, the nick site is downstream of a specific PAM sequence on the non-target strand of the double stranded target DNA, wherein the PAM sequence is specific for recognition by a napDNAbp that associates with the gRNA core of a PEgRNA. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a Streptococcus pyogenes Cas9 nickase, a P. lavamentivorans Cas9 nickase, a C. diphtheriae Cas9 nickase, a N. cinerea Cas9, a S. aureus Cas9, or a N. lari Cas9 nickase. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active HNH domain and a nuclease inactive RuvC domain. In some embodiments, the nickB1195.70180WO00 12418099.1site is 2 base pairs upstream of the PAM sequence, and the PAM sequence is recognized by a S. thermophilus Cas9 nickase.

[0229] In various embodiments of the extended guide RNAs, the cellular repair of the single- strand DNA flap results in installation of the desired nucleotide change, thereby forming a desired product.

[0230] In still other embodiments, the desired nucleotide change is installed in an editing window that is between about -5 to +5 of the nick site, or between about -10 to +10 of the nick site, or between about -20 to +20 of the nick site, or between about -30 to +30 of the nick site, or between about -40 to +40 of the nick site, or between about -50 to +50 of the nick site, or between about -60 to +60 of the nick site, or between about -70 to +70 of the nick site, or between about -80 to +80 of the nick site, or between about -90 to +90 of the nick site, or between about -100 to +100 of the nick site, or between about -200 to +200 of the nick site.

[0231] In other embodiments, the desired nucleotide change is installed in an editing window that is between about +1 to +2 from the nick site, or about +1 to +3, +1 to +4, +1 to +5, +1 to +6, +1 to +7, +1 to +8, +1 to +9, +1 to +10, +1 to +11, +1 to +12, +1 to +13, +1 to +14, +1 to +15, +1 to +16, +1 to +17, +1 to +18, +1 to +19, +1 to +20, +1 to +21, +1 to +22, +1 to +23, +1 to +24, +1 to +25, +1 to +26, +1 to +27, +1 to +28, +1 to +29, +1 to +30, +1 to +31, +1 to +32, +1 to +33, +1 to +34, +1 to +35, +1 to +36, +1 to +37, +1 to +38, +1 to +39, +1 to +40, +1 to +41, +1 to +42, +1 to +43, +1 to +44, +1 to +45, +1 to +46, +1 to +47, +1 to +48, +1 to +49, +1 to +50, +1 to +51, +1 to +52, +1 to +53, +1 to +54, +1 to +55, +1 to +56, +1 to +57, +1 to +58, +1 to +59, +1 to +60, +1 to +61, +1 to +62, +1 to +63, +1 to +64, +1 to +65, +1 to +66, +1 to +67, +1 to +68, +1 to +69, +1 to +70, +1 to +71, +1 to +72, +1 to +73, +1 to +74, +1 to +75, +1 to +76, +1 to +77, +1 to +78, +1 to +79, +1 to +80, +1 to +81, +1 to +82, +1 to +83, +1 to +84, +1 to +85, +1 to +86, +1 to +87, +1 to +88, +1 to +89, +1 to +90, +1 to +90, +1 to +91, +1 to +92, +1 to +93, +1 to +94, +1 to +95, +1 to +96, +1 to +97, +1 to +98, +1 to +99, +1 to +100, +1 to +101, +1 to +102, +1 to +103, +1 to +104, +1 to +105, +1 to +106, +1 to +107, +1 to +108, +1 to +109, +1 to +110, +1 to +111, +1 to +112, +1 to +113, +1 to +114, +1 to +115, +1 to +116, +1 to +117, +1 to +118, +1 to +119, +1 to +120, +1 to +121, +1 to +122, +1 to +123, +1 to +124, or +1 to +125 from the nick site.

[0232] In still other embodiments, the desired nucleotide change is installed in an editing window that is between about +1 to +2 from the nick site, or about +1 to +5, +1 to +10, +1 to +15, +1 to +20, +1 to +25, +1 to +30, +1 to +35, +1 to +40, +1 to +45, +1 to +50, +1 to +55, +1 to +100, +1 to +105, +1 to +110, +1 to +115, +1 to +120, +1 to +125, +1 to +130, +1 to B1195.70180WO00 12418099.1+135, +1 to +140, +1 to +145, +1 to +150, +1 to +155, +1 to +160, +1 to +165, +1 to +170, +1 to +175, +1 to +180, +1 to +185, +1 to +190, +1 to +195, or +1 to +200, from the nick site.

[0233] In various aspects, the extended guide RNAs are modified versions of an extended guide RNA. pegRNAs (i.e., extended guide RNAs) and ngRNAs may be expressed from an encoding nucleic acid, or synthesized chemically. Methods are well known in the art for obtaining or otherwise synthesizing guide RNAs, and for determining the appropriate sequence of the pegRNA, including the protospacer sequence, which interacts and hybridizes with the target strand of a genomic target site of interest.

[0234] In various embodiments, the particular design aspects of a pegRNA sequence and ngRNA sequence will depend upon the nucleotide sequence of a genomic target site of interest (i.e., the desired site to be edited) and the type of napDNAbp (e.g., Cas9 protein) present in the prime editing systems utilized in the methods and compositions described herein, among other factors, such as PAM sequence locations, percent G / C content in the target sequence, the degree of microhomology regions, secondary structures, etc.

[0235] In general, a spacer sequence (i.e., a guide sequence) of a pegRNA or ngRNA can be any polynucleotide sequence having sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a napDNAbp (e.g., a Cas9, Cas9 homolog, or Cas9 variant) to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, a guide sequence is about or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length.

[0236] In some embodiments, a guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. The ability of a guide sequence to direct sequence- specific binding of a prime editor to a target sequence may be assessed by any suitable assay. For example, the components of a prime editor, including the guide sequence to be tested, B1195.70180WO00 12418099.1may be provided to a host cell having the corresponding target sequence, such as by transfection with vectors encoding the components of a prime editor disclosed herein, followed by an assessment of preferential cleavage within the target sequence, such as by Surveyor assay as described herein. Similarly, cleavage of a target polynucleotide sequence may be evaluated in a test tube by providing the target sequence, components of a prime editor, including the guide sequence to be tested and a control guide sequence different from the test guide sequence, and comparing binding or rate of cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible, and will occur to those skilled in the art.

[0237] A guide sequence may be selected to target any target sequence. In some embodiments, the target sequence is a sequence within a genome of a cell. Exemplary target sequences include those that are unique in the target genome. For example, for the S. pyogenes Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGG where NNNNNNNNNNNNXGG (N is A, G, T, or C; and X can be anything). A unique target sequence in a genome may include an S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGG where NNNNNNNNNNNXGG (N is A, G, T, or C; and X can be anything). For the S. thermophilus CRISPR1Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXXAGAAW where NNNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; and W is A or T). A unique target sequence in a genome may include an S. thermophilus CRISPR 1 Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXXAGAAW where NNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; and W is A or T). For the S. pyogenes Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGGXG where NNNNNNNNNNNNXGGXG (N is A, G, T, or C; and X can be anything). A unique target sequence in a genome may include an S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGGXG where NNNNNNNNNNNXGGXG (N is A, G, T, or C; and X can be anything). In each of these sequences “M” may be A, G, T, or C, and need not be considered in identifying a sequence as unique.

[0238] In some embodiments, a guide sequence is selected to reduce the degree of secondary structure within the guide sequence. Secondary structure may be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimal Gibbs free energy. An example of one such algorithm is mFold, as described by Zuker and Stiegler B1195.70180WO00 12418099.1(Nucleic Acids Res.9 (1981), 133-148). Another example folding algorithm is the online webserver RNAfold, developed at Institute for Theoretical Chemistry at the University of Vienna, using the centroid structure prediction algorithm (see, e.g., A. R. Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology1151- 62). Further algorithms may be found in U.S. Application Ser. No.61 / 836,080, incorporated herein by reference. In some embodiments, silent mutations are introduced in a guide sequence in order to alter its secondary structure and increase the efficiency of prime editing.

[0239] In some embodiments, the scaffold or gRNA core portion of a pegRNA comprises sequences corresponding to the tracr sequence and tracr mate sequence of a traditional guide RNA. In general, a tracr mate sequence includes any sequence that has sufficient complementarity with a tracr sequence to promote one or more of: (1) excision of a guide sequence flanked by tracr mate sequences in a cell containing the corresponding tracr sequence; and (2) formation of a complex at a target sequence, wherein the complex comprises the tracr mate sequence hybridized to the tracr sequence. In general, degree of complementarity is with reference to the optimal alignment of the tracr mate sequence and tracr sequence, along the length of the shorter of the two sequences. Optimal alignment may be determined by any suitable alignment algorithm, and may further account for secondary structures, such as self-complementarity within either the tracr sequence or tracr mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and tracr mate sequence along the length of the shorter of the two when optimally aligned is about or more than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or higher. In some embodiments, the tracr sequence is about or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and tracr mate sequence are contained within a single transcript, such that hybridization between the two produces a transcript having a secondary structure, such as a hairpin. Preferred loop forming sequences for use in hairpin structures are four nucleotides in length, and most preferably have the sequence GAAA. However, longer or shorter loop sequences may be used, as may alternative sequences. The sequences preferably include a nucleotide triplet (for example, AAA), and an additional nucleotide (for example C or G). Examples of loop forming sequences include CAAA and AAAG. In an embodiment of the invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins. In preferred embodiments, the transcript has two, three, four or five hairpins. In a further embodiment of the invention, the transcript has at most five hairpins. In some embodiments, the single transcript further includes a transcription B1195.70180WO00 12418099.1termination sequence; preferably this is a polyT sequence, for example six T nucleotides. Further non-limiting examples of single polynucleotides comprising a guide sequence, a tracr mate sequence, and a tracr sequence are as follows (listed 5′ to 3′), where “N” represents a base of a guide sequence, and the final poly-T sequence represents the transcription terminator:

[0240] (1)NNNNNNNNGTTTTTGTACTCTCAAGATTTAGAAATAAATCTTGCAGAAG CTACAAAGATAAGGCTTCATGCCGAAATCAACACCCTGTCATTTTATGGCAGGGTG TTTTCGTTATTTAATTTTTT (SEQ ID NO: 113); (2)NNNNNNNNNNNNNNNNNNGTTTTTGTACTCTCAGAAATGCAGAAGCTACAAA GATAAGGCTTCATGCCGAAATCAACACCCTGTCATTTTATGGCAGGGTGTTTTCGT TATTTAATTTTTT (SEQ ID NO: 114); (3)NNNNNNNNNNNNNNNNNNNNGTTTTTGTACTCTCAGAAATGCAGAAGCTACA AAGATAAGGCTTCATGCCGAAATCAACACCCTGTCATTTTATGGCAGGGTGTTTTT T (SEQ ID NO: 115); (4)NNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAA GGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 116); (5)NNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAA GGCTAGTCCGTTATCAACTTGAAAAAGTGTTTTTTT (SEQ ID NO: 117); and (6) NNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGG CTAGTCCGTTATCATTTTTTTT (SEQ ID NO: 118).

[0241] In some embodiments, sequences (1) to (3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4) to (6) are used in combination with Cas9 from S. pyogenes. In some embodiments, the tracr sequence is a separate transcript from a transcript comprising the tracr mate sequence.

[0242] It will be apparent to those of skill in the art that in order to target any of the fusion proteins comprising a Cas9 domain and a single-stranded DNA binding protein, as disclosed herein, to a target site, e.g., a site at which a recombinase recognition sequence is to be introduced, it is typically necessary to co-express the fusion protein together with a guide RNA, e.g., an sgRNA. As explained in more detail elsewhere herein, a guide RNA typically comprises a tracrRNA framework allowing for Cas9 binding, and a guide sequence, which confers sequence specificity to the Cas9:nucleic acid editing enzyme / domain fusion protein. B1195.70180WO00 12418099.1

[0243] In some embodiments, a pegRNA comprises a structure 5ʹ-[guide sequence]- GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAAGGCUAGUCCGUUAUCAACU UGAAAAAGUGGCACCGAGUCGGUGCUUUUU (SEQ ID NO: 119)-extension arm-3ʹ, wherein the guide sequence comprises a sequence that is complementary to the target sequence. The guide sequence, also referred to herein as the spacer sequence, is typically 20 nucleotides long. The sequences of suitable guide RNAs for targeting Cas9:nucleic acid editing enzyme / domain fusion proteins to specific genomic target sites will be apparent to those of skill in the art based on the instant disclosure. Such suitable guide RNA sequences typically comprise guide sequences that are complementary to a nucleic acid sequence within 50 nucleotides upstream or downstream of the target nucleotide to be edited. Some exemplary guide RNA sequences suitable for targeting any of the provided fusion proteins to specific target sequences are provided herein. Additional guide sequences are well known in the art and can be used with the prime editors utilized in the methods and compositions described herein.

[0244] In some embodiments, a PEgRNA comprises three main component elements ordered in the 5ʹ to 3ʹ direction, namely: a spacer, a gRNA core, and an extension arm at the 3ʹ end. In some embodiments, the extension arm may further be divided into the following structural elements in the 5ʹ to 3ʹ direction, namely: an edit template, a homology arm, and a primer binding site. In some embodiments, the extension arm may further be divided into the following structural elements in the 5ʹ to 3ʹ direction, namely: a homology arm, an edit template, and a primer binding site. In some embodiments, the extension arm may further be divided into the following structural elements in the 5ʹ to 3ʹ direction, namely: a DNA synthesis template (e.g., a RT template), and a primer binding site. In addition, the PEgRNA may comprise an optional 3ʹ end modifier region and an optional 5ʹ end modifier region . Still further, the PEgRNA may comprise a transcriptional termination signal at the 3ʹ end of the PEgRNA. These structural elements are further defined herein. The depiction of the structure of the PEgRNA is not meant to be limiting and embraces variations in the arrangement of the elements. For example, the optional sequence modifiers and could be positioned within or between any of the other regions shown, and not limited to being located at the 3ʹ and 5ʹ ends. PEgRNA modifications

[0245] The PEgRNAs may also include additional design modifications that may alter the properties and / or characteristics of PEgRNAs, thereby improving the efficacy of prime editing. In various embodiments, these modifications may belong to one or more of a number of different categories, including but not limited to: (1) designs to enable efficient expressionB1195.70180WO00 12418099.1of functional PEgRNAs from non-polymerase III (pol III) promoters, which would enable the expression of longer PEgRNAs without burdensome sequence requirements; (2) modifications to the core, Cas9-binding PEgRNA scaffold, which could improve efficacy; (3) modifications to the PEgRNA to improve RT processivity, allowing the insertion of longer sequences at targeted genomic loci; and (4) addition of RNA motifs to the 5ʹ or 3ʹ termini of the PEgRNA that improve PEgRNA stability, enhance RT processivity, prevent misfolding of the PEgRNA, or recruit additional factors important for genome editing. Such modifications are described further, for example, in PCT publication WO 2022 / 067130, which is incorporated herein by reference. Pharmaceutical compositions

[0246] Other aspects of the present disclosure relate to pharmaceutical compositions comprising any of the reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, or complexes provided herein, or any of the polynucleotides or vectors encoding such reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, or complexes provided herein. The term “pharmaceutical composition,” as used herein, refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition comprises additional agents (e.g., for specific delivery, increasing half-life, or other therapeutic compounds). In some embodiments, the pharmaceutical composition further comprises a pegRNA, or a polynucleotide encoding a pegRNA.

[0247] As used herein, the term “pharmaceutically-acceptable carrier” means a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc magnesium, calcium or zinc stearate, or steric acid), or solvent encapsulating material, involved in carrying or transporting the protein, fusion protein, polynucleotide, or vector from one site (e.g., the delivery site) of the body, to another site (e.g., an organ, tissue, or other part of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the tissue of the subject (e.g., physiologically compatible, sterile, physiologic pH, etc.). Some examples of materials that can serve as pharmaceutically-acceptable carriers include: (1) sugars, such as lactose, glucose and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose, and its derivatives, such as sodium carboxymethyl cellulose, methylcellulose, ethyl cellulose, microcrystalline cellulose and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricating agents, B1195.70180WO00 12418099.1such as magnesium stearate, sodium lauryl sulfate and talc; (8) excipients, such as cocoa butter and suppository waxes; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (10) glycols, such as propylene glycol; (11) polyols, such as glycerin, sorbitol, mannitol and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffering agents, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer’s solution; (19) ethyl alcohol; (20) pH buffered solutions; (21) polyesters, polycarbonates and / or polyanhydrides; (22) bulking agents, such as polypeptides and amino acids; (23) serum component, such as serum albumin, HDL, and LDL; (22) C2-C12 alcohols, such as ethanol; and (23) other non-toxic compatible substances employed in pharmaceutical formulations. Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, perfuming agents, preservatives, and antioxidants can also be present in the formulation. The terms such as “excipient,” “carrier,” “pharmaceutically acceptable carrier,” or the like are used interchangeably herein.

[0248] In some embodiments, the pharmaceutical composition is formulated for delivery to a subject, e.g., for gene editing. Suitable routes of administering the pharmaceutical composition described herein include, without limitation: topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, intradental, intracochlear, transtympanic, intraorgan, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseus, periocular, intratumoral, intracerebral, and intracerebroventricular administration.

[0249] In some embodiments, the pharmaceutical composition described herein is administered locally to a diseased site (e.g., tumor site). In some embodiments, the pharmaceutical composition described herein is administered to a subject by injection, by means of a catheter, by means of a suppository, or by means of an implant, the implant being of a porous, non-porous, or gelatinous material, including a membrane, such as a sialastic membrane, or a fiber.

[0250] In some embodiments, the pharmaceutical composition is formulated in accordance with routine procedures as a composition adapted for intravenous or subcutaneous administration to a subject, e.g., a human. In some embodiments, pharmaceutical compositions for administration by injection are solutions in sterile isotonic aqueous buffer. Where necessary, the pharmaceutical composition can also include a solubilizing agent and a local anesthetic such as lignocaine to ease pain at the site of the injection. Generally, the ingredients are supplied either separately or mixed together in unit dosage form, for example, B1195.70180WO00 12418099.1as a dry lyophilized powder or water free concentrate in a hermetically sealed container such as an ampoule or sachette indicating the quantity of active agent. Where the pharmaceutical composition is to be administered by infusion, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. Where the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration.

[0251] A pharmaceutical composition for systemic administration may be a liquid, e.g., sterile saline, lactated Ringer’s or Hank’s solution. In addition, the pharmaceutical composition can be in solid forms and re-dissolved or suspended immediately prior to use. Lyophilized forms are also contemplated.

[0252] The pharmaceutical composition can be contained within a lipid particle or vesicle, such as a liposome or microcrystal, which is also suitable for parenteral administration. The particles can be of any suitable structure, such as unilamellar or plurilamellar, so long as compositions are contained therein. Proteins, fusion proteins, polynucleotides, or vectors can be entrapped in “stabilized plasmid-lipid particles” (SPLP) containing the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipid and stabilized by a polyethyleneglycol (PEG) coating (Zhang Y. P. et al., Gene Ther.1999, 6:1438-47). Positively charged lipids, such as N-[1-(2,3-dioleoyloxi)propyl]-N,N,N- trimethyl-amoniummethylsulfate, or “DOTAP,” are particularly preferred for such particles and vesicles. The preparation of such lipid particles is well known. See, e.g., U.S. Patent Nos. 4,880,635; 4,906,477; 4,911,928; 4,917,951; 4,920,016; and 4,921,757; each of which is incorporated herein by reference.

[0253] The pharmaceutical compositions described herein may be administered or packaged as a unit dose, for example. The term “unit dose” when used in reference to a pharmaceutical composition of the present disclosure refers to physically discrete units suitable as unitary dosage for the subject, each unit containing a predetermined quantity of active material calculated to produce the desired therapeutic effect in association with the required diluent, i.e., carrier or vehicle.

[0254] Further, the pharmaceutical composition can be provided as a pharmaceutical kit comprising (a) a container containing a protein, fusion protein, complex (e.g., ribonucleoprotein complex), polynucleotide, or vector of the invention in lyophilized form and (b) a second container containing a pharmaceutically acceptable diluent (e.g., sterile water) for injection. The pharmaceutically acceptable diluent can be used for reconstitution or dilution of the lyophilized protein, fusion protein, complex (e.g., ribonucleoprotein complex), B1195.70180WO00 12418099.1polynucleotide, or vector of the invention. Optionally associated with such container(s) can be a notice in the form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceuticals or biological products, which notice reflects approval by the agency of manufacture, use, or sale for human administration.

[0255] In another aspect, an article of manufacture containing materials useful for the treatment of the diseases described above is included. In some embodiments, the article of manufacture comprises a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers may be formed from a variety of materials, such as glass or plastic. In some embodiments, the container holds a composition that is effective for treating a disease and may have a sterile access port. For example, the container may be an intravenous solution bag or a vial having a stopper pierce-able by a hypodermic injection needle. The active agent in the composition is a protein, fusion protein, polynucleotide, or vector of the invention. In some embodiments, the label on or associated with the container indicates that the composition is used for treating the disease of choice. The article of manufacture may further comprise a second container comprising a pharmaceutically acceptable buffer, such as phosphate-buffered saline, Ringer’s solution, or dextrose solution. It may further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, syringes, and package inserts with instructions for use. Polynucleotides, Vectors, AAVs, Kits, and Cells

[0256] In some aspects, the present disclosure provides polynucleotides and vectors encoding any of the reverse transcriptase variants, Cas9 variants, fusion proteins, or prime editors provided herein. In some aspects, the present disclosure provides one or more polynucleotides and vectors encoding any of the complexes provided herein. In some embodiments, the polynucleotides and vectors provided herein comprise DNA. In some embodiments, the polynucleotides and vectors provided herein comprise RNA. In some aspects, any of the polynucleotides described herein may be provided in a vector.

[0257] In some embodiments, one or more polynucleotides encoding any of the complexes provided herein are delivered to a cell, e.g., using an AAV. In certain embodiments, two polynucleotides encoding any of the complexes provided herein are delivered to a cell, e.g., using an AAV. In certain embodiments, the two polynucleotides comprise two halves of a prime editor described herein and comprising a split intein capable of reassembling into a prime editor molecule. In some embodiments, the one or more polynucleotides encoding a complex provided herein are delivered to a cell in one or more adeno-associated virus (AAV) B1195.70180WO00 12418099.1particles. Delivery of prime editor complexes has been described, for example, in U.S. Provisional Application, U.S.S.N., 63 / 426,336, filed November 17, 2022, U.S. Provisional Application, U.S.S.N., 63 / 491,013, filed March 17, 2023, and Davis, J. R., et al. Nat. Biotechnol.2023, each of which is incorporated herein by reference. In certain embodiments, the one or more polynucleotides encoding the complex are delivered to the cell in two AAV particles. In some embodiments, one or both of the AAV particles comprise AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9. In certain embodiments, one or both of the AAV particles comprise AAV9.

[0258] In some embodiments, a first and second AAV particle are delivered to a cell. In certain embodiments, the first AAV particle comprises a polynucleotide comprising the structure 5′-[inverted terminal repeat (ITR) sequence]-[promoter]-[napDNAbp N-terminal fragment]-[N-intein]-[terminator sequence]-[ITR sequence]-3′. In certain embodiments, the second AAV particle comprises a polynucleotide comprising the structure 5′-[ITR sequence]- [promoter]-[C-intein]-[napDNAbp C-terminal fragment]-[reverse transcriptase]-[terminator sequence]-[optional nicking gRNA]-[pegRNA]-[ITR]-3′.

[0259] The reverse transcriptase variants, Cas9 variants, prime editors, fusion proteins, or complexes provided herein may also be assembled into kits. In some embodiments, the kit comprises polynucleotides for expression of any of the reverse transcriptase variants, Cas9 variants, prime editors, or complexes provided herein. In other embodiments, the kit further comprises appropriate pegRNAs or nucleic acid vectors for the expression of such pegRNAs.

[0260] The kits described herein may include one or more containers housing components for performing the methods described herein, and optionally instructions for use. In some embodiments, the kits include instructions for editing a particular disease or treating a particular disease by prime editing. In some embodiments, the kits include instructions for editing a particular gene in a cell. Any of the kits described herein may further comprise components needed for performing any of the methods described herein. Each component of the kits, where applicable, may be provided in liquid form (e.g., in solution) or in solid form, (e.g., a dry powder). In certain cases, some of the components may be reconstitutable or otherwise processible (e.g., to an active form), for example, by the addition of a suitable solvent or other species (for example, water), which may or may not be provided with the kit.

[0261] In some embodiments, the kits may optionally include instructions and / or promotion for use of the components provided. As used herein, “instructions” can define a component of instruction and / or promotion, and typically involve written instructions on or associated with packaging of the disclosure. Instructions also can include any oral or electronic instructions B1195.70180WO00 12418099.1provided in any manner such that a user will clearly recognize that the instructions are to be associated with the kit, for example, audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communi...

Claims

CLAIMS What is claimed is:

1. A reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1, wherein the reverse transcriptase variant comprises amino acid substitutions at positions 70, 72, 87, 102, 106, 118, 128, 158, 269, 363, 413, and 492 relative to SEQ ID NO: 1, or corresponding substitutions in a homologous sequence.

2. The reverse transcriptase variant of claim 1, wherein the amino acid substitutions comprise P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N relative to SEQ ID NO:

1.

3. The reverse transcriptase variant of claim 1 or 2 further comprising amino acid substitutions at positions 188, 260, 297, and 288 relative to SEQ ID NO:

1.

4. The reverse transcriptase variant of claim 3, wherein the amino acid substitutions comprise S188K, I260L, S297Q, and R288Q relative to SEQ ID NO:

1.

5. A reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises amino acid substitutions at positions 128 and 200 relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence.

6. The reverse transcriptase variant of claim 5, wherein the amino acid substitutions comprise T128N and D200C relative to SEQ ID NO:

30.

7. The reverse transcriptase variant of claim 5 or 6 further comprising amino acid substitutions at positions 223, 306, 313, and 330 relative to SEQ ID NO:

30. B1195.70180WO00 12418099.

18. The reverse transcriptase variant of claim 7, wherein the amino acid substitutions comprise V223Y, T306K, W313F, and T330P relative to SEQ ID NO:

30.

9. A reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises the amino acid substitutions T128N and V223M; T128N and V223Y; T128F and V223M; or D200C and V223M relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence, optionally wherein the reverse transcriptase variant further comprises one or more of the amino acid substitutions D200N, T306K, W313F, T330P, and L603W relative to SEQ ID NO:

30.

10. A reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises amino acid substitutions at positions 128, 129, 196, 200, and 223 relative to SEQ ID NO: 30, or corresponding substitutions in a homologous sequence.

11. The reverse transcriptase variant of claim 10, wherein the amino acid substitutions comprise T128N; V129A or V129G; P196S, P196T, or P196F; N200S or N200Y; and V223A, V223M, V223L, or V223E relative to SEQ ID NO: 30, optionally wherein the reverse transcriptase variant further comprises one or more of the amino acid substitutions D200N, T306K, W313F, T330P, and L603W relative to SEQ ID NO:

30.

12. The reverse transcriptase variant of any one of claims 5-11, wherein the reverse transcriptase variant comprises a C-terminal truncation of part or all of the RNaseH domain of SEQ ID NO:

30.

13. The reverse transcriptase variant of claim 12, wherein the C-terminal truncation is between amino acid positions D497 and I498 of SEQ ID NO: 30.B1195.70180WO00 12418099.

114. A Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 775 and 918 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

15. The Cas9 variant of claim 14, wherein the amino acid substitutions comprise K775R and K918A relative to SEQ ID NO:

2.

16. A Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 99, 471, 632, 645, and 721 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

17. The Cas9 variant of claim 16, wherein the amino acid substitutions comprise H99R, E471K, I632V, D645N, and H721Y relative to SEQ ID NO:

2.

18. The Cas9 variant of claim 16 or 17 further comprising an amino acid substitution at position 654 relative to SEQ ID NO:

2.

19. The Cas9 variant of claim 18, wherein the amino acid substitution comprises R654C relative to SEQ ID NO:

2.

20. The Cas9 variant of any one of claims 16-19 further comprising an amino acid substitution at position 918 relative to SEQ ID NO:

2.

21. The Cas9 variant of claim 20, wherein the amino acid substitution comprises K918A relative to SEQ ID NO:

2.

22. A Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the B1195.70180WO00 12418099.1Cas9 variant comprises amino acid substitutions at positions 99, 471, and 632 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

23. The Cas9 variant of claim 22, wherein the amino acid substitutions comprise H99R, E471K, and I632V relative to SEQ ID NO:

2.

24. The Cas9 variant of claim 22 or 23 further comprising an amino acid substitution at position 721 relative to SEQ ID NO:

2.

25. The Cas9 variant of claim 24, wherein the amino acid substitution is H721Y relative to SEQ ID NO:

2.

26. A Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 471 and 918 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

27. The Cas9 variant of claim 26, wherein the amino acid substitutions comprise E471K and K918A relative to SEQ ID NO:

2.

28. A Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 753 and 1151 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

29. The Cas9 variant of claim 28, wherein the amino acid substitutions comprise R753G and K1151E relative to SEQ ID NO:

2.

30. A Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises one or more amino acid substitutions at positions selected from the group B1195.70180WO00 12418099.1consisting of 260, 298, 395, 769, 778, 1014, 1034, 1100, 1106, 1138, 1152, and 1320 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

31. The Cas9 variant of claim 30, wherein the one or more amino acid substitutions are selected from the group consisting of E260K, D298N, R395C, T769P, R778Q, K1014E, A1034E, V1100I, S1106F, T1138A, G1152E, and A1320T.

32. The Cas9 variant of claim 30 or 31 further comprising one or more additional amino acid substitutions at positions selected from the group consisting of 102, 753, 804, and 1003 relative to SEQ ID NO:

2.

33. The Cas9 variant of claim 32, wherein the one or more additional amino acid substitutions are selected from the group consisting of E102K, R753G, T804A, and K1003R.

34. The Cas9 variant of any one of claims 30-33 comprising amino acid substitutions at any one of the groups of positions: 102, 395, 753, 778, and 1100; 753, 769, 1034, and 1320; 298, 753, 1034, and 1138; 102, 260, 395, 753, 778, 804, 1003, 1100, 1106, and 1152; or 102, 260, 395, 753, 778, 804, 1003, 1014, 1100, 1106, and 1152; relative to SEQ ID NO:

2.

35. The Cas9 variant of any one of claims 30-34 comprising amino acid substitutions at any one of the groups of positions: E102K, R395C, R753G, R778Q, and V1100I; R753G, T769P, A1034E, and A1320T; D298N, R753G, A1034E, and T1138A; E102K, E260K, R395C, R753C, R778Q, T804A, K1003R, V1100I, S1106F, and G1152E; or B1195.70180WO00 12418099.1E102K, E260K, R395C, R753G, R778Q, T804A, K1003R, K1014E, V1100I, S1106F, and G1152E; relative to SEQ ID NO:

2.

36. A Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises amino acid substitutions at positions 23 and 754 relative to SEQ ID NO: 2, or corresponding substitutions in a homologous sequence.

37. The Cas9 variant of claim 36, wherein the amino acid substitutions are D23G and H754R.

38. A prime editor comprising a reverse transcriptase variant of any one of claims 1-13 and a nucleic acid-programmable DNA-binding protein (napDNAbp).

39. The prime editor of claim 38, wherein the napDNAbp comprises a Cas9 protein.

40. The prime editor of claim 39, wherein the Cas9 protein is a Cas9 nickase.

41. The prime editor of any one of claims 38-40, wherein the napDNAbp comprises a Cas9 variant of any one of claims 14-37.

42. The prime editor of any one of claims 38-40, wherein the napDNAbp comprises a Cas9 protein of any one of SEQ ID NOs: 2, 6, 8, 9, 12-24, or 133, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 2, 6, 8, 9, 12-24, or 133.

43. The prime editor of any one of claims 38-40, wherein the napDNAbp comprises a Cas9 protein of SEQ ID NO: 133, or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 133.B1195.70180WO00 12418099.

144. A prime editor comprising a Cas9 variant of any one of claims 14-37 and a polymerase.

45. The prime editor of claim 44, wherein the polymerase is a reverse transcriptase.

46. The prime editor of claim 45, wherein the reverse transcriptase is a reverse transcriptase variant of any one of claims 1-13.

47. The prime editor of claim 45, wherein the reverse transcriptase comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 7, wherein the reverse transcriptase comprises amino acid substitutions at positions 60, 87, 165, 243, 267, 279, 318, and 343 relative to SEQ ID NO: 7, or corresponding positions in a homologous sequence.

48. The prime editor of claim 47, wherein the amino acid substitutions comprise E60K, K87E, E165D, D243N, R267I, E279K, K318E, and K343N relative to SEQ ID NO:

7.

49. The prime editor of any one of claims 38-48, wherein the napDNAbp and the reverse transcriptase are provided in trans or are not fused to one another.

50. The prime editor of any one of claims 38-48, wherein the napDNAbp and the reverse transcriptase are provided as a fusion protein or are fused to one another.

51. The prime editor of claim 50, wherein the napDNAbp and the reverse transcriptase are fused via a linker.

52. The prime editor of claim 51, wherein the linker comprises any one of SEQ ID NOs: 80- 93.

53. The prime editor of any one of claims 38-52 further comprising a nuclear localization sequence (NLS). B1195.70180WO00 12418099.

154. A prime editor comprising a Cas9 protein and a reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1, wherein the reverse transcriptase variant comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N relative to SEQ ID NO:

1.

55. A prime editor comprising a Cas9 protein and a reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1, wherein the reverse transcriptase variant comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to SEQ ID NO:

1.

56. A prime editor comprising a Cas9 protein and a reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises the amino acid substitutions T128N, D200C, and V223Y relative to SEQ ID NO:

30.

57. The prime editor of any one of claims 54-56, wherein the Cas9 protein comprises a Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises the amino acid substitutions K775R and K918A; H99R, E471K, I632V, D645N, H721Y, and K918A; or H99R, E471K, I632V, D645N, R654C, and H721Y relative to SEQ ID NO:

2.

58. A prime editor comprising a reverse transcriptase and a Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises the amino acid substitutions K775R and K918A relative to SEQ ID NO:

2. B1195.70180WO00 12418099.

159. A prime editor comprising a reverse transcriptase and a Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises the amino acid substitutions H99R, E471K, I632V, D645N, H721Y, and K918A relative to SEQ ID NO:

2.

60. A prime editor comprising a reverse transcriptase and a Cas9 variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 2, wherein the Cas9 variant comprises the amino acid substitutions H99R, E471K, I632V, D645N, R654C, and H721Y relative to SEQ ID NO:

2.

61. The prime editor of any one of claims 58-60, wherein the reverse transcriptase comprises a reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1, wherein the reverse transcriptase variant comprises the amino acid substitutions P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, F269L, A363V, K413E, and S492N; or P70T, G72V, S87G, M102I, K106R, K118R, I128V, L158Q, S188K, I260L, F269L, R288Q, S297Q, A363V, K413E, and S492N relative to SEQ ID NO:

1.

62. The prime editor of any one of claims 58-60, wherein the reverse transcriptase comprises a reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 30, wherein the reverse transcriptase variant comprises the amino acid substitutions T128N, D200C, and V223Y relative to SEQ ID NO:

30.

63. The prime editor of any one of claims 58-60, wherein the reverse transcriptase comprises a reverse transcriptase variant having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 7, wherein the reverse transcriptase comprises the amino acid substitutions E60K, K87E, E165D, D243N, R267I, E279K, K318E, and K343N relative to SEQ ID NO:

7. B1195.70180WO00 12418099.

164. A prime editor comprising a Cas9 variant of any one of SEQ ID NOs: 28, 48, or 49 and a reverse transcriptase variant of any one of SEQ ID NOs: 25-27 or 50, or a Cas9 variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% to any one of SEQ ID NOs: 28, 48, or 49 and a reverse transcriptase variant at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% to any one of SEQ ID NOs: 25-27 or 50.

65. The prime editor of any one of claims 38-64, wherein the prime editor comprises the structure NH2-[bipartite NLS]-[Cas9]-[linker]-[reverse transcriptase]-[bipartite NLS]-[NLS].

66. The prime editor of any one of claims 38-65, wherein the prime editor comprises the fusion protein architecture of PEmax.

67. The prime editor of any one of claims 38-66, wherein the prime editor is smaller in size than PE2, and wherein the prime editor has an editing efficiency comparable to that of PE2.

68. The prime editor of any one of claims 38-67, wherein the prime editor has an increased editing efficiency compared to PEmax for edits that require structured pegRNA reverse transcriptase templates (RTTs).

69. A fusion protein comprising a Cas9 variant of any one of claims 14-37 and an effector domain.

70. The fusion protein of claim 69, wherein the effector domain comprises nuclease activity, nickase activity, recombinase activity, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity, transcriptional repression activity, or polymerase activity.

71. A reverse transcriptase variant comprising the sequence of any one of SEQ ID NOs: 25- 27 or 50, or a sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% to any one of SEQ ID NOs: 25-27 or 50.B1195.70180WO00 12418099.

172. A Cas9 variant comprising the sequence of any one of SEQ ID NOs: 28, 48, 49, or 145, or a sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% to any one of SEQ ID NOs: 28, 48, 49, or 145.

73. A prime editor comprising the sequence of any one of SEQ ID NOs: 155-158, or a sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% to any one of SEQ ID NOs: 155-158.

74. A complex comprising the prime editor of any one of claims 38-68 or the fusion protein of claim 69 or 70 and a pegRNA, optionally wherein the pegRNA is an epegRNA.

75. A polynucleotide encoding the reverse transcriptase variant of any one of claims 1-13 or 71.

76. A polynucleotide encoding the Cas9 variant of any one of claims 14-37 or 72.

77. One or more polynucleotides encoding the prime editor of any one of claims 38-68 or the fusion protein of claim 69 or 70.

78. A vector comprising the polynucleotide of claim 75 and / or the polynucleotide of claim 76.

79. The vector of claim 78 further comprising a polynucleotide encoding a pegRNA.

80. One or more vectors comprising the one or more polynucleotides of claim 77.

81. The one or more vectors of claim 80 further comprising a polynucleotide encoding a pegRNA. B1195.70180WO00 12418099.

182. One or more AAV particles comprising the one or more polynucleotides of any one of claims 75-77 or the one or more vectors of any one of claims 78-81.

83. The one or more AAV particles of claim 82, wherein the AAV particles comprise AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9.

84. The one or more AAV particles of claim 82 or 83, wherein the AAV particles comprise AAV9.

85. The one or more AAV particles of any one of claims 82-84, comprising a first AAV particle and a second AAV particle, wherein the first AAV particle comprises a polynucleotide comprising the structure 5′-[inverted terminal repeat (ITR) sequence]-[promoter]-[Cas9 N- terminal fragment]-[N-intein]-[terminator sequence]-[ITR sequence]-3′, and wherein the second AAV particle comprises a polynucleotide comprising the structure 5′-[ITR sequence]- [promoter]-[C-intein]-[Cas9 C-terminal fragment]-[reverse transcriptase]-[terminator sequence]- [optional nicking gRNA]-[pegRNA]-[ITR]-3′.

86. A cell comprising a reverse transcriptase variant of any one of claims 1-13 or 71, a Cas9 variant of any one of claims 14-37 or 72, a prime editor of any one of claims 38-68 or 73, a fusion protein of claim 69 or 70, a complex of claim 74, the one or more polynucleotides of any one of claims 75-77, the one or more vectors of any one of claims 78-81, or the one or more AAV particles of any one of claims 82-85.

87. A pharmaceutical composition comprising a reverse transcriptase variant of any one of claims 1-13 or 71, a Cas9 variant of any one of claims 14-37 or 72, a prime editor of any one of claims 38-68 or 73, a fusion protein of claim 69 or 70, a complex of claim 74, the one or more polynucleotides of any one of claims 75-77, the one or more vectors of any one of claims 78-81, the one or more AAV particles of any one of claims 82-85, or the cell of claim 86.

88. A method for editing a nucleic acid molecule by prime editing comprising contacting a nucleic acid molecule with a prime editor of any one of claims 38-68 or 73, a complex of claim B1195.70180WO00 12418099.174, one or more polynucleotides of any one of claims 75-77, or one or more vectors of any one of claims 78-81, thereby installing one or more modifications to the nucleic acid molecule at a target site.

89. A method for simultaneously editing both strands of a double-stranded nucleic acid molecule at a target site to be edited comprising contacting the double-stranded nucleic acid molecule with: (a) a prime editor of any one of claims 38-68, or a polynucleotide encoding the prime editor; (b) a first prime editing guide RNA (first pegRNA), or a polynucleotide encoding the first pegRNA that comprises (i) a first spacer sequence that binds to a first binding site on a first strand of the double-stranded DNA sequence upstream of the target site relative to the second strand, (ii) a first gRNA core that is capable of complexing with the prime editor, and (iii) a first DNA synthesis template that encodes a first single-stranded DNA sequence, and (c) a second prime editing guide RNA (second pegRNA), or a polynucleotide encoding the second pegRNA, that comprises (i) a second spacer sequence that binds to a second binding site on a second strand of the double-stranded DNA sequence downstream of the target site relative to the second strand; (ii) a second gRNA core that is capable of complexing with the prime editor, and (iii) a second DNA synthesis template that encodes a second single-stranded DNA sequence.

90. The method of claim 88 or 89, wherein the method further comprises contacting the nucleic acid molecule with one or more second strand nicking gRNAs.

91. The method of any one of claims 88-90, wherein the method comprises installing a recombinase recognition site in the nucleic acid molecule.B1195.70180WO00 12418099.

192. A kit comprising a reverse transcriptase variant of any one of claims 1-13 or 71, a Cas9 variant of any one of claims 14-37 or 72, a prime editor of any one of claims 38-68 or 73, a fusion protein of claim 69 or 70, a complex of claim 74, the one or more polynucleotides of any one of claims 75-77, the one or more vectors of any one of claims 78-81, the one or more AAV particles of any one of claims 82-85, or the cell of claim 86.

93. Use of a reverse transcriptase variant of any one of claims 1-13 or 71, a Cas9 variant of any one of claims 14-37 or 72, a prime editor of any one of claims 38-68 or 73, a fusion protein of claim 69 or 70, a complex of claim 74, the one or more polynucleotides of any one of claims 75-77, the one or more vectors of any one of claims 78-81, the one or more AAV particles of any one of claims 82-85, or the cell of claim 86 in the manufacture of a medicament.

94. The reverse transcriptase variant of any one of claims 1-13 or 71, the Cas9 variant of any one of claims 14-37 or 72, a prime editor of any one of claims 38-68 or 73, a fusion protein of claim 69 or 70, the complex of claim 74, the one or more polynucleotides of any one of claims 75-77, the one or more vectors of any one of claims 78-81, the one or more AAV particles of any one of claims 82-85, or the cell of claim 86 for use in medicine.

95. A system of polynucleotides for phage-assisted continuous and non-continuous evolution of prime editors comprising: i) a first polynucleotide encoding a pegRNA and the gIII gene; ii) a second polynucleotide encoding a Cas9 protein fused to an N-intein; iii) a third polynucleotide encoding an RNA polymerase; iv) a fourth polynucleotide encoding proteins capable of mutagenizing a phage, optionally wherein the fourth polynucleotide comprises the MP6 plasmid; and v) a fifth polynucleotide encoding a reverse transcriptase fused to a C-intein.

96. A system of polynucleotides for phage-assisted continuous and non-continuous evolution of prime editors comprising: i) a first polynucleotide encoding a pegRNA and the gIII gene; B1195.70180WO00 12418099.1ii) a second polynucleotide encoding a prime editor; iii) a third polynucleotide encoding an RNA polymerase; and iv) a fourth polynucleotide encoding proteins capable of mutagenizing a phage, optionally wherein the fourth polynucleotide comprises the MP6 plasmid. B1195.70180WO00 12418099.1