Self-assembling virus-like particles for delivery of prime editors and methods of making and using same

JP2024542790A5Pending Publication Date: 2025-12-10THE BROAD INST INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024533043
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-07
Filing Date
2022-12-02
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Current methods for delivering prime editing agents in vivo face challenges such as off-target editing, carcinogenesis risks, and inefficiencies due to viral delivery systems, particularly in primary cells with varying environmental sensitivity.

Method used

Development of virus-like particles (VLPs) that encapsulate prime editors (PEs) and associated guide RNAs, utilizing a gag-pro polyprotein, nucleocapsid protein, nuclear export sequence, and cleavable linkers to safely and efficiently deliver PE RNPs to target cells and tissues, leveraging natural budding mechanisms for release.

Benefits of technology

The VLPs enable precise and efficient delivery of prime editors to cells, minimizing off-target editing and ensuring high on-target editing efficacy, while avoiding integration into the genome and reducing carcinogenic risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides virus-like particles (VLPs) for delivering prime editors, and systems comprising such prime editor (PE) VLPs. The present disclosure also provides polynucleotides encoding the PE-VLPs described herein, which may be useful for producing the PE-VLPs. Also provided herein are methods for editing the genome of a target cell by introducing the PE-VLPs currently described into the target cell. The present disclosure also provides fusion proteins that constitute the components of the PE-VLPs described herein, as well as polynucleotides, vectors, cells, and kits.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims priority under 35 USC § 119(e) to U.S. provisional application USSN 63 / 285,995, filed December 3, 2021, U.S. provisional application USSN 63 / 298,626, filed January 11, 2022, and U.S. provisional application USSN 63 / 423,372, filed November 7, 2022, each of which is incorporated herein by reference.

[0002] Government support This invention was made with government support under Grant Nos. UG3AI150551, U01AI142756, R35GM118062, RM1HG009490, R01EY009339, and T32GM095450 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]

[0003] Background of the Invention Recently developed gene editing agents enable precise manipulation of genomic DNA in vivo, raising the possibility of treating the underlying causes of many genetic diseases (Anzalone et al., 2020; Doudna, 2020). The recent development of prime editing allows for the insertion, deletion, and replacement of genomic DNA sequences without the need for error-prone double-strand DNA breaks. See Anzalone et al., "Search-and-replace genome editing without double-strand breaks or donor DNA," Nature, 2019, Vol. 576, pp. 149-157, the contents of which are incorporated herein by reference. Prime editing uses an engineered Cas9 nickase-reverse transcriptase fusion protein paired with an engineered prime editing guide RNA (pegRNA), which not only directs Cas9 to the target genomic site but also encodes information for introducing the desired edit. Prime editing involves the following: 1) The Cas9 domain binds and nicks the target genomic DNA site specified by the spacer sequence of the pegRNA; 2) the reverse transcriptase domain initiates synthesis of an edited DNA strand using the nicked genomic DNA as a primer and the engineered extension on the pegRNA as a template for reverse transcription, which generates a single-stranded 3′ flap containing the edited DNA sequence; 3) cellular DNA repair resolves the 3' flap intermediate by displacement of the 5' flap species, which occurs via invasion by the edited 3' flap, excision of the 5' flap containing the original DNA sequence, and ligation of a new 3' flap to incorporate the edited DNA strand, forming a heteroduplex with one edited and one unedited strand; and 4) Cellular DNA repair uses the edited strand as a repair template to replace the unedited strand in the heteroduplex, completing the editing process. proceed through a multi-step editing process.

[0004] Broad therapeutic applications of in vivo prime editing require safe and efficient methods for delivering prime editors (PEs) to multiple tissues and organs. Adeno-associated viruses (AAVs) and lentiviruses (LVs) have been used to deliver DNA encoding gene editing agents to target tissues (Levy et al., 2020; Newby and Liu, 2021). However, viral delivery of editing agent-encoding DNA leads to persistent expression in transduced cells, which increases the frequency of off-target editing (Akcakaya et al., 2018; Davis et al., 2015; Wang et al., 2020; Yeh et al., 2018). In addition, viral delivery of DNA increases the likelihood of viral vector integration into the genome of transduced cells, both of which may promote oncogenesis or other adverse effects (Anzalone et al., 2020; Chandler et al., 2017). Furthermore, despite the constant evolution of transfection methods and the performance of viral delivery vectors (e.g., AAV or LV), the efficiency of these approaches can vary dramatically, especially in primary cells, which are highly sensitive to modifications of their environment and may even change in response to transfection agents and / or vectors.

[0005] One alternative to delivering gene editing agents (e.g., PE) in vivo would be to directly deliver proteins (e.g., PE) or ribonucleoproteins (RNPs) (e.g., PE complexed with pegRNA) instead of DNA. The short lifespan of RNPs within cells limits the opportunity for off-target editing. No generalizable strategy for delivering PE RNPs to multiple tissues and organs in vivo has been reported. Therefore, there is a need for a method to deliver PE ribonucleoproteins (RNPs) to target cells, tissues, or organs in need effectively and in a manner that improves overall safety by limiting and / or avoiding off-target editing without sacrificing targeted editing. Summary of the Invention

[0006] SUMMARY OF THE INVENTION

[0003] The present disclosure describes modifications of virus-like particles (VLPs) that package a prime editor (PE), an associated prime editor guide RNA (pegRNA), and other components that enable efficient prime editing. In one aspect, the present disclosure provides virus-like particles (interchangeably referred to herein as either "VLPs" or "eVLPs" ("engineered virus-like particles")) comprising a group-specific antigen (gag) protease (pro) polyprotein, and one or more fusion proteins, wherein the gag-pro polyprotein and one or more fusion proteins are encapsulated by a lipid membrane and viral envelope glycoproteins, and wherein each of the one or more fusion proteins comprises one of the following: (i) gag nucleocapsid protein; (ii) nuclear export sequence (NES); (iii) a cleavable linker; and (iv) nucleic acid programmable DNA binding protein (napDNAbp) and / or a domain containing RNA-dependent DNA polymerase activity; In some embodiments, the fusion protein comprises both the napDNAbp and the domain comprising RNA-dependent DNA polymerase activity. In some embodiments, the VLP comprises a first fusion protein comprising the napDNAbp and a second fusion protein comprising the domain comprising RNA-dependent DNA polymerase activity. In certain embodiments, the first and second fusion proteins each comprise a portion of a split intein that facilitates fusion of the napDNAbp and the domain comprising RNA-dependent DNA polymerase activity with each other after delivery of the VLP to a target cell. Without being bound by theory, the components of the VLPs provided herein self-assemble at the cell membrane and bud according to a naturally occurring budding mechanism (e.g., that of retroviruses or other enveloped viruses) to release the fully mature VLP from the cell. Once formed, Gag-Pol-Pro cleaves the protease-sensitive linker of Gag-cargo (i.e., [Gag]-[cleavable linker]-[cargo], where cargo can be, for example, PE-RNP), thereby releasing the PE RNP within the VLP. Thus, in various embodiments, the present disclosure also provides VLPs in which the protease-sensitive linker has been cleaved (e.g., to produce two cleaved products comprising (i) a fusion protein comprising a gag nucleocapsid protein and a nuclear export sequence, and (ii) a prime editor). For example, the present disclosure provides: (i) group-specific antigen (gag) protease (pro) polyprotein; (ii) a prime editor protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a domain containing RNA-dependent DNA polymerase activity (e.g., reverse transcriptase); and (iii) a fusion protein encapsulated by a lipid membrane and viral envelope glycoproteins, containing the gag nucleocapsid protein and a nuclear export sequence (NES); In some embodiments, the present disclosure provides a VLP comprising a mixture of cleaved and uncleaved products (i.e., some prime editors have been cleaved and released from the gag protein, while others have not yet been cleaved from the gag protein). In some embodiments, 50% or more, 60% or more, 70% or more, 80% or more, or 90% or more of the prime editors are cleaved from the gag protein within the VLP. Once the VLP is administered to and taken up by a recipient cell, the contents of the VLP are released, for example, PE RNPs. Once inside the cell, the RNPs translocate to the cell's nucleus (particularly where a nuclear localization signal (NLS) is linked to the RNP), where DNA editing may occur at the target site specified by the guide RNA. The present disclosure also provides polynucleotides and vectors encoding various components of the VLPs described herein.

[0007] In another aspect, the present disclosure provides a method for manufacturing a cellular membrane comprising: (i) a first polynucleotide comprising a nucleic acid sequence encoding a viral envelope glycoprotein; (ii) a second polynucleotide comprising a nucleic acid sequence encoding a group-specific antigen (gag) protease (pro) polyprotein; (iii) a third polynucleotide comprising a nucleic acid sequence encoding one or more fusion proteins, wherein each of the one or more fusion proteins comprises: (a) gag nucleocapsid protein; (b) nuclear export sequence (NES); (c) a cleavable linker; and (d) nucleic acid programmable DNA binding protein (napDNAbp) and / or a domain containing RNA-dependent DNA polymerase activity; the polynucleotide comprising (iv) a fourth polynucleotide comprising a nucleic acid sequence encoding a guide RNA (gRNA), wherein the gRNA binds to the napDNAbp of the fusion protein encoded by the third polynucleotide; In some embodiments, the pharmaceutical composition comprises a plurality of polynucleotides comprising: (i) group-specific antigen (gag) protease (pro) polyprotein; (ii) a prime editor protein comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a domain containing RNA-dependent DNA polymerase activity (e.g., reverse transcriptase); and (iii) a fusion protein encapsulated by a lipid membrane and viral envelope glycoproteins, including the gag nucleocapsid protein and a nuclear export sequence (NES).

[0008] In another aspect, the disclosure provides a pharmaceutical composition comprising a virus-like particle (VLP) comprising a group-specific antigen (gag) protease (pro) polyprotein and one or more fusion proteins, wherein the gag-pro polyprotein and the one or more fusion proteins are encapsulated by a lipid membrane and a viral envelope glycoprotein, and wherein each of the one or more fusion proteins comprises: (i) gag nucleocapsid protein; (ii) nuclear export sequence (NES); (iii) a cleavable linker; and (iv) nucleic acid programmable DNA binding protein (napDNAbp) and / or a domain containing RNA-dependent DNA polymerase activity; Includes.

[0009] In another aspect, the present disclosure provides a method for editing a nucleic acid molecule in a target cell by prime editing, comprising contacting the target cell with any of the compositions provided herein, thereby introducing one or more modifications into the nucleic acid molecule at a target site. In some embodiments, the cell is a mammalian cell (e.g., a human cell). In some embodiments, the cell is from an animal relevant to veterinary or agricultural use. In some embodiments, the cell is within a subject. In some embodiments, the subject is a human. In some embodiments, the one or more modifications to the nucleic acid molecule are associated with reducing, alleviating, or preventing symptoms of a disease or disorder.

[0010] In another aspect, the present disclosure provides a method for manufacturing a semiconductor device comprising: (i) gag nucleocapsid protein; (ii) nuclear export sequence (NES); (iii) a cleavable linker; and (iv) nucleic acid programmable DNA binding protein (napDNAbp) and / or a domain containing RNA-dependent DNA polymerase activity; In some embodiments, the fusion protein comprises both a napDNAbp and a domain comprising RNA-dependent DNA polymerase activity. In some embodiments, the disclosure provides a composition comprising a first fusion protein disclosed herein comprising a napDNAbp, and a second fusion protein disclosed herein comprising a domain comprising RNA-dependent DNA polymerase activity. In certain embodiments, the first and second fusion proteins each comprise a portion of a split intein to facilitate fusion of the napDNAbp and the domain comprising RNA-dependent DNA polymerase activity to each other (e.g., after delivery of the fusion protein in a VLP disclosed herein to a target cell).

[0011] In other aspects, the present disclosure also provides methods for producing the PE-VLPs described herein, and methods for prime editing comprising delivering the PE-VLPs described herein to target cells. Polynucleotides, vectors, cells, and kits comprising the PE-VLPs and fusion proteins described herein are also provided.

[0012] In another aspect, the present disclosure provides VLPs produced by transfecting, transducing, electroporating, or otherwise inserting any of the polynucleotides or vectors disclosed herein into a cell, and expressing components of the VLP from the polynucleotide or vector, thereby causing spontaneous assembly of virus-like particles within the cell. In some embodiments, any of the compositions, methods, or cells described herein may be used to produce the VLPs provided herein.

[0013] In another aspect, the present disclosure provides compositions comprising any of the VLPs, polynucleotides, vectors, and fusion proteins provided herein.

[0014] In another aspect, the present disclosure provides methods of editing a nucleic acid molecule in a target cell using any of the VLPs, polynucleotides, compositions, and fusion proteins provided herein.

[0015] In another aspect, the disclosure provides a cell comprising any of the VLPs, polynucleotides, vectors, compositions, and fusion proteins described herein.

[0016] In another aspect, the disclosure provides kits comprising any of the VLPs, polynucleotides, vectors, compositions, and fusion proteins described herein.

[0017] It should be understood that the foregoing concepts, and additional concepts discussed below, may be arranged in any suitable combination, as the disclosure is not limited in this respect. Furthermore, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments when considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0018] BRIEF DESCRIPTION OF THE DRAWINGS The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure and may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.

[0019] [Figure 1] Figure 1: Overview of delivery methods for CRISPR / Cas systems developed to date.

[0020] [Figure 2A] Figure 2A: Overview of the prime editor ribonucleoprotein (PE-RNP) virus-like particle (VLP) delivery strategy. [Figure 2B] Figure 2B: Overview of the prime editor ribonucleoprotein (PE-RNP) virus-like particle (VLP) delivery strategy. [Figures 2C-2D] Figures 2C-2D: Overview of the prime editor ribonucleoprotein (PE-RNP) virus-like particle (VLP) delivery strategy.

[0021] [Figure 3] Figures 3A-3B: PE-RNP VLP optimization of single particle versus two particle systems. The single particle system is shown to be more efficient than the two particle system.

[0022] [Figure 4] Figures 4A-4B: PE-RNP VLP optimization of 1X NLS system vs. 2X NLS system. Incorporation of two NLSs was shown to improve editing efficiency.

[0023] [Figure 5] Figure 5: Optimization contributes to packaging of editors into VLPs. Incorporation of an NES facilitates transport of PE into the cytoplasm of producing cells. Gag fusions direct packaging of the editor into VLPs.

[0024] [Figure 6] Figure 6: Efficiency of HEK3 +1 T>A editing in HEK293T cells using various concentrations compared to plasmid transfection.

[0025] [Figure 7] Figure 7: Schematic diagram of pegRNA and prime editors.

[0026] [Figure 8A] Figures 8A-8C: Assessment of pegRNA packaging. Supplementation of pegRNA by plasmid transfection has been shown to enhance editing efficiency. In contrast, editing using adenosine base editors (ABEs) is not significantly improved by sgRNA transfection. [Figure 8B-8C] Figures 8B-8C: Assessment of pegRNA packaging. Supplementation of pegRNA by plasmid transfection has been shown to enhance editing efficiency. In contrast, editing using adenosine base editors (ABEs) is not significantly improved by sgRNA transfection.

[0027] [Figure 9] Figure 9: Assessment of the binding affinity of pegRNA to PE. pegRNA has been shown to have lower binding affinity for Cas9 compared to sgRNA.

[0028] [Figure 10]Figures 10A-10B: Employing the F+E scaffold to improve pegRNA binding. The F+E scaffold was shown to modestly improve pegRNA binding to Cas9 in a restrictive pegRNA context.

[0029] [Figures 11A-11B] Figures 11A-11B: Incorporation of the MS2 stem-loop for specific packaging of pegRNA. [Figures 11C-11D] Figures 11C-11D: Incorporation of the MS2 stem-loop for specific packaging of pegRNA. [Figure 11E] FIG. 11E: Incorporation of the MS2 stem-loop for specific packaging of pegRNA.

[0030] [Figure 12] Figure 12: Incorporation of PEmax for more robust editing. Delivery of PEmax using VLPs was shown to result in improved editing efficiency.

[0031] [Figure 13] Figure 13: Assessment of PE packaging. Qualitative assessment of Cas9 content by dot blot.

[0032] [Figures 14A-14B] Figures 14A-14B: Trimming the polymerase domain increases cargo space within the VLP. [Figure 14C] FIG. 14C: The polymerase domain is trimmed to increase cargo space within the VLP.

[0033] [Figure 15] Figure 15: PE3max RNP VLP system. The use of a 30% nicking gRNA is shown to lead to the highest editing efficiency. Approximately a 3.5-fold improvement is observed compared to PE2max.

[0034] [Figure 16A]Figure 16A: Comparison of the PE3max RNP VLP separated particle system vs. the all-in-one particle system. Varying ratios of VLP (editor + ngRNA) to VLP (editor + pegRNA) were screened in 50 μl of total VLP. The separated particle system has been shown to have editing efficiency comparable to the all-in-one particle system. [Figure 16B] Figure 16B: Comparison of the PE3max RNP VLP separated particle system vs. the all-in-one particle system. Varying ratios of VLP (editor + ngRNA) to VLP (editor + pegRNA) were screened in 50 μl of total VLP. The separated particle system has been shown to have editing efficiency comparable to the all-in-one particle system.

[0035] [Figure 17A] Figure 17A: PE3max RNP VLP separated particle system with varying transduction timing. The all-in-one particle system shows increased editing efficiency. [Figure 17B] Figure 17B: PE3max RNP VLP separated particle system with varying transduction timing. The all-in-one particle system shows increased editing efficiency.

[0036] [Figure 18] Figure 18: Mismatch repair-privileged editing was shown to lead to higher overall editing in both PE2 and PE3 RNP VLPs, suggesting that introducing silent mutations to bypass MMR may also improve editing efficiency, especially in PE-limited situations such as RNP VLP systems.

[0037] [Figure 19A]Figure 19A: PE4max ribonucleoprotein VLP. MLH1dn protein was packaged into VLPs using both the all-in-one particle system and the separated particle system. Dual transfection-transduction experiments demonstrated the following: 1) MLH1dn plasmid transfection offered a significant improvement to PE2 VLP editing efficiency, indicating that circumventing MMR plays an important role in improving PE-VLP editing efficiency, and 2) MLH1dn was packaged into VLP particles. [Figures 19B-19C] Figures 19B-19C: PE4max ribonucleoprotein VLPs. MLH1dn protein was packaged into VLPs using both the all-in-one particle system and the separated particle system. Dual transfection-transduction experiments demonstrated the following: 1) MLH1dn plasmid transfection offered a significant improvement to PE2 VLP editing efficiency, indicating that circumventing MMR plays an important role in improving PE-VLP editing efficiency; and 2) MLH1dn was packaged into VLP particles. [Figure 19D] Figure 19D: PE4max ribonucleoprotein VLP. MLH1dn protein was packaged into VLPs using both the all-in-one particle system and the separated particle system. Dual transfection-transduction showed the following: 1) MLH1dn plasmid transfection offered a significant improvement to PE2 VLP editing efficiency, indicating that circumventing MMR plays an important role in improving PE-VLP editing efficiency, and 2) MLH1dn was packaged into VLP particles.

[0038] [Figure 20] Figure 20: Introducing silent mutations improves PE RNP VLPs. PE VLPs have similar editing efficiency to plasmid transfection when MMR is sufficiently circumvented.

[0039] [Figure 21] Figure 21: Assessment of PE assembly. Variable expression of Cas9 and RT halves and inefficient intein trans-splicing can lead to poisoning of editing sites.

[0040] [Figure 22A] Figure 22A: Optimization of full-length PE and Cas9 internal cleavage. The pmA97 construct (full-length PE with RT protease site deletion) showed the highest editing efficiency. At the C-terminus of RT, there is a protease cleavage site that can be recognized by the MMLV protease expressed in the system. When the protease recognizes and cleaves this site, the NLS at the C-terminus of RT is also cleaved from the prime editor. Therefore, deleting the RT protease site improves editing efficiency. In Figure 22B, the sequences shown correspond to SEQ ID NOs: 232-234 (top to bottom). [Figure 22B] Figure 22B: Optimization of full-length PE and Cas9 internal cleavage. The pmA97 construct (full-length PE with RT protease site deletion) showed the highest editing efficiency. At the C-terminus of RT, there is a protease cleavage site that can be recognized by the MMLV protease expressed in the system. When the protease recognizes and cleaves this site, the NLS at the C-terminus of RT is also cleaved from the prime editor. Therefore, deleting the RT protease site improves editing efficiency. In Figure 22B, the sequences shown correspond to SEQ ID NOs: 232-234 (top to bottom).

[0041] [Figure 23] Figures 23A-23B: Optimization of full-length PE and Cas9 internal split. Full-length PE showed higher editing efficiency than split PE.

[0042] [Figure 24] Figures 24A-24B: Validation of the Cas9-mRNA VLP strategy.

[0043] [Figure 25]Figures 25A-25B: Editing efficiency of PE2max mRNA VLP version 1.

[0044] [Figure 26] 26A-26B: Whole editor structure shows higher editing efficiency than split editor structure. Splitting the editor structure did not improve editing.

[0045] [Figures 27A-27B] Figures 27A-27B: Editing efficiency of PE2max mRNA VLP version 2. The Psi signal of the pLV vector allows only two copies of the viral genome into the particle. The MS2-stem-loop inserted RNA can also increase pegRNA packaging. [Figure 27C] Figure 27C: Editing efficiency of PE2max mRNA VLP version 2. The Psi signal of the pLV vector allows only two copies of the viral genome into the particle. The MS2-stem-loop inserted RNA can also increase pegRNA packaging.

[0046] [Figure 28A] Figure 28A: In PEmax mRNA VLP design version 2, the HIV capsid was changed to an MMLV capsid. The MMLV capsid led to higher titer production. Expression of the pegRNA in a lentiviral expression vector allows for packaging of a more functional pegRNA than in a traditional plasmid backbone. [Figures 28B-28C] Figures 28B-28C: In PEmax mRNA VLP design version 2, the HIV capsid was replaced with an MMLV capsid. The MMLV capsid led to higher titer production. Expression of the pegRNA in a lentiviral expression vector allows for packaging of a more functional pegRNA than in a conventional plasmid backbone.

[0047] [Figure 29]Figures 29A-29B: Optimization of MCP-fused gag protein in PE2max mRNA VLP version 2. The polymerase domain is important in the virus production process.

[0048] [Figure 30] Figure 30: Additional MCP fusion structures.

[0049] [Figure 31] Figure 31: PE2max mRNA VLP version 2. Features include 6x MS2 stem loops utilized for packaging of the transgene mRNA.

[0050] [Figure 32] Figure 32 shows modifications of the split prime editor for more efficient packaging. The full-length editor construct generally led to higher editing efficiency. Deleting six amino acids at the C-terminus of MMLV reverse transcriptase, removing the endogenous protease cleavage site and preventing cleavage of the NLS on the prime editor, increased editing efficiency in both the full-length and split prime editor constructs.

[0051] [Figure 33] Figure 33 provides a schematic showing that a portion of the prime editor delivered by an eVLP may still retain an NES after protease cleavage.

[0052] [Figure 34] Figures 34A-34B show alterations to the position of the NES to ensure cleavage from the prime editor. Sites with Gag proteins that can tolerate larger insertions were investigated. Inserting a 3xNES before the endogenous protease cleavage site between the p12 and CA domains (NES position 1) resulted in the highest editing efficiency.

[0053] [Figure 35]Figures 35A-35B show the addition of a linker to better expose the protease cleavage site. SEQ ID NO: 163 (SGGSSGGS) is shown.

[0054] [Figure 36] Figure 36 shows the combination of optimized NES positions and linker sequences. The V5 eVLP configurations incorporate these optimized NES positions and linker sequences.

[0055] [Figure 37] Figures 37A-37B show that the mismatch repair (MMR) pathway can be particularly disruptive to PE-eVLP editing efficiency: MMR-privileged editing leads to higher overall editing in both PE2 and PE3 RNP VLPs.

[0056] [Figure 38A-38B] Figures 38A-38B show the packaging of MLHdn in eVLPs. MLHdn-eVLP transduction showed editing efficiency similar to that of PE2 plasmid transfection. The amount of packaged MLHdn may not be sufficient to suppress MMR. [Figure 38C] Figure 38C shows the packaging of MLHdn in eVLPs. MLHdn-eVLP transduction showed editing efficiency similar to that of PE2 plasmid transfection. The amount of packaged MLHdn may not be sufficient to suppress MMR.

[0057] [Figure 39A] Figure 39A shows the introduction of additional sequential mutations to circumvent MMR. Introducing additional sequential mutations is a promising strategy for escaping MMR because no additional components need to be packaged into the eVLP. In Figure 39A, the sequences correspond to SEQ ID NOs: 235-242 (top to bottom). [Figure 39B]Figure 39B shows the introduction of additional sequential mutations to circumvent MMR. Introducing additional sequential mutations is a promising strategy for escaping MMR because no additional components need to be packaged into the eVLP. In Figure 39A, the sequences correspond to SEQ ID NOs: 235-242 (top to bottom).

[0058] [Figure 40A-40B] 40A-40B show the inclusion of the MS2 stem-loop for specific packaging of pegRNA. Insertion of the MS2 aptamer into the scaffold region of pegRNA improves packaging of pegRNA through interaction with MCP-Gag-pol. [Figure 40C-40D] Figures 40C-40D show the inclusion of the MS2 stem-loop for specific packaging of pegRNA. Insertion of the MS2 aptamer into the scaffold region of pegRNA improves packaging of pegRNA through interaction with MCP-Gag-pol.

[0059] [Figure 41A-41B] Figures 41A-41B show the inclusion of an MS2 stem-loop to facilitate nicking guide RNA (ngRNA) packaging for PE3. The MS2 aptamer was shown to improve ngRNA packaging. An all-in-one particle system incorporating both MS2-pegRNA and MS2-ngRNA was demonstrated to provide the highest PE3 editing efficiency. [Figure 41C] Figure 41C shows the inclusion of an MS2 stem-loop to facilitate nicking guide RNA (ngRNA) packaging for PE3. The MS2 aptamer was shown to improve ngRNA packaging. An all-in-one particle system incorporating both MS2-pegRNA and MS2-ngRNA was demonstrated to provide the highest PE3 editing efficiency.

[0060] [Figure 42]Figures 42A-42B show that the use of com proteins and com aptamers is comparable to the MCP-MS2 aptamer system.

[0061] [Figure 43A-43B] 43A-43B show the optimization of plasmid ratios for VLP production. Specifically, the ratios of Gag-pol, MCP-Gag-pol, and Gag-cargo were optimized as shown. [Figure 43C] Figure 43C shows the optimization of plasmid ratios for VLP production. Specifically, the ratios of Gag-pol, MCP-Gag-pol, and Gag-cargo were optimized as shown.

[0062] [Figure 44] Figures 44A-44B show the use of coiled-coil peptides as an additional mechanism for recruitment of prime editors in VLPs. In Figure 44A, the P4 peptide domain is shown upside down, indicating an antiparallel coiled-coil structural design.

[0063] [Figure 45] Figures 45A-45B show that coiled-coil peptide-prime editor constructs improve editing efficiency.

[0064] [Figures 46A-46B] Figures 46A-46B provide a schematic of the coiled-coil peptide-prime editor structure and show that the MCP fusion structure provides superior editing efficiency over the coiled-coil structure. [Figure 46C-46D] Figures 46C-46D provide schematic diagrams of coiled-coil peptide-prime editor structures and show that the MCP fusion structure provides superior editing efficiency over the coiled-coil structure.

[0065] [Figure 47]Figures 47A-47B show in vivo studies in P0 mice by ICV injection with PE VLPs. PE VLPs demonstrated efficient editing in a cell population transducible by VSV-g.

[0066] [Figure 48] Figure 48 shows the testing of PE VLPs in vivo by subretinal injection in rd6 model mice. Correction of the gene encoding retinal disease-associated membrane frizzled-related protein (Mfrp) was observed.

[0067] [Figure 49A-49B] Figures 49A-49B show further testing of PE VLPs in vivo by subretinal injection in rd6 model mice. An average of 15% editing was observed with PE3 VLPs and protein restoration. [Fig. 49C-49D] Figures 49C-49D show further testing of PE VLPs in vivo by subretinal injection in rd6 model mice. An average of 15% editing was observed with PE3 VLPs and protein renaturation.

[0068] [Figure 50] Figures 50A-50B show further optimization of PE VLPs for subretinal injection in rd12 model mice using additional silent mutations in the pegRNA and various concentrations of VLPs containing either PE2 or PE3.

[0069] [Figure 51] Figure 51 shows an additional strategy for recruitment of prime editors to eVLPs via coiled-coil peptides.

[0070] [Figure 52] Figure 52 shows that an evolved small reverse transcriptase (Tf1) can be used in a prime editor delivered by eVLP. DETAILED DESCRIPTION OF THE INVENTION

[0071] definition Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide those of ordinary skill in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless otherwise specified.

[0072] Cas9 The term "Cas9" or "Cas9 nuclease" refers to an RNA-guided nuclease containing a Cas9 domain, or a fragment thereof (e.g., a protein containing an active or inactive DNA cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9). As used herein, a "Cas9 domain" is a protein fragment containing an active or inactive cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9. A "Cas9 protein" is a full-length Cas9 protein. Cas9 nuclease is also referred to as casn1 nuclease or CRISPR ( C Lustered R regularly I Interspaced S hort P alindromicRCRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain a spacer, a sequence complementary to the preceding mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, the correct processing of the pre-crRNA requires a transcoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA acts as a guide for ribonuclease 3-assisted pre-crRNA processing. The Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to the crRNA is first endonucleolytically cleaved and then 3' to 5' exonucleolytically trimmed. In nature, DNA binding and cleavage typically require both protein and RNA. However, single guide RNA ("sgRNA" or simply "gRNA") can be designed to incorporate both crRNA and tracrRNA aspects into a single RNA species. For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence, which helps distinguish self from non-self.The sequence and structure of Cas9 nuclease are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been reported in various species, including, but not limited to, Streptococcus pyogenes and Streptococcus thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.

[0073] Nuclease-inactivated Cas9 domains are sometimes referred to interchangeably as "dCas9" proteins (dead Cas9). Methods for generating Cas9 domains (or fragments thereof) with inactive DNA cleavage domains are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28;152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can suppress the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83 (2013)). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, proteins comprise two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) Cas9 DNA cleavage domain In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant." A Cas9 variant shares homology with Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 37). In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to a wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 37). In some embodiments, the Cas9 variant comprises a fragment of SEQ ID NO: 37Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), wherein the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 37). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 37).

[0074] CRISPR CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent fragments of a previous infection by a virus that has invaded a prokaryote. The DNA fragments are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with an array of CRISPR-associated proteins (including Cas9 and its homologs) and CRISPR-associated RNAs, effectively constitute the prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In some types of CRISPR systems (e.g., type II CRISPR systems), correct processing of the pre-crRNA requires a transcoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA acts as a guide for processing the pre-crRNA, assisted by ribonuclease 3. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets complementary to the RNA. Specifically, the target strand not complementary to the crRNA is first endonucleolytically cleaved and then 3'-5' exonucleolytically trimmed. In nature, DNA binding and cleavage typically require both a protein and an RNA. However, single guide RNAs ("sgRNAs" or simply "gRNAs") can be designed to incorporate aspects of both the crRNA and tracrRNA into a single RNA species (guide RNA). For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence, which helps distinguish self from non-self.The biology of CRISPR and the sequence and structure of Cas9 nuclease are well known to those skilled in the art (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM, See Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference. Cas9 orthologs have been reported in various species, including, but not limited to, Streptococcus pyogenes and Streptococcus thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure, and include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.

[0075] In some types of CRISPR systems (e.g., type II CRISPR systems), proper processing of the pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA functions as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular nucleic acid targets complementary to the RNA. Specifically, the target strand not complementary to the crRNA is first endonucleolytically cleaved and then 3'-5' exonucleolytically trimmed. In nature, DNA binding and cleavage typically require both a protein and both RNAs. However, single guide RNAs ("sgRNAs" or simply "gRNAs") can be designed to incorporate aspects of both the crRNA and tracrRNA into a single RNA species (guide RNA).

[0076] In general, "CRISPR system" refers collectively to the transcripts and other elements involved in directing the expression or activity of CRISPR-associated ("Cas") genes, including sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active portion tracrRNA), tracr mate sequences (encompassing "direct repeats" and tracrRNA processing portion direct repeats in the context of endogenous CRISPR systems), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems), or other sequences and transcripts from the CRISPR locus, where the tracrRNA in the system is complementary (fully or partially) to the tracr mate sequence present on the guide RNA.

[0077] DNA synthesis template As used herein, the term "DNA synthesis template" refers to the region or portion of the extension arm of a PEG RNA that is utilized as a template strand by the polymerase of a prime editor, encodes a 3' single-stranded DNA flap containing the desired edit, and then displaces the corresponding endogenous strand of DNA at the target site through the mechanism of prime editing. The extension arm encompassing the DNA synthesis template may be composed of DNA or RNA. In the case of RNA, the polymerase of the prime editor may be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). In the case of DNA, the polymerase of the prime editor may be a DNA-dependent DNA polymerase. In various embodiments, the DNA synthesis template may include all or a portion of the "editing template" and "homology arm," as well as the optional 5'-end modification region e2. That is, depending on the nature of the e2 region (e.g., whether it contains a hairpin, toe-loop, or stem / loop secondary structure), the polymerase may encode none, part, or all of the e2 region. In other words, in the case of a 3' extension arm, the DNA synthesis template may include the portion of the extension arm spanning from the 5' end of the primer binding site (PBS) to the 3' end of the gRNA core, which may serve as a template for synthesis of a single strand of DNA by a polymerase (e.g., reverse transcriptase). In the case of a 5' extension arm, the DNA synthesis template may include the portion of the extension arm spanning from the 5' end of the PEGRNA molecule to the 3' end of the editing template. Preferably, the DNA synthesis template excludes the primer binding site (PBS) of a PEGRNA having either a 3' extension arm or a 5' extension arm. Certain embodiments described herein refer to an "RT template," which includes the editing template and the homology arm (i.e., the sequence of the PEGRNA extension arm that is actually used as a template during DNA synthesis). The term "RT template" is equivalent to the term "DNA synthesis template."

[0078] Editing Templates The term "editing template" refers to the portion of the extension arm that encodes the desired edit in the single-stranded 3' DNA flap synthesized by a polymerase (e.g., a DNA-dependent DNA polymerase, an RNA-dependent DNA polymerase (e.g., a reverse transcriptase)). Certain embodiments described herein refer to an "RT template," which refers to both the editing template and the homology arm together (i.e., the sequence of the PEG RNA extension arm that is actually used as a template during DNA synthesis). The term "RT editing template" is also equivalent to the term "DNA synthesis template," except that RT editing template reflects the use of a prime editor with a polymerase that is a reverse transcriptase, and DNA synthesis template reflects the use of a prime editor with any polymerase more broadly.

[0079] Extension Arm The term "extension arm" refers to a nucleotide sequence component of a PEG RNA that provides several functions, including a primer binding site and an editing template for reverse transcriptase. In some embodiments, the extension arm is located at the 3' end of the guide RNA. In other embodiments, the extension arm is located at the 5' end of the guide RNA. In some embodiments, the extension arm also includes a homology arm. In various embodiments, the extension arm is located in the 5' to 3' direction as follows: homology arm, Editing templates, and Primer binding site Because the polymerization activity of reverse transcriptase is in a 5' to 3' direction, the preferred arrangement of the homology arms, editing template, and primer binding site is in a 5' to 3' direction, so that once primed by the annealed primer sequence, the reverse transcriptase polymerizes a single strand of DNA using the editing template as a complementary template strand. Further details, such as the length of the extension arm, are described elsewhere herein.

[0080] The extension arm may illustratively be described as generally comprising two regions: a primer binding site (PBS) and a DNA synthesis template. The primer binding site binds to a primer sequence formed from the endogenous DNA strand of the target site when it is nicked by the prime editor complex, thereby exposing a 3' end on the endogenous nicked strand. As described herein, binding of the primer sequence to the primer binding site on the extension arm of the PEGRNA creates a double-stranded region with an exposed 3' end (i.e., 3' of the primer sequence), which then provides a substrate for a polymerase to initiate polymerization of a single strand of DNA from the exposed 3' end along the length of the DNA synthesis template. The sequence of the single-stranded DNA product is the complement of the DNA synthesis template. Polymerization continues 5' toward the DNA synthesis template (or extension arm) until polymerization is terminated. Thus, the DNA synthesis template is encoded by the polymerase of the prime editor complex into a single-stranded DNA product (i.e., a 3' single-stranded DNA flap containing the desired gene editing information), representing the portion of the extension arm that will eventually replace the corresponding endogenous DNA strand at the target site located immediately downstream of the PE-induced nick site. Without being bound by theory, polymerization of the DNA synthesis template continues toward the 5' end of the extension arm until a termination event. Polymerization may terminate in a variety of ways, including, but not limited to, the following: (a) reaching the 5′ end of the PEGRNA (for example, in the case of the 5′ extension arm, where the DNA polymerase simply runs out of template); (b) reaching an intraversable RNA secondary structure (e.g., a hairpin or stem / loop); or (c) reaching a replication termination signal (e.g., a specific nucleotide sequence that blocks or inhibits the polymerase) or a nucleic acid topological signal such as supercoiled DNA or RNA.

[0081] fusion proteins As used herein, the term "fusion protein" refers to a hybrid polypeptide containing protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein, or at the carboxy-terminal (C-terminal) portion of the protein, thus forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. The protein may contain different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9, which directs binding of the protein to a target site), and a nucleic acid cleavage domain or catalytic domain of a nucleic acid editing protein. As another example, any of the proteins provided herein, including a fusion of Cas9 or its equivalent with a reverse transcriptase, may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing a peptide linker. Methods for recombinant protein expression and purification are well known and are described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.

[0082] group-specific antigen (gag) Without being limited by theory, and in the context of the life cycle of a typical enveloped virus, Gag is the primary structural protein responsible for directing the majority of steps in virus assembly, including the budding of a fully formed enveloped virion with (i) an envelope (a lipid membrane formed from the plasma membrane during budding, containing one or more glycoproteins inserted therein) and (ii) an internal protein shell, the capsid. Most of these assembly steps occur through interactions with three Gag subdomains: the matrix (MA), capsid (CA), and nucleocapsid (NC; Figure 1). These three regions have a low level of sequence conservation among different retrovirus genera, which contradicts the high level of structural conservation observed. Outside of these three domains, the Gag protein is highly variable. For example, HIV-1 Gag additionally encodes a C-terminal p6 protein and two spacer proteins, SP1 and SP2, that define the CA-NC and NC-p6 junctions, whereas HTLV-1 does not encompass additional sequences outside of MA, CA, and NC ( Oroszlan and Copeland, 1985 ; Henderson et al., 1992 ).

[0083] Gag is also called a "viral structural protein." As used herein, the term "viral structural protein" refers to a viral protein that contributes to the overall structure of the capsid protein or protein core of the virus. The term "viral structural protein" also encompasses functional fragments or derivatives of such viral proteins that contribute to the structure of the capsid protein or protein core of the virus. An example of a viral structural protein is MMLV Gag. Viral membrane fusion proteins are not considered viral structural proteins. Typically, the viral structural proteins are localized within the viral core.

[0084] Group-specific antigen (gag) nucleocapsid protein The term "group-specific antigen nucleocapsid protein" or "gag nucleocapsid protein" refers to a protein that constitutes the core structural component of the inner coat of many viruses. The gag nucleocapsid protein used in the PE-VLPs of the present disclosure may be MMLV gag nucleocapsid protein, FMLV gag nucleocapsid protein, or a nucleocapsid protein from any other virus that produces such a protein.

[0085] Group-specific antigen (gag) protease (pro) polyprotein "Group-specific antigen (gag) protease (pro) polyprotein" or "gag-pro polyprotein" refers to the gag nucleocapsid protein further comprising a viral protease linked thereto. The gag-pro polyprotein mediates the proteolytic cleavage of the gag and gag-pol polyproteins or nucleocapsid proteins during or shortly after virion release from the cell membrane. In the PE-VLPs described herein, the protease of the gag-pro polyprotein is responsible for cleaving the cleavable linker in the fusion protein to release the prime editor after delivery of the PE-VLP to a target cell. In some embodiments, the gag-pro polyprotein is the MMLV gag-pro polyprotein or the FMLV gag-pro polyprotein.

[0086] Guide RNA ("gRNA") As used herein, the term "guide RNA" refers to a specific type of guide nucleic acid that is primarily generally associated with the Cas protein of CRISPR-Cas9, associates with Cas9, and directs the Cas9 protein to a specific sequence within a DNA molecule that contains a complementary sequence to the protospacer sequence of the guide RNA. However, the term also encompasses equivalent guide nucleic acid molecules that cooperate with Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., modified or recombinant), and otherwise program the Cas9 equivalent to localize to a specific target nucleotide sequence. Cas9 equivalents may also include other napDNAbp from any type of CRISPR system (e.g., type II, type V, type VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector," Science 2016; 353(6299), the contents of which are incorporated herein by reference. Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing suitable guide RNA sequences are provided herein. As used herein, "guide RNA" may also be referred to as "existing guide RNA," and is contrasted with a modified form of guide RNA called a "prime editing guide RNA" (or "PEgRNA").

[0087] A guide RNA or PEGRNA may comprise a variety of structural elements, including but not limited to:

[0088] Spacer sequence - a sequence in the guide RNA or PEgRNA (approximately 20 nt in length) that has the same sequence as the protospacer in the target DNA.

[0089] gRNA core (or gRNA scaffold or backbone sequence) - The sequence within the gRNA responsible for Cas9 binding. It does not encompass the 20 bp spacer / target sequence used to guide Cas9 to the target DNA.

[0090] Extension arm - a single-stranded extension at the 3' or 5' end of the PEGRNA, containing a primer binding site and a DNA synthesis template sequence that encodes a single-stranded DNA flap containing the desired genetic alteration via a polymerase (e.g., reverse transcriptase), which then integrates into the endogenous DNA by displacing the corresponding endogenous strand, thereby introducing the desired genetic alteration.

[0091] Transcription terminator - The guide RNA or PEGRNA may contain a transcription termination sequence at the 3' of the molecule.

[0092] Linker As used herein, the term "linker" refers to a molecule that connects two other molecules or moieties. In the case of a linker joining two fusion proteins, the linker can be an amino acid sequence. For example, Cas9 can be fused to a reverse transcriptase enzyme via an amino acid linker sequence. In the case of joining two nucleotide sequences together (e.g., in a gRNA), the linker can also be a nucleotide sequence. For example, in this case, the existing guide RNA is linked to the RNA extension of the prime-edited guide RNA, which may include an RT template sequence and an RT primer binding site, via a spacer or linker nucleotide sequence. In other embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.

[0093] A "cleavable linker" refers to a linker that can be split or cleaved by any means. The linker can be an amino acid sequence. In some embodiments, the linker between the NES and napDNAbp of the PE-VLPs provided herein comprises a cleavable linker. The cleavable linker can also comprise a self-cleaving peptide (e.g., a 2A peptide such as EGRGSLLTCGDVEENPGP (SEQ ID NO: 1), ATNFSLLKQAGDVEENPGP (SEQ ID NO: 2), QCTNYALLKLAGDVESNPGP (SEQ ID NO: 3), or VKQTLNFDLLKLAGDVESNPGP (SEQ ID NO: 4)). In some embodiments, the cleavable linker comprises a protease cleavage site that is cleaved after contact with a protease. For example, the present disclosure contemplates the use of a cleavable linker comprising a protease cleavage site for the amino acid sequence TSTLL MENSS (SEQ ID NO: 5), PRSSLYPALTP (SEQ ID NO: 6), VQALVLTQ (SEQ ID NO: 7), PLQVLTLNIERR (SEQ ID NO: 8), or an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 5-8. In certain embodiments, the cleavable linker comprises an MMLV protease cleavage site for an FMLV protease cleavage site.

[0094] MLH1 The term "MLH1" refers to the gene encoding the DNA mismatch repair enzyme MLH1 (i.e., MutL Homolog 1). The protein encoded by this gene can heterodimerize with the mismatch repair endonuclease PMS2 to form MutL alpha (MutLα), which is part of the DNA mismatch repair system. MLH1 mediates protein-protein interactions during mismatch recognition, strand discrimination, and strand removal. During mismatch repair, the heterodimer MSH2:MSH6 (MutSα) forms and binds to the mismatch. MLH1 then heterodimerizes with PMS2 (MutLα) and binds to the MSH2:MSH6 heterodimer. The MutLα heterodimer then incises the nicked strands 5' and 3' of the mismatch, followed by excision of the mismatch from the MutLα-generated nick by EXO1. Finally, POLδ resynthesizes the excised strand, followed by LIG1 ligation.

[0095] An exemplary amino acid sequence of MLH1 is human isoform 1, P40692-1: >sp|P40692|MLH1_HUMAN DNA mismatch repair protein Mlh1 OS=Homo sapiens OX=9606 GN=MLH1 PE=1 SV=1:

[0096] (SEQ ID NO: 9), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 9.

[0097] Another exemplary amino acid sequence of MLH1 is human isoform 2, P40692-2 (where amino acids 1-241 of isoform 1 are missing): >sp|P40692-2|MLH1_HUMAN Isoform 2 of DNA mismatch repair protein Mlh1 OS=Homo sapiens OX=9606 GN=MLH1:

[0098] MNGYISNANYSVKKCIFLLFINHRLVESTSLRKAIETVYAAYLPKNTHPFLYLSLEISPQNVDVNVHPTKHEVHFLHEESILERVQQHIESKLLGSNSSRMYFTQTLLPGLAGPSGEMVKSTTSLTSSSTSGSSDKVYAHQMVRTDSREQKLDAFLQPLSKPLS SQPQAIVTEDKTDISSGRARQQDEEMLELPAPAEVAAKNQSLEGDTTKGTSEMSEKRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFA NFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVFERC (SEQ ID NO: 10), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 10.

[0099] Another exemplary amino acid sequence of MLH1 is human isoform 3, P40692-3 (amino acids 1-101 (MSFVAGVIRR...ASISTYGFRG (SEQ ID NO: 9) have been replaced with MAF): >sp|P40692-2|MLH1_HUMAN Isoform 2 of the DNA mismatch repair protein Mlh1 OS=Homo sapiens OX=9606 GN=MLH1:

[0100] (SEQ ID NO: 12), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 12.

[0101] In some embodiments, the present disclosure contemplates using the VLPs described herein to deliver inhibitors of MLH1 and / or MMR pathway components that interact with MLH1 (including any wild-type or naturally occurring mutants of MLH1, including any amino acid sequence having at least 70%, or 75%, or 80%, or 85%, or 90%, or 95%, or 99% or more sequence identity to any of SEQ ID NOS: 9-19 or 203-211, or a nucleic acid molecule encoding any MLH1 or MLH1 variant (e.g., a dominant-negative mutant of MLH1 as described herein)) to inhibit, block, or inactivate the function of wild-type MLH1 in the MMR pathway, thereby inhibiting, blocking, or inactivating the MMR pathway, e.g., during genome editing using a prime editor.

[0102] In some embodiments, inactivation of the MMR pathway involves an inhibitor that disrupts, blocks, interferes with, or otherwise inactivates the wild-type function of the MLH1 protein. In some embodiments, inactivation of the MMR pathway involves a mutant MLH1 protein, for example, by delivering the MLH1 mutant protein to a target cell using a VLP as currently described below. In some embodiments, the MLH1 mutant protein interferes with the function of the wild-type MLH1 protein in the MMR pathway, thereby inactivating it. In some embodiments, the MLH1 mutant is a dominant-negative mutant. In some embodiments, the MLH mutant protein is capable of binding to an MLH1-interacting protein, such as MutS.

[0103] Without being bound by theory, the MLH1 dominant-negative mutant functions by saturating the binding of MutS, thereby blocking the binding of MutS to wild-type MLH1 and preventing the function of the wild-type MLH1 protein in the MMR pathway.

[0104] In various embodiments, a dominant-negative MLH1 can include, for example, MLH1 E34A, based on SEQ ID NO: 13 and having the following amino acid sequence (underlined and bold indicates the E34A mutation):

[0105] [ka] (SEQ ID NO: 13), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 13.

[0106] In various other embodiments, a dominant-negative MLH1 can include, for example, MLH1 Δ756, based on SEQ ID NO: 14 and having the following amino acid sequence (underlined and bold indicates the Δ756 mutation at the C-terminus of the sequence):

[0107] [-](SEQ ID NO:14), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO:14 (where [-] indicates the amino acid residue(s) deleted relative to the parent or wild-type sequence).

[0108] In still other embodiments, a dominant-negative MLH1 can include, for example, MLH1 Δ754-Δ756, based on SEQ ID NO: 15 and having the following amino acid sequence (underlined and bold indicates the Δ754-Δ756 mutation at the C-terminus of the sequence):

[0109] [- - -] (SEQ ID NO: 15), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 15 (where [- - -] indicates the amino acid residue(s) deleted relative to the parent or wild-type sequence).

[0110] In another alternative embodiment, the dominant-negative MLH1 can include, for example, MLH1 E34A Δ754-Δ756 based on SEQ ID NO: 16 and having the following amino acid sequence (underlined and bold indicates the E34A and Δ754-Δ756 mutations):

[0111] [ka] (SEQ ID NO: 16), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 16.

[0112] In some embodiments, a dominant-negative MLH1 can include, for example, MLH1 1-335, which is based on SEQ ID NO: 17 and has the following amino acid sequence (containing amino acids 1-335 of SEQ ID NO: 9):

[0113] MSFVAGVIRRLDETVVNRIAAGEVIQRPANAIKEMIENCLDAKSTSIQVIVKEGGLKLIQIQDNGTGIRKEDLDIVCERFTTSKLQSFEDLASISTYGFRGEALASISHVAHVTITTK TADGKCAYRASYSDGKLKAPPKPCAGNQGTQITVEDLFYNIATRRKALKNPSEEYGKILEVVGRYSVHNAGISFSVKKQGETVADVRTLPNASTVDNIRSIFGNAVSRELIEIGCEDK TLAFKMNGYISNANYSVKKCIFLLFINHRLVESTSLRKAIETVYAAYLPKNTHPFLYLSLEISPQNVDVNVHPTKHEVHFLHEESILERVQQHIESKLL (SEQ ID NO: 17), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 17.

[0114] In other embodiments, a dominant-negative MLH1 can include, for example, MLH1 1-335 E34A, based on SEQ ID NO: 18 and having the following amino acid sequence (containing amino acids 1-335 of SEQ ID NO: 9 and the E34A mutation relative to SEQ ID NO: 204):

[0115] [ka] (SEQ ID NO: 18), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 18.

[0116] In yet other embodiments, the dominant-negative MLH1 is, for example, based on SEQ ID NO: 9 and has the following amino acid sequence (containing amino acids 1-335 of SEQ ID NO: 9 and the SV40 NLS sequence): SV40 (or MLH1dn NTD ) can include:

[0117] [ka] (SEQ ID NO: 19) (the underlined and bolded portion refers to the NLS of SV40), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 19.

[0118] In yet other embodiments, the dominant-negative MLH1 is, for example, an MLH1 NLS alternate(based on SEQ ID NO: 9 and having the following amino acid sequence (containing amino acids 1-335 of SEQ ID NO: 9 and an alternative NLS sequence):

[0119] MSFVAGVIRRLDETVVNRIAAGEVIQRPANAIKEMIENCLDAKSTSIQVIVKEGGLKLIQIQDNGTGIRKEDLDIVCERFTTSKLQSFEDLASISTYGFRGEALASISHVAHVTITTKTADGKCAYRASYSDGKLKAPPKPCAGNQGTQITVEDLFYNIATRRKALKNPSEE YGKILEVVGRYSVHNAGISFSVKKQGETVADVRTLPNASTVDNIRSIFGNAVSRELIEIGCEDKTLAFKMNGYISNANYSVKCCIFLLFINHRLVESTSLRKAIETVYAAYLPKNTHPFLYLSLEISPQNVDVNVHPTKHEVHFLHEESILERVQQHIESKLL-[Alternative NLS Array] (SEQ ID NO: 17) - [alternative NLS sequence], or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 17. The alternative NLS sequence can be any suitable NLS sequence, including but not limited to: [Table 1]

[0120] In yet other embodiments, the dominant negative MLH1 is, for example, amino acids 501-756 of SEQ ID NO:9:

[0121] INLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVFERC (SEQ ID NO: 206), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 206; The C-terminal fragment of SEQ ID NO: 9 may include MLH1 501-756, which corresponds to the C-terminal fragment of SEQ ID NO: 9.

[0122] In still other embodiments, the dominant negative MLH1 is, for example, amino acids 501-753 of SEQ ID NO: 9: INLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVF[- - -] (SEQ ID NO: 207), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 207; The C-terminal fragment of SEQ ID NO: 9 may include MLH1 501-753, which corresponds to the C-terminal fragment of SEQ ID NO: 9.

[0123] In still other embodiments, the dominant negative MLH1 is, for example, amino acids 461-756 of SEQ ID NO: 9:KRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGL PIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVFERC (SEQ ID NO: 208), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 208; The C-terminal fragment of SEQ ID NO: 9, MLH1 461-756, may be included.

[0124] In various embodiments, the dominant negative MLH1 comprises, for example, amino acids 461-753 of SEQ ID NO:9:

[0125] KRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVF[- - -] (SEQ ID NO: 209), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 209; The C-terminal fragment of SEQ ID NO: 9, MLH1 461-753, may be included.

[0126] In various other embodiments, the dominant negative MLH1 is a C-terminal fragment of SEQ ID NO: 9, e.g., corresponding to amino acids 461 to 753 of SEQ ID NO: 9, as well as an N-terminal NLS, e.g., NLS SV40 : [NLS]-KRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPK EGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVF[- - -] (SEQ ID NO: 209), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to SEQ ID NO: 209; The NLS sequence may include MLH1 461-753, which further comprises: The NLS sequence may be any suitable NLS sequence, including, but not limited to, SEQ ID NOs: 20-31 and 77-81.

[0127] napDNAbp As used herein, the term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to a protein, of which Cas9 is an example, that uses RNA:DNA hybridization to target and bind to specific sequences within a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., a guide RNA), which localizes the napDNAbp to a DNA sequence containing a DNA strand (i.e., a target strand) complementary to the guide nucleic acid or a portion thereof (e.g., a protospacer of the guide RNA). In other words, the guide nucleic acid "programs" the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to the complementary sequence.

[0128] Without being bound by theory, the binding mechanism of the napDNAbp-guide RNA complex generally involves the formation of an R-loop, which allows the napDNAbp to induce unwinding of the double-stranded DNA target, thereby separating the strands in the region bound by the napDNAbp. The guide RNA protospacer then hybridizes to the "target strand." This displaces the "non-target strand" that is complementary to the target strand, forming a single-stranded region of the R-loop. In some embodiments, the napDNAbp contains one or more nuclease activities that then cleave the DNA, leaving various types of lesions. For example, the napDNAbp may contain nuclease activities that cleave the non-target strand at a first location and / or the target strand at a second location. Depending on the nuclease activity, the target DNA may be cleaved, forming a "double-stranded break" in which both strands are cleaved. In other embodiments, the target DNA may be cleaved at only a single site, i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with different nuclease activities include "Cas9 nickase" ("nCas9") and inactivated Cas9 with no nuclease activity ("dead Cas9" or "dCas9"). Exemplary sequences of these and other napDNAbps are provided herein.

[0129] Nickase As used herein, "nickase" refers to a napDNAbp (e.g., a Cas protein) that can cleave only one of the two complementary strands of a double-stranded target DNA sequence, thereby generating a nick in that strand. In some embodiments, the nickase cleaves the non-target strand of the double-stranded target DNA sequence. In some embodiments, the nickase comprises an amino acid sequence with one or more mutations in the catalytic domain of a standard napDNAbp (e.g., a Cas protein), wherein the one or more mutations reduce or eliminate the nuclease activity of the catalytic domain. In some embodiments, the nickase is a Cas9 that comprises one or more mutations in the RuvC-like domain relative to the wild-type Cas9 sequence or relative to the equivalent amino acid position in another Cas9 variant or Cas9 equivalent. In some embodiments, the nickase is a Cas9 that comprises one or more mutations in the HNH-like domain relative to the wild-type Cas9 sequence or relative to the equivalent amino acid position in another Cas9 variant or Cas9 equivalent. In some embodiments, the nickase is a Cas9 that contains an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 relative to the equivalent amino acid position in the standard Cas9 sequence or other Cas9 variants or equivalents. In some embodiments, the nickase is a Cas9 that contains an H840A, N854A, and / or N863A mutation relative to the equivalent amino acid position in the standard Cas9 sequence or other Cas9 variants or equivalents. In some embodiments, the term "Cas9 nickase" refers to a Cas9 in which one of the two nuclease domains is inactivated. This enzyme is capable of cleaving only one strand of target DNA. In some embodiments, the nickase is a Cas protein that is not a Cas9 nickase.

[0130] Nuclear export sequence (NES) The term "nuclear export sequence" or "NES" refers to an amino acid sequence that promotes the transport of proteins from the cell nucleus to the cytoplasm, for example, through the nuclear pore complex by nuclear transport. Nuclear export sequences are known in the art and will be apparent to those skilled in the art. For example, NES sequences are described in Xu, D. et al. Sequence and structural analyses of nuclear export signals in the NESdb database. Mol. Biol. Cell. 2012, 23(18) 3677-3693, the contents of which are incorporated herein by reference.

[0131] Nuclear localization sequence (NLS) The term "nuclear localization sequence" or "NLS" refers to an amino acid sequence that promotes the import of proteins into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and will be clear to those skilled in the art. For example, NLS sequences are described in International PCT application PCT / EP2000 / 011690 (Plank et al.), filed November 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001, the contents of which are incorporated herein by reference as disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 30).

[0132] nucleic acid The term "nucleic acid," as used herein, refers to a polymer of nucleotides. The polymer may be composed of natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, C5 bromouridine, C5 fluorouridine, C5 iodouridine, C5 propynyluridine, C5 propynylcytidine, C5 methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-2, 2-diaminoadenosine, 2-diaminothymidine, 2-diamino-2-methyl-1 ... )methylguanine, 4-acetylcytidine, 5-(carboxyhydroxymethyl)uridine, dihydrouridine, methylpseudouridine, 1-methyladenosine, 1-methylguanosine, N6-methyladenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalating bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-O-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioates and 5'N phosphoramidite conjugates).

[0133] PEG-RNA As used herein, the term "prime editing guide RNA" or "PEgRNA" or "extended guide RNA" refers to a specialized form of guide RNA that has been modified to include one or more additional sequences for carrying out the prime editing methods and compositions described herein. As described herein, a prime editing guide RNA includes one or more "extension regions" of nucleic acid sequence. The extension region may comprise, but is not limited to, single-stranded RNA or DNA. Furthermore, the extension region may occur at the 3' end of an existing guide RNA. In other configurations, the extension region may occur at the 5' end of a conventional guide RNA. In still other configurations, the extension region may occur in an intramolecular region of an existing guide RNA, for example, a gRNA core region that associates with and / or binds to napDNAbp. The extension region comprises a "DNA synthesis template" encoding a single-stranded DNA (by the polymerase of the prime editor), which is then designed to be (a) homologous to the endogenous target DNA to be edited and (b) contain at least one desired nucleotide change (e.g., a transition, a translocation, a deletion, an insertion) to be introduced or integrated into the endogenous target DNA. The extension region may also comprise other functional sequence elements, such as, but not limited to, a "primer binding site" and a "spacer or linker" sequence, or other structural elements, such as, but not limited to, an aptamer, a stem-loop, a hairpin, a toe-loop (e.g., a 3' toe-loop), or an RNA-protein recruitment domain (e.g., an MS2 hairpin). As used herein, a "primer binding site" comprises a sequence that hybridizes to a single-stranded DNA sequence having a 3' end generated from the nicked DNA of the R-loop.

[0134] In some embodiments, the PEG RNA comprises a 5' extension arm, a spacer, and a gRNA core. The 5' extension further comprises, in a 5' to 3' direction, a reverse transcriptase template, a primer binding site, and a linker. The reverse transcriptase template may also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editor described herein is not an RT but a different type of polymerase.

[0135] In some other embodiments, the PEG RNA comprises a 5' extension arm, a spacer, and a gRNA core. The 5' extension further comprises, in a 5' to 3' direction, a reverse transcriptase template, a primer binding site, and a linker. The reverse transcriptase template may also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editor described herein is not an RT but a different type of polymerase.

[0136] In yet another embodiment, the PEgRNA comprises, from 5' to 3', a spacer (1), a gRNA core (2), and an extension arm (3). The extension arm (3) is at the 3' end of the PEgRNA. The extension arm (3) further comprises, from 5' to 3', a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) may also comprise optional modification regions at the 3' and 5' ends, which may be the same or different sequences. In addition, the 3' end of the PEgRNA may comprise a transcription terminator sequence. These sequence elements of the PEgRNA are further described and defined herein.

[0137] In yet another embodiment, the PEgRNA has, from 5' to 3', an extension arm (3), a spacer (1), and a gRNA core (2). The extension arm (3) is at the 5' end of the PEgRNA. The extension arm (3) further comprises, from 3' to 5', a "primer binding site" (A), an "editing template" (B), and a "homology arm" (C). The extension arm (3) may also comprise optional modification regions at the 3' and 5' ends, which may be the same or different sequences. The PEgRNA may also comprise a transcription terminator sequence at the 3' end. These sequence elements of the PEgRNA are further described and defined herein.

[0138] PE1 As used herein, "PE1" refers to the following: [NLS]-[Cas9(H840A)]-[Linker]-[MMLV_RT(wt)] + desired PEG RNA wherein the PE fusion has the amino acid sequence of SEQ ID NO: 32, as shown below; [ka]

[0139] PE2 As used herein, "PE2" refers to the following: [NLS]-[Cas9(H840A)]-[Linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)]+desired PEG RNA wherein the PE fusion has the amino acid sequence of SEQ ID NO: 33, as shown below: [ka]

[0140] PE3 As used herein, "PE3" refers to PE2 plus a second strand nicking guide RNA that complexes with PE2 and introduces a nick in the non-edited DNA strand to induce preferential displacement of the edited strand.

[0141] PE3b As used herein, "PE3b" refers to PE3, where the second-strand nicking guide RNA is designed to temporally control the second-strand nick so that nicking is not introduced until after the desired edit has been introduced. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand, not the original allele. Using this strategy, hereafter referred to as PE3b, the mismatch between the protospacer and the unedited allele should prevent nicking by the sgRNA until the PAM strand editing event has occurred.

[0142] PE4 As used herein, "PE4" refers to a system that includes, in addition to PE2, an MLH1 dominant-negative protein (i.e., wild-type MLH1 truncated at amino acids 754-756, sometimes referred to herein as "MLH1 Δ754-756" or "MLH1dn") expressed in trans. In some embodiments, PE4 refers to a fusion protein that includes PE2 and the MLH1 dominant-negative protein linked via an optional linker.

[0143] PE5 As used herein, "PE5" refers to a system that includes, in addition to PE3, an MLH1 dominant-negative protein (i.e., wild-type MLH1 truncated at amino acids 754-756, sometimes referred to as "MLH1 Δ754-756" or "MLH1dn") expressed in trans. In some embodiments, PE5 refers to a fusion protein that includes PE3 and the MLH1 dominant-negative protein linked via an optional linker.

[0144] PEmax As used herein, "PEmax" refers to a polypeptide containing Cas9(R221K N39K H840A) and the following: [Bipartite NLS]-[Cas9(R221K)(N394K)(H840A)]-[Linker]-[MMLV_RT(D200N)(T330P)(L603W)]-[Bipartite NLS]-[NLS]+Desired PEG-RNA wherein the PE fusion has the amino acid sequence of SEQ ID NO: 34, as shown below: [ka]

[0145] PE4max As used herein, "PE4max" refers to PE4, but where the PE2 component has been replaced with PEmax.

[0146] PE5max As used herein, "PE5max" refers to PE5, but where the PE2 component of PE3 has been replaced with PEmax.

[0147] polymerase As used herein, the term "polymerase" refers to an enzyme that synthesizes a nucleotide strand, and may be used in connection with the prime editor delivery systems described herein. A polymerase may be a "template-dependent" polymerase (i.e., a polymerase that synthesizes a nucleotide strand based on the order of nucleotide bases in a template strand). A polymerase may also be a "template-independent" polymerase (i.e., a polymerase that synthesizes a nucleotide strand without requiring a template strand). A polymerase may also be further classified as a "DNA polymerase" or an "RNA polymerase." In various embodiments, a prime editor system comprises a DNA polymerase. In various embodiments, the DNA polymerase may be a "DNA-dependent DNA polymerase" (i.e., the template molecule is a DNA strand). In such cases, the DNA template molecule may be a PEG-RNA, where the extension arm comprises a strand of DNA. In such cases, the PEGRNA may be referred to as a chimeric or hybrid PEGRNA comprising an RNA portion (i.e., a guide RNA component including a spacer and a gRNA core) and a DNA portion (i.e., an extension arm). In various other embodiments, the DNA polymerase may be an "RNA-dependent DNA polymerase" (i.e., the template molecule is a strand of RNA). In such cases, the PEGRNA may be RNA, i.e., an RNA comprising an RNA extension. The term "polymerase" may also refer to an enzyme that catalyzes the polymerization of nucleotides (i.e., polymerase activity). Generally, the enzyme initiates synthesis at the 3' end of a primer annealed to a polynucleotide template sequence (e.g., a primer sequence annealed to the primer binding site of the PEGRNA) and proceeds toward the 5' end of the template strand. A "DNA polymerase" catalyzes the polymerization of deoxynucleotides. When used herein with respect to a DNA polymerase, the term DNA polymerase encompasses "functional fragments thereof."A "functional fragment thereof" refers to any portion of a wild-type or mutant DNA polymerase that encompasses less than the entire amino acid sequence of the polymerase and that retains the ability to catalyze the polymerization of polynucleotides under at least one set of conditions. Such functional fragments may exist as separate entities or may be components of a larger polypeptide, such as a fusion protein.

[0148] Prime Edit As used herein, the term "prime editing" refers to an approach for gene editing using napDNAbps, a polymerase (e.g., reverse transcriptase), and a specialized guide RNA that contains a DNA synthesis template that encodes the desired new genetic information (or deletes genetic information) and is then incorporated into the target DNA sequence. Prime editing is described in Anzalone, AV et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019), the entire contents of which are incorporated herein by reference.

[0149] Prime editing represents a versatile and precise genome editing platform that uses a nucleic acid-programmable DNA-binding protein ("napDNAbp") working in conjunction with a polymerase (i.e., provided in the form of a fusion protein or in trans with the napDNAbp) to directly write new genetic information into a designated DNA site. The prime editing system is programmed with a prime editing (PE) guide RNA ("PEgRNA") to identify the target site and template the synthesis of the desired edit in the form of a replacement DNA strand by an engineered extension (either DNA or RNA) onto the guide RNA (e.g., at the 5' or 3' end or in an internal portion of the guide RNA). The replacement strand, containing the desired edit (e.g., a single nucleobase substitution), shares the same sequence as (or is homologous to) the endogenous strand immediately downstream of the nick site at the target site to be edited (except that it encompasses the desired edit). Through DNA repair and / or replication mechanisms, the endogenous strand downstream of the nick site is replaced with the newly synthesized replacement strand containing the desired edit. In some cases, prime editing may be thought of as a "search and replace" genome editing technique, as the prime editor not only searches for and finds the desired target site to be edited, as described herein, but also simultaneously encodes a replacement strand containing the desired edit that is introduced in place of the endogenous DNA strand at the corresponding target site. The prime editor disclosed herein relates, in part, to the discovery that the mechanism of target prime reverse transcription (TPRT) or "prime editing" can be exploited or adapted to perform precise CRISPR / Cas-based genome editing with high efficiency and genetic flexibility. TPRT is naturally used by mobile DNA elements such as mammalian non-LTR retrotransposons and bacterial group II introns. Cas protein-reverse transcriptase fusions or related systems are used to target specific DNA sequences with a guide RNA, generate a single-stranded nick at the target site, and use the nicked DNA as a primer for reverse transcription of a modified reverse transcriptase template integrated with the guide RNA.However, while this concept begins with a prime editor that uses reverse transcriptase as the DNA polymerase component, the prime editors described herein are not limited to reverse transcriptase and may encompass the use of virtually any DNA polymerase. Indeed, throughout this application, while reference is made to a prime editor having a "reverse transcriptase," it is noted that reverse transcriptase is merely one type of DNA polymerase that may work with a prime editor. Thus, when this specification refers to a "reverse transcriptase," those skilled in the art should understand that any suitable DNA polymerase may be used in place of the reverse transcriptase. Thus, in one aspect, a prime editor may comprise a Cas9 (or equivalently, napDNAbp) programmed to target a DNA sequence by associating with a specialized guide RNA (i.e., a PEgRNA) that contains a spacer sequence that anneals to a complementary protospacer in the target DNA. The specialized guide RNA also contains new genetic information in the form of an extension that encodes a replacement strand of DNA containing the desired genetic alteration, which is used to replace the corresponding endogenous DNA strand at the target site. To transfer information from PEGRNA to target DNA, the prime editing mechanism involves nicking one strand of DNA at the target site, exposing a 3'-hydroxyl group. The exposed 3'-hydroxyl group can then be used to prime DNA polymerization of an extension encoding an edit on the PEGRNA directly to the target site. In various embodiments, the extension (which provides a template for polymerization of the replacement strand containing the edit) can be formed from RNA or DNA. In the case of RNA extension, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (such as a reverse transcriptase). In the case of DNA extension, the polymerase of the prime editor can be a DNA-dependent DNA polymerase.The newly synthesized strand formed by the prime editor (i.e., the replacement DNA strand containing the desired edit) will be homologous to (i.e., have the same sequence as) the genomic target sequence, except for the desired nucleotide change (e.g., a single nucleotide change, a deletion, or an insertion, or a combination thereof). The newly synthesized (or replaced) strand of DNA, sometimes referred to as a single-stranded DNA flap, will compete for hybridization with the complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. In certain embodiments, the system can be combined with the use of an error-prone reverse transcriptase (e.g., provided as a fusion protein with a Cas9 domain or provided in trans to a Cas9 domain). The error-prone reverse transcriptase can introduce changes during the synthesis of the single-stranded DNA flap. Thus, in certain embodiments, an error-prone reverse transcriptase can be utilized to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used in the system, the changes can be random or non-random. Degradation of the hybridized intermediate (containing the single-stranded DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand) can involve removal of the resulting displaced flap from the endogenous DNA (e.g., using the 5'-end DNA flap endonuclease, FEN1), ligation of the synthesized single-stranded DNA flap to the target DNA, and assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because templated DNA synthesis offers single-nucleotide precision for any nucleotide modification, including insertions and deletions, the applicability of this approach is extremely broad, and it is foreseeable that it could be used for countless applications in basic science and therapeutics.

[0150] In various embodiments, prime editing operates by contacting a target DNA molecule (where a change in the nucleotide sequence is desired to be introduced) with a nucleic acid programmable DNA-binding protein (napDNAbp) complexed with a prime editing guide RNA (PEgRNA). In various embodiments, the prime editing guide RNA (PEgRNA) contains an extension at the 3' or 5' end of the guide RNA, or at an intramolecular location within the guide RNA, and encodes the desired nucleotide change (e.g., a single base change, insertion, or deletion). In step (a), the napDNAbp / extended gRNA complex contacts the DNA molecule, and the extended gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) into one of the strands of DNA at the target locus, thereby creating an available 3' end on one of the strands of the target locus. In some embodiments, a nick is created in the strand of DNA corresponding to the R-loop strand, i.e., the strand that does not hybridize to the guide RNA sequence, i.e., the "non-target strand." However, a nick can be introduced into either strand. That is, a nick can be introduced into the "target strand" of the R-loop (i.e., the strand hybridized to the protospacer of the extended gRNA) or the "non-target strand" (i.e., the strand that forms the single-stranded portion of the R-loop and is complementary to the target strand). In step (c), the 3' end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA to prime reverse transcription (i.e., "target primed RT"). In some embodiments, the 3'-terminal DNA strand hybridizes to a specific RT priming sequence on the extended portion of the guide RNA, i.e., a "reverse transcriptase priming sequence" or "primer binding site" on the PEGRNA. In step (d), a reverse transcriptase (or other suitable DNA polymerase) is introduced to synthesize a single strand of DNA from the 3' end of the primed site toward the 5' end of the primed edited guide RNA. A DNA polymerase (eg, reverse transcriptase) may be fused to the napDNAbp or alternatively provided in trans to the napDNAbp.This forms a single-stranded DNA flap that contains the desired nucleotide change (e.g., a single-base change, an insertion, a deletion, or a combination thereof) and is otherwise homologous to the endogenous DNA at or adjacent to the nick site. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) involve degradation of the single-stranded DNA flap so that the desired nucleotide change becomes incorporated into the target locus. This process can be directed toward desired product formation by removing the corresponding 5' endogenous DNA flap that forms once the 3' single-stranded DNA flap invades and hybridizes to the endogenous DNA sequence. Without being bound by theory, the cell's endogenous DNA repair and replication processes degrade the mismatched DNA and incorporate the nucleotide change(s) to form the desired altered product. This process can be directed toward product formation using "second-strand nicking." This process may introduce at least one or more of the following genetic changes: transitions, deletions, and insertions.

[0151] The terms "prime editor (PE) system" or "prime editor (PE)" or "PE system" or "PE editing system" refer to compositions involved in the genome editing methods using target primed reverse transcription (TPRT) described herein, and include, but are not limited to, napDNAbps, reverse transcriptase, a fusion protein (e.g., comprising napDNAbps and reverse transcriptase), a prime editing guide RNA, and a complex comprising the fusion protein and the prime editing guide RNA, as well as accessory elements such as a second strand nicking component (e.g., a second strand sgRNA) and a 5' endogenous DNA flap removal endonuclease (e.g., FEN1) to help guide the prime editing process toward edited product formation.

[0152] In the embodiments described thus far, the PEgRNA constitutes a single molecule comprising a guide RNA (which itself comprises a spacer sequence and a gRNA core or scaffold) and a 5' or 3' extension arm comprising a primer binding site and a DNA synthesis template; however, the PEgRNA can also take the form of two individual molecules comprised of a guide RNA and a trans-prime editor RNA template (tPERT), where the PEgRNA essentially harbors the extension arm (which encompasses, among other things, the primer binding site and the DNA synthesis domain), and an RNA-protein recruitment domain (e.g., an MS2 aptamer or hairpin) in the same molecule, and is co-localized or recruited to a modified prime editor complex comprising a tPERT recruitment protein (e.g., an MS2cp protein that binds to the MS2 aptamer).

[0153] Prime Editor The term "prime editor" refers to a fusion construct comprising a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase, which is capable of performing prime editing on a target nucleotide sequence in the presence of a PEgRNA (or "extended guide RNA"). The term "prime editor" can also refer to a fusion protein, or a fusion protein complexed with a PEgRNA, and / or a fusion protein further complexed with a second-strand nicking sgRNA. In some embodiments, a prime editor can also refer to a complex comprising a fusion protein (reverse transcriptase fused to a napDNAbp), a PEgRNA, and a normal guide RNA capable of directing the second-site nicking step of the non-edited strand as described herein.

[0154] Primer binding site The term "primer binding site" or "PBS" refers to a nucleotide sequence located on the PEgRNA as a component of the extension arm (typically at the 3' end of the extension arm) that serves to bind to a primer sequence formed after Cas9 nicking of the target sequence by the prime editor. As detailed elsewhere, when the Cas9 nickase component of the prime editor nicks one strand of the target DNA sequence, a 3'-terminal ssDNA flap is formed that serves as a primer sequence that anneals to the primer binding site on the PEgRNA to prime reverse transcription.

[0155] Protease cleavage site As used herein, the term "protease cleavage site" refers to an amino acid sequence that is recognized and cleaved by a protease, i.e., an enzyme that catalyzes proteolysis, breaking down proteins into smaller polypeptides or single amino acids. In some embodiments, the protease cleavage site is contained within a cleavable linker within a fusion protein as described herein. In some embodiments, the protease cleavage site is cleaved by a protease of the gag-pro polyprotein. In some embodiments, the protease cleavage site comprises an MMLV protease cleavage site or an FMLV protease cleavage site. In some embodiments, the protease cleavage site comprises one of the amino acid sequences TSTLLMEANSS (SEQ ID NO: 5), PRSSLYPALTP (SEQ ID NO: 6), VQALVLTQ (SEQ ID NO: 7), PLQVLTLNIERR (SEQ ID NO: 8), or an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 5-8.

[0156] Proteins, peptides, and polypeptides The terms "protein," "peptide," and "polypeptide" are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids in length. A protein, peptide, or polypeptide may refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified by the addition of a chemical entity, such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, or the like. A protein, peptide, or polypeptide may also be a single molecule or a complex of multiple molecules. A protein, peptide, or polypeptide may be simply a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced through recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the contents of which are incorporated herein by reference.

[0157] Protospacer As used herein, the term "protospacer" refers to a sequence (~20 bp) in DNA adjacent to a PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, to one of its strands, i.e., the "target strand" versus the "non-target strand" of the target DNA sequence). Cas9 function also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease is derived from Streptococcus pyogenes and recognizes the NGG PAM sequence found on the non-target strand, immediately downstream of the target sequence in genomic DNA. Those skilled in the art will understand that state-of-the-art literature sometimes refers to the ~20 nt target-specific guide sequence of the guide RNA itself rather than to the "spacer." Thus, in some cases, as used herein, the term "protospacer" may be used interchangeably with the term "spacer." The descriptive context surrounding the appearance of "protospacer" or "spacer" will help inform the reader whether the term refers to a gRNA or a DNA target.

[0158] Protospacer adjacent motif (PAM) As used herein, the term "protospacer adjacent sequence" or "PAM" refers to a DNA sequence of approximately 2-6 base pairs that is the key target component of the Cas9 nuclease. Typically, the PAM sequence is located on either strand, downstream from the Cas9 cleavage site in the 5' to 3' direction. The standard PAM sequence (i.e., the PAM sequence associated with the Streptococcus pyogenes Cas9 nuclease, or SpCas9) is 5'-NGG-3', where "N" is any nucleobase, followed by two guanine ("G") nucleobases. Different PAM sequences may be associated with different Cas9 nucleases or equivalent proteins from different organisms. Additionally, any given Cas9 nuclease, e.g., SpCas9, may be modified to alter the nuclease's PAM specificity so that the nuclease recognizes alternative PAM sequences.

[0159] For example, referring to the standard SpCas9 amino acid sequence SEQ ID NO: 37, the PAM sequence can be modified by introducing one or more mutations, including (a) the "VQR variant" of D1135V, R1335Q, and T1337R (which alters PAM specificity to NGAN or NGNG), (b) the "EQR variant" of D1135E, R1335Q, and T1337R (which alters PAM specificity to NGAG), and (c) the "VRER variant" of D1135V, G1218R, R1335E, and T1337R (which alters PAM specificity to NGCG). In addition, the D1135E variant of standard SpCas9 still recognizes NGG, but is more selective than the wild-type SpCas9 protein.

[0160] It will also be understood that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) can have different PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. Additionally, Cas9 from Neisseria meningitidis (NmCas) recognizes NNNNGATT. In another example, Cas9 from Streptococcus thermophilis (StCas9) recognizes NNAGAAW. In yet another example, Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These are examples and are not meant to be limiting. It will further be understood that non-SpCas9s bind to a variety of PAM sequences, making them useful when a suitable SpCas9 PAM sequence is not present at the desired target cleavage site. Furthermore, non-SpCas9s may have other characteristics that make them more useful than SpCas9s. For example, Cas9 from Staphylococcus aureus (SaCas9) is approximately 1 kilobase smaller than SpCas9 and can be packaged into adeno-associated viruses (AAVs). See also Shah et al., "Protospacer recognition motifs: mixed identities and functional diversity," RNA Biology, 10(5): 891-899, which is incorporated herein by reference.

[0161] reverse transcriptase The term "reverse transcriptase" refers to a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize DNA transcripts from an RNA template. Historically, reverse transcriptases have primarily been used to transcribe mRNA into cDNA, which is then cloned into vectors for further manipulation. Avian myoblast virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). This enzyme possesses 5'→3' RNA-directed DNA polymerase activity, 5'→3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a 5' and 3' ribonuclease specific for the RNA strand of an RNA-DNA hybrid (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Known viral reverse transcriptases lack the 3' to 5' exonuclease activity required for proofreading, so errors during transcription cannot be corrected by the reverse transcriptase (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNase H activity is presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase widely used in molecular biology is the reverse transcriptase derived from Moloney murine leukemia virus (M-MLV or "MMLV"). See, e.g., Gerard, GR, DNA 5:271-279 (1986), and Kotewicz, ML, et al., Gene 35:249-258 (1985). M-MLV reverse transcriptases that are substantially devoid of RNase H activity have also been described. See, e.g., U.S. Patent No. 5,244,797. The present invention contemplates the use of any such reverse transcriptase, or variants or mutants thereof.

[0162] In addition, the present invention contemplates the use of error-prone reverse transcriptases, sometimes referred to as error-prone reverse transcriptases or reverse transcriptases that do not support precise incorporation of nucleotides during polymerization. During synthesis of a single-stranded DNA flap based on an RT template integrated with a guide RNA, the error-prone reverse transcriptase may introduce one or more nucleotides that mismatch with the RT template sequence, thereby introducing them into the nucleotide sequence through erroneous polymerization of the single-stranded DNA flap. These errors introduced during synthesis of the single-stranded DNA flap are then integrated into the double-stranded molecule through hybridization to the corresponding endogenous target strand, removal of the endogenous displaced strand, ligation, and another round of endogenous DNA repair and / or sequencing processes. In some embodiments, the present disclosure provides a prime editor fusion protein comprising an MMLV RT.

[0163] Reverse transcription As used herein, the term "reverse transcription" refers to the ability of an enzyme to synthesize a DNA strand (i.e., complementary DNA or cDNA) using RNA as a template. In some embodiments, the reverse transcription can be "error-prone reverse transcription," which refers to the property of certain reverse transcriptase enzymes to be prone to errors in their DNA polymerization activity.

[0164] Spacer sequence As used herein, the term "spacer sequence" in reference to a guide RNA or PEgRNA refers to a portion of the guide RNA or PEgRNA of approximately 20 nucleotides that contains a nucleotide sequence that shares the same sequence as a protospacer sequence in the target DNA sequence. The spacer sequence anneals to the complement of the protospacer sequence, forming a ssRNA / ssDNA hybrid structure at the target site and a corresponding R-loop ssDNA structure on the endogenous DNA strand.

[0165] subject As used herein, the term "subject" refers to an individual organism, e.g., an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate animal. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject may be male or female and at any stage of development. target site The term "target site" refers to a sequence within a nucleic acid molecule that is edited by a prime editor (PE) disclosed herein. Target site also refers to a sequence within a nucleic acid molecule to which a complex of a prime editor (PE) and a gRNA binds.

[0166] treatment The terms "treatment," "treat," and "treating," as described herein, refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder, or one or more symptoms thereof. As used herein, the terms "treatment," "treat," and "treating," as described herein, refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder, or one or more symptoms thereof. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms (e.g., to prevent or delay the onset of symptoms or to inhibit the onset or progression of a disease). For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay their recurrence.

[0167] variant As used herein, the term "variant" should be interpreted to mean exhibiting a pattern of qualities that deviates from those occurring in nature; for example, a variant Cas9 is a Cas9 that contains one or more amino acid residue changes compared to the wild-type Cas9 amino acid sequence. The term "variant" encompasses homologous proteins that have at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity to a reference sequence and have the same or substantially the same functional activity(ies) as the reference sequence. The term also encompasses mutants, truncations, or domains of a reference sequence that exhibit the same or substantially the same functional activity(ies) as the reference sequence.

[0168] vector As used herein, the term "vector" refers to a nucleic acid that can be modified to encode a gene of interest, enter a host cell, mutate, and replicate within the host cell, and then transfer the replicative form of the vector into another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phages, and conjugative plasmids. Additional suitable vectors will be apparent to those skilled in the art based on the present disclosure.

[0169] viral envelope glycoproteins The term "viral envelope glycoprotein" refers to an oligosaccharide-containing protein that forms part of the viral envelope (i.e., the outermost layer of many types of viruses that protects viral genetic material as it travels between host cells). The glycoprotein may also assist in identifying and binding to a receptor on the target cell membrane so that the viral envelope fuses with the membrane, allowing the contents of the viral particle (which may, for example, include the PE-VLPs described herein) to enter the host cell. The viral envelope glycoprotein used in the PE-VLPs of the present disclosure may include any glycoprotein from an enveloped virus. In some embodiments, the viral envelope glycoprotein is an adenovirus envelope glycoprotein, an adeno-associated virus envelope glycoprotein, a retrovirus envelope glycoprotein, or a lentivirus envelope glycoprotein. In some embodiments, the viral envelope glycoprotein is vesicular stomatitis virus G protein (VSV-G), baboon retrovirus envelope glycoprotein (BaEVRless), FuG-B2 envelope glycoprotein, HIV-1 envelope glycoprotein, or ecotropic murine leukemia virus (MLV) envelope glycoprotein.

[0170] Virus-like particles (VLPs) As used herein, a virus-like particle is a (a) (i) a lipid membrane (e.g., a monolayer or bilayer membrane) and (ii) an envelope containing viral envelope glycoproteins; and (b) a multiprotein core region comprising (ii) a Gag protein, (ii) a first fusion protein comprising a Gag protein and Pro-Pol, and (iii) a second fusion protein comprising a Gag protein fused to a cargo protein via a protease-cleavable linker; In various embodiments, the cargo protein is a prime editor. In various other embodiments, the multiprotein core region of the VLP further comprises one or more guide RNA and / or pegRNA molecules that complex with the prime editor to form a ribonucleoprotein (RNP). In various embodiments, the VLP is prepared in a producer cell transiently transformed with plasmid DNA encoding the various protein and nucleic acid (sgRNA) components of the VLP. The components self-assemble at the cell membrane and bud according to the naturally occurring mechanism of retroviral budding, releasing a fully mature VLP from the cell. Once formed, Pol-Pro cleaves the protease-sensitive linker joining the Gag-cargo linker (e.g., the linker joining Gag to the PE RNP or napDNAbp RNP), releasing the PE RNP and / or napDNAbp RNA, as the case may be, within the VLP. Thus, in various embodiments, the present disclosure also provides a VLP in which the prime editor has been cleaved from the gag protein and released within the VLP. For example, the present disclosure provides a VLP comprising (i) a group-specific antigen (gag) protease (pro) polyprotein, (ii) a prime editor, and (iii) a fusion protein comprising a gag nucleocapsid protein and a nuclear export sequence (NES), encapsulated by a lipid membrane and a viral envelope glycoprotein. In some embodiments, the present disclosure provides a VLP comprising a mixture of cleavage products and uncleavage products (i.e., a mixture of the prime editor cleaved from the gag protein and the prime editor that has not yet been cleaved from the gag protein). Once the VLP is administered to a recipient cell and taken up by the cell, the contents of the VLP, including free PE RNP and / or napDNAbp RNA, are released. Once inside the cell, the RNP translocates to the cell's nucleus (particularly where the NLS is packaged in the RNP), where DNA editing may occur at the target site specified by the guide RNA.

[0171] In some embodiments, the VLP comprises an additional agent for targeting the VLP for delivery to a specific cell type. For example, such an additional targeting agent may be incorporated into the outer lipid membrane encapsulation layer of the VLP. In some embodiments, the additional targeting agent is a protein. In some embodiments, the additional targeting agent is an antibody.

[0172] Thus, as used herein, a virus-derived particle includes a virus-like particle formed by one or more virus-derived protein(s), which is substantially free of the viral genome such that the VLP is replication-incompetent when delivered to a recipient cell.

[0173] Wild type As used herein, the term "wild-type" is a term of the art understood by those skilled in the art and means the typical form of an organism, strain, gene, or characteristic occurring in nature, as distinguished from mutant or variant forms.

[0174] Detailed Description The present disclosure is based on the development and application of an engineered VLP platform, referred to herein as prime editor viral-like protein (PE-VLP), for packaging and delivering prime editor ribonucleoproteins in vitro and in vivo. These optimized PE-VLPs enable efficient prime editing in a variety of cell types. In particular, the PE-VLPs described herein are based on the surprising discovery that both a nuclear export sequence (NES) and a nuclear localization sequence (NLS) can be incorporated into the same fusion protein, facilitating trafficking of the fusion protein to different parts of the cell during production and delivery. The PE-VLPs described herein are produced in virus-producing cells and exported from the nucleus due to the presence of one or more NES sequences within the fusion protein within the PE-VLP. After delivery to a target cell, when the prime editor is released from the VLP, the NES is cleaved from the fusion protein, allowing the PE (which may contain one or more NLS sequences) to enter the nucleus of the target cell and edit the genome. The PE-VLPs described herein also contain a protease cleavage site that separates the NES and VLP proteins from the remainder of the prime editor, facilitating highly efficient cleavage and delivery of the PE. Finally, the present disclosure also describes how optimizing the ratios of the various components of a PE-VLP ensures highly efficient PE-VLP production.

[0175] Thus, the present disclosure provides virus-like particles for delivering prime editor fusion proteins (PE-VLPs), and systems comprising such PE-VLPs. The present disclosure also provides polynucleotides encoding the PE-VLPs described herein, which may be useful for producing the VLPs. Also provided herein are methods for editing the genome of a target cell by introducing the PE-VLPs currently described herein into the target cell. The present disclosure also provides fusion proteins that constitute components of the PE-VLPs described herein, as well as polynucleotides, vectors, cells, and kits.

[0176] eVLP In various embodiments, the eVLP (e.g., a PE-VLP) comprises: (a) (i) a lipid membrane (e.g., a monolayer or bilayer membrane) and (ii) an envelope containing a viral envelope glycoprotein (e.g., VSV-G), and (b) a multiprotein core region surrounded by an envelope and comprising one or more Gag-cargo fusion proteins, each comprising (i) a Gag protein, (ii) a Gag-Pro-Pol protein (the "Pro" component refers to a protease), and (iii) a Gag protein fused to a cargo protein (e.g., napDNAbp or PE or split PE) via a cleavable linker (e.g., a protease-cleavable linker (e.g., an MMLV protease-cleavable linker)); In various embodiments, the cargo protein is a napDNAbp (e.g., Cas9). In other embodiments, the cargo protein is a prime editor. In various embodiments (e.g., Figures 2A and 32), the PE may be split into a Cas9 domain and a reverse transcriptase domain, each as separate fusion proteins with Gag. In various embodiments, the split domains of the PE may contain a split intein sequence, allowing the split domains to reform the PE once delivered to a cell. In various other embodiments, the multiprotein core region of the VLP further comprises one or more pegRNA molecules and / or second-site nicking guide RNAs that complex with the napDNAbp or prime editor to form a ribonucleoprotein (RNP). In some embodiments, the pegRNA contains one or more silent mutations to increase editing efficiency by facilitating bypass of the DNA mismatch repair (MMR) pathway.

[0177] In various embodiments, VLPs are prepared in producer cells transiently transformed with plasmid DNA encoding the various protein and nucleic acid (pegRNA and guide RNA) components of the VLP. Without being bound by theory, the components self-assemble at the cell membrane and bud according to a naturally occurring budding mechanism (e.g., retroviral budding or other enveloped viral budding mechanisms) to release fully mature VLPs from the cell. Once formed, Gag-Pol-Pro cleaves the protease-sensitive linker of Gag-cargo (i.e., [Gag]-[cleavable linker]-[cargo], where cargo can be PE-RNP or napDNAbp RNP), thereby releasing, as the case may be, PE RNP and / or napDNAbp RNA within the VLP. Once the VLP is administered to and taken up by a recipient cell, the contents of the VLP are released (e.g., PE RNP and / or napDNAbp RNP). Once inside the cell, the RNP translocates to the cell's nucleus (particularly where the NLS is contained within the RNP), where DNA editing may occur at the target site specified by the guide RNA. Various embodiments include one or more improvements.

[0178] In some embodiments, the reverse transcriptase of the prime editor (e.g., a full-length prime editor, or a split prime editor) delivered by a VLP disclosed herein is an MMLV reverse transcriptase that includes a C-terminal amino acid truncation to remove an endogenous MMLV protease cleavage site. In some embodiments, the C-terminal amino acid truncation is about 1 to 180, about 1 to 170, about 1 to 160, about 1 to 150, about 1 to 140, about 1 to 130, about 1 to 120, about 1 to 110, about 1 to 100, about 1 to 90, about 1 to 80, about 1 to 70, about 1 to 60, about 1 to 50, about 1 to 40, about 1 to 30, about 1 to 20, or about 1 to 10 amino acids in length. In some embodiments, the C-terminal amino acid truncated form is about 1 to 10 amino acids in length (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 amino acids in length). In some embodiments, the C-terminal amino acid truncated form is about 6 amino acids in length. In some embodiments, the C-terminal amino acid truncated form is 6 amino acids in length.

[0179] In one embodiment, the protease-cleavable linker is optimized to improve cleavage efficiency after VLP maturation, as demonstrated herein for v.2 VLPs (i.e., "second generation" VLPs). In some embodiments, one or more additional linkers are N' and / or C' inserted relative to the cleavable linker in the fusion protein(s). Such additional linkers may be useful to better expose the protease-cleavable linker, making it more likely to be cleaved by a protease, facilitating release of the cargo protein.

[0180] In another embodiment, the Gag-cargo fusion (e.g., Gag-PE) further comprises one or more nuclear export signals at one or more locations along the length of the fusion polypeptide protein, which may be attached by a cleavable linker so that during VLP assembly in the producing cell, the Gag-cargo fusion does not accumulate in the nucleus of the producing cell (due to the presence of a competing NLS signal) but instead is available in the cytoplasm to undergo the VLP assembly process at the cell membrane. After release from the producing cell, once inside the mature VLP, the NES is cleaved by Gag-Pro-Pol, thereby separating the cargo (e.g., napDNAbp or PE) from the NES. Thus, upon delivery to the recipient cell, the cargo (e.g., napDNAbp or PE, typically flanked by one or more NLS elements) does not contain an NES element, which would otherwise prohibit the transport of the cargo to nucleases and prevent gene editing activity. This is exemplified herein as a v.3 VLP (or "third generation" VLP). In some embodiments, the NES is inserted within the gag nucleocapsid protein portion of the fusion protein. The gag nucleocapsid protein contains multiple endogenous protease sites, and inserting the NES within the gag nucleocapsid protein (e.g., rather than at one end of the gag nucleocapsid protein) can help ensure that the NES is cleaved from the cargo protein after delivery into the VLP. In some embodiments, the NES is inserted between the p12 and CA domains of the gag nucleocapsid protein. In some embodiments, the NES is inserted within the p12 domain of the gag nucleocapsid protein. In some embodiments, the NES is inserted between the p12 and MA domains of the gag nucleocapsid protein.

[0181] In other embodiments, the eVLPs disclosed herein may comprise split PE domains contained in a single, all-in-one VLP system or a two-particle system in which each PE half-domain is formed into a separate VLP. See Figure 3A and Figure 32.

[0182] In one aspect, the present disclosure provides eVLPs comprising (a) an envelope and (b) a multiprotein core, wherein the envelope comprises a lipid membrane (e.g., a lipid monolayer or bilayer) and a viral envelope glycoprotein, and wherein the multiprotein core comprises Gag (e.g., a retroviral Gag), a group-specific antigen (gag) protease (pro) polyprotein (e.g., "Gag-Pro-Pol"), and one or more fusion proteins comprising a Gag cargo (e.g., Gag-napDNAbp, Gag-reverse transcriptase, or Gag-PE). In various embodiments, the Gag cargo may comprise a ribonucleoprotein cargo (e.g., napDNAbp, reverse transcriptase, or PE complexed with a guide RNA). In yet further embodiments, the Gag cargo (e.g., Gag fused to napDNAbp, reverse transcriptase, or PE) may comprise one or more NLS sequences and / or one or more NES sequences to regulate the cellular location of the cargo within the cell. The NLS sequence would facilitate the transport of the cargo to cellular nucleases, facilitating editing. The NES would do the opposite: transport the cargo out of the nucleus and / or prevent the cargo from being transported into the nucleus. In some embodiments, the NES may be linked to the fusion protein by a cleavable linker (e.g., a protease linker), such that during assembly in the production cell, the NES signal operates to keep the cargo in the cytoplasm and available for the packaging process. However, when the mature VLP is budded or released from the production cell in its mature form, the cleavable linker attached to the NES may be cleaved, thereby releasing the binding between the NES and the cargo. Thus, without the NES, the cargo would be transferred to the nuclease bearing the NLS sequence, thereby facilitating editing. Various napDNAbps may be used in the system of the present disclosure. In some embodiments, the napDNAbp is a Cas9 protein (e.g., Cas9 nickase, dead Cas9 (dCas9), or another Cas9 variant described herein). In some embodiments, the Cas9 protein is bound to a guide RNA (gRNA).The fusion protein may further comprise other protein domains, such as an effector domain. In some embodiments, the fusion protein further comprises a deaminase domain (e.g., an adenosine deaminase domain or a cytosine deaminase domain). In certain embodiments, the fusion protein comprises a prime editor, such as the PE2, PE3, or PEmax prime editor, or any other prime editor described herein or known in the art.

[0183] In some embodiments, the fusion protein comprises one or more NESs (e.g., two, three, four, five, six, seven, eight, nine, or ten or more NESs). In some embodiments, the fusion protein further comprises a nuclear localization sequence (NLS), or one or more NLSs (e.g., two, three, four, five, six, seven, eight, nine, or ten or more NLSs). In some embodiments, the fusion protein may comprise at least one NES and one NLS.

[0184] The Gag-cargo fusion proteins described herein contain one or more cleavable linkers. In one embodiment, the Gag-cargo fusion protein includes a cleavable linker joining Gag to the cargo, such that once the Gag-cargo fusion is packaged into a mature VLP (which will also contain Gag-Pro-Pol), protease activity can cleave the Gag-cargo cleavable linker, thereby releasing the cargo. In some embodiments, the cleavable linker can also be provided in a location such that the NES is separated from the cargo protein when the cleavable linker is cleaved (e.g., by the Gag-Pro-Pol protein). This configuration of the fusion protein allows the fusion protein to be exported from the nucleus of the producing cell during PE-VLP production, and the NES can be later cleaved from the fusion protein after delivery to the target cell, or before delivery to the target cell but after packaging into a VLP, releasing PE (or release of a split PE half-domain from the same or two-particle system) and allowing it to enter the nucleus of the target cell. In some embodiments, the cleavable linker comprises a protease cleavage site (e.g., a Moloney Murine Leukemia Virus (MMLV) protease cleavage site or a Friend Murine Leukemia Virus (FMLV) protease cleavage site). Various protease cleavage sites can be used in the fusion proteins of the present disclosure. In certain embodiments, the protease cleavage site comprises the amino acid sequence TSTLLMEANSS (SEQ ID NO: 5), PRSSLYPALTP (SEQ ID NO: 6), VQALVLTQ (SEQ ID NO: 7), PLQVLTLNIERR (SEQ ID NO: 8), or an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 5-8. In some embodiments, the protease cleavage site comprises the amino acid sequence of any one of SEQ ID NOs: 5-8, including one mutation, two mutations, three mutations, four mutations, five mutations, or more than five mutations relative to one of SEQ ID NOs: 5-8. In some embodiments, the cleavable linker of the fusion protein is cleaved by a protease of the gag-pro polyprotein.In some embodiments, the cleavable linker of the fusion protein is not cleaved by the proteases of the gag-pro polyprotein until after the PE-VLP has been assembled and delivered into the target cell.

[0185] In some embodiments, one or more additional linkers are N'- and / or C'-inserted relative to the cleavable linker in the fusion protein(s). Such additional linkers may be useful for better exposing the protease-cleavable linker, allowing it to be cleaved at a higher rate by the protease, and facilitating release of the cargo protein. In some embodiments, a linker comprising the amino acid sequence G is N'- and / or C'-inserted relative to the cleavable linker. In some embodiments, a linker comprising the amino acid sequence GGS is N'- and / or C'-inserted relative to the cleavable linker. In some embodiments, a linker comprising the amino acid sequence GGS is both N'- and C'-inserted relative to the cleavable linker. In some embodiments, a linker comprising the amino acid sequence SGGSSGGS (SEQ ID NO: 163) is N'- and / or C'-inserted relative to the cleavable linker. In some embodiments, a linker comprising the amino acid sequence SGGSSGGS (SEQ ID NO: 163) is inserted both N' and C' to the cleavable linker.

[0186] In some embodiments, the gag-pro polyprotein of a PE-VLP described herein comprises an MMLV gag-pro polyprotein or an FMLV gag-pro polyprotein. In some embodiments, the gag nucleocapsid protein of a fusion protein in a PE-VLP described herein comprises an MMLV gag nucleocapsid protein or an FMLV gag nucleocapsid protein.

[0187] In some embodiments, the fusion protein delivered by the VLP comprises both a napDNAbp and a domain containing RNA-dependent DNA polymerase activity (e.g., a reverse transcriptase domain). In some embodiments, the fusion protein comprises one of the following non-limiting structures: [gag nucleocapsid protein]-[nap DNAbp]-[RT domain], where each instance of ]-[ includes an optional linker (e.g., an amino acid linker or any of the linkers provided herein); [gag nucleocapsid protein]-[1X-3X NES]-[cleavable linker]-[NLS]-[RT domain]-[napDNAbp]-[NLS], where each instance of ]-[ includes an optional linker (e.g., an amino acid linker or any of the linkers provided herein); [1X-3X NES]-[gag nucleocapsid protein]-[cleavable linker]-[NLS]-[RT domain]-[nap DNAbp]-[NLS], where each occurrence of ]-[ may include an optional linker (e.g., an amino acid linker or any of the linkers provided herein); or [gag nucleocapsid protein]-[1×-3× NES]-[cleavable linker]-[NLS]-[RT domain]-[nap DNAbp]-[NLS]-[cleavable linker]-[1×-3× NES], where each occurrence of ]-[ includes any linker (e.g., an amino acid linker or any of the linkers provided herein).

[0188] In embodiments where the cleavable linker is cleaved by a protease within the VLP, the VLP may comprise a fusion protein comprising the structure [gag nucleocapsid protein]-[1X-3X NES] and a free prime editor. In some embodiments, the prime editor comprises the structure [NLS]-[domain containing RNA-dependent DNA polymerase activity]-[napDNAbp]-[NLS].

[0189] In some embodiments, any of the above constructs comprises a 3X NES.

[0190] In some embodiments, the napDNAbp and the domain comprising RNA-dependent DNA polymerase activity (e.g., a reverse transcriptase domain) are contained on two different fusion proteins, each delivered into the VLP, or each delivered into each VLP. In some embodiments, each of the fusion proteins comprises a split intein to facilitate fusion of the napDNAbp and the domain comprising RNA-dependent DNA polymerase activity. In certain embodiments, two fusion proteins, one comprising napDNAbp and the other comprising the domain comprising RNA-dependent DNA polymerase activity, comprise the following non-limiting structures: [gag nucleocapsid protein]-[nap DNAbp]-[split intein]; and [gag nucleocapsid protein]-[split intein]-[domain comprising RNA-dependent DNA polymerase activity], where each instance of]-[ in each fusion protein includes an optional linker (e.g., an amino acid linker or any of the linkers provided herein). In one embodiment, two fusion proteins, one comprising napDNAbp and the other comprising a domain comprising RNA-dependent DNA polymerase activity, include the following non-limiting structures: [gag nucleocapsid protein]-[first part of nap DNAbp]-[split intein]; and [gag nucleocapsid protein]-[split intein]-[second portion of nap DNAbp]-[domain containing RNA-dependent DNA polymerase activity], wherein each instance of]-[ in each fusion protein includes an optional linker (e.g., an amino acid linker or any of the linkers provided herein).

[0191] The eVLPs (e.g., PE-VLPs) provided by the present disclosure comprise an outer encapsulation layer (or envelope layer) comprising a viral envelope glycoprotein. Any viral envelope glycoprotein described herein or known in the art may be used in the PE-VLPs of the present disclosure. In some embodiments, the viral envelope glycoprotein is an adenovirus envelope glycoprotein, an adeno-associated virus envelope glycoprotein, a retrovirus envelope glycoprotein, or a lentivirus envelope glycoprotein. In certain embodiments, the viral envelope glycoprotein is a retrovirus envelope glycoprotein. In some embodiments, the viral envelope glycoprotein is a vesicular stomatitis virus G protein (VSV-G), a baboon retrovirus envelope glycoprotein (BaEVRless), a FuG-B2 envelope glycoprotein, an HIV-1 envelope glycoprotein, or an ecotropic murine leukemia virus (MLV) envelope glycoprotein. In some embodiments, the viral envelope glycoprotein targets the system to a specific cell type (e.g., immune cells, neural cells, retinal pigment epithelial cells, etc.). For example, using different envelope glycoproteins in the eVLPs described herein alters their cellular tropism, allowing the PE-VLP to target a specific cell type. In some embodiments, the viral envelope glycoprotein is the VSV-G protein, and the VSV-G protein targets the system to retinal pigment epithelial (RPE) cells. In some embodiments, the viral envelope glycoprotein is the HIV-1 envelope glycoprotein, and the HIV-1 envelope glycoprotein targets the system to CD4+ cells. In some embodiments, the viral envelope glycoprotein is the FuG-B2 envelope glycoprotein, and the FuG-B2 envelope glycoprotein targets the system to neurons.

[0192] It will be understood that general methods for producing viral vector particles that generally contain a coding nucleic acid of interest are known in the art, and such methods may also be used in accordance with the present invention to produce virus-derived particles that do not contain a coding nucleic acid of interest, but instead are designed to deliver a protein cargo (e.g., PE RNP).

[0193] Conventional viral vector particles include retrovirus, lentivirus, adenovirus and adeno-associated viral vector particles, which are well known in the art.For a review of the various viral vector particles that can be used, those skilled in the art can refer to Kushnir et al. (2012, Vaccine, Vol. 31: 58-83), Zeltons (2013, Mol Biotechnol, Vol. 53: 92-107), Ludwig et al. (2007, Curr Opin Biotechnol, Vol. 18(no 6): 537-55) and Naskalaska et al. (2015, Vol. 64 (no 1): 3-13). Furthermore, references to various methods of using virus-derived particles to deliver proteins to cells can be found by those skilled in the art in the articles by Maetzig et al. (2012, Current Gene Therapy, Vol. 12: 389-409), as well as Kaczmarczyk et al. (2011, Proc Natl Acad Sci USA, Vol. 108 (no 41): 16998-17003).

[0194] Generally, the virus-like particles used in accordance with the present disclosure (which virus-like particles may also be referred to as "virus-derived particles") are formed by one or more viral structural protein(s) and / or another viral envelope protein(s).

[0195] The virus-like particles used in accordance with the present invention are replication incompetent in the host cell into which they enter.

[0196] In a preferred embodiment, the virus-like particle is formed by one or more retroviral-derived structural protein(s) and, optionally, one or more viral-derived envelope protein(s).

[0197] In a preferred embodiment, the viral structural protein is a retroviral Gag protein, or a peptide fragment thereof. As known in the art, Gag and Gag / pol precursors are expressed as polyproteins from full-length genomic RNA and require proteolytic cleavage mediated by retroviral protease (PR) to acquire a functional conformation. Furthermore, Gag, which is structurally conserved among retroviruses, is composed of at least three protein units: matrix protein (MA), capsid protein (CA), and nucleocapsid protein (NC), while Pol is composed of retroviral protease (PR), retrotranscriptase (RT), and integrase (IN).

[0198] In some embodiments, the virus-derived particle comprises retroviral Gag proteins but does not comprise Pol proteins.

[0199] As is known in the art, the host range of retroviral vectors, including lentiviral vectors, can also be expanded or altered by a process known as pseudoviralization. Pseudo-lentiviral vectors consist of viral vector particles bearing glycoproteins from other enveloped viruses. Such pseudo-viral vector particles retain the tropism of the virus from which the glycoproteins are derived.

[0200] In some embodiments, the virus-like particle is a pseudovirus-like particle that contains one or more viral structural proteins or viral envelope proteins that confer tropism to the virus-like particle for a certain eukaryotic cell.The pseudovirus-like particle as described herein may contain the viral protein used for pseudogenization as a viral envelope protein selected from the group including VSV-G protein, measles virus HA protein, measles virus F protein, influenza virus HA protein, Moloney virus MLV-A protein, Moloney virus MLV-E protein, baboon endogenous retrovirus (BAEV) envelope protein, Ebola virus glycoprotein, and foamy virus envelope protein, or a combination of two or more of these viral envelope proteins.

[0201] A well-known example of pseudogenizing viral vector particles is pseudogenizing viral vector particles with vesicular stomatitis virus glycoprotein (VSV-G). For pseudogenizing viral vector particles, those skilled in the art may refer to Yee et al. (1994, Proc Natl Acad Sci, USA, Vol. 91: 9564-9568), Cronin et al. (2005, Curr Gene Ther, Vol. 5(no. 4): 387-398), which are incorporated herein by reference.

[0202] For the production of virus-like particles, and more precisely VSV-G pseudovirus-like particles, for the delivery of a protein(s) of interest into target cells, those skilled in the art may refer to Mangeot et al. (2011, Molecular Therapy, Vol. 19 (no 9): 1656-1666).

[0203] In some embodiments, the virus-like particle further comprises a viral envelope protein, wherein (i) the viral envelope protein originates from the same virus as the viral structural proteins (e.g., originates from the same virus as the viral Gag protein), or (ii) the viral envelope protein originates from a virus distinct from the virus from which the viral structural proteins originate (e.g., originates from a virus distinct from the virus from which the viral Gag protein originates).

[0204] As will be readily appreciated by those skilled in the art, virus-like particles for use in accordance with the present disclosure may be selected from the group comprising Moloney murine leukemia virus-derived vector particles, bovine immunodeficiency virus-derived particles, simian immunodeficiency virus-derived vector particles, feline immunodeficiency virus-derived vector particles, human immunodeficiency virus-derived vector particles, equine infectious anemia virus-derived vector particles, caprine arthritis-encephalitis virus-derived vector particles, baboon endogenous virus-derived vector particles, rabies virus-derived vector particles, influenza virus-derived vector particles, norovirus-derived vector particles, respiratory syncytial virus-derived vector particles, hepatitis A virus-derived vector particles, hepatitis B virus-derived vector particles, hepatitis E virus-derived vector particles, Newcastle disease virus-derived vector particles, Norwalk virus-derived vector particles, parvovirus-derived vector particles, papillomavirus-derived vector particles, yeast retrotransposon-derived vector particles, measles virus-derived vector particles, and bacteriophage-derived vector particles.

[0205] In particular, the virus-like particles used according to the present invention are particles derived from retroviruses, which may be selected from Moloney murine leukemia virus, bovine immunodeficiency virus, simian immunodeficiency virus, feline immunodeficiency virus, human immunodeficiency virus, equine infectious anemia virus, and caprine arthritis-encephalitis virus.

[0206] In another embodiment, the virus-like particle used in accordance with the present disclosure is a lentivirus-derived particle. Lentiviruses belong to the retroviridae family and have the unique ability to infect non-dividing cells.

[0207] Such lentiviruses may be selected from among bovine immunodeficiency virus, simian immunodeficiency virus, feline immunodeficiency virus, human immunodeficiency virus, equine infectious anemia virus, and caprine arthritis-encephalitis virus.

[0208] For preparing Moloney murine leukemia virus-derived vector particles, those skilled in the art may refer to the methods disclosed by Sharma et al. (1997, Proc Natl Acad Sci USA, Vol. 94: 10803+-10808) and Guibingua et al. (2002, Molecular Therapy, Vol. 5(no. 5): 538-546), which are incorporated herein by reference. Moloney murine leukemia virus-derived (MLV-derived) vector particles may be selected from the group comprising MLV-A-derived vector particles and MLV-E-derived vector particles.

[0209] For the preparation of bovine immunodeficiency virus-derived vector particles, those skilled in the art may refer to the method disclosed by Rasmussen et al. (1990, Virology, Vol. 178(no 2): 435-451), which is incorporated herein by reference.

[0210] For the preparation of simian immunodeficiency virus-derived vector particles, including VSV-G pseudo-SIV virus-derived particles, those skilled in the art may refer in particular to the methods disclosed by Mangeot et al. (2000, Journal of Virology, Vol. 71(no 18): 8307-8315), Negre et al. (2000, Gene Therapy, Vol. 7: 1613-1623), and Mangeot et al. (2004, Nucleic Acids Research, Vol. 32 (no 12), e102), which are incorporated herein by reference.

[0211] For the preparation of feline immunodeficiency virus-derived vector particles, those skilled in the art may refer to the methods disclosed by Saenz et al. (2012, Cold Spring Harb Protoc, (1): 71-76; 2012, Cold Spring Harb Protoc, (1): 124-125; 2012, Cold Spring Harb Protoc, (1): 118-123), which are incorporated herein by reference.

[0212] For preparing human immunodeficiency virus-derived vector particles, those skilled in the art may noteworthy reference to the methods disclosed by Jalaguier et al. (2011, PlosOne, Vol. 6(no 11), e28314), Cervera et al. (J Biotechnol, Vol. 166(no 4): 152-165), and Tang et al. (2012, Journal of Virology, Vol. 86(no 14): 7662-7676), which are incorporated herein by reference.

[0213] For the preparation of equine infectious anemia virus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Olsen (1998, Gene Ther, Vol. 5(no 11): 1481-1487), which is incorporated herein by reference.

[0214] For the preparation of caprine arthritis-encephalitis virus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Mselli-Lakhal et al. (2006, J Virol Methods, Vol. 136(no 1-2): 177-184), which is incorporated herein by reference.

[0215] For the preparation of baboon endogenous virus-derived vector particles, those skilled in the art may noteworthy reference to the method disclosed by Girard-Gagnepain et al. (2014, Blood, Vol. 124(no 8): 1221-1231), which is incorporated herein by reference.

[0216] For preparing rabies virus-derived vector particles, those skilled in the art may refer in detail to the methods disclosed by Kang et al. (2015, Viruses, Vol. 7: 1134-1152, doi:10.3390 / v7031134) and Fontana et al. (2014, Vaccine, Vol. 32(no 24): 2799-27804), which are incorporated herein by reference, or to the PCT application published under No. WO2012 / 0618, which is incorporated herein by reference.

[0217] For preparing influenza virus-derived vector particles, those skilled in the art may refer in particular to the methods disclosed by Quan et al. (2012, Virology, Vol. 430: 127-135) and Latham et al. (2001, Journal of Virology, Vol. 75(no 13): 6154-6155), which are incorporated herein by reference.

[0218] For preparing norovirus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Tome-Amat et al., (2014, Microbial Cell Factories, Vol. 13: 134-142), which is incorporated herein by reference.

[0219] For preparing respiratory syncytial virus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Walpita et al. (2015, PlosOne, DOI: 10.1371 / journal.pone.0130755), which is incorporated herein by reference.

[0220] For preparing Hepatitis B virus-derived vector particles, those skilled in the art may noteworthy reference to the method disclosed by Hong et al. (2013, Journal of Virology, Vol. 87(no 12): 6615-6624), which is incorporated herein by reference.

[0221] For preparing Hepatitis E virus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Li et al. (1997, Journal of Virology, Vol. 71(no 10): 7207-7213), which is incorporated herein by reference.

[0222] For preparing Newcastle disease virus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Murawski et al. (2010, Journal of Virology, Vol. 84(no 2): 1110-1123), which is incorporated herein by reference.

[0223] For the preparation of Norwalk virus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Herbst-Kralovetz et al. (2010, Expert Rev Vaccines, Vol. 9(no 3): 299-307), which is incorporated herein by reference.

[0224] For preparing parvovirus-derived vector particles, those skilled in the art may refer in particular to the method disclosed by Ogasawara et al. (2006, In Vivo, Vol. 20: 319-324), which is incorporated herein by reference.

[0225] For preparing papillomavirus-derived vector particles, those skilled in the art may noteworthy reference to the method disclosed by Wang et al. (2013, Expert Rev Vaccines, Vol. 12(no 2): doi:10.1586 / erv.12.151), which is incorporated herein by reference.

[0226] As used herein, a virus-like particle comprises Gag protein, and most preferably comprises Gag protein originating from a virus selected from the group consisting of Rous sarcoma virus (RSV), feline immunodeficiency virus (FIV), simian immunodeficiency virus (SIV), Moloney murine leukemia virus (MLV), and human immunodeficiency viruses (HIV-1 and HIV-2), particularly human immunodeficiency virus type 1 (HIV-1).

[0227] In some embodiments, the virus-like particle may also contain one or more viral envelope protein(s). The presence of one or more viral envelope protein(s) may confer more specific tropism to the virus-derived particle for targeted cells, as is known in the art. The one or more viral envelope protein(s) may be selected from the group consisting of envelope proteins from retroviruses, envelope proteins from non-retroviral viruses, and chimeras of these viral envelope proteins with other peptides or proteins. An example of an interesting non-lentiviral envelope glycoprotein is the lymphocytic choriomeningitis virus (LCMV) strain WE54 envelope glycoprotein. These envelope glycoproteins increase the range of cells that can be transduced with retrovirus-derived vectors.

[0228] In some embodiments, the prime editing guide RNA (pegRNA) and / or second strand nicking guide RNA (ngRNA) delivered by the VLP disclosed herein comprises an aptamer. In some embodiments, the gag-pro polyprotein is fused to a target molecule that binds to an aptamer inserted into the structure of the pegRNA or ngRNA. Inclusion of such an aptamer and a target molecule that binds to the aptamer can be useful, for example, to facilitate packaging of the pegRNA and / or ngRNA into a VLP. In some embodiments, the aptamer is inserted into the pegRNA backbone sequence and / or the ngRNA backbone sequence. In some embodiments, the target molecule that binds to the aptamer is inserted into the gag-pro polyprotein. In some embodiments, the aptamer comprises an MS2 stem loop, and the target molecule that binds to the aptamer comprises an MS2 coat protein. In some embodiments, the aptamer comprises a Com aptamer, and the target molecule that binds to the aptamer comprises a Com protein. The present disclosure is not limited with respect to the aptamers and target molecules that can be utilized in the VLPs disclosed herein, and any aptamer and corresponding target molecule known in the art can be incorporated into the VLP. In some embodiments, the ratio of wild-type gag-pro polyprotein to target molecule-modified gag-pro polyprotein to one or more fusion proteins in the VLP is approximately 5:2:1. Such a ratio may provide optimal prime editing efficiency upon delivery of a prime editor cargo protein.

[0229] In some embodiments, various components of the VLPs described herein may also be fused to coiled-coil peptides to facilitate assembly of the VLP through interactions of the coiled-coil peptides. For example, in some embodiments, a first coiled-coil peptide may be inserted into the gag-pro polyprotein of the VLP. In some embodiments, a second coiled-coil peptide may be fused to one or more fusion proteins of the VLP (e.g., at the N-terminus, the C-terminus, or at an internal position within one or more fusion proteins). In some embodiments, a coiled-coil peptide is fused to the C-terminus of one or more fusion proteins.

[0230] Any coiled-coil peptide pair known in the art may be used in the VLPs described herein. For example, in some embodiments, the P3 and P4 peptides may be used: P3 peptide sequence: SPEDEIQQLEEEIAQLEQKNAALKEKNQALKYG (SEQ ID NO: 35); P4 peptide sequence: SPEDKIAQLKQKIQALKQENQQLEEENAALEYG (SEQ ID NO: 36).

[0231] In some embodiments, one of the first or second coiled-coil peptides comprises a P3 peptide, and the other of the first or second coiled-coil peptides comprises a P4 peptide. In some embodiments, the first coiled-coil peptide comprises a P3 peptide. In some embodiments, the second coiled-coil peptide comprises a P4 peptide.

[0232] napDNAbp In various embodiments, the PE-VLPs disclosed herein, as well as the prime editor fusion proteins that constitute the core components of the PE-VLPs presently described, comprise a nucleic acid programmable DNA binding protein (napDNAbp).

[0233] In various embodiments, the PE-VLP and prime editor fusion proteins may include a napDNAbp domain with a wild-type Cas9 sequence, including, for example, the canonical Streptococcus pyogenes Cas9 sequence of SEQ ID NO: 37, shown as follows: [Table 2]

[0234] In other embodiments, PE-VLPs and fusion proteins may include napDNAbp domains with modified Cas9 sequences, including, for example, the nickase variant of Streptococcus pyogenes Cas9 of SEQ ID NO: 38, which has an H840A substitution relative to wild-type SpCas9 (of SEQ ID NO: 37), as shown below: [Table 3]

[0235] The PE-VLPs and prime editor fusion proteins described herein may include any of the modified Cas9 sequences described above, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. In some embodiments, the improved prime editor fusion proteins described herein include any of the following other wild-type SpCas9 sequences, which may be modified with one or more mutations described herein at the corresponding amino acid positions: [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] [Table 4-7] [Table 4-8] [Table 4-9]

[0236] The PE-VLPs and prime editor fusion proteins described herein may include any of the above SpCas9 sequences or any variants thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. In other embodiments, the Cas9 protein may be a wild-type Cas9 ortholog from another bacterial species that differs from the standard Cas9 from Streptococcus pyogenes. For example, modified versions of the following Cas9 orthologs may be used in conjunction with the PE-VLPs and fusion proteins described herein by mutating H840A or other amino acids of interest at positions corresponding to wild-type SpCas9. Additionally, any variant Cas9 ortholog having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any of the following orthologs may also be used with a prime editor. [Table 5-1] [Table 5-2] [Table 5-3] [Table 5-4]

Table 5-5

Table 5-6

Table 5-7

Table 5-8

Table 5-9

[0237] The napDNAbp used in the PE-VLP and prime editor fusion proteins described herein may include any suitable homolog and / or ortholog, or naturally occurring enzyme, such as Cas9. Cas9 homologs and / or orthologs have been described in a variety of species, including, but not limited to, Streptococcus pyogenes and Streptococcus thermophilus. The Cas portion may be configured (e.g., obtained from nature by mutagenesis, recombinant engineering, or other means) as a nickase (i.e., capable of cleaving only one strand of target double-stranded DNA). Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain; i.e., the Cas9 is a nickase. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein as provided by any one of the Cas9 orthologs in the table above.

[0238] Reverse transcriptase domain In various embodiments, the prime editor delivered by the PE-VLP described herein comprises a reverse transcriptase domain. In some embodiments, the reverse transcriptase domain is a wild-type MMLV reverse transcriptase. In some embodiments, the reverse transcriptase domain is a variant of the wild-type MMLV reverse transcriptase having the amino acid sequence of SEQ ID NO: 60.

[0239] For example, PE2 and PEmax comprise a variant reverse transcriptase domain of SEQ ID NO:60, which is based on the wild-type MMLV reverse transcriptase domain of SEQ ID NO:59 (and, inter alia, the Genscript codon-optimized MMLV reverse transcriptase having the nucleotide sequence of SEQ ID NO:59), and which contains the amino acid substitutions D200N T306K W313F T330P L603W relative to the wild-type MMLV RT of SEQ ID NO:60. The amino acid sequence of the variant RT of PE2 and PEmax is SEQ ID NO:60.

[0240] PE-VLPs and prime editors may also comprise other variant RTs. In various embodiments, a prime editor delivered by a VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that includes one or more of the following mutations in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence: P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, or D653N.

[0241] In various embodiments, the PE-VLPs and prime editors described herein may comprise an MMLV reverse transcriptase variant,

[0242] Provided below are several exemplary reverse transcriptases that may be fused to napDNAbp proteins or provided as individual proteins according to various embodiments of the present disclosure. Exemplary reverse transcriptases include variants with at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the following wild-type enzymes or partial enzymes: [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5] [Table 6-6] [Table 6-7]

[0243] In various other embodiments, the PE-VLPs and prime editors described herein (with an RT provided either as a fusion partner or in trans) can include variant RTs that include one or more of the following mutations in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence: P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, or D653X, where "X" can be any amino acid.

[0244] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing a P51X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is L.

[0245] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing an S67X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is K.

[0246] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing an E69X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is K.

[0247] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an L139X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is P.

[0248] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains a T197X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is A.

[0249] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing a D200X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is N.

[0250] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing an H204X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is R.

[0251] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an F209X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is N.

[0252] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an E302X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is K.

[0253] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an E302X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is R.

[0254] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing a T306X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is K.

[0255] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing an F309X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is N.

[0256] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing a W313X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is F.

[0257] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing a T330X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is P.

[0258] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an L345X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is G.

[0259] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an L435X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is G.

[0260] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT containing an N454X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is K.

[0261] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains a D524X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is G.

[0262] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an E562X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is Q.

[0263] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains a D583X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is N.

[0264] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an H594X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is Q.

[0265] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an L603X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is W.

[0266] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains an E607X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is K.

[0267] In various other embodiments, a prime editor delivered by a PE-VLP described herein (with an RT provided either as a fusion partner or in trans) can include a variant RT that contains a D653X mutation in the wild-type M-MLV RT of SEQ ID NO: 59, or at the corresponding amino acid position in another wild-type RT polypeptide sequence, where "X" can be any amino acid. In some embodiments, X is N.

[0268] Provided below are several exemplary reverse transcriptases that may be fused to napDNAbp proteins or provided as individual proteins according to various embodiments of the present disclosure. Exemplary reverse transcriptases include variants with at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the wild-type or partial enzymes set forth in SEQ ID NOs:59-76.

[0269] The prime editor (PE) system described herein contemplates any publicly available reverse transcriptase enzymes described or disclosed in any of the following U.S. patents, each of which is incorporated by reference in its entirety, and any variants thereof that can be made using known methods for introducing mutations or evolving proteins: U.S. Patent Nos. 10,202,658; 10,189,831; 10,150,955; 9,932,567; 9,783,791; 9,580,698; 9,534,201; and 9,458,484. The following references describe reverse transcriptases in the art. Each of their disclosures is incorporated herein by reference in its entirety.

[0270] Herzig, E., Voronin, N., Kucherenko, N. &Hizi, A. A Novel Leu92 Mutant of HIV-1 Reverse Transcriptase with a Selective Deficiency in Strand Transfer Causes a Loss of Viral Replication. J. Virol. 89, 8119-8129 (2015).

[0271] Mohr, G. et al. A Reverse Transcriptase-Cas1 Fusion Protein Contains a Cas6 Domain Required for Both CRISPR RNA Biogenesis and RNA Spacer Acquisition. Mol. Cell 72, 700-714.e8 (2018).

[0272] Zhao, C., Liu, F. &Pyle, A. M. An ultraprocessive, accurate reverse transcriptase encoded by a metazoan group II intron. RNA 24, 183-195 (2018).

[0273] Zimmerly, S. &Wu, L. An Unexplored Diversity of Reverse Transcriptases in Bacteria. Microbiol Spectr 3, MDNA3-0058-2014 (2015).

[0274] Ostertag, E. M. & Kazazian Jr, H. H. Biology of Mammalian L1 Retrotransposons. Annual Review of Genetics 35, 501-538 (2001).

[0275] Perach, M. &Hizi, A. Catalytic Features of the Recombinant Reverse Transcriptase of Bovine Leukemia Virus Expressed in Bacteria. Virology 259, 176-189 (1999).

[0276] Lim, D. et al. Crystal structure of the Moloney murine leukemia virus RNase H domain. J. Virol. 80, 8379-8389 (2006).

[0277] Zhao, C. & Pyle, A. M. Crystal structures of a group II intron maturase reveal a missing link in spliceosome evolution. Nature Structural & Molecular Biology 23, 558-565 (2016).

[0278] Griffiths, D. J. Endogenous retroviruses in the human genome sequence. Genome Biol. 2, REVIEWS1017 (2001).

[0279] Baranauskas, A. et al. Generation and characterization of new highly thermostable and processive M-MuLV reverse transcriptase variants. Protein Eng Des Sel 25, 657-668 (2012).

[0280] Zimmerly, S., Guo, H., Perlman, P. S. & Lambowltz, A. M. Group II intron mobility occurs by target DNA-primed reverse transcription. Cell 82, 545-554 (1995).

[0281] Feng, Q., Moran, J. V., Kazazian, H. H. & Boeke, J. D. Human L1 retrotransposon encodes a conserved endonuclease required for retrotransposition. Cell 87, 905-916 (1996).

[0282] Berkhout, B., Jebbink, M. & Zsiros, J. Identification of an Active Reverse Transcriptase Enzyme Encoded by a Human Endogenous HERV-K Retrovirus. Journal of Virology 73, 2365-2375 (1999).

[0283] Kotewicz, M. L., Sampson, C. M., D’Alessio, J. M. & Gerard, G. F. Isolation of cloned Moloney murine leukemia virus reverse transcriptase lacking ribonuclease H activity. Nucleic Acids Res 16, 265-277 (1988).

[0284] Arezi, B. & Hogrefe, H. Novel mutations in Moloney Murine Leukemia Virus reverse transcriptase increase thermostability through tighter binding to template-primer. Nucleic Acids Res 37, 473-481 (2009).

[0285] Blain, S. W. & Goff, S. P. Nuclease activities of Moloney murine leukemia virus reverse transcriptase. Mutants with altered substrate specificities. J. Biol. Chem. 268, 23585-23592 (1993).

[0286] Xiong, Y. & Eickbush, T. H. Origin and evolution of retroelements based upon their reverse transcriptase sequences. EMBO J 9, 3353-3362 (1990).

[0287] Herschhorn, A. & Hizi, A. Retroviral reverse transcriptases. Cell. Mol. Life Sci. 67, 2717-2747 (2010).

[0288] Taube, R., Loya, S., Avidan, O., Perach, M. & Hizi, A. Reverse transcriptase of mouse mammary tumour virus: expression in bacteria, purification and biochemical characterization. Biochem. J. 329 (Pt 3), 579-587 (1998).

[0289] Liu, M. et al. Reverse Transcriptase-Mediated Tropism Switching in Bordetella Bacteriophage. Science 295, 2091-2094 (2002).

[0290] Luan, D. D., Korman, M. H., Jakubczak, J. L. & Eickbush, T. H. Reverse transcription of R2Bm RNA is primed by a nick at the chromosomal target site: a mechanism for non-LTR retrotransposition. Cell 72, 595-605 (1993).

[0291] Nottingham, R. M. et al. RNA-seq of human reference RNA samples using a thermostable group II intron reverse transcriptase. RNA 22, 597-613 (2016).

[0292] Telesnitsky, A. & Goff, S. P. RNase H domain mutations affect the interaction between Moloney murine leukemia virus reverse transcriptase and its primer-template. Proc. Natl. Acad. Sci. U.S.A. 90, 1276-1280 (1993).

[0293] Halvas, E. K., Svarovskaia, E. S. & Pathak, V. K. Role of Murine Leukemia Virus Reverse Transcriptase Deoxyribonucleoside Triphosphate-Binding Site in Retroviral Replication and In Vivo Fidelity. Journal of Virology 74, 10349-10358 (2000).

[0294] Nowak, E. et al. Structural analysis of monomeric retroviral reverse transcriptase in complex with an RNA / DNA hybrid. Nucleic Acids Res 41, 3874-3887 (2013).

[0295] Stamos, J. L., Lentzsch, A. M. & Lambowitz, A. M. Structure of a Thermostable Group II Intron Reverse Transcriptase with Template-Primer and Its Functional and Evolutionary Implications. Molecular Cell 68, 926-939.e4 (2017).

[0296] Das, D. & Georgiadis, M. M. The Crystal Structure of the Monomeric Reverse Transcriptase from Moloney Murine Leukemia Virus. Structure 12, 819-829 (2004).

[0297] Avidan, O., Meer, M. E., Oz, I. & Hizi, A. The processivity and fidelity of DNA synthesis exhibited by the reverse transcriptase of bovine leukemia virus. European Journal of Biochemistry 269, 859-867 (2002).

[0298] Gerard, GF et al. The role of template-primer in protection of reverse transcriptase from thermal inactivation. Nucleic Acids Res 30, 3118-3129 (2002).

[0299] Monot, C. et al. The Specificity and Flexibility of L1 Reverse Transcription Priming at Imperfect T-Tracts. PLOS Genetics 9, e1003499 (2013).

[0300] Mohr, S. et al. Thermostable group II intron reverse transcriptase fusion proteins and their use in cDNA synthesis and next-generation RNA sequencing. RNA 19, 958-970 (2013).

[0301] The above references regarding reverse transcriptase are hereby incorporated by reference in their entirety unless already stated so.

[0302] Nuclear localization sequence (NLS) In various embodiments, the fusion proteins delivered by the PE-VLPs described herein may contain one or more nuclear localization sequences (NLS), which serve to facilitate translocation of the protein into the cell nucleus. Such sequences are well known in the art and may include the following examples: [Table 7]

[0303] The above examples of NLSs are non-limiting. The prime editor fusion proteins delivered by the presently described PE-VLPs may comprise any known NLS sequence, including any of those described in Cokol et al., "Finding nuclear localization signals," EMBO Rep., 2000, 1(5): 411-415, and Freitas et al., "Mechanisms and Signals for the Nuclear Import of Proteins," Current Genomics, 2009, 10(8): 550-7, each of which is incorporated herein by reference.

[0304] In various embodiments, the fusion proteins, constructs encoding the fusion proteins, and PE-VLPs disclosed herein further comprise one or more, preferably at least two, nuclear localization sequences. In some embodiments, the fusion proteins comprise at least two NLSs. In embodiments with at least two NLSs, the NLSs can be the same or different. In some embodiments, one or more NLSs are bipartite NLSs ("bpNLSs"). In some embodiments, the disclosed fusion proteins comprise two bipartite NLSs. In some embodiments, the disclosed fusion proteins comprise two or more bipartite NLSs.

[0305] The location of the NLS fusion can be at the N-terminus, C-terminus, or within the sequence of the fusion protein (e.g., inserted between the encoded napDNAbp component (e.g., Cas9) and the polymerase domain (e.g., reverse transcriptase)).

[0306] The NLS may be any known NLS sequence in the art.The NLS may also be any NLS for nuclear localization that will be discovered in the future.The NLS may also be any naturally occurring NLS or a non-naturally occurring NLS (for example, an NLS with one or more desired mutations).

[0307] The term "nuclear localization sequence" or "NLS" refers to an amino acid sequence that promotes the import of proteins into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in International PCT Application PCT / EP2000 / 011690 by Plank et al., filed November 23, 2000 (published as WO / 2001 / 038547 on May 31, 2001), the contents of which are incorporated herein by reference. In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 30), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 21), KRTADGSEFESPKKKRKV (SEQ ID NO: 31), or KRTADGSEFEPKKKRKV (SEQ ID NO: 77). In other embodiments, the NLS comprises the amino acid sequence NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 78), PAAKRVKLD (SEQ ID NO: 24), RQRRNELKRSF (SEQ ID NO: 80), or NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 80).

[0308] In one aspect of the present disclosure, a prime editor or other fusion protein may be modified with one or more nuclear localization sequences (NLSs), preferably at least two NLSs. In some embodiments, the fusion protein is modified with two or more NLSs. The present disclosure contemplates the use of any nuclear localization sequence known in the art at the time of disclosure, or any nuclear localization sequence identified or otherwise made available in the state of the art after the time of this application. Exemplary nuclear localization sequences are peptide sequences that target a protein to the nucleus of a cell in which the sequence is expressed. Nuclear localization signals are primarily basic and can be located almost anywhere in the amino acid sequence of a protein; they generally contain short sequences of 4 to 8 amino acids (Autieri & Agrawal, (1998) J. Biol. Chem. 273: 14731-37, incorporated herein by reference), and are typically rich in lysine and arginine residues (Magin et al., (2000) Virology 274: 11-16, incorporated herein by reference). Nuclear localization sequences often contain proline residues. Various nuclear localization sequences have been identified and used to transport biological molecules from the cytoplasm to the nucleus of a cell. For example, see Tinland et al., (1992) Proc. Natl. Acad. Sci. USA 89:7442-46; Moede et al., (1999) FEBS Lett. 461:229-34. These are incorporated herein by reference. The translocation is currently thought to involve nuclear pore proteins.

[0309] Most NLSs can be classified into three groups: (i) a monopartite NLS, exemplified by the SV40 large T antigen NLS (PKKKRKV (SEQ ID NO: 30)); (ii) a bipartite motif, consisting of two basic domains separated by a variable number of spacer amino acids, exemplified by the Xenopus nucleoplasmin NLS (KRXXXXXXXXXXKKKL (SEQ ID NO: 81)); and (iii) noncanonical sequences such as M9 of hnRNP Al protein, influenza virus nucleoprotein NLS, and yeast Gal4 protein NLS ( Dingwall and Laskey 1991 ).

[0310] Nuclear localization sequences appear at various points in the amino acid sequence of a protein. NLS has been identified at the N-terminus, C-terminus, and central region of a protein. Thus, the present disclosure provides fusion proteins that may be modified with one or more NLSs at the C-terminus and / or N-terminus of the fusion protein, as well as in the internal region. The residues of longer sequences that do not function as constituent NLS residues should be selected so as not to interfere, for example, persistently or sterically, with the nuclear localization signal itself. Therefore, although there is no strict limit to the composition of the sequence that contains NLS, in practice, such sequences may be functionally limited in length and composition.

[0311] The present disclosure contemplates any suitable means of modifying a fusion protein to include one or more NLSs. In one aspect, a fusion protein may be engineered to express a fusion protein translationally fused at its N-terminus or its C-terminus (or both) to one or more NLSs, i.e., to form a prime editor-NLS fusion construct. In other embodiments, a nucleotide sequence encoding a fusion protein may be genetically modified to incorporate a reading frame encoding one or more NLSs within an internal region of the encoded prime editor. In addition, an NLS may include various amino acid linker or spacer regions encoded between the prime editor and the N-terminus, C-terminus, or internally attached NLS amino acid sequence (e.g., as well as in the central region of the protein). Thus, the present disclosure also provides nucleotide constructs, vectors, and host cells for expressing fusion proteins comprising, among other components, a prime editor and one or more NLSs.

[0312] The prime editor fusion proteins delivered by the PE-VLPs described herein may also comprise a nuclear localization sequence linked to the prime editor through one or more linkers (e.g., polymeric, amino acid, nucleic acid, polysaccharide, chemical, or nucleic acid linker elements). Linkers within the contemplated scope of this disclosure are not intended to be limiting in any way and can be any suitable type of molecule (e.g., polymeric, amino acid, polysaccharide, nucleic acid, lipid, or any synthetic chemical linker domain) and can be joined to the prime editor by any suitable strategy that forms a bond (e.g., covalent bond, hydrogen bond) between the prime editor and one or more NLSs.

[0313] Nuclear export sequence (NES) In various embodiments, the fusion proteins delivered by the PE-VLPs described herein may contain one or more nuclear export sequences (NES), which serve to facilitate translocation of the protein out of the cell nucleus. Such sequences are well known in the art and may include the following examples: [Table 8-1] [Table 8-2] [Table 8-3]

[0314] The above examples of NES are non-limiting. The prime editor fusion proteins delivered by the presently described PE-VLPs may comprise any known NES sequence, including any of those described below, each of which is incorporated herein by reference: Xu, D. et al. Sequence and structural analyzes of nuclear export signals in the NESdb database. Mol. Biol. Cell. 2012, 23(18), 3677-3693; Fung, HYJ et al. Structural determinants of nuclear export signal orientation in binding to exportin CRM1. eLife. 2015, 4:e10034; and Kosugi, S. et al. Nuclear Export Signal Consensus Sequences Defined Using a Localization-based Yeast Selection System. Traffic. 2008, 9(12), 2053-2062.

[0315] In various embodiments, the fusion proteins, constructs encoding the fusion proteins, and PE-VLPs disclosed herein further comprise one or more, preferably at least three, nuclear export sequences. In certain embodiments, the fusion protein comprises at least three NESs. In embodiments with at least three NESs, the NESs can be the same or different. The location of the NES fusion can be at the N-terminus, C-terminus, or within the sequence of the fusion protein (e.g., inserted between the encoded napDNAbp component (e.g., Cas9) and the gag nucleocapsid protein). In certain preferred embodiments, the NES (or multiple NESs (e.g., three NESs)) is positioned between the napDNAbp and the gag nucleocapsid protein so that it can be cleaved from the napDNAbp upon delivery of the fusion protein to a target cell.

[0316] The NES may be any NES sequence known in the art. The NES may also be any NES for nuclear export discovered in the future. The NES may also be a naturally occurring NES or a non-naturally occurring NES (e.g., an NES with one or more desired mutations).

[0317] The term "nuclear export sequence" or "NES" refers to an amino acid sequence that facilitates the export of a protein from the cell nucleus, e.g., by nuclear export. Nuclear export sequences are known in the art and will be apparent to those skilled in the art.

[0318] In one aspect of the present disclosure, a prime editor or other fusion protein may be modified with one or more nuclear export sequences (NESs), preferably at least three NESs. In some embodiments, the fusion protein is modified with two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more NESs. The present disclosure contemplates the use of any nuclear export sequence known in the art at the time of disclosure, or any nuclear export sequence identified or otherwise made available in the state of the art after the date of this application. A representative nuclear export sequence is a peptide sequence that directs a protein out of the nucleus of a cell in which the sequence is expressed. NESs generally contain hydrophobic amino acid residues in the sequence LXXXLXXLXL, where L is a hydrophobic residue (often leucine) and X represents any amino acid. Nuclear export sequences often contain leucine residues.

[0319] Fusion proteins delivered by the PE-VLPs described herein may also comprise a nuclear export sequence linked to the prime editor through one or more linkers (e.g., polymer, amino acid, nucleic acid, polysaccharide, chemical, or nucleic acid linker elements). Linkers within the contemplated scope of the present disclosure are not intended to be limiting in any way and can be any suitable type of molecule (e.g., polymer, amino acid, polysaccharide, nucleic acid, lipid, or any synthetic chemical linker domain) and can be joined to the prime editor by any suitable strategy that forms a bond (e.g., covalent bond, hydrogen bond) between the prime editor and one or more NESs. In some embodiments, the linker joining the one or more NESs and the prime editor is a cleavable linker, as further described herein, such that the one or more NESs can be cleaved from the prime editor, e.g., upon delivery of the prime editor to a target cell.

[0320] Linker The fusion proteins and PE-VLPs described herein may include one or more linkers. As defined above, the term "linker," as used herein, refers to a chemical group or molecule that links two molecules or moieties (e.g., the binding domain and cleavage domain of a nuclease). In some embodiments, a linker joins the gRNA binding domain of an RNA-programmable nuclease with a polymerase (e.g., a reverse transcriptase). In some embodiments, a linker joins a Cas9 nickase with a reverse transcriptase. Typically, a linker is positioned between or flanked by two groups, molecules, or other moieties and is connected to each other via a covalent bond, thereby linking the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.

[0321] The linker can be as simple as a covalent bond, or it can be a polymeric linker many atoms long. In some embodiments, the linker is polypeptide or amino acid based. In other embodiments, the linker is not peptide-like. In some embodiments, the linker is a covalent bond (e.g., carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In some embodiments, the linker is a carbon-nitrogen bond with an amide linkage. In some embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker. In some embodiments, the linker is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In some embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In some embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In some embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In some embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In other embodiments, the linker comprises an amino acid. In some embodiments, the linker comprises a peptide. In some embodiments, the linker comprises an aryl or heteroaryl moiety. In some embodiments, the linker is based on a phenyl ring. The linker may also include a functionalized moiety to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0322] In some other embodiments, the linker comprises the amino acid sequence (GGGGS) n (SEQ ID NO: 164), (G)n (SEQ ID NO: 165), (EAAAK) n (SEQ ID NO: 166), (GGS) n (SEQ ID NO: 167), (SGGS) n (SEQ ID NO: 168), (XP) n (SEQ ID NO: 169), or any combination thereof, wherein n is independently an integer between 1 and 30, and X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS) n (SEQ ID NO: 167), where n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 170). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 171). In some embodiments, the linker comprises the amino acid sequence SGGSGGSGGS (SEQ ID NO: 172). In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 162). In other embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSSGGS (SEQ ID NO: 173, 60AA). In some embodiments, the linker comprises the amino acid sequence GGS, GGSGGS (SEQ ID NO: 174), GGSGGSGGS (SEQ ID NO: 175), SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 161), SGSETPGTSESATPES (SEQ ID NO: 170), or SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSSGG S (SEQ ID NO: 173).

[0323] In certain embodiments, linkers may be used to link any of the peptides or peptide domains or portions of the invention (e.g., napDNAbp linked or fused to a reverse transcriptase domain, and / or napDNAbp linked to one or more NESs). Any of the domains of the fusion proteins described herein may be connected to each other through any of the linkers currently described below.

[0324] In some embodiments, the linker is a cleavable linker (e.g., a linker that can be split or cut by any means). The cleavable linker may be an amino acid sequence. In some embodiments, the linker between one or more NESs and napDNAbp of the fusion proteins and PE-VLPs provided herein comprises a cleavable linker. The cleavable linker may also comprise a self-cleaving peptide (e.g., a 2A peptide such as EGRGSLLTCGDVEENPGP (SEQ ID NO: 1), ATNFSLLKQAGDVEENPGP (SEQ ID NO: 2), QCTNYALLKLAGDVESNPGP (SEQ ID NO: 3), or VKQTLNFDLLKLAGDVESNPGP (SEQ ID NO: 4)). In some embodiments, the cleavable linker comprises a protease cleavage site that is cleaved after contact with a protease. For example, the present disclosure contemplates the use of a cleavable linker comprising a protease cleavage site of the amino acid sequence TSTLLMEANS (SEQ ID NO: 5), PRSSLYPALTP (SEQ ID NO: 6), VQALVLTQ (SEQ ID NO: 7), PLQVLTLNIERR (SEQ ID NO: 8), or an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 5-8. In certain embodiments, the cleavable linker comprises an MMLV protease cleavage site or an FMLV protease cleavage site. In certain embodiments, the fusion proteins and PE-VLPs described herein comprise the cleavable linker TSTLLMEANS (SEQ ID NO: 5), which joins one or more NESs and the napDNAbp. In some embodiments, the linker is cleaved upon delivery of the PE-VLP / fusion protein to a target cell, releasing a free prime editor that is capable of translocating into the nucleus of the target cell.

[0325] The protease cleavage site may be any known in the art or may be an as yet undiscovered sequence, so long as a corresponding protease would be packaged within the eVLP to allow for post-maturation cleavage within the mature eVLP particle. Such cleavage sites and their corresponding proteases include, but are not limited to: (a) Granzyme A that recognizes and cleaves a sequence including ASPRAGGK (SEQ ID NO: 243); (b) Granzyme B that recognizes and cleaves a sequence containing YEADSLEE (SEQ ID NO: 244); (c) Granzyme K that recognizes and cleaves a sequence containing YQYRAL (SEQ ID NO: 246); (d) Cathepsin D, which recognizes and cleaves sequences including LGVLIV (SEQ ID NO: 247). Many other combinations of specific proteases and protease cleavage sites may also be used in connection with the present disclosure by packaging specific proteases together during the eVLP manufacturing process. Such proteases may include, but are not limited to, Arg-C proteinase, Asp-N endopeptidase, caspase 1, caspase 2, caspase 3, caspase 4, caspase 5, caspase 7, caspase 8, caspase 9, caspase 10, chymotrypsin, clostripain, enterokinase, factor Xa, glutamyl endopeptidase, granzyme B, neutrophil elastase, pepsin, prolyl endopeptidase, proteinase K, staphylococcal peptidase I, thermolysin, thrombin, and trypsin. Any protease that pairs with its cognate recognition sequence may be used in the protease-sensitive linkers of the present disclosure, including any serine protease, cysteine ​​protease, aspartic acid protease, threonine protease, glutamic acid protease, metalloprotease, or aspartic acid peptide lyase (which constitute the major classes of known proteases). Specific protease cleavage sites for such enzymes are well known in the art and may be utilized in the linkers herein to provide protease-sensitive linkers.

[0326] Group-specific antigen (gag) proteins and viral envelope glycoproteins The PE-VLPs described herein include various viral envelope and capsid components that are used to encapsulate and deliver the prime editor fusion proteins described herein. The use of viral envelope and capsid components for nucleic acid and protein delivery is known in the art, and those skilled in the art will readily understand the various options known in the art for using or substituting these components in the PE-VLPs described herein. The use of such viral components for nucleic acid and / or protein delivery (e.g., delivery of Cas9) is described, for example, in Mangeot et al., Nat. Commun. 10, 45 (2019); Gutkin, et al. Nat. Biotechnol. (2021); and Hamilton, JR et al. Cell Reports 35(9), 109207 (2021), each of which is incorporated herein by reference.

[0327] In some embodiments, the PE-VLPs described herein comprise a viral envelope glycoprotein layer as the outermost layer of the PE-VLP. Viral envelope glycoproteins are oligosaccharide-containing proteins that form part of the viral envelope (i.e., the outermost layer of many types of viruses that protect viral genetic material as they travel between host cells). The glycoprotein may also assist in identifying and binding to receptors on the target cell membrane, allowing the viral envelope to fuse with the membrane and the contents of the viral particle (which may, for example, include the fusion protein in the PE-VLP as described herein) to enter the host cell.

[0328] The viral envelope glycoprotein used in the PE-VLP of the present disclosure may include any glycoprotein from an enveloped virus. In some embodiments, the viral envelope glycoprotein is an adenovirus envelope glycoprotein, an adeno-associated virus envelope glycoprotein, a retrovirus envelope glycoprotein, or a lentivirus envelope glycoprotein. In some embodiments, the viral envelope glycoprotein is a vesicular stomatitis virus G protein (VSV-G), a baboon retrovirus envelope glycoprotein (BaEVRless), a FuG-B2 envelope glycoprotein, or an ecotropic murine leukemia virus (MLV) envelope glycoprotein.

[0329] Any known viral envelope glycoprotein can be used in the PE-VLPs of the present disclosure. Any viral envelope glycoprotein discovered or characterized in the future can also be used in the PE-VLPs of the present disclosure. Those skilled in the art will be able to easily identify additional viral envelope glycoproteins that can be used in the PE-VLPs described herein. For example, viral envelope glycoproteins are described in Banerjee, V. and Mukhopadhyay, S. VirusDisease (2016), 27(1), 1-11, and Li, Y. et al. Front. Immunol. (2021), 12, 1-12, each of which is incorporated herein by reference.

[0330] In some embodiments, the PE-VLPs described herein further comprise an inner encapsulation layer comprising components from a viral capsid, including the gag-pro polyprotein (e.g., the gag nucleocapsid protein further comprising a viral protease linked thereto) and the gag nucleocapsid protein (e.g., the protein that constitutes the core structural component of the inner shell of many viruses, but lacks the protease of the gag-pro polyprotein), as described herein.

[0331] The Gag-pro polyprotein mediates the proteolytic cleavage of the gag and gag-pol polyproteins or nucleocapsid protein during or immediately after virion release from the plasma membrane. In the PE-VLPs described herein, the protease of the gag-pro polyprotein is responsible for cleaving the cleavable linker in the fusion protein to release the prime editor after delivery of the PE-VLP to a target cell. In some embodiments, the gag-pro polyprotein is the MMLV gag-pro polyprotein or the FMLV gag-pro polyprotein.

[0332] The gag nucleocapsid protein used in the PE-VLPs of the present disclosure can be an MMLV gag nucleocapsid protein, an FMLV gag nucleocapsid protein, or a nucleocapsid protein from any other virus that produces such a protein. In some embodiments, the gag nucleocapsid protein is fused to a napDNAbp (e.g., as part of a prime editor). In some embodiments, the fusion further comprises an NES as described herein. In certain embodiments, the gag nucleocapsid protein and NES are positioned on one side of a cleavable linker, as described herein, and the napDNAbp or prime editor is positioned on the other side of the cleavable linker, such that the prime editor can be released from the gag nucleocapsid protein upon cleavage of the cleavable linker by a protease of the gag-pro polyprotein after delivery of the PE-VLP to a target cell.

[0333] Both the gag-pro polyprotein and the gag nucleocapsid protein form the inner encapsulating layer of the PE-VLPs described herein. Any ratio of gag-pro polyprotein to gag nucleocapsid protein (i.e., as part of a fusion protein described herein) is contemplated in the PE-VLPs of the present disclosure. In some embodiments, the ratio of gag-pro polyprotein to fusion protein comprising gag nucleocapsid protein is approximately 10:1, approximately 9:1, approximately 8:1, approximately 7:1, approximately 6:1, approximately 5:1, approximately 4:1, approximately 3:1, approximately 2:1, approximately 1.5:1, approximately 1:1, or approximately 0.5:1. In some embodiments, the ratio is approximately 3:1.

[0334] Additional Prime Editor Domains A. Flap endonucleases (e.g., FEN1) In various embodiments, the PE fusion proteins delivered by the PE-VLPs described herein may include one or more flap endonucleases (e.g., FEN1), which refers to enzymes that catalyze the removal of 5' single-stranded DNA flaps (provided in trans or fused to the PE fusion protein). These are naturally occurring enzymes that process the removal of 5' flaps formed during cellular processes involving DNA replication. The prime editors delivered by the PE-VLPs described herein may utilize endogenously supplied flap endonucleases or those supplied in trans to remove 5' flaps of endogenous DNA formed at the target site during prime editing. Flap endonucleases are known in the art and are described in Patel et al., "Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends," Nucleic Acids Research, 2012, 40(10): 4507-4519, and Tsutakawa et al., "Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily," Cell, 2011, 145(2): 198-211, each of which is incorporated herein by reference. An exemplary flap endonuclease is FEN1, which can be represented by the following amino acid sequence: [Table 9]

[0335] Flap endonucleases may also include FEN1 variants, mutants, or orthologs, homologs, or variants of other flap endonucleases. Non-limiting examples of FEN1 variants are as follows: [Table 10-1] [Table 10-2] [Table 10-3]

[0336] In various embodiments, prime editor fusion proteins utilized in the methods and compositions contemplated herein may include any flap endonuclease variant of the sequences disclosed above, having an amino acid sequence that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any of the above sequences. Other endonucleases that may be utilized by the present compositions and methods to facilitate removal of 5'-terminal single-stranded DNA flaps include, but are not limited to, (1) trex2, (2) exo1 endonuclease (e.g., Keijzers et al., Biosci Rep. 2015, 35(3): e00206).

[0337] Trex2 Three prime(3') repair exonuclease 2 (TREX2) - human Accession number NM_080701 MSEAPRAETFVFLDLEATGLPSVEPEIAELSLFAVHRSSLENPEHDESGALVLPRVLDKLTLCMCPERPFTAKASEITGLSSEGLARCRKAGFDGAVVRTLQAFLSRQAGPICLVAHNGFDYDFPLLCAELRRLGARLPRDTVCLDTLPALRGLDRAHSHGTRARGRQGYSLGSLFHRYFRAEPSAAHSAEGDVHTLLLIFLHRAAELLAWADEQARGWAHIEPMYLPPDDPSLEA (SEQ ID NO: 182).

[0338] Three-prime(3') repair exonuclease 2 (TREX2)-mouse Accession number NM_011907 MSEPPRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSESLMHCGKAGFNGAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPQDTVCLDTLPALRGLDRAHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVHTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEA (SEQ ID NO: 183).

[0339] Three prime(3') repair exonuclease 2 (TREX2) - Rat Accession number NM_001107580 MSEPLRAETFVFLDLEATGLPNMDPEIAEISLFAVHRSSLENPERDDSGSLVLPRVLDKLTLCMCPERPFTAKASEITGLSSEGLMNCRKAAFNDAVVRTLQGFLSRQEGPICLVAHNGFDYDFPLLCTELQRLGAHLPRDTVCLDTLPALRGLDRVHSHGTRAQGRKSYSLASLFHRYFQAEPSAAHSAEGDVNTLLLIFLHRAPELLAWADEQARSWAHIEPMYVPPDGPSLEA (SEQ ID NO: 184).

[0340] ExoI Human exonuclease 1 (EXO1) has been implicated in many diverse DNA metabolic processes, including DNA mismatch repair (MMR), micromediated end joining, homologous recombination (HR), and replication. Human EXO1 belongs to the eukaryotic nuclease family Rad2 / XPG, which also includes FEN1 and GEN1. The Rad2 / XPG family is conserved in the nuclease domain across species, from phage to humans. The EXO1 gene product exhibits both 5' exonuclease activity and 5' flap activity. In addition, EXO1 contains intrinsic 5' RNase H activity. Human EXO1 has a high affinity for processing double-stranded DNA (dsDNA), nicks, gaps, and pseudo-Y structures, and can resolve Holliday junctions using its inherited flap activity. Human EXO1 is involved in MMR and contains a conserved binding domain that directly interacts with MLH1 and MSH2. The nucleolytic activity of EXO1 is positively activated by PCNA, MutSα (MSH2 / MSH6 complex), 14-3-3, MRN, and the 9-1-1 complex.

[0341] Exonuclease 1 (EXO1) Accession number NM_003686 (Homo sapiens exonuclease 1 (EXO1), transcript variant 3)-isoform A MGIQGLLQFIKEASEPIHVRKYKGQVVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSHGIKPILVFDGCTLPSKKEVERSRRERRQANLLKGKQLLREGKVSEARECFTRSINITHAMAHKVIKAARSQGVDCLVAPYEADAQLAYLNKAGIVQAIITEDSDLLAFGCKKVILKMDQFGNGLEIDQARL GMCRQLGDVFTEEKFRYMCILSGCDYLSSLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLKMNITVPEDYINGFIRANNTFLYQLVFDPIKRKLIPLNAYEDDVDPETLSYAGQYVDDSIALQIALGNKDINTFEQIDDYNPDTAMPAHSRSHSWDDKTCQKSANVSSIWHRNYSPRPESGTVSDAPQLKENPSTVGVER VISTKGLNLPRKSSIVKRPRSAELSEDDLLSQYSLSFTKKTKKNSSEGNKSLSFSEVFVPDLVNGPTNKKSVSTPPRTRNKFATFLQRKNEESGAVVVPGTRSRFFCSSDSTDCVSNKVSIQPLDETAVTDKENNLHESEYGDQEGKRLVDTDVARNSSDDIPNNHIPGDHIPDKATVFTDEESYSFESSKFTRTISPPTLGTLRSCFSWSGGLGDFSRTPSPSPSTALQQFRRKSDSPTSLPENNMSDVSQLKSEESSDDESHPLREEACSSQSQESGEFSLQSSNASKLSQCSSKDSDSEESDCNIKLLDSQSDQTSKLRLSHFSKKDTPLRNKVPGLYKSSSADSLSTTKIKPLGPARASGLSKKPASIQKRKHHNAENKPGLQIKLNELWKNFGFKKF (SEQ ID NO: 185).

[0342] Exonuclease 1 (EXO1) Accession number NM_006027 (Homo sapiens exonuclease 1 (EXO1), transcript variant 3)-isoform B MGIQGLLQFIKEASEPIHVRKYKGQVVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSHGIKPILVFDGCTLPSKKEVERSRRERRQANLLKGKQ LLREGKVSEARECFTRSINITHAMAHKVIKAARSQGVDCLVAPYEADAQLAYLNKAGIVQAIITEDSDLLAFGCKKVILKMDQFGNGLEIDQARLGMCRQLGDVFT EEKFRYMCILSGCDYLSSLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLKMNITVPEDYINGFIRANNTFLYQLVFDPIKRKLIPLNAYEDDVDPETLSYAGQYVDDSIALQIALGNKDINTFEQIDDYNPDTAMPAHSRSHSWDDKTCQKSANVSSIWHRNYSPRPESGTVSDAPQLKENPSTVGVERVISTKGLNLPRKSSIVKRPRSA ELSEDDLLSQYSLSFTKKTKKNSSEGNKSLSFSEVFVPDLVNGPTNKKSVSTPPRTRNKFATFLQRKNEESGAVVVPGTRSRFFCSSDSTDCVSNKVSIQPLDETAVTDKENNLHESEYGDQEGKRLVDTDVARNSSDDIPNNHIPGDHIPDKATVFTDEESYSFESSKFTRTISPPTLGTLRSCFSWSGGLGDFSRTPSPSPSTALQQFR RKSDSPTSLPENNMSDVSQLKSEESSDDESHPLREEACSSQSQESGEFSLQSSNASKLSQCSSKDSDSEESDCNIKLLDSQSDQTSKLRLSHFFSKKDTPLRNKVPG LYKSSSADSLSTTKIKPLGPARASGLSKKPASIQKRKHHNAENKPGLQIKLNELWKNFGFKKDSEKLPPCKKPLSPVRDNIQLTPEAEEDIFNKPECGRVQRAIFQ (SEQ ID NO: 186).

[0343] Exonuclease 1 (EXO1) Accession number NM_001319224 (Homo sapiens exonuclease 1 (EXO1), transcript variant 4)-isoform C MGIQGLLQFIKEASEPIHVRKYKGQVVAVDTYCWLHKGAIACAEKLAKGEPTDRYVGFCMKFVNMLLSHGIKPILVFDGCTLPSKKEVERSRRERRQANLLKGKQLLREGKVSEARECFTRSINITHAMAHKVIKAARSQGVDCLVAPYEADAQLAYLNKAGIVQAIITEDSDLLAFGCKKVILKMDQFGNGLEIDQARLGMCRQLGDVFTEEKFRYMCILSGCDYLSSLRGIGLAKACKVLRLANNPDIVKVIKKIGHYLKMNITVPEDYINGFIRANNTFLYQLVFDPIKRKLIPLNAYEDDVDPETLSYAGQYVDDSIALQIALGNKDINTFEQIDDYNPDTAMPAHSRSHSWDDKTCQKSANVSSIWHRNYSPRPESGTVSDAPQLKENPSTVGVERVISTKGLNLPRKSSIVKRPRSELSEDDLLSQYSLSFTKKTKKNSSEGNKSLSFSEVFVPDLVNGPTNKKSVSTPPRTRNKFATFLQRKNEESGAVVVPGTRSRFFCSSDSTDCVSNKVSIQPLDETAVTDKENNLHESEYGDQEGKRLVDTDVARNSSDDIPNNHIPGDHIPDKATVFTDEESYSFESSKFTRTISPPTLGTLRSCFSWSGGLGDFSRTPSPSPSTALQQFRRKSDSPTSLPENNMSDVSQLKSEESSDDESHPLREEACSSQSQESGEFSLQSSNASKLSQCSSKDSDSEESDCNIKLLDSQSDQTSKLRLSHFSKKDTPLRNKVPGLYKSSSADSLSTTKIKPLGPARASGLSKKPASIQKRKHHNAENKPGLQIKLNELWKNFGFKKDSEKLPPCKKPLSPVRDNIQLTPEAEEDIFNKPECGRVQRAIFQ (SEQ ID NO: 187).

[0344] B. Inteins and split inteins It will be appreciated that in some embodiments (e.g., delivery of a prime editor in vivo), it may be advantageous to split a polypeptide (e.g., reverse transcriptase or napDNAbp) or a fusion protein (e.g., a prime editor) into N- and C-terminal halves, deliver them separately, and then allow them to colocalize to reconstitute the complete protein (or fusion protein, as the case may be) within the cell. The separate halves of the protein or fusion protein may each comprise a split intein tag to facilitate reconstitution of the complete protein or fusion protein by a protein trans-splicing mechanism.

[0345] Protein trans-splicing catalyzed by split inteins provides a fully enzymatic method for protein ligation. Split inteins are essentially continuous inteins (e.g., mini-inteins) that have been split into two parts, designated N-intein and C-intein, respectively. The N-intein and C-intein of a split intein can non-covalently bind to form an active intein and catalyze a splicing reaction in essentially the same manner as a continuous intein does. Split inteins have been found in nature and have also been engineered in laboratories. As used herein, the term "split intein" refers to any intein in which there are one or more peptide bond breaks between the N- and C-terminal amino acid sequences, resulting in separate molecules that can non-covalently recombine or rearrange to form an intein functional in a trans-splicing reaction. Any catalytically active intein or fragment thereof may be used to derive a split intein for use in the methods of the present invention. For example, in one aspect, the split intein may be derived from a eukaryotic intein. In another aspect, the split intein may be derived from a bacterial intein. In another aspect, the split intein may be derived from an archaeal intein. Preferably, such a derived split intein retains only the amino acid sequence essential for catalyzing a trans-splicing reaction.

[0346] As used herein, "N-terminal split intein (In)" refers to any intein sequence that contains an N-terminal amino acid sequence that is functional in a trans-splicing reaction. Thus, In also includes sequences that are spliced ​​when trans-splicing occurs. In can include sequences that are modifications of the N-terminal portion of a naturally occurring intein sequence. For example, In can contain additional amino acid residues and / or mutated residues, as long as the inclusion of such additional and / or mutated residues does not render the In functional in trans-splicing. Preferably, the inclusion of additional and / or mutated residues improves or enhances the trans-splicing activity of In.

[0347] As used herein, "C-terminal split intein (Ic)" refers to any intein sequence that contains a C-terminal amino acid sequence that is functional in a trans-splicing reaction. In one aspect, Ic contains 4 to 7 consecutive amino acid residues, at least 4 of which are from the final β-strand of the intein from which it is derived. Thus, Ic also contains the sequence that is spliced ​​when trans-splicing occurs. Ic can include a sequence that is a modification of the C-terminal portion of a naturally occurring intein sequence. For example, Ic can contain additional amino acid residues and / or mutated residues, as long as the inclusion of such additional and / or mutated residues does not render the In functional in trans-splicing. Preferably, the inclusion of additional and / or mutated residues improves or enhances the trans-splicing activity of Ic.

[0348] In some embodiments of the present invention, the peptide linked to Ic or In may contain additional chemical moieties, including, among others, fluorescent groups, biotin, polyethylene glycol (PEG), amino acid analogs, unnatural amino acids, phosphate groups, glycosyl groups, radioisotope labels, and pharmaceutical molecules. In other embodiments, the peptide linked to Ic may contain one or more chemically reactive groups, including, among others, ketones, aldehydes, Cys residues, and Lys residues. The N-intein and C-intein of a split intein can non-covalently bind to form an active intein and catalyze the splicing reaction when an "intein-splicing polypeptide (ISP)" is present. As used herein, "intein-splicing polypeptide (ISP)" refers to the remaining portion of the amino acid sequence of a split intein when Ic, In, or both are removed from the split intein. In some embodiments, In contains an ISP. In other embodiments, Ic contains an ISP. In another embodiment, the ISP is a separate peptide that is not covalently linked to either In or Ic.

[0349] Split inteins can also be created from consecutive inteins by engineering one or more split sites into the unstructured loops or intervening amino acid sequences between the 12 conserved beta-strands found in the mini-intein structure. The location of the split site within the region between the beta-strands can be somewhat flexible, provided that the creation of the split does not disrupt the intein structure, particularly the structured beta-strands, to a sufficient extent that the splicing activity of the protein is lost.

[0350] In protein trans-splicing, one precursor protein consists of an N-extein portion followed by an N-intein, another precursor protein consists of a C-intein portion followed by a C-extein portion, and the trans-splicing reaction (catalyzed jointly by the N- and C-inteins) excises the two intein sequences and joins the two extein sequences with a peptide bond. Because protein trans-splicing is an enzymatic reaction, it can work even at very low (e.g., micromolar) concentrations of protein and can be carried out under physiological conditions.

[0351] An exemplary sequence is as follows: [Table 11-1] [Table 11-2] [Table 11-3]

[0352] Inteins are most frequently found as a continuous domain, but some naturally exist in split forms, where the two fragments are expressed as separate polypeptides and must be joined before splicing (so-called protein trans-splicing) can occur.

[0353] An exemplary split intein is the Ssp DnaE intein, which contains two subunits, DnaE-N and DnaE-C. The two distinct subunits are encoded by separate genes, dnaE-n and dnaE-c, which encode the DnaE-N and DnaE-C subunits, respectively. DnaE is a naturally occurring split intein in Synechocytis sp. PCC6803 that is capable of directing the trans-splicing of two separate proteins, each containing a fusion with either DnaE-N or DnaE-C.

[0354] Additional naturally occurring or engineered split intein sequences are known in the art or can be generated from the full intein sequences described herein or available in the art. Examples of split intein sequences can be found in Stevens et al., "A promiscuous split intein with expanded protein engineering applications," PNAS, 2017, Vol. 114: 8538-8543; Iwai et al., "Highly efficient protein trans-splicing by a naturally split DnaE intein from Nostoc punctiforme," FEBS Lett, 580: 1853-1858, each of which is incorporated herein by reference. Additional split intein sequences can be found, for example, in WO 2013 / 045632, WO 2014 / 055782, WO 2016 / 069774, and EP2877490, the contents of each of which are incorporated herein by reference. Additionally, in trans protein splicing has been described in vivo and in vitro (Shingledecker, et al., Gene 207:187 (1998); Southworth, et al., EMBO J. 17:918 (1998); Mills, et al., Proc. Natl. Acad. Sci. USA, 95:3543-3548 (1998); Lew, et al., J. Biol. Chem., 273:15887-15890 (1998); Wu, et al., Biochim. Biophys. Acta 35732:1 (1998b); Yamazaki, et al., J. Am. Chem. Soc. 120:5591 (1998); Evans, et al., J. Biol. Chem. 275:9091). (2000); Otomo, et al., Biochemistry 38:16040-16044 (1999); Otomo, et al., J. Biolmol. NMR 14:105-114 (1999); Scott, et al., Proc. Natl. Acad. Sci. USA 96:13638-13643 (1999)), providing the opportunity to express a protein as two inactive fragments followed by ligation to form a functional product.

[0355] RNA-protein interaction domains In various embodiments, two separate protein domains (e.g., a Cas9 domain and a polymerase domain) may be co-localized with each other to form a functional complex (similar to the function of a fusion protein containing two separate protein domains) by using an "RNA-protein recruitment system," such as the "MS2 tagging technique." Such systems generally involve tagging one protein domain with an "RNA-protein interaction domain" (also known as an "RNA-protein recruitment domain"), and tagging the other with an "RNA-binding protein" that specifically recognizes and binds to the RNA-protein interaction domain (e.g., a specific hairpin structure). These types of systems can be utilized to co-localize domains of a prime editor, as well as to recruit additional functions to the prime editor, such as a UGI domain. In one example, the MS2 tagging technique is based on the natural interaction of the MS2 bacteriophage coat protein ("MCP" or "MS2cp") with a stem-loop or hairpin structure, i.e., an "MS2 hairpin," present in the genome of the phage. In the case of the MS2 hairpin, it is recognized and bound by the MS2 bacteriophage coat protein (MCP), and thus, in one exemplary scenario, a reverse transcriptase-MS2 fusion may employ a Cas9-MCP fusion.

[0356] The scope of other modular RNA-protein interaction domains has been described in the art, for example, in Johansson et al., "RNA recognition by the MS2 phage coat protein," Sem. Virol., 1997, Vol. 8(3): 176-185; Delebecque et al., "Organization of intracellular reactions with rationally designed RNA assemblies," Science, 2011, Vol. 333: 470-474; Mali et al., "Cas9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering," Nat. Biotechnol., 2013, Vol. 31: 833-838; and Zalatan et al., "Engineering complex synthetic transcriptional programs with CRISPR RNA scaffolds," Cell, 2015, Vol. 160: 339-350, each of which is incorporated herein by reference. Other systems include the PP7 hairpin, which specifically recruits the PCP protein, and the "com" hairpin, which specifically recruits the Com protein. See Zalatan et al.

[0357] The nucleotide sequence of the MS2 hairpin (or equivalently called the "MS2 aptamer") is as follows: GCCAACATGAGGATCACCCATGTCTGCAGGGCC (SEQ ID NO: 196).

[0358] The amino acid sequence of MCP or MS2cp is as follows: GSASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKVATQTVGGEELPVAGWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY (SEQ ID NO: 197).

[0359] C. UGI domain In other embodiments, the prime editor delivered by the PE-VLP described herein may comprise one or more uracil glycosylase inhibitor domains. The term "uracil glycosylase inhibitor (UGI)" or "UGI domain," as used herein, refers to a protein capable of inhibiting the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or the UGI set forth in SEQ ID NO: 198. In some embodiments, the UGI proteins provided herein encompass fragments of UGI and proteins homologous to UGI or the UGI fragment. For example, in some embodiments, the UGI domain comprises a fragment of the amino acid sequence set forth in SEQ ID NO: 198. In some embodiments, the UGI fragment comprises an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid sequence set forth in SEQ ID NO: 198. In some embodiments, UGI comprises an amino acid sequence homologous to the amino acid sequence set forth in SEQ ID NO: 198, or an amino acid sequence homologous to a fragment of the amino acid sequence set forth in SEQ ID NO: 198. In some embodiments, a protein comprising UGI, or a fragment of UGI, or a homolog of UGI or a UGI fragment, is referred to as a "UGI variant." A UGI variant shares homology with UGI or a fragment thereof. For example, a UGI variant is at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% identical to wild-type UGI or UGI set forth in SEQ ID NO: 198.In some embodiments, the UGI variant comprises a fragment of UGI, such that the fragment is at least 70% identical, at least 80% identical, at least 90% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or at least 99.9% to the corresponding fragment of wild-type UGI or UGI set forth in SEQ ID NO: 198. In some embodiments, UGI comprises the following amino acid sequence: Uracil-DNA glycosylase inhibitors: >sp|P14739|UNGI_BPPB2 MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 198).

[0360] The prime editor utilized in the methods and compositions described herein may comprise two or more UGI domains, optionally separated by one or more linkers, as described herein.

[0361] D. Additional PE elements In some embodiments, the prime editor utilized in the methods and compositions described herein may comprise an inhibitor of base repair. The term "inhibitor of base repair" or "IBR" refers to a protein capable of inhibiting the activity of a nucleic acid repair enzyme, such as a base excision repair enzyme. In some embodiments, the IBR is an inhibitor of OGG base excision repair. In some embodiments, the IBR is an inhibitor of base excision repair ("iBER"). Exemplary inhibitors of base excision repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 EndoI, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is an iBER, which may be a small molecule or peptide inhibitor of a catalytically inactive glycosylase or a catalytically inactive dioxygenase or oxidase, or a variant thereof. In some embodiments, the IBR is an iBER that may be a TDG inhibitor, an MBD4 inhibitor, or an inhibitor of AlkBH enzymes. In some embodiments, the IBR is an iBER that includes catalytically inactive TDG or catalytically inactive MBD4. An exemplary catalytically inactive TDG is the N140A mutant of SEQ ID NO: 202 (human TDG).

[0362] Some exemplary glycosylases are shown below: Catalytically inactivated variants of any of these glycosylase domains are iBERs that may be fused to the napDNAbp or polymerase domains of the prime editors utilized in the methods and compositions provided herein.

[0363] OGG (human) MPARALLPRRMGHRTLASTPALWASIPCPRSELRLDLVLPSGQSFRWREQSPAHWSGVLADQVWTLTQTEEQLHCTVYRGDKSQASRPTPDELEAVRKYFQLDVTLAQLYHHWGSVDSHFQEVAQKFQGVRLLRQDPIECLFSFICSSNNNIARITGMVERLCQAFGPRLIQLDDVTYHGFPSLQALAGPEVEAHLRKLGLGYRARYVSASARAILEEQGGLAWLQQLRESSYEEAHKALCILPGVGTKVADCICLMALDKPQAVPVDVHMWHIAQRDYSWHPTTSQAKGPSPQTNKELGNFFRSLWGPYAGWAQAVLFSADLRQSRHAQEPPAKRRKGSKGPEG (SEQ ID NO: 199)

[0364] MPG (human) MVTPALQMKKPKQFCRRMGQKKQRPARAGQPHSSSDAAQAPAEQPHSSSDAAQAPCPRERCLGPPTTPGPYRSIYFSSPKGHLTRLGLEFFDQPAVPLARAFLGQVLVRRLPNGTELRGRIVETEAYLGPEDEAAHSRGGRQTPRNRGMFMKPGTLYVYIIYGMYFCMNISSQGDGACVLLRALEPLEGLETMRQLRSTLRKGTASRVLKDRELCSGPSKLCQALAINKSFDQRDLAQDEAVWLERGPLEPSEPAVVAAARVGVGHAGEWARKPLRFYVRGSPWVSVVDRVAEQDTQA (SEQ ID NO: 200)

[0365] MBD4 (human) MGTTGLESLSLGDRGAAPTVTSSERLVPDPPNDLRKEDVAMELERVGEDEEQMMIKRSSECNPLLQEPIASAQFGATAGTECRKSVPCGWERVVKQRLFGKTAGRFDVYFISPQGLKFRSKSSLANYLHKNGETSLKPEDFDFTV LSKRGIKSRYKDCSMAALTSHLQNQSNNSNWNLRTRSKCKKDVFMPPSSSSELQESRGLSNFTSTHLLLKEDEGVDDVNFRKVRKPKGKVTILKGIPIKKTKKGCRKSCSGFVQSDSKRESVCNKADAESEPVAQKSQLDRTVCI SDAGACGETLSVTSEENSLVKKKERSLSSGSNFCSEQKTSGIINKFCSAKDSEHNEKYEDTFLESEEIGTKVEVVERKEHLHTDILKRGSEMDNNCSPTRKDFTGEKIFQEDTIPRTQIERRKTSLYFSSKYNKEALSPPRRKAFKKWTPPRSPFNLVQETLFHDPWKLLIATIFLNRTSGKMAIPVLWKFLEKYPSAEVARTADWRDVSELLKPLGLYDLRAKTIVKFSDEYLTKQWKYPIELHGIGKYGNDSYRIFCVNEWKQVHPEDHKLNKYHDWLWENHEKLSLS (SEQ ID NO: 201)

[0366] TDG (human) MEAENAGSYSLQQAQAFYTFPFQQLMAEAPNMAVVNEQQMPEEVPAPAPAQEPVQEAPKGRKRKPRTTEPKQPVEPKKPVESKKSGKSAKSKEKQEKITDTF KVKRKVDRFNGVSEAELLTKTLPDILTFNLDIVIIGINPGLMAAYKGHHYPGPGNHFWKCLFMSGLSEVQLNHMDDHTLPGKYGIGFTNMVERTTPGSKDLSS KEFREGGRILVQKLQKYQPRIAVFNGKCIYEIFSKEVFGVKVKNLEFGLQPHKIPDTETLCYVMPSSSARCAQFPRAQDKVHYYIKLKDLRDQLKGIERNMD VQEVQYTFDLQLAQEDAKKMAVKEEKYDPGYEAAYGGAYGENPCSSEPCGFSSNGLIESVELRGESAFSGIPNGQWMTQSFTDQIPSFSNHCGTQEQEEESHA (SEQ ID NO: 202)

[0367] In some embodiments, the fusion proteins described herein may contain one or more heterologous protein domains (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more domains in addition to the prime editor component). The fusion protein may also contain any additional protein sequences, and optionally, a linker sequence between any two domains. Other exemplary features that may be present are localization sequences, such as a cytoplasmic localization sequence, a transport sequence, such as a nuclear export sequence, or other localization sequence, as well as sequence tags useful for solubilizing, purifying, or detecting the fusion protein.

[0368] Examples of protein domains that may be fused to a prime editor or its components (e.g., napDNAbp domain, polymerase domain, or NLS domain) include, but are not limited to, epitope tags and reporter gene sequences. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins, including blue fluorescent protein (BFP). Prime editors may also be fused to genetic sequences encoding proteins or protein fragments that bind DNA molecules or other cellular molecules, including, but not limited to, maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that may be part of prime editors are described in U.S. Patent Publication No. 2011 / 0059502, published March 10, 2011, and incorporated herein by reference in its entirety.

[0369] In one aspect of the present disclosure, reporter genes, including but not limited to glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins, including blue fluorescent protein (BFP), may be introduced into cells to encode gene products that serve as markers for measuring altered or modified expression of the gene product. In one embodiment of the present disclosure, the gene product is luciferase. In a further embodiment of the present disclosure, the expression of the gene product is reduced.

[0370] Suitable protein tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, polyhistidine tags (also called histidine tags or His tags), maltose-binding protein (MBP) tags, nus tags, glutathione-S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag 1, Softag 3), streptococcal tags, biotin ligase tags, FlAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein comprises one or more His tags.

[0371] In some embodiments of the present disclosure, the activity of the prime editing system delivered by the presently described PE-VLPs may be temporally modulated by adjusting the residence time, amount, and / or activity of the expression components of the PE system. For example, as described herein, the PE may be fused to a protein domain that can modify the intracellular half-life of the PE. In certain embodiments involving two or more vectors (e.g., vector systems in which the components described herein are encoded on two or more separate vectors), the activity of the PE system may be temporally modulated by controlling the timing at which the vectors are delivered. For example, in some embodiments, the vector encoding the nuclease system may deliver the PE prior to the vector encoding the template. In other embodiments, the vector encoding the PEgRNA may deliver the guide prior to the vector encoding the PE system. In some embodiments, the vectors encoding the PE system and the PEgRNA are delivered simultaneously. In some embodiments, the co-delivered vectors temporarily deliver, for example, the PE, PEgRNA, and / or second-strand guide RNA components. In further embodiments, the RNA transcribed from the coding sequence on the vector (e.g., a nuclease transcript) may further comprise at least one element capable of modifying the intracellular half-life of the RNA and / or modulating translational control. In some embodiments, the half-life of the RNA may be increased. In some embodiments, the half-life of the RNA may be decreased. In some embodiments, the element may be capable of increasing the stability of the RNA. In some embodiments, the element may be capable of decreasing the stability of the RNA. In some embodiments, the element may be within the 3'UTR of the RNA. In some embodiments, the element may include a polyadenylation signal (PA). In some embodiments, the element may include a cap (e.g., at the end of an upstream mRNA or PEG RNA). In some embodiments, the RNA may not include a PA, such that it is more rapidly degraded within the cell after transcription.In some embodiments, the element may include at least one AU-rich element (ARE). The ARE may be bound by an ARE-binding protein (ARE-BP) in a manner that is tissue-type, cell-type, timing, cellular localization, and environment-dependent. In some embodiments, the destabilizing element may promote RNA decay, affect RNA stability, or activate translation. In some embodiments, the ARE may comprise 50-150 nucleotides in length. In some embodiments, the ARE may comprise at least one copy of the sequence AUUUA. In some embodiments, at least one ARE may be added to the 3'UTR of the RNA. In some embodiments, the element may be woodchuck hepatitis virus (WHP).

[0372] Posttranscriptional regulatory elements (WPREs) create tertiary structures to enhance expression from transcripts. In further embodiments, the elements are modified and / or truncated WPRE sequences that can enhance expression from transcripts, e.g., as described in Zufferey et al., J Virol, 73(4): 2886-92 (1999) and Flajolet et al., J Virol, 72(7): 6175-80 (1998). In some embodiments, a WPRE or equivalent may be added to the 3'UTR of an RNA. In some embodiments, the elements may be selected from other RNA sequence motifs that are enriched in either fast-decay or slow-decay transcripts.

[0373] In some embodiments, the vector encoding PE or PEgRNA may be self-destructed by the PE system through cleavage of the target sequence present on the vector. Cleavage may prevent continued transcription of PE or PEgRNA from the vector. Although transcription may occur for some time on the linearized vector, the expressed transcript or protein that is subject to intracellular degradation will have less time to produce off-target effects if it is not continuously supplied by the expression of the encoding vector.

[0374] Delivery of MMR inhibitors using PE-VLPs In some embodiments, the present disclosure contemplates the delivery of a mismatch repair (MMR) pathway inhibitor using the PE-VLPs described herein alongside a prime editor to enhance the efficiency of prime editing. Accordingly, the present disclosure contemplates any suitable means for inhibiting MMR. In one embodiment, the present disclosure encompasses the administration of an effective amount of an MMR pathway inhibitor. In various embodiments, the MMR pathway may be inhibited by inhibiting, blocking, or inactivating any one or more MMR proteins or variants at the genetic level (e.g., in genes encoding one or more MMR proteins, such as by introducing a mutation that inactivates the MMR protein or variants thereof), at the transcriptional level (e.g., by knocking down the transcript), at the translational level (e.g., by blocking the translation of one or more MMR proteins (from their cognate transcripts)), or at the protein level (e.g., by applying an inhibitor (e.g., a small molecule, antibody, dominant-negative protein partner), or by targeted proteolysis (e.g., PROTAC-based degradation)). The present disclosure also contemplates a method of prime editing using the PE-VLPs described herein, which are designed to introduce modifications into nucleic acid molecules that avoid MMR pathway correction without the need to provide an MMR inhibitor. Using the PE-VLPs described herein to deliver an MMR inhibitor together with a prime editor or to introduce modifications into nucleic acid molecules that avoid MMR pathway correction results in increased editing efficiency and reduced indel formation. As used herein, "during" prime editing can encompass any suitable order of events, such that the prime editing step can be applied before, simultaneously with, or after a step that blocks, inhibits, or inactivates the MMR pathway (e.g., by targeting inhibition of MLH1). For example, in some embodiments, an inhibitor of the MMR pathway can be delivered simultaneously with the prime editor, either in the same PE-VLP or in a separate PE-VLP.In some embodiments, the inhibitor of the MMR pathway may be delivered before delivery of the prime editor or after delivery of the prime editor.

[0375] In some embodiments, a prime editing system component (e.g., a PEGRNA) is designed to introduce modifications into a target nucleic acid that circumvent the MMR system without the need to provide an inhibitor. In certain embodiments, the DNA mismatch repair (MMR) system can be inhibited, blocked, or otherwise inactivated by inhibiting one or more proteins of the MMR system, including, but not limited to, MLH1, PMS2 (or MutL alpha), PMS1 (or MutL beta), MLH3 (or MutL gamma), MutS alpha (MSH2-MSH6), MutS beta (MSH2-MSH3), MSH2, MSH6, PCNA, RFC, EXO1, POLδ, and PCNA.

[0376] Thus, in one aspect, the present disclosure provides a method of editing a nucleotide molecule (e.g., a genome) by delivering an inhibitor of the MMR pathway and a prime editor using the PE-VLPs described herein.

[0377] In another aspect, the present disclosure provides methods for editing nucleotide molecules (e.g., genomes) by delivering inhibitors of the MMR system (e.g., MLH1, PMS2 (or MutL alpha), PMS1 (or MutL beta), MLH3 (or MutL gamma), MutS alpha (MSH2-MSH6), MutS beta (MSH2-MSH3), MSH2, MSH6, PCNA, RFC, EXO1, POLδ, and PCNA), and prime editors using the PE-VLPs described herein.

[0378] In one aspect, the present disclosure relates to the delivery of a prime editor and an inhibitor of MLH1 or its variants using the PE-VLPs described herein. Without being bound by theory, MLH1 is a key MMR protein that heterodimerizes with PMS2 to form MutL alpha, a component of the post-replicative DNA mismatch repair (MMR) system. DNA repair is initiated by the binding of MutS alpha (MSH2-MSH6) or MutS beta (MSH2-MSH3) to a dsDNA mismatch, followed by the recruitment of MutL alpha to the heteroduplex. The assembly of the MutL-MutS-heteroduplex ternary complex in the presence of RFC and PCNA is sufficient to activate the endonuclease activity of PMS2. This introduces a single-strand break near the mismatch, thereby generating a new entry point for the exonuclease EXO1 to degrade the mismatch-containing strand. DNA methylation prevents cleavage, thus ensuring that only the newly mutated DNA strand is corrected. MutL alpha (MLH1-PMS2) physically interacts with the clamp loader subunit of DNA polymerase III, suggesting that it may play a role in recruiting DNA polymerase III to sites of MMR. It is also involved in DNA damage signaling, a process that induces cell cycle arrest and, in cases of major DNA damage, can lead to apoptosis. MLH1 also heterodimerizes with MLH3 to form MutL gamma, which plays a role in meiosis. The "canonical" human MLH1 amino acid sequence is represented by:

[0379] >sp|P40692|MLH1_HUMAN DNA mismatch repair protein Mlh1 OS=Homo sapiens OX=9606 GN=MLH1 PE=1 SV=1 MSFVAGVIRRLDETVVNRIAAGEVIQRPANAIKEMIENCLDAKSTSIQVIVKEGGLKLIQ IQDNGTGIRKEDLDIVCERFTTSKLQSFEDLASISTYGFRGEALASISHVAHVTITTKTA DGKCAYRASYSDGKLKAPPKPCAGNQGTQITVEDLFYNIATRRKALKNPSEEYGKILEVVGRYSVHNAGISFSVKKQGETVADVRTLPNASTVDNIRSIFGNAVSRELIEIGCEDKTLAF KMNGYISNANYSVKCCIFLLFINHRLVESTSLRKAIETVYAAYLPKNTHPFLYLSLEISP QNVDVNVHPTKHEVHFLHEESILERVQQHIESKLLGSNSSRMYFTQTLLPGLAGPSGEMVKSTTSLTSSSTSGSSDKVYAHQMVRTDSREQKLDAFLQPLSKPLSSQPQAIVTEDKTDIS SGRARQQDEEMLELPAPAEVAAKNQSLEGDTTKGTSEMSEKRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVFERC (SEQ ID NO: 9)

[0380] MLH1 may also include other human isoforms, including P40692-2, which differs from the canonical sequence in that residues 1-241 of the canonical sequence are missing:

[0381] >sp|P40692-2|MLH1_HUMAN Isoform 2 of DNA mismatch repair protein Mlh1 OS=Homo sapiens OX=9606 GN=MLH1 MNGYISNANYSVKCCIFLLFINHRLVESTSLRKAIETVYAAYLPKNTHPFLYLSLEISPQ NVDVNVHPTKHEVHFLHEESILERVQQHIESKLLGSNSSRMYFTQTLLPGLAGPSGEMVKSTTSLTSSSTSGSSDKVYAHQMVRTDSREQKLDAFLQPLSKPLSSQPQAIVTEDKTDISS GRARQQDEEMLELPAPAEVAAKNQSLEGDTTKGTSEMSEKRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGHEVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLSGQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVFERC (SEQ ID NO: 10)

[0382] MLH1 may also include a third known isoform, known as P40692-3, which differs from the canonical sequence in that residues 1-101 (of MSFVAGVIRR...ASISTYGFRG (SEQ ID NO: 9)) are replaced with MAF:

[0383] >sp|P40692-3|MLH1_HUMAN Isoform 3 of DNA mismatch repair protein Mlh1 OS=Homo sapiens OX=9606 GN=MLH1 MAFEALASISHVAHVTITTKTADGKCAYRASYSDGKLKAPPKPCAGNQGTQITVEDLFYNIATRRKALKNPSEEYGKILEVVGRYSVHNAGISFSVKKQGETVADVRTLPNASTVDNIRSIFGNAVSRELIEIGCEDKTLAFKMNGYISNANYSVKCCIFLLFINHRLVESTSLRKAIET VYAAYLPKNTHPFLYLSLEISPQNVDVNVHPTKHEVHFLHEESILERVQQHIESKLLGSN SSRMYFTQTLLPGLAGPSGEMVKSTTSLTSSSTSGSSDKVYAHQMVRTDSREQKLDAFLQPLSKPLSSQPQAIVTEDKTDISSGRARQQDEEMLELPAPAEVAAKNQSLEGDTTKGTSEMSEKRGPTSSNPRKRHREDSDVEMVEDDSRKEMTAACTPRRRIINLTSVLSLQEEINEQGH EVLREMLHNHSFVGCVNPQWALAQHQTKLYLLNTTKLSEELFYQILIYDFANFGVLRLSEPAPLFDLAMLALDSPESGWTEEDGPKEGLAEYIVEFLKKKAEMLADYFSLEIDEEGNLIGLPLLIDNYVPPLEGLPIFILRLATEVNWDEEKECFESLSKECAMFYSIRKQYISEESTLS GQQSEVPGSIPNSWKWTVEHIVYKALRSHILPPKHFTEDGNILQLANLPDLYKVFERC (SEQ ID NO: 12).

[0384] The present disclosure contemplates that inhibitors of any of the following proteins may be delivered using the PE-VLPs described herein to inhibit the MMR pathway during prime editing. In addition, such exemplary proteins may also be used to engineer or create dominant-negative variants, which may be used as a type of inhibitor when administered in an amount effective to block, inactivate, or inhibit MMR. Without being bound by theory, it is believed that MLH1 dominant-negative mutants may saturate MutS binding. Exemplary MLH1 proteins include the following amino acid sequences, or amino acid sequences having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or up to 100% sequence identity to any of the following sequences: [Table 12-1] [Table 12-2]

[0385] The PE-VLPs described herein may also be used to deliver MLH1 mutant or truncated variants. In some embodiments, mutant and truncated variants of the human MLH1 wild-type protein are utilized.

[0386] In one aspect, a truncated variant of human MLH1 is delivered using the PE-VLPs of the present disclosure. In some embodiments, amino acids 754-756 of the wild-type human MLH1 protein are truncated (Δ754-756, hereinafter referred to as MLH1dn). In some embodiments, a truncated variant of human MLH1 comprising only the N-terminal domain (amino acids 1-335) (hereinafter referred to as MLH1dn) is delivered. NTD In various embodiments, the following MLH1 variants are provided in the present disclosure: [Table 13-1] [Table 13-2] [Table 13-3] [Table 13-4]

[0387] In yet another aspect, the present disclosure contemplates delivery of an inhibitor of MLH1 using the PE-VLPs described herein. In various embodiments, the inhibitor can be a small molecule inhibitor. In other embodiments, the inhibitor can be an anti-MLH1 antibody (e.g., a neutralizing antibody that inactivates MLH1). In yet other embodiments, the inhibitor can be a dominant-negative mutant of MLH1. In still other embodiments, the inhibitor can be targeted at the transcriptional level of MLH1 (e.g., an siRNA or other nucleic acid agent that knocks down the level of the transcript encoding MLH1).

[0388] In yet another aspect, the present disclosure provides a method for prime editing, which prevents the modification introduced into a target nucleic acid molecule from being corrected by the MMR pathway without the need to provide an inhibitor of the MMR pathway. A pegRNA designed to have consecutive nucleotide mismatches compared to the target site on the target nucleic acid (e.g., a pegRNA with three or more consecutive mismatched nucleotides) can avoid correction by the MMR pathway and can be delivered using the PE-VLP described herein, resulting in increased prime editing efficiency and / or reduced frequency of indel formation compared to the introduction of a single nucleotide mismatch using prime editing. In addition, insertions and deletions of 10 nucleotides or more in length introduced by prime editing can also avoid correction by the MMR pathway, resulting in increased prime editing efficiency and / or reduced frequency of indel formation compared to the introduction of insertions or deletions of less than 10 nucleotides in length using prime editing.

[0389] Thus, in one aspect, the present disclosure provides a method for editing a nucleic acid molecule by prime editing, comprising using a PE-VLP described herein to deliver a prime editor and a pegRNA comprising a DNA synthesis template on its extension arm that comprises three or more consecutive nucleotide mismatches to a target site on the nucleic acid molecule. At least one of the consecutive nucleotide mismatches results in a change in the amino acid sequence of a protein expressed from the nucleic acid molecule. In some embodiments, one or more of the consecutive nucleotide mismatches results in a change in the amino acid sequence of a protein expressed from the nucleic acid molecule. Meanwhile, at least one of the remaining nucleotide mismatches (i.e., those that do not result in a change in the amino acid sequence of a protein expressed from the nucleic acid molecule) is a silent mutation. A silent mutation may be present in the coding region of a target nucleic acid molecule or in a non-coding region of a target nucleic acid molecule. When a silent mutation is present in a coding region, it introduces one or more alternative codons into the nucleic acid molecule that encode the same amino acid as in the unedited nucleic acid molecule. Alternatively, if the silent mutation is in a non-coding region, the silent mutation may be in a region of the nucleic acid molecule that does not affect splicing, gene regulation, RNA lifespan, or other biological properties of the target site on the nucleic acid molecule.

[0390] Any number of three or more consecutive nucleotide mismatches can be used to realize the advantage of avoiding correction by the MMR pathway.In some embodiments, the DNA synthesis template of the extension arm on the pegRNA comprises 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive nucleotide mismatches compared to the endogenous sequence of the target site in the nucleic acid molecule edited by prime editing.In some embodiments, the DNA synthesis template of the extension arm on the pegRNA comprises 3, 4, or 5 consecutive nucleotide mismatches compared to the endogenous sequence of the target site in the nucleic acid molecule edited by prime editing. In some embodiments, the DNA synthesis template of the extension arm on the pegRNA contains 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleotide mismatches compared to the endogenous sequence of the target site in the nucleic acid molecule edited by prime editing. In some embodiments, the DNA synthesis template of the extension arm on the pegRNA contains 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more consecutive nucleotide mismatches compared to the target site on the nucleic acid molecule.

[0391] In another aspect, the present disclosure provides a method for editing nucleic acid molecules by prime editing, comprising using a PE-VLP as described herein to deliver a prime editor and a pegRNA comprising a DNA synthesis template on its extension arm, the template comprising an insertion or deletion of 10 or more nucleotides at a target site on a nucleic acid molecule. When introduced by prime editing, insertions and deletions of 10 or more nucleotides in length can avoid correction by the MMR pathway, and thus can benefit from the inhibition of the MMR pathway without the need to provide an MMR inhibitor. Any insertion or deletion of more than 10 nucleotides in length can be used to realize the advantage of naturally avoiding correction by the MMR pathway. In some embodiments, the DNA synthesis template comprises an insertion or deletion of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides compared to the endogenous sequence at the target site of the nucleic acid molecule edited by prime editing. In some embodiments, DNA synthesis template comprises the insertion or deletion of 11 or more nucleotides, 12 or more nucleotides, 13 or more nucleotides, 14 or more nucleotides, 15 or more nucleotides, 16 or more nucleotides, 17 or more nucleotides, 18 or more nucleotides, 19 or more nucleotides, 20 or more nucleotides, 21 or more nucleotides, 22 or more nucleotides, 23 or more nucleotides, 24 or more nucleotides, or 25 or more nucleotides compared to the target site on nucleic acid molecule.In some embodiments, DNA synthesis template comprises the insertion or deletion of 15 or more nucleotides compared to the target site on nucleic acid molecule.

[0392] PEG-RNAs The PE-VLP-delivered primed editing system described herein contemplates the use of any suitable PEgRNA.

[0393] PEG-RNA constructs In some embodiments, extended guide RNAs are used in the prime editing system delivered using the PE-VLPs disclosed herein, whereby a conventional guide RNA includes a ~20nt protospacer sequence and a gRNA core region that binds to the napDNAbp. In some embodiments, the guide RNA includes an extended RNA segment at the 5' end, i.e., a 5' extension. In some embodiments, the 5' extension includes a reverse transcription template sequence, a reverse transcription primer binding site, and an optional 5-20 nucleotide linker sequence. The RT primer binding site hybridizes to the free 3' end formed after nicking in the non-target strand of the R-loop, thereby priming reverse transcriptase to polymerize DNA in the 5→3' direction.

[0394] In another embodiment, extended guide RNAs usable in a prime editing system are used in the methods and compositions disclosed herein, where a conventional guide RNA includes a ~20nt protospacer sequence and a gRNA core and binds to the napDNAbp. In some embodiments, the guide RNA includes an extended RNA segment at its 3' end, i.e., a 3' extension. In some embodiments, the 3' extension includes a reverse transcription template sequence and a reverse transcription primer binding site. The RT primer binding site hybridizes to the free 3' end formed after nicking in the non-target strand of the R-loop, thereby priming reverse transcriptase to polymerize DNA in the 5→3' direction.

[0395] In another embodiment, extended guide RNAs usable in a prime editing system are used in the methods and compositions disclosed herein, where a conventional guide RNA includes a ~20nt protospacer sequence and a gRNA core and binds to the napDNAbp. In some embodiments, the guide RNA includes an extended RNA segment at an intermolecular position within the gRNA core, i.e., an intramolecular extension. In some embodiments, the intramolecular extension includes a reverse transcription template sequence and a reverse transcription primer binding site. The RT primer binding site hybridizes to the free 3' end formed after nicking in the non-target strand of the R-loop, thereby priming reverse transcriptase to polymerize DNA in the 5→3' direction.

[0396] In one embodiment, the position of the intermolecular RNA extension is not within the protospacer sequence of the guide RNA. In another embodiment, the position of the intermolecular RNA extension is within the gRNA core. In yet another embodiment, the position of the intermolecular RNA extension is anywhere within the guide RNA molecule except within the protospacer sequence, or at a position that disrupts the protospacer sequence. In one embodiment, the intermolecular RNA extension is inserted downstream from the 3' end of the protospacer sequence. In another embodiment, the intermolecular RNA extension is inserted at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides downstream from the 3' end of the protospacer sequence.

[0397] In other embodiments, an intermolecular RNA extension is inserted into the gRNA, which refers to a portion of the guide RNA that corresponds to or includes the tracrRNA, and binds and / or interacts with the Cas9 protein or its equivalent (i.e., a different napDNAbp). Preferably, the insertion of the intermolecular RNA extension does not disrupt, or only minimally disrupts, the interaction between the tracrRNA portion and the napDNAbp.

[0398] The length of the RNA extension (comprising at least the RT template and the primer binding site) can be any useful length. In various embodiments, the RNA extension is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0399] The RT template sequence can also be of any suitable length. For example, the RT template sequence can be at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0400] In still other embodiments, the reverse transcription primer binding site sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0401] In other embodiments, any linker or spacer sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.

[0402] The RT template sequence, in some embodiments, encodes a single-stranded DNA molecule that is homologous to the non-target strand (and thus complementary to the corresponding portion of the target strand) but that contains one or more nucleotide changes, which may include one or more single-base nucleotide changes, one or more deletions, and / or one or more insertions.

[0403] The synthesized single-stranded DNA product of the RT template sequence is homologous to the non-target strand and contains one or more nucleotide changes. The single-stranded DNA product of the RT template sequence hybridizes in equilibrium with the complementary target strand sequence, thereby displacing the homologous endogenous target strand sequence. In some embodiments, the displaced endogenous strand is sometimes referred to as a 5' endogenous DNA flap species. This 5' endogenous DNA flap species can be removed by a 5' flap endonuclease (e.g., FEN1), and the single-stranded DNA product hybridized to the endogenous target strand can be ligated, thereby creating a mismatch between the endogenous sequence and the newly synthesized strand. The mismatch can be resolved by the cell's innate DNA repair and / or replication processes.

[0404] In various embodiments, the nucleotide sequence of the RT template sequence is displaced as a 5' flap species and corresponds to the nucleotide sequence of the non-target strand that overlaps the site to be edited.

[0405] In various embodiments of the extended guide RNA, the reverse transcription template sequence may encode a single-stranded DNA flap complementary to the endogenous DNA sequence adjacent to the nick site, where the single-stranded DNA flap contains the desired nucleotide change. The single-stranded DNA flap may displace the endogenous single-stranded DNA at the nick site. The displaced endogenous single-stranded DNA at the nick site may have a 5' end, forming an endogenous flap that can be excised by the cell. In various embodiments, excision of the 5'-end endogenous flap may help guide product formation, as removing the 5'-end endogenous flap promotes hybridization of the single-stranded 3' DNA flap to the corresponding complementary DNA strand and incorporation or assimilation of the desired nucleotide change carried by the single-stranded 3' DNA flap into the target DNA.

[0406] In various embodiments of extended guide RNAs, cellular repair of the single-stranded DNA flap results in the introduction of the desired nucleotide change, thereby forming the desired product.

[0407] In still other embodiments, the desired nucleotide change is introduced into an editing window that is between about -5 and +5 of the nick site, or between about -10 and +10 of the nick site, or between about -20 and +20 of the nick site, or between about -30 and +30 of the nick site, or between about -40 and +40 of the nick site, or between about -50 and +50 of the nick site, or between about -60 and +60 of the nick site, or between about -70 and +70 of the nick site, or between about -80 and +80 of the nick site, or between about -90 and +90 of the nick site, or between about -100 and +100 of the nick site, or between about -200 and +200 of the nick site.

[0408] In other embodiments, the desired nucleotide changes are located at about +1 to +2 from the nick site, or about +1 to +3, +1 to +4, +1 to +5, +1 to +6, +1 to +7, +1 to +8, +1 to +9, +1 to +10, +1 to +11, +1 to +12, +1 to +13, +1 to +14, +1 to +15, +1 to +16, +1 to +17, +1 to +18, +1 to +19, +1 to +20, +1 to +21, +1 to +22, +1 to +23, +1 to +24, +1 to +25, +1 to +26, +1 to +27, +1 to +28, +1 to +29, +1 to +30, +1 ~+31, +1~+32, +1~+33, +1~+34, +1~+35, +1~+36, +1~+37, +1~+38, +1~+39, +1~+40, +1~+41, +1~+42, +1~+43, +1~+44, +1~+45, +1~+46, +1~+47, +1~+48, +1~+49, +1~+50, +1~+51, +1~+52, +1~+53, +1~+54, +1~+55, +1~+56, +1~+57, +1~+58, +1~+59, +1~+60, +1~+61, +1~+62, +1~+63, +1~+64, +1~ +65, +1~+66, +1~+67, +1~+68, +1~+69, +1~+70, +1~+71, +1~+72, +1~+73, +1~+74, +1~+75, +1~+76, +1~+77, +1~+78, +1~+79, +1~+80, +1~+81, +1~+82, +1~+83, +1~+84, +1~+85, +1~+86, +1~+87, +1~+88, +1~+89, +1~+90, +1~+90, +1~+91, +1~+92, +1~+93, +1~+94, +1~+95, +1~+96, +1~+97, +1~+ 98, +1~+99, +1~+100, +1~+101, +1~+102, +1~+103, +1~+104, +1~+105, +1~+106, +1~+107, +1~+108, +1~+109, +1~+110, +1~+111, +1~+112, +1~+113, +1~+114, +1~+115, +1~+116, +1~+117, +1~+118, +1~+119, +1~+120, +1~+121, +1~+122, +1~+123, +1~+124, or +1~+125,

[0409] In still other embodiments, the desired nucleotide change is about +1 to +2 from the nick site, or about +1 to +5, +1 to +10, +1 to +15, +1 to +20, +1 to +25, +1 to +30, +1 to +35, +1 to +40, +1 to +45, +1 to +50, +1 to +55, +1 to +100, +1 to +105, +1 to +110, +1 to +115, It is introduced into the editing window between +1~+120, +1~+125, +1~+130, +1~+135, +1~+140, +1~+145, +1~+150, +1~+155, +1~+160, +1~+165, +1~+170, +1~+175, +1~+180, +1~+185, +1~+190, +1~+195, or +1~+200.

[0410] In various aspects, extended guide RNA is a modified version of guide RNA.Guide RNA can be naturally occurring, can be expressed from encoding nucleic acid, or can be chemically synthesized.The method of obtaining or otherwise synthesizing guide RNA, and the method of determining the appropriate sequence of guide RNA, including the protospacer sequence that interacts and hybridizes with the target strand of the target site of genome of interest, are well known in the art.

[0411] In various embodiments, the specific design aspects of the guide RNA sequence will depend on the nucleotide sequence of the genomic target site of interest (i.e., the desired site to be edited), among other factors such as the location of the PAM sequence, the percent G / C content in the target sequence, the extent of regions of microhomology, secondary structure, etc., and the type of napDNAbp (e.g., Cas9 protein) present in the prime editing system utilized in the methods and compositions described herein.

[0412] Generally, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a napDNAbp (e.g., a Cas9, Cas9 homolog, or Cas9 variant) to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is about or greater than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more when optimally aligned using a suitable alignment algorithm. Optimal alignment may be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.)), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about or about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more nucleotides in length.

[0413] In some embodiments, the guide sequence is about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. The ability of a guide sequence to direct sequence-specific binding of a prime editor to a target sequence may be assessed by any suitable assay. For example, a prime editor component comprising the guide sequence to be tested is provided to a host cell harboring the corresponding target sequence, such as by transfection with a vector encoding a prime editor component disclosed herein, followed by assessment of preferential cleavage within the target sequence, such as by a surveyor assay as described herein. Similarly, cleavage of a target polynucleotide sequence may be assessed in a test tube by providing a prime editor component comprising the target sequence, the guide sequence to be tested, and a control guide sequence different from the test guide sequence, and comparing the rate of binding or cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible and will occur to those skilled in the art.

[0414] The guide sequence may be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of a cell. Exemplary target sequences include those that are unique within the target genome. For example, in the case of Streptococcus pyogenes Cas9, a unique target sequence within a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNNXGG, where N is A, G, T, or C; and X can be anything. A unique target sequence within a genome may include a Streptococcus pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGG, where N is A, G, T, or C; and X can be anything. For S. thermophilus CRISPR1Cas9, a unique target sequence in the genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXXAGAAW, where NNNNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; and W is A or T). A unique target sequence in the genome may include a S. thermophilus CRISPR1Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXXAGAAW, where NNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; and W is A or T). For Streptococcus pyogenes Cas9, a unique target sequence in the genome may include a Cas9 target site of the form NNNNNNNNNNNNXGGXG (N is A, G, T, or C; and X can be anything). Unique target sequences in a genome may also include Streptococcus pyogenes Cas9 target sites of the form MMMMMMMMMNNNNNNNNNNNXGGXG, where N is A, G, T, or C; and X can be anything. In each of these sequences, "M" can be A, G, T, or C and need not be considered when identifying a sequence as unique.

[0415] In some embodiments, the guide sequence is selected to reduce the degree of secondary structure within the guide sequence. The secondary structure may be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. One example of such an algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example of a folding algorithm is the online web server RNAfold, which uses a centroid structure prediction algorithm developed at the Institute of Theoretical Chemistry, University of Vienna (see, for example, AR Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12): 1151-62). Further algorithms may be found in U.S. Application No. 61 / 836,080, which is incorporated herein by reference.

[0416] In general, a tracr mate sequence encompasses any sequence that has sufficient complementarity with a tracr sequence to promote one or more of the following: (1) excision of the guide sequence flanking the tracr mate sequence in cells containing the corresponding tracr sequence, and (2) formation of a complex at the target sequence, wherein the complex includes the tracr mate sequence hybridized to the tracr sequence. Generally, the degree of complementarity refers to optimal alignment of the tracr mate sequence and the tracr sequence along the length of the shorter of the two sequences. Optimal alignment may be determined by any suitable alignment algorithm and may further take into account secondary structure, such as self-complementarity within either the tracr sequence or the tracr mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and the tracr mate sequence along the length of the shorter of the two sequences when optimally aligned is about or greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more. In some embodiments, the tracr sequence is about or about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and the tracr mate sequence are contained within a single transcript such that hybridization between them produces a transcript with a secondary structure, such as a hairpin. A preferred loop-forming sequence used in the hairpin structure is four nucleotides in length and most preferably has the sequence GAAA. However, longer or shorter loop sequences may be used, and alternative sequences may also be used. The sequence preferably includes a triplet of nucleotides (e.g., AAA) and an additional nucleotide (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAG. In embodiments of the present invention, a transcript or transcribed polynucleotide sequence has at least two or more hairpins. In preferred embodiments, a transcript has two, three, four, or five hairpins. In a further embodiment of the invention, the transcript has at most five hairpins.In some embodiments, the single transcription product further includes a transcription termination sequence; preferably, this is a poly-T sequence, e.g., 6 T nucleotides. Further non-limiting examples of single polynucleotides comprising a guide sequence, a tracr mate sequence, and a tracr sequence are as follows (listed 5' to 3'), where "N" represents bases of the guide sequence, the first block of lowercase letters represents the tracr mate sequence, and the second block of lowercase letters represents the tracr sequence, and the final poly-T sequence represents the transcription terminator: (1)NNNNNNNNGTTTTTGTACTCTCAAGATTTAGAAATAAAATCTTGCAGAAGCTACAAAGATAAGGCTTCATGCCGAAATCAACACCCTGTCATTTTATGGCAGGGTGTTTTCGTTATTTAATTTTTT (SEQ ID NO: 212); (2) NNNNNNNNNNNNNNNNNNNNGTTTTTGTACTCTCAGAAATGCAGAAGCTACAAAGATAAGGCTTCATGCCGAAATCAACACCCTGTCATTTTATGGCAGGGTGTTTTCGTTATTTAATTTTTT (SEQ ID NO: 213); (3) NNNNNNNNNNNNNNNNNNNNNNGTTTTTGTACTCTCAGAAATGCAGAAGCTACAAAGATAAGGCTTCATGCCGAAATCAACACCCTGTCATTTTATGGCAGGGTGTTTTTT (SEQ ID NO: 214); (4) NNNNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 215); (5) NNNNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGTTTTTTT (SEQ ID NO: 216); and (6) NNNNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCATTTTTTTT (SEQ ID NO: 217).

[0417] In some embodiments, sequences (1)-(3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4)-(6) are used in combination with Cas9 from Streptococcus pyogenes. In some embodiments, the tracr sequence is a separate transcript from the transcript containing the tracr mate sequence.

[0418] As disclosed herein, it will be apparent to one of skill in the art that targeting any of the fusion proteins comprising a Cas9 domain and a single-stranded DNA-binding protein to a target site (e.g., a site containing a point mutation to be edited) typically requires co-expression of the fusion protein with a guide RNA (e.g., an sgRNA). As explained in more detail elsewhere herein, the guide RNA typically comprises a tracrRNA framework that enables Cas9 binding and a guide sequence that confers sequence specificity to the Cas9:nucleic acid editing enzyme / domain fusion protein.

[0419] In some embodiments, the guide RNA comprises the structure 5'-[guide sequence]-GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUU-3' (SEQ ID NO: 218), where the guide sequence comprises a sequence complementary to the target sequence. Guide sequences are typically 20 nucleotides in length. The sequence of a suitable guide RNA for targeting a Cas9:nucleic acid editing enzyme / domain fusion protein to a specific genomic target site will be apparent to one of skill in the art based on this disclosure. Such suitable guide RNA sequences typically comprise a guide sequence complementary to a nucleic acid sequence within 50 nucleotides upstream or downstream of the target nucleotide to be edited. Several exemplary guide RNA sequences suitable for targeting any of the provided fusion proteins to a specific target sequence are provided herein. Additional guide sequences are known in the art and can be used with the prime editors utilized in the methods and compositions described herein.

[0420] In some embodiments, a PEGRNA comprises three major components arranged in a 5' to 3' direction: a spacer, a gRNA core, and an extension arm at the 3' end. The extension arm may be further divided into the following structural elements in a 5' to 3' direction: a primer binding site (A), an editing template (B), and a homology arm (C). In addition, a PEGRNA may comprise an optional 3'-end modifier region (e1) and an optional 5'-end modifier region (e2). Furthermore, a PEGRNA may comprise a transcription termination signal at the 3' end of the PEGRNA. These structural elements are further defined herein. The depiction of the structure of a PEGRNA is not meant to be limiting and encompasses variations in the placement of elements. For example, the optional sequence modifiers (e1) and (e2) can be located within or between any of the other regions shown and are not limited to being located at the 3' and 5' ends.

[0421] PEG-RNA modification The PEgRNA may also incorporate additional design modifications that alter the properties and / or characteristics of the PEgRNA, which may thereby improve the efficacy of prime editing. In various embodiments, these modifications may fall into one or more of a number of different categories, including, but not limited to: (1) Design to enable efficient expression of functional PEgRNA from a non-polymerase III (pol III) promoter, which would allow expression of longer PEgRNAs without burdensome sequence requirements; (2) modifications to the core, Cas9-bound PEGRNA scaffold that may improve efficacy; (3) Modifications of PEG RNA to improve RT processivity and allow for insertion of longer sequences at the target genomic locus; and (4) Addition of RNA motifs to the 5′ or 3′ end of PEgRNA, which improves PEgRNA stability, enhances RT processivity, prevents PEgRNA misfolding, or recruits additional factors important for genome editing.

[0422] In one embodiment, PEGRNAs can be designed using a Pol III promoter to improve expression of longer PEGRNAs with larger extension arms. sgRNAs are typically expressed from the U6 snRNA promoter. This promoter employs Pol III to express related RNAs and is useful for expressing short RNAs retained in the nucleus. However, Pol III lacks high processivity and cannot express RNAs longer than a few hundred nucleotides at the levels required for efficient genome editing. In addition, Pol III can terminate at a series of Us, potentially limiting the diversity of sequences that can be inserted using PEGRNAs. Other promoters employing Polymerase II (e.g., pCMV) or Polymerase I (e.g., the U1 snRNA promoter) have also been investigated for their ability to express longer sgRNAs. However, these promoters are typically partially transcribed, resulting in an extra spacer sequence 5' in the expressed PEGRNA, which has been shown to significantly reduce Cas9:sgRNA activity in a site-dependent manner. Additionally, while Pol III-transcribed PEG RNAs can simply terminate with a series of 6–7 U, Pol II- or Pol I-transcribed PEG RNAs will require different termination signals. Often, such signals also result in polyadenylation, which results in undesired export of the PEG RNA from the nucleus. Similarly, RNAs expressed from Pol II promoters, such as pCMV, are typically 5'-capped, which also results in nuclear export.

[0423] Previously, Rinn and colleagues screened various expression platforms for producing long noncoding RNA (lncRNA)-tagged sgRNAs. These platforms were expressed from pCMV and contained RNAs terminating in the ENE element from the MALAT1 ncRNA from human, the PAN ENE element from KSHV, or the 3' box from U1 snRNA. Notably, the MALAT1 ncRNA and the PAN ENE form a triple helix that protects the poly(A) tail. These structures may also enhance RNA stability. It is contemplated that these expression systems will also enable the expression of longer PEG RNAs.

[0424] In addition, a series of strategies have been designed to cleave the portion of the Pol II promoter that will be transcribed as part of the PEG RNA, either by adding self-cleaving ribozymes such as hammerhead, pistol, hatchet, hairpin, VS, twister, or twister sister ribozymes, or other self-cleaving elements that process the transcribed guide, or by adding hairpins that are recognized by Csy4 and also lead to guide processing. It is also hypothesized that incorporating multiple ENE motifs, as previously demonstrated with KSHV PAN RNA and elements, may lead to improved expression and stability of the PEG RNA. It is also predicted that circularizing the PEG RNA in the form of a circular intron RNA (ciRNA) will lead to enhanced RNA expression, stability, and nuclear localization.

[0425] In various embodiments, the PEGRNA may include various of the above elements, as exemplified in the following sequences:

[0426] Non-limiting Example 1 - PEGRNA Expression Platform Consisting of pCMV, Csy4 Hairpin, PEGRNA, and MALAT1 ENE TAGTTATTAATAGTAATCAAATTACGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAA GTGTATCATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGGTTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTGACGTCAATG GGAGTTTGTTTGGCACCAAAATCAACGGGACTTTCCAAAAATGTCGTAACAACTCCGCCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGTCAGATCGTTCACTGCCGTATAGGCAGGCCCAGACTGAGCACGTGAGTTTTAGACCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGATCGGTCCTCTGCCATCAAAGCGTGCTCAGTCTTTTAGGGTCATGAAGGTTTTTTCTTTCCTGAGAAAAACAACACACGTATTGTTTCTCAGGTTTGCTTTTGCCTTTTTCTAGCTTAAAAAAAAAAAAAAAAGCAAAAGATGCTGGTGTTGGCACTACTCTTTTCCAGGACGGGGTTCAAATCCCTGCGGGCGTCTTTGCTTTGACT (sequence number 219)

[0427] Non-limiting Example 2—PEG-RNA Expression Platform Consisting of pCMV, Csy4 Hairpin, PEG-RNA, and PAN ENE TAGTTATTAATAGTAATCAAATTACGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATCAT ATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGGTTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACCA AAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGTCAGATCGTTCACTGCCGTATAGGCAGGCCCAGACTGAGCACGTGAGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCGTGCTCAGTTGTTTGTTTGGCTGGGTTTTTCCTTGTCGCCACCGGACACCTCCAGTGACCACGGCCAAGGTTTTTATCCCCAGTGTATATTGGAAAAACATGTTATACTTTTGACAATTTAACGTGCCTAGAGCTCAAATTAAACTAATACCATAACGTAATGGAACTTACAACATAAATAAGGTCAATGTTTAATCCATAAA

[0428] Non-limiting example 3 - PEGRNA expression platform consisting of pCMV, Csy4 hairpin, PEGRNA, and 3xPAN ENE

[0429] Non-limiting example 4 - PEG-RNA expression platform consisting of pCMV, Csy4 hairpin, PEG-RNA, and 3' box TAGTTATTAATAGTAATCAATTACGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTA AACTGCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGGTTTTGGCAGTACATCAATGGGCGTGGATAG CGGTTTGACTCACGGGATTTCCAAGTCTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGTCAGATCGTTCACTGCCGTATAG GCAGGGCCCAGACTGAGCACGTGAGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCGTGCTCAGTCTGTTTGTTTCAAAAAGTAGACTGTACGCTAAGGGTCATATCTTTTTTTGTTTGGTTTGTGTCTTGGTTGGCGTCTTAAA (Sequence number 222)

[0430] PEgRNA expression platform consisting of non-limiting example 5 - pU1, Csy4 hairpin, PEgRNA, and 3' box CTAAGGACCAGCTTCTTTGGGAGAGAACAGACGCAGGGGCGGGAGGGAAAAAGGGAGAGGCAGACGTCACTTCCCCTTGGCGGCTCTGGCAGCAGATTGGTCGGTTGAGTGGCAGAAAGGCAGACGGGGACTGGGCAAGGCACTGTCGGTGACATCACGGACAGGGCGACTTCTATGTAGATGAGGCAGCGCAGAGGCTGCTGCTTCGCCACTTGCTGCTTCACCACGAAGGAGTTCCCGTGCCCTGGGAGCGGGTTCAGGACCGCTGATCGGAAGTGAGAATCCCAGCTGTGTGTCAGGGCTGGAAAGGGCTCGGGAGTGCGCGGGGCAAGTGACCGTGTGTGTAAAGAGTGAGGCGTATGAGGCTGTGTCGGGGCAGAGGCCCAAGATCTCAGTTCACTGCCGTATAGGCAGGGCCCAGACTGAGCACGTGAGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCGTGCTCAGTCTGTTTCAGCAAGTTCAGAGAAATCTGAACTTGCTGGATTTTTGGAGCAGGGAGATGGAATAGGAGCTTGCTCCGTCCACTCCACGCATCGACCTGGTATTGCAGTACCTCCAGGAACGGTGCACCCACTTTCTGGAGTTTCAAAAGTAGACTGTACGCTAAGGGTCATATCTTTTTTTGTTTGGTTTGTGTCTTGGTTGGCGTCTTAAA (SEQ ID NO: 223).

[0431] In various other embodiments, PEGRNAs may be improved by introducing modifications to the scaffold or core sequence. The core Cas9-bound PEGRNA scaffold could conceivably be improved to enhance PE activity. Several such approaches have already been demonstrated. As an illustration, the first pairing element of the scaffold (P1) contains a GTTTT-AAAAC (SEQ ID NO: 231) pairing element. Such a run of Ts has been shown to result in pol III pausing and premature termination of RNA transcripts. Rational mutation of one of the TA pairs in this portion of P1 to a GC pair has been shown to enhance sgRNA activity, suggesting that this approach is also feasible for PEGRNAs. Additionally, increasing the length of P1 enhances sgRNA folding, leading to improved activity and suggesting another means of modifying PEGRNA activity. Examples of modifications to the core include:

[0432] PEG-RNA containing a 6-nt extension to P1 GGCCCAGACTGAGCACGTGAGTTTTAGAGCTAGCTCATGAAAATGAGCTAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCGTGCTCAGTCTGTTTTTTT (SEQ ID NO: 224)

[0433] PEgRNA containing a TA to GC mutation in P1 GGCCCAGACTGAGCACGTGAGTTTGAGAGCTAGAAATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCGTGCTCAGTCTGTTTTTTT (SEQ ID NO: 225)

[0434] In various other embodiments, the PEGRNA may be modified in the editing template region. As the size of the insertion templated by the PEGRNA increases, it becomes more likely to be degraded by endonucleases, undergo spontaneous hydrolysis, or fold into a secondary structure that cannot be reverse-transcribed by RT, disrupting the folding of the PEGRNA scaffold and subsequent Cas9-RT binding. Therefore, to affect large-scale insertions, such as the insertion of entire genes, modifications to the PEGRNA template are likely required. Some strategies for doing so include incorporating modified nucleotides into synthetic or semi-synthetic PEGRNA, which make the RNA more resistant to degradation or hydrolysis or less likely to adopt inhibitory secondary structures. Such modifications may include 8-aza-7-deazaguanosine, which reduces RNA secondary structure in G-rich sequences; locked nucleic acids (LNA), which reduce degradation and enhance certain RNA secondary structures; and 2'-O-methyl, 2'-fluoro, or 2'-O-methoxyethoxy modifications, which enhance RNA stability. Such modifications can be incorporated into other parts of the PEgRNA to enhance stability and activity. Alternatively, or in addition, the PEgRNA template can be designed to both encode the desired protein product and be prone to adopting a simple secondary structure that can be unfolded by RT. Such a simple structure is less likely to form a more complex structure, which would act as a thermodynamic sink and prevent reverse transcription. Ultimately, the template can be separated into two separate PEgRNAs. In such a design, PE will be used to initiate transcription and also to recruit the separate template RNAs to the target site via the RNA recognition element of the PEgRNA itself, such as an RNA-binding protein fused to Cas9 or an MS2 aptamer. RT can either directly bind to the separate template RNAs or initiate reverse transcription on the original PEgRNA before switching to the second template.Such an approach may enable long insertions both by preventing the misfolding of PEgRNA that accompanies the addition of long templates and by not requiring dissociation of Cas9 from the genome for long insertions to occur, which may potentially inhibit PE-based long insertions.

[0435] In yet other embodiments, PEG RNAs may be modified by introducing additional RNA motifs at the 5' and 3' ends of the PEG RNA, or even at positions in between (e.g., the gRNA core region or spacer). Several such motifs, such as the PAN ENE from KSHV and the ENE from MALAT1, were discussed above as possible means of terminating expression of long PEG RNAs from non-pol III promoters. These elements form RNA triple helices that envelop the polyA tail, resulting in nuclear retention. However, by forming complex structures at the 3' end of the PEG RNA that occlude the terminal nucleotides, these structures may also help prevent exonuclease-mediated degradation of the PEG RNA.

[0436] Other structural elements inserted at the 3' end can also enhance RNA stability, even though they do not allow termination from non-pol III promoters. Such motifs can include hairpins and RNA quadruplexes that occlude the 3' end, or self-cleaving ribozymes such as HDV, which result in the formation of a 2'-3'-cyclic phosphate at the 3' end and may also make the PEG RNA less susceptible to exonuclease degradation. Inducing the PEG RNA to circularize and form ciRNA through incomplete splicing can also increase the stability of the PEG RNA, resulting in its retention in the nucleus.

[0437] Additional RNA motifs can improve RT processivity or enhance PEG-RNA activity by enhancing RT binding to the DNA-RNA duplex. Adding native sequences bound by RT in the cognate retroviral genome can enhance RT activity. This can include native primer binding sites (PBSs), polypurine tracts (PPTs), or kissing loops involved in retroviral genome dimerization and transcription initiation.

[0438] The addition of dimerization motifs, such as kissing loops or GNRA tetraloop / tetraloop acceptor pairs, at the 5' and 3' ends of the PEgRNA may also result in effective circularization of the PEgRNA, improving its stability. Additionally, it is conceivable that adding these motifs may allow for physical separation of the PEgRNA spacer and primer, preventing spacer blockage that would interfere with PE activity. Short 5' or 3' extensions to the PEgRNA that form small toehold hairpins in the spacer region or along the primer binding site may also favorably compete for annealing within complementary regions along the length of the PEgRNA (e.g., potential interactions between the spacer and primer binding site). Finally, kissing loops may also be used to recruit other template RNAs to genomic sites and enable RT activity to be exchanged from one RNA to another. Numerous secondary RNA structures may be engineered into any region of the PEgRNA, including the terminal portions of the extension arms (i.e., e1 and e2), as shown. Examples of modifications include, but are not limited to:

[0439] PEgRNA-HDV fusion GGCCCAGACTGAGCACGTGAGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCGTGCTCAGTCTGGGCCGGCATGGTCCCAGCCTCCTCGCTGGCGCCGGCTGGGCAACATGCTTCGGCATGGCGAATGGGACTTTTTTT (SEQ ID NO: 226)

[0440] PEGRNA-MMLV kissing loop GGTGGGAGACGTCCCACCGGCCCAGACTGAGCACGTGAGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCTTCGACCGTGCTCAGTCTGGTGGGAGACGTCCCACCTTTTTT (SEQ ID NO: 227)

[0441] PEGRNA-VS ribozyme kissing loop GAGCAGCATGGCGTCGCTGCTCACGGCCCAGACTGAGCACGTGAGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCTTCGACCGTGCTCAGTCTCCATCAGTTGACACCCTGAGGTTTTTTT (SEQ ID NO: 228)

[0442] PEG-RNA-GNRA tetraloop / tetraloop receptor GCAGACCTAAGTGGUGACATATGGTCTGGGCCCAGACTGAGCACGTGAGTTTTAGAGCTAUACGTAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTUACGAAGTGGGACCGAGTCGGTCCTCTGCCATCAAAGCTTCGACCGTGCTCAGTCTGCATGCGATTAGAAATAATCGCATGTTTTTTT (SEQ ID NO: 229)

[0443] PEG RNA template switching secondary RNA-HDV fusion TCTGCCATCAAAGCTGCGACCGTGCTCAGTCTGGTGGGAGACGTCCCACCGGCCGGCATGGTCCCAGCCTCCTCGCTGGCGCCGGCTGGGCAACATGCTTCGGCATGGCGAATGGGACTTTTTTT (SEQ ID NO: 230)

[0444] The PEgRNA scaffold can be further improved through directed evolution, similar to how SpCas9 and prime editors (PEs) have been improved. Directed evolution could enhance PEgRNA recognition by Cas9 or evolved Cas9 variants. In addition, different PEgRNA scaffold sequences are likely optimal at different genomic loci, either enhancing PE activity at the site of interest, reducing off-target activity, or both. Finally, evolution of the PEgRNA scaffold with additional RNA motifs will almost certainly improve the activity of the fusion PEgRNA compared to the unevolved fusion RNA. As an illustration, evolution of an allosteric ribozyme consisting of a c-di-GMP-I aptamer and a hammerhead ribozyme led to dramatic improvements in activity, suggesting that evolution could also improve the activity of hammerhead-PEgRNA fusions. Additionally, while Cas9 currently does not generally tolerate 5' extensions of sgRNAs, directed evolution could generate enabling mutations that alleviate this intolerance, allowing for the utilization of additional RNA motifs. The present disclosure contemplates any such manner of further improving the effectiveness of the prime editing system utilized in the methods and compositions disclosed herein.

[0445] In various embodiments, it may be advantageous to limit the occurrence of consecutive T sequences from the extension arms, as a series of consecutive Ts may limit the ability of the PEgRNA to be transcribed. For example, stretches of at least three consecutive Ts, at least four consecutive Ts, at least five consecutive Ts, at least six consecutive Ts, at least seven consecutive Ts, at least eight consecutive Ts, at least nine consecutive Ts, at least ten consecutive Ts, at least eleven consecutive Ts, at least 12 consecutive Ts, at least 13 consecutive Ts, at least 14 consecutive Ts, or at least 15 consecutive Ts should be avoided when designing the PEgRNA, or at least removed from the final designed sequence. In one embodiment, avoiding target sites rich in consecutive A:T nucleobase pairs can avoid the inclusion of unnecessary stretches of consecutive Ts in the PEgRNA extension arms.

[0446] Methods for producing PE-VLPs In one aspect, the present disclosure relates to methods for producing eVLPs described herein. In some embodiments, the methods for producing eVLPs currently described include transfecting, transducing, electroporating, or otherwise inserting into a producer cell one or more polynucleotides that together encode all of the components of the eVLP (e.g., any of the polynucleotides described herein or any of the vectors described herein). In some embodiments, the present disclosure provides one or more vectors that include one, two, three, or all four of the polynucleotides provided herein. In some embodiments, each of the first, second, third, and fourth polynucleotides is on a separate vector. In some embodiments, one or more of the first, second, third, and fourth polynucleotides is on the same vector.

[0447] In some embodiments, once the producer cells express the polynucleotides, the various components of the eVLP spontaneously self-assemble within the producer cells. Assembly of the eVLP depends on the multimerization of the gag polyprotein encoded on the polynucleotide, as described above. The gag polyprotein (some of which is fused to a gene editing agent, such as a prime editor) multimerizes at the plasma membrane of the producer cells and is subsequently spontaneously released into the producer cell supernatant. Thus, PE-eVLPs may be produced by transient transfection of producer cells (e.g., Gesicle Producer 293T cells), as described in the examples herein. All polynucleotides required for eVLP production may be transfected into the producer cells simultaneously, or each required polynucleotide may be transfected one at a time. In some embodiments, a single polynucleotide encodes all of the components necessary to produce the eVLPs described herein. After transfection and incubation of the production cells (e.g., for about 2 hours, about 3 hours, about 4 hours, about 5 hours, about 6 hours, about 7 hours, about 8 hours, about 9 hours, about 10 hours, about 15 hours, about 24 hours, about 36 hours, about 48 hours, or more than 48 hours), the production cell supernatant is harvested and eVLPs may be purified therefrom.

[0448] Any cell capable of expressing an exogenous polynucleotide may be used to produce the eVLPs described herein. For example, the present disclosure contemplates the use of any of the cells listed in the "Kits and Cells" section herein, or other cells known in the art capable of expressing an exogenous polynucleotide, for the production of eVLPs.

[0449] Pharmaceutical Composition Another aspect of the present disclosure relates to a pharmaceutical composition comprising any of the PE-VLPs, fusion proteins, and polynucleotide(s) described herein. As used herein, the term "pharmaceutical composition" refers to a composition formulated for pharmaceutical use. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition comprises an additional agent (e.g., a specific delivery, half-life enhancing, or other therapeutic compound).

[0450] As used herein, the term "pharmaceutically acceptable carrier" refers to a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, magnesium talc, calcium or zinc stearate, or stearic acid), or solvent encapsulant, that is involved in carrying or transporting a compound from one site in the body (e.g., a delivery site) to another site (e.g., an organ, tissue, or body part). A pharmaceutically acceptable carrier is "acceptable" in the sense of being compatible with the other ingredients of the formulation and not toxic to the tissues of the subject (e.g., physiologically compatible, sterile, physiological pH, etc.). Examples of materials that can serve as pharmaceutically acceptable carriers include: (1) sugars such as lactose, glucose, and sucrose; (2) starches such as corn starch and potato starch; (3) cellulose and its derivatives such as sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants such as magnesium stearate, sodium lauryl sulfate, and talc; (8) excipients such as cocoa butter and suppository wax; (9) oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (10) glycols such as propylene glycol; (11) (12) polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol (PEG); (13) esters such as ethyl oleate and ethyl laurate; (14) buffers such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer's solution; (19) ethyl alcohol; (20) pH buffer solutions; (21) polyesters, polycarbonates, and / or polyanhydrides; (22) bulking agents such as polypeptides and amino acids; (23) serum components such as serum albumin, HDL, and LDL; (22) C2-C12 alcohols such as ethanol; and (23) other non-toxic compatible substances used in pharmaceutical formulations.Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, fragrances, preservatives, and antioxidants may also be present in the formulation. The terms "excipient," "carrier," "pharmaceutically acceptable carrier," and the like, or the like, are used interchangeably herein.

[0451] In some embodiments, the pharmaceutical composition is formulated for delivery to a subject (e.g., for gene editing). Suitable routes for administering the pharmaceutical compositions described herein include, but are not limited to, topical administration, subcutaneous administration, transdermal administration, intradermal administration, intralesional administration, intraarticular administration, intraperitoneal administration, intravesical administration, transmucosal administration, gingival administration, intradental administration, intracochlear administration, intratympanic administration, intravisceral administration, epidural administration, intrathecal administration, intramuscular administration, intravenous administration, intravascular administration, intraosseous administration, periocular administration, intratumoral administration, intracerebral administration, and intraventricular administration.

[0452] In some embodiments, the pharmaceutical compositions described herein are administered locally to the site of disease (e.g., a tumor site). In some embodiments, the pharmaceutical compositions described herein are administered to a subject by injection, using a catheter, using a suppository, or using an implant, which is a porous, non-porous, or gelatinous material that includes a membrane, such as a silastic membrane, or a fiber.

[0453] In other embodiments, the pharmaceutical compositions described herein are delivered in controlled release systems.In one embodiment, pumps can be used (see, for example, Langer, 1990, Science 249:1527-1533; Sefton, 1989, CRC Crit. Ref. Biomed. Eng. 14:201; Buchwald et al., 1980, Surgery 88:507; Saudek et al., 1989, N. Engl. J. Med. 321:574).In another embodi...

Claims

1. (1) a protein comprising a group-specific antigen (gag) protein linked to a viral protease; and (2) Fusion protein wherein the protein and the fusion protein are encapsulated by a lipid membrane and a viral envelope glycoprotein, the protein is fused to a first coiled-coil peptide, the fusion protein is fused to a second coiled-coil peptide, and the fusion protein further comprises a nucleic acid programmable DNA binding protein (napDNAbp) and / or a domain having DNA polymerase activity.

2. The VLP of claim 1 , wherein the second coiled-coil peptide is present at the N-terminus, C-terminus, or internal position of the fusion protein.

3. 2. The VLP of claim 1, wherein one of the first coiled-coil peptide or the second coiled-coil peptide comprises a P3 peptide, and the other comprises a P4 peptide.

4. 4. The VLP of claim 3, wherein the first coiled-coil peptide comprises a P3 peptide and the second coiled-coil peptide comprises a P4 peptide.

5. The VLP of claim 1 , wherein the fusion protein further comprises a cleavable linker and an NES.

6. The VLP of claim 5, wherein the fusion protein comprises at least three NESs and / or the cleavable linker comprises a protease cleavage site.

7. 7. The VLP of claim 6, wherein the protease cleavage site is a Moloney murine leukemia virus (MMLV) protease cleavage site or a Friend murine leukemia virus (FMLV) protease cleavage site.

8. 2. The VLP of claim 1, wherein the fusion protein comprises a gag nucleocapsid protein.

9. 9. The VLP of claim 8, wherein the gag nucleocapsid protein is an MMLV gag nucleocapsid protein or an FMLV gag nucleocapsid protein.

10. (a) the fusion protein comprises a NES within the gag nucleocapsid protein; and / or 9. The VLP of claim 8, wherein (b) the fusion protein comprises a napDNAbp and a domain having DNA polymerase activity, and a cleavable linker is located (i) between the napDNAbp and the DNA polymerase activity domain, and (ii) between the NES.

11. The VLP of claim 10, wherein the NES is located between the p12 domain and the CA domain, within the p12 domain, or between the p12 domain and the MA domain.

12. The VLP of claim 1 , wherein the fusion protein comprises an NLS.

13. 2. The VLP of claim 1, wherein the protein comprises a Rous sarcoma virus (RSV) gag polyprotein, a feline immunodeficiency virus (FIV) gag polyprotein, a simian immunodeficiency virus (SIV) gag polyprotein, a human immunodeficiency virus 1 (HIV-1) gag polyprotein, a human immunodeficiency virus 2 (HIV-2) gag polyprotein, an MMLV gag polyprotein, or an FMLV gag polyprotein.

14. 9. The VLP of claim 8, wherein the fusion protein has the following structure: (i) NH 2 -[gag nucleocapsid protein]-[nap DNAbp]-[DNA polymerase active domain]-COOH; or (ii) NH 2 -[gag nucleocapsid protein]-[1-3 NES]-[cleavable linker]-[NLS]-[nap DNAbp]-[DNA polymerase active domain]-[NLS]-COOH.

15. The VLP of claim 14, wherein each of ]-[ independently comprises a linker.

16. The VLP of claim 1 , wherein the viral envelope glycoprotein is a retroviral envelope glycoprotein.

17. 17. The VLP of claim 16, wherein the viral envelope glycoprotein is a baboon retrovirus envelope glycoprotein.

18. The VLP of claim 1 further comprising a targeting agent.

19. 19. The VLP of claim 18, wherein the targeting agent is an antibody.

20. The VLP of claim 1 , wherein the fusion protein comprises napDNAbp.

21. 21. The VLP of claim 20, wherein the napDNAbp is a Cas9 protein.

22. 22. The VLP of claim 21, wherein the Cas9 protein is Cas9 nickase or nuclease-inactivated Cas9 (dCas9).

23. The VLP of claim 1, wherein the nap DNAbp is linked to a prime editing guide RNA (peg RNA).

24. The VLP of claim 1 , wherein the fusion protein comprises a domain that contains DNA polymerase activity.

25. 25. The VLP of claim 24, wherein the domain comprising DNA polymerase activity comprises DNA-dependent DNA polymerase activity or RNA-dependent DNA polymerase activity.

26. The VLP of claim 1 , wherein the fusion protein comprises a prime editor or a portion thereof.