Modified prime-edited guide RNA

JP2024545947A5Pending Publication Date: 2025-12-04PRIME MEDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024531243
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-20
Filing Date
2022-11-23
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Current prime editing techniques require improved prime editing guide RNAs (PEgRNAs) with enhanced efficiency for making nucleotide substitutions, insertions, and deletions in DNA to correct genetic mutations associated with diseases.

Method used

The development of PEgRNAs comprising specific structural components such as spacers, gRNA cores, extension arms, editing templates, and 3' nucleic acid motifs, including G-quadruplexes, C-quadruplexes, MS2 protein binding sequences, and MMLV reverse transcriptase recruitment sequences, to enhance editing efficiency.

Benefits of technology

The modified PEgRNAs demonstrate improved editing efficiency, enabling more effective correction of genetic mutations and treatment of diseases with genetic components.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Provided herein are compositions and methods relating to modified prime edited guide RNAs.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 283,076, filed November 24, 2021, and U.S. Provisional Application No. 63 / 417,857, filed October 20, 2022, the entire contents of each of which are incorporated herein by reference. [Background technology]

[0002] Prime editing is a gene editing technique that allows researchers to make nucleotide substitutions, insertions, deletions, or combinations thereof in cellular DNA. Prime editing can be used to correct disease-associated genetic mutations and may also be used to treat diseases with a genetic component. Improved prime PEG RNAs with desirable properties, such as the ability to facilitate prime editing with improved efficiency, are needed. Summary of the Invention

[0003] Provided herein are prime editing guide RNAs (PEgRNAs) useful for prime editing, as well as methods of using and making such PEgRNAs.

[0004] In some embodiments, a prime editing guide RNA (PEgRNA) provided herein comprises: (a) a spacer comprising a region complementary to a sequence to be probed in a target strand of a double-stranded target DNA; (b) a guide RNA (gRNA) core capable of binding to a Cas protein; (c) an extension arm comprising: (i) an editing template comprising an intended edit relative to the double-stranded target DNA; and (ii) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in a non-target strand of the double-stranded target DNA; and (d) a 3' nucleic acid motif selected from the group consisting of SEQ ID NOs: 1-15.

[0005] In some embodiments, a prime editing guide RNA (PEgRNA) provided herein comprises: (a) a spacer comprising a region complementary to a sequence to be searched for in a target strand of double-stranded target DNA; (b) a guide RNA (gRNA) core capable of binding to a Cas protein; (c) an extension arm comprising: (i) an editing template comprising an intended edit relative to the double-stranded target DNA; and (ii) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in the non-target strand of the double-stranded target DNA; and (d) a 3' nucleic acid motif comprising: (A) a G-quadruplex or C-quadruplex derived from a VEGF gene promoter, (B) a pseudoknot derived from Potato Leaf Roll Virus (PLRV), (C) an MS2 protein binding sequence, or (D) a Moloney Murine Leukemia Virus (MMLV) reverse transcriptase recruitment sequence, or a Moloney Murine Leukemia Virus (MMLV) replication recognition sequence.

[0006] In some embodiments, the 3' nucleic acid motif is a G-quadruplex or C-quadruplex derived from a VEGF gene promoter (e.g., the G-quadruplex comprises SEQ ID NO: 10 and / or the C-quadruplex comprises SEQ ID NO: 11). In some embodiments, the 3' nucleic acid motif is a pseudoknot derived from Potato Leaf Roll Virus (PLRV) (e.g., the pseudoknot comprises SEQ ID NO: 4). In some embodiments, the 3' nucleic acid motif comprises an MS2 protein binding sequence (e.g., the MS2 protein binding sequence is SEQ ID NO: 9). In some embodiments, the 3' nucleic acid motif comprises an MMLV reverse transcriptase recruitment sequence (e.g., the MMLV reverse transcriptase recruitment sequence comprises SEQ ID NO: 8). In some embodiments, the 3' nucleic acid motif comprises an MMLV replication recognition sequence.

[0007] In some embodiments, the MMLV replication recognition sequence comprises a sequence selected from the group consisting of SEQ ID NOs: 12 to 15. In some embodiments, the 3' nucleic acid motif comprises SEQ ID NO: 1, 2, 3, 5, 6, or 7.

[0008] As provided herein, a PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, a PBS, and a 3' nucleic acid motif. In some embodiments, the PEGRNA further comprises a linker immediately 5' from the 3' nucleic acid motif. The linker can be 2 to 12 nucleotides in length, e.g., 8 nucleotides. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker does not have perfect complementarity to the PBS sequence, editing template, scaffold, and / or extension arm. The linker has 90% or less, 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, or 15% or less complementarity to the extension arm.

[0009] In some embodiments, a prime editing guide RNA (PEgRNA) provided herein comprises: (a) a spacer comprising a region complementary to a sequence to be searched for in a target strand of a double-stranded target DNA, and a guide RNA (gRNA) core capable of binding to a Cas protein; (b) an extension arm comprising: (i) an editing template comprising an intended edit relative to the double-stranded target DNA; and (ii) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in a non-target strand of the double-stranded target DNA; and (c) a guide RNA (gRNA) core comprising at least 80% identity to SEQ ID NO: 16 and containing one or more modifications relative to SEQ ID NO: 16, wherein the one or more modifications are: (A) a first insertion between nucleotides 12 and 13 and a second insertion between nucleotides 16 and 17, wherein the first insertion is reverse-complementary to the second insertion. (B) a first insertion between nucleotides 52 and 53 and a second insertion between nucleotides 56 and 57, where the first insertion is the reverse complement of the second insertion; (C) a complementary substitution of nucleotides 2 and 29, 3 and 28, 4 and 27, 51 and 58, or a combination thereof; (D) a substitution of nucleotides 11-12 with substitution sequence 1 and nucleotides 17-19 with substitution sequence 2; and (B) a guide RNA (gRNA) core comprising one or more modifications, including a substitution of nucleotides 11-12 with replacement sequence 1 and a substitution of nucleotides 17-18 with replacement sequence 2, wherein replacement sequence 1 is at least three nucleotides in length and replacement sequence 2 is the reverse complement of replacement sequence 1; (C) a substitution of nucleotides 11-12 with replacement sequence 1 and nucleotides 17-18 with replacement sequence 2; (D) a substitution of nucleotides 11-12 with replacement sequence 1 and nucleotides 17-18 with replacement sequence 2; (E) a T to G or T to C substitution at nucleotide 5 and a complementary substitution at nucleotide 26; or (F) any combination thereof.

[0010] In some embodiments, the gRNA core comprises a first insertion between nucleotides 12 and 13 and a second insertion between nucleotides 16 and 17. The insertions can be 1 to 6 nucleotides in length. In some embodiments, the first insertion comprises the sequence UGCUG. In some embodiments, the first insertion comprises the sequence CAGCA.

[0011] In some embodiments, the one or more modifications include substitution of nucleotides 49-52 with replacement sequence 1 and substitution of nucleotides 57-60 with replacement sequence 2, where replacement sequence 2 is the reverse complement of replacement sequence 1, and optionally replacement sequence 1 is 7-11 nucleotides in length, and optionally replacement sequence 1 is 7-9 nucleotides in length, and optionally replacement sequence 1 comprises GCGUCUC, GCGUCCC, GCGUCCA, GCGUGA, GCGUAGCC, GCGUGCAGA, GCGUACCCU, or GCGUUGUCG.

[0012] In some embodiments, the first insertion is 1 to 3 nucleotides in length, for example, the first insertion may comprise a sequence selected from the group consisting of C, CC, CA, CG, A, AC, AA, AG, CCC, CCAC, CCAAC, and CCACAC.

[0013] In some embodiments, the gRNA core comprises a first insertion between nucleotides 52 and 53 and a second insertion between nucleotides 56 and 57. The first insertion can be 1 to 8 nucleotides in length. In some embodiments, the gRNA core comprises a complementary substitution of nucleotides 2 and 29, 3 and 28, 4 and 27, 11 and 18, 12 and 17, 51 and 58, or a combination thereof. In some embodiments, the gRNA core comprises a U to A substitution at nucleotide 2. In some embodiments, the gRNA core comprises a U to A substitution at nucleotide 3. In some embodiments, the gRNA core comprises a U to A substitution at nucleotide 4. In some embodiments, the gRNA core comprises a U to G substitution at nucleotide 51 and, optionally, an A to C substitution at nucleotide 58.

[0014] In some embodiments, the gRNA core comprises the substitution of nucleotides 11-12 with replacement sequence 1 and the substitution of nucleotides 17-18 with replacement sequence 2.

[0015] In some embodiments, replacement sequence 1 is 3-5 nucleotides in length. For example, replacement sequence 1 may comprise a sequence selected from the group consisting of CAGC, CCGC, GGAC, UGC, UCC, GAGGC, AGC, GGC, CGCA, GCACA, GGUC, and GGG. In some embodiments, the gRNA core further comprises a U to A substitution at nucleotide 5 and an A to U substitution at nucleotide 26. In some embodiments, the gRNA core comprises complementary substitutions at nucleotides 52 and 57. In some embodiments, the gRNA core comprises a U to G substitution at nucleotide 52 and an A to C substitution at nucleotide 57. In some embodiments, the gRNA core comprises a U to C substitution at nucleotide 52 and an A to G substitution at nucleotide 57. In some embodiments, the gRNA core comprises complementary substitutions at nucleotides 49 and 60. In some embodiments, the gRNA core comprises an A to G substitution at nucleotide 49 and a U to C substitution at nucleotide 60.

[0016] In some embodiments, a prime editing guide RNA (PEgRNA) provided herein comprises: (a) a spacer comprising a region complementary to a sequence to be searched for in a target strand of a double-stranded target DNA; and a guide RNA (gRNA) core capable of binding to a Cas protein; (b) extension arms comprising: (i) an editing template comprising an intended edit relative to the double-stranded target DNA; and (ii) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in a non-target strand of the double-stranded target DNA; and (c) a guide RNA (gRNA) core comprising at least 80% identity to SEQ ID NO: 16 and containing one or more modifications relative to SEQ ID NO: 16, wherein the one or more modifications include: (A) a U to A substitution at nucleotide 5 and an A to U substitution at nucleotide 26; and (B) a. a first insertion between nucleotides 12 and 13 having a sequence of UGCUG and a second insertion between nucleotides 16 and 17 having a sequence of CAGCA; and a. an insertion of substitution sequence 1 at nucleotides 11-12 and substitution sequence 2 at nucleotides 17-18, wherein substitution sequence 1 comprises GGG and substitution sequence 2 comprises UCC; b. an insertion of substitution sequence 1 at nucleotides 11-12 and substitution sequence 2 at nucleotides 17-18, wherein substitution sequence 1 comprises GGG and substitution sequence 2 comprises UCC; c. an insertion of substitution sequence 1 at nucleotides 16-17, wherein substitution sequence 1 comprises GGG and substitution sequence 2 comprises UCC; d. substitution of nucleotides 11-12 with substitution sequence 1 and substitution of nucleotides 17-18 with substitution sequence 2, wherein said substitution sequence 1 comprises GGG and said substitution sequence 2 comprises UCC, an A to G substitution at nucleotide 49, a U to C substitution at nucleotide 60, a U to G substitution at nucleotide 51, and an A to C substitution at nucleotide 58; or e.and a guide RNA (gRNA) core comprising a modification selected from the group consisting of a substitution of nucleotides 11-12 with substitution sequence 1 and a substitution of nucleotides 17-18 with substitution sequence 2, wherein said substitution sequence 1 and substitution sequence 2 comprise CAGC and GCUG, CCGC and GCGG, GGAC and GUCC, GC and GC, CC and GG, GAGGC and GUCUC, AGC and GCU, GGC and GCC, CGCA and UGCG, GCACA and UGUGC, or GGUC and GGCC.

[0017] In some embodiments, the gRNA core comprises nucleotides 62 to 76 of SEQ ID NO:16.

[0018] In some embodiments, the one or more modifications include substitution of nucleotides 49-52 with replacement sequence 1 and substitution of nucleotides 57-60 with replacement sequence 2, where replacement sequence 2 is the reverse complement of replacement sequence 1, and optionally replacement sequence 1 is 7-11 nucleotides in length, and optionally replacement sequence 1 is 7-9 nucleotides in length, and optionally replacement sequence 1 comprises GCGUCUC, GCGUCCC, GCGUCCA, GCGUGA, GCGUAGCC, GCGUGCAGA, GCGUACCCU, or GCGUUGUCG.

[0019] In some embodiments, a prime editing guide RNA (PEgRNA) provided herein comprises: a spacer comprising a region complementary to a sequence to be searched for in a target strand of a double-stranded target DNA; an extension arm comprising an editing template comprising an intended edit relative to the double-stranded target DNA; and a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in a non-target strand of the double-stranded target DNA; and a guide RNA (gRNA) core comprising a sequence selected from the group consisting of SEQ ID NOs: 17-61, 3860-4253, 4255-4349, 4351-4359, and 4452.

[0020] In some embodiments, the gRNA core comprises a sequence selected from the group consisting of SEQ ID NOs: 4294, 4319, 4322, 4286, 4290, 4346, 4271, 4264, 4317, 4330, 4312, 4356, 4280, and 4452. In some embodiments, the gRNA core comprises SEQ ID NO: 4354.

[0021] In some embodiments, the PEgRNA comprises a 3' nucleic acid motif selected from the group consisting of SEQ ID NOs: 1-15, such as SEQ ID NOs: 1, 2, 3, 5, 6, or 7. In some embodiments, the PEgRNA comprises a 3' nucleic acid motif, wherein the 3' nucleic acid motif comprises a sequence selected from the group consisting of a G-quadruplex or a C-quadruplex derived from a VEGF gene promoter, a pseudoknot derived from potato leaf roll virus (PLRV), an MS2 protein binding sequence, a Moloney murine leukemia virus (MMLV) reverse transcriptase recruitment sequence, or a Moloney murine leukemia virus (MMLV) replication recognition sequence.

[0022] In some embodiments, the selected 3' nucleic acid motif is a G-quadruplex or C-quadruplex derived from the VEGF gene promoter. The G-quadruplex may comprise SEQ ID NO: 10. The C-quadruplex may comprise SEQ ID NO: 11. In some embodiments, the selected 3' nucleic acid motif is a pseudoknot derived from Potato Leaf Roll Virus (PLRV). The pseudoknot may comprise SEQ ID NO: 4. In some embodiments, the 3' nucleic acid motif comprises an MS2 protein binding sequence, such as SEQ ID NO: 9. In some embodiments, the 3' nucleic acid motif comprises an MMLV reverse transcriptase recruitment sequence, such as an MMLV reverse transcriptase recruitment sequence comprising SEQ ID NO: 8. In some embodiments, the selected 3' nucleic acid motif comprises an MMLV replication recognition sequence, such as a sequence selected from the group consisting of SEQ ID NOs: 12-15.

[0023] In some embodiments, the PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, a PBS, and a 3' nucleic acid motif. In some embodiments, the PEGRNA further comprises a linker immediately 5' from the 3' nucleic acid motif. The linker is 2 to 12 nucleotides in length, for example, 8 nucleotides in length. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker does not have perfect complementarity to the PBS sequence, the editing template, the scaffold, and / or the extension arm.

[0024] The linker may comprise 90% or less, 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, or 15% or less complementarity to the extension arm.

[0025] In some embodiments, a prime editing guide RNA (PEgRNA) provided herein comprises: (a) a spacer comprising a region complementary to a sequence to be probed in a target strand of a double-stranded target DNA; (b) a guide RNA (gRNA) core capable of binding to a Cas protein; (c) an extension arm comprising: (i) an editing template comprising an intended edit relative to the double-stranded target DNA; and (ii) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in the non-target strand of the double-stranded target DNA; and (d) a tag sequence that is the reverse complement of a sequence in the editing template.

[0026] In any embodiment disclosed herein, the tag sequence can be 4 to 22 nucleotides in length, e.g., 4 to 10 nucleotides in length, 4 to 9 nucleotides in length, 6 to 8 nucleotides in length, 6 nucleotides in length, or 8 nucleotides in length.

[0027] In some embodiments, the tag sequence does not have perfect complementarity to the PBS, gRNA core, and / or spacer. In some embodiments, the PEgRNA comprises, in 5' to 3' order, a spacer, gRNA core, editing template, PBS, and tag sequence. In some embodiments, the PEgRNA comprises, in 5' to 3' order, an editing template, PBS, tag sequence, spacer, and gRNA core.

[0028] In some embodiments, the PEGRNA comprises a linker between the PBS and the tag sequence. The linker can be 2 to 12 nucleotides in length, e.g., 4 to 8 nucleotides, 4 to 6 nucleotides, 8 nucleotides, 6 nucleotides, or 4 nucleotides in length.

[0029] In some embodiments, the linker does not have perfect complementarity to the PBS, gRNA core, and / or spacer. In some embodiments, the linker does not form a secondary structure.

[0030] In some embodiments, the gRNA core comprises a sequence selected from the group consisting of SEQ ID NOs: 16-60, 3860-4359 and 4452.

[0031] Also provided herein are prime editing systems comprising (a) a PEgRNA disclosed herein or one or more polynucleotides encoding a PEgRNA disclosed herein, and (b) a Cas protein and a DNA polymerase, or one or more polynucleotides encoding a prime editor. In some embodiments, the Cas protein has nickase activity. The Cas protein is Cas9, which may contain a mutation in its HNH domain.

[0032] In some embodiments, the Cas9 comprises at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity compared to SEQ ID NO: 4442. The Cas protein is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or Casφ.

[0033] In some embodiments, the DNA polymerase is a reverse transcriptase, such as a retroviral reverse transcriptase. In some embodiments, the reverse transcriptase comprises at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 4444. In some embodiments, the Cas protein and the DNA polymerase are fused or linked to a fusion protein. In some embodiments, the fusion protein comprises the sequence of SEQ ID NO: 4440.

[0034] In some embodiments, the one or more polynucleotides comprise (a) a first sequence encoding an N-terminal portion of a Cas protein and an intein-N, and (b) a second sequence encoding an intein-C, a C-terminal portion of a Cas protein, and a DNA polymerase.

[0035] In some embodiments, the prime editing system comprises one or more vectors comprising one or more polynucleotides encoding the PEGRNA and one or more polynucleotides encoding the prime editor. The one or more vectors may be, for example, AAV vectors. The one or more polynucleotides may be mRNA.

[0036] Also provided herein are lipid nanoparticles (LNPs) or ribonucleoproteins (RNPs) comprising the prime editing systems disclosed herein.

[0037] In some aspects, the present specification provides methods for editing double-stranded target DNA, the methods comprising contacting the target DNA with (a) a prime editor comprising a PEGRNA disclosed herein and a Cas9 nickase and a reverse transcriptase, (b) a prime editing system disclosed herein, or (c) a LNP or RNP disclosed herein. The target DNA disclosed herein can be present in a cell.

[0038] In some embodiments, the editing efficiency for editing double-stranded target DNA is at least 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 2.1-fold, 2.2-fold, 2.3-fold, 2.4-fold, or 2.5-fold higher than the editing efficiency with a control PEG RNA having the same spacer and extension arms, wherein the control PEG RNA comprises a gRNA core having the sequence of SEQ ID NO: 16 and does not comprise a 3' nucleic acid motif or tag.

[0039] In some embodiments, the gRNA core is selected from the group consisting of SEQ ID NOs: 4352, 3860, 3862, 3865, 3908, 3915, 3982, 3991, 4035, 4261, 4262, 4263, 4264, 4265, 4266, 4268, 4277, 4278, 4280, 4283, 4284, 4285, 4286, 4269, 4287, 4288, 4289, 4290, 4291, 4270, 4271, 4272, 4274, 4275, 4276, 4292, 43 4301, 4302, 4304, 4305, 4306, 4309, 4293, 4311, 4312, 4313, 4315, 4316, 4317, 4319, 4320, 4294, 4321, 4322, 4323, 4295, 4296, 4297, 4299, 4324, 4333, 4334, 4338, 4339, 4341, 4342, 4343, 4345, 4346, 4348, 4349, 4328, 4329, 4330, and

[0040] In certain aspects, the PEgRNA provided herein comprises: i) a spacer comprising a region complementary to a sequence to be probed in a target strand of a double-stranded target DNA; ii) a guide RNA (gRNA) core comprising a repeat sequence, a first stem-loop, and a second stem-loop; iii) an editing template comprising an intended edit relative to the double-stranded target DNA; and iv) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in a non-target strand of the double-stranded target DNA, wherein the PEgRNA comprises one or more nucleic acid moieties at its 3' end. In some embodiments, the PEgRNA comprises, in 5' to 3' order, the spacer, gRNA core, editing template, and PBS.

[0041] In some embodiments, the one or more (e.g., two or more, three or more, four or more, or five or more) nucleic acid moieties are selected from the group consisting of a hairpin (e.g., a hairpin comprising a region of self-complementarity, optionally wherein the region of self-complementarity comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive complementary base pairs), a quadruplex (e.g., a G-quadruplex or a C-quadruplex, optionally wherein the G-quadruplex or C-quadruplex is derived from a VEGF gene promoter), a tRNA sequence (e.g., a tRNA sequence, optionally In some embodiments, the nucleic acid moiety comprises a nucleic acid sequence derived from a viral protein binding sequence, optionally wherein the tRNA sequence is a tRNA(proline) sequence, an aptamer (e.g., an aptamer derived from a viral protein binding sequence, optionally wherein the aptamer comprises a viral reverse transcriptase recruitment sequence, optionally wherein the aptamer comprises an MS2 protein binding sequence or a Moloney Murine Leukemia (MMLV) reverse transcriptase recruitment sequence), and / or a pseudoknot (e.g., the pseudoknot is derived from Potato Leaf Roll Virus (PLRV)), or any combination thereof. In some embodiments, the one or more nucleic acid moieties comprise a structure derived from a replication recognition sequence of a retrovirus, optionally wherein the retrovirus is Moloney Murine Leukemia (MMLV). For example, the one or more nucleic acid moieties can comprise a structure set forth in Table 4. In some embodiments, the one or more nucleic acid moieties comprise a nucleic acid sequence selected from SEQ ID NOs: 1-15.

[0042] In some embodiments, the PEGRNA provided herein comprises a linker immediately 5' to one or more nucleic acid moieties. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. In some embodiments, the linker is 2-13 nucleotides in length. In some embodiments, the linker is 8 nucleotides in length. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker has no region complementary to the PBS sequence. In some embodiments, the linker has no region complementary to the editing template.

[0043] In certain embodiments, the gRNA core of a PEG RNA provided herein comprises one or more sequence modifications compared to SEQ ID NO: 16. In some embodiments, the one or more (e.g., two or more, three or more, four or more, or five or more) sequence modifications comprise a gRNA core difference listed in Table 1. In some embodiments, the gRNA core of a PEG RNA comprises a gRNA core sequence listed in Table 1 or Table 2. In some embodiments, the one or more sequence modifications comprise a sequence modification in a repeat sequence. For example, the repeat sequence can comprise at least one inversion of an A / U base pair in the lower stem of the repeat sequence, optionally wherein the lower stem does not contain two, three, four, or more consecutive AU base pairs, and / or the at least one inversion of an A / U base pair in the repeat sequence comprises an inversion of the fourth A / U base pair in the lower stem of the repeat sequence. An example of a gRNA core structure of a PEG RNA with a single AU base pair inversion in the lower stem of the repeat sequence is shown in Figure 12.

[0044] In some embodiments, the sequence modification in the repeat sequence comprises an extension in the upper stem of the repeat sequence. The upper stem extension of the repeat sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 base pairs in length. In some embodiments, the repeat sequence comprises a sequence selected from SEQ ID NOs: 26-37. In some embodiments, the one or more sequence modifications comprise a modification in the second stem loop. In some embodiments, the modification in the second stem loop comprises an inversion of a G / C base pair in the second stem loop. In some embodiments, the gRNA core comprises a nucleic acid sequence selected from SEQ ID NOs: 21, 22, or 25. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 16-61.

[0045] In certain aspects, the PE gRNA provided herein comprises: i) a spacer comprising a region complementary to a sequence to be probed in a target strand of a double-stranded target DNA; ii) a guide RNA (gRNA) core comprising a repeat sequence, a first stem-loop, and a second stem-loop; iii) an editing template comprising an intended edit relative to the double-stranded target DNA; and iv) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in a non-target strand of the double-stranded target DNA, wherein the gRNA core comprises one or more sequence modifications compared to SEQ ID NO: 16.

[0046] In some embodiments, the PEgRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, and a PBS. In some embodiments, one or more (e.g., two or more, three or more, four or more, or five or more) sequence modifications comprise a gRNA core difference listed in Table 1. In some embodiments, the gRNA core of the PEgRNA comprises a gRNA core sequence listed in Table 1 or Table 2. In some embodiments, the one or more sequence modifications comprise a sequence modification in a repeat sequence. For example, the repeat sequence can comprise at least one inversion of an A / U base pair in the lower stem of the repeat sequence, optionally wherein the lower stem does not contain two, three, four, or more consecutive AU base pairs, and / or the at least one inversion of an A / U base pair in the repeat sequence comprises an inversion of a fourth A / U base pair in the lower stem of the repeat sequence. In some embodiments, the sequence modification in the repeat sequence comprises an extension in the upper stem of the repeat sequence. The upper stem extension of the repeat sequence is from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 base pairs. In some embodiments, the repeat sequence comprises a sequence selected from SEQ ID NOs: 26-37. In some embodiments, the one or more sequence modifications comprise a modification in the second stem loop. In some embodiments, the modification in the second stem loop comprises an inversion of a G / C base pair in the second stem loop. In some embodiments, the gRNA core comprises a nucleic acid sequence selected from SEQ ID NOs: 21, 22, or 25. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 16-61.

[0047] In certain embodiments, the PEgRNA provided herein comprises: i) a spacer comprising a region complementary to a sequence to be probed in a target strand of a double-stranded target DNA; ii) a guide RNA (gRNA) core comprising a repeat sequence, a first stem-loop, and a second stem-loop; iii) an editing template comprising an intended edit relative to the double-stranded target DNA; iv) a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in the non-target strand of the double-stranded target DNA; and v) a tag sequence comprising a region complementary to the PBS and / or the editing template. In some embodiments, the PEgRNA comprises, in 5' to 3' order, the spacer, the gRNA core, the editing template, and the PBS. In some embodiments, the PEgRNA comprises, in 5' to 3' order, the editing template, the spacer, the tag sequence, and the gRNA core.

[0048] In some embodiments, the gRNA core comprises a first gRNA core sequence comprising the 5' half of the gRNA core and a second gRNA core sequence comprising the 3' half of the gRNA core, and the PE gRNA comprises, in 5' to 3' order, a spacer, the first gRNA core sequence, an editing template, a PBS, a tag sequence, and a second gRNA core sequence. In some embodiments, the spacer comprises a first spacer sequence comprising the 5' half of the spacer and a second spacer sequence comprising the 3' half of the spacer, and the tag sequence is located between the first spacer sequence and the second spacer sequence. In some embodiments, the tag sequence comprises a region complementary to the editing template. In some embodiments, the tag sequence comprises a region complementary to the PBS. In some embodiments, the tag sequence comprises a region complementary to the editing template and has no substantial complementarity to the PBS. In some embodiments, the tag sequence comprises a region complementary to the editing template and is not complementary to the PBS. In some embodiments, the tag sequence and the editing template each comprise a region complementary to each other, and the 3' half of the region of complementarity of the editing template is located between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 bases 5' to the 3' half of the editing template. In some embodiments, the tag sequence has no substantial complementarity to the spacer. In some embodiments, the tag has no complementarity to the spacer. In some embodiments, the tag sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. In some embodiments, the tag sequence is at least 4, at least 6, or at least 8 nucleotides in length. In some embodiments, the tag sequence comprises a nucleic acid sequence selected from SEQ ID NOs: 62-1960. In some embodiments, the PEG RNA comprises one or more nucleic acid moieties in its 3' half. In some embodiments, the PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, and a PBS.

[0049] In some embodiments, the one or more nucleic acid moieties comprise a hairpin (e.g., the hairpin comprises a region of self-complementarity, optionally the region of self-complementarity comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive complementary base pairs), a quadruplex (e.g., a G-quadruplex or a C-quadruplex, optionally the G-quadruplex or C-quadruplex is derived from a VEGF gene promoter), a tRNA sequence (e.g., a tRNA sequence, optionally the tRNA sequence is a tRNA(proline) sequence), an aptamer (e.g., an aptamer derived from a viral protein binding sequence, optionally the aptamer comprises a viral reverse transcriptase recruitment sequence, optionally the aptamer comprises an MS2 protein binding sequence or a Moloney murine leukemia (MMLV) reverse transcriptase recruitment sequence), and / or a pseudoknot (e.g., the pseudoknot is derived from potato leaf roll virus (PLRV)), or any combination thereof. In some embodiments, one or more nucleic acid moieties comprise a structure derived from a replication recognition sequence of a retrovirus, and optionally, the retrovirus is Moloney Murine Leukemia (MMLV). For example, one or more nucleic acid moieties can comprise a structure set forth in Table 3. In some embodiments, one or more nucleic acid moieties comprise a nucleic acid sequence selected from SEQ ID NOs: 1-15.

[0050] In some embodiments, the PEG RNA comprises a linker. In some embodiments, the linker is i) immediately 5' of one or more nucleic acid moieties, ii) immediately 5' of the tag sequence, iii) immediately 3' of the tag sequence, iv) immediately 3' of the spacer, v) immediately 5' of the spacer, vi) immediately 3' of the gRNA core, and / or vii) immediately 5' of the gRNA core. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. In some embodiments, the linker is 2-12 nucleotides in length. In some embodiments, the linker is 8 nucleotides in length. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker has no region complementary to the PBS sequence. In some embodiments, the linker has no region complementary to the editing template. In some embodiments, the linker comprises a nucleic acid sequence selected from SEQ ID NOs: 1961-3859.

[0051] In certain embodiments, the gRNA core of a PEgRNA provided herein comprises one or more sequence modifications compared to SEQ ID NO: 16. In some embodiments, the one or more (e.g., two or more, three or more, four or more, or five or more) sequence modifications comprise a gRNA core difference listed in Table 1. In some embodiments, the gRNA core of a PEgRNA comprises a gRNA core sequence listed in Table 1 or Table 2. In some embodiments, the one or more sequence modifications comprise a sequence modification in a repeat sequence. For example, the repeat sequence can comprise at least one inversion of an A / U base pair in the lower stem of the repeat sequence, optionally wherein the lower stem does not contain two, three, four, or more consecutive AU base pairs, and / or the at least one inversion of an A / U base pair in the repeat sequence comprises an inversion of a fourth A / U base pair in the lower stem of the repeat sequence. In some embodiments, the sequence modification in the repeat sequence comprises an extension in the upper stem of the repeat sequence. The upper stem extension of the repeat sequence is from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 base pairs. In some embodiments, the repeat sequence comprises a sequence selected from SEQ ID NOs: 26-37. In some embodiments, the one or more sequence modifications comprise a modification in the second stem loop. In some embodiments, the modification in the second stem loop comprises an inversion of a G / C base pair in the second stem loop. In some embodiments, the gRNA core comprises a nucleic acid sequence selected from SEQ ID NOs: 21, 22, or 25. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 16-61.

[0052] In some aspects, the PEgRNA provided herein comprises, in 5' to 3' order: i) a spacer comprising a region complementary to the sequence to be probed in the target strand of a double-stranded target DNA; ii) a 5' portion of a guide RNA (gRNA) core comprising a repeat sequence and a first stem-loop; iii) an editing template comprising the intended edit relative to the double-stranded target DNA; iv) a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA; and v) a 3' portion of the gRNA core comprising a second stem-loop. In some embodiments, the PEgRNA further comprises a tag sequence comprising the region complementary to the PBS and / or the editing template. In some embodiments, the tag sequence is positioned 3' of the 3' portion of the gRNA core.

[0053] In some embodiments, the tag sequence comprises a region complementary to the editing template. In some embodiments, the tag sequence comprises a region complementary to the PBS. In some embodiments, the tag sequence comprises a region complementary to the editing template but does not have substantial complementarity to the PBS. In some embodiments, the tag sequence comprises a region complementary to the editing template but is not complementary to the PBS. In some embodiments, the 5' end of the tag sequence comprises a region complementary to a position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 bases 5' of the 3' end of the editing template. In some embodiments, the tag sequence does not have substantial complementarity to the spacer. In some embodiments, the tag does not have complementarity to the spacer. In some embodiments, the tag sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. The tag sequence can be at least 4 nucleotides, at least 6 nucleotides, or at least 8 nucleotides in length. In some embodiments, the tag sequence comprises a nucleic acid sequence selected from SEQ ID NOs: 62-1960. In some embodiments, the PEGRNA comprises one or more nucleic acid moieties at the 3' end. In some embodiments, the PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, and a PBS.

[0054] In some embodiments, the one or more nucleic acid moieties comprise a hairpin (e.g., the hairpin comprises a region of self-complementarity, optionally the region of self-complementarity comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive complementary base pairs), a quadruplex (e.g., a G-quadruplex or a C-quadruplex, optionally the G-quadruplex or C-quadruplex is derived from a VEGF gene promoter), a tRNA sequence (e.g., a tRNA sequence, optionally the tRNA sequence is a tRNA(proline) sequence), an aptamer (e.g., an aptamer derived from a viral protein binding sequence, optionally the aptamer comprises a viral reverse transcriptase recruitment sequence, optionally the aptamer comprises an MS2 protein binding sequence or a Moloney murine leukemia (MMLV) reverse transcriptase recruitment sequence), and / or a pseudoknot (e.g., the pseudoknot is derived from potato leaf roll virus (PLRV)), or any combination thereof. In some embodiments, one or more nucleic acid moieties comprise a structure derived from a replication recognition sequence of a retrovirus, and optionally, the retrovirus is Moloney Murine Leukemia (MMLV). For example, one or more nucleic acid moieties can comprise a structure set forth in Table 2. In some embodiments, one or more nucleic acid moieties (e.g., a nucleic acid moiety at the 3' end of a PEG RNA) comprise a nucleic acid sequence selected from SEQ ID NOs: 1-15.

[0055] In some embodiments, the PEG RNA comprises a linker. In some embodiments, the linker is located i) immediately 5' of one or more nucleic acid moieties, ii) immediately 5' of the tag sequence, iii) immediately 3' of the tag sequence, iv) immediately 3' of the spacer, v) immediately 5' of the spacer, vi) immediately 3' of the gRNA core, and / or vii) immediately 5' of the gRNA core. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. The linker can be 2 to 12 nucleotides in length. The linker can be 8 nucleotides in length. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker has no region complementary to the PBS sequence. In some embodiments, the linker has no region complementary to the editing template.

[0056] In some embodiments, the 5' portion of the gRNA core and the 3' portion of the guide RNA (gRNA) core comprise one or more sequence modifications compared to SEQ ID NO: 16. In some embodiments, the one or more sequence modifications comprise a gRNA core difference listed in Table 1 or Table 2.

[0057] In some embodiments, the one or more sequence modifications include sequence modifications in the repeat sequence. For example, the repeat sequence can include at least one A / U base pair inversion in the lower stem of the repeat sequence, where optionally the lower stem does not contain two, three, four, or more consecutive AU base pairs, and / or the at least one A / U base pair inversion in the repeat sequence includes an inversion of a fourth A / U base pair in the lower stem of the repeat sequence. In some embodiments, the sequence modification in the repeat sequence includes an extension in the upper stem of the repeat sequence. The extension in the upper stem of the repeat sequence is from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 base pairs. In some embodiments, the repeat sequence includes a sequence selected from SEQ ID NOs: 26-37. In some embodiments, the one or more sequence modifications include a modification in the second stem loop. In some embodiments, the modification in the second stem loop comprises an inversion of a G / C base pair in the second stem loop. In some embodiments, the gRNA core comprises a nucleic acid sequence selected from SEQ ID NOs: 21, 22, or 25. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 16-61. In some embodiments, the linker comprises a nucleic acid sequence selected from SEQ ID NOs: 1961-3859.

[0058] In some embodiments, the PEGRNA provided herein comprises: i) a first sequence comprising a spacer comprising a region complementary to the sequence to be probed in the target strand of a double-stranded target DNA and a first half of a gRNA core; and ii) a second sequence comprising a second half of the RNA core, an editing template comprising the intended edit relative to the double-stranded target DNA, and a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA, wherein the gRNA core comprises a repeat sequence, a first stem-loop, and a second stem-loop.

[0059] In certain aspects, the PEGRNA provided herein comprises: i) a first sequence comprising an editing template comprising an intended edit relative to the double-stranded target DNA, a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA, a spacer comprising a region complementary to a sequence to be probed in the target strand of the double-stranded target DNA, and a first half of a gRNA core; and ii) a second sequence comprising a second half of the gRNA core, wherein the gRNA core comprises a repeat sequence, a first stem-loop, and a second stem-loop.

[0060] In some embodiments, the first sequence is on a first RNA molecule and the second sequence is on a second RNA molecule. In some embodiments, the spacer and the first and second sequences are on the same RNA molecule. In some embodiments, the first half of the gRNA core and the second half of the gRNA core are selected from pairs of gRNA core first and second half sequences shown in Table 2.

[0061] In some embodiments, the PEgRNA further comprises a tag sequence comprising a region complementary to the PBS and / or editing template. In some embodiments, the PEgRNA comprises, in 5' to 3' order, a spacer, the first half of the gRNA core, the second half of the gRNA core, the editing template, the PBS, and the tag sequence. In some embodiments, the PEgRNA comprises, in 5' to 3' order, the editing template, a spacer, the tag sequence, a spacer, the first half of the gRNA core, and the second half of the gRNA core.

[0062] In some embodiments, the tag sequence includes a region complementary to the editing template but does not have substantial complementarity to the PBS. In some embodiments, the tag sequence includes a region complementary to the editing template but does not have substantial complementarity to the PBS. In some embodiments, the 5' end of the tag sequence includes a region complementary to a position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 bases 5' of the 3' end of the editing template. In some embodiments, the tag sequence does not have substantial complementarity to the spacer. In some embodiments, the tag does not have complementarity to the spacer. In some embodiments, the tag sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. The tag sequence can be at least 4 nucleotides, at least 6 nucleotides, or at least 8 nucleotides in length. The tag sequence is 2 to 12 nucleotides in length. In some embodiments, the tag sequence comprises a nucleic acid sequence selected from SEQ ID NOs: 62-1960. In some embodiments, the PEGRNA comprises one or more nucleic acid moieties in the 3' half. In some embodiments, the PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, and a PBS.

[0063] In some embodiments, the one or more nucleic acid moieties comprise a hairpin (e.g., the hairpin comprises a region of self-complementarity, optionally the region of self-complementarity comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive complementary base pairs), a quadruplex (e.g., a G-quadruplex or a C-quadruplex, optionally the G-quadruplex or C-quadruplex is derived from a VEGF gene promoter), a tRNA sequence (e.g., a tRNA sequence, optionally the tRNA sequence is a tRNA(proline) sequence), an aptamer (e.g., an aptamer derived from a viral protein binding sequence, optionally the aptamer comprises a viral reverse transcriptase recruitment sequence, optionally the aptamer comprises an MS2 protein binding sequence or a Moloney murine leukemia (MMLV) reverse transcriptase recruitment sequence), and / or a pseudoknot (e.g., the pseudoknot is derived from potato leaf roll virus (PLRV)), or any combination thereof. In some embodiments, one or more nucleic acid moieties comprise a structure derived from a replication recognition sequence of a retrovirus, and optionally, the retrovirus is Moloney Murine Leukemia (MMLV). For example, one or more nucleic acid moieties can comprise a structure set forth in Table 2. In some embodiments, one or more nucleic acid moieties comprise a nucleic acid sequence selected from SEQ ID NOs: 1-15.

[0064] In some embodiments, the PEG RNA comprises a linker. In some embodiments, the linker is i) immediately 5' of one or more nucleic acid moieties, ii) immediately 5' of the tag sequence, iii) immediately 3' of the tag sequence, iv) immediately 3' of the spacer, v) immediately 5' of the spacer, vi) immediately 3' of the gRNA core, and / or vii) immediately 5' of the gRNA core. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker has no region complementary to the PBS sequence. In some embodiments, the linker has no region complementary to the editing template.

[0065] In some embodiments, the 5' portion of the gRNA core and the 3' portion of the guide RNA (gRNA) core comprise one or more sequence modifications compared to SEQ ID NO: 16. In some embodiments, the one or more sequence modifications comprise a gRNA core difference listed in Table 1 or Table 2. In some embodiments, the linker comprises a nucleic acid sequence selected from SEQ ID NOs: 1961-3859.

[0066] In some embodiments, the one or more sequence modifications include sequence modifications in the repeat sequence. For example, the repeat sequence can include at least one A / U base pair inversion in the lower stem of the repeat sequence, where optionally the lower stem does not contain two, three, four, or more consecutive AU base pairs, and / or the at least one A / U base pair inversion in the repeat sequence includes an inversion of a fourth A / U base pair in the lower stem of the repeat sequence. In some embodiments, the sequence modification in the repeat sequence includes an extension in the upper stem of the repeat sequence. The extension in the upper stem of the repeat sequence is from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 base pairs. In some embodiments, the repeat sequence includes a sequence selected from SEQ ID NOs: 26-37. In some embodiments, the one or more sequence modifications include a modification in the second stem loop. In some embodiments, the modification in the second stem loop comprises an inversion of a G / C base pair in the second stem loop. In some embodiments, the gRNA core comprises a nucleic acid sequence selected from SEQ ID NO: 21, 22, or 25. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 16-61.

[0067] In certain aspects, the description provides methods for preparing a PEGRNA described herein, the methods comprising linking a first sequence to a second sequence.

[0068] In certain aspects, the present specification provides methods for preparing a PEGRNA, the methods comprising synthesizing a polynucleotide comprising a sequence encoding a PEGRNA described herein.

[0069] In certain aspects, the present description provides a PEGRNA system comprising a PEGRNA described herein.

[0070] In certain aspects, the prime editing complex provided herein comprises (i) a PEgRNA or a PEgRNA system of the present disclosure; and (ii) a prime editor comprising a DNA-binding domain and a DNA polymerase domain. In some embodiments, the DNA-binding domain is a CRISPR-associated (Cas) protein domain. In some embodiments, the Cas protein domain has nickase activity. In some embodiments, the Cas protein domain is Cas9. In some embodiments, the Cas9 comprises a mutation in the HNH domain. In some embodiments, the Cas9 comprises an H840A mutation in the HNH domain. In some embodiments, the Cas protein domain is Cas12b. In some embodiments, the Cas protein domain is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or Casφ. In some embodiments, the DNA polymerase domain is a reverse transcriptase. Many reverse transcriptases are capable of DNA-dependent DNA synthesis in addition to RNA-dependent DNA synthesis (i.e., reverse transcription). In some embodiments, the reverse transcriptase is a retroviral reverse transcriptase. In some embodiments, the reverse transcriptase is Moloney murine leukemia virus (M-MLV) reverse transcriptase. In some embodiments, the DNA polymerase and the programmable DNA-binding domain are fused or linked to form a fusion protein. In some embodiments, the fusion protein comprises the sequence of SEQ ID NO: 4440.

[0071] In certain aspects, the present disclosure provides lipid nanoparticles (LNPs) or ribonucleoproteins (RNPs) comprising a prime editing complex or a component thereof. One embodiment provides a polynucleotide encoding a PEgRNA, a PEgRNA system, or a fusion protein of the present disclosure. In some embodiments, the polynucleotide is an mRNA. In some embodiments, the polynucleotide is operably linked to a regulatory element. In some embodiments, the regulatory element is an inducible regulatory element.

[0072] In certain aspects, the present specification provides a vector comprising the polynucleotide of the above embodiments. In some embodiments, the vector is an AAV vector.

[0073] In certain aspects, the description provides an isolated cell comprising a PEGRNA of the present disclosure, a PEGRNA system of the present disclosure, a primed editing complex of the present disclosure, an LNP or RNP of any one of the above embodiments, a polynucleotide of any one of the above embodiments, and / or a vector of any of the above embodiments. In some embodiments, the cell is a human cell.

[0074] In some aspects, the present disclosure includes a pharmaceutical composition comprising: (i) at least one of a PEGRNA of the present disclosure, a PEGRNA system of the present disclosure, a prime editing complex of the present disclosure, an LNP or RNP of an embodiment described herein, a polynucleotide of any one of the above embodiments, a vector of any one of the above embodiments, and / or a cell of any one of the above embodiments; and (ii) at least one pharmaceutically acceptable carrier.

[0075] In certain aspects, the description provides methods of editing a gene, the methods comprising contacting a gene with any one of (i) a PEgRNA or a PEgRNA system of the present disclosure; or (ii) a prime editor comprising a DNA-binding domain and a DNA polymerase domain, wherein the PEgRNA directs the prime editor to incorporate the intended nucleotide edit into the gene, thereby editing the gene.

[0076] In certain aspects, the specification also provides methods of editing a gene, the method comprising contacting a prime editing complex disclosed herein with a gene, wherein the PEGRNA directs the prime editor to incorporate the intended nucleotide edit into the gene, thereby editing the gene.

[0077] In some embodiments, the prime editor synthesizes single-stranded DNA encoded by the editing template, and the single-stranded DNA replaces the sequence to be edited, such that the intended nucleotide edit is incorporated into the region of the gene corresponding to the edit target. In some embodiments, the gene is a cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is in a subject. In some embodiments, the subject is human. In some embodiments, the method further includes administering the cell to a subject after incorporating the intended nucleotide edit.

[0078] In certain embodiments, the present specification provides a cell produced by any one of the above methods. In certain embodiments, the present specification provides a cell population produced by any one of the methods provided herein.

[0079] The novel features of the methods and compositions provided herein are set forth with particularity in the appended claims. A better understanding of the features and advantages of the methods and compositions provided herein will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the methods and compositions provided herein are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0080] [Figure 1] Exemplary nucleic acid moieties (eg, for inclusion at the 3' end of a PEGRNA disclosed herein) are shown.

[0081] [Figure 2-1] Three graphs showing how nucleic acid moieties enhance editing efficiency. Each violin plot represents a combination of 48 unique PEG-RNAs (one spacer, three unique edits, four PBS lengths, and four RTT lengths). [Figure 2-2] Same as above.

[0082] [Figure 3] An exemplary PEG RNA containing a spacer and a PBS is shown.

[0083] [Figure 4] An exemplary PEG RNA with a spacer, RTT, PBS, linker and tag is shown.

[0084] [Figure 5] An exemplary PEG RNA with spacer, RTT, PBS, linker and tag is shown, with the arrow pointing to the "0" position used in Figures 6 and 7.

[0085] [Figure 6] 10 is a graph showing the effect of tag start position on prime edit efficiency.

[0086] [Figure 7-1]10 is a graph showing the effect of tag start position on prime edit efficiency. [Figure 7-2] Same as above.

[0087] [Figure 8] The gRNA core of an exemplary PEG-RNA is shown. The dashed lines and arrows represent two potential positions where the PEG-RNA can be split into two RNA molecules and synthesized separately (e.g., before ligation).

[0088] [Figure 9-1] 1 is a graph showing the editing efficiency by LegRNA. [Figure 9-2] Same as above.

[0089] [Figure 10] FIG. 1 is a schematic diagram showing an exemplary oligonucleotide library design.

[0090] [Figure 11] FIG. 1 is a schematic representation of an exemplary amplified oligonucleotide library.

[0091] [Figure 12] An exemplary structure of the spacer and gRNA core of a pEGRNA with one inversion of an AU base pair in the lower stem of the repeat sequence and a 5-nucleotide extension in the upper stem of the repeat sequence is shown. The remaining parts of the pEGRNA (such as the editing template and PBS) are not shown.

[0092] [Figure 13] Graphs A to C show the effect of tag (Comp tag) binding position on prime editing efficiency. Graph A shows the binding position of a 4-nucleotide tag relative to the prime editing efficiency. Graph B shows the binding position of a 6-nucleotide tag relative to the prime editing efficiency. Graph C shows the binding position of an 8-nucleotide tag relative to the prime editing efficiency.

[0093] [Figure 14]Graph showing the effect of tag (Comp tag) length on primed editing efficiency for a pool of tags that bind perfectly within the editing template of PEG RNA. DETAILED DESCRIPTION OF THE INVENTION

[0094] In some embodiments, the present specification provides compositions and methods related to modified prime editing guide RNAs (PEgRNAs), for example, useful for prime editing applications. In certain embodiments, the present specification provides compositions and methods for introducing a targeted nucleotide edit into target DNA, e.g., correcting a mutation in a gene, including a gene associated with a disease. The compositions provided herein can include a PEgRNA that can guide a prime editor (PE) to a specific DNA target and introduce a nucleotide edit into the target gene.

[0095] The following description and examples illustrate embodiments of the present disclosure in detail. It should be understood that the present disclosure is not limited to the specific embodiments described herein, and that such embodiments may vary. Those skilled in the art will recognize that the present disclosure has numerous variations and modifications that are also within the scope of the present invention. While various features of the present disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, although the present disclosure may be described in the context of separate embodiments for clarity in this specification, the present disclosure may also be implemented in a single embodiment.

[0096] definition Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0097] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, as used herein, the terms "including," "includes," "having," "has," "with," or variations thereof mean "comprising."

[0098] Unless otherwise specified, the terms "comprising," "comprise," "comprises," "having," "have," "has," "including," "includes," "include," "containing," "contains," and "contain" are open-ended and do not exclude additional, unrecited elements or method steps.

[0099] References to "some embodiments," "one embodiment," "one embodiment," or "other embodiments" mean that a particular feature or characteristic described in connection with an embodiment is included in at least one or more embodiments of the present disclosure, but not necessarily in all embodiments.

[0100] The term "about" or "approximately" means within 10% of a specified value, rounded up where only whole numbers are applicable. For example, about 100 means 90 to 110, about 7 means 6 to 8, etc. Where a particular value is recited in the application and claims, unless otherwise specified, it is assumed that the value is modified by the term "about."

[0101] As used herein, "cell" generally refers to a biological cell. A cell is the basic structural, functional, and / or biological unit of an organism. Cells can originate from any organism that has one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of unicellular eukaryotes, protozoan cells, plant cells, animal cells, invertebrate cells (such as fruit flies, cnidarians, echinoderms, and nematodes), cells of vertebrates (such as fish, amphibians, reptiles, birds, and mammals), cells of mammals (such as pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, and humans), and the like. Cells may not be derived from a natural organism (e.g., cells may be synthetically produced and may be referred to as artificial cells).

[0102] In some embodiments, the cells are human cells. The cells may be derived from different tissues, organs, and / or cell types. In some embodiments, the cells are primary cells. In some embodiments, the term primary cells refers to cells isolated from an organism, such as a mammal, and initially grown in tissue culture (i.e., in vitro) before being divided and transferred to subculture. In some non-limiting examples, primary mammalian cells are modified by the introduction (e.g., transfection, transduction, electroporation, etc.) of one or more polynucleotides, polypeptides, and / or prime editing compositions and further passaged. Such modified primary mammalian cells include muscle cells (e.g., cardiomyocytes, smooth muscle cells, myosatellite cells), epithelial cells (e.g., mammary epithelial cells, intestinal epithelial cells, hepatocytes), endothelial cells, glial cells, neural cells, blood-forming components (e.g., lymphocytes, bone marrow cells), and progenitor and stem cells of any of these somatic cell types. In some embodiments, the cells are fibroblasts. In some embodiments, the cells are stem cells. In some embodiments, the cell is a pluripotent stem cell. In some embodiments, the cell is an induced pluripotent stem cell (iPSC). In some embodiments, the cell is a stem cell. In some embodiments, the cell is an embryonic stem cell (ESC). In some embodiments, the cell is a human stem cell. In some embodiments, the cell is a human pluripotent stem cell. In some embodiments, the cell is a human fibroblast. In some embodiments, the cell is an induced human pluripotent stem cell (iPSC). In some embodiments, the cell is a human stem cell. In some embodiments, the cell is a human embryonic stem cell.

[0103] In some embodiments, the cells are not isolated from an organism, but rather form part of a tissue or organ of an organism, such as a mammal. In some non-limiting examples, mammalian cells include muscle cells (e.g., cardiomyocytes, smooth muscle cells, myosatellite cells), epithelial cells (e.g., mammary epithelial cells, intestinal epithelial cells, hepatocytes), endothelial cells, glial cells, neural cells, blood forming components (e.g., lymphocytes, myeloid cells), progenitor cells and stem cells of any of these somatic cell types. In some embodiments, the cells are primary muscle cells. In some embodiments, the cells are myosatellite cells (satellite cells). In some embodiments, the cells are human myosatellite cells (satellite cells). In some embodiments, the cells are stem cells. In some embodiments, the cells are human stem cells.

[0104] In some embodiments, the cell is a differentiated cell. In some embodiments, the cell is a fibroblast. In some embodiments, the cell is a differentiated muscle cell, myosatellite cell, differentiated epithelial cell, or differentiated neuronal cell. In some embodiments, the cell is a skeletal muscle cell. In some embodiments, the skeletal muscle cell is differentiated from an iPSC, an ESC, or a myosatellite cell. In some embodiments, the cell is a differentiated human cell. In some embodiments, the cell is a human fibroblast. In some embodiments, the cell is a differentiated human muscle cell. In some embodiments, the cell is a human myosatellite cell. In some embodiments, the cell is a human skeletal muscle cell. In some embodiments, the human skeletal muscle cell is differentiated from a human iPSC, a human ESC, or a human myosatellite cell. In some embodiments, the cell is differentiated from a human iPSC or a human ESC.

[0105] In some embodiments, the cells comprise a prime editor, a PEgRNA, an ngRNA, a prime editing system, or a prime editing complex. In some embodiments, the cells are from a human subject. In some embodiments, the subject is suffering from a disease or condition associated with a mutation that is modified by prime editing. In some embodiments, the cells are derived from a human and comprise a prime editor, a PEgRNA, an ngRNA, a prime editing system, or a prime editing complex for correcting the mutation. In some embodiments, the cells are from a human subject and the mutation is edited or corrected by prime editing. In some embodiments, the cells are from a human subject and comprise a prime editor, a PEgRNA, an ngRNA, a prime editing system, or a prime editing complex for correcting the mutation. In some embodiments, the cells are from a human subject and comprise a prime editor, a PEgRNA, an ngRNA, a prime editing system, or a prime editing complex for correcting the mutation. In some embodiments, the cells are from a human subject and the mutation is edited or corrected by prime editing.

[0106] As used herein, the term "substantially" can refer to a value approaching 100% of a given value. In some embodiments, the term can refer to an amount that is at least about 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 99.99% of the total amount. In some embodiments, the term can refer to an amount that corresponds to about 100% of the total amount.

[0107] The terms "protein" and "polypeptide" can be used interchangeably to refer to a polymer of two or more amino acids linked by covalent bonds (such as amide bonds) and capable of adopting a three-dimensional structure. In some embodiments, a protein or polypeptide comprises at least 10, 15, 20, 30, or 50 amino acids linked by covalent bonds (e.g., amide bonds). In some embodiments, a protein comprises at least two amide bonds. In some embodiments, a protein comprises multiple amide bonds. In some embodiments, a protein comprises an enzyme, a zymogen protein, a regulatory protein, a structural protein, a receptor, a nucleic acid binding protein, a biomarker, a member of a specific binding pair (e.g., a ligand or aptamer), or an antibody. In some embodiments, a protein can be a full-length protein (e.g., a fully processed protein with a specific biological function). In some embodiments, a protein can be a variant or fragment of a full-length protein. For example, in some embodiments, a Cas9 protein domain comprises an H840A amino acid substitution compared to the naturally occurring S. pyogenes Cas9 protein. Protein or enzyme variants, e.g., variant reverse transcriptases, include polypeptides having an amino acid sequence that is about 60% identical, about 70% identical, about 80% identical, about 90% identical, about 95% identical, about 96% identical, about 97% identical, about 98% identical, about 99% identical, about 99.5% identical, or about 99.9% identical to the amino acid sequence of a reference protein.

[0108] In some embodiments, a protein comprises one or more protein domains or subdomains. As used herein, the terms "polypeptide domain," "protein domain," or "domain" used in the context of a protein or polypeptide refer to a polypeptide chain having one or more biological functions, such as, for example, a catalytic function, a protein-protein binding function, or a protein-DNA function. In some embodiments, a protein comprises multiple protein domains. In some embodiments, a protein comprises multiple naturally occurring protein domains. In some embodiments, a protein comprises multiple protein domains from different naturally occurring proteins. For example, in some embodiments, a prime editor can be a fusion protein comprising a S. pyogenes Cas9 protein domain and a Moloney murine leukemia virus reverse transcriptase protein domain. A protein comprising amino acid sequences from proteins of different origins or naturally occurring may be referred to as a fusion protein or chimeric protein.

[0109] In some embodiments, the protein comprises a functional variant or functional fragment of a full-length wild-type protein. As used herein, a "functional fragment" or "functional portion" refers to any portion of a reference protein (e.g., a wild-type protein) that contains less than the entire amino acid sequence of the reference protein but retains one or more functions (e.g., catalytic or binding functions). For example, a functional fragment of a reverse transcriptase may contain less than the entire amino acid sequence of a wild-type reverse transcriptase but retain the ability to catalyze the polymerization of polynucleotides under at least one set of conditions. When a reference protein is a fusion of multiple functional domains, the functional fragment can retain one or more functions of at least one of the functional domains. For example, a functional fragment of Cas9 contains less than the entire amino acid sequence of wild-type Cas9 but retains DNA binding ability and partially or completely lacks nuclease activity.

[0110] As used herein, a "functional variant" or "functional mutant" refers to any variant or mutant of a reference protein (e.g., a wild-type protein) that contains one or more changes in the amino acid sequence of the reference protein while retaining one or more functions (e.g., catalytic or binding function). In some embodiments, the one or more changes to the amino acid sequence include amino acid substitutions, insertions, or deletions, or any combination thereof. In some embodiments, the one or more changes to the amino acid sequence include amino acid substitutions. For example, a functional variant of a reverse transcriptase may contain one or more amino acid substitutions compared to the amino acid sequence of a wild-type reverse transcriptase, but retain the ability to catalyze the polymerization of polynucleotides under at least one set of conditions. When the reference protein is a fusion of multiple functional domains, the functional variant may retain one or more functions of at least one functional domain. For example, in some embodiments, a functional fragment of Cas9 may contain one or more amino acid substitutions in the nuclease domain compared to the amino acid sequence of a wild-type Cas9 (e.g., an H840A amino acid substitution), but retain DNA binding ability and partially or completely lack nuclease activity.

[0111] As used herein, the term "functional" and its grammatical equivalents can refer to the ability to perform, have, or fulfill an intended purpose. Functionality includes any percentage from baseline up to 100% of the intended purpose. For example, functionality can include or comprise about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or up to about 100% of the intended purpose. In some embodiments, the term functional can mean more than about 100% of normal function, e.g., 125%, 150%, 175%, 200%, 250%, 300%, 400%, 500%, 600%, 700%, or up to about 1000% of the intended purpose. In some embodiments, a protein or polypeptide includes a naturally occurring amino acid (e.g., one of the 20 amino acids commonly found in naturally synthesized peptides, known by the one-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, V). In some embodiments, a protein or polypeptide includes a non-naturally occurring amino acid (e.g., an amino acid outside the 20 amino acids commonly found in naturally synthesized peptides, such as a synthetic amino acid, amino acid analog, or amino acid mimetic). In some embodiments, a protein or polypeptide is modified.

[0112] In some embodiments, a protein includes an isolated polypeptide. The term "isolated" means that the polypeptide is free, to varying degrees, from components that normally accompany it in its natural state or environment. For example, a polypeptide naturally present in a living animal is not isolated; the same polypeptide partially or completely separated from coexisting materials in its natural state is isolated.

[0113] In some embodiments, the protein is present in a cell, tissue, organ, or virus particle. In some embodiments, the protein is present in a cell or part of a cell (such as a bacterial cell, a plant cell, an animal cell, etc.). In some embodiments, the cell is in a tissue, a subject, or in a cell culture. In some embodiments, the cell is a microorganism (such as a bacterium, a fungus, a protozoan, a virus, etc.). In some embodiments, the protein is present in a mixture of analytes (e.g., a lysate). In some embodiments, the protein is present in a lysate from multiple cells or a lysate of a single cell.

[0114] As used herein, the terms "homology," "homology," or "percentage of homology" refer to the degree of sequence identity between an amino acid or polynucleotide sequence and a corresponding reference sequence. "Homology" refers to polymeric sequences, such as similar polypeptide or DNA sequences. Homology refers to, for example, nucleic acid sequences that have at least approximately 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. In other embodiments, a "homologous sequence" of a nucleic acid sequence may exhibit 93%, 95%, or 98% sequence identity to a reference nucleic acid sequence. For example, a "region of homology to a genomic region" is a region of DNA that has a sequence similar to a specific genomic region within a genome. The homologous region can be of any length sufficient to facilitate binding of a spacer, primer binding site, or protospacer sequence to the genomic region. For example, the homologous region can be at least 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1000, 15 ... , 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100 or more bases in length, wherein the homologous region has sufficient homology to bind to the corresponding genomic region.

[0115] In the context of two nucleic acid sequences or two polypeptide sequences, when the percentage of sequence homology or identity is specified, the percentage of homology or identity usually refers to the alignment of the sequences over a portion of the length when two or more sequences are compared and aligned to obtain maximum correspondence.If a position in the compared sequences is occupied by the same base or amino acid, the molecules may be homologous at that position.Unless otherwise stated, sequence homology or identity is evaluated over the specified length of the nucleic acid, polypeptide, or part thereof.In some embodiments, homology or identity is evaluated over a functional or specified portion of the length.

[0116] Alignment of sequences to assess sequence homology can be performed by algorithms known in the art, such as the Basic Local Alignment Search Tool (BLAST) algorithm, described in Altschul et al., J. Mol. Biol. 215:403-410, 1990. A public internet interface for performing BLAST analyses is accessible through the National Center for Biotechnology Information. Additional known algorithms include Smith & Waterman, "Comparison of Biosequences," Adv. Appl. Math. 2:482, 1981; Needleman & Wunsch, "A general method applicable to the search for similarities in the amino acid sequence of two proteins," J. Mol. Biol. 48:443, 1970; Pearson & Lipman, "Improved tools for biological sequence comparison," Proc. Natl. Acad. Sci. USA 85:2444, 1988, or by automated implementations of these or similar algorithms. Global alignment programs can also be used to align similar sequences of approximately the same size. Examples of global alignment programs include NEEDLE (available at www.ebi.ac.uk / Tools / psa / emboss_needle / ), which is part of the EMBOSS package (Rice P et al., Trends Genet., 2000;16:276-277), and the GGSEARCH program https: / / fasta.bioch.virginia.edu / fasta_www2 / , which is part of the FASTA package (Pearson W and Lipman D, 1988, Proc. Natl. Acad. Sci. USA, 85:2444-2448). Both of these programs are based on the Needleman-Wunsch algorithm, which is used to find the optimal alignment (including gaps) along the entire length of two sequences.A detailed discussion of sequence analysis is also provided in Unit 19.3 of Ausubel et al. ("Current Protocols in Molecular Biology" John Wiley & Sons Inc, 1994-1998, Chapter 15, 1998). Unless otherwise specified, the percentage of identity should be determined based on the alignment between the query and reference sequences performed in a Needleman-Wunsch alignment with gap costs set to 11 for presence and 1 for extension, and the percentage of identity is calculated by dividing the number of identities by the length of the alignment.

[0117] Those skilled in the art will understand that the position of an amino acid (or nucleotide) in a homologous sequence can be determined based on the alignment; for example, "H840" in a reference Cas9 sequence can correspond to H839 or another position in a Cas9 homolog.

[0118] The term "polynucleotide" or "nucleic acid molecule" refers to any polymeric form of nucleotides, including DNA, RNA, hybrids thereof, or RNA-DNA chimeric molecules. In some embodiments, the polynucleotide comprises cDNA, genomic DNA, mRNA, tRNA, rRNA, or microRNA. In some embodiments, the polynucleotide is double-stranded, such as double-stranded DNA within a gene. In some embodiments, the polynucleotide is single-stranded or substantially single-stranded, such as single-stranded DNA or mRNA. In some embodiments, the polynucleotide is an extracellular nucleic acid molecule. In some embodiments, the polynucleotide circulates in the blood. In some embodiments, the polynucleotide is an intracellular nucleic acid molecule. In some embodiments, the polynucleotide is an intracellular nucleic acid molecule circulating in the blood.

[0119] Polynucleotides can have any three-dimensional structure. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, ESTs, or SAGE tags), exons, introns, intergenic DNA (including, but not limited to, heterochromatic DNA), messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA, isolated RNA, sgRNA, guide RNA, nucleic acid probes, primers, snRNA, long non-coding RNA, snoRNA, siRNA, miRNA, small tRNA-derived RNA (tsRNA), antisense RNA, shRNA, or small rDNA-derived RNA (srRNA).

[0120] In some embodiments, a polynucleotide comprises deoxyribonucleotides, ribonucleotides, or their analogs. In some embodiments, a polynucleotide comprises modified nucleotides, such as methylated nucleotides or nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the polynucleotide. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide can be further modified after polymerization, such as by conjugation with a labeling component.

[0121] In some embodiments, a polynucleotide is composed of a specific sequence of four nucleotide bases: adenine (A), cytosine (C), guanine (G), thymine (T), and, if the polynucleotide is RNA, uracil (U) in place of thymine. In some embodiments, a polynucleotide may include one or more other nucleotide bases, such as inosine (I), which is read as guanine (G) by the translation machinery. In accordance with the ST.26 standard, RNA sequences (e.g., PEgRNA sequences) provided herein may include "T" in place of "U."

[0122] In some embodiments, a polynucleotide may be modified. As used herein, the term "modified" or "modification" refers to chemical modifications of A, C, G, T, and U nucleotides. In some embodiments, modifications may be made to the nucleoside base and / or sugar moieties of the nucleosides comprising the polynucleotide. In some embodiments, modifications may be made on the internucleoside linkage (e.g., the phosphate backbone). In some embodiments, a modified nucleic acid molecule comprises multiple modifications. In some embodiments, a modified nucleic acid molecule comprises a single modification.

[0123] As used herein, the terms "complementary," "complementary," or "complementarity" refer to the ability of two polynucleotide molecules to base pair with each other. Complementary polynucleotides can base pair through Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonds. For example, an adenine on one polynucleotide molecule base pairs with a thymine or uracil on a second polynucleotide molecule, a cytosine on one polynucleotide molecule base pairs with a guanine on a second polynucleotide molecule, a guanine on one polynucleotide molecule base pairs with a cytosine or uracil on a second polynucleotide molecule, and a thymine or uracil on one polynucleotide molecule base pairs with an adenine or guanine on a second polynucleotide molecule. Two polynucleotide molecules are complementary to each other if the first polynucleotide molecule containing the first nucleotide sequence can base pair with the second polynucleotide molecule containing the second nucleotide sequence. For example, two DNA molecules 5'-ATGC-3' and 5'-GCAT-3' are complementary, and the complementary molecule of the DNA molecule 5'-ATGC-3' is 5'-GCAT-3'. The percentage of complementarity indicates the percentage of nucleotides in a polynucleotide molecule that can base pair with a second polynucleotide molecule (e.g., 5, 6, 7, 8, 9, and 10 out of 10 indicate 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). "Fully complementary" means that all contiguous nucleotides of a polynucleotide molecule base pair with the same number of contiguous nucleotides in a second polynucleotide molecule. As used herein, "substantially complementary" refers to a degree of complementarity that can be 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% across all or a portion of two polynucleotide molecules. In some embodiments, the portion of complementarity can be a region of 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides. "Substantial complementarity" can also refer to portions of two polynucleotide molecules that are 100% complementary.In some embodiments, the portion of complementarity between two polynucleotide molecules is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% of the length of at least one of the two polynucleotide molecules or a functional or defined portion thereof.

[0124] As used herein, "expression" refers to the process by which a polynucleotide is transcribed into mRNA and / or the process by which a polynucleotide (e.g., a transcribed mRNA) is translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell. In some embodiments, expression of a polynucleotide, e.g., a gene or DNA encoding a protein, is determined by the amount of protein encoded by the gene after transcription and translation of the gene. In some embodiments, expression of a polynucleotide, e.g., a gene or DNA encoding a protein, is determined by the amount of functional form of protein encoded by the gene after transcription and translation of the gene. In some embodiments, expression of a gene is determined by the amount of mRNA, i.e., transcript, encoded by the gene after transcription of the gene. In some embodiments, expression of a polynucleotide, e.g., an mRNA, is determined by the amount of protein encoded by the mRNA after translation of the mRNA. In some embodiments, expression of a polynucleotide, e.g., an mRNA or coding RNA, is determined by the amount of functional form of protein encoded by the polypeptide after translation of the polynucleotide.

[0125] As used herein, the term "sequencing" may include capillary sequencing, sulfite-free sequencing, sulfite sequencing, TET-assisted sulfite (TAB) sequencing, ACE sequencing, high-throughput sequencing, Maxam-Gilbert sequencing, massively parallel signature sequencing, polony sequencing, 454 pyrosequencing, Sanger sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, heliscope single molecule sequencing, single molecule real-time (SMRT) sequencing, nanopore sequencing, shotgun sequencing, RNA sequencing, or a combination thereof.

[0126] The terms "equivalent" or "bioequivalent" are used interchangeably when referring to a particular molecule, or biological or cellular material, and refer to a molecule that has minimal homology to another molecule while maintaining a desired structure or function.

[0127] The term "encode" as applied to a polynucleotide refers to a polynucleotide that is said to "encode" another polynucleotide, polypeptide, or amino acid when, either naturally occurring or manipulated by methods well known to those of skill in the art, it can be used as a template for polynucleotide synthesis, e.g., transcribed into RNA, reverse transcribed into DNA or cDNA, and / or translated to produce an amino acid, polypeptide, or fragment thereof. In some embodiments, a polynucleotide consisting of three consecutive nucleotides forms a codon that encodes a particular amino acid. In some embodiments, a polynucleotide includes one or more codons that encode a polypeptide. In some embodiments, a polynucleotide including one or more codons includes a mutation in the codon compared to a wild-type reference polynucleotide. In some embodiments, the codon mutation encodes an amino acid substitution in a polypeptide encoded by the polynucleotide compared to the wild-type reference polypeptide.

[0128] As used herein, the term "mutation" refers to a change and / or alteration in the amino acid sequence of a protein or the nucleic acid sequence of a polynucleotide. Such a change and / or alteration can include a substitution, insertion, deletion, and / or truncation of one or more amino acids in the case of an amino acid sequence, or nucleotides in the case of a nucleic acid sequence, compared to a reference amino acid or nucleic acid sequence. In some embodiments, the reference sequence is a wild-type sequence. In some embodiments, a mutation in the nucleic acid sequence of a polynucleotide encodes a mutation in the amino acid sequence of a polypeptide. In some embodiments, the mutation in the amino acid sequence of a polypeptide or the mutation in the nucleic acid sequence of a polynucleotide is a mutation associated with a disease state.

[0129] As used herein, the term "subject" and its grammatical equivalents can refer to a human or a non-human. The subject can be a mammal. The human concerned is male or female. The human subject can be of any age. The subject can be a human fetus. The human subject can be a newborn, infant, child, adolescent, or adult. The human subject can be up to about 100 years old. The human subject can be in need of treatment for a genetic disease or disorder.

[0130] The terms "treatment" or "treating," and their grammatical equivalents, can refer to medical management administered to a subject with the intent to cure, ameliorate, or alleviate the symptoms of a disease, condition, or disorder. Treatment can include active treatment, i.e., treatment focused on ameliorating the disease, condition, or disorder. Treatment can include causal treatment, i.e., treatment directed at eliminating the cause of the associated disease, condition, or disorder. Additionally, treatment can include palliative treatment, aimed at alleviating symptoms but not curing the disease, condition, or disorder. Treatment can include supportive care, i.e., treatment used to complement another specific therapy directed at ameliorating the disease, condition, or disorder. In some embodiments, the condition can be pathological. In some embodiments, the disease, condition, or disorder may not be completely cured or prevented by treatment. In some embodiments, treatment reduces, but does not completely cure or prevent, the disease, condition, or disorder. In some embodiments, a subject can receive treatment for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or for the life of the subject.

[0131] The term "ameliorate" and its grammatical equivalents mean to relieve, suppress, attenuate, diminish, arrest, or stabilize the onset or progression of a disease.

[0132] The terms "prevent" or "preventing" mean to delay, forestall, or avert the onset or progression of a disease, condition, or disorder for a period of time. Prevention can also mean reducing the risk of developing a disease, disorder, or condition. Prevention includes minimizing, partially, or completely inhibiting the onset of a disease, condition, or disorder. In some embodiments, the composition, e.g., pharmaceutical composition, prevents a disease by delaying the onset of the disease for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or for the lifetime of the subject.

[0133] The term "effective amount" or "therapeutically effective amount" can refer to an amount of a composition, e.g., a composition comprising a construct, as disclosed herein, sufficient to result in a desired activity when introduced into a subject. An effective amount of a prime editing composition can be provided to a target gene or cell, regardless of whether the cell is ex vivo or in vivo.

[0134] An effective amount can be, for example, an amount that induces at least about a two-fold or greater change (increase or decrease) in the amount of target nucleic acid regulation (e.g., expression of a gene to produce a functional protein) observed compared to a negative control. An effective amount or dosage can, for example, induce an about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, about 10-fold, about 25-fold, about 50-fold, about 100-fold, about 200-fold, about 500-fold, about 700-fold, about 1000-fold, about 5000-fold, or about 10,000-fold increase in target gene regulation (e.g., expression of a target gene to produce a functional protein).

[0135] The amount of modulation of a target gene can be measured by any suitable method known in the art. In some embodiments, an "effective amount" or "therapeutically effective amount" is the amount of a composition required to improve disease symptoms compared to an untreated patient. In some embodiments, an effective amount is the amount of a composition sufficient to introduce an alteration into a gene of interest in a cell (e.g., a cell in vitro or in vivo).

[0136] Prime Edit The term "prime editing" refers to programmable editing of target DNA using a prime editor complexed with a PEgRNA to incorporate intended nucleotide edits into target DNA through target-primed DNA synthesis. A target polynucleotide for prime editing, e.g., a target gene, may comprise a double-stranded DNA molecule having two complementary strands, the first strand being referred to as the "target strand" or "non-edited strand" and the second strand being referred to as the "non-target strand" or "edited strand." In some embodiments, in the prime editing guide RNA (PEgRNA), a spacer sequence is complementary or substantially complementary to a specific sequence on the target strand, which may also be referred to as the "search sequence." In some embodiments, the spacer sequence anneals to the target strand at the search sequence. The target strand may also be referred to as the "non-protospacer adjacent motif (non-PAM strand)." In some embodiments, the non-target strand is also referred to as the "PAM strand." In some embodiments, the PAM strand comprises a protospacer sequence and, optionally, a protospacer adjacent motif (PAM) sequence. In prime editing using a Cas protein-based prime editor, the PAM sequence refers to a short DNA sequence immediately adjacent to the protospacer sequence on the PAM strand of the target gene. The PAM sequence can be specifically recognized by a programmable DNA-binding protein, such as a Cas nickase or a Cas nuclease. In some embodiments, a particular PAM is characteristic of a particular programmable DNA-binding protein, such as a Cas nickase or a Cas nuclease. The protospacer sequence refers to a specific sequence within the PAM strand of the target gene that is complementary to the search control sequence. In PEG RNA, the spacer sequence can have a sequence substantially identical to the protospacer sequence on the edited strand of the target gene, but the spacer sequence can contain uracil (U) and the protospacer sequence can contain thymine (T).

[0137] In some embodiments, the double-stranded target DNA contains a nick site on the PAM strand (or non-target strand). As used herein, "nick site" refers to a specific position between two nucleotides or two base pairs in the double-stranded target DNA. In some embodiments, the position of the nick site is determined relative to the position of a specific PAM sequence. In some embodiments, the nick site is a specific position where a nick occurs when the double-stranded target DNA is contacted with a nickase (such as a Cas nickase) that recognizes a specific PAM sequence. In some embodiments, the nick site is upstream of the specific PAM sequence on the PAM strand of the double-stranded target DNA. In some embodiments, the nick site is downstream of the specific PAM sequence on the PAM strand of the double-stranded target DNA. In some embodiments, the nick site is three base pairs upstream of the PAM sequence, and the PAM sequence is recognized by Streptococcus pyogenes Cas9 nickase, P. lavamentivorans Cas9 nickase, C. diphtheriae Cas9 nickase, C. cinerea Cas9, S. aureus Cas9, or N. lari Cas9 nickase. In some embodiments, the nick site is three base pairs upstream of the PAM sequence, and the PAM sequence is recognized by Cas9 nickase, which comprises a nuclease-active HNH domain and a nuclease-inactive RuvC domain. In some embodiments, the nick site is two base pairs upstream of the PAM sequence, and the PAM sequence is recognized by S. thermophilus Cas9 nickase.

[0138] In some embodiments, the PEgRNA forms a complex with the prime editor, directing the prime editor to bind to the target sequence of the target gene. In some embodiments, the bound prime editor generates a nick in the edited strand (PAM strand) of the target gene at the nick site. In some embodiments, the primer binding site (PBS) of the PEgRNA anneals to the free 3' end formed at the nick site, and the prime editor initiates DNA synthesis from the nick site using the free 3' end as a primer. A single-stranded DNA encoded by the PEgRNA editing template is then synthesized. In some embodiments, the newly synthesized single-stranded DNA contains one or more intended nucleotide edits compared to the endogenous target gene sequence. In some embodiments, the PEgRNA editing template is complementary to a sequence in the edited strand except for one or more mismatches at the intended nucleotide editing positions in the editing template; the partially complementary editing template may be referred to as an "editing target sequence." Thus, in some embodiments, the newly synthesized single-stranded DNA has identical or substantial identity to a sequence in the edited sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide editing positions.

[0139] In some embodiments, the newly synthesized single-stranded DNA equilibrates with the editing target in the edited strand of the target gene and pairs with the target strand of the target gene. In some embodiments, the sequence to be edited in the target gene is excised by a flap endonuclease (FEN), e.g., FEN1. In some embodiments, the FEN is, for example, an endogenous FEN in the cell containing the target gene. In some embodiments, the FEN is provided as part of the prime editor, linked to other components of the prime editor, or provided in trans. In some embodiments, the newly synthesized single-stranded DNA containing the intended nucleotide edit replaces the endogenous single-stranded sequence to be edited on the edited strand of the target gene. In some embodiments, the newly synthesized single-stranded DNA and the endogenous DNA on the target strand form a heteroduplex DNA structure in the region corresponding to the sequence to be edited in the target gene. In some embodiments, the newly synthesized single-stranded DNA containing the nucleotide edit pairs in a heteroduplex with the target strand of the target DNA that does not contain the nucleotide edit, thereby creating a mismatch between the two originally complementary strands. In some embodiments, the mismatch is recognized by a DNA repair mechanism, e.g., an endogenous DNA repair mechanism, and in some embodiments, through DNA repair, the intended nucleotide edit is incorporated into the target gene. Modified PEG-RNA In certain aspects, the present specification provides modified PEgRNAs. The term "prime editing guide RNA" or "PEgRNA" refers to a guide polynucleotide containing one or more intended nucleotide edits to be incorporated into a target DNA. In some embodiments, the PEgRNA is associated with a prime editor and directs the prime editor to incorporate one or more (e.g., two or more, three or more, four or more, or five or more) intended nucleotide edits into a target gene via prime editing. A "nucleotide edit" or "intended nucleotide edit" refers to a specific deletion of one or more nucleotides at a specific position, an insertion of one or more nucleotides at a specific position, a substitution of a single nucleotide, or other change at a specific position that is incorporated into the sequence of a target gene. An intended nucleotide edit can refer to an edit on an editing template compared to the sequence on the target strand of a target gene, or it can refer to an edit encoded by the editing template on newly synthesized single-stranded DNA that replaces the editing target sequence compared to the editing target sequence. In some embodiments, the PEgRNA comprises a spacer sequence complementary or substantially complementary to the sequence to be probed on the target strand of the target gene. In some embodiments, the PEgRNA comprises a gRNA core associated with the DNA-binding domain (e.g., a CRISPR-Cas protein domain) of a prime editor. In some embodiments, the PEgRNA further comprises an extended nucleotide sequence comprising one or more intended nucleotide edits relative to the endogenous sequence of the target gene; the extended nucleotide sequence may also be referred to as an extension arm. In certain embodiments, the PEgRNA comprises a primer binding site sequence (PBS) capable of initiating target-primed DNA synthesis. In some embodiments, the PEgRNA comprises an editing template comprising one or more intended nucleotide edits to be incorporated into the target gene by prime editing. In some embodiments, the extension arm comprises a PBS. In some embodiments, the P extension arm comprises an editing template comprising one or more intended nucleotide edits to be incorporated into the target gene by prime editing.

[0140] The "primer binding site" (PBS or primer binding site sequence) is a single-stranded portion of the PEG RNA that includes a region complementary to the PAM strand (i.e., the non-target strand or edited strand). The PBS is complementary or substantially complementary to a sequence on the PAM strand of the double-stranded target DNA immediately upstream of the nick site. In some embodiments, in the process of prime editing, the PEG RNA forms a complex with the prime editor and directs the prime editor to bind to the sequence being probed on the target strand of the double-stranded target DNA, generating a nick at the nick site on the non-target strand of the double-stranded target DNA. In some embodiments, the PBS is complementary or substantially complementary to and can anneal to the free 3' end on the non-target strand of the double-stranded target DNA at the nick site. In some embodiments, the PBS annealed to the free 3' end of the non-target strand can initiate target-primed DNA synthesis.

[0141] The "editing template" of a PEgRNA is a single-stranded portion of the PEgRNA that is 5' to the PBS, includes a region of complementarity to the PAM strand (i.e., the non-target or editing strand), and contains one or more intended nucleotide edits compared to the endogenous sequence of the double-stranded target DNA. In some embodiments, the editing template and the PBS are directly adjacent to each other. Thus, in some embodiments, the PEgRNA during prime editing contains a single-stranded portion that includes the PBS and the editing template directly adjacent to each other. In some embodiments, the single-stranded portion of the PEgRNA that includes both the PBS and the editing template is complementary or substantially complementary to the endogenous sequence on the PAM strand (i.e., the non-target or editing strand) of the double-stranded target DNA, except for one or more non-complementary nucleotides at the intended nucleotide edit positions. As used herein, regardless of their relative 5'-3' arrangement in other contexts, the relative positions between the PBS and the editing template, and between elements of the PEgRNA, are determined by the 5' to 3' order of the PEgRNA as a single molecule, regardless of the location of sequences within the double-stranded target DNA that may share complementarity or identity with elements of the PEgRNA. In some embodiments, the editing template is complementary or substantially complementary to a sequence on the PAM strand immediately downstream of the nick site, except for one or more non-complementary nucleotides at the intended nucleotide editing position. An endogenous, e.g., genomic, sequence that is complementary or substantially complementary to the editing template, except for one or more non-complementary nucleotides at the position corresponding to the intended nucleotide edit, may be referred to as the "sequence to be edited." In some embodiments, the editing template is complementary to the sequence to be edited, or has identity or substantial identity to a sequence on the target strand at the same position in the genome as the sequence to be edited, except for one or more insertions, deletions, or substitutions at the intended nucleotide editing position. In some embodiments, the editing template encodes a single-stranded DNA that is identical or substantially identical to the sequence to be edited except for one or more insertions, deletions, or substitutions at the location(s) of one or more intended nucleotide edits.

[0142] Spacer The spacer can guide the prime editing complex to a genomic locus with the same or substantially the same sequence during prime editing. In some embodiments, the PEGRNA comprises a spacer. In some embodiments, the length of the spacer varies from at least 10 nucleotides to 100 nucleotides. For example, the spacer can be at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, or at least 100 nucleotides. In some embodiments, the spacer is 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides in length. In some embodiments, the spacer is 15 to 30 nucleotides in length, 15 to 25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, 40 to 50 nucleotides in length, 50 to 60 nucleotides in length, 60 to 70 nucleotides in length, 70 to 80 nucleotides in length, or 90 to 100 nucleotides in length. In some embodiments, the spacer is 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length.

[0143] In some embodiments, the spacer sequence comprises a region that is substantially complementary to the sequence to be searched for on the target strand of the double-stranded target DNA. In some embodiments, the spacer sequence of the PEG RNA is identical or substantially identical to the protospacer sequence on the edited strand of the target gene (except that the protospacer sequence may include thymine and the spacer sequence may include uracil). In some embodiments, the spacer sequence is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to the sequence to be searched for in the target gene. In some embodiments, the spacer is substantially complementary to the sequence to be searched for.

[0144] The length of the spacer varies from at least 10 to 100 nucleotides. For example, the spacer can be at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, or at least 100 nucleotides. In some embodiments, the spacer is 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. In some embodiments, the spacer is 15 to 30 nucleotides in length, 15 to 25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, 40 to 50 nucleotides in length, 50 to 60 nucleotides in length, 60 to 70 nucleotides in length, 70 to 80 nucleotides in length, or 90 to 100 nucleotides in length. In some embodiments, the spacer is 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length.

[0145] As used herein, unless otherwise specified, in a PEgRNA or nicked guide RNA sequence, or fragments thereof, such as a spacer, PBS, or RTT sequence, the letter "T" or "thymine" refers to a nucleobase in the DNA sequence encoding the PEgRNA or guide RNA sequence, and is intended to refer to a uracil (U) nucleobase of the PEgRNA or guide RNA, or a chemically modified uracil nucleobase known in the art, such as 5-methoxyuracil.

[0146] Primer binding site (PBS) The PEGRNA comprises a primer binding site (PBS) and an editing template (e.g., RTT). The extension arm of the PEGRNA comprises a PBS and an editing template. In some embodiments, the PBS may be partially complementary to the spacer. In some embodiments, the editing template (e.g., RTT) is partially complementary to the spacer. In some embodiments, the editing template (e.g., RTT) and the primer binding site (PBS) are each partially complementary to the spacer.

[0147] The extension arm of the PEgRNA may contain a primer binding site sequence (PBS, or PBS sequence) that hybridizes to the free 3' end of the single-stranded DNA in the target gene generated by the nick by the prime editor. The length of the PBS sequence may vary depending on, for example, the components of the prime editor, the sequence to be searched, and other components of the PEgRNA. In some embodiments, the length of the primer binding site (PBS) varies from at least 2 nucleotides to 50 nucleotides. For example, the primer binding site (PBS) can be at least 2 nucleotides long, at least 3 nucleotides long, at least 4 nucleotides long, at least 5 nucleotides long, at least 6 nucleotides long, at least 7 nucleotides long, at least 8 nucleotides long, at least 9 nucleotides long, at least 10 nucleotides long, at least 11 nucleotides long, at least 12 nucleotides long, at least 13 nucleotides long, at least 14 nucleotides long, at least 15 nucleotides long, at least 16 nucleotides long, at least 17 nucleotides long, at least 18 nucleotides long, at least 19 nucleotides long, at least 20 nucleotides long, at least 30 nucleotides long, at least 40 nucleotides long, or at least 50 nucleotides long. In some embodiments, the PBS is at least 6 nucleotides in length. In some embodiments, the PBS is about 4-16 nucleotides in length, about 6-16 nucleotides in length, about 6-18 nucleotides in length, about 6-20 nucleotides in length, about 8-20 nucleotides in length, about 10-20 nucleotides in length, about 12-20 nucleotides in length, about 14-20 nucleotides in length, about 16-20 nucleotides in length, or about 18-20 nucleotides in length. In some embodiments, the PBS is about 7-15 nucleotides in length. In some embodiments, the PBS is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. In some embodiments, the PBS is 8, 9, 10, 11, 12, 13, or 14 nucleotides in length.

[0148] The PBS can be complementary or substantially complementary to a DNA sequence in the edited strand of the target gene. The PBS can anneal to a free hydroxy group in the edited strand, for example, the free 3' end generated by the nick of the prime editor, thereby initiating the synthesis of a new single-stranded DNA encoded by the editing template at the nick site. In some embodiments, the PBS is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to a region of the edited strand of the target gene. In some embodiments, the PBS is fully complementary, i.e., 100% complementary, to a region of the edited strand of the target gene.

[0149] The extension arms of the pegRNA can contain an editing template that serves as a DNA synthesis template for DNA polymerase in the prime editor during prime editing.

[0150] The length of the editing template may vary depending, for example, on the components of the prime editor, the sequence to be searched, and other components of the PEGRNA. In some embodiments, the editing template serves as a DNA synthesis template for reverse transcriptase, and the editing template is referred to as a reverse transcription editing template (RTT).

[0151] In some embodiments, the editing template (e.g., RTT) is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. In some embodiments, the RTT is 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. In some embodiments, the RTT is 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in length.

[0152] In some embodiments, the editing template (e.g., RTT) sequence is about 70%, 75%, 80%, 85%, 90%, 95%, or 99% complementary to the editing target sequence on the edited strand of the target gene. In some embodiments, the editing template sequence (e.g., RTT) is substantially complementary to the sequence to be edited. In some embodiments, the editing template sequence (e.g., RTT) is complementary to the sequence to be edited except for the nucleotide edit position to be incorporated into the target gene. In some embodiments, the editing template comprises a nucleotide sequence that is about 85% to about 95% complementary to the sequence to be edited in the edited strand of the target gene. In some embodiments, the editing template comprises about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% complementary to the sequence to be edited in the edited strand of the target gene.

[0153] In some embodiments, the PEgRNA forms an RNA polynucleotide that contains only RNA nucleotides. In some embodiments, the PEgRNA is a chimeric polynucleotide that contains both RNA and DNA nucleotides. For example, the PEgRNA can contain DNA in the spacer sequence, gRNA core, or extension arm. In some embodiments, the PEgRNA contains DNA within the spacer sequence. In some embodiments, the entire spacer sequence of the PEgRNA is a DNA sequence. In some embodiments, the PEgRNA contains DNA in the gRNA core, e.g., the stem region of the gRNA core. In some embodiments, the PEgRNA contains DNA within the extension arm of the editing template. An editing template containing a DNA sequence can serve as a DNA synthesis template for a DNA polymerase, e.g., a DNA-dependent DNA polymerase, in the prime editor. Thus, the PEgRNA can be a chimeric polynucleotide that contains RNA, gRNA core, and / or PBS sequences in the spacer and DNA in the editing template.

[0154] The components of a PEGRNA can be arranged in a modular fashion. In some embodiments, the spacer and extension arm, including the primer binding site sequence (PBS) and editing template (e.g., reverse transcriptase template (RTT)), can be interchangeably positioned in the 5' portion of the PEGRNA, the 3' portion of the PEGRNA, or in the center of the gRNA core. For example, in some embodiments, a PEGRNA comprises, from 5' to 3', a spacer, a gRNA core, an editing template, and a PBS. In some embodiments, a PEGRNA comprises, from 5' to 3', an editing template, a PBS, a spacer, and a gRNA core. In some embodiments, the PBS and / or editing template are located within the gRNA core, i.e., sandwiched between the first half of the gRNA core and the second half of the gRNA core.

[0155] In certain embodiments, the PEgRNA provided herein comprises: i) a spacer comprising a region complementary to the sequence to be probed in the target strand of a double-stranded target DNA; ii) a guide RNA (gRNA) core comprising a repeat sequence, a first stem-loop, and a second stem-loop; iii) an editing template comprising an intended edit relative to the double-stranded target DNA; and iv) a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA, wherein the PEgRNA further comprises one or more nucleic acid moieties at the 3' end.

[0156] In some embodiments, the PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, and a PBS.

[0157] In certain embodiments, the PE gRNA provided herein comprises: i) a spacer comprising a region complementary to the sequence to be probed in the target strand of a double-stranded target DNA; ii) a guide RNA (gRNA) core comprising a repeat sequence, a first stem-loop, and a second stem-loop; iii) an editing template comprising an intended edit relative to the double-stranded target DNA; and iv) a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA, wherein the gRNA core comprises one or more sequence modifications compared to SEQ ID NO: 16.

[0158] In some embodiments, the PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, and a PBS.

[0159] In certain embodiments, the PEgRNA provided herein comprises: i) a spacer comprising a region complementary to the sequence to be probed in the target strand of a double-stranded target DNA; ii) a guide RNA (gRNA) core comprising a repeat sequence, a first stem-loop, and a second stem-loop; iii) an editing template comprising the intended edit relative to the double-stranded target DNA; iv) a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA; and v) a tag sequence comprising a region complementary to the PBS and / or the editing template.

[0160] In some embodiments, the PEGRNA comprises, in 5' to 3' order, a spacer, a gRNA core, an editing template, a PBS, and a tag sequence.

[0161] In some embodiments, the PEGRNA comprises, in 5' to 3' order, an editing template, a PBS, a tag sequence, a spacer, and a gRNA core.

[0162] In certain embodiments, the PE gRNA provided herein comprises, in 5' to 3' order: i) a spacer comprising a region complementary to the sequence to be probed in the target strand of a double-stranded target DNA; ii) a 5' portion of a guide RNA (gRNA) core; iii) an editing template comprising the intended edit relative to the double-stranded target DNA; iv) a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA; and v) a 3' portion of the gRNA core. In some embodiments, the 5' portion of the gRNA core and the 3' portion of the gRNA core form a complete, functional gRNA core capable of binding to a programmable DNA-binding protein (e.g., Cas9 nickase) of a prime editor. In some embodiments, the 5' portion of the gRNA core comprises a repeat sequence, a first stem-loop, and the 5' half of the second stem-loop. In some embodiments, the 3' portion of the gRNA core comprises the 3' half of the second stem-loop and the third stem-loop. In some embodiments, the PEGRNA further comprises a tag sequence that includes a region complementary to the PBS and / or editing template.

[0163] In certain embodiments, the PE gRNA provided herein comprises: i) a first sequence comprising a spacer comprising a region complementary to the sequence to be probed in the target strand of the double-stranded target DNA and a first half of a gRNA core; and ii) a second sequence comprising a second half of the RNA core, an editing template comprising the intended edit relative to the double-stranded target DNA, and a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA, wherein the gRNA core comprises a repeat sequence, a first stem-loop, and a second stem-loop. In certain embodiments, the PE gRNA provided herein comprises: i) a first sequence comprising an editing template comprising an intended edit relative to the double-stranded target DNA; a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA; a spacer comprising a region complementary to a sequence to be probed in the target strand of the double-stranded target DNA; and a first half of a gRNA core; and ii) a second sequence comprising a second half of the gRNA core, wherein the gRNA core comprises a repeat sequence, a first stem-loop, and a second stem-loop. In some embodiments, the first half of the gRNA core comprises a repeat sequence, the first stem-loop, and the 5' half of the second stem-loop. In some embodiments, the second half of the gRNA core comprises the 3' half of the second stem-loop and a third stem-loop. In some embodiments, the first half of the gRNA core comprises the first half of the repeat sequence. In some embodiments, the second half of the gRNA core comprises the second half of the repeat, the first stem loop, the second stem loop, and the third stem loop.

[0164] In some embodiments, the first sequence is on a first molecule and the second sequence is on a second molecule.

[0165] In some embodiments, the first sequence and the second sequence are on the same molecule.

[0166] In some embodiments, the gRNA core first half and the gRNA core second half are selected from pairs of gRNA core first half and second half sequences shown in Table 2.

[0167] In some embodiments, the present specification provides exemplary sequences of a PEGRNA spacer, PBS, RTT, and ngRNA spacer for a prime editing system comprising a nuclease that recognizes the PAM sequence "NGG." In some embodiments, the PAM motif on the edited strand comprises an "NGG" motif, where N is any nucleotide. In some embodiments, a PEGRNA of the present disclosure is part of a prime editing system that recognizes the PAM motif CGG. In some embodiments, a PEGRNA of the present disclosure is part of a prime editing system that recognizes the PAM motif AGG.

[0168] Modified gRNA core In some embodiments, the gRNA core of the PEgRNA binds to a programmable DNA-binding domain in the prime editor. In some embodiments, the gRNA core comprises a repeat sequence, a first stem-loop, and a second stem-loop. In some embodiments, the gRNA core further comprises a third stem-loop. The guide RNA core of the PEgRNA (also referred to as the gRNA core, gRNA scaffold, or gRNA backbone sequence) can comprise a polynucleotide sequence that binds to the DNA-binding domain (e.g., Cas9) of the prime editor. The gRNA core can interact with the prime editor, for example, by binding to the DNA-binding domain, such as a DNA nickase, of the prime editor, as described herein.

[0169] Those skilled in the art will recognize that different prime editors with different DNA-binding domains from different DNA-binding proteins may require different gRNA core sequences specific to the DNA-binding protein. In some embodiments, the gRNA core can bind to a Cas9-based prime editor. In some embodiments, the gRNA core can bind to a Cpf1-based prime editor. In some embodiments, the gRNA core can bind to a Cas12b-based prime editor.

[0170] In some embodiments, the gRNA core comprises a region and secondary structure involved in binding to a specific CRISPR Cas protein. For example, in a Cas9-based prime editing system, the gRNA core of a PE gRNA can comprise one or more base-pairing regions. In some embodiments, a gRNA core capable of binding to Cas9 comprises, from 5' to 3', a repeat sequence, a loop structure, an anti-repeat sequence, a first stem-loop, a second stem-loop, and a third stem-loop. An exemplary structure of a gRNA core is shown in Figure 8, and the sequence in Figure 8 is a standard SpCas9 sgRNA scaffold. As used herein, the repeat sequence and anti-repeat sequence refer to the nucleic acid secondary structure formed by the repetitive sequence regions formed by base pairing between the sequences corresponding to the crRNA and tracrRNA of the Cas9 guide RNA. The repeat and anti-repeat sequences are connected by a loop structure, and the secondary structure formed by base pairing between the repeat and anti-repeat sequences may be referred to as a repeat region (alternatively, the repeat, anti-repeat, and connecting loop structure may be referred to as a tetraloop). In some embodiments, the repeat region of the gRNA core includes one or more base-paired regions: a base-paired "lower stem" adjacent to the spacer sequence (G1-A6 and U25-U30 in Figure 8 ), and a base-paired "upper stem" following the lower stem (G9-A12 and U17-C20 in Figure 8 ), and the lower and upper stems may be connected by a "bulge" containing unpaired RNA. As used herein, the location of modifications in the gRNA core may be referred to in the context of the secondary structure of the gRNA core. For example, "the first base pair of the repeat sequence (or lower stem)" refers to the base pair between the 5'-most nucleotide of the repeat sequence and the complementary nucleotide that is the 3'-most nucleotide of the anti-repeat sequence (G1 and A30 in Figure 8), and "the second base pair of the repeat sequence (or lower stem)" refers to the base pair between the second 5'-most nucleotide of the repeat sequence and the complementary nucleotide of the anti-repeat sequence (U2 and A29 in Figure 8).Similarly, the "start" or "leading" base pair of a second stem-loop refers to the base pair formed between the 5'-terminal nucleotide of the second stem-loop and the complementary nucleotide of the complementary portion of the second stem-loop (A49 and U60 in Figure 8). The "end" or "last" base pair of a second stem-loop refers to the base pair formed between the 3'-most nucleotide of the 5'-portion of the stem and the complementary nucleotide of the complementary 3'-portion of the stem (U52 and A57 in Figure 8), when the second stem-loop is formed by base pairing between the 5'-portion of the stem connected by the loop and the 3'-portion of the stem.

[0171] The gRNA core may further comprise a first stem-loop, a second stem-loop, and a third stem-loop 3' to the repeat sequence. In some embodiments, the gRNA core may comprise a repeat sequence and at least one, at least two, or at least three stem-loops. As used herein, a stem-loop (or hairpin loop) is a base-pairing pattern that can occur in a single-stranded nucleic acid. In some embodiments, a stem-loop can form when two regions of the same nucleic acid strand are at least partially complementary in nucleotide sequence when read in opposite directions; thus, the base pairs can form a double helix containing an unpaired loop. The stem-loops within the gRNA core described herein may be numbered from the 5' end to the 3' end of the gRNA core. For example, the "first stem-loop" refers to the first stem-loop (not including any repeat sequence) at the 5' end of the gRNA core sequence, close to the repeat sequence. A "second stem loop" refers to a second stem loop (not including any repeat sequences) that continues in the 5' to 3' direction from the first stem loop.

[0172] In some embodiments, the gRNA core comprises a nucleotide alteration compared to a wild-type gRNA core, e.g., a standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16. For example, in some embodiments, one or more nucleotides are deleted, inserted, and / or substituted in the gRNA core compared to the standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16. In some embodiments, the gRNA core of the PE gRNA is capable of binding to Cas9 (e.g., nCas9) in a prime editor and comprises one or more nucleotide alterations or modifications compared to a wild-type CRISPR-Cas9 guide RNA scaffold (e.g., a standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16). In some embodiments, the gRNA core comprises one or more nucleotide insertions, deletions, and / or substitutions in the repeat sequence compared to the standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16. Potential advantages associated with such modified gRNA cores include increased prime editing efficiency and improved manufacturing via split-synthesis schemes.

[0173] In some embodiments, the gRNA core comprises one or more nucleotide insertions, deletions, and / or substitutions in the lower or upper stem of the repeat sequence. In some embodiments, the gRNA core comprises one or more nucleotide substitutions in the lower stem of the repeat sequence. In some embodiments, the gRNA core comprises one or more nucleotide insertions in the upper stem of the repeat sequence. In some embodiments, the gRNA core comprises one or more nucleotide insertions, deletions, and / or substitutions in the first stem loop compared to the standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16. In some embodiments, the gRNA core comprises one or more nucleotide insertions, deletions, and / or substitutions in the second stem loop compared to the standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16. In some embodiments, the gRNA core comprises one or more nucleotide insertions in the second stem loop. In some embodiments, the gRNA core comprises one or more nucleotide insertions, deletions, and / or substitutions in the third stem loop compared to the standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16. In some embodiments, the gRNA core comprises one or more nucleotide insertions, deletions, and / or substitutions compared to a wild-type CRISPR-Cas9 guide RNA scaffold, e.g., the standard SpCas9 gRNA scaffold set forth in SEQ ID NO: 16, and comprises a third stem loop having the same sequence as the third stem loop of the wild-type CRISPR-Cas9 guide RNA scaffold.

[0174] In some embodiments, the RNA nucleotides in the lower stem, upper stem, and / or stem-loop regions can be replaced with one or more DNA sequences. In some embodiments, the gRNA core comprises an unmodified or wild-type RNA sequence in the nexus region and / or bulge region. In some embodiments, the gRNA core does not contain long AU pairs, such as GUUUU-AAAAC pairing elements. Exemplary gRNA core structures are shown in Figures 8 and 12.

[0175] In some embodiments, the PEgRNA comprises a guide RNA (gRNA) core associated with a DNA-binding domain of a prime editor (e.g., a CRISPR-Cas protein domain). In some embodiments, the PEgRNA comprises a guide RNA (gRNA) core associated with a DNA-binding domain of a prime editor (e.g., a Cas9 domain). In certain aspects, the gRNA core of a PEgRNA provided herein comprises one or more sequence modifications compared to a PEgRNA. In some embodiments, the one or more (e.g., two or more, three or more, four or more, or five or more) sequence modifications comprise the gRNA core differences listed in Table 1 or Table 2. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 16-61. In some embodiments, the gRNA core comprises a first gRNA core sequence comprising the 5' half of the gRNA core and a second gRNA core sequence comprising the 3' half of the gRNA core, and the PEgRNA comprises, in 5' to 3' order, a spacer, the first gRNA core sequence, an editing template, a PBS, a tag sequence, and the second gRNA core sequence. The 5' and 3' halves can form a functional gRNA core for association / binding with a programmable DNA-binding protein (e.g., a Cas protein). One of skill in the art will recognize that different prime editors with different DNA-binding domains from different DNA-binding proteins may require different gRNA core sequences specific to the DNA-binding protein. In some embodiments, the gRNA core can bind to a Cas9-based prime editor. In some embodiments, the gRNA core can bind to a Cpf1-based prime editor. In some embodiments, the gRNA core can bind to a Cas12b-based prime editor.

[0176] In some embodiments, the gRNA core of a PE gRNA provided herein comprises one or more sequence modifications compared to SEQ ID NO: 16. In some embodiments, the one or more sequence modifications comprise alterations of the gRNA core compared to SEQ ID NO: 16 as set forth in Table 1. In some embodiments, the gRNA core comprises a gRNA core sequence as set forth in Table 1 or Table 2.

[0177] In some embodiments, the one or more sequence modifications include sequence modifications in a repeat sequence. In some embodiments, the sequence modification in the gRNA core of a PEGRNA includes an inversion of one or more nucleotides. As used herein, the term "inversion" refers to altering a sequence so that nucleotide bases that base pair with each other in the stem of a loop or hairpin structure are interchanged. For example, the original, unmodified stem structure may contain an A / U base pair, with A in the first strand (or region) of the stem structure and U in the complementary strand (or region). In an A / U to U / A base pair inversion, adenosine in the first strand (or region) is replaced with uracil, and uracil in the complementary strand (or region) is replaced with adenosine, thereby "inverting" the A / U base pair to a U / A base pair. In some embodiments, nucleotide inversions can be used, for example, to resolve sequences containing repeats of the same base present within a nucleic acid molecule (e.g., a sequence of at least 3, 4, 5, 6, or 7 consecutive A, U, C, or G nucleotides) without disrupting its secondary structure. An example of an A / U inversion, which separates four consecutive A and U nucleotides at the fourth position of the lower stem of a repeat sequence without disrupting the secondary structure of the gRNA core, is shown in Figure 12. In some embodiments, instead of an inversion, the original base pair is replaced with an alternative base pair (e.g., an A / U base pair is replaced with a C / G or G / C base pair).

[0178] In some embodiments, the repeat sequence of the gRNA core can include at least one inversion of an A / U base pair in the lower stem of the repeat sequence, optionally wherein the lower stem does not contain two, three, four or more consecutive AU base pairs, and / or wherein the at least one inversion of an A / U base pair in the repeat sequence includes an inversion of the fourth A / U base pair in the lower stem of the repeat sequence.

[0179] In some embodiments, the sequence modification in the repeat sequence comprises an insertion of one or more nucleotides into the upper stem of the repeat sequence of the gRNA core, resulting in an extension of the upper stem compared to the wild-type gRNA core, for example, as shown in SEQ ID NO: 16. The upper stem extension can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 base pairs. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 26-37.

[0180] In some embodiments, the one or more sequence modifications comprise a sequence modification in the second stem loop.

[0181] In some embodiments, the modification of the second stem loop comprises inverting a G / C base pair. In some embodiments, the modification in the second stem loop comprises inverting an A / U base pair in the second stem loop. In some embodiments, the modification in the second stem loop comprises substituting an A / U base pair with a G / C base pair. In some embodiments, the modification in the second stem loop comprises substituting a U / A base pair with a G / C base pair. In some embodiments, the modification in the second stem loop comprises substituting an A / U base pair with a G / C base pair, and further comprises substituting a U / A base pair with a G / C base pair. In some embodiments, the gRNA core comprises a nucleic acid sequence selected from SEQ ID NO: 21, 22, or 25.

[0182] Exemplary gRNA core sequences and sequence modifications are shown in Tables 1 and 2. In some embodiments, the gRNA core comprises a sequence selected from SEQ ID NOs: 16-61, 3860-4359, and 4452.

[0183] In some embodiments, the one or more sequence modifications comprise a modification in the third stem loop of the gRNA core. In some embodiments, the modification of the third stem loop comprises an inversion of a G / C base pair. In some embodiments, the modification of the third stem loop comprises an inversion of an A / U base pair.

[0184] The gRNA core may include any one of the modifications listed in Table 1 or Table 2, or any combination thereof.

[0185] In some embodiments, the gRNA core has an inverted first AU base pair in the repeat sequence. In some embodiments, the gRNA core has an inverted second AU base pair in the repeat sequence. In some embodiments, the gRNA core has an inverted third AU base pair in the repeat sequence. In some embodiments, the gRNA core has an inverted fourth AU base pair in the repeat sequence.

[0186] In some embodiments, the gRNA core has an AU base pair (bp) substituted with a GC bp at the fourth base pair of the second stem loop. In some embodiments, the gRNA core has an AU bp substituted with a CG bp at the fourth base pair of the second stem loop.

[0187] In some embodiments, the gRNA core contains a 5-base pair extension of the upper stem of the repeat sequence (tgctg and cagca). In some embodiments, the gRNA has an "inversion and extension" (M4 and E5) as described by Nelson, JW, Randolph, PB, Shen, SP et al., designed pegRNAs to improve prime editing efficiency. Nat Biotechnol (2021). The M4 modification inverts the fourth AU base pair within the repeat sequence of the gRNA core. The E5 modification extends the end of the upper stem of the repeat sequence with 5 bp of the sequence (tgctg and cagca).

[0188] In some embodiments, the gRNA core comprises an M4 modification. In some embodiments, the gRNA core comprises an E5 modification. In some embodiments, the gRNA core comprises an M4 modification and an E5 modification.

[0189] In some embodiments, the gRNA core comprises a substitution of an A / U base pair with a G / C base pair in the second stem loop, hi some embodiments, the gRNA core comprises a substitution of an A / U base pair with a G / C base pair at the first base pair of the second stem loop.

[0190] In some embodiments, the gRNA core has a 1 base pair extension on the upper stem of the repeat sequences (c and g). In some embodiments, the gRNA core has a 2 base pair extension on the upper stem of the repeat sequences (cc and gg). In some embodiments, the gRNA core has a 2 base pair extension on the upper stem of the repeat sequences (ca and tg). In some embodiments, the gRNA core has a 2 base pair extension on the upper stem of the repeat sequences (cg and tg). In some embodiments, the gRNA core has a 1 base pair extension on the upper stem of the repeat sequences (a and t). In some embodiments, the gRNA core has a 2 base pair extension on the upper stem of the repeat sequences (ac and gt). In some embodiments, the gRNA core has a 2 base pair extension on the upper stem of the repeat sequences (aa and tt). In some embodiments, the gRNA core has a 2 base pair extension on the upper stem of the repeat sequences (ag and tt). In some embodiments, the gRNA core has a 3 base pair extension on the upper stem of the repeat sequences (ccc and ggg). In some embodiments, the gRNA core has a 4 base pair extension on the upper stem of the repeat sequences (ccac and gtgg). In some embodiments, the gRNA core has a 5 base pair extension on the upper stem of the repeat sequences ccaac and gttgg. In some embodiments, the gRNA core has a 6 base pair extension on the upper stem of the repeat sequences (ccacac and gtgtgg).

[0191] In some embodiments, the gRNA core has a 1 base pair extension in the second stem-loop sequence (c and g). In some embodiments, the gRNA core has a 2 base pair extension in the second stem-loop sequence (cc and gg). In some embodiments, the gRNA core has a 2 base pair extension in the second stem-loop sequence (ca and tg). In some embodiments, the gRNA core has a 2 base pair extension in the second stem-loop sequence (cg and tg). In some embodiments, the gRNA core has a 1 base pair extension in the second stem-loop sequence (a and t). In some embodiments, the gRNA core has a 2 base pair extension in the second stem-loop sequence (ac and gt). In some embodiments, the gRNA core has a 2 base pair extension in the second stem-loop sequence (aa and tt). In some embodiments, the gRNA core has a 2 base pair extension in the second stem-loop sequence (ag and tt). In some embodiments, the gRNA core has a 3 base pair extension in the second stem-loop sequence (ccc and ggg). In some embodiments, the gRNA core has a 4 base pair extension in the second stem-loop sequence (ccac and gtgg). In some embodiments, the gRNA core has a 5 base pair extension in the second stem-loop sequence (ccaac and gttgg). In some embodiments, the gRNA core has a 6 base pair extension in the second stem-loop sequence (ccacac and gtgtgg).

[0192] In some embodiments, the gRNA core has a 1 base pair extension at the third stem-loop sequence (c and g). In some embodiments, the gRNA core has a 2 base pair extension at the third stem-loop sequence (cc and gg). In some embodiments, the gRNA core has a 2 base pair extension at the third stem-loop sequence (ca and tg). In some embodiments, the gRNA core has a 2 base pair extension at the third stem-loop sequence (cg and tg). In some embodiments, the gRNA core has a 1 base pair extension at the third stem-loop sequence (a and t). In some embodiments, the gRNA core has a 2 base pair extension at the third stem-loop sequence (ac and gt). In some embodiments, the gRNA core has a 2 base pair extension at the third stem-loop sequence (aa and tt). In some embodiments, the gRNA core has a 2 base pair extension at the third stem-loop sequence (ag and tt). In some embodiments, the gRNA core has a 3 base pair extension at the third stem-loop sequence (ccc and ggg). In some embodiments, the gRNA core has a 4 base pair extension at the third stem-loop sequence (ccac and gtgg). In some embodiments, the gRNA core has a 5 base pair extension at the third stem-loop sequence (ccaac and gttgg). In some embodiments, the gRNA core has a 6 base pair extension at the third stem-loop sequence (ccacac and gtgtgg).

[0193] In some embodiments, modifications to the gRNA core increase editing efficiency by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 110%, at least 120%, at least 130%, at least 140%, at least 150%, at least 160%, at least 170%, at least 180%, at least 190%, or at least 200% compared to the editing efficiency of a control PEgRNA having an unmodified gRNA core. Exemplary nucleotide sequence modifications in the gRNA core of a PEgRNA are shown in Table 1. Modifications compared to the standard SpCas9 gRNA scaffold sequence are shown in the third column ("Modification Description"). The gRNA core sequences shown in Table 1 are RNA sequences; however, for consistency with the ST.26 standard, "T" is used in the sequences instead of "U." [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8]

[0194] Nucleic acid part In some embodiments, the PEGRNA comprises one or more nucleic acid moieties (e.g., hairpins, pseudoknots, quadruplexes, tRNA sequences, aptamers, etc.) in addition to the spacer, gRNA core, primer binding site, and editing template. In some embodiments, such nucleic acid moieties are located at the 3' end of the PEGRNA.

[0195] In some embodiments, the nucleic acid portion comprises a hairpin. In some embodiments, a hairpin is a secondary structure of a nucleic acid formed by intramolecular base pairing between two regions of the same strand, typically resulting in complementary nucleotide sequences when read in opposite directions. The two regions base pair to form a double helix, terminating in an unpaired loop. As described herein, a hairpin can be 5-50 nucleotides, 10-40 nucleotides, or at least 15-30 nucleotides long. A hairpin can be at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, or at least 30 nucleotides long. In some embodiments, a hairpin is 14 nucleotides long. In some embodiments, a hairpin is 18 nucleotides long. In some embodiments, a hairpin is 22 nucleotides long. In some embodiments, a hairpin contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more consecutive complementary base pairs. In some embodiments, the hairpin comprises 4, 5, 6, 7, 8, 9, or 10 consecutive complementary base pairs. In some embodiments, the hairpin comprises 4 to 8 consecutive complementary base pairs. In some embodiments, the hairpin comprises 5 consecutive complementary base pairs. In some embodiments, the hairpin comprises 7 consecutive complementary base pairs.

[0196] In some embodiments, the nucleic acid portion comprises a pseudoknot. As used herein, pseudoknot includes, but is not limited to, a nucleic acid secondary structure comprising at least two stem-loop structures in which one half of one stem is inserted between the two halves of another stem. Pseudoknots exist in several different folding topologies, including H-type. In an H-type fold, bases within the loop of the hairpin form intramolecular pairs with bases outside the stem. This forms a second stem and loop, resulting in a pseudoknot with two stems and two loops. As described herein, pseudoknots can be 5-50 nucleotides long, 10-40 nucleotides long, or at least 15-30 nucleotides long. Hairpins can be at least 10 nucleotides long, at least 15 nucleotides long, at least 20 nucleotides long, at least 25 nucleotides long, or at least 30 nucleotides long. In some embodiments, the pseudoknot is 22 nucleotides long.

[0197] In some embodiments, the nucleic acid moiety comprises a quadruplex. In some embodiments, a quadruplex is a non-canonical, four-stranded nucleic acid secondary structure that can be formed with guanine-rich or cysteine-rich DNA and RNA sequences, depending on the context. As described herein, a quadruplex can be 5-50 nucleotides in length, 10-40 nucleotides in length, or at least 15-30 nucleotides in length. A hairpin can be at least 10 nucleotides in length, at least 15 nucleotides in length, at least 20 nucleotides in length, at least 25 nucleotides in length, or at least 30 nucleotides in length. In some embodiments, a quadruplex is 18 nucleotides in length. In some embodiments, a quadruplex is guanine-rich (G-quadruplex). In some embodiments, a quadruplex is cytosine-rich (C-quadruplex).

[0198] In some embodiments, the nucleic acid portion comprises an aptamer. In some embodiments, the aptamer comprises a short, single-stranded nucleic acid oligomer capable of binding to a specific target molecule. Aptamers can assume a variety of shapes due to their tendency to form helices and single-stranded loops. As described herein, aptamers can be 5-50 nucleotides in length, 10-40 nucleotides in length, or at least 15-30 nucleotides in length. Hairpins can be at least 10 nucleotides in length, at least 15 nucleotides in length, at least 20 nucleotides in length, at least 25 nucleotides in length, or at least 30 nucleotides in length. In some embodiments, the aptamer is 19 nucleotides in length. In some embodiments, the aptamer is 33 nucleotides in length.

[0199] In some embodiments, the nucleic acid portion comprises a tRNA sequence. The tRNA sequence can be long (e.g., at least 25 nucleotides, at least 30 nucleotides, at least 35 nucleotides, at least 40 nucleotides, at least 45 nucleotides, at least 50 nucleotides, at least 55 nucleotides, at least 60 nucleotides, at least 65 nucleotides, at least 70 nucleotides, or at least 75 nucleotides). In some embodiments, the tRNA sequence can be short (less than 25 nucleotides, less than 20 nucleotides, less than 15 nucleotides, or less than 10 nucleotides). As described herein, the tRNA sequence can be 5-80 nucleotides, 10-70 nucleotides, or at least 15-60 nucleotides long. The hairpin can be at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, or at least 70 nucleotides long. In some embodiments, the aptamer is 18 nucleotides long. In some embodiments, the aptamer is 61 nucleotides long.

[0200] Exemplary moieties are set forth in Table 4. Those skilled in the art will appreciate that the present disclosure is not limited to the sequences and structures of Table 4, as the configurations in Table 4 are examples of a broader class of moieties included in the present disclosure.

[0201] In some embodiments, the one or more nucleic acid moieties comprise a hairpin (e.g., the hairpin comprises a region of self-complementarity, optionally the region of self-complementarity comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more consecutive complementary base pairs), a quadruplex (e.g., a G-quadruplex or C-quadruplex, optionally the G-quadruplex or C-quadruplex is derived from a VEGF gene promoter), a tRNA sequence (e.g., a tRNA sequence, optionally the tRNA sequence is a tRNA(proline) sequence), an aptamer (e.g., an aptamer derived from a viral protein binding sequence, optionally the aptamer comprises a viral reverse transcriptase recruitment sequence, optionally the aptamer comprises an MS2 protein binding sequence or a Moloney murine leukemia (MMLV) reverse transcriptase recruitment sequence), and / or a pseudoknot (e.g., the pseudoknot is derived from potato leaf roll virus (PLRV)), or any combination thereof.

[0202] In some embodiments, one or more nucleic acid moieties comprise a structure derived from a retroviral replication recognition sequence. In some embodiments, the nucleic acid moieties comprise a sequence derived from a Moloney Murine Leukemia Virus (MMLV) replication recognition sequence. In some embodiments, one or more nucleic acid moieties comprise a nucleic acid sequence selected from SEQ ID NOs: 12-15.

[0203] In some embodiments, one or more nucleic acid moieties comprise a hairpin, hi some embodiments, the hairpin comprises the sequence of any one of SEQ ID NOs: 1-3 or 5-7.

[0204] In some embodiments, the one or more nucleic acid moieties comprise a pseudoknot. In some embodiments, the pseudoknot is derived from Potato Leafroll Virus. In some embodiments, the pseudoknot comprises the sequence of SEQ ID NO: 4. In some embodiments, the one or more nucleic acid moieties comprise an MS2 hairpin. In some embodiments, the nucleotide sequence of the MS2 hairpin (also referred to as the "MS2 aptamer") is GCCAACATGAGGATCACCCATGTCTGCAGGGCC (SEQ ID NO: 4446). In some embodiments, the nucleotide sequence of the MS2 aptamer comprises the sequence of SEQ ID NO: 9. In some embodiments, an MS2 coat protein (MCP) recognizes the MS2 hairpin. In some embodiments, the amino acid sequence of the MCP is GSASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKVATQTVGGEELPVAGWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIA ANSGIY (SEQ ID NO: 4447).

[0205] In some embodiments, one or more nucleic acid moieties comprise a G-quadruplex or a C-quadruplex. In some embodiments, one or more nucleic acid moieties comprise a quadruplex from a VEGF gene promoter. In some embodiments, the quadruplex comprises the sequence of SEQ ID NO: 10 or 11.

[0206] In some embodiments, the PEgRNA comprises one or more nucleic acid moieties at the 3' end. In some embodiments, the PEgRNA comprises one or more nucleic acid moieties at the 5' end. [Table 3] [Table 4-1] [Table 4-2]

[0207] Tag sequence In some embodiments, the PEG RNA comprises a tag sequence in addition to the spacer, gRNA core, primer binding site, and editing template. In some embodiments, the tag sequence comprises a region complementary to the editing template. In some embodiments, the tag sequence comprises a region complementary to the PBS. In some embodiments, the tag sequence comprises a region complementary to the editing template and / or the PBS. In some embodiments, the tag sequence comprises a region complementary to the editing template but does not have substantial complementarity to the PBS. In some embodiments, the tag sequence comprises a region complementary to the editing template but is not complementary to the PBS. In some embodiments, the tag sequence and the editing template each comprise a region complementary to each other, and the 3' end of the complementary region of the editing template is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more bases from the 5' end of the 3' half of the editing template. In some embodiments, the complementary region of the tag sequence is in the 5' portion of the tag sequence. In some embodiments, the tag sequence has no substantial complementarity to the spacer. In some embodiments, the tag has no complementarity to the spacer. In some embodiments, the tag sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. In some embodiments, the tag sequence is at least 4, at least 6, or at least 8 nucleotides in length. In some embodiments, the tag sequence comprises a nucleic acid sequence selected from SEQ ID NOs: 62-1960. Exemplary tag sequences are listed in Table 5. [Table 5-1] [Table 5-2] [Table 5-3] [Table 5-4] [Table 5-5] Table 5-6 Table 5-7 Table 5-8 Table 5-9 Table 5-10 Table 5-11 Table 5-12 Table 5-13 Table 5-14 Table 5-15 Table 5-16 Table 5-17 Table 5-18 Table 5-19 Table 5-20 Table 5-21 Table 5-22 Table 5-23 Table 5-24 Table 5-25 Table 5-26 Table 5-27 Table 5-28 Table 5-29 Table 5-30 Table 5-31 Table 5-32 Table 5-33 Table 5-34 Table 5-35 Table 5-36 Table 5-37 Table 5-38 Table 5-39 Table 5-40 Table 5-41 Table 5-42 Table 5-43 Table 5-44 Table 5-45 Table 5-46 Table 5-47 Table 5-48 Table 5-49 Table 5-50 Table 5-51 Table 5-52 Table 5-53 Table 5-54 Table 5-55 Table 5-56 Table 5-57 Table 5-58 Table 5-59 Table 5-60 Table 5-61 Table 5-62 Table 5-63 Table 5-64 Table 5-65 Table 5-66 Table 5-67 Table 5-68 Table 5-69 Table 5-70 Table 5-71 Table 5-72 Table 5-73 Table 5-74 Table 5-75 Table 5-76 Table 5-77 Table 5-78 Table 5-79 Table 5-80 Table 5-81 Table 5-82 Table 5-83 Table 5-84 Table 5-85 Table 5-86 Table 5-87 Table 5-88 Table 5-89 Table 5-90 Table 5-91 Table 5-92 Table 5-93 Table 5-94 Table 5-95 Table 5-96 Table 5-97 Table 5-98 Table 5-99 Table 5-100 Table 5-101 Table 5-102 Table 5-103 Table 5-104 Table 5-105 Table 5-106 Table 5-107 Table 5-108 Table 5-109 Table 5-110 Table 5-111 Table 5-112 Table 5-113 Table 5-114 Table 5-115 Table 5-116 Table 5-117 Table 5-118 Table 5-119 Table 5-120 Table 5-121 Table 5-122 Table 5-123 Table 5-124 Table 5-125 Table 5-126 Table 5-127 Table 5-128 Table 5-129 Table 5-130 Table 5-131 Table 5-132 Table 5-133 Table 5-134 Table 5-135 Table 5-136 Table 5-137 Table 5-138 Table 5-139 Table 5-140 Table 5-141 Table 5-142 Table 5-143 Table 5-144 Table 5-145 Table 5-146 Table 5-147 Table 5-148 Table 5-149 Table 5-150 Table 5-151 Table 5-152 Table 5-153 Table 5-154 Table 5-155 Table 5-156 Table 5-157 Table 5-158 Table 5-159 Table 5-160 Table 5-161 Table 5-162 Table 5-163 Table 5-164 Table 5-165 Table 5-166 Table 5-167 Table 5-168 Table 5-169 Table 5-170 Table 5-171 Table 5-172 Table 5-173 Table 5-174 Table 5-175 Table 5-176 Table 5-177 Table 5-178 Table 5-179 Table 5-180 Table 5-181 Table 5-182 Table 5-183 Table 5-184 Table 5-185 Table 5-186 Table 5-187 Table 5-188 Table 5-189 Table 5-190 Table 5-191 Table 5-192 Table 5-193 Table 5-194 Table 5-195 Table 5-196 Table 5-197 Table 5-198 Table 5-199 Table 5-200 Table 5-201 Table 5-202 Table 5-203 Table 5-204 Table 5-205 Table 5-206 Table 5-207 Table 5-208 Table 5-209 Table 5-210 Table 5-211 Table 5-212 Table 5-213 Table 5-214 Table 5-215 Table 5-216 Table 5-217 Table 5-218 Table 5-219 Table 5-220 Table 5-221 Table 5-222 Table 5-223 Table 5-224 Table 5-225 Table 5-226 Table 5-227 Table 5-228 Table 5-229 Table 5-230 Table 5-231 Table 5-232 Table 5-233 Table 5-234 Table 5-235 Table 5-236 Table 5-237 Table 5-238 Table 5-239 Table 5-240 Table 5-241 Table 5-242 Table 5-243 Table 5-244 Table 5-245 Table 5-246 Table 5-247 Table 5-248 Table 5-249 Table 5-250 Table 5-251 Table 5-252 Table 5-253 Table 5-254 Table 5-255 Table 5-256 [Table 5-257] [Table 5-258] [Table 5-259] [Table 5-260] [Table 5-261] [Table 5-262] [Table 5-263]

[0208] Linker In some embodiments, the PEG-RNA comprises a linker. In some embodiments, the linker is i) immediately 5' of one or more nucleic acid moieties, ii) immediately 5' of the tag sequence, iii) immediately 3' of the tag sequence, iv) immediately 3' of the spacer, v) immediately 5' of the spacer, vi) immediately 3' of the gRNA core, or vii) immediately 5' of the gRNA core. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. In some embodiments, the linker is 2-12 nucleotides in length. In some embodiments, the linker is 5-20 nucleotides in length. In some embodiments, the linker is 3-10, 3-15, 3-20, 3-25, 3-30, 3-35, 3-40, or 3-50 nucleotides in length. In some embodiments, the linker is 8 nucleotides in length. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker has no region complementary to the PBS sequence. In some embodiments, the linker has no region complementary to the editing template. In some embodiments, the linker comprises a sequence selected from SEQ ID NOs: 1961-3859. As used herein, a linker is any chemical group or molecule that links two molecules / moieties, such as components of a PEG RNA.

[0209] LegRNA The present specification also provides legRNAs. In some embodiments, the PEGRNA is a legRNA. As used herein, "legRNA" refers to a PEGRNA comprising a spacer, a gRNA core, a PBS, and an editing template (e.g., an RTT sequence), wherein the PBS and editing template are located within the gRNA core. The legRNAs disclosed herein may include any 3' portion or other modification disclosed herein.

[0210] In certain embodiments, the legRNA comprises, in 5' to 3' order: i) a spacer comprising a region complementary to the sequence to be probed in the target strand of the double-stranded target DNA; ii) a 5' portion of a guide RNA (gRNA) core; iii) an editing template comprising the intended edit relative to the double-stranded target DNA; iv) a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA; and v) a 3' portion of the gRNA core. In some embodiments, the 5' portion of the gRNA core comprises a repeat sequence, a first stem-loop, and a 5' half of the second stem-loop. In some embodiments, the 3' portion of the gRNA core comprises the 3' half of the second stem-loop and a third stem-loop. In some embodiments, the 5' portion of the gRNA core and the 3' portion of the gRNA core are "split" between the 30th and 31st, 31st and 32nd, 32nd and 33rd, 33rd and 34th, 34th and 35th, 35th and 36th, 36th and 37th, 37th and 38th, 38th and 39th, or 39th and 40th nucleotides of the complete gRNA core sequence, where the nucleotide position numbers are as set forth in SEQ ID NO:16. In some embodiments, the 5' portion of the gRNA core and the 3' portion of the gRNA core are "split" between nucleotides 50 and 51, 51 and 52, 52 and 55, 55 and 54, 54 and 55, 55 and 56, 56 and 57, 57 and 58, 58 and 59, or 59 and 60 of the complete gRNA core sequence, where the nucleotide position numbering is as set forth in SEQ ID NO: 16. In some embodiments, the 5' portion of the gRNA core and the 3' portion of the gRNA core are split between nucleotides 54 and 55 of the complete gRNA core sequence, where the nucleotide position numbering is as set forth in SEQ ID NO: 16. In some embodiments, the 5' portion of the gRNA core comprises the sequence GTTTAAGAGCTAGAAATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAGCGTGA. In some embodiments, the 3' portion of the gRNA core comprises the sequence AAACGCGGCACCGAGTCGGTGC.

[0211] The data is shown in Table 6 below.

[0212] In some embodiments, the PEGRNA further comprises a tag sequence that includes a region complementary to the PBS and / or editing template.

[0213] The legRNA can comprise a tag sequence, an aptamer, a hairpin, a quadruplex, a tRNA, a pseudoknot, a linker, or any nucleic acid moiety described herein. In some embodiments, the legRNA comprises a linker. In some embodiments, the linker is i) immediately 5' of one or more nucleic acid moieties, ii) immediately 5' of the tag sequence, iii) immediately 3' of the tag sequence, iv) immediately 3' of the spacer, v) immediately 5' of the spacer, vi) immediately 3' of the gRNA core, and / or vii) immediately 5' of the gRNA core. In some embodiments, the linker is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 nucleotides in length. In some embodiments, the linker does not form a secondary structure. In some embodiments, the linker has no region complementary to the PBS sequence. In some embodiments, the linker has no region complementary to the editing template. In some embodiments, the linker comprises a nucleic acid sequence selected from SEQ ID NOs: 1961-3859. As used herein, a linker is any chemical group or molecule that connects two molecules or moieties, such as components of a legRNA. [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5] [Table 6-6] [Table 6-7]

[0214] Expanded gRNA core In some embodiments, the PE gRNA comprises a gRNA core that contains one or more nucleotide insertions compared to the wild-type CRISPR guide RNA scaffold sequence (e.g., a standard SpCa9 guide RNA scaffold), i.e., a length-extended gRNA core. Potential advantages associated with such extended gRNA cores include increased prime editing efficiency and improved manufacturing via split-synthesis schemes.

[0215] In some embodiments, the gRNA core comprises one or more nucleotide insertions within the repeat sequence compared to the wild-type CRISPR guide RNA scaffold sequence set forth in SEQ ID NO: 16. In some embodiments, the gRNA core comprises one or more nucleotide insertions in the second stem loop compared to the standard SpCas9 guide RNA scaffold sequence set forth in SEQ ID NO: 16. Exemplary extended gRNA cores are shown in Tables 1 and 2. While the gRNA core sequences shown in Tables 1 and 2 are RNA sequences, the sequences use "T" instead of "U" for consistency with the ST.26 standard.

[0216] Components of a PEgRNA, e.g., an extended PEgRNA, can be synthesized by split synthesis, which refers to separately synthesizing two (or more) portions of the PEgRNA (e.g., the 5' half of the PEgRNA and the 3' half of the PEgRNA) and then ligating the first half to the second half to form the full-length PEgRNA. Exemplary "split" locations between the 5' and 3' halves, e.g., within the repeat sequence or second stem loop of the gRNA core, are shown in Figure 8. Exemplary gRNA core sequences and corresponding first and second halves for split synthesis are shown in Table 2.

[0217] In certain embodiments, the PE gRNA provided herein comprises: i) a first sequence comprising a spacer comprising a region complementary to the sequence to be probed in the target strand of the double-stranded target DNA and a first half of a gRNA core; and ii) a second sequence comprising a second half of the RNA core, an editing template comprising the intended edit relative to the double-stranded target DNA, and a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA, wherein the gRNA core comprises a repeat sequence, a first stem-loop, and a second stem-loop.

[0218] In certain embodiments, the PE gRNA provided herein comprises: i) a first sequence comprising an editing template comprising an intended edit relative to the double-stranded target DNA, a primer binding site (PBS) comprising a region complementary to a region upstream of the nick site in the non-target strand of the double-stranded target DNA, a spacer comprising a region complementary to a sequence to be probed in the target strand of the double-stranded target DNA, and a first half of a gRNA core; and ii) a second sequence comprising a second half of the gRNA core, wherein the gRNA core comprises a repeat sequence, a first stem-loop, and a second stem-loop.

[0219] In some embodiments, the first sequence is on a first RNA molecule and the second sequence is on a second RNA molecule. In some embodiments, the spacer and the first and second sequences are on the same RNA molecule. In some embodiments, the first half of the gRNA core and the second half of the gRNA core are selected from pairs of gRNA core first and second half sequences shown in Table 2.

[0220] It should be noted that the first and second halves of the gRNA core may or may not be equal in length. In some embodiments, the first half of the gRNA core is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or at least 75 nucleotides in length. In some embodiments, the second half of the gRNA core is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or at least 75 nucleotides in length.

[0221] In some embodiments, the first half of the gRNA core is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical to a sequence shown in Table 2. In some embodiments, the first half of the gRNA core is identical to a sequence shown in Table 2. In some embodiments, the second half of the gRNA core is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical to a sequence shown in Table 2. In some embodiments, the second half of the gRNA core is identical to a sequence shown in Table 2.

[0222] As described above, the gRNA core comprises a repeat sequence and / or one or more stem-loops. In some embodiments, the gRNA core synthesized using split synthesis comprises a first half of the gRNA core comprising a first half of the repeat sequence and a second half of the gRNA core comprising a second half of the repeat sequence. In some embodiments, the gRNA core synthesized using split synthesis comprises a first half of the gRNA core comprising a first half of a second stem-loop and a second half of the gRNA core comprising a second half of a second stem-loop. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9] [Table 2-10] [Table 2-11] [Table 2-12] [Table 2-13] Table 2-14 Table 2-15 Table 2-16 Table 2-17 Table 2-18 Table 2-19 Table 2-20 Table 2-21 Table 2-22 Table 2-23 Table 2-24 Table 2-25 Table 2-26 Table 2-27 Table 2-28 Table 2-29 Table 2-30 Table 2-31 Table 2-32 Table 2-33 Table 2-34 Table 2-35 Table 2-36 Table 2-37 Table 2-38 Table 2-39 Table 2-40 Table 2-41 Table 2-42 Table 2-43 Table 2-44 Table 2-45 Table 2-46 Table 2-47 Table 2-48 Table 2-49 Table 2-50 Table 2-51 Table 2-52 Table 2-53 Table 2-54 Table 2-55 Table 2-56 Table 2-57 Table 2-58 Table 2-59 Table 2-60 Table 2-61 Table 2-62 Table 2-63 Table 2-64 Table 2-65 Table 2-66 Table 2-67 Table 2-68 Table 2-69 Table 2-70 Table 2-71 Table 2-72 Table 2-73 Table 2-74 Table 2-75 Table 2-76 Table 2-77 Table 2-78 Table 2-79 Table 2-80 Table 2-81 Table 2-82 Table 2-83 Table 2-84 Table 2-85 Table 2-86 Table 2-87 Table 2-88 Table 2-89 Table 2-90 Table 2-91 Table 2-92 Table 2-93 Table 2-94 Table 2-95 Table 2-96 Table 2-97 Table 2-98 Table 2-99 Table 2-100 Table 2-101 Table 2-102

Table 2-103

Table 2-120

Table 2-130

[0223] Nucleotide editing The present specification provides exemplary PEGRNAs having modifications disclosed herein for nucleotide editing. The intended nucleotide edits in the editing template of the PEGRNA can include various types of changes compared to the target gene sequence. In some embodiments, the nucleotide edit is a single nucleotide substitution compared to the target gene sequence. In some embodiments, the nucleotide edit is a deletion compared to the target gene sequence. In some embodiments, the nucleotide edit is an insertion compared to the target gene sequence. In some embodiments, the editing template includes 1 to 10 intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes one or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes two or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes three or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes four or more, five or more, or six or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes two single nucleotide substitutions, insertions, deletions, or any combination thereof compared to the target gene sequence. In some embodiments, the editing template comprises three single nucleotide substitutions, insertions, deletions, or any combination thereof, compared to the target gene sequence. In some embodiments, the editing template comprises four, five, or six single nucleotide substitutions, insertions, deletions, or any combination thereof, compared to the target gene sequence. In some embodiments, the nucleotide substitution comprises an adenine (A) to thymine (T) substitution. In some embodiments, the nucleotide substitution comprises an A to guanine (G) substitution. In some embodiments, the nucleotide substitution comprises an A to cytosine (C) substitution. In some embodiments, the nucleotide substitution comprises a TA substitution. In some embodiments, the nucleotide substitution comprises a TG substitution. In some embodiments, the nucleotide substitution comprises a TC substitution.In some embodiments, the nucleotide substitution comprises a G to A substitution. In some embodiments, the nucleotide substitution comprises a G to T substitution. In some embodiments, the nucleotide substitution comprises a G to C substitution. In some embodiments, the nucleotide substitution comprises a C to A substitution. In some embodiments, the nucleotide substitution comprises a C to T substitution. In some embodiments, the nucleotide substitution comprises a C to G substitution.

[0224] In some embodiments, the nucleotide insertion is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. In some embodiments, the nucleotide insertion is 1-2, 1-3, 1-4, 1-5, 2-5, 3-5, 3-6, 3-8, 4-9, 5-10, 6-11, 7-12, 8-13, 9-14, 10-15, 11-16, 12-17, 13-18, 14-19, or 15-20 nucleotides in length. In some embodiments, the nucleotide insertion is a single nucleotide insertion. In some embodiments, the nucleotide insertion comprises two nucleotide insertions.

[0225] The editing template of the PEgRNA may contain one or more intended nucleotide edits relative to the gene to be edited. The location of the intended nucleotide edit(s) relative to other components of the PEgRNA or specific nucleotides (e.g., mutations) within the target gene may vary. In some embodiments, the nucleotide edits are made in a region of the PEgRNA that corresponds to or is homologous to the protospacer sequence. In some embodiments, the nucleotide edits are made in a region of the PEgRNA that corresponds to a region of the gene outside the protospacer sequence.

[0226] In some embodiments, the location of nucleotide edit incorporation in a target gene can be determined based on the location of the protospacer adjacent motif (PAM). For example, the intended nucleotide edit can be installed at a sequence corresponding to the protospacer adjacent motif (PAM) sequence. In some embodiments, the nucleotide edit in the editing template is at a position corresponding to the 5'-terminal nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in the editing template is at a position corresponding to the 3'-terminal nucleotide of the PAM sequence. In some embodiments, the location of the intended nucleotide edit in the editing template can be referenced by aligning the editing template with a partially complementary edited strand of the target gene and referencing the nucleotide position on the edited strand where the intended nucleotide edit will be incorporated. In some embodiments, the nucleotide edit is incorporated at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 base pairs upstream of the 5'-most nucleotide of the PAM sequence within the edited strand of the target gene. 0 base pairs upstream or downstream of a reference position means that the intended nucleotide is immediately upstream or downstream of the reference position.In some embodiments, the nucleotide editing is performed on about 0 to 2 base pairs, 0 to 4 base pairs, 0 to 6 base pairs, 0 to 8 base pairs, 0 to 10 base pairs, 2 to 4 base pairs, 2 to 6 base pairs, 2 to 8 base pairs, 2 to 10 base pairs, 2 to 12 base pairs, 4 to 6 base pairs, 4 to 8 base pairs, 4 to 10 base pairs, 4 to 12 base pairs, 4 to 14 base pairs, 6 to 8 base pairs, 6 to 10 base pairs, 6 to 12 base pairs, 6 to 14 base pairs, 6 to 16 base pairs, 8 to 10 base pairs, 8 to 12 base pairs, 8 to 14 base pairs, 8 to 16 base pairs, 8 to 18 base pairs, 10 to 12 base pairs, 10 to 14 base pairs, 10 to 16 base pairs, 10 ... The nucleic acid is incorporated at a position corresponding to up to 18 base pairs, 10 to 20 base pairs, 12 to 14 base pairs, 12 to 16 base pairs, 12 to 18 base pairs, 12 to 20 base pairs, 12 to 22 base pairs, 14 to 16 base pairs, 14 to 18 base pairs, 14 to 20 base pairs, 14 to 22 base pairs, 14 to 24 base pairs, 16 to 18 base pairs, 16 to 20 base pairs, 16 to 22 base pairs, 16 to 24 base pairs, 16 to 26 base pairs, 18 to 20 base pairs, 18 to 22 base pairs, 18 to 24 base pairs, 18 to 26 base pairs, 18 to 28 base pairs, 20 to 22 base pairs, 20 to 24 base pairs, 20 to 26 base pairs, 20 to 28 base pairs, or 20 to 30 base pairs upstream. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 3 base pairs upstream of the 5'-most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 4 base pairs upstream of the 5'-most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 5 base pairs upstream of the 5'-most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in the editing template is at a position corresponding to 6 base pairs upstream of the 5'-most nucleotide of the PAM sequence.

[0227] In some embodiments, the intended nucleotide edit is incorporated into the edited strand of the target gene at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 base pairs downstream of the 5'-most nucleotide of the PAM sequence. In some embodiments, the nucleotide editing is performed on about 0 to 2 base pairs, 0 to 4 base pairs, 0 to 6 base pairs, 0 to 8 base pairs, 0 to 10 base pairs, 2 to 4 base pairs, 2 to 6 base pairs, 2 to 8 base pairs, 2 to 10 base pairs, 2 to 12 base pairs, 4 to 6 base pairs, 4 to 8 base pairs, 4 to 10 base pairs, 4 to 12 base pairs, 4 to 14 base pairs, 6 to 8 base pairs, 6 to 10 base pairs, 6 to 12 base pairs, 6 to 14 base pairs, 6 to 16 base pairs, 8 to 10 base pairs, 8 to 12 base pairs, 8 to 14 base pairs, 8 to 16 base pairs, 8 to 18 base pairs, 10 to 12 base pairs, 10 to 14 base pairs, 10 to 16 base pairs, 10 ... The nucleic acid is incorporated at a position corresponding to up to 18 base pairs, 10 to 20 base pairs, 12 to 14 base pairs, 12 to 16 base pairs, 12 to 18 base pairs, 12 to 20 base pairs, 12 to 22 base pairs, 14 to 16 base pairs, 14 to 18 base pairs, 14 to 20 base pairs, 14 to 22 base pairs, 14 to 24 base pairs, 16 to 18 base pairs, 16 to 20 base pairs, 16 to 22 base pairs, 16 to 24 base pairs, 16 to 26 base pairs, 18 to 20 base pairs, 18 to 22 base pairs, 18 to 24 base pairs, 18 to 26 base pairs, 18 to 28 base pairs, 20 to 22 base pairs, 20 to 24 base pairs, 20 to 26 base pairs, 20 to 28 base pairs, or 20 to 30 base pairs downstream. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 3 base pairs downstream of the 5'-most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 4 base pairs downstream of the 5'-most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 5 base pairs downstream of the 5'-most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 6 base pairs downstream of the 5'-most nucleotide of the PAM sequence."Upstream" and "downstream" are intended to define the relative locations of at least two regions or sequences within a nucleic acid molecule oriented in a 5' to 3' direction. For example, a first sequence is upstream of a second sequence within a DNA molecule, and the first sequence is located 5' of the second sequence. Thus, the second sequence is downstream of the first sequence.

[0228] When referenced in a PEGRNA, the location of one or more intended nucleotide edits can be referenced relative to a component of the PEGRNA. For example, the intended nucleotide edit can be 5' or 3' of the PBS. In some embodiments, the PEGRNA includes, from 5' to 3', the following structures: spacer, gRNA core, editing template, and PBS. In some embodiments, the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 base pairs upstream of the 5'-most nucleotide of the PBS. In some embodiments, the intended nucleotide edit is about 0-2 base pairs, 0-4 base pairs, 0-6 base pairs, 0-8 base pairs, 0-10 base pairs, 2-4 base pairs, 2-6 base pairs, 2-8 base pairs, 2-10 base pairs, 2-12 base pairs, 4-6 base pairs, 4-8 base pairs, 4-10 base pairs, 4-12 base pairs, 4-14 base pairs, 6-8 base pairs, 6-10 base pairs, 6-12 base pairs, 6-14 base pairs, 6-16 base pairs, 8-10 base pairs, 8-12 base pairs, 8-14 base pairs, 8-16 base pairs, 8-18 base pairs, 10-12 base pairs, 10-14 base pairs, 10-16 ... The nucleic acid sequence is incorporated at a position corresponding to 0 to 18 base pairs, 10 to 20 base pairs, 12 to 14 base pairs, 12 to 16 base pairs, 12 to 18 base pairs, 12 to 20 base pairs, 12 to 22 base pairs, 14 to 16 base pairs, 14 to 18 base pairs, 14 to 20 base pairs, 14 to 22 base pairs, 14 to 24 base pairs, 16 to 18 base pairs, 16 to 20 base pairs, 16 to 22 base pairs, 16 to 24 base pairs, 16 to 26 base pairs, 18 to 20 base pairs, 18 to 22 base pairs, 18 to 24 base pairs, 18 to 26 base pairs, 18 to 28 base pairs, 20 to 22 base pairs, 20 to 24 base pairs, 20 to 26 base pairs, 20 to 28 base pairs, or 20 to 30 base pairs upstream.

[0229] The corresponding position of the intended nucleotide edit incorporated into the target gene can also be referenced based on the nicking position generated by the prime editor based on sequence homology and complementarity. For example, in embodiments, the distance between the nucleotide edit incorporated into the target gene and the nick generated by the prime editor can be determined when the spacer hybridizes to the sequence to be searched and the extension arm hybridizes to the sequence to be edited. In certain embodiments, the position of the nucleotide edit can be any position downstream of the nick site on the edited strand (or PAM strand) generated by the prime editor, and the distance between the nick site and the intended nucleotide edit will be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some embodiments, the location of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream of the nick site on the edited strand. In some embodiments, the location of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides downstream from the nick site on the edited strand. In some embodiments, the location of the nucleotide edit is 0 base pairs from the nick site on the edited strand, meaning the edit position is at the same position as the nick site. As used herein, the distance between the nick site and the nucleotide edit refers to the 5'-most position of the nucleotide edit of the nick that creates a 3' free end on the edited strand (i.e., the "proximal position" of the nucleotide edit relative to the nick site), e.g., when the nucleotide edit comprises an insertion or deletion. Similarly, as used herein, the distance between the nick site and the PAM position edit refers to the 5'-most position of the nucleotide edit and the 5'-most position of the PAM sequence, e.g., when the nucleotide edit comprises an insertion, deletion, or substitution of two or more consecutive nucleotides.

[0230] A PEgRNA may also include optional modifiers, such as a 3'-end modifier region and / or a 5'-end modifier region. In some embodiments, a PEgRNA includes at least one nucleotide that is not part of the spacer, gRNA core, or extension arm. The optional sequence modifier can be located within or between the other regions shown and is not limited to being located at the 3' and 5' ends. In certain embodiments, a PEgRNA includes a secondary RNA structure, such as (but not limited to) an aptamer, hairpin, stem / loop, toe-loop, and / or an RNA-binding protein recruitment domain (e.g., an MS2 aptamer that recruits and binds MS2cp protein). In some embodiments, a PEgRNA includes a short uracil sequence at the 5' or 3' end. For example, in some embodiments, a PEgRNA including a 3' extension arm includes a "UUU" sequence at the 3' end of the extension arm. In some embodiments, a PEgRNA includes a toe-loop sequence at the 3' end. In some embodiments, the PEgRNA comprises a 3' extension arm and a toe loop sequence at the 3' end of the extension arm. In some embodiments, the PEgRNA comprises a 5' extension arm and a toe loop sequence at the 5' end of the extension arm. In some embodiments, the PEgRNA comprises a loop element having the sequence 5'-GAAANNNNN-3', where N is any nucleobase. In some embodiments, the secondary RNA structure is located within the spacer. In some embodiments, the secondary structure is located within the extension arm. In some embodiments, the secondary structure is located within the gRNA core. In some embodiments, the secondary structure is located between the spacer and the gRNA core, between the gRNA core and the extension arm, or between the spacer and the extension arm. In some embodiments, the secondary structure is located between the PBS and the editing template. In some embodiments, the secondary structure is located at the 3' or 5' end of the PEgRNA. In some embodiments, the PEgRNA comprises a transcription termination signal at the 3' end of the PEgRNA.In addition to the secondary RNA structure, PEGRNAs contain a chemical linker or poly(N) linker or tail, where "N" is any nucleobase. In some embodiments, the chemical linker can act to prevent reverse transcription of the gRNA core.

[0231] In some embodiments, the prime editing system or composition further comprises a nick-guide polynucleotide, such as a nick-guide RNA (ngRNA). Without intending to be bound by any particular theory, the non-edited strand of double-stranded target DNA within a target gene can be nicked by a CRISPR-Cas nickase guided by the ngRNA. In some embodiments, the nick on the non-edited strand can instruct endogenous DNA repair mechanisms to use the edited strand as a template for repair of the non-edited strand, thereby increasing the efficiency of prime editing. In some embodiments, the non-edited strand is nicked by a prime editor localized to the non-edited strand by the ngRNA. Thus, the present specification also provides a PEgRNA system comprising at least one PEgRNA and at least one ngRNA.

[0232] In some embodiments, the ngRNA is a guide RNA comprising a variable spacer sequence and a guide RNA scaffold or core region that interacts with a DNA-binding domain, such as Cas9, of a prime editor. In some embodiments, the ngRNA comprises a spacer sequence (referred to as the ng spacer or second spacer) that is substantially complementary to a second search target sequence (or ng search target sequence) located on the edited or non-target strand. Thus, in some embodiments, the ng search target sequence recognized by the ng spacer and the search target sequence recognized by the spacer sequence of the PEGRNA are present on opposite strands of the double-stranded target DNA of a target gene, e.g., a gene.

[0233] A prime editing system or composition that does not include ngRNA may be referred to as a "PE2" prime editing system. A prime editing system or composition that includes ngRNA may be referred to as a "PE3" prime editing system or PE3 prime editing complex.

[0234] In some embodiments, the ng search sequence is on the non-target strand within 10 to 100 base pairs of the intended nucleotide edit incorporated by the PEgRNA on the edited strand. In some embodiments, the ng target search sequence is within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp of the intended nucleotide edit incorporated by the PEgRNA on the edited strand. In some embodiments, the 5' ends of the ng search sequence and the 5' ends of the PEgRNA search sequence are separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp. In some embodiments, the 5' end of the ng probe sequence and the 5' end of the PEGRNA probe sequence are separated by 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp.

[0235] In some embodiments, the ng spacer sequence is complementary to and can hybridize with the second target sequence only after the intended nucleotide edit has been incorporated into the edited strand by the PEgRNA editing template. Such a prime editing system may be referred to as a "PE3b" prime editing system or configuration. In some embodiments, the ngRNA includes a spacer sequence that matches only the edited strand after the nucleotide edit has been incorporated, but does not match the endogenous target gene sequence on the edited strand. Thus, in some embodiments, the intended nucleotide edit is incorporated within the ng target sequence. In some embodiments, the intended nucleotide edit is incorporated within about 1 to 10 nucleotides of the position corresponding to the PAM in the ng target sequence.

[0236] The PEGRNA and / or ngRNA of the present disclosure, in some embodiments, comprise modified nucleotides, e.g., chemically modified DNA or RNA nucleobases, and may include one or more nucleobase analogs (e.g., modifications that may add functionality, such as temperature tolerance). In some embodiments, the PEGRNA and / or ngRNA described herein may be chemically modified. As used herein, the phrase "chemically modified" may include modifications that introduce chemistry that differs from that found in naturally occurring DNA or RNA, e.g., covalent modifications such as the introduction of modified nucleotides (e.g., nucleotide analogs, or incorporation of non-naturally occurring pendant groups into DNA or RNA molecules).

[0237] In some embodiments, the PEgRNA and / or ngRNA provided herein can be chemically or biologically modified. Modifications can occur at any position within the PEgRNA or ngRNA and can include modifications to the nucleobase or phosphate backbone of the PEgRNA or ngRNA. In some embodiments, the chemical modification can be a structure-guided modification. In some embodiments, the chemical modification is at the 5' and / or 3' end of the PEgRNA. In some embodiments, the chemical modification is at the 5' and / or 3' end of the ngRNA. In some embodiments, the chemical modification can occur within the spacer sequence, extension arm, editing template sequence, or primer binding site of the PEgRNA. In some embodiments, the chemical modification can occur within the spacer sequence or gRNA core of the PEgRNA or ngRNA. In some embodiments, the chemical modification can occur within the 3'-terminal nucleotide of the PEgRNA or ngRNA. In some embodiments, the chemical modification can occur within the 3'-end of the PEgRNA or ngRNA. In some embodiments, the chemical modification can occur within the 5'-end of the PEgRNA or ngRNA. In some embodiments, the PEGRNA or ngRNA comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more chemically modified nucleotides at the 3'-end. In some embodiments, the PEGRNA or ngRNA comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more chemically modified nucleotides at the 5'-end. In some embodiments, the PEGRNA or ngRNA comprises 1, 2, 3, 4, or 5 or more chemically modified nucleotides at the 3'-end. In some embodiments, the PEGRNA or ngRNA comprises 1, 2, 3, 4, or 5 or more chemically modified nucleotides at the 5'-end. In some embodiments, the PEGRNA or ngRNA comprises 1, 2, or 3 or more chemically modified nucleotides at the 3'-end. In some embodiments, the PEGRNA or ngRNA comprises 1, 2, or 3 chemically modified nucleotides at the 5'-end.In some embodiments, the PEgRNA or ngRNA comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more consecutive chemically modified nucleotides at the 3'-end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more consecutive chemically modified nucleotides at the 5'-end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, 3, 4, or 5 consecutive chemically modified nucleotides at the 3'-end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, 3, 4, or 5 consecutive chemically modified nucleotides at the 5'-end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, or 3 consecutive chemically modified nucleotides at the 3'-end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, or 3 consecutive chemically modified nucleotides at the 5'-end. In some embodiments, the PEgRNA or ngRNA comprises three consecutive chemically modified nucleotides at the 3' end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, 3, 4, 5, or more chemically modified nucleotides near the 3' end. In some embodiments, the PEgRNA or ngRNA comprises three consecutive chemically modified nucleotides at the 3' end. In some embodiments, the PEgRNA or ngRNA comprises three consecutive chemically modified nucleotides at the 5' end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, 3, 4, 5, or more chemically modified nucleotides near the 3' end. In some embodiments, the PEgRNA or ngRNA comprises 1, 2, 3, 4, 5, or more consecutive chemically modified nucleotides near the 3' end. In some embodiments, the PEGRNA or ngRNA comprises 1, 2, 3, 4, 5, or more chemically modified nucleotides near the 3' end, wherein the 3' terminal nucleotide is unmodified and the 1, 2, 3, 4, 5, or more chemically modified nucleotides precede the 3' terminal nucleotide in 5' to 3' order.In some embodiments, the PEGRNA or ngRNA comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more chemically modified nucleotides near the 3'-most The nucleotides are unmodified and 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more chemically modified nucleotides precede the 3'-most nucleotide in 5' to 3' order.

[0238] In some embodiments, a PEG-gRNA or ngRNA comprises one or more chemically modified nucleotides in the gRNA core. As illustrated in Figure 8, the gRNA core of a PEG-gRNA may comprise one or more regions of a base-paired lower stem, a base-paired upper stem, and the lower and upper stems may be connected by a bulge containing unpaired RNA. The gRNA core may further comprise a nexus distal to the spacer sequence. In some embodiments, the gRNA core comprises one or more chemically modified nucleotides in the lower stem, the upper stem, and / or the hairpin region. In some embodiments, all nucleotides in the lower stem, the upper stem, and / or the hairpin region are chemically modified.

[0239] Chemical modifications to PEGRNA or ngRNA include 2'-O-thionocarbamate-protected nucleoside phosphoramidites, 2'-O-methyl (M), 2'-O-methyl 3' phosphorothioate (MS), or 2'-O-methyl 3' thioPACE (MSP), or any combination thereof. In some embodiments, chemically modified PEGRNA and / or ngRNA can include 2'-O-methyl (M) RNA, 2'-O-methyl 3' phosphorothioate (MS) RNA, 2'-O-methyl 3' thioPACE (MSP) RNA, 2'-F RNA, phosphorothioate linkage modifications, other chemical modifications known in the art, or any combination thereof. Chemical modifications can also include, for example, the incorporation of non-nucleotide linkages or modified nucleotides into PEGRNA and / or ngRNA (e.g., modifications to one or both of the 3' and 5' ends of the guide RNA molecule). Such modifications include adding bases to the RNA sequence, complexing the RNA with substances (such as proteins or complementary nucleic acid molecules), or incorporating elements that change the structure of the RNA molecule (such as elements that form secondary structures).

[0240] Prime Editor The term "prime editor (PE)" refers to a polypeptide or polypeptide component involved in prime editing, or any polynucleotide(s) encoding a polypeptide or polypeptide component. In various embodiments, a prime editor comprises a polypeptide domain with DNA-binding activity and a polypeptide domain with DNA polymerase activity. In some embodiments, a prime editor further comprises a polypeptide domain with nuclease activity. In some embodiments, the polypeptide domain with DNA-binding activity comprises a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain with nuclease activity comprises a nickase or a fully active nuclease. As used herein, the term "nickase" refers to a nuclease that can cleave only one strand of a double-stranded DNA target. In some embodiments, a prime editor comprises a polypeptide domain that is an inactive nuclease. In some embodiments, the polypeptide domain with programmable DNA-binding activity comprises a nucleic acid-guided DNA-binding domain, e.g., a CRISPR-Cas protein, e.g., Cas9 nickase, Cpf1 nickase, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain having DNA polymerase activity comprises a template-dependent DNA polymerase, e.g., a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase. In some embodiments, the prime editor comprises an additional polypeptide involved in prime editing, e.g., a polypeptide domain having 5' endonuclease activity, e.g., a 5' endogenous DNA flap endonuclease (e.g., FEN1), which helps drive the prime editing process toward the formation of an edited product. In some embodiments, the prime editor further comprises an RNA-protein recruitment polypeptide, e.g., an MS2 coat protein.

[0241] Prime editors can be designed. In some embodiments, the polypeptide components of a prime editor do not naturally occur within the same organism or cellular environment. In some embodiments, the polypeptide components of a prime editor can be of different origins or from different organisms. In some embodiments, a prime editor comprises a DNA binding domain and a DNA polymerase domain from different species. In some embodiments, a prime editor comprises a Cas polypeptide and a reverse transcriptase polypeptide from different species. For example, a prime editor can comprise a S. pyogenes Cas9 polypeptide and a Moloney murine leukemia virus (M-MLV) reverse transcriptase polypeptide.

[0242] In some embodiments, prime editor polypeptide domains can be fused or linked by a peptide linker to form a fusion protein. In other embodiments, a prime editor comprises one or more polypeptide domains provided in trans as separate proteins that can bind to each other via non-peptide bonds or aptamer or recruitment sequences. For example, a prime editor may comprise a DNA-binding domain and a reverse transcriptase domain linked to each other by an RNA-protein recruitment aptamer (e.g., an MS2 aptamer that can be linked to PEgRNA). Prime editor polypeptide components may be wholly or partially encoded by one or more polynucleotides. In some embodiments, a single polynucleotide, construct, or vector encodes a prime editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or portion of a domain of the prime editor, or a portion of the prime editor fusion protein. For example, a prime editor fusion protein may comprise an N-terminal portion fused to intein N and a C-terminal portion fused to intein C, each encoded separately by an AAV vector.

[0243] Prime Editor Nucleotide Polymerase Domain In some embodiments, the prime editor comprises a nucleotide polymerase domain, e.g., a DNA polymerase domain. The DNA polymerase domain can be a wild-type DNA polymerase domain, a full-length DNA polymerase protein domain, or a functional mutant, functional variant, or functional fragment thereof. In some embodiments, the polymerase domain is a template-dependent polymerase domain. For example, a DNA polymerase can rely on a template polynucleotide strand (e.g., an editing template sequence) to synthesize a new strand of DNA. In some embodiments, the prime editor comprises a DNA-dependent DNA polymerase. For example, a prime editor having a DNA-dependent DNA polymerase can synthesize new single-stranded DNA using a PEgRNA editing template that includes a DNA sequence as a template. In such cases, the PEgRNA is a chimeric or hybrid PEgRNA and includes an extension arm that includes a DNA strand. The chimeric or hybrid PEgRNA includes an RNA portion (including a spacer and a gRNA core) and a DNA portion (an extension arm that includes an editing template that includes a DNA strand).

[0244] The DNA polymerase can be a wild-type polymerase from a eukaryotic, prokaryotic, archaeal, or viral organism, and / or the polymerase can be modified by processes based on genetic engineering, mutagenesis, or directed evolution. Polymerases include T7 DNA polymerase, T5 DNA polymerase, T4 DNA polymerase, Klenow fragment DNA polymerase, DNA polymerase III, etc. The polymerases are thermostable and include Taq, Tne, Tma, Pfu, Tfl, Tth, Stoffel fragment, VENT® and DEEPVENT® DNA polymerase, KOD, Tgo, JDF3, and their mutants, variants, and derivatives.

[0245] For the synthesis of longer nucleic acid molecules (e.g., nucleic acid molecules greater than about 3-5 kb in length), at least two DNA polymerases can be used. In certain embodiments, one of the polymerases may substantially lack 3' exonuclease activity, while the other may possess 3' exonuclease activity. Such pairings may involve the same or different polymerases. Examples of DNA polymerases that substantially lack 3' exonuclease activity include, but are not limited to, Taq, Tne(exo-), Tma(exo-), Pfu(exo-), Pwo(exo-), exo-KOD, and Tth DNA polymerases, as well as functional mutants, variants, and fragments thereof.

[0246] In some embodiments, the DNA polymerase is a bacteriophage polymerase, e.g., T4, T7, or phi29 DNA polymerase. In some embodiments, the DNA polymerase is an archaeal polymerase, e.g., a pol I-type archaeal polymerase or a pol II-type archaeal polymerase. In some embodiments, the DNA polymerase comprises a thermostable archaeal DNA polymerase. In some embodiments, the DNA polymerase comprises a eubacterial DNA polymerase, e.g., a Pol I, Pol II, or Pol III polymerase. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is an E. coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is a Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is E. coli Pol IV DNA polymerase.

[0247] In some embodiments, the DNA polymerase comprises a eukaryotic DNA polymerase. In some embodiments, the DNA polymerase is Pol-beta DNA polymerase, Pol-lambda DNA polymerase, Pol-sigma DNA polymerase, or Pol-mu DNA polymerase. In some embodiments, the DNA polymerase is Pol-alpha DNA polymerase. In some embodiments, the DNA polymerase is POLA1 DNA polymerase. In some embodiments, the DNA polymerase is POLA2 DNA polymerase. In some embodiments, the DNA polymerase is Pol-delta DNA polymerase. In some embodiments, the DNA polymerase is POLD1 DNA polymerase. In some embodiments, the DNA polymerase is POLD2 DNA polymerase. In some embodiments, the DNA polymerase is human POLD1 DNA polymerase. In some embodiments, the DNA polymerase is human POLD2 DNA polymerase. In some embodiments, the DNA polymerase is POLD3 DNA polymerase. In some embodiments, the DNA polymerase is a POLD4 DNA polymerase. In some embodiments, the DNA polymerase is a Pol-epsilon DNA polymerase. In some embodiments, the DNA polymerase is a POLE1 DNA polymerase. In some embodiments, the DNA polymerase is a POLE2 DNA polymerase. In some embodiments, the DNA polymerase is a POLE3 DNA polymerase. In some embodiments, the DNA polymerase is a Pol-eta (POLH) DNA polymerase. In some embodiments, the DNA polymerase is a Pol-iota (POLI) DNA polymerase. In some embodiments, the DNA polymerase is a Pol-kappa (POLK) DNA polymerase. In some embodiments, the DNA polymerase is a Rev1 DNA polymerase. In some embodiments, the DNA polymerase is a human Rev1 DNA polymerase.In some embodiments, the DNA polymerase is a viral DNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a family B DNA polymerase. In some embodiments, the DNA polymerase is herpes simplex virus (HSV) UL30 DNA polymerase. In some embodiments, the DNA polymerase is cytomegalovirus (CMV) UL54 DNA polymerase.

[0248] In some embodiments, the DNA polymerase is an archaeal polymerase. In some embodiments, the DNA polymerase is a family B / pol I type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of Pfu from Pyrococcus furiosus. In some embodiments, the DNA polymerase is a pol II type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of the P. furiosus DP1 / DP2 two-subunit polymerase. In some embodiments, the DNA polymerase lacks 5' to 3' nuclease activity. Suitable DNA polymerases (pol I or pol II) can be obtained from archaea with optimal growth temperatures similar to the desired measurement temperature.

[0249] In some embodiments, the DNA polymerase comprises a thermostable archaeal DNA polymerase, hi some embodiments, the thermostable DNA polymerase is isolated or derived from Pyrococcus species (furiosus, species GB-D, woesii, abysii, horikoshii), Thermococcus species (kodakaraensis KOD1, litoralis, species 9 degrees North-7, species JDF-3, gorgonarius), Pyrodictium occultum, and Archaeoglobus fulgidus.

[0250] The polymerase can also be from a eubacterial species. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is an E. coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is a Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol III family DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is an E. coli Pol IV DNA polymerase. In some embodiments, the Pol I DNA polymerase is a DNA polymerase functional mutant lacking or having reduced 5' to 3' exonuclease activity.

[0251] Suitable thermostable pol I DNA polymerases can be isolated from a variety of thermophilic eubacteria, including Thermus species and Thermotoga maritima, such as Thermus aquaticus (Taq), Thermus thermophilus (Tth), and Thermus maritima (Tma UlTma).

[0252] In some embodiments, the prime editor comprises an RNA-dependent DNA polymerase domain, such as a reverse transcriptase (RT). The RT or RT domain can be a wild-type RT domain, a full-length RT domain, or a functional mutant, functional variant, or functional fragment thereof. The RT or RT domain of the prime editor can comprise a wild-type RT or can be designed or evolved to contain specific amino acid substitutions, truncations, or mutations. The designed RT can contain sequence or amino acid changes that differ from a naturally occurring RT. In some embodiments, the designed RT can have improved reverse transcription activity compared to a naturally occurring RT or RT domain. In some embodiments, the designed RT can have improved functionality, such as thermostability, reverse transcription efficiency, or target fidelity, compared to a naturally occurring RT. In some embodiments, a prime editor comprising a designed RT has improved prime editing efficiency compared to a prime editor with a reference naturally occurring RT.

[0253] In some embodiments, the prime editor comprises a viral RT, e.g., a retroviral RT. Non-limiting examples of viral RTs include Moloney murine leukemia virus (M-MLV or MLVRT), human T-cell leukemia virus type 1 (HTLV-1) RT, bovine leukemia virus (BLV) RT, Rous sarcoma virus (RSV) RT, human immunodeficiency virus (HIV) RT, M-MFV RT, avian sarcoma leukemia virus (ASLV) RT, Rous sarcoma virus (RSV) RT, avian myeloblastosis virus (AMV) RT, avian erythroblastosis virus (AEV) helper virus MCAV RT, avian myelocytomatosis virus MC29 helper virus MCAV RT, avian reticuloendotheliosis virus (REV-T) helper virus REV-A RT, avian sarcoma virus UR2 helper virus (UR2AV) RT, and avian sarcoma virus Y73 helper virus YAV. RT, Rous-associated virus (RAV) RT, and myeloblastosis-associated virus (MAV) RT, all of which may be suitably used in the methods and compositions described herein.

[0254] In some embodiments, the prime editor comprises wild-type M-MLV RT. The sequence of an exemplary wild-type M-MLV RT is provided in SEQ ID NO:4448.

[0255] In some embodiments, the prime editor comprises a M-MMLV RT that includes one or more of the amino acid substitutions P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, or D653X compared to the wild-type M-MMLV RT set forth in SEQ ID NO:4448, where X is any amino acid other than the wild-type amino acid. In some embodiments, the prime editor comprises a M-MMLV RT that includes one or more of the amino acid substitutions P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, and D653N compared to the wild-type M-MMLV RT set forth in SEQ ID NO:4448. In some embodiments, the prime editor comprises an M-MLV RT that includes one or more of the amino acid substitutions D200N, T330P, L603W, T306K, and W313F compared to the wild-type M-MMLV RT set forth in SEQ ID NO: 4448. In some embodiments, the prime editor comprises an M-MLV RT that includes the amino acid substitutions D200N, T330P, L603W, T306K, and W313F compared to the wild-type M-MMLV RT set forth in SEQ ID NO: 4448.

[0256] For example, wild-type Moloney murine leukemia virus reverse transcriptase:

[0257] [ka] In some embodiments, an RT variant can be a functional fragment of a reference RT having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500 or more amino acid changes compared to a reference RT, e.g., a wild-type RT. In some embodiments, the RT variant comprises a fragment of a reference RT, e.g., a wild-type RT, that is about 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% identical to the corresponding fragment of the reference RT. In some embodiments, the fragment is 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical in amino acid length to the corresponding wild-type RT (M-MLV reverse transcriptase) (e.g., SEQ ID NO: 4448).

[0258] In some embodiments, an RT functional fragment is at least 100 amino acids in length, hi some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, or up to 600 or more amino acids in length.

[0259] In yet other embodiments, functional RT variants are truncated at the N-terminus or C-terminus, or both, by a specific number of amino acids, resulting in truncation variants that still retain sufficient DNA polymerase function. In some embodiments, RT truncation variants are truncated by at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the N-terminus compared to a reference RT, e.g., wild-type RT. In some embodiments, the reference RT is wild-type M-MLV RT. In other embodiments, the RT truncation mutant is truncated by at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the C-terminus compared to the reference RT, e.g., wild-type RT. In some embodiments, the reference RT is wild-type M-MLV RT. In yet other embodiments, the RT truncation mutant is truncated at the N-terminus and C-terminus compared to a reference RT, e.g., wild-type RT. In some embodiments, the N-terminal and C-terminal truncations are the same length. In some embodiments, the N-terminal and C-terminal truncations are different lengths.

[0260] For example, a prime editor disclosed herein can comprise a functional variant of a wild-type M-MLV reverse transcriptase. In some embodiments, the prime editor comprises a functional variant of a wild-type M-MLV RT, wherein the functional variant of M-MLV RT is truncated after amino acid position 502 compared to the wild-type M-MLV RT set forth in SEQ ID NO:4448. In some embodiments, the functional variant of M-MLV RT further comprises a D200X, T306X, W313X, and / or T330X amino acid substitution compared to the wild-type M-MLV RT set forth in SEQ ID NO:4448, where X is any amino acid other than the original amino acid. In some embodiments, the functional variant of M-MLV RT further comprises a D200N, T306K, W313F, and / or T330P amino acid substitution compared to the wild-type M-MLV RT set forth in SEQ ID NO:4448, where X is any amino acid other than the original amino acid. The nucleic acid sequence encoding a prime editor containing this truncated RT is 522 nt smaller than the nucleic acid sequence encoding a prime editor containing the full-length M-MLV RT, and is therefore potentially useful in applications where delivery of DNA sequences is difficult due to their size (i.e., adeno-associated viral and lentiviral delivery). In some embodiments, the prime editor comprises an M-MLV RT mutant, where the M-MLV RT consists of the following amino acid sequence: [ka]

[0261] In some embodiments, the prime editor comprises a eukaryotic RT, e.g., a yeast, Drosophila, rodent, or primate RT. In some embodiments, the prime editor comprises a group II intron RT, e.g., a Geobacillus stearothermophilus group II intron (GsI-IIC) RT or a Eubacterium rectale group II intron (Eu.re.I2) RT. In some embodiments, the prime editor comprises a retron RT.

[0262] Programmable DNA-binding domain In some embodiments, the DNA-binding domain of the prime editor is a programmable DNA-binding domain. A programmable DNA-binding domain refers to a protein domain designed to bind to a specific nucleic acid sequence, e.g., a target DNA or target RNA. In some embodiments, the DNA-binding domain is a polynucleotide-programmable DNA-binding domain that can bind to a guide polynucleotide (e.g., PEGRNA) that guides the DNA-binding domain to a specific DNA sequence, e.g., a target sequence within a target gene. In some embodiments, the DNA-binding domain comprises a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein. The Cas protein may comprise any Cas protein described herein, or a functional fragment or variant thereof. In some embodiments, the DNA-binding domain may also comprise a zinc finger protein domain. In other cases, the DNA-binding domain comprises a transcription activator-like effector domain (TALE). In some embodiments, the DNA-binding domain comprises a DNA nuclease. For example, the DNA-binding domain of the prime editor may comprise an RNA-guided DNA endonuclease, e.g., a Cas protein. In some embodiments, the DNA binding domain comprises a zinc finger nuclease (ZFN) or a transcription activator-like effector domain nuclease (TALEN), in which one or more zinc finger motifs or TALE motifs are associated with one or more nucleases, e.g., a Fok I nuclease domain.

[0263] In some embodiments, the DNA-binding domain has nuclease activity. In some embodiments, the DNA-binding domain of the prime editor comprises an endonuclease domain with single-stranded DNA cleavage activity. For example, the endonuclease domain may comprise a FokI nuclease domain. In some embodiments, the DNA-binding domain of the prime editor comprises a nuclease with full nuclease activity. In some embodiments, the DNA-binding domain of the prime editor comprises a nuclease with modified or reduced nuclease activity compared to the wild-type endonuclease domain. For example, the endonuclease domain may comprise one or more amino acid substitutions compared to the wild-type endonuclease domain. In some embodiments, the DNA-binding domain of the prime editor has nickase activity. In some embodiments, the DNA-binding domain of the prime editor comprises a Cas protein domain that is a nickase. In some embodiments, compared to the wild-type Cas protein, the Cas nickase comprises one or more amino acid substitutions in the nuclease domain, thereby reducing or eliminating double-stranded nuclease activity but retaining DNA-binding activity. In some embodiments, the Cas nickase comprises an amino acid substitution in the HNH domain. In some embodiments, the Cas nickase comprises an amino acid substitution in the RuvC domain.

[0264] In some embodiments, the DNA binding domain comprises a CRISPR-associated protein (Cas protein) domain. The Cas protein can be a class 1 or class 2 Cas protein. The Cas protein can be a type I, type II, type III, type IV, type V Cas protein, or type VI Cas protein. Non-limiting examples of Cas proteins include Cas9, Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b1 (C2c1), Cas12b2, Cas12c (C2c3), C2c4, C2c8, C2c5, C2c10, C2c9, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, Cns2, CasΦ, and homologs, functional fragments, or modified versions thereof. Cas proteins can also be chimeric Cas proteins fused to other proteins or polypeptides. Cas proteins can also be chimeras of various Cas proteins, for example, comprising domains of Cas proteins from different organisms.

[0265] The Cas protein, such as Cas9, can be from any suitable organism. In some aspects, the microorganism is Streptococcus pyogenes (S. pyogenes). In some aspects, the microorganism is Staphylococcus aureus (S. aureus). In some aspects, the microorganism is Streptococcus thermophilus (S. thermophilus). In some embodiments, the microorganism is Staphylococcus lugdunensis.

[0266] The Cas protein (e.g., Cas9) can be a wild-type or modified version of a Cas protein. The Cas protein (e.g., Cas9) can be a nuclease-active mutant, a nuclease-inactive mutant, a nickase, or a functional mutant or fragment of a wild-type Cas protein. The Cas protein (e.g., Cas9) can contain amino acid changes, such as deletions, insertions, substitutions, fusions, chimeras, or any combination thereof, compared to the wild-type version of the Cas protein. The Cas protein can be a polypeptide having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity to a wild-type exemplary Cas protein.

[0267] Cas proteins (e.g., Cas9) can comprise one or more domains. Non-limiting examples of Cas domains include a guide nucleic acid recognition and / or binding domain, a nuclease domain (e.g., DNase or RNase domain, RuvC, HNH), a DNA-binding domain, an RNA-binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain. In various embodiments, the Cas protein comprises a guide nucleic acid recognition and / or binding domain capable of interacting with a guide nucleic acid and one or more nuclease domains comprising catalytic activity for nucleic acid cleavage.

[0268] In some embodiments, a Cas protein, such as Cas9, comprises one or more nuclease domains. The Cas protein may comprise an amino acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a nuclease domain (e.g., RuvC domain, HNH domain) of a wild-type Cas protein. In some embodiments, the Cas protein comprises a single nuclease domain. For example, Cpfl may comprise a RuvC domain but lack an HNH domain. In some embodiments, the Cas protein comprises two nuclease domains; for example, a Cas9 protein may comprise an HNH nuclease domain and a RuvC nuclease domain.

[0269] In some embodiments, the prime editor comprises a Cas protein, e.g., Cas9, wherein all nuclease domains of the Cas protein are active. In some embodiments, the prime editor comprises a Cas protein with one or more inactive nuclease domains. One or more nuclease domains of the Cas protein (e.g., RuvC, HNH) can be deleted or mutated to render them non-functional or reduce their nuclease activity. In some embodiments, a Cas protein (e.g., Cas9) containing a mutation in its nuclease domain reduces (e.g., nickase) or eliminates nuclease activity while maintaining the ability to target a nucleic acid locus with a target sequence when complexed with a guide nucleic acid (e.g., PEgRNA).

[0270] In some embodiments, the prime editor comprises a Cas nickase that can bind sequence-specifically to a target gene and generate a single-stranded break at a protospacer within the double-stranded DNA of the target gene, but not a double-stranded break. For example, the Cas nickase can cleave either the edited strand or the non-edited strand of the target gene, but not both. In some embodiments, the prime editor comprises a Cas nickase that comprises two nuclease domains (e.g., Cas9), one of which is modified to lack catalytic activity or deleted. In some embodiments, the Cas nickase of the prime editor comprises a nuclease-inactive RuvC domain and a nuclease-active HNH domain. In some embodiments, the Cas nickase of the prime editor comprises a nuclease-inactive HNH domain and a nuclease-active RuvC domain. In some embodiments, the prime editor comprises a Cas9 nickase with an amino acid substitution in the RuvC domain. In some embodiments, the Cas9 nickase comprises a D10X amino acid substitution relative to wild-type S. pyogenes Cas9, where X is any amino acid other than D. In some embodiments, the prime editor comprises a Cas9 nickase with an amino acid substitution in the HNH domain. In some embodiments, the Cas9 nickase comprises a H840X amino acid substitution relative to wild-type S. pyogenes Cas9, where X is any amino acid other than H.

[0271] In some embodiments, the prime editor comprises a Cas protein that can bind to a target gene in a sequence-specific manner but lacks or has lost nuclease activity and is unable to cleave either strand of double-stranded DNA within the target gene. Loss of activity or absence of activity can refer to less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of the enzymatic activity compared to a wild-type exemplary activity (e.g., wild-type Cas9 nuclease activity). In some embodiments, the Cas protein of the prime editor lacks any nuclease activity. Nucleases that lack nuclease activity, such as Cas9, may be referred to as nuclease-inactive or "nuclease-dead" (abbreviated as "d"). Nuclease-inactive Cas proteins (e.g., dCas, dCas9) can bind to a target polynucleotide but may not be able to cleave the target polynucleotide. In some aspects, the dead Cas protein is a dead Cas9 protein. In some embodiments, the prime editor comprises a nuclease-dead Cas protein, in which all of the nuclease domains (e.g., both the RuvC and HNH nuclease domains in a Cas9 protein, the RuvC nuclease domain in a Cpf1 protein) are mutated or deleted to lack catalytic activity.

[0272] Cas proteins can be modified. Cas proteins, such as Cas9, can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to alter other activities or properties of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, the Cas protein can be truncated to remove domains that are not essential for protein function, or the activity of a Cas protein can be optimized (e.g., enhanced or decreased).

[0273] Cas proteins can be fusion proteins. For example, Cas proteins can be fused to a cleavage domain, epigenetic modification domain, transcriptional regulatory domain, or polymerase domain. Cas proteins can also be fused to heterologous polypeptides, which can increase or decrease stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internally of the Cas protein.

[0274] In some embodiments, the Cas protein of the prime editor is a class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, a homolog, mutant, variant, or functional fragment thereof. As used herein, Cas9, Cas9 protein, Cas9 polypeptide, or Cas9 nuclease refers to an RNA-guided nuclease comprising one or more Cas9 nuclease domains and a Cas9 gRNA-binding domain capable of binding to a guide polynucleotide (e.g., a PEG-RNA). A Cas9 protein can refer to a wild-type Cas9 protein from any organism, or a homolog, ortholog, paralog, functional mutant, or variant thereof from any organism, or a functional fragment or domain thereof. In some embodiments, the prime editor comprises a full-length Cas9 protein. In some embodiments, the Cas9 protein may generally comprise at least about 50%, 60%, 70%, 80%, 90%, 100% sequence identity to a wild-type reference Cas9 protein (e.g., Cas9 from S. pyogenes). In some embodiments, the Cas9 comprises amino acid alterations, such as deletions, insertions, substitutions, fusions, chimeras, or any combination thereof, compared to the wild-type reference Cas9 protein.

[0275] In some embodiments, the Cas9 protein may comprise a Cas9 protein from Streptococcus pyogenes (Sp), Staphylococcus aureus (Sa), Streptococcus canis (Sc), Streptococcus thermophilus (St), Staphylococcus lugdunensis (Slu), Neisseria meningitidis (Nm), Campylobacter jejuni (Cj), Francisella novicida (Fn), or Treponema denticola (Td), or any Cas9 homolog or ortholog from an organism known in the art. In some embodiments, the Cas9 polypeptide is an SpCas9 polypeptide. In some embodiments, the Cas9 polypeptide is an SaCas9 polypeptide. In some embodiments, the Cas9 polypeptide is an ScCas9 polypeptide. In some embodiments, the Cas9 polypeptide is an StCas9 polypeptide. In some embodiments, the Cas9 polypeptide is a SluCas9 polypeptide. In some embodiments, the Cas9 polypeptide is an NmCas9 polypeptide. In some embodiments, the Cas9 polypeptide is a CjCas9 polypeptide. In some embodiments, the Cas9 polypeptide is an FnCas9 polypeptide. In some embodiments, the Cas9 polypeptide is a TdCas9 polypeptide. In some embodiments, the Cas9 polypeptide is a chimera comprising domains from two or more organisms described herein or known in the art. In some embodiments, the Cas9 polypeptide is a Cas9 polypeptide from Streptococcus macacae. In some embodiments, the Cas9 polypeptide is a Cas9 polypeptide generated by replacing the PAM-interaction domain of SpCas9 with the PAM-interaction domain of Streptococcus macacae Cas9 (Spy-mac Cas9).

[0276] An exemplary Streptococcus pyogenes Cas9 (SpCas9) amino acid sequence is provided in SEQ ID NO:4449.

[0277] For an exemplary Streptococcus pyogenes Cas9 (SpCas9) amino acid sequence: [ka]

[0278] In some embodiments, the prime editor comprises the Cas9 protein from Staphylococcus lugdunensis (Slu Cas9). An exemplary amino acid sequence of Slu Cas9 is provided in SEQ ID NO:4450.

[0279] For the exemplary Staphylococcus lugdunensis Cas9 (Slu Cas9) amino acid sequence WP_002460848.1: [ka]

[0280] In some embodiments, the Cas9 protein comprises a mutant Cas9 protein comprising one or more amino acid substitutions. In some embodiments, the wild-type Cas9 protein comprises a RuvC domain and an HNH domain. In some embodiments, the prime editor comprises a nuclease-active Cas9 protein capable of cleaving both strands of a double-stranded target DNA sequence. In some embodiments, the nuclease-active Cas9 protein comprises a functional RuvC domain and a functional HNH domain. In some embodiments, the prime editor comprises a Cas9 nickase that can bind to a guide polynucleotide and recognize the target DNA, but can cleave only one strand of the double-stranded target DNA. In some embodiments, the Cas9 nickase comprises only one functional RuvC domain or one functional HNH domain. In some embodiments, the prime editor comprises a Cas9 with a non-functional HNH domain and a functional RuvC domain. In some embodiments, the prime editor can cleave the edited strand (i.e., the PAM strand) but cannot cleave the non-edited strand of a double-stranded target DNA sequence. In some embodiments, the prime editor comprises a Cas9 with a non-functional RuvC domain that can cleave the target strand (i.e., the non-PAM strand) but cannot cleave the edited strand of a double-stranded target DNA sequence. In some embodiments, the prime editor comprises a Cas9 that has neither a functional RuvC domain nor a functional HNH domain, and may not cleave either strand of a double-stranded target DNA sequence.

[0281] In some embodiments, the prime editor comprises a Cas9 with a mutation in the RuvC domain that reduces or abolishes the nuclease activity of the RuvC domain. In some embodiments, the Cas9 comprises a mutation at amino acid D10, or a corresponding mutation, compared to wild-type SpCas9 set forth in SEQ ID NO: 4449. In some embodiments, the Cas9 comprises a D10A mutation, or a corresponding mutation, compared to wild-type SpCas9 set forth in SEQ ID NO: 4449. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acids D10, G12, and / or G17, or a corresponding mutation, compared to wild-type SpCas9 set forth in SEQ ID NO: 4449. In some embodiments, the Cas9 polypeptide comprises a D10A mutation, a G12A mutation, and / or a G17A mutation, compared to wild-type SpCas9 set forth in SEQ ID NO: 4449, or a corresponding mutation.

[0282] In some embodiments, the prime editor comprises a Cas9 polypeptide having a mutation in the HNH domain that reduces or abolishes the nuclease activity of the HNH domain. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid H840, or a corresponding mutation thereof, compared to wild-type SpCas9 set forth in SEQ ID NO: 4449. In some embodiments, the Cas9 polypeptide comprises an H840A mutation, or a corresponding mutation thereof, compared to wild-type SpCas9 set forth in SEQ ID NO: 4449. In some embodiments, the Cas9 polypeptide comprises an amino acid E762, D839, H840, N854, N856, N863, H982, H983, A984, D986, and / or A987 mutation, or a mutation corresponding thereto, compared to wild-type SpCas9 set forth in SEQ ID NO: 4449. In some embodiments, the Cas9 polypeptide comprises an E762A, D839A, H840A, N854A, N856A, N863A, H982A, H983A, A984A, and / or D986A mutation, or a corresponding mutation, relative to the wild-type SpCas9 set forth in SEQ ID NO: 4449.

[0283] In some embodiments, the prime editor comprises a Cas9 with one or more amino acid substitutions in both the HNH and RuvC domains that reduce or abolish the nuclease activity of both the HNH and RuvC domains. In some embodiments, the prime editor comprises a nuclease-inactive Cas9, or a nuclease-inactive Cas9 (dCas9). In some embodiments, the dCas9 comprises an H840X substitution and a D10X mutation compared to the wild-type SpCas9 set forth in SEQ ID NO: 4449, or its corresponding mutation, where X is any amino acid other than H for the H840X substitution and any amino acid other than D for the D10X substitution. In some embodiments, the deadCas9 comprises an H840A and a D10A mutation compared to the wild-type SpCas9 set forth in SEQ ID NO: 4449, or its corresponding mutation.

[0284] In some embodiments, the N-terminal methionine is removed from a Cas9 nickase, or any Cas9 variant, ortholog, or equivalent disclosed or contemplated herein. For example, a methionine-minus Cas9 nickase includes the following sequence, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto:

[0285] In addition to dead Cas9 and Cas9 nickase variants, Cas9 proteins as used herein can also include other Cas9 variants that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference Cas9 protein, including any wild-type Cas9, or mutant Cas9 (e.g., dead Cas9 or Cas9 nickase), or fragment Cas9, or circularly permuted Cas9, or other Cas9 variants disclosed herein or known in the art. In some embodiments, the Cas9 mutant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to a reference Cas9 (e.g., wild-type Cas9). In some embodiments, the Cas9 variant comprises a fragment of a reference Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), wherein the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of the reference Cas9 (e.g., a wild-type Cas9). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9.

[0286] In some embodiments, the Cas9 fragment is a functional fragment that retains one or more Cas9 activities. In some embodiments, the Cas9 fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.

[0287] In some embodiments, the prime editor comprises a Cas protein (e.g., Cas9) containing modifications that enable altered PAM recognition. In prime editing using a Cas protein-based prime editor, the terms "protospacer adjacent motif (PAM)," PAM sequence, or PAM-like motif can be used to refer to a short DNA sequence immediately following the protospacer sequence on the PAM strand of a target gene. In some embodiments, the PAM is recognized by a Cas nuclease within the prime editor during prime editing. In certain embodiments, the PAM is required for target binding of the Cas protein. The specific PAM sequence required for Cas protein recognition may vary depending on the specific type of Cas protein. The PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides in length. In some embodiments, the PAM is 2-6 nucleotides in length. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). In some embodiments, the Cas protein of the prime editor recognizes a canonical PAM, e.g., SpCas9 recognizes 5'-NGG-3' PAM. In some embodiments, the Cas protein of the prime editor has altered or non-canonical PAM specificity. Exemplary PAM sequences and corresponding Cas mutants are set forth in Table 7 below. It should be understood that for each mutant provided, the Cas protein contains one or more amino acid substitutions as indicated compared to the wild-type Cas protein sequence, e.g., Cas9 set forth in SEQ ID NO: 4449. The PAM motifs set forth in Table 7 below are ordered from 5' to 3'. [Table 7]

[0288] In some embodiments, the prime editor has any of the following amino acids compared to the wild-type SpCas9 polypeptide set forth in SEQ ID NO: 4449: A61R, L111R, D1135V, R221K, A262T, R324L, N394K, S409I, S409I, E427G, E480K, M495V, N497A, Y515N, K526E, F539S, E543D, R654L, R661A, R661L, R691A, N692A, M694A, M694I, Q695A, H698A, R753G, M763I, K848A, K890N, Q926A, K1003A, R1060A, L1111R, R1114G, D11 and any combination thereof.

[0289] In some embodiments, the prime editor comprises a SaCas9 polypeptide. In some embodiments, the SaCas9 polypeptide comprises one or more mutations E782K, N968K, and R1015H compared to wild-type SaCas9. In some embodiments, the prime editor comprises an FnCas9 polypeptide, e.g., a wild-type FnCas9 polypeptide, or an FnCas9 polypeptide comprising one or more mutations E1369R, E1449H, or R1556A compared to wild-type FnCas9. In some embodiments, the prime editor comprises an Sc Cas, e.g., a wild-type ScCas9, or an ScCas9 polypeptide comprising one or more mutations I367K, G368D, I369K, H371L, T375S, T376G, and T1227K compared to wild-type ScCas9. In some embodiments, the prime editor comprises a St1 Cas9 polypeptide, a St3 Cas9 polypeptide, or a Slu Cas9 polypeptide.

[0290] In some embodiments, a prime editor comprises a Cas polypeptide comprising a circularly permuted Cas mutant. For example, a Cas9 polypeptide of a prime editor can be designed such that the N- and C-termini of a Cas9 protein (e.g., a wild-type Cas9 protein or a Cas9 nickase) are locally rearranged to retain the ability to bind to DNA when complexed with a guide RNA (gRNA). An exemplary circularly permuted sequence configuration can be N-terminus-[original C-terminus]-[original N-terminus]-C-terminus. The Cas9 proteins described herein can be rearranged as circularly permuted variants, including any mutant, homologous, or naturally occurring Cas9 or equivalent.

[0291] In various embodiments, a circular permutant of a Cas protein, such as Cas9, can have the structure N-terminus-[original C-terminus]-[optional linker]-[original N-terminus]-C-terminus. In some embodiments, a circularly permuted Cas9 comprises any one of the following structures:

[0292] N-terminus-[1268-1368]-[optional linker]-[1-1267]-C-terminus,

[0293] N-terminus-[1168-1368]-[one optional linker]-[1-1167]-C-terminus;

[0294] N-terminus-[1068-1368]-[optional linker]-[1-1067]-C-terminus,

[0295] N-terminus-[968-1368]-[optional linker]-[1-967]-C-terminus,

[0296] N-terminus-[868-1368]-[optional linker]-[1-867]-C-terminus,

[0297] N-terminus-[768-1368]-[optional linker]-[1-767]-C-terminus,

[0298] N-terminus-[668-1368]-[optional linker]-[1-667]-C-terminus,

[0299] N-terminus-[568-1368]-[optional linker]-[1-567]-C-terminus,

[0300] N-terminus-[468-1368]-[optional linker]-[1-467]-C-terminus,

[0301] N-terminus-[368-1368]-[optional linker]-[1-367]-C-terminus,

[0302] N-terminus-[268-1368]-[optional linker]-[1-267]-C-terminus,

[0303] N-terminus-[168-1368]-[optional linker]-[1-167]-C-terminus,

[0304] N-terminus-[68 to 1368]-[optional linker]-[1 to 67]-C-terminus,

[0305] N-terminus-[10-1368]-[optional linker]-[1-9]-C-terminus, or the corresponding circularly permuted sequence of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).

[0306] In some embodiments, the circularly permuted Cas9 comprises any one of the following structures (amino acid position set forth in SEQ ID NO: 4449 - amino acid 1368 of UniProtKB - Q99ZW2).

[0307] N-terminus-[102-1368]-[optional linker]-[1-101]-C-terminus,

[0308] N-terminus-[1028-1368]-[optional linker]-[1-1027]-C-terminus,

[0309] N-terminus-[1041-1368]-[optional linker]-[1-1043]-C-terminus,

[0310] N-terminus-[1249-1368]-[optional linker]-[1-1248]-C-terminus, or

[0311] N-terminus-[1300-1368]-[optional linker]-[1-1299]-C-terminus, or the corresponding circularly permuted sequence of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).

[0312] In some embodiments, the circularly permuted Cas9 comprises any one of the following structures: (amino acid positions set forth in SEQ ID NO: 4449-amino acid 1368 of UniProtKB-N terminus of Q99ZW2-[103-1368]-[optional linker]-[1-102]-C terminus).

[0313] N-terminus-[1029-1368]-[optional linker]-[1-1028]-C-terminus,

[0314] N-terminus-[1042-1368]-[optional linker]-[1-1041]-C-terminus,

[0315] N-terminus-[1250-1368]-[optional linker]-[1-1249]-C-terminus, or

[0316] N-terminus-[1301-1368]-[optional linker]-[1-1300]-C-terminus, or the corresponding circularly permuted sequence of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).

[0317] In some embodiments, circular permutants are formed by joining a C-terminal fragment of Cas9 to an N-terminal fragment of Cas9 directly or using a linker, such as an amino acid linker. In some embodiments, the C-terminal fragment can correspond to 95% or more of the C-terminal amino acids of Cas9 (e.g., about amino acids 1300-1368 of SEQ ID NO: 4449 or their corresponding amino acid positions), or 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the C-terminal amino acids of Cas9 (e.g., SEQ ID NO: 4449). The N-terminal portion can correspond to 95% or more of the N-terminal amino acids of Cas9 (e.g., about amino acids 1-1300 set forth in SEQ ID NO: 4449 or their corresponding amino acid positions), or 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the N-terminal amino acids of Cas9 (e.g., those set forth in SEQ ID NO: 4449 or their corresponding amino acid positions).

[0318] In some embodiments, circular permutants are formed by joining a C-terminal fragment of Cas9 to an N-terminal fragment of Cas9 either directly or using a linker, such as an amino acid linker. In some embodiments, the N-terminally rearranged C-terminal fragment comprises or corresponds to 30% or less of the amino acids at the C-terminus of Cas9 (e.g., amino acids 1012-1368 set forth in SEQ ID NO: 4449, or their corresponding amino acid positions). In some embodiments, the N-terminally relocated C-terminal fragment comprises or corresponds to the C-terminal 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the Cas9 amino acids (e.g., as set forth in SEQ ID NO: 4449 or its corresponding amino acid positions). In some embodiments, the N-terminally relocated C-terminal fragment comprises or corresponds to the C-terminal 410 or fewer residues of Cas9 (e.g., as set forth in SEQ ID NO: 4449 or its corresponding amino acid positions). In some embodiments, the C-terminal portion relocated to the N-terminus comprises or corresponds to the C-terminal 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 residues of Cas9 (e.g., as set forth in SEQ ID NO: 4449 or its corresponding amino acid positions). In some embodiments, the C-terminal portion relocated to the N-terminus comprises or corresponds to the C-terminal 357, 341, 328, 120, or 69 residues of Cas9 (e.g., as set forth in SEQ ID NO: 4449 or their corresponding amino acid positions).

[0319] In other embodiments, circularly permuted Cas9 mutants can be topologically rearranged based on the S. pyogenes Cas9 of SEQ ID NO: 4449: (a) select a circularly permuted (CP) site corresponding to an internal amino acid residue in the Cas9 primary structure and divide the original protein into two halves, the N-terminal and C-terminal regions; and (b) modify (e.g., by genetic engineering) the Cas9 protein sequence to move the original C-terminal region (including the CP site amino acids) before the original N-terminal region, thereby forming a new N-terminus of the Cas9 protein starting with the CP site amino acid residue. The CP site can be located in any domain of the Cas9 protein, such as the helical II domain, RuvCIII domain, or CTD domain. For example, the CP site can be located at the original amino acid residue 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 (as set forth in SEQ ID NO: 4449 or its corresponding amino acid position). Thus, when relocated to the N-terminus, the original amino acid 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 becomes the new N-terminal amino acid. These CP-Cas9 proteins may be referred to as Cas9-CP181, Cas9-CP199, Cas9-CP230, Cas9-CP270, Cas9-CP310, Cas9-CP1010, Cas9-CP1016, Cas9-CP1023, Cas9-CP1029, Cas9-CP1041, Cas9-CP1247, Cas9-CP1249, and Cas9-CP1282, respectively. This description is not intended to be limited to creating CP variants from SEQ ID NO: 18, but can be implemented to create CP variants of any Cas9 sequence at CP sites corresponding to these positions, or at other CP sites throughout. This description is not intended to be limited to a particular CP site. Virtually any CP site can be used to generate CP-Cas9 variants.

[0320] In some embodiments, the prime editor comprises a Cas9 functional variant that has a smaller molecular weight than the wild-type SpCas9 protein. In some embodiments, the smaller Cas9 functional variant may facilitate delivery to cells, for example, via an expression vector, nanoparticle, or other delivery means. In certain embodiments, the smaller Cas9 functional variant is a Class 2 Type II Cas protein. In certain embodiments, the smaller Cas9 functional variant is a Class 2 Type V Cas protein. In certain embodiments, the smaller Cas9 functional variant is a Class 2 Type VI Cas protein.

[0321] In some embodiments, the prime editor comprises an SpCas9 that is 1368 amino acids in length and has a predicted molecular weight of 158 kilodaltons. In some embodiments, the prime editor comprises an SpCas9 that is less than 1300 amino acids, less than 1290 amino acids, less than 1280 amino acids, less than 1270 amino acids, less than 1260 amino acids, less than 1250 amino acids, less than 1240 amino acids, less than 1230 amino acids, less than 1220 amino acids, less than 1210 amino acids, less than 1200 amino acids, less than 1190 amino acids, less than 1180 amino acids, less than 1170 amino acids, less than 1160 amino acids, less than 1150 amino acids, less than 1140 amino acids, less than 1130 amino acids, or less than 1120 amino acids.

[0033] The term "Cas9" includes functional variants or fragments of less than about 400 amino acids, less than 1110 amino acids, less than 1100 amino acids, less than 1050 amino acids, less than 1000 amino acids, less than 950 amino acids, less than 900 amino acids, less than 850 amino acids, less than 800 amino acids, less than 750 amino acids, less than 700 amino acids, less than 650 amino acids, less than 600 amino acids, less than 550 amino acids, or less than 500 amino acids, but greater than about 400 amino acids, that retain one or more functions of the Cas9 protein, e.g., DNA binding.

[0322] In some embodiments, the Cas protein is Cas12a, Cas12b1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also called Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr In various other embodiments, the napDNAbp is any CRISPR-associated protein, including, but not limited to, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, preferably comprising a nickase mutation (e.g., a mutation corresponding to the D10A mutation in the wild-type Cas9 polypeptide of SEQ ID NO: 18). Cas9, Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), GeoCas9, CjCas9, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, circularly permuted Cas9, or Argonaute (Ago) domain, or a functional variant or fragment thereof. [Table 8]

[0323] In some embodiments, the prime editor described herein can comprise a Cas12a (Cpf1) polypeptide or a functional variant thereof. In some embodiments, the Cas12a polypeptide comprises a mutation that reduces or eliminates the endonuclease domain of the Cas12a polypeptide. In some embodiments, the Cas12a polypeptide is a Cas12a nickase. In some embodiments, the Cas protein comprises an amino acid sequence that comprises at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Cas12a polypeptide.

[0324] In some embodiments, the prime editor comprises a Cas protein that is a Cas12b (C2c1) or Cas12c (C2c3) polypeptide. In some embodiments, the Cas protein comprises an amino acid sequence that comprises at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Cas12b (C2c1) or Cas12c (C2c3) protein. In some embodiments, the Cas protein is a Cas12b nickase or a Cas12c nickase. In some embodiments, the Cas protein is a Cas12e, Cas12d, Cas13, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or CasΦ polypeptide. In some embodiments, the Cas protein comprises an amino acid sequence comprising at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Cas12e, Cas12d, Cas13, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or CasΦ protein. In some embodiments, the Cas protein is Cas12e, Cas12d, Cas13, or CasΦ nickase.

[0325] Flap endonucleases In some embodiments, the prime editor further comprises an additional polypeptide component, such as a flap endonuclease (FEN, e.g., FEN1). In some embodiments, the flap endonuclease excises the 5' single-stranded DNA of the edited strand of the target gene, helping to incorporate the intended nucleotide edit into the target gene. In some embodiments, the FEN is linked or fused to another component. In some embodiments, the FEN is provided in trans, for example, as a separate polypeptide or polynucleotide encoding the FEN.

[0326] In some embodiments, the prime editor or prime editing composition comprises a flap nuclease. In some embodiments, the flap nuclease is FEN1, or any FEN1 functional variant, functional mutant, or functional fragment thereof. In some embodiments, the flap nuclease is TREX2, EXO1, or other flap nucleases known in the art, or functional variants, functional mutants, or functional fragments thereof. In some embodiments, the flap nuclease has an amino acid sequence at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any of the flap nucleases described herein or known in the art.

[0327] nuclear localization sequence In some embodiments, the prime editor further comprises one or more nuclear localization sequences (NLSs). In some embodiments, the NLSs facilitate translocation of the protein to the cell nucleus. In some embodiments, the prime editor comprises a fusion protein, e.g., a fusion protein comprising a DNA-binding domain and a DNA polymerase comprising one or more NLSs. In some embodiments, one or more polypeptides of the prime editor are fused or linked to one or more NLSs. In some embodiments, the prime editor comprises a DNA-binding domain and a DNA polymerase domain provided in trans, wherein the DNA-binding domain and / or the DNA polymerase domain are fused or linked to one or more NLSs.

[0328] In certain embodiments, the prime editor or prime editing complex comprises at least one NLS. In some embodiments, the prime editor or prime editing complex comprises at least two NLSs. In embodiments with at least two NLSs, the NLSs may be the same or different.

[0329] In some cases, the prime editor may further comprise at least one nuclear localization sequence (NLS). In some cases, the prime editor may further comprise one NLS. In some cases, the prime editor may further comprise two NLSs. In some cases, the prime editor may further comprise three NLSs. In some cases, the prime editor may further comprise more than four, five, six, seven, eight, nine, or ten NLSs.

[0330] Additionally, an NLS can be expressed as part of the prime editor complex. In some embodiments, the NLS can be located almost anywhere in the amino acid sequence of the protein, typically comprising a short sequence of three or more or four or more amino acids. The location of the NLS fusion can be at the N-terminus, C-terminus, or anywhere within the sequence of the prime editor or its components (e.g., inserted N- to C-terminally or C- to N-terminally between the DNA binding domain and DNA polymerase domain of the prime editor fusion protein, between the DNA binding domain and linker sequence, between the DNA polymerase and linker sequence, or between two linker sequences of the prime editor fusion protein or its components). In some embodiments, the prime editor is a fusion protein comprising an NLS at the N-terminus. In some embodiments, the prime editor is a fusion protein comprising an NLS at the C-terminus. In some embodiments, the prime editor is a fusion protein comprising at least one NLS at both the N- and C-termini. In some embodiments, the prime editor is a fusion protein comprising two NLSs at the N- and / or C-termini.

[0331] Any NLS known in the art is contemplated herein. The NLS can be any naturally occurring NLS or any non-naturally occurring NLS (e.g., an NLS with one or more mutations compared to a wild-type NLS). In some embodiments, one or more NLSs of the prime editor comprise a bipartite NLS. In some embodiments, the nuclear localization signal (NLS) is predominantly basic. In some embodiments, one or more NLSs of the prime editor are rich in lysine and arginine residues. In some embodiments, one or more NLSs of the prime editor comprise proline residues. In some embodiments, the nuclear localization signal (NLS) comprises the sequence MDSLLMNRRKFLYQFKNVRWAKGRRETYLC, KRTADGSEFESPKKKRKV, KRTADGSEFEPKKKRKV, NLSKRPAAIKKAGQAKKKK, RQRRNELKRSF, or NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY.

[0332] In some embodiments, the NLS is a monopartite NLS. For example, in some embodiments, the NLS is the SV40 large T antigen NLS PKKKRKV. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the bipartite NLS comprises two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the bipartite NLS comprises two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, the spacer amino acid sequence comprises the Xenopus nucleoplasmic sequence KRXXXXXXXXXXKKKL (SEQ ID NO: 4451), where X is any amino acid. In some embodiments, the NLS is a non-canonical sequence, such as M9 of the hnRNP A1 protein, influenza virus nucleoprotein NLS, or yeast Gal4 protein NLS.

[0333] Other non-limiting examples of NLS sequences are shown in Table 9 below. [Table 9]

[0334] Additional Prime Editor Ingredients The prime editors described herein may contain additional functional domains, for example, one or more domains that modify the folding, solubility, or charge of the prime editor. In some cases, the prime editor may contain a solubility-enhancing (SET) domain.

[0335] In some embodiments, a split intein comprises two halves of an intein protein, which may be referred to as the N-terminal half of the intein, or intein N, and the C-terminal half of the intein, or intein C, respectively. In some embodiments, intein N and intein C may be fused to protein domains (N-terminal and C-terminal exteins), respectively. Exteins can be any protein or polypeptide, for example, any prime editor polypeptide component. In some embodiments, intein N and intein C of a split intein non-covalently bind to form an active intein, which can catalyze an α-trans-splicing reaction. In some embodiments, the trans-splicing reaction excises the two intein sequences and links the two extein sequences with a peptide bond. As a result, intein N and intein C are spliced ​​together, and the protein domain linked to intein N is fused to the protein domain linked to intein C in essentially the same manner as a continuous intein. In some embodiments, the split intein is derived from a eukaryotic intein, a bacterial intein, or an archaeal intein. Preferably, the resulting split intein contains only the amino acid sequence essential for catalyzing a trans-splicing reaction. In some embodiments, the intein-N or intein-C further comprises one or more amino acid substitutions compared to the wild-type intein-N or wild-type intein-C, e.g., amino acid substitutions that enhance the trans-splicing activity of the split intein. In some embodiments, the intein-C comprises 4 to 7 consecutive amino acid residues, at least 4 of which are derived from the last β-strand of the intein from which it was derived. In some embodiments, the split intein is derived from the Ssp DnaE intein (e.g., Synechocytis sp. PCC6803), or any intein or split intein known in the art, or a functional variant or fragment thereof.

[0336] In some embodiments, the prime editor comprises one or more epitope tags. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, thioredoxin (Trx) tags, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, polyhistidine tags (also called histidine tags or His tags), maltose binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag1, Softag3), strep tags, biotin ligase tags, FlAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein comprises one or more His tags.

[0337] In some embodiments, the prime editor comprises one or more polypeptide domains encoded by one or more reporter genes. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, and autofluorescent proteins such as green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and blue fluorescent protein (BFP).

[0338] In some embodiments, the prime editor comprises one or more polypeptide domains that bind to DNA molecules or other cellular molecules. Examples of binding proteins or domains include, but are not limited to, maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions.

[0339] In some embodiments, the prime editor comprises a protein domain that can modify the intracellular half-life of the prime editor.

[0340] In some embodiments, the prime editing complex comprises a fusion protein comprising a DNA-binding domain (e.g., Cas9(H840A)) and a reverse transcriptase (e.g., a mutant MMLV RT), the structure of which is [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)], and the desired PEG RNA. In some embodiments, the prime editing complex comprises a prime editor fusion protein having the amino acid sequence of SEQ ID NO: 4440.

[0341] Polypeptides comprising the components of a prime editor can be fused via a peptide linker or provided in trans relative to each other. For example, a reverse transcriptase can be expressed, delivered, or provided as a separate component rather than as part of a fusion protein with a DNA-binding domain. In such cases, the components of the prime editor can be associated through a non-peptide bond or colocalization function. In some embodiments, the prime editor further comprises additional components that can interact with, associate with, or recruit the prime editor or other components of the prime editing system. For example, the prime editor can include an RNA-protein recruitment polypeptide that can associate with an RNA-protein recruitment RNA aptamer. In some embodiments, the RNA-protein recruitment polypeptide can recruit or be recruited by a specific RNA sequence. Non-limiting examples of RNA-protein recruitment polypeptide and RNA aptamer pairs include MS2 coat protein and MS2 RNA hairpin, PCP polypeptide and PP7 RNA hairpin, Com polypeptide and Com RNA hairpin, Ku protein and telomerase Ku-binding RNA motif, and Sm7 protein and telomerase Sm7-binding RNA motif. In some embodiments, the prime editor comprises a DNA-binding domain fused or linked to the RNA-protein recruitment polypeptide. In some embodiments, the prime editor comprises a DNA polymerase domain fused or linked to the RNA-protein recruitment polypeptide. In some embodiments, the DNA-binding domain and DNA polymerase domain fused to the RNA-protein recruitment polypeptide, or the DNA-binding domain and DNA polymerase domain fused to the RNA-protein recruitment polypeptide, are co-localized with the corresponding RNA-protein recruitment RNA aptamer of the RNA-protein recruitment polypeptide. In some embodiments, the corresponding RNA-protein recruitment RNA aptamer is fused or linked to a portion of a PEGRNA or ngRNA.For example, an MS2 coat protein fused or linked to a DNA polymerase and an MS2 hairpin installed on a PEG RNA for colocalization of the DNA polymerase and an RNA-guided DNA-binding domain (e.g., Cas9 nickase).

[0342] In some embodiments, the prime editor comprises a polypeptide domain that recognizes an MS2 hairpin, the MS2 coat protein (MCP). In some embodiments, the nucleotide sequence of the MS2 hairpin (or also referred to as the "MS2 aptamer") is: GCCAACATGAGGATCACCCATGTCTGCAGGGCC. In some embodiments, the amino acid sequence of MCP is: GSASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKVATQTVGGEELPVAGWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIA ANSGIY.

[0343] In certain embodiments, the components of the prime editor are directly fused to each other, hi certain embodiments, the components of the prime editor are associated with each other via a linker.

[0344] As used herein, a linker is any chemical group or molecule that connects two molecules or moieties, such as a DNA binding domain and a polymerase domain of a prime editor. In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker comprises a non-peptide moiety. A linker can be as simple as a covalent bond or a polymeric linker that is many atoms in length, such as a polynucleotide sequence. In certain embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, a disulfide bond, a carbon-heteroatom bond, etc.).

[0345] In certain embodiments, two or more components of the prime editor are linked to each other by a peptide linker. In some embodiments, the peptide linker is 5 to 100 amino acids in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. In some embodiments, the peptide linker is 16, 24, 64, or 96 amino acids in length.

[0346] In some embodiments, the linker comprises the amino acid sequence (GGGGS)n, (G)n(, (EAAAK)n, (GGS)n, (SGGS)n, (XP)n, or any combination thereof, where n is independently an integer between 1 and 30, and X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS)n, where n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESAGGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSSGGS.

[0347] In some embodiments, the linker comprises 1 to 100 amino acids. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker comprises the amino acid sequence SGGSGGSGGS. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises the amino acid sequence GGSGGS(,GGSGSGSGGS,SGGSSGGSSGSETPGTSESATPESAGGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSSGGS, or SGGSSGGSSGSETPGTSESATPESSGGSSGGSS.

[0348] In certain embodiments, two or more components of the prime editor are linked to each other by a non-peptide linker. In some embodiments, the linker is a carbon-nitrogen bond of an amide bond. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched, aliphatic or heteroaliphatic linker. In certain embodiments, the linker is a polymer (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, acetic acid, alanine, β-alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic moiety (cyclopentane, cyclohexane, etc.). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In certain embodiments, the linker comprises an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker may comprise a functionalized moiety to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0349] The components of a prime editor can be connected to each other in any order. In some embodiments, the DNA-binding domain and DNA polymerase domain of a prime editor are fused to form a fusion protein or linked by a peptide or protein linker in any order from N-terminus to C-terminus. In some embodiments, a prime editor comprises a DNA-binding domain fused or linked to the C-terminus of a DNA polymerase domain. In some embodiments, a prime editor comprises a DNA-binding domain fused or linked to the N-terminus of a DNA polymerase domain. In some embodiments, a prime editor comprises a fusion protein comprising the structure NH2-[DNA-binding domain]-[polymerase]-COOH, or NH2-[polymerase]-[DNA-binding domain]-COOH, where each "]-[" indicates the presence of an optional linker sequence. In some embodiments, a prime editor comprises a fusion protein and a DNA polymerase domain provided in trans, where the fusion protein comprises the structure NH2-[DNA-binding domain]-[RNA-protein recruiting polypeptide]-COOH. In some embodiments, the prime editor comprises a fusion protein and a DNA binding domain provided in trans, wherein the fusion protein comprises the structure NH2-[DNA polymerase domain]-[RNA protein recruitment polypeptide]-COOH.

[0350] In some embodiments, a prime editor fusion protein, a polypeptide component of a prime editor, or a polynucleotide encoding a prime editor fusion protein or polypeptide component is split into polypeptides encoding N- and C-terminal halves, or N- and C-terminal halves, and separately delivered to target DNA in a cell. For example, in certain embodiments, a prime editor fusion protein is split into N- and C-terminal halves for separate delivery by an AAV vector, which then translates and colocalizes in the target cell to reform the complete polypeptide or prime editor protein. In such cases, the separate halves of the protein or fusion protein each contain a split intein, facilitating colocalization and reformation of the complete protein or fusion protein by the mechanism of intein-facilitated trans-splicing. In some embodiments, the prime editor comprises a polynucleotide or vector (e.g., an AAV vector) encoding the N-terminal half fused to intein N and the C-terminal half fused to intein C, or each of them. Once delivered and / or expressed in target cells, intein N and intein C are excised by protein trans-splicing to generate the complete prime editor fusion protein within the target cell.

[0351] In some embodiments, the prime editor fusion protein comprises a Cas9(H840A) nickase and wild-type M-MLV RT. In some embodiments, the prime editor fusion protein comprises one or more individual components of a prime editor fusion protein comprising a Cas9(H840A) nickase and wild-type M-MLV RT. In some embodiments, the prime editor fusion protein comprises a Cas9(H840A) nickase and M-MLV RT with amino acid substitutions D200N, T330P, T306K, W313F, and L603W compared to wild-type M-MLV RT. The amino acid sequence of an exemplary prime editor fusion protein comprising a Cas9(H840A) nickase and M-MLV RT with amino acid substitutions D200N, T330P, T306K, W313F, and L603W, the individual components of which are shown in Table 4. In some embodiments, the prime editor fusion protein comprises the complete amino acid sequence of Table 4. In some embodiments, the prime editor fusion protein comprises one or more individual components of Table 4.

[0352] In various embodiments, a prime editor fusion protein comprises an amino acid sequence that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to a prime editor fusion protein sequence described herein or known in the art. [Table 10-1] [Table 10-2] [Table 10-3]

[0353] Prime Editing In some embodiments, the present specification discloses compositions, systems, and methods that use prime editing compositions. The terms "prime editing composition" or "prime editing system" refer to compositions involved in the prime editing methods described herein. A prime editing composition may include a prime editor, e.g., a prime editor fusion protein, and a PEG RNA. A prime editing composition may further include additional elements, such as a second-strand nicked ngRNA. Components of a prime editing composition can be combined to form a complex for prime editing, or can be kept separate, e.g., for management purposes.

[0354] In some embodiments, the prime editing composition comprises a prime editor fusion protein complexed with a PEgRNA and, optionally, with an ngRNA. In some embodiments, the prime editing composition comprises a prime editor comprising a DNA-binding domain and a DNA polymerase domain associated with each other via a PEgRNA. For example, the prime editing composition comprises a prime editor comprising a DNA-binding domain and a DNA polymerase domain linked to each other by an RNA-protein recruitment aptamer RNA sequence, which is linked to a PEgRNA. In some embodiments, the prime editing composition comprises a PEgRNA and a polynucleotide, polynucleotide construct, or vector encoding the prime editor fusion protein.

[0355] In some embodiments, the prime editing composition comprises a PEgRNA, an ngRNA, and a polynucleotide, polynucleotide construct, or vector encoding the prime editor fusion protein. In some embodiments, the prime editing composition comprises multiple polynucleotides, polynucleotide constructs, or vectors, each encoding one or more components of the prime editing composition. In some embodiments, the PEgRNA of the prime editing composition is associated with the DNA-binding domain of the prime editor, e.g., Cas9 nickase. In some embodiments, the PEgRNA of the prime editing composition complexes with the DNA-binding domain of the prime editor and guides the prime editor to the target DNA.

[0356] In some embodiments, the prime editing composition comprises one or more polynucleotides encoding components of a prime editor and / or a PEgRNA or ngRNA. In some embodiments, the prime editing composition comprises a polynucleotide encoding a fusion protein comprising a DNA-binding domain and a DNA polymerase domain. In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding a fusion protein comprising a DNA-binding domain and a DNA polymerase domain, and (ii) a polynucleotide encoding a PEgRNA or PEgRNA. In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding a fusion protein comprising a DNA-binding domain and a DNA polymerase domain, (ii) a polynucleotide encoding a PEgRNA or PEgRNA, and (iii) a polynucleotide encoding a ngRNA or ngRNA. In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding a DNA-binding domain of a prime editor (e.g., Cas9 nickase), (ii) a polynucleotide encoding a DNA polymerase domain of a prime editor (e.g., reverse transcriptase), and (iii) a polynucleotide encoding a PEgRNA or PEgRNA. In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding a DNA binding domain of the prime editor (e.g., Cas9 nickase), (ii) a polynucleotide encoding a DNA polymerase domain of the prime editor (e.g., reverse transcriptase), (iii) a PEgRNA or a polynucleotide encoding the PEgRNA, and (iv) a ngRNA or a polynucleotide encoding the ngRNA.

[0357] In some embodiments, the polynucleotide encoding the DNA-binding domain or the polynucleotide encoding the DNA polymerase domain further encodes an additional polypeptide domain, such as an RNA-protein recruitment domain, such as an MS2 coat protein domain. In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding the N-terminal half of a prime editor fusion protein and intein N, and (ii) a polynucleotide encoding the C-terminal half of a prime editor fusion protein and intein C. In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding the N-terminal half of a prime editor fusion protein and intein N, (ii) a polynucleotide encoding the C-terminal half of a prime editor fusion protein and intein C, (iii) a PEgRNA or a polynucleotide encoding a PEgRNA, and / or (iv) a ngRNA or a polynucleotide encoding a ngRNA. In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding the N-terminal portion of the DNA-binding domain and intein N, and (ii) polynucleotides encoding the C-terminal portion of the DNA-binding domain, intein C, and a DNA polymerase domain. In some embodiments, the DNA-binding domain is a Cas protein domain (e.g., Cas9 nickase). In some embodiments, the prime editing composition comprises (i) a polynucleotide encoding an N-terminal portion of the DNA-binding domain and intein N, (ii) a polynucleotide encoding a C-terminal portion of the DNA-binding domain, intein C, and a DNA polymerase domain, (iii) a PEgRNA or a polynucleotide encoding a PEgRNA, and / or (iv) a ngRNA or a polynucleotide encoding a ngRNA.

[0358] In some embodiments, the prime editing system comprises one or more polynucleotides encoding one or more prime editor polypeptides, and the activity of the prime editing system can be temporally regulated by controlling the timing of delivery of the vector. For example, in some embodiments, a polynucleotide encoding a prime editor and a polynucleotide encoding a PEgRNA can be delivered simultaneously. For example, in some embodiments, a polynucleotide encoding a prime editor and a polynucleotide encoding a PEgRNA can be delivered sequentially.

[0359] In some embodiments, a polynucleotide encoding a component of a prime editing system may further comprise an element capable of modifying the intracellular half-life of the polynucleotide and / or regulating translational control. In some embodiments, the polynucleotide is RNA, e.g., mRNA. In some embodiments, the half-life of the polynucleotide, e.g., RNA, may be increased. In some embodiments, the half-life of the polynucleotide, e.g., RNA, may be decreased. In some embodiments, the element may increase the stability of the polynucleotide, e.g., RNA. In some embodiments, the element may decrease the stability of the polynucleotide, e.g., RNA. In some embodiments, the element may be within the 3' UTR of the RNA. In some embodiments, the element may include a polyadenylation signal (PA). In some embodiments, the element may include a cap, such as at the end of an upstream mRNA or PEG RNA. In some embodiments, the RNA may not include a PA, in which case it will be more rapidly degraded within the cell after transcription.

[0360] In some embodiments, the element can include at least one AU-rich element (ARE). The ARE can be bound by an ARE-binding protein (ARE-BP), depending on the tissue type, cell type, timing, cellular localization, and environment. In some embodiments, the destabilizing element can promote RNA degradation, affect RNA stability, or activate translation. In some embodiments, the ARE can include a length of 50-150 nucleotides. In some embodiments, the ARE can include at least one copy of the sequence AUUUA. In some embodiments, at least one ARE can be added to the 3' UTR of the RNA. In some embodiments, the element can be a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE). In further embodiments, the element is a modified and / or truncated WPRE sequence that can enhance expression from the transcript. In some embodiments, a WPRE or equivalent can be added to the 3' UTR of the RNA. In some embodiments, the element can be selected from other RNA sequence motifs that are enriched in fast-decaying or slow-decaying transcripts. In some embodiments, a polynucleotide (e.g., a vector) encoding a PE or PEgRNA can self-destruct by cleavage of a target sequence present on the polynucleotide (e.g., a vector). Cleavage can prevent continued transcription of the PE or PEgRNA.

[0361] The polynucleotide encoding a component of the prime editing composition can be DNA, RNA, or any combination thereof. In some embodiments, the polynucleotide encoding a component of the prime editing composition is an expression construct. In some embodiments, the polynucleotide encoding a component of the prime editing composition is a vector. In some embodiments, the vector is a DNA vector. In some embodiments, the vector is a plasmid. In some embodiments, the vector is a viral vector, such as a retroviral vector, an adenoviral vector, a lentiviral vector, a herpesvirus vector, or an adeno-associated viral vector (AAV).

[0362] In some embodiments, a polynucleotide encoding a polypeptide component of a prime editing composition is codon-optimized by replacing at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) of the native sequence with a codon more frequently or most frequently used in genes of the host cell while maintaining the native amino acid sequence. In some embodiments, a polynucleotide encoding a polypeptide component of a prime editing composition is operably linked to one or more expression control elements, such as a promoter, a 3' UTR, a 5' UTR, or any combination thereof. In some embodiments, a polynucleotide encoding a component of a prime editing composition is messenger RNA (mRNA). In some embodiments, the mRNA comprises a cap at the 5' end and / or a poly-A tail at the 3' end.

[0363] Pharmaceutical Composition The present specification discloses pharmaceutical compositions comprising any of the components of a prime editing composition, such as a prime editor, a fusion protein, a polynucleotide encoding a prime editor polypeptide, a PEGRNA, an ngRNA, and / or a prime editing complex described herein.

[0364] As used herein, the term "pharmaceutical composition" refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition includes additional agents, for example, for specific delivery, half-life extension, or other therapeutic compounds.

[0365] In some embodiments, pharmaceutically acceptable carriers include any medium that is involved in carrying or transporting a compound from one site in the body (e.g., a delivery site) to another (e.g., an organ, tissue, or body part), such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc, magnesium, calcium or zinc stearate, or stearic acid), or solvent encapsulant. A pharmaceutically acceptable carrier is "acceptable" in the sense of being compatible with the other ingredients of the formulation and not harmful to the tissues of interest (e.g., physiologically compatible, sterile, physiological pH, etc.).

[0366] Formulations of the pharmaceutical compositions described herein can be prepared by any method known or hereafter developed in the art of pharmacology. Generally, such preparation methods include combining the active ingredient(s) with an excipient and / or one or more other accessory ingredients, and then, if necessary and / or desirable, shaping and / or packaging the product into a desired single-dose or multi-dose unit. Pharmaceutical formulations may further include pharmaceutically acceptable excipients, which as used herein includes any solvents, dispersion media, diluents, or other liquid vehicles, dispersing or suspending aids, surfactants, isotonic agents, thickeners or emulsifiers, preservatives, solid binders, lubricants, and the like, appropriate for the particular dosage form desired.

[0367] How to edit The methods and compositions disclosed herein can be used to edit a target gene of interest by prime editing.

[0368] In some embodiments, the prime editing method comprises contacting a target gene with a PEgRNA and a prime editor (PE) polypeptide described herein. In some embodiments, the target gene is double-stranded and comprises two DNA strands that are complementary to each other. In some embodiments, contacting with the PEgRNA and contacting with the prime editor are performed sequentially. In some embodiments, contacting with the prime editor is performed after contacting with the PEgRNA. In some embodiments, contacting with the PEgRNA is performed after contacting with the prime editor. In some embodiments, contacting with the PEgRNA and contacting with the prime editor are performed simultaneously. In some embodiments, the PEgRNA and the prime editor form a complex before contacting the target gene.

[0369] In some embodiments, contacting the target gene with a prime editing composition causes the PEgRNA to bind to the target strand of the target gene. In some embodiments, contacting the target gene with a prime editing composition causes the PEgRNA to bind to the sequence of interest on the target strand of the target gene upon contact with the PEgRNA. In some embodiments, contacting the target gene with a prime editing composition causes the spacer sequence of the PEgRNA to bind to the sequence of interest on the target strand of the target gene.

[0370] In some embodiments, contacting a target gene with a prime editing composition results in the prime editor binding to the target gene, e.g., the target gene, when the PE composition contacts the target gene. In some embodiments, the DNA-binding domain of PE binds to the PEgRNA. In some embodiments, PE binds to the target gene directed by the PEgRNA. Thus, in some embodiments, contacting the target gene results in binding of the DNA-binding domain of the prime editor to the target gene directed by the PEgRNA.

[0371] In some embodiments, contacting the target gene with the prime editing composition results in a nick in the edited strand of the target gene, and contacting the prime editor with the target gene generates a nick in the edited strand of the target gene. In some embodiments, contacting the target gene with the prime editing composition results in a single-stranded DNA comprising a free 3' end at the site of the nick in the edited strand of the target gene. In some embodiments, contacting the target gene with the prime editing composition results in a nick in the edited strand of the target gene by the DNA-binding domain of the prime editor, thereby generating a single-stranded DNA comprising a free 3' end at the site of the nick. In some embodiments, the DNA-binding domain of the prime editor is a Cas domain. In some embodiments, the DNA-binding domain of the prime editor is Cas9. In some embodiments, the DNA-binding domain of the prime editor is Cas9 nickase.

[0372] In some embodiments, contacting the target gene with a prime editing composition causes the PEgRNA to hybridize to the 3' end of the nicked single-stranded DNA, thereby priming DNA polymerization by the DNA polymerase domain of the prime editor. In some embodiments, the free 3' end of the single-stranded DNA generated at the nick site hybridizes to the primer binding site sequence (PBS) of the contacted PEgRNA, thereby priming DNA polymerization. In some embodiments, the DNA polymerization is reverse transcription catalyzed by the reverse transcriptase domain of the prime editor. In some embodiments, the method includes contacting the target gene with a DNA polymerase, e.g., a reverse transcriptase, either as part of a prime editor fusion protein or prime editing complex (cis) or as a separate protein (trans).

[0373] In some embodiments, contacting the target gene with a prime editing composition generates edited single-stranded DNA encoded by the PEGRNA editing template by DNA polymerase-mediated polymerization from the 3' free end of the single-stranded DNA at the nick site. In some embodiments, the PEGRNA editing template contains one or more intended nucleotide edits relative to the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit is incorporated into the target gene by excision of the 5' single-stranded DNA of the edited strand of the target gene generated at the nick site and DNA repair. In some embodiments, the intended nucleotide edit is incorporated into the target gene by excision of the 5' single-stranded DNA of the edited strand generated at the nick site and DNA repair. In some embodiments, excision of the 5' single-stranded DNA of the edited strand generated at the nick site is performed by a flap endonuclease. In some embodiments, the flap nuclease is FEN1. In some embodiments, the method further includes contacting the target gene with a flap endonuclease. In some embodiments, the flap endonuclease is provided as part of a prime editor fusion protein. In some embodiments, the flap endonuclease is provided in trans.

[0374] In some embodiments, contacting the target gene with the primed editing composition generates a mismatched heteroduplex comprising an edited strand of the target gene, which comprises the edited single-stranded DNA, and an unedited target strand of the target gene. Without being bound by theory, endogenous DNA repair and replication can resolve the mismatched edited DNA and incorporate the nucleotide change(s) to form the desired edited target gene.

[0375] In some embodiments, the method further includes contacting the target gene with a nick guide (ngRNA) disclosed herein. In some embodiments, the ngRNA includes a spacer that binds to a second target sequence on the edited strand of the target gene. In some embodiments, the contacted ngRNA instructs PE to nick the target strand of the target gene. In some embodiments, the nick in the target strand (non-edited strand) causes endogenous DNA repair machinery to use the edited strand to repair the non-edited strand, thereby incorporating the intended nucleotide edits into both strands of the target gene and modifying the target gene. In some embodiments, the ngRNA includes a spacer sequence that is complementary to and can hybridize to the second target sequence on the edited strand only after the intended nucleotide edit(s) have been incorporated into the edited strand of the target gene.

[0376] In some embodiments, the target gene is contacted simultaneously with the ngRNA, PEgRNA, and PE. In some embodiments, the ngRNA, PEgRNA, and PE form a complex upon contacting the target gene. In some embodiments, the target gene is contacted sequentially with the ngRNA, PEgRNA, and prime editor. In some embodiments, the target gene is contacted with PE, and then the target gene is contacted with the ngRNA and / or PEgRNA. In some embodiments, the target gene is contacted with the ngRNA and / or PEgRNA before contacting the target gene with the prime editor.

[0377] In some embodiments, the target gene is in a cell. Accordingly, also provided herein are methods of modifying cells, such as human cells, human primary cells, and / or human iPSC-derived cells.

[0378] In some embodiments, the prime editing method involves introducing a PEgRNA, a prime editor, and / or an ngRNA into a cell harboring a target gene. In some embodiments, the prime editing method involves introducing a prime editing composition comprising a PEgRNA, a prime editor polypeptide, and / or an ngRNA into a cell harboring a target gene. In some embodiments, the PEgRNA, the prime editor polypeptide, and / or the ngRNA form a complex before being introduced into the cell. In some embodiments, the PEgRNA, the prime editor polypeptide, and / or the ngRNA form a complex after being introduced into the cell. The prime editor, the PEgRNA, and / or the ngRNA, and the prime editing complex can be introduced into a cell by any delivery approach described herein or any delivery approach known in the art, including physical techniques such as ribonucleoprotein (RNP), lipid nanoparticles (LNP), viral vectors, non-viral vectors, mRNA delivery, and cell membrane disruption with a microfluidic device. The prime editor, the PEgRNA and / or the ngRNA, and the prime editing complex can be introduced into a cell simultaneously or sequentially.

[0379] In some embodiments, the prime editing method comprises introducing into a cell a PEgRNA or a polynucleotide encoding the PEgRNA, a prime editor polynucleotide encoding a prime editor polypeptide, and optionally an ngRNA or a polynucleotide encoding the ngRNA. In some embodiments, the method comprises simultaneously introducing into a cell a PEgRNA or a polynucleotide encoding the PEgRNA, a polynucleotide encoding a prime editor polypeptide, and / or an ngRNA or a polynucleotide encoding the ngRNA. In some embodiments, the method comprises sequentially introducing into a cell a PEgRNA or a polynucleotide encoding the PEgRNA, a polynucleotide encoding a prime editor polypeptide, and / or an ngRNA or a polynucleotide encoding the ngRNA. In some embodiments, the method comprises introducing into a cell a polynucleotide encoding a prime editor polypeptide before introducing the PEgRNA or a polynucleotide encoding the PEgRNA and / or the ngRNA or a polynucleotide encoding the ngRNA. In some embodiments, a polynucleotide encoding a prime editor polypeptide is introduced into a cell and expressed in the cell before a PEgRNA or polynucleotide encoding a PEgRNA and / or an ngRNA or polynucleotide encoding an ngRNA is introduced into the cell. In some embodiments, a polynucleotide encoding a prime editor polypeptide is introduced into a cell after a PEgRNA or polynucleotide encoding a PEgRNA and / or an ngRNA or polynucleotide encoding an ngRNA is introduced into the cell.A polynucleotide encoding a prime editor polypeptide, a PEgRNA or a polynucleotide encoding a PEgRNA, and / or a ngRNA or a polynucleotide encoding a ngRNA can be introduced into a cell by any delivery approach described herein or known in the art, e.g., RNP, LNP, viral vector, non-viral vector, mRNA delivery, and physical delivery.

[0380] In some embodiments, a polynucleotide encoding a prime editor polypeptide, a polynucleotide encoding a PEgRNA, and / or a polynucleotide encoding an ngRNA is introduced into a cell and then integrated into the genome of the cell. In some embodiments, a polynucleotide encoding a prime editor polypeptide, a polynucleotide encoding a PEgRNA, and / or a polynucleotide encoding an ngRNA is introduced into a cell for transient expression. Accordingly, cells modified by prime editing are also provided herein.

[0381] In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a non-human primate cell, a bovine cell, a porcine cell, a rodent cell, or a murine cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is a human primary cell. In some embodiments, the cell is a progenitor cell. In some embodiments, the cell is a human progenitor cell. In some embodiments, the cell is a human cell harvested from an organ. In some embodiments, the cell is a primary human cell.

[0382] In some embodiments, the cell is a progenitor cell. In some embodiments, the cell is a stem cell. In some embodiments, the cell is an induced pluripotent stem cell. In some embodiments, the cell is an embryonic stem cell. In some embodiments, the cell is a retinal progenitor cell. In some embodiments, the cell is a retina precursor cell. In some embodiments, the cell is a fibroblast.

[0383] In some embodiments, the cells are human stem cells. In some embodiments, the cells are induced human pluripotent stem cells. In some embodiments, the cells are human embryonic stem cells. In some embodiments, the cells are human retinal progenitor cells. In some embodiments, the cells are human retinal progenitor cells. In some embodiments, the cells are human fibroblasts.

[0384] In some embodiments, the cells are primary cells. In some embodiments, the cells are human primary cells. In some embodiments, the cells are retinal cells. In some embodiments, the cells are photoreceptors. In some embodiments, the cells are rod cells. In some embodiments, the cells are cone cells. In some embodiments, the cells are human cells harvested from the retina. In some embodiments, the cells are human photoreceptors. In some embodiments, the cells are human rod cells. In some embodiments, the cells are human cone cells. In some embodiments, the cells are primary human photoreceptors derived from induced human pluripotent stem cells (iPSCs).

[0385] In some embodiments, the target gene edited by prime editing is within a chromosome of the cell. In some embodiments, the intended nucleotide edit is incorporated into the chromosome of the cell and inherited by progeny cells. In some embodiments, the intended nucleotide edit introduced into the cell by the prime editing compositions and methods is such that the cell and its progeny also contain the intended nucleotide edit. In some embodiments, the cell is autologous, allogeneic, or xenogeneic to the subject. In some embodiments, the cell is from or derived from the subject. In some embodiments, the cell is from or derived from a human subject. In some embodiments, after the intended nucleotide edit is incorporated by prime editing, the cell is returned to the subject (e.g., a human subject).

[0386] In some embodiments, the methods provided herein comprise introducing a prime editor polypeptide or a polynucleotide encoding a prime editor polypeptide, a PEgRNA or a polynucleotide encoding a PEgRNA, and / or an ngRNA or a polynucleotide encoding an ngRNA into a plur...

Claims

1. A prime-edited guide RNA (PEGRNA), (a) a spacer comprising a region complementary to a sequence to be searched for in a target strand of a double-stranded target DNA; (b) a guide RNA (gRNA) core capable of binding to a Cas protein; (c) an extension arm, i. an editing template containing an intended edit relative to the double-stranded target DNA; ii. the extension arm comprising a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in the non-target strand of the double-stranded target DNA; (d) a 3' nucleic acid motif selected from the group consisting of SEQ ID NOs: 1 to 15.

2. The prime editing guide RNA (PEgRNA) described in claim 1, wherein the 3' nucleic acid motif is a hairpin comprising an array selected from the group consisting of SEQ ID NOs: 1, 2, 3, 5, 6 and 7.

3. The prime editing guide RNA (PEgRNA) described in claim 1, wherein the 3' nucleic acid motif is a hairpin comprising the sequence of SEQ ID NO:

1.

4. The prime editing guide RNA (PEgRNA) described in claim 1, wherein the 3' nucleic acid motif is a hairpin comprising the sequence of SEQ ID NO:

2.

5. The prime editing guide RNA (PEgRNA) described in claim 1, wherein the 3' nucleic acid motif is a hairpin comprising the sequence of SEQ ID NO:

3.

6. The prime editing guide RNA (PEgRNA) described in claim 1, wherein the 3' nucleic acid motif is a hairpin comprising the sequence of SEQ ID NO:

5.

7. The prime editing guide RNA (PEgRNA) described in claim 1, wherein the 3' nucleic acid motif is a hairpin comprising the sequence of SEQ ID NO:

6.

8. The prime editing guide RNA (PEgRNA) described in claim 1, wherein the 3' nucleic acid motif is a hairpin comprising the sequence of SEQ ID NO:

7.

9. A prime-edited guide RNA (PEGRNA), (a) a spacer comprising a region complementary to a sequence to be searched for in a target strand of a double-stranded target DNA; (b) a guide RNA (gRNA) core capable of binding to a Cas protein; (c) an extension arm, (i) an editing template that contains an intended edit relative to the double-stranded target DNA; and (ii) the extension arm comprising a primer binding site (PBS) comprising a region complementary to a region upstream of a nick site in the non-target strand of the double-stranded target DNA; (d) a 3' nucleic acid motif, the 3' nucleic acid motif comprising: (A) A G-quadruplex or C-quadruplex derived from the VEGF gene promoter; (B) Pseudoknot derived from Potato Leaf Roll Virus (PLRV); (C) MS2 protein binding sequence; (D) Moloney murine leukemia virus (MMLV) reverse transcriptase recruitment sequence, and (E) the 3' nucleic acid motif comprising a sequence selected from the group consisting of Moloney murine leukemia virus (MMLV) replication recognition sequences; The prime editing guide RNA (PEgRNA) comprising:

10. The selected 3' nucleic acid motif is (i) the G-quadruplex or the C-quadruplex derived from a VEGF gene promoter; The G-quadruplex may comprise SEQ ID NO: 10; The C-quadruplex may comprise SEQ ID NO: 11, or (ii) the pseudoknot derived from Potato Leaf Roll Virus (PLRV); The pseudoknot may comprise SEQ ID NO: 4, or (iii) comprising the MS2 protein binding sequence; The MS2 protein binding sequence may be SEQ ID NO: 9, or (iv) comprising the MMLV reverse transcriptase recruitment sequence; The MMLV reverse transcriptase recruitment sequence may comprise SEQ ID NO: 8, or (v) comprises an MMLV replication recognition sequence; The PEG-RNA of claim 9, wherein the MMLV replication recognition sequence may comprise a sequence selected from the group consisting of SEQ ID NOs: 12 to 15.

11. the PEGRNA comprises, in 5' to 3' order, the spacer, the gRNA core, the editing template, the PBS, and the 3' nucleic acid motif; The PEG-RNA of claim 1, further comprising a linker immediately 5' to the 3' nucleic acid motif.

12. The linker is (a) does not form secondary structures, and / or (b) does not have perfect complementarity to the PBS sequence, and / or (c) does not have an exact complement to the editing template; and / or (d) the PEG-RNA of claim 11, which does not have perfect complementarity to the scaffold.

13. A prime editing system comprising: (a) a PEGRNA of any one of claims 1 to 12, or one or more polynucleotides encoding the PEGRNA; and (b) a Cas protein and a DNA polymerase, or a prime editor comprising one or more polynucleotides encoding the prime editor.

14. A lipid nanoparticle (LNP) or ribonucleoprotein (RNP) comprising the prime editing system of claim 13.

15. 15. A method for editing double-stranded target DNA, the method comprising contacting the target DNA with (a) a prime editor comprising the PEG RNA of any one of claims 1 to 12, and a Cas9 nickase and a reverse transcriptase; (b) a prime editing system of claim 13; or (c) an LNP or RNP of claim 14. The target DNA may be in a cell; The editing efficiency may be at least 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 2.1-fold, 2.2-fold, 2.3-fold, 2.4-fold, or 2.5-fold higher than the editing efficiency with a control PEG-RNA having the same spacer and extension arms, wherein the control PEG-RNA comprises a gRNA core having the sequence of SEQ ID NO: 16 and does not comprise a 3' nucleic acid motif or tag.