Genome editing compositions and methods of use
The prime editing composition with PEgRNAs and a Cas9 nickase system provides precise TRAC gene disruption, addressing graft-versus-host disease in allogeneic T cell therapy by minimizing undesirable genome editing outcomes, enhancing the safety and efficacy of cancer treatment.
Patent Information
- Application Number
- JP2025539898
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-01-05
- Publication Date
- 2026-01-28
AI Technical Summary
Existing allogeneic T cell immunotherapy for cancer treatment faces challenges such as graft-versus-host disease due to endogenous TCRs recognizing non-tumor antigens, and current genome editing methods like CRISPR-Cas9 introduce undesirable outcomes like complex mixtures of insertions and deletions.
A prime editing composition using specific PEgRNAs and a Cas9 nickase with a nuclease-inactivating mutation, combined with a reverse transcriptase and recombinase, enables precise disruption of the TRAC gene without inducing double-strand breaks, reducing alloreactive potential and minimizing graft-versus-host disease.
This approach achieves precise and targeted disruption of the TRAC gene, reducing the risk of graft-versus-host disease and ensuring safe and effective allogeneic T cell therapy for cancer treatment.
Smart Images

Figure 2026503266000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 478,654, filed January 5, 2023, and U.S. Provisional Patent Application No. 63 / 603,479, filed November 28, 2023, each of which is incorporated by reference herein in its entirety. [Background technology]
[0002] Adoptive T cell therapy is an emerging cancer treatment. Adoptive T cell therapy can involve ex vivo engineering of T cells to express engineered T cell receptors (TCRs) or chimeric antigen receptors (CARs) that recognize tumor-specific antigens. T cell therapy can involve autologous (i.e., patient-derived) T cells, thereby avoiding the problems of immunogenicity and intolerance upon reintroduction of ex vivo engineered cells. However, due to manufacturing challenges (e.g., time, cost, and low quality / quantity) of autologous T cell therapy, allogeneic (i.e., donor-derived) T cell immunotherapy has become an attractive alternative. A major obstacle to allogeneic cell therapy is that endogenous TCRs present on infused allogeneic T cells can recognize non-tumor antigens in the recipient, potentially leading to graft-versus-host disease (GvHD). The TCR is encoded by a single T cell receptor alpha constant (TRAC) gene and contains a TCR α chain complexed with TCR β chains encoded by two T cell receptor beta constant (TRBC) genes. The TRAC gene is located at 14q11.2 of the human genome. The TRAC mRNA is approximately 1.5 kb. Disruption of the endogenous TCR can be achieved by eliminating the expression of the TCRα chain, because the TCRαβ dimer is required for full TCR function. Therefore, the alloreactive potential of donor T cells to induce GvHD is expected to be reduced or eliminated by genetically modifying the TRAC gene to reduce or eliminate its expression.
[0003] Engineered cell receptors such as CARs can be used to redirect T cells to mediate tumor rejection.CARs can be transduced into T cells to generate CAR-T cells using retroviruses, lentiviruses, or other integration vectors.However, virus-mediated transduction leads to random integration into genome, so the application of such CAR-T cells may be limited by risks such as variable expression, transcriptional silencing, or even oncogenic transformation.
[0004] Recent advances in genome editing have enabled targeted manipulation of the human genome. For example, programmable nucleases such as CRISPR-Cas9 can create double-stranded DNA breaks (DSBs), which can disrupt genes by inducing a mixture of insertions and deletions (indels) at specific target sites. However, DSBs are associated with undesirable outcomes, including a complex mixture of products and transpositions. There is a need in the art for compositions and methods for targeted delivery of transgenes and precise disruption of TRAC genes without the need to introduce DSBs. Summary of the Invention
[0005] In one aspect, provided herein is a prime editing composition or system comprising: (A) a first prime editing guide RNA (PEgRNA), or one or more polynucleotides encoding the first PEgRNA; and (B) a second PEgRNA, or one or more polynucleotides encoding the second PEgRNA, wherein the first PEgRNA comprises: (i) a first spacer that is complementary to a first interrogation target sequence on a first strand of a TRAC gene; (ii) a first gRNA core that is capable of binding to a Cas9 protein; and (iii) a first editing template. and a first extension arm comprising a template and a first primer binding site (PBS), wherein the first spacer comprises at its 3' end nucleotides 4-20 of a sequence selected from the group consisting of SEQ ID NOs: 4, 61, 88, and 150, and the first PBS comprises at its 5' end a sequence that is the reverse complement of nucleotides 13-17 of the selected sequence, wherein the second PEGRNA comprises (i) a second spacer that is complementary to a second interrogation target sequence on a second strand of the TRAC gene that is complementary to the first strand, and (ii) a second spacer that is capable of binding to a Cas9 protein. and (iii) a second extension arm comprising a second editing template and a second PBS, wherein the second spacer comprises at its 3' end nucleotides 4 to 20 of a sequence selected from the group consisting of SEQ ID NOs: 177, 233, 260, 287, 314, 341, 368, 414, 441, 468, 495, 522, and 566, and the second PBS comprises at its 5' end a sequence that is the reverse complement of nucleotides 13 to 17 of the selected sequence; and (a) the first editing template comprises a region of complementarity to the second editing template; (b) the first editing template comprises nucleotides 8 to 17 of the sequence selected for the second spacer, and the second editing template comprises nucleotides 8 to 17 of the sequence selected for the first spacer, or (c) the first editing template comprises nucleotides 8 to 17 of the sequence selected for the second spacer and a region of complementarity to the second editing template, and the second editing template comprises nucleotides 8 to 17 of the sequence selected for the first spacer and a region of complementarity to the first editing template.
[0006] In some embodiments, the sequence selected for the first spacer is SEQ ID NO: 4 or 88.
[0007] In some embodiments, the sequence selected for the first spacer is SEQ ID NO:88.
[0008] In some embodiments, the sequence selected for the second spacer is SEQ ID NO: 177, 368, or 522.
[0009] In some embodiments, the sequence selected for the second spacer is SEQ ID NO:177.
[0010] In some embodiments, the first spacer and / or the second spacer is 16 to 22 nucleotides in length.
[0011] In some embodiments, the first spacer and / or the second spacer is 20 nucleotides in length and comprises a selected sequence.
[0012] In some embodiments, the first PBS is 8-17 nucleotides in length and comprises at its 5' end a sequence that is the reverse complement of nucleotides 10-17, 9-17, 8-17, 7-17, 6-17, 5-17, 4-17, 3-17, 2-17, or 1-17 of the selected sequence of the first spacer.
[0013] In some embodiments, the first PBS is 8 to 13 nucleotides in length.
[0014] In some embodiments, the first PBS is 10, 11, or 12 nucleotides in length.
[0015] In some embodiments, the second PBS is 7-17 nucleotides in length and comprises at its 5' end a sequence that is the reverse complement of nucleotides 11-17, 10-17, 9-17, 8-17, 7-17, 6-17, 5-17, 4-17, 3-17, 2-17, or 1-17 of the selected sequence of the second spacer.
[0016] In some embodiments, the second PBS is 8 to 13 nucleotides in length.
[0017] In some embodiments, the second PBS is 11, 12, or 13 nucleotides in length.
[0018] In some embodiments, the first gRNA core and the second gRNA core comprise the same sequence.
[0019] In some embodiments, the first gRNA core, the second gRNA core, or both comprise SEQ ID NO:590.
[0020] In some embodiments, the first spacer, the first gRNA core, the first editing template, and the first PBS form a continuous sequence within a single molecule.
[0021] In some embodiments, the first PEGRNA comprises, from 5' to 3', a first spacer, a first gRNA core, a first editing template, and a first PBS.
[0022] In some embodiments, the second spacer, the second gRNA core, the second editing template, and the second PBS form a contiguous sequence within a single molecule.
[0023] In some embodiments, the second PEGRNA comprises, from 5' to 3', a second spacer, a second gRNA core, a second editing template, and a second PBS.
[0024] In some embodiments, the first editing template comprises a region of complementarity to the second editing template.
[0025] In some embodiments, the first editing template and the second editing template each encode all or a fragment of a recombinase recognition sequence (RSS) or its reverse complement, wherein the first editing template encodes at least a 5' portion of the RSS or its reverse complement, and the second editing template encodes at least a 3' portion of the RSS or its reverse complement, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0026] In some embodiments, at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other, and optionally at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0027] In some embodiments, the first editing template encodes the RSS.
[0028] In some embodiments, the second editing template encodes the RSS.
[0029] In some embodiments, the RSS is an attB sequence recognized by Bxb1 recombinase.
[0030] In some embodiments, the RSS is an attP sequence recognized by Bxb1 recombinase.
[0031] In some embodiments, the RSS is the attB sequence recognized by the Pa01 recombinase.
[0032] In some embodiments, the RSS is an attP sequence recognized by Pa01 recombinase.
[0033] In some embodiments, a first editing template includes RTT number 1 from Table 6 and a second editing template includes RTT number 2 from the same RTT pair in Table 6, or alternatively, a first editing template includes RTT number 2 from Table 6 and a second editing template includes RTT number 1 from the same RTT pair in Table 6.
[0034] In some embodiments, the first editing template comprises SEQ ID NO:23 and the second editing template comprises SEQ ID NO:196.
[0035] In some embodiments, the first editing template comprises SEQ ID NO:24 and the second editing template comprises SEQ ID NO:106.
[0036] In some embodiments, the first editing template comprises SEQ ID NO:27 and the second editing template comprises SEQ ID NO:107.
[0037] In some embodiments, the first editing template comprises SEQ ID NO:106 and the second editing template comprises SEQ ID NO:24.
[0038] In some embodiments, the first editing template comprises SEQ ID NO:107 and the second editing template comprises SEQ ID NO:27.
[0039] In some embodiments, the first editing template comprises SEQ ID NO:25 and the second editing template comprises SEQ ID NO:197.
[0040] In some embodiments, the first editing template comprises SEQ ID NO:28 and the second editing template comprises SEQ ID NO:199.
[0041] In some embodiments, the first editing template comprises SEQ ID NO:22 and the second editing template comprises SEQ ID NO:195.
[0042] In some embodiments, the first editing template comprises SEQ ID NO:26 and the second editing template comprises SEQ ID NO:198.
[0043] In some embodiments, the first editing template comprises a 5' fragment of an RTT listed in Table 6, and the second editing template comprises the full length or 5' fragment of the corresponding RTT pair, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0044] In some embodiments, the second editing template comprises a 5' fragment of an RTT listed in Table 6, the first editing template comprises the full length or 5' fragment of the corresponding RTT pair, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0045] In some embodiments, at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other, and optionally at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0046] In some embodiments, the length of the region of complementarity of the first editing template is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90% of the length of the first editing template; optionally, the length of the region of complementarity of the first editing template is at least 52%, at least 53%, or at least 55% of the length of the first editing template.
[0047] In some embodiments, the length of the region of complementarity of the second editing template is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90% of the length of the second editing template, optionally, the length of the region of complementarity of the second editing template is at least 52%, at least 53%, or at least 55% of the length of the second editing template.
[0048] In some embodiments, (a) the first spacer comprises SEQ ID NO:4 and the first PBS comprises SEQ ID NO:13, or the first spacer comprises SEQ ID NO:88 and the first PBS comprises SEQ ID NO:96, or the first spacer comprises SEQ ID NO:88 and the first PBS comprises SEQ ID NO:157, or the first spacer comprises SEQ ID NO:88 and the first PBS comprises SEQ ID NO:99, or the first spacer comprises SEQ ID NO:88 and the first PBS comprises SEQ ID NO:97; (b) the second spacer comprises SEQ ID NO:177 and the second PBS comprises SEQ ID NO:188, or the second spacer comprises SEQ ID NO:368 and the second PBS comprises SEQ ID NO:376.
[0049] In some embodiments, (a) the first spacer comprises SEQ ID NO: 4 and the first PBS has the sequence of SEQ ID NO: 14 or SEQ ID NO: 15, or the first spacer comprises SEQ ID NO: 88 and the first PBS has the sequence of SEQ ID NO: 96 or SEQ ID NO: 98; (b) the second spacer comprises SEQ ID NO: 177 and the second PBS has the sequence of SEQ ID NO: 186 or SEQ ID NO: 188, or the second spacer comprises SEQ ID NO: 368 and the second PBS has the sequence of SEQ ID NO: 373 or SEQ ID NO: 374, or the second spacer comprises SEQ ID NO: 522 and the second PBS has the sequence of SEQ ID NO: 531 or 533.
[0050] In some embodiments, the first spacer comprises SEQ ID NO: 88, the first PBS has SEQ ID NO: 96, the second spacer comprises SEQ ID NO: 177, and the second PBS has the sequence of SEQ ID NO: 186.
[0051] In some embodiments, the first editing template comprises SEQ ID NO:27 and the second editing template comprises SEQ ID NO:107.
[0052] In some embodiments, the first editing template comprises SEQ ID NO:107 and the second editing template comprises SEQ ID NO:27.
[0053] In some embodiments, the first PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 36, 118, 123, and 1132-1134, and the second PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 1135-1142.
[0054] In some embodiments, the first PEgRNA comprises SEQ ID NO:118 and the second PEgRNA comprises SEQ ID NO:1140.
[0055] In some embodiments, the first PEgRNA comprises SEQ ID NO: 136 and the second PEgRNA comprises SEQ ID NO: 224.
[0056] In some embodiments, the first PEgRNA comprises SEQ ID NO: 51 and the second PEgRNA comprises SEQ ID NO: 126 and the second PEgRNA comprises SEQ ID NO: 220.
[0057] In some embodiments, a first PEgRNA comprises SEQ ID NO:44 and a second PEgRNA comprises SEQ ID NO:550.
[0058] In some embodiments, the first PEgRNA comprises SEQ ID NO:111 and the second PEgRNA comprises SEQ ID NO:1251.
[0059] In some embodiments, the first PEGRNA is selected from the group consisting of SEQ ID NOs: 30, 31, 32, 33, 34, 36, 38, 39, 43, 44, 49, 50, 51, 52, 53, 79, 80, 82, 109, 111, 112, 114, 115, 118, 120, 122, 123, 125, 126, 129, 130, 134, 135, 136, 140 , 141, 143, 168, 169, 171, 592, 593, 594, and 1132, and the second PEG RNA comprises a sequence selected from the group consisting of SEQ ID NOs: 203, 207, 210, 211, 212, 215, 217, 219, 220, 221, 222, 224, 226, 228, 251, 252, 254, 278, 279, 281, 305, 306, 308, 332, 333, 335, 359, 360, 362, 388, 390, 392, 395, 398, 397, 400, 401, 403, 404, 405, 408, 410, 432, 433, 435, 459, 460, 462, 486, 487, 489, 513, 514, 516, 541, 542, 543, 545, 546, 547, 549, 550, 552, 554, 555, 556, 558, 561, 562, 584, 585, 587, 591, 595, 597, 599, 601, 1127, 1128, 1135, 1136, 1137, 1138, 1139, 1140, and 1141.
[0060] In some embodiments, the first PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 30, 44, 109, and 126, and the second PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 207, 221, 388, and 400.
[0061] In some embodiments, the first PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 136, 141, 51, and 53, and the second PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 226, 403, 558, 556, and 224.
[0062] In some embodiments, the first PEgRNA comprises SEQ ID NO: 136 and the second PEgRNA comprises SEQ ID NO: 224.
[0063] In some embodiments, the first PEGRNA and / or the second PEGRNA further comprise a 3' motif, optionally connected to the 3' end of the first PBS or the second PBS via a linker.
[0064] In some embodiments, the first PEgRNA and / or the second PEgRNA further comprise 5'mN*mN*mN* and 3'mN*mN*mN*N modifications, where m indicates that the nucleotide comprises a 2'-O-Me modification and a* indicates the presence of a phosphorothioate linkage.
[0065] In some embodiments, the prime editing composition or system described in any one of the embodiments herein further comprises a prime editor, or one or more polynucleotides encoding the prime editor, wherein the prime editor comprises (a) a Cas9 nickase having a nuclease-inactivating mutation in its HNH domain, and (b) a reverse transcriptase.
[0066] In some embodiments, the Cas9 nickase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO:1007.
[0067] In some embodiments, the reverse transcriptase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO:1003.
[0068] In some embodiments, the prime editor is a fusion protein.
[0069] In some embodiments, the fusion protein comprises SEQ ID NO:1033.
[0070] In some embodiments, the one or more polynucleotides encoding the prime editor comprise (a) a first sequence encoding an N-terminal portion of a Cas9 nickase and an intein N, and (b) a second sequence encoding an intein C, a C-terminal portion of a Cas9 nickase, and a reverse transcriptase.
[0071] In some embodiments, the prime editing composition or system described in any one of the embodiments herein further comprises a recombinase that recognizes one or more recombinase recognition sequences (RSSs), or one or more polynucleotides encoding the recombinase.
[0072] In some embodiments, the recombinase is Bxb1 or Pa01.
[0073] In some embodiments, the recombinase is Bxb1 comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO:1131.
[0074] In some embodiments, sequence identity is determined by a Needleman-Wunsch alignment of the two protein sequences with gap costs set to 11 for presence and 1 for extension, where the percent identity is calculated by dividing the number of identities by the length of the alignment.
[0075] In some embodiments, the recombinase is fused or linked to the prime editor.
[0076] The prime editing composition or system described in any one of the embodiments herein further comprises a DNA polynucleotide comprising (a) a donor sequence and (b) a second RSS recognized by the recombinase.
[0077] In some embodiments, (i) the RSS comprises a Bxb1 attB sequence provided in Table 5 and the second RSS comprises a corresponding attP sequence provided in Table 5, or (ii) the RSS comprises a Bxb1 attP sequence provided in Table 5 and the second RSS comprises a corresponding attB sequence provided in Table 5.
[0078] In some embodiments, the RSS sequence comprises SEQ ID NO:1187 and the second RSS comprises SEQ ID NO:1188.
[0079] In some embodiments, the donor sequence comprises an open reading frame encoding a polypeptide.
[0080] In some embodiments, the donor sequence encodes a chimeric antigen receptor (CAR).
[0081] In some embodiments, the donor sequence encodes a CD19CAR.
[0082] In some embodiments, the donor sequence comprises a splice acceptor sequence.
[0083] In some embodiments, the prime editing composition or system described in any one of the embodiments herein further comprises one or more vectors comprising one or more polynucleotides encoding a first PEgRNA, one or more polynucleotides encoding a second PEgRNA, and one or more polynucleotides encoding a prime editor.
[0084] In some embodiments, a prime editing composition or system described in any one of the embodiments herein comprises one or more polynucleotides encoding a first PEgRNA, one or more polynucleotides encoding a second PEgRNA, one or more polynucleotides encoding a prime editor, one or more polynucleotides encoding a recombinase, and one or more vectors comprising a donor sequence.
[0085] In some embodiments, one or more of the vectors is an AAV vector.
[0086] In some embodiments, the one or more polynucleotides encoding the prime editor and / or the one or more polynucleotides encoding the recombinase are mRNA.
[0087] In some embodiments, the prime editing composition or system described in any one of the embodiments herein further comprises a third PEGRNA or a nucleic acid encoding the third PEGRNA, wherein the third PEGRNA comprises: (i) a third spacer that is complementary to an interrogation target sequence on the first strand of a β2-microglobulin (B2M) gene; (ii) a third gRNA core that can bind to a Cas9 protein; and (iii) a third extension arm that comprises: (a) a third editing template that comprises a region of complementarity to the editing target sequence on the second strand of the B2M gene; and (b) a third primer binding site (PBS) that comprises a sequence that is the reverse complement of a portion of the third spacer, wherein the first and second strands of the B2M gene are complementary to each other, and the third editing template encodes one or more nucleotide changes compared to the editing target sequence.
[0088] In some embodiments, the third spacer comprises SEQ ID NO: 1063 at its 3' end.
[0089] In some embodiments, the third PBS comprises at its 5' end a sequence that is the reverse complement of nucleotides 10-14 of SEQ ID NO:1063.
[0090] In some embodiments, the editing template encodes one or more in-frame stop codons or their complements in the B2M gene.
[0091] In some embodiments, the one or more in-frame stop codons comprise a nonsense mutation in the B2M gene.
[0092] In some embodiments, one or more in-frame stop codons comprise an insertion in the B2M gene.
[0093] In some embodiments, the editing template encodes the insertion of two in-frame stop codons in the B2M gene.
[0094] In some embodiments, the insertion is TAATAA or TTATTA.
[0095] In some embodiments, the editing template encodes a frameshift mutation in the B2M gene.
[0096] In some embodiments, the frameshift mutation is an insertion.
[0097] In some embodiments, the insertion is c.50insG or its complement.
[0098] In some embodiments, the frameshift mutation is a deletion.
[0099] In some embodiments, the deletion is c.51delC or its complement.
[0100] In some embodiments, the prime editing composition or system further comprises a third PEGRNA, or a nucleic acid encoding the third PEGRNA, comprising: a. a third spacer comprising SEQ ID NO: 1063 at its 3' end; b. a third gRNA core capable of binding to a Cas9 protein; c. a third editing template comprising, at its 3' end, (A) nucleotides 13-24 of SEQ ID NO: 1079, (B) nucleotides 12-20 of SEQ ID NO: 1085, or (C) nucleotides 7-17 of SEQ ID NO: 1089; and ii. a third extension arm comprising, at its 5' end, a third primer binding site (PBS) comprising a sequence that is the reverse complement of nucleotides 10-14 of SEQ ID NO: 1063.
[0101] In some embodiments, the third spacer is 17 to 22 nucleotides in length.
[0102] In some embodiments, the third spacer comprises any one of SEQ ID NOs: 1060-1063 at its 3' end.
[0103] In some embodiments, the third editing template comprises nucleotides 13-24 of SEQ ID NO:1077 at its 3' end.
[0104] In some embodiments, the third editing template comprises SEQ ID NO: 1077 at its 3' end.
[0105] In some embodiments, the third editing template comprises any one of SEQ ID NOs: 1078 or 1079 at its 3' end.
[0106] In some embodiments, the third editing template comprises nucleotides 12-20 of SEQ ID NO:1085 at its 3' end.
[0107] In some embodiments, the third editing template comprises SEQ ID NO: 1081 at its 3' end.
[0108] In some embodiments, the third editing template comprises any one of SEQ ID NOs: 1080, 1082-1085 at its 3' end.
[0109] In some embodiments, the third editing template comprises nucleotides 7-17 of SEQ ID NO: 1089 at its 3' end.
[0110] In some embodiments, the third editing template comprises any one of SEQ ID NOs: 1087-1089 at its 3' end.
[0111] In some embodiments, the third editing template has a length of 24 nucleotides or less.
[0112] In some embodiments, the third editing template has a length of 20 nucleotides or less.
[0113] In some embodiments, the third editing template has a length of 10 to 20 nucleotides.
[0114] In some embodiments, the third editing template has a length of 12 to 20 nucleotides.
[0115] In some embodiments, the third editing template has a length of 11 to 17 nucleotides.
[0116] In some embodiments, the third editing template is 16 to 24 nucleotides in length.
[0117] In some embodiments, the third PBS comprises the sequence set forth in any one of SEQ ID NOs: 1064-1076 at its 3' end.
[0118] In some embodiments, the third PBS has a length of 17 nucleotides or less.
[0119] In some embodiments, the third PBS is 8 to 15 nucleotides in length.
[0120] In some embodiments, the third PBS is 8 to 14 nucleotides in length.
[0121] In some embodiments, the third PBS is 12 nucleotides in length.
[0122] In some embodiments, the third spacer, the third gRNA core, the third editing template, and the third PBS form a continuous sequence within a single molecule.
[0123] In some embodiments, the single molecule comprises, from 5' to 3', a third spacer, a third gRNA core, a third editing template, and a third PBS.
[0124] In some embodiments, the third PEGRNA comprises a sequence selected from any one of SEQ ID NOs: 1090-1120.
[0125] In some embodiments, the prime editing composition or system described in any one of the embodiments herein further comprises a nick guide RNA (ngRNA), or a nucleic acid encoding the ngRNA, wherein the ngRNA comprises (i) an ngRNA spacer that is complementary to an ngRNA target sequence on the second strand of the B2M gene, and (ii) an ngRNA core that can bind to a Cas9 protein.
[0126] In some embodiments, the ngRNA spacer comprises at its 3' end a sequence corresponding to nucleotides 4 to 20, 3 to 20, 2 to 20, or 1 to 20 of any one of SEQ ID NOs: 1121-1126.
[0127] In some embodiments, the ngRNA spacer comprises any one of SEQ ID NOs: 1121-1126 at its 3' end.
[0128] In some embodiments, the ngRNA spacer comprises SEQ ID NO: 1126 at its 3' end.
[0129] In some embodiments, one or more nucleotides encoded by the editing template is c.51delC or its complement, and the ngRNA spacer comprises at its 3' end a sequence corresponding to nucleotides 4 to 20, 3 to 20, 2 to 20, or 1 to 20 of SEQ ID NO: 1124.
[0130] In some embodiments, one or more nucleotides encoded by the editing template is c.50insG or its complement, and the ngRNA spacer comprises at its 3' end a sequence corresponding to nucleotides 4 to 20, 3 to 20, 2 to 20, or 1 to 20 of SEQ ID NO: 1125 or 1126.
[0131] In some embodiments, the ngRNA comprises SEQ ID NO: 1129.
[0132] In one aspect, provided herein is an LNP comprising a prime editing composition or system described in any one of the embodiments herein.
[0133] In one aspect, provided herein is a pharmaceutical composition comprising a prime editing composition or system described in any one of the embodiments herein, or an LNP described in any one of the embodiments herein, and a pharmaceutically acceptable excipient.
[0134] In one aspect, provided herein is a method of editing a TRAC gene, the method comprising contacting the TRAC gene with (a) a prime editing composition or system described in any one of the embodiments herein, and a prime editor comprising a Cas9 nickase having a nuclease-inactivating mutation in its HNH domain, and a reverse transcriptase, or (b) a prime editing composition or system described in any one of the embodiments herein.
[0135] In some embodiments, the method further comprises contacting the TRAC gene with a recombinase or one or more polynucleotides encoding the recombinase, and a DNA polynucleotide comprising (a) a donor sequence and (b) one or more recombinase recognition sequences recognized by the recombinase.
[0136] In one aspect, provided herein is a method for inserting a donor sequence into a TRAC gene, the method comprising contacting the TRAC gene with a prime editing composition or system described in any one of the embodiments herein, or an LNP described in any one of the embodiments herein.
[0137] In one aspect, provided herein is a method of making a modified cell, the method comprising contacting a cell with a prime editing composition or system described in any one of the embodiments herein, or an LNP described in any one of the embodiments herein.
[0138] In some embodiments, the TRAC gene is cellular.
[0139] In some embodiments, the cell is a mammalian cell.
[0140] In some embodiments, the cells are human cells.
[0141] In some embodiments, the cell is an immune cell, and optionally, the cell is a T cell.
[0142] In some embodiments, the cell is in a subject.
[0143] In some embodiments, the cells are from a subject.
[0144] In some embodiments, the subject is a human.
[0145] In some embodiments, the method of any one of the embodiments herein further comprises editing the B2M gene.
[0146] In some embodiments, the editing comprises contacting the B2M gene with a prime editor comprising a prime editing composition or system described in any one of the embodiments herein, and a Cas9 nickase with a nuclease-inactivating mutation in its HNH domain, and a reverse transcriptase.
[0147] In one aspect, provided herein is a cell produced by a method according to any one of the embodiments herein.
[0148] In one aspect, provided herein is an edited TRAC gene that comprises GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAACC (SEQ ID NO: 1046) and / or GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 1047) relative to a wild-type TRAC gene.
[0149] In some embodiments, the edited TRAC gene comprises, from 5' to 3', an insert sequence comprising GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAACC (SEQ ID NO: 1046), a donor sequence, and GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 1047).
[0150] In some embodiments, the edited TRAC gene comprises, from 5' to 3', an insert sequence comprising GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 1047), a donor sequence, and GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAACC (SEQ ID NO: 1046).
[0151] In some embodiments, the donor sequence encodes a chimeric antigen receptor (CAR), and optionally, the donor encodes a CD19CAR.
[0152] In some embodiments, the insertion sequence is between a first chromosomal location and a second chromosomal location, wherein the first chromosomal location is selected from the group consisting of positions 22547458, 22547457, 22547449, and 22547448 on human chromosome 14, and the second chromosomal location is selected from the group consisting of positions 22547533, 22547523, 22547491, 22547528, 22547497, 22547579, 22547522, 22547485, 22547506, 22547560, 22547505, 22547529, and 22547490 on human chromosome 14.
[0153] In some embodiments, the insertion sequence is between positions 22547458 and 22547533 on human chromosome 14.
[0154] In some embodiments, the insertion sequence is between positions 22547458 and 22547522 on human chromosome 14.
[0155] In some embodiments, the insertion sequence is between positions 22547458 and 22547529 on human chromosome 14.
[0156] In some embodiments, the insertion sequence is between positions 22547449 and 22547533 on human chromosome 14.
[0157] In some embodiments, the insertion sequence is between positions 22547449 and 22547522 on human chromosome 14.
[0158] In some embodiments, the insertion sequence is between positions 22547449 and 22547529 on human chromosome 14.
[0159] In some embodiments, the T cell further comprises a premature stop codon compared to a wild-type B2M gene.
[0160] In some embodiments, the T cell further comprises a B2M gene that comprises a c.51delC edit compared to a wild-type B2M gene.
[0161] In some embodiments, the T cell further comprises a B2M gene comprising a c.50insG edit relative to a wild-type B2M gene.
[0162] In some embodiments, the T cell further comprises a B2M gene comprising a c.54insTAATAA edit relative to a wild-type B2M gene.
[0163] In some embodiments, the human chromosomal location and coding sequence location are as set forth in Genome Reference Consortium Human Build 38 (GrCh38).
[0164] In one aspect, provided herein is a prime editing composition or system comprising: (a) a prime editing guide RNA (PEgRNA), or one or more polynucleotides encoding the PEgRNA, wherein the PEgRNA comprises: (i) a spacer that is complementary to an interrogated target sequence on a first strand of a target gene; (ii) a gRNA core that can bind to a Cas9 nickase; and (iii) an extension arm that comprises a primer binding site (PBS) that comprises a region of complementarity to the second strand of the target gene and an editing template that encodes a recombinase recognition sequence (RSS) recognized by a Pa01 recombinase; (b) a prime editor, or one or more polynucleotides encoding the prime editor, comprising a Cas9 nickase capable of nicking the second strand of the target gene at a nick site and a reverse transcriptase; and (c) a Pa01 recombinase, or one or more polynucleotides encoding the Pa01 recombinase.
[0165] In some embodiments, the PBS comprises a region of complementarity to a region upstream of the nick site.
[0166] In some embodiments, the editing template comprises a region of complementarity to a region downstream of the nick site.
[0167] In some embodiments, the Cas9 nickase comprises a nuclease-inactivating mutation in the HNH domain.
[0168] In one aspect, provided herein is a prime editing composition or system comprising: (A) a first prime editing guide RNA (PEgRNA), or one or more polynucleotides encoding the first PEgRNA; (B) a second PEgRNA, or one or more polynucleotides encoding the second PEgRNA; (C) a prime editor, or one or more polynucleotides encoding the prime editor, comprising a Cas9 nickase having a nuclease-inactivating mutation in an HNH domain and a reverse transcriptase; and (D) a Pa01 recombinase, or one or more polynucleotides encoding the Pa01 recombinase, wherein the first PEgRNA comprises: (i) a first spacer that is complementary to a first interrogated target sequence on a first strand of a target gene; (ii) a first gRNA core that is capable of binding to the Cas9 nickase; and (iii) a first extension region that comprises a first editing template and a first primer binding site (PBS). and a second PEG gRNA comprising (i) a second spacer complementary to a second search target sequence on a second strand of the target gene complementary to the first strand, (ii) a second gRNA core capable of binding to Cas9 nickase, and (iii) a second extension arm comprising a second editing template and a second PBS, wherein the first editing template comprises a region of complementarity to the second editing template, the first editing template and the second editing template each encode all or a fragment of a recombinase recognition sequence (RSS) or its reverse complement, the first editing template encodes at least a 5' portion of the RSS or its reverse complement, and the second editing template encodes at least a 3' portion of the RSS or its reverse complement, at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other, and the RSS is recognized by Pa01 recombinase.
[0169] In some embodiments, at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0170] In some embodiments, the first editing template encodes the RSS.
[0171] In some embodiments, the second editing template encodes the RSS.
[0172] In some embodiments, the first editing template comprises a 5' fragment of an RTT listed in RTT pair 2 or 3 in Table 6, and the second editing template comprises the full length or a 5' fragment of the corresponding RTT pair, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0173] In some embodiments, the second editing template comprises a 5' fragment of an RTT listed in RTT pair 2 or 3 in Table 6, the first editing template comprises the full length or a 5' fragment of the corresponding RTT pair, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
[0174] In some embodiments, the first PBS comprises a region of complementarity to a region upstream of the nick site of the second strand, and the second PBS comprises a region of complementarity to a region upstream of the nick site of the first strand.
[0175] In some embodiments, the prime editing composition or system described in any one of the embodiments herein further comprises a DNA polynucleotide comprising (i) a second RSS recognized by the Pa01 recombinase and (ii) a donor sequence.
[0176] In some embodiments, the target gene is TRAC.
[0177] In one aspect, provided herein is a method of integrating a donor sequence into a target gene, the method comprising contacting the target gene with a prime editing composition or system described in any one of the embodiments herein. Incorporation by Reference
[0178] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0179] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings, in which: [Brief explanation of the drawings]
[0180] [Figure 1] A schematic diagram of a prime editing guide RNA (PEgRNA) binding to a double-stranded target DNA sequence is shown.
[0181] [Figure 2] An overview of PEGRNA architecture is shown with an exemplary schematic of a PEGRNA designed for a prime editor.
[0182] [Figure 3] Schematic diagram showing the spacer and gRNA core portions of an exemplary guide RNA in two separate molecules. The remainder of the PEGRNA structure is not shown.
[0183] [Figure 4A] FIG. 1 shows an exemplary schematic of a dual-prime editing system for editing both strands of double-stranded target DNA. The same color / shading indicates complementarity or identity between sequences.
[0184] [Figure 4B] An exemplary schematic of dual prime editing using a displacement duplex (RD) containing an overlapping duplex (OD) is shown. The same color / shading indicates complementarity or identity between the sequences.
[0185] [Figure 4C]An exemplary schematic of dual prime editing is shown. The same color / shade indicates complementarity or identity between sequences.
[0186] [Figure 4D] An exemplary schematic of dual prime editing is shown. The same color / shade indicates complementarity or identity between sequences.
[0187] [Figure 4E] An exemplary schematic of dual prime editing is shown. The same color / shade indicates complementarity or identity between sequences.
[0188] [Figure 4F] An exemplary schematic of dual prime editing is shown. The same color / shade indicates complementarity or identity between sequences.
[0189] [Figure 4G] An exemplary schematic of dual prime editing is shown. The same color / shade indicates complementarity or identity between sequences. DETAILED DESCRIPTION OF THE INVENTION
[0190] In some embodiments, systems, compositions, and methods are provided herein for editing the target gene T-cell receptor alpha constant (TRAC) by dual prime editing. The compositions provided herein may include a prime editor (PE), which may use an engineered guide polynucleotide, e.g., a prime editing guide RNA (PEgRNA), that can guide the PE to a specific DNA target and encode DNA edits that perform various functions on the target gene TRAC, including disrupting the target gene (e.g., by introducing a mutation or mutations in the target gene), introducing exogenous sequences (e.g., one or more recombinase recognition sequences) into the TRAC gene, and integrating a DNA donor (e.g., an expression cassette, ORF) into the gene. Also provided are compositions containing edited cells produced by the methods disclosed herein.
[0191] The following description and examples illustrate embodiments of the present disclosure in detail. It should be understood that the present disclosure is not limited to the specific embodiments described herein, and that such embodiments may vary. Those skilled in the art will recognize that the present disclosure has numerous variations and modifications that are also within the scope of the present invention. While various features of the present disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, although the present disclosure may be described in the context of separate embodiments for clarity in this specification, the present disclosure may also be implemented in a single embodiment.
[0192] definition Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0193] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, when used herein, the terms "including," "includes," "having," "has," "with," or variations thereof mean "comprising."
[0194] Unless otherwise specified, the words "comprising," "comprise," "comprises," "having," "have," "has," "including," "includes," "include," "containing," "contains," and "contain" are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
[0195] References to "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that a particular feature or characteristic described in connection with an embodiment is included in at least one or more embodiments of the present disclosure, but not necessarily in all embodiments.
[0196] The term "about" or "approximately" in connection with a numerical value means a range of values that falls within 10% more or less than that value. For example, about x means x ±(10%*x).
[0197] The term "between" means a range of numbers inclusive of the first and last numbers in the range.
[0198] As used herein, "cell" generally refers to a biological cell. A cell can be the basic structural, functional, and / or biological unit of an organism. A cell can originate from any organism that has one or more cells. A cell may not be derived from a natural organism (e.g., a cell may be synthetically produced, sometimes referred to as an artificial cell).
[0199] In some embodiments, the cells are human cells. The cells may be derived from different tissues, organs, and / or cell types. In some embodiments, the cells are primary cells. As used herein, the term "primary cells" refers to cells isolated from an organism, such as a mammal, and initially grown in tissue culture (i.e., in vitro) before being divided and transferred to subculture. In some non-limiting examples, mammalian cells, including primary cells and stem cells, can be modified by the introduction of one or more polynucleotides, polypeptides, and / or prime editing compositions (e.g., via transfection, transduction, electroporation, etc.) and further passaged.
[0200] Such modified cells include T cells, e.g., primary T cells, e.g., inflammatory T cells, T helper cells, cytotoxic T cells, CD4+ T cells, CD8+ T cells, memory T cells, regulatory T cells, natural killer T cells, mucosal-associated invariant T cells, gamma delta T cells, alpha beta T cells, naive T cells, or effector T cells, formed elements of the blood (e.g., lymphocytes, myeloid cells), their precursors or progenitors, differentiated or dedifferentiated cells, and stem cells. In some embodiments, the cell is a naive T cell (e.g., a naive CD8+ T cell). In some embodiments, the cell is a transformed T cell. In some embodiments, the cell is an immune cell (e.g., a primary immune cell) or a progenitor or precursor thereof. In some embodiments, the cell is a T cell, or a progenitor or precursor thereof. In some embodiments, the cell is a human T cell, or a progenitor or precursor thereof. In some embodiments, the cell is a T helper cell (e.g., Th1 cell, Th2 cell, Th9 cell, Th17 cell, Th22 cell, and Tfh (follicular helper) cell). In some embodiments, the cell is a cytotoxic T cell. In some embodiments, the cell is a CD8+ T cell. In some embodiments, the cell is a CD4+ T cell. In some embodiments, the cell is a memory T cell (e.g., a central memory T cell (T CM ), stem memory T cells (TSCM), effector memory T cells, tissue-resident memory T cells). In some embodiments, the cells are effector memory T cells (e.g., T EM Cells and T EMR A(CD45RA +) cells). In some embodiments, the cell is a regulatory T cell. In some embodiments, the cell is a natural killer T cell. In some embodiments, the cell is a mucosal-associated invariant T cell. In some embodiments, the cell is a γδ T cell. In some embodiments, the cell is an effector T cell. In some embodiments, the cell is a thymocyte. In some embodiments, the cell is a lymphocyte. In some embodiments, the cell is a common lymphoid progenitor cell. In some embodiments, the cell is an early thymic progenitor cell. In some embodiments, the cell is a CD3+ cell. In some embodiments, the cell is a tumor-infiltrating lymphocyte. In some embodiments, the cell is a myeloid cell. In some embodiments, the cell is a plasma cell. In some embodiments, the cell is an activated T cell.
[0201] In some embodiments, the cell is a stem cell (e.g., adult stem cell, embryonic stem cell, non-embryonic stem cell), umbilical cord blood stem cell, progenitor cell, bone marrow stem cell, induced pluripotent stem cell, totipotent stem cell, CD34+ cell, or hematopoietic stem cell). In some embodiments, the cell is a pluripotent cell (e.g., multipotent stem cell). In some embodiments, the cell (e.g., stem cell) is an embryonic stem cell, a tissue-specific stem cell, a mesenchymal stem cell, or an induced pluripotent stem cell. In some embodiments, the cell is an induced pluripotent stem cell (iPSC). In some embodiments, the cell is a hematopoietic stem cell. In some embodiments, the cell is a hematopoietic stem and progenitor cell. In some embodiments, the cell is a multipotent progenitor cell. In some embodiments, the cell is a T cell progenitor. In some embodiments, the cell is a T cell precursor. In some embodiments, the cell is an embryonic stem cell (ESC). In some embodiments, the cell is a human stem cell. In some embodiments, the cell is a human pluripotent stem cell. In some embodiments, the cell is a non-embryonic stem cell. In some embodiments, the cell is an induced human pluripotent stem cell. In some embodiments, the cell is a human stem cell. In some embodiments, the cell is a human embryonic stem cell. In some embodiments, the cell is a human T cell progenitor. In some embodiments, the cell is a human T cell precursor.
[0202] In some embodiments, the cell is a mammalian cell.
[0203] In some embodiments, the cells are not isolated from an organism, but form part of a tissue or organ of an organism, for example, a mammal.
[0204] In some embodiments, the cell is a differentiated cell. In some embodiments, the cell is differentiated from an induced pluripotent stem cell. In some embodiments, the cell is a T cell differentiated from an iPSC, an ESC, a T cell precursor, or a T cell progenitor, e.g., a primary T cell, e.g., an inflammatory T cell, a T helper cell, a cytotoxic T cell, a CD4+ T cell, a CD8+ T cell, a memory T cell, a regulatory T cell, a natural killer T cell, a mucosal-associated invariant T cell, a gamma delta T cell, an alpha beta T cell, a naive T cell, or an effector T cell.
[0205] In some embodiments, the cells are differentiated human cells. In some embodiments, the cells are differentiated from induced human pluripotent stem cells. In some embodiments, cells edited by prime editing can differentiate into or regenerate cells, e.g., T cells, e.g., primary T cells, e.g., inflammatory T cells, T helper cells, cytotoxic T cells, CD4+ T cells, CD8+ T cells, memory T cells, regulatory T cells, natural killer T cells, mucosal-associated invariant T cells, gamma delta T cells, alpha beta T cells, naive T cells, or effector T cell populations. In some embodiments, the cells are in a subject, e.g., a human subject. In some embodiments, the cells are obtained from a subject prior to editing. For example, in some embodiments, the cells are obtained from a patient with cancer, microbial infection, graft-versus-host disease, or an autoimmune disorder. Prior to editing by the methods and compositions disclosed herein, cells can be obtained from a subject through a variety of non-limiting methods. T cells can be obtained from several sources, including, but not limited to, peripheral blood mononuclear cells, bone marrow, lymph node tissue, umbilical cord blood, thymus tissue, tissue from an infection site, ascites, pleural effusion, spleen tissue, and tumors. Methods for harvesting blood cells, isolating and enriching T cells, and expanding them ex vivo can be by methods known in the art. In some embodiments, cells can be obtained from cell banks, blood banks, cell cultures, or any number of available T cell lines, known to those skilled in the art. Cells can also be obtained from tissue biopsies, surgery, blood, plasma, serum, or other biological fluids. In some embodiments, cells can be obtained from one or more healthy donors, or from patients with cancer, microbial infection, graft-versus-host infection, or autoimmune disorders, prior to editing.
[0206] For example, cells can be obtained (i.e., isolated or purified) from a whole blood sample by lysing red blood cells or a fractionated blood sample and removing peripheral mononuclear blood cells by centrifugation prior to editing. Cells can be further isolated or purified (e.g., by flow cytometry) using selective purification methods to isolate cells based on cell-specific markers such as CD25, CD3, CD4, CD8, CD28, CD45RA, or CD45RO. In one embodiment, CD4+ is used as a marker to select T cells. In one embodiment, CD8+ is used as a marker to select T cells. In one embodiment, CD4+ and CD8+ are used as markers to select regulatory T cells.
[0207] In some embodiments, edited cells produced using the methods and compositions disclosed herein are cultured, grown, expanded, differentiated, and / or dedifferentiated in vitro.
[0208] In some embodiments, the cell comprises a prime editor or a prime editing composition. In some embodiments, the cell comprises a dual prime editing composition or system comprising a prime editor and at least two PEgRNAs that are different from each other. In some embodiments, the cell is derived from a human subject. In some embodiments, the cell is derived from a human subject and comprises a prime editor or a prime editing composition for editing the TRAC gene. In some embodiments, the cell is derived from a human subject and the TRAC gene has been edited by prime editing. In some embodiments, the cell comprises a prime-edited TRAC gene and is administered to a subject (e.g., a human subject). In some embodiments, the cell is in a human subject and comprises a prime editor or a prime editing composition for editing the TRAC gene. In some embodiments, the cell is derived from a human subject and the TRAC gene has been edited or modified by prime editing. In some embodiments, the human subject is a healthy donor or has a disease, disorder, or condition, e.g., cancer, microbial infection, autoimmune disorder, T-cell malignancy, or graft-versus-host disorder. In some embodiments, the human subject is in need of, is undergoing, or will be undergoing immune cell immunotherapy (e.g., T cell therapy such as CAR-T cell therapy).
[0209] As used herein, the term "substantially" can refer to a value approaching 100% of a given value. In some embodiments, the term can refer to an amount that is at least about 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 99.99% of the total amount. In some embodiments, the term can refer to an amount that corresponds to about 100% of the total amount.
[0210] The terms "protein" and "polypeptide" can be used interchangeably to refer to a polymer of two or more amino acids linked by covalent bonds (such as amide bonds) and capable of adopting a three-dimensional structure. In some embodiments, a protein or polypeptide comprises at least 10, 15, 20, 30, or 50 amino acids linked by covalent bonds (e.g., amide bonds). In some embodiments, a protein comprises at least two amide bonds. In some embodiments, a protein comprises multiple amide bonds. In some embodiments, a protein comprises an enzyme, a zymogen protein, a regulatory protein, a structural protein, a receptor, a nucleic acid binding protein, a biomarker, a member of a specific binding pair (e.g., a ligand or aptamer), or an antibody. In some embodiments, a protein can be a full-length protein (e.g., a fully processed protein with a specific biological function). In some embodiments, a protein can be a variant or fragment of a full-length protein. For example, in some embodiments, a Cas9 protein domain comprises an H840A amino acid substitution compared to the naturally occurring S. pyogenes Cas9 protein. Variants of proteins or enzymes, e.g., variant reverse transcriptases, include polypeptides having an amino acid sequence that is about 60% identical, about 70% identical, about 80% identical, about 90% identical, about 95% identical, about 96% identical, about 97% identical, about 98% identical, about 99% identical, about 99.5% identical, or about 99.9% identical to the amino acid sequence of a reference protein.
[0211] In some embodiments, a protein comprises one or more protein domains or subdomains. As used herein, the terms "polypeptide domain," "protein domain," or "domain" used in the context of a protein or polypeptide refer to a polypeptide chain having one or more biological functions, such as, for example, a catalytic function, a protein-protein binding function, or a protein-DNA function. In some embodiments, a protein comprises multiple protein domains. In some embodiments, a protein comprises multiple naturally occurring protein domains. In some embodiments, a protein comprises multiple protein domains derived from different naturally occurring proteins. For example, in some embodiments, a prime editor can be a fusion protein comprising a Cas9 protein domain of S. pyogenes and a reverse transcriptase protein domain of a retrovirus (e.g., Moloney Murine Leukemia Virus) or a variant of said retrovirus. Proteins comprising amino acid sequences derived from proteins of different origins or naturally occurring may be referred to as fusion proteins or chimeric proteins.
[0212] In some embodiments, the protein comprises a functional variant or functional fragment of a full-length wild-type protein. As used herein, a "functional fragment" or "functional portion" refers to any portion of a reference protein (e.g., a wild-type protein) that contains less than the entire amino acid sequence of the reference protein but retains one or more functions, such as catalytic or binding functions. For example, a functional fragment of a reverse transcriptase may contain less than the entire amino acid sequence of a wild-type reverse transcriptase but retain the ability to catalyze the polymerization of polynucleotides under at least one set of conditions. When a reference protein is a fusion of multiple functional domains, the functional fragment can retain one or more functions of at least one functional domain. For example, a functional fragment of Cas9 may contain less than the entire amino acid sequence of wild-type Cas9 but retain DNA binding ability and partially or completely lack nuclease activity.
[0213] As used herein, a "functional variant" or "functional mutant" refers to any variant or mutant of a reference protein (e.g., a wild-type protein) that contains one or more changes to the amino acid sequence of the reference protein while retaining one or more functions, e.g., catalytic or binding functions. In some embodiments, the one or more changes to the amino acid sequence include amino acid substitutions, insertions, or deletions, or any combination thereof. In some embodiments, the one or more changes to the amino acid sequence include amino acid substitutions. For example, a functional variant of a reverse transcriptase may contain one or more amino acid substitutions compared to the amino acid sequence of a wild-type reverse transcriptase, but retain the ability to catalyze the polymerization of polynucleotides under at least one set of conditions. When the reference protein is a fusion of multiple functional domains, the functional variant can retain one or more functions of at least one functional domain. For example, in some embodiments, a functional variant of Cas9 may contain one or more amino acid substitutions in the nuclease domain (e.g., an H840A amino acid substitution) compared to the amino acid sequence of wild-type Cas9, but retain DNA binding ability and partially or completely lack nuclease activity.
[0214] As used herein, the term "functional" and its grammatical equivalents can refer to the ability to perform, have, or fulfill an intended purpose. Functionality can include any percentage from baseline up to 100% of the intended purpose. For example, functionality can include or comprise about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or up to about 100% of the intended purpose. In some embodiments, the term functional can mean greater than about 100% of normal function, e.g., 125%, 150%, 175%, 200%, 250%, 300%, 400%, 500%, 600%, 700%, or up to about 1000% of the intended purpose.
[0215] In some embodiments, a protein or polypeptide includes a naturally occurring amino acid (e.g., one of the 20 amino acids commonly found in naturally synthesized peptides, known by their single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V). In some embodiments, a protein or polypeptide includes a non-naturally occurring amino acid (e.g., an amino acid that is not one of the 20 amino acids commonly found in naturally synthesized peptides, including synthetic amino acids, amino acid analogs, and amino acid mimetics). In some embodiments, a protein or polypeptide is modified.
[0216] In some embodiments, a protein includes an isolated polypeptide. The term "isolated" means that the polypeptide is free, to varying degrees, from components that normally accompany it in its natural state or environment. For example, a polypeptide naturally present in a living animal is not isolated; the same polypeptide partially or completely separated from coexisting materials in its natural state is isolated.
[0217] In some embodiments, the protein is present in a cell, tissue, organ, or virus particle. In some embodiments, the protein is present in a cell or part of a cell (such as a bacterial cell, a plant cell, an animal cell, etc.). In some embodiments, the cell is in a tissue, a subject, or in cell culture. In some embodiments, the cell is a microorganism (such as a bacterium, a fungus, a protozoan, a virus, etc.). In some embodiments, the protein is present in a mixture of analytes (e.g., a lysate). In some embodiments, the protein is present in a lysate from multiple cells or a lysate of a single cell.
[0218] As used herein, the terms "homology," "homology," or "percentage of homology" refer to the degree of sequence identity between an amino acid sequence and a corresponding reference amino acid sequence, or between a polynucleotide sequence and a corresponding reference polynucleotide sequence. "Homology" can refer to similar polypeptide or polymer sequences, such as DNA sequences. Homology can refer to, for example, a nucleic acid sequence having at least approximately 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. In other embodiments, a "homologous sequence" of a nucleic acid sequence can exhibit 93%, 95%, or 98% sequence identity to a reference nucleic acid sequence. For example, a "region of homology to a genomic region" can be a region of DNA that has a sequence similar to a given genomic region in a genome. The homologous region can be of any length sufficient to facilitate binding of a spacer or primer binding site to a complementary sequence of the genomic region. For example, the homologous region can be at least 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5100, 5200, 5300, 5400, 550 , 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, or more bases in length so that the region of homology has sufficient homology to bind with the corresponding genomic region.
[0219] In the context of two nucleic acid sequences or two polypeptide sequences, when the percentage of sequence homology or identity is specified, the percentage of homology or identity usually refers to the alignment of the sequences over a portion of the length when two or more sequences are compared and aligned to obtain maximum correspondence.If a position in the compared sequences is occupied by the same base or amino acid, the molecules may be homologous at that position.Unless otherwise stated, sequence homology or identity is evaluated over the specified length of the nucleic acid, polypeptide, or part thereof.In some embodiments, homology or identity is evaluated over a functional or specified portion of the length.
[0220] Alignment of sequences to assess sequence homology can be performed by algorithms known in the art, such as the Basic Local Alignment Search Tool (BLAST) algorithm, described in Altschul et al., J. Mol. Biol. 215:403-410, 1990. A public internet interface for performing BLAST analyses is accessible through the National Center for Biotechnology Information. Additional known algorithms include Smith & Waterman, "Comparison of Biosequences," Adv. Appl. Math. 2:482, 1981; Needleman & Wunsch, "A general method applicable to the search for similarities in the amino acid sequence of two proteins," J. Mol. Biol. 48:443, 1970; Pearson & Lipman, "Improved tools for biological sequence comparison," Proc. Natl. Acad. Sci. USA 85:2444, 1988, or by automated implementations of these or similar algorithms. Global alignment programs can also be used to align similar sequences of approximately the same size. Examples of global alignment programs include NEEDLE (available at www.ebi.ac.uk / Tools / psa / emboss_needle / ), which is part of the EMBOSS package (Rice P et al., Trends Genet., 2000;16:276-277), and the GGSEARCH program https: / / fasta.bioch.virginia.edu / fasta_www2 / , which is part of the FASTA package (Pearson W and Lipman D, 1988, Proc. Natl. Acad. Sci. USA, 85:2444-2448). Both of these programs are based on the Needleman-Wunsch algorithm, which is used to find the optimal alignment (including gaps) along the entire length of two sequences.A detailed discussion of sequence analysis can also be found in Unit 19.3 of Ausubel et al. ("Current Protocols in Molecular Biology" John Wiley & Sons Inc, 1994-1998, Chapter 15, 1998). In some embodiments, the alignment between the query sequence and the reference sequence is performed by a Needleman-Wunsch alignment of the two protein sequences with gap costs set to 11 for presence and 1 for extension, where the percent identity is calculated by dividing the number of identities by the length of the alignment, as described in more detail in Altschul et al. ("Gapped BLAST and PSI-BLAST: a new generation of protein database search programs", Nucleic Acids Res. 25:3389-3402, 1997) and Altschul et al. ("Protein database searches using compositionally adjusted substitution matrices", FEBS J. 272:5101-5109, 2005).
[0221] Those skilled in the art will understand that amino acid (or nucleotide) positions in homologous sequences can be determined based on alignment, e.g., "H840" in a reference Cas9 sequence may correspond to H839 or another position in a Cas9 homolog.
[0222] The term "polynucleotide" or "nucleic acid molecule" refers to any polymeric form of nucleotides, including DNA, RNA, hybrids thereof, or RNA-DNA chimeric molecules. In some embodiments, the polynucleotide comprises cDNA, genomic DNA, mRNA, tRNA, rRNA, or microRNA. In some embodiments, the polynucleotide is double-stranded, e.g., double-stranded DNA within a gene. In some embodiments, the polynucleotide is single-stranded or substantially single-stranded, e.g., single-stranded DNA or mRNA. In some embodiments, the polynucleotide is a cell-free nucleic acid molecule. In some embodiments, the polynucleotide circulates in the blood. In some embodiments, the polynucleotide is an intracellular nucleic acid molecule. In some embodiments, the polynucleotide is an intracellular nucleic acid molecule circulating in the blood.
[0223] Polynucleotides can have any three-dimensional structure. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, ESTs, or SAGE tags), exons, introns, intergenic DNA (including, but not limited to, heterochromatic DNA), messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA, isolated RNA, sgRNA, guide RNA, nucleic acid probes, primers, snRNA, long non-coding RNA, snoRNA, siRNA, miRNA, small tRNA-derived RNA (tsRNA), antisense RNA, shRNA, or small rDNA-derived RNA (srRNA).
[0224] In some embodiments, a polynucleotide comprises deoxyribonucleotides, ribonucleotides, or their analogs. In some embodiments, a polynucleotide comprises modified nucleotides, such as methylated nucleotides or nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polynucleotide. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.
[0225] In some embodiments, a polynucleotide is composed of a specific sequence of four nucleotide bases: adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) in place of thymine when the polynucleotide is RNA. In some embodiments, a polynucleotide may contain one or more other nucleotide bases, such as inosine (I), which is read as guanine (G) by the translation machinery.
[0226] In some embodiments, a polynucleotide may be modified. As used herein, the term "modified" or "modification" refers to chemical modifications of A, C, G, T, and U nucleotides. In some embodiments, modifications may be made to the nucleoside base and / or sugar moieties of the nucleosides comprising the polynucleotide. In some embodiments, modifications may be made on the internucleoside linkage (e.g., the phosphate backbone). In some embodiments, a modified nucleic acid molecule comprises multiple modifications. In some embodiments, a modified nucleic acid molecule comprises a single modification.
[0227] As used herein, the terms "complementary," "complementary," or "complementarity" refer to the ability of two polynucleotide molecules to base pair with each other. Complementary polynucleotides can base pair through Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonds. For example, adenine on one polynucleotide molecule base pairs with thymine or uracil on a second polynucleotide molecule, and cytosine on one polynucleotide molecule base pairs with guanine on a second polynucleotide molecule. Two polynucleotide molecules are complementary to each other if the first polynucleotide molecule containing the first nucleotide sequence can base pair with the second polynucleotide molecule containing the second nucleotide sequence. For example, two DNA molecules 5'-ATGC-3' and 5'-GCAT-3' are complementary, and the complement of the DNA molecule 5'-ATGC-3' is 5'-GCAT-3'. The percentage of complementarity indicates the percentage of nucleotides in a polynucleotide molecule that can base pair with a second polynucleotide molecule (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). "Fully complementary" means that all contiguous nucleotides of a polynucleotide molecule base pair with the same number of contiguous nucleotides in a second polynucleotide molecule. As used herein, "substantially complementary" refers to a degree of complementarity that can be 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% across all or a portion of two polynucleotide molecules. In some embodiments, the portion of complementarity can be a region of 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides. "Substantial complementarity" can refer to 100% complementarity over a portion or region of two polynucleotide molecules. In some embodiments, the portion or region of complementarity between two polynucleotide molecules is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% of the length of at least one of the two polynucleotide molecules, or a functional or defined portion thereof.
[0228] As used herein, "expression" refers to the process by which a polynucleotide is transcribed into mRNA and / or the process by which a polynucleotide (e.g., a transcribed mRNA) is translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell. In some embodiments, expression of a polynucleotide, e.g., a gene or DNA encoding a protein, is determined by the amount of protein encoded by the gene after transcription and translation of the gene. In some embodiments, expression of a polynucleotide, e.g., a gene or DNA encoding a protein, is determined by the amount of functional form of protein encoded by the gene after transcription and translation of the gene. In some embodiments, expression of a gene is determined by the amount of mRNA, i.e., transcript, encoded by the gene after transcription of the gene. In some embodiments, expression of a polynucleotide, e.g., an mRNA, is determined by the amount of protein encoded by the mRNA after translation of the mRNA. In some embodiments, expression of a polynucleotide, e.g., an mRNA or coding RNA, is determined by the amount of functional form of protein encoded by the polypeptide after translation of the polynucleotide.
[0229] As used herein, the term "sequencing" may include capillary sequencing, sulfite-free sequencing, sulfite sequencing, TET-assisted sulfite (TAB) sequencing, ACE sequencing, high-throughput sequencing, Maxam-Gilbert sequencing, massively parallel signature sequencing, polony sequencing, 454 pyrosequencing, Sanger sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, heliscope single molecule sequencing, single molecule real-time (SMRT) sequencing, nanopore sequencing, shotgun sequencing, RNA sequencing, or a combination thereof.
[0230] The terms "equivalent" or "biological equivalent" are used interchangeably when referring to a particular molecule, or biological or cellular material, and refer to a molecule that has minimal homology to another molecule while maintaining a desired structure or function.
[0231] The term "encode" as applied to a polynucleotide refers to a polynucleotide that is said to "encode" another polynucleotide, polypeptide, or amino acid when, in its natural state or when manipulated by methods well known to those of skill in the art, it can be used as a template for polynucleotide synthesis, e.g., transcribed into RNA, reverse transcribed into DNA or cDNA, and / or translated to produce an amino acid, or polypeptide, or fragment thereof. In some embodiments, a polynucleotide comprising three consecutive nucleotides forms a codon that encodes a particular amino acid. In some embodiments, a polynucleotide comprises one or more codons that encode a polypeptide. In some embodiments, a polynucleotide comprising one or more codons comprises a mutation in the codon compared to a wild-type reference polynucleotide. In some embodiments, the codon mutation encodes an amino acid substitution in a polypeptide encoded by the polynucleotide compared to the wild-type reference polypeptide.
[0232] As used herein, the term "mutation" refers to a change and / or alteration in the amino acid sequence of a protein or the nucleic acid sequence of a polynucleotide. Such a change and / or alteration can include a substitution, insertion, deletion, and / or truncation of one or more amino acids, in the case of an amino acid sequence, and / or nucleotides, in the case of a nucleic acid sequence, compared to a reference amino acid or reference nucleic acid sequence. In some embodiments, the reference sequence is a wild-type sequence. In some embodiments, a mutation in the nucleic acid sequence of a polynucleotide encodes a mutation in the amino acid sequence of a polypeptide. In some embodiments, the mutation in the amino acid sequence of a polypeptide or the mutation in the nucleic acid sequence of a polynucleotide is a mutation associated with a disease state.
[0233] As used herein, the term "subject" and its grammatical equivalents can refer to a human or a non-human. The subject can be a mammal. A human subject can be male or female. A human subject can be of any age. A subject can be a human fetus. A human subject can be a newborn, infant, child, adolescent, or adult. A human subject can be in need of treatment for a genetic disease or disorder. A human subject can be in need of cell therapy, e.g., immune cell immunotherapy, such as T cell therapy. A human subject can be in need of CAR-T cell therapy.
[0234] The terms "treatment" or "treating," and their grammatical equivalents, may refer to the medical management of a subject with the intent to cure, ameliorate, or ameliorate the symptoms of a disease, condition, or disorder. Treatment may include active treatment, i.e., treatment focused on ameliorating the disease, condition, or disorder. Treatment may include causal treatment, i.e., treatment directed at eliminating the cause of the associated disease, condition, or disorder. Additionally, treatment may include palliative treatment, aimed at alleviating symptoms rather than curing the disease, condition, or disorder. Treatment may include supportive treatment, i.e., treatment used to complement another specific therapy directed at ameliorating the disease, condition, or disorder. In some embodiments, the condition may be pathological. In some embodiments, the disease, condition, or disorder may not be completely cured or prevented by treatment. In some embodiments, treatment improves but does not completely cure or prevent the disease, condition, or disorder. In some embodiments, a subject can receive treatment for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or for the life of the subject.
[0235] The term "ameliorate" and its grammatical equivalents mean to relieve, suppress, attenuate, diminish, arrest, or stabilize the onset or progression of a disease.
[0236] The terms "prevent" or "preventing" mean to delay, forestall, or avert the onset or progression of a disease, condition, or disorder for a period of time. Prevention can also mean reducing the risk of developing a disease, disorder, or condition. Prevention includes minimizing, partially, or completely inhibiting the onset of a disease, condition, or disorder. In some embodiments, the composition, e.g., pharmaceutical composition, prevents a disease by delaying the onset of the disease for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or for the lifetime of the subject.
[0237] The term "effective amount" or "therapeutically effective amount" refers to an amount of a composition, e.g., a prime editing composition, containing a construct, as disclosed herein, that may be sufficient to result in a desired activity when introduced into a subject. An effective amount of a prime editing composition can be provided to a target gene or cell, regardless of whether the cell is in vitro, ex vivo, or in vivo.
[0238] An effective amount can be, for example, an amount that induces at least about a 2-fold or greater change (increase or decrease) in the amount of target nucleic acid regulation observed compared to a negative control. An effective amount or dosage can induce, for example, about a 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 25-fold, 50-fold, 100-fold, 200-fold, 500-fold, 700-fold, 1000-fold, 5000-fold, or 10,000-fold increase in target gene regulation.
[0239] The amount of modulation of a target gene can be measured by any suitable method known in the art. In some embodiments, an "effective amount" or "therapeutically effective amount" is the amount of a composition required to improve disease symptoms compared to an untreated patient. In some embodiments, an effective amount is the amount of a composition sufficient to introduce an alteration into a gene of interest (e.g., a TRAC gene) in a cell (e.g., in vitro, ex vivo, or in vivo).
[0240] When referring to edited cells produced by the methods and compositions comprising edited cells disclosed herein, a "therapeutically effective amount" refers to the amount of a composition comprising edited cells that may be sufficient to produce a desired activity upon introduction into a subject (e.g., a human subject).
[0241] The term "construct" refers to a polynucleotide or a portion of a polynucleotide comprising one or more nucleic acid sequences encoding one or more transcription products and / or proteins. A construct can be a recombinant nucleic acid molecule or a portion thereof. In some embodiments, the one or more nucleic acid sequences of the construct are operably linked to one or more regulatory sequences, e.g., transcription initiation regulatory sequences. In some embodiments, the construct is a vector, a plasmid, or a portion thereof. In some embodiments, the construct comprises DNA. In some embodiments, the construct comprises RNA. In some embodiments, the construct is double-stranded. In some embodiments, the construct is single-stranded. In some embodiments, the construct comprises an expression cassette. An expression cassette refers to a polynucleotide comprising a nucleic acid sequence encoding one or more transcription products and operably linked to at least one transcription regulatory sequence, e.g., a promoter.
[0242] The term "exogenous," when used in reference to a biomolecule, e.g., a polynucleotide sequence or a polypeptide sequence, refers to a biomolecule that is not native to a particular biological context, e.g., a gene, a particular chromosome, a particular cell, or a chromosomal site in a cell, tissue, or organism, or, if from the same source, is modified from its original form or is present in a non-native location, e.g., a chromosomal location.
[0243] The term "endogenous," when used with respect to a biomolecule, e.g., a polynucleotide sequence or polypeptide sequence, refers to a biomolecule that is native or naturally occurring in a particular biological context, e.g., a gene, a particular chromosome, a particular cell, or a chromosomal site in a cell, tissue, or organism. For example, an endogenous sequence may be a wild-type sequence or may contain one or more mutations compared to the wild-type sequence. In some embodiments, an endogenous sequence is mutated compared to the wild-type sequence and may cause or be associated with a disease or disorder in a subject. As used herein, in some embodiments, a wild-type sequence is a gene sequence found in healthy individuals for a particular gene and a particular disease, and the wild-type sequence does not contain a mutation that causes the particular disease.
[0244] As used herein, the term "recombinase" refers to a site-specific enzyme that mediates DNA recombination between recombinase recognition sequences, resulting in the excision, integration, inversion, or exchange (e.g., transposition) of DNA fragments between the recombinase recognition sequences. Recombinases can be classified into two distinct families: serine recombinases (e.g., resolvases and invertases) and tyrosine recombinases (e.g., integrases). Examples of serine recombinases include, but are not limited to, Si74, No67, Kp03, Pa01, Nm60, BceINTa, BcytINTd, SscINTd, SacINTd, Hin, Gin, Tn3, I3-six, CinH, ParA, y6, Bxbl, OC31, TP901, TG1, pBT1, R4, pRV1, pFC1, MR11, A118, U153, and gp29. Examples of tyrosine recombinases include, but are not limited to, Cre, FLP, R, lambda, HK101, HK022, and pSAM2. Recombinases have many applications, including the generation of gene knockouts / knockins and gene therapy applications, as described in International Publication No. WO2020191248A1, the entire contents of which are incorporated herein by reference. The recombinases provided herein are not intended to be exclusive examples of recombinases that can be used in embodiments of the present invention. The methods and compositions of the present invention can be expanded by mining databases for novel orthogonal recombinases or by designing synthetic recombinases with defined DNA specificity.
[0245] In some embodiments, the catalytic domain of a recombinase is fused to the programmable DNA-binding domain of a prime editor, such as an RNA-programmable nuclease (e.g., dCas9, Cas9 nickase, or fragments thereof), such that the recombinase domain does not contain a nucleic acid-binding domain or is incapable of binding to a target nucleic acid (e.g., the recombinase domain is engineered so that it does not have specific DNA-binding activity). For example, serine recombinases of the resolvase-invertase family, such as Tn3- and 76-type resolvases and Hin and Gin invertases, have modular structures with autonomous catalytic and DNA-binding domains. The catalytic domains of these recombinases are therefore suitable for combination, e.g., fusion or conjugation, with a prime editor or components thereof, as described herein, e.g., after isolation of "active" recombinase mutants that do not require any additional factors (e.g., DNA-binding activity).
[0246] Furthermore, many other naturally occurring serine recombinases with N-terminal catalytic domains and C-terminal DNA-binding domains are known (e.g., phiC31 integrase, TnpX transposase, IS607 transposase), and their catalytic domains can be incorporated to engineer the programmable site-specific recombinases described herein. Similarly, the core catalytic domains of tyrosine recombinases (e.g., Cre, integrase) are known and can similarly be incorporated to engineer the programmable site-specific recombinases described herein.
[0247] Other examples of recombinases that are useful in the methods and compositions described herein will be known to those of skill in the art, and it is anticipated that any new recombinases discovered or produced may be used in different embodiments of the present invention.
[0248] As used herein, the term "recombinase recognition sequence," or equivalently, "RRS" or "recombinase target sequence" or "recombinase site," refers to a nucleotide sequence that is recognized by a recombinase and undergoes strand exchange with another DNA molecule having an RSS, resulting in the excision, integration, inversion, or exchange of a DNA fragment between the recombinase recognition sequences. In various embodiments, a prime editing composition can introduce one or more recombinase sites within a target sequence or multiple target sequences. When multiple recombinase sites are introduced by prime editing, the recombinase sites can be introduced into adjacent or non-adjacent target sites (e.g., separate chromosomes). In various embodiments, a single introduced recombinase site can be used as a "landing site" for a recombinase-mediated reaction between a genomic recombinase site and a second recombinase site in an exogenously supplied nucleic acid molecule, such as a plasmid or DNA vector. This allows for targeted integration of the desired nucleic acid molecule.
[0249] In the context of nucleic acid modification (e.g., genomic modification), the term "recombining" or "recombination" is used to refer to a process in which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein (e.g., the recombinase fusion proteins of the invention provided herein). Recombination can result in, among other things, the insertion, inversion, excision, or transposition of nucleic acids, for example, within or between one or more nucleic acid molecules.
[0250] Prime Editing and Dual Prime Editing The term "prime editing" refers to programmable editing of target DNA using a prime editor complexed with a PEgRNA to incorporate an intended nucleotide edit (also referred to herein as a nucleotide change) into the target DNA through target-primed DNA synthesis. In prime editing, the target DNA may comprise a double-stranded DNA molecule having two complementary strands. Considered in relation to each specific PEgRNA, the two complementary strands of the double-stranded target DNA may comprise a first strand, which may be referred to as the "target strand" or "non-edited strand," and a second strand, which may be referred to as the "non-target strand" or "edited strand." In some embodiments, in the prime editing guide RNA (PEgRNA), the spacer sequence is complementary or substantially complementary to a specific sequence on the target strand (which may also be referred to as the "search target sequence"). In some embodiments, the spacer sequence anneals to the target strand at the search target sequence. The target strand may also be referred to as the "non-protospacer adjacent motif (non-PAM strand)." In some embodiments, the non-target strand may also be referred to as the "PAM strand." In some embodiments, the PAM strand comprises a protospacer sequence and, optionally, a protospacer adjacent motif (PAM) sequence. In prime editing using a Cas protein-based prime editor, the PAM sequence refers to a short DNA sequence immediately adjacent to the protospacer sequence on the PAM strand of the target gene. The PAM sequence can be specifically recognized by a programmable DNA-binding protein, such as a Cas nickase or a Cas nuclease. In some embodiments, a particular PAM is characteristic of a particular programmable DNA-binding protein, such as a Cas nickase or a Cas nuclease. The protospacer sequence refers to a specific sequence within the PAM strand of the target gene that is complementary to the interrogated target sequence. In PEG RNA, the spacer sequence can have substantially the same sequence as the protospacer sequence on the edited strand of the target gene, except that the spacer sequence can include uracil (U) and the protospacer sequence can include thymine (T).
[0251] In some embodiments, the double-stranded target DNA contains a nick site on the PAM strand (or non-target strand). As used herein, "nick site" refers to a specific position between two nucleotides or two base pairs in the double-stranded target DNA. In some embodiments, the position of the nick site is specific relative to the position of a specific PAM sequence. In some embodiments, the nick site is specific where a nick occurs when the double-stranded target DNA is contacted with a nickase, e.g., a Cas nickase, that recognizes a specific PAM sequence. In some embodiments, the nick site is upstream of the specific PAM sequence on the PAM strand of the double-stranded target DNA. In some embodiments, the nick site is downstream of the specific PAM sequence on the PAM strand of the double-stranded target DNA. In some embodiments, the nick site is upstream of the PAM sequence recognized by Cas9 nickase, which comprises a nuclease-active RuvC domain and a nuclease-inactive HNH domain. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by Streptococcus pyogenes Cas9 nickase, P. lavamentivorans Cas9 nickase, C. diphtheriae Cas9 nickase, N. cinerea Cas9, S. aureus Cas9, or N. lari Cas9 nickase. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a Cas9 nickase, and the Cas9 nickase comprises a nuclease-active RuvC domain and a nuclease-inactive HNH domain. In some embodiments, the nick site is 2 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a S. thermophilus Cas9 nickase, and the S. thermophilus Cas9 nickase comprises a nuclease-active RuvC domain and a nuclease-inactive HNH domain.
[0252] The "primer binding site" (also referred to as the PBS or primer binding site sequence) is a single-stranded portion of the PEG RNA that contains a region of complementarity to the PAM strand (i.e., the non-target strand or edited strand). The PBS is complementary or substantially complementary to a sequence on the PAM strand of the double-stranded target DNA immediately upstream of the nick site. In some embodiments, in the process of prime editing, the PEG RNA forms a complex with the prime editor, directing the prime editor to bind to the interrogating target sequence on the target strand of the double-stranded target DNA and generating a nick at the nick site on the non-target strand of the double-stranded target DNA. In some embodiments, the PBS is complementary or substantially complementary to and can anneal to the free 3' end on the non-target strand of the double-stranded target DNA at the nick site. In some embodiments, the PBS annealed to the free 3' end of the non-target strand can initiate target-priming DNA synthesis.
[0253] The "editing template" of a PEgRNA is the single-stranded portion of the PEgRNA that is 5' to the PBS and encodes one strand of DNA. The editing template may contain a region of complementarity to the PAM strand (i.e., the non-target or edited strand) and contains one or more intended nucleotide edits compared to the endogenous sequence of the double-stranded target DNA. In some embodiments, the editing template and the PBS are immediately adjacent to each other. Thus, in some embodiments, the PEgRNA undergoing prime editing contains a single-stranded portion that includes the PBS and the editing template immediately adjacent to each other. In some embodiments, the single-stranded portion of the PEgRNA that includes both the PBS and the editing template is complementary or substantially complementary to the endogenous sequence on the PAM strand (i.e., the non-target or edited strand) of the double-stranded target DNA, except for one or more non-complementary nucleotides at the intended nucleotide edit positions. As used herein, regardless of their relative 5'-3' arrangement in other contexts, the relative positions between the PBS and the editing template, and between elements of the PEgRNA, are determined by the 5' to 3' order of the PEgRNA as a single molecule, regardless of the location of sequences within the double-stranded target DNA that may have complementarity or identity to elements of the PEgRNA. In some embodiments, the editing template is complementary or substantially complementary to a sequence on the PAM strand immediately downstream of the nick site, except for one or more non-complementary nucleotides at the intended nucleotide editing position. An endogenous, e.g., genomic, sequence that is complementary or substantially complementary to the editing template may be referred to as an "editing target sequence," except for one or more non-complementary nucleotides at the position corresponding to the intended nucleotide edit. In some embodiments, the editing template is complementary to, or has identity or substantial identity with, a sequence on the target strand having the same genomic location as, the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide editing position. In some embodiments, the editing template encodes a single-stranded DNA that has identity or substantial identity to the editing target sequence except for one or more insertions, deletions, or substitutions at the position(s) of one or more intended nucleotide edits.In some embodiments, the editing template may encode a wild-type or non-disease-associated gene sequence (or its complement if the edited strand is the antisense strand of the gene). In some embodiments, the editing template may encode a wild-type or non-disease-associated protein, but may contain one or more synonymous mutations compared to the coding region of the wild-type or non-disease-associated protein. Such synonymous mutations may include, for example, mutations that reduce the ability of the PEGRNA to rebind to the same target sequence once the desired edit is introduced into the genome (e.g., a synonymous mutation that silences the endogenous PAM sequence or a synonymous mutation that edits the endogenous protospacer).
[0254] In some embodiments, the PEgRNA forms a complex with a prime editor, directing the prime editor to bind to the interrogated target sequence of the target gene. In some embodiments, the bound prime editor generates a nick in the edited strand (PAM strand) of the target gene at the nick site. In some embodiments, the primer binding site (PBS) of the PEgRNA anneals to the free 3' end formed at the nick site, and the prime editor initiates DNA synthesis from the nick site using the free 3' end as a primer. A single-stranded DNA encoded by the PEgRNA editing template is then synthesized. In some embodiments, the newly synthesized single-stranded DNA contains one or more intended nucleotide edits compared to the endogenous target gene sequence. Thus, in some embodiments, the PEgRNA editing template is complementary to a sequence in the edited strand, except for one or more mismatches at the intended nucleotide editing positions in the editing template. An endogenous, e.g., genomic, sequence that is partially complementary to the editing template may be referred to as an "editing target sequence." Thus, in some embodiments, the newly synthesized single-stranded DNA has identity or substantial identity to a sequence within the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide editing positions. In some embodiments, the editing template comprises at least four consecutive nucleotides that are complementary to the edited strand, and the at least four consecutive nucleotides are located upstream of the 5'-most edit in the editing template.
[0255] In some embodiments, prime editing can include programmable editing of target DNA using one or more prime editors, each complexed with a PEgRNA ("dual prime editing"). Dual prime editing refers to programmable editing of double-stranded target DNA using two or more PEgRNAs, each of which is complexed with a prime editor to incorporate one or more intended nucleotide edits into the double-stranded target DNA. In some embodiments, dual prime editing incorporates one or more intended nucleotide edits into double-stranded target DNA via excision of endogenous DNA segments and / or replacement of endogenous DNA segments with newly synthesized DNA via target-primed DNA synthesis. In some embodiments, dual prime editing can be used to edit target DNA that is, or is part of, a target gene. In some embodiments, the target gene is a disease-associated gene. In some embodiments, the target gene is a monogenic disease-associated gene. In some embodiments, the target gene is a polygenic disease-associated gene. In some embodiments, the target gene is mutated compared to the wild-type sequence of the same gene and may cause or be associated with a disease or disorder in a subject. In some embodiments, the mutated target gene causes a disease or disorder in a human subject.
[0256] In some embodiments, dual prime editing involves using two different PEgRNAs, each complexed with a prime editor, where each of the two PEgRNAs includes a spacer that is complementary or substantially complementary to a distinct interrogation target sequence. In some embodiments, each of the two PEgRNAs anneals to a distinct interrogation target sequence via its spacer. Thus, references to a "PAM strand," "non-PAM strand," "target strand," "non-target strand," "editing strand," or "non-editing strand" are relative in the context of a particular PEgRNA, e.g., one of the two PEgRNAs in dual prime editing.
[0257] In some embodiments, dual prime editing involves two distinct PEgRNAs, each complexed with a prime editor. In some embodiments, each of the two PEgRNAs comprises a region of complementarity to a distinct interrogation target sequence of the target DNA, where the two distinct interrogation target sequences are on two complementary strands of the target DNA. The terms "region," "portion," and "segment" are used interchangeably to refer to a portion of a molecule, e.g., a polynucleotide or polypeptide. For example, a region of a polynucleotide can be 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the polynucleotide. In some embodiments, the two PEgRNAs can each direct a prime editor to initiate a prime editing process on two complementary strands of the target DNA.
[0258] In some embodiments, dual prime editing involves two PEGRNAs, each complexed with a prime editor. In some embodiments, the first PEGRNA comprises a first spacer complementary to a first interrogation target sequence on the first strand of a double-stranded target DNA, e.g., a double-stranded target gene. In the context of the first PEGRNA, the first strand of the double-stranded target DNA may be referred to as the first target strand, and the complementary strand may be referred to as the first PAM strand.
[0259] In some embodiments, the second PEgRNA comprises a second spacer complementary to a second interrogation target sequence on the second strand of the double-stranded target DNA. In some embodiments, the first and second strands of the double-stranded target DNA, e.g., a double-stranded target gene, are complementary to each other. Thus, in some embodiments, the second PEgRNA and the first PEgRNA bind to opposite strands of the double-stranded target DNA. In the context of the second PEgRNA, the second strand of the double-stranded target DNA may be referred to as the second target strand, and the complementary strand may be referred to as the second PAM strand. In some embodiments, the first target strand is the same strand as the second PAM strand of the double-stranded target DNA. In some embodiments, the second target strand is the same strand as the first PAM strand of the double-stranded target DNA.
[0260] In some embodiments, a first PEgRNA anneals to a first target strand of a double-stranded target DNA via a first spacer of the first PEgRNA. In some embodiments, the first PEgRNA forms a complex with a first prime editor and directs the first prime editor to bind to the double-stranded target DNA at a position corresponding to the first interrogation target sequence. In some embodiments, a second PEgRNA anneals to a second interrogation target sequence on a second target strand of the double-stranded target DNA via a second spacer of the second PEgRNA. In some embodiments, the second PEgRNA forms a complex with a second prime editor and directs the second prime editor to bind to the double-stranded target DNA at a position corresponding to the second interrogation target sequence. In some embodiments, the first prime editor and the second prime editor are the same. In some embodiments, the first prime editor and the second prime editor are different.
[0261] In some embodiments, the first search target sequence recognized by the spacer of the first PEGRNA and the second search target sequence recognized by the spacer of the second PEGRNA have a region of complementarity to each other. In some embodiments, the region of complementarity is 2 to 20 nucleotides in length. In some embodiments, the region of complementarity is 5 to 15 nucleotides in length.
[0262] In some embodiments, the first search target sequence recognized by the spacer of the first PEgRNA and the second search target sequence recognized by the spacer of the second PEgRNA do not have a region of complementarity to each other. In some embodiments, the positions of the first and second search target sequences relative to each other can be determined by their positions in the double-stranded target DNA before editing. In some embodiments, the positions of the first and second search target sequences relative to each other can be determined by their positions in the reference double-stranded target DNA.
[0263] In some embodiments, the first search target sequence is upstream of the second search target sequence. In some embodiments, the first search target sequence is downstream of the second search target sequence. In some embodiments, the 5' end of the first search target sequence is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47 of the 5' end of the second search target sequence. , 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 base pairs upstream. In some embodiments, the 5' end of the first interrogation target sequence is 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 base pairs upstream of the 5' end of the second interrogation target sequence. In some embodiments, the 5' end of the first interrogation target sequence is 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, or more base pairs upstream of the 5' end of the second interrogation target sequence.
[0264] In some embodiments, the 3' end of the first search target sequence is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, , 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 base pairs downstream. In some embodiments, the 3' end of the first interrogation target sequence is 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 base pairs downstream of the 3' end of the second interrogation target sequence. In some embodiments, the 3' end of the first interrogation target sequence is 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, or more base pairs downstream of the 3' end of the second interrogation target sequence.
[0265] In some embodiments, the bound first prime editor generates a first nick on the first PAM strand of the double-stranded target DNA. In some embodiments, the first PEgRNA comprises a first primer binding site (PBS), also referred to herein as a "primer binding site sequence," that is complementary to the sequence of the first PAM strand of the double-stranded target DNA immediately upstream of the first nick site and is capable of annealing to the sequence of the first strand at the free 3' end formed at the first nick site. In some embodiments, the first PEgRNA comprises a first primer binding site (PBS) that anneals to the free 3' end formed at the first nick site, and the first prime editor initiates DNA synthesis from the nick site using the free 3' end as a primer. In some embodiments, the first prime editor generates a first newly synthesized single-stranded DNA encoded by the first editing template of the first PEgRNA.
[0266] In some embodiments, the bound second prime editor generates a second nick on the second PAM strand of the double-stranded target DNA. In some embodiments, the double-stranded target DNA, e.g., the target gene, comprises a double-stranded DNA sequence between the first nick generated by the first prime editor on the second target strand (also referred to as the first PAM strand) and the second nick generated by the second prime editor on the first target strand (also referred to as the second PAM strand), which may be referred to as an inter-nick duplex (IND). In some embodiments, the two strands of the IND are fully complementary to each other. In some embodiments, the two strands of the IND are partially complementary to each other. In some embodiments, the IND is then excised from the double-stranded target DNA, e.g., the target gene.
[0267] In some embodiments, the IND is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more base pairs in length. In some embodiments, the IND is up to 5, up to 10, up to 15, up to 20, up to 25, up to 30, up to 40, or up to 50 base pairs in length. In some embodiments, the IND is 1 to 3000, 1 to 2500, 1 to 2000, 1 to 1500, 1 to 1000, 1 to 900, 1 to 800, 1 to 700, 1 to 600, 1 to 500, 1 to 400, 1 to 300, 1 to 200, 1 to 100, 1 to 50, 1 to 40, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 base pairs in length. In some embodiments, the IND is 500 to 3000, 500 to 2500, 500 to 2000, 500 to 1500, 500 to 1000, 500 to 900, 500 to 800, 500 to 700, or 500 to 600 base pairs in length. In some embodiments, the IND is 30 to 300, 30 to 250, 30 to 200, 30 to 150, 30 to 100, 30 to 75, 30 to 50, 50 to 200, 50 to 150, 50 to 100, 50 to 75, 75 to 100, 75 to 150, 75 to 200, 75 to 250, or 75 to 300 base pairs in length.In some embodiments, the IND is 1 to 3, 1 to 6, 1 to 9, 1 to 12, 1 to 15, 1 to 18, 1 to 21, 1 to 24, 1 to 27, 1 to 30, 1 to 36, 1 to 45, 1 to 60, 1 to 72, 1 to 90, 3 to 6, 3 to 9, 3 to 12, 3 to 15, 3 to 18, 3 to 21, 3 to 24, 3 to 27, 3 to 30, 3 to 36, 3 to 45, 3 to 60, 3 to 72, 3 to 90, 6 to 9, 6-12, 6-15, 6-18, 6-21, 6-24, 6-27, 6-30, 6-36, 6-45, 6-60, 6-72, 6-90, 9-12, 9-15, 9-18, 9-21, 9-24, 9-27, 9-30, 9-36, 9-45, 9-60, 9-72, 9-90, 12-15, 12-18, 12-21, 12-24, 12-27, 12-30, 12-36, 12- 45, 12-60, 12-72, 12-90, 15-18, 15-21, 15-24, 15-27, 15-30, 15-36, 15-45, 15-60, 15-72, 15-90, 18-21, 18-24, 18-27, 18-30, 18-36, 18-45, 18-60, 18-72, 18-90, 21-24, 21-27, 21-30, 21-36, 21-45, 21-60, 21-72, 21-90, 24-27, 24-30, 24-36, 24-45, 24-60, 24-72, 24-90, 27-30, 27-36, 27-45, 27-60, 27-72, 27-90, 30-36, 30-45, 30-60, 30-72, 30-90, 45-60, 45-72, 60-72, 60-90, or 72-90 base pairs.In some embodiments, the IND has a length of 1-3000, 1-2500, 1-2000, 1-1500, 1-1000, 1-900, 1-800, 1-700, 1-600, 1-500, 1-400, 1-300, 1-200, 1-100, 1-50, 1-40, 1-30, 1-25, 1-20, 1-15, 1-10, 1-5, 500-3000, 500-2500, 500-2 000, 500-1500, 500-1000, 500-900, 500-800, 500-700, 500-600, 30-300, 30-250, 30-200, 30-150, 30-100, 30-75, 30-50, 50-200, 50-150, 50-100, 50-75, 75-100, 75-150, 75-200, 75-250, or 75-300 base pairs.In some embodiments, the IND is 1 to 3, 1 to 6, 1 to 9, 1 to 12, 1 to 15, 1 to 18, 1 to 21, 1 to 24, 1 to 27, 1 to 30, 1 to 36, 1 to 45, 1 to 60, 1 to 72, 1 to 90, 3 to 6, 3 to 9, 3 to 12, 3 to 15, 3 to 18, 3 to 21, 3 to 24, 3 to 27, 3 to 30, 3 to 36, 3 to 45, 3 to 60, 3 to 72, 3 to 90, 6 to 9, 6-12, 6-15, 6-18, 6-21, 6-24, 6-27, 6-30, 6-36, 6-45, 6-60, 6-72, 6-90, 9-12, 9-15, 9-18, 9-21, 9-24, 9-27, 9-30, 9-36, 9-45, 9-60, 9-72, 9-90, 12-15, 12-18, 12-21, 12-24, 12-27, 12-30, 12-36, 12- 45, 12-60, 12-72, 12-90, 15-18, 15-21, 15-24, 15-27, 15-30, 15-36, 15-45, 15-60, 15-72, 15-90, 18-21, 18-24, 18-27, 18-30, 18-36, 18-45, 18-60, 18-72, 18-90, 21-24, 21-27, 21-30, 21-36, 21-45, 21-60, 21-72, 21-90, 24-27, 24-30, 24-36, 24-45, 24-60, 24-72, 24-90, 27-30, 27-36, 27-45, 27-60, 27-72, 27-90, 30-36, 30-45, 30-60, 30-72, 30-90, 45-60, 45-72, 60-72, 60-90, or 72-90 base pairs.
[0268] In some embodiments, the double-stranded target DNA is a double-stranded target gene or a portion of a double-stranded target gene, and the IND comprises a portion of the coding sequence of the target gene. In some embodiments, the IND comprises a portion of the non-coding sequence of the target gene. In some embodiments, the IND comprises a portion of an exon. In some embodiments, the IND comprises an entire exon. In some embodiments, the IND comprises a portion of an intron. In some embodiments, the IND comprises an entire intron. In some embodiments, the IND comprises the 3' UTR sequence of the target gene. In some embodiments, the IND comprises the 5' UTR sequence of the target gene. In some embodiments, the IND comprises all or part of the ORF of the target gene. In some embodiments, the IND comprises both coding and non-coding sequence of the target gene. In some embodiments, the IND comprises both intron and exon sequences of the target gene. For example, in some embodiments, the IND comprises the sequence of an exon flanked by intron sequences at the 5' end, 3' end, or both ends. In some embodiments, the IND comprises one or more exons and intervening introns. In some embodiments, the IND comprises two or more exons and intervening introns. In some embodiments, the IND comprises the entire coding region of the target gene, a regulatory sequence of the target gene, or the entire target gene including its exons, introns, and regulatory sequences. In some embodiments, the double-stranded DNA comprises a gene or a portion of a gene, and the IND comprises one or more mutations compared to a wild-type reference sequence of the same gene. In some embodiments, the one or more mutations are associated with a disease.
[0269] In some embodiments, the first PEgRNA comprises a first primer binding site (PBS) complementary to the free 3' end of the second strand of the double-stranded target DNA formed at the first nick site. In some embodiments, the first PBS anneals to the free 3' end formed at the first nick site, and the first prime editor initiates DNA synthesis from the first nick site using the free 3' end of the first nick site as a primer. In some embodiments, the first prime editor synthesizes a first novel single-stranded DNA encoded by the first editing template of the first PEgRNA. In some embodiments, the second PEgRNA comprises a second PBS complementary to the free 3' end of the first strand of the double-stranded target DNA formed at the second nick site. In some embodiments, the second PBS anneals to the free 3' end formed at the second nick site, and the second prime editor initiates DNA synthesis from the nick site using the free 3' end of the second nick site as a primer. In some embodiments, the second prime editor synthesizes a second newly synthesized single-stranded DNA encoded by a second editing template of the second PEgRNA.
[0270] In some embodiments, DNA repair incorporates a first newly synthesized single-stranded DNA sequence encoded by a first editing template and / or a second newly synthesized single-stranded DNA sequence encoded by a second editing template into a double-stranded target DNA, e.g., a target gene, thereby incorporating one or more intended nucleotide edits into the double-stranded target DNA, e.g., a target gene.
[0271] As used herein, "nucleotide editing" or "intended nucleotide editing" refers to a specific edit of double-stranded target DNA. Nucleotide editing or intended nucleotide editing refers to (i) the deletion of one or more consecutive nucleotides at a specific position, (ii) the insertion of one or more consecutive nucleotides at a specific position, (iii) the substitution of one or more consecutive nucleotides, or (iv) a combination of consecutive nucleotide substitution, insertion, and / or deletion of two or more consecutive nucleotides at a specific position that are incorporated into the sequence of double-stranded target DNA, or other changes at a specific position. Intended nucleotide editing can refer to editing on an editing template (e.g., a first editing template or a second editing template) relative to the sequence of the double-stranded target gene, or can refer to editing encoded by an editing template in newly synthesized single-stranded DNA incorporated into double-stranded target DNA, e.g., the TRAC gene, relative to the endogenous sequence of the double-stranded target DNA, e.g., the TRAC gene. In some embodiments, intended nucleotide editing can also refer to editing that results from the incorporation of newly synthesized DNA encoded by an editing template, or the incorporation of two newly synthesized single-stranded DNAs encoded by a first PEgRNA and a second PEgRNA, respectively, in dual-prime editing.
[0272] In some embodiments, the sequence of the first newly synthesized single-stranded DNA and / or the sequence of the second newly synthesized single-stranded DNA are incorporated into a double-stranded target DNA, e.g., a target gene. In some embodiments, the first and / or second newly synthesized single-stranded DNA comprise one or more intended nucleotide edits to be incorporated into the double-stranded target DNA, e.g., a target gene, compared to the endogenous sequence of the double-stranded target DNA, e.g., a target gene. In some embodiments, the sequence of the first newly synthesized single-stranded DNA encoded by the first editing template is incorporated into the double-stranded target DNA, e.g., a target gene, thereby incorporating one or more intended nucleotide edits into the double-stranded target DNA, e.g., a target gene. In some embodiments, the sequence of the second newly synthesized single-stranded DNA encoded by the second editing template is incorporated into the double-stranded target DNA, e.g., a target gene, thereby incorporating one or more intended nucleotide edits into the double-stranded target DNA, e.g., a target gene. In some embodiments, a sequence of a first newly synthesized single-stranded DNA encoded by a first editing template and a sequence of a second newly synthesized single-stranded DNA encoded by a second editing template are incorporated into a double-stranded target DNA, e.g., a target gene, thereby incorporating one or more intended nucleotide edits into the double-stranded target DNA, e.g., a target gene.
[0273] In some embodiments, the intended nucleotide edits include insertions, deletions, nucleotide substitutions, inversions, or any combination thereof, compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edits include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotide substitutions compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edits include up to 5, up to 10, up to 15, up to 20, up to 25, up to 30, up to 40, or up to 50 nucleotide substitutions compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit comprises 1 to 50, 1 to 40, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 nucleotide substitutions relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit comprises 3 to 50, 3 to 40, 3 to 30, 3 to 25, 3 to 20, 3 to 15, 3 to 10, or 3 to 5 nucleotide substitutions relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit comprises 5 to 50, 5 to 40, 5 to 30, 5 to 25, 5 to 20, 5 to 15, or 5 to 10 nucleotide substitutions relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene.
[0274] In some embodiments, the intended nucleotide edit comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotide insertions relative to the double-stranded target DNA, e.g., the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises up to 5, up to 10, up to 15, up to 20, up to 25, up to 30, up to 40, or up to 50 nucleotide insertions relative to the double-stranded target DNA, e.g., the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises a single nucleotide insertion at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more sites in the double-stranded target DNA, e.g., the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises a nucleotide insertion of two or more nucleotides at each site of the double-stranded target DNA compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. As used herein, "site" refers to a specific position in the sequence of the target DNA, e.g., the target gene. In some embodiments, a specific position in the sequence of the double-stranded target DNA, e.g., the target gene, may be referred to by a specific position in a reference sequence, e.g., a wild-type gene sequence. In some embodiments, a nucleotide insertion at position x refers to the insertion of one or more nucleotides between position x and position x+1, as indicated by numbering in the reference sequence. In some embodiments, a nucleotide deletion at position x refers to the deletion of a specific nucleotide at position x, as indicated by numbering in the reference sequence. In some embodiments, a nucleotide deletion at positions x through x+n refers to the deletion of a specific nucleotide starting from nucleotide x through nucleotide x+n, inclusive, as indicated by numbering in the reference sequence. In some embodiments, a nucleotide inversion at position x to x+n refers to an inversion of a particular nucleotide starting from nucleotide x through nucleotide x+n, inclusive, as indicated by numbering in the reference sequence.
[0275] In some embodiments, the intended nucleotide edit comprises a nucleotide insertion of two or more nucleotides at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more sites in the double-stranded target DNA relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit comprises an insertion of 1 to 3000, 1 to 2500, 1 to 2000, 1 to 1500, 1 to 1000, 1 to 900, 1 to 800, 1 to 700, 1 to 600, 1 to 500, 1 to 400, 1 to 300, 1 to 200, 1 to 100, 1 to 50, 1 to 40, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 nucleotides relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit comprises an insertion of 500 to 3000, 500 to 2500, 500 to 2000, 500 to 1500, 500 to 1000, 500 to 900, 500 to 800, 500 to 700, or 500 to 600 nucleotides relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit comprises an insertion of 30 to 300, 30 to 250, 30 to 200, 30 to 150, 30 to 100, 30 to 75, 30 to 50, 50 to 200, 50 to 150, 50 to 100, 50 to 75, 75 to 100, 75 to 150, 75 to 200, 75 to 250, or 75 to 300 nucleotides relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene.In some embodiments, the intended nucleotide edit is at least 1 to 3000, 1 to 2500, 1 to 2000, 1 to 1500, 1 to 1000, 1 to 900, 1 to 800, 1 to 700, 1 to 600, 1 to 500, 1 to 400, 1 to 300, 1 to 200, 1 to 100, 1 to 50, 1 to 40, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 5, 500 to 10 ...50, 1 to 500, 1 to 500, 1 to 600, 1 to 700, 1 to 800, 1 to 900, 1 to 1000, 1 to 1500, 1 to 1500, 1 to 2000, 1 to 2000, 1 to 3000, 1 to 2000, 1 to 1000, 1 to 500, 1 to 400, 1 to 3000, 1 to 2000, 1 to 1000, 1 to 500, 1 to 400, 1 to 3000, 1 to 2500, 1 to 2000, 1 to 1500, 1 to 1000, 1 to 500, 1 to and nucleotide insertions of 3000, 500-2500, 500-2000, 500-1500, 500-1000, 500-900, 500-800, 500-700, 500-600, 30-300, 30-250, 30-200, 30-150, 30-100, 30-75, 30-50, 50-200, 50-150, 50-100, 50-75, 75-100, 75-150, 75-200, 75-250, or 75-300 nucleotides.
[0276] In some embodiments, the intended nucleotide edit comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotide deletions relative to the double-stranded target DNA, e.g., the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises up to 5, up to 10, up to 15, up to 20, up to 25, up to 30, up to 40, or up to 50 nucleotide deletions relative to the double-stranded target DNA, e.g., the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises a single nucleotide deletion at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more sites in the double-stranded target DNA, e.g., the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises a nucleotide deletion of two or more nucleotides at each site of the double-stranded target DNA, e.g., compared to the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises a nucleotide deletion of two or more nucleotides at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more sites of the double-stranded target DNA, e.g., compared to the endogenous sequence of the target gene. In some embodiments, the intended nucleotide edit comprises a deletion of 1 to 3000, 1 to 2500, 1 to 2000, 1 to 1500, 1 to 1000, 1 to 900, 1 to 800, 1 to 700, 1 to 600, 1 to 500, 1 to 400, 1 to 300, 1 to 200, 1 to 100, 1 to 50, 1 to 40, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 nucleotides relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit comprises a deletion of 500 to 3000, 500 to 2500, 500 to 2000, 500 to 1500, 500 to 1000, 500 to 900, 500 to 800, 500 to 700, or 500 to 600 nucleotides relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene.In some embodiments, the intended nucleotide edit comprises a deletion of 30 to 300, 30 to 250, 30 to 200, 30 to 150, 30 to 100, 30 to 75, 30 to 50, 50 to 200, 50 to 150, 50 to 100, 50 to 75, 75 to 100, 75 to 150, 75 to 200, 75 to 250, or 75 to 300 nucleotides relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene.
[0277] In some embodiments, the intended nucleotide edit is 1-3, 1-6, 1-9, 1-12, 1-15, 1-18, 1-21, 1-24, 1-27, 1-30, 1-36, 1-45, 1-60, 1-72, 1-90, 3-6, 3-9, 3-12, 3-15, 3-18, 3-21, 3-24, 3-27, 3-30, 3-3 6, 3-45, 3-60, 3-72, 3-90, 6-9, 6-12, 6-15, 6-18, 6-21, 6-24, 6-27, 6-30, 6-36, 6-45, 6-60, 6-72, 6-90, 9-12, 9-15, 9-18, 9-21, 9-24, 9-27, 9-30, 9-36, 9-45, 9-60, 9-72, 9-90, 12-15, 12-18, 12-21, 12-24, 12-27, 12-30, 12-36, 12-45, 12-60, 12-72, 12-90, 15-18, 15-21, 15-24, 15-27, 15-30, 15-36, 15-45, 15-60, 15-72, 15-90, 18-21, 18-24, 18-27, 18-30, 18-36, 18-45, 18-60, 18-72, 18-90, 21-24, 21-27, 21-30, 21-36, 2 and nucleotide deletions of 1-45, 21-60, 21-72, 21-90, 24-27, 24-30, 24-36, 24-45, 24-60, 24-72, 24-90, 27-30, 27-36, 27-45, 27-60, 27-72, 27-90, 30-36, 30-45, 30-60, 30-72, 30-90, 45-60, 45-72, 60-72, 60-90, or 72-90.In some embodiments, the intended nucleotide edit is at least 1 to 3000, 1 to 2500, 1 to 2000, 1 to 1500, 1 to 1000, 1 to 900, 1 to 800, 1 to 700, 1 to 600, 1 to 500, 1 to 400, 1 to 300, 1 to 200, 1 to 100, 1 to 50, 1 to 40, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 5, 500 to 10 ...50, 1 to 500, 1 to 500, 1 to 600, 1 to 700, 1 to 800, 1 to 900, 1 to 1000, 1 to 1500, 1 to 1500, 1 to 2000, 1 to 2000, 1 to 3000, 1 to 2000, 1 to 1000, 1 to 500, 1 to 400, 1 to 3000, 1 to 2000, 1 to 1000, 1 to 500, 1 to 400, 1 to 3000, 1 to 2500, 1 to 2000, 1 to 1500, 1 to 1000, 1 to 500, 1 to and nucleotide deletions of 3000, 500-2500, 500-2000, 500-1500, 500-1000, 500-900, 500-800, 500-700, 500-600, 30-300, 30-250, 30-200, 30-150, 30-100, 30-75, 30-50, 50-200, 50-150, 50-100, 50-75, 75-100, 75-150, 75-200, 75-250, or 75-300 nucleotides.In some embodiments, the intended nucleotide edit is at each site of the double-stranded target DNA, e.g., 1-3, 1-6, 1-9, 1-12, 1-15, 1-18, 1-21, 1-24, 1-27, 1-30, 1-36, 1-45, 1-60, 1-72, 1-90, 3-6, 3-9, 3-12, 3-15, 3-18, 3-21, 3-24, 3-27, , 3-30, 3-36, 3-45, 3-60, 3-72, 3-90, 6-9, 6-12, 6-15, 6-18, 6-21, 6-24, 6-27, 6-30, 6-36, 6-45, 6-60, 6-72, 6-90, 9-12, 9-15, 9-18, 9-21, 9-24, 9-27, 9-30, 9-36, 9-45, 9-60, 9-72, 9-90, 12-15, 12-18, 12-21, 12-24, 12 ~27, 12~30, 12~36, 12~45, 12~60, 12~72, 12~90, 15~18, 15~21, 15~24, 15~27, 15~30, 15~36, 15~45, 15~60, 15~72, 15~90, 18~21, 18~24, 18~27, 18~30, 18~36, 18~45, 18~60, 18~72, 18~90, 21~24, 21~27, 21~30, 21~36, 21 and nucleotide deletions of ~45, 21-60, 21-72, 21-90, 24-27, 24-30, 24-36, 24-45, 24-60, 24-72, 24-90, 27-30, 27-36, 27-45, 27-60, 27-72, 27-90, 30-36, 30-45, 30-60, 30-72, 30-90, 45-60, 45-72, 60-72, 60-90, or 72-90 nucleotides.
[0278] In some embodiments, the intended nucleotide edit, e.g., nucleotide substitution, insertion, or deletion, is in consecutive or adjacent nucleotides in the double-stranded target DNA sequence compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edit, e.g., nucleotide substitution, insertion, or deletion, is in non-consecutive or non-adjacent nucleotides in the double-stranded target DNA sequence compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene.
[0279] In some embodiments, the intended nucleotide edit comprises an inversion relative to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, a segment of 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 200, 250, 300, or more nucleotides of the endogenous sequence of the double-stranded target DNA is inverted. In some embodiments, a segment of 1 to 50, 1 to 40, 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 nucleotides of the endogenous sequence of the double-stranded target DNA is inverted. In some embodiments, a segment of 3 to 50, 3 to 40, 3 to 30, 3 to 25, 3 to 20, 3 to 15, 3 to 10, 3 to 5, 5 to 50, 5 to 40, 5 to 30, 5 to 25, 5 to 20, 5 to 15, or 5 to 10 nucleotides of the endogenous sequence of the double-stranded target DNA is inverted.
[0280] In some embodiments, the intended nucleotide edits include two or more nucleotide edits in the double-stranded target DNA sequence compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edits include a combination of one or more nucleotide substitutions, one or more nucleotide insertions, one or more nucleotide deletions, and one or more nucleotide inversions compared to the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the intended nucleotide edits include one or more nucleotide substitutions and one or more nucleotide insertions. In some embodiments, the intended nucleotide edits include one or more nucleotide substitutions and one or more nucleotide deletions. In some embodiments, the intended nucleotide edits include one or more nucleotide substitutions and one or more nucleotide inversions. In some embodiments, the intended nucleotide edits include one or more nucleotide insertions and one or more nucleotide deletions. In some embodiments, the intended nucleotide edits include one or more nucleotide insertions and one or more nucleotide inversions. In some embodiments, the intended nucleotide edits include one or more nucleotide deletions and one or more nucleotide inversions. In some embodiments, the intended nucleotide edits include one or more nucleotide substitutions, one or more nucleotide insertions, and one or more nucleotide deletions. In some embodiments, the intended nucleotide edits include one or more nucleotide substitutions, one or more nucleotide insertions, and one or more nucleotide inversions. In some embodiments, the intended nucleotide edits include one or more nucleotide substitutions, one or more nucleotide deletions, and one or more nucleotide inversions. In some embodiments, the intended nucleotide edits include one or more nucleotide insertions, one or more nucleotide deletions, and one or more nucleotide inversions. In some embodiments, the intended nucleotide edits include one or more nucleotide substitutions, one or more nucleotide insertions, one or more nucleotide deletions, and one or more nucleotide inversions.
[0281] In some embodiments, the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA have regions of complementarity to each other. In some embodiments, the first newly synthesized single-stranded DNA has a region of complementarity to a double-stranded target DNA, e.g., an endogenous sequence of a target gene, adjacent to or near the nick site. In some embodiments, the first newly synthesized single-stranded DNA has a region of complementarity to a double-stranded target DNA, e.g., an endogenous sequence of a target gene, on the first strand adjacent to the second nick site. In some embodiments, the first newly synthesized single-stranded DNA has a region of complementarity to a double-stranded target DNA, e.g., an endogenous sequence of a target gene, on the first strand adjacent to and downstream from the second nick site. In some embodiments, the first newly synthesized single-stranded DNA has a region of identity to an endogenous sequence of a double-stranded target on the second strand adjacent to and downstream from the first nick site. In some embodiments, the first newly synthesized single-stranded DNA has a region of identity to the endogenous sequence of the double-stranded target on the second strand adjacent to and upstream of the second nick site.
[0282] In some embodiments, the second newly synthesized single-stranded DNA has a region of complementarity to the double-stranded target DNA, e.g., the endogenous sequence of the target gene, adjacent to or near the nick site. In some embodiments, the second newly synthesized single-stranded DNA has a region of complementarity to the double-stranded target DNA, e.g., the endogenous sequence of the target gene, on the second strand adjacent to the first nick site. In some embodiments, the second newly synthesized single-stranded DNA has a region of complementarity to the double-stranded target DNA, e.g., the endogenous sequence of the target gene, on the second strand adjacent to and downstream of the first nick site. In some embodiments, the second newly synthesized single-stranded DNA has a region of identity to the endogenous sequence of the double-stranded target on the first strand adjacent to and upstream of the second nick site.
[0283] As used herein, a reference to a position in a chromosome or double-stranded polynucleotide, e.g., double-stranded target DNA, includes a position in either of the two strands unless otherwise indicated. For example, the position of a first nick site can be used to refer to the first nick site on the first edited strand and / or the corresponding position on the second edited strand.
[0284] "Upstream" and "downstream" are intended to define the relative positions of at least two regions or sequences within a nucleic acid molecule oriented in the 5' to 3' direction. For example, a first sequence is upstream of a second sequence within a DNA molecule, and the first sequence is located 5' of the second sequence. Thus, the second sequence is downstream, i.e., 3', of the first sequence. In the context of dual-prime editing of double-stranded target DNA, references to upstream or downstream positioning are based on the reference strand, which is the protein-coding strand (also referred to as the sense strand) of double-stranded target DNA, e.g., the HTT gene, in the 5' to 3' direction, regardless of whether the sequence is present in the translated region. In some embodiments, the reference strand is the first edited strand (i.e., the second strand shown in Figure 4A). In embodiments where the two sequences are on different strands of a double-stranded polynucleotide, e.g., the sense and antisense strands of a double-stranded target DNA, the sequence defined as the upstream (or 5') sequence and the sequence defined as the downstream (or 3') sequence are based on the position of the sequence on the sense strand (e.g., the first edited strand illustrated in Figure 4A) relative to the position of the complementary sequence of the sequence on the antisense strand.
[0285] In some embodiments, the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA each have a region of complementarity to a double-stranded target DNA, e.g., an endogenous sequence of a target gene, adjacent to or near the nick site. In some embodiments, the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA each have a region of identity to a double-stranded target DNA, e.g., an endogenous sequence of a target gene, adjacent to or near the nick site. In some embodiments, the first newly synthesized single-stranded DNA has a region of complementarity to a double-stranded target DNA, e.g., an endogenous sequence of a target gene, on the first strand adjacent to and downstream from the second nick site, and the second newly synthesized single-stranded DNA has a region of complementarity to a double-stranded target DNA, e.g., an endogenous sequence of a target gene, on the second strand adjacent to and upstream from the first nick site. In some embodiments, the first newly synthesized single-stranded DNA has a region of identity to an endogenous sequence of the double-stranded target DNA on the second strand adjacent to and downstream of the first nick site and / or an endogenous sequence of the double-stranded target DNA on the second strand adjacent to and upstream of the second nick site, and the second newly synthesized single-stranded DNA has a region of identity to an endogenous sequence of the double-stranded target DNA on the first strand adjacent to and downstream of the first nick site and / or an endogenous sequence of the double-stranded target DNA on the second strand adjacent to and upstream of the second nick site.
[0286] In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template have a region of complementarity to each other. The region of complementarity between the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA may be referred to as an overlapping duplex (OD).
[0287] In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template are complementary or substantially complementary to each other. In some embodiments, the OD is incorporated into a double-stranded target DNA, e.g., a target gene, thereby incorporating one or more intended nucleotide edits encoded by the first and second editing templates into the double-stranded target DNA, e.g., a target gene. In some embodiments, the OD replaces all or part of the IND, thereby incorporating one or more intended nucleotide edits into the double-stranded target DNA, e.g., a target gene. In some embodiments, the IND is excised or degraded, and the OD is incorporated at the location of the IND excision, followed by ligation of nicks on both strands of the double-stranded target DNA, e.g., a target gene, thereby incorporating one or more intended nucleotide edits into the double-stranded target DNA. In some embodiments, the sequence of the OD contains partial identity compared to the sequence of the IND. In some embodiments, the sequence of the OD contains no identity compared to the sequence of the IND. In some embodiments, the sequence of the OD comprises a sequence exogenous to the double-stranded target DNA. In some embodiments, the incorporation of the OD does not alter the reading frame of the double-stranded target DNA.
[0288] In some embodiments, the first editing template and the second editing template contain regions of complementarity or substantial complementarity to each other, but do not have complementarity to either strand of a double-stranded target DNA, e.g., a target gene. Thus, in some embodiments, a first newly synthesized single-stranded DNA encoded by the first editing template and a second newly synthesized single-stranded DNA encoded by the second editing template can anneal to each other to form an OD that does not have nucleotide sequence identity with the endogenous sequence of the double-stranded target DNA, e.g., the target gene. In some embodiments, the sequence of the OD comprises a sequence exogenous to the double-stranded target DNA, e.g., the target gene. In some embodiments, the sequence of the OD consists of a sequence exogenous to the double-stranded target DNA, e.g., the target gene. In some embodiments, the IND is excised, an OD is integrated at the location of the IND excision, and then nicks are ligated on both strands of the target DNA, thereby incorporating the sequence of the OD into the double-stranded target DNA.
[0289] In some embodiments, the OD comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more consecutive complementary or substantially complementary base pairs. In some embodiments, the OD is about 5-10, 5-15, 5-20, 5-25, 5-30, 5-35, 5-40, 5-45, 5-50, 5-55, 5-60, 5-65, 5-70, 5-75, 5-80, 5-85, 5-90, 5-95, 5-100, 5-110, 5-120, 5-130, 5-140, 5-150, 15-20, 15-25, 15-30, 15-35, 15-40, 15-45, 15-50, 15-55, 15-60, 15-65, 15-70, 15-75, 15-80, 15-85, 15-90, 15-95, 15-100, 15-110, 15-120, 15-130, 15-140, 1 ... ~80, 15~85, 15~90, 15~95, 15~100, 15~110, 15~120, 15~130, 15~140, 15~150, 25~30, 25~35, 25~40, 25~45, 25~50, 25~55, 25~60, 25~65, 25~70, 25~75, 25~80, 25~85, 25~90, 25~95, 25~100, 25~110, 25~120, 25~130, 25~140, 25~150, 35~40, 35~45, 35~50, 35~55, 35~ 60, 35-65, 35-70, 35-75, 35-80, 35-85, 35-90, 35-95, 35-100, 35-110, 35-120, 35-130, 35-140, 35-150, 45-50, 45-55, 45-60, 45-65, 45-70, 45-75, 45-80, 45-85, 45-90, 45-95, 45-100, 45-110, 45-120, 45-130, 45-140, 45-150, 55-60, 55-65, 55-70, 55-75, 55-8 0, 55~85, 55~90, 55~95, 55~100, 55~110, 55~120, 55~130, 55~140, 55~150, 65~70, 65~75, 65~80, 65~85, 65~90, 65~95, 65~100, 65~110, 65~120, 65~130, 65~140, 65~150, 75~80, 75~85, 75~90, 75~95, 75~100, 75~110, 75~120, 75~130, 75~140, 75~150, 85~90, 85~95,and 85-100, 85-110, 85-120, 85-130, 85-140, 85-150, 95-100, 95-110, 95-120, 95-130, 95-140, 95-150, 105-110, 105-120, 105-130, 105-140, 105-150, 115-120, 115-130, 115-140, 115-150, 125-130, 125-140, 125-150, 135-140, 135-150, or 145-150 consecutive complementary or substantially complementary base pairs. In some embodiments, the OD comprises 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 contiguous complementary or substantially complementary base pairs. In some embodiments, the OD comprises 30, 35, 40, 50, 60, 70, 80, 90, or 100 contiguous complementary or substantially complementary base pairs. In some embodiments, the OD comprises no more than 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 contiguous complementary or substantially complementary base pairs. In some embodiments, the OD comprises a sufficient number of contiguous complementary base pairs to form a duplex sufficiently stable for displacement of the IND. In some embodiments, the OD comprises at least 10 contiguous complementary or substantially complementary base pairs. In some embodiments, the OD comprises at least 15 consecutive complementary or substantially complementary base pairs. In some embodiments, the OD comprises about 20 consecutive complementary or substantially complementary base pairs.
[0290] In some embodiments, the OD replaces the IND of the target DNA, wherein the double-stranded target DNA is the entire target gene or a portion of the target gene. In some embodiments, the OD replaces a portion of an exon or an entire exon, a portion of an intron or an entire intron, one or more exons and intervening introns, all of the coding region of the target gene, a regulatory sequence of the target gene, or the entire target gene including its exons, introns, and regulatory sequences. In some embodiments, the OD contains a region of identity to an endogenous sequence of the double-stranded target DNA. In some embodiments, the OD does not have sequence identity to an endogenous sequence of the double-stranded target DNA. In some embodiments, the OD is exogenous to the double-stranded target DNA, e.g., the target gene.
[0291] In some embodiments, the OD has a biological function or encodes a polypeptide or portion thereof having a biological function. In some embodiments, the OD comprises an expression cassette. In some embodiments, the OD comprises a nucleotide sequence encoding an expression tag, such as an affinity tag, a His tag, a V5 tag, or a FLAG tag. In some embodiments, the OD comprises a nucleotide sequence encoding a His tag. In some embodiments, the OD comprises a nucleotide sequence encoding a FLAG tag. In some embodiments, the OD comprises a nucleotide sequence encoding an attB or attP sequence. In some embodiments, the OD comprises a nucleotide sequence encoding a reporter protein, such as green fluorescent protein, blue fluorescent protein, cyan fluorescent protein, yellow fluorescent protein, an autofluorescent protein, or luciferase. In some embodiments, the OD comprises a recognition site for an enzyme, such as a recombinase recognition sequence. In some embodiments, the OD comprises a nucleotide sequence encoding a selectable marker, such as an antibiotic resistance marker. In some embodiments, the OD comprises a regulatory sequence, such as a promoter, an enhancer, or an insulator. In some embodiments, the OD comprises a traceable sequence, such as a barcode. In some embodiments, replacement of IND with OD reduces or eliminates the expression or function of the target gene (e.g., the TRAC gene). In some embodiments, replacement of IND with OD results in disruption of the target DNA (e.g., the TRAC gene) and insertion of one or more recombinase recognition sequences encoded by the OD. In some embodiments, the target gene is a disease-associated gene. In some embodiments, the target gene is a monogenic disease-associated gene. In some embodiments, the target gene is a polygenic disease-associated gene. In some embodiments, the target gene is a disease-associated gene containing one or more disease-causing mutations, and replacement of IND with OD corrects the mutations, thereby restoring or partially restoring function of the target gene. In some embodiments, the disease-associated gene containing one or more disease-causing mutations is in a human subject in need of treatment.In some embodiments, the target gene is a mutated gene that causes a disease or disorder in a human subject, and replacement of IND with OD corrects the mutated gene, thereby restoring or partially restoring function of the target gene. In some embodiments, the target gene is a disease-associated gene containing one or more disease-causing mutations, and replacement of IND with OD modifies the target gene, thereby restoring or partially restoring function of the target gene. In some embodiments, the disease-associated gene containing one or more disease-causing mutations is in a human subject in need of treatment. In some embodiments, the target gene is a mutated gene that causes a disease or disorder in a human subject, and replacement of IND with OD modifies the mutated gene, thereby restoring or partially restoring function of the target gene. In some embodiments, the target gene is a wild-type gene, e.g., wild-type TRAC. In some embodiments, replacement of IND with OD modifies the target gene, reducing expression or function of the target gene, mRNA, or protein encoded by the target gene. In some embodiments, the target gene is a TRAC gene. In some embodiments, replacement of IND with OD modifies the TRAC gene, reducing the function of the TRAC gene, TRAC mRNA, and / or the TRAC protein encoded by the TRAC gene. In some embodiments, replacement of IND with OD results in disruption of the target gene (e.g., the TRAC gene) and insertion of one or more exogenous sequences, e.g., one or more recombinase recognition sequences, into the gene.
[0292] In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template contain regions of complementarity to each other and can anneal to each other to form an OD. In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template further contains a region that is not complementary to the second newly synthesized single-stranded DNA encoded by the second editing template (see exemplary schematic in Figure 4B). In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template further contains a region that is not complementary to the first newly synthesized single-stranded DNA encoded by the first editing template. Thus, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template can anneal to each other via partially complementary sequences to form an OD linked to the 5' overhang and / or 3' overhang. In some embodiments, the IND is removed, and the OD, along with the 5' overhang and / or 3' overhang, is incorporated into the double-stranded target DNA, e.g., the target gene, at the position of the IND excision. DNA repair fills and ligates the gap corresponding to the position of the 5' overhang and / or 3' overhang, thereby incorporating one or more intended nucleotide edits into the double-stranded target DNA, e.g., the target gene.
[0293] Thus, in some embodiments, IND is replaced by the sequence (A+C), (B+C), or (A+B+C), where A is the region of a first newly synthesized single-stranded DNA that is not complementary to the second newly synthesized single-stranded DNA and its complementary strand, where B is the region of a second newly synthesized single-stranded DNA that is not complementary to the first newly synthesized single-stranded DNA and its complementary strand, and where C is OD. The double-stranded sequence of (A+C), (B+C), or (A+B+C) that replaces IND is sometimes referred to as a "displacement duplex (RD)."
[0294] Thus, in some embodiments, the RD comprises an OD. In some embodiments, the first editing template and the second editing template are substantially complementary to one another, as illustrated in FIG. 4A . Thus, in some embodiments, the OD comprises all or substantially all of a first newly synthesized single-stranded DNA encoded by the first editing template and a second newly synthesized single-stranded DNA encoded by the second editing template. In some embodiments, the RD consists of an OD. In some embodiments, the RD comprises the OD, a non-complementary region of the first newly synthesized DNA relative to the second newly synthesized DNA and its complement, and / or a non-complementary region of the second newly synthesized DNA relative to the first newly synthesized DNA and its complement, as illustrated in FIG. 4B .
[0295] In some embodiments, the RD comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 or more base pairs. In some embodiments, RD is about 5 to 10, 5 to 15, 5 to 20, 5 to 25, 5 to 30, 5 to 35, 5 to 40, 5 to 45, 5 to 50, 5 to 55, 5 to 60, 5 to 65, 5 to 70, 5 to 75, 5 to 80, 5 to 85, 5 to 90, 5 to 95, 5 to 100, 5 to 110, 5 to 120, 5 to 130, 5 to 140, 5 to 150, 5 to 175, 5 to 200, 5 to 225, 5 to 250, 5 to 275, 5 to 300, 5 to 325, 5 to 350, 5 to 375, 5 to 400, 5 to 42 5, 5-450, 5-475, 5-500, 10-15, 10-20, 10-25, 10-30, 10-35, 10-40, 10-45, 10-50, 10-55, 10-60, 10-65, 10-70, 10-75, 10-80, 10-85, 10-90, 10-95, 10-100, 10-110, 10-120, 10-130, 10-140, 10-150, 10-175, 10-200, 10-225, 10-250, 10-275, 10-300, 10 ~325, 10~350, 10~375, 10~400, 10~425, 10~450, 10~475, 10~500, 15~20, 15~25, 15~30, 15~35, 15~40, 15~45, 15~50, 15~55, 15~60, 15~65, 15~70, 15~75, 15~80, 15~85, 15~90, 15~95, 15~100, 15~110, 15~120, 15~130, 15~140, 15~150, 15~175, 15~200, 1 5~225, 15~250, 15~275, 15~300, 15~325, 15~350, 15~375, 15~400, 15~425, 15~450, 15~475, 15~500, 20~25, 20~30, 20~35, 20~40, 20~45, 20~50, 20~55, 20~60, 20~65, 20~70, 20~75, 20~80, 20~85, 20~90, 20~95, 20~100, 20~110, 20~120, 20~130, 20~140,20~150、20~175、20~200、20~225、20~250、20~275、20~300、20~325、20~350、20~375、20~400、20~425、20~450、20~475、20~500、30~35、30~40、30~45、30~50、30~55、30~60、30~65、30~70、30~75、30~80、30~85、30~90、30~95、30~100、30~110、30~120、30~130、30~140、30~150、30~175、30~200、30~225、30~250、30~275、30~300、30~325、30~350、30~375、30~400、30~425、30~450、30~475、30~500、40~45、40~50、40~55、40~60、40~65、40~70、40~75、40~80、40~85、40~90、40~95、40~100、40~110、40~120、40~130、40~140、40~150、40~175、40~200、40~225、40~250、40~275、40~300、40~325、40~350、40~375、40~400、40~425、40~450、40~475、40~500、50~55、50~60、50~65、50~70、50~75、50~80、50~85、50~90、50~95、50~100、50~110、50~120、50~130、50~140、50~150、50~175、50~200、50~225、50~250、50~275、50~300、50~325、50~350、50~375、50~400、50~425、50~450、50~475、50~500、75~80、75~85、75~90、75~95、75~100、75~110、75~120、75~130、75~140、75~150、75~175、75~200、75~225、75~250、75~275、75~300、75~325、75~350、75~375、75~400、75~425、75~450、75~475、75~500、100~110、100~120、100~130、100~140、100~150、100~175、100~200、100~225、100~250、100~275、100~300、100~325、100~350、100~375、100~400, 100~425, 100~450, 100~475, 100~500, 125~150, 125~175, 125~200, 125~225, 125~250, 125~275, 125~300, 125~325, 125~350, 125~375, 125~400, 125~425, 125~450, 125~475, 125~500, 150~175, 150~200, 150~225, 150~250, 150~275, 150~300, 150~325, 150~350, 150~375, 15 0~400, 150~425, 150~450, 150~475, 150~500, 175~200, 175~225, 175~250, 175~275, 175~300, 175~325, 175~350, 175~375, 175~400, 175~425, 175~450, 175~475, 175~500, 200~250, 200~275, 200~300, 200~325, 200~350, 200~375, 200~400, 200~425, 200~450, 200~475, 200~500, 225~ 250, 225-275, 225-300, 225-325, 225-350, 225-375, 225-400, 225-425, 225-450, 225-475, 225-500, 250-275, 250-300, 275-300, 275-325, 275-350, 275-375, 275-400, 275-425, 275-450, 275-475, 275-500, 300-325, 300-350, 300-375, 300-400, 300-425, 300-450, 300-475, 300-50 and 0, 325-350, 325-375, 325-400, 325-425, 325-450, 325-475, 325-500, 350-375, 350-400, 350-425, 350-450, 350-475, 350-500, 375-400, 375-425, 375-450, 375-475, 375-500, 400-425, 400-450, 400-475, 400-500, 425-450, 425-475, 425-500, 450-475, 450-500, or 475-500 base pairs. In some embodiments, RD is 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28,29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 base pairs. In some embodiments, the RD comprises at least 30, 35, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 base pairs. In some embodiments, the RD comprises no more than 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 base pairs.
[0296] In some embodiments, the RD replaces the IND of the target DNA, where the IND is the entire target gene or a portion of the target gene. In some embodiments, the RD replaces a portion of an exon or an entire exon, a portion of an intron or an entire intron, one or more exons and intervening introns, the entire coding region of the target gene, the regulatory sequence of the target gene, or the entire target gene including its exons, introns, and regulatory sequences, thereby incorporating one or more intended nucleotide edits compared to the endogenous sequence of the double-stranded target DNA, e.g., the TRAC gene. In some embodiments, the RD comprises a region of identity to the endogenous sequence of the double-stranded target DNA. In some embodiments, the RD does not have sequence identity with the endogenous sequence of the double-stranded target DNA. In some embodiments, the RD is exogenous to the double-stranded target DNA, e.g., the target gene. Thus, in some embodiments, the intended nucleotide edit comprises replacing the entire endogenous sequence of the double-stranded target DNA, e.g., the TRAC gene, with the sequence of the RD.
[0297] In some embodiments, the RD or OD can include a recombinase recognition sequence (RSS), e.g., an RSS recognized by a recombinase recognition site corresponding to Bxbl recombinase, Cre recombinase, PaOl recombinase, Si74 recombinase, No67 recombinase, KpO3 recombinase, Nm60 recombinase, BceINTa recombinase, NcytINTd recombinase, SscINTd recombinase, SacINTd recombinase, or any recombinase disclosed herein. In some embodiments, the RD or OD can include one, two, or more recombinase recognition sites corresponding to a recombinase.
[0298] Replacing an IND with an RD or OD containing one or more recombinase sequences using dual-prime editing can result in the insertion of one or more recombinase sequences into a target gene, such as the TRAC gene. Depending on the number and orientation of the RSSs, they can be used as landing sites for recombinase-mediated reactions between the RSSs. For example, a single RSS inserted into a target gene, such as the TRAC gene, can be used to integrate an exogenous DNA donor sequence through recombination between the inserted RSS and a second RSS within the provided exogenous DNA donor. When two recombinase sites are inserted into adjacent DNA regions, they can be used for recombinase-mediated excision or inversion of an intervening sequence, or for recombinase-mediated cassette exchange with exogenous DNA for cargo integration, depending on the orientation of the recombinase sites.
[0299] In some embodiments, the RD has a biological function or encodes a polypeptide having a biological function. In some embodiments, the RD comprises an expression cassette. In some embodiments, the RD comprises a nucleotide sequence encoding an expression tag, e.g., an affinity tag, a His tag, a V5 tag, or a FLAG tag. In some embodiments, the RD comprises a nucleotide sequence encoding a His tag. In some embodiments, the RD comprises a nucleotide sequence encoding a FLAG tag. In some embodiments, the RD comprises a nucleotide sequence encoding an attB or attP sequence. In some embodiments, the RD comprises a nucleotide sequence encoding a reporter protein, e.g., green fluorescent protein, blue fluorescent protein, cyan fluorescent protein, yellow fluorescent protein, an autofluorescent protein, or luciferase. In some embodiments, the RD comprises a recognition site for an enzyme, e.g., a recombinase recognition sequence. In some embodiments, the RD comprises a nucleotide sequence encoding a selectable marker, e.g., an antibiotic resistance marker. In some embodiments, the RD comprises a regulatory sequence, e.g., a promoter, an enhancer, or an insulator. In some embodiments, the RD comprises a traceable sequence, e.g., a barcode. In some embodiments, replacement of IND with the RD restores or partially restores the function of the target gene. In some embodiments, replacement of IND with the RD reduces or eliminates the function or expression of the target gene. In some embodiments, the target gene is the TRAC gene. In some embodiments, replacement of IND with the RD reduces the function of the TRAC gene, TRAC mRNA, and / or TRAC protein. In some embodiments, the target gene is a disease-associated gene. In some embodiments, the target gene is a monogenic disease-associated gene. In some embodiments, the target gene is a polygenic disease-associated gene. In some embodiments, the target gene is a disease-associated gene containing one or more disease-causing mutations, and replacement of IND with the RD corrects the mutations, thereby restoring or partially restoring the function of the target gene.In some embodiments, a disease-associated gene comprising one or more disease-causing mutations is in a human subject in need of treatment. In some embodiments, the target gene is a mutated gene that causes a disease or disorder in the human subject, and replacement of the IND with the RD corrects the mutated gene, thereby restoring or partially restoring function of the target gene. In some embodiments, the target gene is a disease-associated gene comprising one or more disease-causing mutations, and replacement of the IND with the RD modifies the target gene, thereby restoring or partially restoring function of the target gene. In some embodiments, a disease-associated gene comprising one or more disease-causing mutations is in a human subject in need of treatment. In some embodiments, the target gene is a mutated gene that causes a disease or disorder in the human subject, and replacement of the IND with the RD modifies the mutated gene, thereby restoring or partially restoring function of the target gene.
[0300] In some embodiments, the first editing template and the second editing template are partially complementary to each other. As used herein, a first editing template is partially complementary to a second editing template if the first editing template and the second editing template have complementary or substantially complementary regions over a portion of the length of both editing templates. The partially complementary regions within the first editing template and the second editing template can be located anywhere within the first editing template and the second editing template. Thus, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template are partially complementary to each other at any location within the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template. In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to a second newly synthesized single-stranded DNA at or near the 3' end of the first newly synthesized single-stranded DNA. In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to a second newly synthesized single-stranded DNA at or near the 5' end of the first newly synthesized single-stranded DNA. In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to a second newly synthesized single-stranded DNA in the middle of the first newly synthesized single-stranded DNA.
[0301] In some embodiments, the second newly synthesized single-stranded DNA comprises a region of complementarity to the first newly synthesized single-stranded DNA at or near the 3' end of the second newly synthesized single-stranded DNA. In some embodiments, the second newly synthesized single-stranded DNA comprises a region of complementarity to the first newly synthesized single-stranded DNA at or near the 5' end of the second newly synthesized single-stranded DNA. In some embodiments, the second newly synthesized single-stranded DNA comprises a region of complementarity to the first newly synthesized single-stranded DNA in the middle of the second newly synthesized single-stranded DNA.
[0302] In some embodiments, the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA each comprise a region of complementarity to one another at the 3' end of the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA, respectively.
[0303] In some embodiments, the first and second edit templates are the same length. In some embodiments, the first and second edit templates are different lengths.
[0304] In some embodiments, a first editing template includes a region that has complementarity or substantial complementarity to a second editing template (an OD coding region) and further includes a region that does not have complementarity to the second editing template. In some embodiments, a first editing template includes a region that has complementarity or substantial complementarity to a second editing template (an OD coding region), which region is adjacent to one or more regions that do not have complementarity to the second editing template. In some embodiments, the entire first editing template has complementarity or substantial complementarity to a region of a second editing template, where the second editing template includes a region that does not have complementarity to the first editing template.
[0305] In some embodiments, the first editing template comprises a region that has no complementarity to the second editing template, and the region is about 5 to 10, 5 to 15, 5 to 20, 5 to 25, 5 to 30, 5 to 35, 5 to 40, 5 to 45, 5 to 50, 5 to 55, 5 to 60, 5 to 65, 5 to 70, 5 to 75, 5 to 80, 5 to 85, 5 to 90, 5 to 95, 5 to 100, 5 to 110, 5 to 120, 5 to 130, 5 to 140, 5 to 150, 5 to 175, 5 to 200, 5 to 225, 5 to 250, 5 to 275, 5 to 300, 5 to 325, 5 to 350, 5 to 375, 5 to 400, 5 to 5 425, 5-450, 5-475, 5-500, 10-15, 10-20, 10-25, 10-30, 10-35, 10-40, 10-45, 10-50, 10-55, 10-60, 10-65, 10-70, 10-75, 10-80, 10-85, 10-90, 10-95, 10~100, 10~110, 10~120, 10~130, 10~140, 10~150, 10~175, 10~200, 10~225, 10~250, 10~275, 10~300, 10~325, 10~350, 10~375, 10~400, 10~425, 10~450 , 10~475, 10~500, 15~20, 15~25, 15~30, 15~35, 15~40, 15~45, 15~50, 15~55, 15~60, 15~65, 15~70, 15~75, 15~80, 15~85, 15~90, 15~95, 15~100, 15~110 , 15~120, 15~130, 15~140, 15~150, 15~175, 15~200, 15~225, 15~250, 15~275, 15~300, 15~325, 15~350, 15~375, 15~400, 15~425, 15~450, 15~475, 15~50 0, 20-25, 20-30, 20-35, 20-40, 20-45, 20-50, 20-55, 20-60, 20-65, 20-70, 20-75, 20-80, 20-85, 20-90, 20-95, 20-100, 20-110, 20-120, 20-130, 20-140 0, 20~150, 20~175, 20~200, 20~225, 20~250, 20~275, 20~300, 20~325, 20~350, 20~375, 20~400, 20~425, 20~450, 20~475, 20~500, 30~35, 30~40, 30~45,30~50、30~55、30~60、30~65、30~70、30~75、30~80、30~85、30~90、30~95、30~100、30~110、30~120、30~130、30~140、30~150、30~175、30~200、30~225、30~250、30~275、30~300、30~325、30~350、30~375、30~400、30~425、30~450、30~475、30~500、40~45、40~50、40~55、40~60、40~65、40~70、40~75、40~80、40~85、40~90、40~95、40~100、40~110、40~120、40~130、40~140、40~150、40~175、40~200、40~225、40~250、40~275、40~300、40~325、40~350、40~375、40~400、40~425、40~450、40~475、40~500、50~55、50~60、50~65、50~70、50~75、50~80、50~85、50~90、50~95、50~100、50~110、50~120、50~130、50~140、50~150、50~175、50~200、50~225、50~250、50~275、50~300、50~325、50~350、50~375、50~400、50~425、50~450、50~475、50~500、75~80、75~85、75~90、75~95、75~100、75~110、75~120、75~130、75~140、75~150、75~175、75~200、75~225、75~250、75~275、75~300、75~325、75~350、75~375、75~400、75~425、75~450、75~475、75~500、100~110、100~120、100~130、100~140、100~150、100~175、100~200、100~225、100~250、100~275、100~300、100~325、100~350、100~375、100~400、100~425、100~450、100~475、100~500、125~150、125~175、125~200、125~225、125~250、125~275、125~300、125~325、125~350、125~375、125~400, 125~425, 125~450, 125~475, 125~500, 150~175, 150~200, 150~225, 150~250, 150~275, 150~300, 150~325, 150~350, 150~375, 150~400, 150~425, 150~450, 150~475, 150~500, 175~200, 175~225, 175~250, 175~275, 175~300, 175~325, 175~3 50, 175~375, 175~400, 175~425, 175~450, 175~475, 175~500, 200~250, 200~275, 200~300, 200~325, 200~350, 200~375, 200~400, 200~425, 200~450, 200~475, 200~500, 225~250, 225~275, 225~300, 225~325, 225~350, 225~375, 225~400, 225~425, 22 5~450, 225~475, 225~500, 250~275, 250~300, 275~300, 275~325, 275~350, 275~375, 275~400, 275~425, 275~450, 275~475, 275~500, 300~325, 300~350, 300~375, 300~400, 300~425, 300~450, 300~475, 300~500, 325~350, 325~375, 325~400, 325~425 , 325-450, 325-475, 325-500, 350-375, 350-400, 350-425, 350-450, 350-475, 350-500, 375-400, 375-425, 375-450, 375-475, 375-500, 400-425, 400-450, 400-475, 400-500, 425-450, 425-475, 425-500, 450-475, 450-500, or 475-500 nucleotides. In some embodiments, the first editing template comprises a region that has no complementarity to the second editing template, the region having a length of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150,160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 or more nucleotides.
[0306] In some embodiments, the second editing template includes a region that is complementary or substantially complementary to the first editing template and further includes a region that is not complementary to the first editing template. In some embodiments, the second editing template includes a region that is complementary or substantially complementary to the first editing template and is adjacent to one or more regions that are not complementary to the first editing template. The regions in the first editing template and the second editing template may be the same length or different lengths. In some embodiments, the entire second editing template is complementary or substantially complementary to a region of the first editing template, where the first editing template includes a region that is not complementary to the second editing template.
[0307] In some embodiments, the second editing template comprises a region that has no complementarity to the first editing template, and the region is about 5 to 10, 5 to 15, 5 to 20, 5 to 25, 5 to 30, 5 to 35, 5 to 40, 5 to 45, 5 to 50, 5 to 55, 5 to 60, 5 to 65, 5 to 70, 5 to 75, 5 to 80, 5 to 85, 5 to 90, 5 to 95, 5 to 100, 5 to 110, 5 to 120, 5 to 130, 5 to 140, 5 to 150, 5 to 175, 5 to 200, 5 to 225, 5 to 250, 5 to 275, 5 to 300, 5 to 325, 5 to 350, 5 to 375, 5 to 400, 5 to 5 425, 5-450, 5-475, 5-500, 10-15, 10-20, 10-25, 10-30, 10-35, 10-40, 10-45, 10-50, 10-55, 10-60, 10-65, 10-70, 10-75, 10-80, 10-85, 10-90, 10-95, 10~100, 10~110, 10~120, 10~130, 10~140, 10~150, 10~175, 10~200, 10~225, 10~250, 10~275, 10~300, 10~325, 10~350, 10~375, 10~400, 10~425, 10~450 , 10~475, 10~500, 15~20, 15~25, 15~30, 15~35, 15~40, 15~45, 15~50, 15~55, 15~60, 15~65, 15~70, 15~75, 15~80, 15~85, 15~90, 15~95, 15~100, 15~110 , 15~120, 15~130, 15~140, 15~150, 15~175, 15~200, 15~225, 15~250, 15~275, 15~300, 15~325, 15~350, 15~375, 15~400, 15~425, 15~450, 15~475, 15~50 0, 20-25, 20-30, 20-35, 20-40, 20-45, 20-50, 20-55, 20-60, 20-65, 20-70, 20-75, 20-80, 20-85, 20-90, 20-95, 20-100, 20-110, 20-120, 20-130, 20-140 0, 20~150, 20~175, 20~200, 20~225, 20~250, 20~275, 20~300, 20~325, 20~350, 20~375, 20~400, 20~425, 20~450, 20~475, 20~500, 30~35, 30~40, 30~45,30~50、30~55、30~60、30~65、30~70、30~75、30~80、30~85、30~90、30~95、30~100、30~110、30~120、30~130、30~140、30~150、30~175、30~200、30~225、30~250、30~275、30~300、30~325、30~350、30~375、30~400、30~425、30~450、30~475、30~500、40~45、40~50、40~55、40~60、40~65、40~70、40~75、40~80、40~85、40~90、40~95、40~100、40~110、40~120、40~130、40~140、40~150、40~175、40~200、40~225、40~250、40~275、40~300、40~325、40~350、40~375、40~400、40~425、40~450、40~475、40~500、50~55、50~60、50~65、50~70、50~75、50~80、50~85、50~90、50~95、50~100、50~110、50~120、50~130、50~140、50~150、50~175、50~200、50~225、50~250、50~275、50~300、50~325、50~350、50~375、50~400、50~425、50~450、50~475、50~500、75~80、75~85、75~90、75~95、75~100、75~110、75~120、75~130、75~140、75~150、75~175、75~200、75~225、75~250、75~275、75~300、75~325、75~350、75~375、75~400、75~425、75~450、75~475、75~500、100~110、100~120、100~130、100~140、100~150、100~175、100~200、100~225、100~250、100~275、100~300、100~325、100~350、100~375、100~400、100~425、100~450、100~475、100~500、125~150、125~175、125~200、125~225、125~250、125~275、125~300、125~325、125~350、125~375、125~400, 125~425, 125~450, 125~475, 125~500, 150~175, 150~200, 150~225, 150~250, 150~275, 150~300, 150~325, 150~350, 150~375, 150~400, 150~425, 150~450, 150~475, 150~500, 175~200, 175~225, 175~250, 175~275, 175~300, 175~325, 175~3 50, 175~375, 175~400, 175~425, 175~450, 175~475, 175~500, 200~250, 200~275, 200~300, 200~325, 200~350, 200~375, 200~400, 200~425, 200~450, 200~475, 200~500, 225~250, 225~275, 225~300, 225~325, 225~350, 225~375, 225~400, 225~425, 22 5~450, 225~475, 225~500, 250~275, 250~300, 275~300, 275~325, 275~350, 275~375, 275~400, 275~425, 275~450, 275~475, 275~500, 300~325, 300~350, 300~375, 300~400, 300~425, 300~450, 300~475, 300~500, 325~350, 325~375, 325~400, 325~425 , 325-450, 325-475, 325-500, 350-375, 350-400, 350-425, 350-450, 350-475, 350-500, 375-400, 375-425, 375-450, 375-475, 375-500, 400-425, 400-450, 400-475, 400-500, 425-450, 425-475, 425-500, 450-475, 450-500, or 475-500 nucleotides. In some embodiments, the second editing template comprises a region that has no complementarity to the first editing template, the region having a length of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150,160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 or more nucleotides.
[0308] In some embodiments, the RD comprises a region (or a subset) of the sequence of IND. In some embodiments, the RD consists of a region of the sequence of IND. In some embodiments, the RD comprises one or more intended nucleotide edits compared to IND. In some embodiments, the RD comprises a region with substantial sequence identity to the sequence of IND, which region comprises one or more nucleotide edits compared to the sequence of IND. For example, the RD can comprise a region with substantial sequence identity to the sequence of IND, which region comprises one or more nucleotide substitutions, insertions, or deletions. In some embodiments, the RD comprises a region of the sequence of IND and further comprises a region that does not have sequence identity or complementarity to IND. In some embodiments, the RD comprises a region with substantial identity to the sequence of IND that comprises one or more nucleotide edits, and further comprises a region that does not have sequence identity or complementarity to IND. In some embodiments, the region that does not have sequence identity or complementarity to IND has a biological function or encodes a polypeptide or portion thereof that has a biological function. In some embodiments, the RD comprises one or more intended nucleotide edits compared to IND and encodes a polypeptide or portion thereof.
[0309] In some embodiments, the OD comprises a region (or a subset) of the sequence of IND. In some embodiments, the OD consists of a region of the sequence of IND. In some embodiments, the OD comprises one or more intended nucleotide edits compared to IND. In some embodiments, the OD comprises a region with substantial sequence identity to the sequence of IND, which region comprises one or more nucleotide edits compared to the sequence of IND. For example, the OD can comprise a region with substantial sequence identity to the sequence of IND, which region comprises one or more nucleotide substitutions, insertions, or deletions. In some embodiments, the OD comprises a region of the sequence of IND and further comprises a region that does not have sequence identity or complementarity to IND. In some embodiments, the OD comprises a region with substantial identity to the sequence of IND that comprises one or more nucleotide edits, and further comprises a region that does not have sequence identity or complementarity to IND. In some embodiments, the region that does not have sequence identity or complementarity to IND has a biological function or encodes a polypeptide or portion thereof that has a biological function. In some embodiments, the OD comprises one or more intended nucleotide edits compared to IND and encodes a polypeptide or portion thereof.
[0310] In some embodiments, the first editing template comprises a region of identity to a sequence adjacent to the second nick site on the second PAM strand of the double-stranded target DNA, this sequence being outside the IND. In some embodiments, the second editing template comprises a region of identity to a sequence adjacent to the first nick site on the first PAM strand of the double-stranded target DNA, this sequence being outside the IND. Thus, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template comprises a region of complementarity to a sequence adjacent to the second nick site on the second PAM strand of the double-stranded target DNA, this sequence being outside the IND. In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template comprises a region of complementarity to a sequence adjacent to the first nick site on the first PAM strand of the double-stranded target DNA, this sequence being outside the IND.
[0311] In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to a sequence immediately adjacent to the second nick site on the second PAM strand of the double-stranded target DNA, which sequence is outside of the IND. In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template comprises a region of complementarity to a sequence immediately adjacent to the first nick site on the first PAM strand of the double-stranded target DNA, which sequence is outside of the IND (see, e.g., Figure 4F).
[0312] In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to a sequence adjacent to the second nick site on the second PAM strand of the double-stranded target DNA, which sequence is outside the IND and is spaced 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from the second nick site. In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template comprises a region of complementarity to a sequence adjacent to the first nick site on the first PAM strand of the double-stranded target DNA, which sequence is outside the IND and is spaced 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from the first nick site.
[0313] In some embodiments, the first editing template and the second editing template each comprise a region of complementarity or substantial complementarity to each other. In some embodiments, the first editing template comprises a sequence exogenous to the double-stranded target DNA. In some embodiments, the second editing template comprises a sequence exogenous to the double-stranded target DNA. In some embodiments, the sequence in the first editing template exogenous to the double-stranded target DNA comprises a region of complementarity or substantial complementarity to a sequence in the second editing template exogenous to the double-stranded target DNA. In some embodiments, the sequence in the first editing template exogenous to the double-stranded target DNA further comprises a region that is not complementary to a sequence in the second editing template exogenous to the double-stranded target DNA. In some embodiments, the sequence in the second editing template exogenous to the double-stranded target DNA comprises a region of complementarity or substantial complementarity to a sequence in the first editing template exogenous to the double-stranded target DNA. In some embodiments, the sequence in the second editing template exogenous to the double-stranded target DNA further comprises a region that is not complementary to the sequence in the first editing template exogenous to the double-stranded target DNA. In some embodiments, the first editing template comprises a sequence exogenous to the double-stranded target DNA, wherein the sequence exogenous to the double-stranded target DNA comprises a polynucleotide sequence encoding an expression tag, such as an affinity tag, His tag, V5 tag, or FLAG tag. In some embodiments, the second editing template comprises a sequence exogenous to the double-stranded target DNA, wherein the sequence exogenous to the double-stranded target DNA comprises a polynucleotide sequence encoding an expression tag, such as an affinity tag, His tag, V5 tag, or FLAG tag. In some embodiments, the first editing template and / or the second editing template comprise a sequence exogenous to the double-stranded target DNA, wherein the exogenous sequence comprises a recombinase recognition sequence (RSS), such as an attB sequence.
[0314] Thus, in some embodiments, the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA each comprise a region of complementarity or substantial complementarity to each other. In some embodiments, the first newly synthesized single-stranded DNA comprises a sequence exogenous to the double-stranded target DNA. In some embodiments, the second newly synthesized single-stranded DNA comprises a sequence exogenous to the double-stranded target DNA. In some embodiments, the sequence in the first newly synthesized single-stranded DNA exogenous to the double-stranded target DNA comprises a region of complementarity or substantial complementarity to a sequence in the second newly synthesized single-stranded DNA exogenous to the double-stranded target DNA. In some embodiments, the sequence in the first newly synthesized single-stranded DNA exogenous to the double-stranded target DNA further comprises a region that is not complementary to a sequence in the second newly synthesized single-stranded DNA exogenous to the double-stranded target DNA. In some embodiments, the sequence in the second newly synthesized single-stranded DNA exogenous to the double-stranded target DNA comprises a region of complementarity or substantial complementarity to a sequence in the first newly synthesized single-stranded DNA exogenous to the double-stranded target DNA, hi some embodiments, the sequence in the second newly synthesized single-stranded DNA exogenous to the double-stranded target DNA further comprises a region that is not complementary to a sequence in the first newly synthesized single-stranded DNA exogenous to the double-stranded target DNA.
[0315] Thus, in some embodiments, a first newly synthesized single-stranded DNA and a second newly synthesized single-stranded DNA form an OD containing a sequence exogenous to the double-stranded target DNA, e.g., a TRAC gene. In some embodiments, a first newly synthesized single-stranded DNA and a second newly synthesized single-stranded DNA form an RD containing a sequence exogenous to the double-stranded target DNA, e.g., a TRAC gene. Prime editing, in some embodiments, results in the IND being excised and replaced with the RD. In some embodiments, the IND is excised and replaced with the RD. Thus, in some embodiments, the IND in the target gene, e.g., a TRAC gene, is deleted and replaced with an exogenous sequence, e.g., one or more recombinase recognition sequences.
[0316] In some embodiments, the first editing template comprises a sequence that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA, e.g., the TRAC gene. In some embodiments, the second editing template comprises a sequence that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA, e.g., the TRAC gene.
[0317] In some embodiments, the sequence of the first editing template that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA comprises a region of complementarity or substantial complementarity to the sequence of the second editing template that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA. In some embodiments, the sequence of the first editing template that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA further comprises a region that is not complementary to the sequence of the second editing template that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA. In some embodiments, the sequence of the second editing template that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA comprises a region of complementarity or substantial complementarity to the sequence of the first editing template that is complementary or substantially complementary to the endogenous sequence of the double-stranded target DNA. In some embodiments, the sequence of the second editing template that has complementarity or substantial complementarity to the endogenous sequence of the double-stranded target DNA further comprises a region that is not complementary to the sequence of the first editing template that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA.
[0318] Thus, in some embodiments, the first newly synthesized single-stranded DNA comprises a sequence that is identical or substantially identical to the endogenous sequence of the double-stranded target DNA, e.g., the TRAC gene. In some embodiments, the second newly synthesized single-stranded DNA comprises a sequence that is identical or substantially identical to the endogenous sequence of the double-stranded target DNA, e.g., the TRAC gene. In some embodiments, the first newly synthesized single-stranded DNA and / or the second newly synthesized single-stranded DNA comprises a sequence that is identical or substantially identical to the endogenous sequence of the double-stranded target DNA. In some embodiments, the first newly synthesized single-stranded DNA comprises a sequence that is identical or substantially identical to the endogenous sequence on the second strand of the double-stranded target DNA, e.g., the TRAC gene. In some embodiments, the second newly synthesized single-stranded DNA comprises a sequence that is identical or substantially identical to the endogenous sequence on the first strand of the double-stranded target DNA, e.g., the TRAC gene.
[0319] In some embodiments, the sequence of a first newly synthesized single-stranded DNA that has identity or substantial identity to an endogenous sequence of the double-stranded target DNA comprises a region of complementarity or substantial complementarity to a sequence of a second newly synthesized single-stranded DNA that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA. In some embodiments, the sequence of a first newly synthesized single-stranded DNA that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA further comprises a region that is not complementary to the sequence of a second newly synthesized single-stranded DNA that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA. In some embodiments, the sequence of a second newly synthesized single-stranded DNA that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA comprises a region of complementarity or substantial complementarity to the sequence of a first newly synthesized single-stranded DNA that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA. In some embodiments, the sequence of the second newly synthesized single-stranded DNA that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA further comprises a region that is not complementary to the sequence of the first newly synthesized single-stranded DNA that has identity or substantial identity to the endogenous sequence of the double-stranded target DNA.
[0320] Thus, in some embodiments, a first newly synthesized single-stranded DNA and a second newly synthesized single-stranded DNA form a double-stranded target DNA, e.g., an OD that includes the endogenous sequence of the TRAC gene. In some embodiments, a first newly synthesized single-stranded DNA and a second newly synthesized single-stranded DNA form a double-stranded target DNA, e.g., an RD that includes the endogenous sequence of the TRAC gene. In some embodiments, the RD or OD includes the double-stranded target DNA, e.g., the endogenous sequence of the TRAC gene. In some embodiments, the RD or OD includes a sequence that is exogenous compared to the double-stranded target DNA, e.g., the TRAC gene.
[0321] In some embodiments, the first editing template and / or the second editing template are partially complementary, substantially complementary, or identical to a sequence of IND. In some embodiments, for example, the first editing template comprises a region complementary to or identical to a region of the sequence of IND. In some embodiments, the first editing template comprises a region of complementarity to a sequence on the first PAM strand of IND. In some embodiments, the first editing template further comprises a region of complementarity to a second editing template. In some embodiments, the first editing template is partially complementary, substantially complementary, or identical to a sequence of IND and is also substantially complementary to the second editing template. In some embodiments, the second editing template comprises a region complementary to or identical to a region of the sequence of IND. In some embodiments, the second editing template comprises a region of complementarity to a sequence on the second PAM strand of IND. In some embodiments, the second editing template further comprises a region of complementarity to the first editing template. In some embodiments, the second editing template is partially complementary, substantially complementary, or identical to the sequence of IND and is also substantially complementary to the first editing template, hi some embodiments, the first editing template and the second editing template each comprise a region of complementarity to the sequence of IND.
[0322] The partially complementary regions within the first and second editing templates can be located anywhere within the first and second editing templates. Thus, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template are partially complementary to each other at any location within the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template. In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to the first strand of IND at or near the 3' end of the first newly synthesized single-stranded DNA. In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to the first strand of IND at or near the 5' end of the first newly synthesized single-stranded DNA. In some embodiments, the first newly synthesized single-stranded DNA comprises a region of complementarity to the first strand of IND in the center of the first newly synthesized single-stranded DNA. In some embodiments, the second newly synthesized single-stranded DNA comprises a region of complementarity to the second strand of IND at or near the 3' end of the second newly synthesized single-stranded DNA. In some embodiments, the second newly synthesized single-stranded DNA comprises a region of complementarity to the second strand of IND at or near the 5' end of the second newly synthesized single-stranded DNA. In some embodiments, the second newly synthesized single-stranded DNA comprises a region of complementarity to the second strand of IND in the center of the second newly synthesized single-stranded DNA. In some embodiments, the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA each comprise a region of complementarity to each other at their 3' ends.
[0323] Thus, as illustrated in Figures 4C-4D, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template comprises a region identical to a region of sequence on the first PAM strand of IND. In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template comprises a region identical to a region of sequence on the second PAM strand of IND. In some embodiments, the first newly synthesized single-stranded DNA comprises two or more subregions, each identical to a subregion of sequence on the first PAM strand of IND (as illustrated in Figure 4D). The subregions on the first newly synthesized single-stranded DNA and / or the first PAM strand of IND may or may not be contiguous. For example, the first newly synthesized single-stranded DNA may contain two subregions, each identical to a subregion of the sequence on the first PAM strand of IND, where the two subregions of the sequence on the first PAM strand of IND are separated by a region that has no identity or substantial identity to the first newly synthesized single-stranded DNA. In some embodiments, the second newly synthesized single-stranded DNA contains two or more subregions, each identical to a subregion of the sequence on the second PAM strand of IND (as illustrated in Figure 4D). The subregions on the second newly synthesized single-stranded DNA and / or the second PAM strand of IND may or may not be contiguous. For example, the second newly synthesized single-stranded DNA may contain two subregions, each identical to a subregion of the sequence on the second PAM strand of IND, where the two subregions of the sequence on the second PAM strand of IND are separated by a region that has no identity or substantial identity to the second newly synthesized single-stranded DNA.
[0324] In some embodiments, a region of the sequence on the first PAM strand of the IND and a region of the sequence on the second PAM strand of the IND are complementary to each other. In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template are at least partially complementary to each other and can anneal to form an OD. In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template are substantially complementary or complementary to each other and can anneal to form an OD. In some embodiments, the IND is excised and the OD is incorporated into the double-stranded target DNA at the position of the IND excision. As a result, the portion of the IND that is not complementary or identical to the first editing template or the second editing template is deleted from the double-stranded target DNA. In some embodiments, the deletion is at the 3' end of the IND. In some embodiments, the deletion is at the 5' end of the IND. In some embodiments, the deletion is in the center of the IND.
[0325] In some embodiments, the first editing template of the first PEgRNA is at least partially complementary, substantially complementary, at least partially identical, or identical to a sequence of the double-stranded target DNA outside the IND. "Outside the IND" refers to a sequence or location of the double-stranded target DNA that is not between the two nick sites generated by the first prime editor and the second prime editor. In some embodiments, the first editing template of the first PEgRNA comprises a region of identity to a sequence outside the IND on the second PAM strand (or first strand) of the double-stranded target DNA. In some embodiments, the first editing template of the first PEgRNA comprises a region of identity to a sequence on the first strand of the double-stranded target DNA adjacent to the second nick site generated by the second prime editor complexed with the second PEgRNA, this sequence being outside the IND. In some embodiments, the first editing template of the first PEgRNA comprises a region of identity to a sequence on the first strand of the double-stranded target DNA immediately adjacent to the second nick site generated by the second prime editor complexed with the second PEgRNA, where this sequence is outside of the IND.
[0326] Thus, as illustrated in Figure 4E, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template comprises a region of complementarity to a sequence on the first strand of the double-stranded target DNA adjacent to or immediately adjacent to the second nick site generated by the second prime editor complexed with the second PEgRNA. In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template comprises a region of complementarity to a sequence on the first strand of the double-stranded target DNA immediately adjacent to the second nick site generated by the second prime editor complexed with the second PEgRNA, this sequence being outside the IND. In some embodiments, the first newly synthesized single-stranded DNA anneals to a sequence on the first strand of the double-stranded target DNA adjacent to or immediately adjacent to the second nick site generated by the second prime editor. In some embodiments, DNA repair results in the IND being excised and deleted from the double-stranded target DNA, e.g., the target gene.
[0327] In some embodiments, the second editing template of the second PEgRNA is at least partially complementary, substantially complementary, at least partially identical, or identical to a sequence of the double-stranded target DNA outside the IND. In some embodiments, the second editing template of the second PEgRNA comprises a region of identity to a sequence outside the IND on the first PAM strand (or second strand) of the double-stranded target DNA. In some embodiments, the second editing template of the second PEgRNA comprises a region of identity to a sequence on the second strand of the double-stranded target DNA adjacent to the first nick site generated by the first prime editor complexed with the first PEgRNA, this sequence being outside the IND. In some embodiments, the second editing template of the second PEgRNA comprises a region of identity to a sequence on the second strand of the double-stranded target DNA directly adjacent to the first nick site generated by the first prime editor complexed with the first PEgRNA, this sequence being outside the IND.
[0328] Thus, as illustrated in Figure 4E, in some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template comprises a region of complementarity to a sequence on the second strand of the double-stranded target DNA adjacent to or immediately adjacent to the first nick site generated by the first prime editor complexed with the first PEgRNA. In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template comprises a region of complementarity to a sequence on the second strand of the double-stranded target DNA adjacent to or immediately adjacent to the first nick site generated by the first prime editor complexed with the first PEgRNA, this sequence being outside the IND. In some embodiments, the second newly synthesized single-stranded DNA anneals to a sequence on the second strand of the double-stranded target DNA adjacent to or immediately adjacent to the first nick site generated by the first prime editor. In some embodiments, DNA repair results in the IND being excised and deleted from the double-stranded target DNA, e.g., the target gene.
[0329] In some embodiments, the first editing template of the first PEgRNA comprises a region at least partially identical to a sequence on the first strand of the double-stranded target DNA immediately adjacent to the second nick site generated by the second prime editor complexed with the second PEgRNA, where this sequence is outside the IND. In some embodiments, the second editing template of the second PEgRNA comprises a region at least partially identical to a sequence on the second strand of the double-stranded target DNA immediately adjacent to the first nick site generated by the first prime editor complexed with the first PEgRNA, where this sequence is outside the IND. In some embodiments, the first editing template and the second editing template further comprise regions of complementarity or substantial complementarity to each other.
[0330] Thus, as illustrated in Figure 4F, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template comprises a region of complementarity to a sequence on the first strand of the double-stranded target DNA, where the sequence is immediately adjacent to the second nick site generated by the second prime editor complexed with the second PEgRNA and is outside the IND. In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template comprises a region of complementarity or substantial complementarity to a sequence on the second strand of the double-stranded target DNA, where the sequence is immediately adjacent to the first nick site generated by the first prime editor complexed with the first PEgRNA and is outside the IND. In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template further comprise regions of complementarity or substantial complementarity to each other and can anneal to each other to form an OD.
[0331] In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template further comprises a region that is not complementary to the second newly synthesized single-stranded DNA encoded by the second editing template, and has no complementarity or identity to the double-stranded target DNA. In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template further comprises a region that is not complementary to the first newly synthesized single-stranded DNA encoded by the first editing template, and has no complementarity or identity to the double-stranded target DNA, e.g., a target gene.
[0332] Thus, in some embodiments, the RD comprises (i) the OD, (ii) a region of the first newly synthesized single-stranded DNA that is not complementary to the second newly synthesized single-stranded DNA and has no complementarity or identity to the double-stranded target DNA, and its complementary sequence, and (iii) a region of the second newly synthesized single-stranded DNA that is not complementary to the first newly synthesized single-stranded DNA and has no complementarity or identity to the double-stranded target DNA, and its complementary sequence. In some embodiments, DNA repair results in the IND being excised from the double-stranded target DNA, e.g., the TRAC gene, and the RD being incorporated into the double-stranded target DNA.
[0333] In some embodiments, the IND is excised and deleted from the target gene, and the RD is integrated at the site of the excision of the IND. In some embodiments, the IND is excised and deleted from the target gene, and the OD is integrated at the site of the excision of the IND. In some embodiments, the RD comprises a region of identity to an endogenous sequence of the double-stranded target DNA. In some embodiments, the OD comprises a region of identity to an endogenous sequence of the double-stranded target DNA. In some embodiments, the RD does not have sequence identity to an endogenous sequence of the double-stranded target DNA. In some embodiments, the RD is exogenous to the double-stranded target DNA, e.g., the target gene. In some embodiments, the RD has a biological function or encodes a polypeptide with a biological function. In some embodiments, the OD does not have sequence identity to an endogenous sequence of the double-stranded target DNA. In some embodiments, the OD is exogenous to the double-stranded target DNA, e.g., the target gene. In some embodiments, the OD has a biological function or encodes a polypeptide with a biological function.
[0334] In some embodiments, the first editing template of the first PEGRNA comprises a region at least partially identical to a sequence of the double-stranded target DNA that is outside the IND and not immediately adjacent to (also referred to as "distal to") the second nick site on the second PAM strand of the double-stranded target DNA. In some embodiments, the first editing template of the first PEGRNA comprises a region of identity to a sequence of the double-stranded target DNA on the second PAM strand that is outside the IND and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides downstream from the second nick site.
[0335] In some embodiments, the second editing template of the second PEGRNA comprises a region at least partially identical to a sequence of the double-stranded target DNA that is outside the IND and not immediately adjacent to the first nick site on the first PAM strand of the double-stranded target DNA. In some embodiments, the second editing template of the second PEGRNA comprises a region of identity to a sequence of the double-stranded target DNA on the first PAM strand that is outside the IND and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides downstream from the first nick site. In some embodiments, the second editing template of the second PEGRNA comprises a region of identity to a sequence of the double-stranded target DNA that is outside the IND and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides upstream from the first nick site.
[0336] Thus, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template contains a region of complementarity to a sequence of the double-stranded target DNA that is outside the IND and not immediately adjacent to (i.e., distal to) the second nick site on the second PAM strand. In some embodiments, the first newly synthesized DNA encoded by the first editing template is capable of annealing to a sequence that is outside the IND and not immediately adjacent to the second nick site on the second PAM strand of the double-stranded target DNA. In some embodiments, the first newly synthesized DNA encoded by the first editing template contains and is capable of annealing to a region of complementarity to a sequence of the double-stranded target DNA that is outside the IND and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides downstream from the second nick site.
[0337] In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template contains a region of complementarity to a sequence of the first PAM strand of the double-stranded target DNA that is outside the IND and not immediately adjacent to (also referred to as "distal to") the first nick site on the first PAM strand. In some embodiments, the second newly synthesized DNA encoded by the second editing template is capable of annealing to a sequence that is outside the IND and not immediately adjacent to (also referred to as "distal to") the first nick site on the first PAM strand of the double-stranded target DNA. In some embodiments, the second newly synthesized DNA encoded by the second editing template contains and is capable of annealing to a sequence of the double-stranded target DNA that is outside the IND and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides upstream from the first nick site.
[0338] In some embodiments, DNA repair results in the IND being excised and deleted from the double-stranded target DNA, e.g., the target gene. In some embodiments, an endogenous sequence of the double-stranded target DNA between the 3' end of the sequence outside the IND and distal to the second nick site on the second PAM strand and the 3' end of the sequence of the double-stranded target DNA outside the IND and distal to the first nick site on the first PAM strand of the double-stranded target DNA is excised and deleted from the double-stranded target DNA.
[0339] In some embodiments, as illustrated in Figure 4G, the first search target sequence is downstream of the second search target sequence. In some embodiments, the first editing template includes a region that is complementary or substantially complementary to the second editing template, and optionally further includes a region that is not complementary to the second editing template. In some embodiments, the second editing template includes a region that is complementary or substantially complementary to the first editing template, and optionally further includes a region that is not complementary to the first editing template. Thus, in some embodiments, the first newly synthesized single-stranded DNA includes a region that is complementary to the second newly synthesized single-stranded DNA, and optionally further includes a region that is not complementary to the second newly synthesized single-stranded DNA, and the first newly synthesized single-stranded DNA is downstream of the second newly synthesized single-stranded DNA. In some embodiments, the second newly synthesized single-stranded DNA comprises a region of complementarity to the first newly synthesized single-stranded DNA, and optionally further comprises a region that does not have complementarity to the first newly synthesized single-stranded DNA, and the first newly synthesized single-stranded DNA is downstream of the second newly synthesized single-stranded DNA. DNA repair incorporates the OD or RD sequence into the double-stranded target DNA, and the IND sequence is replicated in the double-stranded target DNA.
[0340] Prime Editor The term "prime editor (PE)" refers to a polypeptide or polypeptide component involved in prime editing. In various embodiments, a prime editor comprises a polypeptide domain with DNA-binding activity and a polypeptide domain with DNA polymerase activity. In some embodiments, the polypeptide domain with DNA-binding activity is a polypeptide domain with programmable DNA-binding activity. In some embodiments, a prime editor further comprises a polypeptide domain with nuclease activity. In some embodiments, the polypeptide domain with DNA-binding activity comprises a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain with nuclease activity comprises a nickase or a fully active nuclease. As used herein, the term "nickase" refers to a nuclease that can cleave only one strand of a double-stranded DNA target. In some embodiments, a prime editor comprises a polypeptide domain that is an inactive nuclease. In some embodiments, the polypeptide domain with programmable DNA-binding activity comprises a nucleic acid-guided DNA-binding domain, e.g., a CRISPR-Cas protein, e.g., Cas9 nickase, Cpf1 nickase, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain with DNA polymerase activity comprises a template-dependent DNA polymerase, e.g., a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase. In some embodiments, the prime editor comprises an additional polypeptide or polypeptide domain involved in prime editing, e.g., a polypeptide domain with 5' endonuclease activity, e.g., a 5' endogenous DNA flap endonuclease (e.g., FEN1), which helps drive the prime editing process toward the formation of an edited product. In some embodiments, the prime editor further comprises an RNA protein recruitment polypeptide, e.g., MS2 coat protein.
[0341] Prime editors can be designed. In some embodiments, the polypeptide components of a prime editor do not naturally occur within the same organism or cellular environment. In some embodiments, the polypeptide components of a prime editor can be of different origins or from different organisms. In some embodiments, a prime editor comprises a DNA-binding domain and a DNA polymerase domain from different species. In some embodiments, a prime editor comprises a Cas polypeptide (DNA-binding domain) and a reverse transcriptase polypeptide (DNA polymerase) from different species. For example, a prime editor can comprise a S. pyogenes Cas9 polypeptide and a Moloney murine leukemia virus (M-MLV) reverse transcriptase polypeptide.
[0342] In some embodiments, prime editor polypeptide domains can be fused or linked by a peptide linker to form a fusion protein. In other embodiments, a prime editor comprises one or more polypeptide domains provided in trans as separate proteins that can bind to each other via non-peptide bonds or aptamers or recruitment sequences. For example, a prime editor may comprise a DNA-binding domain and a reverse transcriptase domain linked to each other by an RNA-protein recruitment aptamer (e.g., an MS2 aptamer that can be linked to a PEG RNA). Prime editor polypeptide components can be wholly or partially encoded by one or more polynucleotides. In some embodiments, a single polynucleotide, construct, or vector encodes a prime editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or portion of a domain of a prime editor, or a portion of a prime editor fusion protein. For example, a prime editor fusion protein may comprise an N-terminal portion fused to intein N and a C-terminal portion fused to intein C, each separately encoded by an AAV vector.
[0343] The term "prime editor complex" is used interchangeably with the term "prime editing complex" and refers to a complex comprising one or more prime editor components (e.g., a polypeptide domain having DNA-binding activity and a polypeptide domain having DNA polymerase activity) complexed with a PEGRNA.
[0344] Prime Editor Nucleotide Polymerase Domain In some embodiments, the prime editor comprises a nucleotide polymerase domain, e.g., a DNA polymerase domain. The DNA polymerase domain can be a wild-type DNA polymerase domain, a full-length DNA polymerase protein domain, or a functional mutant, functional variant, or functional fragment thereof. In some embodiments, the polymerase domain is a template-dependent polymerase domain. For example, a DNA polymerase can rely on a template polynucleotide strand (e.g., an edited template sequence) for synthesis of a new strand of DNA. In some embodiments, the prime editor comprises a DNA-dependent DNA polymerase. For example, a prime editor having a DNA-dependent DNA polymerase can synthesize a novel single-stranded DNA using a PEgRNA edited template that includes a DNA sequence as a template. In such cases, the PEgRNA is a chimeric or hybrid PEgRNA and includes an extension arm that includes a DNA strand. As used herein, an "extension arm" is a polynucleotide portion of a PEgRNA that includes an edited template and a primer binding site sequence (PBS). In some embodiments, the extension arm further includes an additional component, e.g., a 3' modifier. A chimeric or hybrid PEGRNA may comprise an RNA portion (including a spacer and a gRNA core) and a DNA portion (an extension arm comprising an editing template comprising a DNA strand).
[0345] The DNA polymerase can be a wild-type polymerase from a eukaryotic, prokaryotic, archaeal, or viral organism, and / or the polymerase can be modified by processes based on genetic engineering, mutagenesis, or directed evolution. The polymerase can be T7 DNA polymerase, T5 DNA polymerase, T4 DNA polymerase, Klenow fragment DNA polymerase, DNA polymerase III, etc. The polymerase can be thermostable and include Taq, Tne, Tma, Pfu, Tfl, Tth, Stoffel fragment, VENT® and DEEPVENT® DNA polymerase, KOD, Tgo, JDF3, and mutants, variants, and derivatives thereof.
[0346] In some embodiments, the DNA polymerase is a bacteriophage polymerase, e.g., T4, T7, or phi29 DNA polymerase. In some embodiments, the DNA polymerase is an archaeal polymerase, e.g., a pol I-type archaeal polymerase or a pol II-type archaeal polymerase. In some embodiments, the DNA polymerase comprises a thermostable archaeal DNA polymerase. In some embodiments, the DNA polymerase comprises a eubacterial DNA polymerase, e.g., a Pol I, Pol II, or Pol III polymerase. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is an E. coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is a Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is E. coli Pol IV DNA polymerase.
[0347] In some embodiments, the DNA polymerase comprises a eukaryotic DNA polymerase. In some embodiments, the DNA polymerase is Pol-beta DNA polymerase, Pol-lambda DNA polymerase, Pol-sigma DNA polymerase, or Pol-mu DNA polymerase. In some embodiments, the DNA polymerase is Pol-alpha DNA polymerase. In some embodiments, the DNA polymerase is POLA1 DNA polymerase. In some embodiments, the DNA polymerase is POLA2 DNA polymerase. In some embodiments, the DNA polymerase is Pol-delta DNA polymerase. In some embodiments, the DNA polymerase is POLD1 DNA polymerase. In some embodiments, the DNA polymerase is POLD2 DNA polymerase. In some embodiments, the DNA polymerase is human POLD1 DNA polymerase. In some embodiments, the DNA polymerase is human POLD2 DNA polymerase. In some embodiments, the DNA polymerase is POLD3 DNA polymerase. In some embodiments, the DNA polymerase is a POLD4 DNA polymerase. In some embodiments, the DNA polymerase is a Pol-epsilon DNA polymerase. In some embodiments, the DNA polymerase is a POLE1 DNA polymerase. In some embodiments, the DNA polymerase is a POLE2 DNA polymerase. In some embodiments, the DNA polymerase is a POLE3 DNA polymerase. In some embodiments, the DNA polymerase is a Pol-eta (POLH) DNA polymerase. In some embodiments, the DNA polymerase is a Pol-iota (POLI) DNA polymerase. In some embodiments, the DNA polymerase is a Pol-kappa (POLK) DNA polymerase. In some embodiments, the DNA polymerase is a Rev1 DNA polymerase. In some embodiments, the DNA polymerase is a human Rev1 DNA polymerase.In some embodiments, the DNA polymerase is a viral DNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a family B DNA polymerase. In some embodiments, the DNA polymerase is herpes simplex virus (HSV) UL30 DNA polymerase. In some embodiments, the DNA polymerase is cytomegalovirus (CMV) UL54 DNA polymerase.
[0348] In some embodiments, the DNA polymerase is an archaeal polymerase. In some embodiments, the DNA polymerase is a family B / pol I type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of Pfu from Pyrococcus furiosus. In some embodiments, the DNA polymerase is a Pol II type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of the P. furiosus DP1 / DP2 two-subunit polymerase. In some embodiments, the DNA polymerase lacks 5' to 3' nuclease activity. Suitable DNA polymerases (pol I or pol II) can be obtained from archaea with optimal growth temperatures close to the desired assay temperature.
[0349] In some embodiments, the DNA polymerase comprises a thermostable archaeal DNA polymerase, hi some embodiments, the thermostable DNA polymerase is isolated or derived from Pyrococcus species (furiosus, species GB-D, woesii, abysii, horikoshii), Thermococcus species (kodakaraensis KOD1, litoralis, species 9 degrees North-7, species JDF-3, gorgonarius), Pyrodictium occultum, and Archaeoglobus fulgidus.
[0350] The polymerase may also be from a eubacterial species. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is an E. coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is a Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol III family DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is an E. coli Pol IV DNA polymerase. In some embodiments, the Pol I DNA polymerase is a functional variant of a DNA polymerase lacking or having reduced 5' to 3' exonuclease activity.
[0351] Suitable thermostable pol I DNA polymerases can be isolated from a variety of thermophilic eubacteria, including Thermus species and Thermotoga maritima, such as Thermus aquaticus (Taq), Thermus thermophilus (Tth), and Thermotoga maritima (Tma UlTma).
[0352] In some embodiments, the prime editor comprises an RNA-dependent DNA polymerase domain, such as a reverse transcriptase (RT). The RT or RT domain can be a wild-type RT domain, a full-length RT domain, or a functional mutant, functional variant, or functional fragment thereof. The RT or RT domain of the prime editor can comprise a wild-type RT or can be designed or evolved to contain specific amino acid substitutions, truncations, or variants. An engineered RT can contain sequence or amino acid changes that differ from a naturally occurring RT. In some embodiments, the engineered RT can have improved reverse transcription activity compared to a naturally occurring RT or RT domain. In some embodiments, the engineered RT can have improved functionality, such as thermostability, reverse transcription efficiency, or target fidelity, compared to a naturally occurring RT. In some embodiments, a prime editor comprising a designed RT has improved prime editing efficiency compared to a prime editor with a reference naturally occurring RT.
[0353] In some embodiments, the prime editor comprises a viral RT, e.g., a retroviral RT. Non-limiting examples of viral RTs include Moloney murine leukemia virus (M-MLV, MMLV RT, M-MLV RT, or MLV RT); human T-cell leukemia virus type 1 (HTLV-1) RT; bovine leukemia virus (BLV) RT; Rous sarcoma virus (RSV) RT; human immunodeficiency virus (HIV) RT, M-MFV RT, avian sarcoma leukosis virus (ASLV) RT, Rous sarcoma virus (RSV) RT, avian myeloblastosis virus (AMV) RT, avian erythroblastosis virus (AEV) helper virus MCAV RT, avian myelocytomatosis virus MC29 helper virus MCAV RT, avian reticuloendotheliosis virus (REV-T) helper virus REV-A RT, avian sarcoma virus UR2 helper virus (UR2AV) RT, and avian sarcoma virus Y73 helper virus YAV. RT, Rous-associated virus (RAV) RT, and myeloblastosis-associated virus (MAV) RT, all of which may be suitably used in the methods and compositions described herein.
[0354] In some embodiments, the prime editor comprises a wild-type M-MLV RT, a functional mutant, functional variant, or functional fragment thereof. Table 1A provides the sequences of exemplary M-MLV RTs suitable for use in the compositions and methods of the disclosure.
[0355] In some embodiments, the prime editor comprises a wild-type M-MLV RT set forth in SEQ ID NO: 1002. In some embodiments, the prime editor comprises a variant M-MLV RT set forth in SEQ ID NO: 1001. In some embodiments, the prime editor comprises a variant M-MLV RT set forth in SEQ ID NO: 1003. [Table 1A-1] [Table 1A-2]
[0356] In some embodiments, the prime editor comprises an M-MLV RT that includes one or more of the amino acid substitutions P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, or D653X relative to a reference M-MLV RT, where X is any amino acid other than the original amino acid in the reference M-MLV RT. In some embodiments, the prime editor comprises an M-MLV RT that includes one or more of the amino acid substitutions P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, or D653N compared to a reference M-MLV RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT set forth in SEQ ID NO: 1001. In some embodiments, the M-MLV RT is a WT M-MLV RT set forth in SEQ ID NO: 1002. In some embodiments, the prime editor comprises an M-MLV RT that includes one or more amino acid substitutions D200N, T330P, L603W, T306K, or W313F compared to a reference M-MLV RT. In some embodiments, the prime editor comprises an M-MLV RT that includes the amino acid substitutions D200N, T330P, L603W, T306K, and W313F compared to a reference M-MLV RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT set forth in SEQ ID NO: 1001. In some embodiments, the reference M-MLV RT is a WT M-MLV RT set forth in SEQ ID NO: 1002.
[0357] In some embodiments, an RT variant can be a functional fragment of a reference RT having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500 or more amino acid changes compared to the reference RT. In some embodiments, the RT variant comprises a fragment of a reference RT, such that the fragment is about 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% identical to the corresponding fragment of the reference RT. In some embodiments, the fragment is 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to the amino acid length of the corresponding reference RT (M-MLV reverse transcriptase). The reference RT may be any one of the RTs listed in Table 1A.
[0358] In some embodiments, an RT functional fragment is at least 100 amino acids in length, hi some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, or up to 600 or more amino acids in length.
[0359] In yet other embodiments, functional RT variants are truncated at the N-terminus or C-terminus, or both, by a specific number of amino acids, resulting in a truncated variant that still retains sufficient DNA polymerase function. In some embodiments, a functional RT variant, e.g., a functional MMLV RT variant, is truncated at the C-terminus to eliminate or reduce RNAse H activity while still retaining DNA polymerase activity. In some embodiments, a functional RT variant is truncated by at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the N-terminus compared to a reference RT, e.g., a wild-type RT. In other embodiments, the RT truncation variant is truncated by at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the C-terminus compared to a reference RT, e.g., wild-type RT. In some embodiments, the reference RT is wild-type M-MLV RT. In yet other embodiments, the RT truncation variant is N-terminally and C-terminally truncated compared to a reference RT, eg, a wild-type RT.In some embodiments, the N-terminal and C-terminal truncations are the same length. In some embodiments, the N-terminal and C-terminal truncations are different lengths.
[0360] For example, a prime editor disclosed herein can comprise a functional variant of a wild-type M-MLV reverse transcriptase. In some embodiments, the prime editor comprises a functional variant of a wild-type M-MLV RT, wherein the functional variant of M-MLV RT is truncated after amino acid position 502 relative to a reference M-MLV RT. In some embodiments, the functional variant of M-MLV RT further comprises a D200X, T306X, W313X, and / or T330X amino acid substitution relative to a reference M-MLV RT, where X is any amino acid other than the original amino acid in the reference M-MLV RT. In some embodiments, the functional variant of M-MLV RT further comprises a D200N, T306K, W313F, and / or T330P amino acid substitution relative to a reference M-MLV RT, where X is any amino acid other than the original amino acid in the reference M-MLV RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT set forth in SEQ ID NO: 1001. In some embodiments, the M-MLV RT is a WT M-MLV RT set forth in SEQ ID NO: 1002. A DNA sequence encoding a prime editor comprising this truncated RT is 522 bp smaller than a prime editor comprising the full-length M-MLV RT, and is therefore potentially useful in applications where delivery of DNA sequences is difficult due to its size (e.g., adeno-associated viral and lentiviral delivery). In some embodiments, the prime editor comprises an M-MLV RT comprising an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to an amino acid sequence set forth in Table 1A. In some embodiments, the prime editor comprises an M-MLV RT comprising an amino acid sequence selected from the group consisting of the amino acid sequences set forth in Table 1A or variants or fragments thereof. In some embodiments, the prime editor comprises a variant M-MLV RT comprising the amino acid sequence set forth in SEQ ID NO: 1003.In some embodiments, the prime editor comprises a variant M-MLV RT comprising the amino acid sequence set forth in SEQ ID NO: 1004.
[0361] In some embodiments, a prime editing composition or prime editing system disclosed herein comprises a polynucleotide (e.g., DNA, RNA, e.g., mRNA) encoding an M-MLV RT. In some embodiments, the polynucleotide encodes an M-MLV RT comprising an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to an amino acid sequence set forth in Table 1A. In some embodiments, the polynucleotide encodes an M-MLV RT comprising an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to an amino acid sequence set forth in SEQ ID NOs: 1001, 1002, 1003, 1004. In some embodiments, the polynucleotide encodes a variant M-MLV RT comprising an amino acid sequence selected from the group consisting of the amino acid sequences set forth in Table 1A. In some embodiments, the polynucleotide encodes a variant M-MLV RT comprising the amino acid sequence set forth in SEQ ID NO: 1003. In some embodiments, the polynucleotide encodes a variant M-MLV RT comprising the amino acid sequence set forth in SEQ ID NO: 1004.
[0362] In some embodiments, the prime editor comprises a eukaryotic RT, such as a yeast, Drosophila, rodent, or primate RT. In some embodiments, the prime editor comprises a group II intron RT, such as a Geobacillus stearothermophilus group II intron (GsI-IIC) RT or a Eubacterium rectale group II intron (Eu.re.I2) RT. In some embodiments, the prime editor comprises a retron RT.
[0363] Programmable DNA-binding domain In some embodiments, the DNA-binding domain of the prime editor is a programmable DNA-binding domain. A programmable DNA-binding domain refers to a protein domain designed to bind to a specific nucleic acid sequence, e.g., a target DNA or target RNA. In some embodiments, the DNA-binding domain is a polynucleotide-programmable DNA-binding domain that can bind to a guide polynucleotide (e.g., PEGRNA) that guides the DNA-binding domain to a specific DNA sequence, e.g., an interrogation target sequence within a target gene. In some embodiments, the DNA-binding domain comprises a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein. The Cas protein may comprise any Cas protein described herein, or a functional fragment or variant thereof. In some embodiments, the DNA-binding domain may also comprise a zinc finger protein domain. In other cases, the DNA-binding domain comprises a transcription activator-like effector domain (TALE). In some embodiments, the DNA-binding domain comprises a DNA nuclease. For example, the DNA-binding domain of the prime editor may comprise an RNA-guided DNA endonuclease, e.g., a Cas protein. In some embodiments, the DNA binding domain comprises a zinc finger nuclease (ZFN) or a transcription activator-like effector domain nuclease (TALEN), in which one or more zinc finger motifs or TALE motifs are associated with one or more nucleases, e.g., a Fok I nuclease domain.
[0364] In some embodiments, the DNA-binding domain has nuclease activity. In some embodiments, the DNA-binding domain of the prime editor comprises an endonuclease domain with single-stranded DNA cleavage activity. For example, the endonuclease domain may comprise a FokI nuclease domain. In some embodiments, the DNA-binding domain of the prime editor comprises a nuclease with full nuclease activity. In some embodiments, the DNA-binding domain of the prime editor comprises a nuclease with modified or reduced nuclease activity compared to a wild-type endonuclease domain. For example, the endonuclease domain may comprise one or more amino acid substitutions compared to a wild-type endonuclease domain. In some embodiments, the DNA-binding domain of the prime editor has nickase activity. In some embodiments, the DNA-binding domain of the prime editor comprises a Cas protein domain that is a nickase. In some embodiments, compared to a wild-type Cas protein, the Cas nickase comprises one or more amino acid substitutions in the nuclease domain, thereby reducing or eliminating double-stranded nuclease activity but retaining DNA-binding activity. In some embodiments, the Cas nickase comprises an amino acid substitution in the HNH domain. In some embodiments, the Cas nickase comprises an amino acid substitution in the RuvC domain.
[0365] In some embodiments, the DNA binding domain comprises a CRISPR-associated protein (Cas protein) domain. The Cas protein can be a class 1 or class 2 Cas protein. The Cas protein can be a type I, type II, type III, type IV, type V Cas protein, or type VI Cas protein.Non-limiting examples of Cas proteins include Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csnl or Csx12), Cas10, CaslOd, Cas12a / Cpfl, Cas12b / C2c1, C as12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csyl, Csy2, Csy3, Csy4, Csel, Cse2 , Cse3, Cse4, Cse5e, Cscl, Csc2, Csa5, Csnl, Csn2, Csml, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr 4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx1S, Csx11, Csfl, Csf2, CsO, Csf4, Csdl, Csd2, Cstl, Cst2, Cshl, Csh2, Csal, Csa2, Csa3, Csa4, Csa5, Cas effector protein type II, Cas effector protein type V, Cas effector protein type VI, CARF, DinG, Cpfl, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, hyper-precise Cas9 variant (HypaCas9), Cas Cas proteins include Cas12b2, Cas12c (C2c3), C2c4, C2c8, C2c5, C2c10, C2c9, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, Cns2, Cas Φ, and homologs, modified or altered variants, mutants, and / or functional fragments thereof. Cas proteins can be chimeric Cas proteins fused to other proteins or polypeptides. Cas proteins can be chimeras of various Cas proteins, for example, containing domains of Cas proteins from different organisms.
[0366] The Cas protein, such as Cas9, can be from any suitable organism. In some aspects, the microorganism is Streptococcus pyogenes (S. pyogenes). In some aspects, the microorganism is Staphylococcus aureus (S. aureus). In some aspects, the microorganism is Streptococcus thermophilus (S. thermophilus). In some embodiments, the organism is Staphylococcus lugdunensis.
[0367] Non-limiting examples of suitable organisms include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas species, Crocosphaera watsonii, Cyanothece species, Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus species, Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter species, Nitrosococcus halophilus, Nitrosococcus watsoni, PseudoalteromonasExamples of suitable organisms include: Bacillus haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc species, Arthrospira maxima, Arthrospira platensis, Arthrospira species, Lyngbya species, Microcoleus chthonoplastes, Oscillatoria species, Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella novicida. In some embodiments, the organism is Streptococcus pyogenes (S. pyogenes). In some embodiments, the organism is Staphylococcus aureus (S. aureus). In some embodiments, the organism is Streptococcus thermophilus (S. thermophilus). In some embodiments, the organism is Staphylococcus lugdunensis (S. lugdunensis).
[0368] In some embodiments, the Cas protein can be derived from various bacterial species including, but not limited to: Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractor salsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonaspalustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinella succinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, and Francisella novicida.
[0369] In some embodiments, the Cas protein, e.g., Cas9, can be a wild-type or modified form of a Cas protein. In some embodiments, the Cas protein, e.g., Cas9, can be a nuclease-active variant, a nuclease-inactive variant, a nickase, or a functional variant or fragment of a wild-type Cas protein. In some embodiments, the Cas protein, e.g., Cas9, can include amino acid changes, such as deletions, insertions, substitutions, fusions, chimeras, or any combination thereof, compared to a wild-type version of the corresponding Cas protein. In some embodiments, the Cas protein can be a polypeptide having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity to a wild-type exemplary Cas protein.
[0370] Cas proteins, such as Cas9, can comprise one or more domains. Non-limiting examples of Cas domains include a guide nucleic acid recognition and / or binding domain, a nuclease domain (e.g., a DNase or RNase domain, RuvC, HNH), a DNA-binding domain, an RNA-binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain. In various embodiments, the Cas protein comprises a guide nucleic acid recognition and / or binding domain capable of interacting with a guide nucleic acid and one or more nuclease domains comprising catalytic activity for nucleic acid cleavage.
[0371] In some embodiments, a Cas protein, such as Cas9, comprises one or more nuclease domains. The Cas protein may comprise an amino acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a nuclease domain (e.g., RuvC domain, HNH domain) of a wild-type Cas protein. In some embodiments, the Cas protein comprises a single nuclease domain. For example, Cpfl may comprise a RuvC domain but lack an HNH domain. In some embodiments, the Cas protein comprises two nuclease domains, for example, a Cas9 protein may comprise an HNH nuclease domain and a RuvC nuclease domain.
[0372] In some embodiments, the prime editor comprises a Cas protein, such as Cas9, wherein all nuclease domains of the Cas protein are active. In some embodiments, the prime editor comprises a Cas protein with one or more inactive nuclease domains. One or more nuclease domains of the Cas protein (e.g., RuvC, HNH) can be deleted or mutated to render them non-functional or reduce their nuclease activity. In some embodiments, a Cas protein (e.g., Cas9) containing a mutation in its nuclease domain reduces (e.g., nickase) or eliminates nuclease activity while retaining the ability to target a nucleic acid locus with an interrogated target sequence when complexed with a guide nucleic acid (e.g., PEgRNA).
[0373] In some embodiments, the prime editor comprises a Cas nickase that binds sequence-specifically to a target gene and can generate a single-stranded break at a protospacer within the double-stranded DNA of the target gene, but cannot generate a double-stranded break. For example, the Cas nickase can cleave the edited strand (i.e., the PAM strand) or the non-edited strand of the target gene, but not both. In some embodiments, the prime editor comprises a Cas nickase (e.g., Cas9) that comprises two nuclease domains, one of which is modified to lack catalytic activity or deleted. In some embodiments, the Cas nickase of the prime editor comprises a nuclease-inactive RuvC domain and a nuclease-active HNH domain. In some embodiments, the Cas nickase of the prime editor comprises a nuclease-inactive HNH domain and a nuclease-active RuvC domain. In some embodiments, the prime editor comprises a Cas9 nickase with an amino acid substitution in the RuvC domain, e.g., an amino acid substitution that reduces or eliminates the nuclease activity of the RuvC domain. In some embodiments, the Cas9 nickase comprises a D10X amino acid substitution relative to wild-type S. pyogenes Cas9, where X is any amino acid other than D. In some embodiments, the prime editor comprises a Cas9 nickase with an amino acid substitution in the HNH domain, e.g., an amino acid substitution that reduces or eliminates the nuclease activity of the HNH domain. In some embodiments, the Cas9 nickase comprises a H840X amino acid substitution relative to wild-type S. pyogenes Cas9, where X is any amino acid other than H.
[0374] In some embodiments, the prime editor comprises a Cas protein that can bind to a target gene in a sequence-specific manner but lacks or has abolished nuclease activity and is unable to cleave either strand of double-stranded DNA within the target gene. Absent or lacking activity can refer to an enzymatic activity that is less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of the activity of a wild-type exemplary activity (e.g., wild-type Cas9 nuclease activity). In some embodiments, the Cas protein of the prime editor lacks any nuclease activity. A nuclease that lacks nuclease activity, such as Cas9, can be referred to as nuclease-inactive or "nuclease-dead" (abbreviated as "d"). Nuclease-dead Cas proteins (e.g., dCas, dCas9) can bind to a target polynucleotide but are unable to cleave the target polynucleotide. In some embodiments, the dead Cas protein is a dead Cas9 protein. In some embodiments, the prime editor comprises a nuclease-dead Cas protein, in which all of the nuclease domains (e.g., both the RuvC and HNH nuclease domains in a Cas9 protein; the RuvC nuclease domain in a Cpfl protein) are mutated to lack catalytic activity or are deleted.
[0375] Cas proteins can be modified. Cas proteins, such as Cas9, can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to alter other activities or properties of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, the Cas protein can be truncated to remove domains that are not essential for protein function, or the activity of a Cas protein can be optimized (e.g., enhanced or decreased).
[0376] Cas proteins can be fusion proteins. For example, Cas proteins can be fused to a cleavage domain, epigenetic modification domain, transcriptional regulatory domain, or polymerase domain. Cas proteins can also be fused to heterologous polypeptides, which can increase or decrease stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internally of the Cas protein.
[0377] In some embodiments, the Cas protein of the prime editor is a class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, a homolog, mutant, variant, or functional fragment thereof. As used herein, Cas9, Cas9 protein, Cas9 polypeptide, or Cas9 nuclease refers to an RNA-guided nuclease comprising one or more Cas9 nuclease domains and a Cas9 gRNA-binding domain capable of binding to a guide polynucleotide, e.g., a PEG-RNA. Cas9 protein may refer to a wild-type Cas9 protein from any organism, or a homolog, ortholog, or paralog from any organism, any functional mutant or variant thereof, or any functional fragment or domain thereof. In some embodiments, the prime editor comprises a full-length Cas9 protein. In some embodiments, the Cas9 protein may typically comprise at least about 50%, 60%, 70%, 80%, 90%, 100% sequence identity to a wild-type reference Cas9 protein (e.g., Cas9 from S. pyogenes). In some embodiments, the Cas9 comprises amino acid changes, e.g., deletions, insertions, substitutions, fusions, chimeras, or any combination thereof, compared to the wild-type reference Cas9 protein.
[0378] In some embodiments, the Cas9 protein can include a Cas9 protein from Streptococcus pyogenes (Sp), Staphylococcus aureus (Sa), Streptococcus canis (Sc), Streptococcus thermophilus (St), Staphylococcus lugdunensis (Slu), Neisseria meningitidis (Nm), Campylobacter jejuni (Cj), Francisella novicida (Fn), or Treponema denticola (Td), or any Cas9 homolog or ortholog from an organism known in the art. In some embodiments, the Cas9 polypeptide is, for example, an SpCas9 polypeptide comprising the amino acid sequence set forth in NCBI Accession No. WP_038431314, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is, for example, an SaCas9 polypeptide comprising the amino acid sequence set forth in Uniprot Accession No. J7RUA5, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is an ScCas9 polypeptide comprising, for example, the amino acid sequence set forth in Uniprot Accession No. A0A3P5YA78, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is an StCas9 polypeptide comprising, for example, the amino acid sequence set forth in NCBI Accession No. WP_007896501.1, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is an SluCas9 polypeptide comprising, for example, the amino acid sequence set forth in any of NCBI Accession Nos. WP_230580236.1, WP_250638315.1, WP_242234150.1, WP_241435384.1, WP_002460848.1, or KAK58371.1, or a fragment or variant thereof.In some embodiments, the Cas9 polypeptide is an NmCas9 polypeptide comprising the amino acid sequence set forth in, for example, any of NCBI Accession Nos. WP_002238326.1 or WP_061704949.1, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is a CjCas9 polypeptide comprising the amino acid sequence set forth in, for example, any of NCBI Accession Nos. WP_100612036.1, WP_116882154.1, WP_116560509.1, WP_116484194.1, WP_116479303.1, WP_115794652.1, or WP_100624872.1, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is an FnCas9 polypeptide comprising the amino acid sequence set forth in, for example, Uniprot Accession No. A0Q5Y3, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is a TdCas9 polypeptide comprising, for example, the amino acid sequence set forth in NCBI Accession No. WP_147625065.1, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is a chimera comprising domains from two or more organisms described herein or known in the art. In some embodiments, the Cas9 polypeptide is a Cas9 polypeptide from Streptococcus macacae comprising, for example, the amino acid sequence set forth in NCBI Accession No. WP_003079701.1, or a fragment or variant thereof. In some embodiments, the Cas9 polypeptide is a Cas9 polypeptide generated by replacing the PAM-interaction domain of SpCas9 with that of Streptococcus macacae Cas9 (Spy-mac Cas9). Exemplary Cas9 and Cas9 nickase variants are listed in Table 1B.
[0379] In some embodiments, the prime editor comprises a DNA-binding domain comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences listed in Table 1B. In some embodiments, the DNA-binding domain comprises an amino acid sequence that has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 differences, e.g., mutations, e.g., deletions, substitutions, and / or insertions, compared to any one of the amino acid sequences listed in Table 1B.
[0380] In some embodiments, the prime editor comprises a Cas9 protein comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences listed in Table 1B. In some embodiments, the prime editor comprises a Cas9 protein that is a Cas9 nickase comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the nickase sequences set forth in Table 1B. In some embodiments, the Cas9 protein comprises an amino acid sequence selected from the group consisting of the sequences set forth in Table 1B. In some embodiments, the prime editor comprises a Cas9 protein comprising an amino acid sequence that lacks an N-terminal methionine compared to the amino acid sequence set forth in Table 1B. In some embodiments, a prime editing composition or prime editing system disclosed herein comprises a polynucleotide (e.g., DNA, or RNA, e.g., mRNA) encoding a Cas9 protein comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences set forth in Table 1B.
[0381] In some embodiments, the Cas9 protein comprises a Cas9 protein from Streptococcus pyogenes (Sp), e.g., according to NC_002737.2:854751-858857, or a protein encoded by UniProt Q99ZW2, e.g., according to SEQ ID NO: 1005. In some embodiments, the prime editor comprises a Cas9 protein (e.g., SpCas9) according to any one of the sequences set forth in SEQ ID NOs: 1005-1008, or a variant thereof. In some embodiments, the Cas9 protein is SpCas9. In some embodiments, the SpCas9 can be wild-type SpCas9, an SpCas9 variant, or a nickase SpCas9. In some embodiments, the SpCas9 lacks an N-terminal methionine compared to the corresponding SpCas9 (e.g., wild-type SpCas9, an SpCas9 variant, or a nickase SpCas9). In some embodiments, the prime editor comprises a Cas9 protein having an amino acid sequence according to SEQ ID NO: 1005, without the N-terminal methionine. In some embodiments, wild-type SpCas9 comprises the amino acid sequence set forth in SEQ ID NO: 1005. In some embodiments, the prime editor comprises a Cas9 protein that comprises one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to a corresponding wild-type Cas9 protein (e.g., wild-type SpCas9). In some embodiments, the Cas9 protein that comprises one or more mutations compared to a wild-type Cas9 (e.g., wild-type SpCas9) protein comprises the amino acid sequence set forth in SEQ ID NO: 1006, 1007, or 1008. Exemplary Streptococcus pyogenes Cas9 (SpCas9) amino acid sequences useful in the prime editors disclosed herein are shown in Table 1B.
[0382] In some embodiments, the prime editor comprises a Cas9 protein (e.g., SluCas9) according to any one of SEQ ID NOs: 1009-1011, or a variant thereof. In some embodiments, the prime editor comprises a Cas9 protein from Staphylococcus lugdunensis (SluCas9), e.g., according to any one of SEQ ID NOs: 1009-1011, or a variant thereof. In some embodiments, the Cas9 protein is SluCas9. In some embodiments, the SluCas9 can be a wild-type SluCas9, a SluCas9 variant, or a nickase SluCas9. In some embodiments, the SluCas9 lacks an N-terminal methionine compared to the corresponding SluCas9 (e.g., a wild-type SluCas9, a SluCas9 variant, or a nickase SluCas9). In some embodiments, the prime editor comprises a Cas9 protein having an amino acid sequence according to SEQ ID NO: 1009, which does not include an N-terminal methionine. In some embodiments, the wild-type SluCas9 comprises the amino acid sequence set forth in SEQ ID NO: 1009. In some embodiments, the prime editor comprises a Cas9 protein that comprises one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to a corresponding wild-type Cas9 protein (e.g., wild-type SluCas9). In some embodiments, the Cas9 protein that comprises one or more mutations compared to a wild-type Cas9 protein comprises the amino acid sequence set forth in SEQ ID NO: 1010 or SEQ ID NO: 1011. Exemplary Staphylococcus lugdunensis Cas9 (SluCas9) amino acid sequences useful in the prime editors disclosed herein are shown in Table 1B.
[0383] In some embodiments, the prime editor comprises a Cas9 protein from Staphylococcus aureus (SaCas9), e.g., according to any of SEQ ID NOs: 1012-1014, or a variant thereof. In some embodiments, the prime editor comprises a Cas9 protein from Staphylococcus aureus (SaCas9), e.g., according to any one of SEQ ID NOs: 1012-1014, or a variant thereof. In some embodiments, the Cas9 protein is SaCas9. In some embodiments, the SaCas9 can be wild-type SaCas9, a SaCas9 variant, or a nickase SaCas9. In some embodiments, the SaCas9 lacks an N-terminal methionine compared to the corresponding SaCas9 (e.g., wild-type SaCas9, SaCas9 variant, or nickase SaCas9). In some embodiments, the prime editor comprises a Cas9 protein having an amino acid sequence according to SEQ ID NO: 1012, which does not include an N-terminal methionine. In some embodiments, the wild-type SaCas9 comprises the amino acid sequence set forth in SEQ ID NO: 1012. In some embodiments, the prime editor comprises a Cas9 protein that comprises one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to a corresponding wild-type Cas9 protein (e.g., wild-type SaCas9). In some embodiments, the Cas9 protein that comprises one or more mutations compared to a wild-type Cas9 protein comprises the amino acid sequence set forth in SEQ ID NO: 1013 or SEQ ID NO: 1014. Exemplary Staphylococcus aureus Cas9 (SaCas9) amino acid sequences useful in the prime editors disclosed herein are shown in Table 1B.
[0384] In some embodiments, the prime editor comprises a Cas9 protein according to any one of the sequences set forth in SEQ ID NOs: 1015-1023, 1030-1032, or a variant thereof. In some embodiments, the Cas9 protein is a Cas9 variant, such as an SpCas9 variant (e.g., SpCas9-NG, SpCas9-NGA, SpRY, or SpG). In some embodiments, the Cas9 protein lacks an N-terminal methionine compared to the corresponding Cas9 protein (e.g., a Cas9 variant set forth in any one of SEQ ID NOs: 1015, 1016, 1018, 1019, 1021, 1022, 1030, or 1031). In some embodiments, the prime editor comprises a Cas9 protein (e.g., a Cas9 variant) having an amino acid sequence according to any one of SEQ ID NOs: 1015, 1018, 1021, or 1030 that does not include an N-terminal methionine. In some embodiments, a prime editor comprises a Cas9 protein that comprises one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to a corresponding Cas9 protein (e.g., a Cas9 protein set forth in any one of SEQ ID NOs: 1015, 1018, 1021, or 1030). In some embodiments, the Cas9 protein that comprises one or more mutations compared to a corresponding Cas9 protein comprises the amino acid sequence set forth in any one of SEQ ID NOs: 1016, 1017, 1019, 1020, 1022, 1023, 1031, or 1032.
[0385] In some embodiments, the Cas9 protein is a chimeric Cas9, e.g., a modified Cas9, e.g., a synthetic RNA-guided nuclease (sRGN) (e.g., modified by DNA family shuffling), e.g., sRGN3.1, sRGN3.3. In some embodiments, the DNA family shuffling involves fragmenting and reassembling one or more parent Cas9 genes, e.g., Cas9s from Staphylococcus hyicus (Shy), Staphylococcus lugdunensis (Slu), Staphylococcus microti (Smi), and Staphylococcus pasteuri (Spa). In some embodiments, the modified SluCas9 exhibits increased editing efficiency and / or specificity compared to unmodified SluCas9. In some embodiments, the modified Cas9, e.g., sRGN, exhibits an increase in editing efficiency of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% compared to unmodified Cas9. In some embodiments, Cas9, e.g., sRGN, exhibits an increase in specificity of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% compared to unmodified Cas9.In some embodiments, Cas9, e.g., sRGN, exhibits at least a 10%, at least a 20%, at least a 30%, at least a 40%, at least a 50%, at least a 60%, at least a 70%, at least a 80%, at least a 90%, at least a 100%, at least a 150%, at least a 200%, at least a 300%, at least a 400%, at least a 500%, at least a 600%, at least a 700%, at least a 800%, at least a 900%, or at least a 1000% increase in cleavage activity compared to unmodified Cas9. In some embodiments, Cas9, e.g., sRGN, exhibits the ability to cleave 5'-NNGG-3' PAM-containing targets. In some embodiments, the prime editor comprises a Cas9 protein (e.g., a chimeric Cas9) according to any one of the sequences set forth in SEQ ID NOs: 1024-1029, or a variant thereof. Exemplary amino acid sequences of Cas9 proteins (e.g., sRGN) useful in the prime editors disclosed herein are set forth below in SEQ ID NOs: 1024-1029. In some embodiments, a prime editor comprises a Cas9 protein that lacks an N-terminal methionine compared to SEQ ID NO: 1024 or SEQ ID NO: 1027. In some embodiments, a prime editor comprises a Cas9 protein that includes one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to a corresponding Cas9 protein (e.g., a Cas9 protein set forth in SEQ ID NO: 1024 or SEQ ID NO: 1027). In some embodiments, a Cas9 protein that includes one or more mutations compared to a corresponding Cas9 protein comprises the amino acid sequence set forth in any one of SEQ ID NOs: 1025, 1026, 1028, or 1029.
[0386] In some embodiments, the Cas9 protein comprises a variant Cas9 protein containing one or more amino acid substitutions. In some embodiments, the wild-type Cas9 protein comprises a RuvC domain and an HNH domain. In some embodiments, the prime editor comprises a nuclease-active Cas9 protein capable of cleaving both strands of a double-stranded target DNA sequence. In some embodiments, the nuclease-active Cas9 protein comprises a functional RuvC domain and a functional HNH domain. In some embodiments, the prime editor comprises a Cas9 nickase that can bind to a guide polynucleotide and recognize the target DNA, but can cleave only one strand of the double-stranded target DNA. In some embodiments, the Cas9 nickase comprises only one functional RuvC domain or one functional HNH domain. In some embodiments, the prime editor comprises a Cas9 with a non-functional HNH domain and a functional RuvC domain. In some embodiments, the prime editor can cleave the edited strand (i.e., the PAM strand) but cannot cleave the non-edited strand of a double-stranded target DNA sequence. In some embodiments, the prime editor comprises a Cas9 with a non-functional RuvC domain that can cleave the target strand (i.e., the non-PAM strand) but cannot cleave the edited strand of a double-stranded target DNA sequence. In some embodiments, the prime editor comprises a Cas9 that does not have a functional RuvC domain or a functional HNH domain and is unable to cleave either strand of a double-stranded target DNA sequence.
[0387] In some embodiments, the prime editor comprises a Cas9 with a mutation in the RuvC domain that reduces or eliminates nuclease activity of the RuvC domain. In some embodiments, the Cas9 comprises a mutation at amino acid D10 compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005, or a corresponding mutation thereof. In some embodiments, the Cas9 comprises a D10A mutation compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises mutations at amino acids D10, G12, and / or G17 compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a D10A mutation, a G12A mutation, and / or a G17A mutation compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005, or a corresponding mutation thereof.
[0388] In some embodiments, the prime editor comprises a Cas9 polypeptide having a mutation in the HNH domain that reduces or eliminates nuclease activity of the HNH domain. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid H840, or a corresponding mutation thereof, compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid E762, D839, H840, N854, N856, N863, H982, H983, A984, D986, and / or A987, or a corresponding mutation thereof, compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005. In some embodiments, the Cas9 polypeptide comprises an E762A, D839A, H840A, N854A, N856A, N863A, H982A, H983A, A984A, and / or D986A mutation, or a corresponding mutation, compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid residues R221, N394, and / or H840, compared to the wild-type SpCas9 (e.g., SEQ ID NO: 1005). In some embodiments, the Cas9 polypeptide comprises an R221K, N394L, and / or H840A mutation, or a corresponding mutation, compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005. In some embodiments, the Cas9 polypeptide comprises mutations at amino acid residues R220, N393, and / or H839, or corresponding mutations, relative to wild-type SpCas9 lacking an N-terminal methionine (e.g., SEQ ID NO: 1005). In some embodiments, the Cas9 polypeptide comprises mutations at amino acid residues R220K, N393K, and / or H839A, or corresponding mutations, relative to wild-type SpCas9 lacking an N-terminal methionine (set forth in SEQ ID NO: 1005).
[0389] In some embodiments, the prime editor comprises a Cas9 with one or more amino acid substitutions in both the HNH domain and the RuvC domain that reduce or eliminate nuclease activity of both the HNH domain and the RuvC domain. In some embodiments, the prime editor comprises a nuclease-inactive Cas9 or a nuclease-dead Cas9 (dCas9). In some embodiments, the dCas9 comprises an H840X substitution and a D10X mutation, or a corresponding mutation, compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005, where X is any amino acid other than H in the case of the H840X substitution and any amino acid other than D in the case of the D10X substitution. In some embodiments, the dead Cas9 comprises an H840A and a D10A mutation, or a corresponding mutation, compared to the wild-type SpCas9 set forth in SEQ ID NO: 1005.
[0390] In some embodiments, the N-terminal methionine is removed from a Cas9 nickase or any Cas9 variant, ortholog, or equivalent disclosed or contemplated herein. For example, a methionine-minus (Met(-)) Cas9 nickase includes any one of the sequences set forth in SEQ ID NOs: 1007, 1008, 1011, 1014, 1017, 1020, 1023, 1026, 1029, 1032, or a variant thereof having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0391] In addition to dead Cas9 and Cas9 nickase variants, Cas9 proteins as used herein can also include other Cas9 variants having at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% sequence identity to any reference Cas9 protein, e.g., any wild-type Cas9, or mutant Cas9 (e.g., dead Cas9 or Cas9 nickase) disclosed herein or known in the art, or fragment Cas9, or circular permutant Cas9, or other variant of Cas9. In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to a reference Cas9, e.g., a wild-type Cas9. In some embodiments, a Cas9 variant comprises a fragment of a reference Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of a reference Cas9, e.g., a wild-type Cas9.In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9.
[0392] In some embodiments, the Cas9 fragment is a functional fragment that retains one or more Cas9 activities. In some embodiments, the Cas9 fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0393] Exemplary wild-type Cas proteins and corresponding nickases are shown in Table 1B below.
[0394] In some embodiments, the prime editor comprises a Cas protein, e.g., Cas9, containing a modification that enables altered PAM recognition. In prime editing using a Cas protein-based prime editor, the terms "protospacer adjacent motif (PAM)," PAM sequence, or PAM-like motif can be used to refer to a short DNA sequence immediately following the protospacer on the PAM strand of a target gene. In some embodiments, the PAM is recognized by a Cas nuclease within the prime editor during prime editing. In certain embodiments, the PAM is required for target binding of the Cas protein. The specific PAM sequence required for Cas protein recognition may vary depending on the specific type of Cas protein. The PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides in length. In some embodiments, the PAM is 2-6 nucleotides in length. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). In some embodiments, the Cas protein of the prime editor recognizes a canonical PAM, e.g., SpCas9 recognizes 5'-NGG-3' PAM. In some embodiments, the Cas protein of the prime editor has altered or non-canonical PAM specificity. Exemplary PAM sequences and corresponding Cas variants are listed in Table 1C below. It should be understood that for each variant provided, the Cas protein includes one or more of the indicated amino acid substitutions compared to a wild-type Cas protein sequence, e.g., Cas9 set forth in SEQ ID NO: 1005. The PAM motifs listed in Table 1C below are in 5' to 3' order. In some embodiments, the Cas proteins of the present disclosure can also be used to direct transcriptional regulation of target sequences, e.g., silencing transcription by sequence-specific binding to target sequences. In some embodiments, the Cas proteins described herein can have one or more mutations in the PAM recognition motif. In some embodiments, the Cas proteins described herein can have altered PAM specificity.
[0395] The nucleotides listed in Table 1C are represented by the base codes provided in the Handbook on Industrial Property Information and Documentation, World Intellectual Property Organization (WIPO) Standard ST.26, Version 1.4. For example, "R" in Table 1C represents nucleotides A or G, "W" in Table 1C represents A or T, "V" refers to any one of the nucleotides A, G, or C, and "N" refers to any one of the nucleotides A, G, C, or T. [Table 1B-1] [Table 1B-2] [Table 1B-3] [Table 1B-4] [Table 1B-5] [Table 1B-6] [Table 1B-7] [Table 1B-8] [Table 1B-9] [Table 1B-10] [Table 1B-11] [Table 1B-12] Table 1B-13 Table 1B-14 Table 1B-15 Table 1B-16 Table 1B-17 Table 1B-18 Table 1B-19 Table 1B-20 Table 1B-21 Table 1B-22 Table 1B-23 Table 1B-24 Table 1B-25 Table 1B-26 Table 1B-27 Table 1B-28 [Table 1C-1] [Table 1C-2]
[0396] In some embodiments, the prime editor has any of the following amino acids compared to the wild-type SpCas9 polypeptide set forth in SEQ ID NO: 1005: A61R, L111R, D1135V, R221K, A262T, R324L, N394K, S409I, S409I, E427G, E480K, M495V, N497A, Y515N, K526E, F539S, E543D, R654L, R661A, R661L, R691A, N692A, M694A, M694I, Q695A, H698A, R753G, M763I, K848A, K890N, Q926A, K1003A, R1060A, L1111R, R1114G, D1115V, D1116A, D1117A, D1118A, D1119A, D1120A, D1121B, D1122B, D1123B, D1124C, D1125C, D1126D, D1127E, D1128C, D1129A, D1130C, D1131D, D1132C, D1133C, D1134C, D1135D, D1136C, D1137C, D1138C, D1139D, D1140C, D1141D, D1142C, D1143C, D1144C, D1145C, D1146C, D1147C, D1148C, D1149C, D1150C, D1151C and any combination thereof.
[0397] In some embodiments, the prime editor comprises a SaCas9 polypeptide. In some embodiments, the SaCas9 polypeptide comprises one or more of the mutations E782K, N968K, and R1015H compared to wild-type SaCas9. In some embodiments, the prime editor comprises an FnCas9 polypeptide, such as a wild-type FnCas9 polypeptide, or an FnCas9 polypeptide comprising one or more of the mutations E1369R, E1449H, or R1556A compared to wild-type FnCas9. In some embodiments, the prime editor comprises Sc Cas9, such as wild-type ScCas9, or an ScCas9 polypeptide comprising one or more of the mutations I367K, G368D, I369K, H371L, T375S, T376G, and T1227K compared to wild-type ScCas9. In some embodiments, the prime editor comprises a St1 Cas9 polypeptide, a St3 Cas9 polypeptide, or a Slu Cas9 polypeptide.
[0398] In some embodiments, a prime editor comprises a Cas polypeptide comprising a circularly permuted Cas variant. For example, a Cas9 polypeptide of a prime editor can be designed such that the N- and C-termini of a Cas9 protein (e.g., a wild-type Cas9 protein or a Cas9 nickase) are locally rearranged to retain the ability to bind to DNA when complexed with a guide RNA (gRNA). An exemplary circularly permuted configuration can be N-terminus-[original C-terminus]-[original N-terminus]-C-terminus. Any of the Cas9 proteins described herein, including any variant, ortholog, or native Cas9, or equivalent thereof, can be reconstituted as a circularly permuted variant.
[0399] In various embodiments, a circularly permuted Cas protein, e.g., Cas9, can have the following structure: N-terminus-[original C-terminus]-[optional linker]-[original N-terminus]-C-terminus. In some embodiments, a circularly permuted Cas9 comprises any one of the following structures (amino acid positions set forth in SEQ ID NO: 1005):
[0400] N-terminus-[1268-1368]-[optional linker]-[1-1267]-C-terminus;
[0401] N-terminus-[1168-1368]-[optional linker]-[1-1167]-C-terminus;
[0402] N-terminus-[1068-1368]-[optional linker]-[1-1067]-C-terminus;
[0403] N-terminus-[968-1368]-[optional linker]-[1-967]-C-terminus;
[0404] N-terminus-[868-1368]-[optional linker]-[1-867]-C-terminus;
[0405] N-terminus-[768-1368]-[optional linker]-[1-767]-C-terminus;
[0406] N-terminus-[668-1368]-[optional linker]-[1-667]-C-terminus;
[0407] N-terminus-[568-1368]-[optional linker]-[1-567]-C-terminus;
[0408] N-terminus-[468-1368]-[optional linker]-[1-467]-C-terminus;
[0409] N-terminus-[368-1368]-[optional linker]-[1-367]-C-terminus;
[0410] N-terminus-[268-1368]-[optional linker]-[1-267]-C-terminus;
[0411] N-terminus-[168-1368]-[optional linker]-[1-167]-C-terminus;
[0412] N-terminus-[68-1368]-[optional linker]-[1-67]-C-terminus;
[0413] N-terminus-[10-1368]-[any linker]-[1-9]-C-terminus; or corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0414] In some embodiments, the circularly permuted Cas9 comprises any one of the following structures (amino acid positions set forth in SEQ ID NO: 1005):
[0415] N-terminus-[102-1368]-[optional linker]-[1-101]-C-terminus;
[0416] N-terminus-[1028-1368]-[optional linker]-[1-1027]-C-terminus;
[0417] N-terminus-[1041-1368]-[optional linker]-[1-1043]-C-terminus;
[0418] N-terminus-[1249-1368]-[optional linker]-[1-1248]-C-terminus;
[0419] N-terminus-[1300-1368]-[any linker]-[1-1299]-C-terminus; or corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0420] In some embodiments, the circularly permuted Cas9 comprises any one of the following structures (amino acid positions set forth in SEQ ID NO: 1005):
[0421] N-terminus-[103-1368]-[optional linker]-[1-102]-C-terminus:
[0422] N-terminus-[1029-1368]-[optional linker]-[1-1028]-C-terminus;
[0423] N-terminus-[1042-1368]-[optional linker]-[1-1041]-C-terminus;
[0424] N-terminus-[1250-1368]-[optional linker]-[1-1249]-C-terminus;
[0425] N-terminus-[1301-1368]-[any linker]-[1-1300]-C-terminus; or corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0426] In some embodiments, circular permutants can be formed by joining a C-terminal fragment of Cas9 to an N-terminal fragment of Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, the C-terminal fragment can correspond to 95% or more of the C-terminal amino acids of Cas9 (e.g., about amino acids 1300-1368 set forth in SEQ ID NO: 1005, or corresponding amino acid positions thereof), or 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the C-terminal amino acids of Cas9 (e.g., SEQ ID NO: 1005, or an ortholog or variant thereof). The N-terminal portion can correspond to 95% or more of the N-terminal amino acids of Cas9 (e.g., about amino acids 1-1300 set forth in SEQ ID NO: 1005, or corresponding amino acid positions thereof), or 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the N-terminal amino acids of Cas9 (e.g., set forth in SEQ ID NO: 1005, or corresponding amino acid positions thereof).
[0427] In some embodiments, circular permutants can be formed by joining a C-terminal fragment of Cas9 to an N-terminal fragment of Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, the N-terminally rearranged C-terminal fragment comprises or corresponds to the C-terminal 30% or less of the amino acids of Cas9 (e.g., amino acids 1012-1368 set forth in SEQ ID NO: 1005, or their corresponding amino acid positions). In some embodiments, the N-terminally rearranged C-terminal fragment comprises or corresponds to the C-terminus of 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the amino acids of Cas9 (e.g., as set forth in SEQ ID NO: 1005 or corresponding amino acid positions thereof). In some embodiments, the N-terminally rearranged C-terminal fragment comprises or corresponds to the C-terminus of 410 or fewer residues of Cas9 (e.g., as set forth in SEQ ID NO: 1005 or corresponding amino acid positions thereof). In some embodiments, the N-terminally rearranged C-terminal portion comprises or corresponds to 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 residues C-terminal of Cas9 (e.g., those set forth in SEQ ID NO: 1005, or corresponding amino acid positions thereof). In some embodiments, the N-terminally rearranged C-terminal portion comprises or corresponds to the 357, 341, 328, 120, or 69 residue C-terminus of Cas9 (e.g., as set forth in SEQ ID NO: 1005, or its corresponding amino acid positions).
[0428] In other embodiments, the circularly permuted Cas9 variant can be a topological rearrangement of the Cas9 primary structure based on the S. pyogenes Cas9 of SEQ ID NO: 1005, based on the following method: (a) selecting a circularly permuted (CP) site corresponding to an internal amino acid residue in the Cas9 primary structure, thereby dividing the original protein into two halves, the N-terminal and C-terminal regions; and (b) modifying the Cas9 protein sequence by (e.g., by genetic engineering techniques) moving the original C-terminal region (containing the CP site amino acids) to the front of the original N-terminal region, thereby forming a new N-terminus of the Cas9 protein starting from the CP site amino acid residue. The CP site can be located in any domain of the Cas9 protein, including, for example, the helical II domain, the RuvCIII domain, or the CTD domain. For example, the CP site can be located at the original amino acid residue (as set forth in SEQ ID NO: 1005, or its corresponding amino acid position) 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282. Thus, when relocated to the N-terminus, the original amino acid 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 becomes the new N-terminal amino acid. These CP-Cas9 proteins may be referred to as Cas9-CP181, Cas9-CP199, Cas9-CP230, Cas9-CP270, Cas9-CP310, Cas9-CP1010, Cas9-CP1016, Cas9-CP1023, Cas9-CP1029, Cas9-CP1041, Cas9-CP1247, Cas9-CP1249, and Cas9-CP1282, respectively. This description is not intended to be limited to the creation of CP variants from SEQ ID NO: 1005, but may be employed to create CP variants of any Cas9 sequence, either at the CP sites corresponding to these positions or at other CP sites throughout. This description is not intended to be limited in any way to a particular CP site. Virtually any CP site can be used to form CP-Cas9 variants.
[0429] In some embodiments, the prime editor comprises a Cas9 functional variant that has a smaller molecular weight than the wild-type SpCas9 protein. In some embodiments, the smaller Cas9 functional variant may facilitate delivery to cells, for example, by an expression vector, nanoparticle, or other delivery means. In certain embodiments, the smaller Cas9 functional variant is a Class 2 Type II Cas protein. In certain embodiments, the smaller Cas9 functional variant is a Class 2 Type V Cas protein. In certain embodiments, the smaller Cas9 functional variant is a Class 2 Type VI Cas protein.
[0430] In some embodiments, the prime editor comprises an SpCas9 that is 1368 amino acids in length and has a predicted molecular weight of 158 kilodaltons. In some embodiments, the prime editor comprises an SpCas9 that is less than 1300 amino acids, less than 1290 amino acids, less than 1280 amino acids, less than 1270 amino acids, less than 1260 amino acids, less than 1250 amino acids, less than 1240 amino acids, less than 1230 amino acids, less than 1220 amino acids, less than 1210 amino acids, less than 1200 amino acids, less than 1190 amino acids, less than 1180 amino acids, less than 1170 amino acids, less than 1160 amino acids, less than 1150 amino acids, less than 1140 amino acids, less than 1130 amino acids, less than 1120 amino acids, This includes Cas9 functional variants or fragments that are less than 1110 amino acids, less than 1100 amino acids, less than 1050 amino acids, less than 1000 amino acids, less than 950 amino acids, less than 900 amino acids, less than 850 amino acids, less than 800 amino acids, less than 750 amino acids, less than 700 amino acids, less than 650 amino acids, less than 600 amino acids, less than 550 amino acids, or less than 500 amino acids, but are larger than at least about 400 amino acids and retain one or more functions of the Cas9 protein, such as DNA binding.
[0431] In some embodiments, the Cas protein is Cas12a, Cas12b1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Csn2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csn1, Csn2, Csm3, Csm4, Csm5, Csm6, Csn1, Csn2, Csn2, Csm3, Csm4, Csm5, Csm6, Csn1, Csn2, Csn2, Csn3, Csm4, Csm5, Csm6, Csn1, Csn2, Csn3, Csn4, Csn5 ...3, Csn4, Csn5, Csn6, Csn1, Csn2, Csn3, Csn4, Csn5, Csn6, Csn1, Csn3, Csn4, Csn5, Csn6, Csn1, Csn2, Csn3, Csn4, Csn5, The CRISPR-associated protein may include any CRISPR-associated protein, including, but not limited to, mr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, preferably containing a nickase mutation (e.g., a mutation corresponding to the D10A mutation in the wild-type SpCas9 polypeptide of SEQ ID NO: 1005). In various other embodiments, the polypeptide domain having DNA-binding activity can be any of the following proteins: Cas9, Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), GeoCas9, CjCas9, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, circularly permuted Cas9, or an Argonaute (Ago) domain, or a functional variant or fragment thereof. Exemplary Cas proteins and nomenclature are shown in Table 2 below. [Table 2]
[0432] In some embodiments, the prime editors described herein may also include a Cas protein other than Cas9. For example, in some embodiments, the prime editors described herein may include a Cas12a (Cpf1) polypeptide or a functional variant thereof. In some embodiments, the Cas12a polypeptide comprises a mutation that reduces or eliminates the endonuclease domain of the Cas12a polypeptide. In some embodiments, the Cas12a polypeptide is a Cas12a nickase. In some embodiments, the Cas protein comprises an amino acid sequence that comprises at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Cas12a polypeptide.
[0433] In some embodiments, the prime editor comprises a Cas protein that is a Cas12b (C2c1) or Cas12c (C2c3) polypeptide. In some embodiments, the Cas protein comprises an amino acid sequence that comprises at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Cas12b (C2c1) or Cas12c (C2c3) protein. In some embodiments, the Cas protein is a Cas12b nickase or a Cas12c nickase. In some embodiments, the Cas protein is a Cas12e, Cas12d, Cas13, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or CasΦ polypeptide. In some embodiments, the Cas protein comprises an amino acid sequence comprising at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Cas12e, Cas12d, Cas13, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or CasΦ protein. In some embodiments, the Cas protein is Cas12e, Cas12d, Cas13, or CasΦ nickase.
[0434] nuclear localization sequence In some embodiments, the prime editor further comprises one or more nuclear localization sequences (NLSs). In some embodiments, the NLSs facilitate translocation of the protein to the cell nucleus. In some embodiments, the prime editor comprises a fusion protein, for example, a fusion protein comprising a DNA-binding domain and a DNA polymerase comprising one or more NLSs. In some embodiments, one or more polypeptides of the prime editor are fused or linked to one or more NLSs. In some embodiments, the prime editor comprises a DNA-binding domain and a DNA polymerase domain provided in trans, wherein the DNA-binding domain and / or the DNA polymerase domain are fused or linked to one or more NLSs.
[0435] In certain embodiments, the prime editor or prime editing composition comprises at least one NLS. In some embodiments, the prime editor or prime editing composition comprises at least two NLSs. In some embodiments, the prime editor or prime editing complex comprises at least three NLSs. In some embodiments, the prime editor or prime editing complex comprises 4, 5, 6, 7, 8, 9, or more than 10 NLSs. In embodiments with at least two NLSs, the NLSs may be the same or different NLSs. In some embodiments, one or more NLSs of the prime editor comprise a bipartite NLS.
[0436] An NLS can be expressed as part of a prime editor or prime editing composition. In some embodiments, the NLS can be located at almost any position in the amino acid sequence of the protein, and generally comprises a short sequence of three or more or four or more amino acids. The position of the NLS fusion may be at the N-terminus, the C-terminus, or anywhere within the sequence of the prime editor or its component (e.g., inserted between the DNA-binding domain and DNA polymerase domain of a prime editor fusion protein, between the DNA-binding domain and a linker sequence, between the DNA polymerase and a linker sequence, or between two linker sequences of a prime editor fusion protein or its component, in either N- to C- or C-terminal order). In some embodiments, the prime editor is a fusion protein comprising an NLS at the N-terminus. In some embodiments, the prime editor is a fusion protein comprising an NLS at the C-terminus. In some embodiments, the prime editor is a fusion protein comprising at least one NLS at both the N- and C-termini. In some embodiments, the prime editor is a fusion protein comprising two NLSs at the N- and / or C-termini.
[0437] Any NLS known in the art is contemplated herein. The NLS can be any naturally occurring NLS or any non-naturally occurring NLS (e.g., an NLS with one or more mutations compared to the wild-type NLS). In some embodiments, the nuclear localization signal (NLS) is predominantly basic. In some embodiments, one or more NLSs of the prime editor are rich in lysine and arginine residues. In some embodiments, one or more NLSs of the prime editor comprise a proline residue.
[0438] Non-limiting examples of NLS sequences suitable for use with the methods and compositions of the present disclosure are provided in Table 3A.
[0439] In some embodiments, the NLS is a monopartite NLS. For example, in some embodiments, the NLS is an SV40 large T antigen NLS comprising the sequence of SEQ ID NO: 1159. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the bipartite NLS comprises two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the bipartite NLS consists of two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, the amino acid sequence of the spacer sequence comprises Xenopus nucleoplasmic NLS SEQ ID NO: 1160, where X is any amino acid. In some embodiments, the NLS comprises the nucleoplasmic NLS sequence SEQ ID NO: 1161. In some embodiments, the NLS is a non-canonical sequence, such as M9 of the hnRNP A1 protein, influenza virus nucleoprotein NLS, and yeast Gal4 protein NLS.
[0440] In some embodiments, the NLS comprises an amino acid sequence at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence provided in Table 3A. In some embodiments, the NLS comprises an amino acid sequence selected from the group consisting of the amino acid sequences provided in Table 3A. In some embodiments, the prime editing composition comprises a polynucleotide encoding an NLS comprising an amino acid sequence at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to an amino acid sequence provided in Table 3A. In some embodiments, the prime editing composition comprises a polynucleotide encoding an NLS comprising an amino acid sequence provided in Table 3A. [Table 3A]
[0441] Additional Prime Editor Components The prime editors described herein may contain additional functional domains, such as one or more domains that modify the folding, solubility, or charge of the prime editor. In some cases, the prime editor may contain a solubility-enhancing (SET) domain.
[0442] In some embodiments, a split intein comprises two halves of an intein protein, which may be referred to as the N-terminal half of the intein, or intein N, and the C-terminal half of the intein, or intein C, respectively. In some embodiments, intein N and intein C may be fused to protein domains (N-terminal and C-terminal exteins), respectively. Exteins can be any protein or polypeptide, for example, any prime editor polypeptide component. In some embodiments, intein N and intein C of a split intein non-covalently associate to form an active intein and can catalyze a trans-splicing reaction. In some embodiments, the trans-splicing reaction excises the two intein sequences and links the two extein sequences with a peptide bond. As a result, intein N and intein C are spliced together, and the protein domain linked to intein N is fused to the protein domain linked to intein C in essentially the same manner as a continuous intein. In some embodiments, the split intein is derived from a eukaryotic intein, a bacterial intein, or an archaeal intein. Preferably, the resulting split intein will have only the amino acid sequence essential for catalyzing a trans-splicing reaction. In some embodiments, intein N or intein C further comprises one or more amino acid substitutions compared to wild-type intein N or wild-type intein C, e.g., amino acid substitutions that enhance the trans-splicing activity of the split intein. In some embodiments, intein C comprises 4 to 7 consecutive amino acid residues, at least 4 of which are derived from the last β-strand of the intein from which it is derived. In some embodiments, the split intein is derived from the Ssp DnaE intein (e.g., Synechocytis sp. PCC6803), or any intein or split intein known in the art, or a functional variant or fragment thereof.
[0443] In some embodiments, the prime editor comprises one or more epitope tags. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, thioredoxin (Trx) tags, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, polyhistidine tags (also called histidine tags or His tags), maltose binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag1, Softag3), strep tags, biotin ligase tags, FlAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein comprises one or more His tags.
[0444] In some embodiments, the prime editor comprises one or more polypeptide domains encoded by one or more reporter genes. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, and autofluorescent proteins such as green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and blue fluorescent protein (BFP).
[0445] In some embodiments, the prime editor comprises one or more polypeptide domains that bind to DNA molecules or other cellular molecules. Examples of binding proteins or domains include, but are not limited to, maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions.
[0446] Polypeptides comprising the components of a prime editor can be fused via a peptide linker or provided in trans relative to each other. For example, a reverse transcriptase can be expressed, delivered, or provided as a separate component rather than as part of a fusion protein with a DNA-binding domain. In such cases, the components of the prime editor can be associated through a non-peptide bond or colocalization function. In some embodiments, the prime editor further comprises additional components capable of interacting with...
Claims
1. A prime editing system comprising: (A) a first prime editing guide RNA (PEGRNA), or one or more polynucleotides encoding the first PEGRNA; and (B) a second PEGRNA, or one or more polynucleotides encoding the second PEGRNA; The first PEG-RNA comprises: (i) a first spacer that is complementary to a first interrogation target sequence on the first strand of the TRAC gene; (ii) a first gRNA core capable of binding to a Cas9 protein; and (iii) a first extension arm comprising a first editing template and a first primer binding site (PBS); the first spacer comprises at its 3' end nucleotides 4 to 20 of a sequence selected from the group consisting of SEQ ID NOs: 4, 61, 88, and 150, and the first PBS comprises at its 5' end a sequence which is the reverse complement of nucleotides 13 to 17 of the selected sequence; wherein the second PEG RNA is (i) a second spacer complementary to a second interrogation target sequence on a second strand of the TRAC gene that is complementary to the first strand; (ii) a second gRNA core capable of binding to a Cas9 protein; and (iii) a second extension arm comprising a second editing template and a second PBS; the second spacer comprises at its 3' end nucleotides 4 to 20 of a sequence selected from the group consisting of SEQ ID NOs: 177, 233, 260, 287, 314, 341, 368, 414, 441, 468, 495, 522, and 566, and the second PBS comprises at its 5' end a sequence that is the reverse complement of nucleotides 13 to 17 of the selected sequence; where: (a) the first editing template includes a region of complementarity to the second editing template; (b) the first editing template comprises nucleotides 8-17 of the selected sequence relative to the second spacer, and the second editing template comprises nucleotides 8-17 of the selected sequence relative to the first spacer; or (c) the first editing template comprises nucleotides 8 to 17 of the selected sequence relative to the second spacer and a region of complementarity to the second editing template, and the second editing template comprises nucleotides 8 to 17 of the selected sequence relative to the first spacer and a region of complementarity to the first editing template.
2. 2. The prime editing system of claim 1, wherein the selected sequence of the first spacer is SEQ ID NO: 4 or 88.
3. 3. The prime editing system of claim 2, wherein the selected sequence of the first spacer is SEQ ID NO:
88.
4. 3. The prime editing system of claim 1 or 2, wherein the selected sequence of the second spacer is SEQ ID NO: 177, 368, or 522.
5. 5. The prime editing system of claim 4, wherein the selected sequence of the second spacer is SEQ ID NO:
177.
6. The prime editing system of any one of claims 1 to 5, wherein the first spacer and / or the second spacer is 16 to 22 nucleotides in length.
7. The prime editing system of any one of claims 1 to 6, wherein the first spacer and / or the second spacer is 20 nucleotides in length and comprises the selected sequence.
8. 8. The prime editing system of claim 1, wherein the first PBS is 8 to 17 nucleotides in length and comprises at its 5' end a sequence that is the reverse complement of nucleotides 10 to 17, 9 to 17, 8 to 17, 7 to 17, 6 to 17, 5 to 17, 4 to 17, 3 to 17, 2 to 17, or 1 to 17 of the selected sequence of the first spacer.
9. The prime editing system of claim 8, wherein the first PBS is 8 to 13 nucleotides in length.
10. 10. The prime editing system of claim 9, wherein the first PBS is 10, 11, or 12 nucleotides in length.
11. 11. The prime editing system of any one of claims 1 to 10, wherein the second PBS is 7 to 17 nucleotides in length and comprises a sequence at its 5' end that is the reverse complement of nucleotides 11 to 17, 10 to 17, 9 to 17, 8 to 17, 7 to 17, 6 to 17, 5 to 17, 4 to 17, 3 to 17, 2 to 17, or 1 to 17 of the selected sequence of the second spacer.
12. The prime editing system of claim 11, wherein the second PBS is 8 to 13 nucleotides in length.
13. 12. The prime editing system of claim 11, wherein the second PBS is 11, 12, or 13 nucleotides in length.
14. The prime editing system of any one of claims 1 to 13, wherein the first gRNA core and the second gRNA core comprise the same sequence.
15. 15. The prime editing system of Claim 14, wherein the first gRNA core, the second gRNA core, or both comprise SEQ ID NO:
590.
16. The prime editing system of any one of claims 1 to 15, wherein the first spacer, the first gRNA core, the first editing template, and the first PBS form a continuous sequence within a single molecule.
17. 17. The prime editing system of claim 16, wherein the first PEGRNA comprises, from 5' to 3', the first spacer, the first gRNA core, the first editing template, and the first PBS.
18. The prime editing system of any one of claims 1 to 17, wherein the second spacer, the second gRNA core, the second editing template, and the second PBS form a continuous sequence within a single molecule.
19. 19. The prime editing system of claim 18, wherein the second PEGRNA comprises, from 5' to 3', the second spacer, the second gRNA core, the second editing template, and the second PBS.
20. The prime editing system of any one of claims 1 to 19, wherein the first editing template comprises a region of complementarity to the second editing template.
21. 21. The prime editing system of claim 20, wherein the first editing template and the second editing template each encode all or a fragment of a recombinase recognition sequence (RSS) or its reverse complement, the first editing template encodes at least a 5' portion of the RSS or its reverse complement, the second editing template encodes at least a 3' portion of the RSS or its reverse complement, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
22. 22. The prime editing system of Claim 21, wherein at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other, and optionally at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
23. 23. The prime editing system of claim 21 or 22, wherein the first editing template encodes the RSS.
24. The prime editing system according to any one of claims 21 to 23, wherein the second editing template encodes the RSS.
25. The prime editing system of any one of claims 21 to 24, wherein the RSS is an attB sequence recognized by Bxb1 recombinase.
26. The prime editing system of any one of claims 21 to 24, wherein the RSS is an attP sequence recognized by Bxb1 recombinase.
27. The prime editing system of any one of claims 20 to 25, wherein the first editing template includes RTT number 1 from Table 6 and the second editing template includes RTT number 2 from the same RTT pair in Table 6, or the first editing template includes RTT number 2 from Table 6 and the second editing template includes RTT number 1 from the same RTT pair in Table 6.
28. The prime editing system of any one of claims 20 to 25, wherein the first editing template includes SEQ ID NO: 24 and the second editing template includes SEQ ID NO:
196.
29. The prime editing system of any one of claims 20 to 25, wherein the first editing template includes sequence number 23 and the second editing template includes sequence number 196.
30. The prime editing system of any one of claims 20 to 25, wherein the first editing template includes SEQ ID NO: 24 and the second editing template includes SEQ ID NO:
106.
31. The prime editing system of any one of claims 20 to 25, wherein the first editing template comprises SEQ ID NO: 27 and the second editing template comprises SEQ ID NO:
107.
32. The prime editing system of any one of claims 20 to 25, wherein the first editing template includes sequence number 106 and the second editing template includes sequence number 24.
33. The prime editing system of any one of claims 20 to 25, wherein the first editing template comprises sequence number 107 and the second editing template comprises sequence number 27.
34. The prime editing system of any one of claims 20 to 25, wherein the first editing template comprises SEQ ID NO: 25 and the second editing template comprises SEQ ID NO:
197.
35. The prime editing system of any one of claims 20 to 25, wherein the first editing template comprises SEQ ID NO: 28 and the second editing template comprises SEQ ID NO:
199.
36. The prime editing system of any one of claims 20 to 25, wherein the first editing template comprises SEQ ID NO: 22 and the second editing template comprises SEQ ID NO:
195.
37. The prime editing system of any one of claims 20 to 25, wherein the first editing template includes sequence number 26 and the second editing template includes sequence number 198.
38. 26. The prime editing system of any one of claims 20 to 25, wherein the first editing template comprises a 5' fragment of an RTT listed in Table 6, and the second editing template comprises a full-length or 5' fragment of a corresponding RTT pair, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
39. 26. The prime editing system of any one of claims 20 to 25, wherein the second editing template comprises a 5' fragment of an RTT listed in Table 6, the first editing template comprises a full-length or 5' fragment of a corresponding RTT pair, and at least 10 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
40. 40. The prime editing system of Claim 38 or 39, wherein at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other, and optionally at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides at the 5' ends of the first and second editing templates have perfect reverse complementarity to each other.
41. 40. The prime editing system of Claim 38 or 39, wherein the length of the region of complementarity of the first editing template is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90% of the length of the first editing template; optionally, the length of the region of complementarity of the first editing template is at least 52%, at least 53%, or at least 55% of the length of the first editing template.
42. 40. The prime editing system of Claim 38 or 39, wherein the length of the region of complementarity of the second editing template is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90% of the length of the second editing template; optionally, the length of the region of complementarity of the second editing template is at least 52%, at least 53%, or at least 55% of the length of the second editing template.
43. (a) the first spacer comprises SEQ ID NO:4 and the first PBS comprises SEQ ID NO:13; or the first spacer comprises SEQ ID NO:88 and the first PBS comprises SEQ ID NO:96; or the first spacer comprises SEQ ID NO:88 and the first PBS comprises SEQ ID NO:157; or the first spacer comprises SEQ ID NO:88 and the first PBS comprises SEQ ID NO:99; or the first spacer comprises SEQ ID NO: 88 and the first PBS comprises SEQ ID NO: 97; and (b) the second spacer comprises SEQ ID NO: 177 and the second PBS comprises SEQ ID NO: 188; or The prime editing system of any one of claims 20 to 42, wherein the second spacer comprises SEQ ID NO: 368 and the second PBS comprises SEQ ID NO:
376.
44. (a) the first spacer comprises SEQ ID NO:4 and the first PBS has the sequence of SEQ ID NO:14 or SEQ ID NO:15; or the first spacer comprises SEQ ID NO: 88, and the first PBS has the sequence of SEQ ID NO: 96 or SEQ ID NO: 98; and (b) the second spacer comprises SEQ ID NO: 177 and the second PBS has the sequence of SEQ ID NO: 186 or SEQ ID NO: 188; the second spacer comprises SEQ ID NO: 368 and the second PBS has the sequence of SEQ ID NO: 373 or SEQ ID NO: 374; or The prime editing system of any one of claims 20 to 42, wherein the second spacer comprises SEQ ID NO: 522 and the second PBS has the sequence of SEQ ID NO: 531 or 533.
45. 45. The prime editing system of Claim 44, wherein the first spacer comprises SEQ ID NO:88, the first PBS has SEQ ID NO:96, the second spacer comprises SEQ ID NO:177, and the second PBS has the sequence of SEQ ID NO:
186.
46. 46. The prime editing system of claim 44 or 45, wherein the first editing template comprises SEQ ID NO: 27 and the second editing template comprises SEQ ID NO:
107.
47. 46. The prime editing system of claim 44 or 45, wherein the first editing template comprises SEQ ID NO: 107 and the second editing template comprises SEQ ID NO:
27.
48. 21. The prime editing system of claim 20, wherein the first PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 36, 118, 123, and 1132-1134, and the second PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 1135-1142.
49. 21. The prime editing system of Claim 20, wherein the first PEGRNA comprises SEQ ID NO: 118 and the second PEGRNA comprises SEQ ID NO: 1140.
50. 21. The prime editing system of Claim 20, wherein the first PEGRNA comprises SEQ ID NO: 136 and the second PEGRNA comprises SEQ ID NO:
224.
51. 21. The prime editing system of Claim 20, wherein the first PEGRNA comprises SEQ ID NO: 51 and the second PEGRNA comprises SEQ ID NO:
556.
52. the first PEGRNA comprises SEQ ID NO: 126; 21. The prime editing system of claim 20, wherein the second PEGRNA comprises SEQ ID NO:
220.
53. the first PEGRNA comprises SEQ ID NO: 44; 21. The prime editing system of Claim 20, wherein the second PEGRNA comprises SEQ ID NO:
550.
54. the first PEGRNA comprises SEQ ID NO: 111; 21. The prime editing system of Claim 20, wherein the second PEGRNA comprises SEQ ID NO: 1251.
55. the first PEGRNA is selected from the group consisting of SEQ ID NOs: 30, 31, 32, 33, 34, 36, 38, 39, 43, 44, 49, 50, 51, 52, 53, 79, 80, 82, 109, 111, 112, 114, 115, 118, 120, 122, 123, 125, 126, 129, 130, 134, 135, 136, 140, 141, 143, 168 , 169, 171, 592, 593, 594, and 1132, and wherein the second PEG RNA comprises a sequence selected from the group consisting of SEQ ID NOs: 203, 207, 210, 211, 212, 215, 217, 219, 220, 221, 222, 224, 226, 228, 251, 252, 254, 278, 279, 281, 305, 306, 308, 332, 333, 335, 359, 360, 362, 388, 390, 392, 395, 398, 397, 400, 401, 403, 404, 405, 408, 410, 432, 433, 435, 459, 460, 462, 486, 487, 489, 513, 514, 516, 541, 542, 543, 545, 546, 54 21. The prime editing system of claim 20, comprising a sequence selected from the group consisting of: 7, 549, 550, 552, 554, 555, 556, 558, 561, 562, 584, 585, 587, 591, 595, 597, 599, 601, 1127, 1128, 1135, 1136, 1137, 1138, 1139, 1140, and 1141.
56. 21. The prime editing system of Claim 20, wherein the first PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 30, 44, 109, and 126, and the second PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 207, 221, 388, and 400.
57. 21. The prime editing system of Claim 20, wherein the first PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 136, 141, 51, and 53, and the second PEGRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 226, 403, 558, 556, and 224.
58. 58. The prime editing system of Claim 57, wherein the first PEGRNA comprises SEQ ID NO: 136 and the second PEGRNA comprises SEQ ID NO:
224.
59. The prime editing system of any one of claims 1 to 58, wherein the first PEG-RNA and / or the second PEG-RNA further comprise a 3' motif, and optionally, the 3' motif is connected to the 3' end of the first PBS or the second PBS via a linker.
60. 60. The prime editing system of any one of claims 1 to 59, wherein the first PEGRNA and / or the second PEGRNA further comprise 5'mN*mN*mN* and 3'mN*mN*mN*N modifications, where m indicates that the nucleotide comprises a 2'-O-Me modification and a* indicates the presence of a phosphorothioate bond.
61. 61. The prime editing system of any one of claims 1 to 60, further comprising a prime editor, or one or more polynucleotides encoding the prime editor, wherein the prime editor comprises (a) a Cas9 nickase having a nuclease-inactivating mutation in its HNH domain, and (b) a reverse transcriptase.
62. 62. The prime editing system of Claim 61, wherein the Cas9 nickase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 1007.
63. 63. The prime editing system of Claim 61 or 62, wherein the reverse transcriptase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 1003.
64. The prime editing system of any one of claims 61 to 63, wherein the prime editor is a fusion protein.
65. 65. The prime editing system of Claim 64, wherein the fusion protein comprises the sequence of SEQ ID NO: 1033.
66. 66. The prime editing system of any one of Claims 60 to 65, wherein the one or more polynucleotides encoding the prime editor comprise (a) a first sequence encoding an N-terminal portion of the Cas9 nickase and intein N, and (b) a second sequence encoding intein C, a C-terminal portion of the Cas9 nickase, and the reverse transcriptase.
67. The prime editing system of any one of claims 21 to 66, further comprising a recombinase that recognizes the one or more RSSs, or one or more polynucleotides encoding the recombinase.
68. 68. The prime editing system of Claim 67, wherein the recombinase is Bxb1 comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 1131.
69. 69. The prime editing system of any one of Claims 62, 63, or 68, wherein the sequence identity is determined by Needleman-Wunsch alignment of the two protein sequences with gap costs set to 11 for presence and 1 for extension, where percent identity is calculated by dividing the number of identities by the length of the alignment.
70. 69. The prime editing system of Claim 67 or 68, wherein the recombinase is fused or linked to the prime editor.
71. The prime editing system of any one of claims 67 to 69, further comprising a DNA polynucleotide comprising (a) a donor sequence and (b) a second RSS recognized by the recombinase.
72. 72. The prime editing system of claim 71, wherein (i) the RSS comprises a Bxb1 attB sequence provided in Table 5 and the second RSS comprises a corresponding attP sequence provided in Table 5, or (ii) the RSS comprises a Bxb1 attP sequence provided in Table 5 and the second RSS comprises a corresponding attB sequence provided in Table 5.
73. 72. The prime editing system of claim 71, wherein the RSS sequence comprises SEQ ID NO: 1187 and the second RSS comprises SEQ ID NO: 1188.
74. 72. The prime editing system of Claim 71, wherein the donor sequence comprises an open reading frame encoding a polypeptide.
75. 75. The prime editing system of Claim 74, wherein the donor sequence encodes a chimeric antigen receptor (CAR).
76. 75. The prime editing system of Claim 74, wherein the donor sequence encodes a CD19CAR.
77. 75. The prime editing system of Claim 74, wherein the donor sequence comprises a splice acceptor sequence.
78. 67. The prime editing system of any one of Claims 61 to 66, comprising one or more vectors comprising the one or more polynucleotides encoding the first PEGRNA, the one or more polynucleotides encoding the second PEGRNA, and the one or more polynucleotides encoding the prime editor.
79. 67. The prime editing system of any one of Claims 61 to 66, comprising one or more vectors comprising the one or more polynucleotides encoding the first PEGRNA, the one or more polynucleotides encoding the second PEGRNA, the one or more polynucleotides encoding the prime editor, the one or more polynucleotides encoding the recombinase, and the donor sequence.
80. 80. The prime editing system of Claim 79, wherein the one or more vectors are AAV vectors.
81. The prime editing system of any one of Claims 61 to 79, wherein the one or more polynucleotides encoding the prime editor and / or the one or more polynucleotides encoding the recombinase are mRNA.
82. An LNP comprising the prime editing system described in any one of claims 1 to 81.
83. A pharmaceutical composition comprising the prime editing system of any one of claims 1 to 81 or the LNP of claim 82, and a pharmaceutically acceptable excipient.
84. 67. A method of editing a TRAC gene, the method comprising contacting the TRAC gene with (a) a prime editing system of any one of claims 1-60, and a prime editor comprising a Cas9 nickase with a nuclease-inactivating mutation in its HNH domain, and a reverse transcriptase; or (b) the prime editing system of any one of claims 61-66.
85. 85. The method of claim 84, further comprising contacting the TRAC gene with a recombinase or one or more polynucleotides encoding the recombinase, and a DNA polynucleotide comprising (a) a donor sequence and (b) one or more recombinase recognition sequences recognized by the recombinase.
86. 82. A method for inserting a donor sequence into a TRAC gene, the method comprising contacting the TRAC gene with a prime editing system described in any one of claims 1 to 81, or an LNP described in claim 82.
87. 82. A method for producing a modified cell, the method comprising contacting a cell with a prime editing system described in any one of claims 1 to 81, or an LNP described in claim 82.
88. 87. The method of claim 86, wherein the TRAC gene is in a cell.
89. 89. The method of claim 87 or 88, wherein the cell is a mammalian cell.
90. 89. The method of claim 87 or 88, wherein the cell is a human cell.
91. 91. The method of any one of claims 88 to 90, wherein the cell is an immune cell, optionally wherein the cell is a T cell.
92. 92. The method of any one of claims 88 to 91, wherein the cell is in a subject.
93. 92. The method of any one of claims 88 to 91, wherein the cells are derived from a subject.
94. 94. The method of claim 91 or 93, wherein the subject is a human.
95. A cell produced by the method of any one of claims 87 to 94.
96. A T cell comprising an edited TRAC gene comprising the sequence GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAAC C (SEQ ID NO: 1046) and / or GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 1047) compared to a wild-type TRAC gene.
97. 97. The T cell of claim 96, wherein the edited TRAC gene comprises, from 5' to 3', an insertion sequence comprising GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAAC C (SEQ ID NO: 1046), a donor sequence, and GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 1047).
98. 97. The T cell of claim 96, wherein the edited TRAC gene comprises, from 5' to 3', an insert sequence comprising GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 1047), a donor sequence, and GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAAC C (SEQ ID NO: 1046).
99. 99. The T cell of claim 97 or 98, wherein the donor sequences encode a chimeric antigen receptor (CAR), and optionally, the donor encodes a CD19 CAR.
100. 99. The T cell of claim 97 or 98, wherein the insertion sequence is between a first and a second chromosomal location, wherein the first chromosomal location is selected from the group consisting of positions 22547458, 22547457, 22547449, and 22547448 on human chromosome 14, and the second chromosomal location is selected from the group consisting of positions 22547533, 22547523, 22547491, 22547528, 22547497, 22547579, 22547522, 22547485, 22547506, 22547560, 22547505, 22547529, and 22547490 on human chromosome 14.
101. 101. The T cell of any one of claims 97-100, wherein the insertion sequence is located on human chromosome 14 between positions 22547458 and 22547533.
102. 101. The T cell of any one of claims 97-100, wherein the insertion sequence is located on human chromosome 14 between positions 22547458 and 22547522.
103. 101. The T cell of any one of claims 97-100, wherein the insertion sequence is located on human chromosome 14 between positions 22547458 and 22547529.
104. 101. The T cell of any one of claims 97-100, wherein the insertion sequence is located on human chromosome 14 between positions 22547449 and 22547533.
105. 101. The T cell of any one of claims 97-100, wherein the insertion sequence is located on human chromosome 14 between positions 22547449 and 22547522.
106. 101. The T cell of any one of claims 97-100, wherein the insertion sequence is located on human chromosome 14 between positions 22547449 and 22547529.
107. 107. The T cell of any one of claims 100 to 106, wherein the human chromosomal location and coding sequence location is as set out in Genome Reference Consortium Human Build 38 (GrCh38).