Genome editing composition targeting B2M genes and method of use

Prime editing guide RNA compositions address allograft rejection in cell therapy by precisely editing the B2M gene, improving compatibility and safety in allogeneic cell therapy without inducing double-strand breaks.

JP2026517403APending Publication Date: 2026-05-29PRIME MEDICINE INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PRIME MEDICINE INC
Filing Date
2024-05-16
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing cell therapy methods, particularly allogeneic cell therapy, face challenges with allograft rejection due to human leukocyte antigen (HLA) expression on donor-derived cells, leading to graft dysfunction and engraftment failure, and current genetic editing techniques like CRISPR-Cas9 introduce undesirable consequences such as complex mixtures of products and transpositions.

Method used

The use of prime editing guide RNA (PEgRNA) compositions that precisely disrupt the B2M gene by introducing targeted nucleotide changes without inducing double-strand breaks, thereby reducing HLA expression and minimizing rejection risks.

Benefits of technology

This approach effectively reduces allograft rejection by precisely editing the B2M gene, enhancing the compatibility and safety of allogeneic cell therapy without the undesirable effects associated with traditional genetic editing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026517403000223
    Figure 2026517403000223
  • Figure 2026517403000224
    Figure 2026517403000224
  • Figure 2026517403000225
    Figure 2026517403000225
Patent Text Reader

Abstract

Programmable nucleases such as CRISPR-Cas9 can disrupt genes by performing double-strand breaks (DSBs), thereby inducing a mixture of insertions and deletions (indels) at the target site. However, DSBs are associated with undesirable consequences, including complex mixtures of products and transpositions. Provided herein are compositions comprising a prime editing system and methods for using the prime editing system to edit the B2M gene. Provided herein are compositions comprising edited cells and methods for generating edited cells. Also provided herein are methods for using the edited cells.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefits of U.S. Provisional Application No. 63 / 502,563, filed on 16 May 2023, and U.S. Provisional Application No. 63 / 603,477, filed on 28 November 2023, each of which is incorporated herein by reference in its entirety. [Background technology]

[0002] Cell therapy offers a potential therapeutic approach to multiple diseases, including cancer, autoimmune diseases, and hematological disorders. For example, adoptive T-cell therapy may involve ex vivo manipulation of T cells to express T cell receptors (TCRs) or chimeric antigen receptors (CARs) engineered to recognize tumor-specific antigens. Cell therapy requires autologous (i.e., patient-derived) cells, such as T cells or hematopoietic stem cells (HSCs), which can avoid immunogenicity and intolerance issues during the reintroduction of ex vivo-engineered cells. However, due to manufacturing challenges of autologous cell therapy (e.g., time, cost, poor quality / quantity of extracted T cells), allogeneic (i.e., donor-derived) therapy has become an attractive alternative. A drawback of allogeneic cell therapy is the expression of endogenous proteins on the cell surface that affect the compatibility of donor-derived cells. For example, human leukocyte antigens (HLA) on the surface of allogeneic T cells or HSCs can trigger rejection by the host immune system, leading to graft dysfunction and engraftment failure. While HLA-matched donor-derived cells and immunosuppressive regimens can reduce the risk of host rejection, the former is often unavailable, and the latter carries significant side effects such as an increased risk of infection.

[0003] Human HLA proteins are heterodimers composed of an α chain encoded by a variant HLA gene and a β chain encoded by the β-2 microglobulin (B2M) gene. The B2M gene is located at 15q.21.1 in the human genome and encodes approximately 360 nucleotides of mRNA. Since the β chain is necessary for the dimerization and structure of the HLA complex, disruption of endogenous HLA can be achieved by knocking out or knocking down the expression of the B2M gene. Therefore, the effects of allograft rejection are expected to be reduced or eliminated by reducing or eliminating the expression of the B2M gene through genetic engineering.

[0004] Programmable nucleases such as CRISPR-Cas9 can disrupt genes by inducing double-strand breaks (DSBs) at target sites, thereby introducing a mixture of insertions and deletions (indels). However, DSBs are associated with undesirable consequences, including complex mixtures of products and transpositions. In the art, there is a need for compositions and methods for precisely disrupting B2M genes without introducing DSBs. [Overview of the Initiative]

[0005] In some embodiments, the foregoing provides methods and compositions for introducing donor DNA into target DNA by prime editing.

[0006] In some embodiments, a prime editing guide RNA (PEgRNA) or one or more polynucleotides encoding the PEgRNA comprises a spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, having Sequence ID No. 205 at its 3' end, a gRNA core capable of binding to the Cas9 protein, an editing template comprising a region complementary to the editing target sequence on the second strand of the B2M gene, and an extension arm comprising a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of Sequence ID No. 205, wherein the first and second strands are complementary to each other, and the editing template encodes one or more nucleotide changes compared to the editing target sequence.

[0007] In some embodiments, a prime-editing guide RNA (PEgRNA) or one or more polynucleotides encoding the PEgRNA comprises a spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, having Sequence ID No. 4 at its 3' end, a gRNA core capable of binding to the Cas9 protein, an editing template comprising a region complementary to the editing target sequence on the second strand of the B2M gene, and an extension arm comprising a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of Sequence ID No. 4, wherein the first and second strands are complementary to each other, and the editing template encodes one or more nucleotide changes compared to the editing target sequence.

[0008] A prime editing guide RNA (PEgRNA) or one or more polynucleotides encoding the PEgRNA comprises a spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, having Sequence ID No. 272 ​​at its 3' end; a gRNA core capable of binding to the Cas9 protein; an editing template having a region complementary to the editing target sequence on the second strand of the B2M gene; and an extension arm having a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of Sequence ID No. 272, wherein the first and second strands are complementary to each other, and the editing template encodes one or more nucleotide changes compared to the editing target sequence.

[0009] In some embodiments, a prime editing guide RNA (PEgRNA) or one or more polynucleotides encoding the PEgRNA comprises a spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, having Sequence ID 330 at its 3' end; a gRNA core capable of binding to the Cas9 protein; an editing template comprising a region complementary to the editing target sequence on the second strand of the B2M gene; and an extension arm comprising a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of Sequence ID 330, wherein the first and second strands are complementary to each other, and the editing template encodes one or more nucleotide changes compared to the editing target sequence.

[0010] In some embodiments, the spacer is 17 to 22 nucleotides long, and optionally, 20 nucleotides long. In some embodiments, the spacer includes one of sequence numbers 202 to 204 at its 3' end. In some embodiments, the spacer includes sequence number 204. In some embodiments, the spacer includes one of sequence numbers 1 to 3 at its 3' end. In some embodiments, the spacer includes sequence number 1. In some embodiments, the spacer includes one of sequence numbers 269 to 271 at its 3' end. In some embodiments, the spacer includes sequence number 269. In some embodiments, the spacer includes one of sequence numbers 327 to 329 at its 3' end. In some embodiments, the spacer includes sequence number 327.

[0011] In some embodiments, the PEgRNA further includes one or more nucleotide changes encoded by an editing template, which includes nonsynonymous editing that alters the mRNA or protein sequence encoded by the B2M gene. In some embodiments, the nonsynonymous editing results in one or more in-frame stop codons in the B2M gene. In some embodiments, the one or more in-frame stop codons include nonsense mutations in the B2M gene. In some embodiments, the nonsynonymous editing includes an insertion in the B2M gene. In some embodiments, the nonsynonymous editing includes one or more substitutions in the B2M gene. In some embodiments, the insertion includes an insertion of an in-frame stop codon in the B2M gene, and optionally, the insertion includes an insertion of two or more consecutive in-frame stop codons in the B2M gene. In some embodiments, the insertion includes an insertion of TAATAA, TTATTA, or TAATAG nucleotides. In some embodiments, the nonsynonymous editing includes a frameshift mutation in the B2M gene. In some embodiments, the frameshift mutation is an insertion of 3x+1 or 3x+2 nucleotides, where x is a non-negative integer. In some embodiments, the frameshift mutation is a deletion of 3x+1 or 3x+2 nucleotides, where x is a non-negative integer. In some embodiments, the insertion is 1, 2, or 4 nucleotides long. In some embodiments, the deletion is 1 nucleotide long.

[0012] In some embodiments, the PEgRNA further includes a non-synonymous edit, which modifies a protospacer adjacent motif (PAM) sequence immediately 3' of a protospacer sequence in the second strand of the B2M gene that is complementary to the search target sequence in the first strand of the B2M gene. In some embodiments, the PAM sequence is NGG, and the non-synonymous edit is an NGG->NGC edit. In some embodiments, the protospacer sequence includes a nick site 3 nucleotides upstream of the furthest 5' nucleotide of the PAM sequence, and the number of nucleotides from the nick site to the position in the second strand of the B2M gene corresponding to the non-synonymous edit is 1 to 19 nucleotides, and the number of nucleotides does not include the furthest 5' nucleotide position in the second strand corresponding to the non-synonymous edit. In some embodiments, the number of nucleotides from the nick site to the position in the second strand of the B2M gene corresponding to the non-synonymous edit is 1, 2, 7, 8, 13, 14, or 19 nucleotides. In some embodiments, the number of nucleotides from the nick site to the position on the second strand of the B2M gene corresponding to the non-synonymous edit is 8 nucleotides or less. In some embodiments, the number of nucleotides from the nick site to the position on the second strand of the B2M gene corresponding to the non-synonymous edit is 1 or 2 nucleotides.

[0013] In some embodiments, the non-synonymous edit is at a chromosomal location corresponding to coding sequence locations c.51, c.54, or c.50 of the wild-type B2M gene. In some embodiments, the non-synonymous edit includes the insertion of c.54insTAATAA. In some embodiments, the non-synonymous edit includes the deletion of c.51delC or the insertion of 50insG. In some embodiments, the non-synonymous edit is at a chromosomal location corresponding to coding sequence locations c.54, c.60, or c.66 of the wild-type B2M gene. In some embodiments, the non-synonymous edit includes the insertion of c.54_55insCC or c.54_55insTAAG. In some embodiments, the non-synonymous edit includes the insertion of c.54_55insTAATAA. In some embodiments, the non-synonymous edit includes the insertion of c.66_67insCC or c.66_67insTAAG. In some embodiments, the non-synonymous edit includes the insertion of c.66_67insTAATAA. In some embodiments, the non-synonymous edit includes the deletion of c.60_65 and the insertion of TAATAG (c.60_65_delinsTAATAG). In some embodiments, the non-synonymous edit is at a chromosomal position corresponding to the coding sequence position c.21 or c.3 of the wild-type B2M gene. In some embodiments, the non-synonymous edit includes the insertion of c.21insTAATAA. In some embodiments, the non-synonymous edit includes the insertion of c.21_22insCC or the editing of c.21_22insTAAG. In some embodiments, the non-synonymous edit includes the insertion of c.3_4insCC or the insertion of c.3_4insTAAG. In some embodiments, the non-synonymous edit includes the deletion of c.3_8 and the insertion of TAATGA (c.3_8delinsTAATGA). In some embodiments, the non-synonymous edit is at a chromosomal location corresponding to coding sequence locations c.21, c.15, or c.3 of the wild-type B2M gene. In some embodiments, the non-synonymous edit includes the insertion of c.15_16insCC or c.15_16insTAAG. In some embodiments, the non-synonymous edit includes the insertion of c.15_16insTAATAA.In some embodiments, the non-synonymous edit includes the insertion of c.3_4insCC or c.3_4insTAAG. In some embodiments, the non-synonymous edit includes the insertion of c.3_4insTAATAA. In some embodiments, the non-synonymous edit includes the deletion of c.3_8 and the insertion of TAATGA (c.3_8delinsTAATGA).

[0014] In some embodiments, the PEgRNA further comprises an editing template, which further encodes a further PAM silencing edit. In some embodiments, the PAM silencing edit is a c.58 G>C edit. In some embodiments, the PAM silencing edit is a c.17 C>G edit. In some embodiments, the PAM silencing edit is a c.11 C>G edit. In some embodiments, the editing template comprises at least 6, 8, or 10 consecutive nucleotides complementary to the editing target sequence, wherein the at least 6, 8, or 10 consecutive nucleotides are upstream of the position of the furthest 5' nucleotide of one or more nucleotide changes encoded in the editing template. In some embodiments, the editing template comprises 4, 6, 8, or 10 consecutive nucleotides complementary to the editing target sequence, wherein the 4, 6, 8, or 10 consecutive nucleotides are upstream of the position of the furthest 5' nucleotide of one or more nucleotide changes encoded in the editing template.

[0015] In some embodiments, prime editing guide RNA (PEgRNA), or the nucleic acid encoding the PEgRNA, comprises a spacer containing SEQ ID NO: 205 at its 3' end, a gRNA core capable of binding to a Cas9 protein, and an elongation arm containing an editing template at its 3' end, comprising (A) nucleotides 13-24 of SEQ ID NO: 221, (B) nucleotides 12-20 of SEQ ID NO: 227, or (C) nucleotides 7-17 of SEQ ID NO: 231, and a primer-binding site (PBS) at its 5' end, comprising the reverse complementary sequence of nucleotides 10-14 of SEQ ID NO: 205.

[0016] In some embodiments, prime editing guide RNA (PEgRNA), or the nucleic acid encoding the PEgRNA, comprises a spacer containing SEQ ID NO: 205 at its 3' end, a gRNA core capable of binding to a Cas9 protein, and an elongation arm at its 3' end containing an editing template comprising (A) nucleotides 13-24 of SEQ ID NO: 221, (B) nucleotides 12-20 of SEQ ID NO: 227, or (C) nucleotides 7-17 of SEQ ID NO: 231, and a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of SEQ ID NO: 205.

[0017] In some embodiments, the PEgRNA further comprises: (i) the editing template having nucleotides 13-24 of SEQ ID NO: 221 at its 3' end, and optionally the editing template having SEQ ID NOs: 219, 220; or (ii) the editing template having nucleotides 12-20 of SEQ ID NO: 227 at its 3' end, and optionally the editing template having one of SEQ ID NOs: 224-227 at its 3' end; or (ii) the editing template having nucleotides 7-17 of SEQ ID NO: 231 at its 3' end, and optionally the editing template having one of SEQ ID NOs: 229-231 at its 3' end.

[0018] In some embodiments, the prime editing guide RNA (PEgRNA), or the nucleic acid encoding the PEgRNA, comprises a spacer containing SEQ ID NO: 1 at its 3' end, a gRNA core capable of binding to a Cas9 protein, and an extension arm containing an editing template at its 3' end comprising (A) nucleotides 5-16 of SEQ ID NO: 19, or (B) sequences selected from the group consisting of SEQ ID NOs: 900, 904, 908, 912, 916, 920, and 924, and a primer-binding site (PBS) at its 5' end comprising the reverse complementary sequence of nucleotides 10-14 of SEQ ID NO: 1.

[0019] In some embodiments, the PEgRNA includes an editing template, and the editing template includes (i) a sequence selected from the group consisting of SEQ ID NOs: 900 to 903, or (ii) a sequence selected from the group consisting of SEQ ID NOs: 904 to 907, or (iii) a sequence selected from the group consisting of SEQ ID NOs: 908 to 911, or (iv) a sequence selected from the group consisting of SEQ ID NOs: 912 to 915, or (v) a sequence selected from the group consisting of SEQ ID NOs: 916 to 919, 928, and 929, or (vi) a sequence selected from the group consisting of SEQ ID NOs: 920 to 923, or (vii) a sequence selected from the group consisting of SEQ ID NOs: 924 to 927, or (viii) a sequence selected from the group consisting of SEQ ID NOs: 18 to 20.

[0020] In some aspects, the prime editing guide RNA (PEgRNA), or the nucleic acid encoding the PEgRNA, includes a spacer containing SEQ ID NO: 269 at the 3'-end, a gRNA core capable of binding to the Cas9 protein, and at the 3'-end, (A) nucleotides 3 to 16 of SEQ ID NO: 286, or (B) an editing template including a sequence selected from the group consisting of SEQ ID NOs: 1033, 1037, 1041, 1045, 1049, 1053, and 1057, and an extension arm including a primer binding site (PBS) containing the reverse complementary sequence of nucleotides 10 to 14 of SEQ ID NO: 269 at the 5'-end.

[0021] In some embodiments, the PEgRNA includes an editing template, and the editing template includes (i) a sequence selected from the group consisting of SEQ ID NOs: 1033 to 1036, or (ii) a sequence selected from the group consisting of SEQ ID NOs: 1037 to 1040, or (iii) a sequence selected from the group consisting of SEQ ID NOs: 1041 to 1044, or (iv) a sequence selected from the group consisting of SEQ ID NOs: 1045 to 1048, or (v) a sequence selected from the group consisting of SEQ ID NOs: 1049 to 1052 and 1061 to 1063, or (vi) a sequence selected from the group consisting of SEQ ID NOs: 1053 to 1056, or (vi) a sequence selected from the group consisting of SEQ ID NOs: 1057 to 1060, or (vii) a sequence selected from the group consisting of SEQ ID NOs: 286 to 288.

[0022] In some embodiments, the prime editing guide RNA (PEgRNA), or the nucleic acid encoding the PEgRNA, comprises a spacer containing SEQ ID NO: 327 at the 3′ end, a gRNA core capable of binding to the Cas9 protein, and an editing template containing at the 3′ end (A) nucleotides 6-16 of SEQ ID NO: 344, or (B) a sequence selected from the group consisting of SEQ ID NOs: 1162, 1166, 1170, 1174, 1178, 1182, and 1190, and an extension arm containing a primer binding site (PBS) containing the reverse complementary sequence of nucleotides 10-14 of SEQ ID NO: 327 at the 5′ end.

[0023] In some embodiments, the PEgRNA further comprises an editing template, and the editing template comprises (i) a sequence selected from the group consisting of SEQ ID NOs: 1162-1165, or (ii) a sequence selected from the group consisting of SEQ ID NOs: 1166-1169, or (iii) a sequence selected from the group consisting of SEQ ID NOs: 1170-1173, or (iv) a sequence selected from the group consisting of SEQ ID NOs: 1174-1177, or (v) a sequence selected from the group consisting of SEQ ID NOs: 1178-1181 and 1191, or (vi) a sequence selected from the group consisting of SEQ ID NOs: 1182-1185, or (vi) a sequence selected from the group consisting of SEQ ID NOs: 1186-1190, or (vii) a sequence selected from the group consisting of SEQ ID NOs: 344-346.

[0024] In some embodiments, the editing template has a length of 24 nucleotides or less, or 20 nucleotides or less. In some embodiments, the editing template has a length of (i) 10-20 nucleotides, (ii) 12-20 nucleotides, or (iii) 11-17 nucleotides. In some embodiments, the editing template has a length of 16-24 nucleotides. In some embodiments, the editing template has a length of 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 二十二, 24, 25, 26, 27, 28, 29, 30, 31, 33, 35 nucleotides.

[0025] In some embodiments, the PBS has a length of 17 nucleotides or less. In some embodiments, the PBS has a length of (i) 8 to 15 nucleotides, (ii) 8 to 14 nucleotides, or (iii) 8 to 12 nucleotides. In some embodiments, the PBS is 8, 10, or 12 nucleotides long. In some embodiments, the PBS contains the sequence described in any one of SEQ ID NOs. 206 to 218. In some embodiments, the PBS contains the sequence described in any one of SEQ ID NOs. 5 to 17. In some embodiments, the PBS contains the sequence described in any one of SEQ ID NOs. 273 to 285. In some embodiments, the PBS contains the sequence described in any one of SEQ ID NOs. 331 to 343.

[0026] In some embodiments, the spacer, the gRNA core, the RTT, and the PBS form a continuous sequence in a single molecule. In some embodiments, the PEgRNA includes the spacer, the gRNA core, the RTT, and the PBS from 5' to 3'. In some embodiments, the gRNA core includes SEQ ID NO: 646. In some embodiments, the gRNA core includes SEQ ID NO: 653.

[0027] In some embodiments, the PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs. 232-262. In some embodiments, the PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs. 21-29 and 930-1016. In some embodiments, the PEgRNA includes the sequence described in SEQ ID NOs. 933, 937, 961, 941, 957, or 936. In some embodiments, the PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs. 289-297 and 1064-1151. In some embodiments, the PEgRNA includes the sequence described in SEQ ID NOs. 1141 or 1143. In some embodiments, the PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs. 347-355 and 1192-1279. In some embodiments, the PEgRNA includes the sequence described in SEQ ID NOs. 1269 or 1265. In some embodiments, the PEgRNA contains a sequence selected from the group consisting of SEQ ID NOs: 957, 961, 965, 980, 1016, 956, 933, 941, 937, 1223, 988, 984, 1225, 1151, 1095, 1091, 964, 960, 940, 1221, 945, 1219, 932, 1015, 1014, 1075, 1222, 1250, 936, 1013, 1119, 1226, and 949.

[0028] In some embodiments, the PEgRNA further includes a 3' motif, which is optionally linked to the 3' end of the PBS via a linker.

[0029] In some embodiments, the PEgRNA further comprises 3'mN*mN*mN*N and / or 5'mN*mN*mN* modifications, where m indicates the nucleotide comprises a 2'-O-Me modification and * indicates the presence of a phosphorothioate bond. In some embodiments, the PEgRNA further comprises 3'mT*mT*mT*T and 5'mN*mN*mN* modifications, where m indicates the nucleotide comprises a 2'-O-Me modification, * indicates the presence of a phosphorothioate bond and T indicates the presence of a further uridine nucleotide.

[0030] In some embodiments, the human chromosome locations and coding sequence locations are as described in the Genome Reference Consortium Human Build 38 (GrCh38).

[0031] In some embodiments, the prime editing system includes the PEgRNA or one or more polynucleotides encoding the PEgRNA.

[0032] In some embodiments, the prime editing system further comprises a nick guide RNA (ngRNA), or a nucleic acid encoding the ngRNA, the ngRNA comprising a) an ngRNA spacer complementary to an ngRNA search target sequence on the second strand of the B2M gene, and b) an ngRNA core capable of binding to the Cas9 protein.

[0033] In some embodiments, the prime editing system includes an ngRNA spacer, which is 17 to 22 nucleotides long, and optionally 20 nucleotides long. In some embodiments, the ngRNA core includes SEQ ID NO: 646 or 653. In some embodiments, the prime editing system includes a PEgRNA spacer, which includes SEQ ID NO: 205 at its 3' end. In some embodiments, the ngRNA spacer includes a sequence at its 3' end corresponding to nucleotides 4 to 20, 3 to 20, 2 to 20, or 1 to 20 of any one of SEQ ID NOs: 263 to 268, and optionally, the ngRNA spacer includes any one of SEQ ID NOs: 263 to 268 at its 3' end. In some embodiments, the ngRNA spacer includes nucleotides 1 to 20 of SEQ ID NO: 268 at its 3' end, and optionally, the ngRNA includes SEQ ID NO: 824 or 825.

[0034] In some embodiments, the prime editing system comprises (i) the non-synonymous edit encoded by the editing template comprising a deletion of c.51delC, and the ngRNA spacer comprising a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 266, or (ii) the non-synonymous edit encoded by the editing template comprising an insertion of c.50insG, and the ngRNA spacer comprising a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 267 or 268.

[0035] In some embodiments, the ngRNA includes a sequence selected from the group consisting of SEQ ID NOs. 824-827. In some embodiments, the PEgRNA spacer includes SEQ ID NOs. 4 at its 3' end. In some embodiments, the ngRNA spacer includes a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of any one of SEQ ID NOs. 1017-1024. In some embodiments, the ngRNA spacer includes any one of SEQ ID NOs. 1017-1024.

[0036] In some embodiments, the prime editing system (i) the editing template encodes the editing of c.54_55insCC and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1018; (ii) the editing template encodes the editing of c.66_67insCC and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1019; (iii) the editing template encodes the editing of c.54_55insTAAG and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1020; (iv) the editing template encodes the editing of c.66_67insTAAG and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1020; (v) The editing template contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1021, or (v) the editing template codes for the editing of c.54_55insTAATAA, and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1022, or (vi) the template codes for the editing of c.66_67insTAATAA (vii) The template is configured such that the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1023, or the template is configured such that the template codes for the editing of c.60_65delinsTAATAG, and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1024.

[0037] In some embodiments, the ngRNA includes a sequence selected from the group consisting of SEQ ID NOs. 1025-1032. In some embodiments, the PEgRNA spacer includes SEQ ID NOs. 272 ​​at its 3' end. In some embodiments, the ngRNA spacer includes a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of any one of SEQ ID NOs. 1152-1156. In some embodiments, the ngRNA spacer includes any one of SEQ ID NOs. 1152-1156.

[0038] In some embodiments, the prime editing system (i) the editing template encodes the editing of c.3_4insCC, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1153, or (ii) the editing template encodes the editing of c.3_4insTAAG, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1154. (iii) The editing template comprises the following: the editing template codes for the editing of c.3_4insTAATAA, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1155; or (iv) the editing template codes for the editing of c.3_8delinsTAATGA, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1156.

[0039] In some embodiments, the ngRNA includes a sequence selected from the group consisting of SEQ ID NOs. 1157-1161. In some embodiments, the PEgRNA spacer includes SEQ ID NOs. 330 at its 3' end. In some embodiments, the ngRNA spacer includes a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of any one of SEQ ID NOs. 1280-1284. In some embodiments, the ngRNA spacer includes any one of SEQ ID NOs. 1280-1284.

[0040] In some embodiments, the prime editing system (i) the editing template encodes the editing of c.3_4insCC, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1281, or (ii) the editing template encodes the editing of c.3_4insTAAG, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1282. (iii) The editing template encodes the editing of c.3_4insTAATAA, and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1283, or (iv) The editing template encodes the editing of c.3_8delinsTAATGA, and the ngRNA spacer contains a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1284. In some embodiments, the ngRNA contains a sequence selected from the group consisting of SEQ ID NOs: 1285-1289.

[0041] In some embodiments, the prime editing system (i) the PEgRNA includes the sequence described in SEQ ID NO: 933 or 937 and the ngRNA includes the sequence described in SEQ ID NO: 1018; (ii) the PEgRNA includes the sequence described in SEQ ID NO: 961 and the ngRNA includes the sequence described in SEQ ID NO: 1020; (iii) the PEgRNA includes the sequence described in SEQ ID NO: 941 and the ngRNA includes the sequence described in SEQ ID NO: 1018; (iv) the PEgRNA includes the sequence described in SEQ ID NO: 957 (v) the ngRNA contains the sequence described in SEQ ID NO: 1020, (v) the PEgRNA contains the sequence described in SEQ ID NO: 936 and the ngRNA contains the sequence described in SEQ ID NO: 1018, (vi) the PEgRNA contains the sequence described in SEQ ID NO: 1141 or 1143 and the ngRNA contains the sequence described in SEQ ID NO: 1156, or (vii) the PEgRNA contains the sequence described in SEQ ID NO: 1269 or 1265 and the ngRNA contains the sequence described in SEQ ID NO: 1284.

[0042] In some embodiments, the ngRNA includes the modification 3'mN*mN*mN*N and / or 5'mN*mN*mN*, where m indicates the nucleotide includes the modification 2'-O-Me and * indicates the presence of a phosphorothioate bond. In some embodiments, the ngRNA includes the modification 3'mT*mT*mT*T and 5'mN*mN*mN*, where m indicates the nucleotide includes the modification 2'-O-Me, * indicates the presence of a phosphorothioate bond and T indicates the presence of a further uridine nucleotide.

[0043] In some embodiments, the prime editing system further comprises a TRAC-PEgRNA or one or more polynucleotides encoding the TRAC-PEgRNA, wherein the TRAC-PEgRNA comprises a TRAC-spacer complementary to a search target sequence on the first strand of the T cell receptor α constant (TRAC) gene, a TRAC-gRNA core capable of binding to the Cas9 protein, and a TRAC-edit template comprising a region complementary to an editing target sequence on the second strand of the TRAC gene, and a TRAC-extension arm comprising a TRAC-primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotide p~(q-3) of the second spacer, where q is the length of the second spacer and p is an integer between 1~(q-6), the first and second strands being complementary to each other, and the editing template encoding one or more nucleotide changes compared to the editing target sequence.

[0044] In some embodiments, the TRAC spacer is 17 to 22 nucleotides long. In some embodiments, the editing template encodes an in-frame stop codon in the TRAC gene or a frameshift mutation in the TRAC gene. In some embodiments, the editing template encodes a recombinase-recognition sequence recognized by the recombinase, or its inverse complementary sequence.

[0045] In some embodiments, the prime editing system further comprises a first TRAC-prime editing guide RNA (PEgRNA) or one or more polynucleotides encoding the first TRAC-PEgRNA, and a second TRAC-PEgRNA or one or more polynucleotides encoding the second TRAC-PEgRNA, wherein the first TRAC-PEgRNA comprises a) a first TRAC-spacer complementary to a first TRAC-search target sequence on the first strand of the TRAC gene, b) a first TRAC-gRNA core capable of binding to the Cas9 protein, and c) a first TRAC-ply having (A) a first TRAC-edit template and (B) a first TRAC-ply having at its 5' end a reverse complementary sequence of nucleotide p~(q-3) of the first TRAC-spacer. A first TRAC-extension arm comprising an IMA binding site (PBS), where q is the length of the first TRAC-spacer and p is an integer from 1 to (q-6), the second TRAC-PEgRNA comprising a second TRAC-spacer complementary to a second TRAC-search target sequence on the second strand of the TRAC gene complementary to the first strand, a second TRAC-gRNA core capable of binding to the Cas9 protein, and a second TRAC-extension arm comprising (A) a second TRAC-editing template and (B) a second TRAC-PBS at its 5' end containing the reverse complementary sequence of nucleotides m to (n-3) of the second TRAC-spacer, where n is the length of the second TRAC-spacer and m is an integer from 1 to (n-6).

[0046] In some embodiments, the first TRAC spacer contains nucleotides 4-20 of a sequence selected from the group consisting of SEQ ID NOs: 1303 and 1353 at its 3' end. In some embodiments, the second TRAC spacer contains nucleotides 4-20 of a sequence selected from the group consisting of SEQ ID NOs: 1417, 1481, and 1532 at its 3' end. In some embodiments, the first TRAC spacer has a length of 17-22 nucleotides, and / or the second TRAC spacer has a length of 17-22 nucleotides. In some embodiments, the first and second TRAC spacers are each 20 nucleotides long. In some embodiments, the first TRAC spacer contains SEQ ID NOs: 1303 or 1353 at its 3' end. In some embodiments, the second TRAC spacer contains SEQ ID NOs: 1417, 1481, or 1532 at its 3' end. In some embodiments, the first TRAC PBS is 7 to 17 nucleotides long and has a reverse complementary sequence of nucleotides 11 to 17, 10 to 17, 9 to 17, 8 to 17, 7 to 17, 6 to 17, 5 to 17, 4 to 17, 3 to 17, 2 to 17, or 1 to 17 of the sequence selected for the first TRAC spacer at its 5' end. In some embodiments, the second TRAC PBS is 7 to 17 nucleotides long and has a reverse complementary sequence of nucleotides 11 to 17, 10 to 17, 9 to 17, 8 to 17, 7 to 17, 6 to 17, 5 to 17, 4 to 17, 3 to 17, 2 to 17, or 1 to 17 of the sequence selected for the second TRAC spacer at its 5' end.

[0047] In some embodiments, the first and / or second TRAC PBS is 8 to 13 nucleotides long. In some embodiments, the first and / or second TRAC PBS is 11, 12, or 13 nucleotides long.

[0048] In some embodiments, the first gRNA core, the second gRNA core, or both include sequence number 646 or 653.

[0049] In some embodiments, the first TRAC editing template includes a region complementary to the second TRAC editing template. In some embodiments, the first and second TRAC editing templates each encode all or a fragment of a recombinase recognition sequence (RRS) or its reverse complementary sequence, the first TRAC editing template encodes at least a 5' portion of the RRS or its reverse complementary sequence, the second TRAC editing template encodes at least a 3' portion of the RRS or its reverse complementary sequence, and the 5' ends of at least 10 nucleotides of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other. In some embodiments, at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other, and optionally, at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides at the 5' ends of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other.

[0050] In some embodiments, the first TRAC editing template encodes the RRS. In some embodiments, the second TRAC editing template encodes the RRS. In some embodiments, the RRS is an attB sequence recognized by Bxb1 recombinase. In some embodiments, the RRS is an attP sequence recognized by Bxb1 recombinase. In some embodiments, the first TRAC editing template includes RTT#1 in Table 39 and the second TRAC editing template includes RTT#2 in Table 39, or the first TRAC editing template includes RTT#2 in Table 39 and the second TRAC editing template includes RTT#1 in Table 39. In some embodiments, the first TRAC editing template includes a 5' fragment of the RTT described in Table 39, and the second TRAC editing template includes the full-length or 5' fragment of the corresponding RTT pair, with at least 10 nucleotides at the 5' ends of the first and second TRAC editing templates having complete reverse complementarity with respect to each other.

[0051] In some embodiments, at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other, and optionally, at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides at the 5' ends of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other. In some embodiments, the length of the complementary region of the first TRAC editing template is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90%, and optionally, the length of the complementary region of the first editing template is at least 52%, at least 53%, or at least 55% of the length of the first TRAC editing template. In some embodiments, the length of the complementary region of the second TRAC editing template is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90%, and optionally, the length of the complementary region of the second TRAC editing template is at least 52%, at least 53%, or at least 55% of the length of the second TRAC editing template.

[0052] In some embodiments, the prime editing system includes a TRAC spacer, in which case (a) the first TRAC spacer includes sequence number 1303 and the first TRAC PBS includes sequence number 1312, or (b) the first TRAC spacer includes sequence number 1353 and the first TRAC PBS includes sequence numbers 1361, 1362, 1363, or 1364. In some embodiments, the second TRAC spacer includes sequence number 1417 and the second TRAC PBS includes sequence number 1428, or the second TRAC spacer includes sequence number 1481 and the second TRAC PBS includes sequence number 1489. In some embodiments, (a) the first TRAC spacer includes sequence number 1303 and the first TRAC PBS has a sequence according to sequence number 1313 or 1314, or the first TRAC spacer includes sequence number 1353 and the first TRAC PBS has a sequence according to sequence number 1361 or 1363, (b) the second TRAC spacer includes sequence number 1417 and the second TRAC PBS has a sequence according to sequence number 1426 or 1428, the second TRAC spacer includes sequence number 1480 and the second TRAC PBS has a sequence according to sequence number 1486 or 1487, or the second TRAC spacer includes sequence number 1532 and the second TRAC PBS has a sequence according to sequence number 1541 or 1543.

[0053] In some embodiments, the first TRAC spacer includes sequence number 1353, the first TRAC PBS has a sequence according to sequence number 1361, the second TRAC spacer includes sequence number 1417, and the second TRAC PBS has a sequence according to sequence number 1426. In some embodiments, the first TRAC editing template includes sequence number 1577, and the second TRAC editing template includes sequence number 1584. In some embodiments, the first editing template includes sequence number 1584, and the second editing template includes sequence number 1577.

[0054] In some embodiments, the first TRAC PEgRNA includes a 5' TRAC PEgRNA sequence selected from any one of Tables 34 and 35, and the second TRAC PEgRNA includes a 3' TRAC PEgRNA sequence selected from any one of Tables 36 to 38. In some embodiments, the first TRAC PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs: 1328, 1382, 1387, and 1413, and the second TRAC PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs: 1475, 1476, 1477, 1525, 1526, 1527, 1573, and 1574. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 1382, and the second TRAC PEgRNA includes SEQ ID NO: 1527. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 1401, and the second TRAC PEgRNA includes SEQ ID NO: 1459. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 1343, and the second TRAC PEgRNA includes SEQ ID NO: 1566. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 1390, and the second TRAC PEgRNA includes SEQ ID NO: 1456. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 1336, and the second TRAC PEgRNA includes SEQ ID NO: 1560. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 1345, and the second TRAC PEgRNA includes SEQ ID NO: 1566. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 374, and the second TRAC PEgRNA includes SEQ ID NO: 1251.

[0055] In some embodiments, the first TRAC PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs: 1322, 1336, 1372, and 126, and the second TRAC PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs: 1442, 1456, 1501, and 1513.

[0056] In some embodiments, the first TRAC PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs: 1401, 1406, 1343, and 1345, and the second TRAC PEgRNA includes a sequence selected from the group consisting of SEQ ID NOs: 1451, 1516, 1568, 1566, and 1459. In some embodiments, the first TRAC PEgRNA includes SEQ ID NO: 1401, and the second TRAC PEgRNA includes SEQ ID NO: 1459.

[0057] In some embodiments, the first TRAC PEgRNA and / or the second TRAC PEgRNA further include a 3' motif, which is optionally linked to the 3' end of the first PBS or the second PBS via a linker.

[0058] In some embodiments, the first TRAC PEgRNA and / or the second TRAC PEgRNA further include 5'mN*mN*mN* and 3'mN*mN*mN*N modifications, where m indicates that the nucleotide includes a 2'-O-Me modification and * indicates the presence of a phosphorothioate bond.

[0059] In some embodiments, the prime editing system further includes the recombinase or a nucleic acid encoding the recombinase. In some embodiments, the recombinase is fused to or linked to the prime editor.

[0060] In some embodiments, the prime editing system further comprises a polynucleotide or a nucleic acid encoding the polynucleotide, wherein the polynucleotide comprises (a) a donor sequence and (b) a second recombinase-recognizing sequence (RRS) recognized by the recombinase. In some embodiments, the donor sequence encodes a chimeric antigen receptor (CAR). In some embodiments, (i) the RRS comprises SEQ ID NO: 1590 and the second RRS comprises SEQ ID NO: 1591, or (ii) the RRS comprises SEQ ID NO: 1591 and the second RRS comprises SEQ ID NO: 1590. In some embodiments, the recombinase is Bxb1.

[0061] In some embodiments, the prime editing system further comprises a prime editor or one or more polynucleotides encoding the prime editor, the prime editor comprising a) a Cas9 niccas having a nuclease-inactivating mutation in its HNH domain, and b) a reverse transcriptase. In some embodiments, the prime editor is a fusion protein.

[0062] In some embodiments, the prime editing system further comprises an N-terminal extaine containing an N-intane of the prime editor fusion protein and an N-terminal extaine or a polynucleotide encoding the N-terminal extaine, a C-terminal extaine containing a C-intane of the prime editor fusion protein and a C-terminal extaine or a polynucleotide encoding the C-terminal extaine, wherein the N-intane and C-intane of the N-terminal and C-terminal extaines are capable of self-cleavage to conjugate the N-terminal and C-terminal extaines to form the prime editor fusion protein, the prime editor fusion protein comprising a Cas9 nickas and a reverse transcriptase (RT) domain having a nuclease-inactivating mutation in the HNH domain. In some embodiments, the Cas9 nickas comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 676 or 677. In some embodiments, the reverse transcriptase contains an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 673. In some embodiments, the sequence identity is determined by a Needleman-Wunsch alignment of the two protein sequences with a gap presence cost of 11 and a gap extension cost of 1, and the identity percentage is calculated by dividing the number of identities by the length of the alignment.

[0063] In some embodiments, the prime editing system comprises a polynucleotide, wherein one or more polynucleotides encoding the prime editor, the polynucleotide encoding the N-terminal extension, or the polynucleotide encoding the C-terminal extension are mRNAs.

[0064] In some embodiments, a population of viral particles collectively comprises one or more polynucleotides or prime-editing systems encoding PEgRNA provided herein. In some embodiments, the viral particles are AAV particles.

[0065] In some embodiments, the LNP comprises a prime editing system provided herein. In some embodiments, the LNP comprises the PEgRNA and optionally the ngRNA, the Cas9 nickase-encoding polynucleotide, and the reverse transcriptase-encoding polynucleotide. In some embodiments, the LNP further comprises the Cas9 nickase-encoding polynucleotide, where the reverse transcriptase-encoding polynucleotide is mRNA. In some embodiments, the Cas9 nickase-encoding polynucleotide and the reverse transcriptase-encoding polynucleotide are in the same molecule.

[0066] In some embodiments, a method for editing the B2M gene comprises contacting the B2M gene with (a) a prime editor comprising PEgRNA and Cas9 nickase and reverse transcriptase having a nuclease-inactivating mutation in the HNH domain, (b) a prime editing system, (c) a population of viral particles, or (d) an LNP, as provided herein. In some embodiments, the B2M gene is intracellular. In some embodiments, a method for generating the engineered cell comprises introducing into a cell or population of cells (a) a prime editor comprising PEgRNA and Cas9 nickase and reverse transcriptase having a nuclease-inactivating mutation in the HNH domain, (b) a prime editing system, (c) a population of viral particles, or (d) an LNP, as provided herein. In some embodiments, the cell or population of cells is within a subject. In some embodiments, the cell or population of cells is exovivo, and optionally, the cell or population of cells is obtained from a subject or cell bank. In some embodiments, the cell or population of cells is human cell. In some embodiments, the cells or population of cells are immune cells. In some embodiments, the cells or population of cells are T cells, and optionally, the cells or population of cells are cytotoxic T cells.

[0067] In some embodiments, the manipulated cells or populations of manipulated cells contain immature stop codons within the B2M gene compared to the wild-type B2M gene.

[0068] In some embodiments, the manipulated cells or population of manipulated cells contain a B2M gene that includes an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to the coding sequence position c.51, c.54, or c.50 of the wild-type B2M gene, compared to the wild-type B2M gene.

[0069] In some embodiments, the manipulated cells or population of manipulated cells include a B2M gene that, compared to the wild-type B2M gene, contains an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.54, c.60, or c.66 of the wild-type B2M gene, and optionally, the B2M gene contains an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.58 of the wild-type B2M gene.

[0070] In some embodiments, the manipulated cells or population of manipulated cells include a B2M gene that, compared to the wild-type B2M gene, contains an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.21 or c.3 of the wild-type B2M gene, and optionally, the B2M gene contains an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.17 of the wild-type B2M gene.

[0071] In some embodiments, the manipulated cells or population of manipulated cells include a B2M gene that, compared to the wild-type B2M gene, has an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.21, c.15, or c.3 of the wild-type B2M gene, and optionally, the B2M gene has an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.11 of the wild-type B2M gene.

[0072] In some embodiments, the cells or population of cells contain a deletion of c.51delC within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells contain an insertion of c.50insG within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells contain an insertion of c.54_55insCC within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells contain an insertion of c.54_55insTAAG within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells contain an insertion of c.54_55insTAATAA within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells are comprised of an insertion of c.66_67insCC within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a substitution of c.58G>C within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells are comprised of an insertion of c.66_67insTAAG within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a substitution of c.58G>C within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells are comprised of an insertion of c.66_67insTAATAA within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a substitution of c.58G>C within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include a c.60_65 deletion and a TAATAG insertion (c.60_64delinsTAATAG) within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a c.58G>C substitution within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include a c.21_22insCC insertion within the B2M gene compared to the wild-type B2M gene.In some embodiments, the cells or population of cells include an insertion of c.21_22insTAAG within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include an insertion of c.21_22insTAATAA within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include an insertion of c.3_4insCC within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a substitution of c.17C>G or c.11C>G within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include an insertion of c.3_4insTAAG within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a substitution of c.17C>G or c.11C>G within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include an insertion of c.3_4insTAATAA within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a substitution of c.17C>G or c.11C>G within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include a deletion of c.3_8 and an insertion of TAATGA (c.3_8delinsTAATGA) within the B2M gene compared to the wild-type B2M gene, and optionally, the cells or population of cells further include a substitution of c.17C>G or c.11C>G within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells include an insertion of c.15_16insCC within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells contain an insertion of c.15_16insTAAG within the B2M gene compared to the wild-type B2M gene. In some embodiments, the cells or population of cells contain an insertion of c.15_16insTAATAA within the B2M gene compared to the wild-type B2M gene.

[0073] In some embodiments, the cells or population of cells contain a TRAC gene comprising the sequence GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAACC (SEQ ID NO: 9999) and / or GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 10000) compared to the wild-type TRAC gene. In some embodiments, the edited TRAC gene comprises, from 5' to 3', an insertion sequence comprising GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAACC (SEQ ID NO: 9999), a donor sequence, and GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 10000). In some embodiments, the edited TRAC gene includes, from 5' to 3', GGTTTGTCTGGTCAACCACCGCGGTCTCCGTCGTCAGGATCAT (SEQ ID NO: 10000), a donor sequence, and an insertion sequence including GGCTTGTCGACGACGGCGGTCTCAGTGGTGTACGGTACAAACC (SEQ ID NO: 9999).

[0074] In some embodiments, the cells or population of cells include a donor sequence which encodes a chimeric antigen receptor (CAR), and optionally, the donor encodes CD19CAR.

[0075] In some embodiments, the cells or population of cells include an insertion sequence located between a first chromosome location and a second chromosome location, the first chromosome location being selected from the group consisting of human chromosome 14 locations 22547458, 22547457, 22547449, and 22547448, and the second chromosome location being selected from the group consisting of human chromosome 14 locations 22547533, 22547523, 22547491, 22547528, 22547497, 22547579, 22547522, 22547485, 22547506, 22547560, 22547505, 22547529, and 22547490. In some embodiments, the insertion sequence is located between human chromosome 14 positions 22547458 and 22547533, between human chromosome 14 positions 22547458 and 22547522, between human chromosome 14 positions 22547458 and 22547529, between human chromosome 14 positions 22547449 and 22547533, between human chromosome 14 positions 22547449 and 22547522, or between human chromosome 14 positions 22547449 and 22547529.

[0076] In some embodiments, the human chromosome location and coding sequence location are as described in Genome Reference Consortium Human Build 38 (GrCh38).

[0077] In some embodiments, the cells or population of cells are within the subject. In some embodiments, the cells or population of cells are exovivo, and optionally, the cells or population of cells are obtained from the subject or a cell bank. In some embodiments, the cells or population of cells are human cells. In some embodiments, the cells or population of cells are immune cells. In some embodiments, the cells or population of cells are T cells, and optionally, the cells or population of cells are cytotoxic T cells.

[0078] In some embodiments, an immunotherapy method comprising administering to a subject (a) a prime editor comprising PEgRNA and Cas9 nickase and reverse transcriptase having a nuclease-inactivating mutation in the HNH domain, (b) a prime editing system, (c) a population of viral particles, (d) LNPs, or (e) cells or a population of cells, as provided herein.

[0079] In some embodiments, the PEgRNA comprises a prime editing guide RNA (PEgRNA), or one or more polynucleotides encoding the PEgRNA, wherein the PEgRNA is a spacer complementary to a search target sequence on the first strand of the β2-microglobulin (B2M) gene, and having a PEgRNA spacer sequence selected from any one of Tables 1 to 21 at its 3' end; b) a gRNA core capable of binding to a Cas9 protein; and c) an elongation arm, wherein the elongation arm comprises i) an editing template having an RTT sequence selected from the same table as the PEgRNA spacer sequence at its 3' end, and ii) a PBS having a primer-binding site (PBS) sequence selected from the same table as the PEgRNA spacer sequence at its 5' end. In some embodiments, the PEgRNA spacer is 17 to 22 nucleotides long. In some embodiments, the PEgRNA spacer is 20 nucleotides long. In some embodiments, the spacer, the gRNA core, the editing template, and the PBS form a continuous sequence in a single molecule. In some embodiments, the PEgRNA comprises the spacer, the gRNA core, the editing template, and the PBS, from 5' to 3'. In some embodiments, the prime editing system comprises the PEgRNA or one or more polynucleotides.

[0080] In some embodiments, the prime editing system further comprises a nick guide RNA (ngRNA), or one or more polynucleotides encoding the ngRNA, the ngRNA comprising (i) an ngRNA spacer containing a region complementary to the second strand of the B2M gene, and (ii) an ngRNA core capable of binding to the Cas9 protein. In some embodiments, the ngRNA spacer is 17 to 22 nucleotides long. In some embodiments, the ngRNA spacer is 20 nucleotides long. In some embodiments, the ngRNA spacer has an ngRNA spacer sequence at its 3' end, selected from the same table as the PEgRNA spacer sequence. In some embodiments, the ngRNA comprises an ngRNA sequence selected from the same table as the PEgRNA spacer sequence.

[0081] In some embodiments, the prime editing system further includes a prime editor comprising Cas9 nickas having a nuclease-inactivating mutation in its HNH domain, or one or more polynucleotides encoding the Cas9 nickas, and a reverse transcriptase, or one or more polynucleotides encoding the reverse transcriptase.

[0082] In some embodiments, the prime editing system further comprises an N-terminal extaine containing an N-intane of the prime editor fusion protein and an N-terminal extaine or a polynucleotide encoding the N-terminal extaine, and a C-terminal extaine containing a C-intane of the prime editor fusion protein and a C-terminal extaine or a polynucleotide encoding the C-terminal extaine, wherein the N-intane and C-intane of the N-terminal and C-terminal extaines are capable of self-cleavage to conjugate the N-terminal and C-terminal fragments to form the prime editor fusion protein, the prime editor fusion protein comprising Cas9 nickase and reverse transcriptase (RT) domains. In some embodiments, the Cas9 nickase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 676 or 677. In some embodiments, the reverse transcriptase contains an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 673. In some embodiments, the sequence identity is determined by a Needleman-Wunsch alignment of the two protein sequences with a gap presence cost of 11 and a gap extension cost of 1, and the identity percentage is calculated by dividing the number of identities by the length of the alignment.

[0083] In some embodiments, the prime editing system further comprises a TRAC-PEgRNA pair, the TRAC-PEgRNA pair comprising: a) a first TRAC-prime editing guide RNA (PEgRNA) or one or more polynucleotides encoding the first TRAC-PEgRNA; and b) a second TRAC-PEgRNA or one or more polynucleotides encoding the second TRAC-PEgRNA, wherein the first TRAC-PEgRNA comprises: i) a first TRAC-spacer having a 5' TRAC-PEgRNA spacer sequence at its 3' end selected from any one of Tables 34 and 35; ii) a first TRAC-gRNA core capable of binding to a Cas9 protein; and iii) (A) a first TRAC-editing template. The second TRAC-PEgRNA comprises (A) a second TRAC-edit template and (B) a second TRAC-PBS having a 5' TRAC-primer binding site (PBS) sequence selected from the same table as the first TRAC-spacer at its 5' end, the second TRAC-PEgRNA comprising i) a second TRAC-spacer having a 5' TRAC-PEgRNA spacer sequence selected from any one of Tables 36-38 at its 3' end, ii) a second TRAC-gRNA core capable of binding to the Cas9 protein, and iii) a second TRAC-extension arm comprising (A) a second TRAC-edit template and (B) a second TRAC-PBS having a 5' TRAC-primer binding site (PBS) sequence selected from the same table as the second TRAC-spacer at its 5' end.

[0084] In some embodiments, the first TRAC editing template includes a region complementary to the second TRAC editing template. In some embodiments, the first and second TRAC editing templates each encode all or a fragment of a recombinase recognition sequence (RRS) or its reverse complementary sequence, the first TRAC editing template encodes at least a 5' portion of the RRS or its reverse complementary sequence, the second TRAC editing template encodes at least a 3' portion of the RRS or its reverse complementary sequence, and the 5' ends of at least 10 nucleotides of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other.

[0085] In some embodiments, at least 15, 20, 25, or 30 nucleotides at the 5' ends of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other, and optionally, at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides at the 5' ends of the first and second TRAC editing templates are in complete reverse complementarity with respect to each other.

[0086] In some embodiments, the first TRAC editing template encodes the RRS, or the second TRAC editing template encodes the RRS, which optionally is an attB sequence recognized by Bxb1 recombinase, or an attP sequence recognized by Bxb1 recombinase.

[0087] In some embodiments, the first TRAC editing template includes RTT#1 in Table 39 and the second TRAC editing template includes RTT#2 in Table 39, or the first TRAC editing template includes RTT#2 in Table 39 and the second TRAC editing template includes RTT#1 in Table 39.

[0088] In some embodiments, the first TRAC editing template includes sequence number 1577 and the second TRAC editing template includes sequence number 1584, or the first editing template includes sequence number 1584 and the second editing template includes sequence number 1577.

[0089] In some embodiments, the first TRAC PEgRNA comprises a 5' TRAC PEgRNA sequence selected from any one of Tables 34 and 35, and the second TRAC PEgRNA comprises a 3' TRAC PEgRNA sequence selected from any one of Tables 36 to 38.

[0090] Embedding by reference All publications, patents, and patent applications described herein are incorporated herein by reference to the same extent that each individual publication, patent, or patent application is explicitly and individually indicated to be incorporated by reference.

[0091] Novel features of this disclosure are described in detail in the appended claims. A better understanding of the features and advantages of this disclosure can be obtained by referring to the following detailed description of exemplary embodiments in which the principles of this disclosure are utilized, and the accompanying drawings. [Brief explanation of the drawing]

[0092] [Figure 1] A schematic diagram of prime editing guide RNA (PEgRNA) that binds to double-stranded target DNA sequences is shown.

[0093] [Figure 2] An illustrative schematic diagram of a PEgRNA designed for prime editors provides an overview of the PEgRNA architecture.

[0094] [Figure 3] This is a schematic diagram showing the spacer and gRNA core portions of an exemplary guide RNA in two separate molecules. The rest of the PEgRNA structure is not shown.

[0095] [Figure 4A] An illustrative schematic diagram of a dual-prime editing system for editing both strands of a double-stranded target DNA is shown. Same color / shading indicates complementarity or identity between sequences.

[0096] [Figure 4B] This diagram illustrates a dual-prime editing process using a substituted double helix (RD) containing an overlapping double helix (OD). The same color / shading indicates complementarity or identity between sequences. [Modes for carrying out the invention]

[0097] Provided herein are compositions and methods for editing a target gene, β2-microglobulin (B2M), by prime editing, in some embodiments. The compositions provided herein may include a prime editor (PE) that can use an engineered guide polynucleotide, such as prime editing guide RNA (PEgRNA). The PEgRNA can direct the PE to a specific DNA target and encode DNA edits on the target gene B2M that perform a variety of functions, including direct disruption of the target gene. The editing can disrupt the B2M gene by, for example, introducing one or more stop codons, introducing frameshift mutations (insertions or deletions), or disrupting splice sites. Also provided herein are compositions comprising edited cells produced by the methods disclosed herein.

[0098] The following description and examples illustrate embodiments of the present disclosure in detail. It should be understood that this disclosure is not limited to the specific embodiments described herein and is modifiable. Those skilled in the art will recognize that there are numerous variations and modifications to this disclosure, which also fall within the scope of the invention. Various features of this disclosure can be described in the context of a single embodiment, but features can also be provided separately or in any suitable combination. Conversely, for clarity, this disclosure may be described in the context of separate embodiments, but it can also be implemented in a single embodiment.

[0099] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art.

[0100] The terms used herein are for the purpose of describing specific embodiments only and are not intended to be limiting. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context otherwise explicitly indicates. Furthermore, as used herein, the terms “including,” “includes,” “having,” “has,” and “with,” or variations thereof, mean “comprising.”

[0101] Unless otherwise specified, the terms “comprising,” “comprise,” “comprises,” “having,” “have,” “has,” “including,” “includes,” “include,” “containing,” “contains,” and “contain” are inclusive or unrestricted and do not exclude additional, unlisted elements or method steps.

[0102] References to “several embodiments,” “embodiments,” “one embodiment,” or “other embodiments” mean that certain features or characteristics described in relation to an embodiment are included in at least one embodiment of this disclosure, but not necessarily in all embodiments.

[0103] The terms "approximately" or "about" in relation to numbers refer to a range of values ​​that fall within or within 10% of the given value. For example, approximately x means x ± (10% * x).

[0104] In some embodiments, the cells are human cells. The cells may be derived from different tissues, organs, and / or cell types. In some embodiments, the cells are primary cells. As used herein, the term “primary cells” means cells isolated from an organism such as a mammal that are first grown in tissue culture (i.e., in vitro) before being divided and transferred to subculture. In some non-limiting embodiments, mammalian cells, including primary cells and stem cells, may be modified by the introduction of one or more polynucleotides, polypeptides, and / or prime-edited compositions (e.g., via transfection, transduction, electroporation, etc.) and further subcultured.

[0105] Examples of such modified cells include T cells, such as primary T cells, inflammatory T cells, T helper cells, cytotoxic T cells, CD4+ T cells, CD8+ T cells, memory T cells, regulatory T cells, natural killer T cells, mucosa-associated invariant T cells, γδ T cells, alpha-beta T cells, naive T cells, or effector T cells, hematopoietic elements (e.g., lymphocytes, bone marrow cells), their precursors or progenitors, their differentiated or dedifferentiated cells, and stem cells. In some embodiments, the cells are naive T cells (e.g., naive CD8+ T cells). In some embodiments, the cells are transformed T cells. In some embodiments, the cells are immune cells (e.g., primary immune cells) or their progenitors or precursors. In some embodiments, the cells are T cells, or their progenitors or precursors. In some embodiments, the cells are human T cells, or their progenitors or precursors. In some embodiments, the cells are T helper cells (e.g., Th1 cells, Th2 cells, Th9 cells, Th17 cells, Th22 cells, and Tfh (follicular helper) cells). In some embodiments, the cells are cytotoxic T cells. In some embodiments, the cells are CD8+ T cells. In some embodiments, the cells are CD4+ T cells. In some embodiments, the cells are memory T cells (e.g., central memory T cells (T)). CM These include stem cell memory T cells (TSCMs), effector memory T cells, and tissue-resident memory T cells. In some embodiments, the cells are effector memory T cells (e.g., T EM Cells and T EMR A(CD45RA +In some embodiments, the cells are regulatory T cells. In some embodiments, the cells are natural killer T cells. In some embodiments, the cells are mucosa-associated invariant T cells. In some embodiments, the cells are γδ T cells. In some embodiments, the cells are effector T cells. In some embodiments, the cells are thymocytes. In some embodiments, the cells are lymphocytes. In some embodiments, the cells are common lymphoid cells. In some embodiments, the cells are early thymic cells. In some embodiments, the cells are CD3+ cells. In some embodiments, the cells are tumor-infiltrating lymphocytes. In some embodiments, the cells are myeloid cells. In some embodiments, the cells are plasma cells. In some embodiments, the cells are activated T cells.

[0106] In some embodiments, the cells are stem cells (e.g., adult stem cells, embryonic stem cells, non-embryonic stem cells), umbilical cord blood stem cells, primordial cells, bone marrow stem cells, induced pluripotent stem cells, totipotent stem cells, CD34+ cells, or hematopoietic stem cells). In some embodiments, the cells are pluripotent cells (e.g., pluripotent stem cells). In some embodiments, the cells (e.g., stem cells) are embryonic stem cells, tissue-specific stem cells, mesenchymal stem cells, or induced pluripotent stem cells. In some embodiments, the cells are induced pluripotent stem cells (iPSCs). In some embodiments, the cells are hematopoietic stem cells. In some embodiments, the cells are hematopoietic stem and progenitor cells. In some embodiments, the cells are pluripotent progenitor cells. In some embodiments, the cells are T cell primordia. In some embodiments, the cells are T cell precursors. In some embodiments, the cells are embryonic stem cells (ESCs). In some embodiments, the cells are human stem cells. In some embodiments, the cells are human pluripotent stem cells. In some embodiments, the cells are non-embryonic stem cells. In some embodiments, the cells are induced human pluripotent stem cells. In some embodiments, the cells are human stem cells. In some embodiments, the cells are human embryonic stem cells. In some embodiments, the cells are human T cell primordia. In some embodiments, the cells are human T cell precursors.

[0107] In some embodiments, the cells are not isolated from an organism but form part of the tissue or organ of an organism, such as a mammal.

[0108] In some embodiments, the cells are differentiated cells. In some embodiments, the cells are differentiated from induced pluripotent stem cells. In some embodiments, the cells are T cells differentiated from iPSCs, ESCs, T cell precursors, or T cell primordia, such as primary T cells, such as inflammatory T cells, T helper cells, cytotoxic T cells, CD4+ T cells, CD8+ T cells, memory T cells, regulatory T cells, natural killer T cells, mucosa-associated invariant T cells, γδ T cells, alpha-beta T cells, naive T cells, or effector T cells.

[0109] In some embodiments, the cells are differentiated human cells. In some embodiments, the cells are differentiated from induced human pluripotent stem cells. In some embodiments, the cells edited by prime editing may differentiate into, or regenerate, cells such as T cells, e.g., primary T cells, e.g., inflammatory T cells, T helper cells, cytotoxic T cells, CD4+ T cells, CD8+ T cells, memory T cells, regulatory T cells, natural killer T cells, mucosa-associated invariant T cells, γδ T cells, alpha-beta T cells, naive T cells, or effector T cell populations. In some embodiments, the cells are in a subject, e.g., a human subject. In some embodiments, the cells are obtained from the subject prior to editing. For example, in some embodiments, the cells are obtained from patients with cancer, microbial infection, graft-versus-host disease, or autoimmune disorders. Before editing by the methods and compositions disclosed herein, the cells may be obtained from the subject by a variety of non-limiting methods. T cells can be obtained from several, not limited, sources, such as peripheral blood mononuclear cells, bone marrow, lymph node tissue, umbilical cord blood, thymic tissue, tissue from infection sites, ascites, pleural fluid, spleen tissue, and tumors. Methods for collecting blood cells, isolating and concentrating T cells, and amplifying them ex vivo may be by methods known in the art. In some embodiments, cells can be obtained from cell banks, blood banks, cell cultures, or any number of available T cell lines, as is known to those skilled in the art. Cells can also be obtained from tissue biopsies, surgeries, blood, plasma, serum, or other biological fluids. In some embodiments, cells can be obtained from one or more healthy donors, patients with cancer, microbial infection, graft-versus-host infection, or autoimmune disorders, prior to editing.

[0110] For example, cells can be obtained (i.e., isolated or purified) from a whole blood sample by lysing the red blood cell or fractionated blood sample and removing peripheral mononuclear blood cells by centrifugation before editing. Cells can further be isolated or purified using selective purification methods (e.g., by flow cytometry) that isolate cells based on cell-specific markers such as CD25, CD3, CD4, CD8, CD28, CD45RA, or CD45RO. In one embodiment, CD4+ is used as a marker for selecting T cells. In one embodiment, CD8+ is used as a marker for selecting T cells. In one embodiment, CD4+ and CD8+ are used as markers for selecting regulatory T cells.

[0111] In some embodiments, edited cells produced using the methods and compositions disclosed herein are cultured, proliferated, amplified, differentiated, and / or dedifferentiated in vitro.

[0112] In some embodiments, the cells comprise a prime editor, PEgRNA, or prime editing composition disclosed herein. In some embodiments, the cells further comprise ngRNA. In some embodiments, the cells are derived from a human subject. In some embodiments, the cells are derived from a human subject and comprise a prime editor, PEgRNA, or prime editing composition for editing the B2M gene. In some embodiments, the cells are derived from a human subject and the B2M gene has been edited by prime editing. In some embodiments, the human subject is a healthy donor. In some embodiments, the human subject has a disease, disorder, or condition, such as cancer, a microbial infection, or an autoimmune disorder. In some embodiments, the human subject needs, is receiving, or will receive immunotherapy (e.g., T-cell therapy such as CAR-T cell therapy).

[0113] As used herein, the term “substantially” may refer to a value close to 100% of a given value. In some embodiments, the term may refer to an amount that is at least about 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 99.99% of the total amount. In some embodiments, the term may refer to an amount that is equivalent to about 100% of the total amount.

[0114] The terms “protein” and “polypeptide” can be used interchangeably to refer to polymers of two or more amino acids linked by covalent bonds (such as amide bonds) that can take on a three-dimensional structure. In some embodiments, a protein or polypeptide contains at least 10, 15, 20, 30, or 50 amino acids linked by covalent bonds (e.g., amide bonds). In some embodiments, a protein contains at least two amide bonds. In some embodiments, a protein contains multiple amide bonds. In some embodiments, a protein includes enzymes, enzyme precursor proteins, regulatory proteins, structural proteins, receptors, nucleic acid-binding proteins, biomarkers, members of specific binding pairs (e.g., ligands or aptamers), or antibodies. In some embodiments, a protein may be a full-length protein (e.g., a fully processed protein with a specific biological function). In some embodiments, a protein may be a variant or fragment of a full-length protein. For example, in some embodiments, the Cas9 protein domain contains the H840A amino acid substitution compared to the naturally occurring S. pyogenes Cas9 protein. A variant of a protein or enzyme, for example, a variant reverse transcriptase, comprises a polypeptide having an amino acid sequence that is approximately 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% identical to the amino acid sequence of a reference protein.

[0115] In some embodiments, a protein comprises one or more protein domains or subdomains. As used herein, the terms “polypeptide domain,” “protein domain,” or “domain,” as used in the context of a protein or polypeptide, refer to a polypeptide chain having one or more biological functions, such as catalytic function, protein-protein binding function, or protein-DNA function. In some embodiments, a protein comprises multiple protein domains. In some embodiments, a protein comprises multiple spontaneously occurring protein domains. In some embodiments, a protein comprises multiple protein domains derived from different spontaneously occurring proteins. For example, in some embodiments, prime editor may be a fusion protein comprising the Cas9 protein domain of S. pyogenes and the reverse transcriptase protein domain of a retrovirus (e.g., Moloney's mouse leukemia virus) or a variant of said retrovirus. A protein comprising amino acid sequences derived from proteins of various origins or spontaneously occurring proteins may be called a fusion protein or chimeric protein.

[0116] In some embodiments, the protein comprises a functional variant or functional fragment of a full-length wild-type protein. As used herein, “functional fragment” or “functional portion” refers to any portion of a reference protein (e.g., a wild-type protein) that contains fewer amino acids than the entire amino acid sequence of the reference protein, while retaining one or more functions, such as catalytic function or binding function. For example, a functional fragment of reverse transcriptase may contain fewer amino acids than the entire amino acid sequence of wild-type reverse transcriptase, but retain the ability to catalyze polynucleotide polymerization under conditions of at least one set. If the reference protein is a fusion of multiple functional domains, its functional fragment may retain one or more functions of at least one functional domain. For example, a functional fragment of Cas9 may contain less than the entire amino acid sequence of wild-type Cas9, but retain DNA binding ability and partially or completely lack nuclease activity.

[0117] As used herein, “functional variant” or “functional variant” refers to any variant or variant of a reference protein (e.g., wild-type protein) that includes one or more changes to the amino acid sequence of the reference protein while retaining one or more functions, such as catalytic function or binding function. In some embodiments, one or more changes to the amino acid sequence include amino acid substitutions, insertions, deletions, or any combination thereof. In some embodiments, one or more changes to the amino acid sequence include amino acid substitutions. For example, a functional variant of reverse transcriptase may include one or more amino acid substitutions compared to the amino acid sequence of wild-type reverse transcriptase, but retain the ability to catalyze polynucleotide polymerization under at least one set of conditions. If the reference protein is a fusion of multiple functional domains, its functional variant may retain one or more functions of at least one functional domain. For example, in some embodiments, a functional fragment of Cas9 may contain one or more amino acid substitutions (e.g., H840A amino acid substitution) in the nuclease domain compared to the amino acid sequence of wild-type Cas9, but retain DNA binding ability and partially or completely lack nuclease activity.

[0118] As used herein, the term “function” and its grammatical synonyms can refer to the ability to perform, have, or accomplish an intended purpose. Function can include any percentage from the baseline up to 100% of the intended purpose. For example, function can include or include about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or up to about 100% of the intended purpose. In some embodiments, the term “function” can mean more than about 100% of the function, for example, 125%, 150%, 175%, 200%, 250%, 300%, 400%, 500%, 600%, 700%, or up to about 1000% of the intended purpose.

[0119] In some embodiments, the protein or polypeptide contains a spontaneously occurring amino acid (e.g., one of the 20 amino acids commonly found in naturally synthesized peptides, known by the single-letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, V). In some embodiments, the protein or polypeptide contains a non-spontaneously occurring amino acid (e.g., an amino acid other than one of the 20 amino acids commonly found in naturally synthesized peptides, including synthetic amino acids, amino acid analogs, and amino acid mimes). In some embodiments, the protein or polypeptide is modified.

[0120] In some embodiments, the protein comprises isolated polypeptides. The term “isolated” means that components normally present in the natural state or environment have been removed to varying degrees or are free. For example, polypeptides naturally present in living animals are not isolated; the same polypeptides are isolated after being partially or completely separated from their naturally occurring coexisting substances.

[0121] In some embodiments, the protein is present within cells, tissues, organs, or viral particles. In some embodiments, the protein is present within cells or in parts of cells (e.g., bacterial cells, plant cells, animal cells). In some embodiments, the cells are present within tissues, subjects, or cell cultures. In some embodiments, the cells are microorganisms (e.g., bacteria, fungi, protists, viruses). In some embodiments, the protein is present in a mixture of analytes (e.g., lysates). In some embodiments, the protein is present in lysates from multiple cells or a single cell.

[0122] As used herein, the terms “homologous,” “homonymy,” or “percentage of homology” refer to the degree of sequence identity between an amino acid and its corresponding reference amino acid sequence, or between a polynucleotide sequence and its corresponding reference polynucleotide sequence. “Homologous” can refer to polymer sequences such as similar polypeptide or DNA sequences. Homologous can mean, for example, nucleic acid sequences having at least approximately 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. In other embodiments, a “homologous sequence” of a nucleic acid sequence may exhibit 93%, 95%, or 98% sequence identity with respect to a reference nucleic acid sequence. For example, a "region homologous to a genomic region" can be a region of DNA that has a sequence similar to a given genomic region within the genome. The homologous region can be of any length sufficient to facilitate the binding of a spacer, primer-binding site, or protospacer sequence to the genomic region. For example, the homologous region may be at least 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100 Since homology regions can contain base lengths of 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, or more, homology regions have sufficient homology to bind to the corresponding genomic region.

[0123] In the context of two nucleic acid sequences or two polypeptide sequences, when a percentage of sequence homology or identity is specified, the percentage of homology or identity usually refers to the alignment of sequences over a portion of their length when two or more sequences are compared and aligned to obtain the greatest possible correspondence. Molecules may be homologous at a position if that position in the sequences being compared is occupied by the same base or amino acid. Unless otherwise specified, sequence homology or identity is evaluated over a specified length of nucleic acid, polypeptide, or a portion thereof. In some embodiments, homology or identity is evaluated over a functional portion or a specified portion of the length.

[0124] Sequence alignment for sequence homology evaluation can be performed using algorithms known in the art, such as the Basic Local Alignment Search Tool (BLAST) algorithm, described in Altschul et al, J.Mol.Biol.215:403-410, 1990. A public internet interface for performing BLAST analysis is accessible through the National Center for Biotechnology Information. Additional known algorithms include Smith & Waterman, “Comparison of Biosequences”, Adv. Appl. Math.2:482, 1981; Needleman & Wunsch, “A general method applicable to the search for similarities in the amino acid sequence of two proteins”, J.Mol.Biol.48:443, 1970; Pearson & Lipman, “Improved tools for biological sequence comparison”, Proc.Natl.Acad.Sci.USA 85:2444, 1988, or by automatically implementing these or similar algorithms. Global alignment programs can also be used to align similar sequences of nearly the same size. Examples of global alignment programs include NEEDLE (available at www.ebi.ac.uk / Tools / psa / emboss_needle / ), which is part of the EMBOSS package (Rice P et al., Trends Genet., 2000;16:276-277), and the GGSEARCH program (https: / / fasta.bioch.virginia.edu / fasta_www2 / ), which is part of the FASTA package (Pearson W and Lipman D, 1988, Proc.Natl.Acad.Sci.USA, 85:2444-2448). Both of these programs are based on the Needleman-Wunsch algorithm, which is used to find the best alignment (including gaps) along the entire length of two sequences.A detailed discussion of sequence analysis can also be found in Unit 19.3 of Ausubel et al. ("Current Protocols in Molecular Biology," John Wiley & Sons Inc, 1994-1998, Chapter 15, 1998). Unless otherwise specified, alignment between the query sequence and the reference sequence is performed using the Needleman-Wunsch alignment with a gap presence cost of 11 and a gap extension cost of 1, where the identity percentage is calculated by dividing the number of identities by the length of the alignment, which is described in more detail in Altschul et al. ("Gapped BLAST and PSI-BLAST: a new generation of protein database search programs," Nucleic Acids Res. 25:3389-3402, 1997) and Altschul et al. ("Protein database searches using compositionally adjusted substitution matrices," FEBS J. 272:5101-5109, 2005).

[0125] Those skilled in the art understand that the positions of amino acids (or nucleotides) in homologous sequences can be determined based on alignment, for example, "H840" in a reference Cas9 sequence may correspond to H839 or another position in a Cas9 homolog.

[0126] The terms “polynucleotide” or “nucleic acid molecule” can refer to any polymeric form of nucleotides, including DNA, RNA, their hybridizations, or RNA-DNA chimeric molecules. In some embodiments, polynucleotides include cDNA, genomic DNA, mRNA, tRNA, rRNA, or microRNA. In some embodiments, polynucleotides are double-stranded, such as double-stranded DNA in a gene. In some embodiments, polynucleotides are single-stranded or substantially single-stranded, such as single-stranded DNA or mRNA. In some embodiments, polynucleotides are cell-free nucleic acid molecules. In some embodiments, polynucleotides circulate in the blood. In some embodiments, polynucleotides are intracellular nucleic acid molecules. In some embodiments, polynucleotides are intracellular cellular nucleic acid molecules circulating in the blood.

[0127] Polynucleotides can have any three-dimensional structure. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, EST or SAGE tags), exons, introns, intergenic DNA (including but not limited to heterochromatin DNA), messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA, isolated RNA, sgRNA, guide RNA, nucleic acid probes, primers, snRNA, long non-coding RNA, snoRNA, siRNA, miRNA, small RNA derived from tRNA (tsRNA), antisense RNA, shRNA, or small RNA derived from rDNA (srRNA).

[0128] In some embodiments, the polynucleotide comprises deoxyribonucleotides, ribonucleotides, or analogs thereof. In some embodiments, the polynucleotide comprises modified nucleotides, such as methylated nucleotides or nucleotide analogs. Where present, modifications to the nucleotide structure can be conferred before or after the construction of the polynucleotide. The nucleotide sequence may be interrupted by non-nucleotide components. After polymerization, the polynucleotide may be further modified, such as by binding with labeling components.

[0129] In some embodiments, the polynucleotide consists of a specific sequence of four nucleotide bases: adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) instead of thymine when the polynucleotide is RNA. In some embodiments, the polynucleotide may contain one or more other nucleotide bases, such as inosine (I), which is read as guanine (G) by the translation mechanism.

[0130] In some embodiments, polynucleotides can be modified. As used herein, the terms “modified” or “modified” refer to chemical modifications to A, C, G, T, and U nucleotides, denoted as mA, mC, mG, mT, and mT. In some embodiments, modifications may be made to the nucleoside bases and / or sugar moieties of the nucleosides constituting the polynucleotide. In some embodiments, modifications may be made on internucleoside bonds (e.g., the phosphate backbone). In some embodiments, a modified nucleic acid molecule may contain multiple modifications. In some embodiments, a modified nucleic acid molecule may contain a single modification.

[0131] As used herein, the terms “complementary,” “complementary,” or “complementarity” refer to the ability of two polynucleotide molecules to form base pairs with each other. Complementary polynucleotides can form base pairs via hydrogen bonds, which may be Watson-Crick, Hoogsteen, or inverse Hoogsteen hydrogen bonds. For example, adenine on one polynucleotide molecule can base pair with thymine or uracil on a second polynucleotide molecule, and cytosine on one polynucleotide molecule can base pair with guanine on a second polynucleotide molecule. Two polynucleotide molecules are complementary if a first polynucleotide molecule containing a first nucleotide sequence can base pair with a second polynucleotide molecule containing a second nucleotide sequence. For example, two DNA molecules 5'-ATGC-3' and 5'-GCAT-3' are complementary, and the complement of the DNA molecule 5'-ATGC-3' is 5'-GCAT-3'. The percentage of complementarity indicates the percentage of nucleotides in a polynucleotide molecule that can form base pairs with the second polynucleotide molecule (for example, 5, 6, 7, 8, 9, and 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). "Fully complementary" means that all consecutive nucleotides in the polynucleotide molecule form base pairs with the same number of consecutive nucleotides in the second polynucleotide molecule. As used herein, "substantially complementary" refers to a degree of complementarity that can be 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% across all or part of the two polynucleotide molecules. In some embodiments, the portion of the complementarity may be a region of 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides. "Substantial complementarity" can refer to 100% complementarity across a portion or region of two polynucleotide molecules.In some embodiments, part or a region of complementarity between two polynucleotide molecules is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% of the length of at least one of the two polynucleotide molecules or its functional or defined portion.

[0132] As used herein, “expression” refers to the process by which a polynucleotide is transcribed into mRNA, and / or the process by which a polynucleotide (e.g., transcribed mRNA) is translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression may include splicing of mRNA in eukaryotic cells. In some embodiments, the expression of a polynucleotide, e.g., DNA encoding a gene or protein, is determined by the amount of protein encoded by the gene after transcription and translation. In some embodiments, the expression of a polynucleotide, e.g., DNA encoding a gene or protein, is determined by the amount of the functional form of the protein encoded by the gene after transcription and translation. In some embodiments, gene expression is determined by the amount of mRNA encoded by the gene after transcription, i.e., the transcript. In some embodiments, the expression of a polynucleotide, e.g., mRNA, is determined by the amount of protein encoded by the mRNA after translation. In some embodiments, the expression of a polynucleotide, e.g., mRNA or coding RNA, is determined by the amount of the functional form of the protein encoded by the polypeptide after translation of the polynucleotide.

[0133] As used herein, the term “sequencing” may include capillary sequencing, bisulfite-free sequencing, bisulfite sequencing, TET-assisted bisulfite (TAB) sequencing, ACE sequencing, high-throughput sequencing, Maxam-Gilbert sequencing, ultra-parallel signature sequencing, Polony sequencing, 454 pyrosequencing, Sanger sequencing, Illumina sequencing, SOLiD sequencing, ion torrent semiconductor sequencing, DNA nanoball sequencing, heliscope single-molecule sequencing, single-molecule real-time (SMRT) sequencing, nanopore sequencing, shotgun sequencing, RNA sequencing, or any combination thereof.

[0134] The terms "equivalent" or "biologically equivalent" are used interchangeably when referring to specific molecules or biological or cellular materials, meaning molecules that have minimal homology to another molecule while maintaining a desired structure or function.

[0135] The term “encodes” as applied to polynucleotides refers to a polynucleotide said to “encode” another polynucleotide, polypeptide, or amino acid if, in its natural state or when manipulated by methods well known to those skilled in the art, it can be used as a polynucleotide synthesis template, for example, transcribed into RNA, reverse transcribed into DNA or cDNA, and / or translated to produce amino acids, polypeptides or fragments thereof. In some embodiments, a polynucleotide consisting of three consecutive nucleotides forms a codon that encodes a particular amino acid. In some embodiments, a polynucleotide contains one or more codons that encode a polypeptide. In some embodiments, a polynucleotide containing one or more codons contains a mutation in the codon compared to a wild-type reference polynucleotide. In some embodiments, the codon mutation encodes an amino acid substitution in the polypeptide encoded by the polynucleotide compared to a wild-type reference polypeptide. In some embodiments, a polynucleotide encodes another polynucleotide that contains one or more desired nucleotide edits introduced into target DNA. For example, in some embodiments, the PEgRNA editing template encodes a single-stranded DNA containing one or more nucleotide changes compared to an endogenous editing target DNA sequence within a target gene, e.g., a target B2M gene, the single-stranded DNA being otherwise identical to the editing target sequence. Thus, in some embodiments, the editing template "encodes" one or more nucleotide changes. When used herein with respect to a specific nucleotide change encoded by a PEgRNA editing template, unless otherwise indicated, the specific nucleotide change encoded refers to a change introduced into the coding strand (sense strand) of the target gene, the editing target sequence may be located on either the sense strand or the antisense strand of the target gene, e.g., a B2M gene.For example, if PEgRNA mediates prime editing that results in the insertion of a TAATAA sequence into the sense strand of a target gene, the editing template encodes the TAATAA insertion, but the single-stranded DNA synthesized using the editing template sequence as a template may contain TTATTA at the corresponding position and has sequence identity with respect to the antisense strand.

[0136] As used herein, the term “mutation” refers to a change and / or alteration of the amino acid sequence of a protein or the nucleic acid sequence of a polynucleotide. Such changes and / or alterations may include substitutions, insertions, deletions and / or cleavages of one or more amino acids in the case of an amino acid sequence, and / or nucleotides in the case of a nucleic acid sequence, compared to a reference amino acid or reference nucleic acid sequence. In some embodiments, the reference sequence is a wild-type sequence. In some embodiments, a mutation in the nucleic acid sequence of a polynucleotide encodes a mutation in the amino acid sequence of a polypeptide. In some embodiments, a mutation in the amino acid sequence of a polypeptide or a mutation in the nucleic acid sequence of a polynucleotide is a mutation associated with a disease condition.

[0137] As used herein, the term “subject” and its grammatical synonyms may refer to a human or a non-human. A subject may be a mammal. A human subject may be male or female. A human subject may be of any age. A subject may be a human fetus. A human subject may be a newborn, infant, child, adolescent, or adult. A human subject may require treatment for a genetic disorder or disability. A human subject may require cell therapy, such as immunotherapy (e.g., T-cell therapy). A human subject may require CAR-T cell therapy. Alternatively, a human subject may be a healthy donor.

[0138] The terms “treatment” or “treating,” and their grammatical equivalents, may refer to the medical management of an object for the purpose of curing, improving, or improving the symptoms of a disease, condition, or disorder. Treatment may include active treatment, i.e., treatment specifically focused on improving a disease, condition, or disorder. Treatment may include causal treatment, i.e., treatment aimed at eliminating the cause of the associated disease, condition, or disorder. Furthermore, treatment may include palliative treatment, which aims at alleviating symptoms rather than curing a disease, condition, or disorder. Treatment may include supportive care, i.e., treatment used to complement another specific therapy aimed at improving a disease, condition, or disorder. In some embodiments, the condition may be pathological. In some embodiments, treatment may not completely cure or prevent a disease, condition, or disorder. In some embodiments, treatment may improve a disease, condition, or disorder, but not completely cure or prevent it. In some embodiments, the subject can receive treatment for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or for the lifetime of the subject.

[0139] The term "improvement" and its grammatical synonyms mean reducing, suppressing, attenuating, reducing, preventing, or stabilizing the onset or progression of a disease.

[0140] The terms “prevent” or “preventing” mean delaying, preventing, or avoiding the onset or progression of a disease, condition, or disorder for a certain period of time. Prevention also means reducing the risk of developing a disease, disorder, or condition. Prevention includes minimizing, partially or completely inhibiting, the onset of a disease, condition, or disorder. In some embodiments, a composition, for example, a pharmaceutical composition, prevents a disorder by delaying the onset of the disease for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or over the lifetime of the subject.

[0141] The terms “effective dose” or “therapeutic effective dose” refer to the amount of a composition, such as a prime editing composition, that, when introduced into a target, is sufficient to produce the desired activity, as disclosed herein. Whether the cells are exovivo or in vivo, an effective dose of the prime editing composition can be delivered to the target gene or cell.

[0142] An effective dose may be, for example, a dose that induces at least a twofold change (increase or decrease) in the observed regulatory level of the target nucleic acid (e.g., the expression of a target gene for producing a functional protein) compared to a negative control. An effective dose or dosage may induce, for example, a twofold decrease, a threefold decrease, a fourfold decrease, a fivefold decrease, a sixfold decrease, a sevenfold decrease, an eightfold decrease, a ninefold decrease, a tenfold decrease, a twenty-fivefold decrease, a fiftyfold decrease, or a hundredfold decrease in the regulation of a target gene (e.g., the expression of a target B2M gene for producing a functional β chain of MHC class I).

[0143] The amount of modification of the target gene can be measured by any suitable method known in the art. In some embodiments, the “effective dose” or “therapeutic effective dose” is the amount of composition required to improve the symptoms of the disease compared to an untreated patient. In some embodiments, the effective dose is the amount of composition sufficient to introduce a change into the gene of interest (e.g., the B2M gene) within a cell (e.g., an in vitro or in vivo cell).

[0144] In some embodiments, an effective dose may be an amount that, when administered to a population of cells, induces a certain percentage of that population to have edits in the target gene (e.g., the B2M gene). For example, in some embodiments, an effective dose may be an amount that, when administered to or introduced to a population of cells, induces the introduction of one or more intended nucleotide edits in the B2M gene in at least about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 99% of that population of cells.

[0145] With respect to edited cells produced by the methods disclosed herein and compositions containing such edited cells, “therapeutic dose” means an amount of the composition containing the edited cells that is sufficient to produce the desired activity upon introduction into a subject (e.g., a human subject).

[0146] As used herein, the term “recombinase” refers to a site-specific enzyme that mediates the recombination of DNA between recombinase-recognized sequences, resulting in the excision, integration, inversion, or exchange (e.g., transposition) of DNA fragments between recombinase-recognized sequences. Recombinases can be classified into two distinct families: serine recombinases (e.g., resolvers and invertases) and tyrosine recombinases (e.g., integrases). Examples of serine recombinases include, but are not limited to, Si74, No67, Kp03, Pa01, Nm60, BceINTa, BcytINTd, SscINTd, SacINTd, Hin, Gin, Tn3, I3-six, CinH, ParA, y6, Bxbl, OC31, TP901, TG1, pBT1, R4, pRV1, pFC1, MR11, A118, U153, and gp29. Examples of tyrosine recombinases include, but are not limited to, Cre, FLP, R, Lambda, HK101, HK022, and pSAM2. Recombinases have many applications, including gene knockout / knock-in generation and gene therapy applications, as described in International Publication No. WO2020191248A1, which is incorporated herein by reference in its entirety. The recombinases provided herein are not intended to be exclusive examples of recombinases that may be used in embodiments of the present invention. The methods and compositions of the present invention can be extended by mining databases of novel orthogonal recombinases or by designing synthetic recombinases with defined DNA specificity.

[0147] In some embodiments, the catalytic domain of a recombinase is fused to the programmable DNA-binding domain of a prime editor, such as an RNA programmable nuclease (e.g., dCas9, Cas9 nickase, or a fragment thereof), resulting in the recombinase domain either lacking a nucleic acid-binding domain or being unable to bind to a target nucleic acid (e.g., the recombinase domain is manipulated to lack specific DNA-binding activity). For example, serine recombinases of the resolverase-invertase group, such as Tn3 and 76 type resolvers and Hin and Gin invertases, have a modular structure with an autonomous catalytic domain and a DNA-binding domain. Therefore, the catalytic domains of these recombinases are suitable for being combined with a prime editor or its components, for example, by fusion or ligation. Furthermore, many other natural serine recombinases possessing an N-terminal catalytic domain and a C-terminal DNA-binding domain are known (e.g., phiC31 integrase, TnpX transposase, IS607 transposase), and their catalytic domains can be combined, for example, fused or complexed with a prime editor or its components for programmable recombination.

[0148] Other examples of recombinases useful in the methods and compositions described herein are known to those skilled in the art, and any new recombinases discovered or produced are expected to be usable in different embodiments of the present invention.

[0149] As used herein, the terms “recombinase-recognition sequence,” or equivalently “RRS,” “recombinase-target sequence,” or “recombinase site,” refer to a nucleotide sequence target recognized by a recombinase, which results in the excision, integration, inversion, or exchange of DNA fragments between the recombinase-recognition sequences via strand exchange with another DNA molecule having a corresponding RRS (e.g., an RRS recognized by the same recombinase). In various embodiments, a prime-edit composition may introduce one or more recombinase sites within a single target sequence or within multiple target sequences. When multiple recombinase sites are introduced by prime-editing, the recombinase sites may be introduced into adjacent or non-adjacent target sites (e.g., separate chromosomes). In various embodiments, a single introduced recombinase site may be used as a “landing site” for a recombinase-mediated reaction between an exogenously supplied nucleic acid molecule, e.g., a genomic recombinase site in a plasmid or DNA vector, and a second recombinase site. This enables the targeted integration of the desired nucleic acid molecule.

[0150] In the context of nucleic acid modification (e.g., genome modification), the terms “recombination” or “recombination” are used to refer to the process by which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can, among other things, result in nucleic acid insertion, inversion, excision, or transposition, for example, in or between one or more nucleic acid molecules.

[0151] Prime Edit The term “prime editing” refers to programmable editing of target DNA, which involves incorporating intended nucleotide edits (also known herein as nucleotide changes) into target DNA through targeted priming DNA synthesis using a prime editor complexed with PEgRNA. The target gene for prime editing may include a double-stranded DNA molecule having two complementary strands, the first strand may be called the “target strand” or “unedited strand,” and the second strand may be called the “untargeted strand” or “edited strand.” In some embodiments, the spacer sequence in the prime editing guide RNA (PEgRNA) is complementary or substantially complementary to a specific sequence on the target strand (which may also be called the “exploratory target sequence”). In some embodiments, the spacer sequence anneals with the target strand at the exploratory target sequence. The target strand may also be called the “non-protospacer adjacent motif (non-PAM strand).” In some embodiments, the non-target strand may also be called the “PAM strand.” In some embodiments, the PAM strand includes a protospacer sequence and optionally a protospacer adjacent motif (PAM) sequence. In prime editing using a Cas protein-based prime editor, the PAM sequence refers to a short DNA sequence immediately adjacent to a protospacer sequence on the PAM strand of the target gene. The PAM sequence can be specifically recognized by a programmable DNA-binding protein, such as Cas nickase or Cas nuclease. In some embodiments, a particular PAM is characteristic of a specific programmable DNA-binding protein, such as Cas nickase or Cas nuclease. The protospacer sequence refers to a specific sequence within the PAM strand of the target gene that is complementary to the search target sequence. In PEgRNA, the spacer sequence may be substantially identical to the protospacer sequence on the edited strand of the target gene, however, the spacer sequence may contain uracil (U) and the protospacer sequence may contain thymine (T).

[0152] In some embodiments, the double-stranded target DNA includes a nick site on the PAM strand (or non-target strand). As used herein, “nick site” refers to a specific location between two nucleotides or two base pairs on the double-stranded target DNA. In some embodiments, the location of the nick site is determined relative to a specific PAM sequence. In some embodiments, the nick site is a specific location where a nick occurs when a nickase (such as Cas nickase) that recognizes a specific PAM sequence comes into contact with the double-stranded target DNA. In some embodiments, the nick site is upstream of a specific PAM sequence on the PAM strand of the double-stranded target DNA. In some embodiments, the nick site is downstream of a specific PAM sequence on the PAM strand of the double-stranded target DNA. In some embodiments, the nick site is upstream of a PAM sequence recognized by Cas9 nickase, which includes a nuclease-active RuvC domain and a nuclease-inactive HNH domain. In some embodiments, the nick site is located 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by Cas9 nickase of Streptococcus pyogenes, P. lavamentivorans, C. diphtheriae, N. cinerea, S. aureus, or N. lari. In some embodiments, the nick site is 3 base pairs upstream of the PAM sequence, and the PAM sequence is recognized by Cas9 nickase, which contains a nuclease-active RuvC domain and a nuclease-inactive HNH domain. In some embodiments, the nick site is 2 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by Cas9 nickase of S. thermophilus, which contains a nuclease-active RuvC domain and a nuclease-inactive HNH domain.

[0153] A "primer-binding site" (also called PBS or primer-binding site sequence) is a single-stranded portion of PEgRNA containing a complementary region to the PAM strand (i.e., the non-target strand or edited strand). The PBS is complementary or substantially complementary to a sequence on the PAM strand of the double-stranded target DNA immediately upstream of the nick site. In some embodiments, during the prime editing process, the PEgRNA complexes with the prime editor, instructing the prime editor to bind to a search target sequence on the target strand of the double-stranded target DNA, thereby generating a nick at the nick site on the non-target strand of the double-stranded target DNA. In some embodiments, the PBS is complementary or substantially complementary to the free 3' end on the non-target strand of the double-stranded target DNA at the nick site and can anneal to that end. In some embodiments, the PBS annealed to the free 3' end of the non-target strand can initiate target priming DNA synthesis.

[0154] The “editing template” of PEgRNA is located at 5' of PBS and is a single-stranded portion of PEgRNA encoding a single strand of DNA. The editing template may include a complementary region to the PAM strand (i.e., the non-target or edited strand) and includes one or more intended nucleotide edits compared to the endogenous sequence of the double-stranded target DNA. In some embodiments, the editing template and PBS are directly adjacent to each other. Therefore, in some embodiments, PEgRNA during prime editing includes a single-stranded portion containing the PBS and editing template directly adjacent to each other. In some embodiments, the single-stranded portion of PEgRNA containing both PBS and editing template is complementary or substantially complementary to the endogenous sequence on the PAM strand (i.e., the non-target or edited strand) of the double-stranded target DNA, except for one or more non-complementary nucleotides at the intended nucleotide editing sites. As used herein, regardless of the relative 5'-3' configuration in other contexts, the relative positions between PBS and the editing template, and between elements of PEgRNA, are determined by the 5'-3' order of PEgRNA as a single molecule, regardless of the position of the sequence in the double-stranded target DNA that may have complementarity or identity with the elements of PEgRNA. In some embodiments, the editing template is complementary or substantially complementary to the sequence on the PAM strand immediately downstream of the nick site, except for one or more nucleotide changes (e.g., one or more non-complementary nucleotides) at the intended nucleotide editing site. An endogenous, e.g., genomic sequence that is complementary or substantially complementary to the editing template may be called the “editing target sequence,” except for one or more non-complementary nucleotides at the position corresponding to the intended nucleotide edit. In some embodiments, the editing template is complementary to the editing target sequence, or includes identity or substantial identity with the sequence on the target strand having the same genomic position, except for one or more nucleotide changes (e.g., one or more insertions, deletions, or substitutions) at the intended nucleotide editing site.

[0155] In some embodiments, the editing target sequence is located within the non-coding region of the target B2M gene. In some embodiments, the editing target sequence is located within the coding region of the target B2M gene. In some embodiments, the editing target sequence is located within the exon of the target B2M gene. In some embodiments, the editing target sequence is located within the exon, intron, exon-intron junction, or regulatory element of the B2M gene. In some embodiments, the editing target sequence is located within the open reading frame of the B2M gene. In some embodiments, the editing target sequence is located within the untranslated region of the B2M gene, e.g., the 3'-UTR or 5'-UTR. In some embodiments, the editing target sequence is located within the regulatory element of the B2M gene. In some embodiments, the editing target sequence is located within the promoter, enhancer, operator, silencer, insulator, terminator, transcription start sequence, translation start sequence (e.g., Kozak sequence), or any combination thereof of the B2M gene. In some embodiments, the editing target sequence is located within the splice acceptor-splice donor (SA-SD) site in the B2M gene.

[0156] In some embodiments, the editing template encodes single-stranded DNA that is identical or substantially identical to the editing target sequence, except for one or more nucleotide changes (e.g., one or more insertions, deletions, or substitutions) at one or more intended nucleotide editing sites. In some embodiments, the methods disclosed herein result in the incorporation of one or more nucleotide changes in the B2M gene. In some embodiments, the methods disclosed herein result in the introduction of a mutation into the B2M gene. In some embodiments, the incorporation of one or more nucleotide changes into the target B2M gene results in a mutation into the target B2M gene (e.g., a missense mutation, a nonsense mutation, a frameshift mutation, a null mutation, a mutation that produces an immature stop codon, or a combination thereof). In some embodiments, the incorporation of one or more nucleotide changes introduces a frameshift mutation into the target gene (e.g., the B2M gene) and / or generates one or more immature stop codons (e.g., at least 1, 2, 3, 4, 5, or more immature stop codons). In some embodiments, the incorporation of one or more nucleotide changes generates at least two immature stop codons in the target gene. In some embodiments, the incorporation of one or more nucleotide changes generates at least two, three, four, five, or more consecutive immature stop codons in the target gene (e.g., the B2M gene). In some embodiments, the immature stop codons are generated in exon 1, exon 2, exon 3, exon 4, or a combination thereof, of the B2M gene. An “immature stop codon” is a mutation (e.g., a nonsense mutation or insertion) within the sequence of the target gene (e.g., the B2M gene) that generates a stop codon at a location not normally found in the wild-type gene (e.g., the wild-type B2M gene). Immature stop codons may result in a truncated and / or non-functional protein compared to the full-length protein encoded by the corresponding wild-type target gene. As used herein, a “frameshift mutation” means a mutation in which the reading frame of a codon in a coding region is altered by one or more nucleotide changes.For example, in some embodiments, a frameshift mutation is an insertion of 3x+1 or 3x+2 nucleotides in the coding region, where x is a non-negative integer. In some embodiments, a frameshift mutation is a deletion of 3x+1 or 3x+2 nucleotides in the coding region, where x is a non-negative integer. In some embodiments, a frameshift mutation results in an immature stop codon in the target gene. In some embodiments, a frameshift mutation results in a non-functional protein encoded by the target gene. As used herein, “missense mutation” refers to a change in the type of amino acid in the protein expressed by the target gene, resulting from a change or substitution of bases in the corresponding target gene (e.g., the B2M gene). As used herein, “nonsense mutation” refers to a mutation in which a sense codon encoding an amino acid is changed to a stop codon. As used herein, “null mutation” refers to a mutation in the target gene (e.g., the B2M gene) that results in a complete loss of functional protein expression from the mutated target gene.

[0157] In some embodiments, the mutation is introduced into the non-coding region of the target B2M gene. In some embodiments, the mutation is introduced into the coding region of the target B2M gene. In some embodiments, the mutation is introduced into the exon of the target B2M gene. In some embodiments, the mutation is introduced into exon 1, exon 2, exon 3, exon 4, or any combination thereof of the target B2M gene. In some embodiments, the mutation is introduced into an exon, intron, exon-intron junction, or regulatory element of the B2M gene. In some embodiments, the mutation is introduced into the open reading frame of the B2M gene. In some embodiments, the mutation is introduced into the untranslated region of the B2M gene, for example, the 3'-UTR or 5'-UTR. In some embodiments, the mutation is introduced into a regulatory element of the B2M gene. In some embodiments, the mutation is introduced into the promoter, enhancer, operator, silencer, insulator, terminator, transcription start sequence, translation start sequence (e.g., Kozak sequence), or any combination thereof of the B2M gene. In some embodiments, the mutation is introduced into a splice acceptor-splice donor (SA-SD) site in the B2M gene. In some embodiments, the mutation is introduced into the B2M gene, and the mutation generates a splice acceptor-splice donor (SA-SD) site in the B2M gene. In some embodiments, the mutation is introduced into a splice site in the B2M gene. In some embodiments, the mutation is introduced into a splice site in the B2M gene, and the splice site is interrupted. In some embodiments, the method disclosed herein generates one of the following edits in the B2M gene to generate a stop codon: CAG to TAG, CAA to TAA, CGA to TGA, TGG to TGA, TGG to TAG, or TGG to TAA. In some embodiments, one or more intended nucleotide edits may be introduced at the 3'-UTR, for example, at a polyadenylation (poly-A) site. In some embodiments, one or more immature in-frame stop codons (e.g., two stop codons) are inserted into the B2M gene.

[0158] In some embodiments, the editing template may encode a wild-type or non-disease-related gene sequence (or its complement if the edited strand is an antisense strand of the gene). In some embodiments, the editing template may encode a wild-type or non-disease-related protein, but may contain one or more synonymous mutations compared to the coding region of the wild-type or non-disease-related protein. Such synonymous mutations may include, for example, mutations that, once the desired edit is introduced into the genome, reduce the ability of PEgRNA to rejoin the same target sequence (e.g., a synonymous mutation that silences an endogenous PAM sequence or a synonymous mutation that edits an endogenous protospacer). As used herein, synonymous editing, synonymous alteration, or synonymous mutation in a gene refers to a nucleotide change that does not alter the protein sequence or mRNA sequence encoded by the gene. Non-synonymous editing, non-synonymous alteration, or non-synonymous mutation in a gene refers to a nucleotide change that results in a change in the protein sequence or mRNA sequence encoded by the gene (e.g., by altering a splice donor or acceptor sequence).

[0159] In some embodiments, PEgRNA forms a complex with a prime editor, instructing the prime editor to bind to a target sequence of the target gene. In some embodiments, the bound prime editor generates a nick on the edited strand (PAM strand) of the target gene at the nick site. In some embodiments, the primer-binding site (PBS) of PEgRNA anneals with the free 3' end formed at the nick site, and the prime editor uses the free 3' end as a primer to initiate DNA synthesis from the nick site. Subsequently, single-stranded DNA encoded by the PEgRNA editing template is synthesized. In some embodiments, the newly synthesized single-stranded DNA contains one or more intended nucleotide edits (i.e., one or more nucleotide changes) compared to the endogenous target gene sequence. In some embodiments, the incorporation of one or more intended nucleotide edits into the target B2M gene results in a mutation in the target B2M gene (e.g., a missense mutation, a nonsense mutation, a frameshift mutation, a null mutation, a mutation producing an immature stop codon, or a combination thereof). In some embodiments, one or more intended nucleotide edits introduce a frameshift mutation into a target gene (e.g., the B2M gene) and / or generate one or more immature stop codons (e.g., at least 1, 2, 3, 4, 5, or more immature stop codons). In some embodiments, one or more intended nucleotide edits generate at least two immature stop codons in the target gene. In some embodiments, one or more intended nucleotide edits generate at least 2, 3, 4, 5, or more consecutive immature stop codons in the target gene (e.g., the B2M gene). In some embodiments, one or more intended nucleotide edits involve the insertion of one or more immature in-frame stop codons (e.g., two stop codons) into the B2M gene. Thus, in some embodiments, the PEgRNA editing template is complementary to the sequence in the edited strand, except for one or more mismatches at the intended nucleotide editing sites within the editing template.An endogenous sequence, such as a genomic sequence, that is partially complementary to the editing template may be called the “editing target sequence.” Thus, in some embodiments, the newly synthesized single-stranded DNA has identity or substantial identity with the sequence in the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide editing sites. In some embodiments, the editing template comprises at least four consecutive nucleotides complementary to the editing strand, the at least four consecutive nucleotides located upstream of the 5' edit in the editing template. In some embodiments, the editing template includes 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive nucleotides complementary to the editing chain, wherein at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive nucleotides are located upstream of the 5' edit in the editing template.

[0160] In some embodiments, newly synthesized single-stranded DNA equilibrates with the editing target on the edited strand of the target gene and pairs with the target strand of the target gene. In some embodiments, the editing target sequence of the target gene is excised by a flap endonuclease (FEN), e.g., FEN1. In some embodiments, the FEN is, for example, an endogenous FEN from within the cell containing the target gene. In some embodiments, the FEN is provided as part of a prime editor, either ligated to other components of the prime editor or provided trans. In some embodiments, newly synthesized single-stranded DNA containing the intended nucleotide edit replaces the endogenous single-stranded editing target sequence on the edited strand of the target gene. In some embodiments, the newly synthesized single-stranded DNA and the endogenous DNA on the target strand form a heteroduplex DNA structure in the region corresponding to the editing target sequence of the target gene. In some embodiments, the newly synthesized single-stranded DNA containing nucleotide edit is heteroduplex-paired with the target strand of the target DNA without nucleotide edit, thereby creating a mismatch between the two originally complementary strands. In some embodiments, the mismatch is recognized by a DNA repair mechanism, such as an endogenous DNA repair mechanism. In some embodiments, DNA repair incorporates the intended nucleotide edit into the target gene.

[0161] In some embodiments, prime editing may include programmable editing of target DNA using one or more prime editors, each complexed with a PEgRNA ("dual-prime editing"). Dual-prime editing refers to programmable editing of double-stranded target DNA using two or more PEgRNAs, each of which complexes with a prime editor for incorporating one or more intended nucleotide edits into the double-stranded target DNA. In some embodiments, dual-prime editing incorporates one or more intended nucleotide edits into the double-stranded target DNA via excision of an endogenous DNA segment and / or replacement of the endogenous DNA segment with newly synthesized DNA via targeted priming DNA synthesis. In some embodiments, dual-prime editing can be used to edit target DNA, which is a target gene or a part thereof.

[0162] In some embodiments, dual-prime editing involves using two different PEgRNAs, each complexed with a prime editor, and each of the two PEgRNAs contains a spacer that is complementary or substantially complementary to a separate search target sequence. In some embodiments, each of the two PEgRNAs anneals to the separate search target sequence via its spacer. Thus, references to “PAM strand,” “non-PAM strand,” “target strand,” “non-target strand,” “edited strand,” or “unedited strand” are relative in the context of a particular PEgRNA, for example, one of the two PEgRNAs in dual-prime editing.

[0163] In some embodiments, dual prime editing involves two distinct PEgRNAs, each forming a complex with a prime editor. In some embodiments, each of the two PEgRNAs contains a complementary region to distinct search target sequences on the target DNA, and the two distinct search target sequences are located on two complementary strands of the target DNA. In some embodiments, each of the two PEgRNAs can instruct the prime editor to initiate a prime editing process on the two complementary strands of the target DNA.

[0164] In some embodiments, dual prime editing involves two PEgRNAs, each forming a complex with a prime editor. In some embodiments, the first PEgRNA includes a first spacer complementary to a first search target sequence on the first strand of a double-stranded target DNA, e.g., a double-stranded target gene. In the context of the first PEgRNA, the first strand of the double-stranded target DNA may be called the first target strand, and the complementary strand may be called the first PAM strand.

[0165] In some embodiments, the second PEgRNA includes a second spacer complementary to a second search target sequence on the second strand of the double-stranded target DNA. In some embodiments, the first and second strands of the double-stranded target DNA, e.g., the double-stranded target gene, are complementary to each other. Therefore, in some embodiments, the second PEgRNA and the first PEgRNA bind to opposite strands of the double-stranded target DNA. In the context of the second PEgRNA, the second strand of the double-stranded target DNA may be called the second target strand, and the complementary strand may be called the second PAM strand. In some embodiments, the first target strand is the same strand as the second PAM strand of the double-stranded target DNA. In some embodiments, the second target strand is the same strand as the first PAM strand of the double-stranded target DNA.

[0166] In some embodiments, the first PEgRNA anneals to the first target strand of the double-stranded target DNA via the first spacer of the first PEgRNA. In some embodiments, the first PEgRNA forms a complex with the first prime editor and instructs the first prime editor to bind to the double-stranded target DNA at a position corresponding to the first search target sequence. In some embodiments, the second PEgRNA anneals to the second search target sequence on the second target strand of the double-stranded target DNA via the second spacer of the second PEgRNA. In some embodiments, the second PEgRNA forms a complex with the second prime editor and instructs the second prime editor to bind to the double-stranded target DNA at a position corresponding to the second search target sequence. In some embodiments, the first and second prime editors are the same. In some embodiments, the first and second prime editors are different.

[0167] In some embodiments, the bound first prime editor generates a first nick on the first PAM strand of the double-stranded target DNA. In some embodiments, the first PEgRNA includes a first primer-binding site (PBS), also referred herein as the “primer-binding site sequence,” which is complementary to the sequence of the first PAM strand of the double-stranded target DNA immediately upstream of the first nick site and can anneal to the sequence of the first strand at the free 3' end formed at the first nick site. In some embodiments, the first PEgRNA includes a first primer-binding site (PBS) that anneals to the free 3' end formed at the first nick site, and the first prime editor initiates DNA synthesis from the nick site using the free 3' end as a primer. In some embodiments, the first prime editor generates a first newly synthesized single-stranded DNA encoded by the first editing template of the first PEgRNA.

[0168] In some embodiments, the fused second prime editor generates a second nick on the second PAM strand of the double-stranded target DNA. In some embodiments, the double-stranded target DNA, e.g., the target gene, contains a double-stranded DNA sequence between a first nick generated by the first prime editor on the second target strand (also called the first PAM strand) and a second nick generated by the second prime editor on the first target strand (also called the second PAM strand), which may be called an inter-nick double strand (IND). In some embodiments, the two strands of the IND are completely complementary to each other. In some embodiments, the two strands of the IND are partially complementary to each other. In some embodiments, the IND is then excised from the double-stranded target DNA, e.g., the target gene.

[0169] In some embodiments, the first PEgRNA includes a first primer-binding site (PBS) complementary to the free 3' end of the second strand of the double-stranded target DNA formed at the first nick site. In some embodiments, the first PBS anneals with the free 3' end formed at the first nick site, and the first prime editor uses the free 3' end of the first nick site as a primer to initiate DNA synthesis from the first nick site. In some embodiments, the first prime editor synthesizes a first novel single-stranded DNA encoded by the first editing template of the first PEgRNA. In some embodiments, the second PEgRNA includes a second PBS complementary to the free 3' end of the first strand of the double-stranded target DNA formed at the second nick site. In some embodiments, the second PBS anneals with the free 3' end formed at the second nick site, and the second prime editor uses the free 3' end of the second nick site as a primer to initiate DNA synthesis from the nick site. In some embodiments, the second prime editor synthesizes a second newly synthesized single-stranded DNA encoded by a second editing template of the second PEgRNA.

[0170] In some embodiments, DNA repair integrates the sequence of a first newly synthesized single-stranded DNA encoded by a first editing template and / or the sequence of a second newly synthesized single-stranded DNA encoded by a second editing template into a double-stranded target DNA, e.g., a target gene, thereby integrating one or more intended nucleotide edits into the double-stranded target DNA, e.g., a target gene.

[0171] In some embodiments, the first and second editing templates include complementary or substantially complementary regions to each other. Thus, in some embodiments, the first newly synthesized single-stranded DNA and the second newly synthesized single-stranded DNA have complementary regions to each other. The complementary region between the first and second newly synthesized single-stranded DNA may be called an overlapping double-stranded DNA (OD). As illustrated in Figure 4A, in some embodiments, DNA repair integrates the OD into a double-stranded target DNA, e.g., a target gene, thereby integrating one or more intended nucleotide edits encoded by the first and second editing templates into the double-stranded target DNA, e.g., a target gene. In some embodiments, the OD replaces all or part of the IND, thereby integrating one or more intended nucleotide edits into the double-stranded target DNA, e.g., a target gene. In some embodiments, the IND is excised or degraded, the OD is incorporated into the site of the IND excision, and subsequently, the nicks on both strands of the double-stranded target DNA, e.g., the target gene, are ligated so that one or more intended nucleotide edits are incorporated into the double-stranded target DNA.

[0172] In some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template further includes regions that are not complementary to the second newly synthesized single-stranded DNA encoded by the second editing template (see illustrative schematic diagram in Figure 4B). In some embodiments, the second newly synthesized single-stranded DNA encoded by the second editing template further includes regions that are not complementary to the first newly synthesized single-stranded DNA encoded by the first editing template. Thus, in some embodiments, the first newly synthesized single-stranded DNA encoded by the first editing template and the second newly synthesized single-stranded DNA encoded by the second editing template can anneal to each other via partially complementary sequences to form an OD linked to the 5' overhang and / or 3' overhang. In some embodiments, IND is removed, and the OD, along with the 5' overhang and / or 3' overhang, is incorporated into the double-stranded target DNA, e.g., the site of IND excision in the target gene. DNA repair fills and ligates gaps corresponding to the 5' overhang and / or 3' overhang, thereby integrating one or more intended nucleotide edits into the double-stranded target DNA, e.g., the target gene.

[0173] Therefore, in some embodiments, IND is replaced by the sequence (A+C), (B+C), or (A+B+C), where A is a region of the first newly synthesized single-stranded DNA and its complementary strand that is not complementary to the second newly synthesized single-stranded DNA, where B is a region of the second newly synthesized single-stranded DNA and its complementary strand that is not complementary to the first newly synthesized single-stranded DNA, and where C is OD. The (A+C), (B+C), or (A+B+C) double-stranded sequence that replaces IND may be called a “substitution double helix (RD)”. In some embodiments, DNA repair replaces IND in the target DNA with the RD.

[0174] In some embodiments, the RD or OD may include an RRS recognized by a recombinase recognition site corresponding to, for example, Bxbl recombinase, Cre recombinase, Pa01 recombinase, Si74 recombinase, No67 recombinase, Kp03 recombinase, Nm60 recombinase, BceINTa recombinase, NcytINTd recombinase, SscINTd recombinase, SacINTd recombinase, or any recombinase disclosed herein. In some embodiments, the RD or OD may include one, two, or more recombinase recognition sites corresponding to a recombinase.

[0175] Replacing an IND with an RD or OD containing one or more RRSs using dual-prime editing can result in the insertion of one or more RRSs into the target gene. Depending on the number and orientation of the RRSs, they can be used as landing sites for recombinase-mediated reactions between RRSs. For example, a single RSS inserted into a target gene such as the TRAC gene can be used for the integration of an exogenous DNA donor sequence via recombination between the inserted RRS and a second RRS in the supplied exogenous DNA donor. If two RRS sites are inserted into adjacent regions of DNA, depending on the orientation of the RRS sites, they can be used for recombinase-mediated excision or inversion of the intervening sequence, or for recombinase-mediated cassette exchange with exogenous DNA for cargo integration.

[0176] Prime Editor The term “prime editor (PE)” refers to a polypeptide or polypeptide component involved in prime editing, or any polynucleotide(s) encoding a polypeptide or polypeptide component. In various embodiments, the prime editor comprises a polypeptide domain having DNA-binding activity and a polypeptide domain having DNA polymerase activity. In some embodiments, the prime editor further comprises a polypeptide domain having nuclease activity. In some embodiments, the polypeptide domain having DNA-binding activity comprises a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain having nuclease activity comprises a nickase or a fully active nuclease. As used herein, the term “nickase” refers to a nuclease capable of cleaving only one strand of a double-stranded DNA target. In some embodiments, the prime editor comprises a polypeptide domain that is an inactive nuclease. In some embodiments, the polypeptide domain having programmable DNA-binding activity comprises a nucleic acid-induced DNA-binding domain, e.g., a CRISPR-Cas protein, e.g., Cas9 nickase, Cpf1 nickase, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain having DNA polymerase activity includes a template-dependent DNA polymerase, e.g., a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase. In some embodiments, the prime editor includes an additional polypeptide involved in prime editing, e.g., a polypeptide domain having 5' endonuclease activity, e.g., a 5' endogenous DNA flap endonuclease (e.g., FEN1), which helps to advance the prime editing process toward the formation of the edited product. In some embodiments, the prime editor further includes an RNA-protein mobilization polypeptide, e.g., an MS2 coat protein.

[0177] Prime editors can be designed. In some embodiments, the polypeptide components of the prime editor do not naturally exist within the same organism or cellular environment. In some embodiments, the polypeptide components of the prime editor may be of different origins or derived from different organisms. In some embodiments, the prime editor includes a DNA-binding domain and a DNA polymerase domain derived from different species. In some embodiments, the prime editor includes a Cas polypeptide (DNA-binding domain) and a reverse transcriptase polypeptide (DNA polymerase) derived from different species. For example, the prime editor may include S. pyogenes Cas9 polypeptide and Moloney's mouse leukemia virus (M-MLV) reverse transcriptase polypeptide.

[0178] In some embodiments, the prime editor polypeptide domains can be fused or linked by a peptide linker to form a fusion protein. In other embodiments, the prime editor comprises one or more polypeptide domains provided trans as separate proteins that can bind to each other via non-peptide bonds or aptamers or recruitment sequences. For example, the prime editor may include a DNA-binding domain and a reverse transcriptase domain linked to each other by an RNA-protein recruitment aptamer (e.g., an MS2 aptamer that can be linked to PEgRNA). The polypeptide components of the prime editor can be encoded whole or partially by one or more polynucleotides. In some embodiments, a single polynucleotide, construct, or vector encodes the prime editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or portion of a domain of the prime editor, or a portion of the prime editor fusion protein. For example, the prime editor fusion protein comprises an N-terminal portion fused to intein N and a C-terminal portion fused to intein C, each individually encoded by an AAV vector.

[0179] Prime Editor nucleotide polymerase domain In some embodiments, the prime editor includes a nucleotide polymerase domain, such as a DNA polymerase domain. The DNA polymerase domain may be a wild-type DNA polymerase domain, a full-length DNA polymerase protein domain, or a functional variant, functional variant, or functional fragment thereof. In some embodiments, the polymerase domain is a template-dependent polymerase domain. For example, the DNA polymerase may depend on a template polynucleotide chain (such as an editing template sequence) for the synthesis of a new strand of DNA. In some embodiments, the prime editor includes a DNA-dependent DNA polymerase. For example, a prime editor having a DNA-dependent DNA polymerase can synthesize a new single-stranded DNA using a PEgRNA editing template that includes a DNA sequence as a template. In such cases, the PEgRNA is a chimeric or hybrid PEgRNA and includes an extension arm containing a DNA strand. The chimeric or hybrid PEgRNA may include an RNA portion (including a spacer and a gRNA core) and a DNA portion (an extension arm containing an editing template that includes a DNA strand).

[0180] In some embodiments, the DNA polymerase may be a wild-type polymerase derived from a eukaryote, prokaryote, archaea, or viral organism, and / or the polymerase may be modified by processes based on genetic engineering, mutagenesis, or directed evolution. The polymerase may be T7 DNA polymerase, T5 DNA polymerase, T4 DNA polymerase, Klenow fragment DNA polymerase, DNA polymerase III, etc. The polymerase may be thermally stable and may include Taq, Tne, Tma, Pfu, Tfl, Tth, Stoffel fragment, VENT® and DEEPVENT® DNA polymerase, KOD, Tgo, JDF3, and their variants, variants, and derivatives.

[0181] In some embodiments, the DNA polymerase is a bacteriophage polymerase, e.g., T4, T7, or phi29 DNA polymerase. In some embodiments, the DNA polymerase is an archaeal polymerase, e.g., Pol I archaeal polymerase or Pol II archaeal polymerase. In some embodiments, the DNA polymerase includes a thermostable archaeal DNA polymerase. In some embodiments, the DNA polymerase includes a bacterial DNA polymerase, e.g., Pol I, Pol II, or Pol III polymerase. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is E. coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is E. coli Pol IV DNA polymerase.

[0182] In some embodiments, the DNA polymerase includes eukaryotic DNA polymerase. In some embodiments, the DNA polymerase is Pol-beta DNA polymerase, Pol-lambda DNA polymerase, Pol-sigma DNA polymerase, or Pol-mu DNA polymerase. In some embodiments, the DNA polymerase is Pol-alpha DNA polymerase. In some embodiments, the DNA polymerase is POLA1 DNA polymerase. In some embodiments, the DNA polymerase is POLA2 DNA polymerase. In some embodiments, the DNA polymerase is Pol-delta DNA polymerase. In some embodiments, the DNA polymerase is POLD1 DNA polymerase. In some embodiments, the DNA polymerase is POLD2 DNA polymerase. In some embodiments, the DNA polymerase is human POLD1 DNA polymerase. In some embodiments, the DNA polymerase is human POLD2 DNA polymerase. In some embodiments, the DNA polymerase is POLD3 DNA polymerase. In some embodiments, the DNA polymerase is POLD4 DNA polymerase. In some embodiments, the DNA polymerase is Pol-epsilon DNA polymerase. In some embodiments, the DNA polymerase is POLE1 DNA polymerase. In some embodiments, the DNA polymerase is POLE2 DNA polymerase. In some embodiments, the DNA polymerase is POLE3 DNA polymerase. In some embodiments, the DNA polymerase is Pol-eta (POLH) DNA polymerase. In some embodiments, the DNA polymerase is Pol-iota (POLI) DNA polymerase. In some embodiments, the DNA polymerase is Pol-kappa (POLK) DNA polymerase. In some embodiments, the DNA polymerase is Rev1 DNA polymerase. In some embodiments, the DNA polymerase is human Rev1 DNA polymerase.In some embodiments, the DNA polymerase is a viral DNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a B-family DNA polymerase. In some embodiments, the DNA polymerase is herpes simplex virus (HSV) UL30 DNA polymerase. In some embodiments, the DNA polymerase is cytomegalovirus (CMV) UL54 DNA polymerase.

[0183] In some embodiments, the DNA polymerase is an archaeal polymerase. In some embodiments, the DNA polymerase is a family B / pol I type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of Pfu derived from Pyrococcus furiosus. In some embodiments, the DNA polymerase is a Pol II type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of P. furiosus DP1 / DP2 2 subunit polymerase. In some embodiments, the DNA polymerase lacks 5' to 3' nuclease activity. A suitable DNA polymerase (pol I or pol II) can be obtained from archaea with an optimal growth temperature close to the desired assay temperature.

[0184] In some embodiments, the DNA polymerase includes a thermostable archaeal DNA polymerase. In some embodiments, the thermostable DNA polymerase is isolated or derived from Pyrococcus species (furiosus, species GB-D, woesii, abysii, horikoshii), Thermococcus species (kodakaraensis KOD1, litoralis, species 9 degrees North-7, species JDF-3, gorgonarius), Pyrodictium occultum, and Archaeoglobus fulgidus.

[0185] The polymerase may also be derived from a bacterial species. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is E. coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol III family DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is E. coli Pol IV DNA polymerase. In some embodiments, the Pol I DNA polymerase is a functional variant of the DNA polymerase in which the 5' to 3' exonuclease activity is lacking or reduced.

[0186] Suitable thermostable pol I DNA polymerases can be isolated from a variety of thermophilic bacteria, including Thermus species and Thermotoga maritima, such as Thermus aquaticus (Taq), Thermus thermophilus (Tth), and Thermotoga maritima (Tma UlTma).

[0187] In some embodiments, the prime editor includes an RNA-dependent DNA polymerase domain, such as reverse transcriptase (RT). The RT or RT domain may be a wild-type RT domain, a full-length RT domain, or a functional mutant, functional variant, or a functional fragment thereof. The RT or RT domain of the prime editor may contain wild-type RT, or may be designed or evolved to contain specific amino acid substitutions, cleavages, or variants. The designed RT may contain sequence or amino acid changes different from naturally occurring RT. In some embodiments, the designed RT may have improved reverse transcription activity compared to naturally occurring RT or RT domains. In some embodiments, the designed RT may have improved properties such as thermal stability, reverse transcription efficiency, and target fidelity compared to naturally occurring RT. In some embodiments, a prime editor containing a designed RT may have improved prime editing efficiency compared to a prime editor with a naturally occurring reference RT.

[0188] In some embodiments, the Prime Editor includes viral RTs, such as retrovirus RTs. Non-limiting examples of viral RTs include Moloney's mouse leukemia virus (M-MLV, MMLVRT, or M-MLV RT), human T-cell leukemia virus type 1 (HTLV-1) RT, bovine leukemia virus (BLV) RT, Roussarcoma virus (RSV) RT, human immunodeficiency virus (HIV) RT, M-MFV RT, avian sarcoma leukemia virus (ASLV) RT, Roussarcoma virus (RSV) RT, avian myeloblastosis virus (AMV) RT, avian erythroblastosis virus (AEV) helper virus MCAV RT, avian myelocytosis virus MC29 helper virus MCAV RT, avian reticuloendotheliosis virus (REV-T) helper virus REV-A RT, avian sarcoma virus UR2 helper virus (UR2AV) RT, and avian sarcoma virus Y73 helper virus YAV This includes RT, Rous-associated virus (RAV) RT, and myeloblastosis-associated virus (MAV) RT, all of which can be appropriately used in the methods and compositions described herein.

[0189] In some embodiments, the prime editor includes wild-type M-MLV RT, its functional variant, functional variant, or functional fragment.

[0190] In some embodiments, the prime editor includes a reference M-MLV RT, its functional variant, functional variant, or functional fragment. In some embodiments, the RT domain or RT is an M-MLV RT (e.g., wild-type M-MLV RT, its functional variant, functional variant, or functional fragment). In some embodiments, the RT domain or RT is an M-MLV RT (e.g., a reference M-MLV RT, its functional variant, functional variant, or functional fragment). In some embodiments, the M-MLV RT, e.g., a reference M-MLV RT, includes any of the amino acid sequences described in SEQ ID NO: 672.

[0191] In some embodiments, the reference M-MLV RT is wild-type M-MLV RT. An exemplary amino acid sequence of the reference M-MLV RT is provided in SEQ ID NO: 671.

[0192] In some embodiments, the prime editor includes wild-type M-MLV RT. An exemplary amino acid sequence of wild-type M-MLV RT is provided in SEQ ID NO: 671.

[0193] (Sequence ID 671).

[0194] In some embodiments, the prime editor includes a reference M-MLV RT. An exemplary amino acid sequence of the reference M-MLV RT is provided in SEQ ID NO: 672.

[0195] (Sequence ID 672).

[0196] In some embodiments, the prime editor includes an M-MLV RT containing one or more amino acid substitutions P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, or D653X, compared to the reference M-MLV RT described in SEQ ID NO: 672, where X is any amino acid other than the original amino acid in the reference M-MLV RT. In some embodiments, the Prime Editor includes an M-MLV RT containing one or more amino acid substitutions P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, and D653N, compared to the reference M-MLV RT described in SEQ ID NO: 672. In some embodiments, the prime editor includes an M-MLV RT containing the amino acid substitutions D200N, T330P, L603W, T306K, and W313F compared to the wild-type M-MMLV RT described in SEQ ID NO: 672. In some embodiments, the prime editor includes the prime editor containing D200N, T330P, L603W, T306K, and W313F compared to the reference M-MLV RT described in SEQ ID NO: 672. In some embodiments, the prime editor includes an M-MLV RT containing one or more of the amino acid substitutions D200N, T330P, L603W, T306K, and W313F compared to the wild-type M-MMLV RT described in SEQ ID NO: 671. In some embodiments, the prime editor may contain the amino acid substitutions D200N, T330P, L603W, T306K, and W313F compared to the reference M-MLV RT described in SEQ ID NO: 672.In some embodiments, the prime editor includes an M-MLV RT containing an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to the amino acid sequence described in any one of SEQ ID NOs. 671, 672, or 673. In some embodiments, the prime editor includes an M-MLV RT containing an amino acid sequence selected from the group consisting of SEQ ID NOs. 671, 672, or 673 or its variants or fragments. In some embodiments, the prime editor includes an M-MLV RT containing the amino acid sequence described in SEQ ID NOs. 673. (Sequence ID 673).

[0197] In some embodiments, the RT variant may be a functional fragment of a reference RT having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500, or more amino acid changes compared to the wild-type RT, for example, SEQ ID NO: 671. In some embodiments, the RT variant comprises a fragment of wild-type RT, e.g., SEQ ID NO: 671, which is approximately 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% identical to the corresponding fragment of wild-type RT, e.g., SEQ ID NO: 671. In some embodiments, the fragment is 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to the amino acid length of the corresponding wild-type RT (M-MLV reverse transcriptase) (e.g., SEQ ID NO: 671).

[0198] In some embodiments, the RT variant may be a functional fragment of the reference RT having an amino acid change of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500, or more, compared to the reference RT, for example, SEQ ID NO: 672. In some embodiments, the RT variant comprises a fragment of a reference RT, the fragment being approximately 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% identical to the wild-type RT, e.g., the corresponding fragment of SEQ ID NO: 672. In some embodiments, the fragment is 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to the amino acid length of the reference RT, e.g., M-MLV RT, e.g., SEQ ID NO: 672.

[0199] In some embodiments, the functional fragment of the RT is at least 100 amino acids long. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, or up to 600 or more amino acids long.

[0200] In yet another embodiment, the functional RT variant is cleaved by a specific number of amino acids at the N-terminus, C-terminus, or both, resulting in a cleaved variant that still retains sufficient DNA polymerase function. In some embodiments, the functional RT variant, for example, the functional MMLV RT variant, is cleaved at the C-terminus, losing or reducing RNAase H activity while still retaining DNA polymerase activity.

[0201] In some embodiments, the prime editing composition or prime editing system disclosed herein comprises a polynucleotide (e.g., DNA, RNA, e.g., mRNA) encoding an M-MLV RT. In some embodiments, the polynucleotide encodes an M-MLV RT comprising an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to the amino acid sequence described in any one of SEQ ID NOs: 671, 672, or 673. In some embodiments, the polynucleotide encodes an M-MLV RT comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 671, 672, and 673. In some embodiments, the polynucleotide encodes an M-MLV RT comprising the amino acid sequence described in SEQ ID NO: 673.

[0202] In some embodiments, the prime editor includes eukaryotic RTs, such as those from yeast, fruit flies, rodents, or primates. In some embodiments, the prime editor includes group II intron RTs, such as those from Geobacillus stearothermophilus group II introns (GsI-IIC) or Eubacterium rectale group II introns (Eu.re.I2). In some embodiments, the prime editor includes retron RTs. In some embodiments, the prime editor includes eukaryotic RTs, such as those from yeast, fruit flies, rodents, or primates. In some embodiments, the prime editor includes group II intron RTs, such as those from Geobacillus stearothermophilus group II introns (GsI-IIC) or Eubacterium rectale group II introns (Eu.re.I2). In some embodiments, the prime editor includes retron RTs.

[0203] Programmable DNA-binding domain In some embodiments, the DNA-binding domain of the prime editor is a programmable DNA-binding domain. In some embodiments, the prime editor includes a DNA-binding domain comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences described in SEQ ID NOs. In some embodiments, the DNA-binding domain includes an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or fewer differences compared to any one of the amino acid sequences described in SEQ ID NOs. 674-701, e.g., mutations, deletions, substitutions, and / or insertions. In some embodiments, the DNA-binding domain of the prime editor is a programmable DNA-binding domain. A programmable DNA-binding domain refers to a protein domain designed to bind to a specific nucleic acid sequence, e.g., target DNA or target RNA. In some embodiments, the DNA-binding domain is a polynucleotide programmable DNA-binding domain that can bind to a guide polynucleotide (e.g., PEgRNA) that directs the DNA-binding domain to a specific DNA sequence, e.g., a search target sequence within a target gene. In some embodiments, the DNA-binding domain includes a short repeat palindrome (CRISPR)-associated (Cas) protein clustered at regular intervals. The Cas protein may include any Cas protein described herein or its functional fragments or functional variants. In some embodiments, the DNA-binding domain may also include a zinc finger protein domain. In other embodiments, the DNA-binding domain includes a transcription activator-like effector domain (TALE). In some embodiments, the DNA-binding domain includes a DNA nuclease.For example, the DNA-binding domain of the prime editor may include an RNA-guided DNA endonuclease, such as a Cas protein. In some embodiments, the DNA-binding domain includes a zinc finger nuclease (ZFN) or a transcription activator-like effector domain nuclease (TALEN), where one or more zinc finger motifs or TALE motifs are associated with one or more nucleases, such as a Fok I nuclease domain.

[0204] In some embodiments, the DNA-binding domain contains nuclease activity. In some embodiments, the DNA-binding domain of the prime editor contains an endonuclease domain having single-strand DNA cleavage activity. For example, the endonuclease domain may contain a FokI nuclease domain. In some embodiments, the DNA-binding domain of the prime editor contains a nuclease having full nuclease activity. In some embodiments, the DNA-binding domain of the prime editor contains a nuclease with modified or reduced nuclease activity compared to a wild-type endonuclease domain. For example, the endonuclease domain may contain one or more amino acid substitutions compared to a wild-type endonuclease domain. In some embodiments, the DNA-binding domain of the prime editor has nickase activity. In some embodiments, the DNA-binding domain of the prime editor contains a Cas protein domain that is a nickase. In some embodiments, compared to a wild-type Cas protein, the Cas nickase contains one or more amino acid substitutions in the nuclease domain, thereby reducing or eliminating double-strand nuclease activity while retaining DNA-binding activity. In some embodiments, Cas nickase includes amino acid substitutions within the HNH domain. In some embodiments, Cas nickase includes amino acid substitutions within the RuvC domain.

[0205] In some embodiments, the DNA-binding domain includes a CRISPR-related protein (Cas protein) domain. The Cas protein may be a class 1 or class 2 Cas protein. The Cas protein may be a type I, type II, type III, type IV, type V Cas protein, or type VI Cas protein. Non-limiting examples of Cas proteins include Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csnl or Csx12), Cas10, CaslOd, Cas12a / Cpfl, Cas12b / C2c1, Cas12c / C2c3, Ca s12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csyl, Csy2, Csy3, Csy4, Csel, Cse2, Cse3, Cse4, Cse5e, Csc l, Csc2, Csa5, Csnl, Csn2, Csml, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx1S, Csx11, Csfl, Csf2, CsO, Csf4, Csdl, Csd2, Cstl, Cst2, Cshl, Csh2, Csal, Csa2, Csa3, Csa4, Csa5, Type II Cas effector proteins, Type V Cas effector proteins, Type VI Cas effector proteins Examples include Cas proteins, CARF, DinG, Cpfl, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, high-precision Cas9 variant (HypaCas9), CasΦ, and their homologs, modified or manipulated variants, mutants, and / or functional fragments. Cas proteins can be chimeric Cas proteins fused to other proteins or polypeptides. Cas proteins can be chimeric of various Cas proteins, for example, containing domains of Cas proteins from different organisms.

[0206] Cas proteins such as Cas9 can be derived from any suitable organism. In some embodiments, this microorganism is Streptococcus pyogenes (S. pyogenes). In some embodiments, this microorganism is Staphylococcus aureus (S. aureus). In some embodiments, this microorganism is Streptococcus thermophilus (S. thermophilus). In some embodiments, this organism is Staphylococcus lugdunensis.

[0207] Non-limiting examples of suitable organisms include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas species, Crocosphaera watsonii, Cyanothece species, Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus species, Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter species, Nitrosococcus halophilus, Nitrosococcus watsoni, PseudoalteromonasExamples include haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc species, Arthrospira maxima, Arthrospira platensis, Arthrospira species, Lyngbya species, Microcoleus chthonoplastes, Oscillatoria species, Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella novicida. In some embodiments, this organism is Streptococcus pyogenes (S. pyogenes). In some embodiments, this organism is Staphylococcus aureus (S. aureus). In some embodiments, this organism is Streptococcus thermophilus (S. thermophilus). In some embodiments, this organism is Staphylococcus lugdunensis (S. lugdunensis).

[0208] In some embodiments, the Cas protein can be derived from various bacterial species including, but not limited to: Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractor salsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonaspalustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinella succinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, and Francisella novicida.

[0209] In some embodiments, the Cas protein, for example Cas9, may be a wild-type or modified form of the Cas protein. In some embodiments, the Cas protein, for example Cas9, may be a nuclease-active variant, a nuclease-inactive variant, a nickase, or a functional variant or functional fragment of the wild-type Cas protein. In some embodiments, the Cas protein, for example Cas9, may be a wild-type or modified form of the Cas protein. The Cas protein, for example Cas9, may be a nuclease-active variant, a nuclease-inactive variant, a nickase, or a functional variant or functional fragment of the wild-type Cas protein. In some embodiments, the Cas protein, for example Cas9, may include amino acid changes such as deletions, insertions, substitutions, fusions, chimeras, or any combination thereof, compared to the wild-type version of the corresponding Cas protein. In some embodiments, the Cas protein may be a polypeptide having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity with a wild-type exemplary Cas protein.

[0210] A Cas protein, such as Cas9, may contain one or more domains. Non-limiting examples of Cas domains include guide nucleic acid recognition and / or binding domains, nuclease domains (e.g., DNase or RNase domains, RuvC, HNH), DNA binding domains, RNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. In various embodiments, a Cas protein may include a guide nucleic acid recognition and / or binding domain capable of interacting with a guide nucleic acid, and one or more nuclease domains having catalytic activity for nucleic acid cleavage.

[0211] In some embodiments, a Cas protein, such as Cas9, contains one or more nuclease domains. A Cas protein may contain an amino acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the nuclease domain of a wild-type Cas protein (e.g., a RuvC domain, an HNH domain). In some embodiments, a Cas protein contains a single nuclease domain. For example, Cpf1 may contain a RuvC domain but lack an HNH domain. In some embodiments, a Cas protein contains two nuclease domains; for example, the Cas9 protein may contain an HNH nuclease domain and a RuvC nuclease domain.

[0212] In some embodiments, the prime editor comprises a Cas protein, e.g., Cas9, where all nuclease domains of the Cas protein are active. In some embodiments, the prime editor comprises a Cas protein having one or more inactive nuclease domains. One or more nuclease domains of a Cas protein (e.g., RuvC, HNH) can be deleted or mutated to render them non-functional or reduce their nuclease activity. In some embodiments, a Cas protein containing a mutation in its nuclease domain (e.g., Cas9), when compounded with a guide nucleic acid (e.g., PEgRNA), reduces (e.g., nickase) or eliminates nuclease activity while maintaining the ability to target nucleic acid loci at the search target sequence.

[0213] In some embodiments, the prime editor includes a Cas nickase that can sequence-specifically bind to a target gene and generate a single-strand break at a protospacer within the double-stranded DNA of the target gene, but not a double-strand break. For example, a Cas nickase can cleave either the edited or unedited strand of a target gene, but not both. In some embodiments, the prime editor includes a Cas nickase (e.g., Cas9) containing two nuclease domains, one of which is modified to lack catalytic activity or is deleted. In some embodiments, the Cas nickase of the prime editor includes a nuclease-inactive RuvC domain and a nuclease-active HNH domain. In some embodiments, the Cas nickase of the prime editor includes a nuclease-inactive HNH domain and a nuclease-active RuvC domain. In some embodiments, the prime editor includes Cas9 nickase having an amino acid substitution in the RuvC domain, for example, an amino acid substitution that reduces or eliminates the nuclease activity of the RuvC domain. In some embodiments, the Cas9 nickase includes the amino acid substitution D10X compared to Cas9 of wild-type S. pyogenes, where X is any amino acid other than D. In some embodiments, the prime editor includes Cas9 nickase having an amino acid substitution in the HNH domain, for example, an amino acid substitution that reduces or eliminates the nuclease activity of the HNH domain. In some embodiments, the Cas9 nickase includes the amino acid substitution H840X compared to Cas9 of wild-type S. pyogenes, where X is any amino acid other than H.

[0214] In some embodiments, the prime editor includes a Cas protein that can bind to a target gene in a sequence-specific manner but lacks or has lost nuclease activity and cannot cleave either strand of the double-stranded DNA within the target gene. Loss of activity or absence of activity may refer to enzymatic activity of less than 1%, less than 2%, less than 3%, less than 4%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, or less than 10% of the exemplary wild-type activity (e.g., wild-type Cas9 nuclease activity). In some embodiments, the Cas protein of the prime editor has no nuclease activity at all. Nucleases lacking nuclease activity, such as Cas9, may be called nuclease-inactive or "nuclease-dead" (abbreviated as "d"). Nuclease-dead Cas proteins (e.g., dCas, dCas9) can bind to a target polynucleotide but cannot cleave it. In some embodiments, the dead Cas protein is the dead Cas9 protein. In some embodiments, the prime editor comprises a nuclease-dead Cas protein, where all of the nuclease domains (e.g., both the RuvC and HNH nuclease domains in the Cas9 protein; the RuvC nuclease domain in the Cpf1 protein) are mutated or deleted to lack catalytic activity.

[0215] Cas proteins can be modified. Cas proteins, such as Cas9, can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to alter other protein activities or properties, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or a Cas protein can be cleaved to remove domains that are not essential for the protein's function, or the activity of a Cas protein can be optimized (e.g., enhanced or reduced).

[0216] Cas proteins can form fusion proteins. For example, a Cas protein can fuse to a cleavage domain, an epigenetic modification domain, a transcriptional regulatory domain, or a polymerase domain. Cas proteins can also fuse to heterologous polypeptides, which can increase or decrease their stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internally within the Cas protein.

[0217] In some embodiments, the Cas protein of Prime Editor is a class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of the Cas9 protein, a homolog of the Cas9 protein, a mutant, a variant, or a functional fragment thereof. As used herein, Cas9, Cas9 protein, Cas9 polypeptide, or Cas9 nuclease refers to an RNA guide nuclease comprising one or more Cas9 nuclease domains and a gRNA-binding domain of Cas9 having the ability to bind to a guide polynucleotide, such as PEgRNA. The Cas9 protein may refer to a wild-type Cas9 protein from any organism, or a homolog, orthologue, paralog, functional mutant or functional variant thereof, or a functional fragment or domain thereof from any organism. In some embodiments, Prime Editor contains a full-length Cas9 protein. In some embodiments, the Cas9 protein may generally have at least about 50%, 60%, 70%, 80%, 90%, or 100% sequence identity with the wild-type reference Cas9 protein (e.g., Cas9 from S. pyogenes). In some embodiments, Cas9 may have amino acid changes such as deletions, insertions, substitutions, fusions, chimeras, or any combination thereof, compared to the wild-type reference Cas9 protein.

[0218] In some embodiments, the Cas9 protein may include Cas9 proteins derived from Streptococcus pyogenes (Sp), Staphylococcus aureus (Sa), Streptococcus canis (Sc), Streptococcus thermophilus (St), Staphylococcus lugdunensis (Slu), Neisseria meningitidis (Nm), Campylobacter jejuni (Cj), Francisella novicida (Fn), or Treponema denticola (Td), or any Cas9 homolog or orthologue from organisms known in the art. In some embodiments, the Cas9 polypeptide is, for example, the SpCas9 polypeptide, or a fragment or variant thereof, comprising the amino acid sequence shown in NCBI accession number WP_038431314. In some embodiments, the Cas9 polypeptide is, for example, the SaCas9 polypeptide, or a fragment or variant thereof, comprising the amino acid sequence described in Uniprot accession number J7RUA5. In some embodiments, the Cas9 polypeptide is an ScCas9 polypeptide, or a fragment or variant thereof, comprising an amino acid sequence described, for example, Uniprot accession number A0A3P5YA78. In some embodiments, the Cas9 polypeptide is an StCas9 polypeptide, or a fragment or variant thereof, comprising an amino acid sequence described, for example, NCBI accession number WP_007896501.1. In some embodiments, the Cas9 polypeptide is an SluCas9 polypeptide, or a fragment or variant thereof, comprising an amino acid sequence described, for example, NCBI accession number WP_230580236.1 or WP_250638315.1 or WP_242234150.1, WP_241435384.1, WP_002460848.1, or KAK58371.1.In some embodiments, the Cas9 polypeptide is an NmCas9 polypeptide, or a fragment or variant thereof, comprising an amino acid sequence described, for example, NCBI accession number WP_002238326.1 or WP_061704949.1. In some embodiments, the Cas9 polypeptide is a CjCas9 polypeptide, or a fragment or variant thereof, comprising an amino acid sequence described, for example, NCBI accession number WP_100612036.1, WP_116882154.1, WP_116560509.1, WP_116484194.1, WP_116479303.1, WP_115794652.1, or WP_100624872.1. In some embodiments, the Cas9 polypeptide is an FnCas9 polypeptide, or a fragment or variant thereof, comprising an amino acid sequence described, for example, Uniprot accession number A0Q5Y3. In some embodiments, the Cas9 polypeptide is a TdCas9 polypeptide, or a fragment or variant thereof, comprising an amino acid sequence, for example, as described in NCBI accession number WP_147625065.1. In some embodiments, the Cas9 polypeptide is a chimera comprising domains derived from two or more organisms described herein or known in the art. In some embodiments, the Cas9 polypeptide is a Cas9 polypeptide derived from Streptococcus macacae, or a fragment or variant thereof, comprising an amino acid sequence, for example, as described in NCBI accession number WP_003079701.1. In some embodiments, the Cas9 polypeptide is a Cas9 polypeptide produced by replacing the PAM interaction domain of SpCas9 with that of Cas9 from Streptococcus macacae (Spy-mac Cas9). Exemplary Cas sequences are shown in Table 22 below.

[0219] In some embodiments, the Cas9 protein contains an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences described in SEQ ID NOs. In some embodiments, the Cas9 protein is a Cas9 nikkase comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences described in SEQ ID NOs. 675, 676, 677, 679, 679, 689, 691 In some embodiments, the prime editor includes a Cas9 protein having an amino acid sequence lacking an N-terminal methionine compared to the amino acid sequence described in any one of SEQ ID NOs: 674, 675, 678, 679, 681, 682, 684, 685, 687, 688, 690, 691, 693, 694, 696, 697, 699, or 700.In some embodiments, the prime editing compositions or prime editing systems disclosed herein include a polynucleotide (e.g., DNA, or RNA, e.g., mRNA) encoding a Cas9 protein, comprising an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences described in SEQ ID NOs. 674 to 701.

[0220] In some embodiments, the Cas9 protein includes, for example, the Cas9 protein derived from Streptococcus pyogenes (Sp) of NC_002737.2:854751-858857, or the protein encoded by UniProt Q99ZW2 of, for example, SEQ ID NO: 674. In some embodiments, the prime editor includes one Cas9 protein (e.g., SpCas9) or a variant thereof from any of the sequences described in SEQ ID NOs: 674-677. In some embodiments, the Cas9 protein is SpCas9. In some embodiments, SpCas9 may be wild-type SpCas9, an SpCas9 variant, or nickase SpCas9. In some embodiments, SpCas9 lacks an N-terminal methionine compared to the corresponding SpCas9 (e.g., wild-type SpCas9, an SpCas9 variant, or nickase SpCas9). In some embodiments, the prime editor comprises a Cas9 protein having the amino acid sequence of SEQ ID NO: 674, which does not contain an N-terminal methionine. In some embodiments, wild-type SpCas9 comprises the amino acid sequence described in SEQ ID NO: 674. In some embodiments, the prime editor comprises a Cas9 protein having one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to the corresponding wild-type Cas9 protein (e.g., wild-type SpCas9). In some embodiments, the prime editor comprises a Cas9 protein having the amino acid sequence described in SEQ ID NO: 675, SEQ ID NO: 676, or SEQ ID NO: 677. The amino acid sequences of exemplary Streptococcus pyogenes Cas9 (SpCas9) useful in the prime editors disclosed herein are shown below in SEQ ID NOs: 674-677.

[0221] In some embodiments, Prime Editor comprises one of the Cas9 proteins (e.g., SluCas9) or a variant thereof from SEQ ID NOs. 678-680. In some embodiments, Prime Editor comprises, for example, one of the Cas9 proteins (SluCas9) or a variant thereof from Staphylococcus lugdunensis from SEQ ID NOs. 678-680. In some embodiments, the Cas9 protein is SluCas9. In some embodiments, SluCas9 may be wild-type SluCas9, a SluCas9 variant, or nickase SluCas9. In some embodiments, SluCas9 lacks an N-terminal methionine compared to the corresponding SluCas9 (e.g., wild-type SluCas9, a SluCas9 variant, or nickase SluCas9). In some embodiments, Prime Editor comprises a Cas9 protein having the amino acid sequence of SEQ ID NOs. 678, which does not contain an N-terminal methionine. In some embodiments, wild-type SluCas9 has the amino acid sequence described in SEQ ID NOs. 678. In some embodiments, the prime editor includes a Cas9 protein containing one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to the corresponding wild-type Cas9 protein (e.g., wild-type SluCas9). In some embodiments, the Cas9 protein containing one or more mutations compared to the wild-type Cas9 protein includes the amino acid sequence described in SEQ ID NO: 679. The amino acid sequences of exemplary Staphylococcus lugdunensis Cas9 (SluCas9) useful in the prime editor disclosed herein are shown below in SEQ ID NOs: 678-680.

[0222] In some embodiments, the prime editor includes, for example, a Cas9 protein (SaCas9) derived from Staphylococcus aureus, one of the sequence numbers 681-683, or a variant thereof. In some embodiments, the prime editor includes, for example, one of the sequence numbers 681-683 derived from Staphylococcus aureus, or a variant thereof. In some embodiments, the Cas9 protein is SaCas9. In some embodiments, SaCas9 may be wild-type SaCas9, a SaCas9 variant, or nickase SaCas9. In some embodiments, SaCas9 lacks an N-terminal methionine compared to the corresponding SaCas9 (e.g., wild-type SaCas9, a SaCas9 variant, or nickase SaCas9). In some embodiments, the prime editor includes a Cas9 protein having the amino acid sequence of sequence number 681, which does not contain an N-terminal methionine. In some embodiments, wild-type SaCas9 has the amino acid sequence described in sequence number 681. In some embodiments, the prime editor includes a Cas9 protein containing one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to the corresponding wild-type Cas9 protein (e.g., wild-type SaCas9). In some embodiments, the Cas9 protein containing one or more mutations compared to the wild-type Cas9 protein includes the amino acid sequence described in SEQ ID NO: 682. The amino acid sequences of exemplary Staphylococcus aureus Cas9 (SaCas9) useful in the prime editor disclosed herein are shown below in SEQ ID NOs: 681-683.

[0223] In some embodiments, the prime editor includes a Cas protein, e.g., a Cas9 variant, that includes modifications enabling altered PAM recognition. Exemplary Cas9 protein amino acid sequences useful in the prime editor of this disclosure (e.g., Cas9 variants with altered PAM recognition specificity) are shown below in SEQ ID NOs. 684-692, 699-701. In some embodiments, the prime editor includes a Cas9 protein or a variant thereof with any one of the sequences described in SEQ ID NOs. 684-692, 699-701. In some embodiments, the Cas9 protein is a Cas9 variant, e.g., an SpCas9 variant (e.g., SpCas9-NG, SpCas9-NGA, SpRY, or SpG). In some embodiments, the Cas9 protein lacks an N-terminal methionine compared to a corresponding Cas9 protein (e.g., a Cas9 variant described in any one of SEQ ID NOs. 684, 685, 687, 688, 690, 691, 699, or 700). In some embodiments, the prime editor includes a Cas9 protein (e.g., a Cas9 variant) having one of the amino acid sequences of SEQ ID NOs. 684, 687, 690, or 699 that does not contain an N-terminal methionine. In some embodiments, the prime editor includes a Cas9 protein containing one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to a corresponding Cas9 protein (e.g., a Cas9 protein described in any one of SEQ ID NOs. 684, 687, 690, or 699). In some embodiments, a Cas9 protein containing one or more mutations compared to a corresponding Cas9 protein includes the amino acid sequence described in any one of SEQ ID NOs. 685, 686, 688, 689, 691, 692, 700, or 701.

[0224] In some embodiments, the Cas9 protein is a chimeric Cas9, e.g., a modified Cas9, e.g., a synthetic RNA guide nuclease (sRGN) (e.g., modified by DNA family shuffling), e.g., sRGN3.1, sRGN3.3. In some embodiments, DNA family shuffling involves the fragmentation and reconstruction of one or more Cas9s derived from parental Cas9 genes, e.g., Staphylococcus hyicus (Shy), Staphylococcus lugdunensis (Slu), Staphylococcus microti (Smi), and Staphylococcus pasteuri (Spa). In some embodiments, modified SluCas9 exhibits increased editing efficiency and / or specificity compared to unmodified SluCas9. In some cases, a modified Cas9, such as sRGN, shows an increase in editing efficiency of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% compared to an unmodified Cas9. In some ways, Cas9, e.g., sRGN, exhibits an increase in singularity of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% compared to unmodified Cas9.In some embodiments, Cas9, e.g., sRGN, exhibits an increase in cleavage activity of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% compared to unmodified Cas9. In some embodiments, Cas9, e.g., sRGN, exhibits the ability to cleave 5'-NNGG-3'PAM-containing targets. In some embodiments, the prime editor comprises a Cas9 protein (e.g., chimeric Cas9) or a variant thereof with any one of the sequences described in SEQ ID NOs.693-698. The amino acid sequences of exemplary Cas9 proteins (e.g., sRGN) useful in the prime editor disclosed herein are shown below in SEQ ID NOs: 693-698. In some embodiments, the prime editor includes a Cas9 protein lacking an N-terminal methionine compared to SEQ ID NOs: 693 or 696. In some embodiments, the prime editor includes a Cas9 protein containing one or more mutations (e.g., amino acid substitutions, insertions, and / or deletions) compared to the corresponding Cas9 protein (e.g., the Cas9 protein described in SEQ ID NOs: 693 or 696). In some embodiments, the Cas9 protein containing one or more mutations compared to the corresponding Cas9 protein includes the amino acid sequence described in any one of SEQ ID NOs: 694, 695, 697, or 698. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] Table 1-5 Table 1-6 Table 1-7 Table 1-8 Table 1-9 Table 1-10 Table 1-11 Table 1-12 Table 1-13 Table 1-14 Table 1-15 Table 1-16 Table 1-17 Table 1-18 Table 1-19 Table 1-20 Table 1-21 Table 1-22 Table 1-23 Table 1-24 Table 1-25 Table 1-26 Table 1-27 Table 1-28

[0225] In some embodiments, the Cas9 protein comprises a variant Cas9 protein containing one or more amino acid substitutions. In some embodiments, the wild-type Cas9 protein comprises a RuvC domain and an HNH domain. In some embodiments, the prime editor comprises a nuclease-active Cas9 protein capable of cleaving both strands of a double-stranded target DNA sequence. In some embodiments, the nuclease-active Cas9 protein comprises a functional RuvC domain and a functional HNH domain. In some embodiments, the prime editor comprises a Cas9 nickasase capable of binding to a guide polynucleotide to recognize target DNA but capable of cleaving only one strand of the double-stranded target DNA. In some embodiments, the Cas9 nickasase comprises only one functional RuvC domain or one functional HNH domain. In some embodiments, the prime editor comprises a Cas9 having a non-functional HNH domain and a functional RuvC domain. In some embodiments, the prime editor can cleave the edited strand (i.e., the PAM strand) but cannot cleave the unedited strand of the double-stranded target DNA sequence. In some embodiments, the prime editor includes Cas9 having a non-functional RuvC domain that can cleave the target strand (i.e., the non-PAM strand) but cannot cleave the edited strand of the double-stranded target DNA sequence. In some embodiments, the prime editor includes Cas9 that has neither a functional RuvC domain nor a functional HNH domain and may not cleave any strand of the double-stranded target DNA sequence.

[0226] In some embodiments, the prime editor includes Cas9 having a mutation in the RuvC domain that reduces or disables the nuclease activity of the RuvC domain. In some embodiments, the Cas9 includes a mutation at amino acid D10, or a corresponding mutation, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 includes a mutation at D10A, or a corresponding mutation, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 polypeptide includes mutations at amino acids D10, G12, and / or G17, or corresponding mutations, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 polypeptide includes mutations at D10A, G12A, and / or G17A, or corresponding mutations, compared to the wild-type SpCas9 described in SEQ ID NO: 674.

[0227] In some embodiments, the prime editor includes a Cas9 polypeptide having a mutation in the HNH domain that reduces or disables the nuclease activity of the HNH domain. In some embodiments, the Cas9 polypeptide includes a mutation at amino acid H840, or a corresponding mutation, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 polypeptide includes a mutation at H840A, or a corresponding mutation, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 polypeptide includes mutations at amino acids E762, D839, H840, N854, N856, N863, H982, H983, A984, D986, and / or A987, or a corresponding mutation, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 polypeptide includes mutations in E762A, D839A, H840A, N854A, N856A, N863A, H982A, H983A, A984A, and / or D986A, or corresponding mutations, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 polypeptide includes mutations in amino acid residues R221, N394, and / or H840, compared to the wild-type SpCas9 (e.g., SEQ ID NO: 674). In some embodiments, the Cas9 polypeptide includes mutations in R221K, N394L, and / or H840A, or corresponding mutations, compared to the wild-type SpCas9 described in SEQ ID NO: 674. In some embodiments, the Cas9 polypeptide includes mutations at amino acid residues R220, N393, and / or H839, or corresponding mutations, compared to wild-type SpCas9 lacking an N-terminal methionine (e.g., SEQ ID NO: 674). In some embodiments, the Cas9 polypeptide includes mutations at R220K, N393K, and / or H839A, or corresponding mutations, compared to wild-type SpCas9 lacking an N-terminal methionine (e.g., as described in SEQ ID NO: 674).

[0228] In some embodiments, the prime editor includes Cas9 having one or more amino acid substitutions in both the HNH domain and the RuvC domain that reduce or disable the nuclease activity of both the HNH domain and the RuvC domain. In some embodiments, the prime editor includes nuclease-inactive Cas9, or nuclease-dead Cas9 (dCas9). In some embodiments, the dCas9 includes the H840X substitution and the D10X mutation, or corresponding mutations, compared to the wild-type SpCas9 described in SEQ ID NO: 674, where X is any amino acid other than H in the case of the H840X substitution, and any amino acid other than D in the case of the D10X substitution. In some embodiments, the dead Cas9 includes the H840A and D10A mutations, or corresponding mutations, compared to the wild-type SpCas9 described in SEQ ID NO: 674.

[0229] In some embodiments, the N-terminal methionine is removed from the amino acid sequence of the Cas9 nickase disclosed herein or intended herein, or from any Cas9 variant, orthologue, or equivalent. For example, methionine-minus (Met(-))Cas9 nickase comprises any one of the sequences described in SEQ ID NOs. 676, 677, 680, 683, 686, 689, 692, 695, 698, or 701, or any one of its variants having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.

[0230] In addition to dead Cas9 and Cas9 nickase variants, the Cas9 proteins used herein may also include other Cas9 variants having at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% sequence identity with any reference Cas9 protein, including any wild-type Cas9, or mutant Cas9 (e.g., dead Cas9 or Cas9 nickase), or Cas9 fragments, or circular permutations of Cas9, or other Cas9 variants disclosed herein or known in the art. In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to a reference Cas9, e.g., wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of reference Cas9 (e.g., a gRNA-binding domain or a DNA-cleaving domain) which is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of reference Cas9, e.g., wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas9.

[0231] In some embodiments, the Cas9 fragment is a functional fragment that retains one or more Cas9 activities. In some embodiments, the Cas9 fragment is at least 100 amino acids long. In some embodiments, the length of the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids long.

[0232] In some embodiments, the prime editor includes a Cas protein, e.g., Cas9, that contains modifications enabling the recognition of the altered PAM. In prime editing using a Cas protein-based prime editor, a “protospacer adjacent motif (PAM),” a PAM sequence, or a PAM-like motif may be used to refer to a short DNA sequence immediately following the protospacer sequence on the PAM strand of the target gene. In some embodiments, the PAM is recognized by a Cas nuclease within the prime editor during prime editing. In certain embodiments, the PAM is required for target binding of the Cas protein. The specific PAM sequence required for Cas protein recognition may vary depending on the particular type of Cas protein. PAMs can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides or longer. In some embodiments, the PAM is 2–6 nucleotides long. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). In some embodiments, the prime editor Cas protein recognizes a standard PAM; for example, SpCas9 recognizes a 5'-NGG-3' PAM. In some embodiments, the prime editor Cas protein has modified or non-standard PAM specificity. Exemplary PAM sequences and corresponding Cas variants are listed in Table 23 below. For each variant provided, it should be understood that the Cas protein contains one or more of the indicated amino acid substitutions compared to a wild-type Cas protein sequence, e.g., Cas9 described in Sequence ID No. 674. The PAM motifs shown in Table 23 below are in 5' to 3' order. In some embodiments, the Cas proteins of this disclosure may also be used to direct transcriptional control of a target sequence, e.g., transcriptional silencing by sequence-specific binding to a target sequence. In some embodiments, the Cas proteins described herein may have one or more mutations in the PAM recognition motif. In some embodiments, the Cas proteins described herein may have modified PAM specificity.

[0233] When used in the PAM sequences in Table 23, "N" refers to one of nucleotides A, G, C, and T; "R" refers to nucleotide A or G; "V" refers to nucleotide A or T; "V" refers to one of nucleotides A, C, or G; and "Y" refers to nucleotide C or T. [Table 2-1] [Table 2-2]

[0234] In some embodiments, the prime editor compares A61R, L111R, D1135V, R221K, A262T, R324L, N394K, S409I, S409I, E427G, E480K, M495V, N497A, Y515N, K526E, F539S, E543D, R654L, R661A, R661L, R691A, N692A, M694A, M694I, Q695A, H698A, R753G, M763I, K848A, K890N, Q926A, K1003A, R1060A, L1111R, R1114G, D11 The Cas9 polypeptide contains one or more mutations selected from the group consisting of 35E, D1135L, D1135N, S1136W, V1139A, D1180G, G1218K, G1218R, G1218S, E1219Q, E1219V, E1219V, Q1221H, P1249S, E1253K, N1317R, A1320V, P1321S, A1322R, I1322V, D1332G, R1332N, A1332R, R1333K, R1333P, R1335L, R1335Q, R1335V, T1337N, T1337R, S1338T, H1349R, and any combination thereof.

[0235] In some embodiments, the prime editor contains a SaCas9 polypeptide. In some embodiments, the SaCas9 polypeptide contains one or more mutations E782K, N968K, and R1015H compared to wild-type SaCas9. In some embodiments, the prime editor contains an FnCas9 polypeptide, e.g., wild-type FnCas9 polypeptide, or an FnCas9 polypeptide containing one or more mutations E1369R, E1449H, or R1556A compared to wild-type FnCas9. In some embodiments, the prime editor contains ScCas9, e.g., wild-type ScCas9, or the ScCas9 polypeptide contains one or more mutations I367K, G368D, I369K, H371L, T375S, T376G, and T1227K compared to wild-type ScCas9. In some embodiments, the prime editor contains an St1 Cas9 polypeptide, an St3 Cas9 polypeptide, or an SluCas9 polypeptide.

[0236] In some embodiments, the prime editor comprises a Cas polypeptide containing a circularly permuted Cas variant. For example, the Cas9 polypeptide of the prime editor may be designed such that the N-terminus and C-terminus of the Cas9 protein (e.g., wild-type Cas9 protein, or Cas9 nickase) are locally rearranged to retain the ability to bind to DNA when compounded with guide RNA (gRNA). An exemplary circularly permuted sequence configuration may be N-terminus-[original C-terminus]-[original N-terminus]-C-terminus. The Cas9 proteins described herein may include any variant, homologous gene, or naturally occurring Cas9 or its equivalent, and may be reconstituted as circularly permuted variants.

[0237] In various embodiments, circular permutation mutants of Cas proteins such as Cas9 can have the structure N-terminus-[original C-terminus]-[optional linker]-[original N-terminus]-C-terminus. In some embodiments, the circular permutation mutant Cas9 includes one of the following structures (amino acid positions described in SEQ ID NO: 674):

[0238] N-terminus - [1268~1368] - [optionally selectable linker] - [1~1267] - C-terminus,

[0239] N-terminus - [1168~1368] - [optionally selectable linker] - [1~1167] - C-terminus,

[0240] N-terminus - [1068~1368] - [optionally selectable linker] - [1~1067] - C-terminus,

[0241] N-terminus - [968~1368] - [optionally selectable linker] - [1~967] - C-terminus,

[0242] N-terminus - [868~1368] - [optionally selectable linker] - [1~867] - C-terminus,

[0243] N-terminus - [768~1368] - [optionally selectable linker] - [1~767] - C-terminus,

[0244] N-terminus - [668~1368] - [optionally selectable linker] - [1~667] - C-terminus,

[0245] N-terminus - [568~1368] - [optionally selectable linker] - [1~567] - C-terminus,

[0246] N-terminus - [468~1368] - [optionally selectable linker] - [1~467] - C-terminus,

[0247] N-terminus - [368~1368] - [optionally selectable linker] - [1~367] - C-terminus,

[0248] N-terminus - [268~1368] - [optionally selectable linker] - [1~267] - C-terminus,

[0249] N-terminus - [168~1368] - [optionally selectable linker] - [1~167] - C-terminus,

[0250] N-terminus - [68~1368] - [Optional linker] - [1~67] - C-terminus,

[0251] N-terminus-[10~1368]-[optional linker]-[1~9]-C-terminus, or the corresponding circular permutation variant of other Cas9 proteins (including other Cas9 orthologues, variants, etc.).

[0252] In some embodiments, the circular permutation mutant Cas9 comprises one of the following structures (amino acid positions as shown in SEQ ID NO: 674):

[0253] N-terminus - [102~1368] - [Optional linker] - [1~101] - C-terminus,

[0254] N-terminus - [1028~1368] - [Optional linker] - [1~1027] - C-terminus,

[0255] N-terminus - [1041~1368] - [Optional linker] - [1~1043] - C-terminus,

[0256] N-terminus-[1249~1368]-[Optional linker]-[1~1248]-C-terminus, or

[0257] N-terminus-[1300~1368]-[optional linker]-[1~1299]-C-terminus, or the corresponding circular permutation sequence of other Cas9 proteins (including other Cas9 orthologues, variants, etc.).

[0258] In some embodiments, the circular permutation mutant Cas9 includes one of the following structures (amino acid positions described in SEQ ID NO: 674 ~ 1368 amino acids of UniProtKB-Q99ZW2): N-terminus-[103~1368]-[optional linker]-[1~102]-C-terminus

[0259] N-terminus - [1029~1368] - [Optional linker] - [1~1028] - C-terminus,

[0260] N-terminus - [1042~1368] - [Optional linker] - [1~1041] - C-terminus,

[0261] N-terminus-[1250~1368]-[Optional linker]-[1~1249]-C-terminus, or

[0262] N-terminus-[1301~1368]-[optional linker]-[1~1300]-C-terminus, or the corresponding circular permutation sequence of other Cas9 proteins (including other Cas9 orthologues, variants, etc.).

[0263] In some embodiments, circular permutation variants can be formed by linking the C-terminal fragment of Cas9 to the N-terminal fragment of Cas9, either directly or by using a linker such as an amino acid linker. In some embodiments, the C-terminal fragment may correspond to 95% or more of the C-terminal amino acids of Cas9 (e.g., amino acids approximately 1300-1368 or the corresponding amino acid positions as described in SEQ ID NO: 674), or 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the C-terminal amino acids of Cas9 (e.g., SEQ ID NO: 674 or its ortholog or variant). The N-terminal portion may correspond to 95% or more of the N-terminal amino acids of Cas9 (for example, amino acids approximately 1 to 1300 listed in SEQ ID NO: 674 or the corresponding amino acid positions), or 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the N-terminal amino acids of Cas9 (for example, those listed in SEQ ID NO: 674 or the corresponding amino acid positions).

[0264] In some embodiments, circular permutation variants can be formed by ligating the C-terminal fragment of Cas9 to the N-terminal fragment of Cas9, either directly or by using a linker such as an amino acid linker. In some embodiments, the C-terminal fragment rearranged to the N-terminus contains or corresponds to 30% or less of the C-terminal amino acids of Cas9 (e.g., amino acids 1012-1368 as described in SEQ ID NO: 674, or the corresponding amino acid positions). In some embodiments, the C-terminal fragment rearranged at the N-terminus contains or corresponds to 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the C-terminal amino acids of Cas9 (e.g., those described in SEQ ID NO: 674 or the corresponding amino acid positions). In some embodiments, the C-terminal fragment rearranged at the N-terminus contains or corresponds to 410 residues or less of the C-terminal of Cas9 (e.g., those described in SEQ ID NO: 674 or the corresponding amino acid positions). In some embodiments, the C-terminal portion rearranged at the N-terminus contains or corresponds to the 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 residues of the C-terminus of Cas9 (for example, those described in SEQ ID NO: 674 or the corresponding amino acid positions). In some embodiments, the C-terminal portion rearranged at the N-terminus includes or corresponds to residues 357, 341, 328, 120, or 69 of the C-terminus of Cas9 (e.g., those described in SEQ ID NO: 674 or the corresponding amino acid positions).

[0265] In other embodiments, the circular permutation Cas9 variants can be topological rearrangements of the Cas9 primary structure based on the following method, which is based on Cas9 of S. pyogenes of SEQ ID NO: 674: (a) selecting a circular permutation variant (CP) site corresponding to an internal amino acid residue of the Cas9 primary structure that bisects the original protein into an N-terminal region and a C-terminal region; (b) modifying the Cas9 protein sequence (e.g., by genetic engineering techniques) by moving the original C-terminal region (including the amino acids of the CP site) forward of the original N-terminal region, thereby forming a new N-terminus of the Cas9 protein starting from the amino acid residue of the CP site. The CP site can be located in any domain of the Cas9 protein, including, for example, the helical II domain, the RuvCIII domain, or the CTD domain. For example, the CP site can be located at the original amino acid residues 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 (as described in SEQ ID NO: 674 or the corresponding amino acid positions). Thus, when rearranged to the N-terminus, the original amino acids 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 become the new N-terminal amino acids. The naming of these CP-Cas9 proteins is Cas9-CP 181 , Cas9-CP 199 , Cas9-CP 230 , Cas9-CP 270 , Cas9-CP 310 , Cas9-CP 1010 , Cas9-CP 1016 , Cas9-CP 1023 , Cas9-CP 1029 , Cas9-CP 1041 , Cas9-CP 1247 , Cas9-CP 1249 , and Cas9-CP 1282It can be called. This description is not intended to be limited to the creation of CP variants from SEQ ID NO: 674, and can be introduced to create CP variants of any Cas9 sequence at either the CP site corresponding to these positions or the entire other CP sites. This description is not intended to limit any particular CP site in any way. In fact, any CP site can be used for the formation of CP-Cas9 variants.

[0266] In some embodiments, the prime editor comprises a Cas9 functional variant having a smaller molecular weight than the wild-type SpCas9 protein. In some embodiments, the smaller Cas9 functional variant may facilitate delivery to cells, for example, by an expression vector, a nanoparticle, or other delivery means. In certain embodiments, the smaller Cas9 functional variant is a class 2 type II Cas protein. In certain embodiments, the smaller Cas9 functional variant is a class 2 type V Cas protein. In certain embodiments, the smaller Cas9 functional variant is a class 2 type VI Cas protein.

[0267] In some embodiments, the prime editor contains SpCas9 with a length of 1368 amino acids and a predicted molecular weight of 158 kilodaltons. In some embodiments, the prime editor contains less than 1300 amino acids, less than 1290 amino acids, less than 1280 amino acids, less than 1270 amino acids, less than 1260 amino acids, less than 1250 amino acids, less than 1240 amino acids, less than 1230 amino acids, less than 1220 amino acids, less than 1210 amino acids, less than 1200 amino acids, less than 1190 amino acids, less than 1180 amino acids, less than 1170 amino acids, less than 1160 amino acids, less than 1150 amino acids, less than 1140 amino acids, less than 1130 amino acids, less than 1120 amino acids. Full, less than 1110 amino acids, less than 1100 amino acids, less than 1050 amino acids, less than 1000 amino acids, less than 950 amino acids, less than 900 amino acids, less than 850 amino acids, less than 800 amino acids, less than 750 amino acids, less than 700 amino acids, less than 650 amino acids, less than 600 amino acids, less than 550 amino acids, or less than 500 amino acids, but greater than at least about 400 amino acids, and containing one or more functions of the Cas9 protein, such as DNA binding function, as a Cas9 functional variant or functional fragment.

[0268] In some embodiments, the Cas protein is Cas12a, Cas12b1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also called Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, The system may include, but is not limited to, any CRISPR-related protein, including Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, their homologs, or modified versions thereof, preferably including nickase mutations (e.g., mutations corresponding to the D10A mutation in the wild-type Cas9 polypeptide of SEQ ID NO: 674). In various other embodiments, napDNAbp may be any of the following proteins: Cas9, Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), GeoCas9, CjCas9, Cas12g, Cas12h, Cas12i, Cas13b, Cas13c, Cas13d, Cas14, Csn2, xCas9, SpCas9-NG, circular permutation Cas9, or Argonaute (Ago) domain, or functional variants or fragments thereof.

[0269] Exemplary Cas proteins and their names are shown in Table 24 below. [Table 3]

[0270] In some embodiments, the prime editor described herein may also include Cas proteins other than Cas9. For example, in some embodiments, the prime editor described herein may include the Cas12a(Cpf1) polypeptide or a functional variant thereof. In some embodiments, the Cas12a polypeptide includes a mutation that reduces or eliminates the endonuclease domain of the Cas12a polypeptide. In some embodiments, the Cas12a polypeptide is Cas12a nickase. In some embodiments, the Cas protein includes an amino acid sequence that has at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the spontaneously occurring Cas12a polypeptide.

[0271] In some embodiments, the prime editor comprises a Cas protein which is a Cas12b(C2c1) or Cas12c(C2c3) polypeptide. In some embodiments, the Cas protein comprises an amino acid sequence which has at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a naturally occurring Cas12b(C2c1) or Cas12c(C2c3) protein. In some embodiments, the Cas protein is Cas12b nickase or Cas12c nickase. In some embodiments, the Cas protein is a Cas12e, Cas12d, Cas13, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or CasΦ polypeptide. In some embodiments, the Cas protein contains an amino acid sequence that has at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a naturally occurring Cas12e, Cas12d, Cas13, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, or CasΦ protein. In some embodiments, the Cas protein is Cas12e, Cas12d, Cas13, or CasΦ nickase.

[0272] nuclear localization sequence In some embodiments, the prime editor further comprises one or more nuclear localization sequences (NLS). In some embodiments, the NLS facilitates the translocation of the protein to the cell nucleus. In some embodiments, the prime editor comprises a fusion protein, e.g., a fusion protein comprising a DNA-binding domain and a DNA polymerase containing one or more NLS. In some embodiments, one or more polypeptides of the prime editor are fused or ligated to one or more NLS. In some embodiments, the prime editor comprises a DNA-binding domain and a DNA polymerase domain provided in trans, and the DNA-binding domain and / or DNA polymerase domain are fused or ligated to one or more NLS.

[0273] In certain embodiments, the prime editor or prime editing complex includes at least one NLS. In some embodiments, the prime editor or prime editing complex includes at least two NLSs. In embodiments having at least two NLSs, the NLSs may be the same NLS or different NLSs.

[0274] In some cases, the prime editor may further contain at least one more nuclear localization sequence (NLS). In some cases, the prime editor may further contain one more NLS. In some cases, the prime editor may further contain two more NLS. In other cases, the prime editor may further contain three more NLS. In some cases, the prime editor may further contain four, five, six, seven, eight, nine, or more than ten NLS.

[0275] Furthermore, NLS can be expressed as part of the Prime Editor complex. In some embodiments, NLS can be located at virtually any position in the amino acid sequence of the protein and generally consist of a short sequence of three or more or four or more amino acids. The NLS fusion site may be at the N-terminus, the C-terminus, or any position in the sequence of the Prime Editor or its components (for example, it may be inserted between the DNA-binding domain and the DNA polymerase domain of the Prime Editor fusion protein, between the DNA-binding domain and the linker sequence, between the DNA polymerase and the linker sequence, or between two linker sequences of the Prime Editor fusion protein or its components, either in an N-terminus to C-terminus or C-terminus order). In some embodiments, the Prime Editor is a fusion protein containing an NLS at the N-terminus. In some embodiments, the Prime Editor is a fusion protein containing an NLS at the C-terminus. In some embodiments, the Prime Editor is a fusion protein containing at least one NLS at both the N-terminus and the C-terminus. In some embodiments, the Prime Editor is a fusion protein containing two NLS at the N-terminus and / or C-terminus.

[0276] Any NLS known in the art is also considered herein. The NLS may be any spontaneously occurring NLS or any non-spontaneously occurring NLS (e.g., an NLS having one or more mutations compared to a wild-type NLS). In some embodiments, one or more NLS of the prime editor include bipartite NLS. In some embodiments, the nuclear localization signal (NLS) is primarily basic. In some embodiments, one or more NLS of the prime editor are rich in lysine and arginine residues. In some embodiments, one or more NLS of the prime editor include proline residues. In some embodiments, the nuclear localization signal (NLS) includes the sequence MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 702), KRTADGSEFESPKKKRKV (SEQ ID NO: 703), KRTADGSEFEPKKKRKV (SEQ ID NO: 704), NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 705), RQRRNELKRSF (SEQ ID NO: 706), or NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 707).

[0277] In some embodiments, the NLS is a unipolar NLS. For example, in some embodiments, the NLS is the SV40 large T antigen NLS PKKKRKV (SEQ ID NO: 708). In some embodiments, the NLS is a bipartite NLS. In some embodiments, the bipartite NLS comprises two basic domains separated by a spacer sequence containing a variable number of amino acids. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the bipartite NLS consists of two basic domains separated by a spacer sequence containing a varying number of amino acids. In some embodiments, the amino acid sequence of the spacer includes the sequence KRXXXXXXXXXXKKKL (African clawed frog nucleoplasmic NLS) (SEQ ID NO: 709), where X is any amino acid. In some embodiments, the NLS includes the nucleoplasmic NLS sequence KRPAATKKAGQAKKKK (SEQ ID NO: 710). In some embodiments, the NLS is a non-standard sequence such as the M9 of the hnRNP Al protein, the influenza virus nucleoprotein NLS, and the yeast Gal4 protein NLS. In some embodiments, the NLS is a non-standard sequence such as M9 of the hnRNP Al protein, the influenza virus nucleoprotein NLS, and the yeast Gal4 protein NLS.

[0278] Other non-limiting examples of NLS sequences are shown in Table 25 below. In some embodiments, a bipartite NLS consists of two basic domains separated by a spacer sequence containing a varying number of amino acids. In some embodiments, the NLS contains an amino acid sequence that is at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs. 702–720. In some embodiments, the NLS contains an amino acid sequence selected from the group consisting of SEQ ID NOs. 702–720. In some embodiments, the prime-edited composition comprises a polynucleotide encoding an NLS containing an amino acid sequence that is at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs.

[0279] Any NLS known in the art is also considered herein. An NLS may be any spontaneously occurring NLS or any non-spontaneously occurring NLS (e.g., an NLS having one or more mutations compared to a wild-type NLS). In some embodiments, one or more NLS of the prime editor comprise a bipartite NLS. In some embodiments, one or more NLS of the prime editor are rich in lysine and arginine residues. In some embodiments, one or more NLS of the prime editor comprise a proline residue. Non-limiting examples of NLS sequences are shown in Table 25 below. [Table 4]

[0280] In some embodiments, the prime editing complex comprises a fusion protein having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)], which includes a DNA-binding domain (e.g., Cas9(H840A)) and a reverse transcriptase (e.g., variant MMLV RT), as well as a desired PEgRNA. In some embodiments, the prime editing complex comprises a prime editor fusion protein having the amino acid sequence of SEQ ID NO: 740. Table 26 shows the sequence and components of an exemplary prime editor fusion protein containing a DNA-binding domain (e.g., Cas9(H840A)) and reverse transcriptase (e.g., variant MMLV RT), having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)].

[0281] In some embodiments, the prime editing complex comprises a fusion protein having the following structure: [NLS]-[Cas9((R221K N394K H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)], including a DNA-binding domain (e.g., Cas9(R221K N394K H840A)) and a reverse transcriptase (e.g., variant MMLV RT), as well as a desired PEgRNA. In some embodiments, the prime editing complex comprises a prime editor fusion protein having the amino acid sequence of SEQ ID NO: 741. The following structure: [NLS]-[Cas9(R221K N394K Table 27 shows the sequence and components of an exemplary prime editor fusion protein containing a DNA-binding domain (e.g., Cas9(H840A)) and a reverse transcriptase (e.g., variant MMLV RT), having H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)].

[0282] Polypeptides containing components of a prime editor may be fused via a peptide linker or provided in trans relative to each other. For example, reverse transcriptase may be expressed, delivered, or provided as a separate component rather than as part of a fusion protein with a DNA-binding domain. In such cases, components of the prime editor may be associated through non-peptide bonds or co-localization functions. In some embodiments, the prime editor further includes additional components capable of interacting with, associating with, or recruiting other components of the prime editor or prime editing system. For example, the prime editor may include an RNA-protein mobilization polypeptide that can associate with an RNA-protein mobilization RNA aptamer. In some embodiments, the RNA-protein mobilization polypeptide may or may be mobilized by recruiting a specific RNA sequence. Non-limiting examples of RNA-protein mobilization polypeptide and RNA aptamer pairs include the MS2 coat protein and MS2 RNA hairpin, the PCP polypeptide and PP7 RNA hairpin, the Com polypeptide and Com RNA hairpin, the Ku protein and telomerase Ku-binding RNA motif, and the Sm7 protein and telomerase Sm7-binding RNA motif. In some embodiments, the prime editor includes a DNA-binding domain fused to or ligated to an RNA-protein-mobilizing polypeptide. In some embodiments, the prime editor includes a DNA polymerase domain fused to or ligated to an RNA-protein-mobilizing polypeptide. In some embodiments, the DNA-binding domain and DNA polymerase domain fused to the RNA-protein-mobilizing polypeptide, or the DNA-binding domain fused to the RNA-protein-mobilizing polypeptide and DNA polymerase domain, are colocalized by the corresponding RNA-protein-mobilizing RNA aptamer of the RNA-protein-mobilizing polypeptide. In some embodiments, the corresponding RNA-protein-mobilizing RNA aptamer is fused to or ligated to a portion of PEgRNA or ngRNA.For example, the MS2 coat protein is fused to or ligated to the DNA polymerase and the MS2 hairpin introduced on the PEgRNA for co-localization of the DNA polymerase and the RNA guide DNA binding domain (e.g., Cas9 nickase). In certain embodiments, the components of the prime editor are fused to each other directly. In certain embodiments, the components of the prime editor associate to each other via a linker.

[0283] In some embodiments, the prime editor includes an MS2 coat protein (MCP), which is a polypeptide domain that recognizes the MS2 hairpin. In some embodiments, the nucleotide sequence of the MS2 hairpin (also called the "MS2 aptamer") is GCCAACATGAGGATCACCCATGTCTGCAGGGCC (SEQ ID NO: 721). In some embodiments, the amino acid sequence of the MCP is GSASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKVATQTVGGEELPVAGWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIA ANSGIY (SEQ ID NO: 722).

[0284] As used herein, a linker can be any chemical group or molecule that links two molecules or parts, for example, the DNA-binding domain and the polymerase domain of a prime editor. In some embodiments, the linker is an organic molecule, group, polymer, or chemical part. In some embodiments, the linker includes a non-peptide part. The linker can be as simple as a covalent bond or a polymer linker consisting of many atoms in length, such as a polynucleotide sequence. In certain embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.).

[0285] In certain embodiments, two or more components of the prime editor are linked together by a peptide linker. In some embodiments, the peptide linker is 5 to 100 amino acid long, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acid long. In some embodiments, the peptide linker is 16 amino acid long, 24 amino acid long, 64 amino acid long, or 96 amino acid long.

[0286] In some embodiments, the linker includes the amino acid sequence (GGGGS)n (SEQ ID NO: 723), (G)n (SEQ ID NO: 724), (EAAAK)n (SEQ ID NO: 725), (GGS)n (SEQ ID NO: 726), (SGGS)n (SEQ ID NO: 727), (XP)n (SEQ ID NO: 728), or any combination thereof, where n is an integer independently between 1 and 30, and X is any amino acid. In some embodiments, the linker includes the amino acid sequence (GGS)n (SEQ ID NO: 726), where n is 1, 3, or 7. In some embodiments, the linker includes the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 729). In some embodiments, the linker includes the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 730). In some embodiments, the linker includes the amino acid sequence SGGSGGSGGS (SEQ ID NO: 731). In some embodiments, the linker includes the amino acid sequence SGGS (SEQ ID NO: 732). In other embodiments, the linker includes the amino acid sequence SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSSGGS (SEQ ID NO: 733).

[0287] In some embodiments, the linker contains 1 to 100 amino acids. In some embodiments, the linker contains the amino acid sequence GGSGGS (SEQ ID NO: 734), GGSGGSGGS (SEQ ID NO: 735), or SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 736).

[0288] In certain embodiments, two or more components of the prime editor are linked to each other by a non-peptide linker. In some embodiments, the linker is a carbon-nitrogen bond of an amide bond. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker. In certain embodiments, the linker is a polymer (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, acetic acid, alanine, β-alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of an aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane, etc.). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In certain embodiments, the linker includes an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker may include a functionalization moiety to facilitate the binding of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0289] The components of the prime editor can be linked to each other in any order. In some embodiments, the DNA-binding domain and DNA polymerase domain of the prime editor may be fused to form a fusion protein, or they may be linked by a peptide or protein linker in any order from the N-terminus to the C-terminus. In some embodiments, the prime editor includes a DNA-binding domain fused or linked to the C-terminus of the DNA polymerase domain. In some embodiments, the prime editor includes a DNA-binding domain fused or linked to the N-terminus of the DNA polymerase domain. In some embodiments, the prime editor includes a fusion protein having the structure NH2-[DNA-binding domain]-[polymerase]-COOH, or NH2-[polymerase]-[DNA-binding domain]-COOH, where each ]-[ indicates the presence of an arbitrary linker sequence. In some embodiments, the prime editor includes a fusion protein and a DNA polymerase domain provided in trans, and the fusion protein has the structure NH2-[DNA-binding domain]-[RNA protein-mobilizing polypeptide]-COOH. In some embodiments, the prime editor comprises a fusion protein and a DNA-binding domain provided in trans, the fusion protein having the structure NH2-[DNA polymerase domain]-[RNA protein-mobilizing polypeptide]-COOH.

[0290] In some embodiments, the prime editor fusion protein, the polypeptide component of the prime editor, or a polynucleotide encoding the prime editor fusion protein or polypeptide component is split into an N-terminal half and a C-terminal half, or a polypeptide encoding the N-terminal half and a C-terminal half, and delivered separately to target DNA in the cell. For example, in a particular embodiment, the prime editor fusion protein is split into N-terminal and C-terminal halves for separate delivery in an AAV vector, and is then translated and co-localized in the target cell to reassemble the complete polypeptide or prime editor protein. In such a case, each of the separate halves of the protein or fusion protein contains a split intein, which facilitates the co-localization and reassembly of the complete protein or fusion protein by an intein-enhanced trans-splicing mechanism. In some embodiments, the prime editor includes an N-terminal half fused to intein N and a C-terminal half fused to intein C, or a polynucleotide or vector (e.g., an AAV vector) encoding each of them. Intein N and Intein C, upon delivery to and / or expression in target cells, can be excised by protein trans-splicing to generate a complete prime editor fusion protein within the target cells. In some embodiments, the exemplary proteins described herein may lack an N-terminal methionine residue.

[0291] In some embodiments, the prime editor fusion protein comprises Cas9(H840A) nickase and wild-type M-MLV RT. In some embodiments, the prime editor fusion protein comprises Cas9(H840A) nickase and M-MLV RT containing amino acid substitutions D200N, T330P, T306K, W313F, and L603W compared to wild-type M-MLV RT. In some embodiments, the prime editor fusion protein comprises Cas9(H840A) nickase and M-MLV RT containing amino acid substitutions D200N, T330P, T306K, W313F, and L603W compared to wild-type M-MLV RT. The amino acid sequences of exemplary prime editor fusion proteins and their individual components are shown in Table 26. In some embodiments, the prime editor fusion protein comprises Cas9(R221K N394K H840A) nickase and M-MLV RT containing amino acid substitutions D200N, T330P, T306K, W313F, and L603W compared to wild-type M-MLV RT. The amino acid sequences of exemplary prime editor fusion proteins and their individual components are shown in Table 27. In some embodiments, the exemplary prime editor protein may contain the amino acid sequence described in either SEQ ID NO: 740 or SEQ ID NO: 741.

[0292] In various embodiments, the prime editor fusion protein contains an amino acid sequence that is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to PE1, PE2, or any of the prime editor fusion sequences described herein or known in the art. [Table 5-1] [Table 5-2] [Table 5-3] [Table 5-4] [Table 6-1] [Table 6-2] [Table 6-3]

[0293] PEgRNA for editing the B2M gene The term “prime edit guide RNA” or “PEgRNA” refers to a guide polynucleotide containing one or more intended nucleotide edits (i.e., one or more nucleotide changes) for incorporation into target DNA. In some embodiments, the PEgRNA associates with a prime editor to instruct it to incorporation one or more intended nucleotide edits into the target gene via prime editing. “Nucleotide edit” or “intended nucleotide edit” refers to a specific deletion of one or more nucleotides at one specific location, an insertion of one or more nucleotides at one specific location, a substitution of a single nucleotide, or any other modification at one specific location that is incorporated into the sequence of the target gene. An intended nucleotide edit may refer to an edit on an editing template compared to the sequence on the target strand of the target gene, or it may refer to an edit encoded by an editing template on newly synthesized single-stranded DNA that replaces the edit target sequence compared to the edit target sequence. In some embodiments, the incorporation of one or more intended nucleotide edits into a target B2M gene results in a mutation in the target B2M gene (e.g., a missense mutation, a nonsense mutation, a frameshift mutation, a null mutation, a mutation that produces an immature stop codon, or a combination thereof). In some embodiments, one or more intended nucleotide edits introduce a frameshift mutation into the target gene (e.g., the B2M gene) and / or generate one or more immature stop codons (e.g., at least 1, 2, 3, 4, 5, or more immature stop codons). In some embodiments, one or more intended nucleotide edits generate at least two immature stop codons in the target gene. In some embodiments, one or more intended nucleotide edits generate at least 2, 3, 4, 5, or more consecutive immature stop codons in the target gene (e.g., the B2M gene). In some embodiments, one or more intended nucleotide edits include the insertion of one or more immature in-frame stop codons (e.g., two stop codons) into the B2M gene.In some embodiments, the PEgRNA includes a spacer sequence that is complementary or substantially complementary to the target sequence on the target strand of the target gene. In some embodiments, the PEgRNA includes a gRNA core associated with the DNA-binding domain of the prime editor, e.g., the CRISPR-Cas protein domain. In some embodiments, the PEgRNA further includes an extended nucleotide sequence containing one or more intended nucleotide edits compared to the endogenous sequence of the target gene, the extended nucleotide sequence may also be called an extension arm.

[0294] In certain embodiments, the extension arm includes a primer-binding site (PBS) sequence on which targeted priming DNA synthesis can be initiated. In some embodiments, the PBS is complementary or substantially complementary to the free 3' end on the edited strand of the target gene at a nick site generated by the prime editor. In some embodiments, the extension arm further includes an editing template containing one or more intended nucleotide edits to be incorporated into the target gene by prime editing. In some embodiments, the editing template is a template of the RNA-dependent DNA polymerase domain or polypeptide of the prime editor, e.g., a reverse transcriptase domain. The reverse transcriptase editing template may also be referred to herein as an RT template or RTT. In some embodiments, the editing template includes partial complementarity to the target editing sequence in the target gene, e.g., the B2M gene. In some embodiments, the editing template includes substantial or partial complementarity to the editing sequence, except for the nucleotide editing sites intended to be incorporated into the target gene. Exemplary structures of PEgRNAs containing its components are shown in Figure 2.

[0295] In some embodiments, PEgRNA contains only RNA nucleotides and forms an RNA polynucleotide. In some embodiments, PEgRNA is a chimeric polynucleotide containing both RNA nucleotides and DNA nucleotides. For example, PEgRNA may contain DNA in a spacer sequence, a gRNA core, or an extension arm. In some embodiments, PEgRNA contains DNA in a spacer sequence. In some embodiments, the entire spacer sequence of PEgRNA is a DNA sequence. In some embodiments, PEgRNA contains DNA in a gRNA core, for example, in the stem region of the gRNA core. In some embodiments, PEgRNA contains DNA in an extension arm, for example, in an editing template. An editing template containing a DNA sequence can function as a DNA synthesis template for a DNA polymerase in a prime editor, for example, a DNA-dependent DNA polymerase. Thus, PEgRNA can be a chimeric polynucleotide containing RNA in a spacer, a gRNA core, and / or a PBS sequence, as well as DNA in an editing template.

[0296] The components of PEgRNA may be arranged in a modular manner. In some embodiments, an extension arm containing a spacer and a primer-binding site sequence (PBS) and an editing template, such as a reverse transcriptase template (RTT), may be interchangeably positioned in the 5' portion of the PEgRNA, the 3' portion of the PEgRNA, or in the center of the gRNA core. In some embodiments, the PEgRNA contains the PBS and editing template sequence in the order from 5' to 3'. In some embodiments, the gRNA core of the PEgRNA of this disclosure may be located between the spacer and the extension arm of the PEgRNA. In some embodiments, the gRNA core of the PEgRNA may be located at the 3' end of the spacer. In some embodiments, the gRNA core of the PEgRNA may be located at the 5' end of the spacer. In some embodiments, the gRNA core of the PEgRNA may be located at the 3' end of the extension arm. In some embodiments, the gRNA core of the PEgRNA may be located at the 5' end of the extension arm. In some embodiments, the PEgRNA contains a spacer, a gRNA core, and an extension arm, arranged from 5' to 3'. In some embodiments, the PEgRNA comprises a spacer, a gRNA core, an editing template, and PBS from 5' to 3'. In some embodiments, the PEgRNA comprises an extension arm, a spacer, and a gRNA core from 5' to 3'. In some embodiments, the PEgRNA comprises an editing target, PBS, a spacer, and a gRNA core from 5' to 3'.

[0297] In some embodiments, PEgRNA comprises a single polynucleotide molecule including a spacer sequence, a gRNA core, and an extension arm. In some embodiments, PEgRNA comprises multiple polynucleotide molecules, for example, two polynucleotide molecules. In some embodiments, PEgRNA comprises a first polynucleotide molecule including a spacer and a portion of the gRNA core, and a second polynucleotide molecule including the remainder of the gRNA core and the extension arm. In some embodiments, the gRNA core portion in the first polynucleotide molecule and the gRNA core portion in the second polynucleotide molecule are at least partially complementary to each other. In some embodiments, PEgRNA may comprise a first polynucleotide including a spacer and a first portion of the gRNA core, which may also be referred to as crRNA. In some embodiments, PEgRNA comprises a second polynucleotide including a second portion of the gRNA core and an extension arm, the second portion of the gRNA core may also be referred to as transactivated crRNA or tracr RNA. In some embodiments, the crRNA portion and the tracr RNA portion of the gRNA core are at least partially complementary to each other. In some embodiments, the partially complementary portions of the crRNA and tracrRNA form a lower stem, bulge, and upper stem, as illustrated in Figure 3. The gRNA core (also called the gRNA scaffold or gRNA skeleton) of PEgRNA is analogous to the gRNA scaffold or skeleton of classical CRISPR-Cas9 guide RNA and CRISPR-Cas9 single guide RNA (sgRNA). gRNA core structures and exemplary sequences described in PCT Publication WO2020191234, Nowak et al., Nucleic Acids Research 44(20):95559564(2016), Doudna et al. Annu. Rev. Biophys. 2017. 46:505-29, as well as any gRNA core structures and sequences known in the art, are incorporated herein by reference as a whole.

[0298] In some embodiments, the spacer sequence includes a region substantially complementary to the target sequence on the target strand of a double-stranded target DNA, e.g., the B2M gene. In some embodiments, the spacer sequence of PEgRNA is identical or substantially identical to the protospacer sequence on the edited strand of the target gene (wherein the protospacer sequence may contain thymine, and the spacer sequence may contain uracil). In some embodiments, the spacer sequence is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to the target sequence in the target gene. In some embodiments, the spacer is substantially complementary to the target sequence.

[0299] In some embodiments, the length of the spacer varies from about 10 to about 100 nucleotides. In some embodiments, the spacer is 16 nucleotides long, 17 nucleotides long, 18 nucleotides long, 19 nucleotides long, 20 nucleotides long, 21 nucleotides long, 22 nucleotides long, 23 nucleotides long, 24 nucleotides long, or 25 nucleotides long. In some embodiments, the spacer is 15 to 30 nucleotides long, 15 to 25 nucleotides long, 18 to 22 nucleotides long, 10 to 20 nucleotides long, or 20 to 30 nucleotides long. In some embodiments, the spacer is 16 to 22 nucleotides long, for example, about 16, 17, 18, 19, 20, 21, or 22 nucleotides long.

[0300] When used herein in PEgRNA or nick guide RNA sequences, or fragments thereof, such as spacers, PBS, or RTT sequences, unless otherwise specified, the letter "T" or "thymine" naturally refers to a nucleic acid base in the DNA sequence encoding the PEgRNA or guide RNA sequence, and is intended to refer to the uracil (U) nucleic acid base of the PEgRNA or guide RNA, or any chemically modified uracil nucleic acid base known in the art, such as 5-methoxyuracil.

[0301] The extension arm of PEgRNA may include a primer-binding site (PBS) and an editing template (e.g., RTT). The extension arm may be partially complementary to the spacer. In some embodiments, the editing template (e.g., RTT) is partially complementary to the spacer. In some embodiments, the editing template (e.g., RTT) and the primer-binding site (PBS) are each partially complementary to the spacer.

[0302] The extension arm of PEgRNA may contain a primer-binding site sequence (PBS, or PBS sequence) that is complementary to and can hybridize with the free 3' end of single-stranded DNA within a target gene (e.g., the B2M gene) generated by nicking by a prime editor at the nicking site of the PAM strand.

[0303] The length of the PBS sequence may vary, for example, depending on the components of the prime editor, the search target sequence, and other components of the PEgRNA.

[0304] In some embodiments, the PBS is about 3 to 19 nucleotides long. In some embodiments, the PBS is about 3 to 17 nucleotides long. In some embodiments, the PBS is about 4 to 16 nucleotides long, about 6 to 16 nucleotides long, about 6 to 18 nucleotides long, about 6 to 20 nucleotides long, about 8 to 20 nucleotides long, about 10 to 20 nucleotides long, about 12 to 20 nucleotides long, about 14 to 20 nucleotides long, about 16 to 20 nucleotides long, or about 18 to 20 nucleotides long. In some embodiments, the PBS is 8 to 17 nucleotides long. In some embodiments, the PBS is 8 to 16 nucleotides long. In some embodiments, the PBS is 8 to 15 nucleotides long. In some embodiments, the PBS is 8 to 14 nucleotides long. In some embodiments, the PBS is 8 to 13 nucleotides long. In some embodiments, the PBS is 8 to 12 nucleotides long. In some embodiments, the PBS is 8 to 11 nucleotides long. In some embodiments, the PBS is 8 to 10 nucleotides long. In some embodiments, the PBS is 8 or 9 nucleotides long. In some embodiments, the PBS is 16 or 17 nucleotides long. In some embodiments, the PBS is 15 to 17 nucleotides long. In some embodiments, the PBS is 14 to 17 nucleotides long. In some embodiments, the PBS is 13 to 17 nucleotides long. In some embodiments, the PBS is 12 to 17 nucleotides long. In some embodiments, the PBS is 11 to 17 nucleotides long. In some embodiments, the PBS is 10 to 17 nucleotides long. In some embodiments, the PBS is 9 to 17 nucleotides long. In some embodiments, the PBS is about 7 to 15 nucleotides long. In some embodiments, the PBS is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 nucleotides long. In some embodiments, the PBS is 8 to 14 nucleotides long. For example, the PBS may have a length of 8, 9, 10, 11, 12, 13, or 14 nucleotides.In some embodiments, the PBS is 11 or 12 nucleotides long. In some embodiments, the PBS is 11 to 13 nucleotides long. In some embodiments, the PBS is 11 to 14 nucleotides long.

[0305] The PBS may be complementary or substantially complementary to the DNA sequence in the edited strand of the target gene. The PBS may initiate the synthesis of new single-stranded DNA encoded by the editing template at the nick site by annealing with the free hydroxyl group of the edited strand, e.g., the free 3' end generated by the nick of the prime editor. In some embodiments, the PBS is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to the region of the edited strand of the target gene (e.g., the B2M gene). In some embodiments, the PBS is perfectly complementary, i.e., 100%, to the region of the edited strand of the target gene (e.g., the B2M gene).

[0306] The extension arm of PEgRNA may contain an editing template that acts as a DNA synthesis template for DNA polymerase within the prime editor during prime editing.

[0307] The length of the editing template may vary depending, for example, on the components of the prime editor, the search target sequence, and other components of the PEgRNA. In some embodiments, the editing template functions as a DNA synthesis template for reverse transcriptase, and the editing template is called a reverse transcription editing template (RTT).

[0308] In some embodiments, the editing template (e.g., RTT) is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. In some embodiments, the RTT is 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. In some embodiments, the RTT is 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides long. In some embodiments, the RTT is 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides long. In some embodiments, the RTT is 10 to 110 nucleotides long. In some embodiments, the round-trip time (RTT) is 10–10⁹, 10–10⁸, 10–10⁷, 10–10⁶, 10–10⁵, 10–10⁴, 10–10⁃, 10–10⁂, or 10–10⁻¹ nucleotides long. In some embodiments, the RTT is at least 8 and up to 50 nucleotides long. In some embodiments, the RTT is at least 8 and up to 25 nucleotides long. In some embodiments, the RTT is about 10–20 nucleotides long. In some embodiments, the RTT is about 11, 12, 13, 14, 15, 16, 17, 18, or 19 nucleotides long. In some embodiments, the RTT is 11–17 nucleotides long. In some embodiments, the RTT is 12–17 nucleotides long. In some embodiments, the RTT is 12–16 nucleotides long. In some embodiments, the RTT is 13–17 nucleotides long. In some embodiments, the RTT is 11, 12, 13, 14, 15, 16, or 17 nucleotides long.In some embodiments, the RTT is 12 nucleotides long. In some embodiments, the RTT is 16 nucleotides long. In some embodiments, the RTT is 17 nucleotides long.

[0309] In some embodiments, the editing template (e.g., RTT) sequence is approximately 70%, 75%, 80%, 85%, 90%, 95%, or 99% complementary to the target sequence on the edited strand of the target gene. In some embodiments, the editing template sequence (e.g., RTT) is substantially complementary to the sequence to be edited. In some embodiments, the editing template sequence (e.g., RTT) is complementary to the sequence to be edited except for the nucleotide editing sites intended to be incorporated into the target gene. In some embodiments, the editing template comprises a nucleotide sequence having approximately 85% to approximately 95% complementarity to the target sequence in the edited strand of the target gene (e.g., the B2M gene). In some embodiments, the editing template includes approximately 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementarity with respect to the sequence to be edited within the edited strand of the target gene (e.g., the B2M gene).

[0310] In some embodiments, the editing template may be configured to introduce one or more recombinase-recognizing sequences into a target gene, e.g., the B2M gene. For example, the editing template may encode, or further encode, one or more recombinase-recognizing sequences (RRS). Such an editing template, or RTT, may enable the insertion of an RRS(or RRS) into the target gene, which may be used as a landing site for recombinase-mediated DNA insertion, deletion, inversion, or substitution. For example, in some embodiments, prime editing insertion of an RRS(e.g., attB sequence) enables the incorporation of a DNA donor sequence mediated by a recombinase (e.g., Bxb1) that recognizes the RRS, which also includes an RRS(e.g., attP sequence) recognized by the recombinase. In some embodiments, prime editing insertions of two RRSs allow for deletion of a target gene sequence between the two RRSs, or inversion of a target gene sequence between the two RRSs mediated by a recombinase that recognizes the two RRSs, depending on the orientation of the two RRSs. In some embodiments, prime editing insertions of two RRSs allow for cassette exchange between a target gene sequence and a DNA donor sequence between the two RRSs mediated by the corresponding recombinase, with two RRSs similarly recognized by the recombinase positioned alongside the DNA donor sequence.

[0311] Table 32 shows exemplary RRS sequences that can be encoded by PEgRNA RTT. Those skilled in the art will understand that RRS recognized by the same recombinase can be used for targeted insertion and other recombination events, for example, by inserting an attB sequence into a target B2M gene via prime editing, and by providing a circular DNA donor construct containing the Bxb1 recombinase and attP sequence for incorporating a DNA donor sequence into the attB site of the B2M gene. In some embodiments, orthogonal recognition can be achieved by altering the central dinucleotide of the RRS. For example, the central dinucleotide of the attB or attP sequence of Bxb1 may be GT or GA, shown in bold in Table 32. In some embodiments, the central dinucleotide of the RRS may be any two nucleotides, where each nucleotide is A, T, G, or C. Additional RRS described herein and those known in the art, as well as their corresponding recombinases, are also contemplated. [Table 7-1] [Table 7-2]

[0312] The intended nucleotide edits in a PEgRNA editing template may include various types of modifications compared to the target gene sequence. In some embodiments, the nucleotide edit is a single nucleotide substitution compared to the target gene sequence. In some embodiments, the nucleotide edit is a deletion compared to the target gene sequence. In some embodiments, the nucleotide edit is an insertion compared to the target gene sequence. In some embodiments, the editing template includes 1 to 10 intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes one or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes two or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes three or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes four or more, five or more, or six or more intended nucleotide edits compared to the target gene sequence. In some embodiments, the editing template includes two single nucleotide substitutions, insertions, deletions, or any combination thereof compared to the target gene sequence. In some embodiments, the editing template includes three single nucleotide substitutions, insertions, deletions, or any combination thereof compared to the target gene sequence. In some embodiments, the editing template includes four, five, or six single nucleotide substitutions, insertions, deletions, or any combination thereof, compared to the target gene sequence. In some embodiments, the nucleotide substitution includes an adenine (A) to thymine (T) substitution. In some embodiments, the nucleotide substitution includes an A to guanine (G) substitution. In some embodiments, the nucleotide substitution includes an A to cytosine (C) substitution. In some embodiments, the nucleotide substitution includes a TA substitution. In some embodiments, the nucleotide substitution includes a TG substitution. In some embodiments, the nucleotide substitution includes a TC substitution. In some embodiments, the nucleotide substitution includes a G to A substitution.In some embodiments, the nucleotide substitution includes a substitution from G to T. In some embodiments, the nucleotide substitution includes a substitution from G to C. In some embodiments, the nucleotide substitution includes a substitution from C to A. In some embodiments, the nucleotide substitution includes a substitution from C to T. In some embodiments, the nucleotide substitution includes a substitution from C to G.

[0313] In some embodiments, the nucleotide insertion is at least 1, at least 2, at least 3, at least 4, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, or at least 20 nucleotides long. In some embodiments, the nucleotide insertion is 1-2 nucleotides long, 1-3 nucleotides long, 1-4 nucleotides long, 1-5 nucleotides long, 2-5 nucleotides long, 3-5 nucleotides long, 3-6 nucleotides long, 3-8 nucleotides long, 4-9 nucleotides long, 5-10 nucleotides long, 6-11 nucleotides long, 7-12 nucleotides long, 8-13 nucleotides long, 9-14 nucleotides long, 10-15 nucleotides long, 11-16 nucleotides long, 12-17 nucleotides long, 13-18 nucleotides long, 14-19 nucleotides long, or 15-20 nucleotides long. In some embodiments, the nucleotide insertion is a single nucleotide insertion. In some embodiments, the nucleotide insertion includes two nucleotide insertions. In some embodiments, one or more intended nucleotide edits introduce a frameshift mutation into a target gene (e.g., the B2M gene) and / or generate one or more immature stop codons (e.g., at least 1, 2, 3, 4, 5, or more immature stop codons). In some embodiments, one or more intended nucleotide edits generate at least two immature stop codons in the target gene. In some embodiments, one or more intended nucleotide edits generate at least 2, 3, 4, 5, or more consecutive immature stop codons in the target gene (e.g., the B2M gene). In some embodiments, one or more intended nucleotide edits include the insertion of one or more immature in-frame stop codons (e.g., two stop codons) into the B2M gene.A PEgRNA editing template may contain one or more intended nucleotide edits compared to the B2M gene being edited. The location of the intended nucleotide edit(s) may differ from that of other components of the PEgRNA or specific nucleotides (e.g., mutations) within the B2M target gene. In some embodiments, the nucleotide edit is performed in a region of the PEgRNA corresponding to or homologous to the protospacer sequence. In some embodiments, the nucleotide edit is performed in a region of the PEgRNA corresponding to the B2M gene region outside the protospacer sequence.

[0314] In some embodiments, the integration site of nucleotide editing in a target gene may be "upstream" and "downstream," which is intended to define the relevant positions of at least two regions or sequences within a nucleic acid molecule oriented in the 5' to 3' direction. For example, if the first sequence is located 5' to the second sequence, then within the DNA molecule, the first sequence is upstream of the second sequence. Thus, the second sequence is downstream of the first sequence.

[0315] In some embodiments, the site of nucleotide editing in the target gene may be determined based on the location of the nick site. In some embodiments, the intended nucleotide editing site is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides away from the nick site. In some embodiments, the intended nucleotide editing site is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides downstream of the nick site on the PAM strand (or non-target strand, or edited strand) of the double-stranded target DNA. In some embodiments, the intended nucleotide edit location within the editing template may be referenced by aligning the editing template with a partially complementary editing target sequence on the edited strand and referring to a nucleotide location on the edited strand in which the intended nucleotide edit is incorporated. Thus, in some embodiments, the nucleotide edit in the editing template corresponds to a location approximately 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides away from the nick site.In some embodiments, nucleotide editing in the editing template is performed from the nick site by approximately 0-2 nucleotides, 0-4 nucleotides, 0-6 nucleotides, 0-8 nucleotides, 0-10 nucleotides, 2-4 nucleotides, 2-6 nucleotides, 2-8 nucleotides, 2-10 nucleotides, 2-12 nucleotides, 4-6 nucleotides, 4-8 nucleotides, 4-10 nucleotides, 4-12 nucleotides, 4-14 nucleotides, 6-8 nucleotides, 6-10 nucleotides, 6-12 nucleotides, 6-14 nucleotides, 6-16 nucleotides, 8-10 nucleotides, 8-12 nucleotides, 8-14 nucleotides, 8-16 nucleotides, 8-18 nucleotides, 10-12 nucleotides, 10-14 nucleotides, 10-16 nucleotides, 10-18 nucleotides, 10-20 nucleotides, 12-14 nucleotides, 12-16 nucleotides, 12-18 nucleotides, 12-20 nucleotides, and 12-22 nucleotides. At positions corresponding to nucleotides, 14-16 nucleotides, 14-18 nucleotides, 14-20 nucleotides, 14-22 nucleotides, 14-24 nucleotides, 16-18 nucleotides, 16-20 nucleotides, 16-22 nucleotides, 16-24 nucleotides, 16-26 nucleotides, 18-20 nucleotides, 18-22 nucleotides, 18-24 nucleotides, 18-26 nucleotides, 18-28 nucleotides, 20-22 nucleotides, 20-24 nucleotides, 20-26 nucleotides, 20-28 nucleotides, 20-30 nucleotides, 30-40 nucleotides, 40-50 nucleotides, 50-60 nucleotides, 60-70 nucleotides, 70-80 nucleotides, 80-90 nucleotides, 90-100 nucleotides, 100-110 nucleotides, 110-120 nucleotides, 120-130 nucleotides, 130-140 nucleotides, or 140-150 nucleotides away.In some embodiments, when referring to the PAM chain (or non-targeted chain, or edited chain), nucleotide editing in the editing template involves approximately 0-2 nucleotides, 0-4 nucleotides, 0-6 nucleotides, 0-8 nucleotides, 0-10 nucleotides, 2-4 nucleotides, 2-6 nucleotides, 2-8 nucleotides, 2-10 nucleotides, 2-12 nucleotides, 4-6 nucleotides, 4-8 nucleotides, 4-10 nucleotides, 4-12 nucleotides, 4-14 nucleotides, 6-8 nucleotides, 6-10 nucleotides, 6-12 nucleotides, 6-14 nucleotides, 6-16 nucleotides, 8-10 nucleotides, 8-12 nucleotides, 8-14 nucleotides, 8-16 nucleotides, 8-18 nucleotides, 10-12 nucleotides, 10-14 nucleotides, 10-16 nucleotides, 10-18 nucleotides, 10-20 nucleotides, 12-14 nucleotides, 12-16 nucleotides, 12-18 nucleotides, 12-20 nucleotides, 12-22 nucleotides, 14-16 nucleotides, 14-18 nucleotides, 14-20 nucleotides, 14-22 nucleotides, 14-24 nucleotides, 16-18 nucleotides, 16-20 nucleotides, 16-22 nucleotides, 16-24 nucleotides, 16-26 nucleotides, 18-20 nucleotides, 18-22 nucleotides, 18-24 nucleotides, 18-26 nucleotides, 18-28 nucleotides, 20-22 nucleotides, 20- At positions corresponding to 24 nucleotides, 20-26 nucleotides, 20-28 nucleotides, 20-30 nucleotides, 30-40 nucleotides, 40-50 nucleotides, 50-60 nucleotides, 60-70 nucleotides, 70-80 nucleotides, 80-90 nucleotides, 90-100 nucleotides, 100-110 nucleotides, 110-120 nucleotides, 120-130 nucleotides, 130-140 nucleotides, or 140-150 nucleotides downstream. The relative positions of the intended nucleotide edit(s) and nick sites can be referred to by numbers. For example, in some embodiments, the nucleotide immediately downstream of the nick site on the PAM strand (or non-target strand, or edited strand) may be referred to as position 0.The nucleotide immediately upstream of a nick site on the PAM chain (or non-target chain, or edited chain) may be referred to as position -1. Nucleotides downstream of position 0 on the PAM chain may be referred to as positions +1, +2, +3, +4, ..., +n, while nucleotides upstream of position -1 on the PAM chain may be referred to as positions -2, -3, -4, ..., -n. Therefore, in some embodiments, when an editing template is aligned with a partially complementary editing target sequence by complementarity, the nucleotide of the editing template corresponding to position 0 may also be referred to as position 0 in the editing template; the nucleotides of the editing template corresponding to positions +1, +2, +3, +4, ..., +n on the PAM strand of the double-stranded target DNA may also be referred to as positions +1, +2, +3, +4, ..., +n in the editing template; the nucleotides of the editing template corresponding to positions -1, -2, -3, -4, ..., -n on the PAM strand of the double-stranded target DNA may also be referred to as positions -1, -2, -3, -4, ..., -n in the editing template; and even when the PEgRNA is considered an independent nucleic acid, positions +1, +2, +3, +4, ..., +n are 5' to position 0 in the editing template, and positions -1, -2, -3, -4, ..., -n are 3' to position 0. In some embodiments, the intended nucleotide edit is at position +n of the editing template relative to position 0. Thus, the intended nucleotide edit can be incorporated by prime editing at position +n of the PAM strand of the double-stranded target DNA (followed by the target strand of the double-stranded target DNA). The corresponding position of the intended nucleotide edit incorporated into the B2M gene may also be referenced based on the nicking position generated by the prime editor based on sequence homology and complementarity. For example, in embodiments, the number of nucleotides from the nucleotide edit incorporated into the B2M gene to the nick site (also referred to as the “nick-edit distance” excluding the nucleotide of the edit) can be determined by identifying, for example, the sequence complementarity between the spacer and the search target sequence, and between the editing template and the editing target sequence, based on the location of the nick site and the position of the nucleotide(s) corresponding to the intended nucleotide edit(s).In certain embodiments, the site of nucleotide editing can be any position downstream of the nick site on the edited strand (or PAM strand). As used herein, the distance between the nick site and the nucleotide edit refers to, for example, the furthest 5' position of the nucleotide edit of the nick that creates a 3' free end on the edited strand (i.e., the “closest position” of the nucleotide edit to the nick site) when the nucleotide edit involves insertion or deletion. For a particular edit, e.g., a non-synonymous edit resulting in an immature stop codon in a target gene, the nick-edit distance may also be referred to as the number of nucleotides from the nick site to the edit site (excluding the furthest 5' nucleotide in the edit). In some embodiments, the nick-edit distance is 2 to 106 nucleotides. In some embodiments, the nick-edit distance is 2 to 105, 2 to 104, 2 to 103, 2 to 102, 2 to 101, 2 to 100, 2 to 99, 2 to 98, or 2 to 97 nucleotides. In some embodiments, the nick-edit distance is 2-90, 2-80, 2-70, 2-60, 2-50, 2-40, or 2-30 nucleotides. In some embodiments, the nick-edit distance is 2-25, 2-20, 2-15, or 2-10 nucleotides. In some embodiments, the nick-edit distance is 2, 3, 4, 5, 6, or 7 nucleotides long. In some embodiments, the nick-edit distance is 28 nucleotides. In some embodiments, the nick-edit distance is 22 nucleotides. In some embodiments, the nick-edit distance is 21 nucleotides. In some embodiments, the nick-edit distance is 17 nucleotides. In some embodiments, the nick-edit distance is 16 nucleotides. In some embodiments, the nick-edit distance is 4 nucleotides. In some embodiments, the nick-edit distance is 16 nucleotides. In some embodiments, the nick-edit distance is 1-19 nucleotides. In some embodiments, the nick-edit distance is 16 nucleotides. In some embodiments, the nick-edit distance is 1, 2, 7, 8, 13, 14, or 19 nucleotides. In some embodiments, the nick-edit distance is 8 nucleotides or less.In some embodiments, the nick-edit distance is 1 or 2 nucleotides.

[0316] The RTT length and the nick-edit distance relate to the length of the portion of the RTT that is upstream (i.e., 5' side) of the edit on the 5' side and complementary to the edited chain. In some embodiments, the edit template comprises at least four consecutive nucleotides complementary to the edited chain, and these at least four consecutive nucleotides are located upstream of the edit on the 5' side in the edit template. In some embodiments, the edit template comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or more consecutive nucleotides complementary to the edited chain, and these at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or more consecutive nucleotides are located upstream of the edit on the 5' side in the edit template. In some embodiments, the editing template includes 20-25, 25-30, 30-35, 35-40, 45-45, or 45-50 consecutive nucleotides complementary to the editing chain, and these 20-25, 25-30, 30-35, 35-40, 45-45, or 45-50 or more consecutive nucleotides are located upstream of the 5' edit in the editing template. In some embodiments, the editing template includes 9-14 consecutive nucleotides complementary to the editing chain, and these 9-14 consecutive nucleotides are located upstream of the 5' edit in the editing template. In some embodiments, the editing template includes 6-10 consecutive nucleotides complementary to the editing chain, and these 6-10 consecutive nucleotides are located upstream of the 5' edit in the editing template. In some embodiments, the editing template includes 10 consecutive nucleotides complementary to the editing chain, the 10 consecutive nucleotides located upstream of the 5' edit in the editing template. In some embodiments, the editing template includes 9 consecutive nucleotides complementary to the editing chain, the 9 consecutive nucleotides located upstream of the 5' edit in the editing template.

[0317] When referenced within a PEgRNA, the location of one or more intended nucleotide edits may be referenced in relation to the components of the PEgRNA. For example, the intended nucleotide edit may be 5' or 3' relative to the PBS. In some embodiments, the PEgRNA includes a spacer, a gRNA core, an editing template, and the structure of the PBS, from 5' to 3'. In some embodiments, the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream of the 5' end nucleotide of the PBS.In some embodiments, the intended nucleotide editing is performed on the nucleotides at the 5' end of the PBS, ranging from 0-2 nucleotides, 0-4 nucleotides, 0-6 nucleotides, 0-8 nucleotides, 0-10 nucleotides, 2-4 nucleotides, 2-6 nucleotides, 2-8 nucleotides, 2-10 nucleotides, 2-12 nucleotides, 4-6 nucleotides, 4-8 nucleotides, 4-10 nucleotides, 4-12 nucleotides, 4-14 nucleotides, 6-8 nucleotides, 6-10 nucleotides, 6-12 nucleotides, 6-14 nucleotides, 6-16 nucleotides, 8-10 nucleotides, 8-12 nucleotides, 8-14 nucleotides, 8-16 nucleotides, 8-18 nucleotides, 10-12 nucleotides, 10-14 nucleotides, 10- It is 16 nucleotides, 10-18 nucleotides, 10-20 nucleotides, 12-14 nucleotides, 12-16 nucleotides, 12-18 nucleotides, 12-20 nucleotides, 12-22 nucleotides, 14-16 nucleotides, 14-18 nucleotides, 14-20 nucleotides, 14-22 nucleotides, 14-24 nucleotides, 16-18 nucleotides, 16-20 nucleotides, 16-22 nucleotides, 16-24 nucleotides, 16-26 nucleotides, 16-26 nucleotides, 18-20 nucleotides, 18-22 nucleotides, 18-24 nucleotides, 18-26 nucleotides, 18-28 nucleotides, 20-22 nucleotides, 20-24 nucleotides, 20-26 nucleotides, 20-28 nucleotides, or 20-30 nucleotides upstream.

[0318] The corresponding location of an intended nucleotide edit incorporated into a target gene may also be referenced based on the nicking location generated by the prime editor based on sequence homology and complementarity. For example, in some embodiments, the distance (i.e., the number of nucleotides) between the nucleotide edit incorporated into the target B2M gene and the nicking site (also referred to as the "nick-edit distance," where the number of nucleotides does not include the 5'-side nucleotide location on the second strand corresponding to the edit) may be determined by identifying, for example, the sequence complementarity between the spacer and the search target sequence, and between the edit template and the edit target sequence, based on the location of the nicking site and the location of the nucleotide(s) corresponding to the intended nucleotide edit(s). In certain embodiments, the location of the nucleotide edit can be any location downstream of the nick site on the edited chain (or PAM chain) generated by the prime editor, and the distance between the nick site and the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotide lengths. In some embodiments, the location of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream of the nick site on the edited strand. In some embodiments, the location of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides downstream of the nick site on the edited strand.In some embodiments, the location of the nucleotide edit is 0 base pairs from the nick site on the edited strand, i.e., the edit site is at the same location as the nick site. As used herein, the distance between the nick site and the nucleotide edit refers to, for example, the furthest 5' position of the nucleotide edit of the nick that creates a 3' free end on the edited strand (i.e., the “closest position” of the nucleotide edit to the nick site) when the nucleotide edit involves an insertion or deletion. Similarly, as used herein, the distance between the nick site and the PAM position edit refers to, for example, the furthest 5' position of the nucleotide edit and the furthest 5' position of the PAM sequence when the nucleotide edit involves the insertion, deletion, or substitution of two or more consecutive nucleotides.

[0319] In some embodiments, the editing template extends beyond nucleotide editing incorporated into the target B2M gene sequence. For example, in some embodiments, the editing template includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79 or 80 nucleotides.

[0320] In some embodiments, the editing template may include a second edit to the target sequence. This second edit may be designed to mutate or otherwise silence the PAM sequence so that the corresponding nucleic acid guide nuclease or CRISPR nuclease can no longer cleave the target sequence (such an edit is referred to as a "PAM silencing edit").

[0321] While we do not wish to be bound by any particular theory, PAM silencing editing may improve prime editing efficiency by preventing Cas, such as Cas9 nickase, from renicking the edited strand before the edit is incorporated into the target strand. In some embodiments, PAM silencing editing modifies the sequence of the transcript or protein sequence encoded by the B2M gene. In some embodiments, PAM silencing editing is synonymous editing that does not modify the amino acid sequence or mRNA sequence encoded by the B2M gene after the edit is incorporated. In some embodiments, PAM silencing editing is performed in the coding region of the B2M gene, e.g., at a location corresponding to an exon. In some embodiments, PAM silencing editing is performed in the non-coding region of the B2M gene, e.g., at a location corresponding to an intron. In some embodiments, editing in an intron of the B2M gene is performed not at a location corresponding to an intron-exon junction, and the editing does not affect the splicing of the transcript.

[0322] In some embodiments, the length of the edit template is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37 , 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides longer. In some embodiments, for example, the nick-edit distance is 8 nucleotides, and the edit template is 10-15, 10-20, 10-25, 10-30, 10-35, 10-40, 10-45, 10-50, 10-55, 10-60, 10-65, 10-70, 10-75, or 10-80 nucleotides longer. In some embodiments, the nick-edit distance is 22 nucleotides, and the edit template is 24-28, 24-30, 24-32, 24-34, 24-36, 24-37, 24-38, 24-40, 24-45, 24-50, 24-55, 24-60, 24-65, 24-70, 24-75, 24-80, 24-85, 24-90, 24-95, 24-100, 24-105, 24-100, 24-105, or 24-110 nucleotides long.

[0323] In some embodiments, the editing template contains adenine at the first nucleic acid base position (for example, in the case of PEgRNA following a 5'-spacer-gRNA core-RTT-PBS-3' orientation, the 5'-side nucleic acid base is the "first base"). In some embodiments, the editing template contains guanine at the first nucleic acid base position (for example, in the case of PEgRNA following a 5'-spacer-gRNA core-RTT-PBS-3' orientation, the 5'-side nucleic acid base is the "first base"). In some embodiments, the editing template contains uracil at the first nucleic acid base position (for example, in the case of PEgRNA following a 5'-spacer-gRNA core-RTT-PBS-3' orientation, the 5'-side nucleic acid base is the "first base"). In some embodiments, the editing template contains cytosine at the first nucleic acid base position (for example, in the case of PEgRNA following a 5'-spacer-gRNA core-RTT-PBS-3' orientation, the 5'-side nucleic acid base is the "first base"). In some embodiments, the editing template does not contain cytosine at the first nucleic acid base position (for example, in the case of PEgRNA following a 5'-spacer-gRNA core-RTT-PBS-3' orientation, the 5'-most nucleic acid base is the "first base").

[0324] A PEgRNA editing template may encode new single-stranded DNA (e.g., by reverse transcription) to replace a target editing sequence within a target gene. In some embodiments, the target editing sequence in the edited strand of the target gene is replaced by the newly synthesized strand, and the nucleotide edit(s) are incorporated into the region of the target gene. In some embodiments, the target gene is a B2M gene. In some embodiments, the PEgRNA editing template encodes newly synthesized single-stranded DNA containing mutations or nucleotide changes compared to the wild-type B2M gene sequence. In some embodiments, the newly synthesized DNA strand replaces the target editing sequence within the target B2M gene, and the target editing sequence (or an endogenous sequence complementary to the target editing sequence on the target strand of the B2M gene) contains the wild-type B2M gene.

[0325] In some embodiments, newly synthesized single-stranded DNA encoded by the editing target sequence replaces the editing target sequence and introduces a mutation into the editing target sequence of the B2M gene.

[0326] In some embodiments, the editing template includes one or more intended nucleotide edits compared to the sequence of the target strand of the B2M gene complementary to the editing target sequence. In some embodiments, the editing template encodes single-stranded DNA containing one or more intended nucleotide edits compared to the editing target sequence. In some embodiments, the single-stranded DNA replaces the editing target sequence by prime editing, thereby incorporating the one or more intended nucleotide edits. In some embodiments, the incorporation of the one or more intended nucleotide edits introduces mutations into the editing target sequence compared to wild-type nucleotides at the corresponding positions in the B2M gene.

[0327] In some embodiments, the editing target sequence includes a mutation located at positions 44,711,517–44,718,145 of human chromosome 15 by GRCh38.

[0328] For example, in some embodiments, the incorporation of one or more intended nucleotide edits results in one or more codons that are different from wild-type codons. In some embodiments, the incorporation of one or more intended nucleotide edits results in one or more codons that encode one or more amino acids that are different from wild-type B2M protein. In some embodiments, the incorporation of one or more intended nucleotide edits results in one or more in-frame immature stop codons, resulting in a cleaved polypeptide compared to wild-type B2M protein. "B2M protein," "β2 microglobulin protein," or "β-chain of MHC class I" means a protein having at least about 85% amino acid sequence identity to NCBI accession number P61769.1 or a fragment thereof and possessing immunomodulatory activity. An exemplary amino acid sequence of wild-type B2M protein is shown in SEQ ID NO: 742 (NCBI accession number P61769.1). "B2M gene" means the nucleic acid that encodes the B2M protein. Exemplary B2M nucleic acid sequences are publicly available, for example, in the UCSC Human Genome Database, Gene ENSG00000166710.23, and these sequences are incorporated herein by reference in their entirety. An exemplary mRNA / cDNA sequence of the wild-type B2M protein is shown in SEQ ID NO: 743.

[0329] Wild-type B2M protein sequence (SEQ ID NO: 742) MSRSVALAVLALLSLSGLEAIQRTPKIQVYSRHPAENGKSNFLNCYVSGFHPSDIEVDLLKNGERIEKVEHSDLSFSKDWSFYLLYYTEFTPTEKDEYACRVVTNHLSQPKIVKWDRDM

[0330] mRNA / cDNA sequence of wild-type B2M (SEQ ID NO: 743) ATTCCTGAAGCTGACAGCATTCGGGCCGAGATGTCTCGCTCCGTGGCCTTAGCTGTGCTCGCGCTACTCTCTCTCTTTCTTTCTGGCCTGGAGGCTATCCAGCGTACTCCAAAGATTCAGGTTTTACTCACGTCATCCAGCAGAGAATGGAAAGTCAAATTTCCTGAATTGCTATGTGTCTGGGTTTTCCATCCGACATTGAAGTTGACTTACTGAAGAATGGAGAGAGAATTGAAAAAGTGGAGCATTCAGACTTGTCTTTCAGCAAGGACTGGTCTTTCTATCTCTTGTACTACACTGAATTCACCCCACTGAAAAAGATGAGTATGCCTGCCGTGTGAACCATGTGACTTTGTCACAGCCCAAGATAGTTAAGTGGGATCGAGACATGTAAGCAGCATCATGGAGGTTTGAAGATGCCGCATTTGGATTGGATGAATTCCAAATTCTGCTTGCTTGCTTTTTAATATTGATA TGCTTATACACTTACACTTTATGCACAAAATGTAGGGTTATAATAATGTTAACATGGACATGATCTTCTTTATAATTCTACTTTGAGTGCTGTCTCCATGTTTGATGTATCTGAGCAGGTTGCTCCACAGGTAGCTCTAGGAGGGCTGGCAACTTAGAGGTGGGGAGCAGAGAATTCTCTTATCCAACATCAACATCTTGGTCAGATTTGAACTCTTCAATCTCTTGCACTCAAAGCTTGTTAAGATAGTTAAGCGTGCATAAGTTAACTTCCAATTTACATACTCTGCTTAGAATTTGGGGGAAAATTTAGAAATATAATTGACAGGATTATTGGAAATTTGTTATAATGAATGAAACATTTTGTCATATAAGATTCATATTTACTTCTTATACATTTGATAAAGTAAGGCATGGTTGTGGTTAATCTGGTTTTATTTTTGTTCCACAAGTTAAATAAATCATAAAACTTGA

[0331] The guide RNA core of the PEgRNA (also referred herein as the gRNA core, gRNA scaffold, or gRNA backbone sequence) may include a polynucleotide sequence that binds to the DNA-binding domain of the prime editor (e.g., Cas9). The gRNA core may interact with the prime editor described herein by association with the DNA-binding domain of the prime editor, for example, a DNA nickase.

[0332] Those skilled in the art will recognize that different prime editors having different DNA-binding domains from different DNA-binding proteins may require different gRNA core sequences specific to the DNA-binding protein. In some embodiments, the gRNA core can bind to a Cas9-based prime editor. In some embodiments, the gRNA core can bind to a Cpf1-based prime editor. In some embodiments, the gRNA core can bind to a Cas12b-based prime editor.

[0333] In some embodiments, the gRNA core includes regions and secondary structures involved in binding to specific CRISPR-Cas proteins. For example, in a Cas9-based prime editing system, the gRNA core of PEgRNA may include one or more regions of a base-paired “lower stem” adjacent to a spacer sequence and a base-paired “upper stem” following the lower stem, where the lower and upper stems may be connected by a “bulge” containing unpaired RNA. The gRNA core may further include a “nexus” distal to the spacer sequence and a subsequent hairpin structure, for example, at the 3' end, as illustrated in Figure 3. In some embodiments, the gRNA core contains modified nucleotides in the lower stem, upper stem, and / or hairpin compared to a wild-type gRNA core. For example, nucleotides in the lower stem, upper stem, and / or hairpin regions may be modified, deleted, or substituted. In some embodiments, RNA nucleotides in the lower stem, upper stem, and / or hairpin regions may be substituted with one or more DNA sequences. In some embodiments, the gRNA core includes unmodified or wild-type RNA sequences in the nexus region and / or bulge region. In some embodiments, the gRNA core does not include long stretches of AT pairs, such as the GUUUU-AAAAC paired element. In some embodiments, the prime editing system includes a prime editor and a PEgRNA, the prime editor includes its SpCas9 nickerse variant, and the gRNA core of the PEgRNA includes the sequence:GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 646), GUUUGAGAGCUAGAAAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGGACCGAGUCGGUCC (SEQ ID NO: 648), or GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 649).In some embodiments, the gRNA core comprises the sequence GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 646). Any gRNA core sequence known in the art is also intended in the prime editing compositions described herein.

[0334] In some embodiments, the PEgRNA and / or ngRNA comprises a gRNA core containing a nucleic acid sequence selected from Table 28 below. In some embodiments, the PEgRNA and / or ngRNA comprises a gRNA core containing a nucleic acid sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 653, 646, 652, 647, 649, 654, or 648. In some embodiments, the PEgRNA and / or ngRNA comprises a gRNA core containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 653, 646, 652, 647, 649, 654, or 648. [Table 8]

[0335] In some embodiments, the PEgRNA includes a linker. In some embodiments, a secondary structure or 3' motif is linked to one or more other components of the PEgRNA via a linker. For example, in some embodiments, the secondary structure is located at the 3' end of the PEgRNA (e.g., RTT, or PBS) and is linked to the 3' end of the PBS via a linker. For example, in some embodiments, the 3' motif is located at the 3' end of the PEgRNA and is linked to the 3' end of the PEgRNA (e.g., RTT, or PBS) via a linker. In some embodiments, the secondary structure or 5' motif is located at the 5' end of the PEgRNA and is linked to the 5' end of a spacer via a linker. In some embodiments, the linker is a nucleotide linker with a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides. In some embodiments, the linker is 5 to 10 nucleotides long. In some embodiments, the linker is 10 to 20 nucleotides long. In some embodiments, the linker is 15 to 25 nucleotides long. In some embodiments, the linker is 8 nucleotides long.

[0336] In some embodiments, the linker is designed to minimize base pairing between the linker and another component of PEgRNA. In some embodiments, the linker is designed to minimize base pairing between the linker and a spacer. In some embodiments, the linker is designed to minimize base pairing between the linker and PBS. In some embodiments, the linker is designed to minimize base pairing between the linker and the editing template. In some embodiments, the linker is designed to minimize base pairing between the linker and the sequence of the RNA secondary structure. In some embodiments, the linker is optimized to minimize base pairing between the linker and another component of PEgRNA, in the order of priority of spacer, PBS, editing template, and then scaffold. In some embodiments, the probability of base pairing is calculated using ViennaRNA2.0 as described in Lorenz, R. et al. ViennaRNA package 2.0. Algorithms Mol. Biol. 6 (which is incorporated herein by reference as a whole) under standard parameters (37°C, 1M NaCl, 0.05M MgCl2).

[0337] PEgRNA may also include optional modifiers, e.g., a 3'-terminal modifier region and / or a 5'-terminal modifier region. In some embodiments, PEgRNA includes at least one nucleotide that is not part of a spacer, gRNA core, or elongation arm. Optional sequence modifiers may be located in or between any of the other regions shown, and are not limited to being located at the 3' and 5' ends. In certain embodiments, PEgRNA includes, but is not limited to, secondary RNA structures such as aptamers, hairpins, stem / loops, toe loops, and / or RNA-binding protein recruitment domains (e.g., an MS2 aptamer that recruits and binds to the MS2cp protein). In some embodiments, PEgRNA includes a short stretch of uracil at the 5' or 3' end. For example, in some embodiments, PEgRNA including a 3' elongation arm includes a "UUU" sequence at the 3' end of the elongation arm. In some embodiments, PEgRNA includes a toe loop sequence at the 3' end. In some embodiments, the PEgRNA includes a 3' elongation arm and a troop sequence at the 3' end of the elongation arm. In some embodiments, the PEgRNA includes a 5' elongation arm and a troop sequence at the 5' end of the elongation arm. In some embodiments, the PEgRNA includes a troop element having the sequence 5'-GAAANNNNN-3', where N is any nucleic acid base. In some embodiments, the secondary RNA structure is located within a spacer. In some embodiments, the secondary structure is located within an elongation arm. In some embodiments, the secondary structure is located within a gRNA core. In some embodiments, the secondary structure is located between a spacer and a gRNA core, between a gRNA core and an elongation arm, or between a spacer and an elongation arm. In some embodiments, the secondary structure is located between PBS and an editing template. In some embodiments, the secondary structure is located at the 3' or 5' end of the PEgRNA. In some embodiments, the PEgRNA includes a transcription termination signal at its 3' end.In addition to the secondary RNA structure, PEgRNA includes a chemical linker or a poly(N) linker or tail, where "N" is any nucleic acid base. In some embodiments, the chemical linker can act to prevent reverse transcription of the gRNA core.

[0338] In some embodiments, the prime editing system or composition further comprises a nick guide polynucleotide, such as a nick guide RNA (ngRNA). In some embodiments, the ngRNA comprises a spacer (referred to as the ngRNA spacer or ng spacer) and a gRNA core, wherein the spacer of the ngRNA comprises a region complementary to the edited strand, and the gRNA core may interact with the prime editor's Cas, e.g., Cas9. While we do not wish to be bound to any particular theory, the ngRNA may instruct the Cas nickase to bind to the edited strand and generate a nick on the unedited strand (or target strand). In some embodiments, the nick on the unedited strand can instruct the endogenous DNA repair mechanism to use the edited strand as a template for repairing the unedited strand, thereby increasing the efficiency of prime editing. In some embodiments, the unedited strand is nicked by a prime editor localized to the unedited strand by the ngRNA. Thus, also provided herein are PEgRNA systems comprising at least one PEgRNA and at least one ngRNA.

[0339] A prime editing system comprising PEgRNA (or one or more polynucleotides encoding PEgRNA) and a prime editor protein (or one or more polynucleotides encoding the prime editor) may be referred to as a PE2 prime editing system, and the corresponding editing approach may be referred to as a PE2 approach or PE2 strategy. The PE2 system does not contain ngRNA. A prime editing system comprising PEgRNA (or one or more polynucleotides encoding PEgRNA), a prime editor protein (or one or more polynucleotides encoding the prime editor), and ngRNA (or one or more polynucleotides encoding ngRNA) may be referred to as a "PE3" prime editing system. In some embodiments, the ngRNA spacer sequence is complementary to the portion of the edited strand containing the intended nucleotide edit and can only hybridize with the edited strand after the edit has been incorporated into the edited strand. Such ngRNA may be referred to as a "PE3b" ngRNA, and the prime editing system may be referred to as a PE3b prime editing system.

[0340] In some embodiments, PEgRNA or nick guide RNA (ngRNA) may be chemically synthesized or assembled or cloned and transcribed from a DNA sequence, e.g., a plasmid DNA sequence, or by any RNA oligonucleotide synthesis method known in the art. In some embodiments, the DNA sequence encoding PEgRNA (or ngRNA) may be designed to have one or more nucleotides added to the 5' or 3' end of the PEgRNA (or nick guide RNA) encoding sequence to facilitate the transcription of PEgRNA. For example, in some embodiments, the DNA sequence encoding PEgRNA (or nick guide RNA) (or ngRNA) may be designed to have nucleotide G added to the 5' end. Thus, in some embodiments, the PEgRNA (or nick guide RNA) may contain nucleotide G added to the 5' end. In some embodiments, the DNA sequence encoding PEgRNA (or nick guide RNA) may be designed to have a transcription-promoting sequence, e.g., a Kozak sequence, added to the 5' end. In some embodiments, the DNA sequence encoding PEgRNA (or nick guide RNA) may be designed to have the sequence CACC or CCACC added to the 5' end. Therefore, in some embodiments, the PEgRNA (or nick guide RNA) may include the sequence CACC or CCACC appended to its 5' end. In some embodiments, the DNA sequence encoding the PEgRNA (or nick guide RNA) may be designed to have the sequence TTT, TTTT, TTTTT, TTTTTT, TTTTTTT appended to its 3' end. Therefore, in some embodiments, the PEgRNA (or nick guide RNA) may include the sequence UUU, UUUU, UUUUU, UUUUUU, or UUUUUUUU. In some embodiments, the PEgRNA or ngRNA may include a modified sequence having the sequence AACAUUGACGCGUCUCUACGUGGGGGCGCG (SEQ ID NO: 745) at its 3' end. In some embodiments, the PEgRNA or ngRNA includes the sequence TTTT at its 3' end.In some embodiments, PEgRNA or ngRNA contains the sequence TTTTTTT at its 3' end. In some embodiments, PEgRNA or ngRNA contains a 3' terminator sequence (e.g., TTTT) at its 3' end. In some embodiments, PEgRNA or ngRNA contains a transcriptional adaptation sequence (e.g., TTTTTTT) at its 3' end. Although the sequences TTTT and TTTTTTT are RNA sequences, "T" is used instead of "U" in these sequences to maintain consistency with the ST.26 standard.

[0341] In some embodiments, the ng search target sequence is located on the non-target strand within 10 to 100 base pairs of the intended nucleotide edit incorporated by the PEgRNA on the edited strand. In some embodiments, the ng target search target sequence is within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp of the intended nucleotide edit incorporated by the PEgRNA on the edited strand. In some embodiments, the 5' ends of the ng search target sequence and the 5' ends of the PEgRNA search target sequence are within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp of each other. In some embodiments, the 5' ends of the ng target sequence and the 5' ends of the PEgRNA target sequence are separated from each other by 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp.

[0342] In some embodiments, the ng spacer sequence is complementary to and can hybridize with the second target sequence only after the intended nucleotide edit has been incorporated into the edited strand by the PEgRNA editing template. In some embodiments, such a prime editing system may be referred to as a “PE3b” prime editing system or configuration. In some embodiments, the ngRNA includes a spacer sequence that matches only the edited strand after the nucleotide edit has been incorporated and does not match the endogenous target gene sequence on the edited strand. Thus, in some embodiments, the intended nucleotide edit is incorporated into the ng target sequence.

[0343] Tables 1-21 show exemplary combinations of PEgRNA components, such as spacers, PBS, and editing templates / RTTs, exemplary full-length PEgRNAs, and combinations of PEgRNAs with their corresponding ngRNAs(s). Each of Tables 1-21 contains three columns. The...

Claims

1. Prime editing guide RNA (PEGRNA) or one or more polynucleotides encoding the PEGRNA, wherein the PEGRNA is a. A spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, wherein the spacer contains sequence number 205 at its 3' end. b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template including a region complementary to the editing target sequence on the second strand of the B2M gene, and ii. The extension arm includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 205, The first chain and the second chain are complementary to each other. The editing template comprises the PEGRNA or one or more polynucleotides encoding the PEGRNA, which encode one or more nucleotide changes compared to the editing target sequence.

2. Prime editing guide RNA (PEGRNA) or one or more polynucleotides encoding the PEGRNA, wherein the PEGRNA is a. A spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, wherein the spacer contains sequence number 4 at its 3' end. b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template including a region complementary to the editing target sequence on the second strand of the B2M gene, and ii. The extension arm includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 4, The first chain and the second chain are complementary to each other. The editing template comprises the PEGRNA or one or more polynucleotides encoding the PEGRNA, which encode one or more nucleotide changes compared to the editing target sequence.

3. Prime editing guide RNA (PEGRNA) or one or more polynucleotides encoding the PEGRNA, wherein the PEGRNA is a. A spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, wherein the spacer contains sequence number 272 at its 3' end. b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template including a region complementary to the editing target sequence on the second strand of the B2M gene, and ii. The extension arm includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 272, The first chain and the second chain are complementary to each other. The editing template comprises the PEGRNA or one or more polynucleotides encoding the PEGRNA, which encode one or more nucleotide changes compared to the editing target sequence.

4. Prime editing guide RNA (PEGRNA) or one or more polynucleotides encoding the PEGRNA, wherein the PEGRNA is a. A spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, wherein the spacer contains sequence number 330 at its 3' end. b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template including a region complementary to the editing target sequence on the second strand of the B2M gene, and ii. The extension arm includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 330, The first chain and the second chain are complementary to each other. The editing template comprises the PEGRNA or one or more polynucleotides encoding the PEGRNA, which encode one or more nucleotide changes compared to the editing target sequence.

5. The PEGRNA according to any one of claims 1 to 4, wherein the spacer is 17 to 22 nucleotides long, and optionally, the spacer is 20 nucleotides long.

6. The PEGRNA according to claim 1, wherein the spacer includes one of sequence numbers 202 to 204 at its 3' end.

7. The PEGRNA according to claim 1, wherein the spacer includes sequence number 204.

8. The PEGRNA according to claim 2, wherein the spacer contains one of sequence numbers 1 to 3 at its 3' end.

9. The PEGRNA according to claim 2, wherein the spacer includes sequence number 1.

10. The PEGRNA according to claim 3, wherein the spacer includes one of sequence numbers 269 to 271 at its 3' end.

11. The PEGRNA according to claim 3, wherein the spacer includes sequence number 269.

12. The PEGRNA according to claim 4, wherein the spacer includes one of sequence numbers 327 to 329 at its 3' end.

13. The PEGRNA according to claim 4, wherein the spacer includes sequence number 327.

14. The PEGRNA according to any one of claims 1 to 13, wherein one or more nucleotide changes encoded by the editing template include non-synonymous editing that alters the mRNA sequence or protein sequence encoded by the B2M gene.

15. The PEGRNA according to claim 14, wherein the non-synonymous editing results in one or more in-frame stop codons being introduced into the B2M gene.

16. The PEGRNA according to claim 15, wherein the one or more in-frame stop codons include a nonsense mutation in the B2M gene.

17. The PEGRNA according to claim 14, wherein the non-synonymous editing includes an insertion in the B2M gene.

18. The PEGRNA according to claim 14, wherein the non-synonymous editing comprises one or more substitutions in the B2M gene.

19. The PEGRNA according to claim 17, wherein the insertion comprises the insertion of an in-frame stop codon in the B2M gene, and optionally, the insertion comprises the insertion of two or more consecutive in-frame stop codons in the B2M gene.

20. The PEGRNA according to claim 19, wherein the insertion includes the insertion of a TAATAA, TTATTA, or TAATAG nucleotide.

21. The PEGRNA according to claim 14, wherein the non-synonymous editing includes a frameshift mutation in the B2M gene.

22. The PEGRNA according to claim 21, wherein the frameshift mutation is an insertion of 3x+1 or 3x+2 nucleotides, in which case x is an integer greater than or equal to 0.

23. The PEGRNA according to claim 21, wherein the frameshift mutation is a deletion of 3x+1 or 3x+2 nucleotides, in which case x is an integer of 0 or more.

24. The PEGRNA according to claim 22, wherein the insertion is 1, 2, or 4 nucleotides long.

25. The PEGRNA according to claim 23, wherein the deletion is 1 nucleotide in length.

26. The PEGRNA according to any one of claims 14 to 25, wherein the non-synonymous editing modifies a protospacer adjacent motif (PAM) sequence immediately 3' of a protospacer sequence in the second strand of the B2M gene that is complementary to the search target sequence in the first strand of the B2M gene.

27. The PEGRNA according to claim 26, wherein the PAM sequence is NGG, and the non-synonymous editing is NGG->NGC editing.

28. The PEGRNA according to claim 26 or 27, wherein the protospacer sequence includes a nick site three nucleotides upstream of the furthest 5' nucleotide of the PAM sequence, the number of nucleotides from the nick site to the position on the second strand of the B2M gene corresponding to the non-synonymous editing is 1 to 19 nucleotides, and the number of nucleotides does not include the furthest 5' nucleotide position on the second strand corresponding to the non-synonymous editing.

29. The PEGRNA according to claim 28, wherein the number of nucleotides from the nick site to the position on the second strand of the B2M gene corresponding to the non-synonymous editing is 1, 2, 7, 8, 13, 14, or 19 nucleotides.

30. The PEGRNA according to claim 29, wherein the number of nucleotides from the nick site to the position on the second strand of the B2M gene corresponding to the non-synonymous editing is 8 nucleotides or less.

31. The PEGRNA according to claim 30, wherein the number of nucleotides from the nick site to the position on the second strand of the B2M gene corresponding to the non-synonymous editing is 1 or 2 nucleotides.

32. The PEGRNA according to any one of claims 1, 5-7, and 14-31, wherein the non-synonymous editing is at a chromosomal position corresponding to coding sequence positions c. 51, c. 54, or c. 50 of the wild-type B2M gene.

33. The PEGRNA according to claim 32, wherein the non-synonymous editing includes the insertion of c. 54insTAATAA.

34. The PEGRNA according to claim 32, wherein the non-synonymous editing includes a deletion of c. 51delC or an insertion of 50insG.

35. The PEGRNA according to any one of claims 2, 5, 8-9, and 14-31, wherein the non-synonymous editing is at a chromosomal position corresponding to coding sequence positions c. 54, c. 60, or c. 66 of the wild-type B2M gene.

36. The PEGRNA according to claim 35, wherein the non-synonymous editing includes the insertion of c. 54_55insCC or c. 54_55insTAAG.

37. The PEGRNA according to claim 35, wherein the non-synonymous editing includes the insertion of c. 54_55insTAATAA.

38. The PEGRNA according to claim 35, wherein the non-synonymous editing includes the insertion of c.66_67insCC or c.66_67insTAAG.

39. The PEGRNA according to claim 32, wherein the non-synonymous editing includes the insertion of c. 66_67insTAATAA.

40. The PEGRNA according to claim 35, wherein the non-synonymous editing includes the deletion of c. 60_65 and the insertion of TAATAG (c. 60_65_delinsTAATAG).

41. The PEGRNA according to claims 3, 5, 10-11, and 14-31, wherein the non-synonymous editing is at a chromosomal position corresponding to coding sequence position c.21 or c.3 of the wild-type B2M gene.

42. The PEGRNA according to claim 41, wherein the non-synonymous editing includes the insertion of c. 21insTAATAA.

43. The PEGRNA according to claim 41, wherein the non-synonymous editing includes the insertion of c.21_22insCC or editing of c.21_22insTAAG.

44. The PEGRNA according to claim 41, wherein the non-synonymous editing includes the insertion of c.3_4insCC or c.3_4insTAAG.

45. The PEGRNA according to claim 41, wherein the non-synonymous editing includes the deletion of c.3_8 and the insertion of TAATGA (c.3_8delinsTAATGA).

46. The PEGRNA according to any one of claims 4-5 and 12-31, wherein the non-synonymous editing is at a chromosomal position corresponding to coding sequence positions c.21, c.15, or c.3 of the wild-type B2M gene.

47. The PEGRNA according to claim 46, wherein the non-synonymous editing includes the insertion of c. 21insTAATAA.

48. The PEGRNA according to claim 46, wherein the non-synonymous editing includes the insertion of c.15_16insCC or c.15_16insTAAG.

49. The PEGRNA according to claim 46, wherein the non-synonymous editing includes the insertion of c. 15_16insTAATAA.

50. The PEGRNA according to claim 46, wherein the non-synonymous editing includes the insertion of c.3_4insCC or c.3_4insTAAG.

51. The PEGRNA according to claim 46, wherein the non-synonymous editing includes the insertion of c. 3_4insTAATAA.

52. The PEGRNA according to claim 46, wherein the non-synonymous editing includes the deletion of c.3_8 and the insertion of TAATGA (c.3_8delinsTAATGA).

53. The PEGRNA according to any one of claims 1 to 52, wherein the editing template further encodes further PAM silencing editing.

54. The PEGRNA according to claim 53, wherein the PAM silencing edit is c. 58G > C edit.

55. The PEGRNA according to claim 53, wherein the PAM silencing edit is c. 17C>G edit.

56. The PEGRNA according to claim 53, wherein the PAM silencing edit is c. 11C>G edit.

57. The PEGRNA according to any one of claims 1 to 56, wherein the editing template comprises at least four consecutive nucleotides complementary to the editing target sequence, and the at least four consecutive nucleotides are upstream of the position of the furthest 5' nucleotide of the one or more nucleotide changes encoded in the editing template.

58. The PEGRNA according to claim 57, wherein the editing template comprises at least 6, 8, or 10 consecutive nucleotides complementary to the editing target sequence, and the at least 6, 8, or 10 consecutive nucleotides are upstream of the position of the furthest 5' nucleotide of the one or more nucleotide changes encoded in the editing template.

59. The PEGRNA according to claim 57, wherein the editing template comprises 4, 6, 8, or 10 consecutive nucleotides complementary to the editing target sequence, and the 4, 6, 8, or 10 consecutive nucleotides are upstream of the position of the furthest 5' nucleotide of the one or more nucleotide changes encoded in the editing template.

60. Prime editing guide RNA (PEGRNA), or a nucleic acid encoding the PEGRNA, a. Spacer containing sequence number 205 at the 3' end, b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template having at its 3' end (A) nucleotides 13-24 of SEQ ID NO: 221, (B) nucleotides 12-20 of SEQ ID NO: 227, or (C) nucleotides 7-17 of SEQ ID NO: 231, and ii. The PEGRNA, or nucleic acid encoding the PEGRNA, comprising the elongated arm, which includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 205.

61. (i) The editing template contains nucleotides 13-24 of SEQ ID NO: 221 at its 3' end, and optionally the editing template contains SEQ ID NOs: 219, 220, or (ii) The editing template contains nucleotides 12 to 20 of sequence number 227 at its 3' end, and optionally the editing template contains one of sequence numbers 224 to 227 at its 3' end, or (ii) The PEGRNA according to any one of claims 1, 5 to 7, 14 to 34, or 53 to 60, wherein the editing template comprises nucleotides 7 to 17 of SEQ ID NO: 231 at its 3' end, and optionally, the editing template comprises one of SEQ ID NOs: 229 to 231 at its 3' end.

62. Prime editing guide RNA (PEGRNA), or a nucleic acid encoding the PEGRNA, a. Spacer containing sequence number 1 at the 3' end, b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template having at its 3' end a sequence selected from the group consisting of (A) nucleotides 5-16 of SEQ ID NO: 19, or (B) SEQ ID NOs: 900, 904, 908, 912, 916, 920, and 924. ii. The PEGRNA, or nucleic acid encoding the PEGRNA, comprising the extension arm, which includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 1.

63. The aforementioned editing template, (i) A sequence selected from the group consisting of sequence numbers 900 to 903, or (ii) A sequence selected from the group consisting of sequence numbers 904 to 907, or (iii) A sequence selected from the group consisting of sequence numbers 908 to 911, or (iv) A sequence selected from the group consisting of sequence numbers 912 to 915, or (v) A sequence selected from the group consisting of sequence numbers 916-919, 928, and 929, or (vi) A sequence selected from the group consisting of sequence numbers 920 to 923, or (vii) A sequence selected from the group consisting of sequence numbers 924 to 927, or (viiii) PEGRNA according to any one of claims 2, 5, 8-9, 14-31, 35-40, 53-59, and 62, comprising a sequence selected from the group consisting of sequence numbers 18-20.

64. Prime editing guide RNA (PEGRNA), or a nucleic acid encoding the PEGRNA, a. Spacer containing sequence number 269 at the 3' end, b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template having at its 3' end a sequence selected from the group consisting of (A) nucleotides 3-16 of SEQ ID NO: 286, or (B) SEQ ID NOs: 1033, 1037, 1041, 1045, 1049, 1053, and 1057. and ii. The PEGRNA, or nucleic acid encoding the PEGRNA, comprising the elongated arm, which includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 269.

65. The aforementioned editing template, (i) A sequence selected from the group consisting of sequence numbers 1033 to 1036, or (ii) A sequence selected from the group consisting of sequence numbers 1037 to 1040, or (iii) A sequence selected from the group consisting of sequence numbers 1041 to 1044, or (iv) A sequence selected from the group consisting of sequence numbers 1045 to 1048, or (v) A sequence selected from the group consisting of sequence numbers 1049-1052 and 1061-1063, or (vi) A sequence selected from the group consisting of sequence numbers 1053 to 1056, or (vi) A sequence selected from the group consisting of sequence numbers 1057 to 1060, or (vii) PEGRNA according to any one of claims 3, 5, 10-11, 14-31, 41-45, 53-59, and 64, comprising a sequence selected from the group consisting of sequence numbers 286-288.

66. Prime editing guide RNA (PEGRNA), or a nucleic acid encoding the PEGRNA, a. Spacer containing sequence number 327 at the 3' end, b. A gRNA core capable of binding to the Cas9 protein, and c. An extendable arm, i. An editing template having at its 3' end (A) nucleotides 6-16 of SEQ ID NO: 344, or (B) a sequence selected from the group consisting of SEQ ID NOs: 1162, 1166, 1170, 1174, 1178, 1182, and 1190, and ii. The PEGRNA, or nucleic acid encoding the PEGRNA, comprising the extension arm, which includes a primer-binding site (PBS) at its 5' end containing the reverse complementary sequence of nucleotides 10-14 of sequence number 327.

67. The aforementioned editing template, (i) A sequence selected from the group consisting of sequence numbers 1162 to 1165, or (ii) A sequence selected from the group consisting of sequence numbers 1166 to 1169, or (iii) A sequence selected from the group consisting of sequence numbers 1170 to 1173, or (iv) A sequence selected from the group consisting of sequence numbers 1174 to 1177, or (v) A sequence selected from the group consisting of sequence numbers 1178-1181 and 1191, or (vi) A sequence selected from the group consisting of sequence numbers 1182 to 1185, or (vi) A sequence selected from the group consisting of sequence numbers 1186 to 1190, or (vii) PEGRNA according to any one of claims 4-5, 12-31, 46-59, and 66, comprising a sequence selected from the group consisting of sequence numbers 344-346.

68. The PEGRNA according to any one of claims 1 to 67, wherein the editing template has a length of 24 nucleotides or less, or a length of 20 nucleotides or less.

69. The PEGRNA according to claim 68, wherein the editing template has (i) a length of 10 to 20 nucleotides, (ii) a length of 12 to 20 nucleotides, or (iii) a length of 11 to 17 nucleotides.

70. The PEGRNA according to claim 68, wherein the editing template is 16 to 24 nucleotides long.

71. PEGRNA according to any one of claims 1 to 67, wherein the editing template has a length of 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 24, 25, 26, 27, 28, 29, 30, 31, 33, or 35 nucleotides.

72. The PEGRNA according to any one of claims 1 to 71, wherein the PBS has a length of 17 nucleotides or less.

73. The PEGRNA according to claim 72, wherein the PBS has a length of (i) 8 to 15 nucleotides, (ii) 8 to 14 nucleotides, or (iii) 8 to 12 nucleotides.

74. The PEGRNA according to claim 30, wherein the PBS is 8, 10, or 12 nucleotides long.

75. The PEGRNA according to any one of claims 1, 5 to 7, 14 to 34, 53 to 61, and 68 to 74, wherein the PBS comprises the sequence described in any one of sequence numbers 206 to 218.

76. The PEGRNA according to any one of claims 2, 5, 8-9, 14-31, 35-40, 53-59, 62-63, and 68-74, wherein the PBS comprises the sequence described in any one of sequence numbers 5-17.

77. The PEGRNA according to any one of claims 3, 5, 10-11, 14-31, 41-45, 53-59, 64-65, and 68-74, wherein the PBS comprises the sequence described in any one of sequence numbers 273-285.

78. The PEGRNA according to any one of claims 4-5, 12-31, 46-59, and 66-74, wherein the PBS comprises the sequence described in any one of sequence numbers 331-343.

79. The PEGRNA according to any one of claims 1 to 78, wherein the spacer, the gRNA core, the RTT, and the PBS form a continuous sequence in a single molecule.

80. The PEGRNA according to claim 79, comprising the spacer, the gRNA core, the RTT, and the PBS from 5' to 3'.

81. The PEGRNA according to any one of claims 1 to 80, wherein the gRNA core includes sequence number 646.

82. The PEGRNA according to any one of claims 1 to 80, wherein the gRNA core includes sequence number 653.

83. PEGRNA according to any one of claims 1, 5-7, 14-34, 53-61, 68-75, and 79-82, comprising a sequence selected from the group consisting of SEQ ID NOs. 232-262.

84. PEGRNA according to any one of claims 2, 5, 8-9, 14-31, 35-40, 53-59, 62-63, 68-74, 76, and 79-82, comprising a sequence selected from the group consisting of SEQ ID NOs: 21-29 and 930-1016.

85. The PEGRNA according to claim 84, comprising the sequence described in SEQ ID NOs: 933, 937, 961, 941, 957, or 936.

86. PEGRNA according to any one of claims 3, 5, 10-11, 14-31, 41-45, 53-59, 64-65, 68-74, 77, and 79-82, comprising a sequence selected from the group consisting of SEQ ID NOs: 289-297 and 1064-1151.

87. The PEGRNA according to claim 86, comprising the sequence described in sequence number 1141 or 1143.

88. PEGRNA according to any one of claims 4-5, 12-31, 46-59, and 66-74, 78-82, comprising a sequence selected from the group consisting of SEQ ID NOs: 347-355 and 1192-1279.

89. The PEGRNA according to claim 88, comprising the sequence described in SEQ ID NO: 1269 or 1265.f.

90. PEGRNA according to any one of claims 1 to 74, comprising a sequence selected from the group consisting of SEQ ID NOs: 957, 961, 965, 980, 1016, 956, 933, 941, 937, 1223, 988, 984, 1225, 1151, 1095, 1091, 964, 960, 940, 1221, 945, 1219, 932, 1015, 1014, 1075, 1222, 1250, 936, 1013, 1119, 1226, and 949.

91. Furthermore, the PEGRNA according to any one of claims 1 to 90, wherein the 3' motif is optionally connected to the 3' end of the PBS via a linker.

92. Furthermore, the PEGRNA according to any one of the prior claims, comprising the modifications 3'mN*mN*mN*N and / or 5'mN*mN*mN*, where m indicates that the nucleotide comprises the modification 2'-O-Me, and * indicates the presence of a phosphorothioate bond.

93. Furthermore, the PEGRNA according to any one of the prior claims, comprising the modifications 3'mT*mT*mT*T and 5'mN*mN*mN*, where m indicates that the nucleotide comprises the modification 2'-O-Me, * indicates the presence of a phosphorothioate bond, and T indicates the presence of further uridine nucleotides.

94. PEGRNA according to any one of the prior claims, wherein the human chromosome location and coding sequence location are as described in Genome Reference Consortium Human Build 38 (GrCh38).

95. A prime editing system comprising a PEGRNA according to any one of the prior claims or one or more polynucleotides encoding the PEGRNA.

96. Furthermore, it includes a nick guide RNA (ngRNA), or a nucleic acid encoding the ngRNA, wherein the ngRNA is a. An ngRNA spacer complementary to the ngRNA search target sequence on the second strand of the B2M gene, and b. The prime editing system according to claim 95, comprising an ngRNA core capable of binding to the Cas9 protein.

97. The prime editing system according to claim 96, wherein the ngRNA spacer is 17 to 22 nucleotides long, and optionally, the ngRNA spacer is 20 nucleotides long.

98. The prime editing system according to claim 96 or 97, wherein the ngRNA core includes sequence number 646 or 653.

99. The prime editing system according to any one of claims 96 to 98, wherein the PEGRNA spacer includes sequence number 205 at its 3' end.

100. The prime editing system according to claim 99, wherein the ngRNA spacer includes a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of any one of sequence numbers 263-268, and optionally, the ngRNA spacer includes any one of sequence numbers 263-268 at its 3' end.

101. The prime editing system according to claim 100, wherein the ngRNA spacer contains nucleotides 1 to 20 of SEQ ID NO: 268 at its 3' end, and optionally the ngRNA contains SEQ ID NO: 824 or 825.

102. (i) The non-synonymous edit encoded by the editing template includes a deletion of c. 51delC, and the ngRNA spacer has at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 266, or (ii) The prime editing system according to claim 99, wherein the non-synonymous edit encoded by the editing template includes the insertion of c. 50insG, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 267 or 268.

103. The prime editing system according to any one of claims 99 to 102, wherein the ngRNA includes a sequence selected from the group consisting of SEQ ID NOs: 824 to 827.

104. The prime editing system according to any one of claims 96 to 98, wherein the PEGRNA spacer includes SEQ ID NO: 4 at its 3' end.

105. The prime editing system according to claim 104, wherein the ngRNA spacer includes a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of any one of sequence numbers 1017-1024.

106. The prime editing system according to claim 104, wherein the ngRNA spacer includes one of sequence numbers 1017 to 1024.

107. (i) The editing template codes for the editing of c. 54_55insCC, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1018, (ii) The editing template codes for the editing of c. 66_67insCC, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1019, (iii) The editing template codes for the editing of c. 54_55insTAAG, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1020, (iv) The editing template codes for the editing of c. 66_67insTAAG, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1021, (v) The editing template codes for the editing of c. 54_55insTAATAA, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1022, (vi) The template codes for the editing of c. 66_67insTAATAA, and the ngRNA spacer has at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1023, or (vii) The prime editing system according to claim 104, wherein the template codes for editing c. 60_65delinsTAATAG, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1024.

108. The prime editing system according to any one of claims 104 to 107, wherein the ngRNA includes a sequence selected from the group consisting of SEQ ID NOs. 1025 to 1032.

109. The prime editing system according to any one of claims 96 to 98, wherein the PEGRNA spacer includes sequence number 272 at its 3' end.

110. The prime editing system according to claim 109, wherein the ngRNA spacer includes at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of any one of sequence numbers 1152-1156.

111. The prime editing system according to claim 109, wherein the ngRNA spacer includes one of sequence numbers 1152 to 1156.

112. (i) The editing template codes for editing c.3_4insCC, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1153, (ii) The editing template codes for the editing of c.3_4insTAAG, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1154, (iii) The editing template codes for the editing of c. 3_4insTAATAA, and the ngRNA spacer has at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1155, or (iv) The prime editing system according to claim 109, wherein the editing template codes for editing c.3_8delinsTAATGA, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1156.

113. The prime editing system according to any one of claims 109 to 112, wherein the ngRNA includes a sequence selected from the group consisting of SEQ ID NOs. 1157 to 1161.

114. The prime editing system according to any one of claims 96 to 98, wherein the PEGRNA spacer includes sequence number 330 at its 3' end.

115. The prime editing system according to claim 114, wherein the ngRNA spacer includes a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of any one of sequence numbers 1280-1284.

116. The prime editing system according to claim 114, wherein the ngRNA spacer includes one of sequence numbers 1280 to 1284.

117. (i) The editing template codes for the editing of c.3_4insCC, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1281, (ii) The editing template codes for the editing of c.3_4insTAAG, and the ngRNA spacer contains at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1282, (iii) The editing template codes for the editing of c. 3_4insTAATAA, and the ngRNA spacer has at its 3' end a sequence corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of SEQ ID NO: 1283, or (iv) The prime editing system according to claim 114, wherein the editing template codes for editing c.3_8delinsTAATGA, and the ngRNA spacer has a sequence at its 3' end corresponding to nucleotides 4-20, 3-20, 2-20, or 1-20 of sequence number 1284.

118. The prime editing system according to any one of claims 115 to 117, wherein the ngRNA includes a sequence selected from the group consisting of sequence numbers 1285 to 1289.

119. (i) The PEGRNA contains the sequence described in SEQ ID NO: 933 or 937, and the ngRNA contains the sequence described in SEQ ID NO: 1018, (ii) The PEGRNA contains the sequence described in SEQ ID NO: 961, and the ngRNA contains the sequence described in SEQ ID NO: 1020, (iii) The PEGRNA contains the sequence described in SEQ ID NO: 941, and the ngRNA contains the sequence described in SEQ ID NO: 1018, (iv) The PEGRNA contains the sequence described in SEQ ID NO: 957, and the ngRNA contains the sequence described in SEQ ID NO: 1020, (v) The PEGRNA contains the sequence described in SEQ ID NO: 936, and the ngRNA contains the sequence described in SEQ ID NO: 1018, (vi) The PEGRNA contains the sequence described in SEQ ID NO: 1141 or 1143, and the ngRNA contains the sequence described in SEQ ID NO: 1156, or (vii) The prime editing system according to any one of claims 96 to 118, wherein the PEGRNA comprises the sequence described in SEQ ID NO: 1269 or 1265, and the ngRNA comprises the sequence described in SEQ ID NO: 1284.

120. The prime editing system according to any one of claims 96 to 119, wherein the ngRNA includes the modification 3'mN*mN*mN*N and / or 5'mN*mN*mN*, where m indicates that the nucleotide includes the modification 2'-O-Me, and * indicates the presence of a phosphorothioate bond.

121. The prime editing system according to any one of claims 96 to 120, wherein the ngRNA includes the modifications 3'mT*mT*mT*T and 5'mN*mN*mN*, where m indicates that the nucleotide includes the modification 2'-O-Me, * indicates the presence of a phosphorothioate bond, and T indicates the presence of a further uridine nucleotide.

122. Furthermore, it includes a prime editor or one or more polynucleotides that code the prime editor, and the prime editor is a. Cas9 niccas having a nuclease-inactivating mutation in the HNH domain, and b. A prime editing system according to any one of claims 95 to 121, comprising a reverse transcriptase.

123. The prime editing system according to claim 122, wherein the prime editor is a fusion protein.

124. moreover, An N-terminal fragment of a prime editor fusion protein and an N-terminal extension containing an N-intane, or a polynucleotide encoding the N-terminal extension, The C-terminal fragment of the prime editor fusion protein and a C-terminal extension containing a C-intane or a polynucleotide encoding the C-terminal extension, The prime editing system according to any one of claims 95 to 121, wherein the N-intane and C-intane of the N-terminus and C-terminus extensions are capable of self-cleavage to join the N-terminus fragment and the C-terminus fragment to form the prime editor fusion protein, and the prime editor fusion protein comprises a Cas9 niccas and a reverse transcriptase (RT) domain having a nuclease-inactivating mutation in the HNH domain.

125. The prime editing system according to any one of claims 122 to 124, wherein the Cas9 nickase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 676 or 677.

126. The prime editing system according to any one of claims 122 to 125, wherein the reverse transcriptase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:

673.

127. The prime editing system according to claim 125 or 126, wherein the sequence identity is determined by a Needleman-Wunsch alignment of two protein sequences with a gap presence cost of 11 and a gap extension cost of 1, and the identity percentage is calculated by dividing the number of identities by the length of the alignment.

128. The prime editing system according to any one of claims 122 to 127, wherein one or more polynucleotides encoding the prime editor, the polynucleotide encoding the N-terminal extension, or the polynucleotide encoding the C-terminal extension is mRNA.

129. A collection of viral particles comprising one or more polynucleotides encoding PEGRNA according to any one of claims 1 to 94, or a prime editing system according to any one of claims 95 to 128.

130. The group of virus particles according to claim 129, wherein the virus particles are AAV particles.

131. An LNP comprising a prime editing system according to any one of claims 95 to 128.

132. The LNP according to claim 131, comprising the PEGRNA and optionally the ngRNA, the polynucleotide encoding the Cas9 nickase, and the polynucleotide encoding the reverse transcriptase.

133. The LNP according to claim 132, wherein the polynucleotide encoding Cas9 nickase and the polynucleotide encoding the reverse transcriptase are mRNA.

134. The LNP according to claim 132 or 133, wherein the polynucleotide encoding Cas9 nickase and the polynucleotide encoding the reverse transcriptase are in the same molecule.

135. A method for editing a B2M gene, comprising contacting the B2M gene with (a) a prime editor comprising a PEGRNA according to any one of claims 1 to 94, and Cas9 nickase and reverse transcriptase having a nuclease-inactivating mutation in the HNH domain; (b) a prime editing system according to any one of claims 95 to 128; (c) a population of viral particles according to claim 129 or 130; or (d) an LNP according to any one of claims 131 to 134.

136. The method according to claim 135, wherein the B2M gene is located inside a cell.

137. A method for generating manipulated cells, comprising introducing into a cell or population of cells (a) a prime editor comprising a PEGRNA according to any one of claims 1 to 94, and Cas9 nickase and reverse transcriptase having a nuclease-inactivating mutation in the HNH domain, (b) a prime editing system according to any one of claims 95 to 128, (c) a population of viral particles according to claim 129 or 130, or (d) an LNP according to any one of claims 131 to 134.

138. The method according to claim 136 or 137, wherein the cells or population of cells are located within the target.

139. The method according to claim 136 or 137, wherein the cells or population of cells are exovivo, and optionally, the cells or population of cells are obtained from a subject or cell bank.

140. The method according to any one of claims 136 to 139, wherein the cells or the population of cells are human cells.

141. The method according to claim 140, wherein the cells or population of cells are immune cells or stem cells.

142. The method according to claim 141, wherein the cells or population of cells are T cells or hematopoietic stem cells (HSCs), and optionally the cells or population of cells are cytotoxic T cells.

143. Cells or populations of cells produced by the method described in any one of claims 136 to 142.

144. Engineered cells or populations of engineered cells that contain immature stop codons within the B2M gene, compared to the wild-type B2M gene.

145. An engineered cell or population of engineered cells containing a B2M gene that includes an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to the coding sequence position c. 51, c. 54, or c. 50 of the wild-type B2M gene, compared to the wild-type B2M gene.

146. A manipulated cell or population of manipulated cells comprising a B2M gene having an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c. 54, c. 60, or c. 66 of the wild-type B2M gene, and optionally, the B2M gene having an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c. 58 of the wild-type B2M gene.

147. A manipulated cell or population of manipulated cells comprising a B2M gene having an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.21 or c.3 of the wild-type B2M gene, and optionally, the B2M gene having an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.17 of the wild-type B2M gene.

148. A manipulated cell or population of manipulated cells comprising a B2M gene having an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.21, c.15, or c.3 of the wild-type B2M gene, and optionally, the B2M gene having an insertion, deletion, substitution, or combination thereof at a chromosomal position corresponding to coding sequence position c.11 of the wild-type B2M gene.

149. The cell or population of cells according to claim 144, wherein the B2M gene contains a deletion of c. 51delC compared to the wild-type B2M gene.

150. The cell or population of cells according to claim 144, wherein the B2M gene contains an insertion of c.50insG compared to the wild-type B2M gene.

151. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c. 54_55insCC compared to the wild-type B2M gene.

152. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c. 54_55insTAAG compared to the wild-type B2M gene.

153. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c. 54_55insTAATAA compared to the wild-type B2M gene.

154. The cell or population of cells according to claim 144, wherein the B2M gene includes an insertion of c.66_67insCC compared to the wild-type B2M gene, and optionally further includes a substitution of c.58G>C compared to the wild-type B2M gene.

155. The cell or population of cells according to claim 144, wherein the B2M gene includes an insertion of c.66_67insTAAG compared to the wild-type B2M gene, and optionally further includes a substitution of c.58G>C compared to the wild-type B2M gene.

156. The cell or population of cells according to claim 144, wherein the B2M gene includes an insertion of c.66_67insTAATAA compared to the wild-type B2M gene, and optionally further includes a substitution of c.58G>C compared to the wild-type B2M gene.

157. The cell or population of cells according to claim 144, wherein, compared to the wild-type B2M gene, the B2M gene contains a deletion of c. 60_65 and an insertion of TAATAG (c. 60_64 delinsTAATAG), and optionally further contains a substitution of c. 58G>C within the B2M gene compared to the wild-type B2M gene.

158. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c.21_22insCC compared to the wild-type B2M gene.

159. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c.21_22insTAAG compared to the wild-type B2M gene.

160. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c.21_22insTAATAA compared to the wild-type B2M gene.

161. The cell or population of cells according to claim 144, wherein the B2M gene includes an insertion of c.3_4insCC compared to the wild-type B2M gene, and optionally further includes a substitution of c.17C>G or a substitution of c.11C>G compared to the wild-type B2M gene.

162. The cell or population of cells according to claim 144, wherein the B2M gene includes an insertion of c.3_4insTAAG compared to the wild-type B2M gene, and optionally further includes a substitution of c.17C>G or c.11C>G compared to the wild-type B2M gene.

163. The cell or population of cells according to claim 144, wherein the B2M gene includes an insertion of c.3_4insTAATAA compared to the wild-type B2M gene, and optionally further includes a substitution of c.17C>G or c.11C>G compared to the wild-type B2M gene.

164. The cell or population of cells according to claim 144, wherein, compared to the wild-type B2M gene, the B2M gene contains a deletion of c.3_8 and an insertion of TAATGA (c.3_8delinsTAATGA), and optionally further contains a substitution of c.17C>G or a substitution of c.11C>G within the B2M gene, compared to the wild-type B2M gene.

165. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c.15_16insCC compared to the wild-type B2M gene.

166. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c.15_16insTAAG compared to the wild-type B2M gene.

167. The cell or population of cells according to claim 144, wherein the B2M gene contains the insertion of c.15_16insTAATAA compared to the wild-type B2M gene.

168. The cell or population of cells according to any one of claims 145 to 167, wherein the human chromosome location and the coding sequence location are as described in Genome Reference Consortium Human Build 38 (GrCh38).

169. A cell or population of cells according to any one of claims 143 to 168, which is located within the subject area.

170. An exovivo cell or population of cells according to any one of claims 143 to 168, optionally obtained from a subject or cell bank.

171. A cell or population of cells according to any one of claims 143 to 168, which is a human cell.

172. The cells or population of cells according to claim 171, which are immune cells or stem cells.

173. The cell or population of cells according to claim 172, which is a T cell or hematopoietic stem cell (HSC), and optionally a cytotoxic T cell.

174. A method of immunotherapy comprising administering to a subject (a) a prime editor comprising a PEGRNA according to any one of claims 1 to 94, and Cas9 nicasse and reverse transcriptase having a nuclease-inactivating mutation in the HNH domain; (b) a prime editing system according to any one of claims 95 to 128; (c) a population of viral particles according to claim 129 or 130; (d) an LNP according to any one of claims 131 to 134; or (e) cells or a population of cells according to any one of claims 143 to 172.

175. Prime editing guide RNA (PEGRNA) or one or more polynucleotides encoding the PEGRNA, a) A spacer complementary to the search target sequence on the first strand of the β2-microglobulin (B2M) gene, wherein the 3' end contains a PEGRNA spacer sequence selected from any one of Tables 1 to 21, b) A gRNA core capable of binding to the Cas9 protein, and c) An extendable arm, i) An editing template having an RTT sequence selected from the same table as the PEGRNA spacer sequence at its 3' end, and ii) The PEGRNA, or one or more polynucleotides encoding the PEGRNA, comprising the elongation arm, which includes a PBS having a primer-binding site (PBS) sequence selected from the same table as the PEGRNA spacer sequence at its 5' end.

176. The PEGRNA according to claim 175, wherein the spacer of the PEGRNA is 17 to 22 nucleotides long.

177. The PEGRNA according to claim 176, wherein the spacer of the PEGRNA is 20 nucleotides long.

178. The PEGRNA according to any one of claims 175 to 177, wherein the spacer, the gRNA core, the editing template, and the PBS form a continuous sequence in a single molecule.

179. PEGRNA according to claim 178, comprising the spacer, the gRNA core, the editing template, and the PBS from 5' to 3'.

180. A prime editing system comprising PEGRNA or one or more polynucleotides according to claims 175 to 179.

181. Furthermore, it comprises a nick guide RNA (ngRNA), or one or more polynucleotides encoding the ngRNA, and the ngRNA is (i) an ngRNA spacer containing a region complementary to the second strand of the B2M gene, and (ii) The prime editing system according to claim 180, comprising an ngRNA core capable of binding to a Cas9 protein.

182. The prime editing system according to claim 181, wherein the spacer of the ngRNA is 17 to 22 nucleotides long.

183. The prime editing system according to claim 182, wherein the spacer of the ngRNA is 20 nucleotides long.

184. The prime editing system according to any one of claims 175 to 182, wherein the ngRNA spacer includes an ngRNA spacer sequence selected from the same table as the PEGRNA spacer sequence at its 3' end.

185. The prime editing system according to any one of claims 175 to 182, wherein the ngRNA includes an ngRNA sequence selected from the same table as the PEGRNA spacer sequence.

186. moreover, Cas9 nicasse having a nuclease-inactivating mutation in the HNH domain, or one or more polynucleotides encoding the Cas9 nicasse, and A prime editing system according to any one of claims 175 to 185, comprising a prime editor comprising a reverse transcriptase or one or more polynucleotides encoding the reverse transcriptase.

187. moreover, An N-terminal fragment of a prime editor fusion protein and an N-terminal extension containing an N-intane, or a polynucleotide encoding the N-terminal extension, and The C-terminal fragment of the prime editor fusion protein and a C-terminal extension containing a C-intane or a polynucleotide encoding the C-terminal extension, The prime editing system according to any one of claims 175 to 185, wherein the N-intane and C-intane of the N-terminus and the C-terminus extensions are capable of self-cleavage to join the N-terminus fragment and the C-terminus fragment to form the prime editor fusion protein, and the prime editor fusion protein comprises a Cas9 niccasse and a reverse transcriptase (RT) domain.

188. The prime editing system according to claim 186 or 187, wherein the Cas9 nickase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 676 or 677.

189. The prime editing system according to any one of claims 186 to 188, wherein the reverse transcriptase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO:

673.

190. The prime editing system according to claim 188 or 189, wherein the sequence identity is determined by a Needleman-Wunsch alignment of two protein sequences with a gap presence cost of 11 and a gap extension cost of 1, and the identity percentage is calculated by dividing the number of identities by the length of the alignment.