Methods and compositions for editing edited nucleotide sequences
Prime editing addresses the challenges of high-precision genome editing by using a nucleic acid-programmable DNA-binding protein and polymerase for precise nucleotide modifications, enhancing editing efficiency and specificity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- THE BROAD INST INC
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-10
AI Technical Summary
Current genome editing technologies face challenges in achieving high-precision editing of single nucleotide mutations, particularly in human cells, with issues such as low efficiency of homology-directed repair, generation of chromosomal rearrangements, and cell type-dependent effects.
A novel genome editing method called 'prime editing' uses a nucleic acid-programmable DNA-binding protein and a polymerase to directly write new genetic information into specific DNA sites through a mechanism of target-primed reverse transcription, enabling precise incorporation of single nucleotide changes and small insertions or deletions without creating double-strand breaks.
Prime editing achieves highly efficient and flexible genome editing with high specificity, overcoming limitations of existing methods by allowing precise nucleotide modifications and reducing unwanted by-products.
Smart Images

Figure 2026062659000488 
Figure 2026062659000489 
Figure 2026062659000490
Abstract
Description
[Technical Field]
[0001] Government support This invention was made with government support under grants U01AI142756, RM1HG009490, R01EB022376, and R35GM118062, awarded by the National Institutes of Health. The government has certain rights in this invention.
[0002] Incorporation by related applications and references This U.S. provisional application is based on the following applications: U.S. provisional application No. 62 / 820,813 filed on March 19, 2019 (Agent reference number B1195.70074US00), U.S. provisional application No. 62 / 858,958 filed on June 7, 2019 (Agent reference number B1195.70074US01), and U.S. provisional application No. 62 / 889,996 filed on August 21, 2019 (Agent reference number B1195.70074US01). (Reference number B1195.70074US02), U.S. Provisional Application No. 62 / 922,654 filed on August 21, 2019 (Agent reference number B1195.70083US00), U.S. Provisional Application No. 62 / 913,553 filed on October 10, 2019 (Agent reference number B1195.70074US03), U.S. Provisional Application No. 62 / 973,558 filed on October 10, 2019 (Agent reference number B1195.70083US01), US provisional application No. 62 / 931,195 filed November 5, 2019 (Agent reference number B1195.70074US04), US provisional application No. 62 / 944,231 filed December 5, 2019 (Agent reference number B1195.70074US05), US provisional application No. 62 / 974 filed December 5, 2019 This refers to and incorporates by reference U.S. Provisional Application No. 537 (Agent Reference Number B1195.70083US02), U.S. Provisional Application No. 62 / 991,069 filed March 17, 2020 (Agent Reference Number B1195.70074US06), and U.S. Provisional Application filed March 17, 2020 (the serial number was not available at the time of filing) (Agent Reference Number B1195.70083US03). [Background technology]
[0003] Background of the present invention Pathogenic single-nucleotide mutations are estimated to contribute to approximately 50% of human diseases with a genetic component. 7 Unfortunately, despite decades of gene therapy exploration, treatment options for patients with these genetic disorders remain extremely limited. 8 Perhaps the simplest solution to this therapeutic challenge is the direct correction of a single nucleotide mutation in the patient's genome, which may address the root cause of the disease and provide lasting benefits. Such a strategy was previously unthinkable, but the CRISRP / Cas system 9 Recent improvements in genome editing capabilities, brought about by the advent of CRISPR, now put this therapeutic approach within reach. By simply designing a guide RNA (gRNA) sequence containing approximately 20 nucleotides complementary to the target DNA sequence, virtually any conceivable genomic region can be specifically accessed by CRISPR-related (Cas) nucleases. 1,2 To date, several monomeric bacterial Cas nuclease systems have been identified and adapted for genome editing applications. 10 This natural diversity of Cas nuclease, along with the ever-growing group of modified variants, 11~14 This will create fertile ground for developing new genome editing technologies.
[0004] Although gene disruption using CRISPR is now a mature technique, high-precision editing of single base pairs in the human genome remains a significant challenge. 3 Homology-directed repair (HDR) has long been used in human cells and other organisms to insert, modify, or exchange DNA sequences at double-strand break (DSB) sites using donor DNA repair templates that encode the desired edits. 15However, existing HDR has extremely low efficiency in most human cell types, specifically in non-dividing cells, and non-homologous end joining (NHEJ), which is incompatible, mainly leads to insertion-deletion (indel) by-products. 16 Another problem relates to the generation of DSBs, which can cause large chromosomal rearrangements and deletions at the target locus 17 or activate the p53 axis leading to growth arrest and apoptosis. 18,19 .
[0005] Several approaches have been explored to address these drawbacks of HDR. For example, repair with oligonucleotide donors of single-strand DNA breaks (nicks) has been shown to reduce indel formation, but the yield of the desired repair product is still low. 20 Other strategies attempt to bias repair towards HDR rather than NHEJ using small molecules and biological reagents. 21~23 However, the effectiveness of these methods can be cell type-dependent, and disruption of the normal cell state can lead to undesirable and unpredictable effects.
[0006] In recent years, the inventors led by Professor David Liu have developed base editing as a technology for editing target nucleotides without creating DSBs or relying on HDR. 4~6,24~27Direct modification of DNA bases by Cas fusion deaminase enables highly efficient C·G→T·A or A·T→G·C base pair conversions in short target windows (~5-7 bases). As a result, base editors have been rapidly adopted by the scientific community. However, the following factors limit their generality for high-precision genome editing: (1) "bystander editing" of non-target C or A bases on the target window is observed; (2) a mixture of target nucleotide products is observed; (3) the target base must be located 15±2 nucleotides upstream of the PAM sequence; and (5) repair of small insertion and deletion mutations is not possible.
[0007] Therefore, the development of programmed editing factors that can flexibly introduce any desired single nucleotide change and / or incorporate base pair insertions or deletions (e.g., insertions or deletions of at least 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more base pairs) and / or alter or modify nucleotide sequences at target sites with high specificity and efficiency would substantially expand the scope and therapeutic potential of CRISPR-based genome editing technologies. [Overview of the Initiative]
[0008] Summary of the present invention This invention describes a novel platform for genome editing called "prime editing." Prime editing is a versatile and precise genome editing method that uses a nucleic acid-programmable DNA-binding protein ("napDNAbp") in cooperation with polymerase (i.e., in the form of a fusion protein or otherwise provided in trans with napDNAbp) to directly write new genetic information into specific DNA sites. The prime editing system is programmed with a prime editing (PE) guide RNA ("PEgRNA") that identifies both the target site and template for the synthesis of the desired edit in the form of a replacement DNA strand via an extension (either DNA or RNA) that is modified onto (e.g., at the 5' or 3' end of the guide RNA, or within the guide RNA). The replacement strand containing the desired edit (e.g., a single nucleic acid base substitution) shares the same sequence as the endogenous strand of the target site to be edited (with the exception that it contains the desired edit). Through DNA repair and / or replication mechanisms, the endogenous strand of the target site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing can be considered a "search-and-replace" genome editing technique, as the prime editing factors described herein not only search for and locate the desired target sites to be edited, but also simultaneously encode a replacement strand containing the desired edits to be incorporated in place of the corresponding target sites on the endogenous DNA strand.
[0009] The prime editing factors of this disclosure relate in part to the discovery that the mechanism of target-primed reverse transcription (TPRT) or "prime editing" can be utilized or employed to perform highly efficient and genetically flexible high-precision CRISPR / Cas-based genome editing (as illustrated, for example, in various embodiments in Figures 1A–1F). TPRT is naturally used by mobile DNA elements such as mammalian non-LTR retrotransposons and bacterial group II introns.28,29The inventors herein use Cas protein-reverse transcriptase fusions or related systems in which a specific DNA sequence is targeted with guide RNA, a single-strand nick is generated at the target site, and the nicked DNA is used as a primer for reverse transcription of a modified reverse transcriptase template incorporated by the guide RNA. However, although this concept began with prime editing factors that use reverse transcriptase as a component of DNA polymerase, the prime editing factors described herein are not limited to reverse transcriptase and may effectively encompass the use of DNA polymerase. In fact, this application sometimes refers to prime editing factors that have “reverse transcriptase” throughout, but it is explained herein that reverse transcriptase is the only type of DNA polymerase that can work in prime editing. Therefore, wherever “reverse transcriptase” is referred to herein, those skilled in the art will understand that any suitable DNA polymerase may be used instead of reverse transcriptase. Therefore, in one aspect, the prime editing factor may include Cas9 (or equivalent napDNAbp) programmed to target a DNA sequence by linking the DNA sequence to a specialized guide RNA (i.e., PEgRNA) containing a spacer sequence that anneals with a complementary protospacer in the target DNA. The specialized guide RNA also contains new genetic information in the form of an elongation encoding a replacement strand of DNA containing the desired genetic modification, which is used to replace the corresponding endogenous DNA strand at the target site. To transfer the information from the PEgRNA to the target DNA, the prime editing mechanism involves nicking a target site in one strand of DNA, exposing a 3' hydroxyl group. The exposed 3' hydroxyl group can then be used to prime the direct DNA polymerization of the edit-encoding elongation onto the PEgRNA into the target site. In various embodiments, the elongation—which provides a template for the polymerization of the replacement strand containing the edit—can be formed from RNA or DNA. In the case of RNA elongation, the polymerase of the prime editing factor may be an RNA-dependent DNA polymerase (such as reverse transcriptase).In the case of DNA elongation, the polymerase of the prime editing factor may be a DNA-dependent DNA polymerase.
[0010] The newly synthesized strand (i.e., the replacement DNA strand containing the desired edit) formed by the prime editing factor disclosed herein will be homologous to the target sequence of the genome (i.e., have the same sequence), except that it contains the desired nucleotide change (e.g., a single nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized (or replacement) DNA strand may also be referred to as a single-strand DNA flap, which competes for hybridization with a complementary homologous endogenous DNA strand, thereby replacing the corresponding endogenous strand. In some embodiments, the system may be combined with the use of an error-prone reverse transcriptase (e.g., provided as a fusion protein with a Cas9 domain, or provided trans with a Cas9 domain). The error-prone reverse transcriptase can introduce the change during single-strand DNA flap synthesis. Thus, in some embodiments, the error-prone reverse transcriptase can be used to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used with the system, the changes may be random or non-random.
[0011] The degradation of hybridized intermediates (including single-strand DNA flaps synthesized by reverse transcriptase hybridized with endogenous DNA strands) can encompass the removal of the replaced flap (e.g., by 5'-terminus DNA flap endonuclease, FEN1) from the resulting endogenous DNA, ligation of the synthesized single-strand DNA flap to target DNA, and assimilation of desired nucleotide changes as a result of the cell's DNA repair and / or replication processes. Because templated DNA synthesis imparts single-nucleotide precision to any nucleotide modification, including insertions and deletions, the breadth of this approach is extremely broad, and it is foreseeable that it could be used for countless applications in basic science and therapeutics.
[0012] In one aspect, this specification provides a fusion protein comprising a nucleic acid-programmed DNA-binding protein (napDNAbp) and a reverse transcriptase. In various embodiments, the fusion protein can perform genome editing by reverse transcription primed to a prime target in the presence of an extended guide RNA.
[0013] In one embodiment, napDNAbp possesses nickase activity. napDNAbp may also be the Cas9 protein or its functional equivalent, such as nuclease-active Cas9, nuclease-inactive Cas9 (dCas9), or Cas9 nickase (nCas9).
[0014] In one embodiment, napDNAbp is selected from the group consisting of Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, Cas12c, and Argonaut, and optionally possesses nickase activity.
[0015] In another embodiment, the fusion protein can bind to a target DNA sequence when complexed with an extended guide RNA.
[0016] In other embodiments, the target DNA sequence includes a target strand and a complementary non-target strand.
[0017] In another embodiment, the binding of the complexed fusion protein to the extended guide RNA forms an R-loop. The R-loop may include (i) an RNA-DNA hybrid comprising the extended guide RNA and the target strand, and (ii) a complementary non-target strand.
[0018] In another embodiment, the complementary non-target chain is nicked to form a reverse transcriptase prime sequence with a free 3' end.
[0019] In various embodiments, the elongated guide RNA comprises (a) a guide RNA and (b) RNA elongation at the 5' or 3' end of the guide RNA or at an intramolecular position of the guide RNA. The RNA elongation may comprise (i) a reverse transcription template sequence containing the desired nucleotide changes, (ii) a reverse transcription primer binding site, and (iii) optionally a linker sequence. In various embodiments, the reverse transcription template sequence may encode a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to the nick site, the single-stranded DNA flap containing the desired nucleotide changes.
[0020] In various embodiments, the RNA elongation is of a length of at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides.
[0021] In another embodiment, the single-stranded DNA flap can hybridize to an endogenous DNA sequence adjacent to the nick site, thereby incorporating a desired nucleotide change. In yet another embodiment, the single-stranded DNA flap replaces an endogenous DNA sequence having a free 5' end and adjacent to the nick site. In one embodiment, the replaced endogenous DNA having a 5' end is excised by the cell.
[0022] In various embodiments, cellular repair of a single-stranded DNA flap results in the incorporation of a desired nucleotide change, thereby forming a desired product.
[0023] In various other embodiments, the desired nucleotide changes are incorporated into an editing window located between approximately -4 and +10 in the PAM sequence.
[0024] In another embodiment, the desired nucleotide change is incorporated into an editing window where the nick site is between approximately -5 and +5, or between approximately -10 and +10, or between approximately -20 and +20, or between approximately -30 and +30, or between approximately -40 and +40, or between approximately -50 and +50, or between approximately -60 and +60, or between approximately -70 and +70, or between approximately -80 and +80, or between approximately -90 and +90, or between approximately -100 and +100, or between approximately -200 and +200.
[0025] In various embodiments, napDNAbp includes the amino acid sequence of SEQ ID NO: 18. In various other embodiments, napDNAbp includes an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs: 26-39, 42-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467, and 482-487 (Cas9); (SpCas9); SEQ ID NOs: 77-86 (CP-Cas9); SEQ ID NOs: 18-25 and 87-88 (SpCas9); and SEQ ID NOs: 62-72 (Cas12).
[0026] In other embodiments, the reverse transcriptase of the disclosed fusion protein and / or composition may comprise any one of the amino acid sequences of SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. In other embodiments, the reverse transcriptase may contain an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. These sequences may be naturally occurring reverse transcriptase sequences from, for example, retroviruses or retrotransposons, or the sequences may be recombinant.
[0027] In various other embodiments, the fusion proteins disclosed herein may include various structural configurations. For example, the fusion protein may include the structure NH2-[napDNAbp]-[reverse transcriptase]-COOH; or NH2-[reverse transcriptase]-[napDNAbp]-COOH, where each of the ]-[ indicates the presence of any linker sequence.
[0028] In various embodiments, the linker sequence includes an amino acid sequence that is at least 80%, 85%, 90%, 95%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs: 127, 165-176, 446, 453, and 767-769, or to any one of the linker amino acid sequences of SEQ ID NOs: 127, 165-176, 446, 453, and 767-769.
[0029] In various embodiments, the desired nucleotide change incorporated into the target DNA may be a single nucleotide change (e.g., a transition or transversion), an insertion of one or more nucleotides, or a deletion of one or more nucleotides.
[0030] In certain cases, the insertion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.
[0031] In some other cases, the deletion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.
[0032] In another aspect, the disclosure provides an elongated guide RNA comprising a guide RNA and at least one RNA elongation. The RNA elongation may be located at the 3' end of the guide RNA. In another embodiment, the RNA elongation may be located at the 5' end of the guide RNA. In yet another embodiment, the RNA elongation may be located at an intramolecular position on the guide RNA. However, preferably, the intramolecular positioning of the elongated portion does not obstruct the function of the protospacer.
[0033] In various embodiments, prime editing factor guide RNA (PEgRNA) can bind to napDNAbp and direct the napDNAbp towards a target DNA sequence. The target DNA sequence may include a target strand and a complementary non-target strand, where the guide RNA hybridizes with the target strand to form an RNA-DNA hybrid and an R-loop.
[0034] In various embodiments of the prime editing factor guide RNA, at least one RNA elongation includes a DNA synthesis template. In various other embodiments, the RNA elongation further includes a reverse transcription primer binding site. In yet another embodiment, the RNA elongation includes a linker or spacer that binds the RNA elongation to the guide RNA.
[0035] In various embodiments, the RNA elongation may have a length of at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0036] In another embodiment, the DNA synthesis template (i.e., the editing template for Figure 27) is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0037] In yet another embodiment, the reverse transcription primer binding site sequence (i.e., primer binding site for Figure 27) is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0038] In another embodiment, any linker or spacer has a length of at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides.
[0039] In various embodiments of the extended guide RNA, the reverse transcription template sequence may encode a single-stranded DNA flap complementary to an endogenous DNA sequence adjacent to a nick site, the single-stranded DNA flap containing the desired nucleotide changes. The single-stranded DNA flap may replace the endogenous single-stranded DNA at the nick site. The replaced endogenous single-stranded DNA at the nick site may have a 5' end, forming an endogenous flap that can be excised by the cell. In various embodiments, excision of the 5' endogenous flap may drive product formation because removing the 5' endogenous flap promotes hybridization of the single-stranded 3' DNA flap to the corresponding complementary DNA strand and the incorporation or assimilation of the desired nucleotide changes carried by the single-stranded 3' DNA flap into the target DNA.
[0040] In various forms of extended guide RNA, cellular repair of single-stranded DNA flaps results in the incorporation of desired nucleotide changes, thereby forming the desired product.
[0041] In one embodiment, PEgRNA is sequence numbers 131, 222, 394, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 641, 642, 643, 644, 645, 646, 647, 648, 649, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 738, 2997, 2998, 2999, 3000, 3001, 3002, 3003, 3004, 3005, 3006, 3007, 3008, 3009, 3010, 3011, 3012, 3013, 3014, 3015, 3016, 3017, 3018, 3019, 3020, 3021, 3022, 3023, 3024, 3025, 3026, 3027, 3028, 3029, 3030, 3031, 3032, 3033, 3034, 3035, 3036, 3037, 3038, 3039, 3040, 3041, 3042, 3043, 3044, 3045, 3046, 3047, 3048, 3049, 3050, 3051, 3052, 3053, 3054, 3055, 3056, 3057, 3058, 3059, 3060, 3061, 3062, 3063, 3064, 3065, 3066, 3067, 3068, 3069, 3070, 3071, 3072, 3073, 3074, 3075, 3076, 3077, 3078, 3079, 3080, 3081, 3082, 3083, 3084, 3085, 3086, 3087, 3088, 3089, 3090, 3091, 3092, 3093, 3094, 3095, 3096, 3097, 3098, 3099, 3100, 3101, 3102, 3103, 3113, 3114, 3115, 3116, 3117, 3118, 3119, 3120, 3121, 3305, 3306, 3307, 3308, 3309, 3310, 3311, 3312, 3313, 3314, 3315, 3316, 3317, 3318, 3319, 3320, 3321, 3322, 3323, 3324, 3325, 3326, 3327, 3328, 3329, 3330, 3331, 3332, 3333, 3334, 3335, 3336, 3337, 3338, 3339, 3340, 3341, 3342, 3343, 3344, 3345, 3346, 3347, 3348, 3349, 3350,3351、3352、3353、3354、3355、3356、3357、3358、3359、3360、3361、3362、3363、3364、3365、3366、3367、3368、3369、3370、3371、3372、3373、3374、3375、3376、3377、3378、3379、3380、3381、3382、3383、3384、3385、3386、3387、3388、3389、3390、3391、3392、3393、3394、3395、3396、3397、3398、3399、3400、3401、3402、3403、3404、3405、3406、3407、3408、3409、3410、3411、3412、3413、3414、3415、3416、3417、3418、3419、3420、3421、3422、3423、3424、3425、3426、3427、3428、3429、3430、3431、3432、3433、3434、3435、3436、3437、3438、3439、3440、3441、3442、3443、3444、3445、3446、3447、3448、3449、3450、3451、3452、3453、3454、3455、3479、3480、3481、3482、3483、3484、3485、3486、3487、3488、3489、3490、3491、3492、3493、3522、3523、3524、3525、3526、3527、3528、3529、3530、3531、3532、3533、3534、3535、3536、3537、3538、3539、3540、3549、3550、3551、3552、3553、3554、3555、3556、3628、3629、3630、3631、3632、3633、3634、3635、3636、3637、3638、3639、3640、3641、3642、3643、3644、3645、3646、3647、3648、3649、3650、3651、3652、3653、3654、3655、3656、3657、3658、3659、3660、3661、3662、3663、3664、3665、3666、3667、3668、3669、3670、3671、3672、3673、3674、3675、3676、3677、3678、3679、3680、3681, 3682, 3683, 3684, 3685, 3686, 3687, 3688, 3689, 3690, 3691, 3692, 3693, 3694, 3695, 3696, 3697, 3698, 3755, 3756, 3757, 3758, 3759, 3760, 3761 ,3762,3763,3764,3765,3766,3767,3768,3769,3770,3771,3772,3773,3774,3775,3776,3777,3778,3779,3780,3781,3782,3783,3784,3785,3786 nucleotide sequences of 3787, 3788, 3789, 3790, 3791, 3792, 3793, 3794, 3795, 3796, 3797, 3798, 3799, 3800, 3801, 3802, 3803, 3804, 3805, 3806, 3807, 3808, 3809, and 3810, or sequence numbers 131, 222, 394, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 641, 642, 643, 644, 645, 646, 647, 648, 649, 678, 6 79, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 738, 2997, 2998, 2999, 3000, 3001, 3002, 3003, 3004, 3005, 3006, 3007, 3008, 3009, 3010, 3011, 3012, 3013, 3014, 3015, 3016, 3017, 3018, 3019, 3020, 3021, 3022, 3023, 3024, 3025, 3026, 3027, 3028, 3029, 3030, 3031, 3032, 3033, 3034, 3035, 3036, 3037, 3038, 3039, 3040, 3041, 3042, 3043, 3044, 3045, 3046, 3047, 3048, 3049, 3050, 3051, 3052, 3053, 3054, 3055, 3056, 3057, 3058, 3059, 3060, 3061, 3062, 3063, 3064, 3065, 3066, 3067, 3068, 3069, 3070, 3071, 3072, 3073, 3074, 3075, 3076, 3077, 3078, 3079, 3080, 3081, 3082, 3083, 3084,3085、3086、3087、3088、3089、3090、3091、3092、3093、3094、3095、3096、3097、3098、3099、3100、3101、3102、3103、3113、3114、3115、3116、3117、3118、3119、3120、3121、3305、3306、3307、3308、3309、3310、3311、3312、3313、3314、3315、3316、3317、3318、3319、3320、3321、3322、3323、3324、3325、3326、3327、3328、3329、3330、3331、3332、3333、3334、3335、3336、3337、3338、3339、3340、3341、3342、3343、3344、3345、3346、3347、3348、3349、3350、3351、3352、3353、3354、3355、3356、3357、3358、3359、3360、3361、3362、3363、3364、3365、3366、3367、3368、3369、3370、3371、3372、3373、3374、3375、3376、3377、3378、3379、3380、3381、3382、3383、3384、3385、3386、3387、3388、3389、3390、3391、3392、3393、3394、3395、3396、3397、3398、3399、3400、3401、3402、3403、3404、3405、3406、3407、3408、3409、3410、3411、3412、3413、3414、3415、3416、3417、3418、3419、3420、3421、3422、3423、3424、3425、3426、3427、3428、3429、3430、3431、3432、3433、3434、3435、3436、3437、3438、3439、3440、3441、3442、3443、3444、3445、3446、3447、3448、3449、3450、3451、3452、3453、3454、3455、3479、3480、3481、3482、3483、3484、3485、3486、3487、3488、3489、3490、3491、3492、3493、3522、3523、3524、3525、3526、3527、3528, 3529, 3530, 3531, 3532, 3533, 3534, 3535, 3536, 3537, 3538, 3539, 3540, 3549, 3550, 3551, 3552, 3553, 3554, 3555, 3556, 3628, 3629, 3630, 3631, 3632, 3633, 3634, 3635, 3636, 3637, 3638, 3639, 3640, 3641, 3642, 3643, 3644, 3645, 3646, 3647, 364 8, 3649, 3650, 3651, 3652, 3653, 3654, 3655, 3656, 3657, 3658, 3659, 3660, 3661, 3662, 3663, 3664, 3665, 3666, 3667, 3668, 3669, 3670, 3671, 3672, 3673, 3674, 3675, 3676, 3677, 3678, 3679, 3680, 3681, 3682, 3683, 3684, 3685, 3686, 3687, 3688, 3689, 3 690, 3691, 3692, 3693, 3694, 3695, 3696, 3697, 3698, 3755, 3756, 3757, 3758, 3759, 3760, 3761, 3762, 3763, 3764, 3765, 3766, 3767, 3768, 3769, 3770, 3771, 3772, 3773, 3774, 3775, 3776, 3777, 3778, 3779, 3780, 3781, 3782, 3783, 3784, 3785, 3786, 3787 It contains a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with any one of 3788, 3789, 3790, 3791, 3792, 3793, 3794, 3795, 3796, 3797, 3798, 3799, 3800, 3801, 3802, 3803, 3804, 3805, 3806, 3807, 3808, 3809, and 3810.
[0042] In another aspect of the present invention, this specification provides a complex comprising the fusion protein described herein and any of the elongated guide RNAs described above.
[0043] In another further aspect of the present invention, this specification provides a complex comprising napDNAbp and an extended guide RNA. napDNAbp may be Cas9 nickase, or may be an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to the amino acid sequences of SEQ ID NOs. 42-57 (Cas9 nickase) and 65 (AsCas12a nickase), or one of the amino acid sequences of SEQ ID NOs. 42-57 (Cas9 nickase) and 65 (AsCas12a nickase).
[0044] In various embodiments involving the complex, the extended guide RNA can direct napDNAbp to the target DNA sequence. In various embodiments, the reverse transcriptase may be supplied trans, i.e., from a source different from the complex itself. For example, the reverse transcriptase may be supplied to the same cell containing the complex by introducing separate vectors that individually encode the reverse transcriptase.
[0045] In yet another aspect, this specification provides polynucleotides. In one embodiment, a polynucleotide may encode any of the fusion proteins disclosed herein. In certain other embodiments, a polynucleotide may encode any of the napDNAbp disclosed herein. In yet another embodiment, a polynucleotide may encode any of the reverse transcriptases disclosed herein. In yet another embodiment, a polynucleotide may encode any of the extended guide RNAs, any of the reverse transcription template sequences, any of the reverse transcription primer sites, or any linker sequence disclosed herein.
[0046] In other aspects, this specification provides vectors comprising the polynucleotides described herein. Therefore, in one embodiment, the vector comprises a polynucleotide encoding a fusion protein comprising napDNAbp and reverse transcriptase. In another embodiment, the vector comprises polynucleotides separately encoding napDNAbp and reverse transcriptase. In yet another embodiment, the vector may comprise a polynucleotide encoding an extended guide RNA. In various embodiments, the vector may comprise one or more polynucleotides encoding napDNAbp, reverse transcriptase and extended guide RNA on the same or separate vector.
[0047] In other respects, this specification provides cells comprising the fusion protein and extended guide RNA described herein. The cells may be transformed by a vector comprising the fusion protein, napDNAbp, reverse transcriptase, and extended guide RNA. These genetic elements may be contained on the same vector or on different vectors.
[0048] In another aspect, this specification provides pharmaceutical compositions. In one embodiment, a pharmaceutical composition comprises one or more of napDNAbp, fusion proteins, reverse transcriptase, and extended guide RNA. In one embodiment, a fusion protein as described herein and a pharmaceutically acceptable excipient. In another embodiment, a pharmaceutical composition comprises any of the extended guide RNAs as described herein and a pharmaceutically acceptable excipient. In yet another embodiment, a pharmaceutical composition comprises any of the extended guide RNAs as described herein in combination with any of the fusion proteins and pharmaceutically acceptable excipients as described herein. In yet another embodiment, a pharmaceutical composition comprises any polynucleotide sequence encoding one or more of napDNAbp, fusion proteins, reverse transcriptase, and extended guide RNA. In yet another embodiment, various components of the disclosed herein may be separated into one or more pharmaceutical compositions. For example, a first pharmaceutical composition may contain a fusion protein or napDNAbp, a second pharmaceutical composition may contain a reverse transcriptase, and a third pharmaceutical composition may contain an extended guide RNA.
[0049] In further aspects, this disclosure provides kits. In one embodiment, a kit comprises one or more polynucleotides encoding one or more components comprising a fusion protein, napDNAbp, reverse transcriptase, and an extended guide RNA. The kit may also comprise isolated preparations of vectors, cells, and polypeptides comprising any of the fusion proteins, napDNAbp, or reverse transcriptases disclosed herein.
[0050] In another aspect, this disclosure provides methods for using the disclosed compositions of matter.
[0051] In one embodiment, the method relates to a method for incorporating a desired nucleotide change onto a double-stranded DNA sequence. The method first involves contacting the double-stranded DNA sequence with a complex comprising a fusion protein and an extended guide RNA, wherein the fusion protein comprises napDNAbp and reverse transcriptase, and the extended guide RNA comprises a reverse transcription template sequence containing the desired nucleotide change. Next, the method involves nicking the double-stranded DNA sequence on a non-target strand, thereby generating a free single-stranded DNA having a 3' end. Then, the method involves hybridizing the 3' end of the free single-stranded DNA with the reverse transcription template sequence, thereby priming the reverse transcriptase domain. Then, the method involves polymerizing the DNA strand from the 3' end, thereby generating a single-stranded DNA flap containing the desired nucleotide change. Then, the method involves replacing the endogenous DNA strand adjacent to the cut site with the single-stranded DNA flap, thereby incorporating the desired nucleotide change onto the double-stranded DNA sequence.
[0052] In another embodiment, the disclosure provides a method for introducing one or more changes into the nucleotide sequence of a DNA molecule at a target locus, comprising contacting the DNA molecule with a nucleic acid-programmed DNA-binding protein (napDNAbp) and a guide RNA that causes the napDNAbp to target the target locus, wherein the guide RNA comprises a reverse transcriptase (RT) template sequence containing at least one desired nucleotide change. The method then involves forming an exposed 3' end in the DNA strand at the target locus, and then priming reverse transcription by hybridizing the exposed 3' end with the RT template sequence. Next, a single-strand DNA flap containing at least one desired nucleotide change based on the RT template sequence is synthesized or polymerized by reverse transcriptase. Finally, at least one desired nucleotide change is incorporated into the corresponding endogenous DNA, thereby introducing one or more changes into the nucleotide sequence of the DNA molecule at the target locus.
[0053] In further other embodiments, the Disclosure provides a method for introducing one or more changes into the nucleotide sequence of a DNA molecule at a target locus by targeted-primed reverse transcription, the method comprising: (a) contacting the DNA molecule at the target locus with a guide RNA comprising i) a fusion protein comprising a nucleic acid-programmed DNA-binding protein (napDNAbp) and reverse transcriptase, and (ii) an RT template containing the desired nucleotide changes; (b) causing targeted-primed reverse transcription of the RT template to generate single-stranded DNA containing the desired nucleotide changes; and (c) incorporating the desired nucleotide changes into the DNA molecule at the target locus through DNA repair and / or replication processes.
[0054] In one embodiment, the step of replacing an endogenous DNA strand includes: (i) creating a sequence mismatch by hybridizing a single-stranded DNA flap with an endogenous DNA strand adjacent to the cleavage site; (ii) excising the endogenous DNA strand; and (iii) repairing the mismatch to form a desired product containing the desired nucleotide changes in both strands of DNA.
[0055] In various embodiments, the desired nucleotide change may be a single nucleotide substitution (e.g., a transition or transversion), deletion, or insertion. For example, the desired nucleotide change may be (1) a G to T substitution, (2) a G to A substitution, (3) a G to C substitution, (4) a T to G substitution, (5) a T to A substitution, (6) a T to C substitution, (7) a C to G substitution, (8) a C to T substitution, (9) a C to A substitution, (10) an A to T substitution, (11) an A to G substitution, or (12) an A to C substitution.
[0056] In other embodiments, the desired nucleotide change may be (1) a G:C base pair to a T:A base pair, (2) a G:C base pair to an A:T base pair, (3) a G:C base pair to a C:G base pair, (4) a T:A base pair to a G:C base pair, (5) a T:A base pair to an A:T base pair, (6) a T:A base pair to a C:G base pair, (7) a C:G base pair to a G:C base pair, (8) a C:G base pair to a T:A base pair, (9) a C:G base pair to an A:T base pair, (10) an A:T base pair to a T:A base pair, (11) an A:T base pair to a G:C base pair, or (12) an A:T base pair to a C:G base pair.
[0057] In another embodiment, the method introduces a desired nucleotide change, which is an insertion. In certain cases, the insertion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.
[0058] In another aspect, the method introduces a desired nucleotide change, which is a deletion. In some other cases, the deletion is of a length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.
[0059] In various embodiments, the desired nucleotide alteration modifies disease-associated genes. Disease-associated genes may be associated with monogenic disorders selected from the group consisting of: adenosine deaminase (ADA) deficiency; alpha-1 antitrypsin deficiency; cystic fibrosis; Duchenne muscular dystrophy; galactosemia; hemochromatosis; Huntington's disease; maple syrup urine disease; Marfan syndrome; neurofibromatosis type 1; onychoplasia; phenylketonuria; severe combined immunodeficiency; sickle cell anemia; Smith-Lemle-Oppitz syndrome; and Tay-Sachs disease. In other embodiments, disease-associated genes may be associated with polygenic disorders selected from the group consisting of: heart disease; hypertension; Alzheimer's disease; arthritis; diabetes mellitus; cancer; and obesity.
[0060] The methods disclosed herein may involve a fusion protein having napDNAbp, which is a nuclease-inactive (dead) Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9. In other embodiments, napDNAbp and reverse transcriptase may be provided in separate constructs rather than encoded as a single fusion protein. Thus, in some embodiments, reverse transcriptase may be provided trans to napDNAbp (rather than as a fusion protein).
[0061] In various embodiments of the method, napDNAbp may include the amino acid sequences of SEQ ID NOs. 26-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467 and 482-487 (Cas9); (SpCas9); SEQ ID NOs. 77-86 (CP-Cas9); SEQ ID NOs. 18-25 and 87-88 (SpCas9); and SEQ ID NOs. 62-72 (Cas12). napDNAbp may also include amino acid sequences that are at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 26-61, 75-76, 126, 130, 137, 141, 147, 153, 157, 445, 460, 467, and 482-487 (Cas9); (SpCas9); SEQ ID NOs. 77-86 (CP-Cas9); SEQ ID NOs. 18-25 and 87-88 (SpCas9); and SEQ ID NOs. 62-72 (Cas12).
[0062] In various embodiments of the method, the reverse transcriptase may contain any one of the amino acid sequences of SEQ ID NOs. 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766. The reverse transcriptase may also contain an amino acid sequence that is at least 80%, 85%, 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700-716, 739-742, and 766.
[0063] The method is as follows: Sequence numbers 131, 222, 394, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 641, 642, 643, 644, 645, 646, 647, 648, 649, 678, 679, 680, 6 81, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 738, 2997, 2998, 2999, 3000, 3001, 3002, 3003, 3004, 3005, 3006, 3007, 3008, 3009, 3010, 3011 ,3012,3013,3014,3015,3016,3017,3018,3019,3020,3021,3022,3023,3024,3025,3026,3027,3028,3029,3030,3031,3032,3033,3034,3035,3036 ,3037,3038,3039,3040,3041,3042,3043,3044,3045,3046,3047,3048,3049,3050,3051,3052,3053,3054,3055,3056,3057,3058,3059,3060,3061 ,3062,3063,3064,3065,3066,3067,3068,3069,3070,3071,3072,3073,3074,3075,3076,3077,3078,3079,3080,3081,3082,3083,3084,3085,3086 ,3087,3088,3089,3090,3091,3092,3093,3094,3095,3096,3097,3098,3099,3100,3101,3102,3103,3113,3114,3115,3116,3117,3118,3119,3120 ,3121,3305,3306,3307,3308,3309,3310,3311,3312,3313,3314,3315,3316,3317,3318,3319,3320,3321,3322,3323,3324,3325,3326,3327,3328 ,3329,3330,3331,3332,3333,3334,3335,3336,3337,3338,3339,3340,3341,3342,3343,3344,3345,3346,3347,3348,3349,3350,3351,3352,3353,3354、3355、3356、3357、3358、3359、3360、3361、3362、3363、3364、3365、3366、3367、3368、3369、3370、3371、3372、3373、3374、3375、3376、3377、3378、3379、3380、3381、3382、3383、3384、3385、3386、3387、3388、3389、3390、3391、3392、3393、3394、3395、3396、3397、3398、3399、3400、3401、3402、3403、3404、3405、3406、3407、3408、3409、3410、3411、3412、3413、3414、3415、3416、3417、3418、3419、3420、3421、3422、3423、3424、3425、3426、3427、3428、3429、3430、3431、3432、3433、3434、3435、3436、3437、3438、3439、3440、3441、3442、3443、3444、3445、3446、3447、3448、3449、3450、3451、3452、3453、3454、3455、3479、3480、3481、3482、3483、3484、3485、3486、3487、3488、3489、3490、3491、3492、3493、3522、3523、3524、3525、3526、3527、3528、3529、3530、3531、3532、3533、3534、3535、3536、3537、3538、3539、3540、3549、3550、3551、3552、3553、3554、3555、3556、3628、3629、3630、3631、3632、3633、3634、3635、3636、3637、3638、3639、3640、3641、3642、3643、3644、3645、3646、3647、3648、3649、3650、3651、3652、3653、3654、3655、3656、3657、3658、3659、3660、3661、3662、3663、3664、3665、3666、3667、3668、3669、3670、3671、3672、3673、3674、3675、3676、3677、3678、3679、3680、3681、3682、3683、3684, 3685, 3686, 3687, 3688, 3689, 3690, 3691, 3692, 3693, 3694, 3695, 3696, 3697, 3698, 3755, 3756, 3757, 3758, 3759, 3760, 3761, 3762, 3763, 3764, 3765, 3766, 3767, 3768, 3769, 3770, 3771, 3772, 3773, 3774, 3775, 3776, 3777, 3778, 3779, 3780, 3781, 3782, 3783, 3784, 3785, 3786, 3 The method may involve the use of PEgRNA containing the nucleotide sequences 787, 3788, 3789, 3790, 3791, 3792, 3793, 3794, 3795, 3796, 3797, 3798, 3799, 3800, 3801, 3802, 3803, 3804, 3805, 3806, 3807, 3808, 3809, and 3810, or nucleotide sequences having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with these. The method may also involve the use of an elongated guide RNA containing RNA elongation at its 3' end, where the RNA elongation contains a reverse transcriptase template sequence.
[0064] The method may involve the use of an elongated guide RNA containing RNA elongation at its 5' end, where the RNA elongation contains a reverse transcription template sequence.
[0065] The method may involve the use of an elongated guide RNA, including RNA elongation, to position the guide RNA within the molecule, and the RNA elongation includes a reverse transcription template sequence.
[0066] The method may include the use of an elongated guide RNA having one or more RNA elongations that are at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides in length.
[0067] It should be understood that the above concepts, and the additional concepts discussed below, can be arranged in any preferred combination (however, this disclosure is not limited thereto). Furthermore, other advantages and novel features of this disclosure will become apparent from the detailed description of various non-limiting aspects below, when considered in conjunction with the attached figures. [Brief explanation of the drawing]
[0068] Simple description of the drawing The following drawings form part of this specification and are included to further demonstrate certain aspects of the disclosure that can be better understood by referring to one or more of these drawings in combination with the detailed description of the particular aspects presented herein.
[0069] [Figure 1A]Figure 1A provides a schematic diagram of an exemplary process for introducing single-nucleotide alterations, and / or insertions, and / or deletions, into a DNA molecule (e.g., a genome) using a fusion protein comprising a reverse transcriptase fused with a Cas9 protein conjugated with an extended guide RNA molecule. In this embodiment, the guide RNA contains the reverse transcriptase template sequence by being extended at its 3' end. The schematic diagram shows how the reverse transcriptase (RT), fused with Cas9 nickase and conjugated with the guide RNA (gRNA), binds to the DNA target site and nicks the PAM-containing DNA strand adjacent to the target nucleotide. The RT enzyme uses the nicked DNA as a primer for DNA synthesis from the gRNA, which is used as a template for the synthesis of a new DNA strand encoding the desired edit. The editing process shown may be referred to as targeted-primed reverse transcription editing (TRT editing), or "primed editing."
[0070] [Figure 1B] Figure 1B provides the same diagram as Figure 1A, with the exception that the prime editing factor complex is more commonly represented as [napDNAbp]-[P]:PEgRNA or [P]-[napDNAbp]:PEgRNA. In the formula, "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid programmed DNA-binding protein (e.g., SpCas9), "PEgRNA" refers to the prime editing guide RNA, and ]-[ refers to any linker. Elsewhere, as shown in Figures 3A-3G, for example, PEgRNA includes a 5' elongation arm containing the primer binding site and the DNA synthesis template. Although not shown, it is intended that the elongation arm of PEgRNA (i.e., this contains the primer binding site and the DNA synthesis template) can be DNA or RNA. The specific polymerase intended in this configuration will depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase may be RNA-dependent DNA polymerase (e.g., reverse transcriptase). When the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase.
[0071] [Figure 1C] Figure 1C provides a schematic diagram of an exemplary process for introducing a single nucleotide change and / or insertion and / or deletion into a DNA molecule (e.g., a genome), using a fusion protein containing reverse transcriptase fused to a Cas9 protein as a complex with an elongated guide RNA molecule. In this embodiment, the guide RNA is elongated at its 5' end to encompass the reverse transcriptase template sequence. The schematic diagram shows how reverse transcriptase (RT) fused to Cas9 nickase, as a complex with guide RNA (gRNA), binds to the DNA target site and nicks the PAM-containing DNA strand adjacent to the target nucleotide. The RT enzyme uses the nicked DNA as a primer for DNA synthesis from the gRNA, which is used as a template for the synthesis of a new DNA strand encoding the desired edit. The editing process shown may be called primed reverse transcription editing (TRT editing) or equivalently "primed editing".
[0072] [Figure 1D]Figure 1D provides the same depiction as Figure 1C, except that the prime editing factor complex is more commonly represented as [napDNAbp]-[P]:PEgRNA or [P]-[napDNAbp]:PEgRNAPEgRNA, where "P" refers to any polymerase (e.g., reverse transcriptase), "napDNAbp" refers to a nucleic acid programmed DNA-binding protein (e.g., SpCas9), "PEgRNA" refers to the prime editing guide RNA, and "[]-[" refers to any linker. Elsewhere, as shown in Figures 3A-3G for example, the PEgRNA includes a 3' elongation arm containing the primer binding site and the DNA synthesis template. Although not shown, it is intended that the elongation arm of the PEgRNA (i.e., containing the primer binding site and the DNA synthesis template) can be DNA or RNA. The specific polymerase intended in this configuration would depend on the nature of the DNA synthesis template. For example, if the DNA synthesis template is RNA, the polymerase can be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). If the DNA synthesis template is DNA, the polymerase can be a DNA-dependent DNA polymerase. In various embodiments, PEgRNA can be modified or synthesized to incorporate DNA-based DNA synthesis templates.
[0073] [Figure 1E] Figure 1E is a schematic diagram illustrating an illustrative process of how a synthesized single DNA strand (containing the desired nucleotide changes) is degraded so that the desired nucleotide changes are incorporated into the DNA. As shown, subsequent synthesis of the edited strand (or "mutagenic strand"), equilibrium with the endogenous strand, flap cleavage of the endogenous strand, and ligation lead to the incorporation of the DNA edit after the degradation of the mismatched DNA double helix through the action of the endogenous DNA repair and / or replication processes.
[0074] [Figure 1F]Figure 1F is a schematic diagram showing how incorporating "opposite strand nicking" into the degradation method shown in Figure 1E can help promote the formation of the desired product-pair-restore product. In opposite strand nicking, a second Cas9 / gRNA complex is used to introduce a second nick onto the opposite strand of the initially nicked strand. This induces the endogenetic DNA repair and / or replication process to preferentially replace the unedited strand (i.e., the strand containing the second nick site).
[0075] [Figure 1G]Figure 1G provides another schematic diagram of an exemplary process for introducing single-nucleotide alterations, and / or insertions, and / or deletions into a DNA molecule (e.g., a genome) at a target locus using a nucleic acid-programmed DNA-binding protein (napDNAbp) complexed with an elongated guide RNA. This process is sometimes referred to as prime editing. The elongated guide RNA involves elongation at the 3' or 5' end of the guide RNA, or at some location within the molecule within the guide RNA. In step (a), the napDNAbp / gRNA complex contacts the DNA molecule, and the gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced (e.g., by a nuclease or chemical agent) to one strand of the DNA at the target locus (R-loop strand, or PAM-containing strand, or non-target DNA strand, or protospacer strand), thereby creating a usable 3' end on one strand of the target locus. In one embodiment, the nick is created on the DNA strand corresponding to the R loop strand, i.e., the strand that does not hybridize with the guide RNA sequence. In step (c), the 3'-end DNA strand interacts with the extended portion of the guide RNA to prime the reverse transcription. In some embodiments, the 3'-end DNA strand hybridizes with a specific RT prime sequence on the extended portion of the guide RNA. In step (d), reverse transcriptase is introduced to synthesize a single strand of DNA from the 3' end of the primed site to the 3' end of the guide RNA. This forms a single-strand DNA flap containing the desired nucleotide change (e.g., a single base change, insertion, deletion, or a combination thereof). In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) relate to the degradation of the single-strand DNA flap so that the desired nucleotide change is incorporated into the target locus. This process can drive the formation of the desired product by removing the corresponding 5' endogenous DNA flap once the 3' single-strand DNA flap has penetrated and hybridized with a complementary sequence on the other strand. The process can also be driven to the formation of a product with a nicked second strand, as illustrated in Figure 1F.This process may introduce at least one or more of the following genetic changes: transversion, transition, deletion, and insertion.
[0076] [Figure 1H] Figure 1H is a schematic diagram illustrating the types of gene changes that can occur in the prime editing process as described herein. The types of nucleotide changes that can be achieved by prime editing include deletions (including short and long deletions), single nucleotide changes (including transitions and transversions), and insertions (including short and long ones).
[0077] [Figure 1I] Figure 1I is a schematic diagram illustrating a temporary nick to the second strand, exemplified by PE3b (PE3b = PE2 prime editing factor fusion protein + PEgRNA + guide RNA for nicking the second strand). Temporary nicking to the second strand is a variant of nicking to the second strand that facilitates the formation of the desired edited product. The term "temporary" refers to the fact that the second strand nick to the unedited strand occurs only after the desired edit has been incorporated into the edited strand. This avoids simultaneous nicking on both strands, which would lead to double-strand DNA breaks.
[0078] [Figure 1J]Figure 1J illustrates a variation of prime editing intended herein, in which a napDNAbp (e.g., SpCas9 nickase) is replaced with any programmed nuclease domain, such as a zinc finger nuclease (ZFN) or a transcription activator-like effector nuclease (TALEN). Therefore, the preferred nuclease does not necessarily need to be "programmed" by the nucleic acid target molecule (e.g., guide RNA), but rather may be programmed by defining the specificity of the DNA-binding domain, such as the nuclease, in particular. Just as with prime editing with the napDNAbp moiety, the alternative programmed nuclease is preferably modified to cleave only one strand of the target DNA. In other words, the programmed nuclease should preferably function as a nickase. Once a programmed nuclease is selected (e.g., a ZFN or TALEN), additional functions may be modified to enable it to operate according to a prime editing-like mechanism. For example, a programmed nuclease may be modified by coupling it with an RNA or DNA elongation arm (e.g., via a chemical linker), where the elongation arm includes a primer-binding site (PBS) and a DNA synthesis template. The programmed nuclease may also be coupled with a polymerase (e.g., via a chemical linker or an amino acid linker), the properties of which will depend on whether the elongation arm is DNA or RNA. In the case of an RNA elongation arm, the polymerase may be an RNA-dependent DNA polymerase (e.g., a reverse transcriptase). In the case of a DNA elongation arm, the polymerase may be a DNA-dependent DNA polymerase (e.g., a prokaryotic polymerase encompassing Pol I, Pol II, or Pol III, or a eukaryotic polymerase encompassing Pol a, Pol b, Pol g, Pol d, Pol e, or Pol z).The system may also include other functions that are added as a fusion with a programmed nuclease or added in trans to facilitate the overall reaction (e.g., (a) a helicase that unwinds the DNA at the cleavage site to create a cleavage strand with a 3' end that can be used as a primer, (b) a flap-end nuclease (e.g., FEN1) that removes the endogenous strand on the cleavage strand to help facilitate the reaction toward the replacement of the endogenous strand with the synthesized strand, or (c) an nCas9:gRNA complex that creates a second-site nick on the opposite strand (which may also help facilitate the incorporation of synthetic repair through favorable cellular repair of the unedited strand)). In a similar manner to priming editing with napDNAbp, such complexes with different programmed nucleases may be used to synthesize and then permanently incorporate the replacement strand of the newly synthesized DNA with the edit of interest into the target site of the DNA.
[0079] [Figure 1K]Figure 1K depicts, in one embodiment, the anatomical features of target DNA that may be edited by prime editing. The target DNA includes a “non-target strand” and a “target strand.” The target strand is the strand that will anneal with the PEgRNA spacer of the prime editing factor complex that recognizes the PAM site (in this case, NGG recognized by a standard SpCas9-based prime editing factor). The target strand may also be referred to as the “non-PAM strand” or “unedited strand.” In contrast, the non-target strand (i.e., the strand containing the protospacer and the PAM sequence of NGG) may also be referred to as the “PAM strand” or “edited strand.” In various embodiments, the nick site of the PE complex (e.g., in SpCas9-based PE) will be located in the protospacer on the PAM strand. The location of the nick will be a feature of the specific Cas9 that forms the PE. For example, in a SpCas9-based PE, the nick site is located in the phosphodiester bond between bases 3 (at position -3 relative to position 1 in the PAM sequence) and 4 (at position -4 relative to position 1 in the PAM sequence). The nick site in the protospacer forms a free 3' hydroxyl group that complexes with the primer binding site of the PEgRNA elongation arm, as shown in the figure below, providing a substrate to initiate the polymerization of a single strand of DNA encoding the DNA synthesis template of the PEgRNA elongation arm. This polymerization reaction is catalyzed in the 5'→3' direction by the polymerase of the PE fusion protein (e.g., reverse transcriptase). Polymerization is terminated before reaching the gRNA core (e.g., by a polymerization termination signal or inclusion of a secondary structure that functions to terminate the polymerization activity of the PE), producing a single-strand DNA flap extended from the original 3' hydroxyl group of the nicked PAM strand. The DNA synthesis template encodes a single strand of DNA homologous to the 5' end of the endogenous DNA immediately following the nicking site on the PAM strand, and incorporates the desired nucleotide changes (e.g., single base substitutions, insertions, deletions, inversions).The desired editing location can be any position downstream of the nick site on the PAM chain, such as positions +1, +2, +3, +4 (start of PAM site), +5 (PAM site position 2), +6 (PAM site position 3), +7, +8, +9, +10, +11, +12, +13, +14, +15, +16, +17, +18, +19, + 20, +21, +22, +23, +24, +25, +26, +27, +28, +29, +30, +31, +32, +33, +34, +35, +36, +37, +38, +39, +40, +41, +42, +43, +44, +45, +46, +47, +48, +49, +50, +51, +52, +53, +54, +55 +56, +57, +58, +59, +60, +61, +62, +63, +64, +65, +66, +67, +68, +69, +70, +71, +72, +73, +74, +75, +76, +77, +78, +79, +80, +81, +82, +83, +84, +85, +86, +87, +88, +89, +90, +91, +92, +93, +94, +95, +96, +97, +98, +99, +100, +101, +102, +103, +104, +105, +106, +107, +108, +109, +110, +111, +112, +113, +114, +115, +116, +117, +118, +119, +120, + +121, +122, +123, +124, +125, +126, +127, +128, +129, +130, +131, +132, +133, +134, +135, +136, +137, +138, +139, +140, +141, +142, +143, +144, +145, +146, +147, +148, +149, or +150, or more (relative to the downstream position of the nick site). Once the 3'-end single strand DNA (containing the edit of interest) replaces the endogenous 5'-end single strand DNA, the DNA repair and replication processes will result in the permanent incorporation of the edit site on the PAM strand, followed by the correction of mismatches on the non-PAM strand present at the edit site. Thus, the edit will spread to both strands of DNA on the target DNA site. It should be understood that the references to "edited strands" and "unedited" strands are simply intended to accurately describe (delineate) the DNA strands involved in the PE mechanism.The "edited strand" is the strand that becomes edited only after the 5' end of the single-stranded DNA immediately downstream of the nick site is replaced with a synthesized 3' end of single-stranded DNA containing the desired edit. The "unedited" strand is the paired strand with the edited strand, but it too becomes edited (in particular the edit of interest) through repair and / or replication to become complementary to the edited strand.
[0080] [Figure 1L]Figure 1L illustrates the mechanism of prime editing, showing the anatomical features of the target DNA, the prime editing factor complex, and the interaction between PEgRNA and the target DNA. Firstly, the prime editing factor, which includes a fusion protein containing a polymerase (e.g., reverse transcriptase) and napDNAbp (e.g., SpCas9 nickase, e.g., SpCas9 with an inactivating mutation in the HNH nuclease domain (e.g., H840A) or an activating mutation in the RuvC nuclease domain (D10A)), is complexed with PEgRNA and DNA containing the target DNA to be edited. PEgRNA includes a spacer, a gRNA core (also known as the gRNA backbone or gRNA main chain) (which binds to the napDNAbp), and an elongation arm. The elongation arm can be at the 3' end, the 5' end, or anywhere else within the PEgRNA molecule. As shown, the elongation arm is at the 3' end of PEgRNA. The elongation arm contains a primer-binding site and a DNA synthesis template (including both the edit of interest and the homologous region (i.e., the homologous arm)) that is homologous to the single-stranded DNA at the 5' end immediately following the nick site on the PAM strand in the 3'→5' direction. As shown, once a nick is introduced, which produces a free 3' hydroxyl group immediately upstream of the nick site, the region immediately upstream of the nick site on the PAM strand anneals with a complementary sequence at the 3' end of the elongation arm, referred to as the "primer-binding site," creating a short double-stranded region with an available 3' hydroxyl end, thereby forming a substrate for the polymerase of the prime editing factor complex. The polymerase (e.g., reverse transcriptase) then polymerizes the DNA strand from the 3' hydroxyl end to the end of the elongation arm. The sequence of the single-stranded DNA is encoded by the DNA synthesis template, which is the portion of the elongation arm (i.e., excluding the primer-binding site) that is "read" by the polymerase that synthesizes the new DNA. This polymerization effectively extends to the original 3' hydroxyl-terminus sequence of the initial nick site. The DNA synthesis template encodes a single strand of DNA that includes not only the desired edit but also a region homologous to the endogenous DNA single strand immediately downstream of the nick site on the PAM strand.Next, the single strand of DNA at the 3' end that is encoded (i.e., the 3' single-strand DNA flap) replaces the corresponding homologous single strand of endogenous 5' end DNA immediately downstream of the nick site on the PAM strand, forming a DNA intermediate with a 5' single-strand DNA flap, which is removed by the cell (e.g., by flap endonuclease). The 3' single-strand DNA flap, which anneals with the complement of the endogenous 5' single-strand DNA flap, is ligated with the endogenous strand after the 5' DNA flap has been removed. The desired edit in the 3' single-strand DNA flap, which has just been annealed and ligated, forms a mismatch with the complementary strand, and after DNA repair and / or a series of replications, the desired edit is permanently incorporated into both strands.
[0081] [Figure 2] Figure 2 shows three Cas complexes (SpCas9, SaCas9, and LbCas12a) that can be used in the prime editing factors described herein, and their PAM, gRNA, and DNA cleavage characteristics. The figure shows the design of the complexes involving SpCas9, SaCas9, and LbCas12a.
[0082] [Figure 3]Figures 3A–3F show designs of the manipulated 5' prime edit factor gRNA (Figure 3A), 3' prime edit factor gRNA (Figure 3B), and intramolecular extension (Figure 3C). The extended guide RNA (or extended gRNA) may also be referred to in this application as PEgRNA or "prime edit guide RNA." Figures 3D and 3E provide additional embodiments of the 3' and 5' prime edit factor gRNAs (PEgRNAs), respectively. Figure 3F illustrates the interaction between the 3'-end prime edit factor guide RNA and the target DNA sequence. The embodiments of Figures 3A–3C illustrate the arrangement of the reverse transcription template sequence (i.e., or more broadly, the DNA synthesis template, as indicated, since RT is the only type of polymerase that can be used in the context of prime edit factors), primer binding sites, and exemplary linker sequences in the extended portions of the 3', 5', and intramolecular versions, as well as the general arrangement of spacers and core regions. The disclosed prime editing process is not limited to these configurations of the extended guide RNA. An aspect of Figure 3D provides the structure of an exemplary PEgRNA intended in this application. The PEgRNA comprises three main component elements ordered in the 5' to 3' direction: a spacer, a gRNA core, and an extension arm at the 3' end. The extension arm may be further divided in the 5' to 3' direction into the following structural elements: a primer binding site (A), an editing template (B), and a homologous arm (C). In addition, the PEgRNA may contain an optional 3' end modification region (e1) and an optional 5' end modification region (e2). Furthermore, the PEgRNA may contain a transcription termination signal at the 3' end of the PEgRNA (not shown). These structural elements are further defined in this application. The illustration of the PEgRNA structure is not intended to be limiting and encompasses variations in the arrangement of elements. For example, the optional sequence modifications (e1) and (e2) may be located within or between any of the other regions shown. It is not limited to being located at the 3' and 5' ends.In some embodiments, PEgRNA may include secondary RNA structures such as, but not limited to, hairpins, stem-loops, toe-loops, and RNA-binding protein recruitment domains (e.g., MS2 aptamers that recruit and bind the MS2cp protein). For example, such secondary structures may be located within spacers, gRNA cores, or elongation arms, particularly within the e1 and / or e2 modification regions. In addition to secondary RNA structures, PEgRNA may include chemical linkers or poly(N) linkers or tails (e.g., within the e1 and / or e2 modification regions), where "N" can be any nucleic acid base. In some embodiments (e.g., as shown in Figure 72(c)), the chemical linkers may function to prevent reverse transcription of the sgRNA backbone or core. In addition, in certain embodiments (see, for example, Figure 72(c)), the elongation arms (3) may consist of RNA or DNA and / or contain one or more nucleic acid base analogs (e.g., this may add functionality such as temperature resilience). Furthermore, the orientation of the elongation arm (3) may be the natural 5' to 3' direction, or it may be synthesized in the opposite direction, 3' to 5' (relative to the orientation of the overall PEgRNA molecule). It should also be noted that those skilled in the art will have the ability to select a suitable DNA polymerase for use in prime editing, depending on the properties of the nucleic acid material of the elongation arm (i.e., DNA or RNA). This may be implemented either as a fusion with napDNAbp or as a separate part in trans, to synthesize a 3' single-stranded DNA flap encoded by the desired template encompassing the desired edit. For example, if the elongation arm is RNA, the DNA polymerase may be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the elongation arm is DNA, the DNA polymerase may be a DNA-dependent DNA polymerase. In various embodiments, the provision of DNA polymerase can be trans, for example, an RNA-protein recruitment domain (e.g., an MS2 hairpin incorporated on PEgRNA (e.g., in the e1 or e2 region or elsewhere) and an MS2cp protein fused to the DNA polymerase).This is achieved by using a primer binding site (which co-localizes the DNA polymerase to PEgRNA). It should also be noted that the primer binding site generally does not form part of the template used by the DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3' single-stranded DNA flap containing the desired edit. Therefore, the designation “DNA synthesis template” refers to a region or portion of the elongation arm (3) used as a template by the DNA polymerase to encode the desired 3' single-stranded DNA flap containing a homologous region to the 5' endogenous single-stranded DNA flap replaced by the 3' single-stranded DNA product of the edit and primed DNA synthesis. In some embodiments, the DNA synthesis template includes an “edit template” and a “homologous arm” or one or more homologous arms, for example, before and after the edit template. The edit template can be as small as a single-nucleotide substitution, or it may be a DNA insertion or inversion. In addition, the edit template may also include a deletion, which can be manipulated by encoding a homologous arm containing the desired deletion. In other embodiments, the DNA synthesis template may also include the e2 region or a portion thereof. For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, it is possible that DNA polymerase function will terminate before any portion of the e2 region is actually encoded on the DNA. It is also possible that some or even all of the e2 region will be encoded on the DNA. How much of the e2 is actually used as a template will depend on its composition and whether that composition disrupts DNA polymerase function.
[0083] [Figure 3E]The aspect of Figure 3E provides a structure of another PEgRNA intended in this application. The PEgRNA comprises three main component elements ordered in the 5' to 3' direction: a spacer, a gRNA core, and an elongation arm at the 3' end. The elongation arm may be further divided in the 5' to 3' direction into the following structural elements: a primer binding site (A), an editing template (B), and a homologous arm (C). In addition, the PEgRNA may include an optional 3' end modification region (e1) and an optional 5' end modification region (e2). Furthermore, the PEgRNA may include a transcription termination signal at the 3' end (not shown). These structural elements are further defined in this application. The illustration of the PEgRNA structure is not intended to be limiting and encompasses variations in the arrangement of elements. For example, any sequence modifications (e1) and (e2) may be located within or between any of the other regions shown, and are not limited to being located at the 3' and 5' ends. In some embodiments, PEgRNA may include secondary RNA structures such as, but not limited to, hairpins, stem-loops, toe-loops, and RNA-binding protein recruitment domains (e.g., MS2 aptamers that recruit and bind the MS2cp protein). These secondary structures can be located anywhere on the PEgRNA molecule. For example, such secondary structures may be located within spacers, the gRNA core, or the elongation arms, particularly within the e1 and / or e2 modification regions. In addition to secondary RNA structures, PEgRNA may include chemical linkers or poly(N) linkers or tails (e.g., within the e1 and / or e2 modification regions), where "N" can be any nucleic acid base. In some embodiments (e.g., as shown in Figure 72(c)), the chemical linkers may function to prevent reverse transcription of the sgRNA backbone or core. In addition, in certain embodiments (see, for example, Figure 72(c)), the elongation arm (3) may consist of RNA or DNA and / or contain one or more nucleic acid base analogs (for example, this may add functionality such as temperature resilience). Furthermore, the orientation of the elongation arm (3) may be the native 5' to 3' direction, or it may be synthesized in the opposite direction, 3' to 5' (relative to the orientation of the overall PEgRNA molecule).It will also be noted that those skilled in the art will have the ability to select a suitable DNA polymerase for use in prime editing, depending on the properties of the nucleic acid material of the elongation arm (i.e., DNA or RNA). This can be implemented either as a fusion with napDNAbp or as a separate part in trans, to synthesize a 3' single-stranded DNA flap encoded by a desired template encompassing the desired edit. For example, if the elongation arm is RNA, the DNA polymerase may be a reverse transcriptase or any other suitable RNA-dependent DNA polymerase. However, if the elongation arm is DNA, the DNA polymerase may be a DNA-dependent DNA polymerase. In various embodiments, the provision of the DNA polymerase may be in trans, for example, by the use of an RNA-protein recruitment domain (e.g., an MS2 hairpin incorporated on PEgRNA (e.g., in or elsewhere in the e1 or e2 region) and an MS2cp protein fused to the DNA polymerase, thereby colocalizing the DNA polymerase to PEgRNA). It should also be noted that primer binding sites generally do not form part of the template used by a DNA polymerase (e.g., reverse transcriptase) to encode the resulting 3' single-stranded DNA flap containing the desired edit. Therefore, the designation “DNA synthesis template” refers to a region or portion of the elongation arm (3) used as a template by a DNA polymerase to encode the desired 3' single-stranded DNA flap, which contains a homologous region to the 5' endogenous single-stranded DNA flap replaced by the edited and primed DNA synthesis product. In some embodiments, the DNA synthesis template includes an “edit template” and a “homologous arm” or one or more homologous arms, for example, before and after the edit template. The edit template can be as small as a single-nucleotide substitution, or it may be a DNA insertion or inversion. In addition, the edit template may also include a deletion, which can be manipulated by encoding a homologous arm containing the desired deletion. In other embodiments, the DNA synthesis template may also include the e2 region or a portion thereof.For example, if the e2 region contains a secondary structure that causes termination of DNA polymerase activity, it is possible that DNA polymerase function will terminate before any portion of the e2 region is actually encoded on the DNA. It is also possible that some or even all of the e2 region will be encoded on the DNA. How much of e2 is actually used as a template will depend on its composition and whether that composition disrupts DNA polymerase function.
[0084] [Figure 3F]The schematic diagram in Figure 3F illustrates the typical interaction of PEgRNA with a target site on double-stranded DNA and the associated generation of a 3' single-stranded DNA flap containing the desired gene alteration. The double-stranded DNA is shown by an upper strand oriented from 3' to 5' (i.e., the target strand) and a lower strand oriented from 5' to 3' (i.e., the PAM strand or non-target strand). The upper strand, containing the complements of the "protospacer" and the PAM sequence, is called the "target strand" because it is the strand targeted by the PEgRNA spacer and anneals to it. The complementary lower strand is called the "non-target strand," "PAM strand," or "protospacer strand" because it contains the PAM sequence (e.g., NGG) and the protospacer. Although not shown, the illustrated PEgRNA would complex with the Cas9 or equivalent domain of a prime editing factor fusion protein. As shown in the schematic diagram, the PEgRNA spacer anneals to the complementary region of the protospacer on the target strand. This interaction forms a DNA / RNA hybrid between the spacer RNA and the complement of the protospacer DNA, inducing the formation of an R-loop on the protospacer. As taught elsewhere in this application, the Cas9 protein (not shown) then induces a nick on the non-target strand as shown. This then leads to the formation of a 3' ssDNA flap region immediately upstream of the nick site, which interacts with the 3' end of the PEgRNA at the primer binding site according to *z*. The 3' end of the ssDNA flap (i.e., the reverse transcriptase primer sequence) anneals to the primer binding site (A) on the PEgRNA, thereby priming the reverse transcriptase. Next, the reverse transcriptase (for example, provided in trans or cis as a fusion protein attached to a Cas9 construct) polymerizes a single strand of DNA encoded by the DNA synthesis template (which includes the editing template (B) and homologous arm (C)). Polymerization continues toward the 5' end of the elongation arm.The polymerized strands of ssDNA form an ssDNA 3' end flap, which, as described elsewhere (for example, as shown in Figure 1G), invades endogenous DNA, displacing the corresponding endogenous strand (which is removed as a DNA flap at the 5' end of the endogenous DNA), and incorporates the desired nucleotide edits (single nucleotide base pair changes, deletions, and insertions (including entire genes)) by the naturally occurring DNA repair / replication rounds.
[0085] [Figure 3G]Figure 3G illustrates yet another embodiment of prime editing as envisioned in this application. In particular, the upper schematic diagram illustrates one embodiment of a prime editing factor (PE). This comprises a fusion protein of napDNAbp (e.g., SpCas9) and polymerase (e.g., reverse transcriptase), which are linked by a linker. The PE forms a complex with PEgRNA by binding to the gRNA core of PEgRNA. In the embodiment shown, the PEgRNA has a 3' extension arm, which, starting at the 3' end, contains a primer-binding site (PBS), followed by a DNA synthesis template. The lower schematic diagram illustrates a variant of the prime editing factor called a “transprime editing factor (tPE)”. In this embodiment, the DNA synthesis template and PBS are decoupled from the PEgRNA and presented by a separate molecule called a transprime editing factor RNA template ("tPERT"). This contains an RNA-protein recruitment domain (e.g., MS2 hairpin). PE itself is further modified to include a fusion with the rPERT recruiting protein ("RP"), which is a protein that specifically recognizes and binds to the RNA-protein recruiting domain. In the example where the RNA-protein recruiting domain is the MS2 hairpin, the corresponding rPERT recruiting protein could be the MS2cp of the MS2 tagging system. The MS2 tagging system is based on the innate interaction of the MS2 bacteriophage coat protein ("MCP" or "MS2cp") with stem-loop or hairpin structures present on the phage genome, i.e., "MS2 hairpins" or "MS2 aptamers". In the case of transprime editing, the RP-PE:gRNA complex "recruits" a tPERT having a suitable RNA-protein recruiting domain to colocalize with the PE:gRNA complex, thereby providing trans-PBS and DNA synthesis templates for use in prime editing, as shown in the example illustrated in Figure 3H.
[0086] [Figure 3H]Figure 3H illustrates the process of transprime editing. In this embodiment, the transprime editing factor comprises a “PE2” prime editing factor (i.e., a fusion of Cas9(H840A) and variant MMLV RT) fused to an MS2cp protein (i.e., a recruiting protein of the type that recognizes and binds to the MS2 aptamer) and complexed with sgRNA (i.e., a standard guide RNA in contrast to PEgRNA). The transprime editing factor binds to target DNA and introduces a nick into the non-target strand. The MS2cp protein recruits tPERT trans-prime by specific interaction with the RNA-protein recruiting domain on the tPERT molecule. tPERT then co-localizes with the transprime editing factor, thereby trans-prime providing PBS and DNA synthesis template functions for use by reverse transcriptase polymerase to synthesize a single-stranded DNA flap having a 3' end and containing the desired genetic information encoded by the DNA synthesis template.
[0087] [Figure 4A] Figures 4A-4E demonstrate the in vitro TPRT assay (i.e., prime editing assay). Figure 4A shows a schematic diagram of the fluorescently labeled DNA substrate, gRNA template extension by the RT enzyme, and PAGE. [Figure 4B-C] Figure 4B shows TPRT (i.e., prime editing) with pre-nicked substrates, dCas9, and 5' elongated gRNAs of different synthetic template lengths. Figure 4C shows the RT reaction with pre-nicked DNA substrates in the absence of Cas9. [Figure 4D-E] Figure 4D shows TPRT (i.e., prime editing) of Cas9(H840A) and 5' extended gRNA to a full-length dsDNA substrate. Figure 4E shows the 3' extended gRNA template with a pre-nicked full-length dsDNA substrate. M-MLV RT is present in all reactions.
[0088] [Figure 5]Figure 5 shows in vitro validation results using 5' elongated gRNA with varying synthetic template lengths. A fluorescently labeled (Cy5) DNA target was used as the substrate and pre-nicked in this experimental setup. The Cas9 used in these experiments was the catalytically inactive Cas9 (dCas9), and the RT used was Superscript III, a commercial RT derived from Moloney's mouse leukemia virus (M-MLV). dCas9:gRNA complexes were formed from purified components. The fluorescently labeled DNA substrate was then added along with dNTPs and the RT enzyme. After incubation at 37°C for 1 hour, the reaction product was analyzed by denatured urea-polyacrylamide gel electrophoresis (PAGE). The gel image shows elongation to a length consistent with the original DNA strand ~ reverse transcription template length.
[0089] [Figure 6] Figure 6 shows the in vitro validation results using 5' elongated gRNA with varying synthetic template lengths, which are very similar to those shown in Figure 5. However, the DNA substrate in this experiment did not contain pre-mixed nicks. The Cas9 used in these experiments was Cas9 nickase (SpyCas9 H840A mutant), and the RT used was Superscript III, a commercial RT derived from Moloney's mouse leukemia virus (M-MLV). The reaction products were analyzed by denatured urea polyacrylamide gel electrophoresis (PAGE). As shown in the gel, nickase efficiently cleaves DNA strands when standard gRNA is used (gRNA_0, lane 3).
[0090] [Figure 7]Figure 7 demonstrates that 3' extension does not significantly induce Cas9 nickase activity, supporting DNA synthesis. Pre-nicked substrates (black arrows) are converted to RT products almost quantitatively when either dCas9 or Cas9 nickase is used (lanes 4 and 5). More than 50% conversion to RT products (red arrow) is observed for full-length substrates (lane 3). Cas9 nickase (SpyCas9 H840A mutant), catalytically inactive Cas9 (dCas9), and Superscript III (commercial RT derived from Moloney's mouse leukemia virus (M-MLV)) are used.
[0091] [Figure 8] Figure 8 demonstrates a dual-color experiment used to determine whether the RT reaction preferentially results in cis (binding within the same complex) to the gRNA. Two separate experiments were performed for 5'-extended gRNA and 3'-extended gRNA. Products were analyzed by PAGE. Product ratios were calculated as (Cy3cis / Cy3trans) / (Cy5trans / Cy5cis).
[0092] [Figure 9A-B] Figures 9A–9D demonstrate the flap model substrate. Figure 9A shows a dual FP reporter for flap-specific (-directed) mutagenesis. Figure 9B shows arrest codon repair in HEK cells. [Figure 9C-D] Figure 9C shows sequenced yeast clones after flap repair. Figure 9D shows testing of different flap characteristics in human cells.
[0093] [Figure 10]Figure 10 demonstrates prime editing on a plasmid substrate. A dual fluorescent reporter plasmid was constructed for expression in yeast (S. cerevisiae). Expression of this construct in yeast produces only GFP. In vitro prime editing induces point mutations and transforms yeast with the parent plasmid or a plasmid nicked in vitro with Cas9(H840A). Colonies are visualized by fluorescence imaging. Yeast dual FP plasmid transformants are shown. Transdevelopment of the parent plasmid or a plasmid nicked in vitro with Cas9(H840A) results in only green GFP-expressing colonies. Prime editing with 5'-extended gRNA or 3'-extended gRNA produces mixed green and yellow colonies. The latter expresses both GFP and mCherry. More yellow colonies are observed with 3'-extended gRNA. Positive controls without stop codons are also not shown.
[0094] [Figure 11] Figure 11 shows prime editing on a plasmid substrate similar to the experiment in Figure 10, but instead of incorporating a point mutation within a stop codon, prime editing incorporates a single nucleotide insertion (left) or deletion (right) that repairs a frameshift mutation, enabling downstream mCherry synthesis. Both experiments used 3' elongated gRNAs.
[0095] [Figure 12] Figure 12 shows the edit products of prime editing on plasmid substrates, characterized by Sanger sequencing. Colonies from TRT transformation were individually selected and analyzed by Sanger sequencing. Precise editing was observed by sequencing the selected colonies. Green colonies contained plasmids with the original DNA sequence, while yellow colonies contained precise mutations designed by the prime editing gRNA. No other point mutations or indels were observed.
[0096] [Figure 13]Figure 13 illustrates the potential scope of the new prime editing technology and shows a comparison with deaminase-mediated base editing technology.
[0097] [Figure 14] Figure 14 shows a schematic diagram of editing in human cells.
[0098] [Figure 15] Figure 15 demonstrates the elongation of the primer binding site in gRNA.
[0099] [Figure 16] Figure 16 shows gRNA truncated for adjacent targeting.
[0100] [Figure 17A-C] Figures 17A-17C are graphs showing the %T→A conversion at target nucleotides after transfection of components in human embryonic kidney (HEK) cells. Figure 17A shows data presenting the results using N-terminal fusion of wild-type MLV reverse transcriptase to Cas9(H840A) nickase (32-amino acid linker). Figure 17B is similar to Figure 17A, except for the C-terminal fusion of the RT enzyme. Figure 17C is similar to Figure 17A, but the linker between MLV RT and Cas9 is 60 amino acids long instead of 32.
[0101] [Figure 18] Figure 18 shows high-purity T→A editing at the HEK3 site by high-throughput amplicon sequencing. The sequencing analysis output displays the most abundant genotype of the edited cells.
[0102] [Figure 19] Figure 19 shows the indel ratio (blue bars) alongside the editing efficiency at the target nucleotide (orange bars). WT refers to the wild-type MLV RT enzyme. Mutant enzymes (M1-M4) contain the mutations listed on the right. The editing ratio was quantified by high-throughput sequencing of genomic DNA amplicons.
[0103] [Figure 20] Figure 20 shows the editing efficiency of a target nucleotide when a single-strand nick is introduced into a complementary DNA strand adjacent to the target nucleotide. Experiments were conducted to introduce the nick at various intervals from the target nucleotide (triangles). Editing efficiency at the target base pair (blue bars) is shown alongside the indel formation rate (orange bars). The "none" example does not contain guide RNA for introducing the nick into the complementary strand. Editing rates were quantified by high-throughput sequencing of genomic DNA amplicons.
[0104] [Figure 21] Figure 21 demonstrates processed high-throughput sequencing data showing the overall absence of desired T→A transversion mutations and other major genome editing byproducts.
[0105] [Figure 22]Figure 22 provides a schematic diagram of an exemplary process for performing targeted mutagenesis on a target locus (i.e., prime editing with an error-prone reverse transcriptase) using a nucleic acid-programmed DNA-binding protein (napDNAbp) complexed with an extended guide RNA. This process is sometimes referred to as a form of prime editing for targeted mutagenesis. The extended guide RNA involves extension at the 3' or 5' end of the guide RNA, or at some intramolecular location within the guide RNA. In step (a), the napDNAbp / gRNA complex contacts a DNA molecule, and the gRNA guides the napDNAbp to bind to the target locus to be mutageneised. In step (b), a nick is introduced into one strand of the DNA at the target locus (e.g., by a nuclease or chemical agent), thereby creating a usable 3' end on one strand of the DNA at the target locus. In one embodiment, the nick is created in the strand of DNA corresponding to the R loop, i.e., the strand that is not hybridized with the guide RNA sequence. In step (c), the 3'-terminus DNA strand interacts with the elongation region of the guide RNA to prime the reverse transcription. In one embodiment, the 3'-terminus DNA strand hybridizes with a specific RT prime sequence on the elongation region of the guide RNA. In step (d), an error-prone reverse transcriptase is introduced, which synthesizes a single strand of mutagenic DNA from the 3'-terminus of the primed site to the 3'-terminus of the guide RNA. Exemplary mutations are indicated by an asterisk "*". This forms a single-strand DNA flap containing the desired mutagenic region. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) relate to the degradation of the single-strand DNA flap (containing the mutagenic region) so that the desired mutagenic region is incorporated into the target locus. This process can be driven to the formation of the desired product by removing the corresponding 5' endogenous DNA flap once the 3' single-strand DNA flap has entered and hybridized with a complementary sequence on the other strand. The process can also be driven to the formation of a product in which nicks are introduced into the second chain, as illustrated in Figure 1F.Following the endogenous DNA repair and / or replication process, the mutagenic region becomes integrated into both strands of DNA at the DNA locus.
[0106] [Figure 23] Figure 23 is a schematic diagram of trinucleotide repeat reduction in gRNA design and TPRT genome editing (i.e., prime editing) for reducing trinucleotide repeat sequences. Trinucleotide repeat expansion is associated with numerous human diseases, including Huntington's disease, fragile X syndrome, and Friedreich's ataxia. The most common trinucleotide repeat contains the CAG triplet, but GAA triplets (Friedreich's ataxia) and CGG triplets (Fragile X syndrome) also occur. Inheriting a predisposition to expansion or acquiring an already expanded parental allele increases the likelihood of developing the disease. The pathogenic expansion of trinucleotide repeats can hypothetically be corrected using prime editing. The region upstream of the repeat region can be nicked by an RNA-guided nuclease and then used to prime the synthesis of a new DNA strand containing a healthy number of repeats (which depends on the specific gene and disease). Following the repeat sequence, a series of short homologous sequences (red strands) are added that match the identity of the adjacent sequence at the other end of the repeat. The intrusion of the newly synthesized strand, followed by the replacement of the newly synthesized flap with endogenous DNA, leads to the reduced repeat allele.
[0107] [Figure 24] Figure 24 is a schematic diagram showing the precise 10-nucleotide deletion in prime editing. The guide RNA targeted by the HEK3 locus was designed with a reverse transcription template encoding a 10-nucleotide deletion after the nick site. Editing efficiency in transfected HEK cells was assessed using amplicon sequencing.
[0108] [Figure 25]Figure 25 is a schematic diagram illustrating gRNA design for peptide tagging of genes at endogenous genomic loci and peptide tagging in TPRT genome editing (i.e., prime editing). The FlAsH and ReAsH tagging systems contain a genetically encoded peptide comprising two parts: (1) a fluorophore-biarsenic probe and (2) a tetracysteine motif, exemplified by the sequence FLNCCPGCCMEP (SEQ ID NO: 1). When expressed in cells, the tetracysteine motif-containing protein can be fluorescently labeled with the fluorophore-biarsenic probe (see reference: J.Am.Chem.Soc., 2002, 124(21), pp6063-6076; DOI: 10.1021 / ja017687n). The "sortagging" system employs a bacterial saltase enzyme to covalently conjugate a labeled peptide probe to a protein containing a suitable peptide substrate (see reference: Nat.Chem.Biol.2007 Nov;3(11):707-8. DOI:10.1038 / nchembio.2007.31). FLAG tags (DYKDDDDK (SEQ ID NO: 2)), V5 tags (GKPIPNPLLGLDST (SEQ ID NO: 3)), GCN4 tags (EELLSKNYHLENEVARLKK (SEQ ID NO: 4)), HA tags (YPYDVPDYA (SEQ ID NO: 5)), and Myc tags (EQKLISEEDL (SEQ ID NO: 6)) are commonly used as epitope tags for immunoassays. The pi-clamp encodes a peptide sequence (FCPF) that can be labeled with a pentafluoro-aromatic substrate (Reference: Nat. Chem. 2016 Feb;8(2):120-8. doi:10.1038 / nchem.2413).
[0109] [Figure 26A]Figure 26A shows the precise insertion of His6 and FLAG tags into genomic DNA. Guide RNAs targeting the HEK3 locus were designed using reverse transcription templates encoding either an 18nt His tag insertion or a 24nt FLAG tag insertion. Editing efficiency in transfected HEK cells was assessed using amplicon sequencing. Note that the full-length 24nt FLAG tag sequence is cut off from the frame (sequencing confirmed the precise insertion of the full length). [Figure 26B] Figure 26B is a schematic diagram summarizing the various applications of protein / peptide tagging, which include (a) solubilizing or insolubilizing proteins, (b) altering or tracking the intracellular localization of proteins, (c) extending the half-life of proteins, (d) facilitating protein purification, and (e) facilitating protein detection.
[0110] [Figure 27] Figure 27 shows an overview of prime editing by incorporating protective mutations into PRNPs that prevent or halt the progression of prion diseases. The PEgRNA sequences correspond to SEQ ID NO: 351 (i.e., 5' of the sgRNA backbone) on the left and SEQ ID NO: 3864 (i.e., 3' of the sgRNA backbone) on the right.
[0111] [Figure 28A] Figure 28A is a schematic diagram of PE-based insertions of sequences encoding RNA motifs. [Figure 28B] Figure 28B lists (not exhaustive) some example motifs that could potentially be inserted, and their functions.
[0112] [Figure 29]Figure 29A illustrates the prime editing factor. Figure 29B shows possible modifications to genome, plasmid, or viral DNA guided by PE. Figure 29C shows an example scheme for inserting a peptide loop library onto a defined protein (GFP in this case) using a library of PEgRNAs. Figure 29D shows an example of possible programmable deletion or N- or C-terminal shortening of protein codons using different PEgRNAs. Deletions are expected to occur with minimal generation of frameshift mutations.
[0113] [Figure 30] Figure 30 shows possible schemes for repetitive codon insertion in continuous evolutionary systems such as PACE.
[0114] [Figure 31] Figure 31 shows a description of the manipulated gRNA. It shows the gRNA core, a ~20nt spacer that matches the sequence of the target gene, a reverse transcription template with an immunogenic epitope nucleotide sequence, and a primer binding site that matches the sequence of the target gene.
[0115] [Figure 32] Figure 32 is a schematic diagram illustrating the use of prime editing as a means of inserting known immunogenic epitopes into endogenous or exogenous genomic DNA, resulting in modification of the corresponding proteins.
[0116] [Figure 33]Figure 33 is a schematic diagram illustrating PEgRNA design for primer-binding sequence insertion and primer-binding insertion into genomic DNA using prime editing to determine off-target editing. In this embodiment, prime editing is performed in living cells, tissues, or animal models. As a first step, a suitable PEgRNA is designed. The upper schematic diagram shows an exemplary PEgRNA that may be used in this aspect. A spacer on the PEgRNA (labeled "protospacer") is complementary to one of the strands of the genomic target. The PE:PEgRNA complex (i.e., the PE complex) incorporates a single-stranded 3' end flap into the nick site, which contains the encoded primer-binding sequence and a homologous region (encoded by the homologous arm of the PEgRNA) that is complementary to the region just downstream of the cut site (indicated in red). Through flap insertion and DNA repair / replication processes, the synthesized strand is incorporated into the DNA, thereby incorporating the primer-binding site. This process can occur not only at the desired genomic target but also at other genomic sites that may interact with PEgRNA in an off-target manner (i.e., PEgRNA guides the PE complex to other off-target sites due to the complementarity of the spacer region to other genomic sites that are not the intended genomic site). Therefore, primer-binding sequences can be incorporated not only at the desired genomic target but also at other off-target genomic sites elsewhere on the genome. To detect the insertion of these primer-binding sequences at both the intended genomic target site and the off-target genomic site, genomic DNA (post-PE) can be isolated, fragmented, and ligated into adapter nucleotides (shown in red). Next, PCR can be performed using PCR oligonucleotides that anneal to the adapter and the inserted primer-binding sequence to amplify the on-target and off-target genomic DNA regions into which the primer-binding sequence was inserted by PE. Then, high-throughput sequencing and sequence alignment can be performed to identify the insertion site of the PE-inserted primer-binding sequence at either the on-target or off-target site.
[0117] [Figure 34] Figure 34 is a schematic diagram illustrating the precise insertion of genes by PE.
[0118] [Figure 35] Figure 35A is a schematic diagram showing the innate insulin signaling pathway. Figure 35B is a schematic diagram showing FKBP12-tagged insulin receptor activation controlled by FK1012.
[0119] [Figure 36] Figure 36 shows low molecular weight monomers. Reference: Bump FK506 Mimic(2)107.
[0120] [Figure 37] Figure 37 shows the low molecular weight dimer. References: FK1012 495,96; FK1012 5108; FK1012 6107; AP1903 7107; Cyclosporine A dimer 898; FK506-Cyclosporine A dimer (FkCsA) 9100.
[0121] [Figure 38]Figures 38A–38F provide an overview of prime editing and feasibility studies in vitro and in yeast cells. Figure 38A shows 75,122 known pathogenic human gene variants from ClinVar (accessed July 2019), classified by type. Figure 38B shows that the prime editing complex consists of a prime editing factor (PE) protein complexed with prime editing guide RNA (PEgRNA) and containing a DNA nicking domain, e.g., Cas9 nickase, which is guided by RNA fused to an engineered reverse transcriptase domain. The PE:PEgRNA complex binds to a target DNA site, enabling large and varied precision DNA editing at a wide range of DNA locations before and after the protospacer adjacent motif (PAM) of the target site. Figure 38C shows that DNA target binding causes the PE:PEgRNA complex to nick the PAM-containing DNA strand. The resulting free 3' end hybridizes to the primer-binding site of the PEgRNA. The reverse transcriptase domain catalyzes primer extension using the PEgRNA RT template, resulting in a newly synthesized DNA strand (3' flap) containing the desired edits. Equilibrium between the edited 3' flap and the unedited 5' flap containing the original DNA, followed by cellular 5' flap cleavage and ligation, and DNA repair or replication to degrade the heterodouble-stranded DNA, results in stably edited DNA. Figure 38D shows an in vitro 5'-extended PEgRNA primer extension assay using a pre-nicked dsDNA substrate containing a 5'Cy5-labeled PAM strand, dCas9, and a commercially available M-MLV RT variant (RT, Superscript III). dCas9 was complexed with PEgRNAs containing RT templates of various lengths, and then added to the DNA substrate along with the indicated components. The reaction was incubated at 37°C for 1 hour and then analyzed by denatured urea PAGE, visualized for Cy5 fluorescence.Figure 38E shows primer extension performed as shown in Figure 38D using 3'-extended PEgRNA pre-complexed with dCas9 or Cas9 H840A nickase and a pre-nicked or unnicked 5'Cy5-labeled dsDNA substrate. Figure 38F shows yeast colonies transformed with PEGRNA, Cas9 nickase, and a GFP-mCherry fusion reporter plasmid edited in vitro by RT. Plasmids containing nonsense or frameshift mutations between GFP and mCherry were edited with 5'-extended or 3'-extended PEgRNA that restored mCherry translation by transversion mutation, 1 bp insertion, or 1 bp deletion. Cells double-positive for GFP and mCherry (yellow) reflect successful editing.
[0122] [Figure 39]Figures 39A–39D show prime editing of human genomic DNA by PE1 and PE2. Figure 39A shows that PEgRNA contains a spacer sequence, an sgRNA backbone, and a 3' extension containing a reverse transcription (RT) template (purple) with a primer binding site (green) and the base(s) to be edited (singular or plural) (red). The primer binding site hybridizes to the PAM-containing DNA strand immediately upstream of the nicking site. With the exception of the encoded edit, the RT template is homologous to the DNA sequence downstream of the nicking. Figure 39B shows the incorporation of T·A to A·T transversion editing at the HEK3 site in HEK293T cells using Cas9 H840A nickase (PE1) fused to wild-type M-MLV reverse transcriptase and PEgRNAs of various primer binding site lengths. Figure 39C shows that the use of engineered quintuplet mutant M-MLV reverse transcriptases (D200N, L603W, T306K, W313F, T330P) to PE2 substantially improves prime editing transversion efficiency at five genomic sites in HEK293T cells and small insertion and small deletion editing in HEK3. Figure 39D compares PE2 editing efficiency at five genomic sites in HEK293T cells with various RT template lengths. Values and error bars reflect the mean and standard deviation of three independent biological replicas.
[0123] [Figure 40]Figures 40A–40C demonstrate that the PE3 and PE3b systems increase prime editing efficiency by nicking the unedited strand. Figure 40A provides an overview of prime editing by PE3. After the initial synthesis of the edited strand, DNA repair will remove either the newly synthesized strand containing the edit (3' flap excision) or the original genomic DNA strand (5' flap excision). The 5' flap excision leaves behind a DNA heteroduplex containing one edited strand and one unedited strand. Mismatch repair mechanisms or DNA replication may degrade the heteroduplex, yielding either an edited or unedited product. Nicking the unedited strand favors the repair of that strand, resulting in the preferred generation of a stable double-strand DNA containing the desired edit. Figure 40B shows the effect of complementary strand nicking on PE3-mediated prime editing efficiency and indel formation. "None" refers to the PE2 control, which does not nick the complementary strand. Figure 40C compares editing efficiency with PE2 (no complementary strand nick), PE3 (general complementary strand nick), and PE3b (edit-specific complementary strand nick). All edit yields reflect the percentage of total sequencing reads containing the intended edit and free of indels among all treated cells without sorting. Values and error bars reflect the mean and standard deviation of three independent biological replicas.
[0124] [Figure 41]Figures 41A–41K show PE3-targeted insertions, deletions, and all 12 types of point mutations at seven endogenous human genome loci in HEK293T cells. Figure 41A is a graph showing all 12 types of nucleotide transitions and transversion edits at HEK3 sites from position +1 to +8 (counting PEgRNA-induced nick positioning as between position +1 and -1) using a 10nt RT template. Figure 41B is a graph showing long-range PE3 transversion edits at HEK3 sites using a 34nt RT template. Figures 41C–41H are graphs showing all 12 types of transitions and transversion edits at various positions on the prime editing window for (Figure 41C) RNF2, (Figure 41D) FANCF, (Figure 41E) EMX1, (Figure 41F) RUNX1, (Figure 41G) VEGFA, and (Figure 41H) DNMT1. Figure 41I is a graph showing targeted 1 and 3 bp insertions and 1 and 3 bp deletions by PE3 at seven endogenous genomic loci. Figure 41J is a graph showing targeted precise deletions from 5 to 80 bp at HEK3 target sites. Figure 41K is a graph showing combined edits of insertions and deletions, insertions and point mutations, deletions and point mutations, and double point mutations at three endogenous genomic loci. All edit yields reflect the percentage of total sequencing reads that contain the intended edits and do not contain indels, out of all treated cells without sorting. Values and error bars reflect the mean and standard deviation of three independent biological replicas.
[0125] [Figure 42]Figures 42A–42H show a comparison of prime editing, base editing, and off-target editing by Cas9 and PE3 at known Cas9 off-target sites. Figure 42A shows the total C·G to T·A editing efficiency at the same target nucleotide for PE2, PE3, BE2max, and BE4max in endogenous HEK3, FANCF, and EMX1 sites of HEK293 T cells. Figure 42B shows the indel frequency from the treatment in Figure 42A. Figure 42C shows the editing efficiency of precise C·G to T·A editing (bystander editing or no indels) for PE2, PE3, BE2max, and BE4max in HEK3, FANCF, and EMX1. In EMX1, precise PE combination editing of all possible combinations of C·G to T·A conversion at three target nucleotides is also shown. Figure 42D shows the total A·T to G·C editing efficiency for PE2, PE3, ABEdmax, and ABEmax in HEK3 and FANCF. Figure 42E shows the precise A·T to G·C editing efficiency without bystander editing or indels for HEK3 and FANCF. Figure 42F shows the indel frequency from the treatment in Figure 42D. Figure 42G shows the mean triplicate editing efficiency (percentage sequencing reads with indels) in HEK293T cells for Cas9 nuclease at four on-target and 16 known off-target sites. The 16 off-target sites examined were the top four previously reported off-target sites for each of the four on-target sites: 118,159. For each on-target site, Cas9 was paired with either an sgRNA or each of four PEgRNAs that recognize the same protospacer. Figure 42H shows the average triplicate on-target and off-target editing efficiencies and indel efficiencies (in parentheses below) for PE2 or PE3 paired with each PEgRNA (Figure 42G) in HEK293T cells. On-target editing yield reflects the percentage of total sequencing reads that contain the intended edit but do not contain indels, out of all treated cells without sorting.Off-target editing yield reflects off-target locus modifications consistent with prime editing. Values and error bars reflect the mean and standard deviation of three independent biological replicas.
[0126] [Figure 43]Figures 43A–43I show prime editing, pathogenic transversion, insertion or deletion mutation incorporation and modification, as well as a comparison of prime editing and HDR in various human cell lines and primary mouse cortical neurons. Figure 43A is a graph showing incorporation (by T·A to A·T transversion) and modification (by A·T to T·A transversion) of the pathogenic E6V mutation in the HBB of HEK293T cells. Modification to either wild-type HBB or HBB containing a silent mutation that blocks PEgRNA PAM is shown. Figure 43B is a graph showing incorporation (by 4bp insertion) and modification (by 4bp deletion) of the pathogenic HEXA 1278+TATC allele in HEK293T cells. Modification to either wild-type HEXA or HEXA containing a silent mutation that blocks PEgRNA PAM is shown. Figure 43C is a graph showing the incorporation of the protective G127V variant of PRNP in HEK293T cells via G·C to T·A transversion. Figure 43D is a graph showing prime editing in other human cell lines, including K562 (leukemia myeloid cells), U2OS (osteosarcoma cells), and HeLa (cervical cancer cells). Figure 43E is a graph showing the incorporation of the G·C to T·A transversion mutation in DNMT1 of mouse primary cortical neurons, using a binary mitotic intein PE3 lentiviral system. In this system, the N-terminal half is Cas9(1-573) fused to the N-intine and GFP-KASH via a P2A autocleavage peptide, and the C-terminal half is the C-intine fused to the remainder of PE2. The PE2 halves are expressed from a human synapsin promoter that is highly specific to mature neurons. Sorted values reflect edits or indels from GFP-positive nuclei, while unsorted values are from all nuclei. Figure 43F shows a comparison of PE3 and Cas9-mediated HDR editing efficiencies at endogenous genomic loci in HEK293T cells. Figure 43G shows a comparison of PE3 and Cas9-mediated HDR editing efficiencies at endogenous genomic loci in K562, U2OS, and HeLa cells.Figure 43H compares PE3 and Cas9-mediated HDR indel byproduct generation in HEK293T, K562, U2OS, and HeLa cells. Figure 43I shows targeted insertions of His6 tag (18 bp), FLAG epitope tag (24 bp), or extended LoxP site (44 bp) in HEK293T cells by PE3. All edit yields reflect the percentage of total sequencing reads out of all treated cells that contain the intended edit but not the indel. Values and error bars reflect the mean and standard deviation of three independent biological replicates.
[0127] [Figure 44]Figures 44A–44G show in vitro prime editing validation studies using fluorescently labeled DNA substrates. Figure 44A shows electrophoretic mobility shift assays with dCas9, 5'-extended PEgRNA, and 5'-Cy5 labeled DNA substrate. PEgRNAs 1–5 contain a 15nt linker sequence (linker A in PEgRNA 1, linker B in PEgRNAs 2–5), a 5nt PBS sequence, and RT templates of 7nt (PEgRNAs 1 and 2), 8nt (PEgRNA 3), 15nt (PEgRNA 4), and 22nt (PEgRNA 5) between the spacer and PBS. The PEgRNAs used are those shown in Figures 44E and 44F; complete sequences are listed in Tables 2A–2C. Figure 44B shows an in vitro nicking assay of Cas9 H840A using 5'-extended and 3'-extended PEgRNAs. Figure 44C shows Cas9-mediated indel formation in HEK293T cells in HEK3 using 5'-extended and 3'-extended PEgRNA. Figure 44D shows an overview of the prime editing in vitro biochemical assay. 5'-Cy5 labeled pre-nicked and unnicked dsDNA substrates were tested. sgRNA, 5'-extended PEgRNA, or 3'-extended PEgRNA were pre-complexed with dCas9 or Cas9 H840A nickase and then combined with dsDNA substrate, M-MLV RT, and dNTPs. The reaction was allowed to proceed for 1 hour at 37°C prior to separation by denatured urea PAGE and visualization by Cy5 fluorescence. Figure 44E shows that the primer extension reaction using 5'-extended PEgRNA, pre-nicked DNA substrate, and dCas9 leads to significant conversion to the RT product. Figure 44F shows the primer extension reaction using an unnicked DNA substrate and Cas9 H840A nickase, and 5'-extended PEgRNA as shown in Figure 44B. Product yield is greatly reduced compared to a pre-nicked substrate. Figure 44G shows, by denatured urea PAGE, that the in vitro primer extension reaction using 3'-PEgRNA produces a single, clear product.The RT product band was excised, eluted from the gel, and then subjected to homopolymer tailing with terminal transferase (TdT) using either dGTP or dATP. The tailed product was extended with poly-T or poly-C primers, and the resulting DNA was sequenced. Sanger traces indicate that three nucleotides derived from the gRNA backbone were reverse transcribed (added to the DNA product as the last 3' nucleotide). Note that PEgRNA backbone insertions are considerably rarer in mammalian cell prime editing experiments than in vitro (Figures 56A-56D). Possible causes include the inability of the tethered reverse transcriptase to access the Cas9-bound guide RNA backbone, and / or cellular excision of the 3' end of a mismatch in the 3' flap containing the PEgRNA backbone sequence.
[0128] [Figure 45]Figures 45A–45G show cellular repair of 3' DNA flaps in yeast from in vitro prime-edit reactions. Figure 45A shows that a binary fluorescent protein reporter plasmid contains GFP and mCherry open reading frames separated by target sites encoding in-frame stop codons, +1 frameshifts, or -1 frameshifts. Prime-edit reactions were performed in vitro with Cas9 H840A nickase, PEgRNA, dNTPs, and M-MLV reverse transcriptase, and then the cells were transformed into yeast. Colonies containing the unedited plasmid produce GFP but not mCherry. Yeast colonies containing the edited plasmid produce both GFP and mCherry as fusion proteins. Figure 45B shows superposition of GFP and mCherry fluorescence in yeast colonies transformed with reporter plasmids containing a stop codon between GFP and mCherry (unedited negative control; top) or without a stop codon or frameshift between GFP and mCherry (pre-edited positive control; bottom). Figures 45C–45F show visualizations of mCherry and GFP fluorescence from yeast colonies transformed by in vitro prime editing reaction products. Figure 45C shows stop codon correction by T·A to A·T transversion using 3'-extended or 5'-extended PEgRNA, as shown in Figure 45D. Figure 45E shows +1 frameshift correction of 1 bp deletion using 3'-extended PEgRNA. Figure 45F shows -1 frameshift correction by 1 bp insertion using 3'-extended PEgRNA. Figure 45G shows Sanger DNA sequencing traces from plasmids isolated from GFP-only colonies in Figure 45B and GFP and mCherry double-positive colonies in Figure 45C.
[0129] [Figure 46]Figures 46A-46F show correct editing versus indel generation by PE1. Figure 46A shows the efficiency of T·A to A·T transversion editing and indel generation by PE1 at the +1 position of HEK3, using PEgRNA containing a 10nt RT template and a PBS sequence in the range of 8-17nt. Figure 46B shows the efficiency of G·C to T·A transversion editing and indel generation by PE1 at the +5 position of EMX1, using PEgRNA containing a 13nt RT template and a PBS sequence in the range of 9-17nt. Figure 46C shows the efficiency of G·C to T·A transversion editing and indel generation by PE1 at the +5 position of FANCF, using PEgRNA containing a 17nt RT template and a PBS sequence in the range of 8-17nt. Figure 46D shows the efficiency of C·G to A·T transversion editing and indel generation by PE1 at the +1 position of RNF2, using PEgRNA containing an 11nt RT template and a PBS sequence in the range of 9–17nt. Figure 46E shows the efficiency of G·C to T·A transversion editing and indel generation by PE1 at the +2 position of HEK4, using PEgRNA containing a 13nt RT template and a PBS sequence in the range of 7–15nt. Figure 46F shows PE1-mediated +1T deletion, +1A insertion, and +1CTT insertion at the HEK3 site, using 13nt PBS and 10nt RT templates. The PEgRNA sequences used are those used in Figure 39C (see Tables 3A–3R). Values and error bars reflect the mean and sd of three independent biological replicas.
[0130] [Figure 47]Figures 47A-47S show the evaluation of M-MLV RT variants for prime editing. Figure 47A shows the abbreviations for the prime editing factor variants used in this figure. Figure 47B shows targeted insertion and deletion editing by PE1 at the HEK3 locus. Figures 47C–47H show a comparison of 18 prime editing factor constructs containing the M-MLV RT variant in their ability to incorporate the following edits: +2G·C to C·G in HEK3 (as shown in Figure 47C), 24bp FLAG insertion in HEK3 (as shown in Figure 47D), +1C·G to A·T in RNF2 (as shown in Figure 47E), +1G·C to C·G in EMX1 (as shown in Figure 47F), +2T·A to A·T in HBB (as shown in Figure 47G), and +1G·C to C·G in FANCF (as shown in Figure 47H). Figures 47I–47N show a comparison of four prime editing factor constructs containing the M-MLV variant in their ability to incorporate the edits shown in Figures 47C–47H in a second round of independent experiments. Figures 47O–47S show PE2 editing efficiency at five genomic loci with varying PBS lengths. Figure 47O shows the +1T·A to A·T variation in HEK3. Figure 47P shows the +5G·C to T·A variation in EMX1. Figure 47Q shows the +5G·C to T·A variation in FANCF. Figure 47R shows the +1C·G to A·T variation in RNF2. Figure 47S shows the +2G·C to T·A variation in HEK4. Values and error bars reflect the mean and standard deviation of three independent biological replicas.
[0131] [Figure 48]Figures 48A–48C illustrate the design features of PEgRNA PBS and RT template sequences. Figure 48A shows the efficiency of PE2-mediated +5G·C to T·A transversion editing in VEGFA of HEK293T cells as a function of RT template length (blue line). Indels (gray line) are plotted for comparison. The sequences below the graph show the last nucleotide for editing that serves as a template for synthesis by PEgRNA. G nucleotides (templated at C on PEgRNA) are highlighted; to maximize prime editing efficiency, RT templates ending in C should be avoided during PEgRNA design. Figure 48B shows the +5G·C to T·A transversion editing and indels of DNMT1 as in Figure 48A. Figure 48C shows the +5G·C to T·A transversion editing and indels of RUNX1 as in Figure 48A. Values and error bars reflect the mean and sd of three independent biological replicas.
[0132] [Figure 49]Figures 49A–49B show the effects of PE2, PE2 R110S K103L, Cas9 H840A nickase, and dCas9 on cell viability. HEK293T cells were transfected with plasmids encoding PE2, PE2 R110S K103L, Cas9 H840A nickase, or dCas9, along with a HEK3-targeted PEgRNA plasmid. Cell viability was measured every 24 hours and over 3 days post-transfection using the CellTiter-Glo2.0 assay (Promega). Figure 49A shows viability measured by luminescence at 1, 2, or 3 days post-transfection. Values and error bars reflect the mean and sem of three independent biological replicas performed using technical triplicates, respectively. Figure 49B shows the percentage edits and indels for PE2, PE2 R110S K103L, Cas9 H840A nickase, or dCas9, along with the HEK3-targeted PEgRNA plasmid encoding editing from +5G to A. Editing efficiency was measured from treated cells on day 3 post-transfection, alongside those used to assay viability in Figure 49A. Values and error bars reflect the mean and standard deviation of three independent biological replicas.
[0133] [Figure 50]Figures 50A–50B show PE3-mediated HBB E6V and HEXA 1278+TATC modification by various PEgRNAs. Figure 50A shows a screen of 14 PEgRNAs for modification of the HBB E6V allele in HEK293T cells by PE3. All PEgRNAs evaluated convert the HBB E6V allele back to wild-type HBB without introducing any silent PAM mutations. Figure 50B shows a screen of 41 PEgRNAs for modification of the HEXA 1278+TATC allele in HEK293T cells by PE3 or PE3b. HEXA-labeled PEgRNAs modify the pathogenic allele by a shifted 4bp deletion that blocks PAM and leaves a silent mutation. HEXA-labeled PEgRNAs modify the pathogenic allele back to wild-type. Entries ending in "b" use an editing-specific nicking sgRNA in combination with PEgRNA (PE3b system). The values and error bars reflect the mean and standard deviation of three independent biological replicas.
[0134] [Figure 51]Figures 51A–51F show PE3 activity and a comparison of PE3 and Cas9-initiated HDRs in human cell lines. Figures 51A (HEK293T cells), 51B (K562 cells), 51C (U2OS cells), and 51D (HeLa cells) show the efficiency of generating correct editing (no indels) and indel frequencies for PE3 and Cas9-initiated HDRs. Each grouped editing comparison incorporates the same editing by PE3 and Cas9-initiated HDRs. The untargeted controls are PE3 and PEgRNAs that target non-target loci. Figure 51E shows control experiments with untargeted PEgRNA + PE3 and dCas9 + sgRNA compared to wild-type Cas9 HDR experiments. We have confirmed that ssDNA donor HDR templates containing common contaminants, which artificially increase apparent HDR efficiency, do not contribute to the HDR measurements in Figures 51A–51D. Figure 51F shows an example HEK3 site allele table from genomic DNA samples isolated from K562 cells after editing with HDR initiated by PE3 or Cas9. Alleles were sequenced by Illumina MiSeq and analyzed by CRISPResso2.178 The reference HEK3 sequence from this region is shown above. The allele table is shown for a negative control of untargeted PEgRNA, +1CTT insertions in HEK3 using PE3, and +1CTT insertions in HEK3 using HDR initiated by Cas9. Allele frequencies and corresponding Illumina sequencing read counts are shown for each allele. All alleles observed with frequencies ≥0.20% are shown. Values and error bars reflect the mean and sd of three independent biological replicas.
[0135] [Figure 52]Figures 52A–52D show the distribution of pathogenic insertions, duplications, deletions, and indel lengths in the ClinVar database. The ClinVar variant summary was downloaded from NCBI on July 15, 2019. The lengths of reported insertions, deletions, and duplications were calculated using appropriate identifying information such as reference and alternative alleles, variant start and stop positions, or variant names. Variants that did not report any of the above information were excluded from the analysis. The length of reported indels (single variants containing both insertions and deletions relative to the reference genome) was calculated by determining the number of mismatches or gaps in the best pairwise alignment between the reference and alternative alleles.
[0136] [Figure 53]Figures 53A - 53B show FACS gating examples for GFP - positive cell sorting. The bottom is an example of the original batch analysis file, outlining the sorting strategy used to generate the HEXA 1278+TATC and HBB E6V HEK293T cell lines. Image data was generated by a Sony LE - MA900 cytometer using Cell Sorter software v.3.0.5. Graphic 1 shows the gating plot for cells that do not express GFP. Graphic 2 shows the sorting of an example of P2A - GFP - expressing cells used to isolate the HBB E6V HEK293T cell line. HEK293T cells were initially gated as a population using FSC - A / BSC - A (gate A), and then sorted for singlets using FSC - A / FSC - H (gate B). Live cells were sorted by gating DAPI - negative cells (gate C). Cells with GFP fluorescence levels above those of the negative control cells were sorted using EGFP as the fluorochrome (gate D). Figure 53A shows HEK293T cells (GFP - negative). Figure 53B shows a representative plot of FACS gating for cells expressing PE2 - P2A - GFP. Figure 53C shows the genotype of HEXA 1278+TATC homozygous HEK293T cells. Figure 53D shows the allele table of the HBB E6V homozygous HEK293T cell line.
[0137] [Figure 54] Figure 54 is a schematic diagram summarizing the PEgRNA cloning procedure.
[0138] [Figure 55]Figures 55A - 55G are schematic diagrams of PEgRNA design. Figure 55A shows a simple diagram of a PEgRNA with domains labeled (left) and bound to nCas9 at the genomic site (right). Figure 55B shows various types of modifications of PEgRNA that are expected to increase activity. Figure 55C shows modifications of PEgRNA for increasing transcription of longer RNAs by promoter selection and 5', 3' processing and termination. Figure 55D shows elongation of the P1 system. This is an example of backbone modification. Figure 55E shows that incorporation of synthetic modifications on or elsewhere on the PEgRNA can increase activity. Figure 55F shows that designed incorporation of minimal secondary structure on the template can prevent formation of longer, more inhibitory secondary structures. Figure 55G shows a split PEgRNA with a second template sequence anchored by an RNA element at the 3' end of the PEgRNA (left). Incorporation of elements at the 5' or 3' end of the PEgRNA can enhance RT binding.
[0139] [Figure 56] Figures 56A - 56D show incorporation of the PEgRNA backbone sequence into the target locus. The PEgRNA backbone sequence insertions were analyzed as described for the HTS data in Figures 60A - 60B. Figure 56A shows the analysis of the EMX1 locus. Percent of sequencing reads containing one or more PEgRNA backbone sequence nucleotides on the insertion adjacent to the RT template (left); percentage of sequencing reads containing a PEgRNA backbone sequence insertion of a defined length (center); and cumulative total percentage of PEgRNA insertions encompassing up to the length defined on the X - axis are shown. Figure 56B shows the same for FANCF as in Figure 56A. Figure 56C shows the same for HEK3 as in Figure 56A. Figure 56D shows the same for RNF2 as in Figure 56A. Values and error bars reflect the mean and s.d. of three independent biological replicates.
[0140] [Figure 57]Figures 57A–57I show the effects of PE2, PE2-dRT, and Cas9 H840A nickase on transcriptome-wide RNA abundance. Analysis of ribosomal RNA-depleted cellular RNA isolated from HEK293T cells expressing PE2, PE2-dRT, or Cas9 H840A nickase and PRNP-targeted or HEXA-targeted PEgRNA. RNA corresponding to 14,410 genes and 14,368 genes were detected in PRNP and HEXA samples, respectively. Figures 57A–57F show volcano plots presenting the ~-fold change in -log10 FDR-adjusted p-value versus log2 transcript abundance for each (Aeach) RNA. (Figure 57A) PE2 vs. PE2-dRT using PRNP-targeted PEgRNA, (Figure 57B) PE2 vs. Cas9 H840A using PRNP-targeted PEgRNA, (Figure 57C) PE2-dRT vs. Cas9 H840A using PRNP-targeted PEgRNA, (Figure 57D) PE2 vs. PE2-dRT using HEXA-targeted PEgRNA, (Figure 57E) PE2 vs. Cas9 H840A using HEXA-targeted PEgRNA, and (Figure 57F) PE2-dRT vs. Cas9 H840A using HEXA-targeted PEgRNA are compared. Red dots indicate genes showing a statistically significant change in relative abundance of ≥2 times (FDR-adjusted p<0.05). Figures 57G-57I are Venn diagrams of upwardly controlled and downwardly controlled transcripts (≥2x change), comparing PRNP and HEXA samples for (Figure 57G) PE2 vs. PE2-dRT, (Figure 57H) PE2 vs. Cas9 H840A, and (Figure 57I) PE2-dRT vs. Cas9 H840A.
[0141] [Figure 58] Figure 58 shows a typical FACS gating for neuronal nucleus sorting. The nuclei were sequentially gated based on the DyeCycle Ruby signal, FSC / SSC ratio, SSC width / SSC height ratio, and GFP / DyeCycle ratio.
[0142] [Figure 59]Figures 59A–59F show the protocol for cloning 3'-extended PEgRNA onto a mammalian U6 expression vector by Golden Gate assembly. Figure 59A shows an overview of the cloning process. Figure 59B shows "Step 1: Digestion of the pU6-PEgRNA-GG-vector plasmid (component 1)". Figure 59C shows "Steps 2 and 3: Ordering and annealing of oligonucleotide parts (components 2, 3, and 4)". Figure 59D shows "Step 2b.ii.: Phosphorylation of the sgRNA backbone (unnecessary if the oligonucleotide is purchased phosphorylated)". Figure 59E shows "Step 4: PEgRNA assembly". Figure 59F shows "Steps 5 and 6: Transformation of the assembled plasmid". Figure 59G shows a diagram summarizing the PEgRNA cloning protocol.
[0143] [Figure 60] Figures 60A–60B show a Python script for quantifying PEgRNA backbone incorporation. A custom Python script was generated to characterize and quantify PEgRNA insertions at target genomic loci. The script iteratively matches an increasing length text string taken from the reference sequence (guide RNA backbone sequence) with sequence readings in a fastq file and counts the number of sequence readings that match the search query. Each sequential text string corresponds to an additional nucleotide in the guide RNA backbone sequence. Strict length incorporation and cumulative incorporation of up to a specified length were calculated in this manner. To ensure alignment and accurate counting of short slices of sgRNA, the reference sequence starts with 5–6 bases from the 3' end of a new DNA strand synthesized by reverse transcriptase.
[0144] [Figure 61] Figure 61 is a graph showing the percentage of total sequencing reads with specified editing for SaCas9(N580A)-MMLV RT HEK3 +6C>A. The correct editing and indel values are shown.
[0145] [Figure 62] Figures 62A–62B illustrate the importance of protospacers for the efficient incorporation of desired edits in precise positioning by prime editing. Figure 62A is a graph showing the percentage of total sequencing reads in which target T·A base pairs were converted to A·T for various HEK3 loci. Figure 62B shows the sequence analysis illustrating this.
[0146] [Figure 63] Figure 63 is a graph showing SpCas9 PAM variants with PAM editing (N=3). The percentage of total sequencing reads with targeted PAM editing is shown for SpCas9(H840A)-VRQR-MMLV RTs where NGA>NTA and for SpCas9(H840A)-VRER-MMLV RTs where NGCG>NTCG. The PEgRNA primer binding site (PBS) length, RT template (RT) length, and PE system used are listed.
[0147] [Figure 64] Figure 64 is a schematic diagram illustrating the introduction of various site-specific recombinase (SSR) targets into the genome using PE. (a) provides a general schematic diagram of the insertion of a recombinase target sequence by a prime editing factor. (b) shows how a single SSR target inserted by PE can be used as a site for genomic incorporation of a DNA donor template. (c) shows how tandem insertion of SSR target sites can be used to delete a portion of the genome. (d) shows how tandem insertion of SSR target sites can be used to invert a portion of the genome. (e) shows how insertion of two SSR target sites in two distal chromosomal regions can result in a chromosomal translocation. (f) shows how insertion of two different SSR target sites on the genome can be used to exchange a cassette from a DNA donor template. See Example 17 for further details.
[0148] [Figure 65] Figure 65 illustrates 1) PE-mediated synthesis of an SSR target site on the human cell genome and 2) the use of that SSR target site to incorporate a DNA donor template containing a GFP expression marker. Once successfully incorporated, GFP causes the cell to fluoresce. See Example 17 for further details.
[0149] [Figure 66] Figure 66 illustrates one embodiment of a prime editing factor provided as two halves of a prime editing factor protein. These regenerate the whole prime editing factor through self-splicing of the halves of the fission intein located at the end or beginning of each half of the prime editing factor protein.
[0150] [Figure 67]Figure 67 illustrates the mechanism of intein removal and peptide bond reformation from a polypeptide sequence between the N-terminal and C-terminal extein sequences. (a) illustrates the general mechanism of two half-proteins each containing half of the intein sequence. When these contact within a cell, they result in a fully functional intein. This then undergoes self-splicing and excision. The excision process results in the formation of a peptide bond between the N-terminal half of the protein (or “N-extein”) and the C-terminal half of the protein (or “C-extein”), forming an entire single polypeptide containing the N-extein and C-extein portions. In various embodiments, the N-extein may correspond to the N-terminal half of a split prime editing factor fusion protein, and the C-extein may correspond to the C-terminal half of a split prime editing factor. (b) shows the chemical mechanism of intein excision and reformation of the peptide bond linking the half of the N-extein (the red half) and the half of the C-extein (the blue half). Since it involves the splicing action of two separate components provided in trans, the excision of a split intein (i.e., the N-intein and C-intein in a split intein configuration) can also be referred to as “trans-splicing”.
[0151] [Figure 68A] Figure 68A demonstrates that delivery of both halves of the split intein of SpPE (SEQ ID NO: 762) in the linker maintains activity at three test sites when co-transfected into HEK293T cells.
[0152] [Figure 68B] Figure 68B demonstrates that delivery of both halves of the split intein of SaPE2 (e.g., SEQ ID NO: 443 and SEQ ID NO: 450) reproduces the activity of full-length SaPE2 (SEQ ID NO: 134) when co-transfected into HEK293T cells. Residues indicated by quotation marks are SaCas
[0153] The sequence consists of amino acids 741-743 of the 9-chain matrix (the first residue of the C-terminal extension), which are crucial for the intein trans-splicing reaction. "SMP" is the native residue. We also mutated these to the "CFN" consensus splicing sequence. As measured by the prime editing percentage, the consensus sequence has been shown to produce the highest rearrangement.
[0154] [Figure 68C] Figure 68C provides data showing that various disclosed PE ribonucleoprotein complexes (high-concentration PE2, high-concentration PE3, and low-concentration PE3) can be delivered in this manner.
[0155] [Figure 69] Figure 69 shows a bacteriophage plaque assay to determine the efficacy of PE in PANCE. Plaques (dark circles) indicate phages capable of successfully infecting E. coli. Increasing the concentration of L-rhamnose resulted in increased PE expression and increased plaque formation. Plaque sequencing revealed the presence of genome editing incorporated by PE.
[0156] [Figure 70]Figures 70A–70I provide examples of edited target sequences as illustrative step-by-step instructions for designing PEgRNA and nicking sgRNA for prime editing. Figure 70A: Step 1. Determine the target sequence and edit. Retrieve the sequence of the target DNA region (~200 bp) centered on the location of the desired edit (point mutation, insertion, deletion, or combination thereof). Figure 70B: Step 2. Position the target PAM. Identify the PAM proximal to the edit location. Take care to look for PAMs on both strands. A PAM close to the edit site is preferred, but it is possible to incorporate the edit using a protospacer and PAM that nick ≥30 nt from the edit site. Figure 70C: Step 3. Position the nick site. For each PAM to be considered, identify the corresponding nick site. In Sp Cas9 H840A nickase, the cleavage occurs between the 3rd and 4th bases of the 5' of the NGG PAM on the PAM-containing strand. All edited nucleotides must be located at 3' of the nick site. Therefore, a suitable PAM must place the nick at 5' of the target edit on the PAM-containing strand. In the example shown below, there are two possible PAMs. For simplicity, the remaining steps demonstrate the design of PEgRNA using only PAM 1. Figure 70D: Step 4. Design the spacer sequence. The protospacer of Sp Cas9 corresponds to the 20 nucleotides at 5' of the NGG PAM on the PAM-containing strand. Efficient Pol III transcription initiation requires that G be the first transcribed nucleotide. If the first nucleotide of the protospacer is G, then the spacer sequence of PEgRNA is simply the protospacer sequence. If the first nucleotide of the protospacer is not G, then the spacer sequence of PEgRNA is G, then the protospacer sequence. Figure 70E: Step 5. Design the primer binding site (PBS). Identify the DNA primers on the PAM-containing strand using the starting allele sequence. The 3' end of the DNA primer is the nucleotide just upstream of the nick site (i.e., the fourth base at the 5' end of the NGG PAM in SpCas9).As a general design principle for use in PE2 and PE3, a PEgRNA primer-binding site (PBS) containing 12 to 13 nucleotides complementary to the DNA primer can be used for sequences with a GC content of ~40-60%. For sequences with a low GC content, longer (14 to 15 nt) PBS should be tested. For sequences with a higher GC content, shorter (8 to 11 nt) PBS should be tested. Regardless of GC content, the optimal PBS sequence should be determined empirically. To design a PBS sequence of length p, take the reverse complement of the first p nucleotide at 5' of the nick site on the PAM-containing strand using the starting allele sequence. Figure 70F: Step 6. Design the RT template. The RT template encodes homology to the designed edit and the sequences adjacent to the edit. The optimal RT template length varies based on the target site. For short-range editing (positions +1 to +6), it is recommended to test short (9 to 12 nt), medium (13 to 16 nt), and long (17 to 20 nt) RT templates. For long-range editing (positions +7 and beyond), it is recommended to use RT templates that extend at least 5 nt (preferably 10 nt or more) beyond the editing site to allow sufficient 3' DNA flap homology. For long-range editing, several RT templates should be screened to identify a functional design. For larger insertions and deletions (≧5 nt), incorporating greater 3' homology (~20 nt or more) onto the RT template is recommended. Editing efficiency is typically impaired when the RT template encodes the synthesis of G (corresponding to C in PEgRNA RT templates) as the last nucleotide on the DNA product being reverse-transcribed. Since many RT templates support efficient prime editing, it is recommended to avoid G as the last nucleotide synthesized when designing RT templates. To design an RT template sequence of length r, the desired allele sequence is used, and the reverse complement of the first r nucleotide at 3' of the nick site on the strand originally containing PAM is taken. Note that, compared to SNP editing, insertion or deletion editing using the same length RT template will not contain the same homology. Figure 70G: Step 7. Assemble the complete PEgRNA sequence.Concatenate the PEgRNA components in the following order (5' to 3'): spacer, backbone, RT template, and PBS. Figure 70H: Step 8. Design the nicking sgRNA for PE3. Identify the PAM on the unedited strand upstream and downstream of the edit. The optimal nicking site is highly locus-dependent and should be determined empirically. Generally, a nick placed 40 to 90 nucleotides 5' opposite the nick induced by the PEgRNA leads to higher edit yields and fewer indels. The nicking sgRNA has a spacer sequence that matches a 20nt protospacer on the starting allele. It has a 5'G addition if the protospacer does not start with G. Figure 70I: Step 9. Design the PE3b nicking sgRNA. If the PAM is present on the complementary strand and its corresponding protospacer overlaps with the sequence targeted for editing, this edit may be a candidate for the PE3b system. In the PE3b system, the spacer sequence of the nicking sgRNA matches the sequence of the desired allele to be edited, rather than the starting allele. The PE3b system operates efficiently when the nucleotide(s) to be edited are located within the seed region (~10nt adjacent to the PAM) of the nicking sgRNA protospacer. This prevents nicking of the complementary strand until after the incorporation of the edited strand, thus preventing competition between the PEgRNA and sgRNA for binding to the target DNA. PE3b also avoids simultaneous nicking on both strands, and therefore significantly reduces indel formation while maintaining high editing efficiency. The PE3b sgRNA should have a spacer sequence that matches the 20nt protospacer of the desired allele, with an additional 5'G if required.
[0157] [Figure 71A]Figure 71A shows the nucleotide sequence of the SpCas9 PEgRNA molecule (top). It terminates with "UUU" at the 3' end and does not contain a troop element. The lower part of the figure illustrates the same SpCas9 PEgRNA molecule, but is further modified to include a troop element with the sequence 5'-"GAAANNNNN"-3' inserted immediately before the "UUU" 3' end. "N" can be any nucleic acid base.
[0158] [Figure 71B] Figure 71B shows the results of Example 18. This demonstrates that the efficiency of prime editing in HEK or EMX cells can be increased using PEgRNA containing a troop element, but the percentage of indel formation does not change significantly.
[0159] [Figure 72]Figure 72 illustrates alternative PEgRNA configurations that can be used for prime editing. (a) illustrates the PE2:PEgRNA configuration for prime editing. This configuration involves PE2 (a fusion protein containing Cas9 and reverse transcriptase) complexed with PEgRNA (as also described in Figure 1A-1I and / or Figure 3A-3E). In this configuration, the reverse transcription template is incorporated onto the 3' elongation arm of the sgRNA to form PEgRNA, and the DNA polymerase enzyme is reverse transcriptase (RT) directly fused to Cas9. (b) illustrates the MS2cp-PE2:sgRNA+tPERT configuration. This configuration includes a PE2 fusion (Cas9+reverse transcriptase), which is further fused to the MS2 bacteriophage coat protein (MS2cp) to form the MS2cp-PE2 fusion protein. To achieve prime editing, the MS2cp-PE2 fusion protein is complexed with the sgRNA, which targets the complex to a specific target site on the DNA. Furthermore, the embodiment involves the introduction of a transprime-editing RNA template ("tPERT"), which acts in place of PEgRNA by providing a primer-binding site (PBS) and a DNA synthesis template by a separate molecule, namely tPERT. This also comprises an MS2 aptamer (stem-loop). The MS2cp protein recruits tPERT by binding to the MS2 aptamer of the molecule. (c) illustrates alternative designs of PEgRNA that can be achieved by known methods for the chemosynthesis of nucleic acid molecules. For example, chemosynthesis may be used to synthesize a hybrid RNA / DNA PEgRNA molecule for use in prime editing, where the elongation arm of the hybrid PEgRNA is DNA instead of RNA. In such an embodiment, DNA-dependent DNA polymerase may be used instead of reverse transcriptase to synthesize the 3' DNA flap containing the desired genetic alteration formed by prime editing. In another embodiment, the elongation arm may be synthesized to encompass a chemical linker, which prevents DNA polymerase (e.g., reverse transcriptase) from using the sgRNA backbone or backbone as a template.In another embodiment, the elongation arm may include a DNA synthesis template oriented in the opposite direction relative to the overall orientation of the PEgRNA molecule. For example, as shown for a PEgRNA having an elongation attached to the 3' end of the sgRNA backbone and oriented from 5' to 3', the DNA synthesis template is oriented in the opposite direction, i.e., from 3' to 5'. This embodiment may be advantageous for PEgRNA embodiments having an elongation arm positioned at the 3' end of the gRNA. By reversing the orientation of the elongation arm, once it reaches the 5' end of the new orientation of the elongation arm, DNA synthesis by polymerase (e.g., reverse transcriptase) will terminate, and therefore there will be no risk of using the gRNA core as a template.
[0160] [Figure 73] Figure 73 demonstrates prime editing using the tPERT and MS2 recruitment system (also known as MS2 tagging technology). An sgRNA that targets the prime editing factor protein (PE2) to a target locus is expressed in combination with tPERT containing a primer binding site (13nt or 17nt PBS), an RT template encoding His6 tag insertion and homologous arms, and an MS2 aptamer (located at the 5' or 3' end of the tPERT molecule). Either the prime editing factor protein (PE2) or a fusion of MS2cp to the N-terminus of PE2 was used. Editing was performed with or without complementary strand nicking sgRNA, as in the previously developed PE3 system (labeled "PE2+nicking" or "PE2" on the x-axis, respectively). This is also called "second strand nicking" as defined herein.
[0161] [Figure 74]Figure 74 demonstrates the MS2 aptamer expression of reverse transcriptase in trans and its recruitment by the MS2 aptamer system. The PEgRNA contains an MS2 RNA aptamer inserted into one of the two sgRNA backbone hairpins. Wild-type M-MLV reverse transcriptase is expressed as an N-terminal or C-terminal fusion to the MS2 coat protein (MCP). Editing occurs at the HEK3 site in HEK293T cells.
[0162] [Figure 75] Figure 75 provides bar graphs comparing the efficiency (i.e., "percentage of total sequencing reads with defined edits or indels") of PE2, PE2-trunc, PE3, and PE3-trunc for different target sites in various cell lines. The data show that prime editing factors containing truncated RT variants were approximately as efficient as prime editing factors containing untruncated RT proteins.
[0163] [Figure 76] Figure 76 demonstrates the editing efficiency of the intein mitotic prime editing factor in Example 20. HEK239T cells were transfected with plasmids encoding full-length PE2 or intein mitotic PE2, PEgRNA, and nicking guide RNA. The consensus sequence (most of the amino-terminal residues of the C-terminal extension) is indicated. Percentage editing at two sites is shown: HEK3 +1CTT insertion and PRNP +6G to T. Replicate n=3 independent transfections. See Example 20.
[0164] [Figure 77]Figure 77 demonstrates the editing efficiency of the intein mitotic prime editing factor in Example 20. Editing was assessed by targeted deep sequencing of bulk cortical and GFP+ subpopulations after delivery of nuclear-localized GFP:KASH at 5E10vg and a small amount of 1E10 per half SpPE3 to P0 mice via ICV injection. The editing factor and GFP were packaged in AAV9 with an EFS promoter. Mice were harvested 3 weeks after injection, and GFP+ nuclei were isolated by flow cytometry. Individual data points are shown for 1-2 mice per condition analyzed. See Example 20.
[0165] [Figure 78] Figure 78 demonstrates the editing efficiency of the intein mitotic prime editing factor in Example 20. Specifically, the figure illustrates the AV mitotic SpPE3 construct used in Example 20. Cotransduction with AAV particles expressing SpPE3-N and SpPE3-C separately reproduces PE3 activity. Note that the N-terminal genome contains a U6-sgRNA cassette expressing nickel sgRNA, and the C-terminal genome contains a U6-PEgRNA cassette expressing PEgRNA. See Example 20.
[0166] [Figure 79]Figure 79 shows the editing efficiency of certain optimized linkers discussed in Example 21. Specifically, the data shows the editing efficiency of the PE2 construct with the current linker (labeled PE2; white box) for representative PEgRNAs for transition, transversion, insertion, and deletion editing at the HEK3, EMX1, FANCF, and RNF2 loci, compared to various versions in which the linker is replaced by the indicated sequence. The replacement linkers are referred to as "1×SGGS", "2×SGGS", "3×SGGS", "1×XTEN", "no linker", "1×Gly", "1×Pro", "1×EAAAK", "2×EAAAK", and "3×EAAAK". Editing efficiency is measured in bar graph format relative to the "control" editing efficiency of PE2. The linker for PE2 is SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (Sequence ID 127). All edits were performed in the context of the PE3 system. In other words, this refers to the addition of a PE2 editing construct plus an optimal secondary sgRNA nicking guide. See Example 21.
[0167] [Figure 80] Taking an average efficiency ~ times relative to PE2 produces the graph shown, indicating that using a 1×XTEN linker array improves editing efficiency by an average of 1.14 times (n=15). See Example 21.
[0168] [Figure 81] Figure 81 illustrates the transcription levels of PEgRNA from different promoters, as described in Example 22.
[0169] [Figure 82] As illustrated in Example 22, the impact of different types of structural modifications on PEgRNA on the relative editing efficiency compared to unmodified PEgRNA.
[0170] [Figure 83]Figure 83 illustrates a PE experiment targeting HEK3 gene editing. PE3 was used to specifically target a 10nt insertion at position +1 relative to the nick site. See Example 22.
[0171] [Figure 84A] Figure 84A illustrates an exemplary PEgRNA having a spacer, gRNA core, and extension arm (RT template + primer binding site). The 3' end of the PEgRNA is modified by a tRNA molecule coupled via a UCU linker. The tRNA can encompass various post-transcriptional modifications; however, these modifications are not required.
[0172] [Figure 84B] Figure 84B illustrates the structure of a tRNA that may be used to modify the PEgRNA structure. See Example 22. P1 may be of variable length. P1 may be elongated to help prevent RNAseP processing of the PEgRNA-tRNA fusion.
[0173] [Figure 85] Figure 85 illustrates a PE experiment targeting editing of the FANCF gene. The conversion from G to T at position +5 relative to the nick site was specifically targeted, using the PE3 construct. See Example 22.
[0174] [Figure 86] Figure 86 illustrates a PE experiment targeting HEK3 gene editing. The PE3 construct was used to specifically target the insertion of a 71nt FLAG tag at position +1 relative to the nick site. See Example 22.
[0175] [Figure 87] Results from screening N2A cells in which pegRNA incorporates 1412Adel, with details of primer-binding site (PBS) length and reverse transcriptase (RT) template length (shown with and without indels). See Example 23.
[0176] [Figure 88] Results from screening N2A cells in which pegRNA incorporates 1412Adel, with details of primer-binding site (PBS) length and reverse transcriptase (RT) template length (shown with and without indels). See Example 23.
[0177] [Figure 89] Figure 89 illustrates the results of editing at the proximal locus and HEK3 of the β-globin gene in healthy HSCs. The concentrations of editing factors versus pegRNA and nicking gRNA were varied. See Example 23.
[0178] definition Unless otherwise specified, all technical and scientific terms used in this application have the meanings commonly understood by practitioners of the art to which this invention pertains. The following references provide many common definitions of terms used in this invention to those skilled in the art: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). Unless otherwise specified, the following terms used in this application have the meanings attributed thereto.
[0179] Antisense chain In genetics, the "antisense" strand of a segment within a double-stranded DNA is considered the template strand, extending in a 3'→5' orientation (runs). Conversely, the "sense" strand is a segment within the double-stranded DNA that extends from 5' to 3' and is complementary to the antisense or template strand of the DNA that extends from 3' to 5'. In the case of a protein-coding DNA segment, the sense strand is the DNA strand that has the same sequence as the mRNA that takes the antisense strand as its template during transcription and eventually (typically, but not always) undergoes translation to become a protein. Thus, the antisense strand carries the RNA that is later translated into a protein, while the sense strand has a makeup that is nearly identical to that of the mRNA. Note that for each segment of dsDNA, there are likely to be two sets of sense and antisense strands, depending on the direction from which one is read (since sense and antisense are relative to the viewpoint). The gene product or mRNA ultimately indicates that one of the strands of one of the segments of the dsDNA is referred to as either sense or antisense.
[0180] Dual-specific ligands As used in this application, the terms “bispecific ligand” or “bispecific moiety” refer to a ligand that binds to two different ligand-binding domains. In one embodiment, the ligand is a small molecule compound, peptide, or polypeptide. In another embodiment, the ligand-binding domain is a “dimerizing domain,” which may be incorporated onto a protein as a peptide tag. In various embodiments, two proteins, each containing the same or different dimerizing domains, may be induced to dimerize by the binding of each dimerizing domain to a bispecific ligand. As used in this application, the term “bispecific ligand” may also be referred to as a “chemical inducer of dimerization” or “CID.”
[0181] Cas9 The term “Cas9” or “Cas9 nuclease” refers to an RNA-guided nuclease containing the Cas9 domain or a fragment thereof (for example, a protein containing the active or inactive DNA cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9). “Cas9 domain” as used herein is a protein fragment containing the active or inactive cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9. “Cas9 protein” is the full-length Cas9 protein. Cas9 nucleases are also sometimes referred to as casn1 nucleases or CRISPR (clustered and regularly arranged short palindromic sequence repeats). C stestered R egularly I nterspaced S hort P alindromic RAlso referred to as epeat))-related nucleases. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the preceding mobile element, and a targeted entry nucleic acid. The CRISPR cluster is transcribed and processed into processed CRISPR RNA (crRNA). In the type II CRISPR system, the modification processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 domain. The tracrRNA acts as a guide for the ribonuclease 3-aided processing of the pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endonuclease-likely cleaves a linear or circular dsDNA target complementary to the spacer. Target strands that are not complementary to crRNA are first cleaved endonucleaseically and then 3'-5' exonucleaseically. Normally, DNA binding and cleavage require both proteins and both RNAs. However, single guide RNAs ("sgRNAs," or simply "gNRAs") can be modified to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012) (this entire content is incorporated herein by reference). Cas9 helps distinguish self versus non-self by recognizing short motifs within CRISPR repeat sequences (PAMs or protospacer-adjacent motifs).Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.”Deltcheva E., Chylinski K., Sharma See CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012) (the entire contents of each of these are incorporated herein by reference). Cas9 orthologs have been described in various species, including S. pyogenes and S. thermophilus, but are not limited to these. Additional preferred Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure.Furthermore, such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems” (2013) RNA Biology 10:5, 726-737 (the entire content of which is incorporated herein by reference). In some embodiments, the Cas9 nuclease contains one or more mutations that partially impair or inactivate the DNA cleavage domain.
[0182] A Cas9 domain with an inactivated nuclease may interchangeably be referred to as the “dCas9” protein (where the nuclease represents the “inactive” form of Cas9). Methods for generating Cas9 domains (or fragments thereof) with an inactive DNA cleavage domain are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5): 1173-83 (the entire contents of each of these are incorporated herein by reference)). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)). In some embodiments, proteins containing fragments of Cas9 are provided. For example, in some embodiments, the protein contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA-cleaving domain of Cas9. In some embodiments, the protein or fragment containing Cas9 is referred to as a "Cas9 variant." The Cas9 variant shares homology with Cas9 or its fragment.For example, a Cas9 variant is at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, at least approximately 99.8% identical, or at least approximately 99.9% identical to a wild-type Cas9 (e.g., SpCas9 with sequence number 18). In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9 (for example, SpCas9 of SEQ ID NO: 18). In some embodiments, a Cas9 variant includes a fragment of Sequence ID X Cas9 (e.g., a gRNA-binding domain or a DNA-cleaving domain) such that its fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9 (e.g., SpCas9 of Sequence ID X). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas9 (e.g., SpCas9 of SEQ ID NO: 18).
[0183] cDNA The term "cDNA" refers to a strand of DNA copied from an RNA template. cDNA is complementary to the RNA template.
[0184] cyclic permutant As used herein, the term “circular permutant” refers to a protein or polypeptide (e.g., Cas9) that contains a circular permutation, which is a change in the structural configuration of a protein involving a change in the order of amino acids in the amino acid sequence of the protein. In other words, a circular permutant is a protein in which the N-terminus and C-terminus are altered compared to the wild-type counterpart, for example, the C-terminal half of the wild-type protein is replaced by a new N-terminal half. A circular permutation (or CP) is essentially a morphological rearrangement of the protein's primary sequence, where its N-terminus and C-terminus are joined, often with a peptide linker, but simultaneously splitting the sequence at different positions to create new adjacent N-terminus and C-terminus. The result is a protein structure that may have a different but often the same or overall similar three-dimensional (3D) shape, possibly including improved or altered features (including reduced sensitivity to proteolysis, improved catalytic activity, altered substrate or ligand binding, and / or improved thermal stability). Circulating substitution proteins can occur naturally (e.g., concanavalin A and lectins). In addition, circulating substitutions can result from post-translational modifications or may be modified using recombinant techniques.
[0185] Circularly replaced Cas9 The term “circularly permuted Cas9” refers to any Cas9 protein or variant that has been generated as a circularly permuted form, thereby undergoing local rearrangement of its N-terminus and C-terminus. Such circularly permuted Cas9 proteins ("CP-Cas9"), or their variants, retain their ability to bind to DNA when complexed with guide RNA (gRNA). See Oakes et al., “Protein Engineering of Cas9 for enhanced function,” Methods Enzymol, 2014, 546:491-511 and Oakes et al., “CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification,” Cell, January 10, 2019, 176:254-267 (each of which is incorporated herein by reference). This disclosure uses any previously known CP-Cas9 or any new CP-Cas9, provided that the resulting circulating replacement protein retains its ability to bind to DNA when complexed with guide RNA (gRNA). Exemplary CP-Cas9 proteins are sequence numbers 77-86.
[0186] CRISPR CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of preceding infections caused by viruses that have invaded prokaryotes. These DNA snippets are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with arrays of CRISPR-related proteins (including Cas9 and its homologs) and CRISPR-related RNAs, they form an effective prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), the correct processing of precrRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA acts as a guide for the ribonuclease 3-assisted processing of precrRNA. Subsequently, Cas9 / crRNA / tracrRNA endo-cleaves linear or circular dsDNA targets complementary to the RNA. Specifically, target strands not complementary to crRNA are first endo-cut and then 3'-5' exo-trimmed. In nature, DNA binding and cleavage typically require both proteins and both RNAs. However, single guide RNAs ("sgRNA" or simply "gNRA") can be manipulated to incorporate aspects of both crRNA and tracrRNA into a single RNA species guide RNA. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). The entire content of this is incorporated here by reference. Cas9 recognizes short motifs (PAM or protospacer adjacency motifs) on CRISPR repeat sequences to help distinguish between self and non-self.CRISPR biology and Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP,Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.”Deltcheva E., Chylinski See K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). The entire contents of each of these are incorporated into this application by reference. Cas9 orthologs have been described in various species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems” (2013) RNA Biology 10:5, 726-737; the entire contents of this publication are incorporated herein by reference.
[0187] In certain types of CRISPR systems (e.g., type II CRISPR systems), the correct processing of precrRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA acts as a guide for the ribonuclease 3-assisted processing of the precrRNA. Subsequently, Cas9 / crRNA / tracrRNA endo-cleaves linear or circular nucleic acid targets complementary to the RNA. Specifically, target strands not complementary to the crRNA are first endo-cut and then 3'-5' exo-trimmed. In nature, DNA binding and cleavage typically require proteins and both RNAs. However, single guide RNAs ("sgRNA" or simply "gRNA") can be manipulated to incorporate aspects of both crRNA and tracrRNA into a single RNA species guide RNA.
[0188] Generally, the “CRISPR system” collectively refers to the transcripts and other elements involved in the expression of CRISPR-related (“Cas”) genes or that lead to their activity, and includes the Cas gene, tracr (trans-activated CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (in the context of endogenous CRISPR systems, encompassing “direct repeats” and partial direct repeats processed by tracrRNA), guide sequences (also called “spacers” in the context of endogenous CRISPR systems), or sequences encoding other sequences and transcripts from the CRISPR locus. The tracrRNA of the system is (fully or partially) complementary to the tracr mate sequence present on the guide RNA.
[0189] DNA synthesis template As used herein, the term “DNA synthesis template” refers to a region or portion of the elongation arm of a PEgRNA that is utilized by the polymerase of a prime editing factor as a template strand (containing the desired edit and then encoding a 3' single-strand DNA flap that replaces the corresponding endogenous DNA strand at a target site through the mechanism of prime editing). In various embodiments, the DNA synthesis template is shown in Figure 3A (in the context of a PEgRNA including a 5' elongation arm), Figure 3B (in the context of a PEgRNA including a 3' elongation arm), Figure 3C (in the context of an internal elongation arm), Figure 3D (in the context of a 3' elongation arm), and Figure 3E (in the context of a 5' elongation arm). The elongation arm contains the DNA synthesis template and may consist of DNA or RNA. In the case of RNA, the polymerase of the prime editing factor may be an RNA-dependent DNA polymerase (e.g., reverse transcriptase). In the case of DNA, the polymerase of the prime editing factor may be a DNA-dependent DNA polymerase. In various embodiments (as depicted, for example, in Figures 3D-3E), the DNA synthesis template (4) may include the “editing template” and the “homologous arm,” as well as all or part of any 5'-end modified region e2. That is, depending on the nature of the e2 region (for example, whether it encompasses a hairpin, toe loop, or stem / loop secondary structure), the polymerase may encode nothing, or part or all of the e2 region. In other words, in the case of the 3' extension arm, the DNA synthesis template (3) may include a portion of the extension arm (3) extending from the 5' end of the primer binding site (PBS) to the 3' end of the gRNA core, which may act as a template for the synthesis of a single strand of DNA by a polymerase (for example, reverse transcriptase). In the case of a 5' elongated arm, the DNA synthesis template (3) may encompass the portion of the elongated arm (3) extending from the 5' end of the PEgRNA molecule to the 3' end of the editing template. Preferably, the DNA synthesis template excludes the primer binding site (PBS) of PEgRNAs having either a 3' elongated arm or a 5' elongated arm.The embodiments described herein (for example, Figure 71A) refer to an “RT template” that encompasses the editing template and homologous arms, i.e., the sequences of the PEgRNA elongation arms that are actually used as templates during DNA synthesis. The term “RT template” is equivalent to the term “DNA synthesis template.”
[0190] In the case of transpriming (for example, Figures 3G and 3H), the primer binding site (PBS) and DNA synthesis template can be manipulated into separate molecules called transpriming factor RNA templates (tPERT).
[0191] Dimerization domain The term “dimerization domain” refers to a ligand-binding domain that binds to the binding domain of a bispecific ligand. The “first” dimerization domain binds to the first binding site of the bispecific ligand, and the “second” dimerization domain binds to the second binding site of the same bispecific ligand. When the first dimerization domain is fused to the first protein (e.g., via PE as discussed in this application) and the second dimerization domain is fused to the second protein (e.g., via PE as discussed in this application), the first and second proteins dimerize in the presence of the bispecific ligand, and the bispecific ligand has at least one portion that binds to the first dimerization domain and at least another portion that binds to the second dimerization domain.
[0192] downstream As used herein, the terms “upstream” and “downstream” are relative terms defining the linear positions of at least two elements located within a nucleic acid molecule (whether single-stranded or double-stranded) oriented in the 5’→3’ direction. In particular, if the first element is located somewhere at 5’ relative to the second element, the first element is upstream of the second element in the nucleic acid molecule. For example, if the SNP is on the 5’ side of a nick site, the SNP is upstream of the Cas9-induced nick site. Conversely, if the first element is located somewhere at 3’ relative to the second element, the first element is downstream of the second element in the nucleic acid molecule. For example, if the SNP is on the 3’ side of a nick site, the SNP is downstream of the Cas9-induced nick site. Nucleic acid molecules can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. The analysis is the same for single-stranded and double-stranded nucleic acid molecules, except where it is necessary to choose which strand of a double-stranded molecule to consider, where the terms upstream and downstream refer only to single-stranded nucleic acid molecules. Often, the strand of double-stranded DNA that can be used to determine the relative positions of at least two elements is the “sense” strand or the “coding” strand. In genetics, the “sense” strand is the intra-double-stranded DNA segment extending from 5' to 3' that is complementary to the antisense or template strand of the DNA extending from 3' to 5'. Thus, for example, if an SNP nucleic acid base is on the 3' side of the promoter on the sense or coding strand, the SNP nucleic acid base is “downstream” of the promoter sequence in the genomic DNA (which is double-stranded).
[0193] Editing template The term “edit template” refers to the portion of the elongation arm on a single-stranded 3' DNA flap that encodes the desired edit, synthesized by a polymerase, e.g., DNA-dependent DNA polymerase or RNA-dependent DNA polymerase (e.g., reverse transcriptase). The embodiments described herein (for example, Figure 71A) refer to the “RT template” which refers to both the edit template and the homologous arm together, i.e., the sequence of the PEgRNA elongation arm that is actually used as a template during DNA synthesis. The term “RT edit template” is also equivalent to the term “DNA synthesis template,” where the RT edit template reflects the use of a prime editing factor with a polymerase that is a reverse transcriptase, while the DNA synthesis template more broadly reflects the use of a prime editing factor with either polymerase.
[0194] Effective amount As used in this application, the term “effective amount” means the amount of a bioactive agent sufficient to induce a desired biological response. For example, in some embodiments, the effective amount of a prime editing factor (PE) may mean the amount of editing factor sufficient to edit a target site nucleotide sequence, e.g., a genome. In some embodiments, the effective amount of a fusion protein of a prime editing factor (PE) provided herein, e.g., comprising a nickase Cas9 domain and a reverse transcriptase, may mean the amount of fusion protein sufficient to induce editing of a target site specifically bound to and edited by the fusion protein. As will be understood by those skilled in the art, the effective amount of an agent, e.g., a fusion protein, nuclease, hybrid protein, protein dimer, complex of a protein (or protein dimer) and a polynucleotide, or polynucleotide may vary depending on various factors, e.g., the desired biological response, e.g., a specific allele, genome, or target site to be edited, the cell or tissue to be targeted, and the agent to be used.
[0195] Reverse transcriptase prone to errors As used herein, the term “error-prone” reverse transcriptase (or more broadly, any polymerase) refers to a reverse transcriptase (or more broadly, any polymerase) that is naturally occurring or derived from another reverse transcriptase (e.g., wild-type M-MLV reverse transcriptase) having an error rate lower than that of wild-type M-MLV reverse transcriptase. The error rate of wild-type M-MLV reverse transcriptase has been reported to be between 15,000 (higher) and 27,000 (lower) for a 1 in 1 error. The error rate of 1 in 15,000 is 6.7 × 10⁻⁶. -5 This corresponds to an error rate of 1 in 27,000. -5 This corresponds to an error rate of 6.7 × 10⁻¹⁰. See Boutabout et al. (2001) “DNA synthesis fidelity by the reverse transcriptase of the yeast retrotransposon Ty1,” Nucleic Acids Res 29(11):2217-2222 (which is incorporated herein by reference). Therefore, for the purposes of this application, the term “error-prone” means an error greater than 1 (6.7 × 10⁻¹⁰) in the incorporation of 15,000 nucleic acid bases. -5 Or higher), for example, 1 error in 14,000 nucleic acid bases (7.14 × 10⁻¹⁴). -5 (or higher), 1 error (7.7 × 10) within 13,000 nucleic acid bases or fewer. -5 (or higher), 1 error (7.7 × 10) within 12,000 nucleic acid bases or less. -5 (or higher), 1 error (9.1 × 10) within 11,000 nucleic acid bases or less. -5 (or higher), 1 error (1 × 10⁻¹⁰) within 10,000 nucleic acid bases or fewer. -4(or 0.0001 or higher), 1 error within 9,000 nucleic acid bases or less (0.00011 or higher), 1 error within 8,000 nucleic acid bases or less (0.00013 or higher), 1 error within 7,000 nucleic acid bases or less (0.00014 or higher), 1 error within 6,000 nucleic acid bases or less (0.00016 or higher), 1 error within 5,000 nucleic acid bases or less (0.0002 or higher), 1 error within 4,000 nucleic acid bases or less This refers to RTs that have an error rate of (0.00025 or higher), 1 error in 3,000 nucleic acid bases or less (0.00033 or higher), 1 error in 2,000 nucleic acid bases or less (0.00050 or higher), 1 error in 1,000 nucleic acid bases or less (0.001 or higher), 1 error in 500 nucleic acid bases or less (0.002 or higher), or 1 error in 250 nucleic acid bases or less (0.004 or higher).
[0196] Extain As used in this application, the term "extine" refers to a polypeptide sequence that is flanked by an intein and ligated to another extine during the protein splicing process to form a mature, spliced protein. Typically, an intein is flanked by two extine sequences, which are ligated together when the intein catalyzes its own excision. An extine is therefore a protein analog to an exon found on mRNA. For example, a polypeptide containing an intein may have the structure extine(N)-intine-extine(C). After excision of the intein and splicing of the two extines, the resulting structure is extine(N)-extine(C) and free intein. In various configurations, the extines may be separate proteins (e.g., half of a Cas9 or PE fusion protein) each fused to a fission intein, and excision of the fission intein triggers splicing of the extine sequences together.
[0197] Extension arm The term “elongation arm” refers to a nucleotide sequence component of PEgRNA that provides several functions, including a primer-binding site and an editing template for reverse transcriptase. In some embodiments, for example, in Figure 3D, the elongation arm is located at the 3' end of the guide RNA. In other embodiments, for example, in Figure 3E, the elongation arm is located at the 5' end of the guide RNA. In some embodiments, the elongation arm also includes a homologous arm. In various embodiments, the elongation arm includes the following components in the 5'→3' direction: a homologous arm, an editing template, and a primer-binding site. Since the polymerization activity of reverse transcriptase is in the 5'→3' direction, the preferred arrangement of the homologous arm, editing template, and primer-binding site is in the 5'→3' direction so that the reverse transcriptase, once primed by the annealed primer sequence, polymerizes single-stranded DNA using the editing template as a complementary template strand. Further details, such as the length of the elongation arm, are described elsewhere in this specification.
[0198] The elongation arm may also be described as comprising two regions in general: a primer-binding site (PBS) and a DNA synthesis template, as illustrated in Figure 3G (top row). The primer-binding site is nicked by the prime-editing factor complex, so that when its 3' end is exposed to the nicked endogenous strand, it binds to the primer sequence formed from the endogenous DNA strand at the target site. As described herein, the binding of the primer sequence to the primer-binding site on the elongation arm of PEgRNA creates a double-stranded region with an exposed 3' end (i.e., 3' of the primer sequence), which then provides a substrate for polymerase to begin polymerizing a single strand of DNA along the length of the DNA synthesis template from the exposed 3' end. The sequence of the single-stranded DNA product is a complement to the DNA synthesis template. Polymerization continues to 5' of the DNA synthesis template (or elongation arm) until polymerization terminates. Therefore, the DNA synthesis template represents a portion of the elongation arm encoded by the polymerase of the prime editing factor complex into a single-strand DNA product (i.e., a 3' single-strand DNA flap containing the desired gene editing information), which replaces the corresponding endogenous DNA strand at the target site immediately downstream of the PE-induced nick site. While not bound by theory, polymerization of the DNA synthesis template continues to the 5' end of the elongation arm until a termination event occurs. Polymerization may terminate in various ways, including, but is not limited to, (a) reaching the 5' end of the PEgRNA (e.g., in the case of the 5' elongation arm, the DNA polymerase simply runs out of the template), (b) reaching an RNA secondary structure that cannot be passed through (e.g., a hairpin or stem / loop), or (c) reaching a replication termination signal, e.g., a specific nucleotide sequence that blocks or inhibits the polymerase, or a signal on the nucleic acid morphology such as supercoiled DNA or RNA.
[0199] Flap end nucleases (for example, FEN1) As used herein, the term “flap endonuclease” refers to an enzyme that catalyzes the removal of a 5' single-strand DNA flap. These are naturally occurring enzymes that handle the removal of 5' flaps formed during cellular processes (including DNA replication). The prime editing methods described herein may utilize endogenously supplied flap endonucleases or those supplied trans to remove the 5' flap of endogenous DNA formed at the target site during prime editing. Flap endonucleases are known in the art and can be found in Patel et al., “Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends,” Nucleic Acids Research, 2012, 40(10):4507-4519 and Tsutakawa et al., “Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily,” Cell, 2011, 145(2):198-211 and Balakrishnan et al., “Flap Endonuclease 1,” Annu Rev Biochem, 2013, Vol 82:119-138 (each of these incorporated herein by reference). An exemplary flap endonuclease is FEN1, which can be represented by the following amino acid sequence: [Table 1]
[0200] functional equivalent The term “functional equivalent” refers to a second biomolecule that is functionally equivalent but structurally not equivalent to a first biomolecule. For example, a “Cas9 equivalent” refers to a protein that has the same or substantially the same function as Cas9 but does not necessarily have the same amino acid sequence. In the context of this disclosure, “protein X, or its functional equivalent” refers throughout this specification. In this regard, a “functional equivalent” of protein X accepts any homolog, paralog, fragment, naturally occurring version, modified version, mutant version, or synthetic version that carries the equivalent function of protein X.
[0201] Fusion protein The term “fusion protein,” as used herein, refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein may be positioned at the amino-terminus (N-terminus) or carboxy-terminus (C-terminus) of the fusion protein, thus forming an “amino-terminus fusion protein” or a “carboxy-terminus fusion protein,” respectively. The protein may contain different domains, such as a nucleic acid-binding domain (e.g., the gRNA-binding domain of Cas9, which directs the protein toward binding to a target site) and a nucleic acid-cleaving domain or catalytic domain of a nucleic acid-editing protein. Another example includes Cas9 or its equivalent fused with reverse transcriptase. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, particularly suitable for fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) (the entire contents of which are incorporated herein by reference)).
[0202] Gene of interest (GOI) The term "GOI" refers to a gene that codes for a biomolecule of interest (e.g., a protein or RNA molecule). The protein of interest may include any intracellular, membrane, or extracellular protein, such as nuclear proteins, transcription factors, nuclear transporters, intracellular organelle-associated proteins, membrane receptors, catalytic proteins, and enzymes, therapeutic proteins, membrane proteins, membrane transport proteins, signaling proteins, or immunological proteins (e.g., IgG or other antibody proteins), etc. The gene of interest may also code for RNA molecules, including, but is not limited to, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), antisense RNA, guide RNA, microRNA (miRNA), small interfering RNA (siRNA), and cell-free RNA (cfRNA).
[0203] Guide RNA ("gRNA") As used herein, the term “guide RNA” mostly refers to a specific type of guide nucleic acid commonly associated with the Cas protein of CRISPR-Cas9, which binds to Cas9 and directs the Cas9 protein to a specific sequence in the DNA molecule (including complementarity with the protospace sequence of the guide RNA). However, the term also accepts equivalent guide nucleic acid molecules that bind to Cas9 equivalents, homologs, orthologues, or paralogs, whether naturally occurring or not (e.g., modified or recombinant), and that are separately programmed to localize the Cas9 equivalent to a specific target nucleotide sequence. The Cas9 equivalent may also include other napDNAbp from any type of CRISPR system (e.g., type II, type V, type VI), encompassing Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Further Cas equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016;353(6299) (this content is incorporated herein by reference). Exemplary sequences and structures of guide RNAs are provided herein. In addition, methods for designing appropriate guide RNA sequences are also provided herein. When used herein, “guide RNA” may also be referred to as “existing guide RNA” to distinguish it from a modified form of guide RNA called “prime editing guide RNA” (or “PEgRNA”) invented for the prime editing methods and compositions disclosed herein.
[0204] Guide RNA or PEgRNA may contain a variety of structural elements, but are not limited to, the following:
[0205] Spacer sequence - A sequence within a guide RNA or PEgRNA (approximately 20 nt in length) that binds to a protospacer in the target DNA.
[0206] The gRNA core (or gRNA backbone or main chain sequence) refers to the sequence within the gRNA responsible for Cas9 binding, and does not include the spacer / targeting sequence used to guide Cas9 to the target DNA.
[0207] Elongation arm - single-strand elongation of the 3' or 5' end of PEgRNA. This includes a primer binding site and a DNA synthesis template sequence encoding a single-strand DNA flap containing the desired gene alteration, which is then incorporated onto endogenous DNA by polymerase (e.g., reverse transcriptase).
[0208] Transcriptional terminator-guide RNA or PEgRNA may contain a transcription termination sequence at the 3' end of the molecule.
[0209] Homologous arms The term “homologous arm” refers to the portion of an elongation arm that codes for the resulting single-strand DNA flap (which is incorporated into the target DNA site by replacing the endogenous strand) encoded by the reverse transcriptase. The single-strand DNA flap portion coded by the homologous arm is complementary to the unedited strand of the target DNA sequence, facilitating its replacement for the endogenous strand and annealing with the single-strand DNA flap at that location, thereby incorporating the edit. This component is further defined elsewhere. By definition, since it is coded by the polymerase of the prime editing factor described herein, the homologous arm is part of the DNA synthesis template.
[0210] host cell When used herein, the term “host cell” refers to a cell capable of hosting, replicating, and expressing a vector described herein, for example, a vector comprising a nucleic acid molecule encoding a fusion protein including Cas9 or a Cas9 equivalent and reverse transcriptase.
[0211] Intein As used in this application, the term "intene" refers to a self-processing polypeptide domain found in organisms from all domains of life. Inteins (intermediate proteins) perform an intrinsic self-processing event known as protein splicing. In this process, they excise themselves from a larger precursor polypeptide by cleaving two peptide bonds and ligate extein (external protein) sequences that flank in the process by forming new peptide bonds. Since intein genes are found embedded in frame with other protein-coding genes, this rearrangement occurs post-translation (or co-translationally). Furthermore, intein-mediated protein splicing is spontaneous; it requires only the folding of the intein domain and no external factors or energy sources. This process is also known as cis-protein splicing, in contrast to the innate process of trans-protein splicing by "fission inteins." Inteins are protein equivalents of self-splicing RNA introns (see Perler et al., Nucleic Acids Res. 22:1125-1127 (1994)), and they catalyze their own excision from precursor proteins and have associated fusions with flanking protein sequences known as exteins (reviewed in Perler et al., Curr. Opin. Chem. Biol. 1:292-299 (1997); Perler, FBCell 92(1):1-4 (1998); Xu et al., EMBO J. 15(19):5146-5153 (1996)).
[0212] The term "protein splicing" as used in this application refers to the process in which the inner region (intein) of a precursor protein is cleaved, and the flanking regions (extines) of the protein are ligated together to form a mature protein. This natural process has been observed in numerous proteins from both prokaryotes and eukaryotes (Perler, FB, Xu, MQ, Paulus, H. Current Opinion in Chemical Biology 1997, 1, 292-299; Perler, FB Nucleic Acids Research 1999, 27, 346-347). An intein unit contains the necessary components required to catalyze protein splicing, and in many cases, it contains an endonuclease domain involved in intein mobility (Perler, FB, Davis, EO, Dean, GE, Gimble, FS, Jack, WE, Neff, N., Noren, CJ, Thomas, J., Belfort, M. Nucleic Acids Research 1994, 22, 1127-1127). However, the resulting protein is ligated and not expressed as a separate protein. Protein splicing can also occur in trans by fission inteins expressed on separate polypeptides. These spontaneously combine to form a single intein, which then undergoes the protein splicing process to ligate to a separate protein.
[0213] The elucidation of the mechanism of protein splicing has led to numerous intein-based applications (Comb, et al., U.S. Patent No. 5,496,714; Comb, et al., U.S. Patent No. 5,834,247; Camarero and Muir, J. Am. Chem. Soc., 121:5597-5598 (1999); Chong, et al., Gene, 192:271-281 (1997); Chong, et al., Nucleic Acids Res., 26:5109-5115 (1998); Chong, et al., J. Biol. Chem., 273:10567-10577 (1998); Cotton, et al. J. Am. Chem. Soc., 121:1100-1101 (1999); Evans, et al. al.,J.Biol.Chem.,274:18359-18363(1999);Evans,et al.,J.Biol.Chem.,274:3923-3926(1999);Evans,et al.,Protein Sci.,7:2256-2264(1998);Evans,et al. al.,J.Biol.Chem.,275:9091-9094(2000);Iwai and Pluckthun,FEBS Lett.459:166-172(1999);Mathys,et al.,Gene,231:1-13(1999);Mills,et al.,Proc.Natl.Acad.Sci.USA 95:3543-3548(1998);Muir,et al.,Proc.Natl.Acad.Sci.USA 95:6705-6710(1998);Otomo,et al.,Biochemistry 38:16040-16044(1999);Otomo,et al.,J.Biolmol.NMR 14:105-114(1999);Scott,et al. al.,Proc.Natl.Acad.Sci.USA 96:13638-13643(1999);Severinov and Muir,J.Biol.Chem.,273:16205-16209(1998);Shingledecker,et al.,Gene,207:187-195(1998);Southworth,et al.,EMBO J.17:918-926(1998);Southworth,et al.,Biotechniques,27:110-120(1999);Wood,et al.,Nat.Biotechnol.,17:889-892(1999);Wu,et al.,Proc.Natl.Acad.Sci.USA 95:9226-9231(1998a);Wu,et al.,Biochim Biophys Acta 1387:422-432(1998b);Xu,et al.,Proc.Natl.Acad.Sci.USA 96:388-393(1999);Yamazaki,et al. al., J. Am. Chem. Soc., 120:5591-5592 (1998)). Each reference is incorporated herein by reference.
[0214] Ligand-dependent inteins As used in this application, the term "ligand-dependent intein" refers to an intein containing a ligand-binding domain. Typically, the ligand-binding domain is inserted into the amino acid sequence of the intein, resulting in a structural intein (N)-ligand-binding domain-intein (C). Typically, ligand-dependent inteins exhibit little to no protein splicing activity in the absence of a suitable ligand, and a significant increase in protein splicing activity in the presence of a ligand. In some embodiments, ligand-dependent inteins exhibit no observable splicing activity in the absence of a ligand, but exhibit splicing activity in the presence of a ligand. In some embodiments, ligand-dependent inteins exhibit observable protein splicing activity in the absence of a ligand, and in the presence of a suitable ligand, exhibit protein splicing activity that is at least 5 times, at least 10 times, at least 50 times, at least 100 times, at least 150 times, at least 200 times, at least 250 times, at least 500 times, at least 1000 times, at least 1500 times, at least 2000 times, at least 2500 times, at least 5000 times, at least 10000 times, at least 20000 times, at least 25000 times, at least 50000 times, at least 100000 times, at least 500000 times, or at least 1000000 times greater than the activity observed in the absence of the ligand. In some embodiments, the increase in activity is dose-dependent by at least one order of magnitude, at least two orders of magnitude, at least three orders of magnitude, at least four orders of magnitude, or at least five orders of magnitude, allowing for fine-tuning of the intein activity by adjusting the ligand concentration.Suitable ligand-dependent inteins are known in the art and are provided below and published in U.S. Patent Application No. US2014 / 0065711A1; Mootz et al., “Protein splicing triggered by a small molecule.” J.Am.Chem.Soc.2002;124,9044-9045; Mootz et al., “Conditional protein splicing: a new tool to control protein structure and function in vitro and in vivo.” J.Am.Chem.Soc.2003;125,10561-10569; Buskirk et al., Proc.Natl.Acad.Sci.USA. 2004;101,10505-10510; Skretas & Wood, “Regulation of protein activity with small-molecule-controlled inteins.” Protein This includes what is described in Sci.2005;14,523-532;Schwartz, et al., “Post-translational enzyme activation in an animal via optimized conditional protein splicing.” Nat.Chem.Biol.2007;3,50-54;Peck et al., Chem.Biol.2011;18(5),619-630; the entire contents of each are incorporated here by reference. An example sequence is as follows: [Table 2-1] [Table 2-2]
[0215] Linker As used in this application, the term "linker" refers to a molecule that links two other molecules or parts. A linker can be an amino acid sequence in the case of a linker linking two fusion proteins. For example, Cas9 can be fused to reverse transcriptase via an amino acid linker sequence. A linker can also be a nucleotide sequence in the case of linking two nucleotide sequences together. For example, in this case, a conventional guide RNA is linked to the RNA elongation of a prime editing guide RNA, which may include an RT template sequence and an RT primer binding site, via a spacer or linker nucleotide sequence. In other embodiments, a linker can be an organic molecule, a group, a polymer, or a chemical part. In some embodiments, the linker is 5-100 amino acids long, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids long. Longer or shorter linkers are also intended.
[0216] isolated "Isolated" means modified or removed from its natural state. For example, a nucleic acid or peptide that is naturally present in a living animal is not "isolated," but the same nucleic acid or peptide that has been partially or completely separated from its natural state in the coexisting material is "isolated." Isolated nucleic acids or proteins may exist in a substantially purified form or in a non-native environment (e.g., a host cell).
[0217] In some embodiments, the gene of interest is encoded by an isolated nucleic acid. As used herein, the term “isolated” refers to the characteristic of a material as provided herein that it has been removed from its original or native environment (e.g., the natural environment if it is naturally occurring). Thus, a naturally occurring polynucleotide or protein or polypeptide present in a living animal is not isolated, but the same polynucleotide or polypeptide that has been separated from some or all of the coexisting material in a natural system by human intervention is isolated. Artificial or modified materials, such as nucleic acid constructs that do not exist naturally, including the expression constructs and vectors described herein, are also consequently referred to as isolated. The material does not need to be purified to be isolated. Consequently, the material may be part of a vector and / or part of a composition, and such vector or composition is still isolated in that it is not part of the environment in which the material is found in nature.
[0218] MS2 Tagging Technology In various embodiments (for example, as illustrated in Figures 72-73 and the embodiments of Example 19), the term “MS2 tagging technique” refers to a combination of an “RNA-protein interaction domain” (also known as an “RNA-protein mobilization domain or protein”) paired with an RNA-binding protein that specifically recognizes and binds to a particular hairpin structure, e.g., a specific hairpin structure. These types of systems can be utilized to mobilize various functionalities to a prime editing factor complex bound to a target site. MS2 tagging techniques are based on the innate interaction of the MS2 bacteriophage coat protein (“MCP” or “MS2cp”) with stem-loop or hairpin structures present on the phage genome, i.e., “MS2 hairpins.” In the case of prime editing, MS2 tagging techniques involve introducing an MS2 hairpin onto a desired RNA molecule involved in prime editing (e.g., PEgRNA or tPERT). This then constitutes a specific interactable binding target for an RNA-binding protein that recognizes and binds to its structure. In the case of the MS2 hairpin, it is recognized and bound by the MS2 bacteriophage coat protein (MCP). Then, when the MCP is fused to another protein (e.g., reverse transcriptase or other DNA polymerase), the MS2 hairpin can be used to trans-"mobilize" the other protein to the target site occupied by the prime editing complex.
[0219] The prime editing factors described herein may, as an aspect, incorporate any known RNA-protein interaction domain to recruit or “colocalize” a specific function of interest to the prime editing factor complex. Reviews of other modular domains of RNA-protein interactions are described in the art for example in Johansson et al., “RNA recognition by the MS2 phage coat protein,” Sem Virol., 1997, Vol.8(3):176-185; Delebecque et al., “Organization of intracellular reactions with rationally designed RNA assemblies,” Science, 2011, Vol.333:470-474; Mali et al., “Cas9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering,” Nat. Biotechnol., 2013, Vol.31:833-838; and Zalatan et al., “Engineering complex synthetic transcriptional programs with CRISPR RNA scaffolds,” Cell, 2015, Vol.160:339-350 (each of these is incorporated herein by reference in its entirety). Other systems also include PP7 hairpins that specifically recruit PCP proteins and "com" hairpins that specifically recruit Com proteins. See Zalatan et al.
[0220] The nucleotide sequence of the MS2 hairpin (i.e., referred to as the "MS2 aptamer") is as follows: GCCAACATGAGGATCACCCATGTCTGCAGGGCC (Sequence ID 763).
[0221] The amino acid sequence of MCP or MS2cp is as follows: GSASNFTQFVLVDNGGTGDVTVAPSNFANGVAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKVATQTVGGEELPVAGWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY (Sequence ID 764).
[0222] The MS2 hairpin (or "MS2 aptamer") may also be referred to as a type of "RNA effector recruitment domain" (or equivalently, an "RNA-binding protein recruitment domain" or simply a "recruitment domain") because it is a physical structure (e.g., a hairpin) incorporated onto PEgRNA or tPERT that effectively recruits other effector functions (e.g., RNA-binding proteins with various functions, such as DNA polymerase or other DNA-modifying enzymes) to the thus modified PEgRNA or rPERT, and thus co-localizes the effector function in trans to the prime editing mechanism. This application is not intended to be limited to any specific RNA effector recruitment domain, but may encompass any available such domain that includes the MS2 hairpin. Example 19 and Figure 72(b) illustrate the use of a prime editing factor containing an MS2 aptamer linked to the DNA synthesis domain (i.e., the tPERT molecule) and an MS2cp protein fused to PE2 to achieve colocalization between the prime editing factor complex (MS2cp-PE2:sgRNA complex) bound to a target DNA site and the DNA synthesis domain of the tPERT molecule.
[0223] napDNAbp As used herein, the term “nucleic acid programmed DNA-binding protein” or “napDNAbp” (with Cas9 being an example) refers to a protein that uses RNA:DNA hybridization to target and bind to a specific sequence in a DNA molecule. Each napDNAbp is bound to at least one guide nucleic acid (e.g., guide RNA) that localizes the napDNAbp to a DNA sequence containing a DNA strand (i.e., the target strand) complementary to the guide nucleic acid or a portion thereof (e.g., a protospacer of guide RNA). In other words, the guide nucleic acid “programs” the napDNAbp (e.g., Cas9 or an equivalent) to localize and bind to the complementary sequence.
[0224] While not strictly theoretical, the binding mechanism of the napDNAbp-guide RNA complex generally involves the step where napDNAbp forms an R-loop that induces unwinding of the double-stranded DNA target (thus separating the strands in the region bound by napDNAbp). The guide RNA protospacer then hybridizes with the "target strand," replacing the complementary "non-target strand," thereby forming the single-stranded region of the R-loop. In some embodiments, napDNAbp contains one or more nuclease activities, which then cleave the DNA to eliminate various types of lesions. For example, napDNAbp may contain nuclease activity that cleaves the non-target strand at a first site and / or the target strand at a second site. Depending on the nuclease activity, the target DNA may be cleaved, forming a "double-strand break" where both strands are cut. In other embodiments, the target DNA can only be cleaved at a single site; i.e., the DNA is "nicked" on one strand. Exemplary napDNAbps with various nuclease activities include “Cas9 nickase” (“nCas9”) and inactivated Cas9 (“inactive Cas9” or “dCas9”) which have no nuclease activity whatsoever. Exemplary sequences of these and other napDNAbps are provided herein.
[0225] Nikkaze The term "nickase" refers to Cas9 with one of its two inactivated nuclease domains. This enzyme is capable of cleaving only one strand of target DNA.
[0226] Nuclear localization sequence (NLS) The term “nuclear localization sequence” or “NLS” refers to an amino acid sequence that facilitates the import of a protein into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, an NLS sequence is described in Plank et al.’s international PCT application PCT / EP2000 / 011690, filed November 23, 2000, and published May 31, 2001, as WO / 2001 / 038547 (this content is incorporated herein by reference with respect to its disclosure of exemplary nuclear localization sequences). In some embodiments, an NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 16) or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 17).
[0227] nucleic acid molecule The term “nucleic acid” as used herein refers to polymers of nucleotides. Polymers include natural nucleosides (i.e., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (for example, 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C5 bromouridine, C5 fluorouridine, C5 iodouridine, C5 propynyluridine, C5 propynylcytidine, C5 methylcytidine, 7 deazaadenosine, 7 deazaguanosine, 8 oxoadenosine, 8 oxoguanosine, O(6) methylguanine, 4-acetylcytidine) These may include 5-(carboxyhydroxymethyl)uridine, dihydrouridine, methylpseudridine, 1-methyladenosine, 1-methylguanosine, N6-methyladenosine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, 2'-O-methylcytidine, arabinose, and hexose), or modified phosphate groups (e.g., phosphorothioate and 5'-N phosphoramidite linkages).
[0228] PEgRNA As used herein, the terms “prime editing guide RNA,” “PEgRNA,” or “extended guide RNA” refer to a specialized form of guide RNA modified to include one or more additional sequences for performing the prime editing methods and compositions described herein. As described herein, prime editing guide RNA includes one or more “extended regions” of nucleic acid sequences. The extended regions may include, but are not limited to, single-stranded RNA or DNA. Furthermore, the extended region may occur at the 3' end of the existing guide RNA. In other configurations, the extended region may occur at the 5' end of the existing guide RNA. In other configurations, the extended region may occur in a region within the terminal molecule of the existing guide RNA (e.g., in a gRNA core region that binds to and / or ligates to a napDNAbp). The extended region contains a “DNA synthesis template” that encodes single-stranded DNA (by the polymerase of a prime editing factor), but the molecule is then designed to be (a) homologous to the endogenous target DNA to be edited, and (b) contain at least one desired nucleotide change (e.g., a transition, transversion, deletion, or insertion) to be induced into or incorporated into the endogenous target DNA. The extended region may also contain other functional sequence elements (but not limited to “primer binding sites” and “spacer or linker” sequences) or other structural elements (but not limited to aptamers, stem-loops, hairpins, toe-loops (e.g., a 3' toe-loop), or RNA-protein recruitment domains (e.g., an MS2 hairpin)). When used herein, “primer binding sites” include a sequence that hybridizes with a single-stranded DNA sequence having a 3' end generated from R-loop nicked DNA.
[0229] In one embodiment, PEgRNA is represented by Figure 3A, showing PEgRNA having a 5' extension arm, a spacer, and a gRNA core. The 5' extension further includes a reverse transcriptase template, a primer binding site, and a linker in the 5'→3' direction. As shown, the reverse transcriptase template may also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editing factor described herein is a different type of polymerase than RT.
[0230] In one other embodiment, PEgRNA is represented by Figure 3B, showing PEgRNA having a 3' extension arm, a spacer, and a gRNA core. The 3' extension further includes a reverse transcriptase template and a primer binding site in the 5'→3' direction. As shown, the reverse transcriptase template may also be more broadly referred to as a "DNA synthesis template," where the polymerase of the prime editing factor described herein is a different type of polymerase than RT.
[0231] In yet another embodiment, PEgRNA is represented by Figure 3D, showing PEgRNA having a spacer (1), a gRNA core (2), and an elongation arm (3) in the 5'→3' direction. The elongation arm (3) is located at the 3' end of the PEgRNA. The elongation arm (3) further includes a "primer binding site" (A), an "editing template" (B), and a "homologous arm" (C) in the 5'→3' direction. The elongation arm (3) may also include any modification regions at the 3' and 5' ends, which may be the same sequence or different sequences. In addition, the 3' end of the PEgRNA may include a transcription terminator sequence. These sequence elements of PEgRNA are further described and defined herein.
[0232] In another embodiment, PEgRNA is represented by Figure 3E, which shows PEgRNA having an elongation arm (3), a spacer (1), and a gRNA core (2) in the 5' to 3' direction. The elongation arm (3) is located at the 5' end of the PEgRNA. The elongation arm (3) further includes a "primer binding site" (A), an "editing template" (B), and a "homologous arm" (C) in the 3' to 5' direction. The elongation arm (3) may also include any modification regions at the 3' and 5' ends, which may be the same sequence or different sequences. The 3' end of the PEgRNA may include a transcription terminator sequence. These sequence elements of PEgRNA are further described and defined herein.
[0233] PE1 In this application, "PE1" refers to a PE complex containing a fusion protein comprising Cas9(H840A) and wild-type MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(wt)]+desired PEgRNA. The PE fusion has the amino acid sequence of SEQ ID NO: 123, which is shown below; [ka] [ka]
[0234] PE2 In this application, "PE2" refers to a PE complex containing a fusion protein comprising Cas9(H840A) and the variant MMLV RT having the following structure: [NLS]-[Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)]+desired PEgRNA. The PE fusion has the amino acid sequence of SEQ ID NO: 134, which is shown as follows: [ka] [ka]
[0235] PE3 In this application, "PE3" refers to PE2, plus a second-strand nicking guide RNA that complexes with PE2, and introduces a nick onto the unedited DNA strand to induce preferential replacement of the strand to be edited. PE3b
[0236] In this application, "PE3b" refers to PE3, but the second strand nicking guide RNA is designed for temporal control so that the second strand nick is not introduced until after the desired edit has been incorporated. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand and not the original allele. Using this strategy, hereafter referred to as PE3b, the mismatch between the protospacer and the unedited allele should not favor nicking by the sgRNA until after the editing event on the PAM strand occurs.
[0237] Short PE (PE-short) The term used in this application Short PE " refers to the PE construct fused to the C-terminal truncated reverse transcriptase, and has the following amino acid sequence: [ka] [ka]
[0238] Peptide tags The term "peptide tag" refers to a peptide amino acid sequence that genetically fuses with a protein sequence to confer one or more functions to the protein, thereby facilitating the manipulation of proteins for various purposes such as visualization, purification, solubilization, and separation. Peptide tags can encompass various types of tags classified by purpose or function, and may include "affinity tags" (to facilitate protein purification), "solubilization tags" (to assist in the correct folding of proteins), "chromatographic tags" (to alter the chromatographic properties of proteins), "epitope tags" (to bind to high-affinity antibodies), and "fluorescent tags" (to facilitate the visualization of proteins in cells or in vitro).
[0239] polymerase As used herein, the term “polymerase” refers to an enzyme that synthesizes nucleotide chains, which may be used in relation to the prime editing factor systems described herein. A polymerase may be a “template-dependent” polymerase (i.e., a polymerase that synthesizes nucleotide chains based on the order of nucleotide bases of a template chain). A polymerase may also be a “template-independent” polymerase (i.e., a polymerase that synthesizes nucleotide chains without the requirement of a template chain). A polymerase may further be classified as a “DNA polymerase” or an “RNA polymerase”. In various embodiments, the prime editing factor system includes a DNA polymerase. In various embodiments, the DNA polymerase may be a “DNA-dependent DNA polymerase” (i.e., the template molecule is a strand of DNA). In such cases, the DNA template molecule may be PEgRNA, and the extension arm includes a strand of DNA. In such cases, PEgRNA may be called a chimeric or hybrid PEgRNA, which comprises an RNA portion (i.e., a guide RNA component encompassing a spacer and a gRNA core) and a DNA portion (i.e., an elongation arm). In various other embodiments, DNA polymerase may be an "RNA-dependent DNA polymerase" (i.e., the template molecule is a strand of RNA). In such cases, PEgRNA is RNA, i.e., encompasses RNA elongation. The term "polymerase" may also refer to an enzyme that catalyzes the polymerization of nucleotides (i.e., polymerase activity). Generally, the enzyme will initiate synthesis at the 3' end of a primer annealed to a polynucleotide template sequence (e.g., a primer sequence annealed to the primer-binding site of PEgRNA, for example) and proceed toward the 5' end of the template strand. "DNA polymerase" catalyzes the polymerization of deoxynucleotides. The term DNA polymerase as used in this application with reference to DNA polymerase encompasses "its functional fragment". "The functional fragment" refers to a portion of either wild-type or mutant DNA polymerase that encapsulates less than the entire amino acid sequence of the polymerase, while retaining the ability to catalyze polynucleotide polymerization under at least one set of conditions.Such functional fragments may exist as separate entities, or they may be components of larger polypeptides such as fusion proteins.
[0240] Prime Edit As used in this application, the term "prime editing" refers to a novel approach to gene editing that utilizes a specialized guide RNA comprising napDNAbp, polymerase (e.g., reverse transcriptase), and a DNA synthesis template for encoding (or deleting) desired new genetic information to be incorporated into a target DNA sequence. Certain aspects of prime editing are described in other figures, specifically in Figures 1A-1H and 72(a)-72(c).
[0241] Prime editing represents a completely new platform for genome editing, a versatile and precise genome editing method that directly writes new genetic information to a defined DNA site, using a nucleic acid-programmable DNA-binding protein ("napDNAbp") that works in conjunction with polymerase (i.e., provided in the form of a fusion protein or otherwise in trans with napDNAbp). The prime editing system is programmed by prime editing (PE) guide RNA ("PEgRNA"), which defines the target site and serves as a template for the synthesis of the desired edit in the form of a replacement DNA strand by an manipulated extension (either DNA or RNA) on the guide RNA (e.g., at the 5' or 3' end or in the interior of the guide RNA). The replacement strand containing the desired edit (e.g., a single nucleic acid substitution) shares (or is homologous to) the same sequence as the endogenous strand of the target site to be edited (immediately downstream of the nick site) (with the exception that it contains the desired edit). By DNA repair and / or replication mechanisms, the endogenous strand downstream of the nick site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing can be considered a “search and replace” genome editing technology, because the prime editing factors described herein not only search for and locate the desired target site to be edited, but also simultaneously encode a replacement strand containing the desired edit that is incorporated in place of the endogenous DNA strand at the corresponding target site. The prime editing factors of this disclosure relate, in part, to the discovery that the mechanism of primed reverse transcription (TPRT) or “prime editing” to a prime target can be utilized or adapted to perform precise CRISPR / Cas-based genome editing with high efficiency and gene flexibility (as illustrated, for example, in various aspects of Figure 1A–1F). In nature, TPRT is used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial group II introns. 28,29In this application, the inventors used a Cas protein-reverse transcriptase fusion or related system to target a specific DNA sequence with a guide RNA, generate a single-stranded nick at the target site, and use the nicked DNA as a primer for reverse transcription of the manipulated reverse transcriptase template incorporated into the guide RNA. However, while the concept begins with a prime editing factor that uses reverse transcriptase as a DNA polymerase component, the prime editing factors described herein are not limited to reverse transcriptase and may substantially encompass the use of any DNA polymerase. Indeed, although this application may refer to a prime editing factor having “reverse transcriptase” everywhere, it is hereby made clear that reverse transcriptase is only one type of DNA polymerase that can function in prime editing. Therefore, whenever this specification refers to “reverse transcriptase,” those skilled in the art should understand that any suitable DNA polymerase may be used instead of reverse transcriptase. Therefore, in one aspect, the prime editing factor may include Cas9 (or equivalent napDNAbp), which is programmed to target a DNA sequence by associating it with a specialized guide RNA (i.e., PEgRNA) containing a spacer sequence that anneals to a complementary protospacer on the target DNA. The specialized guide RNA also contains new genetic information in the form of an elongation encoding a replacement strand of DNA containing the desired genetic alteration, which is used to replace the corresponding endogenous DNA strand at the target site. To transfer the information from the PEgRNA to the target DNA, the prime editing mechanism involves nicking a target site on one strand of DNA to expose a 3' hydroxyl group. The exposed 3' hydroxyl group can then be used to prime the DNA polymerization of the edit-encoding elongation on the PEgRNA to the target site. In various embodiments, the elongation providing a template for the polymerization of the edit-containing replacement strand may be formed from RNA or DNA. In the case of RNA elongation, the polymerase of the prime editing factor may be an RNA-dependent DNA polymerase (e.g., reverse transcriptase).In the case of DNA elongation, the polymerase of the prime editing factor may be a DNA-dependent DNA polymerase. The newly synthesized strand formed by the prime editing factors disclosed herein (i.e., the replacement DNA strand containing the desired edit) will be homologous to the genomic target sequence (i.e., have the same sequence) with the exception of the inclusion of the desired nucleotide change (e.g., a single nucleotide change, deletion, or insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA may also be called a single-stranded DNA flap. This will compete for hybridization with a complementary homologous endogenous DNA strand, thereby replacing the corresponding endogenous strand. In some embodiments, the system may be combined with the use of an error-prone reverse transcriptase enzyme (e.g., provided as a fusion protein with a Cas9 domain, or provided trans to a Cas9 domain). The error-prone reverse transcriptase enzyme may introduce alterations during the synthesis of the single-stranded DNA flap. Therefore, in some embodiments, the error-prone reverse transcriptase enzyme may be used to introduce nucleotide changes into the target DNA. Depending on the error-prone reverse transcriptase used in the system, the changes can be random or non-random. Degradation of the hybridized intermediate (including a single-stranded DNA flap synthesized by a reverse transcriptase hybridized to an endogenous DNA strand) may encompass removal of the replaced flap resulting from the endogenous DNA (e.g., by the 5' end DNA flap endonuclease FEN1), ligation of the synthesized single-stranded DNA flap onto the target DNA, and assimilation of the desired nucleotide changes as a result of cellular DNA repair and / or replication processes. Since template-based DNA synthesis offers one-nucleotide precision for any nucleotide modification, including insertions and deletions, the scope of this approach is extremely broad, and foreseeably, it could have countless applications in basic research and therapeutics.
[0242] In various embodiments, prime editing operates by bringing a target DNA molecule (to which a nucleotide sequence change is to be introduced) into contact with a nucleic acid programmed DNA-binding protein (napDNAbp) complexed with a prime editing guide RNA (PEgRNA). Referring to Figure 1G, the prime editing guide RNA (PEgRNA) includes an extension at the 3' or 5' end of the guide RNA or at an intramolecular position of the guide RNA, encoding the desired nucleotide change (e.g., a single nucleotide change, insertion, or deletion). In step (a), the napDNAbp / extended gRNA complex contacts the DNA molecule, and the extended gRNA guides the napDNAbp to bind to the target locus. In step (b), a nick is introduced into one of the strands of DNA at the target locus (e.g., by a nuclease or chemical agent), thereby generating an available 3' end into one of the strands at the target locus. In one embodiment, the nick is generated on the DNA strand corresponding to the R-loop strand, i.e., the strand that does not hybridize to the guide RNA sequence, i.e., the “non-target strand”. However, the nick can be introduced on either strand. That is, the nick can be introduced on the R-loop “target strand” (i.e., the strand that hybridizes to the protospacer of the extended gRNA) or the “non-target strand” (i.e., the strand that forms the single-stranded portion of the R-loop that is complementary to the target strand). In step (c), the 3' end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA to prime the reverse transcription (i.e., “RT primed to the prime target”). In one embodiment, the 3' end DNA strand hybridizes to a specific RT prime sequence on the extended portion of the guide RNA, i.e., the “reverse transcriptase prime sequence” or “primer binding site” on the PEgRNA. In step (d), the reverse transcriptase (or other preferred DNA polymerase) is introduced. This synthesizes a single strand of DNA from the 3' end of the primed region toward the 5' end of the prime editing guide RNA. DNA polymerase (e.g., reverse transcriptase) can be fused to napDNAbp, or alternatively, provided trans to napDNAbp.This forms a single-stranded DNA flap containing a desired nucleotide change (e.g., a single nucleotide change, insertion, or deletion, or a combination thereof) that is otherwise homologous to the endogenous DNA at or adjacent to the nicking site. In step (e), napDNAbp and guide RNA are released. Steps (f) and (g) relate to the degradation of the single-stranded DNA flap, resulting in the integration of the desired nucleotide change into the target locus. This process can be driven toward the formation of the desired product by removing the corresponding 5' endogenous DNA flap that forms once the 3' single-stranded DNA flap invades and hybridizes with the endogenous DNA sequence. Without being constrained by theory, the cell's endogenous DNA repair and replication processes degrade mismatched DNA and incorporate nucleotide changes (one or more) to form the desired modified product. The process can also be driven toward product formation by "second-strand nicking," as illustrated in Figure 1F. This process can introduce at least one of the following genetic changes: transversion, transition, deletion, and insertion.
[0243] The terms “prime editing factor (PE) system” or “prime editing factor (PE)” or “PE system” or “PE editing system” refer to compositions relating to genome editing methods using primed reverse transcription (TPRT) to a prime target as described herein, and include, but are not limited to, napDNAbp, reverse transcriptase, fusion proteins (e.g., napDNAbp and reverse transcriptase), prime editing guide RNA, and complexes comprising the fusion protein and prime editing guide RNA, as well as auxiliary elements, such as second-strand nicking components (e.g., second-strand sgRNA) and 5' endogenous DNA flap removal endonucleases (e.g., FEN1) to help drive the prime editing process toward the formation of the edited product.
[0244] In the embodiments described so far, PEgRNA constitutes a single molecule comprising a guide RNA (which itself contains a spacer sequence and a gRNA core or backbone) and a 5' or 3' elongation arm containing a primer binding site and a DNA synthesis template (see, for example, Figure 3D). However, PEgRNA can also take the form of two individual molecules, consisting of a guide RNA and a transprime editing factor RNA template (tPERT). This essentially houses the elongation arm (in particular, containing the primer binding site and DNA synthesis domain) and the RNA-protein recruitment domain (e.g., an MS2 aptamer or hairpin) on the same molecule. This then colocalizes or recruits to a modified prime editing factor complex containing a tPERT recruiting protein (e.g., an MS2cp protein that binds to an MS2 aptamer). See Figures 3G and 3H for examples of tPERTs that can be used for prime editing.
[0245] Prime editing factor The term “prime editing factor” refers to the fusion construct described herein, comprising napDNAbp (e.g., Cas9 nickase) and reverse transcriptase, which is capable of performing prime editing on a target nucleotide sequence in the presence of PEgRNA (or “extended guide RNA”). The term “prime editing factor” may also refer to the fusion protein, or the fusion protein complexed with PEgRNA, and / or the fusion protein further complexed with sgRNA to nick the second strand. In some embodiments, the prime editing factor may also refer to a complex comprising the fusion protein (reverse transcriptase fused with napDNAbp), PEgRNA, and regular guide RNA, which is capable of proceeding to the nicking step to the second site of the unedited strand as described herein. In other embodiments, the reverse transcriptase component of the “primer editing factor” may be provided in trans.
[0246] Primer binding site The term “primer binding site” or “PBS” refers to a nucleotide sequence located on PEgRNA as a component of the elongation arm (typically at the 3' end of the elongation arm) that helps bind to a primer sequence formed after Cas9 nicking of the target sequence by a prime editing factor. As detailed elsewhere, when the Cas9 nickase component of the prime editing factor nicks one strand of the target DNA sequence, a 3' ssDNA flap is formed, which anneals to the primer binding site on the PEgRNA and acts as a primer sequence that primes reverse transcription. Figures 27 and 28 show the configurations of primer binding sites located on the 3' and 5' elongation arms, respectively.
[0247] promoter The term “promoter” is recognized in the art and refers to a nucleic acid molecule having a sequence that is recognized by a cellular transcription mechanism and has the ability to initiate the transcription of a downstream gene. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is active only in the presence of specific conditions. For example, a conditional promoter may be active only in the presence of a specific protein that connects a protein associated with a regulatory element on the promoter to the basic transcription mechanism, or only in the absence of an inhibitory molecule. A subclass of conditionally active promoters is inducible promoters, which require the presence of a small molecule “inducer” for activity. Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. Various constitutive, conditional, and inducible promoters are well known to those skilled in the art, and those skilled in the art will be able to identify various such promoters useful for carrying out the present invention. This is not limited to this.
[0248] Protospacer As used in this application, the term "protospacer" refers to a sequence (~20 bp) on DNA adjacent to a PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand of it, i.e., the "target strand" relative to the "non-target strand" of the target DNA sequence). For Cas9 to function, it also requires a specific protospacer adjacent motif (PAM), which varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes a PAM sequence called NGG on the non-target strand, found directly downstream of the target sequence on the genomic DNA. Those skilled in the art will understand that the literature of the current art sometimes refers to "protospacer" as a ~20 nt target-specific guide sequence on the guide RNA itself, rather than simply calling it a "spacer." Therefore, in some cases, the term "protospacer" as used in this application may be interchangeable with the term "spacer." The context surrounding the appearance of either "protospacer" or "spacer" will help inform the reader whether the term refers to a gRNA or DNA target.
[0249] Protospacer adjacent motif (PAM) As used in this application, the term “protospacer adjacent sequence” or “PAM” refers to a DNA sequence of approximately 2–6 base pairs that is a key targeting component of the Cas9 nuclease. Typically, the PAM sequence is located on either strand and downstream in the 5' to 3' direction of the site cut by Cas9. A standard PAM sequence (i.e., the PAM sequence associated with the Cas9 nuclease or SpCas9 of Streptococcus pyogenes) is 5'-NGG-3', where “N” is any nucleic acid base followed by two guanine (”G") nucleic acid bases. Different PAM sequences may be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., SpCas9, may be modified to alter its PAM specificity by causing the nuclease to recognize alternative PAM sequences.
[0250] For example, with reference to the standard SpCas9 amino acid sequence of SEQ ID NO: 18, the PAM sequence can be modified by introducing one or more mutations, including (a) D1135V, R1335Q, and T1337R "VQR variants" that alter PAM specificity to NGAN or NGNG, (b) D1135E, R1335Q, and T1337R "EQR variants" that alter PAM specificity to NGAG, and (c) D1135V, G1218R, R1335E, and T1337R "VRER variants" that alter PAM specificity to NGCG. In addition, the D1135E variant of standard SpCas9 still recognizes NGG, but it is more selective compared to the wild-type SpCas9 protein.
[0251] It will also be understood that Cas9 enzymes from different bacterial species (i.e., Cas9 orthologs) can have varying PAM specificities. For example, Cas9 from Staphylococcus aureus (SaCas9) recognizes NGRRT or NGRRN. In addition, Cas9 from Neisseria meningitis (NmCas) recognizes NNNNGATT. In another example, Cas9 from Streptococcus thermophilis (StCas9) recognizes NNAGAAW. And yet another example, Cas9 from Treponema denticola (TdCas) recognizes NAAAAC. These are examples and not meant to be limiting. Furthermore, it will be understood that non-SpCas9s bind to a variety of PAM sequences. This makes them useful when a suitable SpCas9 PAM sequence is not present at the desired target's cut site. Moreover, non-SpCas9s may possess other characteristics that make them more useful than SpCas9s. For example, Cas9 from Staphylococcus aureus (SaCas9) is about 1 kilobase smaller than SpCas9. Therefore, it can be packaged in adeno-associated viruses (AAVs). Further reference can be made to Shah et al., “Protospacer recognition motifs: mixed identities and functional diversity,” RNA Biology, 10(5):891-899 (which is incorporated into this application by reference). Recombinase
[0252] As used in this application, the term "recombinase" refers to a site-specific enzyme that mediates the recombination of DNA between recombinase-recognition sequences, resulting in the excision, incorporation, inversion, or exchange (e.g., translocation) of DNA fragments between recombinase-recognition sequences. Recombinases can be classified into two distinct families: serine recombinases (e.g., resolvers and invertases) and tyrosine recombinases (e.g., integrases). Examples of serine recombinases include, without limitation, Hin, Gin, Tn3, β-six, CinH, ParA, γδ, Bxb1, φC31, TP901, TG1, φBT1, R4, φRV1, φFC1, MR11, A118, U153, and gp29. Examples of tyrosine recombinases include, without limitation, Cre, FLP, R, Lambda, HK101, HK022, and pSAM2. The names serine and tyrosine recombinases are derived from the conserved nucleophilic amino acid residues that recombinases use to attack DNA and which become covalently ligated to the DNA during strand exchange. Recombinases have numerous applications, including gene knockout / knock-in production and gene therapy applications. For example, Brown et al., “Serine recombinases as tools for genome engineering.”Methods.2011;53(4):372-9;Hirano et al., “Site-specific recombinases as tools for heterologous gene integration.”Appl.Microbiol.Biotechnol.2011;92(2):227-39;Chavez and Calos, “Therapeutic applications of the ΦC31 integrase system.”Curr.Gene Ther.2011;11(5):375-81;Turan and Bode,“Site-specific recombinases:from tag-and-target- to tag-and-exchange-based genomic modifications.”FASEB J.2011;25(12):4088-107;Venken and Bellen,“Genome-wide manipulations of Drosophila melanogaster with transposons, Flp recombinase, and ΦC31 integrase.”Methods Mol.Biol.2012;859:203-28;Murphy,“Phage recombinases and their applications.”Adv.Virus Res.2012;83:367-414;Zhang et al.,“Conditional gene manipulation: Cre-ating a new biological era.”J.Zhejiang Univ.Sci.B.2012;13(7):511-24;Karpenshif and Bernstein,“From yeast to mammals:recent advances in genetic control of homologous recombination.”DNA See Repair(Amst).2012;1;11(10):781-8 (the full contents of each are incorporated herein by reference). The recombinases provided herein are not intended to be exclusive examples of recombinases that may be used in embodiments of the present invention. The methods and compositions of the present invention may be extended by mining a database of novel orthogonal recombinases or by designing synthetic recombinases with defined DNA specificity (for example, Groth et al., “Phage integrases: biology and applications.” J.Mol.Biol.2004;335,667-678; Gordley et al., “Synthesis of programmable integrases.” Proc.Natl.Acad.Sci.US A.See 2009;106,5053-5058 (the full contents of each are incorporated herein by reference). Other examples of recombinases useful for the methods and compositions described herein are known to those skilled in the art, and it is expected that any new recombinases discovered or produced will be capable of being used in different embodiments of the present invention. In some embodiments, the catalytic domain of the recombinase is fused to a nuclease (e.g., dCas9 or a fragment thereof) that is programmable by RNA inactivating the nuclease, and as a result, the recombinase domain does not contain a nucleic acid binding domain or is incapable of binding to a target nucleic acid (e.g., the recombinase domain is manipulated so that it does not have specific DNA binding activity). Recombinases lacking DNA binding activity and methods for modifying them are known, including Klippel et al., "Isolation and characterization of unusual gin mutants." EMBO J. 1988; 7: 3983-3989; Burke et al., "Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation." Mol Microbiol. 2004; 51: 937-948; Olorunniji et al., "Synapsis and catalysis by activated Tn3 resolvase mutants." Nucleic Acids Res. 2008; 36: 7181-7191; Rowland et al., "Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome." Mol Microbiol. 2009; 74: 282-298; Akopian et al., "Chimeric recombinases with designed DNA sequence recognition.”Proc Natl Acad Sci USA.2003;100:8688-8691;Gordley et al.,“Evolution of programmable zinc finger-recombinases with activity in human cells.J Mol Biol.2007;367:802-813;Gordley et al.,“Synthesis of programmable integrases.”Proc Natl Acad Sci USA.2009;106:5053-5058;Arnold et al.,“Mutants of Tn3 resolvase which do not require accessory binding sites for recombination activity.”EMBO J.1999;18:1407-1414;Gaj et al.,“Structure-guided reprogramming of serine recombinase DNA sequence specificity.”Proc Natl Acad Sci USA.2011;108(2):498-503; and Proudfoot et al., “Zinc finger This includes those described in “Recombinases with adaptable DNA sequence specificity.” PLoS One. 2011;6(4):e19537 (the full contents of each are incorporated herein by reference). For example, serine recombinases of the resolverase-invertase group, such as Tn3 and γδ resolvers, and Hin and Gin invertases, have molecular structures with autonomous catalytic and DNA-binding domains (for example, Grindley et al., “Mechanism of site-specific recombination.” Ann Rev Biochem.See 2006;75:567-605 (its entire contents are incorporated by reference). Therefore, after isolating "activated" recombinase mutants that do not require any auxiliary factors (e.g., DNA binding activity), the catalytic domains of these recombinases are compliant with programmable nucleases (e.g., dCas9 or its fragments) by RNA that has inactivated the nuclease, as described herein (for example, Klippel et al., “Isolation and characterisation of unusual gin mutants.” EMBO J.1988;7:3983-3989; Burke et al., “Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation.” Mol Microbiol.2004;51:937-948; Olorunniji et al., “Synapsis and catalysis by activated Tn3 resolvase mutants.” Nucleic Acids Res.2008;36:7181-7191; Rowland et al. al., “Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome.”Mol Microbiol.2009;74:282-298;Akopian et al., “Chimeric recombinases with designed DNA sequence recognition.”Proc Natl Acad Sci USA.See 2003;100:8688-8691). In addition, many other natural serine recombinases are known that have an N-terminal catalytic domain and a C-terminal DNA-binding domain (e.g., Phi C31 integrase, TnpX transposese, IS607 transposese), and their catalytic domains can be selected to operate programmable site-specific recombinases as described herein (see, for example, Smith et al., “Diversity in the serine recombinases.” Mol Microbiol. 2002;44:299-307 (its entire contents are incorporated by reference)). Similarly, core catalytic domains of tyrosine recombinases (e.g., Cre, λ-integrase) are known and can be similarly selected to manipulate programmable site-specific recombinases as described herein (for example, Guo et al., “Structure of Cre recombinase complexed with DNA in a site-specific recombination synapse.” Nature. 1997; 389: 40-46; Hartung et al., “Cre mutants with altered DNA binding properties.” J Biol Chem 1998; 273: 22884-22891; Shaikh et al., “Chimeras of the Flp and Cre recombinases: Tests of the mode of cleavage by Flp and Cre.” J Mol Biol. 2000; 302: 27-48; Rongrong et al., “Effect of deletion mutation on the recombination activity of Cre recombinase.” Acta Biochim Pol.2005;52:541-544;Kilbride et al., “Determinants of product topology in a hybrid Cre-Tn3 resolvase site-specific recombination system.”J Mol Biol.2006;355:185-195;Warren et al.,“A chimeric cre recombinase with regulated directionality.”Proc Natl Acad Sci USA.2008 105:18278-18283;Van Duyne,“Teaching Cre to follow directions.”Proc Natl Acad Sci USA.2009 Jan 6;106(1):4-5;Numrych et al.,“A comparison of the effects of single-base and triple-base changes in the integrase arm-type binding sites on the site-specific recombination of bacteriophage λ.”Nucleic Acids Res.1990;18:3953-3959;Tirumalai et al.,“The recognition of core-type DNA sites by λ integrase.”J Mol See Biol. 1998;279:513-527; Aihara et al., “A conformational switch controls the DNA cleavage activity of λ integrase.” Mol Cell. 2003;12:187-198; Biswas et al., “A structural basis for allosteric control of DNA recombination by λ integrase.” Nature. 2005;435:1059-1066; and Warren et al., “Mutations in the amino-terminal domain of λ-integrase have differential effects on integrative and excisive recombination.” Mol Microbiol. 2005;55:1104-1112 (the full content of each is incorporated by reference)).
[0253] Recombinase recognition sequence As used in this application, the terms “recombinase-recognized sequence” or equivalently “RRS” or “recombinase-targeted sequence” refer to a nucleotide sequence target that is recognized by a recombinase and undergoes strand exchange with another DNA molecule having an RRS. This results in the excision, incorporation, inversion, or exchange of DNA fragments between the recombinase-recognized sequences.
[0254] Recombine or rearrange The term “recombination” is used in the context of nucleic acid modification (e.g., genome modification) to refer to the process by which two or more nucleic acid molecules or two or more regions of a single nucleic acid molecule are modified by the action of a recombinase protein (e.g., the recombinase fusion protein of the present invention provided herein). Recombination can result in, for example, insertion, inversion, excision, or translocation of nucleic acids on or between one or more nucleic acid molecules. Recombinase recognition sequence
[0255] reverse transcriptase The term “reverse transcriptase” describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require primers to synthesize DNA transcripts from an RNA template. Historically, reverse transcriptases have been primarily used to transcribe mRNA into cDNA, which can then be cloned onto vectors for further manipulation. Myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme possesses 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific to the RNA strand of RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Known viral reverse transcriptases lack the 3'-5' exonuclease activity necessary for proofreading, so transcriptional errors cannot be corrected by reverse transcriptase (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). Detailed studies of AMV reverse transcriptase activity and its associated RNase H activity are presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase widely used in molecular biology is that derived from Moloney's mouse leukemia virus (M-MLV). See, for example, Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML, et al., Gene 35:249-258 (1985). M-MLV reverse transcriptases that substantially lack RNase H activity have also been described. For example, see USPat. No. 5,244,797.The present invention intends to utilize any such reverse transcriptase or its variants or variants.
[0256] In addition, the present invention intends to use a reverse transcriptase that is prone to errors, i.e., a reverse transcriptase that may be called an error-prone reverse transcriptase, or a reverse transcriptase that does not support high-fidelity incorporation of nucleotides during polymerization. During the synthesis of a single-stranded DNA flap based on guide RNA and an incorporated RT template, an error-prone reverse transcriptase may introduce one or more nucleotides that are mismatched with the RT template sequence, thereby introducing changes in the nucleotide sequence through the erroneous polymerization of the single-stranded DNA flap. These errors introduced during the synthesis of the single-stranded DNA flap are then incorporated into the double-stranded molecule by hybridization to the corresponding endogenous target strand, removal of the replaced endogenous strand, ligation, and then another round of the endogenous DNA repair and / or sequencing process.
[0257] Reverse transcription As used in this application, the term "reverse transcription" refers to the ability of an enzyme to synthesize a DNA strand (i.e., complementary DNA or cDNA) using RNA as a template. In some embodiments, reverse transcription may be "error-prone reverse transcription." This refers to the characteristics of certain reverse transcriptases in which their DNA polymerization activity is prone to errors.
[0258] PACE The term "phage-assisted sequential evolution (PACE)" as used in this application refers to sequential evolution using a phage as a viral vector. The general concept of PACE technology is illustrated, for example, by the international PCT application PCT / US2009 / 056194 filed on September 8, 2009, published as WO 2010 / 028347 on March 11, 2010; the international PCT application PCT / US2011 / 066747 filed on December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; the US application, US No. 9,023,594 issued on May 5, 2015; and the international PCT application PCT / US2015 / 012022 filed on January 20, 2015, published as WO 2012 / 088381 on September 11, 2015. Published as 2015 / 134121; and described in the international PCT application PCT / US2016 / 027795 filed on 15 April 2016, published as WO 2016 / 168631 on 20 October 2016 (the contents of each of these in whole are incorporated herein by reference).
[0259] Phage In this application, the term "phage," used interchangeably with the term "bacteriophage," refers to a virus that infects bacterial cells. Typically, a phage consists of an outer protein capsid containing genetic material. The genetic material may be linear or circular ssRNA, dsRNA, ssDNA, or dsDNA. Phages and phage vectors are well known to those skilled in the art, and some, though not limited, examples of phages useful for performing the PACE method provided herein are λ (lysogen), T2, T4, T7, T12, R17, M13, MS2, G4, P1, P2, P4, phi X174, N4, Φ6, and Φ29. In one embodiment, the phage utilized in the present invention is M13. Additional suitable phages and host cells will be apparent to those skilled in the art. The present invention is not limited in this respect.For further examples of suitable phages and host cells, please refer to: Elizabeth Kutter and Alexander Sulakvelidze: Bacteriophages: Biology and Applications. CRC Press; 1st edition (December 2004), ISBN: 0849313368; Martha RJ Clokie and Andrew M. Kropinski: Bacteriophages: Methods and Protocols, Volume 1: Isolation, Characterization, and Interactions (Methods in Molecular Biology). Humana Press; 1st edition (December, 2008), ISBN: 1588296822; Martha RJ Clokie and Andrew M. Kropinski: Bacteriophages: Methods and Protocols, Volume 2: Molecular and Applied Aspects (Methods in Molecular Biology). Humana Press; 1st edition (December See (2008), ISBN:1603275649; all of these, with regard to the disclosure of suitable phages and host cells, as well as methods and protocols for the isolation, culture, and manipulation of such phages, are incorporated herein by reference in their entirety).
[0260] Proteins, peptides, and polypeptides The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acid long. A protein, peptide, or polypeptide may refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified by the addition of a chemical entity, such as a carbohydrate group, hydroxyl group, phosphate group, farnesyl group, isofarnesyl group, fatty acid group, linker for conjugation, functionalization, or other modification. A protein, peptide, or polypeptide may also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide may be merely a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or any combination thereof. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012) (the entire contents of which are incorporated herein by reference)).
[0261] Protein splicing As used in this application, the term "protein splicing" refers to the process in which a sequence of intein (or, as may be, fission intein) is excised from an amino acid sequence, and the remaining fragment of the amino acid sequence, the extein, is ligated by an amide bond to form a continuous amino acid sequence. The term "trans" protein splicing refers to the specific case in which the intein is fission intein and they are located on different proteins.
[0262] Second-chain nicking The degradation of heteroduplex DNA (i.e., containing one edited and one unedited strand) formed as a result of prime editing determines the long-term editing outcome. In words, the goal of prime editing is to degrade the heteroduplex DNA (the edited strand paired with the endogenous unedited strand) formed as an intermediate of priming by permanently incorporating the edited strand onto the complementary endogenous strand. To help drive the degradation of heteroduplex DNA in a manner favorable to the permanent incorporation of the edited strand onto the DNA molecule, a “second strand nicking” approach may be used in this application. The concept of “second strand nicking” as used in this application refers to the introduction of a second nick, preferably on the unedited strand, downstream of the first nick (i.e., the initial nick site providing the free 3' end for use of reverse transcriptase on the extended portion of the guide RNA for priming). In one embodiment, the first and second nicks are on opposing strands. In other embodiments, the first and second nicks are on opposing strands. In yet another embodiment, the first nick is on the non-target strand (i.e., the strand forming the single-stranded portion of the R-loop), and the second nick is on the target strand. In yet another embodiment, the first nick is on the strand being edited, and the second nick is on the strand not being edited. The second nick may be located at least 5 nucleotides downstream of the first nick, or at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 nucleotides downstream of the first nick, or more.In one embodiment, the second nick may be introduced on the unedited strand at a distance of approximately 5–150 nucleotides, or approximately 5–140, or approximately 5–130, or approximately 5–120, or approximately 5–110, or approximately 5–100, or approximately 5–90, or approximately 5–80, or approximately 5–70, or approximately 5–60, or approximately 5–50, or approximately 5–40, or approximately 5–30, or approximately 5–20, or approximately 5–10 from the site of the PEgRNA-induced nick. In one embodiment, the second nick is introduced at a distance of 14–116 nucleotides from the PEgRNA-induced nick. Without being constrained by theory, the second nick directs the cell's endogenous DNA repair and replication processes toward replacement or editing of the unedited strand, thereby permanently incorporating the edited sequence on both strands and degrading the heteroduplex formed as a result of PE. In some embodiments, the edited strand is the non-target strand, and the unedited strand is the target strand. In other embodiments, the edited strand is the target strand, and the unedited strand is the non-target strand.
[0263] Sense Chain In genetics, the "sense" strand is a segment of double-stranded DNA extending from 5' to 3' that is complementary to the antisense or template strand of DNA extending from 3' to 5'. In the case of a protein-coding DNA segment, the sense strand is the DNA strand that has the same sequence as the mRNA that takes the antisense strand as its template during transcription and eventually (typically, but not always) undergoes translation to become a protein. Thus, the antisense strand carries the RNA that is later translated into a protein, while the sense strand has a makeup that is nearly identical to that of the mRNA. Note that for each segment of dsDNA, there will likely be two sets of sense and antisense strands, depending on the direction from which one is read (since sense and antisense are relative to viewpoints). The gene product or mRNA will ultimately indicate that one of the strands of one of the segments of the dsDNA is referred to as either sense or antisense.
[0264] In the context of PEgRNA, the first step is the synthesis of a single-strand complementary DNA (i.e., the 3'ssDNA flap that will be incorporated) oriented in the 5'→3' direction, but the synthesized strand is removed from the template of the PEgRNA elongation arm. Whether the 3'ssDNA flap should be considered a sense strand or an antisense strand depends on the direction of transcription, where it is quite permissible for both strands of DNA to function as templates for transcription (but not simultaneously). Therefore, in some embodiments, the 3'ssDNA flap (extending overall in the 5'→3' direction) will function as a sense strand, since it is the coding strand. In other embodiments, the 3'ssDNA flap (extending overall in the 5'→3' direction) will function as an antisense strand and thus be a template for transcription.
[0265] Spacer array As used herein, the term “spacer sequence,” in relation to guide RNA or PEgRNA, refers to a portion of guide RNA or PEgRNA of approximately 20 nucleotides containing a nucleotide sequence complementary to the protospacer sequence in the target DNA sequence. The spacer sequence anneals with the protospacer sequence to form an ssRNA / ssDNA hybrid structure at the target site, and a corresponding R-loop ssDNA structure of the endogenous DNA strand complementary to the protospacer sequence.
[0266] subject The term "subject," as used herein, refers to an individual organism, for example, an individual mammal. In some aspects, the subject is a human. In some aspects, the subject is a non-human mammal. In some aspects, the subject is a non-human animal of the order Primates. In some aspects, the subject is a rodent. In some aspects, the subject is a sheep, goat, cattle, cattle, or dog. In some aspects, the subject is a vertebrate, amphibian, reptile, fish, insect, flying insect (fly), or nematode. In some aspects, the subject is a research animal. In some aspects, the subject is genetically modified, for example, a genetically modified non-human subject. The subject may be of any sex and at any developmental stage.
[0267] Splitting intein Inteins are most frequently found as a single, continuous domain, but some exist naturally in a split form. In this case, the two fragments are expressed as separate polypeptides and must associate before splicing occurs, a process known as protein trans-splicing.
[0268] The fission intein described is the Ssp DnaE intein, which contains two subunits, namely DnaE-N and DnaE-C. The two distinct subunits are encoded by separate genes, namely dnaE-n and dnaE-c, which encode the DnaE-N and DnaE-C subunits, respectively. DnaE is a naturally occurring fission intein in Synechocytis sp. PCC6803, and each can lead to the transsplicing of two distinct proteins, each containing a fusion with either DnaE-N or DnaE-C.
[0269] Additional naturally occurring or engineered split intein sequences are known or can be derived from the whole intein sequences described herein or those available in the art. Examples of split intein sequences can be found in Stevens et al., “A promiscuous split intein with expanded protein engineering applications,” PNAS, 2017, Vol. 114: 8538-8543; Iwai et al., “Highly efficient protein trans-splicing by a naturally split DnaE intein from Nostc punctiforme,” FEBS Lett, 580: 1853-1858. Each of these is incorporated herein by reference. Additional split intein sequences can be found, for example, in WO2013 / 045632, WO2014 / 055782, WO2016 / 069774, and EP2877490. The contents of each of these are incorporated herein by reference.
[0270] In addition, trans protein splicing has been described in vivo and in vitro (Shingledecker, et al.,
[0076] Gene 207:187 (1998), Southworth, et al., EMBO J.17:918 (1998); Mills, et al., Proc. Natl. Acad. Sci. USA, 95:3543-3548 (1998); Lew, et al., J. Biol. Chem., 273:15887-15890 (1998); Wu, et al., Biochim. Biophys. Acta 35732:1 (1998b), Yamazaki, et al., J. Am. Chem. Soc. 120:5591 (1998), Evans, et al. (Al., J. Biol. Chem. 275:9091 (2000); Otomo, et al., Biochemistry 38:16040-16044 (1999); Otomo, et al., J. Biomol. NMR 14:105-114 (1999); Scott, et al., Proc. Natl. Acad. Sci. USA 96:13638-13643 (1999)), provides an opportunity to express a certain protein as two inactive fragments that subsequently undergo ligation to form a functional product. For example, the formation of a complete PE fusion protein from two separately expressed halves is shown in Figures 66 and 67.
[0271] target site The term “target site” refers to a sequence on a nucleic acid molecule edited by a prime editing factor (PE) disclosed herein. Furthermore, the target site refers to a sequence on a nucleic acid molecule to which a complex of prime editing factor (PE) and gRNA binds.
[0272] tPERT Refer to the definition of the trans-prime editing factor "RNA type (tPERT)".
[0273] Temporary nicking into the second strand As used herein, the term “temporary nicking into the second strand” refers to a variant of nicking into the second strand in which the incorporation of the second nick in the unedited strand occurs only after the desired edit has been incorporated into the edited strand. This avoids simultaneous nicking on both strands, which could lead to double-strand breaks. The guide RNA for nicking into the second strand is designed for temporary control to prevent the introduction of the second strand nick until after the incorporation of the desired edit. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand but not the original allele. Using this strategy, the mismatch between the protospacer and the unedited allele should deter nicking by the sgRNA until after the editing event on the PAM strand has occurred.
[0274] Transprime Editing As used herein, the term “trans-prime editing” refers to a modified form of prime editing that utilizes fission PEgRNA, where PEgRNA is separated into two separate molecules: sgRNA and a trans-prime editing RNA template (tPERT). While the sgRNA works to target the prime editing factor (or more generally, the napDNAbp component of the prime editing factor) to a desired genomic target site, once tPERT is recruited trans-to the prime editing factor by the interaction of the prime editing factor and a binding domain located on tPERT, tPERT is used by a polymerase (e.g., reverse transcriptase) to write the new DNA sequence into the target locus. In one embodiment, the binding domain may contain an RNA-protein recruitment moiety, such as an MS2 aptamer located on tPERT and an MS2cp protein fused with the prime editing factor. The advantage of trans-prime editing is that it allows the use of potentially longer templates by separating the DNA synthesis template from the guide RNA.
[0275] The trans-prime editing process is shown in Figures 3G and 3H. Figure 3G, on the left, shows the composition of the trans-prime editing factor complex ("RP-PE:gRNA complex"), which contains napDNAbp fused with polymerase (e.g., reverse transcriptase) and rPERT recruiting protein (e.g., MS2sc), respectively, and is complexed with guide RNA. Figure 3G further shows separate tPERT molecules, including the elongation arm features of PEgRNA containing the DNA synthesis template and primer binding sequence. The tPERT molecule also contains an RNA-protein recruiting domain (which in this case is a stem-loop structure and may be, for example, an MS2 aptamer). As depicted in the process shown in Figure 3H, the RP-PE:gRNA complex binds to the target DNA sequence and introduces the nick. Next, the recruiting protein (RP) recruits tPERT and colocalizes it to the prime editing factor complex bound to the DNA target site, thereby enabling the primer binding site to bind to the nicked primer sequence on the strand. Subsequently, polymerase (e.g., RT) is able to synthesize a single strand of DNA up to 5' of tPERT using the DNA synthesis template.
[0276] Although tPERT is shown in Figures 3G and 3H to contain PBS and DNA synthesis templates at the 5' end of the RNA-protein recruitment domain, tPERTs with other configurations may also be designed to have PBS and DNA synthesis templates located at the 3' end of the RNA-protein recruitment domain. However, tPERTs with a 5' extension have the advantage that single-strand DNA synthesis will naturally terminate at the 5' end of the tPERT, thus eliminating the risk of using any part of the RNA-protein recruitment domain as a template during the DNA synthesis step of prime editing.
[0277] Transprime editing factor RNA template (tPERT) As used herein, “trans prime editing factor RNA template (tPERT)” refers to the component used in trans prime editing. The modified version of prime editing operates by separating PEgRNA into two distinct molecules: a guide RNA and a tPERT molecule. The tPERT molecule is programmed to co-localize with the prime editing factor complex at the target DNA site, bringing the primer binding site and DNA synthesis template trans to the prime editing factor. For example, see Figure 3G for an embodiment of trans prime editing factor (tPE) showing a two-component system including a tPERT containing (1) an RP-PE:gRNA complex and (2) a DNA synthesis template bound to the primer binding site and RNA-protein recruitment domain. Here, the RP (recruiting protein) component of the RP-PE:gRNA complex recruits tPERT to the target site to be edited, thereby trans-binding the PBS and DNA synthesis template to the prime editing factor. In other words, tPERT has been modified to include (all or part of) the elongation arm of PEgRNA, which includes the primer binding site and the DNA synthesis template.
[0278] Transition As used herein, “transition” refers to the interconversion of purine nucleobases (A⇔G) or pyrimidine nucleobases (C⇔T). This class of interconversions involves nucleobases of similar shape. The compositions and methods disclosed herein are capable of inducing one or more transitions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both transitions and transversions in the same target DNA molecule. These changes involve A⇔G, G⇔A, C⇔T, or T⇔C. In the context of double-stranded DNA with Watson-Crick pairs of nucleobases, transversion refers to the following base pair exchanges: A:T⇔G:C, G:G⇔A:T, C:G⇔T:A, or T:A⇔C:G. The compositions and methods disclosed herein are capable of inducing one or more transitions in a target DNA molecule. The compositions and methods disclosed herein can also induce both transitions and transversions, as well as other nucleotide changes, including deletions and insertions, in the same target DNA molecule.
[0279] Transversion As used herein, “transversion” refers to the interconversion of purine nucleic acid bases to pyrimidine nucleic acid bases, or vice versa, and thus involves the interconversion of dissimilar nucleic acid bases. These changes involve T⇔A, T⇔G, C⇔G, C⇔A, A⇔T, A⇔C, G⇔C, and G⇔T. In the context of double-stranded DNA with Watson-Crick pairs of nucleic acid bases, transversion refers to the following base pair exchanges: T:A⇔A:T, T:A⇔G:C, C:G⇔G:C, C:G⇔A:T, A:T⇔T:A, A:T⇔C:G, G:C⇔C:G, and G:C⇔T:A. The compositions and methods disclosed herein are capable of inducing one or more transversions in a target DNA molecule. The compositions and methods disclosed herein can also induce both transitions and transversions, as well as other nucleotide changes, including deletions and insertions, in the same target DNA molecule.
[0280] treatment The terms “treatment,” “to treat,” and “to treat” refer to a clinical intervention aimed at preventing, alleviating, delaying, or inhibiting the onset of a disease or disorder, or one or more of its symptoms, as described herein. When used herein, the terms “treatment,” “to treat,” and “to treat” refer to a clinical intervention aimed at preventing, alleviating, delaying, or inhibiting the onset of a disease or disorder, or one or more of its symptoms, as described herein. In some embodiments, treatment may be administered after the onset of one or more symptoms and / or after a diagnosis of the disease. In other embodiments, treatment may be administered in the absence of symptoms, for example, to prevent or delay the onset of symptoms or to inhibit the onset or progression of the disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (for example, in light of a history of symptoms and / or in light of genetic factors or other susceptibility factors). Treatment may also be continued after the symptoms have subsided, for example, to prevent or delay their recurrence.
[0281] Trinucleotide repeat disorder As used herein, “trinucleotide repeat disorders” (or alternatively, “extended repeat disorders” or “extended repeat disorders”) refer to a range of hereditary disorders caused by “trinucleotide repeat extensions,” which are a type of mutation (a trinucleotide repeat located in a gene or intron). Trinucleotide repeats were once considered common iterations in the genome, but were classified into these disorders in the 1990s. These seemingly “benign” stretches of DNA can sometimes expand and cause disease. Several key features are common to disorders caused by trinucleotide repeat extensions. First, the mutant repeats exhibit instability in both somatic and germline cells, and more frequently, they expand rather than shrink in successive transmissions. Second, early onset and (expected) increased phenotypic severity in subsequent generations generally correlate with longer repeat lengths. Finally, the parental origin of the disease allele can often influence the expectation that paternal inheritance poses a greater risk of expansion for many of these disorders.
[0282] Triplet expansion is thought to be caused by slippage during DNA replication. Due to the repetition of DNA sequences in these regions, "loop-out" structures can also form during DNA replication while maintaining complementary base pairings between the synthesizing parent and daughter strands. If the loop-out structure forms from the daughter strand, this will result in an increase in the number of repeats. However, if the loop-out structure forms on the parent strand, it results in a decrease in the number of repeats. These repeat expansions appear to be more common than reductions. Generally, the larger the expansion, the more likely they are to cause disease or increase the severity of the disease. This characteristic leads to the expected features seen in trinucleotide repeat disorders. The expectation explains the tendency for younger age of onset and increasing symptom severity across generations of affected families due to these repeat expansions.
[0283] Nucleotide repeat defects may include defects in which triple repeats occur in a non-coding region (i.e., non-coding trinucleotide repeat defects) or defects in which they occur in a coding region.
[0284] The prime editing factor (PE) systems described herein may be used to treat nucleotide repetition disorders, which may include, among others, fragile X syndrome (FRAXA), fragile XE MR (FRAXE), Friedreich's ataxia (FRDA), myotonic dystrophy (DM), spinocerebellar ataxia type 8 (SCA8), and spinocerebellar ataxia type 12 (SCA12).
[0285] Upstream As used herein, the terms “upstream” and “downstream” are relative terms defining the linear positions of at least two elements located within a nucleic acid molecule (whether single-stranded or double-stranded) oriented in the 5’→3’ direction. In particular, if the first element is located somewhere at 5’ relative to the second element, the first element is upstream of the second element in the nucleic acid molecule. For example, if the SNP is on the 5’ side of a nick site, the SNP is upstream of the Cas9-induced nick site. Conversely, if the first element is located somewhere at 3’ relative to the second element, the first element is downstream of the second element in the nucleic acid molecule. For example, if the SNP is on the 3’ side of a nick site, the SNP is downstream of the Cas9-induced nick site. Nucleic acid molecules can be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA. While the terms upstream and downstream refer only to single-stranded nucleic acid molecules unless it is necessary to choose which strand of a double-stranded molecule to consider, the analysis is often the same for single-stranded and double-stranded nucleic acid molecules, and can often be used to determine the relative positions of at least two elements. In double-stranded DNA, the strand is the "sense" strand or the "coding" strand. In genetics, the "sense" strand is the double-stranded DNA segment extending from 5' to 3' that is complementary to the antisense or template strand of the DNA extending from 3' to 5'. Thus, for example, if an SNP nucleic acid base is on the 3' side of the promoter on the sense or coding strand, the SNP nucleic acid base is "downstream" of the promoter sequence in genomic DNA (which is double-stranded).
[0286] variant As used herein, the term “variant” should be interpreted as meaning that an attribute exhibiting qualities having a pattern deviating from that which is naturally occurring; for example, variant Cas9 is Cas9 that contains one or more changes in amino acid residues compared to the wild-type Cas9 amino acid sequence. The term “variant” encompasses homologous proteins that have at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% percent identity with the reference sequence and have the same or substantially the same functional activity as the reference sequence. The term also encompasses mutants, truncations, or domains of the reference sequence that exhibit the same or substantially the same functional activity as the reference sequence.
[0287] vector The term “vector,” as used herein, refers to a nucleic acid that can be modified to encode a gene of interest and that can enter a host cell and mutate or replicate within the host cell (then transfer the replicated form of the vector to another host cell). Illustrated preferred vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phages, and conjugative plasmids. Additional preferred vectors will be apparent to those skilled in the art based on this disclosure.
[0288] Wild type As used herein, the term “wild type” is a technical term understood by those skilled in the art and means a typical form of an organism, strain, gene, or characteristic as it exists in nature, distinct from the form of a mutant or variant.
[0289] 5' endogenous DNA flap As used in this application, the term "endogenous 5' DNA flap" refers to the strand of DNA located immediately downstream of a nick site induced by PE on target DNA. Nicking of the target DNA strand by PE exposes the 3' hydroxyl group upstream of the nick site and the 5' hydroxyl group downstream of the nick site. The endogenous strand ending in the 3' hydroxyl group is used to prime the DNA polymerase of the prime editing factor (for example, DNA polymerase is a reverse transcriptase). The endogenous strand beginning in the exposed 5' hydroxyl group downstream of the nick site is called the "5' endogenous DNA flap" and is ultimately removed and replaced by a newly synthesized replacement strand encoded by the elongation of PEgRNA (i.e., the "3' replacement DNA flap").
[0290] 5' endogenous DNA flap removal As used herein, the terms “5' endogenous DNA flap removal” or “5' flap removal” refer to the removal of a 5' endogenous DNA flap formed when a single-strand DNA flap synthesized by RT competitively invades and hybridizes with endogenous DNA, thereby replacing the endogenous strand in the process. Removal of this replaced endogenous strand may drive a reaction that forms a desired product containing the desired nucleotide changes. The cell’s own DNA repair enzymes may catalyze the removal or excision of the 5' endogenous flap (e.g., flap endonucleases such as EXO1 or FEN1). Alternatively, the host cell may be transformed to express one or more enzymes that catalyze the removal of the 5' endogenous flap, thereby driving the process of product formation (e.g., flap endonucleases). Flap endonucleases are known in the art and can be found in Patel et al., “Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends,” Nucleic Acids Research, 2012, 40(10):4507-4519 and Tsutakawa et al., “Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily,” Cell, 2011, 145(2):198-211 (each of which is incorporated herein by reference).
[0291] 3' Replacement DNA Flap As used herein, the term “3' replacement DNA flap,” or simply “replacement DNA flap,” refers to a strand of DNA synthesized by a prime editing factor and encoded by the elongation arm of the prime editing factor PEgRNA. More specifically, the 3' replacement DNA flap is encoded by the polymerase template of PEgRNA. The 3' replacement DNA flap contains the same sequence as the 5' endogenous DNA flap, except that it also contains the edited sequence (e.g., a single nucleotide change). The 3' replacement DNA flap anneals with the target DNA to replace or replace the 5' endogenous DNA flap (which can be excised by a 5' flap endonuclease such as FEN1 or EXO1, for example), and the 3' end of the 3' replacement DNA flap is then ligated to the exposed (exposed after excision of the 5' endogenous DNA flap) 5' hydroxyl end of the endogenous DNA, thereby reforming the phosphodiester bond and incorporating the 3' replacement DNA flap, forming a heteroduplex DNA containing one edited strand and one unedited strand. The DNA repair process degrades the heteroduplex and permanently incorporates the edit into the DNA by copying the information from the edited strand to the complementary strand. This degradation process can be further advanced to completion by nicking the unedited strand, i.e., via "nicking the second strand" as described herein. [Modes for carrying out the invention]
[0292] Detailed description of a certain aspect The adoption of clustered, regularly arranged short palindromic repeat (CRISPR) systems for genome editing has revolutionized life sciences. 1~3While CRISPR-based gene disruption is now routine, precise incorporation of single-nucleotide edits remains a major challenge, despite their necessity in studying or correcting numerous disease-causing mutations. Homologous recombination repair (HDR), while capable of achieving such editing, suffers from low efficiency (often <5%), the requirement for a donor DNA repair template, and the adverse effects of double-strand break (DSB) formation. Recently, Professor David Liu and his colleagues developed a base-editing technique that achieves efficient single-nucleotide editing without DSBs. Base-editing factors (BEs) combine the CRISPR system with base-modifying deaminase enzymes to convert target C·G or A·T base pairs to A·T or G·C, respectively. 4~6 Although already widely used by researchers worldwide, current genome editing (BE) can only perform four of the 12 possible base pair changes and cannot correct small insertions or deletions. Furthermore, the target range of base editing is limited by editing non-target C or A bases adjacent to the target base ("bystander editing") and by the requirement that the PAM sequence be located 15±2 bp from the target base. Therefore, overcoming these shortcomings would greatly expand the basic research and therapeutic applications of genome editing.
[0293] This disclosure proposes a novel high-precision editing approach that provides many of the benefits of base editing—namely, avoidance of double-strand disruption and donor DNA repair templates—while overcoming its major drawbacks. The proposed approach described herein achieves direct incorporation of an edited DNA strand at a site in the target genome using targeted primed reverse transcription (TPRT). In the design considered herein, a CRISPR guide RNA (gRNA) would be modified to have a reverse transcriptase (RT) template sequence encoding single-strand DNA containing the desired nucleotide changes. Target site DNA nicked by a CRISPR nuclease (Cas9) would act as a primer for reverse transcription of the template sequence on the modified gRNA, enabling direct incorporation of any desired nucleotide edit.
[0294] Consequently, the present invention relates in part to the discovery that the mechanism of targeted primed reverse transcription (TPRT) or “prime editing” can be utilized or employed to perform highly efficient and genetically flexible high-precision CRISPR / Cas-based genome editing (as depicted, for example, in various embodiments in Figures 1A–1F). The inventors herein propose the use of a Cas9 protein-reverse transcriptase fusion that targets a specific DNA sequence with a modified guide RNA (“extended guide RNA”), generates a single-strand nick at the target site, and uses the nick-containing DNA as a primer for a modified reverse transcriptase template incorporated into the extended gRNA. The newly synthesized strand will be homologous to the target sequence in the genome, except that it contains the desired nucleotide alteration (e.g., a single nucleotide alteration, deletion, or insertion, or a combination thereof). The newly synthesized DNA strand, sometimes referred to as a single-strand DNA flap, competes for hybridization with a complementary homologous endogenous DNA strand, thereby replacing the corresponding endogenous strand. Degradation of this hybridized intermediate can involve the removal of the replaced flap (e.g., by the 5'-terminus DNA flap endonuclease, FEN1), ligation of the synthesized single-strand DNA flap to target DNA, and assimilation of desired nucleotide changes as a result of the cell's DNA repair and / or replication processes. Because templated DNA synthesis confers high precision to single nucleotides, the breadth of this approach is extremely broad, and it can be foreseen that it could be used in countless applications in basic science and therapeutics.
[0295] [1]napDNAbp The prime editing factors and transpriming editing factors described herein may include nucleic acid programmed DNA-binding proteins (napDNAbp).
[0296] In one aspect, napDNAbp can associate with or complex with at least one guide nucleic acid (e.g., guide RNA or PEgRNA). This localizes napDNAbp to a DNA sequence containing a DNA strand (i.e., the target strand) that is complementary to the guide nucleic acid or a portion of it (e.g., the spacer of the guide RNA that anneals to the protospacer of the DNA target). In other words, the guide nucleic acid "programs" napDNAbp (e.g., Cas9 or its equivalent) to localize and bind to the complementary sequence of the protospacer on the DNA.
[0297] Any suitable napDNAbp may be used in the prime editing factors described herein. In various embodiments, napDNAbp may be any class 2 CRISPR-Cas system and may encompass any type II, V, or VI CRISPR-Cas enzyme. Due to the rapid development of CRISPR-Cas as a tool for genome editing, there has been a certain development of nomenclature used to describe and / or identify CRISPR-Cas enzymes such as Cas9 and Cas9 orthologues. This application refers to CRISPR-Cas enzymes by nomenclature that may be old and / or new. Those skilled in the art will be able to identify the specific CRISPR-Cas enzymes referenced herein based on the nomenclature used, whether it is old (i.e., “legacy”) or new. The CRISPR-Cas nomenclature is described in detail in Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?”, The CRISPR Journal, Vol. 1. No. 5, 2018. The entire content of this document is incorporated into this application by reference. This application does not limit the specific CRISPR-Cas nomenclature used in any given case, and those skilled in the art should be able to identify which CRISPR-Cas enzyme is being referenced.
[0298] For example, the following type II, type V, and type VI class 2 CRISPR-Cas enzymes have the following older (i.e., legacy) and newer names recognized in the art. Each of these enzymes and / or their variants may be used in the prime editing factors described herein: [Table 3]
[0299] Without being constrained by theory, the mechanism of action of certain napDNAbps envisioned in this application involves the step of forming an R-loop. This causes the napDNAbp to induce unwinding of the double-stranded DNA target, thereby separating the strands in the region bound by the napDNAbp. Then, a guide RNA spacer hybridizes to the "target strand" at the protospacer sequence. This displaces the "non-target strand" complementary to the target strand, forming the single-stranded region of the R-loop. In some embodiments, the napDNAbp incorporates one or more nuclease activities, which then cut the DNA, leaving various types of scars. For example, the napDNAbp may include nuclease activity that cuts the non-target strand in a first position and / or the target strand in a second position. Depending on the nuclease activity, the target DNA may be cut to form a "double-strand break," thereby cutting both strands. In other embodiments, the target DNA may be cut at only a single site; that is, the DNA is "nicked" on one strand. Exemplary napDNAbp with different nuclease activities include “Cas9 nickase” (“nCas9”) and inactivated Cas9 (“inactive Cas9” or “dCas9”) which lack nuclease activity.
[0300] The following description of various napDNAbp that may be used with respect to the prime editing factors of this disclosure is not intended to be limiting. Prime editing factors may include standard SpCas9 or any orthologous Cas9 protein or any variant Cas9 protein that are known or can be produced or evolved by directed evolution or other mutation processes, encompassing any naturally occurring variants, mutants or otherwise engineered versions of Cas9. In various embodiments, Cas9 or Cas9 variants have nickase activity, i.e., cleave only one (of) strands of the target DNA sequence. In other embodiments, Cas9 or Cas9 variants have an inactive nuclease, i.e., they are “inactive” Cas9 proteins. Other variant Cas9 proteins that may be used have a smaller molecular weight than standard SpCas9 (e.g., for easier delivery) or have a modified or rearranged primary amino acid structure (e.g., a cyclic substitution format).
[0301] The prime editing factors described herein may also include Cas9 equivalents and encompass the Cas12a(Cpf1) and Cas12b1 proteins, which are the result of convergent evolution. The napDNAbp (e.g., SpCas9, Cas9 variants, or Cas9 equivalents) used herein may also include various modifications that alter / enhance their PAM specificity. Finally, this application intends any Cas9, Cas9 variant, or Cas9 equivalent having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity with respect to a reference Cas9 sequence, e.g., a reference SpCas9 standard sequence or a reference Cas9 equivalent (e.g., Cas12a(Cpf1)).
[0302] napDNAbp may be a CRISPR (Clustered Regular Interval Short Palindromic Repeat)-associated nuclease. As outlined above, CRISPR is an adaptive immune system that provides protection from mobile genetic elements (viruses, transposable elements, and conjugate-transmitted plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the ancestral mobile element, and a target entry nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of precrRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA acts as a guide for the ribonuclease 3-assisted processing of the precrRNA. Subsequently, Cas9 / crRNA / tracrRNA endo-cleaves a linear or circular dsDNA target complementary to the spacer. Target strands that are not complementary to the crRNA are first cut endoscopically and then trimmed 3'-5' exoscopically. In nature, DNA binding and cleavage typically require both proteins and both RNAs. However, a single guide RNA ("sgRNA" or simply "gRNA") can be manipulated to incorporate aspects of both crRNA and tracrRNA onto a single RNA species. See, for example, Jinek M. et al., Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference).
[0303] In some embodiments, napDNAbp leads to the cleavage of one or both strands at the location of the target sequence, for example, on the target sequence and / or on the complement of the target sequence. In some embodiments, napDNAbp leads to the cleavage of one or both strands within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a napDNAbp that is mutated relative to the corresponding wild-type enzyme such that the mutated napDNAbp lacks the ability to cleave one or both strands of the target polynucleotide containing the target sequence. For example, the substitution of aspartate to alanine on the RuvC I catalytic domain of Cas9 from S. pyogenes (D10A) converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of Cas9-nickase mutations include, without limitation, H840A, N854A, and N863A, referring to the standard SpCas9 sequence or equivalent amino acid positions of other Cas9 variants or Cas9 equivalents.
[0304] As used herein, the term “Cas protein” means any full-length Cas protein obtained from nature, a recombinant Cas protein having a sequence different from that of a naturally occurring Cas protein, or any fragment of a Cas protein that still retains all or a significant amount of essential basic functions required for the methods of the present disclosure, namely (i) possession of nucleic acid programmed binding of the Cas protein to target DNA and (ii) the ability to nick a target DNA sequence on a single strand. The Cas protein intended in this application encompasses the CRISPR Cas9 protein and Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9)) homologs, orthologs, or paralogs, whether naturally occurring or not (e.g., engineered or recombinant), and may encompass Cas9 equivalents from any class 2 CRISPR system (e.g., types II, V, VI), including Cas12a (Cpf1), Cas12e (CasX), Cas12b1 (C2c1), Cas12b2, Cas12c (C2c3), C2c4, C2c8, C2c5, C2c10, and C2c9. This includes Cas13a(C2c2), Cas13d, Cas13c(C2c7), Cas13b(C2c6), and Cas13b. Further Cas equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016;353(6299) and Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR Journal, Vol.1.No.5, 2018. These contents are incorporated herein by reference.
[0305] The terms “Cas9,” “Cas9 nuclease,” “Cas9 moiety,” or “Cas9 domain” encompass any naturally occurring Cas9 from any organism, any naturally occurring Cas9 equivalent or its functional fragment, any Cas9 homolog, ortholog, or paralog from any organism, and any variant or variant of any naturally occurring or engineered Cas9. The term Cas9 is not intended to be particularly limiting and may be said as “Cas9 or equivalent.” Exemplary Cas9 proteins are further described in this application and / or in the art and incorporated herein by reference. This disclosure is not limited to specific Cas9s used in the prime editing factors (PEs) of the present invention.
[0306] The Cas9 nuclease sequences and structures described in this application are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP,Qian Y.,Jia HG,Najar FZ,Ren Q.,Zhu H.,Song L.,White J.,Yuan X.,Clifton SW,Roe BA,McLaughlin RE,Proc.Natl.Acad.Sci.USA98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.”Deltcheva E., Chylinski See K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012) (the entire contents of each of these are incorporated herein by reference).
[0307] Examples of Cas9 and Cas9 equivalents are provided below; however, these specific examples are not intended to be limiting. The primer editing factors of this disclosure may use any suitable napDNAbp encompassing any suitable Cas9 or Cas9 equivalent.
[0308] A. Standard wild-type SpCas9 In one embodiment, the primer editing factor constructs described herein may include the “standard SpCas9” nuclease from S. pyogenes. This is widely used as a tool in genome engineering and is classified as a type II subgroup of enzymes in the class 2 CRISPR-Cas system. This Cas9 protein is a large multi-domain protein containing two distinct nuclease domains. Point mutations can be introduced on Cas9 to abolish the activity of one or both nucleases, resulting in nickased Cas9 (nCas9) or inactive Cas9 (dCas9), respectively, which still retains its ability to bind to DNA in a manner programmed by sgRNA. In principle, when fused to another protein or domain, Cas9 or its variant (e.g., nCas9) can target substantially any DNA sequence simply by co-expression with a suitable sgRNA. The standard SpCas9 protein used herein refers to the wild-type protein from Streptococcus pyogenes having the following amino acid sequence: [Table 4-1] [Table 4-2] [Table 4-3]
[0309] The prime editing factors described herein may encompass standard SpCas9 or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with the wild-type Cas9 sequence provided above. These variants may encompass SpCas9 variants containing one or more mutations, and may encompass any known mutations reported in the SwissProt accession No. Q99ZW2 (SEQ ID NO: 18) entry. These are: [Table 5] It includes.
[0310] Other wild-type SpCas9 sequences that may be used in this disclosure include: [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5] [Table 6-6] [Table 6-7] [Table 6-8] It includes.
[0311] The prime editing factors described herein may include any of the above SpCas9 sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0312] B. Wild-type Cas9 ortholog In other embodiments, the Cas9 protein may be a wild-type Cas9 ortholog from a different bacterial species than the standard Cas9 from S. pyogenes. For example, the following Cas9 orthologs may be used in relation to the prime editing factor constructs described herein. In addition, any variant Cas9 ortholog having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with any of the orthologs below may also be used in the prime editing factor. [Table 7-1] [Table 7-2] [Table 7-3] [Table 7-4] [Table 7-5] [Table 7-6] [Table 7-7]
[0313] The prime editing factors described herein may include any of the above Cas9 orthologous sequences, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0314] napDNAbp may encompass a suitable homolog and / or ortholog of a naturally occurring enzyme such as Cas9. Cas9 homologs and / or orthologs have been described in a variety of species, including but not limited to S. pyogenes and S. thermophilus. Preferably, the Cas portion is configured as a nickase (e.g., mutagenerated, recombinant, or otherwise obtained from nature), that is, capable of cleaving only the single strand of the target. Additional (doubtdditional) suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure. Such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems” (2013) RNA Biology 10:5, 726-737; the entire contents of this are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., deactivated) DNA cleavage domain; that is, Cas9 is a nickase. In some embodiments, the Cas9 protein contains an amino acid sequence that is at least 80% identical to the amino acid sequence of the Cas9 protein provided by any one of the variants in Table 3. In some embodiments, the Cas9 protein contains an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of the Cas9 protein provided by any one of the Cas9 orthologs in the table above.
[0315] C. Inactive Cas9 variant In one embodiment, the prime editing factors described herein may include inactive Cas9, for example, inactive SpCas9, which lacks nuclease activity due to one or ...
Claims
1. (i) A fusion protein comprising a nucleic acid programmed DNA-binding protein (napDNAbp) and a domain containing RNA-dependent DNA polymerase activity; and (ii) Prime editing guide RNA (PEgRNA) A complex for prime editing, including [specific feature / feature].
2. The complex according to claim 1, wherein the fusion protein can perform prime editing in the presence of prime editing guide RNA (PEgRNA) to incorporate a desired nucleotide change into a target sequence.
3. The complex according to claim 11, wherein napDNAbp has nickase activity.
4. The complex according to claim 1, wherein napDNAbp is a Cas9 protein or a variant thereof.
5. The complex according to claim 1, wherein napDNAbp is nuclease-active Cas9, nuclease-inactive Cas9 (dCas9), or Cas9 niccas (nCas9).
6. The complex according to claim 1, wherein napDNAbp is Cas9 nickase (nCas9).
7. The complex according to claim 1, wherein napDNAbp is selected from the group consisting of Cas9, Cas12e, Cas12d, Cas12a, Cas12b1, Cas13a, Cas12c, and Argonaut, and optionally has nickase activity.
8. The complex according to claim 1, wherein the domain containing RNA-dependent DNA polymerase activity is a reverse transcriptase and comprises one of the amino acid sequences of SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700, 701-716, 739-741, and 766.
9. The complex according to claim 1, wherein the domain containing RNA-dependent DNA polymerase activity is a reverse transcriptase and comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with any one of the amino acid sequences of SEQ ID NOs: 89-100, 105-122, 128-129, 132, 139, 143, 149, 154, 159, 235, 454, 471, 516, 662, 700, 701-716, 739-741, and 766.
10. The complex according to claim 1, wherein the domain containing RNA-dependent DNA polymerase activity is a naturally occurring reverse transcriptase from a retrovirus or retrotransposon.
11. The complex according to claim 1, wherein the fusion protein can bind to a target DNA sequence when complexed with PEgRNA.
12. The complex according to claim 1, wherein the PEgRNA comprises a guide RNA and at least one nucleic acid elongation arm comprising a DNA synthesis template.
13. The complex according to claim 12, wherein the nucleic acid elongation arm is at the 3' or 5' end of the guide RNA or at an intramolecular position of the guide RNA, and the nucleic acid elongation arm is DNA or RNA.
14. The complex according to claim 12, wherein PEgRNA can bind to napDNAbp and guide napDNAbp to a target DNA sequence.
15. The complex according to claim 14, wherein the target DNA sequence comprises a target strand and a complementary non-target strand.
16. The complex according to claim 12, wherein the guide RNA hybridizes to the target strand to form an RNA-DNA hybrid and an R-loop.
17. The complex according to claim 12, wherein at least one nucleic acid elongation arm further comprises a primer binding site.
18. The complex according to Claim 12, wherein the nucleic acid elongation arm is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 31 nucleotides, at least 32 nucleotides, at least 33 nucleotides, at least 34 nucleotides, at least 35 nucleotides, at least 36 nucleotides, at least 37 nucleotides, at least 38 nucleotides, at least 39 nucleotides, at least 40 nucleotides, at least 41 nucleotides, at least 42 nucleotides, at least 43 nucleotides, at least 44 nucleotides, at least 45 nucleotides, at least 46 nucleotides, at least 47 nucleotides, at least 48 nucleotides, at least 49 nucleotides, or at least 50 nucleotides.
19. The complex according to claim 12, wherein the DNA synthesis template has a length of at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, or at least 15 nucleotides.
20. The complex according to claim 17, wherein the primer binding site has a length of at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, or at least 15 nucleotides.
21. The complex according to claim 12, wherein the PEgRNA further comprises at least one additional structure selected from the group consisting of a linker, stem-loop, hairpin, toe-loop, aptamer, or RNA-protein recruitment domain.
22. The complex according to claim 12, wherein the DNA synthesis template encodes a single-stranded DNA flap that is complementary to an endogenous DNA sequence adjacent to the nick site, and the single-stranded DNA flap contains a desired nucleotide change.
23. The complex according to claim 22, wherein the single-stranded DNA flap displaces endogenous single-stranded DNA having its 5' end on the nicked target DNA sequence, and the endogenous single-stranded DNA is immediately adjacent downstream of the nicked site.
24. The complex according to claim 23, wherein endogenous single-stranded DNA having a free 5' end is excised by the cell.
25. The complex according to claim 23, wherein cellular repair of a single-stranded DNA flap results in the incorporation of a desired nucleotide change, thereby forming a desired product.
26. The complex according to claim 12, wherein the PEgRNA comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with any one of the nucleotide sequences of SEQ ID NOs. 18-36, or SEQ ID NOs. 101-104, 181-183, 223-244, 277, 325-334, 336, 338, 340, 342, 344, 346, 348, 350, 352, 354, 356, 358, 360, 362, 364, 366, 368, 499-505, 735-761, or 776-777.
27. The complex according to claim 12, wherein the DNA synthesis template comprises a nucleotide sequence that is at least 80%, 85%, 90%, 95%, or 99% identical to the endogenous DNA target.
28. The complex according to claim 17, wherein the primer binding site hybridizes with the free 3' end of the cut DNA.
29. The complex according to claim 21, wherein at least one additional structure is located at the 3' or 5' end of the PEgRNA.
30. The complex according to claim 29, wherein the linker comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 127, 165-176, 446, 453, and 767-769.
31. The complex according to claim 29, wherein the stem-loop comprises a nucleotide sequence selected from the stem-loops described herein.
32. The complex according to claim 29, wherein the hairpin comprises a nucleotide sequence selected from the hairpins described herein.
33. The complex according to claim 29, wherein the troop comprises a nucleotide sequence selected from the troops described herein.
34. The complex according to claim 29, wherein the aptamer comprises a nucleotide sequence selected from the aptamers described herein.
35. The complex according to claim 29, wherein the RNA-protein recruitment domain comprises a nucleotide sequence selected from the RNA-protein recruitment domains described herein.
36. The complex according to claim 1, wherein the target DNA sequence comprises a target strand and a complementary non-target strand.
37. The complex according to claim 36, wherein the R loop comprises (i) an RNA-DNA hybrid including PEgRNA and a target strand and (ii) a complementary non-target strand.
38. The complex according to claim 37, wherein a target or complementary non-target chain is nicked to form a prime sequence having a free 3' end.
39. The complex according to claim 38, wherein the nick site is upstream of the PAM sequence on the target chain.
40. The complex according to claim 38, wherein the nick site is upstream of the PAM sequence on the non-target strand.
41. The complex according to claim 38, wherein the nick site is -1, -2, -3, -4, -5, -6, -7, -8, or -9 relative to the 5' end of the PAM sequence.
42. The complex according to claim 22, wherein a single-stranded DNA flap hybridizes to an endogenous DNA sequence adjacent to a nick site, thereby incorporating a desired nucleotide change onto the target strand.
43. The complex according to claim 22, wherein the single-stranded DNA flap replaces an endogenous DNA sequence having a free 5' end and adjacent to a nick site.
44. The complex according to claim 22, wherein an endogenous DNA sequence having a 5' end is excised by the cell.
45. The complex according to claim 44, wherein an endogenous DNA sequence having a 5' end is excised by a flap endonuclease.
46. The complex according to claim 43, wherein cellular repair of a single-stranded DNA flap incorporates a desired nucleotide change into the non-target strand, thereby forming a desired product.
47. The complex according to claim 46, wherein a desired nucleotide change is incorporated into an editing window where the PAM sequence is between approximately -4 and +10, or between approximately -10 and +20, or between approximately -20 and +40, or between approximately -30 and +100.
48. The desired nucleotide change occurs at least at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, The complex according to claim 47, which is incorporated downstream of 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
49. The fusion protein has a structure NH 2 -[napDNAbp]-[domain containing RNA-dependent DNA polymerase activity]-COOH; or NH 2 The complex according to any one of claims 1 to 48, comprising -[domain containing RNA-dependent DNA polymerase activity]-[napDNAbp]-COOH, wherein each of "]-[" indicates the presence of an arbitrary linker sequence.
50. The complex according to claim 49, wherein the linker sequence comprises the amino acid sequences of SEQ ID NOs. 127, 165-176, 446, 453, and 767-769.
51. The complex according to claim 1, wherein the fusion protein further comprises a linker that links a domain containing napDNAbp and RNA-dependent DNA polymerase activity.
52. The complex according to claim 51, wherein the linker sequence comprises the amino acid sequences of SEQ ID NOs. 3887(1×SGGS), 3888(2×SGGS), 3889(3×SGGS), 3890(1×XTEN), 3891(1×EAAAK), 3892(2×EAAAK), and 3893(3×EAAAK).