Compositions and methods for targeted genomic editing of the GNE gene
By employing a nucleic acid programmable DNA binding protein and prime editing guide RNA to correct the M743T allele in the GNE gene, the method enhances enzyme function and stability, providing a potential therapeutic solution for GNE myopathy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- JOHNS HOPKINS UNIVERSITY
- Filing Date
- 2025-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
There is a lack of effective therapies for GNE myopathy, an adult-onset, autosomal recessive genetic disease caused by loss of function pathogenic variants or mutations in the GNE gene, which leads to progressive muscle weakness and loss of ambulation.
The use of a nucleic acid programmable DNA binding protein (napDNAbp) fused to a polymerase, such as a reverse transcriptase, associated with a prime editing guide RNA (pegRNA), to edit the GNE gene locus and correct the M743T allele, thereby modifying the GNE-encoding polynucleotide to produce a modified GNE gene product with increased enzyme function, stability, and/or production.
The method effectively corrects the M743T mutation in the GNE gene, resulting in a modified GNE gene product with improved enzyme activity, stability, and production, addressing the need for therapeutic interventions for GNE myopathy.
Smart Images

Figure US2025054285_15052026_PF_FP_ABST
Abstract
Description
Attorney Docket No. JHV-16925COMPOSITIONS AND METHODS FOR TARGETED GENOMIC EDITING OF THE GNE GENECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application Ser. No. 63 / 717,088 filed November 6, 2024, which is incorporated by reference herein in its entirety.GOVERNMENT SUPPORT
[0002] This invention was made with government support under grants GM007471 and GM148383 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROUND
[0003] GNE myopathy is an adult-onset, autosomal recessive genetic disease characterized by progressive muscle weakness that that can lead to loss of ambulation and loss of independent living. GNE myopathy is caused by loss of function pathogenic variants or mutations in the GNE gene, which encodes a bifunctional UDP-GIcNAc-epimerase / ManNAc-6 kinase, an essential enzyme in the sialic acid biosynthesis pathway. Despite the fact that mutations in the GNE gene were shown to cause GNE myopathy over twenty years ago, there remains a lack of effective therapies for this disease.SUMMARY
[0004] Provided herein are methods, complexes, vectors, compositions, and kits for editing a mutation in a GNE gene. Specifically, provided are technologies for modifying a GNE-encoding polynucleotide (e.g., genomic DNA) to correct sequence encoding a GNE M743T allele. In some embodiments, provided are technologies for modifying a GNE-encoding polynucleotide (e.g., genomic DNA) using a complex comprising a nucleic acid programmable DNA binding protein (napDNAbp) fused to a polymerase (e.g., a reverse transcriptase (RT)) associated with a prime editing guide RNA (pegRNA) to edit a GNE gene.
[0005] In some embodiments, the present disclosure provides methods of prime editing the GNE gene locus using a prime editor that comprises a reverse transcriptase domain and a nickase domain, which, when complexed with a suitable guide RNA (pegRNA), are effective in installing1FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 one or more edits (e.g., insertion, deletion, or replacement of one or more nucleobases) in the GNE gene locus. In some embodiments, the one or more nucleobase edits installed in the GNE gene locus result in a modified GNE gene product having increased enzyme function, stability, and / or production relative to the non-mutated GNE protein. In some embodiments, the genome editing strategies disclosed herein can be implemented to target the editing of specific nucleobase positions in the GNE gene locus, which, when edited, impart the modified GNE gene product with improved enzyme activity, stability, and / or production.
[0006] In some embodiments, the present disclosure provides pegRNAs for prime editing comprising (i) a spacer sequence, (ii) a scaffold, and (iii) an extension arm comprising a reversetranscription template (RTT) and a primer-binding site (PBS), wherein the scaffold binds to a nucleic acid programmable DNA binding protein (napDNAbp), wherein the RTT comprises a mutation to revert a sequence encoding a GNE M743T allele to a wild-type GNE sequence, and wherein: (a) the PBS has a length between 9-14 nt; and (b) the RTT has a length between 34 to 70 nt.
[0007] In some embodiments, the RTT of the pegRNA has a length between 50 to 70 nt. In some embodiments, the RTT of the pegRNA has a length between 58-60 nt.
[0008] In some embodiments, the RTT further comprises at least one silent mutation. In some embodiments, the RTT comprises 1 to 18 silent mutations. In some embodiments, the RTT comprises 1, 2, 3, or 4 silent mutations.
[0009] In some embodiments, the RTT comprises a sequence that differs by no more than 4 nucleotides, by no more than 3 nucleotides, by no more than 2 nucleotides, or by no more than 1 nucleotides from a sequence as set forth in any one of SEQ ID NOs: 327-652. In some embodiments, the RTT comprises a sequence as set forth in any one of SEQ ID NOs: 327-652.
[0010] In some embodiments, the RTT comprises a sequence as set forth in any one of SEQ ID NOs: 327-652. In some embodiments, the RTT consists or consists essentially of a sequence as set forth in any one of SEQ ID NOs: 327-652.
[0011] In some embodiments, the RTT comprises, consists, or consists essentially of a sequence of TGGGTGCTGCCAGCATGGTCCTGGACTACACAACACGCAGGATCTACTAGACCTGCAG GA (SEQ ID NO: 1644) or2FoleyHoagUS13167537.2Attorney Docket No. JHV-16925GGTGCTGCCAGCATGGTCCTGGACTACACAACACGCAGGATCTACTAGACCTGCAGGA(SEQ ID NO: 327).
[0012] In some embodiments, the primer binding sites (PBS) of the pegRNA comprises a 7- 15 nt sequence that is reverse complementary to the spacer sequence. In some embodiments, the PBS of the pegRNA comprises a 12 nt sequence that is reverse complementary to the spacer sequence.
[0013] In some embodiments, the PBS comprises a sequence that differs by no more than 3 nucleotides, by no more than 2 nucleotides, or by no more than 1 nucleotides from a sequence as set forth in any one of SEQ ID NOs: 979-1304. In some embodiments, the PBS comprises a sequence as set forth in any one of SEQ ID NOs: 979-1304. In some embodiments, the PBS consists or consists essentially of a sequence as set forth in any one of SEQ ID NOs: 979-1304.
[0014] In some embodiments, the PBS comprises a sequence of ACAGACATGGAC (SEQ ID NO: 979).
[0015] In some embodiments, the pegRNA comprises a spacer that is 15-25 nt in length. In some embodiments, the spacer is 20 nt in length.
[0016] In some embodiments, the spacer comprises a sequence that differs by no more than 4 nucleotides, by no more than 3 nucleotides, by no more than 2 nucleotides, or by no more than 1 nucleotides from a sequence as set forth in any one of SEQ ID NOs: 653-978. In some embodiments, the spacer comprises a sequence of any one of SEQ ID NOs: 653-978. In some embodiments, the spacer consists or consists essentially of a sequence of any one of SEQ ID NOs: 653-978.
[0017] In some embodiments, the spacer comprises, consists, or consists essentially of a sequence of AGAAGGTCCATGTCTGTTCC (SEQ ID NO: 653).
[0018] In some embodiments, the spacer binds to a sequence in the genome that is 30-110 nt away from the corresponding mutation site to be corrected. In some embodiments, the spacer binds to a sequence in the genome that is 42 nt away from the corresponding mutation site to be corrected.
[0019] In some embodiments, the scaffold of the pegRNA binds to a napDNAbp that is a Cas9 protein or variant thereof.
[0020] In some embodiments, the scaffold comprises a nucleic acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to the sequence: GTTTAAGAGCTAAGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCA3FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 ACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 1631). In some embodiments, the scaffold comprises a sequence that differs by no more than 4 nucleotides, by no more than 3 nucleotides, by no more than 2 nucleotides, or by no more than 1 nucleotides from a sequence as set forth in SEQ ID NO: 1631. In some embodiments, the scaffold comprises, consists, or consists essentially of a nucleic acid sequence as set forth in SEQ ID NO: 1631.
[0021] In some embodiments, the extension arm of the pegRNA is at the 3' end of the pegRNA. In some embodiments, the pegRNA comprises from 5' to 3' the spacer, the scaffold and the extension arm.
[0022] In some embodiments, the extension arm of the pegRNA is at the 5' end of the pegRNA. In some embodiments, the pegRNA comprises from 5' to 3' the extension arm, the scaffold and the spacer.
[0023] In some embodiments, the pegRNA comprises, consists, or consists essentially of a nucleic acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 1-326. In some embodiments, the PBS comprises a sequence that differs by no more than 8 nucleotides, by no more than 7 nucleotides, by no more than 6 nucleotides, by no more than 5 nucleotides, by no more than 4 nucleotides, by no more than 3 nucleotides, by no more than 2 nucleotides, or by no more than 1 nucleotides from a sequence as set forth in any one of SEQ ID NOs: 1-326. In some embodiments, the pegRNA comprises, consists, or consists essentially of a nucleic acid sequence of any one of SEQ ID NOs: 1-326.
[0024] In some embodiments, the present disclosure provides a gene editing complex for editing a GNE gene, the complex comprising a prime editor comprising (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain and (ii) a domain having an RNA- dependent DNA polymerase activity, and a pegRNA as described herein.
[0025] In some embodiments, the napDNAbp has a nickase activity. In some embodiments, the napDNAbp is a Cas9 protein or variant thereof. In some embodiments, the napDNAbp is a nuclease active Cas9, a nuclease inactive Cas9 (dCas9), or a Cas9 nickase (nCas9). In some embodiments, the napDNAbp is Cas9 nickase (nCas9).
[0026] In some embodiments, the domain comprising an RNA-dependent DNA polymerase activity is a reverse transcriptase.
[0027] In some embodiments, the gene editing complex further comprises a nicking guide RNA, a dead guide RNA or a combination thereof.4FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0028] In some embodiments, the present disclosure provides methods of correcting a M743T mutation in a GNE gene by prime editing, comprising contacting a target DNA sequence with a prime editor comprising (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain, and (ii) a domain having an RNA-dependent DNA polymerase activity, and a pegRNA.
[0029] In some embodiments, provided are methods of correcting a M743T mutation in a GNE gene by prime editing comprising contacting a target DNA sequence with the gene editing complex as described herein.
[0030] In some embodiments, provided are methods of correcting a M743T mutation in a GNE gene by cytosine base editing comprising contacting a target DNA sequence with a cytosine base editor comprising (i) a guide RNA, (ii) programmable DNA binding protein (napDNAbp), and (iii) a cytidine deaminase.
[0031] In some embodiments, provided methods of correcting a M743T mutation in a GNE gene by cytosine base editing use a guide RNA that comprises a spacer and a scaffold. In some embodiments, the guide RNA comprises a spacer that differs by no more than 4 nucleotides, by no more than 3 nucleotides, by no more than 2 nucleotides, or by no more than 1 nucleotides from a sequence as set forth in SEQ ID NO: 1668. In some embodiments, the guide RNA comprises a spacer sequence that is at least 90% identical or at least 95% identical to SEQ ID NO: 1668. In some embodiments, the guide RNA comprises a spacer sequence of SEQ ID NO: 1668.
[0032] In some embodiments, the guide RNA comprises a scaffold that binds to a napDNAbp that is a Cas9 protein or variant thereof.
[0033] In some embodiments, the guide RNA comprises a spacer sequence that is at least 90% identical or at least 95% identical to SEQ ID NO: 1668. In some embodiments, the guide RNA for cytosine base editing comprising a spacer sequence of SEQ ID NO: 1668.
[0034] In some embodiments, the guide RNA comprises a scaffold that binds to a napDNAbp that is a Cas9 protein or variant thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The Drawing included herein, which is composed of the following Figures, is for illustration purposes only and not for limitation.5FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0036] FIG. 1A provides a representative karyotype deriving from a HEK293T cell harboring the M743T pathogenic variant. The karyotype displays all chromosomes arranged in pairs from chromosome 1 to 22, along with the sex chromosomes (X and Y). The GNE gene is present on chromosome 9.
[0037] FIG. IB depicts a specific segment of a DNA sequence from chromosome 9, spanning positions 36,217,382 to 36,217,420. It presents the nucleotide sequence (A, T, C, G) alongside a graphical representation of sequencing reads, highlighting the presence of genetic variations or mutations within this region. The substitution of a “T” to “C” leads to the emergence of the M743T pathogenic variant.
[0038] FIG. 2 provides a graph illustrating the editing efficiency of different pegRNAs (prime editing guide RNAs) for correcting a specific DNA mutation in the GNE gene responsible for the M743T pathogenic variant. The top 5 pegRNAs incorporate a 58-nt RTT (reverse transcription template) and 12-nt PBS (primer binding site) with additional silent mutations in the RTT. The bottom 5 pegRNAs vary in RTT and PBS length and lack silent mutations in the RTT. A pegRNA targeting the HEK3 locus was used as a positive control while an empty vector (pucl9) was used as a negative control.
[0039] FIG. 3 provides a graph depicting the efficiency of M743T correction using pegRNAs with varying numbers of silent mutations in the reverse transcription template (RTT). The results are grouped based on the number of silent mutations introduced in the RTT (1, 2, 3, or 4 mutations). The analysis aims to identify the impact of silent mutations in the RTT on the editing efficiency for correcting the M743T mutation.
[0040] FIG. 4A provides sequence alignments for Spacer 1 and Spacer 2 used in M743T correction experiments. For each spacer, the sequences of pegRNAs SE11 and SE18 (for Spacer 1) and SE15 and SE16 (for Spacer 2) are aligned with their target regions. The sequences include silent mutations (highlighted in red) in the reverse transcription template (RTT) to differentiate between the edited and unedited sequences.
[0041] FIG. 4B provides a graph comparing the editing efficiencies of the various transfected pegRNAs described in FIG. 4A. Empty plasmid vectors (“No SE”) were used as a control.6FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0042] FIG. 5A provides a graph comparing the editing efficiencies of various prime editing (PE) variants using Spacer 1. Blue bars indicate the percentage of reads with the specified correction, and the gray bars show the proportion of reads with indels.
[0043] FIG. 5B provides a graph comparing the editing efficiencies of various prime editing (PE) variants using Spacer 2. Blue bars indicate the percentage of reads with the specified correction, and the gray bars show the proportion of reads with indels.
[0044] FIG. 6A provides a graph comparing the editing efficiencies of various nicking sgRNAs (nsgRNAs) used with Spacer 1 for correcting the M743T variant. Blue bars represent the percentage of sequencing reads with the desired nucleotide correction, and gray bars indicate the proportion of reads with indels. Each nsgRNA is labeled with their specific nucleotide position relative to the mutation site. A “no nsgRNA” condition is used as a control.
[0045] FIG. 6B provides a graph comparing the editing efficiencies of various dead sgRNAs (dsgRNAs) used with Spacer 1 for correcting the M743T variant. Blue bars represent the percentage of sequencing reads with the desired nucleotide correction, and gray bars indicate the proportion of reads with indels. Each dsgRNA is labeled with their specific nucleotide position relative to the mutation site. A “no dsgRNA” condition is used as a control.
[0046] FIG. 7 provides a graph comparing the editing efficiencies of the original pegRNA and various pegRNA variants (SEI, SE2, SE3) in the presence (+) or absence (-) of MLHldn (dominant-negative MLH1. Blue bars represent the percentage of sequencing reads with the desired nucleotide correction, and gray bars indicate the proportion of reads with indels.
[0047] FIG. 8 provides a graph comparing the editing efficiencies of various nicking sgRNAs (nsgRNAs) in the presence (+) or absence (-) of MLHldn (dominant-negative MLH1). Blue bars represent the percentage of sequencing reads with the desired nucleotide correction, and gray bars indicate the proportion of reads with indels.
[0048] FIG. 9 provides a graph comparing the editing efficiencies using various combinations of conditions using different nsgRNAs, in the presence or absence of MLHldn, and dsgRNAs. Blue bars represent the percentage of sequencing reads with the desired nucleotide correction, and gray bars indicate the proportion of reads with indels. A pegRNA targeting the HEK3 locus was used as a positive control.
[0049] FIG. 10 provides a graph comparing the best editing efficiencies of the best combinations of nsgRNAs and dsgRNAs using a PE6c editor. Blue bars represent the percentage of7FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 sequencing reads with the desired nucleotide correction, and gray bars indicate the proportion of reads with indels. A pegRNA targeting the HEK3 locus was used as a positive control. Plasmids expressing either pegRNAs (spacer) or PE6c alone were used as a negative control.
[0050] FIG. 11A provides a sequence alignment within the GNE coding sequence (CDS), highlighting the position of the CBE6b-SpRY spacer and two prime editing spacers.
[0051] FIG. 11B provides a graph displaying the efficiency of CBE6b-SpRY-mediated editing, focusing on specific target (C6) and silent bystander mutations (C4 and Cll). A plasmid expressing an empty vector (pucl9) was used as a negative control. Green bars represent the percentage of sequencing reads with the desired nucleotide correction, blue bars represent the percentage of sequencing reads with an edited silent bystander nucleotide, and gray bars indicate the proportion of reads with indels.
[0052] FIG. 12 provides a graph showing prime editing efficiency in mouse Neuro-2a (N2a) cells using pegRNAs adapted from the human GNE M743T editing system.
[0053] FIG. 13A provides a graph showing the results of targeted sequencing analysis of a predicted “Off -target 1” locus in human cells following delivery of the GNE M743T prime editing system.
[0054] FIG. 13B provides a graph showing the results of targeted sequencing analysis of a predicted “Off -target 2” locus in human cells following delivery of the GNE M743T prime editing system.
[0055] FIG. 14 provides sequence alignments and a list showing candidate off-target loci identified by CHANGE-seq analysis using prime editing and base editing guide RNAs.DEFINITIONS
[0056] In order for the present disclosures(s) to be more readily understood, certain terms are first defined below. Additional definitions for the following terms and other terms are set forth throughout the specification. The publications and other reference materials referenced herein to are hereby incorporated by reference.
[0057] Standard art-accepted meanings of terms are used herein unless indicated otherwise. Standard abbreviations for various terms are used herein.
[0058] In this application, unless otherwise clear from context, (i) the terms “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the8FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 article; (ii) the term “or” may be understood to mean “and / or”; (iii) the terms “comprising” and “including” may be understood to encompass itemized components or steps whether presented by themselves or together with one or more additional components or steps; and (iv) where ranges are provided, endpoints are included.
[0059] As used herein, the term “adenosine deaminase” or “adenosine deaminase domain” refers to a protein or enzyme that catalyzes a deamination reaction of an adenosine (or adenine). The terms are used interchangeably. In certain embodiments, the disclosure provides base editors comprising one or more adenosine deaminase domains. For instance, an adenosine deaminase domain may comprise a heterodimer of a first adenosine deaminase and a second deaminase domain, connected by a linker. Adenosine deaminases (e.g., engineered adenosine deaminases or evolved adenosine deaminases) provided herein may be may be enzymes that convert adenine (A) to inosine (I) in DNA or RNA. Such adenosine deaminase can lead to an A:T to G:C base pair conversion. In some embodiments, the deaminase is a variant of a naturally-occurring deaminase from an organism. In some embodiments, the deaminase does not occur in nature. For example, in some embodiments, the deaminase is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring deaminase.
[0060] In some embodiments, the adenosine deaminase is derived from a bacterium, such as, E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is an E. coli TadA deaminase (ecTadA). In some embodiments, the TadA deaminase is a truncated E. coli TadA deaminase. For example, the truncated ecTadA may be missing one or more N-terminal amino acids relative to a full-length ecTadA. In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5 ,6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N- terminal amino acid residues relative to the full length ecTadA. In some embodiments, the truncated ecTadA may be missing 1, 2, 3, 4, 5 ,6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C- terminal amino acid residues relative to the full length ecTadA. In some embodiments, the ecTadA deaminase does not comprise an N-terminal methionine. Reference is made to U.S. Patent Publication No. 2018 / 0073012, published Mar. 15, 2018, which is incorporated herein by reference.
[0061] In genetics, the “antisense” strand of a segment within double-stranded DNA is the template strand, and which is considered to run in the 3' to 5' orientation. By contrast, the “sense”9FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. In the case of a DNA segment that encodes a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription, and eventually undergoes (typically, not always) translation into a protein. The antisense strand is thus responsible for the RNA that is later translated to protein, while the sense strand possesses a nearly identical makeup to that of the mRNA. Note that for each segment of dsDNA, there will possibly be two sets of sense and antisense, depending on which direction one reads (since sense and antisense is relative to perspective). It is ultimately the gene product, or mRNA, that dictates which strand of one segment of dsDNA is referred to as sense or antisense.
[0062] “Base editing” refers to genome editing technology that involves the conversion of a specific nucleic acid base into another at a targeted genomic locus. In certain embodiments, this can be achieved without requiring double-stranded DNA breaks (DSB), or single stranded breaks (i.e., nicking). To date, other genome editing techniques, including CRISPR -based systems, begin with the introduction of a DSB at a locus of interest. Subsequently, cellular DNA repair enzymes mend the break, commonly resulting in random insertions or deletions (indels) of bases at the site of the DSB. However, when the introduction or correction of a point mutation at a target locus is desired rather than stochastic disruption of the entire gene, these genome editing techniques are unsuitable, as correction rates are low (e.g. typically 0.1% to 5%), with the major genome editing products being indels. In order to increase the efficiency of gene correction without simultaneously introducing random indels, the CRISPR / Cas9 system can be modified to directly convert one DNA base into another without inducing double-stranded DNA breaks (DSBs). See, Komor, A. C., et al., Programmable editing of a target base in genomic DNA without double- stranded DNA cleavage. Nature 533, 420-424 (2016), the entire contents of which is incorporated by reference herein.
[0063] The term “base editor (BE)” as used herein, refers to an agent comprising a polypeptide that is capable of making a modification to a base (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA) that converts one base to another (e.g., A to G, AtoC,AtoT,CtoT,CtoG,CtoA,GtoA,GtoC,GtoT,TtoA,TtoC,TtoG). In some embodiments, the base editor is capable of deaminating a base within a nucleic acid such as a base within a DNA molecule. In the case of an adenine base editor, the base editor is capable of deaminating an adenine10FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 (A) in DNA. Such base editors may include a nucleic acid programmable DNA binding protein (napDNAbp) fused to an adenosine deaminase. Some base editors include CRISPR-mediated fusion proteins that are utilized in the base editing methods described herein. In some embodiments, the base editor comprises a nuclease-inactive Cas9 (dCas9) fused to a deaminase which binds a nucleic acid in a guide RNA-programmed manner via the formation of an R-loop, but does not cleave the nucleic acid. For example, the dCas9 domain of the fusion protein may include a D10A and a H840A mutation (which renders Cas9 capable of cleaving only one strand of a nucleic acid duplex), as described in PCT / US2016 / 058344, which published as WO 2017 / 070632, and is incorporated herein by reference in its entirety. The DNA cleavage domain of S. pyogenes Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvCl subdomain. The HNH subdomain cleaves the strand complementary to the gRNA (the “targeted strand”, or the strand in which editing or deamination occurs), whereas the RuvCl subdomain cleaves the non- complementary strand containing the PAM sequence (the “non-edited strand”). The RuvCl mutant D10A generates a nick in the targeted strand, while the HNH mutant H840A generates a nick on the non-edited strand (see Jinek et al., Science, 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83 (2013), each of which are incorporated by reference herein).
[0064] The term “Cas9” or “Cas9 nuclease” refers to an RNA-guided nuclease comprising a Cas9 domain, or a fragment thereof (e.g., a protein comprising an active or inactive DNA cleavage domain of Cas9, and / or the gRNA binding domain of Cas9). A “Cas9 domain” as used herein, is a protein fragment comprising an active or inactive cleavage domain of Cas9 and / or the gRNA binding domain of Cas9. A “Cas9 protein” is a full length Cas9 protein. A Cas9 nuclease is also referred to sometimes as a casnl nuclease or a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and a Cas9 domain. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA target complementary to the spacer. The target strand not complementary to crRNA is first cut endonucleolytically, then11FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 trimmed 3'-5' exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., et al., Science 337:816-821(2012), the entire contents of which are hereby incorporated by reference. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self versus non-self. Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an Ml strain of Streptococcus pyogenes" Ferretti et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III” Deltcheva E., et al., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity” Jinek M., et al., Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus (e.g., StCas9 or StlCas9). Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, a Cas9 nuclease comprises one or more mutations that partially impair or inactivate the DNA cleavage domain.
[0065] A nuclease-inactivated Cas9 domain may interchangeably be referred to as a “dCas9” protein (for nuclease-”dead” Cas9). Methods for generating a Cas9 domain (or a fragment thereof) having an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence- Specific Control of Gene Expression” (2013) Cell. 28;152(5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvCl subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, whereas the RuvCl subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5): 1173-83 (2013)).12FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0066] Cas9 sequences or equivalents thereof are known in the art, for example as described in US 11,795,452, the contents of which is hereby incorporated by reference in their entirety.
[0067] In some embodiments, a protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or fragments thereof are referred to as “Cas9 variants” A Cas9 variant shares homology to Cas9, or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, at least about 99.8% identical, or at least about 99.9% identical to wild type Cas9 (e.g., SpCas9). In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild type Cas9 (e.g., SpCas9). In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA-cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild type Cas9 (e.g., SpCas9). In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild type Cas9 (e.g., SpCas9).
[0068] As used herein, a “cytidine deaminase” encoded by the CDA gene is an enzyme that catalyzes the removal of an amine group from cytidine (i.e., the base cytosine when attached to a ribose ring) to uridine (C to U) and deoxycytidine to deoxyuridine (C to U). A non-limiting example of a cytidine deaminase is APOB EC 1 (“apolipoprotein B mRNA editing enzyme, catalytic polypeptide 1”). Another example is AID (“activation-induced cytidine deaminase”). Under standard Watson-Crick hydrogen bond pairing, a cytosine base hydrogen bonds to a guanine base. When cytidine is converted to uridine (or deoxy cytidine is converted to deoxyuridine), the uridine (or the uracil base of uridine) undergoes hydrogen bond pairing with the base adenine. Thus, a conversion of “C” to uridine (“U”) by cytidine deaminase will cause the insertion of “A” instead of13FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 a “G” during cellular repair and / or replication processes. Since the adenine “A” pairs with thymine “T”, the cytidine deaminase in coordination with DNA replication causes the conversion of an C G pairing to a T A pairing in the double-stranded DNA molecule.
[0069] The term “DNA editing efficiency” as used herein, refers to the number or proportion of intended base pairs that are edited. For example, if a prime editor edits 10% of the base pairs that it is intended to target (e.g., within a cell or within a population of cells), then the prime editor can be described as being 10% efficient. Some aspects of editing efficiency embrace the modification (e.g. deamination) of a specific nucleotide within DNA, without generating a large number or percentage of insertions or deletions (i.e., indels). It is generally accepted that editing while generating less than 5% indels (as measured over total target nucleotide substrates) is high editing efficiency. The generation of more than 20% indels is generally accepted as poor or low editing efficiency. Indel formation may be measured by techniques known in the art, including high-throughput screening of sequencing reads.
[0070] The term “off-target editing frequency” as used herein, refers to the number or proportion of unintended base pairs, e.g. DNA base pairs, that are edited. On-target and off-target editing frequencies may be measured by the methods and assays described herein, further in view of techniques known in the art, including high-throughput sequencing reads. As used herein, high- throughput sequencing involves the hybridization of nucleic acid primers (e.g., DNA primers) with complementarity to nucleic acid (e.g., DNA) regions just upstream or downstream of the target sequence or off-target sequence of interest. Because the DNA target sequence and the Cas9- independent off -target sequences are known a priori in the methods disclosed herein, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the target sequence and Cas9-independent off-target sequences of interest may be designed using techniques known in the art, such as the PhusionU PCR kit (Life Technologies), Phusion HS II kit (Life Technologies), and Illumina MiSeq kit. The number of off -target DNA edits may be measured by techniques known in the art, including high-throughput screening of sequencing reads, EndoV-Seq, GUIDE- Seq, CIRCLE-Seq, Cas-OFFinder, and CRISPResso.
[0071] Since many of the Cas9-dependent off-target sites have high sequence identity to the target site of interest, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9-dependent off-target site may likewise be designed using techniques and kits known in the art. These kits make use of polymerase chain reaction (PCR) amplification, which14FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 produces amplicons as intermediate products. The target and off-target sequences may comprise genomic loci that further comprise protospacers and PAMs. Accordingly, the term “amplicons” as used herein, may refer to nucleic acid molecules that constitute the aggregates of genomic loci, protospacers and PAMs. High-throughput sequencing techniques used herein may further include Sanger sequencing and Illumina-based next-generation genome sequencing (NGS).
[0072] The term “on-target editing” as used herein, refers to the introduction of intended modifications to nucleotides (e.g., cytosine) in a target sequence, such as using the prime editors described herein. The term “off-target DNA editing” as used herein, refers to the introduction of unintended modifications to nucleotides (e.g., cytosine) in a sequence outside the canonical base editor binding window (i.e., from one protospacer position to another, typically 2 to 8 nucleotides long). Off-target DNA editing can result from weak or non-specific binding of the gRNA sequence to the target sequence. As used herein, the term “bystander editing” refers to synonymous off-target point mutations at nucleobases that are near (proximate to) the target base and do not change the outcome of the intended editing method.
[0073] “Engineered” as used herein refers, in general, to the aspect of having been manipulated by the hand of man. For example, in some embodiments, a polynucleotide may be considered to be "engineered" when two or more sequences that are not linked together in that order in nature are manipulated by the hand of man to be directly linked to one another in the engineered polynucleotide. In some embodiments, an engineered polynucleotide may comprise a regulatory sequence that is found in nature in operative association with a first coding sequence but not in operative association with a second coding sequence, is linked by the hand of man so that it is operatively associated with the second coding sequence. Alternatively, or additionally, in some embodiments, first and second nucleic acid sequences that each encode polypeptide elements or domains that in nature are not linked to one another may be linked to one another in a single engineered polynucleotide. Comparably, in some embodiments, a cell or organism may be considered to be "engineered" if it has been manipulated so that its genetic information is altered (e.g., new genetic material not previously present has been introduced, or previously present genetic material has been altered or removed). As is common practice and is understood by persons of skill in the art, progeny of an engineered polynucleotide or cell are typically still referred to as "engineered" even though the actual manipulation was performed on a prior entity. Furthermore, as will be appreciated by persons of skill in the art, a variety of methodologies are available through15FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 which “engineering” as described herein may be achieved. For example, in some embodiments, “engineering” may involve selection or design (e.g., of nucleic acid sequences, polypeptide sequences, cells, tissues, and / or organisms) through use of computer systems programmed to perform analysis or comparison, or otherwise to analyze, recommend, and / or select sequences, alterations, etc.). Alternatively, or additionally, in some embodiments, “engineering” may involve use of in vitro chemical synthesis methodologies and / or recombinant nucleic acid technologies such as, for example, nucleic acid amplification (e.g., via the polymerase chain reaction) hybridization, mutation, transformation, transfection, etc., and / or any of a variety of controlled mating methodologies. As will be appreciated by those skilled in the art, a variety of established such techniques (e.g., for recombinant DNA, oligonucleotide synthesis, and tissue culture and transformation (e.g., electroporation, lipofection, etc.)) are well known in the art and described in various general and more specific references that are cited and / or discussed throughout the present specification. See e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989 and Principles of Gene Manipulation: An Introduction to Genetic Manipulation, 5th Ed., ed. By Old, R.W. and S.B. Primrose, Blackwell Science, Inc., 1994, incorporated herein by reference in their entireties.
[0074] An “expression vector” or “expression construct” refers to a nucleic acid construct, which when introduced into a host cell, results in transcription and / or translation of a RNA and / or polypeptide, respectively. The expression vector may include a nucleic acid comprising a promoter sequence, with or without a sequence containing mRNA polyadenylation signals, and one or more restriction enzyme sites located downstream from the promoter allowing insertion of heterologous gene sequences. In some embodiments, the expression vector is capable of directing the expression of a heterologous pegRNA when the gene encoding the heterologous pegRNA is operably linked to the promoter by insertion into one of the restriction sites. In some embodiments, the recombinant expression construct allows expression of the heterologous pegRNA in a host cell when the xpression construct containing the heterologous pegRNA is introduced into the host cell. Expression constructs can be derived from a variety of sources depending on the host cell to be used for expression. For example, an expression construct can contain components derived from a viral, bacterial, insect, plant, or mammalian source. In the case of both expression of transgenes and inhibition of endogenous genes (e.g., by antisense, or sense suppression) the inserted polynucleotide16FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 sequence need not be identical and can be “substantially identical” to a sequence of the gene from which it was derived.
[0075] The term “extension arm” as used herein refers to a nucleotide sequence component of a pegRNA which provides several functions, including a primer binding site and an edit template for reverse transcriptase. In some embodiments, the extension arm is located at the 3' end of the guide RNA. In other embodiments, the extension arm is located at the 5' end of the guide RNA. In some embodiments, the extension arm also includes a homology arm. In some embodiments, the extension arm comprises the following components in a 5' to 3' direction: the homology arm, the edit template, and the primer binding site. Since polymerization activity of the reverse transcriptase is in the 5' to 3' direction, the preferred arrangement of the homology arm, edit template, and primer binding site is in the 5' to 3' direction such that the reverse transcriptase, once primed by an annealed primer sequence, polymerases a single strand of DNA using the edit template as a complementary template strand.
[0076] The term “fusion protein” as used herein refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy-terminal (C-terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic domain of a nucleic-acid editing protein. Another example includes a Cas9 or equivalent thereof fused to an adenosine deaminae. Any of the proteins described herein may be produced by any method known in the art. For example, the proteins described herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.
[0077] As used herein, the term “gene editing complex” refers to a molecular system designed to introduce specific modifications to a target polynucleotide (e.g., DNA) sequence. The complex typically includes components that recognize a desired target site in the genome and17FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 execute nucleotide alterations, such as substitutions, insertions, or deletions, in a controlled and precise manner. In the context of prime editing, the gene editing complex includes a Cas9 nickase enzyme linked to a reverse transcriptase. In some embodiments, the prime editing complex operates in conjunction with a prime editing guide RNA (pegRNA), which comprises both a targeting sequence for guiding the complex to a specific genomic site and a template for the intended DNA modification. In some embodiments, the prime editing complex induces a single-strand nick at the target site and introduces the desired nucleotide changes through reverse transcription, effectively allowing precise genetic edits without generating double-strand DNA breaks. In base editing, the gene editing complex employs a catalytically inactive or nickase variant of Cas9 fused to a cytidine deaminase or adenine deaminase enzyme, depending on the desired nucleotide conversion. In some embodiments, the base editing complex is guided to a target site by a single-guide RNA (sgRNA). In some embodiments, the cytosine base editors (CBEs), cytidine deaminase converts a target cytosine to uracil, which is then repaired into thymine (C to T conversion). In some embodiments, base editing allows the conversion of a single base pair without introducing double-strand breaks, making it particularly suitable for correcting point mutations.
[0078] As used herein, the term “guide RNA” is a particular type of guide nucleic acid which is mostly commonly associated with a Cas protein of a CRISPR-Cas system and which associates with and directs the Cas protein to a specific sequence in a DNA molecule that includes complementarity to protospacer sequence of the guide RNA. This term includes guide RNAs that associate with any Cas protein, including but not limited to, Cas9, Casl2a (Cpfl), Casl3, Cas 14, as well as equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant). Without being bound by theory, a guide RNA serves to program the Cas protein to localize to a specific target nucleotide sequence. As used herein, the “guide RNA” may also be referred to as a “traditional guide RNA” to contrast it with the modified forms of guide RNA termed “prime editing guide RNAs” (or “pegRNAs”) which may be used in prime editing methods and compositions as described herein.
[0079] As used herein, the term “host cell” refers to a cell into which a nucleic acid or protein has been introduced. Persons of skill upon reading this disclosure will understand that such a term refers not only to the particular subject cell, but also is used to refer to the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still18FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 included within the scope of the phrase “host cell” In some embodiments, a host cell is or comprises a prokaryotic or eukaryotic cell. In general, a host cell is any cell that is suitable for receiving and / or producing a heterologous nucleic acid or protein, regardless of the Kingdom of life to which the cell is designated. Exemplary cells include those of prokaryotes and eukaryotes (single-cell or multiple-cell), bacterial cells (e.g., strains of Escherichia coli, Bacillus spp., Streptomyces spp., etc.), mycobacteria cells, fungal cells, yeast cells (e.g., Saccharomyces cerevisiae, Schizo saccharomyces pombe, Pichia pastoris, Pichia methanolica, etc.), plant cells, insect cells (e.g., SF-9, SF-21, baculovirus-infected insect cells, Trichoplusia ni, etc.), non-human animal cells, human cells, or cell fusions such as, for example, hybridomas or quadromas. In some embodiments, a cell is a human, monkey, ape, hamster, rat, or mouse cell. In some embodiments, a cell is eukaryotic and is selected from the following cells: Chinese Hamster Ovarian (CHO) (e.g., CHO KI, DXB-11 CHO, Veggie-CHO), COS (e.g., COS-7), retinal cell, Vero, CV1, kidney (e.g., HEK293, 293 EBNA, MSR 293, MDCK, HaK, BHK), HeLa, HepG2, WI38, MRC 5, Colo205, HB 8065, HL-60, (e.g., BHK21), Jurkat, Daudi, A431 (epidermal), CV-1, U937, 3T3, L cell, C127 cell, SP2 / 0, NS-0, MMT 060562, Sertoli cell, BRL 3A cell, HT1080 cell, myeloma cell, tumor cell, and a cell line derived from an aforementioned cell. In some embodiments, a cell comprises one or more viral genes, e.g., a retinal cell that expresses a viral gene (e.g., a PER.C6® cell). In some embodiments, a host cell is or comprises an isolated cell. In some embodiments, a host cell is part of a tissue. In some embodiments, a host cell is part of an organism.
[0080] As used herein, the term “linker” refers to a molecule linking two other molecules or moieties. The linker can be an amino acid sequence in the case of a linker joining two fusion proteins. For example, a Cas9 can be fused to a reverse transcriptase by an amino acid linker sequence. The linker can also be a nucleotide sequence in the case of joining two nucleotide sequences together. For example, in the instant case, the traditional guide RNA is linked via a spacer or linker nucleotide sequence to the RNA extension of a prime editing guide RNA which may comprise a RT template sequence and an RT primer binding site. In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90- 100, 100-150, or 150-200 amino acids in length. Eonger or shorter linkers are also contemplated.19FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0081] The term “mutation” as used herein, refers to a substitution of a residue within a sequence, e.g. a nucleic acid or amino acid sequence, with another residue; a deletion or insertion of one or more residues within a sequence; or a substitution of a residue within a sequence of a genome to be corrected. Mutations are typically described herein by identifying the original residue followed by the position of the residue within the sequence and by the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). Mutations can include a variety of categories, such as single base polymorphisms, microduplication regions, indel, and inversions, and is not meant to be limiting in any way. Mutations can include “loss-of-function” mutations which are mutations that reduce or abolish a protein activity. Most loss-of-function mutations are recessive, because in a heterozygote the second chromosome copy carries an unmutated version of the gene coding for a fully functional protein whose presence compensates for the effect of the mutation. There are some exceptions where a loss- of-function mutation is dominant, one example being haploinsufficiency, where the organism is unable to tolerate the approximately 50% reduction in protein activity suffered by the heterozygote. This is the explanation for a few genetic diseases in humans, including Marfan syndrome, which results from a mutation in the gene for the connective tissue protein called fibrillin. Mutations also embrace “gain-of-function” mutations, which is one which confers an abnormal activity on a protein or cell that is otherwise not present in a normal condition. Many gain-of-function mutations are in regulatory sequences rather than in coding regions, and can therefore have a number of consequences. Because of their nature, gain-of-function mutations are usually dominant. Many loss- of-function mutations are recessive, such as autosomal recessive. The GNE mutation for which the presently disclosed prime editing methods aim to correct is autosomal recessive.
[0082] The term “napDNAbp” which stand for “nucleic acid programmable DNA binding protein” refers to any protein that may associate (e.g., form a complex) with one or more nucleic acid molecules (i.e., which may broadly be referred to as a “napDNAbp-programming nucleic acid molecule” and includes, for example, guide RNA in the case of Cas systems) which direct or otherwise program the protein to localize to a specific target nucleotide sequence (e.g., a gene locus of a genome) that is complementary to the one or more nucleic acid molecules (or a portion or region thereof) associated with the protein, thereby causing the protein to bind to the nucleotide20FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 sequence at the specific target site. This term napDNAbp embraces CRISPR-Cas9 proteins, as well as Cas9 equivalents, homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or modified), and may include a Cas9 equivalent from any type of CRISPR system (e.g., type II, V, VI), including Cpfl (a type-V CRISPR-Cas systems), C2cl (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system), C2c3 (a type V CRISPR-Cas system), dCas9, GeoCas9, CjCas9, Casl2a, Casl2b, Casl2c, Casl2d, Casl2g, Casl2h, Casl2i, Casl3d, Casl4, Argonaute, and nCas9. Further Cas-equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector” Science 2016; 353 (6299), the contents of which are incorporated herein by reference. However, the nucleic acid programmable DNA binding protein (napDNAbp) that may be used in connection with the disclosure(s) described herein are not limited to CRISPR-Cas systems. Provided technologies embrace any such programmable protein, such as the Argonaute protein from Natronobacterium gregoryi (NgAgo) which may also be used for DNA-guided genome editing. NgAgo-guide DNA system does not require a PAM sequence or guide RNA molecules, which means genome editing can be performed simply by the expression of generic NgAgo protein and introduction of synthetic oligonucleotides on any genomic sequence. See Gao et al., DNA-guided genome editing using the Natronobacterium gregoryi Argonaute. Nature Biotechnology 2016; 34(7):768-73, which is incorporated herein by reference.
[0083] As used herein, a “nickase” refers to a napDNAbp (e.g., a Cas protein) which is capable of cleaving only one of the two complementary strands of a double- stranded target DNA sequence, thereby generating a nick in that strand. In some embodiments, the nickase cleaves a nontarget strand of a double stranded target DNA sequence. In some embodiments, the nickase comprises an amino acid sequence with one or more mutations in a catalytic domain of a canonical napDNAbp (e.g., a Cas protein), wherein the one or more mutations reduces or abolishes nuclease activity of the catalytic domain. In some embodiments, the nickase is a Cas9 that comprises one or more mutations in a RuvC-like domain relative to a wild type Cas9 sequence or to an equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the nickase is a Cas9 that comprises one or more mutations in a HNH-like domain relative to a wild type Cas9 sequence or to an equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the nickase is a Cas9 that comprises an aspartate-to-alanine substitution (DIO A) in the RuvC I catalytic domain of Cas9 relative to a canonical Cas9 sequence or to an21FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the nickase is a Cas9 that comprises a H840A, N854A, and / or N863A mutation relative to a canonical Cas9 sequence, or to an equivalent amino acid position in other Cas9 variants or Cas9 equivalents. In some embodiments, the term “Cas9 nickase” refers to a Cas9 with one of the two nuclease domains inactivated. This enzyme is capable of cleaving only one strand of a target DNA. In some embodiments, the nickase is a Cas protein that is not a Cas9 nickase.
[0084] As used herein, the term “polymerase” refers to an enzyme that synthesizes a nucleotide strand and which may be used in connection with the prime editor systems described herein. The polymerase can be a “template-dependent” polymerase (i.e., a polymerase which synthesizes a nucleotide strand based on the order of nucleotide bases of a template strand). The polymerase can also be a “template-independent” polymerase (i.e., a polymerase which synthesizes a nucleotide strand without the requirement of a template strand). A polymerase may also be further categorized as a “DNA polymerase” or an “RNA polymerase.” In some embodiments, the prime editor system comprises a DNA polymerase. In some embodiments, the DNA polymerase can be a “DNA-dependent DNA polymerase” (i.e., whereby the template molecule is a strand of DNA). In such cases, the DNA template molecule can be a pegRNA, wherein the extension arm comprises a strand of DNA. In such cases, the pegRNA may be referred to as a chimeric or hybrid pegRNA which comprises an RNA portion (i.e., the guide RNA components, including the spacer and the gRNA core) and a DNA portion (i.e., the extension arm). In some embodiments, the DNA polymerase can be an “RNA-dependent DNA polymerase” (i.e., whereby the template molecule is a strand of RNA). In such cases, the pegRNA is RNA, i.e., including an RNA extension. The term “polymerase” may also refer to an enzyme that catalyzes the polymerization of nucleotide (i.e., the polymerase activity). Generally, the enzyme will initiate synthesis at the 3 '-end of a primer annealed to a polynucleotide template sequence (e.g., such as a primer sequence annealed to the primer binding site of a pegRNA), and will proceed toward the 5' end of the template strand. A “DNA polymerase” catalyzes the polymerization of deoxynucleotides.
[0085] As used herein, the term “prime editing” refers to an approach for gene editing using napDNAbps, a polymerase (e.g., a reverse transcriptase), and specialized guide RNAs that include a Reverse transcription template for encoding desired new genetic information (or deleting genetic information) that is then incorporated into a target DNA sequence.22FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0086] As used herein, the terms “prime editing guide RNA” or “pegRNA” or “extended guide RNA” refer to a specialized form of a guide RNA that has been modified to include one or more additional sequences for implementing the prime editing methods and compositions described herein. As described herein, the prime editing guide RNA comprise one or more “extended regions” of nucleic acid sequence. The extended regions may comprise, but are not limited to, singlestranded RNA or DNA. Further, the extended regions may occur at the 3' end of a traditional guide RNA. In other arrangements, the extended regions may occur at the 5' end of a traditional guide RNA. In still other arrangements, the extended region may occur at an intramolecular region of the traditional guide RNA, for example, in the gRNA core region which associates and / or binds to the napDNAbp. The extended region comprises a “Reverse transcription template ” which encodes (by the polymerase of the prime editor) a single- stranded DNA which, in turn, has been designed to be (a) homologous with the endogenous target DNA to be edited, and (b) which comprises at least one desired nucleotide change (e.g., a transition, a transversion, a deletion, or an insertion) to be introduced or integrated into the endogenous target DNA. The extended region may also comprise other functional sequence elements, such as, but not limited to, a “primer binding site” and a “spacer or linker” sequence, or other structural elements, such as, but not limited to aptamers, stem loops, hairpins, toe loops (e.g., a 3' toeloop), or an RNA-protein recruitment domain (e.g., MS2 hairpin). As used herein the “primer binding site” comprises a sequence that hybridizes to a singlestrand DNA sequence having a 3' end generated from the nicked DNA of the R-loop.
[0087] The term “prime editor” refers to the herein described fusion constructs comprising a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase and is capable of carrying out prime editing on a target nucleotide sequence in the presence of a pegRNA (or “extended guide RNA”). The term “prime editor” may refer to the fusion protein or to the fusion protein complexed with a pegRNA, and / or further complexed with a second-strand nicking sgRNA. In some embodiments, the prime editor may also refer to the complex comprising a fusion protein (reverse transcriptase fused to a napDNAbp), a pegRNA, and a regular guide RNA capable of directing the second-site nicking step of the non-edited strand as described herein. In other embodiments, the reverse transcriptase component of the “primer editor” may be provided in trans.
[0088] The term “primer binding site” or “the PBS” refers to the portion of nucleotide sequence located on a pegRNA as component of the extension arm (typically for example, at the 3' end of the extension arm). The term “primer binding site” refers to a single-stranded portion of the23FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 PEgRNA as a component of the extension arm that comprises a region of complementarity to a sequence on the non-target strand. In some embodiments, the primer binding site is complementary to a region upstream of a nick site in a non-target strand. In some embodiments, the primer binding site is complementary to a region immediately upstream of a nick site in the non-target strand. In some embodiments, the primer binding site is capable of binding to the primer sequence that is formed after nicking of the target sequence by the prime editor. When the prime editor nicks one strand of the target DNA sequence (e.g., by a Cas nickase component of the prime editor), a 3'- ended ssDNA flap is formed, which serves a primer sequence that anneals to the primer binding site on the pegRNA to prime reverse transcription. In some embodiments, the PBS is complementary to or substantially complementary to, and can anneal to a free 3' end on the non-target strand of the double stranded target DNA at the nick site. In some embodiments, the PBS annealed to the free 3' end on the non-target strand can initiate target-primed DNA synthesis.
[0089] The terms “promoter” or “promoter sequence” refer generally to transcriptional regulatory regions of a gene, which may be found at the 5’ or 3’ side of the coding region, or within the coding region, or within introns. Typically, a promoter is a nucleic acid regulatory region capable of binding RNA polymerase in a cell and control initiation and the rate of transcription of a downstream (3’ direction) coding or non-coding sequence. The typical 5’ promoter sequence is bounded at its 3’ terminus by the transcription initiation site and extends upstream (5’ direction) to include the minimum number of bases or elements necessary to initiate transcription at levels detectable above background. Within the promoter sequence is a transcription initiation site (conveniently defined by mapping with nuclease SI), as well as protein binding domains (consensus sequences) responsible for the binding of RNA polymerase. A promoter may also contain subregions at which regulatory proteins and molecules may bind, such as RNA polymerase and other transcription factors. Promoters may be constitutive, inducible, activatable, repressible, tissuespecific or any combination thereof. A promoter drives expression or drives transcription of the nucleic acid sequence that it regulates. A promoter may be one naturally associated with a gene or sequence, as may be obtained by isolating the 5' non-coding sequences located upstream of the coding segment of a given gene or sequence. Such a promoter is referred to as an “endogenous promoter” Examples of promoters that may be used in the context of the present disclosure include, e.g., Pol II and Pol III promoters.24FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0090] As used herein, the term “protospacer” refers to the sequence (e.g., a ~20 bp sequence) in DNA adjacent to the PAM (protospacer adjacent motif) sequence which shares the same sequence as the spacer sequence of the guide RNA, and which is complementary to the target sequence of the non-PAM strand. The spacer sequence of the guide RNA anneals to the target sequence located on the non-PAM strand. In order for Cas9 to function it also requires a specific protospacer adjacent motif (PAM) that varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes a PAM sequence of NGG that is found directly downstream of the protospacer sequence in the genomic DNA, on the non-target strand. The skilled person will appreciate that the literature in the state of the art sometimes refers to the “protospacer” as the ~20-nt target-specific guide sequence on the guide RNA itself, rather than referring to it as a “spacer” (and that the protospacer (DNA) and the spacer (RNA) have the same sequence). Thus, the term “protospacer” as used herein may be used interchangeably with the term “spacer” The context of the description surrounding the appearance of either “protospacer” or “spacer” will help inform the reader as to whether the term is reference to the gRNA or the DNA sequence. Both usages of these terms are acceptable since the state of the art uses both terms in each of these ways.
[0091] As used herein, the term “protospacer adjacent sequence” or “PAM” refers to an approximately 2-6 base pair DNA sequence that is an important targeting component of a Cas9 nuclease. Typically, the PAM sequence is on either strand, and is downstream in the 5' to 3' direction of Cas9 cut site. The canonical PAM sequence (i.e., the PAM sequence that is associated with the Cas9 nuclease of Streptococcus pyogenes or SpCas9) is 5'-NGG-3' wherein “N” is any nucleobase followed by two guanine (“G”) nucleobases. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms. In addition, any given Cas9 nuclease, e.g., SpCas9, may be modified to alter the PAM specificity of the nuclease such that the nuclease recognizes alternative PAM sequence.
[0092] The term “reverse transcriptase” describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize a DNA transcript from an RNA template. Historically, reverse transcriptase has been used primarily to transcribe mRNA into cDNA which can then be cloned into a vector for further manipulation. Avian myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme has 5'-3' RNA-25FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific for the RNA strand for RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Errors in transcription cannot be corrected by reverse transcriptase because known viral reverse transcriptases lack the 3'-5' exonuclease activity necessary for proofreading (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNase H activity has been presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase which is used extensively in molecular biology is reverse transcriptase originating from Moloney murine leukemia virus (M-MLV). See, e.g., Gerard, G. R., DNA 5:271-279 (1986) and Kotewicz, M. L., et al., Gene 35:249-258 (1985). M-MLV reverse transcriptase substantially lacking in RNase H activity has also been described. See, e.g., U.S. Pat. No. 5,244,797. The present disclosure contemplates the use of any such reverse transcriptases, or variants or mutants thereof.
[0093] As used herein, the term “reverse transcription” indicates the capability of an enzyme to synthesize a DNA strand (that is, complementary DNA or cDNA) using RNA as a template. In some embodiments, the reverse transcription can be “error-prone reverse transcription” which refers to the properties of certain reverse transcriptase enzymes which are error-prone in their DNA polymerization activity.
[0094] As used herein, the term “spacer sequence” in connection with a guide RNA or a pegRNA refers to the portion of the guide RNA or pegRNA of about 20 nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23 or 24 nucleotides) which contains a nucleotide sequence that is complementary to the target strand. In some embodiments, the spacer sequence hybridizes to a region on the target strand that is complementary to a protospacer on the non-target strand to form a ssRNA / ssDNA hybrid structure at the target site and a corresponding R loop ssDNA structure of the complementary endogenous DNA strand on the non-target strand.
[0095] As used herein, the “target sequence” refers to the ~20 nucleotides in the target DNA sequence that have complementarity to the protospacer sequence in the PAM strand. The target sequence is the sequence that anneals to or is targeted by the spacer sequence of the guide RNA. The spacer sequence of the guide RNA and the protospacer have the same sequence (except the spacer sequence is RNA, and the protospacer is DNA).26FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0096] As used herein, the “target site” refers to a sequence within a nucleic acid molecule that is edited by a prime editor disclosed herein. The term “target site” in the context of a single strand, also can refer to the “target strand” which anneals or binds to the spacer sequence of the guide RNA. The target site can refer, in certain embodiments, to a segment of double-stranded DNA that includes the protospacer (i.e., the strand of the target site that has the same nucleotide sequence as the spacer sequence of the guide RNA) on the PAM-strand (or non-target strand) and target strand, which is complementary to the protospacer and the spacer alike, and which anneals to the spacer of the guide RNA, thereby targeting or programming a Cas9 base editor to target the target site.
[0097] As used herein, the terms “upstream” and “downstream” are terms of relativity that define the linear position of at least two elements located in a nucleic acid molecule (whether single or double-stranded) that is orientated in a 5'-to-3' direction. In particular, a first element is upstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 5' to the second element. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 3' to the second element. The nucleic acid molecule can be a DNA (double or single stranded). RNA (double or single stranded), or a hybrid of DNA and RNA. The analysis is the same for single strand nucleic acid molecule and a double strand molecule since the terms upstream and downstream are in reference to only a single strand of a nucleic acid molecule, except that one needs to select which strand of the double stranded molecule is being considered. Often, the strand of a double stranded DNA which can be used to determine the positional relativity of at least two elements is the “sense” or “coding” strand. In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'.
[0098] As used herein, the term “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules that comprise one or more free ends, no free ends (e.g. circular); nucleic acid molecules that comprise DNA, RNA, or both; and other varieties of polynucleotides known in the art. One type of vector is a “plasmid” which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques.27FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 Another type of vector is a viral vector, wherein virally derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g. retroviruses, replication defective retroviruses, adenoviruses, replication defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by a virus for transfection into a host cell. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g. bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively-linked. Such vectors are referred to herein as “expression vectors”. Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids. Recombinant expression vectors can comprise a nucleic acid of the presently disclosed subject matter in a form suitable for expression of the nucleic acid in a host cell, which means that the recombinant expression vectors include one or more regulatory elements, which may be selected on the basis of the host cells to be used for expression, that is operably-linked to the nucleic acid sequence to be expressed.
[0099] As used herein, the term “5' endogenous DNA flap” refers to the strand of DNA situated immediately downstream of the PE-induced nick site in the target DNA. The nicking of the target DNA strand by PE exposes a 3' hydroxyl group on the upstream side of the nick site and a 5' hydroxyl group on the downstream side of the nick site. The endogenous strand ending in the 3' hydroxyl group is used to prime the DNA polymerase of the prime editor (e.g., wherein the DNA polymerase is a reverse transcriptase). The endogenous strand on the downstream side of the nick site and which begins with the exposed 5' hydroxyl group is referred to as the “5' endogenous DNA flap” and is ultimately removed and replaced by the newly synthesized replacement strand (i.e., “3' replacement DNA flap”) the encoded by the extension of the PEgRNA.
[0100] As used herein, the term “5' flap removal” refers to the removal of the 5' endogenous DNA flap that forms when the RT-synthesized single-strand DNA flap competitively invades and hybridizes to the endogenous DNA, displacing the endogenous strand in the process. Removing this endogenous displaced strand can drive the reaction towards the formation of the desired product comprising the desired nucleotide change. The cell’s own DNA repair enzymes may catalyze the removal or excision of the 5' endogenous flap (e.g., a flap endonuclease, such as EXO1 or FEN1). Also, host cells may be transformed to express one or more enzymes that catalyze the removal of28FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 said 5' endogenous flaps, thereby driving the process toward product formation (e.g., a flap endonuclease). Flap endonucleases are known in the art and can be found described in Patel et al., “Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends,” Nucleic Acids Research, 2012, 40(10): 4507-4519 and Tsutakawa et al., “Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily,” Cell, 2011, 145(2): 198-211 (each of which are incorporated herein by reference).
[0101] As used herein, the term “3' replacement DNA flap” or simply, “replacement DNA flap,” refers to the strand of DNA that is synthesized by the prime editor and which is encoded by the extension arm of the prime editor PEgRNA. More in particular, the 3' replacement DNA flap is encoded by the polymerase template of the PEgRNA. The 3' replacement DNA flap comprises the same sequence as the 5' endogenous DNA flap except that it also contains the edited sequence (e.g., single nucleotide change). The 3' replacement DNA flap anneals to the target DNA, displacing or replacing the 5' endogenous DNA flap (which can be excised, for example, by a 5' flap endonuclease, such as FEN1 or EXO1) and then is ligated to join the 3' end of the 3' replacement DNA flap to the exposed 5' hydoxyl end of endogenous DNA (exposed after excision of the 5' endogenous DNA flap, thereby reforming a phosophodiester bond and installing the 3' replacement DNA flap to form a heteroduplex DNA containing one edited strand and one unedited strand. DNA repair processes resolve the heteroduplex by copying the information in the edited strand to the complementary strand permanently installs the edit in to the DNA. This resolution process can be driven further to completion by nicking the unedited strand, i.e., by way of “second-strand nicking,” as described herein.DETAILED DESCRIPTIONOverview
[0102] The present disclosure provides, inter alia, compositions and methods for modifying a GNE-encoding polynucleotide (e.g., genomic DNA). GNE myopathy is a rare genetic disorder that primarily affects the skeletal muscles. It is caused by mutations in the GNE gene, which is responsible for producing an enzyme involved in sialic acid biosynthesis, a key component in muscle cell function. The disease leads to progressive muscle weakness and atrophy, typically beginning in the lower limbs (particularly the distal muscles like the foot and ankle) before29FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 progressing to the upper limbs. GNE myopathy has autosomal recessive heredity. In some embodiments, the present disclosure provides compositions, vectors, kits, and methods for modifying a polynucleotide using prime editing strategies that comprise the use of a prime editor, which includes a reverse transcriptase domain and a nickase domain, along with a suitable guide RNA (pegRNA) that targets the GNE gene in order to edit a M743T mutation. In some embodiments, the present disclosure provides compositions, vectors, kits, and methods for modifying a polynucleotide using base editing strategies that comprise the use of a nucleic acid programmable DNA binding protein (“napDNAbp”), a deaminase (e.g., cytidine deaminase) and a suitable guide RNA that targets the GNE gene. In some embodiments, prime editing is used to modify the nucleotide sequence of the GNE gene such that the GNE protein function is restored, either completely or partially. In some embodiments, prime editing or base editing is utilized to correct mutations in the GNE gene sequence by installing targeted edits, such as correcting point mutations or small insertions / deletions. This approach is referred to herein as the “mutation correction approach” In some embodiments, the presently disclosed compositions, kits, and methods utilize highly efficient prime editors or prime editors to install edits in the GNE gene with minimal off-target effects. The compositions also include novel pegRNAs or gRNAs specifically designed for the mutation correction approach. The presently disclosed compositions, kits, and methods make use of recently developed high-efficiency prime editors, such as PE3 or PE3b, or base editors, such as CBE6b, to install targeted edits in the GNE gene with a low frequency of off- target effects. The disclosed compositions and kits also utilize novel pegRNAs or gRNAs to program the prime editor or base editor to the correct site in the GNE gene.
[0103] Mutations in GNE are the primary cause of GNE myopathy, a rare neuromuscular disorder characterized by progressive muscle weakness and atrophy. Over 200 pathogenic mutations in GNE have been identified in humans, many of which are point mutations that can be corrected by prime editing. However, therapeutic applications of prime editing are constrained by the requirement for a protospacer- adjacent motif (PAM) recognition sequence and the potential for nearby bystander edits and off-target effects. Bioinformatic design tools help in identifying prime editing guide sequences with high on-target editing activity and minimal off-target effects, but empirical testing remains essential to evaluate the most effective pegRNA-prime editor combinations.30FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0104] In the examples of the present disclosure, a panel of GNE mutant sequences were integrated into the AAVS1 locus in HEK293 cells, allowing for rapid testing of prime editor activity at these disease-associated mutations using a single cell line. These mutant sequences were selected based on bioinformatic design criteria and their prevalence in humans. Using prime editors, multiple mutations of GNE were identified that showed moderate efficiency of editing (>15%) with a low indel frequency (~3%). Off-target effects at targeted editing sites were empirically evaluated for promising targets. Differences in editing efficiency were observed between various testing methods, including i) integrated editing target versus target plasmid transfection; ii) single-plex versus multiplex guide RNA transfection; and iii) Sanger sequencing versus next-generation sequencing (NGS) editing efficacy evaluation. Testing of prime editors delivered by AAV in vivo is ongoing, with the goal of improving the rate and throughput of prime editor target identification in cultured cells and in vivo.
[0105] In some embodiments, the disclosure provides pegRNA sequences capable of directing prime editors to positions within the GNE gene to correct mutations associated with GNE myopathy. In some embodiments, the disclosure provides prime editing complexes that mediate precise nucleotide changes in the GNE gene to restore GNE enzyme function.
[0106] In some embodiments, the disclosure provides base editors that can be utilized for targeted editing of specific mutations in the GNE gene to correct the genetic defect, restoring the normal reading frame and producing a functional GNE protein.GNE myopathy
[0107] GNE myopathy (also known as hereditary inclusion body myopathy, quadriceps sparing myopathy, distal myopathy with rimmed vacuoles, and Nonaka myopathy) is an adult-onset autosomal recessive genetic disease characterized by progressive muscle weakness that is caused by loss of function pathogenic variants or mutations in the GNE gene. The GNE gene encodes a bifunctional UDP-GIcNAc-epimerase / ManNAc-6 kinase, whose enzymatic activities are essential in sialic acid biosynthetic pathway. Sialic acid is an acidic monosaccharide that modifies nonreducing terminal carbohydrate chains on glycoproteins and glycolipids and plays an important role in different processes such as cell-adhesion and cellular interactions. Sialic acid has been implicated in health and disease and is found in terminal sugar chains of proteins modulating their cellular functions. As UDP-N-acetylglucosamine 2-epimerase / N- acetylmannosamine kinase (GNE) is the31FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 key enzyme for the biosynthesis of sialic acid. Moreover, it has been demonstrated that GNE expression is induced when myofibers are damaged or regenerating, and that GNE plays a role in muscle regeneration.
[0108] Mutations in the GNE gene are the primary cause of GNE myopathy, a rare neuromuscular disorder affecting muscle function. GNE myopathy has an estimated prevalence of 1 in 1,000,000 people globally. Individuals with this disorder experience progressive muscle weakness and atrophy, usually starting in the lower limbs, followed by upper limb involvement. This progression eventually leads to severe disability but does not typically impact heart or respiratory muscles. The condition stems from decreased sialic acid production, a key molecule necessary for muscle cell stability and function.
[0109] The human GNE gene comprises 13 coding exons, with multiple mutations identified, many of which are missense mutations in exons 5-7, commonly affecting the epimerase domain. These mutations hinder sialic acid biosynthesis, resulting in impaired enzyme function and stability. For example, the c.527A>T (p.Metl76Lys) mutation is a recurrent mutation that leads to reduced enzyme function.
[0110] The GNE protein (GenBank Acc. No. NM_001128227.3) is a cytosolic enzyme that catalyzes two key steps in sialic acid biosynthesis. Its stability and activity are critical for effective muscle function. The enzyme structure includes homologous domains, with regions that exhibit high sequence conservation, reflecting its vital role in cellular function.
[0111] Despite the fact that mutations in the GNE gene were shown to cause GNE myopathy in 2001, there are as yet no effective therapies for this disease. Attempts to develop slow- release sialic acid therapy failed in a phase 3 clinical trial, and ManNAc glycan therapy is currently being investigated. While development of a gene therapy approach for GNE gene replacement might seem straightforward, it is in fact complicated by a number of unresolved issues in GNE myopathy research. Such issues include the fact that there are currently no robust and reproducible models for GNE myopathy. While Noguchi and Nishino published several papers on a transgenic GNED176VTg Gne / _mouse model showing clear aspects of disease pathology, other groups, have failed to see the same phenotypes with subsequent breeding, likely the result of genetic drift in the founder transgenic line (Nishino et al., (2015) Journal of Neurology, Neurosurgery & Psychiatry 86(4): 385-392). A GNEM712T variant knock-in mouse model showed premature death in the first few weeks of life due to kidney disease, a clinical phenotype that is not present in GNE myopathy32FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 patients. Other lines of the same model were bred out to show no phenotype at all despite having the same genetic mutation. Another issue is the late onset of disease and the highly variable disease progression, results in a lack of robust short-term clinical milestones.Prime editing
[0112] Prime editing represents a platform for genome editing that is a versatile and precise genome editing method that directly writes new genetic information into a specified DNA site using a nucleic acid programmable DNA binding protein (“napDNAbp”) working in association with a polymerase (i.e., in the form of a fusion protein or otherwise provided in trans with the napDNAbp), wherein the prime editing system is programmed with a prime editing (PE) guide RNA (“pegRNA”) that both specifies the target site and templates the synthesis of the desired edit in the form of a replacement DNA strand by way of an extension (either DNA or RNA) engineered onto a guide RNA (e.g., at the 5' or 3' end, or at an internal portion of a guide RNA). The replacement strand containing the desired edit (e.g., a single nucleobase substitution) shares the same (or is homologous to) sequence as the endogenous strand (immediately downstream of the nick site) of the target site to be edited (with the exception that it includes the desired edit). Through DNA repair and / or replication machinery, the endogenous strand downstream of the nick site is replaced by the newly synthesized replacement strand containing the desired edit. In some embodiments, prime editing may be thought of as a “search-and-replace” genome editing technology since the prime editors, as described herein, not only search and locate the desired target site to be edited, but at the same time, encode a replacement strand containing a desired edit which is installed in place of the corresponding target site endogenous DNA strand.
[0113] In some embodiments, the present disclosure relates to Cas protein-reverse transcriptase fusions or related systems to target a specific DNA sequence with a guide RNA, generate a single strand nick at the target site, and use the nicked DNA as a primer for reverse transcription of an engineered reverse transcriptase template that is integrated with the guide RNA. However, while the concept begins with prime editors that use reverse transcriptase as the DNA polymerase component, the prime editors described herein are not limited to reverse transcriptases but may include the use of virtually any DNA polymerase. Indeed, while the application throughout may refer to prime editors with “reverse transcriptases,” it is set forth here that reverse transcriptases are only one type of DNA polymerase that may work with prime editing. Thus, where33FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 ever the specification mentions a “reverse transcriptase,” the person having ordinary skill in the art should appreciate that any suitable DNA polymerase may be used in place of the reverse transcriptase. Thus, in some embodiments, the prime editors may comprise Cas9 (or an equivalent napDNAbp) which is programmed to target a DNA sequence by associating it with a specialized guide RNA (i.e., pegRNA) containing a spacer sequence that anneals to a complementary protospacer in the target DNA. The specialized guide RNA also contains new genetic information in the form of an extension that encodes a replacement strand of DNA containing a desired genetic alteration which is used to replace a corresponding endogenous DNA strand at the target site. To transfer information from the pegRNA to the target DNA, the mechanism of prime editing involves nicking the target site in one strand of the DNA to expose a 3'-hydroxyl group. The exposed 3'- hydroxyl group can then be used to prime the DNA polymerization of the edit-encoding extension on PEgRNA directly into the target site.
[0114] In some embodiments, the extension — which provides the template for polymerization of the replacement strand containing the edit — can be formed from RNA or DNA. In some embodiments, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (such as, a reverse transcriptase). In some embodiments, the polymerase of the prime editor may be a DNA-dependent DNA polymerase. The newly synthesized strand that is formed by the herein disclosed prime editors would be homologous to the genomic target sequence (i.e., have the same sequence as) except for the inclusion of a desired nucleotide change (e.g., a single nucleotide change, a deletion, or an insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA may also be referred to as a single strand DNA flap, which would compete for hybridization with the complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. In some embodiments, the system can be combined with the use of an error-prone reverse transcriptase enzyme (e.g., provided as a fusion protein with the Cas9 domain, or provided in trans to the Cas9 domain). The error-prone reverse transcriptase enzyme can introduce alterations during synthesis of the single strand DNA flap. In some embodiments, error-prone reverse transcriptase can be utilized to introduce nucleotide changes to the target DNA. Depending on the error-prone reverse transcriptase that is used with the system, the changes can be random or non-random. Resolution of the hybridized intermediate (comprising the single strand DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand) can include removal of the resulting displaced flap of endogenous DNA34FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 (e.g., with a 5' end DNA flap endonuclease, FEN1), ligation of the synthesized single strand DNA flap to the target DNA, and assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because templated DNA synthesis offers single nucleotide precision for the modification of any nucleotide, including insertions and deletions, the scope of this approach is very broad and could foreseeably be used for myriad applications in basic science and therapeutics.
[0115] In some embodiments, prime editing operates by contacting a target DNA molecule (for which a change in the nucleotide sequence is desired to be introduced) with a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editing guide RNA (PEgRNA). In some embodiments, the prime editing guide RNA (PEgRNA) comprises an extension at the 3' or 5' end of the guide RNA, or at an intramolecular location in the guide RNA and encodes the desired nucleotide change (e.g., single nucleotide change, insertion, or deletion In some embodiments, the napDNAbp / extended gRNA complex contacts the DNA molecule and the extended gRNA guides the napDNAbp to bind to a target locus. In some embodiments, a nick in one of the strands of DNA of the target locus is introduced (e.g., by a nuclease or chemical agent), thereby creating an available 3' end in one of the strands of the target locus. In some embodiments, the nick is created in the strand of DNA that corresponds to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence, i.e., the “non-target strand.” The nick, however, could be introduced in either of the strands. That is, the nick could be introduced into the R-loop “target strand” (i.e., the strand hybridized to the protospacer of the extended gRNA) or the “non-target strand” (i.e., the strand forming the single-stranded portion of the R-loop and which is complementary to the target strand). In some embodiments, the 3' end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA in order to prime reverse transcription (i.e., “target-primed RT”). In some embodiments, the 3' end DNA strand hybridizes to a specific RT priming sequence on the extended portion of the guide RNA, i.e., “primer binding site” on the pegRNA. In step (d), a reverse transcriptase (or other suitable DNA polymerase) is introduced which synthesizes a single strand of DNA from the 3' end of the primed site towards the 5' end of the prime editing guide RNA. The DNA polymerase (e.g., reverse transcriptase) can be fused to the napDNAbp or alternatively can be provided in trans to the napDNAbp. This forms a single-strand DNA flap comprising the desired nucleotide change (e.g., the single base change, insertion, or deletion, or a combination thereof) and which is otherwise homologous to the35FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 endogenous DNA at or adjacent to the nick site. In some embodiments, the napDNAbp and guide RNA are released. In some embodiments, the single strand DNA flap becomes incorporated into the target locus. This process can be driven towards the desired product formation by removing the corresponding 5' endogenous DNA flap that forms once the 3' single strand DNA flap invades and hybridizes to the endogenous DNA sequence. Without being bound by theory, the cells endogenous DNA repair and replication processes resolves the mismatched DNA to incorporate the nucleotide change(s) to form the desired altered product. In some embodiments, this process may introduce at least one or more of the following genetic changes: trans versions, transitions, deletions, and insertions.
[0116] The term “prime editor (PE) system” or “prime editor (PE)” or “PE system” or “PE editing system” refers the compositions involved in the method of genome editing using target- primed reverse transcription (TPRT) describe herein, including, but not limited to the napDNAbps, reverse transcriptases, fusion proteins (e.g., comprising napDNAbps and reverse transcriptases), prime editing guide RNAs, and complexes comprising fusion proteins and prime editing guide RNAs, as well as accessory elements, such as second strand nicking components (e.g., second strand sgRNAs) and 5' endogenous DNA flap removal endonucleases (e.g., FEN1) for helping to drive the prime editing process towards the edited product formation.
[0117] Although in the embodiments described thus far the pegRNA constitutes a single molecule comprising a guide RNA (which itself comprises a spacer sequence and a gRNA core or scaffold) and a 5' or 3' extension arm comprising the primer binding site and a Reverse transcription template .Prime editing guide RNA (pegRNA)
[0118] The prime editing system described herein contemplates the use of any suitable pegRNAs. The mechanism of target-primed reverse transcription (TPRT) can be leveraged or adapted for conducting precision and versatile CRISPR / Cas-based genome editing through the use of a specially configured guide RNAs comprising a reverse transcription template (RTT) sequence that codes for the desired nucleotide change. The application refers to this specially configured guide RNA as an “extended guide RNA” or a “pegRNA” since the RTT sequence can be provided as an extension of a standard or traditional guide RNA molecule. The application contemplates any suitable configuration or arrangement for the pegRNA.36FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0119] In some embodiments, the reverse transcription template (RTT) is a single-stranded portion of the pegRNA that is 5' of the PBS and comprises a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand), and comprises one or more nucleotide edits compared to the endogenous sequence of the double stranded target DNA. In some embodiments, the reverse transcription template is complementary or substantially complementary to a sequence on the non-target strand that is downstream of a nick site, except for one or more non- complementary nucleotides at the intended nucleotide edit positions. In some embodiments, the reverse transcription template is complementary or substantially complementary to a sequence on the non-target strand that is immediately downstream (i.e., directly downstream) of a nick site, except for one or more non-complementary nucleotides at the intended nucleotide edit positions. In some embodiments, one or more of the non-complementary nucleotides at the intended nucleotide edit positions are immediately downstream of a nick site. In some embodiments, the reverse transcription template comprises one or more nucleotide edits relative to the double-stranded target DNA sequence. In some embodiments, the reverse transcription template comprises one or more nucleotide edits relative to the non-target strand of the double- stranded target DNA sequence. For each pegRNA described herein, a nick site is characteristic of the particular napDNAbp to which the gRNA core of the pegRNA associates with, and is characteristic of the particular PAM required for recognition and function of the napDNAbp. For example, for a pegRNA that comprises a gRNA core that associates with a SpCas9, the nick site in the phosphodiester bond between bases three (“-3” position relative to the position 1 of the PAM sequence) and four (“-4” position relative to position 1 of the PAM sequence).
[0120] In some embodiments, the reverse transcription template and the primer binding site are immediately adjacent to each other. The terms “nucleotide edit”, “nucleotide change”, “desired nucleotide change”, and “desired nucleotide edit” are used interchangeably to refer to a specific nucleotide edit, e.g., a specific deletion of one or more nucleotides, a specific insertion of one or more nucleotides, a specific substitution(s) of one or more nucleotides, or a combination thereof, at one a specific position in a reverse transcription template of a pegRNA to be incorporated in a target DNA sequence. In some embodiments, the reverse transcription template comprises more than one nucleotide edits relative to the double-stranded target DNA sequence. In some embodiments, each nucleotide edit is a specific nucleotide edit at a specific position in the Reverse transcription template , each nucleotide edit is at a different specific position relative to any of the37FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 other nucleotide edits in the reverse transcription template , and each nucleotide edit is independently selected from a specific deletion of one or more nucleotides, a specific insertion of one or more nucleotides, a specific substitution(s) of one or more nucleotides, or a combination thereof. A nucleotide edit may refer to the edit on the reverse transcription template as compared to the sequence on the target strand of the target gene, or may refer to the edit encoded by the Reverse transcription template on the newly synthesized single stranded DNA that replaces the endogenous target DNA sequence on the non-target strand, in either case, may be refer to as a nucleotide edit compared to the target DNA sequence.
[0121] The extension arm may also be described as comprising generally two regions: a primer binding site (PBS) and a reverse transcription template (RTT), for instance. The primer binding site binds to the primer sequence that is formed from the endogenous DNA strand of the target site when it becomes nicked by the prime editor complex, thereby exposing a 3' end on the endogenous nicked strand. As explained herein, the binding of the primer sequence to the primer binding site on the extension arm of the pegRNA creates a duplex region with an exposed 3' end (i.e., the 3' of the primer sequence), which then provides a substrate for a polymerase to begin polymerizing a single strand of DNA from the exposed 3' end along the length of the Reverse transcription template . The sequence of the single strand DNA product is the complement of the reverse transcription template. Polymerization continues towards the 5' of the reverse transcription template (or extension arm) until polymerization terminates. Thus, the reverse transcription template represents the portion of the extension arm that is encoded into a single strand DNA product (i.e., the 3' single strand DNA flap containing the desired genetic edit information) by the polymerase of the prime editor complex and which ultimately replaces the corresponding endogenous DNA strand of the target site that sits immediately downstream of the PE-induced nick site. Without being bound by theory, polymerization of the reverse transcription template continues towards the 5' end of the extension arm until a termination event. In some embodiments, polymerization may terminate in a variety of ways, including, but not limited to (a) reaching a 5' terminus of the pegRNA (e.g., in the case of the 5' extension arm wherein the DNA polymerase simply runs out of template), (b) reaching an impassable RNA secondary structure (e.g., hairpin or stem / loop), or (c) reaching a replication termination signal, e.g., a specific nucleotide sequence that blocks or inhibits the polymerase, or a nucleic acid topological signal, such as, supercoiled DNA or RNA.38FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0122] In some embodiments, the pegRNA includes a ~20 nt protospacer sequence and a gRNA core region, which binds with the napDNAbp. In some embodiments, the pegRNA includes an extended RNA segment at the 5' end, i.e., a 5' extension. In some embodiments, the 5' extension includes a reverse transcription template sequence, a reverse transcription primer binding site, and an optional 5-20 nucleotide linker sequence. In some embodiments, the RT primer binding site hybrizes to the free 3' end that is formed after a nick is formed in the non-target strand of the R- loop, thereby priming reverse transcriptase for DNA polymerization in the 5'-3' direction.
[0123] In some embodiments, the pegRNA includes an extended RNA segment at an intermolecular position within the gRNA core, i.e., an intramolecular extension. In some embodiments, the intramolecular extension includes a reverse transcription template sequence, and a reverse transcription primer binding site. The RT primer binding site hybrizes to the free 3' end that is formed after a nick is formed in the non-target strand of the R-loop, thereby priming reverse transcriptase for DNA polymerization in the 5'-3' direction. In some embodiments, the position of the intermolecular RNA extension is not in the protospacer sequence of the guide RNA. In some embodiments, the position of the intermolecular RNA extension is in the gRNA core. In some embodiments, the position of the intermolecular RNA extension is any with the guide RNA molecule except within the protospacer sequence, or at a position which disrupts the protospacer sequence.
[0124] In some embodiments, the intermolecular RNA extension is inserted downstream from the 3' end of the protospacer sequence. In some embodiments, the intermolecular RNA extension is inserted at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides downstream of the 3' end of the protospacer sequence.
[0125] In some embodiments, the intermolecular RNA extension is inserted into the gRNA, which refers to the portion of the guide RNA corresponding or comprising the tracrRNA, which binds and / or interacts with the Cas9 protein or equivalent thereof (i.e, a different napDNAbp).39FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 Preferably the insertion of the intermolecular RNA extension does not disrupt or minimally disrupts the interaction between the tracrRNA portion and the napDNAbp.
[0126] The length of the RNA extension can be any useful length. In some embodiments, the RNA extension is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.
[0127] The RTT sequence can also be any suitable length. In some embodiments, the RTT sequence can be at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.
[0128] In some embodiments, wherein the reverse transcription primer binding site (PBS) sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.40FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0129] In some embodiments, the optional linker or spacer sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides in length.
[0130] In some embodiments, the RTT sequence encodes a single-stranded DNA molecule which is homologous to the non-target strand (and thus, complementary to the corresponding site of the target strand) but includes one or more nucleotide changes. In some embodiments, the least one nucleotide change may include one or more single-base nucleotide changes, one or more deletions, and one or more insertions.
[0131] In some embodiments, the desired nucleotide change is installed in an editing window that is between about -5 to +5 of the nick site, or between about -10 to +10 of the nick site, or between about -20 to +20 of the nick site, or between about -30 to +30 of the nick site, or between about -40 to +40 of the nick site, or between about -50 to +50 of the nick site, or between about -60 to +60 of the nick site, or between about -70 to +70 of the nick site, or between about -80 to +80 of the nick site, or between about -90 to +90 of the nick site, or between about -100 to +100 of the nick site, or between about -200 to +200 of the nick site.
[0132] In some embodiments, the desired nucleotide change is installed in an editing window that is between about +1 to +2 from the nick site, or about +1 to +3, +1 to +4, +1 to +5, +1 to +6, +1 to +7, +1 to +8, +1 to +9, +1 to +10, +1 to +11, +1 to +12, +1 to +13, +1 to +14, +1 to +15, +1 to +16, +1 to +17, +1 to +18, +1 to +19, +1 to +20, +1 to +21, +1 to +22, +1 to +23, +1 to+24, +1 to +25, +1 to +26, +1 to +27, +1 to +28, +1 to +29, +1 to +30, +1 to +31, +1 to +32, +1 to+33, +1 to +34, +1 to +35, +1 to +36, +1 to +37, +1 to +38, +1 to +39, +1 to +40, +1 to +41, +1 to+42, +1 to +43, +1 to +44, +1 to +45, +1 to +46, +1 to +47, +1 to +48, +1 to +49, +1 to +50, +1 to+51, +1 to +52, +1 to +53, +1 to +54, +1 to +55, +1 to +56, +1 to +57, +1 to +58, +1 to +59, +1 to+60, +1 to +61, +1 to +62, +1 to +63, +1 to +64, +1 to +65, +1 to +66, +1 to +67, +1 to +68, +1 to+69, +1 to +70, +1 to +71, +1 to +72, +1 to +73, +1 to +74, +1 to +75, +1 to +76, +1 to +77, +1 to41FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 +78, +1 to +79, +1 to +80, +1 to +81, +1 to +82, +1 to +83, +1 to +84, +1 to +85, +1 to +86, +1 to +87, +1 to +88, +1 to +89, +1 to +90, +1 to +90, +1 to +91, +1 to +92, +1 to +93, +1 to +94, +1 to +95, +1 to +96, +1 to +97, +1 to +98, +1 to +99, +1 to +100, +1 to +101, +1 to +102, +1 to +103, +1 to +104, +1 to +105, +1 to +106, +1 to +107, +1 to +108, +1 to +109, +1 to +110, +1 to +111, +1 to +112, +1 to +113, +1 to +114, +1 to +115, +1 to +116, +1 to +117, +1 to +118, +1 to +119, +1 to +120, +1 to +121, +1 to +122, +1 to +123, +1 to +124, or +1 to +125 from the nick site.
[0133] In some embodiments, the desired nucleotide change is installed in an editing window that is between about +1 to +2 from the nick site, or about +1 to +5, +1 to +10, +1 to +15, +1 to +20, +1 to +25, +1 to +30, +1 to +35, +1 to +40, +1 to +45, +1 to +50, +1 to +55, +1 to +100,+1 to +105, +1 to +110, +1 to +115, +1 to +120, +1 to +125, +1 to +130, +1 to +135, +1 to +140,+1 to +145, +1 to +150, +1 to +155, +1 to +160, +1 to +165, +1 to +170, +1 to +175, +1 to +180,+1 to +185, +1 to +190, +1 to +195, or +1 to +200, from the nick site.
[0134] In some embodiments, the pegRNAs are modified versions of a guide RNA. Guide RNAs maybe naturally occurring, expressed from an encoding nucleic acid, or synthesized chemically. Methods are well known in the art for obtaining or otherwise synthesizing guide RNAs and for determining the appropriate sequence of the guide RNA, including the protospacer sequence which interacts and hybridizes with the target strand of a genomic target site of interest.
[0135] In some embodiments, a pegRNA comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, a pegRNA comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, a pegRNA comprises a nucleic acid sequence that is complementary to a target sequence and comprises a sequence selected from SEQ ID NOs: 1-326 described in Table 1.Table 1: Exemplary pegRNA sequences of the disclosure42FoleyHoagUS13167537.2Attorney Docket No. JHV-1692543FoleyHoagUS13167537.2Attorney Docket No. JHV-1692544FoleyHoagUS13167537.2Attorney Docket No. JHV-1692545FoleyHoagUS13167537.2Attorney Docket No. JHV-1692546FoleyHoagUS13167537.2Attorney Docket No. JHV-1692547FoleyHoagUS13167537.2Attorney Docket No. JHV-1692548FoleyHoagUS13167537.2Attorney Docket No. JHV-1692549FoleyHoagUS13167537.2Attorney Docket No. JHV-1692550FoleyHoagUS13167537.2Attorney Docket No. JHV-1692551FoleyHoagUS13167537.2Attorney Docket No. JHV-1692552FoleyHoagUS13167537.2Attorney Docket No. JHV-1692553FoleyHoagUS13167537.2Attorney Docket No. JHV-1692554FoleyHoagUS13167537.2Attorney Docket No. JHV-1692555FoleyHoagUS13167537.2Attorney Docket No. JHV-1692556FoleyHoagUS13167537.2Attorney Docket No. JHV-1692557FoleyHoagUS13167537.2Attorney Docket No. JHV-1692558FoleyHoagUS13167537.2Attorney Docket No. JHV-1692559FoleyHoagUS13167537.2Attorney Docket No. JHV-1692560FoleyHoagUS13167537.2Attorney Docket No. JHV-1692561FoleyHoagUS13167537.2Attorney Docket No. JHV-1692562FoleyHoagUS13167537.2Attorney Docket No. JHV-1692563FoleyHoagUS13167537.2Attorney Docket No. JHV-1692564FoleyHoagUS13167537.2Attorney Docket No. JHV-1692565FoleyHoagUS13167537.2Attorney Docket No. JHV-1692566FoleyHoagUS13167537.2Attorney Docket No. JHV-1692567FoleyHoagUS13167537.2Attorney Docket No. JHV-1692568FoleyHoagUS13167537.2Attorney Docket No. JHV-1692569FoleyHoagUS13167537.2Attorney Docket No. JHV-1692570FoleyHoagUS13167537.2Attorney Docket No. JHV-1692571FoleyHoagUS13167537.2Attorney Docket No. JHV-1692572FoleyHoagUS13167537.2Attorney Docket No. JHV-1692573FoleyHoagUS13167537.2Attorney Docket No. JHV-1692574FoleyHoagUS13167537.2Attorney Docket No. JHV-1692575FoleyHoagUS13167537.2Attorney Docket No. JHV-1692576FoleyHoagUS13167537.2Attorney Docket No. JHV-1692577FoleyHoagUS13167537.2Attorney Docket No. JHV-1692578FoleyHoagUS13167537.2Attorney Docket No. JHV-1692579FoleyHoagUS13167537.2Attorney Docket No. JHV-1692580FoleyHoagUS13167537.2Attorney Docket No. JHV-1692581FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0136] In some embodiments, an reverse transcriptase sequence (RTT) comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, an RTT comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, an RTT consists of a sequence selected from SEQ ID NOs: 327-652 described in Table 2.Table 2: Exemplary RTT sequences of the disclosure82FoleyHoagUS13167537.2Attorney Docket No. JHV-1692583FoleyHoagUS13167537.2Attorney Docket No. JHV-1692584FoleyHoagUS13167537.2Attorney Docket No. JHV-1692585FoleyHoagUS13167537.2Attorney Docket No. JHV-1692586FoleyHoagUS13167537.2Attorney Docket No. JHV-1692587FoleyHoagUS13167537.2Attorney Docket No. JHV-1692588FoleyHoagUS13167537.2Attorney Docket No. JHV-1692589FoleyHoagUS13167537.2Attorney Docket No. JHV-1692590FoleyHoagUS13167537.2Attorney Docket No. JHV-1692591FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0137] In some embodiments, a spacer comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, a spacer comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, a spacer sequence consists of a sequence selected from SEQ ID NOs: 653-978 described in Table 3.Table 3: Exemplary spacer sequences of the disclosure92FoleyHoagUS13167537.2Attorney Docket No. JHV-1692593FoleyHoagUS13167537.2Attorney Docket No. JHV-1692594FoleyHoagUS13167537.2Attorney Docket No. JHV-1692595FoleyHoagUS13167537.2Attorney Docket No. JHV-1692596FoleyHoagUS13167537.2Attorney Docket No. JHV-1692597FoleyHoagUS13167537.2Attorney Docket No. JHV-1692598FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0138] In some embodiments, a primer binding sequences (PBS) comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, a PBS comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, a PBS consists of a sequence selected from SEQ ID NOs: 979-1304 described in Table 4.Table 4: Exemplary primer binding sequences of the disclosure99FoleyHoagUS13167537.2Attorney Docket No. JHV-16925100FoleyHoagUS13167537.2Attorney Docket No. JHV-16925101FoleyHoagUS13167537.2Attorney Docket No. JHV-16925102FoleyHoagUS13167537.2Attorney Docket No. JHV-16925103FoleyHoagUS13167537.2Attorney Docket No. JHV-16925104FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0139] In some embodiments, a linker comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, a liner comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, a linker consists of a sequence selected from SEQ ID NOs: 1305-1630 described in Table 5.Table 5: Exemplary linker sequences of the disclosure105FoleyHoagUS13167537.2Attorney Docket No. JHV-16925106FoleyHoagUS13167537.2Attorney Docket No. JHV-16925107FoleyHoagUS13167537.2Attorney Docket No. JHV-16925108FoleyHoagUS13167537.2Attorney Docket No. JHV-16925109FoleyHoagUS13167537.2Attorney Docket No. JHV-16925110FoleyHoagUS13167537.2Attorney Docket No. JHV-16925IllFoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0140] In some embodiments, a scaffold comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, a scaffold comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, a scaffold consists of a sequence selected from SEQ ID NO: 1631 described in Table 6.
[0141] In some embodiments, a 3' motif comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, a 3' motif comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, a 3' motif consists of a sequence selected from SEQ ID NO: 1632 described in Table 6.
[0142] In some embodiments, a terminator comprises a nucleic acid sequence that is complementary to a target sequence described herein. In some embodiments, a terminator comprises a nucleic acid sequence that is at least 90, 95, 99, or 100% complementary to a target sequence described herein. In some embodiments, a terminator consists of a sequence selected from SEQ ID NO: 1633 described in Table 6.Table 6: Exemplary scaffold, 3' motif , and terminator sequences of the disclosure
[0143] In some embodiments, the pegRNA may be improved by introducing improvements to the scaffold or core sequences. Example improvements consists of a sequence selected from SEQ ID NOs: 1665-1667 described in Table 10. In some embodiments, the pegRNA may be improved112FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 by introducing additional RNA motifs at the 5' and 3' termini of the pegRNAs, or even at positions therein between (e.g., in the gRNA core region, or the spacer).
[0144] In some embodiments, the particular design aspects of a guide RNA sequence will depend upon the nucleotide sequence of a genomic target site of interest (i.e., the desired site to be edited) and the type of napDNAbp (e.g., Cas9 protein) present in prime editing systems described herein, among other factors, such as PAM sequence locations, percent G / C content in the target sequence, the degree of microhomology regions, secondary structures, etc.
[0145] In general, a guide sequence is any polynucleotide sequence having sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a napDNAbp (e.g., a Cas9, Cas9 homolog, or Cas9 variant) to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting example of which include the Smith-Waterman algorithm, the Needleman- Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, a guide sequence is about or more than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length.
[0146] In some embodiments, a guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. The ability of a guide sequence to direct sequencespecific binding of a prime editor (PE) to a target sequence may be assessed by any suitable assay. For example, the components of a prime editor (PE), including the guide sequence to be tested, may be provided to a host cell having the corresponding target sequence, such as by transfection with vectors encoding the components of a prime editor (PE) disclosed herein, followed by an assessment of preferential cleavage within the target sequence, such as by Surveyor assay as described herein. Similarly, cleavage of a target polynucleotide sequence may be evaluated in a test tube by providing the target sequence, components of a prime editor (PE), including the guide sequence to be tested and a control guide sequence different from the test guide sequence, and113FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 comparing binding or rate of cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible, and will occur to those skilled in the art.CRISPR
[0147] In some embodiments provided are CRISPR systems for use with the prime editing and base editing systems as described herein.
[0148] Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) together with Cas (CRISPR-associated) genes comprise an adaptive immune system that provides acquired resistance against invading foreign nucleic acids in bacteria and archaea (Barrangou et al. (2007) Science 315:1709-12). CRISPR consists of arrays of short conserved repeat sequences interspaced by unique variable DNA sequences of similar size called spacers, which often originate from phage or plasmid DNA (Barrangou et al. (2007) Science 315:1709-12; Bolotin et al. (2005) Microbiology 151 :2551-61 ; Mojica et al. (2005) J. Mol. Evol. 60:174-82). The CRISPR-Cas system functions by acquiring short pieces of foreign DNA (spacers) which are inserted into the CRISPR region and provide immunity against subsequent exposures to phages and plasmids that carry matching sequences (Barrangou et al. (2007) Science 315:1709-12; Brouns et al. (2008) Science 321:960-64). It is this CRISPR-Cas interference / immunity that enables crRNA-mediated silencing of foreign nucleic acids (Horvath & Barrangou (2010) Science 327:167-70; Deveau et al. (2010) Annu. Rev. Microbiol. 64:475-93; Marraffini & Sontheimer (2010) Nat. Rev. Genet. 11:181-90; Bhaya et al. (2011) Annu. Rev. Genet. 45:273-97; Wiedenheft et al. (2012) Nature 482:331-338).
[0149] In some embodiments, the term “Cas protein” refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequences that differs from a naturally occurring Cas protein, or any fragment of a Cas protein that nevertheless retains all or a significant amount of the requisite basic functions needed for the disclosed methods, i.e., (i) possession of nucleic-acid programmable binding of the Cas protein to a target DNA, and (ii) ability to nick the target DNA sequence on one strand. The Cas proteins contemplated herein embrace CRISPR Cas 9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease inactive Cas9 (dCas9)) homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and may include a Cas9 equivalent from any type of CRISPR system (e.g., type II, V, VI), including Cpfl (a type-V CRISPR-Cas systems), C2cl (a type V CRISPR-Cas system), C2c2 (a type VI CRISPR-Cas system) and C2c3 (a type V CRISPR-Cas114FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 system). Further Cas-equivalents are described in Makarova et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299), the contents of which are incorporated herein by reference.
[0150] The terms “Cas9” or “Cas9 nuclease” or “Cas9 moiety” or “Cas9 domain” embrace any naturally occurring Cas9 from any organism, any naturally-occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of a Cas9, naturally-occurring or engineered. The term Cas9 is not meant to be particularly limiting and may be referred to as a “Cas9 or equivalent.” Exemplary Cas9 proteins are further described herein and / or are described in the art and are incorporated herein by reference. The present disclosure is unlimited with regard to the particular Cas9 may be employed in prime editing and / or base editing technologies described herein.
[0151] In some embodiments, the presently disclosed methods utilize a composition comprising a non-naturally occurring CRISPR system comprising one or more vectors comprising: a) a mammalian promoter operably linked to at least one nucleotide sequence encoding a CRISPR system prime editing guide RNA (pegRNA), wherein the pegRNA hybridizes with a target sequence of a DNA molecule in a cell, and wherein the DNA molecule encodes one or more gene products expressed in the cell; and b) a regulatory element operable in a cell operably linked to a nucleotide sequence encoding a Cas protein (e.g., Cas9), wherein components (a) and (b) are located on the same or different vectors of the system, wherein the pegRNA targets and hybridizes with the target sequence and the Cas protein cleaves the DNA molecule to alter expression of the one or more gene products.
[0152] In some embodiments, any suitable napDNAbp may be used in the prime editors or base editors described herein. In some embodiments, the napDNAbp may be any Class 2 CRISPR - Cas system, including any type II, type V, or type VI CRISPR-Cas enzyme. Given the rapid development of CRISPR-Cas as a tool for genome editing, there have been constant developments in the nomenclature used to describe and / or identify CRISPR-Cas enzymes, such as Cas9 and Cas9 orthologs. This application references CRISPR-Cas enzymes with nomenclature that may be old and / or new. The skilled person will be able to identify the specific CRISPR-Cas enzyme being referenced in this Application based on the nomenclature that is used, whether it is old (i.e., “legacy”) or new nomenclature. CRISPR-Cas nomenclature is extensively discussed in Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR115FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 Journal, Vol. 1. No. 5, 2018, the entire contents of which are incorporated herein by reference. In some embodiments, the particular CRISPR-Cas nomenclature used in any given instance in this Application is not limiting in any way and the skilled person will be able to identify which CRISPR-Cas enzyme is being referenced.Reverse Transcriptase (RT)
[0153] In some embodiments, the reverse transcriptase (RT) of the prime editors may be an “error-prone” reverse transcriptase variant. Error-prone reverse transcriptases that are known and / or available in the art may be used. It will be appreciated that reverse transcriptases naturally do not have any proofreading function; thus the error rate of reverse transcriptase is generally higher than DNA polymerases comprising a proofreading activity. The error-rate of any particular reverse transcriptase is a property of the enzyme's “fidelity,” which represents the accuracy of template- directed polymerization of DNA against its RNA template. An RT with high fidelity has a low-error rate. Conversely, an RT with low fidelity has a high-error rate. The fidelity of M-MLV-based reverse transcriptases are reported to have an error rate in the range of one error in 15,000 to 27,000 nucleotides synthesized. See Boutabout et al., “DNA synthesis fidelity by the reverse transcriptase of the yeast retrotransposon Tyl, ” Nucleic Acids Res, 2001, 29: 2217-2222, which is incorporated by reference. Thus, for purposes of this application, those reverse transcriptases considered to be “error-prone” or which are considered to have an “error-prone fidelity” are those having an error rate that is less than one error in 15,000 nucleotides synthesized.
[0154] In some embodiments, error-prone reverse transcriptase also may be created through mutagenesis of a starting RT enzyme (e.g., a wild type M-MLV RT). The method of mutagenesis is not limited and may include directed evolution processes, such as phage-assisted continuous evolution (PACE) or phage-assisted noncontinuous evolution (PANCE). The term “phage-assisted continuous evolution (PACE),” as used herein, refers to continuous evolution that employs phage as viral vectors. The general concept of PACE technology has been described, for example, in International PCT Application, PCT / US2009 / 056194, filed Sep. 8, 2009, published as WO 2010 / 028347 on Mar. 11, 2010; International PCT Application, PCT / US2011 / 066747, filed Dec. 22, 2011, published as WO 2012 / 088381 on Jun. 28, 2012; U.S. Pat. No. 9,023,594, issued May 5, 2015, International PCT Application, PCT / US2015 / 012022, filed Jan. 20, 2015, published as WO 2015 / 134121 on Sep. 11, 2015, and International PCT Application, PCT / US2016 / 027795, filed Apr.116FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 15, 2016, published as WO 2016 / 168631 on Oct. 20, 2016, the entire contents of each of which are incorporated herein by reference.
[0155] In some embodiments, error-prone reverse transcriptases may also be obtained by phage-assisted non-continuous evolution (PANCE),” which as used herein, refers to non-continuous evolution that employs phage as viral vectors. PANCE is a simplified technique for rapid in vivo directed evolution using serial flask transfers of evolving ‘selection phage’ (SP), which contain a gene of interest to be evolved, across fresh E. coli host cells, thereby allowing genes inside the host E. coli to be held constant while genes contained in the SP continuously evolve. Serial flask transfers have long served as a widely-accessible approach for laboratory evolution of microbes, and, more recently, analogous approaches have been developed for bacteriophage evolution. The PANCE system features lower stringency than the PACE system.
[0156] In some embodiments, other error-prone reverse transcriptases have been described in the literature, each of which are contemplated for use in the herein methods and compositions. For example, error-prone reverse transcriptases have been described in Bebenek et al., “Error-prone Polymerization by HIV-1 Reverse Transcriptase,” J Biol Chem, 1993, Vol. 268: 10324-10334 and Sebastian-Martin et al., “Transcriptional inaccuracy threshold attenuates differences in RNA- dependent DNA synthesis fidelity between retroviral reverse transcriptases,” Scientific Reports, 2018, Vol. 8: 627, each of which are incorporated by reference. Still further, reverse transcriptases, including error-prone reverse transcriptases can be obtained from a commercial supplier, including ProtoScript® (II) Reverse Transcriptase, AMV Reverse Transcriptase, WarmStart® Reverse Transcriptase, and M-MuLV Reverse Transcriptase, all from NEW ENGLAND BIOLABS®, or AMV Reverse Transcriptase XL, SMARTScribe Reverse Transcriptase, GPR ultra-pure MMLV Reverse Transcriptase, all from TAKARA BIO USA, INC. (formerly CLONTECH).
[0157] In some embodiments, the herein disclosure also contemplates reverse transcriptases having mutations in RNaseH domain. One of the intrinsic properties of reverse transcriptases is the RNase H activity, which cleaves the RNA template of the RNA:cDNA hybrid concurrently with polymerization. The RNase H activity can be undesirable for synthesis of long cDNAs because the RNA template may be degraded before completion of full-length reverse transcription. The RNase H activity may also lower reverse transcription efficiency, presumably due to its competition with117FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 the polymerase activity of the enzyme. In some embodiments, the present disclosure contemplates any reverse transcriptase variants that comprise a modified RNaseH activity.
[0158] In some embodiments, the herein disclosure also contemplates reverse transcriptases having mutations in the RNA-dependent DNA polymerase domain. As mentioned above, one of the intrinsic properties of reverse transcriptases is the RNA-dependent DNA polymerase activity, which incorporates the nucleobases into the nascent cDNA strand as coded by the template RNA strand of the RNA:cDNA hybrid. The RNA-dependent DNA polymerase activity can be increased or decreased (i.e., in terms of its rate of incorporation) to either increase or decrease the processivity of the enzyme. In some embodiments, the present disclosure contemplates any reverse transcriptase variants that comprise a modified RNA-dependent DNA polymerase activity such that the processivity of the enzyme of either increased or decreased relative to an unmodified version.
[0159] In some embodiments, the herein disclosure also contemplates reverse transcriptase variants that have altered thermostability characteristics. The ability of a reverse transcriptase to withstand high temperatures is an important aspect of cDNA synthesis. Elevated reaction temperatures help denature RNA with strong secondary structures and / or high GC content, allowing reverse transcriptases to read through the sequence. As a result, reverse transcription at higher temperatures enables full-length cDNA synthesis and higher yields, which can lead to an improved generation of the 3' flap ssDNA as a result of the prime editing process. For example, Wild type M- MLV reverse transcriptase typically has an optimal temperature in the range of 37-48° C.; however, mutations may be introduced that allow for the reverse transcription activity at higher temperatures of over 48° C„ including 49° C„ 50° C„ 51° C„ 52° C„ 53° C„ 54° C„ 55° C„ 56° C„ 57° C„ 58° C„ 59° C„ 60° C„ 61° C„ 62° C„ 63° C„ 64° C„ 65° C„ 66° C„ and higher.
[0160] In some embodiments, the variant reverse transcriptases contemplated herein, including error-prone RTs, thermostable RTs, increase-processivity RTs, can be engineered by various routine strategies, including mutagenesis or evolutionary processes. In some cases, the variants can be produced by introducing a single mutation. In other cases, the variants may require more than one mutation. For those mutants comprising more than one mutation, the effect of a given mutation may be evaluated by introduction of the identified mutation to the wild-type gene by site-directed mutagenesis in isolation from the other mutations borne by the particular mutant. Screening assays of the single mutant thus produced will then allow the determination of the effect of that mutation alone.118FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0161] In some embodiments, variant RT enzymes used herein may also include other “RT variants” having at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference RT protein, including any wild type RT, or mutant RT, or fragment RT, or other variant of RT disclosed or contemplated herein or known in the art.
[0162] In some embodiments, an RT variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500 or more amino acid changes compared to a reference RT. In some embodiments, the RT variant comprises a fragment of a reference RT, such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of the reference RT. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild type RT (M-MLV reverse transcriptase) or to any of the reverse transcriptase sequences previously described in US 11,795,452, the contents of which are hereby incorporated by reference in their entirety.
[0163] In some embodiments, the disclosure also may utilize RT fragments which retain their functionality and which are fragments of any herein disclosed RT proteins. In some embodiments, the RT fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, or up to 600 or more amino acids in length.
[0164] In some embodiments, the disclosure also may utilize RT variants which are truncated at the N-terminus or the C-terminus, or both, by a certain number of amino acids which results in a truncated variant which still retains sufficient polymerase function. In some embodiments, the RT truncated variant has a truncation of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at119FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the N-terminal end of the protein. In other embodiments, the RT truncated variant has a truncation of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the C- terminal end of the protein. In still other embodiments, the RT truncated variant has a trunction at the N-terminal and the C-terminal end which are the same or different lengths.
[0165] In some embodiments, the prime editors disclosed herein may comprise one of the RT variants described herein, or a RT variant thereof having at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference Cas9 variants.
[0166] In some embodiments, the present methods and compositions may utilize a DNA polymerase that has been evolved into a reverse transcriptase, as described in Effefson et al., “Synthetic evolutionary origin of a proofreading reverse transcriptase,” Science, Jun. 24, 2016, Vol. 352: 1590-1593, the contents of which are incorporated herein by reference.
[0167] In certain other embodiments, the reverse transcriptase is provided as a component of a fusion protein also comprising a napDNAbp. In some embodiments, the reverse transcriptase is fused to a napDNAbp as a fusion protein.
[0168] In some embodiments, the prime editors described herein (with RT provided as either a fusion partner or in trans) can include a variant RT comprising one or more of the following mutations: P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, or D653N in the wild type M-MLV RT or at a corresponding amino acid position in another wild type RT polypeptide sequence.Base editing120FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0169] “Base editing” refers to genome editing technology that involves the conversion of a specific nucleic acid base into another at a targeted genomic locus. In some embodiments, this can be achieved without requiring double-stranded DNA breaks (DSB), or single stranded breaks (i.e., nicking). To date, other genome editing techniques, including CRISPR -based systems, begin with the introduction of a DSB at a locus of interest. Subsequently, cellular DNA repair enzymes mend the break, commonly resulting in random insertions or deletions (indels) of bases at the site of the DSB. However, when the introduction or correction of a point mutation at a target locus is desired rather than stochastic disruption of the entire gene, these genome editing techniques are unsuitable, as correction rates are low (e.g. typically 0.1% to 5%), with the major genome editing products being indels.
[0170] In some embodiments, the base editor constructs described herein may comprise the “canonical SpCas9” nuclease from S. pyogenes, which has been widely used as a tool for genome engineering. This Cas9 protein is a large, multi-domain protein containing two distinct nuclease domains. Point mutations can be introduced into Cas9 to abolish one or both nuclease activities, resulting in a nickase Cas9 (nCas9) or dead Cas9 (dCas9), respectively, that still retains its ability to bind DNA in a sgRNA-programmed manner. In principle, when fused to another protein or domain, Cas9 or variant thereof (e.g., nCas9) can target that protein to virtually any DNA sequence simply by co-expression with an appropriate sgRNA.
[0171] In some embodiments, Cas9 sequences suitable for base editing systems or equivalents thereof have been described in US20230159913A1, the contents of which are hereby incorporated by reference in their entirety.
[0172] The base editors described herein may include any of the above Cas9 ortholog sequences, or any variants thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0173] The napDNAbp may include any suitable homologs and / or orthologs or naturally occurring enzymes, such as, Cas9. Cas9 homologs and / or orthologs have been described in various species, including, but not limited to, 5. pyogenes and S. thermophilus. Preferably, the Cas moiety is configured (e.g., mutagenized, recombinantly engineered, or otherwise obtained from nature) as a nickase, i.e., capable of cleaving only a single strand of the target. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed121FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, a Cas9 nuclease has an inactive (e.g., an inactivated) DNA cleavage domain, that is, the Cas9 is a nickase. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein as provided by any one of the Cas9 orthologs disclosed in US20230159913A1, the contents of which are hereby incorporated by reference in their entirety.
[0174] In some embodiments, the base editors described herein may also comprise Casl2a / Cpfl (dCpf 1) variants that may be used as a guide nucleotide sequence- programmable DNA-binding protein domain. The Casl2a / Cpfl protein has a RuvC-like endonuclease domain that is similar to the RuvC domain of Cas9 but does not have a HNH endonuclease domain, and the N- terminal of Cpfl does not have the alpha-helical recognition lobe of Cas9. It was shown in Zetsche et al., Cell, 163, 759-771, 2015 (which is incorporated herein by reference) that, the RuvC-like domain of Cpfl is responsible for cleaving both DNA strands and inactivation of the RuvC-like domain inactivates Cpfl nuclease activity.
[0175] In some embodiments, the napDNAbp is a nucleic acid programmable DNA binding protein that does not require a canonical (NGG) PAM sequence. In some embodiments, the napDNAbp is an argonaute protein. One example of such a nucleic acid programmable DNA binding protein is an Argonaute protein from Natronobacterium gregoryi (NgAgo). NgAgo is a ssDNA-guided endonuclease. NgAgo binds 5' phosphorylated ssDNA of ~24 nucleotides (gDNA) to guide it to its target site and will make DNA double-strand breaks at the gDNA site. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer-adjacent motif (PAM). Using a nuclease inactive NgAgo (dNgAgo) can greatly expand the bases that may be targeted. The characterization and use of NgAgo have been described in Gao et al., Nat Biotechnol., 2016 Jul; 34(7):768-73. PubMed PMID: 27136078; Swarts et al., Nature. 507(7491) (2014):258-61; and Swarts et al., Nucleic Acids Res. 43(10) (2015):5120-9, each of which is incorporated herein by reference.122FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0176] In some embodiments, the disclosure provides napDNAbp domains that comprise SpCas9 variants that recognize and work best with NRRH, NRCH, and NRTH PAMs. See International Application No. PCT / US2019 / 47996, which published as International Publication No. WO 2020 / 041751 on Feb. 27, 2020, incorporated by reference herein. In some embodiments, the disclosed base editors comprise a napDNAbp domain selected from SpCas9-NRRH, SpCas9- NRTH, and SpCas9-NRCH.
[0177] In some embodiments, the disclosure provides base editors that comprise one or more cytidine deaminase domains. In some embodiments, any of the disclosed base editors are capable of deaminating cytidine in a nucleic acid sequence (e.g., genomic DNA). As one example, any of the base editors described herein may be base editors, (e.g., cytidine base editors).
[0178] In some embodiments, the cytidine deaminase is an apolipoprotein B mRNA- editing complex (APOBEC) family deaminase. In some embodiments, the cytidine deaminase is an APOB EC 1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, an APOBEC3H deaminase, or an APOBEC4 deaminase. In some embodiments, the cytidine deaminase is an activation-induced deaminase (AID). In some embodiments, the deaminase is a Lamprey CDA1 (pmCDAl) deaminase. In some embodiments, the cytidine deaminase is from a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase is from a human. In some embodiments the deaminase is from a rat. In some embodiments, the cytidine deaminase is a human APOBEC1 deaminase. In some embodiments, the cytidine deaminase is pmCDAl. In some embodiments, the deaminase is human APOBEC3G. In some embodiments, the deaminase is a human APOBEC3G variant. In some embodiments, the deaminase is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the APOBEC amino acid sequences set forth in US20230159913A1, the contents of which are hereby incorporated by reference in their entirety.
[0179] Additional exemplary cytidine deaminases from the APOBEC family include Pongo pygmaeus APOB EC 1 (PpAPOBECl), Rhinopithecus roxellana APOBEC3F (RrA3F), Alligator mississippiensis (AmAPOBECl), and Sus scrofa APOBEC3B (SsAPOBEC3B), and variants thereof. For example, any one of the cytidine deaminases PpAPOBECl, PpAPOBECl H122A, PpAPOBECl R33A, RrA3F, RrA3F F130L, AmAPOBECl, SsAPOBEC3B, and SsAPOBEC3B123FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 R54Q, may be used with the editing methods provided herein. See Gaudelli et al. Nat.Commun. (2020) 11:2052 and International Publication No. WO 2021 / 041885, published Mar. 4, 2021, each of which is incorporated by reference herein.
[0180] In some embodiments, any of the base editors described herein comprise a deaminase domain (e.g., a cytidine deaminase domain) that has reduced catalytic deaminase activity. In some embodiments, any of the base editors described herein comprise a deaminase domain (e.g., a cytidine deaminase domain) that has a reduced catalytic deaminase activity as compared to an appropriate control. For example, the appropriate control may be the deaminase activity of the deaminase prior to introducing one or more mutations into the deaminase. In other embodiments, the appropriate control may be a wild-type deaminase. In some embodiments, the appropriate control is a wild-type apolipoprotein B mRNA-editing complex (APOB EC) family deaminase. In some embodiments, the appropriate control is an APOBEC1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, or an APOBEC3H deaminase. In some embodiments, the appropriate control is an activation induced deaminase (AID). In some embodiments, the appropriate control is a cytidine deaminase 1 from Petromyzon marinus (pmCDAl).
[0181] In some embodiments, the deaminase domain may be a deaminase domain that has at least 1%, at least 5%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% less catalytic deaminase activity as compared to an appropriate control. In some embodiments, a base editor recognizes canonical PAMs and therefore can correct the pathogenic G to A or C to T mutations with canonical PAMs, e.g., NGG, respectively, in the flanking sequences.
[0182] In some embodiments, the disclosure provides methods for editing a nucleic acid. In some embodiments, the method is a method for editing a nucleobase of a nucleic acid (e.g., a base pair of a double-stranded DNA sequence). In some embodiments, the method comprises the steps of: a) contacting a target region of a nucleic acid (e.g., a double- stranded DNA sequence) with a complex comprising a base editor (e.g., a Cas9 domain fused to an adenosine deaminase) and a guide nucleic acid (e.g., gRNA), wherein the target region comprises a targeted nucleobase pair, b) inducing strand separation of said target region, c) converting a first nucleobase of said target nucleobase pair in a single strand of the target region to a second nucleobase, and d) cutting no124FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 more than one strand of said target region, where a third nucleobase complementary to the first nucleobase base is replaced by a fourth nucleobase complementary to the second nucleobase. In some embodiments, the method results in less than 20% indel formation in the nucleic acid. It should be appreciated that in some embodiments, step b is omitted. In some embodiments, the first nucleobase is an adenine. In some embodiments, the second nucleobase is a deaminated adenine, or inosine. In some embodiments, the third nucleobase is a thymine. In some embodiments, the fourth nucleobase is a cytosine. In some embodiments, the method results in less than 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, 0.2%, or less than 0.1% indel formation. In some embodiments, the method further comprises replacing the second nucleobase with a fifth nucleobase that is complementary to the fourth nucleobase, thereby generating an intended edited base pair (e.g., A:T to G:C). In some embodiments, the fifth nucleobase is a guanine. In some embodiments, at least 5% of the intended base pairs are edited.
[0183] In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the intended base pairs are edited. In some embodiments, the ratio of intended products to unintended products in the target nucleotide is at least 2:1, 5:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or 200:1, or more. In some embodiments, the ratio of intended point mutation to indel formation is greater than 1:1, 10:1, 50:1, 100:1, 500:1, or 1000:1, or more. In some embodiments, the cut single strand (nicked strand) is hybridized to the guide nucleic acid. In some embodiments, the cut single strand is opposite to the strand comprising the first nucleobase. In some embodiments, the base editor comprises a Cas9 domain. In some embodiments, the first base is adenine, and the second base is not a G, C, A, or T. In some embodiments, the second base is inosine. In some embodiments, the first base is adenine. In some embodiments, the second base is not a G, C, A, or T. In some embodiments, the second base is inosine. In some embodiments, the base editor inhibits base excision repair of the edited strand. In some embodiments, the base editor protects or binds the non-edited strand. In some embodiments, the base editor comprises UGI activity. In some embodiments, the base editor comprises a catalytically inactive inosine-specific nuclease. In some embodiments, the base editor comprises nickase activity. In some embodiments, the intended edited base pair is upstream of a PAM site. In some embodiments, the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream of the PAM site. In some embodiments, the intended edited base pair is downstream of a PAM site. In some embodiments, the intended edited base pair125FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream stream of the PAM site. In some embodiments, the method does not require a canonical (e.g., NGG) PAM site. In some embodiments, the base editor comprises a linker. In some embodiments, the linker is 1-25 amino acids in length. In some embodiments, the linker is 5-20 amino acids in length. In some embodiments, linker is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids in length. In some embodiments, the target region comprises a target window, wherein the target window comprises the target nucleobase pair. In some embodiments, the target window comprises 1-10 nucleotides. In some embodiments, the target window is 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, or 1 nucleotides in length. In some embodiments, the target window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. In some embodiments, the intended edited base pair is within the target window. In some embodiments, the target window comprises the intended edited base pair. In some embodiments, the method is performed using any of the base editors described herein. In some embodiments, a target window is a deamination window.
[0184] In some embodiments, the disclosure provides methods for editing a nucleotide. In some embodiments, the disclosure provides a method for editing a nucleobase pair of a doublestranded DNA sequence. In some embodiments, the method comprises a) contacting a target region of the double-stranded DNA sequence with a complex comprising a base editor and a guide nucleic acid (e.g., gRNA), where the target region comprises a target nucleobase pair, b) inducing strand separation of said target region, c) converting a first nucleobase of said target nucleobase pair in a single strand of the target region to a second nucleobase, d) cutting no more than one strand of said target region, wherein a third nucleobase complementary to the first nucleobase base is replaced by a fourth nucleobase complementary to the second nucleobase, and the second nucleobase is replaced with a fifth nucleobase that is complementary to the fourth nucleobase, thereby generating an intended edited base pair, wherein the efficiency of generating the intended edited base pair is at least 5%. It should be appreciated that in some embodiments, step b is omitted. In some embodiments, at least 5% of the intended base pairs are edited. In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the intended base pairs are edited. In some embodiments, the method causes less than 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, 0.2%, or less than 0.1% indel formation.
[0185] In some embodiments, the ratio of intended product to unintended products at the target nucleotide is at least 2:1, 5:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or126FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 200:1, or more. In some embodiments, the ratio of intended point mutation to indel formation is greater than 1:1, 10:1, 50:1, 100:1, 500:1, or 1000:1, or more. In some embodiments, the cut single strand is hybridized to the guide nucleic acid. In some embodiments, the cut single strand is opposite to the strand comprising the first nucleobase. In some embodiments, the first base is adenine. In some embodiments, the second nucleobase is not G, C, A, or T. In some embodiments, the second base is inosine. In some embodiments, the base editor inhibits base excision repair of the edited strand. In some embodiments, the base editor protects (e.g., form base excision repair) or binds the non-edited strand. In some embodiments, the base editor comprises UGI activity. In some embodiments, the base editor comprises a catalytically inactive inosine-specific nuclease. In some embodiments, the base editor comprises nickase activity.
[0186] In some embodiments, the intended edited base pair is upstream of a PAM site. In some embodiments, the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream of the PAM site. In some embodiments, the intended edited base pair is downstream of a PAM site. In some embodiments, the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream stream of the PAM site. In some embodiments, the method does not require a canonical (e.g., NGG) PAM site. In some embodiments, the base editor comprises a linker.
[0187] In some embodiments, the linker is 1-25 amino acids in length. In some embodiments, the linker is 5-20 amino acids in length. In some embodiments, the linker is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids in length. In some embodiments, the target region comprises a target window, wherein the target window comprises the target nucleobase pair. In some embodiments, the target window comprises 1-10 nucleotides. In some embodiments, the target window is 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, or 1 nucleotides in length. In some embodiments, the target window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. In some embodiments, the intended edited base pair occurs within the target window. In some embodiments, the target window comprises the intended edited base pair. In some embodiments, the base editor is any one of the base editors described herein.
[0188] In some embodiments, this disclosure provide methods of using the base editors, or complexes comprising a guide nucleic acid (e.g., gRNA) and a base editor described herein. In some embodiments, this disclosure provide methods comprising contacting a DNA, or RNA molecule with any of the base editors described herein, and with at least one guide nucleic acid127FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 (e.g., guide RNA), wherein the guide nucleic acid, (e.g., guide RNA) is comprises a sequence (e.g., a guide sequence that binds to a DNA target sequence) of at least 10 (e.g., at least 10, 15, 20, 25, or 30) contiguous nucleotides that is 100% complementary to a target sequence (e.g., any of the target GNE sequences provided herein). In some embodiments, the 3' end of the target sequence is immediately adjacent to a canonical PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is not immediately adjacent to a canonical PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is immediately adjacent to an AGC, GAG, TTT, GTG, or CAA sequence.
[0189] In some embodiments, the disclosure provide methods of using base editors (e.g., any of the base editors described herein) and gRNAs to change a residue in a GNE gene. In some embodiments, the disclosure provides methods of using base editors (e.g., any of the base editors described herein) and gRNAs to generate a T to C mutation in a GNE gene, thereby resulting in a non-functional GNE protein. In some embodiments, the disclosure provides method for deaminating an cytosine nucleobase (“C”) in a GNE gene, the method comprising contacting the GNE gene with a base editor and a guide RNA bound to the base editor, where the guide RNA comprises a guide sequence that is complementary to a target nucleic acid sequence in the GNE gene.
[0190] In some embodiments, the present disclosure provides uses of any one of the base editors described herein and a guide RNA targeting this base editor to a target in the GNE gene in the manufacture of a medicament. In some embodiments, uses of any one of the base editors and guide RNAs described herein are provided in the manufacture of a kit for base editing, wherein the base editing comprises contacting the nucleic acid molecule with the base editor and guide RNA under conditions suitable for the substitution of the cytosine (C) of a C:T nucleobase pair in the target with a thymine (T). In some embodiments, the step of contacting induces separation of the double-stranded DNA at a target region. In some embodiments, the step of contacting thereby comprises the nicking of one strand of the double-stranded DNA, wherein the one strand comprises an unmutated strand that comprises the G of the target C:G nucleobase pair, in cytidine base editing.Cells
[0191] The present disclosure is further directed, in part, to cells comprising an prime editing system or base editing system as described herein. Any cell, e.g., cell line, e.g., a cell line suitable for expression of a pegRNA or sgRNA, known to one of skill in the art can be used.128FoleyHoagUS13167537.2Attorney Docket No. JHV-16925
[0192] In some embodiments, a cell or cell line for expression of a prime editor and / or pegRNA as described herein in a mammalian cell or cell line. In some embodiments, a cell or cell line for expression of a base editor and / or sgRNA as described herein in a mammalian cell or cell line.
[0193] Any suitable mammalian cell line known in the art can be engineered or screened in the context of the present disclosure. Mammalian cells for expression of viral vectors can include any mammalian cell type known in the art. Representative mammalian cells include, but are not limited to, human embryonic kidney (HEK) cells (e.g., HEK 293 cells, HEK 293T cells, Expi293 cells), Chinese hamster ovary (CHO) cells, HeLa cells (e.g., HeLa S3 cells), PER.C6 cells, HKB1 1 cells, CAP cells, Baby Hamster Kidney fibroblasts (BHK cells) (e.g., BHK-21 cells), mouse myeloma cells (e.g., Sp2 / 0 cells and NSO cells), green African monkey kidney cells (e.g., COS cells and Vero cells), A549 cells, rhesus fetal lung cells (e.g., FRhL-2 cells), and any derivatives thereof. In some embodiments, mammalian cells of the present disclosure are highly transfectable.
[0194] In some embodiments, a mammalian cell line of the present disclosure is suitable for manufacturing of biologies. In some embodiments, a mammalian cell line is suitable for use in industrial-scale manufacturing of a biologic product. In some embodiments, a mammalian cell line is suitable for use in a method of manufacture that conforms with local regulatory standards (e.g., FDA and / or EMA regulatory standards). In some embodiments, a mammalian cell line is suitable for manufacturing of biologies (e.g., viral vectors) using current good manufacturing practices (cGMP). In some embodiments, a mammalian cell line is suitable for manufacturing of biologies using good manufacturing practices (GMP). In some embodiments, a mammalian cell line is suitable for manufacturing of biologies using non-good manufacturing practices (non-GMP).Compositions
[0195] In some embodiments, provided herein are compositions comprising prime editing systems of the present disclosure and components thereof (e.g., pegRNAs). In some embodiments, a provided composition is a pharmaceutical composition comprising a pegRNA or prime editor of the present disclosure and a pharmaceutically acceptable carrier.
[0196] In some embodiments, provided herein are compositions comprising base editing systems of the present disclosure and components thereof (e.g., guide RNAs). In some129FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 embodiments, a provided composition is a pharmaceutical composition comprising a guide RNA or base editor of the present disclosure and a pharmaceutically acceptable carrier.
[0197] A “pharmaceutically-acceptable carrier” as used herein means a pharmaceutically- acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, or solvent encapsulating material, involved in carrying or transporting the subject compound from one organ, or portion of the body, to another organ, or portion of the body. Each carrier must be “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient. Some examples of materials which can serve as pharmaceutically- acceptable carriers include: (1) sugars, such as lactose, glucose and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose, and its derivatives, such as sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) talc; (8) excipients, such as cocoa butter and suppository waxes; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil; (10) glycols, such as propylene glycol; (11) polyols, such as glycerin, sorbitol, mannitol and polyethylene glycol; (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffering agents, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer's solution; (19) ethyl alcohol; (20) pH buffered solutions; (21) polyesters, polycarbonates and / or poly anhydrides; and (22) other non-toxic compatible substances employed in pharmaceutical formulations.
[0198] The pharmaceutical formulations of the present disclosure may include, as optional ingredients, pharmaceutically acceptable carriers, diluents, solubilizing or emulsifying agents, and salts of the type that are well-known in the art. Specific non-limiting examples of the carriers and / or diluents that are useful in the pharmaceutical formulations of the present disclsoure include water and physiologically acceptable buffered saline solutions, such as phosphate buffered saline solutions pH 7.0-8.0.Kits
[0199] In some embodiments, provided is a pack or kit comprising one or more containers filled with at least one pegRNA or prime editor as described in the description, examples, and / or figures herein. In some embodiments, provided herein are compositions comprising base editing systems of the present disclosure and components thereof (e.g., guideRNAs). In some130FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 embodiments, a provided composition is a pharmaceutical composition comprising a guide RNA or basee editor of the present disclosure and a pharmaceutically acceptable carrier.
[0200] Kits may be used in any applicable method (e.g., a research method). Optionally associated with such container(s) can be a notice in the form prescribed by a governmental agency regulating the manufacture, use or sale of pharmaceuticals or biological products, which notice reflects (a) approval by the agency of manufacture, use or sale for human administration, (b) directions for use, (c) a contract that governs the transfer of materials and / or biological products between two or more entities and combinations thereof.MethodsMethods of gene editing
[0201] In some embodiments, the present disclosure provides methods of genetically engineering one or more gene products in a cell, wherein the cell comprises a DNA molecule encoding the one or more gene products, the method comprising introducing into the cell a non- naturally occurring CRISPR-Cas system comprising one or more vectors comprising at least one nucleotide sequence encoding a CRISPR-Cas system prime editing guide RNA (pegRNA), wherein the pegRNA hybridizes with a target sequence of the DNA molecule wherein the pegRNA and prime editor are located on the same or different vectors of the system, wherein the pegRNA targets and hybridizes with the target sequence and the Cas protein cleaves the DNA molecule.
[0202] In some embodiments, the Cas protein (e.g., Cas9) is codon optimized for expression in the cell. In some embodiments, the Cas protein is codon optimized for expression in the eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian or human cell. In some embodiments, the editing efficiency of the CRISPR-Cas system prime editing system is increased. In some embodiments, the presently disclosed subject matter provides methods comprising delivering a pegRNA, prime editor, guide RNA and / or base editor as described herein to a host cell. In some embodiments, the presently disclosed subject matter further provides cells produced by such methods, and organisms (such as animals, plants, or fungi) comprising or produced from such cells.131FoleyHoagUS13167537.2Attorney Docket No. JHV-16925Methods of delivery
[0203] In some embodiments, provided herein are methods for delivering the CRISPR / Cas- based editing system for providing genetic constructs and / or proteins of the CRISPR / Cas-based editing system. The delivery of the CRISPR / Cas-based editing system may be the transfection or electroporation of the CRISPR / Cas-based editing system as one or more nucleic acid molecules that is expressed in the cell and delivered to the surface of the cell. The CRISPR / Cas-based editing system protein may be delivered to the cell. The nucleic acid molecules may be electroporated using BioRad Gene Pulser Xcell or Amaxa Nucleofector lib devices or other electroporation device.Several different buffers may be used, including BioRad electroporation solution, Sigma phosphate- buffered saline product # D8537 (PBS), Invitrogen OptiMEM I (OM), or Amaxa Nucleofector solution V (N.V.). Transfections may include a transfection reagent, such as Lipofectamine 2000.
[0204] The vector encoding a CRISPR / Cas-based editing system protein may be delivered to the modified target cell in a tissue or subject by DNA injection (also referred to as DNA vaccination) with and without in vivo electroporation, liposome mediated, nanoparticle facilitated, and / or recombinant vectors. The recombinant vector may be delivered by any viral mode. The viral mode may be recombinant lentivirus, recombinant adenovirus, and / or recombinant adeno-associated virus.Methods of administration
[0205] In some embodiments, the present disclosure provides methods for administering a CRISPR / Cas-based genome editing system to a subject in need thereof. The genome editing system may comprise one or more vectors or compositions encoding or comprising a prime editor or base editor protein, a pegRNA or sgRNA, and optionally one or more nicking guide RNAs, as described herein. In some embodiments, the composition is formulated for systemic or local delivery in a pharmaceutically acceptable carrier.
[0206] In some embodiments, administration may be by intravenous, intraperitoneal, subcutaneous, intramuscular, or intra-arterial injection. In some embodiments, administration is by intrathecal, intranasal, intracerebral, or intra-ocular delivery, depending on the target tissue.
[0207] In some embodiments, the CRISPR / Cas-based system is encapsulated in a lipid nanoparticle (LNP), where the LNP comprises an ionizable lipid, helper lipid, cholesterol, and132FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 polyethylene glycol-lipid component. The LNP may be formulated to deliver mRNA encoding the prime editor and one or more guide RNAs to hepatocytes, skeletal muscle, or other target cells.
[0208] In some embodiments, the CRISPR / Cas-based system is delivered by a recombinant adeno-associated virus (AAV) vector. In such embodiments, the AAV vector may comprise sequences encoding a prime editor, a base editor, or a split prime editor or base editor and pegRNA or sgRNA cassette, operably linked to a promoter active in the target tissue. Exemplary promoters include CMV, EFla, CAG, and tissue-specific promoters such as muscle creatine kinase (MCK).
[0209] In some embodiments, the delivery formulation further comprises one or more stabilizing excipients, cryoprotectants, or buffers suitable for parenteral administration. In certain embodiments, administration may be single-dose or repeated, according to a dosing schedule sufficient to achieve a therapeutically effective level of genome editing in target cells.
[0210] In some embodiments, the methods described herein may be performed in a subject having a GNE myopathy, such as a subject homozygous or heterozygous for the GNE M743T mutation. In some embodiments, the methods further comprise monitoring one or more biomarkers indicative of editing or phenotypic correction (e.g., plasma sialic acid levels, glycosylation profiles, or muscle function).EXAMPLESExample 1: Generation of a Cell Model Harboring the M734T GNE Disease Variant
[0211] HEK293 T cells were transfected by lipofection with plasmids expressing PE6c and pegRNAs that encoded the M743T variant in their RTT. Cells were grown for 72 hours after transfection before being split into two aliquots. One aliquot was used for DNA extraction and subsequent Illumina sequencing to confirm installation of the M743T; the other aliquot was used for limiting dilutions to isolate single clones. Sixteen metaphase spreads obtained from single cell clones were analyzed to evaluate the chromosomal composition of a cell line characterized by a composite karyotype as shown in FIG. 1A. Analysis of the DNA sequence, from three independent cell clones, located between genomic coordinates 36,217,382 to 36,217,420 according to human reference genome hg38 of chr9 indicated the presence of a single nucleotide variation (“T” to “C”) responsible for the emergence of the M743T GNE disease variant (FIG. IB).133FoleyHoagUS13167537.2Attorney Docket No. JHV-16925Example 2: Optimization of M743T Correction Using Different pegRNA Template Lengths and Silent Mutations
[0212] The present example describes the optimization of the correction of the M743T mutation in the GNE gene in the pathogenic human cell line model generated in Example 1 supra. Extensive efforts were undertaken to optimize the pegRNA design, and overall prime editing (PE) system parameters to maximize the correction efficiency of the M743T mutation. Prime editing guide RNAs (pegRNAs) with varying primer binding site (PBS) and template lengths were screened to identify those that enabled the most efficient correction of a point mutation as shown in FIG. 2. The full pegRNA sequences are detailed as set forth in Table 1 (SEQ ID NOs: 1-5, 318-322). The highest efficiency of M743T correction was achieved using the top-performing pegRNAs, including GNE_131, and GNE_138, which displayed about 20% correction rates.
[0213] The optimization process involved testing different pegRNAs with varying numbers of silent mutations in their reverse transcription templates (RTTs), aiming to improve editing efficiency and precision. HEK293T cells harboring the M743T variant mutation were transfected with plasmids expressing PE6c, along with plasmids expressing pegRNAs containing 1, 2, 3, or 4 silent mutations in the RTT, as shown in FIG. 3. The silent mutations did not alter the amino acid sequence of the GNE protein. Editing efficiency was measured as the percentage of sequencing reads containing the desired M743T correction. The full pegRNA sequences containing the different mutations in the RTT, are detailed as set forth in Table 7 below (SEQ ID NOs: 1634-1655) where nucleotides highlighted in bold represent either the mutation correction site (“T”) or the silent mutations introduced in the RTT sequence. The pegRNA SE11 that contained two silent mutations, exhibited the highest editing efficiency, achieving over 20% correction of the mutation.
[0214] These results support that optimization of pegRNAs ensures a more reliable editing outcome, making this approach valuable for applications requiring precise gene correction.Table 7: Exemplary RTT sequences harboring unique silent mutations134FoleyHoagUS13167537.2Attorney Docket No. JHV-16925Example 3: Optimization of Editing Efficiency Using Different PE Variants and SpacerSequences
[0215] The present example describes the optimization of the correction of the M743T mutation using various spacer sequence and prime editing (PE) systems.
[0216] The role of various silent edits in the RTT sequence in enhancing editing efficiency was evaluated alongside the use of Spacer 1 and Spacer 2 sequences (FIG. 4A). HEK293T cells135FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 were transfected with plasmids coding for the PE6c prime editor and pegRNAs SE11, SEI 8 (for Spacer 1), and SE15, SE16 (for Spacer 2). The pegRNAs sequences containing the different silent mutations in the RTT and spacer sequences, are detailed as set forth in Table 8 below (SEQ ID NOs: 1656-1657). The SE11 pegRNA, containing two silent mutations, achieved the highest correction efficiency for Spacer 1, achieving over 20% editing efficiency (FIG. 4B).
[0217] The performance of various prime editor variants (PE2max, PE6a, PE6b, PE6c, PE6d, PE6f, and PE6g) variants was further evaluated using Spacer 1 (FIG. 5A) and Spacer 2 (FIG. 5B) sequences. Each PE variant was tested with the respective pegRNAs, and the editing efficiency was measured as the percentage of sequencing reads containing the desired M743T correction, alongside the proportion of reads with indels. The combination of Spacer 1 with the PE6c variant demonstrated the most superior performance, achieving on average over 10% of editing efficiency. These findings indicate that silent edits in RTT improve editing precision, with Spacer 1 outperforming Spacer 2 in terms of efficiency.Table 8: Comparison of exemplary spacer sequencesExample 4: Optimization of Editing Efficiency Using MLHldn, Nicking Guides, and Dead Guides
[0218] The present example describes the optimization of the correction of the M743T mutation using various combinations of prime editing (PE) systems, MLHldn, different nicking single guide RNAs (nsgRNA) and dead single guide RNAs (dsgRNA).
[0219] The role of various nsgRNAs on the M743T correction efficiency using Spacer 1 was evaluated. The HEK293T cells from Example 1 supra were transfected with the PE6c or PE3b system, along with nsgRNAs positioned at different locations relative to the pegRNA binding site, including positions -59, -81, +34, +76, and -39 bp relative to the mutation site. A “no nsgRNA” treatment was used as a control. The combination of Spacer 1 with nsgRNA -39, which utilizes the PE3b editor strategy, provided the highest correction rate of while maintaining low levels of indels (FIG. 6A). Similarly, HEK293T cells were transfected with the PE6c system along with various136FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 dsgRNAs at positions +33, +21, and -90 bp relative to the mutation site. A “no dsgRNA” treatment was used as a control. The results indicate that the presence of dsgRNAs significantly improves editing efficiency, with dsgRNAs +33, provided the highest correction rate while maintaining low levels of indels (FIG. 6B). The nsgRNA and dsgRNA sequences are detailed as set forth in Table 9 below (SEQ ID NOs: 1657-1664), where nucleotides highlighted in bold represent either the mutation correction site (“T”) or the introduced silent mutations (“C”).
[0220] The role of MLHldn in enhancing the M743T correction efficiency using Spacer 1 was evaluated. The HEK293T cells from Example 1 supra were transfected with the PE6c or PE3b system, along with pegRNAs containing different numbers of silent mutations in the reverse transcription template (RTT). These conditions were tested with and without the addition of MLHldn. A “no MLHldn” treatment was used as a control (FIG. 7). The results indicate that the presence of MLHldn did not significantly alter editing precision. The pegRNAs sequences containing the different silent mutations in the RTT, are detailed as set forth in Table 10 below (SEQ ID NOs: 1665-1667), where nucleotides highlighted in bold represent either the mutation correction site (“T”) or the silent mutations introduced in the RTT sequence.
[0221] The role of various nsgRNAs and MLHldn on the M743T correction efficiency using Spacer 1 was evaluated. The HEK293T cells from Example 1 supra were transfected with the PE6c or PE3b system, along with nsgRNAs positioned at different locations relative to the pegRNA binding site, including positions -39, +34, -81, and +76 bp relative to the mutation site. A “no nsgRNA” treatment was used as a control. The combination of Spacer 1 with nsgRNA -39, which utilizes the PE3b editor strategy, provided the highest correction rate while maintaining low levels of indels (FIG. 8 and FIG. 9). Similarly, as shown above, the presence of MLHldn did not significantly alter editing precision.
[0222] The role of different combinations of nicking sgRNAs (nsgRNAs) and dead sgRNAs (dsgRNAs) on the M743T correction efficiency using Spacer 1 was evaluated. The HEK293T cells from Example 1 supra were transfected with the PE6c or PE3b system and various combinations of nsgRNAs and dsgRNAs, targeting different positions relative to the mutation site. A pegRNA transfection was used as a control. A pegRNA targeting the HEK3 locus was also used as a positive control (FIG. 10) The results showed that the combination of “nickl-deadl” and “nick5-deadl” with the PE6c / PE3b system combination provided the highest correction rate while maintaining low levels of indels.137FoleyHoagUS13167537.2Attorney Docket No. JHV-16925Table 9: Exemplary ngRNA and sgRNA sequencesTable 10: Exemplary extended RTT sequencesExample 5: Base Editing CBE6b System for Targeted M743T Correction
[0223] The present example describes the use of the base-editing CBE6b-SpRY system for targeted correction of the M743T mutation.
[0224] The role of the CBE6b-SpRY system in correcting the M743T mutation was evaluated. HEK293T cells from Example 1 supra were transfected with CBE6b-SpRY mRNA, along with a sgRNA targeting the M743T mutation site (FIG. 11A). Transfection with an empty vector (pucl9) was used as a negative control. As shown in FIG. 11B, transfection with CBE6b led to the introduction of both the desired M743T correction (C6) and silent bystander mutations (C4138FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 and Cl 1) within the GNE coding sequence (CDS) with approximately 60% editing efficiency. The sgRNA sequence is detailed as set forth in Table 11 below (SEQ ID NO: 1668) where the mutagenic nucleotide (“C”) is highlighted in bold.Table 11: Exemplary base editing GNE sgRNA sequenceExample 6: Validation of Prime Editing Strategies in Mouse Cells
[0225] The present example describes the optimization of the correction of the M743T mutation in example mouse cells (mouse Neuro-2a (N2a) cells) using various combinations of prime editing (PE) systems, different nicking single guide RNAs (nsgRNA) and dead single guide RNAs (dsgRNA).
[0226] The pegRNA and nicking-sgRNA components of the human GNE M743T prime editing system described in Examples 2-4 supra, were modified to match the corresponding sequences of the mouse Gne gene while maintaining their relative positioning to the pathogenic variant. The PAM site was conserved between species; however, fourteen nucleotide substitutions were required within the pegRNA spacer to achieve complete homology to the mouse locus. The RTT and PBS regions were similarly adjusted to preserve editing geometry. Sequence differences between human and mouse reagents are summarized in Table 12 below where nucleotides highlighted in bold represent either the mutation correction site (“T”) or the silent mutations introduced in the RTT sequence.
[0227] Mouse N2a cells were transfected by plasmid lipofection with the mouse-optimized prime editing construct. Genomic DNA was harvested 72 hours after transfection, and editing efficiency at the targeted Gne site was quantified by Illumina sequencing.
[0228] As shown in FIG. 12, various combinations of pegRNA were able to successfully edit the Gne gene in example mouse cells. Moreover, the pegRNA combination that produced optimal editing in human cells also yielded the high editing efficiency in mouse N2a cells as well. Specifically, the “nickl-deadl”, “SE2 pegRNA”, “Spacer 1” and “PE6c” combination provided the highest correction rate by achieving approximately 18% editing efficiency. These results indicate that the design principles and pegRNA configuration are transferable across species.Table 12: Exemplary extended RTT sequences139FoleyHoagUS13167537.2Attorney Docket No. JHV-16925Example 7: Off-Target Prediction and Validation of Prime Editing and Base Editing Strategies in Human Cells
[0229] The present example describes the assessment of potential off-target editing associated with the optimized GNE M743T prime editing (PE) or base editing (BE) system described in Examples 2-5 supra. Both computational prediction and experimental validation were performed to evaluate the specificity of the selected pegRNA and nicking-sgRNA combinations in human cells.
[0230] Potential off-target loci for the prime editing reagents were identified using the Cas- OFFinder in silico tool, which searches the human genome for sites with sequence similarity to the pegRNA spacer and nicking-sgRNA sequences. Forty sites with the closest homology and permissive PAM sequences were nominated for targeted validation.
[0231] Human cells were transfected with the optimized prime editing constructs shown in Examples 2^1 supra or with an untreated control plasmid, and genomic DNA was collected 72 hours post-transfection. Editing outcomes at two predicted off-target loci were quantified by deep sequencing.
[0232] As shown in FIG. 13A and FIG. 13B, no difference in editing efficiency or indel frequency was detected at either the “Off-Target 1” or “Off-Target 2” locus in PE-treated cells140FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 compared to untreated controls. The observed variation was within the background error rate of sequencing, confirming the high specificity of the editing system in the human cells.
[0233] In addition to computational prediction, a genome-wide experimental approach (CHANGE-seq) was used to detect potential off-target cleavage events in vitro. In this method, circularized human genomic DNA was incubated with Cas9-guide complexes corresponding to the PE pegRNA, PE nicking-sgRNA, and CBE6b-SpRY sgRNA. Cleavage sites were ligated to sequencing adapters and analyzed by Illumina sequencing.
[0234] FIG. 14 shows candidate off-target loci identified by CHANGE-seq. Each identified site is annotated with the number of sequencing reads supporting cleavage, the number of mismatches relative to the guide RNA, the genomic coordinates, and the corresponding sequence alignment. Each of the identified off-target sites may be assessed in treated and untreated cells.INCORPORATION BY REFERENCE
[0235] All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.
[0236] Also incorporated by reference in their entirety are any polynucleotide and polypeptide sequences which reference an accession number correlating to an entry in a public database, such as those maintained by The Institute for Genomic Research (TIGR) on the world wide web at tigr.org and / or the National Center for Biotechnology Information (NCBI) on the World Wide Web at ncbi.nlm.nih.gov.EQUIVALENTS
[0237] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the disclosure(s) described herein. Such equivalents are intended to be encompassed by the following claims.141FoleyHoagUS13167537.2
Claims
Attorney Docket No. JHV-16925CLAIMS1. A pegRNA for prime editing comprising (i) a spacer sequence, (ii) a scaffold, and (iii) an extension arm comprising a reverse-transcription template (RTT) and a primer-binding site (PBS), wherein the scaffold binds to a nucleic acid programmable DNA binding protein (napDNAbp), wherein the RTT comprises a mutation to revert a sequence encoding a GNE M743T allele to a wild-type GNE sequence, and wherein:(a) the PBS has a length between 9-14 nt; and(b) the RTT has a length between 34 to 70 nt.
2. The pegRNA of claim 1, wherein the RTT has a length between 50 to 70 nt.
3. The pegRNA of claim 2, wherein the RTT has a length between 58-60 nt.
4. The pegRNA of any one of claims 1-3, wherein the RTT further comprises at least one silent mutation.
5. The pegRNA of claim 4, wherein the RTT comprises 1 to 18 silent mutations.
6. The pegRNA of any one of claims 1-5, wherein the RTT comprises a sequence as set forth in any one of SEQ ID NOs: 327-652.
7. The pegRNA of any one of claims 1-5, wherein the RTT comprises a sequence of TGGGTGCTGCCAGCATGGTCCTGGACTACACAACACGCAGGATCTACTAGACCTGCAG GA (SEQ ID NO: 1644) orGGTGCTGCCAGCATGGTCCTGGACTACACAACACGCAGGATCTACTAGACCTGCAGGA (SEQ ID NO: 327).
8. The pegRNA of any one of claims 1-7, wherein PBS comprises a 7-15 nt sequence that is the reverse complement of at least a portion of the spacer sequence.142FoleyHoagUS13167537.2Attorney Docket No. JHV-169259. The pegRNA of any one of claims 1-8, wherein PBS comprises a sequence as set forth in any one of SEQ ID NOs: 979-1304.
10. The pegRNA of any one of claims 1-8, wherein PBS comprises a sequence of ACAGACATGGAC (SEQ ID NO: 979).
11. The pegRNA of any one of claims 1-10, wherein the spacer is 15-25 nt in length.
12. The pegRNA of any one of claims 1-11, wherein the spacer is 20 nt in length.
13. The pegRNA of any one of claims 1-12, wherein the spacer comprises a sequence of any one of SEQ ID NOs: 653-978.
14. The pegRNA of any one of claims 1-12, wherein the spacer comprises a sequence of AGAAGGTCCATGTCTGTTCC (SEQ ID NO: 653).
15. The pegRNA of any one of claims 1-14, wherein the spacer binds to a sequence in the genome that is 30-110 nt away from the corresponding mutation site to be corrected.
16. The pegRNA of any one of claims 1-15, wherein the spacer binds to a sequence in the genome that is 42 nt away from the corresponding mutation site to be corrected.
17. The pegRNA of any one of claims 1-16, wherein the scaffold binds to a napDNAbp that is a Cas9 protein or variant thereof.
18. The pegRNA of any one of claims 1-17, wherein the scaffold comprises a nucleic acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 1631.
19. The pegRNA of any one of claims 1-17, wherein the scaffold comprises a nucleic acid sequence of143FoleyHoagUS13167537.2Attorney Docket No. JHV-16925 GTTTAAGAGCTAAGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCA ACTTGAAAAAGTGGCACCGAGTCGGTGC (SEQ ID NO: 1631).
20. The pegRNA of any one of claims 1-19, wherein the extension arm is at the 3' end of the pegRNA.
21. The pegRNA of claim 20, comprising from 5' to 3' the spacer, the scaffold and the extension arm.
22. The pegRNA of any one of claims 1-19, wherein the extension arm is at the 5' end of the pegRNA.
23. The pegRNA of claim 22, comprising from 5' to 3' the extension arm, the scaffold and the spacer.
24. The pegRNA of any one of claims 1-23, comprising a nucleic acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 1-326.
25. The pegRNA of any one of claims 1-23, comprising or consisting of a nucleic acid sequence of any one of SEQ ID NOs: 1-326.
26. A gene editing complex for editing a GNE gene, the complex comprising a prime editor comprising (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain and (ii) a domain having an RNA-dependent DNA polymerase activity, and a pegRNA of any one of claims 1-25.
27. The complex of claim 26, wherein the napDNAbp has a nickase activity.
28. The complex of claim 26 or 27, wherein the napDNAbp is a Cas9 protein or variant thereof.144FoleyHoagUS13167537.2Attorney Docket No. JHV-1692529. The complex of claim 28, wherein the napDNAbp is a nuclease active Cas9, a nuclease inactive Cas9 (dCas9), or a Cas9 nickase (nCas9).
30. The complex of claim 28, wherein the napDNAbp is Cas9 nickase (nCas9).
31. The complex of any one of claims 26-30, wherein the domain comprising an RNA- dependent DNA polymerase activity is a reverse transcriptase.
32. The complex of any one of claims 26-31, further comprising a nicking guide RNA, a dead guide RNA or a combination thereof.
33. A method of correcting a M743T mutation in a GNE gene by prime editing, comprising contacting a target DNA sequence with a prime editor comprising (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain, and (ii) a domain having an RNA-dependent DNA polymerase activity, and a pegRNA of any one of claims 1-25.
34. A method of correcting a M743T mutation in a GNE gene by prime editing comprising contacting a target DNA sequence with the gene editing complex of any one of claims 26-32.
35. A method of correcting a M743T mutation in a GNE gene by cytosine base editing comprising contacting a target DNA sequence with a cytosine base editor comprising (i) a guide RNA, (ii) programmable DNA binding protein (napDNAbp), and (iii) a cytidine deaminase.
36. The method of claim 35, wherein the guide RNA comprises a spacer sequence that is at least 90% identical or at least 95% identical to SEQ ID NO: 1668.
37. The method of claim 35, wherein the guide RNA comprises a spacer sequence of SEQ ID NO: 1668.
38. The method of any one of claims 35-37, wherein the guide RNA comprises a scaffold that binds to a napDNAbp that is a Cas9 protein or variant thereof.145FoleyHoagUS13167537.2Attorney Docket No. JHV-1692539. A guide RNA for cytosine base editing, wherein the guide RNA comprises a spacer sequence that is at least 90% identical or at least 95% identical to SEQ ID NO: 1668.
40. A guide RNA for cytosine base editing comprising a spacer sequence of SEQ ID NO: 1668.
41. The guide RNA of claim 39, comprising a scaffold that binds to a napDNAbp that is a Cas9 protein or variant thereof.146FoleyHoagUS13167537.2