Prime editor delivery by aav
Patent Information
- Application Number
- EP2023833915
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-17
- Filing Date
- 2023-11-17
- Publication Date
- 2025-09-24
AI Technical Summary
The size limitation of Cas9 (>4 kb) hinders efficient delivery via recombinant adeno-associated virus (rAAV) for genome-editing applications, including prime editing systems, due to packaging constraints.
The prime editor is split into N-terminal and C-terminal portions and fused with intein halves for delivery via separate rAAV vectors, allowing for intein-mediated protein splicing to form a functional prime editor within cells, with optional linker and polymerase domains, including a napDNAbp and SpCas9, and optimized split sites to maintain functionality.
Enables efficient delivery and expression of prime editors, overcoming the size barrier and achieving high editing efficiencies in various tissues, including the central nervous system, with minimal disruption to the editor's function.
Smart Images

Figure 1.1
Abstract
Description
PRIME EDITOR DELIVERY BY AAV RELATED APPLICATIONS
[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application, U.S.S.N.63 / 426,336, filed on November 17, 2022, entitled “Prime editor delivery by AAV,” by David R. Liu et al. and to U.S. Provisional Application, U.S.S.N. 63 / 491,013, filed on March 17, 2023, entitled, “Prime editor delivery by AAV,” by David R. Liu et al., both of which are incorporated herein by reference in their entirety REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] The contents of the electronic sequence listing (B119570173WO00-SEQ- JQM.xml; Size: 596,955 bytes; and Date of Creation: November 16, 2023) is herein incorporated by reference in its entirety. GOVERNMENT FUNDING
[0003] This invention was made with government support under UG3AI150551, U01AI142756, R35GM118062, and RM1HG009490 awarded by US National Institutes of Health (NIH). The government has certain rights in the invention. FIELD
[0004] The present disclosure generally relates to genome editing and, in particular, to the in vivo delivery of genome editors using adeno associated virus (AAVs). BACKGROUND
[0005] Precise genome targeting technologies using the CRISPR / Cas9 system have recently been explored in a wide range of applications, including gene therapy. A major limitation to the application of Cas9 and Cas9-based genome-editing agents, including the prime editing systems, in gene therapy is the size of Cas9 (>4 kb), impeding its efficient delivery via recombinant adeno-associated virus (rAAV). SUMMARY
[0006] Described herein are systems, compositions, kits, and methods for delivering a prime editor to cells, e.g., via recombinant adeno-associated virus vectors (herein AAV). Insome embodiments, the prime editor comprises a fusion protein of a nucleic acid programmable DNA binding protein (herein napDNAbp) and a polymerase (e.g., a reverse transcriptase), optionally connected via a linker. Typically, a prime editor is “split” into an N-terminal portion and a C-terminal portion because the full-length prime editor, which is encoded by a polynucleotide of ~6.3kb, exceeds the packaging limit of rAAV (~4.9 kb). The prime editor comprises a napDNAbp domain, an optional linker, and a polymerase domain. In certain embodiments, the napDNAbp comprises a SpCas9. In certain embodiments the SpCas9 of the prime editor is split into two fragments at a split site located approximately midway through the sequence of the prime editor. In certain embodiments, the split site is located within loop sites that can accommodate fusion to the intein halves with minimal disruption to function of the prime editor. In certain embodiments, the split site is located between residues 844 and 845 or between residues 1024 and 1025 (SEQ ID NOs: 118-124). The N-terminal portion and C-terminal portion of the prime editor may each be fused to one member of the intein system, respectively. For example, the N-terminal portion of the prime editor fused to a C-terminal intein and the C-terminal portion of the prime editor fused to a N-terminal intein. When delivered on separate vectors (e.g., separate rAAV vectors) into one cell and co-expressed, the N-terminal portion of the prime editor fused to a C-terminal intein and the C-terminal portion of the prime editor fused to a N-terminal intein may be joined to form a complete and functional prime editor (e.g., via intein-mediated protein splicing). Further provided herein are empirical testing of regulatory elements in the delivery vectors for high expression levels of the split prime editor.
[0007] In certain embodiments, compositions comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N are described. In certain embodiments, compositions comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor are described. In some embodiments, compositions comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor are described. The split site, in some cases, may be selected so that the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14. In some embodiments, the three N-terminal amino acids of the C-terminal extein may be mutated from the native residues to a consensus Cys-Phe-Asn extein sequence (herein referred to as the 844-CFN split and 1024-CFN split respectively). In otherembodiments, zero, one, or two of the three N-terminal amino acids of the C-terminal extein may be mutated to the corresponding residues of the consensus Cys-Phe-Asn sequence; for example, the 844-SFL split, and 844-SFN split, and the 1024-SEQ split, 1024-SFQ split, 1024-SFN split, and 1024-SEN split, where the native residues for the three N-terminal amino acids of the C-terminal extein are SFL for residues 845-847 and SEQ for residues 1025-1027 in the native SpCas9 of the prime editor.
[0008] In certain embodiments, compositions comprising a rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C- terminus to an intein-N are described. In certain embodiments, compositions comprising a second rAAV particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor are also described. In other embodiments, compositions comprising a rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein- N and a second rAAV particle comprising a second nucleotide sequence encoding a intein-C fused to the N-terminus of a C-terminal portion of the prime editor is also described.
[0009] The split site, in some cases, may be selected so that the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14. In some embodiments, the three N-terminal amino acids of the C-terminal extein may be mutated from the native residues to a consensus Cys-Phe-Asn extein sequence (herein referred to as the 844-CFN split and 1024-CFN split respectively). In other embodiments, zero, one, or two of the three N-terminal amino acids of the C-terminal extein may be mutated to the corresponding residues of the consensus Cys-Phe-Asn sequence; for example, the 844-SFL split, and 844-SFN split, and the 1024-SEQ split, 1024-SFQ split, 1024-SFN split, and 1024- SEN split, where the native residues for the three N-terminal amino acids of the C-terminal extein are SFL for residues 845-847 and SEQ for residues 1025-1027 in the native SpCas9 of the prime editor.
[0010] In some cases, the first or second nucleotide sequence of the composition may further comprise a nucleotide sequence encoding a prime editing guide RNA (herein pegRNA) operably linked to a promoter. In other cases, the composition further comprises a third nucleotide sequence that encodes a pegRNA operably linked to a promoter. In some embodiments, the pegRNA comprises a 3′ stabilizing motif. In some embodiments, the pegRNA encodes one or more edits configured to natively evade an MMR pathway.
[0011] Aspects of the present disclosure relate to compositions comprising a first rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second rAAV particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor. The composition, in some cases, may further comprise a nucleotide sequence encoding a pegRNA operably linked to a promoter. In some embodiments, the pegRNA comprises a 3′ stabilizing motif. In some embodiments, the pegRNA encodes one or more edits configured to natively evade an MMR pathway.
[0012] Other aspects of the present disclosure relate to a first nucleotide sequence encoding a N-terminal portion of a prime editor fused edits C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor. In some embodiments, the first nucleotide sequence or the second nucleotide sequence further comprise a nucleotide encoding a pegRNA operably linked to a promoter. In some embodiments, the prime editor comprises a MMLV reverse transcriptase. In some embodiments the MMLV reverse transcriptase is codon optimized for expression in a mammalian cell.
[0013] Other embodiments of the present disclosure relate to a first rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second rAAV particle comprising a second nucleotide sequence encoding and intein-C fused to the N-terminus of a C-terminal portion of the prime editor. In some embodiments the first nucleotide sequence or the second nucleotide sequence further comprise a nucleotide encoding a peg RNA operably linked to a promoter. In some embodiments, the prime editor comprises a MMLV reverse transcriptase. In some embodiments, the MMLV reverse transcriptase is codon optimized for expression in a mammalian cell.
[0014] Other aspects of the present disclosure relate to a triple rAAV system. For example, in some embodiments, a composition comprises a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor and third nucleotide sequence encoding a pegRNA and a sgRNA.
[0015] In some aspects, the present disclosure relates to a first rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of the prime editor fused at its C- terminus to an intein-N and a second rAAV particle comprising a second nucleotide sequenceencoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor and third rAAV particle comprising third nucleotide sequence encoding a pegRNA and a sgRNA.
[0016] Aspects of the present disclosure further relate to a composition comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N- terminus of a C-terminal portion of the prime editor. In some embodiments, the prime editor further comprises a truncated MMLV reverse transcriptase with at least 80%, 85%, 90%, 95%, 99%, or 99.5% sequence identity to SEQ ID NO: 98.
[0017] In some embodiments, a first rAAV particle comprises a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second rAAV particle comprises a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor. In some embodiments, the prime editor further comprises a truncated MMLV reverse transcriptase with at least 80%, 85%, 90%, 95%, 99%, or 99.5% sequence identity to SEQ ID NO: 98.
[0018] Other aspects of the present disclosure relate to cells with the polynucleotides described herein, the vectors described herein, cells with the prime editors described herein, cells with a portion of the prime editor (spliced or not spliced), and cells with one or more rAAVs described herein.
[0019] The present disclosure also relates to a protein encoded by a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein- N. The present disclosure also relates to a protein encoded by a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor. The present disclosure also relates to the proteins encoded by a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor.
[0020] The present disclosure also relates to vectors for the delivery of a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein- N to cells. The present disclosure further relates to vectors for the delivery of a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor to cells. The present disclosure also relates to vectors for the delivery of a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C- terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor to cells.
[0021] The present disclosure further describes compositions comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N encapsulated with an LNP or VLP. In certain embodiments, compositions comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor encapsulated with an LNP or VLP are described. In some embodiments, compositions comprising a first nucleotide sequence encoding a N- terminal portion of a prime editor fused at its C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor encapsulated with an LNP or VLP are described.
[0022] The present disclosure further describes pharmaceutical compositions comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C- terminus to an intein-N and one or more pharmaceutically acceptable excipients. The present disclosure further describes pharmaceutical compositions comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor and one or more pharmaceutically acceptable excipients. The present disclosure further describes pharmaceutical compositions comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the primer editor and one or more pharmaceutically acceptable excipients. The present disclosure further describes pharmaceutical compositions comprising a protein encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and one or more pharmaceutically acceptable excipients. The present disclosure further describes pharmaceutical compositions comprising a protein encoding an intein-C fused to the N- terminus of a C-terminal portion of the primer editor and one or more pharmaceutically acceptable excipients. The present disclosure further describes pharmaceutical compositions comprising proteins encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and an intein-C fused to the N-terminus of a C-terminal portion of the primer editor and one or more pharmaceutically acceptable excipients.
[0023] Other aspects of the present disclosure relate to kits comprising anyone of the compositions and / or cells described herein.
[0024] Additional aspects of the present disclosure relate to methods of using the compositions, cells, and kits described herein. For example, in some cases, the present disclosure relates to a method for the systemic delivery of prime editors to the central nervous system (CNS) of a subject in need thereof. In some embodiments, the method comprisesadministering to the subject any one of the compositions, cells, or kits described herein. In some cases, the method comprises the use of rAAV particles capable of crossing the blood- brain barrier (e.g., AAV PHP.eB).
[0025] In some aspects, the disclosure relates to a method of editing an organ in a subject. In some embodiments, the method comprises administering to the subject a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein- N; and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C- terminal portion of the prime editor. In some embodiments, the second nucleotide sequence comprises a pegRNA and sgRNA. In some embodiment the pegRNA comprises a 3’ stabilizing motif. In some embodiments, the pegRNA encodes one or more edits configured to natively evade an MMR pathway. In some embodiments, the prime editor comprises a truncated MMLV reverse transcriptase.
[0026] In other aspects, the present disclosure relates to a method of contacting a cell with any one of the compositions or rAAV particles described herein. In some embodiments, contacting results in the delivery of the first nucleotide sequence and the second nucleotide sequence, and optionally third nucleotide sequence, into the cell. Once inside the cell the N- terminal portion of the prime editor and the C-terminal portion of the prime editor are joined to form a prime editor.
[0027] Additional aspects of the present disclosure relate to a method comprising administering to a subject in need thereof a therapeutically effective amount of any one of the compositions or rAAV particles described and a pharmaceutically acceptable
[0028] It should be appreciated that the foregoing concepts, and additional concepts discussed below, may be arranged in any suitable combination, as the present disclosure is not limited in this respect. Further, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non- limiting embodiments when considered in conjunction with the accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Non-limiting embodiments of the present disclosure will be described by way of example with reference to the accompanying figures, which are schematic and are not intended to be drawn to scale. In the figures, each identical or nearly identical component illustrated is typically represented by a single numeral. For purposes of clarity, not every component is labeled in every figure, nor is every component of each embodiment of thedisclosure shown where illustration is not necessary to allow those of ordinary skill in the art to understand the disclosure. In the figures:
[0030] FIGs.1A-1C. Initial development of a split prime editor in tissue culture and for AAV mediated in vivo prime editing a, Editing performance at three genomic loci of intein- split PE3 variants normalized to that of unsplit, canonical PE3 when delivered by plasmid transfection in HEK293T cells. The split site in SpCas9 and the identity of the three most N- terminal residues of the C-terminal extein are indicated below each bar. (FIG.1A) sub- optimal amount of editor plasmid was used in order to avoid saturating editing efficiencies. Dots represent values normalized to canonical PE3 activity and error bars represent mean±SEM of n= 3 biological replicates at three genomic loci. (FIG.1B) Schematic of v1 PE3-AAV architecture. (FIG.1C) In vivo editing activity of v1 PE3-AAV9 with pegRNA encoding the Dnmt1 +5 G-to-T edit delivered to neonatal C57BL / 6 pups by P0 ICV injection at a dose of 1×1011vg total (5×1010vg per half). Cortex (neocortex and hippocampus) was harvested, nuclei were isolated and sorted by FACS into bulk and GFP+ populations, and genomic DNA was harvested and analyzed by high-throughput screening (HTS). For c, dots represent individual mice and error bars represent mean±SEM of n=3-5 mice.
[0031] FIGs.2A-2C. AAV-mediated in vivo prime editing efficiency is dependent on the type of edit. (FIG.2A) Editing activity of PE2 and PE4 in N2a cells by plasmid transfection for four different edits at Dnmt1. N2a cells were transfected with either PE2 and pegRNA or PE4 (PE2 + MLH1dn) and pegRNA. Three days later, genomic DNA was harvested and analyzed by HTS. Dots represent values and error bars represent mean±SEM of n= 2 biological replicates. (FIG.2B), In vivo editing activity of v1 AAV9 PE3 delivered to neonatal C57BL / 6 pups on P0 by ICV at a total dose of 1×1011vg (5×1010vg per half). Cortex (neocortex and hippocampus) was harvested, nuclei were isolated and sorted by FACS into bulk and GFP+ populations, and genomic DNA was analyzed by HTS. (FIG.2C) In vivo editing activity of PE3 delivered via v1 PE3-AAV9 by RO injection to 6- to 8- week-old C57BL / 6 mice at a total dose of 1×1012vg. Three weeks after injection, bulk tissues were harvested and genomic DNA was isolated and analyzed by HTS. For (FIG.2B) and (FIG. 2C), dots represent individual mice and error bars represent mean±SEM of n=3-4 mice.
[0032] FIGs.3A-3E. Evaluation of factors limiting systemic prime editing efficiency. (FIG.3A) Comparison of unmodified pegRNA or epegRNA installing a +2 G-to-C edit in bulk liver. v1 PE3-AAV9 was delivered by RO injection to 6- to 8- week-old C57BL / 6 mice at a dose of 1×1012vg (5×1011vg per N- and C-terminal AAVs). Bulk tissue was harvested three weeks after injection and genomic DNA was isolated and analyzed by HTS. (FIG.3B)Comparison of unmodified pegRNA or epegRNA installing a +1 C-to-G edit in cortical tissue. v1 PE3-AAV9 was delivered via P0 ICV injection at a total dose of 1×1011vg. Neocortex was harvested three weeks after injection. Nuclei were isolated and GFP+ subpopulations were isolated by FACS. Genomic DNA was extracted and analyzed by HTS. (FIG.3C) Schematic of v2 PE3-AAV. (FIG.3D) C57BL / 6 mice were injected retro-orbitally with 1.25×1012vg total of either v2 AAV9-PE3 (5×1011vg each N- & C- terminal AAVs encoding PEmax plus 2.5×1011vg epegRNA / sgRNA AAV encoding the Dnmt1+2 G-to-C edit), or architecture-matched v2 Cas9 nuclease AAV9 (5×1011vg each N- & C- terminal AAVs encoding Cas9 nuclease plus 2.5×1011vg sgRNA AAV). Bulk liver, heart, and muscle tissues were harvested three weeks after injection and genomic DNA was analyzed by HTS. (FIG.3E), Comparison of PE3 and PE3max by v2 AAV9 PE at Dnmt1 installing the +1 C-to- G edit. For (FIG.3A), (FIG.3B), (FIG.3D), and (FIG.3E), dots represent individual mice and error bars represent mean±SEM of n=3 mice.
[0033] FIGs.4A-4E. Development of a dual-AAV system for efficient in vivo prime editing. The liver, heart, or brain. (FIG.4A), C57BL / 6 mice were injected retro-orbitally with 1.25×1012vg total of v2em PE3-AAV9 (5×1011vg each N & C terminal AAVs plus 2.5×1011vg epegRNA / sgRNA AAV) encoding PE3max with its native full- length RT, or PE3max with one of two truncated ∆RNaseH reverse transcriptases to install a Dnmt1 +1 C-to-G edit. Three weeks after injection, liver, heart, and muscle tissues were harvested and analyzed by HTS. (FIG.4B) C57BL / 6 pups were injected on P0 by ICV injection of a total of 5.7×1010vg total of v2em PE3-AAV (2.3×1010vg each N & C terminal AAVs plus 1.1×1010vg epegRNA / sgRNA AAV) encoding PE3max with full-length RT or PE3max with a truncated ∆RNaseH RT to install the Dnmt1 +2 G-to-C edit. Three weeks post injection, mice were harvested, neocortex was dissected, and nuclei were isolated and sorted by FACS. The bulk population (all nuclei) and a GFP-positive subpopulation were lysed and genomic DNA was analyzed by HTS. (FIG.4C) Schematic of v3em PE3max AAV. (FIG.4D) C57BL / 6 mice were injected retro-orbitally with either 1.25×1012vg total of v2em PE3 ∆RNaseH AAV9 (5×1011vg each N & C terminal AAVs plus 2.5×1011vg epegRNA / sgRNA AAV) or 1×1012vg total of v3em PE3-AAV9 (5×1011vg each N & C terminal AAVs), both installing the Dnmt1 +1 C-to-G edit. Three weeks after injection, liver, heart, and muscle tissues were harvested and analyzed by HTS. (FIG.4E) prime editing efficiency in bulk liver of v1em or v3em PE3-AAV9 following a single injection of AAV at two different doses. C57BL / 6 mice were injected retro-orbitally with 1×1011vg or 1×1012vg total of either v1em or v3em PE3- AAV9 containing an epegRNA encoding the Dnmt1 +2 G-to-C edit. Three weeks afterinjection, liver tissue was harvested and analyzed by HTS. For (FIG.4A), (FIG.4B), (FIG. 4D), and (FIG.4E) dots represent individual mice and error bars represent mean±SEM of n=2-5 mice.
[0034] FIGs.5A-5D. Improved prime editing in the mouse CNS. Comparison of prime editing from v1em and v3em PEmax AAVs in the CNS via direct injection in neonatal mice, or systemic injection in adult mice. (FIG.5A) Percent GFP-positive nuclei of capsid-, promoter-, and terminator- matched EGFP:KASH used for enrichment. (FIG.5B) Either v1em or v3em PE3-AAV9 encoding the +1 C-to-G edit at Dnmt1 was delivered via P0 ICV injection at a dose of 1×1011vg. Neocortex was harvested after three weeks and nuclei were isolated. GFP-positive subpopulations were sorted by FACS. Genomic DNA from bulk or GFP-positive nuclei was extracted and analyzed by HTS. (FIG.5C) Adult 6- to 8- weeks old were injected retro-orbitally with a total dose of 1×1012vg AAV PHP.eB with either v1em or v3em PE3-AAV encoding the +2 G-to-C edit at Dnmt1. Neocortex was harvested after three weeks and nuclei were isolated. GFP-positive subpopulations were sorted by FACS. Genomic DNA from bulk or GFP-positive nuclei was extracted and analyzed by HTS. (FIG.5D) Installation of APOE3 R136S (APOE Christchurch) in humanized APOE3 mice.1×1011vg v3em PE3-AAV9 was administered by ICV injection to mice on postnatal day P1 or P3. Genomic DNA or total RNA was isolated from neocortex and hippocampus (matched hemispheres). RNA was converted to cDNA and both genomic DNA and cDNA were analyzed by HTS. For (FIG.5A)-( FIG.5C) dots represent individual mice and error bars represent mean±SEM of n=2-4 mice.
[0035] FIGs.6A-6F. In vivo prime editing to install Pcsk9 Q152H and its effect on plasma cholesterol levels. (FIG.6A) Timing of injections of v3em PE3-AAV9, evaluation of prime editing, and plasma analysis. (FIG.6B) Bulk liver editing efficiencies for installation of Pcsk9 Q152H. (FIG.6C, FIG.6D) Plasma total cholesterol and plasma LDL cholesterol in mice treated with prime editing AAV normalized to untreated mice at the same age. Data are shown as mean±SEM for n=8 mice (4 male and 4 female), *P < 0.05. Significance was calculated using two-way repeated measures ANOVA with Sidak multiple comparisons, and is shown for the eight weeks timepoint. (FIG.6E, FIG.6F) Off-target prime editing for 10 CIRCLE-seq-nominated off-target loci (Off1- Off10) for pegRNA and nicking sgRNA from the livers of v3em PE3-AAV treated and untreated mice. For (FIG.6B), (FIG.6E) and (FIG. 6F), dots represent individual mice and error bars represent mean±SEM of n=8 mice (4 male and 4 female).
[0036] FIGs.7A, 7B. Efficiencies of various prime edits at Dnmt1 and the effect of transient MMR inhibition . (FIG.7A) Mouse N2A cells were transfected with plasmids encoding PE2, a +53 nicking sgRNA, and one pegRNA encoding the edit indicated below each bar. (FIG.7B) The impact of MMR recognition of the installed edit on editing efficiency was assessed by installing one of four edits (Dnmt1 +5 G to T, Dnmt1 +1 CCC insertion, Dnmt1 +1 C to G, or Dnmt1 +2 G to C) as PE2, PE3, PE4 (PE2 + MLH1dn), or PE5 (PE3 + MLH1dn). Data are shown as individual data points and mean±SEM for n=2-3 independent biological replicates.
[0037] FIG.8. Effect of incubation time on systemic in vivo prime editing efficiency. v1 PE-AAV9 was administered at a dose of 1x1012 vg total (5x1011 vg per half) by RO injection to 6- to 8-week-old C57BL / 6 mice and three weeks after, bulk tissues were harvested and isolated genomic DNA was analyzed by HTS. Dots represent individual mice and error bars represent mean±SEM of n=3 different mice.
[0038] FIGs.9A, 9B. Incorporation of epegRNAs and PEmax together substantially enhance prime editing in CNS. (FIG.9A, FIG.9B) PE3-AAVs encoding a Dnmt1 +1 C to G edit were delivered as v1, v1e, or v1em architecture at a dose of 1x1011vg (5x1010vg each N- and C- terminal AAVs) with 1x1010vg promoter-matched AAV9 EGFP:KASH. At three weeks, nuclei from neocortex were isolated and sorted by FACS to either all nuclei (bulk) or GFP+ populations. Genomic DNA was extracted and analyzed by HTS. While improvement from incorporation of epegRNA alone did not reach statistical significance (P=0.17), incorporation of epegRNA and PEmax (v1em) yielded highly statistically significant improvement in prime editing over v1 (for both bulk and GFP+ populations, P<0.0001). Significance was calculated by unpaired t-test with correction for multiple comparisons. Dots represent individual data points and error bars represent mean±SEM for n=3 different mice.
[0039] FIG.10. Activity of full-length and intein-split SaPE2 by plasmid transfection in HEK293T cells with a pegRNA encoding the HEK3 +6 G-to-T edit. An intein split at previously validated position 7401was assessed to maintain on-target editing efficiency compared to full-length (SaPE2). The three N-terminal residues of the C-terminal extein are indicated, with ‘SMP’ being the native residues to SaCas9, and CFN as the consensus residues of Npu intein. Dots represent individual data points and error bars represent mean±SEM for n=3 independent biological replicates.
[0040] FIG.11. PE ∆RNaseH maintains prime editing activity in cultured cells across a variety of edits. HEK293T cells were transfected with plasmids encoding PE2 or PE2 ∆RNaseH and pegRNA (PE2) or pegRNA and nicking sgRNA (PE3). Genomic DNA washarvested after three days and analyzed by HTS. Dots represent individual data points and error bars represent mean±SEM for n=3 independent biological replicates.
[0041] FIGs.12A, 12B. Transduction and RNA expression of PE-AAVs in liver. (FIG. 12A) In vivo comparison of viral genomes per diploid genome by promoter and dose in liver tissue. prime editors were expressed from either the v3em (Cbh promoter) or v1em (EFS promoter) architecture at the doses indicated. Isolated DNA was quantified by ddPCR using primers and a probe specific to either the N- or C-terminal half of SpCas9. Viral genome concentrations were normalized to Gapdh. (FIG.12B) In vivo comparison of PE transcript levels by promoter and dose in liver tissue. prime editors were expressed from either the v3em (Cbh promoter) or v1em (EFS promoter) architecture at the doses indicated. Mature mRNA transcripts were reverse transcribed, and PE transcript quantities from the resulting cDNA were measured by droplet digital PCR (ddPCR) using primers and a probe specific to either the N- or C-terminal half of SpCas9. Transcript concentrations were normalized to Gapdh. For (FIG.12A) and (FIG.12B), dots represent individual mice and error bars represent mean±SEM of n=4 different mice.
[0042] FIG.13. Comparison of v3em PE3-AAV split at two different positions.6- to 8- week-old C57BL / 6 mice were injected with 1x1012vg total of v3em PE3-AAV9 (5x1011vg each N & C terminal AAVs) using the intein-split indicated (residue 1024 or 844, in both cases with extein residues mutated to CFN) installing a Dnmt1 +1 C-to-G edit. Dots represent individual mice and error bars represent mean±SEM of n=3-5 mice.
[0043] FIG.14. Comparison of PE2 and PE3 in vivo. v1em AAV9 with an epegRNA encoding a +1 C-to-G edit or +2 G-to-C edit at Dnmt1 was delivered via P0 ICV injection at a dose of 1x1011vg (5x1010vg each N- and C-terminal AAVs) with 1x1010vg promoter- matched AAV9 EGFP:KASH. At three weeks, nuclei from neocortex were isolated and sorted by FACS to either all nuclei (bulk) or GFP+ populations. Genomic DNA was extracted and analyzed by HTS. Dots represent individual mice and error bars represent mean±SEM of n=2-4 mice.
[0044] FIGs.15A-15D. PE-mediated installation of protective APOE3 R136S Christchurch allele. (FIGs.15A-15C) Screening of pegRNAs with various PBS length (8-nt, 10-nt and 12-nt) and RTT (15-nt, 17-nt, 18-nt, 19-nt and 23-nt) along with three nicking guides (+25, +53 and +83) in HEK293T cells. Data are shown as mean±SEM for n=3 biological replicates. (FIG.15D) PE3 and PE5 (PE3 + MLH1dn) installation of protective APOE3 R136S Christchurch allele in immortalized mouse astrocytes. Data are shown as mean±SEM for n=3 biological replicates.
[0045] FIGs.16A-16D. PE-mediated installation of Pcsk9 Q155H. (FIG.16A, FIG.16B) Screening of pegRNAs with various two protospacers, multiple PBS lengths (8-nt, 9-nt, 10- nt, 11-nt, 12-nt, 13-nt, 14-nt and 15-nt) and RTT lengths (13-nt, 17-nt and 20-nt) in Neuro-2a cells. (FIG.16C) Screening of nicking guides (-45, -66, +30, +42, +63 and +101) in in Neuro-2a cells. (FIG.16D) Further improvements in prime editing efficiencies from using engineered pegRNAs, making a silent PAM edit, and installing silent MMR-evading edits with +64 nicking sgRNA. Data are shown as mean±SEM for n=3 biological replicates.
[0046] FIG.17. Installation of Pcsk9 Q155H in vivo. v3em PE3-AAV9 was injected into 6- to 8-week-old C57BL / 6 mice and liver was harvested eight weeks post injection to assess bulk editing. Dots represent individual mice and error bars represent mean±SEM for n=4 mice.
[0047] FIGs.18A-18E. prime editing and plasma analytes (normalized to untreated) data of either untreated or v3em PE3-AAV9 treated male and female C57BL / 6 mice. (FIG.18A) Installation of Pcsk9 Q152H in mouse liver using v3em PE AAV9. Dots represent individual mice and error bars represent mean±SEM for n=4 mice. (FIG.18B, FIG.18C) Total plasma cholesterol in C57BL / 6 male mice (FIG.18B) and female mice (FIG.18C) normalized to untreated. (FIG.18 D, FIG.18E) Plasma LDL cholesterol levels from C57BL / 6 mice male mice (FIG.18D) and female mice (FIG.18E) normalized to untreated. Data are shown as mean±SEM for n=4 mice.
[0048] FIGs.19A-19D. Raw (unnormalized) levels of plasma analytes of male and female mice treated with either v3em PE3-AAV9 installing Pcsk9 Q152H mutation or untreated. (FIG.19A, FIG.19B) Total plasma cholesterol in C57BL / 6 mice male mice a and female mice b. (FIG.19C, FIG.19D) Plasma LDL cholesterol levels from C57BL / 6 mice male mice c and female mice (FIG.19D). Data are shown as mean±SEM for n=4 mice.
[0049] FIGs.20A, 20B. Assessment of LDL receptor expression. (FIG.20A) Western blot evaluating LDL receptor expression on liver extracts of untreated, v3em PE-AAV9 Pcsk9 Q155H treated, Ldlr– / –and Pcsk9– / –mice. Ldlr– / –and Pcsk9– / –mice samples were used as negative and positive control, respectively. Raw, uncropped membrane images are shown in FIG.23. (FIG.20B) Densitometry-based quantification of LDL receptor expression from western blots. Data are normalized to untreated and shown as individual data points and mean±SEM for n=3-4 mice.
[0050] FIGs.21A-21D. Assessment of liver toxicity following systemic v3em PE-AAV8 injection. Adult mice were systemically injected v3em PE-AAV8 along with dose- and serotype- matched AAV8 encoding EGFP to control for toxicity induced by the AAV vector,dose- and serotype-matched Cas9 nickase to control for toxicity caused by Cas9 nickase and sgRNA portion of the prime editor, and a fourth cohort injected with saline alone. (FIG.21A) Histopathological assessment by hematoxylin and eosin staining of livers at eight weeks post- injection of saline or 1x1012vg total AAV8 encoding either EGFP, Cas9 nuclease, or prime editor. Scale bars, 50 μm, representative images from two separate mice shown per condition. (FIG.21B, FIG.21C) Plasma aspartate transaminase (AST) and alanine transaminase (ALT) levels. Data are shown as mean±SEM for n=4 mice. The dotted lines mark the upper limit of normal physiological range for AST and ALT in adult C57BL / 6 mice. (FIG.21D) Installation of Pcsk9 Q155H in vivo for liver toxicity study. v3em PE3- AAV8 was injected into 6- to 8- week old adult C57BL / 6 mice at a dose of 1x1012vg total by retro-orbital injection and liver was harvested eight weeks post injection to assess bulk editing. Dots represent individual mice and error bars represent mean±SEM for n=4 mice.
[0051] FIG.22. FACS gating strategy for brain nuclei. Nuclei were isolated from fresh or previously frozen brain tissue stained with DyeCycle Ruby. Gates were drawn around nuclei based on forward and side scatter area, then gated on singlets based on DyeCycle Ruby intensity, then sorted based on GFP fluorescence into a “bulk” population containing all nuclei regardless of GFP positivity, and a GFP+ population.
[0052] FIG.23. Raw western blot images. Uncropped images for LDL receptor (top) and ^-Actin (bottom) from liver extracts of v3em PE-AAV9 Pcsk9 Q155H treated, Ldlr– / –and Pcsk9– / –mice. Ldlr– / –and Pcsk9– / –mice samples were used as negative and positive control, respectively.
[0053] FIG.24. Dependence on splicing of split-intein PE2 activity. Full-length or intein-split PE2 were transfected into HEK239T cells with epegRNA and nicking sgRNA then analyzed by HTS three days after transfection. Splicing-inactive inteins contained a Cys1→Ala mutation in NpuN. A sub-optimal amount of editor plasmid was used to avoid saturating editing efficiencies. Two sites per condition are shown (HEK327bp insertion and RNF2 +1 C to A). Dots represent values and error bars represent mean+ / -SEM of n= 3 biological replicates at two genomic loci.
[0054] FIG.25. In vivo comparison of higher ratio of N-terminal v3em PE-AAV to C- terminal half v3em PE-AAV. v3em PE3-AAV9 encoding the Dnmt +2 G to C edit was delivered at a total dose of 1×1012vg at the ratio indicated: either a ratio of 1:1 (5×1011vg each N- and C-terminal halves) or a ratio of 2.5:1 (7.1×1011vg N-terminal half and 2.9×1011vg C-terminal half) and bulk tissues were analyzed three weeks post injection by HTS. Dots represent individual mice and error bars represent mean + / - SEM of n=3-4 different mice. BRIEF DESCRIPTION OF THE SEQUENCESCas9 variants with modified PAM specificities
[0137] In a particular embodiment, the Cas9 variant having expanded PAM capabilities is SpCas9 (H840A) VRQR (SEQ ID NO: 80), which has the following amino acid sequence (with the V, R, Q, R substitutions relative to the SpCas9 (H840A) of SEQ ID NO: 47 being show in bold underline. In addition, the methionine residue in SpCas9 (H840) was removed for SpCas9 (H840A) VRQR):
[0138] In another particular embodiment, the Cas9 variant having expanded PAM capabilities is SpCas9 (H840A) VRER, which has the following amino acid sequence (with the V, R, E, R substitutions relative to the SpCas9 (H840A) of SEQ ID NO: 47 being shown in bold underline . In addition, the methionine residue in SpCas9 (H840) was removed for SpCas9 (H840A) VRER):)
[0139] For example, a napDNAbp domain with altered PAM specificity, such as a domain with at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with wild type Francisella novicida Cpf1 (D917, E1006, and D1255) (SEQ ID NO: 82), which has the following amino acid sequence:
[0140] An additional napDNAbp domain with altered PAM specificity, such as a domain having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with wild type Geobacillus thermodenitrificans Cas9 (SEQ ID NO: 34), which has the following amino acid sequence:
[0141] The disclosed fusion proteins may comprise a napDNAbp domain having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with wild type Natronobacterium gregoryi Argonaute (SEQ ID NO: 84), which has the following amino acid sequence:Wild Type Polymerases
[0143] Reverse transcriptase (M-MLV RT) wild type moloney murine leukemia virus Used in PE1 (prime editor 1 fusion protein disclosed herein)AAV Sequences
[0197] Sequence of N-term v1em PE3-AAV (5' to 3'), 4915 bp ITR-EFS promoter-N-term PEmax (start codon-SV40NLS-SpCAS9)NpuN-SV40NLS-W3- bGH polyA-sgRNA (protospacer in bold)-human U6-ITR (Underlined sequences contain restriction sites for cloning; EFS promoter is in bold, N-term PEmax (start codon-SV40NLS-SpCAS9)NpuN is italicized; SV40NLS is bolded and underlined, W3-bGH polyA is italicized and underlined, sgRNA is bolded and italicized, and human U6 is underlined with dotted line).
[0198] Sequence of C-term v1em PE3-AAV (5' to 3'), 4,880 bp ITR-EFS promoter- SV40NLS- NpuC-C-term PEmax (SpCas9-RT-SV40NLS)-W3-bGH polyA-epegRNA (protospacer in bold)-human U6-ITR (Underlined sequences contain restriction sites for cloning; EFS promoter is in bold, C-term PEmax (SpCas9-RT-SV40NLS) is italicized; NpuC is squiggle underlined; SV40NLS is bolded and underlined, W3-bGH polyA is italicized and underlined, epegRNA is bolded and italicized, and human U6 is underlined with dotted line).
[0199] Sequence of N-term v2em PE3-AAV (5' to 3'), 5083 bp ITR-Cbh promoter-N-term PEmax (start codon-SV40NLS-SpCas9)-NpuN-SV40NLS-W3- bGH polyA-ITR (Underlined sequences contain restriction sites for cloning; Cbh promoter is in bold, N-term PEmax (start codon-SV40NLS-SpCas9 is italicized; NpuN is squiggle underlined; SV40NLS is bolded and underlined, W3-bGH polyA is italicized and underlined).)
[0200] Sequence of C-term v2em PE3-AAV (5' to 3'), 4978 bp ITR-Cbh promoter-SV40NLS-NpuC-C-term PEmax (SpCas9-RT-SV40NLS)-W3- bGHpolyA- ITR (Underlined sequences contain restriction sites for cloning; Cbh promoter is in bold, C-term PEmax (SpCas9-RT-SV40NLS) is italicized; NpuC is squiggle underlined; SV40NLS is bolded and underlined, W3-bGH polyA is italicized and underlined))
[0201] Sequence of v2em EGFP:KASH pegRNA / sgRNA (5' to 3'), 3438 bpITR-Cbh promoter-EGFP:KASH -sgRNA -mouse U6-epegRNA -human U6-ITR (Underlined sequences contain restriction sites for cloning; Cbh promoter is in bold, EGFP:KASH is italicized; sgRNA is bolded and underlined, epegRNA is italicized and underlined, mouse U6 is squiggle underlined, and human U6 is italicized and bolded).
[0202] Sequence of N-term v3em PE3-AAV (5' to 3'), 4,740 bpITR-Cbh promoter-N-term PEmax (start codon-SV40NLS-SpCas9)-NpuN-SV40NLS-SV40 late polyA-ITR (Sequences in grey contain restriction sites for cloning) (Underlined sequences contain restriction sites for cloning; Cbh promoter is in bold, N-term PEmax (start codon-SV40NLS-SpCas9) is italicized; NpuN is squiggle underlined; SV40NLS is bolded and underlined, SV40 late polyA is italicized and underlined).
[0203] Sequence of C-term v3em PE3-AAV (5' to 3'), 4,968 bp ITR-Cbh promoter-SV40NLS-NpuC-C-term PEmax (SpCas9-RT-SV40NLS)-SV40 late polyA-sgRNA (protospacer in bold)-mouse U6-epegRNA (protospacer in bold)-human U6- ITR (Underlined sequences contain restriction sites for cloning; Cbh promoter is in bold (SEQ ID NO: 436), C-term PEmax (SpCas9-RT-SV40NLS) is italicized (SEQ ID NO: 437); NpuC issquiggle underlined (SEQ ID NO: 438); SV40NLS is bolded and underlined (SEQ ID NO: 439), SV40 late polyA is italicized and underlined (SEQ ID NO: 440), epegRNA is bolded and squiggle underlined (SEQ ID NO: 441), sgRNA is dotted underlined (SEQ ID NO: 442), mouse U6 is double underlined (SEQ ID NO: 443), and human U6 is double underline and italicized(SEQ ID NO: 444)).LINKERS
[0225] In some other embodiments, the linker comprises the amino acid sequence (GGGGS)n (SEQ ID NO: 141), (G)n (SEQ ID NO: 142), (EAAAK)n (SEQ ID NO: 143), (GGS)n (SEQ ID NO: 144), (SGGS)n (SEQ ID NO: 145), (XP)n (SEQ ID NO: 146), or any combination thereof, wherein n is independently an integer between 1 and 30, and wherein X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS)n (SEQ ID NO: 144), wherein n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 147). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 148). In some embodiments, the linker comprises the amino acid sequence SGGSGGSGGS (SEQ ID NO: 149). In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 150). In other embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSS GGS (SEQ ID NO: 151, 60AA).
[0226] In particular, the following linkers can be used in various embodiments to join prime editor domains with one another:Nuclear Localization Sequences
[0233] PKKKRKV (SEQ ID NO: 156)
[0234] MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 157)
[0235] Bipartite NLS typically contains two clusters of basic amino acids, separated by a spacer of about 10 amino acids. One non-limiting example of a bipartite NLS is the NLS of nucleoplasmin, KRPAATKKAGQAKKKK (SEQ ID NO: 158) (spacer underlined).
[0236] SV40 bipartite NLS
[0237] KRTADGSEFESPKKKRKV (SEQ ID NO: 159), e.g., as described in Hodel et al., J Biol Chem.2001 Jan 12;276(2):1317-25, incorporated herein by reference)
[0238] Kanadaptin bipartite NLS
[0239] KKTELQTTNAENKTKKL (SEQ ID NO: 160), e.g., as described in Hubner et al., Biochem J.2002 Jan 15;361(Pt 2):287-96, incorporated herein by reference);
[0240] influenza A nucleoprotein bipartite NLS
[0241] KRGINDRNFWRGENGRKTR (SEQ ID NO: 161), e.g., as described in Ketha et al., BMC Cell Biology.2008;9:22, incorporated herein by reference)
[0242] ZO-2 bipartite NLS
[0243] RKSGKIAAIVVKRPRK (SEQ ID NO: 162), e.g., as described in Quiros et al., Nusrat A, ed. Molecular Biology of the Cell.2013;24(16):2528-2543, incorporated herein by reference).
[0296] In some embodiments, N-terminal amino acids at positions 1-3 (e.g., SEQ) of SEQ ID NO: 456 can be mutated to SFQ, SFN, SEN, or CFN. RNA-protein interaction domain Sequences
[0297] The nucleotide sequence of the MS2 hairpin (or equivalently referred to as the “MS2 aptamer”) is: GCCAACATGAGGATCACCCATGTCTGCAGGGCC (SEQ ID NO: 192)Additional PE elements
[0300] Exemplary glycosylases are provided below. The catalytically inactivated variants of any of these glycosylase domains are iBERs that may be fused to the napDNAbp or polymerase domain of the prime editors provided in this disclosure.pegRNAs
[0305] For the S. pyogenes Cas9 target site of the form: MMMMMMMMNNNNNNNNNNNNXGG (SEQ ID NO: 198), where NNNNNNNNNNNNXGG (SEQ ID NO: 199) (N is A, G, T, or C; and X can be anything)
[0306] For the S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGG (SEQ ID NO: 200), where NNNNNNNNNNNXGG (SEQ ID NO: 201) (N is A, G, T, or C; and X can be anything).
[0307] For the S. thermophilus CRISPR1Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXXAGAAW(SEQ ID NO: 202), where NNNNNNNNNNNNXXAGAAW (SEQ ID NO: 203) (N is A, G, T, or C; X can be anything; and W is A or T).
[0308] A unique target sequence in a genome may include an S. thermophilus CRISPR 1 Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXXAGAAW (SEQ ID NO: 204), where NNNNNNNNNNNXXAGAAW (SEQ ID NO: 205) (N is A, G, T, or C; X can be anything; and W is A or T).
[0309] For the S. pyogenes Cas9, a unique target sequence in a genome may include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGGXG (SEQ ID NO: 206) where NNNNNNNNNNNNXGGXG (SEQ ID NO: 207) (N is A, G, T, or C; and X can be anything).
[0310] A unique target sequence in a genome may include an S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGGXG (SEQ ID NO: 208) where NNNNNNNNNNNXGGXG (SEQ ID NO: 209) (N is A, G, T, or C; and X can be anything).
[0311] In each of the above sequences “M” may be A, G, T, or C, and need not be considered in identifying a sequence as unique. listed 5′ to 3′ “N” represents a base of a guide sequence, the first block of lower case letters represent the tracr mate sequence, and the second block of lower case letters represent the tracr sequence, and the final poly-T sequence represents the transcription terminator.
[0312] NNNNNNNNgtttttgtactctcaagatttaGAAAtaaatcttgcagaagctacaaagataaggcttcatgccg aaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO: 210)
[0313] listed 5′ to 3′
[0314] “N” represents a base of a guide sequence, the first block of lower case letters represent the tracr mate sequence, and the second block of lower case letters represent the tracr sequence, and the final poly-T sequence represents the transcription terminator.
[0315] NNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcat gccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT (SEQ ID NO: 211)
[0316] listed 5′ to 3′
[0317] “N” represents a base of a guide sequence, the first block of lower case letters represent the tracr mate sequence, and the second block of lower case letters represent the tracr sequence, and the final poly-T sequence represents the transcription terminator.
[0318] NNNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggct tcatgccgaaatcaacaccctgtcattttatggcagggtgtTTTTT (SEQ ID NO: 212)
[0319] listed 5′ to 3′
[0320] “N” represents a base of a guide sequence, the first block of lower case letters represent the tracr mate sequence, and the second block of lower case letters represent the tracr sequence, and the final poly-T sequence represents the transcription terminator.
[0321] NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAAtagcaagttaaaataaggctagtccg ttatcaacttgaaaaagtggcaccgagtcggtgcTTTTTT (SEQ ID NO: 213)
[0322] listed 5′ to 3′
[0323] “N” represents a base of a guide sequence, the first block of lower case letters represent the tracr mate sequence, and the second block of lower case letters represent the tracr sequence, and the final poly-T sequence represents the transcription terminator.
[0324] NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAATAGcaagttaaaataaggctagtc cgttatcaacttgaaaaagtgTTTTTTT (SEQ ID NO: 214)
[0325] listed 5′ to 3′
[0326] “N” represents a base of a guide sequence, the first block of lower case letters represent the tracr mate sequence, and the second block of lower case letters represent the tracr sequence, and the final poly-T sequence represents the transcription terminator.
[0327] NNNNNNNNNNNNNNNNNNNNgttttagagctagAAATAGcaagttaaaataaggctagtcc gttatcaTTTTTTTT (SEQ ID NO: 215)
[0328] 5ʹ-[guide sequence]- guuuuagagcuagaaauagcaaguuaaaauaaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuu uuu-3ʹ (SEQ ID NO: 216)
[0329] PEgRNA expression platform consisting of pCMV, Csy4 hairpin, the PEgRNA, and MALAT1 ENEDEFINITIONS
[0339] As used herein and in the claims, the singular forms “a,” “an,” and “the” include the singular and the plural reference unless the context clearly indicates otherwise. Thus, for example, a reference to “an agent” includes a single agent and a plurality of such agents. A subject in need thereof
[0340] “A subject in need thereof” refers to an individual who has a disease, a sign and / or symptom of a disease, or a predisposition toward a disease, with the purpose to cure, heal, alleviate, relieve, alter, remedy, ameliorate, improve, or affect the disease, the symptom of the disease, or the predisposition toward the disease. In some embodiments, the subject is a mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is human. In some embodiments, the mammal is a rodent. In some embodiments, the rodent is a mouse. In some embodiments, the rodent is a rat. In some embodiments, the mammal is a companion animal. A “companion animal” refers to pets and other domestic animals. Non-limiting examples of companion animals include dogs and cats; livestock, such as horses, cattle, pigs, sheep, goats, and chickens; and other animals, such as mice, rats, guinea pigs, and hamsters. Adeno-associated virus (AAV)
[0341] An “adeno-associated virus” or “AAV” is a virus which infects humans and some other primate species. The wild-type AAV genome is a single-stranded deoxyribonucleic acid (ssDNA), either positive- or negative-sensed. The genome comprises two inverted terminal repeats (ITRs), one at each end of the DNA strand, and two open reading frames (ORFs): rep and cap between the ITRs. The rep ORF comprises four overlapping genes encoding Rep proteins required for the AAV life cycle. The cap ORF comprises overlapping genes encoding capsid proteins: VP1, VP2 and VP3, which interact together to form the viralcapsid. VP1, VP2 and VP3 are translated from one mRNA transcript, which can be spliced in two different manners: either a longer or shorter intron can be excised resulting in the formation of two isoforms of mRNAs: a ~2.3 kb- and a ~2.6 kb-long mRNA isoform. The capsid forms a supramolecular assembly of approximately 60 individual capsid protein subunits into a non-enveloped, T-1 icosahedral lattice capable of protecting the AAV genome. The mature capsid is composed of VP1, VP2, and VP3 (molecular masses of approximately 87, 73, and 62 kDa respectively) in a ratio of about 1:1:10.
[0342] rAAV particles may comprise a nucleic acid vector (e.g., a recombinant genome), which may comprise at a minimum: (a) one or more heterologous nucleic acid regions comprising a sequence encoding a protein or polypeptide of interest (e.g., a split prime editor) or an RNA of interest (e.g., a gRNA), or one or more nucleic acid regions comprising a sequence encoding a Rep protein; and (b) one or more regions comprising inverted terminal repeat (ITR) sequences (e.g., wild-type ITR sequences or engineered ITR sequences) flanking the one or more nucleic acid regions (e.g., heterologous nucleic acid regions). In some embodiments, the nucleic acid vector is between 4 kb and 5 kb in size (e.g., 4.2 to 4.7 kb in size). In some embodiments, the nucleic acid vector further comprises a region encoding a Rep protein. In some embodiments, the nucleic acid vector is circular. In some embodiments, the nucleic acid vector is single-stranded. In some embodiments, the nucleic acid vector is double-stranded. In some embodiments, a double-stranded nucleic acid vector may be, for example, a self-complimentary vector that contains a region of the nucleic acid vector that is complementary to another region of the nucleic acid vector, initiating the formation of the double-strandedness of the nucleic acid vector. Cas9
[0343] As used herein, the term “Cas9,” “Cas9 protein,” or “Cas9 nuclease” refers to an RNA-guided nuclease comprising a Cas9 protein (e.g., Cas9 nucleases from a variety of bacterial species), a fragment, a variant (e.g., a catalytically inactive Cas9 or a Cas9 nickase), or a fusion protein (e.g., a Cas9 fused to another protein domain) thereof. A Cas9 nuclease is also referred to sometimes as a casn1 nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems correctprocessing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3- aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA target complementary to the spacer. In nature, DNA- binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gNRA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. Non-limiting examples of Cas9 proteins and their respective amino acid sequence are provided in Example 1.
[0344] A nuclease-inactive Cas9 protein may interchangeably be referred to as a “dCas9” protein (for nuclease-“dead” Cas9). Methods for generating a Cas9 protein (or a fragment thereof) having an inactive DNA cleavage domain are known (See, e.g., Jinek et al., Science. 337:816-821(2012); Qi et al., (2013) Cell.28;152(5):1173-83, incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, whereas the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science.337:816-821(2012); Qi et al., Cell.28;152(5):1173-83 (2013). Additional suitable nuclease-inactive dCas9 domains will be apparent to those of skill in the art based on this disclosure and knowledge in the field, and are within the scope of this disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (See, e.g., Prashant et al., Nature Biotechnology.2013; 31(9):833-838, incorporated herein by reference).
[0345] A Cas9 nickase is able to cleave one strand of the double strand DNA. A Cas9 nickase may be generated by introducing an inactivating mutation into either the HNH domain or the RuvC1 domain. For example, an inactivating mutation (D10A) may be introduced in the RuvC1 domain of the S. pyogenes Cas9, while the HNH domain remains active, i.e., the residue at position 840 remains a histidine. Such Cas9 variants are able to generate a single-strand DNA break (nick) at a specific location based on the gRNA-defined target sequence. One skilled in the art is able to identify the catalytic residues in the RuvC1 and HNH domains of any known Cas9 proteins and introduce inactivating mutations to generate a corresponding dCas9 or nCas9.CRISPR
[0346] CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of prior infections by a virus that have invaded the prokaryote. The snippets of DNA are used by the prokaryotic cell to detect and destroy DNA from subsequent attacks by similar viruses and effectively compose, along with an array of CRISPR-associated proteins (including Cas9 and homologs thereof) and CRISPR-associated RNA, a prokaryotic immune defense system. In nature, CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA target complementary to the RNA. Specifically, the target strand not complementary to crRNA is first cut endonucleolytically, then trimmed 3´-5′ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gNRA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species – the guide RNA. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the entire contents of which is hereby incorporated by reference. Cas9 recognizes a short motif in the CRISPR repeat sequences (the PAM or protospacer adjacent motif) to help distinguish self versus non-self. CRISPR biology, as well as Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A.98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA- guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based onthis disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.
[0347] In certain types of CRISPR systems (e.g., type II CRISPR systems), correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3- aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular nucleic acid target complementary to the RNA. Specifically, the target strand not complementary to crRNA is first cut endonucleolytically, then trimmed 3′-5′ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gRNA”) can be engineered so as to incorporate embodiments of both the crRNA and tracrRNA into a single RNA species— the guide RNA.
[0348] In general, a “CRISPR system” refers collectively to transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated (“Cas”) genes, including sequences encoding a Cas gene, a tracr (trans-activating CRISPR) sequence (e.g. tracrRNA or an active partial tracrRNA), a tracr mate sequence (encompassing a “direct repeat” and a tracrRNA-processed partial direct repeat in the context of an endogenous CRISPR system), a guide sequence (also referred to as a “spacer” in the context of an endogenous CRISPR system), or other sequences and transcripts from a CRISPR locus. The tracrRNA of the system is complementary (fully or partially) to the tracr mate sequence present on the guide RNA. Downstream
[0349] As used herein, the terms “upstream” and “downstream” are terms of relativity that define the linear position of at least two elements located in a nucleic acid molecule (whether single or double-stranded) that is orientated in a 5ʹ-to-3ʹ direction. In particular, a first element is upstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 5’ to the second element. For example, a SNP is upstream of a Cas9-induced nick site if the SNP is on the 5’ side of the nick site. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 3’ to the second element. For example, a SNP is downstream of a Cas9-induced nick site if the SNP is on the 3’ side of the nick site. Thenucleic acid molecule can be a DNA (double or single stranded). RNA (double or single stranded), or a hybrid of DNA and RNA. The analysis is the same for single strand nucleic acid molecule and a double strand molecule since the terms upstream and downstream are in reference to only a single strand of a nucleic acid molecule, except that one needs to select which strand of the double stranded molecule is being considered. Often, the strand of a double stranded DNA which can be used to determine the positional relativity of at least two elements is the “sense” or “coding” strand. In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. Thus, as an example, a SNP nucleobase is “downstream” of a promoter sequence in a genomic DNA (which is double-stranded) if the SNP nucleobase is on the 3' side of the promoter on the sense or coding strand. Error-prone reverse transcriptase
[0350] As used herein, the term “error-prone” reverse transcriptase (or more broadly, any polymerase) refers to a reverse transcriptase (or more broadly, any polymerase) that occurs naturally or which has been derived from another reverse transcriptase (e.g., a wild type M- MLV reverse transcriptase) which has an error rate that is less than the error rate of wild type M-MLV reverse transcriptase. The error rate of wild type M-MLV reverse transcriptase is reported to be in the range of one error in 15,000 (higher) to 27,000 (lower). An error rate of 1 in 15,000 corresponds with an error rate of 6.7 x 10-5. An error rate of 1 in 27,000 corresponds with an error rate of 3.7 x 10-5. See Boutabout et al. (2001) “DNA synthesis fidelity by the reverse transcriptase of the yeast retrotransposon Ty1,” Nucleic Acids Res 29(11):2217–2222, which is incorporated herein by reference. Thus, for purposes of this application, the term “error prone” refers to those RT that have an error rate that is greater than one error in 15,000 nucleobase incorporation (6.7 x 10-5or higher), e.g., 1 error in 14,000 nucleobases (7.14 x 10-5or higher), 1 error in 13,000 nucleobases or fewer (7.7 x 10-5or higher), 1 error in 12,000 nucleobases or fewer (7.7 x 10-5or higher), 1 error in 11,000 nucleobases or fewer (9.1 x 10-5or higher), 1 error in 10,000 nucleobases or fewer (1 x 10-4or 0.0001 or higher), 1 error in 9,000 nucleobases or fewer (0.00011 or higher), 1 error in 8,000 nucleobases or fewer (0.00013 or higher) 1 error in 7,000 nucleobases or fewer (0.00014 or higher), 1 error in 6,000 nucleobases or fewer (0.00016 or higher), 1 error in 5,000 nucleobases or fewer (0.0002 or higher), 1 error in 4,000 nucleobases or fewer (0.00025 or higher), 1 error in 3,000 nucleobases or fewer (0.00033 or higher), 1 error in2,000 nucleobase or fewer (0.00050 or higher), or 1 error in 1,000 nucleobases or fewer (0.001 or higher), or 1 error in 500 nucleobases or fewer (0.002 or higher), or 1 error in 250 nucleobases or fewer (0.004 or higher). Extein
[0351] The term “extein,” as used herein, refers to an polypeptide sequence that is flanked by an intein and is ligated to another extein during the process of protein splicing to form a mature, spliced protein. Typically, an intein is flanked by two extein sequences that are ligated together when the intein catalyzes its own excision. Exteins, accordingly, are the protein analog to exons found in mRNA. For example, a polypeptide comprising an intein may be of the structure extein(N) – intein – extein(C). After excision of the intein and splicing of the two exteins, the resulting structures are extein(N) – extein(C) and a free intein. In various configurations, the exteins may be separate proteins (e.g., half of a Cas9 or PE fusion protein), each fused to a split-intein, wherein the excision of the split inteins causes the splicing together of the extein sequences. Extension arm
[0352] The term “extension arm” refers to a nucleotide sequence component of a PEgRNA which provides several functions, including a primer binding site and an edit template for reverse transcriptase. In some embodiments, the extension arm is located at the 3ʹ end of the guide RNA. In other embodiments, the extension arm is located at the 5ʹ end of the guide RNA. In some embodiments, the extension arm also includes a homology arm. In various embodiments, the extension arm comprises the following components in a 5ʹ to 3ʹ direction: the homology arm, the edit template, and the primer binding site. Since polymerization activity of the reverse transcriptase is in the 5ʹ to 3ʹ direction, the preferred arrangement of the homology arm, edit template, and primer binding site is in the 5ʹ to 3ʹ direction such that the reverse transcriptase, once primed by an annealed primer sequence, polymerases a single strand of DNA using the edit template as a complementary template strand. Further details, such as the length of the extension arm, are described elsewhere herein.
[0353] The extension arm may also be described as comprising generally two regions: a primer binding site (PBS) and a DNA synthesis template, for instance. The primer binding site binds to the primer sequence that is formed from the endogenous DNA strand of the target site when it becomes nicked by the prime editor complex, thereby exposing a 3ʹ end on the endogenous nicked strand. As explained herein, the binding of the primer sequence to theprimer binding site on the extension arm of the PEgRNA creates a duplex region with an exposed 3ʹ end (i.e., the 3ʹ of the primer sequence), which then provides a substrate for a polymerase to begin polymerizing a single strand of DNA from the exposed 3ʹ end along the length of the DNA synthesis template. The sequence of the single strand DNA product is the complement of the DNA synthesis template. Polymerization continues towards the 5ʹ of the DNA synthesis template (or extension arm) until polymerization terminates. Thus, the DNA synthesis template represents the portion of the extension arm that is encoded into a single strand DNA product (i.e., the 3ʹ single strand DNA flap containing the desired genetic edit information) by the polymerase of the prime editor complex and which ultimately replaces the corresponding endogenous DNA strand of the target site that sits immediate downstream of the PE-induced nick site. Without being bound by theory, polymerization of the DNA synthesis template continues towards the 5ʹ end of the extension arm until a termination event. Polymerization may terminate in a variety of ways, including, but not limited to (a) reaching a 5ʹ terminus of the PEgRNA (e.g., in the case of the 5ʹ extension arm wherein the DNA polymerase simply runs out of template), (b) reaching an impassable RNA secondary structure (e.g., hairpin or stem / loop), or (c) reaching a replication termination signal, e.g., a specific nucleotide sequence that blocks or inhibits the polymerase, or a nucleic acid topological signal, such as, supercoiled DNA or RNA. Fusion Protein
[0354] Two proteins or protein domains are considered to be “fused” when a peptide bond is formed linking the two proteins or two protein domains. In some embodiments, a linker (e.g., a peptide linker) is present between the two proteins or two protein domains. The term “linker,” as used herein, refers to a chemical group or a molecule linking two molecules or moieties, e.g., two domains of a fusion protein, such as, for example, a nuclease-inactive Cas9 domain and a nucleic acid editing domain (e.g., a recombinase domain). Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.Gene of interest (GOI)
[0355] The term “gene of interest” or “GOI” refers to a gene that encodes a biomolecule of interest (e.g., a protein or an RNA molecule). A protein of interest can include any intracellular protein, membrane protein, or extracellular protein, e.g., a nuclear protein, transcription factor, nuclear membrane transporter, intracellular organelle associated protein, a membrane receptor, a catalytic protein, and enzyme, a therapeutic protein, a membrane protein, a membrane transport protein, a signal transduction protein, or an immunological protein (e.g., an IgG or other antibody protein), etc. The gene of interest may also encode an RNA molecule, including, but not limited to, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), antisense RNA, guide RNA, microRNA (miRNA), small interfering RNA (siRNA), and cell-free RNA (cfRNA). Genetic Elements
[0356] Nucleic acids of the present disclosure may include one or more genetic elements. A “genetic element” refers to a particular nucleotide sequence that has a role in nucleic acid expression (e.g., promoter, enhancer, terminator) or encodes a discrete product of an engineered nucleic acid (e.g., a nucleotide sequence encoding a guide RNA, a protein and / or an RNA interference molecule).
[0357] A “promoter” refers to a control region of a nucleic acid sequence at which initiation and rate of transcription of the remainder of a nucleic acid sequence are controlled. A promoter may also contain sub-regions at which regulatory proteins and molecules may bind, such as RNA polymerase and other transcription factors. Promoters may be constitutive, inducible, activatable, repressible, tissue-specific, or any combination thereof. A promoter drives expression or drives transcription of the nucleic acid sequence that it regulates. Herein, a promoter is considered to be “operably linked” when it is in a correct functional location and orientation in relation to a nucleic acid sequence it regulates to control (“drive”) transcriptional initiation and / or expression of that sequence.
[0358] A promoter may be one naturally associated with a gene or sequence, as may be obtained by isolating the 5' non-coding sequences located upstream of the coding segment of a given gene or sequence. Such a promoter is referred to as an “endogenous promoter.” In some embodiments, a coding nucleic acid sequence may be positioned under the control of a recombinant or heterologous promoter, which refers to a promoter that is not normally associated with the encoded sequence in its natural environment. Such promoters may include promoters of other genes; promoters isolated from any other cell; and syntheticpromoters or enhancers that are not “naturally occurring” such as, for example, those that contain different elements of different transcriptional regulatory regions and / or mutations that alter expression through methods of genetic engineering that are known in the art. In addition to producing nucleic acid sequences of promoters and enhancers synthetically, sequences may be produced using recombinant cloning and / or nucleic acid amplification technology, including polymerase chain reaction (PCR).
[0359] In some embodiments, promoters used in accordance with the present disclosure are “inducible promoters,” which are promoters that are characterized by regulating (e.g., initiating or activating) transcriptional activity when in the presence of, influenced by or contacted by an inducer signal. An inducer signal may be endogenous or a normally exogenous condition (e.g., light), compound (e.g., chemical or non-chemical compound) or protein that contacts an inducible promoter in such a way as to be active in regulating transcriptional activity from the inducible promoter. Thus, a “signal that regulates transcription” of a nucleic acid refers to an inducer signal that acts on an inducible promoter. A signal that regulates transcription may activate or inactivate transcription, depending on the regulatory system used. Activation of transcription may involve directly acting on a promoter to drive transcription or indirectly acting on a promoter by inactivation a repressor that is preventing the promoter from driving transcription. Conversely, deactivation of transcription may involve directly acting on a promoter to prevent transcription or indirectly acting on a promoter by activating a repressor that then acts on the promoter.
[0360] A “transcriptional terminator” is a nucleic acid sequence that causes transcription to stop. A transcriptional terminator may be unidirectional or bidirectional. It is comprised of a DNA sequence involved in specific termination of an RNA transcript by an RNA polymerase. A transcriptional terminator sequence prevents transcriptional activation of downstream nucleic acid sequences by upstream promoters. A transcriptional terminator may be necessary in vivo to achieve desirable expression levels or to avoid transcription of certain sequences. A transcriptional terminator is considered to be “operably linked to” a nucleotide sequence when it is able to terminate the transcription of the sequence it is linked to.
[0361] The most commonly used type of terminator is a forward terminator. When placed downstream of a nucleic acid sequence that is usually transcribed, a forward transcriptional terminator will cause transcription to abort. In some embodiments, bidirectional transcriptional terminators are provided, which usually cause transcription to terminate on both the forward and reverse strand. In some embodiments, reversetranscriptional terminators are provided, which usually terminate transcription on the reverse strand only.
[0362] In prokaryotic systems, terminators usually fall into two categories (1) rho- independent terminators and (2) rho-dependent terminators. Rho-independent terminators are generally composed of palindromic sequence that forms a stem loop rich in G-C base pairs followed by several T bases. Without wishing to be bound by theory, the conventional model of transcriptional termination is that the stem loop causes RNA polymerase to pause, and transcription of the poly-A tail causes the RNA:DNA duplex to unwind and dissociate from RNA polymerase.
[0363] In eukaryotic systems, the terminator region may comprise specific DNA sequences that permit site-specific cleavage of the new transcript so as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3' end of the transcript. RNA molecules modified with this polyA tail appear to more stable and are translated more efficiently. Thus, in some embodiments involving eukaryotes, a terminator may comprise a signal for the cleavage of the RNA. In some embodiments, the terminator signal promotes polyadenylation of the message. The terminator and / or polyadenylation site elements may serve to enhance output nucleic acid levels and / or to minimize read through between nucleic acids.
[0364] Terminators for use in accordance with the present disclosure include any terminator of transcription described herein or known to one of ordinary skill in the art. Examples of terminators include, without limitation, the termination sequences of genes such as, for example, the bovine growth hormone terminator, and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA and arcA terminator. In some embodiments, the termination signal may be a sequence that cannot be transcribed or translated, such as those resulting from a sequence truncation.
[0365] A “Woodchuck Hepatitis Virus (WHP) Posttranscriptional Regulatory Element (WPRE)” is a DNA sequence that, when transcribed creates a tertiary structure enhancing expression. Commonly used in molecular biology to increase expression of genes delivered by viral vectors. WPRE is a tripartite regulatory element with gamma, alpha, and beta components. WPRE sequence listing is in the Description of the Sequences Section. Guide RNA (gRNA)
[0366] A gRNA is a component of the CRISPR / Cas system. A “gRNA” (guide ribonucleic acid) herein refers to a fusion of a CRISPR-targeting RNA (crRNA) and a trans- activation crRNA (tracrRNA), providing both targeting specificity and scaffolding / binding ability for Cas9 nuclease. A “crRNA” is a bacterial RNA that confers target specificity and requires tracrRNA to bind to Cas9. A “tracrRNA” is a bacterial RNA that links the crRNA to the Cas9 nuclease and typically can bind any crRNA. The sequence specificity of a Cas DNA-binding protein is determined by gRNAs, which have nucleotide base-pairing complementarity to target DNA sequences. The native gRNA comprises a 20 nucleotide (nt) Specificity Determining Sequence (SDS), which specifies the DNA sequence to be targeted, and is immediately followed by a 80 nt scaffold sequence, which associates the gRNA with Cas9. In some embodiments, an SDS of the present disclosure has a length of 15 to 100 nucleotides, or more. For example, an SDS may have a length of 15 to 90, 15 to 85, 15 to 80, 15 to 75, 15 to 70, 15 to 65, 15 to 60, 15 to 55, 15 to 50, 15 to 45, 15 to 40, 15 to 35, 15 to 30, or 15 to 20 nucleotides. In some embodiments, the SDS is 20 nucleotides long. For example, the SDS may be 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. At least a portion of the target DNA sequence is complementary to the SDS of the gRNA. For Cas9 to successfully bind to the DNA target sequence, a region of the target sequence is complementary to the SDS of the gRNA sequence and is immediately followed by the correct protospacer adjacent motif (PAM) sequence (e.g., NGG for Cas9 and TTN, TTTN, or YTN for Cpf1). In some embodiments, an SDS is 100% complementary to its target sequence. In some embodiments, the SDS sequence is less than 100% complementary to its target sequence and is, thus, considered to be partially complementary to its target sequence. For example, a targeting sequence may be 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% complementary to its target sequence. In some embodiments, the SDS of template DNA or target DNA may differ from a complementary region of a gRNA by 1, 2, 3, 4 or 5 nucleotides.
[0367] In addition to the SDS, the gRNA comprises a scaffold sequence (corresponding to the tracrRNA in the native CRISPR / Cas system) that is required for its association with Cas9 (referred to herein as the “gRNA handle”). In some embodiments, the gRNA comprises a structure 5′-[SDS] -[gRNA handle]-3′. In some embodiments, the scaffold sequence comprises the nucleotide sequence of 5′-guuuuagagcuagaaauagcaaguuaaaauaaaggcuaguc cguuaucaacuugaaaaaguggcaccgagucggugcuuuuu-3′ (SEQ ID NO: 216). Other non-limiting, suitable gRNA handle sequences that may be used in accordance with the present disclosure are listed in the Description of the Sequences.
[0368] In some embodiments, the guide RNA is about 15-120 nucleotides long and comprises a sequence of at least 10 contiguous nucleotides that is complementary to a target sequence. In some embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 nucleotides long. In some embodiments, the guide RNA comprises a sequence of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more contiguous nucleotides that is complementary to a target sequence. Sequence complementarity refers to distinct interactions between adenine and thymine (DNA) or uracil (RNA), and between guanine and cytosine. Homology arm
[0369] The term “homology arm” refers to a portion of the extension arm that encodes a portion of the resulting reverse transcriptase-encoded single strand DNA flap that is to be integrated into the target DNA site by replacing the endogenous strand. The portion of the single strand DNA flap encoded by the homology arm is complementary to the non-edited strand of the target DNA sequence, which facilitates the displacement of the endogenous strand and annealing of the single strand DNA flap in its place, thereby installing the edit. This component is further defined elsewhere. The homology arm is part of the DNA synthesis template since it is by definition encoded by the polymerase of the prime editors described herein. Host cell
[0370] The term “host cell,” as used herein, refers to a cell that can host, replicate, and express a vector described herein, e.g., a vector comprising a nucleic acid molecule encoding a fusion protein comprising a Cas9 or Cas9 equivalent and a reverse transcriptase. Intein
[0371] An “intein” is a segment of a protein that is able to excise itself and join the remaining portions (the exteins) with a peptide bond in a process known as protein splicing. Inteins are also referred to as “protein introns.” The process of an intein excising itself and joining the remaining portions of the protein is herein termed “protein splicing” or “intein-mediated protein splicing.” In some embodiments, an intein of a precursor protein (an intein containing protein prior to intein-mediated protein splicing) comes from two genes. Such intein is referred to herein as a split intein. For example, in cyanobacteria, DnaE, thecatalytic subunit α of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene is herein referred as “intein-N.” The intein encoded by the dnaE-c gene is herein referred as “intein-C.”
[0372] Other intein systems may also be used. For example, a synthetic intein based on the dnaE intein, the Cfa-N and Cfa-C intein pair, has been described (e.g., in Stevens et al., J Am Chem Soc.2016 Feb 24;138(7):2162-5, incorporated herein by reference). Non-limiting examples of intein pairs that may be used in accordance with the present disclosure include: Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein and Cne Prp8 intein (e.g., as described in US Patent 8,394,604, incorporated herein by reference.
[0373] Intein-N and intein-C may be fused to the N-terminal portion of the split prime editor and the C-terminal portion of the split prime editor, respectively, for the joining of the N-terminal portion of the split prime editor and the C-terminal portion of the split prime editor. For example, in some embodiments, an intein-N is fused to the C-terminus of the N- terminal portion of the split prime editor, i.e., to form a structure of N-[N-terminal portion of the split prime editor]-[intein-N]-C. In some embodiments, an intein-C is fused to the N- terminus of the C-terminal portion of the split prime editor, i.e., to form a structure of N- [intein-C]-[C-terminal portion of the split prime editor]-C. The mechanism of intein-mediated protein splicing for joining the proteins the inteins are fused to (e.g., split prime editor) is known in the art, e.g., as described in Shah et al., Chem Sci.2014; 5(1):446–461, incorporated herein by reference. Ligand-dependent intein
[0374] The term “ligand-dependent intein,” as used herein refers to an intein that comprises a ligand-binding domain. Typically, the ligand-binding domain is inserted into the amino acid sequence of the intein, resulting in a structure intein (N) – ligand-binding domain – intein (C). Typically, ligand-dependent inteins exhibit no or only minimal protein splicing activity in the absence of an appropriate ligand, and a marked increase of protein splicing activity in the presence of the ligand. In some embodiments, the ligand-dependent intein does not exhibit observable splicing activity in the absence of ligand but does exhibit splicing activity in the presence of the ligand. In some embodiments, the ligand-dependent intein exhibits an observable protein splicing activity in the absence of the ligand, and a protein splicing activity in the presence of an appropriate ligand that is at least 5 times, at least 10 times, at least 50 times, at least 100 times, at least 150 times, at least 200 times, at least 250times, at least 500 times, at least 1000 times, at least 1500 times, at least 2000 times, at least 2500 times, at least 5000 times, at least 10000 times, at least 20000 times, at least 25000 times, at least 50000 times, at least 100000 times, at least 500000 times, or at least 1000000 times greater than the activity observed in the absence of the ligand. In some embodiments, the increase in activity is dose dependent over at least 1 order of magnitude, at least 2 orders of magnitude, at least 3 orders of magnitude, at least 4 orders of magnitude, or at least 5 orders of magnitude, allowing for fine-tuning of intein activity by adjusting the concentration of the ligand. Suitable ligand-dependent inteins are known in the art, and in include those provided below and those described in published U.S. Patent Application U.S.2014 / 0065711 A1; Mootz et al., “Protein splicing triggered by a small molecule.” J. Am. Chem. Soc.2002; 124, 9044–9045; Mootz et al., “Conditional protein splicing: a new tool to control protein structure and function in vitro and in vivo.” J. Am. Chem. Soc.2003; 125, 10561–10569; Buskirk et al., Proc. Natl. Acad. Sci. USA.2004; 101, 10505-10510); Skretas & Wood, “Regulation of protein activity with small-molecule-controlled inteins.” Protein Sci.2005; 14, 523-532; Schwartz, et al., “Post-translational enzyme activation in an animal via optimized conditional protein splicing.” Nat. Chem. Biol.2007; 3, 50-54; Peck et al., Chem. Biol.2011; 18 (5), 619-630; the entire contents of each are hereby incorporated by reference. Exemplary sequences are listed in the Description of the Sequences section. Linker
[0375] The term “linker,” as used herein, refers to a molecule linking two other molecules or moieties. The linker can be an amino acid sequence in the case of a linker joining two fusion proteins. For example, a Cas9 can be fused to a reverse transcriptase by an amino acid linker sequence. The linker can also be a nucleotide sequence in the case of joining two nucleotide sequences together. For example, in the instant case, the traditional guide RNA is linked via a spacer or linker nucleotide sequence to the RNA extension of a prime editing guide RNA which may comprise a RT template sequence and an RT primer binding site. In other embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. napDNAbp
[0376] As used herein, the term “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which Cas9 is an example, refer to a proteins which use RNA:DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic-acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to a complementary sequence.
[0377] Without being bound by theory, the binding mechanism of a napDNAbp – guide RNA complex, in general, includes the step of forming an R-loop whereby the napDNAbp induces the unwinding of a double-strand DNA target, thereby separating the strands in the region bound by the napDNAbp. The guide RNA protospacer then hybridizes to the “target strand.” This displaces a “non-target strand” that is complementary to the target strand, which forms the single strand region of the R-loop. In some embodiments, the napDNAbp includes one or more nuclease activities, which then cut the DNA leaving various types of lesions. For example, the napDNAbp may comprises a nuclease activity that cuts the non- target strand at a first location, and / or cuts the target strand at a second location. Depending on the nuclease activity, the target DNA can be cut to form a “double-stranded break” whereby both strands are cut. In other embodiments, the target DNA can be cut at only a single site, i.e., the DNA is “nicked” on one strand. Exemplary napDNAbp with different nuclease activities include “Cas9 nickase” (“nCas9”) and a deactivated Cas9 having no nuclease activities (“dead Cas9” or “dCas9”). Exemplary sequences for these and other napDNAbp are provided herein. Nickase
[0378] The term “nickase” refers to a Cas9 with one of the two nuclease domains inactivated. This enzyme is capable of cleaving only one strand of a target DNA. Nucleic Acid
[0379] The terms “nucleic acid,” and “polynucleotide,” as used herein, refer to a compound comprising a nucleobase and an acidic moiety, e.g., a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides are linear molecules, in which adjacent nucleotides are linked to each other via a phosphodiester linkage. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g. nucleotides and / or nucleosides). In some embodiments,“nucleic acid” refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms “oligonucleotide” and “polynucleotide” can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, “nucleic acid” encompasses RNA as well as single and / or double-stranded DNA. Nucleic acids may be naturally occurring, for example, in the context of a genome, a transcript, an mRNA, tRNA, rRNA, siRNA, snRNA, a plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule may be a non-naturally occurring molecule, e.g., a recombinant DNA or RNA, an artificial chromosome, an engineered genome (e.g., an engineered viral vector), an engineered vector, or fragment thereof, or a synthetic DNA, RNA, or DNA / RNA hybrid, optionally including non-naturally occurring nucleotides or nucleosides. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can comprise nucleoside analogs such as analogs having chemically modified bases or sugars, and backbone modifications. A nucleic acid sequence is presented in the 5′ to 3′ direction unless otherwise indicated. In some embodiments, a nucleic acid is or comprises natural nucleosides (e.g. adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, 5-methylcytidine, 2-aminoadenosine, C5- bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8- oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5′-N-phosphoramidite linkages). Nuclear Localization Signal (NLS)
[0380] A “nuclear localization signal” or “NLS” refers to as an amino acid sequence that “tags” a protein for import into the cell nucleus by nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface. One or more NLS may be added to the N- or C-terminus of aprotein, or internally (e.g., between two protein domains). For example, one or more NLS may be added to the N- or C-terminus of a prime editor, or between the napDNAbp and the polymerase in a prime editor. In some embodiments, 1, 2, 3, 4, 5, or more NLS may be added. Nuclear localization sequences are known in the art and would be apparent to the skilled artisan. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, filed November 23, 2000, the contents of which are incorporated herein by reference for its disclosure of exemplary nuclear localization sequences.
[0381] The nucleotide sequence encoding an NLS is “operably linked” to the nucleotide sequence encoding a protein to which the NLS is fused (e.g., a prime editor) when two coding sequences are “in-frame with each other” and are translated as a single polypeptide fusing two sequences. PegRNAs
[0382] As used herein, the terms “prime editing guide RNA” or “PEgRNA” or “extended guide RNA” refers to a specialized form of a guide RNA that has been modified to include one or more additional sequences for implementing the prime editing methods and compositions described herein. As described herein, the prime editing guide RNA comprise one or more “extended regions” of nucleic acid sequence. The extended regions may comprise, but are not limited to, single-stranded RNA or DNA. Further, the extended regions may occur at the 3´ end of a traditional guide RNA. In other arrangements, the extended regions may occur at the 5´ end of a traditional guide RNA. In still other arrangements, the extended region may occur at an intramolecular region of the traditional guide RNA, for example, in the gRNA core region which associates and / or binds to the napDNAbp. The extended region comprises a “DNA synthesis template” which encodes (by the polymerase of the prime editor) a single-stranded DNA which, in turn, has been designed to be (a) homologous with the endogenous target DNA to be edited, and (b) which comprises at least one desired nucleotide change (e.g., a transition, a transversion, a deletion, or an insertion) to be introduced or integrated into the endogenous target DNA. The extended region may also comprise other functional sequence elements, such as, but not limited to, a “primer binding site” and a “spacer or linker” sequence, or other structural elements, such as, but not limited to aptamers, stem loops, hairpins, toe loops (e.g., a 3’ toeloop), or an RNA- protein recruitment domain (e.g., MS2 hairpin). As used herein the “primer binding site” comprises a sequence that hybridizes to a single-strand DNA sequence having a 3´ end generated from the nicked DNA of the R-loop.
[0383] In certain embodiments, the PEgRNAs are represented by a 5ʹ extension arm, a spacer, and a gRNA core. The 5ʹ extension further comprises in the 5ʹ to 3ʹ direction a reverse transcriptase template, a primer binding site, and a linker. As shown, the reverse transcriptase template may also be referred to more broadly as the “DNA synthesis template” where the polymerase of a prime editor described herein is not an RT, but another type of polymerase.
[0384] In certain other embodiments, the PEgRNAs are represented by a 5ʹ extension arm, a spacer, and a gRNA core. The 5ʹ extension further comprises in the 5ʹ to 3ʹ direction a reverse transcriptase template, a primer binding site, and a linker. As shown, the reverse transcriptase template may also be referred to more broadly as the “DNA synthesis template” where the polymerase of a prime editor described herein is not an RT, but another type of polymerase.
[0385] In still other embodiments, the PEgRNAs are represented having in the 5ʹ to 3ʹ direction a spacer (1), a gRNA core (2), and an extension arm (3). The extension arm (3) is at the 3ʹ end of the PEgRNA. The extension arm (3) further comprises in the 5ʹ to 3ʹ direction a “primer binding site” (A), an “edit template” (B), and a “homology arm” (C). The extension arm (3) may also comprise an optional modifier region at the 3ʹ and 5ʹ ends, which may be the same sequences or different sequences. In addition, the 3ʹ end of the PEgRNA may comprise a transcriptional terminator sequence. These sequence elements of the PEgRNAs are further described and defined herein.
[0386] In still other embodiments, the PEgRNAs are represented by having in the 5ʹ to 3ʹ direction an extension arm (3), a spacer (1), and a gRNA core (2). The extension arm (3) is at the 5ʹ end of the PEgRNA. The extension arm (3) further comprises in the 3ʹ to 5ʹ direction a “primer binding site” (A), an “edit template” (B), and a “homology arm” (C). The extension arm (3) may also comprise an optional modifier region at the 3ʹ and 5ʹ ends, which may be the same sequences or different sequences. The PEgRNAs may also comprise a transcriptional terminator sequence at the 3ʹ end. These sequence elements of the PEgRNAs are further described and defined herein. PE1
[0387] As used herein, “PE1” refers to a PE complex comprising a fusion protein comprising Cas9(H840A) and a wild type MMLV RT having the following structure: [NLS]- [Cas9(H840A)]-[linker]-[MMLV_RT(wt)] + a desired PEgRNA, wherein the PE fusion has the amino acid sequence of SEQ ID NO: 125 (See Description of Sequences)PE2
[0388] As used herein, “PE2” refers to a PE complex comprising a fusion protein comprising Cas9(H840A) and a variant MMLV RT having the following structure: [NLS]- [Cas9(H840A)]-[linker]-[MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)] + a desired PEgRNA, wherein the PE fusion has the amino acid sequence of SEQ ID NO: 129 (See Description of Sequences) PE3
[0389] As used herein, “PE3” refers to PE2 plus a second-strand nicking guide RNA that complexes with the PE2 and introduces a nick in the non-edited DNA strand in order to induce preferential replacement of the edited strand. PE3b
[0390] As used herein, “PE3b” refers to PE3 but wherein the second-strand nicking guide RNA is designed for temporal control such that the second strand nick is not introduced until after the installation of the desired edit. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand, but not the original allele. Using this strategy, referred to hereafter as PE3b, mismatches between the protospacer and the unedited allele should disfavor nicking by the sgRNA until after the editing event on the PAM strand takes place. PE-short
[0391] As used herein, “PE-short” refers to a PE construct that is fused to a C-terminally truncated reverse transcriptase, and has the following SEQ ID NO: 130 (See Description of Sequences) Peptide tag
[0392] The term “peptide tag” refers to a peptide amino acid sequence that is genetically fused to a protein sequence to impart one or more functions onto the proteins that facilitate the manipulation of the protein for various purposes, such as, visualization, purification, solubilization, and separation, etc. Peptide tags can include various types of tags categorized by purpose or function, which may include “affinity tags” (to facilitate protein purification), “solubilization tags” (to assist in proper folding of proteins), “chromatography tags” (to alter chromatographic properties of proteins), “epitope tags” (to bind to high affinity antibodies), “fluorescence tags” (to facilitate visualization of proteins in a cell or in vitro). Polymerase
[0393] As used herein, the term “polymerase” refers to an enzyme that synthesizes a nucleotide strand and which may be used in connection with the prime editor systems described herein. The polymerase can be a “template-dependent” polymerase (i.e., a polymerase which synthesizes a nucleotide strand based on the order of nucleotide bases of a template strand). The polymerase can also be a “template-independent” polymerase (i.e., a polymerase which synthesizes a nucleotide strand without the requirement of a template strand). A polymerase may also be further categorized as a “DNA polymerase” or an “RNA polymerase.” In various embodiments, the prime editor system comprises a DNA polymerase. In various embodiments, the DNA polymerase can be a “DNA-dependent DNA polymerase” (i.e., whereby the template molecule is a strand of DNA). In such cases, the DNA template molecule can be a PEgRNA, wherein the extension arm comprises a strand of DNA. In such cases, the PEgRNA may be referred to as a chimeric or hybrid PEgRNA which comprises an RNA portion (i.e., the guide RNA components, including the spacer and the gRNA core) and a DNA portion (i.e., the extension arm). In various other embodiments, the DNA polymerase can be an “RNA-dependent DNA polymerase” (i.e., whereby the template molecule is a strand of RNA). In such cases, the PEgRNA is RNA, i.e., including an RNA extension. The term “polymerase” may also refer to an enzyme that catalyzes the polymerization of nucleotide (i.e., the polymerase activity). Generally, the enzyme will initiate synthesis at the 3'-end of a primer annealed to a polynucleotide template sequence (e.g., such as a primer sequence annealed to the primer binding site of a PEgRNA), and will proceed toward the 5' end of the template strand. A “DNA polymerase” catalyzes the polymerization of deoxynucleotides. As used herein in reference to a DNA polymerase, the term DNA polymerase includes a “functional fragment thereof”. A “functional fragment thereof” refers to any portion of a wild-type or mutant DNA polymerase that encompasses less than the entire amino acid sequence of the polymerase and which retains the ability, under at least one set of conditions, to catalyze the polymerization of a polynucleotide. Such a functional fragment may exist as a separate entity, or it may be a constituent of a larger polypeptide, such as a fusion protein. Pharmaceutically acceptable Carrier
[0394] The term “pharmaceutically-acceptable carrier” means a pharmaceutically- acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc magnesium, calcium or zinc stearate, or steric acid), or solvent encapsulating material, involved in carrying or transporting the compound from onesite (e.g., the delivery site) of the body, to another site (e.g., organ, tissue or portion of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the tissue of the subject (e.g., physiologically compatible, sterile, physiologic pH, etc.). prime editing
[0395] As used herein, the term “prime editing” refers to a novel approach for gene editing using napDNAbps, a polymerase (e.g., a reverse transcriptase), and specialized guide RNAs that include a DNA synthesis template for encoding desired new genetic information (or deleting genetic information) that is then incorporated into a target DNA sequence.
[0396] prime editing represents an entirely new platform for genome editing that is a versatile and precise genome editing method that directly writes new genetic information into a specified DNA site using a nucleic acid programmable DNA binding protein (“napDNAbp”) working in association with a polymerase (i.e., in the form of a fusion protein or otherwise provided in trans with the napDNAbp), wherein the prime editing system is programmed with a prime editing (PE) guide RNA (“PEgRNA”) that both specifies the target site and templates the synthesis of the desired edit in the form of a replacement DNA strand by way of an extension (either DNA or RNA) engineered onto a guide RNA (e.g., at the 5ʹ or 3ʹ end, or at an internal portion of a guide RNA). The replacement strand containing the desired edit (e.g., a single nucleobase substitution) shares the same (or is homologous to) sequence as the endogenous strand (immediately downstream of the nick site) of the target site to be edited (with the exception that it includes the desired edit). Through DNA repair and / or replication machinery, the endogenous strand downstream of the nick site is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, prime editing may be thought of as a “search-and-replace” genome editing technology since the prime editors, as described herein, not only search and locate the desired target site to be edited, but at the same time, encode a replacement strand containing a desired edit which is installed in place of the corresponding target site endogenous DNA strand. The prime editors of the present disclosure relate, in part, to the discovery that the mechanism of target-primed reverse transcription (TPRT) or “prime editing” can be leveraged or adapted for conducting precision CRISPR / Cas-based genome editing with high efficiency and genetic flexibility. TPRT is naturally used by mobile DNA elements, such as mammalian non-LTR retrotransposons and bacterial Group II introns28,29. The inventors have herein used Cas protein-reverse transcriptase fusions or related systems to target a specific DNA sequencewith a guide RNA, generate a single strand nick at the target site, and use the nicked DNA as a primer for reverse transcription of an engineered reverse transcriptase template that is integrated with the guide RNA. However, while the concept begins with prime editors that use reverse transcriptase as the DNA polymerase component, the prime editors described herein are not limited to reverse transcriptases but may include the use of virtually any DNA polymerase. Indeed, while the application throughout may refer to prime editors with “reverse transcriptases,” it is set forth here that reverse transcriptases are only one type of DNA polymerase that may work with prime editing. Thus, where ever the specification mentions a “reverse transcriptase,” the person having ordinary skill in the art should appreciate that any suitable DNA polymerase may be used in place of the reverse transcriptase. Thus, in one aspect, the prime editors may comprise Cas9 (or an equivalent napDNAbp) which is programmed to target a DNA sequence by associating it with a specialized guide RNA (i.e., PEgRNA) containing a spacer sequence that anneals to a complementary protospacer in the target DNA. The specialized guide RNA also contains new genetic information in the form of an extension that encodes a replacement strand of DNA containing a desired genetic alteration which is used to replace a corresponding endogenous DNA strand at the target site. To transfer information from the PEgRNA to the target DNA, the mechanism of prime editing involves nicking the target site in one strand of the DNA to expose a 3′-hydroxyl group. The exposed 3′-hydroxyl group can then be used to prime the DNA polymerization of the edit-encoding extension on PEgRNA directly into the target site. In various embodiments, the extension—which provides the template for polymerization of the replacement strand containing the edit—can be formed from RNA or DNA. In the case of an RNA extension, the polymerase of the prime editor can be an RNA-dependent DNA polymerase (such as, a reverse transcriptase). In the case of a DNA extension, the polymerase of the prime editor may be a DNA-dependent DNA polymerase. The newly synthesized strand (i.e., the replacement DNA strand containing the desired edit) that is formed by the herein disclosed prime editors would be homologous to the genomic target sequence (i.e., have the same sequence as) except for the inclusion of a desired nucleotide change (e.g., a single nucleotide change, a deletion, or an insertion, or a combination thereof). The newly synthesized (or replacement) strand of DNA may also be referred to as a single strand DNA flap, which would compete for hybridization with the complementary homologous endogenous DNA strand, thereby displacing the corresponding endogenous strand. In certain embodiments, the system can be combined with the use of an error-prone reverse transcriptase enzyme (e.g., provided as a fusion protein with the Cas9 domain, orprovided in trans to the Cas9 domain). The error-prone reverse transcriptase enzyme can introduce alterations during synthesis of the single strand DNA flap. Thus, in certain embodiments, error-prone reverse transcriptase can be utilized to introduce nucleotide changes to the target DNA. Depending on the error-prone reverse transcriptase that is used with the system, the changes can be random or non-random. Resolution of the hybridized intermediate (comprising the single strand DNA flap synthesized by the reverse transcriptase hybridized to the endogenous DNA strand) can include removal of the resulting displaced flap of endogenous DNA (e.g., with a 5ʹ end DNA flap endonuclease, FEN1), ligation of the synthesized single strand DNA flap to the target DNA, and assimilation of the desired nucleotide change as a result of cellular DNA repair and / or replication processes. Because templated DNA synthesis offers single nucleotide precision for the modification of any nucleotide, including insertions and deletions, the scope of this approach is very broad and could foreseeably be used for myriad applications in basic science and therapeutics.
[0397] In various embodiments, prime editing operates by contacting a target DNA molecule (for which a change in the nucleotide sequence is desired to be introduced) with a nucleic acid programmable DNA binding protein (napDNAbp) complexed with a prime editing guide RNA (PEgRNA). In some embodiments, the prime editing guide RNA (PEgRNA) comprises an extension at the 3´ or 5´ end of the guide RNA, or at an intramolecular location in the guide RNA and encodes the desired nucleotide change (e.g., single nucleotide change, insertion, or deletion). In step (a), the napDNAbp / extended gRNA complex contacts the DNA molecule and the extended gRNA guides the napDNAbp to bind to a target locus. In step (b), a nick in one of the strands of DNA of the target locus is introduced (e.g., by a nuclease or chemical agent), thereby creating an available 3´ end in one of the strands of the target locus. In certain embodiments, the nick is created in the strand of DNA that corresponds to the R-loop strand, i.e., the strand that is not hybridized to the guide RNA sequence, i.e., the “non-target strand.” The nick, however, could be introduced in either of the strands. That is, the nick could be introduced into the R-loop “target strand” (i.e., the strand hybridized to the protospacer of the extended gRNA) or the “non-target strand” (i.e., the strand forming the single-stranded portion of the R-loop and which is complementary to the target strand). In step (c), the 3´ end of the DNA strand (formed by the nick) interacts with the extended portion of the guide RNA in order to prime reverse transcription (i.e., “target-primed RT”). In certain embodiments, the 3´ end DNA strand hybridizes to a specific RT priming sequence on the extended portion of the guide RNA, i.e., the “reverse transcriptase priming sequence” or “primer binding site” on the PEgRNA. Instep (d), a reverse transcriptase (or other suitable DNA polymerase) is introduced which synthesizes a single strand of DNA from the 3´ end of the primed site towards the 5´ end of the prime editing guide RNA. The DNA polymerase (e.g., reverse transcriptase) can be fused to the napDNAbp or alternatively can be provided in trans to the napDNAbp. This forms a single-strand DNA flap comprising the desired nucleotide change (e.g., the single base change, insertion, or deletion, or a combination thereof) and which is otherwise homologous to the endogenous DNA at or adjacent to the nick site. In step (e), the napDNAbp and guide RNA are released. Steps (f) and (g) relate to the resolution of the single strand DNA flap such that the desired nucleotide change becomes incorporated into the target locus. This process can be driven towards the desired product formation by removing the corresponding 5´ endogenous DNA flap that forms once the 3´ single strand DNA flap invades and hybridizes to the endogenous DNA sequence. Without being bound by theory, the cells endogenous DNA repair and replication processes resolves the mismatched DNA to incorporate the nucleotide change(s) to form the desired altered product. The process can also be driven towards product formation with “second strand nicking” as is known in the art. This process may introduce at least one or more of the following genetic changes: transversions, transitions, deletions, and insertions. prime editor
[0398] The term “prime editor” refers to the herein described fusion constructs comprising a napDNAbp (e.g., Cas9 nickase) and a reverse transcriptase and is capable of carrying out prime editing on a target nucleotide sequence in the presence of a PEgRNA (or “extended guide RNA”). The term “prime editor” may refer to the fusion protein or to the fusion protein complexed with a PEgRNA, and / or further complexed with a second-strand nicking sgRNA. In some embodiments, the prime editor may also refer to the complex comprising a fusion protein (reverse transcriptase fused to a napDNAbp), a PEgRNA, and a regular guide RNA capable of directing the second-site nicking step of the non-edited strand as described herein. In other embodiments, the reverse transcriptase component of the “primer editor” may be provided in trans. primer binding site
[0399] The term “primer binding site” or “the PBS” refers to the nucleotide sequence located on a PEgRNA as component of the extension arm (typically at the 3ʹ end of the extension arm) and serves to bind to the primer sequence that is formed after Cas9 nicking of the target sequence by the prime editor. As detailed elsewhere, when the Cas9 nickasecomponent of a prime editor nicks one strand of the target DNA sequence, a 3ʹ-ended ssDNA flap is formed, which serves a primer sequence that anneals to the primer binding site on the PEgRNA to prime reverse transcription. Protein, Peptide, and Polypeptides
[0400] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein, and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofarnesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide may also be a single molecule or may be a multi-molecular complex. A protein, peptide, or polypeptide may be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. The term “fusion protein” as used herein refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy- terminal (C-terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy- terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic domain of a nucleic-acid editing protein. In some embodiments, a protein is in a complex with, or is in association with, a nucleic acid, e.g., RNA or DNA. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), which are incorporated herein by reference. Protospacer
[0401] As used herein, the term “protospacer” refers to the sequence (~20 bp) in DNA adjacent to the PAM (protospacer adjacent motif) sequence. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand thereof, i.e., the “target strand” versus the “non-target strand” of the target DNA sequence). In order for Cas9 to function it also requires a specific protospacer adjacent motif (PAM) that varies depending on the bacterial species of the Cas9 gene. The most commonly used Cas9 nuclease, derived from S. pyogenes, recognizes a PAM sequence of NGG that is found directly downstream of the target sequence in the genomic DNA, on the non-target strand. The skilled person will appreciate that the literature in the state of the art sometimes refers to the “protospacer” as the ~20-nt target-specific guide sequence on the guide RNA itself, rather than referring to it as a “spacer.” Thus, in some cases, the term “protospacer” as used herein may be used interchangeably with the term “spacer.” The context of the description surrounding the appearance of either “protospacer” or “spacer” will help inform the reader as to whether the term is in reference to the gRNA or the DNA target. Protospacer adjacent motif (PAM)
[0402] A “protospacer adjacent motif” (PAM) is typically a sequence of nucleotides located adjacent to (e.g., within 10, 9, 8, 7, 6, 5, 4, 3, 3, or 1 nucleotide(s) of a target sequence). A PAM sequence is “immediately adjacent to” a target sequence if the PAM sequence is contiguous with the target sequence (that is, if there are no nucleotides located between the PAM sequence and the target sequence). In some embodiments, a PAM sequence is a wild-type PAM sequence. Examples of PAM sequences include, without limitation, NGG, NGR, NNGRR(T / N), NNNNGATT, NNAGAAW, NGGAG, NAAAAC, AWG, and CC. In some embodiments, a PAM sequence is obtained from Streptococcus pyogenes (e.g., NGG or NGR). In some embodiments, a PAM sequence is obtained from Staphylococcus aureus (e.g., NNGRR(T / N)). In some embodiments, a PAM sequence is obtained from Neisseria meningitidis (e.g., NNNNGATT). In some embodiments, a PAM sequence is obtained from Streptococcus thermophilus (e.g., NNAGAAW or NGGAG). In some embodiments, a PAM sequence is obtained from Treponema denticola (e.g., NAAAAC). In some embodiments, a PAM sequence is obtained from Escherichia coli (e.g., AWG). In some embodiments, a PAM sequence is obtained from Pseudomonas auruginosa (e.g., CC). Other PAM sequences are contemplated. A PAM sequence is typically locateddownstream (i.e., 3′) from the target sequence, although in some embodiments a PAM sequence may be located upstream (i.e., 5′) from the target sequence. Protein splicing
[0403] The term “protein splicing,” as used herein, refers to a process in which a sequence, an intein (or split inteins, as the case may be), is excised from within an amino acid sequence, and the remaining fragments of the amino acid sequence, the exteins, are ligated via an amide bond to form a continuous amino acid sequence. The term “trans” protein splicing refers to the specific case where the inteins are split inteins and they are located on different proteins. Recombinase
[0404] The term “recombinase,” as used herein, refers to a site-specific enzyme that mediates the recombination of DNA between recombinase recognition sequences, which results in the excision, integration, inversion, or exchange (e.g., translocation) of DNA fragments between the recombinase recognition sequences. Recombinases can be classified into two distinct families: serine recombinases (e.g., resolvases and invertases) and tyrosine recombinases (e.g., integrases). Examples of serine recombinases include, without limitation, Hin, Gin, Tn3, β-six, CinH, ParA, γδ, Bxb1, ϕC31, TP901, TG1, φBT1, R4, φRV1, φFC1, MR11, A118, U153, and gp29. Examples of tyrosine recombinases include, without limitation, Cre, FLP, R, Lambda, HK101, HK022, and pSAM2. The serine and tyrosine recombinase names stem from the conserved nucleophilic amino acid residue that the recombinase uses to attack the DNA and which becomes covalently linked to the DNA during strand exchange. Recombinases have numerous applications, including the creation of gene knockouts / knock-ins and gene therapy applications. See, e.g., Brown et al., “Serine recombinases as tools for genome engineering.” Methods.2011;53(4):372-9; Hirano et al., “Site-specific recombinases as tools for heterologous gene integration.” Appl. Microbiol. Biotechnol.2011; 92(2):227-39; Chavez and Calos, “Therapeutic applications of the ΦC31 integrase system.” Curr. Gene Ther.2011;11(5):375-81; Turan and Bode, “Site-specific recombinases: from tag-and-target- to tag-and-exchange-based genomic modifications.” FASEB J.2011; 25(12):4088-107; Venken and Bellen, “Genome-wide manipulations of Drosophila melanogaster with transposons, Flp recombinase, and ΦC31 integrase.” Methods Mol. Biol.2012; 859:203-28; Murphy, “Phage recombinases and their applications.” Adv. Virus Res.2012; 83:367-414; Zhang et al., “Conditional gene manipulation: Cre-ating a new biological era.” J. Zhejiang Univ. Sci. B.2012; 13(7):511-24; Karpenshif and Bernstein,“From yeast to mammals: recent advances in genetic control of homologous recombination.” DNA Repair (Amst).2012; 1;11(10):781-8; the entire contents of each are hereby incorporated by reference in their entirety. The recombinases provided herein are not meant to be exclusive examples of recombinases that can be used in embodiments of the invention. The methods and compositions of the invention can be expanded by mining databases for new orthogonal recombinases or designing synthetic recombinases with defined DNA specificities (See, e.g., Groth et al., “Phage integrases: biology and applications.” J. Mol. Biol.2004; 335, 667-678; Gordley et al., “Synthesis of programmable integrases.” Proc. Natl. Acad. Sci. U S A.2009; 106, 5053-5058; the entire contents of each are hereby incorporated by reference in their entirety). Other examples of recombinases that are useful in the methods and compositions described herein are known to those of skill in the art, and any new recombinase that is discovered or generated is expected to be able to be used in the different embodiments of the invention. In some embodiments, the catalytic domains of a recombinase are fused to a nuclease-inactivated RNA-programmable nuclease (e.g., dCas9, or a fragment thereof), such that the recombinase domain does not comprise a nucleic acid binding domain or is unable to bind to a target nucleic acid (e.g., the recombinase domain is engineered such that it does not have specific DNA binding activity). Recombinases lacking DNA binding activity and methods for engineering such are known, and include those described by Klippel et al., “Isolation and characterisation of unusual gin mutants.” EMBO J.1988; 7: 3983–3989: Burke et al., “Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation. Mol Microbiol.2004; 51: 937–948; Olorunniji et al., “Synapsis and catalysis by activated Tn3 resolvase mutants.” Nucleic Acids Res.2008; 36: 7181–7191; Rowland et al., “Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome.” Mol Microbiol.2009; 74: 282–298; Akopian et al., “Chimeric recombinases with designed DNA sequence recognition.” Proc Natl Acad Sci USA.2003;100: 8688–8691; Gordley et al., “Evolution of programmable zinc finger-recombinases with activity in human cells. J Mol Biol.2007; 367: 802–813; Gordley et al., “Synthesis of programmable integrases.” Proc Natl Acad Sci USA. 2009;106: 5053–5058; Arnold et al., “Mutants of Tn3 resolvase which do not require accessory binding sites for recombination activity.” EMBO J.1999;18: 1407–1414; Gaj et al., “Structure-guided reprogramming of serine recombinase DNA sequence specificity.” Proc Natl Acad Sci USA.2011;108(2):498-503; and Proudfoot et al., “Zinc finger recombinases with adaptable DNA sequence specificity.” PLoS One.2011;6(4):e19537; the entire contents of each are hereby incorporated by reference. For example, serinerecombinases of the resolvase-invertase group, e.g., Tn3 and γδ resolvases and the Hin and Gin invertases, have modular structures with autonomous catalytic and DNA-binding domains (See, e.g., Grindley et al., “Mechanism of site-specific recombination.” Ann Rev Biochem.2006; 75: 567–605, the entire contents of which are incorporated by reference). The catalytic domains of these recombinases are thus amenable to being recombined with nuclease-inactivated RNA-programmable nucleases (e.g., dCas9, or a fragment thereof) as described herein, e.g., following the isolation of ‘activated’ recombinase mutants which do not require any accessory factors (e.g., DNA binding activities) (See, e.g., Klippel et al., “Isolation and characterisation of unusual gin mutants.” EMBO J.1988; 7: 3983–3989: Burke et al., “Activating mutations of Tn3 resolvase marking interfaces important in recombination catalysis and its regulation. Mol Microbiol.2004; 51: 937–948; Olorunniji et al., “Synapsis and catalysis by activated Tn3 resolvase mutants.” Nucleic Acids Res.2008; 36: 7181–7191; Rowland et al., “Regulatory mutations in Sin recombinase support a structure-based model of the synaptosome.” Mol Microbiol.2009; 74: 282–298; Akopian et al., “Chimeric recombinases with designed DNA sequence recognition.” Proc Natl Acad Sci USA.2003;100: 8688–8691). Additionally, many other natural serine recombinases having an N-terminal catalytic domain and a C-terminal DNA binding domain are known (e.g., phiC31 integrase, TnpX transposase, IS607 transposase), and their catalytic domains can be co-opted to engineer programmable site-specific recombinases as described herein (See, e.g., Smith et al., “Diversity in the serine recombinases.” Mol Microbiol.2002;44: 299–307, the entire contents of which are incorporated by reference). Similarly, the core catalytic domains of tyrosine recombinases (e.g., Cre, λ integrase) are known, and can be similarly co-opted to engineer programmable site-specific recombinases as described herein (See, e.g., Guo et al., “Structure of Cre recombinase complexed with DNA in a site-specific recombination synapse.” Nature.1997; 389:40–46; Hartung et al., “Cre mutants with altered DNA binding properties.” J Biol Chem 1998; 273:22884–22891; Shaikh et al., “Chimeras of the Flp and Cre recombinases: Tests of the mode of cleavage by Flp and Cre. J Mol Biol.2000; 302:27– 48; Rongrong et al., “Effect of deletion mutation on the recombination activity of Cre recombinase.” Acta Biochim Pol.2005; 52:541–544; Kilbride et al., “Determinants of product topology in a hybrid Cre-Tn3 resolvase site-specific recombination system.” J Mol Biol.2006; 355:185–195; Warren et al., “A chimeric cre recombinase with regulated directionality.” Proc Natl Acad Sci USA.2008105:18278–18283; Van Duyne, “Teaching Cre to follow directions.” Proc Natl Acad Sci USA.2009 Jan 6;106(1):4-5; Numrych et al., “A comparison of the effects of single-base and triple-base changes in the integrase arm-typebinding sites on the site-specific recombination of bacteriophage λ.” Nucleic Acids Res. 1990; 18:3953–3959; Tirumalai et al., “The recognition of core-type DNA sites by λ integrase.” J Mol Biol.1998; 279:513–527; Aihara et al., “A conformational switch controls the DNA cleavage activity of λ integrase.” Mol Cell.2003; 12:187–198; Biswas et al., “A structural basis for allosteric control of DNA recombination by λ integrase.” Nature.2005; 435:1059–1066; and Warren et al., “Mutations in the amino-terminal domain of λ-integrase have differential effects on integrative and excisive recombination.” Mol Microbiol.2005; 55:1104–1112; the entire contents of each are incorporated by reference). Recombinase recognition sequence
[0405] The term “recombinase recognition sequence”, or equivalently as “RRS” or “recombinase target sequence”, as used herein, refers to a nucleotide sequence target recognized by a recombinase and which undergoes strand exchange with another DNA molecule having a the RRS that results in excision, integration, inversion, or exchange of DNA fragments between the recombinase recognition sequences. Recombine or recombination
[0406] The term “recombine,” or “recombination,” in the context of a nucleic acid modification (e.g., a genomic modification), is used to refer to the process by which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein (e.g., an inventive recombinase fusion protein provided herein). Recombination can result in, inter alia, the insertion, inversion, excision, or translocation of nucleic acids, e.g., in or between one or more nucleic acid molecules. recombinase recognition sequences Reverse transcriptase
[0407] The term "reverse transcriptase" describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize a DNA transcript from an RNA template. Historically, reverse transcriptase has been used primarily to transcribe mRNA into cDNA which can then be cloned into a vector for further manipulation. Avian myoblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme has 5ʹ-3ʹ RNA-directed DNA polymerase activity, 5ʹ-3ʹ DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5ʹ and 3ʹ ribonuclease specific for the RNA strand for RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Errors in transcription cannot becorrected by reverse transcriptase because known viral reverse transcriptases lack the 3ʹ-5ʹ exonuclease activity necessary for proofreading (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNase H activity has been presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase which is used extensively in molecular biology is reverse transcriptase originating from Moloney murine leukemia virus (M-MLV). See, e.g., Gerard, G. R., DNA 5:271-279 (1986) and Kotewicz, M. L., et al., Gene 35:249-258 (1985). M-MLV reverse transcriptase substantially lacking in RNase H activity has also been described. See, e.g., U.S. Pat. No.5,244,797. The invention contemplates the use of any such reverse transcriptases, or variants or mutants thereof.
[0408] In addition, the invention contemplates the use of reverse transcriptases which are error-prone, i.e., which may be referred to as error-prone reverse transcriptases or reverse transcriptases which do not support high fidelity incorporation of nucleotides during polymerization. During synthesis of the single-strand DNA flap based on the RT template integrated with the guide RNA, the error-prone reverse transcriptase can introduce one or more nucleotides which are mismatched with the RT template sequence, thereby introducing changes to the nucleotide sequence through erroneous polymerization of the single-strand DNA flap. These errors introduced during synthesis of the single strand DNA flap then become integrated into the double strand molecule through hybridization to the corresponding endogenous target strand, removal of the endogenous displaced strand, ligation, and then through one more round of endogenous DNA repair and / or sequencing processes. Reverse transcription
[0409] As used herein, the term "reverse transcription" indicates the capability of enzyme to synthesize DNA strand (that is, complementary DNA or cDNA) using RNA as a template. In some embodiments, the reverse transcription can be “error-prone reverse transcription,” which refers to the properties of certain reverse transcriptase enzymes which are error-prone in their DNA polymerization activity. Recombinant
[0410] The term “recombinant” as used herein in the context of proteins or nucleic acids refers to proteins or nucleic acids that do not occur in nature, but are the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that comprises at least one, at leasttwo, at least three, at least four, at least five, at least six, or at least seven mutations as compared to any naturally occurring sequence. The fusion proteins (e.g., base editors) described herein are made recombinantly. Recombinant technology is familiar to those skilled in the art. Sense strand
[0411] In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. In the case of a DNA segment that encodes a protein, the sense strand is the strand of DNA that has the same sequence as the mRNA, which takes the antisense strand as its template during transcription, and eventually undergoes (typically, not always) translation into a protein. The antisense strand is thus responsible for the RNA that is later translated to protein, while the sense strand possesses a nearly identical makeup to that of the mRNA. Note that for each segment of dsDNA, there will possibly be two sets of sense and antisense, depending on which direction one reads (since sense and antisense is relative to perspective). It is ultimately the gene product, or mRNA, that dictates which strand of one segment of dsDNA is referred to as sense or antisense.
[0412] In the context of a PEgRNA, the first step is the synthesis of a single-strand complementary DNA (i.e., the 3ʹ ssDNA flap, which becomes incorporated) oriented in the 5ʹ to 3ʹ direction which is templated off of the PEgRNA extension arm. Whether the 3ʹ ssDNA flap should be regarded as a sense or antisense strand depends on the direction of transcription since it well accepted that both strands of DNA may serve as a template for transcription (but not at the same time). Thus, in some embodiments, the 3ʹ ssDNA flap (which overall runs in the 5ʹ to 3ʹ direction) will serve as the sense strand because it is the coding strand. In other embodiments, the 3ʹ ssDNA flap (which overall runs in the 5ʹ to 3ʹ direction) will serve as the antisense strand and thus, the template for transcription. Spacer sequence
[0413] As used herein, the term “spacer sequence” in connection with a guide RNA or a PEgRNA refers to the portion of the guide RNA or PEgRNA of about 20 nucleotides which contains a nucleotide sequence that is complementary to the protospacer sequence in the target DNA sequence. The spacer sequence anneals to the protospacer sequence to form a ssRNA / ssDNA hybrid structure at the target site and a corresponding R loop ssDNA structure of the endogenous DNA strand that is complementary to the protospacer sequence. Split Cas9
[0414] A “split prime editor protein” or “split Cas9” refers to a fusion protein with a split site located within the Cas9 protein that is provided as an N-terminal portion (also referred to as an N-terminal half) and a C-terminal portion (also referred to as a C-terminal half) encoded by two separate nucleotide sequences. The polypeptides corresponding to the N- terminal portion and the C-terminal portion of the Cas9 protein may be combined (joined) to form a complete Cas9 protein. A Cas9 protein is known to consist of a bi-lobed structure linked by a disordered linker (e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp.935–949, 2014, incorporated herein by reference). In some embodiments, the “split” occurs between the two lobes, generating two portions of a Cas9 protein, each containing one lobe. Split intein
[0415] Although inteins are most frequently found as a contiguous domain, some exist in a naturally split form. In this case, the two fragments are expressed as separate polypeptides and must associate before splicing takes place, so-called protein trans-splicing.
[0416] An exemplary split intein is the Ssp DnaE intein, which comprises two subunits, namely, DnaE-N and DnaE-C. The two different subunits are encoded by separate genes, namely dnaE-n and dnaE-c, which encode the DnaE-N and DnaE-C subunits, respectively. DnaE is a naturally occurring split intein in Synechocytis sp. PCC6803 and is capable of directing trans-splicing of two separate proteins, each comprising a fusion with either DnaE- N or DnaE-C.
[0417] Additional naturally occurring or engineered split-intein sequences are known in the or can be made from whole-intein sequences described herein or those available in the art. Examples of split-intein sequences can be found in Stevens et al., “A promiscuous split intein with expanded protein engineering applications,” PNAS, 2017, Vol.114: 8538-8543; Iwai et al., “Highly efficient protein trans-splicing by a naturally split DnaE intein from Nostc punctiforme, FEBS Lett, 580: 1853-1858, each of which are incorporated herein by reference. Additional split intein sequences can be found, for example, in WO 2013 / 045632, WO 2014 / 055782, WO 2016 / 069774, and EP2877490, the contents each of which are incorporated herein by reference.
[0418] In addition, protein splicing in trans has been described in vivo and in vitro (Shingledecker, et al., Gene 207:187 (1998), Southworth, et al., EMBO J.17:918 (1998); Mills, et al., Proc. Natl. Acad. Sci. USA, 95:3543-3548 (1998); Lew, et al., J. Biol. Chem., 273:15887-15890 (1998); Wu, et al., Biochim. Biophys. Acta 35732:1 (1998b), Yamazaki, etal., J. Am. Chem. Soc.120:5591 (1998), Evans, et al., J. Biol. Chem.275:9091 (2000); Otomo, et al., Biochemistry 38:16040-16044 (1999); Otomo, et al., J. Biolmol. NMR 14:105- 114 (1999); Scott, et al., Proc. Natl. Acad. Sci. USA 96:13638-13643 (1999)) and provides the opportunity to express a protein as to two inactive fragments that subsequently undergo ligation to form a functional product. Split prime Editor
[0419] A “split prime editor” refers to a prime editor that is provided as an N-terminal portion (also referred to as a N-terminal half) and a C-terminal portion (also referred to as a C-terminal half) encoded by two separate nucleic acids. The polypeptides corresponding to the N-terminal portion and the C-terminal portion of the prime editor may be combined to form a complete prime editor. In some embodiments, for a prime editor that comprises a dCas9 or nCas9, the “split” is located in the dCas9 or nCas9 domain, at positions as described herein in the split prime editor. Accordingly, in some embodiments, the N-terminal portion of the prime editor contains the N-terminal portion of the split prime editor, and the C-terminal portion of the prime editor contains the C-terminal portion of the split prime editor. Similarly, intein-N or intein-C may be fused to the N-terminal portion or the C-terminal portion of the prime editor, respectively, for the joining of the N- and C-terminal portions of the prime editor to form a complete prime editor. Subject
[0420] The term “subject,” as used herein, refers to an individual organism, for example, an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent (e.g., mouse, rat). In some embodiments, the subject is a domesticated animal. In some embodiments, the subject is a sheep, a goat, a cow, a cat, or a dog. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, e.g., a genetically engineered non-human subject. The subject may be of either sex and at any stage of development. Target site
[0421] The term “target site” refers to a sequence within a nucleic acid molecule that is edited by a prime editor (PE) disclosed herein. The target site further refers to the sequence within a nucleic acid molecule to which a complex of the prime editor (PE) and gRNA binds. Transitions
[0422] As used herein, “transitions” refer to the interchange of purine nucleobases (A ↔ G) or the interchange of pyrimidine nucleobases (C ↔ T). This class of interchanges involves nucleobases of similar shape. The compositions and methods disclosed herein are capable of inducing one or more transitions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both transitions and transversion in the same target DNA molecule. These changes involve A ↔ G, G ↔ A, C ↔ T, or T ↔ C. In the context of a double-strand DNA with Watson-Crick paired nucleobases, transversions refer to the following base pair exchanges: A:T ↔ G:C, G:G ↔ A:T, C:G ↔ T:A, or T:A↔ C:G. The compositions and methods disclosed herein are capable of inducing one or more transitions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both transitions and transversion in the same target DNA molecule, as well as other nucleotide changes, including deletions and insertions. Transversions
[0423] As used herein, “transversions” refer to the interchange of purine nucleobases for pyrimidine nucleobases, or in the reverse and thus, involve the interchange of nucleobases with dissimilar shape. These changes involve T ↔ A, T↔ G, C ↔ G, C ↔ A, A ↔ T, A ↔ C, G ↔ C, and G ↔ T. In the context of a double-strand DNA with Watson-Crick paired nucleobases, transversions refer to the following base pair exchanges: T:A ↔ A:T, T:A ↔ G:C, C:G ↔ G:C, C:G ↔ A:T, A:T ↔ T:A, A:T ↔ C:G, G:C ↔ C:G, and G:C ↔ T:A. The compositions and methods disclosed herein are capable of inducing one or more transversions in a target DNA molecule. The compositions and methods disclosed herein are also capable of inducing both transitions and transversion in the same target DNA molecule, as well as other nucleotide changes, including deletions and insertions. Treatment, Treat, and Treating
[0424] The terms “treatment,” “treat,” and “treating,” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. As used herein, the terms “treatment,” “treat,” and “treating” refer to a clinical intervention aimed to reverse, alleviate, delay the onset of, or inhibit the progress of a disease or disorder, or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay onset of a symptom or inhibit onset or progression of a disease. For example, treatment may beadministered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, for example, to prevent or delay their recurrence. Therapeutically-effective Amount
[0425] “A therapeutically effective amount” as used herein refers to the amount of each therapeutic agent (e.g., prime editor, rAAV) described in the present disclosure required to confer therapeutic effect on the subject, either alone or in combination with one or more other therapeutic agents. Effective amounts vary, as recognized by those skilled in the art, depending on the particular condition being treated, the severity of the condition, the individual subject parameters including age, physical condition, size, gender, and weight, the duration of the treatment, the nature of concurrent therapy (if any), the specific route of administration and like factors within the knowledge and expertise of the health practitioner. These factors are well known to those of ordinary skill in the art and can be addressed with no more than routine experimentation. It is generally preferred that a maximum dose of the individual components or combinations thereof be used, that is, the highest safe dose according to sound medical judgment. It will be understood by those of ordinary skill in the art, however, that a subject may insist upon a lower dose or tolerable dose for medical reasons, psychological reasons or for virtually any other reasons. Empirical considerations, such as the half-life, generally will contribute to the determination of the dosage. For example, therapeutic agents that are compatible with the human immune system, such as polypeptides comprising regions from humanized antibodies or fully human antibodies, may be used to prolong half-life of the polypeptide and to prevent the polypeptide being attacked by the host's immune system. Upstream
[0426] As used herein, the terms “upstream” and “downstream” are terms of relativity that define the linear position of at least two elements located in a nucleic acid molecule (whether single or double-stranded) that is orientated in a 5ʹ-to-3ʹ direction. In particular, a first element is upstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 5’ to the second element. For example, a SNP is upstream of a Cas9-induced nick site if the SNP is on the 5’ side of the nick site. Conversely, a first element is downstream of a second element in a nucleic acid molecule where the first element is positioned somewhere that is 3’ to the second element. For example, a SNP isdownstream of a Cas9-induced nick site if the SNP is on the 3’ side of the nick site. The nucleic acid molecule can be a DNA (double or single stranded). RNA (double or single stranded), or a hybrid of DNA and RNA. The analysis is the same for single strand nucleic acid molecule and a double strand molecule since the terms upstream and downstream are in reference to only a single strand of a nucleic acid molecule, except that one needs to select which strand of the double stranded molecule is being considered. Often, the strand of a double stranded DNA which can be used to determine the positional relativity of at least two elements is the “sense” or “coding” strand. In genetics, a “sense” strand is the segment within double-stranded DNA that runs from 5' to 3', and which is complementary to the antisense strand of DNA, or template strand, which runs from 3' to 5'. Thus, as an example, a SNP nucleobase is “downstream” of a promoter sequence in a genomic DNA (which is double-stranded) if the SNP nucleobase is on the 3' side of the promoter on the sense or coding strand. Variant
[0427] As used herein the term “variant” should be taken to mean the exhibition of qualities that have a pattern that deviates from what occurs in nature, e.g., a variant Cas9 is a Cas9 comprising one or more changes in amino acid residues as compared to a wild type Cas9 amino acid sequence. The term “variant” encompasses homologous proteins having at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 99% percent identity with a reference sequence and having the same or substantially the same functional activity or activities as the reference sequence. The term also encompasses mutants, truncations, or domains of a reference sequence, and which display the same or substantially the same functional activity or activities as the reference sequence. Vector
[0428] The term “vector,” as used herein, refers to a nucleic acid that can be modified to encode a gene of interest and that is able to enter into a host cell, mutate and replicate within the host cell, and then transfer a replicated form of the vector into another host cell. Exemplary suitable vectors include viral vectors, such as retroviral vectors or bacteriophages and filamentous phage, and conjugative plasmids. Additional suitable vectors will be apparent to those of skill in the art based on the instant disclosure.Wild type
[0429] As used herein the term “wild type” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene or characteristic as it occurs in nature as distinguished from mutant or variant forms. DETAILED DESCRIPTION
[0430] Provided herein, are compositions (e.g.,vectors, recombinant viruses), cells, and kits comprising nucleic acids encoding a prime editor and / or rAAV particles comprising nucleic acids encoding a prime editor. In some embodiments, the prime editor comprises a nucleic acid programmable DNA binding protein (hereinafter “napDNAbp”). The prime editor may, in some cases, comprise a polymerase (e.g., DNA-dependent DNA polymerase or RNA-dependent DNA polymerase, such as, reverse transcriptase), or a variant thereof, which can be provided as a fusion protein with the napDNAbp or other programmable nuclease. The napDNAbp and polymerase may be fused using any technique known to those of skill in the art, for example, using a linker covalently conjugated to each protein.
[0431] In some embodiments, a nucleic acid encoding the fusion protein is split (hereinafter “split prime editor”) into two or more components, for example, to enable its delivery to a cell (e.g., in vitro or in vivo). In some embodiments, the split prime editor is split at a site located within the napDNAbp domian (e.g., SpCas9).
[0432] Without being bound by any particular theory, it is generally believed that napDNAbp domain (e.g., Cas9, SpCas9, etc.), has a N-terminal lobe and a C-terminal lobe linked by a disordered linker (e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp.935–949, 2014, incorporated herein by reference). In some embodiments, the N- terminal portion of the split prime editor comprises the N-terminal lobe of a napDNAbp protein (e.g., SpCas9). In some embodiments, the C-terminal portion of the split prime editor comprises the C-terminal lobe of napDNAbp (e.g., SpCas9). In some embodiments, the N- terminal portion of the split prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 in SEQ ID NO: 14. “1-844” means starting from amino acid 1 and ending at amino acid 844 and “1- 1024” means starting from amino acid 1 and ending at amino acid 1024.
[0433] The C-terminal portion of the split prime editor can be joined with the N-terminal portion of the split prime editor, e.g., via intein-mediated protein splicing, to form a complete fusion protein (e.g., prime editor). In some embodiments, the C-terminal portion of the prime editor starts from where the N-terminal portion of the napDNAbp protein ends. Assuch, in some embodiments, the C-terminal portion of the split prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 845-1368 or 1024-1368 of SEQ ID NO: 14. “845-1368” means starting at amino acid 845 and ending at amino acid 1368.
[0434] napDNAbp variants may also be delivered to cells using the methods described herein. For example, a napDNAbp variant may also be “split” as described herein. A napDNAbp variant may comprise an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the napDNAbp sequences provided herein. In some embodiments, the napDNAbp variant comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 30%, no more than 25%, no more than 20%, no more than 15%, no more than 10%, no more than 5%, no more than 1% longer or shorter) than any of the Cas9 proteins provided herein. In some embodiments, fusion protein comprises a UGI domain. The UGI domain, according to some embodiments, may comprise an amino acid sequence that is shorter or longer in length (e.g., by no more than 200 amino acids, no more than 150 amino acids, no more than 100 amino acids, no more than 50 amino acids, no more than 10 amino acids, no more than 5 amino acids, or no more than 2 amino acids longer or shorter) than any of the Cas9 proteins provided herein.
[0435] In some embodiments, the N-terminal portion of a split prime editor comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the corresponding portion of any one of the napDNAbp sequences provided herein. In some embodiments, the N-terminal portion of the split prime editor comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 30%, no more than 25%, no more than 20%, no more than 15%, no more than 10%, no more than 5%, no more than 1% longer or shorter) than the corresponding portion of any of the napDNAbp sequences provided herein. In some embodiments, the N-terminal portion of the split prime editor comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 200 amino acids, no more than 150 amino acids, no more than 100 amino acids, no more than 50 amino acids, no more than 10 amino acids, no more than 5 amino acids, or no more than 2 amino acids longer or shorter) than the corresponding portion of any of the napDNAbp domains provided herein.
[0436] In some embodiments, the C-terminal portion of a split prime editor comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the corresponding portion of any one of the napDNAbp sequences provided herein. In some embodiments, the C-terminal portion of the split prime editor comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 30%, no more than 25%, no more than 20%, no more than 15%, no more than 10%, no more than 5%, no more than 1% longer or shorter) than the corresponding portion of any of the napDNAbp sequences provided herein. In some embodiments, the C-terminal portion of the split prime editor comprises an amino acid sequence that is shorter or longer in length (e.g., by no more than 200 amino acids, no more than 150 amino acids, no more than 100 amino acids, no more than 50 amino acids, no more than 10 amino acids, no more than 5 amino acids, or no more than 2 amino acids longer or shorter) than the corresponding portion of any of the napDNAbp sequences provided herein.
[0437] In some embodiments, the napDNAbp variant is a SpCas9. In some embodiments, the N-terminal portion of the split prime editor comprises a mutation corresponding to a D10A mutation in SEQ ID NO: 14. In some embodiments, the N-terminal portion of the split prime editor comprises a mutation corresponding to a D10A mutation in SEQ ID NO: 14 and the C-terminal portion of the split prime editor comprises a mutation corresponding to a H840A mutation in SEQ ID NO: 14. In some embodiments, the N-terminal portion of the split prime editor comprises a mutation corresponding to a D10A mutation in SEQ ID NO: 14, and the C-terminal portion of the split prime editor comprises a histidine at the position corresponding to position 840 in SEQ ID NO:14.
[0438] In some embodiments, to join the N-terminal portion of the napDNAbp and the C- terminal portion of the napDNAbp, an intein system may be used. In some embodiments, the N-terminal portion of the napDNAbp is fused to an intein-N. In some embodiments, the intein-N is fused to the C-terminus of the N-terminal portion of the napDNAbp to form a structure of NH2-[N-terminal portion of napDNAbp]-[intein-N]-COOH. In some embodiments, the intein-N is encoded by the dnaE-n gene. In some embodiments, the intein- N comprises the amino acid sequence of SEQ ID NOs: 2-9 , 183-184, or 187-188. In some embodiments, the C-terminal portion of the napDNAbp is fused to an intein-C, and the intein- C is fused to the N-terminus of the C-terminal portion of the napDNAbp to form a structure of NH2-[intein-C]-[C-terminal portion of napDNAbp]-COOH. In some embodiments, the intein-C is encoded by the dnaE-c gene. In some embodiments, the intein-C comprises theamino acid sequence of SEQ ID NO: 2-9, 185-186, or 189-190. Other split intein systems may also be used in the present disclosure and are known in the art.
[0439] In some embodiments, the inteins are catatlytically active (e.g., N-intein and C- intein). As used herein, the phrase “catalytically active” refers to the ability of the inteins to undergo an autocatalytic protein splicing reaction. In other embodiments, inteins have been mutated to decrease the efficiency of the autocatalytic protein splicing reation. In some embodiments, compositions comprising mutated inteins exhibit a decrease in editing efficiency of about 40% compared to compositions comprising catalytically competent constructs. In other words, catalatyically compentent split-intein constructs recombine more efficiently than constructs comprising catalytically inactive inteins domains. Thus, in some embodiments, the spilt-intein constructs of the present disclosure comprise catalytically active inteins.
[0440] In some embodiments, a complete split prime editor is formed when a N-terminal portion and the C-terminal portion of the napDNAbp are joined. In some embodiments, the split prime editor may comprise any one of the following structures:
[0441] NH2-[ polymerase]-[ napDNAbp]-COOH
[0442] NH2-[ napDNAbp]-[ polymerase]-COOH.
[0443] Other architectures are also possible and are herein contemplated.
[0444] In some embodiments, the first nucleotide sequence or the second nucleotide sequence (encoding either the split prime editor protein) is operably linked to a nucleotide sequence encoding at least one bipartite nuclear localization signal (NLS). For example, the first nucleotide sequence may be operably linked to a nucleotide sequence encoding one or more (e.g., 2, 3, 4, 5, or more) bipartite NLS. In some embodiments, the second nucleotide sequence may be operably linked to a nucleotide sequence encoding one or more (e.g., 2, 3, 4, 5, or more) bipartite NLSs. As such, the split prime editor formed by joining the N- terminal portion and the C-terminal portion may comprise one or more bipartite NLSs. For example, the split prime editor may comprise any one of the following structures (bNLS means one or more bipartite nuclear localization signals):
[0445] NH2-bNLS-[napDNAbp]-COOH
[0446] NH2-[napDNAbp]-bNLS-COOH
[0447] NH2-[bNLS]-[napDNAbp]-[ polymerase]-COOH
[0448] NH2-[napDNAbp]-[bNLS]-[ polymerase]-COOH
[0449] NH2-[napDNAbp]-[ polymerase]-[bNLS]-COOH
[0450] NH2-[bNLS]-[napDNAbp]-[bNLS]-[ polymerase]-COOH
[0451] NH2-[napDNAbp]-[bNLS]-[ polymerase]-[bNLS]-COOH
[0452] NH2-[bNLS]-[napDNAbp]-[polymerase]-[bNLS]-COOH
[0453] NH2-[bNLS]-[napDNAbp]-[bNLS]-[polymerase]-[bNLS]-COOH
[0454] NH2-[bNLS]-[ polymerase]-[ napDNAbp]-COOH
[0455] NH2-[ polymerase]-[bNLS]-[ napDNAbp]-COOH
[0456] NH2-[ polymerase]-[ napDNAbp]-[bNLS]-COOH
[0457] NH2-[bNLS]-[ polymerase]-[bNLS]-[ napDNAbp]-COOH
[0458] NH2-[ polymerase]-[bNLS]-[ napDNAbp]-[bNLS]-COOH
[0459] NH2-[bNLS]-[polymerase]-[ napDNAbp]-[bNLS]-COOH
[0460] NH2-[bNLS]-[ polymerase]-[bNLS]-[ napDNAbp]-[bNLS]-COOH
[0461] Aspects of the present disclosure relate to compositions comprising split prime editors configured to optimize in vivo expression in a mammal. In some embodiments, the composition comprises a v1em split prime editor, a v2em split prime editor, and / or a v3em split prime editor. It should be understood by those of skill in the art, however, that the following descriptions are non-limiting and are intended only to provide the skilled artisan exemplary embodiments. Other embodiments (e.g., compositions and / or configurations) are also herein contemplated.
[0462] In some embodiments, a composition comprising a v1em split prime editor comprises a first nucleotide sequence encoding a N-terminal portion with the structure NH2- EFS promoter-N-term PEmax (start codon-SV40NLS-SpCAS9)NpuN-SV40NLS-W3-bGH polyA-sgRNA (protospacer in bold)-human U6-COOH and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion with the structure NH2- EFS promoter- SV40NLS- NpuC-C-term PEmax (SpCas9-RT-SV40NLS)-W3-bGH polyA- epegRNA (protospacer in bold)-human U6-COOH.
[0463] In some embodiments, a composition comprising a v2em split prime editor has a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C- terminus to an intein-N with the structure NH2-Cbh promoter-N-term PEmax (start codon- SV40NLS-SpCas9)-NpuN-SV40NLS-W3- bGH polyA-COOH, a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor with the structure NH2-Cbh promoter-SV40NLS-NpuC-C-term PEmax (SpCas9-RT- SV40NLS)-W3-bGHpolyA-COOH, and a third nucleotide sequence encoding at least one pegRNA and at least one sgRNA with exemplary structure NH2-Cbh promoter-EGFP:KASH -sgRNA (protospacer in bold)-mouse U6-epegRNA (protospacer in bold)-human U6-COOH.
[0464] In some embodiments, a composition comprising a v3em split prime editor has a first nucleotide sequence encoding a N-terminal portion with structure NH2-Cbh promoter-N- term PEmax (start codon-SV40NLS-SpCas9)-NpuN-SV40NLS-SV40 late polyA-COOH and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion with the structure NH2-Cbh promoter-SV40NLS-NpuC-C-term PEmax (SpCas9-RT- SV40NLS)-SV40 late polyA-sgRNA (protospacer in bold)-mouse U6-epegRNA (protospacer in bold)-human U6-COOH
[0465] In some embodiments, the disclosure relates to compositions comprising an rAAV encoding any one of the prime editors disclosed herein (e.g., v1em prime editors, a v2em prime editors, and / or a v3em prime editors).
[0466] In some embodiments, a composition comprising a first rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of a v1em-PE (herein v1em-PE- rAAV) fused at its C-terminus to an intein-N with the structure ITR-EFS promoter-N-term PEmax (start codon-SV40NLS-SpCAS9)NpuN-SV40NLS-W3-bGH polyA-sgRNA (protospacer in bold)-human U6-ITR and a second recombinant rAAV particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of v1em-PE with the structure ITR-EFS promoter- SV40NLS- NpuC-C-term PEmax (SpCas9-RT-SV40NLS)-W3-bGH polyA-epegRNA (protospacer in bold)-human U6-ITR.
[0467] In some embodiments, a composition comprising a first rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of a v2em-PE (herein v2em-PE- rAAV) fused at its C-terminus to an intein-N with the structure ITR-Cbh promoter-N-term PEmax (start codon-SV40NLS-SpCas9)-NpuN-SV40NLS-W3- bGH polyA-ITR, a second recombinant rAAV particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of v2em-PE with the structure ITR-Cbh promoter-SV40NLS-NpuC-C-term PEmax (SpCas9-RT-SV40NLS)-W3-bGHpolyA-ITR and a third nucleotide sequence encoding at least one pegRNA and at least one sgRNA with exemplary structure ITR-Cbh promoter-EGFP:KASH -sgRNA (protospacer in bold)-mouse U6-epegRNA (protospacer in bold)-human U6-ITR.
[0468] In some embodiments, a composition comprising a first rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of a v3em-PE (herein v3em-PE- rAAV) fused at its C-terminus to an intein-N with the structure ITR-Cbh promoter-N-term PEmax (start codon-SV40NLS-SpCas9)-NpuN-SV40NLS-SV40 late polyA-ITR and a second recombinant rAAV particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of v3em-PE with the structure ITR-Cbh promoter-SV40NLS-NpuC-C-term PEmax (SpCas9-RT-SV40NLS)-SV40 late polyA- sgRNA (protospacer in bold)-mouse U6-epegRNA (protospacer in bold)-human U6-ITR.
[0469] In some embodiments, the prime editors disclosed herein require co-transduction of one or more rAAVs to deliver all necessary components for the expression and assembly of the split prime editors. For example, the v1em-PE-AAV and v3em-PE-AAV systems described herein require co-delivery of both a first rAAV particle comprising a first nucleotide sequence encoding a N-terminal portion of the prime editor fused at its C-terminus to an intein-N and a second rAAV particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor.
[0470] The v2em-PE-AAV systems, on the other hand, requires co-delivery of three rAAV particles to enable expression and assembly of the v3em-PE system. Similar to the v1em-PE and v3em-PE systems, the v2em-PE system also requires a first rAAV particle and a second rAAV particle encoding the N-terminal and C-terminal portions of the prime editor, respectively. Contrary to the other described systems, the v2em-PE system is designed to deliver nucleotide constructs that are too large to fit within a typical rAAV. For example, in some embodiments, the v2em-PE system may comprises a Cbh promoter (e.g., 0.7kb), which is significantly larger than other promoters (e.g., compared to EFS promoter in v1em-PE systems), but is known to be a stronger and ubiquitous promoter capable of mediating efficient AAV-mediated based editing when injected systemically. Thus, in some embodiments, the pegRNA and sgRNA cassettes are moved to a third AAV vector. In some cases, the third AAV vector further comprises one or more promoters (e.g., human U6 promoter that drives pegRNA expression and / or a mouse U6 promoter that drives nicking sgRNA expression).
[0471] Alternatively, or additionally, the inventors have discovered that the dual v3em- PE-AAV system may be used to deliver nucleotides with larger promoters by truncating the polymerase domain of the fusion protein. Without wishing to be bound by theory, it is believed that truncated MMLV reverse transcriptase variants that lack the RNAseH domain retain reverse transcriptase activity across of variety of edits (e.g., FIG.11). For instance, in some embodiments, the truncated MMLV RT retain greater than or equal to 50% RT activity, greater than or equal to 60% RT activity, greater than or equal to 70% RT activity, greater than or equal to 80% RT activity, greater than or equal to 90% RT activity, and greater than or equal to 95% RT activity, relative to the full length MMLV RT. In some embodiments, the truncated MMLV RT retains less than or equal to 95% RT activity, less than or equal to 90% RT activity, less than or equal to 80% RT activity, less than or equal to 70% RT activity,less than or equal to 60% RT activity, and less than or equal to 50% RT activity, relative to the full length MMLV RT. Combinations are also possible. Other ranges are also possible.
[0472] In some embodiments, any one of the prime editor systems disclosed herein (e.g., v1em-PE editors, v2em-PE editors, and v3em-PE editors) comprise a split site within the napDNAbp domain of the fusion protein located at either position 844 or 1024. The N- terminal portion of the prime editor, in some embodiments, comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 corresponding to amino acids 1-844 or 1- 1024 of SEQ ID NO: 14.
[0473] In some embodiments, any one of the prime editor systems disclosed herein (e.g., v1em-PE editors, v2em-PE editors, and v3em-PE editors) comprise three N-terminal amino acid mutations at the C-terminal portion of the split prime editor. The mutations may comprise any amino acid known to those of skill in the art and in any combination thereof. In one set of embodiments, the mutation comprises changing the three N-terminal amino acids of the C-terminal extein from SEQ to SFQ, SFN, SEN, or CFN In some embodiments, the preferred composition comprises a napDNAbp split at amino acid 1024, wherein the N- terminal amino acid of the C-terminal portion comprises the CFN e.g., Cys-Phen-Asn) mutation. In other embodiments, the preferred composition comprises a napDNAbp split at amino acid 844, wherein the N-terminal amino acid of the C-terminal portion comprises the CFN e.g., Cys-Phen-Asn) mutation
[0474] In some embodiments, any one of the rAAV compositions disclosed herein (e.g., v1em-PE-AAV, v2em-PE-AAV, and v3em-PE-AAV) comprise a split site within the napDNAbp domain of the fusion protein located at either position 844 or 1024. The N- terminal portion of the prime editor, in some embodiments, comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 corresponding to amino acids 1-844 or 1- 1024 of SEQ ID NO: 14.
[0475] In some embodiments, any one of the rAAV compositions disclosed here in (e.g., v1em-PE-AAV, v2em-PE-AAV, and v3em-PE-AAV) comprise three N-terminal amino acid mutations at the C-terminal portion of the split prime editor. The mutations may comprise any amino acid known to those of skill in the art and in any combination thereof. In one set of embodiments, the mutation comprises changing the three N-terminal amino acids from SEQ to SFQ, SFN, SEN, or CFN. In some embodiments, the preferred composition comprises a napDNAbp split at amino acid 1024, wherein the N-terminal amino acid of the C-terminal portion comprises the CFN e.g., Cys-Phen-Asn) mutation. In other embodiments, the preferred composition comprises a napDNAbp split at amino acid 844, wherein the N-terminal amino acid of the C-terminal portion comprises the CFN e.g., Cys-Phen-Asn) mutation.
[0476] In some embodiments, any one of the prime editor systems disclosed herein (e.g., v1em-PE editors, v2em-PE editors, and v3em-PE editors) further comprise a nucleotide encoding a pegRNA. In some cases, the pegRNA is encoded within the first nucleotide sequence or the second nucleotide sequence, optionally, being operably linked to a promoter. Any pegRNA known to the skilled artisan may be used. A detailed description of pegRNAs may be found elsewhere herein.
[0477] In some embodiments, the pegRNA guides the prime editor to an editing site. Without wishing to be bound by any particular theory, it is generally believed that the identity of the installed edit effects the susceptibility of the edit to repair by the DNA mismatch repair pathway (MMR). Thus, in some embodiments, the pegRNA is selected to encode edits that natively evade the MMR pathway. Any pegRNA sequence known in the art to instill edits that natively evade the MMR pathway are herein contemplated. For instance, in some embodiments, the pegRNA installs the following edits: (1) a +1 C-to-G edit, (2) a +1 C-to-G and +5 G-to-T edit, (3) a +2 G-to-C edit, (4) a +1 CTT insertion, (5) a +1 CCC insertion; and (6) a +1 GCA insertion at a target locus (e.g., Dnmt1 locus in mouse neuro 2a cells). In some embodiments, the pegRNA comprises a 3′ stabilizing motif.
[0478] In some embodiments, the aforementioned pegRNA edits increase the average editing efficiency, from 7.6% to between 14% and 45%, relative to a previously validated +5 G-to-T edit (See FIG.7). In a preferred set of embodiments, the +2 G-to-C edit and +1 CCC insertion edit increase the average editing efficiency between 20% and 52%. Again, without wishing to be bound by theory, it is believed that the prime editing intermediates containing C-C mismatches and multiple contiguous insertions are poor substrates for MMR.
[0479] The skilled artisan will understand that a prime editor comprising a pegRNA may be used to instill edits in a target locus that natively evade one or more repair pathways that decrease editing efficiency (e.g., MMR pathway). The edited site of interest, in some instances, may be the site of a mutation causing a disease; however, in other embodiments, the edited site may be any benign or silent bystander mutation known in the art to enhance in vitro and in vivo prime editing efficiency (e.g., the mammalian brain).
[0480] In some embodiments, any one of the rAAV compositions disclosed herein (e.g., v1em-PE-AAV, v2em-PE-AAV, and v3em-PE-AAV) further comprises a nucleotide encoding a pegRNA. In some cases, the pegRNA is encoded within the first nucleotidesequence of the first rAAV particle or the second nucleotide sequence of the second rAAV particle, optionally, being operably linked to a promoter.
[0481] In some embodiments, any one of the prime editors disclosed herein (e.g., v1em- PE editors, v2em-PE editors, and v3em-PE editors) further comprise a nucleotide encoding a polymerase. In some embodiments, the polymerase is a MMLV reverse transcriptase. In some embodiments, the MMLV reverse transcriptase is codon optimized for expression in a mammalian (e.g., human cell).
[0482] In some embodiments, any one of the rAAV compositions disclosed herein (e.g., v1em-PE-AAV, v2em-PE-AAV, and v3em-PE-AAV) further comprise a nucleotide encoding a polymerase. In some embodiments, the polymerase is a MMLV reverse transcriptase. In some embodiments, the MMLV reverse transcriptase is codon optimized for expression in a mammalian (e.g., human cell).
[0483] In some embodiments, any one of the prime editors disclosed herein (e.g., v1em- PE editors, v2em-PE editors, and v3em-PE editors) and / or any one of the rAAV compositions disclosed herein (e.g., v1em-PE-AAV, v2em-PE-AAV, and v3em-PE-AAV) further comprise a nucleotide encoding a nicking sgRNA. Without be bound by theory, it is believed that inclusion of a nicking sgRNA can increase prime editing efficiencies in cells by biasing cellular repair machinery to repair the non-edited strand.
[0484] In some embodiments, in vivo editing using any of the systems, compositions, cells, and kits described herein may results in off-target effects. In some embodiments, off- target prime editing of off-target loci for pegRNA / sgRNA encoding a target edit using the v3em-PE3 prime editor or v3em-PE3-AAV systems is less than or equal to 0.1%, less than or equal to 0.05%, less than or equal to 0.01%, and less than or equal to 0.005% of the total sequencing reads with a specified target edit. In other embodiments, the percent indels for pegRNA / sgRNAs encoding a target edit using the v3em-PE3 prime editor or v3em-PE3- AAV systems is less than or equal to 1%, less than or equal to 0.5%, less than or equal to 0.1%, less than or equal to 0.05%, less than or equal to 0.01%, and less than 0.005% of the total sequencing reads with indels.
[0485] In some embodiments, off-target prime editing of off-target loci for pegRNA / sgRNA encoding a target edit using the v1em-PE prime editor or v1em-PE-AAV systems described herein is less than or equal to 0.1%, less than or equal to 0.05%, less than or equal to 0.01%, and less than or equal to 0.005% of the total sequencing reads with a specified target edit. In other embodiments, the percent indels for pegRNA / sgRNAs encoding a target edit using the v1em-PE prime editor or v1em-PE-AAV systems is less thanor equal to 1%, less than or equal to 0.5%, less than or equal to 0.1%, less than or equal to 0.05%, less than or equal to 0.01%, and less than 0.005% of the total sequencing reads with indels.
[0486] In some embodiments, off-target prime editing of off-target loci for pegRNA / sgRNA encoding a target edit using the v2em-PE prime editor or v2em-PE-AAV systems described herein is less than or equal to 0.1%, less than or equal to 0.05%, less than or equal to 0.01%, and less than or equal to 0.005% of the total sequencing reads with a specified target edit. In other embodiments, the percent indels for pegRNA / sgRNAs encoding a target edit using the v2em-PE prime editor or v2em-PE-AAV systems is less than or equal to 1%, less than or equal to 0.5%, less than or equal to 0.1%, less than or equal to 0.05%, less than or equal to 0.01%, and less than 0.005% of the total sequencing reads with indels.
[0487] Aspects of the present disclosure relate to a cell comprising any one of the prime editors (e.g., v1em-PE, v2em-PE, and v3em-PE) and / or rAAV particles (e.g., v1em-PE- AAV, v2em-PE-AAV, and v3em-PE-AAV) disclosed herein. In some cases, any one of the prime editors disclosed herein (e.g., v1em-PE, v2em-PE, and v3em-PE) may be complexed with a pharmaceutical excipient known in the art to enhance transfection of the nucleotide sequences into the cell. For example, in some embodiments, a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N and a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor may be encapsulated in a lipid nanoparticle or liposome, or complexed with PEI to form a nanoparticle, as described elsewhere herein (see Pharmaceutical Compositions Section). Thus, in some embodiments, the first nucleotide sequence, second nucleotide sequence, and third nucleotide sequence may be simultaneously encapsulated within a single nanoparticle or complex; or alternatively, each nucleotide sequence may be individually encapsulated and / or complexed to a carrier known to enhance transfection. Other combinations are also possible.
[0488] Combinations of delivery methods are also possible. For example, in some cases, any one of the disclosed split prime editing systems disclosed herein may be delivered using any one of the AAV systems described herein. In some embodiments, a combination of AAV systems and other delivery systems may be used, for example, lipid nanoparticles, liposomes, lipid-polymer complexes and the like.
[0489] Aspects of the present disclosure relate to methods of using any one of the compositions comprising the split prime editors described herein.
[0490] In some embodiments, the method comprises systemically delivering any one of the prime editors disclosed herein to an organ of a subject in need thereof. Any method known to one of skill in the art resulting in systemic delivery may be used. Non-limiting examples, intravenous injection, retroorbital injection (RO), intravenous injection (IV), and intraosseous injection (IO). In other embodiments, the direct injection of the composition into an organ of interest is preferred (e.g., intracerebroventricular injection, ICV, into the brain). In some embodiments, the subject is a mammal (e.g., a human). In some cases, the mammal is a human. The subject may be of any age (e.g., neonatal, an infant, a child, a teenager, or an adult) or gender (e.g., male, female)
[0491] In some embodiments, the method comprises delivery of compositions comprising any one of the rAAV systems described herein. In some cases, the rAAV construct may be configured to cross the blood-brain delivery, for example, to perform editing in the brain (e.g., to treat Alzheimer’s disease).
[0492] In some embodiments, the methods described herein increase the in vivo editing efficiency in one or more organs. For example, in some embodiments, administering to a newborn mammal (e.g., a P0 mouse pup), via ICV injection, the prime editor comprising v1 PE3-AAV9 with W3 encoding +1 C-to-G, +1 CCC insertion, or +2 G-to-C edits results in approximately 6.1%, 25%, and 42% prime editing for +1 C-to-G, +1 CCC insertion and +2 G-to-C edits, respectively in the brain.
[0493] In some embodiments, v3em-PE3-AAV9 systems result in higher editing efficiencies in adult bulk heart and skeletal muscle, relative to v2em-PE3-AAV9 systems. For instance, in some cases, prime editing is increase 2.1-fold and 4.9-fold, respectively for heart and skeletal muscle, for v3em compared to v2em-PE3-AAV9.
[0494] In some embodiments, the v3em-PE3-AAV9 system results in higher editing efficiencies in adult bulk liver tissue, relative to v1em-PE3-AAV9 systems. For instance, in some cases, prime editing is increased between 14% and 46% (depending on the dose administered) for the v3em-PE3-AAV9 system, compared to between 0.1% and 5.7% prime editing for the v1em-PE3-AAV9 system. Other exemplary embodiments are described elsewhere herein (see Examples).
[0495] In some embodiments, the method comprises preferentially editing a neuron or a plurality of neurons. In some cases, the method comprises editing a combination of neurons and astrocytes or a plurality of neurons and a plurality of astrocytes in a subject in need thereof. In some embodiments, the method comprises administering to the subject the v1em- PE-AAV system to preferentially edit a neuron or a plurality of neurons. In otherembodiments, the method comprises administering to the subject the v3em-PE-AAV system to broadly edit multiple CNS cell types (e.g., neurons and astrocytes).
[0496] In some embodiments, the disclosure relates to a method of contacting a cell with the composition of any one of the compositions described herein, wherein the contacting results in the delivery of the first nucleotide sequence and the second nucleotide sequence into the cell, and wherein the N-terminal portion of the prime editor and the C-terminal portion of the prime editor are joined to form a prime editor. In some embodiments, the method comprises delivery of a third nucleotide sequences into the cell.
[0497] In some embodiments, the disclosure relates to methods for editing one or more target genes in an organ of interest. The method, according to some embodiments, comprises administering, to a subject in need, an N-terminal portion of a prime editor. In some embodiments, the method comprises administering, to a subject in need, a C-termial portion of the prime editor. The N-terminal portion and the C-terminal portion both comprise a nucleic acid, in some embodiments.
[0498] In some embodiments, the ratio of the N-terminal portion to the C-terminal portion is 1:1. In other embodiments, the ratio of the N-terminal portion to the C-terminal portion is greater than or equal to 1:10, greater than or equal to 1:9, greater than or equal to 1:8, greater than or equal to 1:7, greater than or equal to 1:6, greater than or equal to 1:5, greater than or equal to 1:4, greater than or equal to 1:2, greater than or equal to 1:1. In some embodiments, the ratio of the N-terminal portion to the C-terminal portion is less than or equal to 1:1, less than or equal to 1:2; less than or equal to 1:3, less than or equal to 1:4, less than or equal to 1:5, less than or equal to 1:6, less than or equal to 1:7, less than or equal to 1:8, less than or equal to 1:9, or less than or equal to 1:10.
[0499] In some embodiments, the ratio of the N-terminal portion to the C-terminal portion is greater than or equal to 1:1, greater than or equal to 2:1, greater than or equal to 3:1, greater than or equal to 4:1, greater than or equal to 5:1, greater than or equal to 6:1, greater than or equal to 7:1, greater than or equal to 8:1, greater than or equal to 9:1, or greater than or equal to 10:1. In other embodiments, the ratioof the N-terminal portion to the C-terminal portion is less than or equal to 10:1, less than or equal to 9:1, less than or equal to 8:1, less than or equal to 7:1, less than or equal to 6:1, less than or equal to 5:1, less than or equal to 4:1, less than or equal to 3:1, less than or equal to 2:1, or less than or equal to 1:1.
[0500] In some embodiments, the disclosure relates to a method comprising administering to a subject in need thereof a therapeutically effective amount of any one of the compositions described and a pharmaceutically acceptable excipient.
[0501] In some embodiments, the disclosure relates to a method of treating a disease (e.g., a neurological disease) in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of any one of the compositions described herein and a pharmaceutically acceptable excipient. For example, in some instances, the compositions described herein may be used to install the putatively protective APOE Christchurch (APOE3 R136S) coding variant, which is a G-to-T transversion mutation that cannot be installed via base editing and would be difficult to install in post-mitotic or slowly dividing cells via HDR. This mutation is of biological and therapeutic interest, as it has been observed in an individual who carried the risk-associated PSEN1 (presenilin 1) E280A mutation but who did not develop cognitive impairment until three decades after the expected age of clinical onset of Alzheimer’s Disease (AD) among PSEN1 E280A carriers. A subsequent study suggests the APOE Christchurch variant may be deleterious in other genetic contexts. The ability to precisely install the APOE Christchurch allele in relevant cells in vivo could help illuminate the mutation’s influence on AD pathology.
[0502] Other exemplary therapeutic targets include, for example, the proprotein convertase subtilisin / kexin type 9 (e.g., Pcsk9), a therapeutically relevant gene involved in cholesterol homeostasis. Without wishing to be bound by theory, it is believed that installing a premature stop codon or splice site mutation in the Pcsk9 results in a reduction of Psck9 protein levels in the liver and a drop in circulating LDL cholesterol. Vectors
[0503] Some aspects of the present disclosure relate to using recombinant virus vectors (e.g., adeno-associated virus vectors, adenovirus vectors, or herpes simplex virus vectors) for the delivery of the prime editors or components thereof described herein, e.g., the split prime editor, into a cell. In the case of a split-PE approach, the N-terminal portion of a PE fusion protein and the C-terminal portion of a PE fusion are delivered by separate recombinant virus vectors (e.g., adeno-associated virus vectors, adenovirus vectors, or herpes simplex virus vectors) into the same cell, since the full-length Cas9 protein or prime editors exceeds the packaging limit of various virus vectors, e.g., rAAV (~4.9 kb).
[0504] Thus, in one embodiment, the disclosure contemplates vectors capable of delivering split prime editor fusion proteins, or split components thereof. In some embodiments, a composition for delivering the split prime editor protein or split prime editor into a cell (e.g., a mammalian cell, a human cell) is provided. In some embodiments, the composition of the present disclosure comprises: (i) a first recombinant adeno-associatedvirus (rAAV) particle comprising a first nucleotide sequence encoding a N-terminal portion of a napDNAbp or prime editor fused at its C-terminus to an intein-N; and (ii) a second recombinant adeno-associated virus (rAAV) particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the napDNAbp or prime editor. The rAAV particles of the present disclosure comprise a rAAV vector (i.e., a recombinant genome of the rAAV) encapsidated in the viral capsid proteins.
[0505] In some embodiments, the rAAV vector comprises: (1) a heterologous nucleic acid region comprising the first or second nucleotide sequence encoding the N-terminal portion or C-terminal portion of a split prime editor protein or a split prime editor in any form as described herein, (2) one or more nucleotide sequences comprising a sequence that facilitates expression of the heterologous nucleic acid region (e.g., a promoter), and (3) one or more nucleic acid regions comprising a sequence that facilitate integration of the heterologous nucleic acid region (optionally with the one or more nucleic acid regions comprising a sequence that facilitates expression) into the genome of a cell. In some embodiments, viral sequences that facilitate integration comprise Inverted Terminal Repeat (ITR) sequences. In some embodiments, the first or second nucleotide sequence encoding the N-terminal portion or C-terminal portion of a split prime editor protein or a split prime editor is flanked on each side by an ITR sequence. In some embodiments, the nucleic acid vector further comprises a region encoding an AAV Rep protein as described herein, either contained within the region flanked by ITRs or outside the region. The ITR sequences can be derived from any AAV serotype (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) or can be derived from more than one serotype. In some embodiments, the ITR sequences are derived from AAV2 or AAV6.
[0506] Thus, in some embodiments, the rAAV particles disclosed herein comprise at least one rAAV2 particle, rAAV6 particle, rAAV8 particle, rPHP.B particle, rPHP.eB particle, or rAAV9 particle, or a variant thereof. In particular embodiments, the disclosed rAAV particles are rPHP.B particles, rPHP.eB particles, rAAV9 particles.
[0507] ITR sequences and plasmids containing ITR sequences are known in the art and commercially available (see, e.g., products and services available from Vector Biolabs, Philadelphia, PA; Cellbiolabs, San Diego, CA; Agilent Technologies, Santa Clara, Ca; and Addgene, Cambridge, MA; and Gene delivery to skeletal muscle results in sustained expression and systemic delivery of a therapeutic protein. Kessler PD, Podsakoff GM, Chen X, McQuiston SA, Colosi PC, Matelis LA, Kurtzman GJ, Byrne BJ. Proc Natl Acad Sci USA.1996 Nov 26;93(24):14082-7; and Curtis A. Machida. Methods in MolecularMedicine™. Viral Vectors for Gene Therapy Methods and Protocols.10.1385 / 1-59259-304- 6:201 © Humana Press Inc.2003. Chapter 10. Targeted Integration by Adeno-Associated Virus. Matthew D. Weitzman, Samuel M. Young Jr., Toni Cathomen and Richard Jude Samulski; U.S. Pat. Nos.5,139,941 and 5,962,313, all of which are incorporated herein by reference).
[0508] In some embodiments, the rAAV vector of the present disclosure comprises one or more regulatory elements to control the expression of the heterologous nucleic acid region (e.g., promoters, transcriptional terminators, and / or other regulatory elements). In some embodiments, the first and / or second nucleotide sequence is operably linked to one or more (e.g., 1, 2, 3, 4, 5, or more) transcriptional terminators. Non-limiting examples of transcriptional terminators that may be used in accordance with the present disclosure include transcription terminators of the bovine growth hormone gene (bGH), human growth hormone gene (hGH), SV40, CW3, ϕ, or combinations thereof. The efficiencies of several transcriptional terminators have been tested to determine their respective effects in the expression level of the split prime editor protein or the split prime editor. In some embodiments, the transcriptional terminator used in the present disclosure is a bGH transcriptional terminator. In some embodiments, the rAAV vector further comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE). In certain embodiments, the WPRE is a truncated WPRE sequence, such as “W3.” In some embodiments, the WPRE is inserted 5´ of the transcriptional terminator. Such sequences, when transcribed, create a tertiary structure which enhances expression, in particular, from viral vectors.
[0509] In some embodiments, the vectors used herein may encode the PE fusion proteins, or any of the components thereof (e.g., napDNAbp, linkers, or polymerases). In addition, the vectors used herein may encode the PEgRNAs, and / or the accessory gRNA for second strand nicking. The vectors may be capable of driving expression of one or more coding sequences in a cell. In some embodiments, the cell may be a prokaryotic cell, such as, e.g., a bacterial cell. In some embodiments, the cell may be a eukaryotic cell, such as, e.g., a yeast, plant, insect, or mammalian cell. In some embodiments, the eukaryotic cell may be a mammalian cell. In some embodiments, the eukaryotic cell may be a rodent cell. In some embodiments, the eukaryotic cell may be a human cell. Suitable promoters to drive expression in different types of cells are known in the art. In some embodiments, the promoter may be wild-type. In other embodiments, the promoter may be modified for more efficient or efficacious expression. In yet other embodiments, the promoter may be truncated yet retain its function.For example, the promoter may have a normal size or a reduced size that is suitable for proper packaging of the vector into a virus.
[0510] In some embodiments, the promoters that may be used in the prime editor vectors may be constitutive, inducible, or tissue-specific. In some embodiments, the promoters may be a constitutive promoters. Non-limiting exemplary constitutive promoters include cytomegalovirus immediate early promoter (CMV), simian virus (SV40) promoter, adenovirus major late (MLP) promoter, Rous sarcoma virus (RSV) promoter, mouse mammary tumor virus (MMTV) promoter, phosphoglycerate kinase (PGK) promoter, elongation factor-alpha (EFla) promoter, ubiquitin promoters, actin promoters, tubulin promoters, immunoglobulin promoters, a functional fragment thereof, or a combination of any of the foregoing. In some embodiments, the promoter may be a CMV promoter. In some embodiments, the promoter may be a truncated CMV promoter. In other embodiments, the promoter may be an EFla promoter. In some embodiments, the promoter may be an inducible promoter. Non-limiting exemplary inducible promoters include those inducible by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. In some embodiments, the inducible promoter may be one that has a low basal (non-induced) expression level, such as, e.g., the Tet-On® promoter (Clontech). In some embodiments, the promoter may be a tissue-specific promoter. In some embodiments, the tissue-specific promoter is exclusively or predominantly expressed in liver tissue. Non-limiting exemplary tissue-specific promoters include B29 promoter, CD14 promoter, CD43 promoter, CD45 promoter, CD68 promoter, desmin promoter, elastase- 1 promoter, endoglin promoter, fibronectin promoter, Flt-1 promoter, GFAP promoter, GPIIb promoter, ICAM- 2 promoter, INF-β promoter, Mb promoter, Nphsl promoter, OG-2 promoter, SP-B promoter, SYN1 promoter, and WASP promoter.
[0511] In some embodiments, the first nucleotide sequence, second nucleotide sequence, and optionally, a third nucleotide sequence are on the same nucleic acid vector. In some embodiments, the first nucleotide sequence, second nucleotide sequence, and optionally, the third nucleotide sequence are on different nucleic acid vectors. In some embodiments, the vector is a plasmid. In some embodiments, the nucleic acid vector is a recombinant genome of a rAAV. In some embodiments, the nucleic acid vector is the genome of an adeno- associated virus packaged in a rAAV particle. In some embodiments, the first and / or the second nucleotide sequence is operably linked to a promoter. In some embodiments, the nucleic acid vector further comprises a nucleotide sequence encoding one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) gRNAs operably linked to a promoter. In some embodiments,the promoter is a constitutive promoter. In some embodiments, the promoter is an inducible promoter.
[0512] An inducible promoter of the present disclosure may be induced by (or repressed by) one or more physiological condition(s), such as changes in light, pH, temperature, radiation, osmotic pressure, saline gradients, cell surface binding, and the concentration of one or more extrinsic or intrinsic inducing agent(s). An extrinsic inducer signal or inducing agent may comprise, without limitation, amino acids and amino acid analogs, saccharides and polysaccharides, nucleic acids, protein transcriptional activators and repressors, cytokines, toxins, petroleum-based compounds, metal containing compounds, salts, ions, enzyme substrate analogs, hormones, or combinations thereof.
[0513] Inducible promoters of the present disclosure include any inducible promoter described herein or known to one of ordinary skill in the art. Examples of inducible promoters include, without limitation, chemically / biochemically-regulated and physically- regulated promoters such as alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc)-responsive promoters and other tetracycline-responsive promoter systems, which include a tetracycline repressor protein (tetR), a tetracycline operator sequence (tetO) and a tetracycline transactivator fusion protein (tTA)), steroid- regulated promoters (e.g., promoters based on the rat glucocorticoid receptor, human estrogen receptor, moth ecdysone receptors, and promoters from the steroid / retinoid / thyroid receptor superfamily), metal-regulated promoters (e.g., promoters derived from metallothionein (proteins that bind and sequester metal ions) genes from yeast, mouse and human), pathogenesis-regulated promoters (e.g., induced by salicylic acid, ethylene or benzothiadiazole (BTH)), temperature / heat-inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light responsive promoters from plant cells). Other inducible promoter systems are known in the art and may be used in accordance with the present disclosure.
[0514] In some embodiments, inducible promoters of the present disclosure function in prokaryotic cells (e.g., bacterial cells). Examples of inducible promoters for use prokaryotic cells include, without limitation, bacteriophage promoters (e.g. Pls1con, T3, T7, SP6, PL) and bacterial promoters (e.g., Pbad, PmgrB, Ptrc2, Plac / ara, Ptac, Pm), or hybrids thereof (e.g. PLlacO, PLtetO). Examples of bacterial promoters for use in accordance with the present disclosure include, without limitation, positively regulated E. coli promoters, such as positively regulated σ70 promoters (e.g., inducible pBad / araC promoter, Lux cassette right promoter, modified lamdba Prm promote, plac Or2-62 (positive), pBad / AraC with extra RENsites, pBad, P(Las) TetO, P(Las) CIO, P(Rhl), Pu, FecA, pRE, cadC, hns, pLas, pLux), σS promoters (e.g., Pdps), σ32 promoters (e.g., heat shock), and σ54 promoters (e.g., glnAp2); negatively regulated E. coli promoters such as negatively regulated σ70 promoters (e.g., Promoter (PRM+), modified lamdba Prm promoter, TetR - TetR-4C P(Las) TetO, P(Las) CIO, P(Lac) IQ, RecA_DlexO_DLacO1, dapAp, FecA, Pspac-hy, pcI, plux-cI, plux-lac, CinR, CinL, glucose controlled, modified Pr, modified Prm+, FecA, Pcya, rec A (SOS), Rec A (SOS), EmrR_regulated, BetI_regulated, pLac_lux, pTet_Lac, pLac / Mnt, pTet / Mnt, LsrA / cI, pLux / cI, LacI, LacIQ, pLacIQ1, pLas / cI, pLas / Lux, pLux / Las, pRecA with LexA binding site, reverse BBa_R0011, pLacI / ara-1, pLacIq, rrnB P1, cadC, hns, PfhuA, pBad / araC, nhaA, OmpF, RcnR), σS promoters (e.g., Lutz-Bujard LacO with alternative sigma factor σ38), σ32 promoters (e.g., Lutz-Bujard LacO with alternative sigma factor σ32), and σ54 promoters (e.g., glnAp2); negatively regulated B. subtilis promoters such as repressible B. subtilis σA promoters (e.g., Gram-positive IPTG-inducible, Xyl, hyper-spank) and σB promoters. Other inducible microbial promoters may be used in accordance with the present disclosure.
[0515] In some embodiments, inducible promoters of the present disclosure function in eukaryotic cells (e.g., mammalian cells). Examples of inducible promoters for use eukaryotic cells include, without limitation, chemically-regulated promoters (e.g., alcohol-regulated promoters, tetracycline-regulated promoters, steroid-regulated promoters, metal-regulated promoters, and pathogenesis-related (PR) promoters) and physically-regulated promoters (e.g., temperature-regulated promoters and light-regulated promoters).
[0516] In some embodiments, the prime editor vectors (e.g., including any vectors encoding the prime editor fusion protein and / or the PEgRNAs, and / or the accessory second strand nicking gRNAs) may comprise inducible promoters to start expression only after it is delivered to a target cell. Non-limiting exemplary inducible promoters include those inducible by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. In some embodiments, the inducible promoter may be one that has a low basal (non-induced) expression level, such as, e.g., the Tet-On® promoter (Clontech).
[0517] In additional embodiments, the prime editor vectors (e.g., including any vectors encoding the prime editor fusion protein and / or the PEgRNAs, and / or the accessory second strand nicking gRNAs) may comprise tissue- specific promoters to start expression only after it is delivered into a specific tissue. Non-limiting exemplary tissue-specific promoters include B29 promoter, CD14 promoter, CD43 promoter, CD45 promoter, CD68 promoter, desmin promoter, elastase- 1 promoter, endoglin promoter, fibronectin promoter, Flt-1 promoter,GFAP promoter, GPIIb promoter, ICAM- 2 promoter, INF-β promoter, Mb promoter, Nphsl promoter, OG-2 promoter, SP-B promoter, SYN1 promoter, and WASP promoter.
[0518] In some embodiments, the nucleotide sequence encoding the PEgRNA (or any guide RNAs used in connection with prime editing) may be operably linked to at least one transcriptional or translational control sequence. In some embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to at least one promoter. In some embodiments, the promoter may be recognized by RNA polymerase III (Pol III). Non- limiting examples of Pol III promoters include U6, HI and tRNA promoters. In some embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to a mouse or human U6 promoter. In other embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to a mouse or human HI promoter. In some embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to a mouse or human tRNA promoter. In embodiments with more than one guide RNA, the promoters used to drive expression may be the same or different. In some embodiments, the nucleotide encoding the crRNA of the guide RNA and the nucleotide encoding the tracr RNA of the guide RNA may be provided on the same vector. In some embodiments, the nucleotide encoding the crRNA and the nucleotide encoding the tracr RNA may be driven by the same promoter. In some embodiments, the crRNA and tracr RNA may be transcribed into a single transcript. For example, the crRNA and tracr RNA may be processed from the single transcript to form a double-molecule guide RNA. Alternatively, the crRNA and tracr RNA may be transcribed into a single-molecule guide RNA.
[0519] In some embodiments, the nucleotide sequence encoding the guide RNA may be located on the same vector comprising the nucleotide sequence encoding the PE fusion protein. In some embodiments, expression of the guide RNA and of the PE fusion protein may be driven by their corresponding promoters. In some embodiments, expression of the guide RNA may be driven by the same promoter that drives expression of the PE fusion protein. In some embodiments, the guide RNA and the PE fusion protein transcript may be contained within a single transcript. For example, the guide RNA may be within an untranslated region (UTR) of the Cas9 protein transcript. In some embodiments, the guide RNA may be within the 5' UTR of the PE fusion protein transcript. In other embodiments, the guide RNA may be within the 3' UTR of the PE fusion protein transcript. In some embodiments, the intracellular half-life of the PE fusion protein transcript may be reduced by containing the guide RNA within its 3' UTR and thereby shortening the length of its 3' UTR. In additional embodiments, the guide RNA may be within an intron of the PE fusion proteintranscript. In some embodiments, suitable splice sites may be added at the intron within which the guide RNA is located such that the guide RNA is properly spliced out of the transcript. In some embodiments, expression of the Cas9 protein and the guide RNA in close proximity on the same vector may facilitate more efficient formation of the CRISPR complex.
[0520] The prime editor vector system may comprise one vector, or two vectors, or three vectors, or four vectors, or five vector, or more. In some embodiments, the vector system may comprise one single vector, which encodes both the PE fusion protein and PEgRNA. In other embodiments, the vector system may comprise two vectors, wherein one vector encodes the PE fusion protein and the other encodes the PEgRNA. In additional embodiments, the vector system may comprise three vectors, wherein the third vector encodes the second strand nicking gRNA used in the herein methods.
[0521] In some embodiments, the composition comprising the rAAV particle (in any form contemplated herein) further comprises a pharmaceutically acceptable carrier. In some embodiments, the composition is formulated in appropriate pharmaceutical vehicles for administration to human or animal subjects.
[0522] Some examples of materials which can serve as pharmaceutically-acceptable carriers include: (1) sugars, such as lactose, glucose and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose, and its derivatives, such as sodium carboxymethyl cellulose, methylcellulose, ethyl cellulose, microcrystalline cellulose and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricating agents, such as magnesium stearate, sodium lauryl sulfate and talc; (8) excipients, such as cocoa butter and suppository waxes; (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil; (10) glycols, such as propylene glycol; (11) polyols, such as glycerin, sorbitol, mannitol and polyethylene glycol (PEG); (12) esters, such as ethyl oleate and ethyl laurate; (13) agar; (14) buffering agents, such as magnesium hydroxide and aluminum hydroxide; (15) alginic acid; (16) pyrogen-free water; (17) isotonic saline; (18) Ringer’s solution; (19) ethyl alcohol; (20) pH buffered solutions; (21) polyesters, polycarbonates and / or polyanhydrides; (22) bulking agents, such as polypeptides and amino acids (23) serum component, such as serum albumin, HDL and LDL; (22) C2-C12 alcohols, such as ethanol; and (23) other non-toxic compatible substances employed in pharmaceutical formulations. Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, perfuming agents, preservative and antioxidants can also be present in theformulation. The terms such as “excipient”, “carrier”, “pharmaceutically acceptable carrier” or the like are used interchangeably herein. Recombinant Adeno-associated Virus (rAAV)
[0523] Some aspects of the present disclosure relate to using recombinant adeno- associated virus vectors for the delivery of a split prime editor into a cell. The N-terminal portion of the prime editor and the C-terminal portion of the prime editor are delivered by separate rAAV vectors or particles into the same cell, since the full-length prime editor exceeds the packaging limit of rAAV (~4.9 kb).
[0524] As such, in some embodiments, a composition for delivering the split prime editor protein into a cell (e.g., a mammalian cell, a human cell) is provided. In some embodiments, the composition of the present disclosure comprises: (i) a first recombinant adeno-associated virus (rAAV) particle comprising a first nucleotide sequence encoding a N-terminal portion of split prime editor fused at its C-terminus to an intein-N; and (ii) a second recombinant adeno-associated virus (rAAV) particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor. The rAAV particles of the present disclosure comprise a rAAV vector (i.e., a recombinant genome of the rAAV) encapsidated in the viral capsid proteins.
[0525] In some embodiments, the rAAV vector comprises: (1) a heterologous nucleic acid region comprising the first or second nucleotide sequence encoding the N-terminal portion or C-terminal portion of a split prime editor in any form as described herein, (2) one or more nucleotide sequences comprising a sequence that facilitates expression of the heterologous nucleic acid region (e.g., a promoter), and (3) one or more nucleic acid regions comprising a sequence that facilitate integration of the heterologous nucleic acid region (optionally with the one or more nucleic acid regions comprising a sequence that facilitates expression) into the genome of a cell. In some embodiments, viral sequences that facilitate integration comprise Inverted Terminal Repeat (ITR) sequences. In some embodiments, the first or second nucleotide sequence encoding the N-terminal portion or C-terminal portion of a split prime editor is flanked on each side by an ITR sequence. In some embodiments, the nucleic acid vector further comprises a region encoding an AAV Rep protein as described herein, either contained within the region flanked by ITRs or outside the region. The ITR sequences can be derived from any AAV serotype (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) or can be derived from more than one serotype. In some embodiments, the ITR sequences are derived from AAV2 or AAV6.
[0526] ITR sequences and plasmids containing ITR sequences are known in the art and commercially available (see, e.g., products and services available from Vector Biolabs, Philadelphia, PA; Cellbiolabs, San Diego, CA; Agilent Technologies, Santa Clara, Ca; and Addgene, Cambridge, MA; and Gene delivery to skeletal muscle results in sustained expression and systemic delivery of a therapeutic protein. Kessler PD, Podsakoff GM, Chen X, McQuiston SA, Colosi PC, Matelis LA, Kurtzman GJ, Byrne BJ. Proc Natl Acad Sci USA.1996 Nov 26;93(24):14082-7; and Curtis A. Machida. Methods in Molecular Medicine™. Viral Vectors for Gene Therapy Methods and Protocols. 10.1385 / 1-59259-304- 6:201 © Humana Press Inc.2003. Chapter 10. Targeted Integration by Adeno-Associated Virus. Matthew D. Weitzman, Samuel M. Young Jr., Toni Cathomen and Richard Jude Samulski; U.S. Pat. Nos.5,139,941 and 5,962,313, all of which are incorporated herein by reference). Exemplary ITR sequences are provided in the Description of the Sequences Section (SEQ ID NOs: 10-13).
[0527] In some embodiments, the rAAV vector of the present disclosure comprises one or more regulatory elements to control the expression of the heterologous nucleic acid region (e.g., promoters, transcriptional terminators, and / or other regulatory elements). In some embodiments, the first and / or second nucleotide sequence is operably linked to one or more (e.g., 1, 2, 3, 4, 5, or more) transcriptional terminators. Non-limiting examples of transcriptional terminators that may be used in accordance with the present disclosure include transcription terminators of the bovine growth hormone gene (bGH), human growth hormone gene (hGH), SV40, CW3, ϕ, or combinations thereof. The efficiencies of several transcriptional terminators have been tested to determine their respective effects in the expression level of the split prime editor protein. In some embodiments, the transcriptional terminator used in the present disclosure is a bGH transcriptional terminator. In some embodiments, the rAAV vector further comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE). In some embodiments, the WPRE is inserted 5´ of the transcriptional terminator. napDNAbp
[0528] The prime editors and trans prime editors described herein may comprise a nucleic acid programmable DNA binding protein (napDNAbp).
[0529] In one aspect, a napDNAbp can be associated with or complexed with at least one guide nucleic acid (e.g., guide RNA or a PEgRNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary tothe guide nucleic acid, or a portion thereof (e.g., the spacer of a guide RNA which anneals to the protospacer of the DNA target). In other words, the guide nucleic-acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to complementary sequence of the protospacer in the DNA.
[0530] Any suitable napDNAbp may be used in the prime editors described herein. In various embodiments, the napDNAbp may be any Class 2 CRISPR-Cas system, including any type II, type V, or type VI CRISPR-Cas enzyme. Given the rapid development of CRISPR-Cas as a tool for genome editing, there have been constant developments in the nomenclature used to describe and / or identify CRISPR-Cas enzymes, such as Cas9 and Cas9 orthologs. This application references CRISPR-Cas enzymes with nomenclature that may be old and / or new. The skilled person will be able to identify the specific CRISPR-Cas enzyme being referenced in this Application based on the nomenclature that is used, whether it is old (i.e., “legacy”) or new nomenclature. CRISPR-Cas nomenclature is extensively discussed in Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR Journal, Vol.1. No.5, 2018, the entire contents of which are incorporated herein by reference. The particular CRISPR-Cas nomenclature used in any given instance in this Application is not limiting in any way and the skilled person will be able to identify which CRISPR-Cas enzyme is being referenced.
[0531] For example, the following type II, type V, and type VI Class 2 CRISPR-Cas enzymes have the following art-recognized old (i.e., legacy) and new names. Each of these enzymes, and / or variants thereof, may be used with the prime editors described herein:
[0569] * See Makarova et al., The CRISPR Journal, Vol.1, No.5, 2018
[0570] Without being bound by theory, the mechanism of action of certain napDNAbp contemplated herein includes the step of forming an R-loop whereby the napDNAbp induces the unwinding of a double-strand DNA target, thereby separating the strands in the region bound by the napDNAbp. The guide RNA spacer then hybridizes to the “target strand” at the protospacer sequence. This displaces a “non-target strand” that is complementary to the target strand, which forms the single strand region of the R-loop. In some embodiments, the napDNAbp includes one or more nuclease activities, which then cut the DNA leaving various types of lesions. For example, the napDNAbp may comprises a nuclease activity that cuts the non-target strand at a first location, and / or cuts the target strand at a second location. Depending on the nuclease activity, the target DNA can be cut to form a “double-stranded break” whereby both strands are cut. In other embodiments, the target DNA can be cut at only a single site, i.e., the DNA is “nicked” on one strand. Exemplary napDNAbp with different nuclease activities include “Cas9 nickase” (“nCas9”) and a deactivated Cas9 having no nuclease activities (“dead Cas9” or “dCas9”).
[0571] The below description of various napDNAbps which can be used in connection with the presently disclose prime editors is not meant to be limiting in any way. The prime editors may comprise the canonical SpCas9, or any ortholog Cas9 protein, or any variant Cas9 protein—including any naturally occurring variant, mutant, or otherwise engineered version of Cas9—that is known or which can be made or evolved through a directed evolutionary or otherwise mutagenic process. In various embodiments, the Cas9 or Cas9 variants have a nickase activity, i.e., only cleave of strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variants have inactive nucleases, i.e., are “dead” Cas9 proteins. Other variant Cas9 proteins that may be used are those having a smaller molecularweight than the canonical SpCas9 (e.g., for easier delivery) or having modified or rearranged primary amino acid structure (e.g., the circular permutant formats).
[0572] The prime editors described herein may also comprise Cas9 equivalents, including Cas12a (Cpf1) and Cas12b1 proteins which are the result of convergent evolution. The napDNAbps used herein (e.g., SpCas9, Cas9 variant, or Cas9 equivalents) may also may also contain various modifications that alter / enhance their PAM specifities. Lastly, the application contemplates any Cas9, Cas9 variant, or Cas9 equivalent which has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as a references SpCas9 canonical sequence or a reference Cas9 equivalent (e.g., Cas12a (Cpf1)).
[0573] The napDNAbp can be a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. As outlined above, CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems correct processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc) and a Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-aided processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA target complementary to the spacer. The target strand not complementary to crRNA is first cut endonucleolytically, then trimmed 3´-5′ exonucleolytically. In nature, DNA-binding and cleavage typically requires protein and both RNAs. However, single guide RNAs (“sgRNA”, or simply “gRNA”) can be engineered so as to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821(2012), the entire contents of which is hereby incorporated by reference.
[0574] In some embodiments, the napDNAbp directs cleavage of one or both strands at the location of a target sequence, such as within the target sequence and / or within the complement of the target sequence. In some embodiments, the napDNAbp directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of a target sequence. In some embodiments, a vector encodes a napDNAbp that is mutated to with respect to a corresponding wild-type enzyme such that the mutated napDNAbp lacks the ability to cleave one or both strands of atarget polynucleotide containing a target sequence. For example, an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from S. pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that render Cas9 a nickase include, without limitation, H840A, N854A, and N863A in reference to the canonical SpCas9 sequence, or to equivalent amino acid positions in other Cas9 variants or Cas9 equivalents.
[0575] As used herein, the term “Cas protein” refers to a full-length Cas protein obtained from nature, a recombinant Cas protein having a sequences that differs from a naturally occurring Cas protein, or any fragment of a Cas protein that nevertheless retains all or a significant amount of the requisite basic functions needed for the disclosed methods, i.e., (i) possession of nucleic-acid programmable binding of the Cas protein to a target DNA, and (ii) ability to nick the target DNA sequence on one strand. The Cas proteins contemplated herein embrace CRISPR Cas 9 proteins, as well as Cas9 equivalents, variants (e.g., Cas9 nickase (nCas9) or nuclease inactive Cas9 (dCas9)) homologs, orthologs, or paralogs, whether naturally occurring or non-naturally occurring (e.g., engineered or recombinant), and may include a Cas9 equivalent from anyClass 2 CRISPR system (e.g., type II, V, VI), including Cas12a (Cpf1), Cas12e (CasX), Cas12b1 (C2c1), Cas12b2, Cas12c (C2c3), C2c4, C2c8, C2c5, C2c10, C2c9 Cas13a (C2c2), Cas13d, Cas13c (C2c7), Cas13b (C2c6), and Cas13b. Further Cas-equivalents are described in Makarova et al., “C2c2 is a single- component programmable RNA-guided RNA-targeting CRISPR effector,” Science 2016; 353(6299) and Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?,” The CRISPR Journal, Vol.1. No.5, 2018, the contents of which are incorporated herein by reference.
[0576] The terms “Cas9” or “Cas9 nuclease” or “Cas9 moiety” or “Cas9 domain” embrace any naturally occurring Cas9 from any organism, any naturally-occurring Cas9 equivalent or functional fragment thereof, any Cas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of a Cas9, naturally-occurring or engineered. The term Cas9 is not meant to be particularly limiting and may be referred to as a “Cas9 or equivalent.” Exemplary Cas9 proteins are further described herein and / or are described in the art and are incorporated herein by reference. The present disclosure is unlimited with regard to the particular Cas9 that is employed in the prime editor (PE) of the invention.
[0577] As noted herein, Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K.,primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A.98:4658-4663(2001); “CRISPR RNA maturation by trans- encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602- 607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012), the entire contents of each of which are incorporated herein by reference).
[0578] Examples of Cas9 and Cas9 equivalents are provided as follows; however, these specific examples are not meant to be limiting. The primer editor of the present disclosure may use any suitable napDNAbp, including any suitable Cas9 or Cas9 equivalent. Wild type canonical SpCas9
[0579] In one embodiment, the primer editor constructs described herein may comprise the “canonical SpCas9” nuclease from S. pyogenes, which has been widely used as a tool for genome engineering and is categorized as the type II subgroup of enzymes of the Class 2 CRISPR-Cas systems. This Cas9 protein is a large, multi-domain protein containing two distinct nuclease domains. Point mutations can be introduced into Cas9 to abolish one or both nuclease activities, resulting in a nickase Cas9 (nCas9) or dead Cas9 (dCas9), respectively, that still retains its ability to bind DNA in a sgRNA-programmed manner. In principle, when fused to another protein or domain, Cas9 or variant thereof (e.g., nCas9) can target that protein to virtually any DNA sequence simply by co-expression with an appropriate sgRNA. As used herein, the canonical SpCas9 protein refers to the wild type protein from Streptococcus pyogenes having SEQ ID NOs: 14-20 (See Description of the Sequences Section).
[0580] The prime editors described herein may include canonical SpCas9, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity with a wild type Cas9 sequence provided above. These variants may include SpCas9 variants containing one or more mutations, including any known mutation reported with the SwissProt Accession No. Q99ZW2 (SEQ ID NO: 14) entry, which include:
[0619] Other wild type SpCas9 sequences that may be used in the present disclosure include any of those with amino acid sequences in SEQ ID NOs: 22-35 (See Description of the Sequences)
[0620] The prime editors described herein may include any SpCas9 with amino acid sequences in SEQ ID NOs 14, 17, 19, or any variant thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0621] Wild type Cas9 orthologs
[0622] In other embodiments, the Cas9 protein can be a wild type Cas9 ortholog from another bacterial species different from the canonical Cas9 from S. pyogenes. For example, Cas9 orthologs with amino acid sequences of any SEQ ID NOs: 22-35 (see Description of the SEQUENCES) can be used in connection with the prime editor constructs described in thisspecification. In addition, any variant Cas9 orthologs having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to any one of SEQ ID NOs: 22-35 may also be used with the present prime editors.
[0623] The prime editors described herein may include any Cas9 ortholog with amino acid sequences in SEQ ID NOs: 22-35, or any variants thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0624] The napDNAbp may include any suitable homologs and / or orthologs or naturally occurring enzymes, such as, Cas9. Cas9 homologs and / or orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Preferably, the Cas moiety is configured (e.g, mutagenized, recombinantly engineered, or otherwise obtained from nature) as a nickase, i.e., capable of cleaving only a single strand of the target doubpdditional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, a Cas9 nuclease has an inactive (e.g., an inactivated) DNA cleavage domain, that is, the Cas9 is a nickase. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein as provided by any one of the variants of Table 3. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein as provided by any one of the Cas9 orthologs listed in the Description of the Sequences (SEQ ID NOs: 22-35). Dead Cas9 variant
[0625] In certain embodiments, the prime editors described herein may include a dead Cas9, e.g., dead SpCas9, which has no nuclease activity due to one or more mutations that inactive both nuclease domains of Cas9, namely the RuvC domain (which cleaves the non- protospacer DNA strand) and HNH domain (which cleaves the protospacer DNA strand). The nuclease inactivation may be due to one or mutations that result in one or more substitutions and / or deletions in the amino acid sequence of the encoded protein, or any variants thereof having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0626] As used herein, the term “dCas9” refers to a nuclease-inactive Cas9 or nuclease- dead Cas9, or a functional fragment thereof, and embraces any naturally occurring dCas9 from any organism, any naturally-occurring dCas9 equivalent or functional fragment thereof, any dCas9 homolog, ortholog, or paralog from any organism, and any mutant or variant of a dCas9, naturally-occurring or engineered. The term dCas9 is not meant to be particularly limiting and may be referred to as a “dCas9 or equivalent.” Exemplary dCas9 proteins and method for making dCas9 proteins are further described herein and / or are described in the art and are incorporated herein by reference.
[0627] In other embodiments, dCas9 corresponds to, or comprises in part or in whole, a Cas9 amino acid sequence having one or more mutations that inactivate the Cas9 nuclease activity. In other embodiments, Cas9 variants having mutations other than D10A and H840A are provided which may result in the full or partial inactivate of the endogneous Cas9 nuclease acivity (e.g., nCas9 or dCas9, respectively). Such mutations, by way of example, include other amino acid substitutions at D10 and H820, or other substitutions within the nuclease domains of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain) with reference to a wild type sequence such as Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1). In some embodiments, variants or homologues of Cas9 (e.g., variants of Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 16 or 17))) are provided which are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to NCBI Reference Sequence: NC_017053.1. In some embodiments, variants of dCas9 (e.g., variants of NCBI Reference Sequence: NC_017053.1 (SEQ ID NO: 17)) are provided having amino acid sequences which are shorter, or longer than NC_017053.1 (SEQ ID NO: 17) by about 5 amino acids, by about 10 amino acids, by about 15 amino acids, by about 20 amino acids, by about 25 amino acids, by about 30 amino acids, by about 40 amino acids, by about 50 amino acids, by about 75 amino acids, by about 100 amino acids or more.
[0628] In one embodiment, the dead Cas9 may be based on the canonical SpCas9 sequence of Q99ZW2 and may have the a sequence comprising an amino acid sequence of SEQ ID NO: 36, which comprises a D10X and an H810X, wherein X may be any amino acid, substitutions (underlined and bolded), or a variant of SEQ ID NO: 36 having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0629] In one embodiment, the dead Cas9 may be based on the canonical SpCas9 sequence of Q99ZW2 and may have any one of SEQ ID NOs: 36-37, which comprises a D10A and an H810A substitutions (underlined and bolded), or be a variant of SEQ ID NOs: 36-37 having at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0630] Cas9 nickase variant
[0631] In one embodment, the prime editors described herein comprise a Cas9 nickase. The term “Cas9 nickase” of “nCas9” refers to a variant of Cas9 which is capable of introducing a single-strand break in a double strand DNA molecule target. In some embodiments, the Cas9 nickase comprises only a single functioning nuclease domain. The wild type Cas9 (e.g., the canonical SpCas9) comprises two separate nuclease domains, namely, the RuvC domain (which cleaves the non-protospacer DNA strand) and HNH domain (which cleaves the protospacer DNA strand). In one embodiment, the Cas9 nickase comprises a mutation in the RuvC domain which inactivates the RuvC nuclease activity. For example, mutations in aspartate (D) 10, histidine (H) 983, aspartate (D) 986, or glutamate (E) 762, have been reported as loss-of-function mutations of the RuvC nuclease domain and the creation of a functional Cas9 nickase (e.g., Nishimasu et al., “Crystal structure of Cas9 in complex with guide RNA and target DNA,” Cell 156(5), 935–949, which is incorporated herein by reference). Thus, nickase mutations in the RuvC domain could include D10X, H983X, D986X, or E762X, wherein X is any amino acid other than the wild type amino acid. In certain embodiments, the nickase could be D10A, of H983A, or D986A, or E762A, or a combination thereof.
[0632] In various embodiments, the Cas9 nickase can have a mutation in the RuvC nuclease domain and have any one of amino acid sequences of SEQ ID NOs: 38-53 (See Description of the Sequences, Cas9 nickase sequences), or a variant thereof having an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0633] In another embodiment, the Cas9 nickase comprises a mutation in the HNH domain which inactivates the HNH nuclease activity. For example, mutations in histidine (H) 840 or asparagine (R) 863 have been reported as loss-of-function mutations of the HNH nuclease domain and the creation of a functional Cas9 nickase (e.g., Nishimasu et al., “Crystal structure of Cas9 in complex with guide RNA and target DNA,” Cell 156(5), 935– 949, which is incorporated herein by reference). Thus, nickase mutations in the HNH domain could include H840X and R863X, wherein X is any amino acid other than the wild typeamino acid. In certain embodiments, the nickase could be H840A or R863A or a combination thereof.
[0634] In various embodiments, the Cas9 nickase can have a mutation in the HNH nuclease domain and have any one of amino acid sequences in SEQ ID NOs: 38-53, or a variant thereof having an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0635] In some embodiments, the N-terminal methionine is removed from a Cas9 nickase, or from any Cas9 variant, ortholog, or equivalent disclosed or contemplated herein. For example, in some embodiments, the methionine-minus Cas9 nickase comprises an amino acid sequence of any one of SEQ ID NOs: 50-53, or a variant thereof having an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. In some embodiments, the methionine-minus Cas9 nickase comprises an amino acid sequence of any one of SEQ ID NOs: 38-49, or a variant thereof having an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto. Other Cas9 variants
[0636] Besides dead Cas9 and Cas9 nickase variants, the Cas9 proteins used herein may also include other “Cas9 variants” having at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference Cas9 protein, including any wild type Cas9, or mutant Cas9 (e.g., a dead Cas9 or Cas9 nickase), or fragment Cas9, or circular permutant Cas9, or other variant of Cas9 disclosed herein or known in the art. In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to a reference Cas9. In some embodiments, the Cas9 variant comprises a fragment of a reference Cas9 (e.g., a gRNA binding domain or a DNA-cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild type Cas9. In some embodiments, the fragment is is at least 30%, at least 35%, at least 40%, at least45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild type Cas9 (e.g., SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84).
[0637] In some embodiments, the disclosure also may utilize Cas9 fragments which retain their functionality and which are fragments of any herein disclosed Cas9 protein. In some embodiments, the Cas9 fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0638] In various embodiments, the prime editors disclosed herein may comprise one of the Cas9 variants described as follows, or a Cas9 variant thereof having at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference Cas9 variants. Small-sized Cas9 variants
[0639] In some embodiments, the prime editors contemplated herein can include a Cas9 protein that is of smaller molecular weight than the canonical SpCas9 sequence. In some embodiments, the smaller-sized Cas9 variants may facilitate delivery to cells, e.g., by an expression vector, nanoparticle, or other means of delivery. In certain embodiments, the smaller-sized Cas9 variants can include enzymes categorized as type II enzymes of the Class 2 CRISPR-Cas systems. In some embodiments, the smaller-sized Cas9 variants can include enzymes categorized as type V enzymes of the Class 2 CRISPR-Cas systems. In other embodiments, the smaller-sized Cas9 variants can include enzymes categorized as type VI enzymes of the Class 2 CRISPR-Cas systems.
[0640] The canonical SpCas9 protein is 1368 amino acids in length and has a predicted molecular weight of 158 kilodaltons. The term “small-sized Cas9 variant”, as used herein, refers to any Cas9 variant—naturally occurring, engineered, or otherwise—that is less than at least 1300 amino acids, or at least less than 1290 amino acids, or than less than 1280 amino acids, or less than 1270 amino acid, or less than 1260 amino acid, or less than 1250 amino acids, or less than 1240 amino acids, or less than 1230 amino acids, or less than 1220 amino acids, or less than 1210 amino acids, or less than 1200 amino acids, or less than 1190 aminoacids, or less than 1180 amino acids, or less than 1170 amino acids, or less than 1160 amino acids, or less than 1150 amino acids, or less than 1140 amino acids, or less than 1130 amino acids, or less than 1120 amino acids, or less than 1110 amino acids, or less than 1100 amino acids, or less than 1050 amino acids, or less than 1000 amino acids, or less than 950 amino acids, or less than 900 amino acids, or less than 850 amino acids, or less than 800 amino acids, or less than 750 amino acids, or less than 700 amino acids, or less than 650 amino acids, or less than 600 amino acids, or less than 550 amino acids, or less than 500 amino acids, but at least larger than about 400 amino acids and retaining the required functions of the Cas9 protein. The Cas9 variants can include those categorized as type II, type V, or type VI enzymes of the Class 2 CRISPR-Cas system.
[0641] In various embodiments, the prime editors disclosed herein may comprise any one of the small-sized Cas9 variants comprising an amino acid sequence of any one of SEQ ID NOs: 54-59 (see Description of the Sequences, small sized Cas9 variants), or a Cas9 variant thereof having at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any reference small-sized Cas9 protein. Cas9 equivalents
[0642] In some embodiments, the prime editors described herein can include any Cas9 equivalent. As used herein, the term “Cas9 equivalent” is a broad term that encompasses any napDNAbp protein that serves the same function as Cas9 in the present prime editors despite that its amino acid primary sequence and / or its three-dimensional structure may be different and / or unrelated from an evolutionary standpoint. Thus, while Cas9 equivalents include any Cas9 ortholog, homolog, mutant, or variant described or embraced herein that are evolutionarily related, the Cas9 equivalents also embrace proteins that may have evolved through convergent evolution processes to have the same or similar function as Cas9, but which do not necessarily have any similarity with regard to amino acid sequence and / or three dimensional structure. The prime editors described here embrace any Cas9 equivalent that would provide the same or similar function as Cas9 despite that the Cas9 equivalent may be based on a protein that arose through convergent evolution. For instance, if Cas9 refers to a type II enzyme of the CRISPR-Cas system, a Cas9 equivalent can refer to a type V or type VI enzyme of the CRISPR-Cas system.
[0643] For example, Cas12e (CasX) is a Cas9 equivalent that reportedly has the same function as Cas9 but which evolved through convergent evolution. Thus, the Cas12e (CasX) protein described in Liu et al., “CasX enzymes comprises a distinct family of RNA-guided genome editors,” Nature, 2019, Vol.566: 218-223, is contemplated to be used with the prime editors described herein. In addition, any variant or modification of Cas12e (CasX) is conceivable and within the scope of the present disclosure.
[0644] Cas9 is a bacterial enzyme that evolved in a wide variety of species. However, the Cas9 equivalents contemplated herein may also be obtained from archaea, which constitute a domain and kingdom of single-celled prokaryotic microbes different from bacteria.
[0645] In some embodiments, Cas9 equivalents may refer to Cas12e (CasX) or Cas12d (CasY), which have been described in, for example, Burstein et al., “New CRISPR–Cas systems from uncultivated microbes.” Cell Res.2017 Feb 21. doi: 10.1038 / cr.2017.21, the entire contents of which is hereby incorporated by reference. Using genome-resolved metagenomics, a number of CRISPR–Cas systems were identified, including the first reported Cas9 in the archaeal domain of life. This divergent Cas9 protein was found in little- studied nanoarchaea as part of an active CRISPR–Cas system. In bacteria, two previously unknown systems were discovered, CRISPR– Cas12e and CRISPR– Cas12d, which are among the most compact systems yet discovered. In some embodiments, Cas9 refers to Cas12e, or a variant of Cas12e. In some embodiments, Cas9 refers to a Cas12d, or a variant of Cas12d. It should be appreciated that other RNA-guided DNA binding proteins may be used as a nucleic acid programmable DNA binding protein (napDNAbp), and are within the scope of this disclosure. Also see Liu et al., “CasX enzymes comprises a distinct family of RNA-guided genome editors,” Nature, 2019, Vol.566: 218-223. Any of these Cas9 equivalents are contemplated.
[0646] In some embodiments, the Cas9 equivalent comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp is a naturally-occurring Cas12e (CasX) or Cas12d (CasY) protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a wild-type Cas moiety or any Cas moiety provided herein.
[0647] In various embodiments, the nucleic acid programmable DNA binding proteins include, without limitation, Cas9 (e.g., dCas9 and nCas9), Cas12e (CasX), Cas12d (CasY), Cas12a (Cpf1), Cas12b1 (C2c1), Cas13a (C2c2), Cas12c (C2c3), Argonaute, , and Cas12b1. One example of a nucleic acid programmable DNA-binding protein that has different PAM specificity than Cas9 is Clustered Regularly Interspaced Short Palindromic Repeats from Prevotella and Francisella 1 (i.e, Cas12a (Cpf1)). Similar to Cas9, Cas12a (Cpf1) is also a Class 2 CRISPR effector, but it is a member of type V subgroup of enzymes, rather than the type II subgroup. It has been shown that Cas12a (Cpf1) mediates robust DNA interference with features distinct from Cas9. Cas12a (Cpf1) is a single RNA-guided endonuclease lacking tracrRNA, and it utilizes a T-rich protospacer-adjacent motif (TTN, TTTN, or YTN). Moreover, Cpf1 cleaves DNA via a staggered DNA double-stranded break. Out of 16 Cpf1- family proteins, two enzymes from Acidaminococcus and Lachnospiraceae are shown to have efficient genome-editing activity in human cells. Cpf1 proteins are known in the art and have been described previously, for example Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, p.949-962; the entire contents of which is hereby incorporated by reference.
[0648] In still other embodiments, the Cas protein may include any CRISPR associated protein, including but not limited to, Cas12a, Cas12b1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, and preferably comprising a nickase mutation (e.g., a mutation corresponding to the D10A mutation of the wild type Cas9 polypeptide of SEQ ID NO: 14).
[0649] In various other embodiments, the napDNAbp can be any of the following proteins: a Cas9, a Cas12a (Cpf1), a Cas12e (CasX), a Cas12d (CasY), a Cas12b1 (C2c1), a Cas13a (C2c2), a Cas12c (C2c3), a GeoCas9, a CjCas9, a Cas12g, a Cas12h, a Cas12i, a Cas13b, a Cas13c, a Cas13d, a Cas14, a Csn2, an xCas9, an SpCas9-NG, a circularly permuted Cas9, or an Argonaute (Ago) domain, or a variant thereof.
[0650] Exemplary Cas9 equivalent protein sequences can be found in the Description of Sequences (SEQ ID NOs: 60-69).
[0651] The prime editors described herein may also comprise Cas12a (Cpf1) (dCpf1) variants that may be used as a guide nucleotide sequence-programmable DNA-binding protein domain. The Cas12a (Cpf1) protein has a RuvC-like endonuclease domain that issimilar to the RuvC domain of Cas9 but does not have a HNH endonuclease domain, and the N-terminal of Cas12a (Cpf1) does not have the alfa-helical recognition lobe of Cas9. It was shown in Zetsche et al., Cell, 163, 759–771, 2015 (which is incorporated herein by reference) that, the RuvC-like domain of Cas12a (Cpf1) is responsible for cleaving both DNA strands and inactivation of the RuvC-like domain inactivates Cas12a (Cpf1) nuclease activity.
[0652] In some embodiments, the napDNAbp is a single effector of a microbial CRISPR- Cas system. Single effectors of microbial CRISPR-Cas systems include, without limitation, Cas9, Cas12a (Cpf1), Cas12b1 (C2c1), Cas13a (C2c2), and Cas12c (C2c3). Typically, microbial CRISPR-Cas systems are divided into Class 1 and Class 2 systems. Class 1 systems have multisubunit effector complexes, while Class 2 systems have a single protein effector. For example, Cas9 and Cas12a (Cpf1) are Class 2 effectors. In addition to Cas9 and Cas12a (Cpf1), three distinct Class 2 CRISPR-Cas systems (Cas12b1, Cas13a, and Cas12c) have been described by Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov 5; 60(3): 385–397, the entire contents of which are hereby incorporated by reference.
[0653] Effectors of two of the systems, Cas12b1 and Cas12c, contain RuvC-like endonuclease domains related to Cas12a. A third system, Cas13a contains an effector with two predicated HEPN RNase domains. Production of mature CRISPR RNA is tracrRNA- independent, unlike production of CRISPR RNA by Cas12b1. Cas12b1 depends on both CRISPR RNA and tracrRNA for DNA cleavage. Bacterial Cas13a has been shown to possess a unique RNase activity for CRISPR RNA maturation distinct from its RNA-activated single- stranded RNA degradation activity. These RNase functions are different from each other and from the CRISPR RNA-processing behavior of Cas12a. See, e.g., East-Seletsky, et al., “Two distinct RNase activities of CRISPR-Cas13a enable guide-RNA processing and RNA detection”, Nature, 2016 Oct 13;538(7624):270-273, the entire contents of which are hereby incorporated by reference. In vitro biochemical analysis of Cas13a in Leptotrichia shahii has shown that Cas13a is guided by a single CRISPR RNA and can be programed to cleave ssRNA targets carrying complementary protospacers. Catalytic residues in the two conserved HEPN domains mediate cleavage. Mutations in the catalytic residues generate catalytically inactive RNA-binding proteins. See e.g., Abudayyeh et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector”, Science, 2016 Aug 5; 353(6299), the entire contents of which are hereby incorporated by reference.
[0654] The crystal structure of Alicyclobaccillus acidoterrastris Cas12b1 (AacC2c1) has been reported in complex with a chimeric single-molecule guide RNA (sgRNA). See e.g.,Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cell, 2017 Jan 19;65(2):310-322, the entire contents of which are hereby incorporated by reference. The crystal structure has also been reported in Alicyclobacillus acidoterrestris C2c1 bound to target DNAs as ternary complexes. See e.g., Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, 2016 Dec 15;167(7):1814-1828, the entire contents of which are hereby incorporated by reference. Catalytically competent conformations of AacC2c1, both with target and non-target DNA strands, have been captured independently positioned within a single RuvC catalytic pocket, with C2c1-mediated cleavage resulting in a staggered seven- nucleotide break of target DNA. Structural comparisons between C2c1 ternary complexes and previously identified Cas9 and Cpf1 counterparts demonstrate the diversity of mechanisms used by CRISPR-Cas9 systems.
[0655] In some embodiments, the napDNAbp may be a C2c1, a C2c2, or a C2c3 protein. In some embodiments, the napDNAbp is a C2c1 protein. In some embodiments, the napDNAbp is a Cas13a protein. In some embodiments, the napDNAbp is a Cas12c protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring Cas12b1 (C2c1), Cas13a (C2c2), or Cas12c (C2c3) protein. In some embodiments, the napDNAbp is a naturally-occurring Cas12b1 (C2c1), Cas13a (C2c2), or Cas12c (C2c3) protein. Cas9 circular permutants
[0656] In various embodiments, the prime editors disclosed herein may comprise a circular permutant of Cas9.
[0657] The term “circularly permuted Cas9” or “circular permutant” of Cas9 or “CP- Cas9”) refers to any Cas9 protein, or variant thereof, that occurs or has been modify to engineered as a circular permutant variant, which means the N-terminus and the C-terminus of a Cas9 protein (e.g., a wild type Cas9 protein) have been topically rearranged. Such circularly permuted Cas9 proteins, or variants thereof, retain the ability to bind DNA when complexed with a guide RNA (gRNA). See, Oakes et al., “Protein Engineering of Cas9 for enhanced function,” Methods Enzymol, 2014, 546: 491–511 and Oakes et al., “CRISPR-Cas9 Circular Permutants as Programmable Scaffolds for Genome Modification,” Cell, January 10, 2019, 176: 254-267, each of are incorporated herein by reference. The instant disclosurecontemplates any previously known CP-Cas9 or use a new CP-Cas9 so long as the resulting circularly permuted protein retains the ability to bind DNA when complexed with a guide RNA (gRNA).
[0658] Any of the Cas9 proteins described herein, including any variant, ortholog, or naturally occurring Cas9 or equivalent thereof, may be reconfigured as a circular permutant variant.
[0659] In various embodiments, the circular permutants of Cas9 may have the following structure:
[0660] N-terminus-[original C-terminus] – [optional linker] – [original N-terminus]-C- terminus.
[0661] As an example, the present disclosure contemplates the following circular permutants of canonical S. pyogenes Cas9 (1368 amino acids of UniProtKB - Q99ZW2 (CAS9_STRP1) (numbering is based on the amino acid position in SEQ ID NO: 14)):
[0662] N-terminus-[1268-1368]-[optional linker]-[1-1267]-C-terminus;
[0663] N-terminus-[1168-1368]-[optional linker]-[1-1167]-C-terminus;
[0664] N-terminus-[1068-1368]-[optional linker]-[1-1067]-C-terminus;
[0665] N-terminus-[968-1368]-[optional linker]-[1-967]-C-terminus;
[0666] N-terminus-[868-1368]-[optional linker]-[1-867]-C-terminus;
[0667] N-terminus-[768-1368]-[optional linker]-[1-767]-C-terminus;
[0668] N-terminus-[668-1368]-[optional linker]-[1-667]-C-terminus;
[0669] N-terminus-[568-1368]-[optional linker]-[1-567]-C-terminus;
[0670] N-terminus-[468-1368]-[optional linker]-[1-467]-C-terminus;
[0671] N-terminus-[368-1368]-[optional linker]-[1-367]-C-terminus;
[0672] N-terminus-[268-1368]-[optional linker]-[1-267]-C-terminus;
[0673] N-terminus-[168-1368]-[optional linker]-[1-167]-C-terminus;
[0674] N-terminus-[68-1368]-[optional linker]-[1-67]-C-terminus; or
[0675] N-terminus-[10-1368]-[optional linker]-[1-9]-C-terminus, or the corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc).
[0676] In particular embodiments, the circular permuant Cas9 has the following structure (based on S. pyogenes Cas9 (1368 amino acids of UniProtKB - Q99ZW2 (CAS9_STRP1) (numbering is based on the amino acid position in SEQ ID NO: 14):
[0677] N-terminus-[102-1368]-[optional linker]-[1-101]-C-terminus;
[0678] N-terminus-[1028-1368]-[optional linker]-[1-1027]-C-terminus;
[0679] N-terminus-[1041-1368]-[optional linker]-[1-1043]-C-terminus;
[0680] N-terminus-[1249-1368]-[optional linker]-[1-1248]-C-terminus; or
[0681] N-terminus-[1300-1368]-[optional linker]-[1-1299]-C-terminus, or the corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc).
[0682] In still other embodiments, the circular permuant Cas9 has the following structure (based on S. pyogenes Cas9 (1368 amino acids of UniProtKB - Q99ZW2 (CAS9_STRP1) (numbering is based on the amino acid position in SEQ ID NO: 14):
[0683] N-terminus-[103-1368]-[optional linker]-[1-102]-C-terminus;
[0684] N-terminus-[1029-1368]-[optional linker]-[1-1028]-C-terminus;
[0685] N-terminus-[1042-1368]-[optional linker]-[1-1041]-C-terminus;
[0686] N-terminus-[1250-1368]-[optional linker]-[1-1249]-C-terminus; or
[0687] N-terminus-[1301-1368]-[optional linker]-[1-1300]-C-terminus, or the corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc).
[0688] In some embodiments, the circular permutant can be formed by linking a C- terminal fragment of a Cas9 to an N-terminal fragment of a Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, The C-terminal fragment may correspond to the C-terminal 95% or more of the amino acids of a Cas9 (e.g., amino acids about 1300-1368), or the C-terminal 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of a Cas9 (e.g., any one of SEQ ID NOs:70-79). The N-terminal portion may correspond to the N-terminal 95% or more of the amino acids of a Cas9 (e.g., amino acids about 1-1300), or the N-terminal 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of a Cas9 (e.g., of SEQ ID NOs: 14).
[0689] In some embodiments, the circular permutant can be formed by linking a C- terminal fragment of a Cas9 to an N-terminal fragment of a Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, the C-terminal fragment that is rearranged to the N-terminus, includes or corresponds to the C-terminal 30% or less of the amino acids of a Cas9 (e.g., amino acids 1012-1368 of SEQ ID NO: 14). In some embodiments, the C-terminal fragment that is rearranged to the N-terminus, includes or corresponds to the C-terminal 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the amino acids of a Cas9 (e.g., the Cas9 of SEQ ID NO: 14). In some embodiments, the C-terminal fragment that is rearranged to the N-terminus, includes orcorresponds to the C-terminal 410 residues or less of a Cas9 (e.g., the Cas9 of SEQ ID NO: 14). In some embodiments, the C-terminal portion that is rearranged to the N-terminus, includes or corresponds to the C-terminal 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 residues of a Cas9 (e.g., the Cas9 of SEQ ID NO: 14). In some embodiments, the C-terminal portion that is rearranged to the N- terminus, includes or corresponds to the C-terminal 357, 341, 328, 120, or 69 residues of a Cas9 (e.g., the Cas9 of SEQ ID NO: 14).
[0690] In other embodiments, circular permutant Cas9 variants may be defined as a topological rearrangement of a Cas9 primary structure based on the following method, which is based on S. pyogenes Cas9 of SEQ ID NO: 14: (a) selecting a circular permutant (CP) site corresponding to an internal amino acid residue of the Cas9 primary structure, which dissects the original protein into two halves: an N-terminal region and a C-terminal region; (b) modifying the Cas9 protein sequence (e.g., by genetic engineering techniques) by moving the original C-terminal region (comprising the CP site amino acid) to preceed the original N- terminal region, thereby forming a new N-terminus of the Cas9 protein that now begins with the CP site amino acid residue. The CP site can be located in any domain of the Cas9 protein, including, for example, the helical-II domain, the RuvCIII domain, or the CTD domain. For example, the CP site may be located (relative the S. pyogenes Cas9 of SEQ ID NO: 14) at original amino acid residue 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282. Thus, once relocated to the N-terminus, original amino acid 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282 would become the new N-terminal amino acid. Nomenclature of these CP-Cas9 proteins may be referred to as Cas9-CP181, Cas9-CP199, Cas9-CP230, Cas9-CP270, Cas9-CP310, Cas9-CP1010, Cas9- CP1016, Cas9-CP1023, Cas9-CP1029, Cas9-CP1041, Cas9-CP1247, Cas9-CP1249, and Cas9-CP1282, respectively. This description is not meant to be limited to making CP variants from SEQ ID NO: 14, but may be implemented to make CP variants in any Cas9 sequence (e.g., SEQ ID NOs: 14), either at CP sites that correspond to these positions, or at other CP sites entirely. This description is not meant to limit the specific CP sites in any way. Virtually any CP site may be used to form a CP-Cas9 variant.
[0691] Exemplary CP-Cas9 amino acid sequences, based on the Cas9 of SEQ ID NO: 14, are shown in SEQ ID NOs: 70-74 (see Description of Sequences, Circular permutants) in which linker sequences are indicated by underlining and optional methionine (M) residues are indicated in bold. It should be appreciated that the disclosure provides CP-Cas9 sequencesthat do not include a linker sequence or that include different linker sequences. It should be appreciated that CP-Cas9 sequences may be based on Cas9 sequences other than that of SEQ ID NO: 14 and any examples provided herein are not meant to be limiting. Exempalry CP- Cas9 sequences may be found in the Description of Sequences (SEQ ID NOs: 70-74, Cas9 Circular Permutants).
[0692] The Cas9 circular permutants that may be useful in the prime editing constructs described herein. Exemplary C-terminal fragments of Cas9, based on the Cas9 of SEQ ID NO: 14, which may be rearranged to an N-terminus of Cas9, are provided in the Description of Sequences (SEQ ID NOs: 75-79, Cas9 circular permutants). It should be appreciated that such C-terminal fragments of Cas9 are exemplary and are not meant to be limiting. These exemplary CP-Cas9 fragments may have an amino acid sequences of any one of SEQ ID NOs: 70-79. Cas9 variants with modified PAM specificities
[0693] The prime editors of the present disclosure may also comprise Cas9 variants with modified PAM specificities. Some aspects of this disclosure provide Cas9 proteins that exhibit activity on a target sequence that does not comprise the canonical PAM (5′-NGG-3′, where N is A, C, G, or T) at its 3′-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5′-NGG-3′ PAM sequence at its 3′-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NNG- 3´ PAM sequence at its 3′-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5′-NNA-3′ PAM sequence at its 3ʹ-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5′-NNC-3′ PAM sequence at its 3ʹ-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NNT-3´ PAM sequence at its 3ʹ-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NGT-3´ PAM sequence at its 3ʹ-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NGA-3´ PAM sequence at its 3ʹ-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NGC-3´ PAM sequence at its 3ʹ-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´- NAA-3´ PAM sequence at its 3´-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NAC-3´ PAM sequence at its 3′-end. In some embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NAT-3´PAM sequence at its 3´-end. In still other embodiments, the Cas9 protein exhibits activity on a target sequence comprising a 5´-NAG-3´ PAM sequence at its 3´-end.
[0694] It should be appreciated that any of the amino acid mutations described herein, (e.g., A262T) from a first amino acid residue (e.g., A) to a second amino acid residue (e.g., T) may also include mutations from the first amino acid residue to an amino acid residue that is similar to (e.g., conserved) the second amino acid residue. For example, mutation of an amino acid with a hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan) may be a mutation to a second amino acid with a different hydrophobic side chain (e.g., alanine, valine, isoleucine, leucine, methionine, phenylalanine, tyrosine, or tryptophan). For example, a mutation of an alanine to a threonine (e.g., a A262T mutation) may also be a mutation from an alanine to an amino acid that is similar in size and chemical properties to a threonine, for example, serine. As another example, mutation of an amino acid with a positively charged side chain (e.g., arginine, histidine, or lysine) may be a mutation to a second amino acid with a different positively charged side chain (e.g., arginine, histidine, or lysine). As another example, mutation of an amino acid with a polar side chain (e.g., serine, threonine, asparagine, or glutamine) may be a mutation to a second amino acid with a different polar side chain (e.g., serine, threonine, asparagine, or glutamine). Additional similar amino acid pairs include, but are not limited to, the following: phenylalanine and tyrosine; asparagine and glutamine; methionine and cysteine; aspartic acid and glutamic acid; and arginine and lysine. The skilled artisan would recognize that such conservative amino acid substitutions will likely have minor effects on protein structure and are likely to be well tolerated without compromising function. In some embodiments, any amino of the amino acid mutations provided herein from one amino acid to a threonine may be an amino acid mutation to a serine. In some embodiments, any amino of the amino acid mutations provided herein from one amino acid to an arginine may be an amino acid mutation to a lysine. In some embodiments, any amino of the amino acid mutations provided herein from one amino acid to an isoleucine, may be an amino acid mutation to an alanine, valine, methionine, or leucine. In some embodiments, any amino of the amino acid mutations provided herein from one amino acid to a lysine may be an amino acid mutation to an arginine. In some embodiments, any amino of the amino acid mutations provided herein from one amino acid to an aspartic acid may be an amino acid mutation to a glutamic acid or asparagine. In some embodiments, any amino of the amino acid mutations provided herein from one amino acid to a valine may be an amino acid mutation to an alanine, isoleucine, methionine, or leucine. In some embodiments, any amino of the aminoacid mutations provided herein from one amino acid to a glycine may be an amino acid mutation to an alanine. It should be appreciated, however, that additional conserved amino acid residues would be recognized by the skilled artisan and any of the amino acid mutations to other conserved amino acid residues are also within the scope of this disclosure.
[0695] In some embodiments, the Cas9 protein comprises a combination of mutations that exhibit activity on a target sequence comprising a 5´-NAA-3´ PAM sequence at its 3´- end. In some embodiments, the combination of mutations are present in any one of the clones listed in Table 1. In some embodiments, the combination of mutations are conservative mutations of the clones listed in Table 1. In some embodiments, the Cas9 protein comprises the combination of mutations of any one of the Cas9 clones listed in Table 1.
[0696] Table 1: NAA PAM Clones
[0697] In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein as provided by any one of the variants of Table 1. In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein as provided by any one of the variants of Table 1.
[0698] In some embodiments, the Cas9 protein exhibits an increased activity on a target sequence that does not comprise the canonical PAM (5´-NGG-3´) at its 3´ end as compared to Streptococcus pyogenes Cas9 as provided by SEQ ID NO: 14. In some embodiments, the Cas9 protein exhibits an activity on a target sequence having a 3ʹ end that is not directly adjacent to the canonical PAM sequence (5´-NGG-3´) that is at least 5-fold increased as compared to the activity of Streptococcus pyogenes Cas9 as provided by SEQ ID NO: 14 on the same target sequence. In some embodiments, the Cas9 protein exhibits an activity on a target sequence that is not directly adjacent to the canonical PAM sequence (5´-NGG-3´) that is at least 10-fold, at least 50-fold, at least 100-fold, at least 500-fold, at least 1,000-fold, at least 5,000-fold, at least 10,000-fold, at least 50,000-fold, at least 100,000-fold, at least 500,000-fold, or at least 1,000,000-fold increased as compared to the activity of Streptococcus pyogenes as provided by SEQ ID NO: 14 on the same target sequence. In some embodiments, the 3ʹ end of the target sequence is directly adjacent to an AAA, GAA, CAA, or TAA sequence. In some embodiments, the Cas9 protein comprises a combination of mutations that exhibit activity on a target sequence comprising a 5´-NAC-3´ PAM sequence at its 3ʹ-end. In some embodiments, the combination of mutations are present in any one of the clones listed in Table 2. In some embodiments, the combination of mutations are conservative mutations of the clones listed in Table 2. In some embodiments, the Cas9 protein comprises the combination of mutations of any one of the Cas9 clones listed in Table 2.
[0699] Table 2: NAC PAM Clones
[0700] In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of a Cas9 protein as provided by any one of the variants of Table 2. In some embodiments, the Cas9 protein comprises an amino acidsequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of a Cas9 protein as provided by any one of the variants of Table 2.
[0701] In some embodiments, the Cas9 protein exhibits an increased activity on a target sequence that does not comprise the canonical PAM (5´-NGG-3´) at its 3ʹ end as compared to Streptococcus pyogenes Cas9 as provided by SEQ ID NO: 14. In some embodiments, the Cas9 protein exhibits an activity on a target sequence having a 3ʹ end that is not directly adjacent to the canonical PAM sequence (5´-NGG-3´) that is at least 5-fold increased as compared to the activity of Streptococcus pyogenes Cas9 as provided by SEQ ID NO: 14 on the same target sequence. In some embodiments, the Cas9 protein exhibits an activity on a target sequence that is not directly adjacent to the canonical PAM sequence (5´-NGG-3´) that is at least 10-fold, at least 50-fold, at least 100-fold, at least 500-fold, at least 1,000-fold, at least 5,000-fold, at least 10,000-fold, at least 50,000-fold, at least 100,000-fold, at least 500,000-fold, or at least 1,000,000-fold increased as compared to the activity of Streptococcus pyogenes as provided by SEQ ID NO: 14 on the same target sequence. In some embodiments, the 3ʹ end of the target sequence is directly adjacent to an AAC, GAC, CAC, or TAC sequence.
[0702] In some embodiments, the Cas9 protein comprises a combination of mutations that exhibit activity on a target sequence comprising a 5´-NAT-3´ PAM sequence at its 3ʹ- end. In some embodiments, the combination of mutations are present in any one of the clones listed in Table 3. In some embodiments, the combination of mutations are conservative mutations of the clones listed in Table 3. In some embodiments, the Cas9 protein comprises the combination of mutations of any one of the Cas9 clones listed in Table 3.
[0703] Table 3: NAT PAM Clones
[0704] The above description of various napDNAbps which can be used in connection with the presently disclose prime editors is not meant to be limiting in any way. The prime editors may comprise the canonical SpCas9, or any ortholog Cas9 protein, or any variant Cas9 protein—including any naturally occurring variant, mutant, or otherwise engineered version of Cas9—that is known or which can be made or evolved through a directed evolutionary or otherwise mutagenic process. In various embodiments, the Cas9 or Cas9 varants have a nickase activity, i.e., only cleave of strand of the target DNA sequence. In other embodiments, the Cas9 or Cas9 variants have inactive nucleases, i.e., are “dead” Cas9 proteins. Other variant Cas9 proteins that may be used are those having a smaller molecular weight than the canonical SpCas9 (e.g., for easier delivery) or having modified or rearrangedprimary amino acid structure (e.g., the circular permutant formats). The prime editors described herein may also comprise Cas9 equivalents, including Cas12a / Cpf1 and Cas12b proteins which are the result of convergent evolution. The napDNAbps used herein (e.g., SpCas9, Cas9 variant, or Cas9 equivalents) may also may also contain various modifications that alter / enhance their PAM specifities. Lastly, the application contemplates any Cas9, Cas9 variant, or Cas9 equivalent which has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.9% sequence identity to a reference Cas9 sequence, such as a references SpCas9 canonical sequences or a reference Cas9 equivalent (e.g., Cas12a / Cpf1).
[0705] In a particular embodiment, the Cas9 variant having expanded PAM capabilities is SpCas9 (H840A) VRQR which comprises an amino acid sequence of SEQ ID NO: 80 (with the V, R, Q, R substitutions relative to the SpCas9 (H840A) of SEQ ID NO: 47). In addition, the methionine residue in SpCas9 (H840) was removed for SpCas9 (H840A) VRQR.
[0706] In another particular embodiment, the Cas9 variant having expanded PAM capabilities is SpCas9 (H840A) VRER which comprises an amino acid sequence of SEQ ID NO: 81 (with the V, R, E, R substitutions relative to the SpCas9 (H840A) of SEQ ID NO: 47). In addition, the methionine residue in SpCas9 (H840) was removed for SpCas9 (H840A) VRER.
[0707] In some embodiments, the napDNAbp that functions with a non-canonical PAM sequence is an Argonaute protein. One example of such a nucleic acid programmable DNA binding protein is an Argonaute protein from Natronobacterium gregoryi (NgAgo). NgAgo is a ssDNA-guided endonuclease. NgAgo binds 5′ phosphorylated ssDNA of ~24 nucleotides (gDNA) to guide it to its target site and will make DNA double-strand breaks at the gDNA site. In contrast to Cas9, the NgAgo–gDNA system does not require a protospacer-adjacent motif (PAM). Using a nuclease inactive NgAgo (dNgAgo) can greatly expand the bases that may be targeted. The characterization and use of NgAgo have been described in Gao et al., Nat Biotechnol., 2016 Jul;34(7):768-73. PubMed PMID: 27136078; Swarts et al., Nature. 507(7491) (2014):258-61; and Swarts et al., Nucleic Acids Res. 43(10) (2015):5120-9, each of which is incorporated herein by reference.
[0708] In some embodiments, the napDNAbp is a prokaryotic homolog of an Argonaute protein. Prokaryotic homologs of Argonaute proteins are known and have been described, for example, in Makarova K., et al., “Prokaryotic homologs of Argonaute proteins are predicted to function as key components of a novel system of defense against mobile geneticelements”, Biol Direct.2009 Aug 25;4:29. doi: 10.1186 / 1745-6150-4-29, the entire contents of which is hereby incorporated by reference. In some embodiments, the napDNAbp is a Marinitoga piezophila Argunaute (MpAgo) protein. The CRISPR-associated Marinitoga piezophila Argunaute (MpAgo) protein cleaves single-stranded target sequences using 5ʹ- phosphorylated guides. The 5ʹ guides are used by all known Argonautes. The crystal structure of an MpAgo-RNA complex shows a guide strand binding site comprising residues that block 5ʹ phosphate interactions. This data suggests the evolution of an Argonaute subclass with noncanonical specificity for a 5ʹ-hydroxylated guide. See, e.g., Kaya et al., “A bacterial Argonaute with noncanonical guide RNA specificity”, Proc Natl Acad Sci U S A.2016 Apr 12;113(15):4057-62, the entire contents of which are hereby incorporated by reference). It should be appreciated that other argonaute proteins may be used, and are within the scope of this disclosure.
[0709] Some aspects of the disclosure provide Cas9 domains that have different PAM specificities. Typically, Cas9 proteins, such as Cas9 from S. pyogenes (spCas9), require a canonical NGG PAM sequence to bind a particular nucleic acid region. This may limit the ability to edit desired bases within a genome. In some embodiments, the base editing fusion proteins provided herein may need to be placed at a precise location, for example where a target base is placed within a 4 base region (e.g., a “editing window”), which is approximately 15 bases upstream of the PAM. See Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016), the entire contents of which are hereby incorporated by reference. Accordingly, in some embodiments, any of the fusion proteins provided herein may contain a Cas9 domain that is capable of binding a nucleotide sequence that does not contain a canonical (e.g., NGG) PAM sequence. Cas9 domains that bind to non-canonical PAM sequences have been described in the art and would be apparent to the skilled artisan. For example, Cas9 domains that bind non-canonical PAM sequences have been described in Kleinstiver, B. P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleinstiver, B. P., et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2015); the entire contents of each are hereby incorporated by reference.
[0710] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a nucleic acid programmable DNA binding protein that does not require a canonical (NGG) PAM sequence. In some embodiments, the napDNAbp is an argonauteprotein. One example of such a nucleic acid programmable DNA binding protein is an Argonaute protein from Natronobacterium gregoryi (NgAgo). NgAgo is a ssDNA-guided endonuclease. NgAgo binds 5′ phosphorylated ssDNA of ~24 nucleotides (gDNA) to guide it to its target site and will make DNA double-strand breaks at the gDNA site. In contrast to Cas9, the NgAgo–gDNA system does not require a protospacer-adjacent motif (PAM). Using a nuclease inactive NgAgo (dNgAgo) can greatly expand the bases that may be targeted. The characterization and use of NgAgo have been described in Gao et al., Nat Biotechnol., 34(7): 768-73 (2016), PubMed PMID: 27136078; Swarts et al., Nature, 507(7491): 258-61 (2014); and Swarts et al., Nucleic Acids Res. 43(10) (2015): 5120-9, each of which is incorporated herein by reference. The sequence of Natronobacterium gregoryi Argonaute is provided in SEQ ID NO:84.
[0711] In addition, any available methods may be utilized to obtain or construct a variant or mutant Cas9 protein. The term “mutation,” as used herein, refers to a substitution of a residue within a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or a deletion or insertion of one or more residues within a sequence. Mutations are typically described herein by identifying the original residue followed by the position of the residue within the sequence and by the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). Mutations can include a variety of categories, such as single base polymorphisms, microduplication regions, indel, and inversions, and is not meant to be limiting in any way. Mutations can include “loss-of-function” mutations which is the normal result of a mutation that reduces or abolishes a protein activity. Most loss-of-function mutations are recessive, because in a heterozygote the second chromosome copy carries an unmutated version of the gene coding for a fully functional protein whose presence compensates for the effect of the mutation. Mutations also embrace “gain-of-function” mutations, which is one which confers an abnormal activity on a protein or cell that is otherwise not present in a normal condition. Many gain-of-function mutations are in regulatory sequences rather than in coding regions, and can therefore have a number of consequences. For example, a mutation might lead to one or more genes being expressed in the wrong tissues, these tissues gaining functions that they normally lack. Because of their nature, gain-of-function mutations are usually dominant.
[0712] Mutations can be introduced into a reference Cas9 protein using site-directed mutagenesis. Older methods of site-directed mutagenesis known in the art rely on sub-cloning of the sequence to be mutated into a vector, such as an M13 bacteriophage vector, that allows the isolation of single-stranded DNA template. In these methods, one anneals a mutagenic primer (i.e., a primer capable of annealing to the site to be mutated but bearing one or more mismatched nucleotides at the site to be mutated) to the single-stranded template and then polymerizes the complement of the template starting from the 3ʹ end of the mutagenic primer. The resulting duplexes are then transformed into host bacteria and plaques are screened for the desired mutation. More recently, site-directed mutagenesis has employed PCR methodologies, which have the advantage of not requiring a single-stranded template. In addition, methods have been developed that do not require sub-cloning. Several issues must be considered when PCR-based site-directed mutagenesis is performed. First, in these methods it is desirable to reduce the number of PCR cycles to prevent expansion of undesired mutations introduced by the polymerase. Second, a selection must be employed in order to reduce the number of non-mutated parental molecules persisting in the reaction. Third, an extended-length PCR method is preferred in order to allow the use of a single PCR primer set. And fourth, because of the non-template-dependent terminal extension activity of some thermostable polymerases it is often necessary to incorporate an end-polishing step into the procedure prior to blunt-end ligation of the PCR-generated mutant product.
[0713] Mutations may also be introduced by directed evolution processes, such as phage- assisted continuous evolution (PACE) or phage-assisted noncontinuous evolution (PANCE). The term “phage-assisted continuous evolution (PACE),” as used herein, refers to continuous evolution that employs phage as viral vectors. The general concept of PACE technology has been described, for example, in International PCT Application, PCT / US2009 / 056194, filed September 8, 2009, published as WO 2010 / 028347 on March 11, 2010; International PCT Application, PCT / US2011 / 066747, filed December 22, 2011, published as WO 2012 / 088381 on June 28, 2012; U.S. Application, U.S. Patent No.9,023,594, issued May 5, 2015, International PCT Application, PCT / US2015 / 012022, filed January 20, 2015, published as WO 2015 / 134121 on September 11, 2015, and International PCT Application, PCT / US2016 / 027795, filed April 15, 2016, published as WO 2016 / 168631 on October 20, 2016, the entire contents of each of which are incorporated herein by reference. Variant Cas9s may also be obtain by phage-assisted non-continuous evolution (PANCE),” which as used herein, refers to non-continuous evolution that employs phage as viral vectors. PANCE is a simplified technique for rapid in vivo directed evolution using serial flask transfers of evolving ‘selection phage’ (SP), which contain a gene of interest to be evolved, across fresh E. coli host cells, thereby allowing genes inside the host E. coli to be held constant whilegenes contained in the SP continuously evolve. Serial flask transfers have long served as a widely-accessible approach for laboratory evolution of microbes, and, more recently, analogous approaches have been developed for bacteriophage evolution. The PANCE system features lower stringency than the PACE system.
[0714] Any of the references noted above which relate to Cas9 or Cas9 equivalents are hereby incorporated by reference in their entireties, if not already stated so. Divided napDNAbp domains for split PE delivery
[0715] In various embodiments, the prime editors described herein may be delivered to cells as two or more fragments which become assembled inside the cell (either by passive assembly, or by active assembly, such as using split intein sequences) into a reconstituted prime editor. In some cases, the self assembly may be passive whereby the two or more prime editor fragments associate inside the cell covalently or non-covalently to reconstitute the prime editor. In other cases, the self-assembly may be catalzyed by dimerization domains installed on each of the fragments. Examples of dimerization domains are described herein. In still other cases, the self-assembly may be catalyzed by split intein sequences installed on each of the prime editor fragments.
[0716] Split PE delivery may be advantageous to address various size constraints of different delivery approaches. For example, delivery approaches may include virus-based delivery methods, messenger RNA-based delivery methods, or RNP-based delivery (ribonucleoprotein-based delivery). And, each of these methods of delivery may be more efficient and / or effective by dividing up the prime editor into smaller pieces. Once inside the cell, the smaller pieces can assemble into a functional prime editor. Depending on the means of splitting, the divided prime editor fragments can be reassembled in a non-covalent manner or a covalent manner to reform the prime editor. In one embodiment, the prime editor can be split at one or more split sites into two or more fragments. The fragments can be unmodified (other than being split). Once the fragments are delivered to the cell (e.g., by direct delivery of a ribonucleoprotein complex or by nucleic delivery – e.g., mRNA delivery or virus vector based delivery), the fragments can reassociate covalently or non-covalently to reconstitute the prime editor. In another embodiment, the prime editor can be split at one or more split sites into two or more fragments. Each of the fragments can be modified to comprise a dimerization domain, whereby each fragment that is formed is coupled to a dimerization domain. Once delivered or expressed within a cell, the dimerization domains of the different fragments associate and bind to one another, bringing the different prime editor fragmentstogether to reform a functional prime editor. In yet another embodiment, the prime editor fragment may be modified to comprise a split intein. Once delivered or expressed within a cell, the split intein domains of the different fragments associate and bind to one another, and then undergo trans-splicing, which results in the excision of the split-intein domains from each of the fragments, and a concomitant formation of a peptide bond between the fragments, thereby restoring the prime editor.
[0717] In one embodiment, the prime editor can be delivered using a split-intein approach.
[0718] The location of the split site can be positioned between any one or more pair of residues in the prime editor and in any domains therein, including within the napDNAbp domain, the polymerase domain (e.g., RT domain), linker domain that joins the napDNAbp domain and the polymerase domain.
[0719] In one embodiment, the prime editor (PE) is divided at a split site within the napDNAbp.
[0720] In certain embodiments, the napDNAbp is a canonical SpCas9 polypeptide of SEQ ID NO: 14.
[0721] In certain embodiments, the SpCas9 is split into two fragments at a split site located between residues 1 and 2, or 2 and 3, or 3 and 4, or 4 and 5, or 5 and 6, or 6 and 7, or 7 and 8, or 8 and 9, or 9 and 10, or between any two pair of residues located anywhere between residues 1-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100- 200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 1000-1100, 1100- 1200, 1200-1300, or 1300-1368 of canonical SpCas9 of SEQ ID NO: 14.
[0722] In certain embodiments, a napDNAbp is split into two fragments at a split site that is located at a pair of residue that corresponds to any two pair of residues located anywhere between positions 1-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100- 200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 1000-1100, 1100- 1200, 1200-1300, or 1300-1368 of canonical SpCas9 of SEQ ID NO: 14.
[0723] In certain embodiments, the SpCas9 is split into two fragments at a split site located between residues 1 and 2, or 2 and 3, or 3 and 4, or 4 and 5, or 5 and 6, or 6 and 7, or 7 and 8, or 8 and 9, or 9 and 10, or between any two pair of residues located anywhere between residues 1-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100- 200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 1000-1100, 1100- 1200, 1200-1300, or 1300-1368 of canonical SpCas9 of SEQ ID NO: 14. In certain embodiments, the split site is located one or more polypeptide bond sites (i.e., a “split site orsplit-intein split site”), fused to a split intein, and then delivered to cells as separately- encoded fusion proteins. Once the split-intein fusion proteins (i.e., protein halves) are expressed within a cell, the proteins undergo trans-splicing to form a complete or whole PE with the concomitant removal of the joined split-intein sequences.
[0724] In some embodiments, the N-terminal extein can be fused to a first split-intein (e.g., N intein) and the C-terminal extein can be fused to a second split-intein (e.g., C intein). The N-terminal extein becomes fused to the C-terminal extein to reform a whole prime editor fusion protein comprising an napDNAbp domain and a polymerase domain (e.g., RT domain) upon the self-association of the N intein and the C intein inside the cell, followed by their self-excision, and the concomitant formation of a peptide bond between the N-terminal extein and C-terminal extein portions of a whole prime editor (PE).
[0725] To take advantage of a split-PE delivery strategy using split-inteins, the prime editor needs to be divided at one or more split sites to create at least two separate halves of a prime editor, each of which may be rejoined inside a cell if each half is fused to a split-intein sequence.
[0726] In certain embodiments, the prime editor is split at a single split site. In certain other embodiments, the prime editor is split at two split sites, or three split sites, or four split sites, or more.
[0727] In a preferred embodiment, the prime editor is split at a single split site to create two separate halves of a prime editor, each of which can be fused to a split intein sequence
[0728] An exemplary split intein is the Ssp DnaE intein, which comprises two subunits, namely, DnaE-N and DnaE-C. The two different subunits are encoded by separate genes, namely dnaE-n and dnaE-c, which encode the DnaE-N and DnaE-C subunits, respectively. DnaE is a naturally occurring split intein in Synechocytis sp. PCC6803 and is capable of directing trans-splicing of two separate proteins, each comprising a fusion with either DnaE- N or DnaE-C.
[0729] Additional naturally occurring or engineered split-intein sequences are known in the or can be made from whole-intein sequences described herein or those available in the art. Examples of split-intein sequences can be found in Stevens et al., “A promiscuous split intein with expanded protein engineering applications,” PNAS, 2017, Vol.114: 8538-8543; Iwai et al., “Highly efficient protein trans-splicing by a naturally split DnaE intein from Nostc punctiforme, FEBS Lett, 580: 1853-1858, each of which are incorporated herein by reference. Additional split intein sequences can be found, for example, in WO 2013 / 045632, WO2014 / 055782, WO 2016 / 069774, and EP2877490, the contents each of which are incorporated herein by reference.
[0730] In addition, protein splicing in trans has been described in vivo and in vitro (Shingledecker, et al.,
[0076] Gene 207:187 (1998), Southworth, et al., EMBO J.17:918 (1998); Mills, et al., Proc. Natl. Acad. Sci. USA, 95:3543-3548 (1998); Lew, et al., J. Biol. Chem., 273:15887-15890 (1998); Wu, et al., Biochim. Biophys. Acta 35732:1 (1998b), Yamazaki, et al., J. Am. Chem. Soc.120:5591 (1998), Evans, et al., J. Biol. Chem.275:9091 (2000); Otomo, et al., Biochemistry 38:16040-16044 (1999); Otomo, et al., J. Biolmol. NMR 14:105-114 (1999); Scott, et al., Proc. Natl. Acad. Sci. USA 96:13638-13643 (1999)) and provides the opportunity to express a protein as to two inactive fragments that subsequently undergo ligation to form a functional product.
[0731] In various embodiments described herein, the continuous evolution methods (e.g., PACE) may be used to evolve a first portion of a base editor. A first portion could include a single component or domain, e.g., a Cas9 domain, a deaminase domain, or a UGI domain. The separately evolved component or domain can be then fused to the remaining portions of the base editor within a cell by separately express both the evolved portion and the remaining non-evolved portions with split-intein polypeptide domains. The first portion could more broadly include any first amino acid portion of a base editor that is desired to be evolved using a continuous evolution method described herein. The second portion would in this embodiment refer to the remaining amino acid portion of the base editor that is not evolved using the herein methods. The evolved first portion and the second portion of the base editor could each be expressed with split-intein polypeptide domains in a cell. The natural protein splicing mechanisms of the cell would reassemble the evolved first portion and the non- evolved second portion to form a single fusion protein evolved base editor. The evolved first portion may comprise either the N- or C-terminal part of the single fusion protein. In an analogous manner, use of a second orthogonal trans-splicing intein pair could allow the evolved first portion to comprise an internal part of the single fusion protein.
[0732] Thus, any of the evolved and non-evolved components of the base editors herein described may be expressed with split-intein tags in order to facilitate the formation of a complete base editor comprising the evolved and non-evolved component within a cell.
[0733] The mechanism of the protein splicing process has been studied in great detail (Chong, et al., J. Biol. Chem.1996, 271, 22159-22168; Xu, M-Q & Perler, F. B. EMBO Journal, 1996, 15, 5146-5153) and conserved amino acids have been found at the intein and extein splicing points (Xu, et al., EMBO Journal, 1994, 135517-522). The constructsdescribed herein contain an intein sequence fused to the 5′-terminus of the first gene (e.g., the evolved portion of the base editor). Suitable intein sequences can be selected from any of the proteins known to contain protein spli...
Claims
CLAIMS What is claimed is:
1. A composition comprising: a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14, wherein the first nucleotide sequence further comprises a nucleotide sequence encoding a sgRNA sequence operably linked to a promoter.
2. A composition comprising: a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, and wherein three terminal amino acids at the N-terminus of the C-terminal portion are mutated relative to SEQ ID NO: 14 wherein the second nucleotide sequence further comprise a nucleotide sequence encoding a pegRNA operably linked to a promoter, wherein the pegRNA comprises at its 3’ end a prequenosine1-1 riboswitch aptamer (evopreQ1) or a variant there of motif, wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase codon optimized for expression in a mammalian cell.
3. A composition comprising: (i) a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14, wherein the first nucleotide sequence further comprises a nucleotide encoding a sgRNA sequence operably linked to a promoter; and (ii) a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C- terminal portion of the prime editor; wherein the C-terminal portion of the prime editorcomprises a portion of any one of SEQ ID NOs: 1417, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, wherein three terminal amino acids at the N-terminus of the C-terminal portion are SEQ, SFQ, SFN, SEN, or CFN, wherein the second nucleotide sequence further comprises a nucleotide encoding a pegRNA operably linked to a promoter, wherein the pegRNA comprises an evopreQ1motif at its 3′ end, and wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase codon optimized for expression in a mammalian cell.
4. A composition comprising: (i) a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 1417, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14; and (ii) a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 1417, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, wherein three terminal amino acids at the N-terminus of the C-terminal portion are SEQ, SFQ, SFN, SEN, or CFN, wherein the first nucleotide sequence further comprises a nucleotide encoding a pegRNA or a second nucleotide sequence further comprise a nucleotide encoding the pegRNA, operably linked to a promoter, wherein the pegRNA comprises an evopreQ1 motif at its 3′ end, and wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase codon optimized for expression in a mammalian cell.
5. A composition comprising:(i) a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14; and (ii) a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, wherein three terminal amino acids at the N-terminus of the C-terminal portion are SEQ,SFQ, SFN, SEN, or CFN, wherein the first nucleotide sequence further comprises a nucleotide encoding a pegRNA and the second nucleotide sequence further comprises a nucleotide encoding a sgRNA, operably linked to a promoter, wherein the pegRNA comprises an evopreQ1motif at its 3′ end, and wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase codon optimized for expression in a mammalian cell.
6. A composition comprising a first recombinant adeno associated virus (rAAV) vector comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N, and wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14, wherein the first nucleotide sequence further comprises a nucleotide encoding a sgRNA sequence operably linked to a promoter.
7. A composition comprising a second recombinant adeno associated virus (rAAV) vector comprising a second nucleotide sequence encoding an intein-C fused to the N- terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, and wherein three terminal amino acids at the N-terminus of the C-terminal portion are mutated from SEQ to SFQ, SFN, SEN, or CFN,wherein the second nucleotide sequence further comprise a nucleotide encoding a pegRNA operably linked to a promoter, wherein the pegRNA comprises an evopreQ1 (tevopreQ1) motif at its 3′ end, wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase codon optimized for expression in a mammalian cell.
8. A composition comprising: (i) a first recombinant adeno associated virus (rAAV) vector comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N, wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84 that corresponds to amino acids 1-844 or 1- 1024 of SEQ ID NO: 14, wherein the first nucleotide sequence further comprises a nucleotide sequence encoding a sgRNA sequence operably linked to a promoter; and (ii) a second recombinant adeno associated virus (rAAV) vector comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NOs: 17, 19, and 21-85 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, wherein three terminal amino acids at the N-terminus of the C-terminal portion are SFQ, SFN, SEN, or CFN, wherein the second nucleotide sequence further comprise a nucleotide encoding a pegRNA operably linked to a promoter, wherein the pegRNA comprises an evopreQ1motif at its 3′ end, wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase codon optimized for expression in a mammalian cell.
9. The composition of claim 2, 3-5, or 7-8, wherein the evopreQ1 comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to CGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAA (SEQ ID NO:447).
10. The composition of claims 6 and 8, wherein the first nucleotide sequence comprises a nucleic acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identical to any one of SEQ ID NOs:15, 16, 18, 20, 118, 120, or 123.
11. The composition of claims 10, wherein the first nucleotide sequence further comprises one or more regulatory elements.
12. The composition of claims 11, wherein the one or more regulatory elements comprises one or more promoters.
13. The composition of claims 12, wherein the one or more promoters is selected from the group consisting of EFS, Cbh, sCAG, hCMV, mPGK, hSYN, and U6.
14. The composition of claim 13, wherein the one or more promoter comprises a Cbh promoter.
15. The composition of any one of claims 12-14, wherein the one or more promoter drives expression of the N-terminal portion of the prime editor.
16. The composition of any one of claims 12-15, wherein the one or more promoters drives expression of the sgRNA.
17. The composition of claim 12, wherein the first nucleotide sequence comprises a first promoter operably linked to the N-terminal portion of the prime editor and a second promoter operably linked to the nucleotide sequence encoding the sgRNA.
18. The composition of claim 17, wherein the first promoter is a Cbh promoter and wherein the second promoter is a U6 promoter.
19. The composition of claim 14 or 18, wherein the Cbh promoter comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to20. The composition of claims 11-19, wherein the one or more regulatory elements comprises a transcriptional terminator.
21. The composition of claim 20, wherein the transcriptional terminator is selected form the group consisting of bovine growth hormone terminator (bGH-polyA), and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA and arcA terminator.
22. The composition of claims 11-21, wherein the one or more regulatory elements comprises at least one SV40 late polyadenylation signal (SV40 late polyA).
23. The composition of claim 22, wherein the SV40 late polyA comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to24. The composition of claims 11-23, wherein the one or more regulatory elements comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE / W3).
25. The composition of claim 24, wherein the WPRE / W3 comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to26. The composition of claims 10-25, wherein the first nucleotide sequence further comprises one or more nucleotides sequences encoding one or more ITR domains.
27. The composition of claims 7-26, wherein the second nucleotide sequence comprises a nucleic acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 15, 16, 18, 20, 119, 121, and 124.
28. The composition of claim 7-27, wherein the second nucleotide sequence further comprises one or more regulatory elements.
29. The composition of claim 28, wherein the one or more regulatory elements comprises one or more promoters.
30. The composition of claims 29, wherein the one or more promoters is selected from the group consisting of EFS, Cbh, sCAG, hCMV, mPGK, hSYN, and U6 31. The composition of claim 29 or 30, wherein the one or more promoters is EFS.
32. The compositions of claim 29 or 30, wherein the one or more promoters is U6 33. The compositions of claim 29 or 30, wherein the one or more promoters comprises a Chb promoter.
34. The composition of claims 29-33, wherein the one or more promoters drives expression of the C-terminal portion of the prime editor.
35. The composition of claims 29-34, wherein the one or more promoters drives expression of the pegRNA.
36. The composition of claim 12, wherein the first nucleotide sequence comprises a first promoter operably linked to the C-terminal portion of the prime editor and a second promoter operably linked to the nucleotide sequence encoding the sgRNA.
37. The composition of claim 36, wherein the first promoter is a Cbh promoter and wherein the second promoter is a U6 promoter.
38. The composition of claim 36 or 37, wherein the Cbh promoter comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to39. The composition of claims 28-38, wherein the one or more regulatory elements comprises a transcriptional terminator.
40. The composition of claim 39, wherein the transcriptional terminator is selected form the group consisting of bovine growth hormone terminator (bGH-polyA), and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA and arcA terminator.
41. The composition of claims 28-40, wherein the one or more regulatory elements comprises at least one SV40 late polyadenylation signal (SV40 late polyA).
42. The composition of claim 41, wherein the SV40 late polyA comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to43. The composition of claims 28-42 wherein the one or more regulatory elements comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE / W3).
44. The composition of claim 24, wherein the WPRE / W3 comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to45. The composition of claims 7-44, wherein the pegRNA encodes for one or more edits selected from the group consisting of +1 C-to-G, +1C-to-G and +5 G-to-T, +2 G-to-C, +1 CTT insertion, and +1 CGA insertion.
46. The composition of claim 45, wherein the one or more edits are not repaired by the DNA mismatch repair pathway (MMR).
47. The composition of claims 7-46, wherein the second nucleotide sequence further comprises one or more nucleotide sequences encoding one or more ITR domains.
48. The composition of any one of claims 6 or 8-47, wherein the first recombinant adeno associated virus (rAAV) particle is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and AAV9.
49. The composition of claims 7-48, wherein the second recombinant adeno associated virus (rAAV) particle is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and AAV9.
50. A composition comprising: (i) a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; (ii) a second nucleotide sequence encoding an intein-C fused to the N- terminus of a C-terminal portion of the prime editor; (iii) a third nucleotide sequence encoding a pegRNA; 51. A composition comprising:(i) a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 1-844 or 1-1024 of SEQ ID NO: 14; (ii) a second nucleotide sequence encoding an intein-C fused to the N- terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, and wherein three terminal amino acids at the N-terminus of the C-terminal portion are mutated from SEQ to SFQ, SFN, SEN, or CFN; and (iii) a third nucleotide sequence encoding a pegRNA; 52. The composition of claim 51, wherein the third nucleotide sequence further encodes a sgRNA.
53. The composition of claim 50- 51, wherein the pegRNA comprises an evopreQ1motif at its 3′ end.
54. The composition of claim 53, wherein the evopreQ1 comprises nucleotide sequence55. The composition of claims 50-54, wherein the pegRNA encodes for one or more edits selected from the group consisting of +1 C-to-G, +1C-to-G and +5 G-to-T, +2 G-to-C, +1 CTT insertion, and +1 CGA insertion.
56. The composition of 50-55, wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase codon optimized for expression in a mammalian cell.
57. A composition comprising: (i) a first recombinant adeno associated virus (rAAV) vector comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; and (ii) a second recombinant adeno associated virus (rAAV) vector comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; and(iii) a third recombinant adeno associated virus (rAAV) vector comprising a third nucleotide sequence encoding a pegRNA.
58. A composition comprising: (i) a first recombinant adeno associated virus (rAAV) vector comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 1-844 or 1- 1024 of SEQ ID NO: 14; and (ii) a second recombinant adeno associated virus (rAAV) vector comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, wherein three terminal amino acids at the N-terminus of the C-terminal portion are SFQ, SFN, SEN, or CFN; and (iii) a third recombinant adeno associated virus (rAAV) vector comprising a third nucleotide sequence encoding a pegRNA.
59. The composition of claim 57 or 58, wherein the third nucleotide sequence further encodes a sgRNA.
60. The composition of claims 57 or 58, wherein the first nucleotide sequence comprises a nucleic acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 15, 16, 18, 20, 118, 120, or 123.
61. The composition of claims 57-60, wherein the first nucleotide sequence further comprises one or more regulatory elements. 62 The composition of claim 61, wherein the one or more regulator elements comprises one or more promoters.
63. The composition of claim 62, wherein the one or more promoters is selected from the group consisting of EFS, Cbh, sCAG, hCMV, mPGK, hSYN, and U6 64. The composition of claims 62 or 63, wherein the one or more promoter drives expression of the N-terminal portion of the prime editor.
65. The composition of claims 61-64, wherein the one or more regulatory elements comprises a transcriptional terminator.
66. The composition of claim 65, wherein the transcriptional terminator is selected form the group consisting of bovine growth hormone terminator (bGH-polyA), and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA and arcA terminator.
67. The composition of claims 61-66, wherein the one or more regulatory elements comprises at least one SV40 late polyadenylation signal (SV40 late polyA).
68. The composition of claim 67, wherein the SV40 late polyA comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to69. The composition of claims 61-68, wherein the one or more regulatory elements comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE / W3).
70. The composition of claims 57-69, wherein the first nucleotide sequence further comprises one or more nucleotide sequences encoding one or more ITR domains.
71. The composition of claims 57-70, wherein the second nucleotide sequence comprises a nucleic acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 15, 16, 18, 20, 119, 121, and 124.
72. The composition of claim 57-71, wherein the second nucleotide sequence further comprises one or more regulatory elements.
73. The composition of claim 72, wherein the one or more regulator elements comprises one or more promoters.
74. The composition of claims 73, wherein the one or more promoters is selected from the group consisting of EFS, Cbh, sCAG, hCMV, mPGK, hSYN, and U6 75. The composition of claim 74, wherein the one or more promoters is EFS.
76. The compositions of claim 74, wherein the one or more promoters is U6 77. The composition of claim 74, wherein the one or more promoters comprises a Chb promoter.
78. The composition of claims 74-77, wherein the one or more promoters drives expression of the C-terminal portion of the prime editor.
79. The composition of claim 12, wherein the first nucleotide sequence comprises a first promoter operably linked to the C-terminal portion of the prime editor and a second promoter operably linked to the nucleotide sequence encoding the sgRNA.
80. The composition of claim 79, wherein the first promoter is a Cbh promoter and wherein the second promoter is a U6 promoter.
81. The composition of claim 79 or 80, wherein the Cbh promoter comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to82. The composition of claims 72-81, wherein the one or more regulatory elements comprises a transcriptional terminator.
83. The composition of claim 82, wherein the transcriptional terminator is selected form the group consisting of bovine growth hormone terminator (bGH-polyA), and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA and arcA terminator.
84. The composition of claims 72-83, wherein the one or more regulatory elements comprises at least one SV40 late polyadenylation signal (SV40 late polyA).
85. The composition of claim 84, wherein the SV40 late polyA comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to CTTTATTTGTGAAATTTGTGATGCTATTGCTTTATTTGTAACCATTATAAGCTGCA ATAAACAAGTTAACAACAACAATTGCATTCATTTTATGTTTCAGGTTCAGGGGGA GATGTGGGAGGTTTTTTAAAGC (SEQ ID NO: 448).
86. The composition of claims 72-85 wherein the one or more regulatory elements comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE / W3).
87. The composition of claim 86, wherein the WPRE / W3 comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to88. The composition of claims 57-87, wherein the second nucleotide sequence further comprises one or more nucleotide sequences encoding one or more ITR domains.
89. The composition of claims 57-88, wherein the third nucleotide sequence comprises a nucleic acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 122 or 198-226.
90. The composition of claim 57-89, wherein the third nucleotide sequence further comprises one or more regulatory elements.
91. The composition of claim 90, wherein the one or more regulator elements comprises one or more promoters.
92. The composition of claims 91, wherein the one or more promoters is selected from the group consisting of EFS, Cbh, sCAG, hCMV, mPGK, hSYN, and U6 93. The composition of claim 91, wherein the one or more promoters is Chb.
94. The compositions of claim 91, wherein the one or more promoters is U6 95. The composition of claims 91-94, wherein the one or more promoters drives expression of the sgRNA.
96. The composition of claims 91-95, wherein the one or more promoters drives expression of pegRNA.
97. The composition of claim 96, wherein the pegRNA encodes for one or more edits selected from the group consisting of +1 C-to-G, +1C-to-G and +5 G-to-T, +2 G-to-C, +1 CTT insertion, and +1 CGA insertion.
98. The composition of claims 97, wherein the one or more edits are not repaired by the DNA mismatch repair pathway (MMR).
99. The composition of claims 90-98, wherein the one or more regulatory elements comprises at least one nuclear localization signal.
100. The composition of claims 57-99, wherein the third nucleotide sequence further comprises a nucleotide sequence encoding an ITR domain.
101. The composition of claims 57-100, wherein the first recombinant adeno associated virus (rAAV) particle is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and AAV9.
102. The composition of claims 57-101, wherein the second recombinant adeno associated virus (rAAV) particle is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and AAV9.
103. The composition of claims 57-102, wherein the third recombinant adeno associated virus (rAAV) particle is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and AAV9.
104. A composition comprising: (i) a first polynucleotide comprising, from 5’ to 3’: a Cbh promoter comprising SEQ ID NO: 449, a first sequence encoding an N-terminal fragment of a prime editor fusion protein, a second sequence encoding an intein-N, a SV40 late polyadenylation signal (SV40 late polyA) comprising SEQ ID NO: 434; (ii) a second polynucleotide comprising, from 5’ to 3’: a Cbh promoter comprising SEQ ID NO:449, a third sequence encoding an intein-C, a fourth sequence encoding a C- terminal fragment of the prime editor fusion protein, a SV40 late polyA comprising SEQ ID NO: 434, wherein the intein-C and the intein N are capable of self-excision to join the N- terminal fragment and the C-terminal fragment to form the prime editor fusion protein, and wherein the prime editor fusion protein comprises a nucleic acid programmable DNA binding protein (napDNAbp)domain and a reverse transcriptase (RT) domain.
105. The composition of claim 104, wherein the second polynucleotide further comprises a fifth sequence encoding a prime editing guide RNA (pegRNA) operably linked to a promoter, wherein the fifth sequence is 3’ to the SV40 late poly A.
106. The composition of claim 104 or 105, wherein the second polynucleotide further comprises a sixth sequence encoding a single guide RNA (sgRNA) operably linked to a promoter, wherein the sixth sequence is 3’ to the SV40 late polyA.
107. A composition comprising: (i) a first polynucleotide comprising, from 5’ to 3’: a Cbh promoter comprising SEQ ID NO: 449, a first sequence encoding an N-terminal fragment of a prime editor fusion protein, a second sequence encoding an intein-N, a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE / W3) comprising SEQ ID NO: 434; (ii) a second polynucleotide comprising, from 5’ to 3’: a Cbh promoter comprising SEQ ID NO: 449, a third sequence encoding an intein-C, a fourth sequence encoding a C- terminal fragment of the prime editor fusion protein, a WPRE / W3 comprising SEQ ID NO: 434, wherein the intein-C and the intein N are capable of self-excision to join the N- terminal fragment and the C-terminal fragment to form the prime editor fusion protein, andwherein the prime editor fusion protein comprises a nucleic acid programmable DNA binding protein (napDNAbp)domain and a reverse transcriptase (RT) domain.
108. The composition of claim 107, further comprising a third polynucleotide, wherein the third polynucleotide comprises a pegRNA operably linked to a promoter and / or a sgRNA operably linked to a promoter.
109. The composition of claim 108, wherein the pegRNA comprises at its 3’ end a prequenosine1-1 riboswitch aptamer (evopreQ1) motif or a variant there of.
110. The composition of claim 109, wherein the evopreQ1 motif comprises a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to111. The composition of any one of claims 108-110, wherein the pegRNA and / or the sgRNA is operably linked to a U6 promoter comprising a nucleotide sequence at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identical to112. The composition of any one of claims 104-111, wherein the napDNAbp is a Cas9 or a variant thereof.
113. The composition of claim 112, wherein the Cas9 comprises one or more of amino acid substitutions selected from the group consisting of S1025C, E1026F, and Q1027N as compared to SEQ ID NO:
14.
114. The composition of claim 112 or 113, wherein the Cas9 comprises amino acid substitutions R221K, N394K, and H840A as compared to SEQ ID NO:
14.
115. The composition of claim 112, wherein the Cas9 comprises at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 14, 17, 19, 24- 53, 55-65, 67-82, 84.
116. The composition of any one of claims 104-115, wherein the RT comprises an amino acid sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity sequence identity with SEQ ID NO:
116.
117. The composition of claim 116, wherein the RT lacks a RnasH domain compared to the RT as set forth in SEQ ID NO:
116.
118. The composition of claim 117, wherein the RT has the amino acid sequence of SEQ ID NO: 98 or119. The composition of any one of claims 104-118, wherein the N-terminal fragment of the prime editor fusion protein comprises a portion of the napDNAbp domain and the C- terminal fragment of the prime editor fusion protein comprises a portion of the napDNA domain and the RT domain.
120. The composition of claim 119, wherein the N-terminal fragment comprises an amino acid sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity to amino acids 1-1024 of a sequence selected from the group consisting of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84, and wherein the C-terminal fragment comprises an amino acid sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity to amino acids 1025-1368 of the selected sequence for the N-terminal fragment.
121. The composition of claim 119, wherein the N-terminal fragment comprises an amino acid sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity to amino acids 1-844 of a sequence selected from the group consisting of SEQ ID NOs: 14, 17, 19, 24-53, 55-65, 67-82, 84, and wherein the C-terminal fragment comprises an amino acid sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity to amino acids 845-1368 of the selected sequence for the N-terminal fragment.
122. The composition of any one of claims 104-121, wherein the intein-N comprises an amino acid sequence having at least 80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity to SEQ ID NO:
184.
123. The composition of claim 122, wherein the intein-C comprises an amino acid sequence having at least 80%,80%, 85%, 90%, 92.5%, 95%, 97%, 98%, 99%, or 100% identity to SEQ ID NO:
186.
124. The composition of any one of claims 104-123, wherein the first polynucleotide is encoded by a first vector, and wherein the second polynucleotide is encoded by a second vector, and / or wherein the third polynucleotide is encoded by a third vector.
125. The composition of claim 124, wherein the first, the second, and / or the third vector is each an AAV vector.
126. The composition of claim 125, wherein the AAV vector is an AAV9 vector.
127. One or more AAV particles comprising the first, the second, and / or the third vector of claim 125 or 126.
128. A composition comprising: (i) a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; and (ii) a second nucleotide sequence encoding an intein-C fused to the N- terminus of a C-terminal portion of the prime editor, and wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase lacking a RNaseH domain.
129. The composition of claim 128, wherein the second nucleotide further comprises a nucleotide sequence encoding for a sgRNA 130. The composition of claim 128 or 129, wherein the second nucleotide further comprises a nucleotide sequence encoding for a pegRNA.
131. The compositions of 128-130, wherein the pegRNA comprises an evopreQ1 (tevopreQ1) motif at its 3′ end.
132. The composition of claims 128-131, wherein the pegRNA encodes for one or more edits selected from the group consisting of +1 C-to-G, +1C-to-G and +5 G-to-T, +2 G-to-C, +1 CTT insertion, and +1 CGA insertion.
133. The composition of 128-132, wherein MMLV reverse transcriptase lacking the RNaseH domain is codon optimized for expression in a mammalian cell.
134. A composition comprising: (i) a first recombinant adeno associated virus (rAAV) particle comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C- terminus to an intein-N; and (ii) a second recombinant adeno associated virus (rAAV) particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; and wherein the second nucleotide sequence further comprises a nucleotide encoding a MMLV reverse transcriptase lacking a RNaseH domain.
135. The composition of claim 134, wherein the first nucleotide sequence comprises a nucleic acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identical to any one of SEQ ID NOs:15, 16, 18, 20, 118, 120, or 123.
136. The composition of claims 134 or 135, wherein the first nucleotide sequence further comprises one or more regulatory elements. 137 The composition of claims 136, wherein the one or more regulator elements comprises one or more promoters.
138. The composition of claims 137, wherein the one or more promoters is selected from the group consisting of EFS, Cbh, sCAG, hCMV, mPGK, hSYN, and U6.
139. The composition of claim 137, wherein the one or more promoters is Cbh.
140. The composition of claims 137-139, wherein the one or more promoter drives expression of the N-terminal portion of the prime editor.
141. The composition of claims 136-140, wherein the one or more regulatory elements comprises a transcriptional terminator.
142. The composition of claim 141, wherein the transcriptional terminator is selected form the group consisting of bovine growth hormone terminator (bGH-polyA), and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA and arcA terminator.
143. The composition of claim 142, wherein the transcriptional terminator is SV40 late- polyA.
144. The composition of claims 136-143, wherein the one or more regulatory elements comprises at least one nuclear localization signal.
145. The composition of claim 144, wherein the nuclear localization signal is SV40NLS.
146. The composition of claims 134-145, wherein the first nucleotide sequence further comprises a nucleotide sequence encoding an ITR domain.
147. The composition of claims 134-146, wherein the second nucleotide sequence comprises a nucleic acid sequence that is at least 85%, 90%, 92.5%, 95%, 97%, 98%, or 99% identical to any one of SEQ ID NOs:15, 16, 18, 20, 119, 121, and 124.
148. The composition of claims 134-147, wherein the second nucleotide further comprises a nucleotide sequence encoding for a sgRNA.
149. The composition of claims 134-148, wherein the second nucleotide further comprises a nucleotide sequence encoding for a pegRNA.
150. The composition of claim 134-147, wherein the second nucleotide sequence further comprises one or more regulatory elements.
151. The composition of claim 150, wherein the one or more regulator elements comprises one or more promoters.
152. The composition of claims 151, wherein the one or more promoters is selected from the group consisting of EFS, Cbh, sCAG, hCMV, mPGK, hSYN, and U6 153. The composition of claim 151, wherein the one or more promoters is Cbh.
154. The compositions of claim 151, wherein the one or more promoters is U6 155. The composition of claims 151-154, wherein the one or more promoters drives expression of the sgRNA.
156. The composition of claims 151-155, wherein the one or more promoters drives expression of the pegRNA.
157. The composition of claims 151-156, wherein the one or more promoters drives expression of the C-terminal portion of the prime editor.
158. The composition of claims 150-157, wherein the one or more regulatory elements comprises a transcriptional terminator.
159. The composition of claim 158, wherein the transcriptional terminator is selected form the group consisting of bovine growth hormone terminator (bGH-polyA), and viral termination sequences such as, for example, the SV40 terminator, spy, yejM, secG-leuU, thrLABC, rrnB T1, hisLGDCBHAFI, metZWV, rrnC, xapR, aspA and arcA terminator.
160. The composition of claim 159, wherein the transcriptional terminator is SV40 late polyA.
161. The composition of claims 150-160, wherein the one or more regulatory elements comprises at least one nuclear localization signal.
162. The composition of claim 161, wherein the nuclear localization signal is SV40NLS.
163. The composition of claims 134-162, wherein the second nucleotide sequence further comprises a nucleotide sequence encoding an ITR domain.
164. The composition of claims 134-163, wherein the nucleotide sequence encoding the truncated MMLV reverse transcriptase lacking the RNaseH domain comprises a nucleotide sequence that is at least 80%, 85%, 90%, 95%, 99% or 99.5% sequence identical to SEQ ID NO:
140.
165. The composition of claims 134-164, wherein the nucleotide sequence truncated MMLV reverse transcriptase lacking the RNaseH is codon optimized for expression in a mammalian cell.
166. The composition of claims 134-165, wherein the first recombinant adeno associated virus (rAAV) particle is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and AAV9.
167. The composition of claims 134-166, wherein the second recombinant adeno associated virus (rAAV) particle is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, and AAV9.
168. A composition comprising a prime editor, the composition comprising: (i) a gene product of the first nucleotide sequence of any one of claims 1, 3-6, 8-26, 128-146, wherein the gene product of the first nucleotide sequence comprises a protein comprising a N-terminal portion of a split prime editor and a sgRNA; and(ii) a gene product of the second nucleotide sequence of any one of claims 2-5, 7, 8, 27-47, 128-134, 147-165, wherein the gene product of the second nucleotide sequence is a protein comprising a C-terminal portion of a split prime editor and a pegRNA, wherein the N-terminal portion of the gene product is fused at its C-terminus to the N- terminus of the C-terminal portion of the gene product to produce the prime editor.
169. A composition comprising a prime editor, the composition comprising: (i) a gene product of the first nucleotide sequence of any one of claims 128- 146, wherein the gene product of the first nucleotide sequence comprises a protein comprising a N-terminal portion of a split prime editor; and (ii) a gene product of the second nucleotide sequence of any one of claims 128-134, 147-165, wherein the gene product of the second nucleotide sequence is a protein comprising a C-terminal portion of a split prime editor, a sgRNA, and a pegRNA, wherein the N-terminal portion of the gene product is fused at its C-terminus to the N- terminus of the C-terminal portion of the gene product to produce the prime editor.
170. A composition comprising a prime editor, the composition comprising: (i) a gene product of the first nucleotide sequence of any one of claims 50-70, wherein the gene product of the first nucleotide sequence is a protein comprising a N- terminal portion of a split prime editor; and (ii) a gene product of the second nucleotide sequence of any one of claims 50- 58 and 71-88, wherein the gene product of the second nucleotide sequence is a protein comprising a C-terminal portion of a split prime editor, and (iii) a gene product of the third nucleotide sequence of any one of claims 50- 58 and 89-100, wherein the gene product of the third nucleotide sequence is a sgRNA and pegRNA, wherein the N-terminal portion of the gene product is fused at its C-terminus to the N- terminus of the C-terminal portion of the gene product to produce the prime editor.
171. A vector comprising a composition of any one of claims 1-5, 50-56, 128-133.
172. An AAV comprising a composition of any one of claims 1-5, 50-56, 128-133.
173. A cell comprising a composition of any one of claims 1-170, a vector of claim 171, or an AAV of claim 172.
174. A pharmaceutical composition comprising a composition of any one of claims 1-170, a vector of claim 171, or an AAV of claim 172 and a pharmaceutically acceptable excipient.
175. A kit comprising the compositions of any one of claims 1-174.
176. A method for editing one or more target genes in an organ of interest of a subject in need thereof, the method comprising administering to the subject a composition of any one of claims 6-49, 57-103, or 134-167.
177. The method of claim 176, wherein the organ of interest is selected from the group consisting of brain, heart, lung, kidney, liver, stomach, spleen, skeletal muscle, eye 178. The method of claim 176, wherein the administration is performed systematically.
179. The method of claim 176 or 177, further comprising an rAAV particle 180. The method of claim 176-179, wherein the rAAV particle is configured to cross the blood-brain barrier.
181. The method of claim 180, wherein the rAAV particle comprises PHP.eB.
182. A method for editing a neuron in a subject in need thereof, comprising administering to the subject a therapeutically effective amount a composition of any one of claims 6-49.
183. A method for editing a neuron and astrocyte in a subject in need thereof, comprising administering to the subject a therapeutically effective amount a composition of any one of claims 134-167.
184. A method of contacting a cell with a composition of any one of claims 6-49, 57-103, or 134-167, wherein the contacting results in the delivery of the first nucleotide sequence and the second nucleotide sequence into the cell, and wherein the N-terminal portion of a split prime editor and the C-terminal portion of a split prime editor are joined to form a prime editor.
185. A method of contacting a cell with a composition of any one of claims 1-5, 50-56, and 128-133, wherein the composition is encapsulated with an LNP or VLP, and wherein contacting results in the delivery of the first nucleotide sequence and the second nucleotide sequence into the cell, and wherein the N-terminal portion of a split prime editor and the C- terminal portion of the a split prime editor are joined to form a prime editor.
186. A method comprising administering to a subject in need thereof a therapeutically effective amount of the composition any one of claims 1-5, 50-56, and 128-133 or 6-49, 57- 103, or 134-167 and a pharmaceutically acceptable excipient.
187. A composition comprising: a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 22, 24-84 that corresponds to amino acids 1-844 or 1-1024 of SEQ ID NO: 14, wherein the first nucleotide sequence further comprises a nucleotide encoding a sgRNA sequence operably linked to a promoter, wherein the intein-N is a catalytically competent intein.
188. A composition comprising: a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C- terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, and 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, wherein the intein-C is a catalytically competent intein.
189. A composition comprising: a first recombinant adeno associated virus (rAAV) particle comprising a first nucleotide sequence encoding a N-terminal portion of a prime editor fused at its C-terminus to an intein-N; and wherein the N-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 1417, 19, 24-53,55-65, 67-82, 84 that corresponds to amino acids 1- 844 or 1-1024 of SEQ ID NO: 14, wherein the first nucleotide sequence further comprises a nucleotide encoding a sgRNA sequence operably linked to a promoter, wherein the intein-N is a catalytically competent intein.
190. A composition comprising: a second recombinant adeno associated virus (rAAV) particle comprising a second nucleotide sequence encoding an intein-C fused to the N-terminus of a C-terminal portion of the prime editor; wherein the C-terminal portion of the prime editor comprises a portion of any one of SEQ ID NO: 14, 17, 19, 24-53, 55-65, 67-82, 84 that correspond to amino acids 845-1368 or 1025-1368 of SEQ ID NO: 14, and wherein three terminal amino acids at the N- terminus of the C-terminal portion are mutated from SEQ to SFQ, SFN, SEN, or CFN,wherein the intein-C is a catalytically competent intein.
191. A vector comprising a composition of any one of claims 187-190.
192. An AAV comprising a composition of any one of claims 187-190.
193. A cell comprising a composition of any one of claims 187-192.
194. A composition comprising any one of claims 187-193 and a pharmaceutically acceptable excipient.
195. A kit comprising the compositions of any one of claims 187-194.
196. The method of claim 176, wherein the ratio of the N-terminal portion to the C- terminal portion is 1:1.