Guide RNA for CRISPR / Cas editing systems
Patent Information
- Application Number
- JP2024504461
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-14
- Filing Date
- 2022-07-22
- Publication Date
- 2025-07-18
AI Technical Summary
The degradation of guide RNA (gRNA) by nucleases and inefficient transport to the nucleus pose challenges in achieving effective gene editing with CRISPR/Cas systems, limiting the persistence and efficiency of therapeutic applications.
Conjugation of gRNA with nuclear localization signals (NLS) via linkers, including cysteine residues, enhances gRNA potency and nuclear delivery, forming NLS-gRNAs that protect against cytoplasmic RNases and increase local concentration for efficient RNP formation.
NLS-gRNAs demonstrate significantly higher potency and editing efficiency compared to conventional gRNAs, with improved stability and resistance to cellular exonucleases, leading to enhanced gene editing outcomes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 255,322, filed July 23, 2021, and U.S. Provisional Patent Application Serial No. 63 / 255,927, filed October 14, 2021, the contents of which are incorporated by reference in their entireties for all purposes.
[0002] Sequence Listing Citation References The contents of the file named "BEM-011WO_ST26.xml", created on July 22, 2022 and having a size of 59.9 kilobytes, are incorporated herein by reference in their entirety. [Background technology]
[0003] CRISPR / Cas editing systems involve the use of guide RNA molecules (gRNAs) and associated enzymes in association with Cas endonucleases for use in gene editing and related systems including base editing. Briefly, one or more gRNA molecules are assembled with Cas proteins in a complex to guide a ribonucleic acid complex (RNP) to a specific DNA (e.g., in Cas9 and Cas12 systems) and / or RNA (e.g., in Cas13 systems) sequence.
[0004] The common form of gRNA used for therapeutic applications is a single non-natural RNA of approximately 100 nucleotides that forms a ribonucleoprotein with a Cas protein, such as Cas9. The ability to adapt CRISPR / Cas editing systems to new technologies (e.g., gene editing) requires that the guide RNA (gRNA) persists long enough in the target cell to allow the desired edits. Degradation of gRNA by nucleases is a key challenge to achieve the desired edits. Furthermore, gRNAs need to be assembled into ribonucleic acids (RNPs) and efficiently transported to the nucleus. Summary of the Invention
[0005] Methods, compositions, and kits are provided herein for enhancing the efficacy of gRNA for use in CRISPR-Cas systems. In some aspects, the present invention provides a method for producing gRNA conjugated to an NLS sequence (NLS-gRNA) with increased efficacy for use in CRISPR-Cas systems, for example, increased frequency of successful editing events. The NLS-gRNA of the present invention can provide better transport of gRNA to the nucleus to protect it from cytoplasmic RNases and increase the higher local concentration of gRNA for the formation of RNPs. The NLS-gRNA of the present invention has significantly higher efficacy compared to the corresponding gRNA that does not contain an NLS sequence, and also shows higher efficacy compared to highly modified gRNAs.
[0006] In one aspect, the invention provides, inter alia, a guide RNA (gRNA), comprising a nuclear localization signal (NLS) linked to the gRNA via a linker, wherein the linker comprises a cysteine residue conjugated to the 3' end of the gRNA.
[0007] In some embodiments, the linker comprises a cysteine residue at the N-terminus. In some embodiments, the linker comprises a cysteine residue at the C-terminus. In some embodiments, the linker comprises a cysteine residue at an internal site within the linker.
[0008] In some embodiments, the linker is conjugated to the 3' end of the gRNA. In some embodiments, the linker is conjugated to the 5' end of the gRNA. In some embodiments, the linker is conjugated to an internal region within the gRNA. In some embodiments, the linker is conjugated to a first hairpin region within the gRNA. In some embodiments, the linker is conjugated to a second hairpin region within the gRNA. In some embodiments, the linker is conjugated to a bulge region within the gRNA. In some embodiments, the gRNA comprises one or more modifications. In some embodiments, the one or more modifications are 2'OMe modifications. In some embodiments, the one or more modifications comprise 2'-fluoro modifications. In some embodiments, the one or more modifications comprise phosphorothioate linkages.
[0009] In some embodiments, the gRNA does not include a backbone modification. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, 6, 7, 8, and 9 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, 6, 7, and 8 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, 6, and 7 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, and 6 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, and 5 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, and 4 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, and 3 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1 and 2 nucleotides from the 3' end of the gRNA. In some embodiments, the one or more modifications occur at 1 nucleotide from the 3' end of the gRNA.
[0010] In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, 6, 7, 8, and 9 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, 6, 7, and 8 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, 6, and 7 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, 5, and 6 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, 4, and 5 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, 3, and 4 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1, 2, and 3 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1 and 2 nucleotides from the 5' end of the gRNA. In some embodiments, the one or more modifications occur at 1 nucleotide from the 5' end of the gRNA.
[0011] In some embodiments, more than 10% of the gRNAs are modified. In some embodiments, more than 20% of the gRNAs are modified. In some embodiments, more than 30% of the gRNAs are modified. In some embodiments, more than 35% of the gRNAs are modified. In some embodiments, more than 40% of the gRNAs are modified. In some embodiments, more than 45% of the gRNAs are modified. In some embodiments, more than 50% of the gRNAs are modified. In some embodiments, more than 55% of the gRNAs are modified. In some embodiments, more than 60% of the gRNAs are modified. In some embodiments, more than 65% of the gRNAs are modified. In some embodiments, more than 70% of the gRNAs are modified. In some embodiments, more than 75% of the gRNAs are modified. In some embodiments, more than 80% of the gRNAs are modified. In some embodiments, more than 85% of the gRNAs are modified. In some embodiments, more than 88% of the gRNAs are modified. In some embodiments, more than 90% of the gRNAs are modified. In some embodiments, more than 95% of the gRNAs are modified.
[0012] In some embodiments, less than 10% of the gRNAs are modified. In some embodiments, less than 20% of the gRNAs are modified. In some embodiments, less than 30% of the gRNAs are modified. In some embodiments, less than 35% of the gRNAs are modified. In some embodiments, less than 40% of the gRNAs are modified. In some embodiments, less than 45% of the gRNAs are modified. In some embodiments, less than 50% of the gRNAs are modified. In some embodiments, less than 55% of the gRNAs are modified. In some embodiments, less than 60% of the gRNAs are modified. In some embodiments, less than 65% of the gRNAs are modified. In some embodiments, less than 70% of the gRNAs are modified. In some embodiments, less than 75% of the gRNAs are modified. In some embodiments, less than 80% of the gRNAs are modified. In some embodiments, less than 85% of the gRNAs are modified. In some embodiments, less than 88% of the gRNAs are modified. In some embodiments, less than 90% of the gRNAs are modified. In some embodiments, less than 95% of the gRNAs are modified.
[0013] In some embodiments, the gRNA is conjugated to one or more NLS sequences. In some embodiments, the gRNA may comprise about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the 3' end, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the 5' end, or combinations thereof (e.g., one or more NLSs at the 3' end and one or more NLSs at the 5' end). When more than one NLS is present, each may be selected independently of the other, such that a single NLS may be present in more than one copy and / or in one or more copies in combination with one or more other NLSs.
[0014] Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 41), an NLS from nucleoplasmin (e.g., the nucleoplasmin bisecting NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 42)), the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 43) or RQRRNELKRSP (SEQ ID NO: 44), the hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 45), the IBB domain from importin-alpha having the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 46), the sequences VSRKRPRP (SEQ ID NO: 47) and PPKKARED (SEQ ID NO: 48) of the sarcoma T protein, the sequence POPKKKPL (SEQ ID NO: 49) of human p53, the mouse c-abl IV sequence SALIKKKKKMAP (sequence number 50), influenza virus NS1 sequences DRLRR (sequence number 51) and PKQKKRK (sequence number 52), hepatitis virus delta antigen sequence RKLKKKIKKL (sequence number 53), mouse Mx1 protein sequence REKKKFLKRR (sequence number 54), human poly(ADP-ribose) polymerase sequence KRKGDEVDGVDEVAKKKSKK (sequence number 55), and steroid hormone receptor (human) glucocorticoid sequence RKCLQAGMNLEARKTKK (sequence number 56).
[0015] In some embodiments, the NLS is derived from Simian Virus 40 (SV40). In some embodiments, the NLS comprises the amino acid sequence of KKKRKV (SEQ ID NO:57). In some embodiments, the NLS comprises a bipartite NLS. In some embodiments, the NLS comprises a bipartite NLS having an SV40 NLS.
[0016] In some embodiments, the linker further comprises a peptide spacer. In some embodiments, the peptide spacer comprises more than 2 amino acids. In some embodiments, the peptide spacer comprises more than 3 amino acids. In some embodiments, the peptide spacer comprises more than 4 amino acids. In some embodiments, the peptide spacer comprises more than 5 amino acids. In some embodiments, the peptide spacer comprises more than 6 amino acids. In some embodiments, the peptide spacer comprises more than 7 amino acids. In some embodiments, the peptide spacer comprises more than 8 amino acids. In some embodiments, the peptide spacer comprises more than 9 amino acids. In some embodiments, the peptide spacer comprises more than 10 amino acids. In some embodiments, the peptide spacer comprises more than 12 amino acids. In some embodiments, the peptide spacer comprises more than 15 amino acids. In some embodiments, the peptide spacer comprises more than 18 amino acids. In some embodiments, the peptide spacer comprises more than 20 amino acids. In some embodiments, the peptide spacer comprises more than 25 amino acids. In some embodiments, the peptide spacer comprises more than 30 amino acids.
[0017] In some embodiments, the peptide spacer comprises 2-30 amino acids. In some embodiments, the peptide spacer comprises 5-25 amino acids. In some embodiments, the peptide spacer comprises 7-20 amino acids. In some embodiments, the peptide spacer comprises 7-15 amino acids. In some embodiments, the peptide spacer comprises 7-12 amino acids.
[0018] In some embodiments, the peptide spacer comprises about 5 amino acids. In some embodiments, the peptide spacer comprises about 7 amino acids. In some embodiments, the peptide spacer comprises about 8 amino acids. In some embodiments, the peptide spacer comprises about 9 amino acids. In some embodiments, the peptide spacer comprises about 10 amino acids. In some embodiments, the peptide spacer comprises about 11 amino acids. In some embodiments, the peptide spacer comprises about 12 amino acids. In some embodiments, the peptide spacer comprises about 13 amino acids. In some embodiments, the peptide spacer comprises about 14 amino acids. In some embodiments, the peptide spacer comprises about 15 amino acids.
[0019] In some embodiments, the peptide spacer comprises the amino acid sequence of KRTADGSEFESP (SEQ ID NO:58). In some embodiments, the peptide spacer is 70% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 75% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 80% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 85% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 90% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 92% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 95% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 97% identical to the amino acid sequence of KRTADGSEFESP. In some embodiments, the peptide spacer is 99% identical to the amino acid sequence of KRTADGSEFESP.
[0020] In some embodiments, the linker further comprises a chemical moiety that conjugates the gRNA to the peptide spacer or the NLS.
[0021] In embodiments, the gRNA is conjugated to the NLS via a linker, which in embodiments comprises a chemical moiety (e.g., L) and / or a peptide moiety (e.g., a peptide spacer).
[0022] In embodiments, the gRNA is directly conjugated to the NLS via a chemical moiety (e.g., L). In embodiments, the chemical moiety (e.g., L) is non-peptidic. In embodiments, the chemical moiety (e.g., L) is covalently attached to both the gRNA and the NLS.
[0023] In embodiments, the gRNA is conjugated to the NLS via a peptide moiety (e.g., a peptide spacer). In embodiments, the peptide moiety (e.g., a peptide spacer) is covalently linked to both the gRNA and the NLS.
[0024] In embodiments, the gRNA is conjugated to the NLS via a linker that includes both a chemical moiety (e.g., L) and a peptide moiety (e.g., a peptide spacer). In embodiments, such a conjugate can have a structure according to formula (I), where the chemical moiety L (e.g., a non-peptidic chemical moiety) is covalently attached to the gRNA and the peptide spacer, and the peptide spacer is covalently attached to the NLS. [ka]
[0025] In embodiments, the gRNA is conjugated to the NLS via a peptide spacer or a chemical moiety (e.g., L) covalently attached to the C-terminus of the NLS amino acid sequence.
[0026] In embodiments, the gRNA is conjugated to the NLS via a peptide spacer or a chemical moiety (e.g., L) covalently attached to the N-terminus of the NLS amino acid sequence.
[0027] In embodiments, the gRNA is conjugated to a peptide spacer or NLS via a chemical moiety (e.g., L) covalently attached to the 3' end of the gRNA.
[0028] In embodiments, the gRNA is conjugated to a peptide spacer or NLS via a chemical moiety (e.g., L) covalently attached to the 5' end of the gRNA.
[0029] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to a thiol-containing residue (eg, a cysteine residue) of the NLS.
[0030] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to a selenium-containing residue (eg, a selenocysteine residue) of the NLS.
[0031] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to an amino-containing residue (eg, a lysine residue) of the NLS.
[0032] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to a phenol-containing residue (eg, a tyrosine residue) of the NLS.
[0033] In embodiments, the amino acid residue used to form the linker (eg, a thiol, selenium, amino, or phenol-containing residue as described herein) comprises a chemical modification.
[0034] In some embodiments, the guide RNA further comprises a nucleic acid linker sequence. In some embodiments, the nucleic acid linker sequence is an RNA sequence.
[0035] In some embodiments, the nucleic acid linker sequence is located at the 5' and / or 3' end of the guide RNA sequence.
[0036] In some embodiments, the nucleic acid linker comprises about 1-50 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-45 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-40 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-35 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-30 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-25 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-20 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-15 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-10 nucleotides. In some embodiments, the nucleic acid linker comprises about 1-5 nucleotides.
[0037] In some embodiments, the nucleic acid linker comprises about 5 nucleotides, about 10 nucleotides, about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 45 nucleotides, or about 50 nucleotides.
[0038] In some embodiments, the guide RNA does not include a nucleic acid linker. In some embodiments, the nucleic acid linker comprises about 1 nucleotide. In some embodiments, the nucleic acid linker comprises about 2 nucleotides. In some embodiments, the nucleic acid linker comprises about 3 nucleotides. In some embodiments, the nucleic acid linker comprises about 4 nucleotides. In some embodiments, the nucleic acid linker comprises about 5 nucleotides. In some embodiments, the nucleic acid linker comprises about 6 nucleotides. In some embodiments, the nucleic acid linker comprises about 7 nucleotides. In some embodiments, the nucleic acid linker comprises about 8 nucleotides. In some embodiments, the nucleic acid linker comprises about 9 nucleotides. In some embodiments, the nucleic acid linker comprises about 10 nucleotides. In some embodiments, the nucleic acid linker comprises about 11 nucleotides. In some embodiments, the nucleic acid linker comprises about 12 nucleotides. In some embodiments, the nucleic acid linker comprises about 13 nucleotides. In some embodiments, the nucleic acid linker comprises about 14 nucleotides. In some embodiments, the nucleic acid linker comprises about 15 nucleotides. In some embodiments, the nucleic acid linker comprises about 16 nucleotides. In some embodiments, the nucleic acid linker comprises about 17 nucleotides. In some embodiments, the nucleic acid linker comprises about 18 nucleotides. In some embodiments, the nucleic acid linker comprises about 19 nucleotides. In some embodiments, the nucleic acid linker comprises about 20 nucleotides. In some embodiments, the nucleic acid linker comprises about 21 nucleotides. In some embodiments, the nucleic acid linker comprises about 22 nucleotides. In some embodiments, the nucleic acid linker comprises about 23 nucleotides. In some embodiments, the nucleic acid linker comprises about 24 nucleotides. In some embodiments, the nucleic acid linker comprises about 25 nucleotides.
[0039] In some embodiments, the nucleic acid linker comprises about 50-100 nucleotides. In some embodiments, the nucleic acid linker comprises about 100-150 nucleotides. In some embodiments, the nucleic acid linker comprises about 150-200 nucleotides. In some embodiments, the nucleic acid linker comprises about 200-500 nucleotides.
[0040] In some embodiments, the nucleic acid linker sequence is a linear linker sequence. In some embodiments, the linker sequence is a non-linear sequence. In some embodiments, the linker sequence comprises an RNA secondary structure.
[0041] In some embodiments, the nucleic acid linker sequence is positioned at the 3' and / or 5' end of the guide RNA sequence.
[0042] In some embodiments, the gRNA containing an NLS improves base editing efficiency compared to a gRNA without an NLS. In some embodiments, the gRNA containing an NLS improves base editing efficiency at least 1.5-fold compared to a gRNA without an NLS. In some embodiments, the gRNA containing an NLS improves base editing efficiency at least 2-fold compared to a gRNA without an NLS. In some embodiments, the gRNA containing an NLS improves base editing efficiency at least 2.5-fold compared to a gRNA without an NLS. In some embodiments, the gRNA containing an NLS improves base editing efficiency at least 3-fold compared to a gRNA without an NLS. In some embodiments, the gRNA containing an NLS improves base editing efficiency at least 4-fold compared to a gRNA without an NLS. In some embodiments, the gRNA containing an NLS improves base editing efficiency at least 5-fold compared to a gRNA without an NLS.
[0043] In some embodiments, the guide RNA further comprises a direct repeat sequence found in a naturally occurring CRISPR system.
[0044] In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is a tracrRNA. In some embodiments, the gRNA is a crRNA.
[0045] In some embodiments, the guide RNA comprises a clustered regularly interspaced short palindromic repeats (CRISPR) RNA (crRNA). In some embodiments, the guide RNA further comprises a trans-activating RNA (tracrRNA).
[0046] In some embodiments, the crRNA is modified. In some embodiments, the tracrRNA is modified. In some embodiments, the crRNA and / or comprises chemically modified nucleotides. In some embodiments, the tracrRNA comprises additional sequences that maintain folding. In some embodiments, the linker comprises chemically modified nucleotides.
[0047] In some embodiments, modifications to the crRNA, tracrRNA, and / or linker include one or more of: 1) chemical modifications; 2) any nucleotide substitutions that preserve secondary structure; 3) alterations in GC content; 4) addition of sequences to maintain the predicted folding of the tracrRNA.
[0048] In some embodiments, a method is provided herein, wherein the NLS-gRNA is an extended guide RNA, or a Cas9 guide RNA, or a Cas13 guide RNA, or a Cas12 guide RNA, such as a Cas12a guide RNA, a Cas12b guide RNA, a Cas12c guide RNA, a Cas12d guide RNA, a Cas12e guide RNA, a Cas12f guide RNA, a Cas12g guide RNA, a Cas12h guide RNA, a Cas12i guide RNA, a Cas12j guide RNA, or a Cas12k guide RNA. Thus, in some embodiments, the NLS-gRNA is an extended guide RNA. In some embodiments, the NLS-gRNA is a Cas9 guide RNA. In some embodiments, the NLS-gRNA is a Cas13 guide RNA. In some embodiments, the NLS-gRNA is a Cas12 guide RNA. In some embodiments, the NLS-gRNA is a Cas12a guide RNA. In some embodiments, the NLS-gRNA is a Cas12b guide RNA. In some embodiments, the NLS-gRNA is a Cas12c guide RNA. In some embodiments, the NLS-gRNA is a Cas12d guide RNA. In some embodiments, the NLS-gRNA is a Cas12e guide RNA. In some embodiments, the NLS-gRNA is a Cas12f guide RNA. In some embodiments, the NLS-gRNA is a Cas12g guide RNA. In some embodiments, the NLS-gRNA is a Cas12h guide RNA. In some embodiments, the NLS-gRNA is a Cas12i guide RNA. In some embodiments, the NLS-gRNA is a Cas12j guide RNA. In some embodiments, the NLS-gRNA is a Cas12k guide RNA.
[0049] In some embodiments, the NLS-gRNA comprises one or more of the following: a spacer, a lower stem, a bulge, an upper stem, a nexus, and a hairpin.
[0050] In some embodiments, the stem loop comprises a GC base pair.
[0051] In some embodiments, methods are provided herein, wherein the NLS-gRNA is produced at a yield of about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more. Thus, in some embodiments, the NLS-gRNA is produced at a yield of about 50%. In some embodiments, the NLS-gRNA is produced at a yield of about 55%. In some embodiments, the NLS-gRNA is produced at a yield of about 60%. In some embodiments, the NLS-gRNA is produced at a yield of about 65%. In some embodiments, the NLS-gRNA is produced at a yield of about 70%. In some embodiments, the NLS-gRNA is produced at a yield of about 75%. In some embodiments, the NLS-gRNA is produced at a yield of about 80%. In some embodiments, the NLS-gRNA is produced at a yield of about 85%. In some embodiments, the NLS-gRNA is produced at a yield of about 90%. In some embodiments, the NLS-gRNA is produced at a yield of about 95%. In some embodiments, the NLS-gRNA is produced at a yield of greater than about 99%.
[0052] In some embodiments, NLS-gRNA is produced with a 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more improved yield compared to traditional synthesis methods. Thus, in some embodiments, NLS-gRNA is produced with a 50% improved yield compared to traditional synthesis methods. In some embodiments, NLS-gRNA is produced with a 55% improved yield compared to traditional synthesis methods. In some embodiments, NLS-gRNA is produced with a 60% improved yield compared to traditional synthesis methods. In some embodiments, NLS-gRNA is produced with a 65% improved yield compared to traditional synthesis methods. In some embodiments, NLS-gRNA is produced with a 70% improved yield compared to traditional synthesis methods. In some embodiments, NLS-gRNA is produced with a 75% improved yield compared to traditional synthesis methods. In some embodiments, NLS-gRNA is produced with an 80% improved yield compared to traditional synthesis methods. In some embodiments, NLS-gRNA is produced with an 85% improved yield compared to conventional synthesis methods. In some embodiments, NLS-gRNA is produced with a 90% improved yield compared to conventional synthesis methods. In some embodiments, NLS-gRNA is produced with a 95% improved yield compared to conventional synthesis methods. In some embodiments, NLS-gRNA is produced with a 99% improved yield compared to conventional synthesis methods. In some embodiments, NLS-gRNA is produced with a greater than 99% improved yield compared to conventional synthesis methods.
[0053] In some embodiments, the NLS-gRNA has a length of about 40 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, or more than about 200 nucleotides. Thus, in some embodiments, the NLS-gRNA has a length of about 40 nucleotides. In some embodiments, the NLS-gRNA has a length of about 100 nucleotides. In some embodiments, the NLS-gRNA has a length of about 125 nucleotides. In some embodiments, the NLS-gRNA has a length of about 150 nucleotides. In some embodiments, the NLS-gRNA has a length of about 175 nucleotides. In some embodiments, the NLS-gRNA has a length of about 200 nucleotides. In some embodiments, the NLS-gRNA has a length of more than about 200 nucleotides.
[0054] In some embodiments, the length of the NLS-gRNA is Cas-dependent. For example, in some embodiments, the length of the NLS-gRNA for Cas12a is greater than 40 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is greater than 123 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is 125-200 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is 125-250 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is 125-300 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is 125-350 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is 125-400 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is 125-450 nucleotides. In some embodiments, the length of the NLS-gRNA for Cas9 is between 125 and 500 nucleotides.
[0055] In some embodiments, the NLS-gRNA comprises one or more backbone modifications.
[0056] In some embodiments, the one or more backbone modifications comprise a 2'O-methyl or phosphorothioate modification. Thus, in some embodiments, the one or more backbone modifications comprise a 2'O-methyl modification. In some embodiments, the one or more backbone modifications comprise a phosphorothioate modification.
[0057] In some embodiments, the one or more backbone modifications are selected from 2'-O-methyl 3'-phosphorothioate, 2'-O-methyl, 2'-ribo 3'-phosphorothioate, 2'-fluoro, 2'-O-methoxyethyl morpholino (PMO), locked nucleic acid (LNA), deoxy, or 5' phosphate modifications. Thus, in some embodiments, the one or more backbone modifications include a 2'-O-methyl 3'-phosphorothioate modification. In some embodiments, the one or more backbone modifications include a 2'-O-methyl modification. In some embodiments, the one or more modifications include a 2'-ribo 3'-phosphorothioate modification. In some embodiments, the one or more modifications include a 2'-fluoro modification. In some embodiments, the one or more modifications include a 2'-O-methoxyethyl morpholino (PMO). In some embodiments, the one or more modifications include a locked nucleic acid (LNA). In some embodiments, the one or more modifications include a deoxy modification. In some embodiments, the one or more modifications include a 5' phosphate modification.
[0058] A variety of modified RNA bases are known in the art and include, for example, 2'-O-methoxy-ethyl bases (2'-MOE), such as 2-methoxyethoxy A, 2-methoxyethoxy MeC, 2-methoxyethoxy G, 2-methoxyethoxy T. Other modified bases include, for example, 2'-O-methyl RNA bases, and fluoro bases. A variety of fluoro bases are known, including, for example, fluoro C, fluoro U, fluoro A, fluoro G bases. A variety of 2'-O methyl modifications can also be used with the methods described herein. For example, the following RNAs containing one or more of the following 2'-O methyl modifications can be used with the described methods: 2'-OMe-5-methyl-rC, 2'-OMe-rT, 2'-OMe-rI, 2'-OMe-2-amino-rA, aminolinker-C6-rC, aminolinker-C6-rU, 2'-OMe-5-Br-rU, 2'-OMe-5-I-rU, 2-OMe-7-deaza-rG.
[0059] In some embodiments, the RNA comprises one or more of the following modifications: phosphorothioate, 2'O-methyl, 2'fluoro (2'F), deoxy. In some embodiments, the RNA comprises a 2'OMe modification at the 3' end. In some embodiments, the RNA comprises a 2'OMe modification at the 5' end. In some embodiments, the RNA comprises a 2'OMe modification at the 3' end and at the 5' end. In some embodiments, the RNA comprises one or more of the following modifications: 2'-O-2-Methoxyethyl (MOE), locked nucleic acid, bridged nucleic acid, unlocked nucleic acid, peptide nucleic acid, morpholino nucleic acid. In some embodiments, the RNA comprises one or more of the following base modifications: 2,6-diaminopurine, 2-aminopurine, pseudouracil, N1-methyl-pseudouracil, 5'methylcytosine, 2'pyrimidinone (zebularine), thymine. Other modified bases include, for example, 2-aminopurine, 5-bromo-dU, deoxyuridine, 2,6-diaminopurine (2-amino-dA), dideoxy-C, deoxyinosine, hydroxymethyl-dC, inverted-dT, iso-dG, iso-dC, inverted-dideoxy-T, 5-methyl-dC, 5-methyl-dC, 5-nitroindole, Super T®, 2'-Fr(C,U), 2'-NH2-r(C,U), 2,2'-anhydro-U, 3'-deoxy-r(A,C,G,U), 3'-O-methyl-r(A,C,G,U), rT, rI, 5-methyl-rC, 2-amino-rA, rspacer (abasic), 7-deaza-rG, 7-deaza-rA, 8-oxo-rG, 5-halogenated-rU, N-alkylated-rN.
[0060] Other chemically modified RNAs can be used herein. For example, the RNA can include modified bases, such as, for example, 5'Int,3'Azide (NHS ester), 5'Hexynyl, 5'Int,3'5-Octadienyl dU, 5',Int Biotin (Azide), 5',Int6-FAM (Azide), and 5',Int5-TAMRA (Azide). Other examples of RNA nucleotide modifications that can be used with the methods described herein include phosphorylation modifications, such as, for example, 5' phosphorylation and 3' phosphorylation. The RNA can also have one or more of the following modifications: amino modification, biotinylation, thiol modification, alkyne modifier, adenylation, azide (NHS ester), cholesterol-TEG, and digoxigenin (NHS ester).
[0061] In some embodiments, the method produces NLS-gRNA with a purity of about 50%, 60%, 70%, 80%, 90%, or greater than 90%. Thus, in some embodiments, the method produces NLS-gRNA with a purity of about 50%. In some embodiments, the method produces NLS-gRNA with a purity of about 60%. In some embodiments, the method produces NLS-gRNA with a purity of about 70%. In some embodiments, the method produces NLS-gRNA with a purity of about 80%. In some embodiments, the method produces NLS-gRNA with a purity of about 90%. In some embodiments, the method produces NLS-gRNA with a purity of about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or greater than 99%. In some embodiments, the method produces NLS-gRNA with a purity of about 91%. In some embodiments, the method produces NLS-gRNA with a purity of about 92%. In some embodiments, the method produces NLS-gRNA with a purity of about 93%. In some embodiments, the method produces NLS-gRNA with a purity of about 94%. In some embodiments, the method produces NLS-gRNA with a purity of about 95%. In some embodiments, the method produces NLS-gRNA with a purity of about 96%. In some embodiments, the method produces NLS-gRNA with a purity of about 97%. In some embodiments, the method produces NLS-gRNA with a purity of about 98%. In some embodiments, the method produces NLS-gRNA with a purity of about 99%. In some embodiments, the method produces NLS-gRNA with a purity of greater than about 99%.
[0062] In one aspect, the present invention provides, inter alia, a composition comprising a guide RNA (gRNA), the gRNA comprising a nuclear localization signal (NLS) linked to the gRNA via a linker, the linker comprising a cysteine residue conjugated to the 3' end of the gRNA, and the NLS-guide RNA is encapsulated in a lipid nanoparticle (LNP). In one aspect, the present invention provides, inter alia, a composition comprising a guide RNA (gRNA), the gRNA comprising a nuclear localization signal (NLS) linked to the gRNA via a linker, the linker comprising a cysteine residue conjugated to the 3' end of the gRNA, and the NLS-guide RNA is associated with a lipid nanoparticle (LNP).
[0063] In some embodiments, the composition comprises a nuclease. In some embodiments, the composition comprises a nucleic acid encoding the nuclease. In some embodiments, the composition comprises an mRNA encoding the nuclease.
[0064] In some embodiments, the nuclease is conjugated to an NLS. In some embodiments, the Cas protein is conjugated to an NLS. In some embodiments, the Cas protein does not comprise an NLS. In some embodiments, the Cas protein is not conjugated to an NLS. In some embodiments, the Cas9 protein does not comprise an NLS. In some embodiments, the Cas9 protein is not conjugated to an NLS.
[0065] In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 1:1. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 2:1. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 3:1. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 4:1. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 5:1. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 6:1. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 7:1. In some embodiments, the composition comprises an mRNA encoding an NLS-gRNA and a nuclease in a weight ratio of 8:1. In some embodiments, the composition comprises an NLS-gRNA and an mRNA encoding a nuclease in a weight ratio of 9:1. In some embodiments, the composition comprises an NLS-gRNA and an mRNA encoding a nuclease in a weight ratio of 10:1. In some embodiments, the composition comprises an NLS-gRNA and an mRNA encoding a nuclease in a weight ratio of 12:1. In some embodiments, the composition comprises an NLS-gRNA and an mRNA encoding a nuclease in a weight ratio of 15:1.
[0066] In some embodiments, the nuclease is a CRISPR class 2 type II enzyme. In some embodiments, the nuclease is a CRISPR class 2 type V enzyme. In some embodiments, the nuclease is a CRISPR class 2 type VI enzyme. In some embodiments, the nuclease is Cas9, Cpf1, SaCas9, Cas12, Cas13, or a modified version thereof. Thus, in some embodiments, the nuclease is Cas9 or a modified version thereof. In some embodiments, the nuclease is Cpf1 or a modified version thereof. In some embodiments, the nuclease is Staphylococcus aureus Cas9 (SaCas9) or a modified version thereof. In some embodiments, the nuclease is Streptococcus thermophilus 1 Cas9 (St1Cas9) or a modified version thereof. In some embodiments, the nuclease is Streptococcus pyogenes Cas9 (SpCas9) or a modified version thereof. In some embodiments, the nuclease is Cas12 or a modified version thereof. In some embodiments, the nuclease is Cas13 or a modified version thereof.
[0067] In some embodiments, the Cas9 comprises a nuclease-dead Cas9 (dCas9). In some embodiments, the Cas9 comprises a Cas9 nickase (nCas9). In some embodiments, the Cas9 comprises a nuclease-active Cas9.
[0068] In some embodiments, the nuclease domain is fused to a heterologous polypeptide. In some embodiments, the heterologous polypeptide comprises an effector domain capable of performing modifications on nucleic acids (e.g., DNA). For example, the DNA effector domain can be a deaminase domain, such as a cytidine deaminase domain, a cytosine domain, or an adenosine deaminase domain. In certain embodiments, the deaminase domain is a cytidine deaminase domain, such as an APOBEC or AID cytidine deaminase. For base editing proteins capable of deaminating cytidine to uridine, for example, to induce a C to T mutation in a DNA molecule, the cytidine deaminase can be a deaminase from the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases. In some embodiments, the heterologous polypeptide is a cytidine or cytosine deaminase domain. In some embodiments, the heterologous polypeptide is a cytosine deaminase domain. In some embodiments, the heterologous polypeptide is a cytidine deaminase domain. In some embodiments, the heterologous polypeptide is an adenosine or adenine deaminase domain. In some embodiments, the heterologous polypeptide is an adenosine domain. In some embodiments, the heterologous polypeptide is an adenine domain.
[0069] In some embodiments, the heterologous polypeptide is an adenosine deaminase variant domain. In some embodiments, the adenosine deaminase variant domain comprises one or more mutations with respect to SEQ ID NO: 3. In some embodiments, the adenosine deaminase variant domain comprises V82G. In some embodiments, the adenosine deaminase variant domain comprises Y147T / D. In some embodiments, the adenosine deaminase variant domain comprises Q154S. In some embodiments, the adenosine deaminase variant domain comprises L36H. In some embodiments, the adenosine deaminase variant domain comprises I76Y. In some embodiments, the adenosine deaminase variant domain comprises F149Y. In some embodiments, the adenosine deaminase variant domain comprises N157K. In some embodiments, the adenosine deaminase variant domain comprises V82G, Y147T / D, and Q154S. In some embodiments, the adenosine deaminase variant domain comprises V82G, Y147T / D, Q154S, and L36H. In some embodiments, the adenosine deaminase variant domain comprises V82G, Y147T / D, Q154S, and I76Y. In some embodiments, the adenosine deaminase variant domain comprises V82G, Y147T / D, Q154S, and F149Y. In some embodiments, the adenosine deaminase variant domain comprises V82G, Y147T / D, Q154S, and N157K. In some embodiments, the adenosine deaminase variant domain comprises V82G, Y147T / D, Q154S, and D167N. In some embodiments, the adenosine deaminase variant domain comprises V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N. In some embodiments, the adenosine deaminase domain comprises the mutations I76Y, V82G, Y147T, and Q154S. In some embodiments, the adenosine deaminase domain comprises the mutations L36H, V82G, Y147T, Q154S, and N157K.In some embodiments, the adenosine deaminase domain comprises the mutations V82G, Y147D, F149Y, Q154S, and D167N. In some embodiments, the adenosine deaminase domain comprises the mutations L36H, V82G, Y147D, F149Y, Q154S, N157K, and D167N. In some embodiments, the adenosine deaminase domain comprises the mutations L36H, I76Y, V82G, Y147T, Q154S, and N157K. In some embodiments, the adenosine deaminase domain comprises the mutations I76Y, V82G, Y147D, F149Y, Q154S, and D167N. In some embodiments, the adenosine deaminase domain comprises the mutations Y147D, F149Y, and D167N. In some embodiments, the adenosine deaminase domain comprises the mutations L36H, I76Y, V82G, Q154S, and N157K. In some embodiments, the adenosine deaminase domain comprises the mutations I76Y, V82G, and Q154S. In some embodiments, the adenosine deaminase domain comprises the mutations L36H, I76Y, V82G, Y147D, F149Y, Q154S, N157K, and D167N.
[0070] In some embodiments, the heterologous polypeptide is fused to the N-terminus of the nuclease domain. In some embodiments, the heterologous polypeptide is fused to the C-terminus of the nuclease domain. In some embodiments, the heterologous polypeptide is internal to the nuclease domain. In some embodiments, the heterologous polypeptide is fused to the N-terminus of Cas9. In some embodiments, the heterologous polypeptide is fused to the C-terminus of Cas9. In some embodiments, the heterologous polypeptide is internal to Cas9. In some embodiments, the adenosine deaminase variant is fused to the N-terminus of Cas9. In some embodiments, the adenosine deaminase variant is fused to the C-terminus of Cas9. In some embodiments, the adenosine deaminase variant is internal to Cas9.
[0071] In some embodiments, the NLS-gRNA is suitable for use with a CRISPR / Cas system. In some embodiments, the NLS-gRNA is suitable for use with a CRISPR class 2 type II enzyme. In some embodiments, the NLS-gRNA is suitable for use with a CRISPR class 2 type V enzyme. In some embodiments, the NLS-gRNA is suitable for use with a CRISPR class 2 type VI enzyme. In some embodiments, the NLS-gRNA is suitable for use with Cas9, Cpf1, SaCas9, Cas12, Cas13, or modified versions thereof. Thus, in some embodiments, the NLS-gRNA is suitable for use with Cas9 or modified versions thereof. In some embodiments, the NLS-gRNA is suitable for use with Cpf1 or modified versions thereof. In some embodiments, the NLS-gRNA is suitable for use with SaCas9 or modified versions thereof. In some embodiments, the NLS-gRNA is suitable for use with Cas12 or modified versions thereof. In some embodiments, the NLS-gRNA is suitable for use with Cas13 or a modified version thereof. In some embodiments, the NLS-gRNA is in a complex with a Cas enzyme.
[0072] In some embodiments, an RNA sequence is included that is cleaved by the endonuclease activity of some Cas, e.g., Cas12a and Cas13, to linearize the gRNA before or during assembly with the Cas protein.
[0073] In some embodiments, the NLS-gRNA provides increased stability and resistance to cellular exonucleases compared to gRNAs that do not contain an NLS sequence, hi some embodiments, the NLS-gRNA provides increased editing events in target cells using a CRISPR / Cas editing system.
[0074] In some embodiments, the NLS-gRNA is in a complex with a CRISPR class 2 type II enzyme. In some embodiments, the NLS-gRNA is in a complex with a CRISPR class 2 type V enzyme. In some embodiments, the NLS-gRNA is in a complex with a CRISPR class 2 type VI enzyme. In some embodiments, the NLS-gRNA is in a complex with Cas9, Cpf1, SaCas9, Cas12, Cas13, or modified versions thereof.
[0075] In some aspects, a Cas protein complex is provided, the complex comprising a Cas nuclease and an NLS-gRNA.
[0076] In some embodiments, the Cas nuclease is a CRISPR class 2 type II enzyme. In some embodiments, the Cas nuclease is a CRISPR class 2 type V enzyme. In some embodiments, the Cas nuclease is a CRISPR class 2 type VI enzyme. In some embodiments, the Cas nuclease is selected from Cas9, Cpf1, SaCas9, Cas12, Cas13, or modified versions thereof.
[0077] In some embodiments, provided herein are methods for targeted transcriptional activation, targeted transcriptional repression, targeted epigenome modification, or targeted genomic modification, the methods comprising introducing into a eukaryotic cell (a) an NLS-conjugated guide RNA (NLS-gRNA); (b) at least one CRISPR / Cas protein or a nucleic acid encoding at least one CRISPR / Cas protein, wherein interaction of (a) and (b) with a target sequence in chromosomal DNA results in targeted transcriptional activation, targeted transcriptional repression, targeted epigenome modification, or targeted genomic modification.
[0078] In some embodiments, provided herein is a method for targeted RNA modification, the method comprising introducing into a eukaryotic cell (a) an NLS-conjugated guide RNA (NLS-gRNA), and (b) at least one CRISPR / Cas protein or a nucleic acid encoding at least one CRISPR / Cas protein, wherein interaction of (a) and (b) with an RNA expressed by chromosomal DNA results in modification of the RNA expressed by the chromosomal DNA.
[0079] In some embodiments, the RNA expressed by the chromosomal DNA is messenger RNA (mRNA).
[0080] In some aspects, the present invention provides a pharmaceutical composition comprising an NLS-gRNA of the present invention and a pharma- ceutically acceptable carrier.
[0081] In one aspect, the invention provides a composition comprising an engineered or non-naturally occurring CRISPR-associated Cas (CRISPR-Cas) system comprising, inter alia, a Cas protein, a gRNA comprising a nuclear localization signal (NLS) linked to the gRNA via a linker, the linker comprising a cysteine residue conjugated to the 3' end of the gRNA, wherein the gRNA is capable of forming a complex with the Cas protein and targeting the Cas9 protein to a target DNA.
[0082] In some embodiments, the gRNA comprises the nucleic acid sequence: 5'-CAGUAUGGACACUGUCCAAA-3' (SEQ ID NO:2).
[0083] In one aspect, the invention provides a composition comprising an engineered or non-naturally occurring CRISPR-associated Cas (CRISPR-Cas) system comprising, inter alia, (a) a saCas9 protein, (b) an adenosine deaminase variant fused to the Cas9 protein, and (c) a gRNA comprising a nuclear localization signal (NLS) linked to the gRNA via a linker, wherein the linker comprises a cysteine residue conjugated to the 3' end of the gRNA, and wherein the gRNA is capable of forming a complex with the saCas9 protein and targeting the saCas9 protein to a target DNA, wherein the adenosine deaminase variant comprises one or more of V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N with respect to SEQ ID NO:3, and wherein the gRNA comprises SEQ ID NO:2.
[0084] In one aspect, the present invention provides, inter alia, a method of treating a genetic disease in a subject in need of such treatment by administering to the subject a composition of the present invention (e.g., an NLS-gRNA).
[0085] In one aspect, the present invention provides, inter alia, a method of treating glycogen storage disease type 1a (GSD1a), the method comprising administering to a subject a composition of the present invention (e.g., an NLS-gRNA).
[0086] In some embodiments, provided herein are compositions comprising a gRNA conjugated to an NLS, wherein nuclear delivery of the composition is increased by about 2-5 fold relative to a composition comprising a gRNA without an NLS. In some embodiments, nuclear delivery of the composition is increased by about 2-fold relative to a composition comprising a gRNA without an NLS. In some embodiments, nuclear delivery of the composition is increased by about 3-fold relative to a composition comprising a gRNA without an NLS. In some embodiments, nuclear delivery in an increase of up to about 4-fold relative to a composition comprising a gRNA without an NLS. In some embodiments, nuclear delivery in an increase of up to about 5-fold relative to a composition comprising a gRNA without an NLS. In some embodiments, nuclear delivery in an increase of more than about 2-fold relative to a composition comprising a gRNA without an NLS. In some embodiments, nuclear delivery in an increase of 1.5-10 fold relative to a composition comprising a gRNA without an NLS. In some embodiments, nuclear delivery in an increase of more than about 10-fold relative to a composition comprising a gRNA without an NLS.
[0087] In some embodiments, the gRNA comprises a sequence having 70%, 80%, 90%, 95%, 99%, or 100% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 70% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 75% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 80% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 85% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 90% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 95% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 99% identity to any one of the sequences in Table 8. In some embodiments, the gRNA comprises a sequence having 100% identity to any one of the sequences in Table 8.
[0088] In some embodiments, provided herein are compositions comprising a gRNA conjugated to an NLS, wherein the gene editing efficiency is increased about 2-5 fold relative to a gRNA that does not contain an NLS. In some embodiments, the gene editing efficiency is increased about 2-fold relative to a gRNA that does not contain an NLS. In some embodiments, the gene editing efficiency is increased about 3-fold relative to a gRNA that does not contain an NLS. In some embodiments, the gene editing efficiency is increased about 4-fold relative to a gRNA that does not contain an NLS. In some embodiments, the gene editing efficiency is increased about 5-fold relative to a gRNA that does not contain an NLS. In some embodiments, the gene editing efficiency is increased about 1.5-10 fold relative to a gRNA that does not contain an NLS.
[0089] In some embodiments, the gRNA target sequence has 70%, 80%, 90%, 95%, 99% or 100% identity to SEQ ID NO:17. In some embodiments, the gRNA target sequence has 70% identity to SEQ ID NO:17. In some embodiments, the gRNA target sequence has 75% identity to SEQ ID NO:17. In some embodiments, the gRNA target sequence has 70%, 80% identity to SEQ ID NO:17. In some embodiments, the gRNA target sequence has 85% identity to SEQ ID NO:17. In some embodiments, the gRNA target sequence has 90% identity to SEQ ID NO:17. In some embodiments, the gRNA target sequence has 95% identity to SEQ ID NO:17. In some embodiments, the gRNA target sequence has 100% identity to SEQ ID NO:17.
[0090] In some embodiments, the gRNA targets one or more organs selected from the liver, kidney, brain, and heart, hi some embodiments, the gRNA targets the liver.
[0091] definition In order that the present invention may be more readily understood, certain terms are first defined below. Further definitions for these terms, as well as other terms, are set forth throughout the specification.
[0092] A or An: The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "an element" means one element or more than one element.
[0093] Approximately or about: As used herein, the term "approximately" or "about" when applied to one or more values of interest refers to a value similar to a stated reference value. In certain embodiments, the term "approximately" or "about" refers to a range of values that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater or less) of the stated reference value (except where such number exceeds 100% of possible values), unless otherwise stated or clear from the context.
[0094] Associated with: As the term is used herein, two events or entities are "associated" with each other when the presence, level and / or form of one correlates with the presence, level and / or form of the other. For example, a particular entity (e.g., a polypeptide) is considered to be associated with a particular disease, disorder or condition when its presence, level and / or form correlates with the incidence and / or susceptibility of the disease, disorder or condition (e.g., across a relevant population). In some embodiments, two or more entities are physically "associated" with each other such that they are in and maintain physical proximity to each other when they directly or indirectly interact. In some embodiments, two or more entities that are physically associated with each other are covalently bound to each other, and in some embodiments, two or more entities that are physically associated with each other are not covalently bound to each other but are non-covalently associated, for example, by hydrogen bonds, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.
[0095] "Adenosine deaminase" or "adenine deaminase" refers to a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be derived from any organism (e.g., eukaryotes, prokaryotes), including, but not limited to, algae, bacteria, fungi, plants, invertebrates (e.g., insects), and vertebrates (e.g., amphibians, mammals). In some embodiments, the adenosine deaminase is an adenosine deaminase variant having one or more modifications and capable of deaminating both adenine and cytosine in a target polynucleotide (e.g., DNA, RNA). In some embodiments, the target polynucleotide is single-stranded or double-stranded. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in single-stranded DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in RNA.
[0096] "Adenosine deaminase activity" means catalyzing the deamination of adenine or adenosine to guanine in a polynucleotide. In some embodiments, the adenosine deaminase variants provided herein maintain adenosine deaminase activity (e.g., at least about 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the activity of a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19)).
[0097] By "adenosine base editor (ABE)" is meant a base editor that comprises an adenosine deaminase.
[0098] By "Adenosine Base Editor 8 (ABE8) polypeptide" or "ABE8" is meant a base editor as defined herein, including an adenosine deaminase variant that comprises a modification at amino acid position 82 and / or 166 of the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 3). In some embodiments, ABE8 comprises further modifications relative to the reference sequence, as described herein.
[0099] By "Adenosine Base Editor 8 (ABE8) polynucleotide" is meant a polynucleotide that encodes an ABE8 polypeptide.
[0100] "Adenosine deaminase polynucleotide" refers to a polynucleotide that encodes an adenosine deaminase polypeptide. In certain embodiments, the adenosine deaminase polynucleotide encodes an adenosine deaminase variant that includes V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N. In some embodiments, the adenosine deaminase polynucleotide encodes an adenosine deaminase variant that includes one of the following combinations of modifications: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154 S+D167N, L36H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, or L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0101] In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is not naturally occurring. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase. In some embodiments, the adenosine deaminase is from a bacteria, such as E. coli, S. aureus, B. subtilis, S. typhi, S. putrefaciens, H. influenzae, C. crescentus, or G. sulfurreducens.
[0102] In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA (ecTadA) deaminase or a fragment thereof. In some embodiments, the ecTadA deaminase is a truncated ecTadA. For example, the truncated ecTadA may lack one or more N-terminal amino acids relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lack 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the ecTadA deaminase does not include an N-terminal methionine. In some embodiments, the TadA deaminase is an N-terminal truncated TadA. In certain embodiments, the TadA is any one of the TadAs described in PCT / US2017 / 045381, which is incorporated herein by reference in its entirety.
[0103] In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*7.10 comprising V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N. In some embodiments, the TadA variant is TadA*7.10 comprising a combination of modifications selected from among the following: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147T+Q154S+N157K, L36H+V ... 36H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, or L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N. In some embodiments, the TadA variant is MSP605, MSP680, MSP823, MSP824, MSP825, MSP827, MSP828, or MSP829.
[0104] Base editor: "Base editor (BE)" or "Nucleic acid base editor (NBE)" refers to an agent that binds to a polynucleotide and has nucleobase modifying activity. In various embodiments, the base editor comprises a nucleobase modifying polypeptide (e.g., a deaminase) and a polynucleotide programmable nucleotide binding domain in conjunction with a guide polynucleotide (e.g., a guide RNA). In various embodiments, the agent is a biomolecular complex that comprises a protein domain with base editing activity, i.e., a domain that can modify a base (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide programmable DNA binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein that comprises one or more domains with base editing activity. In another embodiment, the protein domain with base editing activity is linked to a guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase). In some embodiments, the protein domain with base editing activity is capable of deaminating a base in a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases in a DNA molecule. In some embodiments, a base editor can deaminate nitrogenous bases in DNA. In some embodiments, a base editor can deaminate nitrogenous bases in RNA. In some embodiments, a base editor can deaminate ribonucleosides. In some embodiments, a base editor can deaminate deoxyribonucleosides. In some embodiments, a base editor can deaminate cytosine. In some embodiments, a base editor can deaminate cytidine. In some embodiments, a base editor can deaminate adenosine. In some embodiments, a base editor can deaminate cytosine (C) or adenosine (A) in DNA. In some embodiments, a base editor can deaminate cytosine (C) and adenosine (A) in DNA.In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is a nuclease-inactive Cas9 (dCas9) fused to an adenosine deaminase. In some embodiments, the base editor is fused to an inhibitor of base excision repair (e.g., a UGI domain, or a dISN domain). In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and an inhibitor of base excision repair (e.g., a UGI domain or a dISN domain). In other embodiments, the base editor is an abasic base editor. Details of base editors are described in International PCT Application Nos. 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated by reference in its entirety.Also, Komor, AC, et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016), Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017), Komor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.”Nat Rev Genet.2018 Dec;19(12):770-788. doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.
[0105] Base editing activity: "Base editing activity" means acting to chemically change a base in a polynucleotide (e.g., by deaminating the base). In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is a cytidine deaminase activity, e.g., converting a targeted C·G to a T·A. In another embodiment, the base editing activity is an adenosine or adenine deaminase activity, e.g., converting an A·T to a G·C. In another embodiment, the base editing activity is a cytidine deaminase activity, e.g., converting a targeted C·G to a T·A, and an adenosine or adenine deaminase activity, e.g., converting an A·T to a G·C.
[0106] Base editor system: The term "base editor system" refers to a system for editing nucleobases of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises (1) a polynucleotide programmable nucleotide binding domain (e.g., Cas9), a deaminase domain, and a cytidine deaminase domain for deaminating nucleobases in a target nucleotide sequence, and (2) one or more guide polynucleotides (e.g., guide RNAs) in conjunction with the polynucleotide programmable nucleotide binding domain. In various embodiments, the base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain having a nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises (1) a base editor (BE) comprising a polynucleotide programmable DNA binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence, and (2) one or more guide RNAs in conjunction with the polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).
[0107] Biologically active: As used herein, the phrase "biologically active" refers to the characteristic of any agent that has activity in a biological system, particularly in an organism. For example, an agent that, when administered to an organism, has a biological effect on the organism is considered to be biologically active. In certain embodiments, if a peptide is biologically active, a portion of the peptide that shares at least one biological activity of the peptide is typically referred to as a "biologically active" portion.
[0108] Cleavage: As used herein, cleavage refers to the cleavage of the target nucleic acid produced by the nuclease of the CRISPR system described herein. In some embodiments, the cleavage event is double-stranded DNA cleavage. In some embodiments, the cleavage event is single-stranded DNA cleavage. In some embodiments, the cleavage event is single-stranded RNA cleavage. In some embodiments, the cleavage event is double-stranded RNA cleavage.
[0109] Complementary: "Complementary" or "complementarity" means that a nucleic acid can form hydrogen bond(s) with another nucleic acid sequence by either traditional Watson-Crick or Hoogsteen base pairing. Complementary base pairing includes GC and AT base pairing as well as base pairing with universal bases such as inosine. Percent complementarity indicates the percentage of adjacent residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, or 10 nucleotides out of a total of 10 nucleotides in a first oligonucleotide base paired to a second nucleic acid sequence having 10 nucleotides represents 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). To determine the percentage of complementarity, the percentage of adjacent residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence is calculated and rounded to the nearest integer (e.g., 12, 13, 14, 15, 16, or 17 nucleotides out of a total of 23 nucleotides in a first oligonucleotide that are base paired to a second nucleic acid sequence having 23 nucleotides represents 52%, 57%, 61%, 65%, 70%, and 74%, respectively, and has at least 50%, 50%, 60%, 60%, 70%, and 70% complementarity, respectively). As used herein, "substantially complementary" refers to the complementarity between strands such that they are hybridizable under biological conditions. Substantially complementary sequences have 60%, 70%, 80%, 90%, 95%, or even 100% complementarity. Furthermore, techniques for determining whether two strands are capable of hybridizing under biological conditions by examining their nucleotide sequences are well known in the art.
[0110] Clustered interspaced short palindromic repeats (CRISPR) associated (Cas) system: As used herein, CRISPR-Cas9 system refers to nucleic acids and / or proteins involved in the expression or directing the activity of CRISPR effectors, including sequences encoding CRISPR effectors, RNA guides, and other sequences and transcripts from the CRISPR locus. In some embodiments, the CRISPR system is an engineered non-naturally occurring CRISPR system. In some embodiments, the components of the CRISPR system may include nucleic acid(s) (e.g., vectors) encoding one or more components of the system, component(s) in protein form, or a combination thereof.
[0111] CRISPR array: The term "CRISPR array" as used herein refers to a nucleic acid (e.g., DNA) segment that includes a CRISPR repeat and a spacer. In some embodiments, a CRISPR array includes a CRISPR repeat and a spacer that starts with the first nucleotide of the first CRISPR repeat and ends with the last nucleotide of the last (terminal) CRISPR repeat. Typically, each spacer in a CRISPR array is located between two repeats. The term "CRISPR repeat" or "CRISPR direct repeat" or "direct repeat" as used herein refers to multiple short direct repeat sequences that show little or no sequence variation in a CRISPR array.
[0112] CRISPR-associated protein (Cas): The term "CRISPR-associated protein," "CRISPR effector," "effector," or "CRISPR enzyme," as used herein, refers to a protein that performs an enzymatic activity and / or binds to a target site on a nucleic acid specified by an RNA guide. In various embodiments, a CRISPR effector has endonuclease activity, nickase activity, exonuclease activity, transposase activity, and / or excision activity. In other embodiments, a CRISPR effector is nuclease inactive.
[0113] crRNA: The term "CRISPR RNA" or "crRNA" as used herein refers to an RNA molecule that contains a guide sequence that is used by a CRISPR effector to target a specific nucleic acid sequence. Typically, the crRNA contains a sequence that mediates target recognition and a sequence that forms a duplex with the tracrRNA. In some embodiments, the crRNA:tracrRNA duplex binds to a CRISPR effector.
[0114] Duplex: As used herein, "duplex" refers to a double helical structure formed by the interaction of two single-stranded nucleic acids. A duplex is typically formed by paired hydrogen bonds between two single-stranded nucleic acids oriented antiparallel with respect to each other, i.e., "base pairing". Base pairing in a duplex is generally performed by Watson-Crick base pairing, e.g., guanine (G) base pairs with cytosine (C) in DNA and RNA, adenine (A) base pairs with thymine (T) in DNA, and adenine (A) base pairs with uracil (U) in RNA. Conditions under which base pairs can form include physiological or biologically relevant conditions (e.g., intracellular: pH 7.2, 140 mM potassium ions, extracellular pH 7.4, 145 mM sodium ions). Additionally, duplexes are stabilized by stacking interactions between adjacent nucleotides. As used herein, a duplex can be established or maintained by stacking base pairing or interactions. A duplex is formed by two complementary nucleic acid strands that can be substantially complementary or completely complementary. Single-stranded nucleic acids that base pair over several bases are said to "hybridize."
[0115] Ex vivo: As used herein, the term "ex vivo" refers to events that take place in cells or tissues grown outside, rather than within, a multicellular organism.
[0116] Functional equivalent or analog: As used herein, the term "functional equivalent" or "functional analog" in the context of a functional derivative of an amino acid sequence refers to a molecule that retains substantially similar biological activity (either function or structure) as that of the original sequence. A functional derivative or equivalent may be a natural derivative or may be synthetically prepared. Exemplary functional derivatives include amino acid sequences having one or more amino acid substitutions, deletions, or additions, provided that the biological activity of the protein is preserved. The substituting amino acid desirably has similar chemical and physical properties as the amino acid that is substituted. Desirable similar chemical and physical properties include charge similarity, bulkiness, hydrophobicity, hydrophilicity, and the like.
[0117] Half-life: As used herein, the term "half-life" is the time required for a quantity, such as a protein concentration or activity, to fall to half of its value measured at the beginning of the period.
[0118] Hybridize: "Hybridize" refers to the formation of double-stranded molecules between complementary polynucleotide sequences (e.g., genes described herein) or portions thereof under various stringent conditions. (See, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507). "Hybridization" occurs through hydrogen bonds, which can be Watson-Crick hydrogen bonds, Hoogsteen hydrogen bonds, or reversed Hoogsteen hydrogen bonds, between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.
[0119] Improve, increase, or reduce: As used herein, the terms "improve," "increase," or "reduce," or grammatical equivalents, refer to a value relative to a baseline measurement, such as a measurement in the same individual prior to the initiation of a treatment described herein, or a measurement in a control subject (or control subjects) in the absence of a treatment described herein. A "control subject" is a subject suffering from the same form of disease as the subject being treated, and who is about the same age as the subject being treated.
[0120] Indel: As used herein, the term "indel" refers to an insertion or deletion of a base in a nucleic acid sequence, which generally results in a mutation and is a common form of genetic variation.
[0121] Inhibition: As used herein, the terms "inhibit," "inhibit," and "inhibiting" refer to a process or method of decreasing or reducing the activity and / or expression of a protein or gene of interest. Typically, inhibiting a protein or gene refers to reducing the expression or associated activity of the protein or gene by at least 10% or more, e.g., 20%, 30%, 40%, or 50%, 60%, 70%, 80%, 90% or more, or a greater than 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 50-fold, 100-fold or more decrease in expression or associated activity as measured by one or more methods described herein or recognized in the art.
[0122] In vitro: As used herein, the term "in vitro" refers to events that take place not in a multicellular organism but in an artificial environment, e.g., in a test tube or reaction vessel, in cell culture, etc.
[0123] In vivo: As used herein, the term "in vivo" refers to events that occur within a multicellular organism, such as humans and non-human animals. In the context of cell-based systems, the term can be used to refer to events that occur within a living cell (as opposed to, for example, in vitro systems).
[0124] Linker or Spacer: A linker or spacer is a nucleotide or amino acid sequence that physically separates the terminal positions of the gRNA sequence from the NSL sequence to allow Cas binding and function of the gRNA. In some embodiments, the linker is RNA. In some embodiments, the linker is a chemical moiety. In some embodiments, the linker is a peptide. In some embodiments, the linker is DNA. In some embodiments, the linker is a chemical linker, e.g., PEG9 / 18. In some embodiments, the linker is a DNA linker.
[0125] Oligonucleotide: As used herein, the term "oligonucleotide" generally refers to a polynucleotide of single- or double-stranded DNA of about 5 to about 100 nucleotides. Oligonucleotides are also known as "oligomers" or "oligos" and may be isolated from genes or chemically synthesized.
[0126] PAM: The term "PAM" or "protospacer adjacent motif" refers to a short nucleic acid sequence (usually 2-6 base pairs in length) that follows a nucleic acid region that is targeted for cleavage by a CRISPR system, such as CRISPR-Cas9. A PAM may be required for Cas nucleases to cleave and is generally found 3-4 nucleotides downstream from the cleavage site.
[0127] Polypeptide: The term "polypeptide" as used herein refers to a continuous chain of amino acids linked together via peptide bonds. Although the term is used to refer to an amino acid chain of any length, one of skill in the art will understand that the term is not limited to a long chain and can refer to a minimum chain that includes two amino acids linked together via a peptide bond. As known to those of skill in the art, polypeptides can be processed and / or modified. As used herein, the terms "polypeptide" and "peptide" are used interchangeably.
[0128] Prevent: As used herein, the terms "prevent" or "prevention," when used in relation to the occurrence of a disease, disorder, and / or condition, refer to reducing the risk of developing the disease, disorder, and / or condition.
[0129] Protein: The term "protein" as used herein refers to one or more polypeptides that function as separate units. When a single polypeptide is a separate functional unit and does not require permanent or temporary physical association with other polypeptides to form a separate functional unit, the terms "polypeptide" and "protein" may be used interchangeably. When a separate functional unit consists of more than one polypeptide that is physically associated with one another, the term "protein" refers to multiple polypeptides that are physically coupled and function together as a separate unit.
[0130] Reference: A "reference" entity, system, amount, set of conditions, etc., that is compared to a test entity, system, amount, set of conditions, etc., as described herein. For example, in some embodiments, a "reference" antibody is a control antibody that has not been engineered as described herein.
[0131] RNA guide: The term "RNA guide" or "guide RNA" refers to an RNA molecule that facilitates targeting of a protein to a target nucleic acid as described herein. Exemplary "RNA guide" or "guide RNA" includes, but is not limited to, a crRNA or a crRNA in combination with a cognate tracrRNA. The latter may be an independent RNA or may be fused as a single RNA using a linker (sgRNA). In some embodiments, the RNA guide is engineered to include chemical or biochemical modifications, and in some embodiments, the RNA guide may include one or more nucleotides. The term "RNA guide" or "guide RNA" also refers to an NLS-gRNA.
[0132] Single-stranded ligase: As used herein, the term "single-stranded ligase" refers to a ligase that does not require an oligonucleotide splint or template for its ligation activity.
[0133] Splint or Oligonucleotide Splint: The term "splint" or "oligonucleotide splint" refers to a single-stranded RNA or DNA, or other polymer, that can hybridize with at least two, three, or more single-stranded RNA nucleotides. For example, a splint can refer to an oligonucleotide splint.
[0134] Subject: The term "subject," as used herein, means any subject for which diagnosis, prognosis, or therapy is desired. For example, the subject can be a mammal, such as a human or a non-human primate (such as an ape, monkey, orangutan, or chimpanzee), dog, cat, guinea pig, rabbit, rat, mouse, horse, cattle, or cow.
[0135] sgRNA: The term "sgRNA," "single guide RNA," or "guide RNA" refers to a single guide RNA that contains (i) a guide sequence (crRNA sequence) and (ii) a Cas9 nuclease recruiting sequence (tracrRNA).
[0136] Substantial identity: The phrase "substantial identity" is used herein to refer to a comparison between amino acid or nucleic acid sequences. As will be understood by those skilled in the art, two sequences are generally considered to be "substantially identical" if they contain identical residues at corresponding positions. As is well known in the art, amino acid or nucleic acid sequences can be compared using any of a variety of algorithms, including those available in commercially available computer programs such as BLASTN for nucleotide sequences, and BLASTP, gapped BLAST, and PSI-BLAST for amino acid sequences. Exemplary such programs are described in Altschul, et al., Basic local alignment search tool, J. Mol. Biol., 215(3):403-410, 1990; Altschul, et al., Methods in Enzymology; Altschul et al., Nucleic Acids Res. 25:3389-3402, 1997; Baxevanis et al., Bioinformatics: A Practical Guide to the Analysis of Genes and Proteins, Wiley, 1998; and Misener, et al., (eds.), Bioinformatics Methods and Protocols (Methods in Molecular Biology, Vol. 132), Humana Press, 1999. In addition to identifying identical sequences, the above-mentioned programs typically provide an indication of the degree of identity. In some embodiments, two sequences are considered to be substantially identical if at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of the corresponding residues are identical over the relevant stretch of residues. In some embodiments, the relevant stretch is the complete sequence.In some embodiments, the relevant stretch is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more residues.
[0137] Target nucleic acid: The term "target nucleic acid" as used herein refers to any length of nucleotide (oligonucleotide or polynucleotide), deoxyribonucleotide, ribonucleotide, or any of their analogs to which the CRISPR-Cas9 system binds. The target nucleic acid may have a three-dimensional structure that may include coding or non-coding regions, and may include exons, introns, mRNA, tRNA, rRNA, siRNA, shRNA, miRNA, ribozymes, cDNA, plasmids, vectors, exogenous sequences, endogenous sequences. The target nucleic acid may include modified nucleotides, methylated nucleotides, or nucleotide analogs. The target nucleic acid may be interspersed with non-nucleic acid components. The target nucleic acid may be, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers that include purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0138] Therapeutically effective amount: As used herein, the term "therapeutically effective amount" refers to an amount of a therapeutic molecule (e.g., an engineered antibody described herein) that confers a therapeutic effect on a treated subject at a reasonable benefit / risk ratio applicable to any medical treatment. The therapeutic effect may be objective (i.e., measurable by some test or marker) or subjective (i.e., the subject gives an indication of or feels an effect). In particular, a "therapeutically effective amount" refers to an amount of a therapeutic molecule or composition that is effective to treat, ameliorate, or prevent a particular disease or condition, or that is effective to exhibit a detectable therapeutic or prophylactic effect, such as by improving symptoms associated with the disease, preventing or delaying the onset of the disease, and / or reducing the severity or frequency of symptoms of the disease. Therapeutically effective amounts can be administered in a dosing regimen that may include multiple unit doses. For any particular therapeutic molecule, the therapeutically effective amount (and / or the appropriate unit dose within an effective dosing regimen) may vary depending, for example, on the route of administration, or on the combination with other drugs. In addition, the particular therapeutically effective amount (and / or unit dose) for any particular subject may depend on a variety of factors, including the disorder being treated and the severity of the disorder; the activity of the particular agent used; the particular composition used; the age, weight, general health, sex, and diet of the subject; the time of administration, route of administration, and / or excretion or metabolic rate of the particular therapeutic molecule used; the duration of treatment; and similar factors well known in the medical arts.
[0139] tracrRNA: The term "tracrRNA" or "transactivating crRNA" as used herein refers to an RNA that contains a sequence that forms the structure required for a CRISPR-associated protein to bind to a specified target nucleic acid.
[0140] Treatment: As used herein, the term "treatment" (also "treat" or "treating") refers to any administration of a therapeutic molecule (e.g., a CRISPR-Cas therapeutic protein or system described herein) that partially or completely alleviates, ameliorates, relieves, inhibits, delays the onset of, reduces the severity of, and / or reduces the incidence of, one or more symptoms or characteristics of a particular disease, disorder, and / or condition. Such treatment may be of subjects who do not show signs of the relevant disease, disorder, and / or condition and / or who show only early signs of the disease, disorder, and / or condition. Alternatively or additionally, such treatment may be of subjects who show one or more established signs of the relevant disease, disorder, and / or condition.
[0141] The drawings are for illustration purposes only, and not for limitation. [Brief description of the drawings]
[0142] [Figure 1] FIG. 1 is an exemplary schematic diagram of a gRNA conjugated to an NLS sequence. In this particular design, the 3' end of the gRNA is conjugated to the N-terminus of a peptide spacer followed by an NLS sequence derived from SV40. [Diagram 2] 1 is an exemplary graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with a base editor comprising an adenine deaminase fused to the N-terminus of spCas9. The percentage of A to G conversion (y-axis) is plotted for various guide RNAs with and without NLS at various ratios (1:1, 1:3, and 1:9) of mRNA encoding the base editor. "Lipo control" contains mRNA encoding the base editor gRNA (without NLS) in lipofectamine. "Lipo control" was formulated to serve as a transfection control for the LNP group. [Figure 3A]Exemplary schematic diagrams of gRNAs with different modifications. "EM" (end modified) gRNA has 3 nucleotides at both 3' and 5' ends with 2'OMe modifications. "HM1" (heavy modified 1) has gRNA modified with 47% 2'OMe modifications. "HM2" (heavy modified 2) has gRNA modified with 60% 2'OMe modifications. "HM3" (heavy modified 3) has gRNA modified with 88% 2'OME and 2'F modifications. The NLS-gRNA used in Example 2 contains end modifications. [Figure 3B] 1 is an exemplary graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with a base editor comprising an adenine deaminase fused to the N-terminus of spCas9 in mouse. The percentage of A to G conversion (y-axis) is plotted for different guide RNAs with and without an NLS, and with and without different modifications in the gRNA. [Figure 4A] 1 is an exemplary graph showing the results of base editing efficiency achieved with a base editor comprising an adenine deaminase fused to the N-terminus of spCas9 in non-human primates (NHPs). Base editing efficiency in liver (y-axis) is plotted for various guide RNAs with and without NLS and with and without various modifications in the gRNA. [Figure 4B] 1 is a series of exemplary graphs showing toxicology results, in which AST and ALT levels were measured 24 hours after dosing and fold change is shown compared to pre-dosing AST / ALT levels with different gRNA-containing formulations. [Diagram 5] 1 is an exemplary graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with a base editor comprising an adenine deaminase fused to the N-terminus of saCas9 in mouse. The percentage of A to G conversion (y-axis) for both on-target and bystander editing is plotted for different guide RNAs with different purities and modifications. [Figure 6A]Figure 6 shows in vivo correction of GSD1a mutation in liver extracts of transgenic mouse model heterozygous for huG6PC-R83C. Figure 6A is a schematic diagram showing the in vivo workflow. Lipid nanoparticles (LNPs) carrying base editor mRNA and gRNA were administered via IV injection in transgenic mice heterozygous for huG6PC carrying the R83C mutation (huR83C HET). [Figure 6B] Figure 6B shows in vivo correction of the GSD1a mutation in liver extracts of a transgenic mouse model heterozygous for huG6PC-R83C.Figure 6B is a bar graph showing A to G base editing efficiency of the GSD1a R83C mutation using MSP828 comparing on-target to bystander editing. [Figure 6C] 6C is a bar graph showing the correction of the GSD1a R83C mutation in a transgenic mouse model heterozygous for huG6PC carrying the R83C mutation using TadA adenosine deaminase variants MSP605, MSP824, MSP825, MSP680, MSP828, and MSP820. In vitro screening was performed to select the desired base editor for R83C correction. LNP co-formulations of gRNA and representative base editors were administered in vivo (at a subsaturating dose of 1 mpk) in transgenic mice heterozygous for huG6PC-R83C. The base editing efficacy of the variants for R83C correction in the liver of LNP-treated huG6PC-R83C heterozygous transgenic animals is shown in FIG. 6C. Variant MSP828 produced high levels of on-target activity under these conditions. A to G base editing efficiency is shown for on-target and bystander editing. [Figure 7]A schematic diagram showing normal and loss-of-function g6pc function and associated outcomes is shown. GSD-Ia (or GSD1a herein) is an autosomal recessive disorder caused by mutations in the g6pc gene. R83C, located in the active site of the enzyme, is the most prevalent pathogenic mutation identified in Caucasian GSD-Ia patients and is associated with inactivation of G6Pase. Loss of G6Pase function can result in life-threatening hypoglycemia, seizures, and even death. To mitigate hypoglycemia, patients must maintain strict and frequent adherence to glucose replacement with slow glucose-releasing formulas throughout the day and night. A single missed or delayed dose can cause emergent hypoglycemia. Among many complications, liver enlargement, uric acid, lactic acid, and lipid accumulation are common in GSD-Ia patients. [Figure 8] FIG. 8 shows a schematic diagram showing that the base editors described herein generate persistent, predicted nucleotide substitutions within the editing window. The R83C mutation introduces a single G>A conversion in the g6pc gene. Adenine base editors (ABEs) allow for programmable A to G conversion in genomic DNA and can therefore be used to correct this mutation. FIG. 8 shows the utility of the ABEs and base editing described herein. The ABE binds to the target DNA that is complementary to the guide RNA, exposing a stretch of single-stranded DNA. The deaminase converts the target adenine to inosine, and the Cas enzyme nicks the opposite strand, which is then repaired, completing the conversion of the base pair. Direct repair of point mutations has the potential for restoration of gene function. [Figure 9A]A bar graph is provided showing the delineation of the target nucleotide site, as well as the bystander and PAM nucleotides, and the ABE used in immortalized HEK293 cells results in a significant accurate correction rate of R83C. A base editor for A to G conversion in the g6pc gene was optimized for the correction of R83C. In FIG. 9A, the target DNA sequence (CCACCAGTATGGACACTGTCCAAAGAGAAT (SEQ ID NO: 17)) and the underlying amino acid translation (WWYPCQGFLI, SEQ ID NO: 18) for the GSD-Ia R83C mutation are shown. The target edit is indicated by double underlining at position 12. The editing window also includes a potential bystander, indicated by single underlining at position 6, and an edit that may result in a synonymous conversion is indicated at position 10. For screening, HEK293 cell lines were generated to express the g6pc transgene carrying the R83C mutation and transfected with base editor mRNA and gRNA. Allele frequencies were assessed by high-throughput targeted amplicon next-generation sequencing (NGS). [Figure 9B] A depiction of the targeted nucleotide site, as well as bystander and PAM nucleotides, and a bar graph showing that the ABE used in immortalized HEK293 cells results in a remarkable accurate correction rate of R83C. A base editor for A to G conversion in the g6pc gene was optimized for the correction of R83C. Variants 1-5 represent gRNA and base editor RNA combinations engineered for optimized targeted correction. Variant 5 resulted in approximately 60% targeted base editing efficiency for R83C correction with limited bystander editing (Figure 9B). [Figure 10]Photographic images and bar graphs are presented demonstrating that 3-week-old homozygous huR83C (Hom huR83C) mice exhibited expected growth retardation and metabolic defect characteristics of GSD-1a. For the experiment, GSD-Ia mice expressing the human G6PC-R83C transgene instead of mouse G6PC were generated to verify base editing in vivo. The results shown confirmed that mice homozygous for huR83C exhibited postnatal lethality, they were stillborn or died within 24 hours. On glucose replacement therapy, the animals survived to at least 3 weeks of age and revealed characteristic pathological signs of GSD-Ia with reduced body weight, enlarged liver, marked G6Pase inhibition, and abnormalities of serum metabolites compared to littermate controls. A phenotype consistent with clinical and published reports. [Figure 11A] Dot plots of in vivo correction achieved by the base editors (ABEs) described herein are shown. FIG. 11A shows efficient lipid nanoparticle (LNP)-mediated base editing (huG6PC-R83C correction) in the liver of adult and newborn heterozygous huR83C mice. To validate base editing efficiency for R83C correction in vivo, LNP-mediated delivery was first optimized in less vulnerable transgenic mice heterozygous for huR83C. The schematic in FIG. 6A shows the in vivo workflow for these experiments using lipid nanoparticles (LNPs) or LNP co-formulations of base editor mRNA and gRNA administered via IV injection. Given the neonatal lethality of homozygous mice, LNP administration was used via the temporal vein of heterozygous huR83C mice immediately after birth, and activity was compared to that seen in adult heterozygous huR83C mice that received LNP administered via the tail vein. NGS analysis of whole liver extracts revealed base editing efficiencies of approximately 40% in adults and up to approximately 60% in neonates, with a broader range of efficiencies. Bystander editing remained low in adults and neonates. [Figure 11B]Dot plots of in vivo correction achieved by the base editors (ABEs) described herein are shown. The schematic in Figure 6A shows the in vivo workflow for these experiments using lipid nanoparticles (LNPs) or LNP co-formulations of base editor mRNA and gRNA administered via IV injection. Given the neonatal lethality of homozygous mice, LNP administration was used via the temporal vein of heterozygous huR83C mice shortly after birth, and activity was compared to activity seen in adult heterozygous huR83C mice that received LNP administered via the tail vein. NGS analysis of total liver extracts revealed base editing efficiency of approximately 40% in adults and up to about 60% in neonates, with a broader range of efficiency. Bystander editing remained low in adults and neonates. Figure 11B shows that LNP-mediated R83C correction in the liver correlates with survival of neonatal homozygous huR83C mice and littermate heterozygous huR83C mice. Briefly, newborn mice homozygous for huR83C were treated with LNPs containing guide RNA and mRNA encoding ABE. Treated mice grew normally to 3 weeks of age without hypoglycemia-induced seizures in the absence of glucose therapy. Treated homozygous huR83C mice showed up to 60% editing efficiency (i.e., 60% R83C correction) in total liver extracts, consistent with littermate controls that were heterozygous for huR83C. [Figure 12A]12A shows bar graphs and immunohistochemical staining images showing base editing as described herein in mice homozygous for huG6PC-R83C restores near-normal metabolic function to reverse GSD-Ia pathology. As shown by the darkest bars in the graph in FIG. 12A, at 3 weeks, treated homozygous huR83C mice were verified to exhibit proper metabolic function with restoration of near-normal serum metabolite markers including glucose, triglycerides, cholesterol, lactate, and uric acid. Furthermore, biochemical assays of G6PC activity (assessed biochemically and via lead-phosphate staining) in LNP-treated homozygous huR83C mice were consistent with those of littermate controls. Hepatomegaly, another clinical presentation of GSD-Ia, is primarily caused by excess glycogen and lipid deposition. [Figure 12B] 12B shows bar graphs and immunohistochemical staining images showing base editing described herein in mice homozygous for huG6PC-R83C restores near-normal metabolic function to reverse GSD-Ia pathology. Immunohistochemical analysis revealed normal hepatocyte size and lipid deposition in LNP-treated mice (FIG. 12B). The results demonstrate the potential of base editing to correct the metabolic defects associated with the R83C mutation and GSD-Ia. [Figure 13] 13A-13C depict bar graphs showing that single LNP dose administration maintained normoglycemia in homozygous huG6PC-R83C mice during a 24-hour fasting challenge via base editing as described herein. [Figure 14]Kaplan-Meier survival curves were generated to show the estimated survival of newborn transgenic mice homozygous for huG6PC-R83C, either after ABE mRNA-mediated base editing or untreated. Newborn mice were transfected with the following primers: universal forward primer (5'-ACCTACTGATGATGCACCTTTGATCAATAGAT-3' (SEQ ID NO: 61)), mouse specific reverse primer (5'-CATCACCCCTCGGGATGGTTCTT-3' (SEQ ID NO: 62)), human specific reverse primer 1 (5'-CAGCCCAGAATCCCAACCACAAAAT-3' (SEQ ID NO: 63), and human specific reverse primer 2 (5'-AGACCAGCTCGACTTGGGATGG-3' (SEQ ID NO: 64)). ABE-treated mice were genotyped via PCR analysis on genomic tail DNA using the huG6PC-R83C PCR kit. Survival was noted for transgenic mice homozygous for huG6PC-R83C. Untreated mice were either stillborn (n=6) or died at 8 hours (n=6) and 24 hours (n=1). Administration of 15% glucose injection extended survival to 32 hours (n=5), 48 hours (n=2), and 56 hours (n=2). All ABE-treated mice homozygous for huG6PC-R83C survived until the end of the study at 3 weeks. [Figure 15A] FIG. 1 is a schematic diagram of gRNA fluorescently tagged with Cy5 dye. [Figure 15B] FIG. 1 is a schematic diagram of a gRNA conjugated to an NLS fluorescently tagged with Cy5 dye. [Figure 15C] Nuclear staining with Nuc Blue is shown. [Figure 15D] Nuclear staining with Cy5 and localization of ALAS1 / sg23 gRNA are shown. [Figure 15E] 1 shows enhanced nuclear localization of NLS-gRNA. [Figure 15F] Graph showing nuclear localization of gRNA and NLS-gRNA in PXB cells. When conjugated to NLS, gRNA was efficiently localized to the nucleus. ABE8.8. [Figure 16]A model of the NLS conjugate bound at its 3' end to a saCas9 effector. [Figure 17A] The sequences of an exemplary 5% terminally modified gRNA and an exemplary 25% heavily modified saHM03 gRNA are provided. [Figure 17B] 13A-13C are graphs showing the results of A to G base editing efficiency of exemplary NLS-conjugated gRNAs for terminally modified gRNAs and heavily modified saHM03 gRNAs. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0143] Methods, compositions, and kits are provided herein for enhancing the efficacy of gRNA for use in CRISPR-Cas systems. In some aspects, the present invention provides a method for producing gRNA conjugated to an NLS sequence (NLS-gRNA) with increased efficacy for use in CRISPR-Cas systems, which increases the frequency of successful editing events. The NLS-gRNA of the present invention can provide better transport of gRNA to the nucleus, protecting it from cytoplasmic RNases and increasing the higher local concentration of gRNA for the formation of RNPs. The NLS-gRNA of the present invention has significantly higher efficacy compared to the corresponding gRNA that does not contain an NLS sequence, and also shows higher efficacy compared to highly modified gRNAs.
[0144] A gRNA conjugated to an NLS sequence (NLS-gRNA) has a number of potential advantages, including, for example, increased efficacy. For example, the NLS-gRNA of the present invention provides significantly higher base editing efficiency relative to its corresponding gRNA that does not contain an NLS sequence. Furthermore, NLS-gRNAs with terminal modifications (e.g., 2'OMe modifications at the 3' and / or 5' ends) provide higher efficacy compared to highly modified gRNAs (e.g., more than 40%, more than 60%, or more than 88% modified).
[0145] Various aspects of the invention are described in detail in the following sections. The use of the sections is not intended to limit the invention. Each section may be applicable to any aspect of the invention. In this application, the use of "or" means "and / or" unless otherwise specified.
[0146] Guide RNA (gRNA) As used herein, guide RNA (gRNA) also refers to guide RNA conjugated to an NLS sequence (NLS-gRNA), unless otherwise noted. gRNA comprises a polynucleotide sequence that is complementary to a target sequence. gRNA hybridizes to the target nucleic acid sequence and directs sequence-specific binding of the CRISPR complex to the target nucleic acid. In some embodiments, the RNA guide has 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% complementarity to the target nucleic acid sequence.
[0147] In some embodiments, the gRNA is between about 50 nucleotides and 250 nucleotides. In some embodiments, the gRNA is between about 50 nucleotides and 500 nucleotides. In some embodiments, the gRNA is between about 50 nucleotides and 1,000 nucleotides. In some embodiments, the gRNA is between about 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, or 250 nucleotides in length. In some embodiments, the gRNA is between about 50 and 75 nucleotides in length. In some embodiments, the gRNA is about 75-100 nucleotides in length. In some embodiments, the gRNA is about 100-125 nucleotides in length. In some embodiments, the gRNA is about 125-150 nucleotides in length. In some embodiments, the gRNA is about 150-175 nucleotides in length. In some embodiments, the gRNA is about 175-200 nucleotides in length. In some embodiments, the gRNA is about 200-225 nucleotides in length. In some embodiments, the gRNA is about 225-250 nucleotides in length.
[0148] In some embodiments, the gRNA comprises a ligated crRNA and a tracrRNA. Various crRNA and tracrRNA sequences are known in the art, such as those associated with several type II CRISPR-Cas9 systems (e.g., WO2013 / 176772), Cpf1, SaCas9, Cas12, among others.
[0149] The gRNA can be designed to target any target sequence. Optimal alignment is determined using any algorithm for aligning sequences, including the Needleman-Wunsch algorithm, the Smith-Waterman algorithm, the Burrows-Wheeler algorithm, ClustlW, ClustlX, BLAST, Novoalign, SOAP, Maq, and ELAND.
[0150] In some embodiments, the gRNA is designed to target a unique target sequence in the genome of a cell. In some embodiments, the gRNA is designed to lack a PAM sequence. In some embodiments, the gRNA sequence is designed to have an optimal secondary structure using folding algorithms, including mFold or Geneious. In some embodiments, the expression of the gRNA can be under an inducible promoter, for example, hormone-inducible, tetracycline or doxycycline-inducible, arabinose-inducible, or light-inducible.
[0151] In some embodiments, the gRNA sequence is a "dead crRNA," "dead guide," or "dead guide sequence" that can form a complex with CRISPR-associated proteins and bind to a specific target without any substantial nuclease activity.
[0152] In some embodiments, the gRNA is chemically modified in the sugar phosphate backbone or bases. In some embodiments, the gRNA has one or more of the following modifications 2'O-methyl, 2'-F or locked nucleic acid to improve nuclease resistance or base pairing. In some embodiments, the gRNA may contain modified bases such as 2-thiouridine or N6-methyladenosine.
[0153] In some embodiments, the gRNA is conjugated to other oligonucleotides, peptides, proteins, tags, dyes, or polyethylene glycol.
[0154] In some embodiments, the gRNA comprises an aptamer or riboswitch sequence that binds to a specific target molecule due to its three-dimensional structure.
[0155] In some embodiments, the gRNA has 2, 3, 4, or 5 hairpins.
[0156] In some embodiments, the gRNA comprises a transcription termination sequence comprising a poly-T sequence comprising 6 nucleotides.
[0157] Conjugation of gRNA to nuclear localization (NLS) sequence In one aspect, the present invention provides a gRNA conjugated to an NLS sequence via the 3' end of the gRNA. In one aspect, the present invention provides a gRNA conjugated to an NLS sequence via the 5' end of the gRNA. In one aspect, the present invention provides a gRNA conjugated to an NLS sequence via an internal site of the gRNA.
[0158] In embodiments, the gRNA is conjugated to the NLS via a linker, which in embodiments comprises a chemical moiety (e.g., L) and / or a peptide moiety (e.g., a peptide spacer).
[0159] In embodiments, the gRNA is directly conjugated to the NLS via a chemical moiety (e.g., L). In embodiments, the chemical moiety (e.g., L) is non-peptidic. In embodiments, the chemical moiety (e.g., L) is covalently attached to both the gRNA and the NLS.
[0160] In embodiments, the gRNA is conjugated to the NLS via a peptide moiety (e.g., a peptide spacer). In embodiments, the peptide moiety (e.g., a peptide spacer) is covalently linked to both the gRNA and the NLS.
[0161] In embodiments, the gRNA is conjugated to the NLS via a linker that includes both a chemical moiety (e.g., L) and a peptide moiety (e.g., a peptide spacer). In embodiments, such a conjugate can have a structure according to formula (I), where the chemical moiety L (e.g., a non-peptidic chemical moiety) is covalently attached to the gRNA and the peptide spacer, and the peptide spacer is covalently attached to the NLS. [ka]
[0162] In some embodiments, the N-terminus of the NLS sequence is conjugated to the 3'-terminus of the gRNA via a linker that includes both a chemical moiety (e.g., L) and a peptide moiety (e.g., a peptide spacer). In some embodiments, the C-terminus of the NLS sequence is conjugated to the 5'-terminus of the gRNA via a linker that includes both a chemical moiety (e.g., L) and a peptide moiety (e.g., a peptide spacer). In some embodiments, an internal amino acid in the NLS sequence is conjugated to the 3'-terminus of the gRNA via a linker that includes both a chemical moiety (e.g., L) and a peptide moiety (e.g., a peptide spacer). In some embodiments, an internal amino acid in the NLS sequence is conjugated to the 5'-terminus of the gRNA via a linker that includes both a chemical moiety (e.g., L) and a peptide moiety (e.g., a peptide spacer). In some embodiments, an internal amino acid in the NLS sequence is conjugated to an internal nucleotide of the gRNA via a linker that includes both a chemical moiety (e.g., L) and a peptide moiety (e.g., a peptide spacer).
[0163] In embodiments, the gRNA is conjugated to the NLS via a peptide spacer or a chemical moiety (e.g., L) covalently attached to the C-terminus of the NLS amino acid sequence.
[0164] In embodiments, the gRNA is conjugated to the NLS via a peptide spacer or a chemical moiety (e.g., L) covalently attached to the N-terminus of the NLS amino acid sequence.
[0165] In embodiments, the gRNA is conjugated to a peptide spacer or NLS via a chemical moiety (e.g., L) covalently attached to the 3' end of the gRNA.
[0166] In embodiments, the gRNA is conjugated to a peptide spacer or NLS via a chemical moiety (e.g., L) covalently attached to the 5' end of the gRNA.
[0167] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to a thiol-containing residue (eg, a cysteine residue) of the NLS.
[0168] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to a selenium-containing residue (eg, a selenocysteine residue) of the NLS.
[0169] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to an amino-containing residue (eg, a lysine residue) of the NLS.
[0170] In embodiments, the chemical moiety (eg, L) is covalently attached to the peptide spacer or to a phenol-containing residue (eg, a tyrosine residue) of the NLS.
[0171] In embodiments, the amino acid residue used to form the linker (eg, a thiol, selenium, amino, or phenol-containing residue as described herein) comprises a chemical modification.
[0172] In some embodiments, gRNA is conjugated to NLS via reductive amination. In some embodiments, gRNA is conjugated to NLS native chemical ligation. gRNA is conjugated to NLS via thiol-ene click.
[0173] Exemplary chemicals useful for preparing linkers are described herein. The chemical moieties described herein can be represented by the substructure L 1and / or L 2 and L 1 and L 2 are each independently 1~12 Alkylene or C 2~12 In embodiments, L is an optionally substituted group that is heteroalkylene. 1 and L 2 contains an oxo (=O) substituent (eg, one or two oxo substituents).
[0174] Maleimide-thiol / maleimide-selenol adducts In embodiments, the chemical moiety (eg, L) comprises a maleimide-thiol adduct.
[0175] In embodiments, the gRNA is conjugated to the NLS using an addition reaction between a maleimide group and a thiol group or a thiol-ene click reaction.
[0176] In embodiments, the maleimide-thiol adduct-containing moiety is formed from a gRNA that includes a thiol group, and an NLS (or peptide spacer) that includes a maleimide group. In embodiments, the maleimide-thiol adduct-containing moiety is formed from a gRNA that includes a maleimide group, and an NLS (or peptide spacer) that includes a thiol group.
[0177] In embodiments, the chemical moiety (e.g., L) comprises a maleimide-selenol adduct. In embodiments, the gRNA is conjugated to the NLS using an addition reaction between a maleimide group and a selenol group. In embodiments, the maleimide-selenol adduct-containing moiety is formed from a gRNA comprising a selenol group and an NLS (or a peptide spacer) comprising a maleimide group. In embodiments, the maleimide-selenol adduct-containing moiety is formed from a gRNA comprising a maleimide group and an NLS (or a peptide spacer) comprising a selenol group.
[0178] In embodiments, the chemical moiety (e.g., L) is [ka] wherein Y is S or Se.
[0179] In embodiments, Y is S. In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The moiety is formed from a gRNA that includes a thiol group, and an NLS (or peptide spacer) that includes a maleimide group. In embodiments, the maleimide-thiol adduct-containing moiety is formed from a gRNA that includes a maleimide group, and an NLS (or peptide spacer) that includes a thiol group.
[0180] In embodiments, Y is Se. In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The moiety is formed from a gRNA that includes a selenol group and an NLS (or peptide spacer) that includes a maleimide group. In embodiments, the maleimide-selenol adduct-containing moiety is formed from a gRNA that includes a maleimide group and an NLS (or peptide spacer) that includes a selenol group.
[0181] In an embodiment, the chemical moiety L has the following structure (A): [ka]
[0182] In embodiments, Y is S. In embodiments, * represents a covalent bond to the gRNA. In embodiments, ** represents a covalent bond to the peptide spacer or NLS. In embodiments, ** represents a covalent bond to the peptide spacer.
[0183] Thioethers / Selenoethers In embodiments, the chemical moiety (eg, L) comprises a thioether group.
[0184] In embodiments, the gRNA is conjugated to the NLS using a conjugation reaction between an iodoacetamide group and a thiol group.
[0185] In embodiments, the thioether-containing moiety is formed from a gRNA that includes a thiol group and an NLS (or peptide spacer) that includes an iodoacetamide group. In embodiments, the thioether-containing moiety is formed from a gRNA that includes an iodoacetamide group and an NLS (or peptide spacer) that includes a thiol group.
[0186] In embodiments, the chemical moiety (e.g., L) comprises a selenoether moiety. In embodiments, the gRNA is conjugated to the NLS using a conjugation reaction between an iodoacetamide group and a selenol group. In embodiments, the selenoether-containing moiety is formed from a gRNA comprising a selenol group and an NLS (or peptide spacer) comprising an iodoacetamide group. In embodiments, the selenoether-containing moiety is formed from a gRNA comprising an iodoacetamide group and an NLS (or peptide spacer) comprising a selenol group.
[0187] In embodiments, the chemical moiety (e.g., L) is [ka] wherein Y is S or Se.
[0188] In embodiments, Y is S. In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that contains a thiol group and an NLS (or peptide spacer) that contains an iodoacetamide group. [ka] The moiety is formed from a gRNA that contains an iodoacetamide group, and an NLS (or peptide spacer) that contains a thiol group.
[0189] In embodiments, Y is Se. In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] In an embodiment, the portion is formed from a gRNA that includes a selenol group and an NLS (or peptide spacer) that includes an iodoacetamide group. [ka] The moiety is formed from a gRNA that contains an iodoacetamide group, and an NLS (or peptide spacer) that contains a selenol group.
[0190] Disulfides (thiol-disulfide exchange chemistry) In embodiments, the chemical moiety (e.g., L) comprises a disulfide group. In embodiments, the gRNA is conjugated to the NLS using a thiol-disulfide exchange reaction between a disulfide-containing group and a thiol group.
[0191] In embodiments, the disulfide-containing moiety is formed from a gRNA that includes a thiol group and an NLS (or peptide spacer) that includes a disulfide group. In embodiments, the disulfide-containing moiety is formed from a gRNA that includes a disulfide group and an NLS (or peptide spacer) that includes a thiol group.
[0192] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that contains a thiol group, and an NLS (or peptide spacer) that contains a disulfide group. [ka] The portion is formed from a gRNA that contains a disulfide group, and an NLS (or peptide spacer) that contains a thiol group.
[0193] Oxadiazole Thioether In embodiments, the chemical moiety (e.g., L) comprises an oxadiazole thioether group. In embodiments, the gRNA is conjugated to the NLS using the reaction between a thiol group and a sulfonyloxadiazole group.
[0194] In embodiments, the oxadiazole thioether-containing moiety is formed from a gRNA that includes a sulfonyloxadiazole group and an NLS (or peptide spacer) that includes a thiol group. In embodiments, the oxadiazole thioether-containing moiety is formed from a gRNA that includes a thiol group and an NLS (or peptide spacer) that includes a sulfonyloxadiazole group.
[0195] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] In an embodiment, the portion is formed from a gRNA that includes a sulfonyloxadiazole group and an NLS (or peptide spacer) that includes a thiol group. [ka] The moiety is formed from a gRNA that contains a thiol group, and an NLS (or peptide spacer) that contains a sulfonyloxadiazole group.
[0196] Urea / Thiourea / Dithiocarbamate (Iso(thio)cyanate Chemistry) In embodiments, the chemical moiety (e.g., L) comprises a urea group. In embodiments, the gRNA is conjugated to the NLS using a reaction between an amino (e.g., primary amine) group and an isocyanate group.
[0197] In embodiments, the urea-containing moiety is formed from a gRNA that includes an amino (e.g., primary amine) group and an NLS (or peptide spacer) that includes an isocyanate group. In embodiments, the urea-containing moiety is formed from a gRNA that includes an isocyanate group and an NLS (or peptide spacer) that includes an amino (e.g., primary amine) group.
[0198] In embodiments, the chemical moiety (e.g., L) comprises a thiourea group. In embodiments, the gRNA is conjugated to the NLS using a reaction between an amino (e.g., primary amine) group and an isothiocyanate group. In embodiments, the thiourea-containing moiety is formed from a gRNA that comprises an amino (e.g., primary amine) group and an NLS (or peptide spacer) that comprises an isothiocyanate group. In embodiments, the thiourea-containing moiety is formed from a gRNA that comprises an isothiocyanate group and an NLS (or peptide spacer) that comprises an amino (e.g., primary amine) group.
[0199] In embodiments, the chemical moiety (e.g., L) is [ka] wherein X is S or O.
[0200] In embodiments, X is O. In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that contains an amino (e.g., primary amine) group, and an NLS (or peptide spacer) that contains an isocyanate group. [ka] The moiety is formed from a gRNA that contains an isocyanate group, and an NLS (or peptide spacer) that contains an amino (e.g., primary amine) group.
[0201] In embodiments, X is S. In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that includes an amino (e.g., primary amine) group, and an NLS (or peptide spacer) that includes an isothiocyanate group. [ka] The moiety is formed from a gRNA that contains an isothiocyanate group, and an NLS (or peptide spacer) that contains an amino (e.g., primary amine) group.
[0202] In embodiments, the chemical moiety (e.g., L) comprises a dithiocarbamate group. In embodiments, the gRNA is conjugated to the NLS using a reaction between a thiol group and an isothiocyanate group.
[0203] In embodiments, the dithiocarbamate-containing moiety is formed from a gRNA that includes a thiol group and an NLS (or peptide spacer) that includes an isothiocyanate group. In embodiments, the dithiocarbamate-containing moiety is formed from a gRNA that includes an isothiocyanate group and an NLS (or peptide spacer) that includes a thiol group.
[0204] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that contains a thiol group and an NLS (or peptide spacer) that contains an isothiocyanate group. [ka] The moiety is formed from a gRNA that contains an isothiocyanate group, and an NLS (or peptide spacer) that contains a thiol group.
[0205] Diazenylphenol In embodiments, the chemical moiety (e.g., L) comprises a diazenylphenol group. In embodiments, the gRNA is conjugated to the NLS using the reaction between a phenol group and a diazonium group.
[0206] In embodiments, the diazenylphenol-containing moiety is formed from a gRNA that includes a phenol group and an NLS (or peptide spacer) that includes a diazonium group. In embodiments, the diazenylphenol-containing moiety is formed from a gRNA that includes a diazonium group and an NLS (or peptide spacer) that includes a phenol group.
[0207] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The moiety is formed from a gRNA that contains a phenol group, and an NLS (or peptide spacer) that contains a diazonium group. [ka] The moiety is formed from a gRNA that contains a diazonium group, and an NLS (or peptide spacer) that contains a phenol group.
[0208] Triazolidinedionylphenol In embodiments, the chemical moiety (e.g., L) comprises a triazolidinedione phenol group. In embodiments, the gRNA is conjugated to the NLS using the reaction between the phenol group and a cyclic diazodicarboxamide group.
[0209] In an embodiment, the triazolidinedione phenol-containing moiety is formed from a gRNA that includes a phenol group and an NLS (or peptide spacer) that includes a cyclic diazodicarboxamide group. In an embodiment, the triazolidinedione phenol-containing moiety is formed from a gRNA that includes a cyclic diazodicarboxamide group and an NLS (or peptide spacer) that includes a phenol group.
[0210] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that contains a phenol group, and an NLS (or peptide spacer) that contains a cyclic diazodicarboxamide group. [ka] The moiety is formed from a gRNA that contains a cyclic diazodicarboxamide group, and an NLS (or peptide spacer) that contains a phenol group.
[0211] Triazole (click chemistry) In embodiments, the chemical moiety (e.g., L) comprises a triazole group. In embodiments, the gRNA is conjugated to the NLS using 1,3-dipolar cycloaddition between an alkyne group and an azide group.
[0212] In embodiments, the triazole-containing moiety is formed from a gRNA that includes an alkyne group and an NLS (or a peptide spacer) that includes an azide group. In embodiments, the triazole-containing moiety is formed from a gRNA that includes an azide group and an NLS (or a peptide spacer) that includes an alkyne group. In embodiments, the 1,3-dipolar cycloaddition is a copper-catalyzed cycloaddition. In embodiments, the 1,3-dipolar cycloaddition is a strain-promoted cycloaddition.
[0213] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] In an embodiment, the portion is formed from a gRNA that contains an alkyne group and an NLS (or peptide spacer) that contains an azide group. [ka] The moiety is formed from a gRNA that contains an azide group, and an NLS (or peptide spacer) that contains an alkyne group.
[0214] In embodiments, the chemical moiety (e.g., L) is [ka] wherein each of ring A and ring B is an optionally substituted aryl group. In embodiments, ring A is present. In embodiments, ring A is absent. In embodiments, ring B is present. In embodiments, ring B is absent. In embodiments, both ring A and ring B are present. In embodiments, both ring A and ring B are absent. In embodiments, [ka] In an embodiment, the portion is formed from a gRNA that contains an alkyne group and an NLS (or peptide spacer) that contains an azide group. [ka] The moiety is formed from a gRNA that contains an azide group, and an NLS (or peptide spacer) that contains an alkyne group.
[0215] In embodiments, the chemical moiety (e.g., L) is [ka] wherein each of ring A and ring B is an optionally substituted aryl group. In embodiments, ring A is present. In embodiments, ring A is absent. In embodiments, ring B is present. In embodiments, ring B is absent. In embodiments, both ring A and ring B are present. In embodiments, both ring A and ring B are absent. In embodiments, [ka] In an embodiment, the portion is formed from a gRNA that contains an alkyne group and an NLS (or peptide spacer) that contains an azide group. [ka] The moiety is formed from a gRNA that contains an azide group, and an NLS (or peptide spacer) that contains an alkyne group.
[0216] Diazanorcaradiene In embodiments, the chemical moiety (e.g., L) comprises a diazanorcaradiene group. In embodiments, the gRNA is conjugated to the NLS using a Diels-Alder reaction between a cyclopropene group and a tetrazine group.
[0217] In embodiments, the diazanorcaradiene-containing moiety is formed from a gRNA that includes a cyclopropene group, and an NLS (or peptide spacer) that includes a tetrazine group. In embodiments, the diazanorcaradiene-containing moiety is formed from a gRNA that includes a tetrazine group, and an NLS (or peptide spacer) that includes a cyclopropene group.
[0218] In embodiments, the chemical moiety (e.g., L) is [ka] wherein R is C 1~6 In an embodiment, [ka] The portion is formed from a gRNA that includes a cyclopropene group, and an NLS (or peptide spacer) that includes a tetrazine group. [ka] The portion is formed from a gRNA that contains a tetrazine group, and an NLS (or peptide spacer) that contains a cyclopropene group.
[0219] Amides / Sulfonamides In embodiments, the chemical moiety (e.g., L) comprises an amide group. In embodiments, the gRNA is conjugated to the NLS using a conjugation reaction between a carboxyl group and an amino group (e.g., a primary amine).
[0220] In embodiments, the amide-containing moiety is formed from a gRNA that includes a carboxyl group and an NLS (or peptide spacer) that includes an amino (e.g., primary amine) group. In embodiments, the amide-containing moiety is formed from a gRNA that includes an amino (e.g., primary amine) group and an NLS (or peptide spacer) that includes a carboxyl group. In embodiments, the carboxyl group is an activated carboxyl group. In embodiments, the carboxyl group is activated by a carbodiimide, such as 1-ethyl-3-(3-dimethyl-aminopropyl)carbodiimide (EDC) or dicyclohexylcarbodiimide (DCC). In embodiments, the carboxyl group is activated by an N-hydroxysuccinimide (NHS) derivative (e.g., sulfo-NHS).
[0221] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that contains a carboxyl group and an NLS (or peptide spacer) that contains an amino (e.g., primary amine) group. [ka] The portion is formed from a gRNA that contains an amino group (e.g., a primary amine) and an NLS (or peptide spacer) that contains a carboxyl group.
[0222] In embodiments, the chemical moiety (e.g., L) comprises a sulfonamide group. In embodiments, the gRNA is conjugated to the NLS using a conjugation reaction between a sulfonyl group and an amino (e.g., primary amine) group. In embodiments, the sulfonamide-containing moiety is formed from a gRNA comprising a sulfonyl group and an NLS (or peptide spacer) comprising an amino (e.g., primary amine) group. In embodiments, the amide-containing moiety is formed from a gRNA comprising an amino (e.g., primary amine) group and an NLS (or peptide spacer) comprising a sulfonyl group.
[0223] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that includes a sulfonyl group and an NLS (or peptide spacer) that includes an amino (e.g., primary amine) group. [ka] The portion is formed from a gRNA that contains an amino (e.g., primary amine) group, and an NLS (or peptide spacer) that contains a sulfonyl group.
[0224] Amines (Glutaraldehyde Chemistry) In embodiments, the chemical moiety (e.g., L) comprises an amino group. In embodiments, the gRNA is conjugated to the NLS using a conjugation reaction between an amino group (e.g., a primary amine) and an aldehyde group, followed by a reduction reaction to form the amine-containing moiety.
[0225] In embodiments, the amine-containing moiety is formed from a gRNA that includes an amino group (e.g., a primary amine) and an NLS (or peptide spacer) that includes an aldehyde group. In embodiments, the amine-containing moiety is formed from a gRNA that includes an aldehyde group and an NLS (or peptide spacer) that includes an amino group (e.g., a primary amine).
[0226] In embodiments, the amine-containing moiety is formed from a bifunctional cross-linking reagent (e.g., a dialdehyde such as glutaraldehyde). In embodiments, the amine-containing moiety is formed from a gRNA that includes an amino group (e.g., a primary amine), an NLS (or peptide spacer) that includes an amino group (e.g., a primary amine), and a dialdehyde (e.g., glutaraldehyde). In embodiments, the amine-containing moiety is formed from a gRNA that includes an aldehyde group, an NLS (or peptide spacer) that includes an aldehyde group, and a diaminoalkane.
[0227] In embodiments, the chemical moiety (e.g., L) is [ka] wherein L 3 is C 1~6 In an embodiment, [ka] In some embodiments, the moiety is formed from a gRNA that includes an amino group (e.g., a primary amine), an NLS (or peptide spacer) that includes an amino group (e.g., a primary amine), and a dialdehyde (e.g., glutaraldehyde). [ka] The moiety is formed from a gRNA containing an aldehyde group, an NLS (or peptide spacer) containing an aldehyde group, and a diaminoalkane.
[0228] In embodiments, the chemical moiety (e.g., L) comprises an amino group. In embodiments, the gRNA is conjugated to the NLS using a conjugation reaction between an amino (e.g., primary amine) group and a tresyl (2,2,2-trifluoroethanesulfonyl) group. In embodiments, the amine moiety is formed from a gRNA that comprises an amino (e.g., primary amine) group and an NLS (or peptide spacer) that comprises a tresyl (2,2,2-trifluoroethanesulfonyl) group. In embodiments, the amine-containing moiety is formed from a gRNA that comprises a tresyl (2,2,2-trifluoroethanesulfonyl) group and an NLS (or peptide spacer) that comprises an amino (e.g., primary amine) group.
[0229] In embodiments, the chemical moiety (e.g., L) is [ka] In an embodiment, [ka] The portion is formed from a gRNA that includes an amino (e.g., primary amine) group, and an NLS (or peptide spacer) that includes a tresyl (2,2,2-trifluoroethanesulfonyl) group. [ka] The portion is formed from a gRNA that contains a tresyl (2,2,2-trifluoroethanesulfonyl) group, and an NLS (or peptide spacer) that contains an amino (e.g., primary amine) group.
[0230] In some embodiments, the NLS-gRNA comprises a crRNA. In some embodiments, the NLS-gRNA comprises a tracrRNA. In some embodiments, the NLS-gRNA comprises a crRNA and an NLS-gRNA.
[0231] In some embodiments, linear guide RNA is synthesized first.In this approach, two or more separate RNAs are ligated together.In some embodiments, the first RNA comprises a transactivating RNA (tracrRNA), and the second RNA comprises a clustered regularly interspaced short palindromic repeats (CRISPR) RNA (crRNA).
[0232] In some embodiments, the RNA containing the tracrRNA sequence is synthesized such that a portion of the tracrRNA contains a phosphate at the 5' end. This approach allows for two forms of ligation, both of which are found within the stem-loop region. The first form of ligation occurs within the terminal loop of the hairpin, which is the natural site for T4 RNA ligase 1. The second form of ligation occurs within the duplex, which is the natural product of T4 RNA ligase 2 and DNA ligase. One of the advantages of this form of ligation is that the fragment impurities can be easily removed due to the significant difference in elution times between the fusion gRNA and the fragment impurities.
[0233] Chemically modified NLS-gRNA In some embodiments, the first end of the guide RNA and / or the second end of the guide RNA comprises a chemical modification to its backbone or one or more of its bases. For example, chemically modified RNA can comprise chemical synthesis and can be used to introduce highly modified monomers that contain modified sugars, bases, backbones, or functional groups that are not similar to natural nucleotides.
[0234] Thus, in some embodiments, the first end of the guide RNA and / or the second end of the guide RNA comprises modified bases. In some embodiments, the modified RNA includes one or more of the following 2'-O-methoxy-ethyl bases (2'-MOE), such as 2-methoxyethoxy A, 2-methoxyethoxy MeC, 2-methoxyethoxy G, 2-methoxyethoxy T. Other modified bases include, for example, 2'-O-methyl RNA bases, and fluoro bases. Various fluoro bases are known, including, for example, fluoro C, fluoro U, fluoro A, fluoro G bases. Various 2'-O-methyl modifications can also be used with the methods described herein. For example, the following RNAs containing one or more of the following 2'O methyl modifications can be used with the described methods: 2'-OMe-5-methyl-rC, 2'-OMe-rT, 2'-OMe-rI, 2'-OMe-2-amino-rA, aminolinker-C6-rC, aminolinker-C6-rU, 2'-OMe-5-Br-rU, 2'-OMe-5-I-rU, 2-OMe-7-deaza-rG.
[0235] In some embodiments, the first end of the guide RNA and / or the second end of the guide RNA comprises one or more of the following modifications: phosphorothioate, 2'O-methyl, 2'fluoro (2'F), DNA.
[0236] In some embodiments, the first end of the guide RNA and / or the second end of the guide RNA comprises 2'OMe modifications at the 3' and 5' ends.
[0237] In some embodiments, the first end of the guide RNA and / or the second end of the guide RNA comprises one or more of the following modifications: 2'-O-2-Methoxyethyl (MOE), locked nucleic acid, bridged nucleic acid, unlocked nucleic acid, peptide nucleic acid, morpholino nucleic acid.
[0238] In some embodiments, the first end of the guide RNA and / or the second end of the guide RNA comprises one or more of the following base modifications: 2,6-diaminopurine, 2-aminopurine, pseudouracil, N1-methyl-pseudouracil, 5' methylcytosine, 2' pyrimidinone (zebularine), thymine.
[0239] Other modified bases include, for example, 2-aminopurine, 5-bromo-dU, deoxyuridine, 2,6-diaminopurine (2-amino-dA), dideoxy-C, deoxyinosine, hydroxymethyl-dC, inverted-dT, iso-dG, iso-dC, inverted-dideoxy-T, 5-methyl-dC, 5-methyl-dC, 5-nitroindole, Super T®, 2'-Fr(C,U), 2'-NH2-r(C,U), 2,2'-anhydro-U, 3'-desoxy-r(A,C,G,U), 3'-O-methyl-r(A,C,G,U), rT, rI, 5-methyl-rC, 2-amino-rA, rspacer (abasic), 7-deaza-rG, 7-deaza-rA, 8-oxo-rG, 5-halogenated-rU, N-alkylated-rN.
[0240] Other chemically modified RNAs can be used herein. For example, the first end of the guide RNA and / or the second end of the guide RNA can include modified bases, such as, for example, 5',Int,3' azide (NHS ester), 5' hexynyl, 5',Int,3' 5-octadiynyl dU, 5',Int biotin (azide), 5',Int6-FAM (azide), and 5',Int5-TAMRA (azide). Other examples of RNA nucleotide modifications that can be used with the methods described herein include phosphorylation modifications, such as, for example, 5' phosphorylation and 3' phosphorylation. RNAs can also have one or more of the following modifications: amino modification, biotinylation, thiol modification, alkyne modifier, adenylation, azide (NHS ester), cholesterol-TEG, and digoxigenin (NHS ester).
[0241] Nucleic acid base editor Nucleobase editors that edit, modify, or alter a target nucleotide sequence of a polynucleotide are useful in the methods and compositions described herein. The nucleobase editors described herein typically include a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain (e.g., adenosine deaminase or cytidine deaminase). The polynucleotide programmable nucleotide binding domain, when combined with a bound guide polynucleotide (e.g., gRNA), can specifically bind to a target polynucleotide sequence, thereby localizing the base editor to the target nucleic acid sequence desired to be edited.
[0242] In certain embodiments, the nucleobase editors provided herein include one or more features that improve base editing activity. For example, any of the nucleobase editors provided herein may include a Cas9 domain with reduced nuclease activity. In some embodiments, any of the nucleobase editors provided herein may have a Cas9 domain that does not have nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, referred to as Cas9 nickase (nCas9). Without wishing to be bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of Cas9 to cleave the non-edited (e.g., non-deaminated) strand opposite the targeted nucleobase. Mutation of the catalytic residue (e.g., D10 to A10) prevents cleavage of the edited (e.g., deaminated) strand containing the targeted residue (e.g., A or C). Such Cas9 variants can generate single-stranded DNA breaks (nicks) at specific positions based on the gRNA-defined target sequence, leading to repair of the non-edited strand and ultimately to nucleobase changes on the non-edited strand.
[0243] Polynucleotide Programmable Nucleotide Binding Domains The polynucleotide programmable nucleotide binding domain binds to a polynucleotide (e.g., RNA, DNA). The polynucleotide programmable nucleotide binding domain of a base editor may itself comprise one or more domains (e.g., one or more nuclease domains). In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain may comprise an endonuclease or an exonuclease. An endonuclease can cleave one strand of a double-stranded nucleic acid, or both strands of a double-stranded nucleic acid molecule. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can cleave zero, one, or two strands of a target polynucleotide.
[0244] Fusion Proteins with Internal Insertions Provided herein are fusion proteins that include a heterologous polypeptide fused to a nucleic acid programmable nucleic acid binding protein, such as a nucleic acid programmable DNA binding protein (napDNAbp). The heterologous polypeptide can be a polypeptide that is not found in the native or wild-type napDNAbp polypeptide sequence. The heterologous polypeptide can be fused to the napDNAbp at the C-terminus of the napDNAbp, the N-terminus of the napDNAbp, or can be inserted at an internal position of the napDNAbp.
[0245] In some embodiments, the heterologous polypeptide is inserted at an internal position of the napDNAbp. In some embodiments, the heterologous polypeptide is a deaminase (e.g., adenosine deaminase) or a functional fragment thereof. For example, the fusion protein can include a deaminase (e.g., adenosine deaminase) adjacent to an N-terminal fragment and a C-terminal fragment of a Cas9 polypeptide. The deaminase in the fusion protein can be an adenosine deaminase. In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10 or a variant thereof).
[0246] In some embodiments, the fusion protein has the structure: NH2-[N-terminal fragment of napDNAbp]-[deaminase]-[C-terminal fragment of napDNAbp]-COOH, NH2-[N-terminal fragment of Cas9]-[adenosine deaminase]-[C-terminal fragment of Cas9]-COOH; where each instance of "]-[" is an optional linker.
[0247] The deaminase may be a circular permutant deaminase. For example, the deaminase may be a circular permutant adenosine deaminase. In some embodiments, the deaminase is a circular permutant TadA, circularly permuted at amino acid residue 116, as numbered in the TadA reference sequence. In some embodiments, the deaminase is a circular permutant TadA, circularly permuted at amino acid residue 136, as numbered in the TadA reference sequence. In some embodiments, the deaminase is a circular permutant TadA, circularly permuted at amino acid residue 65, as numbered in the TadA reference sequence.
[0248] The fusion protein may comprise two or more deaminases. The fusion protein may comprise, for example, one, two, three, four, five or more deaminases. In some embodiments, the fusion protein comprises one deaminase. In some embodiments, the fusion protein comprises two deaminases. The two or more deaminases in the fusion protein may be adenosine deaminase, cytidine deaminase, or a combination thereof. The two or more deaminases may be homodimers. The two or more deaminases may be heterodimers. The two or more deaminases may be inserted in tandem in the napDNAbp. In some embodiments, the two or more deaminases may not be in tandem in the napDNAbp.
[0249] In some embodiments, the napDNAbp in the fusion protein is a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a variant Cas9 polypeptide. In some embodiments, the Cas9 polypeptide is a Cas9 nickase (nCas9) polypeptide or a fragment thereof. In some embodiments, the Cas9 polypeptide is a nuclease-dead Cas9 (dCas9) polypeptide or a fragment thereof. The Cas9 polypeptide in the fusion protein can be a full-length Cas9 polypeptide. In some cases, the Cas9 polypeptide in the fusion protein may not be a full-length Cas9 polypeptide. The Cas9 polypeptide can be truncated, for example, at the N-terminus or C-terminus relative to a naturally occurring Cas9 protein. The Cas9 polypeptide can be a circularly permuted Cas9 protein. The Cas9 polypeptide can be a fragment, portion, or domain of a Cas9 polypeptide that is still capable of binding to a target polynucleotide and a guide nucleic acid sequence.
[0250] In some embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus1 Cas9 (St1Cas9), or a fragment or variant thereof.
[0251] Fusion proteins comprising heterologous catalytic domains flanking the N- and C-terminal fragments of Cas9 polypeptide are also useful for base editing in the methods described herein. Fusion proteins comprising Cas9 and one or more deaminase domains, such as adenosine deaminase or comprising an adenosine deaminase domain flanking a Cas9 sequence, are also useful for highly specific and efficient base editing of target sequences. In some embodiments, chimeric Cas9 fusion proteins contain heterologous catalytic domains (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) inserted into a Cas9 polypeptide. In some embodiments, the fusion protein comprises an adenosine deaminase domain and a cytidine deaminase domain inserted into Cas9. In some embodiments, the adenosine deaminase is fused within Cas9 and the cytidine deaminase is fused to the C-terminus. In some embodiments, adenosine deaminase is fused within Cas9 and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused within Cas9 and adenosine deaminase is fused to the C-terminus. In some embodiments, cytidine deaminase is fused within Cas9 and adenosine deaminase is fused to the N-terminus.
[0252] Exemplary structures of fusion proteins with adenosine deaminase and cytidine deaminase and Cas9 are provided below: NH2-[Cas9(adenosine deaminase)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas9(adenosine deaminase)]-COOH, NH2-[Cas9(cytidine deaminase)]-[adenosine deaminase]-COOH, or NH2-[adenosine deaminase]-[Cas9 (cytidine deaminase)]-COOH.
[0253] In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.
[0254] In various embodiments, the catalytic domain has a DNA modifying activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10). In some embodiments, the TadA is a TadA variant. In some embodiments, the TadA variant is fused in Cas9 and the cytidine deaminase is fused to the C-terminus. In some embodiments, the TadA variant is fused in Cas9 and the cytidine deaminase is fused to the N-terminus. In some embodiments, the cytidine deaminase is fused in Cas9 and the TadA variant is fused to the C-terminus. In some embodiments, the cytidine deaminase is fused in Cas9 and the TadA variant is fused to the N-terminus. Exemplary structures of fusion proteins having a TadA variant and a cytidine deaminase and Cas9 are provided below: NH2-[Cas9(TadA variant)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas9(TadA variant)]-COOH, NH2-[Cas9(cytidine deaminase)]-[TadA variant]-COOH, or NH2-[TadA variant]-[Cas9 (cytidine deaminase)]-COOH.
[0255] In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.
[0256] In other embodiments, the fusion protein contains a nuclear localization signal (e.g., a bipartite nuclear localization signal). In other embodiments, the amino acid sequence of the nuclear localization signal is MAPKKKRKVGIHGVPAA (SEQ ID NO: 4). In other embodiments of the above aspects, the nuclear localization signal is encoded by the following sequence: ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCC (SEQ ID NO: 5). In other embodiments, the Cas12b polypeptide contains a mutation that silences the catalytic activity of the RuvC domain. In other embodiments, the Cas12b polypeptide contains a D574A, D829A, and / or D952A mutation. In other embodiments, the fusion protein further contains a tag (e.g., an influenza hemagglutinin tag).
[0257] In some embodiments, the fusion protein comprises a napDNAbp domain (e.g., a domain from Cas12) with an internally fused nucleobase editing domain (e.g., a deaminase domain, e.g., all or a portion of an adenosine deaminase domain). In some embodiments, the napDNAbp is Cas12b.
[0258] As a non-limiting example, an adenosine deaminase (e.g., TadA*8.13) can be inserted into BhCas12b to produce a fusion protein (e.g., TadA*8.13-BhCas12b) that effectively edits a nucleic acid sequence.
[0259] Gene editing using gRNA The NLS-gRNA described herein can be used with a suitable gene editing system for targeted gene editing, which can result in gene silencing events or expression alterations (e.g., increases or decreases) in the expression of a desired target gene. Thus, in some embodiments, the NLS-gRNA described herein can be used in a method for targeted transcriptional activation, targeted transcriptional repression, targeted epigenetic modification, or targeted genomic modification, the method comprising introducing into a eukaryotic cell (a) an NLS-gRNA as defined herein, (b) at least one CRISPR / Cas protein, or a nucleic acid encoding at least one CRISPR / Cas protein, wherein the interaction of (a) and (b) with a target sequence in chromosomal DNA results in targeted transcriptional activation, targeted transcriptional repression, targeted epigenetic modification, or targeted genomic modification.
[0260] In some embodiments, the NLS-gRNAs described herein can be used in a gene editing system comprising: an NLS-gRNA described herein, wherein the RNA guide comprises a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid; a gene editing protein, wherein a gene editing enzyme is capable of binding to the RNA guide and causing cleavage of a target nucleic acid sequence complementary to the RNA guide.
[0261] In some embodiments, the NLS-gRNA described herein can be used in a gene editing system comprising: an NLS-gRNA described herein, where the RNA guide comprises a direct repeat sequence and a spacer sequence capable of hybridizing to a target nucleic acid; and a gene editing protein; where the gene editing protein is fused to a deaminase, and the gene editing protein fusion is capable of binding to the RNA guide and editing a target nucleic acid sequence complementary to the RNA guide.
[0262] In some embodiments, the present invention provides a method of modifying expression of a target nucleic acid in a eukaryotic cell, comprising contacting the cell with a gene-editing protein and an NLS-gRNA as described herein, wherein the NLS-gRNA comprises a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid, and wherein the gene-editing protein is capable of binding to the NLS-gRNA and causing cleavage of the target nucleic acid sequence complementary to the NLS-gRNA.
[0263] In some embodiments, the invention provides a method of modifying expression of a target nucleic acid in a eukaryotic cell, comprising contacting the cell with a gene-editing protein and a synthetic NLS-gRNA as described herein, wherein the NLS-gRNA comprises a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid, and wherein the gene-editing protein is capable of binding to the NLS-gRNA and editing the target nucleic acid sequence complementary to the NLS-gRNA.
[0264] In some embodiments, the invention provides a method of modifying a target nucleic acid in a eukaryotic cell, comprising contacting the cell with a gene-editing protein and an NLS-gRNA described herein, wherein the NLS-gRNA comprises a direct repeat sequence and a spacer sequence capable of hybridizing to the target nucleic acid, and wherein the gene-editing protein is capable of binding to the NLS-gRNA and editing the target nucleic acid sequence complementary to the NLS-gRNA.
[0265] In some embodiments, a gene editing method or system comprises a fusion protein having an effector that modifies target DNA in a site-specific manner, and the modification activity comprises a methyltransferase activity, a demethylase activity, an acetyltransferase activity, a deacetylase activity, a kinase activity, a phosphatase activity, a ubiquitin ligase activity, a deubiquitinating activity, an adenylating activity, a deadenylating activity, a sumoylating activity, a desumoylating activity, a ribosylation activity, a deribosylation activity, a myristoylating activity, a demyristoylating activity, an integrase activity, a transposase activity, a recombinase activity, a polymerase activity, a ligase activity, a helicase activity, or a nuclease activity, any of which can modify DNA or a DNA-associated polypeptide (e.g., a histone or a DNA-binding protein).
[0266] In some embodiments, the gene editing method or system includes a fusion protein with an enzyme that can edit DNA sequences by chemically modifying nucleotide bases, including a deaminase enzyme that can modify adenosine or cytosine bases and function as a site-specific base editor. For example, APOBEC1 cytidine deaminase, which typically uses RNA as a substrate, can target single-stranded and double-stranded DNA when fused to Cas9, directly converting cytidine to uridine, and ADAR enzymes deaminating adenosine to inosine. Thus, "base editing" using deaminases allows for programmable conversion of one target DNA base to another. Various base editors are known in the art and can be used in the methods and systems described herein. Exemplary base editors are described, for example, in Rees and Liu Nature Review Genetics, 2018, 19(12):770-788, the contents of which are incorporated herein.
[0267] In some embodiments, base editing results in the introduction of a stop codon to silence a gene, in some embodiments, base editing results in altered protein function by modifying the amino acid sequence.
[0268] In some embodiments, the NLS-gRNA described herein can be used in gene editing methods or systems to regulate the transcription of target DNA. In some embodiments, the NLS-gRNA can be used in gene editing methods or systems to regulate the expression of target non-coding RNA, including tRNA, rRNA, snoRNA, siRNA, miRNA, and long ncRNA.
[0269] In some embodiments, the NLS-gRNA described herein is used for targeted manipulation of chromatin loop structures using a suitable gene editing system. Targeted manipulation of chromatin loops between regulatory genomic regions provides a means to manipulate endogenous chromatin structure and allow the formation of new enhancer-promoter connections to overcome genetic defects or inhibit aberrant enhancer-promoter connections.
[0270] In some embodiments, the NLS-gRNAs described herein are used in conjunction with gene editing systems for the correction of pathogenic mutations by insertion of beneficial clinical variants or suppressor mutations.
[0271] Editing A to G In some embodiments, the base editors described herein comprise an adenosine deaminase domain. Such an adenosine deaminase domain of a base editor can facilitate the editing of an adenine (A) nucleobase to a guanine (G) nucleobase by deaminating A to form inosine (I), which exhibits the base pairing properties of G. Adenosine deaminase can deaminate (i.e., remove the amine group) the adenine of a deoxyadenosine residue in deoxyribonucleic acid (DNA). In some embodiments, the A to G base editor further comprises an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without wishing to be bound by any particular theory, the UGI domain or catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deaminated adenosine residues (e.g., inosine), which can improve the activity or efficiency of the base editor.
[0272] A base editor comprising an adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In certain embodiments, a base editor comprising an adenosine deaminase can deaminate target A of a polynucleotide comprising RNA. For example, a base editor can comprise an adenosine deaminase domain capable of deaminating target A of an RNA polynucleotide and / or a DNA-RNA hybrid polynucleotide. In certain embodiments, an adenosine deaminase incorporated in a base editor comprises all or a portion of an adenosine deaminase acting on RNA (ADAR, e.g., ADAR1 or ADAR2) or an adenosine deaminase acting on tRNA (ADAT). A base editor comprising an adenosine deaminase domain can also deaminate A nucleobase of a DNA polynucleotide. In certain embodiments, an adenosine deaminase domain of a base editor comprises all or a portion of an ADAT comprising one or more mutations that enable ADAT to deaminate target A in DNA. For example, a base editor can include all or a portion of ADAT from Escherichia coli (EcTadA) that includes one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I156F, or a corresponding mutation in another adenosine deaminase.
[0273] In some embodiments, base editors described herein comprise fusion proteins that include an adenosine deaminase domain (e.g., an adenosine deaminase variant domain). In some embodiments, the adenosine deaminase variant domain contains a combination of modifications in the TadA*7.10 amino acid sequence, the combination being V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N. In some embodiments, the combination of modifications in the TadA*7.10 amino acid sequence is V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147D+F149Y+Q154S+N157K, 4S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, or L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N, or the corresponding modifications in another adenosine deaminase. Such an adenosine deaminase domain of a base editor can facilitate editing of an adenine (A) nucleobase to a guanine (G) nucleobase by deaminating A to form inosine (I), which exhibits the base pairing properties of G. Adenosine deaminase is capable of deaminating (ie, removing the amine group from) the adenine of deoxyadenosine residues in deoxyribonucleic acid (DNA).
[0274] In some embodiments, the nucleobase editors provided herein can be made by fusing one or more protein domains together, thereby generating a fusion protein. In certain embodiments, the fusion proteins provided herein include one or more features that improve the base editing activity (e.g., efficiency, selectivity, and specificity) of the fusion protein. For example, the fusion proteins provided herein can include a Cas9 domain with reduced nuclease activity. In some embodiments, the fusion proteins provided herein can have a Cas9 domain with no nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, referred to as Cas9 nickase (nCas9). Without wishing to be bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of Cas9 to cleave the non-edited (e.g., non-deaminated) strand containing a T opposite the targeted A. Mutation of the catalytic residue of Cas9 (e.g., D10 to A10) prevents cleavage of the edited strand containing the targeted A residue. Such Cas9 variants can generate single-stranded DNA breaks (nicks) at specific positions based on the gRNA-defined target sequence, resulting in repair of the non-edited strand, and ultimately a T to C change on the non-edited strand. In some embodiments, the A to G base editor further comprises an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without wishing to be bound by any particular theory, a UGI domain or a catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deaminated adenosine residues (e.g., inosine), which can improve the activity or efficiency of the base editor.
[0275] The base editor comprising adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In certain embodiments, the base editor comprising adenosine deaminase can deaminate target A of a polynucleotide comprising RNA. For example, the base editor can comprise an adenosine deaminase domain capable of deaminating target A of an RNA polynucleotide and / or a DNA-RNA hybrid polynucleotide. In certain embodiments, the adenosine deaminase incorporated in the base editor comprises all or a portion of an adenosine deaminase acting on RNA (ADAR, e.g., ADAR1 or ADAR2). In another embodiment, the adenosine deaminase incorporated in the base editor comprises all or a portion of an adenosine deaminase acting on tRNA (ADAT). The base editor comprising an adenosine deaminase domain can also deaminate A nucleobase of a DNA polynucleotide. In some embodiments, the adenosine deaminase domain of a base editor comprises all or a portion of ADAT that includes one or more mutations that enable ADAT to deaminate a target A in DNA. For example, a base editor can comprise all or a portion of ADAT from Escherichia coli (EcTadA) that includes one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I156F, or a corresponding mutation in another adenosine deaminase.
[0276] The adenosine deaminase may be from any suitable organism (e.g., E. coli). In some embodiments, the adenosine deaminase is from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from E. coli. In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations that correspond to any of the mutations provided herein (e.g., mutations in ecTadA). Corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determining homologous residues. Mutations in any naturally occurring adenosine deaminase (e.g., with homology to ecTadA) that correspond to any of the mutations described herein (e.g., any of the mutations identified in ecTadA) can be generated accordingly.
[0277] Adenosine deaminase In some embodiments, the fusion proteins described herein comprise one or more adenosine deaminase domains. In some embodiments, the adenosine deaminase provided herein can deaminate adenine. In some embodiments, the adenosine deaminase provided herein can deaminate adenine at deoxyadenosine residues in DNA. The adenosine deaminase can be derived from any suitable organism (e.g., E. coli). In some embodiments, the adenine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations that correspond to any of the mutations provided herein (e.g., mutations in ecTadA). One of skill in the art would be able to identify the corresponding residues in any homologous protein, for example, by sequence alignment and determining homologous residues. Thus, one of skill in the art would be able to generate mutations in any naturally occurring adenosine deaminase (e.g., with homology to ecTadA) that correspond to any of the mutations described herein, for example, any of the mutations identified in ecTadA. In some embodiments, the adenosine deaminase is from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from E. coli.
[0278] Provided and described herein are adenosine deaminase variants with increased efficiency (>50-60%) and specificity. In particular, the adenosine deaminase variants described herein are more likely to edit desired bases in a polynucleotide and less likely to edit bases that are not intended to be modified (i.e., "bystanders").
[0279] In some embodiments, the adenosine deaminase is derived from TadA deaminase. In certain embodiments, TadA is any one of the TadAs described in PCT / US2017 / 045381 (WO2018 / 027078) (incorporated herein by reference in its entirety).
[0280] The wild-type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence): MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 6)
[0281] In some embodiments, the adenosine deaminase is a full-length E. coli TadA deaminase. For example, in certain embodiments, the adenosine deaminase comprises the following amino acid sequence: MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 7).
[0282] In some embodiments, the adenosine deaminase is from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli (E. coli), Staphylococcus aureus (S. aureus), Salmonella typhimurium (S. typhimurium), Shewanella putrefaciens (S. putrefaciens), Haemophilus influenzae (H. influenzae), Caulobacter crescentus (C. crescentus), Geobacter sulfurreducens (G. sulfurreducens), or Bacillus subtilis. In some embodiments, the adenosine deaminase is from E. coli.
[0283] However, it should be understood that additional adenosine deaminases useful in the present application will be apparent to those skilled in the art and are within the scope of the present disclosure. For example, the adenosine deaminase may be a homolog of adenosine deaminase acting on tRNA (ADAT). Without being limited thereto, exemplary amino acid sequences of AD AT homologs include the following:
[0284] Staphylococcus aureus (S. aureus)TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN (SEQ ID NO: 8)
[0285] Bacillus subtilis (B.subtilis)TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE(SEQ ID NO:9)
[0286] Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV(SEQ ID NO:10)
[0287] Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE(SEQ ID NO:11)
[0288] Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK(SEQ ID NO:12)
[0289] Caulobacter crescentus (C.crescentus)TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI (SEQ ID NO: 13)
[0290] Geobacter sulfurreducens(G.sulfurreducens)TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP (SEQ ID NO: 14)
[0291] Some embodiments of E. Coli TadA (ecTadA) include: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 3)
[0292] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described in any of the adenosine deaminases provided herein. It should be understood that the adenosine deaminases provided herein may comprise one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain that has a particular percent identity, as well as any of the mutations described herein or a combination thereof. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to a reference sequence or any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence that has at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical contiguous amino acid residues compared to any one of the amino acid sequences known in the art or described herein.
[0293] It should be understood that any of the mutations provided herein (e.g., based on the TadA reference sequence) can be introduced into other adenosine deaminases, such as E. coli TadA (ecTadA), S. aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). It will be apparent to one of skill in the art that additional deaminases can be similarly aligned to identify homologous amino acid residues that can be mutated as provided herein. Thus, any of the mutations identified in the TadA reference sequence can be made in other adenosine deaminases (e.g., ecTada) that have homologous amino acid residues. It should also be understood that any of the mutations provided herein can be made in the TadA reference sequence or another adenosine deaminase, either individually or in any combination.
[0294] In some embodiments, the adenosine deaminase comprises a D108X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108G, D108N, D108V, D108A, or D108Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. However, it should be understood that additional deaminases can be similarly aligned to identify homologous amino acid residues that can be mutated as provided herein.
[0295] In some embodiments, the adenosine deaminase comprises an A106X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0296] In some embodiments, the adenosine deaminase comprises an E155X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0297] In some embodiments, the adenosine deaminase comprises a D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where the presence of X denotes any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D147Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0298] In some embodiments, the adenosine deaminase comprises an A106X, E155X, or D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation. In some embodiments, the adenosine deaminase comprises D147Y.
[0299] It should be understood that any of the mutations provided herein (e.g., based on the ecTadA amino acid sequence of the TadA reference sequence) can be introduced into other adenosine deaminases, such as S. aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). The degree of homology to the mutated residues in ecTadA will be apparent to one of skill in the art. Thus, any of the mutations identified in ecTadA can be made in other adenosine deaminases that have homologous amino acid residues. It should also be understood that any of the mutations provided herein can be made in ecTadA or another adenosine deaminase, individually or in any combination.
[0300] For example, adenosine deaminase can be expressed by a combination of mutations (e.g., V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147D+F149Y+Q154S+D167N, The TadA reference sequence may contain the following mutations: S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, or L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N) and may contain one or more additional mutations, such as D108N, A106V, E155V, and / or D147Y mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises the following mutations in the TadA reference sequence (mutations are separated by ";"), or corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V, and E155V; D108N, A106V, and D147Y; D108N, E155V, and D147Y; A106V, E155V, and D147Y; and D108N, A106V, E155V, and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be made in an adenosine deaminase (eg, ecTadA).
[0301] In some embodiments, the adenosine deaminase comprises one or more of the H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X, and / or K157X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase is selected from the group consisting of H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E, or A56S, E59G, E85K, or E85G, M94L, I95L, V102A, F104L, A106V, R107C, or R107H, or R107P in the TadA reference sequence; It comprises one or more of the following mutations: D108G, or D108N, or D108V, or D108A, or D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or K157R, or one or more corresponding mutations in another adenosine deaminase.
[0302] In some embodiments, the adenosine deaminase comprises one or more of H8X, D108X, and / or N127X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of H8Y, D108N, and / or N127S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0303] In some embodiments, the adenosine deaminase comprises one or more of the H8X, R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X, K161X, Q163X, and / or T166X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, R26W, M61I, L68Q, M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H, or Q154R, E155G, or E155V, or E155D, K161Q, Q163H, and / or T166P mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0304] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, D108X, N127X, D147X, R152X, and Q154X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0305] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, A106X, and D108X, or a corresponding mutation, or a mutation in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, R26X, L68X, D108X, N127X, D147X, and E155X, or a corresponding mutation, or a mutation in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0306] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of H8X, R126X, L68X, D108X, N127X, D147X, and E155X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase, where X denotes the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, A109X, N127X, and E155X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase, where X denotes the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0307] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, R26W, L68Q, D108N, N127S, D147Y, and E155V in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase (e.g., ecTadA).
[0308] In some embodiments, the adenosine deaminase comprises one or more of the corresponding mutations in another adenosine deaminase or one or more corresponding mutations. In some embodiments, the adenosine deaminase comprises a D108N, D108G, or D108V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V and a D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107C and a D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and Q154H mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises D108N, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, and N127S mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the A106V, D108N, D147Y, and E155V mutations in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0309] In some embodiments, the adenosine deaminase comprises one or more of the S2X, H8X, I49X, L84X, H123X, N127X, I156X, and / or K160X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where the presence of X denotes any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the S2A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0310] In some embodiments, the adenosine deaminase comprises an L84X mutant adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase, In some embodiments, the adenosine deaminase comprises an L84F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0311] In some embodiments, the adenosine deaminase comprises an H123X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H123Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0312] In some embodiments, the adenosine deaminase comprises an I156X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an I156F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0313] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84X, A106X, D108X, H123X, D147X, E155X, and I156X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, A106X, D108X, N127X, and K160X in the TadA reference sequence, or a corresponding mutation, or a mutation in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0314] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F in the TadA reference sequence, or corresponding mutations, or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V in the TadA reference sequence.
[0315] In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, A106T, D108N, N127S, and K160S in the TadA reference sequence, or corresponding mutations, or mutations in another adenosine deaminase.
[0316] In some embodiments, the adenosine deaminase comprises one or more of E25X, R26X, R107X, A142X, and / or A143X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the E25M, E25D, E25A, E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R107K, R107A, R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the mutations described herein corresponding to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0317] In some embodiments, the adenosine deaminase comprises an E25X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E25M, E25D, E25A, E25R, E25V, E25S, or E25Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0318] In some embodiments, the adenosine deaminase comprises an R26X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R26G, R26N, R26Q, R26C, R26L, or R26K mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0319] In some embodiments, the adenosine deaminase comprises a R107X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a R107P, R107K, R107A, R107N, R107W, R107H, or R107S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0320] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N, A142D, A142G mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0321] In some embodiments, the adenosine deaminase comprises an A143X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0322] In some embodiments, the adenosine deaminase comprises one or more of the H36X, N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X, and / or K161X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the H36L, N37T, N37S, P48T, P48L, I49V, R51H, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0323] In some embodiments, the adenosine deaminase comprises an H36X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase, In some embodiments, the adenosine deaminase comprises an H36L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0324] In some embodiments, the adenosine deaminase comprises an N37X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an N37T or N37S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0325] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48T or P48L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0326] In some embodiments, the adenosine deaminase comprises an R51X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R51H or R51L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0327] In some embodiments, the adenosine deaminase comprises a S146X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a S146R or S146C mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0328] In some embodiments, the adenosine deaminase comprises a K157X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a K157N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0329] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48S, P48T, or P48A mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0330] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0331] In some embodiments, the adenosine deaminase comprises a W23X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a W23R or W23L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0332] In some embodiments, the adenosine deaminase comprises a R152X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a R152P or R52H mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0333] In one embodiment, the adenosine deaminase may comprise the mutations H36L, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F, and K157N. In some embodiments, the adenosine deaminase comprises the following combinations of mutations relative to the TadA reference sequence, where each mutation in the combination is separated by an "_" and each combination of mutations is between brackets: (A106V_D108N), (R107C_D108N), (H8Y_D108N_N127S_D147Y_Q154H), (H8Y_D108N_N127S_D147Y_E155V), (D108N_D147Y_E155V), (H8Y_D108N_N127S), (H8Y_D108N_N127S_D147Y_Q154H), (A106V_D108N_D147Y_E155V), (D108Q_D147Y_E155V), (D108M_D147Y_E155V), (D108L_D147Y_E155V), (D108K_D147Y_E155V), (D108I_D147Y_E155V), (D108F_D147Y_E155V), (A106V_D108N_D147Y), (A106V_D108M_D147Y_E155V), (E59A_A106V_D108N_D147Y_E155V)、 (E59A cat dead_A106V_D108N_D147Y_E155V)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (D103A_D104N)、 (G22P_D103A_D104N)、 (D103A_D104N_S138A)、 (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155V_I156F)、(R26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、(R26C_L84F_A106V_R107H_D108N_H123Y_A142N_D147Y_E155V_I156F)、(L84F_A106V_D108N_H123Y_A142N_A143L_D147Y_E155V_I156F)、 (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I156F)、 (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (A106V_D108N_A142N_D147Y_E155V)、 (R26G_A106V_D108N_A142N_D147Y_E155V)、 (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V)、 (R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V)、 (E25D_R26G_A106V_D108N_A142N_D147Y_E155V)、 (A106V_R107K_D108N_A142N_D147Y_E155V)、 (A106V_D108N_A142N_A143G_D147Y_E155V)、 (A106V_D108N_A142N_A143L_D147Y_E155V)、 (H36L_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F)、 (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F)、 (N72S_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F)、 (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N) (H36L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N)、 (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E)、 (H36L_G67V_L84F_A106V_D108N_H123Y_S146T_D147Y_E155V_I156F)、 (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F)、 (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F)、 (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、 (N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E)、(R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F)、 (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (P48S_A142N)、 (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N)、 (P48T_I49V_A142N)、 (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F)、 (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_156F_K157N)、 (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、(H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152H_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_R152P_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_R152P_E155V_I156F_K157N)。
[0334] In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*7.10. In certain embodiments, the fusion protein comprises a single TadA*7.10 domain (e.g., provided as a monomer). In other embodiments, the fusion protein comprises TadA*7.10 and TadA(wt), which can form a heterodimer. In one embodiment, the fusion protein described herein comprises wild-type TadA bound to TadA*7.10, which is bound to a Cas9 nickase.
[0335] In some embodiments, TadA*7.10 comprises at least one modification. In some embodiments, the adenosine deaminase comprises a modification in the following sequence: TadA*7.10
[0336] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 3)
[0337] In some embodiments, TadA*7.10 comprises modifications at amino acids 82 and / or 166. In certain embodiments, TadA*7.10 comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, variants of TadA*7.10 comprise a combination of modifications selected from the following group: Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V82 S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R.
[0338] In some embodiments, the variant of TadA*7.10 comprises one or more modifications selected from the group of L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N. In some embodiments, the variant of TadA*7.10 comprises V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N. In other embodiments, the variant of TadA*7.10 comprises a combination of modifications selected from the following group: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36 H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y +V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0339] In some embodiments, the adenosine deaminase variant (e.g., a TadA variant) comprises a deletion. In some embodiments, the adenosine deaminase variant comprises a C-terminal deletion. In certain embodiments, the adenosine deaminase variant comprises a C-terminal deletion beginning at residues 149, 150, 151, 152, 153, 154, 155, 156, and 157 relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA.
[0340] In other embodiments, the adenosine deaminase variant (e.g., TadA*8) is a monomer that includes one or more of the following modifications relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant (TadA*8) is selected from the group consisting of Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V8 and I76Y+V82S+Y123H+Y147R+Q154R.
[0341] In other embodiments, a base editor of the disclosure comprises an adenosine deaminase variant (e.g., TadA*8) monomer that includes one or more of the following modifications, R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and / or D167N, relative to TadA*7.10, a TadA reference sequence, or a corresponding mutation in another TadA. In another embodiment, the adenosine deaminase variant (TadA*8) monomer is selected from the group consisting of R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A ... N+F149Y+T166I+D167N, R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N.
[0342] In some embodiments, the adenosine deaminase variant (e.g., MSP828) is a monomer that includes one or more of the following modifications, L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N, relative to TadA*7.10, the TadA reference sequence, or the corresponding mutation in another TadA. In some embodiments, the adenosine deaminase variant (e.g., MSP828) is a monomer that includes V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N, relative to TadA*7.10, the TadA reference sequence, or the corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variant (TadA variant) is selected from the group consisting of V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V ... 36H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0343] In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each having one or more of the following modifications relative to TadA*7.10, a TadA reference sequence, or a corresponding mutation in another TadA: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant is a homodimer containing two adenosine deaminase domains (e.g., TadA*8), each of which has one or more of the following mutations relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA: Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V ... S, V82S+Y123H+Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R.
[0344] In other embodiments, a base editor of the disclosure comprises an adenosine deaminase variant (e.g., TadA*8) homodimer that includes two adenosine deaminase domains (e.g., TadA*8) with one or more of the following modifications, R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and / or D167N, each relative to a corresponding mutation in TadA*7.10, a TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a homodimer containing two adenosine deaminase domains (e.g., TadA*8), each of which has the following mutations relative to TadA*7.10, the TadA reference sequence, or the corresponding mutations in another TadA: R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A+A109S, or V90A+A109S. + T111R + D119N + H122N + F149Y + T166I + D167N, R26C + A109S + T111R + D119N + H122N + F149Y + T166I + D167N, V88A + T111R + D119N + F149Y, and A109S + T111R + D119N + H122N + Y147D + F149Y + T166I + D167N.
[0345] In some embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*7.10), each having one or more of the following modifications relative to TadA*7.10, a TadA reference sequence, or a corresponding mutation in another TadA: L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N. In some embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase variant domains (e.g., MSP828), each having one or more of the following modifications relative to TadA*7.10, the TadA reference sequence, or corresponding mutations in another TadA: V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N. In other embodiments, the adenosine deaminase variant is a homodimer that includes two adenosine deaminase domains (e.g., TadA*7.10), each of which has one or more of the following mutations relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147 D+F149Y+Q154S+D167N, L36H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+D167N.
[0346] In other embodiments, the adenosine deaminase variant is a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8) that includes one or more of the following modifications, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variants include the following mutations relative to the wild-type adenosine deaminase domain: Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V82S+Y123H+Y147R ... R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R.
[0347] In other embodiments, the base editor comprises a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain that includes one or more of the following modifications: R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and / or D167N, relative to TadA*7.10, a TadA reference sequence, or a corresponding mutation in another TadA (e.g., TadA*8). In other embodiments, the base editor is a base pair that is modified to include any of the following mutations relative to the wild-type adenosine deaminase domain: R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, R and a heterodimer with an adenosine deaminase variant domain (e.g., TadA*8) comprising a combination of modifications selected from the group consisting of 26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N.
[0348] In other embodiments, the adenosine deaminase variant is a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*7.10) that includes one or more of the following modifications: L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N, relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA. In some embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant domain having one or more of the following modifications, V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N, relative to TadA*7.10, the TadA reference sequence, or corresponding mutations in another TadA (e.g., MSP828). In other embodiments, the adenosine deaminase variant comprises a wild-type adenosine deaminase domain and a corresponding mutation in TadA*7.10, a TadA reference sequence, or another TadA, such as V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147D+ F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0349] In other embodiments, the adenosine deaminase variant is a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8) that includes one or more of the following modifications, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to TadA*7.10, the TadA reference sequence, or the corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variant comprises a TadA*7.10 domain and a corresponding mutation in TadA*7.10, a TadA reference sequence, or another TadA, such as Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V82S+Y123H+Y147R, and a heterodimer with an adenosine deaminase variant domain (e.g., TadA*8) comprising a combination of modifications selected from the group: V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R.
[0350] In other embodiments, the base editor comprises a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8) that includes one or more of the following modifications, R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and / or D167N, relative to TadA*7.10, a TadA reference sequence, or a corresponding mutation in another TadA. In other embodiments, the base editor is a TadA*7.10 domain and a corresponding mutation in TadA*7.10, a TadA reference sequence, or another TadA, such as R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, R26C+A109S+T111R+D119N+H122N ... + A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N.
[0351] In other embodiments, the adenosine deaminase variant is a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*7.10) that includes one or more of the following modifications: L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N, relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA. In some embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., MSP828) having one or more of the following modifications relative to TadA*7.10, the TadA reference sequence, or corresponding mutations in another TadA: V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N. In other embodiments, the adenosine deaminase variant comprises a TadA*7.10 domain and a TadA*7.10 domain that is identical to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA domain, and is selected from the group consisting of V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V ... 9Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0352] In some embodiments, TadA*8 is a variant shown in Tables 8A, 10, 11, or 13. Tables 8A, 10, 11, and 13 show the number of specific amino acid positions in the TadA amino acid sequence and the amino acid present at those positions in TadA-7.10 adenosine deaminase. Tables 8A, 10, 11, and 13 also show the amino acid changes of TadA variants relative to TadA-7.10 after phage-assisted discontinuous evolution (PANCE) and phage-assisted continuous evolution (PACE) as described in M. Richter et al., 2020, Nature Biotechnology, doi.org / 10.1038 / s41587-020-0453-z, which is incorporated herein by reference in its entirety. In some embodiments, TadA*8 is TadA*8a, TadA*8b, TadA*8c, TadA*8d, or TadA*8e. In some embodiments, TadA*8 is TadA*8e.
[0353] In certain embodiments, the adenosine deaminase heterodimer can comprise a TadA*8 domain and an adenosine deaminase domain selected from Staphylococcus aureus (S. aureus) TadA, Bacillus subtilis (B. subtilis) TadA, Salmonella typhimurium (S. typhimurium) TadA, Shewanella putrefaciens (S. putrefaciens) TadA, Haemophilus influenzae F3031 (H. influenzae) TadA, Caulobacter crescentus (C. crescentus) TadA, Geobacter sulfurreducens (G. sulfurreducens) TadA, or TadA*7.10.
[0354] In some embodiments, the adenosine deaminase is TadA*8. In one embodiment, the adenosine deaminase is TadA*8 that comprises or consists essentially of the following sequence, or a fragment thereof, having adenosine deaminase activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 16)
[0355] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated TadA*8 is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full length TadA*8. In some embodiments, the truncated TadA*8 is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full length TadA*8. In some embodiments, the adenosine deaminase variant is full length TadA*8.
[0356] In one embodiment, a fusion protein described and / or exemplified herein comprises wild-type TadA, linked to an adenosine deaminase variant described herein (e.g., TadA*8), and linked to a Cas9 nickase. In certain embodiments, the fusion protein comprises a single TadA*8 domain (e.g., provided as a monomer). In other embodiments, the base editor comprises TadA*8 and TadA(wt), which can form a heterodimer.
[0357] In some embodiments, TadA*8 is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. [Table 1]
[0358] In some embodiments, the TadA variant is a variant shown in Table 6. Table 6 shows the number of specific amino acid positions in the TadA amino acid sequence and the amino acid present at those positions in the TadA*7.10 adenosine deaminase. In some embodiments, the TadA variant is MSP605, MSP680, MSP823, MSP824, MSP825, MSP827, MSP828, or MSP829. In some embodiments, the TadA variant is MSP828. In some embodiments, the TadA variant is MSP829. [Table 2]
[0359] In one embodiment, a fusion protein described herein comprises wild-type TadA, linked to an adenosine deaminase variant described herein, and linked to a Cas9 nickase. In certain embodiments, the fusion protein comprises a single variant TadA domain (e.g., provided as a monomer). In other embodiments, the fusion protein comprises a variant TadA and TadA(wt) capable of forming a heterodimer.
[0360] In some embodiments, the TadA variant is truncated. In some embodiments, the truncated TadA is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length TadA variant. In some embodiments, the truncated TadA variant is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length TadA variant. In some embodiments, the adenosine deaminase variant is a full-length TadA variant.
[0361] In certain embodiments, TadA*8 contains one or more mutations at any of the following positions shown in bold: In other embodiments, TadA*8 contains one or more mutations at any of the positions shown in underline: TIFF2024529425000065.tif32166
[0362] For example, TadA*8 contains modifications at amino acid positions 82 and / or 166 (e.g., V82S, T166R) relative to TadA*7.10, the TadA reference sequence, or the corresponding mutation in another TadA, including any one or more of the following, alone or in combination: Y147T, Y147R, Q154S, Y123H, and / or Q154R.
[0363] In certain embodiments, the combination of modifications is selected from the group consisting of Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V 82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R. In some embodiments, the adenosine deaminase comprises one or more of the following modifications: R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, M94V, P124W, T133K, D139L, D139M, C146R, and A158K. The one or more modifications are shown in the above sequence in underlined and bold font.
[0364] In some embodiments, the adenosine deaminase comprises one or more of the following combinations of modifications: V82S+Q154R+Y147R, V82S+Q154R+Y123H, V82S+Q154R+Y147R+Y123H, Q154R+Y147R+Y123H+I76Y+V82S, V82S+I76Y, V82S+Y147R, V82S+Y147R+Y123H, V82S+Q154R+Y123H, Q154R+Y147R+Y 123H+I76Y, V82S+Y147R, V82S+Y147R+Y123H, V82S+Q154R+Y123H, V82S+Q154R+Y147R, V82S+Q154R+Y147R, Q154R+Y147R+Y123H+I76Y, Q154R+Y147R+Y123H+I76Y+V82S, I76Y_V82S_Y123H_Y147R_Q154R, Y147R+Q154R+H123H, and V82S+Q154R.
[0365] In some embodiments, the adenosine deaminase comprises one or more of the following combinations of modifications: E25F+V82S+Y123H, T133K+Y147R+Q154R, E25F+V82S+Y123H+Y147R+Q154R, L51W+V82S+Y123H+C146R+Y147R+Q154R, Y73S+V82S+Y123H+Y147R+Q154R, P54C+V82S +Y123H+Y147R+Q154R, N38G+V82T+Y123H+Y147R+Q154R, N72K+V82S+Y123H+D139L+Y147R+Q154R, E25F+V82S +Y123H+D139M+Y147R+Q154R, Q71M+V82S+Y123H+Y147R+Q154R, E25F+V82S+Y123H+T133K+Y147R+Q154R, E25 F+V82S+Y123H+Y147R+Q154R, V82S+Y123H+P124W+Y147R+Q154R, L51W+V82S+Y123H+C146R+Y147R+Q154R, P5 4C+V82S+Y123H+Y147R+Q154R, Y73S+V82S+Y123H+Y147R+Q154R, N38G+V82T+Y123H+Y147R+Q154R, R23H+V82 S+Y123H+Y147R+Q154R, R21N+V82S+Y123H+Y147R+Q154R, V82S+Y123H+Y147R+Q154R+A158K, N72K+V82S+Y123H+D139L+Y147R+Q154R, E25F+V82S+Y123H+D139M+Y147R+Q154R, and M70V+V82S+M94V+Y123H+Y147R+Q154R.
[0366] In some embodiments, the adenosine deaminase comprises one or more of the following combinations of modifications: Q71M+V82S+Y123H+Y147R+Q154R, E25F+I76Y+V82S+Y123H+Y147R+Q154R, I76Y+V82T+Y123H+Y147R+Q154R, N38G+I76Y+V82S+Y123H+Y147R+Q154R, R23H+I76Y+V82S+Y123H+Y147R+Q154R, P54C+I76Y+V82S+Y123H+Y147R+Q154R, R21N+ I76Y+V82S+Y123H+Y147R+Q154R, I76Y+V82S+Y123H+D139M+Y147R+Q154 R, Y73S+I76Y+V82S+Y123H+Y147R+Q154R, E25F+I76Y+V82S+Y123H+Y147 R+Q154R, I76Y+V82T+Y123H+Y147R+Q154R, N38G+I76Y+V82S+Y123H+Y14 7R+Q154R, R23H+I76Y+V82S+Y123H+Y147R+Q154R, P54C+I76Y+V82S+Y12 3H+Y147R+Q154R, R21N+I76Y+V82S+Y123H+Y147R+Q154R, I76Y+V82S+Y 123H+D139M+Y147R+Q154R, Y73S+I76Y+V82S+Y123H+Y147R+Q154R, and V8 2S+Q154R, N72K_V82S+Y123H+Y147R+Q154R, Q71M_V82S+Y123H+Y147R+Q 154R, V82S+Y123H+T133K+Y147R+Q154R, V82S+Y123H+T133K+Y147R+Q15 4R+A158K, M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R, N72K_V82S+Y123H+Y147R+Q154R, Q71M_V82S+Y123H+Y147R+Q154R, M70V+V82S+M94V+Y123H+Y147R+Q154R, V82S+Y123H+T133K+Y147R+Q154R, V82S+Y123H+T133K+Y147R+Q154R+A158K, and M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R. In some embodiments, the adenosine deaminase is expressed as a monomer.In other embodiments, adenosine deaminase is expressed as a heterodimer. In some embodiments, the deaminase or other polypeptide sequence lacks methionine, for example, when included as a component of a fusion protein. This may result in a change in the numbering of positions. However, those skilled in the art will understand that such corresponding mutations refer to the same mutation (e.g., Y73S and Y72S, and D139M and D138M).
[0367] In some embodiments, the TadA*9 variant is a monomer. In some embodiments, the TadA*9 variant is a heterodimer with a wild-type TadA adenosine deaminase. In some embodiments, the TadA*9 variant is a heterodimer with another TadA variant (e.g., TadA*8, TadA*9). Additional details of the TadA*9 adenosine deaminase are described in International PCT Application No. PCT / 2020 / 049975, which is incorporated herein by reference in its entirety. In one embodiment, the fusion protein described herein comprises wild-type TadA, linked to an adenosine deaminase variant described herein (e.g., a TadA variant), and linked to a Cas9 nickase. In certain embodiments, the fusion protein comprises a single TadA variant domain (e.g., provided as a monomer). In other embodiments, the base editor comprises TadA*8 and TadA(wt) and can form a heterodimer.
[0368] In certain embodiments, the fusion protein comprises a single TadA variant domain (e.g., provided as a monomer). In some embodiments, the TadA variant is linked to a Cas9 nickase. In some embodiments, the fusion protein described herein comprises wild-type TadA (TadA(wt)) linked to a TadA variant as a heterodimer. In other embodiments, the fusion protein described herein comprises TadA*7.10 linked to a TadA variant as a heterodimer. In some embodiments, the fusion protein comprises a TadA variant monomer. In some embodiments, the fusion protein comprises a heterodimer of a TadA variant and TadA(wt). In some embodiments, the fusion protein comprises a heterodimer of a TadA variant and TadA*7.10. In some embodiments, the fusion protein comprises a heterodimer of two TadA variants. In some embodiments, the TadA variant is selected from Tables 5, 6, below, or any other TadA variant provided herein.
[0369] In some embodiments, the deaminase or other polypeptide sequence, for example when included as a component of a fusion protein, lacks a methionine. This may result in a change in the numbering of positions. However, one of skill in the art will understand that such corresponding mutations refer to the same mutation.
[0370] Any of the mutations provided herein, and any additional mutations (e.g., based on the ecTadA amino acid sequence), can be introduced into any other adenosine deaminase. Any of the mutations provided herein can be made individually or in any combination in the TadA reference sequence or in another adenosine deaminase (e.g., ecTadA).
[0371] Details of the A to G nucleobase editing protein are described in International PCT Application No. PCT / 2017 / 045381 (WO2018 / 027078) and Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature, 551, 464-471 (2017), the entire contents of which are incorporated herein by reference.
[0372] Use of nucleobase editors to target nucleotides in the G6PC gene The suitability of nucleobase editors to target nucleotides in G6PC genes is assessed as described herein.
[0373] The activity of the nucleobase editor is assessed as described herein, i.e., by sequencing the target gene to detect alterations in the target sequence. For Sanger sequencing, purified PCR amplicons are cloned into a plasmid backbone, transformed, miniprepped, and sequenced using a single primer. Sequencing may be performed using next-generation sequencing technology. When using next-generation sequencing, the amplicons may be 300-500 bp, with asymmetrically positioned intended cleavage sites. After PCR, next-generation sequencing adapters and barcodes (e.g., Illumina multiplex adapters and indexes) may be added to the ends of the amplicons, for example, for use in high-throughput sequencing (e.g., Illumina MiSeq).
[0374] In some embodiments, nucleobase editors are used to target a polynucleotide of interest. In one embodiment, a nucleobase editor as described herein is delivered to a cell (e.g., a hepatocyte) in conjunction with a guide RNA that is used to target a nucleic acid sequence, e.g., a G6PC polynucleotide carrying a GSD1a-associated mutation, thereby modifying the target gene, i.e., G6PC.
[0375] In some embodiments, the base editor is targeted by a guide RNA to introduce one or more edits into the sequence of a gene of interest (e.g., G6PC). In some embodiments, the one or more modifications are introduced into the glucose-6-phosphatase (G6PC) gene. In some embodiments, the one or more modifications are R83C. In some embodiments, the one or more modifications are Q347X. In some embodiments, the modifications are introduced into a representative Homo sapiens G6PC protein found under NCBI reference sequence number AAA16222.1. In some embodiments, the modifications are introduced into a representative Homo sapiens G6PC nucleic acid sequence found under GenBank reference sequence number U01120.1.
[0376] therapeutic use The NLS-gRNA described herein can be used in gene editing systems for various therapeutic applications. Thus, in some embodiments, a method of treating a disorder or disease in a subject in need of such treatment is provided, comprising administering to the subject an NLS-gRNA described herein in conjunction with a gene editing system. Various gene editing systems are known in the art, including, for example, CRISPR-Cas9, Cpf1, SaCas9, and Cas12. The NLS-gRNA described herein can be used with any gene editing system. For example, Cas proteins can be used in the treatment of Streptococcus, Campylobacter, Nitratifr. actor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnob acterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospira, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclo The bacteria are derived from organisms from genera including bacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Leptospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, Butyvibrio, Perigrinibacterium, Pareubacterium, Moraxella, Thiomicrospira, or Acidaminococcus.In certain embodiments, the Cpfl effector protein is selected from an organism from a genus selected from Eubacterium, Lachnospiraceae, Leptotrichia, Francisella, Methanomethyophilus, Porphyromonas, Prevotella, Leptospira, Butyvibrio, Perigrinibacterium, Pareubacterium, Moraxella, Thiomicrospira, or Acidaminococcus.
[0377] Non-limiting examples of Cas species include Streptococcus pyogenes, Streptococcus thermophiles, Sterptococcus aureas Neisseria meningitides, Treponema denticola, Francisella tularensis, Campylobacter jejuni, Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma Isolate RUG017, Veillonella parvula, Ezakiella peruensis, Lactobacillus fermentum strain AF15-40LB, and Peptoniphilus sp. Marseille-P3761.
[0378] In some embodiments, the NLS-gRNAs described herein can be used in conjunction with gene editing systems to treat various diseases and disorders, such as genetic disorders (e.g., monogenic diseases), diseases that can be treated by nuclease activity, and various cancers.
[0379] In some embodiments, the NLS-gRNA described herein may be used in conjunction with a gene editing system to edit a target nucleic acid to modify the target nucleic acid (e.g., by inserting, deleting, or mutating one or more nucleic acid residues). For example, in some embodiments, a CRISPR system is used in conjunction with the NLS-gRNA described herein to include an exogenous donor template nucleic acid (e.g., a DNA molecule or an RNA molecule) that includes a desired nucleic acid sequence. Upon resolution of the CRISPR system-induced cleavage event, the cell's molecular machinery will utilize the exogenous donor template nucleic acid in repairing and / or resolving the cleavage event. Alternatively, the cell's molecular machinery can utilize an endogenous template in repairing and / or resolving the cleavage event. In some embodiments, the NLS-gRNA described herein is used in conjunction with a gene editing system to alter the target nucleic acid, resulting in an insertion, deletion, and / or point mutation). In some embodiments, the insertion is a scarless insertion (i.e., insertion of an intended nucleic acid sequence into a target nucleic acid that does not result in additional unintended nucleic acid sequence upon resolution of the cleavage event). The donor template nucleic acid can be a double-stranded or single-stranded nucleic acid molecule (e.g., DNA or RNA).
[0380] In one aspect, the NLS-gRNAs described herein can be used in conjunction with gene editing systems to treat diseases caused by overexpression of RNA, toxic RNA, and / or mutant RNA (e.g., splicing defects or truncations).
[0381] In some embodiments, the NLS-gRNAs described herein can be used in conjunction with gene editing systems to target trans-acting mutations that affect RNA-dependent functions that cause various diseases.
[0382] In some embodiments, the NLS-gRNAs described herein can be used in conjunction with gene editing systems to target mutations that disrupt the cis-acting splicing code that can cause splicing defects and disease.
[0383] The NLS-gRNA described herein can be used in conjunction with gene editing systems for antiviral activity, particularly for RNA viruses.For example, target viral RNA by using suitable NLS-gRNA selected for targeting viral RNA sequence.
[0384] The NLS-gRNAs described herein can be used in conjunction with gene editing systems to treat cancer in a subject (e.g., a human subject). For example, they are found to induce cell death (e.g., via apoptosis) in cancer cells by targeting RNA molecules that are aberrant (e.g., contain point mutations or are alternatively spliced).
[0385] The NLS-gRNA described herein can be used in conjunction with a gene editing system to treat infectious diseases in a subject. For example, by targeting an RNA molecule expressed by an infectious agent (e.g., a bacterium, a virus, a parasite, or a protozoan) to target and induce cell death in infected somatic cells. The synthetic guide RNA described herein can be used in conjunction with a gene editing system to treat diseases in which an intracellular infectious agent infects the cells of a host subject.
[0386] In applications where it is desired to insert a polynucleotide sequence into a target DNA sequence, a polynucleotide comprising the donor sequence to be inserted is also provided to the cell. By "donor sequence" or "donor polynucleotide" is meant a nucleic acid sequence to be inserted at the cleavage site induced by the site-directed modifying polypeptide. The donor polynucleotide contains sufficient homology to the genomic sequence at the cleavage site, e.g., 70%, 80%, 85%, 90%, 95%, or 100% homology to the nucleotide sequence adjacent to the cleavage site, e.g., within about 50 bases or less, e.g., within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases, or immediately adjacent to the cleavage site, to support homology-directed repair between the genomic sequence with which it has homology. About 25, 50, 100, or 200 nucleotides, or more than 200 nucleotides of sequence homology between the donor and the genomic sequence (or any integer value between 10 and 200 nucleotides or more) will support homology-directed repair. The donor sequence can be of any length, e.g., 10 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 250 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 5000 nucleotides or more, etc.
[0387] The donor sequence is typically not identical to the genomic sequence it replaces. Rather, the donor sequence may contain at least one or more single base changes, insertions, deletions, inversions, or translocations with respect to the genomic sequence, so long as there is sufficient homology to support homology-directed repair. In some embodiments, the donor sequence contains two homologous regions and flanking non-homologous sequences, such that homology-directed repair between the target DNA region and the two flanking sequences results in the insertion of the non-homologous sequence into the target region. The donor sequence may also contain a vector backbone that contains sequences that are not homologous to the DNA region of interest and are not intended for insertion into the DNA region of interest. In general, the homologous region(s) of the donor sequence will have at least 50% sequence identity with the genomic sequence with which recombination is desired. In certain embodiments, there is 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.9% sequence identity. Depending on the length of the donor polynucleotide, there can be any value of sequence identity between 1% and 100%.
[0388] The donor sequence may contain certain sequence differences compared to the genomic sequence, such as restriction sites, nucleotide polymorphisms, selectable markers (e.g., drug resistance genes, fluorescent proteins, enzymes, etc.), which may be used to evaluate the successful insertion of the donor sequence at the cleavage site, or in some cases, for other purposes (e.g., to indicate expression at the targeted genomic locus). In some cases, when located in a coding region, such nucleotide sequence differences will not change the amino acid sequence or will make silent amino acid changes (i.e., changes that do not affect the structure or function of the protein). Alternatively, these sequence differences may include adjacent recombination sequences, such as FLP, loxP sequences, that can be activated later to remove the marker sequence.
[0389] The donor sequence may be provided to the cell as single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA. It may be introduced into the cell in linear or circular form. If introduced in linear form, the ends of the donor sequence may be protected (e.g., from exonucleolytic degradation) by methods known to those skilled in the art. For example, one or more dideoxynucleotide residues are added to the 3' end of the linear molecule, and / or a self-complementary oligonucleotide is ligated to one or both ends. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino group(s) and the use of modified internucleotide linkages, such as, for example, phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues. As an alternative to protecting the ends of linear donor sequences, additional lengths of sequence may be included outside the regions of homology that may be degraded without affecting recombination. The donor sequence may be introduced into a cell as part of a vector molecule having additional sequences such as, for example, an origin of replication, a promoter, and genes encoding antibiotic resistance. Additionally, the donor sequence may be introduced as naked nucleic acid, as nucleic acid complexed with an agent such as a liposome or poloxamer, or delivered by a virus (e.g., adenovirus, AAV), as described above for the nucleic acid encoding the DNA-targeting RNA and / or the site-directed modifying polypeptide and / or donor polynucleotide.
[0390] According to the methods described above, the DNA region of interest may be cleaved and modified, i.e., "genetically modified", ex vivo. In some embodiments, as in the case where a selectable marker is inserted into the DNA region of interest, the population of cells may be enriched for those containing the genetic modification by separating the genetically modified cells from the remainder of the population. Prior to enrichment, the "genetically modified" cells may constitute only about 1% or more (e.g., 2% or more, 3% or more, 4% or more, 5% or more, 6% or more, 7% or more, 8% or more, 9% or more, 10% or more, 15% or more, or 20% or more) of the cell population. Separation of the "genetically modified" cells may be achieved by any convenient separation technique appropriate for the selectable marker used. For example, if a fluorescent marker is inserted, the cells may be separated by fluorescence activated cell sorting, and if a cell surface marker is inserted, the cells may be separated from the heterogeneous population by affinity separation techniques, such as magnetic separation, affinity chromatography, "panning" with affinity reagents bound to a solid matrix, or other convenient techniques. Techniques that provide accurate separation include fluorescence-activated cell sorters, which can have various degrees of sophistication, such as multiple color channels, low-angle and obtuse-angle light scattering detection channels, impedance channels, etc. Cells can be selected against dead cells by using a dye that is associated with dead cells (e.g., propidium iodide). Any technique that is not overly detrimental to the viability of the genetically modified cells can be used. A cell composition highly enriched for cells containing modified DNA can be achieved in this manner. By "highly enriched" it is meant that the genetically modified cells are 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, e.g., about 95% or more, or 98% or more of the cell composition. In other words, the composition can be a substantially pure composition of genetically modified cells.
[0391] The genetically modified cells produced by the methods described herein can be used immediately. Alternatively, the cells can be frozen at liquid nitrogen temperatures, stored for extended periods, and thawed and reused. In such cases, the cells are typically frozen in 10% dimethylsulfoxide (DMSO), 50% serum, 40% buffered medium, or some other such solution commonly used in the art to preserve cells at such freezing temperatures, and thawed in a manner commonly known in the art to thaw frozen cultured cells.
[0392] The genetically modified cells may be cultured in vitro under a variety of culture conditions. The cells may be grown in culture, i.e., grown under conditions that promote the growth of the cells. The medium may be liquid or semi-solid, containing, for example, agar, methylcellulose, etc. The cell population is usually cultured in a medium supplemented with fetal bovine serum (about 5-10%),
[0393] They may be suspended in a suitable nutrient medium, such as Iscove's modified DMEM or RPMI 1640, supplemented with L-glutamine, thiols, particularly 2-mercaptoethanol, and antibiotics, such as penicillin and streptomycin. The culture may contain growth factors to which the regulatory T cells are responsive. Growth factors, as defined herein, are molecules that can promote cell survival, proliferation and / or differentiation, either in culture or in intact tissue, through specific effects on transmembrane receptors. Growth factors include polypeptide and non-polypeptide factors.
[0394] Such genetically modified cells may be implanted into a subject, for example, to treat a disease or as an antiviral, antipathogenic, or anticancer therapeutic, for purposes such as gene therapy, for the production of genetically modified organisms in agriculture, or for biological research. The subject may be a neonate, juvenile, or adult. Of particular interest are mammalian subjects. Mammalian species that may be treated with the present method include dogs and cats, horses, cows, sheep, etc., as well as primates, especially humans. Animal models, especially small mammals (e.g., mice, rats, guinea pigs, hamsters, lagomorphs (e.g., rabbits), etc.), may be used for experimental investigations.
[0395] The cells may be provided to a subject alone or, for example, with a suitable substrate or matrix to support the growth and / or organization of the cells within the tissue into which they are implanted. Typically, at least 1×10 3 cells, e.g., 5 x 10 3 cells, 1 x 10 4 cells, 5 x 10 4 cells, 1 x 10 5 cells, 1 x 10 6 One or more cells are administered. The cells may be introduced into the subject via any of the following routes: parenteral, subcutaneous, intravenous, intracranial, intraspinal, intraocular, or into the spinal fluid. The cells may be introduced by injection, catheter, etc. The cells may also be introduced into an embryo (e.g., a blastocyst) for the purpose of generating a transgenic animal (e.g., a transgenic mouse).
[0396] The number of times that the treatment is administered to a subject may vary. Introducing genetically modified cells into a subject may be a one-time event, but in certain circumstances, such treatment may induce improvement over a limited period of time and require a series of ongoing repeated treatments. In other circumstances, multiple administrations of genetically modified cells may be required before an effect is observed. The exact protocol depends on the disease or condition, the stage of the disease, and the parameters of the individual subject being treated.
[0397] In other aspects of the invention, DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides are used to modify cellular DNA in vivo, for example to treat disease or as antiviral, antipathogenic or anticancer therapeutics, for purposes such as gene therapy, for the production of genetically modified organisms in agriculture, or for biological research. In these in vivo embodiments, the DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides are administered directly to an individual. The DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides may be administered by any of several well-known methods in the art for the administration of peptides, small molecules and nucleic acids to a subject. The DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides may be incorporated into various formulations. More specifically, the DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides of the invention may be formulated into a pharmaceutical composition by combination with a suitable pharma-ceutically acceptable carrier or diluent.
[0398] A pharmaceutical preparation is a composition that includes one or more DNA-targeting RNAs and / or site-directed modifying polypeptides and / or donor polynucleotides present in a pharmaceutically acceptable vehicle. A "pharmaceutically acceptable vehicle" can be a vehicle approved by a federal or state government regulatory agency or listed in the United States.
[0399] Pharmacopoeias or other generally accepted pharmacopoeias for use in mammals, such as humans. The term "vehicle" refers to a diluent, adjuvant, excipient, or carrier in which the compound of the present invention is formulated for administration to a mammal. Such pharmaceutical vehicles can be lipids, such as liposomes, e.g., liposomal dendrimers; water, oils, including those of petroleum, animal, vegetable, or synthetic origin, such as peanut oil, soybean oil, mineral oil, sesame oil, and the like, liquids, such as saline; gum acacia, gelatin, starch paste, talc, keratin, colloidal silica, urea, and the like. In addition, auxiliary agents, stabilizers, thickeners, lubricants, and colorants may be used. The pharmaceutical compositions can be formulated into solid, semi-solid, liquid, or gaseous form of preparations, such as tablets, capsules, powders, granules, ointments, solutions, suppositories, injections, inhalants, gels, microparticles, and aerosols. Thus, administration of DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides can be accomplished in a variety of ways, including oral, buccal, rectal, parenteral, intraperitoneal, intradermal, transdermal, intratracheal, intraocular, etc. The active agent may be systemic following administration, or may be localized by use of topical administration, intramural administration, or the use of an implant that acts to retain an active dose at the site of implantation. The active agent may be formulated for immediate activity or for sustained release.
[0400] For some conditions, particularly those of the central nervous system, it may be necessary to formulate drugs to cross the blood-brain barrier (BBB). One strategy for drug delivery through the blood-brain barrier (BBB) involves the disruption of the BBB, either by osmotic means such as mannitol or leukotrienes, or biochemically by the use of vasoactive substances such as bradykinin. The possibility of using BBB opening to target specific drugs to brain tumors is also an option. BBB disrupting agents may be co-administered with the therapeutic compositions of the invention when the compositions are administered by intravascular injection. Other strategies for crossing the BBB may involve the use of endogenous transporter systems, including caveolin 1-mediated transcytosis, carrier-mediated transporters such as glucose and amino acid carriers, receptor-mediated transcytosis of insulin or transferrin, and active efflux transporters such as p-glycoprotein. Active transport moieties may also be conjugated to therapeutic compounds for use in the invention to facilitate transport across the endothelial wall of blood vessels.
[0401] Alternatively, drug delivery of therapeutic agents behind the BBB may be by localized delivery, for example, intrathecal delivery.
[0402] Typically, an effective amount of DNA-targeting RNA and / or site-directed modified polypeptide and / or donor polynucleotide is provided. As discussed above with respect to ex vivo methods, an effective amount or effective dose of DNA-targeting RNA and / or site-directed modified polypeptide and / or donor polynucleotide is an amount that induces a two-fold or greater increase in the amount of recombination observed between two homologous sequences compared to cells contacted with a negative control, e.g., an empty vector or an unrelated polypeptide. The amount of recombination can be measured by any convenient method, e.g., as described above and known in the art. Calculation of the effective amount or effective dose of DNA-targeting RNA and / or site-directed modified polypeptide and / or donor polynucleotide to be administered is within the skill of, and routine for, the skilled artisan. The final amount administered will depend on the route of administration and the nature of the disorder or condition being treated. In some embodiments, an exemplary dose of about 0.01 to 1 mpk is used.
[0403] The effective amount given to a particular patient will depend on a variety of factors, some of which will vary from patient to patient. A competent clinician will be able to determine the effective amount of therapeutic agent to administer to a patient to halt or reverse the progression of a disease state, if necessary. Using LD50 animal data, and other available information about the agent, the clinician can determine the maximum safe dose for an individual, depending on the route of administration. For example, a dose administered intravenously may be higher than a dose administered intrathecally, given the larger volume of fluid into which the therapeutic composition is administered. Similarly, compositions that are rapidly cleared from the body may be administered at higher doses or in repeated doses to maintain therapeutic concentrations. Using routine techniques, a competent clinician will be able to optimize the dosage of a particular therapeutic agent during the course of routine clinical trials.
[0404] For inclusion in the medicament, the DNA-targeting RNA and / or the site-directed modifying polypeptide and / or the donor polynucleotide may be obtained from a suitable commercial source. As a general proposition, the total pharmacologic effective amount of the DNA-targeting RNA and / or the site-directed modifying polypeptide and / or the donor polynucleotide administered parenterally per dose is in a range that can be determined by a dose-response curve.
[0405] Therapies based on DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides, i.e., preparations of DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides used for therapeutic administration, need to be sterile. Sterility is easily achieved by filtration through a sterile filtration membrane (e.g., 0.2 μm membrane). Therapeutic compositions are generally placed in containers with a sterile access port, such as intravenous solution bags or vials with a stopper that can be pierced by a hypodermic needle. Therapies based on DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides can be stored in unit-dose or multi-dose containers, such as sealed ampoules or vials, as aqueous solutions or as lyophilized formulations for reconstitution. As an example of a lyophilized formulation, a 10 mL vial is filled with 5 ml of a sterile-filtered 1% (w / v) aqueous solution of the compound, and the resulting mixture is lyophilized. Infusion solutions are prepared by reconstituting the lyophilized compound using bacteriostatic water for injection.
[0406] The pharmaceutical composition may contain a pharma- ceutically acceptable non-toxic carrier of a diluent, which is defined as a vehicle commonly used to formulate pharmaceutical compositions for animal or human administration, depending on the desired formulation. The diluent is selected so as not to affect the biological activity of the combination. Examples of such diluents are distilled water, buffered water, physiological saline, PBS, Ringer's solution, dextrose solution, and Hank's solution. In addition, the pharmaceutical composition or formulation may contain other carriers, adjuvants, or non-toxic, non-therapeutic, non-immunogenic stabilizers, excipients, and the like. The composition may also contain additional substances that approximate physiological conditions, such as pH adjusting and buffering agents, toxicity adjusting agents, wetting agents, and detergents.
[0407] The composition may also include any of a variety of stabilizing agents, such as, for example, antioxidants. When the pharmaceutical composition includes a polypeptide, the polypeptide may be complexed with a variety of well-known compounds that enhance the in vivo stability of the polypeptide or otherwise enhance its pharmacological properties (e.g., increase the half-life of the polypeptide, reduce its toxicity, enhance solubility or uptake). Examples of such modifying or complexing agents include sulfates, gluconates, citrates, and phosphates. The nucleic acids or polypeptides of the composition may also be complexed with molecules that enhance their in vivo attributes. Such molecules include, for example, carbohydrates, polyamines, amino acids, other peptides, ions (e.g., sodium, potassium, calcium, magnesium, manganese), and lipids.
[0408] The pharmaceutical composition can be administered for prophylactic and / or therapeutic treatment. The toxicity and therapeutic efficacy of the active ingredient can be determined according to standard pharmaceutical procedures in cell cultures and / or experimental animals, including, for example, determining LD50 (the dose lethal to 50% of the population) and ED50 (the dose therapeutically effective for 50% of the population). The dose ratio between the toxic effect and the therapeutic effect is the therapeutic index, which can be expressed as the ratio LD50 / ED50. Therapies that exhibit large therapeutic indices are preferred.
[0409] The data obtained from cell culture and / or animal studies can be used in formulating a range of dosages for humans.The dosage of active ingredient is typically within the range of circulating concentrations that include the ED50 with low toxicity.Dosage can vary within this range depending on the dosage form used and the route of administration utilized.
[0410] The components used to formulate pharmaceutical compositions are preferably of high purity and substantially free of potentially harmful contaminants (e.g., at least National Food (NF) grade, generally at least analytical grade, and more typically at least pharmaceutical grade). Moreover, compositions intended for in vivo use are usually sterile. To the extent that a given compound must be synthesized prior to use, the resulting product is typically substantially free of any potentially toxic agents, particularly any endotoxins, that may be present during the synthesis or purification process. Compositions for parental administration are also sterile, substantially isotonic, and made under GMP conditions.
[0411] delivery system The NLS-gRNA described herein, along with the desired gene editing system components, can be delivered to the cells of interest by various delivery systems, such as vectors, carriers, e.g., lipid nanoparticles.
[0412] The NLS-gRNA described herein can be delivered by nanoparticles, which can be organic or inorganic. Nanoparticles are well known in the art. Any suitable nanoparticle design can be used to deliver the components of the genome editing system, or the nucleic acids encoding such components. For example, organic (e.g., lipid and / or polymer) nanoparticles can be suitable for use as delivery vehicles in certain embodiments of the present disclosure. Exemplary lipids for use in nanoparticle formulations and / or gene transfer are shown in Table 2 (below). [Table 3-1] [Table 3-2]
[0413] Table 3 lists exemplary polymers for use in gene transfer and / or nanoparticle formulations. [Table 4-1] [Table 4-2]
[0414] Table 4 summarizes the delivery methods for the Cas9-encoding polynucleotides described herein. [Table 5-1] [Table 5-2]
[0415] In another aspect, delivery of the genome editing system comprising the NLS-gRNA described herein can be achieved by delivering a ribonucleoprotein (RNP) to a cell. The RNP comprises a nucleic acid binding protein (e.g., Cas9) in a complex with the target gRNA. The RNP can be delivered to a cell using known methods such as electroporation, nucleofection, or cationic lipid-mediated methods, for example, as reported by Zuris, JA et al., 2015, Nat. Biotechnology, 33(1):73-80. RNPs are advantageous for use in CRISPR base editing systems, particularly in cells that are difficult to transfect (e.g., primary cells). In addition, RNPs can also alleviate difficulties that may arise with protein expression in cells, especially when eukaryotic promoters (e.g., CMV or EF1A) that may be used in CRISPR plasmids are not expressed sufficiently. Advantageously, the use of RNPs does not require the delivery of foreign DNA into the cell. Furthermore, the use of RNPs may limit off-target effects, since the RNPs containing the nucleic acid binding protein and gRNA complex are degraded over time. In a manner similar to that of plasmid-based techniques, RNPs can be used to deliver binding proteins (e.g., Cas9 variants) and induce homology-directed repair (HDR).
[0416] The promoter used to drive the CRISPR system (e.g., including the synthetic gRNAs described herein) may include AAV ITRs. This may be advantageous to eliminate the need for additional promoter elements that may occupy space in the vector. The additional space freed up can be used to drive expression of additional elements (e.g., guide nucleic acids or selectable markers). Since ITR activity is relatively weak, it can be used to reduce potential toxicity due to overexpression of selected nucleases.
[0417] Any suitable promoter may be used to drive the expression of Cas9 and, if applicable, guide nucleic acid. For ubiquitous expression, promoters that may be used include CMV, CAG, CBh, PGK, SV40, ferritin heavy or light chain, etc. For expression in brain or other CNS cells, suitable promoters may include synapsin I for all neurons, CaMKIIα for excitatory neurons, GAD67 or GAD65 or VGAT for GABAergic neurons, etc. For expression in hepatocytes, suitable promoters may include albumin promoter. For expression in lung cells, suitable promoters may include SP-B. For endothelial cells, suitable promoters may include ICAM. For hematopoietic cells, suitable promoters may include IFNβ or CD45. For osteoblasts, suitable promoters may include OG-2.
[0418] In some cases, separate promoters drive expression of the base editor and a compatible guide nucleic acid within the same nucleic acid molecule, for example, a vector or viral vector can include a first promoter operably linked to a nucleic acid encoding a base editor and a second promoter operably linked to a guide nucleic acid.
[0419] Promoters used to drive expression of the guide nucleic acid can include PolIII promoters such as U6 or H1. To express the gRNA adeno-associated virus (AAV), a PolII promoter and an intron cassette is used.
[0420] Cas9 can be delivered using adeno-associated virus (AAV), lentivirus, adenovirus, or other plasmid or viral vector types, particularly using formulations and dosages from US Patent No. 8,454,972 (formulations, dosages for adenovirus), US Patent No. 8,404,658 (formulations, dosages for AAV), and US Patent No. 5,846,946 (formulations, dosages for DNA plasmids), as well as clinical trials and publications related to clinical trials involving lentivirus, AAV, and adenovirus. For example, in the case of AAV, the administration route, formulation, and dosage can be similar to US Patent No. 8,454,972 and clinical trials involving AAV. In the case of adenovirus, the administration route, formulation, and dosage can be similar to US Patent No. 8,404,658 and clinical trials involving adenovirus. In the case of plasmid delivery, the administration route, formulation, and dosage can be similar to US Patent No. 5,846,946 and clinical trials involving plasmids. Dosages may be based on or estimated for an average 70 kg individual (e.g., adult male) and may be adjusted for patients, subjects, mammals of different weights and species. Frequency of administration is within the scope of a physician or veterinarian (e.g., physician, veterinarian) depending on the usual factors including age, sex, general health, other conditions of the patient or subject, and the specific condition or symptom being addressed. The viral vector may be injected into the tissue of interest. In the case of cell-type specific base editing, expression of the base editor and optional guide nucleic acid may be driven by a cell-type specific promoter.
[0421] For in vivo delivery, AAV may have advantages over other viral vectors.In some cases, AAV allows low toxicity, which may be due to the purification method that does not require ultracentrifugation of cell particles that may activate immune response.In some cases, AAV is not integrated into host genome, so it is less likely to cause insertional mutagenesis.
[0422] AAV has a packaging limit of 4.5 or 4.75Kb. Constructs that exceed 4.5Kb or 4.75Kb can result in a significant decrease in virus production. For example, SpCas9 is quite large, and the gene itself exceeds 4.1Kb, making it difficult to package into AAV. Therefore, embodiments of the present disclosure include the use of the disclosed Cas9, which is shorter in length than conventional Cas9.
[0423] AAV can be AAV1, AAV2, AAV5, or any combination thereof. The type of AAV can be selected according to the cells to be targeted, for example, AAV serotype 1, 2, 5, or hybrid capsid AAV1, AAV2, AAV5, or any combination thereof for targeting brain or neuronal cells, and AAV4 can be selected when targeting cardiac tissue. AAV8 is useful for delivery to the liver. A tabulation of specific AAV serotypes for these cells can be found in Grimm, D. et al, J. Virol. 82:5887-5911 (2008).
[0424] Lentiviruses are complex retroviruses that have the ability to infect and express their genes in both dividing and post-mitotic cells. The most commonly known lentivirus is the human immunodeficiency virus (HIV), which uses the envelope glycoproteins of other viruses to target a broad range of cell types.
[0425] Lentivirus can be prepared as follows: After cloning pCasES10 (containing the lentiviral transcription plasmid backbone), low passage (p=5) HEK293FT were seeded to 50% confluence in T-75 flasks in DMEM with 10% fetal bovine serum and no antibiotics the day before transfection. After 20 hours, the medium was replaced with OptiMEM (serum-free) medium, and transfection was performed 4 hours later. Cells are transfected with 10 μg of lentiviral transfer plasmid (pCasES10) and the following packaging plasmids: 5 μg of pMD2.G (VSV-g pseudotyped), and 7.5 μg of psPAX2 (gag / pol / rev / tat). Transfection can be performed in 4 mL of OptiMEM containing cationic lipid delivery agent (50 μl of Lipofectamine 2000 and 100 ul of Plus reagent). After 6 hours, the medium is replaced with antibiotic-free DMEM containing 10% fetal bovine serum. Although these methods use serum in cell culture, serum-free methods are preferred.
[0426] Lentivirus can be purified as follows: Viral supernatant is harvested after 48 hours. The supernatant is first removed from debris and filtered through a 0.45 μm low protein binding (PVDF) filter. It is then spun in an ultracentrifuge at 24,000 rpm for 2 hours. The viral pellet is resuspended in 50 μl of DMEM overnight at 4° C. They are then aliquoted and immediately frozen at −80° C.
[0427] In another embodiment, a minimal non-primate lentiviral vector based on the Equine Infectious Anemia Virus (EIAV) is also contemplated. In another embodiment, RetinoStat® is an Equine Infectious Anemia Virus-based lentiviral gene therapy vector that expresses the angiogenesis inhibitory proteins endostatin and angiostatin, which is contemplated to be delivered by subretinal injection. In another embodiment, the use of self-inactivating lentiviral vectors is contemplated.
[0428] Any RNA of the system, such as NLS-gRNA or Cas9-encoding mRNA, can be delivered in the form of RNA. Cas9-encoding mRNA can be generated using in vitro transcription. For example, Cas9 mRNA can be synthesized using a PCR cassette that includes the following elements: T7 promoter, optional kozak sequence (GCCACC), nuclease sequence, and 3'UTR, such as 3'UTR from β-globin-polyA tail. The cassette can be used for transcription by T7 polymerase. Guide polynucleotides (e.g., gRNAs) can also be transcribed using in vitro transcription from a cassette that includes a T7 promoter followed by the sequence "GG" and a guide polynucleotide sequence.
[0429] To enhance expression and reduce potential toxicity, the Cas9 sequence and / or guide nucleic acid can be modified to include one or more modified nucleosides, for example, using pseudo-U or 5-methyl-C.
[0430] The present disclosure encompasses, in some embodiments, a method of modifying a cell or organism. The cell may be a prokaryotic or eukaryotic cell. The cell may be a mammalian cell. Many of the mammalian cells are non-human primate, bovine, porcine, rodent, or murine cells. The modifications introduced into the cell by the base editors, compositions, and methods of the present disclosure can result in the cell and the cell's progeny being altered for improved production of a biological product, such as an antibody, starch, alcohol, or other desired cellular output. The modifications introduced into the cell by the methods of the present disclosure can result in the cell and the cell's progeny containing a modification that alters the biological product produced.
[0431] The system can include one or more different vectors. In one embodiment, Cas9 is codon-optimized for expression in a desired cell type, preferably a eukaryotic cell, preferably a mammalian cell or a human cell.
[0432] Generally, "codon optimization" refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by substituting at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of the native sequence with a codon that is more or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Different species show a certain bias for a particular codon of a particular amino acid. Codon bias (differences in codon usage between organisms) is often thought to correlate with the efficiency of translation of messenger RNA (mRNA), which in turn depends, among other things, on the properties of the codon being translated and the availability of a particular transfer RNA (tRNA) molecule. The dominance of tRNAs selected in a cell is generally a reflection of the codons that are most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, at the "Codon Usage Database" available at www.kazusa.orjp / codon / (accessed July 9, 2002), and these tables can be adapted in several ways. See Nakamura, Y., et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28:292 (2000). Computer algorithms are also available that use codons to optimize a particular sequence for expression in a particular host cell, for example, Gene Forge (Aptagen, Jacobus, Pa.). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons, or all codons) in a sequence encoding an engineered nuclease correspond to the most frequently used codon for a particular amino acid.
[0433] Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells, which package adenovirus, and psi.2 or PA317 cells, which package retrovirus. Viral vectors used in gene therapy are usually generated by producing cell lines that package nucleic acid vectors into viral particles. The vectors typically contain minimal viral sequences necessary for packaging and subsequent integration into the host, with other viral sequences being replaced by an expression cassette for the polynucleotide(s) to be expressed. Missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically only have ITR sequences from the AAV genome necessary for packaging and integration into the host genome. Viral DNA can be packaged into a cell line that contains a helper plasmid that encodes other AAV genes (i.e., rep and cap) but lacks ITR sequences. The cell line can be infected with adenovirus as a helper. The helper virus can facilitate replication of the AAV vector and expression of the AAV genes from the helper plasmid. In some cases, the helper plasmid is not packaged in significant amounts due to the lack of ITR sequences. Adenovirus contamination can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV.
[0434] Pharmaceutical Compositions Another aspect of the present disclosure relates to a pharmaceutical composition comprising a gene editing system (e.g., comprising an NLS-gRNA described herein). The term "pharmaceutical composition" as used herein refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharma- ceutical acceptable carrier. In some embodiments, the pharmaceutical composition comprises an additional agent (e.g., for specific delivery, to increase half-life, or other therapeutic compounds).
[0435] As used herein, the term "pharmaceutically acceptable carrier" means a pharma- ceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, magnesium talc, calcium or zinc stearate, or steric acid), or solvent encapsulant, involved in carrying or transporting a compound from one site in the body (e.g., delivery site) to another (e.g., an organ, tissue, or body part). A pharma- ceutically acceptable carrier is "acceptable" in the sense of being compatible with the other ingredients of the formulation and not deleterious to the tissues of the subject (e.g., physiological compatibility, sterility, physiological pH, etc.).
[0436] Some non-limiting examples of materials that can serve as pharma- ceutically acceptable carriers include: (1) sugars, such as lactose, glucose, and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose and its derivatives, such as sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants, such as magnesium stearate, sodium lauryl sulfate, and talc; (8) excipients, such as cocoa butter and suppository wax; and (9) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil. , (10) glycols such as propylene glycol, (11) polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol (PEG), (12) esters such as ethyl oleate and ethyl laurate, (13) agar, (14) buffers such as magnesium hydroxide and aluminum hydroxide, (15) alginic acid, (16) pyrogen-free water, (17) isotonic saline, (18) Ringer's solution, (19) ethyl alcohol, (20) pH buffer solutions, (21) polyesters, polycarbonates, and / or polyanhydrides, (22) swelling agents such as polypeptides and amino acids, (23) serum alcohols such as ethanol, and (23) other non-toxic compatible substances used in pharmaceutical formulations. Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, fragrances, preservatives, and antioxidants may also be present in the formulation. The terms "excipient," "carrier," "pharmaceutical acceptable carrier," "vehicle," and the like are used interchangeably herein.
[0437] The pharmaceutical composition may contain one or more pH buffer compounds to maintain the pH of the formulation at a predetermined level that reflects physiological pH, for example, in the range of about 5.0 to about 8.0. The pH buffer compound used in the aqueous liquid formulation may be an amino acid, or a mixture of amino acids (e.g., histidine, or a mixture of amino acids such as histidine and glycine). Alternatively, the pH buffer compound is preferably an agent that maintains the pH of the formulation at a predetermined level (e.g., in the range of about 5.0 to about 8.0) and does not chelate calcium ions. Illustrative examples of such pH buffer compounds include, but are not limited to, imidazole and acetate ions. The pH buffer compound may be present in any amount suitable for maintaining the pH of the formulation at a predetermined level.
[0438] The pharmaceutical composition may also include one or more osmotic modifiers, i.e., compounds that adjust the osmotic properties (e.g., tonicity, osmolality, and / or osmolarity) of the formulation to a level that is acceptable to the bloodstream and blood cells of the recipient individual. The osmotic modifier may be an agent that does not chelate calcium ions. The osmotic modifier may be any compound known or available to one of skill in the art that adjusts the osmotic properties of the formulation. One of skill in the art can empirically determine the suitability of a given osmotic modifier for use in the formulations of the invention. Illustrative examples of suitable types of osmotic modifiers include, but are not limited to, salts (e.g., sodium chloride and sodium acetate), sugars (e.g., sucrose, dextrose, and mannitol), amino acids (e.g., glycine), and mixtures of one or more of these agents and / or types of agents. The osmotic modifier(s) may be present in any concentration sufficient to adjust the osmotic properties of the formulation.
[0439] In some embodiments, the pharmaceutical composition is formulated for delivery to a subject, for example, for gene editing. Suitable routes for administering the pharmaceutical compositions described herein include, but are not limited to, topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, intradental, intracochlear, intratympanic, intravisceral, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseous, periocular, intratumoral, intracerebral, and intraventricular administration.
[0440] In some embodiments, the pharmaceutical compositions described herein are administered locally to the disease site. In some embodiments, the pharmaceutical compositions described herein are administered to the subject by injection, by catheter, by suppository, or by implant, the implant being a porous, non-porous, or gel-like material, including membranes such as sialastic membranes or fibers.
[0441] In other embodiments, the pharmaceutical compositions described herein are delivered in a controlled release system.In one embodiment, pumps can be used (see, for example, Langer, 1990, Science 249:1527-1533; Sefton, 1989, CRC Crit. Ref. Biomed. Eng. 14:201; Buchwald et al., 1980, Surgery 88:507; Saudek et al., 1989, N. Engl. J. Med. 321:574).In another embodiment, polymeric materials can be used. (See, e.g., Medical Applications of Controlled Release (Langer and Wise eds., CRC Press, Boca Raton, Fla., 1974); Controlled Drug Bioavailability, Drug Product Design and Performance (Smolen and Ball eds., Wiley, New York, 1984); Ranger and Peppas, 1983, Macromol. Sci. Rev. Macromol. Chem. 23:61; see also, Levy et al., 1985, Science 228:190; During et al., 1989, Ann. Neurol. 25:351; Howard et al., 1989, J. Neurosurg. 71:105.) Other controlled release systems are discussed, e.g., in Langer, supra.
[0442] In some embodiments, the pharmaceutical composition is formulated according to routine procedures as a composition suitable for intravenous or subcutaneous administration to a subject (e.g., a human). In some embodiments, the pharmaceutical composition for administration by injection is a solution in sterile isotonic use as a solubilizing agent, and a local anesthetic, such as lignocaine, to ease pain at the injection site. Generally, the ingredients are supplied either separately or mixed together in a unit dosage form, for example as a dry lyophilized powder or water-free concentrate in a sealed container, such as an ampoule or sachet indicating the quantity of active agent. When the drug product is administered by injection, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. When the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration.
[0443] Pharmaceutical compositions for systemic administration may be liquid, such as sterile saline, lactated Ringer's solution, or Hank's solution. In addition, pharmaceutical compositions may be in solid form and redissolved or suspended immediately prior to use. Lyophilized forms are also contemplated. Pharmaceutical compositions may be contained within lipid particles or vesicles (e.g., liposomes or microcrystals also suitable for parenteral administration). The particles may be of any suitable structure, such as unilamellar or plurilamellar, so long as the composition is contained therein. Compounds may be encapsulated in "stabilized plasmid lipid particles" (SPLPs) containing the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipids, and stabilized by polyethylene glycol (PEG) coating (Zhang YPet ah, Gene Ther. 1999, 6:1438-47). For such particles and vesicles, positively charged lipids such as N-[1-(2,3-dioleoyloxy)propyl]-N,N,N-trimethyl-ammonium methylsulfate or "DOTAP" are particularly preferred.The preparation of such lipid particles is well known.See, for example, U.S. Patent Nos. 4,880,635, 4,906,477, 4,911,928, 4,917,951, 4,920,016, and 4,921,757 (each of which is incorporated herein by reference).
[0444] The pharmaceutical compositions described herein may be administered or packaged, for example, as unit doses. The term "unit dose", when used in reference to the pharmaceutical compositions of the present disclosure, refers to physically discrete units suitable as unitary dosages for subjects, each unit containing a predetermined amount of active material calculated to produce the desired therapeutic effect in association with the required diluent (i.e., carrier or vehicle).
[0445] Further, the pharmaceutical composition may be provided as a pharmaceutical kit, comprising (a) a container containing a compound of the invention in lyophilized form, and (b) a second container containing a pharma- ceutically acceptable diluent (e.g., sterile, for use in reconstituting or diluting the lyophilized compound of the invention). Optionally, a notice in a form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceutical or biological products may be associated with such container(s), the notice reflecting approval by the agency of the manufacture, use, or sale for human administration.
[0446] Another aspect includes an article of manufacture that includes materials useful for treating the diseases described above. In some embodiments, the article of manufacture includes a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers can be formed from a variety of materials, such as glass or plastic. In some embodiments, the container holds a composition effective for treating a disease described herein and can have a sterile access port. For example, the container can be an intravenous infusion bag or a vial with a stopper pierceable by a hypodermic needle. The active agent in the composition is a compound of the present invention. In some embodiments, a label on or associated with the container indicates that the composition is used to treat a selected disease. The article of manufacture can further include a second container that includes a pharma- ceutically acceptable buffer, such as phosphate buffered saline, Ringer's solution, or dextrose solution. It can further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, syringes, and package inserts containing instructions for use.
[0447] In some embodiments, the CRISPR system (e.g., including Cas9 as described herein) is provided as part of a pharmaceutical composition. In some embodiments, the pharmaceutical composition includes any of the fusion proteins provided herein (e.g., including a nucleobase editor as described herein, including LubCas9). In some embodiments, the pharmaceutical composition includes any of the complexes provided herein. In some embodiments, the pharmaceutical composition includes a ribonucleoprotein complex including an RNA-guided nuclease (e.g., Cas9) that forms a complex with a gRNA and a cationic lipid. In some embodiments, the pharmaceutical composition includes a gRNA, a nucleic acid programmable DNA binding protein, a cationic lipid, and a pharma- ceutical acceptable excipient. The pharmaceutical composition can optionally include one or more additional therapeutically active substances.
[0448] kit In one aspect, the NLS-gRNA described herein may be provided and / or produced by a kit containing any one or more of the elements disclosed in the methods and compositions above. For example, the kit may include the NLS-gRNA, a ligase, and suitable buffering reagents.
[0449] In some embodiments, the kit further comprises a nucleobase editor.
[0450] In some embodiments, the kit includes one or more reagents for use in a process utilizing one or more of the elements described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents may be provided in a form that is usable in a particular assay or that requires the addition of one or more other components prior to use (e.g., concentrated or lyophilized form). The buffer may be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to the guide sequence for insertion into a vector such that the guide sequence and the regulatory element are operably linked. In some embodiments, the kit includes a homologous recombination template polynucleotide.
[0451] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety.In addition, the materials, methods, and examples are merely illustrative and are not intended to be limiting.Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs.Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described herein. EXAMPLES
[0452] The following examples illustrate some of the preferred modes of making and practicing the invention, however, it should be understood that these examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0453] Example 1: Ex vivo efficacy of NLS-gRNA This example describes an exemplary gRNA conjugated to an NLS of the present invention (NLS-gRNA) and its ex vivo efficacy. A peptide containing an NLS sequence and a peptide spacer was synthesized by solid-phase peptide synthesis. The synthesized peptide was conjugated to the 3' end of the gRNA via a thiol group as shown in Figure 1. As will be appreciated by those skilled in the art, the linker and peptide spacer can be modified in the practice of the present invention. Additionally, the sequences of the NLS, gRNA, and / or linker can be modified.
[0454] NLS-sgRNA was prepared and formulated in lipid nanoparticles with mRNA encoding the CRISPR-Cas9 base editor. The formulation was delivered to hepatocytes at three different ratios of mRNA:sgRNA (1:1, 3:1, and 9:1). As shown in Figure 2, NLS-sgRNA showed significantly higher base editing efficiency compared to gRNA without NLS sequence.
[0455] The data in this example show that CRISPR-Cas system (e.g., base editing) can be improved by using gRNA conjugated to NLS sequence. Without wishing to be bound by a particular theory, the improvement of CRISPR-Cas system can be partially attributed to better transport of NLS-gRNA to nucleus, which protects gRNA from cytoplasmic RNase, increasing the local concentration of gRNA, thus forming ribonucleic acid complex (RNP), and higher import rate to nucleus. Furthermore, cationic NLS sequence can act partially by promoting endosomal escape.
[0456] Example 2: In vivo efficacy of NLS-gRNA This example shows that NLS-gRNAs significantly improve base editing in vivo, even when compared to highly modified gRNAs. In this example, spCas9 gRNA was used together with an adenine base editor (ABE) that contains spCas9 nickase and adenosine deaminase.
[0457] gRNAs with various modifications were prepared. End-modified (EM) gRNA contains 6% modifications, Heavily Modified 1 (HM1) gRNA contains 47% modifications, Heavily Modified 2 (HM2) gRNA contains 60% modifications, and Heavily Modified 3 (HM3) gRNA contains 88% modifications, as shown in Figure 3A. NLS-gRNA contains an NLS sequence conjugated to the 3' end of the gRNA and 6% modifications. Two different mRNAs were prepared, both encoding the same base editor. Compared to mRNA2, mRNA1 is codon-optimized in the 3' and 5' UTR sequences. Various combinations of gRNAs with either mRNA1 or mRNA2 were formulated in LNPs and delivered to mice at subsaturating doses of 0.03 mpk or 0.01 mpk, as shown in Figure 3B.
[0458] The results show that NLS-gRNA showed higher base editing efficiency compared to all EM, HM1, HM2, or HM3 gRNA. Notably, even at ultra-low doses (0.01 mpk), base editing was visible for NLS-gRNA and was significantly higher than heavily modified (HM1, HM2, and HM3) gRNAs. Furthermore, combining NLS-gRNA with a lower potency mRNA (mRNA2) compensated for the quality of the mRNA. The base editing efficiency of mRNA2 using terminally modified gRNA was about 5%, whereas replacing gRNA with NLS-gRNA increased the base editing efficiency to more than 30%.
[0459] Example 3: Efficacy of NLS-gRNA in Non-Human Primates (NHP) This example shows that the improvement of base editing efficiency by using NLS-gRNA is also observed in NHPs. In this example, spCas9 gRNA was used together with spCas9-based adenine base editor (ABE).
[0460] Various gRNAs and mRNAs encoding base editors were formulated in lipid nanoparticles as shown in FIG. 4A. The formulations were delivered to NHPs at 1.0 mpk, and base editing efficiency was determined in the liver. The results showed that NLS-gRNA with mRNA1 (g5-BVN) and HM3 gRNA with mRNA1 (g4-BVB) showed the highest base editing efficiency, followed by g2-BVI, g3-BVV, and g1-BVE. Notably, NLS-gRNAs with terminal modifications (g1-BVN and g7-BG3IN) showed more than two-fold higher base editing efficiency compared to the respective terminally modified gRNAs without NLS (compared to g1-BVE and g6-BG3IE, respectively).
[0461] Next, toxicology studies were performed in NHPs. Alanine aminotransferase (ALT) and aspartate aminotransferase (AST) levels were measured to assess clinical pathology. Higher levels of ALT and AST correlate with liver damage. As shown in Figure 4B, minimal to mild increases in AST and / or ALT were observed 24 hours after administration for all test articles. Notably, g5-BVN, which contains NLS-gRNA with terminal modifications, showed the lowest increases in AST and ALT. Furthermore, no other significant changes in clinical pathology parameters were observed.
[0462] Overall, the data in this example show that NLS-gRNA improves the CRISPR-Cas system (e.g., base editing efficiency) in NHPs with reduced toxicity.
[0463] Example 4: Use of NLS-gRNA in saCas9 This example shows that NLS-gRNA can be applied to various Cas proteins. In this particular example, Staphylococcus aureus Cas9 (saCas9) was used. Notably, saCas9 requires its own guide, which is not compatible with spCas9 editing shown in the previous example.
[0464] Glycogen storage disease type 1a (GSD1a) is caused by mutations in the glucose-6-phosphatase (G6PC) gene that affect approximately 80% of patients with GSD1a. The R83C mutation affects approximately 900 US patients diagnosed with glycogen storage disease type 1a (GSD1a) each year. This mutation is a single base substitution that introduces a cysteine at position 83 (R83C) of the G6PC protein. Precise correction of R83C would likely restore expression of G6PC and normalize glucose metabolism.
[0465] gRNAs were prepared and their purity was determined. gRNAs with two different backbone chemistries were used in the study (sg029 vs. sg093). Sg093 guide has end modifications with 2'-OMe and phosphothioate modifications). Various gRNAs and mRNAs encoding base editors were formulated in LNPs at a 1:1 ratio of gRNA:mRNA. Adult transgenic mice heterozygous for huG6PC-R83C were administered the LNP formulations at a subsaturating dose of 1 mpk.
[0466] Figure 5 shows the correlation between base editing efficiency and gRNA purity, with 80% purity resulting in maximum base editing levels. Furthermore, NLS-gRNAs showed improved efficacy with spCas9 proteins over other sg093 guides that do not contain NLS sequences, demonstrating that the NLS-gRNAs of the present invention can be applied across multiple Cas proteins.
[0467] Example 5: In vivo base editing correction of metabolic defects in GSD1a R83C mice using NLS-gRNA In this example, a variant of an adenine base editor (ABE) was used in conjunction with an NLS-gRNA to correct the metabolic defect in GSD1a R83C mice. The R83C mutation introduces a single G>A conversion in the g6pc gene. ABEs in combination with the NLS-gRNA described herein result in a programmable A to G conversion in genomic DNA, thus supporting their utility for correcting this mutation.
[0468] The G6PC gRNA sequence hybridizes to the complement of the G6PC target sequence shown below: TIFF2024529425000072.tif14170
[0469] The NNGRRT PAM sequence (i.e., Staphylococcus aureus Cas9 (saCas9)) is underlined above. The gRNA sequence is: CAGUAUGGACACUGUCCAAA (SEQ ID NO:2).
[0470] The base editing efficiency of adenosine deaminase base editors (ABEs) using TadA variants MSP605, MSP824, MSP825, MSP680, MSP828, and MSP829 (see Table 1) and saCas9n was evaluated in vivo using a transgenic mouse model heterozygous for huG6PC and carrying the R83C mutation for glycogen storage disease type 1a (GSD1a) (Figures 6B and 6C). The use of saCas9 for efficient genome editing in vivo and an example of the saCas9 sgRNA scaffold are described in A. Ran et al. (2015, Nature, Vol. 520, pages 186-191). [Table 6]
[0471] Figure 6A shows the in vivo workflow used to introduce the base editor into transgenic mice. Lipid nanoparticles (LNPs) carrying the base editor mRNA and NLS-gRNA were administered to transgenic mice at a dose of 1 mg / kg via intravenous (IV) injection. Next-generation sequencing data from total liver extracts revealed significant correction for R83C (Figures 6B and 6C). TadA variant MSP828 showed approximately 40% precise correction of the R83C mutation, with low bystander editing. This level of mutation correction is expected to restore glucose homeostasis.
[0472] Example 6: In vivo base editing correction of metabolic defects in GSD1a R83C mice Overview of GSD1a As shown diagrammatically in FIG. 7, (GSD1a) is an autosomal recessive disorder caused by mutations in the G6PC gene. The most prevalent pathogenic mutation identified in Caucasian GSD1a patients is R83C, which is located in the active site of the enzyme and is associated with inactivation of G6Pase. Loss of G6Pase function can result in life-threatening hypoglycemia, seizures, and even death. To mitigate hypoglycemia, patients must maintain strict and frequent adherence to glucose replacement with slow glucose-releasing formulas throughout the day and night. A single missed or delayed dose can cause emergent hypoglycemia. Among many complications, liver enlargement, uric acid, lactic acid, and lipid accumulation are common in GSD1a patients.
[0473] Utility of the described base editors for generating durable and predictable single-nucleotide substitutions The R83C mutation introduces a single G>A conversion in the g6pc gene. The adenine base editor (ABE) described herein results in a programmable A to G conversion in genomic DNA, thus supporting their utility for correcting this mutation. As shown diagrammatically in FIG. 8, the adenine base editor is a fusion protein containing an evolved TadA deaminase connected to a CRISPR-Cas enzyme. The base editor binds to the target DNA that is complementary to the guide RNA (overlapped on the CRISPR-Cas9 enzyme) and exposes a stretch of single-stranded DNA. The deaminase converts the target adenine to inosine, and the Cas enzyme nicks the opposite strand, which is then repaired, completing the conversion of the base pair. Thus, direct repair of the point mutation has the potential for restoration of gene function.
[0474] In this example, a base editor for an A>G transition in the g6pc gene was optimized for the correction of R83C. Figure 9A shows the target DNA sequence (CCACCAGTATGGACACTGTCCAAAGAGAAT (SEQ ID NO: 17)) and the underlying amino acid translation for the GSD1a R83C mutation (WWYPCQGFLI, SEQ ID NO: 18). The target nucleobase to be edited is represented by a double underline at position 12. The editing window also includes a potential bystander, shown by a single underline at position 6. An edit that could result in a synonymous transition is shown at position 10.
[0475] For screening, HEK293 cell lines expressing the G6PC transgene carrying the R83C mutation were generated and transfected with base editor mRNA and gRNA. Allele frequencies were assessed by high-throughput targeted amplicon next-generation sequencing. Variants 1-5 represent combinations of gRNA and base editor RNA engineered for optimized targeted correction. Variant 5 resulted in approximately 60% targeted base editing efficiency for R83C correction and limited bystander editing (Figure 9B).
[0476] Murine in vivo disease model and demonstration of in vivo correction of the R83C single nucleotide mutation In vivo correction of R83C base editing To verify the base editing efficiency for R83C correction in vivo, we generated novel GSD1a mice expressing a human G6PC-R83C transgene instead of mouse G6pc. We confirmed that mice homozygous for huR83C exhibited postnatal lethality and rarely survived until weaning (day 21). On glucose replacement therapy, animals survived to at least 3 weeks of age and revealed characteristic pathological signs of GSD1a, including reduced body weight, enlarged liver, marked G6Pase inhibition, and abnormalities in serum metabolites, compared to littermate controls (Figure 7). This phenotype is consistent with published and clinical reports in humans.
[0477] For in vivo experiments, LNP-mediated delivery was tested in transgenic mice that were heterozygous for huR83C due to neonatal lethality of homozygous mice. The schematic in Figure 6A shows the in vivo workflow using lipid nanoparticles or LNP co-formulations of base editor mRNA and gRNA administered via IV injection. Given the neonatal lethality of homozygous mice, LNP administration was performed via the temporal vein immediately after birth and activity was compared to that in adult mice. Next-generation sequencing (NGS) analysis of total liver extracts revealed base editing efficiency of approximately 40% in adults and up to about 60% in neonates, with a wider range of efficiencies (Figure 11A). Bystander editing remained low in adults and neonates. (Figure 11A).
[0478] Newborn mice homozygous for huR83C were treated with lipid nanoparticles (LNPs) containing guide RNA and mRNA encoding ABE. Treated mice were found to survive and grow normally to 3 weeks of age without hypoglycemia-induced seizures in the absence of glucose therapy. Treated homozygous huR83C mice showed up to about 60% editing efficiency in total liver extracts, consistent with littermate controls that were heterozygous for huR83C (Figure 11B). Thus, it was demonstrated that LNP-mediated R83C correction was associated with the survival of homozygous huR83C mice.
[0479] Reversing GSD-1a pathology by base editing to correct R83C in vivo
[0480] At 3 weeks, we verified and confirmed that treated homozygous huR83C mice exhibited proper metabolic function with restoration of near-normal serum metabolites including glucose, triglycerides, chole...