Novel CRISPR enzymes, methods, systems, and uses thereof
Patent Information
- Application Number
- JP2024535360
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-17
- Filing Date
- 2022-12-16
- Publication Date
- 2025-12-11
AI Technical Summary
Existing CRISPR-Cas9 systems lack the diversity and precision needed for effective genetic manipulation across various genomic targets, requiring a broader toolkit for precise gene editing.
Identification and engineering of novel Cas9 enzymes from specific bacterial strains, each recognizing unique protospacer-adjacent motifs (PAMs), enabling targeted gene editing in eukaryotic cells.
Expands the capabilities of CRISPR-Cas9 systems by providing enzymes with diverse PAM recognition, allowing precise modifications in a wide range of target genes, enhancing genetic manipulation efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 291,252, filed December 17, 2021, the contents of which are incorporated by reference in their entirety for all purposes.
[0002] Sequence Listing Citation References The contents of the file named "BEM-013WO1_SL.xml", created on November 10, 2022 and having a size of 239 kilobytes, are incorporated herein by reference in their entirety. [Background technology]
[0003] Enzymes from the prokaryotic clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated proteins (CRISPR-Cas) systems have been utilized as reprogrammable and highly specific genome editing tools for use in eukaryotes. In addition to genome editing and cleavage, CRISPR-Cas9 can be used to localize effector molecules to specific sites on the genome, allowing genetic and epigenetic regulation as well as transcriptional regulation through a variety of mechanisms.
[0004] However, diverse genomes and genomic targets require diverse tools for effective genetic manipulation, and there remains a need to expand the CRISPR toolbox through the discovery and engineering of novel Cas proteins that can recognize and target diverse sequences.
[0005] Although the CRISPR-Cas9 system can be used to knock out genes or modify the expression of genes, certain types of gene editing require precise modifications to target genes, such as editing a single base within a gene. Such precise modifications remain a challenge and require a diverse gene editing toolkit to achieve precise genome modifications in a wide variety of target genes. Summary of the Invention
[0006] The identification of novel Cas9 enzymes with specificity for unique protospacer adjacent motifs (PAMs) allows for the expansion of available tools for gene editing. The present invention provides novel engineered, non-naturally occurring Cas9 enzymes isolated from, inter alia, Streptococcus equinus ATCC 33317, Enterococcus hirae strain F1129E, Streptococcus equinus strain AG46, Staphylococcus simulans strain 19, Streptococcus intermedius B196 strain G1552, Streptococcus sanguinis SK330, Streptococcus sp. C150, Streptococcus oralis subsp. oralis strain RH_1735_08, Streptococcus oralis SK313, Staphylococcus warneri strain 691, Staphylococcus stiuri strain SNUC 2430, Streptococcus gallolyticus strain AM24-4, Lactobacillus kullabergensis strain Biut2, and Streptococcus suis strain LSS83.
[0007] The present invention is based in part on the surprising discovery that novel Cas9 enzymes discovered from different bacteria that recognize specific PAM sequences can be engineered for expression in eukaryotic cells (e.g., humans, plants, etc.). Thus, the described Cas9 enzymes and their variants are functional in eukaryotic organisms. The examples provided herein demonstrate the use of engineered non-naturally occurring Cas9 enzymes in human cells with diverse PAM recognition sequences that target various genomic sites.For example, engineered Cas9 from Streptococcus equinus ATCC 33317 recognizes the consensus PAM sequence 5'-NRGNR-3', Enterococcus hirae strain F1129E recognizes the consensus PAM sequence 5'-NRG-3', Streptococcus equinus strain AG46, Staphylococcus warneri strain 691, and Staphylococcus sciuri strain SNUC 2430 recognize the consensus PAM sequence 5'-NNGR-3', Staphylococcus simulans strain 19 recognizes the consensus PAM sequence 5'-NNGRRT-3', Streptococcus intermedius B196 strain G1552 recognizes the consensus PAM sequence 5'-NNAAAA-3', and Streptococcus sanguinis SK330 recognizes the consensus PAM sequence 5'-NGGNG-3', Streptococcus sp. C150 recognizes the consensus PAM sequence 5'-NNGNRG-3', Streptococcus oralis subsp. oralis strain RH_1735_08 recognizes the consensus PAM sequence 5'-NNAAAC-3', Streptococcus oralis SK313 recognizes the consensus PAM sequence 5'-NNRAAG-3', Streptococcus gallolyticus strain AM24-4 recognizes the consensus PAM sequence 5'-NNAYAA-3', Lactobacillus kullabergensis strain Biut2 recognizes the consensus PAM sequence 5'-NNGAAA-3', and Streptococcus suis strains recognize the consensus PAM sequence 5'-NNAAA-3' (H=A, C or T; R=A or G).
[0008] In one aspect, Streptococcus equinus ATCC 33317 Cas9, Enterococcus hirae strain F1129E Cas9, Streptococcus equinus strain AG46 Cas9, Staphylococcus simulans strain 19 Cas9, Streptococcus intermedius B196 strain G1552, Streptococcus sanguinis SK330 Cas9, Streptococcus sp.C150 Cas9, Streptococcus oralis subsp.oralis strain RH_1735_08 Cas9, Streptococcus oralis SK313 Cas9, Staphylococcus warneri strain 691 Cas9, Staphylococcus sciuri strain SNUC 2430 Cas9, Streptococcus gallolyticus strain AM24-4 Cas9, Lactobacillus kullabergensis strain Biut2 Cas9 and Streptococcus Provided herein is an engineered, non-naturally occurring Cas9 protein modified from C. suis strain LSS83 Cas9.
[0009] In some embodiments, the Streptococcus equinus ATCC 33317 Cas9 protein has at least 80% sequence identity to
[0010] In some embodiments, the Enterococcus hirae strain F1129E Cas9 protein has at least 80% sequence identity to
[0011] In some embodiments, the Streptococcus equinus strain AG46 Cas9 protein has at least 80% sequence identity to
[0012] In some embodiments, the Staphylococcus simulans strain 19 Cas9 protein has at least 80% sequence identity to
[0013] In some embodiments, the Streptococcus_intermedius_strain_B196_G1552 Cas9 protein has at least 80% sequence identity to
[0014] In some embodiments, the Streptococcus sanguinis SK330 Cas9 protein has at least 80% sequence identity to
[0015] In some embodiments, the Streptococcus sp. C150 Cas9 protein has at least 80% sequence identity to
[0016] In some embodiments, the Streptococcus oralis subsp. oralis strain RH_1735_08 Cas9 protein has at least 80% sequence identity to
[0017] In some embodiments, the Streptococcus oralis SK313 Cas9 protein has at least 80% sequence identity to
[0018] In some embodiments, the Staphylococcus warneri strain 691 Cas9 protein has at least 80% sequence identity to
[0019] In some embodiments, the Staphylococcus sciuri strain SNUC 2430 Cas9 protein has at least 80% sequence identity to
[0020] In some embodiments, the Streptococcus gallolyticus strain AM24-4 Cas9 protein has at least 80% sequence identity to
[0021] In some embodiments, the Lactobacillus kullabergensis strain Biut2 Cas9 protein has at least 80% sequence identity to
[0022] In some embodiments, the Streptococcus suis strain LSS83 Cas9 protein has at least 80% sequence identity to
[0023] In some embodiments, the Cas9 protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO:1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14.
[0024] In some embodiments, the Cas9 protein further comprises a nuclear localization sequence (NLS) and / or a FLAG, HIS, or HA tag.
[0025] In some embodiments, the Streptococcus equinus ATCC 33317 Cas9 (Seq2Cas9) has an amino acid sequence at least 80% identical to TIFF2024546841000002.tif41169TIFF2024546841000003.tif114169TIFF2024546841000004.tif9169 (SEQ ID NO: 15) (D10A mutation, bold underlined italics)
[0026] In some embodiments, Enterococcus hirae strain F1129E Cas9 (EhiCas9) has an amino acid sequence at least 80% identical to: TIFF2024546841000005.tif83169TIFF2024546841000006.tif68169TIFF2024546841000007.tif9169 (SEQ ID NO: 16) (D10A mutation, bold underlined italics)
[0027] In some embodiments, Streptococcus equinus strain AG46 (SeqCas9) has an amino acid sequence at least 80% identical to TIFF2024546841000008.tif121169TIFF2024546841000009.tif8169TIFF2024546841000010.tif8169 (SEQ ID NO: 17)
[0028] In some embodiments, Staphylococcus simulans strain 19 (SsiCas9) has an amino acid sequence that is at least 80% identical to TIFF2024546841000011.tif116169TIFF2024546841000012.tif9169(SEQ ID NO:18)
[0029] In some embodiments, Streptococcus intermedius B196 strain G1552 (SinCas9) has an amino acid sequence at least 80% identical to TIFF2024546841000013.tif47169TIFF2024546841000014.tif83169TIFF2024546841000015.tif7169 (SEQ ID NO: 19)
[0030] In some embodiments, the Streptococcus sanguinis SK330 Cas9 has an amino acid sequence that is at least 80% identical to TIFF2024546841000016.tif121169TIFF2024546841000017.tif37169TIFF2024546841000018.tif8169 (SEQ ID NO: 20)
[0031] In some embodiments, the Streptococcus sp. C150 Cas9 has an amino acid sequence that is at least 80% identical to TIFF2024546841000019.tif130169TIFF2024546841000020.tif9169(SEQ ID NO:21)
[0032] In some embodiments, Streptococcus oralis subsp. oralis strain RH_1735_08 Cas9 has an amino acid sequence that is at least 80% identical to TIFF2024546841000021.tif14169TIFF2024546841000022.tif118169(SEQ ID NO:22)
[0033] In some embodiments, the Streptococcus oralis SK313 Cas9 has an amino acid sequence that is at least 80% identical to TIFF2024546841000023.tif89169TIFF2024546841000024.tif43169TIFF2024546841000025.tif9169 (SEQ ID NO: 23)
[0034] In some embodiments, the Streptococcus oralis SK313 Cas9 has an amino acid sequence that is at least 80% identical to: TIFF2024546841000026.tif115169TIFF2024546841000027.tif9169(SEQ ID NO:24)
[0035] In some embodiments, Staphylococcus warneri strain 691 Cas9 has an amino acid sequence at least 80% identical to TIFF2024546841000028.tif116169TIFF2024546841000029.tif10169(SEQ ID NO:25)
[0036] In some embodiments, Streptococcus gallolyticus strain AM24-4 Cas9 has an amino acid sequence at least 80% identical to TIFF2024546841000030.tif83169TIFF2024546841000031.tif49169TIFF2024546841000032.tif7169 (SEQ ID NO: 26)
[0037] In some embodiments, the Lactobacillus kullabergensis strain Biut2 Cas9 has an amino acid sequence at least 80% identical to TIFF2024546841000033.tif149169TIFF2024546841000034.tif8169TIFF2024546841000035.tif9169 (SEQ ID NO: 27)
[0038] In some embodiments, Streptococcus suis strain LSS83 Cas9 has an amino acid sequence that is at least 80% identical to TIFF2024546841000036.tif130169 (SEQ ID NO: 28)
[0039] In some embodiments, the amino acid sequence of the Cas9 protein comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten mutations in SEQ ID NO:1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14.
[0040] In some embodiments, the mutation is an amino acid substitution.
[0041] In some embodiments, the Cas9 protein has nickase activity.
[0042] In some embodiments, at least one mutation results in an inactive Cas9 (dCas9).
[0043] In some embodiments, the Cas9 protein comprises at least one amino acid mutation in the PAM-interacting, HNH domain, and / or RuvC domain.
[0044] In some embodiments, the Cas9 protein further comprises a nuclear localization sequence (NLS) and / or a FLAG, HIS, or HA tag.
[0045] In one aspect, provided herein is an engineered, non-naturally occurring Cas9 fusion protein comprising a Cas9 protein having at least 80% identity to SEQ ID NO:1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14, wherein the Cas9 protein is fused to a histone demethylase, a transcriptional activator, or a deaminase.
[0046] In some embodiments, the Cas9 protein is fused to a cytosine deaminase or an adenosine deaminase.
[0047] In some embodiments, the Cas9 protein is fused to an adenosine deaminase and has an amino acid sequence at least 80% identical to (a) TIFF2024546841000037.tif130169TIFF2024546841000038.tif47169TIFF2024546841000039.tif10169 (sequence number 29). (b) TIFF2024546841000040.tif163169TIFF2024546841000041.tif8169TIFF2024546841000042.tif9169 (SEQ ID NO: 30) (c) TIFF2024546841000043.tif150169TIFF2024546841000044.tif10169 (sequence number 31). (d) TIFF2024546841000045.tif34169TIFF2024546841000046.tif117169TIFF2024546841000047.tif9169 (sequence number 32). (e) TIFF2024546841000048.tif94169TIFF2024546841000049.tif49169TIFF2024546841000050.tif7169 (sequence number 33). (f) TIFF2024546841000051.tif162169TIFF2024546841000052.tif14169TIFF2024546841000053.tif10169 (sequence number 34). (g) TIFF2024546841000054.tif149169TIFF2024546841000055.tif8169 (sequence number 35). (h) TIFF2024546841000056.tif28169TIFF2024546841000057.tif123169TIFF2024546841000058.tif9169 (sequence number 36). (i) TIFF2024546841000059.tif89169TIFF2024546841000060.tif63169TIFF2024546841000061.tif8169 (sequence number 37). (j) TIFF2024546841000062.tif141169TIFF2024546841000063.tif8169 (sequence number 38). (k) TIFF2024546841000064.tif142169TIFF2024546841000065.tif7169 (sequence number 39). (l) TIFF2024546841000066.tif62169TIFF2024546841000067.tif89169TIFF2024546841000068.tif9169 (sequence number 40). (m) TIFF2024546841000069.tif121169TIFF2024546841000070.tif56169TIFF2024546841000071.tif9169 (sequence number 41). (n) TIFF2024546841000072.tif150169TIFF2024546841000073.tif8169 (sequence number 42).
[0048] In some embodiments, the Cas9 protein is fused to a cytosine deaminase.
[0049] In some embodiments, the Streptococcus equinus ATCC 33317 Cas9 protein recognizes a PAM consensus sequence comprising 5'-NRGNR-3'.
[0050] In some embodiments, Enterococcus hirae strain F1129E recognizes a PAM consensus sequence comprising 5'-NRG-3'.
[0051] In some embodiments, Streptococcus equinus strain AG46 recognizes a PAM consensus sequence comprising 5'-NNGR-3'.
[0052] In some embodiments, Staphylococcus warneri strain 691 recognizes a PAM consensus sequence comprising 5'-NNGR-3'.
[0053] In some embodiments, Staphylococcus sciuri strain SNUC 2430 recognizes a PAM consensus sequence comprising 5'-NNGR-3'.
[0054] In some embodiments, Staphylococcus simulans strain 19 recognizes a PAM consensus sequence comprising 5'-NNGRRT-3'.
[0055] In some embodiments, Streptococcus intermedius B196 strain G1552 recognizes a consensus PAM sequence comprising 5'-NNAAAA-3'.
[0056] In some embodiments, Streptococcus sanguinis SK330 recognizes a consensus PAM sequence including 5'-NGGNG-3'.
[0057] In some embodiments, Streptococcus sp. C150 recognizes a consensus PAM sequence comprising 5'-NNGNRG-3'.
[0058] In some embodiments, Streptococcus oralis subsp. oralis strain RH_1735_08 recognizes a consensus PAM sequence including 5'-NNAAAC-3'.
[0059] In some embodiments, Streptococcus oralis SK313 recognizes a consensus PAM sequence comprising 5'-NNRAAG-3'.
[0060] In some embodiments, Streptococcus gallolyticus strain AM24-4 recognizes a consensus PAM sequence including 5'-NNAYAA-3'.
[0061] In some embodiments, Lactobacillus kullabergensis strain Biut2 recognizes a consensus PAM sequence including 5'-NNGAAA-3'.
[0062] In some embodiments, Streptococcus suis strains recognize a consensus PAM sequence comprising 5'-NNAAA-3' (H=A, C, or T; R=A or G).
[0063] In some embodiments, a nucleic acid encoding a Cas9 protein is provided.
[0064] In some embodiments, the nucleic acid is codon optimized for expression in a mammalian cell.
[0065] In some embodiments, the nucleic acid is codon optimized for expression in a human cell.
[0066] In some embodiments, a eukaryotic cell is provided that comprises a Cas9 protein.
[0067] In some embodiments, the cell is a human cell, hi some embodiments, the cell is a plant cell.
[0068] In one aspect, a method is provided for cleaving a target nucleic acid in a eukaryotic cell comprising contacting the cell with a Cas9 as described herein and an RNA guide or a nucleic acid encoding the RNA guide, wherein the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to the target nucleic acid, and wherein the Cas9 protein is capable of binding to the RNA guide and causing cleavage at the target nucleic acid sequence complementary to the RNA guide.
[0069] In one aspect, a method is provided for modifying expression of a target nucleic acid in a eukaryotic cell comprising contacting the cell with a Cas9 as described herein and an RNA guide or a nucleic acid encoding the RNA guide, wherein the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to the target nucleic acid, and wherein the Cas9 protein can bind to the RNA guide and cause cleavage at the target nucleic acid sequence complementary to the RNA guide.
[0070] In one aspect, a method is provided for modifying expression of a target nucleic acid in a eukaryotic cell comprising contacting a cell with a Cas9 as described herein and an RNA guide or a nucleic acid encoding the RNA guide, wherein the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to the target nucleic acid, and wherein the Cas9 protein is capable of binding to the RNA guide and editing the target nucleic acid sequence that is complementary to the RNA guide.
[0071] In one aspect, a method is provided for modifying a target nucleic acid in a eukaryotic cell comprising contacting the cell with a Cas9 as described herein and an RNA guide or a nucleic acid encoding the RNA guide, wherein the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to the target nucleic acid, and wherein the Cas9 protein is capable of binding to the RNA guide and editing the target nucleic acid sequence complementary to the RNA guide.
[0072] In some embodiments, the Cas9 protein is an inactive Cas9 (dCas9).
[0073] In some embodiments, the dCas9 is fused to a deaminase.
[0074] In some embodiments, the RNA guides include crRNA and tracrRNA.
[0075] In some embodiments, the RNA guide comprises an sgRNA.
[0076] In some embodiments, the sgRNA for use in Seq2Cas9 is 5'- TIFF2024546841000074.tif8169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000075.tif10169 (sequence number 43).
[0077] In some embodiments, the sgRNA for use in EhiCas9 is 5'- TIFF2024546841000076.tif8169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000077.tif9169 (sequence number 44).
[0078] In some embodiments, the sgRNA for use in SeqCas9 is 5'- TIFF2024546841000078.tif7169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000079.tif10169 (sequence number 45).
[0079] In some embodiments, the sgRNA for use in SsiCas9 is 5'- TIFF2024546841000080.tif7169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000081.tif9169 (sequence number 46).
[0080] In some embodiments, the sgRNA for use in SinCas9 is 5'- TIFF2024546841000082.tif8169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000083.tif9169 (sequence number 47).
[0081] In some embodiments, the sgRNA for use in SsaCas9 is 5'- TIFF2024546841000084.tif8169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000085.tif9169 (sequence number 48).
[0082] In some embodiments, the sgRNA for use in Ssc2Cas9 is 5'- TIFF2024546841000086.tif7169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000087.tif7169 (sequence number 49).
[0083] In some embodiments, the sgRNA for use in Sor2Cas9 is 5'- TIFF2024546841000088.tif8169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000089.tif10169 (sequence number 50).
[0084] In some embodiments, the sgRNA for use in SorCas9 is 5'- The scaffold comprises a sequence having at least 80% identity to TIFF2024546841000090.tif16169 (SEQ ID NO: 51).
[0085] In some embodiments, the sgRNA for use in SwaCas9 is 5'- TIFF2024546841000091.tif9169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000092.tif10169 (sequence number 52).
[0086] In some embodiments, the sgRNA for use in SscCas9 is 5'- TIFF2024546841000093.tif9169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000094.tif9169 (sequence number 53).
[0087] In some embodiments, the sgRNA for use in SgaCas9 is 5'- 5'- TIFF2024546841000095.tif8169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000096.tif8169 (sequence number 54).
[0088] In some embodiments, the sgRNA for use in LkuCas9 is 5'- 5'- Includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000097.tif8169TIFF2024546841000098.tif8169TIFF2024546841000099.tif10169 (sequence number 55).
[0089] In some embodiments, the sgRNA for use with SsuCas9 is 5'- 5'- TIFF2024546841000100.tif8169 includes a scaffold comprising a sequence having at least 80% identity to TIFF2024546841000101.tif10169 (sequence number 56).
[0090] In the preceding embodiment, for SEQ ID NOs: 43 to 56, Direct Repetition (italics and underline), tetraloop (italics), tracrRNA (underlined).
[0091] In some embodiments, the crRNA comprises a guide sequence about 16-26 nucleotides in length.
[0092] In some embodiments, the crRNA comprises a guide sequence that is 18-24 nucleotides long.
[0093] In some embodiments, the break in the target nucleic acid is a single-stranded break or a double-stranded break.
[0094] In some embodiments, the cleavage in the target nucleic acid is a single-stranded break.
[0095] In some embodiments, the Cas9 protein is a nuclease that cleaves both strands of a target nucleic acid sequence, hi some embodiments, the Cas9 is a nickase that cleaves one strand of a target nucleic acid sequence.
[0096] In some embodiments, the target nucleic acid is 5' to a protospacer adjacent motif (PAM) sequence.
[0097] In some embodiments, the Cas9 is operably linked to a promoter sequence for expression in a eukaryotic cell and the guide RNA is operably linked to a promoter sequence for expression in a eukaryotic cell.
[0098] In some embodiments, the eukaryotic cell is a human cell.
[0099] In some embodiments, the promoter sequence is a eukaryotic promoter or a viral promoter.
[0100] In one aspect, provided herein is an engineered, non-naturally occurring CRISPR-Cas system comprising an RNA guide or a nucleic acid encoding the RNA guide, where the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to a target nucleic acid, and a codon-optimized CRISPR-associated (Cas) protein having at least 80% sequence identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14, where the Cas protein is capable of binding to the RNA guide and causing cleavage in a target nucleic acid sequence complementary to the RNA guide.
[0101] In one aspect, provided herein is an engineered non-naturally occurring CRISPR-Cas system comprising an RNA guide or a nucleic acid encoding the RNA guide, where the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to a target nucleic acid, and a codon-optimized CRISPR-associated (Cas) protein having at least 80% sequence identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14, where the Cas protein is fused to a deaminase, and where the Cas protein fusion is capable of binding to the RNA guide and editing a target nucleic acid sequence complementary to the RNA guide.
[0102] In one embodiment, the Cas9 protein is an inactive Cas9 (dCas9).
[0103] In one embodiment, the RNA guide comprises crRNA and tracrRNA.
[0104] In one embodiment, the RNA guide comprises an sgRNA.
[0105] In one embodiment, the Cas protein is operably linked to a promoter sequence for expression in a eukaryotic cell and the guide RNA is operably linked to a promoter sequence for expression in a eukaryotic cell.
[0106] In one embodiment, the eukaryotic cell is a human cell.
[0107] In one embodiment, the promoter sequence is a eukaryotic promoter sequence.
[0108] In one embodiment, a nucleic acid encoding the system described herein is provided.
[0109] In one embodiment, a vector is provided that comprises the system described herein.
[0110] In one embodiment, the vector is a plasmid vector or a viral vector.
[0111] In one embodiment, the viral vector is an adeno-associated viral (AAV) vector or a lentiviral vector.
[0112] In one embodiment, the viral vector is an AAV vector.
[0113] In one embodiment, two or more AAV vectors are used in the packaging system.
[0114] In one embodiment, a method of treating a disorder or disease in a subject in need thereof comprises administering to the subject a system described herein, wherein a guide RNA is complementary to at least 10 nucleotides of a target nucleic acid associated with the condition or disease, a Cas protein associates with the guide RNA, the guide RNA binds to the target nucleic acid, and the Cas protein causes cleavage in the target nucleic acid, and optionally, the Cas9 is an inactive Cas9 fused to a deaminase (dCas9), effecting one or more base edits in the target nucleic acid, thereby treating the disorder or disease.
[0115] In some embodiments, the guide RNA is complementary to about 18-24 nucleotides.
[0116] In some embodiments, the guide RNA is complementary to 20 nucleotides.
[0117] In some embodiments, the base editor comprises a fusion protein.
[0118] In some embodiments, the base editor comprises an adenosine deaminase domain or a cytidine deaminase domain.
[0119] In some embodiments, provided herein is a method of editing a nucleobase of a polynucleotide, the method comprising contacting the polynucleotide with a base in a complex with one or more guide RNAs, wherein the base editor comprises an adenosine deaminase domain, and the one or more guide RNAs target the base editor to effect an A·T to G·C modification in the polynucleotide.
[0120] In some embodiments, provided herein are methods of editing nucleobases of a polynucleotide, the method comprising contacting the polynucleotide in a complex with a base editor and one or more guide RNAs, wherein the base editor comprises a cytidine deaminase domain, and the one or more guide RNAs target the base editor to result in a C·G to T·A modification in the polynucleotide.
[0121] In some embodiments, the editing results in less than 50% indel formation in the target polynucleotide sequence.
[0122] In some embodiments, the editing generates a point mutation.
[0123] definition In order that the present invention may be more readily understood, certain terms are first defined below. Further definitions for these terms, as well as other terms, are set forth throughout the specification.
[0124] a or an: The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "an element" means one element or more than one element.
[0125] Approximately or about: As used herein, the term "approximately" or "about" when applied to one or more values of interest refers to a value similar to a stated reference value. In certain embodiments, the term "approximately" or "about" refers to a range of values that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater or less) of the stated reference value (except where such number exceeds 100% of possible values), unless otherwise stated or clear from the context.
[0126] Associated with: As the term is used herein, two events or entities are "associated" with each other when the presence, level and / or form of one correlates with the presence, level and / or form of the other. For example, a particular entity (e.g., a polypeptide) is considered to be associated with a particular disease, disorder or condition when its presence, level and / or form correlates with the incidence and / or susceptibility of the disease, disorder or condition (e.g., across a relevant population). In some embodiments, two or more entities are physically "associated" with each other such that they are in and maintain physical proximity to each other when they directly or indirectly interact. In some embodiments, two or more entities that are physically associated with each other are covalently bound to each other, and in some embodiments, two or more entities that are physically associated with each other are not covalently bound to each other, but are non-covalently associated, for example, by hydrogen bonds, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.
[0127] Base editor: "Base editor (BE)" or "Nucleic acid base editor (NBE)" refers to an agent that binds to a polynucleotide and has nucleobase modifying activity. In various embodiments, the base editor comprises a nucleobase modifying polypeptide (e.g., a deaminase) and a polynucleotide programmable nucleotide binding domain in conjunction with a guide polynucleotide (e.g., a guide RNA). In various embodiments, the agent is a biomolecular complex that comprises a protein domain with base editing activity, i.e., a domain that can modify a base (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide programmable DNA binding domain is fused or bound to a deaminase domain. In one embodiment, the agent is a fusion protein that comprises one or more domains with base editing activity. In another embodiment, the protein domain with base editing activity is bound to the guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase). In some embodiments, the protein domain with base editing activity is capable of deaminating a base in a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases in a DNA molecule. In some embodiments, the base editor is capable of deaminating cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor is capable of deaminating cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is a nuclease-inactive Cas9 (dCas9) fused to an adenosine deaminase. In some embodiments, the base editor is fused to an inhibitor of base excision repair (e.g., a UGI domain, or a dISN domain).In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and an inhibitor of base excision repair (e.g., a UGI domain or a dISN domain). In other embodiments, the base editor is an abasic base editor. Details of base editors are described in International PCT Application Nos. 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Also, Komor, AC, et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016), Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017), Komor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.”Nat Rev Genet.2018 Dec;19(12):770-788. doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.
[0128] Base editing activity: "Base editing activity" means acting to chemically modify a base in a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is a cytidine deaminase activity, e.g., converting a targeted C·G to T·A. In another embodiment, the base editing activity is an adenosine or adenine deaminase activity, e.g., converting an A·T to G·C. In another embodiment, the base editing activity is a cytosine or cytidine deaminase activity (e.g., converting a targeted C·G to T·A) and an adenosine or adenine deaminase activity (e.g., converting an A·T to G·C).
[0129] Base editor system: The term "base editor system" refers to a system for editing nucleobases of a target nucleotide sequence. In various embodiments, the base editor (BE) system includes (1) a polynucleotide programmable nucleotide binding domain (e.g., Cas9), a deaminase domain, and a cytidine deaminase domain for deaminating nucleobases in a target nucleotide sequence, and (2) one or more guide polynucleotides (e.g., guide RNAs) in conjunction with the polynucleotide programmable nucleotide binding domain. In various embodiments, the base editor (BE) system includes a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain having nucleic acid sequence-specific binding activity. In some embodiments, the base editor system includes (1) a base editor (BE) including a polynucleotide programmable DNA binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence, and (2) one or more guide RNAs in combination with a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).
[0130] In some embodiments, the polynucleotide programmable nucleotide binding domain can target the deaminase domain to a target nucleotide sequence by non-covalently interacting with or associating with the deaminase domain. For example, in some embodiments, the nucleic acid base editing component (e.g., the deaminase component) can include an additional heterologous moiety or heterologous domain that can interact with, associate with, or form a complex with the additional heterologous moiety or heterologous domain that is part of the polynucleotide programmable nucleotide binding domain. In some embodiments, the additional heterologous moiety can bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments, the additional heterologous moiety can bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, the additional heterologous moiety can bind to a guide polynucleotide. In some embodiments, the additional heterologous moiety can bind to a polypeptide linker. In some embodiments, the additional heterologous moiety can bind to a polynucleotide linker. The additional heterologous moiety can be a protein domain. In some embodiments, the additional heterologous moiety can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a steryl alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.
[0131] Biologically active: As used herein, the phrase "biologically active" refers to the characteristic of any agent that has activity in a biological system, particularly in an organism. For example, an agent that, when administered to an organism, has a biological effect on the organism is considered to be biologically active. In certain embodiments, if a peptide is biologically active, a portion of the peptide that shares at least one biological activity of the peptide is typically referred to as a "biologically active" portion.
[0132] Cleavage: As used herein, cleavage refers to the cleavage of the target nucleic acid made by the nuclease of the CRISPR system described herein. In some embodiments, the cleavage event is double-stranded DNA cleavage. In some embodiments, the cleavage event is single-stranded DNA cleavage. In some embodiments, the cleavage event is single-stranded RNA cleavage. In some embodiments, the cleavage event is double-stranded RNA cleavage.
[0133] Complementary: As used herein, complementary refers to a strand of nucleic acid that forms Watson-Crick base pairing or non-conventional base pairing with bases on a second nucleic acid strand, such that A bases pair with T and C bases pair with G. In other words, nucleic acids that hybridize to each other under appropriate conditions.
[0134] Clustered regularly interspaced short palindromic repeats (CRISPR) associated (Cas) system: As used herein, CRISPR-Cas9 system refers to nucleic acids and / or proteins involved in the expression or directing the activity of CRISPR effectors, including sequences encoding CRISPR effectors, RNA guides, and other sequences and transcripts from the CRISPR locus. In some embodiments, the CRISPR system is an engineered non-naturally occurring CRISPR system. In some embodiments, the components of the CRISPR system may include nucleic acid(s) (e.g., vectors) encoding one or more components of the system, component(s) in protein form, or a combination thereof.
[0135] CRISPR array: The term "CRISPR array" as used herein refers to a nucleic acid (e.g., DNA) segment that includes CRISPR repeats and spacers, starting with the first nucleotide of the first CRISPR repeat and ending with the last nucleotide of the last (terminal) CRISPR repeat. Typically, each spacer in a CRISPR array is located between two repeats. The term "CRISPR repeat" or "CRISPR direct repeat" or "direct repeat" as used herein refers to multiple short direct repeat sequences that show little or no sequence variation in a CRISPR array.
[0136] CRISPR-associated protein (Cas): The term "CRISPR-associated protein," "CRISPR effector," "effector," or "CRISPR enzyme," as used herein, refers to a protein that performs an enzymatic activity or binds to a target site on a nucleic acid specified by an RNA guide. In various embodiments, a CRISPR effector has endonuclease activity, nickase activity, exonuclease activity, transposase activity, and / or excision activity. It should be understood that any of the Cas9 proteins provided herein include naturally occurring Cas9 proteins and non-naturally occurring variants thereof.
[0137] crRNA: The term "CRISPR RNA" or "crRNA" as used herein refers to an RNA molecule that contains a guide sequence that is used by a CRISPR effector to target a specific nucleic acid sequence. Typically, the crRNA contains a sequence that mediates target recognition and a sequence that forms a duplex with the tracrRNA. In some embodiments, the crRNA:tracrRNA duplex binds to a CRISPR effector.
[0138] Ex vivo: As used herein, the term "ex vivo" refers to events that take place in cells or tissues grown outside, rather than within, a multicellular organism.
[0139] Functional equivalent or analog: As used herein, the term "functional equivalent" or "functional analog" in the context of a functional derivative of an amino acid sequence refers to a molecule that retains substantially similar biological activity (either function or structure) as that of the original sequence. A functional derivative or equivalent may be a natural derivative or may be synthetically prepared. Exemplary functional derivatives include amino acid sequences having one or more amino acid substitutions, deletions, or additions, provided that the biological activity of the protein is preserved. The substituting amino acid desirably has similar chemical and physical properties as the amino acid that is substituted. Desirable similar chemical and physical properties include charge similarity, bulkiness, hydrophobicity, hydrophilicity, and the like.
[0140] Half-life: As used herein, the term "half-life" is the time required for a quantity, such as a protein concentration or activity, to fall to half of its value measured at the beginning of the period.
[0141] Improve, increase, or reduce: As used herein, the terms "improve," "increase," or "reduce," or grammatical equivalents, refer to a value relative to a baseline measurement, such as a measurement in the same individual prior to the initiation of a treatment described herein, or a measurement in a control subject (or control subjects) in the absence of a treatment described herein. A "control subject" is a subject suffering from the same form of disease as the subject being treated, and who is about the same age as the subject being treated.
[0142] Inhibition: As used herein, the terms "inhibit," "inhibit," and "inhibiting" refer to a process or method of decreasing or reducing the activity and / or expression of a protein or gene of interest. Typically, inhibiting a protein or gene refers to reducing the expression or associated activity of the protein or gene by at least 10% or more, e.g., 20%, 30%, 40%, or 50%, 60%, 70%, 80%, 90% or more, or a greater than 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 50-fold, 100-fold or more decrease in expression or associated activity as measured by one or more methods described herein or recognized in the art.
[0143] Hybridization: As used herein, the term "hybridization" refers to a reaction in which two or more nucleic acids bind to one another through hydrogen bonding by Watson-Crick pairing, Hoogsteen binding, or other sequence-specific bonds between the bases of the two nucleic acids. A sequence that can hybridize to another sequence is called the "complement" of the sequence and is said to be "complementary" or to exhibit "complementarity."
[0144] Indel: As used herein, the term "indel" refers to an insertion or deletion of a base in a nucleic acid sequence, which generally results in a mutation and is a common form of genetic variation.
[0145] In vitro: As used herein, the term "in vitro" refers to events that take place not in a multicellular organism but in an artificial environment, e.g., in a test tube or reaction vessel, in cell culture, etc.
[0146] In vivo: As used herein, the term "in vivo" refers to events that occur within multicellular organisms, such as humans and non-human animals. In the context of cell-based systems, the term can be used to refer to events that occur within living cells (as opposed to, for example, in vitro systems).
[0147] Linker: The term "linker" refers to any means, entity, or moiety used to connect two or more entities. In some embodiments, the linker is a covalent linker. In some embodiments, the linker is a non-covalent linker. Examples of covalent linkers include covalent bonds or linker moieties covalently attached to one or more of the proteins or domains to be connected. In some embodiments, the linker is a non-covalent bond, e.g., an organometallic bond through a metal center, such as a platinum atom. The bond can be permanent or reversible. For covalent attachment, a variety of functional groups can be used, such as amide groups, including carbonic acid derivatives, esters, including ethers, organic and inorganic esters, amino, urethane, urea, etc. To provide the bond, the domains can be modified by oxidation, hydroxylation, substitution, reduction, etc. to provide sites for coupling. Methods for conjugation are well known by those of skill in the art and are encompassed for use in the present invention. Linker moieties include, but are not limited to, chemical linker moieties, or, for example, peptide linker moieties (linker sequences). It will be appreciated that modifications that do not significantly reduce the function of the RNA binding and effector domains are preferred.
[0148] Mutation: As used herein, the term "mutation" has its normal meaning in the art and includes, for example, point mutations, substitutions, insertions, deletions, inversions, and deletions.
[0149] Oligonucleotide: As used herein, the term "oligonucleotide" generally refers to a polynucleotide of single or double stranded DNA of about 5 to about 100 nucleotides. Oligonucleotides are also known as "oligomers" or "oligos" and may be isolated from genes or chemically synthesized.
[0150] PAM: The term "PAM" or "protospacer adjacent motif" refers to a short nucleic acid sequence (usually 2-6 base pairs in length) that follows a nucleic acid region that is targeted for cleavage by a CRISPR system, such as CRISPR-Cas9. The PAM is required for Cas nucleases to cleave and is generally found 3-4 nucleotides downstream from the cleavage site.
[0151] Polypeptide: The term "polypeptide" as used herein refers to a continuous chain of amino acids linked together via peptide bonds. Although the term is used to refer to an amino acid chain of any length, one of skill in the art will understand that the term is not limited to long chains and can refer to a minimum chain that includes two amino acids linked together via a peptide bond. As known to those of skill in the art, polypeptides can be processed and / or modified. As used herein, the terms "polypeptide" and "peptide" are used interchangeably.
[0152] Prevent: As used herein, the terms "prevent" or "prevention," when used in relation to the occurrence of a disease, disorder, and / or condition, refers to reducing the risk of developing the disease, disorder, and / or condition.
[0153] Protein: The term "protein" as used herein refers to one or more polypeptides that function as separate units. When a single polypeptide is a separate functional unit and does not require permanent or temporary physical association with other polypeptides to form a separate functional unit, the terms "polypeptide" and "protein" may be used interchangeably. When a separate functional unit consists of more than one polypeptide that is physically associated with one another, the term "protein" refers to multiple polypeptides that are physically coupled and function together as a separate unit.
[0154] Reference: A "reference" entity, system, amount, set of conditions, etc., that is compared to a test entity, system, amount, set of conditions, etc., as described herein. For example, in some embodiments, a "reference" antibody is a control antibody that has not been engineered as described herein.
[0155] RNA guide: The term RNA guide refers to an RNA molecule that facilitates targeting of a protein described herein to a target nucleic acid. Exemplary "RNA guides" or "guide RNAs" include, but are not limited to, crRNAs or crRNAs in combination with a cognate tracrRNA. The latter may be independent RNAs or fused as a single RNA using a linker (sgRNA). In some embodiments, the RNA guide is engineered to include chemical or biochemical modifications, and in some embodiments, the RNA guide may include one or more nucleotides.
[0156] Subject: The term "subject," as used herein, means any subject for which diagnosis, prognosis, or therapy is desired. For example, the subject can be a mammal, such as a human or a non-human primate (such as an ape, monkey, orangutan, or chimpanzee), dog, cat, guinea pig, rabbit, rat, mouse, horse, cattle, or cow.
[0157] sgRNA: The term "sgRNA" or "single guide RNA" refers to a single guide RNA that contains (i) a guide sequence (crRNA sequence) and (ii) a Cas9 nuclease recruitment sequence (tracrRNA).
[0158] Substantial identity: The phrase "substantial identity" is used herein to refer to a comparison between amino acid or nucleic acid sequences. As will be understood by those skilled in the art, two sequences are generally considered to be "substantially identical" if they contain identical residues at corresponding positions. As is well known in the art, amino acid or nucleic acid sequences can be compared using any of a variety of algorithms, including those available in commercially available computer programs such as BLASTN for nucleotide sequences, and BLASTP, gapped BLAST, and PSI-BLAST for amino acid sequences. Exemplary such programs are described in Altschul, et al., Basic local alignment search tool, J. Mol. Biol., 215(3):403-410, 1990; Altschul, et al., Methods in Enzymology; Altschul et al., Nucleic Acids Res. 25:3389-3402, 1997; Baxevanis et al., Bioinformatics: A Practical Guide to the Analysis of Genes and Proteins, Wiley, 1998; and Misener, et al., (eds.), Bioinformatics Methods and Protocols (Methods in Molecular Biology, Vol. 132), Humana Press, 1999. In addition to identifying identical sequences, the above-mentioned programs typically provide an indication of the degree of identity. In some embodiments, two sequences are considered to be substantially identical if at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of the corresponding residues are identical over the relevant stretch of residues, in some embodiments the relevant stretch is the complete sequence.In some embodiments the relevant extensions are at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more residues.
[0159] Target nucleic acid: The term "target nucleic acid" as used herein refers to any length of nucleotide (oligonucleotide or polynucleotide), deoxyribonucleotide, ribonucleotide, or any of their analogs to which the CRISPR-Cas9 system binds. Target nucleic acids may have a three-dimensional structure that may include coding or non-coding regions, and may include exons, introns, mRNA, tRNA, rRNA, siRNA, shRNA, miRNA, ribozymes, cDNA, plasmids, vectors, exogenous sequences, endogenous sequences. Target nucleic acids may include modified nucleotides, methylated nucleotides, or nucleotide analogs. Target nucleic acids may be interspersed with non-nucleic acid components. Target nucleic acids are, but are not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers that include purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0160] Therapeutically effective amount: As used herein, the term "therapeutically effective amount" refers to an amount of a therapeutic molecule (e.g., an engineered antibody described herein) that confers a therapeutic effect on a treated subject at a reasonable benefit / risk ratio applicable to any medical treatment. The therapeutic effect may be objective (i.e., measurable by some test or marker) or subjective (i.e., the subject gives an indication of or feels an effect). In particular, a "therapeutically effective amount" refers to an amount of a therapeutic molecule or composition that is effective to treat, ameliorate, or prevent a particular disease or condition, or that is effective to exhibit a detectable therapeutic or prophylactic effect, such as by improving symptoms associated with the disease, preventing or delaying the onset of the disease, and / or reducing the severity or frequency of symptoms of the disease. Therapeutically effective amounts can be administered in a dosing regimen that may include multiple unit doses. For any particular therapeutic molecule, the therapeutically effective amount (and / or the appropriate unit dose within an effective dosing regimen) may vary depending, for example, on the route of administration, combination with other drugs. In addition, the particular therapeutically effective amount (and / or unit dose) for any particular subject may depend on a variety of factors, including the disorder being treated and the severity of the disorder; the activity of the particular agent used; the particular composition used; the age, weight, general health, sex, and diet of the subject; the time of administration, route of administration, and / or excretion or metabolic rate of the particular therapeutic molecule used; the duration of treatment; and similar factors well known in the medical arts.
[0161] tracrRNA: The term "tracrRNA" or "transactivating crRNA" as used herein refers to an RNA that contains a sequence that forms the structure required for a CRISPR-associated protein to bind to a specified target nucleic acid.
[0162] Treatment: As used herein, the term "treatment" (also "treat" or "treating") refers to any administration of a therapeutic molecule (e.g., a CRISPR-Cas therapeutic protein or system described herein) that partially or completely alleviates, ameliorates, relieves, inhibits, delays the onset of, reduces the severity of, and / or reduces the incidence of, one or more symptoms or characteristics of a particular disease, disorder, and / or condition. Such treatment may be of subjects who do not show signs of the relevant disease, disorder, and / or condition and / or who show only early signs of the disease, disorder, and / or condition. Alternatively or additionally, such treatment may be of subjects who show one or more established signs of the relevant disease, disorder, and / or condition.
[0163] The drawings are for illustration purposes only, and not for limitation. [Brief description of the drawings]
[0164] [Figure 1A] 1 is a graph showing the consensus PAM motif recognized by human codon-optimized Streptococcus equinus ATCC 33317 Cas9. [Figure 1B] 1 is a graph showing the consensus PAM motif recognized by human codon-optimized Enterococcus hirae strain F1129E Cas9. [Figure 1C] 1 is a graph showing the consensus PAM motif recognized by human codon-optimized Streptococcus equinus strain AG46 Cas9. [Figure 1D] 1 is a graph showing the consensus PAM motif recognized by human codon-optimized Staphylococcus simulans strain 19. [Figure 1E] 1 is a graph showing the consensus PAM motif recognized by human codon-optimized Streptococcus intermedius B196 strain G1552. [Figure 1F]1 is a graph showing the consensus PAM motif recognized by human codon-optimized Streptococcus sanguinis SK330. [Figure 1G] 1 is a graph showing the consensus PAM motif recognized by human codon-optimized Streptococcus sp. C150. [Figure 1H] 1 is a graph showing the consensus PAM motif recognized by human codon-optimized Streptococcus oralis subsp. oralis strain RH_1735_08. [Figure 1I] 1 is a graph showing the consensus PAM motif recognized by Streptococcus oralis SK313. [Figure 1J] 1 is a graph showing the consensus PAM motifs recognized by Staphylococcus warneri strain 691. [Figure 1K] 1 is a graph showing the consensus PAM motif recognized by Staphylococcus sciuri strain SNUC 2430. [Figure 1L] 1 is a graph showing the consensus PAM motif recognized by Streptococcus gallolyticus strain AM24-4. [Figure 1M] 1 is a graph showing the consensus PAM motifs recognized by Lactobacillus kullabergensis strain Biut2. [Figure 1N] FIG. 1 is a graph showing the consensus PAM motifs recognized by Streptococcus suis strains (N=A, T, G, or C; H=A, C, or T; R=A or G). [Figure 2A] 2A shows a schematic diagram of the predicted RNA folding structure of sgRNA for human codon-optimized Streptococcus equinus ATCC 33317 (Seq2Cas9) using Geneious software. [Figure 2B]FIG 2B is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Enterococcus hirae strain F1129E Cas9 (EhiCas9) using Geneious software. [Figure 2C] FIG 2C is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus equinus strain AG46 Cas9 (SeqCas9) using Geneious software. [Figure 2D] 2A-D are schematic diagrams showing the predicted RNA fold structure of sgRNA for human codon-optimized Staphylococcus simulans strain 19 (SsiCas9) using Geneious software. [Figure 2E] FIG 2E is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus intermedius B196 strain G1552 (SinCas9) using Geneious software. [Figure 2F] FIG 2F is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus sanguinis SK330 (SsaCas9) using Geneious software. [Figure 2G] FIG. 2G is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus sp. C150 (Ssc2Cas9) using Geneious software. FIG. 2G shows the sgRNA comprising SEQ ID NO:49. [Figure 2H]FIG. 2H is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus oralis subsp. oralis strain RH_1735_08 (Sor2Cas9) using Geneious software. FIG. 2H shows the sgRNA comprising SEQ ID NO:50. [Figure 2I] FIG. 2I is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus oralis SK313 (SorCas9) using Geneious software. FIG. 2I shows the sgRNA comprising SEQ ID NO:51. [Figure 2J] FIG. 2J is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Staphylococcus warneri strain 691 (SwaCas9) using Geneious software. FIG. 2J shows the sgRNA comprising SEQ ID NO:52. [Figure 2K] FIG. 2K is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Staphylococcus sciuri strain SNUC 2430 using Geneious software. [Figure 2L] FIG. 2A is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus gallolyticus strain AM24-4 (SgaCas9) using Geneious software. FIG. 2B shows the sgRNA comprising SEQ ID NO:54. [Figure 2M] FIG 2M is a schematic showing the predicted RNA fold structure of the sgRNA for human codon-optimized Lactobacillus kullabergensis strain Biut2 (LkuCas9) using Geneious software. [Figure 2N]FIG 2N shows the predicted RNA fold structure of the sgRNA for human codon-optimized Streptococcus suis strain LSS83 using Geneious software. [Figure 3A] 1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the Seq2Cas9 D10A mutant. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 8) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1A). [Figure 3B] 1 is a graph showing the results of the adenine to guanine base (A to G) conversion percentage achieved with a base editor containing an ABE fused to the N-terminus of the EhiCas9 D10A mutant. The A to G conversion percentage (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 9) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1B). [Figure 3C] 1 is a graph showing the results of the adenine to guanine base (A to G) conversion percentage achieved with a base editor containing an ABE fused to the N-terminus of the SeqCas9 D10A mutant. The A to G conversion percentage (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 10) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1C). [Figure 3D] 1 is a graph showing the results of the adenine to guanine base (A to G) conversion percentage achieved with a base editor containing an ABE fused to the N-terminus of the SsiCas9 D10A mutant. The A to G conversion percentage (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 11) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1D). [Figure 3E]1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the SinCas9 D9A mutant. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 12) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1E). [Figure 3F] 1 is a graph showing the results of the adenine to guanine base (A to G) conversion percentage achieved with a base editor containing an ABE fused to the N-terminus of the SsaCas9 D11A mutant. The A to G conversion percentage (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 13) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1F). [Figure 3G] 1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the Ssc2Cas9 D9A mutant. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 14) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1G). [Figure 3H] 1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the Sor2Cas9 D9A mutant. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 15) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1H). [Figure 3I]1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the SorCas9 D9A mutant. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 16) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1I). [Figure 3J] 1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with base editors containing an ABE fused to the N-terminus of SwaCas9 variants. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 17) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1J). [Figure 3K] 1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with base editors containing an ABE fused to the N-terminus of SscCas9 mutants. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 18) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1K). [Figure 3L] 1 is a graph showing the results of the adenine to guanine base (A to G) conversion percentage achieved with base editors containing an ABE fused to the N-terminus of SgaCas9 variants. The A to G conversion percentage (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 19) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1L). [Figure 3M]1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with base editors containing an ABE fused to the N-terminus of LkuCas9 variants. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 20) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1M). [Figure 3N] 1 is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with base editors containing an ABE fused to the N-terminus of SsuCas9 variants. The percentage A to G conversion (y-axis) is plotted for various guide RNAs targeting A-rich genomic test sites (x-axis, Table 21) adjacent to sequences corresponding to the PAM consensus motif (see FIG. 1N). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0165] Clustered regularly interspaced short palindromic repeats (CRISPR) were first discovered in bacteria and archaea as an adaptive immune system, and then engineered to generate targeted DNA breaks in living cells and organisms. A variety of DNA changes can be introduced during the cellular DNA repair process. The diverse and expanding CRISPR toolbox allows for programmable genome editing, epigenome editing, and transcriptome modulation.
[0166] CRISPR-Cas systems include three main types (I, II, and III) based on their Cas gene organization, as well as the sequences and structures of the component proteins. Each of the three CRISPR systems is characterized by a unique Cas gene: Cas3, a target degradation nuclease / helicase in type I; Cas9, an RNA binding and target degradation nuclease in type II; and Cas10, a large protein for multiple functions in type III. The three CRISPR types also differ in their associated effector complexes. Type I Cas systems are associated with a cascading effector complex, type II effector complexes consist of a single Cas9 and one or more RNA molecules, and type III interference complexes are further divided into type III-A (DNA-targeting Csm complex) and type III-B (RNA-targeting Cmr complex). Cas proteins are key components of effector complexes in all CRISPR-Cas systems.
[0167] Current genome editing technologies focus on class II CRISPR-Cas systems, which contain a single protein effector nuclease for DNA cleavage, specifically, Cas9, a dual RNA-guided nuclease that requires both the CRISPR RNA (crRNA) and tracrRNA and contains both HNH and RuvC nuclease domains, and Cas12a, a single RNA-guided nuclease that requires only the crRNA and contains a single RuvC domain.
[0168] Various aspects of the invention are described in detail in the following sections. The use of the sections is not intended to limit the invention. Each section may be applicable to any aspect of the invention. In this application, the use of "or" means "and / or" unless otherwise specified.
[0169] Engineered, non-naturally occurring Cas9 proteins Streptococcus equinus ATCC 33317 (Seq2Cas9), Enterococcus hirae strain F1129E (EhiCas9), Streptococcus equinus strain AG46 (SeqCas9), Staphylococcus simulans strain 19 (SsiCas9), Streptococcus intermedius B196 strain G1552 (SinCas9), Streptococcus sanguinis SK330 (SsaCas9), Streptococcus sp.C150 (Ssc2Cas9), Streptococcus oralis subsp.oralis strain RH_1735_08 (Sor2Cas9), Streptococcus oralis SK313 (SorCas9), Staphylococcus warneri strain 691 (SwaCas9), Staphylococcus sciuri strain SNUC Described herein are engineered, non-naturally occurring Cas9 proteins modified from WT Cas9 obtained from Lactobacillus kullabergensis strain Biut2 (LkuCas9), and Streptococcus suis strain LSS83 (SsuCas9).
[0170] In some embodiments, the engineered non-naturally occurring Cas9 proteins described herein contain an amino acid sequence at least 60% (e.g., 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. In some embodiments, the Cas9 protein has 80% identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. In some embodiments, the amino acid sequence of the Cas9 protein is identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. Exemplary Cas9 amino acid sequences are provided in Table 1 below. [Table 1] TIFF2024546841000103.tif201170TIFF2024546841000104.tif225170TIFF2024546841000105.tif204170 TIFF2024546841000106.tif199170TIFF2024546841000107.tif199170TIFF2024546841000108.tif225170
[0171] In some embodiments, the Cas9 protein comprises one or more mutations with reference to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. For example, the amino acid sequence of the Cas9 protein comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 mutations in SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14. A variety of mutations are known in the art and include, for example, amino acid substitutions.
[0172] In some embodiments, two or more catalytic domains of Cas9 (RuvC1, RuvCII, RuvCIII) are mutated to produce an inactive or "dead" Cas9 (dCas9) that lacks nucleic acid cleavage activity. In some embodiments, one or more mutations are in the PAM interaction, HNH domain, and / or RuvC domain. In some embodiments, Cas9 is mutated to reduce DNA cleavage activity to about 25%, 15%, 10%, 5%, 1%, 0.1%, 0.01% or less relative to its non-mutated form.
[0173] In some embodiments, a nickase mutant version of Cas9 is provided. In some embodiments, the nickase mutant has one or more amino acid substitutions in the RuvC domain and / or the HNH domain. Various nickase mutations are known for SpCas9 (Streptococcus pyogenes), including, for example, mutations at one or more of amino acid positions 10, 12, 17, 762, 840, 854, 863, 982, 983, 984, 986, 987 of wild-type SpCas9. For example, an aspartic acid to alanine substitution corresponding to D10A in SpCas9 results in the generation of a nickase. In some embodiments, the Cas9 described herein has one or more mutations that result in the generation of a nickase. In some embodiments, the Cas9 described herein has one or more mutations at amino acid positions corresponding to one or more of amino acids 10, 12, 17, 762, 840, 854, 863, 982, 983, 984, 986, 987 of SpCas9.
[0174] In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of Seq2Cas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of EhiCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of SeqCas9 (e.g., corresponding to D10A in SpCas9). In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of SsiCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D9A) in the RuvC domain of SinCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D11A) in the RuvC domain of SsaCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D9A) in the RuvC domain of Ssc2Cas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D9A) in the RuvC domain of Sor2Cas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D9A) in the RuvC domain of SorCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of SwaCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of SscCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of SgaCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D13A) in the RuvC domain of LkuCas9. In some embodiments, the mutation is an aspartic acid to alanine substitution (D10A) in the RuvC domain of SsuCas9.
[0175] In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of Seq2Cas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of EhiCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of SeqCas9 (e.g., corresponding to D10G in SpCas9). In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of SsiCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D9G) in the RuvC domain of SinCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D11G) in the RuvC domain of SsaCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D9G) in the RuvC domain of Ssc2Cas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D9G) in the RuvC domain of Sor2Cas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D9G) in the RuvC domain of SorCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of SwaCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of SscCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of SgaCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D13G) in the RuvC domain of LkuCas9. In some embodiments, the mutation is an aspartic acid to glycine substitution (D10G) in the RuvC domain of SsuCas9.
[0176] In some embodiments, one or more such mutations described herein convert Cas9 into an inactive, or "dead" version of Cas9 (dCas9). Thus, in some embodiments, the Cas9 protein contains one or more mutations that inhibit the ability of Cas9 to cleave both strands of a DNA duplex.
[0177] In some embodiments, when co-expressed with a guide RNA, dead Cas9 generates a DNA recognition complex that can specifically interfere with transcription elongation, RNA polymerase binding, or transcription factor binding. In some embodiments, dead Cas9 is used to specifically target effector proteins of various functions to specific nucleic acid target sites.
[0178] In some embodiments, the engineered non-naturally occurring Cas9 is codon-optimized for human cells. The engineered non-naturally occurring Cas9 is at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) identical to SEQ ID NO:1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14.
[0179] Exemplary Cas9 amino acid sequences with nuclear localization signals (NLS) and linkers are provided in Table 2 below. [Table 2] TIFF2024546841000110.tif239170TIFF2024546841000111.tif217170TIFF2024546841000112.tif246170TIFF2024546841000113.tif216170TIFF2024546841000114.tif206170TIFF2024546841000115.tif240170TIFF2024546841000116.tif104170NLS (bold) can be replaced with a different NLS The linker (underlined) can be removed or extended The 3xHA tag (in italics) can be replaced by a different tag.
[0180] In some embodiments, the engineered non-naturally occurring human codon optimized Cas9 comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99% sequence identity to SEQ ID NO: 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28. In some embodiments, the engineered non-naturally occurring human codon optimized Cas9 comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99% sequence identity to SEQ ID NO: 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28. In some embodiments, the engineered non-naturally occurring human codon optimized Cas9 comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99% sequence identity to SEQ ID NO: 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28. In some embodiments, the engineered non-naturally occurring human codon optimized Cas9 comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99% sequence identity to SEQ ID NO: 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 or 28.
[0181] Various species exhibit codon bias (i.e., differences in codon usage by organisms) that correlate with the efficiency of messenger RNA (mRNA) translation by utilizing codons in mRNAs that correspond to the abundance of tRNA species for that codon in a particular organism. Various methods in the art can be used for computer optimization, including, for example, through the use of software. In some embodiments, codon optimization refers to the modification of a nucleic acid sequence for enhanced expression in a host cell of interest by replacing at least one codon (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of a native sequence with a codon that is more frequently or most frequently used in the host cell's genes while maintaining the native amino acid sequence.
[0182] In some embodiments, the Cas9 proteins described herein are codon-optimized. This type of optimization is known in the art and involves the mutation of exogenous DNA to mimic the codon preferences of the intended host organism or cell while still encoding the same protein. Thus, the codons are altered, but the encoded protein remains unchanged. Codon optimization improves soluble protein levels in a given species, and increases activity and editing efficiency. Codon optimization also results in increased translation and protein expression.
[0183] In some embodiments, the Cas9 protein is codon-optimized for expression in eukaryotic cells. In some embodiments, the Cas9 protein is codon-optimized for expression in human cells.
[0184] Protospacer adjacent motif (PAM) Each Cas endonuclease binds to its target sequence only in the presence of a specific sequence, known as a protospacer adjacent motif (PAM), on the non-targeted, i.e., complementary, DNA strand. Cas nucleases isolated from different bacterial species recognize different PAM sequences. For example, engineered Cas9 from Streptococcus equinus ATCC 33317 cleaves upstream of the consensus PAM sequence 5'-NRGNR-3' (where "N" can be any nucleotide base and R is A or G); Enterococcus hirae strain F1129E recognizes the consensus PAM sequence 5'-NRG-3'; Streptococcus equinus strain AG46, Staphylococcus warneri strain 691, and Staphylococcus sciuri strain SNUC 2430 recognize the consensus PAM sequence 5'-NNGR-3'; Staphylococcus simulans strain 19 recognizes the consensus PAM sequence 5'-NNGRRT-3'; Streptococcus intermedius B196 strain G1552 recognizes the consensus PAM sequence 5'-NNAAAA-3'; and Streptococcus sanguinis SK330 recognizes the consensus PAM sequence 5'-NGGNG-3', Streptococcus sp. C150 recognizes the consensus PAM sequence 5'-NNGNRG-3', Streptococcus oralis subsp. oralis strain RH_1735_08 recognizes the consensus PAM sequence 5'-NNAAAC-3', Streptococcus oralis SK313 recognizes the consensus PAM sequence 5'-NNRAAG-3', Streptococcus gallolyticus strain AM24-4 recognizes the consensus PAM sequence 5'-NNAYAA-3', Lactobacillus kullabergensis strain Biut2 recognizes the consensus PAM sequence 5'-NNGAAA-3', and Streptococcus suis strains recognize the consensus PAM sequence 5'-NNAAA-3' (N=any nucleotide base; H=A, C or T; R=A or G) in the target.Thus, the locations in the genome that can be targeted by different Cas proteins are limited by the location of their unique PAM sequences.
[0185] In some embodiments, the target nucleic acid is 5' or upstream of the PAM sequence. Thus, the Cas9 proteins described herein exhibit activity, e.g., binding, cleavage, modification, or altered gene expression, in the presence of a unique PAM sequence.
[0186] In some embodiments, the Cas9 proteins described herein do not bind to or exhibit activity with any other PAM sequences.
[0187] RNA guide The RNA guide comprises a polynucleotide sequence having complementarity with the target sequence. The RNA guide hybridizes with the target nucleic acid sequence and directs the sequence-specific binding of the CRISPR complex to the target nucleic acid. In some embodiments, the RNA guide has 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% complementarity with the target nucleic acid sequence.
[0188] In some embodiments, the RNA guide is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75 or more nucleotides in length. In some embodiments, the RNA guide is about 18-24 nucleotides in length. In some embodiments, the RNA guide is complementary to about 18-24 nucleotides in the target nucleic acid sequence. For example, the RNA guide is complementary to about 18, 19, 20, 21, 22, 23, or 24 nucleotides in the target nucleic acid sequence. In some embodiments, the RNA guide is complementary to about 18-22 nucleotides. In some embodiments, the RNA guide is complementary to about 18-21 nucleotides. In some embodiments, the RNA guide is complementary to about 18-20 nucleotides. In some embodiments, the RNA guide is complementary to 20 nucleotides in the target nucleic acid sequence.
[0189] The RNA guide can be designed to target any target sequence. Optimal alignment is determined using any algorithm for aligning sequences, including the Needleman-Wunsch algorithm, the Smith-Waterman algorithm, the Burrows-Wheeler algorithm, ClustlW, ClustlX, BLAST, Novoalign, SOAP, Maq, and ELAND.
[0190] In some embodiments, the RNA guide is targeted to a unique target sequence in the genome of the cell. In some embodiments, the RNA guide is designed to lack a PAM sequence. In some embodiments, the RNA guide sequence is designed to have an optimal secondary structure using folding algorithms including mFold or Geneious. In some embodiments, the expression of the RNA guide can be under an inducible promoter, for example, hormone-inducible, tetracycline or doxycycline-inducible, arabinose-inducible, or light-inducible.
[0191] In some embodiments, the CRISPR system comprises one or more RNA guides, such as crRNA, tracrRNA, and / or sgRNA. Thus, in some embodiments, the RNA guide comprises a crRNA. In some embodiments, the RNA guide comprises a tracrRNA. In some embodiments, the RNA guide comprises an sgRNA. In some embodiments, the CRISPR system comprises a plurality of RNA guides, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or more RNA guides.
[0192] In some embodiments, the RNA guide comprises a crRNA. In some embodiments, the CRISPR system comprises a plurality of crRNAs, including 2-15 crRNAs. In some embodiments, the crRNA is a precursor crRNA (pre-crRNA), which comprises a direct repeat sequence, a spacer sequence, and a direct repeat sequence. In some embodiments, the crRNA is a processed or mature crRNA, which comprises a truncated direct repeat sequence.
[0193] In some embodiments, the CRISPR-associated proteins cleave the pre-crRNA to form the processed or mature crRNA.
[0194] In some embodiments, the CRISPR-associated protein forms a complex with the mature crRNA, and the spacer sequence targets the complex to a complementary sequence in the target nucleic acid. In some embodiments, the RNA guide comprises direct repeats and spacer sequences that can hybridize to the target nucleic acid under appropriate conditions.
[0195] In some embodiments, the spacer length of the crRNA can range from about 15 to 50 nucleotides. In some embodiments, the spacer length of the RNA guide is at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, or at least 22 nucleotides. In some embodiments, the spacer length is 15-17 nucleotides (e.g., 15, 16, or 17 nucleotides), 17-20 nucleotides (e.g., 17, 18, 19, or 20 nucleotides), 20-24 nucleotides (e.g., 20, 21, 22, 23, or 24 nucleotides), 23-25 nucleotides (e.g., 23, 24, or 25 nucleotides), 24-27 nucleotides, 27-30 nucleotides, 30-45 nucleotides (e.g., 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 nucleotides), 30 or 35-40 nucleotides, 41-45 nucleotides, 45-50 nucleotides (e.g., 45, 46, 47, 48, 49, or 50 nucleotides) or more.
[0196] In some embodiments, the RNA guide comprises a direct repeat (DR) sequence that is about 16-26 nucleotides in length. For example, in some embodiments, the DR is about 16 nucleotides in length. In some embodiments, the DR is about 17 nucleotides in length. In some embodiments, the DR is about 18 nucleotides in length. In some embodiments, the DR is about 19 nucleotides in length. In some embodiments, the DR is about 20 nucleotides in length. In some embodiments, the DR is about 21 nucleotides in length. In some embodiments, the DR is about 22 nucleotides in length. In some embodiments, the DR is about 23 nucleotides in length. In some embodiments, the DR is about 24 nucleotides in length. In some embodiments, the DR is about 25 nucleotides in length. In some embodiments, the DR is about 26 nucleotides in length.
[0197] In some embodiments, the crRNA comprises a nucleotide guide sequence and a DR sequence. The nucleotide guide sequence can be about 18-24 nucleotides in length. Thus, in some embodiments, the nucleotide guide sequence is about 18 nucleotides in length. In some embodiments, the nucleotide guide sequence is about 19 nucleotides in length. In some embodiments, the nucleotide guide sequence is about 20 nucleotides in length. In some embodiments, the nucleotide guide sequence is about 21 nucleotides in length. In some embodiments, the nucleotide guide sequence is about 22 nucleotides in length. In some embodiments, the crRNA comprises a nucleotide guide sequence that is about 22 nucleotides in length and a direct repeat that is about 22 nucleotides in length.
[0198] In some embodiments, the crRNA sequence may be modified into a "dead crRNA," "dead guide," or "dead guide sequence," which can form a complex with CRISPR-associated proteins and bind to a specific target without any substantial nuclease activity.
[0199] In some embodiments, the crRNA may be chemically modified in the sugar-phosphate backbone or bases. In some embodiments, the crRNA may be modified using 2'O-methyl, 2'-F, or locked nucleic acids to improve nuclease resistance or base pairing. In some embodiments, the crRNA may contain modified bases such as 2-thiouridiene or N6-methyladenosine.
[0200] In some embodiments, the crRNA is conjugated to other oligonucleotides, peptides, proteins, tags, dyes, or polyethylene glycol.
[0201] In some embodiments, the crRNA may contain an aptamer or riboswitch sequence that is capable of binding to a specific target molecule due to its three-dimensional structure.
[0202] In some embodiments, a transactivating RNA (tracrRNA) is associated with the crRNA to facilitate the formation of a complex with the Cas9 protein. In some embodiments, the tracrRNA is about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides in length, or more than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides in length. In some embodiments, the tracrRNA is about 70 nucleotides in length.
[0203] In some embodiments, the tracrRNA and the crRNA are contained in a single transcript called a single guide RNA (sgRNA). In some embodiments, the sgRNA contains a loop between the tracrRNA and the sgRNA.
[0204] In some embodiments, the loop forming sequence is 3, 4, 5 or more nucleotides in length. In some embodiments, the loop has the sequence GAAA, AAAG, CAAA, AAAC, UUUU, UUAUAU, UUA, UUU, and / or AAUCA. In some embodiments, the loop has the sequence GAAA. In some embodiments, the loop has the sequence AAAG. In some embodiments, the loop has the sequence CAAA. In some embodiments, the loop has the sequence AAAC. In some embodiments, the loop has the sequence AAUCA. In some embodiments, the loop has the sequence UUUU. In some embodiments, the loop has the sequence UUAUAU. In some embodiments, the loop has the sequence UUA. In some embodiments, the loop has the sequence UUU. In some embodiments, the loop has the sequence AAUCA.
[0205] In some embodiments, the tracrRNA and the crRNA form a hairpin loop. In some embodiments, the sgRNA has at least two or more hairpins. In some embodiments, the sgRNA has 2, 3, 4, or 5 hairpins.
[0206] In some embodiments, the sgRNA comprises a transcription termination sequence, which comprises a poly-T sequence comprising 6 nucleotides.
[0207] In some embodiments, the sgRNA is 5'- TIFF2024546841000117.tif7170 TIFF2024546841000118.tif7170 (SEQ ID NO: 43) In some embodiments, the sgRNA is 5'- TIFF2024546841000119.tif8170 TIFF2024546841000120.tif9170 (SEQ ID NO: 44) In some embodiments, the sgRNA is 5'- TIFF2024546841000121.tif8170 TIFF2024546841000122.tif8170 (SEQ ID NO: 45) In some embodiments, the sgRNA is 5'- TIFF2024546841000123.tif8170 TIFF2024546841000124.tif9170 (SEQ ID NO: 46) In some embodiments, the sgRNA is 5'- TIFF2024546841000125.tif8170 TIFF2024546841000126.tif9170 (SEQ ID NO: 47) In some embodiments, the sgRNA is 5'- TIFF2024546841000127.tif10170 TIFF2024546841000128.tif9170 (SEQ ID NO: 48) In some embodiments, the sgRNA is 5'- TIFF2024546841000129.tif8170 TIFF2024546841000130.tif8170 (SEQ ID NO: 49) In some embodiments, the sgRNA is 5'- TIFF2024546841000131.tif8170 TIFF2024546841000132.tif10170 (SEQ ID NO: 50) In some embodiments, the sgRNA is 5'- TIFF2024546841000133.tif8170 TIFF2024546841000134.tif8170 (SEQ ID NO: 51) In some embodiments, the sgRNA is 5'- TIFF2024546841000135.tif8170 TIFF2024546841000136.tif8170 (SEQ ID NO: 52) In some embodiments, the sgRNA is 5'- TIFF2024546841000137.tif8170 TIFF2024546841000138.tif7170 (SEQ ID NO: 53) In some embodiments, the sgRNA is 5'- TIFF2024546841000139.tif8170 TIFF2024546841000140.tif10170 (SEQ ID NO: 54) In some embodiments, the sgRNA is 5'- TIFF2024546841000141.tif9170TIFF2024546841000142.tif8170TIFF2024546841000143.tif10170 (SEQ ID NO: 55) In some embodiments, the sgRNA is 5'- TIFF2024546841000144.tif8170 Contains a sequence having at least 80% identity to TIFF2024546841000145.tif10170 (sequence number 56).
[0208] Regarding SEQ ID NOs: 43 to 56, Direct Repetition (italics and underline), tetraloop (italics), tracrRNA (underlined).
[0209] The guide RNA is added to the 5' end of Cas9. In some embodiments, the sgRNA comprises a sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to SEQ ID NO: 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 or 56. In some embodiments, the sgRNA comprises a sequence identical to SEQ ID NO: 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 or 56.
[0210] In some embodiments, the tracrRNA is a separate transcript and is not contained within the same transcript with the crRNA sequence.
[0211] Cas9 fusion protein In some embodiments, the Cas9 enzyme is fused to one or more heterologous protein domains. In some embodiments, the Cas9 enzyme is fused to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more protein domains. In some embodiments, the heterologous protein domain is fused to the C-terminus of the Cas9 enzyme. In some embodiments, the heterologous protein domain is fused to the N-terminus of the Cas9 enzyme. In some embodiments, the heterologous protein domain is fused internally between the C-terminus and N-terminus of the Cas9 enzyme. In some embodiments, the internal fusion is made within the Cas9 RuvCI, RuvC II, RuvCIII, HNH, REC I, or PAM interaction domain.
[0212] The Cas9 protein may be directly or indirectly bound to another protein domain. In some embodiments, a suitable CRISPR system contains a linker or spacer that connects the Cas9 protein and the heterologous protein. The amino acid linker or spacer is generally designed to be flexible or to interpose a structure, such as an alpha-helix, between the two protein moieties. The linker or spacer may be relatively short or may be longer. Typically, the linker or spacer contains, for example, 1 to 100 (e.g., 1 to 100, 5 to 100, 10 to 100, 20 to 100, 30 to 100, 40 to 100, 50 to 100, 60 to 100, 70 to 100, 80 to 100, 90 to 100, 5 to 55, 10 to 50, 10 to 45, 10 to 40, 10 to 35, 10 to 30, 10 to 25, 10 to 20) amino acids in length. In some embodiments, the linker or spacer is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acids long or more. Typically, longer linkers can reduce steric hindrance. In some embodiments, the linker will comprise a mixture of glycine and serine residues. In some embodiments, the linker can further comprise threonine, proline, and / or alanine residues.
[0213] In some embodiments, the Cas9 protein is fused to a cellular localization signal, an epitope tag, a reporter gene, and a protein domain having enzymatic activity, epigenetic modification activity, RNA cleavage activity, nucleic acid binding activity, transcriptional regulation activity, hi some embodiments, the Cas9 protein is fused to a nuclear localization sequence (NLS), a FLAG tag, a HIS tag, and / or an HA tag.
[0214] Suitable fusion partners include, but are not limited to, polypeptides that provide methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, demyristoylating activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, or nuclease activity, any of which can modify DNA or DNA-associated polypeptides (e.g., histones or DNA-binding proteins). In some embodiments, the Cas9 protein is fused to a histone demethylase, transcriptional activator, or deaminase.
[0215] Additional suitable fusion partners include, but are not limited to, boundary elements (e.g., CTCF), proteins and fragments thereof that provide peripheral recruitment (e.g., Lamin A, Lamin B, etc.), and protein docking elements (e.g., FKBP / FRB, Pill / Abyl, etc.).
[0216] In certain embodiments, Cas9 is fused to a cytidine or adenosine deaminase domain, for example, for use in base editing. In some embodiments, the terms "cytidine deaminase" and "cytosine deaminase" can be used interchangeably. In certain embodiments, the cytidine deaminase domain can have 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with any cytidine deaminase described herein. In some embodiments, the cytidine deaminase domain has cytidine deaminase activity (e.g., converting C to U). In certain embodiments, the adenosine deaminase domain may have 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any adenosine deaminase described herein. In some embodiments, the adenosine deaminase domain has adenosine deaminase activity (e.g., converts A to I). In some embodiments, the terms "adenosine deaminase" and "adenine deaminase" can be used interchangeably.
[0217] In some embodiments, the cytidine deaminase can comprise all or part of an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C to U editing enzymes. The N-terminal domain of APOBEC-like proteins is the catalytic domain, and the C-terminal domain is the pseudocatalytic domain. More specifically, the catalytic domain is a zinc-dependent cytidine deaminase domain, which is important for the deamination of cytidine. APOBEC family members include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine or cytosine) deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or part of an APOBEC1 deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC2 deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3 deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3A deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3B deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3C deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3D deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3E deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3F deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3G deaminase.In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC3H deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an APOBEC4 deaminase. In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of an activation-induced deaminase (AID). In some embodiments, the deaminase incorporated into the fusion protein comprises all or a portion of a cytidine deaminase 1 (CDA1). It should be understood that the fusion protein may comprise a deaminase from any suitable organism (e.g., human or rat). In some embodiments, the deaminase domain of the fusion protein is from a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase domain of the fusion protein is from a rat (e.g., rat APOBEC1). In some embodiments, the deaminase domain is human APOBEC1. In some embodiments, the deaminase domain is pmCDA1. Exemplary cytidine deaminase sequences are provided below.
[0218] pmCDA1 (Petromyzon marinus) MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAV (SEQ ID NO: 57) Human AID: MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKAPV (SEQ ID NO: 58) Human AID: (Underlined: nuclear localization sequence, double underlined: nuclear export signal) TIFF2024546841000146.tif15170TIFF2024546841000147.tif9170TIFF2024546841000148.tif7170 (SEQ ID NO: 59) Mouse AID: (Underlined: nuclear localization sequence, double underlined: nuclear export signal) TIFF2024546841000149.tif21170TIFF2024546841000150.tif8170(Sequence number 60) Canine AID: (Underlined: nuclear localization sequence, double underlined: nuclear export signal) TIFF2024546841000151.tif20170TIFF2024546841000152.tif9170 (SEQ ID NO: 61) Bovine AID: (Underlined: nuclear localization sequence, double underlined: nuclear export signal) TIFF2024546841000153.tif21170TIFF2024546841000154.tif9170(SEQ ID NO:62) Rat AID: TIFF2024546841000155.tif30170TIFF2024546841000156.tif9170(SEQ ID NO:63) (Underlined: nuclear localization sequence, double underlined: nuclear export signal) clAID(Canis lupus familiaris): MDSLLMKQRKFLYHFKNVRWAKGRHETYLCYVVKRRDSATSFSLDFGHLRNKSGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGYPNLSLRIFAARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENREKTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 64) btAID(Bos taurus): MDSLLKKQRQFLYQFKNVRWAKGRHETYLCYVVKRRDSPTSFSLDFGHLRNKAGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGYPNLSLRIFTARLYFCDKERKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 65) mAID (Mus musculus): MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL (SEQ ID NO: 66) rAPOBEC-1 (Rattus norvegicus): MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO: 67) maAPOBEC-1 (Mesocricetus auratus): MSSETGPVVVDPTLRRRIEPHEFDAFFDQGELRKETCLLYEIRWGGRHNIWRHTGQNTSRHVEINFIEKFTSERYFYPSTRCSIVWFLSWSPCGECSKAITEFLSGHPNVTLFIYAARLYHHTDQRNRQGLRDLISRGVTIRIMTEQEYCYCWRNFVNYPPSNEVYWPRYPNLWMRLYALELYCIHLGLPPCLKIKRRHQYPLTFFRLNLQSCHYQRIPPHILWATGFI(SEQ ID NO: 68) ppAPOBEC-1(Pongo pygmaeus): MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR(SEQ ID NO: 69) ocAPOBEC1(Oryctolagus cuniculus): MASEKGPSNKDYTLRRRIEPWEFEVFFDPQELRKEACLLYEIKWGASSKTWRSSGKNTTNHVEVNFLEKLTSEGRLGPSTCCSITWFLSWSPCWECSMAIREFLSQHPGVTLIIFVARLFQHMDRRNRQGLKDLVTSGVTVRVMSVSEYCYCWENFVNYPPGKAAQWPRYPPRWMLMYALELYCIILGLPPCLKISRRHQKQLTFFSLTPQYCHYKMIPPYILLATGLLQPSVPWR(SEQ ID NO: 70) mdAPOBEC-1(Monodelphis domestica): MNSKTGPSVGDATLRRRIKPWEFVAFFNPQELRKETCLLYEIKWGNQNIWRHSNQNTSQHAEINFMEKFTAERHFNSSVRCSITWFLSWSPCWECSKAIRKFLDHYPNVTLAIFISRLYWHMDQQHRQGLKELVHSGVTIQIMSYSEYHYCWRNFVDYPQGEEDYWPKYPYLWIMLYVLELHCIILGLPPCLKISGSHSNQLALFSLDLQDCHYQKIPYNVLVATGLVQPFVTWR(SEQ ID NO: 71) ppAPOBEC-2(Pongo pygmaeus): MAQKEEAAAATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKGGQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEELEIQDALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK(SEQ ID NO: 72) btAPOBEC-2(Bos taurus): MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEPEIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK(SEQ ID NO: 73) mAPOBEC-3-(1)(Mus musculus): MQPQRLGPRAGMGPFCLGCSHRKCYSPIRNLISQETFKFHFKNLGYAKGRKDTFLCYEVTRKDCDSPVSLHHGVFKNKDNIHAEICFLYWFHDKVLKVLSPREEFKITWYMS WSPCFECAEQIVRFLATHHNLSLDIFSSRLYNVQDPETQQNLCRLVQEGAQVAAMDLYEFKKCWKKFVDNGGRRFRPWKRLLTNFRYQDSKLQEILRPCYISVPSSSSSTLS NICLTKGLPETRFWVEGRRMDPLSEEEFYSQFYNQRVKHLCYYHRMKPYLCYQLEQFNGQAPLKGCLLSEKGKQHAEILFLDKIRSMELSQVTITCYLTWSPCPNCAWQLAAFKRDRPDLILHIYTSRLYFHWKRPFQKGLCSLWQSGILVDVMDLPQFTDCWTNFVNPKRPFWPWKGLEIISRRTQRRLRRIKESWGLQDLVNDFGNLQLGPPMS (SEQ ID NO: 74) Mouse APOBEC-3-(2): (Italics: nucleic acid editing domain) TIFF2024546841000157.tif49170TIFF2024546841000158.tif10170(SEQ ID NO:75) Rat APOBEC-3: (Italics: nucleic acid editing domain) TIFF2024546841000159.tif49170TIFF2024546841000160.tif9170(SEQ ID NO:76) hAPOBEC-3A (Homo sapiens): MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN (SEQ ID NO: 77) hAPOBEC-3F (Homo sapiens): MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPRLDAKIFRGQVYSQPEHHAEMCFLSWFCGNQLPAYKCFQITWFVSWTPCPDCVAKLAEFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYSEGQPFMPWYKFDDNYAFLHRTLKEILRNPMEAMYPHIFYFHFKNLRKAYGRNESWLCFTMEVVKHHSPVSWKRGVFRNQVDPETHCHAERCFLSWFCDDILSPNTNYEVTWYTSWSPCPEECAGEVAEFLARHSNVNLTIFTARLYYFWDTDYQEGLRSLSQEGASVEIMGYKDFKYCWENFVYNDDEPFKPWKGLKYNFLFLDSKLQEILE (SEQ ID NO: 78) Rhesus APOBEC-3G: (Italics: nucleic acid editing domain, underline: cytoplasmic localization signal) TIFF2024546841000161.tif41170TIFF2024546841000162.tif8170(SEQ ID NO:79) Chimpanzee APOBEC-3G: TIFF2024546841000163.tif42170TIFF2024546841000164.tif8170(sequence number 80) (Italics: nucleic acid editing domain, underline: cytoplasmic localization signal) Green Monkey APOBEC-3G: TIFF2024546841000165.tif41170TIFF2024546841000166.tif7170(SEQ ID NO:81) (Italics: nucleic acid editing domain, underline: cytoplasmic localization signal) Human APOBEC-3G: TIFF2024546841000167.tif41170TIFF2024546841000168.tif7170(SEQ ID NO:82) (Italics: nucleic acid editing domain, underline: cytoplasmic localization signal) Human APOBEC-3F: TIFF2024546841000169.tif35170TIFF2024546841000170.tif7170TIFF2024546841000171.tif8170 (SEQ ID NO: 83) (Italics: nucleic acid editing domain) Human APOBEC-3B: TIFF2024546841000172.tif42170TIFF2024546841000173.tif10170(SEQ ID NO:84) (Italics: nucleic acid editing domain) Rat APOBEC-3B: MQPQGLGPNAGMGPVCLGCSHRRPYSPIRNPLKKLYQQTFYFHFKNVRYAWGRKNNFLCYEVNGMDCALPVPLRQGVFRKQGHIHAELCFIYWFHDKVLRVLSPMEEFKVTWYMSWSPCSKCAEQVARFLAAHRNLSLAIFSSRLYYYLRNPNYQQKLCRLIQEGVHVAAMDLPEFKKCWNKFVDNDGQPFRPWMRLRINFSFYDCKLQEIFSRMNLLREDVFYLQFNNSHRVKPVQNRYYRRKSYLCYQLERANGQEPLKGYLLYKKGEQHVEILFLEKMRSMELSQVRITCYLTWSPCPNCARQLAAFKKDHPDLILRIYTSRLYFWRKKFQKGLCTLWRSGIHVDVMDLPQFADCWTNFVNPQRPFRPWNELEKNSWRIQRRLRRIKESWGL (SEQ ID NO: 85) Bovine APOBEC-3B: MDGWEVAFRSGTVLKAGVLGVSMTEGWAGSGHPGQGACVWTPGTRNTMNLLREVLFKQQFGNQPRVPAPYYRRKTYLCYQLKQRNDLTLDRGCFRNKKQRHAERFIDKINSLDLNPSQSYKIICYITWSPCPNCANELVNFITRNNHLKLEIFASRLYFHWIKSFKMGLQDLQNAGISVAVMTHTEFEDCWEQFVDNQSRPFQPWDKLEQYSASIRRRLQRILTAPI (SEQ ID NO: 86) Chimpanzee APOBEC-3B: MNPQIRNPMEWMYQRTFYYNFENEPILYGRSYTWLCYEVKIRRGHSNLLWDTGVFRGQMYSQPEHHAEMCFLSWFCGNQLSAYKCFQITWFVSWTPCPDCVAKLAKFLAEHPNVTLTISAARLY YYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYNEGQPFMPWYKFDDNYAFLHRTLKEIIRHLMDPDTFTFNFNNDPLVLRRHQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLCGF YGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGQVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQVRASSLCMVPHRPPPPPQSPGPCLPLCSEPPLGSLLPTGRPAPSLPFLLTASFSFPPPASLPPLPSLSLSPGHLPVPSFHSLTSCSIQPPCSSRIRETEGWASVSKEGRDLG (SEQ ID NO: 87) Human APOBEC-3C: TIFF2024546841000174.tif21170TIFF2024546841000175.tif8170(SEQ ID NO:88) (Italics: nucleic acid editing domain) Gorilla APOBEC-3C TIFF2024546841000176.tif21170TIFF2024546841000177.tif8170(SEQ ID NO:89) (Italics: nucleic acid editing domain) Human APOBEC-3A: TIFF2024546841000178.tif22170TIFF2024546841000179.tif9170(Sequence number 90) (Italics: nucleic acid editing domain) Rhesus APOBEC-3A: TIFF2024546841000180.tif21170TIFF2024546841000181.tif10170 (SEQ ID NO: 91) (Italics: nucleic acid editing domain) Bovine APOBEC-3A: TIFF2024546841000182.tif21170TIFF2024546841000183.tif10170(SEQ ID NO:92) (Italics: nucleic acid editing domain) Human APOBEC-3H: TIFF2024546841000184.tif22170TIFF2024546841000185.tif9170 (SEQ ID NO: 93) (Italics: nucleic acid editing domain) Rhesus APOBEC-3H: MALLTAKTFSLQFNNKRRVNKPYYPRKALLCYQLTPQNGSTPTRGHLKNKKKDHAEIRFINKIKSMGLDETQCYQVTCYLTWSPCPSCAGELVDFIKAHRHLNLRIFASRLYYHWRPNYQEGLLLLCGSQVPVEVMGLPEFTDCWENFVDHKEPPSFNPSEKLEELDKNSQAIKRRLERIKSRSVDVLENGLRSLQLGPVTPSSSIRNSR (SEQ ID NO: 94) Human APOBEC-3D: TIFF2024546841000186.tif42170TIFF2024546841000187.tif9170(SEQ ID NO:95) (Italics: nucleic acid editing domain) Human APOBEC-1: MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR (SEQ ID NO: 96) Mouse APOBEC-1: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNTRCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYPPSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK (SEQ ID NO: 97) Rat APOBEC-1: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO: 98) Human APOBEC-2: MAQKEEAAVATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKGGQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEEPEIQAALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK (SEQ ID NO: 99) Mouse APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEVQSKGGQAQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK (SEQ ID NO: 100) Rat APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYLWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK (SEQ ID NO: 101) Bovine APOBEC-2: MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEPEIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK (SEQ ID NO: 102) Petromyzon marinus CDA1(pmCDAl): MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQ LNENRWLEKTLKRAEKRRSELSFMIQVKILHTTKSPAV (SEQ ID NO: 103) Human APOBEC3G D316R D317R: MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKFNYDEFQHCWSKFVYSQRELFEPWNNLPKYYILLHFMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN (SEQ ID NO: 104) Human APOBEC3G A chain: MDPPTFTFNFNNEPWWGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISFTYSEFKHCWDTFVDHQGCPFQPWDGLD EHSQDLSGRLRAILQ (SEQ ID NO: 105) Human APOBEC3G A chain D120R D121R: MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQ (SEQ ID NO: 106) hAPOBEC-4 (Homo sapiens): MEPIYEEYLANHGTIVKPYYWLSFSLDCSNCPYHIRTGEEARVSLTEFCQIFGFPYGTTFPQTKHLTFYELKTSSGSLVQKGHASSCTGNYIHPESMLFEMNGYLDSAIYNNDSIRHIILYSNNSPCNEANHCCISKMYNFLITYPGITLSIYFSQLYHTEMDFPASAWNREALRSLASLWPRVVLSPISGGIWHSVLHSFISGVSGSHVFQPILTGRALADRHNAYEINAITGVKPYFTDVLLQTKRNPNTKAQEALESYPLNNAFPGQFFQMPSGQLQPNLPPDLRAPVVFVLVPLRDLPPMHMGQNPNKPRNIVRHLNMPQMSFQETKDLGRLPTGRSVEIVEITEQFASSKEADEKKKKGKK (SEQ ID NO: 107) mAPOBEC-4 (Mus musculus): MDSLLMKQKKFLYHFKNVRWAKGRHETYLCYVVKRRDSATSCSLDFGHLRNKSGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVAEFLRWNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIGIMTFKDYFYCWNTFVENRERTFKAWEGLHENSVRLTRQLRRILLPLYEVDDLRDAFRMLGF(SEQ ID NO: 108) rAPOBEC-4 (Rattus norvegicus): MEPLYEEYLTHSGTIVKPYYWLSVSLNCTNCPYHIRTGEEARVPYTEFHQTFGFPWSTYPQTKHLTFYELRSSSGNLIQKGLASNCTGSHTHPESMLFERDGYLDSLIFHDSNIRHIILYSNNSPCDEANHCCISKMYNFLMNYPEVTLSVFFSQLYHTENQFPTSAWNREALRGLASLWPQVTLSAISGGIWQSILETFVSGISEGLTAVRPFTAGRTLTDRYNAYEINCITEVKPYFTDALHSWQKENQDQKVWAASENQPLHNTTPAQWQPDMSQDCRTPAVFMLVPYRDLPPIHVNPSPQKPRTVVRHLNTLQLSASKVKALRKSPSGRPVKKEEARKGSTRSQEANETNKSKWKKQTLFIKSNICHLLEREQKKIGILSSWSV(SEQ ID NO: 109) mfAPOBEC-4 (Macaca fascicularis): MEPTYEEYLANHGTIVKPYYWLSFSLDCSNCPYHIRTGEEARVSLTEFCQIFGFPYGTTYPQTKHLTFYELKTSSGSLVQKGHASSCTGNYIHPESMLFEMNGYLDSAIYNNDSIRHIILYCNNSPCNEANHCCISKVYNFLITYPGITLSIYFSQLYHTEMDFPASAWNREALRSLASLWPRVVLSPISGGIWHSVLHSFVSGVSGSHVFQPILTGRALTDRYNAYEINAITGVKPFFTDVLLHTKRNPNTKAQMALESYPLNNAFPGQSFQMTSGIPPDLRAPVVFVLLPLRDLPPMHMGQDPNKPRNIIRHLNMPQMSFQETKDLERLPTRRSVETVEITERFASSKQAEEKTKKKKGKK(SEQ ID NO: 110) pmCDA-1 (Petromyzon marinus): MAGYECVRVSEKLDFDTFEFQFENLHYATERHRTYVIFDVKPQSAGGRSRRLWGYIINNPNVCHAELILMSMIDRHLESNPGVYAMTWYMSWSPCANCSSKLNPWLKNLLEEQGHTLTMHFSRIYDRDREGDHRGLRGLKHVSNSFRMGVVGRAEVKECLAEYVEASRRTLTWLDTTESMAAKMRRKLFCILVRCAGMRESGIPLHLFTLQTPLLSGRVVWWRV(SEQ ID NO: 111) pmCDA-2 (Petromyzon marinus): MELREVVDCALASCVRHEPLSRVAFLRCFAAPSQKPRGTVILFYVEGAGRGVTGGHAVNYNKQGTSIHAEVLLLSAVRAALLRRRRCEDGEEATRGCTLHCYSTYSPCRDCVEYIQEFGASTGVRVVIHCCRLYELDVNRRRSEAEGVLRSLSRLGRDFRLMGPRDAIALLLGGRLANTADGESGASGNAWVTETNVVEPLVDMTGFGDEDLHAQVQRNKQIREAYANYASAVSLMLGELHVDPDKFPFLAEFLAQTSVEPSGTPRETRGRPRGASSRGPEIGRQRPADFERALGAYGLFLHPRIVSREADREEIKRDLIVVMRKHNYQGP (SEQ ID NO: 112) pmCDA-5(Petromyzon marinus): MAGDENVRVSEKLDFDTFEFQFENLHYATERHRTYVIFDVKPQSAGGRSRRLWGYIINNPNVCHAELILMSMIDRHLESNPGVYAMTWYMSWSPCANCSSKLNPWLKNLLEEQGHTLMMHFSRIYDRDREGDHRGLRGLKHVSNSFRMGVVGRAEVKECLAEYVEASRRTLTWLDTTESMAAKMRRKLFCILVRCAGMRESGMPLHLFT (SEQ ID NO: 113) yCD(Saccharomyces cerevisiae): MVTGGMASKWDQKGMDIAYEEAALGYKEGGVPIGGCLINNKDGSVLGRGHNMRFQKGSATLHGEISTLENCGRLEGKVYKDTTLYTTLSPCDMCTGAIIMYGIPRCVVGENVNFKSKGEKYLQTRGHEVVVVDDERCKKIMKQFIDERPQDWFEDIGE (SEQ ID NO: 114) rAPOBEC-1(delta 177-186): MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK (SEQ ID NO: 115) rAPOBEC-1(Delta 202-213): MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQHYQRLPPHILWATGLK (SEQ ID NO: 116) Mouse APOBEC-3: TIFF2024546841000188.tif48170TIFF2024546841000189.tif8170 (SEQ ID NO: 117) (Italics: nucleic acid editing domain)
[0219] In some embodiments, the adenosine deaminase can include all or a portion of the adenosine deaminase ADAR (e.g., ADAR1 or ADAR2). In another embodiment, the adenosine deaminase can include all or a portion of the adenosine deaminase ADAT. In some embodiments, the adenosine deaminase can include all or a portion of the ADAT (EcTadA) from Escherichia coli that includes one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I157F, or a corresponding mutation in another adenosine deaminase. The adenosine deaminase can be derived from any suitable organism (e.g., E. coli). In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from E. coli. In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations corresponding to any of the mutations provided herein (e.g., mutations in ecTadA). Corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determination of homologous residues. Mutations in any naturally occurring adenosine deaminase (e.g., with homology to ecTadA) that correspond to any of the mutations described herein (e.g., any of the mutations identified in ecTadA) can be generated accordingly. In certain embodiments, the TadA is any one of the TadAs described in PCT / US2017 / 045381 (WO2018 / 027078), which is incorporated by reference in its entirety. Mutations with desirable adenosine deaminase activity on single-stranded DNA as shown in Table 3 were identified through rounds of evolution and selection (e.g., TadA*7.10=variant 10 from the 7th round of evolution). [Table 3] TIFF2024546841000191.tif75170
[0220] In some embodiments, TadA is provided as a monomer or a dimer (e.g., a heterodimer of wild-type E. coli TadA and an engineered TadA variant). In some embodiments, the adenosine deaminase is an 8th generation TadA*8 variant as shown in Table 4 below. [Table 4]
[0221] In some embodiments, the adenosine deaminase is a 9th generation TadA*9 variant containing an alteration at an amino acid position selected from: 21, 23, 25, 38, 51, 54, 70, 71, 72, 72, 94, 124, 133, 138, 139, 146, and 158 of the TadA variants shown in the following reference sequences: TIFF2024546841000193.tif58170
[0222] In one embodiment, the adenosine deaminase variant contains modifications at two or more amino acid positions selected from: 21, 23, 25, 38, 51, 54, 70, 71, 72, 94, 124, 133, 138, 139, 146, and 158 of the TadA reference sequence above. In another embodiment, the adenosine deaminase variant contains one or more (e.g., two, three, four) modifications selected from: R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO:1. In other embodiments, the adenosine deaminase variant further contains one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In yet other embodiments, the adenosine deaminase variant contains a combination of modifications to the above TadA reference sequence selected from: E25F+V82S+Y123H, T133K+Y147R+Q154R; E25F+V82S+Y123H+Y147R+Q154R; L51W+V82S+Y123H+C146R+Y147R+Q154R. 4R;Y73S+V82S+Y123H+Y147R+Q154R;P54C+V82S+Y123H+Y147R+Q154R;N38G+V82T+Y123H+Y147 R+Q154R;N72K+V82S+Y123H+D139L+Y147R+Q154R;E25F+V82S+Y123H+D139M+Y147R+Q154R;Q71M +V82S+Y123H+Y147R+Q154R;E25F+V82S+Y123H+T133K+Y147R+Q154R;E25F+V82S+Y123H+Y147R +Q154R;V82S+Y123H+P124W+Y147R+Q154R;L51W+V82S+Y123H+C146R+Y147R+Q154R;P54C+V82S+ Y123H+Y147R+Q154R;Y73S+V82S+Y123H+Y147R+Q154R;N38G+V82T+Y123H+Y147R+Q154R;R23H+ V82S+Y123H+Y147R+Q154R;R21N+V82S+Y123H+Y147R+Q154R;V82S+Y123H+Y147R+Q154R+A158K;N72K+V82S+Y123H+D139L+Y147R+Q154R;E25F+V82S+Y123H+D139M+Y147 R+Q154R;M70V+V82S+M94V+Y123H+Y147R+Q154R;Q71M+V82S+Y123H+Y147 R+Q154R;E25F+I76Y+V82S+Y123H+Y147R+Q154R;I76Y+V82T+Y123H+Y147 R+Q154R;N38G+I76Y+V82S+Y123H+Y147R+Q154R;R23H+I76Y+V82S+Y123H +Y147R+Q154R;P54C+I76Y+V82S+Y123H+Y147R+Q154R;R21N+I76Y+V82S +Y123H+Y147R+Q154R;I76Y+V82S+Y123H+D138M+Y147R+Q154R;Y72S+I76 Y+V82S+Y123H+Y147R+Q154R;E25F+I76Y+V82S+Y123H+Y147R+Q154R;I76 Y+V82T+Y123H+Y147R+Q154R;N38G+I76Y+V82S+Y123H+Y147R+Q154R;R23 H+I76Y+V82S+Y123H+Y147R+Q154R;P54C+I76Y+V82S+Y123H+Y147R+Q154R;R21N+I76Y+V82S+Y123H+Y147R+Q154R;I76Y+V82S+Y123H+D138M+Y147R+Q154R;Y72S+I76Y+V82S+Y123H+Y147R+Q154R;andV82S+Q154R;N72K_V82S+Y123H+Y147R+Q154R;Q71M_V82S+Y123H+Y147R+Q154R;V82S+Y123H+ T133K+Y147R+Q154R;V82S+Y123H+T133K+Y147R+Q154R+A158K;M70V+Q71 M+N72K+V82S+Y123H+Y147R+Q154R;N72K_V82S+Y123H+Y147R+Q154R;Q71 M_V82S+Y123H+Y147R+Q154R;M70V+V82S+M94V+Y123H+Y147R+Q154R;V82 S+Y123H+T133K+Y147R+Q154R;V82S+Y123H+T133K+Y147R+Q154R+A158K;and M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R. In some embodiments, the deaminase or other polypeptide sequence lacks a methionine, for example when included as a component of a fusion protein. This may alter the numbering of the positions. However, one of skill in the art will understand that such corresponding mutations refer to the same mutation (e.g., Y73S and Y72S, and D139M and D138M);
[0223] In some embodiments, Cas9 is fused to a nuclear localization sequence, including the NLS of SV40 large T antigen, nucleoplasmin, c-myc, hRNPA1 M9, the IBB domain from importin-alpha, the NLS of sarcoma T protein, human p53, c-abl IV, influenza virus NS1, hepatitis virus delta antigen, mouse Mx1, human poly(ADP-ribose) polymerase, steroid hormone receptor (human) glucocorticoid.
[0224] In some embodiments, the Cas9 protein is fused to an epitope tag, including, but not limited to, a hemagglutinin (HA) tag, a histidine (His) tag, a FLAG tag, a Myc tag, a V5 tag, a VSV-G tag, a SNAP tag, or a thioredoxin (Trx) tag.
[0225] In some embodiments, Cas9 is fused to a reporter gene, including, but not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol transferase (CAT), HcRed, DsRed, cyan fluorescent protein, yellow fluorescent protein, and blue fluorescent protein, green fluorescent protein (GFP), including enhanced versions or superfolder GFP, and other modified versions of reporter genes.
[0226] In some embodiments, the serum half-life of the engineered Cas9 protein is increased by fusion with a heterologous protein, such as a carboxy-terminal peptide (CTP of chorionic gonadotropin beta chain), human serum albumin protein, transferrin protein, human IgG, and / or a sialylated peptide.
[0227] In some embodiments, the serum half-life of the engineered Cas9 protein is decreased by fusion to a destabilization domain, including, but not limited to, geminin, ubiquitin, FKBP12-L106P, and / or dihydrofolate reductase.
[0228] Suitable fusion partners that provide increased or decreased stability include, but are not limited to, degron sequences. Degrons are readily understood by those skilled in the art to be amino acid sequences that control the stability of the protein of which they are a part. For example, the stability of a protein that includes a degron sequence is controlled at least in part by the degron sequence. In some cases, suitable degrons are constitutive such that the degron exerts its effect on protein stability independent of experimental controls (i.e., the degron is not drug-inducible, temperature-inducible, etc.). In some cases, the degron provides a variant Cas9 polypeptide with controllable stability such that the variant Cas9 polypeptide can be turned "on" (i.e., stable) or "off" (i.e., unstable, degraded) depending on desired conditions. For example, if the degron is a temperature-sensitive degron, the variant Cas9 polypeptide may be functional (i.e., "on," stable) below a threshold temperature (e.g., 42°C, 41°C, 40°C, 39°C, 38°C, 37°C, 36°C, 35°C, 34°C, 33°C, 32°C, 31°C, 30°C, etc.), but may be non-functional (i.e., "off," degraded) above the threshold temperature. As another example, if the degron is a drug-inducible degron, the presence or absence of a drug may switch the protein from an "off (i.e., unstable) state" to an "on" (i.e., stable) state, or vice versa. An exemplary drug-inducible degron is derived from the FKBP12 protein. The stability of the degron is controlled by the presence or absence of a small molecule that binds to the degron.
[0229] Examples of suitable degrons include, but are not limited to, those degrons regulated by Shield-1, DHFR, auxin, and / or temperature. Non-limiting examples of suitable degrons are known in the art (e.g., Dohmen et al., Science, 1994.263(5151):p.1273-1276:Heat-inducible degron: a method for constructing temperature-sensitive mutants; Schoeber et al., Am J Physiol Renal Physiol.2009 Jan;296(l):F204-l l:Conditional fast expression and function of multimeric TRPV5 channels using Shield-1; Chu et al., Bioorg Med Chem Lett.2008 Nov 15;18(22):5941-4:Recent progress with FKBP-derived destabilizing domains; Kanemaki, Pflugers Arch.2012 Dec 28:Frontiers of protein expression control with conditional degrons; Yang et al., Mol Cell.2012 Nov 30;48(4):487-8:Titivated for destruction:the methyl degron, Barbour et al.,Biosci Rep.2013 Jan 18;33(1).:Characterization of the bipartite degron that regulates ubiquitin-independent degradation of thymidylate synthase and Greussing et al.,J Vis Exp.2012 Nov 10;(69):Monitoring of ubiquitin-proteasome activity in living cells using a Degron (dgn)-destabilized green fluorescent protein (GFP)-based reporter protein, all of which are incorporated by reference in their entireties.
[0230] Exemplary degron sequences have been well characterized and tested in both cells and animals, thus fusing a dead Cas9 to a degron sequence produces a "regulatable" and "inducible" dead Cas9 polypeptide.
[0231] Any of the fusion partners described herein can be used in any desired combination. As one non-limiting example to illustrate this point, a Cas9 fusion protein can include a YFP sequence for detection, a degron sequence for stability, and a transcription activator sequence for increasing transcription of target DNA. Furthermore, the number of fusion partners that can be used in a dCas9 fusion protein is unlimited. In some cases, a Cas9 fusion protein includes one or more (e.g., two or more, three or more, four or more, or five or more) heterologous sequences.
[0232] target nucleic acid The target nucleic acid is a DNA molecule, an RNA molecule, which may be single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-naturally occurring, or derivatized nucleotide bases, deoxyribonucleotides, ribonucleotides, or analogs thereof. The target nucleic acid may have a three-dimensional structure that may include coding or non-coding regions, and may include exons, introns, mRNA, tRNA, rRNA, siRNA, shRNA, miRNA, ribozymes, cDNA, plasmids, vectors, exogenous sequences, endogenous sequences. The target nucleic acid may include modified nucleotides, methylated nucleotides, or nucleotide analogs. In some embodiments, the target nucleic acid may be dispersed with non-nucleic acid components.
[0233] The target nucleic acid is recognized by the CRISPR-Cas9 system and binds to Cas9. In some embodiments, it is modified or cleaved or has altered expression due to binding of Cas9. The target nucleic acid contains a specific recognizable PAM motif, such as 5'-NGG-3', 5'-NAGHC-3', 5'-NRHRRH-3', or 5'-NNAAA-3' (H=A, C, or T; R=A or G).
[0234] Recombinant Gene Technology In accordance with the present disclosure there may be employed conventional molecular biology, microbiology, and recombinant DNA techniques within the skill of the art. Such techniques are described in the literature (e.g., Sambrook, Fritsch & Maniatis, Molecular Cloning: A Laboratory Manual, Second Edition (1989) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; DNA Cloning: A Practical Approach, Volumes I and II (DN Glover ed. 1985); Oligonucleotide Synthesis (MJ Gait ed. 1984); Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds. (1985)); Transcription And Translation (B.D. Hames & S.J. Higgins, eds. (1984)); Animal Cell Culture (R.I. Freshney, ed. (1986)); Immobilized Cells and Enzymes (IRL Press, (1986)); B. Perbal, A Practical Guide To Molecular Cloning (1984); F.M.Ausubel et al. (eds.), Current See Protocols in Molecular Biology, John Wiley & Sons, Inc. (1994).
[0235] Recombinant expression of genes, such as nucleic acids encoding polypeptides, such as engineered Cas9 enzymes described herein, can include the construction of expression vectors containing nucleic acids encoding the polypeptides. Once a polynucleotide is obtained, vectors for the production of the polypeptides can be produced by recombinant DNA techniques using techniques known in the art. Using known methods, expression vectors containing polypeptide coding sequences and appropriate transcriptional and translational control signals can be constructed. These methods include, for example, in vitro recombinant DNA techniques, synthetic techniques, and in vivo genetic recombination.
[0236] The expression vector can be transferred to a host cell by conventional techniques and the transfected cells can then be cultured by conventional techniques to produce the polypeptide.
[0237] In some embodiments, the nucleotide sequence encoding the DNA-targeting RNA and / or the Cas9 protein is operably linked to a control element, e.g., a transcription control element, such as a promoter. The transcription control element can be functional in either a eukaryotic cell, e.g., a mammalian cell, or a prokaryotic cell (e.g., a bacterial or archaeal cell). In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the nucleotide sequence encoding the DNA-targeting RNA and / or the novel Cas9 protein is operably linked to multiple control elements that allow for expression of the encoded nucleotide sequence in both prokaryotic and eukaryotic cells.
[0238] The promoter may be a constitutively active promoter (i.e., a promoter that is constitutively active / "ON" state), but may also be an inducible promoter (i.e., a promoter whose state, active / "ON" or inactive / "OFF", is controlled by an external stimulus, e.g., a particular temperature, compound, or the presence of a protein), a spatially restricted promoter (i.e., a transcriptional control element, enhancer, etc.) (e.g., a tissue-specific promoter, a cell type-specific promoter, etc.), or a temporally restricted promoter (i.e., the promoter is in the "ON" or "OFF" state during a particular stage of embryonic development or during a particular stage of a biological process, e.g., the hair follicle cycle in mice).
[0239] Suitable promoters may be derived from a virus, and therefore may be referred to as viral promoters, or they may be derived from any organism, including prokaryotes or eukaryotes. Suitable promoters can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol III). Exemplary promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus long terminal repeat (LTR) promoter, the adenovirus major late promoter (Ad MLP), the herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, e.g., the CMV immediate early promoter region (CMVIE), the Ruth Sarcoma Virus (RSV) promoter, the human U6 micronucleus promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497-500 (2002)), the enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep 1; 31 (17)), and / or the human HI promoter (HI).
[0240] Examples of inducible promoters include, but are not limited to, T7 RNA polymerase promoter, T3 RNA polymerase promoter, isopropyl-beta-D-thiogalactopyranoside (IPTG) regulated promoter, lactose inducible promoter, heat shock promoter, tetracycline regulated promoter (e.g., Tet-ON, Tet-OFF, etc.), steroid regulated promoter, metal regulated promoter, estrogen receptor regulated promoter, etc. Thus, inducible promoters can be regulated by molecules including, but not limited to, doxycycline, RNA polymerase, e.g., T7 RNA polymerase, estrogen receptor and / or estrogen receptor fusions.
[0241] In some embodiments, the promoter is a spatially restricted promoter (i.e., a cell type specific promoter, a tissue specific promoter, etc.), and in a multicellular organism, the promoter is active (i.e., "ON") in a specific subset of cells. A spatially restricted promoter may also be referred to as an enhancer, a transcriptional control element, a control sequence, etc. Any convenient spatially restricted promoter may be used, and the selection of a suitable promoter (e.g., a brain specific promoter, a promoter that drives expression in a subset of neurons, a promoter that drives expression in the germline, a promoter that drives expression in the lung, a promoter that drives expression in muscle, a promoter that drives expression in islet cells of the pancreas, etc.) will depend on the organism. Thus, a spatially restricted promoter can be used to regulate the expression of a nucleic acid encoding a site-specific polypeptide of interest in a wide variety of different tissues and cell types, depending on the organism. Some spatially restricted promoters are also temporally restricted, such that the promoter is in an "ON" or "OFF" state during a particular stage of embryonic development or during a particular stage of a biological process (e.g., the hair follicle cycle).
[0242] For purposes of illustration, examples of spatially restricted promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, photoreceptor-specific promoters, etc. Neuron-specific spatially restricted promoters include, but are not limited to, neuron-specific enolase (NSE) promoter, aromatic amino acid decarboxylase (AADC) promoter, neurofilament promoter, synapsin promoter, thy-1 promoter, serotonin receptor promoter, tyrosine hydroxylase promoter (TH), GnRH promoter, L7 promoter, DNMT promoter, enkephalin promoter, myelin basic protein (MBP) promoter, Ca 2+ Examples include, but are not limited to, the calmodulin-dependent protein kinase II-alpha (CamKIIa) promoter, and / or the CMV enhancer / platelet-derived growth factor-β promoter.
[0243] Adipocyte-specific spatially restricted promoters include, but are not limited to, the aP2 gene promoter / enhancer, e.g., the -5.4 kb to +21 bp region of the human aP2 gene, the glucose transporter-4 (GLUT4) promoter, the fatty acid translocase (FAT / CD36) promoter, the stearoyl-CoA desaturase-1 (SCD1) promoter, the leptin promoter, and the adiponectin promoter, the adipsin promoter and / or the resistin promoter.
[0244] Cardiomyocyte-specific, spatially restricted promoters include, but are not limited to, control sequences derived from the following genes: myosin light chain-2, a-myosin heavy chain, AE3, cardiac troponin C, and / or cardiac actin.
[0245] Smooth muscle-specific spatially restricted promoters include, but are not limited to, the SM22a promoter, the smoothelin promoter, and / or the a-smooth muscle actin promoter.
[0246] Photoreceptor-specific spatially restricted promoters include, but are not limited to, a rhodopsin promoter, a rhodopsin kinase promoter, a beta phosphodiesterase gene promoter, a retinitis pigmentosa gene promoter, an interphotoreceptor retinoid binding protein (IRBP) gene enhancer, and / or an IRBP gene promoter.
[0247] Using CRISPR-Cas9 for gene editing The CRISPR-Cas9 system described herein can be used for gene editing, which can result in gene silencing events, or expression modifications (e.g., increases or decreases) in the expression of a desired target gene. Thus, in some embodiments, the CRISPR-Cas9 system described herein is used in a method for modifying the expression of a target nucleic acid. In some embodiments, the CRISPR-Cas9 system described herein is used in a method for modifying a target nucleic acid in a desired target cell. In some embodiments, the present invention provides a method for site-specific modification of a target nucleic acid in a eukaryotic cell to achieve a desired modification in gene expression.
[0248] In some embodiments, the invention provides an engineered non-naturally occurring CRISPR-Cas system comprising an RNA guide or a nucleic acid encoding the RNA guide, where the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to a target nucleic acid, and a codon-optimized CRISPR-associated (Cas) protein having at least 80% sequence identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14, where the Cas protein is capable of binding to the RNA guide and causing cleavage in a target nucleic acid sequence complementary to the RNA guide.
[0249] In some embodiments, the invention provides an engineered non-naturally occurring CRISPR-Cas system comprising an RNA guide or a nucleic acid encoding the RNA guide, where the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to a target nucleic acid, and a codon-optimized CRISPR-associated (Cas) protein having at least 80% sequence identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14, where the Cas protein is fused to a deaminase, and where the Cas protein fusion is capable of binding to the RNA guide and editing a target nucleic acid sequence complementary to the RNA guide.
[0250] In some embodiments, the invention provides a method of modifying expression of a target nucleic acid in a eukaryotic cell comprising contacting a cell with a Cas9 as described herein and an RNA guide or a nucleic acid encoding the RNA guide, wherein the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to the target nucleic acid, and wherein the Cas9 protein is capable of binding to the RNA guide and causing cleavage at the target nucleic acid sequence complementary to the RNA guide.
[0251] In some embodiments, the invention provides a method of modifying expression of a target nucleic acid in a eukaryotic cell comprising contacting a cell with a Cas9 as described herein and an RNA guide or a nucleic acid encoding the RNA guide, wherein the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to the target nucleic acid, and wherein the Cas9 protein is capable of binding to the RNA guide and editing the target nucleic acid sequence that is complementary to the RNA guide.
[0252] In some embodiments, the invention provides a method of modifying expression of a target nucleic acid in a eukaryotic cell comprising contacting a cell with a Cas9 as described herein and an RNA guide or a nucleic acid encoding the RNA guide, wherein the RNA guide comprises direct repeats and spacer sequences capable of hybridizing to the target nucleic acid, and wherein the Cas9 protein is capable of binding to the RNA guide and editing the target nucleic acid sequence that is complementary to the RNA guide.
[0253] Thus, in some embodiments, the Cas protein has about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 14. In some embodiments, the Cas protein is identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 14.
[0254] Suitable guide RNAs, Cas9 mutations, and fusion proteins for use in the CRISPR-Cas9 systems and methods are as described throughout this disclosure.
[0255] In one aspect, the method includes binding CRISPR-Cas9 to a target nucleic acid and cleaving the target nucleic acid. In some embodiments, the CRISPR-Cas9 system cleaves the target DNA duplex or the target RNA duplex by introducing a double-stranded break. In some embodiments, the CRISPR-Cas9 system cleaves the target DNA or the target RNA by introducing a single-stranded break or nick.
[0256] In some embodiments, a CRISPR-Cas9 method or system comprises a fusion protein having an effector that modifies target DNA in a site-specific manner, the modifying activity comprising a methyltransferase activity, a demethylase activity, an acetyltransferase activity, a deacetylase activity, a kinase activity, a phosphatase activity, a ubiquitin ligase activity, a deubiquitinating activity, an adenylating activity, a deadenylating activity, a sumoylating activity, a desumoylating activity, a ribosylation activity, a deribosylation activity, a myristoylating activity, a demyristoylating activity, an integrase activity, a transposase activity, a recombinase activity, a polymerase activity, a ligase activity, a helicase activity, or a nuclease activity, any of which can modify DNA or a DNA-associated polypeptide (e.g., a histone or a DNA-binding protein).
[0257] In some embodiments, the CRISPR-Cas9 method or system includes a fusion protein with an enzyme that can edit DNA sequences by chemically modifying nucleotide bases, including a deaminase enzyme that can modify adenosine or cytosine bases and function as a site-specific base editor. For example, APOBEC1 cytidine deaminase, which normally uses RNA as a substrate, can target single-stranded and double-stranded DNA when fused to Cas9, directly converting cytidine to uridine, and ADAR enzymes deaminating adenosine to inosine. Thus, "base editing" using deaminases allows for programmable conversion of one target DNA base to another. Various base editors are known in the art and can be used in the methods and systems described herein. Exemplary base editors are described, for example, in Rees and Liu Nature Review Genetics, 2018, 19(12):770-788, the contents of which are incorporated herein. Thus, in some embodiments, the Cas9 enzymes described herein (Seq2Cas9, EhiCas9, SeqCas9, SsiCas9, SinCas9, SsaCas9, Ssc2Cas9, Sor2Cas9, SorCas9, SwaCas9, SscCas9, SgaCas9, LkuCas9, and SsuCas9) are components of a nucleobase editor. In some embodiments, the base editor is the adenine deaminase TadA8 or TadA9.
[0258] In some embodiments, base editing results in the introduction of a stop codon to silence a gene, in some embodiments, base editing results in altered protein function by modifying the amino acid sequence.
[0259] In some embodiments, the CRISPR-Cas9 method or system includes epigenetic modification of the target DNA by fusion with histones. In some embodiments, the CRISPR-Cas9 system includes epigenetic modification of the target DNA by fusion with epigenetic modification enzymes such as reader, writer, or eraser proteins. In some embodiments, the CRISPR-Cas9 system includes fusion with histone modification enzymes that alter the histone modification pattern in selected regions of the target DNA. Histone modifications can occur in many different ways and in many different combinations, including methylation, acetylation, ubiquitination, phosphorylation, resulting in structural changes of the DNA. In some embodiments, the histone modifications result in transcriptional repression or activation.
[0260] In some embodiments, the CRISPR-Cas9 method or system regulates the transcription of target DNA by increasing or decreasing transcription through fusion with transcription activator or repressor proteins, small molecule / drug responsive transcription regulators, inducible transcription regulators. In some embodiments, the CRISPR-Cas9 system is used to control the expression of target coding mRNA (i.e., protein encoding gene), whose binding results in increased or decreased gene expression.
[0261] In some embodiments, the CRISPR-Cas9 method or system is used to control gene regulation by editing gene regulatory elements such as promoters or enhancers.
[0262] In some embodiments, the CRISPR-Cas9 method or system is used to control expression of target non-coding RNAs, including tRNAs, rRNAs, snoRNAs, siRNAs, miRNAs, and long ncRNAs.
[0263] In some embodiments, the CRISPR-Cas9 method or system is used for targeted manipulation of chromatin loop structures. Targeted manipulation of chromatin loops between regulatory genomic regions provides a means to manipulate endogenous chromatin structure and allow the formation of new enhancer-promoter connections to overcome genetic defects or inhibit aberrant enhancer-promoter connections.
[0264] In some embodiments, CRISPR-Cas9 is used for live cell imaging. Fluorescently labeled Cas9 targets repetitive genomic regions such as centromeres and telomeres to track native chromatin loci throughout the cell cycle and determine differential positioning of transcriptionally active and inactive regions in 3D nuclear space.
[0265] In some embodiments, the CRISPR-Cas9 method or system is used for the correction of a pathogenic mutation by the insertion of a beneficial clinical variant or a suppressor mutation.
[0266] Nucleic acid base editor Disclosed herein is a novel base editor or nucleobase editor for editing, modifying, or altering a target nucleotide sequence of a polynucleotide, comprising Cas9. Described herein is a nucleobase editor or base editor, comprising a polynucleotide programmable nucleotide binding domain (e.g., Cas9) and a nucleobase editing domain (e.g., adenosine deaminase). The polynucleotide programmable nucleotide binding domain (e.g., Cas9), when combined with a bound guide polynucleotide (e.g., gRNA), specifically binds to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence), thereby allowing the base editor to localize to the target nucleic acid sequence desired to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. As most of the known genetic variations associated with human disease are point mutations, there is a need for a method that can more efficiently and cleanly create precise point mutations. The base editor systems provided herein offer a new method of providing genome editing that does not generate double-stranded DNA breaks, does not require a donor DNA template, and does not induce excessive stochastic insertions and deletions.
[0267] The base editors provided herein can modify specific nucleotide bases without generating a significant proportion of indels. As used herein, the term "indel(s)" refers to an insertion or deletion of a nucleotide base in a nucleic acid. Such an insertion or deletion can result in a frameshift mutation in the coding region of a gene. In some embodiments, it is desirable to generate a base editor that efficiently modifies (e.g., mutates or deaminates) specific nucleotides in a nucleic acid without generating a large number of insertions or deletions (i.e., indels) in the target nucleotide sequence. In certain embodiments, any of the base editors provided herein can generate a greater proportion of intended modifications (e.g., point mutations or deaminations) relative to indels.
[0268] In some embodiments, any of the base editor systems provided herein result in less than 50%, less than 40%, less than 30%, less than 20%, less than 19%, less than 18%, less than 17%, less than 16%, less than 15%, less than 14%, less than 13%, less than 12%, less than 11%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, less than 0.1%, less than 0.09%, less than 0.08%, less than 0.07%, less than 0.06%, less than 0.05%, less than 0.04%, less than 0.03%, less than 0.02%, or less than 0.01% indel formation in the target polynucleotide sequence.
[0269] Some aspects of the disclosure are based on the recognition that any of the base editors provided herein can efficiently generate intended mutations, e.g., point mutations, in a nucleic acid (e.g., a nucleic acid in a genome of a subject) without generating a significant number of unintended mutations, such as unintended point mutations. In some embodiments, any of the base editors provided herein can generate at least 0.01% intended mutations (i.e., at least 0.01% base editing efficiency). In some embodiments, any of the base editors provided herein can generate at least 0.01%, 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of intended mutations.
[0270] In some embodiments, the base editors provided herein are capable of generating a ratio of intended point mutations to indels of greater than 1:1. In some embodiments, the base editors provided herein are capable of generating a ratio of intended point mutations to indels of greater than 1:1. In some embodiments, the base editors provided herein are capable of generating a ratio of intended point mutations to indels of greater than 1:1. In some embodiments, a ratio of intended point mutations to indels that is at least 13:1, at least 14:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 200:1, at least 300:1, at least 400:1, at least 500:1, at least 600:1, at least 700:1, at least 800:1, at least 900:1, or at least 1000:1 or more can be generated.
[0271] The number of intended mutations and indels may be determined using any suitable method, suitable methods being described, for example, in International PCT Application No. 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), Komor, AC, et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016), Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017), and Komor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), the entire contents of which are incorporated herein by reference.
[0272] In some embodiments, to calculate the indel frequency, sequencing reads are scanned for exact matches with two 10 bp sequences flanking the window in which an indel may occur. If an exact match is not found, the read is excluded from the analysis. If the length of this indel window matches the reference sequence exactly, the read is classified as not containing an indel. If the indel window is more than one base longer or shorter than the reference sequence, the sequencing read is classified as an insertion or deletion, respectively. In some embodiments, the base editors provided herein can limit the formation of indels in a region of a nucleic acid. In some embodiments, the region is at the nucleotide targeted by the base editor or within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nucleotide targeted by the base editor.
[0273] The number of indels formed at a target nucleotide region can depend on the amount of time a nucleic acid (e.g., a nucleic acid in the genome of a cell) is exposed to a base editor. In some embodiments, the number or percentage of indels is determined at least 1 hour, at least 2 hours, at least 6 hours, at least 12 hours, at least 24 hours, at least 36 hours, at least 48 hours, at least 3 days, at least 4 days, at least 5 days, at least 7 days, at least 10 days, or at least 14 days after exposing a target nucleotide sequence (e.g., a nucleic acid in the genome of a cell) to a base editor. It is understood that the features of base editors described herein can be applied to any of the fusion proteins or methods of using the fusion proteins provided herein.
[0274] therapeutic use The CRISPR-Cas9 method or system described herein can have various therapeutic applications. Thus, in some embodiments, a method of treating a disorder or disease in a subject in need thereof is provided, the method comprising administering to the subject a CRISPR-Cas9 system comprising a Cas9 described herein, wherein a guide RNA is complementary to at least 10 nucleotides of a target nucleic acid associated with the condition or disease, a Cas protein associates with the guide RNA, the guide RNA binds to the target nucleic acid, and the Cas protein causes cleavage in the target nucleic acid, and optionally, the Cas9 is an inactive Cas9 fused to a deaminase (dCas9), which causes one or more base edits in the target nucleic acid, thereby treating the disorder or disease.
[0275] In some embodiments, the CRISPR-Cas9 method or system can be used to treat a variety of diseases and disorders, such as genetic disorders (e.g., monogenic diseases), diseases that can be treated by nuclease activity, and various cancers.
[0276] In some embodiments, the CRISPR methods or systems described herein can be used to edit a target nucleic acid to modify the target nucleic acid (e.g., by inserting, deleting, or mutating one or more nucleic acid residues). For example, in some embodiments, the CRISPR systems described herein include an exogenous donor template nucleic acid (e.g., a DNA molecule or an RNA molecule) that includes a desired nucleic acid sequence. After resolution of the cleavage event induced by the CRISPR systems described herein, the molecular machinery of the cell will utilize the exogenous donor template nucleic acid in repairing and / or resolving the cleavage event. Alternatively, the molecular machinery of the cell can utilize an endogenous template in repairing and / or resolving the cleavage event. In some embodiments, the CRISPR systems described herein can be used to modify a target nucleic acid to result in an insertion, deletion, and / or point mutation. In some embodiments, the insertion is a scarless insertion (i.e., insertion of an intended nucleic acid sequence into a target nucleic acid that does not result in additional unintended nucleic acid sequences upon resolution of the cleavage event). The donor template nucleic acid can be a double-stranded or single-stranded nucleic acid molecule (e.g., DNA or RNA). In some embodiments, the CRISPR methods or systems described herein comprise a nucleobase editor. For example, in some embodiments, a Cas9 protein described herein is fused to a polypeptide having nucleobase editing activity.
[0277] In one aspect, the CRISPR methods or systems described herein can be used to treat diseases caused by overexpression of RNA, toxic RNA, and / or mutant RNA (e.g., splicing defects or truncations).
[0278] In some embodiments, the CRISPR methods or systems described herein can also target trans-acting mutations that affect RNA-dependent functions that cause various diseases.
[0279] In some embodiments, the CRISPR methods or systems described herein can also be used to target mutations that disrupt the cis-acting splicing code, which can cause splicing defects and disease.
[0280] The CRISPR methods or systems described herein can further be used for antiviral activity, particularly against RNA viruses. CRISPR-associated proteins can target viral RNA using suitable RNA guides selected to target viral RNA sequences.
[0281] The CRISPR methods or systems described herein can also be used to treat cancer in a subject (e.g., a human subject). For example, the CRISPR-associated proteins described herein can be programmed with a crRNA that targets an RNA molecule that is aberrant (e.g., contains a point mutation or is alternatively spliced) and found in cancer cells to induce cell death (e.g., via apoptosis) in the cancer cells.
[0282] Furthermore, the CRISPR method or system described herein can also be used to treat infectious diseases in a subject. For example, the CRISPR-associated proteins described herein can be programmed with crRNAs that target RNA molecules expressed by infectious agents (e.g., bacteria, viruses, parasites, or protozoa) to target and induce cell death in infected cells. The CRISPR system can also be used to treat diseases in which intracellular infectious agents infect cells of a host subject. By programming the CRISPR-associated proteins to target RNA molecules encoded by infectious agent genes, cells infected with infectious agents can be targeted and induce cell death.
[0283] Furthermore, in vitro RNA sensing assays can be used to detect specific RNA substrates.CRISPR-associated proteins can be used for RNA-based sensing in living cells.An example of application is the diagnosis by sensing, for example, disease-specific RNA.
[0284] In applications where it is desired to insert a polynucleotide sequence into a target DNA sequence, a polynucleotide comprising the donor sequence to be inserted is also provided to the cell. By "donor sequence" or "donor polynucleotide" is meant a nucleic acid sequence to be inserted at the cleavage site induced by the site-directed modifying polypeptide. The donor polynucleotide contains sufficient homology to the genomic sequence at the cleavage site, e.g., 70%, 80%, 85%, 90%, 95%, or 100% homology to the nucleotide sequence adjacent to the cleavage site, e.g., within about 50 bases or less, e.g., within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases, or immediately adjacent to the cleavage site, to support homology-directed repair between the genomic sequence with which it has homology. Approximately 25, 50, 100, or 200 nucleotides, or more than 200 nucleotides of sequence homology between the donor and the genomic sequence (or any integer value between 10 and 200 nucleotides or more) will support homology-directed repair. The donor sequence can be of any length, e.g., 10 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 250 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 5000 nucleotides or more, etc.
[0285] The donor sequence is typically not identical to the genomic sequence it replaces. Rather, the donor sequence may contain at least one or more single base changes, insertions, deletions, inversions, or translocations with respect to the genomic sequence, so long as there is sufficient homology to support homology-directed repair. In some embodiments, the donor sequence contains two regions of homology and flanking non-homologous sequences, such that homology-directed repair between the target DNA region and the two flanking sequences results in the insertion of the non-homologous sequence into the target region. The donor sequence may also include a vector backbone that contains sequences that are not homologous to the DNA region of interest and are not intended for insertion into the DNA region of interest. In general, the homologous region(s) of the donor sequence will have at least 50% sequence identity with the genomic sequence with which recombination is desired. In certain embodiments, there is 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.9% sequence identity. Depending on the length of the donor polynucleotide, there can be any value of sequence identity between 1% and 100%.
[0286] The donor sequence may contain certain sequence differences compared to the genomic sequence, such as restriction sites, nucleotide polymorphisms, selectable markers (e.g., drug resistance genes, fluorescent proteins, enzymes, etc.), which may be used to evaluate the successful insertion of the donor sequence at the cleavage site, or in some cases, for other purposes (e.g., to indicate expression at the targeted genomic locus). In some cases, when located in a coding region, such nucleotide sequence differences will not change the amino acid sequence or will make silent amino acid changes (i.e., changes that do not affect the structure or function of the protein). Alternatively, these sequence differences may include adjacent recombination sequences, such as FLP, loxP sequences, that can be activated later to remove the marker sequence.
[0287] The donor sequence may be provided to the cell as single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA. It may be introduced into the cell in linear or circular form. If introduced in linear form, the ends of the donor sequence may be protected (e.g., from exonucleolytic degradation) by methods known to those skilled in the art. For example, one or more dideoxynucleotide residues are added to the 3' end of the linear molecule and / or a self-complementary oligonucleotide is ligated to one or both ends. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino group(s) and the use of modified internucleotide linkages, such as, for example, phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues. As an alternative to protecting the ends of linear donor sequences, additional lengths of sequence may be included outside the regions of homology that may be degraded without affecting recombination. The donor sequence may be introduced into a cell as part of a vector molecule having additional sequences such as, for example, an origin of replication, a promoter, and genes encoding antibiotic resistance. Additionally, the donor sequence may be introduced as naked nucleic acid, as nucleic acid complexed with an agent such as a liposome or poloxamer, or delivered by a virus (e.g., adenovirus, AAV), as described above for the nucleic acid encoding the DNA-targeting RNA and / or the site-directed modifying polypeptide and / or the donor polynucleotide.
[0288] According to the methods described above, the DNA region of interest may be cleaved and modified, i.e., "genetically modified", ex vivo. In some embodiments, as in the case where a selectable marker is inserted into the DNA region of interest, the population of cells may be enriched for those containing the genetic modification by separating the genetically modified cells from the remaining population. Prior to enrichment, the "genetically modified" cells may constitute only about 1% or more (e.g., 2% or more, 3% or more, 4% or more, 5% or more, 6% or more, 7% or more, 8% or more, 9% or more, 10% or more, 15% or more, or 20% or more) of the cell population. Separation of the "genetically modified" cells may be achieved by any convenient separation technique appropriate for the selectable marker used. For example, if a fluorescent marker is inserted, the cells may be separated by fluorescence activated cell sorting, and if a cell surface marker is inserted, the cells may be separated from the heterogeneous population by affinity separation techniques, such as magnetic separation, affinity chromatography, "panning" with affinity reagents bound to a solid matrix, or other convenient techniques. Techniques that provide accurate separation include fluorescence-activated cell sorters, which can have various degrees of sophistication, such as multiple color channels, low-angle and obtuse-angle light scattering detection channels, impedance channels, etc. Cells can be selected against dead cells by using a dye that is associated with dead cells (e.g., propidium iodide). Any technique that is not overly detrimental to the viability of the genetically modified cells can be used. A cell composition highly enriched for cells containing modified DNA can be achieved in this manner. By "highly enriched" it is meant that the genetically modified cells are 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, e.g., about 95% or more, or 98% or more of the cell composition. In other words, the composition can be a substantially pure composition of genetically modified cells.
[0289] The genetically modified cells produced by the methods described herein can be used immediately. Alternatively, the cells can be frozen at liquid nitrogen temperature, stored for a long time, and thawed and reused. In such cases, the cells are usually frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered medium, or some other such solution commonly used in the art to preserve cells at such freezing temperatures, and thawed in a manner commonly known in the art to thaw frozen cultured cells.
[0290] The genetically modified cells may be cultured in vitro under a variety of culture conditions. The cells may be expanded in culture, i.e., grown under conditions that promote cell proliferation. The culture medium may be liquid or semi-solid, containing, for example, agar, methylcellulose, etc. The cell population may be suspended in a suitable nutrient medium, such as Iscove's modified DMEM or RPMI 1640, usually supplemented with fetal bovine serum (about 5-10%), L-glutamine, thiols, especially 2-mercaptoethanol, and antibiotics, such as penicillin and streptomycin.
[0291] The culture may contain a growth factor to which the regulatory T cells are responsive. A growth factor, as defined herein, is a molecule that can promote the survival, growth and / or differentiation of cells, either in culture or in intact tissue, through a specific effect on a transmembrane receptor. Growth factors include polypeptide and non-polypeptide factors.
[0292] Such genetically modified cells may be implanted into a subject, for example, to treat a disease or as an antiviral, antipathogenic, or anticancer therapeutic, for purposes such as gene therapy, for the production of genetically modified organisms in agriculture, or for biological research. The subject may be a neonate, juvenile, or adult. Of particular interest are mammalian subjects. Mammalian species that may be treated with the present method include dogs and cats, horses, cows, sheep, etc., as well as primates, especially humans. Animal models, especially small mammals (e.g., mice, rats, guinea pigs, hamsters, lagomorphs (e.g., rabbits), etc.), may be used for experimental investigations.
[0293] The cells may be provided to a subject alone or, for example, with a suitable substrate or matrix to support the growth and / or organization of the cells within the tissue into which they are implanted. Typically, at least 1×10 3 cells, e.g., 5 x 10 cells, 1 x 10 4 cells, 5 x 10 4 cells, 1 x 10 5 cells, 1 x 10 6 One or more cells will be administered. The cells may be introduced into the subject via any of the following routes: parenteral, subcutaneous, intravenous, intracranial, intraspinal, intraocular, or into the spinal fluid. The cells may be introduced by injection, catheter, etc. The cells may also be introduced into an embryo (e.g., a blastocyst) for the purpose of generating a transgenic animal (e.g., a transgenic mouse).
[0294] The number of times that the treatment is administered to the subject may vary. Introducing the genetically modified cells into the subject may be a one-time event, but in certain circumstances, such treatment may induce improvement over a limited period of time and require a series of continuous repeated treatments. In other circumstances, multiple administrations of genetically modified cells may be required before an effect is observed. The exact protocol depends on the disease or condition, the stage of the disease, and the parameters of the individual subject being treated.
[0295] In other aspects of the invention, DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides are used to modify cellular DNA in vivo, for example to treat disease or as antiviral, antipathogenic or anticancer therapeutics, for purposes such as gene therapy, for the production of genetically modified organisms in agriculture, or for biological research. In these in vivo embodiments, the DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides are administered directly to an individual. The DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides may be administered by any of several well-known methods in the art for the administration of peptides, small molecules and nucleic acids to a subject. The DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides may be incorporated into various formulations. More specifically, the DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides of the invention may be formulated into a pharmaceutical composition by combination with a suitable pharma-ceutically acceptable carrier or diluent.
[0296] A pharmaceutical preparation is a composition that includes one or more DNA-targeting RNAs and / or site-directed modifying polypeptides and / or donor polynucleotides present in a pharmaceutically acceptable vehicle. A "pharmaceutically acceptable vehicle" can be a vehicle approved by a federal or state government regulatory agency or listed in the United States.
[0297] Pharmacopoeias or other generally accepted pharmacopoeias for use in mammals, such as humans. The term "vehicle" refers to a diluent, adjuvant, excipient, or carrier in which the compound of the present invention is formulated for administration to a mammal. Such pharmaceutical vehicles can be lipids, such as liposomes, e.g., liposomal dendrimers; water, oils, including those of petroleum, animal, vegetable, or synthetic origin, such as peanut oil, soybean oil, mineral oil, sesame oil, and the like, liquids, such as saline; gum acacia, gelatin, starch paste, talc, keratin, colloidal silica, urea, and the like. In addition, auxiliary agents, stabilizers, thickeners, lubricants, and colorants may be used. The pharmaceutical compositions can be formulated into solid, semi-solid, liquid, or gaseous form of preparations, such as tablets, capsules, powders, granules, ointments, solutions, suppositories, injections, inhalants, gels, microparticles, and aerosols. Thus, administration of DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides can be accomplished in a variety of ways, including oral, buccal, rectal, parenteral, intraperitoneal, intradermal, transdermal, intratracheal, intraocular, etc. The active agent may be systemic following administration, or may be localized by use of topical administration, intramural administration, or the use of an implant that acts to retain an active dose at the site of implantation. The active agent may be formulated for immediate activity or for sustained release.
[0298] For some conditions, particularly those of the central nervous system, it may be necessary to formulate drugs to cross the blood-brain barrier (BBB). One strategy for drug delivery through the blood-brain barrier (BBB) involves the disruption of the BBB, either by osmotic means such as mannitol or leukotrienes, or biochemically by the use of vasoactive substances such as bradykinin. The possibility of using BBB opening to target specific drugs to brain tumors is also an option. BBB disrupting agents may be co-administered with the therapeutic compositions of the invention when the compositions are administered by intravascular injection. Other strategies for crossing the BBB may involve the use of endogenous transporter systems, including caveolin 1-mediated transcytosis, carrier-mediated transporters such as glucose and amino acid carriers, receptor-mediated transcytosis of insulin or transferrin, and active efflux transporters such as p-glycoprotein. Active transport moieties may also be conjugated to therapeutic compounds for use in the invention to facilitate transport across the endothelial wall of blood vessels.
[0299] Alternatively, drug delivery of therapeutic agents behind the BBB may be by local delivery, for example, intrathecal delivery.
[0300] Typically, an effective amount of DNA-targeting RNA and / or site-directed modified polypeptide and / or donor polynucleotide is provided. As discussed above with respect to ex vivo methods, an effective amount or effective dose of DNA-targeting RNA and / or site-directed modified polypeptide and / or donor polynucleotide is an amount that induces a two-fold or greater increase in the amount of recombination observed between two homologous sequences, compared to cells contacted with a negative control, such as an empty vector or an unrelated polypeptide. The amount of recombination can be measured by any convenient method, such as those described above and known in the art. Calculation of the effective amount or effective dose of DNA-targeting RNA and / or site-directed modified polypeptide and / or donor polynucleotide to be administered is within the skill of, and routine for, the skilled artisan. The final amount to be administered depends on the route of administration and the nature of the disorder or condition to be treated.
[0301] The effective amount given to a particular patient will depend on a variety of factors, some of which will vary from patient to patient. A competent clinician will be able to determine the effective amount of therapeutic agent to administer to a patient to halt or reverse the progression of a disease state, if necessary. Using LD50 animal data, and other available information about the agent, the clinician can determine the maximum safe dose for an individual, depending on the route of administration. For example, a dose administered intravenously may be higher than a dose administered intrathecally, given the larger volume of fluid into which the therapeutic composition is administered. Similarly, compositions that are rapidly cleared from the body may be administered at higher doses or in repeated doses to maintain therapeutic concentrations. Using routine techniques, a competent clinician will be able to optimize the dosage of a particular therapeutic agent during the course of routine clinical trials.
[0302] For inclusion in the medicament, the DNA-targeting RNA and / or the site-directed modifying polypeptide and / or the donor polynucleotide may be obtained from a suitable commercial source. As a general proposition, the total pharmacologic effective amount of the DNA-targeting RNA and / or the site-directed modifying polypeptide and / or the donor polynucleotide administered parenterally per dose is in a range that can be determined by a dose-response curve.
[0303] Therapies based on DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides, i.e., preparations of DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides used for therapeutic administration, need to be sterile. Sterility is easily achieved by filtration through a sterile filtration membrane (e.g., 0.2 μm membrane). Therapeutic compositions are generally placed in containers with a sterile access port, such as intravenous solution bags or vials with a stopper that can be pierced by a hypodermic needle. Therapies based on DNA-targeting RNA and / or site-directed modified polypeptides and / or donor polynucleotides can be stored in unit-dose or multi-dose containers, such as sealed ampoules or vials, as aqueous solutions or as lyophilized formulations for reconstitution. As an example of a lyophilized formulation, a 10 mL vial is filled with 5 ml of a sterile-filtered 1% (w / v) aqueous solution of the compound, and the resulting mixture is lyophilized. Infusion solutions are prepared by reconstituting the lyophilized compound using bacteriostatic water for injection.
[0304] The pharmaceutical composition may contain a pharma- ceutically acceptable non-toxic carrier of a diluent, which is defined as a vehicle commonly used to formulate pharmaceutical compositions for animal or human administration, depending on the desired formulation. The diluent is selected so as not to affect the biological activity of the combination. Examples of such diluents are distilled water, buffered water, physiological saline, PBS, Ringer's solution, dextrose solution, and Hank's solution. In addition, the pharmaceutical composition or formulation may contain other carriers, adjuvants, or non-toxic, non-therapeutic, non-immunogenic stabilizers, excipients, and the like. The composition may also contain additional substances that approximate physiological conditions, such as pH adjusting and buffering agents, toxicity adjusting agents, wetting agents, and detergents.
[0305] The composition may also include any of a variety of stabilizing agents, such as, for example, antioxidants. When the pharmaceutical composition includes a polypeptide, the polypeptide may be complexed with a variety of well-known compounds that enhance the in vivo stability of the polypeptide or otherwise enhance its pharmacological properties (e.g., increase the half-life of the polypeptide, reduce its toxicity, enhance solubility or uptake). Examples of such modifying or complexing agents include sulfates, gluconates, citrates, and phosphates. The nucleic acids or polypeptides of the composition may also be complexed with molecules that enhance their in vivo attributes. Such molecules include, for example, carbohydrates, polyamines, amino acids, other peptides, ions (e.g., sodium, potassium, calcium, magnesium, manganese), and lipids.
[0306] The pharmaceutical composition can be administered for prophylactic and / or therapeutic treatment. The toxicity and therapeutic efficacy of the active ingredient can be determined according to standard pharmaceutical procedures in cell cultures and / or experimental animals, including, for example, determining LD50 (the dose lethal to 50% of the population) and ED50 (the dose therapeutically effective for 50% of the population). The dose ratio between the toxic effect and the therapeutic effect is the therapeutic index, which can be expressed as the ratio LD50 / ED50. Therapies that exhibit large therapeutic indices are preferred.
[0307] The data obtained from cell culture and / or animal studies can be used in formulating a range of dosages for humans.The dosage of active ingredient is typically within the range of circulating concentrations that include the ED50 with low toxicity.Dosage can vary within this range depending on the dosage form used and the route of administration utilized.
[0308] The components used to formulate pharmaceutical compositions are preferably of high purity and substantially free of potentially harmful contaminants (e.g., at least National Food (NF) grade, generally at least analytical grade, and more typically at least pharmaceutical grade). Moreover, compositions intended for in vivo use are usually sterile. To the extent that a given compound must be synthesized prior to use, the resulting product is typically substantially free of any potentially toxic agents, particularly any endotoxins, that may be present during the synthesis or purification process. Compositions for parental administration are also sterile, substantially isotonic, and made under GMP conditions.
[0309] Delivery System The CRISPR system described herein, or its components, its nucleic acid molecules, and / or nucleic acid molecules encoding or providing its components, CRISPR-associated proteins, or RNA guides can be delivered by various delivery systems, such as vectors, e.g., plasmids and delivery vectors. Exemplary embodiments are described below. The CRISPR system (e.g., including Cas9, which includes the nucleic acid base editor described herein) can be encoded by a nucleic acid contained in a viral vector. Viral vectors can include lentiviruses, adenoviruses, retroviruses, and adeno-associated viruses (AAVs). Viral vectors can be selected based on the application. For example, AAVs are commonly used for in vivo gene delivery due to their mild immunogenicity. Adenoviruses are commonly used as vaccines because they induce a strong immunogenic response. The packaging capacity of the viral vector can limit the size of the base editor that can be packaged into the vector. For example, the packaging capacity of AAV is about 4.5 kb, which includes two 145-base inverted terminal repeats (ITRs).
[0310] AAV is a small, single-stranded DNA-dependent virus belonging to the Parvoviridae family. The 4.7 kb wild-type (wt) AAV genome consists of two genes encoding four replication proteins and three capsid proteins, respectively, flanked on either side by 145 bp inverted terminal repeats (ITRs). Virions are composed of three capsid proteins, Vp1, Vp2, and Vp3, which are generated in a 1:1:10 ratio from the same open reading frame, but from differential splicing (Vp1) and alternative translation initiation sites (Vp2 and Vp3, respectively). Vp3 is the most abundant subunit in the virion and is responsible for receptor recognition at the cell surface that defines the viral tropism. A phospholipase domain that functions in viral infectivity has been identified in the unique N-terminus of Vp1.
[0311] Similar to wt AAV, recombinant AAV (rAAV) utilizes 145 bp of cis-acting ITRs to flank the vector transgene cassette, providing up to 4.5 kb for packaging of foreign DNA. After infection, rAAV can express the fusion protein of the invention and persist without integrating into the host genome by existing episomally in circular head-to-tail concatemers. Although there are many successful examples of rAAV using this system in vitro and in vivo, packaging capacity limitations limit the use of AAV-mediated gene delivery when the length of the gene coding sequence is comparable to or greater than the wt AAV genome.
[0312] The small packaging capacity of AAV vectors makes it difficult to deliver some genes that exceed this size and / or use large physiological regulatory elements. These challenges can be addressed, for example, by splitting the protein(s) to be delivered into two or more fragments, with the N-terminal fragment fused to a split intein-N and the C-terminal fragment fused to a split intein-C. These fragments are then packaged into two or more AAV vectors. As used herein, "intein" refers to a self-splicing protein intron (e.g., a peptide) that ligates adjacent N-terminal and C-terminal exteins (e.g., fragments that join). The use of certain inteins to join heterologous protein fragments is described, for example, in Wood et al., J. Biol. Chem. 289(21);14512-9(2014). The inteins IntN and IntC, for example, when fused to separate protein fragments, recognize each other and splice them out while simultaneously ligating the adjacent N- and C-terminal exteins of the protein fragments to which they are fused, thereby reconstituting a full-length protein from the two protein fragments. Other suitable inteins will be apparent to those skilled in the art.
[0313] In some embodiments, the CRISPR systems of the present invention can vary in length. In some embodiments, the protein fragments range from 2 amino acids to about 1000 amino acids in length. In some embodiments, the protein fragments range from about 5 amino acids to about 500 amino acids in length. In some embodiments, the protein fragments range from about 20 amino acids to about 200 amino acids in length. In some embodiments, the protein fragments range from about 10 amino acids to about 100 amino acids in length. Other lengths of suitable protein fragments will be apparent to one of skill in the art.
[0314] In some embodiments, a portion or fragment of a nuclease (e.g., Cas9) is fused to an intein. The nuclease may be fused to the N-terminus or C-terminus of the intein. In some embodiments, a portion or fragment of a fusion protein is fused to an intein and fused to an AAV capsid protein. The intein, nuclease, and capsid protein may be fused together in any configuration (e.g., nuclease-intein-capsid, intein-nuclease-capsid, capsid-intein-nuclease, etc.). In some embodiments, the N-terminus of the intein is fused to the C-terminus of the fusion protein and the C-terminus of the intein is fused to the N-terminus of the AAV capsid protein.
[0315] In one embodiment, dual AAV vectors are generated by splitting a large transgene expression cassette into two separate halves (5' and 3' ends, or head and tail), and each half of the cassette is packaged into a single AAV vector (less than 5 kb). Reconstitution of the full-length transgene expression cassette is then achieved upon coinfection of the same cells with both dual AAV vectors followed by (1) homologous recombination (HR) between the 5' and 3' genomes (dual AAV overlapping vectors), (2) ITR-mediated tail-to-head concatemerization of the 5' and 3' genomes (dual AAV trans-splicing vectors), or (3) a combination of these two mechanisms (dual AAV hybrid vectors). The use of dual AAV vectors in vivo results in the expression of full-length proteins. The use of a dual AAV vector platform represents an efficient and viable gene transfer strategy for transgenes greater than 4.7 kb in size.
[0316] The disclosed strategies for designing CRISPR systems including Cas9 described herein can be useful for generating CRISPR systems that can be packaged into viral vectors. The use of RNA or DNA virus-based systems for delivery of base editors takes advantage of the highly evolved processes for targeting viruses to specific cells in culture or within a host and transporting the viral payload to the nucleus or host cell genome. Viral vectors can be administered directly to cells in culture, to patients (in vivo), or they can be used to treat cells in vitro, and the modified cells can optionally be administered to patients (ex vivo). Traditional virus-based systems can include retroviral, lentiviral, adenoviral, adeno-associated viral, and herpes simplex viral vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated viral gene transfer methods, often resulting in long-term expression of the inserted transgene. In addition, high transduction efficiency has been observed in many different cell types and target tissues.
[0317] The tropism of retroviruses can be altered by incorporating foreign envelope proteins and expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically generate high viral titers. Thus, the choice of retroviral gene transfer system will depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats with packaging capacity for up to 6-10 kb of foreign sequence. The minimal cis-acting LTRs are sufficient for vector replication and packaging, which are then used to integrate the therapeutic gene into the target cell to provide persistent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731-2739 (1992); Johann et al., J. Virol. 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., J. Virol. 63:2374-2378 (1989); Miller et al., J. Virol. 65:2220-2224 (1991); PCT / US94 / 05700).
[0318] Retroviral vectors, particularly lentiviral vectors, may require a polynucleotide sequence smaller than a given length for efficient integration into target cells. For example, retroviral vectors longer than 9 kb may result in low viral titers compared to smaller sizes. In some aspects, the CRISPR system of the present disclosure (e.g., including Cas9 as disclosed herein) is of sufficient size to allow efficient packaging and delivery to target cells via retroviral vectors. In some embodiments, Cas9 is of such size that it allows efficient packaging and delivery even when expressed together with guide nucleic acid and / or other components of targetable nuclease system.
[0319] For applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors can have very high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained using such vectors. The vectors can be produced in large quantities in a relatively simple system. Adeno-associated virus ("AAV") vectors can also be used to transduce cells with target nucleic acids, for example, in the in vitro production of nucleic acids and peptides, for in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; WO93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994)). The construction of recombinant AAV vectors has been described in several publications, including U.S. Pat. No. 5,173,414, Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985), Tratschin, et al., Mol. Cell. Biol. 4:2072-2081 (1984), Hermonat & Muzyczka, PNAS 81:6466-6470 (1984), and Samulski et al., J. Virol. 63:03822-3828 (1989).
[0320] Thus, the CRISPR system described herein (including, for example, Cas9 disclosed herein) can be delivered with a viral vector. One or more components of the base editor system can be encoded on one or more viral vectors. For example, the base editor and guide nucleic acid can be encoded on a single viral vector. In other cases, the base editor and guide nucleic acid are encoded on different viral vectors. In either case, the base editor and guide nucleic acid can be operably linked to a promoter and a terminator, respectively.
[0321] The combination of components encoded on the viral vector may be dictated by the cargo size constraints of the selected viral vector.
[0322] Non-viral delivery of base editors Non-viral delivery approaches for CRISPR are also available. One important category of non-viral nucleic acid vectors are nanoparticles, which can be organic or inorganic. Nanoparticles are well known in the art. Any suitable nanoparticle design can be used to deliver the components of genome editing systems, or the nucleic acids encoding such components. For example, organic (e.g. lipid and / or polymer) nanoparticles can be suitable for use as delivery vehicles in certain embodiments of the present disclosure. Exemplary lipids for use in nanoparticle formulations and / or gene transfer are shown in Table 5 (below). [Table 5] TIFF2024546841000195.tif193170 Table 6 lists exemplary polymers for use in gene transfer and / or nanoparticle formulations. [Table 6] Table 7 summarizes the delivery methods for the Cas9-encoding polynucleotides described herein. [Table 7]
[0323] In another aspect, delivery of genome editing system components or nucleic acids encoding such components, e.g., a nucleic acid binding protein, e.g., Cas9 or a variant thereof, optionally fused to a biologically active polypeptide (e.g., a nucleic acid base editor), and a gRNA targeting a genomic nucleic acid sequence of interest, can be achieved by delivering a ribonucleoprotein (RNP) to a cell. The RNP comprises a nucleic acid binding protein (e.g., Cas9) in a complex with a target gRNA. The RNP can be delivered to a cell using known methods, such as electroporation, nucleofection, or cationic lipid-mediated methods, as reported, for example, by Zuris, JA et al., 2015, Nat. Biotechnology, 33(1):73-80. The RNP is advantageous for use in the CRISPR base editor system, especially in cells that are difficult to transfect (e.g., primary cells). In addition, RNPs can also alleviate the difficulties that may occur in protein expression in cells, especially when eukaryotic promoters (e.g., CMV or EF1A) that may be used in CRISPR plasmids are not fully expressed. Advantageously, the use of RNPs does not require the delivery of foreign DNA into cells. Furthermore, the use of RNPs may limit off-target effects, since the RNPs containing nucleic acid binding proteins and gRNA complexes are degraded over time. In a manner similar to that of plasmid-based technology, RNPs can be used to deliver binding proteins (e.g., Cas9 variants) and induce homology-directed repair (HDR).
[0324] The promoter used to drive the CRISPR system (including, for example, Cas9 as described herein) can include AAV ITRs. This can be advantageous to eliminate the need for additional promoter elements that may occupy space in the vector. The additional space released can be used to drive the expression of additional elements (e.g., guide nucleic acids or selectable markers). Since ITR activity is relatively weak, it can be used to reduce potential toxicity due to overexpression of the selected nuclease.
[0325] Any suitable promoter may be used to drive the expression of Cas9 and, if applicable, guide nucleic acid. For ubiquitous expression, promoters that may be used include CMV, CAG, CBh, PGK, SV40, ferritin heavy or light chain, etc. For expression in brain or other CNS cells, suitable promoters may include synapsin I for all neurons, CaMKII alpha for excitatory neurons, GAD67 or GAD65 or VGAT for GABAergic neurons, etc. For expression in hepatocytes, suitable promoters may include albumin promoter. For expression in lung cells, suitable promoters may include SP-B. For endothelial cells, suitable promoters may include ICAM. For hematopoietic cells, suitable promoters may include IFN beta or CD45. For osteoblasts, suitable promoters may include OG-2.
[0326] In some cases, the Cas9 of the disclosure is of a small enough size to allow separate promoters to drive expression of a base editor and a compatible guide nucleic acid within the same nucleic acid molecule. For example, a vector or viral vector can include a first promoter operably linked to a nucleic acid encoding a base editor and a second promoter operably linked to a guide nucleic acid.
[0327] Promoters used to drive expression of the guide nucleic acid can include PolIII promoters such as U6 or H1. To express the gRNA adeno-associated virus (AAV), a PolII promoter and an intron cassette is used.
[0328] The Cas9 described herein, with or without one or more guide nucleic acids, can be delivered using adeno-associated virus (AAV), lentivirus, adenovirus, or other plasmid or viral vector types, particularly using formulations and dosages from, for example, US Patent No. 8,454,972 (formulations, dosages for adenovirus), US Patent No. 8,404,658 (formulations, dosages for AAV), and US Patent No. 5,846,946 (formulations, dosages for DNA plasmids), as well as from clinical trials and publications regarding clinical trials involving lentivirus, AAV, and adenovirus. For example, in the case of AAV, the route of administration, formulation, and dosage can be similar to US Patent No. 8,454,972 and clinical trials involving AAV. In the case of adenovirus, the route of administration, formulation, and dosage can be similar to US Patent No. 8,404,658 and clinical trials involving adenovirus. In the case of plasmid delivery, the route of administration, formulation, and dosage can be similar to US Patent No. 5,846,946 and clinical trials involving plasmids. Dosages may be based on or extrapolated to an average 70 kg individual (e.g., adult human male) and may be adjusted for patients, subjects, mammals of different weights and species. Frequency of administration is within the scope of a physician or veterinarian (e.g., physician, veterinarian) depending on the usual factors including age, sex, general health, other conditions of the patient or subject, and the specific condition or symptom being addressed. The viral vector may be injected into the tissue of interest. In the case of cell type specific base editors, expression of the base editor and optional guide nucleic acid may be driven by a cell type specific promoter.
[0329] In the case of in vivo delivery, AAV may be more advantageous than other viral vectors.In some cases, AAV allows low toxicity, which may be due to the purification method that does not require ultracentrifugation of cell particles that may activate immune response.In some cases, AAV allows low probability of causing insertional mutagenesis because it does not integrate into host genome.
[0330] AAV has a packaging limit of 4.5 or 4.75Kb. Constructs that exceed 4.5Kb or 4.75Kb can result in a significant reduction in virus production. For example, SpCas9 is quite large, and the gene itself is over 4.1Kb, which makes it difficult to package into AAV. Therefore, embodiments of the present disclosure include the use of the disclosed Cas9, which is shorter in length than conventional Cas9.
[0331] The AAV may be AAV1, AAV2, AAV5, or any combination thereof. The type of AAV may be selected for the cells to be targeted, for example, AAV serotypes 1, 2, 5 or hybrid capsids AAV1, AAV2, AAV5, or any combination thereof may be selected for targeting brain or neuronal cells, and AAV4 may be selected for targeting cardiac tissue. AAV8 is useful for delivery to the liver. A tabulation of specific AAV serotypes for these cells can be found in Grimm, D. et al, J. Virol. 82:5887-5911 (2008).
[0332] Lentiviruses are complex retroviruses that have the ability to infect and express their genes in both dividing and post-dividing cells. The most commonly known lentivirus is the human immunodeficiency virus (HIV), which uses the envelope glycoproteins of other viruses to target a broad range of cell types.
[0333] Lentivirus can be prepared as follows: After cloning pCasES10 (containing the lentiviral transcription plasmid backbone), low passage (p=5) HEK293FT were seeded to 50% confluence in T-75 flasks in DMEM with 10% fetal bovine serum and no antibiotics the day before transfection. After 20 hours, the medium was replaced with OptiMEM (serum-free) medium, and transfection was performed 4 hours later. Cells are transfected with 10 μg of lentiviral transfer plasmid (pCasES10) and the following packaging plasmids: 5 μg of pMD2.G (VSV-g pseudotyped), and 7.5 μg of psPAX2 (gag / pol / rev / tat). Transfection can be performed in 4 mL of OptiMEM containing cationic lipid delivery agent (50 μl of Lipofectamine 2000 and 100 ul of Plus reagent). After 6 hours, the medium is replaced with antibiotic-free DMEM containing 10% fetal bovine serum. Although these methods use serum in cell culture, serum-free methods are preferred.
[0334] Lentivirus can be purified as follows: Viral supernatant is harvested after 48 hours. The supernatant is first removed from debris and filtered through a 0.45 μm low protein binding (PVDF) filter. It is then spun in an ultracentrifuge at 24,000 rpm for 2 hours. The viral pellet is resuspended in 50 μl of DMEM overnight at 4° C. They are then aliquoted and immediately frozen at −80° C.
[0335] In another embodiment, a minimal non-primate lentiviral vector based on the Equine Infectious Anemia Virus (EIAV) is also contemplated. In another embodiment, RetinoStat® is an Equine Infectious Anemia Virus-based lentiviral gene therapy vector that expresses the angiogenesis inhibitory proteins endostatin and angiostatin, which is contemplated to be delivered by subretinal injection. In another embodiment, the use of self-inactivating lentiviral vectors is contemplated.
[0336] Any RNA of the system, such as guide RNA or Cas9-encoding mRNA, can be delivered in the form of RNA. Cas9-encoding mRNA can be generated using in vitro transcription. For example, Cas9 mRNA can be synthesized using a PCR cassette containing the following elements: T7 promoter, optional kozak sequence (GCCACC), nuclease sequence, and 3'UTR (e.g., 3'UTR from beta globin-polyA tail). The cassette can be used for transcription by T7 polymerase. Guide polynucleotides (e.g., gRNAs) can also be transcribed using in vitro transcription from a cassette containing a T7 promoter followed by the sequence "GG" and a guide polynucleotide sequence.
[0337] To enhance expression and reduce potential toxicity, the Cas9 sequence and / or guide nucleic acid can be modified to include one or more modified nucleosides, for example, using pseudo-U or 5-methyl-C.
[0338] The present disclosure encompasses, in some embodiments, methods of modifying a cell or organism. The cell may be a prokaryotic or eukaryotic cell. The cell may be a mammalian cell. Many of the mammalian cells are non-human primate, bovine, porcine, rodent, or murine cells. The modifications introduced into the cell by the base editors, compositions, and methods of the present disclosure can result in the cell and the cell's progeny being altered for improved production of a biological product, such as an antibody, starch, alcohol, or other desired cellular output. The modifications introduced into the cell by the methods of the present disclosure can result in the cell and the cell's progeny containing a modification that alters the biological product produced.
[0339] The system can include one or more different vectors. In one embodiment, the Cas9 is codon-optimized for expression in a desired cell type, preferably a eukaryotic cell, preferably a mammalian cell or a human cell.
[0340] Generally, "codon optimization" refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by substituting at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) of the native sequence with a codon that is more frequently or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Different species show a particular bias for a particular codon of a particular amino acid. Codon bias (differences in codon usage between organisms) is often thought to correlate with the efficiency of translation of messenger RNA (mRNA), which in turn depends, among other things, on the properties of the codon being translated and the availability of a particular transfer RNA (tRNA) molecule. The dominance of tRNAs selected in a cell is generally a reflection of the codons most frequently used in peptide synthesis. Thus, genes can be adjusted for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database" available at www.kazusa.orjp / codon / (accessed July 9, 2002), and these tables can be adapted in several ways. See Nakamura, Y., et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28: 292 (2000). Computer algorithms are also available for optimizing codons for a particular sequence for expression in a particular host cell, for example, Gene Forge (Aptagen, Jacobus, Pa.). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons, or all codons) in a sequence encoding an engineered nuclease correspond to the most frequently used codon for a particular amino acid.
[0341] Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells, which package adenovirus, and psi.2 or PA317 cells, which package retrovirus. Viral vectors used in gene therapy are usually generated by producing cell lines that package nucleic acid vectors into viral particles. The vectors typically contain minimal viral sequences necessary for packaging and subsequent integration into the host, with other viral sequences being replaced by an expression cassette for the polynucleotide(s) to be expressed. Missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically only have ITR sequences from the AAV genome necessary for packaging and integration into the host genome. Viral DNA can be packaged into a cell line that contains a helper plasmid that encodes other AAV genes (i.e., rep and cap) but lacks ITR sequences. The cell line can be infected with adenovirus as a helper. The helper virus can facilitate replication of the AAV vector and expression of the AAV genes from the helper plasmid. In some cases, the helper plasmid is not packaged in significant amounts due to the lack of ITR sequences. Adenovirus contamination can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV.
[0342] Pharmaceutical Compositions Another aspect of the present disclosure relates to pharmaceutical compositions comprising a CRISPR system (e.g., including Cas9 as disclosed herein). The term "pharmaceutical composition" as used herein refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharma- ceutical acceptable carrier. In some embodiments, the pharmaceutical composition comprises an additional agent (e.g., for specific delivery, to increase half-life, or other therapeutic compounds).
[0343] As used herein, the term "pharmaceutically acceptable carrier" means a pharma- ceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, magnesium talc, calcium or zinc stearate, or steric acid), or solvent encapsulant, involved in carrying or transporting a compound from one site in the body (e.g., delivery site) to another (e.g., organ, tissue, or body part). A pharma- ceutically acceptable carrier is "acceptable" in the sense of being compatible with the other ingredients of the formulation and not deleterious to the tissues of the subject (e.g., physiological compatibility, sterility, physiological pH, etc.).
[0344] Some non-limiting examples of materials that can serve as pharma- ceutically acceptable carriers include: (1) sugars (e.g., lactose, glucose, and sucrose); (2) starches (e.g., corn starch and potato starch); (3) cellulose and its derivatives (e.g., sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose, and cellulose acetate); (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants (e.g., magnesium stearate, sodium lauryl sulfate, and talc); (8) excipients (e.g., cocoa butter and suppository wax); and (9) oils (e.g., peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil). , (10) glycols (e.g., propylene glycol), (11) polyols (e.g., glycerin, sorbitol, mannitol, and polyethylene glycol (PEG)), (12) esters (e.g., ethyl oleate and ethyl laurate), (13) agar, (14) buffers (e.g., magnesium hydroxide and aluminum hydroxide), (15) alginic acid, (16) pyrogen-free water, (17) isotonic saline, (18) Ringer's solution, (19) ethyl alcohol, (20) pH buffer solutions, (21) polyesters, polycarbonates, and / or polyanhydrides, (22) swelling agents (e.g., polypeptides and amino acids), (23) serum alcohols (e.g., ethanol), and (23) other non-toxic compatible substances used in pharmaceutical formulations. Wetting agents, coloring agents, release agents, coating agents, sweetening agents, flavoring agents, fragrances, preservatives, and antioxidants may also be present in the formulation. The terms "excipient," "carrier," "pharmaceutical acceptable carrier," "vehicle," and the like are used interchangeably herein.
[0345] The pharmaceutical composition may include one or more pH buffer compounds to maintain the pH of the formulation at a predetermined level that reflects physiological pH, for example, in the range of about 5.0 to about 8.0. The pH buffer compound used in the aqueous liquid formulation may be an amino acid, or a mixture of amino acids (e.g., histidine, or a mixture of amino acids such as histidine and glycine). Alternatively, the pH buffer compound is preferably an agent that maintains the pH of the formulation at a predetermined level, for example, in the range of about 5.0 to about 8.0, and does not chelate calcium ions. Illustrative examples of such pH buffer compounds include, but are not limited to, imidazole and acetate ions. The pH buffer compound may be present in any amount suitable for maintaining the pH of the formulation at a predetermined level.
[0346] The pharmaceutical composition may also contain one or more osmotic modifiers, i.e., compounds that adjust the osmotic properties (e.g., tonicity, osmolality, and / or osmolarity) of the formulation to a level that is acceptable to the bloodstream and blood cells of the recipient individual. The osmotic modifier may be an agent that does not chelate calcium ions. The osmotic modifier may be any compound known or available to one of skill in the art that adjusts the osmotic properties of the formulation. One of skill in the art can empirically determine the suitability of a given osmotic modifier for use in the formulations of the present invention. Illustrative examples of suitable types of osmotic modifiers include, but are not limited to, salts, such as sodium chloride and sodium acetate; sugars, such as sucrose, dextrose, and mannitol; amino acids, such as glycine; and mixtures of one or more of these agents and / or these types of agents. The osmotic modifier(s) may be present in any concentration sufficient to adjust the osmotic properties of the formulation.
[0347] In some embodiments, the pharmaceutical composition is formulated for delivery to a subject, for example, for gene editing. Suitable routes for administering the pharmaceutical compositions described herein include, but are not limited to, topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, intradental, intracochlear, intratympanic, intravisceral, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseous, periocular, intratumoral, intracerebral, and intraventricular administration.
[0348] In some embodiments, the pharmaceutical compositions described herein are administered locally to the disease site. In some embodiments, the pharmaceutical compositions described herein are administered to the subject by injection, by catheter, by suppository, or by implant, the implant being a porous, non-porous, or gel-like material, including membranes such as sialastic membranes or fibers.
[0349] In other embodiments, the pharmaceutical compositions described herein are delivered in a controlled release system.In one embodiment, pumps can be used (see, for example, Langer, 1990, Science 249:1527-1533; Sefton, 1989, CRC Crit. Ref. Biomed. Eng. 14:201; Buchwald et al., 1980, Surgery 88:507; Saudek et al., 1989, N. Engl. J. Med. 321:574).In another embodiment, polymeric materials can be used. (See, e.g., Medical Applications of Controlled Release (Langer and Wise eds., CRC Press, Boca Raton, Fla., 1974); Controlled Drug Bioavailability, Drug Product Design and Performance (Smolen and Ball eds., Wiley, New York, 1984); Ranger and Peppas, 1983, Macromol. Sci. Rev. Macromol. Chem. 23:61; see also Levy et al., 1985, Science 228:190; During et al., 1989, Ann. Neurol. 25:351; Howard et al., 1989, J. Neurosurg. 71:105.) Other controlled release systems are discussed, e.g., in Langer, supra.
[0350] In some embodiments, the pharmaceutical composition is formulated according to routine procedures as a composition suitable for intravenous or subcutaneous administration to a subject (e.g., a human). In some embodiments, the pharmaceutical composition for administration by injection is a solution in sterile isotonic use as a solubilizing agent, and a local anesthetic, such as lignocaine, to ease pain at the injection site. Generally, the ingredients are supplied separately or mixed together in unit dosage form, for example, as a dry lyophilized powder or water-free concentrate in an air-tight sealed container, such as an ampoule or sachet indicating the quantity of active agent. When the pharmaceutical is administered by injection, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline. When the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration.
[0351] Pharmaceutical compositions for systemic administration may be liquid (e.g., sterile saline, lactated Ringer's solution, or Hank's solution). In addition, pharmaceutical compositions may be in solid form and redissolved or suspended immediately prior to use. Lyophilized forms are also contemplated. Pharmaceutical compositions may be contained within lipid particles or vesicles (e.g., liposomes or microcrystals also suitable for parenteral administration). The particles may be of any suitable structure, such as unilamellar or plurilamellar, so long as the composition is contained therein. Compounds may be encapsulated in "stabilized plasmid lipid particles" (SPLPs) containing the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipids, and stabilized by polyethylene glycol (PEG) coating (see Zhang YPet ah, Gene Ther. 1999, 6:1438-47). For such particles and vesicles, positively charged lipids such as N-[1-(2,3-dioleoyloxy)propyl]-N,N,N-trimethyl-ammonium methylsulfate or "DOTAP" are particularly preferred.The preparation of such lipid particles is well known.See, for example, U.S. Patent Nos. 4,880,635, 4,906,477, 4,911,928, 4,917,951, 4,920,016 and 4,921,757 (each of which is incorporated herein by reference).
[0352] The pharmaceutical compositions described herein may be administered or packaged, for example, as unit doses. The term "unit dose", when used in reference to the pharmaceutical compositions of the present disclosure, refers to a physically discrete unit suitable as a unitary dosage for a subject, each unit containing a predetermined amount of active material calculated to produce a desired therapeutic effect in association with a required diluent (i.e., carrier or vehicle).
[0353] Additionally, the pharmaceutical compositions may be provided as pharmaceutical kits, comprising (a) a container containing a compound of the invention in lyophilized form, and (b) a second container containing a pharma- ceutically acceptable diluent (e.g., sterile, for use in reconstituting or diluting the lyophilized compound of the invention). Optionally, associated with such container(s) may be a notice in a form prescribed by a government agency regulating the manufacture, use, or sale of pharmaceutical or biological products, the notice reflecting approval by the agency of the manufacture, use, or sale for human administration.
[0354] Another aspect includes an article of manufacture containing materials useful for treating the diseases described above. In some embodiments, the article of manufacture includes a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The containers can be formed from a variety of materials, such as glass or plastic. In some embodiments, the container holds a composition effective for treating a disease described herein and can have a sterile access port. For example, the container can be an intravenous infusion bag or a vial with a stopper pierceable by a hypodermic needle. The active agent in the composition is a compound of the present invention. In some embodiments, a label on or associated with the container indicates that the composition is used to treat a selected disease. The article of manufacture can further include a second container containing a pharma- ceutically acceptable buffer, such as phosphate buffered saline, Ringer's solution, or dextrose solution. It can further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, syringes, and package inserts containing instructions for use.
[0355] In some embodiments, the CRISPR system (e.g., including Cas9 as described herein) is provided as part of a pharmaceutical composition. In some embodiments, the pharmaceutical composition includes any of the fusion proteins provided herein (e.g., including a nucleobase editor as described herein, including LubCas9). In some embodiments, the pharmaceutical composition includes any of the complexes provided herein. In some embodiments, the pharmaceutical composition includes a ribonucleoprotein complex including an RNA-guided nuclease (e.g., Cas9) that forms a complex with a gRNA and a cationic lipid. In some embodiments, the pharmaceutical composition includes a gRNA, a nucleic acid programmable DNA binding protein, a cationic lipid, and a pharma- ceutical acceptable excipient. The pharmaceutical composition can optionally include one or more additional therapeutically active substances.
[0356] kit In one aspect, the present invention provides a kit containing any one or more of the elements disclosed in the above methods and compositions. In some embodiments, the kit includes a vector system and instructions for using the kit. In some embodiments, the vector system includes one or more insertion sites for inserting a guide sequence, and when expressed, the guide sequence directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, and the CRISPR complex includes a CRISPR enzyme complexed with (1) a guide sequence hybridized to a target sequence, and (2) a sequence hybridized to a tracrRNA sequence, and / or (b) a second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme, the enzyme coding sequence including a nuclear localization sequence. The elements can be provided individually or in combination and can be provided in any suitable container, such as a vial, bottle, or tube. In some embodiments, the kit includes instructions in one or more languages, for example, in two or more languages.
[0357] In some embodiments, the kit comprises a nucleobase editor. For example, in some embodiments, the kit comprises a nucleobase editor comprising a Cas9 enzyme described herein (ScoCas9, Seq2Cas9, EhiCas9, SeqCas9, SsiCas9, SinCas9, SsaCas9, Ssc2Cas9, Sor2Cas9, SorCas9, SwaCas9, SscCas9, SgaCas9, LkuCas9, and SsuCas9).
[0358] In some embodiments, the kit includes one or more reagents for use in a process utilizing one or more of the elements described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents may be provided in a form that is usable in a particular assay or that requires the addition of one or more other components prior to use (e.g., concentrated or lyophilized form). The buffer may be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to the guide sequence for insertion into a vector such that the guide sequence and the regulatory element are operably linked. In some embodiments, the kit includes a homologous recombination template polynucleotide.
[0359] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety.In addition, the materials, methods, and examples are merely illustrative and are not intended to be limiting.Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs.Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described herein. EXAMPLES
[0360] The following examples illustrate some of the preferred modes of making and practicing the invention, however, it should be understood that these examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0361] Example 1. Screening for novel Cas9 enzymes, discovery and optimization of novel Cas9 enzymes This example describes a screen for the discovery of novel Cas9 enzymes. As described herein, this screen was used to identify and characterize Streptococcus equinus ATCC 33317, Enterococcus hirae strain F1129E, Streptococcus equinus strain AG46, Staphylococcus simulans strain 19, Streptococcus intermedius B196 strain G1552, Streptococcus sanguinis SK330, Streptococcus sp. C150, Streptococcus oralis subsp. oralis strain RH_1735_08, Streptococcus oralis SK313, Staphylococcus warneri strain 691, Staphylococcus stiuri strain SNUC 2430, Streptococcus gallolyticus strain AM24-4, Lactobacillus kullabergensis strain Biut2, and Streptococcus A novel Cas9 enzyme isolated from the bacteria Cas9. suis strain LSS83 was isolated and optimized.
[0362] In a search to discover new Cas9 enzymes that recognize novel PAM sequences, a bioinformatics screen was used to search for additional enzymes that would extend the targeting scope of CRISPR. The screen utilized seed sequences of Cas9 from Lachnospira spp. LubCas9 or Streptococcus constellatus ScoCas9. Bioinformatics was performed using the tblastn variant of BLAST with an e-value threshold of 1e-6 to consider BLAST hits. Briefly, the loci selected for testing were those that remained intact in the presence of Cas9 proteins from other species. Loci were selected that had more than three spacers within the CRISPR array, more than 1 kb of Cas9 endogenous sequence 5', and more than 300 nt 3' of the CRISPR array. Using this approach, novel Cas9 enzymes were identified from different bacterial species and codon-optimized for expression in human cells. The novel engineered Cas9 enzymes were then recombinantly produced and tested.
[0363] Example 2. Identifying 3' PAM consensus motifs for novel Cas9 enzymes from Streptococcus equinus ATCC 33317, Enterococcus hirae strain F1129E, Streptococcus equinus strain AG46, Staphylococcus simulans strain 19, Streptococcus intermedius B196 strain G1552, Streptococcus sanguinis SK330, Streptococcus sp. C150, Streptococcus oralis subsp. oralis strain RH_1735_08, Streptococcus oralis SK313, Staphylococcus warneri strain 691, Staphylococcus stiuri strain SNUC 2430, Streptococcus gallolyticus strain AM24-4, Lactobacillus kullabergensis strain Biut2, and Streptococcus suis strain LSS83 bacteria. This example includes Streptococcus equinus ATCC 33317, Enterococcus hirae strain F1129E, Streptococcus equinus strain AG46, Staphylococcus simulans strain 19, Streptococcus intermedius B196 strain G1552, Streptococcus sanguinis SK330, Streptococcus sp.C150, Streptococcus oralis subsp.oralis strain RH_1735_08, Streptococcus oralis SK313, Staphylococcus warneri strain 691, Staphylococcus stiuri strain SNUC 2430, Streptococcus gallolyticus strain AM24-4, Lactobacillus kullabergensis strain Biut2, and Streptococcus
[0046] Figure 1 shows the identification of a protospacer adjacent motif (PAM) sequence for the human codon-optimized Cas9 originally isolated from C. suis strain LSS83 bacteria.
[0364] Human, codon-optimized Cas9 was tested for its recognition of PAM sequences using an in vitro PAM specific assay. A library of plasmids carrying randomized PAM sequences was incubated with Cas9 isolated from different bacteria. Uncleaved plasmids were purified and sequenced to identify the specific PAM motif that was cleaved.For example, engineered Cas9 from Streptococcus equinus ATCC 33317 recognizes the consensus sequence 5'-NRGNR-3' (Figure 1A), Enterococcus hirae strain F1129E recognizes the consensus PAM sequence 5'-NRG-3' (Figure 1B), Streptococcus equinus strain AG46 (Figure 1C), Staphylococcus warneri strain 691 (Figure 1J), and Staphylococcus sciuri strain SNUC 2430 (Figure 1K) recognize the consensus PAM sequence 5'-NNGR-3', Staphylococcus simulans strain 19 recognizes the consensus PAM sequence 5'-NNGRRT-3' (Figure 1D), and Streptococcus intermedius B196 strain G1552 recognizes the consensus PAM sequence 5'-NNAAAA-3' (Fig. 1E), Streptococcus sanguinis SK330 recognizes the consensus PAM sequence 5'-NGGNG-3' (Fig. 1F), Streptococcus sp. C150 recognizes the consensus PAM sequence 5'-NNGNRG-3' (Fig. 1G), Streptococcus oralis subsp. oralis strain RH_1735_08 recognizes the consensus PAM sequence 5'-NNAAAC-3' (Fig. 1H), Streptococcus oralis SK313 recognizes the consensus PAM sequence 5'-NNRAAG-3' (Fig. 1I), Streptococcus gallolyticus strain AM24-4 recognizes the consensus PAM sequence 5'-NNAYAA-3' (Fig. 1L), and Lactobacillus kullabergensis strain Biut2 recognizes the consensus PAM sequence 5'-NNGAAA-3' (Figure 1M), and Streptococcus suis strains recognize the consensus PAM sequence 5'-NNAAA-3' (Figure 1N) (H=A, C, or T; R=A or G).
[0365] Example 3. Predicting RNA folding structures of sgRNA for novel Cas9 enzymes from Streptococcus constellatus, Sharpea spp. isolate RUG017, Veillonella parvula, Ezakiella peruensis, Lactobacillus fermentum strain AF15-40LB, and Peptoniphilus sp. Marseille-P3761 bacteria This example shows the predicted RNA folding structures of exemplary sgRNAs, including crRNA and tracrRNA, for use with the novel Cas9 enzyme.
[0366] Small RNA sequencing was performed on RNA derived from E. coli strains heterologously expressing the Cas9 Crispr locus. Briefly, RNA was isolated from stationary-phase bacteria by first resuspending E. coli in Trizol and then homogenizing the bacteria with zirconia / silica beads in a homogenizer for three 1-min cycles. Total RNA was purified from the homogenized samples, DNAse treated, 3' dephosphorylated with T4 polynucleotide kinase, and rRNA was removed. RNA libraries were prepared from rRNA-depleted RNA and size-selected for small RNAs.
[0367] For RNA sequencing, transcripts were polyA-tailed with E. coli poly(A) polymerase, ligated with 5' RNA adapters using T4 RNA ligase 1, reverse transcribed, followed by PCR amplification of cDNA with barcoded primers and sequencing on a MiSeq. Reads from each sample were identified based on their associated barcodes and aligned to reference sequences using BWA. Paired-end alignments were used to extract transcript sequences using Picard tools, and sequences were analyzed using Geneious software.
[0368] RNA folding was based on predictions from Geneious 11.1.2 software. A single sgRNA transcript fuses the crRNA to the tracrRNA, mimicking the duplex RNA structure required to induce site-specific Cas9 activity.The predicted RNA fold structure for a chimeric sgRNA for use with Seq2Cas9 from Streptococcus equinus ATCC 33317 is shown in FIG. 2A; an sgRNA for use with EhiCas9 from Enterococcus hirae strain F1129E is shown in FIG. 2B; an sgRNA for use with SeqCas9 from Streptococcus equinus strain AG46 is shown in FIG. 2C; an sgRNA for use with SsiCas9 from Staphylococcus simulans strain 19 is shown in FIG. 2D; an sgRNA for use with SinCas9 from Streptococcus intermedius B196 strain G1552 is shown in FIG. 2E; an sgRNA for use with SsaCas9 from Streptococcus sanguinis SK330 is shown in FIG. 2F; sgRNA for use with Ssc2Cas9 from Streptococcus oralis subsp. oralis strain RH_1735_08 is shown in FIG. 2H; sgRNA for use with Sor2Cas9 from Streptococcus oralis subsp. oralis strain RH_1735_08 is shown in FIG. 2I; sgRNA for use with SwaCas9 from Staphylococcus warneri strain 691 is shown in FIG. 2J; sgRNA for use with SsiCas9 from Staphylococcus sciuri strain SNUC 2430 is shown in FIG. 2K; sgRNA for use with SgaCas9 from Streptococcus gallolyticus strain AM24-4 is shown in FIG. 2L; sgRNAs for use with LkuCas9 from Streptococcus kullabergensis strain Biut2 are shown in Figure 2M, and SsuCas9 from Streptococcus suis strain LSS83 are shown in Figure 2N.
[0369] Example 4. Base editing by Cas9 enzyme using N-terminal fusion of adenine base editor (ABE) or cytidine base editor (CBE) This example shows the base conversion efficiency of Cas9 enzymes fused to an adenine base editor (ABE) or to a cytidine base editor (CBE).
[0370] Briefly, 25,000 HEK293T cells were seeded in a 96-well plate. 100 ng of Cas9 expression plasmid and 100 ng of guide expression plasmid were transfected 24 hours after seeding. Five days after transfection, cells were harvested and DNA was extracted.
[0371] Deep sequencing was performed to characterize A to G conversion in HEK293T cells. Exemplary targets were amplified using a two-round PCR region to add Illumina adapters and unique barcodes to the target amplicons. PCR products were run on 2% gels and gel extracted. Samples were pooled, quantified, cDNA libraries were prepared, and sequenced on a MiSeq. Percent A to G conversion was determined by deep sequencing of N- and C-terminal TadA8 fusion constructs.
[0372] Table 8 shows exemplary guide RNA sequences for use with Seq2Cas9. Figure 3A is a graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the Seq2Cas9 D10A mutant. [Table 8]
[0373] Table 9 shows exemplary guide RNA sequences used in EhiCas9. Figure 3B is a graph showing the results of indel frequency and adenine to guanine base (A to G) conversion percentage achieved with a base editor containing an ABE fused to the N-terminus of the EhiCas9 D10A mutant. [Table 9]
[0374] Table 10 shows exemplary guide RNA sequences for use with SeqCas9. Figure 3C is a graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the SeqCas9 D10A mutant. [Table 10]
[0375] Table 11 shows exemplary guide RNA sequences for use with SsiCas9. Figure 3D is a graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the SsiCas9 D10A mutant. [Table 11]
[0376] Table 12 shows exemplary guide RNA sequences for use with SinCas9. Figure 3E is a graph showing the percentage of adenine to guanine base (A to G) conversion results achieved with a base editor containing an ABE fused to the N-terminus of the SinCas9 D9A mutant. [Table 12]
[0377] Table 13 shows exemplary guide RNA sequences for use with SsaCas9. Figure 3F is a graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the SsaCas9 D11A mutant. [Table 13]
[0378] Table 14 shows exemplary guide RNA sequences for use with Ssc2Cas9. Figure 3G shows the results of indel frequency and adenine to guanine base (A to G) conversion percentage achieved with a base editor containing an ABE fused to the N-terminus of the Ssc2Cas9 D9A mutant. [Table 14]
[0379] Table 15 shows exemplary guide RNA sequences for use with Sor2Cas9. Figure 3H is a graph showing the results of indel frequency and adenine to guanine base (A to G) conversion percentage achieved with a base editor containing an ABE fused to the N-terminus of the Sor2Cas9 D9A mutant. [Table 15]
[0380] Table 16 shows exemplary guide RNA sequences for use with SorCas9. Figure 3I is a graph showing the results of the percentage adenine to guanine base (A to G) conversion achieved with a base editor containing an ABE fused to the N-terminus of the SorCas9 D9A mutant. [Table 16]
[0381] Table 17 shows exemplary guide RNA sequences used in SwaCas9. Figure 3J is a graph showing the percentage of adenine to guanine base (A to G) conversion results achieved with base editors containing ABE fused to the N-terminus of SwaCas9 variants. [Table 17]
[0382] Table 18 shows exemplary guide RNA sequences for use with SscCas9. Figure 3K is a graph showing the percentage of adenine to guanine base (A to G) conversion results achieved with base editors containing ABEs fused to the N-terminus of SscCas9 mutants. [Table 18]
[0383] Table 19 shows exemplary guide RNA sequences used in SgaCas9. Figure 3L is a graph showing the percentage of adenine to guanine base (A to G) conversion results achieved with base editors containing ABE fused to the N-terminus of SgaCas9 mutants. [Table 19]
[0384] Table 20 shows exemplary guide RNA sequences used in LkuCas9. Figure 3M is a graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with base editors containing ABE fused to the N-terminus of LkuCas9 mutants. [Table 20]
[0385] Table 21 shows exemplary guide RNA sequences for use with SsuCas9. Figure 3N is a graph showing the results of the percentage of adenine to guanine base (A to G) conversion achieved with base editors containing ABEs fused to the N-terminus of SsuCas9 mutants. [Table 21] [Table 22] Table 23 discloses exemplary Cas9 adenosine or adenine sequences for base editing functions. [Table 23] TIFF2024546841000214.tif251165TIFF2024546841000215.tif248168TIFF2024546841000216.tif165170TIFF2024546841000217.tif250167TIFF2024546841000218.tif250170TIFF2024546841000219.tif249170TIFF2024546841000220.tif152170TIFF2024546841000221.tif121170Linker (no underline, italics or bold) TadA8(ABE)(italics and underlined) Nickase mutations (bold and italics) The 3xHA tag (in italics) can be replaced by a different tag.
[0386] Equivalence and Scope Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. The scope of the invention is not intended to be limited to the above Description, but rather is as set forth in the following claims.